@hviana/sema 0.7.3 → 0.7.6
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +95 -843
- package/README.md +11 -11
- package/dist/src/mind/graph-search.d.ts +32 -22
- package/dist/src/mind/graph-search.js +97 -54
- package/dist/src/mind/mind.js +14 -0
- package/dist/src/store-sqlite.js +17 -0
- package/dist/src/store.d.ts +18 -0
- package/dist/src/store.js +10 -0
- package/docs/INDEX.md +71 -0
- package/docs/INVARIANTS.md +19 -0
- package/docs/architecture/bounded-reads.md +85 -0
- package/docs/architecture/caches.md +89 -0
- package/docs/architecture/commonality.md +45 -0
- package/docs/architecture/cost-model.md +71 -0
- package/docs/architecture/determinism.md +73 -0
- package/docs/architecture/exact-vs-approximate.md +47 -0
- package/docs/architecture/factored-machinery.md +28 -0
- package/docs/architecture/fold-contract.md +87 -0
- package/docs/architecture/halo-sketch.md +99 -0
- package/docs/architecture/match-project.md +62 -0
- package/docs/architecture/mechanism-market.md +95 -0
- package/docs/architecture/memoization.md +96 -0
- package/docs/architecture/meter.md +55 -0
- package/docs/architecture/saturation.md +92 -0
- package/docs/architecture/store.md +79 -0
- package/docs/architecture/thresholds.md +79 -0
- package/docs/failures/tempting-but-wrong.md +144 -0
- package/docs/harness/gates.md +56 -0
- package/docs/mechanisms/alu.md +75 -0
- package/docs/mechanisms/cast.md +75 -0
- package/docs/mechanisms/confluence.md +36 -0
- package/docs/mechanisms/cover.md +54 -0
- package/docs/mechanisms/extraction.md +53 -0
- package/docs/mechanisms/prefix-completion.md +54 -0
- package/docs/mechanisms/recall.md +69 -0
- package/docs/mechanisms/reference.md +58 -0
- package/jsr.json +1 -1
- package/package.json +1 -1
- package/src/mind/graph-search.ts +101 -55
- package/src/mind/mind.ts +14 -0
- package/src/store-sqlite.ts +19 -0
- package/src/store.ts +22 -0
- package/test/89-completion-recursion.test.mjs +30 -10
- package/test/97-store-seed.test.mjs +105 -0
- package/test/98-completion-chaining.test.mjs +140 -0
- package/HOW_IT_WORKS.md +0 -5836
|
@@ -0,0 +1,62 @@
|
|
|
1
|
+
# Match → Project → Gate
|
|
2
|
+
|
|
3
|
+
Every grounding mechanism is a configuration of one shared operation in
|
|
4
|
+
`src/mind/match.ts`. The family is defined once and imported many times;
|
|
5
|
+
duplicating it forks the corpus contract, moving it hides who owns the gate.
|
|
6
|
+
|
|
7
|
+
## The shared family in `mind/match.ts`
|
|
8
|
+
|
|
9
|
+
The match layer locates structure, the project layer moves along the store, and
|
|
10
|
+
the gate layer decides whether the shape licences voicing. All three are pure
|
|
11
|
+
functions over bytes and the store — no mechanism owns a private copy.
|
|
12
|
+
|
|
13
|
+
## The triple
|
|
14
|
+
|
|
15
|
+
| Role | Symbols | What it does |
|
|
16
|
+
| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------- |
|
|
17
|
+
| **Match** (locate structure) | `locate` (exact → halo → gist ladder), `alignRuns` (literal W-gram weave), `alignGraded` (literal + halo gaps), `alignAround` / `frameSlots` (seeded frame with contracted gaps), `bestHaloMate` (in-list halo), `analogyStrength` / `sharedFrameStrength` (distributional + structural analogy) | Finds where a query sits in a learnt form. |
|
|
18
|
+
| **Project** (direction) | `follow` (forward to fixpoint, first hop may `conceptHop`), `reverseContext` (reverse to context), `project` (forward else reverse), `conceptHop` (halo sibling with edge) | Moves along the store from the match — forward toward answers, reverse toward contexts. |
|
|
19
|
+
| **Gate** (structural licence) | `isSpanShaped` (sparse subsequence — open reading), `carriesFillers` (substitution carriage — strict voicing licence) | Decides whether the shape licences voicing. |
|
|
20
|
+
|
|
21
|
+
Mechanisms declare only `(matcher, direction, gate)`. Thresholds behind gates
|
|
22
|
+
live in `src/geometry.ts` — the match layer never invents a cutoff.
|
|
23
|
+
|
|
24
|
+
The graded ladder inside `locate` is exact → distributional → geometric:
|
|
25
|
+
content-addressed identity first, halo similarity second, gist resonance last.
|
|
26
|
+
Reordering the ladder or letting an approximate score override an exact hit is a
|
|
27
|
+
correctness bug (see `exact-vs-approximate.md`).
|
|
28
|
+
|
|
29
|
+
## Frame reading — matcher reports, gate judges, inventory elects nothing
|
|
30
|
+
|
|
31
|
+
`frameSlots` is the shared frame reader. It contracts every gap via
|
|
32
|
+
`contractGap` to its varying core, tags it `substitution` / `insertion` /
|
|
33
|
+
`deletion`, and attaches `covered` — the bytes the frame accounts for. It
|
|
34
|
+
applies no gate; it reports.
|
|
35
|
+
|
|
36
|
+
`carriesFillers` is the substitution gate. It judges byte-exactly:
|
|
37
|
+
|
|
38
|
+
```
|
|
39
|
+
substituteAll(contA, fillersA → fillersB) == contB
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
If the equality holds, voicing through the slot is a derivation; if not, the
|
|
43
|
+
slot cannot carry. This is the only place that decision is made.
|
|
44
|
+
|
|
45
|
+
`Precomputed.frames` is the inventory. It enumerates every frame pairing the
|
|
46
|
+
match layer finds and elects nothing — ranking and refusal belong to the
|
|
47
|
+
consumer.
|
|
48
|
+
|
|
49
|
+
## Voicing gates belong to the consumer
|
|
50
|
+
|
|
51
|
+
The shared layer never refuses on a consumer's behalf. Reference owns its four
|
|
52
|
+
gates: frame dominates the query, each slot reaches `W` on both sides, no
|
|
53
|
+
insertion/deletion, fillers pairwise distinct — plus `carriesFillers` on the
|
|
54
|
+
chosen pair. CAST, recall, and cover each apply their own gate over the same
|
|
55
|
+
shared inventory. Moving a consumer's gate into `match.ts` would hide who is
|
|
56
|
+
responsible for the refusal.
|
|
57
|
+
|
|
58
|
+
## Pins
|
|
59
|
+
|
|
60
|
+
- `test/47` — frame reading split (matcher vs gate vs inventory).
|
|
61
|
+
- `test/50` — CAST / reference voicing via `carriesFillers`.
|
|
62
|
+
- `test/24` / `test/76` — match/project family and span-shape readings.
|
|
@@ -0,0 +1,95 @@
|
|
|
1
|
+
# Mechanism Market — The Free-Will Architecture
|
|
2
|
+
|
|
3
|
+
Every grounding mechanism — including the ALU and user extensions — speaks one
|
|
4
|
+
interface (`mind/pipeline-mechanism.ts`). The decider in `mind/pipeline.ts`
|
|
5
|
+
(`think`) holds a plain list and never branches on which mechanism it holds.
|
|
6
|
+
|
|
7
|
+
## Interface
|
|
8
|
+
|
|
9
|
+
```ts
|
|
10
|
+
interface PipelineMechanism {
|
|
11
|
+
parse?(query: Uint8Array): Promise<ComputedSpan[]>; // authoritative spans
|
|
12
|
+
floor(ctx, query, pre, worthRunning): Promise<number | null>; // bound or null
|
|
13
|
+
run(ctx, query, pre): Promise<MechanismResult[]>; // candidates
|
|
14
|
+
}
|
|
15
|
+
interface MechanismResult {
|
|
16
|
+
bytes: Uint8Array;
|
|
17
|
+
accounted: [number, number][];
|
|
18
|
+
moves: number;
|
|
19
|
+
unexplained: string;
|
|
20
|
+
scaffolding?: number;
|
|
21
|
+
complete?: boolean;
|
|
22
|
+
}
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
- `parse` is optional; all results are collected into `Precomputed.computed`
|
|
26
|
+
before any `floor`/`run`.
|
|
27
|
+
- `floor` returns `null` when structurally impossible, otherwise an admissible
|
|
28
|
+
lower bound (never overstates cost).
|
|
29
|
+
- `run` returns candidates with travelling evidence (below).
|
|
30
|
+
|
|
31
|
+
## Decider
|
|
32
|
+
|
|
33
|
+
`think` in `mind/pipeline.ts` iterates `defaultMechanisms` in list order:
|
|
34
|
+
|
|
35
|
+
```
|
|
36
|
+
defaultMechanisms = [cover, cast, confluence, extraction, reference, recall,
|
|
37
|
+
prefix-completion] + ALU (`aluToMechanism`) + extensions
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
Weight is one currency: `weight = moves + PASS · unaccountedBytes` where
|
|
41
|
+
`unaccountedBytes = unexplainedSpans(query.length, accounted)`. Comparison is at
|
|
42
|
+
`STEP` grade (`grade = floor(weight/STEP)`); equal grade prefers fewer
|
|
43
|
+
`scaffolding` bytes, then list order.
|
|
44
|
+
|
|
45
|
+
## Four constraints
|
|
46
|
+
|
|
47
|
+
1. **Decoupled** — zero cross-imports between `mind/mechanisms/*`. Adding one
|
|
48
|
+
never touches another; no mechanism asks what already decided.
|
|
49
|
+
2. **Declared competence** — binary structural gates inside `floor`/`run` (query
|
|
50
|
+
length, anchor shape, weave existence). Never a learned score; rationale
|
|
51
|
+
states exactly why a mechanism abstained.
|
|
52
|
+
3. **Visible budget** — every corpus-scale loop is capped at a named constant:
|
|
53
|
+
`√N` via `hubBound`/`hubCap` and `k = 2·recallQueryK` (`Precomputed.k`).
|
|
54
|
+
Enforced at the store level.
|
|
55
|
+
4. **Evidence travels** — every candidate carries `accounted` (query spans
|
|
56
|
+
explained), `moves` (priced on `MICRO/STEP/CONCEPT/PASS`), `unexplained`
|
|
57
|
+
(diagnostic label); optionally `scaffolding` (answer bytes from unrecognised
|
|
58
|
+
spans — equal-grade tie-break) and `complete` (trained-form continuation
|
|
59
|
+
reached via identity; post-grounding must not extend). The decider honours
|
|
60
|
+
both without knowing who set them.
|
|
61
|
+
|
|
62
|
+
## Two disciplines
|
|
63
|
+
|
|
64
|
+
- **Admissible-floor pruning.** `floor` runs for every mechanism in list order
|
|
65
|
+
before any `run`. `run` fires only if `worthRunning(floor)` where
|
|
66
|
+
`worthRunning = (floor) => best === null || grade(floor) < grade(best.weight)`.
|
|
67
|
+
Cover runs first so a near-zero-cost computed span prunes the rest through the
|
|
68
|
+
same mechanism — not a special case.
|
|
69
|
+
|
|
70
|
+
- **Investment discipline.** `worthRunning` is passed _into_ `floor`. A floor
|
|
71
|
+
that would first-touch an expensive shared analysis (`pre.attention()` climb,
|
|
72
|
+
`pre.weave()`, `pre.resonance()`) checks `worthRunning(cheapestBound)` first
|
|
73
|
+
and returns the uninvested bound when it already loses. Never compute a shared
|
|
74
|
+
analysis just to discard it. `cast.ts`/`extraction.ts` are the references.
|
|
75
|
+
|
|
76
|
+
## Accounting
|
|
77
|
+
|
|
78
|
+
- **Extraction:** located frames are always evidence; the span between them
|
|
79
|
+
counts only when _both_ borders were located. An open-ended read is priced by
|
|
80
|
+
exclusion (`PASS`/byte).
|
|
81
|
+
- **Reverse reading:** `reverseContext` produces bytes but explains nothing
|
|
82
|
+
forward: `accounted = []`, weight ≈ `PASS·|query|` — last resort by
|
|
83
|
+
arithmetic, not rule.
|
|
84
|
+
- **Paid acts are accounted:** the bridge's corroborated substitutions cost
|
|
85
|
+
`CONCEPT` each in `moves`, so their spans must be `accounted`; otherwise the
|
|
86
|
+
same act is charged twice (`PASS`/byte dominates).
|
|
87
|
+
|
|
88
|
+
`accounted` is a cost-ladder quantity; `cover.ts` leaves masked computed spans
|
|
89
|
+
out of it so `PASS`-bridged bytes are still charged. `unexplained`,
|
|
90
|
+
`narrowDecision`, `thinGrounding` are observational only.
|
|
91
|
+
|
|
92
|
+
## Pins
|
|
93
|
+
|
|
94
|
+
- `test/01-floor` — floor geometry.
|
|
95
|
+
- `test/04-think` — decider, admissible pruning, investment discipline.
|
|
@@ -0,0 +1,96 @@
|
|
|
1
|
+
# Memoization — Shared Evidence Without Duplication
|
|
2
|
+
|
|
3
|
+
> **Law:** asking never writes, so structural reads are pure during one
|
|
4
|
+
> response. Memoization elides probes, not evidence.
|
|
5
|
+
|
|
6
|
+
Two layers: `Precomputed` (response-scoped shared analyses) and `Mind`
|
|
7
|
+
per-response memos. Both are accelerators that must not change what inference
|
|
8
|
+
computes.
|
|
9
|
+
|
|
10
|
+
## Precomputed — one response, one container
|
|
11
|
+
|
|
12
|
+
`Precomputed` (`src/mind/pipeline-mechanism.ts`) is the sole place a response's
|
|
13
|
+
shared evidence lives. Created by `think` (`src/mind/pipeline.ts`) before the
|
|
14
|
+
mechanism loop.
|
|
15
|
+
|
|
16
|
+
### Eager — populated before any `floor`/`run`
|
|
17
|
+
|
|
18
|
+
- `rec: Recognition` — structural + canonical decomposition (`recognise`)
|
|
19
|
+
- `computed: ComputedSpan[]` — `parse()` results from all mechanisms (e.g. ALU)
|
|
20
|
+
- `guide: Vec` — query gist, the response-wide disambiguation guide
|
|
21
|
+
- `k: number` — `cfg.recallQueryK * 2`, the breadth for resonance/weave/climb
|
|
22
|
+
|
|
23
|
+
### Lazy — computed on first touch, cached by promise
|
|
24
|
+
|
|
25
|
+
Expensive analyses are `async` and cached by promise: the first caller starts
|
|
26
|
+
the work, every later caller awaits the same promise.
|
|
27
|
+
|
|
28
|
+
- `attention()` — `climbAttentionAll` (roots + ranked anchors)
|
|
29
|
+
- `weave()` — `alignGraded` over top-k anchors
|
|
30
|
+
- `resonance()` — `store.resonate(guide, k)` (single ANN query)
|
|
31
|
+
- `frames()` — `frameSlots` inventory from resonance
|
|
32
|
+
- `spanShapedOf(anchor)` / `spanShapedAll()` — per-anchor `skillExemplar`,
|
|
33
|
+
memoised per id
|
|
34
|
+
- `queryWindows` / `queryResolved` / `windowsOf(anchor)` — W-window identities
|
|
35
|
+
- `reachMemo` — `sharedReachMemo(ctx)` (ancestor reach, § below)
|
|
36
|
+
|
|
37
|
+
A mechanism that never asks pays nothing; two mechanisms asking the same
|
|
38
|
+
question pay once. `floor()` must gate on `worthRunning` before first-touching
|
|
39
|
+
an expensive analysis.
|
|
40
|
+
|
|
41
|
+
## Mind memos — `beginResponse` → `endResponse`
|
|
42
|
+
|
|
43
|
+
`Mind` (`src/mind/mind.ts:beginResponse`/`endResponse`) swaps per-response state
|
|
44
|
+
for each inference call. `respond` takes fresh maps; `respondTurn` reuses the
|
|
45
|
+
conversation's persistent ones (content-keyed, cross-turn).
|
|
46
|
+
|
|
47
|
+
| Memo | Key | Scope |
|
|
48
|
+
| ------------------- | ----------------------------------------- | -------------------------------------------- |
|
|
49
|
+
| `perceiveMemo` | `perceiveKey(bytes)` (latin1) | response / conversation |
|
|
50
|
+
| `recogniseMemo` | `latin1Key(bytes)` | response / conversation |
|
|
51
|
+
| `climbMemo` | `latin1Key(bytes)` | response / conversation |
|
|
52
|
+
| `canonMemo` | `latin1Key(bytes)` | response (when `canon` set) |
|
|
53
|
+
| `_resolvedSubtrees` | `WeakMap<Sema, {id,len}>` (node identity) | response / conversation |
|
|
54
|
+
| `_edgeChoice` | `Map<nodeId, pick>` | response only — **cleared** in `endResponse` |
|
|
55
|
+
| `_gistCache` | `BoundedMap<nodeId, Vec>` 32 MB | **session-lifetime** (not per-response) |
|
|
56
|
+
|
|
57
|
+
`_gistCache` (≈ 8K gists at D=1024) survives across responses; all others are
|
|
58
|
+
dropped or cleared at `endResponse`. `_resolvedSubtrees` elides store probes
|
|
59
|
+
when `visit` is absent; with a visitor it still walks in full (see
|
|
60
|
+
`src/mind/primitives.ts:foldTree`).
|
|
61
|
+
|
|
62
|
+
## Trace boundary — what is bypassed
|
|
63
|
+
|
|
64
|
+
Traced responses must emit every step, but must not change the answer.
|
|
65
|
+
|
|
66
|
+
- **Bypassed:** `_edgeChoice` via `guidedNext`
|
|
67
|
+
(`src/mind/traverse.ts:guidedNext`) and `sharedReachMemo`
|
|
68
|
+
(`src/mind/traverse.ts:sharedReachMemo`). Both return fresh empty maps when
|
|
69
|
+
`ctx.trace !== null`; `chooseNext` recomputes identically (pure over store +
|
|
70
|
+
guide).
|
|
71
|
+
- **Always consulted:** `perceiveMemo`, `recogniseMemo`, `climbMemo` (and their
|
|
72
|
+
underlying `perceive`/`recognise`/`climbAttention` caches). Bypassing breaks
|
|
73
|
+
idempotence.
|
|
74
|
+
|
|
75
|
+
`foldTree`'s subtree fast path is taken only when no `visit` is supplied. With a
|
|
76
|
+
visitor (recognition, attention) the walk still descends; the cache elides only
|
|
77
|
+
probes. Bypassing `recogniseMemo` under trace re-ran `recogniseImpl` with a warm
|
|
78
|
+
`_resolvedSubtrees` and emitted fewer sites (observed 31 → 5) — a correctness
|
|
79
|
+
change, not just a slowdown.
|
|
80
|
+
|
|
81
|
+
## Meter — charge work to itself
|
|
82
|
+
|
|
83
|
+
Shared analyses are charged to their own phase (`meter.time(phase, fn)` in
|
|
84
|
+
`Precomputed.shared`), not to the mechanism that first touched them
|
|
85
|
+
(`src/meter.ts:PhaseCost`). Without this, the profile reads "cast.floor costs 2
|
|
86
|
+
s" when the cost was the consensus climb cast paid for on everyone's behalf.
|
|
87
|
+
|
|
88
|
+
## Adding a shared analysis
|
|
89
|
+
|
|
90
|
+
Add one lazy method to `Precomputed`. No new memo map elsewhere. Gate it behind
|
|
91
|
+
`worthRunning` in `floor()`.
|
|
92
|
+
|
|
93
|
+
## Pins
|
|
94
|
+
|
|
95
|
+
- `test/42` — recognition idempotence under trace: traced and untraced
|
|
96
|
+
`recognise` return the same site count and cached object.
|
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
# Meter — Work Accounting
|
|
2
|
+
|
|
3
|
+
`src/meter.ts` is the one computational-usage accounting surface. It counts what
|
|
4
|
+
inference _cost_ so a slow response can be attributed instead of guessed at. The
|
|
5
|
+
rationale says why an answer was chosen; the meter says what it cost to choose
|
|
6
|
+
it. Harness: `bench/profile-inference.mjs`.
|
|
7
|
+
|
|
8
|
+
## Five contracts
|
|
9
|
+
|
|
10
|
+
1. **Write-only from inference.** No counter reaches a decision, a threshold, or
|
|
11
|
+
an ordering. Determinism survives only because the meter is observed, never
|
|
12
|
+
consulted. Every call site is `meter?.x++` on a nullable field.
|
|
13
|
+
|
|
14
|
+
2. **Counters vs hints.** Counters are deterministic and diffable between runs;
|
|
15
|
+
the same query on the same store meters identically, so a regression is
|
|
16
|
+
visible in a diff. Millisecond fields (`elapsedMs`, per-phase `ms`) are
|
|
17
|
+
non-deterministic hints reported separately — never use them to gate
|
|
18
|
+
behaviour.
|
|
19
|
+
|
|
20
|
+
3. **Phases nest, they do not partition.** `think` contains every mechanism
|
|
21
|
+
phase; a mechanism's `floor` contains whatever shared analysis it
|
|
22
|
+
first-touched; `recall.run` contains `substitutionBridge`. Read a phase as
|
|
23
|
+
inclusive wall-clock — never sum phases and expect the total.
|
|
24
|
+
`CostReport.elapsedMs` is the only whole.
|
|
25
|
+
|
|
26
|
+
4. **Count once.** Off by default and free when off
|
|
27
|
+
(`new Mind({ profile:
|
|
28
|
+
true })` to attach). A layer that wants to be
|
|
29
|
+
visible bumps a field in `meter.ts` — it does not grow a private counter.
|
|
30
|
+
(Legacy `danglingReads` / `compactFailures` in `store.ts` are health
|
|
31
|
+
counters, not per-response work.)
|
|
32
|
+
|
|
33
|
+
5. **Shared analyses charged to themselves.** The first toucher pays the wall
|
|
34
|
+
clock, but every later consumer gets the result free. Attribution follows the
|
|
35
|
+
analysis, not the mechanism that first triggered it — otherwise the profile
|
|
36
|
+
misreads which work is expensive (e.g. the consensus climb billed through
|
|
37
|
+
whichever mechanism happened to need it first).
|
|
38
|
+
|
|
39
|
+
## Where counters live
|
|
40
|
+
|
|
41
|
+
`src/meter.ts:Meter` is the only definition of a counter name. Phases are
|
|
42
|
+
charged via `meter.time(phase, fn)` / `meter.timeSync(phase, fn)`, which
|
|
43
|
+
snapshot counters on entry and attribute the delta to the phase. The sync/async
|
|
44
|
+
seam is load-bearing: synchronous layers (perception, recognition, graph search)
|
|
45
|
+
must use `timeSync` so the profiled path does not await where the unprofiled
|
|
46
|
+
path does not.
|
|
47
|
+
|
|
48
|
+
`CostReport` is plain JSON (`version`, `elapsedMs`, `queryBytes`, `counters`,
|
|
49
|
+
`phases`). Zero-valued counters are dropped; `formatReport` renders the three
|
|
50
|
+
heaviest counters per phase.
|
|
51
|
+
|
|
52
|
+
## Pins
|
|
53
|
+
|
|
54
|
+
- `test/55` — `Meter`, `CostReport`, `searchPops` / `searchPushes`, phase
|
|
55
|
+
nesting.
|
|
@@ -0,0 +1,92 @@
|
|
|
1
|
+
# Saturation — Named Stop, Not Cap
|
|
2
|
+
|
|
3
|
+
> **Law:** every walk names a deciding saturation beside its cap. The cap is a
|
|
4
|
+
> safety net; saturation is the derived stop that decides and terminates.
|
|
5
|
+
|
|
6
|
+
A walk with only a cap drifts to the cap. A walk with a saturation stops the
|
|
7
|
+
moment the answer (saturated vs. decided) is known — bounded, exact below the
|
|
8
|
+
bound, and named in the trace.
|
|
9
|
+
|
|
10
|
+
## Cap vs. saturation
|
|
11
|
+
|
|
12
|
+
| Role | Value | Nature |
|
|
13
|
+
| ---------- | ------------------------------------------------------------ | ------------------------------------------------ |
|
|
14
|
+
| Cap | `hubBound = ceil(sqrt(N))` per read; `hubBound * W` per walk | Safety net — prevents corpus-proportional work |
|
|
15
|
+
| Saturation | Named derived stop (`SaturationReason`) | Decision — proves continuing cannot discriminate |
|
|
16
|
+
|
|
17
|
+
`N = corpusN = max(2, edgeSourceCount())`, `W = maxGroup`. Defined once in
|
|
18
|
+
`mind/traverse.ts` (`corpusN`, `hubBound`, `boundFor`); never spelled inline.
|
|
19
|
+
|
|
20
|
+
## edgeAncestors — EXPAND-UNTIL-DECIDED
|
|
21
|
+
|
|
22
|
+
`mind/traverse.ts:edgeAncestors` is the model: it climbs the structural DAG
|
|
23
|
+
(parents + containment) until one of five saturations decides, and is exact
|
|
24
|
+
below every one of them.
|
|
25
|
+
|
|
26
|
+
1. **Predecessor fan-in** — `prevCount(node) > bound` via `store.prevCount`. One
|
|
27
|
+
indexed count; no read. Proves `> bound` distinct contexts on its own.
|
|
28
|
+
2. **Distinct-context limit** — `ctxSeen.size > bound` after
|
|
29
|
+
`prevFirst(node, bound)`. The accumulated set of distinct contexts reachable
|
|
30
|
+
from roots visited so far exceeds `sqrt(N)`.
|
|
31
|
+
3. **Parent fan-out** — `parentsFirst(node, bound+1).length > bound`. One
|
|
32
|
+
`LIMIT bound+1` read distinguishes "exactly bound parents" from "hub". The
|
|
33
|
+
node itself is not expanded; saturated reaches are never voted.
|
|
34
|
+
4. **Lateral-cone cumulative** — `lateral > bound`, where `lateral` sums
|
|
35
|
+
`fresh-1` over every expanded node's extra parents beyond its first. The
|
|
36
|
+
per-node guard catches concentration at one node; this catches the same
|
|
37
|
+
commonness distributed across the cone. A deep chain in one structure accrues
|
|
38
|
+
zero laterals and still reaches its root at any depth.
|
|
39
|
+
5. **Byte-atom commonality** — `atomIsHub(N,W)` when
|
|
40
|
+
`atomReach(N,W) = max(1, ceil(N*W/256)) > bound`. Atoms carry no kid/contain
|
|
41
|
+
rows, so containment is unmeasurable; the uniform-expectation floor replaces
|
|
42
|
+
it. Above the scale the atom abstains as voter (edges remain traversable for
|
|
43
|
+
tier-0 recall).
|
|
44
|
+
|
|
45
|
+
Below every threshold the walk is exact — `prevFirst(bound)` is the full list,
|
|
46
|
+
`parentsFirst(bound+1)` is the full list, `containersSlice` pages are walked in
|
|
47
|
+
full — identical to the unbounded climb. Work is `O(bound)` contexts times local
|
|
48
|
+
structure, never `O(N)`. Container seeding is streamed in `bound`-sized pages
|
|
49
|
+
for the same reason.
|
|
50
|
+
|
|
51
|
+
Trace records the first deciding stop as
|
|
52
|
+
`SaturationStop { reason, node, observed, limit }` plus `visited`/`maxDepth`;
|
|
53
|
+
absent when unsaturated or untraced.
|
|
54
|
+
|
|
55
|
+
## pivotInto — longest-wins
|
|
56
|
+
|
|
57
|
+
`mind/resonance.ts:pivotInto` ranks candidates by `contentLen(id, answerLen+1)`
|
|
58
|
+
descending (first-inserted tie-break). The byte score is length itself, so the
|
|
59
|
+
scan is decided at the first candidate that passes every filter — a shorter
|
|
60
|
+
candidate can never outscore it. At most one winner's bytes are reconstructed;
|
|
61
|
+
every shorter proposal is skipped without a read. Saturation, not cap.
|
|
62
|
+
|
|
63
|
+
## Junction — hub guards vs. budget
|
|
64
|
+
|
|
65
|
+
`mind/junction.ts:junctionContainersFrom` has three disciplines:
|
|
66
|
+
|
|
67
|
+
- **Phrase-scale reads** — `bytesPrefix(maxContainer+1)` per visit; a node
|
|
68
|
+
beyond the cap prunes its branch.
|
|
69
|
+
- **Per-node hub guards** (real saturations) — `parentsFirst(bound+1) > bound`
|
|
70
|
+
not expanded; one `containersSlice(bound+1)` page beyond `bound` not expanded.
|
|
71
|
+
Each is exact below `bound`.
|
|
72
|
+
- **Expansion budget** — at most `bound * W` pops total (shared across a tier's
|
|
73
|
+
walks). Budget exhaustion is an abstention (`junctionBudgetExhausted`) that
|
|
74
|
+
falls through to the resonance tier — a net, not a saturation.
|
|
75
|
+
|
|
76
|
+
Refuted tightening: applying `edgeAncestors`' cumulative lateral-cone limit here
|
|
77
|
+
would discard half the successful junctions (measured lateral 1425/1426 at
|
|
78
|
+
`bound` ~ 570).
|
|
79
|
+
|
|
80
|
+
Refuted early-stop: **one-cone-exhausted** — stopping when one side's upward
|
|
81
|
+
cone empties — is wrong in both hub-guarded and hub-flagged forms. A junction
|
|
82
|
+
can be reachable from only one side when that side's seed is a fold sub-node of
|
|
83
|
+
the container (`test/16`: "cold or hot" reached from window "cold" while 3-byte
|
|
84
|
+
"hot" cone is empty; `test/34` n-ary binding fails the same way). Exhausting one
|
|
85
|
+
cone never proves no junction remains; the walk must keep the `bound*W` net
|
|
86
|
+
after per-node saturations.
|
|
87
|
+
|
|
88
|
+
## Pins
|
|
89
|
+
|
|
90
|
+
- `test/16` — one-cone-exhausted refutation (bridge junction).
|
|
91
|
+
- `test/34` — one-cone-exhausted refutation (n-ary cross-region binding).
|
|
92
|
+
- `test/27` — saturation-drop gate (leading/trailing saturated intervals).
|
|
@@ -0,0 +1,79 @@
|
|
|
1
|
+
# Store — AbstractStore Owns the Domain, Adapters Own the Wires
|
|
2
|
+
|
|
3
|
+
> **Law:** `AbstractStore` (`src/store.ts`) owns every domain decision — dedup,
|
|
4
|
+
> near-dedup, gist/halo indexing, containment, batching, LRU, compaction.
|
|
5
|
+
> `SQliteStore` (`src/store-sqlite.ts`) implements only `_db*`/`_vec*` thin
|
|
6
|
+
> wrappers.
|
|
7
|
+
|
|
8
|
+
## Template method
|
|
9
|
+
|
|
10
|
+
```
|
|
11
|
+
AbstractStore all logic: caches, merge gates, halo schedule,
|
|
12
|
+
buffers, chain transparency, length walks
|
|
13
|
+
└─ SQliteStore SQL + VectorDatabase glue — one statement per method
|
|
14
|
+
└─ <NewBackend> same contract — subclass AbstractStore only
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
A new backend subclasses `AbstractStore`; never re-implements dedup or batching.
|
|
18
|
+
|
|
19
|
+
## IDs and leaves
|
|
20
|
+
|
|
21
|
+
Branch ids are dense non-negative `0,1,2,…` (`_nextId`, never deleted).
|
|
22
|
+
Single-byte leaves are **implicit** negative ids `-256..-1` (`-(byte+1)`), never
|
|
23
|
+
a row. `has(id)` is `id < 0 || id < _nextId`.
|
|
24
|
+
|
|
25
|
+
## Flat branches and bytes
|
|
26
|
+
|
|
27
|
+
A branch whose kids are all leaves is **flat** — stored as raw bytes in `leaf`
|
|
28
|
+
with an empty `kids` blob as marker (`flatKidsBytes`/`flatBytesKids`). Dedup
|
|
29
|
+
probes hash then verify: `hashOf`→`h`→`LIMIT 1` fetch→byte compare (bloom
|
|
30
|
+
negative filter first).
|
|
31
|
+
|
|
32
|
+
`bytes(id)`/`bytesPrefix(id, cap)` are shared with `BoundedMap` caches — callers
|
|
33
|
+
must **never mutate** the returned buffer. `contentLen(id, cap)` walks with
|
|
34
|
+
memo; when `cap` is given it saturates (`>= cap` without finishing) so one huge
|
|
35
|
+
root never costs a full walk.
|
|
36
|
+
|
|
37
|
+
## Gist, halo, dedup
|
|
38
|
+
|
|
39
|
+
On `put*`, content dedup (`hashOf`→probe→mint) gates first. `DedupKey` caches
|
|
40
|
+
short keys (`DEDUP_KEY_MAX` bypass). Near-dedup merges by `mergeThreshold(D)` on
|
|
41
|
+
unit gist cosine. Gists sit in `_pendingGist` (byte-budgeted `BoundedMap`);
|
|
42
|
+
`indexSubtree` & `pourHalo` promote via `_vecContentUpsert`/`_vecHaloUpsert` in
|
|
43
|
+
`batchSize` batches. Buffers flush on cadence, `commit()`, and close. Halo mass
|
|
44
|
+
re-indexes geometrically (`mass<=4 || powerOfTwo`) and encodes 2-bit quantized.
|
|
45
|
+
Canon index is optional: `canonAdd`/ `canonFind`/`canonCount` over 32-bit
|
|
46
|
+
canonical hashes, caller verifies bytes.
|
|
47
|
+
|
|
48
|
+
## Containment, batching, LRU
|
|
49
|
+
|
|
50
|
+
`addContainer(child,parent)` buffers per child; flush appends via
|
|
51
|
+
`_dbAppendContain` (packed pages, geometric merge) — never rewrite the whole
|
|
52
|
+
list. `containersSlice` pages through it. Edges and kids write through the same
|
|
53
|
+
deferred transaction.
|
|
54
|
+
|
|
55
|
+
Every in-memory cache is a `BoundedMap` with byte accounting and eviction (`lru`
|
|
56
|
+
vs `smallest` + `clock`/`reorder` recency). ANN reads
|
|
57
|
+
(`resonate`/`resonateHalo`) are content-addressed (`vecKey`) and dropped on any
|
|
58
|
+
index mutation; `RESonate_CACHE_MAX=4096`.
|
|
59
|
+
|
|
60
|
+
## Maintenance (incremental)
|
|
61
|
+
|
|
62
|
+
- `compactContentIndex(minParents)` — scans only entries since last watermark
|
|
63
|
+
(`_vecContentEntriesSince`), removes indexed-but-isolated nodes (`<minParents`
|
|
64
|
+
parents, no edges/halos), compacts the vector DB.
|
|
65
|
+
- `repairContentIndex(regenerateGist)` — walks `_dbEdgeOrHaloIds()` candidates
|
|
66
|
+
only, re-inserts missing bridge nodes whose gists were evicted before
|
|
67
|
+
indexing.
|
|
68
|
+
- `buildCanonIndex` (`_buildCanonIndex`) — iterates `eachContent(fromId)` and
|
|
69
|
+
`canonAdd`s; `fromId` makes refresh incremental.
|
|
70
|
+
|
|
71
|
+
Full scans (`parents()`, `next()`, `containers()`) are maintenance-only — hot
|
|
72
|
+
paths use `LIMIT`ed probes (`parentsFirst`/`nextFirst`/`prevFirst`), `has*`,
|
|
73
|
+
`prevCount`.
|
|
74
|
+
|
|
75
|
+
## Adding a backend
|
|
76
|
+
|
|
77
|
+
Implement every `protected abstract _db*`/`_vec*` in `src/store.ts` as a thin
|
|
78
|
+
wrapper around your storage. Keep `_dbGet*First`/`Slice` as real `LIMIT` queries
|
|
79
|
+
and `has*`/`COUNT` as point probes — never materialise-then-slice.
|
|
@@ -0,0 +1,79 @@
|
|
|
1
|
+
# Thresholds — derived, never tuned
|
|
2
|
+
|
|
3
|
+
Every decision cutoff is a formula over `D` (vector dimension), `W` (`maxGroup`,
|
|
4
|
+
perception window), or `N` (corpus size). No threshold is tuned or added to
|
|
5
|
+
`src/config.ts`.
|
|
6
|
+
|
|
7
|
+
## Source of truth
|
|
8
|
+
|
|
9
|
+
- `src/geometry.ts` — all similarity/decision thresholds.
|
|
10
|
+
- `src/mind/traverse.ts` — corpus-scale readings that parameterise bounded
|
|
11
|
+
walks.
|
|
12
|
+
- `src/sema.ts` — positional coordinate algebra.
|
|
13
|
+
- `src/config.ts` — capacities and budgets only (cache byte budgets, batch
|
|
14
|
+
sizes, index parameters, query `k`, ALU precision, seed). Never a threshold.
|
|
15
|
+
|
|
16
|
+
## Geometry thresholds (`src/geometry.ts`)
|
|
17
|
+
|
|
18
|
+
| Symbol | Definition | Formula |
|
|
19
|
+
| ----------------------- | ---------------------------------------------------------------------------- | ----------------------------------- |
|
|
20
|
+
| `mergeThreshold(D)` | Store identity bar — cosine at which `intern` treats two gists as same node | `1 - 1/√D` |
|
|
21
|
+
| `identityBar(D,W,len)` | Scale-aware whole-span identity claim | `max(mergeThreshold(D), 1 - W/len)` |
|
|
22
|
+
| `reachThreshold(W)` | Recall confidence floor — half a river quantum | `1 - 1/(2·W)` |
|
|
23
|
+
| `estimatorNoise(D)` | RaBitQ noise floor — 1σ of random cosine | `1/√D` |
|
|
24
|
+
| `significanceBar(D)` | Whole-query relatedness — 3σ above chance | `3/√D` |
|
|
25
|
+
| `conceptThreshold(D)` | Halo concept sharing — structural midpoint + ½σ | `0.5 + 0.5/√D` |
|
|
26
|
+
| `dominates(part,whole)` | Half-dominance predicate | `part*2 > whole` |
|
|
27
|
+
| `profileCapacity(D)` | Superposition capacity — terms before readout collapses | `floor(√D)` (min 1) |
|
|
28
|
+
| `consensusFloor(N)` | Pooled-vote significance floor | `ln(N) + ½` |
|
|
29
|
+
| `coverageBar(_,D)` | Reach-index gating (currently unused hot-path; batch compaction replaces it) | `conceptThreshold(D)` |
|
|
30
|
+
|
|
31
|
+
`N` in `consensusFloor` is `corpusN` (edge-source count, floored at 2).
|
|
32
|
+
|
|
33
|
+
## Corpus-scale readings (`src/mind/traverse.ts`)
|
|
34
|
+
|
|
35
|
+
| Symbol | Definition | Formula |
|
|
36
|
+
| ---------------- | ------------------------------------------------------ | --------------------------- |
|
|
37
|
+
| `corpusN` | Distinct learnt contexts, floored | `max(2, edgeSourceCount())` |
|
|
38
|
+
| `hubBound` | Hub bound, used for every `LIMIT` read | `ceil(√max(2,N))` |
|
|
39
|
+
| `hubCap(ids)` | Fan-out cap — list-side reading of `hubBound` | `ids.slice(0, hubBound)` |
|
|
40
|
+
| `atomReach(N,W)` | Uniform-expectation floor on a byte atom's commonality | `max(1, ceil(N·W/256))` |
|
|
41
|
+
| `atomIsHub(N,W)` | Whether atom abstains as consensus voter | `atomReach > hubBound` |
|
|
42
|
+
|
|
43
|
+
`hubBound` is enforced at the store level (`nextFirst`, `parentsFirst`,
|
|
44
|
+
`containersSlice`, `hasNext`/`hasParents`, `bytesPrefix`, `chainRun`).
|
|
45
|
+
`atomReach` is the honest floor for atoms — they carry no kid/contain rows, so
|
|
46
|
+
their reach is unmeasurable and must not default to "maximally rare".
|
|
47
|
+
|
|
48
|
+
## Seat algebra (`src/sema.ts`)
|
|
49
|
+
|
|
50
|
+
`twoEndedSeat(seatCount, size, index)` — the one positional-coordinate algebra
|
|
51
|
+
shared by perception, `fold`, and every synthetic/canonical fold. First half
|
|
52
|
+
uses low seats, second half uses high seats:
|
|
53
|
+
`index < (size+1)/2 ? index : seatCount-size+index`.
|
|
54
|
+
|
|
55
|
+
## Two derivations that bite
|
|
56
|
+
|
|
57
|
+
**1. `identityBar` is scale-aware.** A fixed cosine `1-1/√D` over a `4·√D`-byte
|
|
58
|
+
span tolerates four whole windows of foreign bytes while still claiming
|
|
59
|
+
"near-identical". An identity claim may tolerate at most one window `W` (the
|
|
60
|
+
perception quantum, same budget as `differsByOneWindow` in near-dedup), so the
|
|
61
|
+
bar must be `1-W/len` floored at `mergeThreshold`. Reusing `mergeThreshold` for
|
|
62
|
+
a whole-span claim silently widens the byte budget with span length.
|
|
63
|
+
|
|
64
|
+
**2. `consensusFloor` is priced for pooled climb votes, not for support
|
|
65
|
+
counts.** Each region contributes at most `ln(N/c) ≤ ln(N)`; `ln(N)+½` demands
|
|
66
|
+
corroboration beyond one maximally-specific region. `chooseNext`'s `bestSupport`
|
|
67
|
+
(`prevCount` of one destination) is N-invariant — bounded by retellings of that
|
|
68
|
+
fact, not by `N`. Gating it against `consensusFloor` guarantees failure once `N`
|
|
69
|
+
is large enough (observed: 2-vs-1-1-1 corroboration refused at N≈325K, falling
|
|
70
|
+
back to a noisy concept-hop).
|
|
71
|
+
|
|
72
|
+
## Pins
|
|
73
|
+
|
|
74
|
+
- `test/40-choosenext-scale-guard` — `consensusFloor` must not gate `chooseNext`
|
|
75
|
+
support counts.
|
|
76
|
+
- `test/64-two-ended-thresholds` — `mergeThreshold` / `identityBar` /
|
|
77
|
+
`reachThreshold` derivations and the scale-aware floor.
|
|
78
|
+
- `test/78-atom-hub-recognition-cliff` — `atomReach` / `atomIsHub` hub
|
|
79
|
+
abstention at scale.
|