@hviana/sema 0.7.2 → 0.7.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +95 -843
- package/README.md +11 -11
- package/dist/src/mind/mind.js +14 -0
- package/dist/src/store-sqlite.js +17 -0
- package/dist/src/store.d.ts +51 -4
- package/dist/src/store.js +81 -14
- package/docs/INDEX.md +71 -0
- package/docs/INVARIANTS.md +19 -0
- package/docs/architecture/bounded-reads.md +85 -0
- package/docs/architecture/caches.md +89 -0
- package/docs/architecture/commonality.md +45 -0
- package/docs/architecture/cost-model.md +71 -0
- package/docs/architecture/determinism.md +73 -0
- package/docs/architecture/exact-vs-approximate.md +47 -0
- package/docs/architecture/factored-machinery.md +28 -0
- package/docs/architecture/fold-contract.md +87 -0
- package/docs/architecture/halo-sketch.md +99 -0
- package/docs/architecture/match-project.md +62 -0
- package/docs/architecture/mechanism-market.md +95 -0
- package/docs/architecture/memoization.md +96 -0
- package/docs/architecture/meter.md +55 -0
- package/docs/architecture/saturation.md +92 -0
- package/docs/architecture/store.md +79 -0
- package/docs/architecture/thresholds.md +79 -0
- package/docs/failures/tempting-but-wrong.md +144 -0
- package/docs/harness/gates.md +56 -0
- package/docs/mechanisms/alu.md +75 -0
- package/docs/mechanisms/cast.md +75 -0
- package/docs/mechanisms/confluence.md +36 -0
- package/docs/mechanisms/cover.md +54 -0
- package/docs/mechanisms/extraction.md +53 -0
- package/docs/mechanisms/prefix-completion.md +54 -0
- package/docs/mechanisms/recall.md +69 -0
- package/docs/mechanisms/reference.md +58 -0
- package/jsr.json +1 -1
- package/package.json +1 -1
- package/src/mind/mind.ts +14 -0
- package/src/store-sqlite.ts +19 -0
- package/src/store.ts +92 -16
- package/test/89-completion-recursion.test.mjs +30 -10
- package/test/96-bytes-walk-termination.test.mjs +115 -0
- package/test/97-store-seed.test.mjs +105 -0
- package/HOW_IT_WORKS.md +0 -5836
|
@@ -0,0 +1,45 @@
|
|
|
1
|
+
# Two Measures of Commonality
|
|
2
|
+
|
|
3
|
+
Sema needs "what is shared" in two different populations. One is corpus-global
|
|
4
|
+
(how widely a structure is reused), the other is weave-local (what a local
|
|
5
|
+
cohort of overlapping forms agrees on). They use different data and different
|
|
6
|
+
formulas and must not be conflated.
|
|
7
|
+
|
|
8
|
+
## Corpus-global — `reachOf` + `dominates`
|
|
9
|
+
|
|
10
|
+
_Defined in `src/mind/traverse.ts` + `src/geometry.ts`; used by climb,
|
|
11
|
+
containment, IDF pooling._
|
|
12
|
+
|
|
13
|
+
For a node id, `reachOf(id, N)` counts how many learnt contexts contain it
|
|
14
|
+
(ancestor reach via capped graph walks, memoised per response in
|
|
15
|
+
`sharedReachMemo`). `dominates(reach, N)` then asks whether that reach is above
|
|
16
|
+
the corpus-determined majority threshold (derived in `geometry.ts` over `N`).
|
|
17
|
+
Intuition: minority reach discriminates (a filler), majority reach is
|
|
18
|
+
scaffolding. Powers the consensus climb, edge following, and vote pooling.
|
|
19
|
+
|
|
20
|
+
## Weave-local — `depth[]` + `MIN_WEAVE` + `dominates`
|
|
21
|
+
|
|
22
|
+
_Defined in `src/mind/match.ts` (`depth[]`, `MIN_WEAVE`, `frame`) and gated in
|
|
23
|
+
`src/mind/match.ts:frame`; used by CAST._
|
|
24
|
+
|
|
25
|
+
For an alignment weave, `depth[i]` counts how many aligned structures cover byte
|
|
26
|
+
`i` of the query. `MIN_WEAVE = 2` requires agreement beyond a pair (pair columns
|
|
27
|
+
are ambiguous with insertions/deletions), and `dominates(depth[i], aligned)`
|
|
28
|
+
requires agreement by a majority of the aligned cohort:
|
|
29
|
+
|
|
30
|
+
```
|
|
31
|
+
frame(i) ⇔ depth[i] > MIN_WEAVE ∧ dominates(depth[i], aligned)
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
This powers CAST's frame gate: what the local cohort shares vs what
|
|
35
|
+
differentiates one member. It never consults corpus reach.
|
|
36
|
+
|
|
37
|
+
The two measures answer different questions over different populations; CAST's
|
|
38
|
+
frame must not be replaced by a reach check and the climb must not be driven by
|
|
39
|
+
weave depth.
|
|
40
|
+
|
|
41
|
+
## Pins
|
|
42
|
+
|
|
43
|
+
- `test/17` — weave-local frame / `MIN_WEAVE` / `dominates` vs corpus-global
|
|
44
|
+
reach.
|
|
45
|
+
- `test/34` — containment and reach-driven disambiguation.
|
|
@@ -0,0 +1,71 @@
|
|
|
1
|
+
# Cost Model — One Currency
|
|
2
|
+
|
|
3
|
+
Every mechanism and every byte competes on one cost ladder defined in
|
|
4
|
+
`src/mind/graph-search.ts`. GraphSearch and `pipeline.ts:think` use the same
|
|
5
|
+
units, so a mechanism-level choice and a byte-level choice are the same kind of
|
|
6
|
+
decision: a lightest derivation.
|
|
7
|
+
|
|
8
|
+
## Ladder (`src/mind/graph-search.ts`)
|
|
9
|
+
|
|
10
|
+
| Cost | Value | Meaning |
|
|
11
|
+
| --------- | ------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
12
|
+
| `MICRO` | `1e-3` | Recognised advance (one `rec` bridge); per-byte unit of the A\* heuristic. A recomposed form's onward edge is also `MICRO`. |
|
|
13
|
+
| `STEP` | `1` | Every edge hop (first or fifth), every computed result, every projection. Charging every hop makes the lightest derivation the shortest chain. |
|
|
14
|
+
| `CONCEPT` | `10` | Halo-mediated act (synonym hop, consensus climb) and abandoning an edge chain early (`CONCEPT` above chain cost — genuine fixpoint at `+0` always beats giving up at same depth). |
|
|
15
|
+
| `PASS` | `1000` / byte | Carrying a byte nothing explains. Dominates everything so the search always prefers to recognise. |
|
|
16
|
+
|
|
17
|
+
Only the **ordering** `MICRO < STEP < CONCEPT < PASS` matters; any constants
|
|
18
|
+
with that order give the same derivations.
|
|
19
|
+
|
|
20
|
+
## Pipeline weighing (`src/mind/pipeline.ts:think`)
|
|
21
|
+
|
|
22
|
+
Mechanism candidates are weighed in the same ladder:
|
|
23
|
+
|
|
24
|
+
```
|
|
25
|
+
weight = moves + PASS * unaccounted_bytes
|
|
26
|
+
grade = floor(weight / STEP)
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
`unaccounted` is the query bytes no `accounted` span covers. Comparison is at
|
|
30
|
+
`STEP` resolution: lowest `grade` wins. At equal grade the candidate with fewer
|
|
31
|
+
`scaffolding` bytes (answer bytes lifted from unrecognised spans) wins; only
|
|
32
|
+
then does mechanism list order decide.
|
|
33
|
+
|
|
34
|
+
## Two semirings
|
|
35
|
+
|
|
36
|
+
- **(min, +) tropical** — lightest derivation in `GraphSearch` via `src/derive`
|
|
37
|
+
(`lightestDerivation`). Cost accumulates with `+`, choice selects `min`.
|
|
38
|
+
Powers `cover`/`form`/`out`, edge following, fusing, and the A\* agenda
|
|
39
|
+
(`g + h`).
|
|
40
|
+
|
|
41
|
+
- **(+, +) arithmetic** — evidence pooling in `src/mind/attention.ts:poolVotes`.
|
|
42
|
+
Each region's vote is an axiom; rules carry `Rule.combine = 'sum'` so costs to
|
|
43
|
+
the same anchor **add** rather than minimise. Powers IDF-weighted consensus,
|
|
44
|
+
`votes`/`votesIdf`/`support`, and `regionSupport`/`regionPeak`.
|
|
45
|
+
|
|
46
|
+
## Admissibility
|
|
47
|
+
|
|
48
|
+
The A\* heuristic is admissible and consistent:
|
|
49
|
+
|
|
50
|
+
```
|
|
51
|
+
h(it) = (queryLen - right) * MICRO
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
`right` is `p` for `cover(p)` or `j` for `form/out [i,j)`. `MICRO` is the
|
|
55
|
+
minimum per-byte cost in the ladder (every real per-byte cost is `>= MICRO`,
|
|
56
|
+
including `PASS`), and only the suffix past `right` is counted, so `h` never
|
|
57
|
+
exceeds the true remaining cost.
|
|
58
|
+
|
|
59
|
+
## Policy is not cost
|
|
60
|
+
|
|
61
|
+
"Computation always wins" is **not** priced into the ladder (a computed result
|
|
62
|
+
costs `STEP`, same as a learned edge). It is enforced by masking: `pipeline.ts`
|
|
63
|
+
removes recognised sites overlapped by a `ComputedResult` so the computation is
|
|
64
|
+
the sole completion there. Keep policy in callers; keep the engine neutral.
|
|
65
|
+
|
|
66
|
+
## Pins
|
|
67
|
+
|
|
68
|
+
- `test/52` — climb consensus instrumentation
|
|
69
|
+
- `test/53` — cross-region probe instrumentation
|
|
70
|
+
- `test/54` — evidence `k` instrumentation
|
|
71
|
+
- `test/55` — cost meter (`Meter`, `CostReport`, `searchPops`/`searchPushes`)
|
|
@@ -0,0 +1,73 @@
|
|
|
1
|
+
# Determinism — Same Seed + Same Deposits + Same Query ⇒ Same Bytes
|
|
2
|
+
|
|
3
|
+
## The law
|
|
4
|
+
|
|
5
|
+
> Same `seed` + same deposit order + same query ⇒ byte-identical answer.
|
|
6
|
+
|
|
7
|
+
Determinism is the product. Every code path that can reach output must be
|
|
8
|
+
deterministic given `(seed, store contents, query bytes)`.
|
|
9
|
+
|
|
10
|
+
## Forbidden
|
|
11
|
+
|
|
12
|
+
No `Math.random` or `Date.now` in behaviour, and no iteration over unordered
|
|
13
|
+
collections where order can reach output. Example-only uses
|
|
14
|
+
(`example/train_base`) are outside the library contract. If a test becomes
|
|
15
|
+
flaky, the contract was broken, not the test.
|
|
16
|
+
|
|
17
|
+
## All randomness flows from `seed`
|
|
18
|
+
|
|
19
|
+
`MindConfig.seed` (`src/config.ts:resolveConfig`, `DEFAULT_CONFIG`) is the sole
|
|
20
|
+
entropy root. Subsystems derive deterministically:
|
|
21
|
+
|
|
22
|
+
- **Alphabet** — `Alphabet` (`src/alphabet.ts`) via `rng` (`src/vec.ts:rng`)
|
|
23
|
+
seeded as `seed ^ seedMask`; builds 16→64→256 vectors by refinement.
|
|
24
|
+
- **Keyring / Space** — `Space.seats` (`src/sema.ts:Space`) via `makeKeyring`
|
|
25
|
+
(`src/vec.ts:makeKeyring`) and `rng` seeded from `seed` in `Mind`
|
|
26
|
+
(`src/mind/mind.ts`); `fold`/`twoEndedSeat`/`companySignature` are pure over
|
|
27
|
+
`Space`.
|
|
28
|
+
- **Vector indexes** — `VectorDatabase` (`src/rabitq-ivf/src/database.ts`) and
|
|
29
|
+
`Prng` (`src/rabitq-ivf/src/rabitq.ts`) seeded from config; insertion order is
|
|
30
|
+
the stored order, not a random choice.
|
|
31
|
+
|
|
32
|
+
No other PRNG source may affect grounding. Thresholds in `geometry.ts` are
|
|
33
|
+
derived from `D`/`W`/`N`, not sampled.
|
|
34
|
+
|
|
35
|
+
## Tie-breaks are corpus-determined
|
|
36
|
+
|
|
37
|
+
Every choice among equals bottoms out in a fixed ordering — insertion order or
|
|
38
|
+
lowest node id. The universal no-evidence fallback is **first-inserted**:
|
|
39
|
+
|
|
40
|
+
- `guidedFirst` (`src/mind/traverse.ts:guidedFirst`) — guided pick via
|
|
41
|
+
`chooseNext` else first-inserted edge (`nextFirst` LIMIT 1).
|
|
42
|
+
- `chooseNext` (`src/mind/traverse.ts:chooseNext`) — capped `nextFirst` read
|
|
43
|
+
(`hubBound`), ranked by `prevCount` then `haloMass`; equal ⇒ first-inserted.
|
|
44
|
+
- `chooseAmong` (`src/mind/traverse.ts:chooseAmong`) — `hubCap` + `argmaxCosine`
|
|
45
|
+
over `candidateGist`; first-inserted on tie via stable scan.
|
|
46
|
+
- `companySignature` (`src/sema.ts:companySignature`) — `rng(id ^ 0x9e3779b9)`,
|
|
47
|
+
i.e. seeded by node id, not observation order.
|
|
48
|
+
|
|
49
|
+
Last-inserted was once used in one place; it was a bug. Never reintroduce it.
|
|
50
|
+
|
|
51
|
+
## Memoization and trace must not break identity
|
|
52
|
+
|
|
53
|
+
Per-response memos (`Precomputed`, `perceiveMemo`, `recogniseMemo`, `climbMemo`,
|
|
54
|
+
`_resolvedSubtrees` via `foldTree`, `_edgeChoice`, `_gistCache` in
|
|
55
|
+
`src/mind/mind.ts` / `src/mind/pipeline-mechanism.ts` /
|
|
56
|
+
`src/mind/primitives.ts`) are sound because asking never writes. Only
|
|
57
|
+
`guidedNext`/`sharedReachMemo` are trace-bypassed;
|
|
58
|
+
`perceiveMemo`/`recogniseMemo`/`climbMemo` are always consulted — `foldTree`'s
|
|
59
|
+
subtree fast path skips `visit` (and thus site emission) for cached subtrees, so
|
|
60
|
+
bypassing makes `recognise` non-idempotent.
|
|
61
|
+
|
|
62
|
+
## Follow it
|
|
63
|
+
|
|
64
|
+
When you add any choice among equals, name the tie-break explicitly and make it
|
|
65
|
+
corpus-determined. Thread new randomness through `seed`-derived `rng`; never
|
|
66
|
+
call `Math.random`/`Date.now` on a behavioural path.
|
|
67
|
+
|
|
68
|
+
## Pins
|
|
69
|
+
|
|
70
|
+
- `test/42` pins recognition idempotence under trace — traced and untraced
|
|
71
|
+
`recognise` must return the same cached object and site count.
|
|
72
|
+
- Determinism suites — `test/03`, `test/04`, `test/08`, `test/20` and others
|
|
73
|
+
assert same seed + same training ⇒ byte-identical answers and stores.
|
|
@@ -0,0 +1,47 @@
|
|
|
1
|
+
# Exact vs Approximate — The Law and Its Five Ladders
|
|
2
|
+
|
|
3
|
+
Vector scores (`resonate` / `resonateHalo`) are RaBitQ **estimates**. They rank
|
|
4
|
+
candidates and gate broad regions; they never decide identity. Identity is
|
|
5
|
+
decided only by content-addressed lookup — `resolve` / `findLeaf` / `findBranch`
|
|
6
|
+
/ `canonResolve` — and by re-folding bytes to verify.
|
|
7
|
+
|
|
8
|
+
## The law
|
|
9
|
+
|
|
10
|
+
> Scores propose, bytes dispose.
|
|
11
|
+
|
|
12
|
+
Even recall's echo decision re-folds the top hit's bytes rather than trusting
|
|
13
|
+
the estimate it already has. No `score >= threshold` path may mint an identity
|
|
14
|
+
claim; thresholds derived in `geometry.ts` gate search breadth, not truth.
|
|
15
|
+
|
|
16
|
+
## Graded evidence ladders
|
|
17
|
+
|
|
18
|
+
Five subsystems share one shape — **exact → distributional → geometric** — with
|
|
19
|
+
earlier tiers strictly preferred. Never reorder tiers; never let an approximate
|
|
20
|
+
tier override an exact one.
|
|
21
|
+
|
|
22
|
+
| # | Site | Ladder (strong → weak) | File |
|
|
23
|
+
| - | ------------------ | ------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------- |
|
|
24
|
+
| 1 | `resolve` | exact content-addressed fold → `canonResolve` (equivalence class, hash-then-verify) | `mind/primitives.ts` |
|
|
25
|
+
| 2 | `locate` | exact bytes → halo role → gist | `mind/match.ts` |
|
|
26
|
+
| 3 | `alignGraded` | literal W-gram runs → halo-matched sites + climb proposals (weave) | `mind/match.ts` / `pipeline-mechanism.ts` |
|
|
27
|
+
| 4 | `bridge` | junction containers → edge → synonym → whole-gist | `mind/resonance.ts` |
|
|
28
|
+
| 5 | `crossRegionVotes` | exact containers → single synonym → double → `structuralResonance` (synthetic gist, gated hardest — no byte containment) | `mind/attention.ts` |
|
|
29
|
+
|
|
30
|
+
## Asymmetries (attention)
|
|
31
|
+
|
|
32
|
+
Two rules in `attention.ts` encode "exact decides" and must not be flattened:
|
|
33
|
+
|
|
34
|
+
- Only the **EXACT** tier may explain ordinary votes away.
|
|
35
|
+
- Only **container-backed** evidence may consume its endpoints.
|
|
36
|
+
|
|
37
|
+
## Pins
|
|
38
|
+
|
|
39
|
+
- `test/51` pins the cross-region tier ladder and its gating.
|
|
40
|
+
- Recognition idempotence under trace (`test/42`) depends on exact identity
|
|
41
|
+
remaining byte-determined, not score-determined.
|
|
42
|
+
|
|
43
|
+
## Adding a matcher
|
|
44
|
+
|
|
45
|
+
Add a tier to the shared family in `mind/match.ts` with a derived gate
|
|
46
|
+
(`geometry.ts`), never a private `score >= k` check. A new mechanism is a
|
|
47
|
+
`(matcher, direction, gate)` configuration over that family (§2.5).
|
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
# Factored Machinery — One Definition, Many Consumers
|
|
2
|
+
|
|
3
|
+
Every shared operation is defined once and imported many times. Duplicating it
|
|
4
|
+
forks the corpus contract; moving it hides who owns the gate.
|
|
5
|
+
|
|
6
|
+
For the match → project → gate family see `match-project.md`; for the two
|
|
7
|
+
commonality measures see `commonality.md`; for work accounting see `meter.md`.
|
|
8
|
+
|
|
9
|
+
## Single-definition contracts
|
|
10
|
+
|
|
11
|
+
| Symbol | Defined in | One fact |
|
|
12
|
+
| ------------------------------------------------------------- | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
13
|
+
| `contentLevels` | `src/geometry.ts` | Single boundary rule: cuts + levels from one rolling hash pass; every segmentation reads it. |
|
|
14
|
+
| `canonicalWindows` / `chainReach` / `leafIdRun` / `windowIds` | `src/mind/canonical.ts` | Write/read contract: training interns `W-1,W` windows, reading chains to `W²` and probes `W`-windows — drift silences recognition. |
|
|
15
|
+
| `junction.ts` + `WalkCache` | `src/mind/junction.ts` | Shared junction ascent (parents + containers) with bounded `√N·W` walk; `WalkCache` memoizes capped reads/parents/containers per response; bridge and attention share it. |
|
|
16
|
+
| `joinWithBridge` | `src/mind/resonance.ts` | One out-of-search assembly: `bridge(left,right)` or bare concat with `bridgeMiss` trace. |
|
|
17
|
+
| `dismissedKnownContent` | `src/mind/bridge.ts` | Pure attestation: any unaccounted `W`-window that resolves as known content — shared gap guard for substitution and CAST. |
|
|
18
|
+
| `sharedReachMemo` | `src/mind/traverse.ts` | One response-scoped `AncestorReach` memo (cleared on write and for traces); every `reachOf`/`edgeAncestors` consumer shares it. |
|
|
19
|
+
| `guidedFirst` | `src/mind/traverse.ts` | Guided-or-first answer bytes: `guidedNext` else first-inserted edge (`LIMIT 1`). |
|
|
20
|
+
| `leadsSomewhere` | `src/mind/traverse.ts` | Admission predicate: `hasNext` (cached) or `hasHalo`; sites that lead nowhere contribute no derivation. |
|
|
21
|
+
| `isChunk` | `src/sema.ts` | `kids !== null && kids.every(k=>k.kids===null)` — smallest grouped unit; governs regions, seams, indexing. |
|
|
22
|
+
| `twoEndedSeat` | `src/sema.ts` | One seat algebra: first half low seats, second half high seats; shared by perception, `fold`, and canonical folds. |
|
|
23
|
+
|
|
24
|
+
## Pins
|
|
25
|
+
|
|
26
|
+
- `test/47` — frame reading split (matcher vs gate vs inventory).
|
|
27
|
+
- `test/50` — CAST analog / consensus floor (dismissed content, `MIN_WEAVE` /
|
|
28
|
+
`dominates` frame, `carriesFillers` refusal).
|
|
@@ -0,0 +1,87 @@
|
|
|
1
|
+
# Fold Contract — One Tree For The Same Bytes
|
|
2
|
+
|
|
3
|
+
> **Law:** `perceiveDeposit` ≡ `perceive` — same bytes ⇒ same tree and same node
|
|
4
|
+
> id. Deposit imposes nothing (no boundaries, no turn convention). Geometry
|
|
5
|
+
> never sees conversation metadata.
|
|
6
|
+
|
|
7
|
+
## The identity
|
|
8
|
+
|
|
9
|
+
Perception is a pure function of the bytes. The deposit path and the inference
|
|
10
|
+
path compute the same content-defined fold for the same input, so a trained
|
|
11
|
+
context node and `resolve(query)` reach the same node. When the two sides
|
|
12
|
+
disagreed, alignment went quadratic (measured 5.2M cells on a 476-byte context
|
|
13
|
+
vs 0 when they agree) and cumulative contexts stopped resolving to what they
|
|
14
|
+
were trained as.
|
|
15
|
+
|
|
16
|
+
## Deposit imposes nothing
|
|
17
|
+
|
|
18
|
+
No boundaries, no turn convention, nothing read out of the bytes. Conversational
|
|
19
|
+
turn offsets are API metadata — they feed `ConversationState`, `answeredSpans`
|
|
20
|
+
and `currentTurnStart`; the geometry never sees them. Passing turn boundaries
|
|
21
|
+
into the fold is a correctness bug, not a tuning choice.
|
|
22
|
+
|
|
23
|
+
## Boundaries vs reuse — two problems
|
|
24
|
+
|
|
25
|
+
`contentFoldIncremental` and `stablePrefixFold` solve different problems;
|
|
26
|
+
conflating them is what once put an imposed boundary set on the inference path.
|
|
27
|
+
|
|
28
|
+
| Mechanism | What it buys | Cost / shape |
|
|
29
|
+
| ------------------------ | --------------------------------------------------------- | ----------------------------------------------------------------------------------------------- |
|
|
30
|
+
| `contentFoldIncremental` | Transparent segment reuse (cost only) | Imposes nothing; tree identical to the cold fold |
|
|
31
|
+
| `stablePrefixFold` | Caller-supplied cuts left-nested for prefix-ROOT identity | One prefix-ROOT per cut becomes an identical subtree (and same node id) inside the grown stream |
|
|
32
|
+
|
|
33
|
+
Both carry the same precondition: `prev` must be a fold of a byte-identical
|
|
34
|
+
prefix — reuse is keyed on `[start,end)` offsets, which cannot witness byte
|
|
35
|
+
agreement. A mismatched `prev` produced a wrong tree on 336 of 400 random
|
|
36
|
+
streams. `perceiveDeposit` discharges this via the prefix bytes as cache key; a
|
|
37
|
+
conversation advances only by append. A caller that cannot prove the prefix must
|
|
38
|
+
pass no `prev`.
|
|
39
|
+
|
|
40
|
+
## Identity must not depend on W or absolute offset
|
|
41
|
+
|
|
42
|
+
`contentLevels` is the single boundary rule (rolling hash over a bounded
|
|
43
|
+
window). Any grouping by index — stride, tile, fixed-arity row — reintroduces
|
|
44
|
+
the grid's phase bug: `riverFold` groups `W`-ary from byte 0, so the same byte
|
|
45
|
+
run is a different subtree at a different offset. The content-defined hash
|
|
46
|
+
removes this: a change upstream moves only the cut it falls inside; downstream
|
|
47
|
+
cuts and segments are unchanged (99.7% cuts preserved on real deposits after
|
|
48
|
+
shifts of 1..7 bytes vs 14.3% for the grid). The `groupByLevel` above the
|
|
49
|
+
segments splits by content level, not by count.
|
|
50
|
+
|
|
51
|
+
## `contentLevels` is single source; `contentBoundaries` is projection
|
|
52
|
+
|
|
53
|
+
`contentBoundaries(space, bytes)` is `contentLevels(space, bytes).cuts`. It once
|
|
54
|
+
carried its own rolling-hash loop, which is how a write side and a read side
|
|
55
|
+
drift without a type error. Levels are read from the hash the cut was accepted
|
|
56
|
+
at — level `L` when `h` vanishes mod `W^(L+1)` — so level-`L` cuts nest inside
|
|
57
|
+
level-`(L-1)` and expected span is `W^(L+1)` bytes.
|
|
58
|
+
|
|
59
|
+
## Optional canonical capability
|
|
60
|
+
|
|
61
|
+
`canonAdd`/`canonFind` (`src/store.ts` — `canonCount`/`eachContent`) is an
|
|
62
|
+
optional backend capability. A backend may omit all four; resolution then has no
|
|
63
|
+
equivalence fallback. The store never learns the equivalence — the canonicalizer
|
|
64
|
+
(`Canon` in `src/canon.ts`, e.g. `textCanon`) is injected by the caller and
|
|
65
|
+
every candidate is hash-then-verified (re-canonicalize stored bytes, compare). A
|
|
66
|
+
hash collision costs a read, never a wrong id.
|
|
67
|
+
|
|
68
|
+
## Cost of changing the cut distribution
|
|
69
|
+
|
|
70
|
+
The cut rate, which bits are read, `minLen`/`maxLen`, and the forced cut at
|
|
71
|
+
`maxLen` set the segment distribution every downstream mechanism is fitted to.
|
|
72
|
+
Each has been changed experimentally and cost 5–21 tests (rate: 15–18, bits:
|
|
73
|
+
19–21, normalized chunking: 5–6). `W-1` is the minimum (one window minus one);
|
|
74
|
+
`seats.length` is the maximum (one flat node folds exactly one segment). The
|
|
75
|
+
forced cut is load-bearing — relaxing it to reduce the current 32% forced rate
|
|
76
|
+
looked like a tidying but broke the same suites. Re-measure the whole suite for
|
|
77
|
+
any change here.
|
|
78
|
+
|
|
79
|
+
## Pins
|
|
80
|
+
|
|
81
|
+
- `test/59` — shift invariance floors (content-defined cuts preserved over
|
|
82
|
+
random binary and prose).
|
|
83
|
+
- `test/63` — offset/W invariance and `contentLevels` distribution expectations.
|
|
84
|
+
|
|
85
|
+
See:
|
|
86
|
+
`src/geometry.ts:contentLevels`/`contentBoundaries`/`contentFoldIncremental`/`stablePrefixFold`;
|
|
87
|
+
`AGENTS.md` bootloader invariants.
|
|
@@ -0,0 +1,99 @@
|
|
|
1
|
+
# Halo & Sketch — Distributional Memory
|
|
2
|
+
|
|
3
|
+
A node's **halo** is its distributional signature: the superposition of
|
|
4
|
+
identity-bound company signatures poured from every episode it participated in.
|
|
5
|
+
Its **gist** is the VSA fold of its own bytes — content, not company.
|
|
6
|
+
|
|
7
|
+
## Two vectors per node, two indexes
|
|
8
|
+
|
|
9
|
+
| Vector | Encodes | Index | Query |
|
|
10
|
+
| ------ | ----------------------------- | ------------- | -------------- |
|
|
11
|
+
| gist | what the node _is made of_ | content index | `resonate` |
|
|
12
|
+
| halo | what _company_ the node keeps | halo index | `resonateHalo` |
|
|
13
|
+
|
|
14
|
+
Both indexes are RaBitQ-IVF (`src/rabitq-ivf/`) — 1-bit ANN over the same
|
|
15
|
+
vectors; scores are estimates, never identity. Halos are also persisted durably
|
|
16
|
+
(see below).
|
|
17
|
+
|
|
18
|
+
## Quantization — 2-bit on disk, float in session
|
|
19
|
+
|
|
20
|
+
A halo is a superposition of quasi-orthogonal signatures, so coordinates are
|
|
21
|
+
Gaussian. Storage exploits this:
|
|
22
|
+
|
|
23
|
+
- **In-session accumulator** (`_haloExact` in `src/store.ts`): `Float32Array` —
|
|
24
|
+
exact, additive, incremented by `pourHalo`.
|
|
25
|
+
- **Durable row** (`_dbUpsertHalo`): 2-bit Lloyd–Max quantizer — decision at
|
|
26
|
+
±0.9816σ, levels at ±0.4528σ / ±1.5104σ, σ derived from the stored norm
|
|
27
|
+
(`norm/√D`). Header is the 4-byte norm; body is 2 bits/coordinate. Keeps ≥0.88
|
|
28
|
+
correlation with the exact vector.
|
|
29
|
+
- **ANN index**: 1-bit RaBitQ, irreversible — answers only "which halos are near
|
|
30
|
+
this query?"
|
|
31
|
+
|
|
32
|
+
Re-indexing is geometric: a halo re-enters the ANN when mass is small or crosses
|
|
33
|
+
a power of two (`geometricMass`), so index writes are O(log mass).
|
|
34
|
+
|
|
35
|
+
## Bottom-k sketch & company profile
|
|
36
|
+
|
|
37
|
+
A whole-partner signature alone records tokens, not types — halos of genuine
|
|
38
|
+
synonyms would be quasi-orthogonal. `companyProfile` (`src/mind/learning.ts`)
|
|
39
|
+
superposes:
|
|
40
|
+
|
|
41
|
+
1. the partner's own identity signature, plus
|
|
42
|
+
2. the bottom-k **constituent sketch** — the
|
|
43
|
+
`k = profileCapacity(D) = floor(√D)` minimal units of its subtree with
|
|
44
|
+
smallest `unitPriority`, deduped.
|
|
45
|
+
|
|
46
|
+
The sketch is composable (bottom-k of a union = bottom-k of children's
|
|
47
|
+
sketches), durable derived state via `sketchGet`/`sketchPut`, and bounded: at
|
|
48
|
+
most `k` constituents are classified, each by one `LIMIT`ed parent read
|
|
49
|
+
(`hubBound`). Beyond `√D` terms a single constituent contributes less than
|
|
50
|
+
`1/√D` — below RaBitQ noise — and extra terms shrink every accepted one; the cap
|
|
51
|
+
is a correctness limit.
|
|
52
|
+
|
|
53
|
+
## Gist vectors
|
|
54
|
+
|
|
55
|
+
Folded by the river (`src/geometry.ts`): leaves are alphabet vectors, groups
|
|
56
|
+
bind by two-ended seats, intermediate gists stay unnormalized (magnitude ∝
|
|
57
|
+
√len), only the root is normalized. Gist resonance reads byte-proportional
|
|
58
|
+
overlap; halo resonance reads distributional overlap — the two are independent.
|
|
59
|
+
|
|
60
|
+
## Thresholds & gating
|
|
61
|
+
|
|
62
|
+
All bars are derived in `src/geometry.ts`; no tunable constant:
|
|
63
|
+
|
|
64
|
+
| Symbol | Formula | Use |
|
|
65
|
+
| ------------------ | -------------- | ------------------------------------------------------------------------------------------------- |
|
|
66
|
+
| `estimatorNoise` | `1/√D` | 1σ RaBitQ noise; contrastive margin must clear it |
|
|
67
|
+
| `significanceBar` | `3/√D` | whole-query relatedness — 3σ above chance; gates consensus climb and `analogyStrength` |
|
|
68
|
+
| `conceptThreshold` | `0.5 + 0.5/√D` | halo concept sharing — structural midpoint + ½σ; gates `haloSiblings`, concept hops, articulation |
|
|
69
|
+
|
|
70
|
+
The significance bar gates the whole query; `conceptThreshold` gates per-pair
|
|
71
|
+
halo cosine.
|
|
72
|
+
|
|
73
|
+
## Probes: `haloMass` and `hasHalo`
|
|
74
|
+
|
|
75
|
+
- `haloMass(id)` — count of poured episodes; evidence weight, tie-breaker in
|
|
76
|
+
`chooseAmong`.
|
|
77
|
+
- `hasHalo(id)` — existence probe (indexed point check, no vector decode);
|
|
78
|
+
mirrors `halo(id) !== null`. One tier of the `leadsSomewhere` admission
|
|
79
|
+
predicate (with `hasNext`/`hasParents`).
|
|
80
|
+
|
|
81
|
+
Both are `meter`-counted probes, not full decodes — `halo(id)` is the bounded
|
|
82
|
+
vector read; `resonateHalo` is the IVF ANN query.
|
|
83
|
+
|
|
84
|
+
## Relation to invariants
|
|
85
|
+
|
|
86
|
+
- **Derived thresholds** — all bars above live in `geometry.ts`.
|
|
87
|
+
- **Exact decides / approximate proposes** — halo scores rank and gate; identity
|
|
88
|
+
is content-addressed. The graded ladder is exact → halo → gist
|
|
89
|
+
(`mind/match.ts`).
|
|
90
|
+
- **Bounded reads** — `hasHalo`/`haloMass` are point probes; `resonateHalo` is
|
|
91
|
+
capped ANN; constituent classification uses `LIMIT hubBound+1` reads. No
|
|
92
|
+
per-query scan grows with corpus.
|
|
93
|
+
|
|
94
|
+
## Pins
|
|
95
|
+
|
|
96
|
+
- `test/08 storage halo` — halo persistence, 2-bit round-trip,
|
|
97
|
+
`haloMass`/`hasHalo` contract, index survival across reopen.
|
|
98
|
+
- `test/35 ivf` — RaBitQ-IVF contract (recall vs brute force, sublinear
|
|
99
|
+
scaling); covers the halo index's own layer.
|
|
@@ -0,0 +1,62 @@
|
|
|
1
|
+
# Match → Project → Gate
|
|
2
|
+
|
|
3
|
+
Every grounding mechanism is a configuration of one shared operation in
|
|
4
|
+
`src/mind/match.ts`. The family is defined once and imported many times;
|
|
5
|
+
duplicating it forks the corpus contract, moving it hides who owns the gate.
|
|
6
|
+
|
|
7
|
+
## The shared family in `mind/match.ts`
|
|
8
|
+
|
|
9
|
+
The match layer locates structure, the project layer moves along the store, and
|
|
10
|
+
the gate layer decides whether the shape licences voicing. All three are pure
|
|
11
|
+
functions over bytes and the store — no mechanism owns a private copy.
|
|
12
|
+
|
|
13
|
+
## The triple
|
|
14
|
+
|
|
15
|
+
| Role | Symbols | What it does |
|
|
16
|
+
| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------- |
|
|
17
|
+
| **Match** (locate structure) | `locate` (exact → halo → gist ladder), `alignRuns` (literal W-gram weave), `alignGraded` (literal + halo gaps), `alignAround` / `frameSlots` (seeded frame with contracted gaps), `bestHaloMate` (in-list halo), `analogyStrength` / `sharedFrameStrength` (distributional + structural analogy) | Finds where a query sits in a learnt form. |
|
|
18
|
+
| **Project** (direction) | `follow` (forward to fixpoint, first hop may `conceptHop`), `reverseContext` (reverse to context), `project` (forward else reverse), `conceptHop` (halo sibling with edge) | Moves along the store from the match — forward toward answers, reverse toward contexts. |
|
|
19
|
+
| **Gate** (structural licence) | `isSpanShaped` (sparse subsequence — open reading), `carriesFillers` (substitution carriage — strict voicing licence) | Decides whether the shape licences voicing. |
|
|
20
|
+
|
|
21
|
+
Mechanisms declare only `(matcher, direction, gate)`. Thresholds behind gates
|
|
22
|
+
live in `src/geometry.ts` — the match layer never invents a cutoff.
|
|
23
|
+
|
|
24
|
+
The graded ladder inside `locate` is exact → distributional → geometric:
|
|
25
|
+
content-addressed identity first, halo similarity second, gist resonance last.
|
|
26
|
+
Reordering the ladder or letting an approximate score override an exact hit is a
|
|
27
|
+
correctness bug (see `exact-vs-approximate.md`).
|
|
28
|
+
|
|
29
|
+
## Frame reading — matcher reports, gate judges, inventory elects nothing
|
|
30
|
+
|
|
31
|
+
`frameSlots` is the shared frame reader. It contracts every gap via
|
|
32
|
+
`contractGap` to its varying core, tags it `substitution` / `insertion` /
|
|
33
|
+
`deletion`, and attaches `covered` — the bytes the frame accounts for. It
|
|
34
|
+
applies no gate; it reports.
|
|
35
|
+
|
|
36
|
+
`carriesFillers` is the substitution gate. It judges byte-exactly:
|
|
37
|
+
|
|
38
|
+
```
|
|
39
|
+
substituteAll(contA, fillersA → fillersB) == contB
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
If the equality holds, voicing through the slot is a derivation; if not, the
|
|
43
|
+
slot cannot carry. This is the only place that decision is made.
|
|
44
|
+
|
|
45
|
+
`Precomputed.frames` is the inventory. It enumerates every frame pairing the
|
|
46
|
+
match layer finds and elects nothing — ranking and refusal belong to the
|
|
47
|
+
consumer.
|
|
48
|
+
|
|
49
|
+
## Voicing gates belong to the consumer
|
|
50
|
+
|
|
51
|
+
The shared layer never refuses on a consumer's behalf. Reference owns its four
|
|
52
|
+
gates: frame dominates the query, each slot reaches `W` on both sides, no
|
|
53
|
+
insertion/deletion, fillers pairwise distinct — plus `carriesFillers` on the
|
|
54
|
+
chosen pair. CAST, recall, and cover each apply their own gate over the same
|
|
55
|
+
shared inventory. Moving a consumer's gate into `match.ts` would hide who is
|
|
56
|
+
responsible for the refusal.
|
|
57
|
+
|
|
58
|
+
## Pins
|
|
59
|
+
|
|
60
|
+
- `test/47` — frame reading split (matcher vs gate vs inventory).
|
|
61
|
+
- `test/50` — CAST / reference voicing via `carriesFillers`.
|
|
62
|
+
- `test/24` / `test/76` — match/project family and span-shape readings.
|
|
@@ -0,0 +1,95 @@
|
|
|
1
|
+
# Mechanism Market — The Free-Will Architecture
|
|
2
|
+
|
|
3
|
+
Every grounding mechanism — including the ALU and user extensions — speaks one
|
|
4
|
+
interface (`mind/pipeline-mechanism.ts`). The decider in `mind/pipeline.ts`
|
|
5
|
+
(`think`) holds a plain list and never branches on which mechanism it holds.
|
|
6
|
+
|
|
7
|
+
## Interface
|
|
8
|
+
|
|
9
|
+
```ts
|
|
10
|
+
interface PipelineMechanism {
|
|
11
|
+
parse?(query: Uint8Array): Promise<ComputedSpan[]>; // authoritative spans
|
|
12
|
+
floor(ctx, query, pre, worthRunning): Promise<number | null>; // bound or null
|
|
13
|
+
run(ctx, query, pre): Promise<MechanismResult[]>; // candidates
|
|
14
|
+
}
|
|
15
|
+
interface MechanismResult {
|
|
16
|
+
bytes: Uint8Array;
|
|
17
|
+
accounted: [number, number][];
|
|
18
|
+
moves: number;
|
|
19
|
+
unexplained: string;
|
|
20
|
+
scaffolding?: number;
|
|
21
|
+
complete?: boolean;
|
|
22
|
+
}
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
- `parse` is optional; all results are collected into `Precomputed.computed`
|
|
26
|
+
before any `floor`/`run`.
|
|
27
|
+
- `floor` returns `null` when structurally impossible, otherwise an admissible
|
|
28
|
+
lower bound (never overstates cost).
|
|
29
|
+
- `run` returns candidates with travelling evidence (below).
|
|
30
|
+
|
|
31
|
+
## Decider
|
|
32
|
+
|
|
33
|
+
`think` in `mind/pipeline.ts` iterates `defaultMechanisms` in list order:
|
|
34
|
+
|
|
35
|
+
```
|
|
36
|
+
defaultMechanisms = [cover, cast, confluence, extraction, reference, recall,
|
|
37
|
+
prefix-completion] + ALU (`aluToMechanism`) + extensions
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
Weight is one currency: `weight = moves + PASS · unaccountedBytes` where
|
|
41
|
+
`unaccountedBytes = unexplainedSpans(query.length, accounted)`. Comparison is at
|
|
42
|
+
`STEP` grade (`grade = floor(weight/STEP)`); equal grade prefers fewer
|
|
43
|
+
`scaffolding` bytes, then list order.
|
|
44
|
+
|
|
45
|
+
## Four constraints
|
|
46
|
+
|
|
47
|
+
1. **Decoupled** — zero cross-imports between `mind/mechanisms/*`. Adding one
|
|
48
|
+
never touches another; no mechanism asks what already decided.
|
|
49
|
+
2. **Declared competence** — binary structural gates inside `floor`/`run` (query
|
|
50
|
+
length, anchor shape, weave existence). Never a learned score; rationale
|
|
51
|
+
states exactly why a mechanism abstained.
|
|
52
|
+
3. **Visible budget** — every corpus-scale loop is capped at a named constant:
|
|
53
|
+
`√N` via `hubBound`/`hubCap` and `k = 2·recallQueryK` (`Precomputed.k`).
|
|
54
|
+
Enforced at the store level.
|
|
55
|
+
4. **Evidence travels** — every candidate carries `accounted` (query spans
|
|
56
|
+
explained), `moves` (priced on `MICRO/STEP/CONCEPT/PASS`), `unexplained`
|
|
57
|
+
(diagnostic label); optionally `scaffolding` (answer bytes from unrecognised
|
|
58
|
+
spans — equal-grade tie-break) and `complete` (trained-form continuation
|
|
59
|
+
reached via identity; post-grounding must not extend). The decider honours
|
|
60
|
+
both without knowing who set them.
|
|
61
|
+
|
|
62
|
+
## Two disciplines
|
|
63
|
+
|
|
64
|
+
- **Admissible-floor pruning.** `floor` runs for every mechanism in list order
|
|
65
|
+
before any `run`. `run` fires only if `worthRunning(floor)` where
|
|
66
|
+
`worthRunning = (floor) => best === null || grade(floor) < grade(best.weight)`.
|
|
67
|
+
Cover runs first so a near-zero-cost computed span prunes the rest through the
|
|
68
|
+
same mechanism — not a special case.
|
|
69
|
+
|
|
70
|
+
- **Investment discipline.** `worthRunning` is passed _into_ `floor`. A floor
|
|
71
|
+
that would first-touch an expensive shared analysis (`pre.attention()` climb,
|
|
72
|
+
`pre.weave()`, `pre.resonance()`) checks `worthRunning(cheapestBound)` first
|
|
73
|
+
and returns the uninvested bound when it already loses. Never compute a shared
|
|
74
|
+
analysis just to discard it. `cast.ts`/`extraction.ts` are the references.
|
|
75
|
+
|
|
76
|
+
## Accounting
|
|
77
|
+
|
|
78
|
+
- **Extraction:** located frames are always evidence; the span between them
|
|
79
|
+
counts only when _both_ borders were located. An open-ended read is priced by
|
|
80
|
+
exclusion (`PASS`/byte).
|
|
81
|
+
- **Reverse reading:** `reverseContext` produces bytes but explains nothing
|
|
82
|
+
forward: `accounted = []`, weight ≈ `PASS·|query|` — last resort by
|
|
83
|
+
arithmetic, not rule.
|
|
84
|
+
- **Paid acts are accounted:** the bridge's corroborated substitutions cost
|
|
85
|
+
`CONCEPT` each in `moves`, so their spans must be `accounted`; otherwise the
|
|
86
|
+
same act is charged twice (`PASS`/byte dominates).
|
|
87
|
+
|
|
88
|
+
`accounted` is a cost-ladder quantity; `cover.ts` leaves masked computed spans
|
|
89
|
+
out of it so `PASS`-bridged bytes are still charged. `unexplained`,
|
|
90
|
+
`narrowDecision`, `thinGrounding` are observational only.
|
|
91
|
+
|
|
92
|
+
## Pins
|
|
93
|
+
|
|
94
|
+
- `test/01-floor` — floor geometry.
|
|
95
|
+
- `test/04-think` — decider, admissible pruning, investment discipline.
|