@hviana/sema 0.7.3 → 0.7.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +95 -843
- package/README.md +11 -11
- package/dist/src/mind/mind.js +14 -0
- package/dist/src/store-sqlite.js +17 -0
- package/dist/src/store.d.ts +18 -0
- package/dist/src/store.js +10 -0
- package/docs/INDEX.md +71 -0
- package/docs/INVARIANTS.md +19 -0
- package/docs/architecture/bounded-reads.md +85 -0
- package/docs/architecture/caches.md +89 -0
- package/docs/architecture/commonality.md +45 -0
- package/docs/architecture/cost-model.md +71 -0
- package/docs/architecture/determinism.md +73 -0
- package/docs/architecture/exact-vs-approximate.md +47 -0
- package/docs/architecture/factored-machinery.md +28 -0
- package/docs/architecture/fold-contract.md +87 -0
- package/docs/architecture/halo-sketch.md +99 -0
- package/docs/architecture/match-project.md +62 -0
- package/docs/architecture/mechanism-market.md +95 -0
- package/docs/architecture/memoization.md +96 -0
- package/docs/architecture/meter.md +55 -0
- package/docs/architecture/saturation.md +92 -0
- package/docs/architecture/store.md +79 -0
- package/docs/architecture/thresholds.md +79 -0
- package/docs/failures/tempting-but-wrong.md +144 -0
- package/docs/harness/gates.md +56 -0
- package/docs/mechanisms/alu.md +75 -0
- package/docs/mechanisms/cast.md +75 -0
- package/docs/mechanisms/confluence.md +36 -0
- package/docs/mechanisms/cover.md +54 -0
- package/docs/mechanisms/extraction.md +53 -0
- package/docs/mechanisms/prefix-completion.md +54 -0
- package/docs/mechanisms/recall.md +69 -0
- package/docs/mechanisms/reference.md +58 -0
- package/jsr.json +1 -1
- package/package.json +1 -1
- package/src/mind/mind.ts +14 -0
- package/src/store-sqlite.ts +19 -0
- package/src/store.ts +22 -0
- package/test/89-completion-recursion.test.mjs +30 -10
- package/test/97-store-seed.test.mjs +105 -0
- package/HOW_IT_WORKS.md +0 -5836
|
@@ -0,0 +1,47 @@
|
|
|
1
|
+
# Exact vs Approximate — The Law and Its Five Ladders
|
|
2
|
+
|
|
3
|
+
Vector scores (`resonate` / `resonateHalo`) are RaBitQ **estimates**. They rank
|
|
4
|
+
candidates and gate broad regions; they never decide identity. Identity is
|
|
5
|
+
decided only by content-addressed lookup — `resolve` / `findLeaf` / `findBranch`
|
|
6
|
+
/ `canonResolve` — and by re-folding bytes to verify.
|
|
7
|
+
|
|
8
|
+
## The law
|
|
9
|
+
|
|
10
|
+
> Scores propose, bytes dispose.
|
|
11
|
+
|
|
12
|
+
Even recall's echo decision re-folds the top hit's bytes rather than trusting
|
|
13
|
+
the estimate it already has. No `score >= threshold` path may mint an identity
|
|
14
|
+
claim; thresholds derived in `geometry.ts` gate search breadth, not truth.
|
|
15
|
+
|
|
16
|
+
## Graded evidence ladders
|
|
17
|
+
|
|
18
|
+
Five subsystems share one shape — **exact → distributional → geometric** — with
|
|
19
|
+
earlier tiers strictly preferred. Never reorder tiers; never let an approximate
|
|
20
|
+
tier override an exact one.
|
|
21
|
+
|
|
22
|
+
| # | Site | Ladder (strong → weak) | File |
|
|
23
|
+
| - | ------------------ | ------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------- |
|
|
24
|
+
| 1 | `resolve` | exact content-addressed fold → `canonResolve` (equivalence class, hash-then-verify) | `mind/primitives.ts` |
|
|
25
|
+
| 2 | `locate` | exact bytes → halo role → gist | `mind/match.ts` |
|
|
26
|
+
| 3 | `alignGraded` | literal W-gram runs → halo-matched sites + climb proposals (weave) | `mind/match.ts` / `pipeline-mechanism.ts` |
|
|
27
|
+
| 4 | `bridge` | junction containers → edge → synonym → whole-gist | `mind/resonance.ts` |
|
|
28
|
+
| 5 | `crossRegionVotes` | exact containers → single synonym → double → `structuralResonance` (synthetic gist, gated hardest — no byte containment) | `mind/attention.ts` |
|
|
29
|
+
|
|
30
|
+
## Asymmetries (attention)
|
|
31
|
+
|
|
32
|
+
Two rules in `attention.ts` encode "exact decides" and must not be flattened:
|
|
33
|
+
|
|
34
|
+
- Only the **EXACT** tier may explain ordinary votes away.
|
|
35
|
+
- Only **container-backed** evidence may consume its endpoints.
|
|
36
|
+
|
|
37
|
+
## Pins
|
|
38
|
+
|
|
39
|
+
- `test/51` pins the cross-region tier ladder and its gating.
|
|
40
|
+
- Recognition idempotence under trace (`test/42`) depends on exact identity
|
|
41
|
+
remaining byte-determined, not score-determined.
|
|
42
|
+
|
|
43
|
+
## Adding a matcher
|
|
44
|
+
|
|
45
|
+
Add a tier to the shared family in `mind/match.ts` with a derived gate
|
|
46
|
+
(`geometry.ts`), never a private `score >= k` check. A new mechanism is a
|
|
47
|
+
`(matcher, direction, gate)` configuration over that family (§2.5).
|
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
# Factored Machinery — One Definition, Many Consumers
|
|
2
|
+
|
|
3
|
+
Every shared operation is defined once and imported many times. Duplicating it
|
|
4
|
+
forks the corpus contract; moving it hides who owns the gate.
|
|
5
|
+
|
|
6
|
+
For the match → project → gate family see `match-project.md`; for the two
|
|
7
|
+
commonality measures see `commonality.md`; for work accounting see `meter.md`.
|
|
8
|
+
|
|
9
|
+
## Single-definition contracts
|
|
10
|
+
|
|
11
|
+
| Symbol | Defined in | One fact |
|
|
12
|
+
| ------------------------------------------------------------- | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
13
|
+
| `contentLevels` | `src/geometry.ts` | Single boundary rule: cuts + levels from one rolling hash pass; every segmentation reads it. |
|
|
14
|
+
| `canonicalWindows` / `chainReach` / `leafIdRun` / `windowIds` | `src/mind/canonical.ts` | Write/read contract: training interns `W-1,W` windows, reading chains to `W²` and probes `W`-windows — drift silences recognition. |
|
|
15
|
+
| `junction.ts` + `WalkCache` | `src/mind/junction.ts` | Shared junction ascent (parents + containers) with bounded `√N·W` walk; `WalkCache` memoizes capped reads/parents/containers per response; bridge and attention share it. |
|
|
16
|
+
| `joinWithBridge` | `src/mind/resonance.ts` | One out-of-search assembly: `bridge(left,right)` or bare concat with `bridgeMiss` trace. |
|
|
17
|
+
| `dismissedKnownContent` | `src/mind/bridge.ts` | Pure attestation: any unaccounted `W`-window that resolves as known content — shared gap guard for substitution and CAST. |
|
|
18
|
+
| `sharedReachMemo` | `src/mind/traverse.ts` | One response-scoped `AncestorReach` memo (cleared on write and for traces); every `reachOf`/`edgeAncestors` consumer shares it. |
|
|
19
|
+
| `guidedFirst` | `src/mind/traverse.ts` | Guided-or-first answer bytes: `guidedNext` else first-inserted edge (`LIMIT 1`). |
|
|
20
|
+
| `leadsSomewhere` | `src/mind/traverse.ts` | Admission predicate: `hasNext` (cached) or `hasHalo`; sites that lead nowhere contribute no derivation. |
|
|
21
|
+
| `isChunk` | `src/sema.ts` | `kids !== null && kids.every(k=>k.kids===null)` — smallest grouped unit; governs regions, seams, indexing. |
|
|
22
|
+
| `twoEndedSeat` | `src/sema.ts` | One seat algebra: first half low seats, second half high seats; shared by perception, `fold`, and canonical folds. |
|
|
23
|
+
|
|
24
|
+
## Pins
|
|
25
|
+
|
|
26
|
+
- `test/47` — frame reading split (matcher vs gate vs inventory).
|
|
27
|
+
- `test/50` — CAST analog / consensus floor (dismissed content, `MIN_WEAVE` /
|
|
28
|
+
`dominates` frame, `carriesFillers` refusal).
|
|
@@ -0,0 +1,87 @@
|
|
|
1
|
+
# Fold Contract — One Tree For The Same Bytes
|
|
2
|
+
|
|
3
|
+
> **Law:** `perceiveDeposit` ≡ `perceive` — same bytes ⇒ same tree and same node
|
|
4
|
+
> id. Deposit imposes nothing (no boundaries, no turn convention). Geometry
|
|
5
|
+
> never sees conversation metadata.
|
|
6
|
+
|
|
7
|
+
## The identity
|
|
8
|
+
|
|
9
|
+
Perception is a pure function of the bytes. The deposit path and the inference
|
|
10
|
+
path compute the same content-defined fold for the same input, so a trained
|
|
11
|
+
context node and `resolve(query)` reach the same node. When the two sides
|
|
12
|
+
disagreed, alignment went quadratic (measured 5.2M cells on a 476-byte context
|
|
13
|
+
vs 0 when they agree) and cumulative contexts stopped resolving to what they
|
|
14
|
+
were trained as.
|
|
15
|
+
|
|
16
|
+
## Deposit imposes nothing
|
|
17
|
+
|
|
18
|
+
No boundaries, no turn convention, nothing read out of the bytes. Conversational
|
|
19
|
+
turn offsets are API metadata — they feed `ConversationState`, `answeredSpans`
|
|
20
|
+
and `currentTurnStart`; the geometry never sees them. Passing turn boundaries
|
|
21
|
+
into the fold is a correctness bug, not a tuning choice.
|
|
22
|
+
|
|
23
|
+
## Boundaries vs reuse — two problems
|
|
24
|
+
|
|
25
|
+
`contentFoldIncremental` and `stablePrefixFold` solve different problems;
|
|
26
|
+
conflating them is what once put an imposed boundary set on the inference path.
|
|
27
|
+
|
|
28
|
+
| Mechanism | What it buys | Cost / shape |
|
|
29
|
+
| ------------------------ | --------------------------------------------------------- | ----------------------------------------------------------------------------------------------- |
|
|
30
|
+
| `contentFoldIncremental` | Transparent segment reuse (cost only) | Imposes nothing; tree identical to the cold fold |
|
|
31
|
+
| `stablePrefixFold` | Caller-supplied cuts left-nested for prefix-ROOT identity | One prefix-ROOT per cut becomes an identical subtree (and same node id) inside the grown stream |
|
|
32
|
+
|
|
33
|
+
Both carry the same precondition: `prev` must be a fold of a byte-identical
|
|
34
|
+
prefix — reuse is keyed on `[start,end)` offsets, which cannot witness byte
|
|
35
|
+
agreement. A mismatched `prev` produced a wrong tree on 336 of 400 random
|
|
36
|
+
streams. `perceiveDeposit` discharges this via the prefix bytes as cache key; a
|
|
37
|
+
conversation advances only by append. A caller that cannot prove the prefix must
|
|
38
|
+
pass no `prev`.
|
|
39
|
+
|
|
40
|
+
## Identity must not depend on W or absolute offset
|
|
41
|
+
|
|
42
|
+
`contentLevels` is the single boundary rule (rolling hash over a bounded
|
|
43
|
+
window). Any grouping by index — stride, tile, fixed-arity row — reintroduces
|
|
44
|
+
the grid's phase bug: `riverFold` groups `W`-ary from byte 0, so the same byte
|
|
45
|
+
run is a different subtree at a different offset. The content-defined hash
|
|
46
|
+
removes this: a change upstream moves only the cut it falls inside; downstream
|
|
47
|
+
cuts and segments are unchanged (99.7% cuts preserved on real deposits after
|
|
48
|
+
shifts of 1..7 bytes vs 14.3% for the grid). The `groupByLevel` above the
|
|
49
|
+
segments splits by content level, not by count.
|
|
50
|
+
|
|
51
|
+
## `contentLevels` is single source; `contentBoundaries` is projection
|
|
52
|
+
|
|
53
|
+
`contentBoundaries(space, bytes)` is `contentLevels(space, bytes).cuts`. It once
|
|
54
|
+
carried its own rolling-hash loop, which is how a write side and a read side
|
|
55
|
+
drift without a type error. Levels are read from the hash the cut was accepted
|
|
56
|
+
at — level `L` when `h` vanishes mod `W^(L+1)` — so level-`L` cuts nest inside
|
|
57
|
+
level-`(L-1)` and expected span is `W^(L+1)` bytes.
|
|
58
|
+
|
|
59
|
+
## Optional canonical capability
|
|
60
|
+
|
|
61
|
+
`canonAdd`/`canonFind` (`src/store.ts` — `canonCount`/`eachContent`) is an
|
|
62
|
+
optional backend capability. A backend may omit all four; resolution then has no
|
|
63
|
+
equivalence fallback. The store never learns the equivalence — the canonicalizer
|
|
64
|
+
(`Canon` in `src/canon.ts`, e.g. `textCanon`) is injected by the caller and
|
|
65
|
+
every candidate is hash-then-verified (re-canonicalize stored bytes, compare). A
|
|
66
|
+
hash collision costs a read, never a wrong id.
|
|
67
|
+
|
|
68
|
+
## Cost of changing the cut distribution
|
|
69
|
+
|
|
70
|
+
The cut rate, which bits are read, `minLen`/`maxLen`, and the forced cut at
|
|
71
|
+
`maxLen` set the segment distribution every downstream mechanism is fitted to.
|
|
72
|
+
Each has been changed experimentally and cost 5–21 tests (rate: 15–18, bits:
|
|
73
|
+
19–21, normalized chunking: 5–6). `W-1` is the minimum (one window minus one);
|
|
74
|
+
`seats.length` is the maximum (one flat node folds exactly one segment). The
|
|
75
|
+
forced cut is load-bearing — relaxing it to reduce the current 32% forced rate
|
|
76
|
+
looked like a tidying but broke the same suites. Re-measure the whole suite for
|
|
77
|
+
any change here.
|
|
78
|
+
|
|
79
|
+
## Pins
|
|
80
|
+
|
|
81
|
+
- `test/59` — shift invariance floors (content-defined cuts preserved over
|
|
82
|
+
random binary and prose).
|
|
83
|
+
- `test/63` — offset/W invariance and `contentLevels` distribution expectations.
|
|
84
|
+
|
|
85
|
+
See:
|
|
86
|
+
`src/geometry.ts:contentLevels`/`contentBoundaries`/`contentFoldIncremental`/`stablePrefixFold`;
|
|
87
|
+
`AGENTS.md` bootloader invariants.
|
|
@@ -0,0 +1,99 @@
|
|
|
1
|
+
# Halo & Sketch — Distributional Memory
|
|
2
|
+
|
|
3
|
+
A node's **halo** is its distributional signature: the superposition of
|
|
4
|
+
identity-bound company signatures poured from every episode it participated in.
|
|
5
|
+
Its **gist** is the VSA fold of its own bytes — content, not company.
|
|
6
|
+
|
|
7
|
+
## Two vectors per node, two indexes
|
|
8
|
+
|
|
9
|
+
| Vector | Encodes | Index | Query |
|
|
10
|
+
| ------ | ----------------------------- | ------------- | -------------- |
|
|
11
|
+
| gist | what the node _is made of_ | content index | `resonate` |
|
|
12
|
+
| halo | what _company_ the node keeps | halo index | `resonateHalo` |
|
|
13
|
+
|
|
14
|
+
Both indexes are RaBitQ-IVF (`src/rabitq-ivf/`) — 1-bit ANN over the same
|
|
15
|
+
vectors; scores are estimates, never identity. Halos are also persisted durably
|
|
16
|
+
(see below).
|
|
17
|
+
|
|
18
|
+
## Quantization — 2-bit on disk, float in session
|
|
19
|
+
|
|
20
|
+
A halo is a superposition of quasi-orthogonal signatures, so coordinates are
|
|
21
|
+
Gaussian. Storage exploits this:
|
|
22
|
+
|
|
23
|
+
- **In-session accumulator** (`_haloExact` in `src/store.ts`): `Float32Array` —
|
|
24
|
+
exact, additive, incremented by `pourHalo`.
|
|
25
|
+
- **Durable row** (`_dbUpsertHalo`): 2-bit Lloyd–Max quantizer — decision at
|
|
26
|
+
±0.9816σ, levels at ±0.4528σ / ±1.5104σ, σ derived from the stored norm
|
|
27
|
+
(`norm/√D`). Header is the 4-byte norm; body is 2 bits/coordinate. Keeps ≥0.88
|
|
28
|
+
correlation with the exact vector.
|
|
29
|
+
- **ANN index**: 1-bit RaBitQ, irreversible — answers only "which halos are near
|
|
30
|
+
this query?"
|
|
31
|
+
|
|
32
|
+
Re-indexing is geometric: a halo re-enters the ANN when mass is small or crosses
|
|
33
|
+
a power of two (`geometricMass`), so index writes are O(log mass).
|
|
34
|
+
|
|
35
|
+
## Bottom-k sketch & company profile
|
|
36
|
+
|
|
37
|
+
A whole-partner signature alone records tokens, not types — halos of genuine
|
|
38
|
+
synonyms would be quasi-orthogonal. `companyProfile` (`src/mind/learning.ts`)
|
|
39
|
+
superposes:
|
|
40
|
+
|
|
41
|
+
1. the partner's own identity signature, plus
|
|
42
|
+
2. the bottom-k **constituent sketch** — the
|
|
43
|
+
`k = profileCapacity(D) = floor(√D)` minimal units of its subtree with
|
|
44
|
+
smallest `unitPriority`, deduped.
|
|
45
|
+
|
|
46
|
+
The sketch is composable (bottom-k of a union = bottom-k of children's
|
|
47
|
+
sketches), durable derived state via `sketchGet`/`sketchPut`, and bounded: at
|
|
48
|
+
most `k` constituents are classified, each by one `LIMIT`ed parent read
|
|
49
|
+
(`hubBound`). Beyond `√D` terms a single constituent contributes less than
|
|
50
|
+
`1/√D` — below RaBitQ noise — and extra terms shrink every accepted one; the cap
|
|
51
|
+
is a correctness limit.
|
|
52
|
+
|
|
53
|
+
## Gist vectors
|
|
54
|
+
|
|
55
|
+
Folded by the river (`src/geometry.ts`): leaves are alphabet vectors, groups
|
|
56
|
+
bind by two-ended seats, intermediate gists stay unnormalized (magnitude ∝
|
|
57
|
+
√len), only the root is normalized. Gist resonance reads byte-proportional
|
|
58
|
+
overlap; halo resonance reads distributional overlap — the two are independent.
|
|
59
|
+
|
|
60
|
+
## Thresholds & gating
|
|
61
|
+
|
|
62
|
+
All bars are derived in `src/geometry.ts`; no tunable constant:
|
|
63
|
+
|
|
64
|
+
| Symbol | Formula | Use |
|
|
65
|
+
| ------------------ | -------------- | ------------------------------------------------------------------------------------------------- |
|
|
66
|
+
| `estimatorNoise` | `1/√D` | 1σ RaBitQ noise; contrastive margin must clear it |
|
|
67
|
+
| `significanceBar` | `3/√D` | whole-query relatedness — 3σ above chance; gates consensus climb and `analogyStrength` |
|
|
68
|
+
| `conceptThreshold` | `0.5 + 0.5/√D` | halo concept sharing — structural midpoint + ½σ; gates `haloSiblings`, concept hops, articulation |
|
|
69
|
+
|
|
70
|
+
The significance bar gates the whole query; `conceptThreshold` gates per-pair
|
|
71
|
+
halo cosine.
|
|
72
|
+
|
|
73
|
+
## Probes: `haloMass` and `hasHalo`
|
|
74
|
+
|
|
75
|
+
- `haloMass(id)` — count of poured episodes; evidence weight, tie-breaker in
|
|
76
|
+
`chooseAmong`.
|
|
77
|
+
- `hasHalo(id)` — existence probe (indexed point check, no vector decode);
|
|
78
|
+
mirrors `halo(id) !== null`. One tier of the `leadsSomewhere` admission
|
|
79
|
+
predicate (with `hasNext`/`hasParents`).
|
|
80
|
+
|
|
81
|
+
Both are `meter`-counted probes, not full decodes — `halo(id)` is the bounded
|
|
82
|
+
vector read; `resonateHalo` is the IVF ANN query.
|
|
83
|
+
|
|
84
|
+
## Relation to invariants
|
|
85
|
+
|
|
86
|
+
- **Derived thresholds** — all bars above live in `geometry.ts`.
|
|
87
|
+
- **Exact decides / approximate proposes** — halo scores rank and gate; identity
|
|
88
|
+
is content-addressed. The graded ladder is exact → halo → gist
|
|
89
|
+
(`mind/match.ts`).
|
|
90
|
+
- **Bounded reads** — `hasHalo`/`haloMass` are point probes; `resonateHalo` is
|
|
91
|
+
capped ANN; constituent classification uses `LIMIT hubBound+1` reads. No
|
|
92
|
+
per-query scan grows with corpus.
|
|
93
|
+
|
|
94
|
+
## Pins
|
|
95
|
+
|
|
96
|
+
- `test/08 storage halo` — halo persistence, 2-bit round-trip,
|
|
97
|
+
`haloMass`/`hasHalo` contract, index survival across reopen.
|
|
98
|
+
- `test/35 ivf` — RaBitQ-IVF contract (recall vs brute force, sublinear
|
|
99
|
+
scaling); covers the halo index's own layer.
|
|
@@ -0,0 +1,62 @@
|
|
|
1
|
+
# Match → Project → Gate
|
|
2
|
+
|
|
3
|
+
Every grounding mechanism is a configuration of one shared operation in
|
|
4
|
+
`src/mind/match.ts`. The family is defined once and imported many times;
|
|
5
|
+
duplicating it forks the corpus contract, moving it hides who owns the gate.
|
|
6
|
+
|
|
7
|
+
## The shared family in `mind/match.ts`
|
|
8
|
+
|
|
9
|
+
The match layer locates structure, the project layer moves along the store, and
|
|
10
|
+
the gate layer decides whether the shape licences voicing. All three are pure
|
|
11
|
+
functions over bytes and the store — no mechanism owns a private copy.
|
|
12
|
+
|
|
13
|
+
## The triple
|
|
14
|
+
|
|
15
|
+
| Role | Symbols | What it does |
|
|
16
|
+
| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------- |
|
|
17
|
+
| **Match** (locate structure) | `locate` (exact → halo → gist ladder), `alignRuns` (literal W-gram weave), `alignGraded` (literal + halo gaps), `alignAround` / `frameSlots` (seeded frame with contracted gaps), `bestHaloMate` (in-list halo), `analogyStrength` / `sharedFrameStrength` (distributional + structural analogy) | Finds where a query sits in a learnt form. |
|
|
18
|
+
| **Project** (direction) | `follow` (forward to fixpoint, first hop may `conceptHop`), `reverseContext` (reverse to context), `project` (forward else reverse), `conceptHop` (halo sibling with edge) | Moves along the store from the match — forward toward answers, reverse toward contexts. |
|
|
19
|
+
| **Gate** (structural licence) | `isSpanShaped` (sparse subsequence — open reading), `carriesFillers` (substitution carriage — strict voicing licence) | Decides whether the shape licences voicing. |
|
|
20
|
+
|
|
21
|
+
Mechanisms declare only `(matcher, direction, gate)`. Thresholds behind gates
|
|
22
|
+
live in `src/geometry.ts` — the match layer never invents a cutoff.
|
|
23
|
+
|
|
24
|
+
The graded ladder inside `locate` is exact → distributional → geometric:
|
|
25
|
+
content-addressed identity first, halo similarity second, gist resonance last.
|
|
26
|
+
Reordering the ladder or letting an approximate score override an exact hit is a
|
|
27
|
+
correctness bug (see `exact-vs-approximate.md`).
|
|
28
|
+
|
|
29
|
+
## Frame reading — matcher reports, gate judges, inventory elects nothing
|
|
30
|
+
|
|
31
|
+
`frameSlots` is the shared frame reader. It contracts every gap via
|
|
32
|
+
`contractGap` to its varying core, tags it `substitution` / `insertion` /
|
|
33
|
+
`deletion`, and attaches `covered` — the bytes the frame accounts for. It
|
|
34
|
+
applies no gate; it reports.
|
|
35
|
+
|
|
36
|
+
`carriesFillers` is the substitution gate. It judges byte-exactly:
|
|
37
|
+
|
|
38
|
+
```
|
|
39
|
+
substituteAll(contA, fillersA → fillersB) == contB
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
If the equality holds, voicing through the slot is a derivation; if not, the
|
|
43
|
+
slot cannot carry. This is the only place that decision is made.
|
|
44
|
+
|
|
45
|
+
`Precomputed.frames` is the inventory. It enumerates every frame pairing the
|
|
46
|
+
match layer finds and elects nothing — ranking and refusal belong to the
|
|
47
|
+
consumer.
|
|
48
|
+
|
|
49
|
+
## Voicing gates belong to the consumer
|
|
50
|
+
|
|
51
|
+
The shared layer never refuses on a consumer's behalf. Reference owns its four
|
|
52
|
+
gates: frame dominates the query, each slot reaches `W` on both sides, no
|
|
53
|
+
insertion/deletion, fillers pairwise distinct — plus `carriesFillers` on the
|
|
54
|
+
chosen pair. CAST, recall, and cover each apply their own gate over the same
|
|
55
|
+
shared inventory. Moving a consumer's gate into `match.ts` would hide who is
|
|
56
|
+
responsible for the refusal.
|
|
57
|
+
|
|
58
|
+
## Pins
|
|
59
|
+
|
|
60
|
+
- `test/47` — frame reading split (matcher vs gate vs inventory).
|
|
61
|
+
- `test/50` — CAST / reference voicing via `carriesFillers`.
|
|
62
|
+
- `test/24` / `test/76` — match/project family and span-shape readings.
|
|
@@ -0,0 +1,95 @@
|
|
|
1
|
+
# Mechanism Market — The Free-Will Architecture
|
|
2
|
+
|
|
3
|
+
Every grounding mechanism — including the ALU and user extensions — speaks one
|
|
4
|
+
interface (`mind/pipeline-mechanism.ts`). The decider in `mind/pipeline.ts`
|
|
5
|
+
(`think`) holds a plain list and never branches on which mechanism it holds.
|
|
6
|
+
|
|
7
|
+
## Interface
|
|
8
|
+
|
|
9
|
+
```ts
|
|
10
|
+
interface PipelineMechanism {
|
|
11
|
+
parse?(query: Uint8Array): Promise<ComputedSpan[]>; // authoritative spans
|
|
12
|
+
floor(ctx, query, pre, worthRunning): Promise<number | null>; // bound or null
|
|
13
|
+
run(ctx, query, pre): Promise<MechanismResult[]>; // candidates
|
|
14
|
+
}
|
|
15
|
+
interface MechanismResult {
|
|
16
|
+
bytes: Uint8Array;
|
|
17
|
+
accounted: [number, number][];
|
|
18
|
+
moves: number;
|
|
19
|
+
unexplained: string;
|
|
20
|
+
scaffolding?: number;
|
|
21
|
+
complete?: boolean;
|
|
22
|
+
}
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
- `parse` is optional; all results are collected into `Precomputed.computed`
|
|
26
|
+
before any `floor`/`run`.
|
|
27
|
+
- `floor` returns `null` when structurally impossible, otherwise an admissible
|
|
28
|
+
lower bound (never overstates cost).
|
|
29
|
+
- `run` returns candidates with travelling evidence (below).
|
|
30
|
+
|
|
31
|
+
## Decider
|
|
32
|
+
|
|
33
|
+
`think` in `mind/pipeline.ts` iterates `defaultMechanisms` in list order:
|
|
34
|
+
|
|
35
|
+
```
|
|
36
|
+
defaultMechanisms = [cover, cast, confluence, extraction, reference, recall,
|
|
37
|
+
prefix-completion] + ALU (`aluToMechanism`) + extensions
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
Weight is one currency: `weight = moves + PASS · unaccountedBytes` where
|
|
41
|
+
`unaccountedBytes = unexplainedSpans(query.length, accounted)`. Comparison is at
|
|
42
|
+
`STEP` grade (`grade = floor(weight/STEP)`); equal grade prefers fewer
|
|
43
|
+
`scaffolding` bytes, then list order.
|
|
44
|
+
|
|
45
|
+
## Four constraints
|
|
46
|
+
|
|
47
|
+
1. **Decoupled** — zero cross-imports between `mind/mechanisms/*`. Adding one
|
|
48
|
+
never touches another; no mechanism asks what already decided.
|
|
49
|
+
2. **Declared competence** — binary structural gates inside `floor`/`run` (query
|
|
50
|
+
length, anchor shape, weave existence). Never a learned score; rationale
|
|
51
|
+
states exactly why a mechanism abstained.
|
|
52
|
+
3. **Visible budget** — every corpus-scale loop is capped at a named constant:
|
|
53
|
+
`√N` via `hubBound`/`hubCap` and `k = 2·recallQueryK` (`Precomputed.k`).
|
|
54
|
+
Enforced at the store level.
|
|
55
|
+
4. **Evidence travels** — every candidate carries `accounted` (query spans
|
|
56
|
+
explained), `moves` (priced on `MICRO/STEP/CONCEPT/PASS`), `unexplained`
|
|
57
|
+
(diagnostic label); optionally `scaffolding` (answer bytes from unrecognised
|
|
58
|
+
spans — equal-grade tie-break) and `complete` (trained-form continuation
|
|
59
|
+
reached via identity; post-grounding must not extend). The decider honours
|
|
60
|
+
both without knowing who set them.
|
|
61
|
+
|
|
62
|
+
## Two disciplines
|
|
63
|
+
|
|
64
|
+
- **Admissible-floor pruning.** `floor` runs for every mechanism in list order
|
|
65
|
+
before any `run`. `run` fires only if `worthRunning(floor)` where
|
|
66
|
+
`worthRunning = (floor) => best === null || grade(floor) < grade(best.weight)`.
|
|
67
|
+
Cover runs first so a near-zero-cost computed span prunes the rest through the
|
|
68
|
+
same mechanism — not a special case.
|
|
69
|
+
|
|
70
|
+
- **Investment discipline.** `worthRunning` is passed _into_ `floor`. A floor
|
|
71
|
+
that would first-touch an expensive shared analysis (`pre.attention()` climb,
|
|
72
|
+
`pre.weave()`, `pre.resonance()`) checks `worthRunning(cheapestBound)` first
|
|
73
|
+
and returns the uninvested bound when it already loses. Never compute a shared
|
|
74
|
+
analysis just to discard it. `cast.ts`/`extraction.ts` are the references.
|
|
75
|
+
|
|
76
|
+
## Accounting
|
|
77
|
+
|
|
78
|
+
- **Extraction:** located frames are always evidence; the span between them
|
|
79
|
+
counts only when _both_ borders were located. An open-ended read is priced by
|
|
80
|
+
exclusion (`PASS`/byte).
|
|
81
|
+
- **Reverse reading:** `reverseContext` produces bytes but explains nothing
|
|
82
|
+
forward: `accounted = []`, weight ≈ `PASS·|query|` — last resort by
|
|
83
|
+
arithmetic, not rule.
|
|
84
|
+
- **Paid acts are accounted:** the bridge's corroborated substitutions cost
|
|
85
|
+
`CONCEPT` each in `moves`, so their spans must be `accounted`; otherwise the
|
|
86
|
+
same act is charged twice (`PASS`/byte dominates).
|
|
87
|
+
|
|
88
|
+
`accounted` is a cost-ladder quantity; `cover.ts` leaves masked computed spans
|
|
89
|
+
out of it so `PASS`-bridged bytes are still charged. `unexplained`,
|
|
90
|
+
`narrowDecision`, `thinGrounding` are observational only.
|
|
91
|
+
|
|
92
|
+
## Pins
|
|
93
|
+
|
|
94
|
+
- `test/01-floor` — floor geometry.
|
|
95
|
+
- `test/04-think` — decider, admissible pruning, investment discipline.
|
|
@@ -0,0 +1,96 @@
|
|
|
1
|
+
# Memoization — Shared Evidence Without Duplication
|
|
2
|
+
|
|
3
|
+
> **Law:** asking never writes, so structural reads are pure during one
|
|
4
|
+
> response. Memoization elides probes, not evidence.
|
|
5
|
+
|
|
6
|
+
Two layers: `Precomputed` (response-scoped shared analyses) and `Mind`
|
|
7
|
+
per-response memos. Both are accelerators that must not change what inference
|
|
8
|
+
computes.
|
|
9
|
+
|
|
10
|
+
## Precomputed — one response, one container
|
|
11
|
+
|
|
12
|
+
`Precomputed` (`src/mind/pipeline-mechanism.ts`) is the sole place a response's
|
|
13
|
+
shared evidence lives. Created by `think` (`src/mind/pipeline.ts`) before the
|
|
14
|
+
mechanism loop.
|
|
15
|
+
|
|
16
|
+
### Eager — populated before any `floor`/`run`
|
|
17
|
+
|
|
18
|
+
- `rec: Recognition` — structural + canonical decomposition (`recognise`)
|
|
19
|
+
- `computed: ComputedSpan[]` — `parse()` results from all mechanisms (e.g. ALU)
|
|
20
|
+
- `guide: Vec` — query gist, the response-wide disambiguation guide
|
|
21
|
+
- `k: number` — `cfg.recallQueryK * 2`, the breadth for resonance/weave/climb
|
|
22
|
+
|
|
23
|
+
### Lazy — computed on first touch, cached by promise
|
|
24
|
+
|
|
25
|
+
Expensive analyses are `async` and cached by promise: the first caller starts
|
|
26
|
+
the work, every later caller awaits the same promise.
|
|
27
|
+
|
|
28
|
+
- `attention()` — `climbAttentionAll` (roots + ranked anchors)
|
|
29
|
+
- `weave()` — `alignGraded` over top-k anchors
|
|
30
|
+
- `resonance()` — `store.resonate(guide, k)` (single ANN query)
|
|
31
|
+
- `frames()` — `frameSlots` inventory from resonance
|
|
32
|
+
- `spanShapedOf(anchor)` / `spanShapedAll()` — per-anchor `skillExemplar`,
|
|
33
|
+
memoised per id
|
|
34
|
+
- `queryWindows` / `queryResolved` / `windowsOf(anchor)` — W-window identities
|
|
35
|
+
- `reachMemo` — `sharedReachMemo(ctx)` (ancestor reach, § below)
|
|
36
|
+
|
|
37
|
+
A mechanism that never asks pays nothing; two mechanisms asking the same
|
|
38
|
+
question pay once. `floor()` must gate on `worthRunning` before first-touching
|
|
39
|
+
an expensive analysis.
|
|
40
|
+
|
|
41
|
+
## Mind memos — `beginResponse` → `endResponse`
|
|
42
|
+
|
|
43
|
+
`Mind` (`src/mind/mind.ts:beginResponse`/`endResponse`) swaps per-response state
|
|
44
|
+
for each inference call. `respond` takes fresh maps; `respondTurn` reuses the
|
|
45
|
+
conversation's persistent ones (content-keyed, cross-turn).
|
|
46
|
+
|
|
47
|
+
| Memo | Key | Scope |
|
|
48
|
+
| ------------------- | ----------------------------------------- | -------------------------------------------- |
|
|
49
|
+
| `perceiveMemo` | `perceiveKey(bytes)` (latin1) | response / conversation |
|
|
50
|
+
| `recogniseMemo` | `latin1Key(bytes)` | response / conversation |
|
|
51
|
+
| `climbMemo` | `latin1Key(bytes)` | response / conversation |
|
|
52
|
+
| `canonMemo` | `latin1Key(bytes)` | response (when `canon` set) |
|
|
53
|
+
| `_resolvedSubtrees` | `WeakMap<Sema, {id,len}>` (node identity) | response / conversation |
|
|
54
|
+
| `_edgeChoice` | `Map<nodeId, pick>` | response only — **cleared** in `endResponse` |
|
|
55
|
+
| `_gistCache` | `BoundedMap<nodeId, Vec>` 32 MB | **session-lifetime** (not per-response) |
|
|
56
|
+
|
|
57
|
+
`_gistCache` (≈ 8K gists at D=1024) survives across responses; all others are
|
|
58
|
+
dropped or cleared at `endResponse`. `_resolvedSubtrees` elides store probes
|
|
59
|
+
when `visit` is absent; with a visitor it still walks in full (see
|
|
60
|
+
`src/mind/primitives.ts:foldTree`).
|
|
61
|
+
|
|
62
|
+
## Trace boundary — what is bypassed
|
|
63
|
+
|
|
64
|
+
Traced responses must emit every step, but must not change the answer.
|
|
65
|
+
|
|
66
|
+
- **Bypassed:** `_edgeChoice` via `guidedNext`
|
|
67
|
+
(`src/mind/traverse.ts:guidedNext`) and `sharedReachMemo`
|
|
68
|
+
(`src/mind/traverse.ts:sharedReachMemo`). Both return fresh empty maps when
|
|
69
|
+
`ctx.trace !== null`; `chooseNext` recomputes identically (pure over store +
|
|
70
|
+
guide).
|
|
71
|
+
- **Always consulted:** `perceiveMemo`, `recogniseMemo`, `climbMemo` (and their
|
|
72
|
+
underlying `perceive`/`recognise`/`climbAttention` caches). Bypassing breaks
|
|
73
|
+
idempotence.
|
|
74
|
+
|
|
75
|
+
`foldTree`'s subtree fast path is taken only when no `visit` is supplied. With a
|
|
76
|
+
visitor (recognition, attention) the walk still descends; the cache elides only
|
|
77
|
+
probes. Bypassing `recogniseMemo` under trace re-ran `recogniseImpl` with a warm
|
|
78
|
+
`_resolvedSubtrees` and emitted fewer sites (observed 31 → 5) — a correctness
|
|
79
|
+
change, not just a slowdown.
|
|
80
|
+
|
|
81
|
+
## Meter — charge work to itself
|
|
82
|
+
|
|
83
|
+
Shared analyses are charged to their own phase (`meter.time(phase, fn)` in
|
|
84
|
+
`Precomputed.shared`), not to the mechanism that first touched them
|
|
85
|
+
(`src/meter.ts:PhaseCost`). Without this, the profile reads "cast.floor costs 2
|
|
86
|
+
s" when the cost was the consensus climb cast paid for on everyone's behalf.
|
|
87
|
+
|
|
88
|
+
## Adding a shared analysis
|
|
89
|
+
|
|
90
|
+
Add one lazy method to `Precomputed`. No new memo map elsewhere. Gate it behind
|
|
91
|
+
`worthRunning` in `floor()`.
|
|
92
|
+
|
|
93
|
+
## Pins
|
|
94
|
+
|
|
95
|
+
- `test/42` — recognition idempotence under trace: traced and untraced
|
|
96
|
+
`recognise` return the same site count and cached object.
|
|
@@ -0,0 +1,55 @@
|
|
|
1
|
+
# Meter — Work Accounting
|
|
2
|
+
|
|
3
|
+
`src/meter.ts` is the one computational-usage accounting surface. It counts what
|
|
4
|
+
inference _cost_ so a slow response can be attributed instead of guessed at. The
|
|
5
|
+
rationale says why an answer was chosen; the meter says what it cost to choose
|
|
6
|
+
it. Harness: `bench/profile-inference.mjs`.
|
|
7
|
+
|
|
8
|
+
## Five contracts
|
|
9
|
+
|
|
10
|
+
1. **Write-only from inference.** No counter reaches a decision, a threshold, or
|
|
11
|
+
an ordering. Determinism survives only because the meter is observed, never
|
|
12
|
+
consulted. Every call site is `meter?.x++` on a nullable field.
|
|
13
|
+
|
|
14
|
+
2. **Counters vs hints.** Counters are deterministic and diffable between runs;
|
|
15
|
+
the same query on the same store meters identically, so a regression is
|
|
16
|
+
visible in a diff. Millisecond fields (`elapsedMs`, per-phase `ms`) are
|
|
17
|
+
non-deterministic hints reported separately — never use them to gate
|
|
18
|
+
behaviour.
|
|
19
|
+
|
|
20
|
+
3. **Phases nest, they do not partition.** `think` contains every mechanism
|
|
21
|
+
phase; a mechanism's `floor` contains whatever shared analysis it
|
|
22
|
+
first-touched; `recall.run` contains `substitutionBridge`. Read a phase as
|
|
23
|
+
inclusive wall-clock — never sum phases and expect the total.
|
|
24
|
+
`CostReport.elapsedMs` is the only whole.
|
|
25
|
+
|
|
26
|
+
4. **Count once.** Off by default and free when off
|
|
27
|
+
(`new Mind({ profile:
|
|
28
|
+
true })` to attach). A layer that wants to be
|
|
29
|
+
visible bumps a field in `meter.ts` — it does not grow a private counter.
|
|
30
|
+
(Legacy `danglingReads` / `compactFailures` in `store.ts` are health
|
|
31
|
+
counters, not per-response work.)
|
|
32
|
+
|
|
33
|
+
5. **Shared analyses charged to themselves.** The first toucher pays the wall
|
|
34
|
+
clock, but every later consumer gets the result free. Attribution follows the
|
|
35
|
+
analysis, not the mechanism that first triggered it — otherwise the profile
|
|
36
|
+
misreads which work is expensive (e.g. the consensus climb billed through
|
|
37
|
+
whichever mechanism happened to need it first).
|
|
38
|
+
|
|
39
|
+
## Where counters live
|
|
40
|
+
|
|
41
|
+
`src/meter.ts:Meter` is the only definition of a counter name. Phases are
|
|
42
|
+
charged via `meter.time(phase, fn)` / `meter.timeSync(phase, fn)`, which
|
|
43
|
+
snapshot counters on entry and attribute the delta to the phase. The sync/async
|
|
44
|
+
seam is load-bearing: synchronous layers (perception, recognition, graph search)
|
|
45
|
+
must use `timeSync` so the profiled path does not await where the unprofiled
|
|
46
|
+
path does not.
|
|
47
|
+
|
|
48
|
+
`CostReport` is plain JSON (`version`, `elapsedMs`, `queryBytes`, `counters`,
|
|
49
|
+
`phases`). Zero-valued counters are dropped; `formatReport` renders the three
|
|
50
|
+
heaviest counters per phase.
|
|
51
|
+
|
|
52
|
+
## Pins
|
|
53
|
+
|
|
54
|
+
- `test/55` — `Meter`, `CostReport`, `searchPops` / `searchPushes`, phase
|
|
55
|
+
nesting.
|