@hviana/sema 0.7.3 → 0.7.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (42) hide show
  1. package/AGENTS.md +95 -843
  2. package/README.md +11 -11
  3. package/dist/src/mind/mind.js +14 -0
  4. package/dist/src/store-sqlite.js +17 -0
  5. package/dist/src/store.d.ts +18 -0
  6. package/dist/src/store.js +10 -0
  7. package/docs/INDEX.md +71 -0
  8. package/docs/INVARIANTS.md +19 -0
  9. package/docs/architecture/bounded-reads.md +85 -0
  10. package/docs/architecture/caches.md +89 -0
  11. package/docs/architecture/commonality.md +45 -0
  12. package/docs/architecture/cost-model.md +71 -0
  13. package/docs/architecture/determinism.md +73 -0
  14. package/docs/architecture/exact-vs-approximate.md +47 -0
  15. package/docs/architecture/factored-machinery.md +28 -0
  16. package/docs/architecture/fold-contract.md +87 -0
  17. package/docs/architecture/halo-sketch.md +99 -0
  18. package/docs/architecture/match-project.md +62 -0
  19. package/docs/architecture/mechanism-market.md +95 -0
  20. package/docs/architecture/memoization.md +96 -0
  21. package/docs/architecture/meter.md +55 -0
  22. package/docs/architecture/saturation.md +92 -0
  23. package/docs/architecture/store.md +79 -0
  24. package/docs/architecture/thresholds.md +79 -0
  25. package/docs/failures/tempting-but-wrong.md +144 -0
  26. package/docs/harness/gates.md +56 -0
  27. package/docs/mechanisms/alu.md +75 -0
  28. package/docs/mechanisms/cast.md +75 -0
  29. package/docs/mechanisms/confluence.md +36 -0
  30. package/docs/mechanisms/cover.md +54 -0
  31. package/docs/mechanisms/extraction.md +53 -0
  32. package/docs/mechanisms/prefix-completion.md +54 -0
  33. package/docs/mechanisms/recall.md +69 -0
  34. package/docs/mechanisms/reference.md +58 -0
  35. package/jsr.json +1 -1
  36. package/package.json +1 -1
  37. package/src/mind/mind.ts +14 -0
  38. package/src/store-sqlite.ts +19 -0
  39. package/src/store.ts +22 -0
  40. package/test/89-completion-recursion.test.mjs +30 -10
  41. package/test/97-store-seed.test.mjs +105 -0
  42. package/HOW_IT_WORKS.md +0 -5836
package/README.md CHANGED
@@ -41,7 +41,7 @@ No weights. No gradients. No training loop. No neural network. No GPU.
41
41
  > Vector Symbolic Architecture (Plate 1995; Kanerva 2009) over a
42
42
  > content-addressable memory, with inference by weighted automated deduction
43
43
  > (Knuth 1977; Felzenszwalb & McAllester 2007). Each term is grounded in
44
- > [HOW_IT_WORKS.md](HOW_IT_WORKS.md).
44
+ > [docs/INDEX.md](docs/INDEX.md).
45
45
 
46
46
  ---
47
47
 
@@ -314,16 +314,16 @@ start talking — no install, no runtime, no API key.
314
314
 
315
315
  ## ✦ Learn more
316
316
 
317
- | Document | What's inside |
318
- | :----------------------------------------------------------------------------------- | :------------------------------------------------------------------------------------------------------------------------------------------------------- |
319
- | 📘 **[HOW_IT_WORKS.md](HOW_IT_WORKS.md)** | The full theory: vector symbolic architectures, the Merkle DAG, distributional halos, weighted deduction — concepts, diagrams, and extensive pseudocode. |
320
- | 🛠️ **[AGENTS.md](AGENTS.md)** | The development manual: repo layout, build/test, internals, invariants, and recipes for extending the system. |
321
- | 🎓 **[CITATION.cff](CITATION.cff)** | How to cite Sema in academic work. |
322
- | ⚖️ **[LICENSE.md](LICENSE.md)** | PolyForm Noncommercial License 1.0.0. |
323
- | 📚 **[DATASETS.md](DATASETS.md)** | Training corpora: provenance, per-corpus attribution, and how a trained memory file is licensed. |
324
- | 💼 **[COMMERCIAL-LICENSE.md](COMMERCIAL-LICENSE.md)** | Commercial licensing terms and contact. |
325
- | 🤗 **[Trained examples](https://huggingface.co/buckets/hviana/sema-trained-v1)** | Pre-trained memory files you can download and use directly. |
326
- | 💿 **[Binary examples](https://huggingface.co/buckets/hviana/sema-binary-examples)** | Ready-to-run web chat apps for Windows, Mac, and Linux — one file, no install. |
317
+ | Document | What's inside |
318
+ | :----------------------------------------------------------------------------------- | :----------------------------------------------------------------------------------------------------------------------------------------------------- |
319
+ | 📘 **[docs/INDEX.md](docs/INDEX.md)** | Architecture docs: single-system laws (docs/architecture/), mechanisms, invariants, and harness gates — the entry point for theory and implementation. |
320
+ | 🛠️ **[AGENTS.md](AGENTS.md)** | The development manual: repo layout, build/test, internals, invariants, and recipes for extending the system. |
321
+ | 🎓 **[CITATION.cff](CITATION.cff)** | How to cite Sema in academic work. |
322
+ | ⚖️ **[LICENSE.md](LICENSE.md)** | PolyForm Noncommercial License 1.0.0. |
323
+ | 📚 **[DATASETS.md](DATASETS.md)** | Training corpora: provenance, per-corpus attribution, and how a trained memory file is licensed. |
324
+ | 💼 **[COMMERCIAL-LICENSE.md](COMMERCIAL-LICENSE.md)** | Commercial licensing terms and contact. |
325
+ | 🤗 **[Trained examples](https://huggingface.co/buckets/hviana/sema-trained-v1)** | Pre-trained memory files you can download and use directly. |
326
+ | 💿 **[Binary examples](https://huggingface.co/buckets/hviana/sema-binary-examples)** | Ready-to-run web chat apps for Windows, Mac, and Linux — one file, no install. |
327
327
 
328
328
  ---
329
329
 
@@ -143,10 +143,24 @@ export class Mind {
143
143
  const { store: optsStore, mechanisms: userMechs, mechanismFactories: userFacts, canon: optsCanon, profile: optsProfile, ...rest } = (optsOrCfg ?? {});
144
144
  this._canonOpt = optsCanon ?? null;
145
145
  this._profile = optsProfile === true;
146
+ // `explicitSeed` is read BEFORE resolveConfig folds the default in, so
147
+ // the store can be consulted only when the caller did not choose.
148
+ const explicitSeed = rest.seed;
146
149
  this.cfg = resolveConfig(rest);
147
150
  this.store = optsStore ?? new SQliteStore({
148
151
  maxGroup: this.cfg.geometry.maxGroup,
149
152
  });
153
+ // THE ARTIFACT'S SEED GOVERNS. `train.seed` is recovered by the store at
154
+ // open, exactly like `train.D` and `geometry.maxGroup`. The seed feeds
155
+ // `makeKeyring`, `Space.rand` and the `Alphabet` below, so folding a
156
+ // query under config.ts's default (42) against a store trained with
157
+ // another seed (e.g. 7) lands in a DIFFERENT vector space than the one
158
+ // the artifact's nodes were folded into: recognition and resonance then
159
+ // read the wrong space and every answer degrades silently. An explicit
160
+ // caller seed still wins — this only replaces the unconfigured default.
161
+ if (explicitSeed === undefined && this.store.trainSeed !== null) {
162
+ this.cfg.seed = this.store.trainSeed;
163
+ }
150
164
  userMechanisms = userMechs ?? [];
151
165
  userFactories = userFacts ?? [];
152
166
  }
@@ -360,6 +360,23 @@ export class SQliteStore extends AbstractStore {
360
360
  this._maxGroup = g;
361
361
  }
362
362
  }
363
+ // Recover the TRAINING seed exactly as D and maxGroup are recovered. The
364
+ // seed seeds the alphabet and the seat keyring (mind.ts), so a Mind that
365
+ // folds a query under any other seed lands in a different vector space
366
+ // than the one this artifact's nodes were folded into — recognition,
367
+ // resonance and every mechanism downstream then read the wrong space. The
368
+ // trainer persists `train.seed` and refuses to resume against a store
369
+ // trained with a different one (example/train_base/main.ts), so the value
370
+ // is authoritative for this artifact. Absent on a store that was never
371
+ // trained, where the caller's configured seed stands.
372
+ {
373
+ const row = this.sqlite.prepare("SELECT val FROM meta WHERE key = 'train.seed'").get();
374
+ if (row) {
375
+ const s = Number(row.val);
376
+ if (Number.isInteger(s) && s >= 0)
377
+ this._trainSeed = s;
378
+ }
379
+ }
363
380
  // Persist maxGroup to meta when opening a FRESH store (no rows yet) so
364
381
  // indexSubtree always sees the training-time value even when the store is
365
382
  // accessed without a Mind / full snapshot.
@@ -114,6 +114,16 @@ export declare class BoundedMap<K, V> {
114
114
  }
115
115
  export interface Store {
116
116
  readonly D: number;
117
+ /** The seed the artifact was TRAINED with, recovered from the store's own
118
+ * `train.seed` metadata, or null for a store that was never trained.
119
+ *
120
+ * This is not decoration: the seed feeds the alphabet and the seat keyring
121
+ * (see the Mind constructor), so folding a query under any other seed lands
122
+ * in a different vector space than the one the artifact's nodes were folded
123
+ * into. A Mind opening a trained store MUST adopt this seed unless the
124
+ * caller explicitly overrides it — the same discipline that recovers
125
+ * `train.D` and `geometry.maxGroup` from the metadata. */
126
+ readonly trainSeed: number | null;
117
127
  /** The work accumulator for the inference call in flight, or null. The
118
128
  * Mind attaches one per profiled response and detaches it after (see
119
129
  * src/meter.ts). A store MUST only ever write to it — no read may reach
@@ -475,6 +485,10 @@ export declare abstract class AbstractStore implements Store {
475
485
  protected efFor(clusterCount: number): number;
476
486
  protected _D: number;
477
487
  protected _maxGroup: number;
488
+ /** `train.seed` recovered by the backend at open, or null when the store was
489
+ * never trained. A backend that omits it simply reports null, which leaves
490
+ * the caller's configured seed in force. */
491
+ protected _trainSeed: number | null;
478
492
  protected readonly minHaloMass: number;
479
493
  protected readonly efSearch: number;
480
494
  protected readonly overfetch: number;
@@ -563,6 +577,10 @@ export declare abstract class AbstractStore implements Store {
563
577
  protected _edgeSrcCount: number;
564
578
  constructor(config: StoreConfig, D: number, maxGroup: number);
565
579
  get D(): number;
580
+ /** The seed the artifact was trained with, recovered from `train.seed` at
581
+ * open. Null for a store that was never trained. See
582
+ * {@link Store.trainSeed} for why this must govern inference. */
583
+ get trainSeed(): number | null;
566
584
  /** Await the async initialisation performed by the concrete constructor. */
567
585
  protected _ensureReady(): Promise<void>;
568
586
  has(id: NodeId): boolean;
package/dist/src/store.js CHANGED
@@ -434,6 +434,10 @@ export class AbstractStore {
434
434
  // ── Config ─────────────────────────────────────────────────────────────
435
435
  _D;
436
436
  _maxGroup;
437
+ /** `train.seed` recovered by the backend at open, or null when the store was
438
+ * never trained. A backend that omits it simply reports null, which leaves
439
+ * the caller's configured seed in force. */
440
+ _trainSeed = null;
437
441
  minHaloMass;
438
442
  efSearch;
439
443
  overfetch;
@@ -552,6 +556,12 @@ export class AbstractStore {
552
556
  get D() {
553
557
  return this._D;
554
558
  }
559
+ /** The seed the artifact was trained with, recovered from `train.seed` at
560
+ * open. Null for a store that was never trained. See
561
+ * {@link Store.trainSeed} for why this must govern inference. */
562
+ get trainSeed() {
563
+ return this._trainSeed;
564
+ }
555
565
  /** Await the async initialisation performed by the concrete constructor. */
556
566
  async _ensureReady() {
557
567
  if (!this._ready)
package/docs/INDEX.md ADDED
@@ -0,0 +1,71 @@
1
+ # Sema Documentation Index
2
+
3
+ Sema is a single system stated three ways: the law lives in `docs/architecture/`
4
+ (what holds), the prescription in `AGENTS.md` bootloader (what to do and where),
5
+ and the proof in `test/` (pins that fail when the law is broken).
6
+
7
+ ## Routing — what to read for each task
8
+
9
+ | Task | Read | Why |
10
+ | --------------------------- | ---------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------- |
11
+ | Add a mechanism | `docs/architecture/mechanism-market.md` + `docs/mechanisms/*.md` | Market contract: decoupled, declared competence, visible budget, evidence travels |
12
+ | Add a threshold | `docs/architecture/thresholds.md` | All cutoffs are formulas over D/W/N in `geometry.ts`; `config.ts` holds only budgets |
13
+ | Debug an answer | `docs/architecture/cost-model.md` + `src/meter.ts` | One cost ladder (`MICRO`/`STEP`/`CONCEPT`/`PASS`) decides every grounding choice |
14
+ | Understand the fold | `docs/architecture/fold-contract.md` | Deposit and inference must compute the same tree; boundaries are not turn metadata |
15
+ | Add a store backend | `docs/architecture/store.md` + `docs/architecture/bounded-reads.md` | `AbstractStore` owns domain logic; backends are thin wrappers with capped reads |
16
+ | Add an ALU operation | `src/alu/README.md` | One `registry.derive` per op composing existing ops; no new `derive` needed |
17
+ | Add a matcher or projection | `docs/architecture/match-project.md` | Mechanisms are `(matcher, direction, gate)` configs over the shared `match.ts` family |
18
+ | Add a deduction rule | `docs/architecture/cost-model.md` + `docs/architecture/determinism.md` | Place cost on the ladder, keep heuristic admissible, extend `classifyMove` |
19
+ | Change vector search | `docs/architecture/exact-vs-approximate.md` + `docs/architecture/bounded-reads.md` | Scores propose, bytes dispose; ANN is bounded by `hubBound` |
20
+ | Profile or bound work | `docs/architecture/meter.md` + `docs/architecture/bounded-reads.md` | `meter.ts` is write-only; counters are product, phases are hints |
21
+
22
+ ## Architecture laws (13)
23
+
24
+ | Law | File | Summary | Pins |
25
+ | --- | ------------------------------------------- | ----------------------------------------------------------------------------------------------------------- | -------------------- |
26
+ | 1 | `docs/architecture/determinism.md` | No `Math.random`/`Date.now` in behaviour; seed-derived randomness; corpus-determined tie-breaks | `test/20` |
27
+ | 2 | `docs/architecture/thresholds.md` | Every decision cutoff derived in `geometry.ts` over D/W/N; no tunable knobs | `test/40`, `test/64` |
28
+ | 3 | `docs/architecture/exact-vs-approximate.md` | Vector scores rank only; identity via content-addressed lookup; five graded ladders | `test/51` |
29
+ | 4 | `docs/architecture/cost-model.md` | Single ladder `MICRO`/`STEP`/`CONCEPT`/`PASS`; weight `moves + PASS·unaccounted`; `STEP`-grade compare | `test/04`, `test/55` |
30
+ | 5 | `docs/architecture/match-project.md` | Shared `match.ts` family (`locate`/`alignGraded`/`frameSlots`/`project`); voicing gates belong to consumers | `test/24`, `test/76` |
31
+ | 6 | `docs/architecture/mechanism-market.md` | `PipelineMechanism` (`floor`/`run`/`parse`); admissible-floor pruning and investment discipline | `test/01`, `test/04` |
32
+ | 7 | `docs/architecture/commonality.md` | Two populations: corpus-global (`reachOf`+`dominates`) vs weave-local (`depth[]`) | `test/17`, `test/34` |
33
+ | 8 | `docs/architecture/bounded-reads.md` | No per-query read grows with N; `hubBound=√N` enforced at store via LIMIT/probe/prefix caps | `test/77`, `test/90` |
34
+ | 9 | `docs/architecture/store.md` | `AbstractStore` owns dedup/indexing/batch; `store-sqlite.ts` is thin wrappers; canon index optional | `test/08` |
35
+ | 10 | `docs/architecture/fold-contract.md` | `perceiveDeposit` and `perceive` agree; `contentLevels` is single boundary rule; no W/offset dependence | `test/59`, `test/63` |
36
+ | 11 | `docs/architecture/memoization.md` | `Precomputed` is per-response lazy cache (promise-cached async); `beginResponse`/`endResponse` lifecycle | `test/42` |
37
+ | 12 | `docs/architecture/saturation.md` | Every walk names a deciding saturation beside its cap; cap is safety net, not decision | `test/27`, `test/16` |
38
+ | 13 | `docs/architecture/meter.md` | `meter.ts` is write-only work accounting; counts are deterministic, phases nest | `test/55` |
39
+
40
+ ## Mechanisms (8)
41
+
42
+ | Mechanism | File | Role |
43
+ | ----------------- | -------------------------------------- | --------------------------------------------------------- |
44
+ | cover | `docs/mechanisms/cover.md` | Exact/computed-span covering via `GraphSearch` |
45
+ | cast | `docs/mechanisms/cast.md` | Weave-local analogy via `depth[]` frame gate |
46
+ | confluence | `docs/mechanisms/confluence.md` | Corpus-global filler/scaffolding gate over climb |
47
+ | extraction | `docs/mechanisms/extraction.md` | Located-frame read-out with anchored span accounting |
48
+ | reference | `docs/mechanisms/reference.md` | Slot-bound voicing of asker-supplied referents |
49
+ | recall | `docs/mechanisms/recall.md` | Nearest stored form; echo tier via substitution bridge |
50
+ | prefix-completion | `docs/mechanisms/prefix-completion.md` | Literal prefix of exactly one trained form |
51
+ | alu | `docs/mechanisms/alu.md` | Authoritative computed spans (`parse` → `aluToMechanism`) |
52
+
53
+ ## Supporting docs
54
+
55
+ | Doc | Role |
56
+ | ----------------------------------------- | ------------------------------------------------------------------------------- |
57
+ | `docs/architecture/caches.md` | Every acceleration is a `BoundedMap`; miss re-derives; budgets in `StoreConfig` |
58
+ | `docs/architecture/halo-sketch.md` | Halo & sketch — distributional memory, quantization, bottom-k profiles |
59
+ | `docs/architecture/factored-machinery.md` | Single-definition contracts table — one owner per shared symbol |
60
+
61
+ ## Cross-cutting
62
+
63
+ - `docs/INVARIANTS.md` — the five invariants (determinism, derived thresholds,
64
+ exact-decides, one cost currency, bounded reads) with file-level routing.
65
+ - `docs/failures/tempting-but-wrong.md` — refuted simplifications that passed
66
+ review but failed pins (e.g. reordering ladders, flattening attention
67
+ asymmetries, one-cone-exhausted stop).
68
+ - `docs/harness/gates.md` — how `AGENTS.md` recipes,
69
+ `bench/profile-inference.mjs`, and `test/*.test.mjs` enforce the laws.
70
+ - `docs/architecture/` — full per-law derivation (each file states why the law
71
+ is this way; no separate HOW).
@@ -0,0 +1,19 @@
1
+ # INVARIANTS — Laws, Proofs, Derivations
2
+
3
+ > Law in `docs/architecture/*.md`, proof in `test/*.test.mjs`.
4
+
5
+ | # | Law | Where defined (src symbol) | Pins (test/N) | Doc (docs/architecture/*.md) |
6
+ | -- | ------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------- | ------------------- | ---------------------------- |
7
+ | 1 | Determinism | `src/config.ts:seed` `src/alphabet.ts:Alphabet` `src/mind/traverse.ts:guidedFirst` | `test/20` `test/42` | `determinism.md` |
8
+ | 2 | Derived thresholds | `src/geometry.ts:mergeThreshold,identityBar,reachThreshold,significanceBar,consensusFloor,dominates` | `test/64` `test/40` | `thresholds.md` |
9
+ | 3 | Exact decides / approximate proposes | `src/mind/primitives.ts:resolve` `src/mind/match.ts:locate,alignGraded` `src/mind/resonance.ts:bridge` | `test/51` | `exact-vs-approximate.md` |
10
+ | 4 | One cost currency | `src/mind/graph-search.ts:MICRO,STEP,CONCEPT,PASS` `src/derive:lightestDerivation` (min,+) `src/mind/attention.ts:poolVotes` (+,+) | `test/55` `test/04` | `cost-model.md` |
11
+ | 5 | Bounded reads | `src/store.ts:AbstractStore:nextFirst,parentsFirst,containersSlice,hasNext,bytesPrefix` `src/mind/traverse.ts:hubBound,hubCap` | `test/90` `test/14` | `bounded-reads.md` |
12
+ | 6 | Fold contract | `src/geometry.ts:contentLevels` `src/mind/canonical.ts:canonicalWindows,chainReach` `src/canon.ts:canonicalizer` | `test/59` `test/63` | `fold-contract.md` |
13
+ | 7 | Mechanism market | `src/mind/pipeline-mechanism.ts:PipelineMechanism,Precomputed` `src/mind/pipeline.ts:think,worthRunning` | `test/01` `test/04` | `mechanism-market.md` |
14
+ | 8 | Two commonality measures | `src/mind/traverse.ts:reachOf,dominates,corpusN` (global) `src/mind/match.ts:depth[],MIN_WEAVE` (weave-local) | `test/17` `test/34` | `commonality.md` |
15
+ | 9 | Memoization idempotence | `src/mind/pipeline-mechanism.ts:Precomputed` `src/mind/mind.ts:beginResponse,endResponse,_resolvedSubtrees` | `test/42` | `memoization.md` |
16
+ | 10 | Caches as budgets | `src/store.ts:BoundedMap` `src/config.ts:StoreConfig:bytesCacheMax,recCacheBytes,haloCacheBytes` | `test/96` `test/91` | `caches.md` |
17
+ | 11 | Honest degradation | `src/mind/pipeline.ts:weight=moves+PASS*unaccounted` `src/store.ts:BoundedMap:miss→re-derive` | `test/28` `test/84` | `store.md`+`caches.md` |
18
+ | 12 | Meter contracts | `src/meter.ts:Meter,PhaseCost,time` `src/mind/pipeline-mechanism.ts:Precomputed.shared` | `test/55` | `meter.md` |
19
+ | 13 | Saturation | `src/mind/traverse.ts:edgeAncestors:SaturationReason` `src/mind/junction.ts:junctionContainersFrom` `src/mind/resonance.ts:pivotInto` | `test/27` `test/16` | `saturation.md` |
@@ -0,0 +1,85 @@
1
+ # Bounded Reads — No Per-Query Read Grows With the Corpus
2
+
3
+ > **Law:** the cost of one query is proportional to the query, not to how much
4
+ > was learned. No per-query read may grow with corpus size N.
5
+
6
+ Every fan-out, walk, and disambiguation is capped at `hubBound` —
7
+ `ceil(sqrt(N))` — derived once from `corpusN` and floored at 2 so `sqrt` and
8
+ `ln` stay meaningful on a near-empty store. There is no second convention; do
9
+ not invent one.
10
+
11
+ ## Scale
12
+
13
+ ```
14
+ corpusN(ctx) = max(2, store.edgeSourceCount()) // distinct learnt contexts
15
+ hubBound(ctx) = ceil(sqrt(corpusN(ctx))) // >= 2, the store cap
16
+ hubCap(ctx, ids) = ids.slice(0, hubBound(ctx)) // list-side reading
17
+ boundFor(n) = ceil(sqrt(max(2, n))) // ctx-free reading
18
+ ```
19
+
20
+ Defined once in `mind/traverse.ts` (`corpusN`, `hubBound`, `hubCap`,
21
+ `boundFor`). Every consumer imports them; never spell `Math.sqrt` inline.
22
+
23
+ ## Enforcement at the store level
24
+
25
+ The cap is not advisory — adapters must make bounded reads bounded in SQL.
26
+
27
+ ### 1. LIMITed reads — real `LIMIT ?`
28
+
29
+ `nextFirst(id, limit)`, `prevFirst(id, limit)`, `parentsFirst(id, limit)`,
30
+ `containersSlice(child, offset, limit)`.
31
+
32
+ Same statement and `ORDER BY` as the full read, with `LIMIT ?`. Never
33
+ "materialise then slice". Reading `hubBound + 1` parents decides "hub or not"
34
+ exactly without reading the rest. Implemented as thin wrappers in
35
+ `store-sqlite.ts` over `AbstractStore` in `store.ts`.
36
+
37
+ ### 2. Existence probes — indexed point probes
38
+
39
+ `hasNext(id)`, `hasParents(id)`, `hasContainers(child)`, `hasHalo(id)`,
40
+ `prevCount(id)`.
41
+
42
+ One indexed `EXISTS` / `COUNT` probe that never decodes vectors or unpacks
43
+ blobs. Use them for every "does this lead anywhere?" question instead of
44
+ `next(id).length > 0` or `prev(id).length`. `prevCount` is the reverse-edge
45
+ support count for `chooseNext`/`chooseAmong`; `hasNext`/`hasHalo` gate the
46
+ `leadsSomewhere` admission predicate in `mind/traverse.ts`.
47
+
48
+ ### 3. Prefix-capped reads — reject without reconstruction
49
+
50
+ `bytesPrefix(id, cap)` and `contentLen(id, cap)`.
51
+
52
+ `contentLen` under a cap returns an exact length below the cap and `>= cap`
53
+ otherwise — an indexed memo walk that stops early, never a full subtree walk.
54
+ `bytesPrefix` stops after `cap` bytes. A candidate exceeding the cap is rejected
55
+ on the length probe alone; the weave, junction walks, and bridge all read this
56
+ way. Uncapped reads there cost seconds per query on a large store.
57
+
58
+ ### 4. Transparent scaffolding — one bounded read
59
+
60
+ `chainRun(id)` climbs a run of transparent nodes (exactly one parent, no edges)
61
+ in a single recursive CTE, cached for the store lifetime and dropped on writes
62
+ that break transparency. The climber hops the whole run where a node-at-a-time
63
+ ascent would pay three probes per node.
64
+
65
+ ## Maintenance only
66
+
67
+ The full materialising reads — `next(id)`, `prev(id)`, `parents(id)`,
68
+ `containers(child)` — exist for inspection, repair, and compaction only
69
+ (`compactContentIndex`, `repairContentIndex`). Keep them off hot paths.
70
+
71
+ ## Adding a walk
72
+
73
+ Any new fan-out walk uses `hubBound`/`hubCap`. Do not call `edgeSourceCount()`
74
+ or `Math.ceil(Math.sqrt(...))` inline, and do not invent a per-walk limit. The
75
+ walk's saturation decision (when to stop) is separate from the cap (the safety
76
+ net); a walk with only a cap drifts to the cap.
77
+
78
+ ## Pins
79
+
80
+ - `test/14` — sublinear inference in corpus size and constant-rate in input
81
+ length; training throughput floor; exact recall at scale.
82
+ - `test/89` — completion recursion stays output-sensitive (nested searches/pops
83
+ sublinear); guards the count of reads, not just per-read size.
84
+ - `test/90` — connector probe (`resolveConnectors`) reads by the query length
85
+ (`QUERY.length + 1`), not by the learnt continuation; per-read size bound.
@@ -0,0 +1,89 @@
1
+ # Caches — Every Acceleration Is a BoundedMap
2
+
3
+ > **Law:** every acceleration is a `BoundedMap` with a byte budget. A miss
4
+ > re-derives from durable state. Degradation order is speed/reach lost, never
5
+ > identity.
6
+
7
+ No cache may change what is stored, what is resolved, or what tree is folded.
8
+ Eviction costs a re-read, a re-walk, or a narrower reach — never a wrong answer
9
+ or a wrong tree.
10
+
11
+ ## `BoundedMap` — the one cache primitive
12
+
13
+ `src/store.ts:BoundedMap<K,V>` — LRU with byte accounting (`maxBytes`, `sizeOf`,
14
+ `evict`, `recency`).
15
+
16
+ - `evict: "lru"` — uniform-cost entries (dedup, vectors, records).
17
+ - `evict: "smallest"` — variable-cost reconstruction (`_bytesCache`): protects
18
+ expensive large branches over cheap leaves.
19
+ - `recency: "reorder"` (default) — exact LRU via `delete+set`; required when
20
+ eviction choice is load-bearing (`_depositTrees` — 8 entries, victim changes
21
+ fold).
22
+ - `recency: "clock"` — bit instead of reorder; only for transparent caches where
23
+ wrong victim costs a re-read (`_bytesCache`, `_recCache`). Measured: same
24
+ entries cached, hot-path time 55% → bit.
25
+
26
+ Persistent cursor over V8 insertion order makes eviction amortised O(1);
27
+ candidate window for `"smallest"` never rescans from front.
28
+
29
+ ## Store caches — budgets in `src/config.ts:StoreConfig`
30
+
31
+ | Cache | Field | Budget | `sizeOf` | Eviction |
32
+ | ------------------- | -------------------------- | ---------------------------- | ------------------- | -------------- |
33
+ | dedup leaf/branch | `_leafKey` / `_branchKey` | `dedupCacheMax` 1M entries | 1 | lru |
34
+ | reconstructed bytes | `_bytesCache` | `bytesCacheMax` 20 MB | `byteLength` | smallest+clock |
35
+ | content length | `_lenCache` | `bytesCacheMax` | 16 | lru |
36
+ | node records | `_recCache` | `recCacheBytes` 10 MB | leaf+4·kids+12 | lru+clock |
37
+ | pending gists | `_pendingGist` | `pendingGistBytes` 16 MB | `byteLength` (D·4) | lru |
38
+ | halo exact / norm | `_haloExact` / `_haloNorm` | `haloCacheBytes` 16 MB each | `byteLength` | lru |
39
+ | skipped interiors | `_coveredIds` | `coveredIdsMax` 100K entries | 1 | lru |
40
+ | indexed ids | `_indexedIds` | `coveredIdsMax` | 1 | lru |
41
+ | transparent chains | `_chainMemo` | `chainCacheBytes` 16 MB | 4·len+32 | lru |
42
+ | ingest memo | `CachedIngest._memo` | `ingestCacheBytes` 50 MB | vector+ids+keyBytes | lru |
43
+
44
+ ANN read caches (`_resonateCache`, `_resonateHaloCache`) are `Map<string,Hit[]>`
45
+ keyed by `vecKey(v)+":"+k`, dropped on any index mutation.
46
+ `vectorCacheMb`/`sqliteCacheMb` are pure page-cache latency knobs.
47
+
48
+ `_bytesCache` only caches complete reconstructions — `bytesPrefix(id,cap)` with
49
+ `got < cap`; a truncated prefix is never stored. `_chainMemo` is dropped on any
50
+ write that could break transparency; `_pendingGist` eviction falls back to DAG
51
+ climb; halo eviction re-decodes the durable 2-bit row.
52
+
53
+ ## Mind caches — session and per-response
54
+
55
+ | Cache | Location | Budget | Scope / invalidation |
56
+ | ------------------- | ------------------------------------------------ | --------- | ------------------------------------------------------------------ |
57
+ | `_gistCache` | `Mind._gistCache` | 32 MB | session-lifetime, never invalidated (perception pure) |
58
+ | `_depositTrees` | `Mind._depositTrees` | 8 entries | session; `perceiveDeposit` only when `conversational` |
59
+ | `_depositLens` | `Mind._depositLens` | — | byte lengths for prefix probes; cleared with map when >64 |
60
+ | `_internIds` | `Mind._internIds: WeakMap<Sema,number>` | — | Mind lifetime; ids permanent |
61
+ | `_resolvedSubtrees` | `Mind._resolvedSubtrees: WeakMap<Sema,{id,len}>` | — | per-response/conversation; fast path only when `visit===undefined` |
62
+
63
+ `REACH_MEMO_MAX` / `STRUCT_MEMO_MAX` 100K (`src/mind/traverse.ts`) — whole-climb
64
+ and per-node structural probes (`hasNext`/`prevCount`/`hasParents`); cleared on
65
+ write or when cap reached. `reachMemo`/`structCaches` keyed by `_structMemoKey`,
66
+ bypassed under trace.
67
+
68
+ ## Deposit caches — offset-keyed, caller-discharged
69
+
70
+ `_depositTrees`/`_depositLens`/`_internIds`/`_resolvedSubtrees` key **offsets**,
71
+ not bytes, for O(1) reuse. Offsets alone cannot witness byte agreement — caller
72
+ must discharge it.
73
+
74
+ - Correct: conversation append — each turn extends the prior cumulative context
75
+ by its own bytes; longest cached proper prefix hit (`L < bytes.length`) reuses
76
+ `contentFoldIncremental` segments bit-identically.
77
+ - Wrong: mismatched `prev` reused by offset produced wrong tree (336 vs 400
78
+ bytes) — a coincidental prefix length aliased unrelated content.
79
+ - Now: `perceiveDeposit` keys by `latin1Key(bytes.subarray(0,L))` (prefix
80
+ bytes), probes longest cached proper prefix first; `_depositTrees` populated
81
+ only for conversational deposits (budget discipline), otherwise cold path
82
+ always correct.
83
+
84
+ ## Pins
85
+
86
+ - `test/91` — `chainRun` via capped `_prefix`: bounded transparent-chain hop,
87
+ not per-node probes.
88
+ - `test/96` — `_bytesCache` is byte-accounted `BoundedMap` that evicts; miss
89
+ re-derives.
@@ -0,0 +1,45 @@
1
+ # Two Measures of Commonality
2
+
3
+ Sema needs "what is shared" in two different populations. One is corpus-global
4
+ (how widely a structure is reused), the other is weave-local (what a local
5
+ cohort of overlapping forms agrees on). They use different data and different
6
+ formulas and must not be conflated.
7
+
8
+ ## Corpus-global — `reachOf` + `dominates`
9
+
10
+ _Defined in `src/mind/traverse.ts` + `src/geometry.ts`; used by climb,
11
+ containment, IDF pooling._
12
+
13
+ For a node id, `reachOf(id, N)` counts how many learnt contexts contain it
14
+ (ancestor reach via capped graph walks, memoised per response in
15
+ `sharedReachMemo`). `dominates(reach, N)` then asks whether that reach is above
16
+ the corpus-determined majority threshold (derived in `geometry.ts` over `N`).
17
+ Intuition: minority reach discriminates (a filler), majority reach is
18
+ scaffolding. Powers the consensus climb, edge following, and vote pooling.
19
+
20
+ ## Weave-local — `depth[]` + `MIN_WEAVE` + `dominates`
21
+
22
+ _Defined in `src/mind/match.ts` (`depth[]`, `MIN_WEAVE`, `frame`) and gated in
23
+ `src/mind/match.ts:frame`; used by CAST._
24
+
25
+ For an alignment weave, `depth[i]` counts how many aligned structures cover byte
26
+ `i` of the query. `MIN_WEAVE = 2` requires agreement beyond a pair (pair columns
27
+ are ambiguous with insertions/deletions), and `dominates(depth[i], aligned)`
28
+ requires agreement by a majority of the aligned cohort:
29
+
30
+ ```
31
+ frame(i) ⇔ depth[i] > MIN_WEAVE ∧ dominates(depth[i], aligned)
32
+ ```
33
+
34
+ This powers CAST's frame gate: what the local cohort shares vs what
35
+ differentiates one member. It never consults corpus reach.
36
+
37
+ The two measures answer different questions over different populations; CAST's
38
+ frame must not be replaced by a reach check and the climb must not be driven by
39
+ weave depth.
40
+
41
+ ## Pins
42
+
43
+ - `test/17` — weave-local frame / `MIN_WEAVE` / `dominates` vs corpus-global
44
+ reach.
45
+ - `test/34` — containment and reach-driven disambiguation.
@@ -0,0 +1,71 @@
1
+ # Cost Model — One Currency
2
+
3
+ Every mechanism and every byte competes on one cost ladder defined in
4
+ `src/mind/graph-search.ts`. GraphSearch and `pipeline.ts:think` use the same
5
+ units, so a mechanism-level choice and a byte-level choice are the same kind of
6
+ decision: a lightest derivation.
7
+
8
+ ## Ladder (`src/mind/graph-search.ts`)
9
+
10
+ | Cost | Value | Meaning |
11
+ | --------- | ------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
12
+ | `MICRO` | `1e-3` | Recognised advance (one `rec` bridge); per-byte unit of the A\* heuristic. A recomposed form's onward edge is also `MICRO`. |
13
+ | `STEP` | `1` | Every edge hop (first or fifth), every computed result, every projection. Charging every hop makes the lightest derivation the shortest chain. |
14
+ | `CONCEPT` | `10` | Halo-mediated act (synonym hop, consensus climb) and abandoning an edge chain early (`CONCEPT` above chain cost — genuine fixpoint at `+0` always beats giving up at same depth). |
15
+ | `PASS` | `1000` / byte | Carrying a byte nothing explains. Dominates everything so the search always prefers to recognise. |
16
+
17
+ Only the **ordering** `MICRO < STEP < CONCEPT < PASS` matters; any constants
18
+ with that order give the same derivations.
19
+
20
+ ## Pipeline weighing (`src/mind/pipeline.ts:think`)
21
+
22
+ Mechanism candidates are weighed in the same ladder:
23
+
24
+ ```
25
+ weight = moves + PASS * unaccounted_bytes
26
+ grade = floor(weight / STEP)
27
+ ```
28
+
29
+ `unaccounted` is the query bytes no `accounted` span covers. Comparison is at
30
+ `STEP` resolution: lowest `grade` wins. At equal grade the candidate with fewer
31
+ `scaffolding` bytes (answer bytes lifted from unrecognised spans) wins; only
32
+ then does mechanism list order decide.
33
+
34
+ ## Two semirings
35
+
36
+ - **(min, +) tropical** — lightest derivation in `GraphSearch` via `src/derive`
37
+ (`lightestDerivation`). Cost accumulates with `+`, choice selects `min`.
38
+ Powers `cover`/`form`/`out`, edge following, fusing, and the A\* agenda
39
+ (`g + h`).
40
+
41
+ - **(+, +) arithmetic** — evidence pooling in `src/mind/attention.ts:poolVotes`.
42
+ Each region's vote is an axiom; rules carry `Rule.combine = 'sum'` so costs to
43
+ the same anchor **add** rather than minimise. Powers IDF-weighted consensus,
44
+ `votes`/`votesIdf`/`support`, and `regionSupport`/`regionPeak`.
45
+
46
+ ## Admissibility
47
+
48
+ The A\* heuristic is admissible and consistent:
49
+
50
+ ```
51
+ h(it) = (queryLen - right) * MICRO
52
+ ```
53
+
54
+ `right` is `p` for `cover(p)` or `j` for `form/out [i,j)`. `MICRO` is the
55
+ minimum per-byte cost in the ladder (every real per-byte cost is `>= MICRO`,
56
+ including `PASS`), and only the suffix past `right` is counted, so `h` never
57
+ exceeds the true remaining cost.
58
+
59
+ ## Policy is not cost
60
+
61
+ "Computation always wins" is **not** priced into the ladder (a computed result
62
+ costs `STEP`, same as a learned edge). It is enforced by masking: `pipeline.ts`
63
+ removes recognised sites overlapped by a `ComputedResult` so the computation is
64
+ the sole completion there. Keep policy in callers; keep the engine neutral.
65
+
66
+ ## Pins
67
+
68
+ - `test/52` — climb consensus instrumentation
69
+ - `test/53` — cross-region probe instrumentation
70
+ - `test/54` — evidence `k` instrumentation
71
+ - `test/55` — cost meter (`Meter`, `CostReport`, `searchPops`/`searchPushes`)
@@ -0,0 +1,73 @@
1
+ # Determinism — Same Seed + Same Deposits + Same Query ⇒ Same Bytes
2
+
3
+ ## The law
4
+
5
+ > Same `seed` + same deposit order + same query ⇒ byte-identical answer.
6
+
7
+ Determinism is the product. Every code path that can reach output must be
8
+ deterministic given `(seed, store contents, query bytes)`.
9
+
10
+ ## Forbidden
11
+
12
+ No `Math.random` or `Date.now` in behaviour, and no iteration over unordered
13
+ collections where order can reach output. Example-only uses
14
+ (`example/train_base`) are outside the library contract. If a test becomes
15
+ flaky, the contract was broken, not the test.
16
+
17
+ ## All randomness flows from `seed`
18
+
19
+ `MindConfig.seed` (`src/config.ts:resolveConfig`, `DEFAULT_CONFIG`) is the sole
20
+ entropy root. Subsystems derive deterministically:
21
+
22
+ - **Alphabet** — `Alphabet` (`src/alphabet.ts`) via `rng` (`src/vec.ts:rng`)
23
+ seeded as `seed ^ seedMask`; builds 16→64→256 vectors by refinement.
24
+ - **Keyring / Space** — `Space.seats` (`src/sema.ts:Space`) via `makeKeyring`
25
+ (`src/vec.ts:makeKeyring`) and `rng` seeded from `seed` in `Mind`
26
+ (`src/mind/mind.ts`); `fold`/`twoEndedSeat`/`companySignature` are pure over
27
+ `Space`.
28
+ - **Vector indexes** — `VectorDatabase` (`src/rabitq-ivf/src/database.ts`) and
29
+ `Prng` (`src/rabitq-ivf/src/rabitq.ts`) seeded from config; insertion order is
30
+ the stored order, not a random choice.
31
+
32
+ No other PRNG source may affect grounding. Thresholds in `geometry.ts` are
33
+ derived from `D`/`W`/`N`, not sampled.
34
+
35
+ ## Tie-breaks are corpus-determined
36
+
37
+ Every choice among equals bottoms out in a fixed ordering — insertion order or
38
+ lowest node id. The universal no-evidence fallback is **first-inserted**:
39
+
40
+ - `guidedFirst` (`src/mind/traverse.ts:guidedFirst`) — guided pick via
41
+ `chooseNext` else first-inserted edge (`nextFirst` LIMIT 1).
42
+ - `chooseNext` (`src/mind/traverse.ts:chooseNext`) — capped `nextFirst` read
43
+ (`hubBound`), ranked by `prevCount` then `haloMass`; equal ⇒ first-inserted.
44
+ - `chooseAmong` (`src/mind/traverse.ts:chooseAmong`) — `hubCap` + `argmaxCosine`
45
+ over `candidateGist`; first-inserted on tie via stable scan.
46
+ - `companySignature` (`src/sema.ts:companySignature`) — `rng(id ^ 0x9e3779b9)`,
47
+ i.e. seeded by node id, not observation order.
48
+
49
+ Last-inserted was once used in one place; it was a bug. Never reintroduce it.
50
+
51
+ ## Memoization and trace must not break identity
52
+
53
+ Per-response memos (`Precomputed`, `perceiveMemo`, `recogniseMemo`, `climbMemo`,
54
+ `_resolvedSubtrees` via `foldTree`, `_edgeChoice`, `_gistCache` in
55
+ `src/mind/mind.ts` / `src/mind/pipeline-mechanism.ts` /
56
+ `src/mind/primitives.ts`) are sound because asking never writes. Only
57
+ `guidedNext`/`sharedReachMemo` are trace-bypassed;
58
+ `perceiveMemo`/`recogniseMemo`/`climbMemo` are always consulted — `foldTree`'s
59
+ subtree fast path skips `visit` (and thus site emission) for cached subtrees, so
60
+ bypassing makes `recognise` non-idempotent.
61
+
62
+ ## Follow it
63
+
64
+ When you add any choice among equals, name the tie-break explicitly and make it
65
+ corpus-determined. Thread new randomness through `seed`-derived `rng`; never
66
+ call `Math.random`/`Date.now` on a behavioural path.
67
+
68
+ ## Pins
69
+
70
+ - `test/42` pins recognition idempotence under trace — traced and untraced
71
+ `recognise` must return the same cached object and site count.
72
+ - Determinism suites — `test/03`, `test/04`, `test/08`, `test/20` and others
73
+ assert same seed + same training ⇒ byte-identical answers and stores.