@hviana/sema 0.7.2 → 0.7.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (43) hide show
  1. package/AGENTS.md +95 -843
  2. package/README.md +11 -11
  3. package/dist/src/mind/mind.js +14 -0
  4. package/dist/src/store-sqlite.js +17 -0
  5. package/dist/src/store.d.ts +51 -4
  6. package/dist/src/store.js +81 -14
  7. package/docs/INDEX.md +71 -0
  8. package/docs/INVARIANTS.md +19 -0
  9. package/docs/architecture/bounded-reads.md +85 -0
  10. package/docs/architecture/caches.md +89 -0
  11. package/docs/architecture/commonality.md +45 -0
  12. package/docs/architecture/cost-model.md +71 -0
  13. package/docs/architecture/determinism.md +73 -0
  14. package/docs/architecture/exact-vs-approximate.md +47 -0
  15. package/docs/architecture/factored-machinery.md +28 -0
  16. package/docs/architecture/fold-contract.md +87 -0
  17. package/docs/architecture/halo-sketch.md +99 -0
  18. package/docs/architecture/match-project.md +62 -0
  19. package/docs/architecture/mechanism-market.md +95 -0
  20. package/docs/architecture/memoization.md +96 -0
  21. package/docs/architecture/meter.md +55 -0
  22. package/docs/architecture/saturation.md +92 -0
  23. package/docs/architecture/store.md +79 -0
  24. package/docs/architecture/thresholds.md +79 -0
  25. package/docs/failures/tempting-but-wrong.md +144 -0
  26. package/docs/harness/gates.md +56 -0
  27. package/docs/mechanisms/alu.md +75 -0
  28. package/docs/mechanisms/cast.md +75 -0
  29. package/docs/mechanisms/confluence.md +36 -0
  30. package/docs/mechanisms/cover.md +54 -0
  31. package/docs/mechanisms/extraction.md +53 -0
  32. package/docs/mechanisms/prefix-completion.md +54 -0
  33. package/docs/mechanisms/recall.md +69 -0
  34. package/docs/mechanisms/reference.md +58 -0
  35. package/jsr.json +1 -1
  36. package/package.json +1 -1
  37. package/src/mind/mind.ts +14 -0
  38. package/src/store-sqlite.ts +19 -0
  39. package/src/store.ts +92 -16
  40. package/test/89-completion-recursion.test.mjs +30 -10
  41. package/test/96-bytes-walk-termination.test.mjs +115 -0
  42. package/test/97-store-seed.test.mjs +105 -0
  43. package/HOW_IT_WORKS.md +0 -5836
package/README.md CHANGED
@@ -41,7 +41,7 @@ No weights. No gradients. No training loop. No neural network. No GPU.
41
41
  > Vector Symbolic Architecture (Plate 1995; Kanerva 2009) over a
42
42
  > content-addressable memory, with inference by weighted automated deduction
43
43
  > (Knuth 1977; Felzenszwalb & McAllester 2007). Each term is grounded in
44
- > [HOW_IT_WORKS.md](HOW_IT_WORKS.md).
44
+ > [docs/INDEX.md](docs/INDEX.md).
45
45
 
46
46
  ---
47
47
 
@@ -314,16 +314,16 @@ start talking — no install, no runtime, no API key.
314
314
 
315
315
  ## ✦ Learn more
316
316
 
317
- | Document | What's inside |
318
- | :----------------------------------------------------------------------------------- | :------------------------------------------------------------------------------------------------------------------------------------------------------- |
319
- | 📘 **[HOW_IT_WORKS.md](HOW_IT_WORKS.md)** | The full theory: vector symbolic architectures, the Merkle DAG, distributional halos, weighted deduction — concepts, diagrams, and extensive pseudocode. |
320
- | 🛠️ **[AGENTS.md](AGENTS.md)** | The development manual: repo layout, build/test, internals, invariants, and recipes for extending the system. |
321
- | 🎓 **[CITATION.cff](CITATION.cff)** | How to cite Sema in academic work. |
322
- | ⚖️ **[LICENSE.md](LICENSE.md)** | PolyForm Noncommercial License 1.0.0. |
323
- | 📚 **[DATASETS.md](DATASETS.md)** | Training corpora: provenance, per-corpus attribution, and how a trained memory file is licensed. |
324
- | 💼 **[COMMERCIAL-LICENSE.md](COMMERCIAL-LICENSE.md)** | Commercial licensing terms and contact. |
325
- | 🤗 **[Trained examples](https://huggingface.co/buckets/hviana/sema-trained-v1)** | Pre-trained memory files you can download and use directly. |
326
- | 💿 **[Binary examples](https://huggingface.co/buckets/hviana/sema-binary-examples)** | Ready-to-run web chat apps for Windows, Mac, and Linux — one file, no install. |
317
+ | Document | What's inside |
318
+ | :----------------------------------------------------------------------------------- | :----------------------------------------------------------------------------------------------------------------------------------------------------- |
319
+ | 📘 **[docs/INDEX.md](docs/INDEX.md)** | Architecture docs: single-system laws (docs/architecture/), mechanisms, invariants, and harness gates — the entry point for theory and implementation. |
320
+ | 🛠️ **[AGENTS.md](AGENTS.md)** | The development manual: repo layout, build/test, internals, invariants, and recipes for extending the system. |
321
+ | 🎓 **[CITATION.cff](CITATION.cff)** | How to cite Sema in academic work. |
322
+ | ⚖️ **[LICENSE.md](LICENSE.md)** | PolyForm Noncommercial License 1.0.0. |
323
+ | 📚 **[DATASETS.md](DATASETS.md)** | Training corpora: provenance, per-corpus attribution, and how a trained memory file is licensed. |
324
+ | 💼 **[COMMERCIAL-LICENSE.md](COMMERCIAL-LICENSE.md)** | Commercial licensing terms and contact. |
325
+ | 🤗 **[Trained examples](https://huggingface.co/buckets/hviana/sema-trained-v1)** | Pre-trained memory files you can download and use directly. |
326
+ | 💿 **[Binary examples](https://huggingface.co/buckets/hviana/sema-binary-examples)** | Ready-to-run web chat apps for Windows, Mac, and Linux — one file, no install. |
327
327
 
328
328
  ---
329
329
 
@@ -143,10 +143,24 @@ export class Mind {
143
143
  const { store: optsStore, mechanisms: userMechs, mechanismFactories: userFacts, canon: optsCanon, profile: optsProfile, ...rest } = (optsOrCfg ?? {});
144
144
  this._canonOpt = optsCanon ?? null;
145
145
  this._profile = optsProfile === true;
146
+ // `explicitSeed` is read BEFORE resolveConfig folds the default in, so
147
+ // the store can be consulted only when the caller did not choose.
148
+ const explicitSeed = rest.seed;
146
149
  this.cfg = resolveConfig(rest);
147
150
  this.store = optsStore ?? new SQliteStore({
148
151
  maxGroup: this.cfg.geometry.maxGroup,
149
152
  });
153
+ // THE ARTIFACT'S SEED GOVERNS. `train.seed` is recovered by the store at
154
+ // open, exactly like `train.D` and `geometry.maxGroup`. The seed feeds
155
+ // `makeKeyring`, `Space.rand` and the `Alphabet` below, so folding a
156
+ // query under config.ts's default (42) against a store trained with
157
+ // another seed (e.g. 7) lands in a DIFFERENT vector space than the one
158
+ // the artifact's nodes were folded into: recognition and resonance then
159
+ // read the wrong space and every answer degrades silently. An explicit
160
+ // caller seed still wins — this only replaces the unconfigured default.
161
+ if (explicitSeed === undefined && this.store.trainSeed !== null) {
162
+ this.cfg.seed = this.store.trainSeed;
163
+ }
150
164
  userMechanisms = userMechs ?? [];
151
165
  userFactories = userFacts ?? [];
152
166
  }
@@ -360,6 +360,23 @@ export class SQliteStore extends AbstractStore {
360
360
  this._maxGroup = g;
361
361
  }
362
362
  }
363
+ // Recover the TRAINING seed exactly as D and maxGroup are recovered. The
364
+ // seed seeds the alphabet and the seat keyring (mind.ts), so a Mind that
365
+ // folds a query under any other seed lands in a different vector space
366
+ // than the one this artifact's nodes were folded into — recognition,
367
+ // resonance and every mechanism downstream then read the wrong space. The
368
+ // trainer persists `train.seed` and refuses to resume against a store
369
+ // trained with a different one (example/train_base/main.ts), so the value
370
+ // is authoritative for this artifact. Absent on a store that was never
371
+ // trained, where the caller's configured seed stands.
372
+ {
373
+ const row = this.sqlite.prepare("SELECT val FROM meta WHERE key = 'train.seed'").get();
374
+ if (row) {
375
+ const s = Number(row.val);
376
+ if (Number.isInteger(s) && s >= 0)
377
+ this._trainSeed = s;
378
+ }
379
+ }
363
380
  // Persist maxGroup to meta when opening a FRESH store (no rows yet) so
364
381
  // indexSubtree always sees the training-time value even when the store is
365
382
  // accessed without a Mind / full snapshot.
@@ -114,6 +114,16 @@ export declare class BoundedMap<K, V> {
114
114
  }
115
115
  export interface Store {
116
116
  readonly D: number;
117
+ /** The seed the artifact was TRAINED with, recovered from the store's own
118
+ * `train.seed` metadata, or null for a store that was never trained.
119
+ *
120
+ * This is not decoration: the seed feeds the alphabet and the seat keyring
121
+ * (see the Mind constructor), so folding a query under any other seed lands
122
+ * in a different vector space than the one the artifact's nodes were folded
123
+ * into. A Mind opening a trained store MUST adopt this seed unless the
124
+ * caller explicitly overrides it — the same discipline that recovers
125
+ * `train.D` and `geometry.maxGroup` from the metadata. */
126
+ readonly trainSeed: number | null;
117
127
  /** The work accumulator for the inference call in flight, or null. The
118
128
  * Mind attaches one per profiled response and detaches it after (see
119
129
  * src/meter.ts). A store MUST only ever write to it — no read may reach
@@ -475,6 +485,10 @@ export declare abstract class AbstractStore implements Store {
475
485
  protected efFor(clusterCount: number): number;
476
486
  protected _D: number;
477
487
  protected _maxGroup: number;
488
+ /** `train.seed` recovered by the backend at open, or null when the store was
489
+ * never trained. A backend that omits it simply reports null, which leaves
490
+ * the caller's configured seed in force. */
491
+ protected _trainSeed: number | null;
478
492
  protected readonly minHaloMass: number;
479
493
  protected readonly efSearch: number;
480
494
  protected readonly overfetch: number;
@@ -563,6 +577,10 @@ export declare abstract class AbstractStore implements Store {
563
577
  protected _edgeSrcCount: number;
564
578
  constructor(config: StoreConfig, D: number, maxGroup: number);
565
579
  get D(): number;
580
+ /** The seed the artifact was trained with, recovered from `train.seed` at
581
+ * open. Null for a store that was never trained. See
582
+ * {@link Store.trainSeed} for why this must govern inference. */
583
+ get trainSeed(): number | null;
566
584
  /** Await the async initialisation performed by the concrete constructor. */
567
585
  protected _ensureReady(): Promise<void>;
568
586
  has(id: NodeId): boolean;
@@ -570,9 +588,6 @@ export declare abstract class AbstractStore implements Store {
570
588
  nodeCount(): number;
571
589
  size(): Promise<number>;
572
590
  get(id: NodeId): NodeRec | null;
573
- /** Reconstruct the bytes a node spans by traversing the DAG bottom-up.
574
- * Iterative post-order on an explicit stack — the call stack never sees the
575
- * tree depth, so even an adversarial chain of nodes stays safe. */
576
591
  /** How many reads hit a MISSING node record this session (a dangling edge
577
592
  * or kid id). Zero in a healthy store; a growing count means references
578
593
  * outlive their records — the read degrades safely to empty bytes, this
@@ -582,6 +597,25 @@ export declare abstract class AbstractStore implements Store {
582
597
  * nothing is profiling. Every read below bumps it through `?.`, so an
583
598
  * unprofiled store pays one null check per read and allocates nothing. */
584
599
  meter: Meter | null;
600
+ /** Reconstruct the bytes a node spans by traversing the DAG bottom-up.
601
+ * Iterative post-order on an explicit stack — the call stack never sees the
602
+ * tree depth, so even an adversarial chain of nodes stays safe.
603
+ *
604
+ * TERMINATION. The walk memoizes into a LOCAL map, and `_bytesCache` is
605
+ * consulted only as a warm hint whose hit is immediately promoted into that
606
+ * map. It used to use `_bytesCache` itself as the memo, which is not a
607
+ * memo at all: it EVICTS, and its `"smallest"` policy prefers precisely the
608
+ * freshly-resolved small children that the pending parents on the stack are
609
+ * waiting for. A parent then finds them uncached again, re-pushes them,
610
+ * they are re-resolved, re-inserted, re-evicted — the loop makes no
611
+ * progress and never exits. Latent until the cache saturates, then
612
+ * unconditional: observed in the wild at 19.9M nodes with the 20 MB cache
613
+ * pinned at 19,999,962/20,000,000 bytes, spinning 8h45m on a node whose
614
+ * whole content was 124 bytes (5 kids, 2 of them perpetually re-evicted).
615
+ * Because the loop is synchronous, no timer could fire — the trainer's stall
616
+ * watchdog never got a turn either. A local map resolves each node at most
617
+ * once per call, so the walk terminates by construction and `_bytesCache`
618
+ * goes back to being a pure speed hint. */
585
619
  bytes(id: NodeId): Uint8Array;
586
620
  /** First `maxLen` bytes of a node. Walks only the leftmost branch,
587
621
  * stopping at `maxLen` — so a 1 MB document root costs the same as a
@@ -665,7 +699,20 @@ export declare abstract class AbstractStore implements Store {
665
699
  * common-prefix / common-suffix trim: whatever remains after both trims is
666
700
  * the single differing span (substitution, insertion or deletion), and both
667
701
  * remainders must fit the budget. Scattered differences leave a wide
668
- * middle and are rejected. */
702
+ * middle and are rejected.
703
+ *
704
+ * Every read here is CAPPED (§2.8). It used to open with
705
+ * `bytesPrefix(k, Number.MAX_SAFE_INTEGER)` — the ALL sentinel, i.e. the
706
+ * full materialising `bytes()` read — on the deposit hot path, and only
707
+ * then compare lengths. So a candidate the length test was about to reject
708
+ * had already been reconstructed byte for byte. The LENGTHS decide first
709
+ * instead, from the `contentLen` memo the interning order has already built
710
+ * bottom-up, and the target's length is itself read under a cap: a target
711
+ * longer than `la + W` is rejected without touching one of its bytes.
712
+ * Same semantics — the old capped `b` read would have produced
713
+ * `a.length + W + 1` here and failed the very same test — strictly fewer
714
+ * byte reads. The `+ 1` on each byte cap keeps `_prefix`'s
715
+ * "complete reconstruction" test true, so the results still cache. */
669
716
  private differsByOneWindow;
670
717
  putLeaf(bytes: Uint8Array, gist: Vec): Promise<NodeId>;
671
718
  putBranch(kids: NodeId[], gist: Vec): Promise<NodeId>;
package/dist/src/store.js CHANGED
@@ -434,6 +434,10 @@ export class AbstractStore {
434
434
  // ── Config ─────────────────────────────────────────────────────────────
435
435
  _D;
436
436
  _maxGroup;
437
+ /** `train.seed` recovered by the backend at open, or null when the store was
438
+ * never trained. A backend that omits it simply reports null, which leaves
439
+ * the caller's configured seed in force. */
440
+ _trainSeed = null;
437
441
  minHaloMass;
438
442
  efSearch;
439
443
  overfetch;
@@ -552,6 +556,12 @@ export class AbstractStore {
552
556
  get D() {
553
557
  return this._D;
554
558
  }
559
+ /** The seed the artifact was trained with, recovered from `train.seed` at
560
+ * open. Null for a store that was never trained. See
561
+ * {@link Store.trainSeed} for why this must govern inference. */
562
+ get trainSeed() {
563
+ return this._trainSeed;
564
+ }
555
565
  /** Await the async initialisation performed by the concrete constructor. */
556
566
  async _ensureReady() {
557
567
  if (!this._ready)
@@ -591,9 +601,6 @@ export class AbstractStore {
591
601
  this._recCache.set(id, rec);
592
602
  return rec;
593
603
  }
594
- /** Reconstruct the bytes a node spans by traversing the DAG bottom-up.
595
- * Iterative post-order on an explicit stack — the call stack never sees the
596
- * tree depth, so even an adversarial chain of nodes stays safe. */
597
604
  /** How many reads hit a MISSING node record this session (a dangling edge
598
605
  * or kid id). Zero in a healthy store; a growing count means references
599
606
  * outlive their records — the read degrades safely to empty bytes, this
@@ -603,6 +610,25 @@ export class AbstractStore {
603
610
  * nothing is profiling. Every read below bumps it through `?.`, so an
604
611
  * unprofiled store pays one null check per read and allocates nothing. */
605
612
  meter = null;
613
+ /** Reconstruct the bytes a node spans by traversing the DAG bottom-up.
614
+ * Iterative post-order on an explicit stack — the call stack never sees the
615
+ * tree depth, so even an adversarial chain of nodes stays safe.
616
+ *
617
+ * TERMINATION. The walk memoizes into a LOCAL map, and `_bytesCache` is
618
+ * consulted only as a warm hint whose hit is immediately promoted into that
619
+ * map. It used to use `_bytesCache` itself as the memo, which is not a
620
+ * memo at all: it EVICTS, and its `"smallest"` policy prefers precisely the
621
+ * freshly-resolved small children that the pending parents on the stack are
622
+ * waiting for. A parent then finds them uncached again, re-pushes them,
623
+ * they are re-resolved, re-inserted, re-evicted — the loop makes no
624
+ * progress and never exits. Latent until the cache saturates, then
625
+ * unconditional: observed in the wild at 19.9M nodes with the 20 MB cache
626
+ * pinned at 19,999,962/20,000,000 bytes, spinning 8h45m on a node whose
627
+ * whole content was 124 bytes (5 kids, 2 of them perpetually re-evicted).
628
+ * Because the loop is synchronous, no timer could fire — the trainer's stall
629
+ * watchdog never got a turn either. A local map resolves each node at most
630
+ * once per call, so the walk terminates by construction and `_bytesCache`
631
+ * goes back to being a pure speed hint. */
606
632
  bytes(id) {
607
633
  if (this.meter) {
608
634
  this.meter.byteReads++;
@@ -619,10 +645,22 @@ export class AbstractStore {
619
645
  return hit;
620
646
  const stack = [id];
621
647
  const cache = this._bytesCache;
648
+ // The walk's own memo. Entries are the same shared arrays `_bytesCache`
649
+ // holds (no extra copy), and it lives exactly as long as this call.
650
+ const done = new Map();
622
651
  while (stack.length > 0) {
623
652
  const nid = stack[stack.length - 1]; // peek
624
- // Already resolved by an earlier traversal.
625
- if (cache.get(nid)) {
653
+ // Already resolved by this walk — the ONLY authority the readiness test
654
+ // below trusts, because it cannot be evicted underneath us.
655
+ if (done.has(nid)) {
656
+ stack.pop();
657
+ continue;
658
+ }
659
+ // Warm hint: a hit is promoted into `done` in the same step, so from
660
+ // here on the entry is pinned for the rest of the walk.
661
+ const warm = cache.get(nid);
662
+ if (warm !== undefined) {
663
+ done.set(nid, warm);
626
664
  stack.pop();
627
665
  continue;
628
666
  }
@@ -634,21 +672,26 @@ export class AbstractStore {
634
672
  // The cache makes the empty read permanent for the session; the
635
673
  // counter survives as the visible trace.
636
674
  this.danglingReads++;
675
+ done.set(nid, _ZERO);
637
676
  cache.set(nid, _ZERO);
638
677
  stack.pop();
639
678
  continue;
640
679
  }
641
680
  if (rec.leaf) {
642
- cache.set(nid, new Uint8Array(rec.leaf));
681
+ // COPY before caching: rec.leaf is the node record's own buffer, and
682
+ // handing it out would let one mutating caller corrupt the record.
683
+ const leaf = new Uint8Array(rec.leaf);
684
+ done.set(nid, leaf);
685
+ cache.set(nid, leaf);
643
686
  stack.pop();
644
687
  continue;
645
688
  }
646
- // Branch — push any uncached children (reverse order so they resolve
647
- // left-to-right). If every child is already cached, concatenate now.
689
+ // Branch — push any unresolved children (reverse order so they resolve
690
+ // left-to-right). If every child is resolved, concatenate now.
648
691
  const kids = rec.kids ?? [];
649
692
  let ready = true;
650
693
  for (let i = kids.length - 1; i >= 0; i--) {
651
- if (!cache.get(kids[i])) {
694
+ if (!done.has(kids[i])) {
652
695
  stack.push(kids[i]);
653
696
  ready = false;
654
697
  }
@@ -656,10 +699,11 @@ export class AbstractStore {
656
699
  if (!ready)
657
700
  continue;
658
701
  stack.pop();
659
- const out = concat(kids.map((k) => cache.get(k)));
702
+ const out = concat(kids.map((k) => done.get(k)));
703
+ done.set(nid, out);
660
704
  cache.set(nid, out);
661
705
  }
662
- const out = cache.get(id) ?? _ZERO;
706
+ const out = done.get(id) ?? _ZERO;
663
707
  if (this.meter)
664
708
  this.meter.bytesRead += out.length;
665
709
  return out;
@@ -1189,10 +1233,33 @@ export class AbstractStore {
1189
1233
  * common-prefix / common-suffix trim: whatever remains after both trims is
1190
1234
  * the single differing span (substitution, insertion or deletion), and both
1191
1235
  * remainders must fit the budget. Scattered differences leave a wide
1192
- * middle and are rejected. */
1236
+ * middle and are rejected.
1237
+ *
1238
+ * Every read here is CAPPED (§2.8). It used to open with
1239
+ * `bytesPrefix(k, Number.MAX_SAFE_INTEGER)` — the ALL sentinel, i.e. the
1240
+ * full materialising `bytes()` read — on the deposit hot path, and only
1241
+ * then compare lengths. So a candidate the length test was about to reject
1242
+ * had already been reconstructed byte for byte. The LENGTHS decide first
1243
+ * instead, from the `contentLen` memo the interning order has already built
1244
+ * bottom-up, and the target's length is itself read under a cap: a target
1245
+ * longer than `la + W` is rejected without touching one of its bytes.
1246
+ * Same semantics — the old capped `b` read would have produced
1247
+ * `a.length + W + 1` here and failed the very same test — strictly fewer
1248
+ * byte reads. The `+ 1` on each byte cap keeps `_prefix`'s
1249
+ * "complete reconstruction" test true, so the results still cache. */
1193
1250
  differsByOneWindow(kids, targetId, W) {
1194
- const a = concat(kids.map((k) => this.bytesPrefix(k, Number.MAX_SAFE_INTEGER)));
1195
- const b = this.bytesPrefix(targetId, a.length + W + 1);
1251
+ const lens = kids.map((k) => this.contentLen(k));
1252
+ let la = 0;
1253
+ for (const n of lens)
1254
+ la += n;
1255
+ const cap = la + W + 1;
1256
+ // `contentLen` under a cap returns a clamped LOWER BOUND once the partial
1257
+ // sum reaches it, so `>= cap` is exactly "longer than la + W".
1258
+ const lb = this.contentLen(targetId, cap);
1259
+ if (lb >= cap || Math.abs(la - lb) > W)
1260
+ return false;
1261
+ const a = concat(kids.map((k, i) => this.bytesPrefix(k, lens[i] + 1)));
1262
+ const b = this.bytesPrefix(targetId, lb + 1);
1196
1263
  if (Math.abs(a.length - b.length) > W)
1197
1264
  return false;
1198
1265
  const n = Math.min(a.length, b.length);
package/docs/INDEX.md ADDED
@@ -0,0 +1,71 @@
1
+ # Sema Documentation Index
2
+
3
+ Sema is a single system stated three ways: the law lives in `docs/architecture/`
4
+ (what holds), the prescription in `AGENTS.md` bootloader (what to do and where),
5
+ and the proof in `test/` (pins that fail when the law is broken).
6
+
7
+ ## Routing — what to read for each task
8
+
9
+ | Task | Read | Why |
10
+ | --------------------------- | ---------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------- |
11
+ | Add a mechanism | `docs/architecture/mechanism-market.md` + `docs/mechanisms/*.md` | Market contract: decoupled, declared competence, visible budget, evidence travels |
12
+ | Add a threshold | `docs/architecture/thresholds.md` | All cutoffs are formulas over D/W/N in `geometry.ts`; `config.ts` holds only budgets |
13
+ | Debug an answer | `docs/architecture/cost-model.md` + `src/meter.ts` | One cost ladder (`MICRO`/`STEP`/`CONCEPT`/`PASS`) decides every grounding choice |
14
+ | Understand the fold | `docs/architecture/fold-contract.md` | Deposit and inference must compute the same tree; boundaries are not turn metadata |
15
+ | Add a store backend | `docs/architecture/store.md` + `docs/architecture/bounded-reads.md` | `AbstractStore` owns domain logic; backends are thin wrappers with capped reads |
16
+ | Add an ALU operation | `src/alu/README.md` | One `registry.derive` per op composing existing ops; no new `derive` needed |
17
+ | Add a matcher or projection | `docs/architecture/match-project.md` | Mechanisms are `(matcher, direction, gate)` configs over the shared `match.ts` family |
18
+ | Add a deduction rule | `docs/architecture/cost-model.md` + `docs/architecture/determinism.md` | Place cost on the ladder, keep heuristic admissible, extend `classifyMove` |
19
+ | Change vector search | `docs/architecture/exact-vs-approximate.md` + `docs/architecture/bounded-reads.md` | Scores propose, bytes dispose; ANN is bounded by `hubBound` |
20
+ | Profile or bound work | `docs/architecture/meter.md` + `docs/architecture/bounded-reads.md` | `meter.ts` is write-only; counters are product, phases are hints |
21
+
22
+ ## Architecture laws (13)
23
+
24
+ | Law | File | Summary | Pins |
25
+ | --- | ------------------------------------------- | ----------------------------------------------------------------------------------------------------------- | -------------------- |
26
+ | 1 | `docs/architecture/determinism.md` | No `Math.random`/`Date.now` in behaviour; seed-derived randomness; corpus-determined tie-breaks | `test/20` |
27
+ | 2 | `docs/architecture/thresholds.md` | Every decision cutoff derived in `geometry.ts` over D/W/N; no tunable knobs | `test/40`, `test/64` |
28
+ | 3 | `docs/architecture/exact-vs-approximate.md` | Vector scores rank only; identity via content-addressed lookup; five graded ladders | `test/51` |
29
+ | 4 | `docs/architecture/cost-model.md` | Single ladder `MICRO`/`STEP`/`CONCEPT`/`PASS`; weight `moves + PASS·unaccounted`; `STEP`-grade compare | `test/04`, `test/55` |
30
+ | 5 | `docs/architecture/match-project.md` | Shared `match.ts` family (`locate`/`alignGraded`/`frameSlots`/`project`); voicing gates belong to consumers | `test/24`, `test/76` |
31
+ | 6 | `docs/architecture/mechanism-market.md` | `PipelineMechanism` (`floor`/`run`/`parse`); admissible-floor pruning and investment discipline | `test/01`, `test/04` |
32
+ | 7 | `docs/architecture/commonality.md` | Two populations: corpus-global (`reachOf`+`dominates`) vs weave-local (`depth[]`) | `test/17`, `test/34` |
33
+ | 8 | `docs/architecture/bounded-reads.md` | No per-query read grows with N; `hubBound=√N` enforced at store via LIMIT/probe/prefix caps | `test/77`, `test/90` |
34
+ | 9 | `docs/architecture/store.md` | `AbstractStore` owns dedup/indexing/batch; `store-sqlite.ts` is thin wrappers; canon index optional | `test/08` |
35
+ | 10 | `docs/architecture/fold-contract.md` | `perceiveDeposit` and `perceive` agree; `contentLevels` is single boundary rule; no W/offset dependence | `test/59`, `test/63` |
36
+ | 11 | `docs/architecture/memoization.md` | `Precomputed` is per-response lazy cache (promise-cached async); `beginResponse`/`endResponse` lifecycle | `test/42` |
37
+ | 12 | `docs/architecture/saturation.md` | Every walk names a deciding saturation beside its cap; cap is safety net, not decision | `test/27`, `test/16` |
38
+ | 13 | `docs/architecture/meter.md` | `meter.ts` is write-only work accounting; counts are deterministic, phases nest | `test/55` |
39
+
40
+ ## Mechanisms (8)
41
+
42
+ | Mechanism | File | Role |
43
+ | ----------------- | -------------------------------------- | --------------------------------------------------------- |
44
+ | cover | `docs/mechanisms/cover.md` | Exact/computed-span covering via `GraphSearch` |
45
+ | cast | `docs/mechanisms/cast.md` | Weave-local analogy via `depth[]` frame gate |
46
+ | confluence | `docs/mechanisms/confluence.md` | Corpus-global filler/scaffolding gate over climb |
47
+ | extraction | `docs/mechanisms/extraction.md` | Located-frame read-out with anchored span accounting |
48
+ | reference | `docs/mechanisms/reference.md` | Slot-bound voicing of asker-supplied referents |
49
+ | recall | `docs/mechanisms/recall.md` | Nearest stored form; echo tier via substitution bridge |
50
+ | prefix-completion | `docs/mechanisms/prefix-completion.md` | Literal prefix of exactly one trained form |
51
+ | alu | `docs/mechanisms/alu.md` | Authoritative computed spans (`parse` → `aluToMechanism`) |
52
+
53
+ ## Supporting docs
54
+
55
+ | Doc | Role |
56
+ | ----------------------------------------- | ------------------------------------------------------------------------------- |
57
+ | `docs/architecture/caches.md` | Every acceleration is a `BoundedMap`; miss re-derives; budgets in `StoreConfig` |
58
+ | `docs/architecture/halo-sketch.md` | Halo & sketch — distributional memory, quantization, bottom-k profiles |
59
+ | `docs/architecture/factored-machinery.md` | Single-definition contracts table — one owner per shared symbol |
60
+
61
+ ## Cross-cutting
62
+
63
+ - `docs/INVARIANTS.md` — the five invariants (determinism, derived thresholds,
64
+ exact-decides, one cost currency, bounded reads) with file-level routing.
65
+ - `docs/failures/tempting-but-wrong.md` — refuted simplifications that passed
66
+ review but failed pins (e.g. reordering ladders, flattening attention
67
+ asymmetries, one-cone-exhausted stop).
68
+ - `docs/harness/gates.md` — how `AGENTS.md` recipes,
69
+ `bench/profile-inference.mjs`, and `test/*.test.mjs` enforce the laws.
70
+ - `docs/architecture/` — full per-law derivation (each file states why the law
71
+ is this way; no separate HOW).
@@ -0,0 +1,19 @@
1
+ # INVARIANTS — Laws, Proofs, Derivations
2
+
3
+ > Law in `docs/architecture/*.md`, proof in `test/*.test.mjs`.
4
+
5
+ | # | Law | Where defined (src symbol) | Pins (test/N) | Doc (docs/architecture/*.md) |
6
+ | -- | ------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------- | ------------------- | ---------------------------- |
7
+ | 1 | Determinism | `src/config.ts:seed` `src/alphabet.ts:Alphabet` `src/mind/traverse.ts:guidedFirst` | `test/20` `test/42` | `determinism.md` |
8
+ | 2 | Derived thresholds | `src/geometry.ts:mergeThreshold,identityBar,reachThreshold,significanceBar,consensusFloor,dominates` | `test/64` `test/40` | `thresholds.md` |
9
+ | 3 | Exact decides / approximate proposes | `src/mind/primitives.ts:resolve` `src/mind/match.ts:locate,alignGraded` `src/mind/resonance.ts:bridge` | `test/51` | `exact-vs-approximate.md` |
10
+ | 4 | One cost currency | `src/mind/graph-search.ts:MICRO,STEP,CONCEPT,PASS` `src/derive:lightestDerivation` (min,+) `src/mind/attention.ts:poolVotes` (+,+) | `test/55` `test/04` | `cost-model.md` |
11
+ | 5 | Bounded reads | `src/store.ts:AbstractStore:nextFirst,parentsFirst,containersSlice,hasNext,bytesPrefix` `src/mind/traverse.ts:hubBound,hubCap` | `test/90` `test/14` | `bounded-reads.md` |
12
+ | 6 | Fold contract | `src/geometry.ts:contentLevels` `src/mind/canonical.ts:canonicalWindows,chainReach` `src/canon.ts:canonicalizer` | `test/59` `test/63` | `fold-contract.md` |
13
+ | 7 | Mechanism market | `src/mind/pipeline-mechanism.ts:PipelineMechanism,Precomputed` `src/mind/pipeline.ts:think,worthRunning` | `test/01` `test/04` | `mechanism-market.md` |
14
+ | 8 | Two commonality measures | `src/mind/traverse.ts:reachOf,dominates,corpusN` (global) `src/mind/match.ts:depth[],MIN_WEAVE` (weave-local) | `test/17` `test/34` | `commonality.md` |
15
+ | 9 | Memoization idempotence | `src/mind/pipeline-mechanism.ts:Precomputed` `src/mind/mind.ts:beginResponse,endResponse,_resolvedSubtrees` | `test/42` | `memoization.md` |
16
+ | 10 | Caches as budgets | `src/store.ts:BoundedMap` `src/config.ts:StoreConfig:bytesCacheMax,recCacheBytes,haloCacheBytes` | `test/96` `test/91` | `caches.md` |
17
+ | 11 | Honest degradation | `src/mind/pipeline.ts:weight=moves+PASS*unaccounted` `src/store.ts:BoundedMap:miss→re-derive` | `test/28` `test/84` | `store.md`+`caches.md` |
18
+ | 12 | Meter contracts | `src/meter.ts:Meter,PhaseCost,time` `src/mind/pipeline-mechanism.ts:Precomputed.shared` | `test/55` | `meter.md` |
19
+ | 13 | Saturation | `src/mind/traverse.ts:edgeAncestors:SaturationReason` `src/mind/junction.ts:junctionContainersFrom` `src/mind/resonance.ts:pivotInto` | `test/27` `test/16` | `saturation.md` |
@@ -0,0 +1,85 @@
1
+ # Bounded Reads — No Per-Query Read Grows With the Corpus
2
+
3
+ > **Law:** the cost of one query is proportional to the query, not to how much
4
+ > was learned. No per-query read may grow with corpus size N.
5
+
6
+ Every fan-out, walk, and disambiguation is capped at `hubBound` —
7
+ `ceil(sqrt(N))` — derived once from `corpusN` and floored at 2 so `sqrt` and
8
+ `ln` stay meaningful on a near-empty store. There is no second convention; do
9
+ not invent one.
10
+
11
+ ## Scale
12
+
13
+ ```
14
+ corpusN(ctx) = max(2, store.edgeSourceCount()) // distinct learnt contexts
15
+ hubBound(ctx) = ceil(sqrt(corpusN(ctx))) // >= 2, the store cap
16
+ hubCap(ctx, ids) = ids.slice(0, hubBound(ctx)) // list-side reading
17
+ boundFor(n) = ceil(sqrt(max(2, n))) // ctx-free reading
18
+ ```
19
+
20
+ Defined once in `mind/traverse.ts` (`corpusN`, `hubBound`, `hubCap`,
21
+ `boundFor`). Every consumer imports them; never spell `Math.sqrt` inline.
22
+
23
+ ## Enforcement at the store level
24
+
25
+ The cap is not advisory — adapters must make bounded reads bounded in SQL.
26
+
27
+ ### 1. LIMITed reads — real `LIMIT ?`
28
+
29
+ `nextFirst(id, limit)`, `prevFirst(id, limit)`, `parentsFirst(id, limit)`,
30
+ `containersSlice(child, offset, limit)`.
31
+
32
+ Same statement and `ORDER BY` as the full read, with `LIMIT ?`. Never
33
+ "materialise then slice". Reading `hubBound + 1` parents decides "hub or not"
34
+ exactly without reading the rest. Implemented as thin wrappers in
35
+ `store-sqlite.ts` over `AbstractStore` in `store.ts`.
36
+
37
+ ### 2. Existence probes — indexed point probes
38
+
39
+ `hasNext(id)`, `hasParents(id)`, `hasContainers(child)`, `hasHalo(id)`,
40
+ `prevCount(id)`.
41
+
42
+ One indexed `EXISTS` / `COUNT` probe that never decodes vectors or unpacks
43
+ blobs. Use them for every "does this lead anywhere?" question instead of
44
+ `next(id).length > 0` or `prev(id).length`. `prevCount` is the reverse-edge
45
+ support count for `chooseNext`/`chooseAmong`; `hasNext`/`hasHalo` gate the
46
+ `leadsSomewhere` admission predicate in `mind/traverse.ts`.
47
+
48
+ ### 3. Prefix-capped reads — reject without reconstruction
49
+
50
+ `bytesPrefix(id, cap)` and `contentLen(id, cap)`.
51
+
52
+ `contentLen` under a cap returns an exact length below the cap and `>= cap`
53
+ otherwise — an indexed memo walk that stops early, never a full subtree walk.
54
+ `bytesPrefix` stops after `cap` bytes. A candidate exceeding the cap is rejected
55
+ on the length probe alone; the weave, junction walks, and bridge all read this
56
+ way. Uncapped reads there cost seconds per query on a large store.
57
+
58
+ ### 4. Transparent scaffolding — one bounded read
59
+
60
+ `chainRun(id)` climbs a run of transparent nodes (exactly one parent, no edges)
61
+ in a single recursive CTE, cached for the store lifetime and dropped on writes
62
+ that break transparency. The climber hops the whole run where a node-at-a-time
63
+ ascent would pay three probes per node.
64
+
65
+ ## Maintenance only
66
+
67
+ The full materialising reads — `next(id)`, `prev(id)`, `parents(id)`,
68
+ `containers(child)` — exist for inspection, repair, and compaction only
69
+ (`compactContentIndex`, `repairContentIndex`). Keep them off hot paths.
70
+
71
+ ## Adding a walk
72
+
73
+ Any new fan-out walk uses `hubBound`/`hubCap`. Do not call `edgeSourceCount()`
74
+ or `Math.ceil(Math.sqrt(...))` inline, and do not invent a per-walk limit. The
75
+ walk's saturation decision (when to stop) is separate from the cap (the safety
76
+ net); a walk with only a cap drifts to the cap.
77
+
78
+ ## Pins
79
+
80
+ - `test/14` — sublinear inference in corpus size and constant-rate in input
81
+ length; training throughput floor; exact recall at scale.
82
+ - `test/89` — completion recursion stays output-sensitive (nested searches/pops
83
+ sublinear); guards the count of reads, not just per-read size.
84
+ - `test/90` — connector probe (`resolveConnectors`) reads by the query length
85
+ (`QUERY.length + 1`), not by the learnt continuation; per-read size bound.
@@ -0,0 +1,89 @@
1
+ # Caches — Every Acceleration Is a BoundedMap
2
+
3
+ > **Law:** every acceleration is a `BoundedMap` with a byte budget. A miss
4
+ > re-derives from durable state. Degradation order is speed/reach lost, never
5
+ > identity.
6
+
7
+ No cache may change what is stored, what is resolved, or what tree is folded.
8
+ Eviction costs a re-read, a re-walk, or a narrower reach — never a wrong answer
9
+ or a wrong tree.
10
+
11
+ ## `BoundedMap` — the one cache primitive
12
+
13
+ `src/store.ts:BoundedMap<K,V>` — LRU with byte accounting (`maxBytes`, `sizeOf`,
14
+ `evict`, `recency`).
15
+
16
+ - `evict: "lru"` — uniform-cost entries (dedup, vectors, records).
17
+ - `evict: "smallest"` — variable-cost reconstruction (`_bytesCache`): protects
18
+ expensive large branches over cheap leaves.
19
+ - `recency: "reorder"` (default) — exact LRU via `delete+set`; required when
20
+ eviction choice is load-bearing (`_depositTrees` — 8 entries, victim changes
21
+ fold).
22
+ - `recency: "clock"` — bit instead of reorder; only for transparent caches where
23
+ wrong victim costs a re-read (`_bytesCache`, `_recCache`). Measured: same
24
+ entries cached, hot-path time 55% → bit.
25
+
26
+ Persistent cursor over V8 insertion order makes eviction amortised O(1);
27
+ candidate window for `"smallest"` never rescans from front.
28
+
29
+ ## Store caches — budgets in `src/config.ts:StoreConfig`
30
+
31
+ | Cache | Field | Budget | `sizeOf` | Eviction |
32
+ | ------------------- | -------------------------- | ---------------------------- | ------------------- | -------------- |
33
+ | dedup leaf/branch | `_leafKey` / `_branchKey` | `dedupCacheMax` 1M entries | 1 | lru |
34
+ | reconstructed bytes | `_bytesCache` | `bytesCacheMax` 20 MB | `byteLength` | smallest+clock |
35
+ | content length | `_lenCache` | `bytesCacheMax` | 16 | lru |
36
+ | node records | `_recCache` | `recCacheBytes` 10 MB | leaf+4·kids+12 | lru+clock |
37
+ | pending gists | `_pendingGist` | `pendingGistBytes` 16 MB | `byteLength` (D·4) | lru |
38
+ | halo exact / norm | `_haloExact` / `_haloNorm` | `haloCacheBytes` 16 MB each | `byteLength` | lru |
39
+ | skipped interiors | `_coveredIds` | `coveredIdsMax` 100K entries | 1 | lru |
40
+ | indexed ids | `_indexedIds` | `coveredIdsMax` | 1 | lru |
41
+ | transparent chains | `_chainMemo` | `chainCacheBytes` 16 MB | 4·len+32 | lru |
42
+ | ingest memo | `CachedIngest._memo` | `ingestCacheBytes` 50 MB | vector+ids+keyBytes | lru |
43
+
44
+ ANN read caches (`_resonateCache`, `_resonateHaloCache`) are `Map<string,Hit[]>`
45
+ keyed by `vecKey(v)+":"+k`, dropped on any index mutation.
46
+ `vectorCacheMb`/`sqliteCacheMb` are pure page-cache latency knobs.
47
+
48
+ `_bytesCache` only caches complete reconstructions — `bytesPrefix(id,cap)` with
49
+ `got < cap`; a truncated prefix is never stored. `_chainMemo` is dropped on any
50
+ write that could break transparency; `_pendingGist` eviction falls back to DAG
51
+ climb; halo eviction re-decodes the durable 2-bit row.
52
+
53
+ ## Mind caches — session and per-response
54
+
55
+ | Cache | Location | Budget | Scope / invalidation |
56
+ | ------------------- | ------------------------------------------------ | --------- | ------------------------------------------------------------------ |
57
+ | `_gistCache` | `Mind._gistCache` | 32 MB | session-lifetime, never invalidated (perception pure) |
58
+ | `_depositTrees` | `Mind._depositTrees` | 8 entries | session; `perceiveDeposit` only when `conversational` |
59
+ | `_depositLens` | `Mind._depositLens` | — | byte lengths for prefix probes; cleared with map when >64 |
60
+ | `_internIds` | `Mind._internIds: WeakMap<Sema,number>` | — | Mind lifetime; ids permanent |
61
+ | `_resolvedSubtrees` | `Mind._resolvedSubtrees: WeakMap<Sema,{id,len}>` | — | per-response/conversation; fast path only when `visit===undefined` |
62
+
63
+ `REACH_MEMO_MAX` / `STRUCT_MEMO_MAX` 100K (`src/mind/traverse.ts`) — whole-climb
64
+ and per-node structural probes (`hasNext`/`prevCount`/`hasParents`); cleared on
65
+ write or when cap reached. `reachMemo`/`structCaches` keyed by `_structMemoKey`,
66
+ bypassed under trace.
67
+
68
+ ## Deposit caches — offset-keyed, caller-discharged
69
+
70
+ `_depositTrees`/`_depositLens`/`_internIds`/`_resolvedSubtrees` key **offsets**,
71
+ not bytes, for O(1) reuse. Offsets alone cannot witness byte agreement — caller
72
+ must discharge it.
73
+
74
+ - Correct: conversation append — each turn extends the prior cumulative context
75
+ by its own bytes; longest cached proper prefix hit (`L < bytes.length`) reuses
76
+ `contentFoldIncremental` segments bit-identically.
77
+ - Wrong: mismatched `prev` reused by offset produced wrong tree (336 vs 400
78
+ bytes) — a coincidental prefix length aliased unrelated content.
79
+ - Now: `perceiveDeposit` keys by `latin1Key(bytes.subarray(0,L))` (prefix
80
+ bytes), probes longest cached proper prefix first; `_depositTrees` populated
81
+ only for conversational deposits (budget discipline), otherwise cold path
82
+ always correct.
83
+
84
+ ## Pins
85
+
86
+ - `test/91` — `chainRun` via capped `_prefix`: bounded transparent-chain hop,
87
+ not per-node probes.
88
+ - `test/96` — `_bytesCache` is byte-accounted `BoundedMap` that evicts; miss
89
+ re-derives.