@hviana/sema 0.9.4 → 0.9.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (45) hide show
  1. package/AGENTS.md +16 -7
  2. package/dist/src/mind/learning.js +11 -12
  3. package/dist/src/mind/mind.js +4 -4
  4. package/dist/src/mind/types.d.ts +10 -15
  5. package/dist/src/store.d.ts +10 -10
  6. package/dist/src/store.js +5 -5
  7. package/docs/INDEX.md +62 -63
  8. package/docs/INVARIANTS.md +36 -18
  9. package/docs/architecture/bounded-reads.md +31 -71
  10. package/docs/architecture/caches.md +61 -81
  11. package/docs/architecture/closure.md +88 -107
  12. package/docs/architecture/commonality.md +47 -38
  13. package/docs/architecture/cost-model.md +57 -79
  14. package/docs/architecture/determinism.md +43 -55
  15. package/docs/architecture/evidence.md +158 -235
  16. package/docs/architecture/exact-vs-approximate.md +41 -35
  17. package/docs/architecture/factored-machinery.md +34 -20
  18. package/docs/architecture/fold-contract.md +110 -118
  19. package/docs/architecture/halo-sketch.md +105 -96
  20. package/docs/architecture/match-project.md +51 -42
  21. package/docs/architecture/mechanism-market.md +87 -91
  22. package/docs/architecture/memoization.md +60 -74
  23. package/docs/architecture/meter.md +37 -47
  24. package/docs/architecture/saturation.md +75 -101
  25. package/docs/architecture/store.md +118 -99
  26. package/docs/architecture/thresholds.md +66 -73
  27. package/docs/failures/tempting-but-wrong.md +139 -165
  28. package/docs/harness/gates.md +27 -32
  29. package/docs/mechanisms/alu.md +22 -69
  30. package/docs/mechanisms/cast.md +76 -71
  31. package/docs/mechanisms/confluence.md +22 -29
  32. package/docs/mechanisms/cover.md +58 -66
  33. package/docs/mechanisms/extraction.md +33 -37
  34. package/docs/mechanisms/prefix-completion.md +36 -39
  35. package/docs/mechanisms/recall.md +60 -53
  36. package/docs/mechanisms/reference.md +63 -49
  37. package/jsr.json +1 -1
  38. package/package.json +1 -1
  39. package/src/alu/README.md +90 -298
  40. package/src/derive/README.md +94 -256
  41. package/src/mind/learning.ts +11 -12
  42. package/src/mind/mind.ts +4 -4
  43. package/src/mind/types.ts +10 -15
  44. package/src/rabitq-ivf/README.md +11 -8
  45. package/src/store.ts +5 -5
@@ -1,87 +1,65 @@
1
1
  # Cost Model — One Currency
2
2
 
3
- Every mechanism and every byte competes on one cost ladder defined in
4
- `src/mind/graph-search.ts`. GraphSearch and `pipeline.ts:think` use the same
5
- units, so a mechanism-level choice and a byte-level choice are the same kind of
6
- decision: a lightest derivation.
7
-
8
- ## Ladder (`src/mind/graph-search.ts`)
9
-
10
- | Cost | Value | Meaning |
11
- | --------- | ------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
12
- | `MICRO` | `1e-3` | Recognised advance (one `rec` bridge); per-byte unit of the A\* heuristic. A recomposed form's onward edge is also `MICRO`. |
13
- | `STEP` | `1` | Every edge hop (first or fifth), every computed result, every projection. Charging every hop makes the lightest derivation the shortest chain. |
14
- | `CONCEPT` | `10` | Halo-mediated act (synonym hop, consensus climb) and abandoning an edge chain early (`CONCEPT` above chain cost — genuine fixpoint at `+0` always beats giving up at same depth). |
15
- | `PASS` | `1000` / byte | Carrying a byte nothing explains. Dominates everything so the search always prefers to recognise. |
16
-
17
- Only the **ordering** `MICRO < STEP < CONCEPT < PASS` matters; any constants
18
- with that order give the same derivations.
19
-
20
- ## Pipeline weighing (`src/mind/pipeline.ts:think`)
21
-
22
- Candidates are weighed in ONE place — a mechanism reports `moves` and
23
- `accounted`, never a price:
24
-
25
- ```
26
- weight = moves + PASS * unaccounted_bytes
27
- grade = floor(weight / STEP)
28
- ```
29
-
30
- `unaccounted` is what no `accounted` span covers. Comparison is at `STEP`
31
- resolution: lowest `grade` wins; at equal grade fewer `scaffolding` bytes
32
- (answer bytes lifted from unrecognised spans) wins; then list order.
33
-
34
- ## Two semirings
35
-
36
- - **(min, +) tropical** — lightest derivation in `GraphSearch` via `src/derive`
37
- (`lightestDerivation`). Cost accumulates with `+`, choice selects `min`.
38
- Powers `cover`/`form`/`out`, edge following, fusing, and the A\* agenda
39
- (`g + h`).
40
-
41
- - **(+, +) arithmetic** — evidence pooling in `src/mind/attention.ts:poolVotes`.
42
- Each region's vote is an axiom; rules carry `Rule.combine = 'sum'` so costs to
43
- the same anchor **add** rather than minimise. Powers IDF-weighted consensus,
44
- `votes`/`votesIdf`/`support`, and `regionSupport`/`regionPeak`.
45
-
46
- ## Admissibility
47
-
48
- The A\* heuristic is admissible and consistent:
49
-
50
- ```
51
- h(it) = (queryLen - right) * MICRO
52
- ```
53
-
54
- `right` is `p` for `cover(p)` or `j` for `form/out [i,j)`. `MICRO` is the
55
- minimum per-byte cost in the ladder (every real per-byte cost is `>= MICRO`,
56
- including `PASS`), and only the suffix past `right` is counted, so `h` never
57
- exceeds the true remaining cost.
58
-
59
- ## Dominance — why `PASS ≫ STEP` does not flood the chart
60
-
61
- The heuristic charges `MICRO` for a byte the goal will pay `PASS` for, so a
62
- cover that leaves bytes unexplained lets the search spend up to `PASS / STEP`
63
- hops looking for one more explained byte. Coverage itself cannot use them: every
64
- recognised completion of `[i, j)` advances the cover from `i` to `j` at the same
65
- `MICRO`. So a form or completion of `[i, j)` whose cost has reached that of a
66
- completion of `[i, j)` already yielded is DOMINATED and fires no rule
67
- (`buildSearch`, metered as `searchDominated`): every completion it could lead to
68
- costs at least as much, and a tie goes to the one yielded first. What a
69
- completion's BYTES could still do — fuse, splice, join — fires from the
70
- completion the search would stand on for that span, the same cure the join
71
- license and `deepen` apply. The first hop's stop-here (`STEP + CONCEPT`) is then
72
- a real horizon: no chain deeper than it is expanded.
3
+ > **Law:** every choice, whether a byte inside the search or a mechanism in the
4
+ > market, is a lightest derivation on one ladder (`mind/graph-search.ts`). The
5
+ > price is the question left unexplained, never confidence.
6
+
7
+ **Why.** Exact identity yields no confidence to choose by. What can be measured
8
+ exactly is how much of the question an answer accounts for. Pricing that makes
9
+ the winner the reading that explains the most, not the one most eager to speak.
10
+ When nothing explains the question, silence is the lightest answer.
11
+
12
+ ## The ladder
13
+
14
+ | Cost | Value | Charged for |
15
+ | --------- | ----------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
16
+ | `MICRO` | `1e-3` | a recognised advance, and a recomposed form's onward edge; it is also the A\* heuristic's unit per byte |
17
+ | `STEP` | `1` | every edge hop, computed result and projection. Charging every hop makes the lightest derivation the shortest chain |
18
+ | `CONCEPT` | `10` | an act mediated by a halo (a synonym hop, the consensus climb), and abandoning a chain early, so that a real fixpoint always beats giving up |
19
+ | `PASS` | `1000`/byte | carrying a byte nothing explains. It dominates everything, so the search always prefers to recognise |
20
+
21
+ Only the order `MICRO < STEP < CONCEPT < PASS` matters. Any constants with that
22
+ order give the same derivations. The market weighs candidates on the same
23
+ ladder, `moves + PASS·unaccounted`, compared at `STEP` grade
24
+ (`mechanism-market.md`).
25
+
26
+ ## Two semirings, one engine (`src/derive`)
27
+
28
+ - **(min, +), the tropical semiring:** the lightest derivation. Costs add along
29
+ a derivation, and the cheapest route to a conclusion wins. This powers
30
+ `cover`, form and continuation rules, edge following, fusion, and the A\*
31
+ agenda.
32
+ - **(+, +), the arithmetic semiring:** pooled evidence. Rules with
33
+ `combine: "sum"` add every independent line of evidence for a conclusion
34
+ instead of keeping the cheapest. This powers the consensus climb's votes
35
+ (`poolVotes`).
36
+
37
+ ## Admissibility and dominance
38
+
39
+ The heuristic `h = (queryLen − right) · MICRO` is admissible and consistent.
40
+ `MICRO` is the smallest cost per byte, and only the suffix past the item is
41
+ counted.
42
+
43
+ Because the heuristic charges `MICRO` for a byte the goal will charge `PASS`
44
+ for, the search could spend up to `PASS/STEP` hops looking for one more
45
+ explained byte. Coverage cannot use them: every recognised completion of
46
+ `[i, j)` advances the cover at the same price. So a form or completion of
47
+ `[i, j)` that costs as much as one already yielded is **dominated**, and fires
48
+ no rule (`searchDominated`). Fusion, splicing and joining fire from the
49
+ completion the search stands on, never from every alternative it reached. That
50
+ makes the first hop's stop-here (`STEP + CONCEPT`) a real horizon.
73
51
 
74
52
  ## Policy is not cost
75
53
 
76
- "Computation always wins" is **not** priced into the ladder (a computed result
77
- costs `STEP`, same as a learned edge). It is enforced by masking: `cover.ts`
78
- removes recognised sites overlapped by a `ComputedResult`, so the computation is
79
- the sole completion there. Keep policy in callers; keep the engine neutral.
54
+ "Computation always wins" is not priced. A computed result costs `STEP`, like a
55
+ learnt edge. It is enforced by masking: `cover.ts` removes any recognised site
56
+ overlapped by a computed span. Keep policy in the callers, and keep the engine
57
+ neutral. Tuning `PASS` to encode a preference breaks the one contract the ladder
58
+ has, its order.
80
59
 
81
60
  ## Pins
82
61
 
83
- - `test/52` — climb consensus instrumentation
84
- - `test/53` — cross-region probe instrumentation
85
- - `test/54` — evidence `k` instrumentation
86
- - `test/55` — cost meter (`Meter`, `CostReport`, `searchPops`/`searchPushes`)
87
- - `test/151` — dominance: a hub's degree generates no chart work
62
+ - `test/04`, `test/55` — the decider's weighing, and the cost meter.
63
+ - `test/151` — dominance: a hub's degree generates no work in the chart.
64
+ - `test/52`, `test/53`, `test/54` — instrumentation of the climb, the
65
+ cross-region probe and evidence `k`.
@@ -1,73 +1,61 @@
1
- # Determinism — Same Seed + Same Deposits + Same Query ⇒ Same Bytes
1
+ # Determinism — Same Seed, Same Deposits, Same Question, Same Bytes
2
2
 
3
- ## The law
3
+ > **Law:** the same `seed`, the same deposit order and the same query give a
4
+ > byte-identical answer. Every path that can reach output is a function of
5
+ > `(seed, store contents, query bytes)`.
4
6
 
5
- > Same `seed` + same deposit order + same query ⇒ byte-identical answer.
7
+ **Why.** An answer meant to be audited, contested or certified must replay.
8
+ Reproducibility is a property of the architecture, not a flag.
6
9
 
7
- Determinism is the product. Every code path that can reach output must be
8
- deterministic given `(seed, store contents, query bytes)`.
10
+ ## Forbidden on a behavioural path
9
11
 
10
- ## Forbidden
12
+ - `Math.random` and `Date.now`.
13
+ - Iteration over an unordered collection whose order can reach output.
11
14
 
12
- No `Math.random` or `Date.now` in behaviour, and no iteration over unordered
13
- collections where order can reach output. Example-only uses
14
- (`example/train_base`) are outside the library contract. If a test becomes
15
- flaky, the contract was broken, not the test.
15
+ If a test becomes flaky, the contract was broken, not the test. Uses inside
16
+ `example/` are outside the library contract.
16
17
 
17
18
  ## All randomness flows from `seed`
18
19
 
19
- `MindConfig.seed` (`src/config.ts:resolveConfig`, `DEFAULT_CONFIG`) is the sole
20
- entropy root. Subsystems derive deterministically:
20
+ `MindConfig.seed` (`config.ts`) is the only entropy root:
21
21
 
22
- - **Alphabet** — `Alphabet` (`src/alphabet.ts`) via `rng` (`src/vec.ts:rng`)
23
- seeded as `seed ^ seedMask`; builds 16→64→256 vectors.
24
- - **Keyring / Space** — `Space.seats` (`src/sema.ts:Space`) via `makeKeyring`
25
- (`src/vec.ts:makeKeyring`) and `rng` seeded from `seed` in `Mind`
26
- (`src/mind/mind.ts`); `fold`/`twoEndedSeat`/`companySignature` are pure over
27
- `Space`.
28
- - **Vector indexes** — `VectorDatabase` (`src/rabitq-ivf/src/database.ts`) and
29
- `Prng` (`src/rabitq-ivf/src/rabitq.ts`) seeded from config; insertion order is
30
- the stored order, not a random choice.
22
+ - the **alphabet** (`alphabet.ts`) is drawn from `rng(seed ^ seedMask)`;
23
+ - the **keyring of seats** (`makeKeyring`) is drawn from the seed in `Mind`;
24
+ - the **vector indexes** (`rabitq-ivf`) are seeded from config, and insertion
25
+ order is their stored order;
26
+ - **company signatures** are seeded by node id (`companySignature`), not by the
27
+ seed or by observation order. That is why halo comparisons survive a change of
28
+ seed while gist comparisons do not (`halo-sketch.md`).
31
29
 
32
- No other PRNG source may affect grounding. Thresholds in `geometry.ts` are
33
- derived from `D`/`W`/`N`, not sampled.
30
+ ## Ties are broken by the corpus
34
31
 
35
- ## Tie-breaks are corpus-determined
32
+ Every choice bottoms out in a fixed order, and the fallback is
33
+ **first-inserted**: the lowest node id, or the `LIMIT 1` insertion order. Never
34
+ last-inserted, because that would make an answer depend on recency instead of
35
+ evidence. The order of teaching is part of what was taught, and a correction
36
+ prevails only by evidence.
36
37
 
37
- Every choice bottoms out in a fixed ordering — insertion order or lowest node id
38
- — not interchangeable (`test/34`). The fallback is **first-inserted**:
38
+ - **`chooseNext` / `guidedFirst`** (`traverse.ts`) try three things in turn: a
39
+ continuation the question names (`evidence.md`); then distributional support,
40
+ `prevCount` and then halo mass; then first-inserted.
41
+ - **`chooseAmong`** takes `argmaxCosine` over `candidateGist`, capped at
42
+ `hubCap`, and resolves ties by stable scan.
43
+ - **Ties in the junction bridge** go to the shortest interior, then the lowest
44
+ node id.
39
45
 
40
- - `guidedFirst` (`src/mind/traverse.ts:guidedFirst`) — guided pick via
41
- `chooseNext` else first-inserted edge (`nextFirst` LIMIT 1).
42
- - `chooseNext` (`src/mind/traverse.ts:chooseNext`) — capped `nextFirst` read
43
- (`hubBound`), ranked by `prevCount` then `haloMass`; equal ⇒ first-inserted.
44
- - `chooseAmong` (`src/mind/traverse.ts:chooseAmong`) — `hubCap` + `argmaxCosine`
45
- over `candidateGist`; first-inserted on tie via stable scan.
46
- - `companySignature` (`src/sema.ts:companySignature`) — `rng(id ^ 0x9e3779b9)`,
47
- i.e. seeded by node id, not observation order.
46
+ When you add a choice among equals, name its tie-break explicitly and make it
47
+ corpus-determined.
48
48
 
49
- Never use last-inserted.
49
+ ## Memoization and tracing must not change the answer
50
50
 
51
- ## Memoization and trace must not break identity
52
-
53
- Per-response memos (`Precomputed`, `perceiveMemo`, `recogniseMemo`, `climbMemo`,
54
- `_resolvedSubtrees` via `foldTree`, `_edgeChoice`, `_gistCache` in
55
- `src/mind/mind.ts` / `src/mind/pipeline-mechanism.ts` /
56
- `src/mind/primitives.ts`) are sound because asking never writes. Only
57
- `guidedNext`/`sharedReachMemo` are trace-bypassed;
58
- `perceiveMemo`/`recogniseMemo`/`climbMemo` are always consulted — `foldTree`'s
59
- subtree fast path skips `visit` (and site emission) for cached subtrees, so
60
- bypassing makes `recognise` non-idempotent.
61
-
62
- ## Follow it
63
-
64
- When you add any choice among equals, name the tie-break explicitly and make it
65
- corpus-determined. Thread new randomness through `seed`-derived `rng`; never
66
- call `Math.random`/`Date.now` on a behavioural path.
51
+ Memos are sound because asking never writes. Which ones a trace bypasses, and
52
+ why the pick memo is cleared when the climb publishes its points, is
53
+ `memoization.md`'s trace boundary.
67
54
 
68
55
  ## Pins
69
56
 
70
- - `test/42` pins recognition idempotence under trace — traced and untraced
71
- `recognise` must return the same cached object and site count.
72
- - Determinism suites — `test/03`, `test/04`, `test/08`, `test/20` — assert same
73
- seed + same training ⇒ byte-identical answers and stores.
57
+ - `test/20`, `test/03`, `test/04`, `test/08` — the same seed and the same
58
+ training give byte-identical answers and stores.
59
+ - `test/42` — recognition is idempotent under trace.
60
+ - `test/155.4` — traced and untraced responses agree after the climb publishes.
61
+ - `test/34` — tie-breaks are corpus-determined, not interchangeable.