gemchat 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,1900 @@
1
+ # GemChat Local — Bundle-Scoped Gem Documentation Index
2
+
3
+ **Repo:** `/Users/zoras/work/gemchat` (new `bundle gem gemchat` skeleton)
4
+ **Gem name:** `gemchat` (RubyGems name confirmed available; overlap with gemchat.org is intentional)
5
+ **Status:** `init`, `index`, `reindex`, `search`, `vsearch`, `query`, `embed`,
6
+ `status`, `version`, and a working trust-gated, sha256-gated auto-index plugin.
7
+ Next up: the lexical arm of the fusion — see §11.4, which is designed and
8
+ eval-gated but not built.
9
+
10
+ ## 1. Summary
11
+
12
+ A local, offline, agent-facing search index over the gems in *your*
13
+ `Gemfile.lock`, modelled on [qmd](https://github.com/tobi/qmd). Ships as a
14
+ Ruby gem with a CLI on `PATH`.
15
+
16
+ It lives in a **separate repo** from the Rails app (`/Users/zoras/work/gemchat_app`)
17
+ because it must:
18
+
19
+ - install and start in milliseconds without booting Rails,
20
+ - have no PostgreSQL/pgvector dependency,
21
+ - version and release independently of the hosted service.
22
+
23
+ The problem it actually solves is **bundle awareness**. A Rails `Gemfile.lock`
24
+ is 500–1500 lines; making an agent read it to discover your stack burns context
25
+ for something a deterministic script does better. Secondary wins: exact-lockfile
26
+ version fidelity, privacy (your dependency graph stays local), and offline use.
27
+
28
+ This is a **retrieval** tool. It returns chunks and citations. It does not
29
+ generate prose answers.
30
+
31
+ ## 2. End-user experience
32
+
33
+ The goal in one sentence: **an agent working in my repo can answer "how do I
34
+ use *my* pinned version of this gem" without me pasting a 1500-line
35
+ `Gemfile.lock` into its context, and without my dependency graph leaving my
36
+ machine.**
37
+
38
+ ### Path A — local (default)
39
+
40
+ ```bash
41
+ $ bundle add gemchat --group development # or: gem install gemchat
42
+ $ gemchat init
43
+ Appends `plugin "gemchat"` to your Gemfile and writes .gemchat.yml. [y/N] y
44
+ Run `bundle install` once to activate it.
45
+
46
+ $ bundle install && gemchat index
47
+ ri: 0/27 found — Bundler installs without documentation. Generating locally…
48
+ [12/27] sidekiq 8.1.7 1,482 chunks (18s)
49
+ Indexed 27 gems · 4,312 chunks · 14MB · ri: 27/27 generated
50
+
51
+ $ gemchat search "retry policy" # instant, BM25, offline, no model
52
+ $ gemchat query "configure sidekiq retries" # hybrid, exact 7.3.9
53
+ → [sidekiq 7.3.9] lib/sidekiq/options.rb:42
54
+ ```
55
+
56
+ ### Path B — hosted, zero local footprint
57
+
58
+ For users who don't want ri generated, a 146MB model, or the disk at all:
59
+
60
+ ```bash
61
+ $ gem install gemchat
62
+ $ gemchat init --hosted # records hosted as this project's default
63
+ $ export GEMCHAT_API_KEY=… # or a `credentials` file in the config home
64
+ $ gemchat query "configure sidekiq retries"
65
+ ```
66
+
67
+ `init` does not prompt for a key: it edits the Gemfile, and the key is read from
68
+ the environment or from `credentials` under the config home. That is deliberate
69
+ — a prompt here would put a secret on a tty in a command that otherwise only
70
+ reads a Gemfile.
71
+
72
+ Three cost tiers on the local path, in the order they arrive:
73
+
74
+ | Moment | What works | Cost |
75
+ |---|---|---|
76
+ | after `gemchat index` | `search` — BM25, exact version, fully offline | seconds + one ri pass |
77
+ | after `gemchat embed` | `vsearch`, and `query`'s vector arm | 146MB, once |
78
+ | always | `query` hybrid with RRF | both arms, and it degrades to BM25 alone if the model is absent |
79
+
80
+ Note the second row. `vsearch` and `query` do **not** download the model; only
81
+ `gemchat embed` does. A search command that silently pulls 146MB is a different
82
+ product from one that does not, and it would be a nasty surprise in CI. When the
83
+ model is absent they say so and answer from BM25 — or, for `vsearch`, exit 1
84
+ rather than answer a different question than the one asked.
85
+
86
+ **The first useful result costs nothing** — no account, no API key, no model
87
+ download.
88
+
89
+ ## 3. How we got here
90
+
91
+ | Thread | Outcome |
92
+ |---|---|
93
+ | Chat markdown table styling | `b3e1d51` — unrelated, done |
94
+ | rubydex native ext unusable on Alpine/musl | `3866330` lazy-load → `372e57e` compile from Rust source; `available? => true` in production |
95
+ | Agent Skills spike (Hyperdrive interop) | `b9d1491` grounded `SKILL.md` generation, spec-validated |
96
+ | **This plan** | Local bundle index, with a hosted tier for users who want neither |
97
+
98
+ Two findings shape the architecture, and neither came from the rubydex work
99
+ itself. Reading `SearchService` showed it is unportable (pgvector + PostgreSQL
100
+ FTS + HNSW), which is *why* a separate local store exists. The hosted app's
101
+ `AGENTS.md` records `RERANKING_ENABLED=false` as the production default, which
102
+ is why no reranker is planned here. What the rubydex work did contribute is the
103
+ lesson that a Ruby dev tool has to install cleanly on musl, which shaped the
104
+ dependency set (§4.4) and the embedder choice (§4.3).
105
+
106
+ The Agent Skills work is reused directly: §16 ships the same
107
+ Hyperdrive-compatible `SKILL.md` format that `b9d1491` validated.
108
+
109
+ ## 4. Decisions
110
+
111
+ | Decision | Choice | Rationale |
112
+ |---|---|---|
113
+ | Language / packaging | Ruby gem, `gemchat` | Reuses RDoc/ri/Ruby-aware parsing; no Bun/GGUF toolchain |
114
+ | Embedding default | `rllama` + downloaded GGUF | Works on a bare machine after `gem install` |
115
+ | ri content | ri is the corpus, always indexed | 98.5% of the hosted corpus; carries the API surface |
116
+ | Prose content | READMEs and `guides/`, indexed under `source_type = "readme"` | ri answers "what does this take"; a README answers "how is this used", which is the question ri is worst at. 1,301 of 86,945 chunks, but that ratio counts chunks, not usefulness |
117
+ | **ri acquisition** | **Generate it ourselves — bundler does not** | Measured: `bundle install` produces no ri (§4.1) |
118
+ | ri pollution | Document `lib/`+`ext/` only, drop `--all` | Structural filter, not a heuristic |
119
+ | Reranker | Excluded from v1 | Hosted app ships with it disabled; eval-gated follow-up |
120
+ | Index scope | Direct deps by default, `--all` opt-in | ~14.2k chunks / ~108MB vs ~105k / ~308MB (§6) |
121
+ | Prism | Unconditional `prism` gem dependency | Identical behaviour on Ruby 3.2 and 3.3+ |
122
+ | **ActiveSupport** | **Not a dependency** | Measured: 13 extra gems (`concurrent-ruby`, `i18n`, `tzinfo`, `drb`, `connection_pool`, …) for a CLI that uses none of them. The single thing wanted — boolean env parsing — is `Gemchat::Env` (~15 lines) |
123
+ | rubydex | **Not a dependency** | Cost is musl-only, not universal; and the join is worth less here because we index the ri and Prism halves independently |
124
+ | Core/stdlib ri | Generate from local `rubylibdir` — **not** `rdoc-data` | See §8.4 |
125
+ | **Hosted tier** | **In v1, explicit opt-in only** | Low-disk / no-ri / no-model users still get value |
126
+ | Degradation | Diagnose and stop; never auto-switch | A local privacy tool must not silently phone home |
127
+ | Staleness trigger | Bundler plugin on `add`/`install`/`update`, trust-gated | All three change the lockfile; sha256-gating handles them uniformly |
128
+ | Setup verb | Separate `gemchat init` | `index` must stay non-interactive and never mutate the Gemfile (CI safety) |
129
+ | Store drift | Refuse on `schema_version`, tolerate `gemchat_version` | A newer layout would return wrong answers silently; a point release must not strand an index |
130
+ | Agent discoverability | Ship `skills/gemchat/SKILL.md` in the gem | An agent cannot use a CLI it does not know exists |
131
+ | Install path | **Both** Gemfile (dev group) and global | Reproducibility for projects, lightness for one-offs |
132
+
133
+ ### 4.1 Measured correction: `bundle install` does not generate ri
134
+
135
+ The original plan assumed "ri documentation is usually installed in the user's
136
+ system unless `--no-document` is set". **That is true for `gem install` and
137
+ false for `bundle install`.** Verified with isolated installs:
138
+
139
+ | Method | ri generated? |
140
+ |---|---|
141
+ | `gem install zeitwerk` (no flags at all) | **yes** |
142
+ | `gem install zeitwerk --document=ri` | yes |
143
+ | `bundle install` via isolated `BUNDLE_PATH` | **no** |
144
+
145
+ Corroborating evidence on a dev machine: 3 of 411 installed gems have ri, and
146
+ all three (`bundler`, `rubygems`, `rubygems-update`) ship with Ruby itself. No
147
+ `~/.gemrc` and no `.bundle/config` — this is not a user setting. Bundler skips
148
+ documentation generation deliberately, because it is slow across hundreds of
149
+ gems.
150
+
151
+ **Consequence:** for bundler users — essentially the whole target audience —
152
+ Tier 0 (§8.1) hits ~0%. ri generation is therefore the **main path, not a
153
+ fallback**, and the resumable generation subsystem (§13) is a core phase rather
154
+ than a conditional one. An earlier draft of this plan got this backwards.
155
+
156
+ The user's original counterpoint still holds in narrower form: `gem rdoc --all
157
+ --ri --no-rdoc` *does* generate ri for every installed gem, so tier 0 is worth
158
+ keeping as a free fast path and the remedy is worth printing.
159
+
160
+ ### 4.2 `rdoc-data` must not be a dependency
161
+
162
+ Looks like the obvious route to core/stdlib ri, but the gem is a dead end:
163
+
164
+ - latest version `4.1.0`, published **2015-12-29**; nothing since
165
+ - runtime dependency `rdoc ~> 4.0`, while current RDoc is `8.0.0` — adding it
166
+ would pin a decade-old RDoc onto a Ruby 3.2+/4.x install
167
+ - payload covers Ruby 1.8–2.3 only, and its own description says that data
168
+ "are not needed for typical installs of C Ruby"
169
+ - `require "rdoc/data"` fails on this machine's Ruby 4.0.7
170
+
171
+ Instead: `RbConfig::CONFIG["rubylibdir"]` already contains **727 stdlib `.rb`
172
+ files** locally, so core/stdlib ri is generated offline with the same RDoc
173
+ recipe. No dependency, no download, no version-manager intervention. See §8.4.
174
+
175
+ ### 4.3 Embedding default flipped from Ollama to rllama
176
+
177
+ The original reasoning was wrong: it optimised for *this* machine, where
178
+ `bin/pull_models` already pulled `nomic-embed-text`. A user who just ran
179
+ `gem install gemchat` has Ruby, their bundle, and nothing else. Defaulting to
180
+ Ollama means a second install, a running daemon, and a port. Sizes are
181
+ near-identical (`nomic-embed-text` 274MB, `embeddinggemma-300M-Q8_0` ~300MB) so
182
+ the differentiator is one step versus two, not bytes.
183
+
184
+ Superseded in one respect: the `rllama` default is now
185
+ `nomic-embed-text-v1.5` rather than `embeddinggemma-300M-Q8_0`, because the
186
+ latter is gated on HuggingFace. See §10.1. The one-step-versus-two argument
187
+ above still holds.
188
+
189
+ ### 4.4 Why rubydex is not a dependency
190
+
191
+ rubydex resolves one specific mismatch: RDoc reports *mixin-flattened* owners
192
+ (`Sidekiq::Job.perform_async`, via `extend ClassMethods`) while Prism reports
193
+ the *lexical* owner (`Sidekiq::Job::ClassMethods`). Measured in the Rails app
194
+ on sidekiq 8.1.7: **+8 of 771 chunks, ~1.1 percentage points** (88.7% → 89.8%).
195
+
196
+ The cost side is weaker than it first appears, and the value side is lower here
197
+ than in the Rails app:
198
+
199
+ - **Distribution cost is musl-specific, not universal.** rubydex publishes
200
+ precompiled builds for `x86_64-linux`, `aarch64-linux`, `x86_64-darwin`,
201
+ `arm64-darwin`, and `x64-mingw-ucrt`, so `gem install rubydex` is a plain
202
+ download on macOS, glibc Linux, and Windows x64 — no Rust, no build. The gap
203
+ is musl: no `-linux-musl` variant is published, so Alpine users would need a
204
+ Rust toolchain — and an unbuildable dependency fails the *entire*
205
+ `bundle install`, not just gemchat. Real severity, but for a minority.
206
+ - **The value is lower here than in the Rails app.** We generate the ri store
207
+ ourselves *and* run Prism over the same source, so both artifacts are ours
208
+ and both are indexed. The Rails app joined them to enrich a third-party
209
+ corpus; here search can surface either half independently. The 1.1% is an
210
+ upper bound, not a floor.
211
+ - **It would not fix the remaining ~10%**, which is mostly
212
+ `attr_reader`/`attr_accessor` (a `CallNode`, not a `DefNode`) and dynamic
213
+ definitions.
214
+ - **The loss is nearly harmless.** Reporting
215
+ `Sidekiq::Job::ClassMethods#perform_async` rather than
216
+ `Sidekiq::Job.perform_async` is arguably *more* precise — it tells the agent
217
+ the method comes from a mixin.
218
+
219
+ Prism alone is pure Ruby, dependency-free, and the sole source of file:line and
220
+ signature metadata in this design.
221
+
222
+ The decision rests on the value argument, not the cost one. An earlier draft of
223
+ this section claimed rubydex would impose a Rust build on every end user; that
224
+ was wrong, and generalised from the Rails app's own Alpine deploy onto a
225
+ population that mostly does not resemble it.
226
+
227
+ ### 4.5 Four defects the test suite found
228
+
229
+ Worth recording because all four passed a manual end-to-end check and a clean
230
+ `standardrb` run. A smoke test proves the happy path works; it does not prove
231
+ the *state* is correct, and these were all state bugs.
232
+
233
+ 1. **`Gemchat.root` was memoised.** It derives from `GEMCHAT_HOME`, and a cached
234
+ environment value means the first reader wins for the life of the process. Any
235
+ later `GEMCHAT_HOME` silently did nothing, and the test suite was writing into
236
+ the developer's real `~/.gemchat`. The join is cheap; do not cache it. This
237
+ class of bug is why the first suite run failed only in one order.
238
+ 2. **`Store#search` returned `{**row, score: …}`.** With `results_as_hash` set,
239
+ every row already carries integer keys, so the splat leaked `0, 1, 2…` into
240
+ results and made `score` a Symbol while every other column was a String.
241
+ Fixed by selecting an explicit column list and building the hash by hand.
242
+ 3. **Blank chunk fields were indexed as dead tokens.** `searchable_text` emitted
243
+ `method_name: ` for class-level chunks. Class chunks outnumber method chunks,
244
+ so this padded a large share of the corpus. Now dropped.
245
+ 4. **`index` reported work it had not done.** A cached ri store plus an unchanged
246
+ content digest printed `generated … 0 chunks`, which reads like success and
247
+ like corruption equally. `Store#replace_gem` now returns an `Outcome`
248
+ (`:indexed` / `:updated` / `:unchanged`) and the CLI prints `unchanged`
249
+ instead of a chunk count. A no-op and a run that lost its content are very
250
+ different events and must not look the same.
251
+
252
+ The generalisable lesson: `:unchanged` is a *status*, and
253
+ "wrote 0 rows" is not a substitute for one. Minitest 6 removing `Object#stub`
254
+ and `minitest/mock` also forced a small local `with_stub` helper — check the
255
+ loaded minitest version before assuming mocking exists.
256
+
257
+ ## 5. What already exists (reuse, don't rebuild)
258
+
259
+ Read-only references in `gemchat_app`:
260
+
261
+ | Asset | Location | Reuse as |
262
+ |---|---|---|
263
+ | `read_ri_from_path` | `app/services/gem_fetcher_service.rb:176-183` | Port the `RDoc::RI::Driver` reader (~50 lines, Rails-free) |
264
+ | ri generation recipe | `app/services/gem_fetcher_service.rb:339-369` | Port verbatim, including both gotchas |
265
+ | ri tuning notes | `docs/ri-doc-generation.md` | Source of the `--all` decision |
266
+ | Prism symbol pass | `RubySourceDefinitionService` | Port for `class_name`/`method_name`/signature |
267
+ | `find_ri_path` | `app/services/gem_fetcher_service.rb:259-266` | **Fix, don't copy** — see §8.1 |
268
+ | Hybrid weights | 0.7 vector / 0.3 FTS, `RRF_K = 60.0` | Identical to qmd's RRF k |
269
+ | Source-prior rebalancing | `RAG_SOURCE_PRIORS_ENABLED` | Reuse the idea for ri/readme balancing |
270
+ | Core/stdlib generation | `app/jobs/ruby_core_stdlib_index_job.rb:92-120` | Reference for the tier-3 path |
271
+ | Skill artifact format | `b9d1491` in this thread | Reuse the `SKILL.md` shape for §16 |
272
+ | Plugin trust pattern | qmd `.qmd/index.yml` gating | Model for §15 + §17 |
273
+ | **Hosted MCP tools** | `query_docs`, `answer_extractive`, `get_index_status`, `read_changelog` | The backend for §12 |
274
+
275
+ **Not reusable:** `SearchService` (806 lines) — pgvector `<=>`,
276
+ `websearch_to_tsquery`, generated `tsvector`, `pg_trgm`, HNSW
277
+ `vector_cosine_ops`. That constraint is precisely *why* a local store exists.
278
+
279
+ ## 6. Research findings (measured, not assumed)
280
+
281
+ Measured against the 124-gem local corpus:
282
+
283
+ | Metric | Value |
284
+ |---|---|
285
+ | ri chunks | **85,644** |
286
+ | readme chunks | 1,301 |
287
+ | ri share | **98.5%** |
288
+ | per-gem chunk counts | median 176 · p90 1,482 · max 12,663 |
289
+ | 150-gem bundle (est.) | ~105,000 chunks · 308MB of vectors |
290
+ | ~25 direct deps (**superseded** — see below) | ~4,400 chunks · ~~14MB~~ of vectors |
291
+
292
+ **Conclusion:** ri is the corpus, not an optional extra. The hosted index
293
+ already answered this empirically. The cost objection dissolves at direct-deps
294
+ scope. The 308MB figure applies only to `--all`.
295
+
296
+ Install-method measurement (§4.1):
297
+
298
+ | Finding | Value |
299
+ |---|---|
300
+ | `gem install` generates ri | yes, no flags needed |
301
+ | `bundle install` generates ri | **no** |
302
+ | Local machine: gems with ri | 3 of 411 (all Ruby-shipped) |
303
+ | `~/.gemrc` / `.bundle/config` | absent — not user-configured |
304
+
305
+ ri generation cost, measured across the **49 resolvable direct dependencies of a
306
+ real Rails app** (`gemchat_app`, 52 direct deps, 223 total specs), public API
307
+ only, `lib/`+`ext/`:
308
+
309
+ | Metric | Value |
310
+ |---|---|
311
+ | generation time | **9.9s total · 0.20s/gem mean · 0.09s/gem median** |
312
+ | chunks produced | **14,229 · 290/gem mean · 101/gem median** |
313
+ | ri store on disk | **66 MB** |
314
+ | vectors @768d f32 | **42 MB** |
315
+ | 4-core wall clock | **~2s** |
316
+ | skipped (no `lib`/`ext`) | `rails` (meta-gem), `rubocop-rails-omakase`, `tzinfo-data` |
317
+ | largest single gems | brakeman 1,597 · yard 1,473 · rouge 1,229 · simplecov 1,214 · capybara 1,019 chunks |
318
+
319
+ This **resolves the plan's largest stated risk**. ri generation is seconds, not
320
+ minutes, so the resumable design (§13) is good hygiene rather than a
321
+ load-bearing necessity. Chunk count is ~3x the plan's earlier 4,400 estimate,
322
+ so the real local footprint is ~108MB (ri + vectors) rather than ~14MB — still
323
+ comfortable, and the 308MB figure remains the `--all`-on-everything case.
324
+
325
+ `--all` cost, measured on a separate 9-gem sample: adds **3–46%** more chunks
326
+ (roo +46%, pg_query +23%, selenium-webdriver +20%, prism +3%) and roughly
327
+ proportional generation time. Dropping it (§9) is worth keeping, though the win
328
+ is smaller than the noise argument alone suggests.
329
+
330
+ RDoc and SQLite behaviour found empirically while building the spike:
331
+
332
+ - **RDoc refuses to write into an output directory that already exists.** Do not
333
+ pre-create the target; create only its parent. This constrains §13's atomic
334
+ rename — RDoc creates the directory, we then move it into place.
335
+ - **`RDoc::RI::Driver` takes `--doc-dir`**, not `--ri`/`--op`.
336
+ - `Dir.chdir(gem_dir)` with relative source dirs is required (§8.3); absolute
337
+ `--op` is fine alongside it.
338
+ - **`benchmark` is not a default gem in Ruby 4.0** — use
339
+ `Process.clock_gettime(Process::CLOCK_MONOTONIC)` for all timing.
340
+ - Gems with no `lib/` and no `ext/` (meta-gems like `rails`) must be skipped
341
+ rather than passed to RDoc, which errors.
342
+
343
+ Environment verification:
344
+
345
+ - `sqlite3` 2.9.6 present, **FTS5 working**, precompiled natives for
346
+ `x86_64-linux-musl`, `x86_64-linux-gnu`, `x86_64-darwin` (confirm
347
+ `aarch64-linux` during the spike)
348
+ - `sqlite-vec` **not** present, and not needed (§11.1)
349
+ - `rllama` 1.2.0 exists; single runtime dependency `ffi >= 1.0`
350
+ - `prism` 1.9.0 exists; no runtime dependencies
351
+ - `rdoc-data` 4.1.0 is legacy (2015) and constrains `rdoc ~> 4.0` (§4.2)
352
+ - `rubydex` 0.4.1 publishes `ruby`, `x86_64-linux`, `aarch64-linux`,
353
+ `x86_64-darwin`, `arm64-darwin`, `x64-mingw-ucrt` — **no `-linux-musl`
354
+ variant**, which is part of why it is excluded (§4.4). For contrast `sqlite3`
355
+ publishes the full matrix including `x86_64-linux-musl` and
356
+ `aarch64-linux-musl`, so Alpine coverage is a packaging choice some gems make
357
+ and rubydex has not.
358
+ - stdlib doc dir **absent** on this machine (mise Ruby without ri data)
359
+ - Bundler 4.0.21 supports the `plugin` DSL and defines `Bundler::Plugin`
360
+
361
+ ## 7. Architecture
362
+
363
+ ```
364
+ ~/.gemchat/index.sqlite # global store: every gem ever indexed, keyed name+version
365
+ ~/.gemchat/models/ # downloaded GGUF embedding models
366
+ ~/.gemchat/ri/ # ri stores WE generate (the main path, §8.3)
367
+ ~/.gemchat/credentials # hosted MCP API key, chmod 600
368
+ ~/.gemchat/trusted.json # trust approvals for project-local config and plugin
369
+ <project>/.gemchat.yml # lockfile sha256 + scope (trust-gated, §17)
370
+ ```
371
+
372
+ **Global store + per-project manifest** mirrors qmd's collection model and
373
+ avoids re-embedding `rails` in every project. Config is XDG-aware
374
+ (`$XDG_CONFIG_HOME` honoured, plus `GEMCHAT_CONFIG_DIR`).
375
+
376
+ Two backends share one CLI: `local` (default, §8–§11) and `hosted`
377
+ (gemchat.org, §12). The backend is a stored setting, never inferred.
378
+
379
+ ### 7.1 Store identity
380
+
381
+ One `meta` table records everything that would make a store unreadable or
382
+ silently wrong:
383
+
384
+ | Key | Guards against |
385
+ |---|---|
386
+ | `schema_version` | store written by a newer/older gemchat |
387
+ | `gemchat_version` | global install drifting from the bundled one (§14) |
388
+ | `embedding_model` | mixing vectors from different models |
389
+ | `embedding_dimensions` | a model swap with a different output width |
390
+ | `backend` | local vs hosted confusion in output |
391
+
392
+ On mismatch the store **refuses to open** and prints the exact remediation
393
+ (`gemchat reindex`, or the specific field that changed). This generalises the
394
+ "never mix embedding models" rule the hosted app already enforces, and is the
395
+ reason dual install paths (§14) are safe.
396
+
397
+ ## 8. ri strategy — four tiers
398
+
399
+ Because `bundle install` produces no ri (§4.1), tier 0 is a rare fast path and
400
+ tier 2 is the main event. The ordering below reflects that honestly.
401
+
402
+ ### 8.1 Tier 0 — opportunistic read (rare fast path)
403
+
404
+ Worth keeping: `gem install` users, anyone who has run `gem rdoc`, and gems
405
+ that ship prebuilt docs all hit this instantly for free.
406
+
407
+ Resolve the **exact locked version**. The hosted app's lookup must *not* be
408
+ copied verbatim:
409
+
410
+ ```ruby
411
+ # WRONG for this tool — gemchat_app:259-266 globs any version
412
+ Dir.glob(File.join(gem_path, "doc", "#{name}-*", "ri"))
413
+
414
+ # RIGHT — exact locked version
415
+ Dir.glob(File.join(gem_path, "doc", "#{name}-#{version}", "ri"))
416
+ ```
417
+
418
+ Search `Gem.path` **plus** `Bundler.bundle_path` (bundler's install path is
419
+ frequently absent from `Gem.path`). Fast path first:
420
+ `Gem::Specification.find_by_name(name, version)&.doc_dir`.
421
+
422
+ ### 8.2 Tier 1 — diagnose, and print the remedy
423
+
424
+ Expected to fire for most bundler users, so it must read as normal rather than
425
+ as breakage. State the cause plainly:
426
+
427
+ > ri: 0/27 found — Bundler installs without documentation by default.
428
+
429
+ Then the remedy, which genuinely works and is cheap to run once:
430
+
431
+ ```bash
432
+ gem rdoc --all --ri --no-rdoc # generate ri for every installed gem
433
+ ```
434
+
435
+ Offer both routes — run that, or let `gemchat index` generate ri per-gem
436
+ itself (§8.3), or switch to the hosted backend (§12). The user picks.
437
+
438
+ ### 8.3 Tier 2 — generate per-gem into a private cache (main path)
439
+
440
+ ```bash
441
+ rdoc --all --ri --no-rdoc --op ~/.gemchat/ri/<name>-<version>
442
+ ```
443
+
444
+ Generates into **our** cache. Never writes into the user's gem directories and
445
+ never installs. This is the expected path for a fresh project, and its
446
+ cost profile drives the resumable design in §13.
447
+
448
+ Two gotchas ported verbatim from `gem_fetcher_service.rb`:
449
+
450
+ 1. **`Dir.chdir(source_dir)` with relative source dirs.** RDoc 8 computes
451
+ `full_names` relative to cwd, then chdirs into `--op` when saving. Absolute
452
+ paths from a different cwd yield escaping relative paths
453
+ (`./../../../../var/...`) and writes land *outside* the output directory.
454
+ 2. **`rescue SystemStackError` and delete the partial output.** RDoc recurses on
455
+ symlink cycles and deeply nested modules. `SystemStackError < Exception`, so
456
+ a bare `rescue => e` misses it, and a truncated ri store must never be read.
457
+
458
+ Also applied here (from §9): document only `lib/` and `ext/`, and **omit
459
+ `--all`** — the latter halves the noise and is the single most important lever
460
+ on output size and speed.
461
+
462
+ ### 8.4 Tier 3 — core and stdlib
463
+
464
+ Do **not** depend on the `rdoc-data` gem (§4.2). Instead:
465
+
466
+ 1. If the system store exists, read it
467
+ (`RbConfig::CONFIG["rubylibdir"]/../doc`).
468
+ 2. Otherwise generate from `RbConfig::CONFIG["rubylibdir"]` — 727 stdlib `.rb`
469
+ files ship with Ruby, so `rdoc --ri --op <cache> <rubylibdir>` is fully
470
+ offline. Cache under a synthetic `ruby-core-stdlib` key.
471
+ 3. Surface the version-manager remedy for users who prefer a prebuilt store
472
+ (rbenv `rdoc` package, mise, brew) rather than generating.
473
+
474
+ ## 9. Fixing ri pollution without excluding ri
475
+
476
+ Two flags, both already established in `gemchat_app`:
477
+
478
+ 1. **Document only `lib/` and `ext/`.** Structurally excludes Rakefiles,
479
+ `spec/`, `.github/`, `bin/` — the "Sidekiq's own HERB linting config"
480
+ failure mode observed in the agent-skills spike.
481
+ 2. **Drop `--all`.** The tuning table in `docs/ri-doc-generation.md` names
482
+ `--all` (private/protected methods) as the noise source and states that
483
+ removing it limits output to the public API.
484
+
485
+ At query time, source priors bias toward README/guides for *how-to* questions
486
+ while preserving ri for *API signature* questions — the same rebalancing the
487
+ hosted app already performs.
488
+
489
+ ## 10. Embedding
490
+
491
+ | Provider | Model | User cost | Status |
492
+ |---|---|---|---|
493
+ | `rllama` | `nomic-embed-text-v1.5` Q8_0 (146MB) | auto-download | **default** |
494
+ | `ollama` | `nomic-embed-text` (274MB) | must install and run Ollama | opt-in |
495
+ | `openai` | `text-embedding-3-small` | key + network + per-query cost | opt-in |
496
+
497
+ `rllama` ships prebuilt llama.cpp binaries in the gem; its only dependency is
498
+ `ffi`. Users compile nothing.
499
+
500
+ **`gemchat search` must work with zero models downloaded.** This is qmd's
501
+ behaviour and the reason it feels instant: BM25 works immediately after
502
+ `index`, and the model is only pulled when semantic recall is wanted — either
503
+ via the explicit `gemchat embed`, or lazily on the first `vsearch`/`query`.
504
+ First vector use downloads with progress, resume, and checksum verification.
505
+
506
+ Model identity and dimensionality live in the store's `meta` table (§7.1).
507
+ **Refuse to mix** and require an explicit re-embed.
508
+
509
+ ### 10.1 The default model changed: embeddinggemma is gated
510
+
511
+ The table above originally named `embeddinggemma-300M-Q8_0`. That model is
512
+ **`gated: manual` on HuggingFace**: fetching it needs an account, a licence
513
+ click, and a personal token. An unattended `gemchat embed` cannot do any of
514
+ that, so it was not a default, it was a wall. `nomic-embed-text-v1.5` is
515
+ ungated, Apache-2.0, 768-dimensional (matching `gemchat_app`'s vectors, so the
516
+ two agree on what good retrieval looks like), and 146MB rather than 300MB.
517
+
518
+ The digest is pinned rather than described, so "checksum verification" means
519
+ something specific:
520
+
521
+ ```
522
+ nomic-ai/nomic-embed-text-v1.5-GGUF / nomic-embed-text-v1.5.Q8_0.gguf
523
+ 146,146,432 bytes
524
+ sha256 3e24342164b3d94991ba9692fdc0dd08e3fd7362e0aacc396a9a5c54a544c3b7
525
+ ```
526
+
527
+ `Models.installed?` compares **size, not digest**, because SHA-256 over 146MB on
528
+ every `gemchat status` and every `gemchat search` is not affordable. The digest
529
+ is enforced in `download` before the file is moved into place, so a wrong file
530
+ cannot be installed by that path.
531
+
532
+ ### 10.2 Prefixes are part of the model identity
533
+
534
+ Nomic models are trained with task prefixes. Measured on the rake corpus, the
535
+ correct hit's similarity for the query *"stop the process when the OS sends a
536
+ signal"* rose from **0.515 to 0.580** once `search_document: ` and
537
+ `search_query: ` were applied to the two sides.
538
+
539
+ The prefixes are declared per model in the registry rather than hardcoded, and
540
+ `+prefixed` is part of the stored identity string — because changing the prefixes
541
+ changes what the vectors *mean*, so an existing index has to be rebuilt rather
542
+ than extended. Omitting them is not a crash, it is a silent quality loss, which
543
+ is the harder failure to notice.
544
+
545
+ ### 10.3 The native embedder has to be supervised, and has a size cliff
546
+
547
+ rllama 1.2.0's prebuilt llama.cpp **aborts the process** on this machine rather
548
+ than returning an error. Measured, not assumed: the same command failed four
549
+ times in a row and later passed 20 of 20 with no configuration change, and
550
+ `GGML_METAL_DISABLE=1` made no reliable difference. A `SIGILL` cannot be
551
+ rescued — it descends from nothing `StandardError` catches — so in-process there
552
+ is no way to detect it, retry it, or degrade. Embedding therefore runs in a
553
+ **supervised child process**, retried once with Metal disabled.
554
+
555
+ There is also a size cliff, and its exact position could not be pinned:
556
+
557
+ | Input | Result |
558
+ |---|---|
559
+ | 128 real chunks, one batch | pass |
560
+ | 216 real chunks, ~90KB, one batch | crash |
561
+ | 160 real chunks, 44,624 bytes | pass |
562
+ | 400 short synthetic chunks, 23,200 bytes | pass |
563
+ | one real chunk at 1634 bytes | pass |
564
+ | one real chunk at 1665 bytes | **crash** |
565
+
566
+ It is not a byte count — 2500 bytes of repeated `"word "` is fine. It is not
567
+ ASCII-related — stripping to pure ASCII does not help. It is not cleanly a token
568
+ count either. The cause was not isolated, so the defence is two-layered rather
569
+ than tuned: a per-item ceiling of **1200 bytes** (margin under the observed
570
+ cliff) and a per-batch budget of **16KB**, with any batch that still fails
571
+ **halved and retried**, down to a single chunk.
572
+
573
+ The cost is real and stated: a long prose chunk is represented by its opening,
574
+ and its tail stays reachable through BM25 but not through `vsearch`. Prose is
575
+ already split at the chunker, so this affects the largest sections only. On the
576
+ 236-chunk rake index the result is **236/236 embedded in ~6s**, from a clean
577
+ auto-download, with nothing left pending.
578
+
579
+ ### 10.4 Partial success is a first-class outcome
580
+
581
+ A batch that cannot be embedded is left **pending**, not skipped and not faked.
582
+ `Supervised#embed` returns `[index, vector]` pairs rather than a bare vector
583
+ list precisely so a failed chunk cannot be written with a neighbour's vector or
584
+ dropped from the count and silently misalign every row after it. A run where
585
+ *nothing* embeds raises rather than reporting a successful no-op.
586
+
587
+ `embed --force` discards the vectors and re-embeds everything. This is not
588
+ cosmetic: `unembedded_chunks` only returns rows where the vector `IS NULL`, so
589
+ without the reset a model change cannot be reconciled by re-running `embed` at
590
+ all, and the mismatch error would be telling the user to run a command that does
591
+ nothing.
592
+
593
+ ## 11. Search (local backend)
594
+
595
+ | Command | Engine | Requires models |
596
+ |---|---|---|
597
+ | `gemchat search` | FTS5 BM25 | no |
598
+ | `gemchat vsearch` | vector cosine | yes |
599
+ | `gemchat query` | hybrid: typed expansion → BM25 + vector → RRF k=60 | yes |
600
+
601
+ ### 11.1 No ANN index needed — but two claims here were wrong, and are corrected
602
+
603
+ **The first version of this section claimed the full scan "ranks in
604
+ milliseconds".** Measured on a real 22,141-chunk index (120 locally installed
605
+ gems, 33 with ri, 21,253 ri chunks) it takes **915ms per vector query**, not
606
+ milliseconds:
607
+
608
+ | Query | Full scan | FTS-narrowed to 200 | to 500 |
609
+ |---|---|---|---|
610
+ | how do I parse a ruby file into an abstract syntax tree | 915.3ms | 9.2ms | 22.2ms |
611
+ | what raises a cop offense for a long line | 913.1ms | 9.2ms | 22.3ms |
612
+
613
+ 915ms is the real number, and it is because a UDF that unpacks two 3072-byte
614
+ blobs into Ruby arrays is called once per row: ~34M array elements allocated per
615
+ query. The 236-chunk index every other measurement in this project used reports
616
+ ~5ms, which is exactly how the wrong claim survived.
617
+
618
+ **The second claim was that two-stage retrieval "cannot be expressed in SQL".**
619
+ That is false, and the correction is the useful part. A candidate CTE joined to
620
+ `chunks` limits which rows the UDF is called for, so "the top N by FTS, then rank
621
+ those" is one statement:
622
+
623
+ ```sql
624
+ WITH candidates AS (
625
+ SELECT rowid FROM chunks_fts WHERE chunks_fts MATCH ?
626
+ ORDER BY bm25(chunks_fts) LIMIT 500
627
+ )
628
+ SELECT c.title, -dot(c.embedding, ?) AS rank
629
+ FROM candidates JOIN chunks c ON c.id = candidates.rowid
630
+ WHERE c.embedding IS NOT NULL
631
+ ORDER BY rank LIMIT 10
632
+ ```
633
+
634
+ 40x faster, and **not free**. Measured against the full scan on eight questions:
635
+
636
+ | Pool | recall@200 of the full-scan top-200 | top-1 identical |
637
+ |---|---|---|
638
+ | 200 | 33.5% | 7/8 |
639
+ | 500 | 48.8% | **8/8** |
640
+ | 1000 | 59.9% | 8/8 |
641
+
642
+ So the answer the user sees is preserved; the tail is not. And there is a cliff
643
+ the numbers do not show: the candidate generator is lexical, so a paraphrase
644
+ sharing no vocabulary with its answer produces few or no candidates and returns
645
+ nothing, where the full scan would have found it. The OR mode is a *terrible*
646
+ ranker (§11.3) and may still be an adequate *recall filter*, because the
647
+ question's rare terms still have to match. That is the trade worth making
648
+ deliberately rather than by default.
649
+
650
+ **Not implemented.** It is recorded here as a measured, specified optimisation
651
+ with its cost stated, rather than slipped in on the strength of a 40x number. The
652
+ `sqlite-vec` conclusion still stands — unnecessary at bundle scale — but for the
653
+ right reason: the bottleneck is a Ruby-level UDF, not the absence of an ANN index.
654
+
655
+ Two details the full scan cost anyway, both silent:
656
+
657
+ - **sqlite3 2.x ignores the block's return value** and reads `FunctionProxy#result`
658
+ instead. Returning the float makes every rank `NULL`, which sorts as `0.0` and
659
+ returns an arbitrary slice of the corpus as if it were ranked output.
660
+ - **A mis-sized query** makes `dot` return `NULL` for every row. `vector_search`
661
+ refuses up front rather than returning junk.
662
+
663
+ Vectors are stored as packed float32 BLOBs, 3072 bytes for 768 dimensions.
664
+
665
+ ### 11.2 Ranking
666
+
667
+ RRF with `k = 60`, matching both `SearchService`'s `RRF_K` and qmd. Typed query
668
+ expansion copied from qmd — `lex` routes to BM25, `vec` and `hyde` route to
669
+ vector search, routed exclusively rather than sent to both.
670
+
671
+ **No reranker in v1.** Corroborated by the hosted app's
672
+ `RERANKING_ENABLED=false` production default. Revisit only if eval shows top-1
673
+ errors concentrated in ri-heavy queries.
674
+
675
+ ### 11.3 There is no usable keyword arm yet, and the fusion is built around that
676
+
677
+ The fusion is implemented and correct, but it is only ever as good as the weaker
678
+ arm, so the weaker arm had to be measured before the feature was called done.
679
+
680
+ **`search` is an exact-phrase matcher wearing BM25's name.** `fts_query` quotes
681
+ each whitespace-separated term — which is load-bearing, because an unquoted
682
+ `Puma::Server` parses as an FTS5 column filter and fails with "no such column",
683
+ a systematic collision with Ruby naming. But joining those quoted terms with a
684
+ *space* is FTS5's AND-of-a-phrase, and that was a side effect of the escaping
685
+ rather than a decision. Measured on a ~22k-chunk corpus (a synthetic lockfile
686
+ built from locally installed gems; see §11.4 for the exact configuration):
687
+
688
+ | Query | `:phrase` (shipped) | `:any` (OR'd) |
689
+ |---|---|---|
690
+ | any natural-language question | **0 rows** | 5 rows |
691
+ | `Rake::Task#enhance` | 1 row | 1 row |
692
+
693
+ So the obvious fix — OR the terms, which is what a keyword search means —
694
+ measures **worse, not better**. On five known-answer queries, each with a
695
+ definitely-existing answer in the corpus, the OR'd arm located the right chunk in
696
+ **1 of 5**:
697
+
698
+ | Query | Expected | OR'd top-1 |
699
+ |---|---|---|
700
+ | write a rubocop cop that checks line length | `RuboCop::Cop::Layout::LineLength` | `RuboCop::Cop::MinBodyLength` |
701
+ | generate documentation from source comments | `RDoc::Task` | `Parser::Source::Comment#document?` |
702
+ | run tests in parallel with processes | `Minitest::Parallel::Executor` | `Minitest::Parallel::Executor` ✓ |
703
+ | parse ruby source into an abstract syntax tree | `Parser::CurrentRuby` | `prism › docs › Ruby API > API` |
704
+ | configure the default rake task | `Rake::Task.define_task` | — |
705
+
706
+ The mechanism is length normalisation: a rare query term inside a short ri
707
+ signature scores enormously under BM25, so `MinBodyLength` beats a long prose
708
+ chunk that actually discusses line length. Corpus size does not save it — the
709
+ 22k-chunk run above has perfectly good IDF statistics.
710
+
711
+ **Therefore `:any` is not the default, and this matters more for fusion than for
712
+ search alone.** RRF trusts rank position, so a confidently wrong arm is not
713
+ ignored — it is *promoted*. Feeding OR'd BM25 to the fusion changed the top-1 on
714
+ 4 of 5 rake questions and made it worse on most, e.g. `syntax highlighting in
715
+ the console` went from `Rakefile Format > Comments` to `Why rake? · 4/4`.
716
+
717
+ What ships instead: `:phrase` everywhere, so the lexical arm contributes only
718
+ when it has an exact hit — which *should* rank first anyway. Measured effect on
719
+ the rake index:
720
+
721
+ - Six natural-language questions: fused output **agrees with the vector arm on
722
+ all 6** (3/3 shared every time), because the phrase arm correctly has nothing
723
+ to contribute. No regression.
724
+ - `Rake::FileList`: the vector arm's top-1 was `Rake::FileList::[]`, a class
725
+ method. The phrase arm found `#==`, `#resolve`, `#*` as literal matches, and
726
+ fusion promoted the doubly-matched `#==` and `#*` above it.
727
+
728
+ That is the honest characterisation: **fusion earns its keep on lexical queries,
729
+ where the two engines disagree, and is a harmless no-op on questions.** Every
730
+ result carries `matched`, so the caller can see which case they are in.
731
+
732
+ A real keyword arm needs term selection before it goes near `:any` — drop the
733
+ common words, and require at least one genuinely rare term so that a lone
734
+ incidental match cannot outrank prose. That is the next piece of work, and it
735
+ should be eval-gated rather than assumed. §11.4 is that work, written down.
736
+
737
+ ### 11.4 A keyword arm: what already exists, what is missing, what to do
738
+
739
+ **Status: designed, not built.** Everything below was verified against the
740
+ SQLite that ships in this environment (3.53.2 via sqlite3 2.9.6) rather than
741
+ assumed from the FTS5 documentation.
742
+
743
+ The question that prompted this section was whether `websearch_to_tsquery` could
744
+ be borrowed from `gemchat_app`, which uses
745
+ `websearch_to_tsquery('english', …)` against a `tsvector` column generated from
746
+ `searchable_text`. It cannot be borrowed — it is PostgreSQL, and gemchat is
747
+ SQLite — but the underlying idea is right, and separating its three parts
748
+ changes what the work actually is:
749
+
750
+ | `websearch_to_tsquery('english', …)` gives you | gemchat status |
751
+ |---|---|
752
+ | stemming / root-word matching | **already present** |
753
+ | stopword removal | **absent** |
754
+ | injection-safe parsing | **already present** |
755
+
756
+ #### Stemming is already on
757
+
758
+ `chunks_fts` is declared `tokenize='porter unicode61'`, and FTS5 applies a
759
+ table's tokenizer to **both** sides of a match. Verified:
760
+
761
+ | Query | Document text | Result |
762
+ |---|---|---|
763
+ | `run` | `running a stream…` | matches — stemmed |
764
+ | `streams` | `stream` | matches — stemmed |
765
+ | `parse` | `parser` | **no match** — over-stemmed |
766
+ | `the`, `of`, `in`, `a` | present | matches — **indexed, not stopwords** |
767
+
768
+ So "switch to stem/root word matching" is already done. The single real gap is
769
+ the last row: FTS5 has no `stopwords=` option, so `the`/`of`/`how` sit in the
770
+ index and match. On an OR'd query those are exactly the terms that hit nearly
771
+ every chunk, which is the noise measured in §11.3 as
772
+ `RuboCop::Cop::MinBodyLength` outranking prose about line length.
773
+
774
+ Two fidelity caveats that no amount of work here removes:
775
+
776
+ - **FTS5's `porter` is Porter-1980; PostgreSQL's `english` is Snowball.** They
777
+ are different algorithms. Local and hosted retrieval will never stem
778
+ identically, and this is not fixable by adopting a list.
779
+ - **Porter over-stems Ruby-shaped words.** `parser` reduces to `pars` while
780
+ `parse` does not, so the two never match. Ruby docs are full of `-er`/`-or`
781
+ nouns, so this is a real recall cost and must be re-flagged in the docs rather
782
+ than left to look solved.
783
+
784
+ The rest of the `websearch_to_tsquery` feature set is either already covered or
785
+ free, and is deliberately **out of scope** for this slice:
786
+
787
+ - *Injection safety* — `fts_query` quotes every term, so `:` and FTS5 operators
788
+ cannot escape. Already have it; the systematic `Puma::Server` collision is
789
+ handled.
790
+ - *Phrase / `OR` / `-exclusion` / prefix* — FTS5 supports all of these natively
791
+ (verified: `OR`, `NOT`, `NEAR`, implicit AND, `parse*`, column filters). They
792
+ are not *exposed*: `fts_query` quotes each whitespace token, so a user typing
793
+ `foo OR bar` gets the literal phrase `"foo OR bar"`. Exposing them is a
794
+ hand-rolled parser with its own escaping edge cases — deferred, see the
795
+ out-of-scope table in §11.4.
796
+ - *Trigram* — available and working (bare-term substring: `FileLis` → 1 hit).
797
+ This is the lever for identifier near-misses, and `gemchat_app` already scores
798
+ on it. It needs a second FTS table plus sync in `replace_gem` — deferred, see
799
+ the out-of-scope table in §11.4.
800
+
801
+ #### Stage 1 — query-side stopwords in `:any` mode only
802
+
803
+ **Do not reuse `gemchat_app`'s stopword list.** At
804
+ `app/services/agentic/gem_similarity_service.rb:132` it is tuned for SEO title
805
+ work and drops `gem`, `ruby`, `library`, `tool`, `management`, `system` — all of
806
+ which are signal in a documentation index. Use Snowball's list instead, vendored
807
+ from PostgreSQL's `src/backend/snowball/stopwords/english.stop`: **127 words**,
808
+ counted, and containing no domain vocabulary.
809
+
810
+ Change:
811
+
812
+ - New `lib/gemchat/stopwords.rb` exposing `Gemchat::Stopwords::ENGLISH` as a
813
+ frozen `Set`. The comment records the source and the rejection above, so the
814
+ list is not quietly swapped later.
815
+ - `Store.fts_query(raw, mode:)` — in `:any` mode only, drop a term when **every**
816
+ alphanumeric token of it is a stopword. That rule rather than a plain
817
+ `include?` check matters: `unicode61` splits on non-alphanumerics, so `how-to`
818
+ and `don't` are multi-token terms whose tokens are all stopwords, and dropping
819
+ them is what actually removes the noise. If *every* term is dropped, fall back
820
+ to the unfiltered set — a stopword-only query must not silently return nothing,
821
+ because the caller asked a real question.
822
+
823
+ `:phrase` is **not** touched. Dropping terms from a literal phrase changes what
824
+ it matches, so `search` and `lex:` keep exact behaviour and the flagship command
825
+ carries zero regression risk. `search` still works with no model and no stopword
826
+ table, which matters for a CLI that must run on a bare `bundle install`.
827
+
828
+ #### Tests
829
+
830
+ New `test/stopwords_test.rb`:
831
+
832
+ - `ENGLISH.size == 127` — pins a truncated or mis-pasted list, which is the most
833
+ likely way this rots silently.
834
+ - Every entry is lowercase and alphanumeric.
835
+ - **The `gemchat_app` guard:** assert `rake`, `task`, `file`, `gem`, `method`,
836
+ `class` are *absent*. This is the test that stops someone "fixing" the list by
837
+ copying the app's.
838
+ - `fts_query(mode: :any)` drops `how do I` from a question and keeps `list`,
839
+ `tasks`.
840
+ - Multi-token handling: `how-to` and `don't` dropped, `task-manager` kept.
841
+ - An all-stopwords query falls back instead of returning `nil`.
842
+ - `fts_query` `:phrase` output is unchanged for a fixed input — an explicit
843
+ regression guard on `search`.
844
+
845
+ Plus one end-to-end check in `test/store_test.rb` that stopwords no longer match
846
+ via `search(query, mode: :any)`.
847
+
848
+ #### Eval gate
849
+
850
+ Build the corpus and run the five known-answer queries from §11.3 against
851
+ `fts_query` in both modes. Configuration, so the number is reproducible:
852
+
853
+ - A synthetic `Gemfile.lock` built from **33 locally installed gems**, ~21s to
854
+ index, **22,141 chunks**.
855
+ - **`gemchat` is excluded, deliberately.** It is `path:`-sourced, so its own
856
+ `docs/plans/*.md` gets indexed as prose — and those documents quote the eval
857
+ queries verbatim. The first run included it and `:phrase` "scored" 0/5 by
858
+ matching the plan that documents the failing queries. A corpus containing the
859
+ project evaluating itself is not a baseline. Excluding it left both numbers
860
+ unchanged, which is worth knowing but should not have needed checking.
861
+
862
+ | Arm | Known-answer result |
863
+ |---|---|
864
+ | `:phrase` | **0/5** |
865
+ | `:any` (current) | **1/5** |
866
+ | `:any` + stopwords | **1/5** |
867
+ | `:all` — AND of the content terms (probe, never shipped) | **0/5** |
868
+
869
+ **The gate was not met. `:any` remains off by default, and `:phrase` stays.**
870
+
871
+ Stopwords changed nothing, which is worth more than a marginal number. It
872
+ rules out the explanation §11.3 offered — that the OR'd arm was drowning in
873
+ common words — and leaves the real one. Look at what actually wins:
874
+
875
+ | Query | Content terms | `:any` top-1 |
876
+ |---|---|---|
877
+ | write a rubocop cop that checks line length | `write rubocop cop checks line length` | `RuboCop::Cop::MinBodyLength` |
878
+ | generate documentation from source comments | `generate documentation source comments` | `Parser::Source::Comment#document?` |
879
+ | run tests in parallel with processes | `run tests parallel processes` | `Minitest::Parallel::Executor` ✓ |
880
+
881
+ `MinBodyLength` and `Comment#document?` are short ri signatures containing one
882
+ query term. They beat the prose chunk that actually answers the question,
883
+ because BM25's length normalisation rewards exactly that shape. Removing `the`
884
+ and `of` never touched the terms doing the damage — `length`, `document`,
885
+ `task` were already the discriminating ones.
886
+
887
+ And the obvious correction makes it worse, not better. Requiring **all** the
888
+ content terms (`:all`, FTS5's implicit AND) drops to **0/5** — it
889
+ over-constrains, and `"parse" "ruby" "source" "abstract" "syntax" "tree"` matches
890
+ nothing in a 22k-chunk corpus at all.
891
+
892
+ #### What this rules out, and what it means
893
+
894
+ Four combinations of FTS5's boolean semantics were measured, and the spread is
895
+ 0/5, 1/5, 1/5, 0/5. **Query-side term selection is exhausted.** Stopwords,
896
+ AND, OR and phrase are the whole vocabulary available, and none of them fixes a
897
+ paraphrase question, because the problem is not which terms are selected — it
898
+ is that the question and the answer share no vocabulary. The chunk that answers
899
+ *generate documentation from source comments* is RDoc's README, which never
900
+ uses the phrase "from source comments".
901
+
902
+ So Stage 2 as originally written — a rare-term requirement, whether via
903
+ `fts5vocab` IDF or a second intersecting `MATCH` — is **falsified**, not merely
904
+ deferred. It was going to attack term *selection* inside a lexical arm, and
905
+ term selection is not the failure. A paraphrase query needs a different
906
+ mechanism entirely: document or query expansion, a learned lexical expansion
907
+ model, or a reranker over a candidate set. That is well beyond a stopword list
908
+ and is not planned here.
909
+
910
+ The honest conclusion, which matches what §11.3 already measured from the other
911
+ direction:
912
+
913
+ - The **lexical arm is for identifiers**, where the query terms appear
914
+ literally. `Rake::FileList` and friends, where fusion genuinely earns its keep.
915
+ - The **vector arm is for paraphrase**. It answers all five of these questions
916
+ correctly, unaided.
917
+ - A hybrid that fuses both is correct for a query that is *partly* one and
918
+ partly the other, and the fusion is currently a harmless no-op when the query
919
+ is purely a question. That is not a defect to be engineered away; it is the
920
+ correct behaviour, and it is now measured rather than assumed.
921
+
922
+ `Stopwords` stays in the codebase. It is ~15 tested lines behind an explicit
923
+ `mode:`, it changed no shipped behaviour, and `:any` is still the mode a future
924
+ expansion or reranking experiment would reach for. It simply is not the fix, and
925
+ the tests now say so by pinning the count and the domain-vocabulary guard rather
926
+ than by implying a quality win it does not deliver.
927
+
928
+ To make the gate enforceable rather than a claim in a commit body, add
929
+ `bin/eval-fusion` alongside `bin/sandbox`: it builds the corpus, runs the query
930
+ set through all three arms, and prints the table.
931
+
932
+ #### Explicitly out of scope
933
+
934
+ | Stage | Why not now |
935
+ |---|---|
936
+ | ~~Rare-term requirement~~ | **Falsified.** Measured above: AND of the content terms scores 0/5, worse than OR's 1/5. Term *selection* is not the failure, so per-term IDF and a second intersecting `MATCH` were aimed at the wrong thing. |
937
+ | Document-side stopword removal | Superseded. Stopwords were not the cause, so removing them from the index would not have helped either. |
938
+ | Document or query expansion | The mechanism that *would* address a paraphrase question — the vocabulary mismatch is the actual blocker. Beyond a stopword list and not planned here. |
939
+ | Expose `OR` / `NOT` / `NEAR` / prefix | A parser with its own escaping edge cases, and the eval above says more boolean surface is not the lever. |
940
+ | Trigram table | Second FTS table plus sync in `replace_gem`. The one deferred item that targets identifiers, which *is* where the lexical arm is strong. |
941
+
942
+ ### 11.5 Source locations from Prism
943
+
944
+ `Gemchat::Symbols` is the sole source of `file:line`, and the only part of this
945
+ design that ri cannot supply. ri reports a class and a method; it does not report
946
+ which file the method is defined in, or on which line. That is the one thing a
947
+ result cannot be used without and cannot be inferred from, so it is printed after
948
+ every ri hit:
949
+
950
+ ```
951
+ 1. [rake 13.4.2 ri ] Rake::FileList#exclude
952
+ exclude(*patterns, &block) (lib/rake/file_list.rb:150)
953
+ ```
954
+
955
+ Ported from the app's `RubySourceDefinitionService`, minus Rails — no
956
+ `Rails.logger`, no `blank?`/`presence`, and the namespace stack is plain array
957
+ manipulation. The visitor logic is unchanged, including its known imprecision: a
958
+ method defined inside `extend SomeModule` is reported under its *lexical* owner
959
+ rather than the mixin. §4.4 argues that is arguably more precise, since it tells
960
+ the reader the method arrives via a mixin.
961
+
962
+ Measured on rake 13.4.2: **146 of 179 ri chunks (82%) resolve to a file and
963
+ line.** The remainder are C extensions, `define_method`, and macros that
964
+ generate methods, none of which are recoverable from parsing alone. An unmatched
965
+ chunk keeps a nil location rather than borrowing a neighbouring chunk's path,
966
+ because an invented citation is worse than an absent one.
967
+
968
+ Three details that cost time:
969
+
970
+ - **Paths are relative to the gem root**, so the indexer passes `gem_dir` and
971
+ paths come out `lib/rake/file_list.rb`. Passing `lib/` yields `file_list.rb` —
972
+ correct for the argument given, and a different contract.
973
+ - **The lookup key normalises a missing owner to `""`.** A top-level `def` has no
974
+ owner, so its `class_name` is `nil`, and an ri chunk for it also has `nil`.
975
+ Both sides `.to_s`, which is what lets the two meet; without it every top-level
976
+ method silently misses.
977
+ - **The content digest now covers the location.** A definition that moves files is
978
+ a changed chunk, and reusing the old row would leave a citation pointing at the
979
+ wrong line.
980
+
981
+ `schema_version` moves 1 → 2, so existing stores need `gemchat reindex` — the
982
+ `Store::IdentityError` path already names that command.
983
+
984
+ ## 12. Hosted tier (gemchat.org)
985
+
986
+ For users who don't want ri generated, a 146MB model, or the disk at all. In v1,
987
+ explicit opt-in only. **Built**; see §12.6 for the two corrections this section
988
+ needed.
989
+
990
+ ### 12.1 Activation is always explicit
991
+
992
+ | Mechanism | Scope |
993
+ |---|---|
994
+ | `gemchat init --hosted` | stores the default for the project (`.gemchat.yml`) |
995
+ | `gemchat query --hosted "..."` | one invocation |
996
+ | `GEMCHAT_HOST=gemchat.org` | environment override |
997
+ | `GEMCHAT_API_KEY` | environment, for CI |
998
+ | `$GEMCHAT_CONFIG_HOME/credentials` | `GEMCHAT_API_KEY=…`, chmod 600 |
999
+
1000
+ The key is read from the **config home**, not the data home, for the same reason
1001
+ trust and settings are: wiping the index must not discard a key.
1002
+
1003
+ **Never automatic.** If the local index is empty or unusable, gemchat
1004
+ **diagnoses and stops** (§12.4). A local-first privacy tool must not silently
1005
+ send queries to a server.
1006
+
1007
+ ### 12.2 Command mapping
1008
+
1009
+ | Local | Transport | Notes |
1010
+ |---|---|---|
1011
+ | `query` | `POST /api/v1/query` | hybrid retrieval, cited, ~$0.000001, no LLM |
1012
+ | `search`, `vsearch` | — | **local-only**, no hosted equivalent |
1013
+ | `status` | — | names the configured backend and, if unset, the settings URL |
1014
+ | — | `POST /api/mcp` (`query_docs`) | the agent-facing surface, unchanged |
1015
+
1016
+ `lex:` is BM25 and therefore local-only; `query --hosted lex: …` refuses and
1017
+ names `gemchat search` rather than quietly returning something else.
1018
+
1019
+ The mapping is imperfect and the docs must say so: `search` and `vsearch` simply
1020
+ do not exist on the hosted path.
1021
+
1022
+ ### 12.3 Honest limits
1023
+
1024
+ State these in the README, in `gemchat status`, and in the shipped skill:
1025
+
1026
+ - **Version drift — but less than this section originally claimed.** The hosted
1027
+ API accepts `gem_versions`, so the local CLI sends this lockfile's exact pins
1028
+ and the answer is about the version actually installed. The response discloses
1029
+ any pin it could not apply (`scoped to 8.0.2; no per-version index for rake`),
1030
+ so drift is visible rather than silent. A gem with no per-version index still
1031
+ falls back to that gem's full docs, which is the residual limit.
1032
+ - **The dependency graph leaves the machine.** The whole point of Path A is that
1033
+ it does not.
1034
+ - **Requires a key and a network.** Not zero-config.
1035
+ - **Different ranking.** Postgres hybrid + source priors vs local RRF. Results
1036
+ will not be identical between backends, and that is expected.
1037
+
1038
+ ### 12.4 Diagnose and stop
1039
+
1040
+ `gemchat status` shows the full tier ladder so the user always knows where they
1041
+ stand:
1042
+
1043
+ ```
1044
+ gemchat: 0.1.0
1045
+ store: ~/.gemchat/index.sqlite
1046
+ index: 27 gems · 4,312 chunks
1047
+ vectors: not downloaded (146MB) → gemchat embed
1048
+ hosted: not configured → gemchat.org/settings
1049
+ ```
1050
+
1051
+ The `ri:` line from an earlier draft is gone. It suggested
1052
+ `gem rdoc --all --ri --no-rdoc` as the remedy for a missing ri store, but §6
1053
+ measured that `--all` adds 3–46% more chunks of private machinery, and
1054
+ §8.3 established that `gemchat index` generates ri itself — so the remedy is
1055
+ simply to index. Showing a command that inflates the index by up to half in
1056
+ order to avoid an index is not a remedy.
1057
+
1058
+ The `backend:` and `MCP_API_KEY` phrasing was also dropped: the tier ladder
1059
+ should name the action (`gemchat embed`, the settings URL), not the variable
1060
+ that gates it.
1061
+
1062
+ When a query cannot be served, gemchat names the tier, states the remedy, and
1063
+ stops. It never escalates on the user's behalf.
1064
+
1065
+ ### 12.5 Product and coupling notes
1066
+
1067
+ This gives the `gemchat` gem a reason to exist for people who never index
1068
+ anything: a first-party, authenticated CLI for the hosted service. That is a
1069
+ real product surface, not just a fallback.
1070
+
1071
+ The cost is coupling — the gem becomes a client of gemchat.org, so hosted
1072
+ breaking changes reach the gem. Treat gemchat.org as a **versioned API**: pin
1073
+ the tool contract, and make the client tolerant of unknown parameters (qmd's own
1074
+ MCP server silently ignores unknown params, which is the right precedent).
1075
+
1076
+ ### 12.6 Built, and two things this section got wrong
1077
+
1078
+ Shipped: `Hosted` (the client), `query --hosted`, `init --hosted`, and the
1079
+ `status` tier line. `GEMCHAT_HOST` accepts a bare host or a full URL and always
1080
+ lands on the versioned path, because §12.1 names a host.
1081
+
1082
+ **The transport is JSON, not MCP.** This section originally mapped local commands
1083
+ onto hosted *MCP* tools. The transport turned out not to be the problem — it is a
1084
+ plain JSON-RPC POST, no SSE, no session, so a client is about thirty lines. The
1085
+ problem is the response: `query_docs` returns **Markdown prose**, which is right
1086
+ for a model to read and wrong for a CLI, because a cosmetic server-side edit
1087
+ breaks every client with no error on either side. So the hosted tier speaks
1088
+ `POST /api/v1/query`, which returns rows.
1089
+
1090
+ MCP is untouched and remains the agent surface. Both transports call one
1091
+ `HostedQuery` service in `gemchat_app`, so the ranking is implemented once rather
1092
+ than twice and drifting. The rendered MCP Markdown is byte-identical to before,
1093
+ which is why its existing tests pass unchanged.
1094
+
1095
+ **Auto-indexing on a miss is off for the API, on for MCP.** That is a policy
1096
+ difference, not a bug: MCP's caller is an agent that is told what happened and
1097
+ can poll `get_index_status` and retry, whereas a CLI that silently enqueued
1098
+ indexing would spend the caller's quota on every miss with nothing in the
1099
+ response to show for it. The API requires `auto_index=true`.
1100
+
1101
+ Two bugs worth recording, both invisible in review because the code reads
1102
+ correctly:
1103
+
1104
+ - **`lex:` was not refused.** `cmd_query` splits the type off the query and then
1105
+ called `split_type` again on the bare question, which of course could not see
1106
+ it, so the guard never fired and `lex:` silently got a hosted answer. The type
1107
+ is now threaded through. This is exactly the failure the skill warns agents
1108
+ about, shipped in the thing that tells agents to avoid it.
1109
+ - **The credentials path printed empty.** The instruction used the *parsed* key
1110
+ path rather than the path itself, so with no file present it said "write a key
1111
+ to " and stopped.
1112
+
1113
+ ## 13. Indexing and resumability
1114
+
1115
+ **Scope:** direct deps by default; `--all` for transitive. ri + README +
1116
+ guides all on.
1117
+
1118
+ **Delta gates**, cheapest first:
1119
+
1120
+ 1. `Gemfile.lock` sha256 unchanged → no-op
1121
+ 2. per-gem `version` + `source_hash` match a completed row → skip
1122
+ 3. per-chunk `content_digest` → re-embed only changed chunks
1123
+
1124
+ **Resumable ri generation** — because §8.3 is the main path, this is a core
1125
+ subsystem, not a contingency. Foreground process, state persisted in SQLite (not
1126
+ a daemon; a daemon buys little and costs a lot of complexity):
1127
+
1128
+ - `ri_jobs(name, version, state, attempts, error, chunk_count, source_hash)`
1129
+ - states `pending → generating → done | failed | skipped`
1130
+ - reclaim stale `generating` rows on startup, recovering Ctrl-C and crashes
1131
+ - generate into a temp directory, then **atomically rename** into place before
1132
+ recording `done`, so a half-written store is never visible
1133
+ - SIGINT returns the in-flight gem to `pending` and exits cleanly
1134
+ - **fork a small RDoc worker pool (`nproc`), not threads** — RDoc is CPU-bound.
1135
+ Never fork across an `rllama`/llama.cpp boundary: load it only in the parent
1136
+ process. FFI state plus `fork` is a segfault waiting to happen.
1137
+
1138
+ ## 14. Install paths
1139
+
1140
+ Both supported, with different tradeoffs.
1141
+
1142
+ | | A. Gemfile (development group) | B. Global |
1143
+ |---|---|---|
1144
+ | Command | `bundle add gemchat --group development` | `gem install gemchat` |
1145
+ | Invocation | `bundle exec gemchat` | `gemchat` |
1146
+ | Version | pinned per project, reproducible | floats with `gem update` |
1147
+ | Store format | locked to the project's gemchat version | may drift |
1148
+ | Best for | teams, CI, reproducibility | solo work, one-off queries |
1149
+
1150
+ - **Never** in the `:default`/runtime group. Development tool; must not ship to
1151
+ production.
1152
+ - When both are present, `bundle exec` disambiguates and the plugin (§15)
1153
+ prefers the bundled one.
1154
+ - Drift is caught by `gemchat_version` in the store identity table (§7.1).
1155
+ - The shipped skill (§16) must say which invocation to prefer.
1156
+
1157
+ ## 15. Bundler plugin (staleness trigger)
1158
+
1159
+ Adding a gem to a Gemfile is inert. Unlike qmd's notes, **the bundle changes on
1160
+ every `bundle add` and `bundle update`**, so the index goes stale constantly.
1161
+
1162
+ ### 15.1 `gem` and `plugin` are different directives
1163
+
1164
+ Both lines are required, and they are not interchangeable:
1165
+
1166
+ ```ruby
1167
+ # Gemfile
1168
+ gem "gemchat", group: :development # → DEPENDENCIES. Provides `bundle exec gemchat`.
1169
+ plugin "gemchat" # → PLUGINS. Provides the bundle-install hook.
1170
+ ```
1171
+
1172
+ | | `gem "gemchat"` | `plugin "gemchat"` |
1173
+ |---|---|---|
1174
+ | Lockfile section | `DEPENDENCIES` | **no lockfile section at all** (§15.6) |
1175
+ | Purpose | bundle dependency, version-pinned | hooks Bundler's own lifecycle |
1176
+ | Gives you | the CLI on the load path | auto-index on `bundle install` |
1177
+ | Required extra | — | the gem must ship a top-level **`plugins.rb`** |
1178
+
1179
+ They are parsed in separate passes: during the plugins-only pre-parse,
1180
+ `Bundler::Plugin::DSL#gem` is a no-op. So `plugin` does not imply `gem`.
1181
+ `plugins.rb` sits at the top level of the gem and registers a hook:
1182
+
1183
+ ```ruby
1184
+ require_relative "lib/gemchat"
1185
+ Bundler::Plugin.add_hook("after-install-all") { |_deps| Gemchat::Bundler.auto_index }
1186
+ ```
1187
+
1188
+ **"Top level of the gem", emphatically — not of the project.** Measured, and the
1189
+ distinction is easy to get backwards because the file name is identical:
1190
+
1191
+ - A `plugins.rb` in the **project root** is **never executed**. Verified by
1192
+ putting a `$stderr.puts` at the top of one and running `bundle install`: zero
1193
+ output, while the installed plugin's own `plugins.rb` ran and printed.
1194
+ - Bundler only ever resolves the file inside the installed gem.
1195
+ `Plugin#load_plugin` reads `path = index.plugin_path(name)` and then
1196
+ `load path.join(PLUGIN_FILE_NAME)`; `validate_plugin!` asserts the file
1197
+ exists at `Pathname.new(spec.full_gem_path)`. No code path consults the app
1198
+ directory.
1199
+ - So a root `plugins.rb` is silently inert — the worst failure mode available,
1200
+ because the hook appears installed and simply never fires. Anyone following
1201
+ the usual "drop a plugins.rb in your app" advice will ship exactly that.
1202
+
1203
+ The event name must be a **string**. `Events::GEM_AFTER_INSTALL_ALL` is
1204
+ literally that same String — `define :GEM_AFTER_INSTALL_ALL, "after-install-all"` —
1205
+ so passing the constant also works. Only the underscored symbol
1206
+ `:after_install_all` is rejected, and it is rejected as `ArgumentError`, which
1207
+ `register_plugin` rescues and rewraps as `MalformattedPlugin`.
1208
+ `after-install-all` is the right event: the `*-require` hooks fire on every
1209
+ `bundle exec`, which is far too often.
1210
+
1211
+ ### 15.1a Four ways a plugin looks installed and does nothing
1212
+
1213
+ All four measured against `gemchat`'s own `plugins.rb`. Every one of them
1214
+ produces a *successful* `bundle install`, which is what makes them dangerous —
1215
+ there is no error to notice.
1216
+
1217
+ 1. **A `plugins.rb` in the project root.** Never executed (§15.1). Bundler only
1218
+ loads `<installed gem>/plugins.rb`.
1219
+ 2. **A `plugins.rb` that only runs, and never calls `add_hook`.** Executed
1220
+ exactly once — during the plugin phase of the install that just installed it —
1221
+ and never again. Bundler records *registered hooks* into
1222
+ `.bundle/plugin/index`, and `hook_plugins` gates every later invocation on
1223
+ that record. Observed directly: the index showed `hooks:` empty while the
1224
+ plugin had already printed a line during `bundle install`, so it looked alive.
1225
+ `bundle add` and `bundle update` were then complete silent no-ops.
1226
+ 3. **A bare `Bundler::Plugin` reference.** Bundler loads the file with
1227
+ `load(path, true)`, wrapping it in an anonymous module, while
1228
+ `register_plugin` is mid-flight inside `module Bundler::Plugin`. A bare
1229
+ constant can resolve against the wrapper and raise
1230
+ `NameError: uninitialized constant …::Bundler::Plugin` — measured, and it
1231
+ aborted the install. Every Bundler reference is root-scoped `::Bundler`.
1232
+ 4. **An `RAILS_ENV == "development"` guard.** Copied from Hyperdrive, where the
1233
+ host is a Rails engine. `gemchat` is not a Rails tool, so in any plain Ruby
1234
+ project `RAILS_ENV` is unset, the guard is permanently false, and the hook is
1235
+ permanently silent. Corrected to a `CI`-unset check only.
1236
+
1237
+ The common thread: **the failure mode of a plugin is silence, not an error.**
1238
+ A shim that appears installed and does nothing is strictly worse than one that
1239
+ refuses to install, because the user has no signal to act on. `test/plugin_shim_test.rb`
1240
+ asserts each of these mechanically, since all four are invisible in review.
1241
+
1242
+ ### 15.2 Setup is a separate `gemchat init`
1243
+
1244
+ Following the Hyperdrive pattern. Its `hyperdrive:install` generator ends with
1245
+ `append_to_file "Gemfile", "plugin \"bundler-rails-hyperdrive\"\n"`, after first
1246
+ checking for an existing line so the write is idempotent.
1247
+
1248
+ - `gemchat init` — the once-per-project setup verb. Appends
1249
+ `plugin "gemchat"` with an interactive confirm, writes `.gemchat.yml`,
1250
+ records trust, prints the `gem rdoc` remedy if ri is missing (§8.2), and can
1251
+ set the hosted backend (§12). **The only command that may modify the
1252
+ Gemfile.**
1253
+ - `gemchat index` — syncs the index. Non-interactive, never writes to the
1254
+ project.
1255
+
1256
+ Separate on purpose: a command named `index` that mutates your Gemfile is a
1257
+ surprise, and in CI or a pre-commit hook it would either block on a prompt —
1258
+ hanging the build — or silently edit a file it has no business touching.
1259
+
1260
+ **Activation is not immediate.** Bundler only activates a newly declared plugin
1261
+ on the next `bundle install` (Hyperdrive documents the same chicken-and-egg). So
1262
+ `init` must tell the user to run `bundle install` once to arm the hook.
1263
+
1264
+ ### 15.3 Fires on `add`, `install`, and `update` — when Bundler installs
1265
+
1266
+ The qualifier is the whole correction. `after-install-all` fires when Bundler
1267
+ performs an install, so `bundle update --all` with nothing to change installs
1268
+ nothing and the hook correctly stays silent. A test that asserted the hook "fires
1269
+ on update" while the update was a no-op was asserting the wrong thing; it now adds
1270
+ a gem to the Gemfile first, which is the ordinary reason anyone runs `update`.
1271
+
1272
+ Earlier this claim was measured against the *published* `bundler-rails-hyperdrive`
1273
+ plugin, watching for that plugin's own output, gated behind an env var so it
1274
+ skipped by default. It never exercised gemchat's hook and it failed when actually
1275
+ run — leaving `bundle add` untested while the plan claimed all three. It now runs
1276
+ gemchat's own path-sourced plugin offline via `test/support/sandbox.rb`, and
1277
+ passes.
1278
+
1279
+ Measured end-to-end against the published `bundler-rails-hyperdrive` plugin, one
1280
+ throwaway project, observing the hook's own output:
1281
+
1282
+ | Command | `after-install-all` fired | Index may need |
1283
+ |---|---|---|
1284
+ | `bundle install` | **yes** | new/changed gems |
1285
+ | `bundle add <gem>` | **yes** | the new gem immediately |
1286
+ | `bundle update` | **yes** | often many changed versions |
1287
+
1288
+ All three mutate `Gemfile.lock`, so the sha256 gate handles them uniformly and
1289
+ there is no need to detect which command ran. The plugin hooks the post-install
1290
+ point and asks one question — did the lockfile digest change?
1291
+
1292
+ `bundle add` is the case that matters: that is exactly when a developer has just
1293
+ adopted a gem. It is now confirmed rather than assumed.
1294
+
1295
+ ### 15.4 The plugin is a stateless shim, not the indexer
1296
+
1297
+ Measured: dual-declaring one gem installs **two physical copies** — one under
1298
+ the bundle's `gems/`, one under the project's `.bundle/plugin/gems/` — and
1299
+ Bundler 4.0.21 writes **no `PLUGINS` section** into `Gemfile.lock`, so nothing
1300
+ ties the plugin copy to the dependency version. Drift is therefore possible.
1301
+
1302
+ Rails Hyperdrive already solves this, and its solution is the reason to copy its
1303
+ shape rather than invent one. `rails-hyperdrive` (the engine) does **not** ship
1304
+ `plugins.rb`; a separate tiny gem, `bundler-rails-hyperdrive` is the plugin.
1305
+ Its two files, read from the installed copy:
1306
+
1307
+ ```ruby
1308
+ # plugins.rb — the only thing Bundler loads, and all of it is registration
1309
+ require_relative "lib/bundler/hyperdrive"
1310
+
1311
+ Bundler::Plugin.add_hook("after-install-all") do |_dependencies|
1312
+ Bundler::Hyperdrive.auto_install
1313
+ end
1314
+ ```
1315
+
1316
+ ```ruby
1317
+ # lib/bundler/hyperdrive.rb — the logic, never executed from the plugin copy
1318
+ require_relative "hyperdrive/version"
1319
+
1320
+ module Bundler
1321
+ module Hyperdrive
1322
+ HOST_GEM = "rails-hyperdrive".freeze
1323
+ SUPPORTED_RAILS_HYPERDRIVE = Gem::Requirement.new(">= 0.2")
1324
+ PREFIX = "[hyperdrive] ".freeze
1325
+ DEVELOPMENT = "development".freeze
1326
+
1327
+ module_function
1328
+
1329
+ # Runs inside `bundle install`: every failure must degrade to a printed
1330
+ # line, never a raised error that fails the install.
1331
+ def auto_install
1332
+ return unless development_environment?
1333
+ # ... locates HOST_GEM in the bundle and requires THAT copy's lib
1334
+ rescue StandardError => e
1335
+ report "auto-install skipped (#{e.class}: #{e.message})"
1336
+ end
1337
+ end
1338
+ end
1339
+ ```
1340
+
1341
+ The split matters and is the actual lesson. `plugins.rb` holds *only* the
1342
+ `add_hook` registration; the logic lives under `lib/` and is reached through the
1343
+ host gem found in the bundle. That is what makes the plugin copy stateless —
1344
+ there is no second implementation of anything, so there is nothing to drift.
1345
+
1346
+ `gemchat/plugins.rb` follows this shape, with one deliberate deviation recorded
1347
+ in §15.1a: the `RAILS_ENV` guard is dropped, because `DEVELOPMENT` is a Rails
1348
+ concept and gemchat is not a Rails tool.
1349
+
1350
+ Three transferable properties:
1351
+
1352
+ 1. **The shim carries no logic.** It locates the host gem *in the bundle* and
1353
+ requires **that** copy. The plugin's own copy of the code never runs, so
1354
+ version drift between the two copies is harmless. `gemchat` should do the
1355
+ same: `plugins.rb` finds the bundled `gemchat` spec via
1356
+ `Bundler.load.specs` and requires its `lib`. The only contract that must
1357
+ stay stable across versions is *locate and require*.
1358
+ 2. **A declared supported range.** `SUPPORTED_RAILS_HYPERDRIVE = ">= 0.2"` lets
1359
+ an old shim politely skip a new engine instead of misbehaving. `gemchat`
1360
+ should declare the same for its own version pair.
1361
+ 3. **A cheap environment pre-filter.** Its `development_environment?` returns
1362
+ early unless the env is development *and* `CI` is unset, with the comment
1363
+ that the hook fires on every install — CI and deploys included — where output
1364
+ would be noise. Adopt that as a first check, before any trust lookup: it
1365
+ costs nothing and removes pointless work.
1366
+
1367
+ ### 15.5 Contract
1368
+
1369
+ 1. **Never fails the bundle.** Any error is swallowed; exit status unaffected.
1370
+ This is not optional-by-choice. Measured on `gemchat`'s own `plugins.rb`, all
1371
+ three of these fail the install:
1372
+
1373
+ | `plugins.rb` | Result |
1374
+ |---|---|
1375
+ | raises | `MalformattedPlugin`, **exit 29**, no lockfile written |
1376
+ | calls `add_hook(:after_install_all)` | `MalformattedPlugin` (from `ArgumentError`), exit 29 |
1377
+ | missing entirely | `MalformattedPlugin: plugins.rb was not found`, exit 29 |
1378
+
1379
+ "Aborts" is imprecise in one respect worth recording: the gems themselves do
1380
+ get installed and `bundle list` shows them. What fails is the run — non-zero
1381
+ exit, no `Gemfile.lock` — so the project is left without a usable bundle.
1382
+ Note also that the *gems* land before the plugin phase, which is why a broken
1383
+ plugin leaves a half-configured project rather than an untouched one.
1384
+ 2. **Non-blocking.** Runs on `after-install-all`, never during install.
1385
+ 3. **Silent by default.** One line at most, prefixed, and only when work was
1386
+ done. Hyperdrive uses the literal prefix `[hyperdrive] `.
1387
+ 4. **sha256-gated.** Compares the lockfile digest against `.gemchat.yml` and
1388
+ no-ops when unchanged. Covers `add`, `install`, and `update` uniformly
1389
+ (§15.3), so no command detection is needed.
1390
+ 5. **Opt-out.** `bundle config set --local no_auto-index true` (Bundler also
1391
+ honours a global `plugins false`, which disables the hook entirely), or
1392
+ `gemchat init --no-auto-index`.
1393
+
1394
+ **The plugin must also be trust-gated.** A Bundler plugin is *checked-in code
1395
+ that executes on `bundle install`* — `git clone && bundle install` on a
1396
+ repository you do not own would run it. Hyperdrive does not address this; qmd's
1397
+ trust model does, and it is the stronger precedent. Same hazard class as the
1398
+ checked-in `.gemchat.yml` in §17.
1399
+
1400
+ The gate is low-friction for a project's own author: **the first explicit
1401
+ gemchat command run inside a project grants trust for that project**, recorded
1402
+ in `~/.gemchat/trusted.json`. Someone who just added `plugin "gemchat"` to
1403
+ their own Gemfile gets auto-indexing with no ceremony. Someone who *clones* that
1404
+ repository gets a no-op that prints one line, because they have not asked.
1405
+
1406
+ The `CI` pre-filter (§15.4) runs first, so unattended contexts exit before any
1407
+ trust lookup. `gemchat trust` pre-approves explicitly — the only way to enable
1408
+ auto-index in CI on purpose.
1409
+
1410
+ ### 15.6 Plugin mechanics — measured, not assumed
1411
+
1412
+ | Question | Answer | How |
1413
+ |---|---|---|
1414
+ | Dual `gem` + `plugin` of one gem: one copy or two? | **two** | `.bundle/plugin/gems/` + bundle `gems/`, published gem |
1415
+ | Does a `#` inside a fenced code block split a section? | it does, unless handled | `Markdown` tracks fence state |
1416
+ | Does Bundler 4 pin the plugin copy? | **no** | no `PLUGINS` section in `Gemfile.lock` |
1417
+ | Is `Bundler::Plugin` a class or module? | **module** | `module Bundler; module Plugin` |
1418
+ | `add_hook` signature | `(event, &block)` | event is a String key of `Events` |
1419
+ | Malformed `plugins.rb` | exit 29, no lockfile | raises / bad event name / missing file |
1420
+ | Project-root `plugins.rb` | **never executed** | `load_plugin` reads `index.plugin_path(name)` only |
1421
+ | Fires on `add`/`install`/`update`? | **all three**, but only when Bundler *installs* | gemchat's own path-sourced plugin, `Sandbox` |
1422
+ | Registration without `add_hook` | fires **once**, then never | `.bundle/plugin/index` records `hooks:` |
1423
+
1424
+ ### 15.6a Testing the plugin against a local checkout
1425
+
1426
+ Before release, the shim has to be exercised for real. The mechanism is a
1427
+ Bundler local override, and the exact form matters:
1428
+
1429
+ ```sh
1430
+ bundle config set --local local.gemchat /path/to/gemchat # does NOT work yet
1431
+ ```
1432
+
1433
+ `local.<gem>` is a *remote-source override*: Bundler still resolves the gem from
1434
+ rubygems.org first and only substitutes the local path for a version it finds
1435
+ there. For an unpublished gem it fails with "Could not find gem 'gemchat'".
1436
+ Pinning a version does not help.
1437
+
1438
+ Two forms that also do not do what they look like:
1439
+
1440
+ - `bundle config --local set path …` is the **deprecated** `config` form. It
1441
+ writes `BUNDLE_SET: "path …"` and silently has no effect; Bundler prints a
1442
+ deprecation warning naming the replacement. The modern
1443
+ `bundle config set --local path …` sets `BUNDLE_PATH`, which is the
1444
+ *install location* — pointing it at the gemchat checkout made even `rake`
1445
+ unresolvable.
1446
+ - `bundle config set --local local.<gem> …`, for the reason above.
1447
+
1448
+ What actually works is a path source in the Gemfile:
1449
+
1450
+ ```ruby
1451
+ gem "gemchat", path: "/path/to/gemchat"
1452
+ plugin "gemchat", path: "/path/to/gemchat"
1453
+ ```
1454
+
1455
+ Verified end to end with this: the plugin installs, `plugins.rb` is loaded from
1456
+ the checkout, `.bundle/plugin/index` records `after-install-all: [gemchat]`, the
1457
+ hook fires on `bundle install`, `bundle add` and `bundle update`, and it stays
1458
+ silent under `CI=1` and `GEMCHAT_NO_AUTO_INDEX=1`.
1459
+
1460
+ **The one thing this cannot test is the dual-copy condition.** A `path:` source is
1461
+ referenced in place — no second physical copy is made, and `.bundle/plugin/gems/`
1462
+ stays empty. So the two-copies finding in §15.4 must be re-verified against a
1463
+ genuinely *published* gem before release; the local path harness exercises every
1464
+ other mechanic but structurally cannot reproduce that one.
1465
+
1466
+ Verified against Bundler 4.0.21 and a real published plugin gem
1467
+ (`bundler-rails-hyperdrive` 0.1.0), not from memory:
1468
+
1469
+ | Question | Answer |
1470
+ |---|---|
1471
+ | Dual declaration: one copy or two? | **Two.** `<project>/.bundle/plugin/gems/<name>-<ver>` and `<bundle>/gems/<name>-<ver>` |
1472
+ | Does the lockfile pin the plugin copy? | **No.** Bundler 4.0.21 writes **no `PLUGINS` section**; only `DEPENDENCIES` |
1473
+ | How does a plugin register? | `Bundler::Plugin.add_hook("<event>") { \|args\| }` in a top-level `plugins.rb` |
1474
+ | What form does `<event>` take? | The **string** key of `Bundler::Plugin::Events` (`"after-install-all"`). The constant or an underscored symbol raise |
1475
+ | `Bundler::Plugin` a class or module? | A **module** — you cannot subclass it |
1476
+ | Available events | `before/after-install`, `before/after-install-all`, `before/after-require`, `before/after-require-all` |
1477
+ | Does a malformed plugin break the bundle? | **Yes.** `MalformattedPlugin` → `PluginInstallError` aborts `bundle install` before anything is written |
1478
+ | Does the hook fire on `bundle add`? | Yes — it fires on the post-install path, which `add` and `update` share |
1479
+ | Can plugins be disabled? | Yes — `Plugin.hook` returns early unless `Bundler.settings[:plugins]`, so `bundle config set --local plugins false` disables it globally |
1480
+
1481
+ Consequence for §7.1: because two copies exist and the lockfile does not tie
1482
+ them, the plugin **must not** be the thing that owns the store. §15.4's
1483
+ stateless-shim design is what makes that safe — the shim requires the bundled
1484
+ copy, so only one version ever runs. Store identity is then checked by the
1485
+ running copy against `meta`, and a mismatch is a clear error rather than silent
1486
+ corruption.
1487
+
1488
+ ## 16. Agent skill
1489
+
1490
+ The gem ships `skills/gemchat/SKILL.md` — a CLI is worthless to an agent that
1491
+ does not know it exists. Reuses the artifact format validated in `b9d1491`
1492
+ against `agentskills.io/specification`: frontmatter with `name` +
1493
+ `description`, body as agent-facing instructions.
1494
+
1495
+ The skill must encode:
1496
+
1497
+ - **When to use it** — questions about how a gem in *this* project works, its
1498
+ API, configuration, version differences.
1499
+ - **Which invocation** — `bundle exec gemchat` when the gem is in the Gemfile,
1500
+ otherwise bare `gemchat` (§14).
1501
+ - **Check `gemchat status` first**, and act on what it reports rather than
1502
+ guessing. It states backend, index size, ri coverage, model state, and hosted
1503
+ availability (§12.4).
1504
+ - **Version fidelity matters here.** If the local backend is active, prefer it;
1505
+ if the user is on the hosted backend, answers are about gemchat.org's indexed
1506
+ version, not the lockfile pin. Say so rather than presenting a drifted answer
1507
+ as exact.
1508
+ - **Division of labor.** Gem documentation → this tool. *This application's*
1509
+ runtime state, routes, schema, or logs → Rails Hyperdrive.
1510
+ - **Cost expectations** — `search` is instant and free with no model on disk;
1511
+ `vsearch` and `query` need `gemchat embed` to have run first and do **not**
1512
+ download anything themselves. The one-time download is **146MB**, not the
1513
+ 300MB this section originally said, and only `embed` performs it.
1514
+
1515
+ Two corrections to this section as written, both made when the skill was
1516
+ authored:
1517
+
1518
+ - The 300MB figure predates §10.1; the shipped default is 146MB.
1519
+ - This section says `status` reports "hosted availability", which it does — as
1520
+ the `hosted: not configured → gemchat.org/settings` line. But §12 is not built,
1521
+ so there is no hosted backend to be available yet. The skill says so rather
1522
+ than describing a tier that does not exist.
1523
+
1524
+ Side benefit: because it ships in a `skills/` directory, it is also consumable
1525
+ by Rails Hyperdrive's `enabled:` opt-in, which installs skills from any bundled
1526
+ gem regardless of native Hyperdrive support.
1527
+
1528
+ ## 17. Configuration and trust
1529
+
1530
+ A checked-in `.gemchat.yml` travels with a `git clone` and must not be trusted
1531
+ unattended. **Four** things can reach outside the project or execute code:
1532
+
1533
+ - command hooks
1534
+ - collection/store paths outside the project
1535
+ - custom model URIs
1536
+ - **the Bundler plugin auto-index hook** (§15)
1537
+
1538
+ Two ways to grant approval, recorded in `~/.gemchat/trusted.json`:
1539
+
1540
+ - **Implicitly, by use** — the first explicit `gemchat` command run inside a
1541
+ project approves that project. The common path, no ceremony.
1542
+ - **Explicitly, via `gemchat trust`** — pre-approves a project you have
1543
+ reviewed, and is the only way to enable auto-index in CI on purpose.
1544
+
1545
+ Non-interactive contexts never reach the implicit grant and therefore **skip**
1546
+ gated behaviour, continuing with in-project indexing. `gemchat trust list` and
1547
+ `gemchat trust revoke` inspect and withdraw approvals. Editing any gated field
1548
+ invalidates a prior approval and re-asks, matching qmd.
1549
+
1550
+ ## 18. CLI surface
1551
+
1552
+ ```
1553
+ gemchat init [--hosted|--no-local] [--no-auto-index] # once per project; only cmd that writes the Gemfile
1554
+ gemchat index [--all] [--dry-run] [--jobs N] [--ri-only]
1555
+ gemchat embed [--force] [--all] # vectors only; downloads the model on first use
1556
+ gemchat search <query> [-n N] [--json|--files] [--min-score] [--explain]
1557
+ gemchat vsearch <query> [-n N] [--json|--files]
1558
+ gemchat query <query> [-n N] [--json|--files] [--explain] [--hosted]
1559
+ gemchat status # backend, delta, disk, store identity, tier ladder
1560
+ gemchat trust [list|revoke]
1561
+ gemchat reindex # force, after a store-identity change
1562
+ gemchat version
1563
+ ```
1564
+
1565
+ `index` and `embed` are separate verbs, following qmd (`qmd update` vs
1566
+ `qmd embed`). That separation is what lets `search` work instantly after
1567
+ `index` without a 300MB download; `embed` is the explicit opt-in to semantic
1568
+ recall. `vsearch` and `query` also trigger it lazily on first use, so the
1569
+ explicit command is a convenience rather than a requirement.
1570
+
1571
+ Agent-first output: `--json`, `--files`, `--explain` (retrieval score traces),
1572
+ `--min-score`, and `gemchat://` URIs. Defaults tuned for agents, not humans.
1573
+
1574
+ ## 19. Repository work
1575
+
1576
+ Now done, and verified by a `gem build` that produces a non-empty package:
1577
+
1578
+ - **`exe/gemchat` exists.** `bindir = "exe"` and `executables` greps
1579
+ `spec.files` for `exe/`, so without it there is no binary. The convention
1580
+ splits `exe/` (ships, lands on `PATH`) from `bin/` (dev scripts, explicitly
1581
+ rejected from `spec.files` — already holds `bin/console` and `bin/setup`).
1582
+ - **Gemspec metadata filled in** — `summary`, `description`, `homepage`,
1583
+ `source_code_uri`, `changelog_uri`, and `rubygems_mfa_required = "true"`.
1584
+ - **The description is unambiguous about local-vs-hosted.** Users will see
1585
+ `gemchat` in a Gemfile and reasonably expect the gemchat.org *service*. It
1586
+ states plainly that this is a local indexer with an optional hosted mode: no
1587
+ account needed for local use.
1588
+ - **Runtime dependencies declared**: `sqlite3` (FTS5, precompiled natives),
1589
+ `rllama` (prebuilt llama.cpp binaries, so no compiler on install), `prism`
1590
+ (unconditional, so the symbol pass behaves identically on Ruby 3.2), `rdoc`.
1591
+ The omission of these from the skeleton gemspec was not caught by the manual
1592
+ smoke test, only by running the suite under `bundle exec` — where
1593
+ `require "sqlite3"` failed, because an undeclared gem happens to be present
1594
+ outside a bundle. **A smoke test that runs outside `bundle exec` will not
1595
+ catch a missing runtime dependency.**
1596
+ - **Linter is `standardrb`** (`.standard.yml`), not rubocop.
1597
+ - `spec.files` derives from `git ls-files -z`, so untracked files are invisible
1598
+ to the built gem. `git add` before `gem build` still matters — this will cover
1599
+ `skills/gemchat/SKILL.md` and `plugins.rb` when they arrive.
1600
+ - **`plugins.rb` is shipped** at the top level of the gem, in the Hyperdrive
1601
+ shape: the file holds only the `add_hook` registration, and the logic lives in
1602
+ `lib/`. Verified against a local `path:` checkout — installs, records
1603
+ `after-install-all` in `.bundle/plugin/index`, fires on install/add/update,
1604
+ silent under `CI` and `GEMCHAT_NO_AUTO_INDEX`. It currently only reports that
1605
+ auto-index is not implemented yet (§20 phase 2). The dual-copy condition is
1606
+ still unverified for a *published* gemchat; see §15.6a.
1607
+ - Optional `spec.metadata["hyperdrive_targets"]` for Hyperdrive `discover`
1608
+ interop (§16).
1609
+ - **Prism** as an unconditional runtime dependency. Ruby 3.3+ bundles it, but
1610
+ depending on the gem unconditionally keeps behaviour identical on 3.2.
1611
+ - **No test framework** declared. Minitest, to match the Rails app. (Done in
1612
+ Phase 1: 77 examples, no skips.)"
1613
+ - Runtime dependencies: `sqlite3`, `rllama`, `prism`, `rdoc`. **Not**
1614
+ `rdoc-data` (§4.2) and **not** `rubydex` (§4.4) — the latter is a marginal
1615
+ owner-resolution gain and would break `bundle install` for musl users.
1616
+
1617
+ ## 20. Phases
1618
+
1619
+ 1. ~~**Spike — ri resolution, chunking, FTS5, and the plugin mechanics.**~~
1620
+ **Done 2026-09-26.** ri generation cost measured across 49 real direct deps
1621
+ (9.9s total, ~2s on 4 cores — see §6); `read_ri_from_path` ported; chunker and
1622
+ FTS5 store built; `gemchat index` + `gemchat search` working end-to-end with
1623
+ exact signatures; all plugin mechanics answered in §15.6, including the
1624
+ dual-copy result that forces the shim design in §15.4.
1625
+ **Test suite added alongside it:** 77 Minitest examples, no skips, covering
1626
+ the FTS quoting rules, chunk caps and digests, the store's
1627
+ indexed/updated/unchanged outcomes, RDoc failure cleanup, and CLI exit codes.
1628
+ Four implementation defects were found by writing those tests: a memoised
1629
+ `Gemchat.root` that made the *first* `GEMCHAT_HOME` reader win for the life of
1630
+ the process (tests were writing to the real `~/.gemchat`), `Store#search`
1631
+ leaking sqlite3's integer row keys into results, blank fields being indexed as
1632
+ dead tokens, and `index` reporting `generated … 0 chunks` on a no-op. Fixed
1633
+ in this phase; see §4.5.
1634
+ 2. ~~**Resumable ri generation** (§13) — **largely superseded.** The phase-1
1635
+ measurement showed generation is seconds, so this was never a necessity. The
1636
+ specific cost it was meant to fix — a repeat run re-reading and re-chunking
1637
+ every gem, 0.19s for irb and 0.04s for rake — is now handled one level up:
1638
+ the `.gemchat.yml` digest gate means `gemchat index` is only reached at all
1639
+ when the lockfile moved, and the per-gem store digest skips unchanged gems
1640
+ after that. Worth revisiting only if a real bundle makes indexing visibly
1641
+ slow, not as planned work.
1642
+ 3. ~~**Delta indexer + `.gemchat.yml` manifest + store identity (§7.1) +
1643
+ trust**~~ — **Done 2026-09-27.** `5dac123` (trust + sha256 gate), `c0f432d`
1644
+ (`gemchat init`, `reindex`, store identity). The detail is in the "Also landed
1645
+ with the plugin" note below.
1646
+ 4. ~~**Embeddings via rllama + `gemchat embed` + model download UX + store-identity
1647
+ guard**~~ — **Done 2026-09-27.** `bfc730f` (model download), `0fa0477` (vector
1648
+ storage and ranking), `8c2f6be` (`embed`, `vsearch`). The default model
1649
+ changed from the plan's `embeddinggemma-300M` to `nomic-embed-text-v1.5`
1650
+ because the former is gated; §10.1 has the reasoning and §10.3 the crash
1651
+ measurements.
1652
+ 5. ~~**Hybrid RRF + source priors + Prism symbol metadata**~~ — **Done
1653
+ 2026-09-27**, and not how the phase was written.
1654
+ - *Hybrid RRF* landed in `34d137b`. §11.3 records what it measured: a no-op on
1655
+ questions, because the lexical arm is a phrase matcher, and §11.4 records
1656
+ the designed-but-falsified fix.
1657
+ - *Source priors* were measured before being built, and **not built**. Against
1658
+ an evenly-split ri/prose eval of 10 known-answer questions, the app's own
1659
+ 1.15 prose prior is neutral on hit@1 (6/10 either way) and worse on MRR
1660
+ (0.744 → 0.705); 1.30 gives 5/10, and a 1.15 ri prior gives 4/10 with MRR
1661
+ 0.547. Recall@10 is 10/10 in every case, so a prior only ever reorders, and
1662
+ always for the worse. A prior shifts a whole class of documents and cannot
1663
+ tell a good prose chunk from a bad one, so the app's constant does not
1664
+ transfer to a bundle-scoped index.
1665
+ - *Prism symbol metadata* landed as `Gemchat::Symbols`.
1666
+ 6. **Hosted tier** (§12) — `--hosted`, command mapping, the tier ladder in
1667
+ `status`
1668
+ 7. ~~**`skills/gemchat/SKILL.md`**~~ — **Done 2026-09-27.** Written against the
1669
+ real CLI rather than this section, which had drifted: it says 146MB, not
1670
+ 300MB, and it does not claim `vsearch` may trigger a download, because only
1671
+ `embed` downloads. The most valuable thing in the file is the routing table
1672
+ and the warning that `search` is a phrase matcher, because that is the one
1673
+ thing an agent cannot infer and the one that silently returns nothing.
1674
+ `test/skill_test.rb` guards it mechanically: the skill must name only
1675
+ commands that exist, must document every command that does, and must not
1676
+ regress to a stale model size.
1677
+ 8. MCP **server** — deferred; the hosted gemchat.org already owns that surface
1678
+
1679
+ Also landed with the plugin: `Gemchat::Lockfile` (the sha256), `Gemchat::Manifest`
1680
+ (`.gemchat.yml`), `Gemchat::Trust` (`trusted.json`), and `Gemchat::Hook` (the
1681
+ logic, reached through the bundled copy). `gemchat index` grants trust for the
1682
+ current project and records the manifest, so opting in is simply running it
1683
+ once; the hook then indexes on any later `install`/`add`/`update` that moves the
1684
+ lockfile digest, and stays completely silent when nothing has.
1685
+
1686
+ Also landed: `gemchat init` (§15.2), the only command that may modify a Gemfile.
1687
+ It adds `gem "gemchat", group: :development` and `plugin "gemchat"`, idempotently
1688
+ and matched per directive, grants trust, and tells the user to run `bundle install`
1689
+ once because Bundler does not activate a newly declared plugin until the next
1690
+ install. `gemchat reindex` is the remediation `Store::IdentityError` names.
1691
+ `gemchat status` now reports the lockfile digest, manifest state, trust, and
1692
+ whether the hook is actually *armed* — read from Bundler's own plugin index,
1693
+ because a Gemfile line only says the plugin was requested, not that it runs.
1694
+
1695
+ Store identity (§7.1) is in: `schema_version` is a hard gate and the store
1696
+ refuses to open on a mismatch, `gemchat_version` drift is tolerated and surfaced
1697
+ in `status`, and `embedding_model`/`embedding_dimensions` are reserved for the
1698
+ embedding phase.
1699
+
1700
+ ### 20.1 Prose ingestion — `Gemchat::Prose`
1701
+
1702
+ `Gemchat::Markdown` splits prose into heading-scoped sections, handling ATX,
1703
+ setext and RDoc headings and treating a `#` inside a fenced block as code rather
1704
+ than a heading. `Gemchat::Prose` discovers the sources below and labels each one,
1705
+ so a gem's rows never overwrite each other and search output says what kind of
1706
+ document a hit came from.
1707
+
1708
+ | Source | Path | `source_type` |
1709
+ |---|---|---|
1710
+ | README | `README.{md,markdown,mkd,rdoc,txt}`, bare `README` | `readme` |
1711
+ | Guides | `guides/**`, `guide/**` | `guides` |
1712
+ | Reference docs | `doc/**` prose | `doc` |
1713
+
1714
+ **`readme`, `guides` and `doc` stay distinct types, grouped only at query time.**
1715
+ This is the vocabulary `gemchat_app` already uses, and it is not incidental: its
1716
+ `search_service.rb` gives `guides` a prior of `1.15` — ranked *up*, because guides
1717
+ are better than baseline for the questions they suit — and derives a
1718
+ `readme_guides` group via `source_group` purely to enforce prose diversity
1719
+ (`WORKFLOW_MIN_README_GUIDES`). Storing them separately and grouping at query time
1720
+ is the model. Folding `guides/` into `readme` would also make local `guides` mean
1721
+ "bundled `guides/`" while hosted `guides` means "web-crawled reference pages" —
1722
+ drift that is miserable to unpick later. `doc/` is likewise kept separate from
1723
+ hosted `guides`: the local tool's whole claim is offline and local, and conflating
1724
+ "shipped in the gem" with "fetched from the internet" would muddy that.
1725
+
1726
+ `guides/` is **0 of 409** gems in the measured population, so its own type is
1727
+ justified by vocabulary and ranking headroom, not by data. `doc/` is 3.9%. Stated
1728
+ plainly so the split is not mistaken for evidence-driven.
1729
+
1730
+ **`doc/` is hand-written prose, and an earlier draft of this plan asserted the
1731
+ opposite.** The claim was that `doc/` holds generated ri and that indexing it
1732
+ would "double the corpus with the same text `ri` already covers". Measured across
1733
+ every installed gem with a `doc/`:
1734
+
1735
+ | | |
1736
+ |---|---|
1737
+ | gems with `doc/` | 12 |
1738
+ | files | 96 |
1739
+ | prose (`.md`/`.rdoc`/`.txt`) | **83** |
1740
+ | generated ri | **0** |
1741
+ | files byte-identical to the README | **0** |
1742
+ | subdirectories | all hand-written: `example`, `irb`, `releases`, `markup_reference`, `ja`, `en` |
1743
+
1744
+ The generated shape, `doc/<name>-<version>/ri`, appears in none of them. What is
1745
+ there is exactly what ri cannot produce: `irb/Index.md` (23KB), `irb/EXTEND_IRB.md`,
1746
+ `irb/Configurations.md`, `rake/glossary.rdoc`, `test/how-to.md`,
1747
+ `rubygems/UPGRADING.md`, `racc/doc/en/grammar.en.rdoc`. The claim could not be
1748
+ traced to any source — not `ri-doc-generation.md`, and `gemchat_app`'s own
1749
+ AGENTS.md lists `guides` as *web-crawled*, not read from `doc/`. It appears to be
1750
+ an over-generalisation from the ri-*generation* context, where RDoc's default
1751
+ target happens to be called `doc/`. Asserting a plausible rule without measuring
1752
+ it is the exact failure mode this document keeps warning about.
1753
+
1754
+ The generated-ri case is still excluded, specifically: any `**/ri/**` path under
1755
+ `doc/`, because `gem install` with documentation does write there, and it costs
1756
+ nothing. So is generated web output (`.html`, `.js`, `.css`). This is the
1757
+ narrow, correct version of the earlier rule rather than a rejection of it.
1758
+
1759
+ **Non-English documents are skipped by default**, with `--all-languages` to opt
1760
+ in. `doc/` carries substantial translations — `irb.rd.ja` (15KB),
1761
+ `grammar.ja.rdoc` (14.5KB), `usage.ja.html` (18KB) — and Japanese text pollutes
1762
+ BM25 with tokens no English query can match, while still costing embedding time
1763
+ later. Detection is by filename and directory marker (`ja`, `zh`, `ko`, and a
1764
+ documented short list), which is predictable and covers every observed case:
1765
+ `racc/doc/ja/`, `irb/doc/irb/irb.rd.ja`, `typeprof/doc.ja.md`. `en` is always
1766
+ kept. This also fixes `racc`, whose *only* README is `README.ja.rdoc` and which
1767
+ therefore currently indexes Japanese into an English index.
1768
+
1769
+ **Long sections are split, not truncated.** `MAX_BODY_CHARS` is 2000, tuned for ri
1770
+ method bodies where that is naturally the whole unit. Applied to prose it silently
1771
+ dropped **13% of all `doc/` prose**, concentrated in files with long sections:
1772
+ `tutorial.rdoc` (40KB, 22% lost), `master_toc.rdoc` (40%), `doc.md` (13%),
1773
+ truncating mid-sentence. Sections over the cap are now split on blank lines into
1774
+ consecutive chunks, titled with a ` · 2/5` suffix so a search returning three
1775
+ parts of one section does not look like three unrelated hits. The cap itself is
1776
+ unchanged, so there is no new constant and ri chunks are unaffected.
1777
+
1778
+ Sections are stored under their own `source_type` so a gem gaining or losing a
1779
+ README does not rewrite its much larger ri row. A gem then has one row per source
1780
+ type, so `Store#stats` counts distinct name+version rather than rows, or every gem
1781
+ with a README would be reported twice.
1782
+
1783
+ Measured after the change, across every installed gem with a `doc/`: **1,406
1784
+ chunks** and **601KB** of prose, `doc` 45 / `readme` 12 / `ri` 179 for a
1785
+ two-gem rake-only bundle. `rake/doc/command_line_usage.rdoc` is now findable —
1786
+ `rake › doc › Rake Command Line Usage` — which was unreachable before.
1787
+
1788
+ Still ahead: vectors, hybrid RRF and Prism metadata, the hosted tier, and
1789
+ `skills/gemchat/SKILL.md`.
1790
+
1791
+ ## 21. Risks
1792
+
1793
+ - **The Bundler plugin is a code-execution surface.** A checked-in
1794
+ `plugin "gemchat"` runs for anyone who clones the repo. Mitigated by the trust
1795
+ gate defaulting to a no-op print (§15.4), but it remains the sharpest edge
1796
+ here and warrants explicit review before any release.
1797
+ - **Plugin copy may drift from the dependency copy.** Resolved: the two copies
1798
+ are *confirmed* and nothing pins them (§15.4), so the answer is not to
1799
+ deduplicate but to make the plugin copy carry no logic — it requires the
1800
+ bundled copy, so drift is harmless by construction. The shim is not written
1801
+ yet (phase 7), so this is a design commitment, not a tested result.
1802
+ - ~~**ri generation cost is unmeasured.**~~ **Resolved in phase 1:** 9.9s
1803
+ across 49 real direct dependencies, 0.09s median, ~2s wall clock on 4 cores
1804
+ (§6). The resumable design is good hygiene rather than a necessity. The
1805
+ remaining open question is different and smaller — whether unchanged gems can
1806
+ be skipped *before* re-reading and re-chunking their ri (§20 phase 2).
1807
+ - **Version fidelity.** Serving the wrong version's ri defeats the tool's
1808
+ purpose; the exact-version glob is a correctness requirement.
1809
+ - **Hosted tier couples the gem to gemchat.org.** Hosted breaking changes reach
1810
+ the gem. Mitigate by treating gemchat.org as a versioned API and tolerating
1811
+ unknown parameters (§12.5).
1812
+ - **Silent escalation risk.** If the local path ever auto-falls-back to hosted,
1813
+ the tool becomes a privacy leak. The "diagnose and stop" rule (§12.4) is a
1814
+ hard requirement, not a default.
1815
+ - **Windows** has no prebuilt `rllama` binary. Requires an explicit "use
1816
+ `--embedder ollama` or `--embedder openai`" message — never a crash.
1817
+ - **300MB model download** is the largest user-visible cost. BM25 without models
1818
+ is the mitigation.
1819
+ - **RDoc `SystemStackError`** on pathological gems.
1820
+ - **Name collision with the hosted service.** Same name, two products. Mitigated
1821
+ only by documentation discipline (§19).
1822
+ - **Drift** from `gemchat_app`'s parsing and ranking logic, which is read-only
1823
+ reference. Accept divergence when the local store's constraints differ;
1824
+ re-sync deliberately.
1825
+
1826
+ ## 22. Out of scope
1827
+
1828
+ - Cross-gem discovery, alternatives, gem guides
1829
+ - MCP **server** (phase 8, deferred)
1830
+ - Answer generation / chat — this tool returns chunks and citations
1831
+ - Cross-project global search — the manifest scopes queries to the bundle
1832
+ - Windows support for the default embedder (opt-in providers only)
1833
+ - Multi-user or shared index stores
1834
+ - Automatic background indexing while no bundle command runs
1835
+
1836
+ ## 23. Open questions
1837
+
1838
+ - ~~Per-gem ri generation cost on a realistic bundle.~~ **Answered in phase 1:**
1839
+ 9.9s for 49 direct deps, 0.09s median, ~2s on 4 cores (§6). Carried forward as
1840
+ a narrower question in phase 2: whether unchanged gems can be skipped before
1841
+ re-reading and re-chunking their ri, rather than after.
1842
+ - ~~Should `vsearch` warn when the store has no vectors yet, or silently fall back
1843
+ to BM25?~~ **Answered by the shipped code: it warns and exits 1.** "no vectors
1844
+ yet → run `gemchat embed`" on stderr, with the reminder that `gemchat search`
1845
+ still works. The silent-fallback option is the one §12 explicitly rejects for
1846
+ the hosted tier — "a user who asked for hosted results and got local ones
1847
+ cannot tell the difference" — and the same reasoning applies here: `vsearch`
1848
+ would be answering a different question than the one asked.
1849
+ - ~~Should a `--all` run hard-cap per-gem ri chunks?~~ **Answered 2026-09-27: no,
1850
+ and the framing was wrong.** Measured across 120 local gems, 33 with ri:
1851
+
1852
+ | | chunks |
1853
+ |---|---|
1854
+ | median gem | 32 |
1855
+ | p90 | 32 |
1856
+ | max (`parser`) | 8,631 |
1857
+ | total ri | 21,253 |
1858
+
1859
+ The distribution is not a long tail, it is bimodal: a median of 32 and a
1860
+ maximum of 40% of the corpus in one gem. A cap of 5,000 would keep 82.7% of
1861
+ chunks, i.e. **delete 3,631 real methods of `parser`** — not "a few percent of
1862
+ one gem" as the question assumed, but most of a gem the user has installed,
1863
+ and silently.
1864
+
1865
+ What the cap was meant to bound, measured on that same 22,141-chunk index:
1866
+
1867
+ | | |
1868
+ |---|---|
1869
+ | `gemchat index` | 20s |
1870
+ | `gemchat embed` | 494s (8.2min), **22,134 of 22,141 embedded, 7 pending** |
1871
+ | vector query | 915ms (§11.1) |
1872
+ | store size | 13.5MB |
1873
+
1874
+ 99.97% embed success on a 94x larger corpus than anything else here was
1875
+ measured on, which is the strongest single validation of the supervised-child
1876
+ design — it was built for a cliff that could not be located, and it absorbed
1877
+ 22k chunks without one.
1878
+
1879
+ So there is no failure for a cap to prevent, and implementing one would trade a
1880
+ large, silent loss of real API surface for a bound nobody needs. **No default
1881
+ cap.** If a pathological gem is ever observed — a vendored mega-gem, say —
1882
+ the trigger is visible in these numbers: a single gem above ~10,000 chunks is
1883
+ 5% of this corpus and worth capping then, loudly and overridably, not silently
1884
+ and now.
1885
+ - **Whether Tier 3 (core/stdlib) belongs in v1.** **Answered 2026-09-27: not in
1886
+ v1.** The argument for is that it is genuinely free — 727 stdlib `.rb` files
1887
+ under `RbConfig::CONFIG["rubylibdir"]`, no download, no dependency, ~20s to
1888
+ index by the same recipe. The argument against is scope: the corpus is defined
1889
+ as *the gems in your bundle*, and stdlib is not in the bundle. Indexing it
1890
+ answers questions about methods the project does not use, diluting every
1891
+ result with 727 files' worth of otherwise-unreferenced API. Answering
1892
+ "what does `FileUtils#split_all` do" for a gem you do not have installed is
1893
+ what the hosted tier exists for. Revisit if someone asks; the recipe is in
1894
+ §8.4 and nothing else has to change.
1895
+ - **Whether to add a `gemchat changelog <gem>` subcommand.** **Deferred, and the
1896
+ plan's reasoning is now wrong.** It said the hosted `read_changelog` tool
1897
+ "supports it for free" — true when the tier spoke MCP, false now. The tier is a
1898
+ row API (§12.6), so a changelog subcommand needs a second endpoint in
1899
+ `HostedQuery` plus a new subcommand, and neither is free. Worth doing when
1900
+ something else needs the endpoint.