pi-gauntlet 5.21.0 → 5.23.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,25 @@
1
1
  # Changelog
2
2
 
3
+ ## v5.23.0 - 2026-09-27
4
+
5
+ ### Added
6
+
7
+ - `gauntlet-spec-index` gains `--corpus specs|docs` (default `specs`, output unchanged). `--corpus docs` indexes project documentation - tracked and untracked non-gitignored `*.md` files matching `**/doc/**/*.md`, `**/docs/**/*.md`, `README.md`, `AGENTS.md`, replaceable through `- docs: <glob>` bullets under `## Spec index` in the gauntlet overrides file - into a second FTS5 table in the same cache, minus `doc(s)/specs`, `doc(s)/plans`, `node_modules`, `.pi/gauntlet`, `.worktrees`, symlinks and context drafts, and prints `score`, `path`, `title`, `snippet`. Both corpora share one confidence rule driven by a corpus descriptor (strong fields `title`/`goal` for specs, `title`/`headings` for docs). The brainstorming scout runs the docs query and renders `Docs touched: <path> - <section heading>` lines that round 2 weighs as documentation-impact candidates.
8
+
9
+ ### Changed
10
+
11
+ - Spec-index cache schema is 3 (the `files` table is keyed by `(corpus, path)`); existing caches rebuild on first query. `#`/`##`/`###` lines inside fenced code blocks no longer count as the title or headings in either corpus.
12
+
13
+ ## v5.22.0 - 2026-09-26
14
+
15
+ ### Added
16
+
17
+ - `gauntlet-spec-index` gains a `state` column after `shipped_at` (`superseded` for a `- fully` banner directly under the title, `live` otherwise) and a repeatable `--exclude <path>` flag. `/skill:brainstorming` runs a second predecessor query at spec self-review from the finished spec's title, goal, and H2 headings and lists new `live` candidates at the review gate; the scout composes its first query from the ticket title and body.
18
+
19
+ ### Changed
20
+
21
+ - `gauntlet-spec-index` applies a confidence rule before `--limit`: only terms in fewer than half the specs count as evidence, a row needs two distinct evidence terms in title or goal, and rows under half the best live score are dropped, so queries with no real match return zero rows. Live rows sort before superseded rows. The cache schema is version 2 and rebuilds on first use.
22
+
3
23
  ## v5.21.0 - 2026-09-26
4
24
 
5
25
  ### Added
package/README.md CHANGED
@@ -137,7 +137,7 @@ Pin an exact release with `npm:pi-gauntlet@X.Y.Z`. See [doc/install-internals.md
137
137
 
138
138
  ## Spec search index
139
139
 
140
- `gauntlet-spec-index` provides lexical search across `doc/specs/*.md` at the repository root and one service level down. From a repository worktree, run `node <pi-gauntlet-package>/bin/gauntlet-spec-index.mjs --query "<text>" [--limit N]`; it requires Node >=24.15.0, refreshes its FTS5 index on every query, and prints tab-separated `score`, `path`, `service`, `title`, `status`, `shipped_at`, `files`, and `snippet` columns. The `files` column is a `;`-separated list of repo-relative paths the spec's shipped change modified and that still exist in the repository, the literal `missing` when the spec's telemetry record has no `derived.modified_files` list, or blank when there is no readable record or no recorded path remains. The per-worktree cache lives at `.pi/gauntlet/index.sqlite`, and its first creation adds `/.pi/gauntlet/index.sqlite*` to Git's `info/exclude` so the database and SQLite sidecars stay out of `git status`.
140
+ `gauntlet-spec-index` provides lexical search over two corpora in one per-worktree cache. From a repository worktree, run `node <pi-gauntlet-package>/bin/gauntlet-spec-index.mjs --query "<text>" [--corpus specs|docs] [--limit N] [--exclude <repo-relative path>]...`; it requires Node >=24.15.0 and refreshes its FTS5 index on every query. `--corpus specs` (the default) searches `doc/specs/*.md` at the repository root and one service level down and prints tab-separated `score`, `path`, `service`, `title`, `status`, `shipped_at`, `state`, `files`, and `snippet` columns. `--corpus docs` searches project documentation and prints `score`, `path`, `title`, and `snippet`; the snippet is the matched `##`/`###` heading line when one holds a query term, else a body fragment. Docs candidates are the tracked and untracked, non-gitignored `*.md` files (`git ls-files -co --exclude-standard`, so nested repositories and submodules are not indexed) matching an include list - default `**/doc/**/*.md`, `**/docs/**/*.md`, `README.md`, `AGENTS.md`, where a slash-free entry also matches one service level down - minus the fixed exclusions `**/doc/specs/**`, `**/docs/specs/**`, `**/doc/plans/**`, `**/docs/plans/**`, `**/node_modules/**`, `.pi/gauntlet/**`, `.worktrees/**`, symlinks, and files whose first line is the context-draft marker; `**` never matches a dot directory, so `.pi/*.md` or `.github/*.md` is indexed only through an explicit include. A project replaces the include list with `- docs: <glob>` bullets under a `## Spec index` heading in its gauntlet overrides file (see [Project-specific overrides](#project-specific-overrides)); the exclusions are not configurable. Both corpora pass the same confidence rule before `--limit` applies: a query term is evidence when it occurs in fewer than half the rows of the queried table (stem variants of one word count once); query tokens split on word boundaries, so `gauntlet-spec-index` searches as three words; a row is returned only when it matches at least two distinct evidence terms in the corpus's strong fields - `title` or `goal` for specs, `title` or `headings` for docs - and rows scoring below half the best live row's score are dropped (every docs row is live). A query with no real match therefore prints the header only and exits 0, and a corpus of two or fewer files never returns rows. `--exclude` (repeatable) removes a path before the rule runs, so the spec under review never sets the bar. `state` is `superseded` when a `> **Superseded by:** ... - fully` banner sits in the block directly under the title and `live` otherwise; named-section banners and consumer-defined banner syntaxes both read as `live`. Live rows sort before superseded rows, then by score. The `files` column is a `;`-separated list of repo-relative paths the spec's shipped change modified and that still exist in the repository, the literal `missing` when the spec's telemetry record has no `derived.modified_files` list, or blank when there is no readable record or no recorded path remains. The cache lives at `.pi/gauntlet/index.sqlite`, and its first creation adds `/.pi/gauntlet/index.sqlite*` to Git's `info/exclude` so the database and SQLite sidecars stay out of `git status`. `/skill:brainstorming` queries the specs corpus twice - the scout at gather time from the request, and the main loop at spec-writing from the finished spec's title, goal, and headings, surfacing new `live` candidates at the review gate - and the scout queries the docs corpus once to render `Docs touched:` lines that round 2 weighs as documentation-impact candidates.
141
141
 
142
142
  ## Performance digest
143
143
 
@@ -279,10 +279,18 @@ Use the project's wrapper: `script/worktree create <name>`. It provisions an iso
279
279
  database and copies `.env.local`. Never call `git worktree add` directly.
280
280
  ```
281
281
 
282
- Beyond `## conventions` and skill-named sections, headings a skill reads by name are documented in their owning skills; other headings retain topic/workflow-convention matching. The override file is read by skill instructions, not by the Pi runtime itself. Missing or empty `## conventions` adds no rules; a lower-priority file cannot supplement the selected file.
282
+ Beyond `## conventions` and skill-named sections, headings a skill reads by name are documented in their owning skills; other headings retain topic/workflow-convention matching. The override file is read by skill instructions and, for its `## Spec index` section only, by `gauntlet-spec-index`; the Pi runtime itself never reads it. Missing or empty `## conventions` adds no rules; a lower-priority file cannot supplement the selected file.
283
283
 
284
284
  **Discovery ladder:** skills check three locations, in order, and use the first one found - never merged: `.pi/gauntlet-overrides.md`, then `<repo root>/gauntlet-overrides.md`, then `<repo root>/doc/gauntlet-overrides.md` (`<repo root>` = `git rev-parse --show-toplevel`, or the current directory outside a repo). Pick one location per repo.
285
285
 
286
+ **`## Spec index` section:** `gauntlet-spec-index --corpus docs` reads `- docs: <glob>` bullets under this heading as the docs include list, replacing the default (`**/doc/**/*.md`, `**/docs/**/*.md`, `README.md`, `AGENTS.md`). Backticks or quotes around the glob are stripped; no other key under the heading is read, and the fixed exclusions always apply. For example:
287
+
288
+ ```markdown
289
+ ## Spec index
290
+ - docs: `**/doc/**/*.md`
291
+ - docs: `handbook/**/*.md`
292
+ ```
293
+
286
294
  **`## Issue tracker` section:** `shape-ticket` resolves tracker access through a capability ladder, and this is its first rung - it overrides the zero-config `gh` (GitHub) / `linearis` (Linear) defaults for any other tracker. Name the CLI's read, search, create, update, and post comment commands explicitly. For a Jira CLI, for example:
287
295
 
288
296
  ```markdown
@@ -1,28 +1,39 @@
1
1
  #!/usr/bin/env node
2
- // Lexical search over the spec corpus. Build-on-query: refresh a per-worktree
3
- // FTS5 cache by mtime+size, then rank with bm25 and join telemetry at output.
4
- import { readFileSync, appendFileSync, existsSync, statSync, readdirSync, mkdirSync, rmSync, realpathSync } from "node:fs";
5
- import { join, dirname, basename, isAbsolute } from "node:path";
2
+ // Lexical search over spec and docs corpora (FTS5, bm25, shared confidence rule); telemetry joined for specs only.
3
+ import { readFileSync, appendFileSync, existsSync, statSync, lstatSync, readdirSync, mkdirSync, rmSync, realpathSync, openSync, readSync, closeSync } from "node:fs";
4
+ import { join, dirname, basename, isAbsolute, matchesGlob } from "node:path";
6
5
  import { execFileSync } from "node:child_process";
7
6
  import process from "node:process";
8
7
  import { parse as parseYaml } from "yaml";
9
8
 
10
- const SCHEMA_VERSION = 1;
9
+ const SCHEMA_VERSION = 3;
10
+ const EVIDENCE_DF_FRACTION = 0.5;
11
+ const MIN_EVIDENCE_TOKENS = 2;
12
+ const SCORE_RATIO = 0.5;
13
+ // Default banner grammar of skills/brainstorming/reference/superseding.md; a named-section
14
+ // scope or anything after `- fully` keeps the spec live.
15
+ const FULLY_BANNER = /^> \*\*Superseded by:\*\* \[.*\]\(.*\) - fully$/;
11
16
  const MIN_NODE = [24, 15, 0];
12
17
  const DRAFT_MARKER = "# CONTEXT DRAFT - NOT A SPEC - fully replaced at spec-writing";
13
18
  const EXCLUDE_LINE = "/.pi/gauntlet/index.sqlite*";
14
19
  const SKIP_DIRS = new Set([".worktrees", "node_modules", "build"]);
15
- const HEADER = ["score", "path", "service", "title", "status", "shipped_at", "files", "snippet"];
20
+ const FENCE = /^\s*(`{3,}|~{3,})/;
21
+ const DOC_INCLUDES = ["**/doc/**/*.md", "**/docs/**/*.md", "README.md", "AGENTS.md"];
22
+ const DOC_EXCLUDES = ["**/doc/specs/**", "**/docs/specs/**", "**/doc/plans/**", "**/docs/plans/**", "**/node_modules/**", ".pi/gauntlet/**", ".worktrees/**"];
23
+ const OVERRIDES_FILES = [".pi/gauntlet-overrides.md", "gauntlet-overrides.md", "doc/gauntlet-overrides.md"];
16
24
  const SCHEMA = `
17
25
  CREATE TABLE IF NOT EXISTS meta (key TEXT PRIMARY KEY, value TEXT);
18
- CREATE TABLE IF NOT EXISTS files (path TEXT PRIMARY KEY, mtime_ms INTEGER, size INTEGER);
26
+ CREATE TABLE IF NOT EXISTS files (corpus TEXT NOT NULL, path TEXT NOT NULL, mtime_ms INTEGER, size INTEGER, PRIMARY KEY (corpus, path));
19
27
  CREATE VIRTUAL TABLE IF NOT EXISTS specs USING fts5(
20
- path UNINDEXED, service UNINDEXED, title, goal, headings, body,
28
+ path UNINDEXED, service UNINDEXED, title, goal, headings, body, state UNINDEXED,
29
+ tokenize = 'porter unicode61');
30
+ CREATE VIRTUAL TABLE IF NOT EXISTS docs USING fts5(
31
+ path UNINDEXED, title, headings, body,
21
32
  tokenize = 'porter unicode61');
22
33
  `;
23
34
 
24
35
  const usage = () => {
25
- process.stderr.write('usage: gauntlet-spec-index --query "<text>" [--limit N]\n');
36
+ process.stderr.write('usage: gauntlet-spec-index --query "<text>" [--corpus specs|docs] [--limit N] [--exclude <repo-relative path>]...\n');
26
37
  process.exit(1);
27
38
  };
28
39
  const die = (msg) => {
@@ -32,14 +43,18 @@ const die = (msg) => {
32
43
 
33
44
  function parseArgs(argv) {
34
45
  let query;
46
+ let corpus = "specs";
35
47
  let limit = 10;
48
+ const exclude = [];
36
49
  for (let i = 0; i < argv.length; i++) {
37
50
  if (argv[i] === "--query" && argv[i + 1] !== undefined) query = argv[++i];
51
+ else if (argv[i] === "--corpus" && Object.hasOwn(CORPORA, argv[i + 1] ?? "")) corpus = argv[++i];
38
52
  else if (argv[i] === "--limit" && /^[1-9]\d*$/.test(argv[i + 1] ?? "")) limit = Number(argv[++i]);
53
+ else if (argv[i] === "--exclude" && argv[i + 1] !== undefined) exclude.push(argv[++i]);
39
54
  else usage();
40
55
  }
41
56
  if (query === undefined) usage();
42
- return { query, limit };
57
+ return { query, corpus, limit, exclude };
43
58
  }
44
59
 
45
60
  function nodeOk() {
@@ -80,11 +95,64 @@ function discover(root) {
80
95
  }
81
96
  };
82
97
  collect("doc/specs", "root");
83
- const entries = readdirSync(root, { withFileTypes: true }).sort((a, b) => a.name.localeCompare(b.name));
84
- for (const e of entries) {
85
- if (e.name.startsWith(".") || SKIP_DIRS.has(e.name)) continue;
86
- if (!e.isDirectory() && !e.isSymbolicLink()) continue;
87
- collect(`${e.name}/doc/specs`, e.name);
98
+ for (const s of services(root).sort((a, b) => a.localeCompare(b))) collect(`${s}/doc/specs`, s);
99
+ return out;
100
+ }
101
+
102
+ // Top-level directories discover() would scan; a slash-free include also matches one level under them.
103
+ function services(root) {
104
+ return readdirSync(root, { withFileTypes: true })
105
+ .filter((e) => (e.isDirectory() || e.isSymbolicLink()) && !e.name.startsWith(".") && !SKIP_DIRS.has(e.name))
106
+ .map((e) => e.name);
107
+ }
108
+
109
+ function includes(root) {
110
+ const file = OVERRIDES_FILES.map((rel) => join(root, rel)).find((abs) => existsSync(abs));
111
+ if (!file) return DOC_INCLUDES;
112
+ const globs = [];
113
+ let inSection = false;
114
+ let text;
115
+ try {
116
+ text = readFileSync(file, "utf8");
117
+ } catch (e) {
118
+ die(`gauntlet-spec-index: unreadable overrides file ${file}: ${e.message}`);
119
+ }
120
+ for (const line of text.split(/\r?\n/)) {
121
+ if (/^## /.test(line)) { inSection = line.trim() === "## Spec index"; continue; }
122
+ if (!inSection) continue;
123
+ const m = /^-\s+docs:\s*(.+?)\s*$/.exec(line);
124
+ if (m) globs.push(m[1].replace(/^[`"']+|[`"']+$/g, ""));
125
+ }
126
+ return globs.length ? globs : DOC_INCLUDES;
127
+ }
128
+
129
+ function firstLine(abs) {
130
+ const fd = openSync(abs, "r");
131
+ try {
132
+ const buf = Buffer.alloc(128);
133
+ const n = readSync(fd, buf, 0, 128, 0);
134
+ return buf.toString("utf8", 0, n).split(/\r?\n/, 1)[0];
135
+ } finally {
136
+ closeSync(fd);
137
+ }
138
+ }
139
+
140
+ function discoverDocs(root, globs) {
141
+ const svc = services(root);
142
+ const patterns = globs.flatMap((g) => (g.includes("/") ? [g] : [g, ...svc.map((s) => `${s}/${g}`)]));
143
+ const listed = execFileSync("git", ["-C", root, "ls-files", "-co", "--exclude-standard", "-z", "--", "*.md"], { encoding: "utf8", stdio: ["ignore", "pipe", "ignore"] }).split("\0").filter(Boolean);
144
+ const out = [];
145
+ for (const rel of listed) {
146
+ if (!patterns.some((p) => matchesGlob(rel, p))) continue;
147
+ if (DOC_EXCLUDES.some((p) => matchesGlob(rel, p))) continue;
148
+ let st;
149
+ try {
150
+ st = lstatSync(join(root, rel));
151
+ if (!st.isFile() || firstLine(join(root, rel)) === DRAFT_MARKER) continue;
152
+ } catch {
153
+ continue;
154
+ }
155
+ out.push({ path: rel, mtime_ms: Math.trunc(st.mtimeMs), size: st.size });
88
156
  }
89
157
  return out;
90
158
  }
@@ -137,33 +205,60 @@ async function openDb(root) {
137
205
  }
138
206
  }
139
207
 
208
+ function fenceMask(lines) {
209
+ const mask = new Array(lines.length).fill(false);
210
+ let open = null;
211
+ for (let i = 0; i < lines.length; i++) {
212
+ const m = FENCE.exec(lines[i]);
213
+ if (open) {
214
+ mask[i] = true;
215
+ if (m && m[1][0] === open.ch && m[1].length >= open.len) open = null;
216
+ } else if (m) {
217
+ mask[i] = true;
218
+ open = { ch: m[1][0], len: m[1].length };
219
+ }
220
+ }
221
+ return mask;
222
+ }
223
+
140
224
  function extract(text, path) {
141
225
  const lines = text.split(/\r?\n/);
142
- const title = lines.find((l) => l.startsWith("# "))?.slice(2).trim() || basename(path, ".md");
226
+ const fenced = fenceMask(lines);
227
+ const h1 = lines.findIndex((l, i) => !fenced[i] && l.startsWith("# "));
228
+ const title = (h1 >= 0 ? lines[h1].slice(2).trim() : "") || basename(path, ".md");
143
229
  const goal = lines.find((l) => l.startsWith("**Goal:**"))?.slice("**Goal:**".length).trim() ?? "";
144
- const headings = lines.filter((l) => /^##{1,2} /.test(l)).map((l) => l.replace(/^#+ /, "")).join("\n");
145
- return { title, goal, headings };
230
+ const headings = lines.filter((l, i) => !fenced[i] && /^##{1,2} /.test(l)).map((l) => l.replace(/^#+ /, "")).join("\n");
231
+ let state = "live";
232
+ if (h1 >= 0) {
233
+ // Title block: everything after the H1 up to the first H2 heading or code fence.
234
+ for (const l of lines.slice(h1 + 1)) {
235
+ if (/^## /.test(l) || FENCE.test(l)) break;
236
+ if (FULLY_BANNER.test(l)) { state = "superseded"; break; }
237
+ }
238
+ }
239
+ return { title, goal, headings, state };
146
240
  }
147
241
 
148
242
  function refresh(db, root, corpus) {
149
- const known = new Map(db.prepare("SELECT path, mtime_ms, size FROM files").all().map((r) => [r.path, r]));
150
- const present = new Set(corpus.map((f) => f.path));
151
- const delSpec = db.prepare("DELETE FROM specs WHERE path = ?");
152
- const delFile = db.prepare("DELETE FROM files WHERE path = ?");
153
- const insSpec = db.prepare("INSERT INTO specs (path, service, title, goal, headings, body) VALUES (?, ?, ?, ?, ?, ?)");
154
- const putFile = db.prepare("INSERT OR REPLACE INTO files (path, mtime_ms, size) VALUES (?, ?, ?)");
243
+ const { table } = corpus;
244
+ const found = corpus.discover(root);
245
+ const known = new Map(db.prepare("SELECT path, mtime_ms, size FROM files WHERE corpus = ?").all(table).map((r) => [r.path, r]));
246
+ const present = new Set(found.map((f) => f.path));
247
+ const delRow = db.prepare(`DELETE FROM ${table} WHERE path = ?`);
248
+ const delFile = db.prepare("DELETE FROM files WHERE corpus = ? AND path = ?");
249
+ const insRow = db.prepare(`INSERT INTO ${table} (${corpus.columns.join(", ")}) VALUES (${corpus.columns.map(() => "?").join(", ")})`);
250
+ const putFile = db.prepare("INSERT OR REPLACE INTO files (corpus, path, mtime_ms, size) VALUES (?, ?, ?, ?)");
155
251
  db.exec("BEGIN IMMEDIATE");
156
252
  try {
157
- for (const path of known.keys()) if (!present.has(path)) { delSpec.run(path); delFile.run(path); }
158
- for (const f of corpus) {
253
+ for (const path of known.keys()) if (!present.has(path)) { delRow.run(path); delFile.run(table, path); }
254
+ for (const f of found) {
159
255
  const k = known.get(f.path);
160
256
  if (k && k.mtime_ms === f.mtime_ms && k.size === f.size) continue;
161
257
  const text = readFileSync(join(root, f.path), "utf8");
162
- delSpec.run(f.path);
163
- if (text.split(/\r?\n/, 1)[0] === DRAFT_MARKER) { delFile.run(f.path); continue; }
164
- const x = extract(text, f.path);
165
- insSpec.run(f.path, f.service, x.title, x.goal, x.headings, text);
166
- putFile.run(f.path, f.mtime_ms, f.size);
258
+ delRow.run(f.path);
259
+ if (text.split(/\r?\n/, 1)[0] === DRAFT_MARKER) { delFile.run(table, f.path); continue; }
260
+ insRow.run(...corpus.row(text, f));
261
+ putFile.run(table, f.path, f.mtime_ms, f.size);
167
262
  }
168
263
  db.exec("COMMIT");
169
264
  } catch (e) {
@@ -172,8 +267,58 @@ function refresh(db, root, corpus) {
172
267
  }
173
268
  }
174
269
 
175
- const toMatch = (query) =>
176
- query.split(/\s+/).filter((t) => t.length >= 2).map((t) => `"${t.replaceAll('"', '""')}"`).join(" OR ");
270
+ const tokens = (query) => query.split(/[^\p{L}\p{N}\p{Co}]+/u).filter((t) => t.length >= 2);
271
+ const quote = (t) => `"${t.replaceAll('"', '""')}"`;
272
+ const toMatch = (toks) => toks.map(quote).join(" OR ");
273
+
274
+ // Porter is not idempotent (manatee -> manate -> manat), so MATCH always receives an
275
+ // original token; the vocab pass only decides which tokens are the same term.
276
+ function representatives(db, toks) {
277
+ const mem = new db.constructor(":memory:");
278
+ try {
279
+ mem.exec("CREATE VIRTUAL TABLE q USING fts5(t, tokenize = 'porter unicode61'); CREATE VIRTUAL TABLE qi USING fts5vocab(q, 'instance')");
280
+ const ins = mem.prepare("INSERT INTO q (rowid, t) VALUES (?, ?)");
281
+ toks.forEach((t, i) => ins.run(i + 1, t));
282
+ const rep = new Map();
283
+ for (const r of mem.prepare("SELECT term, doc FROM qi ORDER BY doc, offset").all()) if (!rep.has(r.term)) rep.set(r.term, toks[r.doc - 1]);
284
+ return rep;
285
+ } finally {
286
+ mem.close();
287
+ }
288
+ }
289
+
290
+ function confidenceFilter(rows, evidence) {
291
+ return rows.filter((r) => {
292
+ const hits = evidence.filter((e) => e.all.has(r.rowid)).length;
293
+ const strong = evidence.filter((e) => e.strong.has(r.rowid)).length;
294
+ return hits >= MIN_EVIDENCE_TOKENS && strong >= MIN_EVIDENCE_TOKENS;
295
+ });
296
+ }
297
+
298
+ function query(db, corpus, toks, { limit, exclude }) {
299
+ const { table } = corpus;
300
+ const n = db.prepare(`SELECT count(*) AS n FROM ${table}`).get().n;
301
+ const rowids = (match) => new Set(db.prepare(`SELECT rowid FROM ${table} WHERE ${table} MATCH ?`).all(match).map((r) => r.rowid));
302
+ const reps = representatives(db, toks);
303
+ const terms = [...new Set(reps.values())];
304
+ const evidence = [];
305
+ for (const tok of terms) {
306
+ const all = rowids(quote(tok));
307
+ if (all.size === 0 || all.size >= n * EVIDENCE_DF_FRACTION) continue;
308
+ evidence.push({ all, strong: rowids(corpus.strong.map((f) => `${f}:${quote(tok)}`).join(" OR ")) });
309
+ }
310
+ if (evidence.length < MIN_EVIDENCE_TOKENS) return [];
311
+ const rows = db.prepare(
312
+ `SELECT rowid, ${corpus.select}, bm25(${table}, ${corpus.weights.map((w) => w.toFixed(1)).join(", ")}) AS score
313
+ FROM ${table} WHERE ${table} MATCH ?`,
314
+ ).all(toMatch(terms)).filter((r) => !exclude.includes(r.path));
315
+ const kept = confidenceFilter(rows, evidence);
316
+ const live = (r) => !corpus.hasState || r.state === "live";
317
+ kept.sort((a, b) => Number(!live(a)) - Number(!live(b)) || a.score - b.score);
318
+ const best = kept.find(live);
319
+ const cut = best ? kept.filter((r) => Math.abs(r.score) >= SCORE_RATIO * Math.abs(best.score)) : kept;
320
+ return cut.slice(0, limit);
321
+ }
177
322
 
178
323
  function telemetry(root, specPath) {
179
324
  const blank = { status: null, shipped_at: null, files: "" };
@@ -198,26 +343,60 @@ const cell = (v) => (v === null || v === undefined ? "" : String(v).replace(/\s+
198
343
  // Paths keep their spaces; only column and row delimiters are neutralised.
199
344
  const filesCell = (v) => v.replace(/[\t\r\n]+/g, " ");
200
345
 
346
+ function docSnippet(r) {
347
+ if (!r.hsnip.includes("\u0001")) return r.bsnip;
348
+ const fragment = r.hsnip.split("\n").find((line) => line.includes("\u0001")).replace(/[\u0001\u0002]/g, "");
349
+ return r.headings.split("\n").find((line) => line.includes(fragment));
350
+ }
351
+
352
+ const CORPORA = {
353
+ specs: {
354
+ table: "specs",
355
+ columns: ["path", "service", "title", "goal", "headings", "body", "state"],
356
+ strong: ["title", "goal"],
357
+ weights: [0, 0, 10, 5, 2, 1, 0],
358
+ hasState: true,
359
+ header: ["score", "path", "service", "title", "status", "shipped_at", "state", "files", "snippet"],
360
+ select: "path, service, title, state, snippet(specs, 5, '', '', '...', 12) AS snippet",
361
+ discover,
362
+ row: (text, f) => {
363
+ const x = extract(text, f.path);
364
+ return [f.path, f.service, x.title, x.goal, x.headings, text, x.state];
365
+ },
366
+ format: (root, r) => {
367
+ const t = telemetry(root, r.path);
368
+ return [...[r.score.toFixed(3), r.path, r.service, r.title, t.status, t.shipped_at, r.state].map(cell), filesCell(t.files), cell(r.snippet)];
369
+ },
370
+ },
371
+ docs: {
372
+ table: "docs",
373
+ columns: ["path", "title", "headings", "body"],
374
+ strong: ["title", "headings"],
375
+ weights: [0, 10, 3, 1],
376
+ hasState: false,
377
+ header: ["score", "path", "title", "snippet"],
378
+ select: "path, title, headings, snippet(docs, 2, '\u0001', '\u0002', '', 12) AS hsnip, snippet(docs, 3, '', '', '...', 12) AS bsnip",
379
+ discover: (root) => discoverDocs(root, includes(root)),
380
+ row: (text, f) => {
381
+ const x = extract(text, f.path);
382
+ return [f.path, x.title, x.headings, text];
383
+ },
384
+ format: (root, r) => [r.score.toFixed(3), r.path, r.title, docSnippet(r)].map(cell),
385
+ },
386
+ };
387
+
201
388
  async function main() {
202
389
  if (!nodeOk()) die(`gauntlet-spec-index needs Node >=24.15.0 (found ${process.versions.node})`);
203
- const { query, limit } = parseArgs(process.argv.slice(2));
204
- const match = toMatch(query);
205
- if (!match) usage();
390
+ const { query: text, corpus: name, limit, exclude } = parseArgs(process.argv.slice(2));
391
+ const toks = tokens(text);
392
+ if (!toks.length) usage();
206
393
  const root = repoRoot();
394
+ const corpus = CORPORA[name];
207
395
  const db = await openDb(root);
208
- refresh(db, root, discover(root));
209
- const rows = db.prepare(
210
- `SELECT path, service, title,
211
- bm25(specs, 0, 0, 10.0, 5.0, 2.0, 1.0) AS score,
212
- snippet(specs, 5, '', '', '...', 12) AS snippet
213
- FROM specs WHERE specs MATCH ?
214
- ORDER BY score LIMIT ?`,
215
- ).all(match, limit);
216
- const out = [HEADER.join("\t")];
217
- for (const r of rows) {
218
- const t = telemetry(root, r.path);
219
- out.push([...[r.score.toFixed(3), r.path, r.service, r.title, t.status, t.shipped_at].map(cell), filesCell(t.files), cell(r.snippet)].join("\t"));
220
- }
396
+ refresh(db, root, corpus);
397
+ const rows = query(db, corpus, toks, { limit, exclude });
398
+ const out = [corpus.header.join("\t")];
399
+ for (const r of rows) out.push(corpus.format(root, r).join("\t"));
221
400
  process.stdout.write(out.join("\n") + "\n");
222
401
  db.close();
223
402
  }
@@ -1,6 +1,6 @@
1
1
  import { test } from "node:test";
2
2
  import assert from "node:assert/strict";
3
- import { mkdtempSync, mkdirSync, writeFileSync, readFileSync, rmSync, utimesSync } from "node:fs";
3
+ import { mkdtempSync, mkdirSync, writeFileSync, readFileSync, rmSync, utimesSync, symlinkSync } from "node:fs";
4
4
  import { join, dirname } from "node:path";
5
5
  import { tmpdir } from "node:os";
6
6
  import { spawnSync } from "node:child_process";
@@ -9,30 +9,58 @@ import { DatabaseSync } from "node:sqlite";
9
9
 
10
10
  const CLI = join(dirname(fileURLToPath(import.meta.url)), "gauntlet-spec-index.mjs");
11
11
  const DRAFT = "# CONTEXT DRAFT - NOT A SPEC - fully replaced at spec-writing";
12
+ const HEADER = ["score", "path", "service", "title", "status", "shipped_at", "state", "files", "snippet"];
13
+ const DOCS_HEADER = ["score", "path", "title", "snippet"];
14
+ const docFiller = (n) => `# Filler ${n}\n\n## Design\n\nnothing of note here\n`;
15
+ const decoy = "# ptarmigan gannet decoy\n\n## ptarmigan gannet\n\nbody\n";
16
+ const docsRepo = () => {
17
+ const root = gitRepo();
18
+ write(root, "doc/guide.md", "# Guide\n\n## Configuring the ptarmigan gannet\n\nprose about setup\n");
19
+ write(root, "README.md", "# Repo\n\nThe body mentions ptarmigan once.\n");
20
+ for (let n = 1; n <= 4; n++) write(root, `doc/filler-${n}.md`, docFiller(n));
21
+ for (const p of ["doc/specs/a.md", "docs/specs/b.md", "doc/plans/p.md", "docs/plans/q.md", "node_modules/x/doc/a.md", "build/doc/b.md", ".pi/gauntlet/n.md", ".worktrees/w/doc/g.md"]) write(root, p, decoy);
22
+ write(root, ".gitignore", "build/\n");
23
+ write(root, "doc/draft.md", `${DRAFT}\n\n## ptarmigan gannet\n`);
24
+ commit(root);
25
+ return root;
26
+ };
27
+ const docsPaths = (root) => withDb(root, (db) => db.prepare("SELECT path FROM docs ORDER BY path").all().map((r) => r.path));
28
+ const overrides = (root, globs) => write(root, ".pi/gauntlet-overrides.md", `## Issue tracker\ntracker: none\n\n## Spec index\n${globs.map((g) => `- docs: \`${g}\``).join("\n")}\n\n## release\nnothing\n`);
12
29
 
13
30
  const write = (root, rel, text) => {
14
31
  mkdirSync(join(root, dirname(rel)), { recursive: true });
15
32
  writeFileSync(join(root, rel), text);
16
33
  };
17
34
 
18
- const repo = () => {
35
+ const gitRepo = () => {
19
36
  const root = mkdtempSync(join(tmpdir(), "gsi-"));
20
37
  spawnSync("git", ["init", "-q"], { cwd: root });
21
- write(root, "doc/specs/a.md", "# Alpha zephyr widget\n\n**Goal:** rank the widget.\n\n## Design\n\nbody text\n");
22
- write(root, "svc-a/doc/specs/b.md", "# Beta service\n\n**Goal:** unrelated.\n\nDeep in the body a zephyr appears.\n");
23
- write(root, ".worktrees/x/doc/specs/decoy1.md", "# zephyr decoy one\n");
24
- write(root, "build/doc/specs/decoy2.md", "# zephyr decoy two\n");
25
- write(root, "apps/svc/doc/specs/decoy3.md", "# zephyr decoy three\n");
26
- write(root, "doc/specs/draft.md", `${DRAFT}\n\nzephyr zephyr zephyr\n`);
38
+ return root;
39
+ };
40
+ const commit = (root) => {
27
41
  spawnSync("git", ["add", "-A"], { cwd: root });
28
42
  spawnSync("git", ["-c", "user.name=t", "-c", "user.email=t@t", "commit", "-q", "-m", "fixture"], { cwd: root });
43
+ };
44
+ // Fillers keep the corpus large enough that two-document probe terms stay below N/2.
45
+ const filler = (n) => `# Filler ${n}\n\n**Goal:** filler goal ${n}.\n\n## Design\n\nnothing of note here\n`;
46
+
47
+ const repo = () => {
48
+ const root = gitRepo();
49
+ write(root, "doc/specs/a.md", "# Alpha zephyr widget\n\n**Goal:** rank the widget.\n\n## Design\n\nbody text\n");
50
+ write(root, "svc-a/doc/specs/b.md", "# Beta service\n\n**Goal:** zephyr widget for the beta service.\n\nDeep in the body a zephyr appears.\n");
51
+ for (const n of ["one", "two", "three", "four"]) write(root, `doc/specs/filler-${n}.md`, filler(n));
52
+ write(root, ".worktrees/x/doc/specs/decoy1.md", "# zephyr widget decoy one\n");
53
+ write(root, "build/doc/specs/decoy2.md", "# zephyr widget decoy two\n");
54
+ write(root, "apps/svc/doc/specs/decoy3.md", "# zephyr widget decoy three\n");
55
+ write(root, "doc/specs/draft.md", `${DRAFT}\n\nzephyr widget zephyr widget\n`);
56
+ commit(root);
29
57
  return root;
30
58
  };
31
59
 
32
60
  const run = (cwd, args) => {
33
61
  const r = spawnSync(process.execPath, [CLI, ...args], { cwd, encoding: "utf8" });
34
62
  const lines = r.stdout.split("\n").filter(Boolean);
35
- return { status: r.status, stderr: r.stderr, header: lines[0]?.split("\t"), rows: lines.slice(1).map((l) => l.split("\t")) };
63
+ return { status: r.status, stderr: r.stderr, stdout: r.stdout, header: lines[0]?.split("\t"), rows: lines.slice(1).map((l) => l.split("\t")) };
36
64
  };
37
65
  const paths = (res) => res.rows.map((r) => r[1]);
38
66
  const withDb = (root, fn) => {
@@ -47,41 +75,39 @@ const withDb = (root, fn) => {
47
75
  test("1: corpus boundary, ordering, draft skip, git status clean, exclude written once", (t) => {
48
76
  const root = repo();
49
77
  t.after(() => rmSync(root, { recursive: true, force: true }));
50
- const r1 = run(root, ["--query", "zephyr"]);
78
+ const r1 = run(root, ["--query", "zephyr widget"]);
51
79
  assert.equal(r1.status, 0, r1.stderr);
52
80
  assert.deepEqual(paths(r1), ["doc/specs/a.md", "svc-a/doc/specs/b.md"]);
53
81
  assert.deepEqual(r1.rows.map((r) => r[2]), ["root", "svc-a"]);
54
82
  const status = spawnSync("git", ["status", "--porcelain"], { cwd: root, encoding: "utf8" }).stdout;
55
83
  assert.equal(status, "");
56
- run(root, ["--query", "zephyr"]);
84
+ run(root, ["--query", "zephyr widget"]);
57
85
  const exclude = readFileSync(join(root, ".git/info/exclude"), "utf8");
58
86
  assert.equal(exclude.split("\n").filter((l) => l === "/.pi/gauntlet/index.sqlite*").length, 1);
59
87
  });
60
88
 
61
- test("2: title outranks body; --limit 1; default up to 10", (t) => {
89
+ test("2: title outranks goal; --limit 1", (t) => {
62
90
  const root = repo();
63
91
  t.after(() => rmSync(root, { recursive: true, force: true }));
64
- for (let i = 0; i < 12; i++) write(root, `doc/specs/many-${i}.md`, `# Spec ${i}\n\nquokka\n`);
65
- assert.deepEqual(paths(run(root, ["--query", "zephyr", "--limit", "1"])), ["doc/specs/a.md"]);
66
- assert.equal(run(root, ["--query", "quokka"]).rows.length, 10);
92
+ assert.deepEqual(paths(run(root, ["--query", "zephyr widget", "--limit", "1"])), ["doc/specs/a.md"]);
67
93
  });
68
94
 
69
95
  test("3: incremental refresh updates only the edited row; delete removes rows", (t) => {
70
96
  const root = repo();
71
97
  t.after(() => rmSync(root, { recursive: true, force: true }));
72
- run(root, ["--query", "zephyr"]);
98
+ run(root, ["--query", "zephyr widget"]);
73
99
  const before = withDb(root, (database) => Object.fromEntries(database.prepare("SELECT path, mtime_ms, size FROM files").all().map((r) => [r.path, `${r.mtime_ms}:${r.size}`])));
74
100
  const b = join(root, "svc-a/doc/specs/b.md");
75
- writeFileSync(b, readFileSync(b, "utf8") + "\nwombat\n");
101
+ writeFileSync(b, "# Beta service wombat kudu\n\n**Goal:** zephyr widget for the beta service.\n\nDeep in the body a zephyr appears.\n");
76
102
  utimesSync(b, new Date(), new Date(Date.now() + 5000));
77
- assert.deepEqual(paths(run(root, ["--query", "wombat"])), ["svc-a/doc/specs/b.md"]);
103
+ assert.deepEqual(paths(run(root, ["--query", "wombat kudu"])), ["svc-a/doc/specs/b.md"]);
78
104
  const after = withDb(root, (database) => Object.fromEntries(database.prepare("SELECT path, mtime_ms, size FROM files").all().map((r) => [r.path, `${r.mtime_ms}:${r.size}`])));
79
105
  assert.equal(after["doc/specs/a.md"], before["doc/specs/a.md"]);
80
106
  assert.notEqual(after["svc-a/doc/specs/b.md"], before["svc-a/doc/specs/b.md"]);
81
107
  const count = (root, sql, path) => withDb(root, (database) => Object.values(database.prepare(sql).get(path))[0]);
82
108
  assert.equal(count(root, "SELECT count(*) FROM specs WHERE path = ?", "svc-a/doc/specs/b.md"), 1);
83
109
  rmSync(b);
84
- run(root, ["--query", "zephyr"]);
110
+ run(root, ["--query", "zephyr widget"]);
85
111
  assert.equal(count(root, "SELECT count(*) FROM specs WHERE path = ?", "svc-a/doc/specs/b.md"), 0);
86
112
  assert.equal(count(root, "SELECT count(*) FROM files WHERE path = ?", "svc-a/doc/specs/b.md"), 0);
87
113
  });
@@ -89,71 +115,70 @@ test("3: incremental refresh updates only the edited row; delete removes rows",
89
115
  test("4: schema_version mismatch rebuilds the db", (t) => {
90
116
  const root = repo();
91
117
  t.after(() => rmSync(root, { recursive: true, force: true }));
92
- run(root, ["--query", "zephyr"]);
118
+ run(root, ["--query", "zephyr widget"]);
93
119
  withDb(root, (database) => database.prepare("UPDATE meta SET value = '999' WHERE key = 'schema_version'").run());
94
- const r = run(root, ["--query", "zephyr"]);
120
+ const r = run(root, ["--query", "zephyr widget"]);
95
121
  assert.equal(r.status, 0, r.stderr);
96
122
  assert.equal(paths(r).length, 2);
97
- assert.equal(withDb(root, (database) => database.prepare("SELECT value FROM meta WHERE key = 'schema_version'").get().value), "1");
123
+ assert.equal(withDb(root, (database) => database.prepare("SELECT value FROM meta WHERE key = 'schema_version'").get().value), "3");
98
124
  });
99
125
 
100
126
  test("5: a non-SQLite database is rebuilt", (t) => {
101
127
  const root = repo();
102
128
  t.after(() => rmSync(root, { recursive: true, force: true }));
103
129
  write(root, ".pi/gauntlet/index.sqlite", "not a sqlite database");
104
- const r = run(root, ["--query", "zephyr"]);
130
+ const r = run(root, ["--query", "zephyr widget"]);
105
131
  assert.equal(r.status, 0, r.stderr);
106
132
  assert.deepEqual(paths(r), ["doc/specs/a.md", "svc-a/doc/specs/b.md"]);
107
- assert.equal(withDb(root, (database) => database.prepare("SELECT value FROM meta WHERE key = 'schema_version'").get().value), "1");
133
+ assert.equal(withDb(root, (database) => database.prepare("SELECT value FROM meta WHERE key = 'schema_version'").get().value), "3");
108
134
  });
109
135
 
110
136
  test("6: telemetry join is output-only and tolerant", (t) => {
111
137
  const root = repo();
112
138
  t.after(() => rmSync(root, { recursive: true, force: true }));
113
- run(root, ["--query", "zephyr"]);
139
+ run(root, ["--query", "zephyr widget"]);
114
140
  const filesBefore = withDb(root, (database) => JSON.stringify(database.prepare("SELECT * FROM files ORDER BY path").all()));
115
141
  write(root, "x", "x\n");
116
142
  write(root, "y", "y\n");
117
143
  write(root, ".pi/gauntlet/telemetry/doc/specs/a.yaml", "status: shipped\nshipped_at: 2026-09-17T10:00:00Z\nderived:\n modified_files:\n - x\n - y\n");
118
- let r = run(root, ["--query", "zephyr"]);
119
- assert.deepEqual(r.rows[0].slice(4, 7), ["shipped", "2026-09-17T10:00:00Z", "x;y"]);
120
- assert.equal(r.rows[0].length, 8);
121
- assert.equal(r.rows[1][6], "");
144
+ let r = run(root, ["--query", "zephyr widget"]);
145
+ assert.deepEqual(r.rows[0].slice(4, 8), ["shipped", "2026-09-17T10:00:00Z", "live", "x;y"]);
146
+ assert.equal(r.rows[0].length, 9);
147
+ assert.equal(r.rows[1][7], "");
122
148
  assert.equal(withDb(root, (database) => JSON.stringify(database.prepare("SELECT * FROM files ORDER BY path").all())), filesBefore);
123
149
  write(root, ".pi/gauntlet/telemetry/doc/specs/a.yaml", "status: in_progress\nshipped_at: 2026-09-18T10:00:00Z\n");
124
- r = run(root, ["--query", "zephyr"]);
125
- assert.deepEqual(r.rows[0].slice(4, 7), ["in_progress", "2026-09-18T10:00:00Z", "missing"]);
150
+ r = run(root, ["--query", "zephyr widget"]);
151
+ assert.deepEqual(r.rows[0].slice(4, 8), ["in_progress", "2026-09-18T10:00:00Z", "live", "missing"]);
126
152
  write(root, ".pi/gauntlet/telemetry/doc/specs/a.yaml", "status: [unclosed\n");
127
- r = run(root, ["--query", "zephyr"]);
153
+ r = run(root, ["--query", "zephyr widget"]);
128
154
  assert.equal(r.status, 0);
129
- assert.deepEqual(r.rows[0].slice(4, 7), ["", "", ""]);
155
+ assert.deepEqual(r.rows[0].slice(4, 8), ["", "", "live", ""]);
130
156
  assert.match(r.stderr, /a\.yaml/);
131
157
  });
132
158
 
133
159
  test("6b: files cell is the AC fixture - present paths kept, gone paths dropped, no list is missing", (t) => {
134
- const root = mkdtempSync(join(tmpdir(), "gsi-"));
160
+ const root = gitRepo();
135
161
  t.after(() => rmSync(root, { recursive: true, force: true }));
136
- spawnSync("git", ["init", "-q"], { cwd: root });
137
- write(root, "doc/specs/a.md", "# Alpha quokka\n\n**Goal:** quokka.\n");
138
- write(root, "doc/specs/b.md", "# Beta quokka\n\n**Goal:** quokka too.\n");
162
+ write(root, "doc/specs/a.md", "# Alpha quokka wallaby\n\n**Goal:** quokka.\n");
163
+ write(root, "doc/specs/b.md", "# Beta quokka wallaby\n\n**Goal:** quokka too.\n");
164
+ for (const n of ["one", "two", "three"]) write(root, `doc/specs/filler-${n}.md`, filler(n));
139
165
  write(root, "src/x.ts", "export {};\n");
140
166
  write(root, ".pi/gauntlet/telemetry/doc/specs/a.yaml", "status: shipped\nderived:\n modified_files:\n - src/x.ts\n - src/gone.ts\n");
141
167
  write(root, ".pi/gauntlet/telemetry/doc/specs/b.yaml", "status: shipped\n");
142
- spawnSync("git", ["add", "-A"], { cwd: root });
143
- spawnSync("git", ["-c", "user.name=t", "-c", "user.email=t@t", "commit", "-q", "-m", "fixture"], { cwd: root });
144
- const r = run(root, ["--query", "quokka"]);
168
+ commit(root);
169
+ const r = run(root, ["--query", "quokka wallaby"]);
145
170
  assert.equal(r.status, 0, r.stderr);
146
- const cells = Object.fromEntries(r.rows.map((row) => [row[1], row[6]]));
171
+ const cells = Object.fromEntries(r.rows.map((row) => [row[1], row[7]]));
147
172
  assert.equal(cells["doc/specs/a.md"], "src/x.ts");
148
173
  assert.equal(cells["doc/specs/b.md"], "missing");
149
- for (const row of r.rows) assert.equal(row.length, 8);
174
+ for (const row of r.rows) assert.equal(row.length, 9);
150
175
  });
151
176
 
152
177
  test("6c: files cell edge cases - order, spaces kept, all gone, non-array, null record", (t) => {
153
178
  const root = repo();
154
179
  t.after(() => rmSync(root, { recursive: true, force: true }));
155
180
  const yaml = ".pi/gauntlet/telemetry/doc/specs/a.yaml";
156
- const files = () => run(root, ["--query", "zephyr"]).rows[0][6];
181
+ const files = () => run(root, ["--query", "zephyr widget"]).rows[0][7];
157
182
  write(root, "b.js", "");
158
183
  write(root, "a.js", "");
159
184
  write(root, "sp aced.txt", "");
@@ -162,7 +187,7 @@ test("6c: files cell edge cases - order, spaces kept, all gone, non-array, null
162
187
  write(root, "tab\there.js", "");
163
188
  write(root, yaml, 'status: shipped\nderived:\n modified_files:\n - a.js\n - "tab\\there.js"\n');
164
189
  assert.equal(files(), "a.js;tab here.js");
165
- assert.equal(run(root, ["--query", "zephyr"]).rows[0].length, 8);
190
+ assert.equal(run(root, ["--query", "zephyr widget"]).rows[0].length, 9);
166
191
  write(root, yaml, "status: shipped\nderived:\n modified_files:\n - 'sp aced.txt'\n - 7\n");
167
192
  assert.equal(files(), "sp aced.txt");
168
193
  write(root, yaml, "status: shipped\nderived:\n modified_files:\n - nope1\n - nope2\n");
@@ -178,45 +203,358 @@ test("6c: files cell edge cases - order, spaces kept, all gone, non-array, null
178
203
  test("7: query sanitising tolerates embedded quotes", (t) => {
179
204
  const root = repo();
180
205
  t.after(() => rmSync(root, { recursive: true, force: true }));
181
- write(root, "doc/specs/q.md", '# Quotes\n\nfoo"bar and "plain" words\n');
182
- let r = run(root, ["--query", 'foo"bar']);
206
+ write(root, "doc/specs/q.md", '# Quotes foo"bar ocelot\n\n**Goal:** "plain" ocelot words\n');
207
+ let r = run(root, ["--query", 'foo"bar ocelot']);
183
208
  assert.equal(r.status, 0, r.stderr);
184
209
  assert.ok(paths(r).includes("doc/specs/q.md"));
185
- r = run(root, ["--query", '"plain"']);
210
+ r = run(root, ["--query", '"plain" ocelot']);
186
211
  assert.equal(r.status, 0, r.stderr);
187
212
  assert.ok(paths(r).includes("doc/specs/q.md"));
188
213
  });
189
214
 
190
- test("8: one line per hit, eight tab-separated fields, whitespace collapsed", (t) => {
215
+ test("8: one line per hit, nine tab-separated fields, whitespace collapsed", (t) => {
191
216
  const root = repo();
192
217
  t.after(() => rmSync(root, { recursive: true, force: true }));
193
- write(root, "doc/specs/t.md", "# Tab\tin\ttitle narwhal\n\nbody narwhal\nline two\twith tab narwhal\r\nmore\n");
194
- const r = run(root, ["--query", "narwhal"]);
195
- assert.equal(r.header.length, 8);
218
+ write(root, "doc/specs/t.md", "# Tab\tin\ttitle narwhal manatee\n\n**Goal:** narwhal manatee.\n\nbody narwhal\nline two\twith tab narwhal\r\nmore\n");
219
+ const r = run(root, ["--query", "narwhal manatee"]);
220
+ assert.deepEqual(r.header, HEADER);
196
221
  assert.equal(r.rows.length, 1);
197
- assert.equal(r.rows[0].length, 8);
198
- assert.equal(r.rows[0][3], "Tab in title narwhal");
222
+ assert.equal(r.rows[0].length, 9);
223
+ assert.equal(r.rows[0][3], "Tab in title narwhal manatee");
199
224
  });
200
225
 
201
- test("9: supersession banner is plain body text", (t) => {
226
+ test("9: state is superseded only for a `- fully` banner in the title block", (t) => {
202
227
  const root = repo();
203
228
  t.after(() => rmSync(root, { recursive: true, force: true }));
204
- write(root, "doc/specs/old.md", "# Old\n\n> **Superseded by:** [doc/specs/a.md](./a.md) - fully\n\nplatypus\n");
205
- const r = run(root, ["--query", "Superseded"]);
206
- assert.deepEqual(paths(r), ["doc/specs/old.md"]);
207
- assert.deepEqual(r.rows[0].slice(4, 7), ["", "", ""]);
208
- assert.deepEqual(r.header, ["score", "path", "service", "title", "status", "shipped_at", "files", "snippet"]);
229
+ const banner = "> **Superseded by:** [doc/specs/a.md](./a.md) - fully";
230
+ write(root, "doc/specs/old-fully.md", `# Old kestrel marmot\n\n${banner}\n\n**Goal:** kestrel marmot.\n`);
231
+ write(root, "doc/specs/old-trailing.md", `# Old trailing\n\n${banner} \n`);
232
+ write(root, "doc/specs/blank-title.md", "# \n\nbody\n");
233
+ write(root, "doc/specs/old-prose.md", `# Old two\n\n> Historical note kept for context.\n${banner}\n`);
234
+ write(root, "doc/specs/old-section.md", '# Old three\n\n> **Superseded by:** [doc/specs/a.md](./a.md) - "Design" section only\n');
235
+ write(root, "doc/specs/old-semicolon.md", `# Old four\n\n${banner}; see also b\n`);
236
+ write(root, "doc/specs/old-fenced.md", `# Old five\n\n\`\`\`markdown\n${banner}\n\`\`\`\n`);
237
+ write(root, "doc/specs/old-h2.md", `# Old six\n\n## History\n\n${banner}\n`);
238
+ write(root, "doc/specs/old-h3.md", `# Old seven\n\n### Detail\n\n${banner}\n`);
239
+ const r = run(root, ["--query", "kestrel marmot"]);
240
+ assert.equal(r.status, 0, r.stderr);
241
+ assert.deepEqual(r.header, HEADER);
242
+ assert.equal(r.header[6], "state");
243
+ assert.deepEqual(paths(r), ["doc/specs/old-fully.md"]);
244
+ assert.equal(r.rows[0][6], "superseded");
245
+ const states = withDb(root, (database) => Object.fromEntries(database.prepare("SELECT path, state FROM specs WHERE path LIKE 'doc/specs/old-%'").all().map((row) => [row.path, row.state])));
246
+ assert.deepEqual(states, {
247
+ "doc/specs/old-fully.md": "superseded",
248
+ "doc/specs/old-trailing.md": "live",
249
+ "doc/specs/old-prose.md": "superseded",
250
+ "doc/specs/old-section.md": "live",
251
+ "doc/specs/old-semicolon.md": "live",
252
+ "doc/specs/old-fenced.md": "live",
253
+ "doc/specs/old-h2.md": "live",
254
+ "doc/specs/old-h3.md": "superseded",
255
+ });
256
+ assert.equal(withDb(root, (database) => database.prepare("SELECT state FROM specs WHERE path = 'doc/specs/a.md'").get().state), "live");
257
+ assert.equal(withDb(root, (database) => database.prepare("SELECT title FROM specs WHERE path = 'doc/specs/blank-title.md'").get().title), "blank-title");
209
258
  });
210
259
 
211
260
  test("10: usage errors exit 1, environment errors exit 2", (t) => {
212
261
  const root = repo();
213
262
  t.after(() => rmSync(root, { recursive: true, force: true }));
214
263
  assert.equal(run(root, []).status, 1);
215
- assert.equal(run(root, ["--query", "zephyr", "--limit", "0"]).status, 1);
216
- assert.equal(run(root, ["--query", "zephyr", "--limit", "x"]).status, 1);
217
- assert.equal(run(root, ["--query", "zephyr", "--json"]).status, 1);
264
+ assert.equal(run(root, ["--query", "zephyr widget", "--limit", "0"]).status, 1);
265
+ assert.equal(run(root, ["--query", "zephyr widget", "--limit", "x"]).status, 1);
266
+ assert.equal(run(root, ["--query", "zephyr widget", "--json"]).status, 1);
267
+ assert.equal(run(root, ["--query", "zephyr widget", "--exclude"]).status, 1);
218
268
  assert.equal(run(root, ["--query", "a"]).status, 1);
219
269
  const bare = mkdtempSync(join(tmpdir(), "gsi-bare-"));
220
270
  t.after(() => rmSync(bare, { recursive: true, force: true }));
221
- assert.equal(run(bare, ["--query", "zephyr"]).status, 2);
271
+ assert.equal(run(bare, ["--query", "zephyr widget"]).status, 2);
272
+ });
273
+
274
+ test("11: AC1 fixture - exactly the strong row; no-match and all-common queries return header only", (t) => {
275
+ const root = gitRepo();
276
+ t.after(() => rmSync(root, { recursive: true, force: true }));
277
+ const common = (n) => `# Plain spec ${n}\n\n**Goal:** plain goal ${n}.\n\nThe body mentions the common shared vocabulary.\n`;
278
+ write(root, "doc/specs/strong.md", "# Strong heron ibis\n\n**Goal:** heron ibis jackal.\n\nbody\n");
279
+ for (let n = 1; n <= 5; n++) write(root, `doc/specs/p${n}.md`, common(n));
280
+ commit(root);
281
+ const strong = run(root, ["--query", "heron ibis jackal common", "--limit", "10"]);
282
+ assert.equal(strong.status, 0, strong.stderr);
283
+ assert.deepEqual(paths(strong), ["doc/specs/strong.md"]);
284
+ for (const q of ["yeti unicorn", "common shared"]) {
285
+ const r = run(root, ["--query", q]);
286
+ assert.equal(r.status, 0, r.stderr);
287
+ assert.deepEqual(r.header, HEADER);
288
+ assert.deepEqual(r.rows, []);
289
+ }
290
+ });
291
+
292
+ test("12: evidence identity - stem variants of one word are one term", (t) => {
293
+ const root = repo();
294
+ t.after(() => rmSync(root, { recursive: true, force: true }));
295
+ write(root, "doc/specs/cfg.md", "# Configuration lemur index\n\n**Goal:** configure the lemur index.\n\nbody\n");
296
+ for (const q of ["configuration configuration", "Configuration configure", "index indexing"]) {
297
+ const r = run(root, ["--query", q]);
298
+ assert.equal(r.status, 0, r.stderr);
299
+ assert.deepEqual(r.rows, [], q);
300
+ }
301
+ assert.deepEqual(paths(run(root, ["--query", "configuration lemur"])), ["doc/specs/cfg.md"]);
302
+ });
303
+
304
+ test("13: ratio - a row far below the best live score is dropped; an equal pair is kept", (t) => {
305
+ const root = repo();
306
+ t.after(() => rmSync(root, { recursive: true, force: true }));
307
+ write(root, "doc/specs/r1.md", "# Otter pangolin vole yak emu dingo\n\n**Goal:** otter pangolin vole yak emu dingo.\n\nbody\n");
308
+ write(root, "doc/specs/r2.md", "# Otter pangolin service\n\n**Goal:** otter pangolin.\n\nbody\n");
309
+ write(root, "doc/specs/weak-super.md", "# Otter pangolin service archive\n\n> **Superseded by:** [doc/specs/r1.md](./r1.md) - fully\n\n**Goal:** otter pangolin.\n\nbody\n");
310
+ write(root, "doc/specs/e1.md", "# Quail raccoon\n\n**Goal:** quail raccoon.\n\nbody\n");
311
+ write(root, "doc/specs/e2.md", "# Quail raccoon\n\n**Goal:** quail raccoon.\n\nbody\n");
312
+ const r = run(root, ["--query", "otter pangolin vole yak emu dingo"]);
313
+ assert.deepEqual(paths(r), ["doc/specs/r1.md"]);
314
+ const pair = run(root, ["--query", "quail raccoon"]);
315
+ assert.deepEqual(paths(pair).sort(), ["doc/specs/e1.md", "doc/specs/e2.md"]);
316
+ assert.equal(pair.rows[0][0], pair.rows[1][0]);
317
+ });
318
+
319
+ test("14: superseded rows sort after live rows and never set the ratio bar", (t) => {
320
+ const root = repo();
321
+ t.after(() => rmSync(root, { recursive: true, force: true }));
322
+ write(root, "doc/specs/old.md", "# Otter pangolin vole yak emu dingo\n\n> **Superseded by:** [doc/specs/new.md](./new.md) - fully\n\n**Goal:** otter pangolin vole yak emu dingo.\n");
323
+ write(root, "doc/specs/new.md", "# Otter pangolin service\n\n**Goal:** otter pangolin.\n\nbody\n");
324
+ const r = run(root, ["--query", "otter pangolin vole yak emu dingo"]);
325
+ assert.deepEqual(paths(r), ["doc/specs/new.md", "doc/specs/old.md"]);
326
+ assert.deepEqual(r.rows.map((row) => row[6]), ["live", "superseded"]);
327
+ assert.ok(Math.abs(Number(r.rows[1][0])) > Math.abs(Number(r.rows[0][0])));
328
+ assert.deepEqual(paths(run(root, ["--query", "otter pangolin vole yak emu dingo", "--limit", "1"])), ["doc/specs/new.md"]);
329
+ assert.deepEqual(paths(run(root, ["--query", "vole yak"])), ["doc/specs/old.md"]);
330
+ });
331
+
332
+ test("15: --exclude removes a path before the filter so it never sets the bar", (t) => {
333
+ const root = repo();
334
+ t.after(() => rmSync(root, { recursive: true, force: true }));
335
+ write(root, "doc/specs/self.md", "# Sloth tapir vole yak emu dingo\n\n**Goal:** sloth tapir vole yak emu dingo.\n\nbody\n");
336
+ write(root, "doc/specs/pred.md", "# Sloth tapir service\n\n**Goal:** sloth tapir.\n\nbody\n");
337
+ const q = ["--query", "sloth tapir vole yak emu dingo"];
338
+ assert.deepEqual(paths(run(root, q)), ["doc/specs/self.md"]);
339
+ assert.deepEqual(paths(run(root, [...q, "--exclude", "doc/specs/self.md"])), ["doc/specs/pred.md"]);
340
+ assert.deepEqual(paths(run(root, [...q, "--exclude", "doc/specs/nope.md"])), ["doc/specs/self.md"]);
341
+ });
342
+
343
+ test("16: punctuation splits into distinct evidence terms and query order does not change rows", (t) => {
344
+ const root = gitRepo();
345
+ t.after(() => rmSync(root, { recursive: true, force: true }));
346
+ write(root, "doc/specs/cfg.md", "# Configuration lemur index\n\n**Goal:** configure the lemur index.\n\nbody\n");
347
+ write(root, "doc/specs/cf.md", "# Config file lynx\n\n**Goal:** config file.\n\nbody\n");
348
+ for (let n = 1; n <= 5; n++) write(root, `doc/specs/p${n}.md`, `# Plain spec ${n}\n\n**Goal:** plain goal ${n}.\n\nthe plain body\n`);
349
+ commit(root);
350
+ const common = run(root, ["--query", "the-lemur"]);
351
+ assert.equal(common.status, 0, common.stderr);
352
+ assert.deepEqual(common.header, HEADER);
353
+ assert.deepEqual(common.rows, []);
354
+ const distinct = run(root, ["--query", "configuration-lemur"]);
355
+ assert.equal(distinct.status, 0, distinct.stderr);
356
+ assert.deepEqual(paths(distinct), ["doc/specs/cfg.md"]);
357
+ assert.deepEqual(paths(run(root, ["--query", "config config-file"])), paths(run(root, ["--query", "config-file config"])));
358
+ });
359
+
360
+ test("D1: --corpus specs and omitted --corpus are byte-identical; nine columns", (t) => {
361
+ const root = repo();
362
+ t.after(() => rmSync(root, { recursive: true, force: true }));
363
+ const plain = run(root, ["--query", "zephyr widget"]);
364
+ const explicit = run(root, ["--query", "zephyr widget", "--corpus", "specs"]);
365
+ assert.equal(explicit.status, 0, explicit.stderr);
366
+ assert.equal(explicit.stdout, plain.stdout);
367
+ assert.deepEqual(explicit.header, HEADER);
368
+ });
369
+
370
+ test("D11: unknown --corpus exits 1 with usage", (t) => {
371
+ const root = repo();
372
+ t.after(() => rmSync(root, { recursive: true, force: true }));
373
+ const r = run(root, ["--query", "zephyr widget", "--corpus", "notes"]);
374
+ assert.equal(r.status, 1);
375
+ assert.match(r.stderr, /usage:/);
376
+ assert.equal(run(root, ["--query", "zephyr widget", "--corpus"]).status, 1);
377
+ });
378
+
379
+ test("D12: a schema 2 cache is rebuilt to 3 with both tables present", (t) => {
380
+ const root = repo();
381
+ t.after(() => rmSync(root, { recursive: true, force: true }));
382
+ run(root, ["--query", "zephyr widget"]);
383
+ withDb(root, (database) => database.prepare("UPDATE meta SET value = '2' WHERE key = 'schema_version'").run());
384
+ const r = run(root, ["--query", "zephyr widget"]);
385
+ assert.equal(r.status, 0, r.stderr);
386
+ assert.equal(withDb(root, (database) => database.prepare("SELECT value FROM meta WHERE key = 'schema_version'").get().value), "3");
387
+ const tables = withDb(root, (database) => database.prepare("SELECT name FROM sqlite_master WHERE name IN ('specs', 'docs') ORDER BY name").all().map((row) => row.name));
388
+ assert.deepEqual(tables, ["docs", "specs"]);
389
+ const cols = withDb(root, (database) => database.prepare("PRAGMA table_info(files)").all().map((row) => row.name));
390
+ assert.deepEqual(cols, ["corpus", "path", "mtime_ms", "size"]);
391
+ });
392
+
393
+ test("D10-specs: a `## ` line inside a code fence is not a heading", (t) => {
394
+ const root = repo();
395
+ t.after(() => rmSync(root, { recursive: true, force: true }));
396
+ write(root, "doc/specs/fenced.md", "# Fenced pelican\n\n**Goal:** pelican.\n\n```markdown\n## Not a heading pelican\n```\n\n## Real heading\n");
397
+ run(root, ["--query", "zephyr widget"]);
398
+ const h = withDb(root, (database) => database.prepare("SELECT headings FROM specs WHERE path = 'doc/specs/fenced.md'").get().headings);
399
+ assert.equal(h, "Real heading");
400
+ });
401
+
402
+ test("D2: docs include defaults and fixed exclusions", (t) => {
403
+ const root = docsRepo();
404
+ t.after(() => rmSync(root, { recursive: true, force: true }));
405
+ const r = run(root, ["--corpus", "docs", "--query", "ptarmigan gannet"]);
406
+ assert.equal(r.status, 0, r.stderr);
407
+ assert.deepEqual(r.header, DOCS_HEADER);
408
+ assert.deepEqual(paths(r), ["doc/guide.md"]);
409
+ assert.equal(r.rows[0].length, 4);
410
+ assert.deepEqual(docsPaths(root), ["README.md", ...[1, 2, 3, 4].map((n) => `doc/filler-${n}.md`), "doc/guide.md"]);
411
+ overrides(root, [".pi/**/*.md", ".worktrees/**/*.md", "doc/**/*.md"]);
412
+ assert.equal(run(root, ["--corpus", "docs", "--query", "ptarmigan gannet"]).status, 0);
413
+ assert.deepEqual(docsPaths(root), [".pi/gauntlet-overrides.md", ...[1, 2, 3, 4].map((n) => `doc/filler-${n}.md`), "doc/guide.md"]);
414
+ const two = gitRepo();
415
+ t.after(() => rmSync(two, { recursive: true, force: true }));
416
+ write(two, "doc/x.md", decoy);
417
+ write(two, "doc/y.md", decoy);
418
+ commit(two);
419
+ const tiny = run(two, ["--corpus", "docs", "--query", "ptarmigan gannet"]);
420
+ assert.equal(tiny.status, 0, tiny.stderr);
421
+ assert.deepEqual(tiny.header, DOCS_HEADER);
422
+ assert.deepEqual(tiny.rows, []);
423
+ });
424
+
425
+ test("D3: shared confidence rejects common body-only hits", (t) => {
426
+ const root = gitRepo();
427
+ t.after(() => rmSync(root, { recursive: true, force: true }));
428
+ write(root, "doc/strong.md", "# Strong heron ibis\n\n## jackal\n\nbody\n");
429
+ for (let n = 1; n <= 5; n++) write(root, `doc/p${n}.md`, `# Plain doc ${n}\n\nThe body mentions the common shared vocabulary.\n`);
430
+ commit(root);
431
+ const r = run(root, ["--corpus", "docs", "--query", "heron ibis jackal common", "--limit", "10"]);
432
+ assert.equal(r.status, 0, r.stderr);
433
+ assert.deepEqual(paths(r), ["doc/strong.md"]);
434
+ });
435
+
436
+ test("D4: no cross-corpus leakage", (t) => {
437
+ const root = repo();
438
+ t.after(() => rmSync(root, { recursive: true, force: true }));
439
+ write(root, "doc/guide.md", "# zephyr widget guide\n\n## zephyr widget\n");
440
+ for (let n = 1; n <= 4; n++) write(root, `doc/filler-${n}.md`, docFiller(n));
441
+ assert.deepEqual(paths(run(root, ["--query", "zephyr widget"])), ["doc/specs/a.md", "svc-a/doc/specs/b.md"]);
442
+ const docs = run(root, ["--corpus", "docs", "--query", "zephyr widget"]);
443
+ assert.deepEqual(paths(docs), ["doc/guide.md"]);
444
+ assert.deepEqual(docs.header, DOCS_HEADER);
445
+ });
446
+
447
+ test("D5: independent cache refresh in both directions", (t) => {
448
+ const root = repo();
449
+ t.after(() => rmSync(root, { recursive: true, force: true }));
450
+ write(root, "doc/guide.md", "# Guide\n\n## zephyr widget\n");
451
+ run(root, ["--query", "zephyr widget"]);
452
+ run(root, ["--corpus", "docs", "--query", "zephyr widget"]);
453
+ const files = () => withDb(root, (db) => Object.fromEntries(db.prepare("SELECT corpus, path, mtime_ms FROM files").all().map((r) => [`${r.corpus}:${r.path}`, r.mtime_ms])));
454
+ const before = files();
455
+ utimesSync(join(root, "doc/specs/a.md"), new Date(), new Date(Date.now() + 5000));
456
+ run(root, ["--corpus", "docs", "--query", "zephyr widget"]);
457
+ assert.deepEqual(files(), before);
458
+ run(root, ["--query", "zephyr widget"]);
459
+ assert.notEqual(files()["specs:doc/specs/a.md"], before["specs:doc/specs/a.md"]);
460
+ assert.equal(files()["docs:doc/guide.md"], before["docs:doc/guide.md"]);
461
+ assert.deepEqual(docsPaths(root), ["doc/guide.md"]);
462
+ });
463
+
464
+ test("D6: overrides replace defaults but cannot bypass exclusions", (t) => {
465
+ const root = docsRepo();
466
+ t.after(() => rmSync(root, { recursive: true, force: true }));
467
+ write(root, "notes/a.md", decoy);
468
+ write(root, "CONTRIBUTING.md", decoy);
469
+ write(root, "svc/CONTRIBUTING.md", decoy);
470
+ write(root, "svc/doc/inner.md", decoy);
471
+ commit(root);
472
+ const search = () => run(root, ["--corpus", "docs", "--query", "ptarmigan gannet"]);
473
+ overrides(root, ["notes/**/*.md"]);
474
+ assert.equal(search().status, 0);
475
+ assert.deepEqual(docsPaths(root), ["notes/a.md"]);
476
+ overrides(root, ["**/*.md"]);
477
+ assert.equal(search().status, 0);
478
+ const all = docsPaths(root);
479
+ for (const p of ["doc/specs/a.md", "docs/specs/b.md", "doc/plans/p.md", "docs/plans/q.md", "node_modules/x/doc/a.md", "build/doc/b.md", "doc/draft.md", ".pi/gauntlet/n.md", ".worktrees/w/doc/g.md"]) assert.ok(!all.includes(p), p);
480
+ assert.ok(all.includes("doc/guide.md"));
481
+ overrides(root, ["CONTRIBUTING.md"]);
482
+ assert.equal(search().status, 0);
483
+ assert.deepEqual(docsPaths(root), ["CONTRIBUTING.md", "svc/CONTRIBUTING.md"]);
484
+ write(root, ".pi/gauntlet-overrides.md", "## Spec index\n\nno bullets here\n");
485
+ assert.equal(search().status, 0);
486
+ assert.ok(docsPaths(root).includes("doc/guide.md"));
487
+ assert.ok(docsPaths(root).includes("svc/doc/inner.md"));
488
+ });
489
+
490
+ test("D6b: overrides file resolution order", (t) => {
491
+ const root = docsRepo();
492
+ t.after(() => rmSync(root, { recursive: true, force: true }));
493
+ write(root, "doc/gauntlet-overrides.md", "## Spec index\n- docs: `notes/**/*.md`\n");
494
+ write(root, "notes/a.md", decoy);
495
+ const search = () => run(root, ["--corpus", "docs", "--query", "ptarmigan gannet"]);
496
+ assert.equal(search().status, 0);
497
+ assert.deepEqual(docsPaths(root), ["notes/a.md"]);
498
+ write(root, "gauntlet-overrides.md", "## Spec index\n- docs: `other/**/*.md`\n");
499
+ write(root, "other/b.md", decoy);
500
+ assert.equal(search().status, 0);
501
+ assert.deepEqual(docsPaths(root), ["other/b.md"]);
502
+ overrides(root, ["doc/**/*.md"]);
503
+ assert.equal(search().status, 0);
504
+ assert.deepEqual(docsPaths(root), ["doc/filler-1.md", "doc/filler-2.md", "doc/filler-3.md", "doc/filler-4.md", "doc/gauntlet-overrides.md", "doc/guide.md"]);
505
+ });
506
+
507
+ test("D7: heading hits are strong, body-only hits are not", (t) => {
508
+ const root = gitRepo();
509
+ t.after(() => rmSync(root, { recursive: true, force: true }));
510
+ write(root, "doc/h.md", "# H\n\n## Setup\n\n## Feeding the kudu wombat\n\nprose\n");
511
+ write(root, "doc/b.md", "# B\n\n## Setup\n\nkudu wombat in the body only\n");
512
+ for (let n = 1; n <= 4; n++) write(root, `doc/filler-${n}.md`, docFiller(n));
513
+ commit(root);
514
+ const r = run(root, ["--corpus", "docs", "--query", "kudu wombat"]);
515
+ assert.equal(r.status, 0, r.stderr);
516
+ assert.deepEqual(paths(r), ["doc/h.md"]);
517
+ });
518
+
519
+ test("D8: snippet selects heading line or body fallback", (t) => {
520
+ const root = gitRepo();
521
+ t.after(() => rmSync(root, { recursive: true, force: true }));
522
+ const longHeading = "Feeding the wandering herds across twelve distant valleys before finally finding the kudu wombat";
523
+ write(root, "doc/h.md", `# H\n\n## Setup\n\n## ${longHeading}\n\nprose\n`);
524
+ write(root, "doc/t.md", "# lynx ocelot title\n\n## Unrelated\n\nbody says lynx ocelot here\n");
525
+ for (let n = 1; n <= 4; n++) write(root, `doc/filler-${n}.md`, docFiller(n));
526
+ commit(root);
527
+ const heading = run(root, ["--corpus", "docs", "--query", "kudu wombat"]);
528
+ assert.deepEqual(paths(heading), ["doc/h.md"]);
529
+ assert.equal(heading.rows[0][3], longHeading);
530
+ const body = run(root, ["--corpus", "docs", "--query", "lynx ocelot"]);
531
+ assert.deepEqual(paths(body), ["doc/t.md"]);
532
+ assert.match(body.rows[0][3], /body says lynx ocelot/);
533
+ assert.doesNotMatch(body.rows[0][3], /[\u0001\u0002]/);
534
+ });
535
+
536
+ test("D9: deleted files and symlinks are skipped, stale rows removed", (t) => {
537
+ const root = docsRepo();
538
+ t.after(() => rmSync(root, { recursive: true, force: true }));
539
+ write(root, "doc/gone.md", "# Gone\n\n## ptarmigan gannet\n");
540
+ commit(root);
541
+ run(root, ["--corpus", "docs", "--query", "ptarmigan gannet"]);
542
+ assert.ok(docsPaths(root).includes("doc/gone.md"));
543
+ rmSync(join(root, "doc/gone.md"));
544
+ symlinkSync(join(root, "doc/specs/a.md"), join(root, "doc/link.md"));
545
+ symlinkSync(join(root, "doc/nowhere.md"), join(root, "doc/dangling.md"));
546
+ const r = run(root, ["--corpus", "docs", "--query", "ptarmigan gannet"]);
547
+ assert.equal(r.status, 0, r.stderr);
548
+ assert.deepEqual(paths(r), ["doc/guide.md"]);
549
+ for (const p of ["doc/gone.md", "doc/link.md", "doc/dangling.md"]) assert.ok(!docsPaths(root).includes(p), p);
550
+ assert.equal(withDb(root, (db) => db.prepare("SELECT count(*) AS c FROM files WHERE corpus = 'docs' AND path = 'doc/gone.md'").get().c), 0);
551
+ });
552
+
553
+ test("D10: fenced headings and shorter nested fences are ignored", (t) => {
554
+ const root = docsRepo();
555
+ t.after(() => rmSync(root, { recursive: true, force: true }));
556
+ write(root, "doc/f.md", "# F\n\n```\n## kudu wombat\n```\n\nbody\n");
557
+ write(root, "doc/nested.md", "# N\n\n````markdown\n```js\n## inner heading\n```\n## still fenced\n````\n");
558
+ run(root, ["--corpus", "docs", "--query", "ptarmigan gannet"]);
559
+ for (const p of ["doc/f.md", "doc/nested.md"]) assert.equal(withDb(root, (db) => db.prepare("SELECT headings FROM docs WHERE path = ?").get(p).headings), "");
222
560
  });
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-gauntlet",
3
- "version": "5.21.0",
3
+ "version": "5.23.0",
4
4
  "description": "Opinionated, gated workflow skills, subagent personas, and runtime extensions for the pi coding agent.",
5
5
  "author": "Jacek Juraszek",
6
6
  "type": "module",
@@ -44,7 +44,7 @@ Work through the items below **in order**. This is your own checklist to follow,
44
44
  4. **Understand the idea against the draft** - see [Understand the idea](#3-understand-the-idea).
45
45
  5. **Propose 2-3 approaches** - see [Explore approaches](#4-explore-approaches).
46
46
  6. **Present the design** - see [Present the design in two rounds](#6-present-the-design-in-two-rounds).
47
- 7. **Write the spec** - follow [Spec Self-Review](#spec-self-review-before-user-review-gate) steps 1-4.
47
+ 7. **Write the spec** - follow [Spec Self-Review](#spec-self-review-before-user-review-gate) steps 1-5.
48
48
  8. **Spec self-review (lint)** - run the inline checks in [Spec Self-Review](#spec-self-review-before-user-review-gate).
49
49
  9. **Critique pass (auto-dispatched)** - use [Spec Council](#spec-council-optional).
50
50
  10. **Re-run placeholder scan** - follow [Spec Self-Review](#spec-self-review-before-user-review-gate).
@@ -104,7 +104,7 @@ Prefer clear testable boundaries, YAGNI, existing conventions, the owning schema
104
104
 
105
105
  ### 6. Present the design in two rounds
106
106
 
107
- Use two rounds, targeting 300-500 words each, with one approval each; revisions remain within that approval point. Ask once per round; round-1 approval without correction confirms the predecessor. Round 1 covers architecture, responsibilities, data flow, and `supersedes <path>, <scope>` when applicable. Round 2 covers errors, edges, tests, and `## Documentation impact`.
107
+ Use two rounds, targeting 300-500 words each, with one approval each; revisions remain within that approval point. Ask once per round; round-1 approval without correction confirms the predecessor. Round 1 covers architecture, responsibilities, data flow, and `supersedes <path>, <scope>` when applicable. Round 2 covers errors, edges, tests, and `## Documentation impact`. In round 2, cite the draft's `Docs touched:` entries as candidates for "Materially amended existing docs"; admit or drop each by the materiality bar in `reference/documentation-impact.md`, never list them automatically.
108
108
 
109
109
  `## Documentation impact` is required. Cite `reference/documentation-impact.md` by relative path, do not restate its categories, and reproduce this template verbatim:
110
110
 
@@ -163,7 +163,21 @@ Spec-writing replaces the context draft, in this exact order:
163
163
  `# CONTEXT DRAFT - NOT A SPEC - fully replaced at spec-writing` — before
164
164
  dispatching lint, critique, council, or summarizer. The phase-tracker commit
165
165
  guard is a backstop, not the primary check.
166
- 4. **After the line-1 check and before the inline lint**, `edit` any known
166
+ 4. **After the line-1 check**, run the second predecessor pass. Compose a query from
167
+ the spec's H1 terms, then the terms of every H2 that is not a template heading
168
+ (Problem, Acceptance criteria, Design, Errors and edge cases, Tests, Documentation
169
+ impact, Out of scope, Open questions), then the `**Goal:**` line's terms - same
170
+ token rules as the scout, deduplicated, cut at 15. Run
171
+ `(cd <abs worktree path> && node <SPEC_INDEX> --query '<terms>' --limit 5 --exclude <spec path relative to the worktree>)`
172
+ with `<SPEC_INDEX>` resolved as in `gatherer.md`. Take the `live` rows and drop every
173
+ path the draft's `Predecessor:` line(s) named (retained from the step-1 read). Carry
174
+ the remainder to the gate as one adjacent line:
175
+ `New predecessor candidates at spec-writing: <path> (<title>), ... - index rows are a hint; code is the source and an absent row proves nothing.`
176
+ or `New predecessor candidates at spec-writing: none.` A non-zero exit degrades to
177
+ `Second predecessor pass unavailable: <first stderr line>`; it never blocks the gate
178
+ or the commit. "yes, <path> is a predecessor" at the gate is a change request that
179
+ adds the `supersedes <path>, <scope>` clause and the banner, then re-presents the gate.
180
+ 5. **After the second predecessor pass and before the inline lint**, `edit` any known
167
181
  predecessor spec to insert its supersession banner (see
168
182
  [Marking superseded specs](reference/superseding.md)). This position is fixed:
169
183
  the banner is written after any slug rename, so it always cites the final path.
@@ -244,6 +258,7 @@ Rejected: [<severity>] <cluster> — raised-by: [<slugs>] -> <one-line reason>
244
258
  (one line per item, exactly as returned by roasting-the-spec - `Applied: none` / `Deferred: none` / `Rejected: none` when a list is empty; omit the audit lines when the worker path ran, not the council)
245
259
 
246
260
  <unresolved ambiguities; every gap-footer entry from the summary>
261
+ New predecessor candidates at spec-writing: <path> (<title>), ... - index rows are a hint; code is the source and an absent row proves nothing.
247
262
 
248
263
  Please review. Approve to proceed, tell me what to change in the spec, or say "revert applied council edit <X>" to undo a specific applied edit. Reply "auto-apply amends" - every later amend-class change in this flow then applies without review, scope changes included; redraws and the spec gate still stop. "approve, auto-apply amends" does both.
249
264
  ```
@@ -288,6 +303,7 @@ One question at a time, YAGNI, 2-3 approaches, two design rounds, clarify freely
288
303
  - Proposed-change execution before approval ([owner](#hard-constraint)).
289
304
  - Plan before approval; brainstorming invocation for an amend ([owner](#user-review-gate)).
290
305
  - Missing predecessor banner; invalid multi-spec split ([owner](#spec-self-review-before-user-review-gate); [owner](#2-scope-check)).
306
+ - Gate reached without the second predecessor pass ([owner](#spec-self-review-before-user-review-gate)).
291
307
  - Approaches while a contradicted premise remains unresolved ([owner](#3-understand-the-idea)).
292
308
  - Amend without `reference/amendment-surface.md`; waiting after an amend grant; auto-applying a redraw ([owner](#amending-an-approved-spec)).
293
309
 
@@ -53,11 +53,13 @@ Scout (always dispatched):
53
53
  > instead; cite the old spec only for its unsuperseded sections (banner contract:
54
54
  > `reference/superseding.md`). Predecessor check: compose a 5-15 term keyword query
55
55
  > from the request (topic nouns, component names, file names - not stop words; if
56
- > the request is only a ticket reference, take the terms from the ticket title via
56
+ > the request is only a ticket reference, take the terms from the ticket title and body via
57
57
  > the tracker CLI when one is available, otherwise use the fallback below). Run
58
58
  > `node <SPEC_INDEX> --query '<keywords>' --limit 10` from the worktree root,
59
59
  > keeping the keywords inside single quotes, and treat its rows as the candidate
60
- > list. If the command fails, fall back to listing the project's spec directory
60
+ > list; zero rows means the index found no evidence, not that no predecessor exists -
61
+ > judge `Predecessor: none` from the code recon. If the command fails, fall back to
62
+ > listing the project's spec directory
61
63
  > and reading titles and `**Goal:**` lines, and write
62
64
  > `Spec index unavailable - predecessor check used directory listing.` in your
63
65
  > handoff. Either way open at
@@ -67,7 +69,8 @@ Scout (always dispatched):
67
69
  > Judge by topic; shared file paths never decide.
68
70
  > The `files` column of each candidate row is the `;`-separated list of repo-relative
69
71
  > paths that predecessor's ship modified and that still exist, or the literal `missing`,
70
- > or blank; do not recompute it from git or telemetry. After the `Predecessor:` line(s),
72
+ > or blank; do not recompute it from git or telemetry. The `state` column is `live` or
73
+ > `superseded`; cite a `superseded` row only through its successor. After the `Predecessor:` line(s),
71
74
  > and only when at least one predecessor is named, render a `Predecessor anchors`
72
75
  > section: list every attributed path exactly once, attributed to the first named
73
76
  > predecessor in index output order whose cell lists it, with no per-path commentary
@@ -79,7 +82,22 @@ Scout (always dispatched):
79
82
  > from the directory-listing fallback) contributes no path and no `missing` line. When
80
83
  > no path and no `missing` line results - including whenever the index was unavailable -
81
84
  > omit the section entirely; `Predecessor: none` produces no anchors section. The anchors
82
- > are a recon hint, never a selection input. End with an
85
+ > are a recon hint, never a selection input.
86
+ > Docs check: run `node <SPEC_INDEX> --corpus docs --query '<the same keywords>' --limit 10`
87
+ > from the worktree root. Its rows are project documentation that already speaks about the
88
+ > request's topic; open at most five whose topic matches and let them inform the "already
89
+ > solves this?" finding and the current-contract recon. If the command fails, write
90
+ > `Docs index unavailable - docs check used recon only.` in your handoff and continue. After
91
+ > the `Predecessor:` line(s) and any `Predecessor anchors` section, render at most
92
+ > five `Docs touched: <path> - <section heading>` lines, one per document that
93
+ > genuinely covers the request's topic - an opened index row or a document your own
94
+ > recon found - where
95
+ > `<section heading>` is the `##`/`###` heading of the covering section as read in the
96
+ > document (the row's `snippet` column is a hint to it), or the document title when no single
97
+ > section applies; or `Docs touched: none`. Judge by topic; a lexical hit alone is not
98
+ > coverage. Zero rows means the index found no evidence, not that no document covers the
99
+ > topic. This query supplements the code and documentation recon you already perform; it
100
+ > never replaces it - keep reading the files the request touches. End with an
83
101
  > "Open questions that matter for the spec"
84
102
  > section. Compact handoff, not a dump.
85
103