pi-gauntlet 5.22.0 → 5.23.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,21 @@
1
1
  # Changelog
2
2
 
3
+ ## v5.23.1 - 2026-09-27
4
+
5
+ ### Changed
6
+
7
+ - `gauntlet-spec-index` prints a one-line stderr note (`<corpus> corpus has <n> documents - too small for the confidence rule; no rows returned`) when a query returns no rows on a corpus under 10 documents; stdout and exit code are unchanged. The brainstorming scout carries the reason as `Predecessor: none (specs corpus too small: <n> documents)` / `Docs touched: none (docs corpus too small: <n> documents)`.
8
+
9
+ ## v5.23.0 - 2026-09-27
10
+
11
+ ### Added
12
+
13
+ - `gauntlet-spec-index` gains `--corpus specs|docs` (default `specs`, output unchanged). `--corpus docs` indexes project documentation - tracked and untracked non-gitignored `*.md` files matching `**/doc/**/*.md`, `**/docs/**/*.md`, `README.md`, `AGENTS.md`, replaceable through `- docs: <glob>` bullets under `## Spec index` in the gauntlet overrides file - into a second FTS5 table in the same cache, minus `doc(s)/specs`, `doc(s)/plans`, `node_modules`, `.pi/gauntlet`, `.worktrees`, symlinks and context drafts, and prints `score`, `path`, `title`, `snippet`. Both corpora share one confidence rule driven by a corpus descriptor (strong fields `title`/`goal` for specs, `title`/`headings` for docs). The brainstorming scout runs the docs query and renders `Docs touched: <path> - <section heading>` lines that round 2 weighs as documentation-impact candidates.
14
+
15
+ ### Changed
16
+
17
+ - Spec-index cache schema is 3 (the `files` table is keyed by `(corpus, path)`); existing caches rebuild on first query. `#`/`##`/`###` lines inside fenced code blocks no longer count as the title or headings in either corpus.
18
+
3
19
  ## v5.22.0 - 2026-09-26
4
20
 
5
21
  ### Added
package/README.md CHANGED
@@ -137,7 +137,7 @@ Pin an exact release with `npm:pi-gauntlet@X.Y.Z`. See [doc/install-internals.md
137
137
 
138
138
  ## Spec search index
139
139
 
140
- `gauntlet-spec-index` provides lexical search across `doc/specs/*.md` at the repository root and one service level down. From a repository worktree, run `node <pi-gauntlet-package>/bin/gauntlet-spec-index.mjs --query "<text>" [--limit N] [--exclude <repo-relative path>]...`; it requires Node >=24.15.0, refreshes its FTS5 index on every query, and prints tab-separated `score`, `path`, `service`, `title`, `status`, `shipped_at`, `state`, `files`, and `snippet` columns. Rows pass a confidence rule before `--limit` applies: a query term is evidence when it occurs in fewer than half the indexed specs (stem variants of one word count once); query tokens split on word boundaries, so `gauntlet-spec-index` searches as three words; a row is returned only when it matches at least two distinct evidence terms in its `title` or `goal`, and rows scoring below half the best live row's score are dropped. A query with no real match therefore prints the header only and exits 0, and a corpus of two or fewer specs never returns rows. `--exclude` (repeatable) removes a path before the rule runs, so the spec under review never sets the bar. `state` is `superseded` when a `> **Superseded by:** ... - fully` banner sits in the block directly under the title and `live` otherwise; named-section banners and consumer-defined banner syntaxes both read as `live`. Live rows sort before superseded rows, then by score. The `files` column is a `;`-separated list of repo-relative paths the spec's shipped change modified and that still exist in the repository, the literal `missing` when the spec's telemetry record has no `derived.modified_files` list, or blank when there is no readable record or no recorded path remains. The per-worktree cache lives at `.pi/gauntlet/index.sqlite`, and its first creation adds `/.pi/gauntlet/index.sqlite*` to Git's `info/exclude` so the database and SQLite sidecars stay out of `git status`. `/skill:brainstorming` queries the index twice: the scout at gather time from the request, and the main loop at spec-writing from the finished spec's title, goal, and headings, surfacing new `live` candidates at the review gate.
140
+ `gauntlet-spec-index` provides lexical search over two corpora in one per-worktree cache. From a repository worktree, run `node <pi-gauntlet-package>/bin/gauntlet-spec-index.mjs --query "<text>" [--corpus specs|docs] [--limit N] [--exclude <repo-relative path>]...`; it requires Node >=24.15.0 and refreshes its FTS5 index on every query. `--corpus specs` (the default) searches `doc/specs/*.md` at the repository root and one service level down and prints tab-separated `score`, `path`, `service`, `title`, `status`, `shipped_at`, `state`, `files`, and `snippet` columns. `--corpus docs` searches project documentation and prints `score`, `path`, `title`, and `snippet`; the snippet is the matched `##`/`###` heading line when one holds a query term, else a body fragment. Docs candidates are the tracked and untracked, non-gitignored `*.md` files (`git ls-files -co --exclude-standard`, so nested repositories and submodules are not indexed) matching an include list - default `**/doc/**/*.md`, `**/docs/**/*.md`, `README.md`, `AGENTS.md`, where a slash-free entry also matches one service level down - minus the fixed exclusions `**/doc/specs/**`, `**/docs/specs/**`, `**/doc/plans/**`, `**/docs/plans/**`, `**/node_modules/**`, `.pi/gauntlet/**`, `.worktrees/**`, symlinks, and files whose first line is the context-draft marker; `**` never matches a dot directory, so `.pi/*.md` or `.github/*.md` is indexed only through an explicit include. A project replaces the include list with `- docs: <glob>` bullets under a `## Spec index` heading in its gauntlet overrides file (see [Project-specific overrides](#project-specific-overrides)); the exclusions are not configurable. Both corpora pass the same confidence rule before `--limit` applies: a query term is evidence when it occurs in fewer than half the rows of the queried table (stem variants of one word count once); query tokens split on word boundaries, so `gauntlet-spec-index` searches as three words; a row is returned only when it matches at least two distinct evidence terms in the corpus's strong fields - `title` or `goal` for specs, `title` or `headings` for docs - and rows scoring below half the best live row's score are dropped (every docs row is live). A query with no real match therefore prints the header only and exits 0 - and when the queried corpus has fewer than 10 documents, also writes one stderr line naming the corpus and its size (`gauntlet-spec-index: docs corpus has 5 documents - too small for the confidence rule; no rows returned`), so a small corpus and a genuine miss stay distinguishable; a corpus of two or fewer files never returns rows. `--exclude` (repeatable) removes a path before the rule runs, so the spec under review never sets the bar. `state` is `superseded` when a `> **Superseded by:** ... - fully` banner sits in the block directly under the title and `live` otherwise; named-section banners and consumer-defined banner syntaxes both read as `live`. Live rows sort before superseded rows, then by score. The `files` column is a `;`-separated list of repo-relative paths the spec's shipped change modified and that still exist in the repository, the literal `missing` when the spec's telemetry record has no `derived.modified_files` list, or blank when there is no readable record or no recorded path remains. The cache lives at `.pi/gauntlet/index.sqlite`, and its first creation adds `/.pi/gauntlet/index.sqlite*` to Git's `info/exclude` so the database and SQLite sidecars stay out of `git status`. `/skill:brainstorming` queries the specs corpus twice - the scout at gather time from the request, and the main loop at spec-writing from the finished spec's title, goal, and headings, surfacing new `live` candidates at the review gate - and the scout queries the docs corpus once to render `Docs touched:` lines that round 2 weighs as documentation-impact candidates.
141
141
 
142
142
  ## Performance digest
143
143
 
@@ -279,10 +279,18 @@ Use the project's wrapper: `script/worktree create <name>`. It provisions an iso
279
279
  database and copies `.env.local`. Never call `git worktree add` directly.
280
280
  ```
281
281
 
282
- Beyond `## conventions` and skill-named sections, headings a skill reads by name are documented in their owning skills; other headings retain topic/workflow-convention matching. The override file is read by skill instructions, not by the Pi runtime itself. Missing or empty `## conventions` adds no rules; a lower-priority file cannot supplement the selected file.
282
+ Beyond `## conventions` and skill-named sections, headings a skill reads by name are documented in their owning skills; other headings retain topic/workflow-convention matching. The override file is read by skill instructions and, for its `## Spec index` section only, by `gauntlet-spec-index`; the Pi runtime itself never reads it. Missing or empty `## conventions` adds no rules; a lower-priority file cannot supplement the selected file.
283
283
 
284
284
  **Discovery ladder:** skills check three locations, in order, and use the first one found - never merged: `.pi/gauntlet-overrides.md`, then `<repo root>/gauntlet-overrides.md`, then `<repo root>/doc/gauntlet-overrides.md` (`<repo root>` = `git rev-parse --show-toplevel`, or the current directory outside a repo). Pick one location per repo.
285
285
 
286
+ **`## Spec index` section:** `gauntlet-spec-index --corpus docs` reads `- docs: <glob>` bullets under this heading as the docs include list, replacing the default (`**/doc/**/*.md`, `**/docs/**/*.md`, `README.md`, `AGENTS.md`). Backticks or quotes around the glob are stripped; no other key under the heading is read, and the fixed exclusions always apply. For example:
287
+
288
+ ```markdown
289
+ ## Spec index
290
+ - docs: `**/doc/**/*.md`
291
+ - docs: `handbook/**/*.md`
292
+ ```
293
+
286
294
  **`## Issue tracker` section:** `shape-ticket` resolves tracker access through a capability ladder, and this is its first rung - it overrides the zero-config `gh` (GitHub) / `linearis` (Linear) defaults for any other tracker. Name the CLI's read, search, create, update, and post comment commands explicitly. For a Jira CLI, for example:
287
295
 
288
296
  ```markdown
@@ -1,16 +1,17 @@
1
1
  #!/usr/bin/env node
2
- // Lexical search over the spec corpus. Build-on-query: refresh a per-worktree
3
- // FTS5 cache by mtime+size, then rank with bm25 and join telemetry at output.
4
- import { readFileSync, appendFileSync, existsSync, statSync, readdirSync, mkdirSync, rmSync, realpathSync } from "node:fs";
5
- import { join, dirname, basename, isAbsolute } from "node:path";
2
+ // Lexical search over spec and docs corpora (FTS5, bm25, shared confidence rule); telemetry joined for specs only.
3
+ import { readFileSync, appendFileSync, existsSync, statSync, lstatSync, readdirSync, mkdirSync, rmSync, realpathSync, openSync, readSync, closeSync } from "node:fs";
4
+ import { join, dirname, basename, isAbsolute, matchesGlob } from "node:path";
6
5
  import { execFileSync } from "node:child_process";
7
6
  import process from "node:process";
8
7
  import { parse as parseYaml } from "yaml";
9
8
 
10
- const SCHEMA_VERSION = 2;
9
+ const SCHEMA_VERSION = 3;
11
10
  const EVIDENCE_DF_FRACTION = 0.5;
12
11
  const MIN_EVIDENCE_TOKENS = 2;
13
12
  const SCORE_RATIO = 0.5;
13
+ // Policy floor, not derived from the rule: below it an empty result says little about coverage.
14
+ const MIN_CORPUS = 10;
14
15
  // Default banner grammar of skills/brainstorming/reference/superseding.md; a named-section
15
16
  // scope or anything after `- fully` keeps the spec live.
16
17
  const FULLY_BANNER = /^> \*\*Superseded by:\*\* \[.*\]\(.*\) - fully$/;
@@ -18,17 +19,23 @@ const MIN_NODE = [24, 15, 0];
18
19
  const DRAFT_MARKER = "# CONTEXT DRAFT - NOT A SPEC - fully replaced at spec-writing";
19
20
  const EXCLUDE_LINE = "/.pi/gauntlet/index.sqlite*";
20
21
  const SKIP_DIRS = new Set([".worktrees", "node_modules", "build"]);
21
- const HEADER = ["score", "path", "service", "title", "status", "shipped_at", "state", "files", "snippet"];
22
+ const FENCE = /^\s*(`{3,}|~{3,})/;
23
+ const DOC_INCLUDES = ["**/doc/**/*.md", "**/docs/**/*.md", "README.md", "AGENTS.md"];
24
+ const DOC_EXCLUDES = ["**/doc/specs/**", "**/docs/specs/**", "**/doc/plans/**", "**/docs/plans/**", "**/node_modules/**", ".pi/gauntlet/**", ".worktrees/**"];
25
+ const OVERRIDES_FILES = [".pi/gauntlet-overrides.md", "gauntlet-overrides.md", "doc/gauntlet-overrides.md"];
22
26
  const SCHEMA = `
23
27
  CREATE TABLE IF NOT EXISTS meta (key TEXT PRIMARY KEY, value TEXT);
24
- CREATE TABLE IF NOT EXISTS files (path TEXT PRIMARY KEY, mtime_ms INTEGER, size INTEGER);
28
+ CREATE TABLE IF NOT EXISTS files (corpus TEXT NOT NULL, path TEXT NOT NULL, mtime_ms INTEGER, size INTEGER, PRIMARY KEY (corpus, path));
25
29
  CREATE VIRTUAL TABLE IF NOT EXISTS specs USING fts5(
26
30
  path UNINDEXED, service UNINDEXED, title, goal, headings, body, state UNINDEXED,
27
31
  tokenize = 'porter unicode61');
32
+ CREATE VIRTUAL TABLE IF NOT EXISTS docs USING fts5(
33
+ path UNINDEXED, title, headings, body,
34
+ tokenize = 'porter unicode61');
28
35
  `;
29
36
 
30
37
  const usage = () => {
31
- process.stderr.write('usage: gauntlet-spec-index --query "<text>" [--limit N] [--exclude <repo-relative path>]...\n');
38
+ process.stderr.write('usage: gauntlet-spec-index --query "<text>" [--corpus specs|docs] [--limit N] [--exclude <repo-relative path>]...\n');
32
39
  process.exit(1);
33
40
  };
34
41
  const die = (msg) => {
@@ -38,16 +45,18 @@ const die = (msg) => {
38
45
 
39
46
  function parseArgs(argv) {
40
47
  let query;
48
+ let corpus = "specs";
41
49
  let limit = 10;
42
50
  const exclude = [];
43
51
  for (let i = 0; i < argv.length; i++) {
44
52
  if (argv[i] === "--query" && argv[i + 1] !== undefined) query = argv[++i];
53
+ else if (argv[i] === "--corpus" && Object.hasOwn(CORPORA, argv[i + 1] ?? "")) corpus = argv[++i];
45
54
  else if (argv[i] === "--limit" && /^[1-9]\d*$/.test(argv[i + 1] ?? "")) limit = Number(argv[++i]);
46
55
  else if (argv[i] === "--exclude" && argv[i + 1] !== undefined) exclude.push(argv[++i]);
47
56
  else usage();
48
57
  }
49
58
  if (query === undefined) usage();
50
- return { query, limit, exclude };
59
+ return { query, corpus, limit, exclude };
51
60
  }
52
61
 
53
62
  function nodeOk() {
@@ -88,11 +97,64 @@ function discover(root) {
88
97
  }
89
98
  };
90
99
  collect("doc/specs", "root");
91
- const entries = readdirSync(root, { withFileTypes: true }).sort((a, b) => a.name.localeCompare(b.name));
92
- for (const e of entries) {
93
- if (e.name.startsWith(".") || SKIP_DIRS.has(e.name)) continue;
94
- if (!e.isDirectory() && !e.isSymbolicLink()) continue;
95
- collect(`${e.name}/doc/specs`, e.name);
100
+ for (const s of services(root).sort((a, b) => a.localeCompare(b))) collect(`${s}/doc/specs`, s);
101
+ return out;
102
+ }
103
+
104
+ // Top-level directories discover() would scan; a slash-free include also matches one level under them.
105
+ function services(root) {
106
+ return readdirSync(root, { withFileTypes: true })
107
+ .filter((e) => (e.isDirectory() || e.isSymbolicLink()) && !e.name.startsWith(".") && !SKIP_DIRS.has(e.name))
108
+ .map((e) => e.name);
109
+ }
110
+
111
+ function includes(root) {
112
+ const file = OVERRIDES_FILES.map((rel) => join(root, rel)).find((abs) => existsSync(abs));
113
+ if (!file) return DOC_INCLUDES;
114
+ const globs = [];
115
+ let inSection = false;
116
+ let text;
117
+ try {
118
+ text = readFileSync(file, "utf8");
119
+ } catch (e) {
120
+ die(`gauntlet-spec-index: unreadable overrides file ${file}: ${e.message}`);
121
+ }
122
+ for (const line of text.split(/\r?\n/)) {
123
+ if (/^## /.test(line)) { inSection = line.trim() === "## Spec index"; continue; }
124
+ if (!inSection) continue;
125
+ const m = /^-\s+docs:\s*(.+?)\s*$/.exec(line);
126
+ if (m) globs.push(m[1].replace(/^[`"']+|[`"']+$/g, ""));
127
+ }
128
+ return globs.length ? globs : DOC_INCLUDES;
129
+ }
130
+
131
+ function firstLine(abs) {
132
+ const fd = openSync(abs, "r");
133
+ try {
134
+ const buf = Buffer.alloc(128);
135
+ const n = readSync(fd, buf, 0, 128, 0);
136
+ return buf.toString("utf8", 0, n).split(/\r?\n/, 1)[0];
137
+ } finally {
138
+ closeSync(fd);
139
+ }
140
+ }
141
+
142
+ function discoverDocs(root, globs) {
143
+ const svc = services(root);
144
+ const patterns = globs.flatMap((g) => (g.includes("/") ? [g] : [g, ...svc.map((s) => `${s}/${g}`)]));
145
+ const listed = execFileSync("git", ["-C", root, "ls-files", "-co", "--exclude-standard", "-z", "--", "*.md"], { encoding: "utf8", stdio: ["ignore", "pipe", "ignore"] }).split("\0").filter(Boolean);
146
+ const out = [];
147
+ for (const rel of listed) {
148
+ if (!patterns.some((p) => matchesGlob(rel, p))) continue;
149
+ if (DOC_EXCLUDES.some((p) => matchesGlob(rel, p))) continue;
150
+ let st;
151
+ try {
152
+ st = lstatSync(join(root, rel));
153
+ if (!st.isFile() || firstLine(join(root, rel)) === DRAFT_MARKER) continue;
154
+ } catch {
155
+ continue;
156
+ }
157
+ out.push({ path: rel, mtime_ms: Math.trunc(st.mtimeMs), size: st.size });
96
158
  }
97
159
  return out;
98
160
  }
@@ -145,17 +207,34 @@ async function openDb(root) {
145
207
  }
146
208
  }
147
209
 
210
+ function fenceMask(lines) {
211
+ const mask = new Array(lines.length).fill(false);
212
+ let open = null;
213
+ for (let i = 0; i < lines.length; i++) {
214
+ const m = FENCE.exec(lines[i]);
215
+ if (open) {
216
+ mask[i] = true;
217
+ if (m && m[1][0] === open.ch && m[1].length >= open.len) open = null;
218
+ } else if (m) {
219
+ mask[i] = true;
220
+ open = { ch: m[1][0], len: m[1].length };
221
+ }
222
+ }
223
+ return mask;
224
+ }
225
+
148
226
  function extract(text, path) {
149
227
  const lines = text.split(/\r?\n/);
150
- const h1 = lines.findIndex((l) => l.startsWith("# "));
228
+ const fenced = fenceMask(lines);
229
+ const h1 = lines.findIndex((l, i) => !fenced[i] && l.startsWith("# "));
151
230
  const title = (h1 >= 0 ? lines[h1].slice(2).trim() : "") || basename(path, ".md");
152
231
  const goal = lines.find((l) => l.startsWith("**Goal:**"))?.slice("**Goal:**".length).trim() ?? "";
153
- const headings = lines.filter((l) => /^##{1,2} /.test(l)).map((l) => l.replace(/^#+ /, "")).join("\n");
232
+ const headings = lines.filter((l, i) => !fenced[i] && /^##{1,2} /.test(l)).map((l) => l.replace(/^#+ /, "")).join("\n");
154
233
  let state = "live";
155
234
  if (h1 >= 0) {
156
235
  // Title block: everything after the H1 up to the first H2 heading or code fence.
157
236
  for (const l of lines.slice(h1 + 1)) {
158
- if (/^## /.test(l) || /^\s*(`{3,}|~{3,})/.test(l)) break;
237
+ if (/^## /.test(l) || FENCE.test(l)) break;
159
238
  if (FULLY_BANNER.test(l)) { state = "superseded"; break; }
160
239
  }
161
240
  }
@@ -163,24 +242,25 @@ function extract(text, path) {
163
242
  }
164
243
 
165
244
  function refresh(db, root, corpus) {
166
- const known = new Map(db.prepare("SELECT path, mtime_ms, size FROM files").all().map((r) => [r.path, r]));
167
- const present = new Set(corpus.map((f) => f.path));
168
- const delSpec = db.prepare("DELETE FROM specs WHERE path = ?");
169
- const delFile = db.prepare("DELETE FROM files WHERE path = ?");
170
- const insSpec = db.prepare("INSERT INTO specs (path, service, title, goal, headings, body, state) VALUES (?, ?, ?, ?, ?, ?, ?)");
171
- const putFile = db.prepare("INSERT OR REPLACE INTO files (path, mtime_ms, size) VALUES (?, ?, ?)");
245
+ const { table } = corpus;
246
+ const found = corpus.discover(root);
247
+ const known = new Map(db.prepare("SELECT path, mtime_ms, size FROM files WHERE corpus = ?").all(table).map((r) => [r.path, r]));
248
+ const present = new Set(found.map((f) => f.path));
249
+ const delRow = db.prepare(`DELETE FROM ${table} WHERE path = ?`);
250
+ const delFile = db.prepare("DELETE FROM files WHERE corpus = ? AND path = ?");
251
+ const insRow = db.prepare(`INSERT INTO ${table} (${corpus.columns.join(", ")}) VALUES (${corpus.columns.map(() => "?").join(", ")})`);
252
+ const putFile = db.prepare("INSERT OR REPLACE INTO files (corpus, path, mtime_ms, size) VALUES (?, ?, ?, ?)");
172
253
  db.exec("BEGIN IMMEDIATE");
173
254
  try {
174
- for (const path of known.keys()) if (!present.has(path)) { delSpec.run(path); delFile.run(path); }
175
- for (const f of corpus) {
255
+ for (const path of known.keys()) if (!present.has(path)) { delRow.run(path); delFile.run(table, path); }
256
+ for (const f of found) {
176
257
  const k = known.get(f.path);
177
258
  if (k && k.mtime_ms === f.mtime_ms && k.size === f.size) continue;
178
259
  const text = readFileSync(join(root, f.path), "utf8");
179
- delSpec.run(f.path);
180
- if (text.split(/\r?\n/, 1)[0] === DRAFT_MARKER) { delFile.run(f.path); continue; }
181
- const x = extract(text, f.path);
182
- insSpec.run(f.path, f.service, x.title, x.goal, x.headings, text, x.state);
183
- putFile.run(f.path, f.mtime_ms, f.size);
260
+ delRow.run(f.path);
261
+ if (text.split(/\r?\n/, 1)[0] === DRAFT_MARKER) { delFile.run(table, f.path); continue; }
262
+ insRow.run(...corpus.row(text, f));
263
+ putFile.run(table, f.path, f.mtime_ms, f.size);
184
264
  }
185
265
  db.exec("COMMIT");
186
266
  } catch (e) {
@@ -212,34 +292,34 @@ function representatives(db, toks) {
212
292
  function confidenceFilter(rows, evidence) {
213
293
  return rows.filter((r) => {
214
294
  const hits = evidence.filter((e) => e.all.has(r.rowid)).length;
215
- const titleGoal = evidence.filter((e) => e.titleGoal.has(r.rowid)).length;
216
- return hits >= MIN_EVIDENCE_TOKENS && titleGoal >= MIN_EVIDENCE_TOKENS;
295
+ const strong = evidence.filter((e) => e.strong.has(r.rowid)).length;
296
+ return hits >= MIN_EVIDENCE_TOKENS && strong >= MIN_EVIDENCE_TOKENS;
217
297
  });
218
298
  }
219
299
 
220
- function query(db, toks, { limit, exclude }) {
221
- const n = db.prepare("SELECT count(*) AS n FROM specs").get().n;
222
- const rowids = (match) => new Set(db.prepare("SELECT rowid FROM specs WHERE specs MATCH ?").all(match).map((r) => r.rowid));
300
+ function query(db, corpus, toks, { limit, exclude }) {
301
+ const { table } = corpus;
302
+ const n = db.prepare(`SELECT count(*) AS n FROM ${table}`).get().n;
303
+ const rowids = (match) => new Set(db.prepare(`SELECT rowid FROM ${table} WHERE ${table} MATCH ?`).all(match).map((r) => r.rowid));
223
304
  const reps = representatives(db, toks);
224
305
  const terms = [...new Set(reps.values())];
225
306
  const evidence = [];
226
307
  for (const tok of terms) {
227
308
  const all = rowids(quote(tok));
228
309
  if (all.size === 0 || all.size >= n * EVIDENCE_DF_FRACTION) continue;
229
- evidence.push({ all, titleGoal: rowids(`title:${quote(tok)} OR goal:${quote(tok)}`) });
310
+ evidence.push({ all, strong: rowids(corpus.strong.map((f) => `${f}:${quote(tok)}`).join(" OR ")) });
230
311
  }
231
- if (evidence.length < MIN_EVIDENCE_TOKENS) return [];
312
+ if (evidence.length < MIN_EVIDENCE_TOKENS) return { rows: [], n };
232
313
  const rows = db.prepare(
233
- `SELECT rowid, path, service, title, state,
234
- bm25(specs, 0, 0, 10.0, 5.0, 2.0, 1.0, 0) AS score,
235
- snippet(specs, 5, '', '', '...', 12) AS snippet
236
- FROM specs WHERE specs MATCH ?`,
314
+ `SELECT rowid, ${corpus.select}, bm25(${table}, ${corpus.weights.map((w) => w.toFixed(1)).join(", ")}) AS score
315
+ FROM ${table} WHERE ${table} MATCH ?`,
237
316
  ).all(toMatch(terms)).filter((r) => !exclude.includes(r.path));
238
317
  const kept = confidenceFilter(rows, evidence);
239
- kept.sort((a, b) => (a.state !== "live") - (b.state !== "live") || a.score - b.score);
240
- const best = kept.find((r) => r.state === "live");
318
+ const live = (r) => !corpus.hasState || r.state === "live";
319
+ kept.sort((a, b) => Number(!live(a)) - Number(!live(b)) || a.score - b.score);
320
+ const best = kept.find(live);
241
321
  const cut = best ? kept.filter((r) => Math.abs(r.score) >= SCORE_RATIO * Math.abs(best.score)) : kept;
242
- return cut.slice(0, limit);
322
+ return { rows: cut.slice(0, limit), n };
243
323
  }
244
324
 
245
325
  function telemetry(root, specPath) {
@@ -265,21 +345,64 @@ const cell = (v) => (v === null || v === undefined ? "" : String(v).replace(/\s+
265
345
  // Paths keep their spaces; only column and row delimiters are neutralised.
266
346
  const filesCell = (v) => v.replace(/[\t\r\n]+/g, " ");
267
347
 
348
+ function docSnippet(r) {
349
+ if (!r.hsnip.includes("\u0001")) return r.bsnip;
350
+ const fragment = r.hsnip.split("\n").find((line) => line.includes("\u0001")).replace(/[\u0001\u0002]/g, "");
351
+ return r.headings.split("\n").find((line) => line.includes(fragment));
352
+ }
353
+
354
+ const CORPORA = {
355
+ specs: {
356
+ table: "specs",
357
+ columns: ["path", "service", "title", "goal", "headings", "body", "state"],
358
+ strong: ["title", "goal"],
359
+ weights: [0, 0, 10, 5, 2, 1, 0],
360
+ hasState: true,
361
+ header: ["score", "path", "service", "title", "status", "shipped_at", "state", "files", "snippet"],
362
+ select: "path, service, title, state, snippet(specs, 5, '', '', '...', 12) AS snippet",
363
+ discover,
364
+ row: (text, f) => {
365
+ const x = extract(text, f.path);
366
+ return [f.path, f.service, x.title, x.goal, x.headings, text, x.state];
367
+ },
368
+ format: (root, r) => {
369
+ const t = telemetry(root, r.path);
370
+ return [...[r.score.toFixed(3), r.path, r.service, r.title, t.status, t.shipped_at, r.state].map(cell), filesCell(t.files), cell(r.snippet)];
371
+ },
372
+ },
373
+ docs: {
374
+ table: "docs",
375
+ columns: ["path", "title", "headings", "body"],
376
+ strong: ["title", "headings"],
377
+ weights: [0, 10, 3, 1],
378
+ hasState: false,
379
+ header: ["score", "path", "title", "snippet"],
380
+ select: "path, title, headings, snippet(docs, 2, '\u0001', '\u0002', '', 12) AS hsnip, snippet(docs, 3, '', '', '...', 12) AS bsnip",
381
+ discover: (root) => discoverDocs(root, includes(root)),
382
+ row: (text, f) => {
383
+ const x = extract(text, f.path);
384
+ return [f.path, x.title, x.headings, text];
385
+ },
386
+ format: (root, r) => [r.score.toFixed(3), r.path, r.title, docSnippet(r)].map(cell),
387
+ },
388
+ };
389
+
268
390
  async function main() {
269
391
  if (!nodeOk()) die(`gauntlet-spec-index needs Node >=24.15.0 (found ${process.versions.node})`);
270
- const { query: text, limit, exclude } = parseArgs(process.argv.slice(2));
392
+ const { query: text, corpus: name, limit, exclude } = parseArgs(process.argv.slice(2));
271
393
  const toks = tokens(text);
272
394
  if (!toks.length) usage();
273
395
  const root = repoRoot();
396
+ const corpus = CORPORA[name];
274
397
  const db = await openDb(root);
275
- refresh(db, root, discover(root));
276
- const rows = query(db, toks, { limit, exclude });
277
- const out = [HEADER.join("\t")];
278
- for (const r of rows) {
279
- const t = telemetry(root, r.path);
280
- out.push([...[r.score.toFixed(3), r.path, r.service, r.title, t.status, t.shipped_at, r.state].map(cell), filesCell(t.files), cell(r.snippet)].join("\t"));
281
- }
398
+ refresh(db, root, corpus);
399
+ const { rows, n } = query(db, corpus, toks, { limit, exclude });
400
+ const out = [corpus.header.join("\t")];
401
+ for (const r of rows) out.push(corpus.format(root, r).join("\t"));
282
402
  process.stdout.write(out.join("\n") + "\n");
403
+ if (rows.length === 0 && n < MIN_CORPUS) {
404
+ process.stderr.write(`gauntlet-spec-index: ${name} corpus has ${n} documents - too small for the confidence rule; no rows returned\n`);
405
+ }
283
406
  db.close();
284
407
  }
285
408
 
@@ -1,6 +1,6 @@
1
1
  import { test } from "node:test";
2
2
  import assert from "node:assert/strict";
3
- import { mkdtempSync, mkdirSync, writeFileSync, readFileSync, rmSync, utimesSync } from "node:fs";
3
+ import { mkdtempSync, mkdirSync, writeFileSync, readFileSync, rmSync, utimesSync, symlinkSync } from "node:fs";
4
4
  import { join, dirname } from "node:path";
5
5
  import { tmpdir } from "node:os";
6
6
  import { spawnSync } from "node:child_process";
@@ -10,6 +10,22 @@ import { DatabaseSync } from "node:sqlite";
10
10
  const CLI = join(dirname(fileURLToPath(import.meta.url)), "gauntlet-spec-index.mjs");
11
11
  const DRAFT = "# CONTEXT DRAFT - NOT A SPEC - fully replaced at spec-writing";
12
12
  const HEADER = ["score", "path", "service", "title", "status", "shipped_at", "state", "files", "snippet"];
13
+ const DOCS_HEADER = ["score", "path", "title", "snippet"];
14
+ const docFiller = (n) => `# Filler ${n}\n\n## Design\n\nnothing of note here\n`;
15
+ const decoy = "# ptarmigan gannet decoy\n\n## ptarmigan gannet\n\nbody\n";
16
+ const docsRepo = () => {
17
+ const root = gitRepo();
18
+ write(root, "doc/guide.md", "# Guide\n\n## Configuring the ptarmigan gannet\n\nprose about setup\n");
19
+ write(root, "README.md", "# Repo\n\nThe body mentions ptarmigan once.\n");
20
+ for (let n = 1; n <= 4; n++) write(root, `doc/filler-${n}.md`, docFiller(n));
21
+ for (const p of ["doc/specs/a.md", "docs/specs/b.md", "doc/plans/p.md", "docs/plans/q.md", "node_modules/x/doc/a.md", "build/doc/b.md", ".pi/gauntlet/n.md", ".worktrees/w/doc/g.md"]) write(root, p, decoy);
22
+ write(root, ".gitignore", "build/\n");
23
+ write(root, "doc/draft.md", `${DRAFT}\n\n## ptarmigan gannet\n`);
24
+ commit(root);
25
+ return root;
26
+ };
27
+ const docsPaths = (root) => withDb(root, (db) => db.prepare("SELECT path FROM docs ORDER BY path").all().map((r) => r.path));
28
+ const overrides = (root, globs) => write(root, ".pi/gauntlet-overrides.md", `## Issue tracker\ntracker: none\n\n## Spec index\n${globs.map((g) => `- docs: \`${g}\``).join("\n")}\n\n## release\nnothing\n`);
13
29
 
14
30
  const write = (root, rel, text) => {
15
31
  mkdirSync(join(root, dirname(rel)), { recursive: true });
@@ -27,6 +43,16 @@ const commit = (root) => {
27
43
  };
28
44
  // Fillers keep the corpus large enough that two-document probe terms stay below N/2.
29
45
  const filler = (n) => `# Filler ${n}\n\n**Goal:** filler goal ${n}.\n\n## Design\n\nnothing of note here\n`;
46
+ const boundaryRepo = (fillers) => {
47
+ const root = gitRepo();
48
+ const common = (n) => `# Plain spec ${n}\n\n**Goal:** plain goal ${n}.\n\nThe body mentions the common shared vocabulary.\n`;
49
+ write(root, "doc/specs/strong.md", "# Strong heron ibis\n\n**Goal:** heron ibis jackal.\n\nbody\n");
50
+ for (let n = 1; n <= 5; n++) write(root, `doc/specs/p${n}.md`, common(n));
51
+ for (const f of fillers) write(root, `doc/specs/filler-${f}.md`, filler(f));
52
+ commit(root);
53
+ return root;
54
+ };
55
+ const NOTE = "too small for the confidence rule; no rows returned";
30
56
 
31
57
  const repo = () => {
32
58
  const root = gitRepo();
@@ -44,7 +70,7 @@ const repo = () => {
44
70
  const run = (cwd, args) => {
45
71
  const r = spawnSync(process.execPath, [CLI, ...args], { cwd, encoding: "utf8" });
46
72
  const lines = r.stdout.split("\n").filter(Boolean);
47
- return { status: r.status, stderr: r.stderr, header: lines[0]?.split("\t"), rows: lines.slice(1).map((l) => l.split("\t")) };
73
+ return { status: r.status, stderr: r.stderr, stdout: r.stdout, header: lines[0]?.split("\t"), rows: lines.slice(1).map((l) => l.split("\t")) };
48
74
  };
49
75
  const paths = (res) => res.rows.map((r) => r[1]);
50
76
  const withDb = (root, fn) => {
@@ -104,7 +130,7 @@ test("4: schema_version mismatch rebuilds the db", (t) => {
104
130
  const r = run(root, ["--query", "zephyr widget"]);
105
131
  assert.equal(r.status, 0, r.stderr);
106
132
  assert.equal(paths(r).length, 2);
107
- assert.equal(withDb(root, (database) => database.prepare("SELECT value FROM meta WHERE key = 'schema_version'").get().value), "2");
133
+ assert.equal(withDb(root, (database) => database.prepare("SELECT value FROM meta WHERE key = 'schema_version'").get().value), "3");
108
134
  });
109
135
 
110
136
  test("5: a non-SQLite database is rebuilt", (t) => {
@@ -114,7 +140,7 @@ test("5: a non-SQLite database is rebuilt", (t) => {
114
140
  const r = run(root, ["--query", "zephyr widget"]);
115
141
  assert.equal(r.status, 0, r.stderr);
116
142
  assert.deepEqual(paths(r), ["doc/specs/a.md", "svc-a/doc/specs/b.md"]);
117
- assert.equal(withDb(root, (database) => database.prepare("SELECT value FROM meta WHERE key = 'schema_version'").get().value), "2");
143
+ assert.equal(withDb(root, (database) => database.prepare("SELECT value FROM meta WHERE key = 'schema_version'").get().value), "3");
118
144
  });
119
145
 
120
146
  test("6: telemetry join is output-only and tolerant", (t) => {
@@ -340,3 +366,262 @@ test("16: punctuation splits into distinct evidence terms and query order does n
340
366
  assert.deepEqual(paths(distinct), ["doc/specs/cfg.md"]);
341
367
  assert.deepEqual(paths(run(root, ["--query", "config config-file"])), paths(run(root, ["--query", "config-file config"])));
342
368
  });
369
+
370
+ test("21: three-spec corpus, no match - header only plus the small-corpus note", (t) => {
371
+ const root = gitRepo();
372
+ t.after(() => rmSync(root, { recursive: true, force: true }));
373
+ for (const n of ["one", "two", "three"]) write(root, `doc/specs/filler-${n}.md`, filler(n));
374
+ commit(root);
375
+ const r = run(root, ["--query", "yeti unicorn"]);
376
+ assert.equal(r.status, 0, r.stderr);
377
+ assert.deepEqual(r.header, HEADER);
378
+ assert.deepEqual(r.rows, []);
379
+ assert.equal(r.stderr, `gauntlet-spec-index: specs corpus has 3 documents - ${NOTE}\n`);
380
+ });
381
+
382
+ test("22: MIN_CORPUS boundary - note at N = 9, silent at N = 10, rows unaffected", (t) => {
383
+ const root = boundaryRepo(["one", "two", "three"]);
384
+ t.after(() => rmSync(root, { recursive: true, force: true }));
385
+ const nine = run(root, ["--query", "yeti unicorn"]);
386
+ assert.equal(nine.status, 0, nine.stderr);
387
+ assert.deepEqual(nine.rows, []);
388
+ assert.equal(nine.stderr, `gauntlet-spec-index: specs corpus has 9 documents - ${NOTE}\n`);
389
+ write(root, "doc/specs/filler-four.md", filler("four"));
390
+ commit(root);
391
+ for (const q of ["yeti unicorn", "common shared"]) {
392
+ const r = run(root, ["--query", q]);
393
+ assert.equal(r.status, 0, r.stderr);
394
+ assert.deepEqual(r.header, HEADER);
395
+ assert.deepEqual(r.rows, []);
396
+ assert.equal(r.stderr, "");
397
+ }
398
+ const strong = run(root, ["--query", "heron ibis jackal common", "--limit", "10"]);
399
+ assert.equal(strong.status, 0, strong.stderr);
400
+ assert.deepEqual(paths(strong), ["doc/specs/strong.md"]);
401
+ assert.equal(strong.stderr, "");
402
+ });
403
+
404
+ test("23: n is the table count before --exclude", (t) => {
405
+ const root = boundaryRepo(["one", "two", "three", "four"]);
406
+ t.after(() => rmSync(root, { recursive: true, force: true }));
407
+ const r = run(root, ["--query", "yeti unicorn", "--exclude", "doc/specs/p1.md"]);
408
+ assert.equal(r.status, 0, r.stderr);
409
+ assert.deepEqual(r.header, HEADER);
410
+ assert.deepEqual(r.rows, []);
411
+ assert.equal(r.stderr, "");
412
+ });
413
+
414
+ test("D1: --corpus specs and omitted --corpus are byte-identical; nine columns", (t) => {
415
+ const root = repo();
416
+ t.after(() => rmSync(root, { recursive: true, force: true }));
417
+ const plain = run(root, ["--query", "zephyr widget"]);
418
+ const explicit = run(root, ["--query", "zephyr widget", "--corpus", "specs"]);
419
+ assert.equal(explicit.status, 0, explicit.stderr);
420
+ assert.equal(explicit.stdout, plain.stdout);
421
+ assert.deepEqual(explicit.header, HEADER);
422
+ });
423
+
424
+ test("D11: unknown --corpus exits 1 with usage", (t) => {
425
+ const root = repo();
426
+ t.after(() => rmSync(root, { recursive: true, force: true }));
427
+ const r = run(root, ["--query", "zephyr widget", "--corpus", "notes"]);
428
+ assert.equal(r.status, 1);
429
+ assert.match(r.stderr, /usage:/);
430
+ assert.equal(run(root, ["--query", "zephyr widget", "--corpus"]).status, 1);
431
+ });
432
+
433
+ test("D12: a schema 2 cache is rebuilt to 3 with both tables present", (t) => {
434
+ const root = repo();
435
+ t.after(() => rmSync(root, { recursive: true, force: true }));
436
+ run(root, ["--query", "zephyr widget"]);
437
+ withDb(root, (database) => database.prepare("UPDATE meta SET value = '2' WHERE key = 'schema_version'").run());
438
+ const r = run(root, ["--query", "zephyr widget"]);
439
+ assert.equal(r.status, 0, r.stderr);
440
+ assert.equal(withDb(root, (database) => database.prepare("SELECT value FROM meta WHERE key = 'schema_version'").get().value), "3");
441
+ const tables = withDb(root, (database) => database.prepare("SELECT name FROM sqlite_master WHERE name IN ('specs', 'docs') ORDER BY name").all().map((row) => row.name));
442
+ assert.deepEqual(tables, ["docs", "specs"]);
443
+ const cols = withDb(root, (database) => database.prepare("PRAGMA table_info(files)").all().map((row) => row.name));
444
+ assert.deepEqual(cols, ["corpus", "path", "mtime_ms", "size"]);
445
+ });
446
+
447
+ test("D10-specs: a `## ` line inside a code fence is not a heading", (t) => {
448
+ const root = repo();
449
+ t.after(() => rmSync(root, { recursive: true, force: true }));
450
+ write(root, "doc/specs/fenced.md", "# Fenced pelican\n\n**Goal:** pelican.\n\n```markdown\n## Not a heading pelican\n```\n\n## Real heading\n");
451
+ run(root, ["--query", "zephyr widget"]);
452
+ const h = withDb(root, (database) => database.prepare("SELECT headings FROM specs WHERE path = 'doc/specs/fenced.md'").get().headings);
453
+ assert.equal(h, "Real heading");
454
+ });
455
+
456
+ test("D2: docs include defaults and fixed exclusions", (t) => {
457
+ const root = docsRepo();
458
+ t.after(() => rmSync(root, { recursive: true, force: true }));
459
+ const r = run(root, ["--corpus", "docs", "--query", "ptarmigan gannet"]);
460
+ assert.equal(r.status, 0, r.stderr);
461
+ assert.deepEqual(r.header, DOCS_HEADER);
462
+ assert.deepEqual(paths(r), ["doc/guide.md"]);
463
+ assert.equal(r.rows[0].length, 4);
464
+ assert.deepEqual(docsPaths(root), ["README.md", ...[1, 2, 3, 4].map((n) => `doc/filler-${n}.md`), "doc/guide.md"]);
465
+ overrides(root, [".pi/**/*.md", ".worktrees/**/*.md", "doc/**/*.md"]);
466
+ assert.equal(run(root, ["--corpus", "docs", "--query", "ptarmigan gannet"]).status, 0);
467
+ assert.deepEqual(docsPaths(root), [".pi/gauntlet-overrides.md", ...[1, 2, 3, 4].map((n) => `doc/filler-${n}.md`), "doc/guide.md"]);
468
+ const two = gitRepo();
469
+ t.after(() => rmSync(two, { recursive: true, force: true }));
470
+ write(two, "doc/x.md", decoy);
471
+ write(two, "doc/y.md", decoy);
472
+ commit(two);
473
+ const tiny = run(two, ["--corpus", "docs", "--query", "ptarmigan gannet"]);
474
+ assert.equal(tiny.status, 0, tiny.stderr);
475
+ assert.deepEqual(tiny.header, DOCS_HEADER);
476
+ assert.deepEqual(tiny.rows, []);
477
+ assert.equal(tiny.stderr, `gauntlet-spec-index: docs corpus has 2 documents - ${NOTE}\n`);
478
+ });
479
+
480
+ test("D3: shared confidence rejects common body-only hits", (t) => {
481
+ const root = gitRepo();
482
+ t.after(() => rmSync(root, { recursive: true, force: true }));
483
+ write(root, "doc/strong.md", "# Strong heron ibis\n\n## jackal\n\nbody\n");
484
+ for (let n = 1; n <= 5; n++) write(root, `doc/p${n}.md`, `# Plain doc ${n}\n\nThe body mentions the common shared vocabulary.\n`);
485
+ commit(root);
486
+ const r = run(root, ["--corpus", "docs", "--query", "heron ibis jackal common", "--limit", "10"]);
487
+ assert.equal(r.status, 0, r.stderr);
488
+ assert.deepEqual(paths(r), ["doc/strong.md"]);
489
+ });
490
+
491
+ test("D4: no cross-corpus leakage", (t) => {
492
+ const root = repo();
493
+ t.after(() => rmSync(root, { recursive: true, force: true }));
494
+ write(root, "doc/guide.md", "# zephyr widget guide\n\n## zephyr widget\n");
495
+ for (let n = 1; n <= 4; n++) write(root, `doc/filler-${n}.md`, docFiller(n));
496
+ assert.deepEqual(paths(run(root, ["--query", "zephyr widget"])), ["doc/specs/a.md", "svc-a/doc/specs/b.md"]);
497
+ const docs = run(root, ["--corpus", "docs", "--query", "zephyr widget"]);
498
+ assert.deepEqual(paths(docs), ["doc/guide.md"]);
499
+ assert.deepEqual(docs.header, DOCS_HEADER);
500
+ });
501
+
502
+ test("D5: independent cache refresh in both directions", (t) => {
503
+ const root = repo();
504
+ t.after(() => rmSync(root, { recursive: true, force: true }));
505
+ write(root, "doc/guide.md", "# Guide\n\n## zephyr widget\n");
506
+ run(root, ["--query", "zephyr widget"]);
507
+ run(root, ["--corpus", "docs", "--query", "zephyr widget"]);
508
+ const files = () => withDb(root, (db) => Object.fromEntries(db.prepare("SELECT corpus, path, mtime_ms FROM files").all().map((r) => [`${r.corpus}:${r.path}`, r.mtime_ms])));
509
+ const before = files();
510
+ utimesSync(join(root, "doc/specs/a.md"), new Date(), new Date(Date.now() + 5000));
511
+ run(root, ["--corpus", "docs", "--query", "zephyr widget"]);
512
+ assert.deepEqual(files(), before);
513
+ run(root, ["--query", "zephyr widget"]);
514
+ assert.notEqual(files()["specs:doc/specs/a.md"], before["specs:doc/specs/a.md"]);
515
+ assert.equal(files()["docs:doc/guide.md"], before["docs:doc/guide.md"]);
516
+ assert.deepEqual(docsPaths(root), ["doc/guide.md"]);
517
+ });
518
+
519
+ test("D6: overrides replace defaults but cannot bypass exclusions", (t) => {
520
+ const root = docsRepo();
521
+ t.after(() => rmSync(root, { recursive: true, force: true }));
522
+ write(root, "notes/a.md", decoy);
523
+ write(root, "CONTRIBUTING.md", decoy);
524
+ write(root, "svc/CONTRIBUTING.md", decoy);
525
+ write(root, "svc/doc/inner.md", decoy);
526
+ commit(root);
527
+ const search = () => run(root, ["--corpus", "docs", "--query", "ptarmigan gannet"]);
528
+ overrides(root, ["notes/**/*.md"]);
529
+ assert.equal(search().status, 0);
530
+ assert.deepEqual(docsPaths(root), ["notes/a.md"]);
531
+ overrides(root, ["**/*.md"]);
532
+ assert.equal(search().status, 0);
533
+ const all = docsPaths(root);
534
+ for (const p of ["doc/specs/a.md", "docs/specs/b.md", "doc/plans/p.md", "docs/plans/q.md", "node_modules/x/doc/a.md", "build/doc/b.md", "doc/draft.md", ".pi/gauntlet/n.md", ".worktrees/w/doc/g.md"]) assert.ok(!all.includes(p), p);
535
+ assert.ok(all.includes("doc/guide.md"));
536
+ overrides(root, ["CONTRIBUTING.md"]);
537
+ assert.equal(search().status, 0);
538
+ assert.deepEqual(docsPaths(root), ["CONTRIBUTING.md", "svc/CONTRIBUTING.md"]);
539
+ write(root, ".pi/gauntlet-overrides.md", "## Spec index\n\nno bullets here\n");
540
+ assert.equal(search().status, 0);
541
+ assert.ok(docsPaths(root).includes("doc/guide.md"));
542
+ assert.ok(docsPaths(root).includes("svc/doc/inner.md"));
543
+ });
544
+
545
+ test("D6b: overrides file resolution order", (t) => {
546
+ const root = docsRepo();
547
+ t.after(() => rmSync(root, { recursive: true, force: true }));
548
+ write(root, "doc/gauntlet-overrides.md", "## Spec index\n- docs: `notes/**/*.md`\n");
549
+ write(root, "notes/a.md", decoy);
550
+ const search = () => run(root, ["--corpus", "docs", "--query", "ptarmigan gannet"]);
551
+ assert.equal(search().status, 0);
552
+ assert.deepEqual(docsPaths(root), ["notes/a.md"]);
553
+ write(root, "gauntlet-overrides.md", "## Spec index\n- docs: `other/**/*.md`\n");
554
+ write(root, "other/b.md", decoy);
555
+ assert.equal(search().status, 0);
556
+ assert.deepEqual(docsPaths(root), ["other/b.md"]);
557
+ overrides(root, ["doc/**/*.md"]);
558
+ assert.equal(search().status, 0);
559
+ assert.deepEqual(docsPaths(root), ["doc/filler-1.md", "doc/filler-2.md", "doc/filler-3.md", "doc/filler-4.md", "doc/gauntlet-overrides.md", "doc/guide.md"]);
560
+ });
561
+
562
+ test("D7: heading hits are strong, body-only hits are not", (t) => {
563
+ const root = gitRepo();
564
+ t.after(() => rmSync(root, { recursive: true, force: true }));
565
+ write(root, "doc/h.md", "# H\n\n## Setup\n\n## Feeding the kudu wombat\n\nprose\n");
566
+ write(root, "doc/b.md", "# B\n\n## Setup\n\nkudu wombat in the body only\n");
567
+ for (let n = 1; n <= 4; n++) write(root, `doc/filler-${n}.md`, docFiller(n));
568
+ commit(root);
569
+ const r = run(root, ["--corpus", "docs", "--query", "kudu wombat"]);
570
+ assert.equal(r.status, 0, r.stderr);
571
+ assert.deepEqual(paths(r), ["doc/h.md"]);
572
+ });
573
+
574
+ test("D8: snippet selects heading line or body fallback", (t) => {
575
+ const root = gitRepo();
576
+ t.after(() => rmSync(root, { recursive: true, force: true }));
577
+ const longHeading = "Feeding the wandering herds across twelve distant valleys before finally finding the kudu wombat";
578
+ write(root, "doc/h.md", `# H\n\n## Setup\n\n## ${longHeading}\n\nprose\n`);
579
+ write(root, "doc/t.md", "# lynx ocelot title\n\n## Unrelated\n\nbody says lynx ocelot here\n");
580
+ for (let n = 1; n <= 4; n++) write(root, `doc/filler-${n}.md`, docFiller(n));
581
+ commit(root);
582
+ const heading = run(root, ["--corpus", "docs", "--query", "kudu wombat"]);
583
+ assert.deepEqual(paths(heading), ["doc/h.md"]);
584
+ assert.equal(heading.rows[0][3], longHeading);
585
+ const body = run(root, ["--corpus", "docs", "--query", "lynx ocelot"]);
586
+ assert.deepEqual(paths(body), ["doc/t.md"]);
587
+ assert.match(body.rows[0][3], /body says lynx ocelot/);
588
+ assert.doesNotMatch(body.rows[0][3], /[\u0001\u0002]/);
589
+ });
590
+
591
+ test("D9: deleted files and symlinks are skipped, stale rows removed", (t) => {
592
+ const root = docsRepo();
593
+ t.after(() => rmSync(root, { recursive: true, force: true }));
594
+ write(root, "doc/gone.md", "# Gone\n\n## ptarmigan gannet\n");
595
+ commit(root);
596
+ run(root, ["--corpus", "docs", "--query", "ptarmigan gannet"]);
597
+ assert.ok(docsPaths(root).includes("doc/gone.md"));
598
+ rmSync(join(root, "doc/gone.md"));
599
+ symlinkSync(join(root, "doc/specs/a.md"), join(root, "doc/link.md"));
600
+ symlinkSync(join(root, "doc/nowhere.md"), join(root, "doc/dangling.md"));
601
+ const r = run(root, ["--corpus", "docs", "--query", "ptarmigan gannet"]);
602
+ assert.equal(r.status, 0, r.stderr);
603
+ assert.deepEqual(paths(r), ["doc/guide.md"]);
604
+ for (const p of ["doc/gone.md", "doc/link.md", "doc/dangling.md"]) assert.ok(!docsPaths(root).includes(p), p);
605
+ assert.equal(withDb(root, (db) => db.prepare("SELECT count(*) AS c FROM files WHERE corpus = 'docs' AND path = 'doc/gone.md'").get().c), 0);
606
+ });
607
+
608
+ test("D10: fenced headings and shorter nested fences are ignored", (t) => {
609
+ const root = docsRepo();
610
+ t.after(() => rmSync(root, { recursive: true, force: true }));
611
+ write(root, "doc/f.md", "# F\n\n```\n## kudu wombat\n```\n\nbody\n");
612
+ write(root, "doc/nested.md", "# N\n\n````markdown\n```js\n## inner heading\n```\n## still fenced\n````\n");
613
+ run(root, ["--corpus", "docs", "--query", "ptarmigan gannet"]);
614
+ for (const p of ["doc/f.md", "doc/nested.md"]) assert.equal(withDb(root, (db) => db.prepare("SELECT headings FROM docs WHERE path = ?").get(p).headings), "");
615
+ });
616
+
617
+ test("D13: five-document docs corpus with rare heading terms returns rows and no note", (t) => {
618
+ const root = gitRepo();
619
+ t.after(() => rmSync(root, { recursive: true, force: true }));
620
+ write(root, "doc/strong.md", "# Strong doc\n\n## heron ibis\n\n## jackal\n\nbody\n");
621
+ for (let n = 1; n <= 4; n++) write(root, `doc/p${n}.md`, `# Plain doc ${n}\n\nThe body mentions the common shared vocabulary.\n`);
622
+ commit(root);
623
+ const r = run(root, ["--corpus", "docs", "--query", "heron ibis jackal common", "--limit", "10"]);
624
+ assert.equal(r.status, 0, r.stderr);
625
+ assert.deepEqual(paths(r), ["doc/strong.md"]);
626
+ assert.equal(r.stderr, "");
627
+ });
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-gauntlet",
3
- "version": "5.22.0",
3
+ "version": "5.23.1",
4
4
  "description": "Opinionated, gated workflow skills, subagent personas, and runtime extensions for the pi coding agent.",
5
5
  "author": "Jacek Juraszek",
6
6
  "type": "module",
@@ -104,7 +104,7 @@ Prefer clear testable boundaries, YAGNI, existing conventions, the owning schema
104
104
 
105
105
  ### 6. Present the design in two rounds
106
106
 
107
- Use two rounds, targeting 300-500 words each, with one approval each; revisions remain within that approval point. Ask once per round; round-1 approval without correction confirms the predecessor. Round 1 covers architecture, responsibilities, data flow, and `supersedes <path>, <scope>` when applicable. Round 2 covers errors, edges, tests, and `## Documentation impact`.
107
+ Use two rounds, targeting 300-500 words each, with one approval each; revisions remain within that approval point. Ask once per round; round-1 approval without correction confirms the predecessor. Round 1 covers architecture, responsibilities, data flow, and `supersedes <path>, <scope>` when applicable. Round 2 covers errors, edges, tests, and `## Documentation impact`. In round 2, cite the draft's `Docs touched:` entries as candidates for "Materially amended existing docs"; admit or drop each by the materiality bar in `reference/documentation-impact.md`, never list them automatically.
108
108
 
109
109
  `## Documentation impact` is required. Cite `reference/documentation-impact.md` by relative path, do not restate its categories, and reproduce this template verbatim:
110
110
 
@@ -58,7 +58,11 @@ Scout (always dispatched):
58
58
  > `node <SPEC_INDEX> --query '<keywords>' --limit 10` from the worktree root,
59
59
  > keeping the keywords inside single quotes, and treat its rows as the candidate
60
60
  > list; zero rows means the index found no evidence, not that no predecessor exists -
61
- > judge `Predecessor: none` from the code recon. If the command fails, fall back to
61
+ > judge `Predecessor: none` from the code recon. When the command output contains a line
62
+ > ending "too small for the confidence rule; no rows returned" (a diagnostic, not a row) and
63
+ > the code recon names no predecessor, write
64
+ > `Predecessor: none (specs corpus too small: <n> documents)`, copying <n> from that
65
+ > line. If the command fails, fall back to
62
66
  > listing the project's spec directory
63
67
  > and reading titles and `**Goal:**` lines, and write
64
68
  > `Spec index unavailable - predecessor check used directory listing.` in your
@@ -82,7 +86,25 @@ Scout (always dispatched):
82
86
  > from the directory-listing fallback) contributes no path and no `missing` line. When
83
87
  > no path and no `missing` line results - including whenever the index was unavailable -
84
88
  > omit the section entirely; `Predecessor: none` produces no anchors section. The anchors
85
- > are a recon hint, never a selection input. End with an
89
+ > are a recon hint, never a selection input.
90
+ > Docs check: run `node <SPEC_INDEX> --corpus docs --query '<the same keywords>' --limit 10`
91
+ > from the worktree root. Its rows are project documentation that already speaks about the
92
+ > request's topic; open at most five whose topic matches and let them inform the "already
93
+ > solves this?" finding and the current-contract recon. If the command fails, write
94
+ > `Docs index unavailable - docs check used recon only.` in your handoff and continue. After
95
+ > the `Predecessor:` line(s) and any `Predecessor anchors` section, render at most
96
+ > five `Docs touched: <path> - <section heading>` lines, one per document that
97
+ > genuinely covers the request's topic - an opened index row or a document your own
98
+ > recon found - where
99
+ > `<section heading>` is the `##`/`###` heading of the covering section as read in the
100
+ > document (the row's `snippet` column is a hint to it), or the document title when no single
101
+ > section applies; or `Docs touched: none`. Judge by topic; a lexical hit alone is not
102
+ > coverage. Zero rows means the index found no evidence, not that no document covers the
103
+ > topic. When the docs command output contains that same diagnostic line and the recon
104
+ > found no covering document, write
105
+ > `Docs touched: none (docs corpus too small: <n> documents)` instead. This query
106
+ > supplements the code and documentation recon you already perform; it
107
+ > never replaces it - keep reading the files the request touches. End with an
86
108
  > "Open questions that matter for the spec"
87
109
  > section. Compact handoff, not a dump.
88
110