woods 2.0.0.beta3 → 2.0.0.beta4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (73) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +77 -0
  3. data/CONTRIBUTING.md +29 -17
  4. data/README.md +92 -177
  5. data/docs/AGENT_GUIDE.md +26 -4
  6. data/docs/AGENT_SETUP.md +18 -8
  7. data/docs/BACKEND_MATRIX.md +5 -0
  8. data/docs/CONFIGURATION_REFERENCE.md +74 -8
  9. data/docs/CONSOLE_MCP_SETUP.md +45 -2
  10. data/docs/DOCKER_SETUP.md +1 -1
  11. data/docs/EXTRACTOR_REFERENCE.md +9 -1
  12. data/docs/INCREMENTAL_EXTRACTION.md +30 -6
  13. data/docs/MCP_SERVERS.md +57 -2
  14. data/docs/MCP_TOOL_COOKBOOK.md +4 -4
  15. data/docs/MCP_WORKTREE_SETUP.md +43 -83
  16. data/docs/PUBLISHED_INDEX.md +17 -0
  17. data/docs/RETRIEVAL_GUIDE.md +24 -5
  18. data/docs/TROUBLESHOOTING.md +12 -13
  19. data/docs/UPGRADING_TO_2.md +6 -2
  20. data/docs/WATCH_DAEMON.md +18 -8
  21. data/exe/woods-mcp-start +14 -9
  22. data/lib/generators/woods/pgvector_generator.rb +8 -2
  23. data/lib/woods/agent_configuration/applier.rb +5 -3
  24. data/lib/woods/agent_configuration/cli.rb +2 -2
  25. data/lib/woods/agent_configuration/layout.rb +13 -0
  26. data/lib/woods/console/credential_scanner.rb +4 -3
  27. data/lib/woods/console/dispatch_pipeline.rb +7 -0
  28. data/lib/woods/console/embedded_executor.rb +31 -9
  29. data/lib/woods/console/sql_noise_stripper.rb +9 -7
  30. data/lib/woods/console/sql_table_scanner.rb +47 -7
  31. data/lib/woods/console/sql_validator.rb +49 -9
  32. data/lib/woods/console/sqlite_read_guard.rb +46 -0
  33. data/lib/woods/coordination/pipeline_lock.rb +3 -2
  34. data/lib/woods/embedding/indexer.rb +24 -14
  35. data/lib/woods/extractor.rb +45 -12
  36. data/lib/woods/extractors/declared_parent.rb +55 -0
  37. data/lib/woods/extractors/graphql_extractor.rb +2 -11
  38. data/lib/woods/extractors/lib_extractor.rb +10 -8
  39. data/lib/woods/extractors/mailer_extractor.rb +6 -10
  40. data/lib/woods/extractors/model_extractor.rb +1 -15
  41. data/lib/woods/extractors/poro_extractor.rb +10 -8
  42. data/lib/woods/extractors/shared_utility_methods.rb +22 -5
  43. data/lib/woods/mcp/bearer_auth.rb +2 -1
  44. data/lib/woods/mcp/bootstrapper.rb +17 -4
  45. data/lib/woods/mcp/config_resolver.rb +2 -1
  46. data/lib/woods/mcp/index_reader.rb +11 -2
  47. data/lib/woods/mcp/renderers/markdown_renderer.rb +14 -8
  48. data/lib/woods/mcp/renderers/plain_renderer.rb +11 -7
  49. data/lib/woods/mcp/server.rb +22 -28
  50. data/lib/woods/mcp/tool_contract.rb +1 -1
  51. data/lib/woods/mcp/tool_response_renderer.rb +16 -0
  52. data/lib/woods/mcp/traversal_evidence_text.rb +1 -1
  53. data/lib/woods/mcp/traversal_response.rb +22 -0
  54. data/lib/woods/path_dispatcher.rb +6 -5
  55. data/lib/woods/published_index/typed_unit_reader.rb +40 -3
  56. data/lib/woods/published_index.rb +2 -2
  57. data/lib/woods/rake_helpers.rb +2 -12
  58. data/lib/woods/retrieval/lexical_assembler.rb +14 -3
  59. data/lib/woods/retrieval/lexical_index.rb +2 -1
  60. data/lib/woods/session_tracer/file_store.rb +6 -1
  61. data/lib/woods/source_inputs/consumer_errors.rb +4 -0
  62. data/lib/woods/storage/pgvector.rb +6 -2
  63. data/lib/woods/temporal/json_snapshot_store.rb +35 -7
  64. data/lib/woods/version.rb +1 -1
  65. data/lib/woods/watch/daemon.rb +18 -4
  66. data/plugin/.claude-plugin/plugin.json +1 -1
  67. data/plugin/hooks/woods-input-rules.sh +4 -4
  68. data/plugin/skills/woods-agent-enable/SKILL.md +7 -1
  69. data/plugin/skills/woods-diagnose/SKILL.md +64 -33
  70. data/plugin/skills/woods-investigate/SKILL.md +51 -12
  71. data/plugin/skills/woods-mcp-config/SKILL.md +11 -11
  72. data/plugin/skills/woods-setup/SKILL.md +14 -11
  73. metadata +8 -5
data/docs/AGENT_GUIDE.md CHANGED
@@ -129,18 +129,38 @@ Do not infer completeness from a full page or missing metadata. See the
129
129
 
130
130
  ## Traverse deliberately
131
131
 
132
- `dependencies` means “what this unit uses.” `dependents` means “what uses this unit.” Both default to bounded breadth-first traversal and accept type or relationship filters.
132
+ `dependencies` shows recorded relationships from this unit; `dependents` shows
133
+ recorded relationships to it. Both default to bounded breadth-first traversal and
134
+ accept type or relationship filters. They are not exhaustive source-reference or
135
+ call graphs: selective scanning can miss arbitrary method-body constant references,
136
+ including generic PORO and library targets. No dependents or test-only dependents
137
+ do not establish absence of production callers. Check source before claiming absence.
138
+ The `graph_coverage` notice makes this scope explicit in supporting responses;
139
+ that metadata is unreleased after `2.0.0.beta3`.
133
140
 
134
141
  Start at depth 1 or 2. A deeper unfiltered traversal can obscure the direct evidence that matters. Common relationship values include associations (`belongs_to`, `has_many`, `has_one`), code references, renders, redirects, form actions, and navigation links.
135
142
 
136
143
  Both return at most 50 nodes by default and say so with a `Showing N of M (truncated)`
137
- line. Narrow with `depth`, `types` and `via` before paging with `limit` and
144
+ line for completed walks. Budget-limited answers in supporting versions instead
145
+ say `Showing N of at least M (total unknown: <budget reason>)`. Narrow with `depth`, `types` and `via` before paging with `limit` and
138
146
  `offset`: narrowing answers the question, paging only splits the same answer
139
147
  across turns. In a multi-database app each row names the unit's database.
140
148
 
149
+ On supporting servers, inspect `structuredContent.data` for traversal nodes,
150
+ `graph_coverage`, exactness, budgets and optional explanation witnesses, regardless
151
+ of the text renderer. This packaged stdio/HTTP payload is unreleased after
152
+ `2.0.0.beta3`; check the actual response and use its text when structured data is
153
+ absent. There is no traversal `format` argument. See the
154
+ [response contract](MCP_SERVERS.md#dependency-graph-coverage).
155
+
141
156
  A traversal can also stop at its independent node or edge budget. Treat
142
157
  `partial`/`partial_reason` as incomplete graph evidence even on the final page;
143
- paging cannot recover nodes the walk never reached. Check the connected schema
158
+ paging cannot recover nodes the walk never reached. Supporting responses include
159
+ `total_is_exact: false` for a cutoff and true for a finished walk, independently of
160
+ pagination. This field is unreleased after `2.0.0.beta3`; on older servers inspect
161
+ `partial` directly. `nodes_total` remains the root-inclusive admitted prefix count,
162
+ not the full reachable total when partial. Even an exact count covers only the
163
+ requested root, depth, filters and published graph generation. Check the connected schema
144
164
  before using `max_nodes`/`max_edges`, and follow the
145
165
  [budget contract](MCP_SERVERS.md#dependency-traversal-budgets).
146
166
 
@@ -149,7 +169,9 @@ recorded source-to-target relationships and a shared shortest witness to each
149
169
  row. Report `direct` relationships separately from `transitive` inferred impact.
150
170
  Follow `parent`/`edge_id` references; `context: true` ancestors are outside the
151
171
  current result page. Unknown labels and ambiguous candidate types stay unknown;
152
- `typed_path_complete: false` does not establish a uniquely typed path. See the
172
+ `typed_path_complete: false` does not establish a uniquely typed path. True means
173
+ only unambiguous witness types, not complete source coverage. Supporting text
174
+ responses label this `witness types unambiguous` (unreleased after `2.0.0.beta3`). See the
153
175
  [explanation contract](MCP_SERVERS.md#traversal-explanations).
154
176
 
155
177
  Use recorded relationship labels as evidence. Do not infer execution or call
data/docs/AGENT_SETUP.md CHANGED
@@ -40,7 +40,9 @@ If the worktree contains unrelated changes, preserve them. Do not overwrite an e
40
40
 
41
41
  Use structural-only setup when the user wants code navigation, runtime Rails structure, dependencies, flows, or blast-radius analysis. Fourteen tools register in the normal packaged launch without an embedding provider.
42
42
 
43
- Discuss semantic retrieval only if the user needs natural-language `codebase_retrieve`. The choice depends on whether they prefer local Ollama or hosted OpenAI and which vector store fits their environment. See [Backend matrix](BACKEND_MATRIX.md).
43
+ If the user wants ranked discovery through `codebase_retrieve`, offer [explicit lexical mode](RETRIEVAL_GUIDE.md#embedding-free-lexical-retrieval) over the published index without a provider or embeddings. Check that the installed version supports it, set `WOODS_RETRIEVAL_MODE=lexical` in the MCP process environment, restart that server, and verify `woods_status.retriever.mode`. Keep structural-only setup as the default unless this mode is requested.
44
+
45
+ For semantic matching, discuss local Ollama or hosted OpenAI and the appropriate vector store separately; adding a provider still requires authorization. See [Backend matrix](BACKEND_MATRIX.md).
44
46
 
45
47
  Do not infer permission to configure Console MCP from a request to “set up Woods” or “set up MCP.” The Index Server reads generated code context; the Console Server can read live data.
46
48
 
@@ -130,7 +132,7 @@ Reconnect the client and call `woods_status`. Confirm a current generation and n
130
132
 
131
133
  ### Managed Claude Code configuration
132
134
 
133
- The development command `woods-agent-config` is unreleased after 2.0.0.beta2.
135
+ `woods-agent-config` is available from Woods `2.0.0.beta3`.
134
136
  Check `bundle exec woods-agent-config --help` in the selected application bundle;
135
137
  use the manual client configuration below when it is absent. The supported
136
138
  client format is Claude Code (tested with 2.1.267).
@@ -188,8 +190,16 @@ receipt for future update/removal. Unrelated servers, hooks, settings,
188
190
  instruction text, permissions, and line-ending conventions are retained;
189
191
  changing JSON may reformat its whitespace.
190
192
 
193
+ Unreleased after `2.0.0.beta3`: apply and recovery coordinate on the actual
194
+ managed file paths, including user configuration and shared instruction files.
195
+ Two application roots sharing those files cannot apply overlapping plans at the
196
+ same time. A competing operation reports a conflict; after it finishes, create a
197
+ fresh preview if the saved plan's snapshots changed. Both applications keep
198
+ their own ownership receipts. Do not delete an active coordination lock.
199
+
191
200
  Writes use atomic replacement per file and a private recovery journal beside
192
- the receipt. The plan summary names the `.lock` and `.pending` runtime paths;
201
+ the receipt. The plan summary names all adjacent `.woods.lock` files, the
202
+ receipt `.lock`, and the `.pending` journal;
193
203
  a lock file may remain after completion. Multiple files are not one atomic
194
204
  transaction. An ordinary write failure restores original files when safe; an
195
205
  interruption or concurrent edit can retain the journal. Resolve reported
@@ -203,11 +213,11 @@ is changed; reduce the selected configuration before applying.
203
213
 
204
214
  Use a class known to exist in the application:
205
215
 
206
- 1. Call `search` to obtain its exact identifier.
207
- 2. Call `lookup` to confirm source and metadata are present.
216
+ 1. Call `search` to obtain its exact identifier and type.
217
+ 2. Call `lookup` with that identifier and type to confirm source and metadata are present.
208
218
  3. Call `dependents` with depth 1 or 2 to confirm graph edges are queryable.
209
219
 
210
- If `codebase_retrieve` reports that semantic search is disabled, that is expected for structural-only setup. Do not configure credentials merely to remove the message.
220
+ If `codebase_retrieve` reports that semantic search is disabled, that is expected for structural-only setup. Do not configure credentials merely to remove the message. If lexical retrieval was requested, verify its mode with `woods_status` and make one `codebase_retrieve` call against the published index.
211
221
 
212
222
  ## 8. Offer automatic index maintenance
213
223
 
@@ -263,7 +273,7 @@ Verified capabilities:
263
273
  - Index Server connected: yes/no
264
274
  - woods_status current: yes/no
265
275
  - search/lookup/dependents checked: yes/no
266
- - semantic retrieval: disabled/enabled (provider)
276
+ - retrieval: disabled/lexical/semantic (provider when semantic)
267
277
  - Console MCP: disabled/enabled (authorization)
268
278
  - automatic structural updates: disabled/enabled (process manager)
269
279
 
@@ -274,7 +284,7 @@ Never report a capability as enabled solely because its schema exists in source.
274
284
 
275
285
  ## Copyable prompt for an installation agent
276
286
 
277
- > Install Woods 2.x in this Rails repository using `docs/AGENT_SETUP.md`. Start with read-only preflight and preserve unrelated changes. Default to the structural Index Server; do not enable embeddings, Console MCP, HTTP transport, secrets, or purge overrides without asking me. Inspect generated files before migrating, run extraction and validation in the app's normal execution environment, configure a project-scoped MCP server in the same filesystem context as the application bundle and index, and verify `woods_status`, `search`, `lookup`, and `dependents`. Finish with the runbook's handoff report.
287
+ > Install Woods 2.x in this Rails repository using https://github.com/lost-in-the/woods/blob/main/docs/AGENT_SETUP.md. Select a published version and follow that version's tag documentation and supported capabilities. Start with read-only preflight and preserve unrelated changes. Default to the structural Index Server; do not enable embeddings, Console MCP, HTTP transport, secrets, or purge overrides without asking me. Inspect generated files before migrating, run extraction and validation in the app's normal execution environment, configure a project-scoped MCP server in the same filesystem context as the application bundle and index, and verify `woods_status`, `search`, `lookup` with the discovered identifier and type, and `dependents`. Finish with the runbook's handoff report.
278
288
 
279
289
  ## Related guides
280
290
 
@@ -96,6 +96,11 @@ CREATE INDEX IF NOT EXISTS idx_woods_vectors_embedding_hnsw
96
96
  ON woods_vectors USING hnsw (embedding vector_cosine_ops);
97
97
  ```
98
98
 
99
+ **Dimension limit:** Woods uses `vector_cosine_ops` HNSW, limited to 2,000 dimensions.
100
+ The default 3,072-dimensional `text-embedding-3-large` output needs an explicit
101
+ smaller provider output width or another backend. See the
102
+ [pgvector configuration contract](CONFIGURATION_REFERENCE.md#pgvector-postgresql).
103
+
99
104
  **Performance notes:**
100
105
  - HNSW: ~5ms search at 10K vectors, ~20ms at 100K. Memory: ~1.5x vector size.
101
106
  - For codebase indexing (~1000-5000 units, potentially 5000-20000 chunks), HNSW is appropriate.
@@ -208,6 +208,25 @@ config.vector_store_options = {
208
208
  }
209
209
  ```
210
210
 
211
+ Woods uses an HNSW index over pgvector's `vector` representation, which supports
212
+ **1–2,000 dimensions** ([pgvector's HNSW limits](https://github.com/pgvector/pgvector#hnsw)).
213
+ Unreleased after `2.0.0.beta3`: the adapter rejects wider dimensions before any
214
+ SQL, and `woods:pgvector` rejects invalid widths before writing a migration.
215
+ There is no automatic vector truncation or half-precision conversion.
216
+
217
+ The default `text-embedding-3-large` output is 3,072 dimensions. With pgvector,
218
+ explicitly request a supported provider output width, for example:
219
+
220
+ ```ruby
221
+ config.embedding_model = 'text-embedding-3-large'
222
+ config.embedding_options = { dimensions: 1536 }
223
+ ```
224
+
225
+ Keep the provider, `vector_store_options[:dimensions]` (when set), and generated
226
+ migration width equal. Changing a stored width requires a compatible new table
227
+ or an intentional index rebuild; changing the setting does not resize old data.
228
+ Use another backend if you need the full 3,072-dimensional output.
229
+
211
230
  Requires the pgvector extension. Run the generator to create migrations:
212
231
 
213
232
  ```bash
@@ -263,7 +282,12 @@ in the promoted dump; the index MCP server loads that snapshot at startup or
263
282
  reload. Incremental embedding publishes changes to paths, dependencies, and
264
283
  other unit metadata even when unchanged source needs no new embedding. A run
265
284
  with no content or metadata changes keeps the existing dump and retention
266
- window. This is a reasonable default for hosts that don't bundle `sqlite3`.
285
+ window. Unreleased after `2.0.0.beta3`: a full `Indexer#index_all` run replaces
286
+ the published corpus even when a custom caller reuses in-memory vector and
287
+ metadata stores. Deleted units, including metadata-only records, are removed;
288
+ an empty full rebuild publishes an empty dump. Failed embedding leaves the
289
+ previous promoted dump and checkpoint intact. Incremental purge guards remain
290
+ unchanged. This is a reasonable default for hosts that don't bundle `sqlite3`.
267
291
 
268
292
  ## Retrieval cache options
269
293
 
@@ -335,7 +359,7 @@ The embed run writes `woods.json` + `dumps/<ISO8601>/vectors.bin` + `metadata.ms
335
359
 
336
360
  Requirements:
337
361
  - `output_dir` must be set and readable by both the embed process and the MCP server.
338
- - The MCP server must know the same `output_dir` (pass via `woods-mcp <DIR>` or set `WOODS_DIR`).
362
+ - The MCP server must know the same `output_dir` (pass via `woods-mcp <DIR>` or set `WOODS_DIR`; see MCP path precedence below).
339
363
 
340
364
  ## Presets
341
365
 
@@ -418,11 +442,26 @@ finding qualifying edges. Run extraction again after changing either setting;
418
442
  MCP reads the published report. These findings remain informational, never a
419
443
  release or architecture gate.
420
444
 
421
- When the JSON snapshot fallback is in use, malformed JSON, top-level values
422
- other than objects, and files that cannot be read (including concurrent retention
423
- removals) are warned about and treated as absent. Snapshot lists and unit history
424
- omit them; direct lookup returns no snapshot, and a diff with an unavailable
425
- snapshot returns empty added, modified, and deleted lists.
445
+ When the JSON snapshot fallback is in use, malformed JSON, invalid snapshot
446
+ shapes, and files that cannot be read (including concurrent retention removals)
447
+ are warned about and treated as absent. A snapshot needs a hexadecimal string
448
+ `git_sha` matching its filename; `extracted_at` may be a string, null, or omitted.
449
+ When present and non-null, `units` must be an object whose records are objects.
450
+ A malformed record invalidates the entire snapshot, rather than exposing partial
451
+ history. Legacy bare identifier keys, omitted/null unit collections, and optional per-unit hash
452
+ fields remain supported; timestamp strings are not restricted to a new format.
453
+ Unit history limits count matching unit records, not the most recent snapshots
454
+ searched (JSON fallback correction unreleased after `2.0.0.beta3`). A unit
455
+ missing from newer snapshots can still have retained history.
456
+ Snapshot lists and unit history omit unusable files; direct lookup returns no
457
+ snapshot, and a diff with an unavailable snapshot returns empty added, modified,
458
+ and deleted lists. An empty diff in this case is not proof that nothing changed.
459
+ New captures compare against the latest usable snapshot. Unusable SHA-named files
460
+ still count toward retention and are pruned first when the limit is exceeded;
461
+ reading alone does not delete them. Valid legacy snapshots with null or omitted
462
+ timestamps are retained ahead of corrupt files, then treated as oldest among
463
+ usable snapshots. Direct lookup and diff still reject invalid
464
+ caller-supplied SHA paths with an argument error.
426
465
 
427
466
  `incremental_blast_radius_depth` is unbounded by default because a unit two hops
428
467
  out really can have content that depends on the changed file. An STI grandchild
@@ -473,6 +512,18 @@ config.session_store = Woods::SessionTracer::FileStore.new(
473
512
  config.session_exclude_paths = ['/health', '/metrics', '/assets']
474
513
  ```
475
514
 
515
+ ### File session retention
516
+
517
+ `FileStore` accepts `ttl:` in seconds (default `nil`, expiration disabled),
518
+ `max_sessions:` (default `1000`), and `max_requests_per_session:` (default `1000`).
519
+ TTL expires a file when the store clock reaches its modification time plus the
520
+ TTL. Recording after expiry starts a fresh history; expired events are discarded
521
+ before appending or migrating legacy filenames, under the same store lock.
522
+ When legacy and encoded files coexist, each expires independently before any
523
+ surviving histories are merged. Clearing a session is idempotent for supported
524
+ IDs, including Unicode and punctuation, and removes both filename formats when
525
+ applicable.
526
+
476
527
  ### Redis session index compatibility
477
528
 
478
529
  `RedisStore` lists and clears both legacy SET indexes and recency ZSET indexes.
@@ -652,7 +703,8 @@ These variables are read by the gem and its MCP servers at runtime. They complem
652
703
  | Variable | Default | Purpose |
653
704
  |----------|---------|---------|
654
705
  | `WOODS_RETRIEVAL_MODE` | `semantic` | Explicit packaged MCP retrieval mode: `semantic` or `lexical`. Lexical reads extraction unit JSON without provider autodetection, credentials or vector artifacts. |
655
- | `WOODS_DIR` | `Dir.pwd` | Path to the extraction output directory. |
706
+ | `WOODS_DIR` | unset | MCP extraction-index path, after a positional argument and before `WOODS_OUTPUT`. See precedence below. |
707
+ | `WOODS_OUTPUT` | unset | MCP index-path fallback when neither a positional path nor `WOODS_DIR` is set; unreleased after `2.0.0.beta3`. |
656
708
  | `WOODS_REQUIRE_INDEX` | unset | Set to `"1"` to fail closed: the server refuses to boot (raises `MissingArtifact`) unless a real index (`woods.json`) is present. By default an extract-only host boots in pattern/structural mode without it. Explicit lexical mode requires a valid published extraction index, not `woods.json`. |
657
709
  | `WOODS_ALLOW_AUTODETECT` | unset | **Deprecated no-op.** Auto-detect is now the default; accepted for backward compatibility only. |
658
710
  | `WOODS_SEARCH_MAX_SCAN` | `500` | Cap on unit files loaded during a phase-2 (metadata/source_code) `search`. Hitting the cap sets `partial: true` in the response. |
@@ -668,6 +720,20 @@ These variables are read by the gem and its MCP servers at runtime. They complem
668
720
  | `WOODS_QDRANT_URL`, `WOODS_QDRANT_COLLECTION`, `WOODS_QDRANT_API_KEY` | n/a | Override/require Qdrant connection settings when a pgvector/Qdrant-backed index is served outside its host application (no `Woods.configuration` available). |
669
721
  | `WOODS_PG_URL` | n/a | Required when a pgvector-backed index is served outside its host application. |
670
722
 
723
+ **MCP index path precedence (unreleased after `2.0.0.beta3`):** positional
724
+ argument → `WOODS_DIR` → `WOODS_OUTPUT` → current directory for `woods-mcp`
725
+ and `woods-mcp-http`. `woods-mcp-start` still requires one of the first three;
726
+ it never silently selects the current directory. An explicitly empty
727
+ `WOODS_DIR` remains an invalid override rather than falling through. Earlier
728
+ versions accept the positional path or `WOODS_DIR`; use an explicit path for
729
+ portable client configuration.
730
+
731
+ Paths are resolved in the MCP process's working directory and filesystem.
732
+ A published index can have `generation.json` pointing to a payload's
733
+ `manifest.json`; a root `manifest.json` is only the legacy flat layout. If
734
+ startup cannot find a manifest, check the reported directory and point at the
735
+ existing index before deciding another extraction is needed.
736
+
671
737
  ### Rake tasks
672
738
 
673
739
  | Variable | Default | Purpose |
@@ -198,6 +198,12 @@ the Rails server environment. The middleware stack registers automatically via
198
198
  the gem's Railtie and requires `Authorization: Bearer <token>` on every Console
199
199
  request. Missing or incorrect tokens receive `401 Unauthorized`.
200
200
 
201
+ The HTTP authentication scheme is ASCII case-insensitive (`Bearer`, `bearer`,
202
+ or `BEARER`); the token remains case-sensitive and must match exactly after one
203
+ space. This applies to both Console HTTP and `woods-mcp-http`. Case-insensitive
204
+ scheme support is unreleased after `2.0.0.beta3`; use the canonical `Bearer`
205
+ spelling in client configuration for compatibility with earlier releases.
206
+
201
207
  For non-loopback access, `console_mcp_allowed_origins` must include the public
202
208
  Rails/MCP host. If a browser-based client sends an `Origin` header from a
203
209
  different host, include that exact origin too. This allow-list controls both
@@ -760,12 +766,12 @@ Each transaction sets a statement timeout before any query runs. The default is
760
766
 
761
767
  `SqlValidator` rejects non-read-only SQL at the string level, before any database interaction.
762
768
 
763
- Validation runs **once**, inside the executor, with the dialect of the live adapter. There is deliberately no earlier dialect-blind pre-check in the tool handler: a validator built without a dialect is the conservative MySQL+PostgreSQL union, and running it first meant a MySQL host rejected statements whose `\'`/backtick grammar produces a spuriously forbidden PostgreSQL view — the adapter-aware acceptance below could never be reached on a real transport. The executor raises `SqlValidationError` for anything it refuses, which the dispatch pipeline renders as a tool error, so nothing is ungated.
769
+ Validation runs **once**, inside the executor, with the dialect of the live adapter. There is deliberately no earlier dialect-blind pre-check in the tool handler: a validator built without a dialect is the conservative union of supported dialects, and running it first meant a MySQL host rejected statements whose `\'`/backtick grammar produces a spuriously forbidden PostgreSQL view — the adapter-aware acceptance below could never be reached on a real transport. The executor raises `SqlValidationError` for anything it refuses, which the dispatch pipeline renders as a tool error, so nothing is ungated.
764
770
 
765
771
 
766
772
  - **Allowed prefixes:** `SELECT`, `WITH...SELECT`, and plain `EXPLAIN`. `EXPLAIN ANALYZE` is rejected, it executes the query rather than just planning it (both the whitespace and `EXPLAIN (ANALYZE, …)` option-list spellings).
767
773
  - **Rejected prefixes:** `INSERT`, `UPDATE`, `DELETE`, `MERGE`, `DROP`, `ALTER`, `TRUNCATE`, `CREATE`, `GRANT`, `REVOKE`
768
- - **Rejected anywhere in query:** `UNION`, `INTO`, `COPY`; row-lock clauses (`FOR UPDATE`, `FOR NO KEY UPDATE`, `FOR SHARE`, `FOR KEY SHARE`, `FOR UPDATE NOWAIT`/`SKIP LOCKED`, MySQL `LOCK IN SHARE MODE`) — these take live row locks even inside the rolled-back transaction. The lock check is adapter-aware: `console_sql` validates with the active adapter's dialect, including MySQL double-quoted strings/backtick identifiers and PostgreSQL quoted identifiers/E-strings. Unknown adapters conservatively scan both normalizations. Every view is scanned under both MySQL executable-comment (`/*!...*/`) semantics, so `#` comments and version-guarded comments cannot split a clause apart.
774
+ - **Rejected anywhere in query:** `UNION`, `INTO`, `COPY`; row-lock clauses (`FOR UPDATE`, `FOR NO KEY UPDATE`, `FOR SHARE`, `FOR KEY SHARE`, `FOR UPDATE NOWAIT`/`SKIP LOCKED`, MySQL `LOCK IN SHARE MODE`) — these take live row locks even inside the rolled-back transaction. The lock check is adapter-aware: `console_sql` validates with the active adapter's dialect, including MySQL double-quoted strings/backtick identifiers and PostgreSQL quoted identifiers/E-strings. Unknown adapters conservatively scan all supported normalizations. Every view is scanned under both MySQL executable-comment (`/*!...*/`) semantics, so `#` comments and version-guarded comments cannot split a clause apart.
769
775
  - **Function allowlist (the authoritative function control):** every function-call-shaped identifier must appear in `ALLOWED_FUNCTIONS`, a conservative set of pure read-only functions (aggregates, window functions, string/number/date/JSON readers) kept portable across MySQL, PostgreSQL, and SQLite. Anything else is rejected by name, quoted forms (`"pg_terminate_backend"(…)`) included. This is an allowlist because a denylist cannot enumerate every side-effecting function (`nextval`, `pg_advisory_lock`, `pg_terminate_backend`, …). A legacy `DANGEROUS_FUNCTIONS` denylist (`pg_sleep`, `lo_import`, `lo_export`, `pg_read_file`, `pg_write_file`, `load_file`, `sleep`, `benchmark`) still runs first as belt-and-suspenders.
770
776
  - **Rejected patterns:** multiple statements (semicolons), writable CTEs (every `AS (...)` body is checked, so a writable CTE in any WITH position is refused — `WITH a AS (SELECT 1), b AS (DELETE FROM users RETURNING *) SELECT * FROM b`), a CTE list attached to top-level DML (`WITH a AS (SELECT 1) DELETE FROM users RETURNING *`), comment-hidden injections
771
777
 
@@ -853,7 +859,44 @@ quote and comment rules. MySQL also reads the executing session's `ANSI_QUOTES`
853
859
  and `NO_BACKSLASH_ESCAPES` settings for validation, protected-column scanning, and
854
860
  table gating; adjacent subtraction operators are not assumed to begin a comment.
855
861
  Direct scanner callers without session settings use conservative quote-mode scans.
862
+ SQLite read SQL accepts simple ASCII bare, double-quoted, or backtick identifiers
863
+ (letters, digits, and underscores, starting with a letter or underscore).
864
+ Use whitespace after `FROM` and `JOIN`, and use `SELECT` subqueries rather than
865
+ parenthesized table groups. Bracket-quoted and string-quoted names, quoted names
866
+ containing punctuation, and unsupported table-reference syntax are refused before
867
+ execution, because they cannot be reliably checked against the configured table
868
+ policy. Ordinary string literals remain supported. This restriction is part of
869
+ 2.0.0.beta4; keep read tools disabled on older versions when this policy is
870
+ needed. Confirm that the release is available before selecting it. The Rails-version
871
+ integration lane checks these boundaries on real SQLite.
872
+
856
873
  The contributor live-backend lane exercises these boundaries
857
874
  through Console requests against PostgreSQL and MySQL. Keep read tools disabled
858
875
  unless live SQL access is needed, and retain the configured blocked-table and
859
876
  redaction policies when diagnosing a rejected request.
877
+
878
+ ## Console policy corrections in 2.0.0.beta4
879
+
880
+ In `2.0.0.beta4`, the default model-reading tools check the resolved relation
881
+ against `console_blocked_tables` before fetching records or counts. This includes
882
+ application-defined default scopes and the parent lookup for association counts.
883
+ The checked relation is reused for execution so a dynamic default scope is not
884
+ resolved twice. These checks do not change which Console tools are enabled.
885
+
886
+ Response handling redacts protected fields before invoking serializers, converts
887
+ the remaining response to JSON-compatible values, then redacts and scans that
888
+ normalized tree before either JSON or Markdown rendering. Symbol values and custom
889
+ JSON serializers therefore receive the same credential checks as ordinary strings.
890
+ Custom values in Markdown now use their JSON-compatible representation. Numbers,
891
+ booleans, nulls, and ordinary record shapes retain their existing meanings.
892
+
893
+ For SQLite SQL, keyword spellings receive function-policy exceptions only where
894
+ supported query grammar requires them. PostgreSQL reserved-keyword grammar is
895
+ preserved. Unsupported parenthesized offset expressions on other dialects may
896
+ require a plain numeric offset. Keep the configured access and credential policies
897
+ in place when adjusting a query.
898
+
899
+ These corrections require `2.0.0.beta4` or a reviewed development revision that
900
+ contains them. Confirm that a patched release is available before selecting it.
901
+ On affected versions, disable Console where these policies are required; Index MCP
902
+ can stay enabled because it reads the published code index separately.
data/docs/DOCKER_SETUP.md CHANGED
@@ -468,6 +468,6 @@ To register `console_sql` and `console_query`, enable
468
468
 
469
469
  ### Woods MCP tools not available in a git worktree
470
470
 
471
- When working in a git worktree, subagents may not find the woods MCP servers because `.mcp.json` discovery is path-based and the worktree has a different root directory. See [MCP_WORKTREE_SETUP.md](MCP_WORKTREE_SETUP.md) for the fix and verification steps.
471
+ A separate Claude Code session launched in a worktree may have different project-scoped MCP registrations. Subagents inherit the parent session's MCP tools, subject to tool restrictions; changing their working directory does not select a different Woods index. Check registration, the Compose service's mounted checkout, and the served index using [MCP worktree setup](MCP_WORKTREE_SETUP.md).
472
472
 
473
473
  See [CONSOLE_MCP_SETUP.md](CONSOLE_MCP_SETUP.md) for detailed console server documentation.
@@ -2,7 +2,7 @@
2
2
 
3
3
  Woods ships **35 extractor classes** producing **39 distinct unit types**: one for each meaningful category of Rails code. This doc covers what each extractor captures, how to configure them, and the shape of the data they produce.
4
4
 
5
- > **Counts explained.** `lib/woods/extractors/` contains 42 files: 35 extractor classes (each ending in `_extractor.rb`) plus 7 supporting utilities (`shared_utility_methods`, `shared_dependency_scanner`, `callback_analyzer`, `behavioral_profile`, `route_helper_resolver`, `ast_source_extraction`, `source_nesting`). The 39 unit types comes from some extractors emitting multiple categories, `GraphQLExtractor` alone produces four (`graphql_type`, `graphql_mutation`, `graphql_resolver`, `graphql_query`), and `RailsSourceExtractor` produces both `rails_source` and `gem_source`. Supporting utilities enrich existing extractors (callback side-effects, behavioral config, AST-based source slicing, nested-namespace resolution) but are not themselves extractors and do not appear in the unit type enumeration. The authoritative mapping is `Woods::Extractor::TYPE_TO_EXTRACTOR_KEY` in `lib/woods/extractor.rb`.
5
+ > **Counts explained.** `lib/woods/extractors/` contains 35 extractor classes (each ending in `_extractor.rb`) plus supporting utilities such as `shared_utility_methods`, `shared_dependency_scanner`, `callback_analyzer`, `behavioral_profile`, `route_helper_resolver`, `ast_source_extraction`, `source_nesting`, and `declared_parent`. The 39 unit types comes from some extractors emitting multiple categories, `GraphQLExtractor` alone produces four (`graphql_type`, `graphql_mutation`, `graphql_resolver`, `graphql_query`), and `RailsSourceExtractor` produces both `rails_source` and `gem_source`. Supporting utilities enrich existing extractors (callback side-effects, behavioral config, AST-based source slicing, nested-namespace resolution) but are not themselves extractors and do not appear in the unit type enumeration. The authoritative mapping is `Woods::Extractor::TYPE_TO_EXTRACTOR_KEY` in `lib/woods/extractor.rb`.
6
6
 
7
7
  ---
8
8
 
@@ -65,6 +65,7 @@ Every extractor returns `Array<ExtractedUnit>`. An `ExtractedUnit` is a self-con
65
65
  - Reads Rails' per-event callback chains (`_save_callbacks`, `_create_callbacks`, and the other lifecycle events), preserving each chain's order and each entry's `kind`, filter and conditions. The public `type` combines kind and event, such as `before_save` or `after_create`; Rails' separate `before_commit` event is reported as `before_commit`, not `before_before_commit`. `callback_count` equals the emitted callback list's length. The list includes framework-registered callbacks; it is runtime metadata, not an application-only filter or a cross-event execution trace. Regenerate the index after upgrading to pick up corrected callback metadata.
66
66
  - Proc/lambda filters, including Rails-generated association callbacks, use stable source-site labels in both metadata and callback chunks: `#<Proc app/models/post.rb:12>` (or `lambda`). App paths are relative to `Rails.root`; external paths are retained and native procs use `native`. Rails 6's numeric filter identity is resolved through `raw_filter`. These labels describe location and callable kind, not captured closure state; callbacks are never executed. Model condition labels retain their existing format.
67
67
  - Default callback-object representations omit process addresses: an instance becomes `#<CleanupCallback>`, an anonymous class becomes `#<Class>`, and its instance becomes `#<#<Class>>`. Anonymous namespace prefixes are normalized too (for example, `#<Module>::CleanupCallback`). Named classes and custom `to_s` labels retain their text. These labels do not distinguish arbitrary object state; separate registered callbacks remain separate entries even when their descriptive labels match. Controller object-filter formatting is unchanged.
68
+ - Direct Proc/lambda validation option values (for example inclusion/exclusion membership or message callables) use the same stable kind/source-site labels without execution. Validation order and duplicates, condition formats (`if`, `unless`, `on`), and non-Proc values are unchanged. Nested arrays/hashes are not recursively normalized, and labels do not serialize captured closure state. Run a full extraction after upgrading to refresh retained validation metadata; the stored schema is unchanged.
68
69
  - Callback side-effects are analyzed via `CallbackAnalyzer`: detects columns written (`self.col =`), jobs enqueued (`perform_later`), and services called
69
70
  - Reflects model class and instance methods after reading the schema, so Rails schema-loading optimizations produce the same method metadata in cold and warmed runs. Application-defined constructors remain visible; Rails versions that install an optimized singleton `new` during schema loading consistently include it in `class_methods`.
70
71
  - Automatically skips HABTM join models and anonymous classes
@@ -255,6 +256,10 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
255
256
  - Discovery and direct extraction accept only mailers backed by an existing app-owned source file, excluding dependency mailers and fabricated convention paths.
256
257
  - Each mailer action corresponds to an email template, template paths are recorded in metadata
257
258
  - Extracts `default from:`, `layout`, and per-action subject patterns
259
+ - Action names are sorted consistently in metadata, the generated header, template discovery and action chunks. Callback chain order and duplicate registrations are preserved.
260
+ - Direct Proc-valued defaults and Proc callback filters use source-location/kind labels without executing them; application paths are relative to `Rails.root`. Default containers, literal strings and non-Proc values retain their existing types. This does not serialize closure captures or recursively normalize arbitrary nested objects.
261
+ - Object callback filters use descriptive labels with addresses removed only from Ruby's default representation, including anonymous classes and namespaces. Custom labels and literal hexadecimal text are preserved. Labels do not serialize callback object state, and extraction never invokes callbacks.
262
+ - After upgrading, run a full extraction to refresh retained mailer units. Stabilized headers and callable labels can cause a one-time source-hash change; stored index schemas are unchanged.
258
263
 
259
264
  ---
260
265
 
@@ -382,6 +387,7 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
382
387
  - Scans `app/models` for files that don't define an `ActiveRecord::Base` descendant
383
388
  - Common examples: value objects, form objects placed in `app/models`, domain structs
384
389
  - Excludes concerns (those go to ConcernExtractor)
390
+ - `parent_class` and the generated Parent annotation describe the selected unit declaration only. Nested or sibling classes cannot supply its parent. An implicit `Object` parent, a dynamic superclass expression, or unparseable source produces `nil`; explicit constant-path parents retain their source names.
385
391
 
386
392
  ---
387
393
 
@@ -421,6 +427,7 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
421
427
  - Produces unit types: `graphql_type`, `graphql_mutation`, `graphql_resolver`, `graphql_query`
422
428
  - Extracts field metadata (types, descriptions, complexity, arguments), authorization patterns (Pundit, CanCan, `authorized?`), and dependencies on models/services
423
429
  - Since all GraphQL units come from one extractor, incremental re-extraction handles them via `extract_graphql_file`
430
+ - `parent_class` and summary chunks describe the selected declaration's explicit constant-path superclass, preserving its written qualification. Nested or sibling declarations and literal text cannot supply a parent. Implicit Object, module interfaces, dynamic superclass expressions, unavailable source, and invalid source have no declared parent (`null` metadata; `unknown` in summaries). This is source declaration metadata, not resolved runtime ancestry.
424
431
 
425
432
  **Example output (abbreviated):**
426
433
 
@@ -677,6 +684,7 @@ Every app-owned unit under a package root carries `metadata[:package]` with the
677
684
  **Key details:**
678
685
  - Excludes `lib/tasks/` (covered by RakeTaskExtractor) and `lib/generators/`
679
686
  - File-based scanning; no assumption about class hierarchy
687
+ - `parent_class` and the generated Parent annotation describe the selected unit declaration only. Nested or sibling classes cannot supply its parent. An implicit `Object` parent, a dynamic superclass expression, or unparseable source produces `nil`; explicit constant-path parents retain their source names.
680
688
 
681
689
  ---
682
690
 
@@ -98,6 +98,23 @@ including inlined code and callback analysis. Multiple runtime mixins sharing a
98
98
  file retain separate identities and refresh all their includers. Run a full extraction after upgrading
99
99
  to populate these previously missing source mappings.
100
100
 
101
+ ## Handled source errors and retry
102
+
103
+ Unreleased after `2.0.0.beta3`: when an incremental extraction or named refresh
104
+ records a handled consumer error (for example malformed locale or schedule
105
+ YAML), it raises `Woods::ExtractionError` before publishing. Empty output from
106
+ that failed consumer does not authorize replacing or deleting its last-good
107
+ units. The published generation and source provenance remain unchanged,
108
+ including when other files in the batch extracted successfully.
109
+
110
+ Fix the source error named in the extraction log, then retry the **complete
111
+ batch**, or the same named refresh. The watch daemon reports degraded and keeps
112
+ the failed batch pending for retry. An error on one file does not mark a later
113
+ successful file as failed, but the batch still cannot publish until all handled
114
+ errors are resolved. This does not change full extraction's existing tolerance
115
+ for handled consumer errors; its source-freshness report marks those scopes
116
+ unverified.
117
+
101
118
  ## What a run does, in order
102
119
 
103
120
  `Extractor#extract_changed` is order-sensitive; each step exists because of the
@@ -109,8 +126,8 @@ step before it.
109
126
  2. **Reconcile changed paths.** Every changed path that still exists is handed
110
127
  to the file-based extractors that claim it (`PathDispatcher`), and units the
111
128
  path no longer produces are dropped. This is what indexes a file the index
112
- has never seen, and what lets a task removed from a multi-task `.rake` file
113
- actually go away.
129
+ has never seen, and what removes definitions deleted from a surviving
130
+ source file. Multi-file Rake tasks use wholesale reconciliation below.
114
131
  3. **Re-extract the rest of the blast radius**: units whose own file did not
115
132
  change but which depend on something that did.
116
133
  4. **Reconcile class-based types** against each extractor's
@@ -208,7 +225,6 @@ automatically.
208
225
  | `config/locales/**/*.yml` | i18n |
209
226
  | `config/initializers`, `config/environments` | configurations |
210
227
  | `db/migrate/*.rb` (top level only) | migrations |
211
- | `lib/tasks/**/*.rake` | rake_tasks |
212
228
  | `lib/**/*.rb` (outside `tasks/`, `generators/`) | libs |
213
229
  | `spec/**/*_spec.rb`, `test/**/*_test.rb` | test_mappings |
214
230
 
@@ -219,12 +235,13 @@ serializer and decorator extractors, and all matching rules run.
219
235
  ### Wholesale re-runs
220
236
 
221
237
  `PathDispatcher.whole_app_rules` → `Extractor::WHOLE_APP_EXTRACTORS`. These
222
- extractors have no per-file entry point: they introspect the runtime or scan a
223
- whole directory in one pass. In an already-booted process re-running them is
238
+ extractors need a complete runtime or directory view, even when a low-level
239
+ per-file reader exists. In an already-booted process re-running them is
224
240
  cheap, which is what makes wholesale replacement the right shape.
225
241
 
226
242
  | Trigger | Re-runs |
227
243
  |---|---|
244
+ | `lib/tasks/**/*.rake` | rake_tasks (all definitions of every task) |
228
245
  | `config/routes.rb`, `config/routes/**` | routes, engines, **and** controllers, mailers, components, view components, view templates |
229
246
  | `Gemfile.lock` | engines, middleware, rails_source (gated by `include_framework_sources`) |
230
247
  | `config/application.rb`, `config/initializers/**`, `config/environments/**` | middleware |
@@ -235,7 +252,14 @@ cheap, which is what makes wholesale replacement the right shape.
235
252
  | `db/views/**/*.sql` | database_views |
236
253
  | any `package.yml`, `packwerk.yml` | packages |
237
254
 
238
- Three of these deserve a note:
255
+ Four of these deserve a note:
256
+
257
+ - **Rake tasks merge definitions across files.** Any changed or deleted `.rake`
258
+ file reruns the task extractor over all task files. Removing the primary
259
+ definition preserves surviving definitions; removing a secondary definition
260
+ drops its source and dependencies from the shared unit. After upgrading from
261
+ the old per-file rules, run a full extraction to establish a source-freshness
262
+ baseline with the new rule fingerprint.
239
263
 
240
264
  - **Routes cascade.** `ROUTE_CONSUMER_EXTRACTORS` embed the route table, controllers write each action's routes into unit metadata and into the action
241
265
  chunks, and everything using `RouteHelperResolver` resolves navigation edges
data/docs/MCP_SERVERS.md CHANGED
@@ -180,7 +180,7 @@ failure boundary; they do not prove uniqueness or become `ambiguous_identity` er
180
180
  Use `depth: 0` for the request timeline, or inspect candidates with typed `lookup`
181
181
  calls. Re-extraction does not remove a legitimate cross-type collision. Successful
182
182
  traces retain their existing identifiers and response shape; target identity has
183
- not been migrated globally. These corrections are unreleased after `2.0.0.beta2`.
183
+ not been migrated globally. These corrections are available in `2.0.0.beta3`.
184
184
 
185
185
  The server also exposes MCP resources and resource templates for indexed units. Tool descriptions returned by MCP are the parameter-level source of truth; [Agent guide](AGENT_GUIDE.md) explains selection strategy.
186
186
 
@@ -192,6 +192,23 @@ error and continues serving the previous aligned generation; it never swaps in a
192
192
  partial or empty replacement. Grant write access for live reloads, or restart the MCP
193
193
  process after publishing a new embedded index.
194
194
 
195
+ ### Graph-analysis pages
196
+
197
+ Unreleased after `2.0.0.beta3`: `graph_analysis` enforces its advertised default
198
+ of 20 rows per section. Pass `limit` and `offset` to page one selected `analysis`
199
+ or each section of `analysis: "all"`. Explicit limits also bound nested hub
200
+ `dependents` lists. Older servers may return every section row when `limit` is
201
+ omitted; pass a limit explicitly when supporting both versions.
202
+
203
+ JSON responses retain `<section>_total`, `<section>_offset` (when positive), and
204
+ `<section>_truncated: true` whenever a page omits rows before or after it. Markdown,
205
+ plain, and Claude responses show the same total and offset on last and empty
206
+ pages. For example, offset 20 with limit 5 over 25 published orphans shows
207
+ `5 of 25 from offset 20`; offset 100 shows `0 of 25 from offset 100`. An empty
208
+ page does not mean the section has no findings. These totals describe the
209
+ published report arrays, which can themselves be bounded during extraction;
210
+ they do not establish complete source-reference coverage.
211
+
195
212
  ### Search completeness
196
213
 
197
214
  Search responses retain `query`, `result_count`, and `results`; `result_count`
@@ -229,6 +246,31 @@ Detected missing, unreadable, or corrupt artifacts remain `isError: true` with
229
246
  `has_more`, `total_matches`, and `matched_lower_bound`; no successful empty
230
247
  result is substituted. Inspect `woods_status` and run `woods:validate`.
231
248
 
249
+ ### Dependency graph coverage
250
+
251
+ `dependencies` and `dependents` return relationships recorded in the published
252
+ index, not an exhaustive call graph or source-reference index. Extraction combines
253
+ runtime reflection with selective source scanning; arbitrary method-body constant
254
+ references (including references to generic PORO and library classes) may have no
255
+ edge. No dependents, a test-only dependent, or a completed traversal does not prove
256
+ there are no production callers. Verify important absence claims in source.
257
+
258
+ Supporting servers expose the annotated, paginated traversal result in
259
+ `structuredContent.data` for every renderer, including the default packaged
260
+ stdio and HTTP servers. Read `data.total_is_exact`, `data.graph_coverage`, budget
261
+ counters and optional explanation witnesses there; `content[0].text` and
262
+ `structuredContent.text` keep the same human-readable rendering. No `format`
263
+ tool argument is needed or accepted. This additive data payload is unreleased
264
+ after `2.0.0.beta3`; verify the installed response before relying on it. Older
265
+ human-renderer responses can carry only text. The structured nodes and witnesses
266
+ cover the same page, not an additional traversal or an unpaginated graph.
267
+
268
+ Successful responses carry `graph_coverage` with `scope: "published_relationships"`,
269
+ `source_references: "not_exhaustive"`, and a human-readable `notice`. Text formats
270
+ show the same notice, including compact, root-only and empty-page responses.
271
+ This response metadata and the total exactness field below are unreleased after
272
+ Woods `2.0.0.beta3`; older servers need the same conservative interpretation.
273
+
232
274
  ### Dependency traversal budgets
233
275
 
234
276
  `dependencies` and `dependents` walk breadth-first in stored graph order. The
@@ -251,7 +293,16 @@ Exact-budget walks that finish all requested work are complete and have no
251
293
  `limit` (default 50) and `offset` only page that discovered result; they never
252
294
  change the walk budget or depth. On a partial traversal, `nodes_total`, when
253
295
  present for pagination, counts the discovered prefix, **not the full reachable
254
- graph**. Paging beyond that prefix stays partial. To explore more, narrow
296
+ graph**. Every successful response includes `total_is_exact`: false for a
297
+ budget cutoff, true when the requested walk finishes, even when its page is
298
+ truncated or empty. It is independent of `limit`/`offset` and is present for
299
+ unpaged answers too. Partial text answers say `Showing N of at least M (total
300
+ unknown: node_budget)` (or `edge_budget`), including when no pagination is needed.
301
+ `M` includes the root and counts the admitted prefix; it is a lower bound for the
302
+ requested root, depth, type/relationship filters and published generation, not a
303
+ count of all application callers. Exactness describes that same recorded-graph
304
+ scope and never implies exhaustive source coverage. Paging beyond that prefix
305
+ stays partial. To explore more, narrow
255
306
  `depth`/`types`/`via`, choose another root, or increase the traversal budget within
256
307
  its maximum. Keep the root, filters, budgets, and published generation unchanged
257
308
  for stable pages. No wall-clock deadline is used, so cutoffs are deterministic.
@@ -289,6 +340,10 @@ A target name shared by several types has `type: null`,
289
340
  record target types, so the response cannot choose among candidates. A witness
290
341
  through an ambiguous or unresolved identity sets `typed_path_complete: false`;
291
342
  it describes identifier-level reachability, never a uniquely typed path.
343
+ A true value means only that identities along this witness have unambiguous
344
+ types. It does not establish source-reference coverage or observed execution.
345
+ Text labels this `witness types unambiguous=yes/no`; the JSON key and its meaning
346
+ remain unchanged. The text label change is unreleased after `2.0.0.beta3`.
292
347
  `types` filters retain the compact traversal's identifier-level semantics: any
293
348
  registered type can qualify a name, while edge evidence keeps its actual source
294
349
  owner. Multiple relationship kinds between the same endpoints remain separate.
@@ -64,7 +64,7 @@ The Index Server defines **29 schemas**: the packaged executable registers **14*
64
64
  | Snapshot (4) | 4 | Extraction with `enable_snapshots = true` normally creates `woods.sqlite3`, which packaged servers discover. If extraction used the JSON fallback, set `WOODS_SNAPSHOTS=true` on the standalone server. Custom embedded servers pass `snapshot_store:`. Internal SQLite migrations are automatic. Tools: `list_snapshots`, `snapshot_diff`, `unit_history`, `snapshot_detail` |
65
65
  | `notion_sync` | 1 | `notion_api_token` + `notion_database_ids` both set |
66
66
 
67
- `codebase_retrieve` is always registered (no `retrieve` alias exists), but only returns results once an embedding provider is configured and `rake woods:embed` has run.
67
+ `codebase_retrieve` is always registered (no `retrieve` alias exists). Default semantic mode requires an embedding provider and a completed `woods:embed` run. Explicit `WOODS_RETRIEVAL_MODE=lexical` ranks published extraction units without a provider or embeddings; set it in the MCP process environment and restart the server. See [embedding-free lexical retrieval](RETRIEVAL_GUIDE.md#embedding-free-lexical-retrieval).
68
68
 
69
69
  If an agent reports a missing tool, compare its request with the connected server's registered list and [MCP server boundaries](MCP_SERVERS.md#conditional-index-capabilities). The normal packaged executable does not wire operator or feedback collaborators. **Console Server tools are not all unconditionally registered**: 31 tool schemas exist as an inventory, but only the 9 Tier 1 tools are executable by default, or 11 with `console_embedded_read_tools: true` (adds `console_sql`/`console_query`). Tier 2, Tier 3, and `console_eval` are schema-only in every supported mode; there is no bridge or confirmation flow that unlocks them. See [MCP servers](MCP_SERVERS.md#console-server) for the supported inventory.
70
70
 
@@ -551,7 +551,7 @@ Static tools miss all of these because they only exist after Rails processes the
551
551
  }
552
552
  ```
553
553
 
554
- **What you'll get:** Units with no dependents, nothing in the codebase references them. Good candidates for removal or investigation.
554
+ **What you'll get:** Units with no recorded dependents in the published graph, excluding types treated as natural entry points. These are candidates for investigation, not proof of dead code. Method-body references and dynamic callers may be missing; verify source references, framework entry points, and runtime usage before removing anything. See [dependency graph coverage](MCP_SERVERS.md#dependency-graph-coverage).
555
555
 
556
556
  ---
557
557
 
@@ -811,7 +811,7 @@ Keys without a recognised suffix fall through to ActiveRecord `where(hash)` equa
811
811
 
812
812
  ### "Find code related to subscription billing"
813
813
 
814
- **Tool:** `codebase_retrieve` (Index Server, requires embedding provider)
814
+ **Tool:** `codebase_retrieve` (Index Server, semantic or explicit lexical mode)
815
815
 
816
816
  ```json
817
817
  {
@@ -820,7 +820,7 @@ Keys without a recognised suffix fall through to ActiveRecord `where(hash)` equa
820
820
  }
821
821
  ```
822
822
 
823
- **What you'll get:** A token-budgeted context string of the most semantically relevant units, ranked by hybrid search (semantic + keyword + PageRank). Requires an embedding provider (`embedding_provider: :openai` or `:ollama`) to be configured.
823
+ **What you'll get:** Ranked context within an estimated text-token budget. Default semantic mode uses configured embeddings and hybrid ranking; explicit lexical mode uses field-aware BM25 over published units, without a provider or `woods:embed`. The same query works in either configured mode, though rankings differ. Confirm the active mode with `woods_status.retriever.mode`; see the [retrieval guide](RETRIEVAL_GUIDE.md#embedding-free-lexical-retrieval) for setup and budget limits.
824
824
 
825
825
  ---
826
826