woods 2.0.0.beta3 → 2.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (120) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +500 -420
  3. data/CONTRIBUTING.md +29 -17
  4. data/README.md +78 -178
  5. data/docs/AGENT_GUIDE.md +52 -11
  6. data/docs/AGENT_SETUP.md +34 -17
  7. data/docs/AUTOMATIC_MAINTENANCE.md +222 -0
  8. data/docs/BACKEND_MATRIX.md +18 -7
  9. data/docs/CLIENT_HOOKS.md +1 -1
  10. data/docs/CONFIGURATION_REFERENCE.md +105 -29
  11. data/docs/CONSOLE_MCP_SETUP.md +54 -9
  12. data/docs/DOCKER_SETUP.md +16 -1
  13. data/docs/EVALUATION.md +10 -4
  14. data/docs/EXTRACTOR_REFERENCE.md +23 -3
  15. data/docs/FAQ.md +14 -3
  16. data/docs/GETTING_STARTED.md +18 -17
  17. data/docs/INCREMENTAL_EXTRACTION.md +37 -8
  18. data/docs/INDEX_LAYOUT.md +2 -2
  19. data/docs/MCP_SERVERS.md +79 -7
  20. data/docs/MCP_TOOL_COOKBOOK.md +5 -5
  21. data/docs/MCP_WORKTREE_SETUP.md +55 -83
  22. data/docs/PUBLISHED_INDEX.md +17 -0
  23. data/docs/README.md +2 -1
  24. data/docs/RETRIEVAL_GUIDE.md +81 -13
  25. data/docs/SOURCE_FRESHNESS.md +1 -1
  26. data/docs/TOKEN_BENCHMARK.md +16 -10
  27. data/docs/TROUBLESHOOTING.md +142 -47
  28. data/docs/UPGRADING_TO_2.md +12 -6
  29. data/docs/WATCH_DAEMON.md +189 -24
  30. data/docs/WHY_WOODS.md +9 -5
  31. data/exe/woods-console +13 -11
  32. data/exe/woods-mcp-start +14 -9
  33. data/exe/woods-watch +5 -0
  34. data/lib/generators/woods/pgvector_generator.rb +8 -2
  35. data/lib/generators/woods/watch_generator.rb +53 -0
  36. data/lib/puma/plugin/woods.rb +10 -0
  37. data/lib/tasks/woods.rake +14 -0
  38. data/lib/woods/agent_configuration/applier.rb +5 -3
  39. data/lib/woods/agent_configuration/cli.rb +2 -2
  40. data/lib/woods/agent_configuration/layout.rb +13 -0
  41. data/lib/woods/cache/cache_middleware.rb +6 -0
  42. data/lib/woods/console/credential_scanner.rb +4 -3
  43. data/lib/woods/console/dispatch_pipeline.rb +7 -0
  44. data/lib/woods/console/embedded_executor.rb +31 -9
  45. data/lib/woods/console/sql_noise_stripper.rb +9 -7
  46. data/lib/woods/console/sql_table_scanner.rb +47 -7
  47. data/lib/woods/console/sql_validator.rb +49 -9
  48. data/lib/woods/console/sqlite_read_guard.rb +46 -0
  49. data/lib/woods/console/stdio_transport.rb +27 -0
  50. data/lib/woods/coordination/pipeline_lock.rb +3 -2
  51. data/lib/woods/embedding/indexer.rb +24 -14
  52. data/lib/woods/extractor.rb +70 -19
  53. data/lib/woods/extractors/declared_parent.rb +55 -0
  54. data/lib/woods/extractors/graphql_extractor.rb +2 -11
  55. data/lib/woods/extractors/lib_extractor.rb +10 -8
  56. data/lib/woods/extractors/mailer_extractor.rb +6 -10
  57. data/lib/woods/extractors/model_extractor.rb +1 -15
  58. data/lib/woods/extractors/poro_extractor.rb +10 -8
  59. data/lib/woods/extractors/shared_utility_methods.rb +22 -5
  60. data/lib/woods/git_command.rb +6 -7
  61. data/lib/woods/git_provenance.rb +4 -6
  62. data/lib/woods/mcp/bearer_auth.rb +2 -1
  63. data/lib/woods/mcp/bootstrapper.rb +20 -5
  64. data/lib/woods/mcp/config_resolver.rb +2 -1
  65. data/lib/woods/mcp/index_reader.rb +11 -2
  66. data/lib/woods/mcp/initialization_guidance.rb +1 -1
  67. data/lib/woods/mcp/renderers/markdown_renderer.rb +14 -8
  68. data/lib/woods/mcp/renderers/plain_renderer.rb +11 -7
  69. data/lib/woods/mcp/server.rb +63 -37
  70. data/lib/woods/mcp/tool_contract.rb +1 -1
  71. data/lib/woods/mcp/tool_response_renderer.rb +16 -0
  72. data/lib/woods/mcp/traversal_evidence_text.rb +1 -1
  73. data/lib/woods/mcp/traversal_response.rb +22 -0
  74. data/lib/woods/path_dispatcher.rb +6 -5
  75. data/lib/woods/published_index/typed_unit_reader.rb +40 -3
  76. data/lib/woods/published_index.rb +2 -2
  77. data/lib/woods/rake_helpers.rb +2 -12
  78. data/lib/woods/retrieval/corpus_status.rb +46 -0
  79. data/lib/woods/retrieval/lexical_assembler.rb +14 -3
  80. data/lib/woods/retrieval/lexical_index.rb +2 -1
  81. data/lib/woods/retriever.rb +19 -7
  82. data/lib/woods/session_tracer/file_store.rb +6 -1
  83. data/lib/woods/source_inputs/consumer_errors.rb +4 -0
  84. data/lib/woods/storage/local_corpus_stats.rb +32 -0
  85. data/lib/woods/storage/metadata_store.rb +20 -0
  86. data/lib/woods/storage/pgvector.rb +6 -2
  87. data/lib/woods/storage/vector_store.rb +10 -0
  88. data/lib/woods/temporal/json_snapshot_store.rb +35 -7
  89. data/lib/woods/version.rb +1 -1
  90. data/lib/woods/watch/child_environment.rb +30 -0
  91. data/lib/woods/watch/cli.rb +91 -0
  92. data/lib/woods/watch/daemon.rb +73 -11
  93. data/lib/woods/watch/event_stream.rb +70 -0
  94. data/lib/woods/watch/guardian.rb +142 -0
  95. data/lib/woods/watch/installation/layout.rb +70 -0
  96. data/lib/woods/watch/installation/options.rb +128 -0
  97. data/lib/woods/watch/installation/planner.rb +128 -0
  98. data/lib/woods/watch/installation/probe.rb +101 -0
  99. data/lib/woods/watch/installation/receipt.rb +77 -0
  100. data/lib/woods/watch/installation/recovery.rb +64 -0
  101. data/lib/woods/watch/installation/templates.rb +58 -0
  102. data/lib/woods/watch/installation.rb +56 -0
  103. data/lib/woods/watch/lifecycle.rb +182 -0
  104. data/lib/woods/watch/managed_child.rb +113 -0
  105. data/lib/woods/watch/managed_cleanup.rb +48 -0
  106. data/lib/woods/watch/managed_process.rb +144 -0
  107. data/lib/woods/watch/puma_adapter.rb +87 -0
  108. data/lib/woods/watch/puma_child.rb +66 -0
  109. data/lib/woods/watch/supervision_records.rb +95 -0
  110. data/lib/woods/watch/supervision_status.rb +104 -0
  111. data/lib/woods/watch/supervisor.rb +161 -0
  112. data/lib/woods/watch/supervisor_reporting.rb +46 -0
  113. data/plugin/.claude-plugin/plugin.json +1 -1
  114. data/plugin/hooks/woods-input-rules.sh +4 -4
  115. data/plugin/skills/woods-agent-enable/SKILL.md +7 -1
  116. data/plugin/skills/woods-diagnose/SKILL.md +134 -34
  117. data/plugin/skills/woods-investigate/SKILL.md +54 -15
  118. data/plugin/skills/woods-mcp-config/SKILL.md +38 -11
  119. data/plugin/skills/woods-setup/SKILL.md +72 -15
  120. metadata +38 -5
@@ -89,8 +89,13 @@ Recovery choices, in the order they are worth trying:
89
89
  3. Run a full `woods:extract` when the range cannot be repaired this run.
90
90
 
91
91
  The diff itself is rooted at the extracted application (`git -C Rails.root`),
92
- so it cannot read whatever checkout the process happened to start in — the
93
- same rooting rule the manifest's git provenance follows.
92
+ independently of the process working directory. An explicit `WOODS_GIT_DIR`
93
+ selects that Git directory's HEAD for both the diff and manifest provenance.
94
+ For a linked worktree, use its worktree-specific directory within the complete
95
+ shared layout; selecting the shared root instead reads the primary checkout's
96
+ HEAD. See [worktree mount verification](TROUBLESHOOTING.md#git-directory-mounts-for-linked-worktrees).
97
+ Git-only changes may not trigger the source-file watcher; see the
98
+ [watcher limitation](WATCH_DAEMON.md#watcher-backends).
94
99
 
95
100
  Named, source-defined app modules included by runtime models are tracked as concern units even
96
101
  outside `concerns/` directories. Changing their source refreshes their includers,
@@ -98,6 +103,23 @@ including inlined code and callback analysis. Multiple runtime mixins sharing a
98
103
  file retain separate identities and refresh all their includers. Run a full extraction after upgrading
99
104
  to populate these previously missing source mappings.
100
105
 
106
+ ## Handled source errors and retry
107
+
108
+ Included in Woods `2.0.0`: when an incremental extraction or named refresh
109
+ records a handled consumer error (for example malformed locale or schedule
110
+ YAML), it raises `Woods::ExtractionError` before publishing. Empty output from
111
+ that failed consumer does not authorize replacing or deleting its last-good
112
+ units. The published generation and source provenance remain unchanged,
113
+ including when other files in the batch extracted successfully.
114
+
115
+ Fix the source error named in the extraction log, then retry the **complete
116
+ batch**, or the same named refresh. The watch daemon reports degraded and keeps
117
+ the failed batch pending for retry. An error on one file does not mark a later
118
+ successful file as failed, but the batch still cannot publish until all handled
119
+ errors are resolved. This does not change full extraction's existing tolerance
120
+ for handled consumer errors; its source-freshness report marks those scopes
121
+ unverified.
122
+
101
123
  ## What a run does, in order
102
124
 
103
125
  `Extractor#extract_changed` is order-sensitive; each step exists because of the
@@ -109,8 +131,8 @@ step before it.
109
131
  2. **Reconcile changed paths.** Every changed path that still exists is handed
110
132
  to the file-based extractors that claim it (`PathDispatcher`), and units the
111
133
  path no longer produces are dropped. This is what indexes a file the index
112
- has never seen, and what lets a task removed from a multi-task `.rake` file
113
- actually go away.
134
+ has never seen, and what removes definitions deleted from a surviving
135
+ source file. Multi-file Rake tasks use wholesale reconciliation below.
114
136
  3. **Re-extract the rest of the blast radius**: units whose own file did not
115
137
  change but which depend on something that did.
116
138
  4. **Reconcile class-based types** against each extractor's
@@ -208,7 +230,6 @@ automatically.
208
230
  | `config/locales/**/*.yml` | i18n |
209
231
  | `config/initializers`, `config/environments` | configurations |
210
232
  | `db/migrate/*.rb` (top level only) | migrations |
211
- | `lib/tasks/**/*.rake` | rake_tasks |
212
233
  | `lib/**/*.rb` (outside `tasks/`, `generators/`) | libs |
213
234
  | `spec/**/*_spec.rb`, `test/**/*_test.rb` | test_mappings |
214
235
 
@@ -219,12 +240,13 @@ serializer and decorator extractors, and all matching rules run.
219
240
  ### Wholesale re-runs
220
241
 
221
242
  `PathDispatcher.whole_app_rules` → `Extractor::WHOLE_APP_EXTRACTORS`. These
222
- extractors have no per-file entry point: they introspect the runtime or scan a
223
- whole directory in one pass. In an already-booted process re-running them is
243
+ extractors need a complete runtime or directory view, even when a low-level
244
+ per-file reader exists. In an already-booted process re-running them is
224
245
  cheap, which is what makes wholesale replacement the right shape.
225
246
 
226
247
  | Trigger | Re-runs |
227
248
  |---|---|
249
+ | `lib/tasks/**/*.rake` | rake_tasks (all definitions of every task) |
228
250
  | `config/routes.rb`, `config/routes/**` | routes, engines, **and** controllers, mailers, components, view components, view templates |
229
251
  | `Gemfile.lock` | engines, middleware, rails_source (gated by `include_framework_sources`) |
230
252
  | `config/application.rb`, `config/initializers/**`, `config/environments/**` | middleware |
@@ -235,7 +257,14 @@ cheap, which is what makes wholesale replacement the right shape.
235
257
  | `db/views/**/*.sql` | database_views |
236
258
  | any `package.yml`, `packwerk.yml` | packages |
237
259
 
238
- Three of these deserve a note:
260
+ Four of these deserve a note:
261
+
262
+ - **Rake tasks merge definitions across files.** Any changed or deleted `.rake`
263
+ file reruns the task extractor over all task files. Removing the primary
264
+ definition preserves surviving definitions; removing a secondary definition
265
+ drops its source and dependencies from the shared unit. After upgrading from
266
+ the old per-file rules, run a full extraction to establish a source-freshness
267
+ baseline with the new rule fingerprint.
239
268
 
240
269
  - **Routes cascade.** `ROUTE_CONSUMER_EXTRACTORS` embed the route table, controllers write each action's routes into unit metadata and into the action
241
270
  chunks, and everything using `RouteHelperResolver` resolves navigation edges
data/docs/INDEX_LAYOUT.md CHANGED
@@ -94,8 +94,8 @@ families and `manifest.provenance.mode`; it is not Rails runtime evidence.
94
94
  ### Semantic graph validation
95
95
 
96
96
  `woods:validate` checks raw graph data against the unit indexes and artifacts in
97
- one pinned published generation. This validation is unreleased after
98
- `2.0.0.beta2`. It checks the shapes of `nodes`, `edges`, `reverse`, `file_map`,
97
+ one pinned published generation. This validation is included in Woods `2.0.0`.
98
+ It checks the shapes of `nodes`, `edges`, `reverse`, `file_map`,
99
99
  `type_index`, and optional `variants`; typed identities must be unique and agree
100
100
  with the actual indexed units. Forward sources must exist. Reverse, file, and
101
101
  type memberships must match the union of primary and variant contributions.
data/docs/MCP_SERVERS.md CHANGED
@@ -56,6 +56,12 @@ Prefer the application's bundle and a project-scoped configuration:
56
56
 
57
57
  `woods-mcp-start` checks that the directory and published manifest exist, then replaces itself with `woods-mcp`. It does not install dependencies or restart a crashed process.
58
58
 
59
+ The MCP client owns this reader process. Automatic index maintenance has its own
60
+ [watcher startup](WATCH_DAEMON.md#managed-development-startup): configuring MCP
61
+ does not enable it. A running reader sees new published generations on subsequent
62
+ calls without reconnecting. Verify a real edit, not only a successful connection;
63
+ see [automatic maintenance](AUTOMATIC_MAINTENANCE.md).
64
+
59
65
  `woods_status.index.woods_version` identifies the last publisher of the served
60
66
  manifest; `server.version` identifies the running MCP reader. Missing writer
61
67
  provenance is `null`. See [manifest writer provenance](PUBLISHED_INDEX.md#manifest-writer-provenance).
@@ -133,12 +139,23 @@ before using `codebase_retrieve`, including after reload. The guidance grants
133
139
  no extraction, configuration-change, or Console authorization. Detailed usage
134
140
  belongs in the [agent guide](AGENT_GUIDE.md).
135
141
 
136
- This addition is unreleased after `2.0.0.beta2`. The SDK omits `instructions`
142
+ This addition is included in Woods `2.0.0`. The SDK omits `instructions`
137
143
  when negotiating protocol `2024-11-05`; that behavior is preserved. Older gems,
138
144
  legacy clients, and clients that do not show server instructions can use the
139
145
  agent guide or investigation skill. Leave protocol negotiation enabled rather
140
146
  than pinning a newer version solely to obtain guidance.
141
147
 
148
+ ### Structural and semantic readiness
149
+
150
+ `ready` describes the published structural index. A reachable embedding provider
151
+ and bootstrap state `hydrated` do not establish that semantic stores contain
152
+ data. Supporting readers also report `retriever.corpus`: locally known vector
153
+ and metadata entry counts, counts by type, and whether those stores are empty,
154
+ populated, or unknown. These are stored records (including chunks), not a count
155
+ of extracted units with verified embedding coverage. Positive counts do not
156
+ prove complete coverage or working provider access. See
157
+ [semantic corpus diagnostics](RETRIEVAL_GUIDE.md#semantic-corpus-diagnostics).
158
+
142
159
  ### Tools (29 — 14 registered in the packaged default)
143
160
 
144
161
  The Index Server defines 29 schemas across core and conditional capabilities. The normal packaged executable registers the 14 tools below; the remaining schemas require the specialized wiring described afterward.
@@ -180,7 +197,7 @@ failure boundary; they do not prove uniqueness or become `ambiguous_identity` er
180
197
  Use `depth: 0` for the request timeline, or inspect candidates with typed `lookup`
181
198
  calls. Re-extraction does not remove a legitimate cross-type collision. Successful
182
199
  traces retain their existing identifiers and response shape; target identity has
183
- not been migrated globally. These corrections are unreleased after `2.0.0.beta2`.
200
+ not been migrated globally. These corrections are available in `2.0.0.beta3`.
184
201
 
185
202
  The server also exposes MCP resources and resource templates for indexed units. Tool descriptions returned by MCP are the parameter-level source of truth; [Agent guide](AGENT_GUIDE.md) explains selection strategy.
186
203
 
@@ -192,12 +209,29 @@ error and continues serving the previous aligned generation; it never swaps in a
192
209
  partial or empty replacement. Grant write access for live reloads, or restart the MCP
193
210
  process after publishing a new embedded index.
194
211
 
212
+ ### Graph-analysis pages
213
+
214
+ Included in Woods `2.0.0`: `graph_analysis` enforces its advertised default
215
+ of 20 rows per section. Pass `limit` and `offset` to page one selected `analysis`
216
+ or each section of `analysis: "all"`. Explicit limits also bound nested hub
217
+ `dependents` lists. Older servers may return every section row when `limit` is
218
+ omitted; pass a limit explicitly when supporting both versions.
219
+
220
+ JSON responses retain `<section>_total`, `<section>_offset` (when positive), and
221
+ `<section>_truncated: true` whenever a page omits rows before or after it. Markdown,
222
+ plain, and Claude responses show the same total and offset on last and empty
223
+ pages. For example, offset 20 with limit 5 over 25 published orphans shows
224
+ `5 of 25 from offset 20`; offset 100 shows `0 of 25 from offset 100`. An empty
225
+ page does not mean the section has no findings. These totals describe the
226
+ published report arrays, which can themselves be bounded during extraction;
227
+ they do not establish complete source-reference coverage.
228
+
195
229
  ### Search completeness
196
230
 
197
231
  Search responses retain `query`, `result_count`, and `results`; `result_count`
198
232
  is the number returned, not an estimated total. The additive `completeness`
199
233
  object describes the requested types, literal filters, and fields in the pinned
200
- generation. This contract is unreleased after `2.0.0.beta2`.
234
+ generation. This contract is included in Woods `2.0.0`.
201
235
 
202
236
  | `reason` | `status` | `has_more` | `total_matches` |
203
237
  |---|---|---|---|
@@ -229,6 +263,31 @@ Detected missing, unreadable, or corrupt artifacts remain `isError: true` with
229
263
  `has_more`, `total_matches`, and `matched_lower_bound`; no successful empty
230
264
  result is substituted. Inspect `woods_status` and run `woods:validate`.
231
265
 
266
+ ### Dependency graph coverage
267
+
268
+ `dependencies` and `dependents` return relationships recorded in the published
269
+ index, not an exhaustive call graph or source-reference index. Extraction combines
270
+ runtime reflection with selective source scanning; arbitrary method-body constant
271
+ references (including references to generic PORO and library classes) may have no
272
+ edge. No dependents, a test-only dependent, or a completed traversal does not prove
273
+ there are no production callers. Verify important absence claims in source.
274
+
275
+ Supporting servers expose the annotated, paginated traversal result in
276
+ `structuredContent.data` for every renderer, including the default packaged
277
+ stdio and HTTP servers. Read `data.total_is_exact`, `data.graph_coverage`, budget
278
+ counters and optional explanation witnesses there; `content[0].text` and
279
+ `structuredContent.text` keep the same human-readable rendering. No `format`
280
+ tool argument is needed or accepted. This additive data payload is included in
281
+ Woods `2.0.0`; verify the installed response before relying on it. Older
282
+ human-renderer responses can carry only text. The structured nodes and witnesses
283
+ cover the same page, not an additional traversal or an unpaginated graph.
284
+
285
+ Successful responses carry `graph_coverage` with `scope: "published_relationships"`,
286
+ `source_references: "not_exhaustive"`, and a human-readable `notice`. Text formats
287
+ show the same notice, including compact, root-only and empty-page responses.
288
+ This response metadata and the total exactness field below are included in Woods
289
+ `2.0.0`; older servers need the same conservative interpretation.
290
+
232
291
  ### Dependency traversal budgets
233
292
 
234
293
  `dependencies` and `dependents` walk breadth-first in stored graph order. The
@@ -251,7 +310,16 @@ Exact-budget walks that finish all requested work are complete and have no
251
310
  `limit` (default 50) and `offset` only page that discovered result; they never
252
311
  change the walk budget or depth. On a partial traversal, `nodes_total`, when
253
312
  present for pagination, counts the discovered prefix, **not the full reachable
254
- graph**. Paging beyond that prefix stays partial. To explore more, narrow
313
+ graph**. Every successful response includes `total_is_exact`: false for a
314
+ budget cutoff, true when the requested walk finishes, even when its page is
315
+ truncated or empty. It is independent of `limit`/`offset` and is present for
316
+ unpaged answers too. Partial text answers say `Showing N of at least M (total
317
+ unknown: node_budget)` (or `edge_budget`), including when no pagination is needed.
318
+ `M` includes the root and counts the admitted prefix; it is a lower bound for the
319
+ requested root, depth, type/relationship filters and published generation, not a
320
+ count of all application callers. Exactness describes that same recorded-graph
321
+ scope and never implies exhaustive source coverage. Paging beyond that prefix
322
+ stays partial. To explore more, narrow
255
323
  `depth`/`types`/`via`, choose another root, or increase the traversal budget within
256
324
  its maximum. Keep the root, filters, budgets, and published generation unchanged
257
325
  for stable pages. No wall-clock deadline is used, so cutoffs are deterministic.
@@ -259,14 +327,14 @@ for stable pages. No wall-clock deadline is used, so cutoffs are deterministic.
259
327
  Budgets cover traversal work after per-generation graph loading and cache
260
328
  preparation (JSON parsing, typed-edge normalization, node types and database
261
329
  metadata). They do not cap that initial load, elapsed time, or total process
262
- memory. These arguments are unreleased in Woods 2.0.0.beta2; check the connected
330
+ memory. These arguments are included in Woods `2.0.0`; check the connected
263
331
  server's tool schema before sending them to an older installation.
264
332
 
265
333
  ### Traversal explanations
266
334
 
267
335
  Supporting development versions accept `explain: true` on `dependencies` and
268
- `dependents`. Check the connected schema first; this option is unreleased after
269
- 2.0.0.beta2. Omitted or false keeps the existing compact response.
336
+ `dependents`. Check the connected schema first; this option is included in Woods
337
+ `2.0.0`. Omitted or false keeps the existing compact response.
270
338
 
271
339
  The additive `explanation` object contains:
272
340
 
@@ -289,6 +357,10 @@ A target name shared by several types has `type: null`,
289
357
  record target types, so the response cannot choose among candidates. A witness
290
358
  through an ambiguous or unresolved identity sets `typed_path_complete: false`;
291
359
  it describes identifier-level reachability, never a uniquely typed path.
360
+ A true value means only that identities along this witness have unambiguous
361
+ types. It does not establish source-reference coverage or observed execution.
362
+ Text labels this `witness types unambiguous=yes/no`; the JSON key and its meaning
363
+ remain unchanged. The text label change is included in Woods `2.0.0`.
292
364
  `types` filters retain the compact traversal's identifier-level semantics: any
293
365
  registered type can qualify a name, while edge evidence keeps its actual source
294
366
  owner. Multiple relationship kinds between the same endpoints remain separate.
@@ -64,7 +64,7 @@ The Index Server defines **29 schemas**: the packaged executable registers **14*
64
64
  | Snapshot (4) | 4 | Extraction with `enable_snapshots = true` normally creates `woods.sqlite3`, which packaged servers discover. If extraction used the JSON fallback, set `WOODS_SNAPSHOTS=true` on the standalone server. Custom embedded servers pass `snapshot_store:`. Internal SQLite migrations are automatic. Tools: `list_snapshots`, `snapshot_diff`, `unit_history`, `snapshot_detail` |
65
65
  | `notion_sync` | 1 | `notion_api_token` + `notion_database_ids` both set |
66
66
 
67
- `codebase_retrieve` is always registered (no `retrieve` alias exists), but only returns results once an embedding provider is configured and `rake woods:embed` has run.
67
+ `codebase_retrieve` is always registered (no `retrieve` alias exists). Default semantic mode requires an embedding provider and a completed `woods:embed` run. Explicit `WOODS_RETRIEVAL_MODE=lexical` ranks published extraction units without a provider or embeddings; set it in the MCP process environment and restart the server. See [embedding-free lexical retrieval](RETRIEVAL_GUIDE.md#embedding-free-lexical-retrieval).
68
68
 
69
69
  If an agent reports a missing tool, compare its request with the connected server's registered list and [MCP server boundaries](MCP_SERVERS.md#conditional-index-capabilities). The normal packaged executable does not wire operator or feedback collaborators. **Console Server tools are not all unconditionally registered**: 31 tool schemas exist as an inventory, but only the 9 Tier 1 tools are executable by default, or 11 with `console_embedded_read_tools: true` (adds `console_sql`/`console_query`). Tier 2, Tier 3, and `console_eval` are schema-only in every supported mode; there is no bridge or confirmation flow that unlocks them. See [MCP servers](MCP_SERVERS.md#console-server) for the supported inventory.
70
70
 
@@ -356,7 +356,7 @@ column.
356
356
  Use `lookup` on the returned identifier for source, actions, and routes. Search
357
357
  `source_code` for textual matches beyond names. Supporting versions distinguish
358
358
  exact totals from a bounded result prefix; `partial` means this is discovery,
359
- not an exhaustive list. Completeness metadata is unreleased after `2.0.0.beta2`;
359
+ not an exhaustive list. Completeness metadata is included in Woods `2.0.0`;
360
360
  see the [search contract](MCP_SERVERS.md#search-completeness).
361
361
 
362
362
  ---
@@ -551,7 +551,7 @@ Static tools miss all of these because they only exist after Rails processes the
551
551
  }
552
552
  ```
553
553
 
554
- **What you'll get:** Units with no dependents, nothing in the codebase references them. Good candidates for removal or investigation.
554
+ **What you'll get:** Units with no recorded dependents in the published graph, excluding types treated as natural entry points. These are candidates for investigation, not proof of dead code. Method-body references and dynamic callers may be missing; verify source references, framework entry points, and runtime usage before removing anything. See [dependency graph coverage](MCP_SERVERS.md#dependency-graph-coverage).
555
555
 
556
556
  ---
557
557
 
@@ -811,7 +811,7 @@ Keys without a recognised suffix fall through to ActiveRecord `where(hash)` equa
811
811
 
812
812
  ### "Find code related to subscription billing"
813
813
 
814
- **Tool:** `codebase_retrieve` (Index Server, requires embedding provider)
814
+ **Tool:** `codebase_retrieve` (Index Server, semantic or explicit lexical mode)
815
815
 
816
816
  ```json
817
817
  {
@@ -820,7 +820,7 @@ Keys without a recognised suffix fall through to ActiveRecord `where(hash)` equa
820
820
  }
821
821
  ```
822
822
 
823
- **What you'll get:** A token-budgeted context string of the most semantically relevant units, ranked by hybrid search (semantic + keyword + PageRank). Requires an embedding provider (`embedding_provider: :openai` or `:ollama`) to be configured.
823
+ **What you'll get:** Ranked context within an estimated text-token budget. Default semantic mode uses configured embeddings and hybrid ranking; explicit lexical mode uses field-aware BM25 over published units, without a provider or `woods:embed`. The same query works in either configured mode, though rankings differ. Confirm the active mode with `woods_status.retriever.mode`; see the [retrieval guide](RETRIEVAL_GUIDE.md#embedding-free-lexical-retrieval) for setup and budget limits.
824
824
 
825
825
  ---
826
826
 
@@ -1,127 +1,99 @@
1
1
  # MCP Registration in Git Worktrees
2
2
 
3
- > **Claude Code specific.** This page covers Claude Code's MCP registration model (`/mcp`, `~/.claude/plugins/`); other MCP clients manage per-directory registration their own way.
3
+ > **Claude Code specific.** Other MCP clients manage registration and subagent tool access differently. Use the client's current documentation alongside the Woods [MCP setup guide](MCP_SERVERS.md).
4
4
 
5
- When you work in a git worktree, a separate directory checked out from the same repository, your MCP tools may not be available to subagents running in that directory. This page explains why, how to fix it, and how to confirm the registration took effect.
5
+ A git worktree is a separate checkout. For Woods, verify both which servers a session can access and which checkout's index those servers serve.
6
6
 
7
- ## Why Worktree Subagents May Not See Woods Tools
7
+ ## Separate sessions and inherited subagents
8
8
 
9
- MCP server registration in Claude Code is controlled by `.mcp.json` files. Claude Code discovers these files by walking up the directory tree from the working directory. It stops at the first `.mcp.json` it finds (or at `~/.claude/settings.json` for global registrations).
9
+ A **new Claude Code session launched in a worktree** can have different MCP registrations from a session in the main checkout:
10
10
 
11
- A git worktree has its own root directory separate from the main repository checkout. When a subagent starts inside the worktree root, it walks up from that path, not from the main repository root. Unless a `.mcp.json` exists inside the worktree directory tree or in an ancestor shared with both checkouts, the subagent sees no MCP servers.
11
+ | Registration scope | Location | Availability |
12
+ |---|---|---|
13
+ | Local | Project entry in `~/.claude.json` | The project where the server was registered |
14
+ | Project | `.mcp.json` in the project root | That project, subject to client approval/settings |
15
+ | User | `mcpServers` in `~/.claude.json` | Across the user's projects |
12
16
 
13
- Example directory layout:
17
+ A tracked `.mcp.json` may already be present in another worktree; a local, untracked file will not be copied by Git. User-level registration can make a server available across checkouts, but its configured paths still determine which index it serves.
14
18
 
15
- ```
16
- ~/work/my-app/ ← main checkout, has .mcp.json here
17
- ~/work/my-app-feature/ ← worktree, no .mcp.json, so MCP tools are missing
18
- ```
19
+ An **inherited subagent** receives the parent conversation's MCP tools, subject to tool restrictions. Changing the subagent's working directory does not by itself launch another Woods server or switch the served index. Distinguish this from starting an independent client process in that directory.
20
+
21
+ See Claude Code's [MCP installation scopes](https://code.claude.com/docs/en/mcp#mcp-installation-scopes) and [subagent tool access](https://code.claude.com/docs/en/sub-agents#available-tools). Do not diagnose registration by assuming the client searches ancestor directories for the nearest `.mcp.json`.
19
22
 
20
- A subagent spawned in `~/work/my-app-feature/` will not find the `.mcp.json` from `~/work/my-app/`.
23
+ ## Configure a server for the intended checkout
21
24
 
22
- ## Fix: Add a `.mcp.json` to the Worktree Root
25
+ First check `/mcp` in the session that needs Woods. If a suitable server is already connected, verify its index before adding a duplicate registration.
23
26
 
24
- Create a `.mcp.json` in the worktree's root directory with the same woods server entries you use in the main checkout:
27
+ For a host with the application's Ruby bundle installed, a project-root `.mcp.json` can select the bundle and index explicitly:
25
28
 
26
29
  ```json
27
30
  {
28
31
  "mcpServers": {
29
32
  "woods": {
30
- "command": "woods-mcp-start",
31
- "args": ["./tmp/woods"]
32
- },
33
- "woods-console": {
34
- "command": "docker",
35
- "args": [
36
- "compose", "exec", "-T", "app",
37
- "bundle", "exec", "rake", "woods:console"
38
- ],
39
- "cwd": "/absolute/host/path/to/worktree"
33
+ "command": "bundle",
34
+ "args": ["exec", "woods-mcp-start", "/absolute/path/to/worktree/tmp/woods"],
35
+ "env": {
36
+ "BUNDLE_GEMFILE": "/absolute/path/to/worktree/Gemfile"
37
+ }
40
38
  }
41
39
  }
42
40
  }
43
41
  ```
44
42
 
45
- Adjust the paths and arguments to match your project's setup. In particular:
46
-
47
- - `./tmp/woods` is a relative path, it resolves against the worktree root, which is correct if extraction output is written into each worktree separately.
48
- - If you share a single extraction output directory between checkouts, use the absolute path to the shared output: `"/absolute/path/to/main-checkout/tmp/woods"`.
49
- - For Docker projects, the `docker compose exec` command works the same from any host path.
50
-
51
- ## How Plugin Discovery Works
52
-
53
- Claude Code also discovers MCP servers registered in plugin manifests. A plugin at `~/.claude/plugins/<plugin-name>/admin-tools/.mcp.json` is loaded globally, its servers are available in any session regardless of working directory.
54
-
55
- If your team distributes the woods MCP registration through a shared plugin, the worktree problem does not apply. Check whether woods is already registered this way:
56
-
57
- ```bash
58
- ls ~/.claude/plugins/
59
- # look for a directory containing admin-tools/.mcp.json
60
- cat ~/.claude/plugins/<plugin-name>/admin-tools/.mcp.json
61
- ```
62
-
63
- If you find the woods servers registered there, subagents in any worktree will have access automatically, you do not need a per-worktree `.mcp.json`.
64
-
65
- ## Verifying MCP Registration for a Subagent
66
-
67
- ### Option 1: List tools from a claude session in the worktree
68
-
69
- Open a new Claude Code session with the worktree as the working directory and run:
70
-
71
- ```
72
- /mcp
73
- ```
74
-
75
- This lists all connected MCP servers and their tools. If `woods` and/or `woods-console` appear, registration is working.
43
+ Replace both paths with the intended Rails worktree. Preserve other server entries. These absolute paths are machine-specific; do not commit them as shared team defaults. Follow the client's approval/reconnection steps, then verify the connection below.
76
44
 
77
- ### Option 2: Check via the woods status tool
45
+ For Docker, use the [container-first setup](DOCKER_SETUP.md#default-run-it-through-the-application-container). Select the intended Compose project and service, and verify that its mounted application directory is the worktree you mean to index. `docker compose exec` targets an existing container; running the command from a different checkout does not change that container's mounts. The index path must be visible inside that container.
78
46
 
79
- Ask Claude to call the status tool:
47
+ Keep a separate output directory for each checkout when you need checkout-specific answers. If you intentionally point at another checkout's index, Woods serves that published content; registration alone does not make it describe the current worktree.
80
48
 
81
- ```
82
- Use woods-console console_status to check what models are available.
83
- ```
49
+ ## Plugin-provided servers
84
50
 
85
- If the tool runs successfully, MCP is registered and the console server is reachable.
51
+ An enabled plugin can supply MCP server definitions, but installing a plugin does not establish that a Woods server is connected to the intended checkout. Availability depends on the plugin's configuration, enabled scope, and the client's tool restrictions. Check `/mcp` rather than assuming a fixed path under `~/.claude/plugins/` or global availability.
86
52
 
87
- ### Option 3: Inspect the MCP config that Claude Code loaded
53
+ The Woods setup/configuration skills help configure a server; their presence alone is not verification that the Index Server is running. See [agent setup](AGENT_SETUP.md).
88
54
 
89
- From the worktree directory, run:
55
+ ## Verify registration and the served index
90
56
 
91
- ```bash
92
- cat .mcp.json
93
- ```
57
+ 1. Open `/mcp` in the relevant Claude Code session and inspect the Woods connection and available tools.
58
+ 2. Call the Index Server's `woods_status`. Check `index_dir`, generation, and available freshness/provenance information against the intended extraction. A successful connection to the wrong index is still the wrong setup.
59
+ 3. Use `search` and typed `lookup` for a known unit from that checkout. When checking a subagent, verify that it can call the inherited tools and is using the same intended index.
94
60
 
95
- If the file exists and contains the woods entries, Claude Code will use it. If the file is missing, check parent directories up to your home directory for any `.mcp.json` that would be discovered.
61
+ The optional Console Server is a separate connection to a booted Rails application. If deliberately enabled, check it with `console_status` and verify its application/container separately. Console access is not required to verify the Index Server.
96
62
 
97
63
  ## Extraction Provenance in Worktrees (`git_branch` / `git_sha`)
98
64
 
99
- `manifest.json` records the `git_branch` and `git_sha` the extraction ran against. In a linked worktree, `.git` is a **file** containing a `gitdir:` pointer to the real git directory, frequently an absolute host path, rather than a `.git` directory.
100
-
101
- Woods resolves provenance with worktree-aware git plumbing (`git -C <root> rev-parse`), so an ordinary worktree reports the correct branch and SHA. When a `.git` is present but the pointed-to git directory **cannot be resolved**: most commonly a worktree extracted inside a container where the host path isn't mounted. Woods records `git_branch: "unknown"` / `git_sha: "unknown"` rather than a stale, misleading value: a baked `GIT_BRANCH`/`GIT_SHA` build arg is **not** trusted here (it could be stale). The env vars are honored only when there is no `.git` at the root at all (a non-repo checkout, e.g. a Docker `COPY` that excludes `.git`) or git is unavailable.
65
+ The published payload's `manifest.json` records the extraction's `git_branch` and `git_sha`. In a linked worktree, `.git` is a file containing a `gitdir:` pointer, often to an absolute host path. That private worktree directory also refers to the parent repository's shared Git data.
102
66
 
103
- To get correct provenance from a containerized worktree, mount the directory the `gitdir:` pointer references (the parent repository's `.git`) into the container so git can resolve it. Temporal snapshots skip an `"unknown"` SHA, so misleading provenance never keys a snapshot.
67
+ Woods uses worktree-aware Git commands. If a present `.git` cannot be resolved, provenance is `"unknown"`; stale `GIT_BRANCH`/`GIT_SHA` values are not substituted. Those environment variables are fallbacks only when the root has no `.git` or Git is unavailable. Temporal snapshots skip an unknown SHA.
104
68
 
105
- ## Troubleshooting
69
+ For extraction in a container, mount the complete shared Git layout at the
70
+ original path so the worktree's `.git` pointer resolves without an override,
71
+ or select `/mounted-common/worktrees/<id>` with `WOODS_GIT_DIR` inside a
72
+ relocated complete mount. Derive `<id>` from Git metadata, not the branch name.
73
+ Selecting the shared root instead uses the primary checkout's HEAD, affecting
74
+ provenance, per-file history, and incremental paths. Follow the
75
+ [worktree mount and verification steps](TROUBLESHOOTING.md#git-directory-mounts-for-linked-worktrees)
76
+ and the [published index layout](INDEX_LAYOUT.md) when locating the manifest.
106
77
 
107
- **"Unknown tool" or "no MCP server named woods"**
78
+ Compare the branch and exact SHA in the extraction environment with the intended
79
+ worktree, then verify the published manifest after full extraction. A commit
80
+ alone may not trigger the source-file watcher; use full `woods:extract` when
81
+ current Git history is required.
108
82
 
109
- The woods MCP server is not registered for this session. Add a `.mcp.json` to the worktree root as shown above, then restart the session.
83
+ ## Troubleshooting
110
84
 
111
- **Console server starts but returns "unsupported: Not yet implemented in embedded mode" for console_sql / console_query**
85
+ **Woods or an expected tool is absent**
112
86
 
113
- The console server is registered and reachable, but `embedded_read_tools` is disabled (the default). To enable console_sql and console_query in embedded mode, see the [Console MCP Setup guide](CONSOLE_MCP_SETUP.md), specifically the `embedded_read_tools: true` option for the Rack middleware.
87
+ Check `/mcp` for connection errors, enabled registration scopes, and project approvals. For a subagent, also check its tool restrictions. Compare the requested tool with the [supported tool surface](MCP_SERVERS.md#conditional-index-capabilities); some tools require optional collaborators and are not registered by the packaged server. Do not add duplicate registrations before establishing which case applies.
114
88
 
115
- **Extraction output is empty or stale in the worktree**
89
+ **The server connects but shows another checkout**
116
90
 
117
- Extraction writes to `tmp/woods/` relative to the Rails application root (inside the container). If the worktree's volume mount points to a different host path than the main checkout, run extraction again from the worktree:
91
+ Inspect its configured bundle, index path, Compose project/service, and container mounts. Re-extract from the intended Rails checkout into its own output directory, then reconnect to that index and repeat `woods_status`. See [Docker setup](DOCKER_SETUP.md).
118
92
 
119
- ```bash
120
- docker compose exec app bundle exec rake woods:extract
121
- ```
93
+ **Two registrations use the same server name**
122
94
 
123
- See [DOCKER_SETUP.md](DOCKER_SETUP.md) for the full Docker workflow.
95
+ Check the client's [scope precedence](https://code.claude.com/docs/en/mcp#scope-hierarchy-and-precedence) and the effective connection shown by `/mcp`. Do not assume the two definitions merge or that a project file overrides every other scope.
124
96
 
125
- **Worktree `.mcp.json` conflicts with main checkout `.mcp.json`**
97
+ **Console SQL/query tools are absent or unsupported**
126
98
 
127
- Each file is independent. Claude Code loads the one closest to the working directory. There is no inheritance or merging between them. Keep both files in sync manually, or move the shared configuration into a global plugin manifest.
99
+ Console read-tool availability is configured separately from registration. Follow [Console MCP setup](CONSOLE_MCP_SETUP.md) for the supported 9-tool default and optional 11-tool mode.
@@ -109,6 +109,23 @@ The reader wraps `Woods::MCP::IndexReader` with `auto_refresh: false`; the unit
109
109
 
110
110
  `Woods::MCP::IndexReader#find_unit` keys its identifier map on identifier alone. If two type directories both list the same identifier (a model and a service both named `Foo`, for example), whichever type sorts last in `Woods::MCP::IndexReader::TYPE_DIRS` silently wins, and `unit(identifier)` returns that one. Pass `type:` to read a specific type's unit file directly and skip the collision entirely; `#table_database_map` always does this internally (`type: 'model'`), so a same-named non-model unit can never shadow a model's `table_name`/`database`.
111
111
 
112
+ ### Actual unit types and directory families
113
+
114
+ Included in Woods `2.0.0`: `unit` and `units` accept actual published
115
+ `graphql_type`, `graphql_mutation`, `graphql_resolver`, `graphql_query`, and
116
+ `gem_source` types. Enumeration preserves each unit's actual type rather than
117
+ labeling every GraphQL member `graphql` or every gem source `rails_source`.
118
+ Check the loaded Git revision when testing a development checkout.
119
+
120
+ The existing `graphql` and `rails_source` filters remain directory-family
121
+ aliases: `graphql` selects all four GraphQL types; `rails_source` selects both
122
+ Rails and gem sources. Returned records keep their actual type. Use
123
+ `units(type: 'gem_source')` for gem sources only; to select Rails sources only,
124
+ filter `units(type: 'rails_source')` on each entry's `"type" == "rails_source"`.
125
+ Unknown types return no records, and an explicit GraphQL subtype never matches
126
+ another subtype. Enumeration continues to return index-entry fields, not full
127
+ unit source bodies.
128
+
112
129
  ### `available_generations`: published means published
113
130
 
114
131
  A generation is listed only when both hold:
data/docs/README.md CHANGED
@@ -10,7 +10,7 @@ Woods extracts runtime-accurate Rails context and serves it to coding agents thr
10
10
  | Ask an agent to install or configure Woods | [Agent setup runbook](AGENT_SETUP.md) | A safe, reviewable install with an agent handoff report |
11
11
  | Configure an MCP client or Docker path | [MCP servers](MCP_SERVERS.md) | A working Index Server and, if authorized, an optional Console Server |
12
12
  | Use Woods tools as an agent | [Agent guide](AGENT_GUIDE.md) | A repeatable query workflow for code context, flows, and blast radius |
13
- | Keep the index current automatically | [Watch daemon](WATCH_DAEMON.md) | A resident development process that catches up changes and republishes the index |
13
+ | Keep the index current automatically | [Automatic maintenance guide](AUTOMATIC_MAINTENANCE.md) | A setup with supervised indexing, automatic reader refresh, and clear hook ownership |
14
14
  | Upgrade from Woods 1.x | [Upgrade to Woods 2.0](UPGRADING_TO_2.md) | A backed-up, re-indexed, verified v2 installation |
15
15
  | Diagnose an error | [Troubleshooting](TROUBLESHOOTING.md) | Symptom-to-cause checks for extraction, MCP, embeddings, storage, and Docker |
16
16
  | Contribute to Woods | [Contributing](../CONTRIBUTING.md) | A tested change with synchronized docs and plugin guidance |
@@ -41,6 +41,7 @@ and stdio or Streamable HTTP endpoints directly.
41
41
 
42
42
  ## Index lifecycle
43
43
 
44
+ - [Automatic maintenance guide](AUTOMATIC_MAINTENANCE.md): recommended development workflow and a section-by-section map of watcher, indexing, MCP, and hook documentation.
44
45
  - [Source freshness](SOURCE_FRESHNESS.md): verify dirty source against a served generation, establish a fresh-process baseline, and understand bounded unknown results.
45
46
 
46
47
  - [Retrieval guide](RETRIEVAL_GUIDE.md): configure embeddings and understand semantic retrieval, ranking, and token budgets.