woods 2.0.0.beta2 → 2.0.0.beta4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (233) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +339 -1
  3. data/CONTRIBUTING.md +188 -12
  4. data/README.md +93 -174
  5. data/SECURITY.md +9 -6
  6. data/docs/AGENT_GUIDE.md +109 -8
  7. data/docs/AGENT_SETUP.md +98 -7
  8. data/docs/BACKEND_MATRIX.md +25 -0
  9. data/docs/CLIENT_HOOKS.md +111 -0
  10. data/docs/CONFIGURATION_REFERENCE.md +267 -16
  11. data/docs/CONSOLE_MCP_SETUP.md +80 -7
  12. data/docs/DOCKER_SETUP.md +22 -3
  13. data/docs/EVALUATION.md +464 -1
  14. data/docs/EXTRACTOR_REFERENCE.md +45 -6
  15. data/docs/FAQ.md +11 -12
  16. data/docs/GETTING_STARTED.md +17 -5
  17. data/docs/INCREMENTAL_EXTRACTION.md +147 -7
  18. data/docs/INDEX_LAYOUT.md +382 -0
  19. data/docs/INTERNALS.md +7 -2
  20. data/docs/MCP_SERVERS.md +276 -5
  21. data/docs/MCP_TOOL_COOKBOOK.md +37 -22
  22. data/docs/MCP_WORKTREE_SETUP.md +43 -83
  23. data/docs/NOTION_INTEGRATION.md +13 -0
  24. data/docs/OBSIDIAN_INTEGRATION.md +57 -9
  25. data/docs/PUBLISHED_INDEX.md +72 -0
  26. data/docs/README.md +7 -0
  27. data/docs/RETRIEVAL_GUIDE.md +273 -12
  28. data/docs/RUNTIME_TRACING.md +71 -0
  29. data/docs/SOURCE_FRESHNESS.md +143 -0
  30. data/docs/TROUBLESHOOTING.md +129 -18
  31. data/docs/UNBLOCKED_INTEGRATION.md +25 -0
  32. data/docs/UPGRADING_TO_2.md +48 -22
  33. data/docs/WATCH_DAEMON.md +277 -67
  34. data/exe/woods-agent-config +6 -0
  35. data/exe/woods-extract +5 -0
  36. data/exe/woods-hook-context +6 -0
  37. data/exe/woods-mcp-start +14 -9
  38. data/lib/generators/woods/pgvector_generator.rb +8 -2
  39. data/lib/generators/woods/templates/woods.rb.tt +1 -3
  40. data/lib/tasks/woods.rake +47 -397
  41. data/lib/woods/agent_configuration/applier.rb +135 -0
  42. data/lib/woods/agent_configuration/cli.rb +101 -0
  43. data/lib/woods/agent_configuration/cli_options.rb +29 -0
  44. data/lib/woods/agent_configuration/document.rb +105 -0
  45. data/lib/woods/agent_configuration/error.rb +7 -0
  46. data/lib/woods/agent_configuration/launcher.rb +75 -0
  47. data/lib/woods/agent_configuration/layout.rb +72 -0
  48. data/lib/woods/agent_configuration/managed_section.rb +62 -0
  49. data/lib/woods/agent_configuration/plan.rb +98 -0
  50. data/lib/woods/agent_configuration/plan_diff.rb +38 -0
  51. data/lib/woods/agent_configuration/planned_files.rb +61 -0
  52. data/lib/woods/agent_configuration/planner.rb +63 -0
  53. data/lib/woods/agent_configuration/planner_validation.rb +77 -0
  54. data/lib/woods/agent_configuration/preflight.rb +100 -0
  55. data/lib/woods/agent_configuration/recovery.rb +49 -0
  56. data/lib/woods/ast/node.rb +2 -0
  57. data/lib/woods/ast/parser.rb +38 -5
  58. data/lib/woods/builder.rb +21 -5
  59. data/lib/woods/cache/cache_middleware.rb +28 -7
  60. data/lib/woods/cache/cache_store.rb +4 -5
  61. data/lib/woods/change_set.rb +5 -4
  62. data/lib/woods/console/credential_index.rb +20 -2
  63. data/lib/woods/console/credential_scanner.rb +18 -17
  64. data/lib/woods/console/credential_scanner_registry.rb +36 -0
  65. data/lib/woods/console/dispatch_pipeline.rb +7 -0
  66. data/lib/woods/console/embedded_executor.rb +32 -10
  67. data/lib/woods/console/encrypted_credential_snapshot.rb +16 -0
  68. data/lib/woods/console/rack_middleware.rb +22 -13
  69. data/lib/woods/console/server.rb +18 -16
  70. data/lib/woods/console/sql_noise_stripper.rb +9 -7
  71. data/lib/woods/console/sql_table_scanner.rb +47 -7
  72. data/lib/woods/console/sql_validator.rb +49 -9
  73. data/lib/woods/console/sqlite_read_guard.rb +46 -0
  74. data/lib/woods/coordination/pipeline_lock.rb +3 -2
  75. data/lib/woods/dependency_graph.rb +65 -13
  76. data/lib/woods/embedding/corpus.rb +94 -0
  77. data/lib/woods/embedding/indexer.rb +114 -60
  78. data/lib/woods/embedding/openai.rb +17 -6
  79. data/lib/woods/evaluation/ablation_executor.rb +6 -1
  80. data/lib/woods/evaluation/ablation_timed_executor.rb +22 -4
  81. data/lib/woods/export/typed_reader.rb +56 -0
  82. data/lib/woods/extractor.rb +277 -149
  83. data/lib/woods/extractors/action_cable_extractor.rb +3 -1
  84. data/lib/woods/extractors/behavioral_profile.rb +9 -7
  85. data/lib/woods/extractors/caching_extractor.rb +3 -1
  86. data/lib/woods/extractors/concern_extractor.rb +64 -6
  87. data/lib/woods/extractors/configuration_extractor.rb +7 -3
  88. data/lib/woods/extractors/controller_extractor.rb +13 -4
  89. data/lib/woods/extractors/database_view_extractor.rb +3 -1
  90. data/lib/woods/extractors/declared_parent.rb +55 -0
  91. data/lib/woods/extractors/decorator_extractor.rb +3 -1
  92. data/lib/woods/extractors/engine_extractor.rb +3 -1
  93. data/lib/woods/extractors/event_extractor.rb +4 -2
  94. data/lib/woods/extractors/factory_extractor.rb +3 -1
  95. data/lib/woods/extractors/graphql_extractor.rb +10 -13
  96. data/lib/woods/extractors/i18n_extractor.rb +3 -1
  97. data/lib/woods/extractors/job_extractor.rb +6 -19
  98. data/lib/woods/extractors/lib_extractor.rb +13 -9
  99. data/lib/woods/extractors/mailer_extractor.rb +26 -15
  100. data/lib/woods/extractors/manager_extractor.rb +3 -1
  101. data/lib/woods/extractors/method_parameters.rb +53 -0
  102. data/lib/woods/extractors/middleware_argument.rb +65 -0
  103. data/lib/woods/extractors/middleware_extractor.rb +9 -3
  104. data/lib/woods/extractors/migration_extractor.rb +3 -1
  105. data/lib/woods/extractors/model_extractor.rb +26 -34
  106. data/lib/woods/extractors/package_extractor.rb +24 -4
  107. data/lib/woods/extractors/phlex_extractor.rb +3 -1
  108. data/lib/woods/extractors/policy_extractor.rb +3 -1
  109. data/lib/woods/extractors/poro_extractor.rb +13 -9
  110. data/lib/woods/extractors/pundit_extractor.rb +3 -1
  111. data/lib/woods/extractors/rails_source_extractor.rb +4 -2
  112. data/lib/woods/extractors/rake_task_extractor.rb +4 -2
  113. data/lib/woods/extractors/route_extractor.rb +3 -1
  114. data/lib/woods/extractors/route_helper_resolver.rb +10 -33
  115. data/lib/woods/extractors/scheduled_job_extractor.rb +41 -15
  116. data/lib/woods/extractors/serializer_extractor.rb +4 -2
  117. data/lib/woods/extractors/service_extractor.rb +3 -1
  118. data/lib/woods/extractors/shared_dependency_scanner.rb +2 -2
  119. data/lib/woods/extractors/shared_utility_methods.rb +48 -19
  120. data/lib/woods/extractors/source_nesting.rb +1 -1
  121. data/lib/woods/extractors/state_machine_extractor.rb +3 -1
  122. data/lib/woods/extractors/test_mapping_extractor.rb +3 -1
  123. data/lib/woods/extractors/validator_extractor.rb +3 -1
  124. data/lib/woods/extractors/view_component_extractor.rb +3 -1
  125. data/lib/woods/extractors/view_template_extractor.rb +3 -1
  126. data/lib/woods/gem_mapper.rb +2 -0
  127. data/lib/woods/git_history.rb +116 -0
  128. data/lib/woods/graph_analyzer.rb +35 -6
  129. data/lib/woods/hooks/context_cli.rb +54 -0
  130. data/lib/woods/hooks/context_event.rb +88 -0
  131. data/lib/woods/hooks/context_hint.rb +73 -0
  132. data/lib/woods/hooks/context_impact.rb +77 -0
  133. data/lib/woods/hooks/context_output.rb +47 -0
  134. data/lib/woods/hooks/context_state.rb +102 -0
  135. data/lib/woods/hooks/refresh.rb +79 -0
  136. data/lib/woods/hooks/rule_projection.rb +78 -0
  137. data/lib/woods/input_rules.rb +19 -0
  138. data/lib/woods/mcp/bearer_auth.rb +22 -13
  139. data/lib/woods/mcp/bootstrapper.rb +79 -4
  140. data/lib/woods/mcp/config_resolver.rb +2 -1
  141. data/lib/woods/mcp/index_reader.rb +334 -162
  142. data/lib/woods/mcp/initialization_guidance.rb +27 -0
  143. data/lib/woods/mcp/origin_guard.rb +17 -9
  144. data/lib/woods/mcp/published_lexical_retriever.rb +115 -0
  145. data/lib/woods/mcp/renderers/markdown_renderer.rb +22 -9
  146. data/lib/woods/mcp/renderers/plain_renderer.rb +18 -8
  147. data/lib/woods/mcp/search_results.rb +74 -0
  148. data/lib/woods/mcp/server.rb +178 -63
  149. data/lib/woods/mcp/tool_contract.rb +3 -1
  150. data/lib/woods/mcp/tool_response_renderer.rb +41 -0
  151. data/lib/woods/mcp/traversal_evidence.rb +113 -0
  152. data/lib/woods/mcp/traversal_evidence_index.rb +100 -0
  153. data/lib/woods/mcp/traversal_evidence_page.rb +41 -0
  154. data/lib/woods/mcp/traversal_evidence_text.rb +52 -0
  155. data/lib/woods/mcp/traversal_response.rb +22 -0
  156. data/lib/woods/notion/exporter.rb +56 -17
  157. data/lib/woods/obsidian/destination_plan.rb +98 -0
  158. data/lib/woods/obsidian/name_mapper.rb +19 -3
  159. data/lib/woods/obsidian/note_builder.rb +19 -10
  160. data/lib/woods/obsidian/vault_exporter.rb +88 -32
  161. data/lib/woods/operator/pipeline_guard.rb +18 -13
  162. data/lib/woods/path_dispatcher.rb +13 -6
  163. data/lib/woods/payload_store.rb +27 -26
  164. data/lib/woods/published_index/typed_unit_reader.rb +40 -3
  165. data/lib/woods/published_index.rb +2 -2
  166. data/lib/woods/railtie.rb +3 -3
  167. data/lib/woods/railtie_support.rb +12 -12
  168. data/lib/woods/rake_helpers.rb +382 -0
  169. data/lib/woods/resilience/graph_invariant_validator/membership_checks.rb +71 -0
  170. data/lib/woods/resilience/graph_invariant_validator/node_checks.rb +61 -0
  171. data/lib/woods/resilience/graph_invariant_validator/reverse_relationship_checks.rb +46 -0
  172. data/lib/woods/resilience/graph_invariant_validator.rb +119 -0
  173. data/lib/woods/resilience/index_validator/graph_checks.rb +80 -0
  174. data/lib/woods/resilience/index_validator.rb +112 -23
  175. data/lib/woods/retrieval/context_assembler.rb +50 -15
  176. data/lib/woods/retrieval/lexical_assembler.rb +84 -0
  177. data/lib/woods/retrieval/lexical_index.rb +120 -0
  178. data/lib/woods/retrieval/ranker.rb +4 -2
  179. data/lib/woods/retrieval/scope.rb +108 -0
  180. data/lib/woods/retrieval/scoped_graph_store.rb +32 -0
  181. data/lib/woods/retrieval/scoped_vector_store.rb +55 -0
  182. data/lib/woods/retrieval/search_executor.rb +86 -27
  183. data/lib/woods/retrieval/source_evidence.rb +200 -0
  184. data/lib/woods/retriever.rb +98 -22
  185. data/lib/woods/ruby_analyzer/trace_enricher.rb +77 -38
  186. data/lib/woods/session_tracer/file_store.rb +6 -1
  187. data/lib/woods/session_tracer/middleware.rb +10 -12
  188. data/lib/woods/session_tracer/redis_store.rb +22 -6
  189. data/lib/woods/session_tracer/session_flow_assembler.rb +23 -17
  190. data/lib/woods/session_tracer/solid_cache_coordination.rb +6 -4
  191. data/lib/woods/session_tracer/unit_resolver.rb +63 -0
  192. data/lib/woods/source_inputs/consumer_errors.rb +31 -0
  193. data/lib/woods/source_inputs/handoff.rb +102 -0
  194. data/lib/woods/source_inputs/launcher.rb +157 -0
  195. data/lib/woods/source_inputs/manifest.rb +124 -0
  196. data/lib/woods/source_inputs/private_key.rb +55 -0
  197. data/lib/woods/source_inputs/scanner.rb +171 -0
  198. data/lib/woods/source_inputs/scopes.rb +71 -0
  199. data/lib/woods/source_inputs/session.rb +214 -0
  200. data/lib/woods/source_inputs/status.rb +84 -0
  201. data/lib/woods/source_inputs/verifier.rb +107 -0
  202. data/lib/woods/storage/metadata_store.rb +25 -25
  203. data/lib/woods/storage/pgvector.rb +35 -10
  204. data/lib/woods/storage/qdrant.rb +17 -7
  205. data/lib/woods/storage/vector_store.rb +18 -6
  206. data/lib/woods/tasks.rb +3 -2
  207. data/lib/woods/temporal/json_snapshot_store.rb +58 -9
  208. data/lib/woods/unblocked/exporter.rb +59 -70
  209. data/lib/woods/version.rb +1 -1
  210. data/lib/woods/watch/boot_snapshot.rb +52 -0
  211. data/lib/woods/watch/daemon.rb +154 -32
  212. data/lib/woods/watch/listen_watcher.rb +4 -0
  213. data/lib/woods/watch/polling_watcher.rb +5 -1
  214. data/lib/woods/watch/status.rb +20 -15
  215. data/lib/woods/watch/tree_scan.rb +21 -13
  216. data/lib/woods/watch/watcher.rb +4 -1
  217. data/lib/woods.rb +50 -11
  218. data/plugin/.claude-plugin/plugin.json +1 -1
  219. data/plugin/hooks/adapters/normalize.jq +15 -0
  220. data/plugin/hooks/adapters/normalize.rb +63 -0
  221. data/plugin/hooks/hooks.json +20 -0
  222. data/plugin/hooks/woods-context.sh +50 -0
  223. data/plugin/hooks/woods-input-rules.sh +159 -0
  224. data/plugin/hooks/woods-opencode.mjs +65 -0
  225. data/plugin/hooks/woods-post-edit.sh +2 -225
  226. data/plugin/hooks/woods-refresh.sh +260 -0
  227. data/plugin/hooks/woods-session-start.sh +47 -55
  228. data/plugin/skills/woods-agent-enable/SKILL.md +19 -0
  229. data/plugin/skills/woods-diagnose/SKILL.md +319 -1
  230. data/plugin/skills/woods-investigate/SKILL.md +145 -0
  231. data/plugin/skills/woods-mcp-config/SKILL.md +90 -2
  232. data/plugin/skills/woods-setup/SKILL.md +110 -6
  233. metadata +87 -5
data/README.md CHANGED
@@ -4,38 +4,23 @@
4
4
 
5
5
  # Woods
6
6
 
7
- **Give AI coding agents a runtime-accurate map of your Rails application.**
7
+ **Give coding agents the Rails context that source files alone leave out.**
8
8
 
9
9
  [![Gem Version](https://img.shields.io/gem/v/woods)](https://rubygems.org/gems/woods)
10
10
  [![CI](https://github.com/lost-in-the/woods/actions/workflows/ci.yml/badge.svg?branch=main)](https://github.com/lost-in-the/woods/actions/workflows/ci.yml)
11
11
  [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE.txt)
12
12
 
13
- <!-- release-state:version-banner -->
14
- > **This tree documents version 2.0.0.** It is a major update from 1.x: read [what changed and how to upgrade](docs/UPGRADING_TO_2.md) before updating. The full history is in the [CHANGELOG](CHANGELOG.md).
15
- >
16
- > `main` is the development branch and can run ahead of the latest published gem. The gem badge above shows the latest published version; documentation for a published version lives on its tag.
17
- >
18
- > ### Version: 2.0.0.beta2 is published as a prerelease; `main` documents 2.0.0
19
- >
20
- > | Line | Version | Documentation |
21
- > |---|---|---|
22
- > | Documented here | **2.0.0**, unreleased | this README and the [documentation index](docs/README.md) |
23
- > | Latest prerelease | **2.0.0.beta2** | [the v2.0.0.beta2 tag](https://github.com/lost-in-the/woods/tree/v2.0.0.beta2) |
24
- > | Latest published gem | **1.6.1** | [the v1.6.1 tag](https://github.com/lost-in-the/woods/tree/v1.6.1) |
25
- >
26
- > RubyGems treats 2.0.0.beta2 as a prerelease, so `gem "woods", "~> 2.0"` does not resolve it. Install it explicitly with `gem "woods", "2.0.0.beta2"`. The released constraint stays `gem "woods", "~> 1.6"`.
27
- <!-- release-state:end -->
13
+ Woods boots your Rails application, extracts its resolved structure, and publishes an index that coding agents can query through the [Model Context Protocol (MCP)](https://modelcontextprotocol.io/). It brings together database schema, associations, callbacks, concerns, routes, and source code so an agent can inspect how Rails assembles your application.
28
14
 
29
- Woods boots your Rails app, extracts the behavior Rails assembles at runtime, and serves it to AI tools through the [Model Context Protocol (MCP)](https://modelcontextprotocol.io/). Agents can inspect resolved routes, schema, associations, callbacks, included concerns, dependencies, and execution flows instead of guessing from source files alone.
15
+ Supports **Ruby 3.0+ and Rails 6.0–8.x**, using a Ruby version supported by your Rails release. The application must boot and connect to its database. Structural queries need no embedding provider or vector database.
30
16
 
31
- Woods 2.0 supports Ruby 3.0 or later and Rails 6.0 through 8.x. It connects AI coding tools and agents through MCP.
17
+ [Get started](#five-minute-setup) · [Documentation](docs/README.md) · [Agent setup](docs/AGENT_SETUP.md) · [Upgrade from 1.x](docs/UPGRADING_TO_2.md)
32
18
 
33
19
  ## What Woods adds
34
20
 
35
- A Rails model rarely lives in one file. Its real behavior can include database schema, generated methods, framework defaults, and concerns loaded from elsewhere:
21
+ Consider a model whose behavior is spread across Rails, the database, and a concern:
36
22
 
37
23
  ```ruby
38
- # app/models/order.rb
39
24
  class Order < ApplicationRecord
40
25
  include Auditable
41
26
  belongs_to :customer
@@ -43,40 +28,58 @@ class Order < ApplicationRecord
43
28
  end
44
29
  ```
45
30
 
46
- Woods turns that runtime class into one connected unit with:
31
+ Woods can give an agent one unit containing its columns and indexes, association metadata, resolved callbacks, and included concern source. Recorded relationships connect that unit to other parts of the application.
47
32
 
48
- - column types, indexes, and foreign keys from the live database;
49
- - associations, validations, scopes, enums, and resolved callbacks;
50
- - source from included concerns, kept beside the owning class;
51
- - callback side effects such as jobs, mailers, and columns written;
52
- - forward dependencies and reverse dependents;
53
- - route, controller, view, job, and service relationships.
33
+ An agent can then ask:
54
34
 
55
- The result is a codebase index an agent can query by exact name, pattern, dependency path, graph structure, or natural language.
35
+ > Find the Order model, inspect its callbacks and associations, and show its recorded dependents. Cite the indexed evidence and check source code for callers the graph may miss.
56
36
 
57
- Still weighing it? [Why Woods](docs/WHY_WOODS.md) makes the case against grep, cloud indexers, and IDE language servers.
37
+ Models are one part of the index: Woods also extracts controllers, routes, jobs, mailers, views, components, GraphQL types, service objects, tests, and more. See the [extractor reference](docs/EXTRACTOR_REFERENCE.md) for coverage and the [agent guide](docs/AGENT_GUIDE.md) for query examples.
58
38
 
59
39
  ## Five-minute setup
60
40
 
61
- The default setup provides structural code intelligence. It does not require an embedding provider, vector database, or access to live application records.
41
+ For agent-led setup, use the [agent installation option](#let-an-agent-install-it) and its runbook. For a new manual installation, follow the steps below.
62
42
 
63
- ### 1. Install Woods
43
+ **Already using Woods?** If you are upgrading from 1.x, follow the [upgrade guide](docs/UPGRADING_TO_2.md). For an existing 2.x installation, go directly to [retrieval modes](#retrieval-with-or-without-embeddings), [MCP configuration](docs/MCP_SERVERS.md), or the [configuration reference](docs/CONFIGURATION_REFERENCE.md). Preserve your initializer, index path, provider settings, and other client entries. Changing only the MCP launch configuration or retrieval mode does not require rerunning the installer or rebuilding the structural index.
64
44
 
65
- ```ruby
66
- # Gemfile
67
- group :development do
68
- gem "woods", "~> 2.0"
69
- end
70
- ```
45
+ Run installation and extraction commands from your Rails application root in its normal development environment. **Using Docker?** Follow [Docker setup](docs/DOCKER_SETUP.md) first: run those commands inside the application container and use paths visible to the process that runs MCP.
46
+
47
+ ### 1. Install and configure
48
+
49
+ These steps are for **Woods 2.x**. Choose a published 2.x version from the release information below and confirm it on [RubyGems](https://rubygems.org/gems/woods/versions). If only prereleases are available, use an exact prerelease pin; `~> 2.0` will not select one. Follow the chosen version's tag documentation rather than assuming every feature on `main` is published. If you choose 1.x, use its tag documentation instead of this quickstart.
50
+
51
+ <details>
52
+ <summary>Release information and Gemfile version constraints</summary>
53
+
54
+ <!-- release-state:version-banner -->
55
+ > **This tree documents version 2.0.0.** It is a major update from 1.x: read [what changed and how to upgrade](docs/UPGRADING_TO_2.md) before updating. The full history is in the [CHANGELOG](CHANGELOG.md).
56
+ >
57
+ > `main` is the development branch and can run ahead of the latest published gem. The gem badge above shows the latest published version; documentation for a published version lives on its tag.
58
+ >
59
+ > ### Version: this tree declares prerelease 2.0.0.beta4; `main` documents 2.0.0
60
+ >
61
+ > | Line | Version | Documentation |
62
+ > |---|---|---|
63
+ > | Documented here | **2.0.0**, unreleased | this README and the [documentation index](docs/README.md) |
64
+ > | Declared prerelease | **2.0.0.beta4** | [the v2.0.0.beta4 tag](https://github.com/lost-in-the/woods/tree/v2.0.0.beta4) |
65
+ > | Latest published gem | **1.6.2** | [the v1.6.2 tag](https://github.com/lost-in-the/woods/tree/v1.6.2) |
66
+ >
67
+ > RubyGems treats 2.0.0.beta4 as a prerelease, so `gem "woods", "~> 2.0"` does not resolve it. Once published, install it explicitly with `gem "woods", "2.0.0.beta4"`. The released constraint stays `gem "woods", "~> 1.6"`.
68
+ <!-- release-state:end -->
69
+
70
+ </details>
71
+
72
+ Expand the release information above, then add its appropriate `gem "woods", …` declaration to your Gemfile's `:development` group and run:
71
73
 
72
74
  ```bash
73
75
  bundle install
76
+ bundle exec ruby -rwoods/version -e 'puts Woods::VERSION'
74
77
  bin/rails generate woods:install
75
78
  ```
76
79
 
77
- **Do not run the generated migration for a new default installation.** The generator creates an annotated `config/initializers/woods.rb` plus a legacy application migration for `woods_units`, `woods_edges`, and `woods_embeddings`. Woods 2's shipped structural index and storage backends do not use those application tables. Remove the migration before continuing; keep and run it only when deliberately preserving an older/custom integration that uses them.
80
+ **For a new default installation, remove the generated `db/migrate/*_create_woods_tables.rb` migration without running it.** Those legacy application tables are unused by the shipped index and storage backends. Keep the generated `config/initializers/woods.rb`; its defaults are sufficient. Only retain the migration for a deliberate older/custom integration. See [Getting started](docs/GETTING_STARTED.md#2-generate-and-review-configuration).
78
81
 
79
- ### 2. Extract and verify the codebase
82
+ ### 2. Extract and validate
80
83
 
81
84
  ```bash
82
85
  bin/rails woods:extract
@@ -84,11 +87,11 @@ bin/rails woods:validate
84
87
  bin/rails woods:stats
85
88
  ```
86
89
 
87
- Extraction must run where Rails can boot. The default index lives at `tmp/woods/`.
90
+ Run these where your Rails application can boot. The default output is `tmp/woods/`; keep this generated directory out of source control.
88
91
 
89
- ### 3. Connect the Index Server
92
+ ### 3. Connect your MCP client
90
93
 
91
- Add this to your MCP client's project configuration. The configuration location varies by client:
94
+ Adapt this example to your MCP client's project configuration format, using your application path and preserving other server entries. See [client configuration locations](docs/MCP_SERVERS.md#client-configuration-locations) for guidance:
92
95
 
93
96
  ```json
94
97
  {
@@ -102,175 +105,91 @@ Add this to your MCP client's project configuration. The configuration location
102
105
  }
103
106
  ```
104
107
 
105
- Restart or reconnect your MCP client, then ask it to call `woods_status`. A ready response with non-zero unit counts confirms the path from Rails extraction to the MCP client.
106
-
107
- The Index Server reads the published index from disk. It does not boot Rails or query application records.
108
-
109
- > **Using Docker?** Run Rails commands inside the application container. If Woods is installed only there, launch the Index Server through that container too; a host-side server requires a host Ruby bundle and host-visible index. Follow [Docker setup](docs/DOCKER_SETUP.md).
110
-
111
- The complete walkthrough, including expected output and first questions to ask, is in [Getting started](docs/GETTING_STARTED.md).
112
-
113
- ## Let an agent install it
114
-
115
- Woods is built to be agent-operated, and the fastest path is handing installation to the coding agent that will use it. Claude Code users can install the distributed skills once — they trigger on install, upgrade, configuration, investigation, and diagnosis on their own:
116
-
117
- ```bash
118
- /plugin marketplace add lost-in-the/plugins
119
- /plugin install woods-plugin@lost-in-the-plugins
120
- ```
121
-
122
- With any coding agent (no plugin needed), paste this into a session opened at your Rails app's root:
123
-
124
- ```text
125
- Install or upgrade the woods gem in this Rails application by following
126
- https://github.com/lost-in-the/woods/blob/main/docs/AGENT_SETUP.md.
127
- Structural setup only: add the gem to the development group, run the
128
- installer, extract and validate the index, and register the Index MCP
129
- server for this app. Do not run the generated legacy migration, and do
130
- not add embedding providers, vector databases, Console/live-data access,
131
- or secrets without asking me first. If woods 1.x is already installed,
132
- follow the upgrade runbook in docs/UPGRADING_TO_2.md instead and plan a
133
- clean re-index. Finish by reporting the installed version, files
134
- changed, commands run, and one verified woods_status call through the
135
- registered MCP server.
136
- ```
137
-
138
- The runbook holds the agent to the same guardrails the skills enforce: a version preflight, minimal diffs, and explicit approval before anything beyond the structural index. Prefer doing it by hand? The five-minute setup above is the same procedure as commands.
139
-
140
- ## Choose your path
141
-
142
- | Goal | Start here |
143
- |---|---|
144
- | Install Woods yourself | [Getting started](docs/GETTING_STARTED.md) |
145
- | Ask a coding agent to install Woods safely | [Agent setup runbook](docs/AGENT_SETUP.md) |
146
- | Configure an MCP client, Docker, or HTTP | [MCP servers](docs/MCP_SERVERS.md) |
147
- | Teach an agent how to query Woods effectively | [Agent guide](docs/AGENT_GUIDE.md) |
148
- | Add semantic search with OpenAI or local Ollama | [Retrieval guide](docs/RETRIEVAL_GUIDE.md) |
149
- | Query live Rails data through the optional Console Server | [Console MCP setup and security](docs/CONSOLE_MCP_SETUP.md) |
150
- | Keep the index current automatically while coding | [Watch daemon](docs/WATCH_DAEMON.md) |
151
- | Upgrade an existing 1.x installation | [Upgrade to Woods 2.0](docs/UPGRADING_TO_2.md) |
152
- | Diagnose a failure | [Troubleshooting](docs/TROUBLESHOOTING.md) |
153
-
154
- ## Upgrading from 1.x
155
-
156
- Woods 2.0 is a major release: identifiers, the on-disk layout, the MCP surface, and task failure posture all changed. [Upgrade to Woods 2.0](docs/UPGRADING_TO_2.md) holds the full what-changed table, the step-by-step runbook with backups and rollback, and an agent-operated upgrade prompt.
108
+ Reconnect the client and ask it to call `woods_status`. Confirm the index path and non-zero unit counts. Then use `search` to discover a known class and `lookup` with its identifier **and type** to inspect it.
157
109
 
158
- ## Optional Claude Code workflows
110
+ The Index Server reads the published index without booting Rails or querying application records. See [MCP servers](docs/MCP_SERVERS.md) for client-specific configuration and HTTP transport.
159
111
 
160
- Woods itself is MCP-client and model independent. The separately packaged Woods plugin (install commands under [Let an agent install it](#let-an-agent-install-it)) gives Claude Code five guided skills: setup and upgrade, MCP configuration, index-driven investigation, repository agent enablement, and diagnosis. Other MCP clients do not need it; follow the human or agent runbooks linked above and configure either stdio or Streamable HTTP directly.
112
+ If Woods is installed only inside Docker, launch MCP through that container too. Host-side launch needs a host bundle and a host-visible index; see the [Docker process and path rule](docs/MCP_SERVERS.md#docker-process-and-path-rule).
161
113
 
162
- ## Two servers, two trust boundaries
114
+ ## Retrieval: with or without embeddings
163
115
 
164
- Woods ships two MCP servers. Most users only need the Index Server.
116
+ Exact lookup, pattern search, and graph queries work immediately after extraction. For ranked retrieval through `codebase_retrieve`, choose a mode:
165
117
 
166
- | | Index Server | Console Server |
118
+ | Mode | Setup | What it searches |
167
119
  |---|---|---|
168
- | Purpose | Query pre-extracted code context | Query live Rails models and schema |
169
- | Data source | Files under `tmp/woods/` | A booted Rails process and its database |
170
- | Default tools | 14 | 9 |
171
- | Optional tools | Semantic retrieval activates after embedding; advanced Ruby embeddings can wire more collaborators | `console_sql` and `console_query` raise the total to 11 when explicitly enabled |
172
- | Default posture | Read-only index | Disabled; live-data access requires deliberate setup |
173
-
174
- The 14 Index tools cover health, exact lookup, search, dependency traversal, flow tracing, graph analysis, framework source, change recency, and optional semantic retrieval. The Console Server exposes nine supported model/schema tools by default. Nineteen Tier 2/3 Console schemas (9 Tier 2, 10 Tier 3) and `console_eval` exist as source inventory but do not register in any supported mode.
175
-
176
- See [MCP servers](docs/MCP_SERVERS.md) for the callable tool lists and client configuration.
120
+ | **Lexical** | Set `WOODS_RETRIEVAL_MODE=lexical` in the MCP process environment and restart the server | Published extraction units, ranked by field-aware keyword matching; no provider or embeddings |
121
+ | **Semantic** (default mode) | Configure a local or hosted embedding provider, then run `bin/rails woods:embed` | Embedded code context, ranked by semantic similarity |
177
122
 
178
- ## Optional semantic search
123
+ For the stdio configuration above, add `"env": {"WOODS_RETRIEVAL_MODE": "lexical"}` inside the `woods` server entry to choose lexical mode. Confirm the active retriever with `woods_status`.
179
124
 
180
- Exact search, lookup, graph traversal, and flow tools work after extraction alone. Natural-language retrieval through `codebase_retrieve` also needs embeddings:
125
+ To switch back to semantic retrieval, remove the lexical environment override or set `WOODS_RETRIEVAL_MODE=semantic`, configure the provider and embedding artifacts, then restart the MCP server and verify `woods_status`. Switching to lexical does not delete existing vectors or provider configuration.
181
126
 
182
- ```ruby
183
- # config/initializers/woods.rb
184
- Woods.configure_with_preset(:local)
185
- ```
186
-
187
- The `:local` preset uses SQLite metadata, in-memory vectors persisted under the index, and a local Ollama service. It needs the `sqlite3` gem in the application bundle plus an installed, running Ollama service, but no cloud API key. Pull the default model before the first embed:
188
-
189
- ```bash
190
- ollama pull nomic-embed-text
191
- ```
192
-
193
- MySQL/PostgreSQL applications that do not bundle `sqlite3` can use `:shared_filesystem` for local persisted stores instead. PostgreSQL/OpenAI, Qdrant/OpenAI, and shared-filesystem configurations are documented in the [backend matrix](docs/BACKEND_MATRIX.md) and [configuration reference](docs/CONFIGURATION_REFERENCE.md).
194
-
195
- For dense Ruby source, add `gem "tokenizers", "~> 0.5"` for exact WordPiece token counting. Without it, Woods uses a character estimate that can over-pack some Ollama chunks.
196
-
197
- ```bash
198
- bin/rails woods:embed
199
- ```
200
-
201
- Reconnect the Index Server after the first embed, then check `woods_status` before using `codebase_retrieve`.
127
+ The [lexical guide](docs/RETRIEVAL_GUIDE.md#embedding-free-lexical-retrieval) and [semantic setup](docs/RETRIEVAL_GUIDE.md#configuring-retrieval) cover configuration, ranking, and response budgets. Lexical matching depends on shared vocabulary; semantic mode requires the configured provider and embedding artifacts.
202
128
 
203
129
  ## Keeping the index current
204
130
 
205
- Run a full extraction after installation or broad configuration changes:
206
-
207
- ```bash
208
- bin/rails woods:extract
209
- ```
210
-
211
- For automatic maintenance during development, run the watcher as a dedicated process:
131
+ Run a watcher alongside your development processes:
212
132
 
213
133
  ```bash
214
134
  bin/rails woods:watch
215
135
  ```
216
136
 
217
- Add it to your development process manager so it starts beside Rails:
137
+ It catches up on changes, publishes complete generations, and lets the Index Server refresh on later tool calls. Use a process supervisor for changes that require the watcher to restart. Without a watcher, run `bin/rails woods:incremental` after edits or `bin/rails woods:extract` for a full rebuild.
218
138
 
219
- ```text
220
- # Procfile.dev
221
- web: bin/rails server
222
- woods: bundle exec rake woods:watch
223
- ```
224
-
225
- The watcher catches up changes made while it was stopped, batches new file changes, reloads Rails code when safe, and publishes complete generations atomically. The Index Server notices a new generation on its next tool call and refreshes itself. **After the initial extraction, ordinary code changes need no manual re-extraction or MCP restart.**
139
+ Incremental cost depends on the affected code and relationships; broad changes can cost as much as a full extraction. Semantic embeddings have a separate update step. See [Watch daemon](docs/WATCH_DAEMON.md), [incremental extraction](docs/INCREMENTAL_EXTRACTION.md), and [source freshness](docs/SOURCE_FRESHNESS.md).
226
140
 
227
- Changes to boot-captured state, including dependencies, initializers, database configuration, credentials, or schema, make the watcher exit with status 75 so a process supervisor can restart it cleanly. If semantic retrieval is enabled, the watcher keeps structural context current; run `bin/rails woods:embed_incremental` to update vectors.
228
-
229
- In CI on Rails 8.1, add one step to `config/ci.rb` so the index the gates read matches the commit under test:
141
+ ## Two servers, two trust boundaries
230
142
 
231
- ```ruby
232
- step "Woods: refresh", "bin/rails woods:incremental"
233
- ```
143
+ | | Index Server | Console Server |
144
+ |---|---|---|
145
+ | Purpose | Inspect extracted code context | Query live Rails models and schema |
146
+ | Reads | Published index files | A booted application and its database |
147
+ | Packaged tools | 14; retrieval usable when configured | 9; 11 with embedded read tools enabled |
148
+ | Setup | The workflow above | Optional, disabled by default |
234
149
 
235
- Claude Code users with the Woods plugin can opt into the same refresh from a `PostToolUse` hook, plus a `SessionStart` warning scoped to commit timestamps (it does not see uncommitted edits or an older checkout). Both ship disabled; set `WOODS_HOOKS_ENABLED=1` to turn them on. See [Watch daemon](docs/WATCH_DAEMON.md#hooks-for-agent-sessions).
150
+ Extraction itself boots and eager-loads your application, so its boot-time behavior still runs. Treat the generated index as confidential application source. Enabling hosted embeddings sends the embedded content to that provider. MCP responses also contain application source, which your client may send to its model provider even when Woods uses lexical retrieval or local embeddings.
236
151
 
237
- Without a resident watcher, run `bin/rails woods:incremental` after changes. See [Watch daemon](docs/WATCH_DAEMON.md) for Docker polling, failure behavior, and restart triggers.
152
+ The optional Console Server can access live data. Review its [setup and security model](docs/CONSOLE_MCP_SETUP.md) before enabling it. Report vulnerabilities privately through [SECURITY.md](SECURITY.md).
238
153
 
239
- ## What gets indexed
154
+ ## What the index can and cannot establish
240
155
 
241
- Woods recognizes the Rails application as a connected system, including:
156
+ - **It is a snapshot.** Check freshness against your working tree before relying on it for a change.
157
+ - **Relationships are recorded evidence, not a complete call graph.** Arbitrary method-body constant references are not exhaustively indexed. No recorded dependents does not prove that a class has no callers or is safe to delete.
158
+ - **A traced flow is not proof of execution.** Follow the tool's evidence and limits, and verify behavior in the application when it matters.
242
159
 
243
- - models, concerns, controllers, routes, middleware, and engines;
244
- - services, interactors, commands, jobs, mailers, and scheduled work;
245
- - ERB views, Phlex components, ViewComponents, and navigation edges;
246
- - GraphQL types, mutations, resolvers, and fields;
247
- - policies, serializers, decorators, validators, state machines, and events;
248
- - migrations, database views, factories, tests, configuration, and installed framework source.
160
+ Use Woods to locate and connect evidence, then confirm the relevant source and tests. The [agent guide](docs/AGENT_GUIDE.md) describes this workflow.
249
161
 
250
- Read the [extractor reference](docs/EXTRACTOR_REFERENCE.md) for the complete per-type contract and [internals](docs/INTERNALS.md) for how extraction, storage, retrieval, and MCP fit together.
162
+ ## Let an agent install it
251
163
 
252
- ## Security boundary
164
+ Use the [agent setup runbook](docs/AGENT_SETUP.md) for a copyable installation prompt and verification checklist. Woods works with MCP-capable clients independently of a specific model or editor.
253
165
 
254
- Woods extraction reads application code, resolved Rails configuration, and database schema. Treat the generated index as source code: do not publish it unless the source itself may be published.
166
+ Claude Code users can optionally install the companion workflows:
255
167
 
256
- The optional Console Server has a larger trust boundary because it can read live application data. It is disabled by default and adds table blocking, credential scanning, column redaction, SQL validation, and rolled-back transactions when enabled. Those controls reduce risk; they do not turn production data access into a harmless default. Review [Console MCP security](docs/CONSOLE_MCP_SETUP.md#safety-model) before enabling it.
168
+ ```text
169
+ /plugin marketplace add lost-in-the/plugins
170
+ /plugin install woods-plugin@lost-in-the-plugins
171
+ ```
257
172
 
258
- Report vulnerabilities privately through [SECURITY.md](SECURITY.md).
173
+ The plugin guides installation, MCP configuration, investigation, repository agent setup, and diagnosis. It is distributed separately from the gem.
259
174
 
260
175
  ## Documentation
261
176
 
262
- Use the [documentation index](docs/README.md) to find guides by task or audience. Frequently used references include:
177
+ | Task | Guide |
178
+ |---|---|
179
+ | Install and verify | [Getting started](docs/GETTING_STARTED.md) |
180
+ | Configure clients, Docker, or HTTP | [MCP servers](docs/MCP_SERVERS.md) |
181
+ | Query effectively | [Agent guide](docs/AGENT_GUIDE.md) and [tool cookbook](docs/MCP_TOOL_COOKBOOK.md) |
182
+ | Configure Woods | [Configuration reference](docs/CONFIGURATION_REFERENCE.md) |
183
+ | Choose retrieval and storage | [Retrieval guide](docs/RETRIEVAL_GUIDE.md) and [backend matrix](docs/BACKEND_MATRIX.md) |
184
+ | Upgrade from 1.x | [Upgrade guide](docs/UPGRADING_TO_2.md) |
185
+ | Diagnose a failure | [Troubleshooting](docs/TROUBLESHOOTING.md) |
263
186
 
264
- - [Configuration reference](docs/CONFIGURATION_REFERENCE.md)
265
- - [MCP tool cookbook](docs/MCP_TOOL_COOKBOOK.md)
266
- - [FAQ](docs/FAQ.md)
267
- - [Troubleshooting](docs/TROUBLESHOOTING.md)
268
- - [Upgrade to Woods 2.0](docs/UPGRADING_TO_2.md)
187
+ See the [documentation index](docs/README.md) for all guides and canonical reference pages.
269
188
 
270
189
  ## Contributing
271
190
 
272
- Read [CONTRIBUTING.md](CONTRIBUTING.md) before opening an issue or pull request. Coding agents working in the source repository should also read [AGENTS.md](https://github.com/lost-in-the/woods/blob/main/AGENTS.md).
191
+ Use [GitHub issues](https://github.com/lost-in-the/woods/issues) for bugs and feature requests. Read [CONTRIBUTING.md](CONTRIBUTING.md) before submitting a pull request; coding agents should also read [AGENTS.md](https://github.com/lost-in-the/woods/blob/main/AGENTS.md).
273
192
 
274
193
  ## License
275
194
 
276
- Woods is available under the [MIT License](LICENSE.txt).
195
+ [MIT](LICENSE.txt).
data/SECURITY.md CHANGED
@@ -60,11 +60,12 @@ Woods runs inside your Rails application and has access to:
60
60
  - **Application source code**: extracted and written to the output directory as JSON
61
61
  - **Database schema**: column names, types, indexes, and foreign keys (no row data)
62
62
  - **Git metadata**: commit history, contributors, file change frequency
63
- - **Runtime state** (Console MCP Server only), live database queries within a rolled-back transaction
63
+ - **Optional session traces**: session/trace identifiers, request paths and controller/action timelines when session tracing is configured
64
+ - **Live database rows** (Console MCP Server), queries within a rolled-back transaction
64
65
 
65
66
  ### Output Directory
66
67
 
67
- Extracted data is written to `tmp/woods/` by default. This directory contains your application's source code and schema in structured JSON format. Treat it with the same sensitivity as your source code, do not expose it to untrusted parties.
68
+ Extracted data is written to `tmp/woods/` by default. This directory contains your application's source code and schema in structured JSON format. Source literals, comments and configuration metadata may contain sensitive values. Treat the index with the same sensitivity as your source code; do not expose it to untrusted parties. Optional session stores can be configured under this directory and add request metadata that may be more sensitive than the source itself. Apply the [session tracing access and retention guidance](docs/CONFIGURATION_REFERENCE.md#session-tracer-options) to those stores too.
68
69
 
69
70
  ### Console Server
70
71
 
@@ -80,13 +81,15 @@ If extraction output leaks, what can an attacker do with it?
80
81
 
81
82
  **What the output contains.** Application source code (inlined concerns, callback-resolved behavior), database schema (column names, types, indexes, foreign keys), route tables, migration history, gem versions, and git metadata (commit history, contributor emails, file change frequency).
82
83
 
83
- **What the output does not contain.** No row-level data from your database. Woods extracts schema only. No environment variables, no `Rails.application.credentials`, no API keys, no session state, no request logs, no customer data.
84
+ **Structural extraction boundaries.** Structural extraction collects schema rather than dumping live database rows, and does not intentionally collect environment variables or the Rails encrypted credential store. This is not a guarantee that output is free of secrets or customer information: extracted source can contain hardcoded values, comments or examples, and Index tools can return that source. Console redaction and export-specific scrubbing do not sanitize the structural index.
85
+
86
+ **Optional session data.** When configured, session tracing records session/trace identifiers, request paths and controller/action timelines. Request paths and identifiers can contain sensitive user information. The configured session store controls where these records live; configured Index session tools can return them. Session storage and access need their own retention and authorization controls. Console MCP is a separate live-data boundary described above.
84
87
 
85
88
  | Leak scenario | Attacker gains | Attacker does not gain |
86
89
  |---|---|---|
87
- | `tmp/woods/` directory exfiltrated | Source code + schema equivalent to a git clone + `rails db:schema:dump` | Database rows, secrets, tokens, customer data |
88
- | MCP Index Server token leaked (HTTP transport) | Read-only query access to the extracted index, no write or execution paths | Shell access, database row data, secrets |
90
+ | `tmp/woods/` directory exfiltrated | Source code, schema and metadata, including sensitive values present in source; session records if their store is configured here | No additional live database or shell access merely from possessing these files |
91
+ | MCP Index Server token leaked (HTTP transport) | Access to published source/metadata and configured session tools, including sensitive content they contain | The packaged default does not boot Rails or provide live database queries or shell execution; custom tool wiring has its own access boundary |
89
92
  | Notion sync database compromised | Model and column summaries synced to Notion | Anything not mirrored, source code stays local |
90
- | Console MCP Server exposed (dev/staging) | Read-only database access through a rolled-back transaction, bounded by TableGate + Redactor + SqlValidator | Write access (rolled back), full credentials (redacted), blocked tables |
93
+ | Console MCP Server exposed (dev/staging) | Live database access constrained by TableGate, credential scanning, Redactor and SqlValidator; this remains admin-trust access | Supported packaged modes expose no write or eval tools; rollback and redaction have the limits described in the Console guide |
91
94
 
92
95
  **Mitigation.** Treat `tmp/woods/` as source-equivalent, keep it out of world-readable directories and public container images. Rotate `WOODS_MCP_HTTP_TOKEN` on compromise. Keep `console_mcp_enabled = false` in production regardless of environment, since the console layers are defense-in-depth and not primary controls.
data/docs/AGENT_GUIDE.md CHANGED
@@ -2,6 +2,14 @@
2
2
 
3
3
  This guide is for coding agents using an already connected Woods MCP server. Woods is evidence from the running Rails application and its extracted graph; it complements file search, tests, git history, and direct source inspection.
4
4
 
5
+ Supporting servers also send a concise version of this workflow in MCP
6
+ initialization/discovery instructions, without requiring an installed plugin.
7
+ Check the connected server's version and registered tools; this feature is
8
+ unreleased after `2.0.0.beta2`, and protocol `2024-11-05` omits the field.
9
+ The [initialization contract](MCP_SERVERS.md#initialization-guidance) describes
10
+ availability. This guide remains the detailed reference when instructions are
11
+ absent or the client does not display them.
12
+
5
13
  ## Start every session with status
6
14
 
7
15
  Call `woods_status` before relying on the index. Check:
@@ -13,6 +21,11 @@ Call `woods_status` before relying on the index. Check:
13
21
 
14
22
  If status is unhealthy, report the evidence and ask the owner to extract or refresh. Do not fill gaps by asserting that Woods found nothing.
15
23
 
24
+ Compare `index.woods_version` (last manifest publisher, unknown for older indexes)
25
+ with `server.version` (the MCP reader). A major-version difference warrants a
26
+ full extraction and upgrade review; a match does not prove every retained unit
27
+ was rewritten. See [manifest writer provenance](PUBLISHED_INDEX.md#manifest-writer-provenance).
28
+
16
29
  ## The default query loop
17
30
 
18
31
  Use this four-step loop for most codebase questions:
@@ -103,22 +116,70 @@ fields: ["identifier", "source_code", "metadata"]
103
116
 
104
117
  Start with identifier search. Add source or metadata only when name discovery fails. Restrict types and keep result limits small enough to inspect.
105
118
 
119
+ On supporting versions, read `completeness` before calling search exhaustive.
120
+ `result_count` counts returned rows. `result_limit` proves one more match exists;
121
+ `scan_budget` and `regex_timeout` leave that unknown. Only `exhausted` establishes
122
+ an exact total within the requested index/query domain. Narrow types, literal
123
+ prefix/suffix filters, or deep fields when `partial` is true. A detected artifact
124
+ failure remains an error with unknown completeness, never proof of no matches.
125
+
126
+ This metadata is unreleased after `2.0.0.beta2`; older servers may omit it.
127
+ Do not infer completeness from a full page or missing metadata. See the
128
+ [search response contract](MCP_SERVERS.md#search-completeness).
129
+
106
130
  ## Traverse deliberately
107
131
 
108
- `dependencies` means “what this unit uses.” `dependents` means “what uses this unit.” Both default to bounded breadth-first traversal and accept type or relationship filters.
132
+ `dependencies` shows recorded relationships from this unit; `dependents` shows
133
+ recorded relationships to it. Both default to bounded breadth-first traversal and
134
+ accept type or relationship filters. They are not exhaustive source-reference or
135
+ call graphs: selective scanning can miss arbitrary method-body constant references,
136
+ including generic PORO and library targets. No dependents or test-only dependents
137
+ do not establish absence of production callers. Check source before claiming absence.
138
+ The `graph_coverage` notice makes this scope explicit in supporting responses;
139
+ that metadata is unreleased after `2.0.0.beta3`.
109
140
 
110
141
  Start at depth 1 or 2. A deeper unfiltered traversal can obscure the direct evidence that matters. Common relationship values include associations (`belongs_to`, `has_many`, `has_one`), code references, renders, redirects, form actions, and navigation links.
111
142
 
112
- Both return at most 50 nodes and say so with a `Showing N of M (truncated)`
113
- line. Narrow with `depth`, `types` and `via` before paging with `limit` and
143
+ Both return at most 50 nodes by default and say so with a `Showing N of M (truncated)`
144
+ line for completed walks. Budget-limited answers in supporting versions instead
145
+ say `Showing N of at least M (total unknown: <budget reason>)`. Narrow with `depth`, `types` and `via` before paging with `limit` and
114
146
  `offset`: narrowing answers the question, paging only splits the same answer
115
147
  across turns. In a multi-database app each row names the unit's database.
116
148
 
117
- Use returned relationship labels as evidence. Do not infer call order from a dependency edge alone.
118
-
119
- ## Use semantic retrieval only when ready
120
-
121
- `codebase_retrieve` answers natural-language questions with token-budgeted context. Use it when `woods_status` reports a configured embedding provider and current vector data.
149
+ On supporting servers, inspect `structuredContent.data` for traversal nodes,
150
+ `graph_coverage`, exactness, budgets and optional explanation witnesses, regardless
151
+ of the text renderer. This packaged stdio/HTTP payload is unreleased after
152
+ `2.0.0.beta3`; check the actual response and use its text when structured data is
153
+ absent. There is no traversal `format` argument. See the
154
+ [response contract](MCP_SERVERS.md#dependency-graph-coverage).
155
+
156
+ A traversal can also stop at its independent node or edge budget. Treat
157
+ `partial`/`partial_reason` as incomplete graph evidence even on the final page;
158
+ paging cannot recover nodes the walk never reached. Supporting responses include
159
+ `total_is_exact: false` for a cutoff and true for a finished walk, independently of
160
+ pagination. This field is unreleased after `2.0.0.beta3`; on older servers inspect
161
+ `partial` directly. `nodes_total` remains the root-inclusive admitted prefix count,
162
+ not the full reachable total when partial. Even an exact count covers only the
163
+ requested root, depth, filters and published graph generation. Check the connected schema
164
+ before using `max_nodes`/`max_edges`, and follow the
165
+ [budget contract](MCP_SERVERS.md#dependency-traversal-budgets).
166
+
167
+ When the connected schema supports `explain`, request `explain: true` to see
168
+ recorded source-to-target relationships and a shared shortest witness to each
169
+ row. Report `direct` relationships separately from `transitive` inferred impact.
170
+ Follow `parent`/`edge_id` references; `context: true` ancestors are outside the
171
+ current result page. Unknown labels and ambiguous candidate types stay unknown;
172
+ `typed_path_complete: false` does not establish a uniquely typed path. True means
173
+ only unambiguous witness types, not complete source coverage. Supporting text
174
+ responses label this `witness types unambiguous` (unreleased after `2.0.0.beta3`). See the
175
+ [explanation contract](MCP_SERVERS.md#traversal-explanations).
176
+
177
+ Use recorded relationship labels as evidence. Do not infer execution or call
178
+ order from a dependency edge alone.
179
+
180
+ ## Use ranked retrieval only when ready
181
+
182
+ `codebase_retrieve` answers natural-language questions with token-budgeted context. Use it when `woods_status` reports explicit lexical mode over a current published index, or a configured embedding provider and current vector data in semantic mode. Lexical mode explains matching terms/fields and does not infer synonyms absent from the text; a no-match response is not proof of missing behavior.
122
183
 
123
184
  Important parameters:
124
185
 
@@ -202,3 +263,43 @@ Verification: <source/test/history checked or still needed>
202
263
  - [Extractor reference](EXTRACTOR_REFERENCE.md): indexed unit and edge contracts.
203
264
  - [Retrieval guide](RETRIEVAL_GUIDE.md): embeddings, ranking, and token budgets.
204
265
  - [Troubleshooting](TROUBLESHOOTING.md): stale indexes, disabled retrieval, and startup failures.
266
+
267
+ ### Explicit retrieval and discovery scope
268
+
269
+ On a server whose tool schema advertises them, `packages` and `source_paths` narrow
270
+ `search` and `codebase_retrieve` before candidate limits. These are per-call
271
+ arguments, not configuration settings. Inspect applied scope and completeness;
272
+ a narrow graph query can omit relevant cross-boundary dependencies. See the
273
+ [scope contract](RETRIEVAL_GUIDE.md#explicit-package-and-source-path-scopes) for
274
+ root/nested ownership, path normalization, errors, storage support, and cost.
275
+
276
+ ### Verify source content before relying on freshness
277
+
278
+ When supported by the installed version, inspect `woods_status.index.source_freshness`.
279
+ Repeated edits can leave the porcelain fingerprint unchanged. A quick-budget
280
+ `unknown` can justify one explicit `source_check: "deep"` call; persistent unknown
281
+ needs the reported limitation resolved, not repeated status polling. Use
282
+ `bundle exec woods-extract full` to establish verified preboot source evidence.
283
+ A named refresh does not certify unrelated consumers or external runtime state.
284
+ Follow [source freshness](SOURCE_FRESHNESS.md) and keep ordinary query scopes narrow.
285
+
286
+ ### Recovering relevant code under a small context budget
287
+
288
+ Check the installed schemas before requesting `evidence: 'compact'` on retrieval
289
+ or lookup, or `evidence: 'outline'` for API orientation. These modes preserve
290
+ complete selected spans and explicitly report omissions. An outline is not proof
291
+ of implementation behavior. Follow the returned typed `full_evidence` call when
292
+ you need full source; its SHA guard refuses changed source instead of validating
293
+ a different publication accidentally. Published-unit line/byte ranges can include
294
+ synthesized or commented concern source and are not physical file coordinates.
295
+ See the [evidence contract](RETRIEVAL_GUIDE.md#compact-published-evidence-and-api-outlines).
296
+
297
+ ### Optional hook hints
298
+
299
+ When explicitly enabled on a supporting installed gem, Claude hook context offers
300
+ a small served-generation orientation and post-edit candidate dependents. Treat
301
+ pre-refresh, unknown freshness, truncation and ambiguous identity labels as limits
302
+ on the evidence. Verify direct and inferred downstream candidates using typed
303
+ lookup and `dependents explain:true`; suggested tests do not prove coverage.
304
+ Silence does not establish no impact. See [bounded context hints](WATCH_DAEMON.md#optional-bounded-context-hints)
305
+ for opt-in, independent refresh controls, limits and repeat suppression.