woods 2.0.0.beta2 → 2.0.0.beta4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (233) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +339 -1
  3. data/CONTRIBUTING.md +188 -12
  4. data/README.md +93 -174
  5. data/SECURITY.md +9 -6
  6. data/docs/AGENT_GUIDE.md +109 -8
  7. data/docs/AGENT_SETUP.md +98 -7
  8. data/docs/BACKEND_MATRIX.md +25 -0
  9. data/docs/CLIENT_HOOKS.md +111 -0
  10. data/docs/CONFIGURATION_REFERENCE.md +267 -16
  11. data/docs/CONSOLE_MCP_SETUP.md +80 -7
  12. data/docs/DOCKER_SETUP.md +22 -3
  13. data/docs/EVALUATION.md +464 -1
  14. data/docs/EXTRACTOR_REFERENCE.md +45 -6
  15. data/docs/FAQ.md +11 -12
  16. data/docs/GETTING_STARTED.md +17 -5
  17. data/docs/INCREMENTAL_EXTRACTION.md +147 -7
  18. data/docs/INDEX_LAYOUT.md +382 -0
  19. data/docs/INTERNALS.md +7 -2
  20. data/docs/MCP_SERVERS.md +276 -5
  21. data/docs/MCP_TOOL_COOKBOOK.md +37 -22
  22. data/docs/MCP_WORKTREE_SETUP.md +43 -83
  23. data/docs/NOTION_INTEGRATION.md +13 -0
  24. data/docs/OBSIDIAN_INTEGRATION.md +57 -9
  25. data/docs/PUBLISHED_INDEX.md +72 -0
  26. data/docs/README.md +7 -0
  27. data/docs/RETRIEVAL_GUIDE.md +273 -12
  28. data/docs/RUNTIME_TRACING.md +71 -0
  29. data/docs/SOURCE_FRESHNESS.md +143 -0
  30. data/docs/TROUBLESHOOTING.md +129 -18
  31. data/docs/UNBLOCKED_INTEGRATION.md +25 -0
  32. data/docs/UPGRADING_TO_2.md +48 -22
  33. data/docs/WATCH_DAEMON.md +277 -67
  34. data/exe/woods-agent-config +6 -0
  35. data/exe/woods-extract +5 -0
  36. data/exe/woods-hook-context +6 -0
  37. data/exe/woods-mcp-start +14 -9
  38. data/lib/generators/woods/pgvector_generator.rb +8 -2
  39. data/lib/generators/woods/templates/woods.rb.tt +1 -3
  40. data/lib/tasks/woods.rake +47 -397
  41. data/lib/woods/agent_configuration/applier.rb +135 -0
  42. data/lib/woods/agent_configuration/cli.rb +101 -0
  43. data/lib/woods/agent_configuration/cli_options.rb +29 -0
  44. data/lib/woods/agent_configuration/document.rb +105 -0
  45. data/lib/woods/agent_configuration/error.rb +7 -0
  46. data/lib/woods/agent_configuration/launcher.rb +75 -0
  47. data/lib/woods/agent_configuration/layout.rb +72 -0
  48. data/lib/woods/agent_configuration/managed_section.rb +62 -0
  49. data/lib/woods/agent_configuration/plan.rb +98 -0
  50. data/lib/woods/agent_configuration/plan_diff.rb +38 -0
  51. data/lib/woods/agent_configuration/planned_files.rb +61 -0
  52. data/lib/woods/agent_configuration/planner.rb +63 -0
  53. data/lib/woods/agent_configuration/planner_validation.rb +77 -0
  54. data/lib/woods/agent_configuration/preflight.rb +100 -0
  55. data/lib/woods/agent_configuration/recovery.rb +49 -0
  56. data/lib/woods/ast/node.rb +2 -0
  57. data/lib/woods/ast/parser.rb +38 -5
  58. data/lib/woods/builder.rb +21 -5
  59. data/lib/woods/cache/cache_middleware.rb +28 -7
  60. data/lib/woods/cache/cache_store.rb +4 -5
  61. data/lib/woods/change_set.rb +5 -4
  62. data/lib/woods/console/credential_index.rb +20 -2
  63. data/lib/woods/console/credential_scanner.rb +18 -17
  64. data/lib/woods/console/credential_scanner_registry.rb +36 -0
  65. data/lib/woods/console/dispatch_pipeline.rb +7 -0
  66. data/lib/woods/console/embedded_executor.rb +32 -10
  67. data/lib/woods/console/encrypted_credential_snapshot.rb +16 -0
  68. data/lib/woods/console/rack_middleware.rb +22 -13
  69. data/lib/woods/console/server.rb +18 -16
  70. data/lib/woods/console/sql_noise_stripper.rb +9 -7
  71. data/lib/woods/console/sql_table_scanner.rb +47 -7
  72. data/lib/woods/console/sql_validator.rb +49 -9
  73. data/lib/woods/console/sqlite_read_guard.rb +46 -0
  74. data/lib/woods/coordination/pipeline_lock.rb +3 -2
  75. data/lib/woods/dependency_graph.rb +65 -13
  76. data/lib/woods/embedding/corpus.rb +94 -0
  77. data/lib/woods/embedding/indexer.rb +114 -60
  78. data/lib/woods/embedding/openai.rb +17 -6
  79. data/lib/woods/evaluation/ablation_executor.rb +6 -1
  80. data/lib/woods/evaluation/ablation_timed_executor.rb +22 -4
  81. data/lib/woods/export/typed_reader.rb +56 -0
  82. data/lib/woods/extractor.rb +277 -149
  83. data/lib/woods/extractors/action_cable_extractor.rb +3 -1
  84. data/lib/woods/extractors/behavioral_profile.rb +9 -7
  85. data/lib/woods/extractors/caching_extractor.rb +3 -1
  86. data/lib/woods/extractors/concern_extractor.rb +64 -6
  87. data/lib/woods/extractors/configuration_extractor.rb +7 -3
  88. data/lib/woods/extractors/controller_extractor.rb +13 -4
  89. data/lib/woods/extractors/database_view_extractor.rb +3 -1
  90. data/lib/woods/extractors/declared_parent.rb +55 -0
  91. data/lib/woods/extractors/decorator_extractor.rb +3 -1
  92. data/lib/woods/extractors/engine_extractor.rb +3 -1
  93. data/lib/woods/extractors/event_extractor.rb +4 -2
  94. data/lib/woods/extractors/factory_extractor.rb +3 -1
  95. data/lib/woods/extractors/graphql_extractor.rb +10 -13
  96. data/lib/woods/extractors/i18n_extractor.rb +3 -1
  97. data/lib/woods/extractors/job_extractor.rb +6 -19
  98. data/lib/woods/extractors/lib_extractor.rb +13 -9
  99. data/lib/woods/extractors/mailer_extractor.rb +26 -15
  100. data/lib/woods/extractors/manager_extractor.rb +3 -1
  101. data/lib/woods/extractors/method_parameters.rb +53 -0
  102. data/lib/woods/extractors/middleware_argument.rb +65 -0
  103. data/lib/woods/extractors/middleware_extractor.rb +9 -3
  104. data/lib/woods/extractors/migration_extractor.rb +3 -1
  105. data/lib/woods/extractors/model_extractor.rb +26 -34
  106. data/lib/woods/extractors/package_extractor.rb +24 -4
  107. data/lib/woods/extractors/phlex_extractor.rb +3 -1
  108. data/lib/woods/extractors/policy_extractor.rb +3 -1
  109. data/lib/woods/extractors/poro_extractor.rb +13 -9
  110. data/lib/woods/extractors/pundit_extractor.rb +3 -1
  111. data/lib/woods/extractors/rails_source_extractor.rb +4 -2
  112. data/lib/woods/extractors/rake_task_extractor.rb +4 -2
  113. data/lib/woods/extractors/route_extractor.rb +3 -1
  114. data/lib/woods/extractors/route_helper_resolver.rb +10 -33
  115. data/lib/woods/extractors/scheduled_job_extractor.rb +41 -15
  116. data/lib/woods/extractors/serializer_extractor.rb +4 -2
  117. data/lib/woods/extractors/service_extractor.rb +3 -1
  118. data/lib/woods/extractors/shared_dependency_scanner.rb +2 -2
  119. data/lib/woods/extractors/shared_utility_methods.rb +48 -19
  120. data/lib/woods/extractors/source_nesting.rb +1 -1
  121. data/lib/woods/extractors/state_machine_extractor.rb +3 -1
  122. data/lib/woods/extractors/test_mapping_extractor.rb +3 -1
  123. data/lib/woods/extractors/validator_extractor.rb +3 -1
  124. data/lib/woods/extractors/view_component_extractor.rb +3 -1
  125. data/lib/woods/extractors/view_template_extractor.rb +3 -1
  126. data/lib/woods/gem_mapper.rb +2 -0
  127. data/lib/woods/git_history.rb +116 -0
  128. data/lib/woods/graph_analyzer.rb +35 -6
  129. data/lib/woods/hooks/context_cli.rb +54 -0
  130. data/lib/woods/hooks/context_event.rb +88 -0
  131. data/lib/woods/hooks/context_hint.rb +73 -0
  132. data/lib/woods/hooks/context_impact.rb +77 -0
  133. data/lib/woods/hooks/context_output.rb +47 -0
  134. data/lib/woods/hooks/context_state.rb +102 -0
  135. data/lib/woods/hooks/refresh.rb +79 -0
  136. data/lib/woods/hooks/rule_projection.rb +78 -0
  137. data/lib/woods/input_rules.rb +19 -0
  138. data/lib/woods/mcp/bearer_auth.rb +22 -13
  139. data/lib/woods/mcp/bootstrapper.rb +79 -4
  140. data/lib/woods/mcp/config_resolver.rb +2 -1
  141. data/lib/woods/mcp/index_reader.rb +334 -162
  142. data/lib/woods/mcp/initialization_guidance.rb +27 -0
  143. data/lib/woods/mcp/origin_guard.rb +17 -9
  144. data/lib/woods/mcp/published_lexical_retriever.rb +115 -0
  145. data/lib/woods/mcp/renderers/markdown_renderer.rb +22 -9
  146. data/lib/woods/mcp/renderers/plain_renderer.rb +18 -8
  147. data/lib/woods/mcp/search_results.rb +74 -0
  148. data/lib/woods/mcp/server.rb +178 -63
  149. data/lib/woods/mcp/tool_contract.rb +3 -1
  150. data/lib/woods/mcp/tool_response_renderer.rb +41 -0
  151. data/lib/woods/mcp/traversal_evidence.rb +113 -0
  152. data/lib/woods/mcp/traversal_evidence_index.rb +100 -0
  153. data/lib/woods/mcp/traversal_evidence_page.rb +41 -0
  154. data/lib/woods/mcp/traversal_evidence_text.rb +52 -0
  155. data/lib/woods/mcp/traversal_response.rb +22 -0
  156. data/lib/woods/notion/exporter.rb +56 -17
  157. data/lib/woods/obsidian/destination_plan.rb +98 -0
  158. data/lib/woods/obsidian/name_mapper.rb +19 -3
  159. data/lib/woods/obsidian/note_builder.rb +19 -10
  160. data/lib/woods/obsidian/vault_exporter.rb +88 -32
  161. data/lib/woods/operator/pipeline_guard.rb +18 -13
  162. data/lib/woods/path_dispatcher.rb +13 -6
  163. data/lib/woods/payload_store.rb +27 -26
  164. data/lib/woods/published_index/typed_unit_reader.rb +40 -3
  165. data/lib/woods/published_index.rb +2 -2
  166. data/lib/woods/railtie.rb +3 -3
  167. data/lib/woods/railtie_support.rb +12 -12
  168. data/lib/woods/rake_helpers.rb +382 -0
  169. data/lib/woods/resilience/graph_invariant_validator/membership_checks.rb +71 -0
  170. data/lib/woods/resilience/graph_invariant_validator/node_checks.rb +61 -0
  171. data/lib/woods/resilience/graph_invariant_validator/reverse_relationship_checks.rb +46 -0
  172. data/lib/woods/resilience/graph_invariant_validator.rb +119 -0
  173. data/lib/woods/resilience/index_validator/graph_checks.rb +80 -0
  174. data/lib/woods/resilience/index_validator.rb +112 -23
  175. data/lib/woods/retrieval/context_assembler.rb +50 -15
  176. data/lib/woods/retrieval/lexical_assembler.rb +84 -0
  177. data/lib/woods/retrieval/lexical_index.rb +120 -0
  178. data/lib/woods/retrieval/ranker.rb +4 -2
  179. data/lib/woods/retrieval/scope.rb +108 -0
  180. data/lib/woods/retrieval/scoped_graph_store.rb +32 -0
  181. data/lib/woods/retrieval/scoped_vector_store.rb +55 -0
  182. data/lib/woods/retrieval/search_executor.rb +86 -27
  183. data/lib/woods/retrieval/source_evidence.rb +200 -0
  184. data/lib/woods/retriever.rb +98 -22
  185. data/lib/woods/ruby_analyzer/trace_enricher.rb +77 -38
  186. data/lib/woods/session_tracer/file_store.rb +6 -1
  187. data/lib/woods/session_tracer/middleware.rb +10 -12
  188. data/lib/woods/session_tracer/redis_store.rb +22 -6
  189. data/lib/woods/session_tracer/session_flow_assembler.rb +23 -17
  190. data/lib/woods/session_tracer/solid_cache_coordination.rb +6 -4
  191. data/lib/woods/session_tracer/unit_resolver.rb +63 -0
  192. data/lib/woods/source_inputs/consumer_errors.rb +31 -0
  193. data/lib/woods/source_inputs/handoff.rb +102 -0
  194. data/lib/woods/source_inputs/launcher.rb +157 -0
  195. data/lib/woods/source_inputs/manifest.rb +124 -0
  196. data/lib/woods/source_inputs/private_key.rb +55 -0
  197. data/lib/woods/source_inputs/scanner.rb +171 -0
  198. data/lib/woods/source_inputs/scopes.rb +71 -0
  199. data/lib/woods/source_inputs/session.rb +214 -0
  200. data/lib/woods/source_inputs/status.rb +84 -0
  201. data/lib/woods/source_inputs/verifier.rb +107 -0
  202. data/lib/woods/storage/metadata_store.rb +25 -25
  203. data/lib/woods/storage/pgvector.rb +35 -10
  204. data/lib/woods/storage/qdrant.rb +17 -7
  205. data/lib/woods/storage/vector_store.rb +18 -6
  206. data/lib/woods/tasks.rb +3 -2
  207. data/lib/woods/temporal/json_snapshot_store.rb +58 -9
  208. data/lib/woods/unblocked/exporter.rb +59 -70
  209. data/lib/woods/version.rb +1 -1
  210. data/lib/woods/watch/boot_snapshot.rb +52 -0
  211. data/lib/woods/watch/daemon.rb +154 -32
  212. data/lib/woods/watch/listen_watcher.rb +4 -0
  213. data/lib/woods/watch/polling_watcher.rb +5 -1
  214. data/lib/woods/watch/status.rb +20 -15
  215. data/lib/woods/watch/tree_scan.rb +21 -13
  216. data/lib/woods/watch/watcher.rb +4 -1
  217. data/lib/woods.rb +50 -11
  218. data/plugin/.claude-plugin/plugin.json +1 -1
  219. data/plugin/hooks/adapters/normalize.jq +15 -0
  220. data/plugin/hooks/adapters/normalize.rb +63 -0
  221. data/plugin/hooks/hooks.json +20 -0
  222. data/plugin/hooks/woods-context.sh +50 -0
  223. data/plugin/hooks/woods-input-rules.sh +159 -0
  224. data/plugin/hooks/woods-opencode.mjs +65 -0
  225. data/plugin/hooks/woods-post-edit.sh +2 -225
  226. data/plugin/hooks/woods-refresh.sh +260 -0
  227. data/plugin/hooks/woods-session-start.sh +47 -55
  228. data/plugin/skills/woods-agent-enable/SKILL.md +19 -0
  229. data/plugin/skills/woods-diagnose/SKILL.md +319 -1
  230. data/plugin/skills/woods-investigate/SKILL.md +145 -0
  231. data/plugin/skills/woods-mcp-config/SKILL.md +90 -2
  232. data/plugin/skills/woods-setup/SKILL.md +110 -6
  233. metadata +87 -5
@@ -1,127 +1,87 @@
1
1
  # MCP Registration in Git Worktrees
2
2
 
3
- > **Claude Code specific.** This page covers Claude Code's MCP registration model (`/mcp`, `~/.claude/plugins/`); other MCP clients manage per-directory registration their own way.
3
+ > **Claude Code specific.** Other MCP clients manage registration and subagent tool access differently. Use the client's current documentation alongside the Woods [MCP setup guide](MCP_SERVERS.md).
4
4
 
5
- When you work in a git worktree, a separate directory checked out from the same repository, your MCP tools may not be available to subagents running in that directory. This page explains why, how to fix it, and how to confirm the registration took effect.
5
+ A git worktree is a separate checkout. For Woods, verify both which servers a session can access and which checkout's index those servers serve.
6
6
 
7
- ## Why Worktree Subagents May Not See Woods Tools
7
+ ## Separate sessions and inherited subagents
8
8
 
9
- MCP server registration in Claude Code is controlled by `.mcp.json` files. Claude Code discovers these files by walking up the directory tree from the working directory. It stops at the first `.mcp.json` it finds (or at `~/.claude/settings.json` for global registrations).
9
+ A **new Claude Code session launched in a worktree** can have different MCP registrations from a session in the main checkout:
10
10
 
11
- A git worktree has its own root directory separate from the main repository checkout. When a subagent starts inside the worktree root, it walks up from that path, not from the main repository root. Unless a `.mcp.json` exists inside the worktree directory tree or in an ancestor shared with both checkouts, the subagent sees no MCP servers.
11
+ | Registration scope | Location | Availability |
12
+ |---|---|---|
13
+ | Local | Project entry in `~/.claude.json` | The project where the server was registered |
14
+ | Project | `.mcp.json` in the project root | That project, subject to client approval/settings |
15
+ | User | `mcpServers` in `~/.claude.json` | Across the user's projects |
12
16
 
13
- Example directory layout:
17
+ A tracked `.mcp.json` may already be present in another worktree; a local, untracked file will not be copied by Git. User-level registration can make a server available across checkouts, but its configured paths still determine which index it serves.
14
18
 
15
- ```
16
- ~/work/my-app/ ← main checkout, has .mcp.json here
17
- ~/work/my-app-feature/ ← worktree, no .mcp.json, so MCP tools are missing
18
- ```
19
+ An **inherited subagent** receives the parent conversation's MCP tools, subject to tool restrictions. Changing the subagent's working directory does not by itself launch another Woods server or switch the served index. Distinguish this from starting an independent client process in that directory.
20
+
21
+ See Claude Code's [MCP installation scopes](https://code.claude.com/docs/en/mcp#mcp-installation-scopes) and [subagent tool access](https://code.claude.com/docs/en/sub-agents#available-tools). Do not diagnose registration by assuming the client searches ancestor directories for the nearest `.mcp.json`.
19
22
 
20
- A subagent spawned in `~/work/my-app-feature/` will not find the `.mcp.json` from `~/work/my-app/`.
23
+ ## Configure a server for the intended checkout
21
24
 
22
- ## Fix: Add a `.mcp.json` to the Worktree Root
25
+ First check `/mcp` in the session that needs Woods. If a suitable server is already connected, verify its index before adding a duplicate registration.
23
26
 
24
- Create a `.mcp.json` in the worktree's root directory with the same woods server entries you use in the main checkout:
27
+ For a host with the application's Ruby bundle installed, a project-root `.mcp.json` can select the bundle and index explicitly:
25
28
 
26
29
  ```json
27
30
  {
28
31
  "mcpServers": {
29
32
  "woods": {
30
- "command": "woods-mcp-start",
31
- "args": ["./tmp/woods"]
32
- },
33
- "woods-console": {
34
- "command": "docker",
35
- "args": [
36
- "compose", "exec", "-T", "app",
37
- "bundle", "exec", "rake", "woods:console"
38
- ],
39
- "cwd": "/absolute/host/path/to/worktree"
33
+ "command": "bundle",
34
+ "args": ["exec", "woods-mcp-start", "/absolute/path/to/worktree/tmp/woods"],
35
+ "env": {
36
+ "BUNDLE_GEMFILE": "/absolute/path/to/worktree/Gemfile"
37
+ }
40
38
  }
41
39
  }
42
40
  }
43
41
  ```
44
42
 
45
- Adjust the paths and arguments to match your project's setup. In particular:
46
-
47
- - `./tmp/woods` is a relative path, it resolves against the worktree root, which is correct if extraction output is written into each worktree separately.
48
- - If you share a single extraction output directory between checkouts, use the absolute path to the shared output: `"/absolute/path/to/main-checkout/tmp/woods"`.
49
- - For Docker projects, the `docker compose exec` command works the same from any host path.
50
-
51
- ## How Plugin Discovery Works
52
-
53
- Claude Code also discovers MCP servers registered in plugin manifests. A plugin at `~/.claude/plugins/<plugin-name>/admin-tools/.mcp.json` is loaded globally, its servers are available in any session regardless of working directory.
43
+ Replace both paths with the intended Rails worktree. Preserve other server entries. These absolute paths are machine-specific; do not commit them as shared team defaults. Follow the client's approval/reconnection steps, then verify the connection below.
54
44
 
55
- If your team distributes the woods MCP registration through a shared plugin, the worktree problem does not apply. Check whether woods is already registered this way:
45
+ For Docker, use the [container-first setup](DOCKER_SETUP.md#default-run-it-through-the-application-container). Select the intended Compose project and service, and verify that its mounted application directory is the worktree you mean to index. `docker compose exec` targets an existing container; running the command from a different checkout does not change that container's mounts. The index path must be visible inside that container.
56
46
 
57
- ```bash
58
- ls ~/.claude/plugins/
59
- # look for a directory containing admin-tools/.mcp.json
60
- cat ~/.claude/plugins/<plugin-name>/admin-tools/.mcp.json
61
- ```
62
-
63
- If you find the woods servers registered there, subagents in any worktree will have access automatically, you do not need a per-worktree `.mcp.json`.
64
-
65
- ## Verifying MCP Registration for a Subagent
47
+ Keep a separate output directory for each checkout when you need checkout-specific answers. If you intentionally point at another checkout's index, Woods serves that published content; registration alone does not make it describe the current worktree.
66
48
 
67
- ### Option 1: List tools from a claude session in the worktree
49
+ ## Plugin-provided servers
68
50
 
69
- Open a new Claude Code session with the worktree as the working directory and run:
51
+ An enabled plugin can supply MCP server definitions, but installing a plugin does not establish that a Woods server is connected to the intended checkout. Availability depends on the plugin's configuration, enabled scope, and the client's tool restrictions. Check `/mcp` rather than assuming a fixed path under `~/.claude/plugins/` or global availability.
70
52
 
71
- ```
72
- /mcp
73
- ```
53
+ The Woods setup/configuration skills help configure a server; their presence alone is not verification that the Index Server is running. See [agent setup](AGENT_SETUP.md).
74
54
 
75
- This lists all connected MCP servers and their tools. If `woods` and/or `woods-console` appear, registration is working.
55
+ ## Verify registration and the served index
76
56
 
77
- ### Option 2: Check via the woods status tool
57
+ 1. Open `/mcp` in the relevant Claude Code session and inspect the Woods connection and available tools.
58
+ 2. Call the Index Server's `woods_status`. Check `index_dir`, generation, and available freshness/provenance information against the intended extraction. A successful connection to the wrong index is still the wrong setup.
59
+ 3. Use `search` and typed `lookup` for a known unit from that checkout. When checking a subagent, verify that it can call the inherited tools and is using the same intended index.
78
60
 
79
- Ask Claude to call the status tool:
80
-
81
- ```
82
- Use woods-console console_status to check what models are available.
83
- ```
84
-
85
- If the tool runs successfully, MCP is registered and the console server is reachable.
86
-
87
- ### Option 3: Inspect the MCP config that Claude Code loaded
88
-
89
- From the worktree directory, run:
90
-
91
- ```bash
92
- cat .mcp.json
93
- ```
94
-
95
- If the file exists and contains the woods entries, Claude Code will use it. If the file is missing, check parent directories up to your home directory for any `.mcp.json` that would be discovered.
61
+ The optional Console Server is a separate connection to a booted Rails application. If deliberately enabled, check it with `console_status` and verify its application/container separately. Console access is not required to verify the Index Server.
96
62
 
97
63
  ## Extraction Provenance in Worktrees (`git_branch` / `git_sha`)
98
64
 
99
- `manifest.json` records the `git_branch` and `git_sha` the extraction ran against. In a linked worktree, `.git` is a **file** containing a `gitdir:` pointer to the real git directory, frequently an absolute host path, rather than a `.git` directory.
65
+ The published payload's `manifest.json` records the extraction's `git_branch` and `git_sha`. In a linked worktree, `.git` is a file containing a `gitdir:` pointer, often to an absolute host path. That private worktree directory also refers to the parent repository's shared Git data.
100
66
 
101
- Woods resolves provenance with worktree-aware git plumbing (`git -C <root> rev-parse`), so an ordinary worktree reports the correct branch and SHA. When a `.git` is present but the pointed-to git directory **cannot be resolved**: most commonly a worktree extracted inside a container where the host path isn't mounted. Woods records `git_branch: "unknown"` / `git_sha: "unknown"` rather than a stale, misleading value: a baked `GIT_BRANCH`/`GIT_SHA` build arg is **not** trusted here (it could be stale). The env vars are honored only when there is no `.git` at the root at all (a non-repo checkout, e.g. a Docker `COPY` that excludes `.git`) or git is unavailable.
67
+ Woods uses worktree-aware Git commands. If a present `.git` cannot be resolved, provenance is `"unknown"`; stale `GIT_BRANCH`/`GIT_SHA` values are not substituted. Those environment variables are fallbacks only when the root has no `.git` or Git is unavailable. Temporal snapshots skip an unknown SHA.
102
68
 
103
- To get correct provenance from a containerized worktree, mount the directory the `gitdir:` pointer references (the parent repository's `.git`) into the container so git can resolve it. Temporal snapshots skip an `"unknown"` SHA, so misleading provenance never keys a snapshot.
69
+ For extraction in a container, make the canonical Git directory and the worktree's pointer resolvable there. Mounting only the private worktree Git directory can leave its shared object store unreachable. Follow the [Git provenance troubleshooting guide](TROUBLESHOOTING.md) for mount and `WOODS_GIT_DIR` guidance, and the [published index layout](INDEX_LAYOUT.md) when locating the manifest.
104
70
 
105
71
  ## Troubleshooting
106
72
 
107
- **"Unknown tool" or "no MCP server named woods"**
108
-
109
- The woods MCP server is not registered for this session. Add a `.mcp.json` to the worktree root as shown above, then restart the session.
73
+ **Woods or an expected tool is absent**
110
74
 
111
- **Console server starts but returns "unsupported: Not yet implemented in embedded mode" for console_sql / console_query**
75
+ Check `/mcp` for connection errors, enabled registration scopes, and project approvals. For a subagent, also check its tool restrictions. Compare the requested tool with the [supported tool surface](MCP_SERVERS.md#conditional-index-capabilities); some tools require optional collaborators and are not registered by the packaged server. Do not add duplicate registrations before establishing which case applies.
112
76
 
113
- The console server is registered and reachable, but `embedded_read_tools` is disabled (the default). To enable console_sql and console_query in embedded mode, see the [Console MCP Setup guide](CONSOLE_MCP_SETUP.md), specifically the `embedded_read_tools: true` option for the Rack middleware.
77
+ **The server connects but shows another checkout**
114
78
 
115
- **Extraction output is empty or stale in the worktree**
79
+ Inspect its configured bundle, index path, Compose project/service, and container mounts. Re-extract from the intended Rails checkout into its own output directory, then reconnect to that index and repeat `woods_status`. See [Docker setup](DOCKER_SETUP.md).
116
80
 
117
- Extraction writes to `tmp/woods/` relative to the Rails application root (inside the container). If the worktree's volume mount points to a different host path than the main checkout, run extraction again from the worktree:
118
-
119
- ```bash
120
- docker compose exec app bundle exec rake woods:extract
121
- ```
81
+ **Two registrations use the same server name**
122
82
 
123
- See [DOCKER_SETUP.md](DOCKER_SETUP.md) for the full Docker workflow.
83
+ Check the client's [scope precedence](https://code.claude.com/docs/en/mcp#scope-hierarchy-and-precedence) and the effective connection shown by `/mcp`. Do not assume the two definitions merge or that a project file overrides every other scope.
124
84
 
125
- **Worktree `.mcp.json` conflicts with main checkout `.mcp.json`**
85
+ **Console SQL/query tools are absent or unsupported**
126
86
 
127
- Each file is independent. Claude Code loads the one closest to the working directory. There is no inheritance or merging between them. Keep both files in sync manually, or move the shared configuration into a global plugin manifest.
87
+ Console read-tool availability is configured separately from registration. Follow [Console MCP setup](CONSOLE_MCP_SETUP.md) for the supported 9-tool default and optional 11-tool mode.
@@ -34,6 +34,19 @@ Sync is incremental. A **sync manifest** (`<output_dir>/notion_sync_manifest.jso
34
34
 
35
35
  Manifest entries for models/columns that vanished from the current extraction are pruned, but **no Notion page is ever deleted** by the sync, there is no delete path. A renamed or removed model just leaves its old page in Notion untouched.
36
36
 
37
+ Models and migration dates are loaded by `(identifier, type)`. A same-named
38
+ PORO, library unit, or other extracted type cannot supply a model's table or
39
+ column payload. Public identifiers, page titles, and manifest keys are unchanged.
40
+ If a listed model or migration cannot be read with its exact identity, the
41
+ affected sync refuses before mapping pages or pruning its manifest. Preserve
42
+ the error, validate the index, and regenerate it before retrying; force sync
43
+ does not bypass identity checks. Each public sync method, including standalone
44
+ model or column sync, uses one validated generation snapshot through mapping
45
+ and API calls. Native readers validate the published index before selecting
46
+ model/migration payloads. Custom readers must provide complete published
47
+ enumeration or complete per-bucket listings plus
48
+ `find_unit(identifier, type:)` returning the requested identity.
49
+
37
50
  If the manifest is missing (first run, or a CI cache miss), the exporter falls back to the full lookup/create path for every page and rebuilds the manifest, correct, just more API calls than a steady-state run.
38
51
 
39
52
  ### Escape hatch
@@ -8,7 +8,7 @@ ways at once:
8
8
  Obsidian [Bases](https://help.obsidian.md/bases) table, and drill into a single unit's note with its
9
9
  dependencies and dependents as clickable wikilinks.
10
10
  - **By agents**: load the entire dependency topology from a single `_woods/` sidecar (one read,
11
- no per-note fan-out), with a stable `id → note path` manifest for navigation.
11
+ no per-note fan-out), with a stable typed-unit → note-path manifest for navigation.
12
12
 
13
13
  Unlike the Notion and Unblocked exporters, this one writes **local files only**: there is no API
14
14
  token, no network call, and no rate limit. An Obsidian vault is just a folder.
@@ -95,6 +95,33 @@ Wikilinks are path-qualified with an alias (`[[models/Account|Account]]`): the t
95
95
  sanitized vault path so the link always resolves, and the alias shows the original identifier. The
96
96
  note's `# H1` carries the clean identifier so the sanitized filename never shows as the title.
97
97
 
98
+ ## Sidecar identity and schema versions
99
+
100
+ A unit keeps its original `id` and `type` in note frontmatter. Different types may
101
+ share an identifier: a database view and a factory called `reports` export to
102
+ `database_views/reports.md` and `factories/reports.md`. Existing unambiguous note
103
+ paths and display aliases stay unchanged.
104
+
105
+ Check `schema_version` before consuming `_woods/manifest.json`:
106
+
107
+ - **Version 1:** the graph has no cross-type identifier collisions. The existing
108
+ `notes[id]` and `paths[path]` maps are unchanged.
109
+ - **Version 2:** `notes[id]` retains the graph's primary type (when exported), and
110
+ `variants` lists each additional exported unit as
111
+ `{ "identifier": "reports", "type": "factory", "path": "factories/reports.md" }`.
112
+ Combine primary entries and variants using **`(identifier, type)`** as identity.
113
+ `paths[path]` still gives the original identifier, so several paths may return
114
+ the same value. Resolve a path against both collections; `notes[id]` alone is
115
+ incomplete. A consumer supporting only version 1 must refuse version 2.
116
+
117
+ The exporter reads each variant's own outgoing graph edges and derives incoming
118
+ links from them. Persisted targets are bare identifiers: if a target names several
119
+ types, the exporter omits that relationship from note links and counts it in a
120
+ progress diagnostic. It does this even when one sibling is excluded or unreadable.
121
+ Human association links use the same ambiguity rule. Bare PageRank and graph-analysis
122
+ annotations are omitted for ambiguous identifiers, rather than assigned to either
123
+ type. The verbatim graph sidecar retains all original records for inspection.
124
+
98
125
  ## The three visualizer surfaces
99
126
 
100
127
  | Surface | Best for | Notes |
@@ -114,7 +141,7 @@ credential):
114
141
  | `WOODS_OUTPUT` | `config.output_dir` (`tmp/woods`) | extraction directory to read from |
115
142
  | `WOODS_OBSIDIAN_VAULT` | `<output>/obsidian_vault` | where to write the vault |
116
143
  | `WOODS_OBSIDIAN_INCLUDE_SOURCE` | off | embed each unit's source code (credential-scrubbed) in its note |
117
- | `WOODS_OBSIDIAN_INCLUDE_FRAMEWORK` | off | include `rails_source` units (large; off by default) |
144
+ | `WOODS_OBSIDIAN_INCLUDE_FRAMEWORK` | off | include `rails_source` and readable `gem_source` units (large; off by default) |
118
145
  | `WOODS_OBSIDIAN_FORCE_PURGE` | off | bypass the 30% mass-deletion guard during the stale-note sweep |
119
146
 
120
147
  ```bash
@@ -141,6 +168,28 @@ The exporter **fully regenerates** the vault on every run (no incremental manife
141
168
  cheap). Output is deterministic: re-running against an unchanged extraction produces byte-identical
142
169
  notes, so unchanged units never show up in a git diff.
143
170
 
171
+ Before writing, Woods renders the output and checks **every destination**. An existing note or index
172
+ must carry `woods_managed: true`; a `.woods-vault` sentinel alone never authorizes replacing your
173
+ notes or settings. Conflicts stop the export before any writes or sweep, appear in `errors`, and make
174
+ `woods:obsidian` exit nonzero. `WOODS_OBSIDIAN_FORCE_PURGE` does not bypass ownership checks.
175
+ Child symlinks, directories in place of files, and special files are refused.
176
+
177
+ Generated machine assets are tracked by SHA-256 in `_woods/ownership.json`: the three sidecar JSON
178
+ files, three `.obsidian/` JSON settings files, and `Units.base`. Existing assets must match their
179
+ recorded bytes or the newly generated bytes. This receipt does not change the public sidecar manifest
180
+ schema. Modified settings are preserved by refusing the export; the sentinel is not an override.
181
+
182
+ **Upgrading an older vault:** without a receipt, byte-identical generated assets can be adopted. Run
183
+ once against the same extraction to establish ownership before updating the index. If the index has
184
+ already changed, Woods may refuse the old sidecars because their ownership cannot be proved. Inspect
185
+ and back up the named files before moving them aside, or export into a new directory. Do not remove
186
+ personal content or manufacture a receipt to bypass a conflict.
187
+
188
+ Preflight prevents known destination conflicts from causing partial exports. This is not a multi-file
189
+ transaction or a lock against concurrent editors: a later I/O failure can leave some generated files
190
+ updated. Such failures report errors, suppress the sweep, and return zero completed-export counts;
191
+ inspect the vault before retrying. Avoid editing the destination while an export runs.
192
+
144
193
  Notes Woods manages carry `woods_managed: true` in their frontmatter. On each run, after all notes are
145
194
  written successfully, a **sweep** removes managed notes whose unit no longer exists, so deletions in
146
195
  your code propagate. Several guards make the sweep safe to point at a real vault:
@@ -150,21 +199,20 @@ your code propagate. Several guards make the sweep safe to point at a real vault
150
199
  - It refuses to delete more than 30% of managed notes at once (the signature of a partial extraction)
151
200
  unless `WOODS_OBSIDIAN_FORCE_PURGE=1` is set.
152
201
  - It resolves symlinks and confirms every deletion target is inside the vault root.
153
- - It is skipped entirely if any note failed to write that run (a stale note is harmless; a deleted
154
- reviewed note is not).
202
+ - It is skipped entirely if any candidate unit is unreadable or returns a mismatched identity,
203
+ or any note failed to write. `force_purge` does not bypass this incomplete-export guard.
155
204
 
156
- The same ownership check guards the `.obsidian/` config: Woods will **not** overwrite an existing,
157
- foreign `.obsidian/` folder, it leaves your Obsidian settings untouched and warns instead.
205
+ In a foreign vault, Woods leaves `.obsidian/` configuration untouched and skips the sweep. Notes and
206
+ sidecars are added only after the same per-file ownership preflight. In a Woods-owned vault, existing
207
+ configuration assets still require a matching ownership receipt or byte-identical generated content.
158
208
 
159
209
  ## Limitations
160
210
 
161
211
  - **Bases needs Obsidian ≥ 1.9** (≥ 1.10 for card/list views). The `.base` file is harmless on older
162
212
  versions, it simply doesn't render.
163
- - **`gem_source` units are not exported.** They aren't reachable through the index reader; only
164
- `rails_source` is covered by `include_framework`.
165
213
  - **Hand-edits diverge.** The vault is meant to be regenerated. Editing a note's properties in
166
214
  Obsidian rewrites its frontmatter, after which a re-export will overwrite your changes.
167
215
  - **Nested vaults ignore the shipped config.** Open the generated folder as its own vault for graph
168
216
  colors, Bases, and link-format settings to apply.
169
217
 
170
- See `docs/AGENT_GUIDE.md` for how an agent consumes the `_woods/` sidecar.
218
+ For MCP-based exploration instead of local vault files, see [the agent guide](AGENT_GUIDE.md).
@@ -1,5 +1,8 @@
1
1
  # Reading a published index from Ruby
2
2
 
3
+ For shell, Python, shipping snapshots, or direct JSON access, use the
4
+ [published filesystem layout contract](INDEX_LAYOUT.md).
5
+
3
6
  `Woods::PublishedIndex` is the stable, read-only API for tools that are not MCP clients: RuboCop cops, CI gate scripts, and the `woods:check:*` tasks. It needs no Rails, opens one published generation, and never moves off it for the life of the reader.
4
7
 
5
8
  ```ruby
@@ -24,6 +27,58 @@ Woods::PublishedIndex.open(Rails.root.join('tmp/woods')) do |index|
24
27
  end
25
28
  ```
26
29
 
30
+ ## Validate a published generation
31
+
32
+ ```ruby
33
+ require 'woods/resilience/index_validator'
34
+
35
+ report = Woods::Resilience::IndexValidator.new(index_dir: 'tmp/woods').validate
36
+ report.valid? # false when artifact or semantic graph errors were found
37
+ report.errors
38
+ report.warnings
39
+ ```
40
+
41
+ `woods:validate` uses the same checker. Supporting development versions validate
42
+ raw graph relationships as well as JSON, content hashes, and indexed files; see
43
+ the [semantic invariants and limits](INDEX_LAYOUT.md#semantic-graph-validation).
44
+ Errors retain typed identifiers and artifact paths. Writer-version and source-path
45
+ warnings keep their existing advisory behavior.
46
+
47
+ Each call resolves and pins one published generation for all checks. Concurrent
48
+ publication can advance the pointer, but retention cannot remove the payload
49
+ being validated. The next call sees the new generation. Locks release on success
50
+ and errors; malformed pointers fail instead of falling back to stale root files.
51
+ Legacy flat layouts remain supported but cannot provide immutable-generation
52
+ isolation against in-place writes. Bare type-directory fixtures without a
53
+ manifest retain structural-only validation.
54
+
55
+ The reusable `Woods::Resilience::GraphInvariantValidator` accepts raw string-keyed
56
+ `graph:` data and typed `index_entries:` and returns an array of errors without
57
+ modifying either input. Callers supplying raw data must keep it within one pinned
58
+ generation and verify the unit artifacts themselves, as `IndexValidator` does.
59
+ Validation is a read-only diagnostic, not automatic repair or proof of runtime
60
+ execution.
61
+
62
+ ## Manifest writer provenance
63
+
64
+ The published `manifest.json` records `woods_version`, a string naming the Woods gem version
65
+ that last published that manifest. Full extraction, changed incremental runs,
66
+ targeted refreshes, and the static Woods self-map write it. A no-op leaves the
67
+ published manifest and its version unchanged. Resolve the manifest through the
68
+ generation pointer, as Woods readers do.
69
+
70
+ Older manifests may omit the field or contain `null`; that means unknown.
71
+ `woods_status.index.woods_version` reports this value from the served manifest,
72
+ while `woods_status.server.version` identifies the running MCP reader. The reader
73
+ never substitutes its own version for missing writer provenance.
74
+
75
+ `woods:validate` warns when a present writer version is malformed or its major
76
+ version differs from the installed validator. The warning is advisory and does
77
+ not invalidate an otherwise structurally valid index. Re-run full extraction
78
+ when investigating a major-version mismatch. Matching versions do not certify
79
+ compatibility: an incremental publisher can retain units written by an older
80
+ version. This field does not replace the [upgrade procedure](UPGRADING_TO_2.md).
81
+
27
82
  ## One generation, pinned for the reader's whole life
28
83
 
29
84
  Unlike `Woods::MCP::IndexReader`, a `PublishedIndex` never refreshes between calls. It resolves one generation at `.new`/`.open` time and every fact it returns, `units`, the table map, `generation_number`, `external_dependency_checksum`, comes from that one generation for as long as the reader is open. There is no `reload` and no auto-refresh: open a new reader to see a later publish.
@@ -54,6 +109,23 @@ The reader wraps `Woods::MCP::IndexReader` with `auto_refresh: false`; the unit
54
109
 
55
110
  `Woods::MCP::IndexReader#find_unit` keys its identifier map on identifier alone. If two type directories both list the same identifier (a model and a service both named `Foo`, for example), whichever type sorts last in `Woods::MCP::IndexReader::TYPE_DIRS` silently wins, and `unit(identifier)` returns that one. Pass `type:` to read a specific type's unit file directly and skip the collision entirely; `#table_database_map` always does this internally (`type: 'model'`), so a same-named non-model unit can never shadow a model's `table_name`/`database`.
56
111
 
112
+ ### Actual unit types and directory families
113
+
114
+ Unreleased after `2.0.0.beta3`: `unit` and `units` accept actual published
115
+ `graphql_type`, `graphql_mutation`, `graphql_resolver`, `graphql_query`, and
116
+ `gem_source` types. Enumeration preserves each unit's actual type rather than
117
+ labeling every GraphQL member `graphql` or every gem source `rails_source`.
118
+ Check the loaded Git revision when testing a development checkout.
119
+
120
+ The existing `graphql` and `rails_source` filters remain directory-family
121
+ aliases: `graphql` selects all four GraphQL types; `rails_source` selects both
122
+ Rails and gem sources. Returned records keep their actual type. Use
123
+ `units(type: 'gem_source')` for gem sources only; to select Rails sources only,
124
+ filter `units(type: 'rails_source')` on each entry's `"type" == "rails_source"`.
125
+ Unknown types return no records, and an explicit GraphQL subtype never matches
126
+ another subtype. Enumeration continues to return index-entry fields, not full
127
+ unit source bodies.
128
+
57
129
  ### `available_generations`: published means published
58
130
 
59
131
  A generation is listed only when both hold:
data/docs/README.md CHANGED
@@ -36,10 +36,13 @@ and stdio or Streamable HTTP endpoints directly.
36
36
  - [MCP tool cookbook](MCP_TOOL_COOKBOOK.md): scenario-based calls with parameters and expected response shapes.
37
37
  - [Console MCP setup](CONSOLE_MCP_SETUP.md): Console transports, blocked tables, credential scanning, redaction, SQL validation, and production safeguards.
38
38
  - [MCP HTTP transport](MCP_HTTP_TRANSPORT.md): shared/remote Index Server transport, authentication, origins, and protocol details.
39
+ - [Edit client adapters](CLIENT_HOOKS.md): opt-in Claude/OpenCode registration, complete path batches, and recovery.
39
40
  - [MCP worktree setup](MCP_WORKTREE_SETUP.md): register Woods correctly when agents work in linked git worktrees.
40
41
 
41
42
  ## Index lifecycle
42
43
 
44
+ - [Source freshness](SOURCE_FRESHNESS.md): verify dirty source against a served generation, establish a fresh-process baseline, and understand bounded unknown results.
45
+
43
46
  - [Retrieval guide](RETRIEVAL_GUIDE.md): configure embeddings and understand semantic retrieval, ranking, and token budgets.
44
47
  - [Embedding models](EMBEDDING_MODELS.md): choose and size local Ollama models.
45
48
  - [Upgrade to Woods 2.0](UPGRADING_TO_2.md): identifier changes, atomic payloads, durable-store reconciliation, and rollback.
@@ -48,8 +51,10 @@ and stdio or Streamable HTTP endpoints directly.
48
51
 
49
52
  - [Why Woods](WHY_WOODS.md): the problems runtime introspection solves.
50
53
  - [Internals](INTERNALS.md): extraction, publication, graph, storage, retrieval, and MCP components.
54
+ - [Ruby runtime trace enrichment](RUNTIME_TRACING.md): record observed Ruby callers and merge trace evidence into method units.
51
55
  - [Extractor reference](EXTRACTOR_REFERENCE.md): what each extractor produces and the edge cases it handles.
52
56
  - [Reading a published index from Ruby](PUBLISHED_INDEX.md): the `Woods::PublishedIndex` Ruby API for cops, gate scripts, and `woods:check:*` tasks (including the moved-message check).
57
+ - [Published index layout](INDEX_LAYOUT.md): the filesystem contract for non-Ruby readers, with Bash/jq and Python examples, retention pins, and snapshot-copy rules.
53
58
  - [Evaluation](EVALUATION.md): retrieval scoring, baselines, and the agent-level index on/off ablation.
54
59
  - [Backend matrix](BACKEND_MATRIX.md): implemented provider/store combinations and their operational requirements.
55
60
  - [Token benchmark](TOKEN_BENCHMARK.md): evidence behind Woods token-estimation defaults.
@@ -89,6 +94,8 @@ Use this map when changing behavior or documentation. Update the owner first; ot
89
94
  | Contributor policy | [CONTRIBUTING.md](../CONTRIBUTING.md) |
90
95
  | Coding-agent repository instructions | [AGENTS.md](https://github.com/lost-in-the/woods/blob/main/AGENTS.md) |
91
96
  | Non-MCP Ruby access to a published index | [PUBLISHED_INDEX.md](PUBLISHED_INDEX.md) |
97
+ | Generation-bound source evidence and fresh extraction | [SOURCE_FRESHNESS.md](SOURCE_FRESHNESS.md) |
98
+ | Published filesystem layout for external readers | [INDEX_LAYOUT.md](INDEX_LAYOUT.md) |
92
99
  | Evaluation harnesses | [EVALUATION.md](EVALUATION.md) |
93
100
 
94
101
  The current public surface is generated from 35 extractors. Counts and capability claims must match `.Codex/release-v2/surface-inventory.json`, which is generated from the code and verified in CI.