woods 2.0.0.beta2 → 2.0.0.beta3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (218) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +262 -1
  3. data/CONTRIBUTING.md +173 -9
  4. data/README.md +7 -3
  5. data/SECURITY.md +9 -6
  6. data/docs/AGENT_GUIDE.md +83 -4
  7. data/docs/AGENT_SETUP.md +82 -1
  8. data/docs/BACKEND_MATRIX.md +20 -0
  9. data/docs/CLIENT_HOOKS.md +111 -0
  10. data/docs/CONFIGURATION_REFERENCE.md +199 -14
  11. data/docs/CONSOLE_MCP_SETUP.md +35 -5
  12. data/docs/DOCKER_SETUP.md +21 -2
  13. data/docs/EVALUATION.md +464 -1
  14. data/docs/EXTRACTOR_REFERENCE.md +36 -5
  15. data/docs/FAQ.md +11 -12
  16. data/docs/GETTING_STARTED.md +17 -5
  17. data/docs/INCREMENTAL_EXTRACTION.md +117 -1
  18. data/docs/INDEX_LAYOUT.md +382 -0
  19. data/docs/INTERNALS.md +7 -2
  20. data/docs/MCP_SERVERS.md +221 -5
  21. data/docs/MCP_TOOL_COOKBOOK.md +33 -18
  22. data/docs/NOTION_INTEGRATION.md +13 -0
  23. data/docs/OBSIDIAN_INTEGRATION.md +57 -9
  24. data/docs/PUBLISHED_INDEX.md +55 -0
  25. data/docs/README.md +7 -0
  26. data/docs/RETRIEVAL_GUIDE.md +253 -11
  27. data/docs/RUNTIME_TRACING.md +71 -0
  28. data/docs/SOURCE_FRESHNESS.md +143 -0
  29. data/docs/TROUBLESHOOTING.md +117 -5
  30. data/docs/UNBLOCKED_INTEGRATION.md +25 -0
  31. data/docs/UPGRADING_TO_2.md +44 -22
  32. data/docs/WATCH_DAEMON.md +259 -59
  33. data/exe/woods-agent-config +6 -0
  34. data/exe/woods-extract +5 -0
  35. data/exe/woods-hook-context +6 -0
  36. data/lib/generators/woods/templates/woods.rb.tt +1 -3
  37. data/lib/tasks/woods.rake +47 -397
  38. data/lib/woods/agent_configuration/applier.rb +133 -0
  39. data/lib/woods/agent_configuration/cli.rb +101 -0
  40. data/lib/woods/agent_configuration/cli_options.rb +29 -0
  41. data/lib/woods/agent_configuration/document.rb +105 -0
  42. data/lib/woods/agent_configuration/error.rb +7 -0
  43. data/lib/woods/agent_configuration/launcher.rb +75 -0
  44. data/lib/woods/agent_configuration/layout.rb +59 -0
  45. data/lib/woods/agent_configuration/managed_section.rb +62 -0
  46. data/lib/woods/agent_configuration/plan.rb +98 -0
  47. data/lib/woods/agent_configuration/plan_diff.rb +38 -0
  48. data/lib/woods/agent_configuration/planned_files.rb +61 -0
  49. data/lib/woods/agent_configuration/planner.rb +63 -0
  50. data/lib/woods/agent_configuration/planner_validation.rb +77 -0
  51. data/lib/woods/agent_configuration/preflight.rb +100 -0
  52. data/lib/woods/agent_configuration/recovery.rb +49 -0
  53. data/lib/woods/ast/node.rb +2 -0
  54. data/lib/woods/ast/parser.rb +38 -5
  55. data/lib/woods/builder.rb +21 -5
  56. data/lib/woods/cache/cache_middleware.rb +28 -7
  57. data/lib/woods/cache/cache_store.rb +4 -5
  58. data/lib/woods/change_set.rb +5 -4
  59. data/lib/woods/console/credential_index.rb +20 -2
  60. data/lib/woods/console/credential_scanner.rb +14 -14
  61. data/lib/woods/console/credential_scanner_registry.rb +36 -0
  62. data/lib/woods/console/embedded_executor.rb +1 -1
  63. data/lib/woods/console/encrypted_credential_snapshot.rb +16 -0
  64. data/lib/woods/console/rack_middleware.rb +22 -13
  65. data/lib/woods/console/server.rb +18 -16
  66. data/lib/woods/dependency_graph.rb +65 -13
  67. data/lib/woods/embedding/corpus.rb +94 -0
  68. data/lib/woods/embedding/indexer.rb +90 -46
  69. data/lib/woods/embedding/openai.rb +17 -6
  70. data/lib/woods/evaluation/ablation_executor.rb +6 -1
  71. data/lib/woods/evaluation/ablation_timed_executor.rb +22 -4
  72. data/lib/woods/export/typed_reader.rb +56 -0
  73. data/lib/woods/extractor.rb +232 -137
  74. data/lib/woods/extractors/action_cable_extractor.rb +3 -1
  75. data/lib/woods/extractors/behavioral_profile.rb +9 -7
  76. data/lib/woods/extractors/caching_extractor.rb +3 -1
  77. data/lib/woods/extractors/concern_extractor.rb +64 -6
  78. data/lib/woods/extractors/configuration_extractor.rb +7 -3
  79. data/lib/woods/extractors/controller_extractor.rb +13 -4
  80. data/lib/woods/extractors/database_view_extractor.rb +3 -1
  81. data/lib/woods/extractors/decorator_extractor.rb +3 -1
  82. data/lib/woods/extractors/engine_extractor.rb +3 -1
  83. data/lib/woods/extractors/event_extractor.rb +4 -2
  84. data/lib/woods/extractors/factory_extractor.rb +3 -1
  85. data/lib/woods/extractors/graphql_extractor.rb +8 -2
  86. data/lib/woods/extractors/i18n_extractor.rb +3 -1
  87. data/lib/woods/extractors/job_extractor.rb +6 -19
  88. data/lib/woods/extractors/lib_extractor.rb +3 -1
  89. data/lib/woods/extractors/mailer_extractor.rb +20 -5
  90. data/lib/woods/extractors/manager_extractor.rb +3 -1
  91. data/lib/woods/extractors/method_parameters.rb +53 -0
  92. data/lib/woods/extractors/middleware_argument.rb +65 -0
  93. data/lib/woods/extractors/middleware_extractor.rb +9 -3
  94. data/lib/woods/extractors/migration_extractor.rb +3 -1
  95. data/lib/woods/extractors/model_extractor.rb +39 -33
  96. data/lib/woods/extractors/package_extractor.rb +24 -4
  97. data/lib/woods/extractors/phlex_extractor.rb +3 -1
  98. data/lib/woods/extractors/policy_extractor.rb +3 -1
  99. data/lib/woods/extractors/poro_extractor.rb +3 -1
  100. data/lib/woods/extractors/pundit_extractor.rb +3 -1
  101. data/lib/woods/extractors/rails_source_extractor.rb +4 -2
  102. data/lib/woods/extractors/rake_task_extractor.rb +4 -2
  103. data/lib/woods/extractors/route_extractor.rb +3 -1
  104. data/lib/woods/extractors/route_helper_resolver.rb +10 -33
  105. data/lib/woods/extractors/scheduled_job_extractor.rb +41 -15
  106. data/lib/woods/extractors/serializer_extractor.rb +4 -2
  107. data/lib/woods/extractors/service_extractor.rb +3 -1
  108. data/lib/woods/extractors/shared_dependency_scanner.rb +2 -2
  109. data/lib/woods/extractors/shared_utility_methods.rb +27 -15
  110. data/lib/woods/extractors/source_nesting.rb +1 -1
  111. data/lib/woods/extractors/state_machine_extractor.rb +3 -1
  112. data/lib/woods/extractors/test_mapping_extractor.rb +3 -1
  113. data/lib/woods/extractors/validator_extractor.rb +3 -1
  114. data/lib/woods/extractors/view_component_extractor.rb +3 -1
  115. data/lib/woods/extractors/view_template_extractor.rb +3 -1
  116. data/lib/woods/gem_mapper.rb +2 -0
  117. data/lib/woods/git_history.rb +116 -0
  118. data/lib/woods/graph_analyzer.rb +35 -6
  119. data/lib/woods/hooks/context_cli.rb +54 -0
  120. data/lib/woods/hooks/context_event.rb +88 -0
  121. data/lib/woods/hooks/context_hint.rb +73 -0
  122. data/lib/woods/hooks/context_impact.rb +77 -0
  123. data/lib/woods/hooks/context_output.rb +47 -0
  124. data/lib/woods/hooks/context_state.rb +102 -0
  125. data/lib/woods/hooks/refresh.rb +79 -0
  126. data/lib/woods/hooks/rule_projection.rb +78 -0
  127. data/lib/woods/input_rules.rb +19 -0
  128. data/lib/woods/mcp/bearer_auth.rb +20 -12
  129. data/lib/woods/mcp/bootstrapper.rb +62 -0
  130. data/lib/woods/mcp/index_reader.rb +323 -160
  131. data/lib/woods/mcp/initialization_guidance.rb +27 -0
  132. data/lib/woods/mcp/origin_guard.rb +17 -9
  133. data/lib/woods/mcp/published_lexical_retriever.rb +115 -0
  134. data/lib/woods/mcp/renderers/markdown_renderer.rb +8 -1
  135. data/lib/woods/mcp/renderers/plain_renderer.rb +7 -1
  136. data/lib/woods/mcp/search_results.rb +74 -0
  137. data/lib/woods/mcp/server.rb +158 -37
  138. data/lib/woods/mcp/tool_contract.rb +2 -0
  139. data/lib/woods/mcp/tool_response_renderer.rb +25 -0
  140. data/lib/woods/mcp/traversal_evidence.rb +113 -0
  141. data/lib/woods/mcp/traversal_evidence_index.rb +100 -0
  142. data/lib/woods/mcp/traversal_evidence_page.rb +41 -0
  143. data/lib/woods/mcp/traversal_evidence_text.rb +52 -0
  144. data/lib/woods/notion/exporter.rb +56 -17
  145. data/lib/woods/obsidian/destination_plan.rb +98 -0
  146. data/lib/woods/obsidian/name_mapper.rb +19 -3
  147. data/lib/woods/obsidian/note_builder.rb +19 -10
  148. data/lib/woods/obsidian/vault_exporter.rb +88 -32
  149. data/lib/woods/operator/pipeline_guard.rb +18 -13
  150. data/lib/woods/path_dispatcher.rb +7 -1
  151. data/lib/woods/payload_store.rb +27 -26
  152. data/lib/woods/railtie.rb +3 -3
  153. data/lib/woods/railtie_support.rb +12 -12
  154. data/lib/woods/rake_helpers.rb +392 -0
  155. data/lib/woods/resilience/graph_invariant_validator/membership_checks.rb +71 -0
  156. data/lib/woods/resilience/graph_invariant_validator/node_checks.rb +61 -0
  157. data/lib/woods/resilience/graph_invariant_validator/reverse_relationship_checks.rb +46 -0
  158. data/lib/woods/resilience/graph_invariant_validator.rb +119 -0
  159. data/lib/woods/resilience/index_validator/graph_checks.rb +80 -0
  160. data/lib/woods/resilience/index_validator.rb +112 -23
  161. data/lib/woods/retrieval/context_assembler.rb +50 -15
  162. data/lib/woods/retrieval/lexical_assembler.rb +73 -0
  163. data/lib/woods/retrieval/lexical_index.rb +119 -0
  164. data/lib/woods/retrieval/ranker.rb +4 -2
  165. data/lib/woods/retrieval/scope.rb +108 -0
  166. data/lib/woods/retrieval/scoped_graph_store.rb +32 -0
  167. data/lib/woods/retrieval/scoped_vector_store.rb +55 -0
  168. data/lib/woods/retrieval/search_executor.rb +86 -27
  169. data/lib/woods/retrieval/source_evidence.rb +200 -0
  170. data/lib/woods/retriever.rb +98 -22
  171. data/lib/woods/ruby_analyzer/trace_enricher.rb +77 -38
  172. data/lib/woods/session_tracer/middleware.rb +10 -12
  173. data/lib/woods/session_tracer/redis_store.rb +22 -6
  174. data/lib/woods/session_tracer/session_flow_assembler.rb +23 -17
  175. data/lib/woods/session_tracer/solid_cache_coordination.rb +6 -4
  176. data/lib/woods/session_tracer/unit_resolver.rb +63 -0
  177. data/lib/woods/source_inputs/consumer_errors.rb +27 -0
  178. data/lib/woods/source_inputs/handoff.rb +102 -0
  179. data/lib/woods/source_inputs/launcher.rb +157 -0
  180. data/lib/woods/source_inputs/manifest.rb +124 -0
  181. data/lib/woods/source_inputs/private_key.rb +55 -0
  182. data/lib/woods/source_inputs/scanner.rb +171 -0
  183. data/lib/woods/source_inputs/scopes.rb +71 -0
  184. data/lib/woods/source_inputs/session.rb +214 -0
  185. data/lib/woods/source_inputs/status.rb +84 -0
  186. data/lib/woods/source_inputs/verifier.rb +107 -0
  187. data/lib/woods/storage/metadata_store.rb +25 -25
  188. data/lib/woods/storage/pgvector.rb +29 -8
  189. data/lib/woods/storage/qdrant.rb +17 -7
  190. data/lib/woods/storage/vector_store.rb +18 -6
  191. data/lib/woods/tasks.rb +3 -2
  192. data/lib/woods/temporal/json_snapshot_store.rb +29 -8
  193. data/lib/woods/unblocked/exporter.rb +59 -70
  194. data/lib/woods/version.rb +1 -1
  195. data/lib/woods/watch/boot_snapshot.rb +52 -0
  196. data/lib/woods/watch/daemon.rb +136 -28
  197. data/lib/woods/watch/listen_watcher.rb +4 -0
  198. data/lib/woods/watch/polling_watcher.rb +5 -1
  199. data/lib/woods/watch/status.rb +20 -15
  200. data/lib/woods/watch/tree_scan.rb +21 -13
  201. data/lib/woods/watch/watcher.rb +4 -1
  202. data/lib/woods.rb +50 -11
  203. data/plugin/.claude-plugin/plugin.json +1 -1
  204. data/plugin/hooks/adapters/normalize.jq +15 -0
  205. data/plugin/hooks/adapters/normalize.rb +63 -0
  206. data/plugin/hooks/hooks.json +20 -0
  207. data/plugin/hooks/woods-context.sh +50 -0
  208. data/plugin/hooks/woods-input-rules.sh +159 -0
  209. data/plugin/hooks/woods-opencode.mjs +65 -0
  210. data/plugin/hooks/woods-post-edit.sh +2 -225
  211. data/plugin/hooks/woods-refresh.sh +260 -0
  212. data/plugin/hooks/woods-session-start.sh +47 -55
  213. data/plugin/skills/woods-agent-enable/SKILL.md +13 -0
  214. data/plugin/skills/woods-diagnose/SKILL.md +288 -1
  215. data/plugin/skills/woods-investigate/SKILL.md +106 -0
  216. data/plugin/skills/woods-mcp-config/SKILL.md +89 -1
  217. data/plugin/skills/woods-setup/SKILL.md +107 -6
  218. metadata +84 -5
data/docs/MCP_SERVERS.md CHANGED
@@ -27,6 +27,15 @@ bin/rails woods:validate
27
27
  bin/rails woods:stats
28
28
  ```
29
29
 
30
+ For embedding-free ranked retrieval, start Index MCP with
31
+ `WOODS_RETRIEVAL_MODE=lexical`. This opt-in reads the published extraction units;
32
+ it does not probe providers or load vectors. `woods_status.retriever.mode` reports
33
+ `lexical`, and inactive embedding fields are `null`. The default semantic mode
34
+ keeps its existing embedding setup and failure behavior. See
35
+ [retrieval modes](RETRIEVAL_GUIDE.md#embedding-free-lexical-retrieval) for scoring,
36
+ query limits and measured tradeoffs. Both packaged stdio and HTTP launches honor
37
+ the setting; put it in the MCP process's environment, not just a Rails initializer.
38
+
30
39
  The stdio server can then run outside Rails. Point it at the index root (`tmp/woods/` by default), not at an internal generation or payload directory.
31
40
 
32
41
  ### Configure a stdio client
@@ -47,6 +56,10 @@ Prefer the application's bundle and a project-scoped configuration:
47
56
 
48
57
  `woods-mcp-start` checks that the directory and published manifest exist, then replaces itself with `woods-mcp`. It does not install dependencies or restart a crashed process.
49
58
 
59
+ `woods_status.index.woods_version` identifies the last publisher of the served
60
+ manifest; `server.version` identifies the running MCP reader. Missing writer
61
+ provenance is `null`. See [manifest writer provenance](PUBLISHED_INDEX.md#manifest-writer-provenance).
62
+
50
63
  You can launch the server directly when the client already handles preflight:
51
64
 
52
65
  ```bash
@@ -57,6 +70,12 @@ Keep stdout reserved for MCP protocol messages. Diagnose startup failures from s
57
70
 
58
71
  ### Client configuration locations
59
72
 
73
+ Supporting development builds offer preview/apply/update/remove ownership for
74
+ Claude Code project or explicit user configuration. See
75
+ [managed configuration](AGENT_SETUP.md#managed-claude-code-configuration) for
76
+ `woods-agent-config`, host/Compose preflight, conflict handling, and recovery.
77
+ Manual configuration remains available for older gems and other clients.
78
+
60
79
  MCP clients expose project or user-level server settings in different locations. Use project scope when available, preserve the `command`, `args`, and absolute `cwd` semantics above, and translate only the surrounding client-specific format. Woods is model-independent: compatibility depends on the client supporting MCP stdio or Streamable HTTP, not on whether the connected model is from OpenAI, Anthropic, Google, xAI, or another provider.
61
80
 
62
81
  Client configuration formats can change independently of Woods. If a client rejects otherwise valid JSON, check that client's current MCP documentation.
@@ -99,6 +118,27 @@ Reconnect the client, then call:
99
118
 
100
119
  Prefer a real MCP client's connection flow over a hand-written JSON-RPC pipe. Modern MCP 2026-07-28 requests carry per-request protocol metadata and can use `server/discover` without an initialization handshake; older clients still use `initialize`. A valid raw smoke test must implement one complete flow rather than sending an isolated `tools/list` or `tools/call` request.
101
120
 
121
+ ### Initialization guidance
122
+
123
+ The Index Server supplies a short, client-neutral `instructions` field through
124
+ the SDK's `initialize` response and modern `server/discover`. It describes the
125
+ status → discovery → inspection → bounded traversal → source-verification
126
+ workflow, and lists only the tools actually registered by this server.
127
+ Instructions are stable for unchanged tool registration and bounded to 2,048
128
+ UTF-8 bytes across supported configurations. Building the text does not probe
129
+ providers, extract code, or write configuration.
130
+
131
+ Registration alone does not establish retrieval readiness: check `woods_status`
132
+ before using `codebase_retrieve`, including after reload. The guidance grants
133
+ no extraction, configuration-change, or Console authorization. Detailed usage
134
+ belongs in the [agent guide](AGENT_GUIDE.md).
135
+
136
+ This addition is unreleased after `2.0.0.beta2`. The SDK omits `instructions`
137
+ when negotiating protocol `2024-11-05`; that behavior is preserved. Older gems,
138
+ legacy clients, and clients that do not show server instructions can use the
139
+ agent guide or investigation skill. Leave protocol negotiation enabled rather
140
+ than pinning a newer version solely to obtain guidance.
141
+
102
142
  ### Tools (29 — 14 registered in the packaged default)
103
143
 
104
144
  The Index Server defines 29 schemas across core and conditional capabilities. The normal packaged executable registers the 14 tools below; the remaining schemas require the specialized wiring described afterward.
@@ -108,8 +148,8 @@ The Index Server defines 29 schemas across core and conditional capabilities. Th
108
148
  | `woods_status` | Index health, generation, counts, and retrieval readiness |
109
149
  | `search` | Discover identifiers by regex, prefix, suffix, source, or metadata |
110
150
  | `lookup` | Fetch one exact unit with source, metadata, and relationships |
111
- | `dependencies` | Traverse what a unit depends on (`depth`, `types`, `via` narrow; `limit`, `offset` page) |
112
- | `dependents` | Traverse what depends on a unit (`depth`, `types`, `via` narrow; `limit`, `offset` page) |
151
+ | `dependencies` | Traverse what a unit depends on (`depth`, `types`, `via` narrow; `max_nodes`, `max_edges` budget work; `limit`, `offset` page) |
152
+ | `dependents` | Traverse what depends on a unit (`depth`, `types`, `via` narrow; `max_nodes`, `max_edges` budget work; `limit`, `offset` page) |
113
153
  | `structure` | Summarize structural relationships around a unit |
114
154
  | `trace_flow` | Follow a request, job, mail, or other execution flow |
115
155
  | `framework` | Inspect relevant Rails or installed gem source |
@@ -118,7 +158,29 @@ The Index Server defines 29 schemas across core and conditional capabilities. Th
118
158
  | `domain_clusters` | Discover connected domains in the graph |
119
159
  | `pagerank` | Find structurally central units |
120
160
  | `reload` | Reload a newly published generation without restarting the client |
121
- | `codebase_retrieve` | Natural-language retrieval; returns a configuration error until embeddings exist |
161
+ | `codebase_retrieve` | Natural-language retrieval with embeddings or explicit lexical mode over extraction output |
162
+
163
+ When an identifier appears in multiple extraction types, `framework` reads its
164
+ framework-source bucket and `recent_changes` reads each selected type bucket.
165
+ Their paths and metadata belong to that selected bucket. When session tracing
166
+ is configured, newly recorded requests use the dispatched controller's runtime
167
+ class name; the fallback for requests without an instance respects Rails acronym
168
+ inflections. Existing trace records are unchanged. Controller lookup and root
169
+ outgoing-edge selection preserve the controller type. Downstream references and the shared context pool still use
170
+ bare identifiers. If a dependency encountered within the requested depth has
171
+ multiple published types, `session_trace` returns an `ambiguous_identity` tool
172
+ error naming the identifier and candidate types, with no partial context. This
173
+ also prevents an earlier dependency from occupying a later controller’s context
174
+ key. Unrelated collisions do not block a trace, and a known controller root keeps
175
+ its controller identity. A controller absent from the index remains in the
176
+ timeline without a source reference, so another type cannot fill that reference.
177
+ Candidate discovery and source reads use one pinned
178
+ generation. Corrupt or missing listed artifacts retain the `internal_error`
179
+ failure boundary; they do not prove uniqueness or become `ambiguous_identity` errors.
180
+ Use `depth: 0` for the request timeline, or inspect candidates with typed `lookup`
181
+ calls. Re-extraction does not remove a legitimate cross-type collision. Successful
182
+ traces retain their existing identifiers and response shape; target identity has
183
+ not been migrated globally. These corrections are unreleased after `2.0.0.beta2`.
122
184
 
123
185
  The server also exposes MCP resources and resource templates for indexed units. Tool descriptions returned by MCP are the parameter-level source of truth; [Agent guide](AGENT_GUIDE.md) explains selection strategy.
124
186
 
@@ -130,6 +192,126 @@ error and continues serving the previous aligned generation; it never swaps in a
130
192
  partial or empty replacement. Grant write access for live reloads, or restart the MCP
131
193
  process after publishing a new embedded index.
132
194
 
195
+ ### Search completeness
196
+
197
+ Search responses retain `query`, `result_count`, and `results`; `result_count`
198
+ is the number returned, not an estimated total. The additive `completeness`
199
+ object describes the requested types, literal filters, and fields in the pinned
200
+ generation. This contract is unreleased after `2.0.0.beta2`.
201
+
202
+ | `reason` | `status` | `has_more` | `total_matches` |
203
+ |---|---|---|---|
204
+ | `exhausted` | `complete` | `false` | Exact count |
205
+ | `result_limit` | `partial` | `true` | `null` (unknown) |
206
+ | `scan_budget` or `regex_timeout` | `partial` | `null` (unknown) | `null` (unknown) |
207
+
208
+ `matched_lower_bound` counts distinct observed `(type, identifier)` matches,
209
+ including at most one lookahead match beyond `limit`. A result-limit response
210
+ therefore establishes another match; an exactly full page can instead be
211
+ complete if the requested domain is exhausted. Deep lookahead shares
212
+ `WOODS_SEARCH_MAX_SCAN` with the initial scan and retains round-robin scanning
213
+ across types. Search does not count the entire omitted tail or offer pagination.
214
+ The existing `types` filter and result labels name directory families:
215
+ `rails_source` includes both Rails and gem source units. Deep reads accept those
216
+ two stored types only in that shared directory; `lookup` and lexical retrieval
217
+ retain the unit's actual `rails_source` or `gem_source` type.
218
+
219
+ All partial responses retain `partial: true` and include a narrowing `hint`.
220
+ JSON exposes these fields; Markdown, plain text, and Claude formats label the
221
+ returned count, stopping reason, known/unknown remainder, and total explicitly.
222
+ Narrow `types`, literal `exact_prefix`/`exact_suffix`, or deep `fields` before
223
+ using discovery as exhaustive evidence. Completeness applies to this index and
224
+ query domain, not to unindexed application code.
225
+
226
+ Detected missing, unreadable, or corrupt artifacts remain `isError: true` with
227
+ `_meta.error_code: "corrupt_artifact"`. Their `_meta.completeness` has
228
+ `status: "unknown"`, `reason: "unreadable_or_corrupt_source"`, and `null` for
229
+ `has_more`, `total_matches`, and `matched_lower_bound`; no successful empty
230
+ result is substituted. Inspect `woods_status` and run `woods:validate`.
231
+
232
+ ### Dependency traversal budgets
233
+
234
+ `dependencies` and `dependents` walk breadth-first in stored graph order. The
235
+ walk defaults to `max_nodes: 1000` (including the root) and `max_edges: 10000`;
236
+ callers can select 1–10,000 nodes and 1–100,000 edge checks. The node budget
237
+ counts distinct nodes admitted after filters. Every candidate edge is charged
238
+ before filtering, including duplicates, cycles, and the forward-edge checks
239
+ needed to match a reverse `via` filter. Thus restrictive filters cannot bypass
240
+ the edge budget. Nodes at the requested `depth` are recorded without reading
241
+ their adjacency lists.
242
+
243
+ When further expansion would exceed a budget, JSON reports `partial: true`,
244
+ `partial_reason: "node_budget"` or `"edge_budget"`, and `traversal_budget` with
245
+ `max_nodes`, `max_edges`, `visited_nodes`, and `visited_edges`. Text renderers
246
+ also identify the partial traversal. Already discovered nodes remain in the
247
+ answer, but an empty `deps` array in a partial answer does not prove a leaf.
248
+ Exact-budget walks that finish all requested work are complete and have no
249
+ `partial` marker.
250
+
251
+ `limit` (default 50) and `offset` only page that discovered result; they never
252
+ change the walk budget or depth. On a partial traversal, `nodes_total`, when
253
+ present for pagination, counts the discovered prefix, **not the full reachable
254
+ graph**. Paging beyond that prefix stays partial. To explore more, narrow
255
+ `depth`/`types`/`via`, choose another root, or increase the traversal budget within
256
+ its maximum. Keep the root, filters, budgets, and published generation unchanged
257
+ for stable pages. No wall-clock deadline is used, so cutoffs are deterministic.
258
+
259
+ Budgets cover traversal work after per-generation graph loading and cache
260
+ preparation (JSON parsing, typed-edge normalization, node types and database
261
+ metadata). They do not cap that initial load, elapsed time, or total process
262
+ memory. These arguments are unreleased in Woods 2.0.0.beta2; check the connected
263
+ server's tool schema before sending them to an older installation.
264
+
265
+ ### Traversal explanations
266
+
267
+ Supporting development versions accept `explain: true` on `dependencies` and
268
+ `dependents`. Check the connected schema first; this option is unreleased after
269
+ 2.0.0.beta2. Omitted or false keeps the existing compact response.
270
+
271
+ The additive `explanation` object contains:
272
+
273
+ - `direction`: `forward` or `reverse`, plus the requested `root` identity.
274
+ - `edges`: records keyed by response-local IDs such as `e0`. Every record keeps
275
+ the original **source → target** direction, even during reverse traversal.
276
+ `source` contains its recorded `identifier` and `type`; `target` contains its
277
+ identifier and the unique type when the published graph establishes one.
278
+ `via`, `through`, `through_db`, and `disable_joins` preserve recorded values;
279
+ absent legacy attributes are null (shown as unknown in text), including an
280
+ unrecorded `disable_joins` rather than an invented false value.
281
+ - `witnesses`: one shortest breadth-first predecessor per admitted identifier,
282
+ keyed by identifier. Each has `parent`, `edge_id`, `impact` (`root`, `direct`,
283
+ or `transitive`), and `typed_path_complete`. Follow parent references to the
284
+ root to reconstruct one witness; alternative paths are not enumerated.
285
+
286
+ A target name shared by several types has `type: null`,
287
+ `resolution: "ambiguous"`, and sorted `candidate_types`. An unresolved target has
288
+ `resolution: "unresolved"` and an empty candidate list. Forward artifacts do not
289
+ record target types, so the response cannot choose among candidates. A witness
290
+ through an ambiguous or unresolved identity sets `typed_path_complete: false`;
291
+ it describes identifier-level reachability, never a uniquely typed path.
292
+ `types` filters retain the compact traversal's identifier-level semantics: any
293
+ registered type can qualify a name, while edge evidence keeps its actual source
294
+ owner. Multiple relationship kinds between the same endpoints remain separate.
295
+
296
+ Direct witnesses establish a recorded root relationship; transitive witnesses
297
+ represent inferred downstream reachability through recorded relationships.
298
+ Neither establishes observed execution, confidence, call order, or test coverage.
299
+
300
+ Node pagination retains required ancestor witnesses once, marked `context: true`
301
+ when outside the page; returned rows have `context: false`. Context records do
302
+ not increase the result-row count. Page evidence retains the witness edges and
303
+ other observed relationships among its visible/context endpoints; an empty page
304
+ has empty edge/witness maps. Edge IDs are local to this traversal response.
305
+
306
+ All examined evidence shares the existing edge budget, before `via`/`types`
307
+ filtering. Current `reverse_via` buckets allow direct reverse evidence lookup;
308
+ legacy recovery charges each reverse candidate and every inspected forward edge.
309
+ The shared predecessor forest and emitted records remain bounded by admitted
310
+ nodes and inspected edges. Per-generation JSON loading and the cached
311
+ O(nodes + variants) ownership/type preparation are outside the walk budget;
312
+ explanation mode never flattens all forward edges as per-request preparation.
313
+ Partial traversal and pagination metadata retain the budget contract above.
314
+
133
315
  ### Conditional Index capabilities
134
316
 
135
317
  The Ruby server builder contains 15 additional schemas for sessions, pipeline operations, retrieval feedback, temporal snapshots, and Notion sync. They register only when their required collaborators or configuration are wired.
@@ -140,6 +322,7 @@ The normal packaged executable does not wire pipeline-operator or feedback-store
140
322
 
141
323
  Use HTTP only for a deliberate shared or remote deployment. It expands the network boundary and requires authentication, origin restrictions, and TLS termination. Follow [MCP HTTP transport](MCP_HTTP_TRANSPORT.md); do not translate the stdio example into an unauthenticated public listener.
142
324
 
325
+
143
326
  ## Console Server
144
327
 
145
328
  The Console Server launches a Rails process through direct, Docker, or SSH connection configuration. It reads live data and must be treated as a separate security decision.
@@ -151,11 +334,14 @@ Console MCP is disabled by default because it reads live application data. Enabl
151
334
  ```ruby
152
335
  Woods.configure do |config|
153
336
  config.console_mcp_enabled = true
154
- config.console_mcp_token = ENV["WOODS_CONSOLE_MCP_TOKEN"]
337
+ config.console_mcp_http_enabled = false # stdio-only
155
338
  end
156
339
  ```
157
340
 
158
- The bearer token authenticates HTTP clients and is not sent over stdio. Rails still validates Console configuration while booting: production requires a token of at least 32 characters whenever Console is enabled, including for a stdio-only client. Keep it in the application's secret store, not in the initializer.
341
+ This explicitly disables HTTP Console while retaining stdio access; no HTTP
342
+ token is needed at boot. Existing configurations default to HTTP enabled.
343
+ For HTTP deployment, enable the HTTP flag and configure its token, origins
344
+ and TLS using the [Console setup guide](CONSOLE_MCP_SETUP.md#option-c-http-rack-middleware).
159
345
 
160
346
  Without a console connection file, the executable then launches the Rails task directly from its `cwd`:
161
347
 
@@ -229,3 +415,33 @@ Report vulnerabilities privately through [SECURITY.md](../SECURITY.md).
229
415
  5. Check [Troubleshooting](TROUBLESHOOTING.md) for the exact stderr message.
230
416
 
231
417
  For agent query behavior after connection, continue to [Agent guide](AGENT_GUIDE.md).
418
+
419
+ ### Explicit retrieval and discovery scope
420
+
421
+ On a server whose tool schema advertises them, `packages` and `source_paths` narrow
422
+ `search` and `codebase_retrieve` before candidate limits. These are per-call
423
+ arguments, not configuration settings. Inspect applied scope and completeness;
424
+ a narrow graph query can omit relevant cross-boundary dependencies. See the
425
+ [scope contract](RETRIEVAL_GUIDE.md#explicit-package-and-source-path-scopes) for
426
+ root/nested ownership, path normalization, errors, storage support, and cost.
427
+
428
+ ### Source freshness in status
429
+
430
+ `woods_status` accepts optional `source_check: "quick"` (default, 250ms scan) or
431
+ `"deep"` (five seconds). `index.source_freshness` describes the served generation
432
+ as `current`, `drifted` or `unknown`; missing source/key and incomplete capture
433
+ never count as current. No Rails initialization or provider call is needed.
434
+ See [source freshness](SOURCE_FRESHNESS.md) for scope, private-key handling and
435
+ fresh-process extraction. Existing HEAD/dirty fields remain separate diagnostics.
436
+
437
+ ### Explicit source evidence modes
438
+
439
+ When advertised by the installed schema, `lookup` and `codebase_retrieve` accept
440
+ `evidence: 'compact'` or `'outline'`; omitted/`'full'` preserves existing behavior.
441
+ Retrieval uses its original query. Compact lookup accepts optional `query` and an
442
+ estimated `budget` (default 2000); full lookup remains complete. `lookup` also
443
+ accepts an actual `type` and a `source_sha256` guard for typed, byte-verified
444
+ follow-up from an excerpt. Compact modes cannot be combined with metadata-only
445
+ lookup controls. Structured provenance stays within the existing closed output
446
+ schema's `data` field. Read the [evidence contract](RETRIEVAL_GUIDE.md#compact-published-evidence-and-api-outlines)
447
+ before interpreting published line ranges as physical source locations.
@@ -255,13 +255,24 @@ The `metadata.inlined_concerns` array lists which concerns were resolved:
255
255
  }
256
256
  ```
257
257
 
258
- **What you'll get:** A BFS tree of everything that references `User`, controllers, services, jobs, mailers, up to 2 hops out. Set `depth: 1` for direct dependents only.
258
+ **What you'll get:** A BFS traversal of units that reference `User`, such as controllers, services, jobs, and mailers, up to 2 hops out. Set `depth: 1` for direct dependents only.
259
259
 
260
- The answer is bounded to 50 nodes. When it is cut, the response ends with a
260
+ The answer is paged to 50 nodes by default. When it is cut, the response ends with a
261
261
  `Showing N of M (truncated)` line, the same one `graph_analysis` prints. Reach
262
262
  for `depth`, `types` and `via` first: they make the answer smaller. `limit` and
263
263
  `offset` only page what those leave, so a hub read one page at a time still
264
- costs every page.
264
+ repeats the walk. A separate traversal budget can return `partial: true`;
265
+ that marker means the reachable graph is incomplete even after the final page.
266
+ See [traversal budgets](MCP_SERVERS.md#dependency-traversal-budgets) before
267
+ changing `max_nodes` or `max_edges`.
268
+
269
+ To explain why a row is affected, check the connected schema and add
270
+ `"explain": true`. The response preserves source-to-target labels even while
271
+ walking dependents. Its predecessor witnesses distinguish direct relationships
272
+ from transitive inferred reachability, retain ancestor context across pages, and
273
+ mark ambiguous types explicitly. See the
274
+ [explanation contract](MCP_SERVERS.md#traversal-explanations); these witnesses are
275
+ not proof of observed execution.
265
276
 
266
277
  To find only which jobs depend on `User`:
267
278
 
@@ -323,26 +334,30 @@ column.
323
334
  }
324
335
  ```
325
336
 
326
- **Example response:**
337
+ **Example JSON data** (inside the MCP response envelope):
327
338
 
328
339
  ```json
329
- [
330
- {
331
- "identifier": "PaymentsController",
332
- "type": "controller",
333
- "file_path": "app/controllers/payments_controller.rb",
334
- "metadata": {
335
- "actions": ["create", "show", "webhook"],
336
- "routes": [
337
- { "verb": "POST", "path": "/payments", "action": "create" },
338
- { "verb": "POST", "path": "/payments/webhook", "action": "webhook" }
339
- ]
340
- }
340
+ {
341
+ "query": "payment",
342
+ "result_count": 1,
343
+ "results": [
344
+ { "identifier": "PaymentsController", "type": "controller", "match_field": "identifier" }
345
+ ],
346
+ "completeness": {
347
+ "status": "complete",
348
+ "reason": "exhausted",
349
+ "has_more": false,
350
+ "total_matches": 1,
351
+ "matched_lower_bound": 1
341
352
  }
342
- ]
353
+ }
343
354
  ```
344
355
 
345
- Search `source_code` when you want semantic matches, not just naming matches.
356
+ Use `lookup` on the returned identifier for source, actions, and routes. Search
357
+ `source_code` for textual matches beyond names. Supporting versions distinguish
358
+ exact totals from a bounded result prefix; `partial` means this is discovery,
359
+ not an exhaustive list. Completeness metadata is unreleased after `2.0.0.beta2`;
360
+ see the [search contract](MCP_SERVERS.md#search-completeness).
346
361
 
347
362
  ---
348
363
 
@@ -34,6 +34,19 @@ Sync is incremental. A **sync manifest** (`<output_dir>/notion_sync_manifest.jso
34
34
 
35
35
  Manifest entries for models/columns that vanished from the current extraction are pruned, but **no Notion page is ever deleted** by the sync, there is no delete path. A renamed or removed model just leaves its old page in Notion untouched.
36
36
 
37
+ Models and migration dates are loaded by `(identifier, type)`. A same-named
38
+ PORO, library unit, or other extracted type cannot supply a model's table or
39
+ column payload. Public identifiers, page titles, and manifest keys are unchanged.
40
+ If a listed model or migration cannot be read with its exact identity, the
41
+ affected sync refuses before mapping pages or pruning its manifest. Preserve
42
+ the error, validate the index, and regenerate it before retrying; force sync
43
+ does not bypass identity checks. Each public sync method, including standalone
44
+ model or column sync, uses one validated generation snapshot through mapping
45
+ and API calls. Native readers validate the published index before selecting
46
+ model/migration payloads. Custom readers must provide complete published
47
+ enumeration or complete per-bucket listings plus
48
+ `find_unit(identifier, type:)` returning the requested identity.
49
+
37
50
  If the manifest is missing (first run, or a CI cache miss), the exporter falls back to the full lookup/create path for every page and rebuilds the manifest, correct, just more API calls than a steady-state run.
38
51
 
39
52
  ### Escape hatch
@@ -8,7 +8,7 @@ ways at once:
8
8
  Obsidian [Bases](https://help.obsidian.md/bases) table, and drill into a single unit's note with its
9
9
  dependencies and dependents as clickable wikilinks.
10
10
  - **By agents**: load the entire dependency topology from a single `_woods/` sidecar (one read,
11
- no per-note fan-out), with a stable `id → note path` manifest for navigation.
11
+ no per-note fan-out), with a stable typed-unit → note-path manifest for navigation.
12
12
 
13
13
  Unlike the Notion and Unblocked exporters, this one writes **local files only**: there is no API
14
14
  token, no network call, and no rate limit. An Obsidian vault is just a folder.
@@ -95,6 +95,33 @@ Wikilinks are path-qualified with an alias (`[[models/Account|Account]]`): the t
95
95
  sanitized vault path so the link always resolves, and the alias shows the original identifier. The
96
96
  note's `# H1` carries the clean identifier so the sanitized filename never shows as the title.
97
97
 
98
+ ## Sidecar identity and schema versions
99
+
100
+ A unit keeps its original `id` and `type` in note frontmatter. Different types may
101
+ share an identifier: a database view and a factory called `reports` export to
102
+ `database_views/reports.md` and `factories/reports.md`. Existing unambiguous note
103
+ paths and display aliases stay unchanged.
104
+
105
+ Check `schema_version` before consuming `_woods/manifest.json`:
106
+
107
+ - **Version 1:** the graph has no cross-type identifier collisions. The existing
108
+ `notes[id]` and `paths[path]` maps are unchanged.
109
+ - **Version 2:** `notes[id]` retains the graph's primary type (when exported), and
110
+ `variants` lists each additional exported unit as
111
+ `{ "identifier": "reports", "type": "factory", "path": "factories/reports.md" }`.
112
+ Combine primary entries and variants using **`(identifier, type)`** as identity.
113
+ `paths[path]` still gives the original identifier, so several paths may return
114
+ the same value. Resolve a path against both collections; `notes[id]` alone is
115
+ incomplete. A consumer supporting only version 1 must refuse version 2.
116
+
117
+ The exporter reads each variant's own outgoing graph edges and derives incoming
118
+ links from them. Persisted targets are bare identifiers: if a target names several
119
+ types, the exporter omits that relationship from note links and counts it in a
120
+ progress diagnostic. It does this even when one sibling is excluded or unreadable.
121
+ Human association links use the same ambiguity rule. Bare PageRank and graph-analysis
122
+ annotations are omitted for ambiguous identifiers, rather than assigned to either
123
+ type. The verbatim graph sidecar retains all original records for inspection.
124
+
98
125
  ## The three visualizer surfaces
99
126
 
100
127
  | Surface | Best for | Notes |
@@ -114,7 +141,7 @@ credential):
114
141
  | `WOODS_OUTPUT` | `config.output_dir` (`tmp/woods`) | extraction directory to read from |
115
142
  | `WOODS_OBSIDIAN_VAULT` | `<output>/obsidian_vault` | where to write the vault |
116
143
  | `WOODS_OBSIDIAN_INCLUDE_SOURCE` | off | embed each unit's source code (credential-scrubbed) in its note |
117
- | `WOODS_OBSIDIAN_INCLUDE_FRAMEWORK` | off | include `rails_source` units (large; off by default) |
144
+ | `WOODS_OBSIDIAN_INCLUDE_FRAMEWORK` | off | include `rails_source` and readable `gem_source` units (large; off by default) |
118
145
  | `WOODS_OBSIDIAN_FORCE_PURGE` | off | bypass the 30% mass-deletion guard during the stale-note sweep |
119
146
 
120
147
  ```bash
@@ -141,6 +168,28 @@ The exporter **fully regenerates** the vault on every run (no incremental manife
141
168
  cheap). Output is deterministic: re-running against an unchanged extraction produces byte-identical
142
169
  notes, so unchanged units never show up in a git diff.
143
170
 
171
+ Before writing, Woods renders the output and checks **every destination**. An existing note or index
172
+ must carry `woods_managed: true`; a `.woods-vault` sentinel alone never authorizes replacing your
173
+ notes or settings. Conflicts stop the export before any writes or sweep, appear in `errors`, and make
174
+ `woods:obsidian` exit nonzero. `WOODS_OBSIDIAN_FORCE_PURGE` does not bypass ownership checks.
175
+ Child symlinks, directories in place of files, and special files are refused.
176
+
177
+ Generated machine assets are tracked by SHA-256 in `_woods/ownership.json`: the three sidecar JSON
178
+ files, three `.obsidian/` JSON settings files, and `Units.base`. Existing assets must match their
179
+ recorded bytes or the newly generated bytes. This receipt does not change the public sidecar manifest
180
+ schema. Modified settings are preserved by refusing the export; the sentinel is not an override.
181
+
182
+ **Upgrading an older vault:** without a receipt, byte-identical generated assets can be adopted. Run
183
+ once against the same extraction to establish ownership before updating the index. If the index has
184
+ already changed, Woods may refuse the old sidecars because their ownership cannot be proved. Inspect
185
+ and back up the named files before moving them aside, or export into a new directory. Do not remove
186
+ personal content or manufacture a receipt to bypass a conflict.
187
+
188
+ Preflight prevents known destination conflicts from causing partial exports. This is not a multi-file
189
+ transaction or a lock against concurrent editors: a later I/O failure can leave some generated files
190
+ updated. Such failures report errors, suppress the sweep, and return zero completed-export counts;
191
+ inspect the vault before retrying. Avoid editing the destination while an export runs.
192
+
144
193
  Notes Woods manages carry `woods_managed: true` in their frontmatter. On each run, after all notes are
145
194
  written successfully, a **sweep** removes managed notes whose unit no longer exists, so deletions in
146
195
  your code propagate. Several guards make the sweep safe to point at a real vault:
@@ -150,21 +199,20 @@ your code propagate. Several guards make the sweep safe to point at a real vault
150
199
  - It refuses to delete more than 30% of managed notes at once (the signature of a partial extraction)
151
200
  unless `WOODS_OBSIDIAN_FORCE_PURGE=1` is set.
152
201
  - It resolves symlinks and confirms every deletion target is inside the vault root.
153
- - It is skipped entirely if any note failed to write that run (a stale note is harmless; a deleted
154
- reviewed note is not).
202
+ - It is skipped entirely if any candidate unit is unreadable or returns a mismatched identity,
203
+ or any note failed to write. `force_purge` does not bypass this incomplete-export guard.
155
204
 
156
- The same ownership check guards the `.obsidian/` config: Woods will **not** overwrite an existing,
157
- foreign `.obsidian/` folder, it leaves your Obsidian settings untouched and warns instead.
205
+ In a foreign vault, Woods leaves `.obsidian/` configuration untouched and skips the sweep. Notes and
206
+ sidecars are added only after the same per-file ownership preflight. In a Woods-owned vault, existing
207
+ configuration assets still require a matching ownership receipt or byte-identical generated content.
158
208
 
159
209
  ## Limitations
160
210
 
161
211
  - **Bases needs Obsidian ≥ 1.9** (≥ 1.10 for card/list views). The `.base` file is harmless on older
162
212
  versions, it simply doesn't render.
163
- - **`gem_source` units are not exported.** They aren't reachable through the index reader; only
164
- `rails_source` is covered by `include_framework`.
165
213
  - **Hand-edits diverge.** The vault is meant to be regenerated. Editing a note's properties in
166
214
  Obsidian rewrites its frontmatter, after which a re-export will overwrite your changes.
167
215
  - **Nested vaults ignore the shipped config.** Open the generated folder as its own vault for graph
168
216
  colors, Bases, and link-format settings to apply.
169
217
 
170
- See `docs/AGENT_GUIDE.md` for how an agent consumes the `_woods/` sidecar.
218
+ For MCP-based exploration instead of local vault files, see [the agent guide](AGENT_GUIDE.md).
@@ -1,5 +1,8 @@
1
1
  # Reading a published index from Ruby
2
2
 
3
+ For shell, Python, shipping snapshots, or direct JSON access, use the
4
+ [published filesystem layout contract](INDEX_LAYOUT.md).
5
+
3
6
  `Woods::PublishedIndex` is the stable, read-only API for tools that are not MCP clients: RuboCop cops, CI gate scripts, and the `woods:check:*` tasks. It needs no Rails, opens one published generation, and never moves off it for the life of the reader.
4
7
 
5
8
  ```ruby
@@ -24,6 +27,58 @@ Woods::PublishedIndex.open(Rails.root.join('tmp/woods')) do |index|
24
27
  end
25
28
  ```
26
29
 
30
+ ## Validate a published generation
31
+
32
+ ```ruby
33
+ require 'woods/resilience/index_validator'
34
+
35
+ report = Woods::Resilience::IndexValidator.new(index_dir: 'tmp/woods').validate
36
+ report.valid? # false when artifact or semantic graph errors were found
37
+ report.errors
38
+ report.warnings
39
+ ```
40
+
41
+ `woods:validate` uses the same checker. Supporting development versions validate
42
+ raw graph relationships as well as JSON, content hashes, and indexed files; see
43
+ the [semantic invariants and limits](INDEX_LAYOUT.md#semantic-graph-validation).
44
+ Errors retain typed identifiers and artifact paths. Writer-version and source-path
45
+ warnings keep their existing advisory behavior.
46
+
47
+ Each call resolves and pins one published generation for all checks. Concurrent
48
+ publication can advance the pointer, but retention cannot remove the payload
49
+ being validated. The next call sees the new generation. Locks release on success
50
+ and errors; malformed pointers fail instead of falling back to stale root files.
51
+ Legacy flat layouts remain supported but cannot provide immutable-generation
52
+ isolation against in-place writes. Bare type-directory fixtures without a
53
+ manifest retain structural-only validation.
54
+
55
+ The reusable `Woods::Resilience::GraphInvariantValidator` accepts raw string-keyed
56
+ `graph:` data and typed `index_entries:` and returns an array of errors without
57
+ modifying either input. Callers supplying raw data must keep it within one pinned
58
+ generation and verify the unit artifacts themselves, as `IndexValidator` does.
59
+ Validation is a read-only diagnostic, not automatic repair or proof of runtime
60
+ execution.
61
+
62
+ ## Manifest writer provenance
63
+
64
+ The published `manifest.json` records `woods_version`, a string naming the Woods gem version
65
+ that last published that manifest. Full extraction, changed incremental runs,
66
+ targeted refreshes, and the static Woods self-map write it. A no-op leaves the
67
+ published manifest and its version unchanged. Resolve the manifest through the
68
+ generation pointer, as Woods readers do.
69
+
70
+ Older manifests may omit the field or contain `null`; that means unknown.
71
+ `woods_status.index.woods_version` reports this value from the served manifest,
72
+ while `woods_status.server.version` identifies the running MCP reader. The reader
73
+ never substitutes its own version for missing writer provenance.
74
+
75
+ `woods:validate` warns when a present writer version is malformed or its major
76
+ version differs from the installed validator. The warning is advisory and does
77
+ not invalidate an otherwise structurally valid index. Re-run full extraction
78
+ when investigating a major-version mismatch. Matching versions do not certify
79
+ compatibility: an incremental publisher can retain units written by an older
80
+ version. This field does not replace the [upgrade procedure](UPGRADING_TO_2.md).
81
+
27
82
  ## One generation, pinned for the reader's whole life
28
83
 
29
84
  Unlike `Woods::MCP::IndexReader`, a `PublishedIndex` never refreshes between calls. It resolves one generation at `.new`/`.open` time and every fact it returns, `units`, the table map, `generation_number`, `external_dependency_checksum`, comes from that one generation for as long as the reader is open. There is no `reload` and no auto-refresh: open a new reader to see a later publish.
data/docs/README.md CHANGED
@@ -36,10 +36,13 @@ and stdio or Streamable HTTP endpoints directly.
36
36
  - [MCP tool cookbook](MCP_TOOL_COOKBOOK.md): scenario-based calls with parameters and expected response shapes.
37
37
  - [Console MCP setup](CONSOLE_MCP_SETUP.md): Console transports, blocked tables, credential scanning, redaction, SQL validation, and production safeguards.
38
38
  - [MCP HTTP transport](MCP_HTTP_TRANSPORT.md): shared/remote Index Server transport, authentication, origins, and protocol details.
39
+ - [Edit client adapters](CLIENT_HOOKS.md): opt-in Claude/OpenCode registration, complete path batches, and recovery.
39
40
  - [MCP worktree setup](MCP_WORKTREE_SETUP.md): register Woods correctly when agents work in linked git worktrees.
40
41
 
41
42
  ## Index lifecycle
42
43
 
44
+ - [Source freshness](SOURCE_FRESHNESS.md): verify dirty source against a served generation, establish a fresh-process baseline, and understand bounded unknown results.
45
+
43
46
  - [Retrieval guide](RETRIEVAL_GUIDE.md): configure embeddings and understand semantic retrieval, ranking, and token budgets.
44
47
  - [Embedding models](EMBEDDING_MODELS.md): choose and size local Ollama models.
45
48
  - [Upgrade to Woods 2.0](UPGRADING_TO_2.md): identifier changes, atomic payloads, durable-store reconciliation, and rollback.
@@ -48,8 +51,10 @@ and stdio or Streamable HTTP endpoints directly.
48
51
 
49
52
  - [Why Woods](WHY_WOODS.md): the problems runtime introspection solves.
50
53
  - [Internals](INTERNALS.md): extraction, publication, graph, storage, retrieval, and MCP components.
54
+ - [Ruby runtime trace enrichment](RUNTIME_TRACING.md): record observed Ruby callers and merge trace evidence into method units.
51
55
  - [Extractor reference](EXTRACTOR_REFERENCE.md): what each extractor produces and the edge cases it handles.
52
56
  - [Reading a published index from Ruby](PUBLISHED_INDEX.md): the `Woods::PublishedIndex` Ruby API for cops, gate scripts, and `woods:check:*` tasks (including the moved-message check).
57
+ - [Published index layout](INDEX_LAYOUT.md): the filesystem contract for non-Ruby readers, with Bash/jq and Python examples, retention pins, and snapshot-copy rules.
53
58
  - [Evaluation](EVALUATION.md): retrieval scoring, baselines, and the agent-level index on/off ablation.
54
59
  - [Backend matrix](BACKEND_MATRIX.md): implemented provider/store combinations and their operational requirements.
55
60
  - [Token benchmark](TOKEN_BENCHMARK.md): evidence behind Woods token-estimation defaults.
@@ -89,6 +94,8 @@ Use this map when changing behavior or documentation. Update the owner first; ot
89
94
  | Contributor policy | [CONTRIBUTING.md](../CONTRIBUTING.md) |
90
95
  | Coding-agent repository instructions | [AGENTS.md](https://github.com/lost-in-the/woods/blob/main/AGENTS.md) |
91
96
  | Non-MCP Ruby access to a published index | [PUBLISHED_INDEX.md](PUBLISHED_INDEX.md) |
97
+ | Generation-bound source evidence and fresh extraction | [SOURCE_FRESHNESS.md](SOURCE_FRESHNESS.md) |
98
+ | Published filesystem layout for external readers | [INDEX_LAYOUT.md](INDEX_LAYOUT.md) |
92
99
  | Evaluation harnesses | [EVALUATION.md](EVALUATION.md) |
93
100
 
94
101
  The current public surface is generated from 35 extractors. Counts and capability claims must match `.Codex/release-v2/surface-inventory.json`, which is generated from the code and verified in CI.