woods 2.0.0.beta2 → 2.0.0.beta4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (233) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +339 -1
  3. data/CONTRIBUTING.md +188 -12
  4. data/README.md +93 -174
  5. data/SECURITY.md +9 -6
  6. data/docs/AGENT_GUIDE.md +109 -8
  7. data/docs/AGENT_SETUP.md +98 -7
  8. data/docs/BACKEND_MATRIX.md +25 -0
  9. data/docs/CLIENT_HOOKS.md +111 -0
  10. data/docs/CONFIGURATION_REFERENCE.md +267 -16
  11. data/docs/CONSOLE_MCP_SETUP.md +80 -7
  12. data/docs/DOCKER_SETUP.md +22 -3
  13. data/docs/EVALUATION.md +464 -1
  14. data/docs/EXTRACTOR_REFERENCE.md +45 -6
  15. data/docs/FAQ.md +11 -12
  16. data/docs/GETTING_STARTED.md +17 -5
  17. data/docs/INCREMENTAL_EXTRACTION.md +147 -7
  18. data/docs/INDEX_LAYOUT.md +382 -0
  19. data/docs/INTERNALS.md +7 -2
  20. data/docs/MCP_SERVERS.md +276 -5
  21. data/docs/MCP_TOOL_COOKBOOK.md +37 -22
  22. data/docs/MCP_WORKTREE_SETUP.md +43 -83
  23. data/docs/NOTION_INTEGRATION.md +13 -0
  24. data/docs/OBSIDIAN_INTEGRATION.md +57 -9
  25. data/docs/PUBLISHED_INDEX.md +72 -0
  26. data/docs/README.md +7 -0
  27. data/docs/RETRIEVAL_GUIDE.md +273 -12
  28. data/docs/RUNTIME_TRACING.md +71 -0
  29. data/docs/SOURCE_FRESHNESS.md +143 -0
  30. data/docs/TROUBLESHOOTING.md +129 -18
  31. data/docs/UNBLOCKED_INTEGRATION.md +25 -0
  32. data/docs/UPGRADING_TO_2.md +48 -22
  33. data/docs/WATCH_DAEMON.md +277 -67
  34. data/exe/woods-agent-config +6 -0
  35. data/exe/woods-extract +5 -0
  36. data/exe/woods-hook-context +6 -0
  37. data/exe/woods-mcp-start +14 -9
  38. data/lib/generators/woods/pgvector_generator.rb +8 -2
  39. data/lib/generators/woods/templates/woods.rb.tt +1 -3
  40. data/lib/tasks/woods.rake +47 -397
  41. data/lib/woods/agent_configuration/applier.rb +135 -0
  42. data/lib/woods/agent_configuration/cli.rb +101 -0
  43. data/lib/woods/agent_configuration/cli_options.rb +29 -0
  44. data/lib/woods/agent_configuration/document.rb +105 -0
  45. data/lib/woods/agent_configuration/error.rb +7 -0
  46. data/lib/woods/agent_configuration/launcher.rb +75 -0
  47. data/lib/woods/agent_configuration/layout.rb +72 -0
  48. data/lib/woods/agent_configuration/managed_section.rb +62 -0
  49. data/lib/woods/agent_configuration/plan.rb +98 -0
  50. data/lib/woods/agent_configuration/plan_diff.rb +38 -0
  51. data/lib/woods/agent_configuration/planned_files.rb +61 -0
  52. data/lib/woods/agent_configuration/planner.rb +63 -0
  53. data/lib/woods/agent_configuration/planner_validation.rb +77 -0
  54. data/lib/woods/agent_configuration/preflight.rb +100 -0
  55. data/lib/woods/agent_configuration/recovery.rb +49 -0
  56. data/lib/woods/ast/node.rb +2 -0
  57. data/lib/woods/ast/parser.rb +38 -5
  58. data/lib/woods/builder.rb +21 -5
  59. data/lib/woods/cache/cache_middleware.rb +28 -7
  60. data/lib/woods/cache/cache_store.rb +4 -5
  61. data/lib/woods/change_set.rb +5 -4
  62. data/lib/woods/console/credential_index.rb +20 -2
  63. data/lib/woods/console/credential_scanner.rb +18 -17
  64. data/lib/woods/console/credential_scanner_registry.rb +36 -0
  65. data/lib/woods/console/dispatch_pipeline.rb +7 -0
  66. data/lib/woods/console/embedded_executor.rb +32 -10
  67. data/lib/woods/console/encrypted_credential_snapshot.rb +16 -0
  68. data/lib/woods/console/rack_middleware.rb +22 -13
  69. data/lib/woods/console/server.rb +18 -16
  70. data/lib/woods/console/sql_noise_stripper.rb +9 -7
  71. data/lib/woods/console/sql_table_scanner.rb +47 -7
  72. data/lib/woods/console/sql_validator.rb +49 -9
  73. data/lib/woods/console/sqlite_read_guard.rb +46 -0
  74. data/lib/woods/coordination/pipeline_lock.rb +3 -2
  75. data/lib/woods/dependency_graph.rb +65 -13
  76. data/lib/woods/embedding/corpus.rb +94 -0
  77. data/lib/woods/embedding/indexer.rb +114 -60
  78. data/lib/woods/embedding/openai.rb +17 -6
  79. data/lib/woods/evaluation/ablation_executor.rb +6 -1
  80. data/lib/woods/evaluation/ablation_timed_executor.rb +22 -4
  81. data/lib/woods/export/typed_reader.rb +56 -0
  82. data/lib/woods/extractor.rb +277 -149
  83. data/lib/woods/extractors/action_cable_extractor.rb +3 -1
  84. data/lib/woods/extractors/behavioral_profile.rb +9 -7
  85. data/lib/woods/extractors/caching_extractor.rb +3 -1
  86. data/lib/woods/extractors/concern_extractor.rb +64 -6
  87. data/lib/woods/extractors/configuration_extractor.rb +7 -3
  88. data/lib/woods/extractors/controller_extractor.rb +13 -4
  89. data/lib/woods/extractors/database_view_extractor.rb +3 -1
  90. data/lib/woods/extractors/declared_parent.rb +55 -0
  91. data/lib/woods/extractors/decorator_extractor.rb +3 -1
  92. data/lib/woods/extractors/engine_extractor.rb +3 -1
  93. data/lib/woods/extractors/event_extractor.rb +4 -2
  94. data/lib/woods/extractors/factory_extractor.rb +3 -1
  95. data/lib/woods/extractors/graphql_extractor.rb +10 -13
  96. data/lib/woods/extractors/i18n_extractor.rb +3 -1
  97. data/lib/woods/extractors/job_extractor.rb +6 -19
  98. data/lib/woods/extractors/lib_extractor.rb +13 -9
  99. data/lib/woods/extractors/mailer_extractor.rb +26 -15
  100. data/lib/woods/extractors/manager_extractor.rb +3 -1
  101. data/lib/woods/extractors/method_parameters.rb +53 -0
  102. data/lib/woods/extractors/middleware_argument.rb +65 -0
  103. data/lib/woods/extractors/middleware_extractor.rb +9 -3
  104. data/lib/woods/extractors/migration_extractor.rb +3 -1
  105. data/lib/woods/extractors/model_extractor.rb +26 -34
  106. data/lib/woods/extractors/package_extractor.rb +24 -4
  107. data/lib/woods/extractors/phlex_extractor.rb +3 -1
  108. data/lib/woods/extractors/policy_extractor.rb +3 -1
  109. data/lib/woods/extractors/poro_extractor.rb +13 -9
  110. data/lib/woods/extractors/pundit_extractor.rb +3 -1
  111. data/lib/woods/extractors/rails_source_extractor.rb +4 -2
  112. data/lib/woods/extractors/rake_task_extractor.rb +4 -2
  113. data/lib/woods/extractors/route_extractor.rb +3 -1
  114. data/lib/woods/extractors/route_helper_resolver.rb +10 -33
  115. data/lib/woods/extractors/scheduled_job_extractor.rb +41 -15
  116. data/lib/woods/extractors/serializer_extractor.rb +4 -2
  117. data/lib/woods/extractors/service_extractor.rb +3 -1
  118. data/lib/woods/extractors/shared_dependency_scanner.rb +2 -2
  119. data/lib/woods/extractors/shared_utility_methods.rb +48 -19
  120. data/lib/woods/extractors/source_nesting.rb +1 -1
  121. data/lib/woods/extractors/state_machine_extractor.rb +3 -1
  122. data/lib/woods/extractors/test_mapping_extractor.rb +3 -1
  123. data/lib/woods/extractors/validator_extractor.rb +3 -1
  124. data/lib/woods/extractors/view_component_extractor.rb +3 -1
  125. data/lib/woods/extractors/view_template_extractor.rb +3 -1
  126. data/lib/woods/gem_mapper.rb +2 -0
  127. data/lib/woods/git_history.rb +116 -0
  128. data/lib/woods/graph_analyzer.rb +35 -6
  129. data/lib/woods/hooks/context_cli.rb +54 -0
  130. data/lib/woods/hooks/context_event.rb +88 -0
  131. data/lib/woods/hooks/context_hint.rb +73 -0
  132. data/lib/woods/hooks/context_impact.rb +77 -0
  133. data/lib/woods/hooks/context_output.rb +47 -0
  134. data/lib/woods/hooks/context_state.rb +102 -0
  135. data/lib/woods/hooks/refresh.rb +79 -0
  136. data/lib/woods/hooks/rule_projection.rb +78 -0
  137. data/lib/woods/input_rules.rb +19 -0
  138. data/lib/woods/mcp/bearer_auth.rb +22 -13
  139. data/lib/woods/mcp/bootstrapper.rb +79 -4
  140. data/lib/woods/mcp/config_resolver.rb +2 -1
  141. data/lib/woods/mcp/index_reader.rb +334 -162
  142. data/lib/woods/mcp/initialization_guidance.rb +27 -0
  143. data/lib/woods/mcp/origin_guard.rb +17 -9
  144. data/lib/woods/mcp/published_lexical_retriever.rb +115 -0
  145. data/lib/woods/mcp/renderers/markdown_renderer.rb +22 -9
  146. data/lib/woods/mcp/renderers/plain_renderer.rb +18 -8
  147. data/lib/woods/mcp/search_results.rb +74 -0
  148. data/lib/woods/mcp/server.rb +178 -63
  149. data/lib/woods/mcp/tool_contract.rb +3 -1
  150. data/lib/woods/mcp/tool_response_renderer.rb +41 -0
  151. data/lib/woods/mcp/traversal_evidence.rb +113 -0
  152. data/lib/woods/mcp/traversal_evidence_index.rb +100 -0
  153. data/lib/woods/mcp/traversal_evidence_page.rb +41 -0
  154. data/lib/woods/mcp/traversal_evidence_text.rb +52 -0
  155. data/lib/woods/mcp/traversal_response.rb +22 -0
  156. data/lib/woods/notion/exporter.rb +56 -17
  157. data/lib/woods/obsidian/destination_plan.rb +98 -0
  158. data/lib/woods/obsidian/name_mapper.rb +19 -3
  159. data/lib/woods/obsidian/note_builder.rb +19 -10
  160. data/lib/woods/obsidian/vault_exporter.rb +88 -32
  161. data/lib/woods/operator/pipeline_guard.rb +18 -13
  162. data/lib/woods/path_dispatcher.rb +13 -6
  163. data/lib/woods/payload_store.rb +27 -26
  164. data/lib/woods/published_index/typed_unit_reader.rb +40 -3
  165. data/lib/woods/published_index.rb +2 -2
  166. data/lib/woods/railtie.rb +3 -3
  167. data/lib/woods/railtie_support.rb +12 -12
  168. data/lib/woods/rake_helpers.rb +382 -0
  169. data/lib/woods/resilience/graph_invariant_validator/membership_checks.rb +71 -0
  170. data/lib/woods/resilience/graph_invariant_validator/node_checks.rb +61 -0
  171. data/lib/woods/resilience/graph_invariant_validator/reverse_relationship_checks.rb +46 -0
  172. data/lib/woods/resilience/graph_invariant_validator.rb +119 -0
  173. data/lib/woods/resilience/index_validator/graph_checks.rb +80 -0
  174. data/lib/woods/resilience/index_validator.rb +112 -23
  175. data/lib/woods/retrieval/context_assembler.rb +50 -15
  176. data/lib/woods/retrieval/lexical_assembler.rb +84 -0
  177. data/lib/woods/retrieval/lexical_index.rb +120 -0
  178. data/lib/woods/retrieval/ranker.rb +4 -2
  179. data/lib/woods/retrieval/scope.rb +108 -0
  180. data/lib/woods/retrieval/scoped_graph_store.rb +32 -0
  181. data/lib/woods/retrieval/scoped_vector_store.rb +55 -0
  182. data/lib/woods/retrieval/search_executor.rb +86 -27
  183. data/lib/woods/retrieval/source_evidence.rb +200 -0
  184. data/lib/woods/retriever.rb +98 -22
  185. data/lib/woods/ruby_analyzer/trace_enricher.rb +77 -38
  186. data/lib/woods/session_tracer/file_store.rb +6 -1
  187. data/lib/woods/session_tracer/middleware.rb +10 -12
  188. data/lib/woods/session_tracer/redis_store.rb +22 -6
  189. data/lib/woods/session_tracer/session_flow_assembler.rb +23 -17
  190. data/lib/woods/session_tracer/solid_cache_coordination.rb +6 -4
  191. data/lib/woods/session_tracer/unit_resolver.rb +63 -0
  192. data/lib/woods/source_inputs/consumer_errors.rb +31 -0
  193. data/lib/woods/source_inputs/handoff.rb +102 -0
  194. data/lib/woods/source_inputs/launcher.rb +157 -0
  195. data/lib/woods/source_inputs/manifest.rb +124 -0
  196. data/lib/woods/source_inputs/private_key.rb +55 -0
  197. data/lib/woods/source_inputs/scanner.rb +171 -0
  198. data/lib/woods/source_inputs/scopes.rb +71 -0
  199. data/lib/woods/source_inputs/session.rb +214 -0
  200. data/lib/woods/source_inputs/status.rb +84 -0
  201. data/lib/woods/source_inputs/verifier.rb +107 -0
  202. data/lib/woods/storage/metadata_store.rb +25 -25
  203. data/lib/woods/storage/pgvector.rb +35 -10
  204. data/lib/woods/storage/qdrant.rb +17 -7
  205. data/lib/woods/storage/vector_store.rb +18 -6
  206. data/lib/woods/tasks.rb +3 -2
  207. data/lib/woods/temporal/json_snapshot_store.rb +58 -9
  208. data/lib/woods/unblocked/exporter.rb +59 -70
  209. data/lib/woods/version.rb +1 -1
  210. data/lib/woods/watch/boot_snapshot.rb +52 -0
  211. data/lib/woods/watch/daemon.rb +154 -32
  212. data/lib/woods/watch/listen_watcher.rb +4 -0
  213. data/lib/woods/watch/polling_watcher.rb +5 -1
  214. data/lib/woods/watch/status.rb +20 -15
  215. data/lib/woods/watch/tree_scan.rb +21 -13
  216. data/lib/woods/watch/watcher.rb +4 -1
  217. data/lib/woods.rb +50 -11
  218. data/plugin/.claude-plugin/plugin.json +1 -1
  219. data/plugin/hooks/adapters/normalize.jq +15 -0
  220. data/plugin/hooks/adapters/normalize.rb +63 -0
  221. data/plugin/hooks/hooks.json +20 -0
  222. data/plugin/hooks/woods-context.sh +50 -0
  223. data/plugin/hooks/woods-input-rules.sh +159 -0
  224. data/plugin/hooks/woods-opencode.mjs +65 -0
  225. data/plugin/hooks/woods-post-edit.sh +2 -225
  226. data/plugin/hooks/woods-refresh.sh +260 -0
  227. data/plugin/hooks/woods-session-start.sh +47 -55
  228. data/plugin/skills/woods-agent-enable/SKILL.md +19 -0
  229. data/plugin/skills/woods-diagnose/SKILL.md +319 -1
  230. data/plugin/skills/woods-investigate/SKILL.md +145 -0
  231. data/plugin/skills/woods-mcp-config/SKILL.md +90 -2
  232. data/plugin/skills/woods-setup/SKILL.md +110 -6
  233. metadata +87 -5
@@ -2,7 +2,7 @@
2
2
 
3
3
  Woods ships **35 extractor classes** producing **39 distinct unit types**: one for each meaningful category of Rails code. This doc covers what each extractor captures, how to configure them, and the shape of the data they produce.
4
4
 
5
- > **Counts explained.** `lib/woods/extractors/` contains 42 files: 35 extractor classes (each ending in `_extractor.rb`) plus 7 supporting utilities (`shared_utility_methods`, `shared_dependency_scanner`, `callback_analyzer`, `behavioral_profile`, `route_helper_resolver`, `ast_source_extraction`, `source_nesting`). The 39 unit types comes from some extractors emitting multiple categories, `GraphQLExtractor` alone produces four (`graphql_type`, `graphql_mutation`, `graphql_resolver`, `graphql_query`), and `RailsSourceExtractor` produces both `rails_source` and `gem_source`. Supporting utilities enrich existing extractors (callback side-effects, behavioral config, AST-based source slicing, nested-namespace resolution) but are not themselves extractors and do not appear in the unit type enumeration. The authoritative mapping is `Woods::Extractor::TYPE_TO_EXTRACTOR_KEY` in `lib/woods/extractor.rb`.
5
+ > **Counts explained.** `lib/woods/extractors/` contains 35 extractor classes (each ending in `_extractor.rb`) plus supporting utilities such as `shared_utility_methods`, `shared_dependency_scanner`, `callback_analyzer`, `behavioral_profile`, `route_helper_resolver`, `ast_source_extraction`, `source_nesting`, and `declared_parent`. The 39 unit types comes from some extractors emitting multiple categories, `GraphQLExtractor` alone produces four (`graphql_type`, `graphql_mutation`, `graphql_resolver`, `graphql_query`), and `RailsSourceExtractor` produces both `rails_source` and `gem_source`. Supporting utilities enrich existing extractors (callback side-effects, behavioral config, AST-based source slicing, nested-namespace resolution) but are not themselves extractors and do not appear in the unit type enumeration. The authoritative mapping is `Woods::Extractor::TYPE_TO_EXTRACTOR_KEY` in `lib/woods/extractor.rb`.
6
6
 
7
7
  ---
8
8
 
@@ -60,9 +60,14 @@ Every extractor returns `Array<ExtractedUnit>`. An `ExtractedUnit` is a self-con
60
60
 
61
61
  **Key details:**
62
62
  - Uses `ActiveRecord::Base.descendants` for discovery (runtime introspection, not static parsing)
63
+ - Named, source-defined app model mixins (for example `Card::Pinnable` in `app/models/card/pinnable.rb`) resolve through runtime source locations, with conventional concern paths as fallbacks. Included mixins also receive `:concern` units, so their actual files map to the includer through dependency edges. Gem-owned modules stay outside this discovery. Conventional concern files retain their existing identity even when nested helpers share the same file. Outside those directories, each included runtime mixin receives its own concern identity even when several share a source file; editing that file refreshes every includer.
63
64
  - Inlines concerns: all `include FooConcern` references are resolved and the concern source is appended to `source_code`. Inlined concern names are recorded in `metadata[:inlined_concerns]`
64
- - Extracts all 19 callback types: `before_validation`, `after_validation`, `before_save`, `after_save`, `around_save`, `before_create`, `after_create`, `around_create`, `before_update`, `after_update`, `around_update`, `before_destroy`, `after_destroy`, `around_destroy`, `after_commit`, `after_rollback`, `after_initialize`, `after_find`, `after_touch`
65
+ - Reads Rails' per-event callback chains (`_save_callbacks`, `_create_callbacks`, and the other lifecycle events), preserving each chain's order and each entry's `kind`, filter and conditions. The public `type` combines kind and event, such as `before_save` or `after_create`; Rails' separate `before_commit` event is reported as `before_commit`, not `before_before_commit`. `callback_count` equals the emitted callback list's length. The list includes framework-registered callbacks; it is runtime metadata, not an application-only filter or a cross-event execution trace. Regenerate the index after upgrading to pick up corrected callback metadata.
66
+ - Proc/lambda filters, including Rails-generated association callbacks, use stable source-site labels in both metadata and callback chunks: `#<Proc app/models/post.rb:12>` (or `lambda`). App paths are relative to `Rails.root`; external paths are retained and native procs use `native`. Rails 6's numeric filter identity is resolved through `raw_filter`. These labels describe location and callable kind, not captured closure state; callbacks are never executed. Model condition labels retain their existing format.
67
+ - Default callback-object representations omit process addresses: an instance becomes `#<CleanupCallback>`, an anonymous class becomes `#<Class>`, and its instance becomes `#<#<Class>>`. Anonymous namespace prefixes are normalized too (for example, `#<Module>::CleanupCallback`). Named classes and custom `to_s` labels retain their text. These labels do not distinguish arbitrary object state; separate registered callbacks remain separate entries even when their descriptive labels match. Controller object-filter formatting is unchanged.
68
+ - Direct Proc/lambda validation option values (for example inclusion/exclusion membership or message callables) use the same stable kind/source-site labels without execution. Validation order and duplicates, condition formats (`if`, `unless`, `on`), and non-Proc values are unchanged. Nested arrays/hashes are not recursively normalized, and labels do not serialize captured closure state. Run a full extraction after upgrading to refresh retained validation metadata; the stored schema is unchanged.
65
69
  - Callback side-effects are analyzed via `CallbackAnalyzer`: detects columns written (`self.col =`), jobs enqueued (`perform_later`), and services called
70
+ - Reflects model class and instance methods after reading the schema, so Rails schema-loading optimizations produce the same method metadata in cold and warmed runs. Application-defined constructors remain visible; Rails versions that install an optimized singleton `new` during schema loading consistently include it in `class_methods`.
66
71
  - Automatically skips HABTM join models and anonymous classes
67
72
  - Chunks every model into semantic sections: `:summary`, `:associations`, `:callbacks`, `:validations`, `:scopes`, `:methods`
68
73
  - **Runtime-generated method detection:** Because extraction runs inside a booted Rails process, `instance_methods(false)` captures every method Rails generates dynamically, enum predicates (`status_active?`, `status_pending?`), association builders (`build_profile`, `create_line_item!`), attribute accessors, and dynamically registered scopes. Static analysis tools cannot see these methods because they only exist after Rails processes the DSL declarations at boot time
@@ -158,6 +163,8 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
158
163
  - Route context is inlined in `source_code` as a comment header, not just in metadata
159
164
  - Chunks per-action: each action becomes a `:action` chunk with its applicable filters and route
160
165
  - Metadata includes permitted params (strong parameters), response formats, and applied filters per action
166
+ - Inline callbacks use stable source-site labels in filter metadata, controller annotations and action chunks: `#<Proc app/controllers/posts_controller.rb:12>` (or `lambda`). The controller filter metadata and annotations use the same labels for `if`/`unless` procs. App paths are relative to `Rails.root`; external paths are retained and native procs use `native`. Labels describe the callable location and kind, not captured closure state, and never execute callbacks.
167
+ - Route helper resolution accepts every live named controller/action route, including `file_path`, `image_url`, `download_path`, and `root_path`. Unknown filesystem/asset helpers produce no edge; matching names are conservative source references, not proof a call executes.
161
168
  - Extracts `redirect_to` navigation edges: named route helpers (`posts_path`, `users_url`) are resolved to controller targets via `RouteHelperResolver`, producing `:redirect_to` dependency edges (gated by `extract_navigation_edges` config)
162
169
 
163
170
  **Edge cases:**
@@ -194,6 +201,7 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
194
201
  - Scans: `app/services`, `app/interactors`, `app/operations`, `app/commands`, `app/use_cases`
195
202
  - Extracts public entry points (`call`, `perform`, `execute`, `run`), custom error classes, and dependency references
196
203
  - File-based discovery (not class introspection), so it catches services with non-standard superclasses
204
+ - `initialize_params` describes declared names, default presence and keyword status from Ruby syntax. Nested/comma-bearing defaults are not evaluated or treated as parameters; named rest, keyword-rest and block parameters retain their names. Anonymous forwarding has no name to report; malformed source produces an empty parameter list.
197
205
 
198
206
  **Example output (abbreviated):**
199
207
 
@@ -218,6 +226,7 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
218
226
  **Key details:**
219
227
  - Scans: `app/jobs`, `app/workers`, `app/sidekiq`
220
228
  - Extracts queue name, retry configuration, concurrency options, perform method arguments, and callbacks
229
+ - `perform_params` uses the same syntax-aware signature parsing as service initializers and preserves its `name`, `splat` (`single`/`double`/null), and `has_default` fields. Keyword defaults do not invent additional argument names.
221
230
  - Records what triggers this job (reverse lookup via dependency graph after extraction)
222
231
  - Supports both ActiveJob and Sidekiq native workers
223
232
 
@@ -243,9 +252,14 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
243
252
  **What it captures:** ActionMailer classes with their mailer actions, defaults, template paths, callbacks, and helper usage.
244
253
 
245
254
  **Key details:**
246
- - Discovers via class introspection (`ActionMailer::Base.descendants`)
255
+ - Discovers `ApplicationMailer.descendants` when that class exists, otherwise `ActionMailer::Base.descendants`; an app without ActionMailer contributes no mailer units.
256
+ - Discovery and direct extraction accept only mailers backed by an existing app-owned source file, excluding dependency mailers and fabricated convention paths.
247
257
  - Each mailer action corresponds to an email template, template paths are recorded in metadata
248
258
  - Extracts `default from:`, `layout`, and per-action subject patterns
259
+ - Action names are sorted consistently in metadata, the generated header, template discovery and action chunks. Callback chain order and duplicate registrations are preserved.
260
+ - Direct Proc-valued defaults and Proc callback filters use source-location/kind labels without executing them; application paths are relative to `Rails.root`. Default containers, literal strings and non-Proc values retain their existing types. This does not serialize closure captures or recursively normalize arbitrary nested objects.
261
+ - Object callback filters use descriptive labels with addresses removed only from Ruby's default representation, including anonymous classes and namespaces. Custom labels and literal hexadecimal text are preserved. Labels do not serialize callback object state, and extraction never invokes callbacks.
262
+ - After upgrading, run a full extraction to refresh retained mailer units. Stabilized headers and callable labels can cause a one-time source-hash change; stored index schemas are unchanged.
249
263
 
250
264
  ---
251
265
 
@@ -294,7 +308,8 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
294
308
 
295
309
  **Key details:**
296
310
  - Extracts the entire stack as one unit (not one per middleware)
297
- - Records middleware class names, insertion order, and any initialization arguments
311
+ - Records middleware class names, insertion order, and initialization arguments as readable strings
312
+ - Argument rendering preserves literal strings and nested array/hash configuration. Procs use source locations; anonymous classes (including Ruby temporary names used by Rails executors/reloaders) use parent names and method source locations. Opaque objects using Ruby's default `to_s` are represented by class, without walking private runtime state. Custom `to_s` output is preserved, so application-defined nondeterministic renderers can still vary. Closure captures and opaque object internals are not serialized.
298
313
  - No per-file mapping, so incremental re-extraction re-runs `MiddlewareExtractor` wholesale when `config/application.rb`, `Gemfile.lock`, or a file under `config/initializers`/`config/environments` changes
299
314
 
300
315
  ---
@@ -331,6 +346,7 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
331
346
  **Key details:**
332
347
  - File-based scanning, no Rails boot needed for the actual file reading
333
348
  - Records which partials a template renders and which instance variables it expects
349
+ - Loads the runtime route collection before caching named helpers, including Rails lazy route sets. Fresh-process incremental view extraction resolves the same navigation targets as full extraction.
334
350
  - Extracts navigation dependencies: `link_to` and `form_with`/`form_for` calls using `_path`/`_url` route helpers are resolved to controller targets via `RouteHelperResolver`
335
351
  - Navigation edges use `:link_to` and `:form_action` via types in the dependency array
336
352
  - Gated by `extract_navigation_edges` config (default: true)
@@ -371,6 +387,7 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
371
387
  - Scans `app/models` for files that don't define an `ActiveRecord::Base` descendant
372
388
  - Common examples: value objects, form objects placed in `app/models`, domain structs
373
389
  - Excludes concerns (those go to ConcernExtractor)
390
+ - `parent_class` and the generated Parent annotation describe the selected unit declaration only. Nested or sibling classes cannot supply its parent. An implicit `Object` parent, a dynamic superclass expression, or unparseable source produces `nil`; explicit constant-path parents retain their source names.
374
391
 
375
392
  ---
376
393
 
@@ -410,6 +427,7 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
410
427
  - Produces unit types: `graphql_type`, `graphql_mutation`, `graphql_resolver`, `graphql_query`
411
428
  - Extracts field metadata (types, descriptions, complexity, arguments), authorization patterns (Pundit, CanCan, `authorized?`), and dependencies on models/services
412
429
  - Since all GraphQL units come from one extractor, incremental re-extraction handles them via `extract_graphql_file`
430
+ - `parent_class` and summary chunks describe the selected declaration's explicit constant-path superclass, preserving its written qualification. Nested or sibling declarations and literal text cannot supply a parent. Implicit Object, module interfaces, dynamic superclass expressions, unavailable source, and invalid source have no declared parent (`null` metadata; `unknown` in summaries). This is source declaration metadata, not resolved runtime ancestry.
413
431
 
414
432
  **Example output (abbreviated):**
415
433
 
@@ -469,6 +487,7 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
469
487
  **Key details:**
470
488
  - Identifier is the package directory relative to `Rails.root` (`.` for the root package), the same name Packwerk uses
471
489
  - Honors `packwerk.yml` `package_paths` and `exclude`; without one, `**/` with the Packwerk default excludes (`bin`, `node_modules`, `script`, `tmp`, `vendor`)
490
+ - With those exact defaults, excluded top-level directories are pruned before discovery descends into them, so an index or snapshots under `tmp/` do not add package-scan work. Custom patterns or exclusions retain their configured glob behavior. Hidden directories and symlink directories are not recursively followed by default.
472
491
  - `metadata`: `name`, `dependencies` (sorted), `enforce_dependencies` (`true`, `false`, or `"strict"`), `enforce_privacy`, `layer` (pks), `public_path`, `owner`
473
492
  - Each declared dependency becomes a `{ type: :package, target: <name>, via: :package_dependency }` edge
474
493
  - Package membership on other units (`metadata[:package]`, below) does not depend on how a unit was discovered: any registered unit with a file path under a package root is annotated. The undeclared cross-package edge report remains a follow-up, not this extractor. Discovery is the separate open gap: a pack-resident file-based unit is not yet found by `PathDispatcher` when only its `package.yml` changes (follow-up B-175), so it carries no membership only because it has no unit at all yet, not because membership skips it
@@ -530,7 +549,17 @@ Every app-owned unit under a package root carries `metadata[:package]` with the
530
549
  **Key details:**
531
550
  - Reads: `config/recurring.yml` (Solid Queue), `config/sidekiq_cron.yml` (Sidekiq Cron), `config/schedule.rb` (Whenever)
532
551
  - Extracts job class name, cron expression, queue, and any arguments
533
- - File-based (static read, no Rails introspection needed)
552
+ - Resolves trusted application `recurring.yml` through Rails' configuration loader,
553
+ including ERB, filename-relative `require_relative`, and YAML aliases. On Rails
554
+ 6.0 (before that loader existed), evaluates ERB with its filename and retains
555
+ safe YAML loading of scalars, hashes, arrays and symbols. ERB runs application
556
+ code in the extraction process; index only applications you trust.
557
+ - Environment-wrapped task maps select the current Rails environment, including
558
+ custom names; an absent environment falls back to the first section, while an
559
+ explicitly empty section stays empty. Flat task maps remain supported.
560
+ - Sidekiq-Cron remains safe-loaded YAML; Whenever remains a static DSL scan.
561
+ Invalid YAML/ERB, missing required files and runtime configuration errors are
562
+ logged and omit that schedule file. Source remains the original file text.
534
563
  - No per-file mapping, so incremental re-extraction re-runs `ScheduledJobExtractor` wholesale whenever one of the schedule files above changes
535
564
 
536
565
  ---
@@ -655,6 +684,7 @@ Every app-owned unit under a package root carries `metadata[:package]` with the
655
684
  **Key details:**
656
685
  - Excludes `lib/tasks/` (covered by RakeTaskExtractor) and `lib/generators/`
657
686
  - File-based scanning; no assumption about class hierarchy
687
+ - `parent_class` and the generated Parent annotation describe the selected unit declaration only. Nested or sibling classes cannot supply its parent. An implicit `Object` parent, a dynamic superclass expression, or unparseable source produces `nil`; explicit constant-path parents retain their source names.
658
688
 
659
689
  ---
660
690
 
@@ -711,7 +741,16 @@ When written to disk, units also include:
711
741
 
712
742
  ### Git enrichment fields (`metadata[:git]`)
713
743
 
714
- If the host app is a git repo, the following are added to `metadata[:git]` after extraction:
744
+ If the host app is a git repo, the following are added to `metadata[:git]` after extraction.
745
+ History is limited to commits reachable from `HEAD` in the past 365 days,
746
+ including merged branch history. Unmerged branches, remote refs, and tool
747
+ checkpoint refs do not contribute. Commands run against the application root;
748
+ when `WOODS_GIT_DIR` is set, `HEAD` belongs to that explicitly selected git
749
+ directory, which may differ from a linked worktree's HEAD.
750
+
751
+ After upgrading from a version that included all refs, run a full
752
+ `woods:extract` to replace previously published git metadata. Incremental
753
+ extraction refreshes only the units it rewrites.
715
754
 
716
755
  | Field | Description |
717
756
  |-------|-------------|
data/docs/FAQ.md CHANGED
@@ -274,7 +274,13 @@ Verify the active v2 generation with `docker compose exec app bundle exec rake w
274
274
 
275
275
  ### How do I configure the Console Server with Docker?
276
276
 
277
- First set `config.console_mcp_enabled = true` in the Rails initializer after reviewing the live-data trust boundary. Stdio does not send a bearer token, but production Rails boot still requires a configured `console_mcp_token` of at least 32 characters whenever Console is enabled. Supply `WOODS_CONSOLE_MCP_TOKEN` through the container's secret mechanism; see [Console MCP setup](CONSOLE_MCP_SETUP.md#option-a-stdio-via-rake-recommended). Then, for the embedded mode (9 Tier 1 tools), point the MCP client at `docker compose exec -T` so Compose does not allocate a pseudo-TTY:
277
+ First enable the master `console_mcp_enabled` switch after reviewing the
278
+ live-data trust boundary. For stdio-only use, explicitly set
279
+ `console_mcp_http_enabled = false`; HTTP remains enabled by default for
280
+ compatibility and requires its bearer token while enabled. Follow
281
+ [Console MCP setup](CONSOLE_MCP_SETUP.md#option-a-stdio-via-rake-recommended)
282
+ for the configuration. Then point the embedded-mode client (9 Tier 1 tools)
283
+ at `docker compose exec -T` so Compose does not allocate a pseudo-TTY:
278
284
 
279
285
  ```json
280
286
  {
@@ -414,7 +420,7 @@ When you run `rake woods:embed`, Woods generates embedding vectors for each extr
414
420
  Several options for tuning retrieval:
415
421
 
416
422
  - **Increase `max_context_tokens`** to include more units per query (at the cost of larger LLM context).
417
- - **Lower `similarity_threshold`** (default 0.7) to include less similar results.
423
+ - **Use explicit retrieval scopes** and inspect ranking evidence. `similarity_threshold` is deprecated and does not filter results; see [retrieval tuning](RETRIEVAL_GUIDE.md#tuning).
418
424
  - **Enable framework sources** (`include_framework_sources: true`) if Rails internals are relevant to your queries.
419
425
  - **Use retrieval feedback only in a custom embedded server** that wires a feedback store. The normal packaged executable does not register feedback tools.
420
426
 
@@ -442,16 +448,9 @@ Snapshots prefer their own SQLite database (`woods.sqlite3` in the output direct
442
448
 
443
449
  The session tracer is middleware that records which Rails actions are invoked during a browser session, assembles the relevant extracted units, and makes that context available via the `session_trace` MCP tool. It is useful for giving an AI tool accurate context about what code path was active during a specific user interaction.
444
450
 
445
- Session tracing is disabled by default. To enable it:
446
-
447
- ```ruby
448
- config.session_tracer_enabled = true
449
- config.session_store = Woods::SessionTracer::FileStore.new(
450
- Rails.root.join('tmp/session_traces')
451
- )
452
- ```
453
-
454
- The `session_store` option is required, there is no default store.
451
+ Session tracing is disabled by default and requires an explicit `session_store`.
452
+ Follow the [canonical configuration example](CONFIGURATION_REFERENCE.md#session-tracer-options)
453
+ for store construction and review trace retention and access controls before enabling it.
455
454
 
456
455
  ---
457
456
 
@@ -12,7 +12,19 @@ If an agent will perform the installation, use the safety and handoff checklist
12
12
 
13
13
  ## 1. Install the gem
14
14
 
15
- Add Woods to the development group:
15
+ Use the [README release table](../README.md) to choose a version, then confirm
16
+ that **exact version is published** on the [RubyGems versions page](https://rubygems.org/gems/woods/versions)
17
+ before editing the Gemfile. A prepared release checkout can update the README
18
+ before its gem is published; if the version is absent, choose an available
19
+ version or wait for publication.
20
+
21
+ If the published 2.x line has only beta or release-candidate versions, use an
22
+ exact pin to the published prerelease, following the README's prerelease
23
+ instructions; `~> 2.0` does not select prereleases. Follow the selected version's
24
+ tag documentation. The `main` guides may describe features absent from the
25
+ published gem.
26
+
27
+ Once a stable 2.x release is published, add Woods to the development group with:
16
28
 
17
29
  ```ruby
18
30
  # Gemfile
@@ -108,13 +120,13 @@ For example:
108
120
 
109
121
  > Use Woods to find `Order`, inspect its resolved callbacks and associations, and list the first two levels of code that depend on it. Cite the Woods identifiers you used.
110
122
 
111
- The Index schema inventory totals 29 schemas. Fourteen register as tools in a normal packaged launch; `codebase_retrieve` is among them but returns a configuration error until embeddings are enabled. The other structural tools work immediately. See [Agent guide](AGENT_GUIDE.md) for a reliable query workflow.
123
+ The Index schema inventory totals 29 schemas. Fourteen register as tools in a normal packaged launch; `codebase_retrieve` is among them but needs embeddings in the default semantic mode, or explicit [lexical retrieval](RETRIEVAL_GUIDE.md#embedding-free-lexical-retrieval) over extraction output. The other structural tools work immediately. See [Agent guide](AGENT_GUIDE.md) for a reliable query workflow.
112
124
 
113
125
  ## Optional next steps
114
126
 
115
127
  ### Add semantic search
116
128
 
117
- Structural search, exact lookup, dependency traversal, graph analysis, and flow tracing do not need embeddings. Add embeddings only when agents need natural-language retrieval.
129
+ Structural search, exact lookup, dependency traversal, graph analysis, and flow tracing do not need embeddings. For ranked natural-language retrieval, choose explicit [lexical mode](RETRIEVAL_GUIDE.md#embedding-free-lexical-retrieval) without providers, or configure embeddings for semantic matching.
118
130
 
119
131
  The local preset uses SQLite metadata, persisted in-memory vectors, and a local Ollama service. Add `gem "sqlite3"` to the application bundle if it is not already present. MySQL/PostgreSQL applications that do not want that dependency can use the `:shared_filesystem` preset instead; it still uses Ollama but persists all stores beneath the Woods output directory.
120
132
 
@@ -154,7 +166,7 @@ When dependencies, initializers, database configuration, credentials, or schema
154
166
 
155
167
  The watcher maintains the structural index. If semantic retrieval is enabled, also run `bin/rails woods:embed_incremental` to update vectors. Without a resident watcher, run `bin/rails woods:incremental` after changes. Use a full `woods:extract` after major upgrades or when validation reports drift. CI and shared-artifact patterns are covered in [Incremental extraction](INCREMENTAL_EXTRACTION.md).
156
168
 
157
- On Rails 8.1, `config/ci.rb` can refresh the index before any gate that reads it: `step "Woods: refresh", "bin/rails woods:incremental"`. With the Claude Code plugin installed, an opt-in `PostToolUse` hook refreshes the index after graph-changing edits and an opt-in `SessionStart` hook warns when it predates the last commit; set `WOODS_HOOKS_ENABLED=1` to turn them on. See [Watch daemon](WATCH_DAEMON.md#hooks-for-agent-sessions).
169
+ On Rails 8.1, `config/ci.rb` can refresh the index before any gate that reads it: `step "Woods: refresh", "bin/rails woods:incremental"`. With the Claude Code plugin installed, an opt-in `PostToolUse` hook refreshes the index after graph-changing edits and an opt-in `SessionStart` hook warns about source-content drift or unknown evidence; set `WOODS_HOOKS_ENABLED=1` to turn them on. See [Watch daemon](WATCH_DAEMON.md#hooks-for-agent-sessions).
158
170
 
159
171
  ### Enable the Console Server
160
172
 
@@ -171,7 +183,7 @@ If live-data queries are necessary, review its allowlists, blocked tables, crede
171
183
  | Rails fails during extraction | Boot and eager-load Rails with the same environment variables | [Troubleshooting](TROUBLESHOOTING.md) |
172
184
  | Validation reports missing or stale units | Run a full extraction, then validate again | [Incremental extraction](INCREMENTAL_EXTRACTION.md) |
173
185
  | MCP reports no index or zero units | Confirm `cwd`, the host-visible `tmp/woods` path, and `woods:stats` output | [MCP servers](MCP_SERVERS.md) |
174
- | `codebase_retrieve` says it is disabled | Configure an embedding provider and run `woods:embed`, or use `search` | [Retrieval guide](RETRIEVAL_GUIDE.md) |
186
+ | `codebase_retrieve` says it is disabled | Choose lexical mode or configure embeddings and run `woods:embed`; `search` also works | [Retrieval guide](RETRIEVAL_GUIDE.md) |
175
187
  | Docker extraction succeeds but MCP cannot see it | Translate the container output path to its host-mounted path | [Docker setup](DOCKER_SETUP.md) |
176
188
 
177
189
  ## Where to go next
@@ -28,12 +28,24 @@ Three differences are tolerated, and nothing else:
28
28
  | Ordering inside a unit's `dependents` | Full extraction appends in extractor order, incremental in graph order. Same multiset. |
29
29
  | PageRank beyond six decimal places | Iterative floating point accumulated in each run's registration order. Scores are compared as values; only the last bits are forgiven. |
30
30
 
31
+ The unit-file write skip ignores only Woods' top-level `extracted_at` stamp.
32
+ A nested metadata field with the same name is application data: changing it
33
+ rewrites the unit in both compact and pretty JSON output.
34
+
31
35
  `graph_analysis.json` used to be a fourth row, tolerating list ordering. It no
32
36
  longer is: the analyzer is order-independent and the oracle compares the file
33
37
  exactly. Tolerating the ordering there meant the harness, the only test that
34
38
  compares a full run against an incremental one, could not see the very
35
39
  dependence the analyzer's determinism work existed to remove.
36
40
 
41
+ Graph targets use string identifiers even when an extractor emits a symbolic
42
+ external target such as `:http_api`. Full and incremental runs retain every
43
+ reverse dependency across JSON restoration and re-registration, including
44
+ contributions from units sharing an identifier under different types. Unit
45
+ dependency metadata retains its extractor-provided values. If an older version
46
+ already lost reverse dependencies on these targets, run a full extraction once
47
+ to restore them; loading the damaged graph cannot recover discarded entries.
48
+
37
49
  This matters most for **incremental CI chains**: restore the previous graph,
38
50
  run `woods:incremental` per merge. There, a unit that goes missing propagates
39
51
  forward run over run instead of being erased by the next full rebuild.
@@ -52,6 +64,10 @@ failure a CI chain cannot afford:
52
64
  | The range fails otherwise | Actionable error naming the range, **exit 1**. |
53
65
  | There is no `git` binary at all | Same two rows as above: the failure reads `git unavailable: …` and takes the daemon-coverage decision, rather than dying with an `Errno::ENOENT` backtrace. |
54
66
 
67
+ Changed paths are normalized lexically before dispatch: trailing root slashes,
68
+ duplicate separators and `.`/`..` segments do not create separate changes or
69
+ bypass matching. Missing files remain representable; symlinks are not resolved.
70
+
55
71
  The range comes from `CI_COMMIT_BEFORE_SHA..CI_COMMIT_SHA` (GitLab),
56
72
  `origin/$GITHUB_BASE_REF...HEAD` (GitHub Actions), or `HEAD~1` (default). An
57
73
  unresolvable range — a GitLab zero-SHA on a new branch, an unfetched base ref,
@@ -76,6 +92,29 @@ The diff itself is rooted at the extracted application (`git -C Rails.root`),
76
92
  so it cannot read whatever checkout the process happened to start in — the
77
93
  same rooting rule the manifest's git provenance follows.
78
94
 
95
+ Named, source-defined app modules included by runtime models are tracked as concern units even
96
+ outside `concerns/` directories. Changing their source refreshes their includers,
97
+ including inlined code and callback analysis. Multiple runtime mixins sharing a source
98
+ file retain separate identities and refresh all their includers. Run a full extraction after upgrading
99
+ to populate these previously missing source mappings.
100
+
101
+ ## Handled source errors and retry
102
+
103
+ Unreleased after `2.0.0.beta3`: when an incremental extraction or named refresh
104
+ records a handled consumer error (for example malformed locale or schedule
105
+ YAML), it raises `Woods::ExtractionError` before publishing. Empty output from
106
+ that failed consumer does not authorize replacing or deleting its last-good
107
+ units. The published generation and source provenance remain unchanged,
108
+ including when other files in the batch extracted successfully.
109
+
110
+ Fix the source error named in the extraction log, then retry the **complete
111
+ batch**, or the same named refresh. The watch daemon reports degraded and keeps
112
+ the failed batch pending for retry. An error on one file does not mark a later
113
+ successful file as failed, but the batch still cannot publish until all handled
114
+ errors are resolved. This does not change full extraction's existing tolerance
115
+ for handled consumer errors; its source-freshness report marks those scopes
116
+ unverified.
117
+
79
118
  ## What a run does, in order
80
119
 
81
120
  `Extractor#extract_changed` is order-sensitive; each step exists because of the
@@ -87,8 +126,8 @@ step before it.
87
126
  2. **Reconcile changed paths.** Every changed path that still exists is handed
88
127
  to the file-based extractors that claim it (`PathDispatcher`), and units the
89
128
  path no longer produces are dropped. This is what indexes a file the index
90
- has never seen, and what lets a task removed from a multi-task `.rake` file
91
- actually go away.
129
+ has never seen, and what removes definitions deleted from a surviving
130
+ source file. Multi-file Rake tasks use wholesale reconciliation below.
92
131
  3. **Re-extract the rest of the blast radius**: units whose own file did not
93
132
  change but which depend on something that did.
94
133
  4. **Reconcile class-based types** against each extractor's
@@ -120,6 +159,11 @@ step before it.
120
159
  string literal, and an unrelated addition in the same batch do not.
121
160
  Idempotent when nothing was pruned.
122
161
 
162
+ Git enrichment uses the same eligibility checks in full and incremental runs:
163
+ existing app-owned files under `Rails.root`, excluding `vendor/`, `node_modules/`,
164
+ and framework/gem source units. Each typed unit resolves its own file history,
165
+ even when its identifier is shared by another type.
166
+
123
167
  Then the second pass: `dependents` and `metadata.git` are refreshed on every
124
168
  touched unit (the incremental equivalents of full extraction's phases 2 and 4),
125
169
  type indexes are regenerated, the graph, `graph_analysis.json` and the
@@ -174,13 +218,13 @@ automatically.
174
218
  | `app/decorators`, `app/presenters`, `app/form_objects` | decorators |
175
219
  | `app/managers` / `app/policies` / `app/validators` | managers / policies + pundit_policies / validators |
176
220
  | `app/**/concerns/**/*.rb` | concerns |
177
- | `app/models/**/*.rb` (outside `concerns/`) | poros, caching |
221
+ | `app/models/**/*.rb` (outside `concerns/`) | poros, caching; concerns when runtime model inclusion confirms a mixin |
222
+ | `app/**/*.rb`, `lib/**/*.rb` (outside `concerns/`) | runtime model mixins also dispatch to concerns |
178
223
  | `app/controllers/**/*.rb` | caching |
179
224
  | `app/views/**/*.erb` | view_templates, caching |
180
225
  | `config/locales/**/*.yml` | i18n |
181
226
  | `config/initializers`, `config/environments` | configurations |
182
227
  | `db/migrate/*.rb` (top level only) | migrations |
183
- | `lib/tasks/**/*.rake` | rake_tasks |
184
228
  | `lib/**/*.rb` (outside `tasks/`, `generators/`) | libs |
185
229
  | `spec/**/*_spec.rb`, `test/**/*_test.rb` | test_mappings |
186
230
 
@@ -191,12 +235,13 @@ serializer and decorator extractors, and all matching rules run.
191
235
  ### Wholesale re-runs
192
236
 
193
237
  `PathDispatcher.whole_app_rules` → `Extractor::WHOLE_APP_EXTRACTORS`. These
194
- extractors have no per-file entry point: they introspect the runtime or scan a
195
- whole directory in one pass. In an already-booted process re-running them is
238
+ extractors need a complete runtime or directory view, even when a low-level
239
+ per-file reader exists. In an already-booted process re-running them is
196
240
  cheap, which is what makes wholesale replacement the right shape.
197
241
 
198
242
  | Trigger | Re-runs |
199
243
  |---|---|
244
+ | `lib/tasks/**/*.rake` | rake_tasks (all definitions of every task) |
200
245
  | `config/routes.rb`, `config/routes/**` | routes, engines, **and** controllers, mailers, components, view components, view templates |
201
246
  | `Gemfile.lock` | engines, middleware, rails_source (gated by `include_framework_sources`) |
202
247
  | `config/application.rb`, `config/initializers/**`, `config/environments/**` | middleware |
@@ -207,7 +252,14 @@ cheap, which is what makes wholesale replacement the right shape.
207
252
  | `db/views/**/*.sql` | database_views |
208
253
  | any `package.yml`, `packwerk.yml` | packages |
209
254
 
210
- Three of these deserve a note:
255
+ Four of these deserve a note:
256
+
257
+ - **Rake tasks merge definitions across files.** Any changed or deleted `.rake`
258
+ file reruns the task extractor over all task files. Removing the primary
259
+ definition preserves surviving definitions; removing a secondary definition
260
+ drops its source and dependencies from the shared unit. After upgrading from
261
+ the old per-file rules, run a full extraction to establish a source-freshness
262
+ baseline with the new rule fingerprint.
211
263
 
212
264
  - **Routes cascade.** `ROUTE_CONSUMER_EXTRACTORS` embed the route table, controllers write each action's routes into unit metadata and into the action
213
265
  chunks, and everything using `RouteHelperResolver` resolves navigation edges
@@ -258,6 +310,34 @@ oracle compares against emits it too, both sides agree, wrongly. The coverage
258
310
  is in `spec/extractor_spec.rb`, driving the reconciler with a shrinking
259
311
  discovery set.
260
312
 
313
+ ### Runtime removals and bundle updates
314
+
315
+ Jobs discovered through `ApplicationJob.descendants` supplement the job-file
316
+ scan, but jobs are not part of `CLASS_BASED_DISCOVERY` removal reconciliation.
317
+ If a dynamically defined or gem-owned job disappears without a tracked source
318
+ path changing, its unit can survive subsequent incremental runs. A full
319
+ extraction in a fresh Rails process removes it; an in-process full extraction
320
+ can still see an old constant retained by that process (B-165).
321
+
322
+ After adding, removing, or updating bundled gems, boot the updated bundle in a
323
+ fresh process and run:
324
+
325
+ ```bash
326
+ bundle exec rake woods:extract woods:validate
327
+ ```
328
+
329
+ A `Gemfile.lock` change refreshes engines, middleware, and optional framework
330
+ sources. It does not refresh every gem-owned model, job, or other runtime unit.
331
+ Their recorded paths or metadata can remain stale, including absolute paths to
332
+ a removed gem version and paths under `vendor/`. A full extraction rebuilds
333
+ those units against the installed bundle (B-166).
334
+
335
+ An absent external source path can also mean the validator runs on a different
336
+ host or mount from extraction. Confirm the bundle and filesystem context before
337
+ rebuilding; a full run in one container does not make its gem paths visible on
338
+ another host. Validation warnings identify missing paths, but do not prove a
339
+ retained unit matches the currently installed gem when its path still exists.
340
+
261
341
  ### Deletion
262
342
 
263
343
  - Paths named in the change set that no longer exist are **authoritative** for
@@ -423,6 +503,66 @@ owns the definition of "the two indexes agree" and documents every exclusion.
423
503
 
424
504
  **Run it before and after any change to the incremental path.**
425
505
 
506
+ ## Profiling fixed costs
507
+
508
+ Set `WOODS_PROFILE=1` to time extraction phases. Git enrichment and unit JSON
509
+ finalization are separate from incremental re-extraction; runtime discovery,
510
+ whole-app reruns and pruning appear under `reconciliation`. `payload sync`,
511
+ `publish` (the generation pointer write), and `payload prune` (retention) are
512
+ separate, additive phases. Older versions included sync and retention inside
513
+ `publish`, so do not sum those older lines without subtracting nested sync.
514
+
515
+ `[profile total]` reports time inside the extraction call, including setup and
516
+ failed runs. It is not another phase to sum. Compare it with phase durations
517
+ to find unaccounted work; each line rounds to hundredths of a second. Rails
518
+ boot, Bundler and watch reload work outside the call require separate wall
519
+ measurements. Do not infer that all unaccounted time is boot.
520
+
521
+ Both full and incremental runs currently seed the prior payload. Cloning uses
522
+ per-file hardlinks (or copies where links are unsupported); updated files are
523
+ replaced atomically, preserving previous generations. Full-run carry-forward
524
+ and generation retention remain unchanged. The seed walker classifies each entry
525
+ once, avoiding duplicate file metadata lookups; it still links each file
526
+ separately. Generation directories remain independent, so pruning an old generation cannot remove a newer one's files.
527
+ Compare repeated runs on the actual index filesystem before attributing latency
528
+ to extraction or changing the index layout.
529
+
530
+ For a repeatable component comparison from a source checkout:
531
+
532
+ ```bash
533
+ # BENCH_ROOT is an existing scratch parent on the index filesystem.
534
+ # The script creates and removes only its own temporary directory there.
535
+ BENCH_ROOT=/path/on/index/filesystem WOODS_SOURCE=/path/to/baseline \
536
+ ruby bench/payload_seed.rb
537
+ BENCH_ROOT=/path/on/index/filesystem WOODS_SOURCE=/path/to/candidate \
538
+ ruby bench/payload_seed.rb
539
+ ```
540
+
541
+ The benchmark clones 8,335 synthetic 2 KiB files across 35 directories, checks
542
+ file counts and bytes outside timing, and reports seven samples plus medians.
543
+ This measures the seed component only; it does not boot Rails or establish an
544
+ end-to-end host improvement. Filesystem metadata latency and the copy fallback
545
+ can dominate differently from local hardlink results. A first full extraction
546
+ has no previous payload, so include a repeated full run when measuring seed cost.
547
+ Full seeding also preserves on-demand framework units when framework extraction
548
+ is disabled, and non-JSON files in existing type directories; dropping the seed
549
+ wholesale would change that behavior.
550
+
551
+ ### Choosing full versus incremental for CI
552
+
553
+ Measure repeated full and representative-day incremental runs with
554
+ `WOODS_PROFILE=1` on the same resulting application tree, configuration and index
555
+ filesystem. Restore the same baseline index before each incremental trial;
556
+ otherwise a second trial may be a no-op. Include Rails boot in both wall times
557
+ when comparing separate task invocations, and compare resident watcher cycles
558
+ separately. Validate each resulting index with `woods:validate`.
559
+
560
+ Use full extraction for that workload when its median wall time is no greater
561
+ than the representative incremental run. There is no universal changed-file
562
+ threshold: shared dependencies and whole-app extractor triggers change the work
563
+ per file. Re-measure after substantial application or Woods changes. A fast leaf
564
+ edit does not establish that a day of commits is below the crossover.
565
+
426
566
  ## Boundaries and open work
427
567
 
428
568
  - **Reloaded deletion is supported.** The resident watcher reloads changed