woods 2.0.1 → 2.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (154) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +94 -7
  3. data/CONTRIBUTING.md +134 -19
  4. data/README.md +1 -1
  5. data/docs/AGENT_GUIDE.md +19 -0
  6. data/docs/AGENT_SETUP.md +22 -2
  7. data/docs/BACKEND_MATRIX.md +7 -0
  8. data/docs/CLIENT_HOOKS.md +6 -0
  9. data/docs/CONFIGURATION_REFERENCE.md +133 -25
  10. data/docs/CONSOLE_MCP_SETUP.md +82 -30
  11. data/docs/EMBEDDING_MODELS.md +16 -19
  12. data/docs/EXTRACTOR_REFERENCE.md +219 -21
  13. data/docs/FAQ.md +11 -25
  14. data/docs/GETTING_STARTED.md +7 -1
  15. data/docs/INCREMENTAL_EXTRACTION.md +261 -19
  16. data/docs/INDEX_LAYOUT.md +5 -0
  17. data/docs/INTERNALS.md +9 -0
  18. data/docs/MCP_HTTP_TRANSPORT.md +20 -15
  19. data/docs/MCP_SERVERS.md +87 -8
  20. data/docs/MCP_TOOL_COOKBOOK.md +13 -55
  21. data/docs/NOTION_INTEGRATION.md +7 -1
  22. data/docs/PUBLISHED_INDEX.md +6 -0
  23. data/docs/README.md +6 -1
  24. data/docs/RETRIEVAL_GUIDE.md +17 -0
  25. data/docs/SOURCE_FRESHNESS.md +157 -5
  26. data/docs/TOKEN_BENCHMARK.md +10 -18
  27. data/docs/TROUBLESHOOTING.md +70 -14
  28. data/docs/UNBLOCKED_INTEGRATION.md +60 -8
  29. data/docs/UPGRADING_TO_2.md +153 -38
  30. data/docs/WATCH_DAEMON.md +97 -14
  31. data/exe/woods-console-mcp +2 -2
  32. data/lib/generators/woods/templates/woods.rb.tt +2 -1
  33. data/lib/tasks/woods.rake +23 -7
  34. data/lib/tasks/woods_checks.rake +2 -2
  35. data/lib/woods/agent_configuration/cli.rb +1 -1
  36. data/lib/woods/agent_configuration/layout.rb +16 -2
  37. data/lib/woods/agent_configuration/plan.rb +13 -3
  38. data/lib/woods/agent_configuration/planner_validation.rb +4 -2
  39. data/lib/woods/agent_configuration/preflight.rb +5 -3
  40. data/lib/woods/builder.rb +17 -57
  41. data/lib/woods/cache/cache_middleware.rb +56 -30
  42. data/lib/woods/chunking/contributor_chunks.rb +119 -0
  43. data/lib/woods/chunking/semantic_chunker.rb +44 -21
  44. data/lib/woods/console/connection_manager.rb +56 -3
  45. data/lib/woods/console/embedded_executor.rb +30 -5
  46. data/lib/woods/console/rack_middleware.rb +29 -1
  47. data/lib/woods/dependency_graph.rb +34 -10
  48. data/lib/woods/embedding/fake.rb +12 -0
  49. data/lib/woods/embedding/indexer.rb +195 -98
  50. data/lib/woods/embedding/input_budget.rb +67 -0
  51. data/lib/woods/embedding/openai.rb +70 -20
  52. data/lib/woods/embedding/provider.rb +37 -25
  53. data/lib/woods/embedding/text_preparer.rb +76 -32
  54. data/lib/woods/embedding/token_counter.rb +18 -81
  55. data/lib/woods/embedding/vector_configuration.rb +48 -0
  56. data/lib/woods/extraction_identities.rb +175 -0
  57. data/lib/woods/extractor.rb +304 -107
  58. data/lib/woods/extractors/action_cable_extractor.rb +8 -3
  59. data/lib/woods/extractors/assigned_value_discovery.rb +74 -0
  60. data/lib/woods/extractors/class_declarations.rb +121 -0
  61. data/lib/woods/extractors/configuration_extractor.rb +11 -3
  62. data/lib/woods/extractors/declaration_ancestry.rb +92 -0
  63. data/lib/woods/extractors/event_extractor.rb +8 -0
  64. data/lib/woods/extractors/graphql_extractor.rb +134 -77
  65. data/lib/woods/extractors/job_extractor.rb +5 -1
  66. data/lib/woods/extractors/lib_extractor.rb +132 -15
  67. data/lib/woods/extractors/mailer_extractor.rb +3 -5
  68. data/lib/woods/extractors/manager_extractor.rb +7 -21
  69. data/lib/woods/extractors/migration_declaration.rb +87 -0
  70. data/lib/woods/extractors/migration_extractor.rb +5 -39
  71. data/lib/woods/extractors/phlex_extractor.rb +6 -2
  72. data/lib/woods/extractors/policy_extractor.rb +9 -5
  73. data/lib/woods/extractors/poro_extractor.rb +112 -53
  74. data/lib/woods/extractors/pundit_extractor.rb +11 -6
  75. data/lib/woods/extractors/scheduled_job_extractor.rb +45 -4
  76. data/lib/woods/extractors/serializer_extractor.rb +34 -22
  77. data/lib/woods/extractors/shared_utility_methods.rb +18 -1
  78. data/lib/woods/extractors/source_nesting.rb +142 -106
  79. data/lib/woods/extractors/standalone_module_discovery.rb +123 -0
  80. data/lib/woods/extractors/state_machine_extractor.rb +46 -40
  81. data/lib/woods/extractors/view_component_extractor.rb +9 -7
  82. data/lib/woods/flow_assembler.rb +4 -1
  83. data/lib/woods/generation.rb +25 -0
  84. data/lib/woods/hooks/context_hint.rb +7 -2
  85. data/lib/woods/mcp/bootstrapper.rb +33 -7
  86. data/lib/woods/mcp/config_resolver.rb +26 -7
  87. data/lib/woods/mcp/index_reader.rb +125 -24
  88. data/lib/woods/mcp/index_reader_pinning.rb +16 -0
  89. data/lib/woods/mcp/renderers/markdown_renderer.rb +7 -1
  90. data/lib/woods/mcp/renderers/plain_renderer.rb +3 -1
  91. data/lib/woods/mcp/search_results.rb +7 -1
  92. data/lib/woods/mcp/server.rb +24 -4
  93. data/lib/woods/module_reconciliation.rb +151 -0
  94. data/lib/woods/path_dispatcher.rb +7 -2
  95. data/lib/woods/rake_helpers.rb +43 -11
  96. data/lib/woods/release.rb +1 -1
  97. data/lib/woods/resilience/index_validator.rb +8 -3
  98. data/lib/woods/resilience/retryable_provider.rb +18 -1
  99. data/lib/woods/resolved_config.rb +68 -8
  100. data/lib/woods/retrieval/context_assembler.rb +3 -3
  101. data/lib/woods/retrieval/lexical_assembler.rb +3 -2
  102. data/lib/woods/retrieval/scope.rb +18 -2
  103. data/lib/woods/retrieval/source_evidence.rb +14 -2
  104. data/lib/woods/source_contributor_validation.rb +78 -0
  105. data/lib/woods/source_contributors.rb +116 -0
  106. data/lib/woods/source_inputs/handoff.rb +37 -0
  107. data/lib/woods/source_inputs/launcher.rb +53 -13
  108. data/lib/woods/source_inputs/manifest.rb +84 -3
  109. data/lib/woods/source_inputs/private_key.rb +44 -12
  110. data/lib/woods/source_inputs/scanner.rb +98 -27
  111. data/lib/woods/source_inputs/scopes.rb +1 -1
  112. data/lib/woods/source_inputs/session.rb +147 -15
  113. data/lib/woods/source_inputs/stable_reader.rb +127 -0
  114. data/lib/woods/source_inputs/status.rb +40 -8
  115. data/lib/woods/source_inputs/verifier.rb +28 -5
  116. data/lib/woods/source_path_encoding.rb +33 -0
  117. data/lib/woods/source_references/cache.rb +284 -0
  118. data/lib/woods/source_references/collector.rb +120 -0
  119. data/lib/woods/source_references/extraction.rb +185 -0
  120. data/lib/woods/source_references/inputs.rb +134 -0
  121. data/lib/woods/source_references/parser_adapter.rb +134 -0
  122. data/lib/woods/source_references/pass.rb +152 -0
  123. data/lib/woods/source_references/prism_adapter.rb +116 -0
  124. data/lib/woods/source_references/registry.rb +178 -0
  125. data/lib/woods/source_references/runtime_lookup.rb +127 -0
  126. data/lib/woods/source_references/value_class.rb +82 -0
  127. data/lib/woods/storage/metadata_store.rb +4 -1
  128. data/lib/woods/storage/qdrant.rb +2 -2
  129. data/lib/woods/unblocked/client.rb +12 -7
  130. data/lib/woods/unblocked/document_builder.rb +4 -1
  131. data/lib/woods/unblocked/exporter.rb +127 -37
  132. data/lib/woods/unblocked/sync_manifest.rb +137 -21
  133. data/lib/woods/unblocked/uri_migration.rb +105 -0
  134. data/lib/woods/util/host_guard.rb +3 -2
  135. data/lib/woods/version.rb +1 -1
  136. data/lib/woods/watch/catch_up.rb +138 -0
  137. data/lib/woods/watch/claim_lease.rb +150 -0
  138. data/lib/woods/watch/cli.rb +26 -2
  139. data/lib/woods/watch/daemon.rb +80 -59
  140. data/lib/woods/watch/installation/options.rb +1 -1
  141. data/lib/woods/watch/installation/receipt.rb +6 -1
  142. data/lib/woods/watch/managed_child.rb +1 -1
  143. data/lib/woods/watch/supervisor.rb +1 -1
  144. data/lib/woods/watch/tree_scan.rb +14 -2
  145. data/plugin/.claude-plugin/plugin.json +1 -1
  146. data/plugin/hooks/adapters/normalize.rb +3 -2
  147. data/plugin/hooks/woods-input-rules.sh +4 -0
  148. data/plugin/hooks/woods-refresh.sh +15 -7
  149. data/plugin/hooks/woods-session-start.sh +60 -3
  150. data/plugin/skills/woods-diagnose/SKILL.md +334 -11
  151. data/plugin/skills/woods-investigate/SKILL.md +11 -0
  152. data/plugin/skills/woods-mcp-config/SKILL.md +79 -8
  153. data/plugin/skills/woods-setup/SKILL.md +53 -6
  154. metadata +32 -5
@@ -32,23 +32,69 @@ Extractors discover code one of two ways:
32
32
 
33
33
  Some extractors combine both (e.g., `JobExtractor` scans directories first, then supplements with `ApplicationJob.descendants`).
34
34
 
35
- Discovery is not exhaustive. A standalone module under `app/models` that is
36
- called through singleton methods is not discovered by the model/PORO paths;
37
- conventional concerns and app modules included by live models have separate
38
- concern discovery. This gap is tracked in [#552](https://github.com/lost-in-the/woods/issues/552).
39
- Separately, dependency scanning does not capture every method-body constant
40
- reference ([#475](https://github.com/lost-in-the/woods/issues/475)). A missing unit
41
- or edge is not proof of unused code; cross-check the application source.
35
+ Discovery is not exhaustive. Woods 2.0.0 does not discover callable standalone
36
+ modules under `app/models` through its model/PORO paths. Woods 2.1
37
+ [standalone-module support](#poroextractor) addresses that case; conventional
38
+ concerns and app modules included by live models retain concern ownership.
39
+ Dependency scanning also remains partial, including with the source-reference
40
+ pass below. A missing unit or edge is not proof of unused code; cross-check the
41
+ application source.
42
+
43
+ ### Constant source references
44
+
45
+ **Included in Woods 2.1.** For Git/path installations,
46
+ record the loaded gem path and exact revision; the development version alone
47
+ does not establish that this change is installed.
48
+
49
+ After discovery and deduplication, extraction adds `code_reference` relationships
50
+ from models, controllers, services, POROs, library units and concerns to verified,
51
+ indexed class/module targets. It parses original Ruby source and attributes reads
52
+ to their declaring owner. Ordinary method bodies, singleton methods and callback
53
+ blocks participate. Comments, plain strings and declaration names do not establish
54
+ references. Rails relationships derived through runtime reflection stay intact.
55
+
56
+ Resolution respects qualified names and supported lexical/runtime context.
57
+ A pending library autoload can supply a target only when its exact registered
58
+ constant and absolute file path match the captured library declaration. Woods
59
+ does not execute the autoload. Pending namespaces, unverified owner scopes,
60
+ mismatched registration paths and non-library autoload targets remain unresolved.
61
+ Dynamic constant lookup, uncertain aliases, ambiguous typed identities and
62
+ unsupported scopes remain unresolved. In particular, references inside
63
+ `class << self` are recorded as candidates but currently skipped during resolution;
64
+ ordinary `def self.method` bodies are supported. A source reference does not prove
65
+ execution, and an absent edge does not prove there are no callers. The internal cache retains parsed candidates, their scope, and collector skip
66
+ reasons; it does not establish complete reference coverage.
67
+
68
+ Forward edges and reverse relationships publish together. Incremental extraction
69
+ and targeted refresh reconsider cached references when targets appear, disappear
70
+ or change resolution, including references in unchanged callers. Larger recorded
71
+ dependency sets can increase the incremental blast radius; no separate limit is
72
+ applied to these edges. See [source-reference baseline and upgrades](INCREMENTAL_EXTRACTION.md#source-reference-baseline-and-upgrades).
42
73
 
43
74
  ### Identifier naming (source-derived units)
44
75
 
45
76
  File-based extractors derive an identifier in three steps, first match wins:
46
77
 
47
- 1. **Zeitwerk-governed naming.** For a file under a managed autoload path, the expected constant path is computed — from `Rails.autoloaders.main` when a Rails autoloader is up, from the same path-to-constant convention offline otherwise — and the source must declare exactly that constant: `app/services/domain/container/parser.rb` is expected to define `Domain::Container::Parser`. This is what lets a file whose namespaces are written as *classes* (`module Domain; class Container; class Parser`) name the file's own constant instead of the wrapper `Domain::Container`, which every sibling under the same wrapper would otherwise collide with.
78
+ 1. **Zeitwerk-governed naming.** For a file under a managed autoload path, the expected constant path comes from its owning Rails loader (`main` or `once`), including that loader's inflections, root namespace, collapsed directories and ignored paths. Without loader information, the offline path convention applies. The source must declare exactly the expected constant: `app/services/domain/container/parser.rb` is expected to define `Domain::Container::Parser`. This is what lets a file whose namespaces are written as *classes* (`module Domain; class Container; class Parser`) name the file's own constant instead of the wrapper `Domain::Container`, which every sibling under the same wrapper would otherwise collide with.
48
79
  2. **Position-aware nesting scan** (`SourceNesting`, the supporting utility above). The first `class` declaration qualified by the namespaces actually open at that position. Compact declarations keep their segments; a helper module nested inside the class and sibling modules that closed earlier do not contribute.
49
80
  3. **Path convention.** Camelize the path under `app/<kind>/` (details vary per extractor).
50
81
 
51
- Unmanaged paths (`lib/`, configured non-autoload roots) and sources that declare nothing matching the expected constant skip step 1 entirely: the source scan and path convention decide, exactly as before the governed lookup existed.
82
+ Unmanaged paths (`lib/` unless configured for autoloading, and configured non-autoload roots) and sources that declare nothing matching the expected constant skip step 1 entirely: the source scan and path convention decide. A namespace file containing only `Acme::VERSION` does not acquire an invented `Acme::Version` declaration.
83
+
84
+ An inline primary module with a body, such as `module Helpers; def self.call; end; end`, keeps its declared name even when an unmanaged directory suggests another namespace. Namespace-only wrappers containing a single nonempty module retain the nested identity; direct behavior, empty inner modules, and `ClassMethods`/`InstanceMethods` plumbing stop that descent. Empty completed modules before the primary declaration remain namespace preludes. This correction is included in Woods 2.1; run a full extraction after upgrading to replace affected path-derived identities and their source-reference edges.
85
+
86
+ One-line class declarations retain the same nesting as their multiline forms:
87
+ `module Acme; class Error < StandardError; end; end` names `Acme::Error`,
88
+ separately from the `Acme` namespace unit. Complete-source parsing keeps
89
+ declaration-like text inside strings and heredocs out of identifier selection.
90
+ Conditional declarations remain source candidates; this scan does not establish
91
+ which branch ran. Method and singleton-class bodies do not supply ordinary
92
+ owners. Invalid source keeps the tolerant class scan, while module selection
93
+ requires valid syntax. These corrections are also included in Woods 2.1;
94
+ re-extract affected indexes in full to replace wrapper identities and restore
95
+ their source-reference edges.
96
+
97
+ Once-loader ownership (#579) is included in Woods 2.1; verify the loaded revision as well as the gem version. It covers declared constants under `config.autoload_lib_once` and other once-managed roots. Direct ownership across both loaders takes precedence over copied-application inference; ambiguous roots and a loader's explicit non-claim remain unmanaged. After upgrading an affected index, run one full extraction before resuming incremental maintenance to replace stale wrapper identifiers. This does not aggregate multiple unmanaged source files reopening the same namespace.
52
98
 
53
99
  ### Eager loading
54
100
 
@@ -234,6 +280,7 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
234
280
  **Key details:**
235
281
  - Scans: `app/jobs`, `app/workers`, `app/sidekiq`
236
282
  - Extracts queue name, retry configuration, concurrency options, perform method arguments, and callbacks
283
+ - File-discovered jobs read queue metadata from source. Runtime `queue_name` supplements only class-discovered jobs, such as nested jobs outside the job directories; a file-discovered `queue_as Settings::QUEUE` can therefore retain a `null` queue even after refresh.
237
284
  - `perform_params` uses the same syntax-aware signature parsing as service initializers and preserves its `name`, `splat` (`single`/`double`/null), and `has_default` fields. Keyword defaults do not invent additional argument names.
238
285
  - Records what triggers this job (reverse lookup via dependency graph after extraction)
239
286
  - Supports both ActiveJob and Sidekiq native workers
@@ -260,7 +307,7 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
260
307
  **What it captures:** ActionMailer classes with their mailer actions, defaults, template paths, callbacks, and helper usage.
261
308
 
262
309
  **Key details:**
263
- - Discovers `ApplicationMailer.descendants` when that class exists, otherwise `ActionMailer::Base.descendants`; an app without ActionMailer contributes no mailer units.
310
+ - **Included in Woods 2.1:** discovers all app-owned `ActionMailer::Base.descendants`, including parallel abstract bases and direct subclasses even when `ApplicationMailer` exists. An app without ActionMailer contributes no mailer units.
264
311
  - Discovery and direct extraction accept only mailers backed by an existing app-owned source file, excluding dependency mailers and fabricated convention paths.
265
312
  - Each mailer action corresponds to an email template, template paths are recorded in metadata
266
313
  - Extracts `default from:`, `layout`, and per-action subject patterns
@@ -326,6 +373,8 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
326
373
 
327
374
  ### PhlexExtractor
328
375
 
376
+ **Included in Woods 2.1:** discovery and direct extraction require an existing app-owned source file; dependency components are excluded.
377
+
329
378
  **What it captures:** Phlex component classes (`Phlex::HTML`, `Phlex::SVG` subclasses) from `app/components`. Extracts slots, initialize parameters, sub-component references, Stimulus controller names, and route helper usage.
330
379
 
331
380
  **Key details:**
@@ -336,10 +385,13 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
336
385
 
337
386
  ### ViewComponentExtractor
338
387
 
388
+ **Included in Woods 2.1:** discovery and direct extraction require an existing app-owned source file and runtime ancestry; dependency components are excluded.
389
+
339
390
  **What it captures:** ViewComponent classes from `app/components`. Extracts slots, template paths, preview class references, and collection rendering support.
340
391
 
341
392
  **Key details:**
342
393
  - Template path is inferred from the component file name (e.g., `ButtonComponent` → `button_component.html.erb`)
394
+ - **Included in Woods 2.1:** `metadata.sidecar_template` uses an application-relative path, such as `app/components/button_component.html.erb`. Detection still checks the actual file under `Rails.root`; extracting from another checkout does not change this metadata. Re-extract existing component units to update their stored paths.
343
395
  - Preview class associations are extracted when `<ComponentName>Preview` is found in `spec/components/previews/` or `test/components/previews/`
344
396
 
345
397
  **Edge cases:**
@@ -389,7 +441,7 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
389
441
 
390
442
  ### PoroExtractor
391
443
 
392
- **What it captures:** Plain Ruby objects in `app/models` that are not ActiveRecord (non-AR classes, excluding concerns).
444
+ **What it captures:** Plain Ruby objects in `app/models` that are not ActiveRecord (non-AR classes, excluding concerns). Woods 2.1 writers also discover callable standalone modules as described below.
393
445
 
394
446
  **Key details:**
395
447
  - Scans `app/models` for files that don't define an `ActiveRecord::Base` descendant
@@ -397,6 +449,60 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
397
449
  - Excludes concerns (those go to ConcernExtractor)
398
450
  - `parent_class` and the generated Parent annotation describe the selected unit declaration only. Nested or sibling classes cannot supply its parent. An implicit `Object` parent, a dynamic superclass expression, or unparseable source produces `nil`; explicit constant-path parents retain their source names.
399
451
 
452
+ #### Assigned value classes
453
+
454
+ **Included in Woods 2.1.** Direct assignments such as
455
+ `Criteria = Struct.new(:value)` and `Page = Data.define(:items)` can own a PORO
456
+ unit inside class or module namespace wrappers. Discovery uses the active
457
+ loader's expected constant where available, then verifies the loaded class,
458
+ core factory, ancestry and canonical assignment file. For example, a reopened
459
+ `Products::Search` containing `Criteria = Struct.new(...)` in its child file
460
+ produces `Products::Search::Criteria`, not another `Products::Search` unit.
461
+
462
+ The same verified naming applies to [library files](#libextractor). Shadowed
463
+ factories, aliases, arbitrary `Class.new` assignments and unloaded nested value
464
+ classes are not promoted by this check. Existing top-level named-Struct lookup
465
+ identities remain compatible, but an alias is not a verified reference target.
466
+ A callable canonical module sharing the file retains its own PORO unit. Library
467
+ extraction preserves its canonical primary module when it matches the file's
468
+ identity or has its own methods defined there; namespace-only child-file wrappers
469
+ do not displace the assigned child. An assigned class can be the target of a
470
+ `code_reference` edge, but references inside its Struct/Data constructor block
471
+ remain unsupported because the block does not establish ordinary lexical class
472
+ nesting. Woods does not execute the constructor or its method bodies.
473
+
474
+ Run a **full extraction** after upgrading the writer. It repairs older wrapper
475
+ identities and rebuilds the source-reference cache in format 3; incremental
476
+ extraction refuses format 1 even when source files are unchanged. Existing
477
+ readers can continue serving the last published index until the full extraction
478
+ succeeds. Genuine collisions between different files still refuse publication.
479
+
480
+ **Standalone modules — included in Woods 2.1.** A named,
481
+ loaded module whose canonical declaration and own methods are defined under
482
+ `app/models` can produce a `poro` unit with `metadata.ruby_kind: "module"` and
483
+ `parent_class: null`. This includes `def self.method`, `module_function`,
484
+ `class << self` methods and plain instance-method mixins. Discovery uses runtime
485
+ ownership and source locations without calling those methods.
486
+
487
+ Namespace-only wrappers, aliases, unloaded constants, gem-owned modules and
488
+ methods supplied only by another file do not establish standalone ownership.
489
+ An app module included by a live model belongs to `ConcernExtractor`; a library
490
+ module remains owned by `LibExtractor`. Separate callable modules sharing a
491
+ source file retain their own identities. Ordinary class PORO identifiers stay unchanged.
492
+
493
+ Incremental extraction and model/PORO/concern refreshes reconcile ownership when
494
+ an includer changes even if the module's file does not. Incomplete eager loading
495
+ cannot prove that a former concern became standalone or that an undiscovered
496
+ module disappeared. Those units remain retained unless positive ownership
497
+ evidence supersedes them; changed source for a retained unit refuses publication
498
+ until a complete run can re-extract it.
499
+
500
+ Module discovery and reference resolution are separate: a module with
501
+ `class << self` methods can be indexed even though references inside that scope
502
+ remain unresolved by the [constant-reference pass](#constant-source-references).
503
+ Run a full extraction when upgrading to establish the new units and their
504
+ reference cache; updating only the reader does not add them.
505
+
400
506
  ---
401
507
 
402
508
  ### SerializerExtractor
@@ -422,19 +528,35 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
422
528
 
423
529
  **What it captures:** `SimpleDelegator` subclasses that wrap a model. Records the wrapped model class, all public methods, and the delegation chain.
424
530
 
531
+ **Included in Woods 2.1:** after eager loading, discovery checks the
532
+ selected class's actual delegator ancestry and source ownership. Application
533
+ base classes and leading `::` therefore work without changing the unit identity;
534
+ an unrelated or foreign same-named class cannot supply that proof. When the class
535
+ is unavailable, the extractor retains its limited direct-declaration fallback
536
+ without triggering autoloads. `delegation_type` reflects resolved SimpleDelegator
537
+ ancestry; unknown delegation mechanisms remain `unknown`.
538
+
425
539
  ---
426
540
 
427
541
  ## API & authorization extractors
428
542
 
429
543
  ### GraphQLExtractor
430
544
 
431
- **What it captures:** graphql-ruby types, mutations, queries, and resolvers. Produces four distinct unit types from one extractor.
545
+ **What it captures:** graphql-ruby schemas, types, mutations, queries, and resolvers. Produces four distinct unit types from one extractor.
546
+
547
+ **Included in Woods 2.1:** install 2.1.0 or later for the discovery fixes below. For Git/path installs, verify the exact revision and loaded gem path; the declared development version alone does not establish availability.
432
548
 
433
549
  **Key details:**
434
- - Scans `app/graphql` with runtime introspection via `GraphQL::Schema.types` when available, falls back to file discovery
435
- - Produces unit types: `graphql_type`, `graphql_mutation`, `graphql_resolver`, `graphql_query`
550
+ - Scans every current, named application `GraphQL::Schema` subclass and unions their runtime type inventories by canonical Ruby constant name. Distinct Ruby classes can share a schema-local GraphQL name. A type used as the query root of any schema has one `graphql_query` identity.
551
+ - Produces unit types: `graphql_type`, `graphql_mutation`, `graphql_resolver`, `graphql_query`. Schema classes use `graphql_type` with `metadata.graphql_kind: "schema"`; their source includes configuration such as `max_complexity` for lookup and source search.
552
+ - Scans governed declarations in `app/graphql`, including unattached resolvers. Loaded declarations qualify through GraphQL ancestry, so application superclass chains, leading `::`, and superclass whitespace do not affect discovery. Reflection reads schema/type metadata without executing field or resolver bodies.
553
+ - File discovery also retains ordinary Ruby subclasses of the loaded `Resolvers::Base` class, including application inheritance chains. These helpers keep their historical `graphql_type` identity even when mutations call them with `new(...).call` rather than registering them as graphql-ruby resolvers. Woods verifies actual ancestry; placing an unrelated class in a `resolvers/` directory does not qualify it.
554
+ - When graphql-ruby or a declaration is unavailable, file discovery retains its limited recognized-source-form fallback; it cannot establish arbitrary application superclass ancestry. Woods does not trigger pending declaration autoloads itself.
436
555
  - Extracts field metadata (types, descriptions, complexity, arguments), authorization patterns (Pundit, CanCan, `authorized?`), and dependencies on models/services
437
- - Since all GraphQL units come from one extractor, incremental re-extraction handles them via `extract_graphql_file`
556
+ - Incremental extraction handles changed files and runtime inventory additions, and reclassifies an unchanged type when its query-root role changes. Runtime inventory absence alone does not delete units: unattached file-defined resolvers remain valid, and runtime-only removals still require a full extraction or `woods:refresh[graphql]` after the application reloads.
557
+ - A failed schema inventory logs the schema name and stops full, incremental, or refresh publication; the prior generation remains active. Fix the introspection error and retry the complete extraction batch.
558
+ - If application eager loading is incomplete, incremental extraction and refresh refuse a change to an existing GraphQL unit's kind rather than demoting a query root whose schema might not have loaded. Retry after a complete application boot.
559
+ - Runtime source locations take precedence over convention-named files. Initializer-built `Class.new` schemas/types can be indexed, but this does not add method-body `code_reference` ownership for generated declarations; see [source-reference coverage](#constant-source-references).
438
560
  - `parent_class` and summary chunks describe the selected declaration's explicit constant-path superclass, preserving its written qualification. Nested or sibling declarations and literal text cannot supply a parent. Implicit Object, module interfaces, dynamic superclass expressions, unavailable source, and invalid source have no declared parent (`null` metadata; `unknown` in summaries). This is source declaration metadata, not resolved runtime ancestry.
439
561
 
440
562
  **Example output (abbreviated):**
@@ -463,14 +585,23 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
463
585
  - Pairs policy units with their corresponding model (e.g., `UserPolicy` → `User`)
464
586
  - Extracts scope class and `resolve` method when present
465
587
 
588
+ **Included in Woods 2.1:** a loaded policy's actual ApplicationPolicy
589
+ ancestry, including namespaced and intermediate application bases, establishes
590
+ Pundit inheritance. The selected constant must belong to the source being read;
591
+ aliases and foreign same-named classes do not qualify. Without a loaded class,
592
+ direct ApplicationPolicy declarations and the existing user/record conventions
593
+ remain supported. Woods does not invoke pending autoloads to classify a policy.
594
+ The generic PolicyExtractor retains its existing policy units and applies the
595
+ same ancestry correction to `metadata.is_pundit`.
596
+
466
597
  ---
467
598
 
468
599
  ### PolicyExtractor
469
600
 
470
- **What it captures:** Domain policy classes (non-Pundit) with decision methods and eligibility rules. Covers plain Ruby objects used for authorization decisions.
601
+ **What it captures:** Policy classes with decision methods and eligibility rules, including plain Ruby objects used for authorization decisions.
471
602
 
472
603
  **Key details:**
473
- - Scans `app/policies` for files not identified as Pundit policies
604
+ - Scans `app/policies`; `metadata.is_pundit` identifies recognized Pundit-style classes
474
605
  - Extracts public predicate methods and their dependencies
475
606
 
476
607
  ---
@@ -525,7 +656,7 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
525
656
 
526
657
  #### Package membership
527
658
 
528
- Every app-owned unit under a package root carries `metadata[:package]` with the package name (longest root wins; `.` when only a root package exists). Framework sources and units with no path never carry it. The graph node carries the same value as `package`. When a `package.yml` changes, an incremental run re-annotates every unit whose package changed in the same run, so `metadata[:package]` never lags behind the file that defines it.
659
+ Every single-file app-owned unit under a package root carries `metadata[:package]` with the package name (longest root wins; `.` when only a root package exists). Framework sources and units with no path never carry it. The graph node carries the same value as `package`. When a `package.yml` changes, an incremental run re-annotates every unit whose package changed in the same run, so `metadata[:package]` never lags behind the file that defines it.
529
660
 
530
661
  ---
531
662
 
@@ -545,13 +676,22 @@ Every app-owned unit under a package root carries `metadata[:package]` with the
545
676
  **What it captures:** ActionCable channel classes with stream subscriptions, subscribed/unsubscribed hooks, broadcast patterns, and action methods.
546
677
 
547
678
  **Key details:**
548
- - Discovers via `ActionCable::Channel::Base.descendants`
679
+ - Discovers via `ActionCable::Channel::Base.descendants`. **Included in Woods 2.1:** discovery and direct extraction exclude dependency-owned source files.
549
680
  - Records stream names, authentication checks in `subscribed`, and any `broadcast_to` calls
550
681
 
551
682
  ---
552
683
 
553
684
  ### ScheduledJobExtractor
554
685
 
686
+ **Included in Woods 2.1:** unique schedule names retain `scheduled:<name>`.
687
+ Names shared across scheduler formats become `scheduled:<format>:<name>`, with
688
+ a deterministic suffix if that identifier is already a literal task name.
689
+ `metadata.task_name` preserves the original name. Ambiguous names within one
690
+ format refuse publication; the last generation stays usable. Run a full
691
+ extraction to remove dependency-owned components/channels and reconcile these
692
+ expanded mailer and schedule identities in an existing index.
693
+
694
+
555
695
  **What it captures:** Scheduled job definitions from cron-style config files. Supports multiple scheduling backends.
556
696
 
557
697
  **Key details:**
@@ -594,6 +734,29 @@ Every app-owned unit under a package root carries `metadata[:package]` with the
594
734
  - Risk indicators: data migrations (manual SQL or bulk updates), irreversible operations (`remove_column` without type), `execute` calls with raw SQL
595
735
  - Rails internal tables (`schema_migrations`, `active_storage_blobs`, etc.) are excluded from model dependency links
596
736
 
737
+ **Included in Woods 2.1:** migration identity comes from the actual Ruby
738
+ declaration, with the filename's conventional class name selecting among eligible
739
+ declarations. Helper classes and closed sibling namespaces cannot rename a
740
+ migration or contribute unrelated DDL metadata. Qualified declaration receivers
741
+ use verified lexical namespace ownership, including an existing outer or root
742
+ namespace; Woods does not prepend the syntactic nesting blindly. A namespace
743
+ established earlier in the same source can supply structural evidence. Unknown
744
+ receivers, pending autoloads, and unavailable lexical constant tables produce an
745
+ explicit ownership diagnostic rather than an invented identity. Leading
746
+ `::ActiveRecord::Migration` remains supported. A custom migration base requires
747
+ already-loaded runtime ancestry or a same-file structural chain to
748
+ ActiveRecord::Migration; unknown custom bases are not guessed. Historical files
749
+ are never required or evaluated for discovery. Ambiguous declarations produce an
750
+ explicit extraction error, and duplicate identities across files retain the
751
+ normal collision guard. Run a full extraction after upgrading to replace any
752
+ previously misidentified migration units.
753
+
754
+ If a qualified receiver cannot be verified, define its namespace before the
755
+ declaration in the same source, or make that namespace available through normal
756
+ application boot. Use a leading `::` when the receiver intentionally belongs to
757
+ the root namespace. Woods will not load a historical migration to discover its
758
+ namespace.
759
+
597
760
  **Example output (abbreviated):**
598
761
 
599
762
  ```json
@@ -631,8 +794,9 @@ Every app-owned unit under a package root carries `metadata[:package]` with the
631
794
  **What it captures:** State machine DSL definitions using AASM, Statesman, or the `state_machines` gem.
632
795
 
633
796
  **Key details:**
634
- - Detects which library is active by checking `defined?` for each DSL constant
635
- - Extracts states, events, transitions, guard conditions, and callbacks
797
+ - Detects literal DSL declarations in model source; it does not evaluate the DSL or run callbacks
798
+ - Extracts states, events, transitions, guard conditions, and callbacks from supported source forms; it does not claim complete dynamic or inherited registry coverage
799
+ - **Included in Woods 2.1:** directly declared `state_machines` calls in the selected model class support the default `state` attribute (`state_machine initial: :pending do`), explicit attributes, and parenthesized calls, including multiline arguments. The default produces `Model::state_machine_state`; existing named identifiers remain `Model::state_machine_<attribute>`. Each declaration uses its own block and literal initial state, keeping multiple machines separate. Nested/sibling classes, singleton scopes and deferred/receiver blocks cannot supply another model's machine. Dynamic attribute expressions are not guessed, and dynamic initial-state functions are not called.
636
800
  - Returns an array from the file method (like `ScheduledJobExtractor`), cannot be used in the incremental file-based dispatch map; incremental re-extraction re-runs it wholesale on any `.rb` change under the model directories it scans
637
801
 
638
802
  ---
@@ -645,6 +809,7 @@ Every app-owned unit under a package root carries `metadata[:package]` with the
645
809
  - Two-pass approach: first collects all `publish`/`instrument` calls, then `subscribe`/`on` calls, then merges them
646
810
  - No single-file extraction method, incremental re-extraction re-runs `EventExtractor` wholesale on any `.rb` change under `app/` (a publish or subscribe site can appear anywhere)
647
811
  - Useful for tracing event-driven flows: "what subscribes to order.created?"
812
+ - **Included in Woods 2.1:** app-owned `metadata.publishers` and `metadata.subscribers` paths, and the same paths in generated source annotations, are relative to `Rails.root`. Their array order, counts and event identifiers stay unchanged. Explicitly scanned paths outside the application remain absolute. Re-extract event units after upgrading: removing the checkout prefix changes existing `source_hash` values once, then identical app sources produce the same annotations and hashes across checkout roots.
648
813
 
649
814
  ---
650
815
 
@@ -694,6 +859,39 @@ Every app-owned unit under a package root carries `metadata[:package]` with the
694
859
  - File-based scanning; no assumption about class hierarchy
695
860
  - `parent_class` and the generated Parent annotation describe the selected unit declaration only. Nested or sibling classes cannot supply its parent. An implicit `Object` parent, a dynamic superclass expression, or unparseable source produces `nil`; explicit constant-path parents retain their source names.
696
861
 
862
+ **Included in Woods 2.1:** compatible reopened library classes/modules across
863
+ files form one typed unit. Matching inferred names alone are insufficient:
864
+ conflicting declaration kinds/parents, unresolved constructors and aliases still
865
+ refuse ambiguous ownership. Extraction does not require or execute unmanaged
866
+ files. The sorted first path remains the compatibility `file_path`; it is a
867
+ primary display path, not a claim that every definition resides there.
868
+
869
+ Aggregates add `metadata.source_contributors_version: 1`, `source_contributors`,
870
+ and sorted `defined_in`. Each contributor records its own path, raw-file SHA256,
871
+ physical line range and half-open byte range inside the published composite.
872
+ Generated headers/separators lie outside those ranges. Per-file facts remain
873
+ separate; sorted paths do not establish runtime override order. The unit's
874
+ `source_hash` describes the composite. One-file units retain their prior shape.
875
+
876
+ The graph maps every contributor to the typed owner. Any relevant library edit,
877
+ deletion or dependency-triggered refresh reconciles the complete library family
878
+ once per batch; direct file extraction also returns the aggregate. This costs
879
+ more than re-reading only the primary file, and preserves full/incremental facts.
880
+ A contributor read failure prevents publishing a partial aggregate. Supporting
881
+ Woods 2.1 writers read library source explicitly as UTF-8 regardless of the
882
+ process locale; invalid UTF-8 is logged with the physical source path and still
883
+ prevents a partial aggregate.
884
+
885
+ Git facts remain per contributor: Woods does not fabricate a summed aggregate
886
+ commit count or churn rank. A unit-level package is emitted only when every
887
+ contributor belongs to the same package. Retrieval package/path scopes require
888
+ **all** contributors to match, avoiding disclosure of out-of-scope source.
889
+ Run one full extraction after upgrading: source-reference cache format 3 must
890
+ be rebuilt, and existing indexes do not yet contain contributor provenance.
891
+
892
+ Loaded Struct/Data assignments use the [assigned value-class ownership rules](#assigned-value-classes)
893
+ to avoid naming sibling files after their shared namespace wrapper.
894
+
697
895
  ---
698
896
 
699
897
  ### RailsSourceExtractor
data/docs/FAQ.md CHANGED
@@ -143,7 +143,7 @@ bundle exec rake woods:incremental
143
143
  docker compose exec app bundle exec rake woods:incremental
144
144
  ```
145
145
 
146
- The default Git range is `HEAD~1`; CI variables, an explicit range, or
146
+ The default Git range is `HEAD~1`; supported CI variables or
147
147
  `CHANGED_FILES` can select another batch. Incremental extraction also refreshes
148
148
  affected concern consumers and re-runs whole-app extractors when their trigger
149
149
  paths change. It can reduce work, but has no universal speedup: Rails boot, graph
@@ -449,7 +449,7 @@ Enable them in your initializer:
449
449
  config.enable_snapshots = true
450
450
  ```
451
451
 
452
- Snapshots prefer their own SQLite database (`woods.sqlite3` in the output directory), separate from your Rails app's database. Extraction falls back to JSON files when the `sqlite3` gem is unavailable; other SQLite open/migration failures are reported and do not capture a snapshot. `Woods::Db::Migrator` runs the internal SQLite migrations automatically during extraction and MCP boot; `bundle exec rails db:migrate` does not touch this store and no manual migration step is needed. The packaged MCP server discovers an existing `woods.sqlite3` automatically. When extraction used the JSON fallback, set `WOODS_SNAPSHOTS=true` on the server so it wires `list_snapshots`, `snapshot_diff`, `unit_history`, and `snapshot_detail`.
452
+ Snapshots prefer their own SQLite database (`woods.sqlite3` in the output directory), separate from your Rails app's database. Extraction falls back to JSON files when the `sqlite3` gem is unavailable; other SQLite open/migration failures are reported and do not capture a snapshot. `Woods::Db::Migrator` runs the internal SQLite migrations automatically during extraction and MCP boot; `bundle exec rails db:migrate` does not touch this store and no manual migration step is needed. The packaged MCP server discovers an existing `woods.sqlite3` automatically. `WOODS_SNAPSHOTS=true` enables snapshot tools but still prefers SQLite; it does not force the JSON store or import its history. If SQLite became available after JSON capture, use an explicitly configured JSON reader for that history. See [snapshot store selection](MCP_SERVERS.md#conditional-index-capabilities).
453
453
 
454
454
  ---
455
455
 
@@ -457,7 +457,11 @@ Snapshots prefer their own SQLite database (`woods.sqlite3` in the output direct
457
457
 
458
458
  ### What does the session tracer do?
459
459
 
460
- The session tracer is middleware that records which Rails actions are invoked during a browser session, assembles the relevant extracted units, and makes that context available via the `session_trace` MCP tool. It is useful for giving an AI tool accurate context about what code path was active during a specific user interaction.
460
+ The session tracer records which Rails actions run during a browser session.
461
+ A custom/embedded Index Server configured with that store can expose the traces
462
+ through `session_trace`. The packaged Index executable does not load Rails
463
+ initializers, so enabling the application middleware alone does not add this tool.
464
+ See [conditional capabilities](MCP_SERVERS.md#conditional-index-capabilities).
461
465
 
462
466
  Session tracing is disabled by default and requires an explicit `session_store`.
463
467
  Follow the [canonical configuration example](CONFIGURATION_REFERENCE.md#session-tracer-options)
@@ -469,28 +473,10 @@ for store construction and review trace retention and access controls before ena
469
473
 
470
474
  ### How do I keep the index in sync in CI?
471
475
 
472
- Use incremental extraction in your CI pipeline. Fetch enough git history for the incremental diff to work:
473
-
474
- ```yaml
475
- # .github/workflows/index.yml
476
- jobs:
477
- index:
478
- steps:
479
- - uses: actions/checkout@v4
480
- with:
481
- fetch-depth: 2
482
- - name: Update index
483
- run: bundle exec rake woods:incremental
484
- env:
485
- GITHUB_BASE_REF: ${{ github.base_ref }}
486
- ```
487
-
488
- For Docker-based CI:
489
-
490
- ```yaml
491
- - name: Update index
492
- run: docker compose exec -T app bundle exec rake woods:incremental
493
- ```
476
+ Restore an index for the exact commit your incremental diff starts from;
477
+ otherwise run full extraction. Fetch the actual pull-request base and its
478
+ history before diffing. Follow the [incremental CI recipe](INCREMENTAL_EXTRACTION.md#github-actions-with-an-exact-baseline)
479
+ for cache selection, cold starts, nested applications, and Docker.
494
480
 
495
481
  ---
496
482
 
@@ -65,6 +65,12 @@ bin/rails woods:extract
65
65
 
66
66
  Woods boots and eager-loads the Rails application, runs its extractors, builds dependency edges, and publishes one complete generation under `tmp/woods/`. A reader stays on the previous complete generation until the new one is published.
67
67
 
68
+ For verified source capture before Rails boots, use the installed
69
+ `bundle exec woods-extract full` launcher; see [source freshness](SOURCE_FRESHNESS.md#establish-a-fresh-baseline)
70
+ for custom roots and output paths. **Included in Woods 2.1:** the launcher
71
+ prefers the application's executable `bin/rake`, falling back to
72
+ `bundle exec rake` when it is absent or not executable.
73
+
68
74
  If Rails only boots with environment variables, provide the same variables here. Do not work around a boot failure inside Woods configuration. First confirm that the application can boot and eager-load with the same environment:
69
75
 
70
76
  ```bash
@@ -142,7 +148,7 @@ ollama pull nomic-embed-text
142
148
  bin/rails woods:embed
143
149
  ```
144
150
 
145
- For dense Ruby source, add `gem "tokenizers", "~> 0.5"` for exact WordPiece token counting. Without it, Woods falls back to character-based estimation, which can over-pack some Ollama chunks.
151
+ Ollama input counts are estimates, not a universal exact tokenizer. Supporting builds after 2.0.0 split complete prefixed inputs and request `truncate: false`; installing a tokenizer gem alone does not select a model-matched tokenizer. See [input sizing and refusal](EMBEDDING_MODELS.md#why-num_ctx-isnt-enough).
146
152
 
147
153
  Reconnect the MCP server and check `woods_status`. For OpenAI, pgvector, Qdrant, model dimensions, and provider changes, read the [Retrieval guide](RETRIEVAL_GUIDE.md) and [Backend matrix](BACKEND_MATRIX.md).
148
154