woods 2.0.0.beta3 → 2.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (120) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +500 -420
  3. data/CONTRIBUTING.md +29 -17
  4. data/README.md +78 -178
  5. data/docs/AGENT_GUIDE.md +52 -11
  6. data/docs/AGENT_SETUP.md +34 -17
  7. data/docs/AUTOMATIC_MAINTENANCE.md +222 -0
  8. data/docs/BACKEND_MATRIX.md +18 -7
  9. data/docs/CLIENT_HOOKS.md +1 -1
  10. data/docs/CONFIGURATION_REFERENCE.md +105 -29
  11. data/docs/CONSOLE_MCP_SETUP.md +54 -9
  12. data/docs/DOCKER_SETUP.md +16 -1
  13. data/docs/EVALUATION.md +10 -4
  14. data/docs/EXTRACTOR_REFERENCE.md +23 -3
  15. data/docs/FAQ.md +14 -3
  16. data/docs/GETTING_STARTED.md +18 -17
  17. data/docs/INCREMENTAL_EXTRACTION.md +37 -8
  18. data/docs/INDEX_LAYOUT.md +2 -2
  19. data/docs/MCP_SERVERS.md +79 -7
  20. data/docs/MCP_TOOL_COOKBOOK.md +5 -5
  21. data/docs/MCP_WORKTREE_SETUP.md +55 -83
  22. data/docs/PUBLISHED_INDEX.md +17 -0
  23. data/docs/README.md +2 -1
  24. data/docs/RETRIEVAL_GUIDE.md +81 -13
  25. data/docs/SOURCE_FRESHNESS.md +1 -1
  26. data/docs/TOKEN_BENCHMARK.md +16 -10
  27. data/docs/TROUBLESHOOTING.md +142 -47
  28. data/docs/UPGRADING_TO_2.md +12 -6
  29. data/docs/WATCH_DAEMON.md +189 -24
  30. data/docs/WHY_WOODS.md +9 -5
  31. data/exe/woods-console +13 -11
  32. data/exe/woods-mcp-start +14 -9
  33. data/exe/woods-watch +5 -0
  34. data/lib/generators/woods/pgvector_generator.rb +8 -2
  35. data/lib/generators/woods/watch_generator.rb +53 -0
  36. data/lib/puma/plugin/woods.rb +10 -0
  37. data/lib/tasks/woods.rake +14 -0
  38. data/lib/woods/agent_configuration/applier.rb +5 -3
  39. data/lib/woods/agent_configuration/cli.rb +2 -2
  40. data/lib/woods/agent_configuration/layout.rb +13 -0
  41. data/lib/woods/cache/cache_middleware.rb +6 -0
  42. data/lib/woods/console/credential_scanner.rb +4 -3
  43. data/lib/woods/console/dispatch_pipeline.rb +7 -0
  44. data/lib/woods/console/embedded_executor.rb +31 -9
  45. data/lib/woods/console/sql_noise_stripper.rb +9 -7
  46. data/lib/woods/console/sql_table_scanner.rb +47 -7
  47. data/lib/woods/console/sql_validator.rb +49 -9
  48. data/lib/woods/console/sqlite_read_guard.rb +46 -0
  49. data/lib/woods/console/stdio_transport.rb +27 -0
  50. data/lib/woods/coordination/pipeline_lock.rb +3 -2
  51. data/lib/woods/embedding/indexer.rb +24 -14
  52. data/lib/woods/extractor.rb +70 -19
  53. data/lib/woods/extractors/declared_parent.rb +55 -0
  54. data/lib/woods/extractors/graphql_extractor.rb +2 -11
  55. data/lib/woods/extractors/lib_extractor.rb +10 -8
  56. data/lib/woods/extractors/mailer_extractor.rb +6 -10
  57. data/lib/woods/extractors/model_extractor.rb +1 -15
  58. data/lib/woods/extractors/poro_extractor.rb +10 -8
  59. data/lib/woods/extractors/shared_utility_methods.rb +22 -5
  60. data/lib/woods/git_command.rb +6 -7
  61. data/lib/woods/git_provenance.rb +4 -6
  62. data/lib/woods/mcp/bearer_auth.rb +2 -1
  63. data/lib/woods/mcp/bootstrapper.rb +20 -5
  64. data/lib/woods/mcp/config_resolver.rb +2 -1
  65. data/lib/woods/mcp/index_reader.rb +11 -2
  66. data/lib/woods/mcp/initialization_guidance.rb +1 -1
  67. data/lib/woods/mcp/renderers/markdown_renderer.rb +14 -8
  68. data/lib/woods/mcp/renderers/plain_renderer.rb +11 -7
  69. data/lib/woods/mcp/server.rb +63 -37
  70. data/lib/woods/mcp/tool_contract.rb +1 -1
  71. data/lib/woods/mcp/tool_response_renderer.rb +16 -0
  72. data/lib/woods/mcp/traversal_evidence_text.rb +1 -1
  73. data/lib/woods/mcp/traversal_response.rb +22 -0
  74. data/lib/woods/path_dispatcher.rb +6 -5
  75. data/lib/woods/published_index/typed_unit_reader.rb +40 -3
  76. data/lib/woods/published_index.rb +2 -2
  77. data/lib/woods/rake_helpers.rb +2 -12
  78. data/lib/woods/retrieval/corpus_status.rb +46 -0
  79. data/lib/woods/retrieval/lexical_assembler.rb +14 -3
  80. data/lib/woods/retrieval/lexical_index.rb +2 -1
  81. data/lib/woods/retriever.rb +19 -7
  82. data/lib/woods/session_tracer/file_store.rb +6 -1
  83. data/lib/woods/source_inputs/consumer_errors.rb +4 -0
  84. data/lib/woods/storage/local_corpus_stats.rb +32 -0
  85. data/lib/woods/storage/metadata_store.rb +20 -0
  86. data/lib/woods/storage/pgvector.rb +6 -2
  87. data/lib/woods/storage/vector_store.rb +10 -0
  88. data/lib/woods/temporal/json_snapshot_store.rb +35 -7
  89. data/lib/woods/version.rb +1 -1
  90. data/lib/woods/watch/child_environment.rb +30 -0
  91. data/lib/woods/watch/cli.rb +91 -0
  92. data/lib/woods/watch/daemon.rb +73 -11
  93. data/lib/woods/watch/event_stream.rb +70 -0
  94. data/lib/woods/watch/guardian.rb +142 -0
  95. data/lib/woods/watch/installation/layout.rb +70 -0
  96. data/lib/woods/watch/installation/options.rb +128 -0
  97. data/lib/woods/watch/installation/planner.rb +128 -0
  98. data/lib/woods/watch/installation/probe.rb +101 -0
  99. data/lib/woods/watch/installation/receipt.rb +77 -0
  100. data/lib/woods/watch/installation/recovery.rb +64 -0
  101. data/lib/woods/watch/installation/templates.rb +58 -0
  102. data/lib/woods/watch/installation.rb +56 -0
  103. data/lib/woods/watch/lifecycle.rb +182 -0
  104. data/lib/woods/watch/managed_child.rb +113 -0
  105. data/lib/woods/watch/managed_cleanup.rb +48 -0
  106. data/lib/woods/watch/managed_process.rb +144 -0
  107. data/lib/woods/watch/puma_adapter.rb +87 -0
  108. data/lib/woods/watch/puma_child.rb +66 -0
  109. data/lib/woods/watch/supervision_records.rb +95 -0
  110. data/lib/woods/watch/supervision_status.rb +104 -0
  111. data/lib/woods/watch/supervisor.rb +161 -0
  112. data/lib/woods/watch/supervisor_reporting.rb +46 -0
  113. data/plugin/.claude-plugin/plugin.json +1 -1
  114. data/plugin/hooks/woods-input-rules.sh +4 -4
  115. data/plugin/skills/woods-agent-enable/SKILL.md +7 -1
  116. data/plugin/skills/woods-diagnose/SKILL.md +134 -34
  117. data/plugin/skills/woods-investigate/SKILL.md +54 -15
  118. data/plugin/skills/woods-mcp-config/SKILL.md +38 -11
  119. data/plugin/skills/woods-setup/SKILL.md +72 -15
  120. metadata +38 -5
@@ -208,6 +208,25 @@ config.vector_store_options = {
208
208
  }
209
209
  ```
210
210
 
211
+ Woods uses an HNSW index over pgvector's `vector` representation, which supports
212
+ **1–2,000 dimensions** ([pgvector's HNSW limits](https://github.com/pgvector/pgvector#hnsw)).
213
+ Included in Woods `2.0.0`: the adapter rejects wider dimensions before any
214
+ SQL, and `woods:pgvector` rejects invalid widths before writing a migration.
215
+ There is no automatic vector truncation or half-precision conversion.
216
+
217
+ The default `text-embedding-3-large` output is 3,072 dimensions. With pgvector,
218
+ explicitly request a supported provider output width, for example:
219
+
220
+ ```ruby
221
+ config.embedding_model = 'text-embedding-3-large'
222
+ config.embedding_options = { dimensions: 1536 }
223
+ ```
224
+
225
+ Keep the provider, `vector_store_options[:dimensions]` (when set), and generated
226
+ migration width equal. Changing a stored width requires a compatible new table
227
+ or an intentional index rebuild; changing the setting does not resize old data.
228
+ Use another backend if you need the full 3,072-dimensional output.
229
+
211
230
  Requires the pgvector extension. Run the generator to create migrations:
212
231
 
213
232
  ```bash
@@ -263,7 +282,12 @@ in the promoted dump; the index MCP server loads that snapshot at startup or
263
282
  reload. Incremental embedding publishes changes to paths, dependencies, and
264
283
  other unit metadata even when unchanged source needs no new embedding. A run
265
284
  with no content or metadata changes keeps the existing dump and retention
266
- window. This is a reasonable default for hosts that don't bundle `sqlite3`.
285
+ window. Included in Woods `2.0.0`: a full `Indexer#index_all` run replaces
286
+ the published corpus even when a custom caller reuses in-memory vector and
287
+ metadata stores. Deleted units, including metadata-only records, are removed;
288
+ an empty full rebuild publishes an empty dump. Failed embedding leaves the
289
+ previous promoted dump and checkpoint intact. Incremental purge guards remain
290
+ unchanged. This is a reasonable default for hosts that don't bundle `sqlite3`.
267
291
 
268
292
  ## Retrieval cache options
269
293
 
@@ -335,7 +359,7 @@ The embed run writes `woods.json` + `dumps/<ISO8601>/vectors.bin` + `metadata.ms
335
359
 
336
360
  Requirements:
337
361
  - `output_dir` must be set and readable by both the embed process and the MCP server.
338
- - The MCP server must know the same `output_dir` (pass via `woods-mcp <DIR>` or set `WOODS_DIR`).
362
+ - The MCP server must know the same `output_dir` (pass via `woods-mcp <DIR>` or set `WOODS_DIR`; see MCP path precedence below).
339
363
 
340
364
  ## Presets
341
365
 
@@ -418,11 +442,26 @@ finding qualifying edges. Run extraction again after changing either setting;
418
442
  MCP reads the published report. These findings remain informational, never a
419
443
  release or architecture gate.
420
444
 
421
- When the JSON snapshot fallback is in use, malformed JSON, top-level values
422
- other than objects, and files that cannot be read (including concurrent retention
423
- removals) are warned about and treated as absent. Snapshot lists and unit history
424
- omit them; direct lookup returns no snapshot, and a diff with an unavailable
425
- snapshot returns empty added, modified, and deleted lists.
445
+ When the JSON snapshot fallback is in use, malformed JSON, invalid snapshot
446
+ shapes, and files that cannot be read (including concurrent retention removals)
447
+ are warned about and treated as absent. A snapshot needs a hexadecimal string
448
+ `git_sha` matching its filename; `extracted_at` may be a string, null, or omitted.
449
+ When present and non-null, `units` must be an object whose records are objects.
450
+ A malformed record invalidates the entire snapshot, rather than exposing partial
451
+ history. Legacy bare identifier keys, omitted/null unit collections, and optional per-unit hash
452
+ fields remain supported; timestamp strings are not restricted to a new format.
453
+ Unit history limits count matching unit records, not the most recent snapshots
454
+ searched (JSON fallback correction included in Woods `2.0.0`). A unit
455
+ missing from newer snapshots can still have retained history.
456
+ Snapshot lists and unit history omit unusable files; direct lookup returns no
457
+ snapshot, and a diff with an unavailable snapshot returns empty added, modified,
458
+ and deleted lists. An empty diff in this case is not proof that nothing changed.
459
+ New captures compare against the latest usable snapshot. Unusable SHA-named files
460
+ still count toward retention and are pruned first when the limit is exceeded;
461
+ reading alone does not delete them. Valid legacy snapshots with null or omitted
462
+ timestamps are retained ahead of corrupt files, then treated as oldest among
463
+ usable snapshots. Direct lookup and diff still reject invalid
464
+ caller-supplied SHA paths with an argument error.
426
465
 
427
466
  `incremental_blast_radius_depth` is unbounded by default because a unit two hops
428
467
  out really can have content that depends on the changed file. An STI grandchild
@@ -473,6 +512,18 @@ config.session_store = Woods::SessionTracer::FileStore.new(
473
512
  config.session_exclude_paths = ['/health', '/metrics', '/assets']
474
513
  ```
475
514
 
515
+ ### File session retention
516
+
517
+ `FileStore` accepts `ttl:` in seconds (default `nil`, expiration disabled),
518
+ `max_sessions:` (default `1000`), and `max_requests_per_session:` (default `1000`).
519
+ TTL expires a file when the store clock reaches its modification time plus the
520
+ TTL. Recording after expiry starts a fresh history; expired events are discarded
521
+ before appending or migrating legacy filenames, under the same store lock.
522
+ When legacy and encoded files coexist, each expires independently before any
523
+ surviving histories are merged. Clearing a session is idempotent for supported
524
+ IDs, including Unicode and punctuation, and removes both filename formats when
525
+ applicable.
526
+
476
527
  ### Redis session index compatibility
477
528
 
478
529
  `RedisStore` lists and clears both legacy SET indexes and recency ZSET indexes.
@@ -652,7 +703,8 @@ These variables are read by the gem and its MCP servers at runtime. They complem
652
703
  | Variable | Default | Purpose |
653
704
  |----------|---------|---------|
654
705
  | `WOODS_RETRIEVAL_MODE` | `semantic` | Explicit packaged MCP retrieval mode: `semantic` or `lexical`. Lexical reads extraction unit JSON without provider autodetection, credentials or vector artifacts. |
655
- | `WOODS_DIR` | `Dir.pwd` | Path to the extraction output directory. |
706
+ | `WOODS_DIR` | unset | MCP extraction-index path, after a positional argument and before `WOODS_OUTPUT`. See precedence below. |
707
+ | `WOODS_OUTPUT` | unset | MCP index-path fallback when neither a positional path nor `WOODS_DIR` is set; included in Woods `2.0.0`. |
656
708
  | `WOODS_REQUIRE_INDEX` | unset | Set to `"1"` to fail closed: the server refuses to boot (raises `MissingArtifact`) unless a real index (`woods.json`) is present. By default an extract-only host boots in pattern/structural mode without it. Explicit lexical mode requires a valid published extraction index, not `woods.json`. |
657
709
  | `WOODS_ALLOW_AUTODETECT` | unset | **Deprecated no-op.** Auto-detect is now the default; accepted for backward compatibility only. |
658
710
  | `WOODS_SEARCH_MAX_SCAN` | `500` | Cap on unit files loaded during a phase-2 (metadata/source_code) `search`. Hitting the cap sets `partial: true` in the response. |
@@ -668,6 +720,20 @@ These variables are read by the gem and its MCP servers at runtime. They complem
668
720
  | `WOODS_QDRANT_URL`, `WOODS_QDRANT_COLLECTION`, `WOODS_QDRANT_API_KEY` | n/a | Override/require Qdrant connection settings when a pgvector/Qdrant-backed index is served outside its host application (no `Woods.configuration` available). |
669
721
  | `WOODS_PG_URL` | n/a | Required when a pgvector-backed index is served outside its host application. |
670
722
 
723
+ **MCP index path precedence (included in Woods `2.0.0`):** positional
724
+ argument → `WOODS_DIR` → `WOODS_OUTPUT` → current directory for `woods-mcp`
725
+ and `woods-mcp-http`. `woods-mcp-start` still requires one of the first three;
726
+ it never silently selects the current directory. An explicitly empty
727
+ `WOODS_DIR` remains an invalid override rather than falling through. Earlier
728
+ versions accept the positional path or `WOODS_DIR`; use an explicit path for
729
+ portable client configuration.
730
+
731
+ Paths are resolved in the MCP process's working directory and filesystem.
732
+ A published index can have `generation.json` pointing to a payload's
733
+ `manifest.json`; a root `manifest.json` is only the legacy flat layout. If
734
+ startup cannot find a manifest, check the reported directory and point at the
735
+ existing index before deciding another extraction is needed.
736
+
671
737
  ### Rake tasks
672
738
 
673
739
  | Variable | Default | Purpose |
@@ -705,10 +771,24 @@ These variables are read by the gem and its MCP servers at runtime. They complem
705
771
  | `WOODS_WATCH_CATCH_UP` | `1` (enabled) | Set to `"0"` to skip generation-watermark catch-up on daemon start. |
706
772
  | `WOODS_WATCH_TRUST_FOREIGN_HOST` | unset (disabled) | Set to `"1"` in each task/MCP reader to trust a foreign daemon's heartbeat for up to 15 minutes, without a local pid check. See [cross-host liveness](WATCH_DAEMON.md#cross-host-liveness) for clock bounds, degraded coverage, and startup limitations. |
707
773
 
774
+ ### Managed watcher startup
775
+
776
+ **Included in Woods `2.0.0`.** `woods-watch` wraps the raw task for native
777
+ Foreman/Puma lifecycle management. `--root PATH` selects the application working
778
+ directory; `--boot-timeout SECONDS` defaults to `300` and bounds boot/handshake,
779
+ not extraction. Explicit child argv follows `--`. Output-directory precedence
780
+ remains `WOODS_OUTPUT`, then `Woods.configuration.output_dir`.
781
+
782
+ Managed mode requires `WOODS_WATCH_IDLE_TIMEOUT` to be unset, preserves one owner,
783
+ and never takes over a conflicting daemon. The optional Puma adapter only starts
784
+ in its finalized development environment. See [startup and installation](WATCH_DAEMON.md#managed-development-startup)
785
+ for the generator's explicit modes, portable receipt, update/removal, and the
786
+ separate supervision status. Raw task settings above remain compatible.
787
+
708
788
  ### Opt-in plugin refresh hooks
709
789
 
710
790
  These settings control the plugin shell worker. Check installed
711
- `woods:hook_refresh` support first; the task is unreleased after 2.0.0.beta2.
791
+ `woods:hook_refresh` support first; the task is included in Woods `2.0.0`.
712
792
  See [hook coverage and retry](WATCH_DAEMON.md#hooks-for-agent-sessions) and
713
793
  [optional context limits](WATCH_DAEMON.md#optional-bounded-context-hints).
714
794
 
@@ -741,26 +821,23 @@ tasks for manual refreshes; hook transport is not a general shell execution API.
741
821
  | `GITHUB_BASE_REF` | unset (GitHub Actions) | Build the diff range `origin/<ref>...HEAD` for `woods:incremental`; an unfetched ref makes the range unresolvable, same exit behavior. |
742
822
  | `RAILS_ENV` | `development` | Rails environment the rake tasks boot in. |
743
823
  | `WOODS_PROFILE` | unset | Set to `"1"` to log disjoint `[Woods] [profile] <phase> in N.NNs` durations, including git enrichment, reconciliation, payload sync, pointer publication (`publish`) and retention (`payload prune`). Separate `[profile total]` lines report whole extraction wall time, including unprofiled setup and failed runs; never add these totals to phase durations. Excludes process/Rails boot before extraction. Off by default. |
744
- | `WOODS_GIT_DIR` | unset | Absolute path to the canonical git directory. Wins over the repository Woods would otherwise find, at all three of its git call sites: per-unit `commit_count`/`change_frequency` (enrichment), `manifest.json`'s `git_branch`/`git_sha` (provenance), and the `woods:incremental` diff range. All three build their command line with `Woods::GitCommand.argv`. |
824
+ | `WOODS_GIT_DIR` | unset | Absolute path passed directly to Git's `--git-dir`, selecting that directory's `HEAD`. For a linked worktree use its `worktrees/<id>` directory inside the complete shared Git layout. Wins over the repository Woods would otherwise find, at all three of its git call sites: per-unit `commit_count`/`change_frequency` (enrichment), `manifest.json`'s `git_branch`/`git_sha` (provenance), and the `woods:incremental` diff range. All three build their command line with `Woods::GitCommand.argv`. |
745
825
  | `GIT_BRANCH`, `GIT_SHA` | unset | Provenance for a checkout with no `.git` at all (a source tarball, a Docker `COPY` that excludes it). Ignored when a `.git` is present but unresolvable, so a stale build arg cannot mask a worktree. |
746
826
 
747
- **`GIT_DIR` alone is not enough for a linked git worktree.** Woods runs git as a
748
- subprocess, so git's own `GIT_DIR` and `GIT_COMMON_DIR` are honored wherever git
749
- honors them. But pointing `GIT_DIR` at a worktree's *private* git directory only
750
- moves the failure: that directory reaches the shared object store through a
751
- relative `commondir` pointer, which still resolves outside a container mount,
752
- and `GIT_COMMON_DIR` does not override it. `git rev-parse --git-dir` then
753
- succeeds while no ref resolves.
754
-
755
- Woods refuses to enrich in that state rather than writing `commit_count: 0` and
756
- `change_frequency: "new"` on every unit: the git keys are omitted, provenance is
757
- `"unknown"`, and one warning names the cause. Point `WOODS_GIT_DIR` at the
758
- canonical git directory (the one a worktree's `gitdir:` pointer ultimately leads
759
- to) and mount it:
760
-
761
- ```bash
762
- WOODS_GIT_DIR=/canonical-git bundle exec rake woods:extract
763
- ```
827
+ For linked worktrees in containers, preserve access to the complete shared
828
+ Git directory and the worktree-specific HEAD. A same-path mount that resolves
829
+ the existing `.git` pointer needs no override. For a relocated complete layout,
830
+ set `WOODS_GIT_DIR=/mounted-common/worktrees/<id>`, deriving `<id>` from Git's
831
+ worktree metadata rather than the branch name. Selecting `/mounted-common`
832
+ itself selects the primary checkout's HEAD and can produce incorrect history,
833
+ provenance, and incremental paths. See the canonical
834
+ [worktree mount and verification steps](TROUBLESHOOTING.md#git-directory-mounts-for-linked-worktrees).
835
+
836
+ Git subprocesses inherit Git's own environment variables, so inspect existing
837
+ `GIT_DIR` and `GIT_COMMON_DIR` settings when resolving layout problems.
838
+ If HEAD cannot be resolved, enrichment omits the Git keys and provenance is
839
+ `"unknown"` for a present but unresolvable `.git`; one warning names the cause.
840
+ After correcting the layout, run a full extraction to refresh retained metadata.
764
841
 
765
842
  ### Exporters
766
843
 
@@ -780,11 +857,10 @@ The `woods-mcp` bootstrapper emits a one-line STDERR banner at startup indicatin
780
857
 
781
858
  ## Git enrichment history
782
859
 
783
- Current source requires **Git 2.31 or newer** for optional per-unit git
860
+ Woods 2.0 requires **Git 2.31 or newer** for optional per-unit git
784
861
  metadata. Extraction still succeeds when git is unavailable or history cannot
785
862
  be read completely. Git enrichment is omitted in either case; a failed or
786
863
  incomplete streamed history read logs a warning.
787
- This requirement and the history policy below are unreleased after 2.0.0.beta2.
788
864
 
789
865
  Per-unit enrichment also requires a non-shallow repository. A shallow checkout
790
866
  or a failed repository-depth probe omits enrichment with one warning per
@@ -48,7 +48,7 @@ its token, allowed origins and TLS as described in [Option C](#option-c-http-rac
48
48
 
49
49
  The rake task does two things before starting the MCP server:
50
50
 
51
- 1. **Captures stdout before Rails boots.** Rails boot emits OpenTelemetry warnings, gem notices, and other output to stdout. An MCP client cannot parse these as JSON-RPC, they break the protocol. The rake task redirects stdout → stderr immediately, saves the real stdout fd, and restores it after boot completes.
51
+ 1. **Captures stdout before Rails boots.** Rails boot emits OpenTelemetry warnings, gem notices, and other output to stdout. An MCP client cannot parse these as JSON-RPC. The rake task saves the protocol output and redirects application stdout to stderr. In the runtime isolation fix (included in Woods `2.0.0`), that redirection remains active throughout the server's lifetime; only the MCP transport writes to the saved protocol pipe. Rails loggers, `puts`, and writes to standard output during queries stay on stderr.
52
52
  2. **Calls `Rails.application.eager_load!`** to load all application models. Without eager loading, only the models that happen to be autoloaded before the first query appear in the registry.
53
53
 
54
54
  ### MCP client configuration
@@ -83,7 +83,7 @@ rake woods:console
83
83
  │ ├─ Rails.application.eager_load!
84
84
  │ ├─ build model registry from ActiveRecord::Base.descendants
85
85
  │ ├─ Server.build_embedded(model_validator:, safe_context:, ...)
86
- │ └─ MCP::Server::Transports::StdioTransport.new(server).open
86
+ │ └─ Woods::Console::StdioTransport.new(server, output: protocol_out).open
87
87
  │
88
88
  └─ MCP server responds to tool calls via stdin/stdout
89
89
  ```
@@ -198,6 +198,12 @@ the Rails server environment. The middleware stack registers automatically via
198
198
  the gem's Railtie and requires `Authorization: Bearer <token>` on every Console
199
199
  request. Missing or incorrect tokens receive `401 Unauthorized`.
200
200
 
201
+ The HTTP authentication scheme is ASCII case-insensitive (`Bearer`, `bearer`,
202
+ or `BEARER`); the token remains case-sensitive and must match exactly after one
203
+ space. This applies to both Console HTTP and `woods-mcp-http`. Case-insensitive
204
+ scheme support is included in Woods `2.0.0`; use the canonical `Bearer`
205
+ spelling in client configuration for compatibility with earlier releases.
206
+
201
207
  For non-loopback access, `console_mcp_allowed_origins` must include the public
202
208
  Rails/MCP host. If a browser-based client sends an `Origin` header from a
203
209
  different host, include that exact origin too. This allow-list controls both
@@ -760,12 +766,12 @@ Each transaction sets a statement timeout before any query runs. The default is
760
766
 
761
767
  `SqlValidator` rejects non-read-only SQL at the string level, before any database interaction.
762
768
 
763
- Validation runs **once**, inside the executor, with the dialect of the live adapter. There is deliberately no earlier dialect-blind pre-check in the tool handler: a validator built without a dialect is the conservative MySQL+PostgreSQL union, and running it first meant a MySQL host rejected statements whose `\'`/backtick grammar produces a spuriously forbidden PostgreSQL view — the adapter-aware acceptance below could never be reached on a real transport. The executor raises `SqlValidationError` for anything it refuses, which the dispatch pipeline renders as a tool error, so nothing is ungated.
769
+ Validation runs **once**, inside the executor, with the dialect of the live adapter. There is deliberately no earlier dialect-blind pre-check in the tool handler: a validator built without a dialect is the conservative union of supported dialects, and running it first meant a MySQL host rejected statements whose `\'`/backtick grammar produces a spuriously forbidden PostgreSQL view — the adapter-aware acceptance below could never be reached on a real transport. The executor raises `SqlValidationError` for anything it refuses, which the dispatch pipeline renders as a tool error, so nothing is ungated.
764
770
 
765
771
 
766
772
  - **Allowed prefixes:** `SELECT`, `WITH...SELECT`, and plain `EXPLAIN`. `EXPLAIN ANALYZE` is rejected, it executes the query rather than just planning it (both the whitespace and `EXPLAIN (ANALYZE, …)` option-list spellings).
767
773
  - **Rejected prefixes:** `INSERT`, `UPDATE`, `DELETE`, `MERGE`, `DROP`, `ALTER`, `TRUNCATE`, `CREATE`, `GRANT`, `REVOKE`
768
- - **Rejected anywhere in query:** `UNION`, `INTO`, `COPY`; row-lock clauses (`FOR UPDATE`, `FOR NO KEY UPDATE`, `FOR SHARE`, `FOR KEY SHARE`, `FOR UPDATE NOWAIT`/`SKIP LOCKED`, MySQL `LOCK IN SHARE MODE`) — these take live row locks even inside the rolled-back transaction. The lock check is adapter-aware: `console_sql` validates with the active adapter's dialect, including MySQL double-quoted strings/backtick identifiers and PostgreSQL quoted identifiers/E-strings. Unknown adapters conservatively scan both normalizations. Every view is scanned under both MySQL executable-comment (`/*!...*/`) semantics, so `#` comments and version-guarded comments cannot split a clause apart.
774
+ - **Rejected anywhere in query:** `UNION`, `INTO`, `COPY`; row-lock clauses (`FOR UPDATE`, `FOR NO KEY UPDATE`, `FOR SHARE`, `FOR KEY SHARE`, `FOR UPDATE NOWAIT`/`SKIP LOCKED`, MySQL `LOCK IN SHARE MODE`) — these take live row locks even inside the rolled-back transaction. The lock check is adapter-aware: `console_sql` validates with the active adapter's dialect, including MySQL double-quoted strings/backtick identifiers and PostgreSQL quoted identifiers/E-strings. Unknown adapters conservatively scan all supported normalizations. Every view is scanned under both MySQL executable-comment (`/*!...*/`) semantics, so `#` comments and version-guarded comments cannot split a clause apart.
769
775
  - **Function allowlist (the authoritative function control):** every function-call-shaped identifier must appear in `ALLOWED_FUNCTIONS`, a conservative set of pure read-only functions (aggregates, window functions, string/number/date/JSON readers) kept portable across MySQL, PostgreSQL, and SQLite. Anything else is rejected by name, quoted forms (`"pg_terminate_backend"(…)`) included. This is an allowlist because a denylist cannot enumerate every side-effecting function (`nextval`, `pg_advisory_lock`, `pg_terminate_backend`, …). A legacy `DANGEROUS_FUNCTIONS` denylist (`pg_sleep`, `lo_import`, `lo_export`, `pg_read_file`, `pg_write_file`, `load_file`, `sleep`, `benchmark`) still runs first as belt-and-suspenders.
770
776
  - **Rejected patterns:** multiple statements (semicolons), writable CTEs (every `AS (...)` body is checked, so a writable CTE in any WITH position is refused — `WITH a AS (SELECT 1), b AS (DELETE FROM users RETURNING *) SELECT * FROM b`), a CTE list attached to top-level DML (`WITH a AS (SELECT 1) DELETE FROM users RETURNING *`), comment-hidden injections
771
777
 
@@ -793,13 +799,15 @@ to enforce their narrower parameterized scope grammar.
793
799
  - **HTTP:** Check that the Rails server is running and listening on the expected port. An unauthenticated `curl http://localhost:3000/mcp/console` should return `401` when the enabled middleware and bearer-auth guard are mounted. A request with the configured bearer token proceeds to MCP protocol handling.
794
800
  - **All modes:** Run `bundle exec rake woods:console` directly in a terminal. It should hang (waiting for MCP protocol input) rather than exit immediately. If it exits, check the error output.
795
801
 
796
- ### Rails boot noise breaks MCP protocol
802
+ <a id="rails-boot-noise-breaks-mcp-protocol"></a>
803
+
804
+ ### Rails logs break MCP protocol
805
+
806
+ Prefer `bundle exec rake woods:console`: it captures output before the Rails environment boots. Direct `rails runner` invocation can redirect output only after Rails has booted; it cannot recover an already contaminated protocol stream.
797
807
 
798
- The rake task redirects stdout to stderr before Rails boots specifically to prevent this. If you see JSON parse errors from the MCP client, check:
808
+ The runtime isolation fix is included in Woods `2.0.0`. Earlier builds restore stdout after boot, so a logger that writes there can interleave SQL or application logs with MCP responses. On those builds, configure the Console process's application logger to write to stderr or a file. Discarding stderr does not fix logs written to stdout.
799
809
 
800
- 1. You are using `bundle exec rake woods:console`, not `rails runner exe/woods-console` directly (the runner path handles this too, but via a different mechanism).
801
- 2. No `puts` or `print` calls run at boot in your initializers before the task can capture stdout.
802
- 3. Try running `bundle exec rake woods:console 2>/dev/null` to isolate, the MCP protocol output goes to stdout, Rails noise goes to stderr.
810
+ With a build containing the fix, application output remains on stderr during tool calls. If protocol contamination persists, inspect wrapper scripts and output emitted before the rake task starts; reserve stdout for MCP and retain stderr for diagnosis. Do not disable Console redaction or credential scanning to troubleshoot transport logging.
803
811
 
804
812
  ### Models not visible to `console_status`
805
813
 
@@ -853,7 +861,44 @@ quote and comment rules. MySQL also reads the executing session's `ANSI_QUOTES`
853
861
  and `NO_BACKSLASH_ESCAPES` settings for validation, protected-column scanning, and
854
862
  table gating; adjacent subtraction operators are not assumed to begin a comment.
855
863
  Direct scanner callers without session settings use conservative quote-mode scans.
864
+ SQLite read SQL accepts simple ASCII bare, double-quoted, or backtick identifiers
865
+ (letters, digits, and underscores, starting with a letter or underscore).
866
+ Use whitespace after `FROM` and `JOIN`, and use `SELECT` subqueries rather than
867
+ parenthesized table groups. Bracket-quoted and string-quoted names, quoted names
868
+ containing punctuation, and unsupported table-reference syntax are refused before
869
+ execution, because they cannot be reliably checked against the configured table
870
+ policy. Ordinary string literals remain supported. This restriction is part of
871
+ 2.0.0.beta4; keep read tools disabled on older versions when this policy is
872
+ needed. Confirm that the release is available before selecting it. The Rails-version
873
+ integration lane checks these boundaries on real SQLite.
874
+
856
875
  The contributor live-backend lane exercises these boundaries
857
876
  through Console requests against PostgreSQL and MySQL. Keep read tools disabled
858
877
  unless live SQL access is needed, and retain the configured blocked-table and
859
878
  redaction policies when diagnosing a rejected request.
879
+
880
+ ## Console policy corrections in 2.0.0.beta4
881
+
882
+ In `2.0.0.beta4`, the default model-reading tools check the resolved relation
883
+ against `console_blocked_tables` before fetching records or counts. This includes
884
+ application-defined default scopes and the parent lookup for association counts.
885
+ The checked relation is reused for execution so a dynamic default scope is not
886
+ resolved twice. These checks do not change which Console tools are enabled.
887
+
888
+ Response handling redacts protected fields before invoking serializers, converts
889
+ the remaining response to JSON-compatible values, then redacts and scans that
890
+ normalized tree before either JSON or Markdown rendering. Symbol values and custom
891
+ JSON serializers therefore receive the same credential checks as ordinary strings.
892
+ Custom values in Markdown now use their JSON-compatible representation. Numbers,
893
+ booleans, nulls, and ordinary record shapes retain their existing meanings.
894
+
895
+ For SQLite SQL, keyword spellings receive function-policy exceptions only where
896
+ supported query grammar requires them. PostgreSQL reserved-keyword grammar is
897
+ preserved. Unsupported parenthesized offset expressions on other dialects may
898
+ require a plain numeric offset. Keep the configured access and credential policies
899
+ in place when adjusting a query.
900
+
901
+ These corrections require `2.0.0.beta4` or a reviewed development revision that
902
+ contains them. Confirm that a patched release is available before selecting it.
903
+ On affected versions, disable Console where these policies are required; Index MCP
904
+ can stay enabled because it reads the published code index separately.
data/docs/DOCKER_SETUP.md CHANGED
@@ -87,9 +87,24 @@ docker compose exec app bundle exec rake woods:extract_framework
87
87
 
88
88
  Run the watcher as its own development service or process-manager entry, not as a one-off terminal command. Docker Desktop bind mounts may not deliver reliable native filesystem events; set `WOODS_WATCH_POLL=1` for polling when needed. The watcher updates structural generations automatically, while semantic vectors still require `woods:embed_incremental`.
89
89
 
90
+ Use the application's actual Rails task entrypoint. If its root `Rakefile` wraps
91
+ Compose, the container may need `bundle exec rails woods:watch`. Keep the raw
92
+ task under one external restart owner: `restart: unless-stopped` recovers after
93
+ Docker restarts while respecting an intentional stop; `on-failure` covers failed
94
+ process exits but not Docker restart. Leave idle TTL unset for continuous work.
95
+
96
+ Validate `docker compose config`: explicit YAML anchor keys can replace inherited
97
+ mounts/environment, and short `depends_on` does not establish database readiness.
98
+ Preserve the existing source/bundle mounts and use the application's healthcheck
99
+ convention. With Grove, include the watcher in the applicable shared or isolated
100
+ services list and align source/index mounts to the selected worktree. Follow
101
+ [Docker and Grove verification](AUTOMATIC_MAINTENANCE.md#docker-verify-the-resolved-service)
102
+ instead of copying a generic service that loses required settings.
103
+
90
104
  When host-side tasks or one-off containers read the daemon's shared index,
91
105
  `WOODS_WATCH_TRUST_FOREIGN_HOST=1` lets those readers trust its recent heartbeat.
92
106
  Set it in each reader process; Docker does not forward host variables by default.
107
+ Ordinary Index MCP reads do not need this trust or the extraction writer lock.
93
108
  See [cross-host liveness](WATCH_DAEMON.md#cross-host-liveness) for the 15-minute
94
109
  crash-detection bound and single-supervisor requirement.
95
110
 
@@ -468,6 +483,6 @@ To register `console_sql` and `console_query`, enable
468
483
 
469
484
  ### Woods MCP tools not available in a git worktree
470
485
 
471
- When working in a git worktree, subagents may not find the woods MCP servers because `.mcp.json` discovery is path-based and the worktree has a different root directory. See [MCP_WORKTREE_SETUP.md](MCP_WORKTREE_SETUP.md) for the fix and verification steps.
486
+ A separate Claude Code session launched in a worktree may have different project-scoped MCP registrations. Subagents inherit the parent session's MCP tools, subject to tool restrictions; changing their working directory does not select a different Woods index. Check registration, the Compose service's mounted checkout, and the served index using [MCP worktree setup](MCP_WORKTREE_SETUP.md).
472
487
 
473
488
  See [CONSOLE_MCP_SETUP.md](CONSOLE_MCP_SETUP.md) for detailed console server documentation.
data/docs/EVALUATION.md CHANGED
@@ -90,8 +90,8 @@ bundle exec ruby -Ilib bench/evaluation/runner.rb
90
90
  strategy selection also fail. The baseline format is developer-only and is
91
91
  **not** the `EVAL_BASELINE_FILE` aggregate-threshold format.
92
92
 
93
- B-190/B-191 recapture on Ruby 4.0.6 (five warmed pipeline repetitions per query; Ruby 3.3.1
94
- and 3.4.10 replay the same answers):
93
+ B-190/B-191 baseline, with the #549 output-label refresh captured on Ruby 4.0.6
94
+ (five warmed pipeline repetitions per query):
95
95
 
96
96
  | Strategy | Queries | Precision@5 | Recall | MRR | Mean actual context tokens |
97
97
  |---|---:|---:|---:|---:|---:|
@@ -99,8 +99,14 @@ and 3.4.10 replay the same answers):
99
99
  | Vector | 4 | 0.313 | 0.375 | 0.625 | 1,041.2 |
100
100
  | Graph | 8 | 0.813 | 0.519 | 1.000 | 1,041.1 |
101
101
  | Hybrid | 4 | 0.750 | 0.396 | 1.000 | 1,058.8 |
102
- | Direct with type filtering | 4 | 0.375 | 0.750 | 0.625 | 551.0 |
103
- | Within-type vector fallback | 4 | 0.400 | 1.000 | 1.000 | 1,250.8 |
102
+ | Direct with type filtering | 4 | 0.375 | 0.750 | 0.625 | 552.0 |
103
+ | Within-type vector fallback | 4 | 0.400 | 1.000 | 1.000 | 1,251.8 |
104
+
105
+ The clearer `Retrieval metadata records` heading adds one exact `cl100k_base`
106
+ token to each of the eight direct/fallback contexts. The other 20 contexts,
107
+ retrieved identifiers and order, annotations, and quality metrics are unchanged.
108
+ The refresh reuses captured vectors and the pinned `tiktoken 0.11.0` tokenizer;
109
+ it performs no new model inference.
104
110
 
105
111
  Precision@5 divides relevant hits by the actual returned slice size (up to five),
106
112
  not always by five. Recall divides retrieved relevant units by all annotated
@@ -2,7 +2,7 @@
2
2
 
3
3
  Woods ships **35 extractor classes** producing **39 distinct unit types**: one for each meaningful category of Rails code. This doc covers what each extractor captures, how to configure them, and the shape of the data they produce.
4
4
 
5
- > **Counts explained.** `lib/woods/extractors/` contains 42 files: 35 extractor classes (each ending in `_extractor.rb`) plus 7 supporting utilities (`shared_utility_methods`, `shared_dependency_scanner`, `callback_analyzer`, `behavioral_profile`, `route_helper_resolver`, `ast_source_extraction`, `source_nesting`). The 39 unit types comes from some extractors emitting multiple categories, `GraphQLExtractor` alone produces four (`graphql_type`, `graphql_mutation`, `graphql_resolver`, `graphql_query`), and `RailsSourceExtractor` produces both `rails_source` and `gem_source`. Supporting utilities enrich existing extractors (callback side-effects, behavioral config, AST-based source slicing, nested-namespace resolution) but are not themselves extractors and do not appear in the unit type enumeration. The authoritative mapping is `Woods::Extractor::TYPE_TO_EXTRACTOR_KEY` in `lib/woods/extractor.rb`.
5
+ > **Counts explained.** `lib/woods/extractors/` contains 35 extractor classes (each ending in `_extractor.rb`) plus supporting utilities such as `shared_utility_methods`, `shared_dependency_scanner`, `callback_analyzer`, `behavioral_profile`, `route_helper_resolver`, `ast_source_extraction`, `source_nesting`, and `declared_parent`. The 39 unit types comes from some extractors emitting multiple categories, `GraphQLExtractor` alone produces four (`graphql_type`, `graphql_mutation`, `graphql_resolver`, `graphql_query`), and `RailsSourceExtractor` produces both `rails_source` and `gem_source`. Supporting utilities enrich existing extractors (callback side-effects, behavioral config, AST-based source slicing, nested-namespace resolution) but are not themselves extractors and do not appear in the unit type enumeration. The authoritative mapping is `Woods::Extractor::TYPE_TO_EXTRACTOR_KEY` in `lib/woods/extractor.rb`.
6
6
 
7
7
  ---
8
8
 
@@ -32,6 +32,14 @@ Extractors discover code one of two ways:
32
32
 
33
33
  Some extractors combine both (e.g., `JobExtractor` scans directories first, then supplements with `ApplicationJob.descendants`).
34
34
 
35
+ Discovery is not exhaustive. A standalone module under `app/models` that is
36
+ called through singleton methods is not discovered by the model/PORO paths;
37
+ conventional concerns and app modules included by live models have separate
38
+ concern discovery. This gap is tracked in [#552](https://github.com/lost-in-the/woods/issues/552).
39
+ Separately, dependency scanning does not capture every method-body constant
40
+ reference ([#475](https://github.com/lost-in-the/woods/issues/475)). A missing unit
41
+ or edge is not proof of unused code; cross-check the application source.
42
+
35
43
  ### Identifier naming (source-derived units)
36
44
 
37
45
  File-based extractors derive an identifier in three steps, first match wins:
@@ -65,6 +73,7 @@ Every extractor returns `Array<ExtractedUnit>`. An `ExtractedUnit` is a self-con
65
73
  - Reads Rails' per-event callback chains (`_save_callbacks`, `_create_callbacks`, and the other lifecycle events), preserving each chain's order and each entry's `kind`, filter and conditions. The public `type` combines kind and event, such as `before_save` or `after_create`; Rails' separate `before_commit` event is reported as `before_commit`, not `before_before_commit`. `callback_count` equals the emitted callback list's length. The list includes framework-registered callbacks; it is runtime metadata, not an application-only filter or a cross-event execution trace. Regenerate the index after upgrading to pick up corrected callback metadata.
66
74
  - Proc/lambda filters, including Rails-generated association callbacks, use stable source-site labels in both metadata and callback chunks: `#<Proc app/models/post.rb:12>` (or `lambda`). App paths are relative to `Rails.root`; external paths are retained and native procs use `native`. Rails 6's numeric filter identity is resolved through `raw_filter`. These labels describe location and callable kind, not captured closure state; callbacks are never executed. Model condition labels retain their existing format.
67
75
  - Default callback-object representations omit process addresses: an instance becomes `#<CleanupCallback>`, an anonymous class becomes `#<Class>`, and its instance becomes `#<#<Class>>`. Anonymous namespace prefixes are normalized too (for example, `#<Module>::CleanupCallback`). Named classes and custom `to_s` labels retain their text. These labels do not distinguish arbitrary object state; separate registered callbacks remain separate entries even when their descriptive labels match. Controller object-filter formatting is unchanged.
76
+ - Direct Proc/lambda validation option values (for example inclusion/exclusion membership or message callables) use the same stable kind/source-site labels without execution. Validation order and duplicates, condition formats (`if`, `unless`, `on`), and non-Proc values are unchanged. Nested arrays/hashes are not recursively normalized, and labels do not serialize captured closure state. Run a full extraction after upgrading to refresh retained validation metadata; the stored schema is unchanged.
68
77
  - Callback side-effects are analyzed via `CallbackAnalyzer`: detects columns written (`self.col =`), jobs enqueued (`perform_later`), and services called
69
78
  - Reflects model class and instance methods after reading the schema, so Rails schema-loading optimizations produce the same method metadata in cold and warmed runs. Application-defined constructors remain visible; Rails versions that install an optimized singleton `new` during schema loading consistently include it in `class_methods`.
70
79
  - Automatically skips HABTM join models and anonymous classes
@@ -255,6 +264,10 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
255
264
  - Discovery and direct extraction accept only mailers backed by an existing app-owned source file, excluding dependency mailers and fabricated convention paths.
256
265
  - Each mailer action corresponds to an email template, template paths are recorded in metadata
257
266
  - Extracts `default from:`, `layout`, and per-action subject patterns
267
+ - Action names are sorted consistently in metadata, the generated header, template discovery and action chunks. Callback chain order and duplicate registrations are preserved.
268
+ - Direct Proc-valued defaults and Proc callback filters use source-location/kind labels without executing them; application paths are relative to `Rails.root`. Default containers, literal strings and non-Proc values retain their existing types. This does not serialize closure captures or recursively normalize arbitrary nested objects.
269
+ - Object callback filters use descriptive labels with addresses removed only from Ruby's default representation, including anonymous classes and namespaces. Custom labels and literal hexadecimal text are preserved. Labels do not serialize callback object state, and extraction never invokes callbacks.
270
+ - After upgrading, run a full extraction to refresh retained mailer units. Stabilized headers and callable labels can cause a one-time source-hash change; stored index schemas are unchanged.
258
271
 
259
272
  ---
260
273
 
@@ -382,6 +395,7 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
382
395
  - Scans `app/models` for files that don't define an `ActiveRecord::Base` descendant
383
396
  - Common examples: value objects, form objects placed in `app/models`, domain structs
384
397
  - Excludes concerns (those go to ConcernExtractor)
398
+ - `parent_class` and the generated Parent annotation describe the selected unit declaration only. Nested or sibling classes cannot supply its parent. An implicit `Object` parent, a dynamic superclass expression, or unparseable source produces `nil`; explicit constant-path parents retain their source names.
385
399
 
386
400
  ---
387
401
 
@@ -421,6 +435,7 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
421
435
  - Produces unit types: `graphql_type`, `graphql_mutation`, `graphql_resolver`, `graphql_query`
422
436
  - Extracts field metadata (types, descriptions, complexity, arguments), authorization patterns (Pundit, CanCan, `authorized?`), and dependencies on models/services
423
437
  - Since all GraphQL units come from one extractor, incremental re-extraction handles them via `extract_graphql_file`
438
+ - `parent_class` and summary chunks describe the selected declaration's explicit constant-path superclass, preserving its written qualification. Nested or sibling declarations and literal text cannot supply a parent. Implicit Object, module interfaces, dynamic superclass expressions, unavailable source, and invalid source have no declared parent (`null` metadata; `unknown` in summaries). This is source declaration metadata, not resolved runtime ancestry.
424
439
 
425
440
  **Example output (abbreviated):**
426
441
 
@@ -677,6 +692,7 @@ Every app-owned unit under a package root carries `metadata[:package]` with the
677
692
  **Key details:**
678
693
  - Excludes `lib/tasks/` (covered by RakeTaskExtractor) and `lib/generators/`
679
694
  - File-based scanning; no assumption about class hierarchy
695
+ - `parent_class` and the generated Parent annotation describe the selected unit declaration only. Nested or sibling classes cannot supply its parent. An implicit `Object` parent, a dynamic superclass expression, or unparseable source produces `nil`; explicit constant-path parents retain their source names.
680
696
 
681
697
  ---
682
698
 
@@ -738,11 +754,15 @@ History is limited to commits reachable from `HEAD` in the past 365 days,
738
754
  including merged branch history. Unmerged branches, remote refs, and tool
739
755
  checkpoint refs do not contribute. Commands run against the application root;
740
756
  when `WOODS_GIT_DIR` is set, `HEAD` belongs to that explicitly selected git
741
- directory, which may differ from a linked worktree's HEAD.
757
+ directory. For a linked worktree, select its `worktrees/<id>` directory within
758
+ the complete shared Git layout; selecting the shared root instead uses the
759
+ primary checkout's HEAD. See the [worktree mount guide](TROUBLESHOOTING.md#git-directory-mounts-for-linked-worktrees).
742
760
 
743
761
  After upgrading from a version that included all refs, run a full
744
762
  `woods:extract` to replace previously published git metadata. Incremental
745
- extraction refreshes only the units it rewrites.
763
+ extraction refreshes only the units it rewrites. A commit alone does not
764
+ necessarily trigger the source-file watcher; run full extraction when current
765
+ Git history is required.
746
766
 
747
767
  | Field | Description |
748
768
  |-------|-------------|
data/docs/FAQ.md CHANGED
@@ -134,7 +134,7 @@ When a model includes a concern, the behavior defined in that concern is part of
134
134
 
135
135
  ### How do I update the index after code changes?
136
136
 
137
- Use incremental mode, which re-extracts only files that have changed since the last run:
137
+ Use incremental mode to dispatch a selected batch of changed paths:
138
138
 
139
139
  ```bash
140
140
  bundle exec rake woods:incremental
@@ -143,7 +143,12 @@ bundle exec rake woods:incremental
143
143
  docker compose exec app bundle exec rake woods:incremental
144
144
  ```
145
145
 
146
- Incremental mode is ideal for CI pipelines and local development workflows. It is typically 5-10× faster than a full extraction. Nine unit types, `route`, `middleware`, `engine`, `scheduled_job`, `state_machine`, `factory`, `event`, `database_view`, and `rails_source`, don't map to individual files, so incremental mode re-runs their extractor **wholesale** whenever the relevant trigger path changes (e.g. `config/routes.rb` for routes, `Gemfile.lock` for middleware/engines/rails_source; `rails_source` participates only when `include_framework_sources` is enabled). You never need to run a full extraction just because one of these changed, see [TROUBLESHOOTING.md](TROUBLESHOOTING.md) for details.
146
+ The default Git range is `HEAD~1`; CI variables, an explicit range, or
147
+ `CHANGED_FILES` can select another batch. Incremental extraction also refreshes
148
+ affected concern consumers and re-runs whole-app extractors when their trigger
149
+ paths change. It can reduce work, but has no universal speedup: Rails boot, graph
150
+ work, and publication still take time. See the [incremental contract](INCREMENTAL_EXTRACTION.md)
151
+ for scope, whole-app triggers, and cases that require a full extraction.
147
152
 
148
153
  ---
149
154
 
@@ -411,7 +416,13 @@ When you run `rake woods:embed`, Woods generates embedding vectors for each extr
411
416
 
412
417
  ### What is the `codebase_retrieve` tool for?
413
418
 
414
- `codebase_retrieve` is the primary semantic search tool on the Index Server. It accepts a natural-language description of what you're looking for ("find where user email validation happens", "which services send Stripe API calls") and returns the most relevant extracted units as formatted context. It requires embedding configuration, without an embedding provider the tool responds with an error (`isError`, code `:not_configured`) and a remediation hint covering provider setup and the `search` tool for pattern-based matching in the meantime. Token budget is controlled by `config.max_context_tokens` (default: 8000).
419
+ `codebase_retrieve` ranks extracted units for a natural-language query, such as
420
+ "find where user email validation happens". Default semantic mode requires an
421
+ embedding provider and vector store. Explicit lexical mode ranks published
422
+ extraction text without either: set `WOODS_RETRIEVAL_MODE=lexical` in the MCP
423
+ process environment and restart it. There is no automatic fallback between
424
+ modes. See the [retrieval guide](RETRIEVAL_GUIDE.md#embedding-free-lexical-retrieval)
425
+ for setup, ranking evidence, and the estimated text-budget contract.
415
426
 
416
427
  ---
417
428
 
@@ -12,15 +12,15 @@ If an agent will perform the installation, use the safety and handoff checklist
12
12
 
13
13
  ## 1. Install the gem
14
14
 
15
- Use the [README release table](../README.md) to choose a version, then confirm
16
- that **exact version is published** on the [RubyGems versions page](https://rubygems.org/gems/woods/versions)
17
- before editing the Gemfile. A prepared release checkout can update the README
15
+ Choose a version from the [RubyGems versions page](https://rubygems.org/gems/woods/versions)
16
+ and confirm that **exact version is published** before editing the Gemfile.
17
+ A prepared release checkout can update the documentation
18
18
  before its gem is published; if the version is absent, choose an available
19
19
  version or wait for publication.
20
20
 
21
21
  If the published 2.x line has only beta or release-candidate versions, use an
22
- exact pin to the published prerelease, following the README's prerelease
23
- instructions; `~> 2.0` does not select prereleases. Follow the selected version's
22
+ exact pin to the published prerelease; `~> 2.0` does not select prereleases.
23
+ Follow the selected version's
24
24
  tag documentation. The `main` guides may describe features absent from the
25
25
  published gem.
26
26
 
@@ -148,21 +148,22 @@ Reconnect the MCP server and check `woods_status`. For OpenAI, pgvector, Qdrant,
148
148
 
149
149
  ### Keep the index current
150
150
 
151
- For automatic maintenance, keep a watcher running beside the Rails development process:
152
-
153
- ```bash
154
- bin/rails woods:watch
155
- ```
156
-
157
- ```text
158
- # Procfile.dev
159
- web: bin/rails server
160
- woods: bundle exec rake woods:watch
161
- ```
151
+ Enable one watcher through the application's normal development startup. The
152
+ [managed startup guide](WATCH_DAEMON.md#managed-development-startup) covers
153
+ opt-in Puma integration for a simple Rails application, an owned Foreman entry
154
+ for existing Procfile workflows, and external supervision for Docker/Grove.
155
+ The managed launcher and generator are **included in Woods `2.0.0`**;
156
+ check the installed commands before using them. Older packages can run the raw
157
+ `bin/rails woods:watch` task under an external restart-capable supervisor.
162
158
 
163
159
  On startup it reconciles changes made since the last successful generation. While running it batches file events, reloads Rails code when safe, extracts affected units, and publishes atomically. The Index Server detects the new generation on its next call and reloads automatically. After the initial extraction, ordinary code edits need no manual extraction or MCP restart.
164
160
 
165
- When dependencies, initializers, database configuration, credentials, or schema change, Rails cannot safely reload all captured state. The watcher records a degraded reason and exits with status 75 so the process manager can restart it. Docker bind mounts may require polling; follow [Watch daemon](WATCH_DAEMON.md).
161
+ When dependencies, initializers, database configuration, credentials, or schema
162
+ change, the raw task exits 75 to request a fresh Rails boot. The managed launcher
163
+ handles that restart internally; an external supervisor must handle it for the
164
+ raw task. Do not add the bare task to Foreman. Docker bind mounts may require
165
+ polling. See the [low-interaction workflow](AUTOMATIC_MAINTENANCE.md) for ownership,
166
+ hooks, and the checks that establish automatic maintenance is active.
166
167
 
167
168
  The watcher maintains the structural index. If semantic retrieval is enabled, also run `bin/rails woods:embed_incremental` to update vectors. Without a resident watcher, run `bin/rails woods:incremental` after changes. Use a full `woods:extract` after major upgrades or when validation reports drift. CI and shared-artifact patterns are covered in [Incremental extraction](INCREMENTAL_EXTRACTION.md).
168
169