woods 2.0.0.beta4 → 2.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (79) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +480 -477
  3. data/CONTRIBUTING.md +2 -2
  4. data/README.md +11 -26
  5. data/docs/AGENT_GUIDE.md +31 -12
  6. data/docs/AGENT_SETUP.md +17 -10
  7. data/docs/AUTOMATIC_MAINTENANCE.md +222 -0
  8. data/docs/BACKEND_MATRIX.md +13 -7
  9. data/docs/CLIENT_HOOKS.md +1 -1
  10. data/docs/CONFIGURATION_REFERENCE.md +36 -26
  11. data/docs/CONSOLE_MCP_SETUP.md +10 -8
  12. data/docs/DOCKER_SETUP.md +15 -0
  13. data/docs/EVALUATION.md +10 -4
  14. data/docs/EXTRACTOR_REFERENCE.md +14 -2
  15. data/docs/FAQ.md +14 -3
  16. data/docs/GETTING_STARTED.md +18 -17
  17. data/docs/INCREMENTAL_EXTRACTION.md +8 -3
  18. data/docs/INDEX_LAYOUT.md +2 -2
  19. data/docs/MCP_SERVERS.md +28 -11
  20. data/docs/MCP_TOOL_COOKBOOK.md +1 -1
  21. data/docs/MCP_WORKTREE_SETUP.md +13 -1
  22. data/docs/PUBLISHED_INDEX.md +1 -1
  23. data/docs/README.md +2 -1
  24. data/docs/RETRIEVAL_GUIDE.md +57 -8
  25. data/docs/SOURCE_FRESHNESS.md +1 -1
  26. data/docs/TOKEN_BENCHMARK.md +16 -10
  27. data/docs/TROUBLESHOOTING.md +133 -37
  28. data/docs/UPGRADING_TO_2.md +9 -7
  29. data/docs/WATCH_DAEMON.md +172 -17
  30. data/docs/WHY_WOODS.md +9 -5
  31. data/exe/woods-console +13 -11
  32. data/exe/woods-watch +5 -0
  33. data/lib/generators/woods/watch_generator.rb +53 -0
  34. data/lib/puma/plugin/woods.rb +10 -0
  35. data/lib/tasks/woods.rake +14 -0
  36. data/lib/woods/cache/cache_middleware.rb +6 -0
  37. data/lib/woods/console/stdio_transport.rb +27 -0
  38. data/lib/woods/extractor.rb +25 -7
  39. data/lib/woods/git_command.rb +6 -7
  40. data/lib/woods/git_provenance.rb +4 -6
  41. data/lib/woods/mcp/bootstrapper.rb +3 -1
  42. data/lib/woods/mcp/initialization_guidance.rb +1 -1
  43. data/lib/woods/mcp/server.rb +41 -9
  44. data/lib/woods/retrieval/corpus_status.rb +46 -0
  45. data/lib/woods/retriever.rb +19 -7
  46. data/lib/woods/storage/local_corpus_stats.rb +32 -0
  47. data/lib/woods/storage/metadata_store.rb +20 -0
  48. data/lib/woods/storage/vector_store.rb +10 -0
  49. data/lib/woods/version.rb +1 -1
  50. data/lib/woods/watch/child_environment.rb +30 -0
  51. data/lib/woods/watch/cli.rb +91 -0
  52. data/lib/woods/watch/daemon.rb +55 -7
  53. data/lib/woods/watch/event_stream.rb +70 -0
  54. data/lib/woods/watch/guardian.rb +142 -0
  55. data/lib/woods/watch/installation/layout.rb +70 -0
  56. data/lib/woods/watch/installation/options.rb +128 -0
  57. data/lib/woods/watch/installation/planner.rb +128 -0
  58. data/lib/woods/watch/installation/probe.rb +101 -0
  59. data/lib/woods/watch/installation/receipt.rb +77 -0
  60. data/lib/woods/watch/installation/recovery.rb +64 -0
  61. data/lib/woods/watch/installation/templates.rb +58 -0
  62. data/lib/woods/watch/installation.rb +56 -0
  63. data/lib/woods/watch/lifecycle.rb +182 -0
  64. data/lib/woods/watch/managed_child.rb +113 -0
  65. data/lib/woods/watch/managed_cleanup.rb +48 -0
  66. data/lib/woods/watch/managed_process.rb +144 -0
  67. data/lib/woods/watch/puma_adapter.rb +87 -0
  68. data/lib/woods/watch/puma_child.rb +66 -0
  69. data/lib/woods/watch/supervision_records.rb +95 -0
  70. data/lib/woods/watch/supervision_status.rb +104 -0
  71. data/lib/woods/watch/supervisor.rb +161 -0
  72. data/lib/woods/watch/supervisor_reporting.rb +46 -0
  73. data/plugin/.claude-plugin/plugin.json +1 -1
  74. data/plugin/skills/woods-agent-enable/SKILL.md +1 -1
  75. data/plugin/skills/woods-diagnose/SKILL.md +77 -8
  76. data/plugin/skills/woods-investigate/SKILL.md +6 -6
  77. data/plugin/skills/woods-mcp-config/SKILL.md +28 -1
  78. data/plugin/skills/woods-setup/SKILL.md +58 -4
  79. metadata +35 -5
@@ -210,7 +210,7 @@ config.vector_store_options = {
210
210
 
211
211
  Woods uses an HNSW index over pgvector's `vector` representation, which supports
212
212
  **1–2,000 dimensions** ([pgvector's HNSW limits](https://github.com/pgvector/pgvector#hnsw)).
213
- Unreleased after `2.0.0.beta3`: the adapter rejects wider dimensions before any
213
+ Included in Woods `2.0.0`: the adapter rejects wider dimensions before any
214
214
  SQL, and `woods:pgvector` rejects invalid widths before writing a migration.
215
215
  There is no automatic vector truncation or half-precision conversion.
216
216
 
@@ -282,7 +282,7 @@ in the promoted dump; the index MCP server loads that snapshot at startup or
282
282
  reload. Incremental embedding publishes changes to paths, dependencies, and
283
283
  other unit metadata even when unchanged source needs no new embedding. A run
284
284
  with no content or metadata changes keeps the existing dump and retention
285
- window. Unreleased after `2.0.0.beta3`: a full `Indexer#index_all` run replaces
285
+ window. Included in Woods `2.0.0`: a full `Indexer#index_all` run replaces
286
286
  the published corpus even when a custom caller reuses in-memory vector and
287
287
  metadata stores. Deleted units, including metadata-only records, are removed;
288
288
  an empty full rebuild publishes an empty dump. Failed embedding leaves the
@@ -451,7 +451,7 @@ A malformed record invalidates the entire snapshot, rather than exposing partial
451
451
  history. Legacy bare identifier keys, omitted/null unit collections, and optional per-unit hash
452
452
  fields remain supported; timestamp strings are not restricted to a new format.
453
453
  Unit history limits count matching unit records, not the most recent snapshots
454
- searched (JSON fallback correction unreleased after `2.0.0.beta3`). A unit
454
+ searched (JSON fallback correction included in Woods `2.0.0`). A unit
455
455
  missing from newer snapshots can still have retained history.
456
456
  Snapshot lists and unit history omit unusable files; direct lookup returns no
457
457
  snapshot, and a diff with an unavailable snapshot returns empty added, modified,
@@ -704,7 +704,7 @@ These variables are read by the gem and its MCP servers at runtime. They complem
704
704
  |----------|---------|---------|
705
705
  | `WOODS_RETRIEVAL_MODE` | `semantic` | Explicit packaged MCP retrieval mode: `semantic` or `lexical`. Lexical reads extraction unit JSON without provider autodetection, credentials or vector artifacts. |
706
706
  | `WOODS_DIR` | unset | MCP extraction-index path, after a positional argument and before `WOODS_OUTPUT`. See precedence below. |
707
- | `WOODS_OUTPUT` | unset | MCP index-path fallback when neither a positional path nor `WOODS_DIR` is set; unreleased after `2.0.0.beta3`. |
707
+ | `WOODS_OUTPUT` | unset | MCP index-path fallback when neither a positional path nor `WOODS_DIR` is set; included in Woods `2.0.0`. |
708
708
  | `WOODS_REQUIRE_INDEX` | unset | Set to `"1"` to fail closed: the server refuses to boot (raises `MissingArtifact`) unless a real index (`woods.json`) is present. By default an extract-only host boots in pattern/structural mode without it. Explicit lexical mode requires a valid published extraction index, not `woods.json`. |
709
709
  | `WOODS_ALLOW_AUTODETECT` | unset | **Deprecated no-op.** Auto-detect is now the default; accepted for backward compatibility only. |
710
710
  | `WOODS_SEARCH_MAX_SCAN` | `500` | Cap on unit files loaded during a phase-2 (metadata/source_code) `search`. Hitting the cap sets `partial: true` in the response. |
@@ -720,7 +720,7 @@ These variables are read by the gem and its MCP servers at runtime. They complem
720
720
  | `WOODS_QDRANT_URL`, `WOODS_QDRANT_COLLECTION`, `WOODS_QDRANT_API_KEY` | n/a | Override/require Qdrant connection settings when a pgvector/Qdrant-backed index is served outside its host application (no `Woods.configuration` available). |
721
721
  | `WOODS_PG_URL` | n/a | Required when a pgvector-backed index is served outside its host application. |
722
722
 
723
- **MCP index path precedence (unreleased after `2.0.0.beta3`):** positional
723
+ **MCP index path precedence (included in Woods `2.0.0`):** positional
724
724
  argument → `WOODS_DIR` → `WOODS_OUTPUT` → current directory for `woods-mcp`
725
725
  and `woods-mcp-http`. `woods-mcp-start` still requires one of the first three;
726
726
  it never silently selects the current directory. An explicitly empty
@@ -771,10 +771,24 @@ existing index before deciding another extraction is needed.
771
771
  | `WOODS_WATCH_CATCH_UP` | `1` (enabled) | Set to `"0"` to skip generation-watermark catch-up on daemon start. |
772
772
  | `WOODS_WATCH_TRUST_FOREIGN_HOST` | unset (disabled) | Set to `"1"` in each task/MCP reader to trust a foreign daemon's heartbeat for up to 15 minutes, without a local pid check. See [cross-host liveness](WATCH_DAEMON.md#cross-host-liveness) for clock bounds, degraded coverage, and startup limitations. |
773
773
 
774
+ ### Managed watcher startup
775
+
776
+ **Included in Woods `2.0.0`.** `woods-watch` wraps the raw task for native
777
+ Foreman/Puma lifecycle management. `--root PATH` selects the application working
778
+ directory; `--boot-timeout SECONDS` defaults to `300` and bounds boot/handshake,
779
+ not extraction. Explicit child argv follows `--`. Output-directory precedence
780
+ remains `WOODS_OUTPUT`, then `Woods.configuration.output_dir`.
781
+
782
+ Managed mode requires `WOODS_WATCH_IDLE_TIMEOUT` to be unset, preserves one owner,
783
+ and never takes over a conflicting daemon. The optional Puma adapter only starts
784
+ in its finalized development environment. See [startup and installation](WATCH_DAEMON.md#managed-development-startup)
785
+ for the generator's explicit modes, portable receipt, update/removal, and the
786
+ separate supervision status. Raw task settings above remain compatible.
787
+
774
788
  ### Opt-in plugin refresh hooks
775
789
 
776
790
  These settings control the plugin shell worker. Check installed
777
- `woods:hook_refresh` support first; the task is unreleased after 2.0.0.beta2.
791
+ `woods:hook_refresh` support first; the task is included in Woods `2.0.0`.
778
792
  See [hook coverage and retry](WATCH_DAEMON.md#hooks-for-agent-sessions) and
779
793
  [optional context limits](WATCH_DAEMON.md#optional-bounded-context-hints).
780
794
 
@@ -807,26 +821,23 @@ tasks for manual refreshes; hook transport is not a general shell execution API.
807
821
  | `GITHUB_BASE_REF` | unset (GitHub Actions) | Build the diff range `origin/<ref>...HEAD` for `woods:incremental`; an unfetched ref makes the range unresolvable, same exit behavior. |
808
822
  | `RAILS_ENV` | `development` | Rails environment the rake tasks boot in. |
809
823
  | `WOODS_PROFILE` | unset | Set to `"1"` to log disjoint `[Woods] [profile] <phase> in N.NNs` durations, including git enrichment, reconciliation, payload sync, pointer publication (`publish`) and retention (`payload prune`). Separate `[profile total]` lines report whole extraction wall time, including unprofiled setup and failed runs; never add these totals to phase durations. Excludes process/Rails boot before extraction. Off by default. |
810
- | `WOODS_GIT_DIR` | unset | Absolute path to the canonical git directory. Wins over the repository Woods would otherwise find, at all three of its git call sites: per-unit `commit_count`/`change_frequency` (enrichment), `manifest.json`'s `git_branch`/`git_sha` (provenance), and the `woods:incremental` diff range. All three build their command line with `Woods::GitCommand.argv`. |
824
+ | `WOODS_GIT_DIR` | unset | Absolute path passed directly to Git's `--git-dir`, selecting that directory's `HEAD`. For a linked worktree use its `worktrees/<id>` directory inside the complete shared Git layout. Wins over the repository Woods would otherwise find, at all three of its git call sites: per-unit `commit_count`/`change_frequency` (enrichment), `manifest.json`'s `git_branch`/`git_sha` (provenance), and the `woods:incremental` diff range. All three build their command line with `Woods::GitCommand.argv`. |
811
825
  | `GIT_BRANCH`, `GIT_SHA` | unset | Provenance for a checkout with no `.git` at all (a source tarball, a Docker `COPY` that excludes it). Ignored when a `.git` is present but unresolvable, so a stale build arg cannot mask a worktree. |
812
826
 
813
- **`GIT_DIR` alone is not enough for a linked git worktree.** Woods runs git as a
814
- subprocess, so git's own `GIT_DIR` and `GIT_COMMON_DIR` are honored wherever git
815
- honors them. But pointing `GIT_DIR` at a worktree's *private* git directory only
816
- moves the failure: that directory reaches the shared object store through a
817
- relative `commondir` pointer, which still resolves outside a container mount,
818
- and `GIT_COMMON_DIR` does not override it. `git rev-parse --git-dir` then
819
- succeeds while no ref resolves.
820
-
821
- Woods refuses to enrich in that state rather than writing `commit_count: 0` and
822
- `change_frequency: "new"` on every unit: the git keys are omitted, provenance is
823
- `"unknown"`, and one warning names the cause. Point `WOODS_GIT_DIR` at the
824
- canonical git directory (the one a worktree's `gitdir:` pointer ultimately leads
825
- to) and mount it:
826
-
827
- ```bash
828
- WOODS_GIT_DIR=/canonical-git bundle exec rake woods:extract
829
- ```
827
+ For linked worktrees in containers, preserve access to the complete shared
828
+ Git directory and the worktree-specific HEAD. A same-path mount that resolves
829
+ the existing `.git` pointer needs no override. For a relocated complete layout,
830
+ set `WOODS_GIT_DIR=/mounted-common/worktrees/<id>`, deriving `<id>` from Git's
831
+ worktree metadata rather than the branch name. Selecting `/mounted-common`
832
+ itself selects the primary checkout's HEAD and can produce incorrect history,
833
+ provenance, and incremental paths. See the canonical
834
+ [worktree mount and verification steps](TROUBLESHOOTING.md#git-directory-mounts-for-linked-worktrees).
835
+
836
+ Git subprocesses inherit Git's own environment variables, so inspect existing
837
+ `GIT_DIR` and `GIT_COMMON_DIR` settings when resolving layout problems.
838
+ If HEAD cannot be resolved, enrichment omits the Git keys and provenance is
839
+ `"unknown"` for a present but unresolvable `.git`; one warning names the cause.
840
+ After correcting the layout, run a full extraction to refresh retained metadata.
830
841
 
831
842
  ### Exporters
832
843
 
@@ -846,11 +857,10 @@ The `woods-mcp` bootstrapper emits a one-line STDERR banner at startup indicatin
846
857
 
847
858
  ## Git enrichment history
848
859
 
849
- Current source requires **Git 2.31 or newer** for optional per-unit git
860
+ Woods 2.0 requires **Git 2.31 or newer** for optional per-unit git
850
861
  metadata. Extraction still succeeds when git is unavailable or history cannot
851
862
  be read completely. Git enrichment is omitted in either case; a failed or
852
863
  incomplete streamed history read logs a warning.
853
- This requirement and the history policy below are unreleased after 2.0.0.beta2.
854
864
 
855
865
  Per-unit enrichment also requires a non-shallow repository. A shallow checkout
856
866
  or a failed repository-depth probe omits enrichment with one warning per
@@ -48,7 +48,7 @@ its token, allowed origins and TLS as described in [Option C](#option-c-http-rac
48
48
 
49
49
  The rake task does two things before starting the MCP server:
50
50
 
51
- 1. **Captures stdout before Rails boots.** Rails boot emits OpenTelemetry warnings, gem notices, and other output to stdout. An MCP client cannot parse these as JSON-RPC, they break the protocol. The rake task redirects stdout → stderr immediately, saves the real stdout fd, and restores it after boot completes.
51
+ 1. **Captures stdout before Rails boots.** Rails boot emits OpenTelemetry warnings, gem notices, and other output to stdout. An MCP client cannot parse these as JSON-RPC. The rake task saves the protocol output and redirects application stdout to stderr. In the runtime isolation fix (included in Woods `2.0.0`), that redirection remains active throughout the server's lifetime; only the MCP transport writes to the saved protocol pipe. Rails loggers, `puts`, and writes to standard output during queries stay on stderr.
52
52
  2. **Calls `Rails.application.eager_load!`** to load all application models. Without eager loading, only the models that happen to be autoloaded before the first query appear in the registry.
53
53
 
54
54
  ### MCP client configuration
@@ -83,7 +83,7 @@ rake woods:console
83
83
  │ ├─ Rails.application.eager_load!
84
84
  │ ├─ build model registry from ActiveRecord::Base.descendants
85
85
  │ ├─ Server.build_embedded(model_validator:, safe_context:, ...)
86
- │ └─ MCP::Server::Transports::StdioTransport.new(server).open
86
+ │ └─ Woods::Console::StdioTransport.new(server, output: protocol_out).open
87
87
  │
88
88
  └─ MCP server responds to tool calls via stdin/stdout
89
89
  ```
@@ -201,7 +201,7 @@ request. Missing or incorrect tokens receive `401 Unauthorized`.
201
201
  The HTTP authentication scheme is ASCII case-insensitive (`Bearer`, `bearer`,
202
202
  or `BEARER`); the token remains case-sensitive and must match exactly after one
203
203
  space. This applies to both Console HTTP and `woods-mcp-http`. Case-insensitive
204
- scheme support is unreleased after `2.0.0.beta3`; use the canonical `Bearer`
204
+ scheme support is included in Woods `2.0.0`; use the canonical `Bearer`
205
205
  spelling in client configuration for compatibility with earlier releases.
206
206
 
207
207
  For non-loopback access, `console_mcp_allowed_origins` must include the public
@@ -799,13 +799,15 @@ to enforce their narrower parameterized scope grammar.
799
799
  - **HTTP:** Check that the Rails server is running and listening on the expected port. An unauthenticated `curl http://localhost:3000/mcp/console` should return `401` when the enabled middleware and bearer-auth guard are mounted. A request with the configured bearer token proceeds to MCP protocol handling.
800
800
  - **All modes:** Run `bundle exec rake woods:console` directly in a terminal. It should hang (waiting for MCP protocol input) rather than exit immediately. If it exits, check the error output.
801
801
 
802
- ### Rails boot noise breaks MCP protocol
802
+ <a id="rails-boot-noise-breaks-mcp-protocol"></a>
803
803
 
804
- The rake task redirects stdout to stderr before Rails boots specifically to prevent this. If you see JSON parse errors from the MCP client, check:
804
+ ### Rails logs break MCP protocol
805
805
 
806
- 1. You are using `bundle exec rake woods:console`, not `rails runner exe/woods-console` directly (the runner path handles this too, but via a different mechanism).
807
- 2. No `puts` or `print` calls run at boot in your initializers before the task can capture stdout.
808
- 3. Try running `bundle exec rake woods:console 2>/dev/null` to isolate, the MCP protocol output goes to stdout, Rails noise goes to stderr.
806
+ Prefer `bundle exec rake woods:console`: it captures output before the Rails environment boots. Direct `rails runner` invocation can redirect output only after Rails has booted; it cannot recover an already contaminated protocol stream.
807
+
808
+ The runtime isolation fix is included in Woods `2.0.0`. Earlier builds restore stdout after boot, so a logger that writes there can interleave SQL or application logs with MCP responses. On those builds, configure the Console process's application logger to write to stderr or a file. Discarding stderr does not fix logs written to stdout.
809
+
810
+ With a build containing the fix, application output remains on stderr during tool calls. If protocol contamination persists, inspect wrapper scripts and output emitted before the rake task starts; reserve stdout for MCP and retain stderr for diagnosis. Do not disable Console redaction or credential scanning to troubleshoot transport logging.
809
811
 
810
812
  ### Models not visible to `console_status`
811
813
 
data/docs/DOCKER_SETUP.md CHANGED
@@ -87,9 +87,24 @@ docker compose exec app bundle exec rake woods:extract_framework
87
87
 
88
88
  Run the watcher as its own development service or process-manager entry, not as a one-off terminal command. Docker Desktop bind mounts may not deliver reliable native filesystem events; set `WOODS_WATCH_POLL=1` for polling when needed. The watcher updates structural generations automatically, while semantic vectors still require `woods:embed_incremental`.
89
89
 
90
+ Use the application's actual Rails task entrypoint. If its root `Rakefile` wraps
91
+ Compose, the container may need `bundle exec rails woods:watch`. Keep the raw
92
+ task under one external restart owner: `restart: unless-stopped` recovers after
93
+ Docker restarts while respecting an intentional stop; `on-failure` covers failed
94
+ process exits but not Docker restart. Leave idle TTL unset for continuous work.
95
+
96
+ Validate `docker compose config`: explicit YAML anchor keys can replace inherited
97
+ mounts/environment, and short `depends_on` does not establish database readiness.
98
+ Preserve the existing source/bundle mounts and use the application's healthcheck
99
+ convention. With Grove, include the watcher in the applicable shared or isolated
100
+ services list and align source/index mounts to the selected worktree. Follow
101
+ [Docker and Grove verification](AUTOMATIC_MAINTENANCE.md#docker-verify-the-resolved-service)
102
+ instead of copying a generic service that loses required settings.
103
+
90
104
  When host-side tasks or one-off containers read the daemon's shared index,
91
105
  `WOODS_WATCH_TRUST_FOREIGN_HOST=1` lets those readers trust its recent heartbeat.
92
106
  Set it in each reader process; Docker does not forward host variables by default.
107
+ Ordinary Index MCP reads do not need this trust or the extraction writer lock.
93
108
  See [cross-host liveness](WATCH_DAEMON.md#cross-host-liveness) for the 15-minute
94
109
  crash-detection bound and single-supervisor requirement.
95
110
 
data/docs/EVALUATION.md CHANGED
@@ -90,8 +90,8 @@ bundle exec ruby -Ilib bench/evaluation/runner.rb
90
90
  strategy selection also fail. The baseline format is developer-only and is
91
91
  **not** the `EVAL_BASELINE_FILE` aggregate-threshold format.
92
92
 
93
- B-190/B-191 recapture on Ruby 4.0.6 (five warmed pipeline repetitions per query; Ruby 3.3.1
94
- and 3.4.10 replay the same answers):
93
+ B-190/B-191 baseline, with the #549 output-label refresh captured on Ruby 4.0.6
94
+ (five warmed pipeline repetitions per query):
95
95
 
96
96
  | Strategy | Queries | Precision@5 | Recall | MRR | Mean actual context tokens |
97
97
  |---|---:|---:|---:|---:|---:|
@@ -99,8 +99,14 @@ and 3.4.10 replay the same answers):
99
99
  | Vector | 4 | 0.313 | 0.375 | 0.625 | 1,041.2 |
100
100
  | Graph | 8 | 0.813 | 0.519 | 1.000 | 1,041.1 |
101
101
  | Hybrid | 4 | 0.750 | 0.396 | 1.000 | 1,058.8 |
102
- | Direct with type filtering | 4 | 0.375 | 0.750 | 0.625 | 551.0 |
103
- | Within-type vector fallback | 4 | 0.400 | 1.000 | 1.000 | 1,250.8 |
102
+ | Direct with type filtering | 4 | 0.375 | 0.750 | 0.625 | 552.0 |
103
+ | Within-type vector fallback | 4 | 0.400 | 1.000 | 1.000 | 1,251.8 |
104
+
105
+ The clearer `Retrieval metadata records` heading adds one exact `cl100k_base`
106
+ token to each of the eight direct/fallback contexts. The other 20 contexts,
107
+ retrieved identifiers and order, annotations, and quality metrics are unchanged.
108
+ The refresh reuses captured vectors and the pinned `tiktoken 0.11.0` tokenizer;
109
+ it performs no new model inference.
104
110
 
105
111
  Precision@5 divides relevant hits by the actual returned slice size (up to five),
106
112
  not always by five. Recall divides retrieved relevant units by all annotated
@@ -32,6 +32,14 @@ Extractors discover code one of two ways:
32
32
 
33
33
  Some extractors combine both (e.g., `JobExtractor` scans directories first, then supplements with `ApplicationJob.descendants`).
34
34
 
35
+ Discovery is not exhaustive. A standalone module under `app/models` that is
36
+ called through singleton methods is not discovered by the model/PORO paths;
37
+ conventional concerns and app modules included by live models have separate
38
+ concern discovery. This gap is tracked in [#552](https://github.com/lost-in-the/woods/issues/552).
39
+ Separately, dependency scanning does not capture every method-body constant
40
+ reference ([#475](https://github.com/lost-in-the/woods/issues/475)). A missing unit
41
+ or edge is not proof of unused code; cross-check the application source.
42
+
35
43
  ### Identifier naming (source-derived units)
36
44
 
37
45
  File-based extractors derive an identifier in three steps, first match wins:
@@ -746,11 +754,15 @@ History is limited to commits reachable from `HEAD` in the past 365 days,
746
754
  including merged branch history. Unmerged branches, remote refs, and tool
747
755
  checkpoint refs do not contribute. Commands run against the application root;
748
756
  when `WOODS_GIT_DIR` is set, `HEAD` belongs to that explicitly selected git
749
- directory, which may differ from a linked worktree's HEAD.
757
+ directory. For a linked worktree, select its `worktrees/<id>` directory within
758
+ the complete shared Git layout; selecting the shared root instead uses the
759
+ primary checkout's HEAD. See the [worktree mount guide](TROUBLESHOOTING.md#git-directory-mounts-for-linked-worktrees).
750
760
 
751
761
  After upgrading from a version that included all refs, run a full
752
762
  `woods:extract` to replace previously published git metadata. Incremental
753
- extraction refreshes only the units it rewrites.
763
+ extraction refreshes only the units it rewrites. A commit alone does not
764
+ necessarily trigger the source-file watcher; run full extraction when current
765
+ Git history is required.
754
766
 
755
767
  | Field | Description |
756
768
  |-------|-------------|
data/docs/FAQ.md CHANGED
@@ -134,7 +134,7 @@ When a model includes a concern, the behavior defined in that concern is part of
134
134
 
135
135
  ### How do I update the index after code changes?
136
136
 
137
- Use incremental mode, which re-extracts only files that have changed since the last run:
137
+ Use incremental mode to dispatch a selected batch of changed paths:
138
138
 
139
139
  ```bash
140
140
  bundle exec rake woods:incremental
@@ -143,7 +143,12 @@ bundle exec rake woods:incremental
143
143
  docker compose exec app bundle exec rake woods:incremental
144
144
  ```
145
145
 
146
- Incremental mode is ideal for CI pipelines and local development workflows. It is typically 5-10× faster than a full extraction. Nine unit types, `route`, `middleware`, `engine`, `scheduled_job`, `state_machine`, `factory`, `event`, `database_view`, and `rails_source`, don't map to individual files, so incremental mode re-runs their extractor **wholesale** whenever the relevant trigger path changes (e.g. `config/routes.rb` for routes, `Gemfile.lock` for middleware/engines/rails_source; `rails_source` participates only when `include_framework_sources` is enabled). You never need to run a full extraction just because one of these changed, see [TROUBLESHOOTING.md](TROUBLESHOOTING.md) for details.
146
+ The default Git range is `HEAD~1`; CI variables, an explicit range, or
147
+ `CHANGED_FILES` can select another batch. Incremental extraction also refreshes
148
+ affected concern consumers and re-runs whole-app extractors when their trigger
149
+ paths change. It can reduce work, but has no universal speedup: Rails boot, graph
150
+ work, and publication still take time. See the [incremental contract](INCREMENTAL_EXTRACTION.md)
151
+ for scope, whole-app triggers, and cases that require a full extraction.
147
152
 
148
153
  ---
149
154
 
@@ -411,7 +416,13 @@ When you run `rake woods:embed`, Woods generates embedding vectors for each extr
411
416
 
412
417
  ### What is the `codebase_retrieve` tool for?
413
418
 
414
- `codebase_retrieve` is the primary semantic search tool on the Index Server. It accepts a natural-language description of what you're looking for ("find where user email validation happens", "which services send Stripe API calls") and returns the most relevant extracted units as formatted context. It requires embedding configuration, without an embedding provider the tool responds with an error (`isError`, code `:not_configured`) and a remediation hint covering provider setup and the `search` tool for pattern-based matching in the meantime. Token budget is controlled by `config.max_context_tokens` (default: 8000).
419
+ `codebase_retrieve` ranks extracted units for a natural-language query, such as
420
+ "find where user email validation happens". Default semantic mode requires an
421
+ embedding provider and vector store. Explicit lexical mode ranks published
422
+ extraction text without either: set `WOODS_RETRIEVAL_MODE=lexical` in the MCP
423
+ process environment and restart it. There is no automatic fallback between
424
+ modes. See the [retrieval guide](RETRIEVAL_GUIDE.md#embedding-free-lexical-retrieval)
425
+ for setup, ranking evidence, and the estimated text-budget contract.
415
426
 
416
427
  ---
417
428
 
@@ -12,15 +12,15 @@ If an agent will perform the installation, use the safety and handoff checklist
12
12
 
13
13
  ## 1. Install the gem
14
14
 
15
- Use the [README release table](../README.md) to choose a version, then confirm
16
- that **exact version is published** on the [RubyGems versions page](https://rubygems.org/gems/woods/versions)
17
- before editing the Gemfile. A prepared release checkout can update the README
15
+ Choose a version from the [RubyGems versions page](https://rubygems.org/gems/woods/versions)
16
+ and confirm that **exact version is published** before editing the Gemfile.
17
+ A prepared release checkout can update the documentation
18
18
  before its gem is published; if the version is absent, choose an available
19
19
  version or wait for publication.
20
20
 
21
21
  If the published 2.x line has only beta or release-candidate versions, use an
22
- exact pin to the published prerelease, following the README's prerelease
23
- instructions; `~> 2.0` does not select prereleases. Follow the selected version's
22
+ exact pin to the published prerelease; `~> 2.0` does not select prereleases.
23
+ Follow the selected version's
24
24
  tag documentation. The `main` guides may describe features absent from the
25
25
  published gem.
26
26
 
@@ -148,21 +148,22 @@ Reconnect the MCP server and check `woods_status`. For OpenAI, pgvector, Qdrant,
148
148
 
149
149
  ### Keep the index current
150
150
 
151
- For automatic maintenance, keep a watcher running beside the Rails development process:
152
-
153
- ```bash
154
- bin/rails woods:watch
155
- ```
156
-
157
- ```text
158
- # Procfile.dev
159
- web: bin/rails server
160
- woods: bundle exec rake woods:watch
161
- ```
151
+ Enable one watcher through the application's normal development startup. The
152
+ [managed startup guide](WATCH_DAEMON.md#managed-development-startup) covers
153
+ opt-in Puma integration for a simple Rails application, an owned Foreman entry
154
+ for existing Procfile workflows, and external supervision for Docker/Grove.
155
+ The managed launcher and generator are **included in Woods `2.0.0`**;
156
+ check the installed commands before using them. Older packages can run the raw
157
+ `bin/rails woods:watch` task under an external restart-capable supervisor.
162
158
 
163
159
  On startup it reconciles changes made since the last successful generation. While running it batches file events, reloads Rails code when safe, extracts affected units, and publishes atomically. The Index Server detects the new generation on its next call and reloads automatically. After the initial extraction, ordinary code edits need no manual extraction or MCP restart.
164
160
 
165
- When dependencies, initializers, database configuration, credentials, or schema change, Rails cannot safely reload all captured state. The watcher records a degraded reason and exits with status 75 so the process manager can restart it. Docker bind mounts may require polling; follow [Watch daemon](WATCH_DAEMON.md).
161
+ When dependencies, initializers, database configuration, credentials, or schema
162
+ change, the raw task exits 75 to request a fresh Rails boot. The managed launcher
163
+ handles that restart internally; an external supervisor must handle it for the
164
+ raw task. Do not add the bare task to Foreman. Docker bind mounts may require
165
+ polling. See the [low-interaction workflow](AUTOMATIC_MAINTENANCE.md) for ownership,
166
+ hooks, and the checks that establish automatic maintenance is active.
166
167
 
167
168
  The watcher maintains the structural index. If semantic retrieval is enabled, also run `bin/rails woods:embed_incremental` to update vectors. Without a resident watcher, run `bin/rails woods:incremental` after changes. Use a full `woods:extract` after major upgrades or when validation reports drift. CI and shared-artifact patterns are covered in [Incremental extraction](INCREMENTAL_EXTRACTION.md).
168
169
 
@@ -89,8 +89,13 @@ Recovery choices, in the order they are worth trying:
89
89
  3. Run a full `woods:extract` when the range cannot be repaired this run.
90
90
 
91
91
  The diff itself is rooted at the extracted application (`git -C Rails.root`),
92
- so it cannot read whatever checkout the process happened to start in — the
93
- same rooting rule the manifest's git provenance follows.
92
+ independently of the process working directory. An explicit `WOODS_GIT_DIR`
93
+ selects that Git directory's HEAD for both the diff and manifest provenance.
94
+ For a linked worktree, use its worktree-specific directory within the complete
95
+ shared layout; selecting the shared root instead reads the primary checkout's
96
+ HEAD. See [worktree mount verification](TROUBLESHOOTING.md#git-directory-mounts-for-linked-worktrees).
97
+ Git-only changes may not trigger the source-file watcher; see the
98
+ [watcher limitation](WATCH_DAEMON.md#watcher-backends).
94
99
 
95
100
  Named, source-defined app modules included by runtime models are tracked as concern units even
96
101
  outside `concerns/` directories. Changing their source refreshes their includers,
@@ -100,7 +105,7 @@ to populate these previously missing source mappings.
100
105
 
101
106
  ## Handled source errors and retry
102
107
 
103
- Unreleased after `2.0.0.beta3`: when an incremental extraction or named refresh
108
+ Included in Woods `2.0.0`: when an incremental extraction or named refresh
104
109
  records a handled consumer error (for example malformed locale or schedule
105
110
  YAML), it raises `Woods::ExtractionError` before publishing. Empty output from
106
111
  that failed consumer does not authorize replacing or deleting its last-good
data/docs/INDEX_LAYOUT.md CHANGED
@@ -94,8 +94,8 @@ families and `manifest.provenance.mode`; it is not Rails runtime evidence.
94
94
  ### Semantic graph validation
95
95
 
96
96
  `woods:validate` checks raw graph data against the unit indexes and artifacts in
97
- one pinned published generation. This validation is unreleased after
98
- `2.0.0.beta2`. It checks the shapes of `nodes`, `edges`, `reverse`, `file_map`,
97
+ one pinned published generation. This validation is included in Woods `2.0.0`.
98
+ It checks the shapes of `nodes`, `edges`, `reverse`, `file_map`,
99
99
  `type_index`, and optional `variants`; typed identities must be unique and agree
100
100
  with the actual indexed units. Forward sources must exist. Reverse, file, and
101
101
  type memberships must match the union of primary and variant contributions.
data/docs/MCP_SERVERS.md CHANGED
@@ -56,6 +56,12 @@ Prefer the application's bundle and a project-scoped configuration:
56
56
 
57
57
  `woods-mcp-start` checks that the directory and published manifest exist, then replaces itself with `woods-mcp`. It does not install dependencies or restart a crashed process.
58
58
 
59
+ The MCP client owns this reader process. Automatic index maintenance has its own
60
+ [watcher startup](WATCH_DAEMON.md#managed-development-startup): configuring MCP
61
+ does not enable it. A running reader sees new published generations on subsequent
62
+ calls without reconnecting. Verify a real edit, not only a successful connection;
63
+ see [automatic maintenance](AUTOMATIC_MAINTENANCE.md).
64
+
59
65
  `woods_status.index.woods_version` identifies the last publisher of the served
60
66
  manifest; `server.version` identifies the running MCP reader. Missing writer
61
67
  provenance is `null`. See [manifest writer provenance](PUBLISHED_INDEX.md#manifest-writer-provenance).
@@ -133,12 +139,23 @@ before using `codebase_retrieve`, including after reload. The guidance grants
133
139
  no extraction, configuration-change, or Console authorization. Detailed usage
134
140
  belongs in the [agent guide](AGENT_GUIDE.md).
135
141
 
136
- This addition is unreleased after `2.0.0.beta2`. The SDK omits `instructions`
142
+ This addition is included in Woods `2.0.0`. The SDK omits `instructions`
137
143
  when negotiating protocol `2024-11-05`; that behavior is preserved. Older gems,
138
144
  legacy clients, and clients that do not show server instructions can use the
139
145
  agent guide or investigation skill. Leave protocol negotiation enabled rather
140
146
  than pinning a newer version solely to obtain guidance.
141
147
 
148
+ ### Structural and semantic readiness
149
+
150
+ `ready` describes the published structural index. A reachable embedding provider
151
+ and bootstrap state `hydrated` do not establish that semantic stores contain
152
+ data. Supporting readers also report `retriever.corpus`: locally known vector
153
+ and metadata entry counts, counts by type, and whether those stores are empty,
154
+ populated, or unknown. These are stored records (including chunks), not a count
155
+ of extracted units with verified embedding coverage. Positive counts do not
156
+ prove complete coverage or working provider access. See
157
+ [semantic corpus diagnostics](RETRIEVAL_GUIDE.md#semantic-corpus-diagnostics).
158
+
142
159
  ### Tools (29 — 14 registered in the packaged default)
143
160
 
144
161
  The Index Server defines 29 schemas across core and conditional capabilities. The normal packaged executable registers the 14 tools below; the remaining schemas require the specialized wiring described afterward.
@@ -194,7 +211,7 @@ process after publishing a new embedded index.
194
211
 
195
212
  ### Graph-analysis pages
196
213
 
197
- Unreleased after `2.0.0.beta3`: `graph_analysis` enforces its advertised default
214
+ Included in Woods `2.0.0`: `graph_analysis` enforces its advertised default
198
215
  of 20 rows per section. Pass `limit` and `offset` to page one selected `analysis`
199
216
  or each section of `analysis: "all"`. Explicit limits also bound nested hub
200
217
  `dependents` lists. Older servers may return every section row when `limit` is
@@ -214,7 +231,7 @@ they do not establish complete source-reference coverage.
214
231
  Search responses retain `query`, `result_count`, and `results`; `result_count`
215
232
  is the number returned, not an estimated total. The additive `completeness`
216
233
  object describes the requested types, literal filters, and fields in the pinned
217
- generation. This contract is unreleased after `2.0.0.beta2`.
234
+ generation. This contract is included in Woods `2.0.0`.
218
235
 
219
236
  | `reason` | `status` | `has_more` | `total_matches` |
220
237
  |---|---|---|---|
@@ -260,16 +277,16 @@ Supporting servers expose the annotated, paginated traversal result in
260
277
  stdio and HTTP servers. Read `data.total_is_exact`, `data.graph_coverage`, budget
261
278
  counters and optional explanation witnesses there; `content[0].text` and
262
279
  `structuredContent.text` keep the same human-readable rendering. No `format`
263
- tool argument is needed or accepted. This additive data payload is unreleased
264
- after `2.0.0.beta3`; verify the installed response before relying on it. Older
280
+ tool argument is needed or accepted. This additive data payload is included in
281
+ Woods `2.0.0`; verify the installed response before relying on it. Older
265
282
  human-renderer responses can carry only text. The structured nodes and witnesses
266
283
  cover the same page, not an additional traversal or an unpaginated graph.
267
284
 
268
285
  Successful responses carry `graph_coverage` with `scope: "published_relationships"`,
269
286
  `source_references: "not_exhaustive"`, and a human-readable `notice`. Text formats
270
287
  show the same notice, including compact, root-only and empty-page responses.
271
- This response metadata and the total exactness field below are unreleased after
272
- Woods `2.0.0.beta3`; older servers need the same conservative interpretation.
288
+ This response metadata and the total exactness field below are included in Woods
289
+ `2.0.0`; older servers need the same conservative interpretation.
273
290
 
274
291
  ### Dependency traversal budgets
275
292
 
@@ -310,14 +327,14 @@ for stable pages. No wall-clock deadline is used, so cutoffs are deterministic.
310
327
  Budgets cover traversal work after per-generation graph loading and cache
311
328
  preparation (JSON parsing, typed-edge normalization, node types and database
312
329
  metadata). They do not cap that initial load, elapsed time, or total process
313
- memory. These arguments are unreleased in Woods 2.0.0.beta2; check the connected
330
+ memory. These arguments are included in Woods `2.0.0`; check the connected
314
331
  server's tool schema before sending them to an older installation.
315
332
 
316
333
  ### Traversal explanations
317
334
 
318
335
  Supporting development versions accept `explain: true` on `dependencies` and
319
- `dependents`. Check the connected schema first; this option is unreleased after
320
- 2.0.0.beta2. Omitted or false keeps the existing compact response.
336
+ `dependents`. Check the connected schema first; this option is included in Woods
337
+ `2.0.0`. Omitted or false keeps the existing compact response.
321
338
 
322
339
  The additive `explanation` object contains:
323
340
 
@@ -343,7 +360,7 @@ it describes identifier-level reachability, never a uniquely typed path.
343
360
  A true value means only that identities along this witness have unambiguous
344
361
  types. It does not establish source-reference coverage or observed execution.
345
362
  Text labels this `witness types unambiguous=yes/no`; the JSON key and its meaning
346
- remain unchanged. The text label change is unreleased after `2.0.0.beta3`.
363
+ remain unchanged. The text label change is included in Woods `2.0.0`.
347
364
  `types` filters retain the compact traversal's identifier-level semantics: any
348
365
  registered type can qualify a name, while edge evidence keeps its actual source
349
366
  owner. Multiple relationship kinds between the same endpoints remain separate.
@@ -356,7 +356,7 @@ column.
356
356
  Use `lookup` on the returned identifier for source, actions, and routes. Search
357
357
  `source_code` for textual matches beyond names. Supporting versions distinguish
358
358
  exact totals from a bounded result prefix; `partial` means this is discovery,
359
- not an exhaustive list. Completeness metadata is unreleased after `2.0.0.beta2`;
359
+ not an exhaustive list. Completeness metadata is included in Woods `2.0.0`;
360
360
  see the [search contract](MCP_SERVERS.md#search-completeness).
361
361
 
362
362
  ---
@@ -66,7 +66,19 @@ The published payload's `manifest.json` records the extraction's `git_branch` an
66
66
 
67
67
  Woods uses worktree-aware Git commands. If a present `.git` cannot be resolved, provenance is `"unknown"`; stale `GIT_BRANCH`/`GIT_SHA` values are not substituted. Those environment variables are fallbacks only when the root has no `.git` or Git is unavailable. Temporal snapshots skip an unknown SHA.
68
68
 
69
- For extraction in a container, make the canonical Git directory and the worktree's pointer resolvable there. Mounting only the private worktree Git directory can leave its shared object store unreachable. Follow the [Git provenance troubleshooting guide](TROUBLESHOOTING.md) for mount and `WOODS_GIT_DIR` guidance, and the [published index layout](INDEX_LAYOUT.md) when locating the manifest.
69
+ For extraction in a container, mount the complete shared Git layout at the
70
+ original path so the worktree's `.git` pointer resolves without an override,
71
+ or select `/mounted-common/worktrees/<id>` with `WOODS_GIT_DIR` inside a
72
+ relocated complete mount. Derive `<id>` from Git metadata, not the branch name.
73
+ Selecting the shared root instead uses the primary checkout's HEAD, affecting
74
+ provenance, per-file history, and incremental paths. Follow the
75
+ [worktree mount and verification steps](TROUBLESHOOTING.md#git-directory-mounts-for-linked-worktrees)
76
+ and the [published index layout](INDEX_LAYOUT.md) when locating the manifest.
77
+
78
+ Compare the branch and exact SHA in the extraction environment with the intended
79
+ worktree, then verify the published manifest after full extraction. A commit
80
+ alone may not trigger the source-file watcher; use full `woods:extract` when
81
+ current Git history is required.
70
82
 
71
83
  ## Troubleshooting
72
84
 
@@ -111,7 +111,7 @@ The reader wraps `Woods::MCP::IndexReader` with `auto_refresh: false`; the unit
111
111
 
112
112
  ### Actual unit types and directory families
113
113
 
114
- Unreleased after `2.0.0.beta3`: `unit` and `units` accept actual published
114
+ Included in Woods `2.0.0`: `unit` and `units` accept actual published
115
115
  `graphql_type`, `graphql_mutation`, `graphql_resolver`, `graphql_query`, and
116
116
  `gem_source` types. Enumeration preserves each unit's actual type rather than
117
117
  labeling every GraphQL member `graphql` or every gem source `rails_source`.
data/docs/README.md CHANGED
@@ -10,7 +10,7 @@ Woods extracts runtime-accurate Rails context and serves it to coding agents thr
10
10
  | Ask an agent to install or configure Woods | [Agent setup runbook](AGENT_SETUP.md) | A safe, reviewable install with an agent handoff report |
11
11
  | Configure an MCP client or Docker path | [MCP servers](MCP_SERVERS.md) | A working Index Server and, if authorized, an optional Console Server |
12
12
  | Use Woods tools as an agent | [Agent guide](AGENT_GUIDE.md) | A repeatable query workflow for code context, flows, and blast radius |
13
- | Keep the index current automatically | [Watch daemon](WATCH_DAEMON.md) | A resident development process that catches up changes and republishes the index |
13
+ | Keep the index current automatically | [Automatic maintenance guide](AUTOMATIC_MAINTENANCE.md) | A setup with supervised indexing, automatic reader refresh, and clear hook ownership |
14
14
  | Upgrade from Woods 1.x | [Upgrade to Woods 2.0](UPGRADING_TO_2.md) | A backed-up, re-indexed, verified v2 installation |
15
15
  | Diagnose an error | [Troubleshooting](TROUBLESHOOTING.md) | Symptom-to-cause checks for extraction, MCP, embeddings, storage, and Docker |
16
16
  | Contribute to Woods | [Contributing](../CONTRIBUTING.md) | A tested change with synchronized docs and plugin guidance |
@@ -41,6 +41,7 @@ and stdio or Streamable HTTP endpoints directly.
41
41
 
42
42
  ## Index lifecycle
43
43
 
44
+ - [Automatic maintenance guide](AUTOMATIC_MAINTENANCE.md): recommended development workflow and a section-by-section map of watcher, indexing, MCP, and hook documentation.
44
45
  - [Source freshness](SOURCE_FRESHNESS.md): verify dirty source against a served generation, establish a fresh-process baseline, and understand bounded unknown results.
45
46
 
46
47
  - [Retrieval guide](RETRIEVAL_GUIDE.md): configure embeddings and understand semantic retrieval, ranking, and token budgets.