woods 2.0.0.beta3 → 2.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (120) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +500 -420
  3. data/CONTRIBUTING.md +29 -17
  4. data/README.md +78 -178
  5. data/docs/AGENT_GUIDE.md +52 -11
  6. data/docs/AGENT_SETUP.md +34 -17
  7. data/docs/AUTOMATIC_MAINTENANCE.md +222 -0
  8. data/docs/BACKEND_MATRIX.md +18 -7
  9. data/docs/CLIENT_HOOKS.md +1 -1
  10. data/docs/CONFIGURATION_REFERENCE.md +105 -29
  11. data/docs/CONSOLE_MCP_SETUP.md +54 -9
  12. data/docs/DOCKER_SETUP.md +16 -1
  13. data/docs/EVALUATION.md +10 -4
  14. data/docs/EXTRACTOR_REFERENCE.md +23 -3
  15. data/docs/FAQ.md +14 -3
  16. data/docs/GETTING_STARTED.md +18 -17
  17. data/docs/INCREMENTAL_EXTRACTION.md +37 -8
  18. data/docs/INDEX_LAYOUT.md +2 -2
  19. data/docs/MCP_SERVERS.md +79 -7
  20. data/docs/MCP_TOOL_COOKBOOK.md +5 -5
  21. data/docs/MCP_WORKTREE_SETUP.md +55 -83
  22. data/docs/PUBLISHED_INDEX.md +17 -0
  23. data/docs/README.md +2 -1
  24. data/docs/RETRIEVAL_GUIDE.md +81 -13
  25. data/docs/SOURCE_FRESHNESS.md +1 -1
  26. data/docs/TOKEN_BENCHMARK.md +16 -10
  27. data/docs/TROUBLESHOOTING.md +142 -47
  28. data/docs/UPGRADING_TO_2.md +12 -6
  29. data/docs/WATCH_DAEMON.md +189 -24
  30. data/docs/WHY_WOODS.md +9 -5
  31. data/exe/woods-console +13 -11
  32. data/exe/woods-mcp-start +14 -9
  33. data/exe/woods-watch +5 -0
  34. data/lib/generators/woods/pgvector_generator.rb +8 -2
  35. data/lib/generators/woods/watch_generator.rb +53 -0
  36. data/lib/puma/plugin/woods.rb +10 -0
  37. data/lib/tasks/woods.rake +14 -0
  38. data/lib/woods/agent_configuration/applier.rb +5 -3
  39. data/lib/woods/agent_configuration/cli.rb +2 -2
  40. data/lib/woods/agent_configuration/layout.rb +13 -0
  41. data/lib/woods/cache/cache_middleware.rb +6 -0
  42. data/lib/woods/console/credential_scanner.rb +4 -3
  43. data/lib/woods/console/dispatch_pipeline.rb +7 -0
  44. data/lib/woods/console/embedded_executor.rb +31 -9
  45. data/lib/woods/console/sql_noise_stripper.rb +9 -7
  46. data/lib/woods/console/sql_table_scanner.rb +47 -7
  47. data/lib/woods/console/sql_validator.rb +49 -9
  48. data/lib/woods/console/sqlite_read_guard.rb +46 -0
  49. data/lib/woods/console/stdio_transport.rb +27 -0
  50. data/lib/woods/coordination/pipeline_lock.rb +3 -2
  51. data/lib/woods/embedding/indexer.rb +24 -14
  52. data/lib/woods/extractor.rb +70 -19
  53. data/lib/woods/extractors/declared_parent.rb +55 -0
  54. data/lib/woods/extractors/graphql_extractor.rb +2 -11
  55. data/lib/woods/extractors/lib_extractor.rb +10 -8
  56. data/lib/woods/extractors/mailer_extractor.rb +6 -10
  57. data/lib/woods/extractors/model_extractor.rb +1 -15
  58. data/lib/woods/extractors/poro_extractor.rb +10 -8
  59. data/lib/woods/extractors/shared_utility_methods.rb +22 -5
  60. data/lib/woods/git_command.rb +6 -7
  61. data/lib/woods/git_provenance.rb +4 -6
  62. data/lib/woods/mcp/bearer_auth.rb +2 -1
  63. data/lib/woods/mcp/bootstrapper.rb +20 -5
  64. data/lib/woods/mcp/config_resolver.rb +2 -1
  65. data/lib/woods/mcp/index_reader.rb +11 -2
  66. data/lib/woods/mcp/initialization_guidance.rb +1 -1
  67. data/lib/woods/mcp/renderers/markdown_renderer.rb +14 -8
  68. data/lib/woods/mcp/renderers/plain_renderer.rb +11 -7
  69. data/lib/woods/mcp/server.rb +63 -37
  70. data/lib/woods/mcp/tool_contract.rb +1 -1
  71. data/lib/woods/mcp/tool_response_renderer.rb +16 -0
  72. data/lib/woods/mcp/traversal_evidence_text.rb +1 -1
  73. data/lib/woods/mcp/traversal_response.rb +22 -0
  74. data/lib/woods/path_dispatcher.rb +6 -5
  75. data/lib/woods/published_index/typed_unit_reader.rb +40 -3
  76. data/lib/woods/published_index.rb +2 -2
  77. data/lib/woods/rake_helpers.rb +2 -12
  78. data/lib/woods/retrieval/corpus_status.rb +46 -0
  79. data/lib/woods/retrieval/lexical_assembler.rb +14 -3
  80. data/lib/woods/retrieval/lexical_index.rb +2 -1
  81. data/lib/woods/retriever.rb +19 -7
  82. data/lib/woods/session_tracer/file_store.rb +6 -1
  83. data/lib/woods/source_inputs/consumer_errors.rb +4 -0
  84. data/lib/woods/storage/local_corpus_stats.rb +32 -0
  85. data/lib/woods/storage/metadata_store.rb +20 -0
  86. data/lib/woods/storage/pgvector.rb +6 -2
  87. data/lib/woods/storage/vector_store.rb +10 -0
  88. data/lib/woods/temporal/json_snapshot_store.rb +35 -7
  89. data/lib/woods/version.rb +1 -1
  90. data/lib/woods/watch/child_environment.rb +30 -0
  91. data/lib/woods/watch/cli.rb +91 -0
  92. data/lib/woods/watch/daemon.rb +73 -11
  93. data/lib/woods/watch/event_stream.rb +70 -0
  94. data/lib/woods/watch/guardian.rb +142 -0
  95. data/lib/woods/watch/installation/layout.rb +70 -0
  96. data/lib/woods/watch/installation/options.rb +128 -0
  97. data/lib/woods/watch/installation/planner.rb +128 -0
  98. data/lib/woods/watch/installation/probe.rb +101 -0
  99. data/lib/woods/watch/installation/receipt.rb +77 -0
  100. data/lib/woods/watch/installation/recovery.rb +64 -0
  101. data/lib/woods/watch/installation/templates.rb +58 -0
  102. data/lib/woods/watch/installation.rb +56 -0
  103. data/lib/woods/watch/lifecycle.rb +182 -0
  104. data/lib/woods/watch/managed_child.rb +113 -0
  105. data/lib/woods/watch/managed_cleanup.rb +48 -0
  106. data/lib/woods/watch/managed_process.rb +144 -0
  107. data/lib/woods/watch/puma_adapter.rb +87 -0
  108. data/lib/woods/watch/puma_child.rb +66 -0
  109. data/lib/woods/watch/supervision_records.rb +95 -0
  110. data/lib/woods/watch/supervision_status.rb +104 -0
  111. data/lib/woods/watch/supervisor.rb +161 -0
  112. data/lib/woods/watch/supervisor_reporting.rb +46 -0
  113. data/plugin/.claude-plugin/plugin.json +1 -1
  114. data/plugin/hooks/woods-input-rules.sh +4 -4
  115. data/plugin/skills/woods-agent-enable/SKILL.md +7 -1
  116. data/plugin/skills/woods-diagnose/SKILL.md +134 -34
  117. data/plugin/skills/woods-investigate/SKILL.md +54 -15
  118. data/plugin/skills/woods-mcp-config/SKILL.md +38 -11
  119. data/plugin/skills/woods-setup/SKILL.md +72 -15
  120. metadata +38 -5
@@ -63,7 +63,7 @@ Keyword results are scored by how many distinct fields matched (identifier, sour
63
63
 
64
64
  ## Configuring Retrieval
65
65
 
66
- Retrieval requires an embedding provider and a vector store. Set these in `config/initializers/woods.rb`.
66
+ Semantic retrieval requires an embedding provider and a vector store. Configure these in `config/initializers/woods.rb` before embedding. For provider-free ranked retrieval over extraction output, use [explicit lexical mode](#embedding-free-lexical-retrieval) in the MCP process environment instead.
67
67
 
68
68
  ### Presets (recommended)
69
69
 
@@ -180,7 +180,12 @@ WOODS_RETRIEVAL_MODE=lexical bundle exec woods-mcp-start ./tmp/woods
180
180
  For a Ruby-built retriever, set `config.retrieval_mode = :lexical` and supply a
181
181
  populated metadata store to `Builder#build_retriever`. The packaged MCP server
182
182
  loads published unit JSON itself; it does not boot Rails or read current source
183
- files. Neither path constructs an embedding provider or vector adapter. The
183
+ files. Neither path constructs an embedding provider or vector adapter. If a
184
+ no-provider error suggests only OpenAI, Ollama or `search`, explicit lexical
185
+ mode is still available from `2.0.0.beta3`: set the variable in the MCP client
186
+ configuration and restart that server. Confirm `woods_status.retriever.mode`
187
+ is `lexical`; setting it only in a Rails initializer does not configure a
188
+ separate MCP process. The
184
189
  existing `:semantic` mode remains the default; provider failures never switch
185
190
  modes automatically. A Rails initializer is not loaded by the standalone MCP
186
191
  process, so set the environment variable in that process's client configuration.
@@ -191,15 +196,25 @@ and validations). Exact full identifiers rank first, with ambiguous typed owners
191
196
  retained. Other ties are deterministic. Responses name the lexical mode and
192
197
  matching fields/terms; runtime-field hits include the selected published runtime
193
198
  values. Lexical Ruby results leave the semantic-only `type_rank_context` table
194
- `nil`; they do not report a global vector rank or vector fallback. The top 20 eligible positive matches are considered for the
195
- output budget; this is ranked discovery, not an exhaustive match listing. Explicit
199
+ `nil`; they do not report a global vector rank or vector fallback. The top 20
200
+ eligible positive matches form the candidate shortlist for the output budget.
201
+ The lexical header reports `sources included`, `candidates considered`, and
202
+ `candidate limit: 20`. Included sources count the actual returned entries;
203
+ considered candidates count the shortlist after filtering and the limit, not
204
+ all matches or all documents examined. This is ranked discovery, not an
205
+ exhaustive match listing. Explicit
196
206
  `types` filters override default exclusions, as in semantic retrieval, and apply
197
207
  before that limit. A query with no lexical evidence returns no matches; unrelated
198
208
  graph hubs are never added. Query-seeded graph ranking is evaluation-only.
199
209
 
200
210
  The budget covers headers, matching explanations and truncation notices using a
201
211
  labelled character-based estimate, not an exact provider tokenizer. Full source
202
- remains available through `lookup`. Very small budgets can omit all sources. The
212
+ remains available through `lookup`. Count text is charged before source selection;
213
+ final counts do not trigger a second selection pass. Very small budgets can omit
214
+ all sources or clip the header itself. Zero candidates means no lexical matches;
215
+ positive candidates with zero included sources means no source entry fit the
216
+ available budget. The same counts apply to full, compact, outline and scoped
217
+ retrieval, agreeing with returned source attribution. The
203
218
  reader pins one published generation for building and querying its immutable
204
219
  lexical snapshot, rebuilding after publication. Corrupt units fail explicitly;
205
220
  they cannot quietly become a successful partial index. Older flat indexes rebuild
@@ -276,16 +291,65 @@ result = retriever.retrieve("what validations does Order have?")
276
291
 
277
292
  ---
278
293
 
279
- ## Degradation Tiers
280
-
281
- Retrieval degrades gracefully when components are unavailable. The Retriever itself does not implement explicit fallback tiers, degradation happens naturally through how each component handles errors:
294
+ ## Semantic corpus diagnostics
295
+
296
+ Extraction and embedding publish different data. `woods_status.ready` describes
297
+ the structural index; neither that flag nor bootstrap `hydrated` proves that
298
+ semantic retrieval has indexed records. Provider detection alone can succeed
299
+ before the first embedding run.
300
+
301
+ Supporting readers report `woods_status.retriever.corpus`:
302
+
303
+ - `state`: `empty`, `metadata_only`, `vectors_only`, `nonempty`, or `unknown`.
304
+ - `vectors` and `metadata`: each has `count`, `by_type`, and `untyped_count`.
305
+ Counts describe stored entries, including chunks; they are not distinct
306
+ extracted-unit counts or a completeness certificate. Missing type labels are
307
+ reported separately rather than assigned to a guessed type.
308
+ - Unknown counts are `null`. Diagnostics use an explicit local-store capability;
309
+ they do not query remote stores to discover their counts. A missing capability
310
+ does not mean the backend is empty or broken.
311
+
312
+ When both semantic stores are known empty, `codebase_retrieve` reports
313
+ `empty_index` with recovery guidance instead of presenting an empty match as
314
+ evidence that application code is absent. Run `woods:embed` in the application
315
+ with the intended provider and storage configuration, then reload or restart
316
+ the reader. Restart after changing provider or store configuration; reload
317
+ refreshes the stores of the existing retriever. Alternatively, explicitly select `WOODS_RETRIEVAL_MODE=lexical`
318
+ in the MCP process and restart for ranked retrieval over the published units.
319
+ Woods does not change retrieval mode automatically.
320
+
321
+ Metadata-only stores can still answer some keyword, direct, and graph queries;
322
+ they are not blocked by this diagnostic. Source-empty units deliberately retain
323
+ metadata without vectors, so the two counts need not match. Nonempty stores can
324
+ still have missing types, stale vectors, or provider failures. Inspect the
325
+ query result and embedding evidence before claiming coverage. The type-rank
326
+ table's metadata count describes the retrieval metadata store, not all units
327
+ in the structural index.
328
+
329
+ Older readers may omit `retriever.corpus`; record the reader revision separately
330
+ from the index writer version and verify the embedding artifacts directly.
331
+ Explicit lexical mode does not use or report semantic corpus counts.
282
332
 
283
- - **Embedding provider unavailable**: `codebase_retrieve` returns a structured configuration error. Check `woods_status` for retrieval readiness.
284
- - **Vector store unavailable**: vector and hybrid strategies fail at query time. Keyword and graph strategies remain available for direct calls to `SearchExecutor`.
285
- - **Metadata store error**: the structural context overview (unit counts by type) is silently omitted; `Retriever#build_structural_context` rescues `StandardError` and returns `nil`. The retrieval result is still returned without the overview.
286
- - **Graph store unavailable**: graph expansion in hybrid strategy produces no graph candidates; vector and keyword candidates are still ranked and returned.
333
+ ## Degradation Tiers
287
334
 
288
- In all cases, errors in individual components produce empty candidate sets for that source rather than raising through the `Retriever`. Configure circuit breakers via `Woods::Resilience::CircuitBreaker` on external providers (Qdrant, OpenAI) for production deployments.
335
+ The MCP boundary distinguishes missing configuration, empty stores, and failed
336
+ stores. A failure is not evidence that no application code matches:
337
+
338
+ - **No embedding provider configured:** `codebase_retrieve` reports a configuration
339
+ error with embedding and explicit lexical-mode options.
340
+ - **Both semantic stores known empty:** the tool reports `empty_index`; see the
341
+ [corpus diagnostics](#semantic-corpus-diagnostics) above.
342
+ - **Failed dump hydration:** the tool reports `degraded_index` rather than serving
343
+ a clean empty result. Repair or regenerate the named embedding artifact.
344
+ - **Query-time storage failures:** vector, metadata, and graph adapter exceptions
345
+ are translated into store errors and reported as `degraded_index` by MCP.
346
+ This includes failures building the metadata overview; that failure is not
347
+ silently omitted. Direct Ruby callers should handle `Woods::Retriever::StoreError`.
348
+
349
+ Provider failures and missing metadata for returned candidates have their own
350
+ error paths. Preserve their diagnostics; do not silently switch modes or treat
351
+ an exception as an empty match. Positive corpus counts do not override these
352
+ checks.
289
353
 
290
354
  ---
291
355
 
@@ -353,6 +417,10 @@ bundle exec rake woods:embed
353
417
  | `text-embedding-3-small` (default) | 1536 |
354
418
  | `text-embedding-3-large` | 3072 |
355
419
 
420
+ Woods' pgvector HNSW adapter supports at most 2,000 dimensions. For the large
421
+ model, request a supported output width explicitly or choose another backend;
422
+ see [pgvector configuration](CONFIGURATION_REFERENCE.md#pgvector-postgresql).
423
+
356
424
  **Ollama default model:** `nomic-embed-text`. Dimensions are detected dynamically on first embed.
357
425
 
358
426
  ---
@@ -4,7 +4,7 @@
4
4
  application inputs with source bytes visible to the reader. It is separate from
5
5
  index age, HEAD equality, daemon liveness and external database/runtime state.
6
6
 
7
- This capability is unreleased after `2.0.0.beta2`. Check the installed gem's
7
+ This capability is included in Woods `2.0.0`. Check the installed gem's
8
8
  `woods-extract --help`, `rake -T woods:source_status`, and `woods_status` schema
9
9
  before using it; upgrading the plugin alone does not upgrade Woods.
10
10
 
@@ -13,7 +13,8 @@
13
13
  > When the optional [`tokenizers`](https://github.com/ankane/tokenizers-ruby)
14
14
  > gem is installed, the Ollama path uses the real BERT WordPiece tokenizer
15
15
  > (`Woods::Embedding::TokenCounter`) instead of this heuristic. The 4.0
16
- > divisor below is what the gem falls back to everywhere else.
16
+ > divisor applies to the OpenAI/default path; Ollama falls back to 1.5 when
17
+ > its tokenizer is unavailable.
17
18
 
18
19
  This is a historical record of the benchmark that picked 4.0 over the
19
20
  original 3.5 divisor. It is cited from five places in `lib/` as the evidence
@@ -35,17 +36,21 @@ for that choice, keep the numbers below intact if you edit this doc.
35
36
  | 3.8 | 16.2% | 42.5% |
36
37
  | **4.0 (shipped)** | **10.6%** | **35.4%** |
37
38
 
38
- Mean chars/token across the corpus was **4.41** (range 3.94–5.42). The
39
- heuristic always overestimated, never underestimated, across all 19 files,
40
- which is what makes it safe for token-limit enforcement even at its worst
41
- case. Code lines and comment/YARD lines had similar ratios (4.38 vs. 4.27
42
- chars/token), no separate handling needed for either.
39
+ Mean chars/token across the corpus was **4.41** (range 3.94–5.42). These
40
+ aggregate results favor the 4.0 divisor for this sample; they do not establish
41
+ an upper bound on token counts. The recorded range includes values below 4.0,
42
+ so the heuristic can underestimate. Code lines and comment/YARD lines had
43
+ similar ratios (4.38 vs. 4.27 chars/token) in this sample.
44
+
45
+ The character estimate covers only the text passed to the counter. It does not
46
+ bound the serialized MCP response: text rendering, structured output, provenance,
47
+ and JSON framing can add bytes and tokens beyond that input.
43
48
 
44
49
  ## What shipped
45
50
 
46
51
  **The divisor changed from 3.5 to 4.0.** It roughly halves the mean
47
- overestimate (26.2% → 10.6%) while keeping the conservative
48
- always-overestimates property, at zero new runtime dependencies. The
52
+ error (26.2% → 10.6%) at zero new runtime dependencies. It remains an estimate,
53
+ not a guarantee that arbitrary input fits a model's token limit. The
49
54
  constant lives in one place now (`Woods::TokenUtils::CHARS_PER_TOKEN_BY_PROVIDER`),
50
55
  not scattered across call sites, see `lib/woods/token_utils.rb` for the
51
56
  current definition and `docs/EMBEDDING_MODELS.md` for the Ollama-side ratio.
@@ -53,8 +58,9 @@ current definition and `docs/EMBEDDING_MODELS.md` for the Ollama-side ratio.
53
58
  **tiktoken_ruby was deliberately not added as a runtime dependency.** A 10.6%
54
59
  mean error is acceptable for chunking decisions, budget estimates, and
55
60
  truncation; a native-extension dependency for marginal accuracy gains wasn't
56
- worth it. The optional `tokenizers` gem covers the case where exact counts
57
- matter more (see above).
61
+ worth it. The optional `tokenizers` gem provides counts for its supported BERT
62
+ WordPiece tokenizer, not every model. Strict token-limit enforcement requires
63
+ the tokenizer used by the target model.
58
64
 
59
65
  ## Reproducing this benchmark
60
66
 
@@ -8,7 +8,7 @@ This guide covers the most common problems encountered when installing, extracti
8
8
 
9
9
  | Error message | Cause | Fix |
10
10
  |---------------|-------|-----|
11
- | `No manifest.json found` | Wrong index path or no published generation | Use the path visible to the server process; run `woods:validate` |
11
+ | `Could not resolve a published Woods index` (older versions: `No manifest.json found`) | Wrong index path or unresolved published generation | Select the existing index using a path visible to the server process; see [startup diagnostics](#index-cannot-be-resolved-at-startup) |
12
12
  | `uninitialized constant Rails` | Not running inside Rails app | Run via `bundle exec rake` in Rails root |
13
13
  | `type "vector" does not exist` | pgvector not installed | `CREATE EXTENSION vector` in PostgreSQL |
14
14
  | `Connection refused (localhost:11434)` | Ollama not running | `ollama serve` |
@@ -23,6 +23,7 @@ This guide covers the most common problems encountered when installing, extracti
23
23
  | `No such container` | Wrong container name | Check with `docker ps --format '{{.Names}}'` |
24
24
  | `JSON parse errors` (MCP) | Rails boot noise on stdout | Remove `puts` calls from initializers |
25
25
  | Query timeout | Large table, no scope | Add scope conditions to narrow results |
26
+ | `Extraction failed for …; the previous generation remains active` | A consumer handled a source error during incremental extraction or refresh (included in Woods `2.0.0`) | Fix the logged source error and retry the [complete batch](INCREMENTAL_EXTRACTION.md#handled-source-errors-and-retry); watch keeps it pending |
26
27
  | Empty extraction output | `eager_load!` failure | Check for `NameError` in boot output |
27
28
  | Git metadata missing | Shallow clone in CI | Use `fetch-depth: 0` for complete history |
28
29
  | Parallel tool calls all fail | MCP client batches calls | Send calls sequentially, validate params first |
@@ -52,13 +53,44 @@ version. See [manifest writer provenance](PUBLISHED_INDEX.md#manifest-writer-pro
52
53
 
53
54
  If a tool call fails with **"Tool not found: … not available in the installed Woods v…"**, the client is asking for a tool a newer gem provides. Run `bundle update woods` and reconnect the MCP server, then retry.
54
55
 
56
+ ### Watcher startup or planned restart fails
57
+
58
+ Managed `woods-watch` startup is **included in Woods `2.0.0`**; record the
59
+ loaded version/path and revision, then verify executable and generator help.
60
+ If changing an initializer stops every Foreman process, replace a bare
61
+ `woods:watch` entry with the [managed setup](WATCH_DAEMON.md#managed-development-startup).
62
+
63
+ Read launcher logs and `woods_status` supervision records separately from daemon
64
+ liveness and index freshness. `retrying` means the last generation remains usable
65
+ while boot is retried. A parked ownership/protocol conflict requires correcting
66
+ the selected owner or installed command and restarting that owner; do not delete
67
+ claim files or kill PIDs taken from status. No index-visible record exists before
68
+ the first boot resolves the application's output directory.
69
+
70
+ Unset `WOODS_WATCH_IDLE_TIMEOUT` in managed modes. If the boot deadline is reached,
71
+ diagnose Bundler/initializer startup before increasing `--boot-timeout`; a valid
72
+ long extraction has a separate readiness state and is not bounded by that clock.
73
+ If setup created a Procfile but normal `bin/dev` still only launches Rails, choose
74
+ Puma or explicitly run the selected Foreman command. The generator never rewrites
75
+ `bin/dev` or starts services during preview.
76
+
77
+ If installation reports a pending transaction, use `woods:watch --operation
78
+ recover` through the Rails generator, initially with `--pretend`; see
79
+ [owned setup recovery](WATCH_DAEMON.md#ownership-updates-and-removal). That Rails
80
+ command boots the application first. For broken initializers use the documented
81
+ direct bundled Ruby helper, which does not boot Rails or require task discovery.
82
+ Both refuse to overwrite intervening edits. A Puma setup refusal for
83
+ `config/puma/development.rb` means the default
84
+ configuration would bypass the generated plugin; select an external/Foreman
85
+ arrangement instead of installing an inactive directive.
86
+
55
87
  ### Semantic graph validation errors
56
88
 
57
89
  In development versions containing #413, `woods:validate` rejects graphs that
58
90
  parse as JSON but disagree with their indexes. Errors name the section and
59
91
  identity, for example `reverse["http_api"]: missing "Order"`, a duplicate typed
60
- variant, or an indexed unit absent from `nodes`. This is unreleased after
61
- `2.0.0.beta2`; check the installed gem before expecting these diagnostics.
92
+ variant, or an indexed unit absent from `nodes`. This is included in Woods
93
+ `2.0.0`; check the installed gem before expecting these diagnostics.
62
94
 
63
95
  Keep the failing generation and report the exact errors. Run a full extraction
64
96
  in a fresh application process with the intended bundle, then validate again.
@@ -87,7 +119,7 @@ Scoped resets leave corrupt state untouched. Valid state keeps any unrelated
87
119
  operation entries, and missing state remains a no-op without creating a file.
88
120
  A permission failure must be corrected before repair can succeed.
89
121
 
90
- This recovery is unreleased after `2.0.0.beta2`; check the installed version.
122
+ This recovery is included in Woods `2.0.0`; check the installed version.
91
123
  Older versions report corrupt state as nothing to repair. Stop pipeline writers,
92
124
  back up the configured guard state's `pipeline_guard.json`, and remove only that
93
125
  file before restarting, or upgrade to a version containing the fix.
@@ -159,7 +191,12 @@ For subsequent runs, use incremental mode instead of full extraction:
159
191
  bundle exec rake woods:incremental
160
192
  ```
161
193
 
162
- Incremental extraction only re-extracts files that changed since the last run. It skips unchanged units and is typically 5-10× faster.
194
+ Incremental extraction dispatches the selected changed paths, including affected
195
+ concern consumers and whole-app extractors whose trigger paths changed. The default
196
+ Git range is `HEAD~1`; pass an explicit range or `CHANGED_FILES` for other batches.
197
+ It can reduce extraction work, but Rails boot, graph rebuilding, and publication
198
+ still contribute to runtime. Measure the improvement in your application; Woods
199
+ does not guarantee a speedup. See the [incremental contract](INCREMENTAL_EXTRACTION.md).
163
200
 
164
201
  ---
165
202
 
@@ -227,7 +264,7 @@ dependents after an incremental run.
227
264
  identities when restoring and updating the graph.
228
265
 
229
266
  **Fix:** Check whether the installed version includes B-193; this fix is
230
- unreleased. After upgrading to a version containing the fix, run
267
+ included in Woods `2.0.0`. After upgrading to a version containing the fix, run
231
268
  `bundle exec rake woods:extract` once to rebuild lost reverse dependencies.
232
269
  Loading an already damaged graph does not restore discarded entries. See the
233
270
  [incremental graph contract](INCREMENTAL_EXTRACTION.md#the-contract).
@@ -238,7 +275,7 @@ Loading an already damaged graph does not restore discarded entries. See the
238
275
  most files as `change_frequency: new` in a shallow CI checkout.
239
276
 
240
277
  **Cause:** A shallow clone truncates HEAD ancestry. The shallow-checkout guard is
241
- unreleased after 2.0.0.beta2: current source omits git enrichment and warns once,
278
+ included in Woods `2.0.0`: Woods omits git enrichment and warns once,
242
279
  rather than treating the truncated history as complete. If repository depth
243
280
  cannot be verified, enrichment is also omitted; check git access and version.
244
281
 
@@ -255,12 +292,33 @@ clone), then run full extraction to replace retained metadata:
255
292
  Two commits can suffice for an incremental diff, but do not establish the full
256
293
  ancestry needed for churn metadata.
257
294
 
295
+ ### Git executable is missing from the extraction environment
296
+
297
+ **Symptom:** Extraction logs `Git history unavailable: git executable was not
298
+ found in PATH`, and newly extracted units have no `metadata.git`. Older builds
299
+ can omit this enrichment silently when the executable is missing.
300
+
301
+ **Fix:** Run `git --version` in the same container and environment that runs
302
+ extraction. Install Git 2.31 or newer there, ensure its executable is on `PATH`,
303
+ then run full `woods:extract` to refresh every unit's history. A working Git
304
+ installation on the host does not provide Git inside an application container.
305
+
306
+ Extraction continues without inventing zero-commit history. The warning appears
307
+ once per extractor instance when the application has a `.git` entry or an
308
+ explicit `WOODS_GIT_DIR`/`GIT_DIR` setting. A source archive with neither remains
309
+ supported and quiet. `GIT_BRANCH`/`GIT_SHA` provenance fallback is unchanged;
310
+ those values identify a build but cannot supply per-file history.
311
+
312
+ This diagnostic is emitted during extraction. `woods_status.ready` and a
313
+ manifest Git SHA do not establish that per-unit history was available, and
314
+ `recent_changes` returning no results does not prove no files changed.
315
+
258
316
  ---
259
317
 
260
318
  ### Git enrichment warns that history could not be read completely
261
319
 
262
- Current source uses an explicit merge-diff mode requiring **Git 2.31 or newer**.
263
- This is unreleased after 2.0.0.beta2: first confirm the installed Woods version.
320
+ Woods 2.0 uses an explicit merge-diff mode requiring **Git 2.31 or newer**.
321
+ First confirm the installed Woods version.
264
322
  Check `git --version` inside the same container/process environment as extraction,
265
323
  and upgrade git if it is older. On a supported version, check that the application's
266
324
  `HEAD` and object store can be read using the same `WOODS_GIT_DIR` setting.
@@ -292,29 +350,62 @@ When it does not, the git keys are omitted from every unit, provenance records
292
350
  `"unknown"`, and one warning names git's own reason. Absent keys mean "not
293
351
  known"; they never mean "brand new".
294
352
 
295
- **Fix:** Point `WOODS_GIT_DIR` at the *canonical* git directory, the one the
296
- worktree's `gitdir:` pointer ultimately leads to, and make sure it is mounted:
353
+ **Fix:** Restore access to both the worktree-specific Git directory and the
354
+ shared objects and refs using the mount layouts below, then run a full
355
+ `woods:extract` to replace retained metadata.
356
+
357
+ ### Git directory mounts for linked worktrees
358
+
359
+ `WOODS_GIT_DIR` is passed directly to Git's `--git-dir`. It selects that
360
+ directory's `HEAD` for manifest provenance, per-file history, and incremental
361
+ diff ranges. **For a linked worktree, pointing it at the shared `.git` root
362
+ selects the primary checkout's HEAD.** A successful Git command alone does
363
+ not prove Woods is reading the intended branch.
364
+
365
+ First inspect Git metadata on the host, from the intended worktree:
297
366
 
298
367
  ```bash
299
- # docker-compose.yml, mounting the parent repository's git directory
300
- # volumes:
301
- # - /path/to/repo/.git:/canonical-git:ro
302
- WOODS_GIT_DIR=/canonical-git bundle exec rake woods:extract
368
+ git -C /path/to/worktree rev-parse --absolute-git-dir
369
+ # Example: /path/to/repo/.git/worktrees/wt
370
+ git -C /path/to/worktree rev-parse --path-format=absolute --git-common-dir
371
+ # Example: /path/to/repo/.git
372
+ git -C /path/to/worktree rev-parse --abbrev-ref HEAD
373
+ git -C /path/to/worktree rev-parse HEAD
303
374
  ```
304
375
 
305
- `WOODS_GIT_DIR` wins over whatever the worktree pointer says, and applies to
306
- every git call Woods makes: per-unit enrichment, `manifest.json` provenance,
307
- and the diff range `woods:incremental` resolves. All three run through
308
- `Woods::GitCommand.argv`, so the override cannot reach two of them and miss the
309
- third.
376
+ The worktree ID in this example is `wt`. Use the ID returned by Git metadata;
377
+ it need not match the branch name. Choose one of these layouts:
378
+
379
+ - **Same-path mount:** mount the complete shared directory read-only at its
380
+ original absolute path (`/path/to/repo/.git:/path/to/repo/.git:ro`). With the
381
+ application's existing `.git` pointer resolvable, leave `WOODS_GIT_DIR`
382
+ unset and remove conflicting Git-directory overrides from the environment.
383
+ - **Relocated mount:** mount that complete directory read-only at a new path
384
+ (`/path/to/repo/.git:/mounted-common:ro`), including `objects`, `refs`, and
385
+ `worktrees`. Select the worktree-specific directory inside it:
386
+
387
+ ```bash
388
+ WOODS_GIT_DIR=/mounted-common/worktrees/wt bundle exec rake woods:extract
389
+ ```
310
390
 
311
- **`GIT_DIR` alone is not enough for a linked worktree.** Woods honors git's own
312
- `GIT_DIR` and `GIT_COMMON_DIR` because git does, but setting `GIT_DIR` to a
313
- worktree's private git directory only moves the failure: the `commondir`
314
- pointer inside it is relative, so it still resolves to a path that is not
315
- mounted, and `GIT_COMMON_DIR` does not override it. Either mount the canonical
316
- git directory at the same absolute path the pointer names, or use
317
- `WOODS_GIT_DIR`.
391
+ Mounting only the private worktree directory can leave its `commondir` pointer
392
+ without access to shared objects and refs. Git's own environment variables are
393
+ inherited by the subprocess; check any existing `GIT_DIR` and `GIT_COMMON_DIR`
394
+ settings when diagnosing the effective layout. The complete layouts above
395
+ preserve both worktree identity and shared storage.
396
+
397
+ In the extraction container, verify the relocated selection against the host
398
+ branch and exact SHA before extracting (replace `/app` and `wt` as needed):
399
+
400
+ ```bash
401
+ git --git-dir=/mounted-common/worktrees/wt --work-tree=/app -C /app rev-parse --abbrev-ref HEAD
402
+ git --git-dir=/mounted-common/worktrees/wt --work-tree=/app -C /app rev-parse HEAD
403
+ ```
404
+
405
+ After fixing the selection, run full `woods:extract` and verify the published
406
+ manifest. Incremental extraction can retain older per-file Git metadata.
407
+ A commit alone does not necessarily trigger the source-file watcher; run a
408
+ full extraction when current history and provenance are required.
318
409
 
319
410
  ---
320
411
 
@@ -374,29 +465,39 @@ retries after a later filesystem event.
374
465
 
375
466
  ### `manifest.json` shows the wrong branch (or `git_branch: "unknown"`) in a worktree
376
467
 
377
- **Symptom:** `git_branch` / `git_sha` in `manifest.json` name a different branch than the worktree is actually on, or report `"unknown"`. The extracted units themselves are correct, only the provenance metadata is off.
468
+ **Symptom:** `git_branch` / `git_sha` in `manifest.json` name a different
469
+ branch or SHA than the intended worktree, or report `"unknown"`.
378
470
 
379
- **Cause:** In a linked git worktree, `.git` is a *file* containing a `gitdir:` pointer to the real git directory, often an absolute host path. When extraction runs where that path can't be resolved (e.g. inside a container where the host path isn't mounted), git can't read the ref. Woods now reports `"unknown"` in that case rather than emitting a stale, misleading value (previously it fell back to a baked `GIT_BRANCH`/`GIT_SHA` build arg).
471
+ **Cause:** An unreachable `.git` file's `gitdir:` pointer prevents Git from
472
+ resolving the worktree's HEAD. An override selecting the shared `.git` root
473
+ instead resolves the primary checkout's HEAD successfully. That wrong selection
474
+ also affects per-file history and HEAD-based incremental ranges.
380
475
 
381
- **Fix:** Make the worktree's git directory reachable from the extraction environment, for example, mount the parent repository (the directory the `gitdir:` pointer references) into the container, or run extraction from a normal (non-worktree) checkout. With the real git directory reachable, `git_branch`/`git_sha` resolve correctly. If the checkout legitimately ships without a `.git` at all (a source tarball, or a Docker `COPY` that excludes it), set `GIT_BRANCH` / `GIT_SHA` explicitly. Woods honors these when there is no `.git` at the root (or no git binary), but suppresses them when a `.git` *is* present but unresolvable (so a stale build arg can't mask a worktree).
476
+ **Fix:** Follow [Git directory mounts for linked worktrees](#git-directory-mounts-for-linked-worktrees),
477
+ compare the selected branch and exact SHA in the extraction environment, then
478
+ run a full extraction. Compare the newly published manifest, not a retained
479
+ generation. A commit without a source edit may leave the watcher idle.
382
480
 
383
- When the canonical git directory is mounted but not at the path the pointer
384
- names, set `WOODS_GIT_DIR` to where it actually is. It wins over the pointer for
385
- provenance and for unit-level git metadata both. Setting git's own `GIT_DIR` to
386
- the worktree's private git directory does not work: its `commondir` pointer is
387
- relative and resolves outside the mount.
481
+ For a checkout legitimately shipped without `.git` (such as a source tarball),
482
+ `GIT_BRANCH` / `GIT_SHA` can supply provenance. They are fallbacks only when
483
+ `.git` is absent or Git is unavailable; a present but unresolvable `.git`
484
+ reports `"unknown"` instead of substituting stale build arguments.
388
485
 
389
486
  ---
390
487
 
391
488
  ## MCP Server Problems
392
489
 
393
- ### "No manifest.json" error when starting the Index Server
490
+ <a id="no-manifestjson-error-when-starting-the-index-server"></a>
491
+
492
+ ### Index cannot be resolved at startup
394
493
 
395
- **Symptom:** `woods-mcp-start` exits with an error like `No manifest.json found at /path/to/...` even though extraction completed.
494
+ **Symptom:** An Index MCP executable exits with `Could not resolve a published Woods index in: /path/to/...` even though extraction completed. This headline is included in Woods `2.0.0`; older versions say `No manifest.json found`. Both mean the selected index could not resolve its manifest, not that an atomic index needs a root manifest.
396
495
 
397
- **Cause:** The Index Server is using the container-internal path rather than the host-side path to the volume-mounted output. The server runs on the host and cannot access container filesystem paths.
496
+ Embedded Index MCP startup through `IndexReader` also raises an `ArgumentError` with the selected directory and layout guidance when the marker cannot resolve a manifest, including malformed marker shapes such as `[]` or a numeric `payload` (included in Woods `2.0.0`). Earlier builds may expose a raw `TypeError` or `NoMethodError` for those shapes. Inspect the marker and preserve the failing index before attempting recovery.
398
497
 
399
- **Fix:** Use the host path in your `.mcp.json`:
498
+ **Cause:** The selected directory is not the published index root, the published generation cannot be resolved, or the path is not visible to the MCP process. A container path is appropriate for a container process; a host process needs the host-visible path.
499
+
500
+ **Fix:** Point at the existing index before extracting again. Check the examined directory in the error and the [MCP path precedence](CONFIGURATION_REFERENCE.md#environment-variables). For a host-side launch whose working directory contains `tmp/woods`, for example:
400
501
 
401
502
  ```json
402
503
  {
@@ -409,13 +510,7 @@ relative and resolves outside the mount.
409
510
  }
410
511
  ```
411
512
 
412
- Verify the output is accessible from the host:
413
-
414
- ```bash
415
- ls ./tmp/woods/manifest.json
416
- ```
417
-
418
- **Since Woods 2.0, a healthy index may not have `manifest.json` at the output root at all.** Extraction publishes each generation into an immutable `payloads/gen-<N>/` directory and points to it from `generation.json`. If the flat path is missing, check the payload path instead before assuming extraction failed:
513
+ **Since Woods 2.0, a healthy index may not have `manifest.json` at the output root at all.** Extraction publishes each generation into an immutable `payloads/gen-<N>/` directory and points to it from `generation.json`. Inspect the marker and its payload in the MCP process's filesystem before assuming extraction failed (use the generation named by your marker):
419
514
 
420
515
  ```bash
421
516
  cat ./tmp/woods/generation.json # {"number": 42, "payload": "payloads/gen-42", ...}
@@ -427,7 +522,7 @@ Update their gate using the [filesystem layout contract](INDEX_LAYOUT.md), which
427
522
  includes Bash/jq and Python readers. An upload must pin and copy one complete
428
523
  payload before publishing its captured pointer; keep a failed copy unpublished.
429
524
 
430
- `woods-mcp-start` and `IndexReader` already resolve this automatically, this is only for manual inspection. If neither path has a manifest, your Docker volume mount is not configured correctly. See [DOCKER_SETUP.md](DOCKER_SETUP.md).
525
+ `woods-mcp-start` and `IndexReader` resolve this automatically; these commands are for manual inspection. Legacy flat indexes use a root `manifest.json`. If neither layout resolves, check the selected path, pointer, payload and any volume mount. See [DOCKER_SETUP.md](DOCKER_SETUP.md) for container launches.
431
526
 
432
527
  ---
433
528
 
@@ -2,12 +2,10 @@
2
2
 
3
3
  Woods 2.0 changes observable index identifiers, publication layout, vector-store reconciliation, and the supported MCP surface. Plan a clean re-index. Do not upgrade a shared or durable index in place without a backup and a rollback window.
4
4
 
5
- This guide assumes the last v1 release, 1.6.1, and targets 2.0.0.
5
+ This guide covers the supported 1.6.x line and targets 2.0.0. Use the latest
6
+ published 1.6.x security patch as the rollback version.
6
7
 
7
8
  <!-- release-state:upgrade-availability -->
8
- > RubyGems lists 2.0.0.beta3 as a prerelease. Pin it explicitly with
9
- > `gem "woods", "2.0.0.beta3"`; `~> 2.0` resolves only once
10
- > 2.0.0 is published.
11
9
  <!-- release-state:end -->
12
10
 
13
11
  ## Upgrade outcome
@@ -95,7 +93,11 @@ Also back up managed Obsidian/Unblocked destinations before allowing a mass stal
95
93
 
96
94
  ### 3. Choose a rollback point
97
95
 
98
- Keep the v1 Gemfile/lockfile commit and all durable-store backups until v2 extraction, MCP calls, retrieval, and exports are verified. Downgrading the gem does not translate v2 identifiers back to v1.
96
+ Record and test a Gemfile/lockfile selecting the latest published 1.6.x security
97
+ patch as the rollback bundle. If the current installation is older, verify that
98
+ patched v1 bundle before beginning the v2 migration. Keep its commit and all
99
+ durable-store backups until v2 extraction, MCP calls, retrieval, and exports are
100
+ verified. Downgrading the gem does not translate v2 identifiers back to v1.
99
101
 
100
102
  ## Upgrade the application
101
103
 
@@ -142,6 +144,10 @@ bin/rails woods:validate
142
144
  bin/rails woods:stats
143
145
  ```
144
146
 
147
+ Included in Woods `2.0.0`: `woods:clean` removes index artifacts but keeps
148
+ the output directory and its hidden extraction guard. This stable guard lets
149
+ concurrent writers coordinate safely; its presence does not mean an index remains.
150
+
145
151
  The clean extract is required for corrected identifier shapes. Do not use an incremental run as the first v2 extraction: after `woods:clean` there is no baseline, and v2 `woods:incremental` refuses that state rather than publishing a near-empty index as the application's complete truth.
146
152
 
147
153
  An interrupted extraction leaves readers on the last complete generation because Woods publishes `generation.json` only after the payload is complete. Re-run the task; do not delete a partial directory speculatively. A run that completes its payload but cannot publish the marker now fails loudly instead of reporting success, so treat a non-zero exit as work to redo rather than as a partial success.
@@ -327,7 +333,7 @@ Complete every applicable check:
327
333
  If verification fails:
328
334
 
329
335
  1. stop v2 MCP, watcher, embedding, and exporter processes;
330
- 2. restore the v1 Gemfile and lockfile or deploy the recorded v1 commit;
336
+ 2. restore the tested, patched v1 Gemfile and lockfile or deploy its recorded commit;
331
337
  3. run the v1 `woods:clean` before restoring anything under the configured output directory;
332
338
  4. either restore the complete pre-upgrade v1 output-directory backup, or run a fresh v1 extraction and then restore its v1 `dumps/` and configuration artifacts;
333
339
  5. restore external vector-store and managed export backups when v2 modified them;