woods 2.0.0.beta4 → 2.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (79) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +480 -477
  3. data/CONTRIBUTING.md +2 -2
  4. data/README.md +11 -26
  5. data/docs/AGENT_GUIDE.md +31 -12
  6. data/docs/AGENT_SETUP.md +17 -10
  7. data/docs/AUTOMATIC_MAINTENANCE.md +222 -0
  8. data/docs/BACKEND_MATRIX.md +13 -7
  9. data/docs/CLIENT_HOOKS.md +1 -1
  10. data/docs/CONFIGURATION_REFERENCE.md +36 -26
  11. data/docs/CONSOLE_MCP_SETUP.md +10 -8
  12. data/docs/DOCKER_SETUP.md +15 -0
  13. data/docs/EVALUATION.md +10 -4
  14. data/docs/EXTRACTOR_REFERENCE.md +14 -2
  15. data/docs/FAQ.md +14 -3
  16. data/docs/GETTING_STARTED.md +18 -17
  17. data/docs/INCREMENTAL_EXTRACTION.md +8 -3
  18. data/docs/INDEX_LAYOUT.md +2 -2
  19. data/docs/MCP_SERVERS.md +28 -11
  20. data/docs/MCP_TOOL_COOKBOOK.md +1 -1
  21. data/docs/MCP_WORKTREE_SETUP.md +13 -1
  22. data/docs/PUBLISHED_INDEX.md +1 -1
  23. data/docs/README.md +2 -1
  24. data/docs/RETRIEVAL_GUIDE.md +57 -8
  25. data/docs/SOURCE_FRESHNESS.md +1 -1
  26. data/docs/TOKEN_BENCHMARK.md +16 -10
  27. data/docs/TROUBLESHOOTING.md +133 -37
  28. data/docs/UPGRADING_TO_2.md +9 -7
  29. data/docs/WATCH_DAEMON.md +172 -17
  30. data/docs/WHY_WOODS.md +9 -5
  31. data/exe/woods-console +13 -11
  32. data/exe/woods-watch +5 -0
  33. data/lib/generators/woods/watch_generator.rb +53 -0
  34. data/lib/puma/plugin/woods.rb +10 -0
  35. data/lib/tasks/woods.rake +14 -0
  36. data/lib/woods/cache/cache_middleware.rb +6 -0
  37. data/lib/woods/console/stdio_transport.rb +27 -0
  38. data/lib/woods/extractor.rb +25 -7
  39. data/lib/woods/git_command.rb +6 -7
  40. data/lib/woods/git_provenance.rb +4 -6
  41. data/lib/woods/mcp/bootstrapper.rb +3 -1
  42. data/lib/woods/mcp/initialization_guidance.rb +1 -1
  43. data/lib/woods/mcp/server.rb +41 -9
  44. data/lib/woods/retrieval/corpus_status.rb +46 -0
  45. data/lib/woods/retriever.rb +19 -7
  46. data/lib/woods/storage/local_corpus_stats.rb +32 -0
  47. data/lib/woods/storage/metadata_store.rb +20 -0
  48. data/lib/woods/storage/vector_store.rb +10 -0
  49. data/lib/woods/version.rb +1 -1
  50. data/lib/woods/watch/child_environment.rb +30 -0
  51. data/lib/woods/watch/cli.rb +91 -0
  52. data/lib/woods/watch/daemon.rb +55 -7
  53. data/lib/woods/watch/event_stream.rb +70 -0
  54. data/lib/woods/watch/guardian.rb +142 -0
  55. data/lib/woods/watch/installation/layout.rb +70 -0
  56. data/lib/woods/watch/installation/options.rb +128 -0
  57. data/lib/woods/watch/installation/planner.rb +128 -0
  58. data/lib/woods/watch/installation/probe.rb +101 -0
  59. data/lib/woods/watch/installation/receipt.rb +77 -0
  60. data/lib/woods/watch/installation/recovery.rb +64 -0
  61. data/lib/woods/watch/installation/templates.rb +58 -0
  62. data/lib/woods/watch/installation.rb +56 -0
  63. data/lib/woods/watch/lifecycle.rb +182 -0
  64. data/lib/woods/watch/managed_child.rb +113 -0
  65. data/lib/woods/watch/managed_cleanup.rb +48 -0
  66. data/lib/woods/watch/managed_process.rb +144 -0
  67. data/lib/woods/watch/puma_adapter.rb +87 -0
  68. data/lib/woods/watch/puma_child.rb +66 -0
  69. data/lib/woods/watch/supervision_records.rb +95 -0
  70. data/lib/woods/watch/supervision_status.rb +104 -0
  71. data/lib/woods/watch/supervisor.rb +161 -0
  72. data/lib/woods/watch/supervisor_reporting.rb +46 -0
  73. data/plugin/.claude-plugin/plugin.json +1 -1
  74. data/plugin/skills/woods-agent-enable/SKILL.md +1 -1
  75. data/plugin/skills/woods-diagnose/SKILL.md +77 -8
  76. data/plugin/skills/woods-investigate/SKILL.md +6 -6
  77. data/plugin/skills/woods-mcp-config/SKILL.md +28 -1
  78. data/plugin/skills/woods-setup/SKILL.md +58 -4
  79. metadata +35 -5
@@ -291,16 +291,65 @@ result = retriever.retrieve("what validations does Order have?")
291
291
 
292
292
  ---
293
293
 
294
- ## Degradation Tiers
295
-
296
- Retrieval degrades gracefully when components are unavailable. The Retriever itself does not implement explicit fallback tiers, degradation happens naturally through how each component handles errors:
294
+ ## Semantic corpus diagnostics
295
+
296
+ Extraction and embedding publish different data. `woods_status.ready` describes
297
+ the structural index; neither that flag nor bootstrap `hydrated` proves that
298
+ semantic retrieval has indexed records. Provider detection alone can succeed
299
+ before the first embedding run.
300
+
301
+ Supporting readers report `woods_status.retriever.corpus`:
302
+
303
+ - `state`: `empty`, `metadata_only`, `vectors_only`, `nonempty`, or `unknown`.
304
+ - `vectors` and `metadata`: each has `count`, `by_type`, and `untyped_count`.
305
+ Counts describe stored entries, including chunks; they are not distinct
306
+ extracted-unit counts or a completeness certificate. Missing type labels are
307
+ reported separately rather than assigned to a guessed type.
308
+ - Unknown counts are `null`. Diagnostics use an explicit local-store capability;
309
+ they do not query remote stores to discover their counts. A missing capability
310
+ does not mean the backend is empty or broken.
311
+
312
+ When both semantic stores are known empty, `codebase_retrieve` reports
313
+ `empty_index` with recovery guidance instead of presenting an empty match as
314
+ evidence that application code is absent. Run `woods:embed` in the application
315
+ with the intended provider and storage configuration, then reload or restart
316
+ the reader. Restart after changing provider or store configuration; reload
317
+ refreshes the stores of the existing retriever. Alternatively, explicitly select `WOODS_RETRIEVAL_MODE=lexical`
318
+ in the MCP process and restart for ranked retrieval over the published units.
319
+ Woods does not change retrieval mode automatically.
320
+
321
+ Metadata-only stores can still answer some keyword, direct, and graph queries;
322
+ they are not blocked by this diagnostic. Source-empty units deliberately retain
323
+ metadata without vectors, so the two counts need not match. Nonempty stores can
324
+ still have missing types, stale vectors, or provider failures. Inspect the
325
+ query result and embedding evidence before claiming coverage. The type-rank
326
+ table's metadata count describes the retrieval metadata store, not all units
327
+ in the structural index.
328
+
329
+ Older readers may omit `retriever.corpus`; record the reader revision separately
330
+ from the index writer version and verify the embedding artifacts directly.
331
+ Explicit lexical mode does not use or report semantic corpus counts.
297
332
 
298
- - **Embedding provider unavailable**: `codebase_retrieve` returns a structured configuration error. Check `woods_status` for retrieval readiness.
299
- - **Vector store unavailable**: vector and hybrid strategies fail at query time. Keyword and graph strategies remain available for direct calls to `SearchExecutor`.
300
- - **Metadata store error**: the structural context overview (unit counts by type) is silently omitted; `Retriever#build_structural_context` rescues `StandardError` and returns `nil`. The retrieval result is still returned without the overview.
301
- - **Graph store unavailable**: graph expansion in hybrid strategy produces no graph candidates; vector and keyword candidates are still ranked and returned.
333
+ ## Degradation Tiers
302
334
 
303
- In all cases, errors in individual components produce empty candidate sets for that source rather than raising through the `Retriever`. Configure circuit breakers via `Woods::Resilience::CircuitBreaker` on external providers (Qdrant, OpenAI) for production deployments.
335
+ The MCP boundary distinguishes missing configuration, empty stores, and failed
336
+ stores. A failure is not evidence that no application code matches:
337
+
338
+ - **No embedding provider configured:** `codebase_retrieve` reports a configuration
339
+ error with embedding and explicit lexical-mode options.
340
+ - **Both semantic stores known empty:** the tool reports `empty_index`; see the
341
+ [corpus diagnostics](#semantic-corpus-diagnostics) above.
342
+ - **Failed dump hydration:** the tool reports `degraded_index` rather than serving
343
+ a clean empty result. Repair or regenerate the named embedding artifact.
344
+ - **Query-time storage failures:** vector, metadata, and graph adapter exceptions
345
+ are translated into store errors and reported as `degraded_index` by MCP.
346
+ This includes failures building the metadata overview; that failure is not
347
+ silently omitted. Direct Ruby callers should handle `Woods::Retriever::StoreError`.
348
+
349
+ Provider failures and missing metadata for returned candidates have their own
350
+ error paths. Preserve their diagnostics; do not silently switch modes or treat
351
+ an exception as an empty match. Positive corpus counts do not override these
352
+ checks.
304
353
 
305
354
  ---
306
355
 
@@ -4,7 +4,7 @@
4
4
  application inputs with source bytes visible to the reader. It is separate from
5
5
  index age, HEAD equality, daemon liveness and external database/runtime state.
6
6
 
7
- This capability is unreleased after `2.0.0.beta2`. Check the installed gem's
7
+ This capability is included in Woods `2.0.0`. Check the installed gem's
8
8
  `woods-extract --help`, `rake -T woods:source_status`, and `woods_status` schema
9
9
  before using it; upgrading the plugin alone does not upgrade Woods.
10
10
 
@@ -13,7 +13,8 @@
13
13
  > When the optional [`tokenizers`](https://github.com/ankane/tokenizers-ruby)
14
14
  > gem is installed, the Ollama path uses the real BERT WordPiece tokenizer
15
15
  > (`Woods::Embedding::TokenCounter`) instead of this heuristic. The 4.0
16
- > divisor below is what the gem falls back to everywhere else.
16
+ > divisor applies to the OpenAI/default path; Ollama falls back to 1.5 when
17
+ > its tokenizer is unavailable.
17
18
 
18
19
  This is a historical record of the benchmark that picked 4.0 over the
19
20
  original 3.5 divisor. It is cited from five places in `lib/` as the evidence
@@ -35,17 +36,21 @@ for that choice, keep the numbers below intact if you edit this doc.
35
36
  | 3.8 | 16.2% | 42.5% |
36
37
  | **4.0 (shipped)** | **10.6%** | **35.4%** |
37
38
 
38
- Mean chars/token across the corpus was **4.41** (range 3.94–5.42). The
39
- heuristic always overestimated, never underestimated, across all 19 files,
40
- which is what makes it safe for token-limit enforcement even at its worst
41
- case. Code lines and comment/YARD lines had similar ratios (4.38 vs. 4.27
42
- chars/token), no separate handling needed for either.
39
+ Mean chars/token across the corpus was **4.41** (range 3.94–5.42). These
40
+ aggregate results favor the 4.0 divisor for this sample; they do not establish
41
+ an upper bound on token counts. The recorded range includes values below 4.0,
42
+ so the heuristic can underestimate. Code lines and comment/YARD lines had
43
+ similar ratios (4.38 vs. 4.27 chars/token) in this sample.
44
+
45
+ The character estimate covers only the text passed to the counter. It does not
46
+ bound the serialized MCP response: text rendering, structured output, provenance,
47
+ and JSON framing can add bytes and tokens beyond that input.
43
48
 
44
49
  ## What shipped
45
50
 
46
51
  **The divisor changed from 3.5 to 4.0.** It roughly halves the mean
47
- overestimate (26.2% → 10.6%) while keeping the conservative
48
- always-overestimates property, at zero new runtime dependencies. The
52
+ error (26.2% → 10.6%) at zero new runtime dependencies. It remains an estimate,
53
+ not a guarantee that arbitrary input fits a model's token limit. The
49
54
  constant lives in one place now (`Woods::TokenUtils::CHARS_PER_TOKEN_BY_PROVIDER`),
50
55
  not scattered across call sites, see `lib/woods/token_utils.rb` for the
51
56
  current definition and `docs/EMBEDDING_MODELS.md` for the Ollama-side ratio.
@@ -53,8 +58,9 @@ current definition and `docs/EMBEDDING_MODELS.md` for the Ollama-side ratio.
53
58
  **tiktoken_ruby was deliberately not added as a runtime dependency.** A 10.6%
54
59
  mean error is acceptable for chunking decisions, budget estimates, and
55
60
  truncation; a native-extension dependency for marginal accuracy gains wasn't
56
- worth it. The optional `tokenizers` gem covers the case where exact counts
57
- matter more (see above).
61
+ worth it. The optional `tokenizers` gem provides counts for its supported BERT
62
+ WordPiece tokenizer, not every model. Strict token-limit enforcement requires
63
+ the tokenizer used by the target model.
58
64
 
59
65
  ## Reproducing this benchmark
60
66
 
@@ -23,7 +23,7 @@ This guide covers the most common problems encountered when installing, extracti
23
23
  | `No such container` | Wrong container name | Check with `docker ps --format '{{.Names}}'` |
24
24
  | `JSON parse errors` (MCP) | Rails boot noise on stdout | Remove `puts` calls from initializers |
25
25
  | Query timeout | Large table, no scope | Add scope conditions to narrow results |
26
- | `Extraction failed for …; the previous generation remains active` | A consumer handled a source error during incremental extraction or refresh (unreleased after `2.0.0.beta3`) | Fix the logged source error and retry the [complete batch](INCREMENTAL_EXTRACTION.md#handled-source-errors-and-retry); watch keeps it pending |
26
+ | `Extraction failed for …; the previous generation remains active` | A consumer handled a source error during incremental extraction or refresh (included in Woods `2.0.0`) | Fix the logged source error and retry the [complete batch](INCREMENTAL_EXTRACTION.md#handled-source-errors-and-retry); watch keeps it pending |
27
27
  | Empty extraction output | `eager_load!` failure | Check for `NameError` in boot output |
28
28
  | Git metadata missing | Shallow clone in CI | Use `fetch-depth: 0` for complete history |
29
29
  | Parallel tool calls all fail | MCP client batches calls | Send calls sequentially, validate params first |
@@ -53,13 +53,44 @@ version. See [manifest writer provenance](PUBLISHED_INDEX.md#manifest-writer-pro
53
53
 
54
54
  If a tool call fails with **"Tool not found: … not available in the installed Woods v…"**, the client is asking for a tool a newer gem provides. Run `bundle update woods` and reconnect the MCP server, then retry.
55
55
 
56
+ ### Watcher startup or planned restart fails
57
+
58
+ Managed `woods-watch` startup is **included in Woods `2.0.0`**; record the
59
+ loaded version/path and revision, then verify executable and generator help.
60
+ If changing an initializer stops every Foreman process, replace a bare
61
+ `woods:watch` entry with the [managed setup](WATCH_DAEMON.md#managed-development-startup).
62
+
63
+ Read launcher logs and `woods_status` supervision records separately from daemon
64
+ liveness and index freshness. `retrying` means the last generation remains usable
65
+ while boot is retried. A parked ownership/protocol conflict requires correcting
66
+ the selected owner or installed command and restarting that owner; do not delete
67
+ claim files or kill PIDs taken from status. No index-visible record exists before
68
+ the first boot resolves the application's output directory.
69
+
70
+ Unset `WOODS_WATCH_IDLE_TIMEOUT` in managed modes. If the boot deadline is reached,
71
+ diagnose Bundler/initializer startup before increasing `--boot-timeout`; a valid
72
+ long extraction has a separate readiness state and is not bounded by that clock.
73
+ If setup created a Procfile but normal `bin/dev` still only launches Rails, choose
74
+ Puma or explicitly run the selected Foreman command. The generator never rewrites
75
+ `bin/dev` or starts services during preview.
76
+
77
+ If installation reports a pending transaction, use `woods:watch --operation
78
+ recover` through the Rails generator, initially with `--pretend`; see
79
+ [owned setup recovery](WATCH_DAEMON.md#ownership-updates-and-removal). That Rails
80
+ command boots the application first. For broken initializers use the documented
81
+ direct bundled Ruby helper, which does not boot Rails or require task discovery.
82
+ Both refuse to overwrite intervening edits. A Puma setup refusal for
83
+ `config/puma/development.rb` means the default
84
+ configuration would bypass the generated plugin; select an external/Foreman
85
+ arrangement instead of installing an inactive directive.
86
+
56
87
  ### Semantic graph validation errors
57
88
 
58
89
  In development versions containing #413, `woods:validate` rejects graphs that
59
90
  parse as JSON but disagree with their indexes. Errors name the section and
60
91
  identity, for example `reverse["http_api"]: missing "Order"`, a duplicate typed
61
- variant, or an indexed unit absent from `nodes`. This is unreleased after
62
- `2.0.0.beta2`; check the installed gem before expecting these diagnostics.
92
+ variant, or an indexed unit absent from `nodes`. This is included in Woods
93
+ `2.0.0`; check the installed gem before expecting these diagnostics.
63
94
 
64
95
  Keep the failing generation and report the exact errors. Run a full extraction
65
96
  in a fresh application process with the intended bundle, then validate again.
@@ -88,7 +119,7 @@ Scoped resets leave corrupt state untouched. Valid state keeps any unrelated
88
119
  operation entries, and missing state remains a no-op without creating a file.
89
120
  A permission failure must be corrected before repair can succeed.
90
121
 
91
- This recovery is unreleased after `2.0.0.beta2`; check the installed version.
122
+ This recovery is included in Woods `2.0.0`; check the installed version.
92
123
  Older versions report corrupt state as nothing to repair. Stop pipeline writers,
93
124
  back up the configured guard state's `pipeline_guard.json`, and remove only that
94
125
  file before restarting, or upgrade to a version containing the fix.
@@ -160,7 +191,12 @@ For subsequent runs, use incremental mode instead of full extraction:
160
191
  bundle exec rake woods:incremental
161
192
  ```
162
193
 
163
- Incremental extraction only re-extracts files that changed since the last run. It skips unchanged units and is typically 5-10× faster.
194
+ Incremental extraction dispatches the selected changed paths, including affected
195
+ concern consumers and whole-app extractors whose trigger paths changed. The default
196
+ Git range is `HEAD~1`; pass an explicit range or `CHANGED_FILES` for other batches.
197
+ It can reduce extraction work, but Rails boot, graph rebuilding, and publication
198
+ still contribute to runtime. Measure the improvement in your application; Woods
199
+ does not guarantee a speedup. See the [incremental contract](INCREMENTAL_EXTRACTION.md).
164
200
 
165
201
  ---
166
202
 
@@ -228,7 +264,7 @@ dependents after an incremental run.
228
264
  identities when restoring and updating the graph.
229
265
 
230
266
  **Fix:** Check whether the installed version includes B-193; this fix is
231
- unreleased. After upgrading to a version containing the fix, run
267
+ included in Woods `2.0.0`. After upgrading to a version containing the fix, run
232
268
  `bundle exec rake woods:extract` once to rebuild lost reverse dependencies.
233
269
  Loading an already damaged graph does not restore discarded entries. See the
234
270
  [incremental graph contract](INCREMENTAL_EXTRACTION.md#the-contract).
@@ -239,7 +275,7 @@ Loading an already damaged graph does not restore discarded entries. See the
239
275
  most files as `change_frequency: new` in a shallow CI checkout.
240
276
 
241
277
  **Cause:** A shallow clone truncates HEAD ancestry. The shallow-checkout guard is
242
- unreleased after 2.0.0.beta2: current source omits git enrichment and warns once,
278
+ included in Woods `2.0.0`: Woods omits git enrichment and warns once,
243
279
  rather than treating the truncated history as complete. If repository depth
244
280
  cannot be verified, enrichment is also omitted; check git access and version.
245
281
 
@@ -256,12 +292,33 @@ clone), then run full extraction to replace retained metadata:
256
292
  Two commits can suffice for an incremental diff, but do not establish the full
257
293
  ancestry needed for churn metadata.
258
294
 
295
+ ### Git executable is missing from the extraction environment
296
+
297
+ **Symptom:** Extraction logs `Git history unavailable: git executable was not
298
+ found in PATH`, and newly extracted units have no `metadata.git`. Older builds
299
+ can omit this enrichment silently when the executable is missing.
300
+
301
+ **Fix:** Run `git --version` in the same container and environment that runs
302
+ extraction. Install Git 2.31 or newer there, ensure its executable is on `PATH`,
303
+ then run full `woods:extract` to refresh every unit's history. A working Git
304
+ installation on the host does not provide Git inside an application container.
305
+
306
+ Extraction continues without inventing zero-commit history. The warning appears
307
+ once per extractor instance when the application has a `.git` entry or an
308
+ explicit `WOODS_GIT_DIR`/`GIT_DIR` setting. A source archive with neither remains
309
+ supported and quiet. `GIT_BRANCH`/`GIT_SHA` provenance fallback is unchanged;
310
+ those values identify a build but cannot supply per-file history.
311
+
312
+ This diagnostic is emitted during extraction. `woods_status.ready` and a
313
+ manifest Git SHA do not establish that per-unit history was available, and
314
+ `recent_changes` returning no results does not prove no files changed.
315
+
259
316
  ---
260
317
 
261
318
  ### Git enrichment warns that history could not be read completely
262
319
 
263
- Current source uses an explicit merge-diff mode requiring **Git 2.31 or newer**.
264
- This is unreleased after 2.0.0.beta2: first confirm the installed Woods version.
320
+ Woods 2.0 uses an explicit merge-diff mode requiring **Git 2.31 or newer**.
321
+ First confirm the installed Woods version.
265
322
  Check `git --version` inside the same container/process environment as extraction,
266
323
  and upgrade git if it is older. On a supported version, check that the application's
267
324
  `HEAD` and object store can be read using the same `WOODS_GIT_DIR` setting.
@@ -293,29 +350,62 @@ When it does not, the git keys are omitted from every unit, provenance records
293
350
  `"unknown"`, and one warning names git's own reason. Absent keys mean "not
294
351
  known"; they never mean "brand new".
295
352
 
296
- **Fix:** Point `WOODS_GIT_DIR` at the *canonical* git directory, the one the
297
- worktree's `gitdir:` pointer ultimately leads to, and make sure it is mounted:
353
+ **Fix:** Restore access to both the worktree-specific Git directory and the
354
+ shared objects and refs using the mount layouts below, then run a full
355
+ `woods:extract` to replace retained metadata.
356
+
357
+ ### Git directory mounts for linked worktrees
358
+
359
+ `WOODS_GIT_DIR` is passed directly to Git's `--git-dir`. It selects that
360
+ directory's `HEAD` for manifest provenance, per-file history, and incremental
361
+ diff ranges. **For a linked worktree, pointing it at the shared `.git` root
362
+ selects the primary checkout's HEAD.** A successful Git command alone does
363
+ not prove Woods is reading the intended branch.
364
+
365
+ First inspect Git metadata on the host, from the intended worktree:
298
366
 
299
367
  ```bash
300
- # docker-compose.yml, mounting the parent repository's git directory
301
- # volumes:
302
- # - /path/to/repo/.git:/canonical-git:ro
303
- WOODS_GIT_DIR=/canonical-git bundle exec rake woods:extract
368
+ git -C /path/to/worktree rev-parse --absolute-git-dir
369
+ # Example: /path/to/repo/.git/worktrees/wt
370
+ git -C /path/to/worktree rev-parse --path-format=absolute --git-common-dir
371
+ # Example: /path/to/repo/.git
372
+ git -C /path/to/worktree rev-parse --abbrev-ref HEAD
373
+ git -C /path/to/worktree rev-parse HEAD
304
374
  ```
305
375
 
306
- `WOODS_GIT_DIR` wins over whatever the worktree pointer says, and applies to
307
- every git call Woods makes: per-unit enrichment, `manifest.json` provenance,
308
- and the diff range `woods:incremental` resolves. All three run through
309
- `Woods::GitCommand.argv`, so the override cannot reach two of them and miss the
310
- third.
376
+ The worktree ID in this example is `wt`. Use the ID returned by Git metadata;
377
+ it need not match the branch name. Choose one of these layouts:
378
+
379
+ - **Same-path mount:** mount the complete shared directory read-only at its
380
+ original absolute path (`/path/to/repo/.git:/path/to/repo/.git:ro`). With the
381
+ application's existing `.git` pointer resolvable, leave `WOODS_GIT_DIR`
382
+ unset and remove conflicting Git-directory overrides from the environment.
383
+ - **Relocated mount:** mount that complete directory read-only at a new path
384
+ (`/path/to/repo/.git:/mounted-common:ro`), including `objects`, `refs`, and
385
+ `worktrees`. Select the worktree-specific directory inside it:
386
+
387
+ ```bash
388
+ WOODS_GIT_DIR=/mounted-common/worktrees/wt bundle exec rake woods:extract
389
+ ```
390
+
391
+ Mounting only the private worktree directory can leave its `commondir` pointer
392
+ without access to shared objects and refs. Git's own environment variables are
393
+ inherited by the subprocess; check any existing `GIT_DIR` and `GIT_COMMON_DIR`
394
+ settings when diagnosing the effective layout. The complete layouts above
395
+ preserve both worktree identity and shared storage.
396
+
397
+ In the extraction container, verify the relocated selection against the host
398
+ branch and exact SHA before extracting (replace `/app` and `wt` as needed):
399
+
400
+ ```bash
401
+ git --git-dir=/mounted-common/worktrees/wt --work-tree=/app -C /app rev-parse --abbrev-ref HEAD
402
+ git --git-dir=/mounted-common/worktrees/wt --work-tree=/app -C /app rev-parse HEAD
403
+ ```
311
404
 
312
- **`GIT_DIR` alone is not enough for a linked worktree.** Woods honors git's own
313
- `GIT_DIR` and `GIT_COMMON_DIR` because git does, but setting `GIT_DIR` to a
314
- worktree's private git directory only moves the failure: the `commondir`
315
- pointer inside it is relative, so it still resolves to a path that is not
316
- mounted, and `GIT_COMMON_DIR` does not override it. Either mount the canonical
317
- git directory at the same absolute path the pointer names, or use
318
- `WOODS_GIT_DIR`.
405
+ After fixing the selection, run full `woods:extract` and verify the published
406
+ manifest. Incremental extraction can retain older per-file Git metadata.
407
+ A commit alone does not necessarily trigger the source-file watcher; run a
408
+ full extraction when current history and provenance are required.
319
409
 
320
410
  ---
321
411
 
@@ -375,17 +465,23 @@ retries after a later filesystem event.
375
465
 
376
466
  ### `manifest.json` shows the wrong branch (or `git_branch: "unknown"`) in a worktree
377
467
 
378
- **Symptom:** `git_branch` / `git_sha` in `manifest.json` name a different branch than the worktree is actually on, or report `"unknown"`. The extracted units themselves are correct, only the provenance metadata is off.
468
+ **Symptom:** `git_branch` / `git_sha` in `manifest.json` name a different
469
+ branch or SHA than the intended worktree, or report `"unknown"`.
379
470
 
380
- **Cause:** In a linked git worktree, `.git` is a *file* containing a `gitdir:` pointer to the real git directory, often an absolute host path. When extraction runs where that path can't be resolved (e.g. inside a container where the host path isn't mounted), git can't read the ref. Woods now reports `"unknown"` in that case rather than emitting a stale, misleading value (previously it fell back to a baked `GIT_BRANCH`/`GIT_SHA` build arg).
471
+ **Cause:** An unreachable `.git` file's `gitdir:` pointer prevents Git from
472
+ resolving the worktree's HEAD. An override selecting the shared `.git` root
473
+ instead resolves the primary checkout's HEAD successfully. That wrong selection
474
+ also affects per-file history and HEAD-based incremental ranges.
381
475
 
382
- **Fix:** Make the worktree's git directory reachable from the extraction environment, for example, mount the parent repository (the directory the `gitdir:` pointer references) into the container, or run extraction from a normal (non-worktree) checkout. With the real git directory reachable, `git_branch`/`git_sha` resolve correctly. If the checkout legitimately ships without a `.git` at all (a source tarball, or a Docker `COPY` that excludes it), set `GIT_BRANCH` / `GIT_SHA` explicitly. Woods honors these when there is no `.git` at the root (or no git binary), but suppresses them when a `.git` *is* present but unresolvable (so a stale build arg can't mask a worktree).
476
+ **Fix:** Follow [Git directory mounts for linked worktrees](#git-directory-mounts-for-linked-worktrees),
477
+ compare the selected branch and exact SHA in the extraction environment, then
478
+ run a full extraction. Compare the newly published manifest, not a retained
479
+ generation. A commit without a source edit may leave the watcher idle.
383
480
 
384
- When the canonical git directory is mounted but not at the path the pointer
385
- names, set `WOODS_GIT_DIR` to where it actually is. It wins over the pointer for
386
- provenance and for unit-level git metadata both. Setting git's own `GIT_DIR` to
387
- the worktree's private git directory does not work: its `commondir` pointer is
388
- relative and resolves outside the mount.
481
+ For a checkout legitimately shipped without `.git` (such as a source tarball),
482
+ `GIT_BRANCH` / `GIT_SHA` can supply provenance. They are fallbacks only when
483
+ `.git` is absent or Git is unavailable; a present but unresolvable `.git`
484
+ reports `"unknown"` instead of substituting stale build arguments.
389
485
 
390
486
  ---
391
487
 
@@ -395,9 +491,9 @@ relative and resolves outside the mount.
395
491
 
396
492
  ### Index cannot be resolved at startup
397
493
 
398
- **Symptom:** An Index MCP executable exits with `Could not resolve a published Woods index in: /path/to/...` even though extraction completed. This headline is unreleased after `2.0.0.beta3`; older versions say `No manifest.json found`. Both mean the selected index could not resolve its manifest, not that an atomic index needs a root manifest.
494
+ **Symptom:** An Index MCP executable exits with `Could not resolve a published Woods index in: /path/to/...` even though extraction completed. This headline is included in Woods `2.0.0`; older versions say `No manifest.json found`. Both mean the selected index could not resolve its manifest, not that an atomic index needs a root manifest.
399
495
 
400
- Embedded Index MCP startup through `IndexReader` also raises an `ArgumentError` with the selected directory and layout guidance when the marker cannot resolve a manifest, including malformed marker shapes such as `[]` or a numeric `payload` (unreleased after `2.0.0.beta3`). Earlier builds may expose a raw `TypeError` or `NoMethodError` for those shapes. Inspect the marker and preserve the failing index before attempting recovery.
496
+ Embedded Index MCP startup through `IndexReader` also raises an `ArgumentError` with the selected directory and layout guidance when the marker cannot resolve a manifest, including malformed marker shapes such as `[]` or a numeric `payload` (included in Woods `2.0.0`). Earlier builds may expose a raw `TypeError` or `NoMethodError` for those shapes. Inspect the marker and preserve the failing index before attempting recovery.
401
497
 
402
498
  **Cause:** The selected directory is not the published index root, the published generation cannot be resolved, or the path is not visible to the MCP process. A container path is appropriate for a container process; a host process needs the host-visible path.
403
499
 
@@ -2,12 +2,10 @@
2
2
 
3
3
  Woods 2.0 changes observable index identifiers, publication layout, vector-store reconciliation, and the supported MCP surface. Plan a clean re-index. Do not upgrade a shared or durable index in place without a backup and a rollback window.
4
4
 
5
- This guide assumes the last v1 release, 1.6.1, and targets 2.0.0.
5
+ This guide covers the supported 1.6.x line and targets 2.0.0. Use the latest
6
+ published 1.6.x security patch as the rollback version.
6
7
 
7
8
  <!-- release-state:upgrade-availability -->
8
- > This tree declares 2.0.0.beta4 as a prerelease. After RubyGems lists it, pin it with
9
- > `gem "woods", "2.0.0.beta4"`; `~> 2.0` resolves only once
10
- > 2.0.0 is published.
11
9
  <!-- release-state:end -->
12
10
 
13
11
  ## Upgrade outcome
@@ -95,7 +93,11 @@ Also back up managed Obsidian/Unblocked destinations before allowing a mass stal
95
93
 
96
94
  ### 3. Choose a rollback point
97
95
 
98
- Keep the v1 Gemfile/lockfile commit and all durable-store backups until v2 extraction, MCP calls, retrieval, and exports are verified. Downgrading the gem does not translate v2 identifiers back to v1.
96
+ Record and test a Gemfile/lockfile selecting the latest published 1.6.x security
97
+ patch as the rollback bundle. If the current installation is older, verify that
98
+ patched v1 bundle before beginning the v2 migration. Keep its commit and all
99
+ durable-store backups until v2 extraction, MCP calls, retrieval, and exports are
100
+ verified. Downgrading the gem does not translate v2 identifiers back to v1.
99
101
 
100
102
  ## Upgrade the application
101
103
 
@@ -142,7 +144,7 @@ bin/rails woods:validate
142
144
  bin/rails woods:stats
143
145
  ```
144
146
 
145
- Unreleased after `2.0.0.beta3`: `woods:clean` removes index artifacts but keeps
147
+ Included in Woods `2.0.0`: `woods:clean` removes index artifacts but keeps
146
148
  the output directory and its hidden extraction guard. This stable guard lets
147
149
  concurrent writers coordinate safely; its presence does not mean an index remains.
148
150
 
@@ -331,7 +333,7 @@ Complete every applicable check:
331
333
  If verification fails:
332
334
 
333
335
  1. stop v2 MCP, watcher, embedding, and exporter processes;
334
- 2. restore the v1 Gemfile and lockfile or deploy the recorded v1 commit;
336
+ 2. restore the tested, patched v1 Gemfile and lockfile or deploy its recorded commit;
335
337
  3. run the v1 `woods:clean` before restoring anything under the configured output directory;
336
338
  4. either restore the complete pre-upgrade v1 output-directory backup, or run a fresh v1 extraction and then restore its v1 `dumps/` and configuration artifacts;
337
339
  5. restore external vector-store and managed export backups when v2 modified them;