woods 2.0.0.beta3 → 2.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +500 -420
- data/CONTRIBUTING.md +29 -17
- data/README.md +78 -178
- data/docs/AGENT_GUIDE.md +52 -11
- data/docs/AGENT_SETUP.md +34 -17
- data/docs/AUTOMATIC_MAINTENANCE.md +222 -0
- data/docs/BACKEND_MATRIX.md +18 -7
- data/docs/CLIENT_HOOKS.md +1 -1
- data/docs/CONFIGURATION_REFERENCE.md +105 -29
- data/docs/CONSOLE_MCP_SETUP.md +54 -9
- data/docs/DOCKER_SETUP.md +16 -1
- data/docs/EVALUATION.md +10 -4
- data/docs/EXTRACTOR_REFERENCE.md +23 -3
- data/docs/FAQ.md +14 -3
- data/docs/GETTING_STARTED.md +18 -17
- data/docs/INCREMENTAL_EXTRACTION.md +37 -8
- data/docs/INDEX_LAYOUT.md +2 -2
- data/docs/MCP_SERVERS.md +79 -7
- data/docs/MCP_TOOL_COOKBOOK.md +5 -5
- data/docs/MCP_WORKTREE_SETUP.md +55 -83
- data/docs/PUBLISHED_INDEX.md +17 -0
- data/docs/README.md +2 -1
- data/docs/RETRIEVAL_GUIDE.md +81 -13
- data/docs/SOURCE_FRESHNESS.md +1 -1
- data/docs/TOKEN_BENCHMARK.md +16 -10
- data/docs/TROUBLESHOOTING.md +142 -47
- data/docs/UPGRADING_TO_2.md +12 -6
- data/docs/WATCH_DAEMON.md +189 -24
- data/docs/WHY_WOODS.md +9 -5
- data/exe/woods-console +13 -11
- data/exe/woods-mcp-start +14 -9
- data/exe/woods-watch +5 -0
- data/lib/generators/woods/pgvector_generator.rb +8 -2
- data/lib/generators/woods/watch_generator.rb +53 -0
- data/lib/puma/plugin/woods.rb +10 -0
- data/lib/tasks/woods.rake +14 -0
- data/lib/woods/agent_configuration/applier.rb +5 -3
- data/lib/woods/agent_configuration/cli.rb +2 -2
- data/lib/woods/agent_configuration/layout.rb +13 -0
- data/lib/woods/cache/cache_middleware.rb +6 -0
- data/lib/woods/console/credential_scanner.rb +4 -3
- data/lib/woods/console/dispatch_pipeline.rb +7 -0
- data/lib/woods/console/embedded_executor.rb +31 -9
- data/lib/woods/console/sql_noise_stripper.rb +9 -7
- data/lib/woods/console/sql_table_scanner.rb +47 -7
- data/lib/woods/console/sql_validator.rb +49 -9
- data/lib/woods/console/sqlite_read_guard.rb +46 -0
- data/lib/woods/console/stdio_transport.rb +27 -0
- data/lib/woods/coordination/pipeline_lock.rb +3 -2
- data/lib/woods/embedding/indexer.rb +24 -14
- data/lib/woods/extractor.rb +70 -19
- data/lib/woods/extractors/declared_parent.rb +55 -0
- data/lib/woods/extractors/graphql_extractor.rb +2 -11
- data/lib/woods/extractors/lib_extractor.rb +10 -8
- data/lib/woods/extractors/mailer_extractor.rb +6 -10
- data/lib/woods/extractors/model_extractor.rb +1 -15
- data/lib/woods/extractors/poro_extractor.rb +10 -8
- data/lib/woods/extractors/shared_utility_methods.rb +22 -5
- data/lib/woods/git_command.rb +6 -7
- data/lib/woods/git_provenance.rb +4 -6
- data/lib/woods/mcp/bearer_auth.rb +2 -1
- data/lib/woods/mcp/bootstrapper.rb +20 -5
- data/lib/woods/mcp/config_resolver.rb +2 -1
- data/lib/woods/mcp/index_reader.rb +11 -2
- data/lib/woods/mcp/initialization_guidance.rb +1 -1
- data/lib/woods/mcp/renderers/markdown_renderer.rb +14 -8
- data/lib/woods/mcp/renderers/plain_renderer.rb +11 -7
- data/lib/woods/mcp/server.rb +63 -37
- data/lib/woods/mcp/tool_contract.rb +1 -1
- data/lib/woods/mcp/tool_response_renderer.rb +16 -0
- data/lib/woods/mcp/traversal_evidence_text.rb +1 -1
- data/lib/woods/mcp/traversal_response.rb +22 -0
- data/lib/woods/path_dispatcher.rb +6 -5
- data/lib/woods/published_index/typed_unit_reader.rb +40 -3
- data/lib/woods/published_index.rb +2 -2
- data/lib/woods/rake_helpers.rb +2 -12
- data/lib/woods/retrieval/corpus_status.rb +46 -0
- data/lib/woods/retrieval/lexical_assembler.rb +14 -3
- data/lib/woods/retrieval/lexical_index.rb +2 -1
- data/lib/woods/retriever.rb +19 -7
- data/lib/woods/session_tracer/file_store.rb +6 -1
- data/lib/woods/source_inputs/consumer_errors.rb +4 -0
- data/lib/woods/storage/local_corpus_stats.rb +32 -0
- data/lib/woods/storage/metadata_store.rb +20 -0
- data/lib/woods/storage/pgvector.rb +6 -2
- data/lib/woods/storage/vector_store.rb +10 -0
- data/lib/woods/temporal/json_snapshot_store.rb +35 -7
- data/lib/woods/version.rb +1 -1
- data/lib/woods/watch/child_environment.rb +30 -0
- data/lib/woods/watch/cli.rb +91 -0
- data/lib/woods/watch/daemon.rb +73 -11
- data/lib/woods/watch/event_stream.rb +70 -0
- data/lib/woods/watch/guardian.rb +142 -0
- data/lib/woods/watch/installation/layout.rb +70 -0
- data/lib/woods/watch/installation/options.rb +128 -0
- data/lib/woods/watch/installation/planner.rb +128 -0
- data/lib/woods/watch/installation/probe.rb +101 -0
- data/lib/woods/watch/installation/receipt.rb +77 -0
- data/lib/woods/watch/installation/recovery.rb +64 -0
- data/lib/woods/watch/installation/templates.rb +58 -0
- data/lib/woods/watch/installation.rb +56 -0
- data/lib/woods/watch/lifecycle.rb +182 -0
- data/lib/woods/watch/managed_child.rb +113 -0
- data/lib/woods/watch/managed_cleanup.rb +48 -0
- data/lib/woods/watch/managed_process.rb +144 -0
- data/lib/woods/watch/puma_adapter.rb +87 -0
- data/lib/woods/watch/puma_child.rb +66 -0
- data/lib/woods/watch/supervision_records.rb +95 -0
- data/lib/woods/watch/supervision_status.rb +104 -0
- data/lib/woods/watch/supervisor.rb +161 -0
- data/lib/woods/watch/supervisor_reporting.rb +46 -0
- data/plugin/.claude-plugin/plugin.json +1 -1
- data/plugin/hooks/woods-input-rules.sh +4 -4
- data/plugin/skills/woods-agent-enable/SKILL.md +7 -1
- data/plugin/skills/woods-diagnose/SKILL.md +134 -34
- data/plugin/skills/woods-investigate/SKILL.md +54 -15
- data/plugin/skills/woods-mcp-config/SKILL.md +38 -11
- data/plugin/skills/woods-setup/SKILL.md +72 -15
- metadata +38 -5
data/docs/RETRIEVAL_GUIDE.md
CHANGED
|
@@ -63,7 +63,7 @@ Keyword results are scored by how many distinct fields matched (identifier, sour
|
|
|
63
63
|
|
|
64
64
|
## Configuring Retrieval
|
|
65
65
|
|
|
66
|
-
|
|
66
|
+
Semantic retrieval requires an embedding provider and a vector store. Configure these in `config/initializers/woods.rb` before embedding. For provider-free ranked retrieval over extraction output, use [explicit lexical mode](#embedding-free-lexical-retrieval) in the MCP process environment instead.
|
|
67
67
|
|
|
68
68
|
### Presets (recommended)
|
|
69
69
|
|
|
@@ -180,7 +180,12 @@ WOODS_RETRIEVAL_MODE=lexical bundle exec woods-mcp-start ./tmp/woods
|
|
|
180
180
|
For a Ruby-built retriever, set `config.retrieval_mode = :lexical` and supply a
|
|
181
181
|
populated metadata store to `Builder#build_retriever`. The packaged MCP server
|
|
182
182
|
loads published unit JSON itself; it does not boot Rails or read current source
|
|
183
|
-
files. Neither path constructs an embedding provider or vector adapter.
|
|
183
|
+
files. Neither path constructs an embedding provider or vector adapter. If a
|
|
184
|
+
no-provider error suggests only OpenAI, Ollama or `search`, explicit lexical
|
|
185
|
+
mode is still available from `2.0.0.beta3`: set the variable in the MCP client
|
|
186
|
+
configuration and restart that server. Confirm `woods_status.retriever.mode`
|
|
187
|
+
is `lexical`; setting it only in a Rails initializer does not configure a
|
|
188
|
+
separate MCP process. The
|
|
184
189
|
existing `:semantic` mode remains the default; provider failures never switch
|
|
185
190
|
modes automatically. A Rails initializer is not loaded by the standalone MCP
|
|
186
191
|
process, so set the environment variable in that process's client configuration.
|
|
@@ -191,15 +196,25 @@ and validations). Exact full identifiers rank first, with ambiguous typed owners
|
|
|
191
196
|
retained. Other ties are deterministic. Responses name the lexical mode and
|
|
192
197
|
matching fields/terms; runtime-field hits include the selected published runtime
|
|
193
198
|
values. Lexical Ruby results leave the semantic-only `type_rank_context` table
|
|
194
|
-
`nil`; they do not report a global vector rank or vector fallback. The top 20
|
|
195
|
-
|
|
199
|
+
`nil`; they do not report a global vector rank or vector fallback. The top 20
|
|
200
|
+
eligible positive matches form the candidate shortlist for the output budget.
|
|
201
|
+
The lexical header reports `sources included`, `candidates considered`, and
|
|
202
|
+
`candidate limit: 20`. Included sources count the actual returned entries;
|
|
203
|
+
considered candidates count the shortlist after filtering and the limit, not
|
|
204
|
+
all matches or all documents examined. This is ranked discovery, not an
|
|
205
|
+
exhaustive match listing. Explicit
|
|
196
206
|
`types` filters override default exclusions, as in semantic retrieval, and apply
|
|
197
207
|
before that limit. A query with no lexical evidence returns no matches; unrelated
|
|
198
208
|
graph hubs are never added. Query-seeded graph ranking is evaluation-only.
|
|
199
209
|
|
|
200
210
|
The budget covers headers, matching explanations and truncation notices using a
|
|
201
211
|
labelled character-based estimate, not an exact provider tokenizer. Full source
|
|
202
|
-
remains available through `lookup`.
|
|
212
|
+
remains available through `lookup`. Count text is charged before source selection;
|
|
213
|
+
final counts do not trigger a second selection pass. Very small budgets can omit
|
|
214
|
+
all sources or clip the header itself. Zero candidates means no lexical matches;
|
|
215
|
+
positive candidates with zero included sources means no source entry fit the
|
|
216
|
+
available budget. The same counts apply to full, compact, outline and scoped
|
|
217
|
+
retrieval, agreeing with returned source attribution. The
|
|
203
218
|
reader pins one published generation for building and querying its immutable
|
|
204
219
|
lexical snapshot, rebuilding after publication. Corrupt units fail explicitly;
|
|
205
220
|
they cannot quietly become a successful partial index. Older flat indexes rebuild
|
|
@@ -276,16 +291,65 @@ result = retriever.retrieve("what validations does Order have?")
|
|
|
276
291
|
|
|
277
292
|
---
|
|
278
293
|
|
|
279
|
-
##
|
|
280
|
-
|
|
281
|
-
|
|
294
|
+
## Semantic corpus diagnostics
|
|
295
|
+
|
|
296
|
+
Extraction and embedding publish different data. `woods_status.ready` describes
|
|
297
|
+
the structural index; neither that flag nor bootstrap `hydrated` proves that
|
|
298
|
+
semantic retrieval has indexed records. Provider detection alone can succeed
|
|
299
|
+
before the first embedding run.
|
|
300
|
+
|
|
301
|
+
Supporting readers report `woods_status.retriever.corpus`:
|
|
302
|
+
|
|
303
|
+
- `state`: `empty`, `metadata_only`, `vectors_only`, `nonempty`, or `unknown`.
|
|
304
|
+
- `vectors` and `metadata`: each has `count`, `by_type`, and `untyped_count`.
|
|
305
|
+
Counts describe stored entries, including chunks; they are not distinct
|
|
306
|
+
extracted-unit counts or a completeness certificate. Missing type labels are
|
|
307
|
+
reported separately rather than assigned to a guessed type.
|
|
308
|
+
- Unknown counts are `null`. Diagnostics use an explicit local-store capability;
|
|
309
|
+
they do not query remote stores to discover their counts. A missing capability
|
|
310
|
+
does not mean the backend is empty or broken.
|
|
311
|
+
|
|
312
|
+
When both semantic stores are known empty, `codebase_retrieve` reports
|
|
313
|
+
`empty_index` with recovery guidance instead of presenting an empty match as
|
|
314
|
+
evidence that application code is absent. Run `woods:embed` in the application
|
|
315
|
+
with the intended provider and storage configuration, then reload or restart
|
|
316
|
+
the reader. Restart after changing provider or store configuration; reload
|
|
317
|
+
refreshes the stores of the existing retriever. Alternatively, explicitly select `WOODS_RETRIEVAL_MODE=lexical`
|
|
318
|
+
in the MCP process and restart for ranked retrieval over the published units.
|
|
319
|
+
Woods does not change retrieval mode automatically.
|
|
320
|
+
|
|
321
|
+
Metadata-only stores can still answer some keyword, direct, and graph queries;
|
|
322
|
+
they are not blocked by this diagnostic. Source-empty units deliberately retain
|
|
323
|
+
metadata without vectors, so the two counts need not match. Nonempty stores can
|
|
324
|
+
still have missing types, stale vectors, or provider failures. Inspect the
|
|
325
|
+
query result and embedding evidence before claiming coverage. The type-rank
|
|
326
|
+
table's metadata count describes the retrieval metadata store, not all units
|
|
327
|
+
in the structural index.
|
|
328
|
+
|
|
329
|
+
Older readers may omit `retriever.corpus`; record the reader revision separately
|
|
330
|
+
from the index writer version and verify the embedding artifacts directly.
|
|
331
|
+
Explicit lexical mode does not use or report semantic corpus counts.
|
|
282
332
|
|
|
283
|
-
|
|
284
|
-
- **Vector store unavailable**: vector and hybrid strategies fail at query time. Keyword and graph strategies remain available for direct calls to `SearchExecutor`.
|
|
285
|
-
- **Metadata store error**: the structural context overview (unit counts by type) is silently omitted; `Retriever#build_structural_context` rescues `StandardError` and returns `nil`. The retrieval result is still returned without the overview.
|
|
286
|
-
- **Graph store unavailable**: graph expansion in hybrid strategy produces no graph candidates; vector and keyword candidates are still ranked and returned.
|
|
333
|
+
## Degradation Tiers
|
|
287
334
|
|
|
288
|
-
|
|
335
|
+
The MCP boundary distinguishes missing configuration, empty stores, and failed
|
|
336
|
+
stores. A failure is not evidence that no application code matches:
|
|
337
|
+
|
|
338
|
+
- **No embedding provider configured:** `codebase_retrieve` reports a configuration
|
|
339
|
+
error with embedding and explicit lexical-mode options.
|
|
340
|
+
- **Both semantic stores known empty:** the tool reports `empty_index`; see the
|
|
341
|
+
[corpus diagnostics](#semantic-corpus-diagnostics) above.
|
|
342
|
+
- **Failed dump hydration:** the tool reports `degraded_index` rather than serving
|
|
343
|
+
a clean empty result. Repair or regenerate the named embedding artifact.
|
|
344
|
+
- **Query-time storage failures:** vector, metadata, and graph adapter exceptions
|
|
345
|
+
are translated into store errors and reported as `degraded_index` by MCP.
|
|
346
|
+
This includes failures building the metadata overview; that failure is not
|
|
347
|
+
silently omitted. Direct Ruby callers should handle `Woods::Retriever::StoreError`.
|
|
348
|
+
|
|
349
|
+
Provider failures and missing metadata for returned candidates have their own
|
|
350
|
+
error paths. Preserve their diagnostics; do not silently switch modes or treat
|
|
351
|
+
an exception as an empty match. Positive corpus counts do not override these
|
|
352
|
+
checks.
|
|
289
353
|
|
|
290
354
|
---
|
|
291
355
|
|
|
@@ -353,6 +417,10 @@ bundle exec rake woods:embed
|
|
|
353
417
|
| `text-embedding-3-small` (default) | 1536 |
|
|
354
418
|
| `text-embedding-3-large` | 3072 |
|
|
355
419
|
|
|
420
|
+
Woods' pgvector HNSW adapter supports at most 2,000 dimensions. For the large
|
|
421
|
+
model, request a supported output width explicitly or choose another backend;
|
|
422
|
+
see [pgvector configuration](CONFIGURATION_REFERENCE.md#pgvector-postgresql).
|
|
423
|
+
|
|
356
424
|
**Ollama default model:** `nomic-embed-text`. Dimensions are detected dynamically on first embed.
|
|
357
425
|
|
|
358
426
|
---
|
data/docs/SOURCE_FRESHNESS.md
CHANGED
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
application inputs with source bytes visible to the reader. It is separate from
|
|
5
5
|
index age, HEAD equality, daemon liveness and external database/runtime state.
|
|
6
6
|
|
|
7
|
-
This capability is
|
|
7
|
+
This capability is included in Woods `2.0.0`. Check the installed gem's
|
|
8
8
|
`woods-extract --help`, `rake -T woods:source_status`, and `woods_status` schema
|
|
9
9
|
before using it; upgrading the plugin alone does not upgrade Woods.
|
|
10
10
|
|
data/docs/TOKEN_BENCHMARK.md
CHANGED
|
@@ -13,7 +13,8 @@
|
|
|
13
13
|
> When the optional [`tokenizers`](https://github.com/ankane/tokenizers-ruby)
|
|
14
14
|
> gem is installed, the Ollama path uses the real BERT WordPiece tokenizer
|
|
15
15
|
> (`Woods::Embedding::TokenCounter`) instead of this heuristic. The 4.0
|
|
16
|
-
> divisor
|
|
16
|
+
> divisor applies to the OpenAI/default path; Ollama falls back to 1.5 when
|
|
17
|
+
> its tokenizer is unavailable.
|
|
17
18
|
|
|
18
19
|
This is a historical record of the benchmark that picked 4.0 over the
|
|
19
20
|
original 3.5 divisor. It is cited from five places in `lib/` as the evidence
|
|
@@ -35,17 +36,21 @@ for that choice, keep the numbers below intact if you edit this doc.
|
|
|
35
36
|
| 3.8 | 16.2% | 42.5% |
|
|
36
37
|
| **4.0 (shipped)** | **10.6%** | **35.4%** |
|
|
37
38
|
|
|
38
|
-
Mean chars/token across the corpus was **4.41** (range 3.94–5.42).
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
chars/token)
|
|
39
|
+
Mean chars/token across the corpus was **4.41** (range 3.94–5.42). These
|
|
40
|
+
aggregate results favor the 4.0 divisor for this sample; they do not establish
|
|
41
|
+
an upper bound on token counts. The recorded range includes values below 4.0,
|
|
42
|
+
so the heuristic can underestimate. Code lines and comment/YARD lines had
|
|
43
|
+
similar ratios (4.38 vs. 4.27 chars/token) in this sample.
|
|
44
|
+
|
|
45
|
+
The character estimate covers only the text passed to the counter. It does not
|
|
46
|
+
bound the serialized MCP response: text rendering, structured output, provenance,
|
|
47
|
+
and JSON framing can add bytes and tokens beyond that input.
|
|
43
48
|
|
|
44
49
|
## What shipped
|
|
45
50
|
|
|
46
51
|
**The divisor changed from 3.5 to 4.0.** It roughly halves the mean
|
|
47
|
-
|
|
48
|
-
|
|
52
|
+
error (26.2% → 10.6%) at zero new runtime dependencies. It remains an estimate,
|
|
53
|
+
not a guarantee that arbitrary input fits a model's token limit. The
|
|
49
54
|
constant lives in one place now (`Woods::TokenUtils::CHARS_PER_TOKEN_BY_PROVIDER`),
|
|
50
55
|
not scattered across call sites, see `lib/woods/token_utils.rb` for the
|
|
51
56
|
current definition and `docs/EMBEDDING_MODELS.md` for the Ollama-side ratio.
|
|
@@ -53,8 +58,9 @@ current definition and `docs/EMBEDDING_MODELS.md` for the Ollama-side ratio.
|
|
|
53
58
|
**tiktoken_ruby was deliberately not added as a runtime dependency.** A 10.6%
|
|
54
59
|
mean error is acceptable for chunking decisions, budget estimates, and
|
|
55
60
|
truncation; a native-extension dependency for marginal accuracy gains wasn't
|
|
56
|
-
worth it. The optional `tokenizers` gem
|
|
57
|
-
|
|
61
|
+
worth it. The optional `tokenizers` gem provides counts for its supported BERT
|
|
62
|
+
WordPiece tokenizer, not every model. Strict token-limit enforcement requires
|
|
63
|
+
the tokenizer used by the target model.
|
|
58
64
|
|
|
59
65
|
## Reproducing this benchmark
|
|
60
66
|
|
data/docs/TROUBLESHOOTING.md
CHANGED
|
@@ -8,7 +8,7 @@ This guide covers the most common problems encountered when installing, extracti
|
|
|
8
8
|
|
|
9
9
|
| Error message | Cause | Fix |
|
|
10
10
|
|---------------|-------|-----|
|
|
11
|
-
| `No manifest.json found` | Wrong index path or
|
|
11
|
+
| `Could not resolve a published Woods index` (older versions: `No manifest.json found`) | Wrong index path or unresolved published generation | Select the existing index using a path visible to the server process; see [startup diagnostics](#index-cannot-be-resolved-at-startup) |
|
|
12
12
|
| `uninitialized constant Rails` | Not running inside Rails app | Run via `bundle exec rake` in Rails root |
|
|
13
13
|
| `type "vector" does not exist` | pgvector not installed | `CREATE EXTENSION vector` in PostgreSQL |
|
|
14
14
|
| `Connection refused (localhost:11434)` | Ollama not running | `ollama serve` |
|
|
@@ -23,6 +23,7 @@ This guide covers the most common problems encountered when installing, extracti
|
|
|
23
23
|
| `No such container` | Wrong container name | Check with `docker ps --format '{{.Names}}'` |
|
|
24
24
|
| `JSON parse errors` (MCP) | Rails boot noise on stdout | Remove `puts` calls from initializers |
|
|
25
25
|
| Query timeout | Large table, no scope | Add scope conditions to narrow results |
|
|
26
|
+
| `Extraction failed for …; the previous generation remains active` | A consumer handled a source error during incremental extraction or refresh (included in Woods `2.0.0`) | Fix the logged source error and retry the [complete batch](INCREMENTAL_EXTRACTION.md#handled-source-errors-and-retry); watch keeps it pending |
|
|
26
27
|
| Empty extraction output | `eager_load!` failure | Check for `NameError` in boot output |
|
|
27
28
|
| Git metadata missing | Shallow clone in CI | Use `fetch-depth: 0` for complete history |
|
|
28
29
|
| Parallel tool calls all fail | MCP client batches calls | Send calls sequentially, validate params first |
|
|
@@ -52,13 +53,44 @@ version. See [manifest writer provenance](PUBLISHED_INDEX.md#manifest-writer-pro
|
|
|
52
53
|
|
|
53
54
|
If a tool call fails with **"Tool not found: … not available in the installed Woods v…"**, the client is asking for a tool a newer gem provides. Run `bundle update woods` and reconnect the MCP server, then retry.
|
|
54
55
|
|
|
56
|
+
### Watcher startup or planned restart fails
|
|
57
|
+
|
|
58
|
+
Managed `woods-watch` startup is **included in Woods `2.0.0`**; record the
|
|
59
|
+
loaded version/path and revision, then verify executable and generator help.
|
|
60
|
+
If changing an initializer stops every Foreman process, replace a bare
|
|
61
|
+
`woods:watch` entry with the [managed setup](WATCH_DAEMON.md#managed-development-startup).
|
|
62
|
+
|
|
63
|
+
Read launcher logs and `woods_status` supervision records separately from daemon
|
|
64
|
+
liveness and index freshness. `retrying` means the last generation remains usable
|
|
65
|
+
while boot is retried. A parked ownership/protocol conflict requires correcting
|
|
66
|
+
the selected owner or installed command and restarting that owner; do not delete
|
|
67
|
+
claim files or kill PIDs taken from status. No index-visible record exists before
|
|
68
|
+
the first boot resolves the application's output directory.
|
|
69
|
+
|
|
70
|
+
Unset `WOODS_WATCH_IDLE_TIMEOUT` in managed modes. If the boot deadline is reached,
|
|
71
|
+
diagnose Bundler/initializer startup before increasing `--boot-timeout`; a valid
|
|
72
|
+
long extraction has a separate readiness state and is not bounded by that clock.
|
|
73
|
+
If setup created a Procfile but normal `bin/dev` still only launches Rails, choose
|
|
74
|
+
Puma or explicitly run the selected Foreman command. The generator never rewrites
|
|
75
|
+
`bin/dev` or starts services during preview.
|
|
76
|
+
|
|
77
|
+
If installation reports a pending transaction, use `woods:watch --operation
|
|
78
|
+
recover` through the Rails generator, initially with `--pretend`; see
|
|
79
|
+
[owned setup recovery](WATCH_DAEMON.md#ownership-updates-and-removal). That Rails
|
|
80
|
+
command boots the application first. For broken initializers use the documented
|
|
81
|
+
direct bundled Ruby helper, which does not boot Rails or require task discovery.
|
|
82
|
+
Both refuse to overwrite intervening edits. A Puma setup refusal for
|
|
83
|
+
`config/puma/development.rb` means the default
|
|
84
|
+
configuration would bypass the generated plugin; select an external/Foreman
|
|
85
|
+
arrangement instead of installing an inactive directive.
|
|
86
|
+
|
|
55
87
|
### Semantic graph validation errors
|
|
56
88
|
|
|
57
89
|
In development versions containing #413, `woods:validate` rejects graphs that
|
|
58
90
|
parse as JSON but disagree with their indexes. Errors name the section and
|
|
59
91
|
identity, for example `reverse["http_api"]: missing "Order"`, a duplicate typed
|
|
60
|
-
variant, or an indexed unit absent from `nodes`. This is
|
|
61
|
-
`2.0.0
|
|
92
|
+
variant, or an indexed unit absent from `nodes`. This is included in Woods
|
|
93
|
+
`2.0.0`; check the installed gem before expecting these diagnostics.
|
|
62
94
|
|
|
63
95
|
Keep the failing generation and report the exact errors. Run a full extraction
|
|
64
96
|
in a fresh application process with the intended bundle, then validate again.
|
|
@@ -87,7 +119,7 @@ Scoped resets leave corrupt state untouched. Valid state keeps any unrelated
|
|
|
87
119
|
operation entries, and missing state remains a no-op without creating a file.
|
|
88
120
|
A permission failure must be corrected before repair can succeed.
|
|
89
121
|
|
|
90
|
-
This recovery is
|
|
122
|
+
This recovery is included in Woods `2.0.0`; check the installed version.
|
|
91
123
|
Older versions report corrupt state as nothing to repair. Stop pipeline writers,
|
|
92
124
|
back up the configured guard state's `pipeline_guard.json`, and remove only that
|
|
93
125
|
file before restarting, or upgrade to a version containing the fix.
|
|
@@ -159,7 +191,12 @@ For subsequent runs, use incremental mode instead of full extraction:
|
|
|
159
191
|
bundle exec rake woods:incremental
|
|
160
192
|
```
|
|
161
193
|
|
|
162
|
-
Incremental extraction
|
|
194
|
+
Incremental extraction dispatches the selected changed paths, including affected
|
|
195
|
+
concern consumers and whole-app extractors whose trigger paths changed. The default
|
|
196
|
+
Git range is `HEAD~1`; pass an explicit range or `CHANGED_FILES` for other batches.
|
|
197
|
+
It can reduce extraction work, but Rails boot, graph rebuilding, and publication
|
|
198
|
+
still contribute to runtime. Measure the improvement in your application; Woods
|
|
199
|
+
does not guarantee a speedup. See the [incremental contract](INCREMENTAL_EXTRACTION.md).
|
|
163
200
|
|
|
164
201
|
---
|
|
165
202
|
|
|
@@ -227,7 +264,7 @@ dependents after an incremental run.
|
|
|
227
264
|
identities when restoring and updating the graph.
|
|
228
265
|
|
|
229
266
|
**Fix:** Check whether the installed version includes B-193; this fix is
|
|
230
|
-
|
|
267
|
+
included in Woods `2.0.0`. After upgrading to a version containing the fix, run
|
|
231
268
|
`bundle exec rake woods:extract` once to rebuild lost reverse dependencies.
|
|
232
269
|
Loading an already damaged graph does not restore discarded entries. See the
|
|
233
270
|
[incremental graph contract](INCREMENTAL_EXTRACTION.md#the-contract).
|
|
@@ -238,7 +275,7 @@ Loading an already damaged graph does not restore discarded entries. See the
|
|
|
238
275
|
most files as `change_frequency: new` in a shallow CI checkout.
|
|
239
276
|
|
|
240
277
|
**Cause:** A shallow clone truncates HEAD ancestry. The shallow-checkout guard is
|
|
241
|
-
|
|
278
|
+
included in Woods `2.0.0`: Woods omits git enrichment and warns once,
|
|
242
279
|
rather than treating the truncated history as complete. If repository depth
|
|
243
280
|
cannot be verified, enrichment is also omitted; check git access and version.
|
|
244
281
|
|
|
@@ -255,12 +292,33 @@ clone), then run full extraction to replace retained metadata:
|
|
|
255
292
|
Two commits can suffice for an incremental diff, but do not establish the full
|
|
256
293
|
ancestry needed for churn metadata.
|
|
257
294
|
|
|
295
|
+
### Git executable is missing from the extraction environment
|
|
296
|
+
|
|
297
|
+
**Symptom:** Extraction logs `Git history unavailable: git executable was not
|
|
298
|
+
found in PATH`, and newly extracted units have no `metadata.git`. Older builds
|
|
299
|
+
can omit this enrichment silently when the executable is missing.
|
|
300
|
+
|
|
301
|
+
**Fix:** Run `git --version` in the same container and environment that runs
|
|
302
|
+
extraction. Install Git 2.31 or newer there, ensure its executable is on `PATH`,
|
|
303
|
+
then run full `woods:extract` to refresh every unit's history. A working Git
|
|
304
|
+
installation on the host does not provide Git inside an application container.
|
|
305
|
+
|
|
306
|
+
Extraction continues without inventing zero-commit history. The warning appears
|
|
307
|
+
once per extractor instance when the application has a `.git` entry or an
|
|
308
|
+
explicit `WOODS_GIT_DIR`/`GIT_DIR` setting. A source archive with neither remains
|
|
309
|
+
supported and quiet. `GIT_BRANCH`/`GIT_SHA` provenance fallback is unchanged;
|
|
310
|
+
those values identify a build but cannot supply per-file history.
|
|
311
|
+
|
|
312
|
+
This diagnostic is emitted during extraction. `woods_status.ready` and a
|
|
313
|
+
manifest Git SHA do not establish that per-unit history was available, and
|
|
314
|
+
`recent_changes` returning no results does not prove no files changed.
|
|
315
|
+
|
|
258
316
|
---
|
|
259
317
|
|
|
260
318
|
### Git enrichment warns that history could not be read completely
|
|
261
319
|
|
|
262
|
-
|
|
263
|
-
|
|
320
|
+
Woods 2.0 uses an explicit merge-diff mode requiring **Git 2.31 or newer**.
|
|
321
|
+
First confirm the installed Woods version.
|
|
264
322
|
Check `git --version` inside the same container/process environment as extraction,
|
|
265
323
|
and upgrade git if it is older. On a supported version, check that the application's
|
|
266
324
|
`HEAD` and object store can be read using the same `WOODS_GIT_DIR` setting.
|
|
@@ -292,29 +350,62 @@ When it does not, the git keys are omitted from every unit, provenance records
|
|
|
292
350
|
`"unknown"`, and one warning names git's own reason. Absent keys mean "not
|
|
293
351
|
known"; they never mean "brand new".
|
|
294
352
|
|
|
295
|
-
**Fix:**
|
|
296
|
-
|
|
353
|
+
**Fix:** Restore access to both the worktree-specific Git directory and the
|
|
354
|
+
shared objects and refs using the mount layouts below, then run a full
|
|
355
|
+
`woods:extract` to replace retained metadata.
|
|
356
|
+
|
|
357
|
+
### Git directory mounts for linked worktrees
|
|
358
|
+
|
|
359
|
+
`WOODS_GIT_DIR` is passed directly to Git's `--git-dir`. It selects that
|
|
360
|
+
directory's `HEAD` for manifest provenance, per-file history, and incremental
|
|
361
|
+
diff ranges. **For a linked worktree, pointing it at the shared `.git` root
|
|
362
|
+
selects the primary checkout's HEAD.** A successful Git command alone does
|
|
363
|
+
not prove Woods is reading the intended branch.
|
|
364
|
+
|
|
365
|
+
First inspect Git metadata on the host, from the intended worktree:
|
|
297
366
|
|
|
298
367
|
```bash
|
|
299
|
-
|
|
300
|
-
#
|
|
301
|
-
|
|
302
|
-
|
|
368
|
+
git -C /path/to/worktree rev-parse --absolute-git-dir
|
|
369
|
+
# Example: /path/to/repo/.git/worktrees/wt
|
|
370
|
+
git -C /path/to/worktree rev-parse --path-format=absolute --git-common-dir
|
|
371
|
+
# Example: /path/to/repo/.git
|
|
372
|
+
git -C /path/to/worktree rev-parse --abbrev-ref HEAD
|
|
373
|
+
git -C /path/to/worktree rev-parse HEAD
|
|
303
374
|
```
|
|
304
375
|
|
|
305
|
-
|
|
306
|
-
|
|
307
|
-
|
|
308
|
-
|
|
309
|
-
|
|
376
|
+
The worktree ID in this example is `wt`. Use the ID returned by Git metadata;
|
|
377
|
+
it need not match the branch name. Choose one of these layouts:
|
|
378
|
+
|
|
379
|
+
- **Same-path mount:** mount the complete shared directory read-only at its
|
|
380
|
+
original absolute path (`/path/to/repo/.git:/path/to/repo/.git:ro`). With the
|
|
381
|
+
application's existing `.git` pointer resolvable, leave `WOODS_GIT_DIR`
|
|
382
|
+
unset and remove conflicting Git-directory overrides from the environment.
|
|
383
|
+
- **Relocated mount:** mount that complete directory read-only at a new path
|
|
384
|
+
(`/path/to/repo/.git:/mounted-common:ro`), including `objects`, `refs`, and
|
|
385
|
+
`worktrees`. Select the worktree-specific directory inside it:
|
|
386
|
+
|
|
387
|
+
```bash
|
|
388
|
+
WOODS_GIT_DIR=/mounted-common/worktrees/wt bundle exec rake woods:extract
|
|
389
|
+
```
|
|
310
390
|
|
|
311
|
-
|
|
312
|
-
|
|
313
|
-
|
|
314
|
-
|
|
315
|
-
|
|
316
|
-
|
|
317
|
-
|
|
391
|
+
Mounting only the private worktree directory can leave its `commondir` pointer
|
|
392
|
+
without access to shared objects and refs. Git's own environment variables are
|
|
393
|
+
inherited by the subprocess; check any existing `GIT_DIR` and `GIT_COMMON_DIR`
|
|
394
|
+
settings when diagnosing the effective layout. The complete layouts above
|
|
395
|
+
preserve both worktree identity and shared storage.
|
|
396
|
+
|
|
397
|
+
In the extraction container, verify the relocated selection against the host
|
|
398
|
+
branch and exact SHA before extracting (replace `/app` and `wt` as needed):
|
|
399
|
+
|
|
400
|
+
```bash
|
|
401
|
+
git --git-dir=/mounted-common/worktrees/wt --work-tree=/app -C /app rev-parse --abbrev-ref HEAD
|
|
402
|
+
git --git-dir=/mounted-common/worktrees/wt --work-tree=/app -C /app rev-parse HEAD
|
|
403
|
+
```
|
|
404
|
+
|
|
405
|
+
After fixing the selection, run full `woods:extract` and verify the published
|
|
406
|
+
manifest. Incremental extraction can retain older per-file Git metadata.
|
|
407
|
+
A commit alone does not necessarily trigger the source-file watcher; run a
|
|
408
|
+
full extraction when current history and provenance are required.
|
|
318
409
|
|
|
319
410
|
---
|
|
320
411
|
|
|
@@ -374,29 +465,39 @@ retries after a later filesystem event.
|
|
|
374
465
|
|
|
375
466
|
### `manifest.json` shows the wrong branch (or `git_branch: "unknown"`) in a worktree
|
|
376
467
|
|
|
377
|
-
**Symptom:** `git_branch` / `git_sha` in `manifest.json` name a different
|
|
468
|
+
**Symptom:** `git_branch` / `git_sha` in `manifest.json` name a different
|
|
469
|
+
branch or SHA than the intended worktree, or report `"unknown"`.
|
|
378
470
|
|
|
379
|
-
**Cause:**
|
|
471
|
+
**Cause:** An unreachable `.git` file's `gitdir:` pointer prevents Git from
|
|
472
|
+
resolving the worktree's HEAD. An override selecting the shared `.git` root
|
|
473
|
+
instead resolves the primary checkout's HEAD successfully. That wrong selection
|
|
474
|
+
also affects per-file history and HEAD-based incremental ranges.
|
|
380
475
|
|
|
381
|
-
**Fix:**
|
|
476
|
+
**Fix:** Follow [Git directory mounts for linked worktrees](#git-directory-mounts-for-linked-worktrees),
|
|
477
|
+
compare the selected branch and exact SHA in the extraction environment, then
|
|
478
|
+
run a full extraction. Compare the newly published manifest, not a retained
|
|
479
|
+
generation. A commit without a source edit may leave the watcher idle.
|
|
382
480
|
|
|
383
|
-
|
|
384
|
-
|
|
385
|
-
|
|
386
|
-
|
|
387
|
-
relative and resolves outside the mount.
|
|
481
|
+
For a checkout legitimately shipped without `.git` (such as a source tarball),
|
|
482
|
+
`GIT_BRANCH` / `GIT_SHA` can supply provenance. They are fallbacks only when
|
|
483
|
+
`.git` is absent or Git is unavailable; a present but unresolvable `.git`
|
|
484
|
+
reports `"unknown"` instead of substituting stale build arguments.
|
|
388
485
|
|
|
389
486
|
---
|
|
390
487
|
|
|
391
488
|
## MCP Server Problems
|
|
392
489
|
|
|
393
|
-
|
|
490
|
+
<a id="no-manifestjson-error-when-starting-the-index-server"></a>
|
|
491
|
+
|
|
492
|
+
### Index cannot be resolved at startup
|
|
394
493
|
|
|
395
|
-
**Symptom:**
|
|
494
|
+
**Symptom:** An Index MCP executable exits with `Could not resolve a published Woods index in: /path/to/...` even though extraction completed. This headline is included in Woods `2.0.0`; older versions say `No manifest.json found`. Both mean the selected index could not resolve its manifest, not that an atomic index needs a root manifest.
|
|
396
495
|
|
|
397
|
-
|
|
496
|
+
Embedded Index MCP startup through `IndexReader` also raises an `ArgumentError` with the selected directory and layout guidance when the marker cannot resolve a manifest, including malformed marker shapes such as `[]` or a numeric `payload` (included in Woods `2.0.0`). Earlier builds may expose a raw `TypeError` or `NoMethodError` for those shapes. Inspect the marker and preserve the failing index before attempting recovery.
|
|
398
497
|
|
|
399
|
-
**
|
|
498
|
+
**Cause:** The selected directory is not the published index root, the published generation cannot be resolved, or the path is not visible to the MCP process. A container path is appropriate for a container process; a host process needs the host-visible path.
|
|
499
|
+
|
|
500
|
+
**Fix:** Point at the existing index before extracting again. Check the examined directory in the error and the [MCP path precedence](CONFIGURATION_REFERENCE.md#environment-variables). For a host-side launch whose working directory contains `tmp/woods`, for example:
|
|
400
501
|
|
|
401
502
|
```json
|
|
402
503
|
{
|
|
@@ -409,13 +510,7 @@ relative and resolves outside the mount.
|
|
|
409
510
|
}
|
|
410
511
|
```
|
|
411
512
|
|
|
412
|
-
|
|
413
|
-
|
|
414
|
-
```bash
|
|
415
|
-
ls ./tmp/woods/manifest.json
|
|
416
|
-
```
|
|
417
|
-
|
|
418
|
-
**Since Woods 2.0, a healthy index may not have `manifest.json` at the output root at all.** Extraction publishes each generation into an immutable `payloads/gen-<N>/` directory and points to it from `generation.json`. If the flat path is missing, check the payload path instead before assuming extraction failed:
|
|
513
|
+
**Since Woods 2.0, a healthy index may not have `manifest.json` at the output root at all.** Extraction publishes each generation into an immutable `payloads/gen-<N>/` directory and points to it from `generation.json`. Inspect the marker and its payload in the MCP process's filesystem before assuming extraction failed (use the generation named by your marker):
|
|
419
514
|
|
|
420
515
|
```bash
|
|
421
516
|
cat ./tmp/woods/generation.json # {"number": 42, "payload": "payloads/gen-42", ...}
|
|
@@ -427,7 +522,7 @@ Update their gate using the [filesystem layout contract](INDEX_LAYOUT.md), which
|
|
|
427
522
|
includes Bash/jq and Python readers. An upload must pin and copy one complete
|
|
428
523
|
payload before publishing its captured pointer; keep a failed copy unpublished.
|
|
429
524
|
|
|
430
|
-
`woods-mcp-start` and `IndexReader`
|
|
525
|
+
`woods-mcp-start` and `IndexReader` resolve this automatically; these commands are for manual inspection. Legacy flat indexes use a root `manifest.json`. If neither layout resolves, check the selected path, pointer, payload and any volume mount. See [DOCKER_SETUP.md](DOCKER_SETUP.md) for container launches.
|
|
431
526
|
|
|
432
527
|
---
|
|
433
528
|
|
data/docs/UPGRADING_TO_2.md
CHANGED
|
@@ -2,12 +2,10 @@
|
|
|
2
2
|
|
|
3
3
|
Woods 2.0 changes observable index identifiers, publication layout, vector-store reconciliation, and the supported MCP surface. Plan a clean re-index. Do not upgrade a shared or durable index in place without a backup and a rollback window.
|
|
4
4
|
|
|
5
|
-
This guide
|
|
5
|
+
This guide covers the supported 1.6.x line and targets 2.0.0. Use the latest
|
|
6
|
+
published 1.6.x security patch as the rollback version.
|
|
6
7
|
|
|
7
8
|
<!-- release-state:upgrade-availability -->
|
|
8
|
-
> RubyGems lists 2.0.0.beta3 as a prerelease. Pin it explicitly with
|
|
9
|
-
> `gem "woods", "2.0.0.beta3"`; `~> 2.0` resolves only once
|
|
10
|
-
> 2.0.0 is published.
|
|
11
9
|
<!-- release-state:end -->
|
|
12
10
|
|
|
13
11
|
## Upgrade outcome
|
|
@@ -95,7 +93,11 @@ Also back up managed Obsidian/Unblocked destinations before allowing a mass stal
|
|
|
95
93
|
|
|
96
94
|
### 3. Choose a rollback point
|
|
97
95
|
|
|
98
|
-
|
|
96
|
+
Record and test a Gemfile/lockfile selecting the latest published 1.6.x security
|
|
97
|
+
patch as the rollback bundle. If the current installation is older, verify that
|
|
98
|
+
patched v1 bundle before beginning the v2 migration. Keep its commit and all
|
|
99
|
+
durable-store backups until v2 extraction, MCP calls, retrieval, and exports are
|
|
100
|
+
verified. Downgrading the gem does not translate v2 identifiers back to v1.
|
|
99
101
|
|
|
100
102
|
## Upgrade the application
|
|
101
103
|
|
|
@@ -142,6 +144,10 @@ bin/rails woods:validate
|
|
|
142
144
|
bin/rails woods:stats
|
|
143
145
|
```
|
|
144
146
|
|
|
147
|
+
Included in Woods `2.0.0`: `woods:clean` removes index artifacts but keeps
|
|
148
|
+
the output directory and its hidden extraction guard. This stable guard lets
|
|
149
|
+
concurrent writers coordinate safely; its presence does not mean an index remains.
|
|
150
|
+
|
|
145
151
|
The clean extract is required for corrected identifier shapes. Do not use an incremental run as the first v2 extraction: after `woods:clean` there is no baseline, and v2 `woods:incremental` refuses that state rather than publishing a near-empty index as the application's complete truth.
|
|
146
152
|
|
|
147
153
|
An interrupted extraction leaves readers on the last complete generation because Woods publishes `generation.json` only after the payload is complete. Re-run the task; do not delete a partial directory speculatively. A run that completes its payload but cannot publish the marker now fails loudly instead of reporting success, so treat a non-zero exit as work to redo rather than as a partial success.
|
|
@@ -327,7 +333,7 @@ Complete every applicable check:
|
|
|
327
333
|
If verification fails:
|
|
328
334
|
|
|
329
335
|
1. stop v2 MCP, watcher, embedding, and exporter processes;
|
|
330
|
-
2. restore the v1 Gemfile and lockfile or deploy
|
|
336
|
+
2. restore the tested, patched v1 Gemfile and lockfile or deploy its recorded commit;
|
|
331
337
|
3. run the v1 `woods:clean` before restoring anything under the configured output directory;
|
|
332
338
|
4. either restore the complete pre-upgrade v1 output-directory backup, or run a fresh v1 extraction and then restore its v1 `dumps/` and configuration artifacts;
|
|
333
339
|
5. restore external vector-store and managed export backups when v2 modified them;
|