woods 2.0.0.beta4 → 2.0.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (95) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +495 -471
  3. data/CONTRIBUTING.md +12 -2
  4. data/README.md +11 -26
  5. data/docs/AGENT_GUIDE.md +31 -12
  6. data/docs/AGENT_SETUP.md +17 -10
  7. data/docs/AUTOMATIC_MAINTENANCE.md +222 -0
  8. data/docs/BACKEND_MATRIX.md +13 -7
  9. data/docs/CLIENT_HOOKS.md +1 -1
  10. data/docs/CONFIGURATION_REFERENCE.md +44 -27
  11. data/docs/CONSOLE_MCP_SETUP.md +95 -16
  12. data/docs/DOCKER_SETUP.md +15 -0
  13. data/docs/EVALUATION.md +10 -4
  14. data/docs/EXTRACTOR_REFERENCE.md +14 -2
  15. data/docs/FAQ.md +14 -3
  16. data/docs/GETTING_STARTED.md +18 -17
  17. data/docs/INCREMENTAL_EXTRACTION.md +8 -3
  18. data/docs/INDEX_LAYOUT.md +2 -2
  19. data/docs/MCP_HTTP_TRANSPORT.md +54 -2
  20. data/docs/MCP_SERVERS.md +28 -11
  21. data/docs/MCP_TOOL_COOKBOOK.md +1 -1
  22. data/docs/MCP_WORKTREE_SETUP.md +13 -1
  23. data/docs/PUBLISHED_INDEX.md +1 -1
  24. data/docs/README.md +2 -1
  25. data/docs/RETRIEVAL_GUIDE.md +57 -8
  26. data/docs/SOURCE_FRESHNESS.md +1 -1
  27. data/docs/TOKEN_BENCHMARK.md +16 -10
  28. data/docs/TROUBLESHOOTING.md +133 -37
  29. data/docs/UPGRADING_TO_2.md +69 -7
  30. data/docs/WATCH_DAEMON.md +172 -17
  31. data/docs/WHY_WOODS.md +9 -5
  32. data/exe/woods-console +13 -11
  33. data/exe/woods-mcp-http +16 -9
  34. data/exe/woods-watch +5 -0
  35. data/lib/generators/woods/watch_generator.rb +53 -0
  36. data/lib/puma/plugin/woods.rb +10 -0
  37. data/lib/tasks/woods.rake +14 -0
  38. data/lib/woods/cache/cache_middleware.rb +18 -11
  39. data/lib/woods/console/adapter_family.rb +39 -0
  40. data/lib/woods/console/credential_index.rb +33 -3
  41. data/lib/woods/console/embedded_executor.rb +401 -43
  42. data/lib/woods/console/model_validator.rb +8 -0
  43. data/lib/woods/console/rack_middleware.rb +39 -10
  44. data/lib/woods/console/redactor.rb +24 -10
  45. data/lib/woods/console/safe_context.rb +44 -7
  46. data/lib/woods/console/sql_noise_stripper.rb +41 -12
  47. data/lib/woods/console/sql_table_scanner.rb +45 -34
  48. data/lib/woods/console/sql_validator.rb +37 -2
  49. data/lib/woods/console/stdio_transport.rb +27 -0
  50. data/lib/woods/extractor.rb +25 -7
  51. data/lib/woods/git_command.rb +6 -7
  52. data/lib/woods/git_provenance.rb +4 -6
  53. data/lib/woods/mcp/bearer_auth.rb +1 -1
  54. data/lib/woods/mcp/bootstrapper.rb +3 -1
  55. data/lib/woods/mcp/initialization_guidance.rb +1 -1
  56. data/lib/woods/mcp/origin_guard.rb +24 -77
  57. data/lib/woods/mcp/origin_policy.rb +124 -0
  58. data/lib/woods/mcp/server.rb +41 -9
  59. data/lib/woods/railtie_support.rb +8 -0
  60. data/lib/woods/retrieval/corpus_status.rb +46 -0
  61. data/lib/woods/retriever.rb +19 -7
  62. data/lib/woods/storage/local_corpus_stats.rb +32 -0
  63. data/lib/woods/storage/metadata_store.rb +20 -0
  64. data/lib/woods/storage/vector_store.rb +10 -0
  65. data/lib/woods/version.rb +1 -1
  66. data/lib/woods/watch/child_environment.rb +30 -0
  67. data/lib/woods/watch/cli.rb +91 -0
  68. data/lib/woods/watch/daemon.rb +55 -7
  69. data/lib/woods/watch/event_stream.rb +70 -0
  70. data/lib/woods/watch/guardian.rb +142 -0
  71. data/lib/woods/watch/installation/layout.rb +70 -0
  72. data/lib/woods/watch/installation/options.rb +128 -0
  73. data/lib/woods/watch/installation/planner.rb +128 -0
  74. data/lib/woods/watch/installation/probe.rb +101 -0
  75. data/lib/woods/watch/installation/receipt.rb +77 -0
  76. data/lib/woods/watch/installation/recovery.rb +64 -0
  77. data/lib/woods/watch/installation/templates.rb +58 -0
  78. data/lib/woods/watch/installation.rb +56 -0
  79. data/lib/woods/watch/lifecycle.rb +182 -0
  80. data/lib/woods/watch/managed_child.rb +113 -0
  81. data/lib/woods/watch/managed_cleanup.rb +48 -0
  82. data/lib/woods/watch/managed_process.rb +144 -0
  83. data/lib/woods/watch/puma_adapter.rb +87 -0
  84. data/lib/woods/watch/puma_child.rb +66 -0
  85. data/lib/woods/watch/supervision_records.rb +95 -0
  86. data/lib/woods/watch/supervision_status.rb +104 -0
  87. data/lib/woods/watch/supervisor.rb +161 -0
  88. data/lib/woods/watch/supervisor_reporting.rb +46 -0
  89. data/plugin/.claude-plugin/plugin.json +1 -1
  90. data/plugin/skills/woods-agent-enable/SKILL.md +1 -1
  91. data/plugin/skills/woods-diagnose/SKILL.md +88 -8
  92. data/plugin/skills/woods-investigate/SKILL.md +6 -6
  93. data/plugin/skills/woods-mcp-config/SKILL.md +43 -1
  94. data/plugin/skills/woods-setup/SKILL.md +66 -4
  95. metadata +37 -5
@@ -86,16 +86,68 @@ Clients must send `Authorization: Bearer $WOODS_MCP_HTTP_TOKEN` on every request
86
86
 
87
87
  ### Browser origins (DNS rebinding defense)
88
88
 
89
- A second middleware, `Woods::MCP::OriginGuard`, rejects requests whose `Origin` header is outside an allow-list. Requests without an `Origin` header (curl, MCP stdio clients, server-to-server) pass through, bearer auth still gates them.
89
+ **Unreleased diagnostic correction:** invalid origin encodings also refuse before
90
+ binding HTTP. `woods-mcp-http` exits 2 with one bounded `ConfigurationError`
91
+ message naming the invalid entry; re-enter that origin using an ASCII hostname
92
+ (or its Punycode form). No allowlist keeps the existing defaults. Explicit entries
93
+ are literal origins, not wildcard patterns.
94
+ Retain the actual request Host through a reverse proxy; forwarded headers do not
95
+ replace it. When authentication is configured, requests without Host still pass
96
+ through bearer authentication.
97
+
98
+ In the 2.0.1 maintenance patch, preflight and SDK dispatch share an immutable
99
+ normalized policy. A configured list replaces the default browser origins;
100
+ include loopback explicitly if needed. Cross-origin ports must match an entry,
101
+ while omitted default HTTP(S) ports match their explicit 80/443 forms. A portless
102
+ entry permits same-authority traffic rather than every cross-port browser origin.
103
+ Console defaults remain HTTP loopback; Index HTTP defaults include HTTP and HTTPS
104
+ loopback. When no list is set, these defaults are unchanged. Invalid entries fail
105
+ at boot, naming the offending entry. Restart after changing the configuration.
106
+
107
+ A second middleware, `Woods::MCP::OriginGuard`, rejects requests whose `Origin` header is outside an allow-list. Requests without an `Origin` header still pass through Host validation and configured bearer authentication.
90
108
 
91
109
  | Scenario | `WOODS_MCP_HTTP_ALLOWED_ORIGINS` | Origins accepted |
92
110
  |--------------------|-----------------------------------------|-------------------------------------------------------------------|
93
- | default | unset | `http(s)://localhost`, `127.0.0.1`, `::1` (any port) |
111
+ | default | unset | HTTP(S) loopback, matching request authority; explicit entries for cross-port origins |
94
112
  | explicit list | `https://app.example.com` | exactly `https://app.example.com`, loopback no longer allowed |
95
113
  | multiple origins | `https://a.example,https://b.example` | each listed origin |
96
114
 
97
115
  `OPTIONS` preflights are answered with the matching `Access-Control-Allow-*` headers; successful responses carry `Access-Control-Allow-Origin` and `Vary: Origin`. `Access-Control-Expose-Headers: Mcp-Session-Id` appears only in legacy session mode (`WOODS_MCP_HTTP_STATELESS=0`).
98
116
 
117
+ ### Origin configuration compatibility
118
+
119
+ These details apply to the supporting security-patch revisions described above.
120
+
121
+ The Index HTTP launcher names `WOODS_MCP_HTTP_ALLOWED_ORIGINS` and the offending
122
+ entry when refusing malformed origin configuration, including under a POSIX
123
+ (`C`) locale. It emits one diagnostic before index resolution or HTTP binding.
124
+
125
+ Configured origins are literal values, with lowercasing, one trailing slash
126
+ removed and HTTP(S) default-port normalization. Wildcard-looking hostnames are
127
+ literal hostnames, not patterns. A nonempty explicit list replaces browser-origin
128
+ defaults, including loopback; add the loopback browser origins you need. Console's
129
+ configured defaults are HTTP loopback; the Index defaults include HTTP and HTTPS
130
+ loopback. Loopback **Host** acceptance is separate from browser-origin acceptance.
131
+
132
+ Ruby-configured 2.x origin entries reject surrounding whitespace; 1.6.4 trims it.
133
+ The `woods-mcp-http` comma-separated environment setting trims each entry on both
134
+ lines. Use whitespace-free values for portable configuration. The exact-port and
135
+ same-authority rules above still apply; literal matching does not imply wildcard
136
+ or arbitrary cross-port access.
137
+
138
+ The guards read the actual `Origin` and `Host`; `Forwarded` and
139
+ `X-Forwarded-*` do not replace them. Preserve the public Host through a proxy.
140
+ A genuinely absent Host is accepted by the Host policy, but any supplied Origin
141
+ must still pass its own check. Console requests and token-configured Index
142
+ requests still require bearer authentication before dispatch. The loopback-only
143
+ Index mode without a token retains its documented unauthenticated behavior;
144
+ CORS preflights do not dispatch tools.
145
+
146
+ Malformed Origin/Host values receive constant `403` responses without reflecting
147
+ the values. Invalid Authorization receives `401` when it reaches the auth guard.
148
+ Malformed HTTP framing can instead receive Puma's `400` before Rack runs. Other
149
+ HTTP servers may reject framing at their own boundary.
150
+
99
151
  ### TLS termination
100
152
 
101
153
  The server speaks plain HTTP. Any deployment beyond a single trusted host should front it with a reverse proxy that handles TLS, HTTP/2, and connection limits.
data/docs/MCP_SERVERS.md CHANGED
@@ -56,6 +56,12 @@ Prefer the application's bundle and a project-scoped configuration:
56
56
 
57
57
  `woods-mcp-start` checks that the directory and published manifest exist, then replaces itself with `woods-mcp`. It does not install dependencies or restart a crashed process.
58
58
 
59
+ The MCP client owns this reader process. Automatic index maintenance has its own
60
+ [watcher startup](WATCH_DAEMON.md#managed-development-startup): configuring MCP
61
+ does not enable it. A running reader sees new published generations on subsequent
62
+ calls without reconnecting. Verify a real edit, not only a successful connection;
63
+ see [automatic maintenance](AUTOMATIC_MAINTENANCE.md).
64
+
59
65
  `woods_status.index.woods_version` identifies the last publisher of the served
60
66
  manifest; `server.version` identifies the running MCP reader. Missing writer
61
67
  provenance is `null`. See [manifest writer provenance](PUBLISHED_INDEX.md#manifest-writer-provenance).
@@ -133,12 +139,23 @@ before using `codebase_retrieve`, including after reload. The guidance grants
133
139
  no extraction, configuration-change, or Console authorization. Detailed usage
134
140
  belongs in the [agent guide](AGENT_GUIDE.md).
135
141
 
136
- This addition is unreleased after `2.0.0.beta2`. The SDK omits `instructions`
142
+ This addition is included in Woods `2.0.0`. The SDK omits `instructions`
137
143
  when negotiating protocol `2024-11-05`; that behavior is preserved. Older gems,
138
144
  legacy clients, and clients that do not show server instructions can use the
139
145
  agent guide or investigation skill. Leave protocol negotiation enabled rather
140
146
  than pinning a newer version solely to obtain guidance.
141
147
 
148
+ ### Structural and semantic readiness
149
+
150
+ `ready` describes the published structural index. A reachable embedding provider
151
+ and bootstrap state `hydrated` do not establish that semantic stores contain
152
+ data. Supporting readers also report `retriever.corpus`: locally known vector
153
+ and metadata entry counts, counts by type, and whether those stores are empty,
154
+ populated, or unknown. These are stored records (including chunks), not a count
155
+ of extracted units with verified embedding coverage. Positive counts do not
156
+ prove complete coverage or working provider access. See
157
+ [semantic corpus diagnostics](RETRIEVAL_GUIDE.md#semantic-corpus-diagnostics).
158
+
142
159
  ### Tools (29 — 14 registered in the packaged default)
143
160
 
144
161
  The Index Server defines 29 schemas across core and conditional capabilities. The normal packaged executable registers the 14 tools below; the remaining schemas require the specialized wiring described afterward.
@@ -194,7 +211,7 @@ process after publishing a new embedded index.
194
211
 
195
212
  ### Graph-analysis pages
196
213
 
197
- Unreleased after `2.0.0.beta3`: `graph_analysis` enforces its advertised default
214
+ Included in Woods `2.0.0`: `graph_analysis` enforces its advertised default
198
215
  of 20 rows per section. Pass `limit` and `offset` to page one selected `analysis`
199
216
  or each section of `analysis: "all"`. Explicit limits also bound nested hub
200
217
  `dependents` lists. Older servers may return every section row when `limit` is
@@ -214,7 +231,7 @@ they do not establish complete source-reference coverage.
214
231
  Search responses retain `query`, `result_count`, and `results`; `result_count`
215
232
  is the number returned, not an estimated total. The additive `completeness`
216
233
  object describes the requested types, literal filters, and fields in the pinned
217
- generation. This contract is unreleased after `2.0.0.beta2`.
234
+ generation. This contract is included in Woods `2.0.0`.
218
235
 
219
236
  | `reason` | `status` | `has_more` | `total_matches` |
220
237
  |---|---|---|---|
@@ -260,16 +277,16 @@ Supporting servers expose the annotated, paginated traversal result in
260
277
  stdio and HTTP servers. Read `data.total_is_exact`, `data.graph_coverage`, budget
261
278
  counters and optional explanation witnesses there; `content[0].text` and
262
279
  `structuredContent.text` keep the same human-readable rendering. No `format`
263
- tool argument is needed or accepted. This additive data payload is unreleased
264
- after `2.0.0.beta3`; verify the installed response before relying on it. Older
280
+ tool argument is needed or accepted. This additive data payload is included in
281
+ Woods `2.0.0`; verify the installed response before relying on it. Older
265
282
  human-renderer responses can carry only text. The structured nodes and witnesses
266
283
  cover the same page, not an additional traversal or an unpaginated graph.
267
284
 
268
285
  Successful responses carry `graph_coverage` with `scope: "published_relationships"`,
269
286
  `source_references: "not_exhaustive"`, and a human-readable `notice`. Text formats
270
287
  show the same notice, including compact, root-only and empty-page responses.
271
- This response metadata and the total exactness field below are unreleased after
272
- Woods `2.0.0.beta3`; older servers need the same conservative interpretation.
288
+ This response metadata and the total exactness field below are included in Woods
289
+ `2.0.0`; older servers need the same conservative interpretation.
273
290
 
274
291
  ### Dependency traversal budgets
275
292
 
@@ -310,14 +327,14 @@ for stable pages. No wall-clock deadline is used, so cutoffs are deterministic.
310
327
  Budgets cover traversal work after per-generation graph loading and cache
311
328
  preparation (JSON parsing, typed-edge normalization, node types and database
312
329
  metadata). They do not cap that initial load, elapsed time, or total process
313
- memory. These arguments are unreleased in Woods 2.0.0.beta2; check the connected
330
+ memory. These arguments are included in Woods `2.0.0`; check the connected
314
331
  server's tool schema before sending them to an older installation.
315
332
 
316
333
  ### Traversal explanations
317
334
 
318
335
  Supporting development versions accept `explain: true` on `dependencies` and
319
- `dependents`. Check the connected schema first; this option is unreleased after
320
- 2.0.0.beta2. Omitted or false keeps the existing compact response.
336
+ `dependents`. Check the connected schema first; this option is included in Woods
337
+ `2.0.0`. Omitted or false keeps the existing compact response.
321
338
 
322
339
  The additive `explanation` object contains:
323
340
 
@@ -343,7 +360,7 @@ it describes identifier-level reachability, never a uniquely typed path.
343
360
  A true value means only that identities along this witness have unambiguous
344
361
  types. It does not establish source-reference coverage or observed execution.
345
362
  Text labels this `witness types unambiguous=yes/no`; the JSON key and its meaning
346
- remain unchanged. The text label change is unreleased after `2.0.0.beta3`.
363
+ remain unchanged. The text label change is included in Woods `2.0.0`.
347
364
  `types` filters retain the compact traversal's identifier-level semantics: any
348
365
  registered type can qualify a name, while edge evidence keeps its actual source
349
366
  owner. Multiple relationship kinds between the same endpoints remain separate.
@@ -356,7 +356,7 @@ column.
356
356
  Use `lookup` on the returned identifier for source, actions, and routes. Search
357
357
  `source_code` for textual matches beyond names. Supporting versions distinguish
358
358
  exact totals from a bounded result prefix; `partial` means this is discovery,
359
- not an exhaustive list. Completeness metadata is unreleased after `2.0.0.beta2`;
359
+ not an exhaustive list. Completeness metadata is included in Woods `2.0.0`;
360
360
  see the [search contract](MCP_SERVERS.md#search-completeness).
361
361
 
362
362
  ---
@@ -66,7 +66,19 @@ The published payload's `manifest.json` records the extraction's `git_branch` an
66
66
 
67
67
  Woods uses worktree-aware Git commands. If a present `.git` cannot be resolved, provenance is `"unknown"`; stale `GIT_BRANCH`/`GIT_SHA` values are not substituted. Those environment variables are fallbacks only when the root has no `.git` or Git is unavailable. Temporal snapshots skip an unknown SHA.
68
68
 
69
- For extraction in a container, make the canonical Git directory and the worktree's pointer resolvable there. Mounting only the private worktree Git directory can leave its shared object store unreachable. Follow the [Git provenance troubleshooting guide](TROUBLESHOOTING.md) for mount and `WOODS_GIT_DIR` guidance, and the [published index layout](INDEX_LAYOUT.md) when locating the manifest.
69
+ For extraction in a container, mount the complete shared Git layout at the
70
+ original path so the worktree's `.git` pointer resolves without an override,
71
+ or select `/mounted-common/worktrees/<id>` with `WOODS_GIT_DIR` inside a
72
+ relocated complete mount. Derive `<id>` from Git metadata, not the branch name.
73
+ Selecting the shared root instead uses the primary checkout's HEAD, affecting
74
+ provenance, per-file history, and incremental paths. Follow the
75
+ [worktree mount and verification steps](TROUBLESHOOTING.md#git-directory-mounts-for-linked-worktrees)
76
+ and the [published index layout](INDEX_LAYOUT.md) when locating the manifest.
77
+
78
+ Compare the branch and exact SHA in the extraction environment with the intended
79
+ worktree, then verify the published manifest after full extraction. A commit
80
+ alone may not trigger the source-file watcher; use full `woods:extract` when
81
+ current Git history is required.
70
82
 
71
83
  ## Troubleshooting
72
84
 
@@ -111,7 +111,7 @@ The reader wraps `Woods::MCP::IndexReader` with `auto_refresh: false`; the unit
111
111
 
112
112
  ### Actual unit types and directory families
113
113
 
114
- Unreleased after `2.0.0.beta3`: `unit` and `units` accept actual published
114
+ Included in Woods `2.0.0`: `unit` and `units` accept actual published
115
115
  `graphql_type`, `graphql_mutation`, `graphql_resolver`, `graphql_query`, and
116
116
  `gem_source` types. Enumeration preserves each unit's actual type rather than
117
117
  labeling every GraphQL member `graphql` or every gem source `rails_source`.
data/docs/README.md CHANGED
@@ -10,7 +10,7 @@ Woods extracts runtime-accurate Rails context and serves it to coding agents thr
10
10
  | Ask an agent to install or configure Woods | [Agent setup runbook](AGENT_SETUP.md) | A safe, reviewable install with an agent handoff report |
11
11
  | Configure an MCP client or Docker path | [MCP servers](MCP_SERVERS.md) | A working Index Server and, if authorized, an optional Console Server |
12
12
  | Use Woods tools as an agent | [Agent guide](AGENT_GUIDE.md) | A repeatable query workflow for code context, flows, and blast radius |
13
- | Keep the index current automatically | [Watch daemon](WATCH_DAEMON.md) | A resident development process that catches up changes and republishes the index |
13
+ | Keep the index current automatically | [Automatic maintenance guide](AUTOMATIC_MAINTENANCE.md) | A setup with supervised indexing, automatic reader refresh, and clear hook ownership |
14
14
  | Upgrade from Woods 1.x | [Upgrade to Woods 2.0](UPGRADING_TO_2.md) | A backed-up, re-indexed, verified v2 installation |
15
15
  | Diagnose an error | [Troubleshooting](TROUBLESHOOTING.md) | Symptom-to-cause checks for extraction, MCP, embeddings, storage, and Docker |
16
16
  | Contribute to Woods | [Contributing](../CONTRIBUTING.md) | A tested change with synchronized docs and plugin guidance |
@@ -41,6 +41,7 @@ and stdio or Streamable HTTP endpoints directly.
41
41
 
42
42
  ## Index lifecycle
43
43
 
44
+ - [Automatic maintenance guide](AUTOMATIC_MAINTENANCE.md): recommended development workflow and a section-by-section map of watcher, indexing, MCP, and hook documentation.
44
45
  - [Source freshness](SOURCE_FRESHNESS.md): verify dirty source against a served generation, establish a fresh-process baseline, and understand bounded unknown results.
45
46
 
46
47
  - [Retrieval guide](RETRIEVAL_GUIDE.md): configure embeddings and understand semantic retrieval, ranking, and token budgets.
@@ -291,16 +291,65 @@ result = retriever.retrieve("what validations does Order have?")
291
291
 
292
292
  ---
293
293
 
294
- ## Degradation Tiers
295
-
296
- Retrieval degrades gracefully when components are unavailable. The Retriever itself does not implement explicit fallback tiers, degradation happens naturally through how each component handles errors:
294
+ ## Semantic corpus diagnostics
295
+
296
+ Extraction and embedding publish different data. `woods_status.ready` describes
297
+ the structural index; neither that flag nor bootstrap `hydrated` proves that
298
+ semantic retrieval has indexed records. Provider detection alone can succeed
299
+ before the first embedding run.
300
+
301
+ Supporting readers report `woods_status.retriever.corpus`:
302
+
303
+ - `state`: `empty`, `metadata_only`, `vectors_only`, `nonempty`, or `unknown`.
304
+ - `vectors` and `metadata`: each has `count`, `by_type`, and `untyped_count`.
305
+ Counts describe stored entries, including chunks; they are not distinct
306
+ extracted-unit counts or a completeness certificate. Missing type labels are
307
+ reported separately rather than assigned to a guessed type.
308
+ - Unknown counts are `null`. Diagnostics use an explicit local-store capability;
309
+ they do not query remote stores to discover their counts. A missing capability
310
+ does not mean the backend is empty or broken.
311
+
312
+ When both semantic stores are known empty, `codebase_retrieve` reports
313
+ `empty_index` with recovery guidance instead of presenting an empty match as
314
+ evidence that application code is absent. Run `woods:embed` in the application
315
+ with the intended provider and storage configuration, then reload or restart
316
+ the reader. Restart after changing provider or store configuration; reload
317
+ refreshes the stores of the existing retriever. Alternatively, explicitly select `WOODS_RETRIEVAL_MODE=lexical`
318
+ in the MCP process and restart for ranked retrieval over the published units.
319
+ Woods does not change retrieval mode automatically.
320
+
321
+ Metadata-only stores can still answer some keyword, direct, and graph queries;
322
+ they are not blocked by this diagnostic. Source-empty units deliberately retain
323
+ metadata without vectors, so the two counts need not match. Nonempty stores can
324
+ still have missing types, stale vectors, or provider failures. Inspect the
325
+ query result and embedding evidence before claiming coverage. The type-rank
326
+ table's metadata count describes the retrieval metadata store, not all units
327
+ in the structural index.
328
+
329
+ Older readers may omit `retriever.corpus`; record the reader revision separately
330
+ from the index writer version and verify the embedding artifacts directly.
331
+ Explicit lexical mode does not use or report semantic corpus counts.
297
332
 
298
- - **Embedding provider unavailable**: `codebase_retrieve` returns a structured configuration error. Check `woods_status` for retrieval readiness.
299
- - **Vector store unavailable**: vector and hybrid strategies fail at query time. Keyword and graph strategies remain available for direct calls to `SearchExecutor`.
300
- - **Metadata store error**: the structural context overview (unit counts by type) is silently omitted; `Retriever#build_structural_context` rescues `StandardError` and returns `nil`. The retrieval result is still returned without the overview.
301
- - **Graph store unavailable**: graph expansion in hybrid strategy produces no graph candidates; vector and keyword candidates are still ranked and returned.
333
+ ## Degradation Tiers
302
334
 
303
- In all cases, errors in individual components produce empty candidate sets for that source rather than raising through the `Retriever`. Configure circuit breakers via `Woods::Resilience::CircuitBreaker` on external providers (Qdrant, OpenAI) for production deployments.
335
+ The MCP boundary distinguishes missing configuration, empty stores, and failed
336
+ stores. A failure is not evidence that no application code matches:
337
+
338
+ - **No embedding provider configured:** `codebase_retrieve` reports a configuration
339
+ error with embedding and explicit lexical-mode options.
340
+ - **Both semantic stores known empty:** the tool reports `empty_index`; see the
341
+ [corpus diagnostics](#semantic-corpus-diagnostics) above.
342
+ - **Failed dump hydration:** the tool reports `degraded_index` rather than serving
343
+ a clean empty result. Repair or regenerate the named embedding artifact.
344
+ - **Query-time storage failures:** vector, metadata, and graph adapter exceptions
345
+ are translated into store errors and reported as `degraded_index` by MCP.
346
+ This includes failures building the metadata overview; that failure is not
347
+ silently omitted. Direct Ruby callers should handle `Woods::Retriever::StoreError`.
348
+
349
+ Provider failures and missing metadata for returned candidates have their own
350
+ error paths. Preserve their diagnostics; do not silently switch modes or treat
351
+ an exception as an empty match. Positive corpus counts do not override these
352
+ checks.
304
353
 
305
354
  ---
306
355
 
@@ -4,7 +4,7 @@
4
4
  application inputs with source bytes visible to the reader. It is separate from
5
5
  index age, HEAD equality, daemon liveness and external database/runtime state.
6
6
 
7
- This capability is unreleased after `2.0.0.beta2`. Check the installed gem's
7
+ This capability is included in Woods `2.0.0`. Check the installed gem's
8
8
  `woods-extract --help`, `rake -T woods:source_status`, and `woods_status` schema
9
9
  before using it; upgrading the plugin alone does not upgrade Woods.
10
10
 
@@ -13,7 +13,8 @@
13
13
  > When the optional [`tokenizers`](https://github.com/ankane/tokenizers-ruby)
14
14
  > gem is installed, the Ollama path uses the real BERT WordPiece tokenizer
15
15
  > (`Woods::Embedding::TokenCounter`) instead of this heuristic. The 4.0
16
- > divisor below is what the gem falls back to everywhere else.
16
+ > divisor applies to the OpenAI/default path; Ollama falls back to 1.5 when
17
+ > its tokenizer is unavailable.
17
18
 
18
19
  This is a historical record of the benchmark that picked 4.0 over the
19
20
  original 3.5 divisor. It is cited from five places in `lib/` as the evidence
@@ -35,17 +36,21 @@ for that choice, keep the numbers below intact if you edit this doc.
35
36
  | 3.8 | 16.2% | 42.5% |
36
37
  | **4.0 (shipped)** | **10.6%** | **35.4%** |
37
38
 
38
- Mean chars/token across the corpus was **4.41** (range 3.94–5.42). The
39
- heuristic always overestimated, never underestimated, across all 19 files,
40
- which is what makes it safe for token-limit enforcement even at its worst
41
- case. Code lines and comment/YARD lines had similar ratios (4.38 vs. 4.27
42
- chars/token), no separate handling needed for either.
39
+ Mean chars/token across the corpus was **4.41** (range 3.94–5.42). These
40
+ aggregate results favor the 4.0 divisor for this sample; they do not establish
41
+ an upper bound on token counts. The recorded range includes values below 4.0,
42
+ so the heuristic can underestimate. Code lines and comment/YARD lines had
43
+ similar ratios (4.38 vs. 4.27 chars/token) in this sample.
44
+
45
+ The character estimate covers only the text passed to the counter. It does not
46
+ bound the serialized MCP response: text rendering, structured output, provenance,
47
+ and JSON framing can add bytes and tokens beyond that input.
43
48
 
44
49
  ## What shipped
45
50
 
46
51
  **The divisor changed from 3.5 to 4.0.** It roughly halves the mean
47
- overestimate (26.2% → 10.6%) while keeping the conservative
48
- always-overestimates property, at zero new runtime dependencies. The
52
+ error (26.2% → 10.6%) at zero new runtime dependencies. It remains an estimate,
53
+ not a guarantee that arbitrary input fits a model's token limit. The
49
54
  constant lives in one place now (`Woods::TokenUtils::CHARS_PER_TOKEN_BY_PROVIDER`),
50
55
  not scattered across call sites, see `lib/woods/token_utils.rb` for the
51
56
  current definition and `docs/EMBEDDING_MODELS.md` for the Ollama-side ratio.
@@ -53,8 +58,9 @@ current definition and `docs/EMBEDDING_MODELS.md` for the Ollama-side ratio.
53
58
  **tiktoken_ruby was deliberately not added as a runtime dependency.** A 10.6%
54
59
  mean error is acceptable for chunking decisions, budget estimates, and
55
60
  truncation; a native-extension dependency for marginal accuracy gains wasn't
56
- worth it. The optional `tokenizers` gem covers the case where exact counts
57
- matter more (see above).
61
+ worth it. The optional `tokenizers` gem provides counts for its supported BERT
62
+ WordPiece tokenizer, not every model. Strict token-limit enforcement requires
63
+ the tokenizer used by the target model.
58
64
 
59
65
  ## Reproducing this benchmark
60
66