woods 2.0.0.beta4 → 2.0.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +495 -471
- data/CONTRIBUTING.md +12 -2
- data/README.md +11 -26
- data/docs/AGENT_GUIDE.md +31 -12
- data/docs/AGENT_SETUP.md +17 -10
- data/docs/AUTOMATIC_MAINTENANCE.md +222 -0
- data/docs/BACKEND_MATRIX.md +13 -7
- data/docs/CLIENT_HOOKS.md +1 -1
- data/docs/CONFIGURATION_REFERENCE.md +44 -27
- data/docs/CONSOLE_MCP_SETUP.md +95 -16
- data/docs/DOCKER_SETUP.md +15 -0
- data/docs/EVALUATION.md +10 -4
- data/docs/EXTRACTOR_REFERENCE.md +14 -2
- data/docs/FAQ.md +14 -3
- data/docs/GETTING_STARTED.md +18 -17
- data/docs/INCREMENTAL_EXTRACTION.md +8 -3
- data/docs/INDEX_LAYOUT.md +2 -2
- data/docs/MCP_HTTP_TRANSPORT.md +54 -2
- data/docs/MCP_SERVERS.md +28 -11
- data/docs/MCP_TOOL_COOKBOOK.md +1 -1
- data/docs/MCP_WORKTREE_SETUP.md +13 -1
- data/docs/PUBLISHED_INDEX.md +1 -1
- data/docs/README.md +2 -1
- data/docs/RETRIEVAL_GUIDE.md +57 -8
- data/docs/SOURCE_FRESHNESS.md +1 -1
- data/docs/TOKEN_BENCHMARK.md +16 -10
- data/docs/TROUBLESHOOTING.md +133 -37
- data/docs/UPGRADING_TO_2.md +69 -7
- data/docs/WATCH_DAEMON.md +172 -17
- data/docs/WHY_WOODS.md +9 -5
- data/exe/woods-console +13 -11
- data/exe/woods-mcp-http +16 -9
- data/exe/woods-watch +5 -0
- data/lib/generators/woods/watch_generator.rb +53 -0
- data/lib/puma/plugin/woods.rb +10 -0
- data/lib/tasks/woods.rake +14 -0
- data/lib/woods/cache/cache_middleware.rb +18 -11
- data/lib/woods/console/adapter_family.rb +39 -0
- data/lib/woods/console/credential_index.rb +33 -3
- data/lib/woods/console/embedded_executor.rb +401 -43
- data/lib/woods/console/model_validator.rb +8 -0
- data/lib/woods/console/rack_middleware.rb +39 -10
- data/lib/woods/console/redactor.rb +24 -10
- data/lib/woods/console/safe_context.rb +44 -7
- data/lib/woods/console/sql_noise_stripper.rb +41 -12
- data/lib/woods/console/sql_table_scanner.rb +45 -34
- data/lib/woods/console/sql_validator.rb +37 -2
- data/lib/woods/console/stdio_transport.rb +27 -0
- data/lib/woods/extractor.rb +25 -7
- data/lib/woods/git_command.rb +6 -7
- data/lib/woods/git_provenance.rb +4 -6
- data/lib/woods/mcp/bearer_auth.rb +1 -1
- data/lib/woods/mcp/bootstrapper.rb +3 -1
- data/lib/woods/mcp/initialization_guidance.rb +1 -1
- data/lib/woods/mcp/origin_guard.rb +24 -77
- data/lib/woods/mcp/origin_policy.rb +124 -0
- data/lib/woods/mcp/server.rb +41 -9
- data/lib/woods/railtie_support.rb +8 -0
- data/lib/woods/retrieval/corpus_status.rb +46 -0
- data/lib/woods/retriever.rb +19 -7
- data/lib/woods/storage/local_corpus_stats.rb +32 -0
- data/lib/woods/storage/metadata_store.rb +20 -0
- data/lib/woods/storage/vector_store.rb +10 -0
- data/lib/woods/version.rb +1 -1
- data/lib/woods/watch/child_environment.rb +30 -0
- data/lib/woods/watch/cli.rb +91 -0
- data/lib/woods/watch/daemon.rb +55 -7
- data/lib/woods/watch/event_stream.rb +70 -0
- data/lib/woods/watch/guardian.rb +142 -0
- data/lib/woods/watch/installation/layout.rb +70 -0
- data/lib/woods/watch/installation/options.rb +128 -0
- data/lib/woods/watch/installation/planner.rb +128 -0
- data/lib/woods/watch/installation/probe.rb +101 -0
- data/lib/woods/watch/installation/receipt.rb +77 -0
- data/lib/woods/watch/installation/recovery.rb +64 -0
- data/lib/woods/watch/installation/templates.rb +58 -0
- data/lib/woods/watch/installation.rb +56 -0
- data/lib/woods/watch/lifecycle.rb +182 -0
- data/lib/woods/watch/managed_child.rb +113 -0
- data/lib/woods/watch/managed_cleanup.rb +48 -0
- data/lib/woods/watch/managed_process.rb +144 -0
- data/lib/woods/watch/puma_adapter.rb +87 -0
- data/lib/woods/watch/puma_child.rb +66 -0
- data/lib/woods/watch/supervision_records.rb +95 -0
- data/lib/woods/watch/supervision_status.rb +104 -0
- data/lib/woods/watch/supervisor.rb +161 -0
- data/lib/woods/watch/supervisor_reporting.rb +46 -0
- data/plugin/.claude-plugin/plugin.json +1 -1
- data/plugin/skills/woods-agent-enable/SKILL.md +1 -1
- data/plugin/skills/woods-diagnose/SKILL.md +88 -8
- data/plugin/skills/woods-investigate/SKILL.md +6 -6
- data/plugin/skills/woods-mcp-config/SKILL.md +43 -1
- data/plugin/skills/woods-setup/SKILL.md +66 -4
- metadata +37 -5
data/docs/MCP_HTTP_TRANSPORT.md
CHANGED
|
@@ -86,16 +86,68 @@ Clients must send `Authorization: Bearer $WOODS_MCP_HTTP_TOKEN` on every request
|
|
|
86
86
|
|
|
87
87
|
### Browser origins (DNS rebinding defense)
|
|
88
88
|
|
|
89
|
-
|
|
89
|
+
**Unreleased diagnostic correction:** invalid origin encodings also refuse before
|
|
90
|
+
binding HTTP. `woods-mcp-http` exits 2 with one bounded `ConfigurationError`
|
|
91
|
+
message naming the invalid entry; re-enter that origin using an ASCII hostname
|
|
92
|
+
(or its Punycode form). No allowlist keeps the existing defaults. Explicit entries
|
|
93
|
+
are literal origins, not wildcard patterns.
|
|
94
|
+
Retain the actual request Host through a reverse proxy; forwarded headers do not
|
|
95
|
+
replace it. When authentication is configured, requests without Host still pass
|
|
96
|
+
through bearer authentication.
|
|
97
|
+
|
|
98
|
+
In the 2.0.1 maintenance patch, preflight and SDK dispatch share an immutable
|
|
99
|
+
normalized policy. A configured list replaces the default browser origins;
|
|
100
|
+
include loopback explicitly if needed. Cross-origin ports must match an entry,
|
|
101
|
+
while omitted default HTTP(S) ports match their explicit 80/443 forms. A portless
|
|
102
|
+
entry permits same-authority traffic rather than every cross-port browser origin.
|
|
103
|
+
Console defaults remain HTTP loopback; Index HTTP defaults include HTTP and HTTPS
|
|
104
|
+
loopback. When no list is set, these defaults are unchanged. Invalid entries fail
|
|
105
|
+
at boot, naming the offending entry. Restart after changing the configuration.
|
|
106
|
+
|
|
107
|
+
A second middleware, `Woods::MCP::OriginGuard`, rejects requests whose `Origin` header is outside an allow-list. Requests without an `Origin` header still pass through Host validation and configured bearer authentication.
|
|
90
108
|
|
|
91
109
|
| Scenario | `WOODS_MCP_HTTP_ALLOWED_ORIGINS` | Origins accepted |
|
|
92
110
|
|--------------------|-----------------------------------------|-------------------------------------------------------------------|
|
|
93
|
-
| default | unset |
|
|
111
|
+
| default | unset | HTTP(S) loopback, matching request authority; explicit entries for cross-port origins |
|
|
94
112
|
| explicit list | `https://app.example.com` | exactly `https://app.example.com`, loopback no longer allowed |
|
|
95
113
|
| multiple origins | `https://a.example,https://b.example` | each listed origin |
|
|
96
114
|
|
|
97
115
|
`OPTIONS` preflights are answered with the matching `Access-Control-Allow-*` headers; successful responses carry `Access-Control-Allow-Origin` and `Vary: Origin`. `Access-Control-Expose-Headers: Mcp-Session-Id` appears only in legacy session mode (`WOODS_MCP_HTTP_STATELESS=0`).
|
|
98
116
|
|
|
117
|
+
### Origin configuration compatibility
|
|
118
|
+
|
|
119
|
+
These details apply to the supporting security-patch revisions described above.
|
|
120
|
+
|
|
121
|
+
The Index HTTP launcher names `WOODS_MCP_HTTP_ALLOWED_ORIGINS` and the offending
|
|
122
|
+
entry when refusing malformed origin configuration, including under a POSIX
|
|
123
|
+
(`C`) locale. It emits one diagnostic before index resolution or HTTP binding.
|
|
124
|
+
|
|
125
|
+
Configured origins are literal values, with lowercasing, one trailing slash
|
|
126
|
+
removed and HTTP(S) default-port normalization. Wildcard-looking hostnames are
|
|
127
|
+
literal hostnames, not patterns. A nonempty explicit list replaces browser-origin
|
|
128
|
+
defaults, including loopback; add the loopback browser origins you need. Console's
|
|
129
|
+
configured defaults are HTTP loopback; the Index defaults include HTTP and HTTPS
|
|
130
|
+
loopback. Loopback **Host** acceptance is separate from browser-origin acceptance.
|
|
131
|
+
|
|
132
|
+
Ruby-configured 2.x origin entries reject surrounding whitespace; 1.6.4 trims it.
|
|
133
|
+
The `woods-mcp-http` comma-separated environment setting trims each entry on both
|
|
134
|
+
lines. Use whitespace-free values for portable configuration. The exact-port and
|
|
135
|
+
same-authority rules above still apply; literal matching does not imply wildcard
|
|
136
|
+
or arbitrary cross-port access.
|
|
137
|
+
|
|
138
|
+
The guards read the actual `Origin` and `Host`; `Forwarded` and
|
|
139
|
+
`X-Forwarded-*` do not replace them. Preserve the public Host through a proxy.
|
|
140
|
+
A genuinely absent Host is accepted by the Host policy, but any supplied Origin
|
|
141
|
+
must still pass its own check. Console requests and token-configured Index
|
|
142
|
+
requests still require bearer authentication before dispatch. The loopback-only
|
|
143
|
+
Index mode without a token retains its documented unauthenticated behavior;
|
|
144
|
+
CORS preflights do not dispatch tools.
|
|
145
|
+
|
|
146
|
+
Malformed Origin/Host values receive constant `403` responses without reflecting
|
|
147
|
+
the values. Invalid Authorization receives `401` when it reaches the auth guard.
|
|
148
|
+
Malformed HTTP framing can instead receive Puma's `400` before Rack runs. Other
|
|
149
|
+
HTTP servers may reject framing at their own boundary.
|
|
150
|
+
|
|
99
151
|
### TLS termination
|
|
100
152
|
|
|
101
153
|
The server speaks plain HTTP. Any deployment beyond a single trusted host should front it with a reverse proxy that handles TLS, HTTP/2, and connection limits.
|
data/docs/MCP_SERVERS.md
CHANGED
|
@@ -56,6 +56,12 @@ Prefer the application's bundle and a project-scoped configuration:
|
|
|
56
56
|
|
|
57
57
|
`woods-mcp-start` checks that the directory and published manifest exist, then replaces itself with `woods-mcp`. It does not install dependencies or restart a crashed process.
|
|
58
58
|
|
|
59
|
+
The MCP client owns this reader process. Automatic index maintenance has its own
|
|
60
|
+
[watcher startup](WATCH_DAEMON.md#managed-development-startup): configuring MCP
|
|
61
|
+
does not enable it. A running reader sees new published generations on subsequent
|
|
62
|
+
calls without reconnecting. Verify a real edit, not only a successful connection;
|
|
63
|
+
see [automatic maintenance](AUTOMATIC_MAINTENANCE.md).
|
|
64
|
+
|
|
59
65
|
`woods_status.index.woods_version` identifies the last publisher of the served
|
|
60
66
|
manifest; `server.version` identifies the running MCP reader. Missing writer
|
|
61
67
|
provenance is `null`. See [manifest writer provenance](PUBLISHED_INDEX.md#manifest-writer-provenance).
|
|
@@ -133,12 +139,23 @@ before using `codebase_retrieve`, including after reload. The guidance grants
|
|
|
133
139
|
no extraction, configuration-change, or Console authorization. Detailed usage
|
|
134
140
|
belongs in the [agent guide](AGENT_GUIDE.md).
|
|
135
141
|
|
|
136
|
-
This addition is
|
|
142
|
+
This addition is included in Woods `2.0.0`. The SDK omits `instructions`
|
|
137
143
|
when negotiating protocol `2024-11-05`; that behavior is preserved. Older gems,
|
|
138
144
|
legacy clients, and clients that do not show server instructions can use the
|
|
139
145
|
agent guide or investigation skill. Leave protocol negotiation enabled rather
|
|
140
146
|
than pinning a newer version solely to obtain guidance.
|
|
141
147
|
|
|
148
|
+
### Structural and semantic readiness
|
|
149
|
+
|
|
150
|
+
`ready` describes the published structural index. A reachable embedding provider
|
|
151
|
+
and bootstrap state `hydrated` do not establish that semantic stores contain
|
|
152
|
+
data. Supporting readers also report `retriever.corpus`: locally known vector
|
|
153
|
+
and metadata entry counts, counts by type, and whether those stores are empty,
|
|
154
|
+
populated, or unknown. These are stored records (including chunks), not a count
|
|
155
|
+
of extracted units with verified embedding coverage. Positive counts do not
|
|
156
|
+
prove complete coverage or working provider access. See
|
|
157
|
+
[semantic corpus diagnostics](RETRIEVAL_GUIDE.md#semantic-corpus-diagnostics).
|
|
158
|
+
|
|
142
159
|
### Tools (29 — 14 registered in the packaged default)
|
|
143
160
|
|
|
144
161
|
The Index Server defines 29 schemas across core and conditional capabilities. The normal packaged executable registers the 14 tools below; the remaining schemas require the specialized wiring described afterward.
|
|
@@ -194,7 +211,7 @@ process after publishing a new embedded index.
|
|
|
194
211
|
|
|
195
212
|
### Graph-analysis pages
|
|
196
213
|
|
|
197
|
-
|
|
214
|
+
Included in Woods `2.0.0`: `graph_analysis` enforces its advertised default
|
|
198
215
|
of 20 rows per section. Pass `limit` and `offset` to page one selected `analysis`
|
|
199
216
|
or each section of `analysis: "all"`. Explicit limits also bound nested hub
|
|
200
217
|
`dependents` lists. Older servers may return every section row when `limit` is
|
|
@@ -214,7 +231,7 @@ they do not establish complete source-reference coverage.
|
|
|
214
231
|
Search responses retain `query`, `result_count`, and `results`; `result_count`
|
|
215
232
|
is the number returned, not an estimated total. The additive `completeness`
|
|
216
233
|
object describes the requested types, literal filters, and fields in the pinned
|
|
217
|
-
generation. This contract is
|
|
234
|
+
generation. This contract is included in Woods `2.0.0`.
|
|
218
235
|
|
|
219
236
|
| `reason` | `status` | `has_more` | `total_matches` |
|
|
220
237
|
|---|---|---|---|
|
|
@@ -260,16 +277,16 @@ Supporting servers expose the annotated, paginated traversal result in
|
|
|
260
277
|
stdio and HTTP servers. Read `data.total_is_exact`, `data.graph_coverage`, budget
|
|
261
278
|
counters and optional explanation witnesses there; `content[0].text` and
|
|
262
279
|
`structuredContent.text` keep the same human-readable rendering. No `format`
|
|
263
|
-
tool argument is needed or accepted. This additive data payload is
|
|
264
|
-
|
|
280
|
+
tool argument is needed or accepted. This additive data payload is included in
|
|
281
|
+
Woods `2.0.0`; verify the installed response before relying on it. Older
|
|
265
282
|
human-renderer responses can carry only text. The structured nodes and witnesses
|
|
266
283
|
cover the same page, not an additional traversal or an unpaginated graph.
|
|
267
284
|
|
|
268
285
|
Successful responses carry `graph_coverage` with `scope: "published_relationships"`,
|
|
269
286
|
`source_references: "not_exhaustive"`, and a human-readable `notice`. Text formats
|
|
270
287
|
show the same notice, including compact, root-only and empty-page responses.
|
|
271
|
-
This response metadata and the total exactness field below are
|
|
272
|
-
|
|
288
|
+
This response metadata and the total exactness field below are included in Woods
|
|
289
|
+
`2.0.0`; older servers need the same conservative interpretation.
|
|
273
290
|
|
|
274
291
|
### Dependency traversal budgets
|
|
275
292
|
|
|
@@ -310,14 +327,14 @@ for stable pages. No wall-clock deadline is used, so cutoffs are deterministic.
|
|
|
310
327
|
Budgets cover traversal work after per-generation graph loading and cache
|
|
311
328
|
preparation (JSON parsing, typed-edge normalization, node types and database
|
|
312
329
|
metadata). They do not cap that initial load, elapsed time, or total process
|
|
313
|
-
memory. These arguments are
|
|
330
|
+
memory. These arguments are included in Woods `2.0.0`; check the connected
|
|
314
331
|
server's tool schema before sending them to an older installation.
|
|
315
332
|
|
|
316
333
|
### Traversal explanations
|
|
317
334
|
|
|
318
335
|
Supporting development versions accept `explain: true` on `dependencies` and
|
|
319
|
-
`dependents`. Check the connected schema first; this option is
|
|
320
|
-
2.0.0
|
|
336
|
+
`dependents`. Check the connected schema first; this option is included in Woods
|
|
337
|
+
`2.0.0`. Omitted or false keeps the existing compact response.
|
|
321
338
|
|
|
322
339
|
The additive `explanation` object contains:
|
|
323
340
|
|
|
@@ -343,7 +360,7 @@ it describes identifier-level reachability, never a uniquely typed path.
|
|
|
343
360
|
A true value means only that identities along this witness have unambiguous
|
|
344
361
|
types. It does not establish source-reference coverage or observed execution.
|
|
345
362
|
Text labels this `witness types unambiguous=yes/no`; the JSON key and its meaning
|
|
346
|
-
remain unchanged. The text label change is
|
|
363
|
+
remain unchanged. The text label change is included in Woods `2.0.0`.
|
|
347
364
|
`types` filters retain the compact traversal's identifier-level semantics: any
|
|
348
365
|
registered type can qualify a name, while edge evidence keeps its actual source
|
|
349
366
|
owner. Multiple relationship kinds between the same endpoints remain separate.
|
data/docs/MCP_TOOL_COOKBOOK.md
CHANGED
|
@@ -356,7 +356,7 @@ column.
|
|
|
356
356
|
Use `lookup` on the returned identifier for source, actions, and routes. Search
|
|
357
357
|
`source_code` for textual matches beyond names. Supporting versions distinguish
|
|
358
358
|
exact totals from a bounded result prefix; `partial` means this is discovery,
|
|
359
|
-
not an exhaustive list. Completeness metadata is
|
|
359
|
+
not an exhaustive list. Completeness metadata is included in Woods `2.0.0`;
|
|
360
360
|
see the [search contract](MCP_SERVERS.md#search-completeness).
|
|
361
361
|
|
|
362
362
|
---
|
data/docs/MCP_WORKTREE_SETUP.md
CHANGED
|
@@ -66,7 +66,19 @@ The published payload's `manifest.json` records the extraction's `git_branch` an
|
|
|
66
66
|
|
|
67
67
|
Woods uses worktree-aware Git commands. If a present `.git` cannot be resolved, provenance is `"unknown"`; stale `GIT_BRANCH`/`GIT_SHA` values are not substituted. Those environment variables are fallbacks only when the root has no `.git` or Git is unavailable. Temporal snapshots skip an unknown SHA.
|
|
68
68
|
|
|
69
|
-
For extraction in a container,
|
|
69
|
+
For extraction in a container, mount the complete shared Git layout at the
|
|
70
|
+
original path so the worktree's `.git` pointer resolves without an override,
|
|
71
|
+
or select `/mounted-common/worktrees/<id>` with `WOODS_GIT_DIR` inside a
|
|
72
|
+
relocated complete mount. Derive `<id>` from Git metadata, not the branch name.
|
|
73
|
+
Selecting the shared root instead uses the primary checkout's HEAD, affecting
|
|
74
|
+
provenance, per-file history, and incremental paths. Follow the
|
|
75
|
+
[worktree mount and verification steps](TROUBLESHOOTING.md#git-directory-mounts-for-linked-worktrees)
|
|
76
|
+
and the [published index layout](INDEX_LAYOUT.md) when locating the manifest.
|
|
77
|
+
|
|
78
|
+
Compare the branch and exact SHA in the extraction environment with the intended
|
|
79
|
+
worktree, then verify the published manifest after full extraction. A commit
|
|
80
|
+
alone may not trigger the source-file watcher; use full `woods:extract` when
|
|
81
|
+
current Git history is required.
|
|
70
82
|
|
|
71
83
|
## Troubleshooting
|
|
72
84
|
|
data/docs/PUBLISHED_INDEX.md
CHANGED
|
@@ -111,7 +111,7 @@ The reader wraps `Woods::MCP::IndexReader` with `auto_refresh: false`; the unit
|
|
|
111
111
|
|
|
112
112
|
### Actual unit types and directory families
|
|
113
113
|
|
|
114
|
-
|
|
114
|
+
Included in Woods `2.0.0`: `unit` and `units` accept actual published
|
|
115
115
|
`graphql_type`, `graphql_mutation`, `graphql_resolver`, `graphql_query`, and
|
|
116
116
|
`gem_source` types. Enumeration preserves each unit's actual type rather than
|
|
117
117
|
labeling every GraphQL member `graphql` or every gem source `rails_source`.
|
data/docs/README.md
CHANGED
|
@@ -10,7 +10,7 @@ Woods extracts runtime-accurate Rails context and serves it to coding agents thr
|
|
|
10
10
|
| Ask an agent to install or configure Woods | [Agent setup runbook](AGENT_SETUP.md) | A safe, reviewable install with an agent handoff report |
|
|
11
11
|
| Configure an MCP client or Docker path | [MCP servers](MCP_SERVERS.md) | A working Index Server and, if authorized, an optional Console Server |
|
|
12
12
|
| Use Woods tools as an agent | [Agent guide](AGENT_GUIDE.md) | A repeatable query workflow for code context, flows, and blast radius |
|
|
13
|
-
| Keep the index current automatically | [
|
|
13
|
+
| Keep the index current automatically | [Automatic maintenance guide](AUTOMATIC_MAINTENANCE.md) | A setup with supervised indexing, automatic reader refresh, and clear hook ownership |
|
|
14
14
|
| Upgrade from Woods 1.x | [Upgrade to Woods 2.0](UPGRADING_TO_2.md) | A backed-up, re-indexed, verified v2 installation |
|
|
15
15
|
| Diagnose an error | [Troubleshooting](TROUBLESHOOTING.md) | Symptom-to-cause checks for extraction, MCP, embeddings, storage, and Docker |
|
|
16
16
|
| Contribute to Woods | [Contributing](../CONTRIBUTING.md) | A tested change with synchronized docs and plugin guidance |
|
|
@@ -41,6 +41,7 @@ and stdio or Streamable HTTP endpoints directly.
|
|
|
41
41
|
|
|
42
42
|
## Index lifecycle
|
|
43
43
|
|
|
44
|
+
- [Automatic maintenance guide](AUTOMATIC_MAINTENANCE.md): recommended development workflow and a section-by-section map of watcher, indexing, MCP, and hook documentation.
|
|
44
45
|
- [Source freshness](SOURCE_FRESHNESS.md): verify dirty source against a served generation, establish a fresh-process baseline, and understand bounded unknown results.
|
|
45
46
|
|
|
46
47
|
- [Retrieval guide](RETRIEVAL_GUIDE.md): configure embeddings and understand semantic retrieval, ranking, and token budgets.
|
data/docs/RETRIEVAL_GUIDE.md
CHANGED
|
@@ -291,16 +291,65 @@ result = retriever.retrieve("what validations does Order have?")
|
|
|
291
291
|
|
|
292
292
|
---
|
|
293
293
|
|
|
294
|
-
##
|
|
295
|
-
|
|
296
|
-
|
|
294
|
+
## Semantic corpus diagnostics
|
|
295
|
+
|
|
296
|
+
Extraction and embedding publish different data. `woods_status.ready` describes
|
|
297
|
+
the structural index; neither that flag nor bootstrap `hydrated` proves that
|
|
298
|
+
semantic retrieval has indexed records. Provider detection alone can succeed
|
|
299
|
+
before the first embedding run.
|
|
300
|
+
|
|
301
|
+
Supporting readers report `woods_status.retriever.corpus`:
|
|
302
|
+
|
|
303
|
+
- `state`: `empty`, `metadata_only`, `vectors_only`, `nonempty`, or `unknown`.
|
|
304
|
+
- `vectors` and `metadata`: each has `count`, `by_type`, and `untyped_count`.
|
|
305
|
+
Counts describe stored entries, including chunks; they are not distinct
|
|
306
|
+
extracted-unit counts or a completeness certificate. Missing type labels are
|
|
307
|
+
reported separately rather than assigned to a guessed type.
|
|
308
|
+
- Unknown counts are `null`. Diagnostics use an explicit local-store capability;
|
|
309
|
+
they do not query remote stores to discover their counts. A missing capability
|
|
310
|
+
does not mean the backend is empty or broken.
|
|
311
|
+
|
|
312
|
+
When both semantic stores are known empty, `codebase_retrieve` reports
|
|
313
|
+
`empty_index` with recovery guidance instead of presenting an empty match as
|
|
314
|
+
evidence that application code is absent. Run `woods:embed` in the application
|
|
315
|
+
with the intended provider and storage configuration, then reload or restart
|
|
316
|
+
the reader. Restart after changing provider or store configuration; reload
|
|
317
|
+
refreshes the stores of the existing retriever. Alternatively, explicitly select `WOODS_RETRIEVAL_MODE=lexical`
|
|
318
|
+
in the MCP process and restart for ranked retrieval over the published units.
|
|
319
|
+
Woods does not change retrieval mode automatically.
|
|
320
|
+
|
|
321
|
+
Metadata-only stores can still answer some keyword, direct, and graph queries;
|
|
322
|
+
they are not blocked by this diagnostic. Source-empty units deliberately retain
|
|
323
|
+
metadata without vectors, so the two counts need not match. Nonempty stores can
|
|
324
|
+
still have missing types, stale vectors, or provider failures. Inspect the
|
|
325
|
+
query result and embedding evidence before claiming coverage. The type-rank
|
|
326
|
+
table's metadata count describes the retrieval metadata store, not all units
|
|
327
|
+
in the structural index.
|
|
328
|
+
|
|
329
|
+
Older readers may omit `retriever.corpus`; record the reader revision separately
|
|
330
|
+
from the index writer version and verify the embedding artifacts directly.
|
|
331
|
+
Explicit lexical mode does not use or report semantic corpus counts.
|
|
297
332
|
|
|
298
|
-
|
|
299
|
-
- **Vector store unavailable**: vector and hybrid strategies fail at query time. Keyword and graph strategies remain available for direct calls to `SearchExecutor`.
|
|
300
|
-
- **Metadata store error**: the structural context overview (unit counts by type) is silently omitted; `Retriever#build_structural_context` rescues `StandardError` and returns `nil`. The retrieval result is still returned without the overview.
|
|
301
|
-
- **Graph store unavailable**: graph expansion in hybrid strategy produces no graph candidates; vector and keyword candidates are still ranked and returned.
|
|
333
|
+
## Degradation Tiers
|
|
302
334
|
|
|
303
|
-
|
|
335
|
+
The MCP boundary distinguishes missing configuration, empty stores, and failed
|
|
336
|
+
stores. A failure is not evidence that no application code matches:
|
|
337
|
+
|
|
338
|
+
- **No embedding provider configured:** `codebase_retrieve` reports a configuration
|
|
339
|
+
error with embedding and explicit lexical-mode options.
|
|
340
|
+
- **Both semantic stores known empty:** the tool reports `empty_index`; see the
|
|
341
|
+
[corpus diagnostics](#semantic-corpus-diagnostics) above.
|
|
342
|
+
- **Failed dump hydration:** the tool reports `degraded_index` rather than serving
|
|
343
|
+
a clean empty result. Repair or regenerate the named embedding artifact.
|
|
344
|
+
- **Query-time storage failures:** vector, metadata, and graph adapter exceptions
|
|
345
|
+
are translated into store errors and reported as `degraded_index` by MCP.
|
|
346
|
+
This includes failures building the metadata overview; that failure is not
|
|
347
|
+
silently omitted. Direct Ruby callers should handle `Woods::Retriever::StoreError`.
|
|
348
|
+
|
|
349
|
+
Provider failures and missing metadata for returned candidates have their own
|
|
350
|
+
error paths. Preserve their diagnostics; do not silently switch modes or treat
|
|
351
|
+
an exception as an empty match. Positive corpus counts do not override these
|
|
352
|
+
checks.
|
|
304
353
|
|
|
305
354
|
---
|
|
306
355
|
|
data/docs/SOURCE_FRESHNESS.md
CHANGED
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
application inputs with source bytes visible to the reader. It is separate from
|
|
5
5
|
index age, HEAD equality, daemon liveness and external database/runtime state.
|
|
6
6
|
|
|
7
|
-
This capability is
|
|
7
|
+
This capability is included in Woods `2.0.0`. Check the installed gem's
|
|
8
8
|
`woods-extract --help`, `rake -T woods:source_status`, and `woods_status` schema
|
|
9
9
|
before using it; upgrading the plugin alone does not upgrade Woods.
|
|
10
10
|
|
data/docs/TOKEN_BENCHMARK.md
CHANGED
|
@@ -13,7 +13,8 @@
|
|
|
13
13
|
> When the optional [`tokenizers`](https://github.com/ankane/tokenizers-ruby)
|
|
14
14
|
> gem is installed, the Ollama path uses the real BERT WordPiece tokenizer
|
|
15
15
|
> (`Woods::Embedding::TokenCounter`) instead of this heuristic. The 4.0
|
|
16
|
-
> divisor
|
|
16
|
+
> divisor applies to the OpenAI/default path; Ollama falls back to 1.5 when
|
|
17
|
+
> its tokenizer is unavailable.
|
|
17
18
|
|
|
18
19
|
This is a historical record of the benchmark that picked 4.0 over the
|
|
19
20
|
original 3.5 divisor. It is cited from five places in `lib/` as the evidence
|
|
@@ -35,17 +36,21 @@ for that choice, keep the numbers below intact if you edit this doc.
|
|
|
35
36
|
| 3.8 | 16.2% | 42.5% |
|
|
36
37
|
| **4.0 (shipped)** | **10.6%** | **35.4%** |
|
|
37
38
|
|
|
38
|
-
Mean chars/token across the corpus was **4.41** (range 3.94–5.42).
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
chars/token)
|
|
39
|
+
Mean chars/token across the corpus was **4.41** (range 3.94–5.42). These
|
|
40
|
+
aggregate results favor the 4.0 divisor for this sample; they do not establish
|
|
41
|
+
an upper bound on token counts. The recorded range includes values below 4.0,
|
|
42
|
+
so the heuristic can underestimate. Code lines and comment/YARD lines had
|
|
43
|
+
similar ratios (4.38 vs. 4.27 chars/token) in this sample.
|
|
44
|
+
|
|
45
|
+
The character estimate covers only the text passed to the counter. It does not
|
|
46
|
+
bound the serialized MCP response: text rendering, structured output, provenance,
|
|
47
|
+
and JSON framing can add bytes and tokens beyond that input.
|
|
43
48
|
|
|
44
49
|
## What shipped
|
|
45
50
|
|
|
46
51
|
**The divisor changed from 3.5 to 4.0.** It roughly halves the mean
|
|
47
|
-
|
|
48
|
-
|
|
52
|
+
error (26.2% → 10.6%) at zero new runtime dependencies. It remains an estimate,
|
|
53
|
+
not a guarantee that arbitrary input fits a model's token limit. The
|
|
49
54
|
constant lives in one place now (`Woods::TokenUtils::CHARS_PER_TOKEN_BY_PROVIDER`),
|
|
50
55
|
not scattered across call sites, see `lib/woods/token_utils.rb` for the
|
|
51
56
|
current definition and `docs/EMBEDDING_MODELS.md` for the Ollama-side ratio.
|
|
@@ -53,8 +58,9 @@ current definition and `docs/EMBEDDING_MODELS.md` for the Ollama-side ratio.
|
|
|
53
58
|
**tiktoken_ruby was deliberately not added as a runtime dependency.** A 10.6%
|
|
54
59
|
mean error is acceptable for chunking decisions, budget estimates, and
|
|
55
60
|
truncation; a native-extension dependency for marginal accuracy gains wasn't
|
|
56
|
-
worth it. The optional `tokenizers` gem
|
|
57
|
-
|
|
61
|
+
worth it. The optional `tokenizers` gem provides counts for its supported BERT
|
|
62
|
+
WordPiece tokenizer, not every model. Strict token-limit enforcement requires
|
|
63
|
+
the tokenizer used by the target model.
|
|
58
64
|
|
|
59
65
|
## Reproducing this benchmark
|
|
60
66
|
|