woods 2.0.0.beta3 → 2.0.0.beta4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +77 -0
- data/CONTRIBUTING.md +29 -17
- data/README.md +92 -177
- data/docs/AGENT_GUIDE.md +26 -4
- data/docs/AGENT_SETUP.md +18 -8
- data/docs/BACKEND_MATRIX.md +5 -0
- data/docs/CONFIGURATION_REFERENCE.md +74 -8
- data/docs/CONSOLE_MCP_SETUP.md +45 -2
- data/docs/DOCKER_SETUP.md +1 -1
- data/docs/EXTRACTOR_REFERENCE.md +9 -1
- data/docs/INCREMENTAL_EXTRACTION.md +30 -6
- data/docs/MCP_SERVERS.md +57 -2
- data/docs/MCP_TOOL_COOKBOOK.md +4 -4
- data/docs/MCP_WORKTREE_SETUP.md +43 -83
- data/docs/PUBLISHED_INDEX.md +17 -0
- data/docs/RETRIEVAL_GUIDE.md +24 -5
- data/docs/TROUBLESHOOTING.md +12 -13
- data/docs/UPGRADING_TO_2.md +6 -2
- data/docs/WATCH_DAEMON.md +18 -8
- data/exe/woods-mcp-start +14 -9
- data/lib/generators/woods/pgvector_generator.rb +8 -2
- data/lib/woods/agent_configuration/applier.rb +5 -3
- data/lib/woods/agent_configuration/cli.rb +2 -2
- data/lib/woods/agent_configuration/layout.rb +13 -0
- data/lib/woods/console/credential_scanner.rb +4 -3
- data/lib/woods/console/dispatch_pipeline.rb +7 -0
- data/lib/woods/console/embedded_executor.rb +31 -9
- data/lib/woods/console/sql_noise_stripper.rb +9 -7
- data/lib/woods/console/sql_table_scanner.rb +47 -7
- data/lib/woods/console/sql_validator.rb +49 -9
- data/lib/woods/console/sqlite_read_guard.rb +46 -0
- data/lib/woods/coordination/pipeline_lock.rb +3 -2
- data/lib/woods/embedding/indexer.rb +24 -14
- data/lib/woods/extractor.rb +45 -12
- data/lib/woods/extractors/declared_parent.rb +55 -0
- data/lib/woods/extractors/graphql_extractor.rb +2 -11
- data/lib/woods/extractors/lib_extractor.rb +10 -8
- data/lib/woods/extractors/mailer_extractor.rb +6 -10
- data/lib/woods/extractors/model_extractor.rb +1 -15
- data/lib/woods/extractors/poro_extractor.rb +10 -8
- data/lib/woods/extractors/shared_utility_methods.rb +22 -5
- data/lib/woods/mcp/bearer_auth.rb +2 -1
- data/lib/woods/mcp/bootstrapper.rb +17 -4
- data/lib/woods/mcp/config_resolver.rb +2 -1
- data/lib/woods/mcp/index_reader.rb +11 -2
- data/lib/woods/mcp/renderers/markdown_renderer.rb +14 -8
- data/lib/woods/mcp/renderers/plain_renderer.rb +11 -7
- data/lib/woods/mcp/server.rb +22 -28
- data/lib/woods/mcp/tool_contract.rb +1 -1
- data/lib/woods/mcp/tool_response_renderer.rb +16 -0
- data/lib/woods/mcp/traversal_evidence_text.rb +1 -1
- data/lib/woods/mcp/traversal_response.rb +22 -0
- data/lib/woods/path_dispatcher.rb +6 -5
- data/lib/woods/published_index/typed_unit_reader.rb +40 -3
- data/lib/woods/published_index.rb +2 -2
- data/lib/woods/rake_helpers.rb +2 -12
- data/lib/woods/retrieval/lexical_assembler.rb +14 -3
- data/lib/woods/retrieval/lexical_index.rb +2 -1
- data/lib/woods/session_tracer/file_store.rb +6 -1
- data/lib/woods/source_inputs/consumer_errors.rb +4 -0
- data/lib/woods/storage/pgvector.rb +6 -2
- data/lib/woods/temporal/json_snapshot_store.rb +35 -7
- data/lib/woods/version.rb +1 -1
- data/lib/woods/watch/daemon.rb +18 -4
- data/plugin/.claude-plugin/plugin.json +1 -1
- data/plugin/hooks/woods-input-rules.sh +4 -4
- data/plugin/skills/woods-agent-enable/SKILL.md +7 -1
- data/plugin/skills/woods-diagnose/SKILL.md +64 -33
- data/plugin/skills/woods-investigate/SKILL.md +51 -12
- data/plugin/skills/woods-mcp-config/SKILL.md +11 -11
- data/plugin/skills/woods-setup/SKILL.md +14 -11
- metadata +8 -5
data/docs/AGENT_GUIDE.md
CHANGED
|
@@ -129,18 +129,38 @@ Do not infer completeness from a full page or missing metadata. See the
|
|
|
129
129
|
|
|
130
130
|
## Traverse deliberately
|
|
131
131
|
|
|
132
|
-
`dependencies`
|
|
132
|
+
`dependencies` shows recorded relationships from this unit; `dependents` shows
|
|
133
|
+
recorded relationships to it. Both default to bounded breadth-first traversal and
|
|
134
|
+
accept type or relationship filters. They are not exhaustive source-reference or
|
|
135
|
+
call graphs: selective scanning can miss arbitrary method-body constant references,
|
|
136
|
+
including generic PORO and library targets. No dependents or test-only dependents
|
|
137
|
+
do not establish absence of production callers. Check source before claiming absence.
|
|
138
|
+
The `graph_coverage` notice makes this scope explicit in supporting responses;
|
|
139
|
+
that metadata is unreleased after `2.0.0.beta3`.
|
|
133
140
|
|
|
134
141
|
Start at depth 1 or 2. A deeper unfiltered traversal can obscure the direct evidence that matters. Common relationship values include associations (`belongs_to`, `has_many`, `has_one`), code references, renders, redirects, form actions, and navigation links.
|
|
135
142
|
|
|
136
143
|
Both return at most 50 nodes by default and say so with a `Showing N of M (truncated)`
|
|
137
|
-
line
|
|
144
|
+
line for completed walks. Budget-limited answers in supporting versions instead
|
|
145
|
+
say `Showing N of at least M (total unknown: <budget reason>)`. Narrow with `depth`, `types` and `via` before paging with `limit` and
|
|
138
146
|
`offset`: narrowing answers the question, paging only splits the same answer
|
|
139
147
|
across turns. In a multi-database app each row names the unit's database.
|
|
140
148
|
|
|
149
|
+
On supporting servers, inspect `structuredContent.data` for traversal nodes,
|
|
150
|
+
`graph_coverage`, exactness, budgets and optional explanation witnesses, regardless
|
|
151
|
+
of the text renderer. This packaged stdio/HTTP payload is unreleased after
|
|
152
|
+
`2.0.0.beta3`; check the actual response and use its text when structured data is
|
|
153
|
+
absent. There is no traversal `format` argument. See the
|
|
154
|
+
[response contract](MCP_SERVERS.md#dependency-graph-coverage).
|
|
155
|
+
|
|
141
156
|
A traversal can also stop at its independent node or edge budget. Treat
|
|
142
157
|
`partial`/`partial_reason` as incomplete graph evidence even on the final page;
|
|
143
|
-
paging cannot recover nodes the walk never reached.
|
|
158
|
+
paging cannot recover nodes the walk never reached. Supporting responses include
|
|
159
|
+
`total_is_exact: false` for a cutoff and true for a finished walk, independently of
|
|
160
|
+
pagination. This field is unreleased after `2.0.0.beta3`; on older servers inspect
|
|
161
|
+
`partial` directly. `nodes_total` remains the root-inclusive admitted prefix count,
|
|
162
|
+
not the full reachable total when partial. Even an exact count covers only the
|
|
163
|
+
requested root, depth, filters and published graph generation. Check the connected schema
|
|
144
164
|
before using `max_nodes`/`max_edges`, and follow the
|
|
145
165
|
[budget contract](MCP_SERVERS.md#dependency-traversal-budgets).
|
|
146
166
|
|
|
@@ -149,7 +169,9 @@ recorded source-to-target relationships and a shared shortest witness to each
|
|
|
149
169
|
row. Report `direct` relationships separately from `transitive` inferred impact.
|
|
150
170
|
Follow `parent`/`edge_id` references; `context: true` ancestors are outside the
|
|
151
171
|
current result page. Unknown labels and ambiguous candidate types stay unknown;
|
|
152
|
-
`typed_path_complete: false` does not establish a uniquely typed path.
|
|
172
|
+
`typed_path_complete: false` does not establish a uniquely typed path. True means
|
|
173
|
+
only unambiguous witness types, not complete source coverage. Supporting text
|
|
174
|
+
responses label this `witness types unambiguous` (unreleased after `2.0.0.beta3`). See the
|
|
153
175
|
[explanation contract](MCP_SERVERS.md#traversal-explanations).
|
|
154
176
|
|
|
155
177
|
Use recorded relationship labels as evidence. Do not infer execution or call
|
data/docs/AGENT_SETUP.md
CHANGED
|
@@ -40,7 +40,9 @@ If the worktree contains unrelated changes, preserve them. Do not overwrite an e
|
|
|
40
40
|
|
|
41
41
|
Use structural-only setup when the user wants code navigation, runtime Rails structure, dependencies, flows, or blast-radius analysis. Fourteen tools register in the normal packaged launch without an embedding provider.
|
|
42
42
|
|
|
43
|
-
|
|
43
|
+
If the user wants ranked discovery through `codebase_retrieve`, offer [explicit lexical mode](RETRIEVAL_GUIDE.md#embedding-free-lexical-retrieval) over the published index without a provider or embeddings. Check that the installed version supports it, set `WOODS_RETRIEVAL_MODE=lexical` in the MCP process environment, restart that server, and verify `woods_status.retriever.mode`. Keep structural-only setup as the default unless this mode is requested.
|
|
44
|
+
|
|
45
|
+
For semantic matching, discuss local Ollama or hosted OpenAI and the appropriate vector store separately; adding a provider still requires authorization. See [Backend matrix](BACKEND_MATRIX.md).
|
|
44
46
|
|
|
45
47
|
Do not infer permission to configure Console MCP from a request to “set up Woods” or “set up MCP.” The Index Server reads generated code context; the Console Server can read live data.
|
|
46
48
|
|
|
@@ -130,7 +132,7 @@ Reconnect the client and call `woods_status`. Confirm a current generation and n
|
|
|
130
132
|
|
|
131
133
|
### Managed Claude Code configuration
|
|
132
134
|
|
|
133
|
-
|
|
135
|
+
`woods-agent-config` is available from Woods `2.0.0.beta3`.
|
|
134
136
|
Check `bundle exec woods-agent-config --help` in the selected application bundle;
|
|
135
137
|
use the manual client configuration below when it is absent. The supported
|
|
136
138
|
client format is Claude Code (tested with 2.1.267).
|
|
@@ -188,8 +190,16 @@ receipt for future update/removal. Unrelated servers, hooks, settings,
|
|
|
188
190
|
instruction text, permissions, and line-ending conventions are retained;
|
|
189
191
|
changing JSON may reformat its whitespace.
|
|
190
192
|
|
|
193
|
+
Unreleased after `2.0.0.beta3`: apply and recovery coordinate on the actual
|
|
194
|
+
managed file paths, including user configuration and shared instruction files.
|
|
195
|
+
Two application roots sharing those files cannot apply overlapping plans at the
|
|
196
|
+
same time. A competing operation reports a conflict; after it finishes, create a
|
|
197
|
+
fresh preview if the saved plan's snapshots changed. Both applications keep
|
|
198
|
+
their own ownership receipts. Do not delete an active coordination lock.
|
|
199
|
+
|
|
191
200
|
Writes use atomic replacement per file and a private recovery journal beside
|
|
192
|
-
the receipt. The plan summary names
|
|
201
|
+
the receipt. The plan summary names all adjacent `.woods.lock` files, the
|
|
202
|
+
receipt `.lock`, and the `.pending` journal;
|
|
193
203
|
a lock file may remain after completion. Multiple files are not one atomic
|
|
194
204
|
transaction. An ordinary write failure restores original files when safe; an
|
|
195
205
|
interruption or concurrent edit can retain the journal. Resolve reported
|
|
@@ -203,11 +213,11 @@ is changed; reduce the selected configuration before applying.
|
|
|
203
213
|
|
|
204
214
|
Use a class known to exist in the application:
|
|
205
215
|
|
|
206
|
-
1. Call `search` to obtain its exact identifier.
|
|
207
|
-
2. Call `lookup` to confirm source and metadata are present.
|
|
216
|
+
1. Call `search` to obtain its exact identifier and type.
|
|
217
|
+
2. Call `lookup` with that identifier and type to confirm source and metadata are present.
|
|
208
218
|
3. Call `dependents` with depth 1 or 2 to confirm graph edges are queryable.
|
|
209
219
|
|
|
210
|
-
If `codebase_retrieve` reports that semantic search is disabled, that is expected for structural-only setup. Do not configure credentials merely to remove the message.
|
|
220
|
+
If `codebase_retrieve` reports that semantic search is disabled, that is expected for structural-only setup. Do not configure credentials merely to remove the message. If lexical retrieval was requested, verify its mode with `woods_status` and make one `codebase_retrieve` call against the published index.
|
|
211
221
|
|
|
212
222
|
## 8. Offer automatic index maintenance
|
|
213
223
|
|
|
@@ -263,7 +273,7 @@ Verified capabilities:
|
|
|
263
273
|
- Index Server connected: yes/no
|
|
264
274
|
- woods_status current: yes/no
|
|
265
275
|
- search/lookup/dependents checked: yes/no
|
|
266
|
-
-
|
|
276
|
+
- retrieval: disabled/lexical/semantic (provider when semantic)
|
|
267
277
|
- Console MCP: disabled/enabled (authorization)
|
|
268
278
|
- automatic structural updates: disabled/enabled (process manager)
|
|
269
279
|
|
|
@@ -274,7 +284,7 @@ Never report a capability as enabled solely because its schema exists in source.
|
|
|
274
284
|
|
|
275
285
|
## Copyable prompt for an installation agent
|
|
276
286
|
|
|
277
|
-
> Install Woods 2.x in this Rails repository using
|
|
287
|
+
> Install Woods 2.x in this Rails repository using https://github.com/lost-in-the/woods/blob/main/docs/AGENT_SETUP.md. Select a published version and follow that version's tag documentation and supported capabilities. Start with read-only preflight and preserve unrelated changes. Default to the structural Index Server; do not enable embeddings, Console MCP, HTTP transport, secrets, or purge overrides without asking me. Inspect generated files before migrating, run extraction and validation in the app's normal execution environment, configure a project-scoped MCP server in the same filesystem context as the application bundle and index, and verify `woods_status`, `search`, `lookup` with the discovered identifier and type, and `dependents`. Finish with the runbook's handoff report.
|
|
278
288
|
|
|
279
289
|
## Related guides
|
|
280
290
|
|
data/docs/BACKEND_MATRIX.md
CHANGED
|
@@ -96,6 +96,11 @@ CREATE INDEX IF NOT EXISTS idx_woods_vectors_embedding_hnsw
|
|
|
96
96
|
ON woods_vectors USING hnsw (embedding vector_cosine_ops);
|
|
97
97
|
```
|
|
98
98
|
|
|
99
|
+
**Dimension limit:** Woods uses `vector_cosine_ops` HNSW, limited to 2,000 dimensions.
|
|
100
|
+
The default 3,072-dimensional `text-embedding-3-large` output needs an explicit
|
|
101
|
+
smaller provider output width or another backend. See the
|
|
102
|
+
[pgvector configuration contract](CONFIGURATION_REFERENCE.md#pgvector-postgresql).
|
|
103
|
+
|
|
99
104
|
**Performance notes:**
|
|
100
105
|
- HNSW: ~5ms search at 10K vectors, ~20ms at 100K. Memory: ~1.5x vector size.
|
|
101
106
|
- For codebase indexing (~1000-5000 units, potentially 5000-20000 chunks), HNSW is appropriate.
|
|
@@ -208,6 +208,25 @@ config.vector_store_options = {
|
|
|
208
208
|
}
|
|
209
209
|
```
|
|
210
210
|
|
|
211
|
+
Woods uses an HNSW index over pgvector's `vector` representation, which supports
|
|
212
|
+
**1–2,000 dimensions** ([pgvector's HNSW limits](https://github.com/pgvector/pgvector#hnsw)).
|
|
213
|
+
Unreleased after `2.0.0.beta3`: the adapter rejects wider dimensions before any
|
|
214
|
+
SQL, and `woods:pgvector` rejects invalid widths before writing a migration.
|
|
215
|
+
There is no automatic vector truncation or half-precision conversion.
|
|
216
|
+
|
|
217
|
+
The default `text-embedding-3-large` output is 3,072 dimensions. With pgvector,
|
|
218
|
+
explicitly request a supported provider output width, for example:
|
|
219
|
+
|
|
220
|
+
```ruby
|
|
221
|
+
config.embedding_model = 'text-embedding-3-large'
|
|
222
|
+
config.embedding_options = { dimensions: 1536 }
|
|
223
|
+
```
|
|
224
|
+
|
|
225
|
+
Keep the provider, `vector_store_options[:dimensions]` (when set), and generated
|
|
226
|
+
migration width equal. Changing a stored width requires a compatible new table
|
|
227
|
+
or an intentional index rebuild; changing the setting does not resize old data.
|
|
228
|
+
Use another backend if you need the full 3,072-dimensional output.
|
|
229
|
+
|
|
211
230
|
Requires the pgvector extension. Run the generator to create migrations:
|
|
212
231
|
|
|
213
232
|
```bash
|
|
@@ -263,7 +282,12 @@ in the promoted dump; the index MCP server loads that snapshot at startup or
|
|
|
263
282
|
reload. Incremental embedding publishes changes to paths, dependencies, and
|
|
264
283
|
other unit metadata even when unchanged source needs no new embedding. A run
|
|
265
284
|
with no content or metadata changes keeps the existing dump and retention
|
|
266
|
-
window.
|
|
285
|
+
window. Unreleased after `2.0.0.beta3`: a full `Indexer#index_all` run replaces
|
|
286
|
+
the published corpus even when a custom caller reuses in-memory vector and
|
|
287
|
+
metadata stores. Deleted units, including metadata-only records, are removed;
|
|
288
|
+
an empty full rebuild publishes an empty dump. Failed embedding leaves the
|
|
289
|
+
previous promoted dump and checkpoint intact. Incremental purge guards remain
|
|
290
|
+
unchanged. This is a reasonable default for hosts that don't bundle `sqlite3`.
|
|
267
291
|
|
|
268
292
|
## Retrieval cache options
|
|
269
293
|
|
|
@@ -335,7 +359,7 @@ The embed run writes `woods.json` + `dumps/<ISO8601>/vectors.bin` + `metadata.ms
|
|
|
335
359
|
|
|
336
360
|
Requirements:
|
|
337
361
|
- `output_dir` must be set and readable by both the embed process and the MCP server.
|
|
338
|
-
- The MCP server must know the same `output_dir` (pass via `woods-mcp <DIR>` or set `WOODS_DIR
|
|
362
|
+
- The MCP server must know the same `output_dir` (pass via `woods-mcp <DIR>` or set `WOODS_DIR`; see MCP path precedence below).
|
|
339
363
|
|
|
340
364
|
## Presets
|
|
341
365
|
|
|
@@ -418,11 +442,26 @@ finding qualifying edges. Run extraction again after changing either setting;
|
|
|
418
442
|
MCP reads the published report. These findings remain informational, never a
|
|
419
443
|
release or architecture gate.
|
|
420
444
|
|
|
421
|
-
When the JSON snapshot fallback is in use, malformed JSON,
|
|
422
|
-
|
|
423
|
-
|
|
424
|
-
|
|
425
|
-
|
|
445
|
+
When the JSON snapshot fallback is in use, malformed JSON, invalid snapshot
|
|
446
|
+
shapes, and files that cannot be read (including concurrent retention removals)
|
|
447
|
+
are warned about and treated as absent. A snapshot needs a hexadecimal string
|
|
448
|
+
`git_sha` matching its filename; `extracted_at` may be a string, null, or omitted.
|
|
449
|
+
When present and non-null, `units` must be an object whose records are objects.
|
|
450
|
+
A malformed record invalidates the entire snapshot, rather than exposing partial
|
|
451
|
+
history. Legacy bare identifier keys, omitted/null unit collections, and optional per-unit hash
|
|
452
|
+
fields remain supported; timestamp strings are not restricted to a new format.
|
|
453
|
+
Unit history limits count matching unit records, not the most recent snapshots
|
|
454
|
+
searched (JSON fallback correction unreleased after `2.0.0.beta3`). A unit
|
|
455
|
+
missing from newer snapshots can still have retained history.
|
|
456
|
+
Snapshot lists and unit history omit unusable files; direct lookup returns no
|
|
457
|
+
snapshot, and a diff with an unavailable snapshot returns empty added, modified,
|
|
458
|
+
and deleted lists. An empty diff in this case is not proof that nothing changed.
|
|
459
|
+
New captures compare against the latest usable snapshot. Unusable SHA-named files
|
|
460
|
+
still count toward retention and are pruned first when the limit is exceeded;
|
|
461
|
+
reading alone does not delete them. Valid legacy snapshots with null or omitted
|
|
462
|
+
timestamps are retained ahead of corrupt files, then treated as oldest among
|
|
463
|
+
usable snapshots. Direct lookup and diff still reject invalid
|
|
464
|
+
caller-supplied SHA paths with an argument error.
|
|
426
465
|
|
|
427
466
|
`incremental_blast_radius_depth` is unbounded by default because a unit two hops
|
|
428
467
|
out really can have content that depends on the changed file. An STI grandchild
|
|
@@ -473,6 +512,18 @@ config.session_store = Woods::SessionTracer::FileStore.new(
|
|
|
473
512
|
config.session_exclude_paths = ['/health', '/metrics', '/assets']
|
|
474
513
|
```
|
|
475
514
|
|
|
515
|
+
### File session retention
|
|
516
|
+
|
|
517
|
+
`FileStore` accepts `ttl:` in seconds (default `nil`, expiration disabled),
|
|
518
|
+
`max_sessions:` (default `1000`), and `max_requests_per_session:` (default `1000`).
|
|
519
|
+
TTL expires a file when the store clock reaches its modification time plus the
|
|
520
|
+
TTL. Recording after expiry starts a fresh history; expired events are discarded
|
|
521
|
+
before appending or migrating legacy filenames, under the same store lock.
|
|
522
|
+
When legacy and encoded files coexist, each expires independently before any
|
|
523
|
+
surviving histories are merged. Clearing a session is idempotent for supported
|
|
524
|
+
IDs, including Unicode and punctuation, and removes both filename formats when
|
|
525
|
+
applicable.
|
|
526
|
+
|
|
476
527
|
### Redis session index compatibility
|
|
477
528
|
|
|
478
529
|
`RedisStore` lists and clears both legacy SET indexes and recency ZSET indexes.
|
|
@@ -652,7 +703,8 @@ These variables are read by the gem and its MCP servers at runtime. They complem
|
|
|
652
703
|
| Variable | Default | Purpose |
|
|
653
704
|
|----------|---------|---------|
|
|
654
705
|
| `WOODS_RETRIEVAL_MODE` | `semantic` | Explicit packaged MCP retrieval mode: `semantic` or `lexical`. Lexical reads extraction unit JSON without provider autodetection, credentials or vector artifacts. |
|
|
655
|
-
| `WOODS_DIR` |
|
|
706
|
+
| `WOODS_DIR` | unset | MCP extraction-index path, after a positional argument and before `WOODS_OUTPUT`. See precedence below. |
|
|
707
|
+
| `WOODS_OUTPUT` | unset | MCP index-path fallback when neither a positional path nor `WOODS_DIR` is set; unreleased after `2.0.0.beta3`. |
|
|
656
708
|
| `WOODS_REQUIRE_INDEX` | unset | Set to `"1"` to fail closed: the server refuses to boot (raises `MissingArtifact`) unless a real index (`woods.json`) is present. By default an extract-only host boots in pattern/structural mode without it. Explicit lexical mode requires a valid published extraction index, not `woods.json`. |
|
|
657
709
|
| `WOODS_ALLOW_AUTODETECT` | unset | **Deprecated no-op.** Auto-detect is now the default; accepted for backward compatibility only. |
|
|
658
710
|
| `WOODS_SEARCH_MAX_SCAN` | `500` | Cap on unit files loaded during a phase-2 (metadata/source_code) `search`. Hitting the cap sets `partial: true` in the response. |
|
|
@@ -668,6 +720,20 @@ These variables are read by the gem and its MCP servers at runtime. They complem
|
|
|
668
720
|
| `WOODS_QDRANT_URL`, `WOODS_QDRANT_COLLECTION`, `WOODS_QDRANT_API_KEY` | n/a | Override/require Qdrant connection settings when a pgvector/Qdrant-backed index is served outside its host application (no `Woods.configuration` available). |
|
|
669
721
|
| `WOODS_PG_URL` | n/a | Required when a pgvector-backed index is served outside its host application. |
|
|
670
722
|
|
|
723
|
+
**MCP index path precedence (unreleased after `2.0.0.beta3`):** positional
|
|
724
|
+
argument → `WOODS_DIR` → `WOODS_OUTPUT` → current directory for `woods-mcp`
|
|
725
|
+
and `woods-mcp-http`. `woods-mcp-start` still requires one of the first three;
|
|
726
|
+
it never silently selects the current directory. An explicitly empty
|
|
727
|
+
`WOODS_DIR` remains an invalid override rather than falling through. Earlier
|
|
728
|
+
versions accept the positional path or `WOODS_DIR`; use an explicit path for
|
|
729
|
+
portable client configuration.
|
|
730
|
+
|
|
731
|
+
Paths are resolved in the MCP process's working directory and filesystem.
|
|
732
|
+
A published index can have `generation.json` pointing to a payload's
|
|
733
|
+
`manifest.json`; a root `manifest.json` is only the legacy flat layout. If
|
|
734
|
+
startup cannot find a manifest, check the reported directory and point at the
|
|
735
|
+
existing index before deciding another extraction is needed.
|
|
736
|
+
|
|
671
737
|
### Rake tasks
|
|
672
738
|
|
|
673
739
|
| Variable | Default | Purpose |
|
data/docs/CONSOLE_MCP_SETUP.md
CHANGED
|
@@ -198,6 +198,12 @@ the Rails server environment. The middleware stack registers automatically via
|
|
|
198
198
|
the gem's Railtie and requires `Authorization: Bearer <token>` on every Console
|
|
199
199
|
request. Missing or incorrect tokens receive `401 Unauthorized`.
|
|
200
200
|
|
|
201
|
+
The HTTP authentication scheme is ASCII case-insensitive (`Bearer`, `bearer`,
|
|
202
|
+
or `BEARER`); the token remains case-sensitive and must match exactly after one
|
|
203
|
+
space. This applies to both Console HTTP and `woods-mcp-http`. Case-insensitive
|
|
204
|
+
scheme support is unreleased after `2.0.0.beta3`; use the canonical `Bearer`
|
|
205
|
+
spelling in client configuration for compatibility with earlier releases.
|
|
206
|
+
|
|
201
207
|
For non-loopback access, `console_mcp_allowed_origins` must include the public
|
|
202
208
|
Rails/MCP host. If a browser-based client sends an `Origin` header from a
|
|
203
209
|
different host, include that exact origin too. This allow-list controls both
|
|
@@ -760,12 +766,12 @@ Each transaction sets a statement timeout before any query runs. The default is
|
|
|
760
766
|
|
|
761
767
|
`SqlValidator` rejects non-read-only SQL at the string level, before any database interaction.
|
|
762
768
|
|
|
763
|
-
Validation runs **once**, inside the executor, with the dialect of the live adapter. There is deliberately no earlier dialect-blind pre-check in the tool handler: a validator built without a dialect is the conservative
|
|
769
|
+
Validation runs **once**, inside the executor, with the dialect of the live adapter. There is deliberately no earlier dialect-blind pre-check in the tool handler: a validator built without a dialect is the conservative union of supported dialects, and running it first meant a MySQL host rejected statements whose `\'`/backtick grammar produces a spuriously forbidden PostgreSQL view — the adapter-aware acceptance below could never be reached on a real transport. The executor raises `SqlValidationError` for anything it refuses, which the dispatch pipeline renders as a tool error, so nothing is ungated.
|
|
764
770
|
|
|
765
771
|
|
|
766
772
|
- **Allowed prefixes:** `SELECT`, `WITH...SELECT`, and plain `EXPLAIN`. `EXPLAIN ANALYZE` is rejected, it executes the query rather than just planning it (both the whitespace and `EXPLAIN (ANALYZE, …)` option-list spellings).
|
|
767
773
|
- **Rejected prefixes:** `INSERT`, `UPDATE`, `DELETE`, `MERGE`, `DROP`, `ALTER`, `TRUNCATE`, `CREATE`, `GRANT`, `REVOKE`
|
|
768
|
-
- **Rejected anywhere in query:** `UNION`, `INTO`, `COPY`; row-lock clauses (`FOR UPDATE`, `FOR NO KEY UPDATE`, `FOR SHARE`, `FOR KEY SHARE`, `FOR UPDATE NOWAIT`/`SKIP LOCKED`, MySQL `LOCK IN SHARE MODE`) — these take live row locks even inside the rolled-back transaction. The lock check is adapter-aware: `console_sql` validates with the active adapter's dialect, including MySQL double-quoted strings/backtick identifiers and PostgreSQL quoted identifiers/E-strings. Unknown adapters conservatively scan
|
|
774
|
+
- **Rejected anywhere in query:** `UNION`, `INTO`, `COPY`; row-lock clauses (`FOR UPDATE`, `FOR NO KEY UPDATE`, `FOR SHARE`, `FOR KEY SHARE`, `FOR UPDATE NOWAIT`/`SKIP LOCKED`, MySQL `LOCK IN SHARE MODE`) — these take live row locks even inside the rolled-back transaction. The lock check is adapter-aware: `console_sql` validates with the active adapter's dialect, including MySQL double-quoted strings/backtick identifiers and PostgreSQL quoted identifiers/E-strings. Unknown adapters conservatively scan all supported normalizations. Every view is scanned under both MySQL executable-comment (`/*!...*/`) semantics, so `#` comments and version-guarded comments cannot split a clause apart.
|
|
769
775
|
- **Function allowlist (the authoritative function control):** every function-call-shaped identifier must appear in `ALLOWED_FUNCTIONS`, a conservative set of pure read-only functions (aggregates, window functions, string/number/date/JSON readers) kept portable across MySQL, PostgreSQL, and SQLite. Anything else is rejected by name, quoted forms (`"pg_terminate_backend"(…)`) included. This is an allowlist because a denylist cannot enumerate every side-effecting function (`nextval`, `pg_advisory_lock`, `pg_terminate_backend`, …). A legacy `DANGEROUS_FUNCTIONS` denylist (`pg_sleep`, `lo_import`, `lo_export`, `pg_read_file`, `pg_write_file`, `load_file`, `sleep`, `benchmark`) still runs first as belt-and-suspenders.
|
|
770
776
|
- **Rejected patterns:** multiple statements (semicolons), writable CTEs (every `AS (...)` body is checked, so a writable CTE in any WITH position is refused — `WITH a AS (SELECT 1), b AS (DELETE FROM users RETURNING *) SELECT * FROM b`), a CTE list attached to top-level DML (`WITH a AS (SELECT 1) DELETE FROM users RETURNING *`), comment-hidden injections
|
|
771
777
|
|
|
@@ -853,7 +859,44 @@ quote and comment rules. MySQL also reads the executing session's `ANSI_QUOTES`
|
|
|
853
859
|
and `NO_BACKSLASH_ESCAPES` settings for validation, protected-column scanning, and
|
|
854
860
|
table gating; adjacent subtraction operators are not assumed to begin a comment.
|
|
855
861
|
Direct scanner callers without session settings use conservative quote-mode scans.
|
|
862
|
+
SQLite read SQL accepts simple ASCII bare, double-quoted, or backtick identifiers
|
|
863
|
+
(letters, digits, and underscores, starting with a letter or underscore).
|
|
864
|
+
Use whitespace after `FROM` and `JOIN`, and use `SELECT` subqueries rather than
|
|
865
|
+
parenthesized table groups. Bracket-quoted and string-quoted names, quoted names
|
|
866
|
+
containing punctuation, and unsupported table-reference syntax are refused before
|
|
867
|
+
execution, because they cannot be reliably checked against the configured table
|
|
868
|
+
policy. Ordinary string literals remain supported. This restriction is part of
|
|
869
|
+
2.0.0.beta4; keep read tools disabled on older versions when this policy is
|
|
870
|
+
needed. Confirm that the release is available before selecting it. The Rails-version
|
|
871
|
+
integration lane checks these boundaries on real SQLite.
|
|
872
|
+
|
|
856
873
|
The contributor live-backend lane exercises these boundaries
|
|
857
874
|
through Console requests against PostgreSQL and MySQL. Keep read tools disabled
|
|
858
875
|
unless live SQL access is needed, and retain the configured blocked-table and
|
|
859
876
|
redaction policies when diagnosing a rejected request.
|
|
877
|
+
|
|
878
|
+
## Console policy corrections in 2.0.0.beta4
|
|
879
|
+
|
|
880
|
+
In `2.0.0.beta4`, the default model-reading tools check the resolved relation
|
|
881
|
+
against `console_blocked_tables` before fetching records or counts. This includes
|
|
882
|
+
application-defined default scopes and the parent lookup for association counts.
|
|
883
|
+
The checked relation is reused for execution so a dynamic default scope is not
|
|
884
|
+
resolved twice. These checks do not change which Console tools are enabled.
|
|
885
|
+
|
|
886
|
+
Response handling redacts protected fields before invoking serializers, converts
|
|
887
|
+
the remaining response to JSON-compatible values, then redacts and scans that
|
|
888
|
+
normalized tree before either JSON or Markdown rendering. Symbol values and custom
|
|
889
|
+
JSON serializers therefore receive the same credential checks as ordinary strings.
|
|
890
|
+
Custom values in Markdown now use their JSON-compatible representation. Numbers,
|
|
891
|
+
booleans, nulls, and ordinary record shapes retain their existing meanings.
|
|
892
|
+
|
|
893
|
+
For SQLite SQL, keyword spellings receive function-policy exceptions only where
|
|
894
|
+
supported query grammar requires them. PostgreSQL reserved-keyword grammar is
|
|
895
|
+
preserved. Unsupported parenthesized offset expressions on other dialects may
|
|
896
|
+
require a plain numeric offset. Keep the configured access and credential policies
|
|
897
|
+
in place when adjusting a query.
|
|
898
|
+
|
|
899
|
+
These corrections require `2.0.0.beta4` or a reviewed development revision that
|
|
900
|
+
contains them. Confirm that a patched release is available before selecting it.
|
|
901
|
+
On affected versions, disable Console where these policies are required; Index MCP
|
|
902
|
+
can stay enabled because it reads the published code index separately.
|
data/docs/DOCKER_SETUP.md
CHANGED
|
@@ -468,6 +468,6 @@ To register `console_sql` and `console_query`, enable
|
|
|
468
468
|
|
|
469
469
|
### Woods MCP tools not available in a git worktree
|
|
470
470
|
|
|
471
|
-
|
|
471
|
+
A separate Claude Code session launched in a worktree may have different project-scoped MCP registrations. Subagents inherit the parent session's MCP tools, subject to tool restrictions; changing their working directory does not select a different Woods index. Check registration, the Compose service's mounted checkout, and the served index using [MCP worktree setup](MCP_WORKTREE_SETUP.md).
|
|
472
472
|
|
|
473
473
|
See [CONSOLE_MCP_SETUP.md](CONSOLE_MCP_SETUP.md) for detailed console server documentation.
|
data/docs/EXTRACTOR_REFERENCE.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Woods ships **35 extractor classes** producing **39 distinct unit types**: one for each meaningful category of Rails code. This doc covers what each extractor captures, how to configure them, and the shape of the data they produce.
|
|
4
4
|
|
|
5
|
-
> **Counts explained.** `lib/woods/extractors/` contains
|
|
5
|
+
> **Counts explained.** `lib/woods/extractors/` contains 35 extractor classes (each ending in `_extractor.rb`) plus supporting utilities such as `shared_utility_methods`, `shared_dependency_scanner`, `callback_analyzer`, `behavioral_profile`, `route_helper_resolver`, `ast_source_extraction`, `source_nesting`, and `declared_parent`. The 39 unit types comes from some extractors emitting multiple categories, `GraphQLExtractor` alone produces four (`graphql_type`, `graphql_mutation`, `graphql_resolver`, `graphql_query`), and `RailsSourceExtractor` produces both `rails_source` and `gem_source`. Supporting utilities enrich existing extractors (callback side-effects, behavioral config, AST-based source slicing, nested-namespace resolution) but are not themselves extractors and do not appear in the unit type enumeration. The authoritative mapping is `Woods::Extractor::TYPE_TO_EXTRACTOR_KEY` in `lib/woods/extractor.rb`.
|
|
6
6
|
|
|
7
7
|
---
|
|
8
8
|
|
|
@@ -65,6 +65,7 @@ Every extractor returns `Array<ExtractedUnit>`. An `ExtractedUnit` is a self-con
|
|
|
65
65
|
- Reads Rails' per-event callback chains (`_save_callbacks`, `_create_callbacks`, and the other lifecycle events), preserving each chain's order and each entry's `kind`, filter and conditions. The public `type` combines kind and event, such as `before_save` or `after_create`; Rails' separate `before_commit` event is reported as `before_commit`, not `before_before_commit`. `callback_count` equals the emitted callback list's length. The list includes framework-registered callbacks; it is runtime metadata, not an application-only filter or a cross-event execution trace. Regenerate the index after upgrading to pick up corrected callback metadata.
|
|
66
66
|
- Proc/lambda filters, including Rails-generated association callbacks, use stable source-site labels in both metadata and callback chunks: `#<Proc app/models/post.rb:12>` (or `lambda`). App paths are relative to `Rails.root`; external paths are retained and native procs use `native`. Rails 6's numeric filter identity is resolved through `raw_filter`. These labels describe location and callable kind, not captured closure state; callbacks are never executed. Model condition labels retain their existing format.
|
|
67
67
|
- Default callback-object representations omit process addresses: an instance becomes `#<CleanupCallback>`, an anonymous class becomes `#<Class>`, and its instance becomes `#<#<Class>>`. Anonymous namespace prefixes are normalized too (for example, `#<Module>::CleanupCallback`). Named classes and custom `to_s` labels retain their text. These labels do not distinguish arbitrary object state; separate registered callbacks remain separate entries even when their descriptive labels match. Controller object-filter formatting is unchanged.
|
|
68
|
+
- Direct Proc/lambda validation option values (for example inclusion/exclusion membership or message callables) use the same stable kind/source-site labels without execution. Validation order and duplicates, condition formats (`if`, `unless`, `on`), and non-Proc values are unchanged. Nested arrays/hashes are not recursively normalized, and labels do not serialize captured closure state. Run a full extraction after upgrading to refresh retained validation metadata; the stored schema is unchanged.
|
|
68
69
|
- Callback side-effects are analyzed via `CallbackAnalyzer`: detects columns written (`self.col =`), jobs enqueued (`perform_later`), and services called
|
|
69
70
|
- Reflects model class and instance methods after reading the schema, so Rails schema-loading optimizations produce the same method metadata in cold and warmed runs. Application-defined constructors remain visible; Rails versions that install an optimized singleton `new` during schema loading consistently include it in `class_methods`.
|
|
70
71
|
- Automatically skips HABTM join models and anonymous classes
|
|
@@ -255,6 +256,10 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
|
|
|
255
256
|
- Discovery and direct extraction accept only mailers backed by an existing app-owned source file, excluding dependency mailers and fabricated convention paths.
|
|
256
257
|
- Each mailer action corresponds to an email template, template paths are recorded in metadata
|
|
257
258
|
- Extracts `default from:`, `layout`, and per-action subject patterns
|
|
259
|
+
- Action names are sorted consistently in metadata, the generated header, template discovery and action chunks. Callback chain order and duplicate registrations are preserved.
|
|
260
|
+
- Direct Proc-valued defaults and Proc callback filters use source-location/kind labels without executing them; application paths are relative to `Rails.root`. Default containers, literal strings and non-Proc values retain their existing types. This does not serialize closure captures or recursively normalize arbitrary nested objects.
|
|
261
|
+
- Object callback filters use descriptive labels with addresses removed only from Ruby's default representation, including anonymous classes and namespaces. Custom labels and literal hexadecimal text are preserved. Labels do not serialize callback object state, and extraction never invokes callbacks.
|
|
262
|
+
- After upgrading, run a full extraction to refresh retained mailer units. Stabilized headers and callable labels can cause a one-time source-hash change; stored index schemas are unchanged.
|
|
258
263
|
|
|
259
264
|
---
|
|
260
265
|
|
|
@@ -382,6 +387,7 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
|
|
|
382
387
|
- Scans `app/models` for files that don't define an `ActiveRecord::Base` descendant
|
|
383
388
|
- Common examples: value objects, form objects placed in `app/models`, domain structs
|
|
384
389
|
- Excludes concerns (those go to ConcernExtractor)
|
|
390
|
+
- `parent_class` and the generated Parent annotation describe the selected unit declaration only. Nested or sibling classes cannot supply its parent. An implicit `Object` parent, a dynamic superclass expression, or unparseable source produces `nil`; explicit constant-path parents retain their source names.
|
|
385
391
|
|
|
386
392
|
---
|
|
387
393
|
|
|
@@ -421,6 +427,7 @@ class PageView < AnalyticsRecord; end # metadata[:database] => "analytics"
|
|
|
421
427
|
- Produces unit types: `graphql_type`, `graphql_mutation`, `graphql_resolver`, `graphql_query`
|
|
422
428
|
- Extracts field metadata (types, descriptions, complexity, arguments), authorization patterns (Pundit, CanCan, `authorized?`), and dependencies on models/services
|
|
423
429
|
- Since all GraphQL units come from one extractor, incremental re-extraction handles them via `extract_graphql_file`
|
|
430
|
+
- `parent_class` and summary chunks describe the selected declaration's explicit constant-path superclass, preserving its written qualification. Nested or sibling declarations and literal text cannot supply a parent. Implicit Object, module interfaces, dynamic superclass expressions, unavailable source, and invalid source have no declared parent (`null` metadata; `unknown` in summaries). This is source declaration metadata, not resolved runtime ancestry.
|
|
424
431
|
|
|
425
432
|
**Example output (abbreviated):**
|
|
426
433
|
|
|
@@ -677,6 +684,7 @@ Every app-owned unit under a package root carries `metadata[:package]` with the
|
|
|
677
684
|
**Key details:**
|
|
678
685
|
- Excludes `lib/tasks/` (covered by RakeTaskExtractor) and `lib/generators/`
|
|
679
686
|
- File-based scanning; no assumption about class hierarchy
|
|
687
|
+
- `parent_class` and the generated Parent annotation describe the selected unit declaration only. Nested or sibling classes cannot supply its parent. An implicit `Object` parent, a dynamic superclass expression, or unparseable source produces `nil`; explicit constant-path parents retain their source names.
|
|
680
688
|
|
|
681
689
|
---
|
|
682
690
|
|
|
@@ -98,6 +98,23 @@ including inlined code and callback analysis. Multiple runtime mixins sharing a
|
|
|
98
98
|
file retain separate identities and refresh all their includers. Run a full extraction after upgrading
|
|
99
99
|
to populate these previously missing source mappings.
|
|
100
100
|
|
|
101
|
+
## Handled source errors and retry
|
|
102
|
+
|
|
103
|
+
Unreleased after `2.0.0.beta3`: when an incremental extraction or named refresh
|
|
104
|
+
records a handled consumer error (for example malformed locale or schedule
|
|
105
|
+
YAML), it raises `Woods::ExtractionError` before publishing. Empty output from
|
|
106
|
+
that failed consumer does not authorize replacing or deleting its last-good
|
|
107
|
+
units. The published generation and source provenance remain unchanged,
|
|
108
|
+
including when other files in the batch extracted successfully.
|
|
109
|
+
|
|
110
|
+
Fix the source error named in the extraction log, then retry the **complete
|
|
111
|
+
batch**, or the same named refresh. The watch daemon reports degraded and keeps
|
|
112
|
+
the failed batch pending for retry. An error on one file does not mark a later
|
|
113
|
+
successful file as failed, but the batch still cannot publish until all handled
|
|
114
|
+
errors are resolved. This does not change full extraction's existing tolerance
|
|
115
|
+
for handled consumer errors; its source-freshness report marks those scopes
|
|
116
|
+
unverified.
|
|
117
|
+
|
|
101
118
|
## What a run does, in order
|
|
102
119
|
|
|
103
120
|
`Extractor#extract_changed` is order-sensitive; each step exists because of the
|
|
@@ -109,8 +126,8 @@ step before it.
|
|
|
109
126
|
2. **Reconcile changed paths.** Every changed path that still exists is handed
|
|
110
127
|
to the file-based extractors that claim it (`PathDispatcher`), and units the
|
|
111
128
|
path no longer produces are dropped. This is what indexes a file the index
|
|
112
|
-
has never seen, and what
|
|
113
|
-
|
|
129
|
+
has never seen, and what removes definitions deleted from a surviving
|
|
130
|
+
source file. Multi-file Rake tasks use wholesale reconciliation below.
|
|
114
131
|
3. **Re-extract the rest of the blast radius**: units whose own file did not
|
|
115
132
|
change but which depend on something that did.
|
|
116
133
|
4. **Reconcile class-based types** against each extractor's
|
|
@@ -208,7 +225,6 @@ automatically.
|
|
|
208
225
|
| `config/locales/**/*.yml` | i18n |
|
|
209
226
|
| `config/initializers`, `config/environments` | configurations |
|
|
210
227
|
| `db/migrate/*.rb` (top level only) | migrations |
|
|
211
|
-
| `lib/tasks/**/*.rake` | rake_tasks |
|
|
212
228
|
| `lib/**/*.rb` (outside `tasks/`, `generators/`) | libs |
|
|
213
229
|
| `spec/**/*_spec.rb`, `test/**/*_test.rb` | test_mappings |
|
|
214
230
|
|
|
@@ -219,12 +235,13 @@ serializer and decorator extractors, and all matching rules run.
|
|
|
219
235
|
### Wholesale re-runs
|
|
220
236
|
|
|
221
237
|
`PathDispatcher.whole_app_rules` → `Extractor::WHOLE_APP_EXTRACTORS`. These
|
|
222
|
-
extractors
|
|
223
|
-
|
|
238
|
+
extractors need a complete runtime or directory view, even when a low-level
|
|
239
|
+
per-file reader exists. In an already-booted process re-running them is
|
|
224
240
|
cheap, which is what makes wholesale replacement the right shape.
|
|
225
241
|
|
|
226
242
|
| Trigger | Re-runs |
|
|
227
243
|
|---|---|
|
|
244
|
+
| `lib/tasks/**/*.rake` | rake_tasks (all definitions of every task) |
|
|
228
245
|
| `config/routes.rb`, `config/routes/**` | routes, engines, **and** controllers, mailers, components, view components, view templates |
|
|
229
246
|
| `Gemfile.lock` | engines, middleware, rails_source (gated by `include_framework_sources`) |
|
|
230
247
|
| `config/application.rb`, `config/initializers/**`, `config/environments/**` | middleware |
|
|
@@ -235,7 +252,14 @@ cheap, which is what makes wholesale replacement the right shape.
|
|
|
235
252
|
| `db/views/**/*.sql` | database_views |
|
|
236
253
|
| any `package.yml`, `packwerk.yml` | packages |
|
|
237
254
|
|
|
238
|
-
|
|
255
|
+
Four of these deserve a note:
|
|
256
|
+
|
|
257
|
+
- **Rake tasks merge definitions across files.** Any changed or deleted `.rake`
|
|
258
|
+
file reruns the task extractor over all task files. Removing the primary
|
|
259
|
+
definition preserves surviving definitions; removing a secondary definition
|
|
260
|
+
drops its source and dependencies from the shared unit. After upgrading from
|
|
261
|
+
the old per-file rules, run a full extraction to establish a source-freshness
|
|
262
|
+
baseline with the new rule fingerprint.
|
|
239
263
|
|
|
240
264
|
- **Routes cascade.** `ROUTE_CONSUMER_EXTRACTORS` embed the route table, controllers write each action's routes into unit metadata and into the action
|
|
241
265
|
chunks, and everything using `RouteHelperResolver` resolves navigation edges
|
data/docs/MCP_SERVERS.md
CHANGED
|
@@ -180,7 +180,7 @@ failure boundary; they do not prove uniqueness or become `ambiguous_identity` er
|
|
|
180
180
|
Use `depth: 0` for the request timeline, or inspect candidates with typed `lookup`
|
|
181
181
|
calls. Re-extraction does not remove a legitimate cross-type collision. Successful
|
|
182
182
|
traces retain their existing identifiers and response shape; target identity has
|
|
183
|
-
not been migrated globally. These corrections are
|
|
183
|
+
not been migrated globally. These corrections are available in `2.0.0.beta3`.
|
|
184
184
|
|
|
185
185
|
The server also exposes MCP resources and resource templates for indexed units. Tool descriptions returned by MCP are the parameter-level source of truth; [Agent guide](AGENT_GUIDE.md) explains selection strategy.
|
|
186
186
|
|
|
@@ -192,6 +192,23 @@ error and continues serving the previous aligned generation; it never swaps in a
|
|
|
192
192
|
partial or empty replacement. Grant write access for live reloads, or restart the MCP
|
|
193
193
|
process after publishing a new embedded index.
|
|
194
194
|
|
|
195
|
+
### Graph-analysis pages
|
|
196
|
+
|
|
197
|
+
Unreleased after `2.0.0.beta3`: `graph_analysis` enforces its advertised default
|
|
198
|
+
of 20 rows per section. Pass `limit` and `offset` to page one selected `analysis`
|
|
199
|
+
or each section of `analysis: "all"`. Explicit limits also bound nested hub
|
|
200
|
+
`dependents` lists. Older servers may return every section row when `limit` is
|
|
201
|
+
omitted; pass a limit explicitly when supporting both versions.
|
|
202
|
+
|
|
203
|
+
JSON responses retain `<section>_total`, `<section>_offset` (when positive), and
|
|
204
|
+
`<section>_truncated: true` whenever a page omits rows before or after it. Markdown,
|
|
205
|
+
plain, and Claude responses show the same total and offset on last and empty
|
|
206
|
+
pages. For example, offset 20 with limit 5 over 25 published orphans shows
|
|
207
|
+
`5 of 25 from offset 20`; offset 100 shows `0 of 25 from offset 100`. An empty
|
|
208
|
+
page does not mean the section has no findings. These totals describe the
|
|
209
|
+
published report arrays, which can themselves be bounded during extraction;
|
|
210
|
+
they do not establish complete source-reference coverage.
|
|
211
|
+
|
|
195
212
|
### Search completeness
|
|
196
213
|
|
|
197
214
|
Search responses retain `query`, `result_count`, and `results`; `result_count`
|
|
@@ -229,6 +246,31 @@ Detected missing, unreadable, or corrupt artifacts remain `isError: true` with
|
|
|
229
246
|
`has_more`, `total_matches`, and `matched_lower_bound`; no successful empty
|
|
230
247
|
result is substituted. Inspect `woods_status` and run `woods:validate`.
|
|
231
248
|
|
|
249
|
+
### Dependency graph coverage
|
|
250
|
+
|
|
251
|
+
`dependencies` and `dependents` return relationships recorded in the published
|
|
252
|
+
index, not an exhaustive call graph or source-reference index. Extraction combines
|
|
253
|
+
runtime reflection with selective source scanning; arbitrary method-body constant
|
|
254
|
+
references (including references to generic PORO and library classes) may have no
|
|
255
|
+
edge. No dependents, a test-only dependent, or a completed traversal does not prove
|
|
256
|
+
there are no production callers. Verify important absence claims in source.
|
|
257
|
+
|
|
258
|
+
Supporting servers expose the annotated, paginated traversal result in
|
|
259
|
+
`structuredContent.data` for every renderer, including the default packaged
|
|
260
|
+
stdio and HTTP servers. Read `data.total_is_exact`, `data.graph_coverage`, budget
|
|
261
|
+
counters and optional explanation witnesses there; `content[0].text` and
|
|
262
|
+
`structuredContent.text` keep the same human-readable rendering. No `format`
|
|
263
|
+
tool argument is needed or accepted. This additive data payload is unreleased
|
|
264
|
+
after `2.0.0.beta3`; verify the installed response before relying on it. Older
|
|
265
|
+
human-renderer responses can carry only text. The structured nodes and witnesses
|
|
266
|
+
cover the same page, not an additional traversal or an unpaginated graph.
|
|
267
|
+
|
|
268
|
+
Successful responses carry `graph_coverage` with `scope: "published_relationships"`,
|
|
269
|
+
`source_references: "not_exhaustive"`, and a human-readable `notice`. Text formats
|
|
270
|
+
show the same notice, including compact, root-only and empty-page responses.
|
|
271
|
+
This response metadata and the total exactness field below are unreleased after
|
|
272
|
+
Woods `2.0.0.beta3`; older servers need the same conservative interpretation.
|
|
273
|
+
|
|
232
274
|
### Dependency traversal budgets
|
|
233
275
|
|
|
234
276
|
`dependencies` and `dependents` walk breadth-first in stored graph order. The
|
|
@@ -251,7 +293,16 @@ Exact-budget walks that finish all requested work are complete and have no
|
|
|
251
293
|
`limit` (default 50) and `offset` only page that discovered result; they never
|
|
252
294
|
change the walk budget or depth. On a partial traversal, `nodes_total`, when
|
|
253
295
|
present for pagination, counts the discovered prefix, **not the full reachable
|
|
254
|
-
graph**.
|
|
296
|
+
graph**. Every successful response includes `total_is_exact`: false for a
|
|
297
|
+
budget cutoff, true when the requested walk finishes, even when its page is
|
|
298
|
+
truncated or empty. It is independent of `limit`/`offset` and is present for
|
|
299
|
+
unpaged answers too. Partial text answers say `Showing N of at least M (total
|
|
300
|
+
unknown: node_budget)` (or `edge_budget`), including when no pagination is needed.
|
|
301
|
+
`M` includes the root and counts the admitted prefix; it is a lower bound for the
|
|
302
|
+
requested root, depth, type/relationship filters and published generation, not a
|
|
303
|
+
count of all application callers. Exactness describes that same recorded-graph
|
|
304
|
+
scope and never implies exhaustive source coverage. Paging beyond that prefix
|
|
305
|
+
stays partial. To explore more, narrow
|
|
255
306
|
`depth`/`types`/`via`, choose another root, or increase the traversal budget within
|
|
256
307
|
its maximum. Keep the root, filters, budgets, and published generation unchanged
|
|
257
308
|
for stable pages. No wall-clock deadline is used, so cutoffs are deterministic.
|
|
@@ -289,6 +340,10 @@ A target name shared by several types has `type: null`,
|
|
|
289
340
|
record target types, so the response cannot choose among candidates. A witness
|
|
290
341
|
through an ambiguous or unresolved identity sets `typed_path_complete: false`;
|
|
291
342
|
it describes identifier-level reachability, never a uniquely typed path.
|
|
343
|
+
A true value means only that identities along this witness have unambiguous
|
|
344
|
+
types. It does not establish source-reference coverage or observed execution.
|
|
345
|
+
Text labels this `witness types unambiguous=yes/no`; the JSON key and its meaning
|
|
346
|
+
remain unchanged. The text label change is unreleased after `2.0.0.beta3`.
|
|
292
347
|
`types` filters retain the compact traversal's identifier-level semantics: any
|
|
293
348
|
registered type can qualify a name, while edge evidence keeps its actual source
|
|
294
349
|
owner. Multiple relationship kinds between the same endpoints remain separate.
|
data/docs/MCP_TOOL_COOKBOOK.md
CHANGED
|
@@ -64,7 +64,7 @@ The Index Server defines **29 schemas**: the packaged executable registers **14*
|
|
|
64
64
|
| Snapshot (4) | 4 | Extraction with `enable_snapshots = true` normally creates `woods.sqlite3`, which packaged servers discover. If extraction used the JSON fallback, set `WOODS_SNAPSHOTS=true` on the standalone server. Custom embedded servers pass `snapshot_store:`. Internal SQLite migrations are automatic. Tools: `list_snapshots`, `snapshot_diff`, `unit_history`, `snapshot_detail` |
|
|
65
65
|
| `notion_sync` | 1 | `notion_api_token` + `notion_database_ids` both set |
|
|
66
66
|
|
|
67
|
-
`codebase_retrieve` is always registered (no `retrieve` alias exists)
|
|
67
|
+
`codebase_retrieve` is always registered (no `retrieve` alias exists). Default semantic mode requires an embedding provider and a completed `woods:embed` run. Explicit `WOODS_RETRIEVAL_MODE=lexical` ranks published extraction units without a provider or embeddings; set it in the MCP process environment and restart the server. See [embedding-free lexical retrieval](RETRIEVAL_GUIDE.md#embedding-free-lexical-retrieval).
|
|
68
68
|
|
|
69
69
|
If an agent reports a missing tool, compare its request with the connected server's registered list and [MCP server boundaries](MCP_SERVERS.md#conditional-index-capabilities). The normal packaged executable does not wire operator or feedback collaborators. **Console Server tools are not all unconditionally registered**: 31 tool schemas exist as an inventory, but only the 9 Tier 1 tools are executable by default, or 11 with `console_embedded_read_tools: true` (adds `console_sql`/`console_query`). Tier 2, Tier 3, and `console_eval` are schema-only in every supported mode; there is no bridge or confirmation flow that unlocks them. See [MCP servers](MCP_SERVERS.md#console-server) for the supported inventory.
|
|
70
70
|
|
|
@@ -551,7 +551,7 @@ Static tools miss all of these because they only exist after Rails processes the
|
|
|
551
551
|
}
|
|
552
552
|
```
|
|
553
553
|
|
|
554
|
-
**What you'll get:** Units with no dependents
|
|
554
|
+
**What you'll get:** Units with no recorded dependents in the published graph, excluding types treated as natural entry points. These are candidates for investigation, not proof of dead code. Method-body references and dynamic callers may be missing; verify source references, framework entry points, and runtime usage before removing anything. See [dependency graph coverage](MCP_SERVERS.md#dependency-graph-coverage).
|
|
555
555
|
|
|
556
556
|
---
|
|
557
557
|
|
|
@@ -811,7 +811,7 @@ Keys without a recognised suffix fall through to ActiveRecord `where(hash)` equa
|
|
|
811
811
|
|
|
812
812
|
### "Find code related to subscription billing"
|
|
813
813
|
|
|
814
|
-
**Tool:** `codebase_retrieve` (Index Server,
|
|
814
|
+
**Tool:** `codebase_retrieve` (Index Server, semantic or explicit lexical mode)
|
|
815
815
|
|
|
816
816
|
```json
|
|
817
817
|
{
|
|
@@ -820,7 +820,7 @@ Keys without a recognised suffix fall through to ActiveRecord `where(hash)` equa
|
|
|
820
820
|
}
|
|
821
821
|
```
|
|
822
822
|
|
|
823
|
-
**What you'll get:**
|
|
823
|
+
**What you'll get:** Ranked context within an estimated text-token budget. Default semantic mode uses configured embeddings and hybrid ranking; explicit lexical mode uses field-aware BM25 over published units, without a provider or `woods:embed`. The same query works in either configured mode, though rankings differ. Confirm the active mode with `woods_status.retriever.mode`; see the [retrieval guide](RETRIEVAL_GUIDE.md#embedding-free-lexical-retrieval) for setup and budget limits.
|
|
824
824
|
|
|
825
825
|
---
|
|
826
826
|
|