woods 2.0.0.beta2 → 2.0.0.beta3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +262 -1
- data/CONTRIBUTING.md +173 -9
- data/README.md +7 -3
- data/SECURITY.md +9 -6
- data/docs/AGENT_GUIDE.md +83 -4
- data/docs/AGENT_SETUP.md +82 -1
- data/docs/BACKEND_MATRIX.md +20 -0
- data/docs/CLIENT_HOOKS.md +111 -0
- data/docs/CONFIGURATION_REFERENCE.md +199 -14
- data/docs/CONSOLE_MCP_SETUP.md +35 -5
- data/docs/DOCKER_SETUP.md +21 -2
- data/docs/EVALUATION.md +464 -1
- data/docs/EXTRACTOR_REFERENCE.md +36 -5
- data/docs/FAQ.md +11 -12
- data/docs/GETTING_STARTED.md +17 -5
- data/docs/INCREMENTAL_EXTRACTION.md +117 -1
- data/docs/INDEX_LAYOUT.md +382 -0
- data/docs/INTERNALS.md +7 -2
- data/docs/MCP_SERVERS.md +221 -5
- data/docs/MCP_TOOL_COOKBOOK.md +33 -18
- data/docs/NOTION_INTEGRATION.md +13 -0
- data/docs/OBSIDIAN_INTEGRATION.md +57 -9
- data/docs/PUBLISHED_INDEX.md +55 -0
- data/docs/README.md +7 -0
- data/docs/RETRIEVAL_GUIDE.md +253 -11
- data/docs/RUNTIME_TRACING.md +71 -0
- data/docs/SOURCE_FRESHNESS.md +143 -0
- data/docs/TROUBLESHOOTING.md +117 -5
- data/docs/UNBLOCKED_INTEGRATION.md +25 -0
- data/docs/UPGRADING_TO_2.md +44 -22
- data/docs/WATCH_DAEMON.md +259 -59
- data/exe/woods-agent-config +6 -0
- data/exe/woods-extract +5 -0
- data/exe/woods-hook-context +6 -0
- data/lib/generators/woods/templates/woods.rb.tt +1 -3
- data/lib/tasks/woods.rake +47 -397
- data/lib/woods/agent_configuration/applier.rb +133 -0
- data/lib/woods/agent_configuration/cli.rb +101 -0
- data/lib/woods/agent_configuration/cli_options.rb +29 -0
- data/lib/woods/agent_configuration/document.rb +105 -0
- data/lib/woods/agent_configuration/error.rb +7 -0
- data/lib/woods/agent_configuration/launcher.rb +75 -0
- data/lib/woods/agent_configuration/layout.rb +59 -0
- data/lib/woods/agent_configuration/managed_section.rb +62 -0
- data/lib/woods/agent_configuration/plan.rb +98 -0
- data/lib/woods/agent_configuration/plan_diff.rb +38 -0
- data/lib/woods/agent_configuration/planned_files.rb +61 -0
- data/lib/woods/agent_configuration/planner.rb +63 -0
- data/lib/woods/agent_configuration/planner_validation.rb +77 -0
- data/lib/woods/agent_configuration/preflight.rb +100 -0
- data/lib/woods/agent_configuration/recovery.rb +49 -0
- data/lib/woods/ast/node.rb +2 -0
- data/lib/woods/ast/parser.rb +38 -5
- data/lib/woods/builder.rb +21 -5
- data/lib/woods/cache/cache_middleware.rb +28 -7
- data/lib/woods/cache/cache_store.rb +4 -5
- data/lib/woods/change_set.rb +5 -4
- data/lib/woods/console/credential_index.rb +20 -2
- data/lib/woods/console/credential_scanner.rb +14 -14
- data/lib/woods/console/credential_scanner_registry.rb +36 -0
- data/lib/woods/console/embedded_executor.rb +1 -1
- data/lib/woods/console/encrypted_credential_snapshot.rb +16 -0
- data/lib/woods/console/rack_middleware.rb +22 -13
- data/lib/woods/console/server.rb +18 -16
- data/lib/woods/dependency_graph.rb +65 -13
- data/lib/woods/embedding/corpus.rb +94 -0
- data/lib/woods/embedding/indexer.rb +90 -46
- data/lib/woods/embedding/openai.rb +17 -6
- data/lib/woods/evaluation/ablation_executor.rb +6 -1
- data/lib/woods/evaluation/ablation_timed_executor.rb +22 -4
- data/lib/woods/export/typed_reader.rb +56 -0
- data/lib/woods/extractor.rb +232 -137
- data/lib/woods/extractors/action_cable_extractor.rb +3 -1
- data/lib/woods/extractors/behavioral_profile.rb +9 -7
- data/lib/woods/extractors/caching_extractor.rb +3 -1
- data/lib/woods/extractors/concern_extractor.rb +64 -6
- data/lib/woods/extractors/configuration_extractor.rb +7 -3
- data/lib/woods/extractors/controller_extractor.rb +13 -4
- data/lib/woods/extractors/database_view_extractor.rb +3 -1
- data/lib/woods/extractors/decorator_extractor.rb +3 -1
- data/lib/woods/extractors/engine_extractor.rb +3 -1
- data/lib/woods/extractors/event_extractor.rb +4 -2
- data/lib/woods/extractors/factory_extractor.rb +3 -1
- data/lib/woods/extractors/graphql_extractor.rb +8 -2
- data/lib/woods/extractors/i18n_extractor.rb +3 -1
- data/lib/woods/extractors/job_extractor.rb +6 -19
- data/lib/woods/extractors/lib_extractor.rb +3 -1
- data/lib/woods/extractors/mailer_extractor.rb +20 -5
- data/lib/woods/extractors/manager_extractor.rb +3 -1
- data/lib/woods/extractors/method_parameters.rb +53 -0
- data/lib/woods/extractors/middleware_argument.rb +65 -0
- data/lib/woods/extractors/middleware_extractor.rb +9 -3
- data/lib/woods/extractors/migration_extractor.rb +3 -1
- data/lib/woods/extractors/model_extractor.rb +39 -33
- data/lib/woods/extractors/package_extractor.rb +24 -4
- data/lib/woods/extractors/phlex_extractor.rb +3 -1
- data/lib/woods/extractors/policy_extractor.rb +3 -1
- data/lib/woods/extractors/poro_extractor.rb +3 -1
- data/lib/woods/extractors/pundit_extractor.rb +3 -1
- data/lib/woods/extractors/rails_source_extractor.rb +4 -2
- data/lib/woods/extractors/rake_task_extractor.rb +4 -2
- data/lib/woods/extractors/route_extractor.rb +3 -1
- data/lib/woods/extractors/route_helper_resolver.rb +10 -33
- data/lib/woods/extractors/scheduled_job_extractor.rb +41 -15
- data/lib/woods/extractors/serializer_extractor.rb +4 -2
- data/lib/woods/extractors/service_extractor.rb +3 -1
- data/lib/woods/extractors/shared_dependency_scanner.rb +2 -2
- data/lib/woods/extractors/shared_utility_methods.rb +27 -15
- data/lib/woods/extractors/source_nesting.rb +1 -1
- data/lib/woods/extractors/state_machine_extractor.rb +3 -1
- data/lib/woods/extractors/test_mapping_extractor.rb +3 -1
- data/lib/woods/extractors/validator_extractor.rb +3 -1
- data/lib/woods/extractors/view_component_extractor.rb +3 -1
- data/lib/woods/extractors/view_template_extractor.rb +3 -1
- data/lib/woods/gem_mapper.rb +2 -0
- data/lib/woods/git_history.rb +116 -0
- data/lib/woods/graph_analyzer.rb +35 -6
- data/lib/woods/hooks/context_cli.rb +54 -0
- data/lib/woods/hooks/context_event.rb +88 -0
- data/lib/woods/hooks/context_hint.rb +73 -0
- data/lib/woods/hooks/context_impact.rb +77 -0
- data/lib/woods/hooks/context_output.rb +47 -0
- data/lib/woods/hooks/context_state.rb +102 -0
- data/lib/woods/hooks/refresh.rb +79 -0
- data/lib/woods/hooks/rule_projection.rb +78 -0
- data/lib/woods/input_rules.rb +19 -0
- data/lib/woods/mcp/bearer_auth.rb +20 -12
- data/lib/woods/mcp/bootstrapper.rb +62 -0
- data/lib/woods/mcp/index_reader.rb +323 -160
- data/lib/woods/mcp/initialization_guidance.rb +27 -0
- data/lib/woods/mcp/origin_guard.rb +17 -9
- data/lib/woods/mcp/published_lexical_retriever.rb +115 -0
- data/lib/woods/mcp/renderers/markdown_renderer.rb +8 -1
- data/lib/woods/mcp/renderers/plain_renderer.rb +7 -1
- data/lib/woods/mcp/search_results.rb +74 -0
- data/lib/woods/mcp/server.rb +158 -37
- data/lib/woods/mcp/tool_contract.rb +2 -0
- data/lib/woods/mcp/tool_response_renderer.rb +25 -0
- data/lib/woods/mcp/traversal_evidence.rb +113 -0
- data/lib/woods/mcp/traversal_evidence_index.rb +100 -0
- data/lib/woods/mcp/traversal_evidence_page.rb +41 -0
- data/lib/woods/mcp/traversal_evidence_text.rb +52 -0
- data/lib/woods/notion/exporter.rb +56 -17
- data/lib/woods/obsidian/destination_plan.rb +98 -0
- data/lib/woods/obsidian/name_mapper.rb +19 -3
- data/lib/woods/obsidian/note_builder.rb +19 -10
- data/lib/woods/obsidian/vault_exporter.rb +88 -32
- data/lib/woods/operator/pipeline_guard.rb +18 -13
- data/lib/woods/path_dispatcher.rb +7 -1
- data/lib/woods/payload_store.rb +27 -26
- data/lib/woods/railtie.rb +3 -3
- data/lib/woods/railtie_support.rb +12 -12
- data/lib/woods/rake_helpers.rb +392 -0
- data/lib/woods/resilience/graph_invariant_validator/membership_checks.rb +71 -0
- data/lib/woods/resilience/graph_invariant_validator/node_checks.rb +61 -0
- data/lib/woods/resilience/graph_invariant_validator/reverse_relationship_checks.rb +46 -0
- data/lib/woods/resilience/graph_invariant_validator.rb +119 -0
- data/lib/woods/resilience/index_validator/graph_checks.rb +80 -0
- data/lib/woods/resilience/index_validator.rb +112 -23
- data/lib/woods/retrieval/context_assembler.rb +50 -15
- data/lib/woods/retrieval/lexical_assembler.rb +73 -0
- data/lib/woods/retrieval/lexical_index.rb +119 -0
- data/lib/woods/retrieval/ranker.rb +4 -2
- data/lib/woods/retrieval/scope.rb +108 -0
- data/lib/woods/retrieval/scoped_graph_store.rb +32 -0
- data/lib/woods/retrieval/scoped_vector_store.rb +55 -0
- data/lib/woods/retrieval/search_executor.rb +86 -27
- data/lib/woods/retrieval/source_evidence.rb +200 -0
- data/lib/woods/retriever.rb +98 -22
- data/lib/woods/ruby_analyzer/trace_enricher.rb +77 -38
- data/lib/woods/session_tracer/middleware.rb +10 -12
- data/lib/woods/session_tracer/redis_store.rb +22 -6
- data/lib/woods/session_tracer/session_flow_assembler.rb +23 -17
- data/lib/woods/session_tracer/solid_cache_coordination.rb +6 -4
- data/lib/woods/session_tracer/unit_resolver.rb +63 -0
- data/lib/woods/source_inputs/consumer_errors.rb +27 -0
- data/lib/woods/source_inputs/handoff.rb +102 -0
- data/lib/woods/source_inputs/launcher.rb +157 -0
- data/lib/woods/source_inputs/manifest.rb +124 -0
- data/lib/woods/source_inputs/private_key.rb +55 -0
- data/lib/woods/source_inputs/scanner.rb +171 -0
- data/lib/woods/source_inputs/scopes.rb +71 -0
- data/lib/woods/source_inputs/session.rb +214 -0
- data/lib/woods/source_inputs/status.rb +84 -0
- data/lib/woods/source_inputs/verifier.rb +107 -0
- data/lib/woods/storage/metadata_store.rb +25 -25
- data/lib/woods/storage/pgvector.rb +29 -8
- data/lib/woods/storage/qdrant.rb +17 -7
- data/lib/woods/storage/vector_store.rb +18 -6
- data/lib/woods/tasks.rb +3 -2
- data/lib/woods/temporal/json_snapshot_store.rb +29 -8
- data/lib/woods/unblocked/exporter.rb +59 -70
- data/lib/woods/version.rb +1 -1
- data/lib/woods/watch/boot_snapshot.rb +52 -0
- data/lib/woods/watch/daemon.rb +136 -28
- data/lib/woods/watch/listen_watcher.rb +4 -0
- data/lib/woods/watch/polling_watcher.rb +5 -1
- data/lib/woods/watch/status.rb +20 -15
- data/lib/woods/watch/tree_scan.rb +21 -13
- data/lib/woods/watch/watcher.rb +4 -1
- data/lib/woods.rb +50 -11
- data/plugin/.claude-plugin/plugin.json +1 -1
- data/plugin/hooks/adapters/normalize.jq +15 -0
- data/plugin/hooks/adapters/normalize.rb +63 -0
- data/plugin/hooks/hooks.json +20 -0
- data/plugin/hooks/woods-context.sh +50 -0
- data/plugin/hooks/woods-input-rules.sh +159 -0
- data/plugin/hooks/woods-opencode.mjs +65 -0
- data/plugin/hooks/woods-post-edit.sh +2 -225
- data/plugin/hooks/woods-refresh.sh +260 -0
- data/plugin/hooks/woods-session-start.sh +47 -55
- data/plugin/skills/woods-agent-enable/SKILL.md +13 -0
- data/plugin/skills/woods-diagnose/SKILL.md +288 -1
- data/plugin/skills/woods-investigate/SKILL.md +106 -0
- data/plugin/skills/woods-mcp-config/SKILL.md +89 -1
- data/plugin/skills/woods-setup/SKILL.md +107 -6
- metadata +84 -5
data/docs/MCP_SERVERS.md
CHANGED
|
@@ -27,6 +27,15 @@ bin/rails woods:validate
|
|
|
27
27
|
bin/rails woods:stats
|
|
28
28
|
```
|
|
29
29
|
|
|
30
|
+
For embedding-free ranked retrieval, start Index MCP with
|
|
31
|
+
`WOODS_RETRIEVAL_MODE=lexical`. This opt-in reads the published extraction units;
|
|
32
|
+
it does not probe providers or load vectors. `woods_status.retriever.mode` reports
|
|
33
|
+
`lexical`, and inactive embedding fields are `null`. The default semantic mode
|
|
34
|
+
keeps its existing embedding setup and failure behavior. See
|
|
35
|
+
[retrieval modes](RETRIEVAL_GUIDE.md#embedding-free-lexical-retrieval) for scoring,
|
|
36
|
+
query limits and measured tradeoffs. Both packaged stdio and HTTP launches honor
|
|
37
|
+
the setting; put it in the MCP process's environment, not just a Rails initializer.
|
|
38
|
+
|
|
30
39
|
The stdio server can then run outside Rails. Point it at the index root (`tmp/woods/` by default), not at an internal generation or payload directory.
|
|
31
40
|
|
|
32
41
|
### Configure a stdio client
|
|
@@ -47,6 +56,10 @@ Prefer the application's bundle and a project-scoped configuration:
|
|
|
47
56
|
|
|
48
57
|
`woods-mcp-start` checks that the directory and published manifest exist, then replaces itself with `woods-mcp`. It does not install dependencies or restart a crashed process.
|
|
49
58
|
|
|
59
|
+
`woods_status.index.woods_version` identifies the last publisher of the served
|
|
60
|
+
manifest; `server.version` identifies the running MCP reader. Missing writer
|
|
61
|
+
provenance is `null`. See [manifest writer provenance](PUBLISHED_INDEX.md#manifest-writer-provenance).
|
|
62
|
+
|
|
50
63
|
You can launch the server directly when the client already handles preflight:
|
|
51
64
|
|
|
52
65
|
```bash
|
|
@@ -57,6 +70,12 @@ Keep stdout reserved for MCP protocol messages. Diagnose startup failures from s
|
|
|
57
70
|
|
|
58
71
|
### Client configuration locations
|
|
59
72
|
|
|
73
|
+
Supporting development builds offer preview/apply/update/remove ownership for
|
|
74
|
+
Claude Code project or explicit user configuration. See
|
|
75
|
+
[managed configuration](AGENT_SETUP.md#managed-claude-code-configuration) for
|
|
76
|
+
`woods-agent-config`, host/Compose preflight, conflict handling, and recovery.
|
|
77
|
+
Manual configuration remains available for older gems and other clients.
|
|
78
|
+
|
|
60
79
|
MCP clients expose project or user-level server settings in different locations. Use project scope when available, preserve the `command`, `args`, and absolute `cwd` semantics above, and translate only the surrounding client-specific format. Woods is model-independent: compatibility depends on the client supporting MCP stdio or Streamable HTTP, not on whether the connected model is from OpenAI, Anthropic, Google, xAI, or another provider.
|
|
61
80
|
|
|
62
81
|
Client configuration formats can change independently of Woods. If a client rejects otherwise valid JSON, check that client's current MCP documentation.
|
|
@@ -99,6 +118,27 @@ Reconnect the client, then call:
|
|
|
99
118
|
|
|
100
119
|
Prefer a real MCP client's connection flow over a hand-written JSON-RPC pipe. Modern MCP 2026-07-28 requests carry per-request protocol metadata and can use `server/discover` without an initialization handshake; older clients still use `initialize`. A valid raw smoke test must implement one complete flow rather than sending an isolated `tools/list` or `tools/call` request.
|
|
101
120
|
|
|
121
|
+
### Initialization guidance
|
|
122
|
+
|
|
123
|
+
The Index Server supplies a short, client-neutral `instructions` field through
|
|
124
|
+
the SDK's `initialize` response and modern `server/discover`. It describes the
|
|
125
|
+
status → discovery → inspection → bounded traversal → source-verification
|
|
126
|
+
workflow, and lists only the tools actually registered by this server.
|
|
127
|
+
Instructions are stable for unchanged tool registration and bounded to 2,048
|
|
128
|
+
UTF-8 bytes across supported configurations. Building the text does not probe
|
|
129
|
+
providers, extract code, or write configuration.
|
|
130
|
+
|
|
131
|
+
Registration alone does not establish retrieval readiness: check `woods_status`
|
|
132
|
+
before using `codebase_retrieve`, including after reload. The guidance grants
|
|
133
|
+
no extraction, configuration-change, or Console authorization. Detailed usage
|
|
134
|
+
belongs in the [agent guide](AGENT_GUIDE.md).
|
|
135
|
+
|
|
136
|
+
This addition is unreleased after `2.0.0.beta2`. The SDK omits `instructions`
|
|
137
|
+
when negotiating protocol `2024-11-05`; that behavior is preserved. Older gems,
|
|
138
|
+
legacy clients, and clients that do not show server instructions can use the
|
|
139
|
+
agent guide or investigation skill. Leave protocol negotiation enabled rather
|
|
140
|
+
than pinning a newer version solely to obtain guidance.
|
|
141
|
+
|
|
102
142
|
### Tools (29 — 14 registered in the packaged default)
|
|
103
143
|
|
|
104
144
|
The Index Server defines 29 schemas across core and conditional capabilities. The normal packaged executable registers the 14 tools below; the remaining schemas require the specialized wiring described afterward.
|
|
@@ -108,8 +148,8 @@ The Index Server defines 29 schemas across core and conditional capabilities. Th
|
|
|
108
148
|
| `woods_status` | Index health, generation, counts, and retrieval readiness |
|
|
109
149
|
| `search` | Discover identifiers by regex, prefix, suffix, source, or metadata |
|
|
110
150
|
| `lookup` | Fetch one exact unit with source, metadata, and relationships |
|
|
111
|
-
| `dependencies` | Traverse what a unit depends on (`depth`, `types`, `via` narrow; `limit`, `offset` page) |
|
|
112
|
-
| `dependents` | Traverse what depends on a unit (`depth`, `types`, `via` narrow; `limit`, `offset` page) |
|
|
151
|
+
| `dependencies` | Traverse what a unit depends on (`depth`, `types`, `via` narrow; `max_nodes`, `max_edges` budget work; `limit`, `offset` page) |
|
|
152
|
+
| `dependents` | Traverse what depends on a unit (`depth`, `types`, `via` narrow; `max_nodes`, `max_edges` budget work; `limit`, `offset` page) |
|
|
113
153
|
| `structure` | Summarize structural relationships around a unit |
|
|
114
154
|
| `trace_flow` | Follow a request, job, mail, or other execution flow |
|
|
115
155
|
| `framework` | Inspect relevant Rails or installed gem source |
|
|
@@ -118,7 +158,29 @@ The Index Server defines 29 schemas across core and conditional capabilities. Th
|
|
|
118
158
|
| `domain_clusters` | Discover connected domains in the graph |
|
|
119
159
|
| `pagerank` | Find structurally central units |
|
|
120
160
|
| `reload` | Reload a newly published generation without restarting the client |
|
|
121
|
-
| `codebase_retrieve` | Natural-language retrieval
|
|
161
|
+
| `codebase_retrieve` | Natural-language retrieval with embeddings or explicit lexical mode over extraction output |
|
|
162
|
+
|
|
163
|
+
When an identifier appears in multiple extraction types, `framework` reads its
|
|
164
|
+
framework-source bucket and `recent_changes` reads each selected type bucket.
|
|
165
|
+
Their paths and metadata belong to that selected bucket. When session tracing
|
|
166
|
+
is configured, newly recorded requests use the dispatched controller's runtime
|
|
167
|
+
class name; the fallback for requests without an instance respects Rails acronym
|
|
168
|
+
inflections. Existing trace records are unchanged. Controller lookup and root
|
|
169
|
+
outgoing-edge selection preserve the controller type. Downstream references and the shared context pool still use
|
|
170
|
+
bare identifiers. If a dependency encountered within the requested depth has
|
|
171
|
+
multiple published types, `session_trace` returns an `ambiguous_identity` tool
|
|
172
|
+
error naming the identifier and candidate types, with no partial context. This
|
|
173
|
+
also prevents an earlier dependency from occupying a later controller’s context
|
|
174
|
+
key. Unrelated collisions do not block a trace, and a known controller root keeps
|
|
175
|
+
its controller identity. A controller absent from the index remains in the
|
|
176
|
+
timeline without a source reference, so another type cannot fill that reference.
|
|
177
|
+
Candidate discovery and source reads use one pinned
|
|
178
|
+
generation. Corrupt or missing listed artifacts retain the `internal_error`
|
|
179
|
+
failure boundary; they do not prove uniqueness or become `ambiguous_identity` errors.
|
|
180
|
+
Use `depth: 0` for the request timeline, or inspect candidates with typed `lookup`
|
|
181
|
+
calls. Re-extraction does not remove a legitimate cross-type collision. Successful
|
|
182
|
+
traces retain their existing identifiers and response shape; target identity has
|
|
183
|
+
not been migrated globally. These corrections are unreleased after `2.0.0.beta2`.
|
|
122
184
|
|
|
123
185
|
The server also exposes MCP resources and resource templates for indexed units. Tool descriptions returned by MCP are the parameter-level source of truth; [Agent guide](AGENT_GUIDE.md) explains selection strategy.
|
|
124
186
|
|
|
@@ -130,6 +192,126 @@ error and continues serving the previous aligned generation; it never swaps in a
|
|
|
130
192
|
partial or empty replacement. Grant write access for live reloads, or restart the MCP
|
|
131
193
|
process after publishing a new embedded index.
|
|
132
194
|
|
|
195
|
+
### Search completeness
|
|
196
|
+
|
|
197
|
+
Search responses retain `query`, `result_count`, and `results`; `result_count`
|
|
198
|
+
is the number returned, not an estimated total. The additive `completeness`
|
|
199
|
+
object describes the requested types, literal filters, and fields in the pinned
|
|
200
|
+
generation. This contract is unreleased after `2.0.0.beta2`.
|
|
201
|
+
|
|
202
|
+
| `reason` | `status` | `has_more` | `total_matches` |
|
|
203
|
+
|---|---|---|---|
|
|
204
|
+
| `exhausted` | `complete` | `false` | Exact count |
|
|
205
|
+
| `result_limit` | `partial` | `true` | `null` (unknown) |
|
|
206
|
+
| `scan_budget` or `regex_timeout` | `partial` | `null` (unknown) | `null` (unknown) |
|
|
207
|
+
|
|
208
|
+
`matched_lower_bound` counts distinct observed `(type, identifier)` matches,
|
|
209
|
+
including at most one lookahead match beyond `limit`. A result-limit response
|
|
210
|
+
therefore establishes another match; an exactly full page can instead be
|
|
211
|
+
complete if the requested domain is exhausted. Deep lookahead shares
|
|
212
|
+
`WOODS_SEARCH_MAX_SCAN` with the initial scan and retains round-robin scanning
|
|
213
|
+
across types. Search does not count the entire omitted tail or offer pagination.
|
|
214
|
+
The existing `types` filter and result labels name directory families:
|
|
215
|
+
`rails_source` includes both Rails and gem source units. Deep reads accept those
|
|
216
|
+
two stored types only in that shared directory; `lookup` and lexical retrieval
|
|
217
|
+
retain the unit's actual `rails_source` or `gem_source` type.
|
|
218
|
+
|
|
219
|
+
All partial responses retain `partial: true` and include a narrowing `hint`.
|
|
220
|
+
JSON exposes these fields; Markdown, plain text, and Claude formats label the
|
|
221
|
+
returned count, stopping reason, known/unknown remainder, and total explicitly.
|
|
222
|
+
Narrow `types`, literal `exact_prefix`/`exact_suffix`, or deep `fields` before
|
|
223
|
+
using discovery as exhaustive evidence. Completeness applies to this index and
|
|
224
|
+
query domain, not to unindexed application code.
|
|
225
|
+
|
|
226
|
+
Detected missing, unreadable, or corrupt artifacts remain `isError: true` with
|
|
227
|
+
`_meta.error_code: "corrupt_artifact"`. Their `_meta.completeness` has
|
|
228
|
+
`status: "unknown"`, `reason: "unreadable_or_corrupt_source"`, and `null` for
|
|
229
|
+
`has_more`, `total_matches`, and `matched_lower_bound`; no successful empty
|
|
230
|
+
result is substituted. Inspect `woods_status` and run `woods:validate`.
|
|
231
|
+
|
|
232
|
+
### Dependency traversal budgets
|
|
233
|
+
|
|
234
|
+
`dependencies` and `dependents` walk breadth-first in stored graph order. The
|
|
235
|
+
walk defaults to `max_nodes: 1000` (including the root) and `max_edges: 10000`;
|
|
236
|
+
callers can select 1–10,000 nodes and 1–100,000 edge checks. The node budget
|
|
237
|
+
counts distinct nodes admitted after filters. Every candidate edge is charged
|
|
238
|
+
before filtering, including duplicates, cycles, and the forward-edge checks
|
|
239
|
+
needed to match a reverse `via` filter. Thus restrictive filters cannot bypass
|
|
240
|
+
the edge budget. Nodes at the requested `depth` are recorded without reading
|
|
241
|
+
their adjacency lists.
|
|
242
|
+
|
|
243
|
+
When further expansion would exceed a budget, JSON reports `partial: true`,
|
|
244
|
+
`partial_reason: "node_budget"` or `"edge_budget"`, and `traversal_budget` with
|
|
245
|
+
`max_nodes`, `max_edges`, `visited_nodes`, and `visited_edges`. Text renderers
|
|
246
|
+
also identify the partial traversal. Already discovered nodes remain in the
|
|
247
|
+
answer, but an empty `deps` array in a partial answer does not prove a leaf.
|
|
248
|
+
Exact-budget walks that finish all requested work are complete and have no
|
|
249
|
+
`partial` marker.
|
|
250
|
+
|
|
251
|
+
`limit` (default 50) and `offset` only page that discovered result; they never
|
|
252
|
+
change the walk budget or depth. On a partial traversal, `nodes_total`, when
|
|
253
|
+
present for pagination, counts the discovered prefix, **not the full reachable
|
|
254
|
+
graph**. Paging beyond that prefix stays partial. To explore more, narrow
|
|
255
|
+
`depth`/`types`/`via`, choose another root, or increase the traversal budget within
|
|
256
|
+
its maximum. Keep the root, filters, budgets, and published generation unchanged
|
|
257
|
+
for stable pages. No wall-clock deadline is used, so cutoffs are deterministic.
|
|
258
|
+
|
|
259
|
+
Budgets cover traversal work after per-generation graph loading and cache
|
|
260
|
+
preparation (JSON parsing, typed-edge normalization, node types and database
|
|
261
|
+
metadata). They do not cap that initial load, elapsed time, or total process
|
|
262
|
+
memory. These arguments are unreleased in Woods 2.0.0.beta2; check the connected
|
|
263
|
+
server's tool schema before sending them to an older installation.
|
|
264
|
+
|
|
265
|
+
### Traversal explanations
|
|
266
|
+
|
|
267
|
+
Supporting development versions accept `explain: true` on `dependencies` and
|
|
268
|
+
`dependents`. Check the connected schema first; this option is unreleased after
|
|
269
|
+
2.0.0.beta2. Omitted or false keeps the existing compact response.
|
|
270
|
+
|
|
271
|
+
The additive `explanation` object contains:
|
|
272
|
+
|
|
273
|
+
- `direction`: `forward` or `reverse`, plus the requested `root` identity.
|
|
274
|
+
- `edges`: records keyed by response-local IDs such as `e0`. Every record keeps
|
|
275
|
+
the original **source → target** direction, even during reverse traversal.
|
|
276
|
+
`source` contains its recorded `identifier` and `type`; `target` contains its
|
|
277
|
+
identifier and the unique type when the published graph establishes one.
|
|
278
|
+
`via`, `through`, `through_db`, and `disable_joins` preserve recorded values;
|
|
279
|
+
absent legacy attributes are null (shown as unknown in text), including an
|
|
280
|
+
unrecorded `disable_joins` rather than an invented false value.
|
|
281
|
+
- `witnesses`: one shortest breadth-first predecessor per admitted identifier,
|
|
282
|
+
keyed by identifier. Each has `parent`, `edge_id`, `impact` (`root`, `direct`,
|
|
283
|
+
or `transitive`), and `typed_path_complete`. Follow parent references to the
|
|
284
|
+
root to reconstruct one witness; alternative paths are not enumerated.
|
|
285
|
+
|
|
286
|
+
A target name shared by several types has `type: null`,
|
|
287
|
+
`resolution: "ambiguous"`, and sorted `candidate_types`. An unresolved target has
|
|
288
|
+
`resolution: "unresolved"` and an empty candidate list. Forward artifacts do not
|
|
289
|
+
record target types, so the response cannot choose among candidates. A witness
|
|
290
|
+
through an ambiguous or unresolved identity sets `typed_path_complete: false`;
|
|
291
|
+
it describes identifier-level reachability, never a uniquely typed path.
|
|
292
|
+
`types` filters retain the compact traversal's identifier-level semantics: any
|
|
293
|
+
registered type can qualify a name, while edge evidence keeps its actual source
|
|
294
|
+
owner. Multiple relationship kinds between the same endpoints remain separate.
|
|
295
|
+
|
|
296
|
+
Direct witnesses establish a recorded root relationship; transitive witnesses
|
|
297
|
+
represent inferred downstream reachability through recorded relationships.
|
|
298
|
+
Neither establishes observed execution, confidence, call order, or test coverage.
|
|
299
|
+
|
|
300
|
+
Node pagination retains required ancestor witnesses once, marked `context: true`
|
|
301
|
+
when outside the page; returned rows have `context: false`. Context records do
|
|
302
|
+
not increase the result-row count. Page evidence retains the witness edges and
|
|
303
|
+
other observed relationships among its visible/context endpoints; an empty page
|
|
304
|
+
has empty edge/witness maps. Edge IDs are local to this traversal response.
|
|
305
|
+
|
|
306
|
+
All examined evidence shares the existing edge budget, before `via`/`types`
|
|
307
|
+
filtering. Current `reverse_via` buckets allow direct reverse evidence lookup;
|
|
308
|
+
legacy recovery charges each reverse candidate and every inspected forward edge.
|
|
309
|
+
The shared predecessor forest and emitted records remain bounded by admitted
|
|
310
|
+
nodes and inspected edges. Per-generation JSON loading and the cached
|
|
311
|
+
O(nodes + variants) ownership/type preparation are outside the walk budget;
|
|
312
|
+
explanation mode never flattens all forward edges as per-request preparation.
|
|
313
|
+
Partial traversal and pagination metadata retain the budget contract above.
|
|
314
|
+
|
|
133
315
|
### Conditional Index capabilities
|
|
134
316
|
|
|
135
317
|
The Ruby server builder contains 15 additional schemas for sessions, pipeline operations, retrieval feedback, temporal snapshots, and Notion sync. They register only when their required collaborators or configuration are wired.
|
|
@@ -140,6 +322,7 @@ The normal packaged executable does not wire pipeline-operator or feedback-store
|
|
|
140
322
|
|
|
141
323
|
Use HTTP only for a deliberate shared or remote deployment. It expands the network boundary and requires authentication, origin restrictions, and TLS termination. Follow [MCP HTTP transport](MCP_HTTP_TRANSPORT.md); do not translate the stdio example into an unauthenticated public listener.
|
|
142
324
|
|
|
325
|
+
|
|
143
326
|
## Console Server
|
|
144
327
|
|
|
145
328
|
The Console Server launches a Rails process through direct, Docker, or SSH connection configuration. It reads live data and must be treated as a separate security decision.
|
|
@@ -151,11 +334,14 @@ Console MCP is disabled by default because it reads live application data. Enabl
|
|
|
151
334
|
```ruby
|
|
152
335
|
Woods.configure do |config|
|
|
153
336
|
config.console_mcp_enabled = true
|
|
154
|
-
config.
|
|
337
|
+
config.console_mcp_http_enabled = false # stdio-only
|
|
155
338
|
end
|
|
156
339
|
```
|
|
157
340
|
|
|
158
|
-
|
|
341
|
+
This explicitly disables HTTP Console while retaining stdio access; no HTTP
|
|
342
|
+
token is needed at boot. Existing configurations default to HTTP enabled.
|
|
343
|
+
For HTTP deployment, enable the HTTP flag and configure its token, origins
|
|
344
|
+
and TLS using the [Console setup guide](CONSOLE_MCP_SETUP.md#option-c-http-rack-middleware).
|
|
159
345
|
|
|
160
346
|
Without a console connection file, the executable then launches the Rails task directly from its `cwd`:
|
|
161
347
|
|
|
@@ -229,3 +415,33 @@ Report vulnerabilities privately through [SECURITY.md](../SECURITY.md).
|
|
|
229
415
|
5. Check [Troubleshooting](TROUBLESHOOTING.md) for the exact stderr message.
|
|
230
416
|
|
|
231
417
|
For agent query behavior after connection, continue to [Agent guide](AGENT_GUIDE.md).
|
|
418
|
+
|
|
419
|
+
### Explicit retrieval and discovery scope
|
|
420
|
+
|
|
421
|
+
On a server whose tool schema advertises them, `packages` and `source_paths` narrow
|
|
422
|
+
`search` and `codebase_retrieve` before candidate limits. These are per-call
|
|
423
|
+
arguments, not configuration settings. Inspect applied scope and completeness;
|
|
424
|
+
a narrow graph query can omit relevant cross-boundary dependencies. See the
|
|
425
|
+
[scope contract](RETRIEVAL_GUIDE.md#explicit-package-and-source-path-scopes) for
|
|
426
|
+
root/nested ownership, path normalization, errors, storage support, and cost.
|
|
427
|
+
|
|
428
|
+
### Source freshness in status
|
|
429
|
+
|
|
430
|
+
`woods_status` accepts optional `source_check: "quick"` (default, 250ms scan) or
|
|
431
|
+
`"deep"` (five seconds). `index.source_freshness` describes the served generation
|
|
432
|
+
as `current`, `drifted` or `unknown`; missing source/key and incomplete capture
|
|
433
|
+
never count as current. No Rails initialization or provider call is needed.
|
|
434
|
+
See [source freshness](SOURCE_FRESHNESS.md) for scope, private-key handling and
|
|
435
|
+
fresh-process extraction. Existing HEAD/dirty fields remain separate diagnostics.
|
|
436
|
+
|
|
437
|
+
### Explicit source evidence modes
|
|
438
|
+
|
|
439
|
+
When advertised by the installed schema, `lookup` and `codebase_retrieve` accept
|
|
440
|
+
`evidence: 'compact'` or `'outline'`; omitted/`'full'` preserves existing behavior.
|
|
441
|
+
Retrieval uses its original query. Compact lookup accepts optional `query` and an
|
|
442
|
+
estimated `budget` (default 2000); full lookup remains complete. `lookup` also
|
|
443
|
+
accepts an actual `type` and a `source_sha256` guard for typed, byte-verified
|
|
444
|
+
follow-up from an excerpt. Compact modes cannot be combined with metadata-only
|
|
445
|
+
lookup controls. Structured provenance stays within the existing closed output
|
|
446
|
+
schema's `data` field. Read the [evidence contract](RETRIEVAL_GUIDE.md#compact-published-evidence-and-api-outlines)
|
|
447
|
+
before interpreting published line ranges as physical source locations.
|
data/docs/MCP_TOOL_COOKBOOK.md
CHANGED
|
@@ -255,13 +255,24 @@ The `metadata.inlined_concerns` array lists which concerns were resolved:
|
|
|
255
255
|
}
|
|
256
256
|
```
|
|
257
257
|
|
|
258
|
-
**What you'll get:** A BFS
|
|
258
|
+
**What you'll get:** A BFS traversal of units that reference `User`, such as controllers, services, jobs, and mailers, up to 2 hops out. Set `depth: 1` for direct dependents only.
|
|
259
259
|
|
|
260
|
-
The answer is
|
|
260
|
+
The answer is paged to 50 nodes by default. When it is cut, the response ends with a
|
|
261
261
|
`Showing N of M (truncated)` line, the same one `graph_analysis` prints. Reach
|
|
262
262
|
for `depth`, `types` and `via` first: they make the answer smaller. `limit` and
|
|
263
263
|
`offset` only page what those leave, so a hub read one page at a time still
|
|
264
|
-
|
|
264
|
+
repeats the walk. A separate traversal budget can return `partial: true`;
|
|
265
|
+
that marker means the reachable graph is incomplete even after the final page.
|
|
266
|
+
See [traversal budgets](MCP_SERVERS.md#dependency-traversal-budgets) before
|
|
267
|
+
changing `max_nodes` or `max_edges`.
|
|
268
|
+
|
|
269
|
+
To explain why a row is affected, check the connected schema and add
|
|
270
|
+
`"explain": true`. The response preserves source-to-target labels even while
|
|
271
|
+
walking dependents. Its predecessor witnesses distinguish direct relationships
|
|
272
|
+
from transitive inferred reachability, retain ancestor context across pages, and
|
|
273
|
+
mark ambiguous types explicitly. See the
|
|
274
|
+
[explanation contract](MCP_SERVERS.md#traversal-explanations); these witnesses are
|
|
275
|
+
not proof of observed execution.
|
|
265
276
|
|
|
266
277
|
To find only which jobs depend on `User`:
|
|
267
278
|
|
|
@@ -323,26 +334,30 @@ column.
|
|
|
323
334
|
}
|
|
324
335
|
```
|
|
325
336
|
|
|
326
|
-
**Example response
|
|
337
|
+
**Example JSON data** (inside the MCP response envelope):
|
|
327
338
|
|
|
328
339
|
```json
|
|
329
|
-
|
|
330
|
-
|
|
331
|
-
|
|
332
|
-
|
|
333
|
-
"
|
|
334
|
-
|
|
335
|
-
|
|
336
|
-
|
|
337
|
-
|
|
338
|
-
|
|
339
|
-
|
|
340
|
-
|
|
340
|
+
{
|
|
341
|
+
"query": "payment",
|
|
342
|
+
"result_count": 1,
|
|
343
|
+
"results": [
|
|
344
|
+
{ "identifier": "PaymentsController", "type": "controller", "match_field": "identifier" }
|
|
345
|
+
],
|
|
346
|
+
"completeness": {
|
|
347
|
+
"status": "complete",
|
|
348
|
+
"reason": "exhausted",
|
|
349
|
+
"has_more": false,
|
|
350
|
+
"total_matches": 1,
|
|
351
|
+
"matched_lower_bound": 1
|
|
341
352
|
}
|
|
342
|
-
|
|
353
|
+
}
|
|
343
354
|
```
|
|
344
355
|
|
|
345
|
-
|
|
356
|
+
Use `lookup` on the returned identifier for source, actions, and routes. Search
|
|
357
|
+
`source_code` for textual matches beyond names. Supporting versions distinguish
|
|
358
|
+
exact totals from a bounded result prefix; `partial` means this is discovery,
|
|
359
|
+
not an exhaustive list. Completeness metadata is unreleased after `2.0.0.beta2`;
|
|
360
|
+
see the [search contract](MCP_SERVERS.md#search-completeness).
|
|
346
361
|
|
|
347
362
|
---
|
|
348
363
|
|
data/docs/NOTION_INTEGRATION.md
CHANGED
|
@@ -34,6 +34,19 @@ Sync is incremental. A **sync manifest** (`<output_dir>/notion_sync_manifest.jso
|
|
|
34
34
|
|
|
35
35
|
Manifest entries for models/columns that vanished from the current extraction are pruned, but **no Notion page is ever deleted** by the sync, there is no delete path. A renamed or removed model just leaves its old page in Notion untouched.
|
|
36
36
|
|
|
37
|
+
Models and migration dates are loaded by `(identifier, type)`. A same-named
|
|
38
|
+
PORO, library unit, or other extracted type cannot supply a model's table or
|
|
39
|
+
column payload. Public identifiers, page titles, and manifest keys are unchanged.
|
|
40
|
+
If a listed model or migration cannot be read with its exact identity, the
|
|
41
|
+
affected sync refuses before mapping pages or pruning its manifest. Preserve
|
|
42
|
+
the error, validate the index, and regenerate it before retrying; force sync
|
|
43
|
+
does not bypass identity checks. Each public sync method, including standalone
|
|
44
|
+
model or column sync, uses one validated generation snapshot through mapping
|
|
45
|
+
and API calls. Native readers validate the published index before selecting
|
|
46
|
+
model/migration payloads. Custom readers must provide complete published
|
|
47
|
+
enumeration or complete per-bucket listings plus
|
|
48
|
+
`find_unit(identifier, type:)` returning the requested identity.
|
|
49
|
+
|
|
37
50
|
If the manifest is missing (first run, or a CI cache miss), the exporter falls back to the full lookup/create path for every page and rebuilds the manifest, correct, just more API calls than a steady-state run.
|
|
38
51
|
|
|
39
52
|
### Escape hatch
|
|
@@ -8,7 +8,7 @@ ways at once:
|
|
|
8
8
|
Obsidian [Bases](https://help.obsidian.md/bases) table, and drill into a single unit's note with its
|
|
9
9
|
dependencies and dependents as clickable wikilinks.
|
|
10
10
|
- **By agents**: load the entire dependency topology from a single `_woods/` sidecar (one read,
|
|
11
|
-
no per-note fan-out), with a stable
|
|
11
|
+
no per-note fan-out), with a stable typed-unit → note-path manifest for navigation.
|
|
12
12
|
|
|
13
13
|
Unlike the Notion and Unblocked exporters, this one writes **local files only**: there is no API
|
|
14
14
|
token, no network call, and no rate limit. An Obsidian vault is just a folder.
|
|
@@ -95,6 +95,33 @@ Wikilinks are path-qualified with an alias (`[[models/Account|Account]]`): the t
|
|
|
95
95
|
sanitized vault path so the link always resolves, and the alias shows the original identifier. The
|
|
96
96
|
note's `# H1` carries the clean identifier so the sanitized filename never shows as the title.
|
|
97
97
|
|
|
98
|
+
## Sidecar identity and schema versions
|
|
99
|
+
|
|
100
|
+
A unit keeps its original `id` and `type` in note frontmatter. Different types may
|
|
101
|
+
share an identifier: a database view and a factory called `reports` export to
|
|
102
|
+
`database_views/reports.md` and `factories/reports.md`. Existing unambiguous note
|
|
103
|
+
paths and display aliases stay unchanged.
|
|
104
|
+
|
|
105
|
+
Check `schema_version` before consuming `_woods/manifest.json`:
|
|
106
|
+
|
|
107
|
+
- **Version 1:** the graph has no cross-type identifier collisions. The existing
|
|
108
|
+
`notes[id]` and `paths[path]` maps are unchanged.
|
|
109
|
+
- **Version 2:** `notes[id]` retains the graph's primary type (when exported), and
|
|
110
|
+
`variants` lists each additional exported unit as
|
|
111
|
+
`{ "identifier": "reports", "type": "factory", "path": "factories/reports.md" }`.
|
|
112
|
+
Combine primary entries and variants using **`(identifier, type)`** as identity.
|
|
113
|
+
`paths[path]` still gives the original identifier, so several paths may return
|
|
114
|
+
the same value. Resolve a path against both collections; `notes[id]` alone is
|
|
115
|
+
incomplete. A consumer supporting only version 1 must refuse version 2.
|
|
116
|
+
|
|
117
|
+
The exporter reads each variant's own outgoing graph edges and derives incoming
|
|
118
|
+
links from them. Persisted targets are bare identifiers: if a target names several
|
|
119
|
+
types, the exporter omits that relationship from note links and counts it in a
|
|
120
|
+
progress diagnostic. It does this even when one sibling is excluded or unreadable.
|
|
121
|
+
Human association links use the same ambiguity rule. Bare PageRank and graph-analysis
|
|
122
|
+
annotations are omitted for ambiguous identifiers, rather than assigned to either
|
|
123
|
+
type. The verbatim graph sidecar retains all original records for inspection.
|
|
124
|
+
|
|
98
125
|
## The three visualizer surfaces
|
|
99
126
|
|
|
100
127
|
| Surface | Best for | Notes |
|
|
@@ -114,7 +141,7 @@ credential):
|
|
|
114
141
|
| `WOODS_OUTPUT` | `config.output_dir` (`tmp/woods`) | extraction directory to read from |
|
|
115
142
|
| `WOODS_OBSIDIAN_VAULT` | `<output>/obsidian_vault` | where to write the vault |
|
|
116
143
|
| `WOODS_OBSIDIAN_INCLUDE_SOURCE` | off | embed each unit's source code (credential-scrubbed) in its note |
|
|
117
|
-
| `WOODS_OBSIDIAN_INCLUDE_FRAMEWORK` | off | include `rails_source` units (large; off by default) |
|
|
144
|
+
| `WOODS_OBSIDIAN_INCLUDE_FRAMEWORK` | off | include `rails_source` and readable `gem_source` units (large; off by default) |
|
|
118
145
|
| `WOODS_OBSIDIAN_FORCE_PURGE` | off | bypass the 30% mass-deletion guard during the stale-note sweep |
|
|
119
146
|
|
|
120
147
|
```bash
|
|
@@ -141,6 +168,28 @@ The exporter **fully regenerates** the vault on every run (no incremental manife
|
|
|
141
168
|
cheap). Output is deterministic: re-running against an unchanged extraction produces byte-identical
|
|
142
169
|
notes, so unchanged units never show up in a git diff.
|
|
143
170
|
|
|
171
|
+
Before writing, Woods renders the output and checks **every destination**. An existing note or index
|
|
172
|
+
must carry `woods_managed: true`; a `.woods-vault` sentinel alone never authorizes replacing your
|
|
173
|
+
notes or settings. Conflicts stop the export before any writes or sweep, appear in `errors`, and make
|
|
174
|
+
`woods:obsidian` exit nonzero. `WOODS_OBSIDIAN_FORCE_PURGE` does not bypass ownership checks.
|
|
175
|
+
Child symlinks, directories in place of files, and special files are refused.
|
|
176
|
+
|
|
177
|
+
Generated machine assets are tracked by SHA-256 in `_woods/ownership.json`: the three sidecar JSON
|
|
178
|
+
files, three `.obsidian/` JSON settings files, and `Units.base`. Existing assets must match their
|
|
179
|
+
recorded bytes or the newly generated bytes. This receipt does not change the public sidecar manifest
|
|
180
|
+
schema. Modified settings are preserved by refusing the export; the sentinel is not an override.
|
|
181
|
+
|
|
182
|
+
**Upgrading an older vault:** without a receipt, byte-identical generated assets can be adopted. Run
|
|
183
|
+
once against the same extraction to establish ownership before updating the index. If the index has
|
|
184
|
+
already changed, Woods may refuse the old sidecars because their ownership cannot be proved. Inspect
|
|
185
|
+
and back up the named files before moving them aside, or export into a new directory. Do not remove
|
|
186
|
+
personal content or manufacture a receipt to bypass a conflict.
|
|
187
|
+
|
|
188
|
+
Preflight prevents known destination conflicts from causing partial exports. This is not a multi-file
|
|
189
|
+
transaction or a lock against concurrent editors: a later I/O failure can leave some generated files
|
|
190
|
+
updated. Such failures report errors, suppress the sweep, and return zero completed-export counts;
|
|
191
|
+
inspect the vault before retrying. Avoid editing the destination while an export runs.
|
|
192
|
+
|
|
144
193
|
Notes Woods manages carry `woods_managed: true` in their frontmatter. On each run, after all notes are
|
|
145
194
|
written successfully, a **sweep** removes managed notes whose unit no longer exists, so deletions in
|
|
146
195
|
your code propagate. Several guards make the sweep safe to point at a real vault:
|
|
@@ -150,21 +199,20 @@ your code propagate. Several guards make the sweep safe to point at a real vault
|
|
|
150
199
|
- It refuses to delete more than 30% of managed notes at once (the signature of a partial extraction)
|
|
151
200
|
unless `WOODS_OBSIDIAN_FORCE_PURGE=1` is set.
|
|
152
201
|
- It resolves symlinks and confirms every deletion target is inside the vault root.
|
|
153
|
-
- It is skipped entirely if any
|
|
154
|
-
|
|
202
|
+
- It is skipped entirely if any candidate unit is unreadable or returns a mismatched identity,
|
|
203
|
+
or any note failed to write. `force_purge` does not bypass this incomplete-export guard.
|
|
155
204
|
|
|
156
|
-
|
|
157
|
-
|
|
205
|
+
In a foreign vault, Woods leaves `.obsidian/` configuration untouched and skips the sweep. Notes and
|
|
206
|
+
sidecars are added only after the same per-file ownership preflight. In a Woods-owned vault, existing
|
|
207
|
+
configuration assets still require a matching ownership receipt or byte-identical generated content.
|
|
158
208
|
|
|
159
209
|
## Limitations
|
|
160
210
|
|
|
161
211
|
- **Bases needs Obsidian ≥ 1.9** (≥ 1.10 for card/list views). The `.base` file is harmless on older
|
|
162
212
|
versions, it simply doesn't render.
|
|
163
|
-
- **`gem_source` units are not exported.** They aren't reachable through the index reader; only
|
|
164
|
-
`rails_source` is covered by `include_framework`.
|
|
165
213
|
- **Hand-edits diverge.** The vault is meant to be regenerated. Editing a note's properties in
|
|
166
214
|
Obsidian rewrites its frontmatter, after which a re-export will overwrite your changes.
|
|
167
215
|
- **Nested vaults ignore the shipped config.** Open the generated folder as its own vault for graph
|
|
168
216
|
colors, Bases, and link-format settings to apply.
|
|
169
217
|
|
|
170
|
-
|
|
218
|
+
For MCP-based exploration instead of local vault files, see [the agent guide](AGENT_GUIDE.md).
|
data/docs/PUBLISHED_INDEX.md
CHANGED
|
@@ -1,5 +1,8 @@
|
|
|
1
1
|
# Reading a published index from Ruby
|
|
2
2
|
|
|
3
|
+
For shell, Python, shipping snapshots, or direct JSON access, use the
|
|
4
|
+
[published filesystem layout contract](INDEX_LAYOUT.md).
|
|
5
|
+
|
|
3
6
|
`Woods::PublishedIndex` is the stable, read-only API for tools that are not MCP clients: RuboCop cops, CI gate scripts, and the `woods:check:*` tasks. It needs no Rails, opens one published generation, and never moves off it for the life of the reader.
|
|
4
7
|
|
|
5
8
|
```ruby
|
|
@@ -24,6 +27,58 @@ Woods::PublishedIndex.open(Rails.root.join('tmp/woods')) do |index|
|
|
|
24
27
|
end
|
|
25
28
|
```
|
|
26
29
|
|
|
30
|
+
## Validate a published generation
|
|
31
|
+
|
|
32
|
+
```ruby
|
|
33
|
+
require 'woods/resilience/index_validator'
|
|
34
|
+
|
|
35
|
+
report = Woods::Resilience::IndexValidator.new(index_dir: 'tmp/woods').validate
|
|
36
|
+
report.valid? # false when artifact or semantic graph errors were found
|
|
37
|
+
report.errors
|
|
38
|
+
report.warnings
|
|
39
|
+
```
|
|
40
|
+
|
|
41
|
+
`woods:validate` uses the same checker. Supporting development versions validate
|
|
42
|
+
raw graph relationships as well as JSON, content hashes, and indexed files; see
|
|
43
|
+
the [semantic invariants and limits](INDEX_LAYOUT.md#semantic-graph-validation).
|
|
44
|
+
Errors retain typed identifiers and artifact paths. Writer-version and source-path
|
|
45
|
+
warnings keep their existing advisory behavior.
|
|
46
|
+
|
|
47
|
+
Each call resolves and pins one published generation for all checks. Concurrent
|
|
48
|
+
publication can advance the pointer, but retention cannot remove the payload
|
|
49
|
+
being validated. The next call sees the new generation. Locks release on success
|
|
50
|
+
and errors; malformed pointers fail instead of falling back to stale root files.
|
|
51
|
+
Legacy flat layouts remain supported but cannot provide immutable-generation
|
|
52
|
+
isolation against in-place writes. Bare type-directory fixtures without a
|
|
53
|
+
manifest retain structural-only validation.
|
|
54
|
+
|
|
55
|
+
The reusable `Woods::Resilience::GraphInvariantValidator` accepts raw string-keyed
|
|
56
|
+
`graph:` data and typed `index_entries:` and returns an array of errors without
|
|
57
|
+
modifying either input. Callers supplying raw data must keep it within one pinned
|
|
58
|
+
generation and verify the unit artifacts themselves, as `IndexValidator` does.
|
|
59
|
+
Validation is a read-only diagnostic, not automatic repair or proof of runtime
|
|
60
|
+
execution.
|
|
61
|
+
|
|
62
|
+
## Manifest writer provenance
|
|
63
|
+
|
|
64
|
+
The published `manifest.json` records `woods_version`, a string naming the Woods gem version
|
|
65
|
+
that last published that manifest. Full extraction, changed incremental runs,
|
|
66
|
+
targeted refreshes, and the static Woods self-map write it. A no-op leaves the
|
|
67
|
+
published manifest and its version unchanged. Resolve the manifest through the
|
|
68
|
+
generation pointer, as Woods readers do.
|
|
69
|
+
|
|
70
|
+
Older manifests may omit the field or contain `null`; that means unknown.
|
|
71
|
+
`woods_status.index.woods_version` reports this value from the served manifest,
|
|
72
|
+
while `woods_status.server.version` identifies the running MCP reader. The reader
|
|
73
|
+
never substitutes its own version for missing writer provenance.
|
|
74
|
+
|
|
75
|
+
`woods:validate` warns when a present writer version is malformed or its major
|
|
76
|
+
version differs from the installed validator. The warning is advisory and does
|
|
77
|
+
not invalidate an otherwise structurally valid index. Re-run full extraction
|
|
78
|
+
when investigating a major-version mismatch. Matching versions do not certify
|
|
79
|
+
compatibility: an incremental publisher can retain units written by an older
|
|
80
|
+
version. This field does not replace the [upgrade procedure](UPGRADING_TO_2.md).
|
|
81
|
+
|
|
27
82
|
## One generation, pinned for the reader's whole life
|
|
28
83
|
|
|
29
84
|
Unlike `Woods::MCP::IndexReader`, a `PublishedIndex` never refreshes between calls. It resolves one generation at `.new`/`.open` time and every fact it returns, `units`, the table map, `generation_number`, `external_dependency_checksum`, comes from that one generation for as long as the reader is open. There is no `reload` and no auto-refresh: open a new reader to see a later publish.
|
data/docs/README.md
CHANGED
|
@@ -36,10 +36,13 @@ and stdio or Streamable HTTP endpoints directly.
|
|
|
36
36
|
- [MCP tool cookbook](MCP_TOOL_COOKBOOK.md): scenario-based calls with parameters and expected response shapes.
|
|
37
37
|
- [Console MCP setup](CONSOLE_MCP_SETUP.md): Console transports, blocked tables, credential scanning, redaction, SQL validation, and production safeguards.
|
|
38
38
|
- [MCP HTTP transport](MCP_HTTP_TRANSPORT.md): shared/remote Index Server transport, authentication, origins, and protocol details.
|
|
39
|
+
- [Edit client adapters](CLIENT_HOOKS.md): opt-in Claude/OpenCode registration, complete path batches, and recovery.
|
|
39
40
|
- [MCP worktree setup](MCP_WORKTREE_SETUP.md): register Woods correctly when agents work in linked git worktrees.
|
|
40
41
|
|
|
41
42
|
## Index lifecycle
|
|
42
43
|
|
|
44
|
+
- [Source freshness](SOURCE_FRESHNESS.md): verify dirty source against a served generation, establish a fresh-process baseline, and understand bounded unknown results.
|
|
45
|
+
|
|
43
46
|
- [Retrieval guide](RETRIEVAL_GUIDE.md): configure embeddings and understand semantic retrieval, ranking, and token budgets.
|
|
44
47
|
- [Embedding models](EMBEDDING_MODELS.md): choose and size local Ollama models.
|
|
45
48
|
- [Upgrade to Woods 2.0](UPGRADING_TO_2.md): identifier changes, atomic payloads, durable-store reconciliation, and rollback.
|
|
@@ -48,8 +51,10 @@ and stdio or Streamable HTTP endpoints directly.
|
|
|
48
51
|
|
|
49
52
|
- [Why Woods](WHY_WOODS.md): the problems runtime introspection solves.
|
|
50
53
|
- [Internals](INTERNALS.md): extraction, publication, graph, storage, retrieval, and MCP components.
|
|
54
|
+
- [Ruby runtime trace enrichment](RUNTIME_TRACING.md): record observed Ruby callers and merge trace evidence into method units.
|
|
51
55
|
- [Extractor reference](EXTRACTOR_REFERENCE.md): what each extractor produces and the edge cases it handles.
|
|
52
56
|
- [Reading a published index from Ruby](PUBLISHED_INDEX.md): the `Woods::PublishedIndex` Ruby API for cops, gate scripts, and `woods:check:*` tasks (including the moved-message check).
|
|
57
|
+
- [Published index layout](INDEX_LAYOUT.md): the filesystem contract for non-Ruby readers, with Bash/jq and Python examples, retention pins, and snapshot-copy rules.
|
|
53
58
|
- [Evaluation](EVALUATION.md): retrieval scoring, baselines, and the agent-level index on/off ablation.
|
|
54
59
|
- [Backend matrix](BACKEND_MATRIX.md): implemented provider/store combinations and their operational requirements.
|
|
55
60
|
- [Token benchmark](TOKEN_BENCHMARK.md): evidence behind Woods token-estimation defaults.
|
|
@@ -89,6 +94,8 @@ Use this map when changing behavior or documentation. Update the owner first; ot
|
|
|
89
94
|
| Contributor policy | [CONTRIBUTING.md](../CONTRIBUTING.md) |
|
|
90
95
|
| Coding-agent repository instructions | [AGENTS.md](https://github.com/lost-in-the/woods/blob/main/AGENTS.md) |
|
|
91
96
|
| Non-MCP Ruby access to a published index | [PUBLISHED_INDEX.md](PUBLISHED_INDEX.md) |
|
|
97
|
+
| Generation-bound source evidence and fresh extraction | [SOURCE_FRESHNESS.md](SOURCE_FRESHNESS.md) |
|
|
98
|
+
| Published filesystem layout for external readers | [INDEX_LAYOUT.md](INDEX_LAYOUT.md) |
|
|
92
99
|
| Evaluation harnesses | [EVALUATION.md](EVALUATION.md) |
|
|
93
100
|
|
|
94
101
|
The current public surface is generated from 35 extractors. Counts and capability claims must match `.Codex/release-v2/surface-inventory.json`, which is generated from the code and verified in CI.
|