woods 1.6.1 → 2.0.0.beta2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +2035 -0
- data/CONTRIBUTING.md +253 -87
- data/README.md +161 -513
- data/SECURITY.md +92 -0
- data/assets/woods-wordmark-white-with-bg.png +0 -0
- data/docs/AGENT_GUIDE.md +204 -0
- data/docs/AGENT_SETUP.md +205 -0
- data/docs/BACKEND_MATRIX.md +470 -0
- data/docs/CONFIGURATION_REFERENCE.md +655 -0
- data/docs/CONSOLE_MCP_SETUP.md +829 -0
- data/docs/DOCKER_SETUP.md +454 -0
- data/docs/EMBEDDING_MODELS.md +136 -0
- data/docs/EVALUATION.md +91 -0
- data/docs/EXTRACTOR_REFERENCE.md +765 -0
- data/docs/FAQ.md +544 -0
- data/docs/GETTING_STARTED.md +183 -0
- data/docs/INCREMENTAL_EXTRACTION.md +455 -0
- data/docs/INTERNALS.md +418 -0
- data/docs/MCP_HTTP_TRANSPORT.md +144 -0
- data/docs/MCP_SERVERS.md +231 -0
- data/docs/MCP_TOOL_COOKBOOK.md +987 -0
- data/docs/MCP_WORKTREE_SETUP.md +127 -0
- data/docs/NOTION_INTEGRATION.md +283 -0
- data/docs/OBSIDIAN_INTEGRATION.md +170 -0
- data/docs/PUBLISHED_INDEX.md +213 -0
- data/docs/README.md +94 -0
- data/docs/RETRIEVAL_GUIDE.md +267 -0
- data/docs/TOKEN_BENCHMARK.md +68 -0
- data/docs/TROUBLESHOOTING.md +841 -0
- data/docs/UNBLOCKED_INTEGRATION.md +279 -0
- data/docs/UPGRADING_TO_2.md +321 -0
- data/docs/WATCH_DAEMON.md +667 -0
- data/docs/WHY_WOODS.md +219 -0
- data/exe/woods-console +40 -4
- data/exe/woods-console-mcp +21 -35
- data/exe/woods-mcp +20 -7
- data/exe/woods-mcp-http +80 -11
- data/exe/woods-mcp-start +57 -52
- data/lib/generators/woods/install_generator.rb +6 -5
- data/lib/generators/woods/pgvector_generator.rb +6 -3
- data/lib/generators/woods/templates/add_pgvector_to_woods.rb.erb +29 -9
- data/lib/generators/woods/templates/create_woods_tables.rb.erb +5 -1
- data/lib/generators/woods/templates/woods.rb.tt +49 -28
- data/lib/tasks/woods.rake +622 -168
- data/lib/tasks/woods_checks.rake +107 -0
- data/lib/tasks/woods_evaluation.rake +164 -80
- data/lib/woods/ast/call_site_extractor.rb +6 -15
- data/lib/woods/ast/method_extractor.rb +19 -9
- data/lib/woods/ast/parser.rb +54 -8
- data/lib/woods/atomic_file.rb +171 -2
- data/lib/woods/builder.rb +310 -22
- data/lib/woods/cache/cache_middleware.rb +7 -2
- data/lib/woods/cache/cache_store.rb +9 -1
- data/lib/woods/cache/solid_cache_store.rb +6 -4
- data/lib/woods/change_set.rb +88 -0
- data/lib/woods/checks/generation_resolution.rb +34 -0
- data/lib/woods/checks/moved_messages.rb +186 -0
- data/lib/woods/chunking/semantic_chunker.rb +160 -18
- data/lib/woods/console/audit_logger.rb +12 -3
- data/lib/woods/console/bridge_protocol.rb +3 -16
- data/lib/woods/console/connection_manager.rb +51 -136
- data/lib/woods/console/dispatch_pipeline.rb +42 -12
- data/lib/woods/console/embedded_executor.rb +806 -149
- data/lib/woods/console/eval_guard.rb +27 -20
- data/lib/woods/console/input_contract.rb +78 -0
- data/lib/woods/console/model_validator.rb +29 -1
- data/lib/woods/console/rack_middleware.rb +65 -42
- data/lib/woods/console/redactor.rb +26 -8
- data/lib/woods/console/safe_context.rb +58 -10
- data/lib/woods/console/scope_predicate_parser.rb +41 -0
- data/lib/woods/console/server.rb +119 -247
- data/lib/woods/console/sql_noise_stripper.rb +125 -16
- data/lib/woods/console/sql_table_scanner.rb +82 -22
- data/lib/woods/console/sql_validator.rb +459 -29
- data/lib/woods/console/table_gate.rb +2 -2
- data/lib/woods/console/tool_specs.rb +463 -90
- data/lib/woods/console/tools/tier1.rb +1 -5
- data/lib/woods/console/tools/tier4.rb +18 -9
- data/lib/woods/coordination/lock_heartbeat.rb +103 -0
- data/lib/woods/coordination/pipeline_lock.rb +263 -53
- data/lib/woods/db/migrations/007_typed_snapshot_units.rb +45 -0
- data/lib/woods/db/migrator.rb +3 -9
- data/lib/woods/db/schema_version.rb +47 -2
- data/lib/woods/dependency_graph.rb +898 -64
- data/lib/woods/embedding/fake.rb +138 -0
- data/lib/woods/embedding/indexer.rb +832 -40
- data/lib/woods/embedding/openai.rb +77 -19
- data/lib/woods/embedding/provider.rb +189 -11
- data/lib/woods/embedding/text_preparer.rb +1 -1
- data/lib/woods/embedding/token_counter.rb +0 -7
- data/lib/woods/evaluation/ablation_agent_payload.rb +38 -0
- data/lib/woods/evaluation/ablation_executor.rb +67 -0
- data/lib/woods/evaluation/ablation_provenance.rb +38 -0
- data/lib/woods/evaluation/ablation_report_writer.rb +43 -0
- data/lib/woods/evaluation/ablation_runner.rb +173 -0
- data/lib/woods/evaluation/ablation_summary.rb +65 -0
- data/lib/woods/evaluation/ablation_task.rb +66 -0
- data/lib/woods/evaluation/ablation_task_set.rb +77 -0
- data/lib/woods/evaluation/ablation_timed_executor.rb +91 -0
- data/lib/woods/evaluation/ablation_worktree.rb +71 -0
- data/lib/woods/evaluation/baseline.rb +60 -0
- data/lib/woods/evaluation/baseline_runner.rb +11 -3
- data/lib/woods/evaluation/evaluator.rb +41 -8
- data/lib/woods/evaluation/query_set.rb +79 -13
- data/lib/woods/evaluation/report_generator.rb +20 -1
- data/lib/woods/export/unit_facts.rb +0 -11
- data/lib/woods/extracted_unit.rb +22 -63
- data/lib/woods/extractor.rb +2783 -238
- data/lib/woods/extractors/action_cable_extractor.rb +9 -4
- data/lib/woods/extractors/ast_source_extraction.rb +20 -2
- data/lib/woods/extractors/caching_extractor.rb +46 -12
- data/lib/woods/extractors/callback_analyzer.rb +39 -9
- data/lib/woods/extractors/component_discovery.rb +123 -0
- data/lib/woods/extractors/concern_extractor.rb +17 -3
- data/lib/woods/extractors/controller_extractor.rb +389 -29
- data/lib/woods/extractors/decorator_extractor.rb +7 -14
- data/lib/woods/extractors/engine_extractor.rb +53 -8
- data/lib/woods/extractors/event_extractor.rb +55 -4
- data/lib/woods/extractors/factory_extractor.rb +49 -11
- data/lib/woods/extractors/graphql_extractor.rb +162 -66
- data/lib/woods/extractors/i18n_extractor.rb +6 -1
- data/lib/woods/extractors/job_extractor.rb +51 -21
- data/lib/woods/extractors/lib_extractor.rb +23 -17
- data/lib/woods/extractors/line_neutralizer.rb +171 -0
- data/lib/woods/extractors/mailer_extractor.rb +9 -1
- data/lib/woods/extractors/manager_extractor.rb +19 -2
- data/lib/woods/extractors/migration_extractor.rb +22 -11
- data/lib/woods/extractors/model_extractor.rb +292 -57
- data/lib/woods/extractors/package_extractor.rb +154 -0
- data/lib/woods/extractors/phlex_extractor.rb +18 -3
- data/lib/woods/extractors/policy_extractor.rb +6 -5
- data/lib/woods/extractors/poro_extractor.rb +13 -14
- data/lib/woods/extractors/pundit_extractor.rb +3 -3
- data/lib/woods/extractors/rails_source_extractor.rb +24 -7
- data/lib/woods/extractors/rake_task_extractor.rb +158 -30
- data/lib/woods/extractors/reference_patterns.rb +38 -0
- data/lib/woods/extractors/route_extractor.rb +58 -2
- data/lib/woods/extractors/scheduled_job_extractor.rb +51 -35
- data/lib/woods/extractors/serializer_extractor.rb +3 -4
- data/lib/woods/extractors/service_extractor.rb +11 -1
- data/lib/woods/extractors/shared_dependency_scanner.rb +24 -34
- data/lib/woods/extractors/shared_utility_methods.rb +36 -6
- data/lib/woods/extractors/source_nesting.rb +560 -0
- data/lib/woods/extractors/state_machine_extractor.rb +30 -18
- data/lib/woods/extractors/test_mapping_extractor.rb +26 -9
- data/lib/woods/extractors/view_component_extractor.rb +28 -3
- data/lib/woods/extractors/view_engines/erb.rb +17 -3
- data/lib/woods/feedback/gap_detector.rb +9 -3
- data/lib/woods/feedback/store.rb +7 -1
- data/lib/woods/filename_utils.rb +29 -1
- data/lib/woods/flow_analysis/operation_extractor.rb +22 -10
- data/lib/woods/flow_assembler.rb +147 -26
- data/lib/woods/flow_document.rb +1 -0
- data/lib/woods/flow_precomputer.rb +175 -22
- data/lib/woods/gem_mapper.rb +285 -0
- data/lib/woods/generation.rb +185 -0
- data/lib/woods/git_command.rb +38 -0
- data/lib/woods/git_provenance.rb +16 -2
- data/lib/woods/graph_analyzer.rb +564 -87
- data/lib/woods/index_artifact.rb +93 -23
- data/lib/woods/mcp/bearer_auth.rb +102 -13
- data/lib/woods/mcp/bootstrap_state.rb +77 -0
- data/lib/woods/mcp/bootstrapper.rb +582 -77
- data/lib/woods/mcp/config_resolver.rb +66 -6
- data/lib/woods/mcp/errors.rb +60 -0
- data/lib/woods/mcp/index_reader.rb +836 -117
- data/lib/woods/mcp/index_reader_pinning.rb +78 -0
- data/lib/woods/mcp/origin_guard.rb +66 -7
- data/lib/woods/mcp/protocol_policy.rb +98 -0
- data/lib/woods/mcp/provider_probe.rb +45 -6
- data/lib/woods/mcp/renderers/markdown_renderer.rb +72 -4
- data/lib/woods/mcp/renderers/plain_renderer.rb +54 -6
- data/lib/woods/mcp/server.rb +898 -152
- data/lib/woods/mcp/tasks/extension.rb +196 -0
- data/lib/woods/mcp/tasks/request_capture.rb +45 -0
- data/lib/woods/mcp/tasks/store.rb +518 -0
- data/lib/woods/mcp/tool_contract.rb +171 -0
- data/lib/woods/mcp/tool_response_renderer.rb +7 -0
- data/lib/woods/model_name_cache.rb +19 -1
- data/lib/woods/notion/client.rb +132 -36
- data/lib/woods/notion/exporter.rb +456 -61
- data/lib/woods/notion/mappers/column_mapper.rb +34 -5
- data/lib/woods/notion/mappers/migration_mapper.rb +32 -8
- data/lib/woods/notion/mappers/model_mapper.rb +21 -6
- data/lib/woods/notion/mappers/shared.rb +45 -3
- data/lib/woods/notion/sync_manifest.rb +258 -0
- data/lib/woods/obsidian/errors.rb +6 -0
- data/lib/woods/obsidian/name_mapper.rb +40 -24
- data/lib/woods/obsidian/vault_exporter.rb +103 -36
- data/lib/woods/operator/pipeline_guard.rb +118 -21
- data/lib/woods/operator/status_reporter.rb +20 -3
- data/lib/woods/path_dispatcher.rb +276 -0
- data/lib/woods/payload_store.rb +236 -0
- data/lib/woods/published_index/edge_shaper.rb +61 -0
- data/lib/woods/published_index/generation_catalog.rb +72 -0
- data/lib/woods/published_index/typed_unit_reader.rb +48 -0
- data/lib/woods/published_index.rb +287 -0
- data/lib/woods/railtie.rb +69 -30
- data/lib/woods/railtie_support.rb +167 -0
- data/lib/woods/release.rb +12 -0
- data/lib/woods/reload_policy.rb +206 -0
- data/lib/woods/resilience/circuit_breaker.rb +47 -8
- data/lib/woods/resilience/index_validator.rb +296 -10
- data/lib/woods/resilience/retryable_provider.rb +71 -6
- data/lib/woods/resolved_config.rb +55 -11
- data/lib/woods/retrieval/context_assembler.rb +132 -40
- data/lib/woods/retrieval/query_classifier.rb +26 -8
- data/lib/woods/retrieval/ranker.rb +193 -28
- data/lib/woods/retrieval/search_executor.rb +206 -39
- data/lib/woods/retriever.rb +317 -71
- data/lib/woods/retry_after.rb +22 -2
- data/lib/woods/ruby_analyzer/class_analyzer.rb +10 -14
- data/lib/woods/ruby_analyzer/fqn_builder.rb +2 -0
- data/lib/woods/ruby_analyzer/mermaid_renderer.rb +14 -4
- data/lib/woods/ruby_analyzer/method_analyzer.rb +1 -1
- data/lib/woods/ruby_analyzer/trace_enricher.rb +3 -0
- data/lib/woods/ruby_analyzer.rb +21 -5
- data/lib/woods/session_tracer/file_store.rb +138 -19
- data/lib/woods/session_tracer/middleware.rb +1 -2
- data/lib/woods/session_tracer/redis_store.rb +122 -12
- data/lib/woods/session_tracer/session_flow_assembler.rb +57 -17
- data/lib/woods/session_tracer/session_flow_document.rb +56 -14
- data/lib/woods/session_tracer/solid_cache_coordination.rb +192 -0
- data/lib/woods/session_tracer/solid_cache_store.rb +560 -91
- data/lib/woods/session_tracer/store.rb +14 -1
- data/lib/woods/storage/metadata_store.rb +230 -26
- data/lib/woods/storage/pgvector.rb +180 -22
- data/lib/woods/storage/qdrant.rb +367 -41
- data/lib/woods/storage/snapshotter/metadata.rb +79 -16
- data/lib/woods/storage/snapshotter/vector.rb +128 -17
- data/lib/woods/storage/snapshotter.rb +23 -5
- data/lib/woods/storage/vector_store.rb +49 -8
- data/lib/woods/storage_identity.rb +28 -0
- data/lib/woods/tasks.rb +53 -2
- data/lib/woods/temporal/json_snapshot_store.rb +112 -42
- data/lib/woods/temporal/snapshot_store.rb +139 -42
- data/lib/woods/unblocked/client.rb +119 -17
- data/lib/woods/unblocked/document_builder.rb +34 -2
- data/lib/woods/unblocked/exporter.rb +63 -27
- data/lib/woods/unblocked/rate_limiter.rb +23 -9
- data/lib/woods/unblocked/sync_manifest.rb +16 -8
- data/lib/woods/update_check.rb +24 -1
- data/lib/woods/util/uuid5.rb +124 -0
- data/lib/woods/version.rb +1 -1
- data/lib/woods/watch/daemon.rb +1345 -0
- data/lib/woods/watch/listen_watcher.rb +81 -0
- data/lib/woods/watch/polling_watcher.rb +137 -0
- data/lib/woods/watch/status.rb +169 -0
- data/lib/woods/watch/tree_scan.rb +163 -0
- data/lib/woods/watch/watcher.rb +100 -0
- data/lib/woods.rb +138 -9
- data/plugin/.claude-plugin/plugin.json +18 -0
- data/plugin/hooks/hooks.json +29 -0
- data/plugin/hooks/woods-post-edit.sh +226 -0
- data/plugin/hooks/woods-session-start.sh +77 -0
- data/plugin/skills/woods-agent-enable/SKILL.md +51 -0
- data/plugin/skills/woods-diagnose/SKILL.md +75 -0
- data/plugin/skills/woods-investigate/SKILL.md +39 -0
- data/plugin/skills/woods-mcp-config/SKILL.md +101 -0
- data/plugin/skills/woods-setup/SKILL.md +99 -0
- metadata +134 -23
- data/lib/woods/console/adapters/cache_adapter.rb +0 -58
- data/lib/woods/console/adapters/good_job_adapter.rb +0 -33
- data/lib/woods/console/adapters/job_adapter.rb +0 -74
- data/lib/woods/console/adapters/sidekiq_adapter.rb +0 -33
- data/lib/woods/console/adapters/solid_queue_adapter.rb +0 -33
- data/lib/woods/console/bridge.rb +0 -210
- data/lib/woods/formatting/claude_adapter.rb +0 -98
- data/lib/woods/formatting/generic_adapter.rb +0 -56
- data/lib/woods/formatting/gpt_adapter.rb +0 -64
- data/lib/woods/notion/mapper.rb +0 -40
- data/lib/woods/observability/health_check.rb +0 -79
- data/lib/woods/observability/instrumentation.rb +0 -34
data/docs/FAQ.md
ADDED
|
@@ -0,0 +1,544 @@
|
|
|
1
|
+
# Frequently Asked Questions
|
|
2
|
+
|
|
3
|
+
---
|
|
4
|
+
|
|
5
|
+
## Question index
|
|
6
|
+
|
|
7
|
+
**General**
|
|
8
|
+
- [Does Woods work without Rails?](#does-woods-work-without-rails)
|
|
9
|
+
- [What Rails versions does Woods support?](#what-rails-versions-does-woods-support)
|
|
10
|
+
- [Does Woods work with MySQL?](#does-woods-work-with-mysql)
|
|
11
|
+
- [How large a codebase can Woods handle?](#how-large-a-codebase-can-woods-handle)
|
|
12
|
+
- [Does extraction modify my database?](#does-extraction-modify-my-database)
|
|
13
|
+
- [Can I run Woods in production?](#can-i-run-woods-in-production)
|
|
14
|
+
**Setup**
|
|
15
|
+
- [How do I install Woods?](#how-do-i-install-woods)
|
|
16
|
+
- [What is the minimum configuration?](#what-is-the-minimum-configuration)
|
|
17
|
+
- [How do I set up Woods in my MCP client?](#how-do-i-set-up-woods-in-my-mcp-client)
|
|
18
|
+
**Extraction**
|
|
19
|
+
- [What does Woods extract?](#what-does-woods-extract)
|
|
20
|
+
- [Why does Woods inline concerns?](#why-does-woods-inline-concerns)
|
|
21
|
+
- [How do I update the index after code changes?](#how-do-i-update-the-index-after-code-changes)
|
|
22
|
+
- [How do I add semantic search with embeddings?](#how-do-i-add-semantic-search-with-embeddings)
|
|
23
|
+
- [Why do some extractor types re-run wholesale on incremental runs?](#why-do-some-extractor-types-re-run-wholesale-on-incremental-runs)
|
|
24
|
+
- [How long does extraction take?](#how-long-does-extraction-take)
|
|
25
|
+
**MCP Servers**
|
|
26
|
+
- [What's the difference between the Index Server and the Console Server?](#whats-the-difference-between-the-index-server-and-the-console-server)
|
|
27
|
+
- [Why do I only see 9 console tools?](#why-do-i-only-see-9-console-tools)
|
|
28
|
+
- [Is the Console Server safe to use?](#is-the-console-server-safe-to-use)
|
|
29
|
+
- [How do I get access to SQL and structured query tools?](#how-do-i-get-access-to-sql-and-structured-query-tools)
|
|
30
|
+
- [Why do my parallel tool calls all fail when only one has a bad argument?](#why-do-my-parallel-tool-calls-all-fail-when-only-one-has-a-bad-argument)
|
|
31
|
+
**Docker**
|
|
32
|
+
- [Does extraction run inside or outside the container?](#does-extraction-run-inside-or-outside-the-container)
|
|
33
|
+
- [Why does the Index Server say no published manifest exists after extraction?](#why-does-the-index-server-say-no-published-manifest-exists-after-extraction)
|
|
34
|
+
- [How do I configure the Console Server with Docker?](#how-do-i-configure-the-console-server-with-docker)
|
|
35
|
+
**Storage and Embeddings**
|
|
36
|
+
- [What storage backends does Woods support?](#what-storage-backends-does-woods-support)
|
|
37
|
+
- [What embedding providers does Woods support?](#what-embedding-providers-does-woods-support)
|
|
38
|
+
- [What are the storage presets?](#what-are-the-storage-presets)
|
|
39
|
+
- [What happens if I change my embedding model after indexing?](#what-happens-if-i-change-my-embedding-model-after-indexing)
|
|
40
|
+
**Retrieval**
|
|
41
|
+
- [How does semantic search work?](#how-does-semantic-search-work)
|
|
42
|
+
- [What is the `codebase_retrieve` tool for?](#what-is-the-codebase_retrieve-tool-for)
|
|
43
|
+
- [How do I improve retrieval quality?](#how-do-i-improve-retrieval-quality)
|
|
44
|
+
**Temporal Snapshots**
|
|
45
|
+
- [What are temporal snapshots?](#what-are-temporal-snapshots)
|
|
46
|
+
**Session Tracing**
|
|
47
|
+
- [What does the session tracer do?](#what-does-the-session-tracer-do)
|
|
48
|
+
**Operations**
|
|
49
|
+
- [How do I keep the index in sync in CI?](#how-do-i-keep-the-index-in-sync-in-ci)
|
|
50
|
+
- [How do I check if the index is healthy?](#how-do-i-check-if-the-index-is-healthy)
|
|
51
|
+
- [Can I add custom extractors?](#can-i-add-custom-extractors)
|
|
52
|
+
- [How do I exclude sensitive directories from extraction?](#how-do-i-exclude-sensitive-directories-from-extraction)
|
|
53
|
+
|
|
54
|
+
## General
|
|
55
|
+
|
|
56
|
+
### Does Woods work without Rails?
|
|
57
|
+
|
|
58
|
+
No. Woods requires a booted Rails environment for extraction. It uses runtime introspection APIs (`ActiveRecord::Base.descendants`, `Rails.application.routes`, reflection APIs) that only exist inside a running Rails application. Static analysis of source files alone cannot produce the accurate, inlined output that Woods generates. The MCP Index Server does *not* require Rails, it reads pre-extracted JSON from disk, but the extraction step itself always does.
|
|
59
|
+
|
|
60
|
+
---
|
|
61
|
+
|
|
62
|
+
### What Rails versions does Woods support?
|
|
63
|
+
|
|
64
|
+
Woods supports **Rails 6.0 and newer**, on **Ruby 3.0 through 4.0**. CI runs an end-to-end extraction against Rails 6.0, 6.1, 7.0, 7.1, 7.2, 8.0, and 8.1 (see the version matrix in [CONTRIBUTING.md](../CONTRIBUTING.md)). The gem declares `railties >= 6.0`; the only 6.1-introduced APIs it touches (`connection_db_config`, `has_many_inversing`) are `respond_to?`-guarded and degrade cleanly on 6.0.
|
|
65
|
+
|
|
66
|
+
---
|
|
67
|
+
|
|
68
|
+
### Does Woods work with MySQL?
|
|
69
|
+
|
|
70
|
+
Yes. MySQL, PostgreSQL, and SQLite are all supported equally as application databases. Woods extraction uses ActiveRecord's database-agnostic reflection APIs and never issues raw SQL during extraction. The only backend-specific requirement is pgvector, which is PostgreSQL-only and optional. All other storage backends (SQLite metadata store, Qdrant, in-memory) work identically with MySQL and PostgreSQL. See [BACKEND_MATRIX.md](BACKEND_MATRIX.md) for the full compatibility matrix.
|
|
71
|
+
|
|
72
|
+
---
|
|
73
|
+
|
|
74
|
+
### How large a codebase can Woods handle?
|
|
75
|
+
|
|
76
|
+
Woods does not publish a supported size ceiling or a universal timing estimate. Extraction cost depends on application boot/eager-load time, enabled framework-source indexing, and codebase shape. Measure a full extraction in your application, then use incremental extraction or the watch daemon for ordinary changes.
|
|
77
|
+
|
|
78
|
+
---
|
|
79
|
+
|
|
80
|
+
### Does extraction modify my database?
|
|
81
|
+
|
|
82
|
+
No. Extraction is entirely read-only. It uses ActiveRecord reflection APIs (`columns`, `reflect_on_all_associations`, `_validators`, etc.) rather than running queries against application data. No records are created, modified, or deleted during extraction.
|
|
83
|
+
|
|
84
|
+
---
|
|
85
|
+
|
|
86
|
+
### Can I run Woods in production?
|
|
87
|
+
|
|
88
|
+
Prefer extraction in development or CI and publish the generated index as a controlled artifact. The Index Server is read-only but its index contains source and schema context; HTTP deployments still require authentication, origin controls, and TLS. Leave the live-data Console Server disabled in production unless a deliberate security review and access policy authorize it. See [MCP HTTP transport](MCP_HTTP_TRANSPORT.md) and [Console MCP setup](CONSOLE_MCP_SETUP.md).
|
|
89
|
+
|
|
90
|
+
---
|
|
91
|
+
|
|
92
|
+
## Setup
|
|
93
|
+
|
|
94
|
+
### How do I install Woods?
|
|
95
|
+
|
|
96
|
+
Add `gem "woods", "~> 2.0"` to your development group, install it, and run the install generator. Review the initializer. For a new default v2 install, remove the generated legacy application migration instead of running it; shipped v2 extraction and retrieval do not use those tables. Keep it only for a known older/custom integration. Then extract and validate. Follow [Getting started](GETTING_STARTED.md) for the canonical commands and expected result. If an agent is doing the work, use [Agent setup](AGENT_SETUP.md).
|
|
97
|
+
|
|
98
|
+
---
|
|
99
|
+
|
|
100
|
+
### What is the minimum configuration?
|
|
101
|
+
|
|
102
|
+
The generated defaults are enough for structural extraction to `tmp/woods/`. Embeddings are optional. See [Getting started](GETTING_STARTED.md) for the minimal path and [Configuration reference](CONFIGURATION_REFERENCE.md) for supported settings.
|
|
103
|
+
|
|
104
|
+
---
|
|
105
|
+
|
|
106
|
+
### How do I set up Woods in my MCP client?
|
|
107
|
+
|
|
108
|
+
Extract first, then configure a project-scoped stdio server using the application's bundle and an index path visible to the server process. Woods works with any model through a client that supports MCP stdio or Streamable HTTP. Use the verified examples in [MCP servers](MCP_SERVERS.md#configure-a-stdio-client), adapting them to the client's configuration location.
|
|
109
|
+
|
|
110
|
+
---
|
|
111
|
+
|
|
112
|
+
## Extraction
|
|
113
|
+
|
|
114
|
+
### What does Woods extract?
|
|
115
|
+
|
|
116
|
+
Woods runs **35 extractor classes** on every full extraction, there is no
|
|
117
|
+
opt-in/opt-out. Between them they produce 39 distinct unit types (some
|
|
118
|
+
extractors emit more than one type. GraphQL alone produces four, and
|
|
119
|
+
`RailsSourceExtractor` produces both `rails_source` and `gem_source`).
|
|
120
|
+
Coverage
|
|
121
|
+
includes models (with inlined concerns and schema), controllers, services,
|
|
122
|
+
view components, jobs, mailers, GraphQL types/mutations/queries, serializers,
|
|
123
|
+
managers, policies, validators, Rails framework source, state machines,
|
|
124
|
+
events, decorators, database views, rake tasks, Action Cable channels, and
|
|
125
|
+
more. See [EXTRACTOR_REFERENCE.md](EXTRACTOR_REFERENCE.md) for the full list.
|
|
126
|
+
|
|
127
|
+
---
|
|
128
|
+
|
|
129
|
+
### Why does Woods inline concerns?
|
|
130
|
+
|
|
131
|
+
When a model includes a concern, the behavior defined in that concern is part of the model's effective API, callbacks fire, validations run, scopes are available. A tool that reports only what's in `app/models/user.rb` misses everything defined in included concerns. Woods inlines concern source directly into each unit's `source_code` field so the full behavioral picture is in one place. This is the key differentiator from file-level tools.
|
|
132
|
+
|
|
133
|
+
---
|
|
134
|
+
|
|
135
|
+
### How do I update the index after code changes?
|
|
136
|
+
|
|
137
|
+
Use incremental mode, which re-extracts only files that have changed since the last run:
|
|
138
|
+
|
|
139
|
+
```bash
|
|
140
|
+
bundle exec rake woods:incremental
|
|
141
|
+
|
|
142
|
+
# Docker:
|
|
143
|
+
docker compose exec app bundle exec rake woods:incremental
|
|
144
|
+
```
|
|
145
|
+
|
|
146
|
+
Incremental mode is ideal for CI pipelines and local development workflows. It is typically 5-10× faster than a full extraction. Nine unit types, `route`, `middleware`, `engine`, `scheduled_job`, `state_machine`, `factory`, `event`, `database_view`, and `rails_source`, don't map to individual files, so incremental mode re-runs their extractor **wholesale** whenever the relevant trigger path changes (e.g. `config/routes.rb` for routes, `Gemfile.lock` for middleware/engines/rails_source; `rails_source` participates only when `include_framework_sources` is enabled). You never need to run a full extraction just because one of these changed, see [TROUBLESHOOTING.md](TROUBLESHOOTING.md) for details.
|
|
147
|
+
|
|
148
|
+
---
|
|
149
|
+
|
|
150
|
+
### How do I add semantic search with embeddings?
|
|
151
|
+
|
|
152
|
+
Configure an embedding provider, then run the embed task:
|
|
153
|
+
|
|
154
|
+
```ruby
|
|
155
|
+
# config/initializers/woods.rb
|
|
156
|
+
Woods.configure do |config|
|
|
157
|
+
# OpenAI (cloud)
|
|
158
|
+
config.embedding_provider = :openai
|
|
159
|
+
config.embedding_model = 'text-embedding-3-small'
|
|
160
|
+
config.embedding_options = { api_key: ENV['OPENAI_API_KEY'] }
|
|
161
|
+
|
|
162
|
+
# Ollama (local, no API key needed)
|
|
163
|
+
# config.embedding_provider = :ollama
|
|
164
|
+
# config.embedding_model = 'nomic-embed-text'
|
|
165
|
+
end
|
|
166
|
+
```
|
|
167
|
+
|
|
168
|
+
```bash
|
|
169
|
+
bundle exec rake woods:embed
|
|
170
|
+
```
|
|
171
|
+
|
|
172
|
+
After embedding, the `codebase_retrieve` MCP tool supports natural-language queries ranked by semantic similarity. See [CONFIGURATION_REFERENCE.md](CONFIGURATION_REFERENCE.md) for vector storage options.
|
|
173
|
+
|
|
174
|
+
---
|
|
175
|
+
|
|
176
|
+
### Why do some extractor types re-run wholesale on incremental runs?
|
|
177
|
+
|
|
178
|
+
Nine unit types, routes, middleware, engines, scheduled jobs, state machines, events, factories, database views, and Rails/gem sources, don't map to individual files. They're extracted by introspecting the entire application at once rather than one file at a time. Incremental mode handles this by re-running the whole extractor when a trigger path changes (`config/routes.rb`, `Gemfile.lock`, the relevant model/factory directories, and so on) instead of skipping the type; Rails/gem sources participate only when `include_framework_sources` is enabled. A full extraction is only needed if you suspect drift, not as routine maintenance:
|
|
179
|
+
|
|
180
|
+
```bash
|
|
181
|
+
bundle exec rake woods:extract
|
|
182
|
+
```
|
|
183
|
+
|
|
184
|
+
---
|
|
185
|
+
|
|
186
|
+
### How long does extraction take?
|
|
187
|
+
|
|
188
|
+
There is no reliable application-independent estimate. Rails boot/eager load, framework-source indexing, and codebase shape dominate the result. Time `woods:extract` in your own environment and use `woods:incremental` or `woods:watch` for ordinary changes.
|
|
189
|
+
|
|
190
|
+
---
|
|
191
|
+
|
|
192
|
+
## MCP Servers
|
|
193
|
+
|
|
194
|
+
### What's the difference between the Index Server and the Console Server?
|
|
195
|
+
|
|
196
|
+
The Index Server reads pre-extracted JSON from disk and does not require Rails.
|
|
197
|
+
The Console Server connects to a live Rails application and registers 9
|
|
198
|
+
read-only model/schema tools by default, with `console_sql` and `console_query`
|
|
199
|
+
available through explicit read-tool configuration.
|
|
200
|
+
|
|
201
|
+
---
|
|
202
|
+
|
|
203
|
+
### Why do I only see 9 console tools?
|
|
204
|
+
|
|
205
|
+
That is the default supported surface: count, sample, find, pluck, aggregate,
|
|
206
|
+
association_count, schema, recent, and status. Enable
|
|
207
|
+
`console_embedded_read_tools` to register SQL and query as well. The remaining
|
|
208
|
+
schemas are inventory only.
|
|
209
|
+
|
|
210
|
+
---
|
|
211
|
+
|
|
212
|
+
### Is the Console Server safe to use?
|
|
213
|
+
|
|
214
|
+
The executable tools enforce the feature gate, blocked-table checks,
|
|
215
|
+
credential scanning, configured column/EAV redaction, and rolled-back
|
|
216
|
+
transactions. `console_sql` also runs through `SqlValidator`. These controls
|
|
217
|
+
are defense in depth, not a primary authorization boundary.
|
|
218
|
+
|
|
219
|
+
---
|
|
220
|
+
|
|
221
|
+
### How do I get access to SQL and structured query tools?
|
|
222
|
+
|
|
223
|
+
Set `config.console_embedded_read_tools = true` in `Woods.configure`, or pass
|
|
224
|
+
`embedded_read_tools: true` when mounting the Rack middleware. Tier 2, Tier 3,
|
|
225
|
+
and `console_eval` are not registered in a supported mode.
|
|
226
|
+
|
|
227
|
+
See [CONSOLE_MCP_SETUP.md](CONSOLE_MCP_SETUP.md) for setup details.
|
|
228
|
+
|
|
229
|
+
---
|
|
230
|
+
|
|
231
|
+
### Why do my parallel tool calls all fail when only one has a bad argument?
|
|
232
|
+
|
|
233
|
+
This is an MCP client behavior, not a server bug. Some clients batch parallel tool calls into a single protocol request. If one call in the batch fails (e.g., a typo in an identifier), the transport layer may reject the entire response. There is no server-side fix, the workaround is to validate arguments first (use `search` to confirm identifiers exist) or send calls sequentially when any call is unreliable. See the [Troubleshooting guide](TROUBLESHOOTING.md#parallel-tool-calls-fail-together-sibling-call-failures) for details.
|
|
234
|
+
|
|
235
|
+
---
|
|
236
|
+
|
|
237
|
+
## Docker
|
|
238
|
+
|
|
239
|
+
### Does extraction run inside or outside the container?
|
|
240
|
+
|
|
241
|
+
Extraction runs **inside** the container because it requires Rails to be booted. By default, launch the Index Server through the same application container so it can use the installed bundle and container index path; it still reads only static files and does not boot Rails. A host-side Index Server is optional when the host has the application bundle and can read the index volume. The Console Server connects to a booted Rails process inside the container.
|
|
242
|
+
|
|
243
|
+
```
|
|
244
|
+
HOST APPLICATION CONTAINER
|
|
245
|
+
───────────────── ─────────────────────────
|
|
246
|
+
MCP client ── compose exec ──▶ woods-mcp (reads JSON)
|
|
247
|
+
Rails commands ──────────────▶ rake extract (writes JSON)
|
|
248
|
+
Console client ──────────────▶ rake console (queries Rails)
|
|
249
|
+
```
|
|
250
|
+
|
|
251
|
+
See [DOCKER_SETUP.md](DOCKER_SETUP.md) for the full Docker architecture guide.
|
|
252
|
+
|
|
253
|
+
---
|
|
254
|
+
|
|
255
|
+
### Why does the Index Server say no published manifest exists after extraction?
|
|
256
|
+
|
|
257
|
+
The server is usually running in a different filesystem context from the path you supplied. A Docker-launched server needs the container path; an optional host-launched server needs a host-visible path and the application bundle installed on the host.
|
|
258
|
+
|
|
259
|
+
```jsonc
|
|
260
|
+
{
|
|
261
|
+
"mcpServers": {
|
|
262
|
+
"codebase": {
|
|
263
|
+
"command": "docker",
|
|
264
|
+
"args": ["compose", "exec", "-T", "app", "bundle", "exec", "woods-mcp", "/app/tmp/woods"],
|
|
265
|
+
"cwd": "/absolute/host/path/to/app"
|
|
266
|
+
}
|
|
267
|
+
}
|
|
268
|
+
}
|
|
269
|
+
```
|
|
270
|
+
|
|
271
|
+
Verify the active v2 generation with `docker compose exec app bundle exec rake woods:validate` and `woods:stats`. Do not require a root `manifest.json`: Woods 2 normally publishes root `generation.json`, which points to the active payload manifest.
|
|
272
|
+
|
|
273
|
+
---
|
|
274
|
+
|
|
275
|
+
### How do I configure the Console Server with Docker?
|
|
276
|
+
|
|
277
|
+
First set `config.console_mcp_enabled = true` in the Rails initializer after reviewing the live-data trust boundary. Stdio does not send a bearer token, but production Rails boot still requires a configured `console_mcp_token` of at least 32 characters whenever Console is enabled. Supply `WOODS_CONSOLE_MCP_TOKEN` through the container's secret mechanism; see [Console MCP setup](CONSOLE_MCP_SETUP.md#option-a-stdio-via-rake-recommended). Then, for the embedded mode (9 Tier 1 tools), point the MCP client at `docker compose exec -T` so Compose does not allocate a pseudo-TTY:
|
|
278
|
+
|
|
279
|
+
```json
|
|
280
|
+
{
|
|
281
|
+
"mcpServers": {
|
|
282
|
+
"codebase-console": {
|
|
283
|
+
"command": "docker",
|
|
284
|
+
"args": ["compose", "exec", "-T", "app",
|
|
285
|
+
"bundle", "exec", "rake", "woods:console"],
|
|
286
|
+
"cwd": "/absolute/host/path/to/app"
|
|
287
|
+
}
|
|
288
|
+
}
|
|
289
|
+
}
|
|
290
|
+
```
|
|
291
|
+
|
|
292
|
+
Compose attaches stdin by default; `-T` disables the pseudo-TTY that would corrupt MCP framing. Plain `docker exec` uses `-i` instead. Enable `console_embedded_read_tools` when SQL/query should also be registered. See [DOCKER_SETUP.md](DOCKER_SETUP.md) for complete examples.
|
|
293
|
+
|
|
294
|
+
---
|
|
295
|
+
|
|
296
|
+
## Storage and Embeddings
|
|
297
|
+
|
|
298
|
+
### What storage backends does Woods support?
|
|
299
|
+
|
|
300
|
+
Woods supports three vector storage backends and two metadata backends:
|
|
301
|
+
|
|
302
|
+
| Backend | Type | Use case |
|
|
303
|
+
|---------|------|----------|
|
|
304
|
+
| `in_memory` | Vector + Metadata | Local dev, no persistence needed |
|
|
305
|
+
| `sqlite` | Metadata | Persistent metadata, simple setup |
|
|
306
|
+
| `pgvector` | Vector | PostgreSQL apps wanting unified storage |
|
|
307
|
+
| `qdrant` | Vector | Production-scale semantic search |
|
|
308
|
+
|
|
309
|
+
All backends work with both MySQL and PostgreSQL application databases. pgvector requires PostgreSQL for the vector store, but your application database can still be MySQL. See [BACKEND_MATRIX.md](BACKEND_MATRIX.md) for the full compatibility matrix.
|
|
310
|
+
|
|
311
|
+
---
|
|
312
|
+
|
|
313
|
+
### What embedding providers does Woods support?
|
|
314
|
+
|
|
315
|
+
Three embedding providers are supported:
|
|
316
|
+
|
|
317
|
+
- **OpenAI**: `text-embedding-3-small` (1536 dimensions, default) or `text-embedding-3-large`. Requires an `OPENAI_API_KEY`. Billed per token.
|
|
318
|
+
- **Ollama**: Any locally installed model (e.g., `nomic-embed-text`, `bge-m3`, `mxbai-embed-large`). Runs locally, no API key or cost. Requires Ollama to be running at `localhost:11434`.
|
|
319
|
+
- **Fake** (`:fake`): deterministic vectors with no network or service, for exercising the embed pipeline in tests and smoke checks. See [Configuration reference](CONFIGURATION_REFERENCE.md).
|
|
320
|
+
|
|
321
|
+
A provider object responding to `#embed` and `#embed_batch` can also be injected directly.
|
|
322
|
+
|
|
323
|
+
```ruby
|
|
324
|
+
# OpenAI
|
|
325
|
+
config.embedding_provider = :openai
|
|
326
|
+
config.embedding_options = {
|
|
327
|
+
api_key: ENV['OPENAI_API_KEY'],
|
|
328
|
+
model: 'text-embedding-3-small'
|
|
329
|
+
}
|
|
330
|
+
|
|
331
|
+
# Ollama: default (2048-token context)
|
|
332
|
+
config.embedding_provider = :ollama
|
|
333
|
+
config.embedding_options = { model: 'nomic-embed-text' }
|
|
334
|
+
|
|
335
|
+
# Ollama: bge-m3 (8192-token context, fewer chunks per unit)
|
|
336
|
+
config.embedding_provider = :ollama
|
|
337
|
+
config.embedding_options = { model: 'bge-m3' }
|
|
338
|
+
```
|
|
339
|
+
|
|
340
|
+
See [EMBEDDING_MODELS.md](EMBEDDING_MODELS.md) for the Ollama model comparison and the procedure for adding a new model.
|
|
341
|
+
|
|
342
|
+
---
|
|
343
|
+
|
|
344
|
+
### What are the storage presets?
|
|
345
|
+
|
|
346
|
+
Presets configure storage and embedding together with a single call:
|
|
347
|
+
|
|
348
|
+
```ruby
|
|
349
|
+
# Local services: in-memory vectors, SQLite metadata, Ollama embeddings
|
|
350
|
+
# Requires the sqlite3 gem, a running Ollama service, and the pulled model.
|
|
351
|
+
Woods.configure_with_preset(:local)
|
|
352
|
+
|
|
353
|
+
# Shared filesystem: in-memory stores persisted under output_dir.
|
|
354
|
+
# Requires a running Ollama service and shared output_dir; no sqlite3 gem.
|
|
355
|
+
Woods.configure_with_preset(:shared_filesystem)
|
|
356
|
+
|
|
357
|
+
# PostgreSQL + OpenAI: pgvector vectors, SQLite metadata, OpenAI embeddings
|
|
358
|
+
Woods.configure_with_preset(:postgresql) do |config|
|
|
359
|
+
config.embedding_options = { api_key: ENV.fetch('OPENAI_API_KEY') }
|
|
360
|
+
config.vector_store_options = { connection: ActiveRecord::Base.connection }
|
|
361
|
+
end
|
|
362
|
+
|
|
363
|
+
# Production scale: Qdrant vectors, SQLite metadata, OpenAI embeddings
|
|
364
|
+
Woods.configure_with_preset(:production) do |config|
|
|
365
|
+
config.embedding_options = { api_key: ENV.fetch('OPENAI_API_KEY') }
|
|
366
|
+
config.vector_store_options = {
|
|
367
|
+
url: ENV.fetch('QDRANT_URL'),
|
|
368
|
+
collection: ENV.fetch('WOODS_QDRANT_COLLECTION', 'woods'),
|
|
369
|
+
allow_private_hosts: true # only when QDRANT_URL is deliberately private
|
|
370
|
+
}
|
|
371
|
+
end
|
|
372
|
+
```
|
|
373
|
+
|
|
374
|
+
Presets can be overridden with a block:
|
|
375
|
+
|
|
376
|
+
```ruby
|
|
377
|
+
Woods.configure_with_preset(:local) do |config|
|
|
378
|
+
config.max_context_tokens = 16000
|
|
379
|
+
config.embedding_model = 'mxbai-embed-large'
|
|
380
|
+
end
|
|
381
|
+
```
|
|
382
|
+
|
|
383
|
+
Start with `:local` for zero-dependency development and upgrade to `:postgresql` or `:production` when you need persistence or scale. See [CONFIGURATION_REFERENCE.md](CONFIGURATION_REFERENCE.md) for what each preset configures.
|
|
384
|
+
|
|
385
|
+
---
|
|
386
|
+
|
|
387
|
+
### What happens if I change my embedding model after indexing?
|
|
388
|
+
|
|
389
|
+
Switching embedding models requires a full re-index. The new model produces vectors with different dimensions or a different embedding space, making old and new vectors incompatible for similarity search. Woods detects a dimension change and raises `Woods::MCP::DimensionMismatch`, `rake woods:embed` refuses before embedding anything, and the MCP server refuses at boot, so you get an actionable error rather than silently wrong results. Re-index with:
|
|
390
|
+
|
|
391
|
+
```bash
|
|
392
|
+
bundle exec rake woods:extract
|
|
393
|
+
bundle exec rake woods:embed
|
|
394
|
+
```
|
|
395
|
+
|
|
396
|
+
---
|
|
397
|
+
|
|
398
|
+
## Retrieval
|
|
399
|
+
|
|
400
|
+
### How does semantic search work?
|
|
401
|
+
|
|
402
|
+
When you run `rake woods:embed`, Woods generates embedding vectors for each extracted unit and stores them in your configured vector store. The `codebase_retrieve` MCP tool accepts a natural-language query, embeds the query using the same provider, and finds the most semantically similar units using cosine similarity. Results are re-ranked using Reciprocal Rank Fusion (RRF) that combines semantic similarity with PageRank importance scores, then assembled into a formatted context block within your configured token budget.
|
|
403
|
+
|
|
404
|
+
---
|
|
405
|
+
|
|
406
|
+
### What is the `codebase_retrieve` tool for?
|
|
407
|
+
|
|
408
|
+
`codebase_retrieve` is the primary semantic search tool on the Index Server. It accepts a natural-language description of what you're looking for ("find where user email validation happens", "which services send Stripe API calls") and returns the most relevant extracted units as formatted context. It requires embedding configuration, without an embedding provider the tool responds with an error (`isError`, code `:not_configured`) and a remediation hint covering provider setup and the `search` tool for pattern-based matching in the meantime. Token budget is controlled by `config.max_context_tokens` (default: 8000).
|
|
409
|
+
|
|
410
|
+
---
|
|
411
|
+
|
|
412
|
+
### How do I improve retrieval quality?
|
|
413
|
+
|
|
414
|
+
Several options for tuning retrieval:
|
|
415
|
+
|
|
416
|
+
- **Increase `max_context_tokens`** to include more units per query (at the cost of larger LLM context).
|
|
417
|
+
- **Lower `similarity_threshold`** (default 0.7) to include less similar results.
|
|
418
|
+
- **Enable framework sources** (`include_framework_sources: true`) if Rails internals are relevant to your queries.
|
|
419
|
+
- **Use retrieval feedback only in a custom embedded server** that wires a feedback store. The normal packaged executable does not register feedback tools.
|
|
420
|
+
|
|
421
|
+
---
|
|
422
|
+
|
|
423
|
+
## Temporal Snapshots
|
|
424
|
+
|
|
425
|
+
### What are temporal snapshots?
|
|
426
|
+
|
|
427
|
+
Temporal snapshots capture the full extraction state at a point in time, tied to a git SHA. They let you compare how the codebase has changed between snapshots, which units were added, modified, or deleted. Snapshots are opt-in and disabled by default.
|
|
428
|
+
|
|
429
|
+
Enable them in your initializer:
|
|
430
|
+
|
|
431
|
+
```ruby
|
|
432
|
+
config.enable_snapshots = true
|
|
433
|
+
```
|
|
434
|
+
|
|
435
|
+
Snapshots prefer their own SQLite database (`woods.sqlite3` in the output directory), separate from your Rails app's database. Extraction falls back to JSON files when the `sqlite3` gem is unavailable; other SQLite open/migration failures are reported and do not capture a snapshot. `Woods::Db::Migrator` runs the internal SQLite migrations automatically during extraction and MCP boot; `bundle exec rails db:migrate` does not touch this store and no manual migration step is needed. The packaged MCP server discovers an existing `woods.sqlite3` automatically. When extraction used the JSON fallback, set `WOODS_SNAPSHOTS=true` on the server so it wires `list_snapshots`, `snapshot_diff`, `unit_history`, and `snapshot_detail`.
|
|
436
|
+
|
|
437
|
+
---
|
|
438
|
+
|
|
439
|
+
## Session Tracing
|
|
440
|
+
|
|
441
|
+
### What does the session tracer do?
|
|
442
|
+
|
|
443
|
+
The session tracer is middleware that records which Rails actions are invoked during a browser session, assembles the relevant extracted units, and makes that context available via the `session_trace` MCP tool. It is useful for giving an AI tool accurate context about what code path was active during a specific user interaction.
|
|
444
|
+
|
|
445
|
+
Session tracing is disabled by default. To enable it:
|
|
446
|
+
|
|
447
|
+
```ruby
|
|
448
|
+
config.session_tracer_enabled = true
|
|
449
|
+
config.session_store = Woods::SessionTracer::FileStore.new(
|
|
450
|
+
Rails.root.join('tmp/session_traces')
|
|
451
|
+
)
|
|
452
|
+
```
|
|
453
|
+
|
|
454
|
+
The `session_store` option is required, there is no default store.
|
|
455
|
+
|
|
456
|
+
---
|
|
457
|
+
|
|
458
|
+
## Operations
|
|
459
|
+
|
|
460
|
+
### How do I keep the index in sync in CI?
|
|
461
|
+
|
|
462
|
+
Use incremental extraction in your CI pipeline. Fetch enough git history for the incremental diff to work:
|
|
463
|
+
|
|
464
|
+
```yaml
|
|
465
|
+
# .github/workflows/index.yml
|
|
466
|
+
jobs:
|
|
467
|
+
index:
|
|
468
|
+
steps:
|
|
469
|
+
- uses: actions/checkout@v4
|
|
470
|
+
with:
|
|
471
|
+
fetch-depth: 2
|
|
472
|
+
- name: Update index
|
|
473
|
+
run: bundle exec rake woods:incremental
|
|
474
|
+
env:
|
|
475
|
+
GITHUB_BASE_REF: ${{ github.base_ref }}
|
|
476
|
+
```
|
|
477
|
+
|
|
478
|
+
For Docker-based CI:
|
|
479
|
+
|
|
480
|
+
```yaml
|
|
481
|
+
- name: Update index
|
|
482
|
+
run: docker compose exec -T app bundle exec rake woods:incremental
|
|
483
|
+
```
|
|
484
|
+
|
|
485
|
+
---
|
|
486
|
+
|
|
487
|
+
### How do I check if the index is healthy?
|
|
488
|
+
|
|
489
|
+
Two rake tasks validate index integrity. Both are `:environment` tasks, they boot Rails, same as extraction:
|
|
490
|
+
|
|
491
|
+
```bash
|
|
492
|
+
bundle exec rake woods:validate
|
|
493
|
+
|
|
494
|
+
# Show unit counts and extraction stats
|
|
495
|
+
bundle exec rake woods:stats
|
|
496
|
+
```
|
|
497
|
+
|
|
498
|
+
The packaged Index Server's `woods_status` tool reports a single-call health snapshot covering generation freshness, counts, retrieval readiness, and configured optional capabilities. `pipeline_status` requires an operator collaborator that the packaged executable does not wire.
|
|
499
|
+
|
|
500
|
+
---
|
|
501
|
+
|
|
502
|
+
### Can I add custom extractors?
|
|
503
|
+
|
|
504
|
+
Not today. `Woods::Extractor::EXTRACTORS` is a frozen constant listing the 34
|
|
505
|
+
built-in extractor classes, and nothing in the extraction path consults
|
|
506
|
+
`config.extractors` to add to that list. `config.extractors` exists only for
|
|
507
|
+
forward compatibility, setting it to anything other than the default warns
|
|
508
|
+
and has no effect on which extractors run. If you need a custom extractor
|
|
509
|
+
today, the extractor interface (`initialize` + `extract_all` returning
|
|
510
|
+
`Array<ExtractedUnit>`) is stable and documented in
|
|
511
|
+
[EXTRACTOR_REFERENCE.md](EXTRACTOR_REFERENCE.md), but wiring one in requires
|
|
512
|
+
patching `EXTRACTORS` directly (a gem fork or monkeypatch), not a config call.
|
|
513
|
+
|
|
514
|
+
---
|
|
515
|
+
|
|
516
|
+
### How do I exclude sensitive directories from extraction?
|
|
517
|
+
|
|
518
|
+
`config.extractors` cannot remove an extractor from the run, see above.
|
|
519
|
+
Exclude a directory from eager loading instead:
|
|
520
|
+
|
|
521
|
+
```ruby
|
|
522
|
+
# config/application.rb
|
|
523
|
+
config.eager_load_paths -= [Rails.root.join('app/internal')]
|
|
524
|
+
```
|
|
525
|
+
|
|
526
|
+
Use `console_redacted_columns` to redact sensitive column values from Console Server results without excluding extraction:
|
|
527
|
+
|
|
528
|
+
```ruby
|
|
529
|
+
config.console_redacted_columns = %w[password_digest api_key ssn token]
|
|
530
|
+
```
|
|
531
|
+
|
|
532
|
+
---
|
|
533
|
+
|
|
534
|
+
## Troubleshooting
|
|
535
|
+
|
|
536
|
+
For detailed problem-specific guidance, see [TROUBLESHOOTING.md](TROUBLESHOOTING.md).
|
|
537
|
+
|
|
538
|
+
Quick links:
|
|
539
|
+
|
|
540
|
+
- Extraction produces empty output → [Extraction Problems](TROUBLESHOOTING.md#extraction-produces-empty-or-incomplete-output)
|
|
541
|
+
- "No manifest.json" error → [MCP Server Problems](TROUBLESHOOTING.md#no-manifestjson-error-when-starting-the-index-server)
|
|
542
|
+
- Only 9 console tools visible → [MCP Server Problems](TROUBLESHOOTING.md#a-console-inventory-tool-is-not-listed)
|
|
543
|
+
- Docker path confusion → [Docker Problems](TROUBLESHOOTING.md#path-confusion-index-server-uses-container-path)
|
|
544
|
+
- Dimension mismatch on embeddings → [Embedding Problems](TROUBLESHOOTING.md#dimension-mismatch-error-when-querying-embeddings)
|