woods 2.0.1 → 2.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (154) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +94 -7
  3. data/CONTRIBUTING.md +134 -19
  4. data/README.md +1 -1
  5. data/docs/AGENT_GUIDE.md +19 -0
  6. data/docs/AGENT_SETUP.md +22 -2
  7. data/docs/BACKEND_MATRIX.md +7 -0
  8. data/docs/CLIENT_HOOKS.md +6 -0
  9. data/docs/CONFIGURATION_REFERENCE.md +133 -25
  10. data/docs/CONSOLE_MCP_SETUP.md +82 -30
  11. data/docs/EMBEDDING_MODELS.md +16 -19
  12. data/docs/EXTRACTOR_REFERENCE.md +219 -21
  13. data/docs/FAQ.md +11 -25
  14. data/docs/GETTING_STARTED.md +7 -1
  15. data/docs/INCREMENTAL_EXTRACTION.md +261 -19
  16. data/docs/INDEX_LAYOUT.md +5 -0
  17. data/docs/INTERNALS.md +9 -0
  18. data/docs/MCP_HTTP_TRANSPORT.md +20 -15
  19. data/docs/MCP_SERVERS.md +87 -8
  20. data/docs/MCP_TOOL_COOKBOOK.md +13 -55
  21. data/docs/NOTION_INTEGRATION.md +7 -1
  22. data/docs/PUBLISHED_INDEX.md +6 -0
  23. data/docs/README.md +6 -1
  24. data/docs/RETRIEVAL_GUIDE.md +17 -0
  25. data/docs/SOURCE_FRESHNESS.md +157 -5
  26. data/docs/TOKEN_BENCHMARK.md +10 -18
  27. data/docs/TROUBLESHOOTING.md +70 -14
  28. data/docs/UNBLOCKED_INTEGRATION.md +60 -8
  29. data/docs/UPGRADING_TO_2.md +153 -38
  30. data/docs/WATCH_DAEMON.md +97 -14
  31. data/exe/woods-console-mcp +2 -2
  32. data/lib/generators/woods/templates/woods.rb.tt +2 -1
  33. data/lib/tasks/woods.rake +23 -7
  34. data/lib/tasks/woods_checks.rake +2 -2
  35. data/lib/woods/agent_configuration/cli.rb +1 -1
  36. data/lib/woods/agent_configuration/layout.rb +16 -2
  37. data/lib/woods/agent_configuration/plan.rb +13 -3
  38. data/lib/woods/agent_configuration/planner_validation.rb +4 -2
  39. data/lib/woods/agent_configuration/preflight.rb +5 -3
  40. data/lib/woods/builder.rb +17 -57
  41. data/lib/woods/cache/cache_middleware.rb +56 -30
  42. data/lib/woods/chunking/contributor_chunks.rb +119 -0
  43. data/lib/woods/chunking/semantic_chunker.rb +44 -21
  44. data/lib/woods/console/connection_manager.rb +56 -3
  45. data/lib/woods/console/embedded_executor.rb +30 -5
  46. data/lib/woods/console/rack_middleware.rb +29 -1
  47. data/lib/woods/dependency_graph.rb +34 -10
  48. data/lib/woods/embedding/fake.rb +12 -0
  49. data/lib/woods/embedding/indexer.rb +195 -98
  50. data/lib/woods/embedding/input_budget.rb +67 -0
  51. data/lib/woods/embedding/openai.rb +70 -20
  52. data/lib/woods/embedding/provider.rb +37 -25
  53. data/lib/woods/embedding/text_preparer.rb +76 -32
  54. data/lib/woods/embedding/token_counter.rb +18 -81
  55. data/lib/woods/embedding/vector_configuration.rb +48 -0
  56. data/lib/woods/extraction_identities.rb +175 -0
  57. data/lib/woods/extractor.rb +304 -107
  58. data/lib/woods/extractors/action_cable_extractor.rb +8 -3
  59. data/lib/woods/extractors/assigned_value_discovery.rb +74 -0
  60. data/lib/woods/extractors/class_declarations.rb +121 -0
  61. data/lib/woods/extractors/configuration_extractor.rb +11 -3
  62. data/lib/woods/extractors/declaration_ancestry.rb +92 -0
  63. data/lib/woods/extractors/event_extractor.rb +8 -0
  64. data/lib/woods/extractors/graphql_extractor.rb +134 -77
  65. data/lib/woods/extractors/job_extractor.rb +5 -1
  66. data/lib/woods/extractors/lib_extractor.rb +132 -15
  67. data/lib/woods/extractors/mailer_extractor.rb +3 -5
  68. data/lib/woods/extractors/manager_extractor.rb +7 -21
  69. data/lib/woods/extractors/migration_declaration.rb +87 -0
  70. data/lib/woods/extractors/migration_extractor.rb +5 -39
  71. data/lib/woods/extractors/phlex_extractor.rb +6 -2
  72. data/lib/woods/extractors/policy_extractor.rb +9 -5
  73. data/lib/woods/extractors/poro_extractor.rb +112 -53
  74. data/lib/woods/extractors/pundit_extractor.rb +11 -6
  75. data/lib/woods/extractors/scheduled_job_extractor.rb +45 -4
  76. data/lib/woods/extractors/serializer_extractor.rb +34 -22
  77. data/lib/woods/extractors/shared_utility_methods.rb +18 -1
  78. data/lib/woods/extractors/source_nesting.rb +142 -106
  79. data/lib/woods/extractors/standalone_module_discovery.rb +123 -0
  80. data/lib/woods/extractors/state_machine_extractor.rb +46 -40
  81. data/lib/woods/extractors/view_component_extractor.rb +9 -7
  82. data/lib/woods/flow_assembler.rb +4 -1
  83. data/lib/woods/generation.rb +25 -0
  84. data/lib/woods/hooks/context_hint.rb +7 -2
  85. data/lib/woods/mcp/bootstrapper.rb +33 -7
  86. data/lib/woods/mcp/config_resolver.rb +26 -7
  87. data/lib/woods/mcp/index_reader.rb +125 -24
  88. data/lib/woods/mcp/index_reader_pinning.rb +16 -0
  89. data/lib/woods/mcp/renderers/markdown_renderer.rb +7 -1
  90. data/lib/woods/mcp/renderers/plain_renderer.rb +3 -1
  91. data/lib/woods/mcp/search_results.rb +7 -1
  92. data/lib/woods/mcp/server.rb +24 -4
  93. data/lib/woods/module_reconciliation.rb +151 -0
  94. data/lib/woods/path_dispatcher.rb +7 -2
  95. data/lib/woods/rake_helpers.rb +43 -11
  96. data/lib/woods/release.rb +1 -1
  97. data/lib/woods/resilience/index_validator.rb +8 -3
  98. data/lib/woods/resilience/retryable_provider.rb +18 -1
  99. data/lib/woods/resolved_config.rb +68 -8
  100. data/lib/woods/retrieval/context_assembler.rb +3 -3
  101. data/lib/woods/retrieval/lexical_assembler.rb +3 -2
  102. data/lib/woods/retrieval/scope.rb +18 -2
  103. data/lib/woods/retrieval/source_evidence.rb +14 -2
  104. data/lib/woods/source_contributor_validation.rb +78 -0
  105. data/lib/woods/source_contributors.rb +116 -0
  106. data/lib/woods/source_inputs/handoff.rb +37 -0
  107. data/lib/woods/source_inputs/launcher.rb +53 -13
  108. data/lib/woods/source_inputs/manifest.rb +84 -3
  109. data/lib/woods/source_inputs/private_key.rb +44 -12
  110. data/lib/woods/source_inputs/scanner.rb +98 -27
  111. data/lib/woods/source_inputs/scopes.rb +1 -1
  112. data/lib/woods/source_inputs/session.rb +147 -15
  113. data/lib/woods/source_inputs/stable_reader.rb +127 -0
  114. data/lib/woods/source_inputs/status.rb +40 -8
  115. data/lib/woods/source_inputs/verifier.rb +28 -5
  116. data/lib/woods/source_path_encoding.rb +33 -0
  117. data/lib/woods/source_references/cache.rb +284 -0
  118. data/lib/woods/source_references/collector.rb +120 -0
  119. data/lib/woods/source_references/extraction.rb +185 -0
  120. data/lib/woods/source_references/inputs.rb +134 -0
  121. data/lib/woods/source_references/parser_adapter.rb +134 -0
  122. data/lib/woods/source_references/pass.rb +152 -0
  123. data/lib/woods/source_references/prism_adapter.rb +116 -0
  124. data/lib/woods/source_references/registry.rb +178 -0
  125. data/lib/woods/source_references/runtime_lookup.rb +127 -0
  126. data/lib/woods/source_references/value_class.rb +82 -0
  127. data/lib/woods/storage/metadata_store.rb +4 -1
  128. data/lib/woods/storage/qdrant.rb +2 -2
  129. data/lib/woods/unblocked/client.rb +12 -7
  130. data/lib/woods/unblocked/document_builder.rb +4 -1
  131. data/lib/woods/unblocked/exporter.rb +127 -37
  132. data/lib/woods/unblocked/sync_manifest.rb +137 -21
  133. data/lib/woods/unblocked/uri_migration.rb +105 -0
  134. data/lib/woods/util/host_guard.rb +3 -2
  135. data/lib/woods/version.rb +1 -1
  136. data/lib/woods/watch/catch_up.rb +138 -0
  137. data/lib/woods/watch/claim_lease.rb +150 -0
  138. data/lib/woods/watch/cli.rb +26 -2
  139. data/lib/woods/watch/daemon.rb +80 -59
  140. data/lib/woods/watch/installation/options.rb +1 -1
  141. data/lib/woods/watch/installation/receipt.rb +6 -1
  142. data/lib/woods/watch/managed_child.rb +1 -1
  143. data/lib/woods/watch/supervisor.rb +1 -1
  144. data/lib/woods/watch/tree_scan.rb +14 -2
  145. data/plugin/.claude-plugin/plugin.json +1 -1
  146. data/plugin/hooks/adapters/normalize.rb +3 -2
  147. data/plugin/hooks/woods-input-rules.sh +4 -0
  148. data/plugin/hooks/woods-refresh.sh +15 -7
  149. data/plugin/hooks/woods-session-start.sh +60 -3
  150. data/plugin/skills/woods-diagnose/SKILL.md +334 -11
  151. data/plugin/skills/woods-investigate/SKILL.md +11 -0
  152. data/plugin/skills/woods-mcp-config/SKILL.md +79 -8
  153. data/plugin/skills/woods-setup/SKILL.md +53 -6
  154. metadata +32 -5
@@ -58,11 +58,11 @@ The Index Server defines **29 schemas**: the packaged executable registers **14*
58
58
  | Tool group | Count | Wiring condition |
59
59
  |------------|-------|------------------|
60
60
  | Always-on | 14 | Always registered, `lookup`, `search`, `dependencies`, `dependents`, `structure`, `graph_analysis`, `domain_clusters`, `pagerank`, `framework`, `recent_changes`, `reload`, `codebase_retrieve`, `trace_flow`, `woods_status` |
61
- | `session_trace` | 1 | `Woods.configuration.session_store` set and session tracer enabled |
61
+ | `session_trace` | 1 | Custom/embedded Index process has a configured `session_store` supporting `read` and `sessions`; the packaged executable does not load Rails initializer configuration |
62
62
  | Operator (5) | 5 | Custom embedded server wires an operator: `pipeline_extract`, `pipeline_embed`, `pipeline_status`, `pipeline_diagnose`, `pipeline_repair` |
63
63
  | Feedback (4) | 4 | Custom embedded server wires a feedback store: `retrieval_rate`, `retrieval_report_gap`, `retrieval_explain`, `retrieval_suggest` |
64
- | Snapshot (4) | 4 | Extraction with `enable_snapshots = true` normally creates `woods.sqlite3`, which packaged servers discover. If extraction used the JSON fallback, set `WOODS_SNAPSHOTS=true` on the standalone server. Custom embedded servers pass `snapshot_store:`. Internal SQLite migrations are automatic. Tools: `list_snapshots`, `snapshot_diff`, `unit_history`, `snapshot_detail` |
65
- | `notion_sync` | 1 | `notion_api_token` + `notion_database_ids` both set |
64
+ | Snapshot (4) | 4 | Extraction with `enable_snapshots = true` normally creates `woods.sqlite3`, which packaged servers discover. `WOODS_SNAPSHOTS=true` enables construction, preferring SQLite; it does not force JSON or import JSON history. Custom builders can pass an explicit JSON `snapshot_store:`; see [store selection](MCP_SERVERS.md#conditional-index-capabilities). Internal SQLite migrations are automatic. Tools: `list_snapshots`, `snapshot_diff`, `unit_history`, `snapshot_detail` |
65
+ | `notion_sync` | 1 | Token and `notion_database_ids` configured inside a custom/embedded Index process; use `bin/rails woods:notion_sync` for ordinary application export |
66
66
 
67
67
  `codebase_retrieve` is always registered (no `retrieve` alias exists). Default semantic mode requires an embedding provider and a completed `woods:embed` run. Explicit `WOODS_RETRIEVAL_MODE=lexical` ranks published extraction units without a provider or embeddings; set it in the MCP process environment and restart the server. See [embedding-free lexical retrieval](RETRIEVAL_GUIDE.md#embedding-free-lexical-retrieval).
68
68
 
@@ -887,6 +887,12 @@ Trigger extraction, then reload the server's in-memory data:
887
887
 
888
888
  **What you'll get:** Lists of added, modified, and deleted units between the two git SHAs. Use `list_snapshots` first to find valid SHA values.
889
889
 
890
+ **Included in Woods 2.1 (#593):** use the exact stored SHAs returned by
891
+ `list_snapshots`; prefixes are not expanded automatically. `snapshot_diff`
892
+ returns `not_found` if either snapshot is absent and `invalid_params` for a
893
+ malformed SHA. Two available snapshots with no differences still return a
894
+ successful empty diff. These checks apply to both JSON and SQLite stores.
895
+
890
896
  ---
891
897
 
892
898
  ### "How has the User model evolved?"
@@ -908,58 +914,10 @@ Trigger extraction, then reload the server's in-memory data:
908
914
 
909
915
  ### GitHub Actions for Incremental Extraction
910
916
 
911
- Run incremental extraction on every push, cache the index between runs:
912
-
913
- ```yaml
914
- # .github/workflows/woods.yml
915
- name: Update Codebase Index
916
-
917
- on:
918
- push:
919
- branches: [main]
920
- pull_request:
921
-
922
- jobs:
923
- index:
924
- runs-on: ubuntu-latest
925
- steps:
926
- - uses: actions/checkout@v4
927
- with:
928
- fetch-depth: 2 # needed for incremental diff
929
-
930
- - name: Set up Ruby
931
- uses: ruby/setup-ruby@v1
932
- with:
933
- bundler-cache: true
934
-
935
- - name: Restore index cache
936
- uses: actions/cache@v4
937
- with:
938
- path: tmp/woods
939
- key: woods-${{ github.ref }}-${{ github.sha }}
940
- restore-keys: |
941
- woods-${{ github.ref }}-
942
- woods-
943
-
944
- - name: Run database migrations
945
- run: bundle exec rails db:migrate RAILS_ENV=test
946
-
947
- - name: Update codebase index
948
- run: bundle exec rake woods:incremental
949
- env:
950
- RAILS_ENV: test
951
- GITHUB_BASE_REF: ${{ github.base_ref }}
952
-
953
- - name: Validate index
954
- run: bundle exec rake woods:validate
955
- ```
956
-
957
- For Docker-based CI:
958
-
959
- ```yaml
960
- - name: Update codebase index
961
- run: docker compose exec -T app bundle exec rake woods:incremental
962
- ```
917
+ Use the [canonical incremental CI recipe](INCREMENTAL_EXTRACTION.md#github-actions-with-an-exact-baseline).
918
+ It fetches the actual base ref, restores an index for the exact baseline commit,
919
+ and runs full extraction when that cache is absent. The same guide covers
920
+ nested Rails applications and forwarding the selected range into Docker.
963
921
 
964
922
  ---
965
923
 
@@ -171,7 +171,13 @@ steps:
171
171
 
172
172
  ### MCP Server
173
173
 
174
- If using the MCP Index Server, the `notion_sync` tool is available:
174
+ The packaged Index Server does not load Rails initializers and does not register
175
+ `notion_sync` merely because the application has Notion configured. Use
176
+ `bin/rails woods:notion_sync` in the Rails application for the standard workflow.
177
+
178
+ A custom/embedded Index Server can register the tool when the token and database
179
+ IDs are configured in that server process. Confirm it appears in `tools/list`
180
+ before calling it:
175
181
 
176
182
  ```json
177
183
  {
@@ -249,6 +249,12 @@ Retention is `WOODS_PAYLOAD_RETENTION` generations (default 3). Raise it on a CI
249
249
 
250
250
  `woods:check:moved_messages` opens two retained generations through `Woods::PublishedIndex` (block form, so both readers' retention locks always release) and reports every public method name that looks like it moved from one unit to another while a `:test_coverage` edge did not follow.
251
251
 
252
+ The task does not boot Rails. Pass `WOODS_OUTPUT` for a custom index; an explicit
253
+ relative value is relative to the command's working directory. Without it, the
254
+ Woods 2.1 #591 fix selects `tmp/woods` beside the loaded Rakefile, including
255
+ `rake -f /app/Rakefile` from another directory. Earlier builds used the invoking
256
+ directory instead; pass an absolute output path on those builds.
257
+
252
258
  Each row is a **candidate move into a unit without mapped tests**, never a proven coverage loss: the check matches on method name and kind alone, so two unrelated methods that happen to share both look identical to a real move. Reported only when the source unit was covered before and the destination is not covered after; a method that was never covered is not a regression this check owns. `WOODS_CHECK_STRICT=1` exits 1 on any finding, but the check stays heuristic either way, strict mode changes the exit code, not the confidence of a row.
253
259
 
254
260
  ```bash
data/docs/README.md CHANGED
@@ -69,7 +69,12 @@ and stdio or Streamable HTTP endpoints directly.
69
69
 
70
70
  ## Maintainer material
71
71
 
72
- Historical build-phase documents are not user guides. Source checkouts also contain the [MCP protocol decision record](https://github.com/lost-in-the/woods/blob/main/docs/design/MCP_2026_STRATEGY.md), [generated self-analysis diagrams](https://github.com/lost-in-the/woods/tree/main/docs/self-analysis), and the maintainer work ledger in `backlog.json`; these maintainer-only paths are not packaged with the gem.
72
+ Historical build-phase documents are not user guides. Source checkouts also contain the [MCP protocol decision record](https://github.com/lost-in-the/woods/blob/main/docs/design/MCP_2026_STRATEGY.md) and [generated self-analysis diagrams](https://github.com/lost-in-the/woods/tree/main/docs/self-analysis); these maintainer-only paths are not packaged with the gem.
73
+
74
+ [GitHub issues](https://github.com/lost-in-the/woods/issues) track active work
75
+ and pending decisions. The old local backlog is retired; historical records,
76
+ approved decisions and issue migration links are archived under `.Codex/`.
77
+ See [work tracking](../CONTRIBUTING.md#work-tracking) for ownership and release labels.
73
78
 
74
79
  Contract records that read as maintainer reference rather than guides:
75
80
 
@@ -167,6 +167,23 @@ unit still prepares no text before accepting a matching no-vector checkpoint.
167
167
  Empty-vector reconciliation waits until all batches succeed, so a later provider
168
168
  failure does not retire those old vectors.
169
169
 
170
+ ### Prepared-input checkpoints
171
+
172
+ **Included in Woods 2.1:** checkpoint schema 2 records preparation policy and
173
+ ordered complete input fingerprints, including chunk/storage IDs. A change to a
174
+ prepared prefix or chunks re-embeds even if the source hash is unchanged; metadata
175
+ that does not affect prepared text does not force an embedding. Old checkpoints
176
+ invalidate once, so budget for one re-embed after upgrading. Complete input
177
+ bounds and lossless splitting are described in [Embedding Models](EMBEDDING_MODELS.md).
178
+
179
+ Successful units advance their checkpoint only after all their inputs are
180
+ embedded, stored and obsolete chunks reconciled. A provider failure can leave
181
+ earlier successful batches committed; it does not mark the failed unit complete.
182
+ Built-in durable stores refuse changed replacements when existing IDs cannot be
183
+ enumerated. Custom stores without enumeration retain adapter-owned cleanup and
184
+ cannot claim exact obsolete-chunk reconciliation. Keep backups for rollback;
185
+ older readers do not understand the new checkpoint policy.
186
+
170
187
  ---
171
188
 
172
189
  ## Embedding-free lexical retrieval
@@ -14,6 +14,7 @@ before using it; upgrading the plugin alone does not upgrade Woods.
14
14
  |---|---|---|
15
15
  | `current` | All covered inputs match their consumer baselines, with a verified fresh boot boundary and complete checks. A dirty checkout can be current. | Use the indexed facts within the coverage below. |
16
16
  | `drifted` | At least one captured input differs, was removed, or a relevant input was added. | Inspect the changed paths; choose a full run or a justified targeted refresh. |
17
+ | `unavailable` | **Included in Woods 2.1:** source evidence exceeded its serialized-size limit; the code index was published successfully. | Inspect `unavailable.size_bytes` and `unavailable.limit_bytes`; another full extraction will not reduce this size. |
17
18
  | `unknown` | Evidence is incomplete: for example an old index, missing source/key, scan limit, opaque symlink directory, or unproved boot/consumer boundary. | Inspect `reasons`; use a deep check or fresh full capture as appropriate. |
18
19
 
19
20
  Drift can coexist with incomplete coverage. `reasons` reports both; absence of a
@@ -22,6 +23,13 @@ capture/traversal completion, while boot and consumer qualifications remain in
22
23
  `reasons`. `counts` contains total observed added/changed/removed paths; each
23
24
  `changes` list contains at most 30 paths and `truncated` marks longer lists.
24
25
 
26
+ **Included in Woods 2.1:** `comparison_complete` also identifies whether the
27
+ previous consumer baseline can establish additions. A missing or incompatible
28
+ baseline stays unknown; it does not make every visible file an addition. Files
29
+ not reached by an incomplete scan are not reported as deleted. Matching paths
30
+ with different captured identities still prove drift. An unreadable entry leaves
31
+ coverage incomplete while scanning continues through accessible siblings.
32
+
25
33
  ```json
26
34
  {"source_check":"deep"}
27
35
  ```
@@ -37,6 +45,19 @@ The result is tied to the **served** generation, including a reader holding an
37
45
  older generation during a concurrent publication. It is recomputed on each call;
38
46
  a source edit does not require an index generation change to become visible.
39
47
 
48
+ **Included in Woods 2.1:** `recommendations` separates recovery actions:
49
+
50
+ - `deep_check`: the quick reader reached its time limit; request one deep check.
51
+ - `inspect_source_scan`: inspect `verification_reasons`, permissions, source
52
+ mapping and scan limits; a deeper scan cannot repair unavailable inputs.
53
+ - `inspect_source_limits`: source freshness is unavailable because evidence exceeds its size limit; the code index remains usable.
54
+ - `fresh_capture`: the captured baseline, boot or consumer evidence is incomplete;
55
+ fix the reported cause and run the launcher in a fresh process.
56
+
57
+ Several recommendations can apply together. `verification_reasons` describes the
58
+ current reader scan; `reasons` also retains capture failures. A quick reader limit
59
+ alone does not require rebuilding the index.
60
+
40
61
  ## Establish a fresh baseline
41
62
 
42
63
  Run the launcher through the application's installed bundle:
@@ -56,12 +77,37 @@ commas and newlines. Refresh accepts known extractor names. Invalid arguments
56
77
  fail before extraction; the launcher propagates the child's failure or daemon
57
78
  stand-down exit 75. Split oversized incremental batches or choose full.
58
79
 
80
+ For an application with a custom `config.output_dir`, **always pass the matching
81
+ `--output PATH` or `WOODS_OUTPUT`**. The launcher cannot evaluate Rails configuration
82
+ before capturing its boot inputs. In Woods 2.0.0, its implicit `tmp/woods` default
83
+ overrides custom configuration and can create a second index.
84
+
85
+ **Included in Woods 2.1 ([#591](https://github.com/lost-in-the/woods/issues/591)):**
86
+ an implicit default no longer injects `WOODS_OUTPUT` into Rails boot. The task
87
+ compares finalized configuration with that preboot default and refuses a mismatch
88
+ before extraction, naming the configured path and explicit output remedy. An
89
+ initializer using `ENV.fetch('WOODS_OUTPUT', custom_path)` therefore retains its
90
+ configured default for this check. Explicit CLI/environment output still overrides
91
+ configuration and binds capture, key and publication to the same directory.
92
+ The refused launch may already have created a private key under `tmp/woods`, but
93
+ publishes no generation there and preserves the configured index's generation.
94
+
59
95
  The launcher captures source before a **fresh child** evaluates its Gemfile,
60
96
  Rakefile, Rails boot and eager loading. A private one-use handoff binds the capture
61
97
  to root, output, action, rules, nonce and parent process. It waits for the child
62
98
  and removes the handoff when the child exits. Source is checked again before
63
- publication; edits during boot/extraction retain the earlier identity and are
64
- reported, rather than being silently adopted as a current baseline.
99
+ publication. In Woods 2.0.0, edits during boot/extraction retain the earlier
100
+ identity and are reported in the new generation, rather than being silently
101
+ adopted as a current baseline.
102
+
103
+ **Included in Woods 2.1:** the
104
+ [constant-reference writer](EXTRACTOR_REFERENCE.md#constant-source-references)
105
+ requires verified source for its graph evidence. Source changes after capture,
106
+ during boot or extraction, refuse publication while reference enrichment
107
+ participates. The final verification also checks covered non-Ruby inputs. The
108
+ preceding generation remains active instead of publishing a new drifted
109
+ generation. Retry against stable source in a fresh process. The existing
110
+ generation can still report drift through `woods_status`.
65
111
 
66
112
  Existing Rake tasks, direct `Extractor` calls and the watch daemon remain usable.
67
113
  Their post-boot capture is marked `unverified_boot_boundary`; current bytes alone
@@ -80,22 +126,71 @@ scanner's exclusions. Explicit source roots override generic exclusions; the
80
126
  index output is always excluded. Contained file symlinks are checked for stable
81
127
  resolution; directory symlinks and escaping/unreadable inputs leave uncertainty.
82
128
 
129
+ **Included in Woods 2.1 (#653):** supporting writers scan
130
+ directory links under their logical application paths, preserving independent
131
+ aliases. Static files with no input consumer do not prevent publication merely
132
+ because a parent directory is a symlink. Entries in an external linked directory
133
+ may be enumerated to determine scope, but an input file must resolve inside the
134
+ application root before Woods opens its contents. Package files are inputs even
135
+ outside `app/` and `lib/`; an external package file still refuses verification.
136
+ Directory cycles, links that change during scanning, unreadable entries and scan
137
+ limits keep evidence incomplete. Output-directory aliases are pruned. Run a fresh
138
+ full extraction after upgrading to establish the updated capture rules; do not
139
+ replace valid links or disable source verification to work around an older writer.
140
+
141
+ **Included in Woods 2.1:** source paths use UTF-8 bytes even
142
+ when the process locale is `C`. A filename with invalid UTF-8 bytes produces
143
+ `undecodable_source_path` and incomplete coverage; directory entries with those
144
+ names are pruned. Diagnostic labels escape the original bytes and are bounded.
145
+ The watcher skips these entries, and reference verification retains the preceding
146
+ generation. **Included in Woods 2.1:** the publication refusal names the reason
147
+ and up to three escaped, bounded path labels, including entries outside source
148
+ consumer scopes; it never includes file contents. Rename the affected entries to
149
+ valid UTF-8 names before a fresh capture. Explicit root/output paths with invalid UTF-8 bytes are rejected as
150
+ configuration errors.
151
+
83
152
  Every consuming scope keeps its own identities. An events scan can reread a
84
153
  service file while its service unit remains untouched; refreshing events does
85
154
  not certify the retained service unit. Successful file/whole-extractor work
86
155
  updates only its scopes, including negative results and confirmed deletion.
87
- Unchanged scopes keep their earlier baseline. Boot inputs advance only on a full
88
- run. Partial runtime changes retain an explicit `runtime_consumption` uncertainty
156
+ Unchanged scopes keep their earlier baseline. A file can therefore remain
157
+ `drifted` after an incremental task or hook re-extracts it: its file-extractor
158
+ scope may be current while retained `runtime` evidence still refers to the old
159
+ source. Known differences take precedence over the accompanying uncertainty;
160
+ `unverified_boot_boundary` does not hide them. Inspect the reported scopes and
161
+ use a fresh verified full launcher run when all reflected facts need a new
162
+ baseline. An ordinary unverified full run cannot establish verified freshness.
163
+ Boot inputs advance only on a full run. Partial runtime changes retain an explicit `runtime_consumption` uncertainty
89
164
  when Woods cannot prove every retained reflected fact was re-serialized. Named
90
165
  framework refreshes do not certify unrelated application inputs. A handled
91
166
  extractor error retains an explicit `extractor:<name>` uncertainty even if the
92
167
  extractor returns an empty result. Successful consumers keep their own evidence;
93
168
  a later full run without that failure can replace the uncertainty.
169
+ In the Woods 2.1 reference writer, successfully returned models
170
+ also retain their individual source evidence when another model fails. An
171
+ optional framework model with a missing table therefore does not prevent
172
+ unrelated incremental work; the model-extractor uncertainty remains visible.
173
+ If an older writer already published a baseline without that individual evidence,
174
+ run one full extraction after upgrading to rebuild it.
175
+
176
+ With the **Woods 2.1 reference writer**, an edited Ruby service
177
+ retained by an events-only refresh instead prevents publication: its current
178
+ source cannot certify its older runtime facts or cached reference resolution.
179
+ See [reference baseline recovery](INCREMENTAL_EXTRACTION.md#source-reference-baseline-and-upgrades).
180
+ Omitted non-reference inputs retain their earlier consumer identity. The watcher
181
+ keeps failed work for retry; it does not expand an incomplete batch or rebuild a
182
+ missing reference baseline automatically.
94
183
 
95
184
  Custom loader source outside captured roots, missing eager-load coverage and
96
185
  uncaptured application-owned unit paths remain unknown. This is application
97
186
  source evidence: it does not certify external database schemas/data, remote
98
187
  configuration, installed gem bytes, provider state or live runtime services.
188
+ **Included in Woods 2.1:** RubyGems installation metadata can establish that
189
+ loaded/unit source inside the application directory belongs to an installed gem,
190
+ including gems installed under `vendor/bundle`. Such files do not create false
191
+ application-coverage errors. A vendor-shaped directory alone is insufficient;
192
+ local path gems and custom loaders still need captured source roots. Explicit
193
+ `--source-root` declarations retain coverage even for installed gem directories.
99
194
 
100
195
  ## Containers and hooks
101
196
 
@@ -107,6 +202,16 @@ same verifier without Rails initialization or provider work. Its optional
107
202
  Base64-encoded JSON transport supports `output`, `root` (an explicit reader-side
108
203
  source mapping), and `mode` (`quick` or `deep`).
109
204
 
205
+ **Included in Woods 2.1:** results expose `recorded_root` (writer location),
206
+ `checked_root` (the directory actually scanned), and `root_source` (`recorded`,
207
+ `explicit`, or `working_directory`). `current` applies only to that checked root.
208
+ The MCP IndexReader keeps using the recorded root; it does not infer a checkout
209
+ from the index's location. For a copied index, a current result about the original
210
+ root says nothing about edits in the copy. Run `woods:source_status` in the copy,
211
+ or provide an explicit `root` mapping. The task defaults to its process working
212
+ directory, including inside containers; launch it from the application root, not
213
+ an unrelated directory or a monorepo parent. Missing source remains unknown.
214
+
110
215
  The opt-in SessionStart hook uses `WOODS_HOOK_RAKE` and `WOODS_OUTPUT`, including a
111
216
  Docker command prefix without requiring a host application bundle. It prints
112
217
  an actionable drift or unknown warning and stays quiet for current evidence.
@@ -114,6 +219,10 @@ Its ten-second process deadline includes command startup; the scan uses quick
114
219
  mode. Missing older tasks and failed/timed-out commands report unknown. Cancelling
115
220
  Docker exec does not itself prove the process inside the container stopped.
116
221
  A quiet session hook does not acknowledge deferred PostToolUse queue entries.
222
+ **Included in Woods 2.1:** unknown warnings use the returned recommendations,
223
+ including every applicable recovery action. A quick reader timeout advises a deep
224
+ check without requiring a rebuild. Older tasks without recommendations retain
225
+ the conservative generic unknown warning.
117
226
 
118
227
  ## Artifact and cost
119
228
 
@@ -128,11 +237,53 @@ unknown; Woods does not silently repair permissions or rotate keys. Independent
128
237
  outputs have different identities and must be compared using their own keys and
129
238
  consumer semantics, not raw manifest equality.
130
239
 
240
+ ### Identity-key recovery
241
+
242
+ Verified launching refuses an unavailable, insecure or malformed
243
+ `<output>/.source-inputs.key`; it cannot establish verified capture without that
244
+ key. In Woods 2.1 diagnostics, the error names the path and safe
245
+ file requirements while reader reason codes remain unchanged. No key bytes are
246
+ printed and Woods does not chmod, chown, replace or rotate the file automatically.
247
+
248
+ Inspect the file and mount from the application environment. It must be a regular
249
+ file, not a symlink/FIFO, contain exactly 32 bytes, belong to the process's UID,
250
+ and grant no group/other permissions (normally mode `0600`). A host/container UID
251
+ mismatch requires correcting the selected runtime user or deliberately repairing
252
+ ownership of the known original key. Confirm the file belongs to this index before
253
+ changing permissions. Do not expose the key in logs or copy it into payloads.
254
+
255
+ If the original key is lost, replaced or cannot be trusted, retain the previous
256
+ index and establish a fresh full baseline in a new empty output directory with
257
+ the intended application user. Configure the writer and readers together.
258
+ Do not replace a key and assume the old generation now has valid source evidence.
259
+
131
260
  Failed/no-op extraction does not advance the artifact's published generation.
132
261
  Flat fallback and older indexes lack verified atomic source evidence.
133
262
  `WOODS_PROFILE=1` reports `source capture` and `source verification` separately.
134
263
  Capture/recheck each allow up to ten seconds with the same file/byte caps.
135
264
 
265
+ **Included in Woods 2.1:** the published manifest and private launcher handoff
266
+ have a 16 MiB serialized-size limit. If size alone exceeds that limit, extraction
267
+ continues and publishes the code index successfully with a loud warning and
268
+ bounded evidence: `state: "unavailable"`, reason `source_manifest_too_large`, and
269
+ `unavailable.size_bytes` / `unavailable.limit_bytes`. A launcher handoff above the
270
+ limit continues in a fresh child without claiming verified preboot capture.
271
+ Status, validation and hooks retain this limitation; they do not recommend an
272
+ identical full rebuild. Source freshness remains unavailable on subsequent
273
+ incremental runs until a full capture fits the limit.
274
+
275
+ Normal changed-file and Git-based incremental work continues. Unchanged
276
+ source-reference facts require both unchanged keyed source identities and exact
277
+ typed ownership in the validated, hash-bound reference cache from the same
278
+ published generation. Missing, corrupt or mismatched proof still refuses reuse.
279
+ The watcher retains a compact captured-tree fingerprint for catch-up only; that
280
+ fingerprint never establishes runtime freshness. Invalid evidence, failed writes
281
+ and source instability retain their existing publication safeguards.
282
+
283
+ Version-1 manifests remain readable. The optional `comparison_complete` field
284
+ supplements existing coverage errors; an unavailable manifest never claims
285
+ complete comparison coverage.
286
+
136
287
  September 2026 fixture measurements: a pinned Writebook source tree (456 visited
137
288
  files, 223 hashed, 233KB) completed quick scans in median 36ms native / 45ms on a
138
289
  Linux container bind mount. Discourse (26,133 visited, 6,209 hashed, 56.4MB) needed
@@ -140,4 +291,5 @@ about 1.6–2 seconds; quick scans returned unknown, while five-second scans com
140
291
  traversal and still reported an opaque directory symlink. Native Ruby 4.0 and
141
292
  container Ruby 3.4 differed, so these are a budget envelope, not a filesystem
142
293
  speed comparison or a macOS virtiofs benchmark. Measure your own application;
143
- a large or slow source tree may need a full extraction rather than a longer read.
294
+ a large or slow source tree may need a deep check. Rebuilding a verified baseline
295
+ does not remove the reader's time, file or byte limits.
@@ -1,20 +1,12 @@
1
1
  # Token Estimation Benchmark
2
2
 
3
- > **Single source of truth:** `Woods::TokenUtils.chars_per_token_for(provider)`
4
- > in `lib/woods/token_utils.rb`. Production code uses **4.0 chars/token** for
5
- > the OpenAI path and **1.5 chars/token** for the Ollama / WordPiece path.
6
- > Both are applied consistently by `Woods::Builder#chars_per_token_for`,
7
- > `ContextAssembler`, `TextPreparer`, and `ExtractedUnit#estimated_tokens`.
8
- > The cost-model layer (`lib/woods/cost_model/`) is the one exception, it
9
- > uses its own pre-aggregated `TOKENS_PER_CHUNK` constant (450) rather than
10
- > a per-string ratio, since it measures a different thing (per-chunk average
11
- > vs. per-string chars/token).
12
- >
13
- > When the optional [`tokenizers`](https://github.com/ankane/tokenizers-ruby)
14
- > gem is installed, the Ollama path uses the real BERT WordPiece tokenizer
15
- > (`Woods::Embedding::TokenCounter`) instead of this heuristic. The 4.0
16
- > divisor applies to the OpenAI/default path; Ollama falls back to 1.5 when
17
- > its tokenizer is unavailable.
3
+ > **Historical sizing evidence, not an input-limit guarantee.**
4
+ > `Woods::TokenUtils.chars_per_token_for(provider)` supplies character estimates
5
+ > for retrieval assembly and initial sizing. Woods 2.1 uses
6
+ > a conservative UTF-8 byte bound for known OpenAI embedding models and honest
7
+ > estimates for Ollama/custom models. Prefixes count toward admission; oversized
8
+ > inputs are split without dropping source. The optional tokenizer gem no longer
9
+ > downloads or implicitly selects BERT. See [Embedding Models](EMBEDDING_MODELS.md).
18
10
 
19
11
  This is a historical record of the benchmark that picked 4.0 over the
20
12
  original 3.5 divisor. It is cited from five places in `lib/` as the evidence
@@ -58,9 +50,9 @@ current definition and `docs/EMBEDDING_MODELS.md` for the Ollama-side ratio.
58
50
  **tiktoken_ruby was deliberately not added as a runtime dependency.** A 10.6%
59
51
  mean error is acceptable for chunking decisions, budget estimates, and
60
52
  truncation; a native-extension dependency for marginal accuracy gains wasn't
61
- worth it. The optional `tokenizers` gem provides counts for its supported BERT
62
- WordPiece tokenizer, not every model. Strict token-limit enforcement requires
63
- the tokenizer used by the target model.
53
+ worth it. An explicitly injected local tokenizer must match the target model. A conservative
54
+ upper bound can also establish admission for a known encoding; a character
55
+ average cannot.
64
56
 
65
57
  ## Reproducing this benchmark
66
58
 
@@ -32,6 +32,17 @@ This guide covers the most common problems encountered when installing, extracti
32
32
  | Tool returns `error_code: :not_configured` | Feature flag or credential not set | Check `config_key` in `_meta` and the linked `doc_link` |
33
33
  | Tool returns `error_code: :rate_limited` | `PipelineGuard` 5-min cooldown hit | Wait `retry_after_seconds` from `_meta`, then retry |
34
34
 
35
+ ### Source-reference baseline needs a full extraction
36
+
37
+ **Included in Woods 2.1.** The related initial diagnostic is
38
+ `Source-reference baseline is missing or incompatible`. Both mean the writer
39
+ cannot safely reuse its reference cache or verify the source consumed by retained
40
+ units. It can follow an older-index upgrade, missing cache artifacts, or source
41
+ changes omitted from the refresh batch. Confirm the loaded gem revision and
42
+ output directory, then follow the [full baseline rebuild](INCREMENTAL_EXTRACTION.md#source-reference-baseline-and-upgrades).
43
+ The failed operation leaves the published generation unchanged. Preserve pending
44
+ watcher work; increasing traversal budgets cannot repair extraction coverage.
45
+
35
46
  ### First-Pass Diagnostics
36
47
 
37
48
  For a single-call health snapshot, call the Index Server's `woods_status` tool. It reports:
@@ -67,6 +78,12 @@ the selected owner or installed command and restarting that owner; do not delete
67
78
  claim files or kill PIDs taken from status. No index-visible record exists before
68
79
  the first boot resolves the application's output directory.
69
80
 
81
+ For an abandoned foreign-container claim, follow
82
+ [ownership-verified claim recovery](WATCH_DAEMON.md#recovering-an-abandoned-managed-claim).
83
+ The Woods 2.1 command requires the exact token and a free lifetime
84
+ lease; old claims require the documented legacy recovery. Age is not proof of
85
+ abandonment. `woods:clean` preserves ownership sidecars and does not reset them.
86
+
70
87
  Unset `WOODS_WATCH_IDLE_TIMEOUT` in managed modes. If the boot deadline is reached,
71
88
  diagnose Bundler/initializer startup before increasing `--boot-timeout`; a valid
72
89
  long extraction has a separate readiness state and is not bounded by that clock.
@@ -74,6 +91,10 @@ If setup created a Procfile but normal `bin/dev` still only launches Rails, choo
74
91
  Puma or explicitly run the selected Foreman command. The generator never rewrites
75
92
  `bin/dev` or starts services during preview.
76
93
 
94
+ An empty idle-timeout variable crashes older tasks even though managed validation
95
+ accepts it. Unset it on those builds; the Woods 2.1 #591 fix consistently treats
96
+ empty or whitespace-only values as unset.
97
+
77
98
  If installation reports a pending transaction, use `woods:watch --operation
78
99
  recover` through the Rails generator, initially with `--pretend`; see
79
100
  [owned setup recovery](WATCH_DAEMON.md#ownership-updates-and-removal). That Rails
@@ -415,7 +436,7 @@ full extraction when current history and provenance are required.
415
436
 
416
437
  **Cause:** The range git was asked to diff does not resolve in the checkout: a GitLab `CI_COMMIT_BEFORE_SHA` of all zeros (new branch), a GitHub base ref that was never fetched, a shallow clone with no `HEAD~1`, or a typo'd revision. This used to read as "no relevant files changed" and the task exited 0 while the sync never ran; it now fails closed, because a green job hiding a skipped sync lets index drift grow unbounded. The one stand-down: a *running* watch daemon maintaining the same index exits 0 with a printed reason, since its start-up catch-up covers the changes. A degraded daemon covers nothing and still exits 1.
417
438
 
418
- **Fix:** Repair or provide the range (fetch the base ref, e.g. `fetch-depth: 2` or more, or correct the CI environment variables), set `CHANGED_FILES` explicitly to bypass range resolution, or run a full extraction:
439
+ **Fix:** Fetch the actual base ref and enough history to resolve both endpoints (for example `fetch-depth: 0` with an explicit base-ref fetch in GitHub Actions), or correct the CI environment variables. A depth of two does not establish an arbitrary CI base. Set `CHANGED_FILES` explicitly to bypass range resolution, or run a full extraction:
419
440
 
420
441
  ```bash
421
442
  bundle exec rake woods:extract
@@ -461,6 +482,31 @@ the unreachable payload. A resident `woods:watch` process handles the same
461
482
  failure differently: it reports `degraded`, carries the changed paths, and
462
483
  retries after a later filesystem event.
463
484
 
485
+ For the optional embedded `pipeline_extract` tool, a client using the Tasks
486
+ extension sees the task become `failed` for this publication refusal on a
487
+ revision containing #584 (included in Woods 2.1). A background-start response
488
+ alone does not mean extraction completed. The packaged Index Server does not
489
+ register this tool.
490
+
491
+ ### Incremental extraction or refresh reports "Extraction failed for ..."
492
+
493
+ A selected whole-app extractor raised before returning a complete result, or
494
+ its initialization failed. In Woods 2.1, a successful sibling extractor cannot
495
+ turn that failed batch into a successful publication (#584). Readers retain the prior generation. Inspect the
496
+ earlier log line naming the failed extractor, fix its cause, then retry the
497
+ complete changed-file list or refresh selection. Preserve the published index;
498
+ deleting it does not repair the failing extractor.
499
+
500
+ ### Incremental extraction reports "Restart-sensitive inputs"
501
+
502
+ On revisions containing #588 (included in Woods 2.1), the direct incremental
503
+ API and optional embedded pipeline refuse schema or boot-configuration inputs.
504
+ Apply required migrations, then run `bin/rails woods:extract` in a fresh Rails
505
+ process. Reusing an embedded server's old Rails configuration or schema cache
506
+ cannot establish current runtime facts. The fresh one-shot `woods:incremental`
507
+ task selects a full run automatically for these inputs. See the
508
+ [runtime input contract](INCREMENTAL_EXTRACTION.md#what-a-change-actually-requires-reload-restart-or-neither).
509
+
464
510
  ---
465
511
 
466
512
  ### `manifest.json` shows the wrong branch (or `git_branch: "unknown"`) in a worktree
@@ -737,6 +783,15 @@ Woods detects the dimension mismatch and raises `Woods::MCP::DimensionMismatch`
737
783
 
738
784
  **A dimension mismatch is never silently tolerated.** If you are getting poor results without seeing this error, the cause is something else.
739
785
 
786
+ For an unsupported `dimensions` request or a wrong-width cached vector, compare
787
+ the installed embedding and reader revisions as well as the model, endpoint,
788
+ and explicit width configuration. The request/cache consistency fix (#586) is
789
+ included in Woods 2.1: it separates stored widths from requested reductions,
790
+ keeps fixed-width ada requests compatible, and separates embedding cache entries
791
+ by provider configuration. See [embedding options](CONFIGURATION_REFERENCE.md#embedding-options)
792
+ and [cache identity](CONFIGURATION_REFERENCE.md#retrieval-cache-options). Do not
793
+ remove a width guard to accept mismatched vectors.
794
+
740
795
  ---
741
796
 
742
797
  ### OpenAI API errors during embedding
@@ -786,20 +841,16 @@ config.embedding_options = { host: 'http://localhost:11434' }
786
841
 
787
842
  **Symptom:** `rake woods:embed` fails with `Ollama API error: 400 {"error":"the input length exceeds the context length"}`. Individual chunks may look smaller than the configured `num_ctx`.
788
843
 
789
- **Cause:** Ollama's `/api/embed` endpoint enforces the model's **native** `context_length`, not the `options.num_ctx` override (see [ollama/ollama#14186](https://github.com/ollama/ollama/issues/14186)). For `nomic-embed-text` that's 2048 tokens, regardless of what `num_ctx` is set to. Separately, without the `tokenizers` gem, Woods estimates token counts from character length, which under-counts dense Ruby source, so chunks that look safe by char count still trip the 2048-token ceiling.
790
-
791
- **Fix:** Use Woods 2.0 and install the `tokenizers` gem:
792
-
793
- ```ruby
794
- # Gemfile
795
- gem 'woods', '~> 2.0'
796
- gem 'tokenizers', '~> 0.5' # exact BERT WordPiece token counting
797
- ```
844
+ **Cause:** the server enforces the selected model's actual context window.
845
+ Character estimates can undercount dense Ruby or multilingual source; a larger
846
+ `num_ctx` setting does not establish that the model accepts a larger input.
798
847
 
799
- Woods now:
800
-
801
- 1. Advertises the native context ceiling per model (2048 for `nomic-embed-text`, 8192 for `bge-m3`/`snowflake-arctic-embed2`, etc.) so the chunker sizes inputs correctly.
802
- 2. Uses the real BERT tokenizer to verify every chunk, catching the 10–20% gap between char-based estimates and Ollama's internal count.
848
+ **Diagnosis:** record the loaded Woods revision, model and configured context.
849
+ In supporting builds after 2.0.0, Woods counts the full metadata prefix plus
850
+ source, splits without truncating source, and requests Ollama `truncate: false`.
851
+ An overflow is a visible refusal; read the unit/model/limit diagnostic rather
852
+ than enabling truncation or installing an unrelated BERT tokenizer. See the
853
+ [input counting contract](EMBEDDING_MODELS.md#why-num_ctx-isnt-enough).
803
854
 
804
855
  If you want fewer chunks per unit and have the disk space, switch to a larger-context model:
805
856
 
@@ -1046,3 +1097,8 @@ source root/private key, a quick scan limit and an unverified boot boundary are
1046
1097
  different causes. Try `source_check: "deep"` for a budget limit; use the fresh
1047
1098
  launcher for a new verified baseline. Do not delete pending hook events or alter
1048
1099
  key permissions just to suppress a warning. See [source freshness](SOURCE_FRESHNESS.md).
1100
+
1101
+ If `woods-extract` refuses an identity key, use
1102
+ [identity-key recovery](SOURCE_FRESHNESS.md#identity-key-recovery). If it reports a
1103
+ configured-output mismatch, rerun with an explicit matching `--output` or
1104
+ `WOODS_OUTPUT`; do not bypass the check or move a capture/key to another index.