woods 2.0.1 → 2.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (154) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +94 -7
  3. data/CONTRIBUTING.md +134 -19
  4. data/README.md +1 -1
  5. data/docs/AGENT_GUIDE.md +19 -0
  6. data/docs/AGENT_SETUP.md +22 -2
  7. data/docs/BACKEND_MATRIX.md +7 -0
  8. data/docs/CLIENT_HOOKS.md +6 -0
  9. data/docs/CONFIGURATION_REFERENCE.md +133 -25
  10. data/docs/CONSOLE_MCP_SETUP.md +82 -30
  11. data/docs/EMBEDDING_MODELS.md +16 -19
  12. data/docs/EXTRACTOR_REFERENCE.md +219 -21
  13. data/docs/FAQ.md +11 -25
  14. data/docs/GETTING_STARTED.md +7 -1
  15. data/docs/INCREMENTAL_EXTRACTION.md +261 -19
  16. data/docs/INDEX_LAYOUT.md +5 -0
  17. data/docs/INTERNALS.md +9 -0
  18. data/docs/MCP_HTTP_TRANSPORT.md +20 -15
  19. data/docs/MCP_SERVERS.md +87 -8
  20. data/docs/MCP_TOOL_COOKBOOK.md +13 -55
  21. data/docs/NOTION_INTEGRATION.md +7 -1
  22. data/docs/PUBLISHED_INDEX.md +6 -0
  23. data/docs/README.md +6 -1
  24. data/docs/RETRIEVAL_GUIDE.md +17 -0
  25. data/docs/SOURCE_FRESHNESS.md +157 -5
  26. data/docs/TOKEN_BENCHMARK.md +10 -18
  27. data/docs/TROUBLESHOOTING.md +70 -14
  28. data/docs/UNBLOCKED_INTEGRATION.md +60 -8
  29. data/docs/UPGRADING_TO_2.md +153 -38
  30. data/docs/WATCH_DAEMON.md +97 -14
  31. data/exe/woods-console-mcp +2 -2
  32. data/lib/generators/woods/templates/woods.rb.tt +2 -1
  33. data/lib/tasks/woods.rake +23 -7
  34. data/lib/tasks/woods_checks.rake +2 -2
  35. data/lib/woods/agent_configuration/cli.rb +1 -1
  36. data/lib/woods/agent_configuration/layout.rb +16 -2
  37. data/lib/woods/agent_configuration/plan.rb +13 -3
  38. data/lib/woods/agent_configuration/planner_validation.rb +4 -2
  39. data/lib/woods/agent_configuration/preflight.rb +5 -3
  40. data/lib/woods/builder.rb +17 -57
  41. data/lib/woods/cache/cache_middleware.rb +56 -30
  42. data/lib/woods/chunking/contributor_chunks.rb +119 -0
  43. data/lib/woods/chunking/semantic_chunker.rb +44 -21
  44. data/lib/woods/console/connection_manager.rb +56 -3
  45. data/lib/woods/console/embedded_executor.rb +30 -5
  46. data/lib/woods/console/rack_middleware.rb +29 -1
  47. data/lib/woods/dependency_graph.rb +34 -10
  48. data/lib/woods/embedding/fake.rb +12 -0
  49. data/lib/woods/embedding/indexer.rb +195 -98
  50. data/lib/woods/embedding/input_budget.rb +67 -0
  51. data/lib/woods/embedding/openai.rb +70 -20
  52. data/lib/woods/embedding/provider.rb +37 -25
  53. data/lib/woods/embedding/text_preparer.rb +76 -32
  54. data/lib/woods/embedding/token_counter.rb +18 -81
  55. data/lib/woods/embedding/vector_configuration.rb +48 -0
  56. data/lib/woods/extraction_identities.rb +175 -0
  57. data/lib/woods/extractor.rb +304 -107
  58. data/lib/woods/extractors/action_cable_extractor.rb +8 -3
  59. data/lib/woods/extractors/assigned_value_discovery.rb +74 -0
  60. data/lib/woods/extractors/class_declarations.rb +121 -0
  61. data/lib/woods/extractors/configuration_extractor.rb +11 -3
  62. data/lib/woods/extractors/declaration_ancestry.rb +92 -0
  63. data/lib/woods/extractors/event_extractor.rb +8 -0
  64. data/lib/woods/extractors/graphql_extractor.rb +134 -77
  65. data/lib/woods/extractors/job_extractor.rb +5 -1
  66. data/lib/woods/extractors/lib_extractor.rb +132 -15
  67. data/lib/woods/extractors/mailer_extractor.rb +3 -5
  68. data/lib/woods/extractors/manager_extractor.rb +7 -21
  69. data/lib/woods/extractors/migration_declaration.rb +87 -0
  70. data/lib/woods/extractors/migration_extractor.rb +5 -39
  71. data/lib/woods/extractors/phlex_extractor.rb +6 -2
  72. data/lib/woods/extractors/policy_extractor.rb +9 -5
  73. data/lib/woods/extractors/poro_extractor.rb +112 -53
  74. data/lib/woods/extractors/pundit_extractor.rb +11 -6
  75. data/lib/woods/extractors/scheduled_job_extractor.rb +45 -4
  76. data/lib/woods/extractors/serializer_extractor.rb +34 -22
  77. data/lib/woods/extractors/shared_utility_methods.rb +18 -1
  78. data/lib/woods/extractors/source_nesting.rb +142 -106
  79. data/lib/woods/extractors/standalone_module_discovery.rb +123 -0
  80. data/lib/woods/extractors/state_machine_extractor.rb +46 -40
  81. data/lib/woods/extractors/view_component_extractor.rb +9 -7
  82. data/lib/woods/flow_assembler.rb +4 -1
  83. data/lib/woods/generation.rb +25 -0
  84. data/lib/woods/hooks/context_hint.rb +7 -2
  85. data/lib/woods/mcp/bootstrapper.rb +33 -7
  86. data/lib/woods/mcp/config_resolver.rb +26 -7
  87. data/lib/woods/mcp/index_reader.rb +125 -24
  88. data/lib/woods/mcp/index_reader_pinning.rb +16 -0
  89. data/lib/woods/mcp/renderers/markdown_renderer.rb +7 -1
  90. data/lib/woods/mcp/renderers/plain_renderer.rb +3 -1
  91. data/lib/woods/mcp/search_results.rb +7 -1
  92. data/lib/woods/mcp/server.rb +24 -4
  93. data/lib/woods/module_reconciliation.rb +151 -0
  94. data/lib/woods/path_dispatcher.rb +7 -2
  95. data/lib/woods/rake_helpers.rb +43 -11
  96. data/lib/woods/release.rb +1 -1
  97. data/lib/woods/resilience/index_validator.rb +8 -3
  98. data/lib/woods/resilience/retryable_provider.rb +18 -1
  99. data/lib/woods/resolved_config.rb +68 -8
  100. data/lib/woods/retrieval/context_assembler.rb +3 -3
  101. data/lib/woods/retrieval/lexical_assembler.rb +3 -2
  102. data/lib/woods/retrieval/scope.rb +18 -2
  103. data/lib/woods/retrieval/source_evidence.rb +14 -2
  104. data/lib/woods/source_contributor_validation.rb +78 -0
  105. data/lib/woods/source_contributors.rb +116 -0
  106. data/lib/woods/source_inputs/handoff.rb +37 -0
  107. data/lib/woods/source_inputs/launcher.rb +53 -13
  108. data/lib/woods/source_inputs/manifest.rb +84 -3
  109. data/lib/woods/source_inputs/private_key.rb +44 -12
  110. data/lib/woods/source_inputs/scanner.rb +98 -27
  111. data/lib/woods/source_inputs/scopes.rb +1 -1
  112. data/lib/woods/source_inputs/session.rb +147 -15
  113. data/lib/woods/source_inputs/stable_reader.rb +127 -0
  114. data/lib/woods/source_inputs/status.rb +40 -8
  115. data/lib/woods/source_inputs/verifier.rb +28 -5
  116. data/lib/woods/source_path_encoding.rb +33 -0
  117. data/lib/woods/source_references/cache.rb +284 -0
  118. data/lib/woods/source_references/collector.rb +120 -0
  119. data/lib/woods/source_references/extraction.rb +185 -0
  120. data/lib/woods/source_references/inputs.rb +134 -0
  121. data/lib/woods/source_references/parser_adapter.rb +134 -0
  122. data/lib/woods/source_references/pass.rb +152 -0
  123. data/lib/woods/source_references/prism_adapter.rb +116 -0
  124. data/lib/woods/source_references/registry.rb +178 -0
  125. data/lib/woods/source_references/runtime_lookup.rb +127 -0
  126. data/lib/woods/source_references/value_class.rb +82 -0
  127. data/lib/woods/storage/metadata_store.rb +4 -1
  128. data/lib/woods/storage/qdrant.rb +2 -2
  129. data/lib/woods/unblocked/client.rb +12 -7
  130. data/lib/woods/unblocked/document_builder.rb +4 -1
  131. data/lib/woods/unblocked/exporter.rb +127 -37
  132. data/lib/woods/unblocked/sync_manifest.rb +137 -21
  133. data/lib/woods/unblocked/uri_migration.rb +105 -0
  134. data/lib/woods/util/host_guard.rb +3 -2
  135. data/lib/woods/version.rb +1 -1
  136. data/lib/woods/watch/catch_up.rb +138 -0
  137. data/lib/woods/watch/claim_lease.rb +150 -0
  138. data/lib/woods/watch/cli.rb +26 -2
  139. data/lib/woods/watch/daemon.rb +80 -59
  140. data/lib/woods/watch/installation/options.rb +1 -1
  141. data/lib/woods/watch/installation/receipt.rb +6 -1
  142. data/lib/woods/watch/managed_child.rb +1 -1
  143. data/lib/woods/watch/supervisor.rb +1 -1
  144. data/lib/woods/watch/tree_scan.rb +14 -2
  145. data/plugin/.claude-plugin/plugin.json +1 -1
  146. data/plugin/hooks/adapters/normalize.rb +3 -2
  147. data/plugin/hooks/woods-input-rules.sh +4 -0
  148. data/plugin/hooks/woods-refresh.sh +15 -7
  149. data/plugin/hooks/woods-session-start.sh +60 -3
  150. data/plugin/skills/woods-diagnose/SKILL.md +334 -11
  151. data/plugin/skills/woods-investigate/SKILL.md +11 -0
  152. data/plugin/skills/woods-mcp-config/SKILL.md +79 -8
  153. data/plugin/skills/woods-setup/SKILL.md +53 -6
  154. metadata +32 -5
@@ -30,7 +30,10 @@ Three differences are tolerated, and nothing else:
30
30
 
31
31
  The unit-file write skip ignores only Woods' top-level `extracted_at` stamp.
32
32
  A nested metadata field with the same name is application data: changing it
33
- rewrites the unit in both compact and pretty JSON output.
33
+ rewrites the unit in both compact and pretty JSON output. On supporting
34
+ Woods 2.1 writers, an already registered unit whose complete serialized bytes
35
+ match also stays out of the touched set, Git enrichment and dependents updates.
36
+ Derived metadata or dependency differences still use normal registration.
34
37
 
35
38
  `graph_analysis.json` used to be a fourth row, tolerating list ordering. It no
36
39
  longer is: the analyzer is order-independent and the oracle compares the file
@@ -58,22 +61,35 @@ failure a CI chain cannot afford:
58
61
 
59
62
  | Situation | Behavior |
60
63
  |---|---|
61
- | `CHANGED_FILES` is set | Used verbatim (comma-separated paths); git is not consulted. |
64
+ | `CHANGED_FILES` is set | Comma-separated application-relative or contained absolute paths; normalized before filtering. Git is not consulted. |
62
65
  | The git range resolves | Current behavior: extract the changed paths, or exit 0 with `No relevant files changed` when nothing relevant changed. |
66
+ | A changed input requires a restart | On revisions containing #588, the fresh one-shot task selects full extraction before relevance filtering, including schema-only changes. |
63
67
  | The range fails **and** a `:running` watch daemon maintains the index | Stand down with a printed reason, exit 0 — the daemon's start-up catch-up covers whatever changed. |
64
68
  | The range fails otherwise | Actionable error naming the range, **exit 1**. |
65
69
  | There is no `git` binary at all | Same two rows as above: the failure reads `git unavailable: …` and takes the daemon-coverage decision, rather than dying with an `Errno::ENOENT` backtrace. |
66
70
 
67
- Changed paths are normalized lexically before dispatch: trailing root slashes,
68
- duplicate separators and `.`/`..` segments do not create separate changes or
69
- bypass matching. Missing files remain representable; symlinks are not resolved.
71
+ Changed paths are normalized lexically before the task's relevance filter:
72
+ trailing root slashes, duplicate separators and `.`/`..` segments do not create
73
+ separate changes or bypass matching. Paths outside `Rails.root` are excluded.
74
+ Missing files remain representable; symlinks are not resolved. The task-boundary
75
+ normalization and nested-application Git paths are included in Woods 2.1;
76
+ check the installed revision before relying on them.
77
+
78
+ **Included in Woods 2.1:** incremental extraction and targeted refresh refuse
79
+ legacy flat artifacts with a manifest and generations whose manifest records a
80
+ writer major version below 2. Run a full `woods:extract` before resuming partial
81
+ updates. Missing writer provenance in a generation remains compatible with early
82
+ v2 betas. Reading old indexes and rebuilding them fully remain supported; see
83
+ [the v1 upgrade procedure](UPGRADING_TO_2.md).
70
84
 
71
85
  The range comes from `CI_COMMIT_BEFORE_SHA..CI_COMMIT_SHA` (GitLab),
72
- `origin/$GITHUB_BASE_REF...HEAD` (GitHub Actions), or `HEAD~1` (default). An
73
- unresolvable range — a GitLab zero-SHA on a new branch, an unfetched base ref,
74
- a shallow clone with no `HEAD~1` — reads as "nothing changed" to git, which is
75
- why a failed range must not be mistaken for an empty one: the sync never ran,
76
- and CI drift would stay unbounded. A degraded daemon covers nothing, so it
86
+ `origin/$GITHUB_BASE_REF...HEAD` (GitHub Actions), or `HEAD~1` (default).
87
+ Whitespace-only CI variables are ignored; a nonempty GitLab before-SHA with
88
+ no current SHA compares against `HEAD`. Nonempty invalid revisions still fail.
89
+ An unresolvable range — a GitLab zero-SHA on a new branch, an unfetched base ref,
90
+ a shallow clone with no `HEAD~1` — cannot establish which files changed.
91
+ The task must not mistake that failed diff for an empty change set.
92
+ A degraded daemon covers nothing, so it
77
93
  does not stand the run down. `WOODS_IGNORE_WATCH=1` removes daemon coverage
78
94
  too — with it set, a failed range exits 1. A slim image with no `git` binary
79
95
  resolves to the same decision rather than a raw `Errno::ENOENT`: the failure is
@@ -82,14 +98,18 @@ and an uncovered one still gets the remediation text.
82
98
 
83
99
  Recovery choices, in the order they are worth trying:
84
100
 
85
- 1. Repair or provide the range: fetch the base ref (`fetch-depth: 2` or more),
86
- or correct the CI environment variables that build it.
101
+ 1. Repair or provide the range: fetch the actual base ref and enough history
102
+ to find its merge base, or correct the CI environment variables that build
103
+ it. Depth two alone does not fetch a pull request's base branch.
87
104
  2. Set `CHANGED_FILES` explicitly from your CI platform, bypassing git range
88
105
  resolution entirely.
89
106
  3. Run a full `woods:extract` when the range cannot be repaired this run.
90
107
 
91
108
  The diff itself is rooted at the extracted application (`git -C Rails.root`),
92
- independently of the process working directory. An explicit `WOODS_GIT_DIR`
109
+ independently of the process working directory. Paths are application-relative
110
+ even when Rails lives below the repository root. Deletions and both sides of
111
+ renames within the application remain in the change set; sibling applications
112
+ are excluded. An explicit `WOODS_GIT_DIR`
93
113
  selects that Git directory's HEAD for both the diff and manifest provenance.
94
114
  For a linked worktree, use its worktree-specific directory within the complete
95
115
  shared layout; selecting the shared root instead reads the primary checkout's
@@ -103,6 +123,101 @@ including inlined code and callback analysis. Multiple runtime mixins sharing a
103
123
  file retain separate identities and refresh all their includers. Run a full extraction after upgrading
104
124
  to populate these previously missing source mappings.
105
125
 
126
+ ### GitHub Actions with an exact baseline
127
+
128
+ The restored index must describe the first commit in the selected diff.
129
+ A cache from an unrelated branch or older commit is not a valid baseline for
130
+ `HEAD~1` or a pull request's merge base. The recipe below restores only the
131
+ selected commit's exact cache key and runs full extraction on a cache miss.
132
+ It fetches complete history and the actual pull-request base ref; see
133
+ [checkout's history setting](https://github.com/actions/checkout#usage) and
134
+ [cache restore's exact-hit output](https://github.com/actions/cache/blob/v4/restore/README.md#outputs).
135
+
136
+ Adapt database setup, Ruby configuration, and the index path to the host app.
137
+ This example runs from a Rails app at the repository root. For a nested app,
138
+ set the run steps' working directory and adjust the cache path, hash paths,
139
+ and cache namespace to identify that app. Include every extraction-affecting
140
+ configuration input in the cache namespace; change it after a Woods upgrade
141
+ that needs a full baseline. An unverified or incomplete prior index needs a
142
+ full extraction even when a cache key matches.
143
+
144
+ ```yaml
145
+ # .github/workflows/woods.yml
146
+ name: Update Codebase Index
147
+ on:
148
+ push:
149
+ branches: [main]
150
+ pull_request:
151
+
152
+ jobs:
153
+ index:
154
+ runs-on: ubuntu-latest
155
+ env:
156
+ RAILS_ENV: test
157
+ WOODS_IGNORE_WATCH: "1"
158
+ steps:
159
+ - uses: actions/checkout@v4
160
+ with:
161
+ fetch-depth: 0
162
+ - uses: ruby/setup-ruby@v1
163
+ with:
164
+ bundler-cache: true
165
+ - name: Select the baseline commit
166
+ id: base
167
+ env:
168
+ WOODS_BASE_REF: ${{ github.base_ref }}
169
+ WOODS_BEFORE_SHA: ${{ github.event.before }}
170
+ run: |
171
+ if [ -n "$WOODS_BASE_REF" ]; then
172
+ git fetch --no-tags origin "+refs/heads/$WOODS_BASE_REF:refs/remotes/origin/$WOODS_BASE_REF"
173
+ base="$(git merge-base "origin/$WOODS_BASE_REF" HEAD)"
174
+ elif [ -n "$WOODS_BEFORE_SHA" ] && git rev-parse --verify "$WOODS_BEFORE_SHA^{commit}" >/dev/null 2>&1; then
175
+ base="$WOODS_BEFORE_SHA"
176
+ else
177
+ base=""
178
+ fi
179
+ printf 'sha=%s\n' "$base" >> "$GITHUB_OUTPUT"
180
+ - name: Restore exactly that baseline
181
+ id: index-cache
182
+ if: steps.base.outputs.sha != ''
183
+ uses: actions/cache/restore@v4
184
+ with:
185
+ path: tmp/woods
186
+ key: woods-v2-app-${{ runner.os }}-${{ hashFiles('Gemfile.lock', 'config/initializers/woods.rb') }}-${{ steps.base.outputs.sha }}
187
+ - name: Prepare the application database
188
+ run: bin/rails db:prepare
189
+ - name: Update the index
190
+ env:
191
+ WOODS_EXACT_BASELINE: ${{ steps.index-cache.outputs.cache-hit }}
192
+ CI_COMMIT_BEFORE_SHA: ${{ steps.base.outputs.sha }}
193
+ CI_COMMIT_SHA: ${{ github.sha }}
194
+ run: |
195
+ if [ "$WOODS_EXACT_BASELINE" = true ]; then
196
+ bin/rails woods:incremental
197
+ else
198
+ bin/rails woods:extract
199
+ fi
200
+ - name: Validate the index
201
+ run: bin/rails woods:validate
202
+ - name: Save the validated current index
203
+ uses: actions/cache/save@v4
204
+ with:
205
+ path: tmp/woods
206
+ key: woods-v2-app-${{ runner.os }}-${{ hashFiles('Gemfile.lock', 'config/initializers/woods.rb') }}-${{ github.sha }}
207
+ ```
208
+
209
+ There are deliberately no `restore-keys`: a partial match selects full
210
+ extraction. A first push, an unavailable before-SHA, or an absent baseline
211
+ cache also selects full extraction. A failed explicit PR-base fetch stops the
212
+ job with its Git error. The validated publication is saved under the current
213
+ checkout's SHA, never under the old baseline key.
214
+
215
+ For Docker CI, run database preparation and Woods tasks through the application
216
+ service, forward the selected CI variables into that container, and cache the
217
+ host-visible mount of the same index. Fetching the base only on a host whose
218
+ Git object store is absent from the container does not make that range usable
219
+ inside the application.
220
+
106
221
  ## Handled source errors and retry
107
222
 
108
223
  Included in Woods `2.0.0`: when an incremental extraction or named refresh
@@ -133,6 +248,19 @@ step before it.
133
248
  path no longer produces are dropped. This is what indexes a file the index
134
249
  has never seen, and what removes definitions deleted from a surviving
135
250
  source file. Multi-file Rake tasks use wholesale reconciliation below.
251
+
252
+ **Included in Woods 2.1:** changed-path candidates are collected and checked
253
+ before registration or pruning. Moving an identity out of a surviving file
254
+ works in either changed-path order only when completed extraction proves
255
+ the old file no longer produces it and the new owner is unique. Failed or
256
+ unsupported extraction cannot release an owner. For runtime classes, a
257
+ complete eager load, an authoritative discovery inventory with one current
258
+ class, and that class's canonical source location can establish the move.
259
+ Jobs and serializers additionally require their complete file/runtime
260
+ inventory to rule out a source-only duplicate. GraphQL's mixed runtime/file
261
+ inventory does not grant this authority.
262
+ Incomplete eager loading cannot establish a surviving-file ownership move;
263
+ genuine simultaneous source owners still abort before publication.
136
264
  3. **Re-extract the rest of the blast radius**: units whose own file did not
137
265
  change but which depend on something that did.
138
266
  4. **Reconcile class-based types** against each extractor's
@@ -140,6 +268,16 @@ step before it.
140
268
  classes the graph still holds that the set no longer contains. Exact by
141
269
  construction: it is the same discovery code a full extraction uses, so
142
270
  there is no path-to-constant guessing.
271
+
272
+ **Included in Woods 2.1:** jobs and serializers reconcile their combined
273
+ file/runtime inventory whenever a Ruby source path changes. Resolved metadata
274
+ can depend on another file without a recorded graph edge, such as a nested,
275
+ class-discovered job's `queue_as Settings::QUEUE`. Runtime queue enrichment
276
+ applies only to class-discovered jobs; file-discovered jobs keep source-based
277
+ queue metadata. Non-Ruby batches refresh only hybrid families
278
+ already reached by the pre-change dependency graph. The application must have
279
+ loaded the changed runtime before extraction, through a fresh boot or reload.
280
+
143
281
  5. **Re-run whole-app extractors** whose trigger paths changed, replacing that
144
282
  unit type wholesale.
145
283
  6. **Prune vanished units**, so anything steps 2–5 resurrected against a
@@ -164,6 +302,14 @@ step before it.
164
302
  string literal, and an unrelated addition in the same batch do not.
165
303
  Idempotent when nothing was pruned.
166
304
 
305
+ 8. **Reconcile source references** on Woods 2.1 writers described in
306
+ [constant source references](EXTRACTOR_REFERENCE.md#constant-source-references).
307
+ Resolve cached candidates against the complete current typed unit registry,
308
+ update callers' forward relationships, and refresh targets' reverse
309
+ relationships. Unresolved candidates allow an unchanged caller to gain an edge
310
+ when its target becomes indexed. Parsing is reused only when the captured
311
+ source identity still matches.
312
+
167
313
  Git enrichment uses the same eligibility checks in full and incremental runs:
168
314
  existing app-owned files under `Rails.root`, excluding `vendor/`, `node_modules/`,
169
315
  and framework/gem source units. Each typed unit resolves its own file history,
@@ -317,12 +463,27 @@ discovery set.
317
463
 
318
464
  ### Runtime removals and bundle updates
319
465
 
320
- Jobs discovered through `ApplicationJob.descendants` supplement the job-file
321
- scan, but jobs are not part of `CLASS_BASED_DISCOVERY` removal reconciliation.
322
- If a dynamically defined or gem-owned job disappears without a tracked source
323
- path changing, its unit can survive subsequent incremental runs. A full
324
- extraction in a fresh Rails process removes it; an in-process full extraction
325
- can still see an old constant retained by that process (B-165).
466
+ Hybrid runtime reconciliation (#588) is included in Woods 2.1; check the
467
+ loaded writer revision. Jobs and serializers use the union of their directory
468
+ scan and current runtime descendants. A Ruby-file change, or a blast radius
469
+ containing a job or serializer, reruns both families wholesale. This costs a
470
+ complete job and serializer scan for that batch, preserving the same file-pass
471
+ precedence, nested identities, JSON and graph facts as full extraction. A
472
+ serializer nested in a controller retains its own identity; the controller does
473
+ not become a serializer unit. Runtime objects left behind by a reload cannot
474
+ claim a name now owned by another constant.
475
+
476
+ All members of those two families count as refreshed and participate in graph,
477
+ source-reference and metadata maintenance, even when only one file changed.
478
+ Unit-file writes retain the existing byte-equality check; this does not promise
479
+ a single-file write cost or an unchanged manifest for such a batch.
480
+
481
+ Discovery-based removal requires a complete eager load. Failed or reported-partial discovery
482
+ refuses publication, preserving the previous generation; incomplete eager
483
+ loading retains missing units. An empty change set does not discover arbitrary
484
+ runtime-only additions or removals. Use `refresh(:jobs, :serializers)` after
485
+ updating the runtime explicitly, or full extraction in a fresh process after a
486
+ bundle change. An old constant still bound in that process remains observable.
326
487
 
327
488
  After adding, removing, or updating bundled gems, boot the updated bundle in a
328
489
  fresh process and run:
@@ -373,6 +534,13 @@ The family has three parts: `flows/flow_index.json` (entry point → relative
373
534
  document path), one document per controller action, and
374
535
  `metadata[:flow_paths]` on the controller units.
375
536
 
537
+ On revisions containing #588 (included in Woods 2.1), switching the gate off
538
+ withdraws the seeded `flows/` family and removes controller flow annotations on
539
+ the next full, incremental or targeted-refresh publication. An incremental run
540
+ with no changed files still publishes this withdrawal. The preceding generation
541
+ is unchanged; any withdrawal failure prevents publication. `trace_flow` then
542
+ uses its normal query-time assembly instead of an obsolete precomputed document.
543
+
376
544
  A full extraction computes all three in one pass. An incremental run computes
377
545
  them for its **delta**:
378
546
 
@@ -440,6 +608,11 @@ cascades to `ROUTE_CONSUMER_EXTRACTORS` for the reason given above. Like an
440
608
  incremental run, `refresh` rewrites the graph, `graph_analysis.json`, the
441
609
  affected type indexes and the manifest, so the result is durable.
442
610
 
611
+ `refresh(:configurations)` derives `BehavioralProfile` from resolved
612
+ `Rails.application.config`. Its nominal `config/application.rb` source path is
613
+ not a separate configuration unit. Refresh assumes the caller already updated
614
+ the runtime; it does not rerun Rails initialization or apply database migrations.
615
+
443
616
  ## What a change actually requires: reload, restart, or neither
444
617
 
445
618
  `Woods::ReloadPolicy` answers the question a resident process has to ask before
@@ -482,8 +655,31 @@ it, and belong to whoever implements the reload step:
482
655
  the batch demands, and `paths_requiring(:restart)` names the offending paths in
483
656
  the restart message a supervisor sees. See `docs/WATCH_DAEMON.md`.
484
657
 
658
+ The one-shot `woods:incremental` task also uses the shared `InputRules` decision
659
+ before relevance filtering on revisions containing #588. Run it in a fresh
660
+ Rails process with migrations already applied: `db/schema.rb`,
661
+ `db/structure.sql`, and other restart inputs select full extraction, including
662
+ when mixed with ordinary model edits. A migration source file alone retains
663
+ its existing file-extraction classification; Woods does not apply migrations.
664
+
665
+ The direct `Extractor#extract_changed` API and optional embedded
666
+ `pipeline_extract` incremental mode cannot establish a fresh boot. They refuse
667
+ restart-sensitive batches before creating a payload and instruct the caller to
668
+ run `woods:extract` in a fresh Rails process. With the MCP Tasks extension this
669
+ is a failed task, and the prior generation stays active. Calling a full
670
+ extraction on an already stale embedded runtime does not refresh that runtime.
671
+
485
672
  ## Running the differential harness
486
673
 
674
+ The real ActiveJob/AMS and fresh schema-task probes have a separate optional
675
+ bundle and run in their own processes (also in the Rails 7.2 CI row):
676
+
677
+ ```bash
678
+ BUNDLE_GEMFILE=gemfiles/serializers.gemfile bundle install
679
+ WOODS_RUN_BOOTED_APP=1 BUNDLE_GEMFILE=gemfiles/serializers.gemfile \
680
+ bin/rspec spec/integration/incremental_runtime_spec.rb
681
+ ```
682
+
487
683
  `spec/integration/incremental_equivalence_spec.rb` is the oracle. It boots the
488
684
  `spec/dummy` app against a tmpdir copy, applies randomized
489
685
  create/modify/delete/rename sequences, and compares the maintained index to a
@@ -579,6 +775,15 @@ edit does not establish that a day of commits is below the crossover.
579
775
  same type+identifier are not representable; full extraction fails closed
580
776
  with both source paths instead of publishing a glob-order tie-break (resolved
581
777
  B-063). Same-file re-derivation remains a legitimate deduplication case.
778
+ **Included in Woods 2.1:** incremental extraction and targeted refresh enforce
779
+ the same collision refusal (#561), preserving the prior published generation.
780
+ Distinct unit types in separate extractor directories remain independent.
781
+ A retained source can move when the old file is confirmed absent and the new
782
+ file exists; two conflicting sources produced within one run always refuse.
783
+ Complete wholesale replacement can relocate an owner, such as framework
784
+ source after a gem upgrade. Partial runtime discovery cannot grant that authority.
785
+ Correct the producer or declarations, then retry the complete failed batch.
786
+ A watcher retains failed paths and reports degraded status until corrected.
582
787
  - **Class-based units are never swept**: see [Deletion](#deletion) above for
583
788
  why (the `SchemaMigration`/`InternalMetadata` convention-path case).
584
789
  Deleting a class-based unit therefore requires either the caller naming the
@@ -598,3 +803,40 @@ edit does not establish that a day of commits is below the crossover.
598
803
  contract, and multi-worktree operation, all landed alongside this work
599
804
  (B-064, resolved). `docs/WATCH_DAEMON.md` covers them, including the parts
600
805
  that remain unmeasured.
806
+
807
+ ## Source-reference baseline and upgrades
808
+
809
+ **Included in Woods 2.1.** Older indexes remain readable.
810
+ Writers with the [source-reference expansion](EXTRACTOR_REFERENCE.md#constant-source-references)
811
+ require one full `bin/rails woods:extract` before incremental extraction or
812
+ targeted refresh can update an older index without its reference cache. Follow
813
+ with `bin/rails woods:validate` using the application's normal task launcher.
814
+
815
+ `source_references.json` is internal writer state in the published payload.
816
+ Preserve it with the complete payload when copying an index. Missing/incompatible
817
+ cache state or unverified retained source stops publication with a full-extraction
818
+ diagnostic. Reconstruct it by running a full extraction; do not manufacture a
819
+ cache or delete it to bypass the check. The preceding generation remains active.
820
+ The watcher preserves failed batches but does not automatically repair this
821
+ baseline: establish the full baseline before resuming incremental maintenance.
822
+
823
+ This also tightens scoped refreshes: if a retained reference-bearing Ruby unit's
824
+ source changed outside the selected batch, the writer refuses to combine its old
825
+ runtime facts with references resolved from the new source. For example,
826
+ `woods:refresh[events]` cannot adopt an edited service unit. The prior generation
827
+ stays active and can report `drifted`; use a full extraction to establish a
828
+ consistent baseline. Omitted non-reference inputs, such as a view changed before
829
+ capture, retain the existing per-consumer freshness behavior.
830
+
831
+ A source change during reference analysis or final verification also prevents
832
+ publication. Correct source errors and retry the complete batch against a stable
833
+ tree. Reference-only edge updates do not refresh a retained unit's runtime
834
+ metadata, extraction timestamp or Git history.
835
+
836
+ Whole-extractor failure reporting (#584) is included in Woods 2.1; verify the
837
+ writer revision before relying on it. If a selected whole-app extractor cannot
838
+ be initialized or raises before replacing any units, incremental extraction and
839
+ targeted refresh fail without advancing the published generation, even when
840
+ other extractors succeeded. Correct the logged cause and retry the complete
841
+ changed-file batch or refresh selection. The existing refusal after a
842
+ replacement has begun writing remains in place.
data/docs/INDEX_LAYOUT.md CHANGED
@@ -63,10 +63,15 @@ gates below deliberately fail more strictly.
63
63
 
64
64
  ### Payload artifacts
65
65
 
66
+ Read JSON artifacts as UTF-8 regardless of the reader process's locale. Unit
67
+ identifiers, source code, paths, and metadata can contain non-ASCII characters.
68
+ Reject invalid UTF-8 rather than replacing bytes in published source evidence.
69
+
66
70
  | Artifact | Presence and meaning |
67
71
  |---|---|
68
72
  | `manifest.json` | Required for a complete structural publication. Counts by extractor directory, totals, extraction timestamp and provenance. Optional fields vary by writer/version; see [writer provenance](PUBLISHED_INDEX.md#manifest-writer-provenance). |
69
73
  | `source_inputs.json` | Versioned source-input identities and per-consumer provenance for this exact generation; old indexes may omit it. See [source freshness](SOURCE_FRESHNESS.md). |
74
+ | `source_references.json` | Internal, versioned candidate and edge-ownership cache for this generation. Included in Woods 2.1. Older indexes remain readable without it; supporting incremental writers may require a [full baseline rebuild](INCREMENTAL_EXTRACTION.md#source-reference-baseline-and-upgrades). Preserve it when copying payloads. Not a public query schema. |
70
75
  | `dependency_graph.json` | Required for a complete structural publication. Typed graph data; an empty graph is valid. |
71
76
  | `<type>/_index.json` and unit JSON | Present for extracted families. `_index.json` is an array of unit summaries; an empty array is valid. Disabled/unavailable families may be absent. Do not infer completeness from a fixed count of directories. |
72
77
  | `graph_analysis.json` | Derived graph analysis when produced. Treat absence as unavailable analysis, not an empty or corrupt unit index. |
data/docs/INTERNALS.md CHANGED
@@ -388,6 +388,15 @@ All Console Server queries run inside a **rolled-back transaction** (`SafeContex
388
388
 
389
389
  Large units are split into semantic chunks before embedding. The `SemanticChunker` is type-aware, it doesn't split on arbitrary token counts.
390
390
 
391
+ **Included in Woods 2.1:** class-like method chunks retain their full
392
+ method name, including `self.`, `?`, `!`, setters and operators. For example,
393
+ `Callable#method_call` and `Callable#method_self.call` contain the instance and
394
+ class implementations separately, in source order. Plain instance-method chunk
395
+ identities stay the same; formerly truncated class or punctuation-bearing names
396
+ change to their distinct identities. Rebuild embeddings for affected units to
397
+ replace old collapsed chunks. This remains a line-based chunker, rather than a
398
+ complete Ruby parser.
399
+
391
400
  ### Model chunking
392
401
 
393
402
  Models are split into purpose-specific sections:
@@ -86,32 +86,37 @@ Clients must send `Authorization: Bearer $WOODS_MCP_HTTP_TOKEN` on every request
86
86
 
87
87
  ### Browser origins (DNS rebinding defense)
88
88
 
89
- **Unreleased diagnostic correction:** invalid origin encodings also refuse before
89
+ **Included in Woods 2.1:** invalid origin encodings also refuse before
90
90
  binding HTTP. `woods-mcp-http` exits 2 with one bounded `ConfigurationError`
91
91
  message naming the invalid entry; re-enter that origin using an ASCII hostname
92
92
  (or its Punycode form). No allowlist keeps the existing defaults. Explicit entries
93
93
  are literal origins, not wildcard patterns.
94
- Retain the actual request Host through a reverse proxy; forwarded headers do not
95
- replace it. When authentication is configured, requests without Host still pass
96
- through bearer authentication.
97
94
 
98
- In the 2.0.1 maintenance patch, preflight and SDK dispatch share an immutable
99
- normalized policy. A configured list replaces the default browser origins;
100
- include loopback explicitly if needed. Cross-origin ports must match an entry,
101
- while omitted default HTTP(S) ports match their explicit 80/443 forms. A portless
102
- entry permits same-authority traffic rather than every cross-port browser origin.
103
- Console defaults remain HTTP loopback; Index HTTP defaults include HTTP and HTTPS
104
- loopback. When no list is set, these defaults are unchanged. Invalid entries fail
105
- at boot, naming the offending entry. Restart after changing the configuration.
106
-
107
- A second middleware, `Woods::MCP::OriginGuard`, rejects requests whose `Origin` header is outside an allow-list. Requests without an `Origin` header still pass through Host validation and configured bearer authentication.
95
+ A second middleware, `Woods::MCP::OriginGuard`, rejects requests whose `Origin` header is outside an allow-list. Requests without an `Origin` header still validate `Host`; bearer authentication remains required when configured.
108
96
 
109
97
  | Scenario | `WOODS_MCP_HTTP_ALLOWED_ORIGINS` | Origins accepted |
110
98
  |--------------------|-----------------------------------------|-------------------------------------------------------------------|
111
- | default | unset | HTTP(S) loopback, matching request authority; explicit entries for cross-port origins |
99
+ | default | unset | loopback HTTP(S) origins matching the request Host authority |
112
100
  | explicit list | `https://app.example.com` | exactly `https://app.example.com`, loopback no longer allowed |
113
101
  | multiple origins | `https://a.example,https://b.example` | each listed origin |
114
102
 
103
+ **Included in Woods 2.1:** preflight and SDK dispatch share one immutable policy.
104
+ Explicit cross-origin entries match the complete origin; default HTTP(S) ports
105
+ (`:80` and `:443`) are equivalent to their omitted form.
106
+ A portless entry additionally permits same-authority requests on other ports;
107
+ it does not grant arbitrary cross-port CORS. For example, a browser at
108
+ `http://localhost:3000` calling an endpoint on port 9292 must list the browser's
109
+ actual origin. A custom list replaces default browser-origin rules. Configure
110
+ non-loopback endpoint authorities too, and retain the request's real Host through
111
+ any reverse proxy. IPv6 origins use brackets, e.g. `http://[::1]:3000`.
112
+
113
+ Both the preflight and SDK transport use this default-port equivalence. A genuinely
114
+ absent Origin is allowed after Host validation; blank or malformed headers are
115
+ rejected. Invalid configured origin URLs refuse once at boot with a bounded diagnostic
116
+ naming the offending entry; this is stricter than 2.0.0 startup behavior.
117
+ Restart after changing origins: the middleware and transport capture the same
118
+ policy. SDK DNS-rebinding checks remain enabled.
119
+
115
120
  `OPTIONS` preflights are answered with the matching `Access-Control-Allow-*` headers; successful responses carry `Access-Control-Allow-Origin` and `Vary: Origin`. `Access-Control-Expose-Headers: Mcp-Session-Id` appears only in legacy session mode (`WOODS_MCP_HTTP_STATELESS=0`).
116
121
 
117
122
  ### Origin configuration compatibility
data/docs/MCP_SERVERS.md CHANGED
@@ -147,7 +147,10 @@ than pinning a newer version solely to obtain guidance.
147
147
 
148
148
  ### Structural and semantic readiness
149
149
 
150
- `ready` describes the published structural index. A reachable embedding provider
150
+ `ready` describes the published structural index, not corpus size or extraction
151
+ coverage. An intentionally empty application can have a valid, ready zero-unit
152
+ index; inspect `structure` counts and extraction diagnostics when zero is
153
+ unexpected. A reachable embedding provider
151
154
  and bootstrap state `hydrated` do not establish that semantic stores contain
152
155
  data. Supporting readers also report `retriever.corpus`: locally known vector
153
156
  and metadata entry counts, counts by type, and whether those stores are empty,
@@ -209,6 +212,31 @@ error and continues serving the previous aligned generation; it never swaps in a
209
212
  partial or empty replacement. Grant write access for live reloads, or restart the MCP
210
213
  process after publishing a new embedded index.
211
214
 
215
+ **Included in Woods 2.1:** the `:local` preset combines snapshot vectors with
216
+ SQLite metadata, which cannot be refreshed together atomically in the running
217
+ server. Its `reload` returns a degraded error and retains the previous aligned
218
+ state even when the directory is writable. Restart `woods-mcp` after
219
+ `woods:embed` to load the new vectors. See the [backend matrix](BACKEND_MATRIX.md#persistence-story).
220
+
221
+ ### Resource identity and damaged generation markers
222
+
223
+ **Included in Woods 2.1 (#593):** unit resource identifiers may contain
224
+ slashes, such as `views/posts/show.html.erb` or `GET /posts/:id`. Encode the
225
+ whole identifier as one URI segment (`%2F` for `/`) in
226
+ `codebase://unit/{identifier}`. Raw multi-segment paths, dot traversal segments,
227
+ double encoding and backslashes remain invalid. Type resource names cannot
228
+ contain slashes.
229
+
230
+ A held-open server returns `corrupt_artifact` for a malformed published
231
+ generation marker or an unavailable/outside-root payload pointer. Repeated
232
+ calls retain that error rather than presenting the cached generation as current.
233
+ An already pinned request can finish against its known generation. New requests
234
+ resume after a valid marker is restored; a leftover flat root manifest is not
235
+ substituted for a missing or invalid named payload. Run
236
+ `woods:validate` and restore a known-good publication or complete a fresh
237
+ extraction; do not point the marker at an arbitrary payload directory.
238
+ Legacy flat indexes without a generation marker remain supported.
239
+
212
240
  ### Graph-analysis pages
213
241
 
214
242
  Included in Woods `2.0.0`: `graph_analysis` enforces its advertised default
@@ -238,6 +266,7 @@ generation. This contract is included in Woods `2.0.0`.
238
266
  | `exhausted` | `complete` | `false` | Exact count |
239
267
  | `result_limit` | `partial` | `true` | `null` (unknown) |
240
268
  | `scan_budget` or `regex_timeout` | `partial` | `null` (unknown) | `null` (unknown) |
269
+ | `unreadable_or_corrupt_source` (included in Woods 2.1) | `partial` | `null` (unknown) | `null` (unknown) |
241
270
 
242
271
  `matched_lower_bound` counts distinct observed `(type, identifier)` matches,
243
272
  including at most one lookahead match beyond `limit`. A result-limit response
@@ -245,10 +274,19 @@ therefore establishes another match; an exactly full page can instead be
245
274
  complete if the requested domain is exhausted. Deep lookahead shares
246
275
  `WOODS_SEARCH_MAX_SCAN` with the initial scan and retains round-robin scanning
247
276
  across types. Search does not count the entire omitted tail or offer pagination.
248
- The existing `types` filter and result labels name directory families:
249
- `rails_source` includes both Rails and gem source units. Deep reads accept those
250
- two stored types only in that shared directory; `lookup` and lexical retrieval
251
- retain the unit's actual `rails_source` or `gem_source` type.
277
+ **Included in Woods 2.1 (#593):** `search.types` accepts concrete unit types
278
+ such as `graphql_mutation` and `gem_source`, plus the directory-family aliases
279
+ `graphql` (all four GraphQL types) and `rails_source` (Rails and gem sources).
280
+ Aliases expand the same way with or without a package/source-path scope.
281
+ Results always carry the actual stored type, so their `(type, identifier)` can
282
+ be passed to `lookup`; an unknown type returns `invalid_params` instead of an
283
+ apparently complete empty result. This corrects the older unscoped family
284
+ labels. Retrieval tools continue to use their own documented concrete-type
285
+ filters; search aliases do not change those contracts.
286
+
287
+ **Included in Woods 2.1:** typed `lookup` also accepts the `graphql` family
288
+ alias. The returned unit keeps its concrete type, such as `graphql_mutation`;
289
+ prefer that concrete type for follow-up identity checks.
252
290
 
253
291
  All partial responses retain `partial: true` and include a narrowing `hint`.
254
292
  JSON exposes these fields; Markdown, plain text, and Claude formats label the
@@ -257,8 +295,17 @@ Narrow `types`, literal `exact_prefix`/`exact_suffix`, or deep `fields` before
257
295
  using discovery as exhaustive evidence. Completeness applies to this index and
258
296
  query domain, not to unindexed application code.
259
297
 
260
- Detected missing, unreadable, or corrupt artifacts remain `isError: true` with
261
- `_meta.error_code: "corrupt_artifact"`. Their `_meta.completeness` has
298
+ **Included in Woods 2.1:** an individual unit that unscoped search needs but
299
+ cannot decode or open is skipped. Search retains readable matches and returns
300
+ successful partial completeness with reason `unreadable_or_corrupt_source`;
301
+ an empty partial answer does not establish absence. Identifier-only matches can
302
+ use published summaries without opening unit bodies, so search is not an
303
+ artifact-integrity check. Inspect `woods_status` and run `woods:validate`.
304
+ Explicit package/source-path scope first reads the full unit set and still
305
+ returns an error if that preflight encounters a damaged body.
306
+
307
+ Damaged index-wide artifacts, such as a manifest or type index, remain
308
+ `isError: true` with `_meta.error_code: "corrupt_artifact"`. Their `_meta.completeness` has
262
309
  `status: "unknown"`, `reason: "unreadable_or_corrupt_source"`, and `null` for
263
310
  `has_more`, `total_matches`, and `matched_lower_bound`; no successful empty
264
311
  result is substituted. Inspect `woods_status` and run `woods:validate`.
@@ -272,6 +319,10 @@ references (including references to generic PORO and library classes) may have n
272
319
  edge. No dependents, a test-only dependent, or a completed traversal does not prove
273
320
  there are no production callers. Verify important absence claims in source.
274
321
 
322
+ A supporting post-2.0 writer recovers additional [constant source references](EXTRACTOR_REFERENCE.md#constant-source-references)
323
+ (included in Woods 2.1). Upgrading the reader alone cannot add relationships
324
+ to an old index. The coverage warning still applies.
325
+
275
326
  Supporting servers expose the annotated, paginated traversal result in
276
327
  `structuredContent.data` for every renderer, including the default packaged
277
328
  stdio and HTTP servers. Read `data.total_is_exact`, `data.graph_coverage`, budget
@@ -388,7 +439,31 @@ Partial traversal and pagination metadata retain the budget contract above.
388
439
 
389
440
  The Ruby server builder contains 15 additional schemas for sessions, pipeline operations, retrieval feedback, temporal snapshots, and Notion sync. They register only when their required collaborators or configuration are wired.
390
441
 
391
- The normal packaged executable does not wire pipeline-operator or feedback-store collaborators. Do not tell users to call those tools after a standard `woods-mcp` launch. Snapshot, session, and Notion capabilities are specialized configurations; document and test the exact embedded server construction when enabling them.
442
+ The normal packaged executable does not boot Rails or load application
443
+ initializers. It does not wire pipeline-operator or feedback-store collaborators.
444
+ `session_trace` needs an in-process configured store supporting `read` and
445
+ `sessions`; enabling the Rails middleware alone does not add it to a separate
446
+ Index process. `notion_sync` needs the API token and database IDs configured in
447
+ that process. These require an explicitly configured custom/embedded builder;
448
+ verify `tools/list`. For ordinary application Notion export, use
449
+ `bin/rails woods:notion_sync`.
450
+
451
+ Snapshots have a supported packaged path: an existing `woods.sqlite3` is
452
+ auto-discovered. `WOODS_SNAPSHOTS=true` enables store construction but does **not**
453
+ force JSON: bootstrap prefers SQLite and falls back to JSON only if SQLite is
454
+ unavailable or fails. It does not import old JSON history into SQLite. To read
455
+ retained JSON history after SQLite becomes available, use a custom builder with
456
+ an explicit `Woods::Temporal::JsonSnapshotStore` passed as `snapshot_store:`, or
457
+ keep a separate historical reader using that store. See the
458
+ [tool wiring table](MCP_TOOL_COOKBOOK.md#conditional-tools--wiring).
459
+
460
+ For an embedded builder that supplies `operator:`, `pipeline_extract` checks
461
+ publication after both full and incremental runs. With the Tasks extension,
462
+ a generation-marker or final source-verification failure marks the task
463
+ `failed`; it cannot report `completed` merely because extraction returned.
464
+ Clients without the extension still receive a background-start acknowledgement,
465
+ which does not establish completion. This correction (#584) is included in Woods 2.1; check the loaded server revision. See the
466
+ [publication failure recovery](TROUBLESHOOTING.md#extraction-exits-non-zero-after-could-not-publish-generation).
392
467
 
393
468
  ### HTTP transport
394
469
 
@@ -503,6 +578,10 @@ root/nested ownership, path normalization, errors, storage support, and cost.
503
578
  `"deep"` (five seconds). `index.source_freshness` describes the served generation
504
579
  as `current`, `drifted` or `unknown`; missing source/key and incomplete capture
505
580
  never count as current. No Rails initialization or provider call is needed.
581
+ **Included in Woods 2.1:** evidence above the serialized-size limit produces
582
+ `unavailable` with `source_manifest_too_large` while the code index remains
583
+ usable. Follow `inspect_source_limits` and inspect the reported byte counts;
584
+ repeating an identical full extraction cannot remove this limitation.
506
585
  See [source freshness](SOURCE_FRESHNESS.md) for scope, private-key handling and
507
586
  fresh-process extraction. Existing HEAD/dirty fields remain separate diagnostics.
508
587