woods 2.0.1 → 2.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +94 -7
- data/CONTRIBUTING.md +134 -19
- data/README.md +1 -1
- data/docs/AGENT_GUIDE.md +19 -0
- data/docs/AGENT_SETUP.md +22 -2
- data/docs/BACKEND_MATRIX.md +7 -0
- data/docs/CLIENT_HOOKS.md +6 -0
- data/docs/CONFIGURATION_REFERENCE.md +133 -25
- data/docs/CONSOLE_MCP_SETUP.md +82 -30
- data/docs/EMBEDDING_MODELS.md +16 -19
- data/docs/EXTRACTOR_REFERENCE.md +219 -21
- data/docs/FAQ.md +11 -25
- data/docs/GETTING_STARTED.md +7 -1
- data/docs/INCREMENTAL_EXTRACTION.md +261 -19
- data/docs/INDEX_LAYOUT.md +5 -0
- data/docs/INTERNALS.md +9 -0
- data/docs/MCP_HTTP_TRANSPORT.md +20 -15
- data/docs/MCP_SERVERS.md +87 -8
- data/docs/MCP_TOOL_COOKBOOK.md +13 -55
- data/docs/NOTION_INTEGRATION.md +7 -1
- data/docs/PUBLISHED_INDEX.md +6 -0
- data/docs/README.md +6 -1
- data/docs/RETRIEVAL_GUIDE.md +17 -0
- data/docs/SOURCE_FRESHNESS.md +157 -5
- data/docs/TOKEN_BENCHMARK.md +10 -18
- data/docs/TROUBLESHOOTING.md +70 -14
- data/docs/UNBLOCKED_INTEGRATION.md +60 -8
- data/docs/UPGRADING_TO_2.md +153 -38
- data/docs/WATCH_DAEMON.md +97 -14
- data/exe/woods-console-mcp +2 -2
- data/lib/generators/woods/templates/woods.rb.tt +2 -1
- data/lib/tasks/woods.rake +23 -7
- data/lib/tasks/woods_checks.rake +2 -2
- data/lib/woods/agent_configuration/cli.rb +1 -1
- data/lib/woods/agent_configuration/layout.rb +16 -2
- data/lib/woods/agent_configuration/plan.rb +13 -3
- data/lib/woods/agent_configuration/planner_validation.rb +4 -2
- data/lib/woods/agent_configuration/preflight.rb +5 -3
- data/lib/woods/builder.rb +17 -57
- data/lib/woods/cache/cache_middleware.rb +56 -30
- data/lib/woods/chunking/contributor_chunks.rb +119 -0
- data/lib/woods/chunking/semantic_chunker.rb +44 -21
- data/lib/woods/console/connection_manager.rb +56 -3
- data/lib/woods/console/embedded_executor.rb +30 -5
- data/lib/woods/console/rack_middleware.rb +29 -1
- data/lib/woods/dependency_graph.rb +34 -10
- data/lib/woods/embedding/fake.rb +12 -0
- data/lib/woods/embedding/indexer.rb +195 -98
- data/lib/woods/embedding/input_budget.rb +67 -0
- data/lib/woods/embedding/openai.rb +70 -20
- data/lib/woods/embedding/provider.rb +37 -25
- data/lib/woods/embedding/text_preparer.rb +76 -32
- data/lib/woods/embedding/token_counter.rb +18 -81
- data/lib/woods/embedding/vector_configuration.rb +48 -0
- data/lib/woods/extraction_identities.rb +175 -0
- data/lib/woods/extractor.rb +304 -107
- data/lib/woods/extractors/action_cable_extractor.rb +8 -3
- data/lib/woods/extractors/assigned_value_discovery.rb +74 -0
- data/lib/woods/extractors/class_declarations.rb +121 -0
- data/lib/woods/extractors/configuration_extractor.rb +11 -3
- data/lib/woods/extractors/declaration_ancestry.rb +92 -0
- data/lib/woods/extractors/event_extractor.rb +8 -0
- data/lib/woods/extractors/graphql_extractor.rb +134 -77
- data/lib/woods/extractors/job_extractor.rb +5 -1
- data/lib/woods/extractors/lib_extractor.rb +132 -15
- data/lib/woods/extractors/mailer_extractor.rb +3 -5
- data/lib/woods/extractors/manager_extractor.rb +7 -21
- data/lib/woods/extractors/migration_declaration.rb +87 -0
- data/lib/woods/extractors/migration_extractor.rb +5 -39
- data/lib/woods/extractors/phlex_extractor.rb +6 -2
- data/lib/woods/extractors/policy_extractor.rb +9 -5
- data/lib/woods/extractors/poro_extractor.rb +112 -53
- data/lib/woods/extractors/pundit_extractor.rb +11 -6
- data/lib/woods/extractors/scheduled_job_extractor.rb +45 -4
- data/lib/woods/extractors/serializer_extractor.rb +34 -22
- data/lib/woods/extractors/shared_utility_methods.rb +18 -1
- data/lib/woods/extractors/source_nesting.rb +142 -106
- data/lib/woods/extractors/standalone_module_discovery.rb +123 -0
- data/lib/woods/extractors/state_machine_extractor.rb +46 -40
- data/lib/woods/extractors/view_component_extractor.rb +9 -7
- data/lib/woods/flow_assembler.rb +4 -1
- data/lib/woods/generation.rb +25 -0
- data/lib/woods/hooks/context_hint.rb +7 -2
- data/lib/woods/mcp/bootstrapper.rb +33 -7
- data/lib/woods/mcp/config_resolver.rb +26 -7
- data/lib/woods/mcp/index_reader.rb +125 -24
- data/lib/woods/mcp/index_reader_pinning.rb +16 -0
- data/lib/woods/mcp/renderers/markdown_renderer.rb +7 -1
- data/lib/woods/mcp/renderers/plain_renderer.rb +3 -1
- data/lib/woods/mcp/search_results.rb +7 -1
- data/lib/woods/mcp/server.rb +24 -4
- data/lib/woods/module_reconciliation.rb +151 -0
- data/lib/woods/path_dispatcher.rb +7 -2
- data/lib/woods/rake_helpers.rb +43 -11
- data/lib/woods/release.rb +1 -1
- data/lib/woods/resilience/index_validator.rb +8 -3
- data/lib/woods/resilience/retryable_provider.rb +18 -1
- data/lib/woods/resolved_config.rb +68 -8
- data/lib/woods/retrieval/context_assembler.rb +3 -3
- data/lib/woods/retrieval/lexical_assembler.rb +3 -2
- data/lib/woods/retrieval/scope.rb +18 -2
- data/lib/woods/retrieval/source_evidence.rb +14 -2
- data/lib/woods/source_contributor_validation.rb +78 -0
- data/lib/woods/source_contributors.rb +116 -0
- data/lib/woods/source_inputs/handoff.rb +37 -0
- data/lib/woods/source_inputs/launcher.rb +53 -13
- data/lib/woods/source_inputs/manifest.rb +84 -3
- data/lib/woods/source_inputs/private_key.rb +44 -12
- data/lib/woods/source_inputs/scanner.rb +98 -27
- data/lib/woods/source_inputs/scopes.rb +1 -1
- data/lib/woods/source_inputs/session.rb +147 -15
- data/lib/woods/source_inputs/stable_reader.rb +127 -0
- data/lib/woods/source_inputs/status.rb +40 -8
- data/lib/woods/source_inputs/verifier.rb +28 -5
- data/lib/woods/source_path_encoding.rb +33 -0
- data/lib/woods/source_references/cache.rb +284 -0
- data/lib/woods/source_references/collector.rb +120 -0
- data/lib/woods/source_references/extraction.rb +185 -0
- data/lib/woods/source_references/inputs.rb +134 -0
- data/lib/woods/source_references/parser_adapter.rb +134 -0
- data/lib/woods/source_references/pass.rb +152 -0
- data/lib/woods/source_references/prism_adapter.rb +116 -0
- data/lib/woods/source_references/registry.rb +178 -0
- data/lib/woods/source_references/runtime_lookup.rb +127 -0
- data/lib/woods/source_references/value_class.rb +82 -0
- data/lib/woods/storage/metadata_store.rb +4 -1
- data/lib/woods/storage/qdrant.rb +2 -2
- data/lib/woods/unblocked/client.rb +12 -7
- data/lib/woods/unblocked/document_builder.rb +4 -1
- data/lib/woods/unblocked/exporter.rb +127 -37
- data/lib/woods/unblocked/sync_manifest.rb +137 -21
- data/lib/woods/unblocked/uri_migration.rb +105 -0
- data/lib/woods/util/host_guard.rb +3 -2
- data/lib/woods/version.rb +1 -1
- data/lib/woods/watch/catch_up.rb +138 -0
- data/lib/woods/watch/claim_lease.rb +150 -0
- data/lib/woods/watch/cli.rb +26 -2
- data/lib/woods/watch/daemon.rb +80 -59
- data/lib/woods/watch/installation/options.rb +1 -1
- data/lib/woods/watch/installation/receipt.rb +6 -1
- data/lib/woods/watch/managed_child.rb +1 -1
- data/lib/woods/watch/supervisor.rb +1 -1
- data/lib/woods/watch/tree_scan.rb +14 -2
- data/plugin/.claude-plugin/plugin.json +1 -1
- data/plugin/hooks/adapters/normalize.rb +3 -2
- data/plugin/hooks/woods-input-rules.sh +4 -0
- data/plugin/hooks/woods-refresh.sh +15 -7
- data/plugin/hooks/woods-session-start.sh +60 -3
- data/plugin/skills/woods-diagnose/SKILL.md +334 -11
- data/plugin/skills/woods-investigate/SKILL.md +11 -0
- data/plugin/skills/woods-mcp-config/SKILL.md +79 -8
- data/plugin/skills/woods-setup/SKILL.md +53 -6
- metadata +32 -5
data/docs/MCP_TOOL_COOKBOOK.md
CHANGED
|
@@ -58,11 +58,11 @@ The Index Server defines **29 schemas**: the packaged executable registers **14*
|
|
|
58
58
|
| Tool group | Count | Wiring condition |
|
|
59
59
|
|------------|-------|------------------|
|
|
60
60
|
| Always-on | 14 | Always registered, `lookup`, `search`, `dependencies`, `dependents`, `structure`, `graph_analysis`, `domain_clusters`, `pagerank`, `framework`, `recent_changes`, `reload`, `codebase_retrieve`, `trace_flow`, `woods_status` |
|
|
61
|
-
| `session_trace` | 1 | `
|
|
61
|
+
| `session_trace` | 1 | Custom/embedded Index process has a configured `session_store` supporting `read` and `sessions`; the packaged executable does not load Rails initializer configuration |
|
|
62
62
|
| Operator (5) | 5 | Custom embedded server wires an operator: `pipeline_extract`, `pipeline_embed`, `pipeline_status`, `pipeline_diagnose`, `pipeline_repair` |
|
|
63
63
|
| Feedback (4) | 4 | Custom embedded server wires a feedback store: `retrieval_rate`, `retrieval_report_gap`, `retrieval_explain`, `retrieval_suggest` |
|
|
64
|
-
| Snapshot (4) | 4 | Extraction with `enable_snapshots = true` normally creates `woods.sqlite3`, which packaged servers discover.
|
|
65
|
-
| `notion_sync` | 1 | `
|
|
64
|
+
| Snapshot (4) | 4 | Extraction with `enable_snapshots = true` normally creates `woods.sqlite3`, which packaged servers discover. `WOODS_SNAPSHOTS=true` enables construction, preferring SQLite; it does not force JSON or import JSON history. Custom builders can pass an explicit JSON `snapshot_store:`; see [store selection](MCP_SERVERS.md#conditional-index-capabilities). Internal SQLite migrations are automatic. Tools: `list_snapshots`, `snapshot_diff`, `unit_history`, `snapshot_detail` |
|
|
65
|
+
| `notion_sync` | 1 | Token and `notion_database_ids` configured inside a custom/embedded Index process; use `bin/rails woods:notion_sync` for ordinary application export |
|
|
66
66
|
|
|
67
67
|
`codebase_retrieve` is always registered (no `retrieve` alias exists). Default semantic mode requires an embedding provider and a completed `woods:embed` run. Explicit `WOODS_RETRIEVAL_MODE=lexical` ranks published extraction units without a provider or embeddings; set it in the MCP process environment and restart the server. See [embedding-free lexical retrieval](RETRIEVAL_GUIDE.md#embedding-free-lexical-retrieval).
|
|
68
68
|
|
|
@@ -887,6 +887,12 @@ Trigger extraction, then reload the server's in-memory data:
|
|
|
887
887
|
|
|
888
888
|
**What you'll get:** Lists of added, modified, and deleted units between the two git SHAs. Use `list_snapshots` first to find valid SHA values.
|
|
889
889
|
|
|
890
|
+
**Included in Woods 2.1 (#593):** use the exact stored SHAs returned by
|
|
891
|
+
`list_snapshots`; prefixes are not expanded automatically. `snapshot_diff`
|
|
892
|
+
returns `not_found` if either snapshot is absent and `invalid_params` for a
|
|
893
|
+
malformed SHA. Two available snapshots with no differences still return a
|
|
894
|
+
successful empty diff. These checks apply to both JSON and SQLite stores.
|
|
895
|
+
|
|
890
896
|
---
|
|
891
897
|
|
|
892
898
|
### "How has the User model evolved?"
|
|
@@ -908,58 +914,10 @@ Trigger extraction, then reload the server's in-memory data:
|
|
|
908
914
|
|
|
909
915
|
### GitHub Actions for Incremental Extraction
|
|
910
916
|
|
|
911
|
-
|
|
912
|
-
|
|
913
|
-
|
|
914
|
-
|
|
915
|
-
name: Update Codebase Index
|
|
916
|
-
|
|
917
|
-
on:
|
|
918
|
-
push:
|
|
919
|
-
branches: [main]
|
|
920
|
-
pull_request:
|
|
921
|
-
|
|
922
|
-
jobs:
|
|
923
|
-
index:
|
|
924
|
-
runs-on: ubuntu-latest
|
|
925
|
-
steps:
|
|
926
|
-
- uses: actions/checkout@v4
|
|
927
|
-
with:
|
|
928
|
-
fetch-depth: 2 # needed for incremental diff
|
|
929
|
-
|
|
930
|
-
- name: Set up Ruby
|
|
931
|
-
uses: ruby/setup-ruby@v1
|
|
932
|
-
with:
|
|
933
|
-
bundler-cache: true
|
|
934
|
-
|
|
935
|
-
- name: Restore index cache
|
|
936
|
-
uses: actions/cache@v4
|
|
937
|
-
with:
|
|
938
|
-
path: tmp/woods
|
|
939
|
-
key: woods-${{ github.ref }}-${{ github.sha }}
|
|
940
|
-
restore-keys: |
|
|
941
|
-
woods-${{ github.ref }}-
|
|
942
|
-
woods-
|
|
943
|
-
|
|
944
|
-
- name: Run database migrations
|
|
945
|
-
run: bundle exec rails db:migrate RAILS_ENV=test
|
|
946
|
-
|
|
947
|
-
- name: Update codebase index
|
|
948
|
-
run: bundle exec rake woods:incremental
|
|
949
|
-
env:
|
|
950
|
-
RAILS_ENV: test
|
|
951
|
-
GITHUB_BASE_REF: ${{ github.base_ref }}
|
|
952
|
-
|
|
953
|
-
- name: Validate index
|
|
954
|
-
run: bundle exec rake woods:validate
|
|
955
|
-
```
|
|
956
|
-
|
|
957
|
-
For Docker-based CI:
|
|
958
|
-
|
|
959
|
-
```yaml
|
|
960
|
-
- name: Update codebase index
|
|
961
|
-
run: docker compose exec -T app bundle exec rake woods:incremental
|
|
962
|
-
```
|
|
917
|
+
Use the [canonical incremental CI recipe](INCREMENTAL_EXTRACTION.md#github-actions-with-an-exact-baseline).
|
|
918
|
+
It fetches the actual base ref, restores an index for the exact baseline commit,
|
|
919
|
+
and runs full extraction when that cache is absent. The same guide covers
|
|
920
|
+
nested Rails applications and forwarding the selected range into Docker.
|
|
963
921
|
|
|
964
922
|
---
|
|
965
923
|
|
data/docs/NOTION_INTEGRATION.md
CHANGED
|
@@ -171,7 +171,13 @@ steps:
|
|
|
171
171
|
|
|
172
172
|
### MCP Server
|
|
173
173
|
|
|
174
|
-
|
|
174
|
+
The packaged Index Server does not load Rails initializers and does not register
|
|
175
|
+
`notion_sync` merely because the application has Notion configured. Use
|
|
176
|
+
`bin/rails woods:notion_sync` in the Rails application for the standard workflow.
|
|
177
|
+
|
|
178
|
+
A custom/embedded Index Server can register the tool when the token and database
|
|
179
|
+
IDs are configured in that server process. Confirm it appears in `tools/list`
|
|
180
|
+
before calling it:
|
|
175
181
|
|
|
176
182
|
```json
|
|
177
183
|
{
|
data/docs/PUBLISHED_INDEX.md
CHANGED
|
@@ -249,6 +249,12 @@ Retention is `WOODS_PAYLOAD_RETENTION` generations (default 3). Raise it on a CI
|
|
|
249
249
|
|
|
250
250
|
`woods:check:moved_messages` opens two retained generations through `Woods::PublishedIndex` (block form, so both readers' retention locks always release) and reports every public method name that looks like it moved from one unit to another while a `:test_coverage` edge did not follow.
|
|
251
251
|
|
|
252
|
+
The task does not boot Rails. Pass `WOODS_OUTPUT` for a custom index; an explicit
|
|
253
|
+
relative value is relative to the command's working directory. Without it, the
|
|
254
|
+
Woods 2.1 #591 fix selects `tmp/woods` beside the loaded Rakefile, including
|
|
255
|
+
`rake -f /app/Rakefile` from another directory. Earlier builds used the invoking
|
|
256
|
+
directory instead; pass an absolute output path on those builds.
|
|
257
|
+
|
|
252
258
|
Each row is a **candidate move into a unit without mapped tests**, never a proven coverage loss: the check matches on method name and kind alone, so two unrelated methods that happen to share both look identical to a real move. Reported only when the source unit was covered before and the destination is not covered after; a method that was never covered is not a regression this check owns. `WOODS_CHECK_STRICT=1` exits 1 on any finding, but the check stays heuristic either way, strict mode changes the exit code, not the confidence of a row.
|
|
253
259
|
|
|
254
260
|
```bash
|
data/docs/README.md
CHANGED
|
@@ -69,7 +69,12 @@ and stdio or Streamable HTTP endpoints directly.
|
|
|
69
69
|
|
|
70
70
|
## Maintainer material
|
|
71
71
|
|
|
72
|
-
Historical build-phase documents are not user guides. Source checkouts also contain the [MCP protocol decision record](https://github.com/lost-in-the/woods/blob/main/docs/design/MCP_2026_STRATEGY.md)
|
|
72
|
+
Historical build-phase documents are not user guides. Source checkouts also contain the [MCP protocol decision record](https://github.com/lost-in-the/woods/blob/main/docs/design/MCP_2026_STRATEGY.md) and [generated self-analysis diagrams](https://github.com/lost-in-the/woods/tree/main/docs/self-analysis); these maintainer-only paths are not packaged with the gem.
|
|
73
|
+
|
|
74
|
+
[GitHub issues](https://github.com/lost-in-the/woods/issues) track active work
|
|
75
|
+
and pending decisions. The old local backlog is retired; historical records,
|
|
76
|
+
approved decisions and issue migration links are archived under `.Codex/`.
|
|
77
|
+
See [work tracking](../CONTRIBUTING.md#work-tracking) for ownership and release labels.
|
|
73
78
|
|
|
74
79
|
Contract records that read as maintainer reference rather than guides:
|
|
75
80
|
|
data/docs/RETRIEVAL_GUIDE.md
CHANGED
|
@@ -167,6 +167,23 @@ unit still prepares no text before accepting a matching no-vector checkpoint.
|
|
|
167
167
|
Empty-vector reconciliation waits until all batches succeed, so a later provider
|
|
168
168
|
failure does not retire those old vectors.
|
|
169
169
|
|
|
170
|
+
### Prepared-input checkpoints
|
|
171
|
+
|
|
172
|
+
**Included in Woods 2.1:** checkpoint schema 2 records preparation policy and
|
|
173
|
+
ordered complete input fingerprints, including chunk/storage IDs. A change to a
|
|
174
|
+
prepared prefix or chunks re-embeds even if the source hash is unchanged; metadata
|
|
175
|
+
that does not affect prepared text does not force an embedding. Old checkpoints
|
|
176
|
+
invalidate once, so budget for one re-embed after upgrading. Complete input
|
|
177
|
+
bounds and lossless splitting are described in [Embedding Models](EMBEDDING_MODELS.md).
|
|
178
|
+
|
|
179
|
+
Successful units advance their checkpoint only after all their inputs are
|
|
180
|
+
embedded, stored and obsolete chunks reconciled. A provider failure can leave
|
|
181
|
+
earlier successful batches committed; it does not mark the failed unit complete.
|
|
182
|
+
Built-in durable stores refuse changed replacements when existing IDs cannot be
|
|
183
|
+
enumerated. Custom stores without enumeration retain adapter-owned cleanup and
|
|
184
|
+
cannot claim exact obsolete-chunk reconciliation. Keep backups for rollback;
|
|
185
|
+
older readers do not understand the new checkpoint policy.
|
|
186
|
+
|
|
170
187
|
---
|
|
171
188
|
|
|
172
189
|
## Embedding-free lexical retrieval
|
data/docs/SOURCE_FRESHNESS.md
CHANGED
|
@@ -14,6 +14,7 @@ before using it; upgrading the plugin alone does not upgrade Woods.
|
|
|
14
14
|
|---|---|---|
|
|
15
15
|
| `current` | All covered inputs match their consumer baselines, with a verified fresh boot boundary and complete checks. A dirty checkout can be current. | Use the indexed facts within the coverage below. |
|
|
16
16
|
| `drifted` | At least one captured input differs, was removed, or a relevant input was added. | Inspect the changed paths; choose a full run or a justified targeted refresh. |
|
|
17
|
+
| `unavailable` | **Included in Woods 2.1:** source evidence exceeded its serialized-size limit; the code index was published successfully. | Inspect `unavailable.size_bytes` and `unavailable.limit_bytes`; another full extraction will not reduce this size. |
|
|
17
18
|
| `unknown` | Evidence is incomplete: for example an old index, missing source/key, scan limit, opaque symlink directory, or unproved boot/consumer boundary. | Inspect `reasons`; use a deep check or fresh full capture as appropriate. |
|
|
18
19
|
|
|
19
20
|
Drift can coexist with incomplete coverage. `reasons` reports both; absence of a
|
|
@@ -22,6 +23,13 @@ capture/traversal completion, while boot and consumer qualifications remain in
|
|
|
22
23
|
`reasons`. `counts` contains total observed added/changed/removed paths; each
|
|
23
24
|
`changes` list contains at most 30 paths and `truncated` marks longer lists.
|
|
24
25
|
|
|
26
|
+
**Included in Woods 2.1:** `comparison_complete` also identifies whether the
|
|
27
|
+
previous consumer baseline can establish additions. A missing or incompatible
|
|
28
|
+
baseline stays unknown; it does not make every visible file an addition. Files
|
|
29
|
+
not reached by an incomplete scan are not reported as deleted. Matching paths
|
|
30
|
+
with different captured identities still prove drift. An unreadable entry leaves
|
|
31
|
+
coverage incomplete while scanning continues through accessible siblings.
|
|
32
|
+
|
|
25
33
|
```json
|
|
26
34
|
{"source_check":"deep"}
|
|
27
35
|
```
|
|
@@ -37,6 +45,19 @@ The result is tied to the **served** generation, including a reader holding an
|
|
|
37
45
|
older generation during a concurrent publication. It is recomputed on each call;
|
|
38
46
|
a source edit does not require an index generation change to become visible.
|
|
39
47
|
|
|
48
|
+
**Included in Woods 2.1:** `recommendations` separates recovery actions:
|
|
49
|
+
|
|
50
|
+
- `deep_check`: the quick reader reached its time limit; request one deep check.
|
|
51
|
+
- `inspect_source_scan`: inspect `verification_reasons`, permissions, source
|
|
52
|
+
mapping and scan limits; a deeper scan cannot repair unavailable inputs.
|
|
53
|
+
- `inspect_source_limits`: source freshness is unavailable because evidence exceeds its size limit; the code index remains usable.
|
|
54
|
+
- `fresh_capture`: the captured baseline, boot or consumer evidence is incomplete;
|
|
55
|
+
fix the reported cause and run the launcher in a fresh process.
|
|
56
|
+
|
|
57
|
+
Several recommendations can apply together. `verification_reasons` describes the
|
|
58
|
+
current reader scan; `reasons` also retains capture failures. A quick reader limit
|
|
59
|
+
alone does not require rebuilding the index.
|
|
60
|
+
|
|
40
61
|
## Establish a fresh baseline
|
|
41
62
|
|
|
42
63
|
Run the launcher through the application's installed bundle:
|
|
@@ -56,12 +77,37 @@ commas and newlines. Refresh accepts known extractor names. Invalid arguments
|
|
|
56
77
|
fail before extraction; the launcher propagates the child's failure or daemon
|
|
57
78
|
stand-down exit 75. Split oversized incremental batches or choose full.
|
|
58
79
|
|
|
80
|
+
For an application with a custom `config.output_dir`, **always pass the matching
|
|
81
|
+
`--output PATH` or `WOODS_OUTPUT`**. The launcher cannot evaluate Rails configuration
|
|
82
|
+
before capturing its boot inputs. In Woods 2.0.0, its implicit `tmp/woods` default
|
|
83
|
+
overrides custom configuration and can create a second index.
|
|
84
|
+
|
|
85
|
+
**Included in Woods 2.1 ([#591](https://github.com/lost-in-the/woods/issues/591)):**
|
|
86
|
+
an implicit default no longer injects `WOODS_OUTPUT` into Rails boot. The task
|
|
87
|
+
compares finalized configuration with that preboot default and refuses a mismatch
|
|
88
|
+
before extraction, naming the configured path and explicit output remedy. An
|
|
89
|
+
initializer using `ENV.fetch('WOODS_OUTPUT', custom_path)` therefore retains its
|
|
90
|
+
configured default for this check. Explicit CLI/environment output still overrides
|
|
91
|
+
configuration and binds capture, key and publication to the same directory.
|
|
92
|
+
The refused launch may already have created a private key under `tmp/woods`, but
|
|
93
|
+
publishes no generation there and preserves the configured index's generation.
|
|
94
|
+
|
|
59
95
|
The launcher captures source before a **fresh child** evaluates its Gemfile,
|
|
60
96
|
Rakefile, Rails boot and eager loading. A private one-use handoff binds the capture
|
|
61
97
|
to root, output, action, rules, nonce and parent process. It waits for the child
|
|
62
98
|
and removes the handoff when the child exits. Source is checked again before
|
|
63
|
-
publication
|
|
64
|
-
reported, rather than being silently
|
|
99
|
+
publication. In Woods 2.0.0, edits during boot/extraction retain the earlier
|
|
100
|
+
identity and are reported in the new generation, rather than being silently
|
|
101
|
+
adopted as a current baseline.
|
|
102
|
+
|
|
103
|
+
**Included in Woods 2.1:** the
|
|
104
|
+
[constant-reference writer](EXTRACTOR_REFERENCE.md#constant-source-references)
|
|
105
|
+
requires verified source for its graph evidence. Source changes after capture,
|
|
106
|
+
during boot or extraction, refuse publication while reference enrichment
|
|
107
|
+
participates. The final verification also checks covered non-Ruby inputs. The
|
|
108
|
+
preceding generation remains active instead of publishing a new drifted
|
|
109
|
+
generation. Retry against stable source in a fresh process. The existing
|
|
110
|
+
generation can still report drift through `woods_status`.
|
|
65
111
|
|
|
66
112
|
Existing Rake tasks, direct `Extractor` calls and the watch daemon remain usable.
|
|
67
113
|
Their post-boot capture is marked `unverified_boot_boundary`; current bytes alone
|
|
@@ -80,22 +126,71 @@ scanner's exclusions. Explicit source roots override generic exclusions; the
|
|
|
80
126
|
index output is always excluded. Contained file symlinks are checked for stable
|
|
81
127
|
resolution; directory symlinks and escaping/unreadable inputs leave uncertainty.
|
|
82
128
|
|
|
129
|
+
**Included in Woods 2.1 (#653):** supporting writers scan
|
|
130
|
+
directory links under their logical application paths, preserving independent
|
|
131
|
+
aliases. Static files with no input consumer do not prevent publication merely
|
|
132
|
+
because a parent directory is a symlink. Entries in an external linked directory
|
|
133
|
+
may be enumerated to determine scope, but an input file must resolve inside the
|
|
134
|
+
application root before Woods opens its contents. Package files are inputs even
|
|
135
|
+
outside `app/` and `lib/`; an external package file still refuses verification.
|
|
136
|
+
Directory cycles, links that change during scanning, unreadable entries and scan
|
|
137
|
+
limits keep evidence incomplete. Output-directory aliases are pruned. Run a fresh
|
|
138
|
+
full extraction after upgrading to establish the updated capture rules; do not
|
|
139
|
+
replace valid links or disable source verification to work around an older writer.
|
|
140
|
+
|
|
141
|
+
**Included in Woods 2.1:** source paths use UTF-8 bytes even
|
|
142
|
+
when the process locale is `C`. A filename with invalid UTF-8 bytes produces
|
|
143
|
+
`undecodable_source_path` and incomplete coverage; directory entries with those
|
|
144
|
+
names are pruned. Diagnostic labels escape the original bytes and are bounded.
|
|
145
|
+
The watcher skips these entries, and reference verification retains the preceding
|
|
146
|
+
generation. **Included in Woods 2.1:** the publication refusal names the reason
|
|
147
|
+
and up to three escaped, bounded path labels, including entries outside source
|
|
148
|
+
consumer scopes; it never includes file contents. Rename the affected entries to
|
|
149
|
+
valid UTF-8 names before a fresh capture. Explicit root/output paths with invalid UTF-8 bytes are rejected as
|
|
150
|
+
configuration errors.
|
|
151
|
+
|
|
83
152
|
Every consuming scope keeps its own identities. An events scan can reread a
|
|
84
153
|
service file while its service unit remains untouched; refreshing events does
|
|
85
154
|
not certify the retained service unit. Successful file/whole-extractor work
|
|
86
155
|
updates only its scopes, including negative results and confirmed deletion.
|
|
87
|
-
Unchanged scopes keep their earlier baseline.
|
|
88
|
-
|
|
156
|
+
Unchanged scopes keep their earlier baseline. A file can therefore remain
|
|
157
|
+
`drifted` after an incremental task or hook re-extracts it: its file-extractor
|
|
158
|
+
scope may be current while retained `runtime` evidence still refers to the old
|
|
159
|
+
source. Known differences take precedence over the accompanying uncertainty;
|
|
160
|
+
`unverified_boot_boundary` does not hide them. Inspect the reported scopes and
|
|
161
|
+
use a fresh verified full launcher run when all reflected facts need a new
|
|
162
|
+
baseline. An ordinary unverified full run cannot establish verified freshness.
|
|
163
|
+
Boot inputs advance only on a full run. Partial runtime changes retain an explicit `runtime_consumption` uncertainty
|
|
89
164
|
when Woods cannot prove every retained reflected fact was re-serialized. Named
|
|
90
165
|
framework refreshes do not certify unrelated application inputs. A handled
|
|
91
166
|
extractor error retains an explicit `extractor:<name>` uncertainty even if the
|
|
92
167
|
extractor returns an empty result. Successful consumers keep their own evidence;
|
|
93
168
|
a later full run without that failure can replace the uncertainty.
|
|
169
|
+
In the Woods 2.1 reference writer, successfully returned models
|
|
170
|
+
also retain their individual source evidence when another model fails. An
|
|
171
|
+
optional framework model with a missing table therefore does not prevent
|
|
172
|
+
unrelated incremental work; the model-extractor uncertainty remains visible.
|
|
173
|
+
If an older writer already published a baseline without that individual evidence,
|
|
174
|
+
run one full extraction after upgrading to rebuild it.
|
|
175
|
+
|
|
176
|
+
With the **Woods 2.1 reference writer**, an edited Ruby service
|
|
177
|
+
retained by an events-only refresh instead prevents publication: its current
|
|
178
|
+
source cannot certify its older runtime facts or cached reference resolution.
|
|
179
|
+
See [reference baseline recovery](INCREMENTAL_EXTRACTION.md#source-reference-baseline-and-upgrades).
|
|
180
|
+
Omitted non-reference inputs retain their earlier consumer identity. The watcher
|
|
181
|
+
keeps failed work for retry; it does not expand an incomplete batch or rebuild a
|
|
182
|
+
missing reference baseline automatically.
|
|
94
183
|
|
|
95
184
|
Custom loader source outside captured roots, missing eager-load coverage and
|
|
96
185
|
uncaptured application-owned unit paths remain unknown. This is application
|
|
97
186
|
source evidence: it does not certify external database schemas/data, remote
|
|
98
187
|
configuration, installed gem bytes, provider state or live runtime services.
|
|
188
|
+
**Included in Woods 2.1:** RubyGems installation metadata can establish that
|
|
189
|
+
loaded/unit source inside the application directory belongs to an installed gem,
|
|
190
|
+
including gems installed under `vendor/bundle`. Such files do not create false
|
|
191
|
+
application-coverage errors. A vendor-shaped directory alone is insufficient;
|
|
192
|
+
local path gems and custom loaders still need captured source roots. Explicit
|
|
193
|
+
`--source-root` declarations retain coverage even for installed gem directories.
|
|
99
194
|
|
|
100
195
|
## Containers and hooks
|
|
101
196
|
|
|
@@ -107,6 +202,16 @@ same verifier without Rails initialization or provider work. Its optional
|
|
|
107
202
|
Base64-encoded JSON transport supports `output`, `root` (an explicit reader-side
|
|
108
203
|
source mapping), and `mode` (`quick` or `deep`).
|
|
109
204
|
|
|
205
|
+
**Included in Woods 2.1:** results expose `recorded_root` (writer location),
|
|
206
|
+
`checked_root` (the directory actually scanned), and `root_source` (`recorded`,
|
|
207
|
+
`explicit`, or `working_directory`). `current` applies only to that checked root.
|
|
208
|
+
The MCP IndexReader keeps using the recorded root; it does not infer a checkout
|
|
209
|
+
from the index's location. For a copied index, a current result about the original
|
|
210
|
+
root says nothing about edits in the copy. Run `woods:source_status` in the copy,
|
|
211
|
+
or provide an explicit `root` mapping. The task defaults to its process working
|
|
212
|
+
directory, including inside containers; launch it from the application root, not
|
|
213
|
+
an unrelated directory or a monorepo parent. Missing source remains unknown.
|
|
214
|
+
|
|
110
215
|
The opt-in SessionStart hook uses `WOODS_HOOK_RAKE` and `WOODS_OUTPUT`, including a
|
|
111
216
|
Docker command prefix without requiring a host application bundle. It prints
|
|
112
217
|
an actionable drift or unknown warning and stays quiet for current evidence.
|
|
@@ -114,6 +219,10 @@ Its ten-second process deadline includes command startup; the scan uses quick
|
|
|
114
219
|
mode. Missing older tasks and failed/timed-out commands report unknown. Cancelling
|
|
115
220
|
Docker exec does not itself prove the process inside the container stopped.
|
|
116
221
|
A quiet session hook does not acknowledge deferred PostToolUse queue entries.
|
|
222
|
+
**Included in Woods 2.1:** unknown warnings use the returned recommendations,
|
|
223
|
+
including every applicable recovery action. A quick reader timeout advises a deep
|
|
224
|
+
check without requiring a rebuild. Older tasks without recommendations retain
|
|
225
|
+
the conservative generic unknown warning.
|
|
117
226
|
|
|
118
227
|
## Artifact and cost
|
|
119
228
|
|
|
@@ -128,11 +237,53 @@ unknown; Woods does not silently repair permissions or rotate keys. Independent
|
|
|
128
237
|
outputs have different identities and must be compared using their own keys and
|
|
129
238
|
consumer semantics, not raw manifest equality.
|
|
130
239
|
|
|
240
|
+
### Identity-key recovery
|
|
241
|
+
|
|
242
|
+
Verified launching refuses an unavailable, insecure or malformed
|
|
243
|
+
`<output>/.source-inputs.key`; it cannot establish verified capture without that
|
|
244
|
+
key. In Woods 2.1 diagnostics, the error names the path and safe
|
|
245
|
+
file requirements while reader reason codes remain unchanged. No key bytes are
|
|
246
|
+
printed and Woods does not chmod, chown, replace or rotate the file automatically.
|
|
247
|
+
|
|
248
|
+
Inspect the file and mount from the application environment. It must be a regular
|
|
249
|
+
file, not a symlink/FIFO, contain exactly 32 bytes, belong to the process's UID,
|
|
250
|
+
and grant no group/other permissions (normally mode `0600`). A host/container UID
|
|
251
|
+
mismatch requires correcting the selected runtime user or deliberately repairing
|
|
252
|
+
ownership of the known original key. Confirm the file belongs to this index before
|
|
253
|
+
changing permissions. Do not expose the key in logs or copy it into payloads.
|
|
254
|
+
|
|
255
|
+
If the original key is lost, replaced or cannot be trusted, retain the previous
|
|
256
|
+
index and establish a fresh full baseline in a new empty output directory with
|
|
257
|
+
the intended application user. Configure the writer and readers together.
|
|
258
|
+
Do not replace a key and assume the old generation now has valid source evidence.
|
|
259
|
+
|
|
131
260
|
Failed/no-op extraction does not advance the artifact's published generation.
|
|
132
261
|
Flat fallback and older indexes lack verified atomic source evidence.
|
|
133
262
|
`WOODS_PROFILE=1` reports `source capture` and `source verification` separately.
|
|
134
263
|
Capture/recheck each allow up to ten seconds with the same file/byte caps.
|
|
135
264
|
|
|
265
|
+
**Included in Woods 2.1:** the published manifest and private launcher handoff
|
|
266
|
+
have a 16 MiB serialized-size limit. If size alone exceeds that limit, extraction
|
|
267
|
+
continues and publishes the code index successfully with a loud warning and
|
|
268
|
+
bounded evidence: `state: "unavailable"`, reason `source_manifest_too_large`, and
|
|
269
|
+
`unavailable.size_bytes` / `unavailable.limit_bytes`. A launcher handoff above the
|
|
270
|
+
limit continues in a fresh child without claiming verified preboot capture.
|
|
271
|
+
Status, validation and hooks retain this limitation; they do not recommend an
|
|
272
|
+
identical full rebuild. Source freshness remains unavailable on subsequent
|
|
273
|
+
incremental runs until a full capture fits the limit.
|
|
274
|
+
|
|
275
|
+
Normal changed-file and Git-based incremental work continues. Unchanged
|
|
276
|
+
source-reference facts require both unchanged keyed source identities and exact
|
|
277
|
+
typed ownership in the validated, hash-bound reference cache from the same
|
|
278
|
+
published generation. Missing, corrupt or mismatched proof still refuses reuse.
|
|
279
|
+
The watcher retains a compact captured-tree fingerprint for catch-up only; that
|
|
280
|
+
fingerprint never establishes runtime freshness. Invalid evidence, failed writes
|
|
281
|
+
and source instability retain their existing publication safeguards.
|
|
282
|
+
|
|
283
|
+
Version-1 manifests remain readable. The optional `comparison_complete` field
|
|
284
|
+
supplements existing coverage errors; an unavailable manifest never claims
|
|
285
|
+
complete comparison coverage.
|
|
286
|
+
|
|
136
287
|
September 2026 fixture measurements: a pinned Writebook source tree (456 visited
|
|
137
288
|
files, 223 hashed, 233KB) completed quick scans in median 36ms native / 45ms on a
|
|
138
289
|
Linux container bind mount. Discourse (26,133 visited, 6,209 hashed, 56.4MB) needed
|
|
@@ -140,4 +291,5 @@ about 1.6–2 seconds; quick scans returned unknown, while five-second scans com
|
|
|
140
291
|
traversal and still reported an opaque directory symlink. Native Ruby 4.0 and
|
|
141
292
|
container Ruby 3.4 differed, so these are a budget envelope, not a filesystem
|
|
142
293
|
speed comparison or a macOS virtiofs benchmark. Measure your own application;
|
|
143
|
-
a large or slow source tree may need a
|
|
294
|
+
a large or slow source tree may need a deep check. Rebuilding a verified baseline
|
|
295
|
+
does not remove the reader's time, file or byte limits.
|
data/docs/TOKEN_BENCHMARK.md
CHANGED
|
@@ -1,20 +1,12 @@
|
|
|
1
1
|
# Token Estimation Benchmark
|
|
2
2
|
|
|
3
|
-
> **
|
|
4
|
-
>
|
|
5
|
-
>
|
|
6
|
-
>
|
|
7
|
-
>
|
|
8
|
-
>
|
|
9
|
-
>
|
|
10
|
-
> a per-string ratio, since it measures a different thing (per-chunk average
|
|
11
|
-
> vs. per-string chars/token).
|
|
12
|
-
>
|
|
13
|
-
> When the optional [`tokenizers`](https://github.com/ankane/tokenizers-ruby)
|
|
14
|
-
> gem is installed, the Ollama path uses the real BERT WordPiece tokenizer
|
|
15
|
-
> (`Woods::Embedding::TokenCounter`) instead of this heuristic. The 4.0
|
|
16
|
-
> divisor applies to the OpenAI/default path; Ollama falls back to 1.5 when
|
|
17
|
-
> its tokenizer is unavailable.
|
|
3
|
+
> **Historical sizing evidence, not an input-limit guarantee.**
|
|
4
|
+
> `Woods::TokenUtils.chars_per_token_for(provider)` supplies character estimates
|
|
5
|
+
> for retrieval assembly and initial sizing. Woods 2.1 uses
|
|
6
|
+
> a conservative UTF-8 byte bound for known OpenAI embedding models and honest
|
|
7
|
+
> estimates for Ollama/custom models. Prefixes count toward admission; oversized
|
|
8
|
+
> inputs are split without dropping source. The optional tokenizer gem no longer
|
|
9
|
+
> downloads or implicitly selects BERT. See [Embedding Models](EMBEDDING_MODELS.md).
|
|
18
10
|
|
|
19
11
|
This is a historical record of the benchmark that picked 4.0 over the
|
|
20
12
|
original 3.5 divisor. It is cited from five places in `lib/` as the evidence
|
|
@@ -58,9 +50,9 @@ current definition and `docs/EMBEDDING_MODELS.md` for the Ollama-side ratio.
|
|
|
58
50
|
**tiktoken_ruby was deliberately not added as a runtime dependency.** A 10.6%
|
|
59
51
|
mean error is acceptable for chunking decisions, budget estimates, and
|
|
60
52
|
truncation; a native-extension dependency for marginal accuracy gains wasn't
|
|
61
|
-
worth it.
|
|
62
|
-
|
|
63
|
-
|
|
53
|
+
worth it. An explicitly injected local tokenizer must match the target model. A conservative
|
|
54
|
+
upper bound can also establish admission for a known encoding; a character
|
|
55
|
+
average cannot.
|
|
64
56
|
|
|
65
57
|
## Reproducing this benchmark
|
|
66
58
|
|
data/docs/TROUBLESHOOTING.md
CHANGED
|
@@ -32,6 +32,17 @@ This guide covers the most common problems encountered when installing, extracti
|
|
|
32
32
|
| Tool returns `error_code: :not_configured` | Feature flag or credential not set | Check `config_key` in `_meta` and the linked `doc_link` |
|
|
33
33
|
| Tool returns `error_code: :rate_limited` | `PipelineGuard` 5-min cooldown hit | Wait `retry_after_seconds` from `_meta`, then retry |
|
|
34
34
|
|
|
35
|
+
### Source-reference baseline needs a full extraction
|
|
36
|
+
|
|
37
|
+
**Included in Woods 2.1.** The related initial diagnostic is
|
|
38
|
+
`Source-reference baseline is missing or incompatible`. Both mean the writer
|
|
39
|
+
cannot safely reuse its reference cache or verify the source consumed by retained
|
|
40
|
+
units. It can follow an older-index upgrade, missing cache artifacts, or source
|
|
41
|
+
changes omitted from the refresh batch. Confirm the loaded gem revision and
|
|
42
|
+
output directory, then follow the [full baseline rebuild](INCREMENTAL_EXTRACTION.md#source-reference-baseline-and-upgrades).
|
|
43
|
+
The failed operation leaves the published generation unchanged. Preserve pending
|
|
44
|
+
watcher work; increasing traversal budgets cannot repair extraction coverage.
|
|
45
|
+
|
|
35
46
|
### First-Pass Diagnostics
|
|
36
47
|
|
|
37
48
|
For a single-call health snapshot, call the Index Server's `woods_status` tool. It reports:
|
|
@@ -67,6 +78,12 @@ the selected owner or installed command and restarting that owner; do not delete
|
|
|
67
78
|
claim files or kill PIDs taken from status. No index-visible record exists before
|
|
68
79
|
the first boot resolves the application's output directory.
|
|
69
80
|
|
|
81
|
+
For an abandoned foreign-container claim, follow
|
|
82
|
+
[ownership-verified claim recovery](WATCH_DAEMON.md#recovering-an-abandoned-managed-claim).
|
|
83
|
+
The Woods 2.1 command requires the exact token and a free lifetime
|
|
84
|
+
lease; old claims require the documented legacy recovery. Age is not proof of
|
|
85
|
+
abandonment. `woods:clean` preserves ownership sidecars and does not reset them.
|
|
86
|
+
|
|
70
87
|
Unset `WOODS_WATCH_IDLE_TIMEOUT` in managed modes. If the boot deadline is reached,
|
|
71
88
|
diagnose Bundler/initializer startup before increasing `--boot-timeout`; a valid
|
|
72
89
|
long extraction has a separate readiness state and is not bounded by that clock.
|
|
@@ -74,6 +91,10 @@ If setup created a Procfile but normal `bin/dev` still only launches Rails, choo
|
|
|
74
91
|
Puma or explicitly run the selected Foreman command. The generator never rewrites
|
|
75
92
|
`bin/dev` or starts services during preview.
|
|
76
93
|
|
|
94
|
+
An empty idle-timeout variable crashes older tasks even though managed validation
|
|
95
|
+
accepts it. Unset it on those builds; the Woods 2.1 #591 fix consistently treats
|
|
96
|
+
empty or whitespace-only values as unset.
|
|
97
|
+
|
|
77
98
|
If installation reports a pending transaction, use `woods:watch --operation
|
|
78
99
|
recover` through the Rails generator, initially with `--pretend`; see
|
|
79
100
|
[owned setup recovery](WATCH_DAEMON.md#ownership-updates-and-removal). That Rails
|
|
@@ -415,7 +436,7 @@ full extraction when current history and provenance are required.
|
|
|
415
436
|
|
|
416
437
|
**Cause:** The range git was asked to diff does not resolve in the checkout: a GitLab `CI_COMMIT_BEFORE_SHA` of all zeros (new branch), a GitHub base ref that was never fetched, a shallow clone with no `HEAD~1`, or a typo'd revision. This used to read as "no relevant files changed" and the task exited 0 while the sync never ran; it now fails closed, because a green job hiding a skipped sync lets index drift grow unbounded. The one stand-down: a *running* watch daemon maintaining the same index exits 0 with a printed reason, since its start-up catch-up covers the changes. A degraded daemon covers nothing and still exits 1.
|
|
417
438
|
|
|
418
|
-
**Fix:**
|
|
439
|
+
**Fix:** Fetch the actual base ref and enough history to resolve both endpoints (for example `fetch-depth: 0` with an explicit base-ref fetch in GitHub Actions), or correct the CI environment variables. A depth of two does not establish an arbitrary CI base. Set `CHANGED_FILES` explicitly to bypass range resolution, or run a full extraction:
|
|
419
440
|
|
|
420
441
|
```bash
|
|
421
442
|
bundle exec rake woods:extract
|
|
@@ -461,6 +482,31 @@ the unreachable payload. A resident `woods:watch` process handles the same
|
|
|
461
482
|
failure differently: it reports `degraded`, carries the changed paths, and
|
|
462
483
|
retries after a later filesystem event.
|
|
463
484
|
|
|
485
|
+
For the optional embedded `pipeline_extract` tool, a client using the Tasks
|
|
486
|
+
extension sees the task become `failed` for this publication refusal on a
|
|
487
|
+
revision containing #584 (included in Woods 2.1). A background-start response
|
|
488
|
+
alone does not mean extraction completed. The packaged Index Server does not
|
|
489
|
+
register this tool.
|
|
490
|
+
|
|
491
|
+
### Incremental extraction or refresh reports "Extraction failed for ..."
|
|
492
|
+
|
|
493
|
+
A selected whole-app extractor raised before returning a complete result, or
|
|
494
|
+
its initialization failed. In Woods 2.1, a successful sibling extractor cannot
|
|
495
|
+
turn that failed batch into a successful publication (#584). Readers retain the prior generation. Inspect the
|
|
496
|
+
earlier log line naming the failed extractor, fix its cause, then retry the
|
|
497
|
+
complete changed-file list or refresh selection. Preserve the published index;
|
|
498
|
+
deleting it does not repair the failing extractor.
|
|
499
|
+
|
|
500
|
+
### Incremental extraction reports "Restart-sensitive inputs"
|
|
501
|
+
|
|
502
|
+
On revisions containing #588 (included in Woods 2.1), the direct incremental
|
|
503
|
+
API and optional embedded pipeline refuse schema or boot-configuration inputs.
|
|
504
|
+
Apply required migrations, then run `bin/rails woods:extract` in a fresh Rails
|
|
505
|
+
process. Reusing an embedded server's old Rails configuration or schema cache
|
|
506
|
+
cannot establish current runtime facts. The fresh one-shot `woods:incremental`
|
|
507
|
+
task selects a full run automatically for these inputs. See the
|
|
508
|
+
[runtime input contract](INCREMENTAL_EXTRACTION.md#what-a-change-actually-requires-reload-restart-or-neither).
|
|
509
|
+
|
|
464
510
|
---
|
|
465
511
|
|
|
466
512
|
### `manifest.json` shows the wrong branch (or `git_branch: "unknown"`) in a worktree
|
|
@@ -737,6 +783,15 @@ Woods detects the dimension mismatch and raises `Woods::MCP::DimensionMismatch`
|
|
|
737
783
|
|
|
738
784
|
**A dimension mismatch is never silently tolerated.** If you are getting poor results without seeing this error, the cause is something else.
|
|
739
785
|
|
|
786
|
+
For an unsupported `dimensions` request or a wrong-width cached vector, compare
|
|
787
|
+
the installed embedding and reader revisions as well as the model, endpoint,
|
|
788
|
+
and explicit width configuration. The request/cache consistency fix (#586) is
|
|
789
|
+
included in Woods 2.1: it separates stored widths from requested reductions,
|
|
790
|
+
keeps fixed-width ada requests compatible, and separates embedding cache entries
|
|
791
|
+
by provider configuration. See [embedding options](CONFIGURATION_REFERENCE.md#embedding-options)
|
|
792
|
+
and [cache identity](CONFIGURATION_REFERENCE.md#retrieval-cache-options). Do not
|
|
793
|
+
remove a width guard to accept mismatched vectors.
|
|
794
|
+
|
|
740
795
|
---
|
|
741
796
|
|
|
742
797
|
### OpenAI API errors during embedding
|
|
@@ -786,20 +841,16 @@ config.embedding_options = { host: 'http://localhost:11434' }
|
|
|
786
841
|
|
|
787
842
|
**Symptom:** `rake woods:embed` fails with `Ollama API error: 400 {"error":"the input length exceeds the context length"}`. Individual chunks may look smaller than the configured `num_ctx`.
|
|
788
843
|
|
|
789
|
-
**Cause:**
|
|
790
|
-
|
|
791
|
-
|
|
792
|
-
|
|
793
|
-
```ruby
|
|
794
|
-
# Gemfile
|
|
795
|
-
gem 'woods', '~> 2.0'
|
|
796
|
-
gem 'tokenizers', '~> 0.5' # exact BERT WordPiece token counting
|
|
797
|
-
```
|
|
844
|
+
**Cause:** the server enforces the selected model's actual context window.
|
|
845
|
+
Character estimates can undercount dense Ruby or multilingual source; a larger
|
|
846
|
+
`num_ctx` setting does not establish that the model accepts a larger input.
|
|
798
847
|
|
|
799
|
-
Woods
|
|
800
|
-
|
|
801
|
-
|
|
802
|
-
|
|
848
|
+
**Diagnosis:** record the loaded Woods revision, model and configured context.
|
|
849
|
+
In supporting builds after 2.0.0, Woods counts the full metadata prefix plus
|
|
850
|
+
source, splits without truncating source, and requests Ollama `truncate: false`.
|
|
851
|
+
An overflow is a visible refusal; read the unit/model/limit diagnostic rather
|
|
852
|
+
than enabling truncation or installing an unrelated BERT tokenizer. See the
|
|
853
|
+
[input counting contract](EMBEDDING_MODELS.md#why-num_ctx-isnt-enough).
|
|
803
854
|
|
|
804
855
|
If you want fewer chunks per unit and have the disk space, switch to a larger-context model:
|
|
805
856
|
|
|
@@ -1046,3 +1097,8 @@ source root/private key, a quick scan limit and an unverified boot boundary are
|
|
|
1046
1097
|
different causes. Try `source_check: "deep"` for a budget limit; use the fresh
|
|
1047
1098
|
launcher for a new verified baseline. Do not delete pending hook events or alter
|
|
1048
1099
|
key permissions just to suppress a warning. See [source freshness](SOURCE_FRESHNESS.md).
|
|
1100
|
+
|
|
1101
|
+
If `woods-extract` refuses an identity key, use
|
|
1102
|
+
[identity-key recovery](SOURCE_FRESHNESS.md#identity-key-recovery). If it reports a
|
|
1103
|
+
configured-output mismatch, rerun with an explicit matching `--output` or
|
|
1104
|
+
`WOODS_OUTPUT`; do not bypass the check or move a capture/key to another index.
|