woods 2.0.0.beta2 → 2.0.0.beta3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +262 -1
- data/CONTRIBUTING.md +173 -9
- data/README.md +7 -3
- data/SECURITY.md +9 -6
- data/docs/AGENT_GUIDE.md +83 -4
- data/docs/AGENT_SETUP.md +82 -1
- data/docs/BACKEND_MATRIX.md +20 -0
- data/docs/CLIENT_HOOKS.md +111 -0
- data/docs/CONFIGURATION_REFERENCE.md +199 -14
- data/docs/CONSOLE_MCP_SETUP.md +35 -5
- data/docs/DOCKER_SETUP.md +21 -2
- data/docs/EVALUATION.md +464 -1
- data/docs/EXTRACTOR_REFERENCE.md +36 -5
- data/docs/FAQ.md +11 -12
- data/docs/GETTING_STARTED.md +17 -5
- data/docs/INCREMENTAL_EXTRACTION.md +117 -1
- data/docs/INDEX_LAYOUT.md +382 -0
- data/docs/INTERNALS.md +7 -2
- data/docs/MCP_SERVERS.md +221 -5
- data/docs/MCP_TOOL_COOKBOOK.md +33 -18
- data/docs/NOTION_INTEGRATION.md +13 -0
- data/docs/OBSIDIAN_INTEGRATION.md +57 -9
- data/docs/PUBLISHED_INDEX.md +55 -0
- data/docs/README.md +7 -0
- data/docs/RETRIEVAL_GUIDE.md +253 -11
- data/docs/RUNTIME_TRACING.md +71 -0
- data/docs/SOURCE_FRESHNESS.md +143 -0
- data/docs/TROUBLESHOOTING.md +117 -5
- data/docs/UNBLOCKED_INTEGRATION.md +25 -0
- data/docs/UPGRADING_TO_2.md +44 -22
- data/docs/WATCH_DAEMON.md +259 -59
- data/exe/woods-agent-config +6 -0
- data/exe/woods-extract +5 -0
- data/exe/woods-hook-context +6 -0
- data/lib/generators/woods/templates/woods.rb.tt +1 -3
- data/lib/tasks/woods.rake +47 -397
- data/lib/woods/agent_configuration/applier.rb +133 -0
- data/lib/woods/agent_configuration/cli.rb +101 -0
- data/lib/woods/agent_configuration/cli_options.rb +29 -0
- data/lib/woods/agent_configuration/document.rb +105 -0
- data/lib/woods/agent_configuration/error.rb +7 -0
- data/lib/woods/agent_configuration/launcher.rb +75 -0
- data/lib/woods/agent_configuration/layout.rb +59 -0
- data/lib/woods/agent_configuration/managed_section.rb +62 -0
- data/lib/woods/agent_configuration/plan.rb +98 -0
- data/lib/woods/agent_configuration/plan_diff.rb +38 -0
- data/lib/woods/agent_configuration/planned_files.rb +61 -0
- data/lib/woods/agent_configuration/planner.rb +63 -0
- data/lib/woods/agent_configuration/planner_validation.rb +77 -0
- data/lib/woods/agent_configuration/preflight.rb +100 -0
- data/lib/woods/agent_configuration/recovery.rb +49 -0
- data/lib/woods/ast/node.rb +2 -0
- data/lib/woods/ast/parser.rb +38 -5
- data/lib/woods/builder.rb +21 -5
- data/lib/woods/cache/cache_middleware.rb +28 -7
- data/lib/woods/cache/cache_store.rb +4 -5
- data/lib/woods/change_set.rb +5 -4
- data/lib/woods/console/credential_index.rb +20 -2
- data/lib/woods/console/credential_scanner.rb +14 -14
- data/lib/woods/console/credential_scanner_registry.rb +36 -0
- data/lib/woods/console/embedded_executor.rb +1 -1
- data/lib/woods/console/encrypted_credential_snapshot.rb +16 -0
- data/lib/woods/console/rack_middleware.rb +22 -13
- data/lib/woods/console/server.rb +18 -16
- data/lib/woods/dependency_graph.rb +65 -13
- data/lib/woods/embedding/corpus.rb +94 -0
- data/lib/woods/embedding/indexer.rb +90 -46
- data/lib/woods/embedding/openai.rb +17 -6
- data/lib/woods/evaluation/ablation_executor.rb +6 -1
- data/lib/woods/evaluation/ablation_timed_executor.rb +22 -4
- data/lib/woods/export/typed_reader.rb +56 -0
- data/lib/woods/extractor.rb +232 -137
- data/lib/woods/extractors/action_cable_extractor.rb +3 -1
- data/lib/woods/extractors/behavioral_profile.rb +9 -7
- data/lib/woods/extractors/caching_extractor.rb +3 -1
- data/lib/woods/extractors/concern_extractor.rb +64 -6
- data/lib/woods/extractors/configuration_extractor.rb +7 -3
- data/lib/woods/extractors/controller_extractor.rb +13 -4
- data/lib/woods/extractors/database_view_extractor.rb +3 -1
- data/lib/woods/extractors/decorator_extractor.rb +3 -1
- data/lib/woods/extractors/engine_extractor.rb +3 -1
- data/lib/woods/extractors/event_extractor.rb +4 -2
- data/lib/woods/extractors/factory_extractor.rb +3 -1
- data/lib/woods/extractors/graphql_extractor.rb +8 -2
- data/lib/woods/extractors/i18n_extractor.rb +3 -1
- data/lib/woods/extractors/job_extractor.rb +6 -19
- data/lib/woods/extractors/lib_extractor.rb +3 -1
- data/lib/woods/extractors/mailer_extractor.rb +20 -5
- data/lib/woods/extractors/manager_extractor.rb +3 -1
- data/lib/woods/extractors/method_parameters.rb +53 -0
- data/lib/woods/extractors/middleware_argument.rb +65 -0
- data/lib/woods/extractors/middleware_extractor.rb +9 -3
- data/lib/woods/extractors/migration_extractor.rb +3 -1
- data/lib/woods/extractors/model_extractor.rb +39 -33
- data/lib/woods/extractors/package_extractor.rb +24 -4
- data/lib/woods/extractors/phlex_extractor.rb +3 -1
- data/lib/woods/extractors/policy_extractor.rb +3 -1
- data/lib/woods/extractors/poro_extractor.rb +3 -1
- data/lib/woods/extractors/pundit_extractor.rb +3 -1
- data/lib/woods/extractors/rails_source_extractor.rb +4 -2
- data/lib/woods/extractors/rake_task_extractor.rb +4 -2
- data/lib/woods/extractors/route_extractor.rb +3 -1
- data/lib/woods/extractors/route_helper_resolver.rb +10 -33
- data/lib/woods/extractors/scheduled_job_extractor.rb +41 -15
- data/lib/woods/extractors/serializer_extractor.rb +4 -2
- data/lib/woods/extractors/service_extractor.rb +3 -1
- data/lib/woods/extractors/shared_dependency_scanner.rb +2 -2
- data/lib/woods/extractors/shared_utility_methods.rb +27 -15
- data/lib/woods/extractors/source_nesting.rb +1 -1
- data/lib/woods/extractors/state_machine_extractor.rb +3 -1
- data/lib/woods/extractors/test_mapping_extractor.rb +3 -1
- data/lib/woods/extractors/validator_extractor.rb +3 -1
- data/lib/woods/extractors/view_component_extractor.rb +3 -1
- data/lib/woods/extractors/view_template_extractor.rb +3 -1
- data/lib/woods/gem_mapper.rb +2 -0
- data/lib/woods/git_history.rb +116 -0
- data/lib/woods/graph_analyzer.rb +35 -6
- data/lib/woods/hooks/context_cli.rb +54 -0
- data/lib/woods/hooks/context_event.rb +88 -0
- data/lib/woods/hooks/context_hint.rb +73 -0
- data/lib/woods/hooks/context_impact.rb +77 -0
- data/lib/woods/hooks/context_output.rb +47 -0
- data/lib/woods/hooks/context_state.rb +102 -0
- data/lib/woods/hooks/refresh.rb +79 -0
- data/lib/woods/hooks/rule_projection.rb +78 -0
- data/lib/woods/input_rules.rb +19 -0
- data/lib/woods/mcp/bearer_auth.rb +20 -12
- data/lib/woods/mcp/bootstrapper.rb +62 -0
- data/lib/woods/mcp/index_reader.rb +323 -160
- data/lib/woods/mcp/initialization_guidance.rb +27 -0
- data/lib/woods/mcp/origin_guard.rb +17 -9
- data/lib/woods/mcp/published_lexical_retriever.rb +115 -0
- data/lib/woods/mcp/renderers/markdown_renderer.rb +8 -1
- data/lib/woods/mcp/renderers/plain_renderer.rb +7 -1
- data/lib/woods/mcp/search_results.rb +74 -0
- data/lib/woods/mcp/server.rb +158 -37
- data/lib/woods/mcp/tool_contract.rb +2 -0
- data/lib/woods/mcp/tool_response_renderer.rb +25 -0
- data/lib/woods/mcp/traversal_evidence.rb +113 -0
- data/lib/woods/mcp/traversal_evidence_index.rb +100 -0
- data/lib/woods/mcp/traversal_evidence_page.rb +41 -0
- data/lib/woods/mcp/traversal_evidence_text.rb +52 -0
- data/lib/woods/notion/exporter.rb +56 -17
- data/lib/woods/obsidian/destination_plan.rb +98 -0
- data/lib/woods/obsidian/name_mapper.rb +19 -3
- data/lib/woods/obsidian/note_builder.rb +19 -10
- data/lib/woods/obsidian/vault_exporter.rb +88 -32
- data/lib/woods/operator/pipeline_guard.rb +18 -13
- data/lib/woods/path_dispatcher.rb +7 -1
- data/lib/woods/payload_store.rb +27 -26
- data/lib/woods/railtie.rb +3 -3
- data/lib/woods/railtie_support.rb +12 -12
- data/lib/woods/rake_helpers.rb +392 -0
- data/lib/woods/resilience/graph_invariant_validator/membership_checks.rb +71 -0
- data/lib/woods/resilience/graph_invariant_validator/node_checks.rb +61 -0
- data/lib/woods/resilience/graph_invariant_validator/reverse_relationship_checks.rb +46 -0
- data/lib/woods/resilience/graph_invariant_validator.rb +119 -0
- data/lib/woods/resilience/index_validator/graph_checks.rb +80 -0
- data/lib/woods/resilience/index_validator.rb +112 -23
- data/lib/woods/retrieval/context_assembler.rb +50 -15
- data/lib/woods/retrieval/lexical_assembler.rb +73 -0
- data/lib/woods/retrieval/lexical_index.rb +119 -0
- data/lib/woods/retrieval/ranker.rb +4 -2
- data/lib/woods/retrieval/scope.rb +108 -0
- data/lib/woods/retrieval/scoped_graph_store.rb +32 -0
- data/lib/woods/retrieval/scoped_vector_store.rb +55 -0
- data/lib/woods/retrieval/search_executor.rb +86 -27
- data/lib/woods/retrieval/source_evidence.rb +200 -0
- data/lib/woods/retriever.rb +98 -22
- data/lib/woods/ruby_analyzer/trace_enricher.rb +77 -38
- data/lib/woods/session_tracer/middleware.rb +10 -12
- data/lib/woods/session_tracer/redis_store.rb +22 -6
- data/lib/woods/session_tracer/session_flow_assembler.rb +23 -17
- data/lib/woods/session_tracer/solid_cache_coordination.rb +6 -4
- data/lib/woods/session_tracer/unit_resolver.rb +63 -0
- data/lib/woods/source_inputs/consumer_errors.rb +27 -0
- data/lib/woods/source_inputs/handoff.rb +102 -0
- data/lib/woods/source_inputs/launcher.rb +157 -0
- data/lib/woods/source_inputs/manifest.rb +124 -0
- data/lib/woods/source_inputs/private_key.rb +55 -0
- data/lib/woods/source_inputs/scanner.rb +171 -0
- data/lib/woods/source_inputs/scopes.rb +71 -0
- data/lib/woods/source_inputs/session.rb +214 -0
- data/lib/woods/source_inputs/status.rb +84 -0
- data/lib/woods/source_inputs/verifier.rb +107 -0
- data/lib/woods/storage/metadata_store.rb +25 -25
- data/lib/woods/storage/pgvector.rb +29 -8
- data/lib/woods/storage/qdrant.rb +17 -7
- data/lib/woods/storage/vector_store.rb +18 -6
- data/lib/woods/tasks.rb +3 -2
- data/lib/woods/temporal/json_snapshot_store.rb +29 -8
- data/lib/woods/unblocked/exporter.rb +59 -70
- data/lib/woods/version.rb +1 -1
- data/lib/woods/watch/boot_snapshot.rb +52 -0
- data/lib/woods/watch/daemon.rb +136 -28
- data/lib/woods/watch/listen_watcher.rb +4 -0
- data/lib/woods/watch/polling_watcher.rb +5 -1
- data/lib/woods/watch/status.rb +20 -15
- data/lib/woods/watch/tree_scan.rb +21 -13
- data/lib/woods/watch/watcher.rb +4 -1
- data/lib/woods.rb +50 -11
- data/plugin/.claude-plugin/plugin.json +1 -1
- data/plugin/hooks/adapters/normalize.jq +15 -0
- data/plugin/hooks/adapters/normalize.rb +63 -0
- data/plugin/hooks/hooks.json +20 -0
- data/plugin/hooks/woods-context.sh +50 -0
- data/plugin/hooks/woods-input-rules.sh +159 -0
- data/plugin/hooks/woods-opencode.mjs +65 -0
- data/plugin/hooks/woods-post-edit.sh +2 -225
- data/plugin/hooks/woods-refresh.sh +260 -0
- data/plugin/hooks/woods-session-start.sh +47 -55
- data/plugin/skills/woods-agent-enable/SKILL.md +13 -0
- data/plugin/skills/woods-diagnose/SKILL.md +288 -1
- data/plugin/skills/woods-investigate/SKILL.md +106 -0
- data/plugin/skills/woods-mcp-config/SKILL.md +89 -1
- data/plugin/skills/woods-setup/SKILL.md +107 -6
- metadata +84 -5
data/docs/GETTING_STARTED.md
CHANGED
|
@@ -12,7 +12,19 @@ If an agent will perform the installation, use the safety and handoff checklist
|
|
|
12
12
|
|
|
13
13
|
## 1. Install the gem
|
|
14
14
|
|
|
15
|
-
|
|
15
|
+
Use the [README release table](../README.md) to choose a version, then confirm
|
|
16
|
+
that **exact version is published** on the [RubyGems versions page](https://rubygems.org/gems/woods/versions)
|
|
17
|
+
before editing the Gemfile. A prepared release checkout can update the README
|
|
18
|
+
before its gem is published; if the version is absent, choose an available
|
|
19
|
+
version or wait for publication.
|
|
20
|
+
|
|
21
|
+
If the published 2.x line has only beta or release-candidate versions, use an
|
|
22
|
+
exact pin to the published prerelease, following the README's prerelease
|
|
23
|
+
instructions; `~> 2.0` does not select prereleases. Follow the selected version's
|
|
24
|
+
tag documentation. The `main` guides may describe features absent from the
|
|
25
|
+
published gem.
|
|
26
|
+
|
|
27
|
+
Once a stable 2.x release is published, add Woods to the development group with:
|
|
16
28
|
|
|
17
29
|
```ruby
|
|
18
30
|
# Gemfile
|
|
@@ -108,13 +120,13 @@ For example:
|
|
|
108
120
|
|
|
109
121
|
> Use Woods to find `Order`, inspect its resolved callbacks and associations, and list the first two levels of code that depend on it. Cite the Woods identifiers you used.
|
|
110
122
|
|
|
111
|
-
The Index schema inventory totals 29 schemas. Fourteen register as tools in a normal packaged launch; `codebase_retrieve` is among them but
|
|
123
|
+
The Index schema inventory totals 29 schemas. Fourteen register as tools in a normal packaged launch; `codebase_retrieve` is among them but needs embeddings in the default semantic mode, or explicit [lexical retrieval](RETRIEVAL_GUIDE.md#embedding-free-lexical-retrieval) over extraction output. The other structural tools work immediately. See [Agent guide](AGENT_GUIDE.md) for a reliable query workflow.
|
|
112
124
|
|
|
113
125
|
## Optional next steps
|
|
114
126
|
|
|
115
127
|
### Add semantic search
|
|
116
128
|
|
|
117
|
-
Structural search, exact lookup, dependency traversal, graph analysis, and flow tracing do not need embeddings.
|
|
129
|
+
Structural search, exact lookup, dependency traversal, graph analysis, and flow tracing do not need embeddings. For ranked natural-language retrieval, choose explicit [lexical mode](RETRIEVAL_GUIDE.md#embedding-free-lexical-retrieval) without providers, or configure embeddings for semantic matching.
|
|
118
130
|
|
|
119
131
|
The local preset uses SQLite metadata, persisted in-memory vectors, and a local Ollama service. Add `gem "sqlite3"` to the application bundle if it is not already present. MySQL/PostgreSQL applications that do not want that dependency can use the `:shared_filesystem` preset instead; it still uses Ollama but persists all stores beneath the Woods output directory.
|
|
120
132
|
|
|
@@ -154,7 +166,7 @@ When dependencies, initializers, database configuration, credentials, or schema
|
|
|
154
166
|
|
|
155
167
|
The watcher maintains the structural index. If semantic retrieval is enabled, also run `bin/rails woods:embed_incremental` to update vectors. Without a resident watcher, run `bin/rails woods:incremental` after changes. Use a full `woods:extract` after major upgrades or when validation reports drift. CI and shared-artifact patterns are covered in [Incremental extraction](INCREMENTAL_EXTRACTION.md).
|
|
156
168
|
|
|
157
|
-
On Rails 8.1, `config/ci.rb` can refresh the index before any gate that reads it: `step "Woods: refresh", "bin/rails woods:incremental"`. With the Claude Code plugin installed, an opt-in `PostToolUse` hook refreshes the index after graph-changing edits and an opt-in `SessionStart` hook warns
|
|
169
|
+
On Rails 8.1, `config/ci.rb` can refresh the index before any gate that reads it: `step "Woods: refresh", "bin/rails woods:incremental"`. With the Claude Code plugin installed, an opt-in `PostToolUse` hook refreshes the index after graph-changing edits and an opt-in `SessionStart` hook warns about source-content drift or unknown evidence; set `WOODS_HOOKS_ENABLED=1` to turn them on. See [Watch daemon](WATCH_DAEMON.md#hooks-for-agent-sessions).
|
|
158
170
|
|
|
159
171
|
### Enable the Console Server
|
|
160
172
|
|
|
@@ -171,7 +183,7 @@ If live-data queries are necessary, review its allowlists, blocked tables, crede
|
|
|
171
183
|
| Rails fails during extraction | Boot and eager-load Rails with the same environment variables | [Troubleshooting](TROUBLESHOOTING.md) |
|
|
172
184
|
| Validation reports missing or stale units | Run a full extraction, then validate again | [Incremental extraction](INCREMENTAL_EXTRACTION.md) |
|
|
173
185
|
| MCP reports no index or zero units | Confirm `cwd`, the host-visible `tmp/woods` path, and `woods:stats` output | [MCP servers](MCP_SERVERS.md) |
|
|
174
|
-
| `codebase_retrieve` says it is disabled |
|
|
186
|
+
| `codebase_retrieve` says it is disabled | Choose lexical mode or configure embeddings and run `woods:embed`; `search` also works | [Retrieval guide](RETRIEVAL_GUIDE.md) |
|
|
175
187
|
| Docker extraction succeeds but MCP cannot see it | Translate the container output path to its host-mounted path | [Docker setup](DOCKER_SETUP.md) |
|
|
176
188
|
|
|
177
189
|
## Where to go next
|
|
@@ -28,12 +28,24 @@ Three differences are tolerated, and nothing else:
|
|
|
28
28
|
| Ordering inside a unit's `dependents` | Full extraction appends in extractor order, incremental in graph order. Same multiset. |
|
|
29
29
|
| PageRank beyond six decimal places | Iterative floating point accumulated in each run's registration order. Scores are compared as values; only the last bits are forgiven. |
|
|
30
30
|
|
|
31
|
+
The unit-file write skip ignores only Woods' top-level `extracted_at` stamp.
|
|
32
|
+
A nested metadata field with the same name is application data: changing it
|
|
33
|
+
rewrites the unit in both compact and pretty JSON output.
|
|
34
|
+
|
|
31
35
|
`graph_analysis.json` used to be a fourth row, tolerating list ordering. It no
|
|
32
36
|
longer is: the analyzer is order-independent and the oracle compares the file
|
|
33
37
|
exactly. Tolerating the ordering there meant the harness, the only test that
|
|
34
38
|
compares a full run against an incremental one, could not see the very
|
|
35
39
|
dependence the analyzer's determinism work existed to remove.
|
|
36
40
|
|
|
41
|
+
Graph targets use string identifiers even when an extractor emits a symbolic
|
|
42
|
+
external target such as `:http_api`. Full and incremental runs retain every
|
|
43
|
+
reverse dependency across JSON restoration and re-registration, including
|
|
44
|
+
contributions from units sharing an identifier under different types. Unit
|
|
45
|
+
dependency metadata retains its extractor-provided values. If an older version
|
|
46
|
+
already lost reverse dependencies on these targets, run a full extraction once
|
|
47
|
+
to restore them; loading the damaged graph cannot recover discarded entries.
|
|
48
|
+
|
|
37
49
|
This matters most for **incremental CI chains**: restore the previous graph,
|
|
38
50
|
run `woods:incremental` per merge. There, a unit that goes missing propagates
|
|
39
51
|
forward run over run instead of being erased by the next full rebuild.
|
|
@@ -52,6 +64,10 @@ failure a CI chain cannot afford:
|
|
|
52
64
|
| The range fails otherwise | Actionable error naming the range, **exit 1**. |
|
|
53
65
|
| There is no `git` binary at all | Same two rows as above: the failure reads `git unavailable: …` and takes the daemon-coverage decision, rather than dying with an `Errno::ENOENT` backtrace. |
|
|
54
66
|
|
|
67
|
+
Changed paths are normalized lexically before dispatch: trailing root slashes,
|
|
68
|
+
duplicate separators and `.`/`..` segments do not create separate changes or
|
|
69
|
+
bypass matching. Missing files remain representable; symlinks are not resolved.
|
|
70
|
+
|
|
55
71
|
The range comes from `CI_COMMIT_BEFORE_SHA..CI_COMMIT_SHA` (GitLab),
|
|
56
72
|
`origin/$GITHUB_BASE_REF...HEAD` (GitHub Actions), or `HEAD~1` (default). An
|
|
57
73
|
unresolvable range — a GitLab zero-SHA on a new branch, an unfetched base ref,
|
|
@@ -76,6 +92,12 @@ The diff itself is rooted at the extracted application (`git -C Rails.root`),
|
|
|
76
92
|
so it cannot read whatever checkout the process happened to start in — the
|
|
77
93
|
same rooting rule the manifest's git provenance follows.
|
|
78
94
|
|
|
95
|
+
Named, source-defined app modules included by runtime models are tracked as concern units even
|
|
96
|
+
outside `concerns/` directories. Changing their source refreshes their includers,
|
|
97
|
+
including inlined code and callback analysis. Multiple runtime mixins sharing a source
|
|
98
|
+
file retain separate identities and refresh all their includers. Run a full extraction after upgrading
|
|
99
|
+
to populate these previously missing source mappings.
|
|
100
|
+
|
|
79
101
|
## What a run does, in order
|
|
80
102
|
|
|
81
103
|
`Extractor#extract_changed` is order-sensitive; each step exists because of the
|
|
@@ -120,6 +142,11 @@ step before it.
|
|
|
120
142
|
string literal, and an unrelated addition in the same batch do not.
|
|
121
143
|
Idempotent when nothing was pruned.
|
|
122
144
|
|
|
145
|
+
Git enrichment uses the same eligibility checks in full and incremental runs:
|
|
146
|
+
existing app-owned files under `Rails.root`, excluding `vendor/`, `node_modules/`,
|
|
147
|
+
and framework/gem source units. Each typed unit resolves its own file history,
|
|
148
|
+
even when its identifier is shared by another type.
|
|
149
|
+
|
|
123
150
|
Then the second pass: `dependents` and `metadata.git` are refreshed on every
|
|
124
151
|
touched unit (the incremental equivalents of full extraction's phases 2 and 4),
|
|
125
152
|
type indexes are regenerated, the graph, `graph_analysis.json` and the
|
|
@@ -174,7 +201,8 @@ automatically.
|
|
|
174
201
|
| `app/decorators`, `app/presenters`, `app/form_objects` | decorators |
|
|
175
202
|
| `app/managers` / `app/policies` / `app/validators` | managers / policies + pundit_policies / validators |
|
|
176
203
|
| `app/**/concerns/**/*.rb` | concerns |
|
|
177
|
-
| `app/models/**/*.rb` (outside `concerns/`) | poros, caching |
|
|
204
|
+
| `app/models/**/*.rb` (outside `concerns/`) | poros, caching; concerns when runtime model inclusion confirms a mixin |
|
|
205
|
+
| `app/**/*.rb`, `lib/**/*.rb` (outside `concerns/`) | runtime model mixins also dispatch to concerns |
|
|
178
206
|
| `app/controllers/**/*.rb` | caching |
|
|
179
207
|
| `app/views/**/*.erb` | view_templates, caching |
|
|
180
208
|
| `config/locales/**/*.yml` | i18n |
|
|
@@ -258,6 +286,34 @@ oracle compares against emits it too, both sides agree, wrongly. The coverage
|
|
|
258
286
|
is in `spec/extractor_spec.rb`, driving the reconciler with a shrinking
|
|
259
287
|
discovery set.
|
|
260
288
|
|
|
289
|
+
### Runtime removals and bundle updates
|
|
290
|
+
|
|
291
|
+
Jobs discovered through `ApplicationJob.descendants` supplement the job-file
|
|
292
|
+
scan, but jobs are not part of `CLASS_BASED_DISCOVERY` removal reconciliation.
|
|
293
|
+
If a dynamically defined or gem-owned job disappears without a tracked source
|
|
294
|
+
path changing, its unit can survive subsequent incremental runs. A full
|
|
295
|
+
extraction in a fresh Rails process removes it; an in-process full extraction
|
|
296
|
+
can still see an old constant retained by that process (B-165).
|
|
297
|
+
|
|
298
|
+
After adding, removing, or updating bundled gems, boot the updated bundle in a
|
|
299
|
+
fresh process and run:
|
|
300
|
+
|
|
301
|
+
```bash
|
|
302
|
+
bundle exec rake woods:extract woods:validate
|
|
303
|
+
```
|
|
304
|
+
|
|
305
|
+
A `Gemfile.lock` change refreshes engines, middleware, and optional framework
|
|
306
|
+
sources. It does not refresh every gem-owned model, job, or other runtime unit.
|
|
307
|
+
Their recorded paths or metadata can remain stale, including absolute paths to
|
|
308
|
+
a removed gem version and paths under `vendor/`. A full extraction rebuilds
|
|
309
|
+
those units against the installed bundle (B-166).
|
|
310
|
+
|
|
311
|
+
An absent external source path can also mean the validator runs on a different
|
|
312
|
+
host or mount from extraction. Confirm the bundle and filesystem context before
|
|
313
|
+
rebuilding; a full run in one container does not make its gem paths visible on
|
|
314
|
+
another host. Validation warnings identify missing paths, but do not prove a
|
|
315
|
+
retained unit matches the currently installed gem when its path still exists.
|
|
316
|
+
|
|
261
317
|
### Deletion
|
|
262
318
|
|
|
263
319
|
- Paths named in the change set that no longer exist are **authoritative** for
|
|
@@ -423,6 +479,66 @@ owns the definition of "the two indexes agree" and documents every exclusion.
|
|
|
423
479
|
|
|
424
480
|
**Run it before and after any change to the incremental path.**
|
|
425
481
|
|
|
482
|
+
## Profiling fixed costs
|
|
483
|
+
|
|
484
|
+
Set `WOODS_PROFILE=1` to time extraction phases. Git enrichment and unit JSON
|
|
485
|
+
finalization are separate from incremental re-extraction; runtime discovery,
|
|
486
|
+
whole-app reruns and pruning appear under `reconciliation`. `payload sync`,
|
|
487
|
+
`publish` (the generation pointer write), and `payload prune` (retention) are
|
|
488
|
+
separate, additive phases. Older versions included sync and retention inside
|
|
489
|
+
`publish`, so do not sum those older lines without subtracting nested sync.
|
|
490
|
+
|
|
491
|
+
`[profile total]` reports time inside the extraction call, including setup and
|
|
492
|
+
failed runs. It is not another phase to sum. Compare it with phase durations
|
|
493
|
+
to find unaccounted work; each line rounds to hundredths of a second. Rails
|
|
494
|
+
boot, Bundler and watch reload work outside the call require separate wall
|
|
495
|
+
measurements. Do not infer that all unaccounted time is boot.
|
|
496
|
+
|
|
497
|
+
Both full and incremental runs currently seed the prior payload. Cloning uses
|
|
498
|
+
per-file hardlinks (or copies where links are unsupported); updated files are
|
|
499
|
+
replaced atomically, preserving previous generations. Full-run carry-forward
|
|
500
|
+
and generation retention remain unchanged. The seed walker classifies each entry
|
|
501
|
+
once, avoiding duplicate file metadata lookups; it still links each file
|
|
502
|
+
separately. Generation directories remain independent, so pruning an old generation cannot remove a newer one's files.
|
|
503
|
+
Compare repeated runs on the actual index filesystem before attributing latency
|
|
504
|
+
to extraction or changing the index layout.
|
|
505
|
+
|
|
506
|
+
For a repeatable component comparison from a source checkout:
|
|
507
|
+
|
|
508
|
+
```bash
|
|
509
|
+
# BENCH_ROOT is an existing scratch parent on the index filesystem.
|
|
510
|
+
# The script creates and removes only its own temporary directory there.
|
|
511
|
+
BENCH_ROOT=/path/on/index/filesystem WOODS_SOURCE=/path/to/baseline \
|
|
512
|
+
ruby bench/payload_seed.rb
|
|
513
|
+
BENCH_ROOT=/path/on/index/filesystem WOODS_SOURCE=/path/to/candidate \
|
|
514
|
+
ruby bench/payload_seed.rb
|
|
515
|
+
```
|
|
516
|
+
|
|
517
|
+
The benchmark clones 8,335 synthetic 2 KiB files across 35 directories, checks
|
|
518
|
+
file counts and bytes outside timing, and reports seven samples plus medians.
|
|
519
|
+
This measures the seed component only; it does not boot Rails or establish an
|
|
520
|
+
end-to-end host improvement. Filesystem metadata latency and the copy fallback
|
|
521
|
+
can dominate differently from local hardlink results. A first full extraction
|
|
522
|
+
has no previous payload, so include a repeated full run when measuring seed cost.
|
|
523
|
+
Full seeding also preserves on-demand framework units when framework extraction
|
|
524
|
+
is disabled, and non-JSON files in existing type directories; dropping the seed
|
|
525
|
+
wholesale would change that behavior.
|
|
526
|
+
|
|
527
|
+
### Choosing full versus incremental for CI
|
|
528
|
+
|
|
529
|
+
Measure repeated full and representative-day incremental runs with
|
|
530
|
+
`WOODS_PROFILE=1` on the same resulting application tree, configuration and index
|
|
531
|
+
filesystem. Restore the same baseline index before each incremental trial;
|
|
532
|
+
otherwise a second trial may be a no-op. Include Rails boot in both wall times
|
|
533
|
+
when comparing separate task invocations, and compare resident watcher cycles
|
|
534
|
+
separately. Validate each resulting index with `woods:validate`.
|
|
535
|
+
|
|
536
|
+
Use full extraction for that workload when its median wall time is no greater
|
|
537
|
+
than the representative incremental run. There is no universal changed-file
|
|
538
|
+
threshold: shared dependencies and whole-app extractor triggers change the work
|
|
539
|
+
per file. Re-measure after substantial application or Woods changes. A fast leaf
|
|
540
|
+
edit does not establish that a day of commits is below the crossover.
|
|
541
|
+
|
|
426
542
|
## Boundaries and open work
|
|
427
543
|
|
|
428
544
|
- **Reloaded deletion is supported.** The resident watcher reloads changed
|
|
@@ -0,0 +1,382 @@
|
|
|
1
|
+
# Published index layout for shell and Python readers
|
|
2
|
+
|
|
3
|
+
Read `generation.json` at the configured index root, then read every structural
|
|
4
|
+
artifact from the payload it names. Do not require `dependency_graph.json` or
|
|
5
|
+
`manifest.json` at the root, and do not choose the highest directory in `payloads/`.
|
|
6
|
+
A directory can be present before its generation is published.
|
|
7
|
+
|
|
8
|
+
This is the filesystem contract for consumers without the Woods gem. Ruby callers
|
|
9
|
+
can use [Woods::PublishedIndex](PUBLISHED_INDEX.md). These examples require a
|
|
10
|
+
Woods 2.x payload index and a filesystem that supports the shared `flock` protocol
|
|
11
|
+
below. They deliberately reject legacy flat indexes instead of claiming an atomic
|
|
12
|
+
snapshot from them.
|
|
13
|
+
|
|
14
|
+
## Entry point and files
|
|
15
|
+
|
|
16
|
+
A typical structural index looks like this; optional entries need not exist:
|
|
17
|
+
|
|
18
|
+
```text
|
|
19
|
+
<index root>/
|
|
20
|
+
├── generation.json
|
|
21
|
+
├── payloads/
|
|
22
|
+
│ └── gen-42/
|
|
23
|
+
│ ├── manifest.json
|
|
24
|
+
│ ├── source_inputs.json # optional on older indexes
|
|
25
|
+
│ ├── dependency_graph.json
|
|
26
|
+
│ ├── graph_analysis.json
|
|
27
|
+
│ ├── SUMMARY.md
|
|
28
|
+
│ ├── models/
|
|
29
|
+
│ │ ├── _index.json
|
|
30
|
+
│ │ └── <unit filename>.json
|
|
31
|
+
│ └── flows/
|
|
32
|
+
│ ├── flow_index.json
|
|
33
|
+
│ └── <flow filename>.json
|
|
34
|
+
├── woods.json # optional configuration artifact
|
|
35
|
+
├── dumps/ # optional semantic-store snapshots
|
|
36
|
+
└── ... # operational state and locks
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
The configured root can differ between a container and the host. Resolve the
|
|
40
|
+
relative payload against the root visible to the reader, not the writer's path.
|
|
41
|
+
|
|
42
|
+
`generation.json` is one JSON object, for example:
|
|
43
|
+
|
|
44
|
+
```json
|
|
45
|
+
{"number":42,"token":"978e69b524ac74fa","updated_at":"2026-09-15T12:00:00Z","reason":"incremental","payload":"payloads/gen-42"}
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
| Field | Contract |
|
|
49
|
+
|---|---|
|
|
50
|
+
| `number` | Positive integer, incremented on publication. It is local to this index; deleting/recreating the index can restart it. |
|
|
51
|
+
| `token` | Opaque string, changed on each publication. Compare it as well as `number`; do not depend on its length or encoding. |
|
|
52
|
+
| `updated_at` | ISO8601 publication timestamp, distinct from source modification times. |
|
|
53
|
+
| `reason` | String or null describing the publication, such as `full`, `incremental`, or `refresh`. Treat values as extensible. |
|
|
54
|
+
| `payload` | Relative directory name for this generation. Current writers use `payloads/gen-N`; follow the field instead of constructing it. Omitted/null means a flat layout. |
|
|
55
|
+
|
|
56
|
+
Reject an absolute payload path or one whose resolved real path escapes the index
|
|
57
|
+
root, including through a symlink. Unknown fields can be ignored. A missing
|
|
58
|
+
pointer can mean an older flat index or no index at all; it is not evidence that a
|
|
59
|
+
payload directory is published. An existing malformed pointer or a missing named
|
|
60
|
+
payload is an error to investigate, not permission to silently serve root files.
|
|
61
|
+
Some gem readers have permissive flat fallback for these failures; the publication
|
|
62
|
+
gates below deliberately fail more strictly.
|
|
63
|
+
|
|
64
|
+
### Payload artifacts
|
|
65
|
+
|
|
66
|
+
| Artifact | Presence and meaning |
|
|
67
|
+
|---|---|
|
|
68
|
+
| `manifest.json` | Required for a complete structural publication. Counts by extractor directory, totals, extraction timestamp and provenance. Optional fields vary by writer/version; see [writer provenance](PUBLISHED_INDEX.md#manifest-writer-provenance). |
|
|
69
|
+
| `source_inputs.json` | Versioned source-input identities and per-consumer provenance for this exact generation; old indexes may omit it. See [source freshness](SOURCE_FRESHNESS.md). |
|
|
70
|
+
| `dependency_graph.json` | Required for a complete structural publication. Typed graph data; an empty graph is valid. |
|
|
71
|
+
| `<type>/_index.json` and unit JSON | Present for extracted families. `_index.json` is an array of unit summaries; an empty array is valid. Disabled/unavailable families may be absent. Do not infer completeness from a fixed count of directories. |
|
|
72
|
+
| `graph_analysis.json` | Derived graph analysis when produced. Treat absence as unavailable analysis, not an empty or corrupt unit index. |
|
|
73
|
+
| `flows/flow_index.json` and flow documents | Optional precomputed flows. Use the index's relative paths; do not invent flow filenames. |
|
|
74
|
+
| `SUMMARY.md` | Generated human-readable summary, not a machine schema. |
|
|
75
|
+
|
|
76
|
+
Unit summaries identify units, but do not carry an artifact filename. Their
|
|
77
|
+
`file_path` names the application's source file, not the unit JSON file. To avoid
|
|
78
|
+
reimplementing filename normalization, scan JSON files in a needed type directory
|
|
79
|
+
(excluding `_index.json`) and match the JSON `identifier` **and** `type`. The same
|
|
80
|
+
identifier can exist in multiple types. The full unit fields are documented in
|
|
81
|
+
[Extractor reference](EXTRACTOR_REFERENCE.md#extractedunit-field-reference).
|
|
82
|
+
|
|
83
|
+
The graph's `nodes`, typed variants, forward/reverse relationships and relationship
|
|
84
|
+
metadata belong to the graph format; preserve them when transporting the index.
|
|
85
|
+
Extractor directories can contain multiple unit types: `graphql/` contains
|
|
86
|
+
`graphql_type`, `graphql_mutation`, `graphql_resolver`, and `graphql_query`, while
|
|
87
|
+
`rails_source/` contains `rails_source` and `gem_source`. Validate and retain the
|
|
88
|
+
artifact's actual type instead of deriving it by singularizing the directory.
|
|
89
|
+
|
|
90
|
+
Do not flatten typed variants into a single node per textual identifier. The
|
|
91
|
+
static Woods self-map has the same publication envelope but different type
|
|
92
|
+
families and `manifest.provenance.mode`; it is not Rails runtime evidence.
|
|
93
|
+
|
|
94
|
+
### Semantic graph validation
|
|
95
|
+
|
|
96
|
+
`woods:validate` checks raw graph data against the unit indexes and artifacts in
|
|
97
|
+
one pinned published generation. This validation is unreleased after
|
|
98
|
+
`2.0.0.beta2`. It checks the shapes of `nodes`, `edges`, `reverse`, `file_map`,
|
|
99
|
+
`type_index`, and optional `variants`; typed identities must be unique and agree
|
|
100
|
+
with the actual indexed units. Forward sources must exist. Reverse, file, and
|
|
101
|
+
type memberships must match the union of primary and variant contributions.
|
|
102
|
+
When present, `reverse_via` must preserve the typed forward records and their
|
|
103
|
+
relationship attributes, including duplicate-record multiplicity.
|
|
104
|
+
|
|
105
|
+
Cycles, recursion, shared paths, cross-type identifier collisions, nil file paths,
|
|
106
|
+
and legacy bare-string edges are valid. Absent optional legacy fields remain
|
|
107
|
+
legal. A target absent from the nodes is an unresolved reference, not automatically
|
|
108
|
+
corruption: current metadata cannot distinguish an intentional external target
|
|
109
|
+
from an internal node and unit that are both missing. Validation cannot certify a
|
|
110
|
+
unique target type when several types share its name. A missing node whose typed
|
|
111
|
+
unit remains indexed is detectable and is an error.
|
|
112
|
+
|
|
113
|
+
The checker does not repair data, rerun extraction, or become a publication gate.
|
|
114
|
+
It reports identity/path diagnostics through the existing validation report and
|
|
115
|
+
nonzero task exit. It also supports the Woods static source map's explicitly
|
|
116
|
+
recorded type families; that does not give the map Rails runtime fidelity.
|
|
117
|
+
See [running validation from Ruby](PUBLISHED_INDEX.md#validate-a-published-generation)
|
|
118
|
+
and [semantic error recovery](TROUBLESHOOTING.md#semantic-graph-validation-errors).
|
|
119
|
+
|
|
120
|
+
### File profiles and file membership
|
|
121
|
+
|
|
122
|
+
`file_map[path]` lists units associated with a source file, including whole-file
|
|
123
|
+
profiles. It does not promise that every identifier names a Ruby constant. New
|
|
124
|
+
writers mark graph nodes of types `caching`, `configuration`, `test_mapping`,
|
|
125
|
+
`rails_source`, and `gem_source` with `"kind": "file_profile"`. The same field
|
|
126
|
+
appears on non-primary typed variants.
|
|
127
|
+
|
|
128
|
+
For example, `app/controllers/things_controller.rb` can map to both
|
|
129
|
+
`ThingsController` and a caching unit named `app/controllers/things_controller.rb`.
|
|
130
|
+
Inspect each typed node's `kind` to distinguish the profile; both retain their
|
|
131
|
+
file membership so a source edit refreshes both units. Do not classify units by
|
|
132
|
+
comparing the identifier with the path, and do not interpret a missing `kind` as
|
|
133
|
+
proof that the unit names a constant.
|
|
134
|
+
|
|
135
|
+
Older indexes omit the marker. Woods derives it from these known extractor types
|
|
136
|
+
when loading and republishing a graph, including unchanged incremental nodes.
|
|
137
|
+
Raw consumers of older indexes must treat absent markers as unclassified or use
|
|
138
|
+
the documented type list. Identifiers, `file_map`, and type membership retain
|
|
139
|
+
their existing shapes; older readers can ignore `kind`.
|
|
140
|
+
|
|
141
|
+
### Reverse relationship records
|
|
142
|
+
|
|
143
|
+
New writers add `reverse_via` to `dependency_graph.json`. Each target identifier
|
|
144
|
+
maps to its incoming relationship records, including the owning source type:
|
|
145
|
+
|
|
146
|
+
```json
|
|
147
|
+
{
|
|
148
|
+
"reverse": { "Gadget": ["Widget"] },
|
|
149
|
+
"reverse_via": {
|
|
150
|
+
"Gadget": [{ "source": "Widget", "source_type": "service", "via": "render" }]
|
|
151
|
+
}
|
|
152
|
+
}
|
|
153
|
+
```
|
|
154
|
+
|
|
155
|
+
The existing `reverse` arrays retain their bare identifiers. `reverse_via` includes
|
|
156
|
+
edges from primary nodes and typed variants; several records can share a source
|
|
157
|
+
and target while differing in type, relationship, or association attributes.
|
|
158
|
+
Optional `through`, `through_db`, and `disable_joins` values match the forward
|
|
159
|
+
edge. A `null` relationship means unknown legacy evidence. Target type remains
|
|
160
|
+
unresolved when the target identifier belongs to multiple types; the source type
|
|
161
|
+
does not resolve that ambiguity.
|
|
162
|
+
|
|
163
|
+
These are recorded dependencies, not proof that changing a target breaks every
|
|
164
|
+
source. For example, `factory_for` and migration `reference` edges describe
|
|
165
|
+
different relationships from a runtime `render` edge. Consumers can inspect one
|
|
166
|
+
target bucket without scanning the whole forward graph. Buckets and records are
|
|
167
|
+
deterministically ordered, but consumers should treat the ordering as incidental.
|
|
168
|
+
|
|
169
|
+
Older graphs omit `reverse_via`; absence means relationship detail must be derived
|
|
170
|
+
from forward edges and variants, not that there are no dependents. Woods rebuilds
|
|
171
|
+
this derived index from forward evidence when loading and republishing a graph.
|
|
172
|
+
Existing readers can ignore the additive field; a subsequent changed extraction
|
|
173
|
+
or full run publishes it. A no-op leaves the previous generation unchanged.
|
|
174
|
+
|
|
175
|
+
### What is outside this structural snapshot
|
|
176
|
+
|
|
177
|
+
`woods.json`, `dumps/`, embedding checkpoints, temporal snapshots, watch status,
|
|
178
|
+
pending paths, MCP task records, exporter state and extraction/startup lock files
|
|
179
|
+
have separate lifecycles. Their presence is configuration-dependent. Do not glob
|
|
180
|
+
them into a structural payload or assume `generation.json` commits them together.
|
|
181
|
+
A semantic/MCP deployment also needs its configured stores and artifacts; copying
|
|
182
|
+
a structural payload alone does not clone that deployment.
|
|
183
|
+
|
|
184
|
+
`<output>/.source-inputs.key` is private operational state used to verify source
|
|
185
|
+
identities. Never include it in a published/exported structural snapshot; without
|
|
186
|
+
it, a recipient can still read the index but source freshness is unknown.
|
|
187
|
+
|
|
188
|
+
Temporary filenames, abandoned payloads, lock sidecars, retained-directory counts,
|
|
189
|
+
and summary formatting are implementation details. Never remove or modify locks
|
|
190
|
+
to make a reader proceed.
|
|
191
|
+
|
|
192
|
+
## Atomic publication is not indefinite retention
|
|
193
|
+
|
|
194
|
+
For a payload publication, Woods writes the payload, flushes it, and atomically
|
|
195
|
+
replaces `generation.json` last. Capturing that pointer once selects a complete,
|
|
196
|
+
immutable structural generation. Opening each file through a freshly reread pointer
|
|
197
|
+
can mix generations and defeats that guarantee. A failed/no-op run does not advance
|
|
198
|
+
the pointer. [Durability details](PUBLISHED_INDEX.md#durability-the-pointer-is-the-commit-point)
|
|
199
|
+
explain the flush boundary.
|
|
200
|
+
|
|
201
|
+
Retention can delete an older payload after the reader selects it. The default
|
|
202
|
+
retains three generations, not three minutes. To keep a multi-file read or copy
|
|
203
|
+
safe from Woods retention:
|
|
204
|
+
|
|
205
|
+
1. Read and validate the pointer; resolve its payload inside the root.
|
|
206
|
+
2. Open that payload's existing `manifest.json` **read-only** and take a shared
|
|
207
|
+
advisory `flock` on that open file. Do not create a new lock file.
|
|
208
|
+
3. Re-read the pointer and confirm it is unchanged; verify the pathname still
|
|
209
|
+
names the open manifest inode. A pruner may have won between steps 1 and 2.
|
|
210
|
+
4. Keep the handle/lock open for **all** reads or the complete copy. Use the single
|
|
211
|
+
captured payload path throughout. Once pinned, a later pointer advance is fine.
|
|
212
|
+
5. Close the handle when done. On a race, discard partial results and retry the
|
|
213
|
+
entire operation from step 1, with a bounded retry count.
|
|
214
|
+
|
|
215
|
+
Woods retention attempts a nonblocking **exclusive** `flock` on that same manifest
|
|
216
|
+
before deletion and skips a payload held by readers. These are advisory filesystem
|
|
217
|
+
locks, not Woods' writer-coordination locks. Ordinary readers need no exclusive
|
|
218
|
+
lock and need not take `extraction.lock` or its guard. A lock-free reader must
|
|
219
|
+
accept disappearance and restart the whole read; it must never silently replace
|
|
220
|
+
missing files with files from another generation.
|
|
221
|
+
|
|
222
|
+
This protocol protects against cooperating retention only. It does not protect
|
|
223
|
+
against `woods:clean`, manual deletion, index replacement, or a filesystem that
|
|
224
|
+
does not coordinate `flock` across its clients. Stop writers/cleanup and read an
|
|
225
|
+
immutable snapshot if that protection is unavailable. Flat indexes (including
|
|
226
|
+
full-extraction fallback when a payload cannot be created) have individually
|
|
227
|
+
replaced files, not multi-file atomicity; read/copy them only with writers stopped.
|
|
228
|
+
|
|
229
|
+
## Bash and jq: read one pinned generation
|
|
230
|
+
|
|
231
|
+
Requires Bash, jq, GNU `realpath`, util-linux `flock`, and Linux `/proc`. Save as
|
|
232
|
+
`read-woods.sh`, then run `bash read-woods.sh '/path with spaces/tmp/woods'`.
|
|
233
|
+
The final command reads manifest and graph under the same lock. Substitute other
|
|
234
|
+
reads or a complete copy **inside** the script before it exits. Printing a path
|
|
235
|
+
and consuming it after the script exits does not keep it pinned.
|
|
236
|
+
|
|
237
|
+
Exit 75 means a possible publication/retention race: retry the whole script a
|
|
238
|
+
bounded number of times (for example three), discarding any previous output.
|
|
239
|
+
Other nonzero exits need investigation. A persistent 75 can mean a broken index.
|
|
240
|
+
|
|
241
|
+
```bash
|
|
242
|
+
#!/usr/bin/env bash
|
|
243
|
+
set -euo pipefail
|
|
244
|
+
fail() { echo "$*" >&2; exit 1; }
|
|
245
|
+
retry() { echo "$*; retry the whole read" >&2; exit 75; }
|
|
246
|
+
root=$(realpath -e -- "${1:?provide the index root}")
|
|
247
|
+
[[ -d "$root" ]] || fail 'Index root is not a directory'
|
|
248
|
+
[[ -f "$root/generation.json" ]] || fail 'Missing pointer: legacy or unpublished index'
|
|
249
|
+
marker=$(cat -- "$root/generation.json") || retry 'Cannot read pointer'
|
|
250
|
+
jq -e '
|
|
251
|
+
if type != "object" then false
|
|
252
|
+
elif (.number | type) != "number" then false
|
|
253
|
+
else (.number >= 1 and (.number | floor) == .number)
|
|
254
|
+
and (.token | type == "string" and length > 0)
|
|
255
|
+
and (.payload | type == "string" and length > 0)
|
|
256
|
+
and (.payload | explode | all(. >= 32 and . != 127))
|
|
257
|
+
end
|
|
258
|
+
' <<<"$marker" >/dev/null || fail 'Invalid pointer or flat layout'
|
|
259
|
+
relative=$(jq -r '.payload' <<<"$marker")
|
|
260
|
+
[[ "$relative" != /* ]] || fail 'Absolute payload path'
|
|
261
|
+
payload=$(realpath -e -- "$root/$relative") || retry 'Missing payload'
|
|
262
|
+
[[ -d "$payload" && "$payload" == "${root%/}/"* && "$payload" != "$root" ]] \
|
|
263
|
+
|| fail 'Payload must resolve inside the index root'
|
|
264
|
+
exec 9< "$payload/manifest.json" || retry 'Missing manifest'
|
|
265
|
+
flock -sn 9 || retry 'Cannot pin manifest'
|
|
266
|
+
current=$(cat -- "$root/generation.json") || retry 'Pointer disappeared'
|
|
267
|
+
[[ "$current" == "$marker" && "$payload/manifest.json" -ef /proc/self/fd/9 ]] \
|
|
268
|
+
|| retry 'Publication changed or retention removed the payload'
|
|
269
|
+
jq -n --argjson generation "$marker" \
|
|
270
|
+
--slurpfile manifest "$payload/manifest.json" \
|
|
271
|
+
--slurpfile graph "$payload/dependency_graph.json" \
|
|
272
|
+
'if ($manifest | length) == 1 and ($manifest[0] | type) == "object"
|
|
273
|
+
and ($graph | length) == 1 and ($graph[0] | type) == "object"
|
|
274
|
+
then {generation: $generation, manifest: $manifest[0], dependency_graph: $graph[0]}
|
|
275
|
+
else error("Manifest and graph must each contain exactly one JSON object") end'
|
|
276
|
+
# Descriptor 9 closes on exit, releasing the retention pin.
|
|
277
|
+
```
|
|
278
|
+
|
|
279
|
+
## Python: keep the pin while using the payload
|
|
280
|
+
|
|
281
|
+
Requires Python 3.9+ on Unix with working `fcntl.flock`. Save as `read_woods.py`
|
|
282
|
+
and run `python3 read_woods.py '/path with spaces/tmp/woods'`. The context manager
|
|
283
|
+
retries acquisition three times. Errors during the read/copy propagate: discard
|
|
284
|
+
partial output before retrying the entire operation. The scripts assume a trusted
|
|
285
|
+
Woods-owned index; path containment is not a sandbox for hostile filesystem changes.
|
|
286
|
+
|
|
287
|
+
```python
|
|
288
|
+
import contextlib
|
|
289
|
+
import fcntl
|
|
290
|
+
import json
|
|
291
|
+
import os
|
|
292
|
+
from pathlib import Path
|
|
293
|
+
import sys
|
|
294
|
+
import time
|
|
295
|
+
|
|
296
|
+
|
|
297
|
+
@contextlib.contextmanager
|
|
298
|
+
def pinned_payload(index_root):
|
|
299
|
+
root = Path(index_root).resolve(strict=True)
|
|
300
|
+
pointer = root / "generation.json"
|
|
301
|
+
if not pointer.is_file():
|
|
302
|
+
raise ValueError("Missing pointer: legacy or unpublished index")
|
|
303
|
+
for attempt in range(3):
|
|
304
|
+
handle = None
|
|
305
|
+
try:
|
|
306
|
+
raw = pointer.read_bytes()
|
|
307
|
+
marker = json.loads(raw)
|
|
308
|
+
if not isinstance(marker, dict):
|
|
309
|
+
raise ValueError("Pointer must be an object")
|
|
310
|
+
number, token, name = (marker.get(k) for k in ("number", "token", "payload"))
|
|
311
|
+
if type(number) is not int or number < 1 or not isinstance(token, str) or not token:
|
|
312
|
+
raise ValueError("Invalid generation identity")
|
|
313
|
+
if not isinstance(name, str) or not name or Path(name).is_absolute():
|
|
314
|
+
raise ValueError("Invalid payload path or flat layout")
|
|
315
|
+
if any(ord(char) < 32 or ord(char) == 127 for char in name):
|
|
316
|
+
raise ValueError("Control character in payload path")
|
|
317
|
+
payload = (root / name).resolve(strict=True)
|
|
318
|
+
if not payload.is_dir() or root not in payload.parents:
|
|
319
|
+
raise ValueError("Payload must resolve inside the index root")
|
|
320
|
+
manifest = payload / "manifest.json"
|
|
321
|
+
handle = manifest.open("rb")
|
|
322
|
+
fcntl.flock(handle, fcntl.LOCK_SH | fcntl.LOCK_NB)
|
|
323
|
+
if pointer.read_bytes() != raw or not os.path.samestat(os.fstat(handle.fileno()), manifest.stat()):
|
|
324
|
+
raise BlockingIOError("Publication changed or payload was removed")
|
|
325
|
+
except (FileNotFoundError, BlockingIOError):
|
|
326
|
+
if handle is not None:
|
|
327
|
+
handle.close()
|
|
328
|
+
if attempt == 2:
|
|
329
|
+
raise
|
|
330
|
+
time.sleep(0.05)
|
|
331
|
+
except BaseException:
|
|
332
|
+
if handle is not None:
|
|
333
|
+
handle.close()
|
|
334
|
+
raise
|
|
335
|
+
else:
|
|
336
|
+
break
|
|
337
|
+
try:
|
|
338
|
+
yield marker, payload
|
|
339
|
+
finally:
|
|
340
|
+
handle.close()
|
|
341
|
+
|
|
342
|
+
|
|
343
|
+
if __name__ == "__main__":
|
|
344
|
+
with pinned_payload(sys.argv[1]) as (generation, payload):
|
|
345
|
+
manifest = json.loads((payload / "manifest.json").read_text(encoding="utf-8"))
|
|
346
|
+
graph = json.loads((payload / "dependency_graph.json").read_text(encoding="utf-8"))
|
|
347
|
+
if not isinstance(manifest, dict) or not isinstance(graph, dict):
|
|
348
|
+
raise ValueError("Manifest and graph must each contain exactly one JSON object")
|
|
349
|
+
# Read more artifacts, or copy the whole payload, before leaving this block.
|
|
350
|
+
print(json.dumps({"generation": generation, "manifest": manifest, "dependency_graph": graph}))
|
|
351
|
+
```
|
|
352
|
+
|
|
353
|
+
## Shipping a structural snapshot
|
|
354
|
+
|
|
355
|
+
While the pin is held, copy the selected payload into an unpublished staging
|
|
356
|
+
location. Preserve the payload's relative path and pair it with the **captured**
|
|
357
|
+
`generation.json`, not a later pointer reread from the live index. For example,
|
|
358
|
+
a captured `payloads/gen-42` must still resolve to that directory in the exported
|
|
359
|
+
root. Validate the copied manifest and graph, then publish the staged copy as a
|
|
360
|
+
whole, or upload its payload first and switch the destination pointer last.
|
|
361
|
+
Do not upload `generation.json` first or advertise a failed/partial copy.
|
|
362
|
+
|
|
363
|
+
The retention lock is local coordination; do not copy its file descriptor or
|
|
364
|
+
writer lock files. Copying `manifest.json` normally copies content, which is correct.
|
|
365
|
+
Objects in a destination store do not inherit the source's locking guarantees;
|
|
366
|
+
protect any later destination pruning with that destination's own reader protocol.
|
|
367
|
+
|
|
368
|
+
## Compatibility within Woods 2.x
|
|
369
|
+
|
|
370
|
+
Consumers may rely on the pointer field meanings, relative payload resolution,
|
|
371
|
+
the complete-payload publication boundary, manifest/graph locations, and JSON unit
|
|
372
|
+
identity described here. Additive JSON fields, new extractor families, new reason
|
|
373
|
+
values, optional artifacts and additional root-level state may appear in 2.x.
|
|
374
|
+
Ignore unknown fields and inspect available types instead of hardcoding a directory
|
|
375
|
+
count. Existing field meanings and required-file locations are compatibility
|
|
376
|
+
surfaces; incompatible changes require an explicit migration contract.
|
|
377
|
+
|
|
378
|
+
Do not bind to temporary names, a fixed token length, directory listing order,
|
|
379
|
+
retention count, internal lock sidecars, or Markdown summary formatting. Check the
|
|
380
|
+
installed Woods version and its release's documentation when consuming older
|
|
381
|
+
prereleases; this guide describes the current source contract, not a promise that
|
|
382
|
+
every prerelease contains every optional field or retention improvement.
|
data/docs/INTERNALS.md
CHANGED
|
@@ -145,6 +145,11 @@ The `DependencyGraph` is a directed graph where nodes are `ExtractedUnit` identi
|
|
|
145
145
|
- **Forward edges** (`@edges`): what each unit depends on, populated when units are registered
|
|
146
146
|
- **Reverse edges** (`@reverse`): what depends on each unit, built during registration and in the resolve phase
|
|
147
147
|
|
|
148
|
+
The published graph also carries additive `reverse_via` target buckets with
|
|
149
|
+
typed source identities, relationship labels and association attributes. The
|
|
150
|
+
bare-name `reverse` map remains compatible. See the
|
|
151
|
+
[reverse relationship format](INDEX_LAYOUT.md#reverse-relationship-records).
|
|
152
|
+
|
|
148
153
|
```ruby
|
|
149
154
|
graph = DependencyGraph.new
|
|
150
155
|
graph.register(user_unit) # adds User to nodes, adds User→Order edge (from belongs_to)
|
|
@@ -175,7 +180,7 @@ Scores feed into the retrieval ranker as one signal in the final ranking formula
|
|
|
175
180
|
| **Cycles** | Circular dependencies, A→B→C→A. Detected via DFS, and capped: `graph_cycle_limit` (default 500) bounds how many are enumerated and `graph_cycle_max_length` (default 50) skips one longer than that. Either cap firing sets `stats.cycle_limit_reached`. Set both to `nil` for exhaustive enumeration. |
|
|
176
181
|
| **Bridges** | Edges whose removal would disconnect the graph, high-risk structural connections |
|
|
177
182
|
| **Cross-database edges** | Association or foreign-key edges whose two ends resolve to different databases. A `has_many :through` is reported as `join_through_across_databases` when `disable_joins` is false and `from_db`, `through_db` (the join model's database), or `to_db` disagree. A foreign key never resolves to an owner in the source database, even when another database also claims the table; when every owner sits elsewhere and they span more than one database, the entry comes back with `to: nil` and an `ambiguous_owners` list instead of guessing. Read from graph node and edge attributes, so full and incremental runs agree. Scoped to primary nodes (units registered in the graph), not variants. |
|
|
178
|
-
| **Volatile dependencies** | Edges that point at a unit changing at least `volatile_dependency_ratio` times more often than the dependent (POODR: depend on things that change less often than you do). Dependencies with fewer than 5 commits or a `new` change frequency are skipped. Ranked by the dependency's PageRank;
|
|
183
|
+
| **Volatile dependencies** | Edges that point at a unit changing at least `volatile_dependency_ratio` times more often than the dependent (POODR: depend on things that change less often than you do). Dependencies with fewer than 5 commits or a `new` change frequency are skipped. Ranked by the dependency's PageRank; an optional `volatile_dependency_limit_per_target` selects edges per typed dependency before the persisted top 20, while `stats.volatile_dependency_count` reports the full qualifying count and `stats.volatile_dependencies_limit` reports the cap. |
|
|
179
184
|
| **Undeclared package edges** | Edges that cross a Packwerk package boundary the source package does not list in `dependencies`. Membership comes from each unit's `package` node attribute, declarations from the package unit's own `package_dependency` edges. Woods reports the boundary; enforcement stays with `packwerk check` / `pks check`. |
|
|
180
185
|
|
|
181
186
|
Analysis results are written to `graph_analysis.json` and surfaced in `SUMMARY.md`.
|
|
@@ -194,7 +199,7 @@ persisted rather than every hub in the graph.
|
|
|
194
199
|
| `cycles` | `stats.cycle_count`, `stats.cycle_limit_reached` | yes, `graph_cycle_limit` cycles of at most `graph_cycle_max_length` nodes | the count is the persisted array; the flag says whether either cap fired |
|
|
195
200
|
| `bridges` | none | yes, top 10 by score | not counted |
|
|
196
201
|
| `cross_database_edges` | `stats.cross_database_edge_count` | no | every crossing edge |
|
|
197
|
-
| `volatile_dependencies` | `stats.volatile_dependency_count`, `stats.volatile_dependencies_limit` | yes, top 20 by the dependency's PageRank |
|
|
202
|
+
| `volatile_dependencies` | `stats.volatile_dependency_count`, `stats.volatile_dependencies_limit`; optional `stats.volatile_dependencies_limit_per_target`, `stats.volatile_dependency_reported_count` | yes, optional per-target cap followed by top 20 by the dependency's PageRank | count is every qualifying edge before caps; reported count is the final array length and appears only when the per-target cap is enabled |
|
|
198
203
|
| `undeclared_package_edges` | `stats.undeclared_package_edge_count` | no | every undeclared crossing |
|
|
199
204
|
|
|
200
205
|
A capped section means the array on disk is a page, not the population.
|