woods 2.0.0.beta2 → 2.0.0.beta4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +339 -1
- data/CONTRIBUTING.md +188 -12
- data/README.md +93 -174
- data/SECURITY.md +9 -6
- data/docs/AGENT_GUIDE.md +109 -8
- data/docs/AGENT_SETUP.md +98 -7
- data/docs/BACKEND_MATRIX.md +25 -0
- data/docs/CLIENT_HOOKS.md +111 -0
- data/docs/CONFIGURATION_REFERENCE.md +267 -16
- data/docs/CONSOLE_MCP_SETUP.md +80 -7
- data/docs/DOCKER_SETUP.md +22 -3
- data/docs/EVALUATION.md +464 -1
- data/docs/EXTRACTOR_REFERENCE.md +45 -6
- data/docs/FAQ.md +11 -12
- data/docs/GETTING_STARTED.md +17 -5
- data/docs/INCREMENTAL_EXTRACTION.md +147 -7
- data/docs/INDEX_LAYOUT.md +382 -0
- data/docs/INTERNALS.md +7 -2
- data/docs/MCP_SERVERS.md +276 -5
- data/docs/MCP_TOOL_COOKBOOK.md +37 -22
- data/docs/MCP_WORKTREE_SETUP.md +43 -83
- data/docs/NOTION_INTEGRATION.md +13 -0
- data/docs/OBSIDIAN_INTEGRATION.md +57 -9
- data/docs/PUBLISHED_INDEX.md +72 -0
- data/docs/README.md +7 -0
- data/docs/RETRIEVAL_GUIDE.md +273 -12
- data/docs/RUNTIME_TRACING.md +71 -0
- data/docs/SOURCE_FRESHNESS.md +143 -0
- data/docs/TROUBLESHOOTING.md +129 -18
- data/docs/UNBLOCKED_INTEGRATION.md +25 -0
- data/docs/UPGRADING_TO_2.md +48 -22
- data/docs/WATCH_DAEMON.md +277 -67
- data/exe/woods-agent-config +6 -0
- data/exe/woods-extract +5 -0
- data/exe/woods-hook-context +6 -0
- data/exe/woods-mcp-start +14 -9
- data/lib/generators/woods/pgvector_generator.rb +8 -2
- data/lib/generators/woods/templates/woods.rb.tt +1 -3
- data/lib/tasks/woods.rake +47 -397
- data/lib/woods/agent_configuration/applier.rb +135 -0
- data/lib/woods/agent_configuration/cli.rb +101 -0
- data/lib/woods/agent_configuration/cli_options.rb +29 -0
- data/lib/woods/agent_configuration/document.rb +105 -0
- data/lib/woods/agent_configuration/error.rb +7 -0
- data/lib/woods/agent_configuration/launcher.rb +75 -0
- data/lib/woods/agent_configuration/layout.rb +72 -0
- data/lib/woods/agent_configuration/managed_section.rb +62 -0
- data/lib/woods/agent_configuration/plan.rb +98 -0
- data/lib/woods/agent_configuration/plan_diff.rb +38 -0
- data/lib/woods/agent_configuration/planned_files.rb +61 -0
- data/lib/woods/agent_configuration/planner.rb +63 -0
- data/lib/woods/agent_configuration/planner_validation.rb +77 -0
- data/lib/woods/agent_configuration/preflight.rb +100 -0
- data/lib/woods/agent_configuration/recovery.rb +49 -0
- data/lib/woods/ast/node.rb +2 -0
- data/lib/woods/ast/parser.rb +38 -5
- data/lib/woods/builder.rb +21 -5
- data/lib/woods/cache/cache_middleware.rb +28 -7
- data/lib/woods/cache/cache_store.rb +4 -5
- data/lib/woods/change_set.rb +5 -4
- data/lib/woods/console/credential_index.rb +20 -2
- data/lib/woods/console/credential_scanner.rb +18 -17
- data/lib/woods/console/credential_scanner_registry.rb +36 -0
- data/lib/woods/console/dispatch_pipeline.rb +7 -0
- data/lib/woods/console/embedded_executor.rb +32 -10
- data/lib/woods/console/encrypted_credential_snapshot.rb +16 -0
- data/lib/woods/console/rack_middleware.rb +22 -13
- data/lib/woods/console/server.rb +18 -16
- data/lib/woods/console/sql_noise_stripper.rb +9 -7
- data/lib/woods/console/sql_table_scanner.rb +47 -7
- data/lib/woods/console/sql_validator.rb +49 -9
- data/lib/woods/console/sqlite_read_guard.rb +46 -0
- data/lib/woods/coordination/pipeline_lock.rb +3 -2
- data/lib/woods/dependency_graph.rb +65 -13
- data/lib/woods/embedding/corpus.rb +94 -0
- data/lib/woods/embedding/indexer.rb +114 -60
- data/lib/woods/embedding/openai.rb +17 -6
- data/lib/woods/evaluation/ablation_executor.rb +6 -1
- data/lib/woods/evaluation/ablation_timed_executor.rb +22 -4
- data/lib/woods/export/typed_reader.rb +56 -0
- data/lib/woods/extractor.rb +277 -149
- data/lib/woods/extractors/action_cable_extractor.rb +3 -1
- data/lib/woods/extractors/behavioral_profile.rb +9 -7
- data/lib/woods/extractors/caching_extractor.rb +3 -1
- data/lib/woods/extractors/concern_extractor.rb +64 -6
- data/lib/woods/extractors/configuration_extractor.rb +7 -3
- data/lib/woods/extractors/controller_extractor.rb +13 -4
- data/lib/woods/extractors/database_view_extractor.rb +3 -1
- data/lib/woods/extractors/declared_parent.rb +55 -0
- data/lib/woods/extractors/decorator_extractor.rb +3 -1
- data/lib/woods/extractors/engine_extractor.rb +3 -1
- data/lib/woods/extractors/event_extractor.rb +4 -2
- data/lib/woods/extractors/factory_extractor.rb +3 -1
- data/lib/woods/extractors/graphql_extractor.rb +10 -13
- data/lib/woods/extractors/i18n_extractor.rb +3 -1
- data/lib/woods/extractors/job_extractor.rb +6 -19
- data/lib/woods/extractors/lib_extractor.rb +13 -9
- data/lib/woods/extractors/mailer_extractor.rb +26 -15
- data/lib/woods/extractors/manager_extractor.rb +3 -1
- data/lib/woods/extractors/method_parameters.rb +53 -0
- data/lib/woods/extractors/middleware_argument.rb +65 -0
- data/lib/woods/extractors/middleware_extractor.rb +9 -3
- data/lib/woods/extractors/migration_extractor.rb +3 -1
- data/lib/woods/extractors/model_extractor.rb +26 -34
- data/lib/woods/extractors/package_extractor.rb +24 -4
- data/lib/woods/extractors/phlex_extractor.rb +3 -1
- data/lib/woods/extractors/policy_extractor.rb +3 -1
- data/lib/woods/extractors/poro_extractor.rb +13 -9
- data/lib/woods/extractors/pundit_extractor.rb +3 -1
- data/lib/woods/extractors/rails_source_extractor.rb +4 -2
- data/lib/woods/extractors/rake_task_extractor.rb +4 -2
- data/lib/woods/extractors/route_extractor.rb +3 -1
- data/lib/woods/extractors/route_helper_resolver.rb +10 -33
- data/lib/woods/extractors/scheduled_job_extractor.rb +41 -15
- data/lib/woods/extractors/serializer_extractor.rb +4 -2
- data/lib/woods/extractors/service_extractor.rb +3 -1
- data/lib/woods/extractors/shared_dependency_scanner.rb +2 -2
- data/lib/woods/extractors/shared_utility_methods.rb +48 -19
- data/lib/woods/extractors/source_nesting.rb +1 -1
- data/lib/woods/extractors/state_machine_extractor.rb +3 -1
- data/lib/woods/extractors/test_mapping_extractor.rb +3 -1
- data/lib/woods/extractors/validator_extractor.rb +3 -1
- data/lib/woods/extractors/view_component_extractor.rb +3 -1
- data/lib/woods/extractors/view_template_extractor.rb +3 -1
- data/lib/woods/gem_mapper.rb +2 -0
- data/lib/woods/git_history.rb +116 -0
- data/lib/woods/graph_analyzer.rb +35 -6
- data/lib/woods/hooks/context_cli.rb +54 -0
- data/lib/woods/hooks/context_event.rb +88 -0
- data/lib/woods/hooks/context_hint.rb +73 -0
- data/lib/woods/hooks/context_impact.rb +77 -0
- data/lib/woods/hooks/context_output.rb +47 -0
- data/lib/woods/hooks/context_state.rb +102 -0
- data/lib/woods/hooks/refresh.rb +79 -0
- data/lib/woods/hooks/rule_projection.rb +78 -0
- data/lib/woods/input_rules.rb +19 -0
- data/lib/woods/mcp/bearer_auth.rb +22 -13
- data/lib/woods/mcp/bootstrapper.rb +79 -4
- data/lib/woods/mcp/config_resolver.rb +2 -1
- data/lib/woods/mcp/index_reader.rb +334 -162
- data/lib/woods/mcp/initialization_guidance.rb +27 -0
- data/lib/woods/mcp/origin_guard.rb +17 -9
- data/lib/woods/mcp/published_lexical_retriever.rb +115 -0
- data/lib/woods/mcp/renderers/markdown_renderer.rb +22 -9
- data/lib/woods/mcp/renderers/plain_renderer.rb +18 -8
- data/lib/woods/mcp/search_results.rb +74 -0
- data/lib/woods/mcp/server.rb +178 -63
- data/lib/woods/mcp/tool_contract.rb +3 -1
- data/lib/woods/mcp/tool_response_renderer.rb +41 -0
- data/lib/woods/mcp/traversal_evidence.rb +113 -0
- data/lib/woods/mcp/traversal_evidence_index.rb +100 -0
- data/lib/woods/mcp/traversal_evidence_page.rb +41 -0
- data/lib/woods/mcp/traversal_evidence_text.rb +52 -0
- data/lib/woods/mcp/traversal_response.rb +22 -0
- data/lib/woods/notion/exporter.rb +56 -17
- data/lib/woods/obsidian/destination_plan.rb +98 -0
- data/lib/woods/obsidian/name_mapper.rb +19 -3
- data/lib/woods/obsidian/note_builder.rb +19 -10
- data/lib/woods/obsidian/vault_exporter.rb +88 -32
- data/lib/woods/operator/pipeline_guard.rb +18 -13
- data/lib/woods/path_dispatcher.rb +13 -6
- data/lib/woods/payload_store.rb +27 -26
- data/lib/woods/published_index/typed_unit_reader.rb +40 -3
- data/lib/woods/published_index.rb +2 -2
- data/lib/woods/railtie.rb +3 -3
- data/lib/woods/railtie_support.rb +12 -12
- data/lib/woods/rake_helpers.rb +382 -0
- data/lib/woods/resilience/graph_invariant_validator/membership_checks.rb +71 -0
- data/lib/woods/resilience/graph_invariant_validator/node_checks.rb +61 -0
- data/lib/woods/resilience/graph_invariant_validator/reverse_relationship_checks.rb +46 -0
- data/lib/woods/resilience/graph_invariant_validator.rb +119 -0
- data/lib/woods/resilience/index_validator/graph_checks.rb +80 -0
- data/lib/woods/resilience/index_validator.rb +112 -23
- data/lib/woods/retrieval/context_assembler.rb +50 -15
- data/lib/woods/retrieval/lexical_assembler.rb +84 -0
- data/lib/woods/retrieval/lexical_index.rb +120 -0
- data/lib/woods/retrieval/ranker.rb +4 -2
- data/lib/woods/retrieval/scope.rb +108 -0
- data/lib/woods/retrieval/scoped_graph_store.rb +32 -0
- data/lib/woods/retrieval/scoped_vector_store.rb +55 -0
- data/lib/woods/retrieval/search_executor.rb +86 -27
- data/lib/woods/retrieval/source_evidence.rb +200 -0
- data/lib/woods/retriever.rb +98 -22
- data/lib/woods/ruby_analyzer/trace_enricher.rb +77 -38
- data/lib/woods/session_tracer/file_store.rb +6 -1
- data/lib/woods/session_tracer/middleware.rb +10 -12
- data/lib/woods/session_tracer/redis_store.rb +22 -6
- data/lib/woods/session_tracer/session_flow_assembler.rb +23 -17
- data/lib/woods/session_tracer/solid_cache_coordination.rb +6 -4
- data/lib/woods/session_tracer/unit_resolver.rb +63 -0
- data/lib/woods/source_inputs/consumer_errors.rb +31 -0
- data/lib/woods/source_inputs/handoff.rb +102 -0
- data/lib/woods/source_inputs/launcher.rb +157 -0
- data/lib/woods/source_inputs/manifest.rb +124 -0
- data/lib/woods/source_inputs/private_key.rb +55 -0
- data/lib/woods/source_inputs/scanner.rb +171 -0
- data/lib/woods/source_inputs/scopes.rb +71 -0
- data/lib/woods/source_inputs/session.rb +214 -0
- data/lib/woods/source_inputs/status.rb +84 -0
- data/lib/woods/source_inputs/verifier.rb +107 -0
- data/lib/woods/storage/metadata_store.rb +25 -25
- data/lib/woods/storage/pgvector.rb +35 -10
- data/lib/woods/storage/qdrant.rb +17 -7
- data/lib/woods/storage/vector_store.rb +18 -6
- data/lib/woods/tasks.rb +3 -2
- data/lib/woods/temporal/json_snapshot_store.rb +58 -9
- data/lib/woods/unblocked/exporter.rb +59 -70
- data/lib/woods/version.rb +1 -1
- data/lib/woods/watch/boot_snapshot.rb +52 -0
- data/lib/woods/watch/daemon.rb +154 -32
- data/lib/woods/watch/listen_watcher.rb +4 -0
- data/lib/woods/watch/polling_watcher.rb +5 -1
- data/lib/woods/watch/status.rb +20 -15
- data/lib/woods/watch/tree_scan.rb +21 -13
- data/lib/woods/watch/watcher.rb +4 -1
- data/lib/woods.rb +50 -11
- data/plugin/.claude-plugin/plugin.json +1 -1
- data/plugin/hooks/adapters/normalize.jq +15 -0
- data/plugin/hooks/adapters/normalize.rb +63 -0
- data/plugin/hooks/hooks.json +20 -0
- data/plugin/hooks/woods-context.sh +50 -0
- data/plugin/hooks/woods-input-rules.sh +159 -0
- data/plugin/hooks/woods-opencode.mjs +65 -0
- data/plugin/hooks/woods-post-edit.sh +2 -225
- data/plugin/hooks/woods-refresh.sh +260 -0
- data/plugin/hooks/woods-session-start.sh +47 -55
- data/plugin/skills/woods-agent-enable/SKILL.md +19 -0
- data/plugin/skills/woods-diagnose/SKILL.md +319 -1
- data/plugin/skills/woods-investigate/SKILL.md +145 -0
- data/plugin/skills/woods-mcp-config/SKILL.md +90 -2
- data/plugin/skills/woods-setup/SKILL.md +110 -6
- metadata +87 -5
|
@@ -0,0 +1,382 @@
|
|
|
1
|
+
# Published index layout for shell and Python readers
|
|
2
|
+
|
|
3
|
+
Read `generation.json` at the configured index root, then read every structural
|
|
4
|
+
artifact from the payload it names. Do not require `dependency_graph.json` or
|
|
5
|
+
`manifest.json` at the root, and do not choose the highest directory in `payloads/`.
|
|
6
|
+
A directory can be present before its generation is published.
|
|
7
|
+
|
|
8
|
+
This is the filesystem contract for consumers without the Woods gem. Ruby callers
|
|
9
|
+
can use [Woods::PublishedIndex](PUBLISHED_INDEX.md). These examples require a
|
|
10
|
+
Woods 2.x payload index and a filesystem that supports the shared `flock` protocol
|
|
11
|
+
below. They deliberately reject legacy flat indexes instead of claiming an atomic
|
|
12
|
+
snapshot from them.
|
|
13
|
+
|
|
14
|
+
## Entry point and files
|
|
15
|
+
|
|
16
|
+
A typical structural index looks like this; optional entries need not exist:
|
|
17
|
+
|
|
18
|
+
```text
|
|
19
|
+
<index root>/
|
|
20
|
+
├── generation.json
|
|
21
|
+
├── payloads/
|
|
22
|
+
│ └── gen-42/
|
|
23
|
+
│ ├── manifest.json
|
|
24
|
+
│ ├── source_inputs.json # optional on older indexes
|
|
25
|
+
│ ├── dependency_graph.json
|
|
26
|
+
│ ├── graph_analysis.json
|
|
27
|
+
│ ├── SUMMARY.md
|
|
28
|
+
│ ├── models/
|
|
29
|
+
│ │ ├── _index.json
|
|
30
|
+
│ │ └── <unit filename>.json
|
|
31
|
+
│ └── flows/
|
|
32
|
+
│ ├── flow_index.json
|
|
33
|
+
│ └── <flow filename>.json
|
|
34
|
+
├── woods.json # optional configuration artifact
|
|
35
|
+
├── dumps/ # optional semantic-store snapshots
|
|
36
|
+
└── ... # operational state and locks
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
The configured root can differ between a container and the host. Resolve the
|
|
40
|
+
relative payload against the root visible to the reader, not the writer's path.
|
|
41
|
+
|
|
42
|
+
`generation.json` is one JSON object, for example:
|
|
43
|
+
|
|
44
|
+
```json
|
|
45
|
+
{"number":42,"token":"978e69b524ac74fa","updated_at":"2026-09-15T12:00:00Z","reason":"incremental","payload":"payloads/gen-42"}
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
| Field | Contract |
|
|
49
|
+
|---|---|
|
|
50
|
+
| `number` | Positive integer, incremented on publication. It is local to this index; deleting/recreating the index can restart it. |
|
|
51
|
+
| `token` | Opaque string, changed on each publication. Compare it as well as `number`; do not depend on its length or encoding. |
|
|
52
|
+
| `updated_at` | ISO8601 publication timestamp, distinct from source modification times. |
|
|
53
|
+
| `reason` | String or null describing the publication, such as `full`, `incremental`, or `refresh`. Treat values as extensible. |
|
|
54
|
+
| `payload` | Relative directory name for this generation. Current writers use `payloads/gen-N`; follow the field instead of constructing it. Omitted/null means a flat layout. |
|
|
55
|
+
|
|
56
|
+
Reject an absolute payload path or one whose resolved real path escapes the index
|
|
57
|
+
root, including through a symlink. Unknown fields can be ignored. A missing
|
|
58
|
+
pointer can mean an older flat index or no index at all; it is not evidence that a
|
|
59
|
+
payload directory is published. An existing malformed pointer or a missing named
|
|
60
|
+
payload is an error to investigate, not permission to silently serve root files.
|
|
61
|
+
Some gem readers have permissive flat fallback for these failures; the publication
|
|
62
|
+
gates below deliberately fail more strictly.
|
|
63
|
+
|
|
64
|
+
### Payload artifacts
|
|
65
|
+
|
|
66
|
+
| Artifact | Presence and meaning |
|
|
67
|
+
|---|---|
|
|
68
|
+
| `manifest.json` | Required for a complete structural publication. Counts by extractor directory, totals, extraction timestamp and provenance. Optional fields vary by writer/version; see [writer provenance](PUBLISHED_INDEX.md#manifest-writer-provenance). |
|
|
69
|
+
| `source_inputs.json` | Versioned source-input identities and per-consumer provenance for this exact generation; old indexes may omit it. See [source freshness](SOURCE_FRESHNESS.md). |
|
|
70
|
+
| `dependency_graph.json` | Required for a complete structural publication. Typed graph data; an empty graph is valid. |
|
|
71
|
+
| `<type>/_index.json` and unit JSON | Present for extracted families. `_index.json` is an array of unit summaries; an empty array is valid. Disabled/unavailable families may be absent. Do not infer completeness from a fixed count of directories. |
|
|
72
|
+
| `graph_analysis.json` | Derived graph analysis when produced. Treat absence as unavailable analysis, not an empty or corrupt unit index. |
|
|
73
|
+
| `flows/flow_index.json` and flow documents | Optional precomputed flows. Use the index's relative paths; do not invent flow filenames. |
|
|
74
|
+
| `SUMMARY.md` | Generated human-readable summary, not a machine schema. |
|
|
75
|
+
|
|
76
|
+
Unit summaries identify units, but do not carry an artifact filename. Their
|
|
77
|
+
`file_path` names the application's source file, not the unit JSON file. To avoid
|
|
78
|
+
reimplementing filename normalization, scan JSON files in a needed type directory
|
|
79
|
+
(excluding `_index.json`) and match the JSON `identifier` **and** `type`. The same
|
|
80
|
+
identifier can exist in multiple types. The full unit fields are documented in
|
|
81
|
+
[Extractor reference](EXTRACTOR_REFERENCE.md#extractedunit-field-reference).
|
|
82
|
+
|
|
83
|
+
The graph's `nodes`, typed variants, forward/reverse relationships and relationship
|
|
84
|
+
metadata belong to the graph format; preserve them when transporting the index.
|
|
85
|
+
Extractor directories can contain multiple unit types: `graphql/` contains
|
|
86
|
+
`graphql_type`, `graphql_mutation`, `graphql_resolver`, and `graphql_query`, while
|
|
87
|
+
`rails_source/` contains `rails_source` and `gem_source`. Validate and retain the
|
|
88
|
+
artifact's actual type instead of deriving it by singularizing the directory.
|
|
89
|
+
|
|
90
|
+
Do not flatten typed variants into a single node per textual identifier. The
|
|
91
|
+
static Woods self-map has the same publication envelope but different type
|
|
92
|
+
families and `manifest.provenance.mode`; it is not Rails runtime evidence.
|
|
93
|
+
|
|
94
|
+
### Semantic graph validation
|
|
95
|
+
|
|
96
|
+
`woods:validate` checks raw graph data against the unit indexes and artifacts in
|
|
97
|
+
one pinned published generation. This validation is unreleased after
|
|
98
|
+
`2.0.0.beta2`. It checks the shapes of `nodes`, `edges`, `reverse`, `file_map`,
|
|
99
|
+
`type_index`, and optional `variants`; typed identities must be unique and agree
|
|
100
|
+
with the actual indexed units. Forward sources must exist. Reverse, file, and
|
|
101
|
+
type memberships must match the union of primary and variant contributions.
|
|
102
|
+
When present, `reverse_via` must preserve the typed forward records and their
|
|
103
|
+
relationship attributes, including duplicate-record multiplicity.
|
|
104
|
+
|
|
105
|
+
Cycles, recursion, shared paths, cross-type identifier collisions, nil file paths,
|
|
106
|
+
and legacy bare-string edges are valid. Absent optional legacy fields remain
|
|
107
|
+
legal. A target absent from the nodes is an unresolved reference, not automatically
|
|
108
|
+
corruption: current metadata cannot distinguish an intentional external target
|
|
109
|
+
from an internal node and unit that are both missing. Validation cannot certify a
|
|
110
|
+
unique target type when several types share its name. A missing node whose typed
|
|
111
|
+
unit remains indexed is detectable and is an error.
|
|
112
|
+
|
|
113
|
+
The checker does not repair data, rerun extraction, or become a publication gate.
|
|
114
|
+
It reports identity/path diagnostics through the existing validation report and
|
|
115
|
+
nonzero task exit. It also supports the Woods static source map's explicitly
|
|
116
|
+
recorded type families; that does not give the map Rails runtime fidelity.
|
|
117
|
+
See [running validation from Ruby](PUBLISHED_INDEX.md#validate-a-published-generation)
|
|
118
|
+
and [semantic error recovery](TROUBLESHOOTING.md#semantic-graph-validation-errors).
|
|
119
|
+
|
|
120
|
+
### File profiles and file membership
|
|
121
|
+
|
|
122
|
+
`file_map[path]` lists units associated with a source file, including whole-file
|
|
123
|
+
profiles. It does not promise that every identifier names a Ruby constant. New
|
|
124
|
+
writers mark graph nodes of types `caching`, `configuration`, `test_mapping`,
|
|
125
|
+
`rails_source`, and `gem_source` with `"kind": "file_profile"`. The same field
|
|
126
|
+
appears on non-primary typed variants.
|
|
127
|
+
|
|
128
|
+
For example, `app/controllers/things_controller.rb` can map to both
|
|
129
|
+
`ThingsController` and a caching unit named `app/controllers/things_controller.rb`.
|
|
130
|
+
Inspect each typed node's `kind` to distinguish the profile; both retain their
|
|
131
|
+
file membership so a source edit refreshes both units. Do not classify units by
|
|
132
|
+
comparing the identifier with the path, and do not interpret a missing `kind` as
|
|
133
|
+
proof that the unit names a constant.
|
|
134
|
+
|
|
135
|
+
Older indexes omit the marker. Woods derives it from these known extractor types
|
|
136
|
+
when loading and republishing a graph, including unchanged incremental nodes.
|
|
137
|
+
Raw consumers of older indexes must treat absent markers as unclassified or use
|
|
138
|
+
the documented type list. Identifiers, `file_map`, and type membership retain
|
|
139
|
+
their existing shapes; older readers can ignore `kind`.
|
|
140
|
+
|
|
141
|
+
### Reverse relationship records
|
|
142
|
+
|
|
143
|
+
New writers add `reverse_via` to `dependency_graph.json`. Each target identifier
|
|
144
|
+
maps to its incoming relationship records, including the owning source type:
|
|
145
|
+
|
|
146
|
+
```json
|
|
147
|
+
{
|
|
148
|
+
"reverse": { "Gadget": ["Widget"] },
|
|
149
|
+
"reverse_via": {
|
|
150
|
+
"Gadget": [{ "source": "Widget", "source_type": "service", "via": "render" }]
|
|
151
|
+
}
|
|
152
|
+
}
|
|
153
|
+
```
|
|
154
|
+
|
|
155
|
+
The existing `reverse` arrays retain their bare identifiers. `reverse_via` includes
|
|
156
|
+
edges from primary nodes and typed variants; several records can share a source
|
|
157
|
+
and target while differing in type, relationship, or association attributes.
|
|
158
|
+
Optional `through`, `through_db`, and `disable_joins` values match the forward
|
|
159
|
+
edge. A `null` relationship means unknown legacy evidence. Target type remains
|
|
160
|
+
unresolved when the target identifier belongs to multiple types; the source type
|
|
161
|
+
does not resolve that ambiguity.
|
|
162
|
+
|
|
163
|
+
These are recorded dependencies, not proof that changing a target breaks every
|
|
164
|
+
source. For example, `factory_for` and migration `reference` edges describe
|
|
165
|
+
different relationships from a runtime `render` edge. Consumers can inspect one
|
|
166
|
+
target bucket without scanning the whole forward graph. Buckets and records are
|
|
167
|
+
deterministically ordered, but consumers should treat the ordering as incidental.
|
|
168
|
+
|
|
169
|
+
Older graphs omit `reverse_via`; absence means relationship detail must be derived
|
|
170
|
+
from forward edges and variants, not that there are no dependents. Woods rebuilds
|
|
171
|
+
this derived index from forward evidence when loading and republishing a graph.
|
|
172
|
+
Existing readers can ignore the additive field; a subsequent changed extraction
|
|
173
|
+
or full run publishes it. A no-op leaves the previous generation unchanged.
|
|
174
|
+
|
|
175
|
+
### What is outside this structural snapshot
|
|
176
|
+
|
|
177
|
+
`woods.json`, `dumps/`, embedding checkpoints, temporal snapshots, watch status,
|
|
178
|
+
pending paths, MCP task records, exporter state and extraction/startup lock files
|
|
179
|
+
have separate lifecycles. Their presence is configuration-dependent. Do not glob
|
|
180
|
+
them into a structural payload or assume `generation.json` commits them together.
|
|
181
|
+
A semantic/MCP deployment also needs its configured stores and artifacts; copying
|
|
182
|
+
a structural payload alone does not clone that deployment.
|
|
183
|
+
|
|
184
|
+
`<output>/.source-inputs.key` is private operational state used to verify source
|
|
185
|
+
identities. Never include it in a published/exported structural snapshot; without
|
|
186
|
+
it, a recipient can still read the index but source freshness is unknown.
|
|
187
|
+
|
|
188
|
+
Temporary filenames, abandoned payloads, lock sidecars, retained-directory counts,
|
|
189
|
+
and summary formatting are implementation details. Never remove or modify locks
|
|
190
|
+
to make a reader proceed.
|
|
191
|
+
|
|
192
|
+
## Atomic publication is not indefinite retention
|
|
193
|
+
|
|
194
|
+
For a payload publication, Woods writes the payload, flushes it, and atomically
|
|
195
|
+
replaces `generation.json` last. Capturing that pointer once selects a complete,
|
|
196
|
+
immutable structural generation. Opening each file through a freshly reread pointer
|
|
197
|
+
can mix generations and defeats that guarantee. A failed/no-op run does not advance
|
|
198
|
+
the pointer. [Durability details](PUBLISHED_INDEX.md#durability-the-pointer-is-the-commit-point)
|
|
199
|
+
explain the flush boundary.
|
|
200
|
+
|
|
201
|
+
Retention can delete an older payload after the reader selects it. The default
|
|
202
|
+
retains three generations, not three minutes. To keep a multi-file read or copy
|
|
203
|
+
safe from Woods retention:
|
|
204
|
+
|
|
205
|
+
1. Read and validate the pointer; resolve its payload inside the root.
|
|
206
|
+
2. Open that payload's existing `manifest.json` **read-only** and take a shared
|
|
207
|
+
advisory `flock` on that open file. Do not create a new lock file.
|
|
208
|
+
3. Re-read the pointer and confirm it is unchanged; verify the pathname still
|
|
209
|
+
names the open manifest inode. A pruner may have won between steps 1 and 2.
|
|
210
|
+
4. Keep the handle/lock open for **all** reads or the complete copy. Use the single
|
|
211
|
+
captured payload path throughout. Once pinned, a later pointer advance is fine.
|
|
212
|
+
5. Close the handle when done. On a race, discard partial results and retry the
|
|
213
|
+
entire operation from step 1, with a bounded retry count.
|
|
214
|
+
|
|
215
|
+
Woods retention attempts a nonblocking **exclusive** `flock` on that same manifest
|
|
216
|
+
before deletion and skips a payload held by readers. These are advisory filesystem
|
|
217
|
+
locks, not Woods' writer-coordination locks. Ordinary readers need no exclusive
|
|
218
|
+
lock and need not take `extraction.lock` or its guard. A lock-free reader must
|
|
219
|
+
accept disappearance and restart the whole read; it must never silently replace
|
|
220
|
+
missing files with files from another generation.
|
|
221
|
+
|
|
222
|
+
This protocol protects against cooperating retention only. It does not protect
|
|
223
|
+
against `woods:clean`, manual deletion, index replacement, or a filesystem that
|
|
224
|
+
does not coordinate `flock` across its clients. Stop writers/cleanup and read an
|
|
225
|
+
immutable snapshot if that protection is unavailable. Flat indexes (including
|
|
226
|
+
full-extraction fallback when a payload cannot be created) have individually
|
|
227
|
+
replaced files, not multi-file atomicity; read/copy them only with writers stopped.
|
|
228
|
+
|
|
229
|
+
## Bash and jq: read one pinned generation
|
|
230
|
+
|
|
231
|
+
Requires Bash, jq, GNU `realpath`, util-linux `flock`, and Linux `/proc`. Save as
|
|
232
|
+
`read-woods.sh`, then run `bash read-woods.sh '/path with spaces/tmp/woods'`.
|
|
233
|
+
The final command reads manifest and graph under the same lock. Substitute other
|
|
234
|
+
reads or a complete copy **inside** the script before it exits. Printing a path
|
|
235
|
+
and consuming it after the script exits does not keep it pinned.
|
|
236
|
+
|
|
237
|
+
Exit 75 means a possible publication/retention race: retry the whole script a
|
|
238
|
+
bounded number of times (for example three), discarding any previous output.
|
|
239
|
+
Other nonzero exits need investigation. A persistent 75 can mean a broken index.
|
|
240
|
+
|
|
241
|
+
```bash
|
|
242
|
+
#!/usr/bin/env bash
|
|
243
|
+
set -euo pipefail
|
|
244
|
+
fail() { echo "$*" >&2; exit 1; }
|
|
245
|
+
retry() { echo "$*; retry the whole read" >&2; exit 75; }
|
|
246
|
+
root=$(realpath -e -- "${1:?provide the index root}")
|
|
247
|
+
[[ -d "$root" ]] || fail 'Index root is not a directory'
|
|
248
|
+
[[ -f "$root/generation.json" ]] || fail 'Missing pointer: legacy or unpublished index'
|
|
249
|
+
marker=$(cat -- "$root/generation.json") || retry 'Cannot read pointer'
|
|
250
|
+
jq -e '
|
|
251
|
+
if type != "object" then false
|
|
252
|
+
elif (.number | type) != "number" then false
|
|
253
|
+
else (.number >= 1 and (.number | floor) == .number)
|
|
254
|
+
and (.token | type == "string" and length > 0)
|
|
255
|
+
and (.payload | type == "string" and length > 0)
|
|
256
|
+
and (.payload | explode | all(. >= 32 and . != 127))
|
|
257
|
+
end
|
|
258
|
+
' <<<"$marker" >/dev/null || fail 'Invalid pointer or flat layout'
|
|
259
|
+
relative=$(jq -r '.payload' <<<"$marker")
|
|
260
|
+
[[ "$relative" != /* ]] || fail 'Absolute payload path'
|
|
261
|
+
payload=$(realpath -e -- "$root/$relative") || retry 'Missing payload'
|
|
262
|
+
[[ -d "$payload" && "$payload" == "${root%/}/"* && "$payload" != "$root" ]] \
|
|
263
|
+
|| fail 'Payload must resolve inside the index root'
|
|
264
|
+
exec 9< "$payload/manifest.json" || retry 'Missing manifest'
|
|
265
|
+
flock -sn 9 || retry 'Cannot pin manifest'
|
|
266
|
+
current=$(cat -- "$root/generation.json") || retry 'Pointer disappeared'
|
|
267
|
+
[[ "$current" == "$marker" && "$payload/manifest.json" -ef /proc/self/fd/9 ]] \
|
|
268
|
+
|| retry 'Publication changed or retention removed the payload'
|
|
269
|
+
jq -n --argjson generation "$marker" \
|
|
270
|
+
--slurpfile manifest "$payload/manifest.json" \
|
|
271
|
+
--slurpfile graph "$payload/dependency_graph.json" \
|
|
272
|
+
'if ($manifest | length) == 1 and ($manifest[0] | type) == "object"
|
|
273
|
+
and ($graph | length) == 1 and ($graph[0] | type) == "object"
|
|
274
|
+
then {generation: $generation, manifest: $manifest[0], dependency_graph: $graph[0]}
|
|
275
|
+
else error("Manifest and graph must each contain exactly one JSON object") end'
|
|
276
|
+
# Descriptor 9 closes on exit, releasing the retention pin.
|
|
277
|
+
```
|
|
278
|
+
|
|
279
|
+
## Python: keep the pin while using the payload
|
|
280
|
+
|
|
281
|
+
Requires Python 3.9+ on Unix with working `fcntl.flock`. Save as `read_woods.py`
|
|
282
|
+
and run `python3 read_woods.py '/path with spaces/tmp/woods'`. The context manager
|
|
283
|
+
retries acquisition three times. Errors during the read/copy propagate: discard
|
|
284
|
+
partial output before retrying the entire operation. The scripts assume a trusted
|
|
285
|
+
Woods-owned index; path containment is not a sandbox for hostile filesystem changes.
|
|
286
|
+
|
|
287
|
+
```python
|
|
288
|
+
import contextlib
|
|
289
|
+
import fcntl
|
|
290
|
+
import json
|
|
291
|
+
import os
|
|
292
|
+
from pathlib import Path
|
|
293
|
+
import sys
|
|
294
|
+
import time
|
|
295
|
+
|
|
296
|
+
|
|
297
|
+
@contextlib.contextmanager
|
|
298
|
+
def pinned_payload(index_root):
|
|
299
|
+
root = Path(index_root).resolve(strict=True)
|
|
300
|
+
pointer = root / "generation.json"
|
|
301
|
+
if not pointer.is_file():
|
|
302
|
+
raise ValueError("Missing pointer: legacy or unpublished index")
|
|
303
|
+
for attempt in range(3):
|
|
304
|
+
handle = None
|
|
305
|
+
try:
|
|
306
|
+
raw = pointer.read_bytes()
|
|
307
|
+
marker = json.loads(raw)
|
|
308
|
+
if not isinstance(marker, dict):
|
|
309
|
+
raise ValueError("Pointer must be an object")
|
|
310
|
+
number, token, name = (marker.get(k) for k in ("number", "token", "payload"))
|
|
311
|
+
if type(number) is not int or number < 1 or not isinstance(token, str) or not token:
|
|
312
|
+
raise ValueError("Invalid generation identity")
|
|
313
|
+
if not isinstance(name, str) or not name or Path(name).is_absolute():
|
|
314
|
+
raise ValueError("Invalid payload path or flat layout")
|
|
315
|
+
if any(ord(char) < 32 or ord(char) == 127 for char in name):
|
|
316
|
+
raise ValueError("Control character in payload path")
|
|
317
|
+
payload = (root / name).resolve(strict=True)
|
|
318
|
+
if not payload.is_dir() or root not in payload.parents:
|
|
319
|
+
raise ValueError("Payload must resolve inside the index root")
|
|
320
|
+
manifest = payload / "manifest.json"
|
|
321
|
+
handle = manifest.open("rb")
|
|
322
|
+
fcntl.flock(handle, fcntl.LOCK_SH | fcntl.LOCK_NB)
|
|
323
|
+
if pointer.read_bytes() != raw or not os.path.samestat(os.fstat(handle.fileno()), manifest.stat()):
|
|
324
|
+
raise BlockingIOError("Publication changed or payload was removed")
|
|
325
|
+
except (FileNotFoundError, BlockingIOError):
|
|
326
|
+
if handle is not None:
|
|
327
|
+
handle.close()
|
|
328
|
+
if attempt == 2:
|
|
329
|
+
raise
|
|
330
|
+
time.sleep(0.05)
|
|
331
|
+
except BaseException:
|
|
332
|
+
if handle is not None:
|
|
333
|
+
handle.close()
|
|
334
|
+
raise
|
|
335
|
+
else:
|
|
336
|
+
break
|
|
337
|
+
try:
|
|
338
|
+
yield marker, payload
|
|
339
|
+
finally:
|
|
340
|
+
handle.close()
|
|
341
|
+
|
|
342
|
+
|
|
343
|
+
if __name__ == "__main__":
|
|
344
|
+
with pinned_payload(sys.argv[1]) as (generation, payload):
|
|
345
|
+
manifest = json.loads((payload / "manifest.json").read_text(encoding="utf-8"))
|
|
346
|
+
graph = json.loads((payload / "dependency_graph.json").read_text(encoding="utf-8"))
|
|
347
|
+
if not isinstance(manifest, dict) or not isinstance(graph, dict):
|
|
348
|
+
raise ValueError("Manifest and graph must each contain exactly one JSON object")
|
|
349
|
+
# Read more artifacts, or copy the whole payload, before leaving this block.
|
|
350
|
+
print(json.dumps({"generation": generation, "manifest": manifest, "dependency_graph": graph}))
|
|
351
|
+
```
|
|
352
|
+
|
|
353
|
+
## Shipping a structural snapshot
|
|
354
|
+
|
|
355
|
+
While the pin is held, copy the selected payload into an unpublished staging
|
|
356
|
+
location. Preserve the payload's relative path and pair it with the **captured**
|
|
357
|
+
`generation.json`, not a later pointer reread from the live index. For example,
|
|
358
|
+
a captured `payloads/gen-42` must still resolve to that directory in the exported
|
|
359
|
+
root. Validate the copied manifest and graph, then publish the staged copy as a
|
|
360
|
+
whole, or upload its payload first and switch the destination pointer last.
|
|
361
|
+
Do not upload `generation.json` first or advertise a failed/partial copy.
|
|
362
|
+
|
|
363
|
+
The retention lock is local coordination; do not copy its file descriptor or
|
|
364
|
+
writer lock files. Copying `manifest.json` normally copies content, which is correct.
|
|
365
|
+
Objects in a destination store do not inherit the source's locking guarantees;
|
|
366
|
+
protect any later destination pruning with that destination's own reader protocol.
|
|
367
|
+
|
|
368
|
+
## Compatibility within Woods 2.x
|
|
369
|
+
|
|
370
|
+
Consumers may rely on the pointer field meanings, relative payload resolution,
|
|
371
|
+
the complete-payload publication boundary, manifest/graph locations, and JSON unit
|
|
372
|
+
identity described here. Additive JSON fields, new extractor families, new reason
|
|
373
|
+
values, optional artifacts and additional root-level state may appear in 2.x.
|
|
374
|
+
Ignore unknown fields and inspect available types instead of hardcoding a directory
|
|
375
|
+
count. Existing field meanings and required-file locations are compatibility
|
|
376
|
+
surfaces; incompatible changes require an explicit migration contract.
|
|
377
|
+
|
|
378
|
+
Do not bind to temporary names, a fixed token length, directory listing order,
|
|
379
|
+
retention count, internal lock sidecars, or Markdown summary formatting. Check the
|
|
380
|
+
installed Woods version and its release's documentation when consuming older
|
|
381
|
+
prereleases; this guide describes the current source contract, not a promise that
|
|
382
|
+
every prerelease contains every optional field or retention improvement.
|
data/docs/INTERNALS.md
CHANGED
|
@@ -145,6 +145,11 @@ The `DependencyGraph` is a directed graph where nodes are `ExtractedUnit` identi
|
|
|
145
145
|
- **Forward edges** (`@edges`): what each unit depends on, populated when units are registered
|
|
146
146
|
- **Reverse edges** (`@reverse`): what depends on each unit, built during registration and in the resolve phase
|
|
147
147
|
|
|
148
|
+
The published graph also carries additive `reverse_via` target buckets with
|
|
149
|
+
typed source identities, relationship labels and association attributes. The
|
|
150
|
+
bare-name `reverse` map remains compatible. See the
|
|
151
|
+
[reverse relationship format](INDEX_LAYOUT.md#reverse-relationship-records).
|
|
152
|
+
|
|
148
153
|
```ruby
|
|
149
154
|
graph = DependencyGraph.new
|
|
150
155
|
graph.register(user_unit) # adds User to nodes, adds User→Order edge (from belongs_to)
|
|
@@ -175,7 +180,7 @@ Scores feed into the retrieval ranker as one signal in the final ranking formula
|
|
|
175
180
|
| **Cycles** | Circular dependencies, A→B→C→A. Detected via DFS, and capped: `graph_cycle_limit` (default 500) bounds how many are enumerated and `graph_cycle_max_length` (default 50) skips one longer than that. Either cap firing sets `stats.cycle_limit_reached`. Set both to `nil` for exhaustive enumeration. |
|
|
176
181
|
| **Bridges** | Edges whose removal would disconnect the graph, high-risk structural connections |
|
|
177
182
|
| **Cross-database edges** | Association or foreign-key edges whose two ends resolve to different databases. A `has_many :through` is reported as `join_through_across_databases` when `disable_joins` is false and `from_db`, `through_db` (the join model's database), or `to_db` disagree. A foreign key never resolves to an owner in the source database, even when another database also claims the table; when every owner sits elsewhere and they span more than one database, the entry comes back with `to: nil` and an `ambiguous_owners` list instead of guessing. Read from graph node and edge attributes, so full and incremental runs agree. Scoped to primary nodes (units registered in the graph), not variants. |
|
|
178
|
-
| **Volatile dependencies** | Edges that point at a unit changing at least `volatile_dependency_ratio` times more often than the dependent (POODR: depend on things that change less often than you do). Dependencies with fewer than 5 commits or a `new` change frequency are skipped. Ranked by the dependency's PageRank;
|
|
183
|
+
| **Volatile dependencies** | Edges that point at a unit changing at least `volatile_dependency_ratio` times more often than the dependent (POODR: depend on things that change less often than you do). Dependencies with fewer than 5 commits or a `new` change frequency are skipped. Ranked by the dependency's PageRank; an optional `volatile_dependency_limit_per_target` selects edges per typed dependency before the persisted top 20, while `stats.volatile_dependency_count` reports the full qualifying count and `stats.volatile_dependencies_limit` reports the cap. |
|
|
179
184
|
| **Undeclared package edges** | Edges that cross a Packwerk package boundary the source package does not list in `dependencies`. Membership comes from each unit's `package` node attribute, declarations from the package unit's own `package_dependency` edges. Woods reports the boundary; enforcement stays with `packwerk check` / `pks check`. |
|
|
180
185
|
|
|
181
186
|
Analysis results are written to `graph_analysis.json` and surfaced in `SUMMARY.md`.
|
|
@@ -194,7 +199,7 @@ persisted rather than every hub in the graph.
|
|
|
194
199
|
| `cycles` | `stats.cycle_count`, `stats.cycle_limit_reached` | yes, `graph_cycle_limit` cycles of at most `graph_cycle_max_length` nodes | the count is the persisted array; the flag says whether either cap fired |
|
|
195
200
|
| `bridges` | none | yes, top 10 by score | not counted |
|
|
196
201
|
| `cross_database_edges` | `stats.cross_database_edge_count` | no | every crossing edge |
|
|
197
|
-
| `volatile_dependencies` | `stats.volatile_dependency_count`, `stats.volatile_dependencies_limit` | yes, top 20 by the dependency's PageRank |
|
|
202
|
+
| `volatile_dependencies` | `stats.volatile_dependency_count`, `stats.volatile_dependencies_limit`; optional `stats.volatile_dependencies_limit_per_target`, `stats.volatile_dependency_reported_count` | yes, optional per-target cap followed by top 20 by the dependency's PageRank | count is every qualifying edge before caps; reported count is the final array length and appears only when the per-target cap is enabled |
|
|
198
203
|
| `undeclared_package_edges` | `stats.undeclared_package_edge_count` | no | every undeclared crossing |
|
|
199
204
|
|
|
200
205
|
A capped section means the array on disk is a page, not the population.
|