woods 2.0.0.beta2 → 2.0.0.beta4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (233) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +339 -1
  3. data/CONTRIBUTING.md +188 -12
  4. data/README.md +93 -174
  5. data/SECURITY.md +9 -6
  6. data/docs/AGENT_GUIDE.md +109 -8
  7. data/docs/AGENT_SETUP.md +98 -7
  8. data/docs/BACKEND_MATRIX.md +25 -0
  9. data/docs/CLIENT_HOOKS.md +111 -0
  10. data/docs/CONFIGURATION_REFERENCE.md +267 -16
  11. data/docs/CONSOLE_MCP_SETUP.md +80 -7
  12. data/docs/DOCKER_SETUP.md +22 -3
  13. data/docs/EVALUATION.md +464 -1
  14. data/docs/EXTRACTOR_REFERENCE.md +45 -6
  15. data/docs/FAQ.md +11 -12
  16. data/docs/GETTING_STARTED.md +17 -5
  17. data/docs/INCREMENTAL_EXTRACTION.md +147 -7
  18. data/docs/INDEX_LAYOUT.md +382 -0
  19. data/docs/INTERNALS.md +7 -2
  20. data/docs/MCP_SERVERS.md +276 -5
  21. data/docs/MCP_TOOL_COOKBOOK.md +37 -22
  22. data/docs/MCP_WORKTREE_SETUP.md +43 -83
  23. data/docs/NOTION_INTEGRATION.md +13 -0
  24. data/docs/OBSIDIAN_INTEGRATION.md +57 -9
  25. data/docs/PUBLISHED_INDEX.md +72 -0
  26. data/docs/README.md +7 -0
  27. data/docs/RETRIEVAL_GUIDE.md +273 -12
  28. data/docs/RUNTIME_TRACING.md +71 -0
  29. data/docs/SOURCE_FRESHNESS.md +143 -0
  30. data/docs/TROUBLESHOOTING.md +129 -18
  31. data/docs/UNBLOCKED_INTEGRATION.md +25 -0
  32. data/docs/UPGRADING_TO_2.md +48 -22
  33. data/docs/WATCH_DAEMON.md +277 -67
  34. data/exe/woods-agent-config +6 -0
  35. data/exe/woods-extract +5 -0
  36. data/exe/woods-hook-context +6 -0
  37. data/exe/woods-mcp-start +14 -9
  38. data/lib/generators/woods/pgvector_generator.rb +8 -2
  39. data/lib/generators/woods/templates/woods.rb.tt +1 -3
  40. data/lib/tasks/woods.rake +47 -397
  41. data/lib/woods/agent_configuration/applier.rb +135 -0
  42. data/lib/woods/agent_configuration/cli.rb +101 -0
  43. data/lib/woods/agent_configuration/cli_options.rb +29 -0
  44. data/lib/woods/agent_configuration/document.rb +105 -0
  45. data/lib/woods/agent_configuration/error.rb +7 -0
  46. data/lib/woods/agent_configuration/launcher.rb +75 -0
  47. data/lib/woods/agent_configuration/layout.rb +72 -0
  48. data/lib/woods/agent_configuration/managed_section.rb +62 -0
  49. data/lib/woods/agent_configuration/plan.rb +98 -0
  50. data/lib/woods/agent_configuration/plan_diff.rb +38 -0
  51. data/lib/woods/agent_configuration/planned_files.rb +61 -0
  52. data/lib/woods/agent_configuration/planner.rb +63 -0
  53. data/lib/woods/agent_configuration/planner_validation.rb +77 -0
  54. data/lib/woods/agent_configuration/preflight.rb +100 -0
  55. data/lib/woods/agent_configuration/recovery.rb +49 -0
  56. data/lib/woods/ast/node.rb +2 -0
  57. data/lib/woods/ast/parser.rb +38 -5
  58. data/lib/woods/builder.rb +21 -5
  59. data/lib/woods/cache/cache_middleware.rb +28 -7
  60. data/lib/woods/cache/cache_store.rb +4 -5
  61. data/lib/woods/change_set.rb +5 -4
  62. data/lib/woods/console/credential_index.rb +20 -2
  63. data/lib/woods/console/credential_scanner.rb +18 -17
  64. data/lib/woods/console/credential_scanner_registry.rb +36 -0
  65. data/lib/woods/console/dispatch_pipeline.rb +7 -0
  66. data/lib/woods/console/embedded_executor.rb +32 -10
  67. data/lib/woods/console/encrypted_credential_snapshot.rb +16 -0
  68. data/lib/woods/console/rack_middleware.rb +22 -13
  69. data/lib/woods/console/server.rb +18 -16
  70. data/lib/woods/console/sql_noise_stripper.rb +9 -7
  71. data/lib/woods/console/sql_table_scanner.rb +47 -7
  72. data/lib/woods/console/sql_validator.rb +49 -9
  73. data/lib/woods/console/sqlite_read_guard.rb +46 -0
  74. data/lib/woods/coordination/pipeline_lock.rb +3 -2
  75. data/lib/woods/dependency_graph.rb +65 -13
  76. data/lib/woods/embedding/corpus.rb +94 -0
  77. data/lib/woods/embedding/indexer.rb +114 -60
  78. data/lib/woods/embedding/openai.rb +17 -6
  79. data/lib/woods/evaluation/ablation_executor.rb +6 -1
  80. data/lib/woods/evaluation/ablation_timed_executor.rb +22 -4
  81. data/lib/woods/export/typed_reader.rb +56 -0
  82. data/lib/woods/extractor.rb +277 -149
  83. data/lib/woods/extractors/action_cable_extractor.rb +3 -1
  84. data/lib/woods/extractors/behavioral_profile.rb +9 -7
  85. data/lib/woods/extractors/caching_extractor.rb +3 -1
  86. data/lib/woods/extractors/concern_extractor.rb +64 -6
  87. data/lib/woods/extractors/configuration_extractor.rb +7 -3
  88. data/lib/woods/extractors/controller_extractor.rb +13 -4
  89. data/lib/woods/extractors/database_view_extractor.rb +3 -1
  90. data/lib/woods/extractors/declared_parent.rb +55 -0
  91. data/lib/woods/extractors/decorator_extractor.rb +3 -1
  92. data/lib/woods/extractors/engine_extractor.rb +3 -1
  93. data/lib/woods/extractors/event_extractor.rb +4 -2
  94. data/lib/woods/extractors/factory_extractor.rb +3 -1
  95. data/lib/woods/extractors/graphql_extractor.rb +10 -13
  96. data/lib/woods/extractors/i18n_extractor.rb +3 -1
  97. data/lib/woods/extractors/job_extractor.rb +6 -19
  98. data/lib/woods/extractors/lib_extractor.rb +13 -9
  99. data/lib/woods/extractors/mailer_extractor.rb +26 -15
  100. data/lib/woods/extractors/manager_extractor.rb +3 -1
  101. data/lib/woods/extractors/method_parameters.rb +53 -0
  102. data/lib/woods/extractors/middleware_argument.rb +65 -0
  103. data/lib/woods/extractors/middleware_extractor.rb +9 -3
  104. data/lib/woods/extractors/migration_extractor.rb +3 -1
  105. data/lib/woods/extractors/model_extractor.rb +26 -34
  106. data/lib/woods/extractors/package_extractor.rb +24 -4
  107. data/lib/woods/extractors/phlex_extractor.rb +3 -1
  108. data/lib/woods/extractors/policy_extractor.rb +3 -1
  109. data/lib/woods/extractors/poro_extractor.rb +13 -9
  110. data/lib/woods/extractors/pundit_extractor.rb +3 -1
  111. data/lib/woods/extractors/rails_source_extractor.rb +4 -2
  112. data/lib/woods/extractors/rake_task_extractor.rb +4 -2
  113. data/lib/woods/extractors/route_extractor.rb +3 -1
  114. data/lib/woods/extractors/route_helper_resolver.rb +10 -33
  115. data/lib/woods/extractors/scheduled_job_extractor.rb +41 -15
  116. data/lib/woods/extractors/serializer_extractor.rb +4 -2
  117. data/lib/woods/extractors/service_extractor.rb +3 -1
  118. data/lib/woods/extractors/shared_dependency_scanner.rb +2 -2
  119. data/lib/woods/extractors/shared_utility_methods.rb +48 -19
  120. data/lib/woods/extractors/source_nesting.rb +1 -1
  121. data/lib/woods/extractors/state_machine_extractor.rb +3 -1
  122. data/lib/woods/extractors/test_mapping_extractor.rb +3 -1
  123. data/lib/woods/extractors/validator_extractor.rb +3 -1
  124. data/lib/woods/extractors/view_component_extractor.rb +3 -1
  125. data/lib/woods/extractors/view_template_extractor.rb +3 -1
  126. data/lib/woods/gem_mapper.rb +2 -0
  127. data/lib/woods/git_history.rb +116 -0
  128. data/lib/woods/graph_analyzer.rb +35 -6
  129. data/lib/woods/hooks/context_cli.rb +54 -0
  130. data/lib/woods/hooks/context_event.rb +88 -0
  131. data/lib/woods/hooks/context_hint.rb +73 -0
  132. data/lib/woods/hooks/context_impact.rb +77 -0
  133. data/lib/woods/hooks/context_output.rb +47 -0
  134. data/lib/woods/hooks/context_state.rb +102 -0
  135. data/lib/woods/hooks/refresh.rb +79 -0
  136. data/lib/woods/hooks/rule_projection.rb +78 -0
  137. data/lib/woods/input_rules.rb +19 -0
  138. data/lib/woods/mcp/bearer_auth.rb +22 -13
  139. data/lib/woods/mcp/bootstrapper.rb +79 -4
  140. data/lib/woods/mcp/config_resolver.rb +2 -1
  141. data/lib/woods/mcp/index_reader.rb +334 -162
  142. data/lib/woods/mcp/initialization_guidance.rb +27 -0
  143. data/lib/woods/mcp/origin_guard.rb +17 -9
  144. data/lib/woods/mcp/published_lexical_retriever.rb +115 -0
  145. data/lib/woods/mcp/renderers/markdown_renderer.rb +22 -9
  146. data/lib/woods/mcp/renderers/plain_renderer.rb +18 -8
  147. data/lib/woods/mcp/search_results.rb +74 -0
  148. data/lib/woods/mcp/server.rb +178 -63
  149. data/lib/woods/mcp/tool_contract.rb +3 -1
  150. data/lib/woods/mcp/tool_response_renderer.rb +41 -0
  151. data/lib/woods/mcp/traversal_evidence.rb +113 -0
  152. data/lib/woods/mcp/traversal_evidence_index.rb +100 -0
  153. data/lib/woods/mcp/traversal_evidence_page.rb +41 -0
  154. data/lib/woods/mcp/traversal_evidence_text.rb +52 -0
  155. data/lib/woods/mcp/traversal_response.rb +22 -0
  156. data/lib/woods/notion/exporter.rb +56 -17
  157. data/lib/woods/obsidian/destination_plan.rb +98 -0
  158. data/lib/woods/obsidian/name_mapper.rb +19 -3
  159. data/lib/woods/obsidian/note_builder.rb +19 -10
  160. data/lib/woods/obsidian/vault_exporter.rb +88 -32
  161. data/lib/woods/operator/pipeline_guard.rb +18 -13
  162. data/lib/woods/path_dispatcher.rb +13 -6
  163. data/lib/woods/payload_store.rb +27 -26
  164. data/lib/woods/published_index/typed_unit_reader.rb +40 -3
  165. data/lib/woods/published_index.rb +2 -2
  166. data/lib/woods/railtie.rb +3 -3
  167. data/lib/woods/railtie_support.rb +12 -12
  168. data/lib/woods/rake_helpers.rb +382 -0
  169. data/lib/woods/resilience/graph_invariant_validator/membership_checks.rb +71 -0
  170. data/lib/woods/resilience/graph_invariant_validator/node_checks.rb +61 -0
  171. data/lib/woods/resilience/graph_invariant_validator/reverse_relationship_checks.rb +46 -0
  172. data/lib/woods/resilience/graph_invariant_validator.rb +119 -0
  173. data/lib/woods/resilience/index_validator/graph_checks.rb +80 -0
  174. data/lib/woods/resilience/index_validator.rb +112 -23
  175. data/lib/woods/retrieval/context_assembler.rb +50 -15
  176. data/lib/woods/retrieval/lexical_assembler.rb +84 -0
  177. data/lib/woods/retrieval/lexical_index.rb +120 -0
  178. data/lib/woods/retrieval/ranker.rb +4 -2
  179. data/lib/woods/retrieval/scope.rb +108 -0
  180. data/lib/woods/retrieval/scoped_graph_store.rb +32 -0
  181. data/lib/woods/retrieval/scoped_vector_store.rb +55 -0
  182. data/lib/woods/retrieval/search_executor.rb +86 -27
  183. data/lib/woods/retrieval/source_evidence.rb +200 -0
  184. data/lib/woods/retriever.rb +98 -22
  185. data/lib/woods/ruby_analyzer/trace_enricher.rb +77 -38
  186. data/lib/woods/session_tracer/file_store.rb +6 -1
  187. data/lib/woods/session_tracer/middleware.rb +10 -12
  188. data/lib/woods/session_tracer/redis_store.rb +22 -6
  189. data/lib/woods/session_tracer/session_flow_assembler.rb +23 -17
  190. data/lib/woods/session_tracer/solid_cache_coordination.rb +6 -4
  191. data/lib/woods/session_tracer/unit_resolver.rb +63 -0
  192. data/lib/woods/source_inputs/consumer_errors.rb +31 -0
  193. data/lib/woods/source_inputs/handoff.rb +102 -0
  194. data/lib/woods/source_inputs/launcher.rb +157 -0
  195. data/lib/woods/source_inputs/manifest.rb +124 -0
  196. data/lib/woods/source_inputs/private_key.rb +55 -0
  197. data/lib/woods/source_inputs/scanner.rb +171 -0
  198. data/lib/woods/source_inputs/scopes.rb +71 -0
  199. data/lib/woods/source_inputs/session.rb +214 -0
  200. data/lib/woods/source_inputs/status.rb +84 -0
  201. data/lib/woods/source_inputs/verifier.rb +107 -0
  202. data/lib/woods/storage/metadata_store.rb +25 -25
  203. data/lib/woods/storage/pgvector.rb +35 -10
  204. data/lib/woods/storage/qdrant.rb +17 -7
  205. data/lib/woods/storage/vector_store.rb +18 -6
  206. data/lib/woods/tasks.rb +3 -2
  207. data/lib/woods/temporal/json_snapshot_store.rb +58 -9
  208. data/lib/woods/unblocked/exporter.rb +59 -70
  209. data/lib/woods/version.rb +1 -1
  210. data/lib/woods/watch/boot_snapshot.rb +52 -0
  211. data/lib/woods/watch/daemon.rb +154 -32
  212. data/lib/woods/watch/listen_watcher.rb +4 -0
  213. data/lib/woods/watch/polling_watcher.rb +5 -1
  214. data/lib/woods/watch/status.rb +20 -15
  215. data/lib/woods/watch/tree_scan.rb +21 -13
  216. data/lib/woods/watch/watcher.rb +4 -1
  217. data/lib/woods.rb +50 -11
  218. data/plugin/.claude-plugin/plugin.json +1 -1
  219. data/plugin/hooks/adapters/normalize.jq +15 -0
  220. data/plugin/hooks/adapters/normalize.rb +63 -0
  221. data/plugin/hooks/hooks.json +20 -0
  222. data/plugin/hooks/woods-context.sh +50 -0
  223. data/plugin/hooks/woods-input-rules.sh +159 -0
  224. data/plugin/hooks/woods-opencode.mjs +65 -0
  225. data/plugin/hooks/woods-post-edit.sh +2 -225
  226. data/plugin/hooks/woods-refresh.sh +260 -0
  227. data/plugin/hooks/woods-session-start.sh +47 -55
  228. data/plugin/skills/woods-agent-enable/SKILL.md +19 -0
  229. data/plugin/skills/woods-diagnose/SKILL.md +319 -1
  230. data/plugin/skills/woods-investigate/SKILL.md +145 -0
  231. data/plugin/skills/woods-mcp-config/SKILL.md +90 -2
  232. data/plugin/skills/woods-setup/SKILL.md +110 -6
  233. metadata +87 -5
@@ -0,0 +1,382 @@
1
+ # Published index layout for shell and Python readers
2
+
3
+ Read `generation.json` at the configured index root, then read every structural
4
+ artifact from the payload it names. Do not require `dependency_graph.json` or
5
+ `manifest.json` at the root, and do not choose the highest directory in `payloads/`.
6
+ A directory can be present before its generation is published.
7
+
8
+ This is the filesystem contract for consumers without the Woods gem. Ruby callers
9
+ can use [Woods::PublishedIndex](PUBLISHED_INDEX.md). These examples require a
10
+ Woods 2.x payload index and a filesystem that supports the shared `flock` protocol
11
+ below. They deliberately reject legacy flat indexes instead of claiming an atomic
12
+ snapshot from them.
13
+
14
+ ## Entry point and files
15
+
16
+ A typical structural index looks like this; optional entries need not exist:
17
+
18
+ ```text
19
+ <index root>/
20
+ ├── generation.json
21
+ ├── payloads/
22
+ │ └── gen-42/
23
+ │ ├── manifest.json
24
+ │ ├── source_inputs.json # optional on older indexes
25
+ │ ├── dependency_graph.json
26
+ │ ├── graph_analysis.json
27
+ │ ├── SUMMARY.md
28
+ │ ├── models/
29
+ │ │ ├── _index.json
30
+ │ │ └── <unit filename>.json
31
+ │ └── flows/
32
+ │ ├── flow_index.json
33
+ │ └── <flow filename>.json
34
+ ├── woods.json # optional configuration artifact
35
+ ├── dumps/ # optional semantic-store snapshots
36
+ └── ... # operational state and locks
37
+ ```
38
+
39
+ The configured root can differ between a container and the host. Resolve the
40
+ relative payload against the root visible to the reader, not the writer's path.
41
+
42
+ `generation.json` is one JSON object, for example:
43
+
44
+ ```json
45
+ {"number":42,"token":"978e69b524ac74fa","updated_at":"2026-09-15T12:00:00Z","reason":"incremental","payload":"payloads/gen-42"}
46
+ ```
47
+
48
+ | Field | Contract |
49
+ |---|---|
50
+ | `number` | Positive integer, incremented on publication. It is local to this index; deleting/recreating the index can restart it. |
51
+ | `token` | Opaque string, changed on each publication. Compare it as well as `number`; do not depend on its length or encoding. |
52
+ | `updated_at` | ISO8601 publication timestamp, distinct from source modification times. |
53
+ | `reason` | String or null describing the publication, such as `full`, `incremental`, or `refresh`. Treat values as extensible. |
54
+ | `payload` | Relative directory name for this generation. Current writers use `payloads/gen-N`; follow the field instead of constructing it. Omitted/null means a flat layout. |
55
+
56
+ Reject an absolute payload path or one whose resolved real path escapes the index
57
+ root, including through a symlink. Unknown fields can be ignored. A missing
58
+ pointer can mean an older flat index or no index at all; it is not evidence that a
59
+ payload directory is published. An existing malformed pointer or a missing named
60
+ payload is an error to investigate, not permission to silently serve root files.
61
+ Some gem readers have permissive flat fallback for these failures; the publication
62
+ gates below deliberately fail more strictly.
63
+
64
+ ### Payload artifacts
65
+
66
+ | Artifact | Presence and meaning |
67
+ |---|---|
68
+ | `manifest.json` | Required for a complete structural publication. Counts by extractor directory, totals, extraction timestamp and provenance. Optional fields vary by writer/version; see [writer provenance](PUBLISHED_INDEX.md#manifest-writer-provenance). |
69
+ | `source_inputs.json` | Versioned source-input identities and per-consumer provenance for this exact generation; old indexes may omit it. See [source freshness](SOURCE_FRESHNESS.md). |
70
+ | `dependency_graph.json` | Required for a complete structural publication. Typed graph data; an empty graph is valid. |
71
+ | `<type>/_index.json` and unit JSON | Present for extracted families. `_index.json` is an array of unit summaries; an empty array is valid. Disabled/unavailable families may be absent. Do not infer completeness from a fixed count of directories. |
72
+ | `graph_analysis.json` | Derived graph analysis when produced. Treat absence as unavailable analysis, not an empty or corrupt unit index. |
73
+ | `flows/flow_index.json` and flow documents | Optional precomputed flows. Use the index's relative paths; do not invent flow filenames. |
74
+ | `SUMMARY.md` | Generated human-readable summary, not a machine schema. |
75
+
76
+ Unit summaries identify units, but do not carry an artifact filename. Their
77
+ `file_path` names the application's source file, not the unit JSON file. To avoid
78
+ reimplementing filename normalization, scan JSON files in a needed type directory
79
+ (excluding `_index.json`) and match the JSON `identifier` **and** `type`. The same
80
+ identifier can exist in multiple types. The full unit fields are documented in
81
+ [Extractor reference](EXTRACTOR_REFERENCE.md#extractedunit-field-reference).
82
+
83
+ The graph's `nodes`, typed variants, forward/reverse relationships and relationship
84
+ metadata belong to the graph format; preserve them when transporting the index.
85
+ Extractor directories can contain multiple unit types: `graphql/` contains
86
+ `graphql_type`, `graphql_mutation`, `graphql_resolver`, and `graphql_query`, while
87
+ `rails_source/` contains `rails_source` and `gem_source`. Validate and retain the
88
+ artifact's actual type instead of deriving it by singularizing the directory.
89
+
90
+ Do not flatten typed variants into a single node per textual identifier. The
91
+ static Woods self-map has the same publication envelope but different type
92
+ families and `manifest.provenance.mode`; it is not Rails runtime evidence.
93
+
94
+ ### Semantic graph validation
95
+
96
+ `woods:validate` checks raw graph data against the unit indexes and artifacts in
97
+ one pinned published generation. This validation is unreleased after
98
+ `2.0.0.beta2`. It checks the shapes of `nodes`, `edges`, `reverse`, `file_map`,
99
+ `type_index`, and optional `variants`; typed identities must be unique and agree
100
+ with the actual indexed units. Forward sources must exist. Reverse, file, and
101
+ type memberships must match the union of primary and variant contributions.
102
+ When present, `reverse_via` must preserve the typed forward records and their
103
+ relationship attributes, including duplicate-record multiplicity.
104
+
105
+ Cycles, recursion, shared paths, cross-type identifier collisions, nil file paths,
106
+ and legacy bare-string edges are valid. Absent optional legacy fields remain
107
+ legal. A target absent from the nodes is an unresolved reference, not automatically
108
+ corruption: current metadata cannot distinguish an intentional external target
109
+ from an internal node and unit that are both missing. Validation cannot certify a
110
+ unique target type when several types share its name. A missing node whose typed
111
+ unit remains indexed is detectable and is an error.
112
+
113
+ The checker does not repair data, rerun extraction, or become a publication gate.
114
+ It reports identity/path diagnostics through the existing validation report and
115
+ nonzero task exit. It also supports the Woods static source map's explicitly
116
+ recorded type families; that does not give the map Rails runtime fidelity.
117
+ See [running validation from Ruby](PUBLISHED_INDEX.md#validate-a-published-generation)
118
+ and [semantic error recovery](TROUBLESHOOTING.md#semantic-graph-validation-errors).
119
+
120
+ ### File profiles and file membership
121
+
122
+ `file_map[path]` lists units associated with a source file, including whole-file
123
+ profiles. It does not promise that every identifier names a Ruby constant. New
124
+ writers mark graph nodes of types `caching`, `configuration`, `test_mapping`,
125
+ `rails_source`, and `gem_source` with `"kind": "file_profile"`. The same field
126
+ appears on non-primary typed variants.
127
+
128
+ For example, `app/controllers/things_controller.rb` can map to both
129
+ `ThingsController` and a caching unit named `app/controllers/things_controller.rb`.
130
+ Inspect each typed node's `kind` to distinguish the profile; both retain their
131
+ file membership so a source edit refreshes both units. Do not classify units by
132
+ comparing the identifier with the path, and do not interpret a missing `kind` as
133
+ proof that the unit names a constant.
134
+
135
+ Older indexes omit the marker. Woods derives it from these known extractor types
136
+ when loading and republishing a graph, including unchanged incremental nodes.
137
+ Raw consumers of older indexes must treat absent markers as unclassified or use
138
+ the documented type list. Identifiers, `file_map`, and type membership retain
139
+ their existing shapes; older readers can ignore `kind`.
140
+
141
+ ### Reverse relationship records
142
+
143
+ New writers add `reverse_via` to `dependency_graph.json`. Each target identifier
144
+ maps to its incoming relationship records, including the owning source type:
145
+
146
+ ```json
147
+ {
148
+ "reverse": { "Gadget": ["Widget"] },
149
+ "reverse_via": {
150
+ "Gadget": [{ "source": "Widget", "source_type": "service", "via": "render" }]
151
+ }
152
+ }
153
+ ```
154
+
155
+ The existing `reverse` arrays retain their bare identifiers. `reverse_via` includes
156
+ edges from primary nodes and typed variants; several records can share a source
157
+ and target while differing in type, relationship, or association attributes.
158
+ Optional `through`, `through_db`, and `disable_joins` values match the forward
159
+ edge. A `null` relationship means unknown legacy evidence. Target type remains
160
+ unresolved when the target identifier belongs to multiple types; the source type
161
+ does not resolve that ambiguity.
162
+
163
+ These are recorded dependencies, not proof that changing a target breaks every
164
+ source. For example, `factory_for` and migration `reference` edges describe
165
+ different relationships from a runtime `render` edge. Consumers can inspect one
166
+ target bucket without scanning the whole forward graph. Buckets and records are
167
+ deterministically ordered, but consumers should treat the ordering as incidental.
168
+
169
+ Older graphs omit `reverse_via`; absence means relationship detail must be derived
170
+ from forward edges and variants, not that there are no dependents. Woods rebuilds
171
+ this derived index from forward evidence when loading and republishing a graph.
172
+ Existing readers can ignore the additive field; a subsequent changed extraction
173
+ or full run publishes it. A no-op leaves the previous generation unchanged.
174
+
175
+ ### What is outside this structural snapshot
176
+
177
+ `woods.json`, `dumps/`, embedding checkpoints, temporal snapshots, watch status,
178
+ pending paths, MCP task records, exporter state and extraction/startup lock files
179
+ have separate lifecycles. Their presence is configuration-dependent. Do not glob
180
+ them into a structural payload or assume `generation.json` commits them together.
181
+ A semantic/MCP deployment also needs its configured stores and artifacts; copying
182
+ a structural payload alone does not clone that deployment.
183
+
184
+ `<output>/.source-inputs.key` is private operational state used to verify source
185
+ identities. Never include it in a published/exported structural snapshot; without
186
+ it, a recipient can still read the index but source freshness is unknown.
187
+
188
+ Temporary filenames, abandoned payloads, lock sidecars, retained-directory counts,
189
+ and summary formatting are implementation details. Never remove or modify locks
190
+ to make a reader proceed.
191
+
192
+ ## Atomic publication is not indefinite retention
193
+
194
+ For a payload publication, Woods writes the payload, flushes it, and atomically
195
+ replaces `generation.json` last. Capturing that pointer once selects a complete,
196
+ immutable structural generation. Opening each file through a freshly reread pointer
197
+ can mix generations and defeats that guarantee. A failed/no-op run does not advance
198
+ the pointer. [Durability details](PUBLISHED_INDEX.md#durability-the-pointer-is-the-commit-point)
199
+ explain the flush boundary.
200
+
201
+ Retention can delete an older payload after the reader selects it. The default
202
+ retains three generations, not three minutes. To keep a multi-file read or copy
203
+ safe from Woods retention:
204
+
205
+ 1. Read and validate the pointer; resolve its payload inside the root.
206
+ 2. Open that payload's existing `manifest.json` **read-only** and take a shared
207
+ advisory `flock` on that open file. Do not create a new lock file.
208
+ 3. Re-read the pointer and confirm it is unchanged; verify the pathname still
209
+ names the open manifest inode. A pruner may have won between steps 1 and 2.
210
+ 4. Keep the handle/lock open for **all** reads or the complete copy. Use the single
211
+ captured payload path throughout. Once pinned, a later pointer advance is fine.
212
+ 5. Close the handle when done. On a race, discard partial results and retry the
213
+ entire operation from step 1, with a bounded retry count.
214
+
215
+ Woods retention attempts a nonblocking **exclusive** `flock` on that same manifest
216
+ before deletion and skips a payload held by readers. These are advisory filesystem
217
+ locks, not Woods' writer-coordination locks. Ordinary readers need no exclusive
218
+ lock and need not take `extraction.lock` or its guard. A lock-free reader must
219
+ accept disappearance and restart the whole read; it must never silently replace
220
+ missing files with files from another generation.
221
+
222
+ This protocol protects against cooperating retention only. It does not protect
223
+ against `woods:clean`, manual deletion, index replacement, or a filesystem that
224
+ does not coordinate `flock` across its clients. Stop writers/cleanup and read an
225
+ immutable snapshot if that protection is unavailable. Flat indexes (including
226
+ full-extraction fallback when a payload cannot be created) have individually
227
+ replaced files, not multi-file atomicity; read/copy them only with writers stopped.
228
+
229
+ ## Bash and jq: read one pinned generation
230
+
231
+ Requires Bash, jq, GNU `realpath`, util-linux `flock`, and Linux `/proc`. Save as
232
+ `read-woods.sh`, then run `bash read-woods.sh '/path with spaces/tmp/woods'`.
233
+ The final command reads manifest and graph under the same lock. Substitute other
234
+ reads or a complete copy **inside** the script before it exits. Printing a path
235
+ and consuming it after the script exits does not keep it pinned.
236
+
237
+ Exit 75 means a possible publication/retention race: retry the whole script a
238
+ bounded number of times (for example three), discarding any previous output.
239
+ Other nonzero exits need investigation. A persistent 75 can mean a broken index.
240
+
241
+ ```bash
242
+ #!/usr/bin/env bash
243
+ set -euo pipefail
244
+ fail() { echo "$*" >&2; exit 1; }
245
+ retry() { echo "$*; retry the whole read" >&2; exit 75; }
246
+ root=$(realpath -e -- "${1:?provide the index root}")
247
+ [[ -d "$root" ]] || fail 'Index root is not a directory'
248
+ [[ -f "$root/generation.json" ]] || fail 'Missing pointer: legacy or unpublished index'
249
+ marker=$(cat -- "$root/generation.json") || retry 'Cannot read pointer'
250
+ jq -e '
251
+ if type != "object" then false
252
+ elif (.number | type) != "number" then false
253
+ else (.number >= 1 and (.number | floor) == .number)
254
+ and (.token | type == "string" and length > 0)
255
+ and (.payload | type == "string" and length > 0)
256
+ and (.payload | explode | all(. >= 32 and . != 127))
257
+ end
258
+ ' <<<"$marker" >/dev/null || fail 'Invalid pointer or flat layout'
259
+ relative=$(jq -r '.payload' <<<"$marker")
260
+ [[ "$relative" != /* ]] || fail 'Absolute payload path'
261
+ payload=$(realpath -e -- "$root/$relative") || retry 'Missing payload'
262
+ [[ -d "$payload" && "$payload" == "${root%/}/"* && "$payload" != "$root" ]] \
263
+ || fail 'Payload must resolve inside the index root'
264
+ exec 9< "$payload/manifest.json" || retry 'Missing manifest'
265
+ flock -sn 9 || retry 'Cannot pin manifest'
266
+ current=$(cat -- "$root/generation.json") || retry 'Pointer disappeared'
267
+ [[ "$current" == "$marker" && "$payload/manifest.json" -ef /proc/self/fd/9 ]] \
268
+ || retry 'Publication changed or retention removed the payload'
269
+ jq -n --argjson generation "$marker" \
270
+ --slurpfile manifest "$payload/manifest.json" \
271
+ --slurpfile graph "$payload/dependency_graph.json" \
272
+ 'if ($manifest | length) == 1 and ($manifest[0] | type) == "object"
273
+ and ($graph | length) == 1 and ($graph[0] | type) == "object"
274
+ then {generation: $generation, manifest: $manifest[0], dependency_graph: $graph[0]}
275
+ else error("Manifest and graph must each contain exactly one JSON object") end'
276
+ # Descriptor 9 closes on exit, releasing the retention pin.
277
+ ```
278
+
279
+ ## Python: keep the pin while using the payload
280
+
281
+ Requires Python 3.9+ on Unix with working `fcntl.flock`. Save as `read_woods.py`
282
+ and run `python3 read_woods.py '/path with spaces/tmp/woods'`. The context manager
283
+ retries acquisition three times. Errors during the read/copy propagate: discard
284
+ partial output before retrying the entire operation. The scripts assume a trusted
285
+ Woods-owned index; path containment is not a sandbox for hostile filesystem changes.
286
+
287
+ ```python
288
+ import contextlib
289
+ import fcntl
290
+ import json
291
+ import os
292
+ from pathlib import Path
293
+ import sys
294
+ import time
295
+
296
+
297
+ @contextlib.contextmanager
298
+ def pinned_payload(index_root):
299
+ root = Path(index_root).resolve(strict=True)
300
+ pointer = root / "generation.json"
301
+ if not pointer.is_file():
302
+ raise ValueError("Missing pointer: legacy or unpublished index")
303
+ for attempt in range(3):
304
+ handle = None
305
+ try:
306
+ raw = pointer.read_bytes()
307
+ marker = json.loads(raw)
308
+ if not isinstance(marker, dict):
309
+ raise ValueError("Pointer must be an object")
310
+ number, token, name = (marker.get(k) for k in ("number", "token", "payload"))
311
+ if type(number) is not int or number < 1 or not isinstance(token, str) or not token:
312
+ raise ValueError("Invalid generation identity")
313
+ if not isinstance(name, str) or not name or Path(name).is_absolute():
314
+ raise ValueError("Invalid payload path or flat layout")
315
+ if any(ord(char) < 32 or ord(char) == 127 for char in name):
316
+ raise ValueError("Control character in payload path")
317
+ payload = (root / name).resolve(strict=True)
318
+ if not payload.is_dir() or root not in payload.parents:
319
+ raise ValueError("Payload must resolve inside the index root")
320
+ manifest = payload / "manifest.json"
321
+ handle = manifest.open("rb")
322
+ fcntl.flock(handle, fcntl.LOCK_SH | fcntl.LOCK_NB)
323
+ if pointer.read_bytes() != raw or not os.path.samestat(os.fstat(handle.fileno()), manifest.stat()):
324
+ raise BlockingIOError("Publication changed or payload was removed")
325
+ except (FileNotFoundError, BlockingIOError):
326
+ if handle is not None:
327
+ handle.close()
328
+ if attempt == 2:
329
+ raise
330
+ time.sleep(0.05)
331
+ except BaseException:
332
+ if handle is not None:
333
+ handle.close()
334
+ raise
335
+ else:
336
+ break
337
+ try:
338
+ yield marker, payload
339
+ finally:
340
+ handle.close()
341
+
342
+
343
+ if __name__ == "__main__":
344
+ with pinned_payload(sys.argv[1]) as (generation, payload):
345
+ manifest = json.loads((payload / "manifest.json").read_text(encoding="utf-8"))
346
+ graph = json.loads((payload / "dependency_graph.json").read_text(encoding="utf-8"))
347
+ if not isinstance(manifest, dict) or not isinstance(graph, dict):
348
+ raise ValueError("Manifest and graph must each contain exactly one JSON object")
349
+ # Read more artifacts, or copy the whole payload, before leaving this block.
350
+ print(json.dumps({"generation": generation, "manifest": manifest, "dependency_graph": graph}))
351
+ ```
352
+
353
+ ## Shipping a structural snapshot
354
+
355
+ While the pin is held, copy the selected payload into an unpublished staging
356
+ location. Preserve the payload's relative path and pair it with the **captured**
357
+ `generation.json`, not a later pointer reread from the live index. For example,
358
+ a captured `payloads/gen-42` must still resolve to that directory in the exported
359
+ root. Validate the copied manifest and graph, then publish the staged copy as a
360
+ whole, or upload its payload first and switch the destination pointer last.
361
+ Do not upload `generation.json` first or advertise a failed/partial copy.
362
+
363
+ The retention lock is local coordination; do not copy its file descriptor or
364
+ writer lock files. Copying `manifest.json` normally copies content, which is correct.
365
+ Objects in a destination store do not inherit the source's locking guarantees;
366
+ protect any later destination pruning with that destination's own reader protocol.
367
+
368
+ ## Compatibility within Woods 2.x
369
+
370
+ Consumers may rely on the pointer field meanings, relative payload resolution,
371
+ the complete-payload publication boundary, manifest/graph locations, and JSON unit
372
+ identity described here. Additive JSON fields, new extractor families, new reason
373
+ values, optional artifacts and additional root-level state may appear in 2.x.
374
+ Ignore unknown fields and inspect available types instead of hardcoding a directory
375
+ count. Existing field meanings and required-file locations are compatibility
376
+ surfaces; incompatible changes require an explicit migration contract.
377
+
378
+ Do not bind to temporary names, a fixed token length, directory listing order,
379
+ retention count, internal lock sidecars, or Markdown summary formatting. Check the
380
+ installed Woods version and its release's documentation when consuming older
381
+ prereleases; this guide describes the current source contract, not a promise that
382
+ every prerelease contains every optional field or retention improvement.
data/docs/INTERNALS.md CHANGED
@@ -145,6 +145,11 @@ The `DependencyGraph` is a directed graph where nodes are `ExtractedUnit` identi
145
145
  - **Forward edges** (`@edges`): what each unit depends on, populated when units are registered
146
146
  - **Reverse edges** (`@reverse`): what depends on each unit, built during registration and in the resolve phase
147
147
 
148
+ The published graph also carries additive `reverse_via` target buckets with
149
+ typed source identities, relationship labels and association attributes. The
150
+ bare-name `reverse` map remains compatible. See the
151
+ [reverse relationship format](INDEX_LAYOUT.md#reverse-relationship-records).
152
+
148
153
  ```ruby
149
154
  graph = DependencyGraph.new
150
155
  graph.register(user_unit) # adds User to nodes, adds User→Order edge (from belongs_to)
@@ -175,7 +180,7 @@ Scores feed into the retrieval ranker as one signal in the final ranking formula
175
180
  | **Cycles** | Circular dependencies, A→B→C→A. Detected via DFS, and capped: `graph_cycle_limit` (default 500) bounds how many are enumerated and `graph_cycle_max_length` (default 50) skips one longer than that. Either cap firing sets `stats.cycle_limit_reached`. Set both to `nil` for exhaustive enumeration. |
176
181
  | **Bridges** | Edges whose removal would disconnect the graph, high-risk structural connections |
177
182
  | **Cross-database edges** | Association or foreign-key edges whose two ends resolve to different databases. A `has_many :through` is reported as `join_through_across_databases` when `disable_joins` is false and `from_db`, `through_db` (the join model's database), or `to_db` disagree. A foreign key never resolves to an owner in the source database, even when another database also claims the table; when every owner sits elsewhere and they span more than one database, the entry comes back with `to: nil` and an `ambiguous_owners` list instead of guessing. Read from graph node and edge attributes, so full and incremental runs agree. Scoped to primary nodes (units registered in the graph), not variants. |
178
- | **Volatile dependencies** | Edges that point at a unit changing at least `volatile_dependency_ratio` times more often than the dependent (POODR: depend on things that change less often than you do). Dependencies with fewer than 5 commits or a `new` change frequency are skipped. Ranked by the dependency's PageRank; the persisted list keeps the top 20, while `stats.volatile_dependency_count` reports the full qualifying count and `stats.volatile_dependencies_limit` reports the cap. |
183
+ | **Volatile dependencies** | Edges that point at a unit changing at least `volatile_dependency_ratio` times more often than the dependent (POODR: depend on things that change less often than you do). Dependencies with fewer than 5 commits or a `new` change frequency are skipped. Ranked by the dependency's PageRank; an optional `volatile_dependency_limit_per_target` selects edges per typed dependency before the persisted top 20, while `stats.volatile_dependency_count` reports the full qualifying count and `stats.volatile_dependencies_limit` reports the cap. |
179
184
  | **Undeclared package edges** | Edges that cross a Packwerk package boundary the source package does not list in `dependencies`. Membership comes from each unit's `package` node attribute, declarations from the package unit's own `package_dependency` edges. Woods reports the boundary; enforcement stays with `packwerk check` / `pks check`. |
180
185
 
181
186
  Analysis results are written to `graph_analysis.json` and surfaced in `SUMMARY.md`.
@@ -194,7 +199,7 @@ persisted rather than every hub in the graph.
194
199
  | `cycles` | `stats.cycle_count`, `stats.cycle_limit_reached` | yes, `graph_cycle_limit` cycles of at most `graph_cycle_max_length` nodes | the count is the persisted array; the flag says whether either cap fired |
195
200
  | `bridges` | none | yes, top 10 by score | not counted |
196
201
  | `cross_database_edges` | `stats.cross_database_edge_count` | no | every crossing edge |
197
- | `volatile_dependencies` | `stats.volatile_dependency_count`, `stats.volatile_dependencies_limit` | yes, top 20 by the dependency's PageRank | the count is every qualifying edge; the limit is the cap |
202
+ | `volatile_dependencies` | `stats.volatile_dependency_count`, `stats.volatile_dependencies_limit`; optional `stats.volatile_dependencies_limit_per_target`, `stats.volatile_dependency_reported_count` | yes, optional per-target cap followed by top 20 by the dependency's PageRank | count is every qualifying edge before caps; reported count is the final array length and appears only when the per-target cap is enabled |
198
203
  | `undeclared_package_edges` | `stats.undeclared_package_edge_count` | no | every undeclared crossing |
199
204
 
200
205
  A capped section means the array on disk is a page, not the population.