rag-your-code 0.4.1__tar.gz → 0.5.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (46) hide show
  1. rag_your_code-0.5.0/PKG-INFO +358 -0
  2. rag_your_code-0.5.0/README.md +331 -0
  3. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/pyproject.toml +1 -1
  4. rag_your_code-0.5.0/src/rag_your_code.egg-info/PKG-INFO +358 -0
  5. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/src/rag_your_code.egg-info/SOURCES.txt +4 -0
  6. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/src/ragyourcode/__init__.py +1 -1
  7. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/src/ragyourcode/agentic.py +7 -0
  8. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/src/ragyourcode/annotate.py +11 -0
  9. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/src/ragyourcode/cli.py +100 -9
  10. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/src/ragyourcode/config.py +51 -0
  11. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/src/ragyourcode/descriptions.py +125 -11
  12. rag_your_code-0.5.0/src/ragyourcode/document.py +261 -0
  13. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/src/ragyourcode/embeddings.py +22 -0
  14. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/src/ragyourcode/graph.py +40 -0
  15. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/src/ragyourcode/indexer.py +96 -17
  16. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/src/ragyourcode/models.py +23 -0
  17. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/src/ragyourcode/parser.py +148 -0
  18. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/src/ragyourcode/search.py +20 -0
  19. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/tests/test_config.py +1 -1
  20. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/tests/test_descriptions.py +5 -1
  21. rag_your_code-0.5.0/tests/test_doc_comments.py +215 -0
  22. rag_your_code-0.5.0/tests/test_document.py +211 -0
  23. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/tests/test_metadata.py +30 -0
  24. rag_your_code-0.5.0/tests/test_repo_queries.py +72 -0
  25. rag_your_code-0.4.1/PKG-INFO +0 -237
  26. rag_your_code-0.4.1/README.md +0 -210
  27. rag_your_code-0.4.1/src/rag_your_code.egg-info/PKG-INFO +0 -237
  28. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/LICENSE +0 -0
  29. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/setup.cfg +0 -0
  30. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/src/rag_your_code.egg-info/dependency_links.txt +0 -0
  31. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/src/rag_your_code.egg-info/entry_points.txt +0 -0
  32. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/src/rag_your_code.egg-info/requires.txt +0 -0
  33. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/src/rag_your_code.egg-info/top_level.txt +0 -0
  34. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/src/ragyourcode/py.typed +0 -0
  35. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/tests/test_agent_protocol.py +0 -0
  36. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/tests/test_agentic.py +0 -0
  37. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/tests/test_e2e_cli.py +0 -0
  38. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/tests/test_golden.py +0 -0
  39. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/tests/test_graph_incremental.py +0 -0
  40. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/tests/test_language_fixtures.py +0 -0
  41. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/tests/test_large_repo.py +0 -0
  42. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/tests/test_multilanguage.py +0 -0
  43. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/tests/test_parser_edges.py +0 -0
  44. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/tests/test_ragyourcode.py +0 -0
  45. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/tests/test_resilience.py +0 -0
  46. {rag_your_code-0.4.1 → rag_your_code-0.5.0}/tests/test_retrieval_correctness.py +0 -0
@@ -0,0 +1,358 @@
1
+ Metadata-Version: 2.4
2
+ Name: rag-your-code
3
+ Version: 0.5.0
4
+ Summary: A local, explainable RAG index for codebases and coding agents
5
+ Author: rag-your-code contributors
6
+ License-Expression: MIT
7
+ Keywords: rag,code-search,retrieval,indexing,graphrag,offline,explainable,agent
8
+ Classifier: Development Status :: 4 - Beta
9
+ Classifier: Environment :: Console
10
+ Classifier: Intended Audience :: Developers
11
+ Classifier: Operating System :: OS Independent
12
+ Classifier: Programming Language :: Python :: 3
13
+ Classifier: Programming Language :: Python :: 3.10
14
+ Classifier: Programming Language :: Python :: 3.11
15
+ Classifier: Programming Language :: Python :: 3.12
16
+ Classifier: Programming Language :: Python :: 3.13
17
+ Classifier: Topic :: Software Development :: Libraries
18
+ Classifier: Topic :: Text Processing :: Indexing
19
+ Classifier: Typing :: Typed
20
+ Requires-Python: >=3.10
21
+ Description-Content-Type: text/markdown
22
+ License-File: LICENSE
23
+ Provides-Extra: dev
24
+ Requires-Dist: pytest>=7; extra == "dev"
25
+ Requires-Dist: tomli>=2.0; python_version < "3.11" and extra == "dev"
26
+ Dynamic: license-file
27
+
28
+ # RAG Your Code
29
+
30
+ **A local code-retrieval index for coding agents.** Ask a question in plain
31
+ language, get back the functions that answer it — each with its file, its exact
32
+ line range, the words that matched, and its source.
33
+
34
+ [![PyPI](https://img.shields.io/badge/PyPI-rag--your--code-blue)](https://pypi.org/project/rag-your-code/)
35
+ [![License](https://img.shields.io/badge/license-MIT-green)](LICENSE)
36
+ [![Python](https://img.shields.io/badge/python-3.10--3.13-blue)](pyproject.toml)
37
+
38
+ No network calls. No runtime dependencies. No model. It is built to run over a
39
+ private repository on a machine with the network switched off, and to produce
40
+ an index a human can read.
41
+
42
+ ---
43
+
44
+ ## The problem it solves
45
+
46
+ An agent that needs to find something in an unfamiliar codebase has two bad
47
+ options. It can grep — fast, but it only finds the string you already guessed.
48
+ Or it can read whole files into context — thorough, but a handful of them
49
+ exhausts the budget and most of what it read was irrelevant.
50
+
51
+ This sits in between. It indexes every function, method and class once, then
52
+ answers a question with the eight units most likely to be relevant, at roughly
53
+ a hundred lines instead of ten thousand. Every result carries its provenance,
54
+ so the agent can open the real code before it edits anything, and you can see
55
+ why each one came back.
56
+
57
+ It is the **R** in RAG. There is no generation here — your agent is the G.
58
+
59
+ ## Install
60
+
61
+ **As a Claude Code plugin** (this is the primary way to use it):
62
+
63
+ ```
64
+ /plugin marketplace add skymanbp/rag-your-code
65
+ /plugin install rag-your-code@rag-your-code
66
+ /reload-plugins
67
+ ```
68
+
69
+ The plugin is one skill and nothing else — no hooks, no agents, no MCP server.
70
+ Measured with `claude plugin details`: **~39 tokens added to every session**,
71
+ and ~1.4k only when the skill actually fires. The skill installs the Python
72
+ package itself on first use.
73
+
74
+ **Or as a plain CLI:**
75
+
76
+ ```bash
77
+ pip install rag-your-code
78
+
79
+ rag-your-code index .
80
+ rag-your-code search "where are HTTP retries handled" --json
81
+ rag-your-code search "what calls the retry handler" --graph --hops 1 --json
82
+ ```
83
+
84
+ The index is written under `.rag-your-code/`; your source files are never
85
+ modified. Later `index` runs reuse unchanged files. For a large repository,
86
+ prefer `rag-your-code index . --compact`.
87
+
88
+ ## How it works
89
+
90
+ ```
91
+ your repository
92
+ → walk source files (configurable ignores, suffixes, size cap)
93
+ → parse declarations Python via its own AST; 14 other languages via a
94
+ line scanner + per-language rule table
95
+ → one CodeUnit each id, signature, exact line range, source, calls,
96
+ imports, a stable serial number, a description
97
+ → embed description + source into a deterministic local vector
98
+ → inverted word index + hybrid ranking
99
+ → optional graph expansion over calls / imports / contains
100
+ → results, or a JSON-lines protocol for an agent subprocess
101
+ ```
102
+
103
+ **Parsing.** Python goes through the standard-library syntax tree, so nesting,
104
+ qualified names, call lists and line ranges are exact. Every other language
105
+ goes through three separated layers: a scanner that reads one line at a time,
106
+ a rule table per language, and a span closer that follows brace depth, Ruby's
107
+ `end`, or the next declaration. Because a pattern never sees a second line, a
108
+ reported line number *is* the scanner's loop index and cannot drift, and a
109
+ declaration cannot swallow the ones after it.
110
+
111
+ Fifteen languages: Python, JavaScript, TypeScript, Go, Rust, Java, Kotlin,
112
+ Scala, C#, C, C++, Ruby, PHP, Swift, shell.
113
+
114
+ **Graph.** `calls`, `imports` and `contains` edges, each conservative: an
115
+ unresolved or ambiguous reference produces no edge rather than a guessed one,
116
+ and every expanded result carries the exact edge path as evidence.
117
+
118
+ ## What the embedding does — and what it does not
119
+
120
+ This matters more than any feature list, so it is here rather than in a
121
+ footnote.
122
+
123
+ The embedder is a **signed feature hash**: it hashes words into 384 buckets.
124
+ Cosine similarity over those vectors is therefore a normalised measure of
125
+ *shared words*, and it carries no semantics whatsoever:
126
+
127
+ | pair | cosine |
128
+ |---|---|
129
+ | `retry failed card charge` vs itself | 1.0000 |
130
+ | `sum two numbers` vs `add a pair of integers` | **0.0000** |
131
+ | `计算两个数的和` vs `sum two numbers` | **0.0000** |
132
+ | `sum two numbers` vs `delete the user database table` | 0.0000 |
133
+
134
+ A trained embedding model scores row 2 at around 0.8. Here a synonym pair and
135
+ an unrelated pair are indistinguishable, because no shared word is no shared
136
+ word either way.
137
+
138
+ Retrieval works regardless, because **identifiers and docstrings are already
139
+ natural language** — `retry_charge` contains *retry* and *charge*. But it
140
+ reaches only concepts somebody wrote down. Two things close the rest of the
141
+ gap, and neither is a model:
142
+
143
+ - **Your agent rewrites the query.** It has the conversation; turning
144
+ "重试扣款" into `retry charge payment gateway` costs it nothing.
145
+ - **Your agent writes the descriptions**, which puts the missing vocabulary
146
+ into the index once instead of into every query.
147
+
148
+ ## Agent-authored descriptions
149
+
150
+ Every unit carries a description, and that description is indexed. By default
151
+ it is generated without a model: the identifier humanised, the parameters and
152
+ callees listed, the docstring appended. It introduces no vocabulary the source
153
+ did not already have — which is exactly why retrieval cannot reach a concept
154
+ nobody wrote down.
155
+
156
+ **First, the documentation you already wrote is indexed.** Fourteen of the
157
+ fifteen supported languages put documentation immediately above a declaration
158
+ — JSDoc, Javadoc, KDoc, rustdoc, Go doc comments, XML doc comments, PHPDoc —
159
+ and a unit's span begins at the declaration, so all of it used to sit outside
160
+ the index. The same sentence reached thirteen searchable words as a Python
161
+ docstring and two as a JavaScript comment. Now both reach thirteen. Commented-
162
+ out code, separator rules and licence headers are deliberately left out.
163
+
164
+ **Where there is none, the agent can write it:**
165
+
166
+ ```bash
167
+ rag-your-code describe status # coverage, and what is pending
168
+ rag-your-code describe export --limit 20 # a batch, with source and a brief
169
+ rag-your-code describe import written.json # store what the agent wrote
170
+ rag-your-code index . # apply it
171
+ ```
172
+
173
+ or, in the protocol, `describe_pending` and `describe_put` — which take effect
174
+ in the same session, with no refresh.
175
+
176
+ **And you can move it into the code**, where it needs no bookkeeping at all:
177
+
178
+ ```bash
179
+ rag-your-code describe promote | git apply # review it first
180
+ ```
181
+
182
+ That emits a unified diff adding a doc comment in each language's own
183
+ convention, for declarations that have none. The tool still never writes your
184
+ source. Only the half meant for a reader is promoted, so a bilingual
185
+ description leaves its second language in the store where retrieval still uses
186
+ it — measured, promoting all 68 on this repository discarded no description
187
+ and left Chinese retrieval unchanged.
188
+
189
+ ### Measured on this repository
190
+
191
+ This project describes its own implementation: every unit under `src/` carries
192
+ an agent-written bilingual description, committed to the repo, and 68 of them
193
+ have been promoted into the source as doc comments.
194
+
195
+ Seventy natural-language questions about this codebase, in English and
196
+ Chinese, each listing every unit that genuinely answers it
197
+ ([`benchmarks/repo_queries.json`](benchmarks/repo_queries.json)):
198
+
199
+ | | generated descriptions | agent-written |
200
+ |---|---|---|
201
+ | hit@1 | 0.171 | **0.500** |
202
+ | hit@3 | 0.314 | **0.729** |
203
+ | MRR | 0.240 | **0.605** |
204
+ | answered with no shared word at all | 15.7% | **0%** |
205
+
206
+ Roughly a threefold improvement in first-place accuracy. Nineteen questions
207
+ still fail, which is what makes the set usable for measuring the next change;
208
+ `tests/test_repo_queries.py` asserts that some question always does.
209
+
210
+ One failure is worth naming: a query saying `catastrophic backtracking` does
211
+ not reach a description saying `backtracks catastrophically`. There is no
212
+ stemming — exactly the limit documented above.
213
+
214
+ **What this is:** it moves the semantic work from query time to index time.
215
+ Matching stays lexical. It is LLM-authored keyword expansion, and its reach is
216
+ bounded by how many ways of saying the thing the agent thought to write down.
217
+
218
+ Descriptions live in `rag-your-code.descriptions.json` at the repository root
219
+ and are meant to be committed, so one person's pass benefits everyone who
220
+ clones. Each is keyed by unit id **and a digest of the unit's source**: when
221
+ the code changes, the description is not applied, the unit returns to the
222
+ pending queue, and retrieval falls back to the generated sentence. A
223
+ description that outlived its code would be a confident wrong answer, which is
224
+ the one thing this index is built not to give. When code merely *moves* — an
225
+ import added above it — the description follows it by digest.
226
+
227
+ ## Measured
228
+
229
+ **Parsing**, against source-controlled fixtures in `tests/fixtures/languages/`
230
+ (15 fixture files, 96 expected units, 237 negative cases, 89 constructs the
231
+ spec deliberately excludes):
232
+
233
+ | | |
234
+ |---|---|
235
+ | core declarations found | **91 / 91** |
236
+ | with the correct `start_line` | **91 / 91** |
237
+ | with a usable signature | **91 / 91** |
238
+ | units that do not exist | **0** |
239
+
240
+ A 441-byte JavaScript file that once took **12.6 s** to parse now takes
241
+ **0.36 ms**, and 10 KB takes 2.1 ms — growth is linear again.
242
+
243
+ **Scale**, on a synthetic 10,000-unit repository (500 files):
244
+
245
+ | | |
246
+ |---|---|
247
+ | full build | 1.84 s |
248
+ | incremental rebuild after one file changes | 0.207 s (**8.9x**) |
249
+ | compact storage vs readable JSON | 35.6% |
250
+ | index load, in a fresh process | 45.4 ms |
251
+ | inverted index build | 117.7 ms |
252
+ | resident memory | 58.7 MiB |
253
+ | query, mean of 200 warmed samples | 3.90 ms |
254
+
255
+ Directional local measurements, not service levels; the archived run is
256
+ [`large-benchmark-result.json`](large-benchmark-result.json).
257
+
258
+ **Suite:** Python 3.10 – 3.13 on Linux and Windows, plus a job that installs
259
+ the built wheel into a clean environment and runs every command the
260
+ documentation prescribes, and another that runs the skill's own install line
261
+ verbatim. 248 tests as of 0.5.0 — the count is version-stamped rather than
262
+ maintained, because a bare figure in a living document is a claim that rots;
263
+ per-release counts are in [CHANGELOG.md](CHANGELOG.md).
264
+
265
+ ## Configuration
266
+
267
+ Twelve settings in `rag-your-code.toml` at the repository root:
268
+
269
+ ```bash
270
+ rag-your-code config init # a commented file, all defaults
271
+ rag-your-code config list # effective values and their source
272
+ rag-your-code config set index.ignore '["vendor", "generated"]'
273
+ rag-your-code config set search.vector_weight 0.25
274
+ ```
275
+
276
+ | section | settings |
277
+ |---|---|
278
+ | `[index]` | `ignore`, `suffixes`, `max_file_bytes` |
279
+ | `[embedding]` | `dimensions` |
280
+ | `[search]` | `vector_weight`, `limit`, `max_chars` |
281
+ | `[agent]` | `max_open_bytes`, `max_open_chars` |
282
+ | `[describe]` | `languages`, `batch`, `max_chars` |
283
+
284
+ Resolution is CLI flag > file > built-in default. There is no environment
285
+ layer: an index is an artifact of a repository, not of a shell.
286
+
287
+ An unknown key or an out-of-range value is an error, not a shrug — a setting
288
+ silently dropped is indistinguishable from one that had no effect.
289
+ `index.suffixes` may only name suffixes the parser has rules for, because a
290
+ suffix it cannot read is walked, parsed to nothing, and reported as a clean
291
+ index of zero units.
292
+
293
+ The four settings under `[index]` and `[embedding]` decide what an index
294
+ *contains*, so a digest of them is stored in the index and a change forces a
295
+ full rebuild. The rest take effect immediately and invalidate nothing.
296
+
297
+ ## Agent protocol
298
+
299
+ `rag-your-code agent --root PATH` reads one JSON request per line and writes
300
+ one reply per line:
301
+
302
+ ```json
303
+ {"action":"search","query":"database transaction rollback","limit":5}
304
+ {"action":"research","query":"trace payment retry behavior","max_steps":2}
305
+ {"action":"neighbors","id":"payments.py:4:retry_charge","hops":1}
306
+ {"action":"open","path":"payments.py","start_line":1,"end_line":80}
307
+ {"action":"describe_pending","limit":20}
308
+ {"action":"describe_put","descriptions":[{"id":"payments.py:4:retry_charge","text":"..."}]}
309
+ {"action":"refresh"}
310
+ {"action":"stats"}
311
+ ```
312
+
313
+ **No single request can end the session.** Numeric fields saturate at their
314
+ bounds, `open` is bounded in both lines and bytes, and anything unanticipated
315
+ is reported in-band with its exception type. Streams are pinned to UTF-8
316
+ rather than following the console codepage.
317
+
318
+ `research` is a deliberately bounded two-step controller: retrieve, then at
319
+ most one graph expansion when confidence is low, reporting each step and why
320
+ it stopped.
321
+
322
+ ## What lives where
323
+
324
+ | path | authored or generated | commit it? |
325
+ |---|---|---|
326
+ | `rag-your-code.toml` | authored | yes |
327
+ | `rag-your-code.descriptions.json` | authored by your agent | yes |
328
+ | `.rag-your-code/` (index, vectors, annotations) | generated | no |
329
+
330
+ Nothing authored lives under `.rag-your-code/` — that directory is what people
331
+ delete to clear the cache.
332
+
333
+ ## Not here yet
334
+
335
+ Provider-backed embeddings, Tree-sitter parsing, and a SQLite/ANN storage layer
336
+ for repositories past the measured JSON envelope. Agent-authored descriptions
337
+ are deliberately the cheaper answer to the same problem provider embeddings
338
+ solve: they keep the zero-dependency, offline, reproducible-index properties,
339
+ and produce text a human can read and correct rather than opaque floats. See
340
+ [docs/ROADMAP.md](docs/ROADMAP.md).
341
+
342
+ ## Development
343
+
344
+ ```bash
345
+ python -m pip install -e ".[dev]"
346
+ pytest -q
347
+ ```
348
+
349
+ No runtime dependencies; `pytest` and, below Python 3.11, `tomli` come from the
350
+ `dev` extra.
351
+
352
+ - [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) — how each stage works and why
353
+ - [docs/TESTING.md](docs/TESTING.md) — what the suites are protecting
354
+ - [docs/ROADMAP.md](docs/ROADMAP.md) — what shipped, what is still open
355
+ - [CONTRIBUTING.md](CONTRIBUTING.md) — ground rules, and how to add a language
356
+ - [CHANGELOG.md](CHANGELOG.md) — every release, with its measurements
357
+
358
+ MIT licensed.
@@ -0,0 +1,331 @@
1
+ # RAG Your Code
2
+
3
+ **A local code-retrieval index for coding agents.** Ask a question in plain
4
+ language, get back the functions that answer it — each with its file, its exact
5
+ line range, the words that matched, and its source.
6
+
7
+ [![PyPI](https://img.shields.io/badge/PyPI-rag--your--code-blue)](https://pypi.org/project/rag-your-code/)
8
+ [![License](https://img.shields.io/badge/license-MIT-green)](LICENSE)
9
+ [![Python](https://img.shields.io/badge/python-3.10--3.13-blue)](pyproject.toml)
10
+
11
+ No network calls. No runtime dependencies. No model. It is built to run over a
12
+ private repository on a machine with the network switched off, and to produce
13
+ an index a human can read.
14
+
15
+ ---
16
+
17
+ ## The problem it solves
18
+
19
+ An agent that needs to find something in an unfamiliar codebase has two bad
20
+ options. It can grep — fast, but it only finds the string you already guessed.
21
+ Or it can read whole files into context — thorough, but a handful of them
22
+ exhausts the budget and most of what it read was irrelevant.
23
+
24
+ This sits in between. It indexes every function, method and class once, then
25
+ answers a question with the eight units most likely to be relevant, at roughly
26
+ a hundred lines instead of ten thousand. Every result carries its provenance,
27
+ so the agent can open the real code before it edits anything, and you can see
28
+ why each one came back.
29
+
30
+ It is the **R** in RAG. There is no generation here — your agent is the G.
31
+
32
+ ## Install
33
+
34
+ **As a Claude Code plugin** (this is the primary way to use it):
35
+
36
+ ```
37
+ /plugin marketplace add skymanbp/rag-your-code
38
+ /plugin install rag-your-code@rag-your-code
39
+ /reload-plugins
40
+ ```
41
+
42
+ The plugin is one skill and nothing else — no hooks, no agents, no MCP server.
43
+ Measured with `claude plugin details`: **~39 tokens added to every session**,
44
+ and ~1.4k only when the skill actually fires. The skill installs the Python
45
+ package itself on first use.
46
+
47
+ **Or as a plain CLI:**
48
+
49
+ ```bash
50
+ pip install rag-your-code
51
+
52
+ rag-your-code index .
53
+ rag-your-code search "where are HTTP retries handled" --json
54
+ rag-your-code search "what calls the retry handler" --graph --hops 1 --json
55
+ ```
56
+
57
+ The index is written under `.rag-your-code/`; your source files are never
58
+ modified. Later `index` runs reuse unchanged files. For a large repository,
59
+ prefer `rag-your-code index . --compact`.
60
+
61
+ ## How it works
62
+
63
+ ```
64
+ your repository
65
+ → walk source files (configurable ignores, suffixes, size cap)
66
+ → parse declarations Python via its own AST; 14 other languages via a
67
+ line scanner + per-language rule table
68
+ → one CodeUnit each id, signature, exact line range, source, calls,
69
+ imports, a stable serial number, a description
70
+ → embed description + source into a deterministic local vector
71
+ → inverted word index + hybrid ranking
72
+ → optional graph expansion over calls / imports / contains
73
+ → results, or a JSON-lines protocol for an agent subprocess
74
+ ```
75
+
76
+ **Parsing.** Python goes through the standard-library syntax tree, so nesting,
77
+ qualified names, call lists and line ranges are exact. Every other language
78
+ goes through three separated layers: a scanner that reads one line at a time,
79
+ a rule table per language, and a span closer that follows brace depth, Ruby's
80
+ `end`, or the next declaration. Because a pattern never sees a second line, a
81
+ reported line number *is* the scanner's loop index and cannot drift, and a
82
+ declaration cannot swallow the ones after it.
83
+
84
+ Fifteen languages: Python, JavaScript, TypeScript, Go, Rust, Java, Kotlin,
85
+ Scala, C#, C, C++, Ruby, PHP, Swift, shell.
86
+
87
+ **Graph.** `calls`, `imports` and `contains` edges, each conservative: an
88
+ unresolved or ambiguous reference produces no edge rather than a guessed one,
89
+ and every expanded result carries the exact edge path as evidence.
90
+
91
+ ## What the embedding does — and what it does not
92
+
93
+ This matters more than any feature list, so it is here rather than in a
94
+ footnote.
95
+
96
+ The embedder is a **signed feature hash**: it hashes words into 384 buckets.
97
+ Cosine similarity over those vectors is therefore a normalised measure of
98
+ *shared words*, and it carries no semantics whatsoever:
99
+
100
+ | pair | cosine |
101
+ |---|---|
102
+ | `retry failed card charge` vs itself | 1.0000 |
103
+ | `sum two numbers` vs `add a pair of integers` | **0.0000** |
104
+ | `计算两个数的和` vs `sum two numbers` | **0.0000** |
105
+ | `sum two numbers` vs `delete the user database table` | 0.0000 |
106
+
107
+ A trained embedding model scores row 2 at around 0.8. Here a synonym pair and
108
+ an unrelated pair are indistinguishable, because no shared word is no shared
109
+ word either way.
110
+
111
+ Retrieval works regardless, because **identifiers and docstrings are already
112
+ natural language** — `retry_charge` contains *retry* and *charge*. But it
113
+ reaches only concepts somebody wrote down. Two things close the rest of the
114
+ gap, and neither is a model:
115
+
116
+ - **Your agent rewrites the query.** It has the conversation; turning
117
+ "重试扣款" into `retry charge payment gateway` costs it nothing.
118
+ - **Your agent writes the descriptions**, which puts the missing vocabulary
119
+ into the index once instead of into every query.
120
+
121
+ ## Agent-authored descriptions
122
+
123
+ Every unit carries a description, and that description is indexed. By default
124
+ it is generated without a model: the identifier humanised, the parameters and
125
+ callees listed, the docstring appended. It introduces no vocabulary the source
126
+ did not already have — which is exactly why retrieval cannot reach a concept
127
+ nobody wrote down.
128
+
129
+ **First, the documentation you already wrote is indexed.** Fourteen of the
130
+ fifteen supported languages put documentation immediately above a declaration
131
+ — JSDoc, Javadoc, KDoc, rustdoc, Go doc comments, XML doc comments, PHPDoc —
132
+ and a unit's span begins at the declaration, so all of it used to sit outside
133
+ the index. The same sentence reached thirteen searchable words as a Python
134
+ docstring and two as a JavaScript comment. Now both reach thirteen. Commented-
135
+ out code, separator rules and licence headers are deliberately left out.
136
+
137
+ **Where there is none, the agent can write it:**
138
+
139
+ ```bash
140
+ rag-your-code describe status # coverage, and what is pending
141
+ rag-your-code describe export --limit 20 # a batch, with source and a brief
142
+ rag-your-code describe import written.json # store what the agent wrote
143
+ rag-your-code index . # apply it
144
+ ```
145
+
146
+ or, in the protocol, `describe_pending` and `describe_put` — which take effect
147
+ in the same session, with no refresh.
148
+
149
+ **And you can move it into the code**, where it needs no bookkeeping at all:
150
+
151
+ ```bash
152
+ rag-your-code describe promote | git apply # review it first
153
+ ```
154
+
155
+ That emits a unified diff adding a doc comment in each language's own
156
+ convention, for declarations that have none. The tool still never writes your
157
+ source. Only the half meant for a reader is promoted, so a bilingual
158
+ description leaves its second language in the store where retrieval still uses
159
+ it — measured, promoting all 68 on this repository discarded no description
160
+ and left Chinese retrieval unchanged.
161
+
162
+ ### Measured on this repository
163
+
164
+ This project describes its own implementation: every unit under `src/` carries
165
+ an agent-written bilingual description, committed to the repo, and 68 of them
166
+ have been promoted into the source as doc comments.
167
+
168
+ Seventy natural-language questions about this codebase, in English and
169
+ Chinese, each listing every unit that genuinely answers it
170
+ ([`benchmarks/repo_queries.json`](benchmarks/repo_queries.json)):
171
+
172
+ | | generated descriptions | agent-written |
173
+ |---|---|---|
174
+ | hit@1 | 0.171 | **0.500** |
175
+ | hit@3 | 0.314 | **0.729** |
176
+ | MRR | 0.240 | **0.605** |
177
+ | answered with no shared word at all | 15.7% | **0%** |
178
+
179
+ Roughly a threefold improvement in first-place accuracy. Nineteen questions
180
+ still fail, which is what makes the set usable for measuring the next change;
181
+ `tests/test_repo_queries.py` asserts that some question always does.
182
+
183
+ One failure is worth naming: a query saying `catastrophic backtracking` does
184
+ not reach a description saying `backtracks catastrophically`. There is no
185
+ stemming — exactly the limit documented above.
186
+
187
+ **What this is:** it moves the semantic work from query time to index time.
188
+ Matching stays lexical. It is LLM-authored keyword expansion, and its reach is
189
+ bounded by how many ways of saying the thing the agent thought to write down.
190
+
191
+ Descriptions live in `rag-your-code.descriptions.json` at the repository root
192
+ and are meant to be committed, so one person's pass benefits everyone who
193
+ clones. Each is keyed by unit id **and a digest of the unit's source**: when
194
+ the code changes, the description is not applied, the unit returns to the
195
+ pending queue, and retrieval falls back to the generated sentence. A
196
+ description that outlived its code would be a confident wrong answer, which is
197
+ the one thing this index is built not to give. When code merely *moves* — an
198
+ import added above it — the description follows it by digest.
199
+
200
+ ## Measured
201
+
202
+ **Parsing**, against source-controlled fixtures in `tests/fixtures/languages/`
203
+ (15 fixture files, 96 expected units, 237 negative cases, 89 constructs the
204
+ spec deliberately excludes):
205
+
206
+ | | |
207
+ |---|---|
208
+ | core declarations found | **91 / 91** |
209
+ | with the correct `start_line` | **91 / 91** |
210
+ | with a usable signature | **91 / 91** |
211
+ | units that do not exist | **0** |
212
+
213
+ A 441-byte JavaScript file that once took **12.6 s** to parse now takes
214
+ **0.36 ms**, and 10 KB takes 2.1 ms — growth is linear again.
215
+
216
+ **Scale**, on a synthetic 10,000-unit repository (500 files):
217
+
218
+ | | |
219
+ |---|---|
220
+ | full build | 1.84 s |
221
+ | incremental rebuild after one file changes | 0.207 s (**8.9x**) |
222
+ | compact storage vs readable JSON | 35.6% |
223
+ | index load, in a fresh process | 45.4 ms |
224
+ | inverted index build | 117.7 ms |
225
+ | resident memory | 58.7 MiB |
226
+ | query, mean of 200 warmed samples | 3.90 ms |
227
+
228
+ Directional local measurements, not service levels; the archived run is
229
+ [`large-benchmark-result.json`](large-benchmark-result.json).
230
+
231
+ **Suite:** Python 3.10 – 3.13 on Linux and Windows, plus a job that installs
232
+ the built wheel into a clean environment and runs every command the
233
+ documentation prescribes, and another that runs the skill's own install line
234
+ verbatim. 248 tests as of 0.5.0 — the count is version-stamped rather than
235
+ maintained, because a bare figure in a living document is a claim that rots;
236
+ per-release counts are in [CHANGELOG.md](CHANGELOG.md).
237
+
238
+ ## Configuration
239
+
240
+ Twelve settings in `rag-your-code.toml` at the repository root:
241
+
242
+ ```bash
243
+ rag-your-code config init # a commented file, all defaults
244
+ rag-your-code config list # effective values and their source
245
+ rag-your-code config set index.ignore '["vendor", "generated"]'
246
+ rag-your-code config set search.vector_weight 0.25
247
+ ```
248
+
249
+ | section | settings |
250
+ |---|---|
251
+ | `[index]` | `ignore`, `suffixes`, `max_file_bytes` |
252
+ | `[embedding]` | `dimensions` |
253
+ | `[search]` | `vector_weight`, `limit`, `max_chars` |
254
+ | `[agent]` | `max_open_bytes`, `max_open_chars` |
255
+ | `[describe]` | `languages`, `batch`, `max_chars` |
256
+
257
+ Resolution is CLI flag > file > built-in default. There is no environment
258
+ layer: an index is an artifact of a repository, not of a shell.
259
+
260
+ An unknown key or an out-of-range value is an error, not a shrug — a setting
261
+ silently dropped is indistinguishable from one that had no effect.
262
+ `index.suffixes` may only name suffixes the parser has rules for, because a
263
+ suffix it cannot read is walked, parsed to nothing, and reported as a clean
264
+ index of zero units.
265
+
266
+ The four settings under `[index]` and `[embedding]` decide what an index
267
+ *contains*, so a digest of them is stored in the index and a change forces a
268
+ full rebuild. The rest take effect immediately and invalidate nothing.
269
+
270
+ ## Agent protocol
271
+
272
+ `rag-your-code agent --root PATH` reads one JSON request per line and writes
273
+ one reply per line:
274
+
275
+ ```json
276
+ {"action":"search","query":"database transaction rollback","limit":5}
277
+ {"action":"research","query":"trace payment retry behavior","max_steps":2}
278
+ {"action":"neighbors","id":"payments.py:4:retry_charge","hops":1}
279
+ {"action":"open","path":"payments.py","start_line":1,"end_line":80}
280
+ {"action":"describe_pending","limit":20}
281
+ {"action":"describe_put","descriptions":[{"id":"payments.py:4:retry_charge","text":"..."}]}
282
+ {"action":"refresh"}
283
+ {"action":"stats"}
284
+ ```
285
+
286
+ **No single request can end the session.** Numeric fields saturate at their
287
+ bounds, `open` is bounded in both lines and bytes, and anything unanticipated
288
+ is reported in-band with its exception type. Streams are pinned to UTF-8
289
+ rather than following the console codepage.
290
+
291
+ `research` is a deliberately bounded two-step controller: retrieve, then at
292
+ most one graph expansion when confidence is low, reporting each step and why
293
+ it stopped.
294
+
295
+ ## What lives where
296
+
297
+ | path | authored or generated | commit it? |
298
+ |---|---|---|
299
+ | `rag-your-code.toml` | authored | yes |
300
+ | `rag-your-code.descriptions.json` | authored by your agent | yes |
301
+ | `.rag-your-code/` (index, vectors, annotations) | generated | no |
302
+
303
+ Nothing authored lives under `.rag-your-code/` — that directory is what people
304
+ delete to clear the cache.
305
+
306
+ ## Not here yet
307
+
308
+ Provider-backed embeddings, Tree-sitter parsing, and a SQLite/ANN storage layer
309
+ for repositories past the measured JSON envelope. Agent-authored descriptions
310
+ are deliberately the cheaper answer to the same problem provider embeddings
311
+ solve: they keep the zero-dependency, offline, reproducible-index properties,
312
+ and produce text a human can read and correct rather than opaque floats. See
313
+ [docs/ROADMAP.md](docs/ROADMAP.md).
314
+
315
+ ## Development
316
+
317
+ ```bash
318
+ python -m pip install -e ".[dev]"
319
+ pytest -q
320
+ ```
321
+
322
+ No runtime dependencies; `pytest` and, below Python 3.11, `tomli` come from the
323
+ `dev` extra.
324
+
325
+ - [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) — how each stage works and why
326
+ - [docs/TESTING.md](docs/TESTING.md) — what the suites are protecting
327
+ - [docs/ROADMAP.md](docs/ROADMAP.md) — what shipped, what is still open
328
+ - [CONTRIBUTING.md](CONTRIBUTING.md) — ground rules, and how to add a language
329
+ - [CHANGELOG.md](CHANGELOG.md) — every release, with its measurements
330
+
331
+ MIT licensed.