rag-your-code 1.4.2__tar.gz → 1.5.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (55) hide show
  1. rag_your_code-1.5.0/PKG-INFO +780 -0
  2. rag_your_code-1.5.0/README.md +750 -0
  3. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/pyproject.toml +5 -1
  4. rag_your_code-1.5.0/src/rag_your_code.egg-info/PKG-INFO +780 -0
  5. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/src/rag_your_code.egg-info/SOURCES.txt +1 -0
  6. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/src/rag_your_code.egg-info/requires.txt +1 -0
  7. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/src/ragyourcode/__init__.py +1 -1
  8. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/src/ragyourcode/cli.py +1 -0
  9. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/src/ragyourcode/config.py +14 -0
  10. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/src/ragyourcode/descriptions.py +41 -3
  11. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/src/ragyourcode/workflow.py +13 -4
  12. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_absent_queries.py +10 -2
  13. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_descriptions.py +83 -0
  14. rag_your_code-1.5.0/tests/test_diagrams.py +151 -0
  15. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_repo_queries.py +20 -3
  16. rag_your_code-1.4.2/PKG-INFO +0 -779
  17. rag_your_code-1.4.2/README.md +0 -750
  18. rag_your_code-1.4.2/src/rag_your_code.egg-info/PKG-INFO +0 -779
  19. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/LICENSE +0 -0
  20. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/setup.cfg +0 -0
  21. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/src/rag_your_code.egg-info/dependency_links.txt +0 -0
  22. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/src/rag_your_code.egg-info/entry_points.txt +0 -0
  23. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/src/rag_your_code.egg-info/top_level.txt +0 -0
  24. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/src/ragyourcode/agentic.py +0 -0
  25. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/src/ragyourcode/annotate.py +0 -0
  26. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/src/ragyourcode/document.py +0 -0
  27. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/src/ragyourcode/embeddings.py +0 -0
  28. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/src/ragyourcode/graph.py +0 -0
  29. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/src/ragyourcode/indexer.py +0 -0
  30. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/src/ragyourcode/models.py +0 -0
  31. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/src/ragyourcode/parser.py +0 -0
  32. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/src/ragyourcode/providers.py +0 -0
  33. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/src/ragyourcode/py.typed +0 -0
  34. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/src/ragyourcode/search.py +0 -0
  35. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_agent_protocol.py +0 -0
  36. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_agentic.py +0 -0
  37. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_config.py +0 -0
  38. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_doc_comments.py +0 -0
  39. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_document.py +0 -0
  40. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_e2e_cli.py +0 -0
  41. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_evidence.py +0 -0
  42. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_golden.py +0 -0
  43. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_graph_incremental.py +0 -0
  44. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_language_fixtures.py +0 -0
  45. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_large_repo.py +0 -0
  46. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_local_model.py +0 -0
  47. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_metadata.py +0 -0
  48. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_multilanguage.py +0 -0
  49. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_parser_edges.py +0 -0
  50. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_providers.py +0 -0
  51. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_ragyourcode.py +0 -0
  52. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_ranking.py +0 -0
  53. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_resilience.py +0 -0
  54. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_retrieval_correctness.py +0 -0
  55. {rag_your_code-1.4.2 → rag_your_code-1.5.0}/tests/test_workflow.py +0 -0
@@ -0,0 +1,780 @@
1
+ Metadata-Version: 2.4
2
+ Name: rag-your-code
3
+ Version: 1.5.0
4
+ Summary: A local, explainable RAG index for codebases and coding agents
5
+ Author: rag-your-code contributors
6
+ License-Expression: MIT
7
+ Keywords: rag,code-search,retrieval,indexing,graphrag,offline,explainable,agent
8
+ Classifier: Development Status :: 4 - Beta
9
+ Classifier: Environment :: Console
10
+ Classifier: Intended Audience :: Developers
11
+ Classifier: Operating System :: OS Independent
12
+ Classifier: Programming Language :: Python :: 3
13
+ Classifier: Programming Language :: Python :: 3.10
14
+ Classifier: Programming Language :: Python :: 3.11
15
+ Classifier: Programming Language :: Python :: 3.12
16
+ Classifier: Programming Language :: Python :: 3.13
17
+ Classifier: Topic :: Software Development :: Libraries
18
+ Classifier: Topic :: Text Processing :: Indexing
19
+ Classifier: Typing :: Typed
20
+ Requires-Python: >=3.10
21
+ Description-Content-Type: text/markdown
22
+ License-File: LICENSE
23
+ Provides-Extra: dev
24
+ Requires-Dist: pytest>=7; extra == "dev"
25
+ Requires-Dist: pytest-cov>=4; extra == "dev"
26
+ Requires-Dist: tomli>=2.0; python_version < "3.11" and extra == "dev"
27
+ Provides-Extra: sentence-transformers
28
+ Requires-Dist: sentence-transformers>=2.2; extra == "sentence-transformers"
29
+ Dynamic: license-file
30
+
31
+ # RAG Your Code
32
+
33
+ [![PyPI](https://img.shields.io/badge/PyPI-rag--your--code-blue)](https://pypi.org/project/rag-your-code/)
34
+ [![License](https://img.shields.io/badge/license-MIT-green)](LICENSE)
35
+ [![Python](https://img.shields.io/badge/python-3.10--3.13-blue)](pyproject.toml)
36
+
37
+ **A local code-retrieval index for coding agents.** Ask a question in plain
38
+ language; get back the declarations that answer it — each with its file, its
39
+ exact line range, the words it matched on, and its source.
40
+
41
+ Zero runtime dependencies. No network calls. No model required. It runs over a
42
+ private repository on a machine with the network switched off, and produces an
43
+ index a human can read.
44
+
45
+ It is the **R** in RAG. There is no generation here — your agent is the G.
46
+
47
+ ```bash
48
+ pip install rag-your-code
49
+ rag-your-code bootstrap .
50
+ rag-your-code search "where does it decide whether to answer at all" --json
51
+ ```
52
+
53
+ ---
54
+
55
+ ## 1 · The problem
56
+
57
+ An agent looking for something in an unfamiliar codebase has two bad options.
58
+
59
+ **Grep** is fast and exact, and it only finds the string you already guessed.
60
+ Ask "where does it decide whether to answer at all" and there is no string to
61
+ grep for. **Reading whole files** is thorough and blows the context budget.
62
+
63
+ Retrieval sits in between and brings a third problem neither has: Grep can say
64
+ it found nothing, and **a ranking cannot.** It always produces a least-bad
65
+ candidate and returns it with a score and rank that read exactly like an answer,
66
+ whether or not the repository holds anything relevant.
67
+
68
+ ## 2 · What it does
69
+
70
+ | | |
71
+ |---|---|
72
+ | **Index** | Every function, method and class in 15 languages becomes one `CodeUnit`: id, signature, exact line range, source, calls, imports, description. |
73
+ | **Retrieve** | BM25F over five weighted fields, blended with vector similarity. Results carry the terms they matched on. |
74
+ | **Refuse** | Two evidence tests decide whether *any* result is an answer; when neither is met, retrieval returns nothing plus a machine-readable diagnosis. |
75
+ | **Expand** | Optional bounded walk over `calls` / `imports` / `contains`, each hop carrying its edge path as evidence. |
76
+ | **Describe** | Your agent writes the vocabulary the source never contained, stored in a committed sidecar or promoted into the code as a reviewable diff. |
77
+ | **Serve** | A CLI, and a JSON-lines protocol for a long-lived agent subprocess. |
78
+
79
+ **Scope.** Retrieval over source declarations — not a code-understanding model,
80
+ not a generation step, not an IDE index. Questions are answered in vocabulary
81
+ somebody wrote down: in the code, its documentation, or an agent's description.
82
+
83
+ ## 3 · What is actually hard here
84
+
85
+ Three things, and all three are measured rather than argued.
86
+
87
+ ### 3.1 · Ranking cannot say "no answer"
88
+
89
+ Eight releases measured how well retrieval *finds* the answer. None could see
90
+ what it does when there is none, because every question graded had one. A
91
+ ruler of its own — thirty questions about subjects no graded repository
92
+ implements — settled it in one run: **all thirty answered**, both languages,
93
+ every corpus.
94
+
95
+ | asked of a repository containing no such code | answered with | on the evidence of |
96
+ |---|---|---|
97
+ | `where are CUDA kernels dispatched to the device` | a test about word counting | `are` `the` `to` `where` |
98
+ | `准入钩子为什么会拒绝没有资源限额的工作负载` | the UTF-8 console setup | `拒绝` `没有` `为什么` |
99
+ | `how is the OAuth refresh token rotated` | a description-store method | `before` `is` `refresh` `the` |
100
+
101
+ Not a Chinese problem and not a ranking problem — a **missing question**: nothing
102
+ in the pipeline ever asked *is any of this evidence*.
103
+
104
+ Retrieval now asks two questions that ranking cannot:
105
+
106
+ **Coverage** — what share of the query's *discriminating* words occur in the
107
+ index at all. Words the repository uses everywhere are dropped from both sides
108
+ of the fraction, and that is the part that does the work: half of `where are
109
+ CUDA kernels dispatched to the device` matches, and it looks like evidence
110
+ until you notice which half.
111
+
112
+ **Concentration** — what share of the query's *rarity* lands inside a single
113
+ declaration. Coverage alone asks whether each word occurs somewhere, which a
114
+ question about an unimplemented subject can satisfy entirely out of unrelated
115
+ units: four of six words in four declarations with nothing to do with the
116
+ question or with one another. Rarity-weighted rather than counted, because two
117
+ ordinary words are not better evidence than the rare word asked about.
118
+
119
+ Both are **ratios inside the query**, never thresholds on a score: a score
120
+ threshold is tied to whatever scale the ranking produces, and one here silently
121
+ stopped existing the moment BM25F changed that scale.
122
+
123
+ ### 3.2 · The vector was carrying nothing, and here is why
124
+
125
+ The default embedder is a signed feature hash. Ablating it entirely moves the
126
+ three positive rulers by **±1 question in either direction** while the vectors
127
+ occupy **72.1%** of the index. That was known since 0.6.0 and left unexplained.
128
+ The explanation, measured here:
129
+
130
+ - **Not saturation.** Median 56 distinct tokens per unit into 384 buckets, 0.4%
131
+ over the width; widening to 16,384 raises fidelity from r=0.40 to r=0.56 and
132
+ buys no ranking.
133
+ - **Not redundancy.** Its cosine correlates only **+0.45** with BM25F over
134
+ 26,490 scored candidates, so it does carry variance of its own.
135
+ - **The variance is the wrong variance.** A signed hash counts every token
136
+ equally, so the independent part of what it measures is precisely the
137
+ contribution of words that are everywhere — the part rarity weighting exists
138
+ to discard. Independent *noise*, not independent signal.
139
+ - **And it can only reorder.** Candidates come from the lexical half, so a
140
+ vector cannot make anything retrievable. Six of thirty-five foreign-ruler
141
+ questions have an accepted answer sharing **no token at all** with the query.
142
+
143
+ Eight replacement schemes were measured across releases — character n-grams,
144
+ random indexing, truncated SVD, posting-list signatures, a rarity-weighted hash,
145
+ call-graph diffusion, postings expansion, authored-fields-only. None beat using
146
+ no vector: **a vector computed from the same words cannot know anything the
147
+ words do not already say.** Making it useful takes a model, which is an
148
+ installable option and is measured below.
149
+
150
+ ### 3.3 · Retrieval reaches only what somebody wrote down
151
+
152
+ `retry_charge` tokenizes to one opaque term, not to *retry* and *charge*.
153
+ Splitting identifiers was measured with query and stored vectors rebuilt
154
+ together: equal or worse on every ruler, because the pieces are `get`, `find`
155
+ and `check`, which rarity weighting discounts.
156
+
157
+ So the vocabulary ladder is the answer, cheapest rung first:
158
+
159
+ | source | who wrote it | lives in | survives a refactor | cost |
160
+ |---|---|---|---|---|
161
+ | identifier, signature, body | author | the code | by construction | free |
162
+ | docstring / doc comment, 15 languages | author | the code | by construction | free |
163
+ | promoted description | agent | the code | by construction | one review |
164
+ | agent description | agent | a sidecar | needs machinery | tokens |
165
+
166
+ ## 4 · How it works
167
+
168
+ Drawn out, with the refusal path and the surfaces: **[docs/FLOW.md](docs/FLOW.md)**.
169
+
170
+ ```
171
+ your repository
172
+ → walk source files configurable ignores, suffixes, size cap
173
+ → parse declarations Python via its own AST; 14 languages via a
174
+ 3-layer line scanner + per-language rule table
175
+ → one CodeUnit each id, qualified name, signature, exact span,
176
+ source, calls, imports, serial, description
177
+ → embed signed hash (default) · local model · endpoint
178
+ → inverted index BM25F over name/signature/description/
179
+ relations/body, IDF derived from your corpus
180
+ → assess coverage + concentration → answer, or refuse
181
+ → rank lexical + weighted cosine
182
+ → optional graph expansion calls / imports / contains, evidence per hop
183
+ → results, or JSON-lines to an agent subprocess
184
+ ```
185
+
186
+ **Parsing.** Python uses the standard-library syntax tree, so nesting,
187
+ qualified names, call lists and spans are exact. Every other language goes
188
+ through three separated layers: a scanner reading one line at a time, a rule
189
+ table per language, and a span closer following brace depth, Ruby's `end`, or
190
+ the next declaration. Because a pattern never sees a second line, a reported
191
+ line number **is** the loop index and cannot drift, and no declaration can
192
+ swallow the ones after it. A 530-byte JavaScript file that took 12.6 s to parse
193
+ now takes 0.37 ms.
194
+
195
+ Qualified names come from the spans the closer already produced: nested inside
196
+ another's span *is* nested in it. One mechanism, so there is no second one to
197
+ disagree with it.
198
+
199
+ **Ranking.** BM25F with per-field length normalisation, which is the part that
200
+ matters: against one length for the whole unit, a body repeating a word forty
201
+ times beat the declaration named after it, raw count cancelling length penalty.
202
+
203
+ | field | weight | why |
204
+ |---|---|---|
205
+ | `name` | 8 | what the author called the thing |
206
+ | `signature` | 4 | what it takes and returns |
207
+ | `description` | 3 | what somebody said it does |
208
+ | `relations` | 2 | what it calls and imports |
209
+ | `body` | 1 | a mention |
210
+
211
+ **Rarity comes from your corpus, not a stopword list.** `the` and `calls` earn
212
+ their low weight the same way a Chinese bigram does — by being everywhere — so
213
+ no list is maintained and an unanticipated language works. It is also where the
214
+ design degrades: see the refusal table in section 6.
215
+
216
+ **Safety.** A scanned repository is untrusted input, including any
217
+ `.rag-your-code/index.json` it ships, so nothing read out of an index may name
218
+ a path to act on: superseded sidecars are enumerated from the writer's own
219
+ naming scheme. A crafted index once made `index` delete an in-tree file.
220
+
221
+ ## 5 · Before and after
222
+
223
+ A question with no lexical shortcut, asked of this repository:
224
+
225
+ ````console
226
+ $ rag-your-code search "where does it decide whether to answer at all" --limit 1
227
+ [src/ragyourcode/search.py:117:Evidence] score=0.447
228
+ The verdict on whether a question reached this index at all, kept separate from
229
+ how results rank. ... 中文:判定一个提问究竟有没有够到索引的结论。...
230
+ ```python
231
+ class Evidence:
232
+ """Whether a query reached this index at all, kept apart from ..."""
233
+ ```
234
+ ````
235
+
236
+ There is no string here to grep for: *decide* occurs nowhere in that
237
+ declaration and matched nothing. What ranked it first is ordinary words —
238
+ *answer*, *whether*, *where* — rare enough in this corpus to tell declarations
239
+ apart. What the agent-written description adds is the other language:
240
+ 「在哪里判定一个提问有没有答案」 returns the same declaration first, at 0.395,
241
+ sharing not one character with its source.
242
+
243
+ Now the case that motivated 1.0.0 — a question with no answer here at all:
244
+
245
+ ```console
246
+ $ rag-your-code search "why does the print spooler leave a duplex job stuck"
247
+ No matching code units.
248
+ The words that matched occur in this repository, but never together in one
249
+ place, so no single declaration is about what you asked. This is usually a
250
+ question about something the repository does not implement, described in words
251
+ it happens to use elsewhere.
252
+ ```
253
+
254
+ `--json` carries the same answer in a form an agent can branch on:
255
+
256
+ ```json
257
+ {"results": [],
258
+ "diagnosis": {"reason": "matched_terms_are_scattered",
259
+ "query_terms": 10,
260
+ "distinctive_terms": ["duplex","job","leave","print","spooler","stuck"],
261
+ "matched_terms": ["job","leave","print"],
262
+ "ubiquitous_terms": ["a","does","the","why"],
263
+ "coverage": 0.5, "min_coverage": 0.4,
264
+ "concentration": 0.1691, "min_concentration": 0.28,
265
+ "applied_min_coverage": 0.4, "applied_min_concentration": 0.28,
266
+ "hint": "..."}}
267
+ ```
268
+
269
+ Read `coverage: 0.5` against `concentration: 0.1691`. Half the distinctive
270
+ words are here — `job`, `leave`, `print` — and spread thin enough that no
271
+ declaration holds a fifth of what was asked, against a bar of 0.28. Before
272
+ 1.1.0 it came back with a confident-looking result.
273
+
274
+ Four reasons, because each is recovered by a different move:
275
+
276
+ | `reason` | what it means | what to do |
277
+ |---|---|---|
278
+ | `no_query_term_in_index` | no word of the question occurs anywhere | ask in the code's vocabulary |
279
+ | `only_ubiquitous_terms_matched` | only words the repository uses throughout | add a distinctive term |
280
+ | `too_little_of_the_query_matched` | most of the question is absent | rephrase, or write descriptions |
281
+ | `matched_terms_are_scattered` | the words are here, never together | the subject is probably not here |
282
+
283
+ ## 6 · Benchmark dashboard
284
+
285
+ Five rulers, 175 distinct questions in English and Chinese, graded 305 times —
286
+ one of them runs against all three corpora. Four grade whether the answer is
287
+ **found**; the fifth grades whether silence is **kept**. Every report carries a
288
+ fingerprint of the corpus it graded, because between two runs of an unchanged
289
+ `search.py` the foreign ruler moved 0.257 → 0.229 purely because that
290
+ repository had grown by ninety units.
291
+
292
+ **Accuracy — default embedder, zero dependencies**
293
+
294
+ | ruler | what it represents | n | hit@1 | hit@3 | MRR |
295
+ |---|---|---|---|---|---|
296
+ | **E** cobra v1.9.1, Go, no descriptions | a foreign repo in a foreign language | 40 | 0.075 | 0.150 | 0.108 |
297
+ | **A** Flask 3.1.3, no descriptions | what a first-time Python user gets | 35 | 0.200 | 0.286 | 0.238 |
298
+ | **B** this repo, generated descriptions only | a cold index of familiar code | 70 | 0.314 | 0.471 | 0.381 |
299
+ | **C** this repo, agent-written descriptions | the warmest case supported | 70 | 0.429 | 0.600 | 0.498 |
300
+
301
+ The corpora, without which none of the above is reproducible — **E** 602 units,
302
+ `3eabaa705477`; **A** 1,572 units, `5fd51169eacc`; **B** 601 units,
303
+ `566616fbe1e7`; **C** 601 units, `ac3ae43a33e7`. Both foreign subjects are
304
+ carried in this repository at pinned tags, under
305
+ [`benchmarks/corpus/`](benchmarks/corpus/), and CI runs both as ordinary jobs.
306
+
307
+ **The spread across those four rows is the honest headline.** The same code
308
+ scores 0.075 and 0.429 depending on nothing but which repository it is asked
309
+ about and whether anyone described it. Ruler E is 1.5.0's third corpus and the
310
+ first that is not Python: Go documents *above* the declaration in one terse
311
+ sentence beginning with the identifier, which the parser picks up correctly and
312
+ which shares almost nothing with the words a user asks in. Prose density, not
313
+ language, is what a cold number tracks.
314
+
315
+ **Refusal — the fifth ruler, 30 questions with no answer anywhere**
316
+
317
+ | | this repo | Flask | cobra |
318
+ |---|---|---|---|
319
+ | correctly met with silence | **0.967** | **0.833** | **0.900** |
320
+ | English only | **0.933** | 0.667 | 0.800 |
321
+ | Chinese only | **1.000** | **1.000** | **1.000** |
322
+ | results resting on no lexical evidence | **0.000** | **0.000** | **0.000** |
323
+
324
+ Silence is lower on both foreign corpora than on this one, and the cause is a
325
+ limit of the design rather than a defect. A word counts as evidence unless it
326
+ occurs in more than 5% of units — a stopword list derived from the corpus, so
327
+ that it needs no list and works in any language. Here `how`, `when`, `does` and
328
+ `are` are everywhere, because 314 units carry written English prose. Across a
329
+ corpus of short, tersely documented declarations they occur in 1–5% of them and
330
+ start counting as evidence: on cobra, three English questions get through on
331
+ sets like `[a, is, the, how, after]`.
332
+
333
+ **What each bar costs and buys** — every corpus, gate varied alone:
334
+
335
+ | gate | A | B | C | E | silence own / Flask / cobra |
336
+ |---|---|---|---|---|---|
337
+ | neither (pre-1.0.0) | 0.200/0.286/0.238 | 0.314/0.486/0.388 | 0.486/0.686/0.567 | 0.100/0.200/0.146 | 0.000 / 0.000 / 0.000 |
338
+ | coverage only (1.0.0) | 0.200/0.286/0.238 | 0.314/0.471/0.381 | 0.486/0.671/0.559 | 0.075/0.175/0.121 | 0.500 / 0.733 / 0.767 |
339
+ | concentration only | 0.200/0.286/0.238 | 0.314/0.471/0.381 | 0.429/0.600/0.498 | 0.075/0.150/0.108 | 0.967 / 0.800 / 0.833 |
340
+ | **both (1.1.0)** | **0.200/0.286/0.238** | **0.314/0.471/0.381** | 0.429/0.600/0.498 | 0.075/0.150/0.108 | **0.967 / 0.833 / 0.900** |
341
+
342
+ Ruler A is **unmoved by either bar**; B loses one hit@3 question to either bar
343
+ alone and nothing further when both apply. The rest of the cost is four of
344
+ seventy at hit@1 on the warmest ruler and six at hit@3, plus one of forty and
345
+ two of forty on the Go one.
346
+
347
+ **Both bars together give the best silence on all three corpora.** Through
348
+ 1.3.0 this section said concentration subsumes coverage; on Flask it does not
349
+ (0.833 against 0.800), and the third corpus, which arrived long after the
350
+ defaults were fixed, says the same (0.900 against 0.833). Four constants fitted
351
+ on two repositories, tested on a third in a language nobody here chose, and not
352
+ one of them moved.
353
+
354
+ Raising the bar buys the remaining silence out of the answers, and is refused:
355
+ at 0.50 the foreign absent ruler is silent on all thirty while A falls to 0.086
356
+ hit@1, B to 0.214 and C to 0.329. **0.28 was chosen before either foreign
357
+ corpus existed and survived meeting both**, which is the only kind of evidence
358
+ a default can have.
359
+
360
+ **Latency** — warm corpus, 601 units `ac3ae43a33e7`, one
361
+ `python -m benchmarks.query_latency` (5 repeats × 420 samples), idle machine:
362
+
363
+ | | | across the repeats |
364
+ |---|---|---|
365
+ | query, median | **0.49 ms** | 0.45 – 0.58 |
366
+ | query, p95 | 0.85 ms | 0.72 – 1.07 |
367
+ | refusing an unanswerable query | **0.016 ms** | 0.015 – 0.016 |
368
+ | refusal cheaper than answering by | **~30×** | 29 – 37 |
369
+
370
+ *Idle* is load-bearing: the same corpus at the same commit measured 0.99 ms
371
+ median while a coverage run was in progress and 0.49 ms once it finished.
372
+
373
+ Two significant figures and a spread, because that is the precision the
374
+ measurement has. Across twenty invocations over four releases on the same idle
375
+ machine the median has landed anywhere from 0.49 to 1.44 ms and p95 from 0.85
376
+ to 7.34 ms — a band wider than any change the code has made to this number.
377
+ Releases before 1.3.0 published `0.83 ms / p95 1.68 ms` to three figures from a
378
+ script that was never committed; both sit inside that band, which is the point:
379
+ they were unfalsifiable rather than wrong. Refusal is cheap structurally rather
380
+ than by tuning — an unanswerable query touches only the posting lists of its
381
+ own distinctive words and never reaches ranking.
382
+
383
+ **Scale**, synthetic 10,000-unit repository (500 files), re-measured in 1.5.0
384
+ — the previous row of figures was optimistic by more than noise:
385
+
386
+ | | | previously published |
387
+ |---|---|---|
388
+ | full build | 3.45 s | 1.84 s |
389
+ | incremental rebuild after one file changes | 0.286 s (**12.1×**) | 0.207 s |
390
+ | compact storage vs readable JSON | 35.6% | 35.6% |
391
+ | index load, fresh process | 79.3 ms | 45.4 ms |
392
+ | resident memory | 72.2 MiB | 58.7 MiB |
393
+ | mean query, full recall | 15.6 ms | 3.90 ms |
394
+
395
+ That last row had been carried since before BM25F replaced the scoring it was
396
+ taken under. Six releases, a committed script, and nothing that made anyone run
397
+ it again.
398
+
399
+ **Parsing**, against source-controlled fixtures (15 files, 237 negative cases):
400
+
401
+ | | |
402
+ |---|---|
403
+ | core declarations found | **91 / 91** |
404
+ | with the correct `start_line` | **91 / 91** |
405
+ | with a usable signature | **91 / 91** |
406
+ | units invented that do not exist | **0** |
407
+
408
+ Directional local measurements, not service levels — but each is a command
409
+ rather than a memory, which two of them were not before. Each
410
+ prints the corpus fingerprint beside its score; quote both or neither.
411
+ [`benchmarks/README.md`](benchmarks/README.md) lists the six scripts and what
412
+ each is for, and the corpus one of them grades is now carried here too.
413
+
414
+ ## 7 · `rag-your-code search` vs a Grep loop
415
+
416
+ The fair baseline is not one `grep`. An agent handed Grep picks the content
417
+ words out of the question, runs one search per word, and ranks files by how
418
+ many hit. That is what this reproduces — same corpus, same questions, same
419
+ ruler, scored at **file** granularity so Grep is not penalised for lacking
420
+ declaration spans.
421
+
422
+ **On an undescribed repository, which side wins depends on the repository.**
423
+ Three subjects, three answers:
424
+
425
+ | undescribed · Grep loop → rag-your-code | first | top 3 | answered | characters |
426
+ |---|---|---|---|---|
427
+ | **Flask** 35 q · 1,572 units `5fd51169eacc` | 22.9% → **37.1%** | 45.7% → **57.1%** | 30 → 30 | 1,415,656 → **258,236** |
428
+ | **cobra** 40 q · 602 units `3eabaa705477` | 17.5% → 17.5% | 20.0% → **30.0%** | 22 → 17 | 799,475 → **156,336** |
429
+ | a hook-heavy tool, retired in 1.4.0 | 34.3% → 31.4% | — | — | — |
430
+
431
+ Flask wins for this side, cobra ties on first place, and the retired subject
432
+ lost. A cold index retrieves against a generated sentence plus whatever the
433
+ author documented, so the outcome is set by how much prose the repository
434
+ already carries — Flask documents most public methods in paragraphs, cobra in
435
+ one terse line, the retired tool barely at all. What holds on all three is the
436
+ payload: this side hands back a fifth to a half of the text, ranked and spanned.
437
+
438
+ **Once the vocabulary exists, it is not close.**
439
+
440
+ | this repository · 70 questions · 601 units `ac3ae43a33e7` · 314 described | Grep loop | rag-your-code |
441
+ |---|---|---|
442
+ | right file first | 22.9% | **58.6%** |
443
+ | right file in top 3 | 54.3% | **78.6%** |
444
+ | lines it hands back, all questions | 12,421 | — |
445
+ | characters returned, all questions | 1,179,431 | **617,305** |
446
+ | questions it answers | **61** | 60 |
447
+
448
+ That is section 3.3's argument measured rather than asserted, and it is the one
449
+ thing the subject does not change: first-place accuracy more than doubles
450
+ Grep's.
451
+
452
+ **Every table comes from `python -m benchmarks.grep_baseline`**, new in 1.3.0.
453
+ Until then this section — the strongest claim the project makes — came from an
454
+ uncommitted script, so nothing here could be checked and "Grep loop" had no
455
+ precise meaning. The committed version defines it: take the query's words, drop
456
+ the ones the corpus itself shows are everywhere, run one substring search per
457
+ remaining word over exactly the files the index was built from, rank each file
458
+ by how many distinct words hit it, break ties on path.
459
+
460
+ Qualifications, because the tables would otherwise flatter both sides:
461
+
462
+ - **Scored at file granularity**, which understates this side. A Grep hit is a
463
+ file; a hit here is a declaration with an exact span, a score, and the words
464
+ it matched on. The agent that reads the result opens 40 lines, not a file.
465
+ - **Dropping the corpus-common words is generous to Grep**, and is what makes
466
+ the baseline fair rather than a straw man: an agent that greps `the` gets
467
+ every file back in no order. It is also why Grep declines nine of the seventy
468
+ questions here — no word was left that this corpus does not use everywhere.
469
+ - **Payload is counted in characters on both sides.** Grep hands back roughly
470
+ 19,300 characters per question it answers here, unranked and without spans,
471
+ against 10,300 ranked and capped by `search.max_chars` — a factor of 1.9,
472
+ 5.5 on Flask and 5.1 on cobra, where a framework repeats its vocabulary
473
+ across files and Grep cannot rank what it finds. 1.4.1 changed what fits in
474
+ that cap: the block had been reprinting the docstring the code below already
475
+ showed, so the same budget now carries **119 declarations instead of 92** on
476
+ Flask.
477
+ - **Chinese is the corpus's limit, not the tool's, when a corpus is
478
+ monolingual.** Both sides decline the same five of Flask's 35 — every Chinese
479
+ one. A Chinese word is neither a substring of English source nor a token in
480
+ an index built from it.
481
+ - **Grep wins outright when you know the string.** `grep -rn "COMMON_TERM"` is
482
+ exact, instant and complete, and nothing here replaces it.
483
+
484
+ The two are complementary, and the honest summary is narrow: this earns its
485
+ place on questions phrased as questions, over a repository somebody described.
486
+
487
+ ## 8 · Design principles
488
+
489
+ **Build the ruler before reshaping the thing measured.** Four candidate scoring
490
+ changes once landed between five and six correct over an eight-question set —
491
+ the instrument's resolution limit, not a ranking. There are 175 questions now
492
+ across five rulers, and every claim here is a number from one of them.
493
+
494
+ **Measure somewhere it can fail.** Every ruler this project had once graded a
495
+ repository its own authors wrote; cold against a foreign one the same code
496
+ scored 0.086 hit@1 against a self-reported 0.457. Rulers A and E exist so that
497
+ cannot be comfortable again — stemming, which helps both own-repo rulers, was
498
+ rejected on A, and E publishes a 0.075 this project would rather not print.
499
+
500
+ **Make the error structurally impossible rather than checking for it.** A line
501
+ number that *is* the loop index cannot drift; a description keyed by a digest
502
+ of its own code cannot outlive it. **And a ratio inside the query, never a
503
+ threshold on a score** — scales move, ratios do not.
504
+
505
+ **Derive figures from data; a hand-maintained number is a claim nobody checks.**
506
+ This README's settings table is asserted against `config.py` in both directions
507
+ — it had drifted nine settings behind before that test existed — and the four
508
+ diagrams in `docs/FLOW.md` are parsed by a test that refuses a label mermaid
509
+ would silently fail to render.
510
+
511
+ **The contract does not move.** `CodeUnit`, index schema 2 and the JSON-lines
512
+ protocol are unchanged across every release: new information arrives in new
513
+ fields, never by widening an enumeration callers branch on.
514
+
515
+ **Publish what was measured and rejected.** Twelve changes were implemented,
516
+ measured and dropped, with their numbers, in [docs/ROADMAP.md](docs/ROADMAP.md)
517
+ — "we tried that and it cost 3 of 35" beats an unexplored idea.
518
+
519
+ ## 9 · Bringing your own model
520
+
521
+ Everything above works with no model. Three embedders, and the difference
522
+ between them is what the vector is *able* to know.
523
+
524
+ ```toml
525
+ # rag-your-code.toml — a model that runs on your machine
526
+ [embedding]
527
+ provider = "sentence-transformers"
528
+ model = "sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2"
529
+ dimensions = 384
530
+ ```
531
+
532
+ ```bash
533
+ pip install "rag-your-code[sentence-transformers]"
534
+ ```
535
+
536
+ The extra is optional by construction: `dependencies = []` is a default
537
+ install, the import happens inside the constructor, and a test asserts the
538
+ default provider imports none of it.
539
+
540
+ **Measured on the four rulers that then existed, both arms against one
541
+ corpus.** 1.1.0 published this comparison and read it as a win; its largest
542
+ gain was on the foreign ruler, whose two arms turned out to have been taken
543
+ against two states of a repository being edited while the script ran. Repeated
544
+ against a pinned corpus:
545
+
546
+ | ruler | corpus | signed hash (default) | MiniLM, local |
547
+ |---|---|---|---|
548
+ | **A** foreign, cold | 1,572 `5fd51169eacc` | **0.200 / 0.286 / 0.238** | 0.171 / 0.257 / 0.214 |
549
+ | **B** own, cold | 581 `8e1e71942c1c` | 0.314 / 0.471 / 0.383 | 0.314 / 0.471 / 0.383 |
550
+ | **C** own, described | 581 `978a1d48a82a` | **0.443 / 0.614 / 0.507** | 0.429 / 0.600 / 0.500 |
551
+ | **D** silence, own / foreign | as above | 0.967 / 0.833 | 0.967 / 0.833 |
552
+
553
+ **Worse or identical on every ruler.** The 581-unit stamps are the corpus both
554
+ arms shared, kept rather than refreshed — that is what a stamp is for. It ships
555
+ anyway for the one thing the hash cannot do and these rulers cannot see: reach
556
+ a unit sharing no word with the question. The pairs it scores zero on:
557
+
558
+ | pair | signed hash | MiniLM |
559
+ |---|---|---|
560
+ | `retry a failed card charge` vs `resend a payment after a transient error` | 0.298 | **0.583** |
561
+ | `retry a failed card charge` vs `delete every row of the user table` | 0.000 | 0.073 |
562
+ | `计算两个数的和` vs `sum two numbers` | **0.000** | **0.822** |
563
+ | `刷新索引` vs `rebuild the index` | **0.000** | **0.684** |
564
+
565
+ A semantic embedder is **not** exempt from the evidence bars, correcting
566
+ 1.0.0. Exempting one was reasoned — a paraphrase sharing no word with its
567
+ answer is exactly what a model is for — and wrong: exempt and asked no other
568
+ question, the model answered all sixty unanswerable questions. Two vector-space
569
+ replacements were then measured and rejected: a similarity floor is a threshold
570
+ on a score and the distributions overlap (0.469 vs 0.418 median), and a
571
+ scale-free standout metric took ruler B from 0.329 to 0.186 for two thirds of
572
+ the silence. Applying the lexical bars costs ruler A nothing.
573
+
574
+ Install it for the cross-language and paraphrase cases above, not for the hit rates.
575
+
576
+ **A hosted endpoint** is the third option, and the only one that sends your
577
+ source anywhere:
578
+
579
+ ```toml
580
+ provider = "openai-compatible"
581
+ endpoint = "https://api.example.com/v1/embeddings"
582
+ model = "text-embedding-3-small"
583
+ dimensions = 1536
584
+ api_key_env = "OPENAI_API_KEY" # the NAME of the variable, never the key
585
+ ```
586
+
587
+ The key is never a setting: `rag-your-code.toml` is meant to be committed so
588
+ everyone who clones sees what shaped the index, and a credential is the one
589
+ value with the opposite requirement. Sending a key over plain `http://` to
590
+ anything but your own machine is refused rather than warned about, and a
591
+ failure stops the build rather than falling back — a mixed index is two vector
592
+ spaces, and a cosine across them is a meaningless number ranking would act on.
593
+
594
+ With a semantic embedder, similarity may also **add** candidates rather than
595
+ only reorder them (`search.vector_recall`) — the one thing that can reach a
596
+ unit sharing no word with the question. Under the hash it measured worse, so it
597
+ stays off there.
598
+
599
+ ## 10 · Install and use
600
+
601
+ **As a Claude Code plugin** (the primary way):
602
+
603
+ ```
604
+ /plugin marketplace add skymanbp/rag-your-code
605
+ /plugin install rag-your-code@rag-your-code
606
+ /reload-plugins
607
+ ```
608
+
609
+ Updating needs the full id and a marketplace refresh first; the bare name is
610
+ refused with `Plugin "rag-your-code" not found`, which reads like it is gone:
611
+
612
+ ```bash
613
+ claude plugin marketplace update rag-your-code
614
+ claude plugin update rag-your-code@rag-your-code # then restart
615
+ ```
616
+
617
+ Four commands and one skill. No hooks, no agents, no MCP server:
618
+
619
+ | | |
620
+ |---|---|
621
+ | `/rag-your-code:index` | index, and say which rung this repository is on |
622
+ | `/rag-your-code:search` | ask in plain language; cite `path:line` |
623
+ | `/rag-your-code:describe` | write the vocabulary the source does not contain |
624
+ | `/rag-your-code:status` | stale? coverage? which embedder? what next? |
625
+
626
+ Measured with `claude plugin details` on an installed copy: **~249 tokens added
627
+ to every session** (skill ~30, each command ~50–60), and 590–2,400 only when
628
+ one fires. Up from ~39 in 1.1.0, and the increase is the price of being
629
+ findable: a skill fires only when a model decides it should, which left the
630
+ plugin with no entry point a person could discover.
631
+
632
+ **As a CLI:**
633
+
634
+ ```bash
635
+ rag-your-code bootstrap . # index, then say what is missing
636
+ rag-your-code search "how are stale indexes detected" --json
637
+ rag-your-code search "what calls the retry handler" --graph --hops 1
638
+ rag-your-code describe status # description coverage
639
+ rag-your-code describe promote | git apply # move descriptions into the code
640
+ ```
641
+
642
+ `bootstrap` exists because indexing is not the same as being searchable: a
643
+ fresh index retrieves against a generated sentence that adds no word the source
644
+ did not have. It reports which rung the repository is on and hands over that
645
+ rung's work; run it again after each round. The index is written under
646
+ `.rag-your-code/` and **your source files are never modified** — `describe
647
+ promote` emits a diff to review rather than writing source.
648
+
649
+ ### Configuration
650
+
651
+ 23 settings in `rag-your-code.toml`:
652
+
653
+ | section | settings |
654
+ |---|---|
655
+ | `[index]` | `ignore`, `suffixes`, `max_file_bytes` |
656
+ | `[embedding]` | `dimensions`, `provider`, `endpoint`, `model`, `api_key_env`, `batch`, `timeout`, `retries` |
657
+ | `[search]` | `min_coverage`, `min_concentration`, `vector_weight`, `vector_recall`, `limit`, `max_chars` |
658
+ | `[agent]` | `max_open_bytes`, `max_open_chars` |
659
+ | `[describe]` | `languages`, `batch`, `max_chars`, `skip` |
660
+
661
+ This table is asserted against `config.py` in both directions by
662
+ `tests/test_metadata.py`. Resolution is CLI flag > file > built-in default;
663
+ there is deliberately no environment layer, because an index is an artifact of
664
+ a repository rather than of a shell. An unknown key or out-of-range value is an
665
+ error, not a shrug.
666
+
667
+ ### Agent protocol
668
+
669
+ `rag-your-code agent --root PATH` reads one JSON request per line, writes one
670
+ reply per line:
671
+
672
+ ```json
673
+ {"action":"search","query":"database transaction rollback","limit":5}
674
+ {"action":"research","query":"trace payment retry behavior","max_steps":2}
675
+ {"action":"neighbors","id":"payments.py:4:retry_charge","hops":1}
676
+ {"action":"open","path":"payments.py","start_line":1,"end_line":80}
677
+ {"action":"describe_pending","limit":20}
678
+ {"action":"describe_put","descriptions":[{"id":"payments.py:4:retry_charge","text":"..."}]}
679
+ ```
680
+
681
+ **A result is navigation, not the file.** The code arrives once, in `context`,
682
+ trimmed to `max_chars`, with `omitted_for_budget` saying how many results it
683
+ did not reach. Carrying source per result is what let one `search --json` reply
684
+ reach 65,025 characters against a stated budget of 12,000.
685
+
686
+ **No single request can end the session.** Numeric fields saturate at their
687
+ bounds, `open` is bounded in lines and bytes, and anything unanticipated is
688
+ reported in-band with its exception type. Streams are pinned to UTF-8 rather
689
+ than following the console codepage.
690
+
691
+ ### What lives where
692
+
693
+ | path | authored or generated | commit it? |
694
+ |---|---|---|
695
+ | `rag-your-code.toml` | authored | yes |
696
+ | `rag-your-code.descriptions.json` | authored by your agent | yes |
697
+ | `.rag-your-code/` | generated | no |
698
+
699
+ ## 11 · Known limits
700
+
701
+ Named because they are measured, not because they are excuses.
702
+
703
+ **One English question in fifteen is answered when it should not be** on this
704
+ repository, and five in fifteen on Flask. `how is a hostname resolved when the
705
+ nameserver times out` finds `hostname`, `resolved` and `times` genuinely
706
+ co-occurring in one unrelated declaration. No lexical rule separates a real
707
+ vocabulary collision from a real answer. Chinese sits at 1.000 silence on both.
708
+
709
+ **Chinese cold-start hit@1 is 0.000** on rulers A and B. Chinese reaches a
710
+ repository through descriptions or not at all: the code contains no Chinese, so
711
+ a cold index has no Chinese vocabulary to match. `describe` is the fix and it
712
+ works — ruler C is 0.250 on its twelve Chinese questions — but there is no free
713
+ rung of the ladder for it. A Grep loop scores 0.000 there too, on the same
714
+ questions: it is the corpus's limit, not this tool's.
715
+
716
+ **There is no stemming.** `catastrophic backtracking` does not reach
717
+ `backtracks catastrophically`. A light suffix stripper was implemented and
718
+ measured on every ruler that then existed: it improves both own-repository
719
+ rulers and costs the foreign one 3 of 35 hit@3, so it was rejected.
720
+
721
+ **A test declaration sometimes outranks real code** — 9 of 175 questions across
722
+ three rulers, a test at rank 1 displacing an accepted answer at rank 2–3, and
723
+ none of them on Flask. The long-standing explanation, that a test outranks the
724
+ code it *tests*, is wrong: five of the nine are unrelated tests winning on
725
+ prose. A callee-before-caller rerank fires on zero questions, and the `name`
726
+ field weight moves nothing because an underscored test name is one token.
727
+
728
+ **Describing a declaration that already has a good docstring loses ground.** An
729
+ authored description *replaces* the generated sentence, which is the only route
730
+ by which the author's own docstring reaches the weight-3 description field — so
731
+ writing one demotes that docstring to the weight-1 body. Measured on
732
+ `parser.py::_generic_units`: a long description cost one graded question, a
733
+ short one cost three, and appending the docstring to all 314 descriptions
734
+ instead cost 0.443 → 0.414 hit@1. `describe.skip` records the decision.
735
+
736
+ **The vectors are 72.1% of the index and earn ±1 question** under the default
737
+ embedder — 74.8% on Flask and 79.7% on cobra. Kept: the same storage is what
738
+ makes an optional model work.
739
+
740
+ **`search.vector_recall` scans every vector per query** — under a semantic
741
+ embedder. The default hash never widens at all. Affordable at the measured
742
+ envelope, and exactly the work an ANN index would replace.
743
+
744
+ **Tree-sitter parsing and a SQLite/ANN storage layer are not here.** Both need
745
+ a dependency, and the policy is settled: they follow the embedding provider's
746
+ pattern — optional, user-selected, never in a default install. Full reasoning
747
+ in [docs/ROADMAP.md](docs/ROADMAP.md).
748
+
749
+ **Whether the skill fires unprompted is not measured**, and it is the only
750
+ claim here with no command behind it. `claude plugin eval` grades exactly this
751
+ — `tool_used: Skill` as the indicator, against a no-plugin baseline arm — but
752
+ it is gated behind an account-level early access this project does not have,
753
+ and a suite written from `--help` fragments could not be run once to see
754
+ whether it loads. Since 1.2.0 the four commands give an entry path that does
755
+ not depend on it.
756
+
757
+ ## 12 · Development
758
+
759
+ ```bash
760
+ python -m pip install -e ".[dev]"
761
+ pytest -q
762
+ ```
763
+
764
+ Per-release test counts and line coverage are in
765
+ [CHANGELOG.md](CHANGELOG.md); a bare figure in a living document is a claim
766
+ that rots. `pytest --cov=ragyourcode` is the command behind the coverage one.
767
+ CI runs Python 3.10–3.13 on Linux and Windows, installs the built wheel into a
768
+ clean environment and runs the documented CLI end to end — `bootstrap` through
769
+ `describe promote` — plus the skill's own install line verbatim, and grades
770
+ every ruler including both vendored corpora.
771
+
772
+ - [docs/FLOW.md](docs/FLOW.md) — the whole thing in four diagrams
773
+ - [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) — how each stage works and why
774
+ - [docs/TESTING.md](docs/TESTING.md) — what the suites protect, and how covered
775
+ - [docs/ROADMAP.md](docs/ROADMAP.md) — what shipped, what was rejected and why
776
+ - [CONTRIBUTING.md](CONTRIBUTING.md) — ground rules, and how to add a language
777
+ - [CHANGELOG.md](CHANGELOG.md) — every release with its measurements
778
+ - [benchmarks/corpus/](benchmarks/corpus/) — the two vendored repositories
779
+
780
+ MIT licensed.