rag-your-code 1.3.0__tar.gz → 1.4.1__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {rag_your_code-1.3.0/src/rag_your_code.egg-info → rag_your_code-1.4.1}/PKG-INFO +121 -72
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/README.md +120 -71
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/pyproject.toml +1 -1
- {rag_your_code-1.3.0 → rag_your_code-1.4.1/src/rag_your_code.egg-info}/PKG-INFO +121 -72
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/src/ragyourcode/__init__.py +1 -1
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/src/ragyourcode/annotate.py +9 -1
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/src/ragyourcode/document.py +1 -1
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/src/ragyourcode/parser.py +2 -2
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/src/ragyourcode/search.py +36 -1
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_absent_queries.py +42 -9
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_ranking.py +54 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/LICENSE +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/setup.cfg +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/src/rag_your_code.egg-info/SOURCES.txt +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/src/rag_your_code.egg-info/dependency_links.txt +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/src/rag_your_code.egg-info/entry_points.txt +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/src/rag_your_code.egg-info/requires.txt +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/src/rag_your_code.egg-info/top_level.txt +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/src/ragyourcode/agentic.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/src/ragyourcode/cli.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/src/ragyourcode/config.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/src/ragyourcode/descriptions.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/src/ragyourcode/embeddings.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/src/ragyourcode/graph.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/src/ragyourcode/indexer.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/src/ragyourcode/models.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/src/ragyourcode/providers.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/src/ragyourcode/py.typed +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/src/ragyourcode/workflow.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_agent_protocol.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_agentic.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_config.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_descriptions.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_doc_comments.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_document.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_e2e_cli.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_evidence.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_golden.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_graph_incremental.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_language_fixtures.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_large_repo.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_local_model.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_metadata.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_multilanguage.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_parser_edges.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_providers.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_ragyourcode.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_repo_queries.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_resilience.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_retrieval_correctness.py +0 -0
- {rag_your_code-1.3.0 → rag_your_code-1.4.1}/tests/test_workflow.py +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: rag-your-code
|
|
3
|
-
Version: 1.
|
|
3
|
+
Version: 1.4.1
|
|
4
4
|
Summary: A local, explainable RAG index for codebases and coding agents
|
|
5
5
|
Author: rag-your-code contributors
|
|
6
6
|
License-Expression: MIT
|
|
@@ -293,60 +293,83 @@ repository had grown by ninety units.
|
|
|
293
293
|
|
|
294
294
|
| ruler | what it represents | n | hit@1 | hit@3 | MRR |
|
|
295
295
|
|---|---|---|---|---|---|
|
|
296
|
-
| **A**
|
|
296
|
+
| **A** Flask 3.1.3, no descriptions | what a first-time user gets | 35 | 0.200 | 0.286 | 0.238 |
|
|
297
297
|
| **B** this repo, generated descriptions only | a cold index of familiar code | 70 | 0.314 | 0.471 | 0.383 |
|
|
298
|
-
| **C** this repo, agent-written descriptions | the warmest case supported | 70 | 0.443 | 0.614 | 0.
|
|
299
|
-
|
|
300
|
-
The corpora, without which none of the above is reproducible — **A**
|
|
301
|
-
|
|
302
|
-
|
|
303
|
-
|
|
304
|
-
|
|
305
|
-
|
|
306
|
-
0.
|
|
298
|
+
| **C** this repo, agent-written descriptions | the warmest case supported | 70 | 0.443 | 0.614 | 0.509 |
|
|
299
|
+
|
|
300
|
+
The corpora, without which none of the above is reproducible — **A** 1,572
|
|
301
|
+
units, `5fd51169eacc`; **B** 584 units, `fb1f841fa43a`; **C** 584 units,
|
|
302
|
+
`c9df00350cbd`. Ruler A grades **a copy of Flask 3.1.3 carried in this
|
|
303
|
+
repository**, at [`benchmarks/corpus/flask`](benchmarks/corpus/flask), pinned
|
|
304
|
+
to commit `22d9247`. Through 1.3.0 it graded a checkout on one machine, and
|
|
305
|
+
that cost three things: two questions pointed at a declaration the subject had
|
|
306
|
+
renamed, a published score moved 0.257 → 0.229 with no code change because the
|
|
307
|
+
subject had grown, and the model comparison below was taken against two
|
|
308
|
+
different states of it. All three are now a `git clone` away from being
|
|
309
|
+
checked, and CI runs this ruler as an ordinary job.
|
|
307
310
|
|
|
308
311
|
**Refusal — the fourth ruler, 30 questions with no answer anywhere**
|
|
309
312
|
|
|
310
|
-
| | this repo |
|
|
313
|
+
| | this repo | Flask |
|
|
311
314
|
|---|---|---|
|
|
312
|
-
| correctly met with silence | **0.967** | **0.
|
|
313
|
-
| English only | **0.933** |
|
|
315
|
+
| correctly met with silence | **0.967** | **0.833** |
|
|
316
|
+
| English only | **0.933** | 0.667 |
|
|
314
317
|
| Chinese only | **1.000** | **1.000** |
|
|
315
318
|
| results resting on no lexical evidence, rulers A–C | **0.000** | **0.000** |
|
|
316
319
|
|
|
320
|
+
Foreign silence fell from 0.933 to 0.833 when the subject changed, and the
|
|
321
|
+
cause is a limit of the design rather than a defect. A word counts as evidence
|
|
322
|
+
unless it occurs in more than 5% of units — a stopword list derived from the
|
|
323
|
+
corpus, so that it needs no list and works in any language. Here `how`, `when`,
|
|
324
|
+
`does` and `are` are everywhere, because 304 units carry written English prose.
|
|
325
|
+
Across 1,572 units of mostly short, undocumented methods they occur in 1–5% of
|
|
326
|
+
them and start counting as evidence. Five English questions about subjects
|
|
327
|
+
Flask does not implement get through on exactly that.
|
|
328
|
+
|
|
317
329
|
**What each bar costs and buys** — one corpus, gate varied alone:
|
|
318
330
|
|
|
319
331
|
| gate | A hit@1/3/MRR | B hit@1/3/MRR | C hit@1/3/MRR | silence own / foreign |
|
|
320
332
|
|---|---|---|---|---|
|
|
321
|
-
| neither (pre-1.0.0) | 0.
|
|
322
|
-
| coverage only (1.0.0) | 0.
|
|
323
|
-
|
|
|
324
|
-
|
|
325
|
-
|
|
326
|
-
|
|
327
|
-
|
|
328
|
-
|
|
329
|
-
|
|
330
|
-
|
|
331
|
-
|
|
333
|
+
| neither (pre-1.0.0) | 0.200/0.286/0.238 | 0.314/0.486/0.391 | 0.486/0.700/0.569 | 0.000 / 0.000 |
|
|
334
|
+
| coverage only (1.0.0) | 0.200/0.286/0.238 | 0.314/0.471/0.383 | 0.486/0.686/0.564 | 0.567 / 0.733 |
|
|
335
|
+
| concentration only | 0.200/0.286/0.238 | 0.314/0.471/0.383 | 0.443/0.614/0.509 | 0.967 / 0.800 |
|
|
336
|
+
| **both (1.1.0)** | **0.200/0.286/0.238** | **0.314/0.471/0.383** | 0.443/0.614/0.509 | **0.967 / 0.833** |
|
|
337
|
+
|
|
338
|
+
Ruler A is **unmoved by either bar**, and B by concentration. The whole cost is
|
|
339
|
+
three questions of seventy at hit@1 on the warmest ruler, and six at hit@3.
|
|
340
|
+
|
|
341
|
+
Through 1.3.0 this section said concentration subsumes coverage. **On a corpus
|
|
342
|
+
this project did not choose, it does not.** Both bars together silence 0.833 of
|
|
343
|
+
the foreign absent questions, against 0.800 for concentration alone and 0.733
|
|
344
|
+
for coverage alone. One question — and the first time in four releases that
|
|
345
|
+
keeping both has been worth a measurable amount rather than worth a different
|
|
346
|
+
diagnosis.
|
|
347
|
+
|
|
348
|
+
Raising the concentration bar buys the remaining silence, and is refused,
|
|
349
|
+
because it is bought out of the answers: at 0.50 the foreign absent ruler is
|
|
350
|
+
silent on all thirty while ruler A falls to 0.086 hit@1 from 0.200, B to 0.214
|
|
351
|
+
from 0.314 and C to 0.329 from 0.443. **0.28 was chosen before this corpus
|
|
352
|
+
existed and survived meeting it**, which is the only kind of evidence a default
|
|
353
|
+
can have.
|
|
354
|
+
|
|
355
|
+
**Latency** — warm corpus, 584 units `c9df00350cbd`, five consecutive
|
|
332
356
|
invocations of `python -m benchmarks.query_latency --repeats 10` (420 samples
|
|
333
357
|
each):
|
|
334
358
|
|
|
335
359
|
| | | across the five |
|
|
336
360
|
|---|---|---|
|
|
337
|
-
| query, median | **0.
|
|
338
|
-
| query, p95 | 1.
|
|
339
|
-
| refusing an unanswerable query | **0.
|
|
340
|
-
| refusal cheaper than answering by | **~
|
|
361
|
+
| query, median | **0.61 ms** | 0.60 – 0.64 |
|
|
362
|
+
| query, p95 | 1.09 ms | 1.06 – 1.14 |
|
|
363
|
+
| refusing an unanswerable query | **0.017 ms** | 0.015 – 0.019 |
|
|
364
|
+
| refusal cheaper than answering by | **~36×** | 33 – 41 |
|
|
341
365
|
|
|
342
366
|
Two significant figures and a spread, because that is the precision the
|
|
343
|
-
measurement has.
|
|
344
|
-
machine
|
|
345
|
-
|
|
346
|
-
|
|
347
|
-
|
|
348
|
-
|
|
349
|
-
were unfalsifiable rather than wrong.
|
|
367
|
+
measurement has. Across twenty invocations over three releases on the same idle
|
|
368
|
+
machine the median has landed anywhere from 0.51 to 1.44 ms and p95 from 0.85
|
|
369
|
+
to 7.34 ms — a band wider than any change the code has ever made to this
|
|
370
|
+
number. Releases before 1.3.0 published `0.83 ms / p95 1.68 ms` to three
|
|
371
|
+
figures from a script that was never committed; both values sit inside that
|
|
372
|
+
band, which is the point: they were unfalsifiable rather than wrong.
|
|
350
373
|
|
|
351
374
|
Refusal is cheap for a structural reason, not a tuned one: an unanswerable
|
|
352
375
|
query touches only the posting lists of its own distinctive words, and never
|
|
@@ -375,8 +398,8 @@ reaches ranking at all.
|
|
|
375
398
|
Directional local measurements, not service levels — but every one of them is
|
|
376
399
|
now a command rather than a memory, which two of them were not before. Each
|
|
377
400
|
prints the corpus fingerprint beside its score; quote both or neither.
|
|
378
|
-
[`benchmarks/README.md`](benchmarks/README.md) lists the
|
|
379
|
-
each is for.
|
|
401
|
+
[`benchmarks/README.md`](benchmarks/README.md) lists the six scripts and what
|
|
402
|
+
each is for, and the corpus one of them grades is now carried here too.
|
|
380
403
|
|
|
381
404
|
## 7 · `rag-your-code search` vs a Grep loop
|
|
382
405
|
|
|
@@ -386,34 +409,42 @@ many hit. That is what this reproduces — same corpus, same questions, same
|
|
|
386
409
|
ruler, scored at **file** granularity so Grep is not penalised for lacking
|
|
387
410
|
declaration spans.
|
|
388
411
|
|
|
389
|
-
**
|
|
390
|
-
|
|
412
|
+
**Which side wins on an undescribed repository depends on the repository.**
|
|
413
|
+
Through 1.3.0 this section said flatly that Grep wins there, because the one
|
|
414
|
+
undescribed repository ever measured was a hook-heavy tool whose questions were
|
|
415
|
+
answerable by matching identifiers. Swapping the subject for a public web
|
|
416
|
+
framework reversed it. The honest claim is narrower than either table alone:
|
|
391
417
|
|
|
392
|
-
|
|
|
418
|
+
| Flask 3.1.3 · 35 questions · 1,572 units `5fd51169eacc` · no descriptions | Grep loop | rag-your-code |
|
|
393
419
|
|---|---|---|
|
|
394
|
-
| right file first |
|
|
395
|
-
| right file in top 3 |
|
|
396
|
-
| lines it hands back, all questions |
|
|
397
|
-
| characters returned, all questions |
|
|
398
|
-
| questions it answers |
|
|
420
|
+
| right file first | 22.9% | **37.1%** |
|
|
421
|
+
| right file in top 3 | 45.7% | **57.1%** |
|
|
422
|
+
| lines it hands back, all questions | 17,641 | — |
|
|
423
|
+
| characters returned, all questions | 1,415,656 | **258,236** |
|
|
424
|
+
| questions it answers | 30 | 30 |
|
|
399
425
|
|
|
400
426
|
**Once the vocabulary exists, it is not close.**
|
|
401
427
|
|
|
402
|
-
| this repository · 70 questions ·
|
|
428
|
+
| this repository · 70 questions · 584 units `c9df00350cbd` · 304 described | Grep loop | rag-your-code |
|
|
403
429
|
|---|---|---|
|
|
404
430
|
| right file first | 22.9% | **58.6%** |
|
|
405
|
-
| right file in top 3 |
|
|
406
|
-
| lines it hands back, all questions | 11,
|
|
407
|
-
| characters returned, all questions | 1,
|
|
431
|
+
| right file in top 3 | 54.3% | **77.1%** |
|
|
432
|
+
| lines it hands back, all questions | 11,959 | — |
|
|
433
|
+
| characters returned, all questions | 1,135,411 | **615,673** |
|
|
408
434
|
| questions it answers | **61** | 60 |
|
|
409
435
|
|
|
410
436
|
Those two tables are the whole argument of section 3.3, measured against a real
|
|
411
437
|
baseline instead of asserted. A cold index retrieves against a sentence the
|
|
412
|
-
parser generated from identifiers the author already chose
|
|
413
|
-
|
|
414
|
-
|
|
415
|
-
|
|
416
|
-
|
|
438
|
+
parser generated from identifiers the author already chose, plus whatever
|
|
439
|
+
docstrings the author wrote — so how it fares against Grep is decided by how
|
|
440
|
+
much prose the repository already contains. Flask has a written docstring on
|
|
441
|
+
most public methods, and the cold index beats Grep there without a single
|
|
442
|
+
description being added. On the previous subject, a tool with terse comments
|
|
443
|
+
and long identifiers, the same cold index lost to Grep by the same margin.
|
|
444
|
+
|
|
445
|
+
What does not depend on the subject is what descriptions buy: on this
|
|
446
|
+
repository first-place accuracy goes to **more than double** Grep's, and the
|
|
447
|
+
payload comes back ranked, spanned, and roughly half the size.
|
|
417
448
|
|
|
418
449
|
**Both tables come from `python -m benchmarks.grep_baseline`**, which is what
|
|
419
450
|
changed in 1.3.0. Until then this section — the strongest claim the project
|
|
@@ -436,12 +467,22 @@ Four qualifications, because the table would otherwise flatter both sides:
|
|
|
436
467
|
`the` gets every file back in no order. It is also why Grep declines nine of
|
|
437
468
|
the seventy questions here — those had no word left that this corpus does not
|
|
438
469
|
use everywhere.
|
|
439
|
-
- **Payload is counted in characters on both sides
|
|
440
|
-
|
|
441
|
-
|
|
442
|
-
|
|
443
|
-
|
|
444
|
-
|
|
470
|
+
- **Payload is counted in characters on both sides**, and 1.4.1 changed what
|
|
471
|
+
fits in it. A generated description ends with the author's docstring so the
|
|
472
|
+
docstring is searchable, and the source printed below it said the same thing
|
|
473
|
+
again — 2,381 of 3,382 characters of prose header on Flask were a verbatim
|
|
474
|
+
repeat of the code beneath. The block no longer prints what the code shows,
|
|
475
|
+
and the same 12,000-character budget now carries **119 declarations instead
|
|
476
|
+
of 92** there, 323 instead of 305 here. Grep
|
|
477
|
+
hands back 18,600 characters per question it answers, unranked and without
|
|
478
|
+
spans, against 10,300 here, ranked and capped by `search.max_chars` — a
|
|
479
|
+
factor of 1.8. On Flask it is a factor of 5.5 — 47,200 characters against
|
|
480
|
+
8,600 — because a framework repeats its own vocabulary across many files and
|
|
481
|
+
Grep has no way to rank what it finds. Both sides decline the same five of
|
|
482
|
+
those 35 — and they are the five Chinese ones, all of them. A Chinese word is
|
|
483
|
+
not a substring of English source and it is not a token in an index built
|
|
484
|
+
from English source, so on a repository written in one language the cold
|
|
485
|
+
cross-language case is not this tool's failure but the corpus's.
|
|
445
486
|
- **Grep wins outright when you know the string.** `grep -rn "COMMON_TERM"` is
|
|
446
487
|
exact, instant and complete, and nothing here replaces it.
|
|
447
488
|
|
|
@@ -506,19 +547,23 @@ The extra is optional by construction: `dependencies = []` is what a default
|
|
|
506
547
|
install gets, the import happens inside the constructor, and a test asserts the
|
|
507
548
|
default provider imports none of it.
|
|
508
549
|
|
|
509
|
-
**Measured
|
|
510
|
-
|
|
511
|
-
|
|
550
|
+
**Measured on the same four rulers, both arms against one corpus.** 1.1.0
|
|
551
|
+
published this comparison and read it as a win. Its largest gain was on the
|
|
552
|
+
foreign ruler, whose two arms turned out to have been taken against two
|
|
553
|
+
different states of a repository being edited while the script ran. Repeated
|
|
554
|
+
against a pinned corpus:
|
|
512
555
|
|
|
513
|
-
| ruler | signed hash (default) | MiniLM, local |
|
|
514
|
-
|
|
515
|
-
| **A** foreign, cold | 0.
|
|
516
|
-
| **B** own, cold | 0.314 / 0.471 / 0.383 |
|
|
517
|
-
| **C** own, described | 0.443 / 0.614 / 0.507 | 0.
|
|
518
|
-
| **D** silence, own / foreign | 0.967 / 0.
|
|
556
|
+
| ruler | corpus | signed hash (default) | MiniLM, local |
|
|
557
|
+
|---|---|---|---|
|
|
558
|
+
| **A** foreign, cold | 1,572 `5fd51169eacc` | **0.200 / 0.286 / 0.238** | 0.171 / 0.257 / 0.214 |
|
|
559
|
+
| **B** own, cold | 581 `8e1e71942c1c` | 0.314 / 0.471 / 0.383 | 0.314 / 0.471 / 0.383 |
|
|
560
|
+
| **C** own, described | 581 `978a1d48a82a` | **0.443 / 0.614 / 0.507** | 0.429 / 0.600 / 0.500 |
|
|
561
|
+
| **D** silence, own / foreign | as above | 0.967 / 0.833 | 0.967 / 0.833 |
|
|
519
562
|
|
|
520
|
-
|
|
521
|
-
|
|
563
|
+
**Worse or identical on every ruler.** It is shipped anyway, as an extra nobody
|
|
564
|
+
has to install, because it does one thing the hash cannot do at all and these
|
|
565
|
+
rulers cannot see: reach a unit that shares no word with the question. The
|
|
566
|
+
pairs the hash scores exactly zero:
|
|
522
567
|
|
|
523
568
|
| pair | signed hash | MiniLM |
|
|
524
569
|
|---|---|---|
|
|
@@ -536,6 +581,10 @@ a threshold on a score and the distributions overlap (0.469 vs 0.418 median),
|
|
|
536
581
|
and a scale-free standout metric took ruler B from 0.329 to 0.186 for two
|
|
537
582
|
thirds of the silence. Applying the lexical bars costs ruler A nothing.
|
|
538
583
|
|
|
584
|
+
If you install it expecting the hit rates above to move, they will not. Install
|
|
585
|
+
it for the cross-language and paraphrase cases in the table above, which is
|
|
586
|
+
where the difference between the two columns actually lives.
|
|
587
|
+
|
|
539
588
|
**A hosted endpoint** is the third option, and the only one that sends your
|
|
540
589
|
source anywhere:
|
|
541
590
|
|
|
@@ -264,60 +264,83 @@ repository had grown by ninety units.
|
|
|
264
264
|
|
|
265
265
|
| ruler | what it represents | n | hit@1 | hit@3 | MRR |
|
|
266
266
|
|---|---|---|---|---|---|
|
|
267
|
-
| **A**
|
|
267
|
+
| **A** Flask 3.1.3, no descriptions | what a first-time user gets | 35 | 0.200 | 0.286 | 0.238 |
|
|
268
268
|
| **B** this repo, generated descriptions only | a cold index of familiar code | 70 | 0.314 | 0.471 | 0.383 |
|
|
269
|
-
| **C** this repo, agent-written descriptions | the warmest case supported | 70 | 0.443 | 0.614 | 0.
|
|
270
|
-
|
|
271
|
-
The corpora, without which none of the above is reproducible — **A**
|
|
272
|
-
|
|
273
|
-
|
|
274
|
-
|
|
275
|
-
|
|
276
|
-
|
|
277
|
-
0.
|
|
269
|
+
| **C** this repo, agent-written descriptions | the warmest case supported | 70 | 0.443 | 0.614 | 0.509 |
|
|
270
|
+
|
|
271
|
+
The corpora, without which none of the above is reproducible — **A** 1,572
|
|
272
|
+
units, `5fd51169eacc`; **B** 584 units, `fb1f841fa43a`; **C** 584 units,
|
|
273
|
+
`c9df00350cbd`. Ruler A grades **a copy of Flask 3.1.3 carried in this
|
|
274
|
+
repository**, at [`benchmarks/corpus/flask`](benchmarks/corpus/flask), pinned
|
|
275
|
+
to commit `22d9247`. Through 1.3.0 it graded a checkout on one machine, and
|
|
276
|
+
that cost three things: two questions pointed at a declaration the subject had
|
|
277
|
+
renamed, a published score moved 0.257 → 0.229 with no code change because the
|
|
278
|
+
subject had grown, and the model comparison below was taken against two
|
|
279
|
+
different states of it. All three are now a `git clone` away from being
|
|
280
|
+
checked, and CI runs this ruler as an ordinary job.
|
|
278
281
|
|
|
279
282
|
**Refusal — the fourth ruler, 30 questions with no answer anywhere**
|
|
280
283
|
|
|
281
|
-
| | this repo |
|
|
284
|
+
| | this repo | Flask |
|
|
282
285
|
|---|---|---|
|
|
283
|
-
| correctly met with silence | **0.967** | **0.
|
|
284
|
-
| English only | **0.933** |
|
|
286
|
+
| correctly met with silence | **0.967** | **0.833** |
|
|
287
|
+
| English only | **0.933** | 0.667 |
|
|
285
288
|
| Chinese only | **1.000** | **1.000** |
|
|
286
289
|
| results resting on no lexical evidence, rulers A–C | **0.000** | **0.000** |
|
|
287
290
|
|
|
291
|
+
Foreign silence fell from 0.933 to 0.833 when the subject changed, and the
|
|
292
|
+
cause is a limit of the design rather than a defect. A word counts as evidence
|
|
293
|
+
unless it occurs in more than 5% of units — a stopword list derived from the
|
|
294
|
+
corpus, so that it needs no list and works in any language. Here `how`, `when`,
|
|
295
|
+
`does` and `are` are everywhere, because 304 units carry written English prose.
|
|
296
|
+
Across 1,572 units of mostly short, undocumented methods they occur in 1–5% of
|
|
297
|
+
them and start counting as evidence. Five English questions about subjects
|
|
298
|
+
Flask does not implement get through on exactly that.
|
|
299
|
+
|
|
288
300
|
**What each bar costs and buys** — one corpus, gate varied alone:
|
|
289
301
|
|
|
290
302
|
| gate | A hit@1/3/MRR | B hit@1/3/MRR | C hit@1/3/MRR | silence own / foreign |
|
|
291
303
|
|---|---|---|---|---|
|
|
292
|
-
| neither (pre-1.0.0) | 0.
|
|
293
|
-
| coverage only (1.0.0) | 0.
|
|
294
|
-
|
|
|
295
|
-
|
|
296
|
-
|
|
297
|
-
|
|
298
|
-
|
|
299
|
-
|
|
300
|
-
|
|
301
|
-
|
|
302
|
-
|
|
304
|
+
| neither (pre-1.0.0) | 0.200/0.286/0.238 | 0.314/0.486/0.391 | 0.486/0.700/0.569 | 0.000 / 0.000 |
|
|
305
|
+
| coverage only (1.0.0) | 0.200/0.286/0.238 | 0.314/0.471/0.383 | 0.486/0.686/0.564 | 0.567 / 0.733 |
|
|
306
|
+
| concentration only | 0.200/0.286/0.238 | 0.314/0.471/0.383 | 0.443/0.614/0.509 | 0.967 / 0.800 |
|
|
307
|
+
| **both (1.1.0)** | **0.200/0.286/0.238** | **0.314/0.471/0.383** | 0.443/0.614/0.509 | **0.967 / 0.833** |
|
|
308
|
+
|
|
309
|
+
Ruler A is **unmoved by either bar**, and B by concentration. The whole cost is
|
|
310
|
+
three questions of seventy at hit@1 on the warmest ruler, and six at hit@3.
|
|
311
|
+
|
|
312
|
+
Through 1.3.0 this section said concentration subsumes coverage. **On a corpus
|
|
313
|
+
this project did not choose, it does not.** Both bars together silence 0.833 of
|
|
314
|
+
the foreign absent questions, against 0.800 for concentration alone and 0.733
|
|
315
|
+
for coverage alone. One question — and the first time in four releases that
|
|
316
|
+
keeping both has been worth a measurable amount rather than worth a different
|
|
317
|
+
diagnosis.
|
|
318
|
+
|
|
319
|
+
Raising the concentration bar buys the remaining silence, and is refused,
|
|
320
|
+
because it is bought out of the answers: at 0.50 the foreign absent ruler is
|
|
321
|
+
silent on all thirty while ruler A falls to 0.086 hit@1 from 0.200, B to 0.214
|
|
322
|
+
from 0.314 and C to 0.329 from 0.443. **0.28 was chosen before this corpus
|
|
323
|
+
existed and survived meeting it**, which is the only kind of evidence a default
|
|
324
|
+
can have.
|
|
325
|
+
|
|
326
|
+
**Latency** — warm corpus, 584 units `c9df00350cbd`, five consecutive
|
|
303
327
|
invocations of `python -m benchmarks.query_latency --repeats 10` (420 samples
|
|
304
328
|
each):
|
|
305
329
|
|
|
306
330
|
| | | across the five |
|
|
307
331
|
|---|---|---|
|
|
308
|
-
| query, median | **0.
|
|
309
|
-
| query, p95 | 1.
|
|
310
|
-
| refusing an unanswerable query | **0.
|
|
311
|
-
| refusal cheaper than answering by | **~
|
|
332
|
+
| query, median | **0.61 ms** | 0.60 – 0.64 |
|
|
333
|
+
| query, p95 | 1.09 ms | 1.06 – 1.14 |
|
|
334
|
+
| refusing an unanswerable query | **0.017 ms** | 0.015 – 0.019 |
|
|
335
|
+
| refusal cheaper than answering by | **~36×** | 33 – 41 |
|
|
312
336
|
|
|
313
337
|
Two significant figures and a spread, because that is the precision the
|
|
314
|
-
measurement has.
|
|
315
|
-
machine
|
|
316
|
-
|
|
317
|
-
|
|
318
|
-
|
|
319
|
-
|
|
320
|
-
were unfalsifiable rather than wrong.
|
|
338
|
+
measurement has. Across twenty invocations over three releases on the same idle
|
|
339
|
+
machine the median has landed anywhere from 0.51 to 1.44 ms and p95 from 0.85
|
|
340
|
+
to 7.34 ms — a band wider than any change the code has ever made to this
|
|
341
|
+
number. Releases before 1.3.0 published `0.83 ms / p95 1.68 ms` to three
|
|
342
|
+
figures from a script that was never committed; both values sit inside that
|
|
343
|
+
band, which is the point: they were unfalsifiable rather than wrong.
|
|
321
344
|
|
|
322
345
|
Refusal is cheap for a structural reason, not a tuned one: an unanswerable
|
|
323
346
|
query touches only the posting lists of its own distinctive words, and never
|
|
@@ -346,8 +369,8 @@ reaches ranking at all.
|
|
|
346
369
|
Directional local measurements, not service levels — but every one of them is
|
|
347
370
|
now a command rather than a memory, which two of them were not before. Each
|
|
348
371
|
prints the corpus fingerprint beside its score; quote both or neither.
|
|
349
|
-
[`benchmarks/README.md`](benchmarks/README.md) lists the
|
|
350
|
-
each is for.
|
|
372
|
+
[`benchmarks/README.md`](benchmarks/README.md) lists the six scripts and what
|
|
373
|
+
each is for, and the corpus one of them grades is now carried here too.
|
|
351
374
|
|
|
352
375
|
## 7 · `rag-your-code search` vs a Grep loop
|
|
353
376
|
|
|
@@ -357,34 +380,42 @@ many hit. That is what this reproduces — same corpus, same questions, same
|
|
|
357
380
|
ruler, scored at **file** granularity so Grep is not penalised for lacking
|
|
358
381
|
declaration spans.
|
|
359
382
|
|
|
360
|
-
**
|
|
361
|
-
|
|
383
|
+
**Which side wins on an undescribed repository depends on the repository.**
|
|
384
|
+
Through 1.3.0 this section said flatly that Grep wins there, because the one
|
|
385
|
+
undescribed repository ever measured was a hook-heavy tool whose questions were
|
|
386
|
+
answerable by matching identifiers. Swapping the subject for a public web
|
|
387
|
+
framework reversed it. The honest claim is narrower than either table alone:
|
|
362
388
|
|
|
363
|
-
|
|
|
389
|
+
| Flask 3.1.3 · 35 questions · 1,572 units `5fd51169eacc` · no descriptions | Grep loop | rag-your-code |
|
|
364
390
|
|---|---|---|
|
|
365
|
-
| right file first |
|
|
366
|
-
| right file in top 3 |
|
|
367
|
-
| lines it hands back, all questions |
|
|
368
|
-
| characters returned, all questions |
|
|
369
|
-
| questions it answers |
|
|
391
|
+
| right file first | 22.9% | **37.1%** |
|
|
392
|
+
| right file in top 3 | 45.7% | **57.1%** |
|
|
393
|
+
| lines it hands back, all questions | 17,641 | — |
|
|
394
|
+
| characters returned, all questions | 1,415,656 | **258,236** |
|
|
395
|
+
| questions it answers | 30 | 30 |
|
|
370
396
|
|
|
371
397
|
**Once the vocabulary exists, it is not close.**
|
|
372
398
|
|
|
373
|
-
| this repository · 70 questions ·
|
|
399
|
+
| this repository · 70 questions · 584 units `c9df00350cbd` · 304 described | Grep loop | rag-your-code |
|
|
374
400
|
|---|---|---|
|
|
375
401
|
| right file first | 22.9% | **58.6%** |
|
|
376
|
-
| right file in top 3 |
|
|
377
|
-
| lines it hands back, all questions | 11,
|
|
378
|
-
| characters returned, all questions | 1,
|
|
402
|
+
| right file in top 3 | 54.3% | **77.1%** |
|
|
403
|
+
| lines it hands back, all questions | 11,959 | — |
|
|
404
|
+
| characters returned, all questions | 1,135,411 | **615,673** |
|
|
379
405
|
| questions it answers | **61** | 60 |
|
|
380
406
|
|
|
381
407
|
Those two tables are the whole argument of section 3.3, measured against a real
|
|
382
408
|
baseline instead of asserted. A cold index retrieves against a sentence the
|
|
383
|
-
parser generated from identifiers the author already chose
|
|
384
|
-
|
|
385
|
-
|
|
386
|
-
|
|
387
|
-
|
|
409
|
+
parser generated from identifiers the author already chose, plus whatever
|
|
410
|
+
docstrings the author wrote — so how it fares against Grep is decided by how
|
|
411
|
+
much prose the repository already contains. Flask has a written docstring on
|
|
412
|
+
most public methods, and the cold index beats Grep there without a single
|
|
413
|
+
description being added. On the previous subject, a tool with terse comments
|
|
414
|
+
and long identifiers, the same cold index lost to Grep by the same margin.
|
|
415
|
+
|
|
416
|
+
What does not depend on the subject is what descriptions buy: on this
|
|
417
|
+
repository first-place accuracy goes to **more than double** Grep's, and the
|
|
418
|
+
payload comes back ranked, spanned, and roughly half the size.
|
|
388
419
|
|
|
389
420
|
**Both tables come from `python -m benchmarks.grep_baseline`**, which is what
|
|
390
421
|
changed in 1.3.0. Until then this section — the strongest claim the project
|
|
@@ -407,12 +438,22 @@ Four qualifications, because the table would otherwise flatter both sides:
|
|
|
407
438
|
`the` gets every file back in no order. It is also why Grep declines nine of
|
|
408
439
|
the seventy questions here — those had no word left that this corpus does not
|
|
409
440
|
use everywhere.
|
|
410
|
-
- **Payload is counted in characters on both sides
|
|
411
|
-
|
|
412
|
-
|
|
413
|
-
|
|
414
|
-
|
|
415
|
-
|
|
441
|
+
- **Payload is counted in characters on both sides**, and 1.4.1 changed what
|
|
442
|
+
fits in it. A generated description ends with the author's docstring so the
|
|
443
|
+
docstring is searchable, and the source printed below it said the same thing
|
|
444
|
+
again — 2,381 of 3,382 characters of prose header on Flask were a verbatim
|
|
445
|
+
repeat of the code beneath. The block no longer prints what the code shows,
|
|
446
|
+
and the same 12,000-character budget now carries **119 declarations instead
|
|
447
|
+
of 92** there, 323 instead of 305 here. Grep
|
|
448
|
+
hands back 18,600 characters per question it answers, unranked and without
|
|
449
|
+
spans, against 10,300 here, ranked and capped by `search.max_chars` — a
|
|
450
|
+
factor of 1.8. On Flask it is a factor of 5.5 — 47,200 characters against
|
|
451
|
+
8,600 — because a framework repeats its own vocabulary across many files and
|
|
452
|
+
Grep has no way to rank what it finds. Both sides decline the same five of
|
|
453
|
+
those 35 — and they are the five Chinese ones, all of them. A Chinese word is
|
|
454
|
+
not a substring of English source and it is not a token in an index built
|
|
455
|
+
from English source, so on a repository written in one language the cold
|
|
456
|
+
cross-language case is not this tool's failure but the corpus's.
|
|
416
457
|
- **Grep wins outright when you know the string.** `grep -rn "COMMON_TERM"` is
|
|
417
458
|
exact, instant and complete, and nothing here replaces it.
|
|
418
459
|
|
|
@@ -477,19 +518,23 @@ The extra is optional by construction: `dependencies = []` is what a default
|
|
|
477
518
|
install gets, the import happens inside the constructor, and a test asserts the
|
|
478
519
|
default provider imports none of it.
|
|
479
520
|
|
|
480
|
-
**Measured
|
|
481
|
-
|
|
482
|
-
|
|
521
|
+
**Measured on the same four rulers, both arms against one corpus.** 1.1.0
|
|
522
|
+
published this comparison and read it as a win. Its largest gain was on the
|
|
523
|
+
foreign ruler, whose two arms turned out to have been taken against two
|
|
524
|
+
different states of a repository being edited while the script ran. Repeated
|
|
525
|
+
against a pinned corpus:
|
|
483
526
|
|
|
484
|
-
| ruler | signed hash (default) | MiniLM, local |
|
|
485
|
-
|
|
486
|
-
| **A** foreign, cold | 0.
|
|
487
|
-
| **B** own, cold | 0.314 / 0.471 / 0.383 |
|
|
488
|
-
| **C** own, described | 0.443 / 0.614 / 0.507 | 0.
|
|
489
|
-
| **D** silence, own / foreign | 0.967 / 0.
|
|
527
|
+
| ruler | corpus | signed hash (default) | MiniLM, local |
|
|
528
|
+
|---|---|---|---|
|
|
529
|
+
| **A** foreign, cold | 1,572 `5fd51169eacc` | **0.200 / 0.286 / 0.238** | 0.171 / 0.257 / 0.214 |
|
|
530
|
+
| **B** own, cold | 581 `8e1e71942c1c` | 0.314 / 0.471 / 0.383 | 0.314 / 0.471 / 0.383 |
|
|
531
|
+
| **C** own, described | 581 `978a1d48a82a` | **0.443 / 0.614 / 0.507** | 0.429 / 0.600 / 0.500 |
|
|
532
|
+
| **D** silence, own / foreign | as above | 0.967 / 0.833 | 0.967 / 0.833 |
|
|
490
533
|
|
|
491
|
-
|
|
492
|
-
|
|
534
|
+
**Worse or identical on every ruler.** It is shipped anyway, as an extra nobody
|
|
535
|
+
has to install, because it does one thing the hash cannot do at all and these
|
|
536
|
+
rulers cannot see: reach a unit that shares no word with the question. The
|
|
537
|
+
pairs the hash scores exactly zero:
|
|
493
538
|
|
|
494
539
|
| pair | signed hash | MiniLM |
|
|
495
540
|
|---|---|---|
|
|
@@ -507,6 +552,10 @@ a threshold on a score and the distributions overlap (0.469 vs 0.418 median),
|
|
|
507
552
|
and a scale-free standout metric took ruler B from 0.329 to 0.186 for two
|
|
508
553
|
thirds of the silence. Applying the lexical bars costs ruler A nothing.
|
|
509
554
|
|
|
555
|
+
If you install it expecting the hit rates above to move, they will not. Install
|
|
556
|
+
it for the cross-language and paraphrase cases in the table above, which is
|
|
557
|
+
where the difference between the two columns actually lives.
|
|
558
|
+
|
|
510
559
|
**A hosted endpoint** is the third option, and the only one that sends your
|
|
511
560
|
source anywhere:
|
|
512
561
|
|