rag-your-code 0.7.0__tar.gz → 1.0.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {rag_your_code-0.7.0/src/rag_your_code.egg-info → rag_your_code-1.0.0}/PKG-INFO +153 -24
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/README.md +152 -23
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/pyproject.toml +1 -1
- {rag_your_code-0.7.0 → rag_your_code-1.0.0/src/rag_your_code.egg-info}/PKG-INFO +153 -24
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/src/rag_your_code.egg-info/SOURCES.txt +4 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/src/ragyourcode/__init__.py +1 -1
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/src/ragyourcode/agentic.py +27 -4
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/src/ragyourcode/cli.py +49 -15
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/src/ragyourcode/config.py +93 -1
- rag_your_code-1.0.0/src/ragyourcode/embeddings.py +194 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/src/ragyourcode/graph.py +10 -3
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/src/ragyourcode/indexer.py +49 -6
- rag_your_code-1.0.0/src/ragyourcode/providers.py +174 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/src/ragyourcode/search.py +213 -5
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/src/ragyourcode/workflow.py +12 -5
- rag_your_code-1.0.0/tests/test_absent_queries.py +114 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/tests/test_config.py +1 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/tests/test_descriptions.py +65 -0
- rag_your_code-1.0.0/tests/test_evidence.py +201 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/tests/test_metadata.py +51 -0
- rag_your_code-1.0.0/tests/test_providers.py +363 -0
- rag_your_code-0.7.0/src/ragyourcode/embeddings.py +0 -82
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/LICENSE +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/setup.cfg +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/src/rag_your_code.egg-info/dependency_links.txt +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/src/rag_your_code.egg-info/entry_points.txt +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/src/rag_your_code.egg-info/requires.txt +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/src/rag_your_code.egg-info/top_level.txt +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/src/ragyourcode/annotate.py +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/src/ragyourcode/descriptions.py +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/src/ragyourcode/document.py +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/src/ragyourcode/models.py +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/src/ragyourcode/parser.py +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/src/ragyourcode/py.typed +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/tests/test_agent_protocol.py +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/tests/test_agentic.py +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/tests/test_doc_comments.py +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/tests/test_document.py +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/tests/test_e2e_cli.py +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/tests/test_golden.py +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/tests/test_graph_incremental.py +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/tests/test_language_fixtures.py +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/tests/test_large_repo.py +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/tests/test_multilanguage.py +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/tests/test_parser_edges.py +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/tests/test_ragyourcode.py +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/tests/test_ranking.py +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/tests/test_repo_queries.py +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/tests/test_resilience.py +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/tests/test_retrieval_correctness.py +0 -0
- {rag_your_code-0.7.0 → rag_your_code-1.0.0}/tests/test_workflow.py +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: rag-your-code
|
|
3
|
-
Version: 0.
|
|
3
|
+
Version: 1.0.0
|
|
4
4
|
Summary: A local, explainable RAG index for codebases and coding agents
|
|
5
5
|
Author: rag-your-code contributors
|
|
6
6
|
License-Expression: MIT
|
|
@@ -226,9 +226,9 @@ and left Chinese retrieval unchanged.
|
|
|
226
226
|
|
|
227
227
|
### Measured on this repository
|
|
228
228
|
|
|
229
|
-
This project describes its own implementation:
|
|
230
|
-
an agent-written bilingual description, committed to the repo, and
|
|
231
|
-
|
|
229
|
+
This project describes its own implementation: all 163 units under `src/` carry
|
|
230
|
+
an agent-written bilingual description, committed to the repo, and 155 of those
|
|
231
|
+
163 declarations also carry the author's own documentation in the source.
|
|
232
232
|
|
|
233
233
|
Seventy natural-language questions about this codebase, in English and
|
|
234
234
|
Chinese, each listing every unit that genuinely answers it
|
|
@@ -236,13 +236,13 @@ Chinese, each listing every unit that genuinely answers it
|
|
|
236
236
|
|
|
237
237
|
| | generated descriptions | agent-written |
|
|
238
238
|
|---|---|---|
|
|
239
|
-
| hit@1 | 0.
|
|
240
|
-
| hit@3 | 0.
|
|
241
|
-
| MRR | 0.
|
|
242
|
-
|
|
|
239
|
+
| hit@1 | 0.314 | **0.500** |
|
|
240
|
+
| hit@3 | 0.457 | **0.729** |
|
|
241
|
+
| MRR | 0.379 | **0.583** |
|
|
242
|
+
| questions declined for want of evidence | 13.0% | **4.3%** |
|
|
243
243
|
|
|
244
|
-
|
|
245
|
-
|
|
244
|
+
Half again the first-place accuracy. Nineteen questions still fail, which is
|
|
245
|
+
what makes the set usable for measuring the next change;
|
|
246
246
|
`tests/test_repo_queries.py` asserts that some question always does, and that
|
|
247
247
|
the written column beats the generated one.
|
|
248
248
|
|
|
@@ -283,6 +283,77 @@ the code it tests, because it repeats that code's vocabulary and adds its own.
|
|
|
283
283
|
Matching stays lexical. It is LLM-authored keyword expansion, and its reach is
|
|
284
284
|
bounded by how many ways of saying the thing the agent thought to write down.
|
|
285
285
|
|
|
286
|
+
### The question a ranking cannot answer
|
|
287
|
+
|
|
288
|
+
Both rulers above ask questions that *have* an answer, so both can only score
|
|
289
|
+
whether it was found. Neither can see the opposite failure. A ranking always
|
|
290
|
+
produces a least-bad unit and hands it back with a score and a rank that read
|
|
291
|
+
exactly like an answer — and it does that whether or not the repository
|
|
292
|
+
contains anything relevant at all.
|
|
293
|
+
|
|
294
|
+
So there is a third ruler: thirty questions about subjects neither repository
|
|
295
|
+
implements, where the only correct reply is nothing
|
|
296
|
+
([`benchmarks/absent_queries.json`](benchmarks/absent_queries.json)). Before
|
|
297
|
+
1.0.0 it scored **zero**. All thirty answered, on both repositories, in both
|
|
298
|
+
languages:
|
|
299
|
+
|
|
300
|
+
| asked of a repository with no such code | answered with | on the evidence of |
|
|
301
|
+
|---|---|---|
|
|
302
|
+
| `where are CUDA kernels dispatched to the device` | a test about word counting | `are` `the` `to` `where` |
|
|
303
|
+
| `准入控制为什么会拒绝没有资源限额的容器组` | the UTF-8 console setup | `拒绝` `控制` `没有` |
|
|
304
|
+
| `how is the OAuth refresh token rotated` | a description-store method | `before` `is` `refresh` `the` |
|
|
305
|
+
|
|
306
|
+
Not a Chinese problem and not a ranking problem — a missing question. Nothing
|
|
307
|
+
in the pipeline asked *is any of this evidence*; it only asked which ranks
|
|
308
|
+
highest. Retrieval now asks both, and returns nothing when the answer to the
|
|
309
|
+
first is no:
|
|
310
|
+
|
|
311
|
+
| | before | now |
|
|
312
|
+
|---|---|---|
|
|
313
|
+
| unanswerable questions correctly met with silence, this repository | 0.000 | **0.733** |
|
|
314
|
+
| the same, on the foreign repository | 0.000 | **0.800** |
|
|
315
|
+
| results resting on no lexical evidence at all, all three rulers | 0.029 – 0.129 | **0.000** |
|
|
316
|
+
| hit@1 / hit@3 / MRR on the foreign ruler | 0.257 / 0.400 / 0.314 | **unchanged** |
|
|
317
|
+
| hit@1 / hit@3 / MRR on this repository | 0.471 / 0.686 / 0.557 | **unchanged** |
|
|
318
|
+
|
|
319
|
+
The bar is the share of a question's **discriminating** words that occur in
|
|
320
|
+
the index — words the repository uses everywhere are dropped from both sides
|
|
321
|
+
of that fraction, which is the part that does the work. Half of `where are
|
|
322
|
+
CUDA kernels dispatched to the device` matches, and it looks like evidence
|
|
323
|
+
until you notice which half. Counting only words that distinguish silenced 18
|
|
324
|
+
of 30 unanswerable English questions that no plain coverage threshold reached
|
|
325
|
+
at all, at identical cost in real answers — 97 of 98 either way.
|
|
326
|
+
|
|
327
|
+
It is a ratio inside the query rather than a threshold on a score, because a
|
|
328
|
+
score threshold is tied to whatever scale the ranking currently produces —
|
|
329
|
+
this project has already had one of those stop meaning anything the moment
|
|
330
|
+
BM25F changed the scale. `search.min_coverage` sets it; `0` restores the old
|
|
331
|
+
behaviour exactly.
|
|
332
|
+
|
|
333
|
+
**What it costs:** one question of the 158 measured. `控制台编码不是 UTF-8
|
|
334
|
+
会怎么样` was reaching `_use_utf8_streams`, and the only words in it this
|
|
335
|
+
repository contains are `utf` and `8`, both of which it uses everywhere. The
|
|
336
|
+
gate says too little of that question is distinctive, which is defensible, and
|
|
337
|
+
it was getting the right answer on a coincidence.
|
|
338
|
+
|
|
339
|
+
**What it does not fix:** an English question whose words genuinely occur here
|
|
340
|
+
in another sense. `how is the OAuth refresh token rotated` matches `refresh`
|
|
341
|
+
because this repository refreshes *indexes*, and no threshold separates those.
|
|
342
|
+
Six of fifteen English absent questions still get answered for that reason.
|
|
343
|
+
That is the case a real embedding model exists for, and it is measurable now
|
|
344
|
+
that the ruler exists.
|
|
345
|
+
|
|
346
|
+
An empty answer says which kind of empty it is, because each is recovered by a
|
|
347
|
+
different move:
|
|
348
|
+
|
|
349
|
+
```json
|
|
350
|
+
{"results": [],
|
|
351
|
+
"diagnosis": {"reason": "only_ubiquitous_terms_matched",
|
|
352
|
+
"matched_terms": [], "ubiquitous_terms": ["the", "to"],
|
|
353
|
+
"coverage": 0.0, "min_coverage": 0.4,
|
|
354
|
+
"hint": "The only words that matched are ones this repository uses throughout ..."}}
|
|
355
|
+
```
|
|
356
|
+
|
|
286
357
|
Descriptions live in `rag-your-code.descriptions.json` at the repository root
|
|
287
358
|
and are meant to be committed, so one person's pass benefits everyone who
|
|
288
359
|
clones. Each is keyed by unit id **and a digest of the unit's source**: when
|
|
@@ -332,23 +403,28 @@ per-release counts are in [CHANGELOG.md](CHANGELOG.md).
|
|
|
332
403
|
|
|
333
404
|
## Configuration
|
|
334
405
|
|
|
335
|
-
|
|
406
|
+
21 settings in `rag-your-code.toml` at the repository root:
|
|
336
407
|
|
|
337
408
|
```bash
|
|
338
409
|
rag-your-code config init # a commented file, all defaults
|
|
339
410
|
rag-your-code config list # effective values and their source
|
|
340
411
|
rag-your-code config set index.ignore '["vendor", "generated"]'
|
|
341
|
-
rag-your-code config set search.
|
|
412
|
+
rag-your-code config set search.min_coverage 0.25
|
|
342
413
|
```
|
|
343
414
|
|
|
344
415
|
| section | settings |
|
|
345
416
|
|---|---|
|
|
346
417
|
| `[index]` | `ignore`, `suffixes`, `max_file_bytes` |
|
|
347
|
-
| `[embedding]` | `dimensions` |
|
|
348
|
-
| `[search]` | `vector_weight`, `limit`, `max_chars` |
|
|
418
|
+
| `[embedding]` | `dimensions`, `provider`, `endpoint`, `model`, `api_key_env`, `batch`, `timeout`, `retries` |
|
|
419
|
+
| `[search]` | `min_coverage`, `vector_weight`, `vector_recall`, `limit`, `max_chars` |
|
|
349
420
|
| `[agent]` | `max_open_bytes`, `max_open_chars` |
|
|
350
421
|
| `[describe]` | `languages`, `batch`, `max_chars` |
|
|
351
422
|
|
|
423
|
+
This table is asserted against the settings table in `config.py`, in both
|
|
424
|
+
directions, by `tests/test_metadata.py` — it had already fallen nine settings
|
|
425
|
+
behind by 1.0.0, and a section listing three quarters of what exists is worse
|
|
426
|
+
than none, because it reads as complete.
|
|
427
|
+
|
|
352
428
|
Resolution is CLI flag > file > built-in default. There is no environment
|
|
353
429
|
layer: an index is an artifact of a repository, not of a shell.
|
|
354
430
|
|
|
@@ -358,9 +434,10 @@ silently dropped is indistinguishable from one that had no effect.
|
|
|
358
434
|
suffix it cannot read is walked, parsed to nothing, and reported as a clean
|
|
359
435
|
index of zero units.
|
|
360
436
|
|
|
361
|
-
The
|
|
362
|
-
*contains
|
|
363
|
-
|
|
437
|
+
The settings under `[index]` and `[embedding]` that decide what an index
|
|
438
|
+
*contains* — including which provider and model computed its vectors — have a
|
|
439
|
+
digest stored in the index, and changing one forces a full rebuild. The rest
|
|
440
|
+
take effect immediately and invalidate nothing.
|
|
364
441
|
|
|
365
442
|
## Agent protocol
|
|
366
443
|
|
|
@@ -407,14 +484,66 @@ it stopped.
|
|
|
407
484
|
Nothing authored lives under `.rag-your-code/` — that directory is what people
|
|
408
485
|
delete to clear the cache.
|
|
409
486
|
|
|
410
|
-
##
|
|
487
|
+
## Bringing your own model
|
|
488
|
+
|
|
489
|
+
Everything above works with no model at all. If you would rather have real
|
|
490
|
+
semantics, point the index at any OpenAI-compatible embeddings endpoint —
|
|
491
|
+
which includes a model server on your own machine:
|
|
492
|
+
|
|
493
|
+
```toml
|
|
494
|
+
# rag-your-code.toml
|
|
495
|
+
[embedding]
|
|
496
|
+
provider = "openai-compatible"
|
|
497
|
+
endpoint = "http://localhost:11434/v1/embeddings" # ollama, LM Studio, vLLM…
|
|
498
|
+
model = "nomic-embed-text"
|
|
499
|
+
dimensions = 768 # must match the model
|
|
500
|
+
```
|
|
501
|
+
|
|
502
|
+
A hosted service is the same three lines with an `https://` endpoint, plus the
|
|
503
|
+
name of the environment variable holding your key:
|
|
504
|
+
|
|
505
|
+
```toml
|
|
506
|
+
api_key_env = "OPENAI_API_KEY" # the NAME of the variable, never the key
|
|
507
|
+
```
|
|
411
508
|
|
|
412
|
-
|
|
413
|
-
|
|
414
|
-
|
|
415
|
-
|
|
416
|
-
|
|
417
|
-
|
|
509
|
+
**The key is never a setting.** `rag-your-code.toml` is meant to be committed
|
|
510
|
+
so everyone who clones can see what shaped the index; a credential is the one
|
|
511
|
+
value with the opposite requirement, so the file only ever names the variable
|
|
512
|
+
it lives in. Sending a key over plain `http://` to anything but your own
|
|
513
|
+
machine is refused rather than warned about.
|
|
514
|
+
|
|
515
|
+
Three things follow from turning this on, and it is worth knowing all three
|
|
516
|
+
before you do:
|
|
517
|
+
|
|
518
|
+
- **Your source leaves the machine**, unless the endpoint is local. That is
|
|
519
|
+
the whole reason the local case is written first here.
|
|
520
|
+
- **Similarity may now find things, not just order them.** With the local
|
|
521
|
+
hash a cosine shortlist is measurably noise, so it is confined to
|
|
522
|
+
re-ranking. A real model earns the right to add candidates the words never
|
|
523
|
+
reached, which is the one gap no amount of ranking closes:
|
|
524
|
+
`search.vector_recall` sets how many. Lexical evidence still dominates — a
|
|
525
|
+
unit found by similarity alone scores at most `search.vector_weight`.
|
|
526
|
+
- **A failure stops the build.** Falling back to the local hash would leave an
|
|
527
|
+
index whose vectors come from two incompatible spaces, and ranking would act
|
|
528
|
+
on the meaningless cosine between them with full confidence.
|
|
529
|
+
|
|
530
|
+
Switching provider, model or width discards the old vectors and rebuilds, so
|
|
531
|
+
an index can never be a mixture. An incremental run over unchanged files makes
|
|
532
|
+
no request at all.
|
|
533
|
+
|
|
534
|
+
**What is not measured:** whether this helps *your* repository, and by how
|
|
535
|
+
much. No number here is from a real model — this project has no key, and a
|
|
536
|
+
figure produced by a stub would be fiction. The instrument ships instead:
|
|
537
|
+
point `benchmarks/repo_queries.py --index` at your own index and grade it. You
|
|
538
|
+
will probably also want a higher `search.vector_weight` than the 0.15 tuned
|
|
539
|
+
for a hash that carries no meaning.
|
|
540
|
+
|
|
541
|
+
## Still not here
|
|
542
|
+
|
|
543
|
+
Tree-sitter parsing, and a SQLite/ANN storage layer for repositories past the
|
|
544
|
+
measured JSON envelope. Note also that `search.vector_recall` scans every
|
|
545
|
+
unit's vector on every query, which is fine at the measured envelope and is
|
|
546
|
+
the thing an ANN index would replace. See [docs/ROADMAP.md](docs/ROADMAP.md).
|
|
418
547
|
|
|
419
548
|
## Development
|
|
420
549
|
|
|
@@ -199,9 +199,9 @@ and left Chinese retrieval unchanged.
|
|
|
199
199
|
|
|
200
200
|
### Measured on this repository
|
|
201
201
|
|
|
202
|
-
This project describes its own implementation:
|
|
203
|
-
an agent-written bilingual description, committed to the repo, and
|
|
204
|
-
|
|
202
|
+
This project describes its own implementation: all 163 units under `src/` carry
|
|
203
|
+
an agent-written bilingual description, committed to the repo, and 155 of those
|
|
204
|
+
163 declarations also carry the author's own documentation in the source.
|
|
205
205
|
|
|
206
206
|
Seventy natural-language questions about this codebase, in English and
|
|
207
207
|
Chinese, each listing every unit that genuinely answers it
|
|
@@ -209,13 +209,13 @@ Chinese, each listing every unit that genuinely answers it
|
|
|
209
209
|
|
|
210
210
|
| | generated descriptions | agent-written |
|
|
211
211
|
|---|---|---|
|
|
212
|
-
| hit@1 | 0.
|
|
213
|
-
| hit@3 | 0.
|
|
214
|
-
| MRR | 0.
|
|
215
|
-
|
|
|
212
|
+
| hit@1 | 0.314 | **0.500** |
|
|
213
|
+
| hit@3 | 0.457 | **0.729** |
|
|
214
|
+
| MRR | 0.379 | **0.583** |
|
|
215
|
+
| questions declined for want of evidence | 13.0% | **4.3%** |
|
|
216
216
|
|
|
217
|
-
|
|
218
|
-
|
|
217
|
+
Half again the first-place accuracy. Nineteen questions still fail, which is
|
|
218
|
+
what makes the set usable for measuring the next change;
|
|
219
219
|
`tests/test_repo_queries.py` asserts that some question always does, and that
|
|
220
220
|
the written column beats the generated one.
|
|
221
221
|
|
|
@@ -256,6 +256,77 @@ the code it tests, because it repeats that code's vocabulary and adds its own.
|
|
|
256
256
|
Matching stays lexical. It is LLM-authored keyword expansion, and its reach is
|
|
257
257
|
bounded by how many ways of saying the thing the agent thought to write down.
|
|
258
258
|
|
|
259
|
+
### The question a ranking cannot answer
|
|
260
|
+
|
|
261
|
+
Both rulers above ask questions that *have* an answer, so both can only score
|
|
262
|
+
whether it was found. Neither can see the opposite failure. A ranking always
|
|
263
|
+
produces a least-bad unit and hands it back with a score and a rank that read
|
|
264
|
+
exactly like an answer — and it does that whether or not the repository
|
|
265
|
+
contains anything relevant at all.
|
|
266
|
+
|
|
267
|
+
So there is a third ruler: thirty questions about subjects neither repository
|
|
268
|
+
implements, where the only correct reply is nothing
|
|
269
|
+
([`benchmarks/absent_queries.json`](benchmarks/absent_queries.json)). Before
|
|
270
|
+
1.0.0 it scored **zero**. All thirty answered, on both repositories, in both
|
|
271
|
+
languages:
|
|
272
|
+
|
|
273
|
+
| asked of a repository with no such code | answered with | on the evidence of |
|
|
274
|
+
|---|---|---|
|
|
275
|
+
| `where are CUDA kernels dispatched to the device` | a test about word counting | `are` `the` `to` `where` |
|
|
276
|
+
| `准入控制为什么会拒绝没有资源限额的容器组` | the UTF-8 console setup | `拒绝` `控制` `没有` |
|
|
277
|
+
| `how is the OAuth refresh token rotated` | a description-store method | `before` `is` `refresh` `the` |
|
|
278
|
+
|
|
279
|
+
Not a Chinese problem and not a ranking problem — a missing question. Nothing
|
|
280
|
+
in the pipeline asked *is any of this evidence*; it only asked which ranks
|
|
281
|
+
highest. Retrieval now asks both, and returns nothing when the answer to the
|
|
282
|
+
first is no:
|
|
283
|
+
|
|
284
|
+
| | before | now |
|
|
285
|
+
|---|---|---|
|
|
286
|
+
| unanswerable questions correctly met with silence, this repository | 0.000 | **0.733** |
|
|
287
|
+
| the same, on the foreign repository | 0.000 | **0.800** |
|
|
288
|
+
| results resting on no lexical evidence at all, all three rulers | 0.029 – 0.129 | **0.000** |
|
|
289
|
+
| hit@1 / hit@3 / MRR on the foreign ruler | 0.257 / 0.400 / 0.314 | **unchanged** |
|
|
290
|
+
| hit@1 / hit@3 / MRR on this repository | 0.471 / 0.686 / 0.557 | **unchanged** |
|
|
291
|
+
|
|
292
|
+
The bar is the share of a question's **discriminating** words that occur in
|
|
293
|
+
the index — words the repository uses everywhere are dropped from both sides
|
|
294
|
+
of that fraction, which is the part that does the work. Half of `where are
|
|
295
|
+
CUDA kernels dispatched to the device` matches, and it looks like evidence
|
|
296
|
+
until you notice which half. Counting only words that distinguish silenced 18
|
|
297
|
+
of 30 unanswerable English questions that no plain coverage threshold reached
|
|
298
|
+
at all, at identical cost in real answers — 97 of 98 either way.
|
|
299
|
+
|
|
300
|
+
It is a ratio inside the query rather than a threshold on a score, because a
|
|
301
|
+
score threshold is tied to whatever scale the ranking currently produces —
|
|
302
|
+
this project has already had one of those stop meaning anything the moment
|
|
303
|
+
BM25F changed the scale. `search.min_coverage` sets it; `0` restores the old
|
|
304
|
+
behaviour exactly.
|
|
305
|
+
|
|
306
|
+
**What it costs:** one question of the 158 measured. `控制台编码不是 UTF-8
|
|
307
|
+
会怎么样` was reaching `_use_utf8_streams`, and the only words in it this
|
|
308
|
+
repository contains are `utf` and `8`, both of which it uses everywhere. The
|
|
309
|
+
gate says too little of that question is distinctive, which is defensible, and
|
|
310
|
+
it was getting the right answer on a coincidence.
|
|
311
|
+
|
|
312
|
+
**What it does not fix:** an English question whose words genuinely occur here
|
|
313
|
+
in another sense. `how is the OAuth refresh token rotated` matches `refresh`
|
|
314
|
+
because this repository refreshes *indexes*, and no threshold separates those.
|
|
315
|
+
Six of fifteen English absent questions still get answered for that reason.
|
|
316
|
+
That is the case a real embedding model exists for, and it is measurable now
|
|
317
|
+
that the ruler exists.
|
|
318
|
+
|
|
319
|
+
An empty answer says which kind of empty it is, because each is recovered by a
|
|
320
|
+
different move:
|
|
321
|
+
|
|
322
|
+
```json
|
|
323
|
+
{"results": [],
|
|
324
|
+
"diagnosis": {"reason": "only_ubiquitous_terms_matched",
|
|
325
|
+
"matched_terms": [], "ubiquitous_terms": ["the", "to"],
|
|
326
|
+
"coverage": 0.0, "min_coverage": 0.4,
|
|
327
|
+
"hint": "The only words that matched are ones this repository uses throughout ..."}}
|
|
328
|
+
```
|
|
329
|
+
|
|
259
330
|
Descriptions live in `rag-your-code.descriptions.json` at the repository root
|
|
260
331
|
and are meant to be committed, so one person's pass benefits everyone who
|
|
261
332
|
clones. Each is keyed by unit id **and a digest of the unit's source**: when
|
|
@@ -305,23 +376,28 @@ per-release counts are in [CHANGELOG.md](CHANGELOG.md).
|
|
|
305
376
|
|
|
306
377
|
## Configuration
|
|
307
378
|
|
|
308
|
-
|
|
379
|
+
21 settings in `rag-your-code.toml` at the repository root:
|
|
309
380
|
|
|
310
381
|
```bash
|
|
311
382
|
rag-your-code config init # a commented file, all defaults
|
|
312
383
|
rag-your-code config list # effective values and their source
|
|
313
384
|
rag-your-code config set index.ignore '["vendor", "generated"]'
|
|
314
|
-
rag-your-code config set search.
|
|
385
|
+
rag-your-code config set search.min_coverage 0.25
|
|
315
386
|
```
|
|
316
387
|
|
|
317
388
|
| section | settings |
|
|
318
389
|
|---|---|
|
|
319
390
|
| `[index]` | `ignore`, `suffixes`, `max_file_bytes` |
|
|
320
|
-
| `[embedding]` | `dimensions` |
|
|
321
|
-
| `[search]` | `vector_weight`, `limit`, `max_chars` |
|
|
391
|
+
| `[embedding]` | `dimensions`, `provider`, `endpoint`, `model`, `api_key_env`, `batch`, `timeout`, `retries` |
|
|
392
|
+
| `[search]` | `min_coverage`, `vector_weight`, `vector_recall`, `limit`, `max_chars` |
|
|
322
393
|
| `[agent]` | `max_open_bytes`, `max_open_chars` |
|
|
323
394
|
| `[describe]` | `languages`, `batch`, `max_chars` |
|
|
324
395
|
|
|
396
|
+
This table is asserted against the settings table in `config.py`, in both
|
|
397
|
+
directions, by `tests/test_metadata.py` — it had already fallen nine settings
|
|
398
|
+
behind by 1.0.0, and a section listing three quarters of what exists is worse
|
|
399
|
+
than none, because it reads as complete.
|
|
400
|
+
|
|
325
401
|
Resolution is CLI flag > file > built-in default. There is no environment
|
|
326
402
|
layer: an index is an artifact of a repository, not of a shell.
|
|
327
403
|
|
|
@@ -331,9 +407,10 @@ silently dropped is indistinguishable from one that had no effect.
|
|
|
331
407
|
suffix it cannot read is walked, parsed to nothing, and reported as a clean
|
|
332
408
|
index of zero units.
|
|
333
409
|
|
|
334
|
-
The
|
|
335
|
-
*contains
|
|
336
|
-
|
|
410
|
+
The settings under `[index]` and `[embedding]` that decide what an index
|
|
411
|
+
*contains* — including which provider and model computed its vectors — have a
|
|
412
|
+
digest stored in the index, and changing one forces a full rebuild. The rest
|
|
413
|
+
take effect immediately and invalidate nothing.
|
|
337
414
|
|
|
338
415
|
## Agent protocol
|
|
339
416
|
|
|
@@ -380,14 +457,66 @@ it stopped.
|
|
|
380
457
|
Nothing authored lives under `.rag-your-code/` — that directory is what people
|
|
381
458
|
delete to clear the cache.
|
|
382
459
|
|
|
383
|
-
##
|
|
460
|
+
## Bringing your own model
|
|
461
|
+
|
|
462
|
+
Everything above works with no model at all. If you would rather have real
|
|
463
|
+
semantics, point the index at any OpenAI-compatible embeddings endpoint —
|
|
464
|
+
which includes a model server on your own machine:
|
|
465
|
+
|
|
466
|
+
```toml
|
|
467
|
+
# rag-your-code.toml
|
|
468
|
+
[embedding]
|
|
469
|
+
provider = "openai-compatible"
|
|
470
|
+
endpoint = "http://localhost:11434/v1/embeddings" # ollama, LM Studio, vLLM…
|
|
471
|
+
model = "nomic-embed-text"
|
|
472
|
+
dimensions = 768 # must match the model
|
|
473
|
+
```
|
|
474
|
+
|
|
475
|
+
A hosted service is the same three lines with an `https://` endpoint, plus the
|
|
476
|
+
name of the environment variable holding your key:
|
|
477
|
+
|
|
478
|
+
```toml
|
|
479
|
+
api_key_env = "OPENAI_API_KEY" # the NAME of the variable, never the key
|
|
480
|
+
```
|
|
384
481
|
|
|
385
|
-
|
|
386
|
-
|
|
387
|
-
|
|
388
|
-
|
|
389
|
-
|
|
390
|
-
|
|
482
|
+
**The key is never a setting.** `rag-your-code.toml` is meant to be committed
|
|
483
|
+
so everyone who clones can see what shaped the index; a credential is the one
|
|
484
|
+
value with the opposite requirement, so the file only ever names the variable
|
|
485
|
+
it lives in. Sending a key over plain `http://` to anything but your own
|
|
486
|
+
machine is refused rather than warned about.
|
|
487
|
+
|
|
488
|
+
Three things follow from turning this on, and it is worth knowing all three
|
|
489
|
+
before you do:
|
|
490
|
+
|
|
491
|
+
- **Your source leaves the machine**, unless the endpoint is local. That is
|
|
492
|
+
the whole reason the local case is written first here.
|
|
493
|
+
- **Similarity may now find things, not just order them.** With the local
|
|
494
|
+
hash a cosine shortlist is measurably noise, so it is confined to
|
|
495
|
+
re-ranking. A real model earns the right to add candidates the words never
|
|
496
|
+
reached, which is the one gap no amount of ranking closes:
|
|
497
|
+
`search.vector_recall` sets how many. Lexical evidence still dominates — a
|
|
498
|
+
unit found by similarity alone scores at most `search.vector_weight`.
|
|
499
|
+
- **A failure stops the build.** Falling back to the local hash would leave an
|
|
500
|
+
index whose vectors come from two incompatible spaces, and ranking would act
|
|
501
|
+
on the meaningless cosine between them with full confidence.
|
|
502
|
+
|
|
503
|
+
Switching provider, model or width discards the old vectors and rebuilds, so
|
|
504
|
+
an index can never be a mixture. An incremental run over unchanged files makes
|
|
505
|
+
no request at all.
|
|
506
|
+
|
|
507
|
+
**What is not measured:** whether this helps *your* repository, and by how
|
|
508
|
+
much. No number here is from a real model — this project has no key, and a
|
|
509
|
+
figure produced by a stub would be fiction. The instrument ships instead:
|
|
510
|
+
point `benchmarks/repo_queries.py --index` at your own index and grade it. You
|
|
511
|
+
will probably also want a higher `search.vector_weight` than the 0.15 tuned
|
|
512
|
+
for a hash that carries no meaning.
|
|
513
|
+
|
|
514
|
+
## Still not here
|
|
515
|
+
|
|
516
|
+
Tree-sitter parsing, and a SQLite/ANN storage layer for repositories past the
|
|
517
|
+
measured JSON envelope. Note also that `search.vector_recall` scans every
|
|
518
|
+
unit's vector on every query, which is fine at the measured envelope and is
|
|
519
|
+
the thing an ANN index would replace. See [docs/ROADMAP.md](docs/ROADMAP.md).
|
|
391
520
|
|
|
392
521
|
## Development
|
|
393
522
|
|