tunarag-python 0.2.1__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,1164 @@
1
+ Metadata-Version: 2.5
2
+ Name: tunarag-python
3
+ Version: 0.2.1
4
+ Summary: Typed, durable optimization workflows for existing RAG systems
5
+ Project-URL: Repository, https://github.com/shivamshinde123/tunarag-python
6
+ Project-URL: Issues, https://github.com/shivamshinde123/tunarag-python/issues
7
+ Project-URL: Changelog, https://github.com/shivamshinde123/tunarag-python/blob/main/CHANGELOG.md
8
+ Author: Shivam Shinde
9
+ License-Expression: MIT
10
+ Keywords: evaluation,llm,optimization,rag,retrieval
11
+ Classifier: Development Status :: 3 - Alpha
12
+ Classifier: Intended Audience :: Developers
13
+ Classifier: License :: OSI Approved :: MIT License
14
+ Classifier: Operating System :: OS Independent
15
+ Classifier: Programming Language :: Python :: 3
16
+ Classifier: Programming Language :: Python :: 3.10
17
+ Classifier: Programming Language :: Python :: 3.11
18
+ Classifier: Programming Language :: Python :: 3.12
19
+ Classifier: Programming Language :: Python :: 3.13
20
+ Classifier: Typing :: Typed
21
+ Requires-Python: >=3.10
22
+ Requires-Dist: tomli>=2.0; python_version < '3.11'
23
+ Provides-Extra: dev
24
+ Requires-Dist: build>=1.2; extra == 'dev'
25
+ Requires-Dist: hypothesis>=6.100; extra == 'dev'
26
+ Requires-Dist: mypy>=1.11; extra == 'dev'
27
+ Requires-Dist: pytest-asyncio>=0.23; extra == 'dev'
28
+ Requires-Dist: pytest>=8.0; extra == 'dev'
29
+ Requires-Dist: ruff>=0.6; extra == 'dev'
30
+ Provides-Extra: frameworks
31
+ Requires-Dist: langchain<2,>=1; extra == 'frameworks'
32
+ Requires-Dist: langgraph<2,>=1; extra == 'frameworks'
33
+ Provides-Extra: langchain
34
+ Requires-Dist: langchain<2,>=1; extra == 'langchain'
35
+ Provides-Extra: langgraph
36
+ Requires-Dist: langgraph<2,>=1; extra == 'langgraph'
37
+ Provides-Extra: mlflow
38
+ Requires-Dist: mlflow<4,>=3; extra == 'mlflow'
39
+ Provides-Extra: ragas
40
+ Requires-Dist: ragas<0.5,>=0.4; extra == 'ragas'
41
+ Description-Content-Type: text/markdown
42
+
43
+ # TunaRAG Python
44
+
45
+ **Find the best configuration for your existing RAG application.**
46
+
47
+ TunaRAG is an open-source, Python-first optimization toolkit for retrieval-augmented generation
48
+ (RAG) systems. It connects to a RAG pipeline you already own, evaluates different configurations on
49
+ your data, and identifies the best tradeoff between **answer quality, cost, and latency**.
50
+
51
+ Instead of manually changing retrieval settings, prompts, or models and comparing a few examples by
52
+ eye, TunaRAG runs structured, repeatable optimization studies. It records every trial in SQLite,
53
+ isolates failures, supports resumable execution, and produces results you can inspect and export.
54
+
55
+ ### What can you use it for?
56
+
57
+ - Find the best `top_k`, chunking, prompt, model, reranking, or generation settings for your RAG.
58
+ - Compare configurations with custom metrics or optional RAGAS evaluators.
59
+ - Optimize a weighted objective across quality, provider cost, and response latency.
60
+ - Search with deterministic random search or provider-neutral LLM-guided search.
61
+ - Tune existing custom, LangChain, or LangGraph RAG pipelines without rebuilding them.
62
+ - Preserve reproducible experiments with caching, retries, callbacks, reports, and SQLite resume.
63
+
64
+ ### Who is it for?
65
+
66
+ TunaRAG is designed for developers and AI teams that already have a working RAG pipeline and want a
67
+ reliable way to evaluate and improve it. Your application continues to own retrieval, generation,
68
+ indexes, provider clients, and credentials; TunaRAG owns the optimization workflow around it.
69
+
70
+ > Install the distribution as `tunarag-python` and import it as `tunarag`. Python 3.10-3.13 is
71
+ > supported.
72
+
73
+ ## Comprehensive Usage Guide
74
+
75
+ The sections below explain every public TunaRAG Python V0.2 feature with executable examples.
76
+
77
+ ## Contents
78
+
79
+ 1. [Purpose and mental model](#purpose-and-mental-model)
80
+ 2. [Installation](#installation)
81
+ 3. [End-to-end quickstart](#end-to-end-quickstart)
82
+ 4. [RAG adapters](#rag-adapters)
83
+ 5. [LangChain integration](#langchain-integration)
84
+ 6. [LangGraph integration](#langgraph-integration)
85
+ 7. [Evaluation datasets](#evaluation-datasets)
86
+ 8. [Synthetic QA generation](#synthetic-qa-generation)
87
+ 9. [Search spaces and strategies](#search-spaces-and-strategies)
88
+ 10. [Evaluators and RAGAS](#evaluators-and-ragas)
89
+ 11. [Objectives](#objectives)
90
+ 12. [Configuration](#configuration)
91
+ 13. [Running optimization](#running-optimization)
92
+ 14. [Stopping, retries, and concurrency](#stopping-retries-and-concurrency)
93
+ 15. [Caching](#caching)
94
+ 16. [SQLite persistence and resume](#sqlite-persistence-and-resume)
95
+ 17. [Callbacks and MLflow](#callbacks-and-mlflow)
96
+ 18. [Results and reports](#results-and-reports)
97
+ 19. [Errors, privacy, and secrets](#errors-privacy-and-secrets)
98
+ 20. [Extension contracts](#extension-contracts)
99
+ 21. [Production checklist](#production-checklist)
100
+ 22. [Limitations and glossary](#limitations-and-glossary)
101
+
102
+ ## Purpose and mental model
103
+
104
+ TunaRAG compares configurations of an **existing** RAG pipeline. You provide the pipeline, an
105
+ evaluation dataset, a search space, evaluators, and an objective. TunaRAG schedules trials, isolates
106
+ failures, persists state, and returns an inspectable result.
107
+
108
+ ```mermaid
109
+ flowchart LR
110
+ D[Versioned dataset] --> O[Optimizer]
111
+ S[Search strategy] -->|Candidate| O
112
+ O -->|Candidate + example| A[RAG adapter]
113
+ A -->|Output + usage| E[Evaluators]
114
+ E -->|Metrics| J[Objective]
115
+ J -->|Benefit score| O
116
+ O -->|Observe| S
117
+ O --> DB[(SQLite store)]
118
+ O <--> C[(Evaluation cache)]
119
+ DB --> CB[Post-commit callbacks]
120
+ O --> R[OptimizationResult]
121
+ ```
122
+
123
+ A **study** contains ordered **trials**. Each trial evaluates one complete candidate configuration
124
+ against every dataset example. Same-named sample metrics are aggregated into trial metrics, then the
125
+ objective converts those metrics into one larger-is-better score.
126
+
127
+ TunaRAG does not create or host a RAG application, provision indexes, manage provider credentials,
128
+ or expose a hosted API. Your application keeps ownership of those concerns.
129
+
130
+ ## Installation
131
+
132
+ ```bash
133
+ python -m pip install tunarag-python
134
+ ```
135
+
136
+ | Optional feature | Installation |
137
+ |---|---|
138
+ | LangChain | `python -m pip install "tunarag-python[langchain]"` |
139
+ | LangGraph | `python -m pip install "tunarag-python[langgraph]"` |
140
+ | Both frameworks | `python -m pip install "tunarag-python[frameworks]"` |
141
+ | RAGAS | `python -m pip install "tunarag-python[ragas]"` |
142
+ | MLflow | `python -m pip install "tunarag-python[mlflow]"` |
143
+
144
+ Optional dependencies load lazily; the core package does not require them.
145
+
146
+ ## End-to-end quickstart
147
+
148
+ This deterministic example shows the complete public workflow without external providers.
149
+
150
+ ```python
151
+ from tunarag import (
152
+ EvaluationDataset,
153
+ IntegerRange,
154
+ MetricEvaluator,
155
+ Objective,
156
+ ObjectiveTerm,
157
+ Optimizer,
158
+ RandomSearch,
159
+ SearchSpace,
160
+ UsageRecord,
161
+ )
162
+
163
+ DOCUMENTS = {
164
+ "python": "Python is a high-level programming language.",
165
+ "sqlite": "SQLite is a local embedded relational database.",
166
+ "rag": "RAG combines retrieval with generation.",
167
+ }
168
+
169
+
170
+ class SmallRAG:
171
+ async def run(self, candidate, example):
172
+ query_words = set(example.query.lower().split())
173
+ ranked = sorted(
174
+ DOCUMENTS.values(),
175
+ key=lambda text: len(query_words & set(text.lower().split())),
176
+ reverse=True,
177
+ )
178
+ contexts = ranked[: candidate.parameters["top_k"]]
179
+ return {"answer": " ".join(contexts)}, (UsageRecord("small-rag", latency_seconds=0.001),)
180
+
181
+
182
+ def answer_recall(output, example):
183
+ expected = set(example.reference_answer.lower().split())
184
+ actual = set(output["answer"].lower().split())
185
+ return len(expected & actual) / len(expected)
186
+
187
+
188
+ dataset = EvaluationDataset.from_records(
189
+ [
190
+ {
191
+ "id": "q1",
192
+ "query": "What is SQLite?",
193
+ "reference_answer": "a local embedded relational database",
194
+ },
195
+ {
196
+ "id": "q2",
197
+ "query": "What does RAG combine?",
198
+ "reference_answer": "retrieval with generation",
199
+ },
200
+ ]
201
+ )
202
+ optimizer = Optimizer(
203
+ adapter=SmallRAG(),
204
+ strategy=RandomSearch(SearchSpace({"top_k": IntegerRange(1, 3)}), seed=7),
205
+ evaluators=(MetricEvaluator("answer_recall", answer_recall),),
206
+ objective=Objective((ObjectiveTerm("answer_recall"),)),
207
+ )
208
+ try:
209
+ result = optimizer.optimize(dataset, study_name="small-rag")
210
+ print(result.best_candidate.parameters if result.best_candidate else "no winner")
211
+ finally:
212
+ optimizer.close()
213
+ ```
214
+
215
+ The default store is in memory. Use `SQLiteStore(path)` for a durable study.
216
+
217
+ ## RAG adapters
218
+
219
+ A custom adapter implements:
220
+
221
+ ```python
222
+ from collections.abc import Sequence
223
+
224
+ from tunarag import UsageRecord
225
+
226
+
227
+ async def run(candidate, example) -> tuple[object, Sequence[UsageRecord]]: ...
228
+ ```
229
+
230
+ `candidate.parameters` contains the proposed configuration. The output can be any object understood
231
+ by your evaluators. Usage may be empty. Adapter exceptions enter TunaRAG's retry and trial-failure
232
+ boundary.
233
+
234
+ ```python
235
+ from tunarag import UsageRecord, ValueStatus
236
+
237
+
238
+ class ExistingRAGAdapter:
239
+ def __init__(self, service):
240
+ self.service = service
241
+
242
+ async def run(self, candidate, example):
243
+ response = await self.service.answer(
244
+ query=example.query,
245
+ top_k=candidate.parameters["top_k"],
246
+ model=candidate.parameters["model"],
247
+ )
248
+ return response, (
249
+ UsageRecord(
250
+ component="generator",
251
+ input_tokens=response.input_tokens,
252
+ output_tokens=response.output_tokens,
253
+ cost=response.cost,
254
+ latency_seconds=response.latency_seconds,
255
+ status=ValueStatus.EXACT,
256
+ pricing_version="provider-2026-10",
257
+ ),
258
+ )
259
+
260
+ def fingerprint(self):
261
+ return {
262
+ "adapter_version": 2,
263
+ "index_revision": "docs-2026-10-04",
264
+ "prompt_revision": "qa-v3",
265
+ }
266
+ ```
267
+
268
+ Use `ValueStatus.ESTIMATED` for inferred usage and `UNAVAILABLE` when it cannot be determined.
269
+ Implement `fingerprint()` when class identity alone does not represent semantics. Fingerprints affect
270
+ cache and resume compatibility; include stable versions, never credentials.
271
+
272
+ ## LangChain integration
273
+
274
+ `LangChainAdapter` accepts any asynchronous runnable with `ainvoke(input, config=...)`.
275
+
276
+ ```python
277
+ from langchain_core.runnables import RunnableLambda
278
+ from tunarag import LangChainAdapter
279
+
280
+
281
+ async def answer(inputs, config):
282
+ top_k = config["configurable"]["top_k"]
283
+ return {"answer": f"Used top_k={top_k} for {inputs['question']}"}
284
+
285
+
286
+ chain = RunnableLambda(answer)
287
+ adapter = LangChainAdapter(
288
+ chain,
289
+ output_mapper=lambda output: output["answer"],
290
+ identity={"chain": "support-rag", "version": 1},
291
+ )
292
+ ```
293
+
294
+ For `EvaluationExample`, default input is `{"question": example.query}`. Candidate parameters are
295
+ placed in `config["configurable"]`. Other example types pass through unchanged.
296
+
297
+ Use a factory when candidates change construction-time components:
298
+
299
+ ```python
300
+ def build_chain(candidate):
301
+ retriever = vector_store.as_retriever(search_kwargs={"k": candidate.parameters["top_k"]})
302
+ return create_retrieval_chain(
303
+ retriever,
304
+ build_qa_chain(candidate.parameters["model"]),
305
+ )
306
+
307
+
308
+ adapter = LangChainAdapter(
309
+ factory=build_chain,
310
+ output_mapper=lambda output: output["answer"],
311
+ identity={"factory_schema": 1},
312
+ )
313
+ ```
314
+
315
+ Customize all boundaries when required:
316
+
317
+ ```python
318
+ adapter = LangChainAdapter(
319
+ chain,
320
+ input_mapper=lambda example, candidate: {
321
+ "input": example.query,
322
+ "retrieval_limit": candidate.parameters["top_k"],
323
+ },
324
+ config_mapper=lambda candidate: {
325
+ "configurable": {"model": candidate.parameters["model"]},
326
+ "tags": ["tunarag-trial"],
327
+ },
328
+ output_mapper=lambda output: output["answer"],
329
+ usage_mapper=lambda output: (
330
+ UsageRecord(
331
+ "langchain-provider",
332
+ input_tokens=output["usage"]["input_tokens"],
333
+ output_tokens=output["usage"]["output_tokens"],
334
+ ),
335
+ ),
336
+ )
337
+ ```
338
+
339
+ The default usage mapper recursively discovers standard `usage_metadata` in mappings, sequences,
340
+ and message objects. The adapter always adds `langchain.runnable` latency. Set
341
+ `capture_message_usage=False` to disable automatic token extraction. A custom `usage_mapper` takes
342
+ precedence.
343
+
344
+ ## LangGraph integration
345
+
346
+ `LangGraphAdapter` has the same factory and mapping options.
347
+
348
+ ```python
349
+ from typing import TypedDict
350
+
351
+ from langgraph.graph import END, START, StateGraph
352
+ from tunarag import LangGraphAdapter
353
+
354
+
355
+ class State(TypedDict, total=False):
356
+ question: str
357
+ answer: str
358
+ top_k: int
359
+
360
+
361
+ async def answer_node(state: State, config):
362
+ top_k = state["top_k"] if "top_k" in state else config["configurable"]["top_k"]
363
+ return {"answer": f"answer for {state['question']} with {top_k} documents"}
364
+
365
+
366
+ builder = StateGraph(State)
367
+ builder.add_node("answer", answer_node)
368
+ builder.add_edge(START, "answer")
369
+ builder.add_edge("answer", END)
370
+ graph = builder.compile()
371
+ adapter = LangGraphAdapter(
372
+ graph,
373
+ output_mapper=lambda state: state["answer"],
374
+ identity={"graph": "qa", "revision": 1},
375
+ )
376
+ ```
377
+
378
+ To place candidate values in graph state instead of runnable config:
379
+
380
+ ```python
381
+ adapter = LangGraphAdapter(
382
+ graph,
383
+ input_mapper=lambda example, candidate: {
384
+ "question": example.query,
385
+ "top_k": candidate.parameters["top_k"],
386
+ },
387
+ config_mapper=lambda candidate: None,
388
+ output_mapper=lambda state: state["answer"],
389
+ )
390
+ ```
391
+
392
+ The adapter always adds `langgraph.runnable` latency. Framework errors remain visible to the engine's
393
+ retry and failure handling.
394
+
395
+ ## Evaluation datasets
396
+
397
+ ### Canonical model
398
+
399
+ ```python
400
+ from tunarag import DatasetSplit, EvaluationExample
401
+
402
+ example = EvaluationExample(
403
+ id="billing-001",
404
+ query="How can I update my billing address?",
405
+ reference_answer="Open Settings, then Billing.",
406
+ reference_contexts=("Billing addresses are under Settings > Billing.",),
407
+ relevant_document_ids=("billing-guide",),
408
+ tags=("billing", "how-to"),
409
+ metadata={"tenant": "demo", "source": "support-review"},
410
+ split=DatasetSplit.OPTIMIZE,
411
+ synthetic=False,
412
+ )
413
+ ```
414
+
415
+ Text is Unicode-normalized and line endings are canonicalized. IDs must be unique. Metadata must be
416
+ finite JSON-compatible data. Duplicate tags and relevant document IDs are rejected.
417
+
418
+ ### Ingestion
419
+
420
+ ```python
421
+ from tunarag import DatasetFieldMap, EvaluationDataset
422
+
423
+ dataset = EvaluationDataset.from_records(
424
+ [
425
+ {
426
+ "question_id": "q-1",
427
+ "question": "What is TunaRAG?",
428
+ "expected": "A RAG optimization library.",
429
+ "contexts": ["TunaRAG optimizes existing RAG systems."],
430
+ "metadata": {"tenant": "demo"},
431
+ }
432
+ ],
433
+ field_map=DatasetFieldMap(
434
+ id="question_id",
435
+ query="question",
436
+ reference_answer="expected",
437
+ reference_contexts="contexts",
438
+ ),
439
+ )
440
+ json_dataset = EvaluationDataset.from_json("evaluation.json")
441
+ jsonl_dataset = EvaluationDataset.from_jsonl("evaluation.jsonl")
442
+ csv_dataset = EvaluationDataset.from_csv("evaluation.csv")
443
+ ```
444
+
445
+ JSON contains an array; JSONL contains one object per nonblank line. CSV collection and metadata
446
+ cells are JSON-encoded. Missing IDs are deterministically derived. Every `EvaluationDataset` has a
447
+ content-derived `version` used by caches and resume validation.
448
+
449
+ ### Deterministic splitting
450
+
451
+ ```python
452
+ from tunarag import DatasetSplit, SplitRatios
453
+
454
+ split_dataset = dataset.with_splits(
455
+ seed=42,
456
+ ratios=SplitRatios(optimize=0.7, validation=0.2, test=0.1),
457
+ group_metadata_key="tenant",
458
+ )
459
+ optimization_data = split_dataset.select(DatasetSplit.OPTIMIZE)
460
+ ```
461
+
462
+ Splits are hash-based and stable across input order. Group splitting keeps related examples
463
+ together. `select()` rejects an empty split.
464
+
465
+ For custom examples, use a user-managed version:
466
+
467
+ ```python
468
+ from tunarag import InMemoryDataset
469
+
470
+ dataset = InMemoryDataset.from_iterable(
471
+ [{"question": "Q1"}, {"question": "Q2"}],
472
+ version="support-set-v4",
473
+ )
474
+ ```
475
+
476
+ Change the version whenever custom example semantics change.
477
+
478
+ ## Synthetic QA generation
479
+
480
+ Generation is provider-neutral. A provider must return exactly the requested count.
481
+
482
+ ```python
483
+ from tunarag import (
484
+ SourceDocument,
485
+ SourceSpan,
486
+ SyntheticDatasetBuilder,
487
+ SyntheticGenerationResponse,
488
+ SyntheticQA,
489
+ UsageRecord,
490
+ )
491
+
492
+
493
+ class QAProvider:
494
+ async def generate(self, request):
495
+ text = request.document.text
496
+ items = tuple(
497
+ SyntheticQA(
498
+ query=f"What does {request.document.id} explain?",
499
+ answer=text,
500
+ source_spans=(SourceSpan(0, len(text)),),
501
+ confidence=0.9,
502
+ )
503
+ for _ in range(request.count)
504
+ )
505
+ return SyntheticGenerationResponse(
506
+ items=items,
507
+ usage=(UsageRecord("qa-generator", output_tokens=40),),
508
+ )
509
+
510
+
511
+ builder = SyntheticDatasetBuilder(
512
+ QAProvider(),
513
+ samples_per_document=2,
514
+ concurrency=4,
515
+ seed=7,
516
+ generator_version="qa-provider-v1",
517
+ )
518
+ generated = await builder.generate(
519
+ [SourceDocument("guide", "TunaRAG optimizes existing RAG pipelines.")]
520
+ )
521
+ dataset = generated.dataset
522
+ generation_usage = generated.usage
523
+ ```
524
+
525
+ `SourceSpan` is a zero-based half-open range. TunaRAG validates spans, confidence, cardinality,
526
+ source IDs, and usage. Generated examples carry provenance and `synthetic=True`. Review them before
527
+ using them as gold data. Generation usage is separate from optimization usage.
528
+
529
+ ## Search spaces and strategies
530
+
531
+ ### Typed space
532
+
533
+ ```python
534
+ from tunarag import Categorical, FloatRange, IntegerRange, SearchSpace
535
+
536
+ space = SearchSpace(
537
+ {
538
+ "top_k": IntegerRange(1, 10), # inclusive
539
+ "temperature": FloatRange(0.0, 0.8),
540
+ "model": Categorical(("small", "large")),
541
+ "prompt": Categorical(("concise", "grounded")),
542
+ }
543
+ )
544
+ ```
545
+
546
+ `SearchSpace.validate()` requires exactly the declared keys. Prefer JSON-like immutable categorical
547
+ values for portable persistence. Secrets should not be search parameters.
548
+
549
+ ### RandomSearch
550
+
551
+ ```python
552
+ from tunarag import RandomSearch
553
+
554
+ strategy = RandomSearch(space, seed=42)
555
+ ```
556
+
557
+ Candidate order is reproducible for the same seed, space, and reservation history. Resume replays
558
+ and verifies the random generator state.
559
+
560
+ ### LLMSearch
561
+
562
+ ```python
563
+ from tunarag import LLMSearch
564
+
565
+
566
+ class SuggestionProvider:
567
+ def __init__(self, client):
568
+ self.client = client
569
+
570
+ async def suggest(self, request):
571
+ return await self.client.generate_json(
572
+ schema=request.search_space,
573
+ context={
574
+ "suggestion_number": request.suggestion_number,
575
+ "attempt_number": request.attempt_number,
576
+ "observations": [
577
+ {
578
+ "candidate": dict(item.candidate.parameters),
579
+ "metrics": [
580
+ {"name": metric.name, "value": metric.value} for metric in item.metrics
581
+ ],
582
+ }
583
+ for item in request.observations
584
+ ],
585
+ },
586
+ )
587
+
588
+
589
+ strategy = LLMSearch(
590
+ space,
591
+ SuggestionProvider(llm_client),
592
+ max_suggestion_attempts=3,
593
+ fallback_seed=42,
594
+ fallback=True,
595
+ )
596
+ ```
597
+
598
+ Proposals are strictly validated. Provider exceptions and invalid values consume a bounded attempt.
599
+ After exhaustion, deterministic random fallback runs when enabled; otherwise `SearchSpaceError` is
600
+ raised. Provider-facing observations are redacted.
601
+
602
+ ## Evaluators and RAGAS
603
+
604
+ ### Simple and custom evaluators
605
+
606
+ ```python
607
+ from tunarag import MetricEvaluator, MetricValue, ValueStatus
608
+
609
+
610
+ def exact_match(output, example):
611
+ if example.reference_answer is None:
612
+ return None
613
+ return float(output["answer"].strip() == example.reference_answer.strip())
614
+
615
+
616
+ exact_match_evaluator = MetricEvaluator("exact_match", exact_match)
617
+
618
+
619
+ class RetrievalEvaluator:
620
+ async def evaluate(self, output, example):
621
+ retrieved = set(output["document_ids"])
622
+ relevant = set(example.relevant_document_ids)
623
+ recall = len(retrieved & relevant) / len(relevant) if relevant else None
624
+ return (
625
+ MetricValue(
626
+ "retrieval_recall",
627
+ recall,
628
+ ValueStatus.EXACT,
629
+ coverage=1.0 if relevant else 0.0,
630
+ ),
631
+ MetricValue("context_count", float(len(retrieved))),
632
+ )
633
+
634
+ def fingerprint(self):
635
+ return {"evaluator": "retrieval", "version": 1}
636
+ ```
637
+
638
+ `MetricEvaluator` accepts sync or async scorers. `None`, NaN, and infinity become unavailable;
639
+ nonnumeric values raise `EvaluationError`. Coverage must be `[0, 1]`. `tunarag.objective` is a
640
+ reserved metric name.
641
+
642
+ ### RAGAS
643
+
644
+ ```python
645
+ from ragas.metrics import AnswerRelevancy, Faithfulness
646
+ from tunarag import RagasEvaluator
647
+
648
+ ragas_evaluator = RagasEvaluator(
649
+ metrics=(Faithfulness(), AnswerRelevancy()),
650
+ sample_mapper=lambda output, example: {
651
+ "user_input": example.query,
652
+ "response": output["answer"],
653
+ "retrieved_contexts": output["contexts"],
654
+ "reference": example.reference_answer,
655
+ },
656
+ metric_prefix="ragas.",
657
+ llm=evaluation_llm,
658
+ embeddings=evaluation_embeddings,
659
+ raise_exceptions=False,
660
+ )
661
+ ```
662
+
663
+ TunaRAG constructs one RAGAS `SingleTurnSample` and calls the async evaluation API. Names receive
664
+ the configured prefix. Mapping and RAGAS failures become `EvaluationError`. Credentials and provider
665
+ configuration remain owned by your RAGAS components.
666
+
667
+ ## Objectives
668
+
669
+ ```python
670
+ from tunarag import Objective, ObjectiveDirection, ObjectiveTerm
671
+
672
+ objective = Objective(
673
+ (
674
+ ObjectiveTerm("answer_quality", weight=0.70, minimum=0.0, maximum=1.0),
675
+ ObjectiveTerm(
676
+ "cost_usd",
677
+ weight=0.20,
678
+ direction=ObjectiveDirection.MINIMIZE,
679
+ minimum=0.0,
680
+ maximum=0.05,
681
+ ),
682
+ ObjectiveTerm(
683
+ "latency_seconds",
684
+ weight=0.10,
685
+ direction=ObjectiveDirection.MINIMIZE,
686
+ minimum=0.0,
687
+ maximum=5.0,
688
+ ),
689
+ )
690
+ )
691
+ ```
692
+
693
+ Bounded metrics are normalized and clamped to `[0, 1]`; bounded minimize terms become
694
+ `1 - normalized`. Unbounded maximize terms use raw values and unbounded minimize terms use negative
695
+ values. Weights are contributions and are not automatically normalized. A missing or unavailable
696
+ required metric makes the objective unavailable.
697
+
698
+ Usage does not automatically become an objective metric. Emit `cost_usd` or `latency_seconds` from
699
+ an evaluator when you want those values in the objective.
700
+
701
+ ## Configuration
702
+
703
+ ```python
704
+ from tunarag import Settings
705
+
706
+ settings = Settings.resolve(
707
+ config_path="tunarag.toml",
708
+ overrides={"max_trials": 50, "concurrency": 4},
709
+ )
710
+ ```
711
+
712
+ Precedence is defaults < TOML `[tunarag]` < environment < overrides. `.env` files are not loaded.
713
+
714
+ ```toml
715
+ [tunarag]
716
+ home = ".tunarag"
717
+ db_path = ".tunarag/experiments.db"
718
+ cache_dir = ".tunarag/cache"
719
+ artifact_dir = ".tunarag/artifacts"
720
+ max_trials = 25
721
+ concurrency = 4
722
+ sample_concurrency = 8
723
+ seed = 0
724
+ max_retries = 2
725
+ sample_timeout_seconds = 60.0
726
+ trial_timeout_seconds = 300.0
727
+ cache_policy = "read_write"
728
+ ```
729
+
730
+ | Environment variable | Default / meaning |
731
+ |---|---|
732
+ | `TUNARAG_HOME` | Platform application-data directory |
733
+ | `TUNARAG_DB_PATH` | `<home>/experiments.db` |
734
+ | `TUNARAG_CACHE_DIR` | `<home>/cache` |
735
+ | `TUNARAG_ARTIFACT_DIR` | `<home>/artifacts` |
736
+ | `TUNARAG_MAX_TRIALS` | `25` |
737
+ | `TUNARAG_CONCURRENCY` | `4` trials |
738
+ | `TUNARAG_SAMPLE_CONCURRENCY` | `8` samples per trial |
739
+ | `TUNARAG_SEED` | `0` |
740
+ | `TUNARAG_MAX_RETRIES` | `2` |
741
+ | `TUNARAG_SAMPLE_TIMEOUT_SECONDS` | `60.0` |
742
+ | `TUNARAG_TRIAL_TIMEOUT_SECONDS` | `300.0` |
743
+ | `TUNARAG_CACHE_POLICY` | `read_write` |
744
+
745
+ Cache policies: `off`, `read_only`, `write_only`, `read_write`. Settings contain no provider keys.
746
+
747
+ ## Running optimization
748
+
749
+ Use async in async programs:
750
+
751
+ ```python
752
+ result = await optimizer.aoptimize(dataset, study_name="support-rag-v4")
753
+ ```
754
+
755
+ Use the sync wrapper only when no event loop is active:
756
+
757
+ ```python
758
+ result = optimizer.optimize(dataset, study_name="support-rag-v4")
759
+ ```
760
+
761
+ Calling sync methods inside an event loop raises `SyncInAsyncContextError`.
762
+
763
+ Durable composition:
764
+
765
+ ```python
766
+ from pathlib import Path
767
+ from tunarag import (
768
+ MaxFailures,
769
+ MaxTrials,
770
+ NoImprovement,
771
+ Optimizer,
772
+ RetryPolicy,
773
+ SQLiteCache,
774
+ SQLiteStore,
775
+ Settings,
776
+ TargetScore,
777
+ )
778
+
779
+ settings = Settings.resolve(
780
+ overrides={
781
+ "home": Path(".tunarag"),
782
+ "db_path": Path(".tunarag/experiments.db"),
783
+ "cache_dir": Path(".tunarag/cache"),
784
+ "artifact_dir": Path(".tunarag/artifacts"),
785
+ "concurrency": 2,
786
+ "sample_concurrency": 8,
787
+ },
788
+ environ={},
789
+ )
790
+ settings.home.mkdir(parents=True, exist_ok=True)
791
+ settings.cache_dir.mkdir(parents=True, exist_ok=True)
792
+ store = SQLiteStore(settings.db_path)
793
+ cache = SQLiteCache(settings.cache_dir / "evaluations.db")
794
+ optimizer = Optimizer(
795
+ adapter=adapter,
796
+ strategy=strategy,
797
+ evaluators=(quality_evaluator, retrieval_evaluator),
798
+ objective=objective,
799
+ store=store,
800
+ cache=cache,
801
+ settings=settings,
802
+ stop_conditions=(
803
+ MaxTrials(30),
804
+ TargetScore(0.90),
805
+ NoImprovement(patience=6, min_delta=0.01, warmup_trials=8),
806
+ MaxFailures(5),
807
+ ),
808
+ retry_policy=RetryPolicy(
809
+ max_retries=2,
810
+ sample_timeout_seconds=45.0,
811
+ trial_timeout_seconds=300.0,
812
+ base_delay_seconds=0.5,
813
+ max_delay_seconds=10.0,
814
+ ),
815
+ )
816
+ try:
817
+ result = await optimizer.aoptimize(dataset, study_name="support-rag-v4")
818
+ finally:
819
+ cache.close()
820
+ store.close()
821
+ ```
822
+
823
+ When you pass a store, you own it. `optimizer.close()` closes only its internally created store.
824
+
825
+ ## Stopping, retries, and concurrency
826
+
827
+ Stop conditions are OR-composed:
828
+
829
+ | Condition | Stops when |
830
+ |---|---|
831
+ | `MaxTrials(limit)` | Completed trial count reaches the limit |
832
+ | `CostBudget(limit, count_estimated=True)` | Exact plus optional estimated cost reaches limit |
833
+ | `TimeBudget(seconds)` | Wall-clock duration reaches limit |
834
+ | `TargetScore(target)` | Any benefit-oriented objective reaches target |
835
+ | `MaxFailures(limit, consecutive=False)` | Total or trailing consecutive failures reach limit |
836
+ | `NoImprovement(patience, min_delta=0, warmup_trials=1)` | Best score stops improving |
837
+
838
+ Always include a hard bound. Conditions are checked between trial batches, so active trials can
839
+ finish after a cost, time, score, failure, or stagnation threshold is reached. Use `concurrency=1`
840
+ for strict between-trial stopping.
841
+
842
+ ```python
843
+ from tunarag import RetryPolicy
844
+
845
+ policy = RetryPolicy(
846
+ max_retries=3,
847
+ sample_timeout_seconds=30.0,
848
+ trial_timeout_seconds=180.0,
849
+ base_delay_seconds=0.25,
850
+ max_delay_seconds=8.0,
851
+ )
852
+ ```
853
+
854
+ Three retries means four total attempts. Timeouts and connection errors are retryable. A
855
+ `TunaRAGError` retries only when `retryable=True`; other exceptions fail that trial. Backoff is
856
+ deterministic and capped. Attempt usage is append-only, including failed attempts.
857
+
858
+ `concurrency` bounds simultaneous trials; `sample_concurrency` bounds samples inside each trial.
859
+ Their product approximates maximum concurrent adapter calls. Completion may be concurrent, but
860
+ results and strategy observations preserve trial-sequence order.
861
+
862
+ ## Caching
863
+
864
+ ```python
865
+ from pathlib import Path
866
+
867
+ from tunarag import CachePolicy, SQLiteCache, Settings
868
+
869
+ Path(".tunarag").mkdir(parents=True, exist_ok=True)
870
+ cache = SQLiteCache(".tunarag/evaluations.db")
871
+ settings = Settings.resolve(overrides={"cache_policy": CachePolicy.READ_WRITE}, environ={})
872
+ optimizer = Optimizer(
873
+ adapter=adapter,
874
+ strategy=strategy,
875
+ evaluators=evaluators,
876
+ objective=objective,
877
+ cache=cache,
878
+ settings=settings,
879
+ )
880
+ ```
881
+
882
+ Identity includes dataset version, candidate, adapter, and evaluators. A hit skips adapter and
883
+ evaluator execution but still calculates the objective and records the trial.
884
+
885
+ Direct API:
886
+
887
+ ```python
888
+ from tunarag import CacheKey, CachePrivacy, SQLiteCache
889
+
890
+ cache = SQLiteCache("cache.db")
891
+ key = CacheKey.from_parts("my-feature", {"query": "hello"}, format_version=1)
892
+ cache.put(key, b'{"answer":"world"}', privacy=CachePrivacy.SENSITIVE, ttl_seconds=3600)
893
+ lookup = cache.get(key)
894
+ if lookup.is_hit:
895
+ payload = lookup.payload
896
+ cache.prune()
897
+ cache.close()
898
+ ```
899
+
900
+ Statuses are `hit`, `miss`, `expired`, `corrupt`, and `quarantined`. Checksums detect corruption and
901
+ bad rows are quarantined. `prune(include_quarantined=True)` can remove quarantined rows. Privacy
902
+ labels are metadata, not encryption.
903
+
904
+ ## SQLite persistence and resume
905
+
906
+ ```python
907
+ from pathlib import Path
908
+
909
+ from tunarag import SQLiteStore
910
+
911
+ Path(".tunarag").mkdir(parents=True, exist_ok=True)
912
+ store = SQLiteStore(".tunarag/experiments.db")
913
+ optimizer = build_optimizer(store)
914
+ result = await optimizer.aoptimize(dataset, study_name="production-rag")
915
+
916
+ # Rebuild stateful components so replay begins from their initial state.
917
+ resumed_optimizer = build_optimizer(store)
918
+ resumed = await resumed_optimizer.aresume(dataset, study_id=result.study_id)
919
+ # Sync equivalent: resumed_optimizer.resume(dataset, study_id=result.study_id)
920
+ ```
921
+
922
+ Resume validates semantic configuration and dataset identity, returns completed studies without
923
+ provider execution, abandons expired leases, replays committed strategy history, continues
924
+ unfinished work, and schedules new candidates. Unknown, failed, cancelled, or incompatible studies
925
+ raise `ResumeError`. An active lease produces a retryable error; retry later.
926
+
927
+ Read persisted state:
928
+
929
+ ```python
930
+ study = store.get_study(result.study_id)
931
+ trials = store.list_trials(result.study_id)
932
+ events = store.list_events(result.study_id, after_sequence=-1)
933
+ attempts = store.list_attempts(trials[0].id)
934
+ metrics = store.list_metrics(trials[0].id)
935
+ usage = store.list_usage(trials[0].id)
936
+ ```
937
+
938
+ Normally let `Optimizer` own lifecycle mutations and use store read methods for inspection.
939
+
940
+ ## Callbacks and MLflow
941
+
942
+ Callbacks receive ordered durable events after commit:
943
+
944
+ ```python
945
+ class AuditCallback:
946
+ async def on_event(self, event):
947
+ print(event.sequence, event.type, event.study_id, dict(event.payload))
948
+
949
+
950
+ optimizer = Optimizer(
951
+ adapter=adapter,
952
+ strategy=strategy,
953
+ evaluators=evaluators,
954
+ objective=objective,
955
+ callbacks=(AuditCallback(),),
956
+ )
957
+ ```
958
+
959
+ Events cover study, trial, and attempt lifecycles plus callback failures. Handle unknown future event
960
+ types defensively. Callback failures never roll back committed work; TunaRAG records them as durable
961
+ events. Make callbacks idempotent.
962
+
963
+ ```mermaid
964
+ sequenceDiagram
965
+ participant O as Optimizer
966
+ participant S as SQLite
967
+ participant C as Callback
968
+ O->>S: Commit terminal trial and event
969
+ S-->>O: Committed
970
+ O->>C: on_event(event)
971
+ alt callback fails
972
+ O->>S: Append callback.failed
973
+ end
974
+ ```
975
+
976
+ MLflow is an optional post-commit projection; SQLite remains authoritative:
977
+
978
+ ```python
979
+ from tunarag.integrations import MLflowCallback
980
+
981
+ callback = MLflowCallback(
982
+ store,
983
+ experiment_id="42",
984
+ tracking_uri="https://mlflow.example.com",
985
+ tags={"team": "search", "environment": "staging"},
986
+ )
987
+ optimizer = Optimizer(
988
+ adapter=adapter,
989
+ strategy=strategy,
990
+ evaluators=evaluators,
991
+ objective=objective,
992
+ store=store,
993
+ callbacks=(callback,),
994
+ )
995
+ ```
996
+
997
+ One MLflow run is created per study. Candidate parameters, committed metrics, statuses, and event
998
+ position are mirrored. An MLflow outage does not invalidate local results. Automatic external
999
+ reconciliation is not included in V0.2.
1000
+
1001
+ ## Results and reports
1002
+
1003
+ `OptimizationResult` exposes:
1004
+
1005
+ ```python
1006
+ result.study_id
1007
+ result.study_status
1008
+ result.stop_reason
1009
+ result.trials
1010
+ result.successful_trials
1011
+ result.failed_trials
1012
+ result.best_trial
1013
+ result.best_candidate
1014
+ result.ranked_trials
1015
+ result.metric_names
1016
+ result.usage
1017
+ ```
1018
+
1019
+ Rank by an individual metric:
1020
+
1021
+ ```python
1022
+ from tunarag import ObjectiveDirection
1023
+
1024
+ quality_leaders = result.metric_leaders("answer_quality")
1025
+ latency_leaders = result.metric_leaders("latency_seconds", direction=ObjectiveDirection.MINIMIZE)
1026
+ ```
1027
+
1028
+ Inspect and export:
1029
+
1030
+ ```python
1031
+ from pathlib import Path
1032
+
1033
+ if result.ranked_trials:
1034
+ trial = result.ranked_trials[0]
1035
+ print(trial.candidate.parameters)
1036
+ print(trial.objective.score)
1037
+ print(trial.metric("answer_quality"))
1038
+ print(trial.attempts, trial.cached, trial.error)
1039
+
1040
+ print(result.to_json())
1041
+ Path("artifacts").mkdir(parents=True, exist_ok=True)
1042
+ result.write_json("artifacts/result.json")
1043
+ result.write_csv("artifacts/trials.csv")
1044
+ ```
1045
+
1046
+ Reports are deterministic and redacted. Usage separates exact and estimated cost. A completed result
1047
+ may have no best candidate if every trial failed or every objective was unavailable.
1048
+
1049
+ ## Errors, privacy, and secrets
1050
+
1051
+ ```python
1052
+ from tunarag import TunaRAGError
1053
+
1054
+ try:
1055
+ result = optimizer.optimize(dataset, study_name="example")
1056
+ except TunaRAGError as error:
1057
+ diagnostic = error.as_dict()
1058
+ print(diagnostic["code"], diagnostic["stage"], diagnostic["retryable"])
1059
+ ```
1060
+
1061
+ Public subclasses include `AdapterError`, `BudgetError`, `ConfigurationError`, `DatasetError`,
1062
+ `EvaluationError`, `MissingOptionalDependencyError`, `ResumeError`, `SearchSpaceError`, `StoreError`,
1063
+ and `SyncInAsyncContextError`.
1064
+
1065
+ ```python
1066
+ from tunarag import Secret
1067
+
1068
+ credential = Secret("do-not-log-this")
1069
+ print(credential) # ***redacted***
1070
+ print(repr(credential)) # Secret(***redacted***)
1071
+ ```
1072
+
1073
+ TunaRAG redacts `Secret` and common sensitive keys from diagnostics, fingerprints, events, and
1074
+ reports. This is defense in depth, not a secret manager. SQLite/cache files can contain prompts and
1075
+ metadata; privacy labels do not encrypt. LLMSearch, synthetic generation, RAGAS, and callbacks may
1076
+ send data to external providers. Apply filesystem permissions, encrypted storage, data minimization,
1077
+ and provider-specific retention controls.
1078
+
1079
+ ## Extension contracts
1080
+
1081
+ Protocols are structural; inheritance is not required.
1082
+
1083
+ | Contract | Required behavior |
1084
+ |---|---|
1085
+ | `RAGAdapter` | `async run(candidate, example) -> (output, usage)` |
1086
+ | `Dataset` | `.version` and async iteration |
1087
+ | `Evaluator` | `async evaluate(output, example) -> metrics` |
1088
+ | `SearchStrategy` | `async suggest()` and `async observe(...)` |
1089
+ | `ReplayableSearchStrategy` | `async replay(...)` for resume |
1090
+ | `Fingerprintable` | stable secret-safe `fingerprint()` |
1091
+ | `EventCallback` | `async on_event(event)` |
1092
+ | `Cache` | synchronous `get()` and `put()` |
1093
+
1094
+ Custom search-space parameters implement `sample(generator)`, `accepts(value)`, and `describe()`:
1095
+
1096
+ ```python
1097
+ class EvenInteger:
1098
+ def __init__(self, minimum, maximum):
1099
+ self.minimum = minimum
1100
+ self.maximum = maximum
1101
+
1102
+ def sample(self, generator):
1103
+ values = tuple(range(self.minimum + self.minimum % 2, self.maximum + 1, 2))
1104
+ return generator.choice(values)
1105
+
1106
+ def accepts(self, value):
1107
+ return (
1108
+ isinstance(value, int)
1109
+ and not isinstance(value, bool)
1110
+ and self.minimum <= value <= self.maximum
1111
+ and value % 2 == 0
1112
+ )
1113
+
1114
+ def describe(self):
1115
+ return {
1116
+ "type": "integer",
1117
+ "minimum": self.minimum,
1118
+ "maximum": self.maximum,
1119
+ "multiple_of": 2,
1120
+ }
1121
+ ```
1122
+
1123
+ Provider descriptions and component fingerprints must not expose secrets.
1124
+
1125
+ ## Production checklist
1126
+
1127
+ - Use a file-backed `SQLiteStore` and back it up.
1128
+ - Keep a human-reviewed test split outside optimization.
1129
+ - Fingerprint custom adapters and evaluators.
1130
+ - Include a hard stopping bound and conservative concurrency.
1131
+ - Track exact usage when available and label estimates.
1132
+ - Configure sample and trial timeouts.
1133
+ - Use deterministic seeds for reproducibility.
1134
+ - Protect database and cache files; keep credentials outside TunaRAG objects.
1135
+ - Make callbacks idempotent and monitor callback-failure events.
1136
+ - Retain study IDs and exported reports with releases.
1137
+ - Test resume behavior before expensive studies.
1138
+ - Re-evaluate the winner on validation and test sets.
1139
+
1140
+ ## Limitations and glossary
1141
+
1142
+ V0.2 does not provide a hosted service/dashboard, automatic RAG construction, vector-database
1143
+ provisioning, distributed workers, Pareto optimization, Bayesian/evolutionary search, encrypted
1144
+ payloads, MLflow reconciliation, provider pricing management, or a TypeScript implementation.
1145
+ LangChain and LangGraph are official integrations; other frameworks use `RAGAdapter`.
1146
+
1147
+ | Term | Meaning |
1148
+ |---|---|
1149
+ | Adapter | Runs the user-owned RAG for one candidate and example |
1150
+ | Attempt | One trial execution, including retries |
1151
+ | Candidate | Complete parameter mapping proposed by a strategy |
1152
+ | Dataset version | Identity used for cache and resume compatibility |
1153
+ | Evaluator | Converts one output/example pair into metrics |
1154
+ | Fingerprint | Secret-safe semantic component identity |
1155
+ | Objective | Converts trial metrics into one benefit score |
1156
+ | Sample | One candidate evaluated against one example |
1157
+ | Strategy | Proposes candidates and observes outcomes |
1158
+ | Study | Durable optimization run containing trials |
1159
+ | Trial | One candidate evaluated over the whole dataset |
1160
+ | Usage | Tokens, cost, latency, confidence, and pricing provenance |
1161
+
1162
+ For deeper contracts and architecture, see [`docs/api-specification.md`](docs/api-specification.md),
1163
+ [`docs/system-architecture-document.md`](docs/system-architecture-document.md), and
1164
+ [`docs/technical-design-document.md`](docs/technical-design-document.md).