tunarag-python 0.2.1__py3-none-any.whl
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- tunarag/__init__.py +166 -0
- tunarag/cache.py +326 -0
- tunarag/config.py +257 -0
- tunarag/contracts.py +88 -0
- tunarag/dataset.py +481 -0
- tunarag/domain.py +83 -0
- tunarag/engine.py +979 -0
- tunarag/errors.py +179 -0
- tunarag/evaluators.py +186 -0
- tunarag/integrations/__init__.py +25 -0
- tunarag/integrations/mlflow.py +155 -0
- tunarag/integrations/runnables.py +270 -0
- tunarag/objective.py +75 -0
- tunarag/py.typed +1 -0
- tunarag/result.py +309 -0
- tunarag/retry.py +70 -0
- tunarag/search.py +286 -0
- tunarag/serialization.py +78 -0
- tunarag/stopping.py +193 -0
- tunarag/store.py +941 -0
- tunarag/synthetic.py +257 -0
- tunarag_python-0.2.1.dist-info/METADATA +1164 -0
- tunarag_python-0.2.1.dist-info/RECORD +24 -0
- tunarag_python-0.2.1.dist-info/WHEEL +4 -0
|
@@ -0,0 +1,1164 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: tunarag-python
|
|
3
|
+
Version: 0.2.1
|
|
4
|
+
Summary: Typed, durable optimization workflows for existing RAG systems
|
|
5
|
+
Project-URL: Repository, https://github.com/shivamshinde123/tunarag-python
|
|
6
|
+
Project-URL: Issues, https://github.com/shivamshinde123/tunarag-python/issues
|
|
7
|
+
Project-URL: Changelog, https://github.com/shivamshinde123/tunarag-python/blob/main/CHANGELOG.md
|
|
8
|
+
Author: Shivam Shinde
|
|
9
|
+
License-Expression: MIT
|
|
10
|
+
Keywords: evaluation,llm,optimization,rag,retrieval
|
|
11
|
+
Classifier: Development Status :: 3 - Alpha
|
|
12
|
+
Classifier: Intended Audience :: Developers
|
|
13
|
+
Classifier: License :: OSI Approved :: MIT License
|
|
14
|
+
Classifier: Operating System :: OS Independent
|
|
15
|
+
Classifier: Programming Language :: Python :: 3
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
19
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
20
|
+
Classifier: Typing :: Typed
|
|
21
|
+
Requires-Python: >=3.10
|
|
22
|
+
Requires-Dist: tomli>=2.0; python_version < '3.11'
|
|
23
|
+
Provides-Extra: dev
|
|
24
|
+
Requires-Dist: build>=1.2; extra == 'dev'
|
|
25
|
+
Requires-Dist: hypothesis>=6.100; extra == 'dev'
|
|
26
|
+
Requires-Dist: mypy>=1.11; extra == 'dev'
|
|
27
|
+
Requires-Dist: pytest-asyncio>=0.23; extra == 'dev'
|
|
28
|
+
Requires-Dist: pytest>=8.0; extra == 'dev'
|
|
29
|
+
Requires-Dist: ruff>=0.6; extra == 'dev'
|
|
30
|
+
Provides-Extra: frameworks
|
|
31
|
+
Requires-Dist: langchain<2,>=1; extra == 'frameworks'
|
|
32
|
+
Requires-Dist: langgraph<2,>=1; extra == 'frameworks'
|
|
33
|
+
Provides-Extra: langchain
|
|
34
|
+
Requires-Dist: langchain<2,>=1; extra == 'langchain'
|
|
35
|
+
Provides-Extra: langgraph
|
|
36
|
+
Requires-Dist: langgraph<2,>=1; extra == 'langgraph'
|
|
37
|
+
Provides-Extra: mlflow
|
|
38
|
+
Requires-Dist: mlflow<4,>=3; extra == 'mlflow'
|
|
39
|
+
Provides-Extra: ragas
|
|
40
|
+
Requires-Dist: ragas<0.5,>=0.4; extra == 'ragas'
|
|
41
|
+
Description-Content-Type: text/markdown
|
|
42
|
+
|
|
43
|
+
# TunaRAG Python
|
|
44
|
+
|
|
45
|
+
**Find the best configuration for your existing RAG application.**
|
|
46
|
+
|
|
47
|
+
TunaRAG is an open-source, Python-first optimization toolkit for retrieval-augmented generation
|
|
48
|
+
(RAG) systems. It connects to a RAG pipeline you already own, evaluates different configurations on
|
|
49
|
+
your data, and identifies the best tradeoff between **answer quality, cost, and latency**.
|
|
50
|
+
|
|
51
|
+
Instead of manually changing retrieval settings, prompts, or models and comparing a few examples by
|
|
52
|
+
eye, TunaRAG runs structured, repeatable optimization studies. It records every trial in SQLite,
|
|
53
|
+
isolates failures, supports resumable execution, and produces results you can inspect and export.
|
|
54
|
+
|
|
55
|
+
### What can you use it for?
|
|
56
|
+
|
|
57
|
+
- Find the best `top_k`, chunking, prompt, model, reranking, or generation settings for your RAG.
|
|
58
|
+
- Compare configurations with custom metrics or optional RAGAS evaluators.
|
|
59
|
+
- Optimize a weighted objective across quality, provider cost, and response latency.
|
|
60
|
+
- Search with deterministic random search or provider-neutral LLM-guided search.
|
|
61
|
+
- Tune existing custom, LangChain, or LangGraph RAG pipelines without rebuilding them.
|
|
62
|
+
- Preserve reproducible experiments with caching, retries, callbacks, reports, and SQLite resume.
|
|
63
|
+
|
|
64
|
+
### Who is it for?
|
|
65
|
+
|
|
66
|
+
TunaRAG is designed for developers and AI teams that already have a working RAG pipeline and want a
|
|
67
|
+
reliable way to evaluate and improve it. Your application continues to own retrieval, generation,
|
|
68
|
+
indexes, provider clients, and credentials; TunaRAG owns the optimization workflow around it.
|
|
69
|
+
|
|
70
|
+
> Install the distribution as `tunarag-python` and import it as `tunarag`. Python 3.10-3.13 is
|
|
71
|
+
> supported.
|
|
72
|
+
|
|
73
|
+
## Comprehensive Usage Guide
|
|
74
|
+
|
|
75
|
+
The sections below explain every public TunaRAG Python V0.2 feature with executable examples.
|
|
76
|
+
|
|
77
|
+
## Contents
|
|
78
|
+
|
|
79
|
+
1. [Purpose and mental model](#purpose-and-mental-model)
|
|
80
|
+
2. [Installation](#installation)
|
|
81
|
+
3. [End-to-end quickstart](#end-to-end-quickstart)
|
|
82
|
+
4. [RAG adapters](#rag-adapters)
|
|
83
|
+
5. [LangChain integration](#langchain-integration)
|
|
84
|
+
6. [LangGraph integration](#langgraph-integration)
|
|
85
|
+
7. [Evaluation datasets](#evaluation-datasets)
|
|
86
|
+
8. [Synthetic QA generation](#synthetic-qa-generation)
|
|
87
|
+
9. [Search spaces and strategies](#search-spaces-and-strategies)
|
|
88
|
+
10. [Evaluators and RAGAS](#evaluators-and-ragas)
|
|
89
|
+
11. [Objectives](#objectives)
|
|
90
|
+
12. [Configuration](#configuration)
|
|
91
|
+
13. [Running optimization](#running-optimization)
|
|
92
|
+
14. [Stopping, retries, and concurrency](#stopping-retries-and-concurrency)
|
|
93
|
+
15. [Caching](#caching)
|
|
94
|
+
16. [SQLite persistence and resume](#sqlite-persistence-and-resume)
|
|
95
|
+
17. [Callbacks and MLflow](#callbacks-and-mlflow)
|
|
96
|
+
18. [Results and reports](#results-and-reports)
|
|
97
|
+
19. [Errors, privacy, and secrets](#errors-privacy-and-secrets)
|
|
98
|
+
20. [Extension contracts](#extension-contracts)
|
|
99
|
+
21. [Production checklist](#production-checklist)
|
|
100
|
+
22. [Limitations and glossary](#limitations-and-glossary)
|
|
101
|
+
|
|
102
|
+
## Purpose and mental model
|
|
103
|
+
|
|
104
|
+
TunaRAG compares configurations of an **existing** RAG pipeline. You provide the pipeline, an
|
|
105
|
+
evaluation dataset, a search space, evaluators, and an objective. TunaRAG schedules trials, isolates
|
|
106
|
+
failures, persists state, and returns an inspectable result.
|
|
107
|
+
|
|
108
|
+
```mermaid
|
|
109
|
+
flowchart LR
|
|
110
|
+
D[Versioned dataset] --> O[Optimizer]
|
|
111
|
+
S[Search strategy] -->|Candidate| O
|
|
112
|
+
O -->|Candidate + example| A[RAG adapter]
|
|
113
|
+
A -->|Output + usage| E[Evaluators]
|
|
114
|
+
E -->|Metrics| J[Objective]
|
|
115
|
+
J -->|Benefit score| O
|
|
116
|
+
O -->|Observe| S
|
|
117
|
+
O --> DB[(SQLite store)]
|
|
118
|
+
O <--> C[(Evaluation cache)]
|
|
119
|
+
DB --> CB[Post-commit callbacks]
|
|
120
|
+
O --> R[OptimizationResult]
|
|
121
|
+
```
|
|
122
|
+
|
|
123
|
+
A **study** contains ordered **trials**. Each trial evaluates one complete candidate configuration
|
|
124
|
+
against every dataset example. Same-named sample metrics are aggregated into trial metrics, then the
|
|
125
|
+
objective converts those metrics into one larger-is-better score.
|
|
126
|
+
|
|
127
|
+
TunaRAG does not create or host a RAG application, provision indexes, manage provider credentials,
|
|
128
|
+
or expose a hosted API. Your application keeps ownership of those concerns.
|
|
129
|
+
|
|
130
|
+
## Installation
|
|
131
|
+
|
|
132
|
+
```bash
|
|
133
|
+
python -m pip install tunarag-python
|
|
134
|
+
```
|
|
135
|
+
|
|
136
|
+
| Optional feature | Installation |
|
|
137
|
+
|---|---|
|
|
138
|
+
| LangChain | `python -m pip install "tunarag-python[langchain]"` |
|
|
139
|
+
| LangGraph | `python -m pip install "tunarag-python[langgraph]"` |
|
|
140
|
+
| Both frameworks | `python -m pip install "tunarag-python[frameworks]"` |
|
|
141
|
+
| RAGAS | `python -m pip install "tunarag-python[ragas]"` |
|
|
142
|
+
| MLflow | `python -m pip install "tunarag-python[mlflow]"` |
|
|
143
|
+
|
|
144
|
+
Optional dependencies load lazily; the core package does not require them.
|
|
145
|
+
|
|
146
|
+
## End-to-end quickstart
|
|
147
|
+
|
|
148
|
+
This deterministic example shows the complete public workflow without external providers.
|
|
149
|
+
|
|
150
|
+
```python
|
|
151
|
+
from tunarag import (
|
|
152
|
+
EvaluationDataset,
|
|
153
|
+
IntegerRange,
|
|
154
|
+
MetricEvaluator,
|
|
155
|
+
Objective,
|
|
156
|
+
ObjectiveTerm,
|
|
157
|
+
Optimizer,
|
|
158
|
+
RandomSearch,
|
|
159
|
+
SearchSpace,
|
|
160
|
+
UsageRecord,
|
|
161
|
+
)
|
|
162
|
+
|
|
163
|
+
DOCUMENTS = {
|
|
164
|
+
"python": "Python is a high-level programming language.",
|
|
165
|
+
"sqlite": "SQLite is a local embedded relational database.",
|
|
166
|
+
"rag": "RAG combines retrieval with generation.",
|
|
167
|
+
}
|
|
168
|
+
|
|
169
|
+
|
|
170
|
+
class SmallRAG:
|
|
171
|
+
async def run(self, candidate, example):
|
|
172
|
+
query_words = set(example.query.lower().split())
|
|
173
|
+
ranked = sorted(
|
|
174
|
+
DOCUMENTS.values(),
|
|
175
|
+
key=lambda text: len(query_words & set(text.lower().split())),
|
|
176
|
+
reverse=True,
|
|
177
|
+
)
|
|
178
|
+
contexts = ranked[: candidate.parameters["top_k"]]
|
|
179
|
+
return {"answer": " ".join(contexts)}, (UsageRecord("small-rag", latency_seconds=0.001),)
|
|
180
|
+
|
|
181
|
+
|
|
182
|
+
def answer_recall(output, example):
|
|
183
|
+
expected = set(example.reference_answer.lower().split())
|
|
184
|
+
actual = set(output["answer"].lower().split())
|
|
185
|
+
return len(expected & actual) / len(expected)
|
|
186
|
+
|
|
187
|
+
|
|
188
|
+
dataset = EvaluationDataset.from_records(
|
|
189
|
+
[
|
|
190
|
+
{
|
|
191
|
+
"id": "q1",
|
|
192
|
+
"query": "What is SQLite?",
|
|
193
|
+
"reference_answer": "a local embedded relational database",
|
|
194
|
+
},
|
|
195
|
+
{
|
|
196
|
+
"id": "q2",
|
|
197
|
+
"query": "What does RAG combine?",
|
|
198
|
+
"reference_answer": "retrieval with generation",
|
|
199
|
+
},
|
|
200
|
+
]
|
|
201
|
+
)
|
|
202
|
+
optimizer = Optimizer(
|
|
203
|
+
adapter=SmallRAG(),
|
|
204
|
+
strategy=RandomSearch(SearchSpace({"top_k": IntegerRange(1, 3)}), seed=7),
|
|
205
|
+
evaluators=(MetricEvaluator("answer_recall", answer_recall),),
|
|
206
|
+
objective=Objective((ObjectiveTerm("answer_recall"),)),
|
|
207
|
+
)
|
|
208
|
+
try:
|
|
209
|
+
result = optimizer.optimize(dataset, study_name="small-rag")
|
|
210
|
+
print(result.best_candidate.parameters if result.best_candidate else "no winner")
|
|
211
|
+
finally:
|
|
212
|
+
optimizer.close()
|
|
213
|
+
```
|
|
214
|
+
|
|
215
|
+
The default store is in memory. Use `SQLiteStore(path)` for a durable study.
|
|
216
|
+
|
|
217
|
+
## RAG adapters
|
|
218
|
+
|
|
219
|
+
A custom adapter implements:
|
|
220
|
+
|
|
221
|
+
```python
|
|
222
|
+
from collections.abc import Sequence
|
|
223
|
+
|
|
224
|
+
from tunarag import UsageRecord
|
|
225
|
+
|
|
226
|
+
|
|
227
|
+
async def run(candidate, example) -> tuple[object, Sequence[UsageRecord]]: ...
|
|
228
|
+
```
|
|
229
|
+
|
|
230
|
+
`candidate.parameters` contains the proposed configuration. The output can be any object understood
|
|
231
|
+
by your evaluators. Usage may be empty. Adapter exceptions enter TunaRAG's retry and trial-failure
|
|
232
|
+
boundary.
|
|
233
|
+
|
|
234
|
+
```python
|
|
235
|
+
from tunarag import UsageRecord, ValueStatus
|
|
236
|
+
|
|
237
|
+
|
|
238
|
+
class ExistingRAGAdapter:
|
|
239
|
+
def __init__(self, service):
|
|
240
|
+
self.service = service
|
|
241
|
+
|
|
242
|
+
async def run(self, candidate, example):
|
|
243
|
+
response = await self.service.answer(
|
|
244
|
+
query=example.query,
|
|
245
|
+
top_k=candidate.parameters["top_k"],
|
|
246
|
+
model=candidate.parameters["model"],
|
|
247
|
+
)
|
|
248
|
+
return response, (
|
|
249
|
+
UsageRecord(
|
|
250
|
+
component="generator",
|
|
251
|
+
input_tokens=response.input_tokens,
|
|
252
|
+
output_tokens=response.output_tokens,
|
|
253
|
+
cost=response.cost,
|
|
254
|
+
latency_seconds=response.latency_seconds,
|
|
255
|
+
status=ValueStatus.EXACT,
|
|
256
|
+
pricing_version="provider-2026-10",
|
|
257
|
+
),
|
|
258
|
+
)
|
|
259
|
+
|
|
260
|
+
def fingerprint(self):
|
|
261
|
+
return {
|
|
262
|
+
"adapter_version": 2,
|
|
263
|
+
"index_revision": "docs-2026-10-04",
|
|
264
|
+
"prompt_revision": "qa-v3",
|
|
265
|
+
}
|
|
266
|
+
```
|
|
267
|
+
|
|
268
|
+
Use `ValueStatus.ESTIMATED` for inferred usage and `UNAVAILABLE` when it cannot be determined.
|
|
269
|
+
Implement `fingerprint()` when class identity alone does not represent semantics. Fingerprints affect
|
|
270
|
+
cache and resume compatibility; include stable versions, never credentials.
|
|
271
|
+
|
|
272
|
+
## LangChain integration
|
|
273
|
+
|
|
274
|
+
`LangChainAdapter` accepts any asynchronous runnable with `ainvoke(input, config=...)`.
|
|
275
|
+
|
|
276
|
+
```python
|
|
277
|
+
from langchain_core.runnables import RunnableLambda
|
|
278
|
+
from tunarag import LangChainAdapter
|
|
279
|
+
|
|
280
|
+
|
|
281
|
+
async def answer(inputs, config):
|
|
282
|
+
top_k = config["configurable"]["top_k"]
|
|
283
|
+
return {"answer": f"Used top_k={top_k} for {inputs['question']}"}
|
|
284
|
+
|
|
285
|
+
|
|
286
|
+
chain = RunnableLambda(answer)
|
|
287
|
+
adapter = LangChainAdapter(
|
|
288
|
+
chain,
|
|
289
|
+
output_mapper=lambda output: output["answer"],
|
|
290
|
+
identity={"chain": "support-rag", "version": 1},
|
|
291
|
+
)
|
|
292
|
+
```
|
|
293
|
+
|
|
294
|
+
For `EvaluationExample`, default input is `{"question": example.query}`. Candidate parameters are
|
|
295
|
+
placed in `config["configurable"]`. Other example types pass through unchanged.
|
|
296
|
+
|
|
297
|
+
Use a factory when candidates change construction-time components:
|
|
298
|
+
|
|
299
|
+
```python
|
|
300
|
+
def build_chain(candidate):
|
|
301
|
+
retriever = vector_store.as_retriever(search_kwargs={"k": candidate.parameters["top_k"]})
|
|
302
|
+
return create_retrieval_chain(
|
|
303
|
+
retriever,
|
|
304
|
+
build_qa_chain(candidate.parameters["model"]),
|
|
305
|
+
)
|
|
306
|
+
|
|
307
|
+
|
|
308
|
+
adapter = LangChainAdapter(
|
|
309
|
+
factory=build_chain,
|
|
310
|
+
output_mapper=lambda output: output["answer"],
|
|
311
|
+
identity={"factory_schema": 1},
|
|
312
|
+
)
|
|
313
|
+
```
|
|
314
|
+
|
|
315
|
+
Customize all boundaries when required:
|
|
316
|
+
|
|
317
|
+
```python
|
|
318
|
+
adapter = LangChainAdapter(
|
|
319
|
+
chain,
|
|
320
|
+
input_mapper=lambda example, candidate: {
|
|
321
|
+
"input": example.query,
|
|
322
|
+
"retrieval_limit": candidate.parameters["top_k"],
|
|
323
|
+
},
|
|
324
|
+
config_mapper=lambda candidate: {
|
|
325
|
+
"configurable": {"model": candidate.parameters["model"]},
|
|
326
|
+
"tags": ["tunarag-trial"],
|
|
327
|
+
},
|
|
328
|
+
output_mapper=lambda output: output["answer"],
|
|
329
|
+
usage_mapper=lambda output: (
|
|
330
|
+
UsageRecord(
|
|
331
|
+
"langchain-provider",
|
|
332
|
+
input_tokens=output["usage"]["input_tokens"],
|
|
333
|
+
output_tokens=output["usage"]["output_tokens"],
|
|
334
|
+
),
|
|
335
|
+
),
|
|
336
|
+
)
|
|
337
|
+
```
|
|
338
|
+
|
|
339
|
+
The default usage mapper recursively discovers standard `usage_metadata` in mappings, sequences,
|
|
340
|
+
and message objects. The adapter always adds `langchain.runnable` latency. Set
|
|
341
|
+
`capture_message_usage=False` to disable automatic token extraction. A custom `usage_mapper` takes
|
|
342
|
+
precedence.
|
|
343
|
+
|
|
344
|
+
## LangGraph integration
|
|
345
|
+
|
|
346
|
+
`LangGraphAdapter` has the same factory and mapping options.
|
|
347
|
+
|
|
348
|
+
```python
|
|
349
|
+
from typing import TypedDict
|
|
350
|
+
|
|
351
|
+
from langgraph.graph import END, START, StateGraph
|
|
352
|
+
from tunarag import LangGraphAdapter
|
|
353
|
+
|
|
354
|
+
|
|
355
|
+
class State(TypedDict, total=False):
|
|
356
|
+
question: str
|
|
357
|
+
answer: str
|
|
358
|
+
top_k: int
|
|
359
|
+
|
|
360
|
+
|
|
361
|
+
async def answer_node(state: State, config):
|
|
362
|
+
top_k = state["top_k"] if "top_k" in state else config["configurable"]["top_k"]
|
|
363
|
+
return {"answer": f"answer for {state['question']} with {top_k} documents"}
|
|
364
|
+
|
|
365
|
+
|
|
366
|
+
builder = StateGraph(State)
|
|
367
|
+
builder.add_node("answer", answer_node)
|
|
368
|
+
builder.add_edge(START, "answer")
|
|
369
|
+
builder.add_edge("answer", END)
|
|
370
|
+
graph = builder.compile()
|
|
371
|
+
adapter = LangGraphAdapter(
|
|
372
|
+
graph,
|
|
373
|
+
output_mapper=lambda state: state["answer"],
|
|
374
|
+
identity={"graph": "qa", "revision": 1},
|
|
375
|
+
)
|
|
376
|
+
```
|
|
377
|
+
|
|
378
|
+
To place candidate values in graph state instead of runnable config:
|
|
379
|
+
|
|
380
|
+
```python
|
|
381
|
+
adapter = LangGraphAdapter(
|
|
382
|
+
graph,
|
|
383
|
+
input_mapper=lambda example, candidate: {
|
|
384
|
+
"question": example.query,
|
|
385
|
+
"top_k": candidate.parameters["top_k"],
|
|
386
|
+
},
|
|
387
|
+
config_mapper=lambda candidate: None,
|
|
388
|
+
output_mapper=lambda state: state["answer"],
|
|
389
|
+
)
|
|
390
|
+
```
|
|
391
|
+
|
|
392
|
+
The adapter always adds `langgraph.runnable` latency. Framework errors remain visible to the engine's
|
|
393
|
+
retry and failure handling.
|
|
394
|
+
|
|
395
|
+
## Evaluation datasets
|
|
396
|
+
|
|
397
|
+
### Canonical model
|
|
398
|
+
|
|
399
|
+
```python
|
|
400
|
+
from tunarag import DatasetSplit, EvaluationExample
|
|
401
|
+
|
|
402
|
+
example = EvaluationExample(
|
|
403
|
+
id="billing-001",
|
|
404
|
+
query="How can I update my billing address?",
|
|
405
|
+
reference_answer="Open Settings, then Billing.",
|
|
406
|
+
reference_contexts=("Billing addresses are under Settings > Billing.",),
|
|
407
|
+
relevant_document_ids=("billing-guide",),
|
|
408
|
+
tags=("billing", "how-to"),
|
|
409
|
+
metadata={"tenant": "demo", "source": "support-review"},
|
|
410
|
+
split=DatasetSplit.OPTIMIZE,
|
|
411
|
+
synthetic=False,
|
|
412
|
+
)
|
|
413
|
+
```
|
|
414
|
+
|
|
415
|
+
Text is Unicode-normalized and line endings are canonicalized. IDs must be unique. Metadata must be
|
|
416
|
+
finite JSON-compatible data. Duplicate tags and relevant document IDs are rejected.
|
|
417
|
+
|
|
418
|
+
### Ingestion
|
|
419
|
+
|
|
420
|
+
```python
|
|
421
|
+
from tunarag import DatasetFieldMap, EvaluationDataset
|
|
422
|
+
|
|
423
|
+
dataset = EvaluationDataset.from_records(
|
|
424
|
+
[
|
|
425
|
+
{
|
|
426
|
+
"question_id": "q-1",
|
|
427
|
+
"question": "What is TunaRAG?",
|
|
428
|
+
"expected": "A RAG optimization library.",
|
|
429
|
+
"contexts": ["TunaRAG optimizes existing RAG systems."],
|
|
430
|
+
"metadata": {"tenant": "demo"},
|
|
431
|
+
}
|
|
432
|
+
],
|
|
433
|
+
field_map=DatasetFieldMap(
|
|
434
|
+
id="question_id",
|
|
435
|
+
query="question",
|
|
436
|
+
reference_answer="expected",
|
|
437
|
+
reference_contexts="contexts",
|
|
438
|
+
),
|
|
439
|
+
)
|
|
440
|
+
json_dataset = EvaluationDataset.from_json("evaluation.json")
|
|
441
|
+
jsonl_dataset = EvaluationDataset.from_jsonl("evaluation.jsonl")
|
|
442
|
+
csv_dataset = EvaluationDataset.from_csv("evaluation.csv")
|
|
443
|
+
```
|
|
444
|
+
|
|
445
|
+
JSON contains an array; JSONL contains one object per nonblank line. CSV collection and metadata
|
|
446
|
+
cells are JSON-encoded. Missing IDs are deterministically derived. Every `EvaluationDataset` has a
|
|
447
|
+
content-derived `version` used by caches and resume validation.
|
|
448
|
+
|
|
449
|
+
### Deterministic splitting
|
|
450
|
+
|
|
451
|
+
```python
|
|
452
|
+
from tunarag import DatasetSplit, SplitRatios
|
|
453
|
+
|
|
454
|
+
split_dataset = dataset.with_splits(
|
|
455
|
+
seed=42,
|
|
456
|
+
ratios=SplitRatios(optimize=0.7, validation=0.2, test=0.1),
|
|
457
|
+
group_metadata_key="tenant",
|
|
458
|
+
)
|
|
459
|
+
optimization_data = split_dataset.select(DatasetSplit.OPTIMIZE)
|
|
460
|
+
```
|
|
461
|
+
|
|
462
|
+
Splits are hash-based and stable across input order. Group splitting keeps related examples
|
|
463
|
+
together. `select()` rejects an empty split.
|
|
464
|
+
|
|
465
|
+
For custom examples, use a user-managed version:
|
|
466
|
+
|
|
467
|
+
```python
|
|
468
|
+
from tunarag import InMemoryDataset
|
|
469
|
+
|
|
470
|
+
dataset = InMemoryDataset.from_iterable(
|
|
471
|
+
[{"question": "Q1"}, {"question": "Q2"}],
|
|
472
|
+
version="support-set-v4",
|
|
473
|
+
)
|
|
474
|
+
```
|
|
475
|
+
|
|
476
|
+
Change the version whenever custom example semantics change.
|
|
477
|
+
|
|
478
|
+
## Synthetic QA generation
|
|
479
|
+
|
|
480
|
+
Generation is provider-neutral. A provider must return exactly the requested count.
|
|
481
|
+
|
|
482
|
+
```python
|
|
483
|
+
from tunarag import (
|
|
484
|
+
SourceDocument,
|
|
485
|
+
SourceSpan,
|
|
486
|
+
SyntheticDatasetBuilder,
|
|
487
|
+
SyntheticGenerationResponse,
|
|
488
|
+
SyntheticQA,
|
|
489
|
+
UsageRecord,
|
|
490
|
+
)
|
|
491
|
+
|
|
492
|
+
|
|
493
|
+
class QAProvider:
|
|
494
|
+
async def generate(self, request):
|
|
495
|
+
text = request.document.text
|
|
496
|
+
items = tuple(
|
|
497
|
+
SyntheticQA(
|
|
498
|
+
query=f"What does {request.document.id} explain?",
|
|
499
|
+
answer=text,
|
|
500
|
+
source_spans=(SourceSpan(0, len(text)),),
|
|
501
|
+
confidence=0.9,
|
|
502
|
+
)
|
|
503
|
+
for _ in range(request.count)
|
|
504
|
+
)
|
|
505
|
+
return SyntheticGenerationResponse(
|
|
506
|
+
items=items,
|
|
507
|
+
usage=(UsageRecord("qa-generator", output_tokens=40),),
|
|
508
|
+
)
|
|
509
|
+
|
|
510
|
+
|
|
511
|
+
builder = SyntheticDatasetBuilder(
|
|
512
|
+
QAProvider(),
|
|
513
|
+
samples_per_document=2,
|
|
514
|
+
concurrency=4,
|
|
515
|
+
seed=7,
|
|
516
|
+
generator_version="qa-provider-v1",
|
|
517
|
+
)
|
|
518
|
+
generated = await builder.generate(
|
|
519
|
+
[SourceDocument("guide", "TunaRAG optimizes existing RAG pipelines.")]
|
|
520
|
+
)
|
|
521
|
+
dataset = generated.dataset
|
|
522
|
+
generation_usage = generated.usage
|
|
523
|
+
```
|
|
524
|
+
|
|
525
|
+
`SourceSpan` is a zero-based half-open range. TunaRAG validates spans, confidence, cardinality,
|
|
526
|
+
source IDs, and usage. Generated examples carry provenance and `synthetic=True`. Review them before
|
|
527
|
+
using them as gold data. Generation usage is separate from optimization usage.
|
|
528
|
+
|
|
529
|
+
## Search spaces and strategies
|
|
530
|
+
|
|
531
|
+
### Typed space
|
|
532
|
+
|
|
533
|
+
```python
|
|
534
|
+
from tunarag import Categorical, FloatRange, IntegerRange, SearchSpace
|
|
535
|
+
|
|
536
|
+
space = SearchSpace(
|
|
537
|
+
{
|
|
538
|
+
"top_k": IntegerRange(1, 10), # inclusive
|
|
539
|
+
"temperature": FloatRange(0.0, 0.8),
|
|
540
|
+
"model": Categorical(("small", "large")),
|
|
541
|
+
"prompt": Categorical(("concise", "grounded")),
|
|
542
|
+
}
|
|
543
|
+
)
|
|
544
|
+
```
|
|
545
|
+
|
|
546
|
+
`SearchSpace.validate()` requires exactly the declared keys. Prefer JSON-like immutable categorical
|
|
547
|
+
values for portable persistence. Secrets should not be search parameters.
|
|
548
|
+
|
|
549
|
+
### RandomSearch
|
|
550
|
+
|
|
551
|
+
```python
|
|
552
|
+
from tunarag import RandomSearch
|
|
553
|
+
|
|
554
|
+
strategy = RandomSearch(space, seed=42)
|
|
555
|
+
```
|
|
556
|
+
|
|
557
|
+
Candidate order is reproducible for the same seed, space, and reservation history. Resume replays
|
|
558
|
+
and verifies the random generator state.
|
|
559
|
+
|
|
560
|
+
### LLMSearch
|
|
561
|
+
|
|
562
|
+
```python
|
|
563
|
+
from tunarag import LLMSearch
|
|
564
|
+
|
|
565
|
+
|
|
566
|
+
class SuggestionProvider:
|
|
567
|
+
def __init__(self, client):
|
|
568
|
+
self.client = client
|
|
569
|
+
|
|
570
|
+
async def suggest(self, request):
|
|
571
|
+
return await self.client.generate_json(
|
|
572
|
+
schema=request.search_space,
|
|
573
|
+
context={
|
|
574
|
+
"suggestion_number": request.suggestion_number,
|
|
575
|
+
"attempt_number": request.attempt_number,
|
|
576
|
+
"observations": [
|
|
577
|
+
{
|
|
578
|
+
"candidate": dict(item.candidate.parameters),
|
|
579
|
+
"metrics": [
|
|
580
|
+
{"name": metric.name, "value": metric.value} for metric in item.metrics
|
|
581
|
+
],
|
|
582
|
+
}
|
|
583
|
+
for item in request.observations
|
|
584
|
+
],
|
|
585
|
+
},
|
|
586
|
+
)
|
|
587
|
+
|
|
588
|
+
|
|
589
|
+
strategy = LLMSearch(
|
|
590
|
+
space,
|
|
591
|
+
SuggestionProvider(llm_client),
|
|
592
|
+
max_suggestion_attempts=3,
|
|
593
|
+
fallback_seed=42,
|
|
594
|
+
fallback=True,
|
|
595
|
+
)
|
|
596
|
+
```
|
|
597
|
+
|
|
598
|
+
Proposals are strictly validated. Provider exceptions and invalid values consume a bounded attempt.
|
|
599
|
+
After exhaustion, deterministic random fallback runs when enabled; otherwise `SearchSpaceError` is
|
|
600
|
+
raised. Provider-facing observations are redacted.
|
|
601
|
+
|
|
602
|
+
## Evaluators and RAGAS
|
|
603
|
+
|
|
604
|
+
### Simple and custom evaluators
|
|
605
|
+
|
|
606
|
+
```python
|
|
607
|
+
from tunarag import MetricEvaluator, MetricValue, ValueStatus
|
|
608
|
+
|
|
609
|
+
|
|
610
|
+
def exact_match(output, example):
|
|
611
|
+
if example.reference_answer is None:
|
|
612
|
+
return None
|
|
613
|
+
return float(output["answer"].strip() == example.reference_answer.strip())
|
|
614
|
+
|
|
615
|
+
|
|
616
|
+
exact_match_evaluator = MetricEvaluator("exact_match", exact_match)
|
|
617
|
+
|
|
618
|
+
|
|
619
|
+
class RetrievalEvaluator:
|
|
620
|
+
async def evaluate(self, output, example):
|
|
621
|
+
retrieved = set(output["document_ids"])
|
|
622
|
+
relevant = set(example.relevant_document_ids)
|
|
623
|
+
recall = len(retrieved & relevant) / len(relevant) if relevant else None
|
|
624
|
+
return (
|
|
625
|
+
MetricValue(
|
|
626
|
+
"retrieval_recall",
|
|
627
|
+
recall,
|
|
628
|
+
ValueStatus.EXACT,
|
|
629
|
+
coverage=1.0 if relevant else 0.0,
|
|
630
|
+
),
|
|
631
|
+
MetricValue("context_count", float(len(retrieved))),
|
|
632
|
+
)
|
|
633
|
+
|
|
634
|
+
def fingerprint(self):
|
|
635
|
+
return {"evaluator": "retrieval", "version": 1}
|
|
636
|
+
```
|
|
637
|
+
|
|
638
|
+
`MetricEvaluator` accepts sync or async scorers. `None`, NaN, and infinity become unavailable;
|
|
639
|
+
nonnumeric values raise `EvaluationError`. Coverage must be `[0, 1]`. `tunarag.objective` is a
|
|
640
|
+
reserved metric name.
|
|
641
|
+
|
|
642
|
+
### RAGAS
|
|
643
|
+
|
|
644
|
+
```python
|
|
645
|
+
from ragas.metrics import AnswerRelevancy, Faithfulness
|
|
646
|
+
from tunarag import RagasEvaluator
|
|
647
|
+
|
|
648
|
+
ragas_evaluator = RagasEvaluator(
|
|
649
|
+
metrics=(Faithfulness(), AnswerRelevancy()),
|
|
650
|
+
sample_mapper=lambda output, example: {
|
|
651
|
+
"user_input": example.query,
|
|
652
|
+
"response": output["answer"],
|
|
653
|
+
"retrieved_contexts": output["contexts"],
|
|
654
|
+
"reference": example.reference_answer,
|
|
655
|
+
},
|
|
656
|
+
metric_prefix="ragas.",
|
|
657
|
+
llm=evaluation_llm,
|
|
658
|
+
embeddings=evaluation_embeddings,
|
|
659
|
+
raise_exceptions=False,
|
|
660
|
+
)
|
|
661
|
+
```
|
|
662
|
+
|
|
663
|
+
TunaRAG constructs one RAGAS `SingleTurnSample` and calls the async evaluation API. Names receive
|
|
664
|
+
the configured prefix. Mapping and RAGAS failures become `EvaluationError`. Credentials and provider
|
|
665
|
+
configuration remain owned by your RAGAS components.
|
|
666
|
+
|
|
667
|
+
## Objectives
|
|
668
|
+
|
|
669
|
+
```python
|
|
670
|
+
from tunarag import Objective, ObjectiveDirection, ObjectiveTerm
|
|
671
|
+
|
|
672
|
+
objective = Objective(
|
|
673
|
+
(
|
|
674
|
+
ObjectiveTerm("answer_quality", weight=0.70, minimum=0.0, maximum=1.0),
|
|
675
|
+
ObjectiveTerm(
|
|
676
|
+
"cost_usd",
|
|
677
|
+
weight=0.20,
|
|
678
|
+
direction=ObjectiveDirection.MINIMIZE,
|
|
679
|
+
minimum=0.0,
|
|
680
|
+
maximum=0.05,
|
|
681
|
+
),
|
|
682
|
+
ObjectiveTerm(
|
|
683
|
+
"latency_seconds",
|
|
684
|
+
weight=0.10,
|
|
685
|
+
direction=ObjectiveDirection.MINIMIZE,
|
|
686
|
+
minimum=0.0,
|
|
687
|
+
maximum=5.0,
|
|
688
|
+
),
|
|
689
|
+
)
|
|
690
|
+
)
|
|
691
|
+
```
|
|
692
|
+
|
|
693
|
+
Bounded metrics are normalized and clamped to `[0, 1]`; bounded minimize terms become
|
|
694
|
+
`1 - normalized`. Unbounded maximize terms use raw values and unbounded minimize terms use negative
|
|
695
|
+
values. Weights are contributions and are not automatically normalized. A missing or unavailable
|
|
696
|
+
required metric makes the objective unavailable.
|
|
697
|
+
|
|
698
|
+
Usage does not automatically become an objective metric. Emit `cost_usd` or `latency_seconds` from
|
|
699
|
+
an evaluator when you want those values in the objective.
|
|
700
|
+
|
|
701
|
+
## Configuration
|
|
702
|
+
|
|
703
|
+
```python
|
|
704
|
+
from tunarag import Settings
|
|
705
|
+
|
|
706
|
+
settings = Settings.resolve(
|
|
707
|
+
config_path="tunarag.toml",
|
|
708
|
+
overrides={"max_trials": 50, "concurrency": 4},
|
|
709
|
+
)
|
|
710
|
+
```
|
|
711
|
+
|
|
712
|
+
Precedence is defaults < TOML `[tunarag]` < environment < overrides. `.env` files are not loaded.
|
|
713
|
+
|
|
714
|
+
```toml
|
|
715
|
+
[tunarag]
|
|
716
|
+
home = ".tunarag"
|
|
717
|
+
db_path = ".tunarag/experiments.db"
|
|
718
|
+
cache_dir = ".tunarag/cache"
|
|
719
|
+
artifact_dir = ".tunarag/artifacts"
|
|
720
|
+
max_trials = 25
|
|
721
|
+
concurrency = 4
|
|
722
|
+
sample_concurrency = 8
|
|
723
|
+
seed = 0
|
|
724
|
+
max_retries = 2
|
|
725
|
+
sample_timeout_seconds = 60.0
|
|
726
|
+
trial_timeout_seconds = 300.0
|
|
727
|
+
cache_policy = "read_write"
|
|
728
|
+
```
|
|
729
|
+
|
|
730
|
+
| Environment variable | Default / meaning |
|
|
731
|
+
|---|---|
|
|
732
|
+
| `TUNARAG_HOME` | Platform application-data directory |
|
|
733
|
+
| `TUNARAG_DB_PATH` | `<home>/experiments.db` |
|
|
734
|
+
| `TUNARAG_CACHE_DIR` | `<home>/cache` |
|
|
735
|
+
| `TUNARAG_ARTIFACT_DIR` | `<home>/artifacts` |
|
|
736
|
+
| `TUNARAG_MAX_TRIALS` | `25` |
|
|
737
|
+
| `TUNARAG_CONCURRENCY` | `4` trials |
|
|
738
|
+
| `TUNARAG_SAMPLE_CONCURRENCY` | `8` samples per trial |
|
|
739
|
+
| `TUNARAG_SEED` | `0` |
|
|
740
|
+
| `TUNARAG_MAX_RETRIES` | `2` |
|
|
741
|
+
| `TUNARAG_SAMPLE_TIMEOUT_SECONDS` | `60.0` |
|
|
742
|
+
| `TUNARAG_TRIAL_TIMEOUT_SECONDS` | `300.0` |
|
|
743
|
+
| `TUNARAG_CACHE_POLICY` | `read_write` |
|
|
744
|
+
|
|
745
|
+
Cache policies: `off`, `read_only`, `write_only`, `read_write`. Settings contain no provider keys.
|
|
746
|
+
|
|
747
|
+
## Running optimization
|
|
748
|
+
|
|
749
|
+
Use async in async programs:
|
|
750
|
+
|
|
751
|
+
```python
|
|
752
|
+
result = await optimizer.aoptimize(dataset, study_name="support-rag-v4")
|
|
753
|
+
```
|
|
754
|
+
|
|
755
|
+
Use the sync wrapper only when no event loop is active:
|
|
756
|
+
|
|
757
|
+
```python
|
|
758
|
+
result = optimizer.optimize(dataset, study_name="support-rag-v4")
|
|
759
|
+
```
|
|
760
|
+
|
|
761
|
+
Calling sync methods inside an event loop raises `SyncInAsyncContextError`.
|
|
762
|
+
|
|
763
|
+
Durable composition:
|
|
764
|
+
|
|
765
|
+
```python
|
|
766
|
+
from pathlib import Path
|
|
767
|
+
from tunarag import (
|
|
768
|
+
MaxFailures,
|
|
769
|
+
MaxTrials,
|
|
770
|
+
NoImprovement,
|
|
771
|
+
Optimizer,
|
|
772
|
+
RetryPolicy,
|
|
773
|
+
SQLiteCache,
|
|
774
|
+
SQLiteStore,
|
|
775
|
+
Settings,
|
|
776
|
+
TargetScore,
|
|
777
|
+
)
|
|
778
|
+
|
|
779
|
+
settings = Settings.resolve(
|
|
780
|
+
overrides={
|
|
781
|
+
"home": Path(".tunarag"),
|
|
782
|
+
"db_path": Path(".tunarag/experiments.db"),
|
|
783
|
+
"cache_dir": Path(".tunarag/cache"),
|
|
784
|
+
"artifact_dir": Path(".tunarag/artifacts"),
|
|
785
|
+
"concurrency": 2,
|
|
786
|
+
"sample_concurrency": 8,
|
|
787
|
+
},
|
|
788
|
+
environ={},
|
|
789
|
+
)
|
|
790
|
+
settings.home.mkdir(parents=True, exist_ok=True)
|
|
791
|
+
settings.cache_dir.mkdir(parents=True, exist_ok=True)
|
|
792
|
+
store = SQLiteStore(settings.db_path)
|
|
793
|
+
cache = SQLiteCache(settings.cache_dir / "evaluations.db")
|
|
794
|
+
optimizer = Optimizer(
|
|
795
|
+
adapter=adapter,
|
|
796
|
+
strategy=strategy,
|
|
797
|
+
evaluators=(quality_evaluator, retrieval_evaluator),
|
|
798
|
+
objective=objective,
|
|
799
|
+
store=store,
|
|
800
|
+
cache=cache,
|
|
801
|
+
settings=settings,
|
|
802
|
+
stop_conditions=(
|
|
803
|
+
MaxTrials(30),
|
|
804
|
+
TargetScore(0.90),
|
|
805
|
+
NoImprovement(patience=6, min_delta=0.01, warmup_trials=8),
|
|
806
|
+
MaxFailures(5),
|
|
807
|
+
),
|
|
808
|
+
retry_policy=RetryPolicy(
|
|
809
|
+
max_retries=2,
|
|
810
|
+
sample_timeout_seconds=45.0,
|
|
811
|
+
trial_timeout_seconds=300.0,
|
|
812
|
+
base_delay_seconds=0.5,
|
|
813
|
+
max_delay_seconds=10.0,
|
|
814
|
+
),
|
|
815
|
+
)
|
|
816
|
+
try:
|
|
817
|
+
result = await optimizer.aoptimize(dataset, study_name="support-rag-v4")
|
|
818
|
+
finally:
|
|
819
|
+
cache.close()
|
|
820
|
+
store.close()
|
|
821
|
+
```
|
|
822
|
+
|
|
823
|
+
When you pass a store, you own it. `optimizer.close()` closes only its internally created store.
|
|
824
|
+
|
|
825
|
+
## Stopping, retries, and concurrency
|
|
826
|
+
|
|
827
|
+
Stop conditions are OR-composed:
|
|
828
|
+
|
|
829
|
+
| Condition | Stops when |
|
|
830
|
+
|---|---|
|
|
831
|
+
| `MaxTrials(limit)` | Completed trial count reaches the limit |
|
|
832
|
+
| `CostBudget(limit, count_estimated=True)` | Exact plus optional estimated cost reaches limit |
|
|
833
|
+
| `TimeBudget(seconds)` | Wall-clock duration reaches limit |
|
|
834
|
+
| `TargetScore(target)` | Any benefit-oriented objective reaches target |
|
|
835
|
+
| `MaxFailures(limit, consecutive=False)` | Total or trailing consecutive failures reach limit |
|
|
836
|
+
| `NoImprovement(patience, min_delta=0, warmup_trials=1)` | Best score stops improving |
|
|
837
|
+
|
|
838
|
+
Always include a hard bound. Conditions are checked between trial batches, so active trials can
|
|
839
|
+
finish after a cost, time, score, failure, or stagnation threshold is reached. Use `concurrency=1`
|
|
840
|
+
for strict between-trial stopping.
|
|
841
|
+
|
|
842
|
+
```python
|
|
843
|
+
from tunarag import RetryPolicy
|
|
844
|
+
|
|
845
|
+
policy = RetryPolicy(
|
|
846
|
+
max_retries=3,
|
|
847
|
+
sample_timeout_seconds=30.0,
|
|
848
|
+
trial_timeout_seconds=180.0,
|
|
849
|
+
base_delay_seconds=0.25,
|
|
850
|
+
max_delay_seconds=8.0,
|
|
851
|
+
)
|
|
852
|
+
```
|
|
853
|
+
|
|
854
|
+
Three retries means four total attempts. Timeouts and connection errors are retryable. A
|
|
855
|
+
`TunaRAGError` retries only when `retryable=True`; other exceptions fail that trial. Backoff is
|
|
856
|
+
deterministic and capped. Attempt usage is append-only, including failed attempts.
|
|
857
|
+
|
|
858
|
+
`concurrency` bounds simultaneous trials; `sample_concurrency` bounds samples inside each trial.
|
|
859
|
+
Their product approximates maximum concurrent adapter calls. Completion may be concurrent, but
|
|
860
|
+
results and strategy observations preserve trial-sequence order.
|
|
861
|
+
|
|
862
|
+
## Caching
|
|
863
|
+
|
|
864
|
+
```python
|
|
865
|
+
from pathlib import Path
|
|
866
|
+
|
|
867
|
+
from tunarag import CachePolicy, SQLiteCache, Settings
|
|
868
|
+
|
|
869
|
+
Path(".tunarag").mkdir(parents=True, exist_ok=True)
|
|
870
|
+
cache = SQLiteCache(".tunarag/evaluations.db")
|
|
871
|
+
settings = Settings.resolve(overrides={"cache_policy": CachePolicy.READ_WRITE}, environ={})
|
|
872
|
+
optimizer = Optimizer(
|
|
873
|
+
adapter=adapter,
|
|
874
|
+
strategy=strategy,
|
|
875
|
+
evaluators=evaluators,
|
|
876
|
+
objective=objective,
|
|
877
|
+
cache=cache,
|
|
878
|
+
settings=settings,
|
|
879
|
+
)
|
|
880
|
+
```
|
|
881
|
+
|
|
882
|
+
Identity includes dataset version, candidate, adapter, and evaluators. A hit skips adapter and
|
|
883
|
+
evaluator execution but still calculates the objective and records the trial.
|
|
884
|
+
|
|
885
|
+
Direct API:
|
|
886
|
+
|
|
887
|
+
```python
|
|
888
|
+
from tunarag import CacheKey, CachePrivacy, SQLiteCache
|
|
889
|
+
|
|
890
|
+
cache = SQLiteCache("cache.db")
|
|
891
|
+
key = CacheKey.from_parts("my-feature", {"query": "hello"}, format_version=1)
|
|
892
|
+
cache.put(key, b'{"answer":"world"}', privacy=CachePrivacy.SENSITIVE, ttl_seconds=3600)
|
|
893
|
+
lookup = cache.get(key)
|
|
894
|
+
if lookup.is_hit:
|
|
895
|
+
payload = lookup.payload
|
|
896
|
+
cache.prune()
|
|
897
|
+
cache.close()
|
|
898
|
+
```
|
|
899
|
+
|
|
900
|
+
Statuses are `hit`, `miss`, `expired`, `corrupt`, and `quarantined`. Checksums detect corruption and
|
|
901
|
+
bad rows are quarantined. `prune(include_quarantined=True)` can remove quarantined rows. Privacy
|
|
902
|
+
labels are metadata, not encryption.
|
|
903
|
+
|
|
904
|
+
## SQLite persistence and resume
|
|
905
|
+
|
|
906
|
+
```python
|
|
907
|
+
from pathlib import Path
|
|
908
|
+
|
|
909
|
+
from tunarag import SQLiteStore
|
|
910
|
+
|
|
911
|
+
Path(".tunarag").mkdir(parents=True, exist_ok=True)
|
|
912
|
+
store = SQLiteStore(".tunarag/experiments.db")
|
|
913
|
+
optimizer = build_optimizer(store)
|
|
914
|
+
result = await optimizer.aoptimize(dataset, study_name="production-rag")
|
|
915
|
+
|
|
916
|
+
# Rebuild stateful components so replay begins from their initial state.
|
|
917
|
+
resumed_optimizer = build_optimizer(store)
|
|
918
|
+
resumed = await resumed_optimizer.aresume(dataset, study_id=result.study_id)
|
|
919
|
+
# Sync equivalent: resumed_optimizer.resume(dataset, study_id=result.study_id)
|
|
920
|
+
```
|
|
921
|
+
|
|
922
|
+
Resume validates semantic configuration and dataset identity, returns completed studies without
|
|
923
|
+
provider execution, abandons expired leases, replays committed strategy history, continues
|
|
924
|
+
unfinished work, and schedules new candidates. Unknown, failed, cancelled, or incompatible studies
|
|
925
|
+
raise `ResumeError`. An active lease produces a retryable error; retry later.
|
|
926
|
+
|
|
927
|
+
Read persisted state:
|
|
928
|
+
|
|
929
|
+
```python
|
|
930
|
+
study = store.get_study(result.study_id)
|
|
931
|
+
trials = store.list_trials(result.study_id)
|
|
932
|
+
events = store.list_events(result.study_id, after_sequence=-1)
|
|
933
|
+
attempts = store.list_attempts(trials[0].id)
|
|
934
|
+
metrics = store.list_metrics(trials[0].id)
|
|
935
|
+
usage = store.list_usage(trials[0].id)
|
|
936
|
+
```
|
|
937
|
+
|
|
938
|
+
Normally let `Optimizer` own lifecycle mutations and use store read methods for inspection.
|
|
939
|
+
|
|
940
|
+
## Callbacks and MLflow
|
|
941
|
+
|
|
942
|
+
Callbacks receive ordered durable events after commit:
|
|
943
|
+
|
|
944
|
+
```python
|
|
945
|
+
class AuditCallback:
|
|
946
|
+
async def on_event(self, event):
|
|
947
|
+
print(event.sequence, event.type, event.study_id, dict(event.payload))
|
|
948
|
+
|
|
949
|
+
|
|
950
|
+
optimizer = Optimizer(
|
|
951
|
+
adapter=adapter,
|
|
952
|
+
strategy=strategy,
|
|
953
|
+
evaluators=evaluators,
|
|
954
|
+
objective=objective,
|
|
955
|
+
callbacks=(AuditCallback(),),
|
|
956
|
+
)
|
|
957
|
+
```
|
|
958
|
+
|
|
959
|
+
Events cover study, trial, and attempt lifecycles plus callback failures. Handle unknown future event
|
|
960
|
+
types defensively. Callback failures never roll back committed work; TunaRAG records them as durable
|
|
961
|
+
events. Make callbacks idempotent.
|
|
962
|
+
|
|
963
|
+
```mermaid
|
|
964
|
+
sequenceDiagram
|
|
965
|
+
participant O as Optimizer
|
|
966
|
+
participant S as SQLite
|
|
967
|
+
participant C as Callback
|
|
968
|
+
O->>S: Commit terminal trial and event
|
|
969
|
+
S-->>O: Committed
|
|
970
|
+
O->>C: on_event(event)
|
|
971
|
+
alt callback fails
|
|
972
|
+
O->>S: Append callback.failed
|
|
973
|
+
end
|
|
974
|
+
```
|
|
975
|
+
|
|
976
|
+
MLflow is an optional post-commit projection; SQLite remains authoritative:
|
|
977
|
+
|
|
978
|
+
```python
|
|
979
|
+
from tunarag.integrations import MLflowCallback
|
|
980
|
+
|
|
981
|
+
callback = MLflowCallback(
|
|
982
|
+
store,
|
|
983
|
+
experiment_id="42",
|
|
984
|
+
tracking_uri="https://mlflow.example.com",
|
|
985
|
+
tags={"team": "search", "environment": "staging"},
|
|
986
|
+
)
|
|
987
|
+
optimizer = Optimizer(
|
|
988
|
+
adapter=adapter,
|
|
989
|
+
strategy=strategy,
|
|
990
|
+
evaluators=evaluators,
|
|
991
|
+
objective=objective,
|
|
992
|
+
store=store,
|
|
993
|
+
callbacks=(callback,),
|
|
994
|
+
)
|
|
995
|
+
```
|
|
996
|
+
|
|
997
|
+
One MLflow run is created per study. Candidate parameters, committed metrics, statuses, and event
|
|
998
|
+
position are mirrored. An MLflow outage does not invalidate local results. Automatic external
|
|
999
|
+
reconciliation is not included in V0.2.
|
|
1000
|
+
|
|
1001
|
+
## Results and reports
|
|
1002
|
+
|
|
1003
|
+
`OptimizationResult` exposes:
|
|
1004
|
+
|
|
1005
|
+
```python
|
|
1006
|
+
result.study_id
|
|
1007
|
+
result.study_status
|
|
1008
|
+
result.stop_reason
|
|
1009
|
+
result.trials
|
|
1010
|
+
result.successful_trials
|
|
1011
|
+
result.failed_trials
|
|
1012
|
+
result.best_trial
|
|
1013
|
+
result.best_candidate
|
|
1014
|
+
result.ranked_trials
|
|
1015
|
+
result.metric_names
|
|
1016
|
+
result.usage
|
|
1017
|
+
```
|
|
1018
|
+
|
|
1019
|
+
Rank by an individual metric:
|
|
1020
|
+
|
|
1021
|
+
```python
|
|
1022
|
+
from tunarag import ObjectiveDirection
|
|
1023
|
+
|
|
1024
|
+
quality_leaders = result.metric_leaders("answer_quality")
|
|
1025
|
+
latency_leaders = result.metric_leaders("latency_seconds", direction=ObjectiveDirection.MINIMIZE)
|
|
1026
|
+
```
|
|
1027
|
+
|
|
1028
|
+
Inspect and export:
|
|
1029
|
+
|
|
1030
|
+
```python
|
|
1031
|
+
from pathlib import Path
|
|
1032
|
+
|
|
1033
|
+
if result.ranked_trials:
|
|
1034
|
+
trial = result.ranked_trials[0]
|
|
1035
|
+
print(trial.candidate.parameters)
|
|
1036
|
+
print(trial.objective.score)
|
|
1037
|
+
print(trial.metric("answer_quality"))
|
|
1038
|
+
print(trial.attempts, trial.cached, trial.error)
|
|
1039
|
+
|
|
1040
|
+
print(result.to_json())
|
|
1041
|
+
Path("artifacts").mkdir(parents=True, exist_ok=True)
|
|
1042
|
+
result.write_json("artifacts/result.json")
|
|
1043
|
+
result.write_csv("artifacts/trials.csv")
|
|
1044
|
+
```
|
|
1045
|
+
|
|
1046
|
+
Reports are deterministic and redacted. Usage separates exact and estimated cost. A completed result
|
|
1047
|
+
may have no best candidate if every trial failed or every objective was unavailable.
|
|
1048
|
+
|
|
1049
|
+
## Errors, privacy, and secrets
|
|
1050
|
+
|
|
1051
|
+
```python
|
|
1052
|
+
from tunarag import TunaRAGError
|
|
1053
|
+
|
|
1054
|
+
try:
|
|
1055
|
+
result = optimizer.optimize(dataset, study_name="example")
|
|
1056
|
+
except TunaRAGError as error:
|
|
1057
|
+
diagnostic = error.as_dict()
|
|
1058
|
+
print(diagnostic["code"], diagnostic["stage"], diagnostic["retryable"])
|
|
1059
|
+
```
|
|
1060
|
+
|
|
1061
|
+
Public subclasses include `AdapterError`, `BudgetError`, `ConfigurationError`, `DatasetError`,
|
|
1062
|
+
`EvaluationError`, `MissingOptionalDependencyError`, `ResumeError`, `SearchSpaceError`, `StoreError`,
|
|
1063
|
+
and `SyncInAsyncContextError`.
|
|
1064
|
+
|
|
1065
|
+
```python
|
|
1066
|
+
from tunarag import Secret
|
|
1067
|
+
|
|
1068
|
+
credential = Secret("do-not-log-this")
|
|
1069
|
+
print(credential) # ***redacted***
|
|
1070
|
+
print(repr(credential)) # Secret(***redacted***)
|
|
1071
|
+
```
|
|
1072
|
+
|
|
1073
|
+
TunaRAG redacts `Secret` and common sensitive keys from diagnostics, fingerprints, events, and
|
|
1074
|
+
reports. This is defense in depth, not a secret manager. SQLite/cache files can contain prompts and
|
|
1075
|
+
metadata; privacy labels do not encrypt. LLMSearch, synthetic generation, RAGAS, and callbacks may
|
|
1076
|
+
send data to external providers. Apply filesystem permissions, encrypted storage, data minimization,
|
|
1077
|
+
and provider-specific retention controls.
|
|
1078
|
+
|
|
1079
|
+
## Extension contracts
|
|
1080
|
+
|
|
1081
|
+
Protocols are structural; inheritance is not required.
|
|
1082
|
+
|
|
1083
|
+
| Contract | Required behavior |
|
|
1084
|
+
|---|---|
|
|
1085
|
+
| `RAGAdapter` | `async run(candidate, example) -> (output, usage)` |
|
|
1086
|
+
| `Dataset` | `.version` and async iteration |
|
|
1087
|
+
| `Evaluator` | `async evaluate(output, example) -> metrics` |
|
|
1088
|
+
| `SearchStrategy` | `async suggest()` and `async observe(...)` |
|
|
1089
|
+
| `ReplayableSearchStrategy` | `async replay(...)` for resume |
|
|
1090
|
+
| `Fingerprintable` | stable secret-safe `fingerprint()` |
|
|
1091
|
+
| `EventCallback` | `async on_event(event)` |
|
|
1092
|
+
| `Cache` | synchronous `get()` and `put()` |
|
|
1093
|
+
|
|
1094
|
+
Custom search-space parameters implement `sample(generator)`, `accepts(value)`, and `describe()`:
|
|
1095
|
+
|
|
1096
|
+
```python
|
|
1097
|
+
class EvenInteger:
|
|
1098
|
+
def __init__(self, minimum, maximum):
|
|
1099
|
+
self.minimum = minimum
|
|
1100
|
+
self.maximum = maximum
|
|
1101
|
+
|
|
1102
|
+
def sample(self, generator):
|
|
1103
|
+
values = tuple(range(self.minimum + self.minimum % 2, self.maximum + 1, 2))
|
|
1104
|
+
return generator.choice(values)
|
|
1105
|
+
|
|
1106
|
+
def accepts(self, value):
|
|
1107
|
+
return (
|
|
1108
|
+
isinstance(value, int)
|
|
1109
|
+
and not isinstance(value, bool)
|
|
1110
|
+
and self.minimum <= value <= self.maximum
|
|
1111
|
+
and value % 2 == 0
|
|
1112
|
+
)
|
|
1113
|
+
|
|
1114
|
+
def describe(self):
|
|
1115
|
+
return {
|
|
1116
|
+
"type": "integer",
|
|
1117
|
+
"minimum": self.minimum,
|
|
1118
|
+
"maximum": self.maximum,
|
|
1119
|
+
"multiple_of": 2,
|
|
1120
|
+
}
|
|
1121
|
+
```
|
|
1122
|
+
|
|
1123
|
+
Provider descriptions and component fingerprints must not expose secrets.
|
|
1124
|
+
|
|
1125
|
+
## Production checklist
|
|
1126
|
+
|
|
1127
|
+
- Use a file-backed `SQLiteStore` and back it up.
|
|
1128
|
+
- Keep a human-reviewed test split outside optimization.
|
|
1129
|
+
- Fingerprint custom adapters and evaluators.
|
|
1130
|
+
- Include a hard stopping bound and conservative concurrency.
|
|
1131
|
+
- Track exact usage when available and label estimates.
|
|
1132
|
+
- Configure sample and trial timeouts.
|
|
1133
|
+
- Use deterministic seeds for reproducibility.
|
|
1134
|
+
- Protect database and cache files; keep credentials outside TunaRAG objects.
|
|
1135
|
+
- Make callbacks idempotent and monitor callback-failure events.
|
|
1136
|
+
- Retain study IDs and exported reports with releases.
|
|
1137
|
+
- Test resume behavior before expensive studies.
|
|
1138
|
+
- Re-evaluate the winner on validation and test sets.
|
|
1139
|
+
|
|
1140
|
+
## Limitations and glossary
|
|
1141
|
+
|
|
1142
|
+
V0.2 does not provide a hosted service/dashboard, automatic RAG construction, vector-database
|
|
1143
|
+
provisioning, distributed workers, Pareto optimization, Bayesian/evolutionary search, encrypted
|
|
1144
|
+
payloads, MLflow reconciliation, provider pricing management, or a TypeScript implementation.
|
|
1145
|
+
LangChain and LangGraph are official integrations; other frameworks use `RAGAdapter`.
|
|
1146
|
+
|
|
1147
|
+
| Term | Meaning |
|
|
1148
|
+
|---|---|
|
|
1149
|
+
| Adapter | Runs the user-owned RAG for one candidate and example |
|
|
1150
|
+
| Attempt | One trial execution, including retries |
|
|
1151
|
+
| Candidate | Complete parameter mapping proposed by a strategy |
|
|
1152
|
+
| Dataset version | Identity used for cache and resume compatibility |
|
|
1153
|
+
| Evaluator | Converts one output/example pair into metrics |
|
|
1154
|
+
| Fingerprint | Secret-safe semantic component identity |
|
|
1155
|
+
| Objective | Converts trial metrics into one benefit score |
|
|
1156
|
+
| Sample | One candidate evaluated against one example |
|
|
1157
|
+
| Strategy | Proposes candidates and observes outcomes |
|
|
1158
|
+
| Study | Durable optimization run containing trials |
|
|
1159
|
+
| Trial | One candidate evaluated over the whole dataset |
|
|
1160
|
+
| Usage | Tokens, cost, latency, confidence, and pricing provenance |
|
|
1161
|
+
|
|
1162
|
+
For deeper contracts and architecture, see [`docs/api-specification.md`](docs/api-specification.md),
|
|
1163
|
+
[`docs/system-architecture-document.md`](docs/system-architecture-document.md), and
|
|
1164
|
+
[`docs/technical-design-document.md`](docs/technical-design-document.md).
|