TriCacheLLM-MMA 0.1.0__py3-none-any.whl
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- portable_cache_Ai/portable_cache_rerankAi.py +132 -0
- portable_cache_bgWorkers/portable_cache_celery_conf.py +51 -0
- portable_cache_bgWorkers/portable_cache_workers.py +163 -0
- portable_cache_schemas/portable_cache_dbBase.py +3 -0
- portable_cache_schemas/portable_cache_dbConf.py +147 -0
- portable_cache_schemas/portable_cache_schemas.py +8 -0
- portable_cache_utils/portable_cache_embedding_model.py +7 -0
- portable_cache_utils/protable_cache_DynamicEnv_maker.py +152 -0
- tricachellm_mma-0.1.0.dist-info/METADATA +949 -0
- tricachellm_mma-0.1.0.dist-info/RECORD +13 -0
- tricachellm_mma-0.1.0.dist-info/WHEEL +5 -0
- tricachellm_mma-0.1.0.dist-info/licenses/LICENSE +73 -0
- tricachellm_mma-0.1.0.dist-info/top_level.txt +4 -0
|
@@ -0,0 +1,949 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: TriCacheLLM_MMA
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: A robust three-tier LLM-aware caching SDK using Redis, ChromaDB, and automatic SQLite bootstrap.
|
|
5
|
+
Author-email: Mohib Ashfaq Butt <inboxmohib@gmail.com>
|
|
6
|
+
Project-URL: Homepage, https://github.com/mohib-ash/TriCacheLLM_MMA
|
|
7
|
+
Project-URL: Repository, https://github.com/mohib-ash/TriCacheLLM_MMA
|
|
8
|
+
Project-URL: Issues, https://github.com/mohib-ash/TriCacheLLM_MMA/issues
|
|
9
|
+
Keywords: llm,cache,caching,semantic-cache,three-tier-cache,redis,chromadb,rag,rag-cache,ai,nlp
|
|
10
|
+
Classifier: Development Status :: 3 - Alpha
|
|
11
|
+
Classifier: Intended Audience :: Developers
|
|
12
|
+
Classifier: Intended Audience :: Information Technology
|
|
13
|
+
Classifier: Programming Language :: Python :: 3
|
|
14
|
+
Classifier: Programming Language :: Python :: 3 :: Only
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
18
|
+
Classifier: Operating System :: OS Independent
|
|
19
|
+
Classifier: Topic :: Software Development :: Libraries
|
|
20
|
+
Classifier: Topic :: Software Development :: Libraries :: Python Modules
|
|
21
|
+
Classifier: Topic :: Database
|
|
22
|
+
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
|
|
23
|
+
Requires-Python: >=3.11
|
|
24
|
+
Description-Content-Type: text/markdown
|
|
25
|
+
License-File: LICENSE
|
|
26
|
+
Requires-Dist: SQLAlchemy>=2.0.0
|
|
27
|
+
Requires-Dist: aiosqlite>=0.19.0
|
|
28
|
+
Requires-Dist: redis>=7.0.0
|
|
29
|
+
Requires-Dist: chromadb>=1.5.0
|
|
30
|
+
Requires-Dist: langchain-chroma>=0.2.0
|
|
31
|
+
Requires-Dist: langchain-core>=0.3.0
|
|
32
|
+
Requires-Dist: langchain-huggingface>=0.3.0
|
|
33
|
+
Requires-Dist: sentence-transformers>=5.0.0
|
|
34
|
+
Requires-Dist: cohere>=5.0.0
|
|
35
|
+
Requires-Dist: celery>=5.0.0
|
|
36
|
+
Requires-Dist: pydantic>=2.0.0
|
|
37
|
+
Requires-Dist: pydantic-settings>=2.0.0
|
|
38
|
+
Requires-Dist: numpy>=2.0.0
|
|
39
|
+
Requires-Dist: python-dotenv>=1.0.0
|
|
40
|
+
Dynamic: license-file
|
|
41
|
+
|
|
42
|
+
# TriCacheLLM_MMA
|
|
43
|
+
|
|
44
|
+
## Portable 3-Tier Semantic Cache for LLM Applications
|
|
45
|
+
|
|
46
|
+
**TriCacheLLM_MMA** is a portable asynchronous caching system designed for Ai based ***web applications*** that repeatedly ask LLMs similar or identical questions.
|
|
47
|
+
|
|
48
|
+
It combines:
|
|
49
|
+
|
|
50
|
+
* **Exact Redis caching**
|
|
51
|
+
* **Semantic Redis HNSW search**
|
|
52
|
+
* **Persistent vector-database caching**
|
|
53
|
+
* **Cohere reranking**
|
|
54
|
+
* **Celery background workers**
|
|
55
|
+
* **Per-user / per-tenant cache isolation**
|
|
56
|
+
* **SQLite-backed cache infrastructure state**
|
|
57
|
+
* **Automatic cache promotion between tiers**
|
|
58
|
+
|
|
59
|
+
The goal is simple:
|
|
60
|
+
|
|
61
|
+
> Avoid paying the cost of expensive LLM inference when an equivalent or sufficiently similar answer has already been generated.
|
|
62
|
+
|
|
63
|
+
The package is designed so that the consuming application only needs to initialize the cache infrastructure and use three simple operations:
|
|
64
|
+
|
|
65
|
+
```python
|
|
66
|
+
create_cache_system(...)
|
|
67
|
+
check_cache(...)
|
|
68
|
+
populate_cache(...)
|
|
69
|
+
check_tenant_creation_status(...)
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
The internal infrastructure handles the multi-tier lookup and promotion logic.
|
|
73
|
+
|
|
74
|
+
---
|
|
75
|
+
|
|
76
|
+
# Architecture
|
|
77
|
+
|
|
78
|
+
TriCacheLLM_MMA uses three cache tiers.
|
|
79
|
+
|
|
80
|
+
```text
|
|
81
|
+
USER QUESTION
|
|
82
|
+
│
|
|
83
|
+
▼
|
|
84
|
+
┌─────────────────────┐
|
|
85
|
+
│ TIER 1 │
|
|
86
|
+
│ Exact Redis KV │
|
|
87
|
+
│ │
|
|
88
|
+
│ Fastest lookup │
|
|
89
|
+
│ Exact question │
|
|
90
|
+
└──────────┬──────────┘
|
|
91
|
+
│ MISS
|
|
92
|
+
▼
|
|
93
|
+
┌─────────────────────┐
|
|
94
|
+
│ TIER 2 │
|
|
95
|
+
│ Redis HNSW Vector │
|
|
96
|
+
│ Search │
|
|
97
|
+
│ │
|
|
98
|
+
│ Semantic similarity │
|
|
99
|
+
└──────────┬──────────┘
|
|
100
|
+
│ MISS
|
|
101
|
+
▼
|
|
102
|
+
┌─────────────────────┐
|
|
103
|
+
│ TIER 3 │
|
|
104
|
+
│ Persistent VDB │
|
|
105
|
+
│ + Cohere │
|
|
106
|
+
│ Reranking │
|
|
107
|
+
└──────────┬──────────┘
|
|
108
|
+
│
|
|
109
|
+
▼
|
|
110
|
+
Persistent Vector Store
|
|
111
|
+
```
|
|
112
|
+
|
|
113
|
+
## Tier 1: Exact Redis
|
|
114
|
+
|
|
115
|
+
The first lookup is an exact question lookup.
|
|
116
|
+
|
|
117
|
+
This is the fastest path.
|
|
118
|
+
|
|
119
|
+
If the same question was previously cached, the system can immediately return the stored response without embedding generation, vector search, VDB access, or LLM inference.
|
|
120
|
+
|
|
121
|
+
```text
|
|
122
|
+
Question
|
|
123
|
+
↓
|
|
124
|
+
Exact Redis
|
|
125
|
+
↓
|
|
126
|
+
HIT
|
|
127
|
+
↓
|
|
128
|
+
Cached Answer
|
|
129
|
+
```
|
|
130
|
+
|
|
131
|
+
---
|
|
132
|
+
|
|
133
|
+
## Tier 2: Semantic Redis HNSW
|
|
134
|
+
|
|
135
|
+
If Tier 1 misses, the question is embedded and searched against a Redis HNSW vector index.
|
|
136
|
+
|
|
137
|
+
This allows small variations in wording to hit the cache.
|
|
138
|
+
|
|
139
|
+
For example:
|
|
140
|
+
|
|
141
|
+
```text
|
|
142
|
+
Cached:
|
|
143
|
+
How can I run Celery tasks asynchronously using Redis?
|
|
144
|
+
|
|
145
|
+
New:
|
|
146
|
+
How can I run Celery tasks asynchronously using Redis ?
|
|
147
|
+
```
|
|
148
|
+
|
|
149
|
+
The second question is not an exact string match, but it can still be recognized as semantically equivalent.
|
|
150
|
+
|
|
151
|
+
```text
|
|
152
|
+
T1 MISS
|
|
153
|
+
↓
|
|
154
|
+
T2 semantic search
|
|
155
|
+
↓
|
|
156
|
+
HIT
|
|
157
|
+
↓
|
|
158
|
+
Cached Answer
|
|
159
|
+
```
|
|
160
|
+
|
|
161
|
+
Tier 2 is intentionally optimized for fast semantic cache retrieval.
|
|
162
|
+
|
|
163
|
+
---
|
|
164
|
+
|
|
165
|
+
# Tier 3: Persistent VDB-Backed Cache
|
|
166
|
+
|
|
167
|
+
Tier 3 is the persistent cache layer.
|
|
168
|
+
|
|
169
|
+
The persistent vector database acts as the long-term backing store for cached Q&A entries.
|
|
170
|
+
|
|
171
|
+
When Tier 1 and Tier 2 miss, Tier 3 performs the persistent lookup.
|
|
172
|
+
|
|
173
|
+
The persistent lookup can use:
|
|
174
|
+
|
|
175
|
+
* vector similarity
|
|
176
|
+
* metadata
|
|
177
|
+
* provenance information
|
|
178
|
+
* additional filtering
|
|
179
|
+
* Cohere reranking
|
|
180
|
+
|
|
181
|
+
The exact persistent retrieval logic lives inside the VDB cache implementation.
|
|
182
|
+
|
|
183
|
+
A Tier 3 hit is not simply returned and forgotten.
|
|
184
|
+
|
|
185
|
+
The response is promoted upward:
|
|
186
|
+
|
|
187
|
+
```text
|
|
188
|
+
Tier 3 / VDB HIT
|
|
189
|
+
│
|
|
190
|
+
▼
|
|
191
|
+
Populate Tier 2
|
|
192
|
+
│
|
|
193
|
+
▼
|
|
194
|
+
Populate Tier 1
|
|
195
|
+
│
|
|
196
|
+
▼
|
|
197
|
+
Return cached response
|
|
198
|
+
```
|
|
199
|
+
|
|
200
|
+
This means frequently accessed answers naturally migrate toward the faster cache tiers.
|
|
201
|
+
|
|
202
|
+
---
|
|
203
|
+
|
|
204
|
+
# Cache Promotion
|
|
205
|
+
|
|
206
|
+
One of the central design principles of TriCacheLLM_MMA is **upward cache promotion**.
|
|
207
|
+
|
|
208
|
+
```text
|
|
209
|
+
┌──────────────┐
|
|
210
|
+
│ TIER 1 │
|
|
211
|
+
│ Exact Redis │
|
|
212
|
+
└──────▲───────┘
|
|
213
|
+
│
|
|
214
|
+
│ promotion
|
|
215
|
+
│
|
|
216
|
+
┌──────┴───────┐
|
|
217
|
+
│ TIER 2 │
|
|
218
|
+
│ Redis HNSW │
|
|
219
|
+
└──────▲───────┘
|
|
220
|
+
│
|
|
221
|
+
│ promotion
|
|
222
|
+
│
|
|
223
|
+
┌──────┴───────┐
|
|
224
|
+
│ TIER 3 │
|
|
225
|
+
│ Persistent │
|
|
226
|
+
│ VDB-backed │
|
|
227
|
+
└──────────────┘
|
|
228
|
+
```
|
|
229
|
+
|
|
230
|
+
Examples:
|
|
231
|
+
|
|
232
|
+
### Exact hit
|
|
233
|
+
|
|
234
|
+
```text
|
|
235
|
+
T1 HIT
|
|
236
|
+
↓
|
|
237
|
+
Return immediately
|
|
238
|
+
```
|
|
239
|
+
|
|
240
|
+
### Semantic hit
|
|
241
|
+
|
|
242
|
+
```text
|
|
243
|
+
T1 MISS
|
|
244
|
+
↓
|
|
245
|
+
T2 HIT
|
|
246
|
+
↓
|
|
247
|
+
Return response
|
|
248
|
+
↓
|
|
249
|
+
Promote / fill T1
|
|
250
|
+
```
|
|
251
|
+
|
|
252
|
+
### Persistent hit
|
|
253
|
+
|
|
254
|
+
```text
|
|
255
|
+
T1 MISS
|
|
256
|
+
↓
|
|
257
|
+
T2 MISS
|
|
258
|
+
↓
|
|
259
|
+
T3 HIT
|
|
260
|
+
↓
|
|
261
|
+
Promote to T2
|
|
262
|
+
↓
|
|
263
|
+
Promote to T1
|
|
264
|
+
↓
|
|
265
|
+
Return response
|
|
266
|
+
```
|
|
267
|
+
|
|
268
|
+
### Complete miss
|
|
269
|
+
|
|
270
|
+
```text
|
|
271
|
+
T1 MISS
|
|
272
|
+
↓
|
|
273
|
+
T2 MISS
|
|
274
|
+
↓
|
|
275
|
+
T3 MISS
|
|
276
|
+
↓
|
|
277
|
+
Your LLM / AI pipeline
|
|
278
|
+
↓
|
|
279
|
+
populate_cache()
|
|
280
|
+
↓
|
|
281
|
+
Persistent cache seeded
|
|
282
|
+
```
|
|
283
|
+
|
|
284
|
+
---
|
|
285
|
+
|
|
286
|
+
# Important V1 Semantics
|
|
287
|
+
|
|
288
|
+
`populate_cache()` and `check_cache()` have intentionally different responsibilities.
|
|
289
|
+
|
|
290
|
+
|
|
291
|
+
|
|
292
|
+
`populate_cache()` seeds the persistent cache VDB.
|
|
293
|
+
|
|
294
|
+
It does **not** directly populate every cache tier.
|
|
295
|
+
|
|
296
|
+
```text
|
|
297
|
+
populate_cache()
|
|
298
|
+
↓
|
|
299
|
+
Persistent VDB
|
|
300
|
+
```
|
|
301
|
+
|
|
302
|
+
The next `check_cache()` can discover that entry through Tier 3 and promote it upward.
|
|
303
|
+
|
|
304
|
+
This separation keeps the cache population path simple while allowing `check_cache()` to control cache promotion.
|
|
305
|
+
|
|
306
|
+
---
|
|
307
|
+
|
|
308
|
+
# Multi-Tenant Architecture
|
|
309
|
+
|
|
310
|
+
The cache system supports multiple users / tenants.
|
|
311
|
+
|
|
312
|
+
Each consumer user receives an isolated persistent cache VDB.
|
|
313
|
+
|
|
314
|
+
Conceptually:
|
|
315
|
+
|
|
316
|
+
```text
|
|
317
|
+
Consumer Application
|
|
318
|
+
│
|
|
319
|
+
├── User 1
|
|
320
|
+
│ └── Cache VDB 1
|
|
321
|
+
│
|
|
322
|
+
├── User 2
|
|
323
|
+
│ └── Cache VDB 2
|
|
324
|
+
│
|
|
325
|
+
└── User N
|
|
326
|
+
└── Cache VDB N
|
|
327
|
+
```
|
|
328
|
+
|
|
329
|
+
The package itself does not own the consumer application's user table.
|
|
330
|
+
|
|
331
|
+
Instead, the consumer application provides the users.
|
|
332
|
+
|
|
333
|
+
The cache package maintains its own cache infrastructure state.
|
|
334
|
+
|
|
335
|
+
```text
|
|
336
|
+
Consumer Database
|
|
337
|
+
│
|
|
338
|
+
├── users
|
|
339
|
+
│
|
|
340
|
+
├── paths
|
|
341
|
+
│
|
|
342
|
+
└── cache_vdb_resources
|
|
343
|
+
```
|
|
344
|
+
|
|
345
|
+
`cache_vdb_resources.user_id` references the consumer application's:
|
|
346
|
+
|
|
347
|
+
```text
|
|
348
|
+
users.user_id
|
|
349
|
+
```
|
|
350
|
+
|
|
351
|
+
Therefore, the consumer application must provide a compatible `users` table with a unique `user_id` primary key.
|
|
352
|
+
If you don't want your web app to have Multiple Tenants then asside form user_id=0, keep providing same user_id.
|
|
353
|
+
For a complete integration example, see:
|
|
354
|
+
`details_and_examples.py`.
|
|
355
|
+
|
|
356
|
+
---
|
|
357
|
+
|
|
358
|
+
# Internal Database
|
|
359
|
+
|
|
360
|
+
TriCacheLLM_MMA uses SQLite for its internal cache infrastructure state.
|
|
361
|
+
|
|
362
|
+
The package tracks information such as:
|
|
363
|
+
|
|
364
|
+
* configured Redis location
|
|
365
|
+
* Chroma storage location
|
|
366
|
+
* registry database path
|
|
367
|
+
* Cohere configuration
|
|
368
|
+
* per-user VDB status
|
|
369
|
+
* VDB version
|
|
370
|
+
* VDB path
|
|
371
|
+
* VDB creation failures
|
|
372
|
+
|
|
373
|
+
The runtime state is represented by tables including:
|
|
374
|
+
|
|
375
|
+
```text
|
|
376
|
+
paths
|
|
377
|
+
cache_vdb_resources
|
|
378
|
+
```
|
|
379
|
+
|
|
380
|
+
The consumer application's own database remains separate from the cache system's infrastructure state.
|
|
381
|
+
|
|
382
|
+
---
|
|
383
|
+
|
|
384
|
+
# Requirements
|
|
385
|
+
|
|
386
|
+
## Python
|
|
387
|
+
|
|
388
|
+
Recommended:
|
|
389
|
+
|
|
390
|
+
```text
|
|
391
|
+
Python 3.11+
|
|
392
|
+
```
|
|
393
|
+
|
|
394
|
+
## Infrastructure
|
|
395
|
+
|
|
396
|
+
TriCacheLLM_MMA currently expects:
|
|
397
|
+
|
|
398
|
+
* Redis
|
|
399
|
+
* Celery
|
|
400
|
+
* SQLite
|
|
401
|
+
* ChromaDB
|
|
402
|
+
* Cohere API access
|
|
403
|
+
|
|
404
|
+
Your application can use any framework.
|
|
405
|
+
|
|
406
|
+
FastAPI is used in the included example because it provides a convenient demonstration.
|
|
407
|
+
|
|
408
|
+
---
|
|
409
|
+
|
|
410
|
+
# Installation
|
|
411
|
+
|
|
412
|
+
Install from PyPI:
|
|
413
|
+
|
|
414
|
+
```bash
|
|
415
|
+
pip install TriCacheLLM_MMA
|
|
416
|
+
```
|
|
417
|
+
|
|
418
|
+
Or install the development version directly from the repository:
|
|
419
|
+
|
|
420
|
+
```bash
|
|
421
|
+
pip install .
|
|
422
|
+
```
|
|
423
|
+
|
|
424
|
+
---
|
|
425
|
+
|
|
426
|
+
# Redis
|
|
427
|
+
|
|
428
|
+
Start a Redis server before using the cache.
|
|
429
|
+
|
|
430
|
+
The default example uses:
|
|
431
|
+
|
|
432
|
+
```text
|
|
433
|
+
redis://localhost:6379/0
|
|
434
|
+
```
|
|
435
|
+
|
|
436
|
+
You can provide your Redis URL through:
|
|
437
|
+
|
|
438
|
+
```python
|
|
439
|
+
redis_url="your-redis-url"
|
|
440
|
+
```
|
|
441
|
+
|
|
442
|
+
---
|
|
443
|
+
|
|
444
|
+
# Celery Worker
|
|
445
|
+
|
|
446
|
+
TriCacheLLM_MMA uses Celery for background operations such as:
|
|
447
|
+
|
|
448
|
+
* persistent cache VDB creation
|
|
449
|
+
* persistent cache population
|
|
450
|
+
|
|
451
|
+
After installing the package, start the cache worker in a **separate terminal**:
|
|
452
|
+
|
|
453
|
+
```bash
|
|
454
|
+
celery -A portable_cache_bgWorkers.portable_cache_celery_conf.celery_app worker --loglevel=info -Q ai
|
|
455
|
+
```
|
|
456
|
+
|
|
457
|
+
You do **not** need to navigate into the package's `site-packages` directory.
|
|
458
|
+
|
|
459
|
+
The command imports the installed package through the active Python environment.
|
|
460
|
+
|
|
461
|
+
A typical setup therefore looks like:
|
|
462
|
+
|
|
463
|
+
```text
|
|
464
|
+
Terminal 1
|
|
465
|
+
└── Your application
|
|
466
|
+
└── FastAPI / Flask / Django / custom service
|
|
467
|
+
|
|
468
|
+
Terminal 2
|
|
469
|
+
└── TriCacheLLM_MMA Celery worker
|
|
470
|
+
```
|
|
471
|
+
|
|
472
|
+
Keep the Celery worker running while using the cache.
|
|
473
|
+
|
|
474
|
+
---
|
|
475
|
+
|
|
476
|
+
# Cohere
|
|
477
|
+
|
|
478
|
+
The persistent cache tier uses Cohere reranking.
|
|
479
|
+
|
|
480
|
+
Create a Cohere API key and provide it during initialization.
|
|
481
|
+
|
|
482
|
+
```python
|
|
483
|
+
cohere_api_key="YOUR_COHERE_API_KEY"
|
|
484
|
+
```
|
|
485
|
+
|
|
486
|
+
Do not commit your API key to Git.
|
|
487
|
+
|
|
488
|
+
For production deployments, use environment variables or your application's secret-management system.
|
|
489
|
+
|
|
490
|
+
---
|
|
491
|
+
|
|
492
|
+
# Quick Start
|
|
493
|
+
|
|
494
|
+
A minimal integration looks like this:
|
|
495
|
+
|
|
496
|
+
```python
|
|
497
|
+
from fastapi import FastAPI, Depends
|
|
498
|
+
from portable_cache_main import (
|
|
499
|
+
create_cache_system,
|
|
500
|
+
check_cache,
|
|
501
|
+
populate_cache,
|
|
502
|
+
)
|
|
503
|
+
|
|
504
|
+
app = FastAPI()
|
|
505
|
+
|
|
506
|
+
# 1. ADMIN SETUP (Run once on startup / first deploy with user_id=0)
|
|
507
|
+
@app.on_event("startup")
|
|
508
|
+
async def startup_event():
|
|
509
|
+
await create_cache_system(
|
|
510
|
+
redis_url="redis://localhost:6379/0",
|
|
511
|
+
cohere_api_key="YOUR_COHERE_API_KEY",
|
|
512
|
+
chroma_db_dir="./chroma_db",
|
|
513
|
+
db_path="./cache.db",
|
|
514
|
+
user_id=0,
|
|
515
|
+
)
|
|
516
|
+
|
|
517
|
+
|
|
518
|
+
# 2. TENANT ENDPOINT (Zero boilerplate for subsequent users)
|
|
519
|
+
@app.post("/xyz")
|
|
520
|
+
async def ask_question(
|
|
521
|
+
user_payload: QuestionRequest,
|
|
522
|
+
user_jwt_payload: TokenDataSchema = Depends(get_user_jwt_payload)
|
|
523
|
+
):
|
|
524
|
+
user_id = user_jwt_payload.user_id
|
|
525
|
+
question = user_payload.user_question # or user_payload.question
|
|
526
|
+
|
|
527
|
+
# Lazy-initialize tenant's isolated cache partition (no keys needed!)
|
|
528
|
+
await create_cache_system(user_id=user_id)
|
|
529
|
+
|
|
530
|
+
# Check cache before hitting your expensive AI pipeline
|
|
531
|
+
if cached := await check_cache(user_id=user_id, user_input=question):
|
|
532
|
+
return {"source": "cache", "response": cached}
|
|
533
|
+
|
|
534
|
+
# Run your heavy AI / RAG pipeline
|
|
535
|
+
llm_response = await your_ai_pipeline(question)
|
|
536
|
+
|
|
537
|
+
# Populate the persistent cache asynchronously
|
|
538
|
+
await populate_cache(
|
|
539
|
+
user_id=user_id,
|
|
540
|
+
to_cache_question=question,
|
|
541
|
+
to_cache_answer=llm_response,
|
|
542
|
+
)
|
|
543
|
+
|
|
544
|
+
return {"source": "ai", "response": llm_response}
|
|
545
|
+
```
|
|
546
|
+
The core cache operations are asynchronous Python APIs and are not inherently tied to FastAPI.
|
|
547
|
+
The included integration example uses FastAPI because it provides a convenient demonstration of multi-user request handling.
|
|
548
|
+
|
|
549
|
+
V1 is primarily demonstrated in a web-service architecture, while future versions aim to make standalone application integration equally straightforward.
|
|
550
|
+
|
|
551
|
+
---
|
|
552
|
+
|
|
553
|
+
|
|
554
|
+
|
|
555
|
+
|
|
556
|
+
|
|
557
|
+
# Cache Data Model
|
|
558
|
+
|
|
559
|
+
Cached entries conceptually contain information similar to:
|
|
560
|
+
|
|
561
|
+
```python
|
|
562
|
+
cache_metadata = {
|
|
563
|
+
"user_id": user_id,
|
|
564
|
+
"question": question,
|
|
565
|
+
"llm_response": response_json_str,
|
|
566
|
+
"created_at": datetime.now(timezone.utc).isoformat(),
|
|
567
|
+
"timestamp": time.time(),
|
|
568
|
+
}
|
|
569
|
+
```
|
|
570
|
+
|
|
571
|
+
The question is stored as the vector document content, while the answer and supporting information are stored as metadata.
|
|
572
|
+
|
|
573
|
+
This allows the persistent VDB to perform semantic retrieval while retaining the original response.
|
|
574
|
+
|
|
575
|
+
---
|
|
576
|
+
|
|
577
|
+
# Extending Metadata Filtering
|
|
578
|
+
|
|
579
|
+
The persistent VDB layer can be customized if your application requires additional constraints.
|
|
580
|
+
|
|
581
|
+
For example, you may want to restrict cache retrieval to:
|
|
582
|
+
|
|
583
|
+
* a specific time period
|
|
584
|
+
* a specific document version
|
|
585
|
+
* a specific tenant resource
|
|
586
|
+
* a specific source
|
|
587
|
+
* a specific application state
|
|
588
|
+
|
|
589
|
+
The persistent retrieval logic can be extended around the VDB lookup.
|
|
590
|
+
|
|
591
|
+
|
|
592
|
+
|
|
593
|
+
|
|
594
|
+
This allows applications to combine semantic retrieval with deterministic metadata constraints.
|
|
595
|
+
|
|
596
|
+
---
|
|
597
|
+
|
|
598
|
+
# Adding Custom Metadata
|
|
599
|
+
|
|
600
|
+
Additional metadata can be added to the cache payload.
|
|
601
|
+
|
|
602
|
+
The cache population path constructs metadata similar to:
|
|
603
|
+
|
|
604
|
+
|
|
605
|
+
|
|
606
|
+
For example:
|
|
607
|
+
|
|
608
|
+
```python
|
|
609
|
+
cache_metadata = {
|
|
610
|
+
"user_id": user_id,
|
|
611
|
+
"question": question,
|
|
612
|
+
"llm_response": response_json_str,
|
|
613
|
+
"created_at": datetime.now(timezone.utc).isoformat(),
|
|
614
|
+
"timestamp": time.time(),
|
|
615
|
+
"document_id": document_id,
|
|
616
|
+
"latest_version": latest_version,
|
|
617
|
+
}
|
|
618
|
+
```
|
|
619
|
+
|
|
620
|
+
Applications can then use those fields as part of their cache validity and retrieval constraints.
|
|
621
|
+
|
|
622
|
+
---
|
|
623
|
+
|
|
624
|
+
# Embedding Model
|
|
625
|
+
|
|
626
|
+
The default embedding model is:
|
|
627
|
+
|
|
628
|
+
```text
|
|
629
|
+
sentence-transformers/all-MiniLM-L6-v2
|
|
630
|
+
```
|
|
631
|
+
|
|
632
|
+
The embedding model is loaded once per Python process and reused by the cache operations handled by that process.
|
|
633
|
+
|
|
634
|
+
The same model instance is shared across users handled by that process.
|
|
635
|
+
|
|
636
|
+
Multiple independent worker processes may naturally maintain their own model instance.
|
|
637
|
+
|
|
638
|
+
This avoids repeatedly loading the embedding model for every user or request.
|
|
639
|
+
|
|
640
|
+
---
|
|
641
|
+
|
|
642
|
+
# Persistence
|
|
643
|
+
|
|
644
|
+
The package separates code from runtime cache state.
|
|
645
|
+
|
|
646
|
+
After installation, the Python package lives inside the consumer's Python environment.
|
|
647
|
+
|
|
648
|
+
Runtime state can live outside the installed package:
|
|
649
|
+
|
|
650
|
+
```text
|
|
651
|
+
your_project/
|
|
652
|
+
│
|
|
653
|
+
├── .portable_cache_internal/
|
|
654
|
+
│ └── .env_protable_cache
|
|
655
|
+
│
|
|
656
|
+
├── cache.db
|
|
657
|
+
│
|
|
658
|
+
├── chroma_db/
|
|
659
|
+
│
|
|
660
|
+
├── your_app/
|
|
661
|
+
│
|
|
662
|
+
└── .venv/
|
|
663
|
+
└── site-packages/
|
|
664
|
+
└── TriCacheLLM_MMA/
|
|
665
|
+
```
|
|
666
|
+
|
|
667
|
+
The cache registry records the important absolute runtime paths so the cache infrastructure does not have to depend on where the package itself was installed.
|
|
668
|
+
|
|
669
|
+
---
|
|
670
|
+
|
|
671
|
+
# Runtime Configuration
|
|
672
|
+
|
|
673
|
+
The package creates an internal configuration area:
|
|
674
|
+
|
|
675
|
+
```text
|
|
676
|
+
.portable_cache_internal/
|
|
677
|
+
```
|
|
678
|
+
|
|
679
|
+
The configuration file contains values such as:
|
|
680
|
+
|
|
681
|
+
```text
|
|
682
|
+
PORTABLE_CACHE_REDIS_URL
|
|
683
|
+
PORTABLE_CACHE_COHERE_API_KEY
|
|
684
|
+
PORTABLE_CACHE_CHROMA_DB_DIR
|
|
685
|
+
PORTABLE_CACHE_REGISTRY_DB
|
|
686
|
+
```
|
|
687
|
+
|
|
688
|
+
Do not commit this directory.
|
|
689
|
+
|
|
690
|
+
It is included in the repository's `.gitignore`.
|
|
691
|
+
|
|
692
|
+
---
|
|
693
|
+
|
|
694
|
+
# Security
|
|
695
|
+
|
|
696
|
+
## Never commit API keys
|
|
697
|
+
|
|
698
|
+
Do not place real API keys into:
|
|
699
|
+
|
|
700
|
+
```text
|
|
701
|
+
example_user_experice.py
|
|
702
|
+
```
|
|
703
|
+
|
|
704
|
+
Use:
|
|
705
|
+
|
|
706
|
+
```python
|
|
707
|
+
cohere_api_key="YOUR_COHERE_API_KEY"
|
|
708
|
+
```
|
|
709
|
+
|
|
710
|
+
or load secrets through your application's environment/secret-management system.
|
|
711
|
+
|
|
712
|
+
## Runtime state
|
|
713
|
+
|
|
714
|
+
The following should remain local to the consumer environment:
|
|
715
|
+
|
|
716
|
+
```text
|
|
717
|
+
.portable_cache_internal/
|
|
718
|
+
*.db
|
|
719
|
+
chroma_db/
|
|
720
|
+
```
|
|
721
|
+
|
|
722
|
+
These are runtime artifacts, not source code.
|
|
723
|
+
|
|
724
|
+
---
|
|
725
|
+
|
|
726
|
+
# Project Structure
|
|
727
|
+
|
|
728
|
+
The V1 repository is intentionally lightweight.
|
|
729
|
+
|
|
730
|
+
```text
|
|
731
|
+
TriCacheLLM_MMA/
|
|
732
|
+
│
|
|
733
|
+
├── portable_cache_main.py
|
|
734
|
+
├── portable_cache_redis.py
|
|
735
|
+
├── portable_cache_dbSchema.py
|
|
736
|
+
│
|
|
737
|
+
├── portable_cache_Ai/
|
|
738
|
+
│ └── portable_cache_rerankAi.py
|
|
739
|
+
│
|
|
740
|
+
├── portable_cache_bgWorkers/
|
|
741
|
+
│ ├── portable_cache_celery_conf.py
|
|
742
|
+
│ └── portable_cache_workers.py
|
|
743
|
+
│
|
|
744
|
+
├── portable_cache_schemas/
|
|
745
|
+
│ ├── portable_cache_dbBase.py
|
|
746
|
+
│ ├── portable_cache_dbConf.py
|
|
747
|
+
│ └── portable_cache_schemas.py
|
|
748
|
+
│
|
|
749
|
+
├── portable_cache_utils/
|
|
750
|
+
│ ├── portable_cache_embedding_model.py
|
|
751
|
+
│ └── protable_cache_DynamicEnv_maker.py
|
|
752
|
+
│
|
|
753
|
+
├── example_user_experice.py
|
|
754
|
+
├── pyproject.toml
|
|
755
|
+
├── req.txt
|
|
756
|
+
└── README.md
|
|
757
|
+
```
|
|
758
|
+
|
|
759
|
+
---
|
|
760
|
+
|
|
761
|
+
# Design Philosophy
|
|
762
|
+
|
|
763
|
+
TriCacheLLM_MMA follows a simple principle:
|
|
764
|
+
|
|
765
|
+
> **Cheap exact lookup first. Cheap semantic lookup second. Expensive persistent retrieval last. LLM inference only after the cache has genuinely missed.**
|
|
766
|
+
|
|
767
|
+
This creates a natural latency hierarchy:
|
|
768
|
+
|
|
769
|
+
```text
|
|
770
|
+
T1
|
|
771
|
+
│
|
|
772
|
+
├── Exact
|
|
773
|
+
├── Lowest overhead
|
|
774
|
+
└── Fastest
|
|
775
|
+
|
|
776
|
+
T2
|
|
777
|
+
│
|
|
778
|
+
├── Semantic
|
|
779
|
+
├── Embedding + HNSW
|
|
780
|
+
└── Still lightweight
|
|
781
|
+
|
|
782
|
+
T3
|
|
783
|
+
│
|
|
784
|
+
├── Persistent
|
|
785
|
+
├── Vector search
|
|
786
|
+
├── Metadata constraints
|
|
787
|
+
└── Reranking
|
|
788
|
+
|
|
789
|
+
LLM
|
|
790
|
+
│
|
|
791
|
+
└── Expensive generation
|
|
792
|
+
```
|
|
793
|
+
|
|
794
|
+
The cache therefore attempts to stop a request as early as possible.
|
|
795
|
+
|
|
796
|
+
---
|
|
797
|
+
|
|
798
|
+
# Why Three Tiers?
|
|
799
|
+
|
|
800
|
+
A single semantic vector database is powerful, but using it for every request introduces unnecessary work.
|
|
801
|
+
|
|
802
|
+
For repeated exact questions, performing:
|
|
803
|
+
|
|
804
|
+
```text
|
|
805
|
+
embedding
|
|
806
|
+
→ vector search
|
|
807
|
+
→ reranking
|
|
808
|
+
```
|
|
809
|
+
|
|
810
|
+
is unnecessary.
|
|
811
|
+
|
|
812
|
+
Similarly, using only exact Redis cannot handle:
|
|
813
|
+
|
|
814
|
+
```text
|
|
815
|
+
"What is Redis used for?"
|
|
816
|
+
|
|
817
|
+
vs.
|
|
818
|
+
|
|
819
|
+
"Can you explain what Redis is used for?"
|
|
820
|
+
```
|
|
821
|
+
|
|
822
|
+
A multi-tier architecture allows each retrieval mechanism to handle the workload it is best suited for.
|
|
823
|
+
|
|
824
|
+
---
|
|
825
|
+
|
|
826
|
+
# V1 Scope
|
|
827
|
+
|
|
828
|
+
This release intentionally focuses on the core cache architecture.
|
|
829
|
+
|
|
830
|
+
Included:
|
|
831
|
+
|
|
832
|
+
* Exact Redis cache
|
|
833
|
+
* Redis HNSW semantic cache
|
|
834
|
+
* Persistent Chroma cache
|
|
835
|
+
* Cohere reranking
|
|
836
|
+
* Celery background processing
|
|
837
|
+
* Multi-user cache isolation
|
|
838
|
+
* SQLite infrastructure registry
|
|
839
|
+
* Cache promotion
|
|
840
|
+
* Portable initialization
|
|
841
|
+
* Async API
|
|
842
|
+
|
|
843
|
+
Not included as first-class abstractions:
|
|
844
|
+
|
|
845
|
+
* automatic Celery process management
|
|
846
|
+
* distributed task orchestration beyond Celery
|
|
847
|
+
* cloud-specific deployment
|
|
848
|
+
* automatic secret management
|
|
849
|
+
* advanced configuration framework
|
|
850
|
+
* full SDK-style class abstraction
|
|
851
|
+
* production observability platform
|
|
852
|
+
* automatic infrastructure provisioning
|
|
853
|
+
|
|
854
|
+
These can be considered for future versions.
|
|
855
|
+
|
|
856
|
+
---
|
|
857
|
+
|
|
858
|
+
|
|
859
|
+
|
|
860
|
+
# Consumer Responsibility
|
|
861
|
+
|
|
862
|
+
The consuming application is responsible for:
|
|
863
|
+
|
|
864
|
+
* running Redis
|
|
865
|
+
* running the Celery worker
|
|
866
|
+
* providing Cohere credentials
|
|
867
|
+
* maintaining its own users
|
|
868
|
+
* executing the actual LLM / AI pipeline
|
|
869
|
+
* deciding when a question should be cached
|
|
870
|
+
* deciding what constitutes an acceptable cache hit for its application
|
|
871
|
+
|
|
872
|
+
TriCacheLLM_MMA is responsible for:
|
|
873
|
+
|
|
874
|
+
* cache infrastructure
|
|
875
|
+
* multi-tier retrieval
|
|
876
|
+
* semantic cache search
|
|
877
|
+
* persistent cache storage
|
|
878
|
+
* cache promotion
|
|
879
|
+
* background cache operations
|
|
880
|
+
* cache VDB lifecycle state
|
|
881
|
+
|
|
882
|
+
This separation allows the package to remain independent of any specific LLM provider or application framework.
|
|
883
|
+
|
|
884
|
+
---
|
|
885
|
+
|
|
886
|
+
# Roadmap
|
|
887
|
+
|
|
888
|
+
## V2: Runtime and Developer Experience
|
|
889
|
+
|
|
890
|
+
Potential V2 improvements may include:
|
|
891
|
+
|
|
892
|
+
* pluggable background execution backends, including Celery, asyncio and synchronous execution
|
|
893
|
+
* remove the requirement for consumers to manually start a Celery worker when using simpler backends
|
|
894
|
+
* configurable cache thresholds and TTLs
|
|
895
|
+
* pluggable rerankers and embedding providers
|
|
896
|
+
* richer metadata filtering
|
|
897
|
+
* improved observability
|
|
898
|
+
* stronger automated test coverage
|
|
899
|
+
* cleaner class-based SDK API
|
|
900
|
+
* packaging and deployment improvements
|
|
901
|
+
|
|
902
|
+
## V3: Standalone and Broader Application Support
|
|
903
|
+
|
|
904
|
+
Potential V3 improvements may include:
|
|
905
|
+
|
|
906
|
+
* remove the current FastAPI-oriented integration assumptions
|
|
907
|
+
* support standalone Python applications and services
|
|
908
|
+
* support broader application architectures beyond web-based APIs
|
|
909
|
+
* provide more flexible integration patterns for single-user and multi-user applications
|
|
910
|
+
* simplified deployment for standalone environments
|
|
911
|
+
|
|
912
|
+
The current V1 intentionally keeps the architecture close to the underlying implementation rather than hiding every component behind abstractions.
|
|
913
|
+
|
|
914
|
+
---
|
|
915
|
+
|
|
916
|
+
|
|
917
|
+
# License
|
|
918
|
+
|
|
919
|
+
This project is licensed under the **GNU Lesser General Public License v3.0 (LGPLv3)**.
|
|
920
|
+
See the [LICENSE](LICENSE) file for details.
|
|
921
|
+
|
|
922
|
+
---
|
|
923
|
+
|
|
924
|
+
# Author
|
|
925
|
+
|
|
926
|
+
**TriCacheLLM_MMA built by Mohib Ashfaq**
|
|
927
|
+
|
|
928
|
+
A portable three-tier semantic caching system for LLM applications.
|
|
929
|
+
|
|
930
|
+
Built around:
|
|
931
|
+
|
|
932
|
+
```text
|
|
933
|
+
Redis
|
|
934
|
+
+
|
|
935
|
+
Redis HNSW
|
|
936
|
+
+
|
|
937
|
+
Chroma
|
|
938
|
+
+
|
|
939
|
+
Cohere
|
|
940
|
+
+
|
|
941
|
+
Celery
|
|
942
|
+
```
|
|
943
|
+
|
|
944
|
+
with the goal of making expensive LLM inference the **last resort rather than the default path**.
|
|
945
|
+
|
|
946
|
+
---
|
|
947
|
+
|
|
948
|
+
# Disclaimer:
|
|
949
|
+
This software is provided "as is" without warranty of any kind. The author is not responsible for any data loss, system failures, or damages arising from its use.
|