adaptq 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- adaptq-0.1.0/MANIFEST.in +7 -0
- adaptq-0.1.0/PKG-INFO +92 -0
- adaptq-0.1.0/README.md +82 -0
- adaptq-0.1.0/adapters/adapter_llamacpp.h +38 -0
- adaptq-0.1.0/adapters/adapter_python.cpp +93 -0
- adaptq-0.1.0/adaptq/__init__.py +4 -0
- adaptq-0.1.0/adaptq/core.py +73 -0
- adaptq-0.1.0/adaptq/torch_adapter.py +48 -0
- adaptq-0.1.0/adaptq.egg-info/PKG-INFO +92 -0
- adaptq-0.1.0/adaptq.egg-info/SOURCES.txt +33 -0
- adaptq-0.1.0/adaptq.egg-info/dependency_links.txt +1 -0
- adaptq-0.1.0/adaptq.egg-info/not-zip-safe +1 -0
- adaptq-0.1.0/adaptq.egg-info/requires.txt +1 -0
- adaptq-0.1.0/adaptq.egg-info/top_level.txt +2 -0
- adaptq-0.1.0/attention/attention.cpp +543 -0
- adaptq-0.1.0/cache/ring_buffer.cpp +106 -0
- adaptq-0.1.0/core/adaptq_backend_vtable.cpp +28 -0
- adaptq-0.1.0/core/adaptq_c_api.cpp +206 -0
- adaptq-0.1.0/core/codebook.cpp +102 -0
- adaptq-0.1.0/core/fwht.cpp +113 -0
- adaptq-0.1.0/core/quantizer.cpp +205 -0
- adaptq-0.1.0/include/adaptq.h +113 -0
- adaptq-0.1.0/include/adaptq_backend.h +84 -0
- adaptq-0.1.0/include/adaptq_mha_backend.h +48 -0
- adaptq-0.1.0/include/attention.h +37 -0
- adaptq-0.1.0/include/codebook.h +32 -0
- adaptq-0.1.0/include/fwht.h +21 -0
- adaptq-0.1.0/include/quantizer.h +43 -0
- adaptq-0.1.0/include/ring_buffer.h +92 -0
- adaptq-0.1.0/include/timer.h +13 -0
- adaptq-0.1.0/pyproject.toml +19 -0
- adaptq-0.1.0/setup.cfg +4 -0
- adaptq-0.1.0/setup.py +31 -0
- adaptq-0.1.0/tests/test_accuracy.py +90 -0
- adaptq-0.1.0/utils/timer.cpp +34 -0
adaptq-0.1.0/MANIFEST.in
ADDED
adaptq-0.1.0/PKG-INFO
ADDED
|
@@ -0,0 +1,92 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: adaptq
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Adaptive Streaming Vector Quantization KV Cache for LLMs
|
|
5
|
+
Author-email: "Lakshmikanthan K (letchupkt)" <letchupkt.dev@gmail.com>
|
|
6
|
+
Project-URL: Homepage, https://github.com/l3tchupkt/adaptq
|
|
7
|
+
Requires-Python: >=3.7
|
|
8
|
+
Description-Content-Type: text/markdown
|
|
9
|
+
Requires-Dist: numpy
|
|
10
|
+
|
|
11
|
+
# AdapTQ: Adaptive Streaming Vector Quantization
|
|
12
|
+
|
|
13
|
+
**AdapTQ** is a production-grade C++17 KV cache quantization engine for LLM inference on edge and memory-constrained systems.
|
|
14
|
+
|
|
15
|
+
**The Integration Pitch:** AdapTQ is an **optional KV-cache backend**. It runs entirely on the CPU, requires **no model changes**, and fits into existing inference pipelines with minimal adapter-style wrapper logic.
|
|
16
|
+
|
|
17
|
+
## 🚀 Quickstart
|
|
18
|
+
|
|
19
|
+
```python
|
|
20
|
+
import torch; from adaptq import AdaptQAttention
|
|
21
|
+
# 1. Initialize drop-in PyTorch wrapper (4-bit default)
|
|
22
|
+
layer = AdaptQAttention(dim=128, heads=4)
|
|
23
|
+
# 2. Forward pass dynamically routes continuous BxHxD generation tensors
|
|
24
|
+
out = layer(q=torch.randn(1, 4, 128), k=torch.randn(1, 4, 128), v=torch.randn(1, 4, 128))
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
## 📊 Real-world Benchmarks
|
|
28
|
+
|
|
29
|
+
Tested on standard AVX2 desktop hardware (4 heads, `dim=128`, caching up to 4096 tokens).
|
|
30
|
+
*Note: We ignore sequence lengths `< 256` in these claims, as short sequences are explicitly routed to standard FP32 execution via our hybrid fallback.*
|
|
31
|
+
|
|
32
|
+
| Metric | Result (Seq ≥ 256) |
|
|
33
|
+
| --- | --- |
|
|
34
|
+
| **Latencies** | p50: `877.9 µs` \| p95: `2161.8 µs` |
|
|
35
|
+
| **Stable Speedup** | ~10.18x vs NumPy FP32 equivalent |
|
|
36
|
+
| **Throughput** | ~1,139 tokens/sec |
|
|
37
|
+
| **Memory** | 2.10 MB vs FP16's 8.39 MB (**4.0x smaller**) |
|
|
38
|
+
|
|
39
|
+
### Quantization Fidelity (Honest Metrics)
|
|
40
|
+
We use a targeted $\pm 3\sigma$ variance soft-clipping on FWHT distributions without altering Max-Lloyd codebooks. Our strictly measured empirical quality against baseline FP32:
|
|
41
|
+
- **Cosine Similarity**: ~0.947 (1.000 = exact identical match)
|
|
42
|
+
- **Mean Squared Error (MSE)**: ~1.8e-04
|
|
43
|
+
|
|
44
|
+

|
|
45
|
+
|
|
46
|
+
## 🏗 Architecture & Features
|
|
47
|
+
|
|
48
|
+
- **Unified SIMD Pipeline**: 2, 3, and 4-bit decoding share a single, quad-unrolled branchless loop using AVX2 intrinsics. No scalar fallbacks in the hot path.
|
|
49
|
+
- **Fast Hadamard Rotation (HAR)**: $O(d \log d)$ fully in-place rotation minimizes outliers gracefully before codebook matching.
|
|
50
|
+
- **Zero Heap Allocations**: Pure stack/thread-local memory buffers in the hot path.
|
|
51
|
+
- **Pre-Compute LUTs**: Dot products execute directly against packed indices in SIMD registers—avoiding full dequantization inside the attention kernel.
|
|
52
|
+
|
|
53
|
+
## 🛠 Installation & Integration
|
|
54
|
+
|
|
55
|
+
```bash
|
|
56
|
+
git clone https://github.com/l3tchupkt/adaptq.git
|
|
57
|
+
cd adaptq
|
|
58
|
+
|
|
59
|
+
# Build python PyBind modules natively
|
|
60
|
+
pip install .
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
### Python Native Application
|
|
64
|
+
|
|
65
|
+
Use the native generic python API to bypass neural-network tensors explicitly:
|
|
66
|
+
|
|
67
|
+
```python
|
|
68
|
+
import numpy as np
|
|
69
|
+
from adaptq import Engine
|
|
70
|
+
|
|
71
|
+
engine = Engine(dim=128, heads=4, bits=4, capacity=2048)
|
|
72
|
+
k, v, q = np.random.randn(4, 128), np.random.randn(4, 128), np.random.randn(4, 128)
|
|
73
|
+
|
|
74
|
+
engine.append(k, v)
|
|
75
|
+
output = engine.compute(q)
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
### llama.cpp Adapter
|
|
79
|
+
|
|
80
|
+
Using `AdapTQ` as the native KV Cache replacement during computation phase over GGML.
|
|
81
|
+
*(Requires using `llm_build_kqv` hooks. See `/integration/llama_cpp_patch.md` for full unified patch details.)*
|
|
82
|
+
|
|
83
|
+
```cpp
|
|
84
|
+
#include "adapters/adapter_llamacpp.h"
|
|
85
|
+
LlamaCppAdaptQAdapter adapter(n_heads, head_dim, bits, capacity, seed, v_mass, hybrid_thr);
|
|
86
|
+
adapter.feed_kv(head, key_array, val_array, token_pos);
|
|
87
|
+
adapter.attention(head, query_array, out_array);
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
## 📝 License
|
|
91
|
+
|
|
92
|
+
See active repository license policies. Developed based on *AdapTQ: Adaptive Streaming Vector Quantization for Edge-Deployed Large Language Models*.
|
adaptq-0.1.0/README.md
ADDED
|
@@ -0,0 +1,82 @@
|
|
|
1
|
+
# AdapTQ: Adaptive Streaming Vector Quantization
|
|
2
|
+
|
|
3
|
+
**AdapTQ** is a production-grade C++17 KV cache quantization engine for LLM inference on edge and memory-constrained systems.
|
|
4
|
+
|
|
5
|
+
**The Integration Pitch:** AdapTQ is an **optional KV-cache backend**. It runs entirely on the CPU, requires **no model changes**, and fits into existing inference pipelines with minimal adapter-style wrapper logic.
|
|
6
|
+
|
|
7
|
+
## 🚀 Quickstart
|
|
8
|
+
|
|
9
|
+
```python
|
|
10
|
+
import torch; from adaptq import AdaptQAttention
|
|
11
|
+
# 1. Initialize drop-in PyTorch wrapper (4-bit default)
|
|
12
|
+
layer = AdaptQAttention(dim=128, heads=4)
|
|
13
|
+
# 2. Forward pass dynamically routes continuous BxHxD generation tensors
|
|
14
|
+
out = layer(q=torch.randn(1, 4, 128), k=torch.randn(1, 4, 128), v=torch.randn(1, 4, 128))
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
## 📊 Real-world Benchmarks
|
|
18
|
+
|
|
19
|
+
Tested on standard AVX2 desktop hardware (4 heads, `dim=128`, caching up to 4096 tokens).
|
|
20
|
+
*Note: We ignore sequence lengths `< 256` in these claims, as short sequences are explicitly routed to standard FP32 execution via our hybrid fallback.*
|
|
21
|
+
|
|
22
|
+
| Metric | Result (Seq ≥ 256) |
|
|
23
|
+
| --- | --- |
|
|
24
|
+
| **Latencies** | p50: `877.9 µs` \| p95: `2161.8 µs` |
|
|
25
|
+
| **Stable Speedup** | ~10.18x vs NumPy FP32 equivalent |
|
|
26
|
+
| **Throughput** | ~1,139 tokens/sec |
|
|
27
|
+
| **Memory** | 2.10 MB vs FP16's 8.39 MB (**4.0x smaller**) |
|
|
28
|
+
|
|
29
|
+
### Quantization Fidelity (Honest Metrics)
|
|
30
|
+
We use a targeted $\pm 3\sigma$ variance soft-clipping on FWHT distributions without altering Max-Lloyd codebooks. Our strictly measured empirical quality against baseline FP32:
|
|
31
|
+
- **Cosine Similarity**: ~0.947 (1.000 = exact identical match)
|
|
32
|
+
- **Mean Squared Error (MSE)**: ~1.8e-04
|
|
33
|
+
|
|
34
|
+

|
|
35
|
+
|
|
36
|
+
## 🏗 Architecture & Features
|
|
37
|
+
|
|
38
|
+
- **Unified SIMD Pipeline**: 2, 3, and 4-bit decoding share a single, quad-unrolled branchless loop using AVX2 intrinsics. No scalar fallbacks in the hot path.
|
|
39
|
+
- **Fast Hadamard Rotation (HAR)**: $O(d \log d)$ fully in-place rotation minimizes outliers gracefully before codebook matching.
|
|
40
|
+
- **Zero Heap Allocations**: Pure stack/thread-local memory buffers in the hot path.
|
|
41
|
+
- **Pre-Compute LUTs**: Dot products execute directly against packed indices in SIMD registers—avoiding full dequantization inside the attention kernel.
|
|
42
|
+
|
|
43
|
+
## 🛠 Installation & Integration
|
|
44
|
+
|
|
45
|
+
```bash
|
|
46
|
+
git clone https://github.com/l3tchupkt/adaptq.git
|
|
47
|
+
cd adaptq
|
|
48
|
+
|
|
49
|
+
# Build python PyBind modules natively
|
|
50
|
+
pip install .
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
### Python Native Application
|
|
54
|
+
|
|
55
|
+
Use the native generic python API to bypass neural-network tensors explicitly:
|
|
56
|
+
|
|
57
|
+
```python
|
|
58
|
+
import numpy as np
|
|
59
|
+
from adaptq import Engine
|
|
60
|
+
|
|
61
|
+
engine = Engine(dim=128, heads=4, bits=4, capacity=2048)
|
|
62
|
+
k, v, q = np.random.randn(4, 128), np.random.randn(4, 128), np.random.randn(4, 128)
|
|
63
|
+
|
|
64
|
+
engine.append(k, v)
|
|
65
|
+
output = engine.compute(q)
|
|
66
|
+
```
|
|
67
|
+
|
|
68
|
+
### llama.cpp Adapter
|
|
69
|
+
|
|
70
|
+
Using `AdapTQ` as the native KV Cache replacement during computation phase over GGML.
|
|
71
|
+
*(Requires using `llm_build_kqv` hooks. See `/integration/llama_cpp_patch.md` for full unified patch details.)*
|
|
72
|
+
|
|
73
|
+
```cpp
|
|
74
|
+
#include "adapters/adapter_llamacpp.h"
|
|
75
|
+
LlamaCppAdaptQAdapter adapter(n_heads, head_dim, bits, capacity, seed, v_mass, hybrid_thr);
|
|
76
|
+
adapter.feed_kv(head, key_array, val_array, token_pos);
|
|
77
|
+
adapter.attention(head, query_array, out_array);
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
## 📝 License
|
|
81
|
+
|
|
82
|
+
See active repository license policies. Developed based on *AdapTQ: Adaptive Streaming Vector Quantization for Edge-Deployed Large Language Models*.
|
|
@@ -0,0 +1,38 @@
|
|
|
1
|
+
#pragma once
|
|
2
|
+
#include "../include/adaptq_mha_backend.h"
|
|
3
|
+
#include <cstddef>
|
|
4
|
+
|
|
5
|
+
/* -------------------------------------------------------------------------
|
|
6
|
+
* LlamaCppAdaptQAdapter
|
|
7
|
+
*
|
|
8
|
+
* Thin adapter between llama.cpp's float tensor data and the AdapTQ core.
|
|
9
|
+
* Include this file in llama.cpp builds. Zero GGML headers required here.
|
|
10
|
+
* ------------------------------------------------------------------------- */
|
|
11
|
+
class LlamaCppAdaptQAdapter {
|
|
12
|
+
public:
|
|
13
|
+
LlamaCppAdaptQAdapter(int n_heads, int head_dim, int bits, int capacity,
|
|
14
|
+
uint64_t seed = 0, float v_mass = 0.95f,
|
|
15
|
+
int hybrid_thresh = 512);
|
|
16
|
+
|
|
17
|
+
/** Call once per incoming token, once per head, after graph compute. */
|
|
18
|
+
void feed_kv(int head, const float *key, const float *val, int token_pos);
|
|
19
|
+
|
|
20
|
+
/** Single-query attention forward (decode step). */
|
|
21
|
+
int attention(int head, const float *query, float *out);
|
|
22
|
+
|
|
23
|
+
/** Multi-query batch attention (prefill step). */
|
|
24
|
+
int attention_batch(int head, const float *queries, int n_queries,
|
|
25
|
+
float *outs);
|
|
26
|
+
|
|
27
|
+
/** Clear all cached KV (e.g. on context reset). */
|
|
28
|
+
void reset();
|
|
29
|
+
|
|
30
|
+
/** Compressed bytes in the KV cache. */
|
|
31
|
+
size_t kv_bytes() const;
|
|
32
|
+
|
|
33
|
+
/** Logs "KV: X MB vs FP16: Y MB" to stderr. */
|
|
34
|
+
void log_memory(int capacity) const;
|
|
35
|
+
|
|
36
|
+
private:
|
|
37
|
+
AdaptQMHABackend _backend;
|
|
38
|
+
};
|
|
@@ -0,0 +1,93 @@
|
|
|
1
|
+
/* adapter_python.cpp
|
|
2
|
+
*
|
|
3
|
+
* Pluggable adapter: Python (pybind11) → AdapTQ core
|
|
4
|
+
*
|
|
5
|
+
* Build with:
|
|
6
|
+
* pip install pybind11
|
|
7
|
+
* g++ -O3 -march=native -mavx2 -shared -fPIC \
|
|
8
|
+
* $(python3 -m pybind11 --includes) \
|
|
9
|
+
* -Iinclude \
|
|
10
|
+
* adapters/adapter_python.cpp \
|
|
11
|
+
* -L. -ladaptq -Wl,-rpath,. \
|
|
12
|
+
* -o adaptq_py$(python3-config --extension-suffix)
|
|
13
|
+
*
|
|
14
|
+
* Python usage:
|
|
15
|
+
* import adaptq_py as aq
|
|
16
|
+
* ctx = aq.MHAContext(n_heads=32, head_dim=128, bits=4, capacity=8192)
|
|
17
|
+
* ctx.append(head=0, key=k_np, val=v_np, pos=0)
|
|
18
|
+
* out = ctx.compute(head=0, query=q_np)
|
|
19
|
+
* batch_out = ctx.compute_batch(head=0, queries=q_batch_np)
|
|
20
|
+
*/
|
|
21
|
+
|
|
22
|
+
#include "../include/adaptq_mha_backend.h"
|
|
23
|
+
#include <pybind11/numpy.h>
|
|
24
|
+
#include <pybind11/pybind11.h>
|
|
25
|
+
#include <stdexcept>
|
|
26
|
+
|
|
27
|
+
|
|
28
|
+
namespace py = pybind11;
|
|
29
|
+
|
|
30
|
+
class PyMHAContext {
|
|
31
|
+
public:
|
|
32
|
+
PyMHAContext(int n_heads, int head_dim, int bits, int capacity,
|
|
33
|
+
uint64_t seed = 0, float v_mass = 0.95f, int hybrid_thresh = 512)
|
|
34
|
+
: _backend(n_heads, head_dim, bits, capacity, seed, v_mass,
|
|
35
|
+
hybrid_thresh),
|
|
36
|
+
_head_dim(head_dim), _n_heads(n_heads) {}
|
|
37
|
+
|
|
38
|
+
void append(int head, py::array_t<float> key, py::array_t<float> val,
|
|
39
|
+
int pos) {
|
|
40
|
+
auto k = key.unchecked<1>();
|
|
41
|
+
auto v = val.unchecked<1>();
|
|
42
|
+
if (k.shape(0) != _head_dim || v.shape(0) != _head_dim)
|
|
43
|
+
throw std::invalid_argument("key/val dim mismatch");
|
|
44
|
+
_backend.append_kv(head, k.data(0), v.data(0), pos);
|
|
45
|
+
}
|
|
46
|
+
|
|
47
|
+
py::array_t<float> compute(int head, py::array_t<float> query) {
|
|
48
|
+
auto q = query.unchecked<1>();
|
|
49
|
+
if (q.shape(0) != _head_dim)
|
|
50
|
+
throw std::invalid_argument("query dim mismatch");
|
|
51
|
+
auto out = py::array_t<float>(_head_dim);
|
|
52
|
+
_backend.compute(head, q.data(0), out.mutable_data(0));
|
|
53
|
+
return out;
|
|
54
|
+
}
|
|
55
|
+
|
|
56
|
+
py::array_t<float> compute_batch(int head, py::array_t<float> queries) {
|
|
57
|
+
auto q = queries.unchecked<2>();
|
|
58
|
+
int nq = (int)q.shape(0);
|
|
59
|
+
if (q.shape(1) != _head_dim)
|
|
60
|
+
throw std::invalid_argument("queries dim mismatch");
|
|
61
|
+
auto out = py::array_t<float>({nq, _head_dim});
|
|
62
|
+
_backend.compute_batch(head, q.data(0, 0), nq, out.mutable_data(0, 0));
|
|
63
|
+
return out;
|
|
64
|
+
}
|
|
65
|
+
|
|
66
|
+
void reset() { _backend.reset(); }
|
|
67
|
+
size_t kv_bytes() const { return _backend.kv_bytes(); }
|
|
68
|
+
int n_heads() const { return _n_heads; }
|
|
69
|
+
int head_dim() const { return _head_dim; }
|
|
70
|
+
|
|
71
|
+
private:
|
|
72
|
+
AdaptQMHABackend _backend;
|
|
73
|
+
int _head_dim, _n_heads;
|
|
74
|
+
};
|
|
75
|
+
|
|
76
|
+
PYBIND11_MODULE(adaptq_py, m) {
|
|
77
|
+
m.doc() = "AdapTQ — Quantized KV-cache attention, Python bindings";
|
|
78
|
+
|
|
79
|
+
py::class_<PyMHAContext>(m, "MHAContext")
|
|
80
|
+
.def(py::init<int, int, int, int, uint64_t, float, int>(),
|
|
81
|
+
py::arg("n_heads"), py::arg("head_dim"), py::arg("bits"),
|
|
82
|
+
py::arg("capacity"), py::arg("seed") = 0, py::arg("v_mass") = 0.95f,
|
|
83
|
+
py::arg("hybrid_thresh") = 512)
|
|
84
|
+
.def("append", &PyMHAContext::append, py::arg("head"), py::arg("key"),
|
|
85
|
+
py::arg("val"), py::arg("pos"))
|
|
86
|
+
.def("compute", &PyMHAContext::compute, py::arg("head"), py::arg("query"))
|
|
87
|
+
.def("compute_batch", &PyMHAContext::compute_batch, py::arg("head"),
|
|
88
|
+
py::arg("queries"))
|
|
89
|
+
.def("reset", &PyMHAContext::reset)
|
|
90
|
+
.def("kv_bytes", &PyMHAContext::kv_bytes)
|
|
91
|
+
.def_property_readonly("n_heads", &PyMHAContext::n_heads)
|
|
92
|
+
.def_property_readonly("head_dim", &PyMHAContext::head_dim);
|
|
93
|
+
}
|
|
@@ -0,0 +1,73 @@
|
|
|
1
|
+
import numpy as np
|
|
2
|
+
try:
|
|
3
|
+
import adaptq_py
|
|
4
|
+
except ImportError:
|
|
5
|
+
raise ImportError("AdapTQ C++ extension not built. Run 'pip install .' to build from source.")
|
|
6
|
+
|
|
7
|
+
class Engine:
|
|
8
|
+
"""
|
|
9
|
+
API for Adaptive Streaming Vector Quantization.
|
|
10
|
+
Usage:
|
|
11
|
+
engine = adaptq.Engine(dim=128, heads=4, bits=4)
|
|
12
|
+
engine.append(k, v)
|
|
13
|
+
out = engine.compute(q)
|
|
14
|
+
"""
|
|
15
|
+
def __init__(self, dim: int, heads: int, bits: int = 4,
|
|
16
|
+
capacity: int = 4096, seed: int = 42,
|
|
17
|
+
v_mass: float = 0.95, hybrid_thresh: int = 512):
|
|
18
|
+
self.dim = dim
|
|
19
|
+
self.heads = heads
|
|
20
|
+
self.bits = bits
|
|
21
|
+
self.capacity = capacity
|
|
22
|
+
|
|
23
|
+
# Internal state
|
|
24
|
+
self._ctx = adaptq_py.MHAContext(
|
|
25
|
+
n_heads=heads,
|
|
26
|
+
head_dim=dim,
|
|
27
|
+
bits=bits,
|
|
28
|
+
capacity=capacity,
|
|
29
|
+
seed=seed,
|
|
30
|
+
v_mass=v_mass,
|
|
31
|
+
hybrid_thresh=hybrid_thresh
|
|
32
|
+
)
|
|
33
|
+
self.pos = 0
|
|
34
|
+
|
|
35
|
+
def append(self, k: np.ndarray, v: np.ndarray):
|
|
36
|
+
"""
|
|
37
|
+
Append k, v for the current step.
|
|
38
|
+
Supports inputs of shape (heads, dim).
|
|
39
|
+
"""
|
|
40
|
+
if k.shape != (self.heads, self.dim) or v.shape != (self.heads, self.dim):
|
|
41
|
+
raise ValueError(f"k and v must be shape ({self.heads}, {self.dim})")
|
|
42
|
+
|
|
43
|
+
for h in range(self.heads):
|
|
44
|
+
# Using astype to guarantee contiguous float32 alignment to C++ backend
|
|
45
|
+
k_h = np.ascontiguousarray(k[h], dtype=np.float32)
|
|
46
|
+
v_h = np.ascontiguousarray(v[h], dtype=np.float32)
|
|
47
|
+
self._ctx.append(h, k_h, v_h, self.pos)
|
|
48
|
+
|
|
49
|
+
self.pos += 1
|
|
50
|
+
|
|
51
|
+
def compute(self, q: np.ndarray) -> np.ndarray:
|
|
52
|
+
"""
|
|
53
|
+
Compute attention output for query q.
|
|
54
|
+
Supports inputs of shape (heads, dim). Returns shape (heads, dim).
|
|
55
|
+
"""
|
|
56
|
+
if q.shape != (self.heads, self.dim):
|
|
57
|
+
raise ValueError(f"q must be shape ({self.heads}, {self.dim})")
|
|
58
|
+
|
|
59
|
+
out = np.zeros((self.heads, self.dim), dtype=np.float32)
|
|
60
|
+
for h in range(self.heads):
|
|
61
|
+
q_h = np.ascontiguousarray(q[h], dtype=np.float32)
|
|
62
|
+
out[h] = self._ctx.compute(h, q_h)
|
|
63
|
+
|
|
64
|
+
return out
|
|
65
|
+
|
|
66
|
+
def reset(self):
|
|
67
|
+
"""Clear the KV cache."""
|
|
68
|
+
self._ctx.reset()
|
|
69
|
+
self.pos = 0
|
|
70
|
+
|
|
71
|
+
@property
|
|
72
|
+
def kv_bytes(self) -> int:
|
|
73
|
+
return self._ctx.kv_bytes()
|
|
@@ -0,0 +1,48 @@
|
|
|
1
|
+
try:
|
|
2
|
+
import torch
|
|
3
|
+
import torch.nn as nn
|
|
4
|
+
except ImportError:
|
|
5
|
+
raise ImportError("PyTorch is required to use the AdaptQ PyTorch wrapper.")
|
|
6
|
+
|
|
7
|
+
from .core import Engine
|
|
8
|
+
import numpy as np
|
|
9
|
+
|
|
10
|
+
class AdaptQAttention(nn.Module):
|
|
11
|
+
"""
|
|
12
|
+
PyTorch Wrapper for AdapTQ.
|
|
13
|
+
Integrates directly with PyTorch's nn.Module API.
|
|
14
|
+
"""
|
|
15
|
+
def __init__(self, dim: int, heads: int, bits: int = 4, capacity: int = 4096, seed: int = 42):
|
|
16
|
+
super().__init__()
|
|
17
|
+
self.dim = dim
|
|
18
|
+
self.heads = heads
|
|
19
|
+
self.bits = bits
|
|
20
|
+
self.engine = Engine(dim=dim, heads=heads, bits=bits, capacity=capacity, seed=seed)
|
|
21
|
+
|
|
22
|
+
def forward(self, q: torch.Tensor, k: torch.Tensor, v: torch.Tensor) -> torch.Tensor:
|
|
23
|
+
"""
|
|
24
|
+
Expects inputs of shape (batch_size, heads, dim).
|
|
25
|
+
Note: Currently AdapTQ handles a batch_size of 1 for append_kv internally,
|
|
26
|
+
but we can loop over the batch size or just assert batch_size == 1 for generation step.
|
|
27
|
+
"""
|
|
28
|
+
assert q.dim() == 3 and k.dim() == 3 and v.dim() == 3, "Expected 3D tensors: (batch, heads, dim)"
|
|
29
|
+
batch_size = q.size(0)
|
|
30
|
+
|
|
31
|
+
# Move to CPU numpy for backend (AdapTQ backend is pure CPU C++)
|
|
32
|
+
k_np = k.detach().cpu().numpy()
|
|
33
|
+
v_np = v.detach().cpu().numpy()
|
|
34
|
+
q_np = q.detach().cpu().numpy()
|
|
35
|
+
|
|
36
|
+
out = np.zeros_like(q_np)
|
|
37
|
+
|
|
38
|
+
# Currently, AdapTQ C++ context is single-instance per `Engine`.
|
|
39
|
+
# For batch inference, context is strictly sequential in Generation step.
|
|
40
|
+
for b in range(batch_size):
|
|
41
|
+
self.engine.append(k_np[b], v_np[b])
|
|
42
|
+
out[b] = self.engine.compute(q_np[b])
|
|
43
|
+
|
|
44
|
+
# Returning tensor on the same device as query
|
|
45
|
+
return torch.from_numpy(out).to(q.device)
|
|
46
|
+
|
|
47
|
+
def reset_cache(self):
|
|
48
|
+
self.engine.reset()
|
|
@@ -0,0 +1,92 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: adaptq
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Adaptive Streaming Vector Quantization KV Cache for LLMs
|
|
5
|
+
Author-email: "Lakshmikanthan K (letchupkt)" <letchupkt.dev@gmail.com>
|
|
6
|
+
Project-URL: Homepage, https://github.com/l3tchupkt/adaptq
|
|
7
|
+
Requires-Python: >=3.7
|
|
8
|
+
Description-Content-Type: text/markdown
|
|
9
|
+
Requires-Dist: numpy
|
|
10
|
+
|
|
11
|
+
# AdapTQ: Adaptive Streaming Vector Quantization
|
|
12
|
+
|
|
13
|
+
**AdapTQ** is a production-grade C++17 KV cache quantization engine for LLM inference on edge and memory-constrained systems.
|
|
14
|
+
|
|
15
|
+
**The Integration Pitch:** AdapTQ is an **optional KV-cache backend**. It runs entirely on the CPU, requires **no model changes**, and fits into existing inference pipelines with minimal adapter-style wrapper logic.
|
|
16
|
+
|
|
17
|
+
## 🚀 Quickstart
|
|
18
|
+
|
|
19
|
+
```python
|
|
20
|
+
import torch; from adaptq import AdaptQAttention
|
|
21
|
+
# 1. Initialize drop-in PyTorch wrapper (4-bit default)
|
|
22
|
+
layer = AdaptQAttention(dim=128, heads=4)
|
|
23
|
+
# 2. Forward pass dynamically routes continuous BxHxD generation tensors
|
|
24
|
+
out = layer(q=torch.randn(1, 4, 128), k=torch.randn(1, 4, 128), v=torch.randn(1, 4, 128))
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
## 📊 Real-world Benchmarks
|
|
28
|
+
|
|
29
|
+
Tested on standard AVX2 desktop hardware (4 heads, `dim=128`, caching up to 4096 tokens).
|
|
30
|
+
*Note: We ignore sequence lengths `< 256` in these claims, as short sequences are explicitly routed to standard FP32 execution via our hybrid fallback.*
|
|
31
|
+
|
|
32
|
+
| Metric | Result (Seq ≥ 256) |
|
|
33
|
+
| --- | --- |
|
|
34
|
+
| **Latencies** | p50: `877.9 µs` \| p95: `2161.8 µs` |
|
|
35
|
+
| **Stable Speedup** | ~10.18x vs NumPy FP32 equivalent |
|
|
36
|
+
| **Throughput** | ~1,139 tokens/sec |
|
|
37
|
+
| **Memory** | 2.10 MB vs FP16's 8.39 MB (**4.0x smaller**) |
|
|
38
|
+
|
|
39
|
+
### Quantization Fidelity (Honest Metrics)
|
|
40
|
+
We use a targeted $\pm 3\sigma$ variance soft-clipping on FWHT distributions without altering Max-Lloyd codebooks. Our strictly measured empirical quality against baseline FP32:
|
|
41
|
+
- **Cosine Similarity**: ~0.947 (1.000 = exact identical match)
|
|
42
|
+
- **Mean Squared Error (MSE)**: ~1.8e-04
|
|
43
|
+
|
|
44
|
+

|
|
45
|
+
|
|
46
|
+
## 🏗 Architecture & Features
|
|
47
|
+
|
|
48
|
+
- **Unified SIMD Pipeline**: 2, 3, and 4-bit decoding share a single, quad-unrolled branchless loop using AVX2 intrinsics. No scalar fallbacks in the hot path.
|
|
49
|
+
- **Fast Hadamard Rotation (HAR)**: $O(d \log d)$ fully in-place rotation minimizes outliers gracefully before codebook matching.
|
|
50
|
+
- **Zero Heap Allocations**: Pure stack/thread-local memory buffers in the hot path.
|
|
51
|
+
- **Pre-Compute LUTs**: Dot products execute directly against packed indices in SIMD registers—avoiding full dequantization inside the attention kernel.
|
|
52
|
+
|
|
53
|
+
## 🛠 Installation & Integration
|
|
54
|
+
|
|
55
|
+
```bash
|
|
56
|
+
git clone https://github.com/l3tchupkt/adaptq.git
|
|
57
|
+
cd adaptq
|
|
58
|
+
|
|
59
|
+
# Build python PyBind modules natively
|
|
60
|
+
pip install .
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
### Python Native Application
|
|
64
|
+
|
|
65
|
+
Use the native generic python API to bypass neural-network tensors explicitly:
|
|
66
|
+
|
|
67
|
+
```python
|
|
68
|
+
import numpy as np
|
|
69
|
+
from adaptq import Engine
|
|
70
|
+
|
|
71
|
+
engine = Engine(dim=128, heads=4, bits=4, capacity=2048)
|
|
72
|
+
k, v, q = np.random.randn(4, 128), np.random.randn(4, 128), np.random.randn(4, 128)
|
|
73
|
+
|
|
74
|
+
engine.append(k, v)
|
|
75
|
+
output = engine.compute(q)
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
### llama.cpp Adapter
|
|
79
|
+
|
|
80
|
+
Using `AdapTQ` as the native KV Cache replacement during computation phase over GGML.
|
|
81
|
+
*(Requires using `llm_build_kqv` hooks. See `/integration/llama_cpp_patch.md` for full unified patch details.)*
|
|
82
|
+
|
|
83
|
+
```cpp
|
|
84
|
+
#include "adapters/adapter_llamacpp.h"
|
|
85
|
+
LlamaCppAdaptQAdapter adapter(n_heads, head_dim, bits, capacity, seed, v_mass, hybrid_thr);
|
|
86
|
+
adapter.feed_kv(head, key_array, val_array, token_pos);
|
|
87
|
+
adapter.attention(head, query_array, out_array);
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
## 📝 License
|
|
91
|
+
|
|
92
|
+
See active repository license policies. Developed based on *AdapTQ: Adaptive Streaming Vector Quantization for Edge-Deployed Large Language Models*.
|
|
@@ -0,0 +1,33 @@
|
|
|
1
|
+
MANIFEST.in
|
|
2
|
+
README.md
|
|
3
|
+
pyproject.toml
|
|
4
|
+
setup.py
|
|
5
|
+
adapters/adapter_llamacpp.h
|
|
6
|
+
adapters/adapter_python.cpp
|
|
7
|
+
adaptq/__init__.py
|
|
8
|
+
adaptq/core.py
|
|
9
|
+
adaptq/torch_adapter.py
|
|
10
|
+
adaptq.egg-info/PKG-INFO
|
|
11
|
+
adaptq.egg-info/SOURCES.txt
|
|
12
|
+
adaptq.egg-info/dependency_links.txt
|
|
13
|
+
adaptq.egg-info/not-zip-safe
|
|
14
|
+
adaptq.egg-info/requires.txt
|
|
15
|
+
adaptq.egg-info/top_level.txt
|
|
16
|
+
attention/attention.cpp
|
|
17
|
+
cache/ring_buffer.cpp
|
|
18
|
+
core/adaptq_backend_vtable.cpp
|
|
19
|
+
core/adaptq_c_api.cpp
|
|
20
|
+
core/codebook.cpp
|
|
21
|
+
core/fwht.cpp
|
|
22
|
+
core/quantizer.cpp
|
|
23
|
+
include/adaptq.h
|
|
24
|
+
include/adaptq_backend.h
|
|
25
|
+
include/adaptq_mha_backend.h
|
|
26
|
+
include/attention.h
|
|
27
|
+
include/codebook.h
|
|
28
|
+
include/fwht.h
|
|
29
|
+
include/quantizer.h
|
|
30
|
+
include/ring_buffer.h
|
|
31
|
+
include/timer.h
|
|
32
|
+
tests/test_accuracy.py
|
|
33
|
+
utils/timer.cpp
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
numpy
|