adaptq 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,7 @@
1
+ include include/*.h
2
+ include adapters/*.h
3
+ include core/*.h
4
+ include cache/*.h
5
+ include utils/*.h
6
+ include README.md
7
+ include LICENSE
adaptq-0.1.0/PKG-INFO ADDED
@@ -0,0 +1,92 @@
1
+ Metadata-Version: 2.4
2
+ Name: adaptq
3
+ Version: 0.1.0
4
+ Summary: Adaptive Streaming Vector Quantization KV Cache for LLMs
5
+ Author-email: "Lakshmikanthan K (letchupkt)" <letchupkt.dev@gmail.com>
6
+ Project-URL: Homepage, https://github.com/l3tchupkt/adaptq
7
+ Requires-Python: >=3.7
8
+ Description-Content-Type: text/markdown
9
+ Requires-Dist: numpy
10
+
11
+ # AdapTQ: Adaptive Streaming Vector Quantization
12
+
13
+ **AdapTQ** is a production-grade C++17 KV cache quantization engine for LLM inference on edge and memory-constrained systems.
14
+
15
+ **The Integration Pitch:** AdapTQ is an **optional KV-cache backend**. It runs entirely on the CPU, requires **no model changes**, and fits into existing inference pipelines with minimal adapter-style wrapper logic.
16
+
17
+ ## 🚀 Quickstart
18
+
19
+ ```python
20
+ import torch; from adaptq import AdaptQAttention
21
+ # 1. Initialize drop-in PyTorch wrapper (4-bit default)
22
+ layer = AdaptQAttention(dim=128, heads=4)
23
+ # 2. Forward pass dynamically routes continuous BxHxD generation tensors
24
+ out = layer(q=torch.randn(1, 4, 128), k=torch.randn(1, 4, 128), v=torch.randn(1, 4, 128))
25
+ ```
26
+
27
+ ## 📊 Real-world Benchmarks
28
+
29
+ Tested on standard AVX2 desktop hardware (4 heads, `dim=128`, caching up to 4096 tokens).
30
+ *Note: We ignore sequence lengths `< 256` in these claims, as short sequences are explicitly routed to standard FP32 execution via our hybrid fallback.*
31
+
32
+ | Metric | Result (Seq ≥ 256) |
33
+ | --- | --- |
34
+ | **Latencies** | p50: `877.9 µs` \| p95: `2161.8 µs` |
35
+ | **Stable Speedup** | ~10.18x vs NumPy FP32 equivalent |
36
+ | **Throughput** | ~1,139 tokens/sec |
37
+ | **Memory** | 2.10 MB vs FP16's 8.39 MB (**4.0x smaller**) |
38
+
39
+ ### Quantization Fidelity (Honest Metrics)
40
+ We use a targeted $\pm 3\sigma$ variance soft-clipping on FWHT distributions without altering Max-Lloyd codebooks. Our strictly measured empirical quality against baseline FP32:
41
+ - **Cosine Similarity**: ~0.947 (1.000 = exact identical match)
42
+ - **Mean Squared Error (MSE)**: ~1.8e-04
43
+
44
+ ![AdapTQ Benchmarks](adaptq_realtime_bench.png)
45
+
46
+ ## 🏗 Architecture & Features
47
+
48
+ - **Unified SIMD Pipeline**: 2, 3, and 4-bit decoding share a single, quad-unrolled branchless loop using AVX2 intrinsics. No scalar fallbacks in the hot path.
49
+ - **Fast Hadamard Rotation (HAR)**: $O(d \log d)$ fully in-place rotation minimizes outliers gracefully before codebook matching.
50
+ - **Zero Heap Allocations**: Pure stack/thread-local memory buffers in the hot path.
51
+ - **Pre-Compute LUTs**: Dot products execute directly against packed indices in SIMD registers—avoiding full dequantization inside the attention kernel.
52
+
53
+ ## 🛠 Installation & Integration
54
+
55
+ ```bash
56
+ git clone https://github.com/l3tchupkt/adaptq.git
57
+ cd adaptq
58
+
59
+ # Build python PyBind modules natively
60
+ pip install .
61
+ ```
62
+
63
+ ### Python Native Application
64
+
65
+ Use the native generic python API to bypass neural-network tensors explicitly:
66
+
67
+ ```python
68
+ import numpy as np
69
+ from adaptq import Engine
70
+
71
+ engine = Engine(dim=128, heads=4, bits=4, capacity=2048)
72
+ k, v, q = np.random.randn(4, 128), np.random.randn(4, 128), np.random.randn(4, 128)
73
+
74
+ engine.append(k, v)
75
+ output = engine.compute(q)
76
+ ```
77
+
78
+ ### llama.cpp Adapter
79
+
80
+ Using `AdapTQ` as the native KV Cache replacement during computation phase over GGML.
81
+ *(Requires using `llm_build_kqv` hooks. See `/integration/llama_cpp_patch.md` for full unified patch details.)*
82
+
83
+ ```cpp
84
+ #include "adapters/adapter_llamacpp.h"
85
+ LlamaCppAdaptQAdapter adapter(n_heads, head_dim, bits, capacity, seed, v_mass, hybrid_thr);
86
+ adapter.feed_kv(head, key_array, val_array, token_pos);
87
+ adapter.attention(head, query_array, out_array);
88
+ ```
89
+
90
+ ## 📝 License
91
+
92
+ See active repository license policies. Developed based on *AdapTQ: Adaptive Streaming Vector Quantization for Edge-Deployed Large Language Models*.
adaptq-0.1.0/README.md ADDED
@@ -0,0 +1,82 @@
1
+ # AdapTQ: Adaptive Streaming Vector Quantization
2
+
3
+ **AdapTQ** is a production-grade C++17 KV cache quantization engine for LLM inference on edge and memory-constrained systems.
4
+
5
+ **The Integration Pitch:** AdapTQ is an **optional KV-cache backend**. It runs entirely on the CPU, requires **no model changes**, and fits into existing inference pipelines with minimal adapter-style wrapper logic.
6
+
7
+ ## 🚀 Quickstart
8
+
9
+ ```python
10
+ import torch; from adaptq import AdaptQAttention
11
+ # 1. Initialize drop-in PyTorch wrapper (4-bit default)
12
+ layer = AdaptQAttention(dim=128, heads=4)
13
+ # 2. Forward pass dynamically routes continuous BxHxD generation tensors
14
+ out = layer(q=torch.randn(1, 4, 128), k=torch.randn(1, 4, 128), v=torch.randn(1, 4, 128))
15
+ ```
16
+
17
+ ## 📊 Real-world Benchmarks
18
+
19
+ Tested on standard AVX2 desktop hardware (4 heads, `dim=128`, caching up to 4096 tokens).
20
+ *Note: We ignore sequence lengths `< 256` in these claims, as short sequences are explicitly routed to standard FP32 execution via our hybrid fallback.*
21
+
22
+ | Metric | Result (Seq ≥ 256) |
23
+ | --- | --- |
24
+ | **Latencies** | p50: `877.9 µs` \| p95: `2161.8 µs` |
25
+ | **Stable Speedup** | ~10.18x vs NumPy FP32 equivalent |
26
+ | **Throughput** | ~1,139 tokens/sec |
27
+ | **Memory** | 2.10 MB vs FP16's 8.39 MB (**4.0x smaller**) |
28
+
29
+ ### Quantization Fidelity (Honest Metrics)
30
+ We use a targeted $\pm 3\sigma$ variance soft-clipping on FWHT distributions without altering Max-Lloyd codebooks. Our strictly measured empirical quality against baseline FP32:
31
+ - **Cosine Similarity**: ~0.947 (1.000 = exact identical match)
32
+ - **Mean Squared Error (MSE)**: ~1.8e-04
33
+
34
+ ![AdapTQ Benchmarks](adaptq_realtime_bench.png)
35
+
36
+ ## 🏗 Architecture & Features
37
+
38
+ - **Unified SIMD Pipeline**: 2, 3, and 4-bit decoding share a single, quad-unrolled branchless loop using AVX2 intrinsics. No scalar fallbacks in the hot path.
39
+ - **Fast Hadamard Rotation (HAR)**: $O(d \log d)$ fully in-place rotation minimizes outliers gracefully before codebook matching.
40
+ - **Zero Heap Allocations**: Pure stack/thread-local memory buffers in the hot path.
41
+ - **Pre-Compute LUTs**: Dot products execute directly against packed indices in SIMD registers—avoiding full dequantization inside the attention kernel.
42
+
43
+ ## 🛠 Installation & Integration
44
+
45
+ ```bash
46
+ git clone https://github.com/l3tchupkt/adaptq.git
47
+ cd adaptq
48
+
49
+ # Build python PyBind modules natively
50
+ pip install .
51
+ ```
52
+
53
+ ### Python Native Application
54
+
55
+ Use the native generic python API to bypass neural-network tensors explicitly:
56
+
57
+ ```python
58
+ import numpy as np
59
+ from adaptq import Engine
60
+
61
+ engine = Engine(dim=128, heads=4, bits=4, capacity=2048)
62
+ k, v, q = np.random.randn(4, 128), np.random.randn(4, 128), np.random.randn(4, 128)
63
+
64
+ engine.append(k, v)
65
+ output = engine.compute(q)
66
+ ```
67
+
68
+ ### llama.cpp Adapter
69
+
70
+ Using `AdapTQ` as the native KV Cache replacement during computation phase over GGML.
71
+ *(Requires using `llm_build_kqv` hooks. See `/integration/llama_cpp_patch.md` for full unified patch details.)*
72
+
73
+ ```cpp
74
+ #include "adapters/adapter_llamacpp.h"
75
+ LlamaCppAdaptQAdapter adapter(n_heads, head_dim, bits, capacity, seed, v_mass, hybrid_thr);
76
+ adapter.feed_kv(head, key_array, val_array, token_pos);
77
+ adapter.attention(head, query_array, out_array);
78
+ ```
79
+
80
+ ## 📝 License
81
+
82
+ See active repository license policies. Developed based on *AdapTQ: Adaptive Streaming Vector Quantization for Edge-Deployed Large Language Models*.
@@ -0,0 +1,38 @@
1
+ #pragma once
2
+ #include "../include/adaptq_mha_backend.h"
3
+ #include <cstddef>
4
+
5
+ /* -------------------------------------------------------------------------
6
+ * LlamaCppAdaptQAdapter
7
+ *
8
+ * Thin adapter between llama.cpp's float tensor data and the AdapTQ core.
9
+ * Include this file in llama.cpp builds. Zero GGML headers required here.
10
+ * ------------------------------------------------------------------------- */
11
+ class LlamaCppAdaptQAdapter {
12
+ public:
13
+ LlamaCppAdaptQAdapter(int n_heads, int head_dim, int bits, int capacity,
14
+ uint64_t seed = 0, float v_mass = 0.95f,
15
+ int hybrid_thresh = 512);
16
+
17
+ /** Call once per incoming token, once per head, after graph compute. */
18
+ void feed_kv(int head, const float *key, const float *val, int token_pos);
19
+
20
+ /** Single-query attention forward (decode step). */
21
+ int attention(int head, const float *query, float *out);
22
+
23
+ /** Multi-query batch attention (prefill step). */
24
+ int attention_batch(int head, const float *queries, int n_queries,
25
+ float *outs);
26
+
27
+ /** Clear all cached KV (e.g. on context reset). */
28
+ void reset();
29
+
30
+ /** Compressed bytes in the KV cache. */
31
+ size_t kv_bytes() const;
32
+
33
+ /** Logs "KV: X MB vs FP16: Y MB" to stderr. */
34
+ void log_memory(int capacity) const;
35
+
36
+ private:
37
+ AdaptQMHABackend _backend;
38
+ };
@@ -0,0 +1,93 @@
1
+ /* adapter_python.cpp
2
+ *
3
+ * Pluggable adapter: Python (pybind11) → AdapTQ core
4
+ *
5
+ * Build with:
6
+ * pip install pybind11
7
+ * g++ -O3 -march=native -mavx2 -shared -fPIC \
8
+ * $(python3 -m pybind11 --includes) \
9
+ * -Iinclude \
10
+ * adapters/adapter_python.cpp \
11
+ * -L. -ladaptq -Wl,-rpath,. \
12
+ * -o adaptq_py$(python3-config --extension-suffix)
13
+ *
14
+ * Python usage:
15
+ * import adaptq_py as aq
16
+ * ctx = aq.MHAContext(n_heads=32, head_dim=128, bits=4, capacity=8192)
17
+ * ctx.append(head=0, key=k_np, val=v_np, pos=0)
18
+ * out = ctx.compute(head=0, query=q_np)
19
+ * batch_out = ctx.compute_batch(head=0, queries=q_batch_np)
20
+ */
21
+
22
+ #include "../include/adaptq_mha_backend.h"
23
+ #include <pybind11/numpy.h>
24
+ #include <pybind11/pybind11.h>
25
+ #include <stdexcept>
26
+
27
+
28
+ namespace py = pybind11;
29
+
30
+ class PyMHAContext {
31
+ public:
32
+ PyMHAContext(int n_heads, int head_dim, int bits, int capacity,
33
+ uint64_t seed = 0, float v_mass = 0.95f, int hybrid_thresh = 512)
34
+ : _backend(n_heads, head_dim, bits, capacity, seed, v_mass,
35
+ hybrid_thresh),
36
+ _head_dim(head_dim), _n_heads(n_heads) {}
37
+
38
+ void append(int head, py::array_t<float> key, py::array_t<float> val,
39
+ int pos) {
40
+ auto k = key.unchecked<1>();
41
+ auto v = val.unchecked<1>();
42
+ if (k.shape(0) != _head_dim || v.shape(0) != _head_dim)
43
+ throw std::invalid_argument("key/val dim mismatch");
44
+ _backend.append_kv(head, k.data(0), v.data(0), pos);
45
+ }
46
+
47
+ py::array_t<float> compute(int head, py::array_t<float> query) {
48
+ auto q = query.unchecked<1>();
49
+ if (q.shape(0) != _head_dim)
50
+ throw std::invalid_argument("query dim mismatch");
51
+ auto out = py::array_t<float>(_head_dim);
52
+ _backend.compute(head, q.data(0), out.mutable_data(0));
53
+ return out;
54
+ }
55
+
56
+ py::array_t<float> compute_batch(int head, py::array_t<float> queries) {
57
+ auto q = queries.unchecked<2>();
58
+ int nq = (int)q.shape(0);
59
+ if (q.shape(1) != _head_dim)
60
+ throw std::invalid_argument("queries dim mismatch");
61
+ auto out = py::array_t<float>({nq, _head_dim});
62
+ _backend.compute_batch(head, q.data(0, 0), nq, out.mutable_data(0, 0));
63
+ return out;
64
+ }
65
+
66
+ void reset() { _backend.reset(); }
67
+ size_t kv_bytes() const { return _backend.kv_bytes(); }
68
+ int n_heads() const { return _n_heads; }
69
+ int head_dim() const { return _head_dim; }
70
+
71
+ private:
72
+ AdaptQMHABackend _backend;
73
+ int _head_dim, _n_heads;
74
+ };
75
+
76
+ PYBIND11_MODULE(adaptq_py, m) {
77
+ m.doc() = "AdapTQ — Quantized KV-cache attention, Python bindings";
78
+
79
+ py::class_<PyMHAContext>(m, "MHAContext")
80
+ .def(py::init<int, int, int, int, uint64_t, float, int>(),
81
+ py::arg("n_heads"), py::arg("head_dim"), py::arg("bits"),
82
+ py::arg("capacity"), py::arg("seed") = 0, py::arg("v_mass") = 0.95f,
83
+ py::arg("hybrid_thresh") = 512)
84
+ .def("append", &PyMHAContext::append, py::arg("head"), py::arg("key"),
85
+ py::arg("val"), py::arg("pos"))
86
+ .def("compute", &PyMHAContext::compute, py::arg("head"), py::arg("query"))
87
+ .def("compute_batch", &PyMHAContext::compute_batch, py::arg("head"),
88
+ py::arg("queries"))
89
+ .def("reset", &PyMHAContext::reset)
90
+ .def("kv_bytes", &PyMHAContext::kv_bytes)
91
+ .def_property_readonly("n_heads", &PyMHAContext::n_heads)
92
+ .def_property_readonly("head_dim", &PyMHAContext::head_dim);
93
+ }
@@ -0,0 +1,4 @@
1
+ from .core import Engine
2
+ from .torch_adapter import AdaptQAttention
3
+
4
+ __all__ = ["Engine", "AdaptQAttention"]
@@ -0,0 +1,73 @@
1
+ import numpy as np
2
+ try:
3
+ import adaptq_py
4
+ except ImportError:
5
+ raise ImportError("AdapTQ C++ extension not built. Run 'pip install .' to build from source.")
6
+
7
+ class Engine:
8
+ """
9
+ API for Adaptive Streaming Vector Quantization.
10
+ Usage:
11
+ engine = adaptq.Engine(dim=128, heads=4, bits=4)
12
+ engine.append(k, v)
13
+ out = engine.compute(q)
14
+ """
15
+ def __init__(self, dim: int, heads: int, bits: int = 4,
16
+ capacity: int = 4096, seed: int = 42,
17
+ v_mass: float = 0.95, hybrid_thresh: int = 512):
18
+ self.dim = dim
19
+ self.heads = heads
20
+ self.bits = bits
21
+ self.capacity = capacity
22
+
23
+ # Internal state
24
+ self._ctx = adaptq_py.MHAContext(
25
+ n_heads=heads,
26
+ head_dim=dim,
27
+ bits=bits,
28
+ capacity=capacity,
29
+ seed=seed,
30
+ v_mass=v_mass,
31
+ hybrid_thresh=hybrid_thresh
32
+ )
33
+ self.pos = 0
34
+
35
+ def append(self, k: np.ndarray, v: np.ndarray):
36
+ """
37
+ Append k, v for the current step.
38
+ Supports inputs of shape (heads, dim).
39
+ """
40
+ if k.shape != (self.heads, self.dim) or v.shape != (self.heads, self.dim):
41
+ raise ValueError(f"k and v must be shape ({self.heads}, {self.dim})")
42
+
43
+ for h in range(self.heads):
44
+ # Using astype to guarantee contiguous float32 alignment to C++ backend
45
+ k_h = np.ascontiguousarray(k[h], dtype=np.float32)
46
+ v_h = np.ascontiguousarray(v[h], dtype=np.float32)
47
+ self._ctx.append(h, k_h, v_h, self.pos)
48
+
49
+ self.pos += 1
50
+
51
+ def compute(self, q: np.ndarray) -> np.ndarray:
52
+ """
53
+ Compute attention output for query q.
54
+ Supports inputs of shape (heads, dim). Returns shape (heads, dim).
55
+ """
56
+ if q.shape != (self.heads, self.dim):
57
+ raise ValueError(f"q must be shape ({self.heads}, {self.dim})")
58
+
59
+ out = np.zeros((self.heads, self.dim), dtype=np.float32)
60
+ for h in range(self.heads):
61
+ q_h = np.ascontiguousarray(q[h], dtype=np.float32)
62
+ out[h] = self._ctx.compute(h, q_h)
63
+
64
+ return out
65
+
66
+ def reset(self):
67
+ """Clear the KV cache."""
68
+ self._ctx.reset()
69
+ self.pos = 0
70
+
71
+ @property
72
+ def kv_bytes(self) -> int:
73
+ return self._ctx.kv_bytes()
@@ -0,0 +1,48 @@
1
+ try:
2
+ import torch
3
+ import torch.nn as nn
4
+ except ImportError:
5
+ raise ImportError("PyTorch is required to use the AdaptQ PyTorch wrapper.")
6
+
7
+ from .core import Engine
8
+ import numpy as np
9
+
10
+ class AdaptQAttention(nn.Module):
11
+ """
12
+ PyTorch Wrapper for AdapTQ.
13
+ Integrates directly with PyTorch's nn.Module API.
14
+ """
15
+ def __init__(self, dim: int, heads: int, bits: int = 4, capacity: int = 4096, seed: int = 42):
16
+ super().__init__()
17
+ self.dim = dim
18
+ self.heads = heads
19
+ self.bits = bits
20
+ self.engine = Engine(dim=dim, heads=heads, bits=bits, capacity=capacity, seed=seed)
21
+
22
+ def forward(self, q: torch.Tensor, k: torch.Tensor, v: torch.Tensor) -> torch.Tensor:
23
+ """
24
+ Expects inputs of shape (batch_size, heads, dim).
25
+ Note: Currently AdapTQ handles a batch_size of 1 for append_kv internally,
26
+ but we can loop over the batch size or just assert batch_size == 1 for generation step.
27
+ """
28
+ assert q.dim() == 3 and k.dim() == 3 and v.dim() == 3, "Expected 3D tensors: (batch, heads, dim)"
29
+ batch_size = q.size(0)
30
+
31
+ # Move to CPU numpy for backend (AdapTQ backend is pure CPU C++)
32
+ k_np = k.detach().cpu().numpy()
33
+ v_np = v.detach().cpu().numpy()
34
+ q_np = q.detach().cpu().numpy()
35
+
36
+ out = np.zeros_like(q_np)
37
+
38
+ # Currently, AdapTQ C++ context is single-instance per `Engine`.
39
+ # For batch inference, context is strictly sequential in Generation step.
40
+ for b in range(batch_size):
41
+ self.engine.append(k_np[b], v_np[b])
42
+ out[b] = self.engine.compute(q_np[b])
43
+
44
+ # Returning tensor on the same device as query
45
+ return torch.from_numpy(out).to(q.device)
46
+
47
+ def reset_cache(self):
48
+ self.engine.reset()
@@ -0,0 +1,92 @@
1
+ Metadata-Version: 2.4
2
+ Name: adaptq
3
+ Version: 0.1.0
4
+ Summary: Adaptive Streaming Vector Quantization KV Cache for LLMs
5
+ Author-email: "Lakshmikanthan K (letchupkt)" <letchupkt.dev@gmail.com>
6
+ Project-URL: Homepage, https://github.com/l3tchupkt/adaptq
7
+ Requires-Python: >=3.7
8
+ Description-Content-Type: text/markdown
9
+ Requires-Dist: numpy
10
+
11
+ # AdapTQ: Adaptive Streaming Vector Quantization
12
+
13
+ **AdapTQ** is a production-grade C++17 KV cache quantization engine for LLM inference on edge and memory-constrained systems.
14
+
15
+ **The Integration Pitch:** AdapTQ is an **optional KV-cache backend**. It runs entirely on the CPU, requires **no model changes**, and fits into existing inference pipelines with minimal adapter-style wrapper logic.
16
+
17
+ ## 🚀 Quickstart
18
+
19
+ ```python
20
+ import torch; from adaptq import AdaptQAttention
21
+ # 1. Initialize drop-in PyTorch wrapper (4-bit default)
22
+ layer = AdaptQAttention(dim=128, heads=4)
23
+ # 2. Forward pass dynamically routes continuous BxHxD generation tensors
24
+ out = layer(q=torch.randn(1, 4, 128), k=torch.randn(1, 4, 128), v=torch.randn(1, 4, 128))
25
+ ```
26
+
27
+ ## 📊 Real-world Benchmarks
28
+
29
+ Tested on standard AVX2 desktop hardware (4 heads, `dim=128`, caching up to 4096 tokens).
30
+ *Note: We ignore sequence lengths `< 256` in these claims, as short sequences are explicitly routed to standard FP32 execution via our hybrid fallback.*
31
+
32
+ | Metric | Result (Seq ≥ 256) |
33
+ | --- | --- |
34
+ | **Latencies** | p50: `877.9 µs` \| p95: `2161.8 µs` |
35
+ | **Stable Speedup** | ~10.18x vs NumPy FP32 equivalent |
36
+ | **Throughput** | ~1,139 tokens/sec |
37
+ | **Memory** | 2.10 MB vs FP16's 8.39 MB (**4.0x smaller**) |
38
+
39
+ ### Quantization Fidelity (Honest Metrics)
40
+ We use a targeted $\pm 3\sigma$ variance soft-clipping on FWHT distributions without altering Max-Lloyd codebooks. Our strictly measured empirical quality against baseline FP32:
41
+ - **Cosine Similarity**: ~0.947 (1.000 = exact identical match)
42
+ - **Mean Squared Error (MSE)**: ~1.8e-04
43
+
44
+ ![AdapTQ Benchmarks](adaptq_realtime_bench.png)
45
+
46
+ ## 🏗 Architecture & Features
47
+
48
+ - **Unified SIMD Pipeline**: 2, 3, and 4-bit decoding share a single, quad-unrolled branchless loop using AVX2 intrinsics. No scalar fallbacks in the hot path.
49
+ - **Fast Hadamard Rotation (HAR)**: $O(d \log d)$ fully in-place rotation minimizes outliers gracefully before codebook matching.
50
+ - **Zero Heap Allocations**: Pure stack/thread-local memory buffers in the hot path.
51
+ - **Pre-Compute LUTs**: Dot products execute directly against packed indices in SIMD registers—avoiding full dequantization inside the attention kernel.
52
+
53
+ ## 🛠 Installation & Integration
54
+
55
+ ```bash
56
+ git clone https://github.com/l3tchupkt/adaptq.git
57
+ cd adaptq
58
+
59
+ # Build python PyBind modules natively
60
+ pip install .
61
+ ```
62
+
63
+ ### Python Native Application
64
+
65
+ Use the native generic python API to bypass neural-network tensors explicitly:
66
+
67
+ ```python
68
+ import numpy as np
69
+ from adaptq import Engine
70
+
71
+ engine = Engine(dim=128, heads=4, bits=4, capacity=2048)
72
+ k, v, q = np.random.randn(4, 128), np.random.randn(4, 128), np.random.randn(4, 128)
73
+
74
+ engine.append(k, v)
75
+ output = engine.compute(q)
76
+ ```
77
+
78
+ ### llama.cpp Adapter
79
+
80
+ Using `AdapTQ` as the native KV Cache replacement during computation phase over GGML.
81
+ *(Requires using `llm_build_kqv` hooks. See `/integration/llama_cpp_patch.md` for full unified patch details.)*
82
+
83
+ ```cpp
84
+ #include "adapters/adapter_llamacpp.h"
85
+ LlamaCppAdaptQAdapter adapter(n_heads, head_dim, bits, capacity, seed, v_mass, hybrid_thr);
86
+ adapter.feed_kv(head, key_array, val_array, token_pos);
87
+ adapter.attention(head, query_array, out_array);
88
+ ```
89
+
90
+ ## 📝 License
91
+
92
+ See active repository license policies. Developed based on *AdapTQ: Adaptive Streaming Vector Quantization for Edge-Deployed Large Language Models*.
@@ -0,0 +1,33 @@
1
+ MANIFEST.in
2
+ README.md
3
+ pyproject.toml
4
+ setup.py
5
+ adapters/adapter_llamacpp.h
6
+ adapters/adapter_python.cpp
7
+ adaptq/__init__.py
8
+ adaptq/core.py
9
+ adaptq/torch_adapter.py
10
+ adaptq.egg-info/PKG-INFO
11
+ adaptq.egg-info/SOURCES.txt
12
+ adaptq.egg-info/dependency_links.txt
13
+ adaptq.egg-info/not-zip-safe
14
+ adaptq.egg-info/requires.txt
15
+ adaptq.egg-info/top_level.txt
16
+ attention/attention.cpp
17
+ cache/ring_buffer.cpp
18
+ core/adaptq_backend_vtable.cpp
19
+ core/adaptq_c_api.cpp
20
+ core/codebook.cpp
21
+ core/fwht.cpp
22
+ core/quantizer.cpp
23
+ include/adaptq.h
24
+ include/adaptq_backend.h
25
+ include/adaptq_mha_backend.h
26
+ include/attention.h
27
+ include/codebook.h
28
+ include/fwht.h
29
+ include/quantizer.h
30
+ include/ring_buffer.h
31
+ include/timer.h
32
+ tests/test_accuracy.py
33
+ utils/timer.cpp
@@ -0,0 +1 @@
1
+
@@ -0,0 +1 @@
1
+ numpy
@@ -0,0 +1,2 @@
1
+ adaptq
2
+ adaptq_py