ramp-llm 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- ramp_llm-0.1.0/LICENSE +21 -0
- ramp_llm-0.1.0/PKG-INFO +375 -0
- ramp_llm-0.1.0/README.md +337 -0
- ramp_llm-0.1.0/pyproject.toml +66 -0
- ramp_llm-0.1.0/setup.cfg +4 -0
- ramp_llm-0.1.0/src/ramp/__init__.py +10 -0
- ramp_llm-0.1.0/src/ramp/__main__.py +276 -0
- ramp_llm-0.1.0/src/ramp/autoconfig.py +167 -0
- ramp_llm-0.1.0/src/ramp/backend.py +280 -0
- ramp_llm-0.1.0/src/ramp/config.py +181 -0
- ramp_llm-0.1.0/src/ramp/controller.py +265 -0
- ramp_llm-0.1.0/src/ramp/doctor.py +194 -0
- ramp_llm-0.1.0/src/ramp/metrics.py +174 -0
- ramp_llm-0.1.0/src/ramp/mock_llm.py +124 -0
- ramp_llm-0.1.0/src/ramp/monitor.py +177 -0
- ramp_llm-0.1.0/src/ramp/policy.py +187 -0
- ramp_llm-0.1.0/src/ramp/server.py +223 -0
- ramp_llm-0.1.0/src/ramp_llm.egg-info/PKG-INFO +375 -0
- ramp_llm-0.1.0/src/ramp_llm.egg-info/SOURCES.txt +25 -0
- ramp_llm-0.1.0/src/ramp_llm.egg-info/dependency_links.txt +1 -0
- ramp_llm-0.1.0/src/ramp_llm.egg-info/entry_points.txt +2 -0
- ramp_llm-0.1.0/src/ramp_llm.egg-info/requires.txt +10 -0
- ramp_llm-0.1.0/src/ramp_llm.egg-info/top_level.txt +1 -0
- ramp_llm-0.1.0/tests/test_autoconfig.py +130 -0
- ramp_llm-0.1.0/tests/test_config.py +56 -0
- ramp_llm-0.1.0/tests/test_integration.py +388 -0
- ramp_llm-0.1.0/tests/test_policy.py +248 -0
ramp_llm-0.1.0/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Shivapreetham
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
ramp_llm-0.1.0/PKG-INFO
ADDED
|
@@ -0,0 +1,375 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: ramp-llm
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: RAMP: Resource-Aware Model Proxy - an elastic local LLM daemon that auto-calibrates model size to available RAM, VRAM and disk
|
|
5
|
+
Author: Shivapreetham
|
|
6
|
+
License: MIT
|
|
7
|
+
Project-URL: Homepage, https://github.com/shivapreetham/resource-aware-model-proxy
|
|
8
|
+
Project-URL: Repository, https://github.com/shivapreetham/resource-aware-model-proxy
|
|
9
|
+
Project-URL: Issues, https://github.com/shivapreetham/resource-aware-model-proxy/issues
|
|
10
|
+
Project-URL: Changelog, https://github.com/shivapreetham/resource-aware-model-proxy/blob/main/CHANGELOG.md
|
|
11
|
+
Project-URL: Concepts, https://github.com/shivapreetham/resource-aware-model-proxy/blob/main/docs/CONCEPTS.md
|
|
12
|
+
Keywords: llm,local-llm,llama.cpp,ollama,elastic,quantization,proxy,openai-api,memory-management,inference
|
|
13
|
+
Classifier: Development Status :: 3 - Alpha
|
|
14
|
+
Classifier: Intended Audience :: Developers
|
|
15
|
+
Classifier: License :: OSI Approved :: MIT License
|
|
16
|
+
Classifier: Operating System :: OS Independent
|
|
17
|
+
Classifier: Programming Language :: Python :: 3
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
19
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
20
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
21
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
22
|
+
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
|
|
23
|
+
Classifier: Topic :: System :: Monitoring
|
|
24
|
+
Classifier: Topic :: Internet :: Proxy Servers
|
|
25
|
+
Requires-Python: >=3.10
|
|
26
|
+
Description-Content-Type: text/markdown
|
|
27
|
+
License-File: LICENSE
|
|
28
|
+
Requires-Dist: fastapi>=0.110
|
|
29
|
+
Requires-Dist: uvicorn>=0.29
|
|
30
|
+
Requires-Dist: httpx>=0.27
|
|
31
|
+
Requires-Dist: psutil>=5.9
|
|
32
|
+
Requires-Dist: PyYAML>=6.0
|
|
33
|
+
Provides-Extra: dev
|
|
34
|
+
Requires-Dist: pytest>=8; extra == "dev"
|
|
35
|
+
Requires-Dist: pytest-asyncio>=0.23; extra == "dev"
|
|
36
|
+
Requires-Dist: ruff>=0.6; extra == "dev"
|
|
37
|
+
Dynamic: license-file
|
|
38
|
+
|
|
39
|
+
# RAMP — Resource-Aware Model Proxy
|
|
40
|
+
|
|
41
|
+
**Your local LLM should get out of the way when you need the RAM back.**
|
|
42
|
+
|
|
43
|
+
[](https://github.com/shivapreetham/resource-aware-model-proxy/actions/workflows/ci.yml)
|
|
44
|
+
[](LICENSE)
|
|
45
|
+
[](https://www.python.org/downloads/)
|
|
46
|
+
|
|
47
|
+
RAMP is an elastic daemon that sits in front of llama.cpp or Ollama. It
|
|
48
|
+
watches RAM, VRAM, and disk continuously, and moves the loaded model up and
|
|
49
|
+
down a *quality ladder* as your machine gets busy and quiets down:
|
|
50
|
+
|
|
51
|
+
- You open Chrome with 40 tabs → RAMP swaps your 7B model for a 3B one.
|
|
52
|
+
- You close them → after a stable recovery window, the 7B comes back.
|
|
53
|
+
- Your assistant never dies with an OOM, and you never reconfigure anything.
|
|
54
|
+
|
|
55
|
+
It speaks the OpenAI API, so **Open-LLM-VTuber, Khoj, SillyTavern, LangChain,
|
|
56
|
+
Continue, or `curl` just point `base_url` at RAMP and work unchanged** — no
|
|
57
|
+
code, no client plugin, nothing to adopt. Swaps cost ~2 seconds warm
|
|
58
|
+
(measured), and in-flight requests are drained rather than dropped.
|
|
59
|
+
|
|
60
|
+
```diff
|
|
61
|
+
- client = OpenAI(base_url="http://localhost:11434/v1") # straight to Ollama
|
|
62
|
+
+ client = OpenAI(base_url="http://localhost:8090/v1") # through RAMP
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
That's the entire integration. Verified against the official `openai` Python
|
|
66
|
+
SDK: model listing, streaming, and hard-coded model names all work untouched.
|
|
67
|
+
Ollama's native `/api/*` routes are proxied too, for tools that use those.
|
|
68
|
+
|
|
69
|
+
> **New here?** [docs/CONCEPTS.md](docs/CONCEPTS.md) explains what's actually
|
|
70
|
+
> going on: what consumes memory when an LLM runs, why VRAM pressure silently
|
|
71
|
+
> becomes RAM pressure, and the control-theory ideas (hysteresis, damping)
|
|
72
|
+
> behind the policy.
|
|
73
|
+
|
|
74
|
+
## Why this doesn't already exist
|
|
75
|
+
|
|
76
|
+
Everyone running a local model has made the same bad trade: pick a big model
|
|
77
|
+
and let the machine choke, or pick a small one and pay for the worst case all
|
|
78
|
+
day. The tools make you choose **once, up front, forever**.
|
|
79
|
+
|
|
80
|
+
- **Ollama / LM Studio** pick a quantization *once*, at download/load time.
|
|
81
|
+
Nothing adapts after that ([ollama#14674](https://github.com/ollama/ollama/issues/14674)).
|
|
82
|
+
- **[llama-swap](https://github.com/mostlygeek/llama-swap)** swaps models based
|
|
83
|
+
on *which model the client requests* — the system's resource state is never
|
|
84
|
+
consulted.
|
|
85
|
+
- **FlexQuant / LSAQ / Voltron / Any-Precision LLM** validated elastic
|
|
86
|
+
execution academically; none ship as a usable daemon.
|
|
87
|
+
|
|
88
|
+
RAMP is the missing controller: policy-driven, hysteresis-damped,
|
|
89
|
+
backend-agnostic, and measured — it reports its own swap rate so you can prove
|
|
90
|
+
it isn't thrashing.
|
|
91
|
+
|
|
92
|
+
## Architecture
|
|
93
|
+
|
|
94
|
+
```mermaid
|
|
95
|
+
flowchart LR
|
|
96
|
+
C[Any OpenAI client] -->|/v1/*| P[Proxy - FastAPI]
|
|
97
|
+
P -->|forward + x-ramp-tier| B[llama-server child process]
|
|
98
|
+
M[Memory monitor - psutil] -->|raw + EMA available MB| PE[Policy engine]
|
|
99
|
+
PE -->|stay / switch / unload| CT[Controller]
|
|
100
|
+
CT -->|drain, stop, start, health-check| B
|
|
101
|
+
CT -->|gate requests during swap| P
|
|
102
|
+
```
|
|
103
|
+
|
|
104
|
+
| Module | Role |
|
|
105
|
+
|---|---|
|
|
106
|
+
| `monitor.py` | Samples available RAM, GPU/VRAM (NVIDIA, via `nvidia-smi`), and free disk — each with raw + smoothed readings. |
|
|
107
|
+
| `policy.py` | Pure state machine: fast downgrades under pressure, slow damped upgrades, critical-floor emergency handling. Fully unit-tested, no I/O. |
|
|
108
|
+
| `controller.py` | Executes decisions: drain in-flight requests → stop old backend → start new → reopen the gate. Restarts crashed backends. Records an event log. |
|
|
109
|
+
| `backend.py` | Child-process lifecycle for `llama-server` (real) or `ramp.mock_llm` (demo/tests). |
|
|
110
|
+
| `server.py` | OpenAI-compatible passthrough (incl. SSE streaming) + control API. |
|
|
111
|
+
|
|
112
|
+
### The decision policy
|
|
113
|
+
|
|
114
|
+
Every tier declares its footprint (`est_ram_mb`, and `est_vram_mb` when it
|
|
115
|
+
runs on the GPU), and the policy holds **RAM, VRAM, and disk** accountable:
|
|
116
|
+
|
|
117
|
+
- **Downgrade (fast):** free RAM below `safety_margin_mb` — or, for a
|
|
118
|
+
GPU-resident tier, free VRAM below `vram_safety_margin_mb` — for
|
|
119
|
+
`downgrade_after_samples` consecutive polls → switch to the largest smaller
|
|
120
|
+
tier that fits *every* projected post-swap budget. A GPU-starved ladder can
|
|
121
|
+
land on a CPU-only tier (`est_vram_mb: 0`).
|
|
122
|
+
- **Critical (instant):** free RAM below `critical_free_mb` → act immediately;
|
|
123
|
+
if even the smallest tier can't fit, unload entirely (requests queue/503
|
|
124
|
+
until memory returns).
|
|
125
|
+
- **Upgrade (slow):** a bigger tier must fit both RAM and VRAM budgets with
|
|
126
|
+
*extra* headroom (`upgrade_extra_mb` / `vram_upgrade_extra_mb`), sustained
|
|
127
|
+
for `upgrade_after_s` seconds, before it's loaded — so RAMP never thrashes
|
|
128
|
+
when you alt-tab.
|
|
129
|
+
- **Disk (gate):** free disk below `disk_min_free_mb` blocks upgrades — a
|
|
130
|
+
bigger model can't free disk, so it gates rather than downgrades — and is
|
|
131
|
+
flagged in `/ramp/status` (`disk.low: true`).
|
|
132
|
+
- No NVIDIA GPU (or no `nvidia-smi`)? VRAM constraints are simply not
|
|
133
|
+
enforced; everything else works unchanged.
|
|
134
|
+
- Swaps happen **between** requests: in-flight generations are drained first
|
|
135
|
+
(up to `drain_timeout_s`), and incoming requests wait on the swap gate
|
|
136
|
+
instead of failing.
|
|
137
|
+
|
|
138
|
+
## Install
|
|
139
|
+
|
|
140
|
+
RAMP is a Python CLI, so it installs like one. No config file needed — it
|
|
141
|
+
inspects your machine and your installed Ollama models and builds a ladder
|
|
142
|
+
itself.
|
|
143
|
+
|
|
144
|
+
```bash
|
|
145
|
+
uvx ramp-llm # run without installing anything (needs uv)
|
|
146
|
+
pipx install ramp-llm # or install it properly
|
|
147
|
+
pip install ramp-llm # or into the current environment
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
<details>
|
|
151
|
+
<summary>Docker</summary>
|
|
152
|
+
|
|
153
|
+
```bash
|
|
154
|
+
docker build -t ramp .
|
|
155
|
+
docker run --rm -p 8090:8090 \
|
|
156
|
+
--add-host=host.docker.internal:host-gateway \
|
|
157
|
+
-e RAMP_OLLAMA_URL=http://host.docker.internal:11434 \
|
|
158
|
+
ramp
|
|
159
|
+
```
|
|
160
|
+
|
|
161
|
+
Add `--gpus all` (with the NVIDIA container toolkit) for VRAM awareness. Note
|
|
162
|
+
that inside a container RAMP reads the *container's* memory limit — which is
|
|
163
|
+
usually what you want, but means `--memory` shapes its decisions.
|
|
164
|
+
</details>
|
|
165
|
+
|
|
166
|
+
Then:
|
|
167
|
+
|
|
168
|
+
```bash
|
|
169
|
+
ramp doctor # check this machine can run it, and what to fix if not
|
|
170
|
+
ramp # start, with an auto-detected ladder
|
|
171
|
+
ramp status # what's loaded, why, and what it's costing
|
|
172
|
+
```
|
|
173
|
+
|
|
174
|
+
`ramp` prints the endpoint to point your tools at. That's the whole setup.
|
|
175
|
+
|
|
176
|
+
<details>
|
|
177
|
+
<summary>Other commands</summary>
|
|
178
|
+
|
|
179
|
+
```bash
|
|
180
|
+
ramp init # write the auto-detected ladder to ramp.yaml to tune
|
|
181
|
+
ramp run -c ramp.yaml # start from an explicit config
|
|
182
|
+
ramp status --json # raw status for scripting
|
|
183
|
+
ramp --help
|
|
184
|
+
```
|
|
185
|
+
|
|
186
|
+
`ramp run` with no `-c` uses `./ramp.yaml` if present, otherwise auto-detects.
|
|
187
|
+
</details>
|
|
188
|
+
|
|
189
|
+
## Try it without any models
|
|
190
|
+
|
|
191
|
+
The `mock` backend runs the whole daemon against fake OpenAI servers that tag
|
|
192
|
+
their replies with the tier name — so you can watch the laddering behaviour
|
|
193
|
+
in about thirty seconds, with nothing to download.
|
|
194
|
+
|
|
195
|
+
The example configs live in this repo, so grab one first if you installed
|
|
196
|
+
from PyPI:
|
|
197
|
+
|
|
198
|
+
```bash
|
|
199
|
+
curl -O https://raw.githubusercontent.com/shivapreetham/resource-aware-model-proxy/main/examples/ramp.mock.yaml
|
|
200
|
+
curl -O https://raw.githubusercontent.com/shivapreetham/resource-aware-model-proxy/main/scripts/stress_ram.py
|
|
201
|
+
ramp run -c ramp.mock.yaml
|
|
202
|
+
```
|
|
203
|
+
|
|
204
|
+
Then, in another terminal:
|
|
205
|
+
|
|
206
|
+
```bash
|
|
207
|
+
curl http://127.0.0.1:8090/ramp/status
|
|
208
|
+
python stress_ram.py --mb 6000 --hold-s 60
|
|
209
|
+
```
|
|
210
|
+
|
|
211
|
+
Watch the reply tag flip from `[large-mock]` to a smaller tier while the
|
|
212
|
+
stress script runs, and back ~20s after it releases. Every response carries
|
|
213
|
+
an `x-ramp-tier` header naming the tier that produced it.
|
|
214
|
+
|
|
215
|
+
## Real models (Ollama backend)
|
|
216
|
+
|
|
217
|
+
If you already run [Ollama](https://ollama.com) (or your OS blocks unsigned
|
|
218
|
+
binaries — e.g. Windows Smart App Control), use `backend: ollama`
|
|
219
|
+
(see [examples/ramp.ollama.yaml](examples/ramp.ollama.yaml)). Tiers reference
|
|
220
|
+
Ollama model tags; RAMP loads/unloads them via `keep_alive`, rewrites each
|
|
221
|
+
request's `model` field to the active tier, and proxies to Ollama's OpenAI
|
|
222
|
+
API. Import your own GGUFs with `ollama create <tag> -f Modelfile`
|
|
223
|
+
(`FROM ./model.gguf`). If no Ollama server is running, RAMP spawns one.
|
|
224
|
+
|
|
225
|
+
## Real models (llama.cpp)
|
|
226
|
+
|
|
227
|
+
1. Install [llama.cpp](https://github.com/ggml-org/llama.cpp) so `llama-server`
|
|
228
|
+
is on your PATH (or set `llama_server_bin` to its full path).
|
|
229
|
+
2. Download 2–3 GGUF models of different sizes (e.g. Qwen2.5 7B/3B/0.5B
|
|
230
|
+
instruct, Q4_K_M).
|
|
231
|
+
3. Edit `examples/ramp.yaml` — paths, `ctx`, and `est_ram_mb` per tier.
|
|
232
|
+
To calibrate `est_ram_mb`: pin the tier, look at the llama-server process
|
|
233
|
+
in Task Manager, round up.
|
|
234
|
+
4. `ramp -c examples/ramp.yaml`, then point any OpenAI client at
|
|
235
|
+
`http://127.0.0.1:8090/v1`.
|
|
236
|
+
|
|
237
|
+
## Control API
|
|
238
|
+
|
|
239
|
+
| Endpoint | Purpose |
|
|
240
|
+
|---|---|
|
|
241
|
+
| `GET /ramp/status` | Current tier, RAM/GPU/disk readings, `last_decision` (why RAMP is holding — e.g. `disk-low`, `upgrade-pending`), tier ladder, metrics, and event log. |
|
|
242
|
+
| `GET /ramp/metrics` | Prometheus text format for scraping. |
|
|
243
|
+
| `POST /ramp/pin/{name}` | Force a tier; disables auto-calibration. |
|
|
244
|
+
| `DELETE /ramp/pin` | Resume auto-calibration. |
|
|
245
|
+
|
|
246
|
+
## Monitoring
|
|
247
|
+
|
|
248
|
+
An elastic daemon lives or dies on one number: **how often it actually
|
|
249
|
+
swaps.** RAMP measures its own behaviour rather than asking you to trust it —
|
|
250
|
+
swap rate, time lost to swapping, per-tier occupancy, cooldown suppressions,
|
|
251
|
+
and requests that waited or were rejected.
|
|
252
|
+
|
|
253
|
+
```bash
|
|
254
|
+
curl -s http://127.0.0.1:8090/ramp/metrics
|
|
255
|
+
prometheus --config.file=examples/monitoring/prometheus.yml
|
|
256
|
+
```
|
|
257
|
+
|
|
258
|
+
```promql
|
|
259
|
+
sum(rate(ramp_swaps_total[15m])) * 3600 # swaps per hour — the headline
|
|
260
|
+
rate(ramp_swap_seconds_total[30m]) # fraction of time spent swapping
|
|
261
|
+
rate(ramp_tier_seconds_total[1h]) # is your ladder calibrated?
|
|
262
|
+
```
|
|
263
|
+
|
|
264
|
+
Under 2 swaps/hour means RAMP is invisible, which is the goal; above ~12 means
|
|
265
|
+
churn worth tuning. [docs/MONITORING.md](docs/MONITORING.md) explains each
|
|
266
|
+
number, what to alert on, and how to tune from what you see. Alert rules ship
|
|
267
|
+
in [examples/monitoring/alerts.yml](examples/monitoring/alerts.yml).
|
|
268
|
+
|
|
269
|
+
### What RAMP itself costs
|
|
270
|
+
|
|
271
|
+
Fair question for any watchdog: is the watcher eating the memory it claims to
|
|
272
|
+
save? Measured — **~65 MB resident, and flat** (+1.9 MB across 300 requests,
|
|
273
|
+
200 metric polls and 12 swaps; 0.9 s of CPU for the whole run).
|
|
274
|
+
|
|
275
|
+
Two things make that a non-issue. RAMP **never budgets memory it is itself
|
|
276
|
+
using** — `virtual_memory().available` already excludes the daemon's own
|
|
277
|
+
footprint, so the policy reasons about genuinely free memory. And 65 MB is
|
|
278
|
+
roughly 1–3% of a single tier, which is measured in gigabytes.
|
|
279
|
+
|
|
280
|
+
You don't have to take that on faith: the daemon reports its own footprint in
|
|
281
|
+
`/ramp/status` (a `self` block) and as `ramp_self_rss_bytes` /
|
|
282
|
+
`ramp_backend_rss_bytes` in Prometheus. Check it on your machine.
|
|
283
|
+
|
|
284
|
+
## Tests
|
|
285
|
+
|
|
286
|
+
```bash
|
|
287
|
+
pytest
|
|
288
|
+
```
|
|
289
|
+
|
|
290
|
+
Unit tests cover the policy state machine exhaustively (RAM pressure, VRAM
|
|
291
|
+
pressure, the disk gate, hysteresis in both directions); integration tests
|
|
292
|
+
run the full daemon against real child processes with scripted resource
|
|
293
|
+
readings, driving the ladder down, up, through critical unload, streaming,
|
|
294
|
+
manual pinning, and a swap/poll race regression.
|
|
295
|
+
|
|
296
|
+
To reproduce the resource behaviours against real models, see
|
|
297
|
+
[examples/ramp.ollama.yaml](examples/ramp.ollama.yaml) (normal ladder) and
|
|
298
|
+
[examples/ramp.vram-test.yaml](examples/ramp.vram-test.yaml) (isolates VRAM
|
|
299
|
+
pressure: occupy the GPU with another model and watch RAMP fall back to a
|
|
300
|
+
CPU-only tier while system RAM stays healthy).
|
|
301
|
+
|
|
302
|
+
## Roadmap
|
|
303
|
+
|
|
304
|
+
- **Context-length scaling** — shrink the KV cache before swapping models
|
|
305
|
+
(cheaper first response to pressure).
|
|
306
|
+
- **Any-Precision backend** — one weight file servable at 3/4/8-bit
|
|
307
|
+
([paper](https://arxiv.org/abs/2402.10517)) makes downgrades near-free.
|
|
308
|
+
- **OS pressure signals** — Windows memory notifications / Linux PSI instead
|
|
309
|
+
of pure polling.
|
|
310
|
+
- **Multi-GPU and non-NVIDIA GPUs** — the monitor currently reads the first
|
|
311
|
+
NVIDIA GPU via `nvidia-smi`.
|
|
312
|
+
|
|
313
|
+
## Prior art & acknowledgements
|
|
314
|
+
|
|
315
|
+
RAMP didn't invent elastic inference. It productizes an idea that several
|
|
316
|
+
research groups established and that nobody shipped as a usable daemon.
|
|
317
|
+
Credit where it's due.
|
|
318
|
+
|
|
319
|
+
**Research that established the idea**
|
|
320
|
+
|
|
321
|
+
- **[FlexQuant](https://arxiv.org/abs/2501.07139)** (Chai et al., 2025) —
|
|
322
|
+
elastic quantization ensembles for edge devices with fluctuating unified
|
|
323
|
+
memory. The closest academic statement of RAMP's exact problem, and the
|
|
324
|
+
clearest argument that memory elasticity is the right framing.
|
|
325
|
+
- **[Any-Precision LLM](https://arxiv.org/abs/2402.10517)** (Park et al.,
|
|
326
|
+
2024) — one overlaid weight file servable at 3/4/…/n bits. This is the
|
|
327
|
+
engine that would make RAMP's downgrades nearly free; it's on the roadmap
|
|
328
|
+
precisely because of this paper.
|
|
329
|
+
- **[LSAQ](https://arxiv.org/abs/2412.18135)** (2024) — layer-specific
|
|
330
|
+
adaptive quantization chosen per memory budget, the source of the idea that
|
|
331
|
+
a *memory budget* should be the primary input to the decision.
|
|
332
|
+
- **[Voltron](https://arxiv.org/abs/2607.07046)** (2026) — monitors KV-cache
|
|
333
|
+
growth and free memory *during* generation, scaling precision mid-stream.
|
|
334
|
+
RAMP does the cruder between-requests version; Voltron shows where this
|
|
335
|
+
ends up.
|
|
336
|
+
- **[MoBiQuant](https://arxiv.org/abs/2602.20191)** (2026) — token-adaptive
|
|
337
|
+
any-precision inference with efficient runtime bit-width switching.
|
|
338
|
+
- **[PowerInfer](https://arxiv.org/abs/2312.12456)** and
|
|
339
|
+
**[AirLLM](https://github.com/lyogavin/airllm)** — strategies for *fitting*
|
|
340
|
+
oversized models (hot/cold neuron offloading, layer streaming). A different
|
|
341
|
+
problem from elasticity, but complementary and worth knowing.
|
|
342
|
+
|
|
343
|
+
**Tools that shaped the design**
|
|
344
|
+
|
|
345
|
+
- **[llama-swap](https://github.com/mostlygeek/llama-swap)** — the direct
|
|
346
|
+
inspiration for the *shape* of the solution: a transparent proxy in front of
|
|
347
|
+
local inference servers, swapping backends behind a stable endpoint. RAMP
|
|
348
|
+
differs in what triggers a swap (system resource state rather than the
|
|
349
|
+
client's requested model name), but the invisible-proxy architecture is
|
|
350
|
+
llama-swap's idea and it is the right one.
|
|
351
|
+
- **[Ollama](https://ollama.com)** and **[LM Studio](https://lmstudio.ai)** —
|
|
352
|
+
proved local LLM tooling must be zero-configuration to get adopted. Their
|
|
353
|
+
one-time hardware detection is the limitation RAMP addresses, and
|
|
354
|
+
[ollama#14674](https://github.com/ollama/ollama/issues/14674) is the demand
|
|
355
|
+
for it in users' own words.
|
|
356
|
+
|
|
357
|
+
**Built on**
|
|
358
|
+
|
|
359
|
+
[llama.cpp](https://github.com/ggml-org/llama.cpp) and Ollama do the actual
|
|
360
|
+
inference; [psutil](https://github.com/giampaolo/psutil) reads system memory;
|
|
361
|
+
[FastAPI](https://fastapi.tiangolo.com), [httpx](https://www.python-httpx.org)
|
|
362
|
+
and [uvicorn](https://www.uvicorn.org) carry the proxy. RAMP is a controller —
|
|
363
|
+
it deliberately owns none of the hard parts of inference.
|
|
364
|
+
|
|
365
|
+
## Contributing
|
|
366
|
+
|
|
367
|
+
See [CONTRIBUTING.md](CONTRIBUTING.md). You don't need any model files to
|
|
368
|
+
develop or test — the `mock` backend runs the whole daemon end-to-end in
|
|
369
|
+
seconds. Real-world `/ramp/metrics` reports from your own machine are
|
|
370
|
+
especially welcome: whether the tuning defaults are right is an empirical
|
|
371
|
+
question, and more data settles it.
|
|
372
|
+
|
|
373
|
+
## License
|
|
374
|
+
|
|
375
|
+
MIT — see [LICENSE](LICENSE).
|