ramp-llm 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
ramp_llm-0.1.0/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Shivapreetham
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,375 @@
1
+ Metadata-Version: 2.4
2
+ Name: ramp-llm
3
+ Version: 0.1.0
4
+ Summary: RAMP: Resource-Aware Model Proxy - an elastic local LLM daemon that auto-calibrates model size to available RAM, VRAM and disk
5
+ Author: Shivapreetham
6
+ License: MIT
7
+ Project-URL: Homepage, https://github.com/shivapreetham/resource-aware-model-proxy
8
+ Project-URL: Repository, https://github.com/shivapreetham/resource-aware-model-proxy
9
+ Project-URL: Issues, https://github.com/shivapreetham/resource-aware-model-proxy/issues
10
+ Project-URL: Changelog, https://github.com/shivapreetham/resource-aware-model-proxy/blob/main/CHANGELOG.md
11
+ Project-URL: Concepts, https://github.com/shivapreetham/resource-aware-model-proxy/blob/main/docs/CONCEPTS.md
12
+ Keywords: llm,local-llm,llama.cpp,ollama,elastic,quantization,proxy,openai-api,memory-management,inference
13
+ Classifier: Development Status :: 3 - Alpha
14
+ Classifier: Intended Audience :: Developers
15
+ Classifier: License :: OSI Approved :: MIT License
16
+ Classifier: Operating System :: OS Independent
17
+ Classifier: Programming Language :: Python :: 3
18
+ Classifier: Programming Language :: Python :: 3.10
19
+ Classifier: Programming Language :: Python :: 3.11
20
+ Classifier: Programming Language :: Python :: 3.12
21
+ Classifier: Programming Language :: Python :: 3.13
22
+ Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
23
+ Classifier: Topic :: System :: Monitoring
24
+ Classifier: Topic :: Internet :: Proxy Servers
25
+ Requires-Python: >=3.10
26
+ Description-Content-Type: text/markdown
27
+ License-File: LICENSE
28
+ Requires-Dist: fastapi>=0.110
29
+ Requires-Dist: uvicorn>=0.29
30
+ Requires-Dist: httpx>=0.27
31
+ Requires-Dist: psutil>=5.9
32
+ Requires-Dist: PyYAML>=6.0
33
+ Provides-Extra: dev
34
+ Requires-Dist: pytest>=8; extra == "dev"
35
+ Requires-Dist: pytest-asyncio>=0.23; extra == "dev"
36
+ Requires-Dist: ruff>=0.6; extra == "dev"
37
+ Dynamic: license-file
38
+
39
+ # RAMP — Resource-Aware Model Proxy
40
+
41
+ **Your local LLM should get out of the way when you need the RAM back.**
42
+
43
+ [![CI](https://github.com/shivapreetham/resource-aware-model-proxy/actions/workflows/ci.yml/badge.svg)](https://github.com/shivapreetham/resource-aware-model-proxy/actions/workflows/ci.yml)
44
+ [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
45
+ [![Python 3.10+](https://img.shields.io/badge/python-3.10%2B-blue.svg)](https://www.python.org/downloads/)
46
+
47
+ RAMP is an elastic daemon that sits in front of llama.cpp or Ollama. It
48
+ watches RAM, VRAM, and disk continuously, and moves the loaded model up and
49
+ down a *quality ladder* as your machine gets busy and quiets down:
50
+
51
+ - You open Chrome with 40 tabs → RAMP swaps your 7B model for a 3B one.
52
+ - You close them → after a stable recovery window, the 7B comes back.
53
+ - Your assistant never dies with an OOM, and you never reconfigure anything.
54
+
55
+ It speaks the OpenAI API, so **Open-LLM-VTuber, Khoj, SillyTavern, LangChain,
56
+ Continue, or `curl` just point `base_url` at RAMP and work unchanged** — no
57
+ code, no client plugin, nothing to adopt. Swaps cost ~2 seconds warm
58
+ (measured), and in-flight requests are drained rather than dropped.
59
+
60
+ ```diff
61
+ - client = OpenAI(base_url="http://localhost:11434/v1") # straight to Ollama
62
+ + client = OpenAI(base_url="http://localhost:8090/v1") # through RAMP
63
+ ```
64
+
65
+ That's the entire integration. Verified against the official `openai` Python
66
+ SDK: model listing, streaming, and hard-coded model names all work untouched.
67
+ Ollama's native `/api/*` routes are proxied too, for tools that use those.
68
+
69
+ > **New here?** [docs/CONCEPTS.md](docs/CONCEPTS.md) explains what's actually
70
+ > going on: what consumes memory when an LLM runs, why VRAM pressure silently
71
+ > becomes RAM pressure, and the control-theory ideas (hysteresis, damping)
72
+ > behind the policy.
73
+
74
+ ## Why this doesn't already exist
75
+
76
+ Everyone running a local model has made the same bad trade: pick a big model
77
+ and let the machine choke, or pick a small one and pay for the worst case all
78
+ day. The tools make you choose **once, up front, forever**.
79
+
80
+ - **Ollama / LM Studio** pick a quantization *once*, at download/load time.
81
+ Nothing adapts after that ([ollama#14674](https://github.com/ollama/ollama/issues/14674)).
82
+ - **[llama-swap](https://github.com/mostlygeek/llama-swap)** swaps models based
83
+ on *which model the client requests* — the system's resource state is never
84
+ consulted.
85
+ - **FlexQuant / LSAQ / Voltron / Any-Precision LLM** validated elastic
86
+ execution academically; none ship as a usable daemon.
87
+
88
+ RAMP is the missing controller: policy-driven, hysteresis-damped,
89
+ backend-agnostic, and measured — it reports its own swap rate so you can prove
90
+ it isn't thrashing.
91
+
92
+ ## Architecture
93
+
94
+ ```mermaid
95
+ flowchart LR
96
+ C[Any OpenAI client] -->|/v1/*| P[Proxy - FastAPI]
97
+ P -->|forward + x-ramp-tier| B[llama-server child process]
98
+ M[Memory monitor - psutil] -->|raw + EMA available MB| PE[Policy engine]
99
+ PE -->|stay / switch / unload| CT[Controller]
100
+ CT -->|drain, stop, start, health-check| B
101
+ CT -->|gate requests during swap| P
102
+ ```
103
+
104
+ | Module | Role |
105
+ |---|---|
106
+ | `monitor.py` | Samples available RAM, GPU/VRAM (NVIDIA, via `nvidia-smi`), and free disk — each with raw + smoothed readings. |
107
+ | `policy.py` | Pure state machine: fast downgrades under pressure, slow damped upgrades, critical-floor emergency handling. Fully unit-tested, no I/O. |
108
+ | `controller.py` | Executes decisions: drain in-flight requests → stop old backend → start new → reopen the gate. Restarts crashed backends. Records an event log. |
109
+ | `backend.py` | Child-process lifecycle for `llama-server` (real) or `ramp.mock_llm` (demo/tests). |
110
+ | `server.py` | OpenAI-compatible passthrough (incl. SSE streaming) + control API. |
111
+
112
+ ### The decision policy
113
+
114
+ Every tier declares its footprint (`est_ram_mb`, and `est_vram_mb` when it
115
+ runs on the GPU), and the policy holds **RAM, VRAM, and disk** accountable:
116
+
117
+ - **Downgrade (fast):** free RAM below `safety_margin_mb` — or, for a
118
+ GPU-resident tier, free VRAM below `vram_safety_margin_mb` — for
119
+ `downgrade_after_samples` consecutive polls → switch to the largest smaller
120
+ tier that fits *every* projected post-swap budget. A GPU-starved ladder can
121
+ land on a CPU-only tier (`est_vram_mb: 0`).
122
+ - **Critical (instant):** free RAM below `critical_free_mb` → act immediately;
123
+ if even the smallest tier can't fit, unload entirely (requests queue/503
124
+ until memory returns).
125
+ - **Upgrade (slow):** a bigger tier must fit both RAM and VRAM budgets with
126
+ *extra* headroom (`upgrade_extra_mb` / `vram_upgrade_extra_mb`), sustained
127
+ for `upgrade_after_s` seconds, before it's loaded — so RAMP never thrashes
128
+ when you alt-tab.
129
+ - **Disk (gate):** free disk below `disk_min_free_mb` blocks upgrades — a
130
+ bigger model can't free disk, so it gates rather than downgrades — and is
131
+ flagged in `/ramp/status` (`disk.low: true`).
132
+ - No NVIDIA GPU (or no `nvidia-smi`)? VRAM constraints are simply not
133
+ enforced; everything else works unchanged.
134
+ - Swaps happen **between** requests: in-flight generations are drained first
135
+ (up to `drain_timeout_s`), and incoming requests wait on the swap gate
136
+ instead of failing.
137
+
138
+ ## Install
139
+
140
+ RAMP is a Python CLI, so it installs like one. No config file needed — it
141
+ inspects your machine and your installed Ollama models and builds a ladder
142
+ itself.
143
+
144
+ ```bash
145
+ uvx ramp-llm # run without installing anything (needs uv)
146
+ pipx install ramp-llm # or install it properly
147
+ pip install ramp-llm # or into the current environment
148
+ ```
149
+
150
+ <details>
151
+ <summary>Docker</summary>
152
+
153
+ ```bash
154
+ docker build -t ramp .
155
+ docker run --rm -p 8090:8090 \
156
+ --add-host=host.docker.internal:host-gateway \
157
+ -e RAMP_OLLAMA_URL=http://host.docker.internal:11434 \
158
+ ramp
159
+ ```
160
+
161
+ Add `--gpus all` (with the NVIDIA container toolkit) for VRAM awareness. Note
162
+ that inside a container RAMP reads the *container's* memory limit — which is
163
+ usually what you want, but means `--memory` shapes its decisions.
164
+ </details>
165
+
166
+ Then:
167
+
168
+ ```bash
169
+ ramp doctor # check this machine can run it, and what to fix if not
170
+ ramp # start, with an auto-detected ladder
171
+ ramp status # what's loaded, why, and what it's costing
172
+ ```
173
+
174
+ `ramp` prints the endpoint to point your tools at. That's the whole setup.
175
+
176
+ <details>
177
+ <summary>Other commands</summary>
178
+
179
+ ```bash
180
+ ramp init # write the auto-detected ladder to ramp.yaml to tune
181
+ ramp run -c ramp.yaml # start from an explicit config
182
+ ramp status --json # raw status for scripting
183
+ ramp --help
184
+ ```
185
+
186
+ `ramp run` with no `-c` uses `./ramp.yaml` if present, otherwise auto-detects.
187
+ </details>
188
+
189
+ ## Try it without any models
190
+
191
+ The `mock` backend runs the whole daemon against fake OpenAI servers that tag
192
+ their replies with the tier name — so you can watch the laddering behaviour
193
+ in about thirty seconds, with nothing to download.
194
+
195
+ The example configs live in this repo, so grab one first if you installed
196
+ from PyPI:
197
+
198
+ ```bash
199
+ curl -O https://raw.githubusercontent.com/shivapreetham/resource-aware-model-proxy/main/examples/ramp.mock.yaml
200
+ curl -O https://raw.githubusercontent.com/shivapreetham/resource-aware-model-proxy/main/scripts/stress_ram.py
201
+ ramp run -c ramp.mock.yaml
202
+ ```
203
+
204
+ Then, in another terminal:
205
+
206
+ ```bash
207
+ curl http://127.0.0.1:8090/ramp/status
208
+ python stress_ram.py --mb 6000 --hold-s 60
209
+ ```
210
+
211
+ Watch the reply tag flip from `[large-mock]` to a smaller tier while the
212
+ stress script runs, and back ~20s after it releases. Every response carries
213
+ an `x-ramp-tier` header naming the tier that produced it.
214
+
215
+ ## Real models (Ollama backend)
216
+
217
+ If you already run [Ollama](https://ollama.com) (or your OS blocks unsigned
218
+ binaries — e.g. Windows Smart App Control), use `backend: ollama`
219
+ (see [examples/ramp.ollama.yaml](examples/ramp.ollama.yaml)). Tiers reference
220
+ Ollama model tags; RAMP loads/unloads them via `keep_alive`, rewrites each
221
+ request's `model` field to the active tier, and proxies to Ollama's OpenAI
222
+ API. Import your own GGUFs with `ollama create <tag> -f Modelfile`
223
+ (`FROM ./model.gguf`). If no Ollama server is running, RAMP spawns one.
224
+
225
+ ## Real models (llama.cpp)
226
+
227
+ 1. Install [llama.cpp](https://github.com/ggml-org/llama.cpp) so `llama-server`
228
+ is on your PATH (or set `llama_server_bin` to its full path).
229
+ 2. Download 2–3 GGUF models of different sizes (e.g. Qwen2.5 7B/3B/0.5B
230
+ instruct, Q4_K_M).
231
+ 3. Edit `examples/ramp.yaml` — paths, `ctx`, and `est_ram_mb` per tier.
232
+ To calibrate `est_ram_mb`: pin the tier, look at the llama-server process
233
+ in Task Manager, round up.
234
+ 4. `ramp -c examples/ramp.yaml`, then point any OpenAI client at
235
+ `http://127.0.0.1:8090/v1`.
236
+
237
+ ## Control API
238
+
239
+ | Endpoint | Purpose |
240
+ |---|---|
241
+ | `GET /ramp/status` | Current tier, RAM/GPU/disk readings, `last_decision` (why RAMP is holding — e.g. `disk-low`, `upgrade-pending`), tier ladder, metrics, and event log. |
242
+ | `GET /ramp/metrics` | Prometheus text format for scraping. |
243
+ | `POST /ramp/pin/{name}` | Force a tier; disables auto-calibration. |
244
+ | `DELETE /ramp/pin` | Resume auto-calibration. |
245
+
246
+ ## Monitoring
247
+
248
+ An elastic daemon lives or dies on one number: **how often it actually
249
+ swaps.** RAMP measures its own behaviour rather than asking you to trust it —
250
+ swap rate, time lost to swapping, per-tier occupancy, cooldown suppressions,
251
+ and requests that waited or were rejected.
252
+
253
+ ```bash
254
+ curl -s http://127.0.0.1:8090/ramp/metrics
255
+ prometheus --config.file=examples/monitoring/prometheus.yml
256
+ ```
257
+
258
+ ```promql
259
+ sum(rate(ramp_swaps_total[15m])) * 3600 # swaps per hour — the headline
260
+ rate(ramp_swap_seconds_total[30m]) # fraction of time spent swapping
261
+ rate(ramp_tier_seconds_total[1h]) # is your ladder calibrated?
262
+ ```
263
+
264
+ Under 2 swaps/hour means RAMP is invisible, which is the goal; above ~12 means
265
+ churn worth tuning. [docs/MONITORING.md](docs/MONITORING.md) explains each
266
+ number, what to alert on, and how to tune from what you see. Alert rules ship
267
+ in [examples/monitoring/alerts.yml](examples/monitoring/alerts.yml).
268
+
269
+ ### What RAMP itself costs
270
+
271
+ Fair question for any watchdog: is the watcher eating the memory it claims to
272
+ save? Measured — **~65 MB resident, and flat** (+1.9 MB across 300 requests,
273
+ 200 metric polls and 12 swaps; 0.9 s of CPU for the whole run).
274
+
275
+ Two things make that a non-issue. RAMP **never budgets memory it is itself
276
+ using** — `virtual_memory().available` already excludes the daemon's own
277
+ footprint, so the policy reasons about genuinely free memory. And 65 MB is
278
+ roughly 1–3% of a single tier, which is measured in gigabytes.
279
+
280
+ You don't have to take that on faith: the daemon reports its own footprint in
281
+ `/ramp/status` (a `self` block) and as `ramp_self_rss_bytes` /
282
+ `ramp_backend_rss_bytes` in Prometheus. Check it on your machine.
283
+
284
+ ## Tests
285
+
286
+ ```bash
287
+ pytest
288
+ ```
289
+
290
+ Unit tests cover the policy state machine exhaustively (RAM pressure, VRAM
291
+ pressure, the disk gate, hysteresis in both directions); integration tests
292
+ run the full daemon against real child processes with scripted resource
293
+ readings, driving the ladder down, up, through critical unload, streaming,
294
+ manual pinning, and a swap/poll race regression.
295
+
296
+ To reproduce the resource behaviours against real models, see
297
+ [examples/ramp.ollama.yaml](examples/ramp.ollama.yaml) (normal ladder) and
298
+ [examples/ramp.vram-test.yaml](examples/ramp.vram-test.yaml) (isolates VRAM
299
+ pressure: occupy the GPU with another model and watch RAMP fall back to a
300
+ CPU-only tier while system RAM stays healthy).
301
+
302
+ ## Roadmap
303
+
304
+ - **Context-length scaling** — shrink the KV cache before swapping models
305
+ (cheaper first response to pressure).
306
+ - **Any-Precision backend** — one weight file servable at 3/4/8-bit
307
+ ([paper](https://arxiv.org/abs/2402.10517)) makes downgrades near-free.
308
+ - **OS pressure signals** — Windows memory notifications / Linux PSI instead
309
+ of pure polling.
310
+ - **Multi-GPU and non-NVIDIA GPUs** — the monitor currently reads the first
311
+ NVIDIA GPU via `nvidia-smi`.
312
+
313
+ ## Prior art & acknowledgements
314
+
315
+ RAMP didn't invent elastic inference. It productizes an idea that several
316
+ research groups established and that nobody shipped as a usable daemon.
317
+ Credit where it's due.
318
+
319
+ **Research that established the idea**
320
+
321
+ - **[FlexQuant](https://arxiv.org/abs/2501.07139)** (Chai et al., 2025) —
322
+ elastic quantization ensembles for edge devices with fluctuating unified
323
+ memory. The closest academic statement of RAMP's exact problem, and the
324
+ clearest argument that memory elasticity is the right framing.
325
+ - **[Any-Precision LLM](https://arxiv.org/abs/2402.10517)** (Park et al.,
326
+ 2024) — one overlaid weight file servable at 3/4/…/n bits. This is the
327
+ engine that would make RAMP's downgrades nearly free; it's on the roadmap
328
+ precisely because of this paper.
329
+ - **[LSAQ](https://arxiv.org/abs/2412.18135)** (2024) — layer-specific
330
+ adaptive quantization chosen per memory budget, the source of the idea that
331
+ a *memory budget* should be the primary input to the decision.
332
+ - **[Voltron](https://arxiv.org/abs/2607.07046)** (2026) — monitors KV-cache
333
+ growth and free memory *during* generation, scaling precision mid-stream.
334
+ RAMP does the cruder between-requests version; Voltron shows where this
335
+ ends up.
336
+ - **[MoBiQuant](https://arxiv.org/abs/2602.20191)** (2026) — token-adaptive
337
+ any-precision inference with efficient runtime bit-width switching.
338
+ - **[PowerInfer](https://arxiv.org/abs/2312.12456)** and
339
+ **[AirLLM](https://github.com/lyogavin/airllm)** — strategies for *fitting*
340
+ oversized models (hot/cold neuron offloading, layer streaming). A different
341
+ problem from elasticity, but complementary and worth knowing.
342
+
343
+ **Tools that shaped the design**
344
+
345
+ - **[llama-swap](https://github.com/mostlygeek/llama-swap)** — the direct
346
+ inspiration for the *shape* of the solution: a transparent proxy in front of
347
+ local inference servers, swapping backends behind a stable endpoint. RAMP
348
+ differs in what triggers a swap (system resource state rather than the
349
+ client's requested model name), but the invisible-proxy architecture is
350
+ llama-swap's idea and it is the right one.
351
+ - **[Ollama](https://ollama.com)** and **[LM Studio](https://lmstudio.ai)** —
352
+ proved local LLM tooling must be zero-configuration to get adopted. Their
353
+ one-time hardware detection is the limitation RAMP addresses, and
354
+ [ollama#14674](https://github.com/ollama/ollama/issues/14674) is the demand
355
+ for it in users' own words.
356
+
357
+ **Built on**
358
+
359
+ [llama.cpp](https://github.com/ggml-org/llama.cpp) and Ollama do the actual
360
+ inference; [psutil](https://github.com/giampaolo/psutil) reads system memory;
361
+ [FastAPI](https://fastapi.tiangolo.com), [httpx](https://www.python-httpx.org)
362
+ and [uvicorn](https://www.uvicorn.org) carry the proxy. RAMP is a controller —
363
+ it deliberately owns none of the hard parts of inference.
364
+
365
+ ## Contributing
366
+
367
+ See [CONTRIBUTING.md](CONTRIBUTING.md). You don't need any model files to
368
+ develop or test — the `mock` backend runs the whole daemon end-to-end in
369
+ seconds. Real-world `/ramp/metrics` reports from your own machine are
370
+ especially welcome: whether the tuning defaults are right is an empirical
371
+ question, and more data settles it.
372
+
373
+ ## License
374
+
375
+ MIT — see [LICENSE](LICENSE).