turbollm 1.8.7 → 1.9.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +675 -672
- package/dist/cli.js +497 -41
- package/dist/webdist/assets/{AgentEditPage-F6wQ-woa.js → AgentEditPage-CmdThoLF.js} +1 -1
- package/dist/webdist/assets/{ChatScreen-JNejy4c5.js → ChatScreen-ChVYoy7Z.js} +1 -1
- package/dist/webdist/assets/CodeComposer-Bj-OBoKJ.js +1 -0
- package/dist/webdist/assets/{CodeHomeScreen-Bvv8dpoH.js → CodeHomeScreen-D1VHqWAF.js} +1 -1
- package/dist/webdist/assets/CodeSessionScreen-BqvzH7ux.js +5 -0
- package/dist/webdist/assets/{CustomizeScreen-CfQYShcO.js → CustomizeScreen-C053iQLd.js} +4 -4
- package/dist/webdist/assets/DeveloperScreen-BYk0rndu.js +1 -0
- package/dist/webdist/assets/{EnginesScreen-CCzC8Dth.js → EnginesScreen-DkLzY4sp.js} +6 -6
- package/dist/webdist/assets/{FsBrowser-d8UO9bVy.js → FsBrowser-CNit4h3B.js} +1 -1
- package/dist/webdist/assets/ModelLoadMenu-CDHS7gAw.js +2 -0
- package/dist/webdist/assets/{ModelsScreen-DHqnZlFL.js → ModelsScreen-5TvL8kCj.js} +2 -2
- package/dist/webdist/assets/SettingsScreen-BhqvsFo_.js +1 -0
- package/dist/webdist/assets/{SkillEditPage-CKfpniwp.js → SkillEditPage-DmJWJjhQ.js} +1 -1
- package/dist/webdist/assets/{TokensScreen-CSxwCWhD.js → TokensScreen-Dyd5sSGT.js} +1 -1
- package/dist/webdist/assets/{ToolApprovalBar-CDRoIK9R.js → ToolApprovalBar-0Rv_eC33.js} +7 -7
- package/dist/webdist/assets/WorkspaceScreen-BlWA58me.js +1 -0
- package/dist/webdist/assets/{agent-api-DaWwB_7n.js → agent-api-C6XrgjWw.js} +1 -1
- package/dist/webdist/assets/{arc-Bw5r0gQ8.js → arc-V_ckaw1n.js} +1 -1
- package/dist/webdist/assets/{architectureDiagram-3BPJPVTR-D_qgeydj.js → architectureDiagram-3BPJPVTR-BaKHxCMv.js} +1 -1
- package/dist/webdist/assets/{arrow-left-Bs4KhSqC.js → arrow-left-BRhcceVg.js} +1 -1
- package/dist/webdist/assets/{badge-CPXRwtV2.js → badge-Df-MaEdJ.js} +1 -1
- package/dist/webdist/assets/{blockDiagram-GPEHLZMM-BbzqG8AR.js → blockDiagram-GPEHLZMM-W3tqVN_9.js} +1 -1
- package/dist/webdist/assets/{bot-DUgi89DH.js → bot-D2sMRmrf.js} +1 -1
- package/dist/webdist/assets/{c4Diagram-AAUBKEIU-Bltku6hD.js → c4Diagram-AAUBKEIU-Ct_fQADG.js} +1 -1
- package/dist/webdist/assets/channel-Cg9TpvpM.js +1 -0
- package/dist/webdist/assets/{chat-api-DnoBopyh.js → chat-api-DS1CWqFc.js} +1 -1
- package/dist/webdist/assets/chat-queries-BNnEQoUl.js +1 -0
- package/dist/webdist/assets/check-d_QTpKyE.js +1 -0
- package/dist/webdist/assets/{chevron-left-CFIaBRPc.js → chevron-left-Cg7-M3fC.js} +1 -1
- package/dist/webdist/assets/{chevron-right-DT5tldLF.js → chevron-right-B44Qnalv.js} +1 -1
- package/dist/webdist/assets/{chunk-2J33WTMH-jRmwQd0-.js → chunk-2J33WTMH-BhQKciZ6.js} +1 -1
- package/dist/webdist/assets/{chunk-4BX2VUAB-BkQR5P6X.js → chunk-4BX2VUAB-CPVQRcKK.js} +1 -1
- package/dist/webdist/assets/{chunk-55IACEB6-dHYKdZGl.js → chunk-55IACEB6-BeD2m5Zd.js} +1 -1
- package/dist/webdist/assets/{chunk-727SXJPM-Bac05mkQ.js → chunk-727SXJPM-1eXg0vd4.js} +1 -1
- package/dist/webdist/assets/{chunk-AQP2D5EJ-DQSeJwV8.js → chunk-AQP2D5EJ-BI_WWvfb.js} +1 -1
- package/dist/webdist/assets/{chunk-FMBD7UC4-fw71OGhl.js → chunk-FMBD7UC4-DmkZ5E8i.js} +1 -1
- package/dist/webdist/assets/{chunk-ND2GUHAM-xRyemSmx.js → chunk-ND2GUHAM-DgKanzz6.js} +1 -1
- package/dist/webdist/assets/{chunk-QZHKN3VN-BgNNqsB9.js → chunk-QZHKN3VN-C0-EOMS6.js} +1 -1
- package/dist/webdist/assets/{circle-check-CLNNkErq.js → circle-check-CXWe3y-e.js} +1 -1
- package/dist/webdist/assets/{circle-x-DHTcnBYg.js → circle-x-BYUS86Mo.js} +1 -1
- package/dist/webdist/assets/classDiagram-4FO5ZUOK-CUIGvGnh.js +1 -0
- package/dist/webdist/assets/classDiagram-v2-Q7XG4LA2-CUIGvGnh.js +1 -0
- package/dist/webdist/assets/{common-Bh_ckDae.js → common-ClITqWVP.js} +1 -1
- package/dist/webdist/assets/{copy-button-B0AUQv8-.js → copy-button-DOnlD7sA.js} +1 -1
- package/dist/webdist/assets/{cose-bilkent-S5V4N54A-BpeA37l_.js → cose-bilkent-S5V4N54A-CJ-OWUMU.js} +1 -1
- package/dist/webdist/assets/{dagre-BM42HDAG-Dhe-FxnK.js → dagre-BM42HDAG-CP41T21v.js} +1 -1
- package/dist/webdist/assets/{diagram-2AECGRRQ-CsOhE15t.js → diagram-2AECGRRQ-CniIf6PF.js} +1 -1
- package/dist/webdist/assets/{diagram-5GNKFQAL-Dj4NKMxD.js → diagram-5GNKFQAL-C81Uep65.js} +1 -1
- package/dist/webdist/assets/{diagram-KO2AKTUF-BVeDuCrX.js → diagram-KO2AKTUF-BcWsZVCH.js} +1 -1
- package/dist/webdist/assets/{diagram-LMA3HP47-CrQt3n5J.js → diagram-LMA3HP47-D4gsQ5wi.js} +1 -1
- package/dist/webdist/assets/{diagram-OG6HWLK6-CVHDvio_.js → diagram-OG6HWLK6-CKw2XvNh.js} +1 -1
- package/dist/webdist/assets/{dialog-BPhXn1HL.js → dialog-CwglpOgK.js} +1 -1
- package/dist/webdist/assets/{erDiagram-TEJ5UH35-DmEuGD8S.js → erDiagram-TEJ5UH35-DiMRASCw.js} +1 -1
- package/dist/webdist/assets/{external-link-DgElSULb.js → external-link-vubZvvSL.js} +1 -1
- package/dist/webdist/assets/{flowDiagram-I6XJVG4X-D3ES6R-Y.js → flowDiagram-I6XJVG4X-OYvJz7NK.js} +1 -1
- package/dist/webdist/assets/{ganttDiagram-6RSMTGT7-DzovqV_S.js → ganttDiagram-6RSMTGT7-TNBmycEa.js} +1 -1
- package/dist/webdist/assets/{git-branch-DZp9oCCA.js → git-branch-CpH7bdLO.js} +1 -1
- package/dist/webdist/assets/{gitGraphDiagram-PVQCEYII-BV0Hvb-m.js → gitGraphDiagram-PVQCEYII-BhEsrTFm.js} +1 -1
- package/dist/webdist/assets/{index-BivjMfQ5.js → index-BGFp5nvZ.js} +7 -7
- package/dist/webdist/assets/index-mhXL7TR_.css +1 -0
- package/dist/webdist/assets/{infoDiagram-5YYISTIA-DHmj_pKX.js → infoDiagram-5YYISTIA-BDEgDbkv.js} +1 -1
- package/dist/webdist/assets/{ishikawaDiagram-YF4QCWOH-DjAmacek.js → ishikawaDiagram-YF4QCWOH-JTmVtiNo.js} +1 -1
- package/dist/webdist/assets/{journeyDiagram-JHISSGLW-InYx9cWN.js → journeyDiagram-JHISSGLW-CSlH-DCp.js} +1 -1
- package/dist/webdist/assets/{kanban-definition-UN3LZRKU-BO9O2UDq.js → kanban-definition-UN3LZRKU-BIfzX5mh.js} +1 -1
- package/dist/webdist/assets/{linear-B6IwpKfz.js → linear-Csu7LYv_.js} +1 -1
- package/dist/webdist/assets/lock-Dx47ssSz.js +1 -0
- package/dist/webdist/assets/{mermaid.core-MlIRHLxX.js → mermaid.core-BF2_JyC4.js} +4 -4
- package/dist/webdist/assets/{mindmap-definition-RKZ34NQL-kvrr7ftb.js → mindmap-definition-RKZ34NQL-C4HOUf0T.js} +1 -1
- package/dist/webdist/assets/{pencil-k4AQARf_.js → pencil-DiDWTtDe.js} +1 -1
- package/dist/webdist/assets/{personas-CL0Kc_Kd.js → personas-K7vGADki.js} +11 -9
- package/dist/webdist/assets/{pieDiagram-4H26LBE5-DaE8CAe0.js → pieDiagram-4H26LBE5-Civ4ky7C.js} +1 -1
- package/dist/webdist/assets/{plus-CY4mhBkY.js → plus-rMCCWOjM.js} +1 -1
- package/dist/webdist/assets/{quadrantDiagram-W4KKPZXB-Duu8s5l-.js → quadrantDiagram-W4KKPZXB-Cu5aaIeO.js} +1 -1
- package/dist/webdist/assets/{refresh-cw-DCf4F0oQ.js → refresh-cw-zhCl4OJc.js} +1 -1
- package/dist/webdist/assets/{requirementDiagram-4Y6WPE33-Cbl7MhTK.js → requirementDiagram-4Y6WPE33-owXfcooA.js} +1 -1
- package/dist/webdist/assets/{rotate-ccw-Bk6FjMT9.js → rotate-ccw-DSFeEfs1.js} +1 -1
- package/dist/webdist/assets/{sankeyDiagram-5OEKKPKP-BiK0JGus.js → sankeyDiagram-5OEKKPKP-B7o6U2uP.js} +1 -1
- package/dist/webdist/assets/{sequenceDiagram-3UESZ5HK-Dr7D9A4w.js → sequenceDiagram-3UESZ5HK-D0g36FR_.js} +1 -1
- package/dist/webdist/assets/{skeleton-Cy7DsAmm.js → skeleton-B6viaKty.js} +1 -1
- package/dist/webdist/assets/{sliders-horizontal-CVf2FXz2.js → sliders-horizontal-YIE_L0IY.js} +1 -1
- package/dist/webdist/assets/{sparkles-Bvrl94s5.js → sparkles-FpLiZDCb.js} +1 -1
- package/dist/webdist/assets/{star-DLAcO0R_.js → star-D0lj4klT.js} +1 -1
- package/dist/webdist/assets/{stateDiagram-AJRCARHV-BthDRItX.js → stateDiagram-AJRCARHV-BGbyDa8P.js} +1 -1
- package/dist/webdist/assets/stateDiagram-v2-BHNVJYJU-CFpzxmQe.js +1 -0
- package/dist/webdist/assets/{switch-CI322tBg.js → switch-D7jMdkCs.js} +1 -1
- package/dist/webdist/assets/{terminal-Bwy3mmYf.js → terminal-Dfyn1hK2.js} +1 -1
- package/dist/webdist/assets/{timeline-definition-PNZ67QCA-DkR-Byrz.js → timeline-definition-PNZ67QCA-DlGxVI3e.js} +1 -1
- package/dist/webdist/assets/{trash-2-B9m1cz2x.js → trash-2-DpcKQJdU.js} +1 -1
- package/dist/webdist/assets/{useIsDesktop-xqMNa55J.js → useIsDesktop-BuJChlQF.js} +1 -1
- package/dist/webdist/assets/{vennDiagram-CIIHVFJN-aOYXm4nP.js → vennDiagram-CIIHVFJN-D8cIiNWM.js} +1 -1
- package/dist/webdist/assets/{wardley-L42UT6IY-DptJ0PbQ.js → wardley-L42UT6IY-wEcDRtc5.js} +1 -1
- package/dist/webdist/assets/{wardleyDiagram-YWT4CUSO-CN66GNqR.js → wardleyDiagram-YWT4CUSO-E26Pag8Q.js} +1 -1
- package/dist/webdist/assets/{wrench-CLvjdoSH.js → wrench-CNEP_hb3.js} +1 -1
- package/dist/webdist/assets/{xychartDiagram-2RQKCTM6-CbOWShHZ.js → xychartDiagram-2RQKCTM6-lyn53sKq.js} +1 -1
- package/dist/webdist/index.html +2 -2
- package/package.json +1 -1
- package/dist/webdist/assets/CodeComposer-BMivBZTt.js +0 -1
- package/dist/webdist/assets/CodeSessionScreen-BZ_n-Q8C.js +0 -5
- package/dist/webdist/assets/DeveloperScreen-DlLMgh2N.js +0 -1
- package/dist/webdist/assets/ModelLoadMenu-BcSALweE.js +0 -2
- package/dist/webdist/assets/SettingsScreen-Byi4wqlL.js +0 -1
- package/dist/webdist/assets/WorkspaceScreen-CY5EZKSy.js +0 -1
- package/dist/webdist/assets/channel-Cob2-fKv.js +0 -1
- package/dist/webdist/assets/chat-queries-EOqarO2W.js +0 -1
- package/dist/webdist/assets/check-C-_wc2AD.js +0 -1
- package/dist/webdist/assets/classDiagram-4FO5ZUOK-emMvNtz-.js +0 -1
- package/dist/webdist/assets/classDiagram-v2-Q7XG4LA2-emMvNtz-.js +0 -1
- package/dist/webdist/assets/index-S6YWlmeh.css +0 -1
- package/dist/webdist/assets/stateDiagram-v2-BHNVJYJU--IOHWZLa.js +0 -1
package/README.md
CHANGED
|
@@ -1,672 +1,675 @@
|
|
|
1
|
-
<p align="center">
|
|
2
|
-
<img src="https://raw.githubusercontent.com/mohitsoni48/TurboLLM/main/turbollm/web/public/brand/turbollm-icon-512.jpeg?v=2" width="92" height="92" alt="TurboLLM" />
|
|
3
|
-
</p>
|
|
4
|
-
|
|
5
|
-
<h1 align="center">TurboLLM</h1>
|
|
6
|
-
|
|
7
|
-
<p align="center">
|
|
8
|
-
<strong>Run <em>any</em> local LLM engine, auto-tuned to your GPU — with a polished web UI
|
|
9
|
-
and an OpenAI/Anthropic-compatible API.</strong><br/>
|
|
10
|
-
Bring your own llama.cpp fork. No compiling. No Electron. No Python. Point Claude Code at
|
|
11
|
-
your own machine in one command — fully offline.
|
|
12
|
-
</p>
|
|
13
|
-
|
|
14
|
-
<p align="center">
|
|
15
|
-
<a href="https://turbollm.dev"><img src="https://img.shields.io/badge/turbollm.dev-docs%20·%20models%20·%20guides-e2552e" alt="turbollm.dev" /></a>
|
|
16
|
-
<a href="https://www.npmjs.com/package/turbollm"><img src="https://img.shields.io/npm/v/turbollm.svg?color=e2552e" alt="npm version" /></a>
|
|
17
|
-
<a href="https://www.npmjs.com/package/turbollm"><img src="https://img.shields.io/npm/dm/turbollm.svg?color=e2552e" alt="npm downloads" /></a>
|
|
18
|
-
<img src="https://img.shields.io/badge/node-%E2%89%A522-3c873a.svg" alt="node >= 22" />
|
|
19
|
-
<img src="https://img.shields.io/badge/license-FSL--1.1--ALv2-blue.svg" alt="license" />
|
|
20
|
-
<img src="https://img.shields.io/badge/platform-Windows%20%C2%B7%20macOS%20%C2%B7%20Linux-555.svg" alt="platforms" />
|
|
21
|
-
<a href="https://ko-fi.com/mohitsoni"><img src="https://img.shields.io/badge/Ko--fi-support%20us-FF5E5B?logo=kofi&logoColor=white" alt="Ko-fi" /></a>
|
|
22
|
-
<a href="https://github.com/sponsors/mohitsoni48"><img src="https://img.shields.io/badge/GitHub%20Sponsors-support%20us-EA4AAA?logo=githubsponsors&logoColor=white" alt="GitHub Sponsors" /></a>
|
|
23
|
-
<a href="https://discord.gg/v6kRbV7nC"><img src="https://img.shields.io/badge/Discord-join%20chat-5865F2?logo=discord&logoColor=white" alt="Discord" /></a>
|
|
24
|
-
</p>
|
|
25
|
-
|
|
26
|
-
<p align="center">
|
|
27
|
-
<sub><strong>Source-available</strong> under <a href="#license">FSL-1.1</a> — free for personal and
|
|
28
|
-
internal business use; every release converts to <strong>Apache-2.0</strong> two years after it ships.</sub>
|
|
29
|
-
</p>
|
|
30
|
-
|
|
31
|
-
<!-- Brand: shipped app icon web/public/brand/turbollm-icon-512.jpeg · high-res masters web/brand-assets/ (unshipped) · in-app mark web/src/components/Logo.tsx · favicon web/public/favicon.svg -->
|
|
32
|
-
|
|
33
|
-
```bash
|
|
34
|
-
npx turbollm
|
|
35
|
-
```
|
|
36
|
-
|
|
37
|
-
That one command starts a local daemon, opens a browser UI, and serves your models over an
|
|
38
|
-
API any tool can talk to. TurboLLM is the **performance & bleeding-edge layer for local
|
|
39
|
-
LLMs** — built for people who today hand-compile forks and hunt forums for the right flags.
|
|
40
|
-
The speed is measured, not promised: **74.7 vs 61 t/s on the *same* official llama.cpp as
|
|
41
|
-
LM Studio — and 2.2× on forks it can't load** ([the numbers](#speed-turbollm-vs-lm-studio)).
|
|
42
|
-
|
|
43
|
-
<p align="center">
|
|
44
|
-
<strong>📖 Full docs · what runs on your GPU · model picks → <a href="https://turbollm.dev">turbollm.dev</a></strong>
|
|
45
|
-
</p>
|
|
46
|
-
|
|
47
|
-
<p align="center">
|
|
48
|
-
<img src="https://raw.githubusercontent.com/mohitsoni48/TurboLLM/main/assets/how-it-works.svg?v=2" width="860" alt="How TurboLLM works: clients -> one lightweight daemon -> any engine on your GPU" />
|
|
49
|
-
</p>
|
|
50
|
-
|
|
51
|
-
---
|
|
52
|
-
|
|
53
|
-
## Contents
|
|
54
|
-
|
|
55
|
-
- [Why TurboLLM](#why-turbollm)
|
|
56
|
-
- [Who this is for — and who it isn't](#who-this-is-for--and-who-it-isnt)
|
|
57
|
-
- [Speed: TurboLLM vs LM Studio](#speed-turbollm-vs-lm-studio)
|
|
58
|
-
- [Features](#features)
|
|
59
|
-
- [Quick start](#quick-start)
|
|
60
|
-
- [⭐ Bring any engine — the headline feature](#-bring-any-engine--the-headline-feature)
|
|
61
|
-
- [Run Claude Code on your own GPU](#run-claude-code-on-your-own-gpu)
|
|
62
|
-
- [Use it from any device on your network](#use-it-from-any-device-on-your-network)
|
|
63
|
-
- [Command-line reference](#command-line-reference)
|
|
64
|
-
- [Configuration & data](#configuration--data)
|
|
65
|
-
- [Requirements](#requirements)
|
|
66
|
-
- [Privacy](#privacy)
|
|
67
|
-
- [How TurboLLM compares](#how-turbollm-compares)
|
|
68
|
-
- [Troubleshooting](#troubleshooting)
|
|
69
|
-
- [Develop from source](#develop-from-source)
|
|
70
|
-
- [Community](#community)
|
|
71
|
-
- [License](#license)
|
|
72
|
-
|
|
73
|
-
---
|
|
74
|
-
|
|
75
|
-
## Why TurboLLM
|
|
76
|
-
|
|
77
|
-
Local-LLM tools make two choices for you, and both cost you performance:
|
|
78
|
-
|
|
79
|
-
1. **They pick the engine.** LM Studio ships one blessed runtime; Ollama hides the engine
|
|
80
|
-
entirely. The fastest community innovations — new quant formats, speculative decoding,
|
|
81
|
-
low-bit KV cache — land in **forks** first, and you can't use them without compiling.
|
|
82
|
-
2. **They don't tell you what speed to expect**, and they don't tune the dozens of launch
|
|
83
|
-
flags (`-c`, `-ngl`, `--n-cpu-moe`, KV type, threads, flash-attn, draft models) that make
|
|
84
|
-
the difference between 20 and 80 tokens/sec on the *same* hardware.
|
|
85
|
-
|
|
86
|
-
TurboLLM does the opposite:
|
|
87
|
-
|
|
88
|
-
- **🔌 Any engine, including forks.** Point it at any `llama-server`-compatible binary — a
|
|
89
|
-
build you compiled, a community fork, or the one it auto-provisions for your GPU. It probes
|
|
90
|
-
the binary's real capabilities and adapts the UI to them. **This is the whole point.**
|
|
91
|
-
- **⚡ Auto-tuned to your hardware.** It benchmarks on load, derives fast defaults, and shows
|
|
92
|
-
a **VRAM-fit verdict before you load** — no more flag guessing.
|
|
93
|
-
- **📊 Real tokens/sec, never faked.** Speed in the model list is *measured on your machine*
|
|
94
|
-
from actual generation — live while you chat, and remembered per model.
|
|
95
|
-
- **🪶 Lightweight.** A ~7 MB npm package on Node — **no Electron, no bundled Chromium, no
|
|
96
|
-
Python**. It downloads only the engine your GPU actually needs (Vulkan ≈ 38 MB).
|
|
97
|
-
- **🔌 Drop-in APIs.** OpenAI **and** Anthropic-compatible — so Claude Code and every existing
|
|
98
|
-
tool work unchanged.
|
|
99
|
-
- **🔀 A gateway that loads models for you.** Name any model in your API request and TurboLLM
|
|
100
|
-
loads it on demand, keeping your favorites hot in a small pool — so an agent that hops between
|
|
101
|
-
models just works, with nothing to pre-wire.
|
|
102
|
-
- **🔒 Offline-first & private.** No account, no backend, no internet, **no telemetry.**
|
|
103
|
-
|
|
104
|
-
---
|
|
105
|
-
|
|
106
|
-
## Who this is for — and who it isn't
|
|
107
|
-
|
|
108
|
-
**It's for you if:**
|
|
109
|
-
|
|
110
|
-
- you're tired of hand-tuning `-ngl` / KV / offload flags for every model and never trusting
|
|
111
|
-
the speed number;
|
|
112
|
-
- you want community forks (low-bit KV cache, NextN speculative decoding, new quant formats)
|
|
113
|
-
on day 0 — without compiling anything;
|
|
114
|
-
- you're pointing agents or **Claude Code** at a local model and want that to be one command;
|
|
115
|
-
- you already have folders of GGUFs from LM Studio / Ollama and refuse to re-download them.
|
|
116
|
-
|
|
117
|
-
**It's probably not for you if:**
|
|
118
|
-
|
|
119
|
-
- you want a zero-terminal, one-click desktop installer — TurboLLM currently needs
|
|
120
|
-
[Node ≥22](#requirements) and one terminal command (a packaged installer is planned);
|
|
121
|
-
- you're happy with a single blessed runtime and stock settings — LM Studio and Ollama
|
|
122
|
-
already serve that well;
|
|
123
|
-
- you need training or fine-tuning — TurboLLM is inference only.
|
|
124
|
-
|
|
125
|
-
---
|
|
126
|
-
|
|
127
|
-
## Speed: TurboLLM vs LM Studio
|
|
128
|
-
|
|
129
|
-
Same GPU (RTX 5070 Ti 16 GB), same model, same 200K context — measured generation speed.
|
|
130
|
-
**TurboLLM is faster than LM Studio on the very same official llama.cpp, and faster still when you
|
|
131
|
-
run a community fork LM Studio can't.**
|
|
132
|
-
|
|
133
|
-
**① On official llama.cpp, TurboLLM is faster.** It auto-provisions a GPU-native engine build (CUDA
|
|
134
|
-
13 for Blackwell here) and tunes expert-offload to the layer, so at the *same* KV-cache quant it
|
|
135
|
-
beats LM Studio's bundled runtime:
|
|
136
|
-
|
|
137
|
-
| Qwen3.6-35B-A3B · 200K | TurboLLM | LM Studio | Speed-up |
|
|
138
|
-
|---|:---:|:---:|:---:|
|
|
139
|
-
| official llama.cpp — `q4_0` | **74.7 t/s** | 61.0 t/s | **1.2×** |
|
|
140
|
-
| official llama.cpp — `q8_0` | **72.3 t/s** | ~66 t/s\* | **1.1×** |
|
|
141
|
-
|
|
142
|
-
**② Run a faster engine and pull far ahead.** Because TurboLLM runs *any* engine, you can drop in
|
|
143
|
-
the **TurboQuant** fork — a llama.cpp fork with a low-bit `turbo4` KV cache that LM Studio simply
|
|
144
|
-
can't load — in one click. On a large-KV model it delivers `q8_0`-level quality at **more than
|
|
145
|
-
double the speed**:
|
|
146
|
-
|
|
147
|
-
| Qwen3.6-27B · 200K · matched quality | TurboLLM + TurboQuant | LM Studio | Speed-up |
|
|
148
|
-
|---|:---:|:---:|:---:|
|
|
149
|
-
| `turbo4` vs `q8_0` | **24.6 t/s** | 11.4 t/s | **2.2×** |
|
|
150
|
-
|
|
151
|
-
Same run, **1.7× faster prefill** too (1288 vs 757 tok/s).
|
|
152
|
-
|
|
153
|
-
<sub>\*LM Studio's `q8_0` mildly spilled VRAM at its best offload. A low-bit KV cache helps most
|
|
154
|
-
when the cache is large; TurboLLM's auto-tuner and on-screen measured t/s pick the fastest engine +
|
|
155
|
-
config for each model, so you don't have to.</sub>
|
|
156
|
-
|
|
157
|
-
---
|
|
158
|
-
|
|
159
|
-
## Features
|
|
160
|
-
|
|
161
|
-
The headline — **[running any engine, including community forks](#-bring-any-engine--the-headline-feature)** —
|
|
162
|
-
has its own section below. Everything else is grouped here; each summary is the gist, expand for
|
|
163
|
-
the detail:
|
|
164
|
-
|
|
165
|
-
<details>
|
|
166
|
-
<summary><strong>📦 Models — bring your own, or browse Hugging Face</strong></summary>
|
|
167
|
-
|
|
168
|
-
<br/>
|
|
169
|
-
|
|
170
|
-
- **Use the folders you already have.** Point TurboLLM at any directory of GGUFs — your
|
|
171
|
-
existing LM Studio folders or manual downloads — **no re-downloading.** It parses GGUF
|
|
172
|
-
metadata (arch, params, quant, context, vision) for every file. (Models pulled by
|
|
173
|
-
`ollama pull` live in Ollama's own blob store, which TurboLLM doesn't read.)
|
|
174
|
-
- **Browse & download from Hugging Face**, in-app: a live, sortable list (trending / downloads
|
|
175
|
-
/ likes / recently updated / newest) alongside a permanent detail pane — pick a quant, read
|
|
176
|
-
the **rendered model card**, and download with **resume + SHA-256 verification**. Each model
|
|
177
|
-
lands in its own folder (mirroring Hugging Face's own layout) with its **vision projector and
|
|
178
|
-
every shard of a split/multipart quant fetched alongside it automatically**. Gated models
|
|
179
|
-
(Llama, Gemma) work via your own HF token, which **never leaves your machine**.
|
|
180
|
-
- **Import from any URL** — not just Hugging Face. Paste a direct `.gguf` link (model-author
|
|
181
|
-
sites, mirrors, private servers) and it disk-space-checks and downloads through the same
|
|
182
|
-
manager; paste a Hugging Face **model page** link instead and it opens that repo's quant
|
|
183
|
-
picker so you can choose which one to download.
|
|
184
|
-
- **Quant recommendation per GPU** and a **VRAM-fit verdict** so you pick a quant that
|
|
185
|
-
actually fits before you commit.
|
|
186
|
-
- **Primary download folder**, real-time **measured t/s per model**, **delete-from-disk**, and
|
|
187
|
-
**pin your favourites to the top of the list**.
|
|
188
|
-
|
|
189
|
-
</details>
|
|
190
|
-
|
|
191
|
-
<details>
|
|
192
|
-
<summary><strong>⚡ Auto-tuning & performance</strong></summary>
|
|
193
|
-
|
|
194
|
-
<br/>
|
|
195
|
-
|
|
196
|
-
- **Auto-benchmark on load** derives fast defaults for your exact GPU.
|
|
197
|
-
- **Recommended sampling from Hugging Face** — auto-tune checks a repo's structured params /
|
|
198
|
-
`generation_config.json` sidecar first when the quantizer publishes one (exact values, no
|
|
199
|
-
guessing), then falls back to reading the model's card (and the original model behind a
|
|
200
|
-
requant) and prefills the author's recommended `temperature / top_k / top_p / min_p`. No
|
|
201
|
-
recommendation → your sampling is left untouched.
|
|
202
|
-
- **Real measured tokens/sec** in the model list — **live** while generating, **last-session**
|
|
203
|
-
when idle (never a synthetic estimate).
|
|
204
|
-
- **Full load-parameter UI**, a superset of what other tools expose: context length, GPU offload
|
|
205
|
-
(`-ngl`), **MoE CPU-offload (`--n-cpu-moe`)**, parallel slots, **independent K and V cache-quant
|
|
206
|
-
type** (incl. low-bit on supporting forks), CPU threads, flash attention, **speculative decoding
|
|
207
|
-
(NextN / MTP / draft, with a configurable draft min/max window)**, a **pinned engine port** per
|
|
208
|
-
model, and a **custom/raw flags** field for anything not exposed as its own control.
|
|
209
|
-
- **Copy the exact launch command** for a loaded model — runs the same config standalone with
|
|
210
|
-
llama.cpp/the fork directly, no TurboLLM required.
|
|
211
|
-
- **Auto-fit GPU layers / MoE offload** — an optional toggle (off by default) that hands the
|
|
212
|
-
GPU/CPU split decision to llama.cpp's own memory-fitting logic at load time instead of a fixed
|
|
213
|
-
number, honored by Auto-tune too. Useful when a large context or model doesn't fit your usual
|
|
214
|
-
fixed setting.
|
|
215
|
-
- **Fast by default:** flash attention on, NextN self-speculative decoding on for models that
|
|
216
|
-
carry a draft head, threads auto — safely gated to what your engine actually accepts.
|
|
217
|
-
- **Multi-GPU, per model** — split a model across cards (layer/row split + main-GPU pick on
|
|
218
|
-
llama.cpp, tensor-parallel on vLLM). Defaults are no-ops, so single-GPU rigs are untouched.
|
|
219
|
-
- **Saved per-model profiles, per engine** — tune once per (model, engine) pair, so
|
|
220
|
-
switching engines (or between two installs of the same engine, e.g. a fork) never
|
|
221
|
-
overwrites another engine's tuning for the same model.
|
|
222
|
-
- **Configurable VRAM headroom** (Settings → Models & loading → Advanced, 300 MB–2 GB, default
|
|
223
|
-
1 GB) — tell auto-tune how much VRAM to keep free for other GPU workloads instead of a fixed margin.
|
|
224
|
-
Drag it to 0 to opt into an experimental **MoE "VRAM-spill" search** — auto-tune keeps pushing more
|
|
225
|
-
experts onto the GPU past the safe margin as long as both generation and prompt-processing speed
|
|
226
|
-
keep improving.
|
|
227
|
-
|
|
228
|
-
</details>
|
|
229
|
-
|
|
230
|
-
<details>
|
|
231
|
-
<summary><strong>💬 Chat & agentic tools — a genuinely good UI, not an afterthought</strong></summary>
|
|
232
|
-
|
|
233
|
-
<br/>
|
|
234
|
-
|
|
235
|
-
- **Streaming** with a **stop** button, **live tokens/sec**, **prompt-processing %** and
|
|
236
|
-
**prefill t/s**, **time-to-first-token**, **total time**, exact **token counts**, and a
|
|
237
|
-
**context-usage meter** (filled / max) on every reply.
|
|
238
|
-
- **Thinking control** — toggle reasoning **off** for a direct answer, or leave it **on** with
|
|
239
|
-
collapsible, timed "thought for N s" blocks.
|
|
240
|
-
- **Markdown + syntax-highlighted code** with one-click copy — plus **inline Unicode charts**
|
|
241
|
-
the model draws when a comparison, trend, or hierarchy is genuinely worth a visual.
|
|
242
|
-
- **Live artifacts** — `html`, `svg`, and `mermaid` replies render as **sandboxed, offline
|
|
243
|
-
previews** shown as an image, with one-click export to **PNG / JPEG / SVG / animated GIF / HTML**.
|
|
244
|
-
- **Agents** — pick a style (Default · **Designer** · Concise · Detailed · Blunt · Formal · Tutor ·
|
|
245
|
-
Creative · Research · **Lite** · **Code**) per conversation, no prompt-wrangling required. The
|
|
246
|
-
**Designer** persona produces polished, self-contained, previewable designs by default; **Lite**
|
|
247
|
-
strips the hidden prompt to the bare minimum for the fastest responses. See below for editing
|
|
248
|
-
built-ins or creating your own.
|
|
249
|
-
- **Edit or regenerate any message without losing history** — both branch instead of
|
|
250
|
-
overwriting, with a **‹ 1/2 › switcher** to flip between versions (including nested branch
|
|
251
|
-
points); delete and copy still work as before. **Persistent, searchable conversations** with
|
|
252
|
-
rename, delete, and **auto-generated titles**, organized into **drag-resizable, collapsible
|
|
253
|
-
folders** you create/rename/move conversations into.
|
|
254
|
-
- **Switching chats never cancels a reply** — an in-flight generation keeps running in the
|
|
255
|
-
background and saves normally; the sidebar shows a live indicator on any chat still generating,
|
|
256
|
-
and a dot on one that finished while you were elsewhere.
|
|
257
|
-
- **Preserve thinking across turns** (on by default) — the model's past reasoning is resent on
|
|
258
|
-
later turns, not just its final answers, so follow-ups have real context to work with.
|
|
259
|
-
- **Per-chat system prompt** and **per-chat sampling** overrides — temperature, top-p/k, min-p,
|
|
260
|
-
repeat/presence/frequency penalties, and **stop strings** — prefilled from the loaded model's
|
|
261
|
-
own recommended values, available even before you send the first message.
|
|
262
|
-
- **Image input** for vision models, **PDF and code/text attachments** (real extracted text,
|
|
263
|
-
not raw bytes), and **TurboLLM Expert** — a built-in assistant that knows the app and your
|
|
264
|
-
hardware for onboarding and troubleshooting without leaving the UI.
|
|
265
|
-
- **Agentic tools** — built-in `web_search` (Tavily), `fetch_url`, and sandboxed `run_code`, plus
|
|
266
|
-
an **MCP marketplace** in Customize: one-click connect for hosted MCPs (GitHub, Linear, Stripe,
|
|
267
|
-
Atlassian, Neon, Supabase, Cloudflare, Zapier, Apify, Mixpanel) and open-source local MCPs
|
|
268
|
-
(filesystem, git, postgres, playwright, …), plus your own custom servers. Connected tools appear
|
|
269
|
-
in every chat with no restart. A **Research** persona forces multi-step web search and cites sources inline.
|
|
270
|
-
- **Tool-call approval gate** — every tool call asks for your approval by default before it runs,
|
|
271
|
-
with **Deny**, **Allow**, **Allow for this chat**, or **Always Allow** on an inline bar above the
|
|
272
|
-
composer. Set per-tool defaults globally from Settings → Tools & safety, or flip **Auto-allow
|
|
273
|
-
all** to skip prompts entirely — a tool set to Deny still stays blocked either way.
|
|
274
|
-
- **Usage dashboard** — a GitHub-style activity heatmap of your local generation history (adaptive
|
|
275
|
-
1h/12h/24h boxes for the 7-day/30-day/all-time views), streaks, peak hour, a per-model
|
|
276
|
-
breakdown, a lifetime token-milestone tracker, and a separate **API tab** tracking tokens hitting
|
|
277
|
-
the gateway from external tools like Claude Code — not just in-app chat.
|
|
278
|
-
- **Auto-memory** *(experimental, off by default)* — silently extracts durable facts you mention
|
|
279
|
-
in chat (name, preferences, hardware) using your own loaded model, and carries them into future
|
|
280
|
-
new conversations. Nothing leaves your device; the full fact list is reviewable and deletable
|
|
281
|
-
from Settings → Memory, and turning the toggle off stops new chats from seeing it immediately.
|
|
282
|
-
- **Thinking-budget control** — a graduated slider, not just on/off: cap reasoning to a specific
|
|
283
|
-
token count, disable it entirely, or leave it unlimited. Works in Chat and Code alike.
|
|
284
|
-
|
|
285
|
-
</details>
|
|
286
|
-
|
|
287
|
-
<details>
|
|
288
|
-
<summary><strong>🧑💻 Code — a local coding agent, in a real project directory</strong></summary>
|
|
289
|
-
|
|
290
|
-
|
|
291
|
-
|
|
292
|
-
|
|
293
|
-
|
|
294
|
-
|
|
295
|
-
|
|
296
|
-
|
|
297
|
-
|
|
298
|
-
|
|
299
|
-
|
|
300
|
-
-
|
|
301
|
-
|
|
302
|
-
|
|
303
|
-
|
|
304
|
-
|
|
305
|
-
|
|
306
|
-
|
|
307
|
-
`
|
|
308
|
-
|
|
309
|
-
|
|
310
|
-
|
|
311
|
-
|
|
312
|
-
|
|
313
|
-
|
|
314
|
-
|
|
315
|
-
|
|
316
|
-
|
|
317
|
-
<
|
|
318
|
-
|
|
319
|
-
|
|
320
|
-
|
|
321
|
-
|
|
322
|
-
|
|
323
|
-
|
|
324
|
-
- **
|
|
325
|
-
|
|
326
|
-
|
|
327
|
-
|
|
328
|
-
|
|
329
|
-
|
|
330
|
-
|
|
331
|
-
|
|
332
|
-
<
|
|
333
|
-
|
|
334
|
-
|
|
335
|
-
|
|
336
|
-
|
|
337
|
-
|
|
338
|
-
|
|
339
|
-
|
|
340
|
-
|
|
341
|
-
|
|
342
|
-
|
|
343
|
-
-
|
|
344
|
-
|
|
345
|
-
|
|
346
|
-
- **
|
|
347
|
-
|
|
348
|
-
|
|
349
|
-
- **
|
|
350
|
-
|
|
351
|
-
**
|
|
352
|
-
|
|
353
|
-
|
|
354
|
-
|
|
355
|
-
|
|
356
|
-
|
|
357
|
-
**
|
|
358
|
-
|
|
359
|
-
|
|
360
|
-
|
|
361
|
-
|
|
362
|
-
|
|
363
|
-
|
|
364
|
-
|
|
365
|
-
|
|
366
|
-
|
|
367
|
-
|
|
368
|
-
<
|
|
369
|
-
|
|
370
|
-
|
|
371
|
-
|
|
372
|
-
|
|
373
|
-
|
|
374
|
-
|
|
375
|
-
|
|
376
|
-
|
|
377
|
-
|
|
378
|
-
|
|
379
|
-
**
|
|
380
|
-
|
|
381
|
-
|
|
382
|
-
|
|
383
|
-
|
|
384
|
-
|
|
385
|
-
|
|
386
|
-
|
|
387
|
-
|
|
388
|
-
|
|
389
|
-
<
|
|
390
|
-
|
|
391
|
-
|
|
392
|
-
|
|
393
|
-
|
|
394
|
-
|
|
395
|
-
|
|
396
|
-
|
|
397
|
-
|
|
398
|
-
|
|
399
|
-
|
|
400
|
-
|
|
401
|
-
|
|
402
|
-
|
|
403
|
-
|
|
404
|
-
|
|
405
|
-
#
|
|
406
|
-
|
|
407
|
-
|
|
408
|
-
|
|
409
|
-
|
|
410
|
-
|
|
411
|
-
|
|
412
|
-
|
|
413
|
-
|
|
414
|
-
|
|
415
|
-
|
|
416
|
-
|
|
417
|
-
|
|
418
|
-
|
|
419
|
-
|
|
420
|
-
|
|
421
|
-
|
|
422
|
-
|
|
423
|
-
|
|
424
|
-
|
|
425
|
-
|
|
426
|
-
|
|
427
|
-
|
|
428
|
-
|
|
429
|
-
|
|
430
|
-
|
|
431
|
-
|
|
432
|
-
|
|
433
|
-
|
|
434
|
-
|
|
435
|
-
|
|
436
|
-
|
|
437
|
-
|
|
438
|
-
|
|
439
|
-
|
|
440
|
-
|
|
441
|
-
|
|
442
|
-
|
|
443
|
-
|
|
444
|
-
|
|
445
|
-
|
|
446
|
-
|
|
447
|
-
|
|
448
|
-
|
|
449
|
-
|
|
450
|
-
|
|
451
|
-
|
|
452
|
-
|
|
453
|
-
|
|
454
|
-
|
|
455
|
-
**
|
|
456
|
-
|
|
457
|
-
|
|
458
|
-
|
|
459
|
-
|
|
460
|
-
|
|
461
|
-
|
|
462
|
-
|
|
463
|
-
|
|
464
|
-
|
|
465
|
-
|
|
466
|
-
|
|
467
|
-
|
|
468
|
-
Code
|
|
469
|
-
|
|
470
|
-
|
|
471
|
-
|
|
472
|
-
|
|
473
|
-
|
|
474
|
-
```
|
|
475
|
-
|
|
476
|
-
|
|
477
|
-
|
|
478
|
-
|
|
479
|
-
|
|
480
|
-
|
|
481
|
-
|
|
482
|
-
|
|
483
|
-
|
|
484
|
-
|
|
485
|
-
|
|
486
|
-
`turbollm launch
|
|
487
|
-
|
|
488
|
-
|
|
489
|
-
|
|
490
|
-
|
|
491
|
-
|
|
492
|
-
|
|
493
|
-
|
|
494
|
-
|
|
495
|
-
|
|
496
|
-
|
|
497
|
-
|
|
498
|
-
|
|
499
|
-
|
|
500
|
-
|
|
501
|
-
|
|
502
|
-
|
|
503
|
-
|
|
504
|
-
|
|
505
|
-
|
|
506
|
-
|
|
507
|
-
|
|
508
|
-
|
|
509
|
-
|
|
510
|
-
|
|
511
|
-
|
|
512
|
-
|
|
513
|
-
turbollm
|
|
514
|
-
turbollm --
|
|
515
|
-
turbollm
|
|
516
|
-
turbollm
|
|
517
|
-
turbollm
|
|
518
|
-
|
|
519
|
-
|
|
520
|
-
|
|
521
|
-
|
|
522
|
-
|
|
523
|
-
|
|
|
524
|
-
|
|
525
|
-
| `--
|
|
526
|
-
| `--
|
|
527
|
-
| `--
|
|
528
|
-
|
|
529
|
-
`
|
|
530
|
-
|
|
531
|
-
|
|
532
|
-
`turbollm
|
|
533
|
-
|
|
534
|
-
|
|
535
|
-
|
|
536
|
-
|
|
537
|
-
|
|
538
|
-
|
|
539
|
-
|
|
540
|
-
|
|
541
|
-
|
|
542
|
-
|
|
543
|
-
|
|
544
|
-
|
|
545
|
-
|
|
546
|
-
|
|
547
|
-
|
|
548
|
-
|
|
549
|
-
|
|
550
|
-
|
|
551
|
-
|
|
552
|
-
<
|
|
553
|
-
|
|
554
|
-
|
|
555
|
-
|
|
556
|
-
|
|
557
|
-
|
|
558
|
-
|
|
559
|
-
|
|
560
|
-
|
|
561
|
-
|
|
562
|
-
|
|
563
|
-
|
|
564
|
-
|
|
565
|
-
|
|
566
|
-
|
|
567
|
-
|
|
568
|
-
|
|
569
|
-
|
|
570
|
-
|
|
571
|
-
|
|
572
|
-
|
|
573
|
-
|
|
574
|
-
|
|
575
|
-
|
|
576
|
-
|
|
577
|
-
|
|
578
|
-
|
|
579
|
-
|
|
580
|
-
|
|
581
|
-
|
|
582
|
-
|
|
583
|
-
|
|
584
|
-
|
|
585
|
-
|
|
586
|
-
|
|
|
587
|
-
|
|
588
|
-
| **
|
|
589
|
-
|
|
|
590
|
-
|
|
|
591
|
-
|
|
|
592
|
-
|
|
|
593
|
-
|
|
|
594
|
-
|
|
|
595
|
-
|
|
|
596
|
-
|
|
597
|
-
|
|
598
|
-
|
|
599
|
-
|
|
600
|
-
|
|
601
|
-
|
|
602
|
-
|
|
603
|
-
|
|
604
|
-
|
|
605
|
-
|
|
606
|
-
|
|
607
|
-
|
|
608
|
-
|
|
609
|
-
|
|
610
|
-
|
|
611
|
-
-
|
|
612
|
-
|
|
613
|
-
|
|
614
|
-
- **
|
|
615
|
-
|
|
616
|
-
|
|
617
|
-
|
|
618
|
-
|
|
619
|
-
|
|
620
|
-
|
|
621
|
-
|
|
622
|
-
|
|
623
|
-
|
|
624
|
-
|
|
625
|
-
npm
|
|
626
|
-
|
|
627
|
-
|
|
628
|
-
npm run build
|
|
629
|
-
|
|
630
|
-
|
|
631
|
-
|
|
632
|
-
|
|
633
|
-
|
|
634
|
-
|
|
635
|
-
|
|
636
|
-
|
|
637
|
-
|
|
638
|
-
|
|
639
|
-
|
|
640
|
-
|
|
641
|
-
|
|
642
|
-
|
|
643
|
-
|
|
644
|
-
|
|
645
|
-
|
|
646
|
-
|
|
647
|
-
|
|
648
|
-
|
|
649
|
-
|
|
650
|
-
|
|
651
|
-
|
|
652
|
-
|
|
653
|
-
|
|
654
|
-
**License
|
|
655
|
-
|
|
656
|
-
|
|
657
|
-
|
|
658
|
-
|
|
659
|
-
- **
|
|
660
|
-
|
|
661
|
-
|
|
662
|
-
|
|
663
|
-
|
|
664
|
-
|
|
665
|
-
|
|
666
|
-
|
|
667
|
-
|
|
668
|
-
|
|
669
|
-
|
|
670
|
-
|
|
671
|
-
|
|
672
|
-
|
|
1
|
+
<p align="center">
|
|
2
|
+
<img src="https://raw.githubusercontent.com/mohitsoni48/TurboLLM/main/turbollm/web/public/brand/turbollm-icon-512.jpeg?v=2" width="92" height="92" alt="TurboLLM" />
|
|
3
|
+
</p>
|
|
4
|
+
|
|
5
|
+
<h1 align="center">TurboLLM</h1>
|
|
6
|
+
|
|
7
|
+
<p align="center">
|
|
8
|
+
<strong>Run <em>any</em> local LLM engine, auto-tuned to your GPU — with a polished web UI
|
|
9
|
+
and an OpenAI/Anthropic-compatible API.</strong><br/>
|
|
10
|
+
Bring your own llama.cpp fork. No compiling. No Electron. No Python. Point Claude Code at
|
|
11
|
+
your own machine in one command — fully offline.
|
|
12
|
+
</p>
|
|
13
|
+
|
|
14
|
+
<p align="center">
|
|
15
|
+
<a href="https://turbollm.dev"><img src="https://img.shields.io/badge/turbollm.dev-docs%20·%20models%20·%20guides-e2552e" alt="turbollm.dev" /></a>
|
|
16
|
+
<a href="https://www.npmjs.com/package/turbollm"><img src="https://img.shields.io/npm/v/turbollm.svg?color=e2552e" alt="npm version" /></a>
|
|
17
|
+
<a href="https://www.npmjs.com/package/turbollm"><img src="https://img.shields.io/npm/dm/turbollm.svg?color=e2552e" alt="npm downloads" /></a>
|
|
18
|
+
<img src="https://img.shields.io/badge/node-%E2%89%A522-3c873a.svg" alt="node >= 22" />
|
|
19
|
+
<img src="https://img.shields.io/badge/license-FSL--1.1--ALv2-blue.svg" alt="license" />
|
|
20
|
+
<img src="https://img.shields.io/badge/platform-Windows%20%C2%B7%20macOS%20%C2%B7%20Linux-555.svg" alt="platforms" />
|
|
21
|
+
<a href="https://ko-fi.com/mohitsoni"><img src="https://img.shields.io/badge/Ko--fi-support%20us-FF5E5B?logo=kofi&logoColor=white" alt="Ko-fi" /></a>
|
|
22
|
+
<a href="https://github.com/sponsors/mohitsoni48"><img src="https://img.shields.io/badge/GitHub%20Sponsors-support%20us-EA4AAA?logo=githubsponsors&logoColor=white" alt="GitHub Sponsors" /></a>
|
|
23
|
+
<a href="https://discord.gg/v6kRbV7nC"><img src="https://img.shields.io/badge/Discord-join%20chat-5865F2?logo=discord&logoColor=white" alt="Discord" /></a>
|
|
24
|
+
</p>
|
|
25
|
+
|
|
26
|
+
<p align="center">
|
|
27
|
+
<sub><strong>Source-available</strong> under <a href="#license">FSL-1.1</a> — free for personal and
|
|
28
|
+
internal business use; every release converts to <strong>Apache-2.0</strong> two years after it ships.</sub>
|
|
29
|
+
</p>
|
|
30
|
+
|
|
31
|
+
<!-- Brand: shipped app icon web/public/brand/turbollm-icon-512.jpeg · high-res masters web/brand-assets/ (unshipped) · in-app mark web/src/components/Logo.tsx · favicon web/public/favicon.svg -->
|
|
32
|
+
|
|
33
|
+
```bash
|
|
34
|
+
npx turbollm
|
|
35
|
+
```
|
|
36
|
+
|
|
37
|
+
That one command starts a local daemon, opens a browser UI, and serves your models over an
|
|
38
|
+
API any tool can talk to. TurboLLM is the **performance & bleeding-edge layer for local
|
|
39
|
+
LLMs** — built for people who today hand-compile forks and hunt forums for the right flags.
|
|
40
|
+
The speed is measured, not promised: **74.7 vs 61 t/s on the *same* official llama.cpp as
|
|
41
|
+
LM Studio — and 2.2× on forks it can't load** ([the numbers](#speed-turbollm-vs-lm-studio)).
|
|
42
|
+
|
|
43
|
+
<p align="center">
|
|
44
|
+
<strong>📖 Full docs · what runs on your GPU · model picks → <a href="https://turbollm.dev">turbollm.dev</a></strong>
|
|
45
|
+
</p>
|
|
46
|
+
|
|
47
|
+
<p align="center">
|
|
48
|
+
<img src="https://raw.githubusercontent.com/mohitsoni48/TurboLLM/main/assets/how-it-works.svg?v=2" width="860" alt="How TurboLLM works: clients -> one lightweight daemon -> any engine on your GPU" />
|
|
49
|
+
</p>
|
|
50
|
+
|
|
51
|
+
---
|
|
52
|
+
|
|
53
|
+
## Contents
|
|
54
|
+
|
|
55
|
+
- [Why TurboLLM](#why-turbollm)
|
|
56
|
+
- [Who this is for — and who it isn't](#who-this-is-for--and-who-it-isnt)
|
|
57
|
+
- [Speed: TurboLLM vs LM Studio](#speed-turbollm-vs-lm-studio)
|
|
58
|
+
- [Features](#features)
|
|
59
|
+
- [Quick start](#quick-start)
|
|
60
|
+
- [⭐ Bring any engine — the headline feature](#-bring-any-engine--the-headline-feature)
|
|
61
|
+
- [Run Claude Code on your own GPU](#run-claude-code-on-your-own-gpu)
|
|
62
|
+
- [Use it from any device on your network](#use-it-from-any-device-on-your-network)
|
|
63
|
+
- [Command-line reference](#command-line-reference)
|
|
64
|
+
- [Configuration & data](#configuration--data)
|
|
65
|
+
- [Requirements](#requirements)
|
|
66
|
+
- [Privacy](#privacy)
|
|
67
|
+
- [How TurboLLM compares](#how-turbollm-compares)
|
|
68
|
+
- [Troubleshooting](#troubleshooting)
|
|
69
|
+
- [Develop from source](#develop-from-source)
|
|
70
|
+
- [Community](#community)
|
|
71
|
+
- [License](#license)
|
|
72
|
+
|
|
73
|
+
---
|
|
74
|
+
|
|
75
|
+
## Why TurboLLM
|
|
76
|
+
|
|
77
|
+
Local-LLM tools make two choices for you, and both cost you performance:
|
|
78
|
+
|
|
79
|
+
1. **They pick the engine.** LM Studio ships one blessed runtime; Ollama hides the engine
|
|
80
|
+
entirely. The fastest community innovations — new quant formats, speculative decoding,
|
|
81
|
+
low-bit KV cache — land in **forks** first, and you can't use them without compiling.
|
|
82
|
+
2. **They don't tell you what speed to expect**, and they don't tune the dozens of launch
|
|
83
|
+
flags (`-c`, `-ngl`, `--n-cpu-moe`, KV type, threads, flash-attn, draft models) that make
|
|
84
|
+
the difference between 20 and 80 tokens/sec on the *same* hardware.
|
|
85
|
+
|
|
86
|
+
TurboLLM does the opposite:
|
|
87
|
+
|
|
88
|
+
- **🔌 Any engine, including forks.** Point it at any `llama-server`-compatible binary — a
|
|
89
|
+
build you compiled, a community fork, or the one it auto-provisions for your GPU. It probes
|
|
90
|
+
the binary's real capabilities and adapts the UI to them. **This is the whole point.**
|
|
91
|
+
- **⚡ Auto-tuned to your hardware.** It benchmarks on load, derives fast defaults, and shows
|
|
92
|
+
a **VRAM-fit verdict before you load** — no more flag guessing.
|
|
93
|
+
- **📊 Real tokens/sec, never faked.** Speed in the model list is *measured on your machine*
|
|
94
|
+
from actual generation — live while you chat, and remembered per model.
|
|
95
|
+
- **🪶 Lightweight.** A ~7 MB npm package on Node — **no Electron, no bundled Chromium, no
|
|
96
|
+
Python**. It downloads only the engine your GPU actually needs (Vulkan ≈ 38 MB).
|
|
97
|
+
- **🔌 Drop-in APIs.** OpenAI **and** Anthropic-compatible — so Claude Code and every existing
|
|
98
|
+
tool work unchanged.
|
|
99
|
+
- **🔀 A gateway that loads models for you.** Name any model in your API request and TurboLLM
|
|
100
|
+
loads it on demand, keeping your favorites hot in a small pool — so an agent that hops between
|
|
101
|
+
models just works, with nothing to pre-wire.
|
|
102
|
+
- **🔒 Offline-first & private.** No account, no backend, no internet, **no telemetry.**
|
|
103
|
+
|
|
104
|
+
---
|
|
105
|
+
|
|
106
|
+
## Who this is for — and who it isn't
|
|
107
|
+
|
|
108
|
+
**It's for you if:**
|
|
109
|
+
|
|
110
|
+
- you're tired of hand-tuning `-ngl` / KV / offload flags for every model and never trusting
|
|
111
|
+
the speed number;
|
|
112
|
+
- you want community forks (low-bit KV cache, NextN speculative decoding, new quant formats)
|
|
113
|
+
on day 0 — without compiling anything;
|
|
114
|
+
- you're pointing agents or **Claude Code** at a local model and want that to be one command;
|
|
115
|
+
- you already have folders of GGUFs from LM Studio / Ollama and refuse to re-download them.
|
|
116
|
+
|
|
117
|
+
**It's probably not for you if:**
|
|
118
|
+
|
|
119
|
+
- you want a zero-terminal, one-click desktop installer — TurboLLM currently needs
|
|
120
|
+
[Node ≥22](#requirements) and one terminal command (a packaged installer is planned);
|
|
121
|
+
- you're happy with a single blessed runtime and stock settings — LM Studio and Ollama
|
|
122
|
+
already serve that well;
|
|
123
|
+
- you need training or fine-tuning — TurboLLM is inference only.
|
|
124
|
+
|
|
125
|
+
---
|
|
126
|
+
|
|
127
|
+
## Speed: TurboLLM vs LM Studio
|
|
128
|
+
|
|
129
|
+
Same GPU (RTX 5070 Ti 16 GB), same model, same 200K context — measured generation speed.
|
|
130
|
+
**TurboLLM is faster than LM Studio on the very same official llama.cpp, and faster still when you
|
|
131
|
+
run a community fork LM Studio can't.**
|
|
132
|
+
|
|
133
|
+
**① On official llama.cpp, TurboLLM is faster.** It auto-provisions a GPU-native engine build (CUDA
|
|
134
|
+
13 for Blackwell here) and tunes expert-offload to the layer, so at the *same* KV-cache quant it
|
|
135
|
+
beats LM Studio's bundled runtime:
|
|
136
|
+
|
|
137
|
+
| Qwen3.6-35B-A3B · 200K | TurboLLM | LM Studio | Speed-up |
|
|
138
|
+
|---|:---:|:---:|:---:|
|
|
139
|
+
| official llama.cpp — `q4_0` | **74.7 t/s** | 61.0 t/s | **1.2×** |
|
|
140
|
+
| official llama.cpp — `q8_0` | **72.3 t/s** | ~66 t/s\* | **1.1×** |
|
|
141
|
+
|
|
142
|
+
**② Run a faster engine and pull far ahead.** Because TurboLLM runs *any* engine, you can drop in
|
|
143
|
+
the **TurboQuant** fork — a llama.cpp fork with a low-bit `turbo4` KV cache that LM Studio simply
|
|
144
|
+
can't load — in one click. On a large-KV model it delivers `q8_0`-level quality at **more than
|
|
145
|
+
double the speed**:
|
|
146
|
+
|
|
147
|
+
| Qwen3.6-27B · 200K · matched quality | TurboLLM + TurboQuant | LM Studio | Speed-up |
|
|
148
|
+
|---|:---:|:---:|:---:|
|
|
149
|
+
| `turbo4` vs `q8_0` | **24.6 t/s** | 11.4 t/s | **2.2×** |
|
|
150
|
+
|
|
151
|
+
Same run, **1.7× faster prefill** too (1288 vs 757 tok/s).
|
|
152
|
+
|
|
153
|
+
<sub>\*LM Studio's `q8_0` mildly spilled VRAM at its best offload. A low-bit KV cache helps most
|
|
154
|
+
when the cache is large; TurboLLM's auto-tuner and on-screen measured t/s pick the fastest engine +
|
|
155
|
+
config for each model, so you don't have to.</sub>
|
|
156
|
+
|
|
157
|
+
---
|
|
158
|
+
|
|
159
|
+
## Features
|
|
160
|
+
|
|
161
|
+
The headline — **[running any engine, including community forks](#-bring-any-engine--the-headline-feature)** —
|
|
162
|
+
has its own section below. Everything else is grouped here; each summary is the gist, expand for
|
|
163
|
+
the detail:
|
|
164
|
+
|
|
165
|
+
<details>
|
|
166
|
+
<summary><strong>📦 Models — bring your own, or browse Hugging Face</strong></summary>
|
|
167
|
+
|
|
168
|
+
<br/>
|
|
169
|
+
|
|
170
|
+
- **Use the folders you already have.** Point TurboLLM at any directory of GGUFs — your
|
|
171
|
+
existing LM Studio folders or manual downloads — **no re-downloading.** It parses GGUF
|
|
172
|
+
metadata (arch, params, quant, context, vision) for every file. (Models pulled by
|
|
173
|
+
`ollama pull` live in Ollama's own blob store, which TurboLLM doesn't read.)
|
|
174
|
+
- **Browse & download from Hugging Face**, in-app: a live, sortable list (trending / downloads
|
|
175
|
+
/ likes / recently updated / newest) alongside a permanent detail pane — pick a quant, read
|
|
176
|
+
the **rendered model card**, and download with **resume + SHA-256 verification**. Each model
|
|
177
|
+
lands in its own folder (mirroring Hugging Face's own layout) with its **vision projector and
|
|
178
|
+
every shard of a split/multipart quant fetched alongside it automatically**. Gated models
|
|
179
|
+
(Llama, Gemma) work via your own HF token, which **never leaves your machine**.
|
|
180
|
+
- **Import from any URL** — not just Hugging Face. Paste a direct `.gguf` link (model-author
|
|
181
|
+
sites, mirrors, private servers) and it disk-space-checks and downloads through the same
|
|
182
|
+
manager; paste a Hugging Face **model page** link instead and it opens that repo's quant
|
|
183
|
+
picker so you can choose which one to download.
|
|
184
|
+
- **Quant recommendation per GPU** and a **VRAM-fit verdict** so you pick a quant that
|
|
185
|
+
actually fits before you commit.
|
|
186
|
+
- **Primary download folder**, real-time **measured t/s per model**, **delete-from-disk**, and
|
|
187
|
+
**pin your favourites to the top of the list**.
|
|
188
|
+
|
|
189
|
+
</details>
|
|
190
|
+
|
|
191
|
+
<details>
|
|
192
|
+
<summary><strong>⚡ Auto-tuning & performance</strong></summary>
|
|
193
|
+
|
|
194
|
+
<br/>
|
|
195
|
+
|
|
196
|
+
- **Auto-benchmark on load** derives fast defaults for your exact GPU.
|
|
197
|
+
- **Recommended sampling from Hugging Face** — auto-tune checks a repo's structured params /
|
|
198
|
+
`generation_config.json` sidecar first when the quantizer publishes one (exact values, no
|
|
199
|
+
guessing), then falls back to reading the model's card (and the original model behind a
|
|
200
|
+
requant) and prefills the author's recommended `temperature / top_k / top_p / min_p`. No
|
|
201
|
+
recommendation → your sampling is left untouched.
|
|
202
|
+
- **Real measured tokens/sec** in the model list — **live** while generating, **last-session**
|
|
203
|
+
when idle (never a synthetic estimate).
|
|
204
|
+
- **Full load-parameter UI**, a superset of what other tools expose: context length, GPU offload
|
|
205
|
+
(`-ngl`), **MoE CPU-offload (`--n-cpu-moe`)**, parallel slots, **independent K and V cache-quant
|
|
206
|
+
type** (incl. low-bit on supporting forks), CPU threads, flash attention, **speculative decoding
|
|
207
|
+
(NextN / MTP / draft, with a configurable draft min/max window)**, a **pinned engine port** per
|
|
208
|
+
model, and a **custom/raw flags** field for anything not exposed as its own control.
|
|
209
|
+
- **Copy the exact launch command** for a loaded model — runs the same config standalone with
|
|
210
|
+
llama.cpp/the fork directly, no TurboLLM required.
|
|
211
|
+
- **Auto-fit GPU layers / MoE offload** — an optional toggle (off by default) that hands the
|
|
212
|
+
GPU/CPU split decision to llama.cpp's own memory-fitting logic at load time instead of a fixed
|
|
213
|
+
number, honored by Auto-tune too. Useful when a large context or model doesn't fit your usual
|
|
214
|
+
fixed setting.
|
|
215
|
+
- **Fast by default:** flash attention on, NextN self-speculative decoding on for models that
|
|
216
|
+
carry a draft head, threads auto — safely gated to what your engine actually accepts.
|
|
217
|
+
- **Multi-GPU, per model** — split a model across cards (layer/row split + main-GPU pick on
|
|
218
|
+
llama.cpp, tensor-parallel on vLLM). Defaults are no-ops, so single-GPU rigs are untouched.
|
|
219
|
+
- **Saved per-model profiles, per engine** — tune once per (model, engine) pair, so
|
|
220
|
+
switching engines (or between two installs of the same engine, e.g. a fork) never
|
|
221
|
+
overwrites another engine's tuning for the same model.
|
|
222
|
+
- **Configurable VRAM headroom** (Settings → Models & loading → Advanced, 300 MB–2 GB, default
|
|
223
|
+
1 GB) — tell auto-tune how much VRAM to keep free for other GPU workloads instead of a fixed margin.
|
|
224
|
+
Drag it to 0 to opt into an experimental **MoE "VRAM-spill" search** — auto-tune keeps pushing more
|
|
225
|
+
experts onto the GPU past the safe margin as long as both generation and prompt-processing speed
|
|
226
|
+
keep improving.
|
|
227
|
+
|
|
228
|
+
</details>
|
|
229
|
+
|
|
230
|
+
<details>
|
|
231
|
+
<summary><strong>💬 Chat & agentic tools — a genuinely good UI, not an afterthought</strong></summary>
|
|
232
|
+
|
|
233
|
+
<br/>
|
|
234
|
+
|
|
235
|
+
- **Streaming** with a **stop** button, **live tokens/sec**, **prompt-processing %** and
|
|
236
|
+
**prefill t/s**, **time-to-first-token**, **total time**, exact **token counts**, and a
|
|
237
|
+
**context-usage meter** (filled / max) on every reply.
|
|
238
|
+
- **Thinking control** — toggle reasoning **off** for a direct answer, or leave it **on** with
|
|
239
|
+
collapsible, timed "thought for N s" blocks.
|
|
240
|
+
- **Markdown + syntax-highlighted code** with one-click copy — plus **inline Unicode charts**
|
|
241
|
+
the model draws when a comparison, trend, or hierarchy is genuinely worth a visual.
|
|
242
|
+
- **Live artifacts** — `html`, `svg`, and `mermaid` replies render as **sandboxed, offline
|
|
243
|
+
previews** shown as an image, with one-click export to **PNG / JPEG / SVG / animated GIF / HTML**.
|
|
244
|
+
- **Agents** — pick a style (Default · **Designer** · Concise · Detailed · Blunt · Formal · Tutor ·
|
|
245
|
+
Creative · Research · **Lite** · **Code**) per conversation, no prompt-wrangling required. The
|
|
246
|
+
**Designer** persona produces polished, self-contained, previewable designs by default; **Lite**
|
|
247
|
+
strips the hidden prompt to the bare minimum for the fastest responses. See below for editing
|
|
248
|
+
built-ins or creating your own.
|
|
249
|
+
- **Edit or regenerate any message without losing history** — both branch instead of
|
|
250
|
+
overwriting, with a **‹ 1/2 › switcher** to flip between versions (including nested branch
|
|
251
|
+
points); delete and copy still work as before. **Persistent, searchable conversations** with
|
|
252
|
+
rename, delete, and **auto-generated titles**, organized into **drag-resizable, collapsible
|
|
253
|
+
folders** you create/rename/move conversations into.
|
|
254
|
+
- **Switching chats never cancels a reply** — an in-flight generation keeps running in the
|
|
255
|
+
background and saves normally; the sidebar shows a live indicator on any chat still generating,
|
|
256
|
+
and a dot on one that finished while you were elsewhere.
|
|
257
|
+
- **Preserve thinking across turns** (on by default) — the model's past reasoning is resent on
|
|
258
|
+
later turns, not just its final answers, so follow-ups have real context to work with.
|
|
259
|
+
- **Per-chat system prompt** and **per-chat sampling** overrides — temperature, top-p/k, min-p,
|
|
260
|
+
repeat/presence/frequency penalties, and **stop strings** — prefilled from the loaded model's
|
|
261
|
+
own recommended values, available even before you send the first message.
|
|
262
|
+
- **Image input** for vision models, **PDF and code/text attachments** (real extracted text,
|
|
263
|
+
not raw bytes), and **TurboLLM Expert** — a built-in assistant that knows the app and your
|
|
264
|
+
hardware for onboarding and troubleshooting without leaving the UI.
|
|
265
|
+
- **Agentic tools** — built-in `web_search` (Tavily), `fetch_url`, and sandboxed `run_code`, plus
|
|
266
|
+
an **MCP marketplace** in Customize: one-click connect for hosted MCPs (GitHub, Linear, Stripe,
|
|
267
|
+
Atlassian, Neon, Supabase, Cloudflare, Zapier, Apify, Mixpanel) and open-source local MCPs
|
|
268
|
+
(filesystem, git, postgres, playwright, …), plus your own custom servers. Connected tools appear
|
|
269
|
+
in every chat with no restart. A **Research** persona forces multi-step web search and cites sources inline.
|
|
270
|
+
- **Tool-call approval gate** — every tool call asks for your approval by default before it runs,
|
|
271
|
+
with **Deny**, **Allow**, **Allow for this chat**, or **Always Allow** on an inline bar above the
|
|
272
|
+
composer. Set per-tool defaults globally from Settings → Tools & safety, or flip **Auto-allow
|
|
273
|
+
all** to skip prompts entirely — a tool set to Deny still stays blocked either way.
|
|
274
|
+
- **Usage dashboard** — a GitHub-style activity heatmap of your local generation history (adaptive
|
|
275
|
+
1h/12h/24h boxes for the 7-day/30-day/all-time views), streaks, peak hour, a per-model
|
|
276
|
+
breakdown, a lifetime token-milestone tracker, and a separate **API tab** tracking tokens hitting
|
|
277
|
+
the gateway from external tools like Claude Code — not just in-app chat.
|
|
278
|
+
- **Auto-memory** *(experimental, off by default)* — silently extracts durable facts you mention
|
|
279
|
+
in chat (name, preferences, hardware) using your own loaded model, and carries them into future
|
|
280
|
+
new conversations. Nothing leaves your device; the full fact list is reviewable and deletable
|
|
281
|
+
from Settings → Memory, and turning the toggle off stops new chats from seeing it immediately.
|
|
282
|
+
- **Thinking-budget control** — a graduated slider, not just on/off: cap reasoning to a specific
|
|
283
|
+
token count, disable it entirely, or leave it unlimited. Works in Chat and Code alike.
|
|
284
|
+
|
|
285
|
+
</details>
|
|
286
|
+
|
|
287
|
+
<details>
|
|
288
|
+
<summary><strong>🧑💻 Code — a local coding agent, in a real project directory</strong></summary>
|
|
289
|
+
|
|
290
|
+
<br/>
|
|
291
|
+
|
|
292
|
+
Workspace → Code hands a task to an agent running on the same model you already have loaded —
|
|
293
|
+
point it at a repo folder, describe what you want, and it reads, edits, and runs commands to get
|
|
294
|
+
there. Entirely local; nothing leaves your machine.
|
|
295
|
+
|
|
296
|
+
- **Real repo access** — plans and edits end-to-end (or asks before mutating, or plan-only,
|
|
297
|
+
depending on mode), with an optional isolated git worktree so your actual checkout stays
|
|
298
|
+
untouched, and a real diff summary of what changed.
|
|
299
|
+
- **Persistent sessions** — archive/filter past runs, revert to any earlier message (with
|
|
300
|
+
optional real file-edit reversal), attach files as context, and a "Coding activity" dashboard
|
|
301
|
+
(sessions, tasks shipped, files touched, diff shipped, streaks) built from real history, not
|
|
302
|
+
mock data.
|
|
303
|
+
- **Real LSP integration** for TypeScript/JavaScript and Python — detects the language, installs
|
|
304
|
+
the language server if needed, and uses it for edits.
|
|
305
|
+
- **The same tools Chat gets from Customize** — any connected MCP server, plus the sandboxed
|
|
306
|
+
`run_code` tool — are available to Code too, alongside honest skill invocation and
|
|
307
|
+
`AGENTS.md`/`agents.md` support.
|
|
308
|
+
- **Independent access control** — gated behind its own API key on non-host devices, separate
|
|
309
|
+
from Chat's gate.
|
|
310
|
+
- **Expose it to other tools over MCP** — `turbollm mcp-server` runs a stdio MCP server so any
|
|
311
|
+
MCP-compatible tool (Claude Desktop, Cursor, Windsurf, Cline, Claude Code, and more) can
|
|
312
|
+
delegate a real coding task to Code, not just chat completions. Setup snippets for any host are
|
|
313
|
+
in the Developer tab.
|
|
314
|
+
|
|
315
|
+
</details>
|
|
316
|
+
|
|
317
|
+
<details>
|
|
318
|
+
<summary><strong>🤖 Customize → Agents — edit any built-in, or build your own</strong></summary>
|
|
319
|
+
|
|
320
|
+
<br/>
|
|
321
|
+
|
|
322
|
+
- **Edit any built-in agent in place** — system prompt, which shared skills it uses, and which
|
|
323
|
+
tools it may call — with a one-click **Reset** back to the original.
|
|
324
|
+
- **Create your own agent** from scratch: a name, description, and system prompt, plus a
|
|
325
|
+
checklist of which shared skills and which tools it's allowed to use (everything on by
|
|
326
|
+
default). Pick it from the same in-chat agent picker as the built-ins.
|
|
327
|
+
- **MCP tools grouped by server** in the tool checklist — one toggle selects or deselects an
|
|
328
|
+
entire server's tools at once, or expand it to pick individually.
|
|
329
|
+
|
|
330
|
+
</details>
|
|
331
|
+
|
|
332
|
+
<details>
|
|
333
|
+
<summary><strong>🔌 APIs & integrations — OpenAI + Anthropic, plus a model-loading gateway</strong></summary>
|
|
334
|
+
|
|
335
|
+
<br/>
|
|
336
|
+
|
|
337
|
+
With a model loaded, TurboLLM serves two compatible APIs on the same port:
|
|
338
|
+
|
|
339
|
+
```bash
|
|
340
|
+
# OpenAI-compatible
|
|
341
|
+
curl http://127.0.0.1:6996/v1/chat/completions \
|
|
342
|
+
-H "Content-Type: application/json" \
|
|
343
|
+
-d '{"model":"local","messages":[{"role":"user","content":"hello"}]}'
|
|
344
|
+
```
|
|
345
|
+
|
|
346
|
+
- **OpenAI-compatible** `/v1/chat/completions`, `/v1/embeddings`, … — point any OpenAI client
|
|
347
|
+
or tool at it. Embedding models are auto-detected and pooled separately, so a RAG pipeline and
|
|
348
|
+
a chat model can stay loaded side by side.
|
|
349
|
+
- **Anthropic-compatible** `/v1/messages` — including **tool use and streaming** — which powers
|
|
350
|
+
Claude Code below. No other local host offers this.
|
|
351
|
+
- **Structured output** — constrain any response to a **GBNF grammar** (or JSON shape).
|
|
352
|
+
- **API-key auth** you can require when sharing over a LAN (Settings → Network & sharing).
|
|
353
|
+
|
|
354
|
+
**The gateway loads models for you.** Most local hosts make you load a model first, then call it.
|
|
355
|
+
TurboLLM's gateway reads the `model` field of any incoming request, **fuzzy-matches it to your
|
|
356
|
+
library, and loads it on the fly** if it isn't already running — then keeps up to **four models
|
|
357
|
+
hot** in an LRU pool so the next switch is instant. An agent (or Claude Code) that hops between a
|
|
358
|
+
coding model, a vision model, and an embedder just names each one and it works — no pre-wiring.
|
|
359
|
+
|
|
360
|
+
**One command wires up more than just Claude Code.** `turbollm launch <cli>` also supports
|
|
361
|
+
`opencode`, `kilo`, `openclaw`, and `hermes` — each gets pointed at TurboLLM the way that tool
|
|
362
|
+
expects (config file merge, or its own CLI command) instead of a manual copy-paste setup. Inside
|
|
363
|
+
Claude Code, `/model` lists your local models directly (with **Auto Model Swap** on in Settings →
|
|
364
|
+
Models & loading) so you can switch mid-session instead of only at launch.
|
|
365
|
+
|
|
366
|
+
</details>
|
|
367
|
+
|
|
368
|
+
<details>
|
|
369
|
+
<summary><strong>🎨 Share the GPU with ComfyUI</strong></summary>
|
|
370
|
+
|
|
371
|
+
<br/>
|
|
372
|
+
|
|
373
|
+
If you run **ComfyUI** on the same GPU, an LLM holding VRAM while ComfyUI renders means both
|
|
374
|
+
fight for memory (and one usually OOMs). TurboLLM can hand the GPU over automatically:
|
|
375
|
+
|
|
376
|
+
- The instant ComfyUI starts a render, TurboLLM **unloads its model and pauses new loads**.
|
|
377
|
+
- When ComfyUI's queue drains, TurboLLM **reloads the exact model it unloaded**.
|
|
378
|
+
|
|
379
|
+
It's **push-based, not polling** — ComfyUI signals TurboLLM the moment a job starts/ends, so the
|
|
380
|
+
handoff is immediate and deterministic (the model is gone *before* ComfyUI executes).
|
|
381
|
+
|
|
382
|
+
**One-time setup** (Settings → Network & sharing → ComfyUI): turn on **Pause for ComfyUI**, enter your ComfyUI folder
|
|
383
|
+
(the one containing `custom_nodes`), click **Install gate** (it writes a small custom node wired to
|
|
384
|
+
this daemon), then **restart ComfyUI** once. The panel shows a live indicator (rendering / idle /
|
|
385
|
+
connected); **Remove** undoes it.
|
|
386
|
+
|
|
387
|
+
</details>
|
|
388
|
+
|
|
389
|
+
<details>
|
|
390
|
+
<summary><strong>🪶 Platform — tiny, offline, private</strong></summary>
|
|
391
|
+
|
|
392
|
+
<br/>
|
|
393
|
+
|
|
394
|
+
- A **~7 MB npm package** on Node — no Electron, no bundled Chromium, no Python.
|
|
395
|
+
- **Offline-first** — no account, no backend, no internet, no telemetry.
|
|
396
|
+
- **Windows · macOS · Linux**, with a CPU fallback when there's no GPU.
|
|
397
|
+
|
|
398
|
+
</details>
|
|
399
|
+
|
|
400
|
+
---
|
|
401
|
+
|
|
402
|
+
## Quick start
|
|
403
|
+
|
|
404
|
+
```bash
|
|
405
|
+
# run without installing (recommended for first try)
|
|
406
|
+
npx turbollm
|
|
407
|
+
|
|
408
|
+
# or install globally
|
|
409
|
+
npm install -g turbollm
|
|
410
|
+
turbollm
|
|
411
|
+
```
|
|
412
|
+
|
|
413
|
+
**On first run** the daemon:
|
|
414
|
+
|
|
415
|
+
1. Detects your GPU and **downloads a matching `llama-server` build** (CUDA for NVIDIA, ROCm
|
|
416
|
+
for AMD, Metal for Apple, SYCL for Intel, Vulkan otherwise — with a CPU fallback).
|
|
417
|
+
2. Starts on <http://127.0.0.1:6996> and opens your browser.
|
|
418
|
+
3. Drops you on the **Chat** screen, ready to load a model.
|
|
419
|
+
|
|
420
|
+
Then open **Models**, download or pick a GGUF, click **Load**, and start chatting. Stop the
|
|
421
|
+
daemon any time with **Ctrl+C**.
|
|
422
|
+
|
|
423
|
+
---
|
|
424
|
+
|
|
425
|
+
## ⭐ Bring any engine — the headline feature
|
|
426
|
+
|
|
427
|
+
No other local-LLM app lets you run **whatever inference engine you want**. TurboLLM treats
|
|
428
|
+
the engine as a swappable component.
|
|
429
|
+
|
|
430
|
+
**Add a custom engine** (Engines screen → **Add your own engine**):
|
|
431
|
+
|
|
432
|
+
1. Compile or download any `llama-server`-compatible binary — stock
|
|
433
|
+
[llama.cpp](https://github.com/ggml-org/llama.cpp), a community fork, or your own build.
|
|
434
|
+
2. Point TurboLLM at the **folder** — it scans for the `llama-server` binary, runs a
|
|
435
|
+
**capability probe**, and learns exactly which flags and features that build supports.
|
|
436
|
+
*(Optional: paste the source repo URL so TurboLLM flags when a newer build ships.)*
|
|
437
|
+
3. Activate it. The load-parameter UI **adapts to that engine** — features the build doesn't
|
|
438
|
+
support are hidden; ones it adds (e.g. low-bit KV cache, NextN) light up.
|
|
439
|
+
|
|
440
|
+
No prebuilt for your OS? The **build-from-source guide** checks your toolchain (git / CMake /
|
|
441
|
+
CUDA / a compiler — MSVC on Windows, gcc/clang on Linux), hands you the exact build commands
|
|
442
|
+
(or a 1-click **"Build it for me"** on Windows and Linux), then drops you into the folder scan
|
|
443
|
+
above.
|
|
444
|
+
|
|
445
|
+
**Or skip the manual clone entirely** — Engines screen → **Add via git repo**: paste any
|
|
446
|
+
llama.cpp-compatible fork's git URL (+ optional branch, defaults to the repo's own default) and
|
|
447
|
+
build it in-app with the same 1-click flow, no separate "point at a folder" step needed.
|
|
448
|
+
|
|
449
|
+
**Auto-provisioned default.** Don't want to fetch anything? On first run TurboLLM downloads
|
|
450
|
+
the right upstream prebuilt for your GPU automatically — and a **backend picker** lets you
|
|
451
|
+
switch between CUDA / ROCm / Metal / SYCL / Vulkan / CPU at any time (it downloads the variant
|
|
452
|
+
you choose, LM Studio-style).
|
|
453
|
+
|
|
454
|
+
**Engine types.** **llama.cpp / GGUF**, **KoboldCpp** and **llamafile** (GGUF, every OS),
|
|
455
|
+
**MLX** (macOS), and **vLLM** (Linux + NVIDIA) are all first-class engine kinds — install from
|
|
456
|
+
the curated catalog, pick the right one per model, and switch from a single dropdown.
|
|
457
|
+
|
|
458
|
+
**Fully supervised.** Every engine runs under a real state machine: health-gated readiness,
|
|
459
|
+
graceful stop, an **idle auto-stop** watchdog, and **live logs + clear error surfacing** in
|
|
460
|
+
the UI when something fails to load.
|
|
461
|
+
|
|
462
|
+
> Why it matters: fork-exclusive features — **speculative decoding (NextN / MTP / draft)**,
|
|
463
|
+
> low-bit KV cache, new quant formats — are usable on day 0, with **zero compiler knowledge**
|
|
464
|
+
> on your part beyond producing the binary (and often not even that).
|
|
465
|
+
|
|
466
|
+
---
|
|
467
|
+
|
|
468
|
+
## Run Claude Code on your own GPU
|
|
469
|
+
|
|
470
|
+
TurboLLM's Anthropic-compatible endpoint means [Claude
|
|
471
|
+
Code](https://www.npmjs.com/package/@anthropic-ai/claude-code) can run against whatever model
|
|
472
|
+
you've loaded — no cloud key, fully offline. One command wires it up:
|
|
473
|
+
|
|
474
|
+
```bash
|
|
475
|
+
turbollm launch claude # auto-loads a model if none is running, then opens Claude Code
|
|
476
|
+
turbollm launch claude --model qwen3-8b # load a specific model first, then launch
|
|
477
|
+
```
|
|
478
|
+
|
|
479
|
+
It sets Claude Code's `ANTHROPIC_BASE_URL` and pins `ANTHROPIC_MODEL` to whatever model is
|
|
480
|
+
loaded (so the status line, `/status`, and context-window sizing all reflect the real model),
|
|
481
|
+
then execs `claude`; extra args are forwarded. If no model is loaded it auto-loads your
|
|
482
|
+
last-used one (or the first in your library); `--model` picks a specific one by key or name.
|
|
483
|
+
If `claude` isn't installed, it tells you how.
|
|
484
|
+
|
|
485
|
+
Other coding CLIs work too — `turbollm launch opencode`, `turbollm launch kilo`, and
|
|
486
|
+
`turbollm launch openclaw` write a `turbollm` provider into that tool's own config file
|
|
487
|
+
(preserving anything you already configured) and then launch it. If an existing config can't
|
|
488
|
+
be parsed, TurboLLM refuses to overwrite it and prints how to add the provider by hand.
|
|
489
|
+
`turbollm launch hermes` does the equivalent for [Hermes Agent](https://github.com/NousResearch/hermes-agent)
|
|
490
|
+
via its own `hermes config set` command (its config is YAML, so we drive its CLI instead of
|
|
491
|
+
hand-editing that file). The in-app **Developer** screen also shows copy-paste snippets for
|
|
492
|
+
any OpenAI- or Anthropic-compatible tool (Open WebUI, Kilo Code, opencode, …).
|
|
493
|
+
|
|
494
|
+
---
|
|
495
|
+
|
|
496
|
+
## Use it from any device on your network
|
|
497
|
+
|
|
498
|
+
The UI runs in the browser, so any phone, tablet, or laptop on your LAN can use the model on
|
|
499
|
+
your GPU box — every screen is responsive down to phone widths, with the desktop layout
|
|
500
|
+
untouched:
|
|
501
|
+
|
|
502
|
+
```bash
|
|
503
|
+
turbollm --addr 0.0.0.0:6996 # bind all interfaces, then open http://<your-ip>:6996
|
|
504
|
+
```
|
|
505
|
+
|
|
506
|
+
Turn on **Require API key** in Settings → Network & sharing when you expose it.
|
|
507
|
+
|
|
508
|
+
---
|
|
509
|
+
|
|
510
|
+
## Command-line reference
|
|
511
|
+
|
|
512
|
+
```bash
|
|
513
|
+
turbollm # start on :6996, open browser
|
|
514
|
+
turbollm --port 9000 # listen on a specific port
|
|
515
|
+
turbollm --no-open # start without opening a browser
|
|
516
|
+
turbollm --addr 0.0.0.0:6996 # bind all interfaces (LAN sharing)
|
|
517
|
+
turbollm --stop # stop a running daemon (any terminal)
|
|
518
|
+
turbollm launch claude # start Claude Code (auto-loads a model if none is running)
|
|
519
|
+
turbollm launch claude --model qwen3-8b # load a specific model, then launch
|
|
520
|
+
turbollm launch opencode # wire opencode (or kilo / openclaw) to TurboLLM, then launch
|
|
521
|
+
```
|
|
522
|
+
|
|
523
|
+
| Flag | Description |
|
|
524
|
+
|------|-------------|
|
|
525
|
+
| `--port <n>` | Listen on a specific port (default: `6996`) |
|
|
526
|
+
| `--addr <host:port>` | Full host:port override, e.g. `0.0.0.0:6996` for LAN sharing |
|
|
527
|
+
| `--no-open` | Start without opening a browser window |
|
|
528
|
+
| `--config <file>` | Path to a custom config file |
|
|
529
|
+
| `--stop` | Stop a running TurboLLM daemon (reads `~/.turbollm/daemon.pid`) and exit |
|
|
530
|
+
| `--help`, `-h` | Show usage and exit |
|
|
531
|
+
|
|
532
|
+
`turbollm launch claude` also accepts `--model <key|name>` to load a specific model before
|
|
533
|
+
launching; without it, an already-loaded model is used, or the last-used / first model is
|
|
534
|
+
auto-loaded. The same command works with `opencode`, `kilo`, and `openclaw` — each gets a
|
|
535
|
+
`turbollm` provider merged into its own config file before it starts — and with `hermes`,
|
|
536
|
+
which gets configured via its own `hermes config set` command instead.
|
|
537
|
+
|
|
538
|
+
---
|
|
539
|
+
|
|
540
|
+
## Configuration & data
|
|
541
|
+
|
|
542
|
+
Everything lives under **`~/.turbollm/`** on every OS — `config.json`, the SQLite chat
|
|
543
|
+
database, downloaded engines, models cache, and logs. Back it up or delete it to reset.
|
|
544
|
+
Use `--config <file>` to point at an alternate config (its directory becomes the data dir).
|
|
545
|
+
|
|
546
|
+
---
|
|
547
|
+
|
|
548
|
+
## Requirements
|
|
549
|
+
|
|
550
|
+
- **Node.js 22.13.0 or newer** — enforced at startup with a clear message.
|
|
551
|
+
|
|
552
|
+
<details>
|
|
553
|
+
<summary><strong>Don't have Node?</strong> One command installs it.</summary>
|
|
554
|
+
|
|
555
|
+
<br/>
|
|
556
|
+
|
|
557
|
+
- **Windows:** `winget install OpenJS.NodeJS.LTS` (or download from <https://nodejs.org>)
|
|
558
|
+
- **macOS:** `brew install node` (or download from <https://nodejs.org>)
|
|
559
|
+
- **Linux:** use your distro's package manager or <https://nodejs.org> — make sure it's
|
|
560
|
+
v22+ (`node --version`)
|
|
561
|
+
|
|
562
|
+
Then open a **new terminal** (Windows needs one for `PATH` to refresh) and run `npx turbollm`.
|
|
563
|
+
|
|
564
|
+
</details>
|
|
565
|
+
|
|
566
|
+
- **Windows, macOS, or Linux.**
|
|
567
|
+
- A GPU is recommended but **not required** — a CPU build is provisioned as a fallback.
|
|
568
|
+
- On Windows, the first time the auto-downloaded `llama-server` runs, SmartScreen/Defender may
|
|
569
|
+
prompt (it's an upstream binary). Allow it once.
|
|
570
|
+
|
|
571
|
+
---
|
|
572
|
+
|
|
573
|
+
## Privacy
|
|
574
|
+
|
|
575
|
+
TurboLLM is **offline-first**: core local use needs no account, no backend, and no internet.
|
|
576
|
+
**No analytics or telemetry are collected.** Your prompts, chats, files, and keys never leave
|
|
577
|
+
your machine.
|
|
578
|
+
|
|
579
|
+
---
|
|
580
|
+
|
|
581
|
+
## How TurboLLM compares
|
|
582
|
+
|
|
583
|
+
Focused on the differences that matter — all four are good tools, and the others move fast.
|
|
584
|
+
Marks reflect mid-2026; verify the moving rows against each tool's current docs.
|
|
585
|
+
|
|
586
|
+
| | **TurboLLM** | LM Studio | Ollama | Open WebUI |
|
|
587
|
+
|---|:---:|:---:|:---:|:---:|
|
|
588
|
+
| Run **any engine / community forks** | ✅ | ❌ llama.cpp/MLX only | ❌ hidden | ❌ frontend |
|
|
589
|
+
| **Benchmark-based auto-tune** of launch flags | ✅ | ◐ basic offload | ◐ basic offload | ❌ |
|
|
590
|
+
| **Measured** t/s in the model list | ✅ | ◐ per-run | ◐ `--verbose` | ❌ |
|
|
591
|
+
| **Anthropic** API (`/v1/messages`) → Claude Code | ✅ | ✅ 0.4.1+ | ✅ v0.14+ | ❌ |
|
|
592
|
+
| OpenAI-compatible API | ✅ | ✅ | ✅ | ◐ proxy |
|
|
593
|
+
| Auto-load the requested model / multi-model pool | ✅ | ✅ JIT | ✅ | ❌ |
|
|
594
|
+
| Use existing model folders (no re-download) | ✅ | ◐ import | ◐ import | ❌ frontend |
|
|
595
|
+
| Speculative decoding (draft / MTP) | ✅ | ✅ | ◐ env flag | ❌ |
|
|
596
|
+
| Web UI from any LAN device | ✅ | ❌ | ❌ | ✅ |
|
|
597
|
+
| **Lightweight** (no Electron / no Python) | ✅ npm | ❌ Electron | ✅ Go | ❌ Python |
|
|
598
|
+
| Offline-first · **no telemetry** | ✅ verifiable (source-available) | ◐ closed app — not verifiable | ✅ | ✅ |
|
|
599
|
+
|
|
600
|
+
LM Studio and Ollama both added Anthropic `/v1/messages` endpoints in 2026, so the API rows are
|
|
601
|
+
now parity — Claude Code works against any of them. TurboLLM's durable edges are **any engine
|
|
602
|
+
including community forks**, **benchmark-based auto-tuning with a VRAM-fit verdict + measured t/s
|
|
603
|
+
before you commit**, and **zero telemetry**.
|
|
604
|
+
|
|
605
|
+
Prefer Open WebUI's chat breadth? It works great pointed at TurboLLM's OpenAI endpoint.
|
|
606
|
+
|
|
607
|
+
---
|
|
608
|
+
|
|
609
|
+
## Troubleshooting
|
|
610
|
+
|
|
611
|
+
- **`TurboLLM requires Node.js 22.13.0 or newer`** — upgrade Node: <https://nodejs.org>.
|
|
612
|
+
- **Model won't load / OOM** — pick a smaller quant (the VRAM verdict warns you), lower GPU
|
|
613
|
+
offload, or close other GPU apps. Failures surface in the Engines screen with the engine log.
|
|
614
|
+
- **Windows Defender / SmartScreen prompt** — that's the upstream `llama-server` binary on
|
|
615
|
+
first run; allow it once.
|
|
616
|
+
- **Port already in use** — `turbollm --port 9000`.
|
|
617
|
+
- **Slow generation** — open the model's load params; ensure GPU offload is high and flash
|
|
618
|
+
attention / NextN are on for supported models.
|
|
619
|
+
|
|
620
|
+
---
|
|
621
|
+
|
|
622
|
+
## Develop from source
|
|
623
|
+
|
|
624
|
+
```bash
|
|
625
|
+
npm install # daemon deps
|
|
626
|
+
cd web && npm install && cd ..
|
|
627
|
+
|
|
628
|
+
npm run build:web # build the React UI -> src/webdist
|
|
629
|
+
npm run start # run the daemon in dev (hot TS via tsx) -> :6996
|
|
630
|
+
|
|
631
|
+
npm run build # production bundle -> dist/cli.js (web assets included)
|
|
632
|
+
node dist/cli.js --port 6996
|
|
633
|
+
```
|
|
634
|
+
|
|
635
|
+
Frontend hot-reload: `cd web && npm run dev` (proxies `/api` and `/v1` to the daemon on
|
|
636
|
+
:6996).
|
|
637
|
+
|
|
638
|
+
**Stack:** Node ≥22.13 · TypeScript · Hono · `node:sqlite` · tsup — and a React 19 + Tailwind v4 +
|
|
639
|
+
shadcn/ui frontend. One TypeScript codebase, shipped as an npm package.
|
|
640
|
+
|
|
641
|
+
---
|
|
642
|
+
|
|
643
|
+
## Community
|
|
644
|
+
|
|
645
|
+
Questions, ideas, and show-and-tell — join the [Discord](https://discord.gg/v6kRbV7nC).
|
|
646
|
+
Want to contribute? See [CONTRIBUTING.md](https://github.com/mohitsoni48/TurboLLM/blob/main/CONTRIBUTING.md).
|
|
647
|
+
|
|
648
|
+
If TurboLLM saved you flag-hunting time, a ⭐ on the repo helps others find it.
|
|
649
|
+
|
|
650
|
+
---
|
|
651
|
+
|
|
652
|
+
## License
|
|
653
|
+
|
|
654
|
+
Source-available under the **Functional Source License 1.1 (Apache-2.0 future grant)** — SPDX
|
|
655
|
+
**`FSL-1.1-ALv2`**. Full text: [LICENSE.md](https://github.com/mohitsoni48/TurboLLM/blob/main/turbollm/LICENSE.md).
|
|
656
|
+
|
|
657
|
+
**License FAQ** (plain-language summary — the license text is what's binding):
|
|
658
|
+
|
|
659
|
+
- **Can I use it for free?** Yes. Personal use, internal business use, education, research,
|
|
660
|
+
and self-hosting — including commercially, inside your company — are all free. The license
|
|
661
|
+
permits *any* use except the one below.
|
|
662
|
+
- **What's the one restriction?** You can't take TurboLLM and ship it (or its functionality)
|
|
663
|
+
as a **competing commercial product or service**. That's the entire boundary.
|
|
664
|
+
- **Can I fork it or redistribute it?** Yes — modify, fork, and redistribute freely for any
|
|
665
|
+
permitted purpose; keep the license text and copyright notices with it.
|
|
666
|
+
- **When does it become fully open source?** The license includes an **irrevocable** grant:
|
|
667
|
+
each release converts to **Apache-2.0 exactly two years** after that release is published.
|
|
668
|
+
The first public release shipped June 2026, so it converts in June 2028 — and every release
|
|
669
|
+
after it follows on its own two-year clock.
|
|
670
|
+
- **Why not MIT/Apache from day one?** TurboLLM is built by a solo maintainer; FSL prevents a
|
|
671
|
+
large vendor from re-skinning it as their own product while it's young, and the future grant
|
|
672
|
+
guarantees the full open-source outcome anyway. (Sentry created the FSL; GitButler, PowerSync,
|
|
673
|
+
and Codecov ship under it too.)
|
|
674
|
+
|
|
675
|
+
<p align="center"><sub>Built for people who refuse to wait for the mainstream to bless the fast path. ⚡</sub></p>
|