litserve 0.2.0__tar.gz → 0.2.0.dev0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {litserve-0.2.0/src/litserve.egg-info → litserve-0.2.0.dev0}/PKG-INFO +101 -85
- {litserve-0.2.0 → litserve-0.2.0.dev0}/README.md +100 -83
- {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/__about__.py +1 -1
- {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/api.py +1 -1
- {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/examples/openai_spec_example.py +0 -8
- {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/server.py +2 -1
- {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/specs/openai.py +1 -27
- {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/utils.py +21 -1
- {litserve-0.2.0 → litserve-0.2.0.dev0/src/litserve.egg-info}/PKG-INFO +101 -85
- {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve.egg-info/requires.txt +0 -1
- {litserve-0.2.0 → litserve-0.2.0.dev0}/LICENSE +0 -0
- {litserve-0.2.0 → litserve-0.2.0.dev0}/MANIFEST.in +0 -0
- {litserve-0.2.0 → litserve-0.2.0.dev0}/requirements.txt +0 -0
- {litserve-0.2.0 → litserve-0.2.0.dev0}/setup.cfg +0 -0
- {litserve-0.2.0 → litserve-0.2.0.dev0}/setup.py +0 -0
- {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/__init__.py +0 -0
- {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/connector.py +0 -0
- {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/examples/__init__.py +0 -0
- {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/examples/simple_example.py +0 -0
- {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/python_client.py +0 -0
- {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/specs/__init__.py +0 -0
- {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/specs/base.py +0 -0
- {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve.egg-info/SOURCES.txt +0 -0
- {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve.egg-info/dependency_links.txt +0 -0
- {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve.egg-info/not-zip-safe +0 -0
- {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve.egg-info/top_level.txt +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.1
|
|
2
2
|
Name: litserve
|
|
3
|
-
Version: 0.2.0
|
|
3
|
+
Version: 0.2.0.dev0
|
|
4
4
|
Summary: Lightweight AI server.
|
|
5
5
|
Home-page: https://github.com/Lightning-AI/litserve
|
|
6
6
|
Download-URL: https://github.com/Lightning-AI/litserve
|
|
@@ -37,7 +37,6 @@ Requires-Dist: lightning>2.0.0; extra == "test"
|
|
|
37
37
|
Requires-Dist: mypy==1.11.1; extra == "test"
|
|
38
38
|
Requires-Dist: numpy<2.0; extra == "test"
|
|
39
39
|
Requires-Dist: openai>=1.12.0; extra == "test"
|
|
40
|
-
Requires-Dist: pillow; extra == "test"
|
|
41
40
|
Requires-Dist: psutil; extra == "test"
|
|
42
41
|
Requires-Dist: pytest-asyncio; extra == "test"
|
|
43
42
|
Requires-Dist: pytest-cov; extra == "test"
|
|
@@ -49,28 +48,26 @@ Requires-Dist: transformers; extra == "test"
|
|
|
49
48
|
|
|
50
49
|
<div align='center'>
|
|
51
50
|
|
|
52
|
-
# LitServe:
|
|
51
|
+
# LitServe: Deploy AI models Lightning fast ⚡
|
|
53
52
|
|
|
54
53
|
<img alt="Lightning" src="https://pl-bolts-doc-images.s3.us-east-2.amazonaws.com/app-2/ls_banner2.png" width="800px" style="max-width: 100%;">
|
|
55
54
|
|
|
56
55
|
|
|
57
56
|
|
|
58
|
-
<strong>
|
|
57
|
+
<strong>High-throughput serving engine for AI models.</strong>
|
|
59
58
|
Friendly interface. Enterprise scale.
|
|
60
59
|
</div>
|
|
61
60
|
|
|
62
61
|
----
|
|
63
62
|
|
|
64
|
-
**LitServe** is
|
|
65
|
-
|
|
66
|
-
LitServe is at least [2x faster](#performance) than plain FastAPI.
|
|
63
|
+
**LitServe** is an engine for scalable AI model deployment built on FastAPI. Features like batching, streaming, and GPU autoscaling eliminate the need to rebuild a FastAPI server for each model.
|
|
67
64
|
|
|
68
65
|
<div align='center'>
|
|
69
66
|
|
|
70
67
|
<pre>
|
|
71
|
-
✅
|
|
72
|
-
✅ Multi-modal
|
|
73
|
-
✅
|
|
68
|
+
✅ Batching ✅ Streaming ✅ Auto-GPU, multi-GPU
|
|
69
|
+
✅ Multi-modal ✅ PyTorch/JAX/TF ✅ Full control
|
|
70
|
+
✅ Auth ✅ Built on Fast API ✅ Custom specs (Open AI)
|
|
74
71
|
</pre>
|
|
75
72
|
|
|
76
73
|
<div align='center'>
|
|
@@ -84,10 +81,11 @@ LitServe is at least [2x faster](#performance) than plain FastAPI.
|
|
|
84
81
|
<div align="center">
|
|
85
82
|
<div style="text-align: center;">
|
|
86
83
|
<a href="#quick-start" style="margin: 0 10px;">Quick start</a> •
|
|
84
|
+
<a href="https://lightning.ai/" style="margin: 0 10px;">Lightning AI</a> •
|
|
87
85
|
<a href="#featured-examples" style="margin: 0 10px;">Examples</a> •
|
|
86
|
+
<a href="#deployment-options" style="margin: 0 10px;">Deploy</a> •
|
|
88
87
|
<a href="#features" style="margin: 0 10px;">Features</a> •
|
|
89
|
-
<a href="#performance" style="margin: 0 10px;">
|
|
90
|
-
<a href="#hosting-options" style="margin: 0 10px;">Hosting</a> •
|
|
88
|
+
<a href="#performance" style="margin: 0 10px;">Benchmarks</a> •
|
|
91
89
|
<a href="https://lightning.ai/docs/litserve" style="margin: 0 10px;">Docs</a>
|
|
92
90
|
</div>
|
|
93
91
|
</div>
|
|
@@ -102,16 +100,66 @@ LitServe is at least [2x faster](#performance) than plain FastAPI.
|
|
|
102
100
|
|
|
103
101
|
|
|
104
102
|
|
|
103
|
+
## Performance
|
|
104
|
+
Our benchmarks show that LitServe (built on FastAPI) handles more simultaneous requests than FastAPI and TorchServe (higher is better).
|
|
105
|
+
|
|
106
|
+
Reproduce the full benchmarks [here](https://lightning.ai/docs/litserve/home/benchmarks).
|
|
107
|
+
|
|
108
|
+
<div align="center">
|
|
109
|
+
<img alt="LitServe" src="https://pl-bolts-doc-images.s3.us-east-2.amazonaws.com/app-2/ls_charts_v6.png" width="1000px" style="max-width: 100%;">
|
|
110
|
+
</div>
|
|
111
|
+
|
|
112
|
+
These results are for image and text classification ML tasks. The performance relationships hold for other ML tasks (embedding, LLM serving, audio, segmentation, object detection, summarization etc...).
|
|
113
|
+
|
|
114
|
+
***💡 Note on LLM serving:*** For high-performance LLM serving (like Ollama/VLLM), use [LitGPT](https://github.com/Lightning-AI/litgpt?tab=readme-ov-file#deploy-an-llm) or build your custom VLLM-like server with LitServe. Optimizations like kv-caching, which can be done with LitServe, are needed to maximize LLM performance.
|
|
115
|
+
|
|
116
|
+
|
|
117
|
+
|
|
118
|
+
## Featured examples
|
|
119
|
+
|
|
120
|
+
Use LitServe to deploy any type of model or AI service (embeddings, LLMs, vision, audio, multi-modal, etc).
|
|
121
|
+
|
|
122
|
+
<table>
|
|
123
|
+
<tr>
|
|
124
|
+
<td style="vertical-align: top;">
|
|
125
|
+
<pre>
|
|
126
|
+
<strong>Featured examples</strong><br>
|
|
127
|
+
<strong>Toy model:</strong> <a href="#define-a-server">Hello world</a>
|
|
128
|
+
<strong>LLMs:</strong> <a href="https://lightning.ai/lightning-ai/studios/openai-fault-tolerant-proxy-server">Llama 3 (8B)</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-voice-clone-api-coqui-xtts-v2-model">LLM Proxy server</a>
|
|
129
|
+
<strong>NLP:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-any-hugging-face-model-instantly">Hugging face</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-hugging-face-bert-model">BERT</a>
|
|
130
|
+
<strong>Multimodal:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-clip-with-litserve">OpenAI Clip</a>
|
|
131
|
+
<strong>Audio:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-s-whisper-model">Whisper</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-music-generation-api-with-meta-s-audio-craft">AudioCraft</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-audio-generation-api">StableAudio</a>
|
|
132
|
+
<strong>Vision:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-private-api-for-stable-diffusion-2">Stable diffusion 2</a>
|
|
133
|
+
<strong>Speech:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-voice-clone-api-coqui-xtts-v2-model">Text-speech (XTTS V2)</a>
|
|
134
|
+
</pre>
|
|
135
|
+
</td>
|
|
136
|
+
<td style="vertical-align: top;">
|
|
137
|
+
<pre>
|
|
138
|
+
<strong>Key features</strong><br>
|
|
139
|
+
✅ <strong>Serve all models:</strong> LLMs, vision, etc
|
|
140
|
+
✅ <strong>All frameworks: </strong> PyTorch/Jax/sklearn/..
|
|
141
|
+
✅ <strong>Dev friendly: </strong> build AI, not infra
|
|
142
|
+
✅ <strong>Easy interface: </strong> no abstractions
|
|
143
|
+
✅ <strong>Enterprise scale:</strong> scale huge models
|
|
144
|
+
✅ <strong>Auto GPU scaling:</strong> zero code changes
|
|
145
|
+
✅ <strong>Self host: </strong> or run on Studios
|
|
146
|
+
</pre>
|
|
147
|
+
</td>
|
|
148
|
+
</tr>
|
|
149
|
+
</table>
|
|
150
|
+
|
|
151
|
+
|
|
152
|
+
|
|
105
153
|
# Quick start
|
|
106
154
|
|
|
107
|
-
Install LitServe via pip ([
|
|
155
|
+
Install LitServe via pip (or [advanced installs](https://lightning.ai/docs/litserve/home/install)):
|
|
108
156
|
|
|
109
157
|
```bash
|
|
110
158
|
pip install litserve
|
|
111
159
|
```
|
|
112
160
|
|
|
113
161
|
### Define a server
|
|
114
|
-
Here's a hello world example ([explore real examples](
|
|
162
|
+
Here's a hello world example ([explore real examples](https://lightning.ai/docs/litserve/examples)):
|
|
115
163
|
|
|
116
164
|
```python
|
|
117
165
|
# server.py
|
|
@@ -153,14 +201,24 @@ python server.py
|
|
|
153
201
|
|
|
154
202
|
### Query the server
|
|
155
203
|
|
|
156
|
-
Use the automatically generated LitServe client:
|
|
204
|
+
Use the automatically generated LitServe client or write your own:
|
|
157
205
|
|
|
206
|
+
<table>
|
|
207
|
+
<tr>
|
|
208
|
+
<td style="vertical-align: top;">
|
|
209
|
+
<pre>
|
|
210
|
+
<strong>Option A - Use generated client: </strong><br>
|
|
211
|
+
|
|
158
212
|
```bash
|
|
159
213
|
python client.py
|
|
160
214
|
```
|
|
215
|
+
<br>
|
|
161
216
|
|
|
162
|
-
|
|
163
|
-
|
|
217
|
+
</pre>
|
|
218
|
+
</td>
|
|
219
|
+
<td style="vertical-align: top;">
|
|
220
|
+
<pre>
|
|
221
|
+
<strong>Option B - Custom client example: </strong><br>
|
|
164
222
|
|
|
165
223
|
```python
|
|
166
224
|
import requests
|
|
@@ -169,79 +227,18 @@ response = requests.post(
|
|
|
169
227
|
json={"input": 4.0}
|
|
170
228
|
)
|
|
171
229
|
```
|
|
172
|
-
|
|
173
|
-
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
|
|
177
|
-
# Featured examples
|
|
178
|
-
Use LitServe to deploy any model or AI service: (Gen AI, classical ML, embedding servers, LLMs, vision, audio, multi-modal systems, etc...)
|
|
179
|
-
|
|
180
|
-
<div align='center'>
|
|
181
|
-
<div width='200px'>
|
|
182
|
-
<video src="https://github.com/user-attachments/assets/56655727-f5d7-4109-b60d-efc816e148c9" width='200px' controls></video>
|
|
183
|
-
</div>
|
|
184
|
-
</div>
|
|
185
|
-
|
|
186
|
-
<pre>
|
|
187
|
-
<strong>Featured examples</strong><br>
|
|
188
|
-
<strong>Toy model:</strong> <a href="#define-a-server">Hello world</a>
|
|
189
|
-
<strong>LLMs:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-private-llama-3-8b-api">Llama 3 (8B)</a>, <a href="https://lightning.ai/lightning-ai/studios/openai-fault-tolerant-proxy-server">LLM Proxy server</a>
|
|
190
|
-
<strong>NLP:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-any-hugging-face-model-instantly">Hugging face</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-hugging-face-bert-model">BERT</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-text-embedding-api-with-litserve">Text embedding API</a>
|
|
191
|
-
<strong>Multimodal:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-clip-with-litserve">OpenAI Clip</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-multi-modal-llm-with-minicpm">MiniCPM</a>, <a href="https://lightning.ai/lightning-ai/studios/run-meta-s-chameleon-30b">Chameleon 30B</a>
|
|
192
|
-
<strong>Audio:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-s-whisper-model">Whisper</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-music-generation-api-with-meta-s-audio-craft">AudioCraft</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-audio-generation-api">StableAudio</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-noise-cancellation-api-with-deepfilternet">Noise cancellation (DeepFilterNet)</a>
|
|
193
|
-
<strong>Vision:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-private-api-for-stable-diffusion-2">Stable diffusion 2</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-image-generation-api-with-auraflow">AuroraFlow</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-image-generation-api-with-flux">Flux</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-super-resolution-image-api-with-aura-sr">Image super resolution (Aura SR)</a>
|
|
194
|
-
<strong>Speech:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-voice-clone-api-coqui-xtts-v2-model">Text-speech (XTTS V2)</a>
|
|
195
|
-
<strong>Classical ML:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-random-forest-with-litserve">Random forest</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-xgboost-with-litserve">XGBoost</a>
|
|
196
|
-
<strong>Miscellaneous:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-an-media-conversion-api-with-ffmpeg">Media conversion API (ffmpeg)</a>
|
|
230
|
+
<br>
|
|
197
231
|
</pre>
|
|
198
|
-
|
|
199
|
-
|
|
232
|
+
</td>
|
|
233
|
+
</tr>
|
|
234
|
+
</table>
|
|
200
235
|
|
|
201
236
|
|
|
202
237
|
|
|
203
|
-
#
|
|
204
|
-
LitServe
|
|
238
|
+
# Deployment options
|
|
239
|
+
Self-manage LitServe deployments (just run it on any machine!), or deploy with one click on [Lightning AI](https://lightning.ai/).
|
|
205
240
|
|
|
206
|
-
|
|
207
|
-
✅ [Self host on your own machines](https://lightning.ai/docs/litserve/features/hosting-methods#host-on-your-own)
|
|
208
|
-
✅ [Host fully managed on Lightning AI](https://lightning.ai/docs/litserve/features/hosting-methods#host-on-lightning-studios)
|
|
209
|
-
✅ [Serve all models: LLMs, vision, time series, etc...](https://lightning.ai/docs/litserve/examples)
|
|
210
|
-
✅ [Auto-GPU scaling](https://lightning.ai/docs/litserve/features/gpu-inference)
|
|
211
|
-
✅ [Authentication](https://lightning.ai/docs/litserve/features/authentication)
|
|
212
|
-
✅ [Autoscaling](https://lightning.ai/docs/litserve/features/autoscaling)
|
|
213
|
-
✅ [Batching](https://lightning.ai/docs/litserve/features/batching)
|
|
214
|
-
✅ [Streaming](https://lightning.ai/docs/litserve/features/streaming)
|
|
215
|
-
✅ [Scale to zero (serverless)](https://lightning.ai/docs/litserve/features/streaming)
|
|
216
|
-
✅ [All ML frameworks: PyTorch, Jax, Tensorflow, Hugging Face...](https://lightning.ai/docs/litserve/features/full-control)
|
|
217
|
-
✅ [Open AI compatibility](https://lightning.ai/docs/litserve/features/open-ai-spec)
|
|
218
|
-
|
|
219
|
-
[10+ features...](https://lightning.ai/docs/litserve/features)
|
|
220
|
-
|
|
221
|
-
**Note:** Our goal is not to jump on every hype train, but instead support features that scale
|
|
222
|
-
under the most demanding enterprise deployments.
|
|
223
|
-
|
|
224
|
-
|
|
225
|
-
|
|
226
|
-
# Performance
|
|
227
|
-
LitServe is highly optimized for parallel execution with native features optimized to scale AI workloads. Our benchmarks show that LitServe (built on FastAPI) handles more simultaneous requests than FastAPI and TorchServe (higher is better).
|
|
228
|
-
|
|
229
|
-
Reproduce the full benchmarks [here](https://lightning.ai/docs/litserve/home/benchmarks).
|
|
230
|
-
|
|
231
|
-
<div align="center">
|
|
232
|
-
<img alt="LitServe" src="https://pl-bolts-doc-images.s3.us-east-2.amazonaws.com/app-2/ls_charts_v6.png" width="1000px" style="max-width: 100%;">
|
|
233
|
-
</div>
|
|
234
|
-
|
|
235
|
-
These results are for image and text classification ML tasks. The performance relationships hold for other ML tasks (embedding, LLM serving, audio, segmentation, object detection, summarization etc...).
|
|
236
|
-
|
|
237
|
-
***💡 Note on LLM serving:*** For high-performance LLM serving (like Ollama/VLLM), use [LitGPT](https://github.com/Lightning-AI/litgpt?tab=readme-ov-file#deploy-an-llm) or build your custom VLLM-like server with LitServe. Optimizations like kv-caching, which can be done with LitServe, are needed to maximize LLM performance.
|
|
238
|
-
|
|
239
|
-
|
|
240
|
-
|
|
241
|
-
# Hosting options
|
|
242
|
-
LitServe can be hosted independently on your own machines or fully managed via Lightning Studios.
|
|
243
|
-
|
|
244
|
-
Self-hosting is ideal for hackers, students, and DIY developers, while fully managed hosting is ideal for enterprise developers needing easy autoscaling, security, release management, and 99.995% uptime and observability.
|
|
241
|
+
LitServe is developed by [Lightning AI](https://lightning.ai/) which provides infrastructure for deploying AI models.
|
|
245
242
|
|
|
246
243
|
|
|
247
244
|
|
|
@@ -271,6 +268,25 @@ Self-hosting is ideal for hackers, students, and DIY developers, while fully man
|
|
|
271
268
|
|
|
272
269
|
|
|
273
270
|
|
|
271
|
+
# Features
|
|
272
|
+
LitServe supports multiple advanced state-of-the-art features.
|
|
273
|
+
|
|
274
|
+
✅ [All model types: LLMs, vision, time series, etc...](https://lightning.ai/docs/litserve/examples)
|
|
275
|
+
✅ [Auto-GPU scaling](https://lightning.ai/docs/litserve/features/gpu-inference)
|
|
276
|
+
✅ [Authentication](https://lightning.ai/docs/litserve/features/authentication)
|
|
277
|
+
✅ [Autoscaling](https://lightning.ai/docs/litserve/features/autoscaling)
|
|
278
|
+
✅ [Batching](https://lightning.ai/docs/litserve/features/batching)
|
|
279
|
+
✅ [Streaming](https://lightning.ai/docs/litserve/features/streaming)
|
|
280
|
+
✅ [All ML frameworks: PyTorch, Jax, Tensorflow, Hugging Face...](https://lightning.ai/docs/litserve/features/full-control)
|
|
281
|
+
✅ [Open AI spec](https://lightning.ai/docs/litserve/features/open-ai-spec)
|
|
282
|
+
|
|
283
|
+
[10+ features...](https://lightning.ai/docs/litserve/features)
|
|
284
|
+
|
|
285
|
+
**Note:** Our goal is not to jump on every hype train, but instead support features that scale
|
|
286
|
+
under the most demanding enterprise deployments.
|
|
287
|
+
|
|
288
|
+
|
|
289
|
+
|
|
274
290
|
# Community
|
|
275
291
|
LitServe is a [community project accepting contributions](https://lightning.ai/docs/litserve/community) - Let's make the world's most advanced AI inference engine.
|
|
276
292
|
|
|
@@ -1,27 +1,25 @@
|
|
|
1
1
|
<div align='center'>
|
|
2
2
|
|
|
3
|
-
# LitServe:
|
|
3
|
+
# LitServe: Deploy AI models Lightning fast ⚡
|
|
4
4
|
|
|
5
5
|
<img alt="Lightning" src="https://pl-bolts-doc-images.s3.us-east-2.amazonaws.com/app-2/ls_banner2.png" width="800px" style="max-width: 100%;">
|
|
6
6
|
|
|
7
7
|
|
|
8
8
|
|
|
9
|
-
<strong>
|
|
9
|
+
<strong>High-throughput serving engine for AI models.</strong>
|
|
10
10
|
Friendly interface. Enterprise scale.
|
|
11
11
|
</div>
|
|
12
12
|
|
|
13
13
|
----
|
|
14
14
|
|
|
15
|
-
**LitServe** is
|
|
16
|
-
|
|
17
|
-
LitServe is at least [2x faster](#performance) than plain FastAPI.
|
|
15
|
+
**LitServe** is an engine for scalable AI model deployment built on FastAPI. Features like batching, streaming, and GPU autoscaling eliminate the need to rebuild a FastAPI server for each model.
|
|
18
16
|
|
|
19
17
|
<div align='center'>
|
|
20
18
|
|
|
21
19
|
<pre>
|
|
22
|
-
✅
|
|
23
|
-
✅ Multi-modal
|
|
24
|
-
✅
|
|
20
|
+
✅ Batching ✅ Streaming ✅ Auto-GPU, multi-GPU
|
|
21
|
+
✅ Multi-modal ✅ PyTorch/JAX/TF ✅ Full control
|
|
22
|
+
✅ Auth ✅ Built on Fast API ✅ Custom specs (Open AI)
|
|
25
23
|
</pre>
|
|
26
24
|
|
|
27
25
|
<div align='center'>
|
|
@@ -35,10 +33,11 @@ LitServe is at least [2x faster](#performance) than plain FastAPI.
|
|
|
35
33
|
<div align="center">
|
|
36
34
|
<div style="text-align: center;">
|
|
37
35
|
<a href="#quick-start" style="margin: 0 10px;">Quick start</a> •
|
|
36
|
+
<a href="https://lightning.ai/" style="margin: 0 10px;">Lightning AI</a> •
|
|
38
37
|
<a href="#featured-examples" style="margin: 0 10px;">Examples</a> •
|
|
38
|
+
<a href="#deployment-options" style="margin: 0 10px;">Deploy</a> •
|
|
39
39
|
<a href="#features" style="margin: 0 10px;">Features</a> •
|
|
40
|
-
<a href="#performance" style="margin: 0 10px;">
|
|
41
|
-
<a href="#hosting-options" style="margin: 0 10px;">Hosting</a> •
|
|
40
|
+
<a href="#performance" style="margin: 0 10px;">Benchmarks</a> •
|
|
42
41
|
<a href="https://lightning.ai/docs/litserve" style="margin: 0 10px;">Docs</a>
|
|
43
42
|
</div>
|
|
44
43
|
</div>
|
|
@@ -53,16 +52,66 @@ LitServe is at least [2x faster](#performance) than plain FastAPI.
|
|
|
53
52
|
|
|
54
53
|
|
|
55
54
|
|
|
55
|
+
## Performance
|
|
56
|
+
Our benchmarks show that LitServe (built on FastAPI) handles more simultaneous requests than FastAPI and TorchServe (higher is better).
|
|
57
|
+
|
|
58
|
+
Reproduce the full benchmarks [here](https://lightning.ai/docs/litserve/home/benchmarks).
|
|
59
|
+
|
|
60
|
+
<div align="center">
|
|
61
|
+
<img alt="LitServe" src="https://pl-bolts-doc-images.s3.us-east-2.amazonaws.com/app-2/ls_charts_v6.png" width="1000px" style="max-width: 100%;">
|
|
62
|
+
</div>
|
|
63
|
+
|
|
64
|
+
These results are for image and text classification ML tasks. The performance relationships hold for other ML tasks (embedding, LLM serving, audio, segmentation, object detection, summarization etc...).
|
|
65
|
+
|
|
66
|
+
***💡 Note on LLM serving:*** For high-performance LLM serving (like Ollama/VLLM), use [LitGPT](https://github.com/Lightning-AI/litgpt?tab=readme-ov-file#deploy-an-llm) or build your custom VLLM-like server with LitServe. Optimizations like kv-caching, which can be done with LitServe, are needed to maximize LLM performance.
|
|
67
|
+
|
|
68
|
+
|
|
69
|
+
|
|
70
|
+
## Featured examples
|
|
71
|
+
|
|
72
|
+
Use LitServe to deploy any type of model or AI service (embeddings, LLMs, vision, audio, multi-modal, etc).
|
|
73
|
+
|
|
74
|
+
<table>
|
|
75
|
+
<tr>
|
|
76
|
+
<td style="vertical-align: top;">
|
|
77
|
+
<pre>
|
|
78
|
+
<strong>Featured examples</strong><br>
|
|
79
|
+
<strong>Toy model:</strong> <a href="#define-a-server">Hello world</a>
|
|
80
|
+
<strong>LLMs:</strong> <a href="https://lightning.ai/lightning-ai/studios/openai-fault-tolerant-proxy-server">Llama 3 (8B)</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-voice-clone-api-coqui-xtts-v2-model">LLM Proxy server</a>
|
|
81
|
+
<strong>NLP:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-any-hugging-face-model-instantly">Hugging face</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-hugging-face-bert-model">BERT</a>
|
|
82
|
+
<strong>Multimodal:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-clip-with-litserve">OpenAI Clip</a>
|
|
83
|
+
<strong>Audio:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-s-whisper-model">Whisper</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-music-generation-api-with-meta-s-audio-craft">AudioCraft</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-audio-generation-api">StableAudio</a>
|
|
84
|
+
<strong>Vision:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-private-api-for-stable-diffusion-2">Stable diffusion 2</a>
|
|
85
|
+
<strong>Speech:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-voice-clone-api-coqui-xtts-v2-model">Text-speech (XTTS V2)</a>
|
|
86
|
+
</pre>
|
|
87
|
+
</td>
|
|
88
|
+
<td style="vertical-align: top;">
|
|
89
|
+
<pre>
|
|
90
|
+
<strong>Key features</strong><br>
|
|
91
|
+
✅ <strong>Serve all models:</strong> LLMs, vision, etc
|
|
92
|
+
✅ <strong>All frameworks: </strong> PyTorch/Jax/sklearn/..
|
|
93
|
+
✅ <strong>Dev friendly: </strong> build AI, not infra
|
|
94
|
+
✅ <strong>Easy interface: </strong> no abstractions
|
|
95
|
+
✅ <strong>Enterprise scale:</strong> scale huge models
|
|
96
|
+
✅ <strong>Auto GPU scaling:</strong> zero code changes
|
|
97
|
+
✅ <strong>Self host: </strong> or run on Studios
|
|
98
|
+
</pre>
|
|
99
|
+
</td>
|
|
100
|
+
</tr>
|
|
101
|
+
</table>
|
|
102
|
+
|
|
103
|
+
|
|
104
|
+
|
|
56
105
|
# Quick start
|
|
57
106
|
|
|
58
|
-
Install LitServe via pip ([
|
|
107
|
+
Install LitServe via pip (or [advanced installs](https://lightning.ai/docs/litserve/home/install)):
|
|
59
108
|
|
|
60
109
|
```bash
|
|
61
110
|
pip install litserve
|
|
62
111
|
```
|
|
63
112
|
|
|
64
113
|
### Define a server
|
|
65
|
-
Here's a hello world example ([explore real examples](
|
|
114
|
+
Here's a hello world example ([explore real examples](https://lightning.ai/docs/litserve/examples)):
|
|
66
115
|
|
|
67
116
|
```python
|
|
68
117
|
# server.py
|
|
@@ -104,14 +153,24 @@ python server.py
|
|
|
104
153
|
|
|
105
154
|
### Query the server
|
|
106
155
|
|
|
107
|
-
Use the automatically generated LitServe client:
|
|
156
|
+
Use the automatically generated LitServe client or write your own:
|
|
108
157
|
|
|
158
|
+
<table>
|
|
159
|
+
<tr>
|
|
160
|
+
<td style="vertical-align: top;">
|
|
161
|
+
<pre>
|
|
162
|
+
<strong>Option A - Use generated client: </strong><br>
|
|
163
|
+
|
|
109
164
|
```bash
|
|
110
165
|
python client.py
|
|
111
166
|
```
|
|
167
|
+
<br>
|
|
112
168
|
|
|
113
|
-
|
|
114
|
-
|
|
169
|
+
</pre>
|
|
170
|
+
</td>
|
|
171
|
+
<td style="vertical-align: top;">
|
|
172
|
+
<pre>
|
|
173
|
+
<strong>Option B - Custom client example: </strong><br>
|
|
115
174
|
|
|
116
175
|
```python
|
|
117
176
|
import requests
|
|
@@ -120,79 +179,18 @@ response = requests.post(
|
|
|
120
179
|
json={"input": 4.0}
|
|
121
180
|
)
|
|
122
181
|
```
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
# Featured examples
|
|
129
|
-
Use LitServe to deploy any model or AI service: (Gen AI, classical ML, embedding servers, LLMs, vision, audio, multi-modal systems, etc...)
|
|
130
|
-
|
|
131
|
-
<div align='center'>
|
|
132
|
-
<div width='200px'>
|
|
133
|
-
<video src="https://github.com/user-attachments/assets/56655727-f5d7-4109-b60d-efc816e148c9" width='200px' controls></video>
|
|
134
|
-
</div>
|
|
135
|
-
</div>
|
|
136
|
-
|
|
137
|
-
<pre>
|
|
138
|
-
<strong>Featured examples</strong><br>
|
|
139
|
-
<strong>Toy model:</strong> <a href="#define-a-server">Hello world</a>
|
|
140
|
-
<strong>LLMs:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-private-llama-3-8b-api">Llama 3 (8B)</a>, <a href="https://lightning.ai/lightning-ai/studios/openai-fault-tolerant-proxy-server">LLM Proxy server</a>
|
|
141
|
-
<strong>NLP:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-any-hugging-face-model-instantly">Hugging face</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-hugging-face-bert-model">BERT</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-text-embedding-api-with-litserve">Text embedding API</a>
|
|
142
|
-
<strong>Multimodal:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-clip-with-litserve">OpenAI Clip</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-multi-modal-llm-with-minicpm">MiniCPM</a>, <a href="https://lightning.ai/lightning-ai/studios/run-meta-s-chameleon-30b">Chameleon 30B</a>
|
|
143
|
-
<strong>Audio:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-s-whisper-model">Whisper</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-music-generation-api-with-meta-s-audio-craft">AudioCraft</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-audio-generation-api">StableAudio</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-noise-cancellation-api-with-deepfilternet">Noise cancellation (DeepFilterNet)</a>
|
|
144
|
-
<strong>Vision:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-private-api-for-stable-diffusion-2">Stable diffusion 2</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-image-generation-api-with-auraflow">AuroraFlow</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-image-generation-api-with-flux">Flux</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-super-resolution-image-api-with-aura-sr">Image super resolution (Aura SR)</a>
|
|
145
|
-
<strong>Speech:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-voice-clone-api-coqui-xtts-v2-model">Text-speech (XTTS V2)</a>
|
|
146
|
-
<strong>Classical ML:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-random-forest-with-litserve">Random forest</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-xgboost-with-litserve">XGBoost</a>
|
|
147
|
-
<strong>Miscellaneous:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-an-media-conversion-api-with-ffmpeg">Media conversion API (ffmpeg)</a>
|
|
182
|
+
<br>
|
|
148
183
|
</pre>
|
|
149
|
-
|
|
150
|
-
|
|
184
|
+
</td>
|
|
185
|
+
</tr>
|
|
186
|
+
</table>
|
|
151
187
|
|
|
152
188
|
|
|
153
189
|
|
|
154
|
-
#
|
|
155
|
-
LitServe
|
|
190
|
+
# Deployment options
|
|
191
|
+
Self-manage LitServe deployments (just run it on any machine!), or deploy with one click on [Lightning AI](https://lightning.ai/).
|
|
156
192
|
|
|
157
|
-
|
|
158
|
-
✅ [Self host on your own machines](https://lightning.ai/docs/litserve/features/hosting-methods#host-on-your-own)
|
|
159
|
-
✅ [Host fully managed on Lightning AI](https://lightning.ai/docs/litserve/features/hosting-methods#host-on-lightning-studios)
|
|
160
|
-
✅ [Serve all models: LLMs, vision, time series, etc...](https://lightning.ai/docs/litserve/examples)
|
|
161
|
-
✅ [Auto-GPU scaling](https://lightning.ai/docs/litserve/features/gpu-inference)
|
|
162
|
-
✅ [Authentication](https://lightning.ai/docs/litserve/features/authentication)
|
|
163
|
-
✅ [Autoscaling](https://lightning.ai/docs/litserve/features/autoscaling)
|
|
164
|
-
✅ [Batching](https://lightning.ai/docs/litserve/features/batching)
|
|
165
|
-
✅ [Streaming](https://lightning.ai/docs/litserve/features/streaming)
|
|
166
|
-
✅ [Scale to zero (serverless)](https://lightning.ai/docs/litserve/features/streaming)
|
|
167
|
-
✅ [All ML frameworks: PyTorch, Jax, Tensorflow, Hugging Face...](https://lightning.ai/docs/litserve/features/full-control)
|
|
168
|
-
✅ [Open AI compatibility](https://lightning.ai/docs/litserve/features/open-ai-spec)
|
|
169
|
-
|
|
170
|
-
[10+ features...](https://lightning.ai/docs/litserve/features)
|
|
171
|
-
|
|
172
|
-
**Note:** Our goal is not to jump on every hype train, but instead support features that scale
|
|
173
|
-
under the most demanding enterprise deployments.
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
|
|
177
|
-
# Performance
|
|
178
|
-
LitServe is highly optimized for parallel execution with native features optimized to scale AI workloads. Our benchmarks show that LitServe (built on FastAPI) handles more simultaneous requests than FastAPI and TorchServe (higher is better).
|
|
179
|
-
|
|
180
|
-
Reproduce the full benchmarks [here](https://lightning.ai/docs/litserve/home/benchmarks).
|
|
181
|
-
|
|
182
|
-
<div align="center">
|
|
183
|
-
<img alt="LitServe" src="https://pl-bolts-doc-images.s3.us-east-2.amazonaws.com/app-2/ls_charts_v6.png" width="1000px" style="max-width: 100%;">
|
|
184
|
-
</div>
|
|
185
|
-
|
|
186
|
-
These results are for image and text classification ML tasks. The performance relationships hold for other ML tasks (embedding, LLM serving, audio, segmentation, object detection, summarization etc...).
|
|
187
|
-
|
|
188
|
-
***💡 Note on LLM serving:*** For high-performance LLM serving (like Ollama/VLLM), use [LitGPT](https://github.com/Lightning-AI/litgpt?tab=readme-ov-file#deploy-an-llm) or build your custom VLLM-like server with LitServe. Optimizations like kv-caching, which can be done with LitServe, are needed to maximize LLM performance.
|
|
189
|
-
|
|
190
|
-
|
|
191
|
-
|
|
192
|
-
# Hosting options
|
|
193
|
-
LitServe can be hosted independently on your own machines or fully managed via Lightning Studios.
|
|
194
|
-
|
|
195
|
-
Self-hosting is ideal for hackers, students, and DIY developers, while fully managed hosting is ideal for enterprise developers needing easy autoscaling, security, release management, and 99.995% uptime and observability.
|
|
193
|
+
LitServe is developed by [Lightning AI](https://lightning.ai/) which provides infrastructure for deploying AI models.
|
|
196
194
|
|
|
197
195
|
|
|
198
196
|
|
|
@@ -222,6 +220,25 @@ Self-hosting is ideal for hackers, students, and DIY developers, while fully man
|
|
|
222
220
|
|
|
223
221
|
|
|
224
222
|
|
|
223
|
+
# Features
|
|
224
|
+
LitServe supports multiple advanced state-of-the-art features.
|
|
225
|
+
|
|
226
|
+
✅ [All model types: LLMs, vision, time series, etc...](https://lightning.ai/docs/litserve/examples)
|
|
227
|
+
✅ [Auto-GPU scaling](https://lightning.ai/docs/litserve/features/gpu-inference)
|
|
228
|
+
✅ [Authentication](https://lightning.ai/docs/litserve/features/authentication)
|
|
229
|
+
✅ [Autoscaling](https://lightning.ai/docs/litserve/features/autoscaling)
|
|
230
|
+
✅ [Batching](https://lightning.ai/docs/litserve/features/batching)
|
|
231
|
+
✅ [Streaming](https://lightning.ai/docs/litserve/features/streaming)
|
|
232
|
+
✅ [All ML frameworks: PyTorch, Jax, Tensorflow, Hugging Face...](https://lightning.ai/docs/litserve/features/full-control)
|
|
233
|
+
✅ [Open AI spec](https://lightning.ai/docs/litserve/features/open-ai-spec)
|
|
234
|
+
|
|
235
|
+
[10+ features...](https://lightning.ai/docs/litserve/features)
|
|
236
|
+
|
|
237
|
+
**Note:** Our goal is not to jump on every hype train, but instead support features that scale
|
|
238
|
+
under the most demanding enterprise deployments.
|
|
239
|
+
|
|
240
|
+
|
|
241
|
+
|
|
225
242
|
# Community
|
|
226
243
|
LitServe is a [community project accepting contributions](https://lightning.ai/docs/litserve/community) - Let's make the world's most advanced AI inference engine.
|
|
227
244
|
|
|
@@ -11,7 +11,7 @@
|
|
|
11
11
|
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
12
12
|
# See the License for the specific language governing permissions and
|
|
13
13
|
# limitations under the License.
|
|
14
|
-
__version__ = "0.2.0"
|
|
14
|
+
__version__ = "0.2.0.dev0"
|
|
15
15
|
__author__ = "Lightning-AI et al."
|
|
16
16
|
__author_email__ = "community@lightning.ai"
|
|
17
17
|
__license__ = "Apache-2.0"
|
|
@@ -147,7 +147,7 @@ class LitAPI(ABC):
|
|
|
147
147
|
def device(self, value):
|
|
148
148
|
self._device = value
|
|
149
149
|
|
|
150
|
-
def
|
|
150
|
+
def sanitize(self, max_batch_size: int, spec: LitSpec):
|
|
151
151
|
if self.stream:
|
|
152
152
|
self._default_unbatch = self._unbatch_stream
|
|
153
153
|
else:
|
|
@@ -45,14 +45,6 @@ class TestAPIWithToolCalls(TestAPI):
|
|
|
45
45
|
)
|
|
46
46
|
|
|
47
47
|
|
|
48
|
-
class TestAPIWithStructuredOutput(TestAPI):
|
|
49
|
-
def encode_response(self, output):
|
|
50
|
-
yield ChatMessage(
|
|
51
|
-
role="assistant",
|
|
52
|
-
content='{"name": "Science Fair", "date": "Friday", "participants": ["Alice", "Bob"]}',
|
|
53
|
-
)
|
|
54
|
-
|
|
55
|
-
|
|
56
48
|
class OpenAIBatchContext(ls.LitAPI):
|
|
57
49
|
def setup(self, device: str) -> None:
|
|
58
50
|
self.model = None
|
|
@@ -449,7 +449,7 @@ class LitServer:
|
|
|
449
449
|
self.api_path = api_path
|
|
450
450
|
lit_api.stream = stream
|
|
451
451
|
lit_api.request_timeout = timeout
|
|
452
|
-
lit_api.
|
|
452
|
+
lit_api.sanitize(max_batch_size, spec=spec)
|
|
453
453
|
self.app = FastAPI(lifespan=self.lifespan)
|
|
454
454
|
self.app.response_queue_id = None
|
|
455
455
|
self.response_queue_id = None
|
|
@@ -463,6 +463,7 @@ class LitServer:
|
|
|
463
463
|
self.lit_spec = spec
|
|
464
464
|
self.workers_per_device = workers_per_device
|
|
465
465
|
self.max_batch_size = max_batch_size
|
|
466
|
+
self.timeout = timeout
|
|
466
467
|
self.batch_timeout = batch_timeout
|
|
467
468
|
self.stream = stream
|
|
468
469
|
self._connector = _Connector(accelerator=accelerator, devices=devices)
|
|
@@ -20,7 +20,7 @@ import typing
|
|
|
20
20
|
import uuid
|
|
21
21
|
from collections import deque
|
|
22
22
|
from enum import Enum
|
|
23
|
-
from typing import
|
|
23
|
+
from typing import AsyncGenerator, Dict, Iterator, List, Literal, Optional, Union
|
|
24
24
|
|
|
25
25
|
from fastapi import BackgroundTasks, HTTPException, Request, Response
|
|
26
26
|
from fastapi.responses import StreamingResponse
|
|
@@ -105,31 +105,6 @@ class ToolCall(BaseModel):
|
|
|
105
105
|
function: FunctionCall
|
|
106
106
|
|
|
107
107
|
|
|
108
|
-
class ResponseFormatText(BaseModel):
|
|
109
|
-
type: Literal["text"]
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
class ResponseFormatJSONObject(BaseModel):
|
|
113
|
-
type: Literal["json_object"]
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
class JSONSchema(BaseModel):
|
|
117
|
-
name: str
|
|
118
|
-
description: Optional[str] = None
|
|
119
|
-
schema_def: Optional[Dict[str, object]] = Field(None, alias="schema")
|
|
120
|
-
strict: Optional[bool] = False
|
|
121
|
-
|
|
122
|
-
|
|
123
|
-
class ResponseFormatJSONSchema(BaseModel):
|
|
124
|
-
json_schema: JSONSchema
|
|
125
|
-
type: Literal["json_schema"]
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
ResponseFormat = Annotated[
|
|
129
|
-
Union[ResponseFormatText, ResponseFormatJSONObject, ResponseFormatJSONSchema], "ResponseFormat"
|
|
130
|
-
]
|
|
131
|
-
|
|
132
|
-
|
|
133
108
|
class ChatMessage(BaseModel):
|
|
134
109
|
role: str
|
|
135
110
|
content: Union[str, List[Union[TextContent, ImageContent]]]
|
|
@@ -163,7 +138,6 @@ class ChatCompletionRequest(BaseModel):
|
|
|
163
138
|
user: Optional[str] = None
|
|
164
139
|
tools: Optional[List[Tool]] = None
|
|
165
140
|
tool_choice: Optional[ToolChoice] = ToolChoice.auto
|
|
166
|
-
response_format: Optional[ResponseFormat] = None
|
|
167
141
|
|
|
168
142
|
|
|
169
143
|
class ChatCompletionResponseChoice(BaseModel):
|
|
@@ -14,10 +14,12 @@
|
|
|
14
14
|
import asyncio
|
|
15
15
|
import logging
|
|
16
16
|
import pickle
|
|
17
|
-
|
|
17
|
+
import uuid
|
|
18
|
+
from typing import Coroutine, Optional
|
|
18
19
|
from contextlib import contextmanager
|
|
19
20
|
from typing import TYPE_CHECKING
|
|
20
21
|
|
|
22
|
+
|
|
21
23
|
from fastapi import HTTPException
|
|
22
24
|
from starlette.middleware.base import BaseHTTPMiddleware
|
|
23
25
|
|
|
@@ -35,6 +37,24 @@ class LitAPIStatus:
|
|
|
35
37
|
FINISH_STREAMING = "FINISH_STREAMING"
|
|
36
38
|
|
|
37
39
|
|
|
40
|
+
async def wait_for_queue_timeout(coro: Coroutine, timeout: Optional[float], uid: uuid.UUID, request_buffer: dict):
|
|
41
|
+
if timeout == -1 or timeout is False:
|
|
42
|
+
return await coro
|
|
43
|
+
|
|
44
|
+
task = asyncio.create_task(coro)
|
|
45
|
+
shield = asyncio.shield(task)
|
|
46
|
+
try:
|
|
47
|
+
return await asyncio.wait_for(shield, timeout)
|
|
48
|
+
except asyncio.TimeoutError:
|
|
49
|
+
if uid in request_buffer:
|
|
50
|
+
logger.error(
|
|
51
|
+
f"Request was waiting in the queue for too long ({timeout} seconds) and has been timed out. "
|
|
52
|
+
"You can adjust the timeout by providing the `timeout` argument to LitServe(..., timeout=30)."
|
|
53
|
+
)
|
|
54
|
+
raise HTTPException(504, "Request timed out")
|
|
55
|
+
return await task
|
|
56
|
+
|
|
57
|
+
|
|
38
58
|
def load_and_raise(response):
|
|
39
59
|
try:
|
|
40
60
|
exception = pickle.loads(response) if isinstance(response, bytes) else response
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.1
|
|
2
2
|
Name: litserve
|
|
3
|
-
Version: 0.2.0
|
|
3
|
+
Version: 0.2.0.dev0
|
|
4
4
|
Summary: Lightweight AI server.
|
|
5
5
|
Home-page: https://github.com/Lightning-AI/litserve
|
|
6
6
|
Download-URL: https://github.com/Lightning-AI/litserve
|
|
@@ -37,7 +37,6 @@ Requires-Dist: lightning>2.0.0; extra == "test"
|
|
|
37
37
|
Requires-Dist: mypy==1.11.1; extra == "test"
|
|
38
38
|
Requires-Dist: numpy<2.0; extra == "test"
|
|
39
39
|
Requires-Dist: openai>=1.12.0; extra == "test"
|
|
40
|
-
Requires-Dist: pillow; extra == "test"
|
|
41
40
|
Requires-Dist: psutil; extra == "test"
|
|
42
41
|
Requires-Dist: pytest-asyncio; extra == "test"
|
|
43
42
|
Requires-Dist: pytest-cov; extra == "test"
|
|
@@ -49,28 +48,26 @@ Requires-Dist: transformers; extra == "test"
|
|
|
49
48
|
|
|
50
49
|
<div align='center'>
|
|
51
50
|
|
|
52
|
-
# LitServe:
|
|
51
|
+
# LitServe: Deploy AI models Lightning fast ⚡
|
|
53
52
|
|
|
54
53
|
<img alt="Lightning" src="https://pl-bolts-doc-images.s3.us-east-2.amazonaws.com/app-2/ls_banner2.png" width="800px" style="max-width: 100%;">
|
|
55
54
|
|
|
56
55
|
|
|
57
56
|
|
|
58
|
-
<strong>
|
|
57
|
+
<strong>High-throughput serving engine for AI models.</strong>
|
|
59
58
|
Friendly interface. Enterprise scale.
|
|
60
59
|
</div>
|
|
61
60
|
|
|
62
61
|
----
|
|
63
62
|
|
|
64
|
-
**LitServe** is
|
|
65
|
-
|
|
66
|
-
LitServe is at least [2x faster](#performance) than plain FastAPI.
|
|
63
|
+
**LitServe** is an engine for scalable AI model deployment built on FastAPI. Features like batching, streaming, and GPU autoscaling eliminate the need to rebuild a FastAPI server for each model.
|
|
67
64
|
|
|
68
65
|
<div align='center'>
|
|
69
66
|
|
|
70
67
|
<pre>
|
|
71
|
-
✅
|
|
72
|
-
✅ Multi-modal
|
|
73
|
-
✅
|
|
68
|
+
✅ Batching ✅ Streaming ✅ Auto-GPU, multi-GPU
|
|
69
|
+
✅ Multi-modal ✅ PyTorch/JAX/TF ✅ Full control
|
|
70
|
+
✅ Auth ✅ Built on Fast API ✅ Custom specs (Open AI)
|
|
74
71
|
</pre>
|
|
75
72
|
|
|
76
73
|
<div align='center'>
|
|
@@ -84,10 +81,11 @@ LitServe is at least [2x faster](#performance) than plain FastAPI.
|
|
|
84
81
|
<div align="center">
|
|
85
82
|
<div style="text-align: center;">
|
|
86
83
|
<a href="#quick-start" style="margin: 0 10px;">Quick start</a> •
|
|
84
|
+
<a href="https://lightning.ai/" style="margin: 0 10px;">Lightning AI</a> •
|
|
87
85
|
<a href="#featured-examples" style="margin: 0 10px;">Examples</a> •
|
|
86
|
+
<a href="#deployment-options" style="margin: 0 10px;">Deploy</a> •
|
|
88
87
|
<a href="#features" style="margin: 0 10px;">Features</a> •
|
|
89
|
-
<a href="#performance" style="margin: 0 10px;">
|
|
90
|
-
<a href="#hosting-options" style="margin: 0 10px;">Hosting</a> •
|
|
88
|
+
<a href="#performance" style="margin: 0 10px;">Benchmarks</a> •
|
|
91
89
|
<a href="https://lightning.ai/docs/litserve" style="margin: 0 10px;">Docs</a>
|
|
92
90
|
</div>
|
|
93
91
|
</div>
|
|
@@ -102,16 +100,66 @@ LitServe is at least [2x faster](#performance) than plain FastAPI.
|
|
|
102
100
|
|
|
103
101
|
|
|
104
102
|
|
|
103
|
+
## Performance
|
|
104
|
+
Our benchmarks show that LitServe (built on FastAPI) handles more simultaneous requests than FastAPI and TorchServe (higher is better).
|
|
105
|
+
|
|
106
|
+
Reproduce the full benchmarks [here](https://lightning.ai/docs/litserve/home/benchmarks).
|
|
107
|
+
|
|
108
|
+
<div align="center">
|
|
109
|
+
<img alt="LitServe" src="https://pl-bolts-doc-images.s3.us-east-2.amazonaws.com/app-2/ls_charts_v6.png" width="1000px" style="max-width: 100%;">
|
|
110
|
+
</div>
|
|
111
|
+
|
|
112
|
+
These results are for image and text classification ML tasks. The performance relationships hold for other ML tasks (embedding, LLM serving, audio, segmentation, object detection, summarization etc...).
|
|
113
|
+
|
|
114
|
+
***💡 Note on LLM serving:*** For high-performance LLM serving (like Ollama/VLLM), use [LitGPT](https://github.com/Lightning-AI/litgpt?tab=readme-ov-file#deploy-an-llm) or build your custom VLLM-like server with LitServe. Optimizations like kv-caching, which can be done with LitServe, are needed to maximize LLM performance.
|
|
115
|
+
|
|
116
|
+
|
|
117
|
+
|
|
118
|
+
## Featured examples
|
|
119
|
+
|
|
120
|
+
Use LitServe to deploy any type of model or AI service (embeddings, LLMs, vision, audio, multi-modal, etc).
|
|
121
|
+
|
|
122
|
+
<table>
|
|
123
|
+
<tr>
|
|
124
|
+
<td style="vertical-align: top;">
|
|
125
|
+
<pre>
|
|
126
|
+
<strong>Featured examples</strong><br>
|
|
127
|
+
<strong>Toy model:</strong> <a href="#define-a-server">Hello world</a>
|
|
128
|
+
<strong>LLMs:</strong> <a href="https://lightning.ai/lightning-ai/studios/openai-fault-tolerant-proxy-server">Llama 3 (8B)</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-voice-clone-api-coqui-xtts-v2-model">LLM Proxy server</a>
|
|
129
|
+
<strong>NLP:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-any-hugging-face-model-instantly">Hugging face</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-hugging-face-bert-model">BERT</a>
|
|
130
|
+
<strong>Multimodal:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-clip-with-litserve">OpenAI Clip</a>
|
|
131
|
+
<strong>Audio:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-s-whisper-model">Whisper</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-music-generation-api-with-meta-s-audio-craft">AudioCraft</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-audio-generation-api">StableAudio</a>
|
|
132
|
+
<strong>Vision:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-private-api-for-stable-diffusion-2">Stable diffusion 2</a>
|
|
133
|
+
<strong>Speech:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-voice-clone-api-coqui-xtts-v2-model">Text-speech (XTTS V2)</a>
|
|
134
|
+
</pre>
|
|
135
|
+
</td>
|
|
136
|
+
<td style="vertical-align: top;">
|
|
137
|
+
<pre>
|
|
138
|
+
<strong>Key features</strong><br>
|
|
139
|
+
✅ <strong>Serve all models:</strong> LLMs, vision, etc
|
|
140
|
+
✅ <strong>All frameworks: </strong> PyTorch/Jax/sklearn/..
|
|
141
|
+
✅ <strong>Dev friendly: </strong> build AI, not infra
|
|
142
|
+
✅ <strong>Easy interface: </strong> no abstractions
|
|
143
|
+
✅ <strong>Enterprise scale:</strong> scale huge models
|
|
144
|
+
✅ <strong>Auto GPU scaling:</strong> zero code changes
|
|
145
|
+
✅ <strong>Self host: </strong> or run on Studios
|
|
146
|
+
</pre>
|
|
147
|
+
</td>
|
|
148
|
+
</tr>
|
|
149
|
+
</table>
|
|
150
|
+
|
|
151
|
+
|
|
152
|
+
|
|
105
153
|
# Quick start
|
|
106
154
|
|
|
107
|
-
Install LitServe via pip ([
|
|
155
|
+
Install LitServe via pip (or [advanced installs](https://lightning.ai/docs/litserve/home/install)):
|
|
108
156
|
|
|
109
157
|
```bash
|
|
110
158
|
pip install litserve
|
|
111
159
|
```
|
|
112
160
|
|
|
113
161
|
### Define a server
|
|
114
|
-
Here's a hello world example ([explore real examples](
|
|
162
|
+
Here's a hello world example ([explore real examples](https://lightning.ai/docs/litserve/examples)):
|
|
115
163
|
|
|
116
164
|
```python
|
|
117
165
|
# server.py
|
|
@@ -153,14 +201,24 @@ python server.py
|
|
|
153
201
|
|
|
154
202
|
### Query the server
|
|
155
203
|
|
|
156
|
-
Use the automatically generated LitServe client:
|
|
204
|
+
Use the automatically generated LitServe client or write your own:
|
|
157
205
|
|
|
206
|
+
<table>
|
|
207
|
+
<tr>
|
|
208
|
+
<td style="vertical-align: top;">
|
|
209
|
+
<pre>
|
|
210
|
+
<strong>Option A - Use generated client: </strong><br>
|
|
211
|
+
|
|
158
212
|
```bash
|
|
159
213
|
python client.py
|
|
160
214
|
```
|
|
215
|
+
<br>
|
|
161
216
|
|
|
162
|
-
|
|
163
|
-
|
|
217
|
+
</pre>
|
|
218
|
+
</td>
|
|
219
|
+
<td style="vertical-align: top;">
|
|
220
|
+
<pre>
|
|
221
|
+
<strong>Option B - Custom client example: </strong><br>
|
|
164
222
|
|
|
165
223
|
```python
|
|
166
224
|
import requests
|
|
@@ -169,79 +227,18 @@ response = requests.post(
|
|
|
169
227
|
json={"input": 4.0}
|
|
170
228
|
)
|
|
171
229
|
```
|
|
172
|
-
|
|
173
|
-
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
|
|
177
|
-
# Featured examples
|
|
178
|
-
Use LitServe to deploy any model or AI service: (Gen AI, classical ML, embedding servers, LLMs, vision, audio, multi-modal systems, etc...)
|
|
179
|
-
|
|
180
|
-
<div align='center'>
|
|
181
|
-
<div width='200px'>
|
|
182
|
-
<video src="https://github.com/user-attachments/assets/56655727-f5d7-4109-b60d-efc816e148c9" width='200px' controls></video>
|
|
183
|
-
</div>
|
|
184
|
-
</div>
|
|
185
|
-
|
|
186
|
-
<pre>
|
|
187
|
-
<strong>Featured examples</strong><br>
|
|
188
|
-
<strong>Toy model:</strong> <a href="#define-a-server">Hello world</a>
|
|
189
|
-
<strong>LLMs:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-private-llama-3-8b-api">Llama 3 (8B)</a>, <a href="https://lightning.ai/lightning-ai/studios/openai-fault-tolerant-proxy-server">LLM Proxy server</a>
|
|
190
|
-
<strong>NLP:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-any-hugging-face-model-instantly">Hugging face</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-hugging-face-bert-model">BERT</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-text-embedding-api-with-litserve">Text embedding API</a>
|
|
191
|
-
<strong>Multimodal:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-clip-with-litserve">OpenAI Clip</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-multi-modal-llm-with-minicpm">MiniCPM</a>, <a href="https://lightning.ai/lightning-ai/studios/run-meta-s-chameleon-30b">Chameleon 30B</a>
|
|
192
|
-
<strong>Audio:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-s-whisper-model">Whisper</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-music-generation-api-with-meta-s-audio-craft">AudioCraft</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-audio-generation-api">StableAudio</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-noise-cancellation-api-with-deepfilternet">Noise cancellation (DeepFilterNet)</a>
|
|
193
|
-
<strong>Vision:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-private-api-for-stable-diffusion-2">Stable diffusion 2</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-image-generation-api-with-auraflow">AuroraFlow</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-image-generation-api-with-flux">Flux</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-super-resolution-image-api-with-aura-sr">Image super resolution (Aura SR)</a>
|
|
194
|
-
<strong>Speech:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-voice-clone-api-coqui-xtts-v2-model">Text-speech (XTTS V2)</a>
|
|
195
|
-
<strong>Classical ML:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-random-forest-with-litserve">Random forest</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-xgboost-with-litserve">XGBoost</a>
|
|
196
|
-
<strong>Miscellaneous:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-an-media-conversion-api-with-ffmpeg">Media conversion API (ffmpeg)</a>
|
|
230
|
+
<br>
|
|
197
231
|
</pre>
|
|
198
|
-
|
|
199
|
-
|
|
232
|
+
</td>
|
|
233
|
+
</tr>
|
|
234
|
+
</table>
|
|
200
235
|
|
|
201
236
|
|
|
202
237
|
|
|
203
|
-
#
|
|
204
|
-
LitServe
|
|
238
|
+
# Deployment options
|
|
239
|
+
Self-manage LitServe deployments (just run it on any machine!), or deploy with one click on [Lightning AI](https://lightning.ai/).
|
|
205
240
|
|
|
206
|
-
|
|
207
|
-
✅ [Self host on your own machines](https://lightning.ai/docs/litserve/features/hosting-methods#host-on-your-own)
|
|
208
|
-
✅ [Host fully managed on Lightning AI](https://lightning.ai/docs/litserve/features/hosting-methods#host-on-lightning-studios)
|
|
209
|
-
✅ [Serve all models: LLMs, vision, time series, etc...](https://lightning.ai/docs/litserve/examples)
|
|
210
|
-
✅ [Auto-GPU scaling](https://lightning.ai/docs/litserve/features/gpu-inference)
|
|
211
|
-
✅ [Authentication](https://lightning.ai/docs/litserve/features/authentication)
|
|
212
|
-
✅ [Autoscaling](https://lightning.ai/docs/litserve/features/autoscaling)
|
|
213
|
-
✅ [Batching](https://lightning.ai/docs/litserve/features/batching)
|
|
214
|
-
✅ [Streaming](https://lightning.ai/docs/litserve/features/streaming)
|
|
215
|
-
✅ [Scale to zero (serverless)](https://lightning.ai/docs/litserve/features/streaming)
|
|
216
|
-
✅ [All ML frameworks: PyTorch, Jax, Tensorflow, Hugging Face...](https://lightning.ai/docs/litserve/features/full-control)
|
|
217
|
-
✅ [Open AI compatibility](https://lightning.ai/docs/litserve/features/open-ai-spec)
|
|
218
|
-
|
|
219
|
-
[10+ features...](https://lightning.ai/docs/litserve/features)
|
|
220
|
-
|
|
221
|
-
**Note:** Our goal is not to jump on every hype train, but instead support features that scale
|
|
222
|
-
under the most demanding enterprise deployments.
|
|
223
|
-
|
|
224
|
-
|
|
225
|
-
|
|
226
|
-
# Performance
|
|
227
|
-
LitServe is highly optimized for parallel execution with native features optimized to scale AI workloads. Our benchmarks show that LitServe (built on FastAPI) handles more simultaneous requests than FastAPI and TorchServe (higher is better).
|
|
228
|
-
|
|
229
|
-
Reproduce the full benchmarks [here](https://lightning.ai/docs/litserve/home/benchmarks).
|
|
230
|
-
|
|
231
|
-
<div align="center">
|
|
232
|
-
<img alt="LitServe" src="https://pl-bolts-doc-images.s3.us-east-2.amazonaws.com/app-2/ls_charts_v6.png" width="1000px" style="max-width: 100%;">
|
|
233
|
-
</div>
|
|
234
|
-
|
|
235
|
-
These results are for image and text classification ML tasks. The performance relationships hold for other ML tasks (embedding, LLM serving, audio, segmentation, object detection, summarization etc...).
|
|
236
|
-
|
|
237
|
-
***💡 Note on LLM serving:*** For high-performance LLM serving (like Ollama/VLLM), use [LitGPT](https://github.com/Lightning-AI/litgpt?tab=readme-ov-file#deploy-an-llm) or build your custom VLLM-like server with LitServe. Optimizations like kv-caching, which can be done with LitServe, are needed to maximize LLM performance.
|
|
238
|
-
|
|
239
|
-
|
|
240
|
-
|
|
241
|
-
# Hosting options
|
|
242
|
-
LitServe can be hosted independently on your own machines or fully managed via Lightning Studios.
|
|
243
|
-
|
|
244
|
-
Self-hosting is ideal for hackers, students, and DIY developers, while fully managed hosting is ideal for enterprise developers needing easy autoscaling, security, release management, and 99.995% uptime and observability.
|
|
241
|
+
LitServe is developed by [Lightning AI](https://lightning.ai/) which provides infrastructure for deploying AI models.
|
|
245
242
|
|
|
246
243
|
|
|
247
244
|
|
|
@@ -271,6 +268,25 @@ Self-hosting is ideal for hackers, students, and DIY developers, while fully man
|
|
|
271
268
|
|
|
272
269
|
|
|
273
270
|
|
|
271
|
+
# Features
|
|
272
|
+
LitServe supports multiple advanced state-of-the-art features.
|
|
273
|
+
|
|
274
|
+
✅ [All model types: LLMs, vision, time series, etc...](https://lightning.ai/docs/litserve/examples)
|
|
275
|
+
✅ [Auto-GPU scaling](https://lightning.ai/docs/litserve/features/gpu-inference)
|
|
276
|
+
✅ [Authentication](https://lightning.ai/docs/litserve/features/authentication)
|
|
277
|
+
✅ [Autoscaling](https://lightning.ai/docs/litserve/features/autoscaling)
|
|
278
|
+
✅ [Batching](https://lightning.ai/docs/litserve/features/batching)
|
|
279
|
+
✅ [Streaming](https://lightning.ai/docs/litserve/features/streaming)
|
|
280
|
+
✅ [All ML frameworks: PyTorch, Jax, Tensorflow, Hugging Face...](https://lightning.ai/docs/litserve/features/full-control)
|
|
281
|
+
✅ [Open AI spec](https://lightning.ai/docs/litserve/features/open-ai-spec)
|
|
282
|
+
|
|
283
|
+
[10+ features...](https://lightning.ai/docs/litserve/features)
|
|
284
|
+
|
|
285
|
+
**Note:** Our goal is not to jump on every hype train, but instead support features that scale
|
|
286
|
+
under the most demanding enterprise deployments.
|
|
287
|
+
|
|
288
|
+
|
|
289
|
+
|
|
274
290
|
# Community
|
|
275
291
|
LitServe is a [community project accepting contributions](https://lightning.ai/docs/litserve/community) - Let's make the world's most advanced AI inference engine.
|
|
276
292
|
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|