litserve 0.2.0__tar.gz → 0.2.0.dev0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (26) hide show
  1. {litserve-0.2.0/src/litserve.egg-info → litserve-0.2.0.dev0}/PKG-INFO +101 -85
  2. {litserve-0.2.0 → litserve-0.2.0.dev0}/README.md +100 -83
  3. {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/__about__.py +1 -1
  4. {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/api.py +1 -1
  5. {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/examples/openai_spec_example.py +0 -8
  6. {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/server.py +2 -1
  7. {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/specs/openai.py +1 -27
  8. {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/utils.py +21 -1
  9. {litserve-0.2.0 → litserve-0.2.0.dev0/src/litserve.egg-info}/PKG-INFO +101 -85
  10. {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve.egg-info/requires.txt +0 -1
  11. {litserve-0.2.0 → litserve-0.2.0.dev0}/LICENSE +0 -0
  12. {litserve-0.2.0 → litserve-0.2.0.dev0}/MANIFEST.in +0 -0
  13. {litserve-0.2.0 → litserve-0.2.0.dev0}/requirements.txt +0 -0
  14. {litserve-0.2.0 → litserve-0.2.0.dev0}/setup.cfg +0 -0
  15. {litserve-0.2.0 → litserve-0.2.0.dev0}/setup.py +0 -0
  16. {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/__init__.py +0 -0
  17. {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/connector.py +0 -0
  18. {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/examples/__init__.py +0 -0
  19. {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/examples/simple_example.py +0 -0
  20. {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/python_client.py +0 -0
  21. {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/specs/__init__.py +0 -0
  22. {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve/specs/base.py +0 -0
  23. {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve.egg-info/SOURCES.txt +0 -0
  24. {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve.egg-info/dependency_links.txt +0 -0
  25. {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve.egg-info/not-zip-safe +0 -0
  26. {litserve-0.2.0 → litserve-0.2.0.dev0}/src/litserve.egg-info/top_level.txt +0 -0
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.1
2
2
  Name: litserve
3
- Version: 0.2.0
3
+ Version: 0.2.0.dev0
4
4
  Summary: Lightweight AI server.
5
5
  Home-page: https://github.com/Lightning-AI/litserve
6
6
  Download-URL: https://github.com/Lightning-AI/litserve
@@ -37,7 +37,6 @@ Requires-Dist: lightning>2.0.0; extra == "test"
37
37
  Requires-Dist: mypy==1.11.1; extra == "test"
38
38
  Requires-Dist: numpy<2.0; extra == "test"
39
39
  Requires-Dist: openai>=1.12.0; extra == "test"
40
- Requires-Dist: pillow; extra == "test"
41
40
  Requires-Dist: psutil; extra == "test"
42
41
  Requires-Dist: pytest-asyncio; extra == "test"
43
42
  Requires-Dist: pytest-cov; extra == "test"
@@ -49,28 +48,26 @@ Requires-Dist: transformers; extra == "test"
49
48
 
50
49
  <div align='center'>
51
50
 
52
- # LitServe: Easily serve AI models Lightning fast ⚡
51
+ # LitServe: Deploy AI models Lightning fast ⚡
53
52
 
54
53
  <img alt="Lightning" src="https://pl-bolts-doc-images.s3.us-east-2.amazonaws.com/app-2/ls_banner2.png" width="800px" style="max-width: 100%;">
55
54
 
56
55
  &nbsp;
57
56
 
58
- <strong>Flexible, high-throughput serving engine for AI models.</strong>
57
+ <strong>High-throughput serving engine for AI models.</strong>
59
58
  Friendly interface. Enterprise scale.
60
59
  </div>
61
60
 
62
61
  ----
63
62
 
64
- **LitServe** is a flexible serving engine for AI models built on FastAPI. Features like batching, streaming, and GPU autoscaling eliminate the need to rebuild a FastAPI server per model.
65
-
66
- LitServe is at least [2x faster](#performance) than plain FastAPI.
63
+ **LitServe** is an engine for scalable AI model deployment built on FastAPI. Features like batching, streaming, and GPU autoscaling eliminate the need to rebuild a FastAPI server for each model.
67
64
 
68
65
  <div align='center'>
69
66
 
70
67
  <pre>
71
- ✅ (2x)+ faster serving ✅ Self-host or fully managed ✅ Auto-GPU, multi-GPU
72
- ✅ Multi-modal ✅ PyTorch/JAX/TF ✅ Full control
73
- ✅ Batching ✅ Built on Fast API ✅ Streaming
68
+ ✅ Batching ✅ Streaming ✅ Auto-GPU, multi-GPU
69
+ ✅ Multi-modal ✅ PyTorch/JAX/TF ✅ Full control
70
+ ✅ Auth ✅ Built on Fast API ✅ Custom specs (Open AI)
74
71
  </pre>
75
72
 
76
73
  <div align='center'>
@@ -84,10 +81,11 @@ LitServe is at least [2x faster](#performance) than plain FastAPI.
84
81
  <div align="center">
85
82
  <div style="text-align: center;">
86
83
  <a href="#quick-start" style="margin: 0 10px;">Quick start</a> •
84
+ <a href="https://lightning.ai/" style="margin: 0 10px;">Lightning AI</a> •
87
85
  <a href="#featured-examples" style="margin: 0 10px;">Examples</a> •
86
+ <a href="#deployment-options" style="margin: 0 10px;">Deploy</a> •
88
87
  <a href="#features" style="margin: 0 10px;">Features</a> •
89
- <a href="#performance" style="margin: 0 10px;">Performance</a> •
90
- <a href="#hosting-options" style="margin: 0 10px;">Hosting</a> •
88
+ <a href="#performance" style="margin: 0 10px;">Benchmarks</a> •
91
89
  <a href="https://lightning.ai/docs/litserve" style="margin: 0 10px;">Docs</a>
92
90
  </div>
93
91
  </div>
@@ -102,16 +100,66 @@ LitServe is at least [2x faster](#performance) than plain FastAPI.
102
100
 
103
101
  &nbsp;
104
102
 
103
+ ## Performance
104
+ Our benchmarks show that LitServe (built on FastAPI) handles more simultaneous requests than FastAPI and TorchServe (higher is better).
105
+
106
+ Reproduce the full benchmarks [here](https://lightning.ai/docs/litserve/home/benchmarks).
107
+
108
+ <div align="center">
109
+ <img alt="LitServe" src="https://pl-bolts-doc-images.s3.us-east-2.amazonaws.com/app-2/ls_charts_v6.png" width="1000px" style="max-width: 100%;">
110
+ </div>
111
+
112
+ These results are for image and text classification ML tasks. The performance relationships hold for other ML tasks (embedding, LLM serving, audio, segmentation, object detection, summarization etc...).
113
+
114
+ ***💡 Note on LLM serving:*** For high-performance LLM serving (like Ollama/VLLM), use [LitGPT](https://github.com/Lightning-AI/litgpt?tab=readme-ov-file#deploy-an-llm) or build your custom VLLM-like server with LitServe. Optimizations like kv-caching, which can be done with LitServe, are needed to maximize LLM performance.
115
+
116
+ &nbsp;
117
+
118
+ ## Featured examples
119
+
120
+ Use LitServe to deploy any type of model or AI service (embeddings, LLMs, vision, audio, multi-modal, etc).
121
+
122
+ <table>
123
+ <tr>
124
+ <td style="vertical-align: top;">
125
+ <pre>
126
+ <strong>Featured examples</strong><br>
127
+ <strong>Toy model:</strong> <a href="#define-a-server">Hello world</a>
128
+ <strong>LLMs:</strong> <a href="https://lightning.ai/lightning-ai/studios/openai-fault-tolerant-proxy-server">Llama 3 (8B)</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-voice-clone-api-coqui-xtts-v2-model">LLM Proxy server</a>
129
+ <strong>NLP:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-any-hugging-face-model-instantly">Hugging face</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-hugging-face-bert-model">BERT</a>
130
+ <strong>Multimodal:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-clip-with-litserve">OpenAI Clip</a>
131
+ <strong>Audio:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-s-whisper-model">Whisper</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-music-generation-api-with-meta-s-audio-craft">AudioCraft</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-audio-generation-api">StableAudio</a>
132
+ <strong>Vision:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-private-api-for-stable-diffusion-2">Stable diffusion 2</a>
133
+ <strong>Speech:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-voice-clone-api-coqui-xtts-v2-model">Text-speech (XTTS V2)</a>
134
+ </pre>
135
+ </td>
136
+ <td style="vertical-align: top;">
137
+ <pre>
138
+ <strong>Key features</strong><br>
139
+ ✅ <strong>Serve all models:</strong> LLMs, vision, etc
140
+ ✅ <strong>All frameworks: </strong> PyTorch/Jax/sklearn/..
141
+ ✅ <strong>Dev friendly: </strong> build AI, not infra
142
+ ✅ <strong>Easy interface: </strong> no abstractions
143
+ ✅ <strong>Enterprise scale:</strong> scale huge models
144
+ ✅ <strong>Auto GPU scaling:</strong> zero code changes
145
+ ✅ <strong>Self host: </strong> or run on Studios
146
+ </pre>
147
+ </td>
148
+ </tr>
149
+ </table>
150
+
151
+ &nbsp;
152
+
105
153
  # Quick start
106
154
 
107
- Install LitServe via pip ([other install options](https://lightning.ai/docs/litserve/home/install)):
155
+ Install LitServe via pip (or [advanced installs](https://lightning.ai/docs/litserve/home/install)):
108
156
 
109
157
  ```bash
110
158
  pip install litserve
111
159
  ```
112
160
 
113
161
  ### Define a server
114
- Here's a hello world example ([explore real examples](#featured-examples)):
162
+ Here's a hello world example ([explore real examples](https://lightning.ai/docs/litserve/examples)):
115
163
 
116
164
  ```python
117
165
  # server.py
@@ -153,14 +201,24 @@ python server.py
153
201
 
154
202
  ### Query the server
155
203
 
156
- Use the automatically generated LitServe client:
204
+ Use the automatically generated LitServe client or write your own:
157
205
 
206
+ <table>
207
+ <tr>
208
+ <td style="vertical-align: top;">
209
+ <pre>
210
+ <strong>Option A - Use generated client: </strong><br>
211
+
158
212
  ```bash
159
213
  python client.py
160
214
  ```
215
+ <br>
161
216
 
162
- <details>
163
- <summary>Write a custom client</summary>
217
+ </pre>
218
+ </td>
219
+ <td style="vertical-align: top;">
220
+ <pre>
221
+ <strong>Option B - Custom client example: </strong><br>
164
222
 
165
223
  ```python
166
224
  import requests
@@ -169,79 +227,18 @@ response = requests.post(
169
227
  json={"input": 4.0}
170
228
  )
171
229
  ```
172
- </details>
173
-
174
- &nbsp;
175
-
176
-
177
- # Featured examples
178
- Use LitServe to deploy any model or AI service: (Gen AI, classical ML, embedding servers, LLMs, vision, audio, multi-modal systems, etc...)
179
-
180
- <div align='center'>
181
- <div width='200px'>
182
- <video src="https://github.com/user-attachments/assets/56655727-f5d7-4109-b60d-efc816e148c9" width='200px' controls></video>
183
- </div>
184
- </div>
185
-
186
- <pre>
187
- <strong>Featured examples</strong><br>
188
- <strong>Toy model:</strong> <a href="#define-a-server">Hello world</a>
189
- <strong>LLMs:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-private-llama-3-8b-api">Llama 3 (8B)</a>, <a href="https://lightning.ai/lightning-ai/studios/openai-fault-tolerant-proxy-server">LLM Proxy server</a>
190
- <strong>NLP:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-any-hugging-face-model-instantly">Hugging face</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-hugging-face-bert-model">BERT</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-text-embedding-api-with-litserve">Text embedding API</a>
191
- <strong>Multimodal:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-clip-with-litserve">OpenAI Clip</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-multi-modal-llm-with-minicpm">MiniCPM</a>, <a href="https://lightning.ai/lightning-ai/studios/run-meta-s-chameleon-30b">Chameleon 30B</a>
192
- <strong>Audio:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-s-whisper-model">Whisper</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-music-generation-api-with-meta-s-audio-craft">AudioCraft</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-audio-generation-api">StableAudio</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-noise-cancellation-api-with-deepfilternet">Noise cancellation (DeepFilterNet)</a>
193
- <strong>Vision:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-private-api-for-stable-diffusion-2">Stable diffusion 2</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-image-generation-api-with-auraflow">AuroraFlow</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-image-generation-api-with-flux">Flux</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-super-resolution-image-api-with-aura-sr">Image super resolution (Aura SR)</a>
194
- <strong>Speech:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-voice-clone-api-coqui-xtts-v2-model">Text-speech (XTTS V2)</a>
195
- <strong>Classical ML:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-random-forest-with-litserve">Random forest</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-xgboost-with-litserve">XGBoost</a>
196
- <strong>Miscellaneous:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-an-media-conversion-api-with-ffmpeg">Media conversion API (ffmpeg)</a>
230
+ <br>
197
231
  </pre>
198
-
199
- [Browse 100s of community-built templates](https://lightning.ai/studios?section=serving).
232
+ </td>
233
+ </tr>
234
+ </table>
200
235
 
201
236
  &nbsp;
202
237
 
203
- # Features
204
- LitServe supports multiple advanced state-of-the-art features.
238
+ # Deployment options
239
+ Self-manage LitServe deployments (just run it on any machine!), or deploy with one click on [Lightning AI](https://lightning.ai/).
205
240
 
206
- ✅ [(2x)+ faster serving than plain FastAPI](#performance)
207
- ✅ [Self host on your own machines](https://lightning.ai/docs/litserve/features/hosting-methods#host-on-your-own)
208
- ✅ [Host fully managed on Lightning AI](https://lightning.ai/docs/litserve/features/hosting-methods#host-on-lightning-studios)
209
- ✅ [Serve all models: LLMs, vision, time series, etc...](https://lightning.ai/docs/litserve/examples)
210
- ✅ [Auto-GPU scaling](https://lightning.ai/docs/litserve/features/gpu-inference)
211
- ✅ [Authentication](https://lightning.ai/docs/litserve/features/authentication)
212
- ✅ [Autoscaling](https://lightning.ai/docs/litserve/features/autoscaling)
213
- ✅ [Batching](https://lightning.ai/docs/litserve/features/batching)
214
- ✅ [Streaming](https://lightning.ai/docs/litserve/features/streaming)
215
- ✅ [Scale to zero (serverless)](https://lightning.ai/docs/litserve/features/streaming)
216
- ✅ [All ML frameworks: PyTorch, Jax, Tensorflow, Hugging Face...](https://lightning.ai/docs/litserve/features/full-control)
217
- ✅ [Open AI compatibility](https://lightning.ai/docs/litserve/features/open-ai-spec)
218
-
219
- [10+ features...](https://lightning.ai/docs/litserve/features)
220
-
221
- **Note:** Our goal is not to jump on every hype train, but instead support features that scale
222
- under the most demanding enterprise deployments.
223
-
224
- &nbsp;
225
-
226
- # Performance
227
- LitServe is highly optimized for parallel execution with native features optimized to scale AI workloads. Our benchmarks show that LitServe (built on FastAPI) handles more simultaneous requests than FastAPI and TorchServe (higher is better).
228
-
229
- Reproduce the full benchmarks [here](https://lightning.ai/docs/litserve/home/benchmarks).
230
-
231
- <div align="center">
232
- <img alt="LitServe" src="https://pl-bolts-doc-images.s3.us-east-2.amazonaws.com/app-2/ls_charts_v6.png" width="1000px" style="max-width: 100%;">
233
- </div>
234
-
235
- These results are for image and text classification ML tasks. The performance relationships hold for other ML tasks (embedding, LLM serving, audio, segmentation, object detection, summarization etc...).
236
-
237
- ***💡 Note on LLM serving:*** For high-performance LLM serving (like Ollama/VLLM), use [LitGPT](https://github.com/Lightning-AI/litgpt?tab=readme-ov-file#deploy-an-llm) or build your custom VLLM-like server with LitServe. Optimizations like kv-caching, which can be done with LitServe, are needed to maximize LLM performance.
238
-
239
- &nbsp;
240
-
241
- # Hosting options
242
- LitServe can be hosted independently on your own machines or fully managed via Lightning Studios.
243
-
244
- Self-hosting is ideal for hackers, students, and DIY developers, while fully managed hosting is ideal for enterprise developers needing easy autoscaling, security, release management, and 99.995% uptime and observability.
241
+ LitServe is developed by [Lightning AI](https://lightning.ai/) which provides infrastructure for deploying AI models.
245
242
 
246
243
  &nbsp;
247
244
 
@@ -271,6 +268,25 @@ Self-hosting is ideal for hackers, students, and DIY developers, while fully man
271
268
 
272
269
  &nbsp;
273
270
 
271
+ # Features
272
+ LitServe supports multiple advanced state-of-the-art features.
273
+
274
+ ✅ [All model types: LLMs, vision, time series, etc...](https://lightning.ai/docs/litserve/examples)
275
+ ✅ [Auto-GPU scaling](https://lightning.ai/docs/litserve/features/gpu-inference)
276
+ ✅ [Authentication](https://lightning.ai/docs/litserve/features/authentication)
277
+ ✅ [Autoscaling](https://lightning.ai/docs/litserve/features/autoscaling)
278
+ ✅ [Batching](https://lightning.ai/docs/litserve/features/batching)
279
+ ✅ [Streaming](https://lightning.ai/docs/litserve/features/streaming)
280
+ ✅ [All ML frameworks: PyTorch, Jax, Tensorflow, Hugging Face...](https://lightning.ai/docs/litserve/features/full-control)
281
+ ✅ [Open AI spec](https://lightning.ai/docs/litserve/features/open-ai-spec)
282
+
283
+ [10+ features...](https://lightning.ai/docs/litserve/features)
284
+
285
+ **Note:** Our goal is not to jump on every hype train, but instead support features that scale
286
+ under the most demanding enterprise deployments.
287
+
288
+ &nbsp;
289
+
274
290
  # Community
275
291
  LitServe is a [community project accepting contributions](https://lightning.ai/docs/litserve/community) - Let's make the world's most advanced AI inference engine.
276
292
 
@@ -1,27 +1,25 @@
1
1
  <div align='center'>
2
2
 
3
- # LitServe: Easily serve AI models Lightning fast ⚡
3
+ # LitServe: Deploy AI models Lightning fast ⚡
4
4
 
5
5
  <img alt="Lightning" src="https://pl-bolts-doc-images.s3.us-east-2.amazonaws.com/app-2/ls_banner2.png" width="800px" style="max-width: 100%;">
6
6
 
7
7
  &nbsp;
8
8
 
9
- <strong>Flexible, high-throughput serving engine for AI models.</strong>
9
+ <strong>High-throughput serving engine for AI models.</strong>
10
10
  Friendly interface. Enterprise scale.
11
11
  </div>
12
12
 
13
13
  ----
14
14
 
15
- **LitServe** is a flexible serving engine for AI models built on FastAPI. Features like batching, streaming, and GPU autoscaling eliminate the need to rebuild a FastAPI server per model.
16
-
17
- LitServe is at least [2x faster](#performance) than plain FastAPI.
15
+ **LitServe** is an engine for scalable AI model deployment built on FastAPI. Features like batching, streaming, and GPU autoscaling eliminate the need to rebuild a FastAPI server for each model.
18
16
 
19
17
  <div align='center'>
20
18
 
21
19
  <pre>
22
- ✅ (2x)+ faster serving ✅ Self-host or fully managed ✅ Auto-GPU, multi-GPU
23
- ✅ Multi-modal ✅ PyTorch/JAX/TF ✅ Full control
24
- ✅ Batching ✅ Built on Fast API ✅ Streaming
20
+ ✅ Batching ✅ Streaming ✅ Auto-GPU, multi-GPU
21
+ ✅ Multi-modal ✅ PyTorch/JAX/TF ✅ Full control
22
+ ✅ Auth ✅ Built on Fast API ✅ Custom specs (Open AI)
25
23
  </pre>
26
24
 
27
25
  <div align='center'>
@@ -35,10 +33,11 @@ LitServe is at least [2x faster](#performance) than plain FastAPI.
35
33
  <div align="center">
36
34
  <div style="text-align: center;">
37
35
  <a href="#quick-start" style="margin: 0 10px;">Quick start</a> •
36
+ <a href="https://lightning.ai/" style="margin: 0 10px;">Lightning AI</a> •
38
37
  <a href="#featured-examples" style="margin: 0 10px;">Examples</a> •
38
+ <a href="#deployment-options" style="margin: 0 10px;">Deploy</a> •
39
39
  <a href="#features" style="margin: 0 10px;">Features</a> •
40
- <a href="#performance" style="margin: 0 10px;">Performance</a> •
41
- <a href="#hosting-options" style="margin: 0 10px;">Hosting</a> •
40
+ <a href="#performance" style="margin: 0 10px;">Benchmarks</a> •
42
41
  <a href="https://lightning.ai/docs/litserve" style="margin: 0 10px;">Docs</a>
43
42
  </div>
44
43
  </div>
@@ -53,16 +52,66 @@ LitServe is at least [2x faster](#performance) than plain FastAPI.
53
52
 
54
53
  &nbsp;
55
54
 
55
+ ## Performance
56
+ Our benchmarks show that LitServe (built on FastAPI) handles more simultaneous requests than FastAPI and TorchServe (higher is better).
57
+
58
+ Reproduce the full benchmarks [here](https://lightning.ai/docs/litserve/home/benchmarks).
59
+
60
+ <div align="center">
61
+ <img alt="LitServe" src="https://pl-bolts-doc-images.s3.us-east-2.amazonaws.com/app-2/ls_charts_v6.png" width="1000px" style="max-width: 100%;">
62
+ </div>
63
+
64
+ These results are for image and text classification ML tasks. The performance relationships hold for other ML tasks (embedding, LLM serving, audio, segmentation, object detection, summarization etc...).
65
+
66
+ ***💡 Note on LLM serving:*** For high-performance LLM serving (like Ollama/VLLM), use [LitGPT](https://github.com/Lightning-AI/litgpt?tab=readme-ov-file#deploy-an-llm) or build your custom VLLM-like server with LitServe. Optimizations like kv-caching, which can be done with LitServe, are needed to maximize LLM performance.
67
+
68
+ &nbsp;
69
+
70
+ ## Featured examples
71
+
72
+ Use LitServe to deploy any type of model or AI service (embeddings, LLMs, vision, audio, multi-modal, etc).
73
+
74
+ <table>
75
+ <tr>
76
+ <td style="vertical-align: top;">
77
+ <pre>
78
+ <strong>Featured examples</strong><br>
79
+ <strong>Toy model:</strong> <a href="#define-a-server">Hello world</a>
80
+ <strong>LLMs:</strong> <a href="https://lightning.ai/lightning-ai/studios/openai-fault-tolerant-proxy-server">Llama 3 (8B)</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-voice-clone-api-coqui-xtts-v2-model">LLM Proxy server</a>
81
+ <strong>NLP:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-any-hugging-face-model-instantly">Hugging face</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-hugging-face-bert-model">BERT</a>
82
+ <strong>Multimodal:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-clip-with-litserve">OpenAI Clip</a>
83
+ <strong>Audio:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-s-whisper-model">Whisper</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-music-generation-api-with-meta-s-audio-craft">AudioCraft</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-audio-generation-api">StableAudio</a>
84
+ <strong>Vision:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-private-api-for-stable-diffusion-2">Stable diffusion 2</a>
85
+ <strong>Speech:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-voice-clone-api-coqui-xtts-v2-model">Text-speech (XTTS V2)</a>
86
+ </pre>
87
+ </td>
88
+ <td style="vertical-align: top;">
89
+ <pre>
90
+ <strong>Key features</strong><br>
91
+ ✅ <strong>Serve all models:</strong> LLMs, vision, etc
92
+ ✅ <strong>All frameworks: </strong> PyTorch/Jax/sklearn/..
93
+ ✅ <strong>Dev friendly: </strong> build AI, not infra
94
+ ✅ <strong>Easy interface: </strong> no abstractions
95
+ ✅ <strong>Enterprise scale:</strong> scale huge models
96
+ ✅ <strong>Auto GPU scaling:</strong> zero code changes
97
+ ✅ <strong>Self host: </strong> or run on Studios
98
+ </pre>
99
+ </td>
100
+ </tr>
101
+ </table>
102
+
103
+ &nbsp;
104
+
56
105
  # Quick start
57
106
 
58
- Install LitServe via pip ([other install options](https://lightning.ai/docs/litserve/home/install)):
107
+ Install LitServe via pip (or [advanced installs](https://lightning.ai/docs/litserve/home/install)):
59
108
 
60
109
  ```bash
61
110
  pip install litserve
62
111
  ```
63
112
 
64
113
  ### Define a server
65
- Here's a hello world example ([explore real examples](#featured-examples)):
114
+ Here's a hello world example ([explore real examples](https://lightning.ai/docs/litserve/examples)):
66
115
 
67
116
  ```python
68
117
  # server.py
@@ -104,14 +153,24 @@ python server.py
104
153
 
105
154
  ### Query the server
106
155
 
107
- Use the automatically generated LitServe client:
156
+ Use the automatically generated LitServe client or write your own:
108
157
 
158
+ <table>
159
+ <tr>
160
+ <td style="vertical-align: top;">
161
+ <pre>
162
+ <strong>Option A - Use generated client: </strong><br>
163
+
109
164
  ```bash
110
165
  python client.py
111
166
  ```
167
+ <br>
112
168
 
113
- <details>
114
- <summary>Write a custom client</summary>
169
+ </pre>
170
+ </td>
171
+ <td style="vertical-align: top;">
172
+ <pre>
173
+ <strong>Option B - Custom client example: </strong><br>
115
174
 
116
175
  ```python
117
176
  import requests
@@ -120,79 +179,18 @@ response = requests.post(
120
179
  json={"input": 4.0}
121
180
  )
122
181
  ```
123
- </details>
124
-
125
- &nbsp;
126
-
127
-
128
- # Featured examples
129
- Use LitServe to deploy any model or AI service: (Gen AI, classical ML, embedding servers, LLMs, vision, audio, multi-modal systems, etc...)
130
-
131
- <div align='center'>
132
- <div width='200px'>
133
- <video src="https://github.com/user-attachments/assets/56655727-f5d7-4109-b60d-efc816e148c9" width='200px' controls></video>
134
- </div>
135
- </div>
136
-
137
- <pre>
138
- <strong>Featured examples</strong><br>
139
- <strong>Toy model:</strong> <a href="#define-a-server">Hello world</a>
140
- <strong>LLMs:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-private-llama-3-8b-api">Llama 3 (8B)</a>, <a href="https://lightning.ai/lightning-ai/studios/openai-fault-tolerant-proxy-server">LLM Proxy server</a>
141
- <strong>NLP:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-any-hugging-face-model-instantly">Hugging face</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-hugging-face-bert-model">BERT</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-text-embedding-api-with-litserve">Text embedding API</a>
142
- <strong>Multimodal:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-clip-with-litserve">OpenAI Clip</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-multi-modal-llm-with-minicpm">MiniCPM</a>, <a href="https://lightning.ai/lightning-ai/studios/run-meta-s-chameleon-30b">Chameleon 30B</a>
143
- <strong>Audio:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-s-whisper-model">Whisper</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-music-generation-api-with-meta-s-audio-craft">AudioCraft</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-audio-generation-api">StableAudio</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-noise-cancellation-api-with-deepfilternet">Noise cancellation (DeepFilterNet)</a>
144
- <strong>Vision:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-private-api-for-stable-diffusion-2">Stable diffusion 2</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-image-generation-api-with-auraflow">AuroraFlow</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-image-generation-api-with-flux">Flux</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-super-resolution-image-api-with-aura-sr">Image super resolution (Aura SR)</a>
145
- <strong>Speech:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-voice-clone-api-coqui-xtts-v2-model">Text-speech (XTTS V2)</a>
146
- <strong>Classical ML:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-random-forest-with-litserve">Random forest</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-xgboost-with-litserve">XGBoost</a>
147
- <strong>Miscellaneous:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-an-media-conversion-api-with-ffmpeg">Media conversion API (ffmpeg)</a>
182
+ <br>
148
183
  </pre>
149
-
150
- [Browse 100s of community-built templates](https://lightning.ai/studios?section=serving).
184
+ </td>
185
+ </tr>
186
+ </table>
151
187
 
152
188
  &nbsp;
153
189
 
154
- # Features
155
- LitServe supports multiple advanced state-of-the-art features.
190
+ # Deployment options
191
+ Self-manage LitServe deployments (just run it on any machine!), or deploy with one click on [Lightning AI](https://lightning.ai/).
156
192
 
157
- ✅ [(2x)+ faster serving than plain FastAPI](#performance)
158
- ✅ [Self host on your own machines](https://lightning.ai/docs/litserve/features/hosting-methods#host-on-your-own)
159
- ✅ [Host fully managed on Lightning AI](https://lightning.ai/docs/litserve/features/hosting-methods#host-on-lightning-studios)
160
- ✅ [Serve all models: LLMs, vision, time series, etc...](https://lightning.ai/docs/litserve/examples)
161
- ✅ [Auto-GPU scaling](https://lightning.ai/docs/litserve/features/gpu-inference)
162
- ✅ [Authentication](https://lightning.ai/docs/litserve/features/authentication)
163
- ✅ [Autoscaling](https://lightning.ai/docs/litserve/features/autoscaling)
164
- ✅ [Batching](https://lightning.ai/docs/litserve/features/batching)
165
- ✅ [Streaming](https://lightning.ai/docs/litserve/features/streaming)
166
- ✅ [Scale to zero (serverless)](https://lightning.ai/docs/litserve/features/streaming)
167
- ✅ [All ML frameworks: PyTorch, Jax, Tensorflow, Hugging Face...](https://lightning.ai/docs/litserve/features/full-control)
168
- ✅ [Open AI compatibility](https://lightning.ai/docs/litserve/features/open-ai-spec)
169
-
170
- [10+ features...](https://lightning.ai/docs/litserve/features)
171
-
172
- **Note:** Our goal is not to jump on every hype train, but instead support features that scale
173
- under the most demanding enterprise deployments.
174
-
175
- &nbsp;
176
-
177
- # Performance
178
- LitServe is highly optimized for parallel execution with native features optimized to scale AI workloads. Our benchmarks show that LitServe (built on FastAPI) handles more simultaneous requests than FastAPI and TorchServe (higher is better).
179
-
180
- Reproduce the full benchmarks [here](https://lightning.ai/docs/litserve/home/benchmarks).
181
-
182
- <div align="center">
183
- <img alt="LitServe" src="https://pl-bolts-doc-images.s3.us-east-2.amazonaws.com/app-2/ls_charts_v6.png" width="1000px" style="max-width: 100%;">
184
- </div>
185
-
186
- These results are for image and text classification ML tasks. The performance relationships hold for other ML tasks (embedding, LLM serving, audio, segmentation, object detection, summarization etc...).
187
-
188
- ***💡 Note on LLM serving:*** For high-performance LLM serving (like Ollama/VLLM), use [LitGPT](https://github.com/Lightning-AI/litgpt?tab=readme-ov-file#deploy-an-llm) or build your custom VLLM-like server with LitServe. Optimizations like kv-caching, which can be done with LitServe, are needed to maximize LLM performance.
189
-
190
- &nbsp;
191
-
192
- # Hosting options
193
- LitServe can be hosted independently on your own machines or fully managed via Lightning Studios.
194
-
195
- Self-hosting is ideal for hackers, students, and DIY developers, while fully managed hosting is ideal for enterprise developers needing easy autoscaling, security, release management, and 99.995% uptime and observability.
193
+ LitServe is developed by [Lightning AI](https://lightning.ai/) which provides infrastructure for deploying AI models.
196
194
 
197
195
  &nbsp;
198
196
 
@@ -222,6 +220,25 @@ Self-hosting is ideal for hackers, students, and DIY developers, while fully man
222
220
 
223
221
  &nbsp;
224
222
 
223
+ # Features
224
+ LitServe supports multiple advanced state-of-the-art features.
225
+
226
+ ✅ [All model types: LLMs, vision, time series, etc...](https://lightning.ai/docs/litserve/examples)
227
+ ✅ [Auto-GPU scaling](https://lightning.ai/docs/litserve/features/gpu-inference)
228
+ ✅ [Authentication](https://lightning.ai/docs/litserve/features/authentication)
229
+ ✅ [Autoscaling](https://lightning.ai/docs/litserve/features/autoscaling)
230
+ ✅ [Batching](https://lightning.ai/docs/litserve/features/batching)
231
+ ✅ [Streaming](https://lightning.ai/docs/litserve/features/streaming)
232
+ ✅ [All ML frameworks: PyTorch, Jax, Tensorflow, Hugging Face...](https://lightning.ai/docs/litserve/features/full-control)
233
+ ✅ [Open AI spec](https://lightning.ai/docs/litserve/features/open-ai-spec)
234
+
235
+ [10+ features...](https://lightning.ai/docs/litserve/features)
236
+
237
+ **Note:** Our goal is not to jump on every hype train, but instead support features that scale
238
+ under the most demanding enterprise deployments.
239
+
240
+ &nbsp;
241
+
225
242
  # Community
226
243
  LitServe is a [community project accepting contributions](https://lightning.ai/docs/litserve/community) - Let's make the world's most advanced AI inference engine.
227
244
 
@@ -11,7 +11,7 @@
11
11
  # WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
12
12
  # See the License for the specific language governing permissions and
13
13
  # limitations under the License.
14
- __version__ = "0.2.0"
14
+ __version__ = "0.2.0.dev0"
15
15
  __author__ = "Lightning-AI et al."
16
16
  __author_email__ = "community@lightning.ai"
17
17
  __license__ = "Apache-2.0"
@@ -147,7 +147,7 @@ class LitAPI(ABC):
147
147
  def device(self, value):
148
148
  self._device = value
149
149
 
150
- def _sanitize(self, max_batch_size: int, spec: LitSpec):
150
+ def sanitize(self, max_batch_size: int, spec: LitSpec):
151
151
  if self.stream:
152
152
  self._default_unbatch = self._unbatch_stream
153
153
  else:
@@ -45,14 +45,6 @@ class TestAPIWithToolCalls(TestAPI):
45
45
  )
46
46
 
47
47
 
48
- class TestAPIWithStructuredOutput(TestAPI):
49
- def encode_response(self, output):
50
- yield ChatMessage(
51
- role="assistant",
52
- content='{"name": "Science Fair", "date": "Friday", "participants": ["Alice", "Bob"]}',
53
- )
54
-
55
-
56
48
  class OpenAIBatchContext(ls.LitAPI):
57
49
  def setup(self, device: str) -> None:
58
50
  self.model = None
@@ -449,7 +449,7 @@ class LitServer:
449
449
  self.api_path = api_path
450
450
  lit_api.stream = stream
451
451
  lit_api.request_timeout = timeout
452
- lit_api._sanitize(max_batch_size, spec=spec)
452
+ lit_api.sanitize(max_batch_size, spec=spec)
453
453
  self.app = FastAPI(lifespan=self.lifespan)
454
454
  self.app.response_queue_id = None
455
455
  self.response_queue_id = None
@@ -463,6 +463,7 @@ class LitServer:
463
463
  self.lit_spec = spec
464
464
  self.workers_per_device = workers_per_device
465
465
  self.max_batch_size = max_batch_size
466
+ self.timeout = timeout
466
467
  self.batch_timeout = batch_timeout
467
468
  self.stream = stream
468
469
  self._connector = _Connector(accelerator=accelerator, devices=devices)
@@ -20,7 +20,7 @@ import typing
20
20
  import uuid
21
21
  from collections import deque
22
22
  from enum import Enum
23
- from typing import Annotated, AsyncGenerator, Dict, Iterator, List, Literal, Optional, Union
23
+ from typing import AsyncGenerator, Dict, Iterator, List, Literal, Optional, Union
24
24
 
25
25
  from fastapi import BackgroundTasks, HTTPException, Request, Response
26
26
  from fastapi.responses import StreamingResponse
@@ -105,31 +105,6 @@ class ToolCall(BaseModel):
105
105
  function: FunctionCall
106
106
 
107
107
 
108
- class ResponseFormatText(BaseModel):
109
- type: Literal["text"]
110
-
111
-
112
- class ResponseFormatJSONObject(BaseModel):
113
- type: Literal["json_object"]
114
-
115
-
116
- class JSONSchema(BaseModel):
117
- name: str
118
- description: Optional[str] = None
119
- schema_def: Optional[Dict[str, object]] = Field(None, alias="schema")
120
- strict: Optional[bool] = False
121
-
122
-
123
- class ResponseFormatJSONSchema(BaseModel):
124
- json_schema: JSONSchema
125
- type: Literal["json_schema"]
126
-
127
-
128
- ResponseFormat = Annotated[
129
- Union[ResponseFormatText, ResponseFormatJSONObject, ResponseFormatJSONSchema], "ResponseFormat"
130
- ]
131
-
132
-
133
108
  class ChatMessage(BaseModel):
134
109
  role: str
135
110
  content: Union[str, List[Union[TextContent, ImageContent]]]
@@ -163,7 +138,6 @@ class ChatCompletionRequest(BaseModel):
163
138
  user: Optional[str] = None
164
139
  tools: Optional[List[Tool]] = None
165
140
  tool_choice: Optional[ToolChoice] = ToolChoice.auto
166
- response_format: Optional[ResponseFormat] = None
167
141
 
168
142
 
169
143
  class ChatCompletionResponseChoice(BaseModel):
@@ -14,10 +14,12 @@
14
14
  import asyncio
15
15
  import logging
16
16
  import pickle
17
- from typing import Optional
17
+ import uuid
18
+ from typing import Coroutine, Optional
18
19
  from contextlib import contextmanager
19
20
  from typing import TYPE_CHECKING
20
21
 
22
+
21
23
  from fastapi import HTTPException
22
24
  from starlette.middleware.base import BaseHTTPMiddleware
23
25
 
@@ -35,6 +37,24 @@ class LitAPIStatus:
35
37
  FINISH_STREAMING = "FINISH_STREAMING"
36
38
 
37
39
 
40
+ async def wait_for_queue_timeout(coro: Coroutine, timeout: Optional[float], uid: uuid.UUID, request_buffer: dict):
41
+ if timeout == -1 or timeout is False:
42
+ return await coro
43
+
44
+ task = asyncio.create_task(coro)
45
+ shield = asyncio.shield(task)
46
+ try:
47
+ return await asyncio.wait_for(shield, timeout)
48
+ except asyncio.TimeoutError:
49
+ if uid in request_buffer:
50
+ logger.error(
51
+ f"Request was waiting in the queue for too long ({timeout} seconds) and has been timed out. "
52
+ "You can adjust the timeout by providing the `timeout` argument to LitServe(..., timeout=30)."
53
+ )
54
+ raise HTTPException(504, "Request timed out")
55
+ return await task
56
+
57
+
38
58
  def load_and_raise(response):
39
59
  try:
40
60
  exception = pickle.loads(response) if isinstance(response, bytes) else response
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.1
2
2
  Name: litserve
3
- Version: 0.2.0
3
+ Version: 0.2.0.dev0
4
4
  Summary: Lightweight AI server.
5
5
  Home-page: https://github.com/Lightning-AI/litserve
6
6
  Download-URL: https://github.com/Lightning-AI/litserve
@@ -37,7 +37,6 @@ Requires-Dist: lightning>2.0.0; extra == "test"
37
37
  Requires-Dist: mypy==1.11.1; extra == "test"
38
38
  Requires-Dist: numpy<2.0; extra == "test"
39
39
  Requires-Dist: openai>=1.12.0; extra == "test"
40
- Requires-Dist: pillow; extra == "test"
41
40
  Requires-Dist: psutil; extra == "test"
42
41
  Requires-Dist: pytest-asyncio; extra == "test"
43
42
  Requires-Dist: pytest-cov; extra == "test"
@@ -49,28 +48,26 @@ Requires-Dist: transformers; extra == "test"
49
48
 
50
49
  <div align='center'>
51
50
 
52
- # LitServe: Easily serve AI models Lightning fast ⚡
51
+ # LitServe: Deploy AI models Lightning fast ⚡
53
52
 
54
53
  <img alt="Lightning" src="https://pl-bolts-doc-images.s3.us-east-2.amazonaws.com/app-2/ls_banner2.png" width="800px" style="max-width: 100%;">
55
54
 
56
55
  &nbsp;
57
56
 
58
- <strong>Flexible, high-throughput serving engine for AI models.</strong>
57
+ <strong>High-throughput serving engine for AI models.</strong>
59
58
  Friendly interface. Enterprise scale.
60
59
  </div>
61
60
 
62
61
  ----
63
62
 
64
- **LitServe** is a flexible serving engine for AI models built on FastAPI. Features like batching, streaming, and GPU autoscaling eliminate the need to rebuild a FastAPI server per model.
65
-
66
- LitServe is at least [2x faster](#performance) than plain FastAPI.
63
+ **LitServe** is an engine for scalable AI model deployment built on FastAPI. Features like batching, streaming, and GPU autoscaling eliminate the need to rebuild a FastAPI server for each model.
67
64
 
68
65
  <div align='center'>
69
66
 
70
67
  <pre>
71
- ✅ (2x)+ faster serving ✅ Self-host or fully managed ✅ Auto-GPU, multi-GPU
72
- ✅ Multi-modal ✅ PyTorch/JAX/TF ✅ Full control
73
- ✅ Batching ✅ Built on Fast API ✅ Streaming
68
+ ✅ Batching ✅ Streaming ✅ Auto-GPU, multi-GPU
69
+ ✅ Multi-modal ✅ PyTorch/JAX/TF ✅ Full control
70
+ ✅ Auth ✅ Built on Fast API ✅ Custom specs (Open AI)
74
71
  </pre>
75
72
 
76
73
  <div align='center'>
@@ -84,10 +81,11 @@ LitServe is at least [2x faster](#performance) than plain FastAPI.
84
81
  <div align="center">
85
82
  <div style="text-align: center;">
86
83
  <a href="#quick-start" style="margin: 0 10px;">Quick start</a> •
84
+ <a href="https://lightning.ai/" style="margin: 0 10px;">Lightning AI</a> •
87
85
  <a href="#featured-examples" style="margin: 0 10px;">Examples</a> •
86
+ <a href="#deployment-options" style="margin: 0 10px;">Deploy</a> •
88
87
  <a href="#features" style="margin: 0 10px;">Features</a> •
89
- <a href="#performance" style="margin: 0 10px;">Performance</a> •
90
- <a href="#hosting-options" style="margin: 0 10px;">Hosting</a> •
88
+ <a href="#performance" style="margin: 0 10px;">Benchmarks</a> •
91
89
  <a href="https://lightning.ai/docs/litserve" style="margin: 0 10px;">Docs</a>
92
90
  </div>
93
91
  </div>
@@ -102,16 +100,66 @@ LitServe is at least [2x faster](#performance) than plain FastAPI.
102
100
 
103
101
  &nbsp;
104
102
 
103
+ ## Performance
104
+ Our benchmarks show that LitServe (built on FastAPI) handles more simultaneous requests than FastAPI and TorchServe (higher is better).
105
+
106
+ Reproduce the full benchmarks [here](https://lightning.ai/docs/litserve/home/benchmarks).
107
+
108
+ <div align="center">
109
+ <img alt="LitServe" src="https://pl-bolts-doc-images.s3.us-east-2.amazonaws.com/app-2/ls_charts_v6.png" width="1000px" style="max-width: 100%;">
110
+ </div>
111
+
112
+ These results are for image and text classification ML tasks. The performance relationships hold for other ML tasks (embedding, LLM serving, audio, segmentation, object detection, summarization etc...).
113
+
114
+ ***💡 Note on LLM serving:*** For high-performance LLM serving (like Ollama/VLLM), use [LitGPT](https://github.com/Lightning-AI/litgpt?tab=readme-ov-file#deploy-an-llm) or build your custom VLLM-like server with LitServe. Optimizations like kv-caching, which can be done with LitServe, are needed to maximize LLM performance.
115
+
116
+ &nbsp;
117
+
118
+ ## Featured examples
119
+
120
+ Use LitServe to deploy any type of model or AI service (embeddings, LLMs, vision, audio, multi-modal, etc).
121
+
122
+ <table>
123
+ <tr>
124
+ <td style="vertical-align: top;">
125
+ <pre>
126
+ <strong>Featured examples</strong><br>
127
+ <strong>Toy model:</strong> <a href="#define-a-server">Hello world</a>
128
+ <strong>LLMs:</strong> <a href="https://lightning.ai/lightning-ai/studios/openai-fault-tolerant-proxy-server">Llama 3 (8B)</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-voice-clone-api-coqui-xtts-v2-model">LLM Proxy server</a>
129
+ <strong>NLP:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-any-hugging-face-model-instantly">Hugging face</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-hugging-face-bert-model">BERT</a>
130
+ <strong>Multimodal:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-clip-with-litserve">OpenAI Clip</a>
131
+ <strong>Audio:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-s-whisper-model">Whisper</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-music-generation-api-with-meta-s-audio-craft">AudioCraft</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-audio-generation-api">StableAudio</a>
132
+ <strong>Vision:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-private-api-for-stable-diffusion-2">Stable diffusion 2</a>
133
+ <strong>Speech:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-voice-clone-api-coqui-xtts-v2-model">Text-speech (XTTS V2)</a>
134
+ </pre>
135
+ </td>
136
+ <td style="vertical-align: top;">
137
+ <pre>
138
+ <strong>Key features</strong><br>
139
+ ✅ <strong>Serve all models:</strong> LLMs, vision, etc
140
+ ✅ <strong>All frameworks: </strong> PyTorch/Jax/sklearn/..
141
+ ✅ <strong>Dev friendly: </strong> build AI, not infra
142
+ ✅ <strong>Easy interface: </strong> no abstractions
143
+ ✅ <strong>Enterprise scale:</strong> scale huge models
144
+ ✅ <strong>Auto GPU scaling:</strong> zero code changes
145
+ ✅ <strong>Self host: </strong> or run on Studios
146
+ </pre>
147
+ </td>
148
+ </tr>
149
+ </table>
150
+
151
+ &nbsp;
152
+
105
153
  # Quick start
106
154
 
107
- Install LitServe via pip ([other install options](https://lightning.ai/docs/litserve/home/install)):
155
+ Install LitServe via pip (or [advanced installs](https://lightning.ai/docs/litserve/home/install)):
108
156
 
109
157
  ```bash
110
158
  pip install litserve
111
159
  ```
112
160
 
113
161
  ### Define a server
114
- Here's a hello world example ([explore real examples](#featured-examples)):
162
+ Here's a hello world example ([explore real examples](https://lightning.ai/docs/litserve/examples)):
115
163
 
116
164
  ```python
117
165
  # server.py
@@ -153,14 +201,24 @@ python server.py
153
201
 
154
202
  ### Query the server
155
203
 
156
- Use the automatically generated LitServe client:
204
+ Use the automatically generated LitServe client or write your own:
157
205
 
206
+ <table>
207
+ <tr>
208
+ <td style="vertical-align: top;">
209
+ <pre>
210
+ <strong>Option A - Use generated client: </strong><br>
211
+
158
212
  ```bash
159
213
  python client.py
160
214
  ```
215
+ <br>
161
216
 
162
- <details>
163
- <summary>Write a custom client</summary>
217
+ </pre>
218
+ </td>
219
+ <td style="vertical-align: top;">
220
+ <pre>
221
+ <strong>Option B - Custom client example: </strong><br>
164
222
 
165
223
  ```python
166
224
  import requests
@@ -169,79 +227,18 @@ response = requests.post(
169
227
  json={"input": 4.0}
170
228
  )
171
229
  ```
172
- </details>
173
-
174
- &nbsp;
175
-
176
-
177
- # Featured examples
178
- Use LitServe to deploy any model or AI service: (Gen AI, classical ML, embedding servers, LLMs, vision, audio, multi-modal systems, etc...)
179
-
180
- <div align='center'>
181
- <div width='200px'>
182
- <video src="https://github.com/user-attachments/assets/56655727-f5d7-4109-b60d-efc816e148c9" width='200px' controls></video>
183
- </div>
184
- </div>
185
-
186
- <pre>
187
- <strong>Featured examples</strong><br>
188
- <strong>Toy model:</strong> <a href="#define-a-server">Hello world</a>
189
- <strong>LLMs:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-private-llama-3-8b-api">Llama 3 (8B)</a>, <a href="https://lightning.ai/lightning-ai/studios/openai-fault-tolerant-proxy-server">LLM Proxy server</a>
190
- <strong>NLP:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-any-hugging-face-model-instantly">Hugging face</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-hugging-face-bert-model">BERT</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-text-embedding-api-with-litserve">Text embedding API</a>
191
- <strong>Multimodal:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-clip-with-litserve">OpenAI Clip</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-multi-modal-llm-with-minicpm">MiniCPM</a>, <a href="https://lightning.ai/lightning-ai/studios/run-meta-s-chameleon-30b">Chameleon 30B</a>
192
- <strong>Audio:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-open-ai-s-whisper-model">Whisper</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-music-generation-api-with-meta-s-audio-craft">AudioCraft</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-audio-generation-api">StableAudio</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-noise-cancellation-api-with-deepfilternet">Noise cancellation (DeepFilterNet)</a>
193
- <strong>Vision:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-private-api-for-stable-diffusion-2">Stable diffusion 2</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-image-generation-api-with-auraflow">AuroraFlow</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-an-image-generation-api-with-flux">Flux</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-a-super-resolution-image-api-with-aura-sr">Image super resolution (Aura SR)</a>
194
- <strong>Speech:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-a-voice-clone-api-coqui-xtts-v2-model">Text-speech (XTTS V2)</a>
195
- <strong>Classical ML:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-random-forest-with-litserve">Random forest</a>, <a href="https://lightning.ai/lightning-ai/studios/deploy-xgboost-with-litserve">XGBoost</a>
196
- <strong>Miscellaneous:</strong> <a href="https://lightning.ai/lightning-ai/studios/deploy-an-media-conversion-api-with-ffmpeg">Media conversion API (ffmpeg)</a>
230
+ <br>
197
231
  </pre>
198
-
199
- [Browse 100s of community-built templates](https://lightning.ai/studios?section=serving).
232
+ </td>
233
+ </tr>
234
+ </table>
200
235
 
201
236
  &nbsp;
202
237
 
203
- # Features
204
- LitServe supports multiple advanced state-of-the-art features.
238
+ # Deployment options
239
+ Self-manage LitServe deployments (just run it on any machine!), or deploy with one click on [Lightning AI](https://lightning.ai/).
205
240
 
206
- ✅ [(2x)+ faster serving than plain FastAPI](#performance)
207
- ✅ [Self host on your own machines](https://lightning.ai/docs/litserve/features/hosting-methods#host-on-your-own)
208
- ✅ [Host fully managed on Lightning AI](https://lightning.ai/docs/litserve/features/hosting-methods#host-on-lightning-studios)
209
- ✅ [Serve all models: LLMs, vision, time series, etc...](https://lightning.ai/docs/litserve/examples)
210
- ✅ [Auto-GPU scaling](https://lightning.ai/docs/litserve/features/gpu-inference)
211
- ✅ [Authentication](https://lightning.ai/docs/litserve/features/authentication)
212
- ✅ [Autoscaling](https://lightning.ai/docs/litserve/features/autoscaling)
213
- ✅ [Batching](https://lightning.ai/docs/litserve/features/batching)
214
- ✅ [Streaming](https://lightning.ai/docs/litserve/features/streaming)
215
- ✅ [Scale to zero (serverless)](https://lightning.ai/docs/litserve/features/streaming)
216
- ✅ [All ML frameworks: PyTorch, Jax, Tensorflow, Hugging Face...](https://lightning.ai/docs/litserve/features/full-control)
217
- ✅ [Open AI compatibility](https://lightning.ai/docs/litserve/features/open-ai-spec)
218
-
219
- [10+ features...](https://lightning.ai/docs/litserve/features)
220
-
221
- **Note:** Our goal is not to jump on every hype train, but instead support features that scale
222
- under the most demanding enterprise deployments.
223
-
224
- &nbsp;
225
-
226
- # Performance
227
- LitServe is highly optimized for parallel execution with native features optimized to scale AI workloads. Our benchmarks show that LitServe (built on FastAPI) handles more simultaneous requests than FastAPI and TorchServe (higher is better).
228
-
229
- Reproduce the full benchmarks [here](https://lightning.ai/docs/litserve/home/benchmarks).
230
-
231
- <div align="center">
232
- <img alt="LitServe" src="https://pl-bolts-doc-images.s3.us-east-2.amazonaws.com/app-2/ls_charts_v6.png" width="1000px" style="max-width: 100%;">
233
- </div>
234
-
235
- These results are for image and text classification ML tasks. The performance relationships hold for other ML tasks (embedding, LLM serving, audio, segmentation, object detection, summarization etc...).
236
-
237
- ***💡 Note on LLM serving:*** For high-performance LLM serving (like Ollama/VLLM), use [LitGPT](https://github.com/Lightning-AI/litgpt?tab=readme-ov-file#deploy-an-llm) or build your custom VLLM-like server with LitServe. Optimizations like kv-caching, which can be done with LitServe, are needed to maximize LLM performance.
238
-
239
- &nbsp;
240
-
241
- # Hosting options
242
- LitServe can be hosted independently on your own machines or fully managed via Lightning Studios.
243
-
244
- Self-hosting is ideal for hackers, students, and DIY developers, while fully managed hosting is ideal for enterprise developers needing easy autoscaling, security, release management, and 99.995% uptime and observability.
241
+ LitServe is developed by [Lightning AI](https://lightning.ai/) which provides infrastructure for deploying AI models.
245
242
 
246
243
  &nbsp;
247
244
 
@@ -271,6 +268,25 @@ Self-hosting is ideal for hackers, students, and DIY developers, while fully man
271
268
 
272
269
  &nbsp;
273
270
 
271
+ # Features
272
+ LitServe supports multiple advanced state-of-the-art features.
273
+
274
+ ✅ [All model types: LLMs, vision, time series, etc...](https://lightning.ai/docs/litserve/examples)
275
+ ✅ [Auto-GPU scaling](https://lightning.ai/docs/litserve/features/gpu-inference)
276
+ ✅ [Authentication](https://lightning.ai/docs/litserve/features/authentication)
277
+ ✅ [Autoscaling](https://lightning.ai/docs/litserve/features/autoscaling)
278
+ ✅ [Batching](https://lightning.ai/docs/litserve/features/batching)
279
+ ✅ [Streaming](https://lightning.ai/docs/litserve/features/streaming)
280
+ ✅ [All ML frameworks: PyTorch, Jax, Tensorflow, Hugging Face...](https://lightning.ai/docs/litserve/features/full-control)
281
+ ✅ [Open AI spec](https://lightning.ai/docs/litserve/features/open-ai-spec)
282
+
283
+ [10+ features...](https://lightning.ai/docs/litserve/features)
284
+
285
+ **Note:** Our goal is not to jump on every hype train, but instead support features that scale
286
+ under the most demanding enterprise deployments.
287
+
288
+ &nbsp;
289
+
274
290
  # Community
275
291
  LitServe is a [community project accepting contributions](https://lightning.ai/docs/litserve/community) - Let's make the world's most advanced AI inference engine.
276
292
 
@@ -9,7 +9,6 @@ lightning>2.0.0
9
9
  mypy==1.11.1
10
10
  numpy<2.0
11
11
  openai>=1.12.0
12
- pillow
13
12
  psutil
14
13
  pytest-asyncio
15
14
  pytest-cov
File without changes
File without changes
File without changes
File without changes
File without changes