asr-concurrent-stream 1.0.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- asr_concurrent_stream-1.0.0/PKG-INFO +168 -0
- asr_concurrent_stream-1.0.0/README.md +147 -0
- asr_concurrent_stream-1.0.0/asr_concurrent_stream.egg-info/PKG-INFO +168 -0
- asr_concurrent_stream-1.0.0/asr_concurrent_stream.egg-info/SOURCES.txt +21 -0
- asr_concurrent_stream-1.0.0/asr_concurrent_stream.egg-info/dependency_links.txt +1 -0
- asr_concurrent_stream-1.0.0/asr_concurrent_stream.egg-info/entry_points.txt +3 -0
- asr_concurrent_stream-1.0.0/asr_concurrent_stream.egg-info/requires.txt +12 -0
- asr_concurrent_stream-1.0.0/asr_concurrent_stream.egg-info/top_level.txt +3 -0
- asr_concurrent_stream-1.0.0/client/__init__.py +1 -0
- asr_concurrent_stream-1.0.0/client/example_client.py +409 -0
- asr_concurrent_stream-1.0.0/proto/__init__.py +46 -0
- asr_concurrent_stream-1.0.0/proto/asr_pb2.py +60 -0
- asr_concurrent_stream-1.0.0/proto/asr_pb2_grpc.py +269 -0
- asr_concurrent_stream-1.0.0/pyproject.toml +40 -0
- asr_concurrent_stream-1.0.0/server/__init__.py +1 -0
- asr_concurrent_stream-1.0.0/server/_launch_server.py +73 -0
- asr_concurrent_stream-1.0.0/server/asr_grpc_server.py +459 -0
- asr_concurrent_stream-1.0.0/server/asr_model.py +213 -0
- asr_concurrent_stream-1.0.0/server/asr_processor.py +39 -0
- asr_concurrent_stream-1.0.0/server/asr_utils.py +133 -0
- asr_concurrent_stream-1.0.0/server/inference_coordinator.py +480 -0
- asr_concurrent_stream-1.0.0/server/stream_manager.py +337 -0
- asr_concurrent_stream-1.0.0/setup.cfg +4 -0
|
@@ -0,0 +1,168 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: asr-concurrent-stream
|
|
3
|
+
Version: 1.0.0
|
|
4
|
+
Summary: Concurrent streaming gRPC server and client for Qwen3-ASR
|
|
5
|
+
Author: Alibaba Qwen Team
|
|
6
|
+
License: Apache-2.0
|
|
7
|
+
Project-URL: Homepage, https://github.com/Qwen/Qwen3-ASR
|
|
8
|
+
Keywords: asr,speech-recognition,grpc,streaming,qwen3
|
|
9
|
+
Requires-Python: >=3.9
|
|
10
|
+
Description-Content-Type: text/markdown
|
|
11
|
+
Requires-Dist: grpcio
|
|
12
|
+
Requires-Dist: numpy
|
|
13
|
+
Requires-Dist: protobuf
|
|
14
|
+
Requires-Dist: vllm
|
|
15
|
+
Requires-Dist: transformers
|
|
16
|
+
Provides-Extra: client
|
|
17
|
+
Requires-Dist: soundfile; extra == "client"
|
|
18
|
+
Requires-Dist: torch; extra == "client"
|
|
19
|
+
Provides-Extra: dev
|
|
20
|
+
Requires-Dist: grpcio-tools; extra == "dev"
|
|
21
|
+
|
|
22
|
+
# ASR Concurrent Stream Server
|
|
23
|
+
|
|
24
|
+
Concurrent streaming gRPC server and client for Qwen3-ASR.
|
|
25
|
+
|
|
26
|
+
## Installation
|
|
27
|
+
|
|
28
|
+
```bash
|
|
29
|
+
pip install asr-concurrent-stream
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
For client-side audio loading and VAD segmentation:
|
|
33
|
+
|
|
34
|
+
```bash
|
|
35
|
+
pip install asr-concurrent-stream[client]
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
For development (regenerating protobuf code):
|
|
39
|
+
|
|
40
|
+
```bash
|
|
41
|
+
pip install asr-concurrent-stream[dev]
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
**Note:** The server requires `vllm` and `transformers` to run the Qwen3-ASR model. These are installed automatically with the package.
|
|
45
|
+
|
|
46
|
+
## Quick Start
|
|
47
|
+
|
|
48
|
+
### Start the server
|
|
49
|
+
|
|
50
|
+
```bash
|
|
51
|
+
asr-concurrent-server
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
The server reads configuration from environment variables:
|
|
55
|
+
|
|
56
|
+
| Variable | Default | Description |
|
|
57
|
+
|----------|---------|-------------|
|
|
58
|
+
| `ASR_PORT` | `8000` | gRPC server port |
|
|
59
|
+
| `ASR_MODEL_PATH` | `Qwen/Qwen3-ASR-1.7B` | Model path or HuggingFace ID |
|
|
60
|
+
| `ASR_GPU_MEMORY_UTILIZATION` | `0.30` | GPU memory fraction |
|
|
61
|
+
| `ASR_MAX_MODEL_LEN` | `4096` | Maximum model sequence length |
|
|
62
|
+
| `ASR_MAX_CONCURRENT_STREAMS` | `15` | Maximum concurrent streams |
|
|
63
|
+
| `ASR_MAX_BATCH_SIZE` | `8` | Maximum batch size |
|
|
64
|
+
| `ASR_BATCH_TIMEOUT_MS` | `50` | Batch timeout in milliseconds |
|
|
65
|
+
| `ASR_WORKER_THREADS` | `4` | Number of worker threads |
|
|
66
|
+
| `ASR_HEALTH_PORT` | `8080` | HTTP health check port |
|
|
67
|
+
|
|
68
|
+
### Run the client
|
|
69
|
+
|
|
70
|
+
```bash
|
|
71
|
+
asr-concurrent-client --audio path/to/audio.wav --language Italian
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
## Dependencies
|
|
75
|
+
|
|
76
|
+
| Package | Purpose |
|
|
77
|
+
|---------|---------|
|
|
78
|
+
| `grpcio` | gRPC framework |
|
|
79
|
+
| `numpy` | Audio array handling |
|
|
80
|
+
| `protobuf` | Protocol buffer serialization |
|
|
81
|
+
| `vllm` | vLLM inference engine |
|
|
82
|
+
| `transformers` | Model processor and tokenizer |
|
|
83
|
+
|
|
84
|
+
Optional client dependencies: `soundfile`, `torch` (for Silero VAD segmentation).
|
|
85
|
+
|
|
86
|
+
---
|
|
87
|
+
|
|
88
|
+
*Docker image for the Qwen3-ASR concurrent streaming gRPC server.*
|
|
89
|
+
|
|
90
|
+
## Image Composition
|
|
91
|
+
|
|
92
|
+
| Layer | Detail |
|
|
93
|
+
|-------|--------|
|
|
94
|
+
| **Base OS** | Ubuntu 24.04 LTS |
|
|
95
|
+
| **CUDA** | NVIDIA CUDA 12.8.0 (`nvidia/cuda:12.8.0-runtime-ubuntu24.04`) |
|
|
96
|
+
| **Python** | 3.12 (deadsnakes PPA), installed into `/opt/venv` |
|
|
97
|
+
| **Model** | `Qwen/Qwen3-ASR-1.7B` baked into the image |
|
|
98
|
+
| **Server port** | Container `8002` (map to any host port at runtime) |
|
|
99
|
+
|
|
100
|
+
## Installation Sequence
|
|
101
|
+
|
|
102
|
+
1. **System dependencies** — `curl`, `git`, `wget`, `build-essential`, `ffmpeg`, `libsox-*`, `libsndfile1`, `libopus0`, `libffi-dev`.
|
|
103
|
+
2. **Python 3.12** — installed via deadsnakes PPA, venv created at `/opt/venv`, binaries symlinked to `/usr/local/bin`.
|
|
104
|
+
3. **pip** — upgraded to latest `pip`, `setuptools`, `wheel`.
|
|
105
|
+
4. **`qwen-asr[vllm]`** — installs the full Qwen3-ASR stack with vLLM extras (transformers, vllm, torch, accelerate, librosa, soundfile, etc. with pinned versions).
|
|
106
|
+
5. **flash-attention** — prebuilt wheel for CUDA 12.8 + PyTorch 2.9 + Python 3.12.
|
|
107
|
+
6. **Project files** — `ASR_Concurrent_Stream/` copied to `/app/ASR_Concurrent_Stream/`.
|
|
108
|
+
7. **Model download** — `Qwen/Qwen3-ASR-1.7B` downloaded from HuggingFace to `/app/models/Qwen3-ASR-1.7B`.
|
|
109
|
+
8. **Launcher patch** — `run_asr_grpc_server.sh` is updated:
|
|
110
|
+
- `PORT` → `8002`
|
|
111
|
+
- `MODEL_PATH` → `/app/models/Qwen3-ASR-1.7B`
|
|
112
|
+
- `QWEN3_ASR_PATH` → site-packages `qwen_asr` path
|
|
113
|
+
|
|
114
|
+
## Build & Run
|
|
115
|
+
|
|
116
|
+
```bash
|
|
117
|
+
# Build
|
|
118
|
+
docker build -t qwen3-asr-server:latest .
|
|
119
|
+
|
|
120
|
+
# Run (map host port 8002 to container port 8002)
|
|
121
|
+
docker run -d \
|
|
122
|
+
--name qwen3-asr-server \
|
|
123
|
+
--gpus all \
|
|
124
|
+
-p 8002:8002 \
|
|
125
|
+
--shm-size=16g \
|
|
126
|
+
--ulimit memlock=-1 \
|
|
127
|
+
--restart unless-stopped \
|
|
128
|
+
qwen3-asr-server:latest
|
|
129
|
+
|
|
130
|
+
# View logs
|
|
131
|
+
docker logs -f qwen3-asr-server
|
|
132
|
+
|
|
133
|
+
# Stop & remove
|
|
134
|
+
docker stop qwen3-asr-server && docker rm qwen3-asr-server
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
## Server Configuration
|
|
138
|
+
|
|
139
|
+
Hard-coded defaults in `server/run_asr_grpc_server.sh`:
|
|
140
|
+
|
|
141
|
+
| Parameter | Value |
|
|
142
|
+
|-----------|-------|
|
|
143
|
+
| Port | `8002` |
|
|
144
|
+
| Model | `/app/models/Qwen3-ASR-1.7B` |
|
|
145
|
+
| GPU memory utilization | `0.35` |
|
|
146
|
+
| Max model length | `4096` |
|
|
147
|
+
| Max concurrent streams | `10` |
|
|
148
|
+
| Max batch size | `8` |
|
|
149
|
+
| Batch timeout | `50` ms |
|
|
150
|
+
| Worker threads | `4` |
|
|
151
|
+
|
|
152
|
+
## Project Structure (in image)
|
|
153
|
+
|
|
154
|
+
```
|
|
155
|
+
/app/
|
|
156
|
+
├── ASR_Concurrent_Stream/
|
|
157
|
+
│ ├── proto/ # gRPC protobuf definitions
|
|
158
|
+
│ ├── server/
|
|
159
|
+
│ │ ├── asr_grpc_server.py
|
|
160
|
+
│ │ ├── _launch_server.py
|
|
161
|
+
│ │ ├── stream_manager.py
|
|
162
|
+
│ │ ├── inference_coordinator.py
|
|
163
|
+
│ │ └── run_asr_grpc_server.sh
|
|
164
|
+
│ └── client/
|
|
165
|
+
│ └── example_client.py
|
|
166
|
+
└── models/
|
|
167
|
+
└── Qwen3-ASR-1.7B/ # Pre-downloaded model
|
|
168
|
+
```
|
|
@@ -0,0 +1,147 @@
|
|
|
1
|
+
# ASR Concurrent Stream Server
|
|
2
|
+
|
|
3
|
+
Concurrent streaming gRPC server and client for Qwen3-ASR.
|
|
4
|
+
|
|
5
|
+
## Installation
|
|
6
|
+
|
|
7
|
+
```bash
|
|
8
|
+
pip install asr-concurrent-stream
|
|
9
|
+
```
|
|
10
|
+
|
|
11
|
+
For client-side audio loading and VAD segmentation:
|
|
12
|
+
|
|
13
|
+
```bash
|
|
14
|
+
pip install asr-concurrent-stream[client]
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
For development (regenerating protobuf code):
|
|
18
|
+
|
|
19
|
+
```bash
|
|
20
|
+
pip install asr-concurrent-stream[dev]
|
|
21
|
+
```
|
|
22
|
+
|
|
23
|
+
**Note:** The server requires `vllm` and `transformers` to run the Qwen3-ASR model. These are installed automatically with the package.
|
|
24
|
+
|
|
25
|
+
## Quick Start
|
|
26
|
+
|
|
27
|
+
### Start the server
|
|
28
|
+
|
|
29
|
+
```bash
|
|
30
|
+
asr-concurrent-server
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
The server reads configuration from environment variables:
|
|
34
|
+
|
|
35
|
+
| Variable | Default | Description |
|
|
36
|
+
|----------|---------|-------------|
|
|
37
|
+
| `ASR_PORT` | `8000` | gRPC server port |
|
|
38
|
+
| `ASR_MODEL_PATH` | `Qwen/Qwen3-ASR-1.7B` | Model path or HuggingFace ID |
|
|
39
|
+
| `ASR_GPU_MEMORY_UTILIZATION` | `0.30` | GPU memory fraction |
|
|
40
|
+
| `ASR_MAX_MODEL_LEN` | `4096` | Maximum model sequence length |
|
|
41
|
+
| `ASR_MAX_CONCURRENT_STREAMS` | `15` | Maximum concurrent streams |
|
|
42
|
+
| `ASR_MAX_BATCH_SIZE` | `8` | Maximum batch size |
|
|
43
|
+
| `ASR_BATCH_TIMEOUT_MS` | `50` | Batch timeout in milliseconds |
|
|
44
|
+
| `ASR_WORKER_THREADS` | `4` | Number of worker threads |
|
|
45
|
+
| `ASR_HEALTH_PORT` | `8080` | HTTP health check port |
|
|
46
|
+
|
|
47
|
+
### Run the client
|
|
48
|
+
|
|
49
|
+
```bash
|
|
50
|
+
asr-concurrent-client --audio path/to/audio.wav --language Italian
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
## Dependencies
|
|
54
|
+
|
|
55
|
+
| Package | Purpose |
|
|
56
|
+
|---------|---------|
|
|
57
|
+
| `grpcio` | gRPC framework |
|
|
58
|
+
| `numpy` | Audio array handling |
|
|
59
|
+
| `protobuf` | Protocol buffer serialization |
|
|
60
|
+
| `vllm` | vLLM inference engine |
|
|
61
|
+
| `transformers` | Model processor and tokenizer |
|
|
62
|
+
|
|
63
|
+
Optional client dependencies: `soundfile`, `torch` (for Silero VAD segmentation).
|
|
64
|
+
|
|
65
|
+
---
|
|
66
|
+
|
|
67
|
+
*Docker image for the Qwen3-ASR concurrent streaming gRPC server.*
|
|
68
|
+
|
|
69
|
+
## Image Composition
|
|
70
|
+
|
|
71
|
+
| Layer | Detail |
|
|
72
|
+
|-------|--------|
|
|
73
|
+
| **Base OS** | Ubuntu 24.04 LTS |
|
|
74
|
+
| **CUDA** | NVIDIA CUDA 12.8.0 (`nvidia/cuda:12.8.0-runtime-ubuntu24.04`) |
|
|
75
|
+
| **Python** | 3.12 (deadsnakes PPA), installed into `/opt/venv` |
|
|
76
|
+
| **Model** | `Qwen/Qwen3-ASR-1.7B` baked into the image |
|
|
77
|
+
| **Server port** | Container `8002` (map to any host port at runtime) |
|
|
78
|
+
|
|
79
|
+
## Installation Sequence
|
|
80
|
+
|
|
81
|
+
1. **System dependencies** — `curl`, `git`, `wget`, `build-essential`, `ffmpeg`, `libsox-*`, `libsndfile1`, `libopus0`, `libffi-dev`.
|
|
82
|
+
2. **Python 3.12** — installed via deadsnakes PPA, venv created at `/opt/venv`, binaries symlinked to `/usr/local/bin`.
|
|
83
|
+
3. **pip** — upgraded to latest `pip`, `setuptools`, `wheel`.
|
|
84
|
+
4. **`qwen-asr[vllm]`** — installs the full Qwen3-ASR stack with vLLM extras (transformers, vllm, torch, accelerate, librosa, soundfile, etc. with pinned versions).
|
|
85
|
+
5. **flash-attention** — prebuilt wheel for CUDA 12.8 + PyTorch 2.9 + Python 3.12.
|
|
86
|
+
6. **Project files** — `ASR_Concurrent_Stream/` copied to `/app/ASR_Concurrent_Stream/`.
|
|
87
|
+
7. **Model download** — `Qwen/Qwen3-ASR-1.7B` downloaded from HuggingFace to `/app/models/Qwen3-ASR-1.7B`.
|
|
88
|
+
8. **Launcher patch** — `run_asr_grpc_server.sh` is updated:
|
|
89
|
+
- `PORT` → `8002`
|
|
90
|
+
- `MODEL_PATH` → `/app/models/Qwen3-ASR-1.7B`
|
|
91
|
+
- `QWEN3_ASR_PATH` → site-packages `qwen_asr` path
|
|
92
|
+
|
|
93
|
+
## Build & Run
|
|
94
|
+
|
|
95
|
+
```bash
|
|
96
|
+
# Build
|
|
97
|
+
docker build -t qwen3-asr-server:latest .
|
|
98
|
+
|
|
99
|
+
# Run (map host port 8002 to container port 8002)
|
|
100
|
+
docker run -d \
|
|
101
|
+
--name qwen3-asr-server \
|
|
102
|
+
--gpus all \
|
|
103
|
+
-p 8002:8002 \
|
|
104
|
+
--shm-size=16g \
|
|
105
|
+
--ulimit memlock=-1 \
|
|
106
|
+
--restart unless-stopped \
|
|
107
|
+
qwen3-asr-server:latest
|
|
108
|
+
|
|
109
|
+
# View logs
|
|
110
|
+
docker logs -f qwen3-asr-server
|
|
111
|
+
|
|
112
|
+
# Stop & remove
|
|
113
|
+
docker stop qwen3-asr-server && docker rm qwen3-asr-server
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
## Server Configuration
|
|
117
|
+
|
|
118
|
+
Hard-coded defaults in `server/run_asr_grpc_server.sh`:
|
|
119
|
+
|
|
120
|
+
| Parameter | Value |
|
|
121
|
+
|-----------|-------|
|
|
122
|
+
| Port | `8002` |
|
|
123
|
+
| Model | `/app/models/Qwen3-ASR-1.7B` |
|
|
124
|
+
| GPU memory utilization | `0.35` |
|
|
125
|
+
| Max model length | `4096` |
|
|
126
|
+
| Max concurrent streams | `10` |
|
|
127
|
+
| Max batch size | `8` |
|
|
128
|
+
| Batch timeout | `50` ms |
|
|
129
|
+
| Worker threads | `4` |
|
|
130
|
+
|
|
131
|
+
## Project Structure (in image)
|
|
132
|
+
|
|
133
|
+
```
|
|
134
|
+
/app/
|
|
135
|
+
├── ASR_Concurrent_Stream/
|
|
136
|
+
│ ├── proto/ # gRPC protobuf definitions
|
|
137
|
+
│ ├── server/
|
|
138
|
+
│ │ ├── asr_grpc_server.py
|
|
139
|
+
│ │ ├── _launch_server.py
|
|
140
|
+
│ │ ├── stream_manager.py
|
|
141
|
+
│ │ ├── inference_coordinator.py
|
|
142
|
+
│ │ └── run_asr_grpc_server.sh
|
|
143
|
+
│ └── client/
|
|
144
|
+
│ └── example_client.py
|
|
145
|
+
└── models/
|
|
146
|
+
└── Qwen3-ASR-1.7B/ # Pre-downloaded model
|
|
147
|
+
```
|
|
@@ -0,0 +1,168 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: asr-concurrent-stream
|
|
3
|
+
Version: 1.0.0
|
|
4
|
+
Summary: Concurrent streaming gRPC server and client for Qwen3-ASR
|
|
5
|
+
Author: Alibaba Qwen Team
|
|
6
|
+
License: Apache-2.0
|
|
7
|
+
Project-URL: Homepage, https://github.com/Qwen/Qwen3-ASR
|
|
8
|
+
Keywords: asr,speech-recognition,grpc,streaming,qwen3
|
|
9
|
+
Requires-Python: >=3.9
|
|
10
|
+
Description-Content-Type: text/markdown
|
|
11
|
+
Requires-Dist: grpcio
|
|
12
|
+
Requires-Dist: numpy
|
|
13
|
+
Requires-Dist: protobuf
|
|
14
|
+
Requires-Dist: vllm
|
|
15
|
+
Requires-Dist: transformers
|
|
16
|
+
Provides-Extra: client
|
|
17
|
+
Requires-Dist: soundfile; extra == "client"
|
|
18
|
+
Requires-Dist: torch; extra == "client"
|
|
19
|
+
Provides-Extra: dev
|
|
20
|
+
Requires-Dist: grpcio-tools; extra == "dev"
|
|
21
|
+
|
|
22
|
+
# ASR Concurrent Stream Server
|
|
23
|
+
|
|
24
|
+
Concurrent streaming gRPC server and client for Qwen3-ASR.
|
|
25
|
+
|
|
26
|
+
## Installation
|
|
27
|
+
|
|
28
|
+
```bash
|
|
29
|
+
pip install asr-concurrent-stream
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
For client-side audio loading and VAD segmentation:
|
|
33
|
+
|
|
34
|
+
```bash
|
|
35
|
+
pip install asr-concurrent-stream[client]
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
For development (regenerating protobuf code):
|
|
39
|
+
|
|
40
|
+
```bash
|
|
41
|
+
pip install asr-concurrent-stream[dev]
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
**Note:** The server requires `vllm` and `transformers` to run the Qwen3-ASR model. These are installed automatically with the package.
|
|
45
|
+
|
|
46
|
+
## Quick Start
|
|
47
|
+
|
|
48
|
+
### Start the server
|
|
49
|
+
|
|
50
|
+
```bash
|
|
51
|
+
asr-concurrent-server
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
The server reads configuration from environment variables:
|
|
55
|
+
|
|
56
|
+
| Variable | Default | Description |
|
|
57
|
+
|----------|---------|-------------|
|
|
58
|
+
| `ASR_PORT` | `8000` | gRPC server port |
|
|
59
|
+
| `ASR_MODEL_PATH` | `Qwen/Qwen3-ASR-1.7B` | Model path or HuggingFace ID |
|
|
60
|
+
| `ASR_GPU_MEMORY_UTILIZATION` | `0.30` | GPU memory fraction |
|
|
61
|
+
| `ASR_MAX_MODEL_LEN` | `4096` | Maximum model sequence length |
|
|
62
|
+
| `ASR_MAX_CONCURRENT_STREAMS` | `15` | Maximum concurrent streams |
|
|
63
|
+
| `ASR_MAX_BATCH_SIZE` | `8` | Maximum batch size |
|
|
64
|
+
| `ASR_BATCH_TIMEOUT_MS` | `50` | Batch timeout in milliseconds |
|
|
65
|
+
| `ASR_WORKER_THREADS` | `4` | Number of worker threads |
|
|
66
|
+
| `ASR_HEALTH_PORT` | `8080` | HTTP health check port |
|
|
67
|
+
|
|
68
|
+
### Run the client
|
|
69
|
+
|
|
70
|
+
```bash
|
|
71
|
+
asr-concurrent-client --audio path/to/audio.wav --language Italian
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
## Dependencies
|
|
75
|
+
|
|
76
|
+
| Package | Purpose |
|
|
77
|
+
|---------|---------|
|
|
78
|
+
| `grpcio` | gRPC framework |
|
|
79
|
+
| `numpy` | Audio array handling |
|
|
80
|
+
| `protobuf` | Protocol buffer serialization |
|
|
81
|
+
| `vllm` | vLLM inference engine |
|
|
82
|
+
| `transformers` | Model processor and tokenizer |
|
|
83
|
+
|
|
84
|
+
Optional client dependencies: `soundfile`, `torch` (for Silero VAD segmentation).
|
|
85
|
+
|
|
86
|
+
---
|
|
87
|
+
|
|
88
|
+
*Docker image for the Qwen3-ASR concurrent streaming gRPC server.*
|
|
89
|
+
|
|
90
|
+
## Image Composition
|
|
91
|
+
|
|
92
|
+
| Layer | Detail |
|
|
93
|
+
|-------|--------|
|
|
94
|
+
| **Base OS** | Ubuntu 24.04 LTS |
|
|
95
|
+
| **CUDA** | NVIDIA CUDA 12.8.0 (`nvidia/cuda:12.8.0-runtime-ubuntu24.04`) |
|
|
96
|
+
| **Python** | 3.12 (deadsnakes PPA), installed into `/opt/venv` |
|
|
97
|
+
| **Model** | `Qwen/Qwen3-ASR-1.7B` baked into the image |
|
|
98
|
+
| **Server port** | Container `8002` (map to any host port at runtime) |
|
|
99
|
+
|
|
100
|
+
## Installation Sequence
|
|
101
|
+
|
|
102
|
+
1. **System dependencies** — `curl`, `git`, `wget`, `build-essential`, `ffmpeg`, `libsox-*`, `libsndfile1`, `libopus0`, `libffi-dev`.
|
|
103
|
+
2. **Python 3.12** — installed via deadsnakes PPA, venv created at `/opt/venv`, binaries symlinked to `/usr/local/bin`.
|
|
104
|
+
3. **pip** — upgraded to latest `pip`, `setuptools`, `wheel`.
|
|
105
|
+
4. **`qwen-asr[vllm]`** — installs the full Qwen3-ASR stack with vLLM extras (transformers, vllm, torch, accelerate, librosa, soundfile, etc. with pinned versions).
|
|
106
|
+
5. **flash-attention** — prebuilt wheel for CUDA 12.8 + PyTorch 2.9 + Python 3.12.
|
|
107
|
+
6. **Project files** — `ASR_Concurrent_Stream/` copied to `/app/ASR_Concurrent_Stream/`.
|
|
108
|
+
7. **Model download** — `Qwen/Qwen3-ASR-1.7B` downloaded from HuggingFace to `/app/models/Qwen3-ASR-1.7B`.
|
|
109
|
+
8. **Launcher patch** — `run_asr_grpc_server.sh` is updated:
|
|
110
|
+
- `PORT` → `8002`
|
|
111
|
+
- `MODEL_PATH` → `/app/models/Qwen3-ASR-1.7B`
|
|
112
|
+
- `QWEN3_ASR_PATH` → site-packages `qwen_asr` path
|
|
113
|
+
|
|
114
|
+
## Build & Run
|
|
115
|
+
|
|
116
|
+
```bash
|
|
117
|
+
# Build
|
|
118
|
+
docker build -t qwen3-asr-server:latest .
|
|
119
|
+
|
|
120
|
+
# Run (map host port 8002 to container port 8002)
|
|
121
|
+
docker run -d \
|
|
122
|
+
--name qwen3-asr-server \
|
|
123
|
+
--gpus all \
|
|
124
|
+
-p 8002:8002 \
|
|
125
|
+
--shm-size=16g \
|
|
126
|
+
--ulimit memlock=-1 \
|
|
127
|
+
--restart unless-stopped \
|
|
128
|
+
qwen3-asr-server:latest
|
|
129
|
+
|
|
130
|
+
# View logs
|
|
131
|
+
docker logs -f qwen3-asr-server
|
|
132
|
+
|
|
133
|
+
# Stop & remove
|
|
134
|
+
docker stop qwen3-asr-server && docker rm qwen3-asr-server
|
|
135
|
+
```
|
|
136
|
+
|
|
137
|
+
## Server Configuration
|
|
138
|
+
|
|
139
|
+
Hard-coded defaults in `server/run_asr_grpc_server.sh`:
|
|
140
|
+
|
|
141
|
+
| Parameter | Value |
|
|
142
|
+
|-----------|-------|
|
|
143
|
+
| Port | `8002` |
|
|
144
|
+
| Model | `/app/models/Qwen3-ASR-1.7B` |
|
|
145
|
+
| GPU memory utilization | `0.35` |
|
|
146
|
+
| Max model length | `4096` |
|
|
147
|
+
| Max concurrent streams | `10` |
|
|
148
|
+
| Max batch size | `8` |
|
|
149
|
+
| Batch timeout | `50` ms |
|
|
150
|
+
| Worker threads | `4` |
|
|
151
|
+
|
|
152
|
+
## Project Structure (in image)
|
|
153
|
+
|
|
154
|
+
```
|
|
155
|
+
/app/
|
|
156
|
+
├── ASR_Concurrent_Stream/
|
|
157
|
+
│ ├── proto/ # gRPC protobuf definitions
|
|
158
|
+
│ ├── server/
|
|
159
|
+
│ │ ├── asr_grpc_server.py
|
|
160
|
+
│ │ ├── _launch_server.py
|
|
161
|
+
│ │ ├── stream_manager.py
|
|
162
|
+
│ │ ├── inference_coordinator.py
|
|
163
|
+
│ │ └── run_asr_grpc_server.sh
|
|
164
|
+
│ └── client/
|
|
165
|
+
│ └── example_client.py
|
|
166
|
+
└── models/
|
|
167
|
+
└── Qwen3-ASR-1.7B/ # Pre-downloaded model
|
|
168
|
+
```
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
README.md
|
|
2
|
+
pyproject.toml
|
|
3
|
+
asr_concurrent_stream.egg-info/PKG-INFO
|
|
4
|
+
asr_concurrent_stream.egg-info/SOURCES.txt
|
|
5
|
+
asr_concurrent_stream.egg-info/dependency_links.txt
|
|
6
|
+
asr_concurrent_stream.egg-info/entry_points.txt
|
|
7
|
+
asr_concurrent_stream.egg-info/requires.txt
|
|
8
|
+
asr_concurrent_stream.egg-info/top_level.txt
|
|
9
|
+
client/__init__.py
|
|
10
|
+
client/example_client.py
|
|
11
|
+
proto/__init__.py
|
|
12
|
+
proto/asr_pb2.py
|
|
13
|
+
proto/asr_pb2_grpc.py
|
|
14
|
+
server/__init__.py
|
|
15
|
+
server/_launch_server.py
|
|
16
|
+
server/asr_grpc_server.py
|
|
17
|
+
server/asr_model.py
|
|
18
|
+
server/asr_processor.py
|
|
19
|
+
server/asr_utils.py
|
|
20
|
+
server/inference_coordinator.py
|
|
21
|
+
server/stream_manager.py
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
# ASR Concurrent Stream Client
|