asr-concurrent-stream 1.0.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,168 @@
1
+ Metadata-Version: 2.4
2
+ Name: asr-concurrent-stream
3
+ Version: 1.0.0
4
+ Summary: Concurrent streaming gRPC server and client for Qwen3-ASR
5
+ Author: Alibaba Qwen Team
6
+ License: Apache-2.0
7
+ Project-URL: Homepage, https://github.com/Qwen/Qwen3-ASR
8
+ Keywords: asr,speech-recognition,grpc,streaming,qwen3
9
+ Requires-Python: >=3.9
10
+ Description-Content-Type: text/markdown
11
+ Requires-Dist: grpcio
12
+ Requires-Dist: numpy
13
+ Requires-Dist: protobuf
14
+ Requires-Dist: vllm
15
+ Requires-Dist: transformers
16
+ Provides-Extra: client
17
+ Requires-Dist: soundfile; extra == "client"
18
+ Requires-Dist: torch; extra == "client"
19
+ Provides-Extra: dev
20
+ Requires-Dist: grpcio-tools; extra == "dev"
21
+
22
+ # ASR Concurrent Stream Server
23
+
24
+ Concurrent streaming gRPC server and client for Qwen3-ASR.
25
+
26
+ ## Installation
27
+
28
+ ```bash
29
+ pip install asr-concurrent-stream
30
+ ```
31
+
32
+ For client-side audio loading and VAD segmentation:
33
+
34
+ ```bash
35
+ pip install asr-concurrent-stream[client]
36
+ ```
37
+
38
+ For development (regenerating protobuf code):
39
+
40
+ ```bash
41
+ pip install asr-concurrent-stream[dev]
42
+ ```
43
+
44
+ **Note:** The server requires `vllm` and `transformers` to run the Qwen3-ASR model. These are installed automatically with the package.
45
+
46
+ ## Quick Start
47
+
48
+ ### Start the server
49
+
50
+ ```bash
51
+ asr-concurrent-server
52
+ ```
53
+
54
+ The server reads configuration from environment variables:
55
+
56
+ | Variable | Default | Description |
57
+ |----------|---------|-------------|
58
+ | `ASR_PORT` | `8000` | gRPC server port |
59
+ | `ASR_MODEL_PATH` | `Qwen/Qwen3-ASR-1.7B` | Model path or HuggingFace ID |
60
+ | `ASR_GPU_MEMORY_UTILIZATION` | `0.30` | GPU memory fraction |
61
+ | `ASR_MAX_MODEL_LEN` | `4096` | Maximum model sequence length |
62
+ | `ASR_MAX_CONCURRENT_STREAMS` | `15` | Maximum concurrent streams |
63
+ | `ASR_MAX_BATCH_SIZE` | `8` | Maximum batch size |
64
+ | `ASR_BATCH_TIMEOUT_MS` | `50` | Batch timeout in milliseconds |
65
+ | `ASR_WORKER_THREADS` | `4` | Number of worker threads |
66
+ | `ASR_HEALTH_PORT` | `8080` | HTTP health check port |
67
+
68
+ ### Run the client
69
+
70
+ ```bash
71
+ asr-concurrent-client --audio path/to/audio.wav --language Italian
72
+ ```
73
+
74
+ ## Dependencies
75
+
76
+ | Package | Purpose |
77
+ |---------|---------|
78
+ | `grpcio` | gRPC framework |
79
+ | `numpy` | Audio array handling |
80
+ | `protobuf` | Protocol buffer serialization |
81
+ | `vllm` | vLLM inference engine |
82
+ | `transformers` | Model processor and tokenizer |
83
+
84
+ Optional client dependencies: `soundfile`, `torch` (for Silero VAD segmentation).
85
+
86
+ ---
87
+
88
+ *Docker image for the Qwen3-ASR concurrent streaming gRPC server.*
89
+
90
+ ## Image Composition
91
+
92
+ | Layer | Detail |
93
+ |-------|--------|
94
+ | **Base OS** | Ubuntu 24.04 LTS |
95
+ | **CUDA** | NVIDIA CUDA 12.8.0 (`nvidia/cuda:12.8.0-runtime-ubuntu24.04`) |
96
+ | **Python** | 3.12 (deadsnakes PPA), installed into `/opt/venv` |
97
+ | **Model** | `Qwen/Qwen3-ASR-1.7B` baked into the image |
98
+ | **Server port** | Container `8002` (map to any host port at runtime) |
99
+
100
+ ## Installation Sequence
101
+
102
+ 1. **System dependencies** — `curl`, `git`, `wget`, `build-essential`, `ffmpeg`, `libsox-*`, `libsndfile1`, `libopus0`, `libffi-dev`.
103
+ 2. **Python 3.12** — installed via deadsnakes PPA, venv created at `/opt/venv`, binaries symlinked to `/usr/local/bin`.
104
+ 3. **pip** — upgraded to latest `pip`, `setuptools`, `wheel`.
105
+ 4. **`qwen-asr[vllm]`** — installs the full Qwen3-ASR stack with vLLM extras (transformers, vllm, torch, accelerate, librosa, soundfile, etc. with pinned versions).
106
+ 5. **flash-attention** — prebuilt wheel for CUDA 12.8 + PyTorch 2.9 + Python 3.12.
107
+ 6. **Project files** — `ASR_Concurrent_Stream/` copied to `/app/ASR_Concurrent_Stream/`.
108
+ 7. **Model download** — `Qwen/Qwen3-ASR-1.7B` downloaded from HuggingFace to `/app/models/Qwen3-ASR-1.7B`.
109
+ 8. **Launcher patch** — `run_asr_grpc_server.sh` is updated:
110
+ - `PORT` → `8002`
111
+ - `MODEL_PATH` → `/app/models/Qwen3-ASR-1.7B`
112
+ - `QWEN3_ASR_PATH` → site-packages `qwen_asr` path
113
+
114
+ ## Build & Run
115
+
116
+ ```bash
117
+ # Build
118
+ docker build -t qwen3-asr-server:latest .
119
+
120
+ # Run (map host port 8002 to container port 8002)
121
+ docker run -d \
122
+ --name qwen3-asr-server \
123
+ --gpus all \
124
+ -p 8002:8002 \
125
+ --shm-size=16g \
126
+ --ulimit memlock=-1 \
127
+ --restart unless-stopped \
128
+ qwen3-asr-server:latest
129
+
130
+ # View logs
131
+ docker logs -f qwen3-asr-server
132
+
133
+ # Stop & remove
134
+ docker stop qwen3-asr-server && docker rm qwen3-asr-server
135
+ ```
136
+
137
+ ## Server Configuration
138
+
139
+ Hard-coded defaults in `server/run_asr_grpc_server.sh`:
140
+
141
+ | Parameter | Value |
142
+ |-----------|-------|
143
+ | Port | `8002` |
144
+ | Model | `/app/models/Qwen3-ASR-1.7B` |
145
+ | GPU memory utilization | `0.35` |
146
+ | Max model length | `4096` |
147
+ | Max concurrent streams | `10` |
148
+ | Max batch size | `8` |
149
+ | Batch timeout | `50` ms |
150
+ | Worker threads | `4` |
151
+
152
+ ## Project Structure (in image)
153
+
154
+ ```
155
+ /app/
156
+ ├── ASR_Concurrent_Stream/
157
+ │ ├── proto/ # gRPC protobuf definitions
158
+ │ ├── server/
159
+ │ │ ├── asr_grpc_server.py
160
+ │ │ ├── _launch_server.py
161
+ │ │ ├── stream_manager.py
162
+ │ │ ├── inference_coordinator.py
163
+ │ │ └── run_asr_grpc_server.sh
164
+ │ └── client/
165
+ │ └── example_client.py
166
+ └── models/
167
+ └── Qwen3-ASR-1.7B/ # Pre-downloaded model
168
+ ```
@@ -0,0 +1,147 @@
1
+ # ASR Concurrent Stream Server
2
+
3
+ Concurrent streaming gRPC server and client for Qwen3-ASR.
4
+
5
+ ## Installation
6
+
7
+ ```bash
8
+ pip install asr-concurrent-stream
9
+ ```
10
+
11
+ For client-side audio loading and VAD segmentation:
12
+
13
+ ```bash
14
+ pip install asr-concurrent-stream[client]
15
+ ```
16
+
17
+ For development (regenerating protobuf code):
18
+
19
+ ```bash
20
+ pip install asr-concurrent-stream[dev]
21
+ ```
22
+
23
+ **Note:** The server requires `vllm` and `transformers` to run the Qwen3-ASR model. These are installed automatically with the package.
24
+
25
+ ## Quick Start
26
+
27
+ ### Start the server
28
+
29
+ ```bash
30
+ asr-concurrent-server
31
+ ```
32
+
33
+ The server reads configuration from environment variables:
34
+
35
+ | Variable | Default | Description |
36
+ |----------|---------|-------------|
37
+ | `ASR_PORT` | `8000` | gRPC server port |
38
+ | `ASR_MODEL_PATH` | `Qwen/Qwen3-ASR-1.7B` | Model path or HuggingFace ID |
39
+ | `ASR_GPU_MEMORY_UTILIZATION` | `0.30` | GPU memory fraction |
40
+ | `ASR_MAX_MODEL_LEN` | `4096` | Maximum model sequence length |
41
+ | `ASR_MAX_CONCURRENT_STREAMS` | `15` | Maximum concurrent streams |
42
+ | `ASR_MAX_BATCH_SIZE` | `8` | Maximum batch size |
43
+ | `ASR_BATCH_TIMEOUT_MS` | `50` | Batch timeout in milliseconds |
44
+ | `ASR_WORKER_THREADS` | `4` | Number of worker threads |
45
+ | `ASR_HEALTH_PORT` | `8080` | HTTP health check port |
46
+
47
+ ### Run the client
48
+
49
+ ```bash
50
+ asr-concurrent-client --audio path/to/audio.wav --language Italian
51
+ ```
52
+
53
+ ## Dependencies
54
+
55
+ | Package | Purpose |
56
+ |---------|---------|
57
+ | `grpcio` | gRPC framework |
58
+ | `numpy` | Audio array handling |
59
+ | `protobuf` | Protocol buffer serialization |
60
+ | `vllm` | vLLM inference engine |
61
+ | `transformers` | Model processor and tokenizer |
62
+
63
+ Optional client dependencies: `soundfile`, `torch` (for Silero VAD segmentation).
64
+
65
+ ---
66
+
67
+ *Docker image for the Qwen3-ASR concurrent streaming gRPC server.*
68
+
69
+ ## Image Composition
70
+
71
+ | Layer | Detail |
72
+ |-------|--------|
73
+ | **Base OS** | Ubuntu 24.04 LTS |
74
+ | **CUDA** | NVIDIA CUDA 12.8.0 (`nvidia/cuda:12.8.0-runtime-ubuntu24.04`) |
75
+ | **Python** | 3.12 (deadsnakes PPA), installed into `/opt/venv` |
76
+ | **Model** | `Qwen/Qwen3-ASR-1.7B` baked into the image |
77
+ | **Server port** | Container `8002` (map to any host port at runtime) |
78
+
79
+ ## Installation Sequence
80
+
81
+ 1. **System dependencies** — `curl`, `git`, `wget`, `build-essential`, `ffmpeg`, `libsox-*`, `libsndfile1`, `libopus0`, `libffi-dev`.
82
+ 2. **Python 3.12** — installed via deadsnakes PPA, venv created at `/opt/venv`, binaries symlinked to `/usr/local/bin`.
83
+ 3. **pip** — upgraded to latest `pip`, `setuptools`, `wheel`.
84
+ 4. **`qwen-asr[vllm]`** — installs the full Qwen3-ASR stack with vLLM extras (transformers, vllm, torch, accelerate, librosa, soundfile, etc. with pinned versions).
85
+ 5. **flash-attention** — prebuilt wheel for CUDA 12.8 + PyTorch 2.9 + Python 3.12.
86
+ 6. **Project files** — `ASR_Concurrent_Stream/` copied to `/app/ASR_Concurrent_Stream/`.
87
+ 7. **Model download** — `Qwen/Qwen3-ASR-1.7B` downloaded from HuggingFace to `/app/models/Qwen3-ASR-1.7B`.
88
+ 8. **Launcher patch** — `run_asr_grpc_server.sh` is updated:
89
+ - `PORT` → `8002`
90
+ - `MODEL_PATH` → `/app/models/Qwen3-ASR-1.7B`
91
+ - `QWEN3_ASR_PATH` → site-packages `qwen_asr` path
92
+
93
+ ## Build & Run
94
+
95
+ ```bash
96
+ # Build
97
+ docker build -t qwen3-asr-server:latest .
98
+
99
+ # Run (map host port 8002 to container port 8002)
100
+ docker run -d \
101
+ --name qwen3-asr-server \
102
+ --gpus all \
103
+ -p 8002:8002 \
104
+ --shm-size=16g \
105
+ --ulimit memlock=-1 \
106
+ --restart unless-stopped \
107
+ qwen3-asr-server:latest
108
+
109
+ # View logs
110
+ docker logs -f qwen3-asr-server
111
+
112
+ # Stop & remove
113
+ docker stop qwen3-asr-server && docker rm qwen3-asr-server
114
+ ```
115
+
116
+ ## Server Configuration
117
+
118
+ Hard-coded defaults in `server/run_asr_grpc_server.sh`:
119
+
120
+ | Parameter | Value |
121
+ |-----------|-------|
122
+ | Port | `8002` |
123
+ | Model | `/app/models/Qwen3-ASR-1.7B` |
124
+ | GPU memory utilization | `0.35` |
125
+ | Max model length | `4096` |
126
+ | Max concurrent streams | `10` |
127
+ | Max batch size | `8` |
128
+ | Batch timeout | `50` ms |
129
+ | Worker threads | `4` |
130
+
131
+ ## Project Structure (in image)
132
+
133
+ ```
134
+ /app/
135
+ ├── ASR_Concurrent_Stream/
136
+ │ ├── proto/ # gRPC protobuf definitions
137
+ │ ├── server/
138
+ │ │ ├── asr_grpc_server.py
139
+ │ │ ├── _launch_server.py
140
+ │ │ ├── stream_manager.py
141
+ │ │ ├── inference_coordinator.py
142
+ │ │ └── run_asr_grpc_server.sh
143
+ │ └── client/
144
+ │ └── example_client.py
145
+ └── models/
146
+ └── Qwen3-ASR-1.7B/ # Pre-downloaded model
147
+ ```
@@ -0,0 +1,168 @@
1
+ Metadata-Version: 2.4
2
+ Name: asr-concurrent-stream
3
+ Version: 1.0.0
4
+ Summary: Concurrent streaming gRPC server and client for Qwen3-ASR
5
+ Author: Alibaba Qwen Team
6
+ License: Apache-2.0
7
+ Project-URL: Homepage, https://github.com/Qwen/Qwen3-ASR
8
+ Keywords: asr,speech-recognition,grpc,streaming,qwen3
9
+ Requires-Python: >=3.9
10
+ Description-Content-Type: text/markdown
11
+ Requires-Dist: grpcio
12
+ Requires-Dist: numpy
13
+ Requires-Dist: protobuf
14
+ Requires-Dist: vllm
15
+ Requires-Dist: transformers
16
+ Provides-Extra: client
17
+ Requires-Dist: soundfile; extra == "client"
18
+ Requires-Dist: torch; extra == "client"
19
+ Provides-Extra: dev
20
+ Requires-Dist: grpcio-tools; extra == "dev"
21
+
22
+ # ASR Concurrent Stream Server
23
+
24
+ Concurrent streaming gRPC server and client for Qwen3-ASR.
25
+
26
+ ## Installation
27
+
28
+ ```bash
29
+ pip install asr-concurrent-stream
30
+ ```
31
+
32
+ For client-side audio loading and VAD segmentation:
33
+
34
+ ```bash
35
+ pip install asr-concurrent-stream[client]
36
+ ```
37
+
38
+ For development (regenerating protobuf code):
39
+
40
+ ```bash
41
+ pip install asr-concurrent-stream[dev]
42
+ ```
43
+
44
+ **Note:** The server requires `vllm` and `transformers` to run the Qwen3-ASR model. These are installed automatically with the package.
45
+
46
+ ## Quick Start
47
+
48
+ ### Start the server
49
+
50
+ ```bash
51
+ asr-concurrent-server
52
+ ```
53
+
54
+ The server reads configuration from environment variables:
55
+
56
+ | Variable | Default | Description |
57
+ |----------|---------|-------------|
58
+ | `ASR_PORT` | `8000` | gRPC server port |
59
+ | `ASR_MODEL_PATH` | `Qwen/Qwen3-ASR-1.7B` | Model path or HuggingFace ID |
60
+ | `ASR_GPU_MEMORY_UTILIZATION` | `0.30` | GPU memory fraction |
61
+ | `ASR_MAX_MODEL_LEN` | `4096` | Maximum model sequence length |
62
+ | `ASR_MAX_CONCURRENT_STREAMS` | `15` | Maximum concurrent streams |
63
+ | `ASR_MAX_BATCH_SIZE` | `8` | Maximum batch size |
64
+ | `ASR_BATCH_TIMEOUT_MS` | `50` | Batch timeout in milliseconds |
65
+ | `ASR_WORKER_THREADS` | `4` | Number of worker threads |
66
+ | `ASR_HEALTH_PORT` | `8080` | HTTP health check port |
67
+
68
+ ### Run the client
69
+
70
+ ```bash
71
+ asr-concurrent-client --audio path/to/audio.wav --language Italian
72
+ ```
73
+
74
+ ## Dependencies
75
+
76
+ | Package | Purpose |
77
+ |---------|---------|
78
+ | `grpcio` | gRPC framework |
79
+ | `numpy` | Audio array handling |
80
+ | `protobuf` | Protocol buffer serialization |
81
+ | `vllm` | vLLM inference engine |
82
+ | `transformers` | Model processor and tokenizer |
83
+
84
+ Optional client dependencies: `soundfile`, `torch` (for Silero VAD segmentation).
85
+
86
+ ---
87
+
88
+ *Docker image for the Qwen3-ASR concurrent streaming gRPC server.*
89
+
90
+ ## Image Composition
91
+
92
+ | Layer | Detail |
93
+ |-------|--------|
94
+ | **Base OS** | Ubuntu 24.04 LTS |
95
+ | **CUDA** | NVIDIA CUDA 12.8.0 (`nvidia/cuda:12.8.0-runtime-ubuntu24.04`) |
96
+ | **Python** | 3.12 (deadsnakes PPA), installed into `/opt/venv` |
97
+ | **Model** | `Qwen/Qwen3-ASR-1.7B` baked into the image |
98
+ | **Server port** | Container `8002` (map to any host port at runtime) |
99
+
100
+ ## Installation Sequence
101
+
102
+ 1. **System dependencies** — `curl`, `git`, `wget`, `build-essential`, `ffmpeg`, `libsox-*`, `libsndfile1`, `libopus0`, `libffi-dev`.
103
+ 2. **Python 3.12** — installed via deadsnakes PPA, venv created at `/opt/venv`, binaries symlinked to `/usr/local/bin`.
104
+ 3. **pip** — upgraded to latest `pip`, `setuptools`, `wheel`.
105
+ 4. **`qwen-asr[vllm]`** — installs the full Qwen3-ASR stack with vLLM extras (transformers, vllm, torch, accelerate, librosa, soundfile, etc. with pinned versions).
106
+ 5. **flash-attention** — prebuilt wheel for CUDA 12.8 + PyTorch 2.9 + Python 3.12.
107
+ 6. **Project files** — `ASR_Concurrent_Stream/` copied to `/app/ASR_Concurrent_Stream/`.
108
+ 7. **Model download** — `Qwen/Qwen3-ASR-1.7B` downloaded from HuggingFace to `/app/models/Qwen3-ASR-1.7B`.
109
+ 8. **Launcher patch** — `run_asr_grpc_server.sh` is updated:
110
+ - `PORT` → `8002`
111
+ - `MODEL_PATH` → `/app/models/Qwen3-ASR-1.7B`
112
+ - `QWEN3_ASR_PATH` → site-packages `qwen_asr` path
113
+
114
+ ## Build & Run
115
+
116
+ ```bash
117
+ # Build
118
+ docker build -t qwen3-asr-server:latest .
119
+
120
+ # Run (map host port 8002 to container port 8002)
121
+ docker run -d \
122
+ --name qwen3-asr-server \
123
+ --gpus all \
124
+ -p 8002:8002 \
125
+ --shm-size=16g \
126
+ --ulimit memlock=-1 \
127
+ --restart unless-stopped \
128
+ qwen3-asr-server:latest
129
+
130
+ # View logs
131
+ docker logs -f qwen3-asr-server
132
+
133
+ # Stop & remove
134
+ docker stop qwen3-asr-server && docker rm qwen3-asr-server
135
+ ```
136
+
137
+ ## Server Configuration
138
+
139
+ Hard-coded defaults in `server/run_asr_grpc_server.sh`:
140
+
141
+ | Parameter | Value |
142
+ |-----------|-------|
143
+ | Port | `8002` |
144
+ | Model | `/app/models/Qwen3-ASR-1.7B` |
145
+ | GPU memory utilization | `0.35` |
146
+ | Max model length | `4096` |
147
+ | Max concurrent streams | `10` |
148
+ | Max batch size | `8` |
149
+ | Batch timeout | `50` ms |
150
+ | Worker threads | `4` |
151
+
152
+ ## Project Structure (in image)
153
+
154
+ ```
155
+ /app/
156
+ ├── ASR_Concurrent_Stream/
157
+ │ ├── proto/ # gRPC protobuf definitions
158
+ │ ├── server/
159
+ │ │ ├── asr_grpc_server.py
160
+ │ │ ├── _launch_server.py
161
+ │ │ ├── stream_manager.py
162
+ │ │ ├── inference_coordinator.py
163
+ │ │ └── run_asr_grpc_server.sh
164
+ │ └── client/
165
+ │ └── example_client.py
166
+ └── models/
167
+ └── Qwen3-ASR-1.7B/ # Pre-downloaded model
168
+ ```
@@ -0,0 +1,21 @@
1
+ README.md
2
+ pyproject.toml
3
+ asr_concurrent_stream.egg-info/PKG-INFO
4
+ asr_concurrent_stream.egg-info/SOURCES.txt
5
+ asr_concurrent_stream.egg-info/dependency_links.txt
6
+ asr_concurrent_stream.egg-info/entry_points.txt
7
+ asr_concurrent_stream.egg-info/requires.txt
8
+ asr_concurrent_stream.egg-info/top_level.txt
9
+ client/__init__.py
10
+ client/example_client.py
11
+ proto/__init__.py
12
+ proto/asr_pb2.py
13
+ proto/asr_pb2_grpc.py
14
+ server/__init__.py
15
+ server/_launch_server.py
16
+ server/asr_grpc_server.py
17
+ server/asr_model.py
18
+ server/asr_processor.py
19
+ server/asr_utils.py
20
+ server/inference_coordinator.py
21
+ server/stream_manager.py
@@ -0,0 +1,3 @@
1
+ [console_scripts]
2
+ asr-concurrent-client = client.example_client:main
3
+ asr-concurrent-server = server._launch_server:main
@@ -0,0 +1,12 @@
1
+ grpcio
2
+ numpy
3
+ protobuf
4
+ vllm
5
+ transformers
6
+
7
+ [client]
8
+ soundfile
9
+ torch
10
+
11
+ [dev]
12
+ grpcio-tools
@@ -0,0 +1,3 @@
1
+ client
2
+ proto
3
+ server
@@ -0,0 +1 @@
1
+ # ASR Concurrent Stream Client