tensorcodec 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- tensorcodec-0.1.0/LICENSE +21 -0
- tensorcodec-0.1.0/PKG-INFO +149 -0
- tensorcodec-0.1.0/README.md +124 -0
- tensorcodec-0.1.0/docs/compatibility.md +58 -0
- tensorcodec-0.1.0/docs/playback_semantics.md +24 -0
- tensorcodec-0.1.0/docs/releasing.md +60 -0
- tensorcodec-0.1.0/licenses/FFmpeg-GPL-3.0.txt +674 -0
- tensorcodec-0.1.0/licenses/FFmpeg-LGPL-3.0.txt +165 -0
- tensorcodec-0.1.0/licenses/FFmpeg-NOTICE.md +129 -0
- tensorcodec-0.1.0/licenses/OpenSSL.txt +177 -0
- tensorcodec-0.1.0/licenses/README.md +16 -0
- tensorcodec-0.1.0/licenses/Zstandard.txt +121 -0
- tensorcodec-0.1.0/native/Cargo.lock +487 -0
- tensorcodec-0.1.0/native/Cargo.toml +18 -0
- tensorcodec-0.1.0/native/src/ffmpeg.rs +861 -0
- tensorcodec-0.1.0/native/src/lib.rs +131 -0
- tensorcodec-0.1.0/pyproject.toml +59 -0
- tensorcodec-0.1.0/scripts/build_ffmpeg.sh +22 -0
- tensorcodec-0.1.0/scripts/build_linux_wheel.sh +11 -0
- tensorcodec-0.1.0/scripts/build_openssl.sh +15 -0
- tensorcodec-0.1.0/scripts/configure_oracle_ffmpeg.py +19 -0
- tensorcodec-0.1.0/src/tensorcodec/__init__.py +7 -0
- tensorcodec-0.1.0/src/tensorcodec/_frame.py +87 -0
- tensorcodec-0.1.0/src/tensorcodec/_metadata.py +45 -0
- tensorcodec-0.1.0/src/tensorcodec/decoders/__init__.py +4 -0
- tensorcodec-0.1.0/src/tensorcodec/decoders/_decoder.py +389 -0
- tensorcodec-0.1.0/src/tensorcodec/py.typed +0 -0
- tensorcodec-0.1.0/tests/__init__.py +1 -0
- tensorcodec-0.1.0/tests/conftest.py +235 -0
- tensorcodec-0.1.0/tests/test_audio_contract.py +34 -0
- tensorcodec-0.1.0/tests/test_differential.py +102 -0
- tensorcodec-0.1.0/tests/test_runtime.py +115 -0
- tensorcodec-0.1.0/tests/test_video_contract.py +205 -0
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Suhwan Choi
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,149 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: tensorcodec
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Classifier: Development Status :: 3 - Alpha
|
|
5
|
+
Classifier: Operating System :: POSIX :: Linux
|
|
6
|
+
Classifier: Programming Language :: Python :: 3
|
|
7
|
+
Classifier: Programming Language :: Rust
|
|
8
|
+
Classifier: Topic :: Multimedia :: Video
|
|
9
|
+
Requires-Dist: numpy>=1.26
|
|
10
|
+
License-File: LICENSE
|
|
11
|
+
License-File: licenses/FFmpeg-GPL-3.0.txt
|
|
12
|
+
License-File: licenses/FFmpeg-LGPL-3.0.txt
|
|
13
|
+
License-File: licenses/FFmpeg-NOTICE.md
|
|
14
|
+
License-File: licenses/OpenSSL.txt
|
|
15
|
+
License-File: licenses/README.md
|
|
16
|
+
License-File: licenses/Zstandard.txt
|
|
17
|
+
Summary: NumPy audio/video decoding with TorchCodec-compatible playback semantics
|
|
18
|
+
Author-email: Suhwan Choi <milkclouds00@gmail.com>
|
|
19
|
+
License-Expression: MIT
|
|
20
|
+
Requires-Python: >=3.10
|
|
21
|
+
Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM
|
|
22
|
+
Project-URL: Issues, https://github.com/MilkClouds/tensorcodec/issues
|
|
23
|
+
Project-URL: Repository, https://github.com/MilkClouds/tensorcodec
|
|
24
|
+
|
|
25
|
+
# TensorCodec
|
|
26
|
+
|
|
27
|
+
**TorchCodec-style video and audio decoding, without PyTorch.**
|
|
28
|
+
|
|
29
|
+
- Use the CPU decoder API and playback rules of **TorchCodec 0.17.0**.
|
|
30
|
+
- Get **NumPy arrays** instead of `torch.Tensor`.
|
|
31
|
+
- Install **NumPy + TensorCodec**. No Torch, PyAV, or FFmpeg CLI at runtime.
|
|
32
|
+
|
|
33
|
+
The goal is predictable frame selection, timestamps and audio ranges with a small
|
|
34
|
+
runtime dependency set. This is a **CPU decoding subset**, not the entire
|
|
35
|
+
TorchCodec package. It does not promise a speedup over PyAV or TorchCodec.
|
|
36
|
+
|
|
37
|
+
## What is available?
|
|
38
|
+
|
|
39
|
+
| Capability | TorchCodec 0.17.0 | TensorCodec 0.1.0 |
|
|
40
|
+
| --- | --- | --- |
|
|
41
|
+
| Python runtime dependency | PyTorch | NumPy |
|
|
42
|
+
| Output arrays | `torch.Tensor` | `numpy.ndarray`; array interface + DLPack |
|
|
43
|
+
| Video index/slice/batch access | Supported | Supported |
|
|
44
|
+
| Playback time/range access | Supported | Supported |
|
|
45
|
+
| CFR, VFR, offset PTS, B-frames | Supported | Tested |
|
|
46
|
+
| Request ordering and duplicates | Preserved | Preserved |
|
|
47
|
+
| Exact / approximate seeking | Supported | Supported; exact is the default |
|
|
48
|
+
| NCHW / NHWC RGB | Supported | Supported |
|
|
49
|
+
| uint8 / float32 video | Supported | Supported for SDR |
|
|
50
|
+
| FPS sampling, custom frame mappings | Supported | Supported |
|
|
51
|
+
| Audio ranges, resampling, channel mixing | Supported | Supported; float32 output |
|
|
52
|
+
| Paths, URLs, bytes, seekable file objects | Supported | Supported |
|
|
53
|
+
| Encoded tensor input | `torch.Tensor` | 1-D uint8 NumPy arrays |
|
|
54
|
+
| CUDA decoding | Supported | **Not implemented** |
|
|
55
|
+
| Decoder transforms | Supported | **Not implemented** |
|
|
56
|
+
| HDR inputs / display rotation | Supported | **Rejected explicitly** |
|
|
57
|
+
| Other modules, including samplers/encoders | Available | **Outside the initial scope** |
|
|
58
|
+
|
|
59
|
+
### Compatibility means
|
|
60
|
+
|
|
61
|
+
- Match the supported CPU API's frame selection, ordering, timing and metadata.
|
|
62
|
+
- Check behavior independently **and** against pinned TorchCodec 0.17.0.
|
|
63
|
+
- Allow color-conversion rounding: at most 1 uint8 unit or 1/65535 for float32
|
|
64
|
+
in the tested cases. Do not claim identical pixels across every FFmpeg build.
|
|
65
|
+
- Accept empty index lists, including the case affected by the reference's
|
|
66
|
+
empty-list dtype inference bug.
|
|
67
|
+
|
|
68
|
+
Details and the tested scope: [compatibility contract](docs/compatibility.md).
|
|
69
|
+
|
|
70
|
+
### Current limits
|
|
71
|
+
|
|
72
|
+
- Binary wheels: **Linux x86_64, glibc 2.28+, CPython 3.10+**.
|
|
73
|
+
- No macOS or Windows wheels yet; free-threaded Python is not a release target.
|
|
74
|
+
- Exact video seeking scans packet timestamps when opening the decoder.
|
|
75
|
+
- Audio range queries currently decode from the beginning; late ranges can be
|
|
76
|
+
expensive.
|
|
77
|
+
- NumPy return types require caller changes where code expects Torch tensors.
|
|
78
|
+
- Historical avdec benchmarks are not TensorCodec performance results.
|
|
79
|
+
|
|
80
|
+
## Install
|
|
81
|
+
|
|
82
|
+
```sh
|
|
83
|
+
python -m pip install tensorcodec
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
Linux wheels bundle shared FFmpeg libraries. Source builds need Rust, libclang
|
|
87
|
+
and FFmpeg 7 development headers/libraries.
|
|
88
|
+
|
|
89
|
+
Before the first PyPI upload, install a local wheel:
|
|
90
|
+
|
|
91
|
+
```sh
|
|
92
|
+
python -m pip install dist/tensorcodec-*.whl
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
## Use
|
|
96
|
+
|
|
97
|
+
```python
|
|
98
|
+
from tensorcodec.decoders import VideoDecoder, AudioDecoder
|
|
99
|
+
|
|
100
|
+
with VideoDecoder("video.mp4") as video:
|
|
101
|
+
frame = video.get_frame_played_at(1.25) # frame playing at this time
|
|
102
|
+
print(frame.data.shape) # CHW NumPy array
|
|
103
|
+
batch = video.get_frames_at([4, 0, 4]) # order and duplicates preserved
|
|
104
|
+
clip = video.get_frames_played_in_range(0, 1, fps=8)
|
|
105
|
+
|
|
106
|
+
with AudioDecoder("audio.wav", sample_rate=16000, num_channels=1) as audio:
|
|
107
|
+
samples = audio.get_samples_played_in_range(0, 1)
|
|
108
|
+
print(samples.data.shape) # channels × samples, float32
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
Decoded arrays keep their storage after the decoder closes. Input file objects
|
|
112
|
+
remain caller-owned.
|
|
113
|
+
|
|
114
|
+
## Implementation
|
|
115
|
+
|
|
116
|
+
| Layer | Responsibility |
|
|
117
|
+
| --- | --- |
|
|
118
|
+
| Python | Public API, frame/time selection, validation, result objects |
|
|
119
|
+
| Rust + PyO3 | FFmpeg handles, seeking/decoding, color conversion, resampling |
|
|
120
|
+
| FFmpeg | Codec and container implementations |
|
|
121
|
+
|
|
122
|
+
A batch crosses the Python/Rust boundary once. Native decoding releases the GIL.
|
|
123
|
+
|
|
124
|
+
## Development and verification
|
|
125
|
+
|
|
126
|
+
```sh
|
|
127
|
+
# Requires Rust, Clang/libclang, pkg-config and FFmpeg 7 development libraries.
|
|
128
|
+
python -m venv .venv
|
|
129
|
+
. .venv/bin/activate
|
|
130
|
+
python -m pip install numpy pytest ruff 'maturin>=1.8,<2'
|
|
131
|
+
maturin develop --locked
|
|
132
|
+
|
|
133
|
+
# Reference dependencies are for tests only.
|
|
134
|
+
python -m pip install torch==2.14.1 torchcodec==0.17.0 \
|
|
135
|
+
--index-url https://download.pytorch.org/whl/cpu
|
|
136
|
+
pytest tests/test_video_contract.py tests/test_audio_contract.py --backend torchcodec
|
|
137
|
+
pytest --compare
|
|
138
|
+
```
|
|
139
|
+
|
|
140
|
+
Tests generate fixtures with FFmpeg/ffprobe and Python's `wave` module.
|
|
141
|
+
`--compare` requires the exact oracle version; otherwise differential tests skip.
|
|
142
|
+
The old avdec decoder and tests are never executed.
|
|
143
|
+
|
|
144
|
+
- [Playback rules](docs/playback_semantics.md)
|
|
145
|
+
- [Release builds and PyPI publishing](docs/releasing.md)
|
|
146
|
+
- [Native dependency licenses and source/build notices](licenses/README.md)
|
|
147
|
+
|
|
148
|
+
TensorCodec's own code is MIT licensed.
|
|
149
|
+
|
|
@@ -0,0 +1,124 @@
|
|
|
1
|
+
# TensorCodec
|
|
2
|
+
|
|
3
|
+
**TorchCodec-style video and audio decoding, without PyTorch.**
|
|
4
|
+
|
|
5
|
+
- Use the CPU decoder API and playback rules of **TorchCodec 0.17.0**.
|
|
6
|
+
- Get **NumPy arrays** instead of `torch.Tensor`.
|
|
7
|
+
- Install **NumPy + TensorCodec**. No Torch, PyAV, or FFmpeg CLI at runtime.
|
|
8
|
+
|
|
9
|
+
The goal is predictable frame selection, timestamps and audio ranges with a small
|
|
10
|
+
runtime dependency set. This is a **CPU decoding subset**, not the entire
|
|
11
|
+
TorchCodec package. It does not promise a speedup over PyAV or TorchCodec.
|
|
12
|
+
|
|
13
|
+
## What is available?
|
|
14
|
+
|
|
15
|
+
| Capability | TorchCodec 0.17.0 | TensorCodec 0.1.0 |
|
|
16
|
+
| --- | --- | --- |
|
|
17
|
+
| Python runtime dependency | PyTorch | NumPy |
|
|
18
|
+
| Output arrays | `torch.Tensor` | `numpy.ndarray`; array interface + DLPack |
|
|
19
|
+
| Video index/slice/batch access | Supported | Supported |
|
|
20
|
+
| Playback time/range access | Supported | Supported |
|
|
21
|
+
| CFR, VFR, offset PTS, B-frames | Supported | Tested |
|
|
22
|
+
| Request ordering and duplicates | Preserved | Preserved |
|
|
23
|
+
| Exact / approximate seeking | Supported | Supported; exact is the default |
|
|
24
|
+
| NCHW / NHWC RGB | Supported | Supported |
|
|
25
|
+
| uint8 / float32 video | Supported | Supported for SDR |
|
|
26
|
+
| FPS sampling, custom frame mappings | Supported | Supported |
|
|
27
|
+
| Audio ranges, resampling, channel mixing | Supported | Supported; float32 output |
|
|
28
|
+
| Paths, URLs, bytes, seekable file objects | Supported | Supported |
|
|
29
|
+
| Encoded tensor input | `torch.Tensor` | 1-D uint8 NumPy arrays |
|
|
30
|
+
| CUDA decoding | Supported | **Not implemented** |
|
|
31
|
+
| Decoder transforms | Supported | **Not implemented** |
|
|
32
|
+
| HDR inputs / display rotation | Supported | **Rejected explicitly** |
|
|
33
|
+
| Other modules, including samplers/encoders | Available | **Outside the initial scope** |
|
|
34
|
+
|
|
35
|
+
### Compatibility means
|
|
36
|
+
|
|
37
|
+
- Match the supported CPU API's frame selection, ordering, timing and metadata.
|
|
38
|
+
- Check behavior independently **and** against pinned TorchCodec 0.17.0.
|
|
39
|
+
- Allow color-conversion rounding: at most 1 uint8 unit or 1/65535 for float32
|
|
40
|
+
in the tested cases. Do not claim identical pixels across every FFmpeg build.
|
|
41
|
+
- Accept empty index lists, including the case affected by the reference's
|
|
42
|
+
empty-list dtype inference bug.
|
|
43
|
+
|
|
44
|
+
Details and the tested scope: [compatibility contract](docs/compatibility.md).
|
|
45
|
+
|
|
46
|
+
### Current limits
|
|
47
|
+
|
|
48
|
+
- Binary wheels: **Linux x86_64, glibc 2.28+, CPython 3.10+**.
|
|
49
|
+
- No macOS or Windows wheels yet; free-threaded Python is not a release target.
|
|
50
|
+
- Exact video seeking scans packet timestamps when opening the decoder.
|
|
51
|
+
- Audio range queries currently decode from the beginning; late ranges can be
|
|
52
|
+
expensive.
|
|
53
|
+
- NumPy return types require caller changes where code expects Torch tensors.
|
|
54
|
+
- Historical avdec benchmarks are not TensorCodec performance results.
|
|
55
|
+
|
|
56
|
+
## Install
|
|
57
|
+
|
|
58
|
+
```sh
|
|
59
|
+
python -m pip install tensorcodec
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
Linux wheels bundle shared FFmpeg libraries. Source builds need Rust, libclang
|
|
63
|
+
and FFmpeg 7 development headers/libraries.
|
|
64
|
+
|
|
65
|
+
Before the first PyPI upload, install a local wheel:
|
|
66
|
+
|
|
67
|
+
```sh
|
|
68
|
+
python -m pip install dist/tensorcodec-*.whl
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
## Use
|
|
72
|
+
|
|
73
|
+
```python
|
|
74
|
+
from tensorcodec.decoders import VideoDecoder, AudioDecoder
|
|
75
|
+
|
|
76
|
+
with VideoDecoder("video.mp4") as video:
|
|
77
|
+
frame = video.get_frame_played_at(1.25) # frame playing at this time
|
|
78
|
+
print(frame.data.shape) # CHW NumPy array
|
|
79
|
+
batch = video.get_frames_at([4, 0, 4]) # order and duplicates preserved
|
|
80
|
+
clip = video.get_frames_played_in_range(0, 1, fps=8)
|
|
81
|
+
|
|
82
|
+
with AudioDecoder("audio.wav", sample_rate=16000, num_channels=1) as audio:
|
|
83
|
+
samples = audio.get_samples_played_in_range(0, 1)
|
|
84
|
+
print(samples.data.shape) # channels × samples, float32
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
Decoded arrays keep their storage after the decoder closes. Input file objects
|
|
88
|
+
remain caller-owned.
|
|
89
|
+
|
|
90
|
+
## Implementation
|
|
91
|
+
|
|
92
|
+
| Layer | Responsibility |
|
|
93
|
+
| --- | --- |
|
|
94
|
+
| Python | Public API, frame/time selection, validation, result objects |
|
|
95
|
+
| Rust + PyO3 | FFmpeg handles, seeking/decoding, color conversion, resampling |
|
|
96
|
+
| FFmpeg | Codec and container implementations |
|
|
97
|
+
|
|
98
|
+
A batch crosses the Python/Rust boundary once. Native decoding releases the GIL.
|
|
99
|
+
|
|
100
|
+
## Development and verification
|
|
101
|
+
|
|
102
|
+
```sh
|
|
103
|
+
# Requires Rust, Clang/libclang, pkg-config and FFmpeg 7 development libraries.
|
|
104
|
+
python -m venv .venv
|
|
105
|
+
. .venv/bin/activate
|
|
106
|
+
python -m pip install numpy pytest ruff 'maturin>=1.8,<2'
|
|
107
|
+
maturin develop --locked
|
|
108
|
+
|
|
109
|
+
# Reference dependencies are for tests only.
|
|
110
|
+
python -m pip install torch==2.14.1 torchcodec==0.17.0 \
|
|
111
|
+
--index-url https://download.pytorch.org/whl/cpu
|
|
112
|
+
pytest tests/test_video_contract.py tests/test_audio_contract.py --backend torchcodec
|
|
113
|
+
pytest --compare
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
Tests generate fixtures with FFmpeg/ffprobe and Python's `wave` module.
|
|
117
|
+
`--compare` requires the exact oracle version; otherwise differential tests skip.
|
|
118
|
+
The old avdec decoder and tests are never executed.
|
|
119
|
+
|
|
120
|
+
- [Playback rules](docs/playback_semantics.md)
|
|
121
|
+
- [Release builds and PyPI publishing](docs/releasing.md)
|
|
122
|
+
- [Native dependency licenses and source/build notices](licenses/README.md)
|
|
123
|
+
|
|
124
|
+
TensorCodec's own code is MIT licensed.
|
|
@@ -0,0 +1,58 @@
|
|
|
1
|
+
# TensorCodec compatibility contract
|
|
2
|
+
|
|
3
|
+
Reference: **TorchCodec 0.17.0**, CPU audio/video decoding. Python operations,
|
|
4
|
+
frame selection, ordering, timestamps, durations, stream selection and metadata
|
|
5
|
+
are tested independently and against this pinned version. Native decoding uses
|
|
6
|
+
FFmpeg; arrays are returned as NumPy instead of torch.Tensor.
|
|
7
|
+
|
|
8
|
+
## Public surface
|
|
9
|
+
|
|
10
|
+
- `tensorcodec.Frame`, `FrameBatch`, `AudioSamples`
|
|
11
|
+
- `tensorcodec.decoders.VideoDecoder`: constructor, `len`, integer/slice indexing,
|
|
12
|
+
`get_frame_at`, `get_frames_at`, `get_frames_in_range`, `get_frame_played_at`,
|
|
13
|
+
`get_frames_played_at`, `get_frames_played_in_range`, metadata, stream_index.
|
|
14
|
+
Also `get_all_frames`, FPS resampling and custom JSON frame mappings.
|
|
15
|
+
- `tensorcodec.decoders.AudioDecoder`: constructor, metadata, stream_index,
|
|
16
|
+
`get_all_samples`, `get_samples_played_in_range`, resampling and channel mixing.
|
|
17
|
+
- NumPy uint8 or float32 video and float32 audio; float64 batch timestamps/durations.
|
|
18
|
+
Float32 RGB uses 16-bit color conversion rather than scaling uint8 output.
|
|
19
|
+
- NCHW/NHWC, paths/URLs, encoded bytes, uint8 arrays and seekable file-like input.
|
|
20
|
+
- Exact and approximate video seeking; exact is the default.
|
|
21
|
+
|
|
22
|
+
CUDA, torch inputs, torchvision transforms, HDR tone mapping, rotated video,
|
|
23
|
+
encoders and samplers are outside the
|
|
24
|
+
initial CPU decoding contract. Unsupported device/transform options fail explicitly.
|
|
25
|
+
Do not advertise full-package or torch.Tensor type compatibility.
|
|
26
|
+
|
|
27
|
+
## Playback rules
|
|
28
|
+
|
|
29
|
+
Timestamp retrieval selects frame i for `pts[i] <= t < pts[i+1]`; the last frame
|
|
30
|
+
ends at the content-derived stream end. This does not imply that returned frame
|
|
31
|
+
duration equals the next PTS gap. In 0.17.0 the implementation of time ranges
|
|
32
|
+
includes the frame playing at start, even when its PTS is before start. It excludes
|
|
33
|
+
the frame starting exactly at stop. This differs from the reference docstring;
|
|
34
|
+
we follow the tested implementation, with empty output when start equals stop.
|
|
35
|
+
Indices preserve input order and duplicates. Shape/layout, bounds, defaults and
|
|
36
|
+
exception classes follow the reference, including its restrictions on slice steps.
|
|
37
|
+
|
|
38
|
+
## Tests first
|
|
39
|
+
|
|
40
|
+
Fixtures are generated with FFmpeg from known grayscale frame identities and
|
|
41
|
+
explicit timestamp schedules, and with Python's wave module from known PCM.
|
|
42
|
+
ffprobe validates encoded packet metadata independently. The original avdec
|
|
43
|
+
decoder and tests are not executed. Run the same contract tests against the
|
|
44
|
+
reference before adding implementation:
|
|
45
|
+
|
|
46
|
+
```sh
|
|
47
|
+
pytest tests/test_video_contract.py tests/test_audio_contract.py --backend torchcodec
|
|
48
|
+
pytest --compare
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
Comparisons check timing separately from pixels. One uint8 RGB unit or 1/65535
|
|
52
|
+
for float32 RGB is allowed for conversion rounding; frame identities have independent
|
|
53
|
+
expectations. Tests requiring the oracle must fail on a missing/wrong reference
|
|
54
|
+
when `--compare` is requested. Production installation does not require torch.
|
|
55
|
+
|
|
56
|
+
Intentional fix: TensorCodec accepts empty index lists. TorchCodec 0.17.0 infers
|
|
57
|
+
float for empty index lists; oracle tests use explicitly typed input tensors to
|
|
58
|
+
isolate playback semantics from that conversion bug.
|
|
@@ -0,0 +1,24 @@
|
|
|
1
|
+
# TensorCodec playback semantics
|
|
2
|
+
|
|
3
|
+
The supported behavior is specified in [compatibility.md](compatibility.md) and
|
|
4
|
+
executable tests in `tests/test_video_contract.py` and `tests/test_audio_contract.py`.
|
|
5
|
+
The reference is TorchCodec 0.17.0. These replace the previous avdec/PyAV design.
|
|
6
|
+
|
|
7
|
+
Exact video seeking scans presentation timestamps and preceding key-frame PTS.
|
|
8
|
+
Index requests use that map; time requests select the frame playing at the requested
|
|
9
|
+
time. The native decoder seeks to a key frame and decodes forward to the exact
|
|
10
|
+
PTS, restoring caller order and duplicates. It fails if the mapped frame cannot
|
|
11
|
+
be found rather than silently returning another frame.
|
|
12
|
+
|
|
13
|
+
Approximate seeking derives indices and time selection from header FPS. It does
|
|
14
|
+
not provide the content-derived VFR guarantees of exact mode.
|
|
15
|
+
|
|
16
|
+
For a time range, include the frame playing at its start and exclude a frame
|
|
17
|
+
starting at its stop. With `fps`, sample a uniform grid starting at the requested
|
|
18
|
+
start and return grid timestamps/durations, matching the reference implementation.
|
|
19
|
+
Packet duration and the next frame's PTS gap are distinct quantities.
|
|
20
|
+
|
|
21
|
+
Audio decoding preserves codec priming and encoder delay, then resamples with
|
|
22
|
+
FFmpeg. Range selection rounds offsets at the output sample rate. Current audio
|
|
23
|
+
queries replay decoding from the beginning; this trades speed for deterministic
|
|
24
|
+
range output and is a future optimization boundary.
|
|
@@ -0,0 +1,60 @@
|
|
|
1
|
+
# Publishing TensorCodec
|
|
2
|
+
|
|
3
|
+
Release version: `0.1.0`. Distribution and import name: `tensorcodec`.
|
|
4
|
+
The first binary release targets Linux x86_64, glibc 2.28+, CPython 3.10+.
|
|
5
|
+
The wheel bundles shared FFmpeg 7.1.5 and OpenSSL 3.5.9 LTS; its only Python
|
|
6
|
+
runtime dependency is NumPy. macOS/Windows wheels are not yet provided.
|
|
7
|
+
|
|
8
|
+
## One-time account setup
|
|
9
|
+
|
|
10
|
+
1. Rename `MilkClouds/avdec` to `tensorcodec` in GitHub repository Settings.
|
|
11
|
+
Preserve the current visibility; publishing does not require making it public.
|
|
12
|
+
2. Open https://pypi.org/manage/account/publishing/ and add a **pending GitHub
|
|
13
|
+
publisher** for a new project with these values:
|
|
14
|
+
|
|
15
|
+
| Field | Value |
|
|
16
|
+
| --- | --- |
|
|
17
|
+
| PyPI project name | `tensorcodec` |
|
|
18
|
+
| GitHub owner | `MilkClouds` |
|
|
19
|
+
| Repository | `tensorcodec` |
|
|
20
|
+
| Workflow filename | `publish.yml` |
|
|
21
|
+
| Environment | `pypi` |
|
|
22
|
+
|
|
23
|
+
A pending publisher creates the project on its first successful upload. If the
|
|
24
|
+
project already exists, register the publisher in that project's Publishing
|
|
25
|
+
settings instead. Do not put an API token in the repository or chat.
|
|
26
|
+
|
|
27
|
+
## Release
|
|
28
|
+
|
|
29
|
+
Run the **Publish to PyPI** workflow on `main`. It builds the portable Linux wheel
|
|
30
|
+
and source distribution, checks package metadata, validates the pinned oracle
|
|
31
|
+
and compares playback before uploading through PyPI Trusted Publishing. It uses
|
|
32
|
+
the existing GitHub `pypi` environment. Publication fails if authorization is
|
|
33
|
+
missing, tests fail, or the version has already been uploaded.
|
|
34
|
+
|
|
35
|
+
```sh
|
|
36
|
+
gh workflow run publish.yml --repo MilkClouds/tensorcodec --ref main
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
For a build and full validation without uploading, pass `--field publish=false`.
|
|
40
|
+
|
|
41
|
+
Check the workflow and https://pypi.org/project/tensorcodec/0.1.0/ before reporting
|
|
42
|
+
success. Verify a fresh `pip install tensorcodec==0.1.0` and a decode without
|
|
43
|
+
Torch/PyAV. Update the version before subsequent releases; PyPI versions cannot
|
|
44
|
+
be overwritten.
|
|
45
|
+
|
|
46
|
+
The local Linux build is reproducible using `scripts/build_linux_wheel.sh` inside
|
|
47
|
+
`quay.io/pypa/manylinux_2_28_x86_64` with Rust, maturin, libclang, NASM and Perl.
|
|
48
|
+
Both native source archives are version- and checksum-pinned. Their licensing
|
|
49
|
+
and source links are recorded in `licenses/README.md`.
|
|
50
|
+
|
|
51
|
+
## CI versus release builds
|
|
52
|
+
|
|
53
|
+
- Ordinary CI uses prebuilt conda-forge FFmpeg 7.1.1 through Pixi, including its
|
|
54
|
+
headers and shared libraries. It builds only the TensorCodec extension.
|
|
55
|
+
- PyPI wheels use the smaller LGPL FFmpeg 7.1.5 build plus OpenSSL 3.5.9.
|
|
56
|
+
Their native prefix is cached by the build-script checksums. This preserves the
|
|
57
|
+
wheel's codec set, dependency size and licensing rather than bundling the full
|
|
58
|
+
conda-forge dependency graph.
|
|
59
|
+
- Release validation still tests the installed repaired wheel. The fixture CLI
|
|
60
|
+
can be FFmpeg 6 or 7; fixtures explicitly remove auxiliary sentinel packets.
|