seqpro 0.15.0__tar.gz → 0.15.2__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (117) hide show
  1. seqpro-0.15.2/.github/workflows/bench.yaml +31 -0
  2. seqpro-0.15.2/.github/workflows/lint.yaml +18 -0
  3. {seqpro-0.15.0 → seqpro-0.15.2}/.github/workflows/publish.yaml +12 -0
  4. {seqpro-0.15.0 → seqpro-0.15.2}/.pre-commit-config.yaml +6 -0
  5. {seqpro-0.15.0 → seqpro-0.15.2}/CHANGELOG.md +17 -0
  6. {seqpro-0.15.0 → seqpro-0.15.2}/PKG-INFO +1 -1
  7. seqpro-0.15.2/docs/superpowers/plans/2026-06-12-tokenize-lut-codspeed.md +450 -0
  8. seqpro-0.15.2/docs/superpowers/specs/2026-06-12-tokenize-lut-codspeed-design.md +203 -0
  9. {seqpro-0.15.0 → seqpro-0.15.2}/pixi.lock +119 -0
  10. {seqpro-0.15.0 → seqpro-0.15.2}/pixi.toml +4 -0
  11. {seqpro-0.15.0 → seqpro-0.15.2}/pyproject.toml +7 -1
  12. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/_analyzers.py +22 -22
  13. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/_cleaners.py +4 -2
  14. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/_coords.py +6 -6
  15. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/_encoders.py +55 -25
  16. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/_modifiers.py +24 -23
  17. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/_numba.py +30 -10
  18. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/alphabets/_alphabets.py +18 -4
  19. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/bed.py +11 -7
  20. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/experimental/_experimental.py +2 -0
  21. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/experimental/_visualizers.py +5 -0
  22. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/rag/_array.py +69 -21
  23. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/rag/_ops.py +21 -13
  24. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/rag/_utils.py +2 -1
  25. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/transforms/augmentation.py +15 -13
  26. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/transforms/tmm.py +11 -9
  27. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/xr/__init__.py +3 -0
  28. seqpro-0.15.2/tests/test_bench_tokenize.py +121 -0
  29. {seqpro-0.15.0 → seqpro-0.15.2}/tests/test_rag_to_packed.py +59 -0
  30. seqpro-0.15.2/tests/test_tokenize.py +262 -0
  31. seqpro-0.15.0/.github/workflows/lint.yaml +0 -9
  32. seqpro-0.15.0/tests/test_tokenize.py +0 -147
  33. {seqpro-0.15.0 → seqpro-0.15.2}/.claude/skills/zensical/SKILL.md +0 -0
  34. {seqpro-0.15.0 → seqpro-0.15.2}/.gitattributes +0 -0
  35. {seqpro-0.15.0 → seqpro-0.15.2}/.github/workflows/bump.yaml +0 -0
  36. {seqpro-0.15.0 → seqpro-0.15.2}/.github/workflows/docs.yml +0 -0
  37. {seqpro-0.15.0 → seqpro-0.15.2}/.github/workflows/merge.yaml +0 -0
  38. {seqpro-0.15.0 → seqpro-0.15.2}/.github/workflows/release-pipeline.yaml +0 -0
  39. {seqpro-0.15.0 → seqpro-0.15.2}/.github/workflows/release.yaml +0 -0
  40. {seqpro-0.15.0 → seqpro-0.15.2}/.github/workflows/test.yaml +0 -0
  41. {seqpro-0.15.0 → seqpro-0.15.2}/.gitignore +0 -0
  42. {seqpro-0.15.0 → seqpro-0.15.2}/CLAUDE.md +0 -0
  43. {seqpro-0.15.0 → seqpro-0.15.2}/Cargo.lock +0 -0
  44. {seqpro-0.15.0 → seqpro-0.15.2}/Cargo.toml +0 -0
  45. {seqpro-0.15.0 → seqpro-0.15.2}/LICENSE +0 -0
  46. {seqpro-0.15.0 → seqpro-0.15.2}/README.md +0 -0
  47. {seqpro-0.15.0 → seqpro-0.15.2}/benches/kshuffle.rs +0 -0
  48. {seqpro-0.15.0 → seqpro-0.15.2}/benchmarks/bench_to_packed.py +0 -0
  49. {seqpro-0.15.0 → seqpro-0.15.2}/docs/api/alphabets.md +0 -0
  50. {seqpro-0.15.0 → seqpro-0.15.2}/docs/api/bed.md +0 -0
  51. {seqpro-0.15.0 → seqpro-0.15.2}/docs/api/gtf.md +0 -0
  52. {seqpro-0.15.0 → seqpro-0.15.2}/docs/api/index.md +0 -0
  53. {seqpro-0.15.0 → seqpro-0.15.2}/docs/api/ragged.md +0 -0
  54. {seqpro-0.15.0 → seqpro-0.15.2}/docs/api/types.md +0 -0
  55. {seqpro-0.15.0 → seqpro-0.15.2}/docs/index.md +0 -0
  56. {seqpro-0.15.0 → seqpro-0.15.2}/docs/ragged.md +0 -0
  57. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/plans/2026-05-04-ragged-record-array.md +0 -0
  58. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/plans/2026-05-05-documentation-site.md +0 -0
  59. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/plans/2026-05-05-narwhals-coord-schema.md +0 -0
  60. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/plans/2026-05-05-ragged-zip-and-record-introspection.md +0 -0
  61. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/plans/2026-05-20-kshuffle-optimization.md +0 -0
  62. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/plans/2026-05-20-kshuffle-pooled-buffers-and-k2-fast-path.md +0 -0
  63. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/plans/2026-05-20-kshuffle-wilson-single-pass.md +0 -0
  64. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/plans/2026-05-20-release-pipeline.md +0 -0
  65. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/plans/2026-05-28-translate-lut-validation.md +0 -0
  66. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/plans/2026-05-31-flat-buffer-to-padded.md +0 -0
  67. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/plans/2026-05-31-rag-to-packed.md +0 -0
  68. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/plans/2026-06-05-translate-unknown-codon-policy.md +0 -0
  69. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/specs/2026-05-04-ragged-record-array-design.md +0 -0
  70. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/specs/2026-05-05-documentation-site-design.md +0 -0
  71. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/specs/2026-05-05-narwhals-coord-schema-design.md +0 -0
  72. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/specs/2026-05-05-ragged-zip-and-record-introspection-design.md +0 -0
  73. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/specs/2026-05-20-kshuffle-optimization-design.md +0 -0
  74. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/specs/2026-05-20-kshuffle-pooled-buffers-and-k2-fast-path-design.md +0 -0
  75. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/specs/2026-05-20-kshuffle-wilson-single-pass-design.md +0 -0
  76. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/specs/2026-05-20-release-pipeline-design.md +0 -0
  77. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/specs/2026-05-28-translate-lut-and-validation-design.md +0 -0
  78. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/specs/2026-05-31-flat-buffer-to-padded-design.md +0 -0
  79. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/specs/2026-05-31-rag-to-packed-design.md +0 -0
  80. {seqpro-0.15.0 → seqpro-0.15.2}/docs/superpowers/specs/2026-06-05-translate-unknown-codon-policy-design.md +0 -0
  81. {seqpro-0.15.0 → seqpro-0.15.2}/meta.yaml +0 -0
  82. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/__init__.py +0 -0
  83. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/_types.py +0 -0
  84. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/_utils.py +0 -0
  85. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/alphabets/__init__.py +0 -0
  86. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/gtf.py +0 -0
  87. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/py.typed +0 -0
  88. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/rag/__init__.py +0 -0
  89. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/rag/_gufuncs.py +0 -0
  90. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/rag/_types.py +0 -0
  91. {seqpro-0.15.0 → seqpro-0.15.2}/python/seqpro/transforms/__init__.py +0 -0
  92. {seqpro-0.15.0 → seqpro-0.15.2}/scratch_bench_rc.py +0 -0
  93. {seqpro-0.15.0 → seqpro-0.15.2}/scratch_bench_to_padded.py +0 -0
  94. {seqpro-0.15.0 → seqpro-0.15.2}/skills/seqpro/SKILL.md +0 -0
  95. {seqpro-0.15.0 → seqpro-0.15.2}/src/kmer_encode.rs +0 -0
  96. {seqpro-0.15.0 → seqpro-0.15.2}/src/kshuffle.rs +0 -0
  97. {seqpro-0.15.0 → seqpro-0.15.2}/src/kshuffle_ref.rs +0 -0
  98. {seqpro-0.15.0 → seqpro-0.15.2}/src/lib.rs +0 -0
  99. {seqpro-0.15.0 → seqpro-0.15.2}/tests/_shape_fixtures.py +0 -0
  100. {seqpro-0.15.0 → seqpro-0.15.2}/tests/bed/test_pyranges.py +0 -0
  101. {seqpro-0.15.0 → seqpro-0.15.2}/tests/bed/test_read.py +0 -0
  102. {seqpro-0.15.0 → seqpro-0.15.2}/tests/bed/test_sort.py +0 -0
  103. {seqpro-0.15.0 → seqpro-0.15.2}/tests/bed/test_with_length.py +0 -0
  104. {seqpro-0.15.0 → seqpro-0.15.2}/tests/bench_translate_lut.py +0 -0
  105. {seqpro-0.15.0 → seqpro-0.15.2}/tests/conftest.py +0 -0
  106. {seqpro-0.15.0 → seqpro-0.15.2}/tests/test_analyzers.py +0 -0
  107. {seqpro-0.15.0 → seqpro-0.15.2}/tests/test_coords.py +0 -0
  108. {seqpro-0.15.0 → seqpro-0.15.2}/tests/test_encoders.py +0 -0
  109. {seqpro-0.15.0 → seqpro-0.15.2}/tests/test_modifiers.py +0 -0
  110. {seqpro-0.15.0 → seqpro-0.15.2}/tests/test_ohe.py +0 -0
  111. {seqpro-0.15.0 → seqpro-0.15.2}/tests/test_ragged.py +0 -0
  112. {seqpro-0.15.0 → seqpro-0.15.2}/tests/test_ragged_rc.py +0 -0
  113. {seqpro-0.15.0 → seqpro-0.15.2}/tests/test_ragged_to_padded.py +0 -0
  114. {seqpro-0.15.0 → seqpro-0.15.2}/tests/test_shape_matrix.py +0 -0
  115. {seqpro-0.15.0 → seqpro-0.15.2}/tests/test_transforms.py +0 -0
  116. {seqpro-0.15.0 → seqpro-0.15.2}/tests/test_translate.py +0 -0
  117. {seqpro-0.15.0 → seqpro-0.15.2}/zensical.toml +0 -0
@@ -0,0 +1,31 @@
1
+ name: Benchmarks
2
+
3
+ on:
4
+ pull_request:
5
+ branches:
6
+ - main
7
+ workflow_dispatch:
8
+
9
+ jobs:
10
+ benchmarks:
11
+ runs-on: ubuntu-latest
12
+ name: "Run CodSpeed microbenchmarks"
13
+ steps:
14
+ - name: Check out
15
+ uses: actions/checkout@v4
16
+ with:
17
+ fetch-depth: 0
18
+ - name: Setup pixi
19
+ uses: prefix-dev/setup-pixi@v0.9.5
20
+ with:
21
+ pixi-version: v0.67.2
22
+ cache: true
23
+ environments: bench
24
+ locked: false
25
+ - name: Build extension
26
+ run: pixi run -e bench maturin develop
27
+ - name: Run benchmarks
28
+ uses: CodSpeedHQ/action@v3
29
+ with:
30
+ token: ${{ secrets.CODSPEED_TOKEN }}
31
+ run: pixi run -e bench bench
@@ -0,0 +1,18 @@
1
+ name: Prek checks
2
+ on: [push, pull_request]
3
+
4
+ jobs:
5
+ prek:
6
+ runs-on: ubuntu-latest
7
+ steps:
8
+ - uses: actions/checkout@v6
9
+ # The `pyrefly` hook's entry is `pixi run typecheck`, so pixi and the
10
+ # `default` environment must be available to the prek runner.
11
+ - name: Setup pixi
12
+ uses: prefix-dev/setup-pixi@v0.9.5
13
+ with:
14
+ pixi-version: v0.67.2
15
+ cache: true
16
+ environments: default
17
+ locked: false
18
+ - uses: j178/prek-action@v2
@@ -74,6 +74,9 @@ jobs:
74
74
  sccache: ${{ !startsWith(github.ref, 'refs/tags/') }}
75
75
  manylinux: auto
76
76
  - name: Build free-threaded wheels
77
+ # cp314t free-threaded cross-builds are flaky (QEMU emulation); don't let
78
+ # a free-threaded build failure block the release of the stable wheels.
79
+ continue-on-error: true
77
80
  uses: PyO3/maturin-action@v1
78
81
  with:
79
82
  target: ${{ matrix.platform.target }}
@@ -114,6 +117,9 @@ jobs:
114
117
  sccache: ${{ !startsWith(github.ref, 'refs/tags/') }}
115
118
  manylinux: musllinux_1_2
116
119
  - name: Build free-threaded wheels
120
+ # cp314t free-threaded cross-builds are flaky (QEMU emulation); don't let
121
+ # a free-threaded build failure block the release of the stable wheels.
122
+ continue-on-error: true
117
123
  uses: PyO3/maturin-action@v1
118
124
  with:
119
125
  target: ${{ matrix.platform.target }}
@@ -159,6 +165,9 @@ jobs:
159
165
  python-version: '3.14t'
160
166
  architecture: ${{ matrix.platform.python_arch }}
161
167
  - name: Build free-threaded wheels
168
+ # cp314t free-threaded cross-builds are flaky (QEMU emulation); don't let
169
+ # a free-threaded build failure block the release of the stable wheels.
170
+ continue-on-error: true
162
171
  uses: PyO3/maturin-action@v1
163
172
  with:
164
173
  target: ${{ matrix.platform.target }}
@@ -196,6 +205,9 @@ jobs:
196
205
  with:
197
206
  python-version: '3.14t'
198
207
  - name: Build free-threaded wheels
208
+ # cp314t free-threaded cross-builds are flaky (QEMU emulation); don't let
209
+ # a free-threaded build failure block the release of the stable wheels.
210
+ continue-on-error: true
199
211
  uses: PyO3/maturin-action@v1
200
212
  with:
201
213
  target: ${{ matrix.platform.target }}
@@ -21,6 +21,12 @@ repos:
21
21
  stages: [commit-msg]
22
22
  - repo: local
23
23
  hooks:
24
+ - id: pyrefly
25
+ name: Type-check with pyrefly
26
+ entry: pixi run typecheck
27
+ language: system
28
+ types: [python]
29
+ pass_filenames: false
24
30
  - id: pixi lock
25
31
  name: Lock pixi environment
26
32
  entry: pixi lock
@@ -1,3 +1,20 @@
1
+ ## 0.15.2 (2026-06-13)
2
+
3
+ ### Fix
4
+
5
+ - **tokenize**: accept readonly input on parallel path; guard int32 out=
6
+
7
+ ### Perf
8
+
9
+ - **tokenize**: parallel Numba LUT gather with small-input np.take fast path
10
+ - **tokenize**: use 256-entry LUT gather instead of linear scan
11
+
12
+ ## 0.15.1 (2026-06-08)
13
+
14
+ ### Fix
15
+
16
+ - **rag**: traverse IndexedArray in unbox/_extract_list_offsets
17
+
1
18
  ## 0.15.0 (2026-06-07)
2
19
 
3
20
  ### Feat
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.4
2
2
  Name: seqpro
3
- Version: 0.15.0
3
+ Version: 0.15.2
4
4
  Classifier: Programming Language :: Rust
5
5
  Classifier: Programming Language :: Python :: Implementation :: CPython
6
6
  Classifier: Programming Language :: Python :: Implementation :: PyPy
@@ -0,0 +1,450 @@
1
+ # Tokenize LUT Optimization + CodSpeed Microbenchmarks Implementation Plan
2
+
3
+ > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
4
+
5
+ **Goal:** Replace `tokenize`'s per-character linear-scan kernel with a 256-entry NumPy lookup-table gather, and add `pytest-codspeed` microbenchmarks wired into a CodSpeed CI workflow.
6
+
7
+ **Architecture:** `tokenize` builds a 256-entry `int32` LUT (`unknown_token` fill, scattered with the token map) and tokenizes via `np.take(lut, seqs.view(uint8), out=out)` for both dense and Ragged inputs. `gufunc_tokenize` is left untouched because `decode_tokens` (out of scope) still uses it. Benchmarks live in a `test_*` file collected by the normal suite and timed under `--codspeed`. A precomputed-DNA LUT fast path is added only if a benchmark proves it faster.
8
+
9
+ **Tech Stack:** Python, NumPy, Numba (existing, untouched), pytest, pytest-codspeed, pixi, GitHub Actions, CodSpeed.
10
+
11
+ **Spec:** `docs/superpowers/specs/2026-06-12-tokenize-lut-codspeed-design.md`
12
+
13
+ ---
14
+
15
+ ## File Structure
16
+
17
+ - `python/seqpro/_encoders.py` — modify `tokenize` only (the LUT change). `decode_tokens` and all overloads/signatures unchanged.
18
+ - `python/seqpro/_numba.py` — unchanged (`gufunc_tokenize` stays for `decode_tokens`).
19
+ - `tests/test_tokenize.py` — add an equivalence/characterization test guarding the refactor.
20
+ - `tests/test_bench_tokenize.py` — new; `pytest-codspeed` benchmarks (dense + 3 ragged profiles) + the DNA-fast-path comparison.
21
+ - `pixi.toml` — add `pytest-codspeed` to the `bench` feature and a `bench` task.
22
+ - `.github/workflows/bench.yaml` — new CodSpeed CI workflow.
23
+
24
+ ---
25
+
26
+ ## Task 1: Characterization test locking tokenize output (guards the refactor)
27
+
28
+ This test pins current `tokenize` output to a direct `gufunc_tokenize` reference. It passes against the current implementation AND must keep passing after the LUT change — that is the safety net.
29
+
30
+ **Files:**
31
+ - Test: `tests/test_tokenize.py` (append)
32
+
33
+ - [ ] **Step 1: Write the test**
34
+
35
+ Append to `tests/test_tokenize.py`:
36
+
37
+ ```python
38
+ def test_tokenize_matches_gufunc_reference():
39
+ """LUT output must be byte-for-byte identical to the linear-scan gufunc."""
40
+ from seqpro._numba import gufunc_tokenize
41
+
42
+ token_map = {"A": 0, "C": 1, "G": 2, "T": 3}
43
+ unknown_token = 4
44
+
45
+ def reference(cast_seq):
46
+ source = np.array([c.encode("ascii") for c in token_map]).view(np.uint8)
47
+ target = np.array(list(token_map.values()), dtype=np.int32)
48
+ return gufunc_tokenize(
49
+ cast_seq.view(np.uint8), source, target, np.int32(unknown_token)
50
+ )
51
+
52
+ # Dense 2-D, with known + unknown ("N", "x") characters.
53
+ seqs = ["ACGTN", "TTxAC", "GGGGG"]
54
+ cast = sp.cast_seqs(seqs) # (3, 5) S1
55
+ expected = reference(cast)
56
+ result = sp.tokenize(cast, token_map, unknown_token=unknown_token)
57
+ np.testing.assert_array_equal(result, expected)
58
+ assert result.dtype == np.int32
59
+
60
+ # out= path: result written in place, equals expected, returns same buffer.
61
+ out = np.empty(cast.shape, dtype=np.int32)
62
+ returned = sp.tokenize(cast, token_map, unknown_token=unknown_token, out=out)
63
+ np.testing.assert_array_equal(out, expected)
64
+ np.testing.assert_array_equal(returned, expected)
65
+
66
+ # Ragged path.
67
+ rag_seqs = ["ACGTN", "TTxAC", "GGGGG"]
68
+ data = np.frombuffer("".join(rag_seqs).encode("ascii"), dtype="S1")
69
+ lengths = np.array([len(s) for s in rag_seqs])
70
+ rag = Ragged.from_lengths(data, lengths)
71
+ rag_result = sp.tokenize(rag, token_map, unknown_token=unknown_token)
72
+ flat_expected = reference(np.frombuffer(b"".join(s.encode() for s in rag_seqs), dtype="S1"))
73
+ np.testing.assert_array_equal(rag_result.data, flat_expected)
74
+ np.testing.assert_array_equal(rag_result.lengths.ravel(), lengths)
75
+ ```
76
+
77
+ - [ ] **Step 2: Run the test to verify it passes against the CURRENT implementation**
78
+
79
+ Run: `pixi run -e dev pytest tests/test_tokenize.py::test_tokenize_matches_gufunc_reference -v`
80
+ Expected: PASS (current code already produces this output — this confirms the reference is correct).
81
+
82
+ - [ ] **Step 3: Commit**
83
+
84
+ ```bash
85
+ git add tests/test_tokenize.py
86
+ git commit -m "test: characterize tokenize output against gufunc reference"
87
+ ```
88
+
89
+ ---
90
+
91
+ ## Task 2: Replace tokenize's kernel with a 256-entry LUT gather
92
+
93
+ **Files:**
94
+ - Modify: `python/seqpro/_encoders.py` (function `tokenize`, lines ~239-276 — the body after the docstring only)
95
+ - Test: `tests/test_tokenize.py` (the Task 1 test + existing tests)
96
+
97
+ - [ ] **Step 1: Verify the existing + characterization tests pass (baseline green)**
98
+
99
+ Run: `pixi run -e dev pytest tests/test_tokenize.py -v`
100
+ Expected: PASS (all tokenize tests, including `test_tokenize_matches_gufunc_reference`).
101
+
102
+ - [ ] **Step 2: Rewrite the `tokenize` body**
103
+
104
+ In `python/seqpro/_encoders.py`, replace the body of `tokenize` BELOW the docstring (currently lines ~264-276, starting at `source = np.array(...)`) with:
105
+
106
+ ```python
107
+ # Build a 256-entry lookup table: lut[byte] -> token. Input is uint8 (0-255)
108
+ # after cast_seqs, so a single gather replaces a per-character linear scan.
109
+ keys = np.array([c.encode("ascii") for c in token_map]).view(np.uint8)
110
+ vals = np.array(list(token_map.values()), dtype=np.int32)
111
+ lut = np.full(256, np.int32(unknown_token), dtype=np.int32)
112
+ lut[keys] = vals
113
+
114
+ if isinstance(seqs, Ragged):
115
+ seqs = seqs.to_packed()
116
+ n = len(seqs.lengths.ravel())
117
+ trailing = seqs.data.shape[1:]
118
+ flat = np.take(lut, seqs.data.view(np.uint8))
119
+ return Ragged.from_offsets(flat, (n, None, *trailing), seqs.offsets)
120
+
121
+ _seqs = cast_seqs(seqs)
122
+ return np.take(lut, _seqs.view(np.uint8), out=out)
123
+ ```
124
+
125
+ Do NOT modify the docstring, the `@overload` signatures, or `decode_tokens`. Leave the `gufunc_tokenize` import in place (still used by `decode_tokens`).
126
+
127
+ - [ ] **Step 3: Run the characterization + existing tokenize tests**
128
+
129
+ Run: `pixi run -e dev pytest tests/test_tokenize.py -v`
130
+ Expected: PASS (identical output to the reference; roundtrip and ragged tests still green).
131
+
132
+ - [ ] **Step 4: Run the full suite to catch downstream callers (e.g. transforms, ohe roundtrips)**
133
+
134
+ Run: `pixi run -e dev pytest tests/ -q`
135
+ Expected: PASS (no regressions).
136
+
137
+ - [ ] **Step 5: Commit**
138
+
139
+ ```bash
140
+ git add python/seqpro/_encoders.py
141
+ git commit -m "perf(tokenize): use 256-entry LUT gather instead of linear scan"
142
+ ```
143
+
144
+ ---
145
+
146
+ ## Task 3: Add pytest-codspeed dependency and a bench task
147
+
148
+ **Files:**
149
+ - Modify: `pixi.toml` (`[feature.bench.dependencies]` and `[feature.bench.tasks]`)
150
+
151
+ - [ ] **Step 1: Add the dependency**
152
+
153
+ In `pixi.toml`, under `[feature.bench.dependencies]` (currently `marimo`, `seaborn`, `statsmodels`), add:
154
+
155
+ ```toml
156
+ pytest-codspeed = "*"
157
+ ```
158
+
159
+ - [ ] **Step 2: Add the bench task**
160
+
161
+ In `pixi.toml`, under `[feature.bench.tasks]` (currently has `i-kernel`), add:
162
+
163
+ ```toml
164
+ bench = "pytest tests/test_bench_tokenize.py --codspeed"
165
+ ```
166
+
167
+ - [ ] **Step 3: Install the bench environment**
168
+
169
+ Run: `pixi install -e bench`
170
+ Expected: resolves and installs `pytest-codspeed` (and its `pytest-benchmark`-compatible API) into the `bench` env.
171
+
172
+ - [ ] **Step 4: Commit**
173
+
174
+ ```bash
175
+ git add pixi.toml pixi.lock
176
+ git commit -m "build(bench): add pytest-codspeed and bench task"
177
+ ```
178
+
179
+ ---
180
+
181
+ ## Task 4: Write the tokenize microbenchmarks (dense + 3 ragged profiles)
182
+
183
+ `pytest-codspeed` provides a `benchmark` fixture (pytest-benchmark compatible). Named `test_*` so the default suite collects them — under plain pytest the body runs once (cheap); under `--codspeed` it is instrumented and timed.
184
+
185
+ **Files:**
186
+ - Create: `tests/test_bench_tokenize.py`
187
+
188
+ - [ ] **Step 1: Write the benchmark file**
189
+
190
+ Create `tests/test_bench_tokenize.py`:
191
+
192
+ ```python
193
+ """Microbenchmarks for ``seqpro.tokenize`` (pytest-codspeed).
194
+
195
+ Collected by the normal test suite (runs each body once, ~free) and timed under
196
+ ``pytest --codspeed`` (see the ``bench`` pixi task / bench.yaml CI workflow).
197
+ """
198
+
199
+ from __future__ import annotations
200
+
201
+ import numpy as np
202
+ import pytest
203
+ import seqpro as sp
204
+ from seqpro.rag import Ragged
205
+
206
+ DNA_TOKEN_MAP = {"A": 0, "C": 1, "G": 2, "T": 3}
207
+ UNKNOWN_TOKEN = 4
208
+ _BASES = np.frombuffer(b"ACGT", dtype="S1")
209
+
210
+
211
+ def _rng():
212
+ # Argless default_rng would be nondeterministic; pin the seed.
213
+ return np.random.default_rng(0)
214
+
215
+
216
+ def _dense(batch: int, length: int) -> np.ndarray:
217
+ rng = _rng()
218
+ idx = rng.integers(0, 4, size=(batch, length))
219
+ return _BASES[idx] # (batch, length) S1
220
+
221
+
222
+ def _ragged(n: int, low: int, high: int) -> Ragged:
223
+ rng = _rng()
224
+ lengths = rng.integers(low, high + 1, size=n).astype(np.int64)
225
+ total = int(lengths.sum())
226
+ data = _BASES[rng.integers(0, 4, size=total)]
227
+ return Ragged.from_lengths(data, lengths)
228
+
229
+
230
+ def test_bench_dense_batch(benchmark):
231
+ """Realistic training batch (512, 1024) DNA."""
232
+ seqs = _dense(512, 1024)
233
+ benchmark(lambda: sp.tokenize(seqs, DNA_TOKEN_MAP, unknown_token=UNKNOWN_TOKEN))
234
+
235
+
236
+ def test_bench_ragged_short_alleles(benchmark):
237
+ """Thousands of very short sequences (both alleles)."""
238
+ seqs = _ragged(8000, 1, 4)
239
+ benchmark(lambda: sp.tokenize(seqs, DNA_TOKEN_MAP, unknown_token=UNKNOWN_TOKEN))
240
+
241
+
242
+ def test_bench_ragged_flanked_alleles(benchmark):
243
+ """Thousands of >10 bp sequences (alleles with flank nucleotides)."""
244
+ seqs = _ragged(8000, 11, 60)
245
+ benchmark(lambda: sp.tokenize(seqs, DNA_TOKEN_MAP, unknown_token=UNKNOWN_TOKEN))
246
+
247
+
248
+ def test_bench_ragged_cres(benchmark):
249
+ """Hundreds of 100-200 bp sequences (CREs)."""
250
+ seqs = _ragged(500, 100, 200)
251
+ benchmark(lambda: sp.tokenize(seqs, DNA_TOKEN_MAP, unknown_token=UNKNOWN_TOKEN))
252
+ ```
253
+
254
+ - [ ] **Step 2: Verify the benchmarks are collected and pass as plain tests**
255
+
256
+ Run: `pixi run -e bench pytest tests/test_bench_tokenize.py -v`
257
+ Expected: PASS — 4 tests collected and run (the `benchmark` fixture executes each callable once without `--codspeed`).
258
+
259
+ - [ ] **Step 3: Verify they run under codspeed instrumentation locally**
260
+
261
+ Run: `pixi run -e bench pytest tests/test_bench_tokenize.py --codspeed -v`
262
+ Expected: PASS — codspeed reports timing for the 4 benchmarks (a "running in walltime mode / not in CI" notice is fine locally).
263
+
264
+ - [ ] **Step 4: Confirm the default test suite still passes (benchmarks run as plain tests there too)**
265
+
266
+ Run: `pixi run -e dev pytest tests/test_bench_tokenize.py -q`
267
+ Expected: PASS (the `benchmark` fixture is provided by pytest-codspeed; if the `dev` env lacks it this file is skipped/errors — if so, the bench file should only be collected in the bench env: see note).
268
+
269
+ Note: if `dev` env errors on the missing `benchmark` fixture, add `tests/test_bench_tokenize.py` to a `--ignore` for the default `test` task, or guard the import with `pytest.importorskip("pytest_codspeed")` at module top. Prefer `pytest.importorskip` so the file self-skips cleanly:
270
+
271
+ ```python
272
+ pytest.importorskip("pytest_codspeed")
273
+ ```
274
+
275
+ (Place it directly after the imports.)
276
+
277
+ - [ ] **Step 5: Commit**
278
+
279
+ ```bash
280
+ git add tests/test_bench_tokenize.py
281
+ git commit -m "test(bench): add tokenize microbenchmarks (dense + ragged profiles)"
282
+ ```
283
+
284
+ ---
285
+
286
+ ## Task 5: DNA fast-path comparison + conditional integration
287
+
288
+ Add a benchmark comparing the per-call generic LUT against a **precomputed module-level DNA LUT** (built once at import, reused when `token_map` matches canonical DNA). Integrate the precomputed branch into `tokenize` ONLY if the benchmark shows a real, repeatable win; otherwise document the non-win and stop.
289
+
290
+ **Files:**
291
+ - Modify: `tests/test_bench_tokenize.py` (add comparison benchmarks)
292
+ - Modify (CONDITIONAL on benchmark result): `python/seqpro/_encoders.py`
293
+
294
+ - [ ] **Step 1: Add the comparison benchmarks**
295
+
296
+ Append to `tests/test_bench_tokenize.py`:
297
+
298
+ ```python
299
+ # Candidate DNA fast path: a LUT built once at import, reused across calls,
300
+ # avoiding the per-call np.full(256)+scatter. Compared head-to-head below.
301
+ _DNA_LUT = np.full(256, np.int32(UNKNOWN_TOKEN), dtype=np.int32)
302
+ _DNA_LUT[np.frombuffer(b"ACGT", dtype="S1").view(np.uint8)] = np.arange(4, dtype=np.int32)
303
+
304
+
305
+ def _generic_tokenize(u8: np.ndarray) -> np.ndarray:
306
+ keys = np.array([c.encode("ascii") for c in DNA_TOKEN_MAP]).view(np.uint8)
307
+ vals = np.array(list(DNA_TOKEN_MAP.values()), dtype=np.int32)
308
+ lut = np.full(256, np.int32(UNKNOWN_TOKEN), dtype=np.int32)
309
+ lut[keys] = vals
310
+ return np.take(lut, u8)
311
+
312
+
313
+ def test_bench_dna_generic_lut(benchmark):
314
+ u8 = _dense(512, 1024).view(np.uint8)
315
+ benchmark(lambda: _generic_tokenize(u8))
316
+
317
+
318
+ def test_bench_dna_precomputed_lut(benchmark):
319
+ u8 = _dense(512, 1024).view(np.uint8)
320
+ benchmark(lambda: np.take(_DNA_LUT, u8))
321
+ ```
322
+
323
+ - [ ] **Step 2: Run the comparison and record the numbers**
324
+
325
+ Run: `pixi run -e bench pytest tests/test_bench_tokenize.py -k "dna" --codspeed -v`
326
+ Expected: PASS. Record the two timings. The precomputed LUT only saves the O(256) build; on a (512,1024) batch the gather dominates, so a win is expected to be negligible.
327
+
328
+ - [ ] **Step 3: Decision gate**
329
+
330
+ - If `test_bench_dna_precomputed_lut` is **NOT meaningfully faster** (e.g. <5% and within noise): STOP integration. The generic LUT stands. Add a one-line note to the spec's "DNA-specific fast path" section recording the measured non-win, then commit only the benchmark additions:
331
+
332
+ ```bash
333
+ git add tests/test_bench_tokenize.py docs/superpowers/specs/2026-06-12-tokenize-lut-codspeed-design.md
334
+ git commit -m "test(bench): compare generic vs precomputed DNA LUT (no meaningful win)"
335
+ ```
336
+
337
+ - If it **IS meaningfully faster** (repeatable, well outside noise): proceed to Step 4.
338
+
339
+ - [ ] **Step 4 (CONDITIONAL): Integrate the precomputed DNA LUT**
340
+
341
+ Only if Step 3 found a real win. In `python/seqpro/_encoders.py`, add a module-level constant near the top (after imports):
342
+
343
+ ```python
344
+ # Precomputed LUT for the canonical DNA token map (fast path for tokenize).
345
+ _DNA_TOKEN_MAP = {"A": 0, "C": 1, "G": 2, "T": 3}
346
+ _DNA_LUT = np.full(256, np.int32(4), dtype=np.int32)
347
+ _DNA_LUT[np.frombuffer(b"ACGT", dtype="S1").view(np.uint8)] = np.arange(4, dtype=np.int32)
348
+ ```
349
+
350
+ Then in `tokenize`, replace the LUT-build block with a fast-path check:
351
+
352
+ ```python
353
+ if token_map == _DNA_TOKEN_MAP and unknown_token == 4:
354
+ lut = _DNA_LUT
355
+ else:
356
+ keys = np.array([c.encode("ascii") for c in token_map]).view(np.uint8)
357
+ vals = np.array(list(token_map.values()), dtype=np.int32)
358
+ lut = np.full(256, np.int32(unknown_token), dtype=np.int32)
359
+ lut[keys] = vals
360
+ ```
361
+
362
+ - [ ] **Step 5 (CONDITIONAL): Verify correctness unchanged**
363
+
364
+ Run: `pixi run -e dev pytest tests/test_tokenize.py -v`
365
+ Expected: PASS (`test_tokenize_matches_gufunc_reference` confirms the fast path matches the reference, since it uses the canonical DNA map + unknown_token=4).
366
+
367
+ - [ ] **Step 6 (CONDITIONAL): Commit**
368
+
369
+ ```bash
370
+ git add python/seqpro/_encoders.py tests/test_bench_tokenize.py
371
+ git commit -m "perf(tokenize): precomputed LUT fast path for canonical DNA token map"
372
+ ```
373
+
374
+ ---
375
+
376
+ ## Task 6: CodSpeed CI workflow
377
+
378
+ **Files:**
379
+ - Create: `.github/workflows/bench.yaml`
380
+
381
+ - [ ] **Step 1: Write the workflow**
382
+
383
+ Create `.github/workflows/bench.yaml`:
384
+
385
+ ```yaml
386
+ name: Benchmarks
387
+
388
+ on:
389
+ pull_request:
390
+ branches:
391
+ - main
392
+ workflow_dispatch:
393
+
394
+ jobs:
395
+ benchmarks:
396
+ runs-on: ubuntu-latest
397
+ name: "Run CodSpeed microbenchmarks"
398
+ steps:
399
+ - name: Check out
400
+ uses: actions/checkout@v4
401
+ with:
402
+ fetch-depth: 0
403
+ - name: Setup pixi
404
+ uses: prefix-dev/setup-pixi@v0.9.5
405
+ with:
406
+ pixi-version: v0.67.2
407
+ cache: true
408
+ environments: bench
409
+ locked: false
410
+ - name: Build extension
411
+ run: pixi run -e bench maturin develop
412
+ - name: Run benchmarks
413
+ uses: CodSpeedHQ/action@v3
414
+ with:
415
+ token: ${{ secrets.CODSPEED_TOKEN }}
416
+ run: pixi run -e bench bench
417
+ ```
418
+
419
+ - [ ] **Step 2: Validate the workflow YAML syntax**
420
+
421
+ Run: `python -c "import yaml,sys; yaml.safe_load(open('.github/workflows/bench.yaml')); print('valid')"`
422
+ Expected: prints `valid`.
423
+
424
+ - [ ] **Step 3: Commit**
425
+
426
+ ```bash
427
+ git add .github/workflows/bench.yaml
428
+ git commit -m "ci(bench): add CodSpeed benchmark workflow"
429
+ ```
430
+
431
+ - [ ] **Step 4: Manual handoff note (out of repo — user action)**
432
+
433
+ After merge, the user must: install the **CodSpeed GitHub App** on the repo and add a **`CODSPEED_TOKEN`** repository secret (from codspeed.io). Until then the workflow runs but cannot upload results / comment on PRs. This is not a code step — surface it in the PR description.
434
+
435
+ ---
436
+
437
+ ## Self-Review
438
+
439
+ - **Spec coverage:**
440
+ - LUT optimization of `tokenize` only → Task 2. ✓
441
+ - `gufunc_tokenize`/`decode_tokens` untouched → Task 2 Step 2 explicitly preserves them. ✓
442
+ - Equivalence/characterization test (dense, ragged, `out=`) → Task 1. ✓
443
+ - DNA fast path, benchmarked & gated on a real win → Task 5. ✓
444
+ - Microbenchmarks in `tests/test_bench_tokenize.py` (dense + 3 ragged profiles) → Task 4. ✓
445
+ - `pytest-codspeed` dep + `bench` task in pixi → Task 3. ✓
446
+ - `bench.yaml` CodSpeed CI workflow + maturin build → Task 6. ✓
447
+ - User wires App + `CODSPEED_TOKEN` → Task 6 Step 4. ✓
448
+ - No `SKILL.md` change (signature/behavior unchanged) → confirmed; no task needed. ✓
449
+ - **Placeholder scan:** No TBD/TODO; all code steps contain full code. The only conditional content (Task 5 Steps 4-6) is explicitly gated and fully specified. ✓
450
+ - **Type/name consistency:** `DNA_TOKEN_MAP`/`UNKNOWN_TOKEN` (bench module) vs `_DNA_TOKEN_MAP`/`_DNA_LUT` (encoders module) are intentionally distinct namespaces; the canonical map `{"A":0,"C":1,"G":2,"T":3}` and `unknown_token == 4` are consistent across Tasks 4-5. `np.take(lut, ..., out=out)` signature consistent with the `out` param of `tokenize`. ✓