seqpro 0.21.2__tar.gz → 0.23.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {seqpro-0.21.2 → seqpro-0.23.0}/CHANGELOG.md +32 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/PKG-INFO +1 -1
- seqpro-0.23.0/docs/superpowers/plans/2026-07-27-ragged-string-getitem.md +546 -0
- seqpro-0.23.0/docs/superpowers/specs/2026-07-27-ragged-string-getitem-design.md +203 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/pyproject.toml +1 -1
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/rag/_core.py +340 -156
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/rag/_ingest.py +18 -18
- {seqpro-0.21.2 → seqpro-0.23.0}/skills/seqpro/SKILL.md +9 -4
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_ragged_core.py +95 -6
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_ragged_core_records.py +65 -2
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_ragged_nested_diff.py +1 -1
- seqpro-0.23.0/tests/test_ragged_nested_fixed_axis.py +154 -0
- seqpro-0.23.0/tests/test_ragged_tuple_index.py +203 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/.claude/skills/zensical/SKILL.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/.gitattributes +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/.github/workflows/bench.yaml +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/.github/workflows/bump.yaml +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/.github/workflows/docs.yml +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/.github/workflows/lint.yaml +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/.github/workflows/merge.yaml +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/.github/workflows/publish.yaml +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/.github/workflows/release-pipeline.yaml +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/.github/workflows/release.yaml +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/.github/workflows/test.yaml +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/.gitignore +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/.pre-commit-config.yaml +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/CLAUDE.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/Cargo.lock +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/Cargo.toml +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/LICENSE +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/README.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/benches/bench_tokenize_translate.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/benches/kshuffle.rs +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/benchmarks/bench_ragged_backends.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/benchmarks/bench_ragged_gather.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/benchmarks/bench_to_packed.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/crates/seqpro-core/Cargo.toml +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/crates/seqpro-core/src/lib.rs +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/crates/seqpro-core/src/ragged.rs +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/api/alphabets.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/api/bed.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/api/gtf.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/api/index.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/api/ragged.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/api/types.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/index.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/ragged.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/roadmap/rust-ragged.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/plans/2026-05-04-ragged-record-array.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/plans/2026-05-05-documentation-site.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/plans/2026-05-05-narwhals-coord-schema.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/plans/2026-05-05-ragged-zip-and-record-introspection.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/plans/2026-05-20-kshuffle-optimization.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/plans/2026-05-20-kshuffle-pooled-buffers-and-k2-fast-path.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/plans/2026-05-20-kshuffle-wilson-single-pass.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/plans/2026-05-20-release-pipeline.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/plans/2026-05-28-translate-lut-validation.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/plans/2026-05-31-flat-buffer-to-padded.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/plans/2026-05-31-rag-to-packed.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/plans/2026-06-05-translate-unknown-codon-policy.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/plans/2026-06-12-tokenize-lut-codspeed.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/plans/2026-06-18-rust-tokenize-translate.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/plans/2026-06-19-rust-ragged-core.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/plans/2026-06-20-rust-ragged-records.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/plans/2026-06-20-rust-ragged-spec-c.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/plans/2026-06-21-ragged-throughput-gate.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/plans/2026-06-21-rust-ragged-consumer-audit.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/plans/2026-06-23-ragged-subclass-getitem.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/plans/2026-06-24-ragged-getitem-fastpaths.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/plans/2026-06-25-ragged-string-hashing.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/plans/2026-07-21-ragged-gather-ok.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-05-04-ragged-record-array-design.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-05-05-documentation-site-design.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-05-05-narwhals-coord-schema-design.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-05-05-ragged-zip-and-record-introspection-design.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-05-20-kshuffle-optimization-design.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-05-20-kshuffle-pooled-buffers-and-k2-fast-path-design.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-05-20-kshuffle-wilson-single-pass-design.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-05-20-release-pipeline-design.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-05-28-translate-lut-and-validation-design.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-05-31-flat-buffer-to-padded-design.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-05-31-rag-to-packed-design.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-06-05-translate-unknown-codon-policy-design.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-06-12-tokenize-lut-codspeed-design.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-06-18-rust-tokenize-translate-design.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-06-19-rust-ragged-core-design.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-06-20-rust-ragged-nested-design.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-06-20-rust-ragged-records-design.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-06-21-ragged-throughput-gate-design.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-06-21-rust-ragged-audit-ledger.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-06-21-rust-ragged-consumer-audit-design.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-06-23-ragged-subclass-getitem-design.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-06-24-ragged-getitem-fastpaths-design.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-06-25-ragged-string-hashing-design.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/docs/superpowers/specs/2026-07-21-ragged-gather-ok-design.md +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/meta.yaml +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/pixi.lock +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/pixi.toml +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/__init__.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/_analyzers.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/_cleaners.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/_coords.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/_encoders.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/_modifiers.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/_numba.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/_types.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/_utils.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/alphabets/__init__.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/alphabets/_alphabets.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/bed.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/experimental/_experimental.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/experimental/_visualizers.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/gtf.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/py.typed +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/rag/__init__.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/rag/_ak_interop.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/rag/_layout.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/rag/_ops.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/rag/_types.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/rag/_utils.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/transforms/__init__.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/transforms/augmentation.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/transforms/tmm.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/python/seqpro/xr/__init__.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/scratch_bench_rc.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/scratch_bench_to_padded.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/src/hashing.rs +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/src/kmer_encode.rs +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/src/kshuffle.rs +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/src/kshuffle_ref.rs +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/src/lib.rs +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/src/ragged.rs +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/src/translate.rs +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/_shape_fixtures.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/bed/test_pyranges.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/bed/test_read.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/bed/test_sort.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/bed/test_with_length.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/bench_ragged_getitem.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/bench_translate_lut.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/conftest.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_analyzers.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_bench_tokenize.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_concatenate.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_coords.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_core_ragged_surface.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_encoders.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_ingest.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_modifiers.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_ohe.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_rag_to_packed.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_ragged.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_ragged_hash.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_ragged_nested_consumers.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_ragged_rc.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_ragged_record_indexing.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_ragged_slice_fastpath.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_ragged_subclass.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_ragged_to_padded.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_shape_matrix.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_tokenize.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_transforms.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_translate.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/tests/test_translate_rust.py +0 -0
- {seqpro-0.21.2 → seqpro-0.23.0}/zensical.toml +0 -0
|
@@ -1,3 +1,35 @@
|
|
|
1
|
+
## 0.23.0 (2026-09-26)
|
|
2
|
+
|
|
3
|
+
### BREAKING CHANGE
|
|
4
|
+
|
|
5
|
+
- rag[:, k] and similar tuple keys now follow NumPy and
|
|
6
|
+
awkward semantics; code relying on the old axis-0 behavior must index
|
|
7
|
+
rows directly (rag[k]).
|
|
8
|
+
- Ragged record to_ak() for arrays with one leading axis
|
|
9
|
+
now returns records at the leaf (e.g. 2 * var * {a, b}) instead of one
|
|
10
|
+
record of lists per row.
|
|
11
|
+
|
|
12
|
+
### Fix
|
|
13
|
+
|
|
14
|
+
- **rag**: apply each tuple key to the next output axis
|
|
15
|
+
- **rag**: emit leaf-level records from to_ak() for every rag_dim
|
|
16
|
+
- **rag**: honor fixed axes between the outer axis and nested ragged axes
|
|
17
|
+
|
|
18
|
+
## 0.22.0 (2026-07-27)
|
|
19
|
+
|
|
20
|
+
### BREAKING CHANGE
|
|
21
|
+
|
|
22
|
+
- peeling a record row returns opaque-string fields as a
|
|
23
|
+
Ragged of strings, not a concatenated S1 array.
|
|
24
|
+
- Ragged.__getitem__ with an integer on a string-under-axis
|
|
25
|
+
array returns a Ragged of strings, not one concatenated bytes. Use
|
|
26
|
+
b"".join(s[i]) for the old value.
|
|
27
|
+
|
|
28
|
+
### Fix
|
|
29
|
+
|
|
30
|
+
- **rag**: preserve string boundaries in peeled record rows
|
|
31
|
+
- **rag**: preserve string boundaries when indexing string-under-axis
|
|
32
|
+
|
|
1
33
|
## 0.21.2 (2026-07-22)
|
|
2
34
|
|
|
3
35
|
### Fix
|
|
@@ -0,0 +1,546 @@
|
|
|
1
|
+
# String-under-axis integer indexing (issue #71) Implementation Plan
|
|
2
|
+
|
|
3
|
+
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
|
4
|
+
|
|
5
|
+
**Goal:** Make integer indexing on a string-under-axis `Ragged` return one element per string instead of silently concatenating the whole group into one blob.
|
|
6
|
+
|
|
7
|
+
**Architecture:** Integer indexing peels one real ragged level. For a string-under-axis leaf (`offsets` non-empty **and** `str_offsets` set) the peel lands on the standalone opaque-string layout (`offsets == []`, `str_offsets` set, `shape == (k,)`) — which Spec C already defines as the zero-real-level special case of the same layout. Two call sites construct this: the plain `__getitem__` integer branch and the record-row integer branch. Both mirror the narrowing that `_slice_contig_string` already performs for slices.
|
|
8
|
+
|
|
9
|
+
**Tech Stack:** Python 3.9+, NumPy, pytest. Pure Python layer — no Rust (`src/`, `crates/`) changes.
|
|
10
|
+
|
|
11
|
+
## Global Constraints
|
|
12
|
+
|
|
13
|
+
- Spec: `docs/superpowers/specs/2026-07-27-ragged-string-getitem-design.md`.
|
|
14
|
+
- Data must stay **zero-copy**: the result's `data` is a view of the parent's buffer. Only the small `(k+1,)` offsets slice is copied (to rebase to zero, which `is_contiguous` requires for opaque strings — `python/seqpro/rag/_core.py:319-321`).
|
|
15
|
+
- The **standalone/flat** opaque-string case (`offsets == []`) is unchanged: `flat[0]` still returns `bytes`. The `self._layout.offsets` guard distinguishes it.
|
|
16
|
+
- No Python loops over elements — this is a per-batch-item accessor (repo rule: "No naive NumPy in hot paths").
|
|
17
|
+
- Breaking change. `major_version_zero = true` in `pyproject.toml:64`, so a `!` conventional commit on 0.x produces a **minor** bump (0.21.2 → 0.22.0), which is what the spec calls for. Do **not** hand-edit `CHANGELOG.md` — commitizen generates it on bump.
|
|
18
|
+
- Test command (run from the worktree root; the worktree has no `.pixi` of its own, so borrow the main checkout's dev env and point `PYTHONPATH` at the worktree source):
|
|
19
|
+
|
|
20
|
+
```bash
|
|
21
|
+
PYTHONPATH=python /carter/users/dlaub/projects/ML4GLand/SeqPro/.pixi/envs/dev/bin/python -m pytest tests/test_ragged_core.py -q
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
## File Structure
|
|
25
|
+
|
|
26
|
+
| File | Responsibility | Change |
|
|
27
|
+
|---|---|---|
|
|
28
|
+
| `python/seqpro/rag/_core.py:722-731` | `Ragged.__getitem__` integer branch — plain string-under-axis | Modify (Task 1) |
|
|
29
|
+
| `python/seqpro/rag/_core.py:1216-1226` | `_getitem_record_rows` integer branch — opaque-string record fields | Modify (Task 2) |
|
|
30
|
+
| `tests/test_ragged_core.py` | Existing home of the string-under-axis test section (`test_string_under_axis_integer_index` at line 733) | Modify + add (Tasks 1, 2) |
|
|
31
|
+
| `skills/seqpro/SKILL.md` | Public API skill doc; CLAUDE.md requires an update for any breaking change | Modify (Task 3) |
|
|
32
|
+
|
|
33
|
+
`_getitem_record_rows_r2` (`_core.py:1243`) delegates to `Ragged(fl)[where]` and inherits Task 1's fix — no separate change, but Task 2 covers it with a test.
|
|
34
|
+
|
|
35
|
+
---
|
|
36
|
+
|
|
37
|
+
### Task 1: Plain string-under-axis integer indexing
|
|
38
|
+
|
|
39
|
+
**Files:**
|
|
40
|
+
- Modify: `python/seqpro/rag/_core.py:722-731`
|
|
41
|
+
- Test: `tests/test_ragged_core.py` (replace `test_string_under_axis_integer_index` at line 733; add new tests after it)
|
|
42
|
+
|
|
43
|
+
**Interfaces:**
|
|
44
|
+
- Consumes: `RaggedLayout(data=..., offsets=..., shape=..., str_offsets=...)` from `python/seqpro/rag/_layout.py` (already imported at `_core.py:10`).
|
|
45
|
+
- Produces: `Ragged.__getitem__(int)` on a string-under-axis `Ragged` returns `Ragged` with `offsets == []`, `str_offsets` set, `shape == (k,)`, `is_string is True`. Task 2 relies on this same construction shape.
|
|
46
|
+
|
|
47
|
+
- [ ] **Step 1: Replace the test that pins the old behavior**
|
|
48
|
+
|
|
49
|
+
`tests/test_ragged_core.py:732-741` currently reads:
|
|
50
|
+
|
|
51
|
+
```python
|
|
52
|
+
def test_string_under_axis_integer_index():
|
|
53
|
+
rag = Ragged.from_offsets(
|
|
54
|
+
np.frombuffer(b"TTGG", "S1"),
|
|
55
|
+
(2, None),
|
|
56
|
+
np.array([0, 1, 2]),
|
|
57
|
+
str_offsets=np.array([0, 2, 4]),
|
|
58
|
+
)
|
|
59
|
+
assert rag[0] == b"TT"
|
|
60
|
+
assert rag[1] == b"GG"
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
It uses exactly one string per group, so it can never distinguish concatenation from per-string indexing — that is why the bug survived. Replace it with:
|
|
64
|
+
|
|
65
|
+
```python
|
|
66
|
+
def test_string_under_axis_integer_index():
|
|
67
|
+
"""Peeling one group yields the standalone opaque-string layout (Spec C)."""
|
|
68
|
+
rag = Ragged.from_offsets(
|
|
69
|
+
np.frombuffer(b"TTGG", "S1"),
|
|
70
|
+
(2, None),
|
|
71
|
+
np.array([0, 1, 2]),
|
|
72
|
+
str_offsets=np.array([0, 2, 4]),
|
|
73
|
+
)
|
|
74
|
+
row = rag[0]
|
|
75
|
+
assert isinstance(row, Ragged)
|
|
76
|
+
assert row.is_string and row.shape == (1,)
|
|
77
|
+
assert len(row) == 1
|
|
78
|
+
assert row[0] == b"TT"
|
|
79
|
+
assert rag[1][0] == b"GG"
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
- [ ] **Step 2: Add the failing tests from the issue**
|
|
83
|
+
|
|
84
|
+
Append immediately after the test above:
|
|
85
|
+
|
|
86
|
+
```python
|
|
87
|
+
def _issue71_pair():
|
|
88
|
+
"""String-under-axis and numeric Ragged sharing one offsets object.
|
|
89
|
+
|
|
90
|
+
Groups: 0 -> ('A', 'GG'), 1 -> ('TC',).
|
|
91
|
+
"""
|
|
92
|
+
data = np.frombuffer(b"AGGTC", dtype="S1")
|
|
93
|
+
outer = np.array([0, 2, 3], dtype=OFFSET_TYPE) # group -> string index
|
|
94
|
+
inner = np.array([0, 1, 3, 5], dtype=OFFSET_TYPE) # string -> byte index
|
|
95
|
+
s = Ragged.from_offsets(data, (2, None), outer, str_offsets=inner)
|
|
96
|
+
n = Ragged.from_offsets(np.array([10, 20, 30], dtype=np.int32), (2, None), outer)
|
|
97
|
+
return s, n
|
|
98
|
+
|
|
99
|
+
|
|
100
|
+
def test_string_under_axis_index_preserves_boundaries():
|
|
101
|
+
"""Issue #71: interior str_offsets boundaries must survive an integer index."""
|
|
102
|
+
s, _ = _issue71_pair()
|
|
103
|
+
row = s[0]
|
|
104
|
+
assert len(row) == 2
|
|
105
|
+
assert row[0] == b"A"
|
|
106
|
+
assert row[1] == b"GG"
|
|
107
|
+
assert list(s[1]) == [b"TC"]
|
|
108
|
+
|
|
109
|
+
|
|
110
|
+
def test_string_under_axis_index_matches_lengths():
|
|
111
|
+
"""len(s[i]) must equal s.lengths[i] and the numeric row length."""
|
|
112
|
+
s, n = _issue71_pair()
|
|
113
|
+
for i in range(len(s)):
|
|
114
|
+
assert len(s[i]) == int(s.lengths[i])
|
|
115
|
+
assert len(s[i]) == len(n[i])
|
|
116
|
+
|
|
117
|
+
|
|
118
|
+
def test_string_under_axis_index_is_zero_copy():
|
|
119
|
+
s, _ = _issue71_pair()
|
|
120
|
+
assert np.shares_memory(s[0].data, s.data)
|
|
121
|
+
|
|
122
|
+
|
|
123
|
+
def test_string_under_axis_index_empty_group():
|
|
124
|
+
"""A group holding zero strings peels to a length-0 result."""
|
|
125
|
+
rag = Ragged.from_offsets(
|
|
126
|
+
np.frombuffer(b"AC", "S1"),
|
|
127
|
+
(2, None),
|
|
128
|
+
np.array([0, 0, 2], dtype=OFFSET_TYPE), # group 0 empty
|
|
129
|
+
str_offsets=np.array([0, 1, 2], dtype=OFFSET_TYPE),
|
|
130
|
+
)
|
|
131
|
+
assert len(rag[0]) == 0
|
|
132
|
+
assert len(rag[1]) == 2
|
|
133
|
+
|
|
134
|
+
|
|
135
|
+
def test_string_under_axis_index_agrees_with_to_chars():
|
|
136
|
+
s, _ = _issue71_pair()
|
|
137
|
+
chars = s.to_chars()
|
|
138
|
+
for i in range(len(s)):
|
|
139
|
+
for j in range(len(s[i])):
|
|
140
|
+
assert s[i][j] == chars[i][j].tobytes()
|
|
141
|
+
|
|
142
|
+
|
|
143
|
+
def test_string_under_axis_index_negative_and_oob():
|
|
144
|
+
s, _ = _issue71_pair()
|
|
145
|
+
assert list(s[-1]) == [b"TC"]
|
|
146
|
+
with pytest.raises(IndexError):
|
|
147
|
+
s[5]
|
|
148
|
+
|
|
149
|
+
|
|
150
|
+
def test_string_under_axis_index_multidim():
|
|
151
|
+
"""(batch, ploidy, ~variants): reaches the flat branch via _getitem_multidim."""
|
|
152
|
+
data = np.frombuffer(b"AGGTCNNAC", dtype="S1")
|
|
153
|
+
o0 = np.array([0, 2, 3, 4, 6], dtype=OFFSET_TYPE) # 4 segments -> string idx
|
|
154
|
+
i0 = np.array([0, 1, 3, 5, 7, 9], dtype=OFFSET_TYPE) # 6 boundaries -> bytes
|
|
155
|
+
rag = Ragged.from_offsets(data, (2, 2, None), o0, str_offsets=i0)
|
|
156
|
+
row = rag[0] # -> (2, None) string-under-axis
|
|
157
|
+
assert list(row[0]) == [b"A", b"GG"]
|
|
158
|
+
assert list(row[1]) == [b"TC"]
|
|
159
|
+
|
|
160
|
+
|
|
161
|
+
def test_standalone_string_index_still_returns_bytes():
|
|
162
|
+
"""Regression guard: the zero-real-level case is the terminal peel."""
|
|
163
|
+
flat = Ragged.from_offsets(
|
|
164
|
+
np.frombuffer(b"cathithere", "S1"), (3,), np.array([0, 3, 5, 10])
|
|
165
|
+
)
|
|
166
|
+
assert flat[0] == b"cat"
|
|
167
|
+
assert flat[-1] == b"there"
|
|
168
|
+
assert list(flat) == [b"cat", b"hi", b"there"]
|
|
169
|
+
```
|
|
170
|
+
|
|
171
|
+
- [ ] **Step 3: Run the tests to verify they fail**
|
|
172
|
+
|
|
173
|
+
```bash
|
|
174
|
+
PYTHONPATH=python /carter/users/dlaub/projects/ML4GLand/SeqPro/.pixi/envs/dev/bin/python -m pytest tests/test_ragged_core.py -q -k "string_under_axis or standalone_string"
|
|
175
|
+
```
|
|
176
|
+
|
|
177
|
+
Expected: `test_string_under_axis_index_preserves_boundaries` fails (`len(row)` raises `TypeError`/returns 3 rather than 2 — the current return is `bytes`), `test_string_under_axis_integer_index` fails on `isinstance(row, Ragged)`, and the other new per-string tests fail. `test_standalone_string_index_still_returns_bytes` must PASS already.
|
|
178
|
+
|
|
179
|
+
- [ ] **Step 4: Implement the fix**
|
|
180
|
+
|
|
181
|
+
In `python/seqpro/rag/_core.py`, replace the integer branch at lines 722-731:
|
|
182
|
+
|
|
183
|
+
```python
|
|
184
|
+
if isinstance(where, (int, np.integer)):
|
|
185
|
+
lo, hi = int(starts[where]), int(stops[where])
|
|
186
|
+
if self._rl.str_offsets is not None and self._layout.offsets:
|
|
187
|
+
# string-under-axis: outer offsets index variants -> map to bytes via str_offsets
|
|
188
|
+
so = self._rl.str_offsets
|
|
189
|
+
return self._rl.data[int(so[lo]) : int(so[hi])].tobytes()
|
|
190
|
+
row = self._rl.data[lo:hi]
|
|
191
|
+
if self._rl.is_string:
|
|
192
|
+
return row.tobytes()
|
|
193
|
+
return row
|
|
194
|
+
```
|
|
195
|
+
|
|
196
|
+
with:
|
|
197
|
+
|
|
198
|
+
```python
|
|
199
|
+
if isinstance(where, (int, np.integer)):
|
|
200
|
+
lo, hi = int(starts[where]), int(stops[where])
|
|
201
|
+
if self._rl.str_offsets is not None and self._layout.offsets:
|
|
202
|
+
# string-under-axis: peel the real level -> standalone opaque
|
|
203
|
+
# string (k,), preserving the per-string boundaries that live in
|
|
204
|
+
# str_offsets. Concatenating here would drop them (issue #71).
|
|
205
|
+
return Ragged(_peel_string_row(self._rl, lo, hi))
|
|
206
|
+
row = self._rl.data[lo:hi]
|
|
207
|
+
if self._rl.is_string:
|
|
208
|
+
return row.tobytes()
|
|
209
|
+
return row
|
|
210
|
+
```
|
|
211
|
+
|
|
212
|
+
Then add this module-level helper next to the other layout helpers — place it immediately above `class Ragged` in `python/seqpro/rag/_core.py` (Task 2 reuses it, which is why it is a free function rather than a method: `_getitem_record_rows` operates on per-field `RaggedLayout`s, not on `self`):
|
|
213
|
+
|
|
214
|
+
```python
|
|
215
|
+
def _peel_string_row(
|
|
216
|
+
rl: "RaggedLayout[Any]", lo: int, hi: int
|
|
217
|
+
) -> "RaggedLayout[Any]":
|
|
218
|
+
"""Peel strings ``[lo, hi)`` off a string-under-axis leaf.
|
|
219
|
+
|
|
220
|
+
Returns the standalone opaque-string layout (``offsets == []``,
|
|
221
|
+
``str_offsets`` set, ``shape == (hi - lo,)``) — the zero-real-level special
|
|
222
|
+
case of string-under-axis (Spec C Section 2). The data buffer is a view;
|
|
223
|
+
only the ``(k + 1,)`` offsets slice is copied, rebased to zero as
|
|
224
|
+
``is_contiguous`` requires.
|
|
225
|
+
"""
|
|
226
|
+
so = rl.str_offsets
|
|
227
|
+
assert so is not None # caller guarantees a string leaf
|
|
228
|
+
b0 = int(so[lo])
|
|
229
|
+
return RaggedLayout(
|
|
230
|
+
data=rl.data[b0 : int(so[hi])],
|
|
231
|
+
offsets=[],
|
|
232
|
+
shape=(hi - lo,),
|
|
233
|
+
str_offsets=so[lo : hi + 1] - b0,
|
|
234
|
+
)
|
|
235
|
+
```
|
|
236
|
+
|
|
237
|
+
Note `lo`/`hi` are indices in **string** space, so `so[lo : hi + 1]` is the right slice whether `offsets[0]` is 1-D canonical or a lazy `(2, M)` gather layout — `_starts_stops()` normalizes both.
|
|
238
|
+
|
|
239
|
+
- [ ] **Step 5: Run the tests to verify they pass**
|
|
240
|
+
|
|
241
|
+
```bash
|
|
242
|
+
PYTHONPATH=python /carter/users/dlaub/projects/ML4GLand/SeqPro/.pixi/envs/dev/bin/python -m pytest tests/test_ragged_core.py -q -k "string_under_axis or standalone_string"
|
|
243
|
+
```
|
|
244
|
+
|
|
245
|
+
Expected: all PASS.
|
|
246
|
+
|
|
247
|
+
- [ ] **Step 6: Run the full ragged suite for regressions**
|
|
248
|
+
|
|
249
|
+
```bash
|
|
250
|
+
PYTHONPATH=python /carter/users/dlaub/projects/ML4GLand/SeqPro/.pixi/envs/dev/bin/python -m pytest tests/ -q
|
|
251
|
+
```
|
|
252
|
+
|
|
253
|
+
Expected: no new failures. If another test asserts the old concatenating behavior, judge it the same way as `test_string_under_axis_integer_index` — update it to index one level deeper, and note it in the commit body.
|
|
254
|
+
|
|
255
|
+
- [ ] **Step 7: Commit**
|
|
256
|
+
|
|
257
|
+
```bash
|
|
258
|
+
git add python/seqpro/rag/_core.py tests/test_ragged_core.py
|
|
259
|
+
git commit -m "fix(rag)!: preserve string boundaries when indexing string-under-axis
|
|
260
|
+
|
|
261
|
+
Integer indexing on a string-under-axis Ragged concatenated the whole
|
|
262
|
+
group into one bytes, dropping the per-string boundaries already held in
|
|
263
|
+
str_offsets. It now peels to the standalone opaque-string layout (k,),
|
|
264
|
+
so len(s[i]) == s.lengths[i] and s[i][j] is one string.
|
|
265
|
+
|
|
266
|
+
BREAKING CHANGE: Ragged.__getitem__ with an integer on a string-under-axis
|
|
267
|
+
array returns a Ragged of strings, not one concatenated bytes. Use
|
|
268
|
+
b\"\".join(s[i]) for the old value.
|
|
269
|
+
|
|
270
|
+
Refs #71"
|
|
271
|
+
```
|
|
272
|
+
|
|
273
|
+
---
|
|
274
|
+
|
|
275
|
+
### Task 2: Record-layout string fields
|
|
276
|
+
|
|
277
|
+
**Files:**
|
|
278
|
+
- Modify: `python/seqpro/rag/_core.py:1216-1226`
|
|
279
|
+
- Test: `tests/test_ragged_core_records.py`
|
|
280
|
+
|
|
281
|
+
**Interfaces:**
|
|
282
|
+
- Consumes: `_peel_string_row(rl, lo, hi) -> RaggedLayout` from Task 1.
|
|
283
|
+
- Produces: `_getitem_record_rows` integer branch returns `dict[str, NDArray | Ragged]` where opaque-string fields are `Ragged` (standalone opaque-string layout) and numeric/char fields stay `ndarray`; every entry has length `hi - lo`.
|
|
284
|
+
|
|
285
|
+
- [ ] **Step 1: Write the failing tests**
|
|
286
|
+
|
|
287
|
+
Append to `tests/test_ragged_core_records.py`:
|
|
288
|
+
|
|
289
|
+
```python
|
|
290
|
+
def _issue71_record():
|
|
291
|
+
"""Record with an opaque-string field and a numeric field sharing offsets.
|
|
292
|
+
|
|
293
|
+
Groups: 0 -> ('A', 'GG') / starts (1, 2), 1 -> ('TC',) / starts (3,).
|
|
294
|
+
Group 0's second allele is multi-byte, so a 1-byte coincidence cannot
|
|
295
|
+
mask a regression.
|
|
296
|
+
"""
|
|
297
|
+
outer = np.array([0, 2, 3], dtype=OFFSET_TYPE)
|
|
298
|
+
alt = Ragged.from_offsets(
|
|
299
|
+
np.frombuffer(b"AGGTC", dtype="S1"),
|
|
300
|
+
(2, None),
|
|
301
|
+
outer,
|
|
302
|
+
str_offsets=np.array([0, 1, 3, 5], dtype=OFFSET_TYPE),
|
|
303
|
+
)
|
|
304
|
+
start = Ragged.from_offsets(np.array([1, 2, 3], dtype=np.int32), (2, None), outer)
|
|
305
|
+
return Ragged.from_fields({"alt": alt, "start": start})
|
|
306
|
+
|
|
307
|
+
|
|
308
|
+
def test_record_row_string_field_preserves_boundaries():
|
|
309
|
+
"""Issue #71: a string field peeled from a record row keeps its boundaries."""
|
|
310
|
+
row = _issue71_record()[0]
|
|
311
|
+
assert isinstance(row, dict)
|
|
312
|
+
assert list(row["alt"]) == [b"A", b"GG"]
|
|
313
|
+
np.testing.assert_array_equal(row["start"], np.array([1, 2], dtype=np.int32))
|
|
314
|
+
|
|
315
|
+
|
|
316
|
+
def test_record_row_fields_have_matching_lengths():
|
|
317
|
+
"""Every field of a peeled row must have the same length, so zip aligns."""
|
|
318
|
+
rec = _issue71_record()
|
|
319
|
+
for i in range(len(rec)):
|
|
320
|
+
row = rec[i]
|
|
321
|
+
assert len(row["alt"]) == len(row["start"])
|
|
322
|
+
row0 = rec[0]
|
|
323
|
+
assert list(zip(row0["start"], row0["alt"])) == [(1, b"A"), (2, b"GG")]
|
|
324
|
+
|
|
325
|
+
|
|
326
|
+
def test_record_row_string_field_is_zero_copy():
|
|
327
|
+
rec = _issue71_record()
|
|
328
|
+
assert np.shares_memory(rec[0]["alt"].data, rec["alt"].data)
|
|
329
|
+
|
|
330
|
+
|
|
331
|
+
def test_record_multidim_row_string_field_preserves_boundaries():
|
|
332
|
+
"""(batch, ploidy, ~variants) record: rec[0][h] routes via _getitem_record_rows_r2."""
|
|
333
|
+
outer = np.array([0, 2, 3, 4, 6], dtype=OFFSET_TYPE)
|
|
334
|
+
alt = Ragged.from_offsets(
|
|
335
|
+
np.frombuffer(b"AGGTCNNAC", dtype="S1"),
|
|
336
|
+
(2, 2, None),
|
|
337
|
+
outer,
|
|
338
|
+
str_offsets=np.array([0, 1, 3, 5, 7, 9], dtype=OFFSET_TYPE),
|
|
339
|
+
)
|
|
340
|
+
start = Ragged.from_offsets(
|
|
341
|
+
np.arange(6, dtype=np.int32), (2, 2, None), outer
|
|
342
|
+
)
|
|
343
|
+
rec = Ragged.from_fields({"alt": alt, "start": start})
|
|
344
|
+
row = rec[0][0]
|
|
345
|
+
assert list(row["alt"]) == [b"A", b"GG"]
|
|
346
|
+
assert len(row["alt"]) == len(row["start"])
|
|
347
|
+
```
|
|
348
|
+
|
|
349
|
+
- [ ] **Step 2: Run the tests to verify they fail**
|
|
350
|
+
|
|
351
|
+
```bash
|
|
352
|
+
PYTHONPATH=python /carter/users/dlaub/projects/ML4GLand/SeqPro/.pixi/envs/dev/bin/python -m pytest tests/test_ragged_core_records.py -q -k "issue71 or record_row or record_multidim"
|
|
353
|
+
```
|
|
354
|
+
|
|
355
|
+
Expected: FAIL. Today `rec[0]["alt"]` is `array([b'A', b'G', b'G'], dtype='|S1')` — a 3-element `S1` char array next to a 2-element numeric array.
|
|
356
|
+
|
|
357
|
+
- [ ] **Step 3: Implement the fix**
|
|
358
|
+
|
|
359
|
+
In `python/seqpro/rag/_core.py`, replace the integer branch at lines 1216-1226:
|
|
360
|
+
|
|
361
|
+
```python
|
|
362
|
+
if isinstance(where, (int, np.integer)):
|
|
363
|
+
lo, hi = int(starts[where]), int(stops[where])
|
|
364
|
+
out: dict[str, Any] = {}
|
|
365
|
+
for name, fl in rec.fields.items():
|
|
366
|
+
if fl.str_offsets is not None:
|
|
367
|
+
so = fl.str_offsets
|
|
368
|
+
row = fl.data[int(so[lo]) : int(so[hi])]
|
|
369
|
+
else:
|
|
370
|
+
row = fl.data[lo:hi]
|
|
371
|
+
out[name] = row
|
|
372
|
+
return out
|
|
373
|
+
```
|
|
374
|
+
|
|
375
|
+
with:
|
|
376
|
+
|
|
377
|
+
```python
|
|
378
|
+
if isinstance(where, (int, np.integer)):
|
|
379
|
+
lo, hi = int(starts[where]), int(stops[where])
|
|
380
|
+
out: dict[str, Any] = {}
|
|
381
|
+
for name, fl in rec.fields.items():
|
|
382
|
+
if fl.str_offsets is not None:
|
|
383
|
+
# Each field carries its own str_offsets (Spec C Section 5);
|
|
384
|
+
# peel it against the shared lo/hi so every field of the row
|
|
385
|
+
# has the same length (issue #71).
|
|
386
|
+
out[name] = Ragged(_peel_string_row(fl, lo, hi))
|
|
387
|
+
else:
|
|
388
|
+
out[name] = fl.data[lo:hi]
|
|
389
|
+
return out
|
|
390
|
+
```
|
|
391
|
+
|
|
392
|
+
- [ ] **Step 4: Run the tests to verify they pass**
|
|
393
|
+
|
|
394
|
+
```bash
|
|
395
|
+
PYTHONPATH=python /carter/users/dlaub/projects/ML4GLand/SeqPro/.pixi/envs/dev/bin/python -m pytest tests/test_ragged_core_records.py -q -k "issue71 or record_row or record_multidim"
|
|
396
|
+
```
|
|
397
|
+
|
|
398
|
+
Expected: all PASS.
|
|
399
|
+
|
|
400
|
+
- [ ] **Step 5: Run the full suite**
|
|
401
|
+
|
|
402
|
+
```bash
|
|
403
|
+
PYTHONPATH=python /carter/users/dlaub/projects/ML4GLand/SeqPro/.pixi/envs/dev/bin/python -m pytest tests/ -q
|
|
404
|
+
```
|
|
405
|
+
|
|
406
|
+
Expected: no new failures.
|
|
407
|
+
|
|
408
|
+
- [ ] **Step 6: Commit**
|
|
409
|
+
|
|
410
|
+
```bash
|
|
411
|
+
git add python/seqpro/rag/_core.py tests/test_ragged_core_records.py
|
|
412
|
+
git commit -m "fix(rag)!: preserve string boundaries in peeled record rows
|
|
413
|
+
|
|
414
|
+
_getitem_record_rows returned a string field as the raw concatenated S1
|
|
415
|
+
buffer, so a peeled row mixed a 3-char array with a 2-element numeric
|
|
416
|
+
array. String fields now peel to the standalone opaque-string layout,
|
|
417
|
+
giving every field of the row the same length.
|
|
418
|
+
|
|
419
|
+
BREAKING CHANGE: peeling a record row returns opaque-string fields as a
|
|
420
|
+
Ragged of strings, not a concatenated S1 array.
|
|
421
|
+
|
|
422
|
+
Refs #71"
|
|
423
|
+
```
|
|
424
|
+
|
|
425
|
+
---
|
|
426
|
+
|
|
427
|
+
### Task 3: Lint, typecheck, and skill docs
|
|
428
|
+
|
|
429
|
+
**Files:**
|
|
430
|
+
- Modify: `skills/seqpro/SKILL.md`
|
|
431
|
+
|
|
432
|
+
**Interfaces:**
|
|
433
|
+
- Consumes: the public behavior established in Tasks 1 and 2.
|
|
434
|
+
- Produces: nothing other tasks depend on.
|
|
435
|
+
|
|
436
|
+
- [ ] **Step 1: Document the behavior in the "do this, not that" table**
|
|
437
|
+
|
|
438
|
+
In `skills/seqpro/SKILL.md`, add this row to the `### Working with `Ragged` — do this, not that` table (the table starting at line ~92), after the `Count top-level rows` row:
|
|
439
|
+
|
|
440
|
+
```markdown
|
|
441
|
+
| Index one group of an opaque-string `Ragged` | `rag[i]` → `Ragged` of `bytes`, one per string (`len(rag[i]) == rag.lengths[i]`); `rag[i][j]` is one `bytes` | `b"".join(rag[i])`-style concatenation — that was the pre-0.22 behavior and it dropped the per-string boundaries |
|
|
442
|
+
```
|
|
443
|
+
|
|
444
|
+
- [ ] **Step 2: Document the layout rule in the record section**
|
|
445
|
+
|
|
446
|
+
In `skills/seqpro/SKILL.md`, append to the bullet list at the end of the `### Record-layout `Ragged` (multi-field)` section (after the `view` and `apply` bullet):
|
|
447
|
+
|
|
448
|
+
```markdown
|
|
449
|
+
- Peeling a row (`rag[i]` where `i` is an integer) returns a **dict** whose entries all have the same length: numeric/char fields as `ndarray`, opaque-string fields as a `Ragged` of `bytes`. This is what makes `zip(row["start"], row["alt"])` correct.
|
|
450
|
+
```
|
|
451
|
+
|
|
452
|
+
- [ ] **Step 3: Run lint and typecheck**
|
|
453
|
+
|
|
454
|
+
```bash
|
|
455
|
+
cd /carter/users/dlaub/projects/ML4GLand/SeqPro/.claude/worktrees/issue-71-string-under-axis-getitem
|
|
456
|
+
PY=/carter/users/dlaub/projects/ML4GLand/SeqPro/.pixi/envs/dev/bin
|
|
457
|
+
$PY/ruff check python/ tests/ && $PY/ruff format --check python/ tests/
|
|
458
|
+
$PY/pyrefly check python
|
|
459
|
+
```
|
|
460
|
+
|
|
461
|
+
Expected: clean. Fix anything reported before committing.
|
|
462
|
+
|
|
463
|
+
- [ ] **Step 4: Run the full suite one last time**
|
|
464
|
+
|
|
465
|
+
```bash
|
|
466
|
+
PYTHONPATH=python /carter/users/dlaub/projects/ML4GLand/SeqPro/.pixi/envs/dev/bin/python -m pytest tests/ -q
|
|
467
|
+
```
|
|
468
|
+
|
|
469
|
+
Expected: all pass.
|
|
470
|
+
|
|
471
|
+
- [ ] **Step 5: Commit**
|
|
472
|
+
|
|
473
|
+
```bash
|
|
474
|
+
git add skills/seqpro/SKILL.md
|
|
475
|
+
git commit -m "docs(skill): document string-under-axis integer indexing
|
|
476
|
+
|
|
477
|
+
Refs #71"
|
|
478
|
+
```
|
|
479
|
+
|
|
480
|
+
- [ ] **Step 6: Push and open a draft PR**
|
|
481
|
+
|
|
482
|
+
```bash
|
|
483
|
+
git push -u origin worktree-issue-71-string-under-axis-getitem
|
|
484
|
+
gh pr create --draft --title "fix(rag)!: preserve string boundaries when indexing a string-under-axis Ragged" --body "$(cat <<'EOF'
|
|
485
|
+
Closes #71.
|
|
486
|
+
|
|
487
|
+
Integer indexing on a string-under-axis `Ragged` concatenated a whole group
|
|
488
|
+
into one blob, dropping the per-string boundaries already sitting in
|
|
489
|
+
`str_offsets`. `s.lengths[0] == 2` but `len(s[0]) == 3`.
|
|
490
|
+
|
|
491
|
+
Two sites had the defect:
|
|
492
|
+
|
|
493
|
+
- `Ragged.__getitem__` integer branch — the one reported.
|
|
494
|
+
- `_getitem_record_rows` integer branch — worse, and the one GenVarLoader
|
|
495
|
+
hits: a peeled row mixed a concatenated `S1` char array (not even `bytes`)
|
|
496
|
+
with a per-variant numeric array in the same dict.
|
|
497
|
+
|
|
498
|
+
Both now peel to the standalone opaque-string layout (`offsets == []`,
|
|
499
|
+
`str_offsets` set, `shape == (k,)`), which Spec C already defines as the
|
|
500
|
+
zero-real-level special case of string-under-axis. Zero-copy on data; only
|
|
501
|
+
the small offsets slice is copied.
|
|
502
|
+
|
|
503
|
+
This makes `zip(rv.start[0][h], rv.alt[0][h])` correct in GenVarLoader
|
|
504
|
+
(mcvickerlab/GenVarLoader#330).
|
|
505
|
+
|
|
506
|
+
**Breaking:** the integer index now returns a `Ragged` of strings rather than
|
|
507
|
+
one concatenated `bytes`. `b"".join(s[i])` recovers the old value.
|
|
508
|
+
`major_version_zero` is set, so this bumps 0.21.2 → 0.22.0.
|
|
509
|
+
|
|
510
|
+
The existing `test_string_under_axis_integer_index` pinned the old behavior
|
|
511
|
+
but used one string per group, so it could never distinguish the two
|
|
512
|
+
interpretations — that is why the bug survived. It has been rewritten.
|
|
513
|
+
|
|
514
|
+
Spec: `docs/superpowers/specs/2026-07-27-ragged-string-getitem-design.md`
|
|
515
|
+
Plan: `docs/superpowers/plans/2026-07-27-ragged-string-getitem.md`
|
|
516
|
+
EOF
|
|
517
|
+
)"
|
|
518
|
+
```
|
|
519
|
+
|
|
520
|
+
---
|
|
521
|
+
|
|
522
|
+
## Self-Review
|
|
523
|
+
|
|
524
|
+
**Spec coverage:**
|
|
525
|
+
|
|
526
|
+
| Spec item | Task |
|
|
527
|
+
|---|---|
|
|
528
|
+
| Site 1 — `__getitem__` integer branch | 1 |
|
|
529
|
+
| Site 2 — `_getitem_record_rows` | 2 |
|
|
530
|
+
| `_getitem_record_rows_r2` inherits site 1 | 2 (Step 1, `test_record_multidim_row_string_field_preserves_boundaries`) |
|
|
531
|
+
| Standalone/flat case unchanged | 1 (`test_standalone_string_index_still_returns_bytes`) |
|
|
532
|
+
| Test 1 — issue repro | 1 |
|
|
533
|
+
| Test 2 — numeric/string parity | 1 |
|
|
534
|
+
| Test 3 — multi-dim | 1 |
|
|
535
|
+
| Test 4 — record row with indel | 2 |
|
|
536
|
+
| Test 5 — zero-copy | 1 and 2 |
|
|
537
|
+
| Test 6 — empty group | 1 |
|
|
538
|
+
| Test 7 — `to_chars()` agreement | 1 |
|
|
539
|
+
| Test 8 — negative index / OOB | 1 |
|
|
540
|
+
| Test 9 — standalone regression | 1 |
|
|
541
|
+
| Minor bump via `!` commit | 1, 2 (commit messages) |
|
|
542
|
+
| `skills/seqpro/SKILL.md` update | 3 |
|
|
543
|
+
|
|
544
|
+
**Type consistency:** `_peel_string_row(rl: RaggedLayout, lo: int, hi: int) -> RaggedLayout` is defined in Task 1 Step 4 and used with that exact signature in Task 2 Step 3. Both call sites wrap it in `Ragged(...)`.
|
|
545
|
+
|
|
546
|
+
**Imports:** the new tests use `Ragged`, `OFFSET_TYPE`, `np`, and `pytest`. `tests/test_ragged_core.py` already imports all four. `tests/test_ragged_core_records.py` does **not** import `OFFSET_TYPE` — Task 2 Step 1 must extend its existing `from seqpro.rag._utils import lengths_to_offsets` (line 6) to `from seqpro.rag._utils import OFFSET_TYPE, lengths_to_offsets`.
|