@platforma-open/milaboratories.mixcr-clonotyping-2 2.23.4 → 2.23.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,5 +1,66 @@
1
1
  # @platforma-open/milaboratories.mixcr-clonotyping
2
2
 
3
+ ## 2.23.5
4
+
5
+ ### Patch Changes
6
+
7
+ - fd14efa: fix: aggregate the cohort clonotype table in shards
8
+
9
+ `aggregate-by-clonotype-key` groups every sample's clonotype rows by `clonotypeKey` and keeps,
10
+ per column, the value from the most abundant sample. ptabler lowers that `maxBy` to
11
+ `top_k_by(k=1).first()`, which polars runs as an in-memory group-by, so one run held every
12
+ group of the cohort at once. The peak followed the number of input rows and no grant could
13
+ change it: a 113-sample, 130 M-clonotype cohort needed on the order of 500 GiB and was killed
14
+ at whatever the cluster's ceiling was. The step also asked for `max(samples, 32)` cores, and the
15
+ memory need of this plan shape rises with the polars thread count.
16
+
17
+ The aggregation now runs in shards. Rows are bucketed by the first letter of `clonotypeKey`
18
+ after digits are removed, upper-cased — the same first letter the clonotype label's five-letter
19
+ prefix uses — so every key lands in exactly one shard, every group is complete inside its shard
20
+ and the label is computed per shard. The shard outputs are disjoint and a final streaming run
21
+ concatenates them. A cohort small enough for one shard runs unfiltered, as before.
22
+
23
+ The shard count is chosen from the input volume. The template reads the blob size of every
24
+ input TSV through the backend's `getBlobSize` (the same call `f.size()` resolves through) and
25
+ takes the smallest shard count whose largest shard stays under 64 GiB at 6 GiB of RAM per GiB
26
+ of TSV plus a 2 GiB intercept -- the concat + `maxBy` law measured under MILAB-6874 (4.94 at
27
+ eight threads), at the slope the SDK's default ptabler sizing uses. The memory override raises
28
+ every shard's grant but never the shard count: a request the backend cannot satisfy is clamped
29
+ without notice, so a larger target would only recreate the single oversized run. Each shard
30
+ runs on 8 cores. A backend without `getBlobSize` gets a single shard.
31
+
32
+ The `byCloneKey` Parquet import that follows was a flat 24 GiB, which the measured `write_frame`
33
+ law (`4.13 x^0.68` GiB for x GiB of TSV) says holds about 13 GiB of aggregated TSV. Its memory
34
+ is now left to the SDK, which sizes the ptabler run from the blob size of the aggregated TSV
35
+ (`2 GiB + 6 x size`, capped at 256 GiB in workflow-tengo 6.11.0, above the measured need at
36
+ every size it can express). The memory override does not apply to this import: the Xsv output
37
+ passes no floor through, and a fixed request would replace the formula.
38
+
39
+ The six other Parquet imports of the block had flat grants of 12, 16 or 24 GiB: the per-sample
40
+ `byCloneKeyBySample` table and the single-cell abundance, aggregates, properties, cell-linker
41
+ and SHM tables. A 16 GiB grant holds about 7 GiB of TSV under the same law, and one deep sample
42
+ can export twice that. Their memory is now left to the same SDK sizing. All seven imports run in
43
+ the medium queue; the Xsv import default is the light queue.
44
+
45
+ The QC report run in `export-report` framed every sample's full clonotype TSV and every filter
46
+ TSV in one 8 GiB ptabler run to count clonotypes, reads, out-of-frame and stop-codon clones per
47
+ sample. Those counts are now computed by one ptabler run per sample, sized by the SDK from that
48
+ sample's files with an 8 GiB floor, each reading only that sample's files and writing a one-row
49
+ table; the cohort run frames those rows and the qc
50
+ report, so its input no longer grows with the clonotype or cell count. In single-cell mode the
51
+ per-sample run also computes that sample's cell-pairing statistics from its single-cell chain
52
+ TSVs. The `exportClones` filter runs are unchanged.
53
+
54
+ The single-cell per-cell preprocessing run asked for one core and one GiB per sample, with
55
+ floors of 16 and 32. Its memory is now left to the SDK, which sizes the run from the blob size
56
+ of its input TSVs with a 32 GiB floor, and its cpu is pinned at 12, which the formula's slope
57
+ still covers.
58
+
59
+ The block moves to workflow-tengo 6.11.0, which ships that sizing formula and `memFloor`.
60
+
61
+ The hash override of `aggregate-by-clonotype-key` is new, so a failed aggregation is not
62
+ recovered from its old identity.
63
+
3
64
  ## 2.23.4
4
65
 
5
66
  ### Patch Changes
Binary file
@@ -1 +1 @@
1
- {"schema":"v2","description":{"id":{"organization":"milaboratories","name":"mixcr-clonotyping-2","version":"2.23.4"},"components":{"workflow":{"type":"workflow-v1","main":{"type":"relative","path":"main.plj.gz"}},"model":{"type":"relative","path":"model.json"},"ui":{"type":"relative","path":"ui.tgz"}},"meta":{"title":"MiXCR Clonotyping","description":"Extract TCR / BCR clonotypes from next-generation sequencing data","longDescription":{"type":"relative","path":"description.md"},"changelog":{"type":"relative","path":"CHANGELOG.md"},"logo":{"type":"relative","path":"block-logo.png"},"url":"https://github.com/platforma-open/mixcr-clonotyping-2","support":"mailto:support@milaboratories.com","tags":["upstream","airr","vdj","single-cell"],"organization":{"name":"MiLaboratories Inc","url":"https://milaboratories.com/","logo":{"type":"relative","path":"organization-logo.png"}},"marketplaceRanking":16900},"featureFlags":{"supportsLazyState":true,"supportsPframeQueryRanking":true,"requiresUIAPIVersion":3,"requiresModelAPIVersion":2,"requiresCreatePTable":2,"requiresPFramesVersion":1001031,"requiresPFrameSpec":true,"requiresPFrame":true,"requiresDialog":true,"requiresColumnsCollection":true},"kind":"@platforma-open/milaboratories.mixcr-clonotyping-2.kind@1.1.0"},"timestamp":1789730014033,"files":[{"name":"main.plj.gz","size":1526933,"sha256":"59637FD772BDB74973A2FD61076FEA4171007B16B59B0A7515BC89B5198F1646"},{"name":"model.json","size":575939,"sha256":"ABE8B2BA1F3818F65E7E58D81B8C72FB3274CFEDA89BC2F0F5D9AF8FE0D6C416"},{"name":"ui.tgz","size":4040808,"sha256":"A2CC355A7610E003D183B1CF4F829306718A2A47F7873DEE52E13223EE3E59F1"},{"name":"organization-logo.png","size":24439,"sha256":"FA71390C77C91E4B7FAAE5640D00F92F1E3F2869296F68B6040DD7CC549A50B5"},{"name":"description.md","size":1148,"sha256":"B319CBECC5055A89194C4D7B5768E1ABDE225053885800179439839408ECBA54"},{"name":"CHANGELOG.md","size":48838,"sha256":"A187A4D75CED7871F211C1B9D1A502BE598669E86893B643FD0AA75A2E30F7FB"},{"name":"block-logo.png","size":21527,"sha256":"6BB33BAF0CD039549661B51AE490373BE60D1811EC71F5023400928293BC2427"}]}
1
+ {"schema":"v2","description":{"id":{"organization":"milaboratories","name":"mixcr-clonotyping-2","version":"2.23.5"},"components":{"workflow":{"type":"workflow-v1","main":{"type":"relative","path":"main.plj.gz"}},"model":{"type":"relative","path":"model.json"},"ui":{"type":"relative","path":"ui.tgz"}},"meta":{"title":"MiXCR Clonotyping","description":"Extract TCR / BCR clonotypes from next-generation sequencing data","longDescription":{"type":"relative","path":"description.md"},"changelog":{"type":"relative","path":"CHANGELOG.md"},"logo":{"type":"relative","path":"block-logo.png"},"url":"https://github.com/platforma-open/mixcr-clonotyping-2","support":"mailto:support@milaboratories.com","tags":["upstream","airr","vdj","single-cell"],"organization":{"name":"MiLaboratories Inc","url":"https://milaboratories.com/","logo":{"type":"relative","path":"organization-logo.png"}},"marketplaceRanking":16900},"featureFlags":{"supportsLazyState":true,"supportsPframeQueryRanking":true,"requiresUIAPIVersion":3,"requiresModelAPIVersion":2,"requiresCreatePTable":2,"requiresPFramesVersion":1001031,"requiresPFrameSpec":true,"requiresPFrame":true,"requiresDialog":true,"requiresColumnsCollection":true},"kind":"@platforma-open/milaboratories.mixcr-clonotyping-2.kind@1.1.0"},"timestamp":1789802264605,"files":[{"name":"main.plj.gz","size":1547133,"sha256":"3E54F762850AF809CC22E04C3208D2CB422FC3B20992D01E81264FC1A669779B"},{"name":"model.json","size":575939,"sha256":"ABE8B2BA1F3818F65E7E58D81B8C72FB3274CFEDA89BC2F0F5D9AF8FE0D6C416"},{"name":"ui.tgz","size":4040839,"sha256":"DAE992453EA6CF16D1142360DEF0015651FB96FE2BA790E63048A01B273854B1"},{"name":"organization-logo.png","size":24439,"sha256":"FA71390C77C91E4B7FAAE5640D00F92F1E3F2869296F68B6040DD7CC549A50B5"},{"name":"description.md","size":1148,"sha256":"B319CBECC5055A89194C4D7B5768E1ABDE225053885800179439839408ECBA54"},{"name":"CHANGELOG.md","size":53008,"sha256":"086AD4532F8018BA5450E0E1C9BA99B1EB277AE95DD10B6123FE97D26130580F"},{"name":"block-logo.png","size":21527,"sha256":"6BB33BAF0CD039549661B51AE490373BE60D1811EC71F5023400928293BC2427"}]}
@@ -3,6 +3,6 @@
3
3
  "id": {
4
4
  "organization": "milaboratories",
5
5
  "name": "mixcr-clonotyping-2",
6
- "version": "2.23.4"
6
+ "version": "2.23.5"
7
7
  }
8
8
  }
package/block-pack/ui.tgz CHANGED
Binary file
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@platforma-open/milaboratories.mixcr-clonotyping-2",
3
- "version": "2.23.4",
3
+ "version": "2.23.5",
4
4
  "files": [
5
5
  "dist",
6
6
  "block-pack"
@@ -25,9 +25,9 @@
25
25
  "shx": "^0.4.0",
26
26
  "typescript": "~5.6.3",
27
27
  "@platforma-open/milaboratories.mixcr-clonotyping-2.kind": "1.1.0",
28
- "@platforma-open/milaboratories.mixcr-clonotyping-2.model": "1.28.0",
28
+ "@platforma-open/milaboratories.mixcr-clonotyping-2.workflow": "3.29.4",
29
29
  "@platforma-open/milaboratories.mixcr-clonotyping-2.ui": "1.27.3",
30
- "@platforma-open/milaboratories.mixcr-clonotyping-2.workflow": "3.29.3"
30
+ "@platforma-open/milaboratories.mixcr-clonotyping-2.model": "1.28.0"
31
31
  },
32
32
  "block": {
33
33
  "components": {