@pikaa-ai/pikaa 0.3.22 → 0.3.24
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
- package/assets/brand/orbit-logo.jpg +0 -0
- package/assets/brand/orbit-logo.png +0 -0
- package/assets/brand/orbit-logo.svg +3 -0
- package/dist/cli.js +448 -181
- package/dist/index.js +22 -2
- package/package.json +1 -2
- package/skills/adaptyv/SKILL.md +0 -240
- package/skills/aeon/SKILL.md +0 -402
- package/skills/analytical-method-validation/SKILL.md +0 -299
- package/skills/anndata/SKILL.md +0 -431
- package/skills/arbor/SKILL.md +0 -152
- package/skills/arboreto/SKILL.md +0 -267
- package/skills/astropy/SKILL.md +0 -353
- package/skills/autoskill/SKILL.md +0 -233
- package/skills/benchling-integration/SKILL.md +0 -229
- package/skills/bgpt-paper-search/SKILL.md +0 -75
- package/skills/bids/SKILL.md +0 -237
- package/skills/biopython/SKILL.md +0 -472
- package/skills/bioservices/SKILL.md +0 -399
- package/skills/bulk-rnaseq/SKILL.md +0 -198
- package/skills/cellxgene-census/SKILL.md +0 -283
- package/skills/cirq/SKILL.md +0 -370
- package/skills/citation-management/SKILL.md +0 -329
- package/skills/clinical-decision-support/SKILL.md +0 -238
- package/skills/clinical-decision-support/references/README.md +0 -62
- package/skills/clinical-reports/SKILL.md +0 -248
- package/skills/clinical-reports/references/README.md +0 -34
- package/skills/cobrapy/SKILL.md +0 -496
- package/skills/consciousness-council/SKILL.md +0 -151
- package/skills/dask/SKILL.md +0 -482
- package/skills/database-lookup/SKILL.md +0 -386
- package/skills/datamol/SKILL.md +0 -200
- package/skills/deepchem/SKILL.md +0 -244
- package/skills/deepspot-m/SKILL.md +0 -175
- package/skills/deeptools/SKILL.md +0 -412
- package/skills/depmap/SKILL.md +0 -301
- package/skills/dhdna-profiler/SKILL.md +0 -184
- package/skills/diffdock/SKILL.md +0 -488
- package/skills/dnanexus-integration/SKILL.md +0 -325
- package/skills/docx/SKILL.md +0 -99
- package/skills/esm/SKILL.md +0 -334
- package/skills/etetoolkit/SKILL.md +0 -327
- package/skills/exa-search/SKILL.md +0 -102
- package/skills/executing-plans/SKILL.md +0 -14
- package/skills/experimental-design/SKILL.md +0 -234
- package/skills/exploratory-data-analysis/SKILL.md +0 -280
- package/skills/flowio/SKILL.md +0 -310
- package/skills/fluidsim/SKILL.md +0 -279
- package/skills/frontend-design/SKILL.md +0 -100
- package/skills/generate-image/SKILL.md +0 -304
- package/skills/geniml/SKILL.md +0 -310
- package/skills/genomic-coordinates/SKILL.md +0 -189
- package/skills/genomic-intelligence/SKILL.md +0 -243
- package/skills/geomaster/README.md +0 -105
- package/skills/geomaster/SKILL.md +0 -366
- package/skills/geopandas/SKILL.md +0 -250
- package/skills/get-available-resources/SKILL.md +0 -260
- package/skills/gget/SKILL.md +0 -153
- package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
- package/skills/glycoengineering/SKILL.md +0 -339
- package/skills/gtars/SKILL.md +0 -282
- package/skills/guardian-rails/SKILL.md +0 -54
- package/skills/histolab/SKILL.md +0 -243
- package/skills/hugging-science/SKILL.md +0 -132
- package/skills/hypogenic/SKILL.md +0 -290
- package/skills/hypothesis-generation/SKILL.md +0 -264
- package/skills/imaging-data-commons/SKILL.md +0 -496
- package/skills/infographics/SKILL.md +0 -315
- package/skills/iso-standards-readiness/SKILL.md +0 -352
- package/skills/lab-hardware-cad/SKILL.md +0 -372
- package/skills/labarchive-integration/SKILL.md +0 -216
- package/skills/lamindb/SKILL.md +0 -408
- package/skills/latchbio-integration/SKILL.md +0 -227
- package/skills/latex-posters/SKILL.md +0 -369
- package/skills/latex-posters/references/README.md +0 -439
- package/skills/liteparse/SKILL.md +0 -295
- package/skills/literature-review/SKILL.md +0 -263
- package/skills/markdown-mermaid-writing/SKILL.md +0 -322
- package/skills/market-research-reports/SKILL.md +0 -337
- package/skills/markitdown/SKILL.md +0 -264
- package/skills/matchms/SKILL.md +0 -276
- package/skills/matlab/SKILL.md +0 -274
- package/skills/matplotlib/SKILL.md +0 -378
- package/skills/medchem/SKILL.md +0 -321
- package/skills/modal/SKILL.md +0 -468
- package/skills/molecular-dynamics/SKILL.md +0 -458
- package/skills/molfeat/SKILL.md +0 -348
- package/skills/ncats-arax/SKILL.md +0 -178
- package/skills/networkx/SKILL.md +0 -440
- package/skills/neurokit2/SKILL.md +0 -323
- package/skills/neuropixels-analysis/SKILL.md +0 -412
- package/skills/nextflow/SKILL.md +0 -195
- package/skills/omero-integration/SKILL.md +0 -222
- package/skills/onekgpd/SKILL.md +0 -371
- package/skills/ontology-term-resolution/SKILL.md +0 -147
- package/skills/open-notebook/SKILL.md +0 -297
- package/skills/openpiv/SKILL.md +0 -469
- package/skills/opentrons-integration/SKILL.md +0 -322
- package/skills/optimize-for-gpu/SKILL.md +0 -176
- package/skills/owasp-top10/SKILL.md +0 -48
- package/skills/pacsomatic/LICENSE +0 -21
- package/skills/pacsomatic/SKILL.md +0 -150
- package/skills/paper-lookup/SKILL.md +0 -263
- package/skills/paperclip/SKILL.md +0 -413
- package/skills/paperzilla/SKILL.md +0 -159
- package/skills/parallel-web/SKILL.md +0 -128
- package/skills/pathml/SKILL.md +0 -222
- package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
- package/skills/pathway-enrichment/SKILL.md +0 -194
- package/skills/pdf/SKILL.md +0 -322
- package/skills/peer-review/SKILL.md +0 -288
- package/skills/penetration-testing/SKILL.md +0 -31
- package/skills/pennylane/SKILL.md +0 -240
- package/skills/phylogenetics/SKILL.md +0 -409
- package/skills/pi-agent/SKILL.md +0 -83
- package/skills/pkpd-modeling/SKILL.md +0 -381
- package/skills/polars/SKILL.md +0 -393
- package/skills/polars-bio/SKILL.md +0 -379
- package/skills/ponytail/SKILL.md +0 -31
- package/skills/ponytail-audit/SKILL.md +0 -18
- package/skills/pptx/SKILL.md +0 -246
- package/skills/pptx-posters/SKILL.md +0 -258
- package/skills/primekg/SKILL.md +0 -99
- package/skills/protocolsio-integration/SKILL.md +0 -236
- package/skills/pufferlib/SKILL.md +0 -328
- package/skills/pydeseq2/SKILL.md +0 -369
- package/skills/pydicom/SKILL.md +0 -381
- package/skills/pyhealth/SKILL.md +0 -124
- package/skills/pylabrobot/SKILL.md +0 -216
- package/skills/pymatgen/SKILL.md +0 -404
- package/skills/pymc/SKILL.md +0 -310
- package/skills/pymoo/SKILL.md +0 -276
- package/skills/pyopenms/SKILL.md +0 -179
- package/skills/pysam/SKILL.md +0 -330
- package/skills/pytdc/SKILL.md +0 -297
- package/skills/pytorch-lightning/SKILL.md +0 -191
- package/skills/pyzotero/SKILL.md +0 -137
- package/skills/qiskit/SKILL.md +0 -259
- package/skills/qutip/SKILL.md +0 -317
- package/skills/rdkit/SKILL.md +0 -94
- package/skills/relsa-severity-assessment/SKILL.md +0 -354
- package/skills/research-grants/SKILL.md +0 -296
- package/skills/research-grants/references/README.md +0 -287
- package/skills/research-lookup/README.md +0 -106
- package/skills/research-lookup/SKILL.md +0 -338
- package/skills/rowan/SKILL.md +0 -398
- package/skills/scanpy/SKILL.md +0 -303
- package/skills/scholar-evaluation/SKILL.md +0 -296
- package/skills/scientific-brainstorming/SKILL.md +0 -282
- package/skills/scientific-critical-thinking/SKILL.md +0 -180
- package/skills/scientific-schematics/SKILL.md +0 -370
- package/skills/scientific-slides/SKILL.md +0 -379
- package/skills/scientific-visualization/SKILL.md +0 -285
- package/skills/scientific-writing/SKILL.md +0 -356
- package/skills/scikit-bio/SKILL.md +0 -470
- package/skills/scikit-learn/SKILL.md +0 -324
- package/skills/scikit-survival/SKILL.md +0 -313
- package/skills/scvelo/SKILL.md +0 -328
- package/skills/scvi-tools/SKILL.md +0 -201
- package/skills/seaborn/SKILL.md +0 -254
- package/skills/security-auditor/SKILL.md +0 -37
- package/skills/shap/SKILL.md +0 -282
- package/skills/simpy/SKILL.md +0 -283
- package/skills/stable-baselines3/SKILL.md +0 -325
- package/skills/statistical-analysis/SKILL.md +0 -446
- package/skills/statistical-power/SKILL.md +0 -200
- package/skills/statsmodels/SKILL.md +0 -238
- package/skills/sympy/SKILL.md +0 -354
- package/skills/systematic-debugging/SKILL.md +0 -35
- package/skills/tamarind/SKILL.md +0 -285
- package/skills/tdd/SKILL.md +0 -26
- package/skills/tiledbvcf/SKILL.md +0 -456
- package/skills/timesfm-forecasting/SKILL.md +0 -408
- package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
- package/skills/torch-geometric/SKILL.md +0 -458
- package/skills/torchdrug/SKILL.md +0 -241
- package/skills/transformers/SKILL.md +0 -195
- package/skills/treatment-plans/SKILL.md +0 -174
- package/skills/treatment-plans/references/README.md +0 -19
- package/skills/umap-learn/SKILL.md +0 -488
- package/skills/uncertainty-and-units/SKILL.md +0 -384
- package/skills/usfiscaldata/SKILL.md +0 -171
- package/skills/vaex/SKILL.md +0 -204
- package/skills/venue-templates/SKILL.md +0 -269
- package/skills/verification-before-completion/SKILL.md +0 -22
- package/skills/waypoint-bio/SKILL.md +0 -273
- package/skills/what-if-oracle/SKILL.md +0 -184
- package/skills/writing-plans/SKILL.md +0 -15
- package/skills/xlsx/SKILL.md +0 -110
- package/skills/zarr-python/SKILL.md +0 -241
package/skills/pydicom/SKILL.md
DELETED
|
@@ -1,381 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: pydicom
|
|
3
|
-
description: Use pydicom to read, inspect, write, transform, and safely preflight local DICOM datasets and pixel data. Applies to DICOM metadata, transfer syntaxes, compression plugins, frames, private elements, JSON, and bounded de-identification review.
|
|
4
|
-
license: MIT
|
|
5
|
-
compatibility: Python 3.10+ with pydicom 3.0.2; optional pinned NumPy, Pillow, and pixel plugins. Helper CLIs are local-only and require authorized data.
|
|
6
|
-
metadata:
|
|
7
|
-
version: "1.1"
|
|
8
|
-
skill-author: "K-Dense Inc."
|
|
9
|
-
last-reviewed: "2026-07-23"
|
|
10
|
-
---
|
|
11
|
-
|
|
12
|
-
# pydicom
|
|
13
|
-
|
|
14
|
-
Use pydicom for DICOM dataset I/O and pixel processing. Version 3.0.2 is the
|
|
15
|
-
current stable release reviewed here. It fixes CVE-2026-32711, a crafted
|
|
16
|
-
DICOMDIR path-traversal issue. pydicom 3.0.2 declares Python `>=3.10`; its
|
|
17
|
-
bundled DICOM dictionary is 2024c, while the live DICOM Standard may be newer.
|
|
18
|
-
|
|
19
|
-
## Mandatory safety boundary
|
|
20
|
-
|
|
21
|
-
- Work only with local data that the user is authorized to access.
|
|
22
|
-
- DICOM metadata, file names, private elements, overlays, structured content,
|
|
23
|
-
and pixels may contain protected health information (PHI).
|
|
24
|
-
- Never print `Dataset`, export full metadata/JSON, or log element values by
|
|
25
|
-
default. Use a documented allowlist and aggregate output.
|
|
26
|
-
- pydicom is a general DICOM framework, not a diagnostic viewer. Pixel output,
|
|
27
|
-
validation, conversion, and plugin availability are not diagnostic claims.
|
|
28
|
-
- De-identification is profile-, purpose-, recipient-, jurisdiction-, and
|
|
29
|
-
threat-context-specific. It requires privacy/DICOM expert verification.
|
|
30
|
-
- Never claim that a tag-removal script is DICOM PS3.15, HIPAA, GDPR, or other
|
|
31
|
-
compliance. Preserve originals and audit derived outputs.
|
|
32
|
-
- Treat deterministic pseudonymization keys and UID maps as re-identification
|
|
33
|
-
secrets: use least privilege and encrypted/managed secret storage, never
|
|
34
|
-
commit, sync, log, or share them with derivatives, and define backup,
|
|
35
|
-
rotation, revocation, and destruction procedures. A leaked key invalidates
|
|
36
|
-
the intended separation; rotation also changes deterministic mappings.
|
|
37
|
-
- Set explicit input-file, file-count, frame-count, decoded-byte, and output
|
|
38
|
-
limits before parsing untrusted or unusually large datasets.
|
|
39
|
-
|
|
40
|
-
## Installation
|
|
41
|
-
|
|
42
|
-
Create or activate an isolated environment, then install the exact reviewed
|
|
43
|
-
release:
|
|
44
|
-
|
|
45
|
-
```bash
|
|
46
|
-
uv pip install "pydicom==3.0.2"
|
|
47
|
-
```
|
|
48
|
-
|
|
49
|
-
Uncompressed pixel arrays and image rendering:
|
|
50
|
-
|
|
51
|
-
```bash
|
|
52
|
-
uv pip install "pydicom==3.0.2" "numpy==2.5.1" "Pillow==12.3.0"
|
|
53
|
-
```
|
|
54
|
-
|
|
55
|
-
Install only the transfer-syntax plugins required by the deployment:
|
|
56
|
-
|
|
57
|
-
```bash
|
|
58
|
-
# JPEG/JPEG-LS, JPEG 2000/HTJ2K, and faster RLE through pylibjpeg
|
|
59
|
-
uv pip install "numpy==2.5.1" "pylibjpeg==2.1.0" \
|
|
60
|
-
"pylibjpeg-libjpeg==2.4.0" "pylibjpeg-openjpeg==2.5.0" \
|
|
61
|
-
"pylibjpeg-rle==2.2.0"
|
|
62
|
-
|
|
63
|
-
# JPEG-LS encoder/decoder
|
|
64
|
-
uv pip install "numpy==2.5.1" "pyjpegls==1.5.1"
|
|
65
|
-
|
|
66
|
-
# Alternative decoder with platform-specific wheels
|
|
67
|
-
uv pip install "python-gdcm==3.2.6"
|
|
68
|
-
```
|
|
69
|
-
|
|
70
|
-
Plugin licenses and wheels differ by package/platform; review them before
|
|
71
|
-
deployment. Pillow has documented decoding limitations and pydicom cautions
|
|
72
|
-
that plugin output must be independently checked.
|
|
73
|
-
|
|
74
|
-
Native codec wheels widen the supply-chain and memory-safety boundary. For a
|
|
75
|
-
controlled deployment, resolve these exact pins on a trusted build host, lock
|
|
76
|
-
and verify wheel hashes/provenance, mirror approved artifacts internally, scan
|
|
77
|
-
them, and install with hash enforcement rather than resolving from the public
|
|
78
|
-
index at runtime.
|
|
79
|
-
|
|
80
|
-
## Choose the workflow
|
|
81
|
-
|
|
82
|
-
1. Need an aggregate overview: run `scripts/extract_metadata.py`.
|
|
83
|
-
2. Need bounded technical checks: run `scripts/dicom_inventory.py`.
|
|
84
|
-
3. Need codec deployment preflight: run
|
|
85
|
-
`scripts/transfer_syntax_inspector.py`.
|
|
86
|
-
4. Need frame/memory planning: run `scripts/pixel_frame_planner.py`.
|
|
87
|
-
5. Need one non-diagnostic rendered frame: run
|
|
88
|
-
`scripts/dicom_to_image.py`.
|
|
89
|
-
6. Need a pseudonymized derivative: read the de-identification section, create
|
|
90
|
-
a site-reviewed action profile, then run `scripts/anonymize_dicom.py` and
|
|
91
|
-
`scripts/deidentification_audit.py`.
|
|
92
|
-
7. Need to check a sensitive UID map: run
|
|
93
|
-
`scripts/uid_mapping_validator.py`.
|
|
94
|
-
|
|
95
|
-
## Read datasets safely
|
|
96
|
-
|
|
97
|
-
`dcmread()` returns a `FileDataset`, a `Dataset` subclass with File Format
|
|
98
|
-
state such as `file_meta`, preamble, and original encoding.
|
|
99
|
-
|
|
100
|
-
```python
|
|
101
|
-
from pathlib import Path
|
|
102
|
-
import pydicom
|
|
103
|
-
|
|
104
|
-
path = Path("authorized/input.dcm")
|
|
105
|
-
ds = pydicom.dcmread(
|
|
106
|
-
path,
|
|
107
|
-
stop_before_pixels=True,
|
|
108
|
-
specific_tags=[
|
|
109
|
-
"SOPClassUID",
|
|
110
|
-
"Modality",
|
|
111
|
-
"Rows",
|
|
112
|
-
"Columns",
|
|
113
|
-
"NumberOfFrames",
|
|
114
|
-
],
|
|
115
|
-
)
|
|
116
|
-
|
|
117
|
-
technical = {
|
|
118
|
-
"sop_class": ds.get("SOPClassUID"),
|
|
119
|
-
"modality": ds.get("Modality"),
|
|
120
|
-
"rows": ds.get("Rows"),
|
|
121
|
-
"columns": ds.get("Columns"),
|
|
122
|
-
}
|
|
123
|
-
```
|
|
124
|
-
|
|
125
|
-
Use:
|
|
126
|
-
|
|
127
|
-
- `stop_before_pixels=True` for metadata-only work.
|
|
128
|
-
- `specific_tags=[...]` for a minimum allowlist.
|
|
129
|
-
- `defer_size="1 MiB"` when a later write must preserve large values.
|
|
130
|
-
- `force=False` (default). `force=True` only bypasses the File Format header
|
|
131
|
-
check; it does not prove the bytes are valid DICOM.
|
|
132
|
-
|
|
133
|
-
Do not call `print(ds)`, `repr(ds)`, or iterate values into logs on clinical
|
|
134
|
-
data.
|
|
135
|
-
|
|
136
|
-
## Dataset, DataElement, and sequences
|
|
137
|
-
|
|
138
|
-
Access standard elements by keyword and check for absence:
|
|
139
|
-
|
|
140
|
-
```python
|
|
141
|
-
modality = ds.get("Modality", "UNSPECIFIED")
|
|
142
|
-
if "ReferencedImageSequence" in ds:
|
|
143
|
-
for item in ds.ReferencedImageSequence:
|
|
144
|
-
referenced_class = item.get("ReferencedSOPClassUID")
|
|
145
|
-
```
|
|
146
|
-
|
|
147
|
-
Tag access, such as `ds[0x0010, 0x0010]`, returns a `DataElement`; its `.value`
|
|
148
|
-
is separate. `Sequence` behaves like a list of nested `Dataset` items. Privacy
|
|
149
|
-
actions must recurse through every sequence item, not only the top level.
|
|
150
|
-
|
|
151
|
-
When creating a file, use `FileMetaDataset` for group `0002`, keep dataset and
|
|
152
|
-
file-meta SOP UIDs consistent, set a Transfer Syntax UID, and write in enforced
|
|
153
|
-
File Format:
|
|
154
|
-
|
|
155
|
-
```python
|
|
156
|
-
from pydicom import dcmwrite
|
|
157
|
-
from pydicom.dataset import FileDataset, FileMetaDataset
|
|
158
|
-
from pydicom.uid import CTImageStorage, ExplicitVRLittleEndian, generate_uid
|
|
159
|
-
|
|
160
|
-
meta = FileMetaDataset()
|
|
161
|
-
meta.MediaStorageSOPClassUID = CTImageStorage
|
|
162
|
-
meta.MediaStorageSOPInstanceUID = generate_uid()
|
|
163
|
-
meta.TransferSyntaxUID = ExplicitVRLittleEndian
|
|
164
|
-
|
|
165
|
-
ds = FileDataset(None, {}, file_meta=meta, preamble=b"\0" * 128)
|
|
166
|
-
ds.SOPClassUID = meta.MediaStorageSOPClassUID
|
|
167
|
-
ds.SOPInstanceUID = meta.MediaStorageSOPInstanceUID
|
|
168
|
-
# Add all attributes required by the selected IOD before writing.
|
|
169
|
-
dcmwrite("new.dcm", ds, enforce_file_format=True, overwrite=False)
|
|
170
|
-
```
|
|
171
|
-
|
|
172
|
-
`write_like_original` is deprecated in pydicom 3.0; use
|
|
173
|
-
`enforce_file_format`. A successful write is not full PS3.3 IOD conformance.
|
|
174
|
-
|
|
175
|
-
## UIDs and transfer syntax
|
|
176
|
-
|
|
177
|
-
The File Meta Information Transfer Syntax UID controls dataset encoding and
|
|
178
|
-
pixel compression:
|
|
179
|
-
|
|
180
|
-
```python
|
|
181
|
-
ts = ds.file_meta.TransferSyntaxUID
|
|
182
|
-
summary = {
|
|
183
|
-
"uid": str(ts),
|
|
184
|
-
"name": ts.name,
|
|
185
|
-
"compressed": ts.is_compressed,
|
|
186
|
-
"implicit_vr": ts.is_implicit_VR,
|
|
187
|
-
"little_endian": ts.is_little_endian,
|
|
188
|
-
}
|
|
189
|
-
```
|
|
190
|
-
|
|
191
|
-
pydicom 3.0 chooses write encoding from the Transfer Syntax UID before legacy
|
|
192
|
-
dataset flags. Do not replace structural UIDs (Transfer Syntax, SOP Class, or
|
|
193
|
-
coding-scheme UIDs) during pseudonymization. Instance/reference UID replacement
|
|
194
|
-
must be one-to-one and consistent across the complete declared scope.
|
|
195
|
-
|
|
196
|
-
Read [references/transfer_syntaxes.md](references/transfer_syntaxes.md) before
|
|
197
|
-
compression, decompression, or encapsulation.
|
|
198
|
-
|
|
199
|
-
## Pixel data and frames
|
|
200
|
-
|
|
201
|
-
The stable `pydicom.pixels` API supports path-based, frame-specific decoding:
|
|
202
|
-
|
|
203
|
-
```python
|
|
204
|
-
from pydicom.pixels import pixel_array
|
|
205
|
-
|
|
206
|
-
# Reads only the selected frame where the source permits it.
|
|
207
|
-
frame = pixel_array("authorized/image.dcm", index=0, raw=False)
|
|
208
|
-
```
|
|
209
|
-
|
|
210
|
-
Shape semantics:
|
|
211
|
-
|
|
212
|
-
- grayscale single frame: `(rows, columns)`
|
|
213
|
-
- grayscale multi-frame: `(frames, rows, columns)`
|
|
214
|
-
- color single frame: `(rows, columns, samples)`
|
|
215
|
-
- color multi-frame: `(frames, rows, columns, samples)`
|
|
216
|
-
|
|
217
|
-
`raw=False` converts YCbCr pixel data to RGB when possible; `raw=True` retains
|
|
218
|
-
the decoded color space after mandatory minimal processing. Use
|
|
219
|
-
`iter_pixels(path, indices=[...])` for bounded multi-frame iteration.
|
|
220
|
-
|
|
221
|
-
For grayscale display, apply transforms in this order:
|
|
222
|
-
|
|
223
|
-
```python
|
|
224
|
-
from pydicom.pixels import apply_modality_lut, apply_voi_lut
|
|
225
|
-
|
|
226
|
-
modality_values = apply_modality_lut(frame, ds)
|
|
227
|
-
display_values = apply_voi_lut(modality_values, ds, index=0)
|
|
228
|
-
```
|
|
229
|
-
|
|
230
|
-
Modality LUT/rescale and VOI/windowing change display/value semantics.
|
|
231
|
-
MONOCHROME1 may require presentation inversion. Palette Color requires
|
|
232
|
-
`apply_color_lut()`. Presentation states and ICC behavior may require a
|
|
233
|
-
validated viewer. Never use per-frame min/max normalization for quantitative
|
|
234
|
-
analysis.
|
|
235
|
-
|
|
236
|
-
## Compression, decompression, and encapsulation
|
|
237
|
-
|
|
238
|
-
- Accessing `pixel_array` decodes as needed but does not change the dataset.
|
|
239
|
-
- `Dataset.decompress()` changes Pixel Data in place, sets Explicit VR Little
|
|
240
|
-
Endian, updates image metadata, and generates a new SOP Instance UID by
|
|
241
|
-
default.
|
|
242
|
-
- `Dataset.compress(uid)` changes Pixel Data and Transfer Syntax in place and
|
|
243
|
-
generates a new SOP Instance UID by default.
|
|
244
|
-
- pydicom 3.0 built-in/found encoders cover RLE Lossless, JPEG-LS, and JPEG
|
|
245
|
-
2000 combinations documented in the stable plugin matrix.
|
|
246
|
-
- Each compressed frame is separately encoded and then encapsulated. Use
|
|
247
|
-
`encapsulate()` or `encapsulate_extended()` for externally encoded frames.
|
|
248
|
-
- Read frames with current `pydicom.encaps.generate_frames()` or `get_frame()`;
|
|
249
|
-
legacy encapsulation generator names are deprecated for pydicom 4.
|
|
250
|
-
|
|
251
|
-
Always inspect capabilities first, limit decoded bytes/frames, and verify pixel
|
|
252
|
-
correctness independently. Lossy compression acceptability is outside pydicom
|
|
253
|
-
and the DICOM encoding specification.
|
|
254
|
-
|
|
255
|
-
## DICOM JSON and private elements
|
|
256
|
-
|
|
257
|
-
`Dataset.to_json()`, `to_json_dict()`, and `Dataset.from_json()` implement the
|
|
258
|
-
DICOM JSON Model, but pydicom documents JSON support as beta. Full JSON may
|
|
259
|
-
inline binary data and expose every identifier and pixel payload. Do not emit
|
|
260
|
-
it as a metadata report. A `BulkDataURI` handler introduces separate storage,
|
|
261
|
-
authorization, and retrieval obligations.
|
|
262
|
-
|
|
263
|
-
Private elements are not standardized and may contain PHI:
|
|
264
|
-
|
|
265
|
-
```python
|
|
266
|
-
# Recursive removal, but not sufficient de-identification by itself.
|
|
267
|
-
ds.remove_private_tags()
|
|
268
|
-
```
|
|
269
|
-
|
|
270
|
-
Retain private elements only under an explicit reviewed safe-private policy.
|
|
271
|
-
Read [references/common_tags.md](references/common_tags.md) for tag access,
|
|
272
|
-
privacy classes, and standard pointers.
|
|
273
|
-
|
|
274
|
-
## De-identification workflow
|
|
275
|
-
|
|
276
|
-
DICOM PS3.15 Annex E explicitly states that confidentiality profiles do not
|
|
277
|
-
guarantee removal of all identifying information and do not replace a complete
|
|
278
|
-
de-identification process.
|
|
279
|
-
|
|
280
|
-
1. Define purpose, recipients, linkage needs, regulations, threat model, and
|
|
281
|
-
acceptable re-identification risk.
|
|
282
|
-
2. Select the Basic Application Level Confidentiality Profile and needed
|
|
283
|
-
options (pixel, recognizable visual features, graphics, structured content,
|
|
284
|
-
descriptors, temporal information, patient characteristics, devices,
|
|
285
|
-
institutions, UIDs, and safe private data).
|
|
286
|
-
3. Preserve source objects unchanged in controlled storage.
|
|
287
|
-
4. Apply every action recursively, including nested sequences.
|
|
288
|
-
5. Replace instance/reference UIDs consistently across the complete scope;
|
|
289
|
-
preserve structural UIDs.
|
|
290
|
-
6. Decide date/time handling explicitly. A fixed shift can preserve intervals
|
|
291
|
-
but partial dates, time zones, standalone times, leap days, longitudinal
|
|
292
|
-
linkage, and external events require reviewed policy.
|
|
293
|
-
7. Inspect pixels, overlays, graphics, structured content, and recognizable
|
|
294
|
-
visual features. Do not infer clean pixels from missing metadata or set
|
|
295
|
-
`BurnedInAnnotation=NO` without verification.
|
|
296
|
-
8. Rebuild File Meta Information and preamble to prevent leakage.
|
|
297
|
-
9. Run technical validation and a de-identification audit, then perform expert
|
|
298
|
-
verification and documented risk review.
|
|
299
|
-
|
|
300
|
-
The bundled script intentionally sets `PatientIdentityRemoved` to `NO` because
|
|
301
|
-
it cannot establish successful de-identification.
|
|
302
|
-
|
|
303
|
-
## Helper CLIs
|
|
304
|
-
|
|
305
|
-
All `--help` paths are dependency-free. The tools perform no network access and
|
|
306
|
-
emit no DICOM values beyond narrow technical allowlists.
|
|
307
|
-
|
|
308
|
-
Bundled content consists of the two linked references, the documented helper
|
|
309
|
-
scripts, and synthetic tests. The pydicom runtime dependency is installed from
|
|
310
|
-
the pinned PyPI release.
|
|
311
|
-
|
|
312
|
-
```bash
|
|
313
|
-
# Redacted aggregate metadata
|
|
314
|
-
python scripts/extract_metadata.py authorized/ --recursive
|
|
315
|
-
|
|
316
|
-
# Metadata-only technical inventory
|
|
317
|
-
python scripts/dicom_inventory.py authorized/ --recursive
|
|
318
|
-
|
|
319
|
-
# Installed codec/plugin capabilities
|
|
320
|
-
python scripts/transfer_syntax_inspector.py --input authorized/image.dcm
|
|
321
|
-
|
|
322
|
-
# Frame shape, byte, and transform plan
|
|
323
|
-
python scripts/pixel_frame_planner.py authorized/image.dcm --frames 0,2-4
|
|
324
|
-
|
|
325
|
-
# One non-diagnostic frame
|
|
326
|
-
python scripts/dicom_to_image.py authorized/image.dcm frame.png \
|
|
327
|
-
--acknowledge-pixel-phi
|
|
328
|
-
|
|
329
|
-
# Create a secret key, then a scoped pseudonymized derivative plus audit
|
|
330
|
-
python scripts/anonymize_dicom.py --generate-uid-key project.key
|
|
331
|
-
python scripts/anonymize_dicom.py authorized/in.dcm derived/out.dcm \
|
|
332
|
-
--uid-key-file project.key --uid-scope export-v1 \
|
|
333
|
-
--audit-report derived/out.audit.json
|
|
334
|
-
|
|
335
|
-
# Audit candidate metadata; no pixel decompression
|
|
336
|
-
python scripts/deidentification_audit.py derived/out.dcm
|
|
337
|
-
|
|
338
|
-
# Validate an explicitly requested sensitive UID mapping
|
|
339
|
-
python scripts/uid_mapping_validator.py derived/uid-map.json \
|
|
340
|
-
--uid-key-file project.key --uid-scope export-v1
|
|
341
|
-
```
|
|
342
|
-
|
|
343
|
-
The generated raw key file is a controlled-local convenience and is created
|
|
344
|
-
with owner-only permissions. For production, materialize key bytes from an
|
|
345
|
-
approved secret manager into a locked ephemeral file, restrict access to the
|
|
346
|
-
de-identification service, and securely remove it afterward. Store any optional
|
|
347
|
-
UID map separately from derivatives; it directly links original and replacement
|
|
348
|
-
identifiers.
|
|
349
|
-
|
|
350
|
-
## pydicom 3.0 migration notes
|
|
351
|
-
|
|
352
|
-
- `read_file()` and `write_file()` were removed; use `dcmread()` and
|
|
353
|
-
`dcmwrite()`.
|
|
354
|
-
- `write_like_original` is deprecated; use `enforce_file_format`.
|
|
355
|
-
- `pydicom.pixel_data_handlers` is deprecated for removal in v4; use
|
|
356
|
-
`pydicom.pixels`.
|
|
357
|
-
- `Dataset.pixel_array` uses the new pixels backend by default and converts
|
|
358
|
-
YCbCr to RGB when possible.
|
|
359
|
-
- `JPEGLossless` now means UID `1.2.840.10008.1.2.4.57`;
|
|
360
|
-
`JPEGLosslessSV1` is `.70`.
|
|
361
|
-
- `Dataset.is_little_endian` and `is_implicit_VR` are deprecated for v4.
|
|
362
|
-
|
|
363
|
-
## Sources (verified 2026-07-23)
|
|
364
|
-
|
|
365
|
-
- [pydicom 3.0.2 on PyPI](https://pypi.org/project/pydicom/) — released
|
|
366
|
-
2026-03-19; Python `>=3.10`.
|
|
367
|
-
- [pydicom releases](https://github.com/pydicom/pydicom/releases) — 3.0.2 and
|
|
368
|
-
CVE-2026-32711 details.
|
|
369
|
-
- [Stable release notes](https://pydicom.github.io/pydicom/stable/release_notes/index.html)
|
|
370
|
-
- [Stable installation guide](https://pydicom.github.io/pydicom/stable/tutorials/installation.html)
|
|
371
|
-
- [Dataset basics](https://pydicom.github.io/pydicom/stable/tutorials/dataset_basics.html)
|
|
372
|
-
- [Stable pixel tutorial](https://pydicom.github.io/pydicom/stable/tutorials/pixel_data/introduction.html)
|
|
373
|
-
- [Stable pixel plugins](https://pydicom.github.io/pydicom/stable/guides/user/image_data_handlers.html)
|
|
374
|
-
- [Stable compression tutorial](https://pydicom.github.io/pydicom/stable/tutorials/pixel_data/compressing.html)
|
|
375
|
-
- [Stable DICOM JSON tutorial](https://pydicom.github.io/pydicom/stable/tutorials/dicom_json.html)
|
|
376
|
-
- [Stable private-element guide](https://pydicom.github.io/pydicom/stable/guides/user/private_data_elements.html)
|
|
377
|
-
- [Current DICOM Standard](https://www.dicomstandard.org/current)
|
|
378
|
-
- [DICOM PS3.3](https://dicom.nema.org/medical/dicom/current/output/chtml/part03/PS3.3.html),
|
|
379
|
-
[PS3.5](https://dicom.nema.org/medical/dicom/current/output/chtml/part05/PS3.5.html),
|
|
380
|
-
[PS3.6](https://dicom.nema.org/medical/dicom/current/output/chtml/part06/PS3.6.html),
|
|
381
|
-
and [PS3.15](https://dicom.nema.org/medical/dicom/current/output/html/part15.html)
|
package/skills/pyhealth/SKILL.md
DELETED
|
@@ -1,124 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: pyhealth
|
|
3
|
-
description: Build clinical/healthcare deep-learning pipelines with PyHealth — loading EHR/signal/imaging datasets (MIMIC-III/IV, eICU, OMOP, SleepEDF, ChestXray14, EHRShot), defining tasks (mortality, readmission, length-of-stay, drug recommendation, sleep staging, ICD coding, EEG events), instantiating models (Transformer, RETAIN, GAMENet, SafeDrug, MICRON, StageNet, AdaCare, CNN/RNN/MLP), training with the PyHealth Trainer, computing clinical metrics, and using medical code utilities (ICD/ATC/NDC/RxNorm lookup and cross-mapping). Use this skill whenever the user mentions PyHealth, MIMIC, eICU, OMOP, EHR modeling, clinical prediction, drug recommendation, sleep staging, medical code mapping, ICD/ATC codes, or any healthcare ML pipeline that fits the dataset → task → model → trainer → metrics pattern, even if "PyHealth" isn't named explicitly.
|
|
4
|
-
metadata:
|
|
5
|
-
version: "1.0"
|
|
6
|
-
skill-author: K-Dense Inc.
|
|
7
|
-
---
|
|
8
|
-
|
|
9
|
-
# PyHealth
|
|
10
|
-
|
|
11
|
-
PyHealth (https://pyhealth.dev/) is a Python toolkit for clinical deep learning. It provides a unified, modular pipeline across electronic health records (EHR), physiological signals, and medical imaging.
|
|
12
|
-
|
|
13
|
-
The library is built around a **5-stage pipeline** — `Dataset → Task → Model → Trainer → Metrics` — where each stage is replaceable and the interfaces between stages are stable. Code that follows this pipeline shape composes well; code that bypasses it usually fights the library.
|
|
14
|
-
|
|
15
|
-
## When to use this skill
|
|
16
|
-
|
|
17
|
-
Use this skill whenever the user is doing clinical/healthcare ML and any of the following are true:
|
|
18
|
-
|
|
19
|
-
- They mention PyHealth, MIMIC-III/IV, eICU, OMOP-CDM, EHRShot, SleepEDF, SHHS, ISRUC, COVID19-CXR, ChestX-ray14, TUEV/TUAB.
|
|
20
|
-
- They want to predict mortality, readmission, length of stay, drug recommendations, sleep stages, ICD codes, EEG events, or de-identification.
|
|
21
|
-
- They need to look up or cross-map medical codes (ICD-9-CM, ICD-10-CM, ATC, NDC, RxNorm, CCS).
|
|
22
|
-
- They have EHR-shaped data and want to train a clinical model without writing the plumbing themselves.
|
|
23
|
-
|
|
24
|
-
PyHealth is the right tool when the workflow fits its 5 stages. If the user just wants generic PyTorch on tabular data, this skill is not necessary.
|
|
25
|
-
|
|
26
|
-
## Installation (uv)
|
|
27
|
-
|
|
28
|
-
PyHealth 2.0 requires Python ≥ 3.12, < 3.14. Use `uv` for environment management — it's faster and reproducible.
|
|
29
|
-
|
|
30
|
-
```bash
|
|
31
|
-
# Create a project with the right Python
|
|
32
|
-
uv init my-pyhealth-project
|
|
33
|
-
cd my-pyhealth-project
|
|
34
|
-
uv python pin 3.12
|
|
35
|
-
|
|
36
|
-
# Add PyHealth (this also pulls in PyTorch and friends)
|
|
37
|
-
uv add pyhealth
|
|
38
|
-
|
|
39
|
-
# Run scripts inside the env
|
|
40
|
-
uv run python train.py
|
|
41
|
-
```
|
|
42
|
-
|
|
43
|
-
For a one-off script without a project, use `uv run --with pyhealth python script.py`. For the legacy 1.x line (Python 3.9+), `uv add pyhealth==1.16`. Detailed install notes, MIMIC access, and GPU/CPU device tips are in `references/installation.md`.
|
|
44
|
-
|
|
45
|
-
## The 5-stage pipeline
|
|
46
|
-
|
|
47
|
-
A complete pipeline is typically <20 lines. This is the canonical shape — start here and modify pieces:
|
|
48
|
-
|
|
49
|
-
```python
|
|
50
|
-
from pyhealth.datasets import MIMIC3Dataset, split_by_patient, get_dataloader
|
|
51
|
-
from pyhealth.tasks import MortalityPredictionMIMIC3
|
|
52
|
-
from pyhealth.models import Transformer
|
|
53
|
-
from pyhealth.trainer import Trainer
|
|
54
|
-
from pyhealth.metrics.binary import binary_metrics_fn
|
|
55
|
-
|
|
56
|
-
# 1. Dataset — raw patient registry
|
|
57
|
-
base = MIMIC3Dataset(
|
|
58
|
-
root="https://storage.googleapis.com/pyhealth/Synthetic_MIMIC-III/",
|
|
59
|
-
tables=["DIAGNOSES_ICD", "PROCEDURES_ICD", "PRESCRIPTIONS"],
|
|
60
|
-
)
|
|
61
|
-
|
|
62
|
-
# 2. Task — converts patients into supervised samples
|
|
63
|
-
samples = base.set_task(MortalityPredictionMIMIC3())
|
|
64
|
-
|
|
65
|
-
# 3. Split + DataLoaders (split by patient to avoid leakage)
|
|
66
|
-
train_ds, val_ds, test_ds = split_by_patient(samples, [0.8, 0.1, 0.1])
|
|
67
|
-
train_loader = get_dataloader(train_ds, batch_size=32, shuffle=True)
|
|
68
|
-
val_loader = get_dataloader(val_ds, batch_size=32, shuffle=False)
|
|
69
|
-
test_loader = get_dataloader(test_ds, batch_size=32, shuffle=False)
|
|
70
|
-
|
|
71
|
-
# 4. Model — must be passed the SampleDataset, not the BaseDataset
|
|
72
|
-
model = Transformer(dataset=samples)
|
|
73
|
-
|
|
74
|
-
# 5. Train + evaluate
|
|
75
|
-
trainer = Trainer(model=model)
|
|
76
|
-
trainer.train(
|
|
77
|
-
train_dataloader=train_loader,
|
|
78
|
-
val_dataloader=val_loader,
|
|
79
|
-
epochs=50,
|
|
80
|
-
monitor="pr_auc",
|
|
81
|
-
)
|
|
82
|
-
|
|
83
|
-
y_true, y_prob, _ = trainer.inference(test_loader)
|
|
84
|
-
print(binary_metrics_fn(y_true, y_prob, metrics=["pr_auc", "roc_auc"]))
|
|
85
|
-
```
|
|
86
|
-
|
|
87
|
-
A copy-pasteable starter is in `assets/starter_pipeline.py`.
|
|
88
|
-
|
|
89
|
-
## Critical things to get right
|
|
90
|
-
|
|
91
|
-
These are the mistakes that PyHealth code most commonly trips on. Internalize them before writing pipelines:
|
|
92
|
-
|
|
93
|
-
1. **Models take a `SampleDataset`, not a `BaseDataset`.** `MIMIC3Dataset(...)` returns a `BaseDataset` (a queryable patient registry). Only after `.set_task(task)` do you get a `SampleDataset`, which is what models, splitters, and DataLoaders expect. If you pass `base` to a model, it will fail or behave wrong.
|
|
94
|
-
|
|
95
|
-
2. **Always split by patient (or visit), not by sample.** Random sample-level splits leak information across train/test because the same patient can appear in both. Use `split_by_patient` for patient-level prediction, `split_by_visit` only when visits are independent.
|
|
96
|
-
|
|
97
|
-
3. **Match the task to the dataset.** Tasks are dataset-specific: `MortalityPredictionMIMIC3` won't work on MIMIC-IV — use `MortalityPredictionMIMIC4` or `InHospitalMortalityMIMIC4`. The full mapping is in `references/tasks.md`.
|
|
98
|
-
|
|
99
|
-
4. **Pick `monitor` to match the task type.** For binary classification use `"pr_auc"` or `"roc_auc"`. For multilabel (drug rec) use `"pr_auc_samples"` or `"jaccard_samples"`. For multiclass use `"accuracy"` or `"f1_macro"`. Wrong monitor → checkpoint selection saves the wrong epoch.
|
|
100
|
-
|
|
101
|
-
5. **MIMIC-IV uses `ehr_root=`, not `root=`.** This is the one inconsistency in the dataset constructors.
|
|
102
|
-
|
|
103
|
-
6. **For reproducible work, point `cache_dir=` somewhere persistent.** PyHealth caches the parsed dataset; without `cache_dir`, you re-parse every run.
|
|
104
|
-
|
|
105
|
-
## How to use this skill
|
|
106
|
-
|
|
107
|
-
PyHealth has a large API surface — there's no point loading it all at once. Read the reference file that matches the user's task:
|
|
108
|
-
|
|
109
|
-
| If the user is asking about… | Read |
|
|
110
|
-
|---|---|
|
|
111
|
-
| Installing, env setup, MIMIC access, GPU | `references/installation.md` |
|
|
112
|
-
| Which dataset class to use, loading patterns, splitting | `references/datasets.md` |
|
|
113
|
-
| What prediction task to choose (mortality, readmission, drug rec, sleep…) | `references/tasks.md` |
|
|
114
|
-
| Picking a model architecture, model-specific arguments | `references/models.md` |
|
|
115
|
-
| Looking up or cross-mapping ICD/ATC/NDC/RxNorm/CCS codes, tokenizers | `references/medcode.md` |
|
|
116
|
-
| End-to-end recipes for common scenarios | `references/examples.md` |
|
|
117
|
-
|
|
118
|
-
For multi-step tasks (e.g., "build a drug recommendation pipeline on MIMIC-IV"), read `tasks.md` + `models.md` + `examples.md` together — they cross-reference each other.
|
|
119
|
-
|
|
120
|
-
## A note on style
|
|
121
|
-
|
|
122
|
-
Write minimal, idiomatic PyHealth. The library is opinionated; lean into its abstractions instead of reimplementing them in raw PyTorch. If you find yourself writing a custom training loop, ask whether `Trainer` would do the job — it almost always will, and it handles checkpointing, logging, and best-model selection for free.
|
|
123
|
-
|
|
124
|
-
When the user has private MIMIC access, point them at the local CSV root; for demos and learning, the synthetic MIMIC-III bucket (`https://storage.googleapis.com/pyhealth/Synthetic_MIMIC-III/`) is fine and works without credentialing.
|