@pikaa-ai/pikaa 0.3.23 → 0.3.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +337 -162
  6. package/dist/index.js +1 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,381 +0,0 @@
1
- ---
2
- name: pydicom
3
- description: Use pydicom to read, inspect, write, transform, and safely preflight local DICOM datasets and pixel data. Applies to DICOM metadata, transfer syntaxes, compression plugins, frames, private elements, JSON, and bounded de-identification review.
4
- license: MIT
5
- compatibility: Python 3.10+ with pydicom 3.0.2; optional pinned NumPy, Pillow, and pixel plugins. Helper CLIs are local-only and require authorized data.
6
- metadata:
7
- version: "1.1"
8
- skill-author: "K-Dense Inc."
9
- last-reviewed: "2026-07-23"
10
- ---
11
-
12
- # pydicom
13
-
14
- Use pydicom for DICOM dataset I/O and pixel processing. Version 3.0.2 is the
15
- current stable release reviewed here. It fixes CVE-2026-32711, a crafted
16
- DICOMDIR path-traversal issue. pydicom 3.0.2 declares Python `>=3.10`; its
17
- bundled DICOM dictionary is 2024c, while the live DICOM Standard may be newer.
18
-
19
- ## Mandatory safety boundary
20
-
21
- - Work only with local data that the user is authorized to access.
22
- - DICOM metadata, file names, private elements, overlays, structured content,
23
- and pixels may contain protected health information (PHI).
24
- - Never print `Dataset`, export full metadata/JSON, or log element values by
25
- default. Use a documented allowlist and aggregate output.
26
- - pydicom is a general DICOM framework, not a diagnostic viewer. Pixel output,
27
- validation, conversion, and plugin availability are not diagnostic claims.
28
- - De-identification is profile-, purpose-, recipient-, jurisdiction-, and
29
- threat-context-specific. It requires privacy/DICOM expert verification.
30
- - Never claim that a tag-removal script is DICOM PS3.15, HIPAA, GDPR, or other
31
- compliance. Preserve originals and audit derived outputs.
32
- - Treat deterministic pseudonymization keys and UID maps as re-identification
33
- secrets: use least privilege and encrypted/managed secret storage, never
34
- commit, sync, log, or share them with derivatives, and define backup,
35
- rotation, revocation, and destruction procedures. A leaked key invalidates
36
- the intended separation; rotation also changes deterministic mappings.
37
- - Set explicit input-file, file-count, frame-count, decoded-byte, and output
38
- limits before parsing untrusted or unusually large datasets.
39
-
40
- ## Installation
41
-
42
- Create or activate an isolated environment, then install the exact reviewed
43
- release:
44
-
45
- ```bash
46
- uv pip install "pydicom==3.0.2"
47
- ```
48
-
49
- Uncompressed pixel arrays and image rendering:
50
-
51
- ```bash
52
- uv pip install "pydicom==3.0.2" "numpy==2.5.1" "Pillow==12.3.0"
53
- ```
54
-
55
- Install only the transfer-syntax plugins required by the deployment:
56
-
57
- ```bash
58
- # JPEG/JPEG-LS, JPEG 2000/HTJ2K, and faster RLE through pylibjpeg
59
- uv pip install "numpy==2.5.1" "pylibjpeg==2.1.0" \
60
- "pylibjpeg-libjpeg==2.4.0" "pylibjpeg-openjpeg==2.5.0" \
61
- "pylibjpeg-rle==2.2.0"
62
-
63
- # JPEG-LS encoder/decoder
64
- uv pip install "numpy==2.5.1" "pyjpegls==1.5.1"
65
-
66
- # Alternative decoder with platform-specific wheels
67
- uv pip install "python-gdcm==3.2.6"
68
- ```
69
-
70
- Plugin licenses and wheels differ by package/platform; review them before
71
- deployment. Pillow has documented decoding limitations and pydicom cautions
72
- that plugin output must be independently checked.
73
-
74
- Native codec wheels widen the supply-chain and memory-safety boundary. For a
75
- controlled deployment, resolve these exact pins on a trusted build host, lock
76
- and verify wheel hashes/provenance, mirror approved artifacts internally, scan
77
- them, and install with hash enforcement rather than resolving from the public
78
- index at runtime.
79
-
80
- ## Choose the workflow
81
-
82
- 1. Need an aggregate overview: run `scripts/extract_metadata.py`.
83
- 2. Need bounded technical checks: run `scripts/dicom_inventory.py`.
84
- 3. Need codec deployment preflight: run
85
- `scripts/transfer_syntax_inspector.py`.
86
- 4. Need frame/memory planning: run `scripts/pixel_frame_planner.py`.
87
- 5. Need one non-diagnostic rendered frame: run
88
- `scripts/dicom_to_image.py`.
89
- 6. Need a pseudonymized derivative: read the de-identification section, create
90
- a site-reviewed action profile, then run `scripts/anonymize_dicom.py` and
91
- `scripts/deidentification_audit.py`.
92
- 7. Need to check a sensitive UID map: run
93
- `scripts/uid_mapping_validator.py`.
94
-
95
- ## Read datasets safely
96
-
97
- `dcmread()` returns a `FileDataset`, a `Dataset` subclass with File Format
98
- state such as `file_meta`, preamble, and original encoding.
99
-
100
- ```python
101
- from pathlib import Path
102
- import pydicom
103
-
104
- path = Path("authorized/input.dcm")
105
- ds = pydicom.dcmread(
106
- path,
107
- stop_before_pixels=True,
108
- specific_tags=[
109
- "SOPClassUID",
110
- "Modality",
111
- "Rows",
112
- "Columns",
113
- "NumberOfFrames",
114
- ],
115
- )
116
-
117
- technical = {
118
- "sop_class": ds.get("SOPClassUID"),
119
- "modality": ds.get("Modality"),
120
- "rows": ds.get("Rows"),
121
- "columns": ds.get("Columns"),
122
- }
123
- ```
124
-
125
- Use:
126
-
127
- - `stop_before_pixels=True` for metadata-only work.
128
- - `specific_tags=[...]` for a minimum allowlist.
129
- - `defer_size="1 MiB"` when a later write must preserve large values.
130
- - `force=False` (default). `force=True` only bypasses the File Format header
131
- check; it does not prove the bytes are valid DICOM.
132
-
133
- Do not call `print(ds)`, `repr(ds)`, or iterate values into logs on clinical
134
- data.
135
-
136
- ## Dataset, DataElement, and sequences
137
-
138
- Access standard elements by keyword and check for absence:
139
-
140
- ```python
141
- modality = ds.get("Modality", "UNSPECIFIED")
142
- if "ReferencedImageSequence" in ds:
143
- for item in ds.ReferencedImageSequence:
144
- referenced_class = item.get("ReferencedSOPClassUID")
145
- ```
146
-
147
- Tag access, such as `ds[0x0010, 0x0010]`, returns a `DataElement`; its `.value`
148
- is separate. `Sequence` behaves like a list of nested `Dataset` items. Privacy
149
- actions must recurse through every sequence item, not only the top level.
150
-
151
- When creating a file, use `FileMetaDataset` for group `0002`, keep dataset and
152
- file-meta SOP UIDs consistent, set a Transfer Syntax UID, and write in enforced
153
- File Format:
154
-
155
- ```python
156
- from pydicom import dcmwrite
157
- from pydicom.dataset import FileDataset, FileMetaDataset
158
- from pydicom.uid import CTImageStorage, ExplicitVRLittleEndian, generate_uid
159
-
160
- meta = FileMetaDataset()
161
- meta.MediaStorageSOPClassUID = CTImageStorage
162
- meta.MediaStorageSOPInstanceUID = generate_uid()
163
- meta.TransferSyntaxUID = ExplicitVRLittleEndian
164
-
165
- ds = FileDataset(None, {}, file_meta=meta, preamble=b"\0" * 128)
166
- ds.SOPClassUID = meta.MediaStorageSOPClassUID
167
- ds.SOPInstanceUID = meta.MediaStorageSOPInstanceUID
168
- # Add all attributes required by the selected IOD before writing.
169
- dcmwrite("new.dcm", ds, enforce_file_format=True, overwrite=False)
170
- ```
171
-
172
- `write_like_original` is deprecated in pydicom 3.0; use
173
- `enforce_file_format`. A successful write is not full PS3.3 IOD conformance.
174
-
175
- ## UIDs and transfer syntax
176
-
177
- The File Meta Information Transfer Syntax UID controls dataset encoding and
178
- pixel compression:
179
-
180
- ```python
181
- ts = ds.file_meta.TransferSyntaxUID
182
- summary = {
183
- "uid": str(ts),
184
- "name": ts.name,
185
- "compressed": ts.is_compressed,
186
- "implicit_vr": ts.is_implicit_VR,
187
- "little_endian": ts.is_little_endian,
188
- }
189
- ```
190
-
191
- pydicom 3.0 chooses write encoding from the Transfer Syntax UID before legacy
192
- dataset flags. Do not replace structural UIDs (Transfer Syntax, SOP Class, or
193
- coding-scheme UIDs) during pseudonymization. Instance/reference UID replacement
194
- must be one-to-one and consistent across the complete declared scope.
195
-
196
- Read [references/transfer_syntaxes.md](references/transfer_syntaxes.md) before
197
- compression, decompression, or encapsulation.
198
-
199
- ## Pixel data and frames
200
-
201
- The stable `pydicom.pixels` API supports path-based, frame-specific decoding:
202
-
203
- ```python
204
- from pydicom.pixels import pixel_array
205
-
206
- # Reads only the selected frame where the source permits it.
207
- frame = pixel_array("authorized/image.dcm", index=0, raw=False)
208
- ```
209
-
210
- Shape semantics:
211
-
212
- - grayscale single frame: `(rows, columns)`
213
- - grayscale multi-frame: `(frames, rows, columns)`
214
- - color single frame: `(rows, columns, samples)`
215
- - color multi-frame: `(frames, rows, columns, samples)`
216
-
217
- `raw=False` converts YCbCr pixel data to RGB when possible; `raw=True` retains
218
- the decoded color space after mandatory minimal processing. Use
219
- `iter_pixels(path, indices=[...])` for bounded multi-frame iteration.
220
-
221
- For grayscale display, apply transforms in this order:
222
-
223
- ```python
224
- from pydicom.pixels import apply_modality_lut, apply_voi_lut
225
-
226
- modality_values = apply_modality_lut(frame, ds)
227
- display_values = apply_voi_lut(modality_values, ds, index=0)
228
- ```
229
-
230
- Modality LUT/rescale and VOI/windowing change display/value semantics.
231
- MONOCHROME1 may require presentation inversion. Palette Color requires
232
- `apply_color_lut()`. Presentation states and ICC behavior may require a
233
- validated viewer. Never use per-frame min/max normalization for quantitative
234
- analysis.
235
-
236
- ## Compression, decompression, and encapsulation
237
-
238
- - Accessing `pixel_array` decodes as needed but does not change the dataset.
239
- - `Dataset.decompress()` changes Pixel Data in place, sets Explicit VR Little
240
- Endian, updates image metadata, and generates a new SOP Instance UID by
241
- default.
242
- - `Dataset.compress(uid)` changes Pixel Data and Transfer Syntax in place and
243
- generates a new SOP Instance UID by default.
244
- - pydicom 3.0 built-in/found encoders cover RLE Lossless, JPEG-LS, and JPEG
245
- 2000 combinations documented in the stable plugin matrix.
246
- - Each compressed frame is separately encoded and then encapsulated. Use
247
- `encapsulate()` or `encapsulate_extended()` for externally encoded frames.
248
- - Read frames with current `pydicom.encaps.generate_frames()` or `get_frame()`;
249
- legacy encapsulation generator names are deprecated for pydicom 4.
250
-
251
- Always inspect capabilities first, limit decoded bytes/frames, and verify pixel
252
- correctness independently. Lossy compression acceptability is outside pydicom
253
- and the DICOM encoding specification.
254
-
255
- ## DICOM JSON and private elements
256
-
257
- `Dataset.to_json()`, `to_json_dict()`, and `Dataset.from_json()` implement the
258
- DICOM JSON Model, but pydicom documents JSON support as beta. Full JSON may
259
- inline binary data and expose every identifier and pixel payload. Do not emit
260
- it as a metadata report. A `BulkDataURI` handler introduces separate storage,
261
- authorization, and retrieval obligations.
262
-
263
- Private elements are not standardized and may contain PHI:
264
-
265
- ```python
266
- # Recursive removal, but not sufficient de-identification by itself.
267
- ds.remove_private_tags()
268
- ```
269
-
270
- Retain private elements only under an explicit reviewed safe-private policy.
271
- Read [references/common_tags.md](references/common_tags.md) for tag access,
272
- privacy classes, and standard pointers.
273
-
274
- ## De-identification workflow
275
-
276
- DICOM PS3.15 Annex E explicitly states that confidentiality profiles do not
277
- guarantee removal of all identifying information and do not replace a complete
278
- de-identification process.
279
-
280
- 1. Define purpose, recipients, linkage needs, regulations, threat model, and
281
- acceptable re-identification risk.
282
- 2. Select the Basic Application Level Confidentiality Profile and needed
283
- options (pixel, recognizable visual features, graphics, structured content,
284
- descriptors, temporal information, patient characteristics, devices,
285
- institutions, UIDs, and safe private data).
286
- 3. Preserve source objects unchanged in controlled storage.
287
- 4. Apply every action recursively, including nested sequences.
288
- 5. Replace instance/reference UIDs consistently across the complete scope;
289
- preserve structural UIDs.
290
- 6. Decide date/time handling explicitly. A fixed shift can preserve intervals
291
- but partial dates, time zones, standalone times, leap days, longitudinal
292
- linkage, and external events require reviewed policy.
293
- 7. Inspect pixels, overlays, graphics, structured content, and recognizable
294
- visual features. Do not infer clean pixels from missing metadata or set
295
- `BurnedInAnnotation=NO` without verification.
296
- 8. Rebuild File Meta Information and preamble to prevent leakage.
297
- 9. Run technical validation and a de-identification audit, then perform expert
298
- verification and documented risk review.
299
-
300
- The bundled script intentionally sets `PatientIdentityRemoved` to `NO` because
301
- it cannot establish successful de-identification.
302
-
303
- ## Helper CLIs
304
-
305
- All `--help` paths are dependency-free. The tools perform no network access and
306
- emit no DICOM values beyond narrow technical allowlists.
307
-
308
- Bundled content consists of the two linked references, the documented helper
309
- scripts, and synthetic tests. The pydicom runtime dependency is installed from
310
- the pinned PyPI release.
311
-
312
- ```bash
313
- # Redacted aggregate metadata
314
- python scripts/extract_metadata.py authorized/ --recursive
315
-
316
- # Metadata-only technical inventory
317
- python scripts/dicom_inventory.py authorized/ --recursive
318
-
319
- # Installed codec/plugin capabilities
320
- python scripts/transfer_syntax_inspector.py --input authorized/image.dcm
321
-
322
- # Frame shape, byte, and transform plan
323
- python scripts/pixel_frame_planner.py authorized/image.dcm --frames 0,2-4
324
-
325
- # One non-diagnostic frame
326
- python scripts/dicom_to_image.py authorized/image.dcm frame.png \
327
- --acknowledge-pixel-phi
328
-
329
- # Create a secret key, then a scoped pseudonymized derivative plus audit
330
- python scripts/anonymize_dicom.py --generate-uid-key project.key
331
- python scripts/anonymize_dicom.py authorized/in.dcm derived/out.dcm \
332
- --uid-key-file project.key --uid-scope export-v1 \
333
- --audit-report derived/out.audit.json
334
-
335
- # Audit candidate metadata; no pixel decompression
336
- python scripts/deidentification_audit.py derived/out.dcm
337
-
338
- # Validate an explicitly requested sensitive UID mapping
339
- python scripts/uid_mapping_validator.py derived/uid-map.json \
340
- --uid-key-file project.key --uid-scope export-v1
341
- ```
342
-
343
- The generated raw key file is a controlled-local convenience and is created
344
- with owner-only permissions. For production, materialize key bytes from an
345
- approved secret manager into a locked ephemeral file, restrict access to the
346
- de-identification service, and securely remove it afterward. Store any optional
347
- UID map separately from derivatives; it directly links original and replacement
348
- identifiers.
349
-
350
- ## pydicom 3.0 migration notes
351
-
352
- - `read_file()` and `write_file()` were removed; use `dcmread()` and
353
- `dcmwrite()`.
354
- - `write_like_original` is deprecated; use `enforce_file_format`.
355
- - `pydicom.pixel_data_handlers` is deprecated for removal in v4; use
356
- `pydicom.pixels`.
357
- - `Dataset.pixel_array` uses the new pixels backend by default and converts
358
- YCbCr to RGB when possible.
359
- - `JPEGLossless` now means UID `1.2.840.10008.1.2.4.57`;
360
- `JPEGLosslessSV1` is `.70`.
361
- - `Dataset.is_little_endian` and `is_implicit_VR` are deprecated for v4.
362
-
363
- ## Sources (verified 2026-07-23)
364
-
365
- - [pydicom 3.0.2 on PyPI](https://pypi.org/project/pydicom/) — released
366
- 2026-03-19; Python `>=3.10`.
367
- - [pydicom releases](https://github.com/pydicom/pydicom/releases) — 3.0.2 and
368
- CVE-2026-32711 details.
369
- - [Stable release notes](https://pydicom.github.io/pydicom/stable/release_notes/index.html)
370
- - [Stable installation guide](https://pydicom.github.io/pydicom/stable/tutorials/installation.html)
371
- - [Dataset basics](https://pydicom.github.io/pydicom/stable/tutorials/dataset_basics.html)
372
- - [Stable pixel tutorial](https://pydicom.github.io/pydicom/stable/tutorials/pixel_data/introduction.html)
373
- - [Stable pixel plugins](https://pydicom.github.io/pydicom/stable/guides/user/image_data_handlers.html)
374
- - [Stable compression tutorial](https://pydicom.github.io/pydicom/stable/tutorials/pixel_data/compressing.html)
375
- - [Stable DICOM JSON tutorial](https://pydicom.github.io/pydicom/stable/tutorials/dicom_json.html)
376
- - [Stable private-element guide](https://pydicom.github.io/pydicom/stable/guides/user/private_data_elements.html)
377
- - [Current DICOM Standard](https://www.dicomstandard.org/current)
378
- - [DICOM PS3.3](https://dicom.nema.org/medical/dicom/current/output/chtml/part03/PS3.3.html),
379
- [PS3.5](https://dicom.nema.org/medical/dicom/current/output/chtml/part05/PS3.5.html),
380
- [PS3.6](https://dicom.nema.org/medical/dicom/current/output/chtml/part06/PS3.6.html),
381
- and [PS3.15](https://dicom.nema.org/medical/dicom/current/output/html/part15.html)
@@ -1,124 +0,0 @@
1
- ---
2
- name: pyhealth
3
- description: Build clinical/healthcare deep-learning pipelines with PyHealth — loading EHR/signal/imaging datasets (MIMIC-III/IV, eICU, OMOP, SleepEDF, ChestXray14, EHRShot), defining tasks (mortality, readmission, length-of-stay, drug recommendation, sleep staging, ICD coding, EEG events), instantiating models (Transformer, RETAIN, GAMENet, SafeDrug, MICRON, StageNet, AdaCare, CNN/RNN/MLP), training with the PyHealth Trainer, computing clinical metrics, and using medical code utilities (ICD/ATC/NDC/RxNorm lookup and cross-mapping). Use this skill whenever the user mentions PyHealth, MIMIC, eICU, OMOP, EHR modeling, clinical prediction, drug recommendation, sleep staging, medical code mapping, ICD/ATC codes, or any healthcare ML pipeline that fits the dataset → task → model → trainer → metrics pattern, even if "PyHealth" isn't named explicitly.
4
- metadata:
5
- version: "1.0"
6
- skill-author: K-Dense Inc.
7
- ---
8
-
9
- # PyHealth
10
-
11
- PyHealth (https://pyhealth.dev/) is a Python toolkit for clinical deep learning. It provides a unified, modular pipeline across electronic health records (EHR), physiological signals, and medical imaging.
12
-
13
- The library is built around a **5-stage pipeline** — `Dataset → Task → Model → Trainer → Metrics` — where each stage is replaceable and the interfaces between stages are stable. Code that follows this pipeline shape composes well; code that bypasses it usually fights the library.
14
-
15
- ## When to use this skill
16
-
17
- Use this skill whenever the user is doing clinical/healthcare ML and any of the following are true:
18
-
19
- - They mention PyHealth, MIMIC-III/IV, eICU, OMOP-CDM, EHRShot, SleepEDF, SHHS, ISRUC, COVID19-CXR, ChestX-ray14, TUEV/TUAB.
20
- - They want to predict mortality, readmission, length of stay, drug recommendations, sleep stages, ICD codes, EEG events, or de-identification.
21
- - They need to look up or cross-map medical codes (ICD-9-CM, ICD-10-CM, ATC, NDC, RxNorm, CCS).
22
- - They have EHR-shaped data and want to train a clinical model without writing the plumbing themselves.
23
-
24
- PyHealth is the right tool when the workflow fits its 5 stages. If the user just wants generic PyTorch on tabular data, this skill is not necessary.
25
-
26
- ## Installation (uv)
27
-
28
- PyHealth 2.0 requires Python ≥ 3.12, < 3.14. Use `uv` for environment management — it's faster and reproducible.
29
-
30
- ```bash
31
- # Create a project with the right Python
32
- uv init my-pyhealth-project
33
- cd my-pyhealth-project
34
- uv python pin 3.12
35
-
36
- # Add PyHealth (this also pulls in PyTorch and friends)
37
- uv add pyhealth
38
-
39
- # Run scripts inside the env
40
- uv run python train.py
41
- ```
42
-
43
- For a one-off script without a project, use `uv run --with pyhealth python script.py`. For the legacy 1.x line (Python 3.9+), `uv add pyhealth==1.16`. Detailed install notes, MIMIC access, and GPU/CPU device tips are in `references/installation.md`.
44
-
45
- ## The 5-stage pipeline
46
-
47
- A complete pipeline is typically <20 lines. This is the canonical shape — start here and modify pieces:
48
-
49
- ```python
50
- from pyhealth.datasets import MIMIC3Dataset, split_by_patient, get_dataloader
51
- from pyhealth.tasks import MortalityPredictionMIMIC3
52
- from pyhealth.models import Transformer
53
- from pyhealth.trainer import Trainer
54
- from pyhealth.metrics.binary import binary_metrics_fn
55
-
56
- # 1. Dataset — raw patient registry
57
- base = MIMIC3Dataset(
58
- root="https://storage.googleapis.com/pyhealth/Synthetic_MIMIC-III/",
59
- tables=["DIAGNOSES_ICD", "PROCEDURES_ICD", "PRESCRIPTIONS"],
60
- )
61
-
62
- # 2. Task — converts patients into supervised samples
63
- samples = base.set_task(MortalityPredictionMIMIC3())
64
-
65
- # 3. Split + DataLoaders (split by patient to avoid leakage)
66
- train_ds, val_ds, test_ds = split_by_patient(samples, [0.8, 0.1, 0.1])
67
- train_loader = get_dataloader(train_ds, batch_size=32, shuffle=True)
68
- val_loader = get_dataloader(val_ds, batch_size=32, shuffle=False)
69
- test_loader = get_dataloader(test_ds, batch_size=32, shuffle=False)
70
-
71
- # 4. Model — must be passed the SampleDataset, not the BaseDataset
72
- model = Transformer(dataset=samples)
73
-
74
- # 5. Train + evaluate
75
- trainer = Trainer(model=model)
76
- trainer.train(
77
- train_dataloader=train_loader,
78
- val_dataloader=val_loader,
79
- epochs=50,
80
- monitor="pr_auc",
81
- )
82
-
83
- y_true, y_prob, _ = trainer.inference(test_loader)
84
- print(binary_metrics_fn(y_true, y_prob, metrics=["pr_auc", "roc_auc"]))
85
- ```
86
-
87
- A copy-pasteable starter is in `assets/starter_pipeline.py`.
88
-
89
- ## Critical things to get right
90
-
91
- These are the mistakes that PyHealth code most commonly trips on. Internalize them before writing pipelines:
92
-
93
- 1. **Models take a `SampleDataset`, not a `BaseDataset`.** `MIMIC3Dataset(...)` returns a `BaseDataset` (a queryable patient registry). Only after `.set_task(task)` do you get a `SampleDataset`, which is what models, splitters, and DataLoaders expect. If you pass `base` to a model, it will fail or behave wrong.
94
-
95
- 2. **Always split by patient (or visit), not by sample.** Random sample-level splits leak information across train/test because the same patient can appear in both. Use `split_by_patient` for patient-level prediction, `split_by_visit` only when visits are independent.
96
-
97
- 3. **Match the task to the dataset.** Tasks are dataset-specific: `MortalityPredictionMIMIC3` won't work on MIMIC-IV — use `MortalityPredictionMIMIC4` or `InHospitalMortalityMIMIC4`. The full mapping is in `references/tasks.md`.
98
-
99
- 4. **Pick `monitor` to match the task type.** For binary classification use `"pr_auc"` or `"roc_auc"`. For multilabel (drug rec) use `"pr_auc_samples"` or `"jaccard_samples"`. For multiclass use `"accuracy"` or `"f1_macro"`. Wrong monitor → checkpoint selection saves the wrong epoch.
100
-
101
- 5. **MIMIC-IV uses `ehr_root=`, not `root=`.** This is the one inconsistency in the dataset constructors.
102
-
103
- 6. **For reproducible work, point `cache_dir=` somewhere persistent.** PyHealth caches the parsed dataset; without `cache_dir`, you re-parse every run.
104
-
105
- ## How to use this skill
106
-
107
- PyHealth has a large API surface — there's no point loading it all at once. Read the reference file that matches the user's task:
108
-
109
- | If the user is asking about… | Read |
110
- |---|---|
111
- | Installing, env setup, MIMIC access, GPU | `references/installation.md` |
112
- | Which dataset class to use, loading patterns, splitting | `references/datasets.md` |
113
- | What prediction task to choose (mortality, readmission, drug rec, sleep…) | `references/tasks.md` |
114
- | Picking a model architecture, model-specific arguments | `references/models.md` |
115
- | Looking up or cross-mapping ICD/ATC/NDC/RxNorm/CCS codes, tokenizers | `references/medcode.md` |
116
- | End-to-end recipes for common scenarios | `references/examples.md` |
117
-
118
- For multi-step tasks (e.g., "build a drug recommendation pipeline on MIMIC-IV"), read `tasks.md` + `models.md` + `examples.md` together — they cross-reference each other.
119
-
120
- ## A note on style
121
-
122
- Write minimal, idiomatic PyHealth. The library is opinionated; lean into its abstractions instead of reimplementing them in raw PyTorch. If you find yourself writing a custom training loop, ask whether `Trainer` would do the job — it almost always will, and it handles checkpointing, logging, and best-model selection for free.
123
-
124
- When the user has private MIMIC access, point them at the local CSV root; for demos and learning, the synthetic MIMIC-III bucket (`https://storage.googleapis.com/pyhealth/Synthetic_MIMIC-III/`) is fine and works without credentialing.