@pikaa-ai/pikaa 0.3.22 → 0.3.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +448 -181
  6. package/dist/index.js +22 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,297 +0,0 @@
1
- ---
2
- name: pytdc
3
- description: Use Therapeutics Data Commons through the PyTDC Python package for registry discovery, approved dataset access, task-aware splits, evaluator metrics, benchmark groups, and bounded molecular-oracle workflows.
4
- license: MIT
5
- allowed-tools: Read Write Edit Bash
6
- compatibility: Requires uv, CPython 3.11, PyTDC 1.1.15, and setuptools 80.9.0 for its legacy pkg_resources runtime import. Dataset, benchmark, checkpoint, and remote-oracle operations require network/storage review and explicit user approval.
7
- metadata:
8
- version: "1.1"
9
- skill-author: K-Dense Inc.
10
- ---
11
-
12
- # PyTDC (Therapeutics Data Commons)
13
-
14
- Use the official `PyTDC` distribution (`import tdc`) to discover therapeutic ML
15
- tasks, load approved datasets, apply task-appropriate splits, evaluate predictions,
16
- and work with curated benchmark groups. Prefer package metadata over copied dataset
17
- lists, and plan network/storage effects before constructing any loader.
18
-
19
- ## Verified snapshot
20
-
21
- - Research date: **2026-07-23**
22
- - PyPI stable: **PyTDC 1.1.15**, released 2025-03-31
23
- - Package/source repository: `mims-harvard/TDC`
24
- - Code license: MIT
25
- - PyPI supplies only a source distribution and declares no `Requires-Python`
26
- - The dependency graph makes **CPython 3.11** the reproducible target used here:
27
- `cellxgene-census==1.15.0` excludes Python 3.12, and PyTDC's constrained
28
- RDKit release has no CPython 3.13 wheel
29
- - PyTDC imports deprecated `pkg_resources` at runtime. Setuptools 82 removed that
30
- module; pin the verified compatibility release **setuptools 80.9.0**.
31
- - `tdc.readthedocs.io` still identifies itself as TDC 0.4.1; use it as API
32
- cross-reference, not as release-version evidence
33
- - Upstream publishes no GitHub tags/releases or maintained changelog. Treat
34
- undocumented migration claims as uncertainty and verify against the installed
35
- 1.1.15 source/metadata.
36
-
37
- See [references/sources.md](references/sources.md) for dated evidence and known
38
- documentation conflicts.
39
-
40
- ## Installation
41
-
42
- Use an isolated CPython 3.11 environment and pin the reviewed snapshot:
43
-
44
- ```bash
45
- uv venv --python 3.11 .venv-pytdc
46
- uv pip install --dry-run --python .venv-pytdc/bin/python \
47
- "setuptools==80.9.0" "PyTDC==1.1.15"
48
- uv pip install --python .venv-pytdc/bin/python \
49
- "setuptools==80.9.0" "PyTDC==1.1.15"
50
- ```
51
-
52
- The tested macOS ARM64 resolution installed 123 packages, including large
53
- scientific/ML dependencies, so the environment itself can transfer and occupy
54
- hundreds of megabytes before any dataset is downloaded. Review the dry run and
55
- available disk first. The direct pins identify the reviewed API snapshot; generate
56
- a platform-specific `uv.lock` in the user's project when every transitive version
57
- must also be frozen.
58
-
59
- For an ephemeral command:
60
-
61
- ```bash
62
- uv run --python 3.11 \
63
- --with "setuptools==80.9.0" --with "PyTDC==1.1.15" \
64
- python scripts/discover_metadata.py --kind tasks
65
- ```
66
-
67
- To check for a newer release, inspect the PyPI release history at
68
- <https://pypi.org/project/pytdc/>. Before changing the pin, compare its source
69
- distribution, dependencies, official repository, task registries, and smoke tests;
70
- do not silently substitute the separate `pytdc-nextml` package.
71
-
72
- ## Non-negotiable data and network policy
73
-
74
- 1. **Discover first.** Reading `tdc.metadata` or using
75
- `scripts/discover_metadata.py` does not instantiate a loader or download data.
76
- 2. **Plan second.** Record the exact task/dataset, official task page, license,
77
- expected size, cache directory, split, metric, and reproducibility seed.
78
- 3. **Ask the user before downloading.** Loader constructors fetch missing data.
79
- Some datasets and benchmark-group archives are large; model-backed oracles can
80
- fetch checkpoints; remote/docking oracles can transmit molecular structures.
81
- 4. **Execute only after approval.** In bundled CLIs, `--execute` acknowledges
82
- execution and `--download` is additionally required for MolGen corpora or
83
- supported oracle checkpoints.
84
- 5. **Keep outputs bounded.** Emit counts, schema, and small previews rather than
85
- full datasets, sequences, prediction arrays, or molecule corpora.
86
-
87
- ### Cache and cost behavior
88
-
89
- - Ordinary loaders default to `path="./data"` and save files beneath that path.
90
- The bundled scripts instead default to explicit `.pytdc-*` directories.
91
- - Core downloads use Harvard Dataverse file endpoints when a local filename is
92
- absent. Newer resource classes may use other upstream services.
93
- - `admet_group(path=...)` and other benchmark-group constructors download and
94
- extract the group archive when `<path>/<group>` is absent.
95
- - Download-backed `Oracle(...)` construction uses `./oracle` internally. The
96
- bundled oracle CLI changes into a safe runtime directory before approved calls.
97
- - PyTDC 1.1.15 does not provide a universal cache quota, eviction policy, or
98
- dataset-wide checksum manifest. Use `scripts/cache_audit.py` and manage disk
99
- retention explicitly.
100
- - Network transfer, local storage, decompression, parsing, feature generation,
101
- docking, and external service calls can all incur time or monetary cost.
102
-
103
- The PyTDC **code** is MIT. Dataset/task licenses are heterogeneous: official task
104
- pages include per-dataset terms ranging from Creative Commons licenses to
105
- non-commercial restrictions or “Not Specified.” Verify the exact dataset's page and
106
- original source terms before download, redistribution, publication, or commercial
107
- use. Cite both TDC and the original dataset.
108
-
109
- ## Start with metadata-only discovery
110
-
111
- From this skill directory:
112
-
113
- ```bash
114
- uv run --python 3.11 --with "setuptools==80.9.0" --with "PyTDC==1.1.15" \
115
- python scripts/discover_metadata.py --kind datasets --task ADME --limit 50
116
-
117
- uv run --python 3.11 --with "setuptools==80.9.0" --with "PyTDC==1.1.15" \
118
- python scripts/discover_metadata.py --kind benchmarks --limit 50
119
-
120
- uv run --python 3.11 --with "setuptools==80.9.0" --with "PyTDC==1.1.15" \
121
- python scripts/discover_metadata.py --kind evaluators --limit 100
122
- ```
123
-
124
- The package API is also metadata-only:
125
-
126
- ```python
127
- from tdc.utils import retrieve_dataset_names, retrieve_benchmark_names
128
-
129
- adme_names = retrieve_dataset_names("ADME")
130
- admet_benchmarks = retrieve_benchmark_names("admet_group")
131
- ```
132
-
133
- Use exact returned names. PyTDC performs fuzzy matching internally, but explicit
134
- matching avoids silently selecting the wrong dataset/oracle.
135
-
136
- ## Dataset workflow
137
-
138
- Plan a split without downloading:
139
-
140
- ```bash
141
- uv run --python 3.11 --with "setuptools==80.9.0" --with "PyTDC==1.1.15" \
142
- python scripts/load_and_split_data.py \
143
- --task ADME --dataset Caco2_Wang --method scaffold \
144
- --seed 42 --data-dir .pytdc-data
145
- ```
146
-
147
- After the user approves the dataset, license, transfer, and storage:
148
-
149
- ```bash
150
- uv run --python 3.11 --with "setuptools==80.9.0" --with "PyTDC==1.1.15" \
151
- python scripts/load_and_split_data.py \
152
- --task ADME --dataset Caco2_Wang --method scaffold \
153
- --seed 42 --data-dir .pytdc-data --execute
154
- ```
155
-
156
- Verified public import patterns include:
157
-
158
- ```python
159
- from tdc.single_pred import ADME, Tox
160
- from tdc.multi_pred import DDI, DTI
161
- from tdc.generation import MolGen, Reaction, RetroSyn
162
- ```
163
-
164
- Constructors perform data access, so do not run them before approval:
165
-
166
- ```python
167
- data = ADME(name="Caco2_Wang", path=".pytdc-data")
168
- frame = data.get_data(format="df")
169
- split = data.get_split(
170
- method="scaffold",
171
- seed=42,
172
- frac=[0.7, 0.1, 0.2],
173
- )
174
- # split keys are: train, valid, test
175
- ```
176
-
177
- Read [references/datasets.md](references/datasets.md) before choosing a task or
178
- dataset.
179
-
180
- ## Split selection without overclaiming leakage control
181
-
182
- - `random`: default for loaders; default seed 42 and fractions 0.7/0.1/0.2.
183
- - `scaffold`: documented generic support for molecule-based ADME, Tox, and HTS.
184
- PyTDC groups RDKit Bemis–Murcko scaffold strings (chirality disabled), but that
185
- does **not** prove absence of analog, duplicate, label, temporal, or provenance
186
- leakage.
187
- - `cold_split`: multi-instance API. Pass exact dataframe columns, for example
188
- `method="cold_split", column_name=["Drug", "Target"]`. Multi-column splitting can
189
- discard cross-partition rows and need not preserve requested row fractions.
190
- - `combination`: built-in DrugSyn combination split.
191
- - `time`: pair-loader API requiring `time_column`; the verified built-in case is
192
- `BindingDB_Patent` with its `Year` column. The API spelling is `time`, not
193
- `temporal`.
194
-
195
- Do not use undocumented `cold_drug_target`, `temporal`, or `stratified=True`
196
- examples. For every split, record PyTDC version, parameters, row counts, and exact
197
- entity overlap audits. PyTDC 1.1.15's random splitter uses the supplied seed for
198
- test sampling but a fixed `random_state=1` for validation sampling; do not describe
199
- all partitions as independently varying with the seed.
200
-
201
- Detailed semantics and caveats are in
202
- [references/utilities.md](references/utilities.md).
203
-
204
- ## Evaluators
205
-
206
- Use exact names from the installed evaluator registry:
207
-
208
- ```python
209
- from tdc import Evaluator
210
-
211
- mae = Evaluator(name="MAE")(y_true, y_pred)
212
- auroc = Evaluator(name="ROC-AUC")(y_true_binary, predicted_scores)
213
- pcc = Evaluator(name="PCC")(y_true, y_pred)
214
- ```
215
-
216
- `PCC` is the registered Pearson-correlation name; `Pearson` is not. Multi-class
217
- registry names are `micro-f1`, `macro-f1`, and `kappa`. Thresholded binary metrics
218
- default to 0.5. Metric direction and input shape are metric-specific; use the
219
- official task/benchmark metric rather than choosing from task type alone.
220
-
221
- ## Benchmark groups
222
-
223
- Use specialized classes. Top-level `from tdc import BenchmarkGroup` is retained
224
- only as a deprecated compatibility path in 1.1.15.
225
-
226
- ```python
227
- from tdc.benchmark_group import admet_group
228
-
229
- # Run only after approval: construction may download the group archive.
230
- group = admet_group(path=".pytdc-benchmarks")
231
- benchmark = group.get("Caco2_Wang")
232
- train_val = benchmark["train_val"]
233
- test = benchmark["test"]
234
- train, valid = group.get_train_valid_split(
235
- seed=1,
236
- benchmark=benchmark["name"],
237
- split_type="default",
238
- )
239
- ```
240
-
241
- For one run, `group.evaluate({name: test_predictions})` returns metric results.
242
- For leaderboard aggregation, pass a **list of at least five prediction
243
- dictionaries** to `group.evaluate_many(...)`. Do not index `group.get(...)` by
244
- seed, and do not derive dummy predictions from test labels.
245
-
246
- Use `scripts/benchmark_evaluation.py` to validate a bounded JSON prediction plan
247
- before any group download. See [references/utilities.md](references/utilities.md)
248
- for the exact JSON shape and API behavior.
249
-
250
- ## Molecular generation and oracles
251
-
252
- PyTDC supplies molecule corpora, evaluators, and oracles; it does not train or
253
- provide a generic molecule generator in the core workflow. Discover current names:
254
-
255
- ```bash
256
- uv run --python 3.11 --with "setuptools==80.9.0" --with "PyTDC==1.1.15" \
257
- python scripts/discover_metadata.py --kind oracles --limit 100
258
- ```
259
-
260
- Plan bounded local QED scoring:
261
-
262
- ```bash
263
- uv run --python 3.11 --with "setuptools==80.9.0" --with "PyTDC==1.1.15" \
264
- python scripts/molecular_generation.py score --oracle QED --smiles CCO
265
- ```
266
-
267
- Add `--execute` only after review. LogP and SA call the downloadable `fpscores`
268
- artifact in 1.1.15; they and DRD2/GSK3B/JNK3/CYP3A4_Veith also require
269
- `--download`. The helper intentionally refuses remote services, docking,
270
- distribution, and composite oracles. It preserves input order and never assumes
271
- score direction.
272
-
273
- Read [references/oracles.md](references/oracles.md) before any oracle call.
274
-
275
- ## Bundled resources
276
-
277
- ### Scripts
278
-
279
- - `scripts/discover_metadata.py` — download-free package registry discovery
280
- - `scripts/load_and_split_data.py` — task-aware split plan/explicit execution
281
- - `scripts/benchmark_evaluation.py` — prediction validation and explicit evaluation
282
- - `scripts/molecular_generation.py` — bounded local/checkpoint scoring and MolGen plan
283
- - `scripts/cache_audit.py` — read-only bounded cache manifest
284
-
285
- Every CLI uses lazy optional imports, safe relative output/cache paths, JSON
286
- summaries, bounded output, and no implicit dataset/model download.
287
-
288
- ### References
289
-
290
- - [references/datasets.md](references/datasets.md) — task discovery, data access,
291
- cache behavior, and licensing
292
- - [references/utilities.md](references/utilities.md) — splits, evaluators, and
293
- benchmark-group APIs
294
- - [references/oracles.md](references/oracles.md) — oracle categories, side effects,
295
- and safe execution
296
- - [references/sources.md](references/sources.md) — dated authoritative sources and
297
- unresolved upstream gaps
@@ -1,191 +0,0 @@
1
- ---
2
- name: pytorch-lightning
3
- description: Deep learning framework (PyTorch Lightning / lightning package). Organize PyTorch code into LightningModules, configure Trainers for multi-GPU/TPU, implement data pipelines, callbacks, logging (W&B, TensorBoard, MLflow), distributed training (DDP, FSDP, DeepSpeed), for scalable neural network training.
4
- allowed-tools: Read Write Edit Bash
5
- license: Apache-2.0 license
6
- compatibility: Requires Python 3.10+ and lightning 2.6+ (or pytorch-lightning 2.6+). GPU training needs CUDA-capable PyTorch. Optional loggers (wandb, mlflow, comet-ml) and DeepSpeed require separate installs.
7
- metadata:
8
- version: "1.1"
9
- skill-author: K-Dense Inc.
10
- ---
11
-
12
- # PyTorch Lightning
13
-
14
- ## Overview
15
-
16
- PyTorch Lightning is a deep learning framework that organizes PyTorch code to eliminate boilerplate while maintaining full flexibility. Automate training workflows, multi-device orchestration, and implement best practices for neural network training and scaling across multiple GPUs/TPUs.
17
-
18
- **Current upstream:** lightning 2.6.4 (PyPI, May 2026). Docs: [lightning.ai/docs/pytorch/stable](https://lightning.ai/docs/pytorch/stable/). Use `import lightning as L` (the `pytorch-lightning` package name still installs the same library).
19
-
20
- ## Installation
21
-
22
- ```bash
23
- uv pip install lightning
24
- ```
25
-
26
- Optional extras:
27
-
28
- ```bash
29
- uv pip install lightning[extra] # loggers, strategies, etc.
30
- uv pip install wandb mlflow # specific loggers as needed
31
- ```
32
-
33
- ## When to Use This Skill
34
-
35
- This skill should be used when:
36
- - Building, training, or deploying neural networks using PyTorch Lightning
37
- - Organizing PyTorch code into LightningModules
38
- - Configuring Trainers for multi-GPU/TPU training
39
- - Implementing data pipelines with LightningDataModules
40
- - Working with callbacks, logging, and distributed training strategies (DDP, FSDP, DeepSpeed)
41
- - Structuring deep learning projects professionally
42
-
43
- ## Core Capabilities
44
-
45
- ### 1. LightningModule - Model Definition
46
-
47
- Organize PyTorch models into six logical sections:
48
-
49
- 1. **Initialization** - `__init__()` and `setup()`
50
- 2. **Training Loop** - `training_step(batch, batch_idx)`
51
- 3. **Validation Loop** - `validation_step(batch, batch_idx)`
52
- 4. **Test Loop** - `test_step(batch, batch_idx)`
53
- 5. **Prediction** - `predict_step(batch, batch_idx)`
54
- 6. **Optimizer Configuration** - `configure_optimizers()`
55
-
56
- **Quick template reference:** See `scripts/template_lightning_module.py` for a complete boilerplate.
57
-
58
- **Detailed documentation:** Read `references/lightning_module.md` for comprehensive method documentation, hooks, properties, and best practices.
59
-
60
- ### 2. Trainer - Training Automation
61
-
62
- The Trainer automates the training loop, device management, gradient operations, and callbacks. Key features:
63
-
64
- - Multi-GPU/TPU support with strategy selection (DDP, FSDP, DeepSpeed)
65
- - Automatic mixed precision training
66
- - Gradient accumulation and clipping
67
- - Checkpointing and early stopping
68
- - Progress bars and logging
69
-
70
- **Quick setup reference:** See `scripts/quick_trainer_setup.py` for common Trainer configurations.
71
-
72
- **Detailed documentation:** Read `references/trainer.md` for all parameters, methods, and configuration options.
73
-
74
- ### 3. LightningDataModule - Data Pipeline Organization
75
-
76
- Encapsulate all data processing steps in a reusable class:
77
-
78
- 1. `prepare_data()` - Download and process data (single-process)
79
- 2. `setup()` - Create datasets and apply transforms (per-GPU)
80
- 3. `train_dataloader()` - Return training DataLoader
81
- 4. `val_dataloader()` - Return validation DataLoader
82
- 5. `test_dataloader()` - Return test DataLoader
83
-
84
- **Quick template reference:** See `scripts/template_datamodule.py` for a complete boilerplate.
85
-
86
- **Detailed documentation:** Read `references/data_module.md` for method details and usage patterns.
87
-
88
- ### 4. Callbacks - Extensible Training Logic
89
-
90
- Add custom functionality at specific training hooks without modifying your LightningModule. Built-in callbacks include:
91
-
92
- - **ModelCheckpoint** - Save best/latest models
93
- - **EarlyStopping** - Stop when metrics plateau
94
- - **LearningRateMonitor** - Track LR scheduler changes
95
- - **BatchSizeFinder** - Auto-determine optimal batch size
96
-
97
- **Detailed documentation:** Read `references/callbacks.md` for built-in callbacks and custom callback creation.
98
-
99
- ### 5. Logging - Experiment Tracking
100
-
101
- Integrate with multiple logging platforms:
102
-
103
- - TensorBoard (default)
104
- - Weights & Biases (WandbLogger)
105
- - MLflow (MLFlowLogger)
106
- - Comet (CometLogger)
107
- - CSV (CSVLogger)
108
-
109
- Note: `NeptuneLogger` was removed in lightning 2.6.4. Use W&B, MLflow, or TensorBoard instead.
110
-
111
- Log metrics using `self.log("metric_name", value)` in any LightningModule method.
112
-
113
- **Detailed documentation:** Read `references/logging.md` for logger setup and configuration.
114
-
115
- ### 6. Distributed Training - Scale to Multiple Devices
116
-
117
- Choose the right strategy based on model size:
118
-
119
- - **DDP** - For models <500M parameters (ResNet, smaller transformers)
120
- - **FSDP** - For models 500M+ parameters (large transformers, recommended for Lightning users)
121
- - **DeepSpeed** - For cutting-edge features and fine-grained control
122
-
123
- Configure with: `Trainer(strategy="ddp", accelerator="gpu", devices=4)`
124
-
125
- **Detailed documentation:** Read `references/distributed_training.md` for strategy comparison and configuration.
126
-
127
- ### 7. Best Practices
128
-
129
- - Device agnostic code - Use `self.device` instead of `.cuda()`
130
- - Hyperparameter saving - Use `self.save_hyperparameters()` in `__init__()`
131
- - Metric logging - Use `self.log()` for automatic aggregation across devices
132
- - Reproducibility - Use `seed_everything()` and `Trainer(deterministic=True)`
133
- - Debugging - Use `Trainer(fast_dev_run=True)` to test with 1 batch
134
-
135
- **Detailed documentation:** Read `references/best_practices.md` for common patterns and pitfalls.
136
-
137
- ## Quick Workflow
138
-
139
- 1. **Define model:**
140
- ```python
141
- class MyModel(L.LightningModule):
142
- def __init__(self):
143
- super().__init__()
144
- self.save_hyperparameters()
145
- self.model = YourNetwork()
146
-
147
- def training_step(self, batch, batch_idx):
148
- x, y = batch
149
- loss = F.cross_entropy(self.model(x), y)
150
- self.log("train_loss", loss)
151
- return loss
152
-
153
- def configure_optimizers(self):
154
- return torch.optim.Adam(self.parameters())
155
- ```
156
-
157
- 2. **Prepare data:**
158
- ```python
159
- # Option 1: Direct DataLoaders
160
- train_loader = DataLoader(train_dataset, batch_size=32)
161
-
162
- # Option 2: LightningDataModule (recommended for reusability)
163
- dm = MyDataModule(batch_size=32)
164
- ```
165
-
166
- 3. **Train:**
167
- ```python
168
- trainer = L.Trainer(max_epochs=10, accelerator="gpu", devices=2)
169
- trainer.fit(model, train_loader) # or trainer.fit(model, datamodule=dm)
170
- ```
171
-
172
- ## Resources
173
-
174
- ### scripts/
175
- Executable Python templates for common PyTorch Lightning patterns:
176
-
177
- - `template_lightning_module.py` - Complete LightningModule boilerplate
178
- - `template_datamodule.py` - Complete LightningDataModule boilerplate
179
- - `quick_trainer_setup.py` - Common Trainer configuration examples
180
-
181
- ### references/
182
- Detailed documentation for each PyTorch Lightning component:
183
-
184
- - `lightning_module.md` - Comprehensive LightningModule guide (methods, hooks, properties)
185
- - `trainer.md` - Trainer configuration and parameters
186
- - `data_module.md` - LightningDataModule patterns and methods
187
- - `callbacks.md` - Built-in and custom callbacks
188
- - `logging.md` - Logger integrations and usage
189
- - `distributed_training.md` - DDP, FSDP, DeepSpeed comparison and setup
190
- - `best_practices.md` - Common patterns, tips, and pitfalls
191
-
@@ -1,137 +0,0 @@
1
- ---
2
- name: pyzotero
3
- description: Interact with Zotero reference management libraries using the pyzotero Python client. Retrieve, create, update, and delete items, collections, tags, and attachments via the Zotero Web API v3. Use this skill when working with Zotero libraries programmatically, managing bibliographic references, exporting citations, searching library contents, uploading PDF attachments, or building research automation workflows that integrate with Zotero.
4
- allowed-tools: Read Write Edit Bash
5
- license: MIT License
6
- compatibility: Requires Python 3.10+ and pyzotero 1.13+. Web API access needs a Zotero API key. Optional CLI and MCP extras require Zotero 7 with local API access enabled.
7
- metadata:
8
- version: "1.1"
9
- skill-author: K-Dense Inc.
10
- openclaw:
11
- primaryEnv: ZOTERO_API_KEY
12
- envVars:
13
- - name: ZOTERO_API_KEY
14
- required: true
15
- description: Zotero API key.
16
- - name: ZOTERO_LIBRARY_ID
17
- required: true
18
- description: Zotero library id.
19
- - name: ZOTERO_LIBRARY_TYPE
20
- required: false
21
- description: 'Zotero library type: ''user'' or ''group'' (default ''user'').'
22
- ---
23
-
24
- # Pyzotero
25
-
26
- Pyzotero is a Python wrapper for the [Zotero API v3](https://www.zotero.org/support/dev/web_api/v3/start). Use it to programmatically manage Zotero libraries: read items and collections, create and update references, upload attachments, manage tags, and export citations.
27
-
28
- **Current upstream:** pyzotero 1.13.0 (PyPI, May 2026). Docs: [pyzotero.readthedocs.io](https://pyzotero.readthedocs.io/en/latest/).
29
-
30
- ## Authentication Setup
31
-
32
- **Required credentials** — get from https://www.zotero.org/settings/keys:
33
- - **User ID**: shown as "Your userID for use in API calls"
34
- - **API Key**: create at https://www.zotero.org/settings/keys/new
35
- - **Library ID**: for group libraries, the integer after `/groups/` in the group URL
36
-
37
- Store credentials in environment variables or a `.env` file:
38
- ```
39
- ZOTERO_LIBRARY_ID=your_user_id
40
- ZOTERO_API_KEY=your_api_key
41
- ZOTERO_LIBRARY_TYPE=user # or "group"
42
- ```
43
-
44
- See [references/authentication.md](references/authentication.md) for full setup details.
45
-
46
- ## Installation
47
-
48
- ```bash
49
- uv add pyzotero # Web API client
50
- uv add "pyzotero[cli]" # + local CLI (Zotero 7)
51
- uv add "pyzotero[mcp]" # + MCP server for LLM clients (Zotero 7)
52
- ```
53
-
54
- ## Quick Start
55
-
56
- ```python
57
- import os
58
- from pyzotero import Zotero
59
-
60
- zot = Zotero(
61
- library_id=os.environ['ZOTERO_LIBRARY_ID'],
62
- library_type=os.environ.get('ZOTERO_LIBRARY_TYPE', 'user'),
63
- api_key=os.environ['ZOTERO_API_KEY'],
64
- )
65
-
66
- # Retrieve top-level items (returns 100 by default)
67
- items = zot.top(limit=10)
68
- for item in items:
69
- print(item['data']['title'], item['data']['itemType'])
70
-
71
- # Search by keyword
72
- results = zot.items(q='machine learning', limit=20)
73
-
74
- # Retrieve all items (use everything() for complete results)
75
- all_items = zot.everything(zot.items())
76
- ```
77
-
78
- ## Core Concepts
79
-
80
- - A `Zotero` instance is bound to a single library (user or group). All methods operate on that library.
81
- - Item data lives in `item['data']`. Access fields like `item['data']['title']`, `item['data']['creators']`.
82
- - Pyzotero returns 100 items by default (API default is 25). Use `zot.everything(zot.items())` to get all items.
83
- - Write methods return `True` on success or raise a `ZoteroError`.
84
-
85
- ## Reference Files
86
-
87
- | File | Contents |
88
- |------|----------|
89
- | [references/authentication.md](references/authentication.md) | Credentials, library types, local mode |
90
- | [references/read-api.md](references/read-api.md) | Retrieving items, collections, tags, groups |
91
- | [references/search-params.md](references/search-params.md) | Filtering, sorting, search parameters |
92
- | [references/write-api.md](references/write-api.md) | Creating, updating, deleting items |
93
- | [references/collections.md](references/collections.md) | Collection CRUD operations |
94
- | [references/tags.md](references/tags.md) | Tag access and management |
95
- | [references/files-attachments.md](references/files-attachments.md) | File download and attachment uploads |
96
- | [references/exports.md](references/exports.md) | BibTeX, CSL-JSON, bibliography export |
97
- | [references/pagination.md](references/pagination.md) | follow(), everything(), generators |
98
- | [references/full-text.md](references/full-text.md) | Full-text content indexing and access |
99
- | [references/saved-searches.md](references/saved-searches.md) | Saved search management |
100
- | [references/cli.md](references/cli.md) | Command-line interface (local Zotero 7) |
101
- | [references/mcp.md](references/mcp.md) | MCP server for LLM clients (local Zotero 7) |
102
- | [references/error-handling.md](references/error-handling.md) | Errors and exception handling |
103
-
104
- ## Common Patterns
105
-
106
- ### Fetch and modify an item
107
- ```python
108
- item = zot.item('ITEMKEY')
109
- item['data']['title'] = 'New Title'
110
- zot.update_item(item)
111
- ```
112
-
113
- ### Create an item from a template
114
- ```python
115
- template = zot.item_template('journalArticle')
116
- template['title'] = 'My Paper'
117
- template['creators'][0] = {'creatorType': 'author', 'firstName': 'Jane', 'lastName': 'Doe'}
118
- zot.create_items([template])
119
- ```
120
-
121
- ### Export as BibTeX
122
- ```python
123
- zot.add_parameters(format='bibtex')
124
- bibtex = zot.top(limit=50)
125
- # bibtex is a bibtexparser BibDatabase object
126
- print(bibtex.entries)
127
- ```
128
-
129
- ### Local mode (read-only, no API key needed)
130
- ```python
131
- zot = Zotero(library_id='123456', library_type='user', local=True)
132
- items = zot.items()
133
- ```
134
-
135
- ### Local Zotero 7 (CLI or MCP, no API key)
136
-
137
- For searching a locally running Zotero desktop app (including full-text PDF search), use the CLI or MCP server instead of the Web API. Both require Zotero 7 with local API access enabled. See [references/cli.md](references/cli.md) and [references/mcp.md](references/mcp.md).