@pikaa-ai/pikaa 0.3.22 → 0.3.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +448 -181
  6. package/dist/index.js +22 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,324 +0,0 @@
1
- ---
2
- name: scikit-learn
3
- description: Machine learning in Python with scikit-learn. Use when working with supervised learning (classification, regression), unsupervised learning (clustering, dimensionality reduction), model evaluation, hyperparameter tuning, preprocessing, or building ML pipelines. Provides comprehensive reference documentation for algorithms, preprocessing techniques, pipelines, and best practices.
4
- license: BSD-3-Clause license
5
- allowed-tools: Read Write Edit Bash
6
- compatibility: Requires Python 3.11+ and scikit-learn 1.7+. NumPy and SciPy are required dependencies. Optional matplotlib/seaborn for bundled example scripts that save plots.
7
- metadata:
8
- version: "1.2"
9
- skill-author: K-Dense Inc.
10
- ---
11
-
12
- # Scikit-learn
13
-
14
- ## Overview
15
-
16
- This skill provides comprehensive guidance for machine learning tasks using scikit-learn, the industry-standard Python library for classical machine learning. Use this skill for classification, regression, clustering, dimensionality reduction, preprocessing, model evaluation, and building production-ready ML pipelines.
17
-
18
- ## Installation
19
-
20
- Tested against **scikit-learn 1.8.0** (stable; December 2025). Requires **Python 3.11–3.14** (free-threaded CPython 3.14 wheels available in 1.8+).
21
-
22
- Install the PyPI package **`scikit-learn`** (not the deprecated `sklearn` package on PyPI). Import in code as `sklearn`.
23
-
24
- ```bash
25
- # Install scikit-learn using uv
26
- uv pip install "scikit-learn>=1.7"
27
-
28
- # Optional: plotting utilities and bundled script dependencies
29
- uv pip install "scikit-learn[plots]" matplotlib seaborn
30
-
31
- # Commonly used with
32
- uv pip install pandas numpy
33
- ```
34
-
35
- Check your version:
36
-
37
- ```python
38
- import sklearn
39
- print(sklearn.__version__)
40
- ```
41
-
42
- ## When to Use This Skill
43
-
44
- Use the scikit-learn skill when:
45
-
46
- - Building classification or regression models
47
- - Performing clustering or dimensionality reduction
48
- - Preprocessing and transforming data for machine learning
49
- - Evaluating model performance with cross-validation
50
- - Tuning hyperparameters with grid or random search
51
- - Creating ML pipelines for production workflows
52
- - Comparing different algorithms for a task
53
- - Working with both structured (tabular) and text data
54
- - Need interpretable, classical machine learning approaches
55
-
56
- ## Quick Start
57
-
58
- ### Classification Example
59
-
60
- ```python
61
- from sklearn.model_selection import train_test_split
62
- from sklearn.preprocessing import StandardScaler
63
- from sklearn.ensemble import RandomForestClassifier
64
- from sklearn.metrics import classification_report
65
-
66
- # Split data
67
- X_train, X_test, y_train, y_test = train_test_split(
68
- X, y, test_size=0.2, stratify=y, random_state=42
69
- )
70
-
71
- # Preprocess
72
- scaler = StandardScaler()
73
- X_train_scaled = scaler.fit_transform(X_train)
74
- X_test_scaled = scaler.transform(X_test)
75
-
76
- # Train model
77
- model = RandomForestClassifier(n_estimators=100, random_state=42)
78
- model.fit(X_train_scaled, y_train)
79
-
80
- # Evaluate
81
- y_pred = model.predict(X_test_scaled)
82
- print(classification_report(y_test, y_pred))
83
- ```
84
-
85
- ### Complete Pipeline with Mixed Data
86
-
87
- ```python
88
- from sklearn.pipeline import Pipeline
89
- from sklearn.compose import ColumnTransformer
90
- from sklearn.preprocessing import StandardScaler, OneHotEncoder
91
- from sklearn.impute import SimpleImputer
92
- from sklearn.ensemble import GradientBoostingClassifier
93
-
94
- # Define feature types
95
- numeric_features = ['age', 'income']
96
- categorical_features = ['gender', 'occupation']
97
-
98
- # Create preprocessing pipelines
99
- numeric_transformer = Pipeline([
100
- ('imputer', SimpleImputer(strategy='median')),
101
- ('scaler', StandardScaler())
102
- ])
103
-
104
- categorical_transformer = Pipeline([
105
- ('imputer', SimpleImputer(strategy='most_frequent')),
106
- ('onehot', OneHotEncoder(handle_unknown='ignore'))
107
- ])
108
-
109
- # Combine transformers
110
- preprocessor = ColumnTransformer([
111
- ('num', numeric_transformer, numeric_features),
112
- ('cat', categorical_transformer, categorical_features)
113
- ])
114
-
115
- # Full pipeline
116
- model = Pipeline([
117
- ('preprocessor', preprocessor),
118
- ('classifier', GradientBoostingClassifier(random_state=42))
119
- ])
120
-
121
- # Fit and predict
122
- model.fit(X_train, y_train)
123
- y_pred = model.predict(X_test)
124
- ```
125
-
126
- ## Core Capabilities
127
-
128
- Five capability areas are documented in
129
- [references/core_capabilities.md](references/core_capabilities.md), with per-topic detail
130
- in [references/supervised_learning.md](references/supervised_learning.md),
131
- [references/unsupervised_learning.md](references/unsupervised_learning.md),
132
- [references/model_evaluation.md](references/model_evaluation.md),
133
- [references/preprocessing.md](references/preprocessing.md), and
134
- [references/pipelines_and_composition.md](references/pipelines_and_composition.md):
135
-
136
- 1. **Supervised learning** — classification and regression estimator families.
137
- 2. **Unsupervised learning** — clustering, decomposition, and manifold learning.
138
- 3. **Model evaluation and selection** — metrics, cross-validation, and hyperparameter search.
139
- 4. **Data preprocessing** — scaling, encoding, imputation, and feature selection.
140
- 5. **Pipelines and composition** — `Pipeline` and `ColumnTransformer`.
141
-
142
- Always fit preprocessing inside a `Pipeline` so it is refit per cross-validation fold;
143
- scaling or imputing before splitting leaks test information into training.
144
-
145
- Two worked workflows are in
146
- [references/common_workflows.md](references/common_workflows.md).
147
-
148
- ## Example Scripts
149
-
150
- ### Classification Pipeline
151
-
152
- Run a complete classification workflow with preprocessing, model comparison, hyperparameter tuning, and evaluation:
153
-
154
- ```bash
155
- uv run python scripts/classification_pipeline.py
156
- ```
157
-
158
- This script demonstrates:
159
- - Handling mixed data types (numeric and categorical)
160
- - Model comparison using cross-validation
161
- - Hyperparameter tuning with GridSearchCV
162
- - Comprehensive evaluation with multiple metrics
163
- - Feature importance analysis
164
-
165
- ### Clustering Analysis
166
-
167
- Perform clustering analysis with algorithm comparison and visualization:
168
-
169
- ```bash
170
- uv run python scripts/clustering_analysis.py
171
- ```
172
-
173
- This script demonstrates:
174
- - Finding optimal number of clusters (elbow method, silhouette analysis)
175
- - Comparing multiple clustering algorithms (K-Means, DBSCAN, Agglomerative, Gaussian Mixture)
176
- - Evaluating clustering quality without ground truth
177
- - Visualizing results with PCA projection
178
-
179
- ## Reference Documentation
180
-
181
- This skill includes comprehensive reference files for deep dives into specific topics:
182
-
183
- ### Quick Reference
184
- **File:** `references/quick_reference.md`
185
- - Common import patterns and installation instructions
186
- - Quick workflow templates for common tasks
187
- - Algorithm selection cheat sheets
188
- - Common patterns and gotchas
189
- - Performance optimization tips
190
-
191
- ### Supervised Learning
192
- **File:** `references/supervised_learning.md`
193
- - Linear models (regression and classification)
194
- - Support Vector Machines
195
- - Decision Trees and ensemble methods
196
- - K-Nearest Neighbors, Naive Bayes, Neural Networks
197
- - Algorithm selection guide
198
-
199
- ### Unsupervised Learning
200
- **File:** `references/unsupervised_learning.md`
201
- - All clustering algorithms with parameters and use cases
202
- - Dimensionality reduction techniques
203
- - Outlier and novelty detection
204
- - Gaussian Mixture Models
205
- - Method selection guide
206
-
207
- ### Model Evaluation
208
- **File:** `references/model_evaluation.md`
209
- - Cross-validation strategies
210
- - Hyperparameter tuning methods
211
- - Classification, regression, and clustering metrics
212
- - Learning and validation curves
213
- - Best practices for model selection
214
-
215
- ### Preprocessing
216
- **File:** `references/preprocessing.md`
217
- - Feature scaling and normalization
218
- - Encoding categorical variables
219
- - Missing value imputation
220
- - Feature engineering techniques
221
- - Custom transformers
222
-
223
- ### Pipelines and Composition
224
- **File:** `references/pipelines_and_composition.md`
225
- - Pipeline construction and usage
226
- - ColumnTransformer for mixed data types
227
- - FeatureUnion for parallel transformations
228
- - Complete end-to-end examples
229
- - Best practices
230
-
231
- ## Best Practices
232
-
233
- ### Always Use Pipelines
234
- Pipelines prevent data leakage and ensure consistency:
235
- ```python
236
- # Good: Preprocessing in pipeline
237
- pipeline = Pipeline([
238
- ('scaler', StandardScaler()),
239
- ('model', LogisticRegression())
240
- ])
241
-
242
- # Bad: Preprocessing outside (can leak information)
243
- X_scaled = StandardScaler().fit_transform(X)
244
- ```
245
-
246
- ### Fit on Training Data Only
247
- Never fit on test data:
248
- ```python
249
- # Good
250
- scaler = StandardScaler()
251
- X_train_scaled = scaler.fit_transform(X_train)
252
- X_test_scaled = scaler.transform(X_test) # Only transform
253
-
254
- # Bad
255
- scaler = StandardScaler()
256
- X_all_scaled = scaler.fit_transform(np.vstack([X_train, X_test]))
257
- ```
258
-
259
- ### Use Stratified Splitting for Classification
260
- Preserve class distribution:
261
- ```python
262
- X_train, X_test, y_train, y_test = train_test_split(
263
- X, y, test_size=0.2, stratify=y, random_state=42
264
- )
265
- ```
266
-
267
- ### Set Random State for Reproducibility
268
- ```python
269
- model = RandomForestClassifier(n_estimators=100, random_state=42)
270
- ```
271
-
272
- ### Choose Appropriate Metrics
273
- - Balanced data: Accuracy, F1-score
274
- - Imbalanced data: Precision, Recall, ROC AUC, Balanced Accuracy
275
- - Cost-sensitive: Define custom scorer
276
-
277
- ### Scale Features When Required
278
- Algorithms requiring feature scaling:
279
- - SVM, KNN, Neural Networks
280
- - PCA, Linear/Logistic Regression with regularization
281
- - K-Means clustering
282
-
283
- Algorithms not requiring scaling:
284
- - Tree-based models (Decision Trees, Random Forest, Gradient Boosting)
285
- - Naive Bayes
286
-
287
- ## Troubleshooting Common Issues
288
-
289
- ### ConvergenceWarning
290
- **Issue:** Model didn't converge
291
- **Solution:** Increase `max_iter` or scale features
292
- ```python
293
- model = LogisticRegression(max_iter=1000)
294
- ```
295
-
296
- ### Poor Performance on Test Set
297
- **Issue:** Overfitting
298
- **Solution:** Use regularization, cross-validation, or simpler model
299
- ```python
300
- # Add regularization
301
- model = Ridge(alpha=1.0)
302
-
303
- # Use cross-validation
304
- scores = cross_val_score(model, X, y, cv=5)
305
- ```
306
-
307
- ### Memory Error with Large Datasets
308
- **Solution:** Use algorithms designed for large data
309
- ```python
310
- # Use SGD for large datasets
311
- from sklearn.linear_model import SGDClassifier
312
- model = SGDClassifier()
313
-
314
- # Or MiniBatchKMeans for clustering
315
- from sklearn.cluster import MiniBatchKMeans
316
- model = MiniBatchKMeans(n_clusters=8, batch_size=100)
317
- ```
318
-
319
- ## Additional Resources
320
-
321
- - Official Documentation: https://scikit-learn.org/stable/
322
- - User Guide: https://scikit-learn.org/stable/user_guide.html
323
- - API Reference: https://scikit-learn.org/stable/api/index.html
324
- - Examples Gallery: https://scikit-learn.org/stable/auto_examples/index.html
@@ -1,313 +0,0 @@
1
- ---
2
- name: scikit-survival
3
- description: Build, evaluate, and audit right-censored or competing-risk survival workflows with scikit-survival, including leakage-safe preprocessing, model selection, probability prediction, and censoring-aware metrics.
4
- license: MIT
5
- compatibility: Requires Python 3.11+, uv, and the pinned scikit-survival 0.28.0 stack for executable examples. Bundled CLIs are local and network-free by default.
6
- allowed-tools: Read Write Edit Bash
7
- metadata:
8
- version: "1.1"
9
- skill-author: K-Dense Inc.
10
- ---
11
-
12
- # scikit-survival
13
-
14
- ## Scope
15
-
16
- Use this skill for scikit-survival 0.28.0 workflows involving:
17
-
18
- - right-censored structured outcomes;
19
- - Cox PH, Coxnet, IPC ridge, survival trees, forests, boosting, and SVMs;
20
- - discrimination, prediction error, calibration-oriented checks, and time-dependent prediction;
21
- - nonparametric cumulative incidence with competing risks;
22
- - scikit-learn pipelines, nested model selection, and reproducible reports.
23
-
24
- scikit-survival primarily models right-censored outcomes. Its built-in competing-risk
25
- support is nonparametric cumulative incidence; it does not provide Fine-Gray regression.
26
- Do not present model output as clinical advice, causal evidence, or proof of clinical
27
- utility.
28
-
29
- ## Current release and installation
30
-
31
- Verified 2026-07-23:
32
-
33
- - Latest stable: **scikit-survival 0.28.0**, released 2026-07-05.
34
- - Python: **3.11 or later**; PyPI wheels cover CPython 3.11-3.14 on Linux
35
- x86-64, macOS x86-64/ARM64, and Windows x86-64.
36
- - Runtime bounds: NumPy >=2.0.0, pandas >=2.2.0, SciPy >=1.13.0,
37
- scikit-learn >=1.9.0,<1.10, OSQP >=1.0.2, narwhals >=2.0.1.
38
- - 0.28 adds pandas/Polars estimator support through narwhals and removes
39
- `criterion` from `GradientBoostingSurvivalAnalysis`.
40
-
41
- Create an isolated environment and install the tested snapshot:
42
-
43
- ```bash
44
- uv venv --python 3.11
45
- source .venv/bin/activate
46
- uv pip install \
47
- "scikit-survival==0.28.0" \
48
- "scikit-learn==1.9.0" \
49
- "numpy==2.4.6" \
50
- "pandas==3.0.5" \
51
- "scipy==1.17.1" \
52
- "ecos==2.0.14" \
53
- "osqp==1.1.3" \
54
- "joblib==1.5.3" \
55
- "numexpr==2.14.2" \
56
- "narwhals==2.24.0"
57
- ```
58
-
59
- Binary wheels are preferred. A source build requires a C/C++ compiler; OSQP may
60
- also require CMake. This skill is MIT-licensed; the upstream scikit-survival package
61
- is GPL-3.0-or-later, so review upstream licensing before redistribution.
62
-
63
- ## Non-negotiable workflow
64
-
65
- 1. **Define the estimand and event coding.** Decide whether the target is
66
- all-event survival, cause-specific hazard, or cause-specific cumulative incidence.
67
- 2. **Validate outcomes.** Standard estimators need a two-field structured array:
68
- boolean event first, observed time second. Competing-risk CIF instead needs a
69
- separate integer event vector: 0=censored, 1..K=causes.
70
- 3. **Split before learned preprocessing.** Never fit imputers, encoders, scalers,
71
- feature selectors, or alpha choices on all rows before splitting.
72
- 4. **Fit preprocessing inside a pipeline.** Unknown categories and missingness must
73
- be handled using training-fold state only.
74
- 5. **Tune without reusing evaluation data.** Use nested CV when reporting
75
- cross-validated tuned performance, or reserve a truly untouched final holdout.
76
- 6. **Fit censoring distributions on training data.** IPCW concordance, dynamic AUC,
77
- and Brier metrics receive `survival_train`, never a pooled train+test outcome.
78
- 7. **Restrict evaluation times.** Use a strictly increasing grid inside test
79
- follow-up and below the end of training support where the estimated censoring
80
- survival remains positive.
81
- 8. **Match predictions to metrics.** Concordance/dynamic AUC consume higher-is-riskier
82
- scores. Brier metrics consume survival probabilities with shape
83
- `(n_test, n_times)`, not risk scores or unevaluated step functions.
84
- 9. **Handle competing causes explicitly.** Standard survival probabilities and CIFs
85
- answer different questions. Never estimate event-specific probability with
86
- `1 - Kaplan-Meier` while censoring competing events.
87
- 10. **Report limits.** Separate discrimination, calibration, prediction error,
88
- and cumulative incidence. None alone establishes decision or clinical utility.
89
-
90
- ## Outcome construction
91
-
92
- ```python
93
- from sksurv.util import Surv
94
-
95
- y = Surv.from_arrays(event=event_bool, time=observed_time)
96
- # Equivalent for pandas or Polars:
97
- y = Surv.from_dataframe("event", "time", frame)
98
- ```
99
-
100
- The first field is boolean (`True`=event, `False`=right-censored); the second is
101
- floating-point time. Field names may vary, but field order and meaning may not.
102
- Use `references/data-handling.md` before loading custom or competing-risk data.
103
-
104
- ## Leakage-safe pipeline
105
-
106
- ```python
107
- from sklearn.compose import ColumnTransformer
108
- from sklearn.impute import SimpleImputer
109
- from sklearn.model_selection import train_test_split
110
- from sklearn.pipeline import make_pipeline
111
- from sklearn.preprocessing import OneHotEncoder, StandardScaler
112
- from sksurv.linear_model import CoxPHSurvivalAnalysis
113
-
114
- X_train, X_test, y_train, y_test = train_test_split(
115
- X, y, test_size=0.25, stratify=y["event"], random_state=20260723
116
- )
117
-
118
- preprocess = ColumnTransformer(
119
- [
120
- ("num", make_pipeline(SimpleImputer(strategy="median"), StandardScaler()), numeric),
121
- (
122
- "cat",
123
- make_pipeline(
124
- SimpleImputer(strategy="most_frequent"),
125
- OneHotEncoder(handle_unknown="ignore", drop="first", sparse_output=False),
126
- ),
127
- categorical,
128
- ),
129
- ],
130
- sparse_threshold=0.0,
131
- )
132
- model = make_pipeline(preprocess, CoxPHSurvivalAnalysis(alpha=0.1, ties="efron"))
133
- model.fit(X_train, y_train)
134
- risk = model.predict(X_test)
135
- ```
136
-
137
- The split precedes every learned transformation. For repeated or grouped records,
138
- use a group-aware split; for temporal deployment, use a time-respecting split.
139
-
140
- ## Model choice
141
-
142
- - `CoxPHSurvivalAnalysis`: interpretable log-hazard coefficients under proportional
143
- hazards; `alpha` is ridge shrinkage and `ties` is `"breslow"` or `"efron"`.
144
- - `CoxnetSurvivalAnalysis`: LASSO/elastic-net path for high-dimensional data.
145
- `l1_ratio` is in `(0, 1]`; use `fit_baseline_model=True` before requesting
146
- survival or cumulative-hazard functions.
147
- - `IPCRidge`: IPC-weighted ridge AFT model; prediction is on a time/log-time scale,
148
- not a Cox risk score.
149
- - `RandomSurvivalForest` / `ExtraSurvivalTrees`: nonlinear survival and cumulative
150
- hazard predictions; use permutation importance, not impurity importance.
151
- - `GradientBoostingSurvivalAnalysis`: tree boosting with `"coxph"`, `"squared"`,
152
- or `"ipcwls"` loss. `criterion` was removed in 0.28.
153
- - `ComponentwiseGradientBoostingSurvivalAnalysis`: sparse linear componentwise
154
- boosting.
155
- - `FastSurvivalSVM` / `FastKernelSurvivalSVM`: ranking or regression objectives.
156
- Only `rank_ratio=1` directly returns higher-is-riskier scores; SVMs do not yield
157
- survival probabilities for Brier metrics.
158
-
159
- Read the model-specific reference before interpreting coefficients or predictions:
160
- `references/cox-models.md`, `references/ensemble-models.md`, or
161
- `references/svm-models.md`.
162
-
163
- ## Prediction and metric contracts
164
-
165
- ```python
166
- import numpy as np
167
- from sksurv.metrics import (
168
- brier_score,
169
- concordance_index_ipcw,
170
- cumulative_dynamic_auc,
171
- integrated_brier_score,
172
- )
173
-
174
- risk = model.predict(X_test) # (n_test,), higher means higher event risk
175
- uno_c = concordance_index_ipcw(y_train, y_test, risk, tau=times[-1])[0]
176
- auc_t, mean_auc = cumulative_dynamic_auc(y_train, y_test, risk, times)
177
-
178
- surv_fns = model.predict_survival_function(X_test)
179
- surv_prob = np.vstack([fn(times) for fn in surv_fns]) # (n_test, n_times)
180
- _, brier_t = brier_score(y_train, y_test, surv_prob, times)
181
- ibs = integrated_brier_score(y_train, y_test, surv_prob, times)
182
- ```
183
-
184
- - Harrell C and Uno C measure rank discrimination, not calibration.
185
- - Cumulative/dynamic AUC measures discrimination at selected horizons and accepts
186
- 1D or time-dependent 2D risk scores; it rejects survival probabilities.
187
- - Brier score is censoring-weighted probability error and reflects both
188
- discrimination and calibration. It is not a standalone calibration curve.
189
- - Calibration requires horizon-specific predicted-versus-observed checks on
190
- independent data. scikit-survival 0.28 has no dedicated calibration-curve API.
191
-
192
- See `references/evaluation-metrics.md` for assumptions, primary literature, safe
193
- time-grid construction, and scorer wrappers.
194
-
195
- ## Pipelines, metadata routing, and tuning
196
-
197
- Ordinary `Pipeline.fit(X, y)` needs no metadata-routing setup. Metric wrappers such
198
- as `as_concordance_index_ipcw_scorer` are estimator wrappers, not `scoring=`
199
- callables:
200
-
201
- ```python
202
- from sklearn.model_selection import GridSearchCV
203
- from sksurv.metrics import as_concordance_index_ipcw_scorer
204
-
205
- wrapped = as_concordance_index_ipcw_scorer(model, tau=tau)
206
- search = GridSearchCV(
207
- wrapped,
208
- {"estimator__coxphsurvivalanalysis__alpha": [0.01, 0.1, 1.0]},
209
- cv=inner_splits,
210
- )
211
- ```
212
-
213
- The wrapper learns the censoring distribution from each fit fold. Prefix wrapped
214
- parameters with `estimator__`. Enable scikit-learn metadata routing only when
215
- passing extra metadata through a meta-estimator. For example, Coxnet's
216
- `set_predict_request(alpha=True)` matters only when routing the `alpha` prediction
217
- argument with `sklearn.set_config(enable_metadata_routing=True)`.
218
-
219
- Use an outer CV loop for an unbiased CV performance estimate after inner tuning.
220
- Do not select parameters and report performance from the same folds as if external.
221
-
222
- ## Competing risks
223
-
224
- ```python
225
- from sksurv.nonparametric import cumulative_incidence_competing_risks
226
-
227
- # status: integer array, 0=censored, 1..K=mutually exclusive causes
228
- time, cif = cumulative_incidence_competing_risks(status, observed_time)
229
- total_cif = cif[0]
230
- cause_1_cif = cif[1]
231
- ```
232
-
233
- `cif` has shape `(K + 1, n_times)`; row 0 is total risk and rows 1..K are
234
- cause-specific cumulative incidence. Cause-specific Cox models treat other causes
235
- as censored to estimate cause-specific hazards, but one such model's
236
- `1 - survival` is not the cause-specific CIF. See `references/competing-risks.md`.
237
-
238
- ## Bundled local CLIs
239
-
240
- All helpers use deterministic synthetic data when no input is given. They make no
241
- network calls, reject URLs and symlinks, bound files/rows/features, avoid unsafe
242
- pickle loading, and lazily import scientific packages.
243
-
244
- ```bash
245
- python skills/scikit-survival/scripts/validate_survival_csv.py --help
246
- python skills/scikit-survival/scripts/train_survival_model.py --help
247
- python skills/scikit-survival/scripts/evaluate_survival_metrics.py --help
248
- python skills/scikit-survival/scripts/competing_risk_cif.py --help
249
- python skills/scikit-survival/scripts/model_report.py --help
250
- ```
251
-
252
- Typical local flow:
253
-
254
- ```bash
255
- python skills/scikit-survival/scripts/validate_survival_csv.py \
256
- --input data.csv --event-column event --time-column time \
257
- --feature-columns age,group,measurement --structured-output outcome.npy
258
-
259
- python skills/scikit-survival/scripts/train_survival_model.py \
260
- --input data.csv --event-column event --time-column time \
261
- --numeric-columns age,measurement --categorical-columns group \
262
- --model coxph --tune --prediction-output predictions.npz \
263
- --output training-summary.json
264
-
265
- python skills/scikit-survival/scripts/evaluate_survival_metrics.py \
266
- --input predictions.npz --output metrics-summary.json
267
-
268
- python skills/scikit-survival/scripts/model_report.py \
269
- --training-summary training-summary.json \
270
- --metrics-summary metrics-summary.json --output model-report.md
271
- ```
272
-
273
- Use only de-identified, authorized local data. The bundled tests contain synthetic
274
- records only and no patient data or PHI.
275
-
276
- ## Security triage
277
-
278
- `SECURITY.md` previously claimed this skill bundled package-shadowing files named
279
- `sklearn.py` and `sksurv.py`. The 2026-07-23 inventory confirmed those files did
280
- not exist; the claim was a phantom analyzer finding. This refresh adds only
281
- descriptively named helpers and no shadow modules, environment reads, or network
282
- calls.
283
-
284
- Never name a project script after an imported package (including `sklearn.py`,
285
- `sksurv.py`, `numpy.py`, or `pandas.py`), because Python may import the local file
286
- instead of the installed library. Inspect the working directory before executing
287
- examples copied from untrusted sources.
288
-
289
- ## Reference files
290
-
291
- - `references/data-handling.md` — structured arrays, datasets, schema validation,
292
- pandas/Polars preprocessing, and leakage-safe splitting.
293
- - `references/cox-models.md` — Cox PH, Coxnet, IPCRidge, assumptions, and tuning.
294
- - `references/ensemble-models.md` — forests, trees, boosting, predictions, and
295
- permutation importance.
296
- - `references/svm-models.md` — SVM objectives, prediction direction, scaling,
297
- kernels, and limitations.
298
- - `references/evaluation-metrics.md` — metric inputs, censoring assumptions,
299
- time grids, calibration, nested CV, and primary literature.
300
- - `references/competing-risks.md` — integer event coding, CIF API, built-in
301
- datasets, cause-specific hazards, and unsupported Fine-Gray regression.
302
-
303
- ## Dated sources
304
-
305
- Official API and compatibility sources, checked 2026-07-23:
306
-
307
- - [PyPI 0.28.0](https://pypi.org/project/scikit-survival/) — released 2026-07-05.
308
- - [GitHub v0.28.0 release](https://github.com/sebp/scikit-survival/releases/tag/v0.28.0)
309
- — published 2026-07-05.
310
- - [0.28 release notes](https://scikit-survival.readthedocs.io/en/stable/release_notes/v0.28.html).
311
- - [Installation guide](https://scikit-survival.readthedocs.io/en/stable/install.html).
312
- - [Stable user guide](https://scikit-survival.readthedocs.io/en/stable/user_guide/index.html).
313
- - [Stable API reference](https://scikit-survival.readthedocs.io/en/stable/api/index.html).