@pikaa-ai/pikaa 0.3.22 → 0.3.24

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (191) hide show
  1. package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
  2. package/assets/brand/orbit-logo.jpg +0 -0
  3. package/assets/brand/orbit-logo.png +0 -0
  4. package/assets/brand/orbit-logo.svg +3 -0
  5. package/dist/cli.js +448 -181
  6. package/dist/index.js +22 -2
  7. package/package.json +1 -2
  8. package/skills/adaptyv/SKILL.md +0 -240
  9. package/skills/aeon/SKILL.md +0 -402
  10. package/skills/analytical-method-validation/SKILL.md +0 -299
  11. package/skills/anndata/SKILL.md +0 -431
  12. package/skills/arbor/SKILL.md +0 -152
  13. package/skills/arboreto/SKILL.md +0 -267
  14. package/skills/astropy/SKILL.md +0 -353
  15. package/skills/autoskill/SKILL.md +0 -233
  16. package/skills/benchling-integration/SKILL.md +0 -229
  17. package/skills/bgpt-paper-search/SKILL.md +0 -75
  18. package/skills/bids/SKILL.md +0 -237
  19. package/skills/biopython/SKILL.md +0 -472
  20. package/skills/bioservices/SKILL.md +0 -399
  21. package/skills/bulk-rnaseq/SKILL.md +0 -198
  22. package/skills/cellxgene-census/SKILL.md +0 -283
  23. package/skills/cirq/SKILL.md +0 -370
  24. package/skills/citation-management/SKILL.md +0 -329
  25. package/skills/clinical-decision-support/SKILL.md +0 -238
  26. package/skills/clinical-decision-support/references/README.md +0 -62
  27. package/skills/clinical-reports/SKILL.md +0 -248
  28. package/skills/clinical-reports/references/README.md +0 -34
  29. package/skills/cobrapy/SKILL.md +0 -496
  30. package/skills/consciousness-council/SKILL.md +0 -151
  31. package/skills/dask/SKILL.md +0 -482
  32. package/skills/database-lookup/SKILL.md +0 -386
  33. package/skills/datamol/SKILL.md +0 -200
  34. package/skills/deepchem/SKILL.md +0 -244
  35. package/skills/deepspot-m/SKILL.md +0 -175
  36. package/skills/deeptools/SKILL.md +0 -412
  37. package/skills/depmap/SKILL.md +0 -301
  38. package/skills/dhdna-profiler/SKILL.md +0 -184
  39. package/skills/diffdock/SKILL.md +0 -488
  40. package/skills/dnanexus-integration/SKILL.md +0 -325
  41. package/skills/docx/SKILL.md +0 -99
  42. package/skills/esm/SKILL.md +0 -334
  43. package/skills/etetoolkit/SKILL.md +0 -327
  44. package/skills/exa-search/SKILL.md +0 -102
  45. package/skills/executing-plans/SKILL.md +0 -14
  46. package/skills/experimental-design/SKILL.md +0 -234
  47. package/skills/exploratory-data-analysis/SKILL.md +0 -280
  48. package/skills/flowio/SKILL.md +0 -310
  49. package/skills/fluidsim/SKILL.md +0 -279
  50. package/skills/frontend-design/SKILL.md +0 -100
  51. package/skills/generate-image/SKILL.md +0 -304
  52. package/skills/geniml/SKILL.md +0 -310
  53. package/skills/genomic-coordinates/SKILL.md +0 -189
  54. package/skills/genomic-intelligence/SKILL.md +0 -243
  55. package/skills/geomaster/README.md +0 -105
  56. package/skills/geomaster/SKILL.md +0 -366
  57. package/skills/geopandas/SKILL.md +0 -250
  58. package/skills/get-available-resources/SKILL.md +0 -260
  59. package/skills/gget/SKILL.md +0 -153
  60. package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
  61. package/skills/glycoengineering/SKILL.md +0 -339
  62. package/skills/gtars/SKILL.md +0 -282
  63. package/skills/guardian-rails/SKILL.md +0 -54
  64. package/skills/histolab/SKILL.md +0 -243
  65. package/skills/hugging-science/SKILL.md +0 -132
  66. package/skills/hypogenic/SKILL.md +0 -290
  67. package/skills/hypothesis-generation/SKILL.md +0 -264
  68. package/skills/imaging-data-commons/SKILL.md +0 -496
  69. package/skills/infographics/SKILL.md +0 -315
  70. package/skills/iso-standards-readiness/SKILL.md +0 -352
  71. package/skills/lab-hardware-cad/SKILL.md +0 -372
  72. package/skills/labarchive-integration/SKILL.md +0 -216
  73. package/skills/lamindb/SKILL.md +0 -408
  74. package/skills/latchbio-integration/SKILL.md +0 -227
  75. package/skills/latex-posters/SKILL.md +0 -369
  76. package/skills/latex-posters/references/README.md +0 -439
  77. package/skills/liteparse/SKILL.md +0 -295
  78. package/skills/literature-review/SKILL.md +0 -263
  79. package/skills/markdown-mermaid-writing/SKILL.md +0 -322
  80. package/skills/market-research-reports/SKILL.md +0 -337
  81. package/skills/markitdown/SKILL.md +0 -264
  82. package/skills/matchms/SKILL.md +0 -276
  83. package/skills/matlab/SKILL.md +0 -274
  84. package/skills/matplotlib/SKILL.md +0 -378
  85. package/skills/medchem/SKILL.md +0 -321
  86. package/skills/modal/SKILL.md +0 -468
  87. package/skills/molecular-dynamics/SKILL.md +0 -458
  88. package/skills/molfeat/SKILL.md +0 -348
  89. package/skills/ncats-arax/SKILL.md +0 -178
  90. package/skills/networkx/SKILL.md +0 -440
  91. package/skills/neurokit2/SKILL.md +0 -323
  92. package/skills/neuropixels-analysis/SKILL.md +0 -412
  93. package/skills/nextflow/SKILL.md +0 -195
  94. package/skills/omero-integration/SKILL.md +0 -222
  95. package/skills/onekgpd/SKILL.md +0 -371
  96. package/skills/ontology-term-resolution/SKILL.md +0 -147
  97. package/skills/open-notebook/SKILL.md +0 -297
  98. package/skills/openpiv/SKILL.md +0 -469
  99. package/skills/opentrons-integration/SKILL.md +0 -322
  100. package/skills/optimize-for-gpu/SKILL.md +0 -176
  101. package/skills/owasp-top10/SKILL.md +0 -48
  102. package/skills/pacsomatic/LICENSE +0 -21
  103. package/skills/pacsomatic/SKILL.md +0 -150
  104. package/skills/paper-lookup/SKILL.md +0 -263
  105. package/skills/paperclip/SKILL.md +0 -413
  106. package/skills/paperzilla/SKILL.md +0 -159
  107. package/skills/parallel-web/SKILL.md +0 -128
  108. package/skills/pathml/SKILL.md +0 -222
  109. package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
  110. package/skills/pathway-enrichment/SKILL.md +0 -194
  111. package/skills/pdf/SKILL.md +0 -322
  112. package/skills/peer-review/SKILL.md +0 -288
  113. package/skills/penetration-testing/SKILL.md +0 -31
  114. package/skills/pennylane/SKILL.md +0 -240
  115. package/skills/phylogenetics/SKILL.md +0 -409
  116. package/skills/pi-agent/SKILL.md +0 -83
  117. package/skills/pkpd-modeling/SKILL.md +0 -381
  118. package/skills/polars/SKILL.md +0 -393
  119. package/skills/polars-bio/SKILL.md +0 -379
  120. package/skills/ponytail/SKILL.md +0 -31
  121. package/skills/ponytail-audit/SKILL.md +0 -18
  122. package/skills/pptx/SKILL.md +0 -246
  123. package/skills/pptx-posters/SKILL.md +0 -258
  124. package/skills/primekg/SKILL.md +0 -99
  125. package/skills/protocolsio-integration/SKILL.md +0 -236
  126. package/skills/pufferlib/SKILL.md +0 -328
  127. package/skills/pydeseq2/SKILL.md +0 -369
  128. package/skills/pydicom/SKILL.md +0 -381
  129. package/skills/pyhealth/SKILL.md +0 -124
  130. package/skills/pylabrobot/SKILL.md +0 -216
  131. package/skills/pymatgen/SKILL.md +0 -404
  132. package/skills/pymc/SKILL.md +0 -310
  133. package/skills/pymoo/SKILL.md +0 -276
  134. package/skills/pyopenms/SKILL.md +0 -179
  135. package/skills/pysam/SKILL.md +0 -330
  136. package/skills/pytdc/SKILL.md +0 -297
  137. package/skills/pytorch-lightning/SKILL.md +0 -191
  138. package/skills/pyzotero/SKILL.md +0 -137
  139. package/skills/qiskit/SKILL.md +0 -259
  140. package/skills/qutip/SKILL.md +0 -317
  141. package/skills/rdkit/SKILL.md +0 -94
  142. package/skills/relsa-severity-assessment/SKILL.md +0 -354
  143. package/skills/research-grants/SKILL.md +0 -296
  144. package/skills/research-grants/references/README.md +0 -287
  145. package/skills/research-lookup/README.md +0 -106
  146. package/skills/research-lookup/SKILL.md +0 -338
  147. package/skills/rowan/SKILL.md +0 -398
  148. package/skills/scanpy/SKILL.md +0 -303
  149. package/skills/scholar-evaluation/SKILL.md +0 -296
  150. package/skills/scientific-brainstorming/SKILL.md +0 -282
  151. package/skills/scientific-critical-thinking/SKILL.md +0 -180
  152. package/skills/scientific-schematics/SKILL.md +0 -370
  153. package/skills/scientific-slides/SKILL.md +0 -379
  154. package/skills/scientific-visualization/SKILL.md +0 -285
  155. package/skills/scientific-writing/SKILL.md +0 -356
  156. package/skills/scikit-bio/SKILL.md +0 -470
  157. package/skills/scikit-learn/SKILL.md +0 -324
  158. package/skills/scikit-survival/SKILL.md +0 -313
  159. package/skills/scvelo/SKILL.md +0 -328
  160. package/skills/scvi-tools/SKILL.md +0 -201
  161. package/skills/seaborn/SKILL.md +0 -254
  162. package/skills/security-auditor/SKILL.md +0 -37
  163. package/skills/shap/SKILL.md +0 -282
  164. package/skills/simpy/SKILL.md +0 -283
  165. package/skills/stable-baselines3/SKILL.md +0 -325
  166. package/skills/statistical-analysis/SKILL.md +0 -446
  167. package/skills/statistical-power/SKILL.md +0 -200
  168. package/skills/statsmodels/SKILL.md +0 -238
  169. package/skills/sympy/SKILL.md +0 -354
  170. package/skills/systematic-debugging/SKILL.md +0 -35
  171. package/skills/tamarind/SKILL.md +0 -285
  172. package/skills/tdd/SKILL.md +0 -26
  173. package/skills/tiledbvcf/SKILL.md +0 -456
  174. package/skills/timesfm-forecasting/SKILL.md +0 -408
  175. package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
  176. package/skills/torch-geometric/SKILL.md +0 -458
  177. package/skills/torchdrug/SKILL.md +0 -241
  178. package/skills/transformers/SKILL.md +0 -195
  179. package/skills/treatment-plans/SKILL.md +0 -174
  180. package/skills/treatment-plans/references/README.md +0 -19
  181. package/skills/umap-learn/SKILL.md +0 -488
  182. package/skills/uncertainty-and-units/SKILL.md +0 -384
  183. package/skills/usfiscaldata/SKILL.md +0 -171
  184. package/skills/vaex/SKILL.md +0 -204
  185. package/skills/venue-templates/SKILL.md +0 -269
  186. package/skills/verification-before-completion/SKILL.md +0 -22
  187. package/skills/waypoint-bio/SKILL.md +0 -273
  188. package/skills/what-if-oracle/SKILL.md +0 -184
  189. package/skills/writing-plans/SKILL.md +0 -15
  190. package/skills/xlsx/SKILL.md +0 -110
  191. package/skills/zarr-python/SKILL.md +0 -241
@@ -1,254 +0,0 @@
1
- ---
2
- name: seaborn
3
- description: Statistical visualization with pandas integration. Use for quick exploration of distributions, relationships, and categorical comparisons with attractive defaults. Best for box plots, violin plots, pair plots, heatmaps. Built on matplotlib. For interactive plots use plotly; for publication styling use scientific-visualization.
4
- license: BSD-3-Clause license
5
- allowed-tools: Read Write Edit Bash
6
- compatibility: Requires Python 3.8+ and seaborn 0.13.2-compatible dependencies. Install with uv pip install seaborn==0.13.2; use seaborn[stats]==0.13.2 when advanced regression or clustering examples need scipy/statsmodels.
7
- metadata:
8
- version: "1.2"
9
- skill-author: K-Dense Inc.
10
- ---
11
-
12
- # Seaborn Statistical Visualization
13
-
14
- ## Overview
15
-
16
- Seaborn is a Python visualization library for creating publication-quality statistical graphics. Use this skill for dataset-oriented plotting, multivariate analysis, automatic statistical estimation, and complex multi-panel figures with minimal code.
17
-
18
- ## Environment and Installation
19
-
20
- Current upstream documentation is for seaborn 0.13.2. Official docs support Python 3.8+ with mandatory NumPy, pandas, and matplotlib dependencies; scipy, statsmodels, and fastcluster are optional for some advanced statistics and clustering workflows.
21
-
22
- ```bash
23
- # Reproducible install for examples in this skill
24
- uv pip install "seaborn==0.13.2"
25
-
26
- # Include optional statistical dependencies when needed
27
- uv pip install "seaborn[stats]==0.13.2"
28
- ```
29
-
30
- Recommended imports:
31
-
32
- ```python
33
- import numpy as np
34
- import pandas as pd
35
- import matplotlib.pyplot as plt
36
- import seaborn as sns
37
- import seaborn.objects as so
38
- ```
39
-
40
- `sns.load_dataset()` downloads public example data when it is not cached. For private, regulated, or offline work, load local files explicitly with pandas and pass the resulting DataFrame to seaborn.
41
-
42
- ## Design Philosophy
43
-
44
- Seaborn follows these core principles:
45
-
46
- 1. **Dataset-oriented**: Work directly with DataFrames and named variables rather than abstract coordinates
47
- 2. **Semantic mapping**: Automatically translate data values into visual properties (colors, sizes, styles)
48
- 3. **Statistical awareness**: Built-in aggregation, error estimation, and confidence intervals
49
- 4. **Aesthetic defaults**: Publication-ready themes and color palettes out of the box
50
- 5. **Matplotlib integration**: Full compatibility with matplotlib customization when needed
51
-
52
- ## Quick Start
53
-
54
- ```python
55
- import seaborn as sns
56
- import matplotlib.pyplot as plt
57
- import pandas as pd
58
-
59
- # Load example dataset
60
- df = sns.load_dataset('tips')
61
-
62
- # Create a simple visualization
63
- sns.scatterplot(data=df, x='total_bill', y='tip', hue='day')
64
- plt.show()
65
- ```
66
-
67
- ## Core Plotting Interfaces
68
-
69
- ### Function Interface (Traditional)
70
-
71
- The function interface provides specialized plotting functions organized by visualization type. Each category has **axes-level** functions (plot to single axes) and **figure-level** functions (manage entire figure with faceting).
72
-
73
- **When to use:**
74
- - Quick exploratory analysis
75
- - Single-purpose visualizations
76
- - When you need a specific plot type
77
-
78
- ### Objects Interface (Modern)
79
-
80
- The `seaborn.objects` interface provides a declarative, composable API similar to ggplot2. Build visualizations by chaining methods to specify data mappings, marks, transformations, and scales. Upstream still describes this interface as experimental and incomplete in 0.13.2, although stable enough for serious use; prefer the function interface for conservative production code unless the compositional API materially simplifies the plot.
81
-
82
- **When to use:**
83
- - Complex layered visualizations
84
- - When you need fine-grained control over transformations
85
- - Building custom plot types
86
- - Programmatic plot generation
87
-
88
- ```python
89
- from seaborn import objects as so
90
-
91
- # Declarative syntax
92
- (
93
- so.Plot(data=df, x='total_bill', y='tip')
94
- .add(so.Dot(), color='day')
95
- .add(so.Line(), so.PolyFit())
96
- )
97
- ```
98
-
99
- ## Current API Notes
100
-
101
- Seaborn 0.12 and 0.13 changed several common plotting patterns:
102
-
103
- - Most plotting functions now require keyword arguments for variables. Prefer `sns.scatterplot(data=df, x="x", y="y")` over positional `sns.scatterplot(df["x"], df["y"])`.
104
- - `errorbar` replaces the old `ci` parameter in `lineplot()`, `barplot()`, and `pointplot()`. Regression functions such as `regplot()` and `lmplot()` still use `ci`.
105
- - Categorical plots were rewritten in 0.13. Use `native_scale=True` when numeric or datetime categories should keep their original scale instead of ordinal positions.
106
- - Passing `palette` without assigning `hue` is deprecated for categorical functions. If each category should get its own color, assign a redundant hue such as `hue="day"` and set `legend=False`.
107
- - Prefer renamed parameters: `violinplot(density_norm=..., common_norm=...)` instead of `scale`/`scale_hue`, `boxenplot(width_method=...)` instead of `scale`, and `barplot(err_kws=...)` instead of `errcolor`/`errwidth`.
108
-
109
- ## Data Structure Requirements
110
-
111
- ### Long-Form Data (Preferred)
112
-
113
- Each variable is a column, each observation is a row. This "tidy" format provides maximum flexibility:
114
-
115
- ```python
116
- # Long-form structure
117
- subject condition measurement
118
- 0 1 control 10.5
119
- 1 1 treatment 12.3
120
- 2 2 control 9.8
121
- 3 2 treatment 13.1
122
- ```
123
-
124
- **Advantages:**
125
- - Works with all seaborn functions
126
- - Easy to remap variables to visual properties
127
- - Supports arbitrary complexity
128
- - Natural for DataFrame operations
129
-
130
- ### Wide-Form Data
131
-
132
- Variables are spread across columns. Useful for simple rectangular data:
133
-
134
- ```python
135
- # Wide-form structure
136
- control treatment
137
- 0 10.5 12.3
138
- 1 9.8 13.1
139
- ```
140
-
141
- **Use cases:**
142
- - Simple time series
143
- - Correlation matrices
144
- - Heatmaps
145
- - Quick plots of array data
146
-
147
- **Converting wide to long:**
148
- ```python
149
- df_long = df.melt(var_name='condition', value_name='measurement')
150
- ```
151
-
152
- ## Plotting Functions, Grids, Palettes, and Patterns
153
-
154
- - [references/plotting_functions.md](references/plotting_functions.md): relational,
155
- distribution, categorical, regression, and matrix plots by category.
156
- - [references/grids_and_levels.md](references/grids_and_levels.md): `FacetGrid`,
157
- `PairGrid`, `JointGrid`, and the figure-level vs axes-level distinction.
158
- - [references/palettes_and_theming.md](references/palettes_and_theming.md): palette
159
- choice (including colorblind-safe options), themes, contexts, and styles.
160
- - [references/patterns_and_troubleshooting.md](references/patterns_and_troubleshooting.md):
161
- common recipes and what seaborn's errors actually mean.
162
- - [references/objects_interface.md](references/objects_interface.md): the `seaborn.objects`
163
- interface. [references/function_reference.md](references/function_reference.md) and
164
- [references/examples.md](references/examples.md): full signatures and more examples.
165
-
166
- ## Best Practices
167
-
168
- ### 1. Data Preparation
169
-
170
- Always use well-structured DataFrames with meaningful column names:
171
-
172
- ```python
173
- # Good: Named columns in DataFrame
174
- df = pd.DataFrame({'bill': bills, 'tip': tips, 'day': days})
175
- sns.scatterplot(data=df, x='bill', y='tip', hue='day')
176
-
177
- # Avoid: Unnamed arrays
178
- sns.scatterplot(x=x_array, y=y_array) # Loses axis labels
179
- ```
180
-
181
- ### 2. Choose the Right Plot Type
182
-
183
- **Continuous x, continuous y:** `scatterplot`, `lineplot`, `kdeplot`, `regplot`
184
- **Continuous x, categorical y:** `violinplot`, `boxplot`, `stripplot`, `swarmplot`
185
- **One continuous variable:** `histplot`, `kdeplot`, `ecdfplot`
186
- **Correlations/matrices:** `heatmap`, `clustermap`
187
- **Pairwise relationships:** `pairplot`, `jointplot`
188
-
189
- ### 3. Use Figure-Level Functions for Faceting
190
-
191
- ```python
192
- # Instead of manual subplot creation
193
- sns.relplot(data=df, x='x', y='y', col='category', col_wrap=3)
194
-
195
- # Not: Creating subplots manually for simple faceting
196
- ```
197
-
198
- ### 4. Leverage Semantic Mappings
199
-
200
- Use `hue`, `size`, and `style` to encode additional dimensions:
201
-
202
- ```python
203
- sns.scatterplot(data=df, x='x', y='y',
204
- hue='category', # Color by category
205
- size='importance', # Size by continuous variable
206
- style='type') # Marker style by type
207
- ```
208
-
209
- ### 5. Control Statistical Estimation
210
-
211
- Many functions compute statistics automatically. Understand and customize:
212
-
213
- ```python
214
- # Lineplot computes mean and 95% CI by default
215
- sns.lineplot(data=df, x='time', y='value',
216
- errorbar='sd') # Use standard deviation instead
217
-
218
- # Barplot computes mean by default
219
- sns.barplot(data=df, x='category', y='value',
220
- estimator='median', # Use median instead
221
- errorbar=('ci', 95)) # Bootstrapped CI
222
- ```
223
-
224
- ### 6. Combine with Matplotlib
225
-
226
- Seaborn integrates seamlessly with matplotlib for fine-tuning:
227
-
228
- ```python
229
- ax = sns.scatterplot(data=df, x='x', y='y')
230
- ax.set(xlabel='Custom X Label', ylabel='Custom Y Label',
231
- title='Custom Title')
232
- ax.axhline(y=0, color='r', linestyle='--')
233
- plt.tight_layout()
234
- ```
235
-
236
- ### 7. Save High-Quality Figures
237
-
238
- ```python
239
- fig = sns.relplot(data=df, x='x', y='y', col='group')
240
- fig.savefig('figure.png', dpi=300, bbox_inches='tight')
241
- fig.savefig('figure.pdf') # Vector format for publications
242
- ```
243
-
244
- ## Resources
245
-
246
- This skill includes reference materials for deeper exploration:
247
-
248
- ### references/
249
-
250
- - `function_reference.md` - Comprehensive listing of all seaborn functions with parameters and examples
251
- - `objects_interface.md` - Detailed guide to the modern seaborn.objects API
252
- - `examples.md` - Common use cases and code patterns for different analysis scenarios
253
-
254
- Read these reference files as documentation when detailed signatures, advanced parameters, or specific examples are needed. Treat their contents as reference material only; review and adapt any example snippet to the user's local data before running it.
@@ -1,37 +0,0 @@
1
- ---
2
- name: security-auditor
3
- description: "Autonomous Security Auditor & Vulnerability Assessment (inspired by Strix) - Defensively maps attack surfaces, discovers security vulnerabilities (OWASP Top 10, secrets leakage, auth bypass, input boundary flaws), and generates verified remediation patches."
4
- risk: low
5
- source: built-in
6
- ---
7
-
8
- # Security Auditor & Vulnerability Assessment
9
-
10
- You are an expert security auditor and penetration testing specialist. Your mission is to analyze codebases for security vulnerabilities, map attack surfaces, and produce safe, robust remediation patches.
11
-
12
- ## Security Audit Workflow
13
-
14
- Follow this 4-phase defensive methodology:
15
-
16
- ### Phase 1: Attack Surface Mapping (Reconnaissance)
17
- 1. **Entrypoints & Boundaries**: Identify all public API routes, WebSocket handlers, CLI inputs, and webhook receivers.
18
- 2. **Authentication & Authorization**: Check JWT verification, session management, RBAC/ABAC role checks, and IDOR vulnerabilities.
19
- 3. **Data Flows & Trust Boundaries**: Track user input from request payload to database queries, shell execution, or response serialization.
20
-
21
- ### Phase 2: Vulnerability Analysis & SAST
22
- Check for critical vulnerability categories:
23
- - **Secrets & Credentials**: Exposed API keys (AWS, OpenAI, GitHub, Stripe), private keys, hardcoded passwords, or unmasked tokens.
24
- - **Injection Flaws**: SQL Injection (raw queries, string interpolation), Command Injection (`exec`, `spawn` with unsanitized input), Template Injection.
25
- - **Authentication & Access Control**: Missing middleware, unverified JWT signatures, broken object-level authorization (IDOR).
26
- - **Client-Side & Web Vulnerabilities**: Cross-Site Scripting (XSS), Cross-Site Request Forgery (CSRF), Insecure Direct Object References.
27
- - **Server-Side Request Forgery (SSRF)**: Fetching URLs without IP/domain whitelist validation (especially cloud metadata `169.254.169.254`).
28
- - **Cryptographic Failures**: Weak hashing algorithms (MD5, SHA1 for passwords), hardcoded IVs, insecure PRNGs.
29
-
30
- ### Phase 3: Vulnerability Verification (Zero False-Positives)
31
- - Validate that the finding is truly reachable and exploitable in the current codebase context.
32
- - Verify whether existing framework protections (e.g. ORM parameterization, auto-escaping, middleware) already mitigate the issue.
33
-
34
- ### Phase 4: Defensive Patching & Remediation
35
- - Formulate minimal, secure code patches using `apply_patch` or `write_file`.
36
- - Ensure the fix closes the vulnerability at the root cause without breaking existing features.
37
- - Add regression tests to verify that invalid/malicious input is properly rejected.
@@ -1,282 +0,0 @@
1
- ---
2
- name: shap
3
- description: Explain and audit machine-learning predictions with SHAP. Use for selecting SHAP explainers and maskers, computing and validating feature attributions, handling multi-output explanations, and producing local or global SHAP visualizations.
4
- license: MIT
5
- compatibility: Requires Python 3.12+ and uv for SHAP 0.52.0; model-specific libraries are optional.
6
- allowed-tools: "Read Bash"
7
- metadata:
8
- version: "2.0"
9
- skill-author: K-Dense Inc.
10
- ---
11
-
12
- # SHAP
13
-
14
- Use SHAP to describe how a fitted predictive model maps inputs to outputs. Work from the modern `shap.Explanation` API, make the explained output and background distribution explicit, and validate every explanation before interpreting it.
15
-
16
- This skill is aligned with **SHAP 0.52.0** (released 2026-05-28). That release requires Python 3.12 or newer.
17
-
18
- ## Operating Rules
19
-
20
- 1. Explain a fixed, evaluated model; do not use SHAP as a substitute for predictive validation.
21
- 2. Use held-out or clearly labeled analysis rows for explanations. Choose background rows only from an appropriate training or reference population.
22
- 3. State the explained output: regression value, raw margin, probability, log loss, logit, or another model method.
23
- 4. Keep explanations as `shap.Explanation` objects. Call `explainer(X)`; use `.shap_values(X)` only when maintaining legacy code.
24
- 5. For multi-output models, select one output before using tabular plots: `explanation[..., output_index]`.
25
- 6. Check `base_values + values.sum(...)` against the exact model output being explained.
26
- 7. Treat SHAP as a description of model behavior under a masking/background choice. It does not establish causality, fairness, recourse, or scientific mechanism.
27
- 8. Never silence an additivity failure until input shape, preprocessing, model version, output space, and row ordering have been checked.
28
- 9. Do not load untrusted pickle, joblib, model, or explainer artifacts; those formats can execute code during deserialization.
29
-
30
- ## Install
31
-
32
- Create an isolated environment and pin the documented release:
33
-
34
- ```bash
35
- uv venv --python 3.12
36
- source .venv/bin/activate
37
- uv pip install "shap[plots]==0.52.0"
38
- ```
39
-
40
- `shap[plots]` installs the plotting dependencies. Add the fitted model's package at a version compatible with the project. For older Python compatibility, read [references/migration.md](references/migration.md) instead of silently installing a different SHAP release.
41
-
42
- Confirm the environment before debugging an API mismatch:
43
-
44
- ```python
45
- import platform
46
- import shap
47
-
48
- print("Python:", platform.python_version())
49
- print("SHAP:", shap.__version__)
50
- ```
51
-
52
- ## Standard Workflow
53
-
54
- ### 1. Define the explanation target
55
-
56
- Record:
57
-
58
- - model and preprocessing version;
59
- - exact callable or model method being explained;
60
- - output name/index and units;
61
- - evaluation rows;
62
- - background/reference population;
63
- - masker and explainer algorithm;
64
- - SHAP and model-library versions.
65
-
66
- For classifiers, decide whether the task needs raw margins or probabilities. Defaults differ by model family; never infer units from the plot color or sign.
67
-
68
- ### 2. Select an explainer and masker
69
-
70
- Start with `shap.Explainer(model, masker)` when automatic dispatch is sufficient. Instantiate a specialized explainer when its assumptions or output controls matter.
71
-
72
- | Situation | Preferred choice | Important constraint |
73
- |---|---|---|
74
- | Supported tree ensemble | `TreeExplainer` | `model_output="probability"` and `"log_loss"` require interventional masking and background data |
75
- | Linear model | `LinearExplainer` | The masker determines interventional versus correlation-aware behavior |
76
- | Small feature space | `ExactExplainer` | Cost grows quickly with unconstrained feature count |
77
- | General tabular callable | `PermutationExplainer` | Budget at least one full forward/reverse permutation |
78
- | Hierarchical feature groups, text, or image | `PartitionExplainer` | The partition tree changes the cooperative game |
79
- | Differentiable neural network | `DeepExplainer` or `GradientExplainer` | Framework support, output shape, and background choice require testing |
80
- | Legacy Kernel SHAP workflow | `KernelExplainer` | Usually much slower than model-specific methods |
81
-
82
- Use the detailed decision guide in [references/explainers.md](references/explainers.md). Use [references/data-maskers.md](references/data-maskers.md) when features are correlated, structured, sparse, or semantically grouped.
83
-
84
- ### 3. Compute a modern `Explanation`
85
-
86
- This complete binary-classification example uses an explicit background and selects the positive-class output:
87
-
88
- ```python
89
- import numpy as np
90
- import shap
91
- from sklearn.datasets import load_breast_cancer
92
- from sklearn.ensemble import RandomForestClassifier
93
- from sklearn.model_selection import train_test_split
94
-
95
- X, y = load_breast_cancer(as_frame=True, return_X_y=True)
96
- X_train, X_test, y_train, y_test = train_test_split(
97
- X,
98
- y,
99
- test_size=0.2,
100
- stratify=y,
101
- random_state=7,
102
- )
103
-
104
- model = RandomForestClassifier(
105
- n_estimators=200,
106
- min_samples_leaf=3,
107
- random_state=7,
108
- n_jobs=-1,
109
- ).fit(X_train, y_train)
110
-
111
- background = shap.sample(X_train, 100, random_state=7)
112
- explainer = shap.Explainer(model, background, algorithm="tree")
113
- all_outputs = explainer(X_test)
114
-
115
- # sklearn tree classifiers expose one output per class.
116
- positive = all_outputs[..., 1]
117
- assert positive.values.shape == X_test.shape
118
-
119
- reconstructed = np.asarray(positive.base_values) + positive.values.sum(axis=1)
120
- expected = model.predict_proba(X_test)[:, 1]
121
- np.testing.assert_allclose(reconstructed, expected, rtol=1e-5, atol=1e-6)
122
-
123
- shap.plots.beeswarm(positive, max_display=15)
124
- shap.plots.waterfall(positive[0], max_display=15)
125
- ```
126
-
127
- Output shape is model-dependent:
128
-
129
- - one tabular output: `(samples, features)`;
130
- - multiple tabular outputs: `(samples, features, outputs)`;
131
- - multiple model inputs: often a list of arrays or explanations;
132
- - image/text explanations: feature axes follow the input representation, with output selection on the final axis when present.
133
-
134
- Do not use the pre-0.45 pattern `values[class_index]` for a modern multi-output array. Use `values[..., class_index]` or slice the `Explanation` itself.
135
-
136
- ### 4. Control tree output semantics when needed
137
-
138
- For a supported tree classifier, probability-space explanations must be explicit:
139
-
140
- ```python
141
- background = shap.sample(X_train, 200, random_state=7)
142
-
143
- explainer = shap.TreeExplainer(
144
- model,
145
- data=background,
146
- feature_perturbation="interventional",
147
- model_output="probability",
148
- )
149
- probability_exp = explainer(X_test)
150
- ```
151
-
152
- In SHAP 0.52:
153
-
154
- - `feature_perturbation="auto"` uses interventional semantics when background data is supplied and tree-path-dependent semantics otherwise;
155
- - probability and log-loss output modes are supported only with interventional semantics;
156
- - pass `approximate=True` to `explainer(X, approximate=True)` if deliberately using the lower-fidelity tree approximation; do not pass it to the constructor.
157
-
158
- ### 5. Use a model-agnostic callable deliberately
159
-
160
- Pass the exact callable whose outputs will be interpreted:
161
-
162
- ```python
163
- masker = shap.maskers.Independent(background, max_samples=100)
164
- explainer = shap.Explainer(
165
- model.predict_proba,
166
- masker,
167
- algorithm="permutation",
168
- output_names=[str(label) for label in model.classes_],
169
- seed=7,
170
- )
171
-
172
- budget = 2 * X_test.shape[1] + 1
173
- all_outputs = explainer(X_test.iloc[:20], max_evals=budget)
174
- positive = all_outputs[..., 1]
175
- ```
176
-
177
- Increase `max_evals` to average over more permutations when estimates are unstable. Keep the seed, background sample, and evaluation budget in the report.
178
-
179
- ### 6. Visualize the question, not merely the available plot
180
-
181
- | Question | Plot |
182
- |---|---|
183
- | Which features have the largest average attribution magnitude? | `shap.plots.bar(exp)` |
184
- | How do direction, magnitude, and observed values vary globally? | `shap.plots.beeswarm(exp)` |
185
- | Why did one prediction differ from its baseline? | `shap.plots.waterfall(exp[i])` |
186
- | How does one feature's attribution vary over its values? | `shap.plots.scatter(exp[:, feature])` |
187
- | Do explanations form sample-level patterns? | `shap.plots.heatmap(exp)` |
188
- | How do predefined cohorts differ descriptively? | `shap.plots.bar(exp.cohorts(labels).abs.mean(0))` |
189
- | Which tokens or image regions contribute to an output? | `shap.plots.text(exp)` or `shap.plots.image(exp)` |
190
-
191
- Read [references/plots.md](references/plots.md) before customizing or saving figures.
192
-
193
- ### 7. Report limitations with results
194
-
195
- At minimum, report:
196
-
197
- - output and units;
198
- - baseline/reference population;
199
- - explainer and masker;
200
- - sample count and selection;
201
- - output index/name;
202
- - additivity error or applicable approximation diagnostics;
203
- - known correlated/grouped features;
204
- - whether results are local, aggregated, or cohort-specific;
205
- - a clear non-causal statement.
206
-
207
- ## Common Tasks
208
-
209
- ### Global and local analysis
210
-
211
- Use global plots to locate important patterns, scatter plots to inspect those patterns, and local plots to investigate selected rows. Do not select only visually dramatic rows without documenting the selection rule.
212
-
213
- ### Multiclass models
214
-
215
- Set `output_names` where possible, inspect `explanation.output_names`, and slice an output before plotting:
216
-
217
- ```python
218
- class_exp = explanation[..., "class_name"]
219
- # or
220
- class_exp = explanation[..., class_index]
221
- ```
222
-
223
- Never average signed attributions across classes. For cross-class comparison, preserve the same model, rows, background, output space, and aggregation.
224
-
225
- ### Cohorts, subgroup analysis, and fairness
226
-
227
- SHAP can compare how a model uses features across cohorts, but this is not a fairness test. A protected feature with small SHAP magnitude does not rule out proxy discrimination, and removing a protected feature does not establish fairness. Pair attribution analysis with performance, calibration, error-rate, and domain-appropriate fairness metrics.
228
-
229
- See [references/workflows.md](references/workflows.md) for cohort construction, model comparison, error analysis, log-loss explanations, monitoring, and production records.
230
-
231
- ### Text and images
232
-
233
- Use domain maskers rather than treating tokens or pixels as ordinary independent columns:
234
-
235
- - `shap.maskers.Text(tokenizer)` with `PartitionExplainer` for token groups;
236
- - `shap.maskers.Image(...)` with `PartitionExplainer` for image regions;
237
- - restrict expensive multi-output models with `outputs=...`.
238
-
239
- Read [references/modalities.md](references/modalities.md) for current examples and output-shape guidance.
240
-
241
- ## Troubleshooting Order
242
-
243
- 1. Print Python, SHAP, model-library, NumPy, and framework versions.
244
- 2. Verify the model receives exactly the same transformed columns, order, dtype, and missing-value representation used during fitting.
245
- 3. Print `values.shape`, `base_values.shape`, `data.shape`, `feature_names`, and `output_names`.
246
- 4. Confirm the selected output and output units.
247
- 5. Recompute predictions on the same rows in the same order.
248
- 6. Test a smaller batch and representative background.
249
- 7. Only then investigate package-specific compatibility or approximation settings.
250
-
251
- Use [references/troubleshooting.md](references/troubleshooting.md) for additivity failures, shape mismatches, categorical features, pipelines, deep-learning frameworks, plotting, and performance.
252
-
253
- ## Bundled Script
254
-
255
- Run a deterministic, self-contained tabular example that writes importance data, metadata, and plots:
256
-
257
- ```bash
258
- uv run --no-project --python 3.12 --with "shap[plots]==0.52.0" \
259
- skills/shap/scripts/tabular_report.py --output-dir /tmp/shap-report
260
- ```
261
-
262
- The script does not download data or deserialize models. Read it as a template, then replace the built-in dataset and model while preserving output selection and additivity validation.
263
-
264
- ## Reference Map
265
-
266
- | File | Load when |
267
- |---|---|
268
- | [references/explainers.md](references/explainers.md) | Selecting or configuring explainers |
269
- | [references/data-maskers.md](references/data-maskers.md) | Choosing background data, masking semantics, or feature groups |
270
- | [references/plots.md](references/plots.md) | Selecting, composing, or saving visualizations |
271
- | [references/workflows.md](references/workflows.md) | Running audits, comparisons, cohorts, monitoring, or production workflows |
272
- | [references/modalities.md](references/modalities.md) | Explaining text, images, or deep models |
273
- | [references/migration.md](references/migration.md) | Updating legacy SHAP code or supporting older Python |
274
- | [references/theory.md](references/theory.md) | Explaining estimands, guarantees, dependence, interactions, and limitations |
275
- | [references/troubleshooting.md](references/troubleshooting.md) | Diagnosing runtime, shape, additivity, and compatibility problems |
276
-
277
- ## Primary Sources
278
-
279
- - Documentation: https://shap.readthedocs.io/en/latest/
280
- - API reference: https://shap.readthedocs.io/en/latest/api.html
281
- - Release notes: https://shap.readthedocs.io/en/latest/release_notes.html
282
- - Repository: https://github.com/shap/shap