@pikaa-ai/pikaa 0.3.23 → 0.3.25
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
- package/assets/brand/orbit-logo.jpg +0 -0
- package/assets/brand/orbit-logo.png +0 -0
- package/assets/brand/orbit-logo.svg +3 -0
- package/dist/cli.js +407 -219
- package/dist/index.js +7 -2
- package/package.json +1 -2
- package/skills/adaptyv/SKILL.md +0 -240
- package/skills/aeon/SKILL.md +0 -402
- package/skills/analytical-method-validation/SKILL.md +0 -299
- package/skills/anndata/SKILL.md +0 -431
- package/skills/arbor/SKILL.md +0 -152
- package/skills/arboreto/SKILL.md +0 -267
- package/skills/astropy/SKILL.md +0 -353
- package/skills/autoskill/SKILL.md +0 -233
- package/skills/benchling-integration/SKILL.md +0 -229
- package/skills/bgpt-paper-search/SKILL.md +0 -75
- package/skills/bids/SKILL.md +0 -237
- package/skills/biopython/SKILL.md +0 -472
- package/skills/bioservices/SKILL.md +0 -399
- package/skills/bulk-rnaseq/SKILL.md +0 -198
- package/skills/cellxgene-census/SKILL.md +0 -283
- package/skills/cirq/SKILL.md +0 -370
- package/skills/citation-management/SKILL.md +0 -329
- package/skills/clinical-decision-support/SKILL.md +0 -238
- package/skills/clinical-decision-support/references/README.md +0 -62
- package/skills/clinical-reports/SKILL.md +0 -248
- package/skills/clinical-reports/references/README.md +0 -34
- package/skills/cobrapy/SKILL.md +0 -496
- package/skills/consciousness-council/SKILL.md +0 -151
- package/skills/dask/SKILL.md +0 -482
- package/skills/database-lookup/SKILL.md +0 -386
- package/skills/datamol/SKILL.md +0 -200
- package/skills/deepchem/SKILL.md +0 -244
- package/skills/deepspot-m/SKILL.md +0 -175
- package/skills/deeptools/SKILL.md +0 -412
- package/skills/depmap/SKILL.md +0 -301
- package/skills/dhdna-profiler/SKILL.md +0 -184
- package/skills/diffdock/SKILL.md +0 -488
- package/skills/dnanexus-integration/SKILL.md +0 -325
- package/skills/docx/SKILL.md +0 -99
- package/skills/esm/SKILL.md +0 -334
- package/skills/etetoolkit/SKILL.md +0 -327
- package/skills/exa-search/SKILL.md +0 -102
- package/skills/executing-plans/SKILL.md +0 -14
- package/skills/experimental-design/SKILL.md +0 -234
- package/skills/exploratory-data-analysis/SKILL.md +0 -280
- package/skills/flowio/SKILL.md +0 -310
- package/skills/fluidsim/SKILL.md +0 -279
- package/skills/frontend-design/SKILL.md +0 -100
- package/skills/generate-image/SKILL.md +0 -304
- package/skills/geniml/SKILL.md +0 -310
- package/skills/genomic-coordinates/SKILL.md +0 -189
- package/skills/genomic-intelligence/SKILL.md +0 -243
- package/skills/geomaster/README.md +0 -105
- package/skills/geomaster/SKILL.md +0 -366
- package/skills/geopandas/SKILL.md +0 -250
- package/skills/get-available-resources/SKILL.md +0 -260
- package/skills/gget/SKILL.md +0 -153
- package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
- package/skills/glycoengineering/SKILL.md +0 -339
- package/skills/gtars/SKILL.md +0 -282
- package/skills/guardian-rails/SKILL.md +0 -54
- package/skills/histolab/SKILL.md +0 -243
- package/skills/hugging-science/SKILL.md +0 -132
- package/skills/hypogenic/SKILL.md +0 -290
- package/skills/hypothesis-generation/SKILL.md +0 -264
- package/skills/imaging-data-commons/SKILL.md +0 -496
- package/skills/infographics/SKILL.md +0 -315
- package/skills/iso-standards-readiness/SKILL.md +0 -352
- package/skills/lab-hardware-cad/SKILL.md +0 -372
- package/skills/labarchive-integration/SKILL.md +0 -216
- package/skills/lamindb/SKILL.md +0 -408
- package/skills/latchbio-integration/SKILL.md +0 -227
- package/skills/latex-posters/SKILL.md +0 -369
- package/skills/latex-posters/references/README.md +0 -439
- package/skills/liteparse/SKILL.md +0 -295
- package/skills/literature-review/SKILL.md +0 -263
- package/skills/markdown-mermaid-writing/SKILL.md +0 -322
- package/skills/market-research-reports/SKILL.md +0 -337
- package/skills/markitdown/SKILL.md +0 -264
- package/skills/matchms/SKILL.md +0 -276
- package/skills/matlab/SKILL.md +0 -274
- package/skills/matplotlib/SKILL.md +0 -378
- package/skills/medchem/SKILL.md +0 -321
- package/skills/modal/SKILL.md +0 -468
- package/skills/molecular-dynamics/SKILL.md +0 -458
- package/skills/molfeat/SKILL.md +0 -348
- package/skills/ncats-arax/SKILL.md +0 -178
- package/skills/networkx/SKILL.md +0 -440
- package/skills/neurokit2/SKILL.md +0 -323
- package/skills/neuropixels-analysis/SKILL.md +0 -412
- package/skills/nextflow/SKILL.md +0 -195
- package/skills/omero-integration/SKILL.md +0 -222
- package/skills/onekgpd/SKILL.md +0 -371
- package/skills/ontology-term-resolution/SKILL.md +0 -147
- package/skills/open-notebook/SKILL.md +0 -297
- package/skills/openpiv/SKILL.md +0 -469
- package/skills/opentrons-integration/SKILL.md +0 -322
- package/skills/optimize-for-gpu/SKILL.md +0 -176
- package/skills/owasp-top10/SKILL.md +0 -48
- package/skills/pacsomatic/LICENSE +0 -21
- package/skills/pacsomatic/SKILL.md +0 -150
- package/skills/paper-lookup/SKILL.md +0 -263
- package/skills/paperclip/SKILL.md +0 -413
- package/skills/paperzilla/SKILL.md +0 -159
- package/skills/parallel-web/SKILL.md +0 -128
- package/skills/pathml/SKILL.md +0 -222
- package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
- package/skills/pathway-enrichment/SKILL.md +0 -194
- package/skills/pdf/SKILL.md +0 -322
- package/skills/peer-review/SKILL.md +0 -288
- package/skills/penetration-testing/SKILL.md +0 -31
- package/skills/pennylane/SKILL.md +0 -240
- package/skills/phylogenetics/SKILL.md +0 -409
- package/skills/pi-agent/SKILL.md +0 -83
- package/skills/pkpd-modeling/SKILL.md +0 -381
- package/skills/polars/SKILL.md +0 -393
- package/skills/polars-bio/SKILL.md +0 -379
- package/skills/ponytail/SKILL.md +0 -31
- package/skills/ponytail-audit/SKILL.md +0 -18
- package/skills/pptx/SKILL.md +0 -246
- package/skills/pptx-posters/SKILL.md +0 -258
- package/skills/primekg/SKILL.md +0 -99
- package/skills/protocolsio-integration/SKILL.md +0 -236
- package/skills/pufferlib/SKILL.md +0 -328
- package/skills/pydeseq2/SKILL.md +0 -369
- package/skills/pydicom/SKILL.md +0 -381
- package/skills/pyhealth/SKILL.md +0 -124
- package/skills/pylabrobot/SKILL.md +0 -216
- package/skills/pymatgen/SKILL.md +0 -404
- package/skills/pymc/SKILL.md +0 -310
- package/skills/pymoo/SKILL.md +0 -276
- package/skills/pyopenms/SKILL.md +0 -179
- package/skills/pysam/SKILL.md +0 -330
- package/skills/pytdc/SKILL.md +0 -297
- package/skills/pytorch-lightning/SKILL.md +0 -191
- package/skills/pyzotero/SKILL.md +0 -137
- package/skills/qiskit/SKILL.md +0 -259
- package/skills/qutip/SKILL.md +0 -317
- package/skills/rdkit/SKILL.md +0 -94
- package/skills/relsa-severity-assessment/SKILL.md +0 -354
- package/skills/research-grants/SKILL.md +0 -296
- package/skills/research-grants/references/README.md +0 -287
- package/skills/research-lookup/README.md +0 -106
- package/skills/research-lookup/SKILL.md +0 -338
- package/skills/rowan/SKILL.md +0 -398
- package/skills/scanpy/SKILL.md +0 -303
- package/skills/scholar-evaluation/SKILL.md +0 -296
- package/skills/scientific-brainstorming/SKILL.md +0 -282
- package/skills/scientific-critical-thinking/SKILL.md +0 -180
- package/skills/scientific-schematics/SKILL.md +0 -370
- package/skills/scientific-slides/SKILL.md +0 -379
- package/skills/scientific-visualization/SKILL.md +0 -285
- package/skills/scientific-writing/SKILL.md +0 -356
- package/skills/scikit-bio/SKILL.md +0 -470
- package/skills/scikit-learn/SKILL.md +0 -324
- package/skills/scikit-survival/SKILL.md +0 -313
- package/skills/scvelo/SKILL.md +0 -328
- package/skills/scvi-tools/SKILL.md +0 -201
- package/skills/seaborn/SKILL.md +0 -254
- package/skills/security-auditor/SKILL.md +0 -37
- package/skills/shap/SKILL.md +0 -282
- package/skills/simpy/SKILL.md +0 -283
- package/skills/stable-baselines3/SKILL.md +0 -325
- package/skills/statistical-analysis/SKILL.md +0 -446
- package/skills/statistical-power/SKILL.md +0 -200
- package/skills/statsmodels/SKILL.md +0 -238
- package/skills/sympy/SKILL.md +0 -354
- package/skills/systematic-debugging/SKILL.md +0 -35
- package/skills/tamarind/SKILL.md +0 -285
- package/skills/tdd/SKILL.md +0 -26
- package/skills/tiledbvcf/SKILL.md +0 -456
- package/skills/timesfm-forecasting/SKILL.md +0 -408
- package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
- package/skills/torch-geometric/SKILL.md +0 -458
- package/skills/torchdrug/SKILL.md +0 -241
- package/skills/transformers/SKILL.md +0 -195
- package/skills/treatment-plans/SKILL.md +0 -174
- package/skills/treatment-plans/references/README.md +0 -19
- package/skills/umap-learn/SKILL.md +0 -488
- package/skills/uncertainty-and-units/SKILL.md +0 -384
- package/skills/usfiscaldata/SKILL.md +0 -171
- package/skills/vaex/SKILL.md +0 -204
- package/skills/venue-templates/SKILL.md +0 -269
- package/skills/verification-before-completion/SKILL.md +0 -22
- package/skills/waypoint-bio/SKILL.md +0 -273
- package/skills/what-if-oracle/SKILL.md +0 -184
- package/skills/writing-plans/SKILL.md +0 -15
- package/skills/xlsx/SKILL.md +0 -110
- package/skills/zarr-python/SKILL.md +0 -241
|
@@ -1,446 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: statistical-analysis
|
|
3
|
-
description: Guided statistical analysis for research data - test selection, assumption checking, effect sizes, power analysis, Bayesian alternatives, and APA-formatted reporting. Use whenever a user wants to compare groups, test a hypothesis, analyze experimental or survey data, check statistical assumptions, compute required sample sizes, or write up results - even if they never name a specific test. Covers t-tests, ANOVA, chi-square, correlation, regression, non-parametric and Bayesian methods. For low-level model APIs, see the statsmodels and pymc skills.
|
|
4
|
-
license: MIT license
|
|
5
|
-
metadata:
|
|
6
|
-
version: "1.1"
|
|
7
|
-
skill-author: K-Dense Inc.
|
|
8
|
-
---
|
|
9
|
-
|
|
10
|
-
# Statistical Analysis
|
|
11
|
-
|
|
12
|
-
## Overview
|
|
13
|
-
|
|
14
|
-
Conduct hypothesis tests (t-tests, ANOVA, chi-square), regression, correlation, and Bayesian analyses with systematic assumption checking, effect sizes, and APA-style reporting. The goal is an analysis a reviewer could not tear apart: the right test, verified assumptions, honest effect sizes, and a complete write-up.
|
|
15
|
-
|
|
16
|
-
## When to Use This Skill
|
|
17
|
-
|
|
18
|
-
Use this skill when:
|
|
19
|
-
- Conducting statistical hypothesis tests (t-tests, ANOVA, chi-square, non-parametric)
|
|
20
|
-
- Performing regression or correlation analyses
|
|
21
|
-
- Running Bayesian statistical analyses
|
|
22
|
-
- Checking statistical assumptions and diagnostics
|
|
23
|
-
- Calculating effect sizes and conducting power analyses
|
|
24
|
-
- Reporting statistical results in APA format
|
|
25
|
-
- Analyzing experimental or observational data for research
|
|
26
|
-
|
|
27
|
-
---
|
|
28
|
-
|
|
29
|
-
## Installation
|
|
30
|
-
|
|
31
|
-
Use **uv** to install the libraries used in this skill. Pin versions in production; unpinned installs are fine for exploration.
|
|
32
|
-
|
|
33
|
-
```bash
|
|
34
|
-
# Core frequentist stack (Python 3.10+; 3.12+ recommended for latest SciPy/ArviZ)
|
|
35
|
-
uv pip install "pingouin>=0.6" "scipy>=1.11" "statsmodels>=0.14.6" pandas matplotlib seaborn
|
|
36
|
-
|
|
37
|
-
# Bayesian modeling (PyMC 5 + ArviZ)
|
|
38
|
-
uv pip install "pymc>=5.0" "arviz>=1.0"
|
|
39
|
-
```
|
|
40
|
-
|
|
41
|
-
**Compatibility notes (verified against pingouin 0.6.1, statsmodels 0.14.6, arviz 1.2, 2026):**
|
|
42
|
-
|
|
43
|
-
- **Pingouin 0.6.0** renamed output columns to remove special characters: `p_val`, `cohen_d`, `CI95`, `p_unc` (previously `p-val`, `cohen-d`, `CI95%`, `p-unc` in 0.5.x). Examples below use the current names; if stuck on 0.5.x, use the hyphenated forms.
|
|
44
|
-
- **statsmodels + SciPy**: use `statsmodels>=0.14.6` with `scipy>=1.11` to avoid `_lazywhere` import errors on SciPy 1.16+.
|
|
45
|
-
- **ArviZ 1.x**: `az.summary()` now defaults to **89% intervals** (`eti89` columns) and the width parameter is `ci_prob` (not `hdi_prob`). To report a conventional 95% credible interval, pass `az.summary(trace, ci_prob=0.95)`.
|
|
46
|
-
- **One-sided Bayes Factors are gone from Pingouin**: `pg.ttest(..., alternative='greater')` silently drops the `BF10` column, and `pg.bayesfactor_ttest` raises on one-sided alternatives. For one-sided Bayesian tests, use PyMC directly (compute the posterior probability of the directional hypothesis) or JASP/R's BayesFactor.
|
|
47
|
-
|
|
48
|
-
For model-specific APIs (OLS, GLM, ARIMA), see the **statsmodels** skill. For PyMC workflows, see the **pymc** skill.
|
|
49
|
-
|
|
50
|
-
---
|
|
51
|
-
|
|
52
|
-
## Analysis Workflow
|
|
53
|
-
|
|
54
|
-
Every sound analysis follows the same arc. Skipping steps is how analyses end up retracted, so work through them in order and say what you did at each one.
|
|
55
|
-
|
|
56
|
-
1. **Frame the question before touching the data.** State the hypothesis, the outcome and predictor variables, and the design (independent vs. paired, number of groups). Commit to a planned test now — choosing the test after peeking at results is p-hacking, even when done innocently.
|
|
57
|
-
2. **Inspect the data.** Per group: n, mean, SD, median, missing values. Plot the raw data (histograms or box plots) before any test. Unequal group sizes, missingness, floor/ceiling effects, and outliers all change what test is appropriate — surface them to the user rather than silently working around them.
|
|
58
|
-
3. **Select the test** using the quick reference below, or `references/test_selection_guide.md` for designs beyond the basics (counts, time-to-event, reliability, factorial).
|
|
59
|
-
4. **Check assumptions** with `scripts/assumption_checks.py`. If an assumption fails, switch to the remedial test (table below) and report both the plan and the change.
|
|
60
|
-
5. **Run the test** and always compute the effect size alongside it — a p-value says an effect exists; the effect size says whether anyone should care.
|
|
61
|
-
6. **Report** using the APA templates below, including descriptives, exact statistics, effect sizes with CIs, and the assumption checks performed.
|
|
62
|
-
|
|
63
|
-
If the user only needs one step (e.g., "how many participants do I need?"), jump straight to that section — but still confirm the design assumptions the calculation rests on.
|
|
64
|
-
|
|
65
|
-
---
|
|
66
|
-
|
|
67
|
-
## Test Selection Guide
|
|
68
|
-
|
|
69
|
-
### Quick Reference: Choosing the Right Test
|
|
70
|
-
|
|
71
|
-
Use `references/test_selection_guide.md` for comprehensive guidance (counts, survival, reliability, factorial designs). Quick reference:
|
|
72
|
-
|
|
73
|
-
**Comparing Two Groups:**
|
|
74
|
-
- Independent, continuous, normal → Independent t-test
|
|
75
|
-
- Independent, continuous, non-normal → Mann-Whitney U test
|
|
76
|
-
- Paired, continuous, normal → Paired t-test
|
|
77
|
-
- Paired, continuous, non-normal → Wilcoxon signed-rank test
|
|
78
|
-
- Binary outcome → Chi-square or Fisher's exact test
|
|
79
|
-
|
|
80
|
-
**Comparing 3+ Groups:**
|
|
81
|
-
- Independent, continuous, normal → One-way ANOVA
|
|
82
|
-
- Independent, continuous, non-normal → Kruskal-Wallis test
|
|
83
|
-
- Paired, continuous, normal → Repeated measures ANOVA
|
|
84
|
-
- Paired, continuous, non-normal → Friedman test
|
|
85
|
-
|
|
86
|
-
**Relationships:**
|
|
87
|
-
- Two continuous variables → Pearson (normal) or Spearman correlation (non-normal)
|
|
88
|
-
- Continuous outcome with predictor(s) → Linear regression
|
|
89
|
-
- Binary outcome with predictor(s) → Logistic regression
|
|
90
|
-
|
|
91
|
-
**Bayesian Alternatives:**
|
|
92
|
-
All tests have Bayesian versions providing direct probability statements about hypotheses, Bayes Factors quantifying evidence, and the ability to support the null. See `references/bayesian_statistics.md`.
|
|
93
|
-
|
|
94
|
-
---
|
|
95
|
-
|
|
96
|
-
## Assumption Checking
|
|
97
|
-
|
|
98
|
-
**Always check assumptions before interpreting test results**, and report the checks — reviewers look for them.
|
|
99
|
-
|
|
100
|
-
Use the bundled `scripts/assumption_checks.py` module. Run Python from the skill directory (`skills/statistical-analysis/`) or add `scripts/` to `sys.path`:
|
|
101
|
-
|
|
102
|
-
```python
|
|
103
|
-
from assumption_checks import comprehensive_assumption_check
|
|
104
|
-
|
|
105
|
-
# Outliers + normality (per group) + homogeneity of variance, with plots
|
|
106
|
-
results = comprehensive_assumption_check(
|
|
107
|
-
data=df,
|
|
108
|
-
value_col='score',
|
|
109
|
-
group_col='group', # Optional: for group comparisons
|
|
110
|
-
alpha=0.05
|
|
111
|
-
)
|
|
112
|
-
```
|
|
113
|
-
|
|
114
|
-
For targeted checks, import individual functions:
|
|
115
|
-
|
|
116
|
-
```python
|
|
117
|
-
from assumption_checks import (
|
|
118
|
-
check_normality, # Shapiro-Wilk + Q-Q plot + histogram
|
|
119
|
-
check_normality_per_group,
|
|
120
|
-
check_homogeneity_of_variance, # Levene's test + box plots
|
|
121
|
-
check_linearity, # scatter + residual plot for simple regression
|
|
122
|
-
check_regression_diagnostics, # full OLS diagnostics (see Regression below)
|
|
123
|
-
detect_outliers # IQR or z-score methods
|
|
124
|
-
)
|
|
125
|
-
|
|
126
|
-
result = check_normality(data=df['score'], name='Test Score', alpha=0.05, plot=True)
|
|
127
|
-
print(result['interpretation'])
|
|
128
|
-
print(result['recommendation'])
|
|
129
|
-
```
|
|
130
|
-
|
|
131
|
-
### What to Do When Assumptions Are Violated
|
|
132
|
-
|
|
133
|
-
**Normality violated:**
|
|
134
|
-
- Mild violation + n > 30 per group → Proceed with parametric test (robust)
|
|
135
|
-
- Moderate violation → Use non-parametric alternative
|
|
136
|
-
- Severe violation → Transform data or use non-parametric test
|
|
137
|
-
|
|
138
|
-
**Homogeneity of variance violated:**
|
|
139
|
-
- For t-test → Use Welch's t-test (`pg.ttest` applies it automatically with `correction='auto'`)
|
|
140
|
-
- For ANOVA → Use Welch's ANOVA (`pg.welch_anova`) or Brown-Forsythe
|
|
141
|
-
- For regression → Use robust standard errors or weighted least squares
|
|
142
|
-
|
|
143
|
-
**Linearity violated (regression):**
|
|
144
|
-
- Add polynomial terms, transform variables, or use non-linear models / GAM
|
|
145
|
-
|
|
146
|
-
Formal tests get oversensitive as n grows: for n ≥ 100, weigh the Q-Q plot more heavily than the Shapiro-Wilk p-value. See `references/assumptions_and_diagnostics.md` for comprehensive guidance.
|
|
147
|
-
|
|
148
|
-
---
|
|
149
|
-
|
|
150
|
-
## Running Statistical Tests
|
|
151
|
-
|
|
152
|
-
Primary libraries:
|
|
153
|
-
- **pingouin**: user-friendly tests that return effect sizes by default — prefer it for standard tests
|
|
154
|
-
- **scipy.stats**: core statistical tests
|
|
155
|
-
- **statsmodels**: regression, diagnostics, power analysis
|
|
156
|
-
- **pymc** + **arviz**: Bayesian modeling and diagnostics
|
|
157
|
-
|
|
158
|
-
### T-Test with Complete Reporting
|
|
159
|
-
|
|
160
|
-
```python
|
|
161
|
-
import pingouin as pg
|
|
162
|
-
|
|
163
|
-
# correction='auto' applies Welch's correction when variances are unequal
|
|
164
|
-
result = pg.ttest(group_a, group_b, correction='auto')
|
|
165
|
-
|
|
166
|
-
# Pingouin >= 0.6 column names
|
|
167
|
-
t_stat = result['T'].values[0]
|
|
168
|
-
df = result['dof'].values[0]
|
|
169
|
-
p_value = result['p_val'].values[0]
|
|
170
|
-
cohens_d = result['cohen_d'].values[0]
|
|
171
|
-
ci_lower, ci_upper = result['CI95'].values[0] # CI for the mean difference
|
|
172
|
-
|
|
173
|
-
print(f"t({df:.0f}) = {t_stat:.2f}, p = {p_value:.3f}, d = {cohens_d:.2f}")
|
|
174
|
-
```
|
|
175
|
-
|
|
176
|
-
### ANOVA with Post-Hoc Tests
|
|
177
|
-
|
|
178
|
-
```python
|
|
179
|
-
import pingouin as pg
|
|
180
|
-
|
|
181
|
-
aov = pg.anova(dv='score', between='group', data=df, detailed=True)
|
|
182
|
-
print(aov)
|
|
183
|
-
|
|
184
|
-
# Effect size: partial eta-squared
|
|
185
|
-
eta_p2 = aov['np2'].values[0]
|
|
186
|
-
|
|
187
|
-
# If significant, conduct post-hoc tests (Tukey HSD controls family-wise error)
|
|
188
|
-
if aov['p_unc'].values[0] < 0.05:
|
|
189
|
-
posthoc = pg.pairwise_tukey(dv='score', between='group', data=df)
|
|
190
|
-
print(posthoc) # includes Hedges' g per pair
|
|
191
|
-
```
|
|
192
|
-
|
|
193
|
-
### Linear Regression with Diagnostics
|
|
194
|
-
|
|
195
|
-
```python
|
|
196
|
-
import statsmodels.api as sm
|
|
197
|
-
from assumption_checks import check_regression_diagnostics
|
|
198
|
-
|
|
199
|
-
X = sm.add_constant(X_predictors) # Add intercept
|
|
200
|
-
model = sm.OLS(y, X).fit()
|
|
201
|
-
print(model.summary())
|
|
202
|
-
|
|
203
|
-
# 4-panel residual plot + Shapiro-Wilk, Breusch-Pagan, Durbin-Watson, VIF
|
|
204
|
-
diag = check_regression_diagnostics(model)
|
|
205
|
-
print(diag['interpretation'])
|
|
206
|
-
print(diag['vif'])
|
|
207
|
-
|
|
208
|
-
# If heteroscedasticity was flagged, report robust standard errors instead
|
|
209
|
-
robust = model.get_robustcov_results('HC3')
|
|
210
|
-
```
|
|
211
|
-
|
|
212
|
-
### Bayesian T-Test
|
|
213
|
-
|
|
214
|
-
```python
|
|
215
|
-
import pymc as pm
|
|
216
|
-
import arviz as az
|
|
217
|
-
import numpy as np
|
|
218
|
-
|
|
219
|
-
with pm.Model() as model:
|
|
220
|
-
# Priors
|
|
221
|
-
mu1 = pm.Normal('mu_group1', mu=0, sigma=10)
|
|
222
|
-
mu2 = pm.Normal('mu_group2', mu=0, sigma=10)
|
|
223
|
-
sigma = pm.HalfNormal('sigma', sigma=10)
|
|
224
|
-
|
|
225
|
-
# Likelihood
|
|
226
|
-
y1 = pm.Normal('y1', mu=mu1, sigma=sigma, observed=group_a)
|
|
227
|
-
y2 = pm.Normal('y2', mu=mu2, sigma=sigma, observed=group_b)
|
|
228
|
-
|
|
229
|
-
# Derived quantity
|
|
230
|
-
diff = pm.Deterministic('difference', mu1 - mu2)
|
|
231
|
-
|
|
232
|
-
trace = pm.sample(2000, tune=1000)
|
|
233
|
-
|
|
234
|
-
# ArviZ 1.x defaults to 89% intervals; request 95% explicitly for reporting
|
|
235
|
-
print(az.summary(trace, var_names=['difference'], ci_prob=0.95))
|
|
236
|
-
|
|
237
|
-
# Direct probability statement (this is what one-sided questions become)
|
|
238
|
-
prob_greater = np.mean(trace.posterior['difference'].values > 0)
|
|
239
|
-
print(f"P(mu1 > mu2 | data) = {prob_greater:.3f}")
|
|
240
|
-
|
|
241
|
-
# ArviZ 1.x removed az.plot_posterior; use plot_dist (on 0.x, plot_posterior still works)
|
|
242
|
-
az.plot_dist(trace, var_names=['difference'], ci_prob=0.95)
|
|
243
|
-
```
|
|
244
|
-
|
|
245
|
-
Scale priors to the data (e.g., `sigma=10` suits outcomes with SD near 10; use the observed SD as a guide) and state the priors in the report.
|
|
246
|
-
|
|
247
|
-
---
|
|
248
|
-
|
|
249
|
-
## Effect Sizes
|
|
250
|
-
|
|
251
|
-
**Effect sizes quantify magnitude; p-values only indicate existence.** Report one for every test. See `references/effect_sizes_and_power.md` for the full guide.
|
|
252
|
-
|
|
253
|
-
### Quick Reference: Common Effect Sizes
|
|
254
|
-
|
|
255
|
-
| Test | Effect Size | Small | Medium | Large |
|
|
256
|
-
|------|-------------|-------|--------|-------|
|
|
257
|
-
| T-test | Cohen's d | 0.20 | 0.50 | 0.80 |
|
|
258
|
-
| ANOVA | η²_p | 0.01 | 0.06 | 0.14 |
|
|
259
|
-
| Correlation | r | 0.10 | 0.30 | 0.50 |
|
|
260
|
-
| Regression | R² | 0.02 | 0.13 | 0.26 |
|
|
261
|
-
| Chi-square | Cramér's V | 0.07 | 0.21 | 0.35 |
|
|
262
|
-
|
|
263
|
-
Benchmarks are conventions, not laws — a "small" effect can matter enormously (drug side effects) and a "large" one can be trivial. Interpret in context.
|
|
264
|
-
|
|
265
|
-
### Calculating Effect Sizes
|
|
266
|
-
|
|
267
|
-
Pingouin returns effect sizes with its tests (`cohen_d` from `pg.ttest`, `np2` from `pg.anova`, `hedges` from `pg.pairwise_tukey`; `r` from `pg.corr` is already an effect size).
|
|
268
|
-
|
|
269
|
-
### Confidence Intervals for Effect Sizes
|
|
270
|
-
|
|
271
|
-
Report a CI for the effect size to show its precision. Use `pg.compute_esci` (note: `pg.compute_effsize_from_t` returns only the point estimate — it does **not** return a CI):
|
|
272
|
-
|
|
273
|
-
```python
|
|
274
|
-
import pingouin as pg
|
|
275
|
-
|
|
276
|
-
d = pg.compute_effsize(group_a, group_b, eftype='cohen')
|
|
277
|
-
ci_lower, ci_upper = pg.compute_esci(stat=d, nx=len(group_a), ny=len(group_b),
|
|
278
|
-
eftype='cohen', confidence=0.95)
|
|
279
|
-
print(f"d = {d:.2f}, 95% CI [{ci_lower:.2f}, {ci_upper:.2f}]")
|
|
280
|
-
```
|
|
281
|
-
|
|
282
|
-
---
|
|
283
|
-
|
|
284
|
-
## Power Analysis
|
|
285
|
-
|
|
286
|
-
### A Priori Power Analysis (Study Planning)
|
|
287
|
-
|
|
288
|
-
Determine required sample size before data collection:
|
|
289
|
-
|
|
290
|
-
```python
|
|
291
|
-
from statsmodels.stats.power import tt_ind_solve_power, FTestAnovaPower
|
|
292
|
-
|
|
293
|
-
# T-test: What n per group is needed to detect d = 0.5?
|
|
294
|
-
n_required = tt_ind_solve_power(
|
|
295
|
-
effect_size=0.5,
|
|
296
|
-
alpha=0.05,
|
|
297
|
-
power=0.80,
|
|
298
|
-
ratio=1.0,
|
|
299
|
-
alternative='two-sided'
|
|
300
|
-
)
|
|
301
|
-
print(f"Required n per group: {n_required:.0f}")
|
|
302
|
-
|
|
303
|
-
# One-way ANOVA: What n is needed to detect Cohen's f = 0.25?
|
|
304
|
-
# Notes: the parameter is k_groups; effect_size is Cohen's f (f = sqrt(eta2/(1-eta2)));
|
|
305
|
-
# and solve_power returns the TOTAL sample size, not n per group.
|
|
306
|
-
import math
|
|
307
|
-
anova_power = FTestAnovaPower()
|
|
308
|
-
n_total = anova_power.solve_power(
|
|
309
|
-
effect_size=0.25,
|
|
310
|
-
k_groups=3,
|
|
311
|
-
alpha=0.05,
|
|
312
|
-
power=0.80
|
|
313
|
-
)
|
|
314
|
-
print(f"Required total N: {math.ceil(n_total)} ({math.ceil(n_total / 3)} per group)")
|
|
315
|
-
```
|
|
316
|
-
|
|
317
|
-
### Sensitivity Analysis (Post-Study)
|
|
318
|
-
|
|
319
|
-
Determine what effect size the study could detect:
|
|
320
|
-
|
|
321
|
-
```python
|
|
322
|
-
# With n=50 per group, what effect could we detect at 80% power?
|
|
323
|
-
detectable_d = tt_ind_solve_power(
|
|
324
|
-
effect_size=None, # Solve for this
|
|
325
|
-
nobs1=50,
|
|
326
|
-
alpha=0.05,
|
|
327
|
-
power=0.80,
|
|
328
|
-
ratio=1.0,
|
|
329
|
-
alternative='two-sided'
|
|
330
|
-
)
|
|
331
|
-
print(f"Study could detect d >= {detectable_d:.2f}")
|
|
332
|
-
```
|
|
333
|
-
|
|
334
|
-
**Note**: Post-hoc "observed power" (computing power from the observed effect) is circular and misleading — it is a deterministic function of the p-value. If a study is done and someone asks about power, run a sensitivity analysis instead.
|
|
335
|
-
|
|
336
|
-
See `references/effect_sizes_and_power.md` for detailed guidance.
|
|
337
|
-
|
|
338
|
-
---
|
|
339
|
-
|
|
340
|
-
## Reporting Results
|
|
341
|
-
|
|
342
|
-
Follow `references/reporting_standards.md` for APA style. Every report needs:
|
|
343
|
-
|
|
344
|
-
1. **Descriptive statistics**: M, SD, n for all groups/variables
|
|
345
|
-
2. **Test statistics**: Test name, statistic, df, exact p-value (`p = .034`, not `p < .05`; use `p < .001` only below .001)
|
|
346
|
-
3. **Effect sizes**: With confidence intervals
|
|
347
|
-
4. **Assumption checks**: Which tests were run, results, and actions taken
|
|
348
|
-
5. **All planned analyses**: Including non-significant findings — omitting them is cherry-picking
|
|
349
|
-
|
|
350
|
-
### Example Report Templates
|
|
351
|
-
|
|
352
|
-
#### Independent T-Test
|
|
353
|
-
|
|
354
|
-
```
|
|
355
|
-
Group A (n = 48, M = 75.2, SD = 8.5) scored significantly higher than
|
|
356
|
-
Group B (n = 52, M = 68.3, SD = 9.2), t(98) = 3.82, p < .001, d = 0.77,
|
|
357
|
-
95% CI [0.36, 1.18], two-tailed. Assumptions of normality (Shapiro-Wilk:
|
|
358
|
-
Group A W = 0.97, p = .18; Group B W = 0.96, p = .12) and homogeneity
|
|
359
|
-
of variance (Levene's F(1, 98) = 1.23, p = .27) were satisfied.
|
|
360
|
-
```
|
|
361
|
-
|
|
362
|
-
#### One-Way ANOVA
|
|
363
|
-
|
|
364
|
-
```
|
|
365
|
-
A one-way ANOVA revealed a significant main effect of treatment condition
|
|
366
|
-
on test scores, F(2, 147) = 8.45, p < .001, η²_p = .10. Post hoc
|
|
367
|
-
comparisons using Tukey's HSD indicated that Condition A (M = 78.2,
|
|
368
|
-
SD = 7.3) scored significantly higher than Condition B (M = 71.5,
|
|
369
|
-
SD = 8.1, p = .002, d = 0.87) and Condition C (M = 70.1, SD = 7.9,
|
|
370
|
-
p < .001, d = 1.07). Conditions B and C did not differ significantly
|
|
371
|
-
(p = .52, d = 0.18).
|
|
372
|
-
```
|
|
373
|
-
|
|
374
|
-
#### Multiple Regression
|
|
375
|
-
|
|
376
|
-
```
|
|
377
|
-
Multiple linear regression was conducted to predict exam scores from
|
|
378
|
-
study hours, prior GPA, and attendance. The overall model was significant,
|
|
379
|
-
F(3, 146) = 45.2, p < .001, R² = .48, adjusted R² = .47. Study hours
|
|
380
|
-
(B = 1.80, SE = 0.31, β = .35, t = 5.78, p < .001, 95% CI [1.18, 2.42])
|
|
381
|
-
and prior GPA (B = 8.52, SE = 1.95, β = .28, t = 4.37, p < .001,
|
|
382
|
-
95% CI [4.66, 12.38]) were significant predictors, while attendance was
|
|
383
|
-
not (B = 0.15, SE = 0.12, β = .08, t = 1.25, p = .21, 95% CI [-0.09, 0.39]).
|
|
384
|
-
Multicollinearity was not a concern (all VIF < 1.5).
|
|
385
|
-
```
|
|
386
|
-
|
|
387
|
-
#### Bayesian Analysis
|
|
388
|
-
|
|
389
|
-
```
|
|
390
|
-
A Bayesian independent samples t-test was conducted using weakly
|
|
391
|
-
informative priors (Normal(0, 10) for group means). The posterior
|
|
392
|
-
distribution indicated that Group A scored higher than Group B
|
|
393
|
-
(M_diff = 6.8, 95% credible interval [3.2, 10.4]), with a 99.8%
|
|
394
|
-
posterior probability that Group A's mean exceeded Group B's mean.
|
|
395
|
-
Convergence diagnostics were satisfactory (all R-hat < 1.01, ESS > 1000).
|
|
396
|
-
```
|
|
397
|
-
|
|
398
|
-
If a non-parametric test was used, report medians rather than means, the U/W/H statistic, and a rank-based effect size (e.g., rank-biserial correlation, returned by `pg.mwu` as `RBC`).
|
|
399
|
-
|
|
400
|
-
---
|
|
401
|
-
|
|
402
|
-
## Bayesian Statistics
|
|
403
|
-
|
|
404
|
-
Consider Bayesian approaches when:
|
|
405
|
-
- You have prior information to incorporate
|
|
406
|
-
- You want direct probability statements about hypotheses ("there is a 95% probability the effect lies in this interval")
|
|
407
|
-
- Sample size is small or data collection is sequential (no correction needed for optional stopping)
|
|
408
|
-
- You need to quantify evidence *for* the null hypothesis
|
|
409
|
-
- The model is complex (hierarchical structure, missing data)
|
|
410
|
-
|
|
411
|
-
See `references/bayesian_statistics.md` for prior specification, Bayes Factors, credible intervals, hierarchical models, and convergence checking (R-hat < 1.01, sufficient ESS, posterior predictive checks).
|
|
412
|
-
|
|
413
|
-
---
|
|
414
|
-
|
|
415
|
-
## Bundled Resources
|
|
416
|
-
|
|
417
|
-
### References (`references/`)
|
|
418
|
-
|
|
419
|
-
- **test_selection_guide.md**: Decision tree covering group comparisons, relationships, counts, time-to-event, agreement/reliability, and categorical analysis
|
|
420
|
-
- **assumptions_and_diagnostics.md**: Detailed guidance on checking and handling assumption violations
|
|
421
|
-
- **effect_sizes_and_power.md**: Calculating, interpreting, and reporting effect sizes; power analysis
|
|
422
|
-
- **bayesian_statistics.md**: Priors, Bayes Factors, credible intervals, hierarchical models, diagnostics
|
|
423
|
-
- **reporting_standards.md**: APA-style reporting guidelines with worked examples
|
|
424
|
-
|
|
425
|
-
### Scripts (`scripts/`)
|
|
426
|
-
|
|
427
|
-
- **assumption_checks.py**: Automated assumption checking with visualizations
|
|
428
|
-
- `comprehensive_assumption_check()`: outliers + normality + variance homogeneity in one call
|
|
429
|
-
- `check_normality()`, `check_normality_per_group()`: Shapiro-Wilk with Q-Q plots
|
|
430
|
-
- `check_homogeneity_of_variance()`: Levene's test with box plots
|
|
431
|
-
- `check_regression_diagnostics()`: 4-panel residual plots + Shapiro-Wilk, Breusch-Pagan, Durbin-Watson, VIF for fitted OLS models
|
|
432
|
-
- `check_linearity()`, `detect_outliers()`
|
|
433
|
-
|
|
434
|
-
---
|
|
435
|
-
|
|
436
|
-
## Statistical Integrity
|
|
437
|
-
|
|
438
|
-
These are the practices that keep an analysis defensible. They matter because the most common statistical failures are not computational errors — they are silent flexibility (testing until something works) and selective reporting.
|
|
439
|
-
|
|
440
|
-
1. **Distinguish confirmatory from exploratory.** State the planned analysis before running it; label anything discovered along the way as exploratory.
|
|
441
|
-
2. **Don't shop for significance.** If the planned test is non-significant, that is the result. Trying alternative tests, subgroups, or outlier-removal schemes until p < .05 invalidates the p-value.
|
|
442
|
-
3. **Correct for multiple comparisons** when running families of tests (Tukey HSD for post-hoc ANOVA; Holm or Benjamini-Hochberg FDR for other families) and say which correction was used.
|
|
443
|
-
4. **A non-significant result is not evidence of no effect.** With small n, the study may simply have been underpowered — run a sensitivity analysis, or use a Bayesian analysis / equivalence test to actually quantify support for the null.
|
|
444
|
-
5. **Statistical significance is not practical importance.** With large n, trivial effects reach p < .001. Lead the interpretation with the effect size.
|
|
445
|
-
6. **Understand missing data before dropping rows.** Listwise deletion is only safe when data are missing completely at random; otherwise consider multiple imputation and say what was done.
|
|
446
|
-
7. **Make it reproducible.** Set random seeds, report library versions for simulation-based methods, and keep the analysis in a runnable script.
|
|
@@ -1,200 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: statistical-power
|
|
3
|
-
description: Sample-size and statistical power calculations for planning studies. Use whenever someone asks "how many subjects/samples/replicates do I need", wants an a priori power analysis, a minimum detectable effect (MDE), a power curve, or needs to justify a sample size for a grant, IRB protocol, or pre-registration. Covers closed-form power for t-tests, ANOVA, proportions, correlations, chi-square, and regression, plus simulation-based (Monte Carlo) power for designs with no formula — logistic/Poisson regression, mixed models, cluster-randomized trials, survival, and interactions. Use this skill even when the request only mentions an effect size, alpha, or "80% power" without saying "power analysis" explicitly. For laying out the study (randomization, blocking, factorial/DOE, crossover, sequential designs) use experimental-design; for analyzing data already collected and reporting it use statistical-analysis.
|
|
4
|
-
allowed-tools: Read Write Edit Bash
|
|
5
|
-
compatibility: Requires Python >=3.10. Examples target statsmodels >=0.14.6, scipy >=1.11, pingouin >=0.6, numpy >=1.26, and matplotlib. Optional extras are statsmodels mixed models and lifelines for simulation-based power.
|
|
6
|
-
license: MIT license
|
|
7
|
-
metadata:
|
|
8
|
-
version: "1.0"
|
|
9
|
-
skill-author: K-Dense Inc.
|
|
10
|
-
---
|
|
11
|
-
|
|
12
|
-
# Statistical Power & Sample Size
|
|
13
|
-
|
|
14
|
-
## Overview
|
|
15
|
-
|
|
16
|
-
Power analysis answers one of the most consequential questions in study planning: **how large a sample do you need to reliably detect an effect of a given size, and what could you detect with the sample you can afford?** An underpowered study wastes resources and produces inconclusive or irreproducible results; an overpowered one wastes participants, money, and (in clinical work) exposes more people to risk than necessary. Getting this right *before* data collection is the single highest-leverage statistical decision in a project.
|
|
17
|
-
|
|
18
|
-
Four quantities are locked together for any given test: **sample size (n)**, **effect size**, **significance level (α)**, and **power (1 − β)**. Fix any three and the fourth is determined. Every calculation in this skill is some rearrangement of that relationship.
|
|
19
|
-
|
|
20
|
-
This skill covers the two ways to do power analysis:
|
|
21
|
-
- **Closed-form** formulas (fast, exact for standard tests) — see `references/closed_form_recipes.md`.
|
|
22
|
-
- **Simulation / Monte Carlo** (works for *any* design or model you can simulate and analyze) — see `references/simulation_based_power.md`.
|
|
23
|
-
|
|
24
|
-
For choosing and converting effect sizes — usually the hardest part — see `references/effect_sizes.md`.
|
|
25
|
-
|
|
26
|
-
## When to Use This Skill
|
|
27
|
-
|
|
28
|
-
- Determining required sample size before collecting data (a priori power analysis)
|
|
29
|
-
- Finding the minimum detectable effect (MDE) for a fixed, already-determined sample size
|
|
30
|
-
- Producing power curves (power vs. n, or power vs. effect size) for a grant or protocol
|
|
31
|
-
- Justifying a sample size for an IRB submission, grant, or pre-registration
|
|
32
|
-
- Powering designs with unequal group sizes or non-1:1 allocation
|
|
33
|
-
- Powering anything without a textbook formula (mixed models, logistic/Poisson regression, cluster-randomized trials, survival analysis, mediation, interactions) via simulation
|
|
34
|
-
- Accounting for multiple comparisons, attrition/dropout, or clustering in the sample-size estimate
|
|
35
|
-
|
|
36
|
-
## Installation
|
|
37
|
-
|
|
38
|
-
Use **uv**. Pin versions in production; unpinned is fine for exploration.
|
|
39
|
-
|
|
40
|
-
```bash
|
|
41
|
-
uv pip install "statsmodels>=0.14.6" "scipy>=1.11" "pingouin>=0.6" "numpy>=1.26" matplotlib pandas
|
|
42
|
-
# For simulation-based power of advanced models (optional, add as needed):
|
|
43
|
-
uv pip install lifelines # survival
|
|
44
|
-
# mixed models and GLMs come with statsmodels
|
|
45
|
-
```
|
|
46
|
-
|
|
47
|
-
**Compatibility note:** use `statsmodels>=0.14.6` with `scipy>=1.11` to avoid `_lazywhere` import errors on SciPy 1.16+. Pingouin 0.5+ renamed power-function arguments to match the names used below.
|
|
48
|
-
|
|
49
|
-
---
|
|
50
|
-
|
|
51
|
-
## The one decision that drives everything: the effect size
|
|
52
|
-
|
|
53
|
-
Power calculations are only as trustworthy as the effect size you feed them. **Do not invent a number.** Use, in rough order of preference:
|
|
54
|
-
|
|
55
|
-
1. A **minimally important effect** — the smallest effect that would actually change a decision or matter scientifically/clinically (the "smallest effect size of interest", SESOI). This is the most defensible basis: you power to detect what matters, not what you hope to see.
|
|
56
|
-
2. A **pilot or prior-study estimate**, but shrink it — published and pilot effects are inflated by publication bias and the winner's curse. Powering on a raw pilot estimate routinely underpowers the real study.
|
|
57
|
-
3. A **convention** (Cohen's small/medium/large) only as a last resort, and say so explicitly.
|
|
58
|
-
|
|
59
|
-
Whatever you pick, run a **sensitivity analysis**: report how required n changes across a plausible range of effect sizes, not a single point. A power analysis presented as one number hides its biggest source of uncertainty. See `references/effect_sizes.md` for benchmarks and conversions between d, f, r, η², odds ratios, and Cohen's h/w.
|
|
60
|
-
|
|
61
|
-
> **Avoid post-hoc ("observed") power.** Computing power from the effect size you just estimated is circular: it is a deterministic function of the p-value and tells you nothing new. If a study is already done and you want to know what it could have detected, report a **sensitivity analysis** (MDE at the achieved n) or, better, the confidence interval around the observed effect. This is a common reviewer complaint — do not produce observed power even if asked without flagging the issue.
|
|
62
|
-
|
|
63
|
-
---
|
|
64
|
-
|
|
65
|
-
## Quick recipes (closed-form)
|
|
66
|
-
|
|
67
|
-
The bundled `scripts/power.py` wraps statsmodels into one consistent interface so you don't have to remember which solver belongs to which test. Run from the skill directory or add `scripts/` to `sys.path`.
|
|
68
|
-
|
|
69
|
-
```python
|
|
70
|
-
from power import sample_size, power, mde, power_curve
|
|
71
|
-
|
|
72
|
-
# 1. How many per group to detect Cohen's d = 0.5, two-sided, 80% power?
|
|
73
|
-
sample_size(test="t_ind", effect_size=0.5, power=0.80, alpha=0.05)
|
|
74
|
-
# -> required n per group
|
|
75
|
-
|
|
76
|
-
# 2. Two groups, 3:1 allocation (e.g. more controls than cases)
|
|
77
|
-
sample_size(test="t_ind", effect_size=0.5, power=0.80, ratio=3.0)
|
|
78
|
-
|
|
79
|
-
# 3. Fixed n=30/group — what's the minimum detectable d at 80% power?
|
|
80
|
-
mde(test="t_ind", nobs1=30, power=0.80, alpha=0.05)
|
|
81
|
-
|
|
82
|
-
# 4. One-way ANOVA, 4 groups, detect Cohen's f = 0.25
|
|
83
|
-
sample_size(test="anova", effect_size=0.25, k_groups=4, power=0.80)
|
|
84
|
-
|
|
85
|
-
# 5. Two proportions: 0.40 vs 0.55 (auto-converts to Cohen's h)
|
|
86
|
-
sample_size(test="two_proportions", prop1=0.40, prop2=0.55, power=0.80)
|
|
87
|
-
|
|
88
|
-
# 6. Correlation: detect r = 0.30
|
|
89
|
-
sample_size(test="correlation", effect_size=0.30, power=0.80)
|
|
90
|
-
|
|
91
|
-
# 7. Power curve for the grant figure
|
|
92
|
-
power_curve(test="t_ind", effect_size=0.5, n_range=range(10, 120, 5),
|
|
93
|
-
save="power_curve.png")
|
|
94
|
-
```
|
|
95
|
-
|
|
96
|
-
Supported `test=` values: `t_ind` (two independent means), `t_paired`/`t_one` (paired or one-sample mean), `anova` (one-way), `two_proportions`, `one_proportion`, `correlation`, `chi2` (goodness-of-fit / contingency via effect size *w*), `linear_regression` (R² increment / f²). Full argument tables and the underlying statsmodels calls are in `references/closed_form_recipes.md`.
|
|
97
|
-
|
|
98
|
-
---
|
|
99
|
-
|
|
100
|
-
## When there is no formula: simulate
|
|
101
|
-
|
|
102
|
-
Closed-form power exists only for a handful of simple tests. For **logistic/Poisson regression, mixed-effects / repeated-measures models, cluster-randomized trials, survival analysis, mediation, multi-way interactions, or any non-standard analysis**, the right tool is simulation. The logic is always the same three steps:
|
|
103
|
-
|
|
104
|
-
1. **Simulate** a dataset from your assumed truth (the effect you want to detect, plus realistic noise, baseline rates, cluster structure, etc.).
|
|
105
|
-
2. **Analyze** it with the *exact* test/model you plan to use on the real data.
|
|
106
|
-
3. **Repeat** many times (≥1,000; 5,000–10,000 for a stable estimate near 80%). Power is the fraction of replicates in which the test is significant.
|
|
107
|
-
|
|
108
|
-
`scripts/simulate_power.py` provides a reusable harness plus worked examples (two-group difference, logistic regression, cluster-randomized trial with an ICC, and a linear mixed model). The core is just:
|
|
109
|
-
|
|
110
|
-
```python
|
|
111
|
-
from simulate_power import simulate_power
|
|
112
|
-
|
|
113
|
-
def gen_and_test(n, rng):
|
|
114
|
-
# build a dataset of size n under the assumed effect, run the planned test,
|
|
115
|
-
# return True if the result is significant
|
|
116
|
-
...
|
|
117
|
-
|
|
118
|
-
est = simulate_power(gen_and_test, n=200, n_sims=2000, alpha=0.05)
|
|
119
|
-
print(f"Power at n=200: {est.power:.3f} (95% CI {est.ci_low:.3f}-{est.ci_high:.3f})")
|
|
120
|
-
```
|
|
121
|
-
|
|
122
|
-
Report the **Monte Carlo confidence interval** on the estimate (the harness returns it) so the reader knows whether 0.81 vs. 0.79 is signal or simulation noise. See `references/simulation_based_power.md` for the full patterns, including how to search for the n that hits target power and how to model dropout and clustering.
|
|
123
|
-
|
|
124
|
-
---
|
|
125
|
-
|
|
126
|
-
## Adjustments people forget
|
|
127
|
-
|
|
128
|
-
These routinely make the difference between an adequately powered study and an underpowered one. Apply them explicitly and state that you did.
|
|
129
|
-
|
|
130
|
-
- **Multiple comparisons.** If the analysis tests *m* hypotheses with a Bonferroni-style correction, power each test at the corrected α (e.g. α/m), which raises n. Better: power on the family-wise or FDR-controlled procedure directly via simulation. Ignoring this silently underpowers every secondary endpoint.
|
|
131
|
-
- **Attrition / dropout / unusable samples.** Power gives the n you need *analyzed*. Inflate the *enrolled* n: `n_enroll = ceil(n_analyzed / (1 − dropout_rate))`. A 20% dropout rate means enrolling 25% more than the formula returns.
|
|
132
|
-
- **Clustering (design effect).** When observations are nested (patients within clinics, cells within animals, repeated measures within subject), the effective sample size is smaller than the raw count. Inflate by the design effect `DEFF = 1 + (m − 1)·ICC`, where *m* is cluster size and ICC the intraclass correlation. Treating clustered data as independent is **pseudoreplication** and badly overstates power — for cluster-randomized designs, simulate instead.
|
|
133
|
-
- **One- vs. two-sided.** Two-sided is the default and almost always the right choice; a one-sided test buys power only by refusing to detect an effect in the unexpected direction. Justify any one-sided test.
|
|
134
|
-
- **Unequal allocation.** Equal groups are most efficient for a fixed total n. If allocation is fixed by design (e.g. 2:1 treatment:control), pass `ratio=` so the calculation reflects it.
|
|
135
|
-
|
|
136
|
-
---
|
|
137
|
-
|
|
138
|
-
## Workflow
|
|
139
|
-
|
|
140
|
-
1. **State the design and the planned analysis.** The test you will run determines the power method. If the analysis is a mixed model or GLM, go straight to simulation.
|
|
141
|
-
2. **Choose the effect size** on a defensible basis (SESOI > shrunk pilot > convention) and write down the justification.
|
|
142
|
-
3. **Set α and target power.** Conventional defaults are α = 0.05 (two-sided) and power = 0.80; 0.90 is common for confirmatory/clinical work. State them.
|
|
143
|
-
4. **Compute** with `scripts/power.py` (closed-form) or `scripts/simulate_power.py` (simulation).
|
|
144
|
-
5. **Sensitivity analysis.** Recompute across a range of plausible effect sizes and produce a power curve. This is the deliverable, not a single number.
|
|
145
|
-
6. **Apply adjustments** for dropout, clustering, and multiplicity.
|
|
146
|
-
7. **Report** following the template below.
|
|
147
|
-
|
|
148
|
-
---
|
|
149
|
-
|
|
150
|
-
## Reporting template
|
|
151
|
-
|
|
152
|
-
A defensible power statement contains every input, so a reader could reproduce it. Adapt:
|
|
153
|
-
|
|
154
|
-
```
|
|
155
|
-
A priori power analysis was conducted to determine the sample size needed to detect
|
|
156
|
-
a [between-group difference of Cohen's d = 0.50], which we considered the smallest
|
|
157
|
-
effect of clinical interest. With α = .05 (two-sided) and power = .80, a two-sample
|
|
158
|
-
t-test requires n = 64 per group (128 total; computed with statsmodels 0.14).
|
|
159
|
-
Allowing for 20% attrition, we will enrol 160 participants. A sensitivity analysis
|
|
160
|
-
showed required n ranges from 45 to 105 per group across plausible effects
|
|
161
|
-
d = 0.40–0.60 (Figure X).
|
|
162
|
-
```
|
|
163
|
-
|
|
164
|
-
For simulation: also state the data-generating assumptions (baseline rate, residual SD, ICC, cluster sizes), the number of simulations, and the Monte Carlo CI.
|
|
165
|
-
|
|
166
|
-
---
|
|
167
|
-
|
|
168
|
-
## Common pitfalls
|
|
169
|
-
|
|
170
|
-
1. **Inventing the effect size** or copying an inflated pilot estimate — the most common way power analyses go wrong.
|
|
171
|
-
2. **Reporting a single n** instead of a sensitivity range / power curve.
|
|
172
|
-
3. **Post-hoc / observed power** — circular and uninformative; use sensitivity analysis or the effect-size CI instead.
|
|
173
|
-
4. **Ignoring clustering** (pseudoreplication) — counting cells/measurements as if they were independent subjects.
|
|
174
|
-
5. **Forgetting dropout** — powering the analyzed n but enrolling the same number.
|
|
175
|
-
6. **Confusing α with power**, or one-sided with two-sided.
|
|
176
|
-
7. **Powering only the primary endpoint** while reporting secondary/interaction tests that need far larger n.
|
|
177
|
-
8. **Using a t-test formula for a model you won't actually fit** (e.g. planning a logistic regression with a means-based calculation) — match the power method to the planned analysis.
|
|
178
|
-
|
|
179
|
-
---
|
|
180
|
-
|
|
181
|
-
## Resources
|
|
182
|
-
|
|
183
|
-
### Scripts
|
|
184
|
-
- `scripts/power.py` — unified closed-form interface (`sample_size`, `power`, `mde`, `power_curve`) over statsmodels/pingouin for all standard tests.
|
|
185
|
-
- `scripts/simulate_power.py` — Monte Carlo power harness with `simulate_power()` and `find_sample_size()`, plus worked examples (two-group, logistic regression, cluster-randomized, linear mixed model).
|
|
186
|
-
|
|
187
|
-
### References
|
|
188
|
-
- `references/closed_form_recipes.md` — per-test argument tables and exact statsmodels/pingouin calls, including proportions, chi-square, and regression.
|
|
189
|
-
- `references/simulation_based_power.md` — full simulation patterns for GLMs, mixed models, cluster designs, survival, and dropout.
|
|
190
|
-
- `references/effect_sizes.md` — choosing effect sizes (SESOI), Cohen's benchmarks, and conversions between d, f, r, η²/f², OR, h, and w.
|
|
191
|
-
|
|
192
|
-
### Related skills
|
|
193
|
-
- **experimental-design** — once you know n, lay out the actual study (randomization, blocking, factorial/DOE, crossover, sequential designs).
|
|
194
|
-
- **statistical-analysis** — assumption checks, running the test, effect sizes, and APA reporting after data collection.
|
|
195
|
-
- **statsmodels** / **pymc** — fitting the models referenced here.
|
|
196
|
-
|
|
197
|
-
### Key references
|
|
198
|
-
- Cohen, J. (1988). *Statistical Power Analysis for the Behavioral Sciences* (2nd ed.).
|
|
199
|
-
- Lakens, D. (2022). *Sample Size Justification*. Collabra: Psychology, 8(1).
|
|
200
|
-
- Arnold, B. F. et al. (2011). Simulation methods to estimate design power. *BMC Medical Research Methodology*, 11:94.
|