clearai-dsh 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +26 -0
- package/LICENSE +201 -0
- package/README.md +138 -0
- package/README.zh-CN.md +138 -0
- package/bin/clearai.mjs +224 -0
- package/brand/README.md +41 -0
- package/brand/logo-512-dark.png +0 -0
- package/brand/logo-512.png +0 -0
- package/brand/logo-lockup-dark.png +0 -0
- package/brand/logo-lockup.png +0 -0
- package/brand/logo-lockup.svg +12 -0
- package/brand/logo-wordmark.svg +6 -0
- package/brand/logo.svg +19 -0
- package/cordis.patch.yml +39 -0
- package/lib/client.js +3071 -0
- package/lib/fold.js +1576 -0
- package/lib/host.js +605 -0
- package/package.json +65 -0
- package/presets/clearai/agent.cordis.yml +226 -0
- package/presets/clearai/plugins/brain.js +547 -0
- package/presets/clearai/plugins/clearai-kernel.js +5485 -0
- package/presets/clearai/plugins/ontology.js +306 -0
- package/presets/clearai/plugins/prompts.js +312 -0
- package/presets/clearai/preset.yml +5 -0
- package/presets/clearai/skills/clearai-loop/SKILL.md +89 -0
- package/presets/clearai/template/knowledge/README.md +25 -0
- package/presets/clearai/template/memory/README.md +34 -0
- package/presets/clearai/template/project.md +49 -0
- package/presets/clearai/template/skills/README.md +37 -0
- package/presets/clearai/template/skills/chart-diagram-qa/SKILL.md +43 -0
- package/presets/clearai/template/skills/citation-management/SKILL.md +73 -0
- package/presets/clearai/template/skills/citation-management/references/bibtex_formatting.md +908 -0
- package/presets/clearai/template/skills/citation-management/references/citation_validation.md +794 -0
- package/presets/clearai/template/skills/citation-management/references/google_scholar_search.md +725 -0
- package/presets/clearai/template/skills/citation-management/references/metadata_extraction.md +870 -0
- package/presets/clearai/template/skills/citation-management/references/pubmed_search.md +839 -0
- package/presets/clearai/template/skills/citation-management/scripts/doi_to_bibtex.py +204 -0
- package/presets/clearai/template/skills/citation-management/scripts/extract_metadata.py +569 -0
- package/presets/clearai/template/skills/citation-management/scripts/format_bibtex.py +349 -0
- package/presets/clearai/template/skills/citation-management/scripts/generate_schematic.py +139 -0
- package/presets/clearai/template/skills/citation-management/scripts/generate_schematic_ai.py +817 -0
- package/presets/clearai/template/skills/citation-management/scripts/search_google_scholar.py +282 -0
- package/presets/clearai/template/skills/citation-management/scripts/search_pubmed.py +398 -0
- package/presets/clearai/template/skills/citation-management/scripts/validate_citations.py +497 -0
- package/presets/clearai/template/skills/data-analysis/SKILL.md +92 -0
- package/presets/clearai/template/skills/data-analysis/checklists/readiness_check.md +23 -0
- package/presets/clearai/template/skills/data-analysis/templates/analysis_report.md.tpl +63 -0
- package/presets/clearai/template/skills/data-analysis/templates/cleaning_rules_draft.yaml.tpl +32 -0
- package/presets/clearai/template/skills/data-analysis/templates/data_dictionary.md.tpl +12 -0
- package/presets/clearai/template/skills/data-analysis/templates/domain_knowledge_template.md.tpl +316 -0
- package/presets/clearai/template/skills/data-analysis/templates/feature_candidates.json.tpl +20 -0
- package/presets/clearai/template/skills/data-analysis/templates/quality_scorecard.md.tpl +30 -0
- package/presets/clearai/template/skills/data-analysis/workflows/01-data-profiling.md +42 -0
- package/presets/clearai/template/skills/data-analysis/workflows/02-quality-audit.md +36 -0
- package/presets/clearai/template/skills/data-analysis/workflows/03-physical-correlation.md +25 -0
- package/presets/clearai/template/skills/data-analysis/workflows/04-unstructured-mining.md +26 -0
- package/presets/clearai/template/skills/data-qa-analysis/SKILL.md +102 -0
- package/presets/clearai/template/skills/data-qa-analysis/checklists/readiness_check.md +62 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/best_in_class_report.md.tpl +56 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/cleaning_rules_draft.yaml.tpl +56 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/data_dictionary.md.tpl +13 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/data_source_inventory_and_lineage.md.tpl +146 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/data_status_report.md.tpl +60 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/steady_state_rules.yaml.tpl +41 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/subsystem_registry.md.tpl +101 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/unified_execution_plan.md.tpl +100 -0
- package/presets/clearai/template/skills/data-qa-analysis/workflows/01-data-source-inventory-and-lineage.md +194 -0
- package/presets/clearai/template/skills/data-qa-analysis/workflows/02-data-alignment-and-tag-semantics.md +122 -0
- package/presets/clearai/template/skills/data-qa-analysis/workflows/03-steady-state-identification.md +126 -0
- package/presets/clearai/template/skills/data-qa-analysis/workflows/04-consumption-analysis.md +152 -0
- package/presets/clearai/template/skills/data-qa-analysis/workflows/05-best-in-class-and-optimization-space.md +78 -0
- package/presets/clearai/template/skills/domain-presearch/SKILL.md +131 -0
- package/presets/clearai/template/skills/domain-presearch/checklists/domain_checklist.md +24 -0
- package/presets/clearai/template/skills/domain-presearch/references/figure_code.md +78 -0
- package/presets/clearai/template/skills/domain-presearch/references/strategic_frameworks.md +38 -0
- package/presets/clearai/template/skills/exploration-loop/SKILL.md +81 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/SKILL.md +77 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/references/bioinformatics_genomics_formats.md +664 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/references/chemistry_molecular_formats.md +664 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/references/general_scientific_formats.md +518 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/references/microscopy_imaging_formats.md +620 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/references/proteomics_metabolomics_formats.md +517 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/references/spectroscopy_analytical_formats.md +633 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/scripts/eda_analyzer.py +547 -0
- package/presets/clearai/template/skills/hypothesis-generation/SKILL.md +73 -0
- package/presets/clearai/template/skills/hypothesis-generation/references/experimental_design_patterns.md +329 -0
- package/presets/clearai/template/skills/hypothesis-generation/references/hypothesis_quality_criteria.md +198 -0
- package/presets/clearai/template/skills/hypothesis-generation/references/literature_search_strategies.md +622 -0
- package/presets/clearai/template/skills/hypothesis-generation/scripts/generate_schematic.py +139 -0
- package/presets/clearai/template/skills/hypothesis-generation/scripts/generate_schematic_ai.py +817 -0
- package/presets/clearai/template/skills/literature-review/SKILL.md +72 -0
- package/presets/clearai/template/skills/literature-review/references/citation_styles.md +166 -0
- package/presets/clearai/template/skills/literature-review/references/database_strategies.md +455 -0
- package/presets/clearai/template/skills/literature-review/scripts/generate_pdf.py +176 -0
- package/presets/clearai/template/skills/literature-review/scripts/generate_schematic.py +139 -0
- package/presets/clearai/template/skills/literature-review/scripts/generate_schematic_ai.py +817 -0
- package/presets/clearai/template/skills/literature-review/scripts/search_databases.py +303 -0
- package/presets/clearai/template/skills/literature-review/scripts/verify_citations.py +221 -0
- package/presets/clearai/template/skills/paper-lookup/SKILL.md +59 -0
- package/presets/clearai/template/skills/paper-lookup/references/arxiv.md +161 -0
- package/presets/clearai/template/skills/paper-lookup/references/biorxiv.md +118 -0
- package/presets/clearai/template/skills/paper-lookup/references/core.md +150 -0
- package/presets/clearai/template/skills/paper-lookup/references/crossref.md +181 -0
- package/presets/clearai/template/skills/paper-lookup/references/medrxiv.md +104 -0
- package/presets/clearai/template/skills/paper-lookup/references/openalex.md +174 -0
- package/presets/clearai/template/skills/paper-lookup/references/pmc.md +152 -0
- package/presets/clearai/template/skills/paper-lookup/references/pubmed.md +124 -0
- package/presets/clearai/template/skills/paper-lookup/references/semantic-scholar.md +203 -0
- package/presets/clearai/template/skills/paper-lookup/references/unpaywall.md +127 -0
- package/presets/clearai/template/skills/process-presearch/SKILL.md +196 -0
- package/presets/clearai/template/skills/process-presearch/checklists/process_checklist.md +18 -0
- package/presets/clearai/template/skills/process-presearch/references/figure_code.md +107 -0
- package/presets/clearai/template/skills/process-presearch/references/source_attribution_example.md +22 -0
- package/presets/clearai/template/skills/process-understanding-extraction/SKILL.md +69 -0
- package/presets/clearai/template/skills/process-understanding-extraction/checklists/readiness_check.md +34 -0
- package/presets/clearai/template/skills/process-understanding-extraction/templates/docx_raw_dump_extractor.py.tpl +132 -0
- package/presets/clearai/template/skills/process-understanding-extraction/templates/entity_map_unit_topology.json.tpl +86 -0
- package/presets/clearai/template/skills/process-understanding-extraction/templates/process_brief.md.tpl +89 -0
- package/presets/clearai/template/skills/process-understanding-extraction/templates/process_brief_builder_from_raw_dump.py.tpl +203 -0
- package/presets/clearai/template/skills/process-understanding-extraction/templates/process_flow_mermaid.md.tpl +41 -0
- package/presets/clearai/template/skills/process-understanding-extraction/templates/unified_execution_plan.md.tpl +53 -0
- package/presets/clearai/template/skills/process-understanding-extraction/workflows/01-process-doc-discovery.md +173 -0
- package/presets/clearai/template/skills/process-understanding-extraction/workflows/02-process-understanding-and-diagramming.md +106 -0
- package/presets/clearai/template/skills/scientific-brainstorming/SKILL.md +64 -0
- package/presets/clearai/template/skills/scientific-brainstorming/references/brainstorming_methods.md +326 -0
- package/presets/clearai/template/skills/scientific-critical-thinking/SKILL.md +72 -0
- package/presets/clearai/template/skills/scientific-critical-thinking/references/common_biases.md +364 -0
- package/presets/clearai/template/skills/scientific-critical-thinking/references/evidence_hierarchy.md +485 -0
- package/presets/clearai/template/skills/scientific-critical-thinking/references/experimental_design.md +496 -0
- package/presets/clearai/template/skills/scientific-critical-thinking/references/logical_fallacies.md +478 -0
- package/presets/clearai/template/skills/scientific-critical-thinking/references/scientific_method.md +169 -0
- package/presets/clearai/template/skills/scientific-critical-thinking/references/statistical_pitfalls.md +506 -0
- package/presets/clearai/template/skills/skill-creator/SKILL.md +109 -0
- package/presets/clearai/template/skills/skill-creator/references/authoring-guide.md +89 -0
- package/presets/clearai/template/skills/statistical-analysis/SKILL.md +79 -0
- package/presets/clearai/template/skills/statistical-analysis/references/assumptions_and_diagnostics.md +369 -0
- package/presets/clearai/template/skills/statistical-analysis/references/bayesian_statistics.md +653 -0
- package/presets/clearai/template/skills/statistical-analysis/references/effect_sizes_and_power.md +578 -0
- package/presets/clearai/template/skills/statistical-analysis/references/reporting_standards.md +469 -0
- package/presets/clearai/template/skills/statistical-analysis/references/test_selection_guide.md +129 -0
- package/presets/clearai/template/skills/statistical-analysis/scripts/assumption_checks.py +538 -0
- package/presets/clearai/template/skills/web-artifact/SKILL.md +165 -0
- package/presets/clearai/template/skills/web-artifact/assets/renderer/renderer.css +229 -0
- package/presets/clearai/template/skills/web-artifact/assets/renderer/renderer.js +373 -0
- package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/LICENSE +263 -0
- package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/UPSTREAM.md +26 -0
- package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/elk.bundled.js +6605 -0
- package/presets/clearai/template/skills/web-artifact/references/when-drawing-a-topology.md +150 -0
- package/presets/clearai/template/skills/web-artifact/references/when-the-page-must-work-offline.md +62 -0
- package/presets/clearai/template/skills/web-artifact/scripts/check_artifact.py +167 -0
- package/presets/clearai/template/skills/web-artifact/scripts/render_topology.js +272 -0
- package/presets/clearai/template/skills/what-if-oracle/LICENSE.txt +5 -0
- package/presets/clearai/template/skills/what-if-oracle/SKILL.md +72 -0
- package/presets/clearai/template/skills/what-if-oracle/references/scenario-templates.md +154 -0
package/presets/clearai/template/skills/statistical-analysis/references/effect_sizes_and_power.md
ADDED
|
@@ -0,0 +1,578 @@
|
|
|
1
|
+
# Effect Sizes and Power Analysis
|
|
2
|
+
|
|
3
|
+
This document provides guidance on calculating, interpreting, and reporting effect sizes, as well as conducting power analyses for study planning.
|
|
4
|
+
|
|
5
|
+
## Why Effect Sizes Matter
|
|
6
|
+
|
|
7
|
+
1. **Statistical significance ≠ practical significance**: p-values only tell if an effect exists, not how large it is
|
|
8
|
+
2. **Sample size dependent**: With large samples, trivial effects become "significant"
|
|
9
|
+
3. **Interpretation**: Effect sizes provide magnitude and practical importance
|
|
10
|
+
4. **Meta-analysis**: Effect sizes enable combining results across studies
|
|
11
|
+
5. **Power analysis**: Required for sample size determination
|
|
12
|
+
|
|
13
|
+
**Golden rule**: ALWAYS report effect sizes alongside p-values.
|
|
14
|
+
|
|
15
|
+
---
|
|
16
|
+
|
|
17
|
+
## Effect Sizes by Analysis Type
|
|
18
|
+
|
|
19
|
+
### T-Tests and Mean Differences
|
|
20
|
+
|
|
21
|
+
#### Cohen's d (Standardized Mean Difference)
|
|
22
|
+
|
|
23
|
+
**Formula**:
|
|
24
|
+
- Independent groups: d = (M₁ - M₂) / SD_pooled
|
|
25
|
+
- Paired groups: d = M_diff / SD_diff
|
|
26
|
+
|
|
27
|
+
**Interpretation** (Cohen, 1988):
|
|
28
|
+
- Small: |d| = 0.20
|
|
29
|
+
- Medium: |d| = 0.50
|
|
30
|
+
- Large: |d| = 0.80
|
|
31
|
+
|
|
32
|
+
**Context-dependent interpretation**:
|
|
33
|
+
- In education: d = 0.40 is typical for successful interventions
|
|
34
|
+
- In psychology: d = 0.40 is considered meaningful
|
|
35
|
+
- In medicine: Small effect sizes can be clinically important
|
|
36
|
+
|
|
37
|
+
**Python calculation**:
|
|
38
|
+
```python
|
|
39
|
+
import pingouin as pg
|
|
40
|
+
import numpy as np
|
|
41
|
+
|
|
42
|
+
# Independent t-test with effect size
|
|
43
|
+
result = pg.ttest(group1, group2, correction=False)
|
|
44
|
+
cohens_d = result['cohen_d'].values[0]
|
|
45
|
+
|
|
46
|
+
# Manual calculation
|
|
47
|
+
mean_diff = np.mean(group1) - np.mean(group2)
|
|
48
|
+
pooled_std = np.sqrt((np.var(group1, ddof=1) + np.var(group2, ddof=1)) / 2)
|
|
49
|
+
cohens_d = mean_diff / pooled_std
|
|
50
|
+
|
|
51
|
+
# Paired t-test
|
|
52
|
+
result = pg.ttest(pre, post, paired=True)
|
|
53
|
+
cohens_d = result['cohen_d'].values[0]
|
|
54
|
+
```
|
|
55
|
+
|
|
56
|
+
**Confidence intervals for d**:
|
|
57
|
+
```python
|
|
58
|
+
from pingouin import compute_effsize_from_t
|
|
59
|
+
|
|
60
|
+
d, ci = compute_effsize_from_t(t_statistic, nx=n1, ny=n2, eftype='cohen')
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
---
|
|
64
|
+
|
|
65
|
+
#### Hedges' g (Bias-Corrected d)
|
|
66
|
+
|
|
67
|
+
**Why use it**: Cohen's d has slight upward bias with small samples (n < 20)
|
|
68
|
+
|
|
69
|
+
**Formula**: g = d × correction_factor, where correction_factor = 1 - 3/(4df - 1)
|
|
70
|
+
|
|
71
|
+
**Python calculation**:
|
|
72
|
+
```python
|
|
73
|
+
result = pg.ttest(group1, group2, correction=False)
|
|
74
|
+
hedges_g = result['hedges'].values[0]
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
**Use Hedges' g when**:
|
|
78
|
+
- Sample sizes are small (n < 20 per group)
|
|
79
|
+
- Conducting meta-analyses (standard in meta-analysis)
|
|
80
|
+
|
|
81
|
+
---
|
|
82
|
+
|
|
83
|
+
#### Glass's Δ (Delta)
|
|
84
|
+
|
|
85
|
+
**When to use**: When one group is a control with known variability
|
|
86
|
+
|
|
87
|
+
**Formula**: Δ = (M₁ - M₂) / SD_control
|
|
88
|
+
|
|
89
|
+
**Use cases**:
|
|
90
|
+
- Clinical trials (use control group SD)
|
|
91
|
+
- When treatment affects variability
|
|
92
|
+
|
|
93
|
+
---
|
|
94
|
+
|
|
95
|
+
### ANOVA
|
|
96
|
+
|
|
97
|
+
#### Eta-squared (η²)
|
|
98
|
+
|
|
99
|
+
**What it measures**: Proportion of total variance explained by factor
|
|
100
|
+
|
|
101
|
+
**Formula**: η² = SS_effect / SS_total
|
|
102
|
+
|
|
103
|
+
**Interpretation**:
|
|
104
|
+
- Small: η² = 0.01 (1% of variance)
|
|
105
|
+
- Medium: η² = 0.06 (6% of variance)
|
|
106
|
+
- Large: η² = 0.14 (14% of variance)
|
|
107
|
+
|
|
108
|
+
**Limitation**: Biased with multiple factors (sums to > 1.0)
|
|
109
|
+
|
|
110
|
+
**Python calculation**:
|
|
111
|
+
```python
|
|
112
|
+
import pingouin as pg
|
|
113
|
+
|
|
114
|
+
# One-way ANOVA
|
|
115
|
+
aov = pg.anova(dv='value', between='group', data=df)
|
|
116
|
+
eta_squared = aov['SS'][0] / aov['SS'].sum()
|
|
117
|
+
|
|
118
|
+
# Or use pingouin directly
|
|
119
|
+
aov = pg.anova(dv='value', between='group', data=df, detailed=True)
|
|
120
|
+
eta_squared = aov['np2'][0] # Note: pingouin reports partial eta-squared
|
|
121
|
+
```
|
|
122
|
+
|
|
123
|
+
---
|
|
124
|
+
|
|
125
|
+
#### Partial Eta-squared (η²_p)
|
|
126
|
+
|
|
127
|
+
**What it measures**: Proportion of variance explained by factor, excluding other factors
|
|
128
|
+
|
|
129
|
+
**Formula**: η²_p = SS_effect / (SS_effect + SS_error)
|
|
130
|
+
|
|
131
|
+
**Interpretation**: Same benchmarks as η²
|
|
132
|
+
|
|
133
|
+
**When to use**: Multi-factor ANOVA (standard in factorial designs)
|
|
134
|
+
|
|
135
|
+
**Python calculation**:
|
|
136
|
+
```python
|
|
137
|
+
aov = pg.anova(dv='value', between=['factor1', 'factor2'], data=df)
|
|
138
|
+
# pingouin reports partial eta-squared by default
|
|
139
|
+
partial_eta_sq = aov['np2']
|
|
140
|
+
```
|
|
141
|
+
|
|
142
|
+
---
|
|
143
|
+
|
|
144
|
+
#### Omega-squared (ω²)
|
|
145
|
+
|
|
146
|
+
**What it measures**: Less biased estimate of population variance explained
|
|
147
|
+
|
|
148
|
+
**Why use it**: η² overestimates effect size; ω² provides better population estimate
|
|
149
|
+
|
|
150
|
+
**Formula**: ω² = (SS_effect - df_effect × MS_error) / (SS_total + MS_error)
|
|
151
|
+
|
|
152
|
+
**Interpretation**: Same benchmarks as η², but typically smaller values
|
|
153
|
+
|
|
154
|
+
**Python calculation**:
|
|
155
|
+
```python
|
|
156
|
+
def omega_squared(aov_table):
|
|
157
|
+
ss_effect = aov_table.loc[0, 'SS']
|
|
158
|
+
ss_total = aov_table['SS'].sum()
|
|
159
|
+
ms_error = aov_table.loc[aov_table.index[-1], 'MS'] # Residual MS
|
|
160
|
+
df_effect = aov_table.loc[0, 'DF']
|
|
161
|
+
|
|
162
|
+
omega_sq = (ss_effect - df_effect * ms_error) / (ss_total + ms_error)
|
|
163
|
+
return omega_sq
|
|
164
|
+
```
|
|
165
|
+
|
|
166
|
+
---
|
|
167
|
+
|
|
168
|
+
#### Cohen's f
|
|
169
|
+
|
|
170
|
+
**What it measures**: Effect size for ANOVA (analogous to Cohen's d)
|
|
171
|
+
|
|
172
|
+
**Formula**: f = √(η² / (1 - η²))
|
|
173
|
+
|
|
174
|
+
**Interpretation**:
|
|
175
|
+
- Small: f = 0.10
|
|
176
|
+
- Medium: f = 0.25
|
|
177
|
+
- Large: f = 0.40
|
|
178
|
+
|
|
179
|
+
**Python calculation**:
|
|
180
|
+
```python
|
|
181
|
+
eta_squared = 0.06 # From ANOVA
|
|
182
|
+
cohens_f = np.sqrt(eta_squared / (1 - eta_squared))
|
|
183
|
+
```
|
|
184
|
+
|
|
185
|
+
**Use in power analysis**: Required for ANOVA power calculations
|
|
186
|
+
|
|
187
|
+
---
|
|
188
|
+
|
|
189
|
+
### Correlation
|
|
190
|
+
|
|
191
|
+
#### Pearson's r / Spearman's ρ
|
|
192
|
+
|
|
193
|
+
**Interpretation**:
|
|
194
|
+
- Small: |r| = 0.10
|
|
195
|
+
- Medium: |r| = 0.30
|
|
196
|
+
- Large: |r| = 0.50
|
|
197
|
+
|
|
198
|
+
**Important notes**:
|
|
199
|
+
- r² = coefficient of determination (proportion of variance explained)
|
|
200
|
+
- r = 0.30 means 9% shared variance (0.30² = 0.09)
|
|
201
|
+
- Consider direction (positive/negative) and context
|
|
202
|
+
|
|
203
|
+
**Python calculation**:
|
|
204
|
+
```python
|
|
205
|
+
import pingouin as pg
|
|
206
|
+
|
|
207
|
+
# Pearson correlation with CI
|
|
208
|
+
result = pg.corr(x, y, method='pearson')
|
|
209
|
+
r = result['r'].values[0]
|
|
210
|
+
ci = result['CI95'].values[0] # Pingouin 0.5+: was CI95%
|
|
211
|
+
|
|
212
|
+
# Spearman correlation
|
|
213
|
+
result = pg.corr(x, y, method='spearman')
|
|
214
|
+
rho = result['r'].values[0]
|
|
215
|
+
```
|
|
216
|
+
|
|
217
|
+
---
|
|
218
|
+
|
|
219
|
+
### Regression
|
|
220
|
+
|
|
221
|
+
#### R² (Coefficient of Determination)
|
|
222
|
+
|
|
223
|
+
**What it measures**: Proportion of variance in Y explained by model
|
|
224
|
+
|
|
225
|
+
**Interpretation**:
|
|
226
|
+
- Small: R² = 0.02
|
|
227
|
+
- Medium: R² = 0.13
|
|
228
|
+
- Large: R² = 0.26
|
|
229
|
+
|
|
230
|
+
**Context-dependent**:
|
|
231
|
+
- Physical sciences: R² > 0.90 expected
|
|
232
|
+
- Social sciences: R² > 0.30 considered good
|
|
233
|
+
- Behavior prediction: R² > 0.10 may be meaningful
|
|
234
|
+
|
|
235
|
+
**Python calculation**:
|
|
236
|
+
```python
|
|
237
|
+
from sklearn.metrics import r2_score
|
|
238
|
+
from statsmodels.api import OLS
|
|
239
|
+
|
|
240
|
+
# Using statsmodels
|
|
241
|
+
model = OLS(y, X).fit()
|
|
242
|
+
r_squared = model.rsquared
|
|
243
|
+
adjusted_r_squared = model.rsquared_adj
|
|
244
|
+
|
|
245
|
+
# Manual
|
|
246
|
+
r_squared = 1 - (SS_residual / SS_total)
|
|
247
|
+
```
|
|
248
|
+
|
|
249
|
+
---
|
|
250
|
+
|
|
251
|
+
#### Adjusted R²
|
|
252
|
+
|
|
253
|
+
**Why use it**: R² artificially increases when adding predictors; adjusted R² penalizes model complexity
|
|
254
|
+
|
|
255
|
+
**Formula**: R²_adj = 1 - (1 - R²) × (n - 1) / (n - k - 1)
|
|
256
|
+
|
|
257
|
+
**When to use**: Always report alongside R² for multiple regression
|
|
258
|
+
|
|
259
|
+
---
|
|
260
|
+
|
|
261
|
+
#### Standardized Regression Coefficients (β)
|
|
262
|
+
|
|
263
|
+
**What it measures**: Effect of one-SD change in predictor on outcome (in SD units)
|
|
264
|
+
|
|
265
|
+
**Interpretation**: Similar to Cohen's d
|
|
266
|
+
- Small: |β| = 0.10
|
|
267
|
+
- Medium: |β| = 0.30
|
|
268
|
+
- Large: |β| = 0.50
|
|
269
|
+
|
|
270
|
+
**Python calculation**:
|
|
271
|
+
```python
|
|
272
|
+
from scipy import stats
|
|
273
|
+
|
|
274
|
+
# Standardize variables first
|
|
275
|
+
X_std = (X - X.mean()) / X.std()
|
|
276
|
+
y_std = (y - y.mean()) / y.std()
|
|
277
|
+
|
|
278
|
+
model = OLS(y_std, X_std).fit()
|
|
279
|
+
beta = model.params
|
|
280
|
+
```
|
|
281
|
+
|
|
282
|
+
---
|
|
283
|
+
|
|
284
|
+
#### f² (Cohen's f-squared for Regression)
|
|
285
|
+
|
|
286
|
+
**What it measures**: Effect size for individual predictors or model comparison
|
|
287
|
+
|
|
288
|
+
**Formula**: f² = R²_AB - R²_A / (1 - R²_AB)
|
|
289
|
+
|
|
290
|
+
Where:
|
|
291
|
+
- R²_AB = R² for full model with predictor
|
|
292
|
+
- R²_A = R² for reduced model without predictor
|
|
293
|
+
|
|
294
|
+
**Interpretation**:
|
|
295
|
+
- Small: f² = 0.02
|
|
296
|
+
- Medium: f² = 0.15
|
|
297
|
+
- Large: f² = 0.35
|
|
298
|
+
|
|
299
|
+
**Python calculation**:
|
|
300
|
+
```python
|
|
301
|
+
# Compare two nested models
|
|
302
|
+
model_full = OLS(y, X_full).fit()
|
|
303
|
+
model_reduced = OLS(y, X_reduced).fit()
|
|
304
|
+
|
|
305
|
+
r2_full = model_full.rsquared
|
|
306
|
+
r2_reduced = model_reduced.rsquared
|
|
307
|
+
|
|
308
|
+
f_squared = (r2_full - r2_reduced) / (1 - r2_full)
|
|
309
|
+
```
|
|
310
|
+
|
|
311
|
+
---
|
|
312
|
+
|
|
313
|
+
### Categorical Data Analysis
|
|
314
|
+
|
|
315
|
+
#### Cramér's V
|
|
316
|
+
|
|
317
|
+
**What it measures**: Association strength for χ² test (works for any table size)
|
|
318
|
+
|
|
319
|
+
**Formula**: V = √(χ² / (n × (k - 1)))
|
|
320
|
+
|
|
321
|
+
Where k = min(rows, columns)
|
|
322
|
+
|
|
323
|
+
**Interpretation** (for k > 2):
|
|
324
|
+
- Small: V = 0.07
|
|
325
|
+
- Medium: V = 0.21
|
|
326
|
+
- Large: V = 0.35
|
|
327
|
+
|
|
328
|
+
**For 2×2 tables**: Use phi coefficient (φ)
|
|
329
|
+
|
|
330
|
+
**Python calculation**:
|
|
331
|
+
```python
|
|
332
|
+
from scipy.stats.contingency import association
|
|
333
|
+
|
|
334
|
+
# Cramér's V
|
|
335
|
+
cramers_v = association(contingency_table, method='cramer')
|
|
336
|
+
|
|
337
|
+
# Phi coefficient (for 2x2)
|
|
338
|
+
phi = association(contingency_table, method='pearson')
|
|
339
|
+
```
|
|
340
|
+
|
|
341
|
+
---
|
|
342
|
+
|
|
343
|
+
#### Odds Ratio (OR) and Risk Ratio (RR)
|
|
344
|
+
|
|
345
|
+
**For 2×2 contingency tables**:
|
|
346
|
+
|
|
347
|
+
| | Outcome + | Outcome - |
|
|
348
|
+
|-----------|-----------|-----------|
|
|
349
|
+
| Exposed | a | b |
|
|
350
|
+
| Unexposed | c | d |
|
|
351
|
+
|
|
352
|
+
**Odds Ratio**: OR = (a/b) / (c/d) = ad / bc
|
|
353
|
+
|
|
354
|
+
**Interpretation**:
|
|
355
|
+
- OR = 1: No association
|
|
356
|
+
- OR > 1: Positive association (increased odds)
|
|
357
|
+
- OR < 1: Negative association (decreased odds)
|
|
358
|
+
- OR = 2: Twice the odds
|
|
359
|
+
- OR = 0.5: Half the odds
|
|
360
|
+
|
|
361
|
+
**Risk Ratio**: RR = (a/(a+b)) / (c/(c+d))
|
|
362
|
+
|
|
363
|
+
**When to use**:
|
|
364
|
+
- Cohort studies: Use RR (more interpretable)
|
|
365
|
+
- Case-control studies: Use OR (RR not available)
|
|
366
|
+
- Logistic regression: OR is natural output
|
|
367
|
+
|
|
368
|
+
**Python calculation**:
|
|
369
|
+
```python
|
|
370
|
+
import statsmodels.api as sm
|
|
371
|
+
|
|
372
|
+
# From contingency table
|
|
373
|
+
odds_ratio = (a * d) / (b * c)
|
|
374
|
+
|
|
375
|
+
# Confidence interval
|
|
376
|
+
table = np.array([[a, b], [c, d]])
|
|
377
|
+
oddsratio, pvalue = stats.fisher_exact(table)
|
|
378
|
+
|
|
379
|
+
# From logistic regression
|
|
380
|
+
model = sm.Logit(y, X).fit()
|
|
381
|
+
odds_ratios = np.exp(model.params) # Exponentiate coefficients
|
|
382
|
+
ci = np.exp(model.conf_int()) # Exponentiate CIs
|
|
383
|
+
```
|
|
384
|
+
|
|
385
|
+
---
|
|
386
|
+
|
|
387
|
+
### Bayesian Effect Sizes
|
|
388
|
+
|
|
389
|
+
#### Bayes Factor (BF)
|
|
390
|
+
|
|
391
|
+
**What it measures**: Ratio of evidence for alternative vs. null hypothesis
|
|
392
|
+
|
|
393
|
+
**Interpretation**:
|
|
394
|
+
- BF₁₀ = 1: Equal evidence for H₁ and H₀
|
|
395
|
+
- BF₁₀ = 3: H₁ is 3× more likely than H₀ (moderate evidence)
|
|
396
|
+
- BF₁₀ = 10: H₁ is 10× more likely than H₀ (strong evidence)
|
|
397
|
+
- BF₁₀ = 100: H₁ is 100× more likely than H₀ (decisive evidence)
|
|
398
|
+
- BF₁₀ = 0.33: H₀ is 3× more likely than H₁
|
|
399
|
+
- BF₁₀ = 0.10: H₀ is 10× more likely than H₁
|
|
400
|
+
|
|
401
|
+
**Classification** (Jeffreys, 1961):
|
|
402
|
+
- 1-3: Anecdotal evidence
|
|
403
|
+
- 3-10: Moderate evidence
|
|
404
|
+
- 10-30: Strong evidence
|
|
405
|
+
- 30-100: Very strong evidence
|
|
406
|
+
- >100: Decisive evidence
|
|
407
|
+
|
|
408
|
+
**Python calculation**:
|
|
409
|
+
```python
|
|
410
|
+
import pingouin as pg
|
|
411
|
+
|
|
412
|
+
# Pingouin 0.5+: two-sided BF10 on independent t-tests; use BayesFactor/JASP/PyMC for full inference
|
|
413
|
+
result = pg.ttest(group1, group2, correction=False)
|
|
414
|
+
bf10 = result['BF10'].values[0]
|
|
415
|
+
```
|
|
416
|
+
|
|
417
|
+
---
|
|
418
|
+
|
|
419
|
+
## Power Analysis
|
|
420
|
+
|
|
421
|
+
### Concepts
|
|
422
|
+
|
|
423
|
+
**Statistical power**: Probability of detecting an effect if it exists (1 - β)
|
|
424
|
+
|
|
425
|
+
**Conventional standards**:
|
|
426
|
+
- Power = 0.80 (80% chance of detecting effect)
|
|
427
|
+
- α = 0.05 (5% Type I error rate)
|
|
428
|
+
|
|
429
|
+
**Four interconnected parameters** (given 3, can solve for 4th):
|
|
430
|
+
1. Sample size (n)
|
|
431
|
+
2. Effect size (d, f, etc.)
|
|
432
|
+
3. Significance level (α)
|
|
433
|
+
4. Power (1 - β)
|
|
434
|
+
|
|
435
|
+
---
|
|
436
|
+
|
|
437
|
+
### A Priori Power Analysis (Planning)
|
|
438
|
+
|
|
439
|
+
**Purpose**: Determine required sample size before study
|
|
440
|
+
|
|
441
|
+
**Steps**:
|
|
442
|
+
1. Specify expected effect size (from literature, pilot data, or minimum meaningful effect)
|
|
443
|
+
2. Set α level (typically 0.05)
|
|
444
|
+
3. Set desired power (typically 0.80)
|
|
445
|
+
4. Calculate required n
|
|
446
|
+
|
|
447
|
+
**Python implementation**:
|
|
448
|
+
```python
|
|
449
|
+
from statsmodels.stats.power import (
|
|
450
|
+
tt_ind_solve_power,
|
|
451
|
+
zt_ind_solve_power,
|
|
452
|
+
FTestAnovaPower,
|
|
453
|
+
NormalIndPower
|
|
454
|
+
)
|
|
455
|
+
|
|
456
|
+
# T-test power analysis
|
|
457
|
+
n_required = tt_ind_solve_power(
|
|
458
|
+
effect_size=0.5, # Cohen's d
|
|
459
|
+
alpha=0.05,
|
|
460
|
+
power=0.80,
|
|
461
|
+
ratio=1.0, # Equal group sizes
|
|
462
|
+
alternative='two-sided'
|
|
463
|
+
)
|
|
464
|
+
|
|
465
|
+
# ANOVA power analysis
|
|
466
|
+
anova_power = FTestAnovaPower()
|
|
467
|
+
n_per_group = anova_power.solve_power(
|
|
468
|
+
effect_size=0.25, # Cohen's f
|
|
469
|
+
ngroups=3,
|
|
470
|
+
alpha=0.05,
|
|
471
|
+
power=0.80
|
|
472
|
+
)
|
|
473
|
+
|
|
474
|
+
# Correlation power analysis
|
|
475
|
+
from pingouin import power_corr
|
|
476
|
+
n_required = power_corr(r=0.30, power=0.80, alpha=0.05)
|
|
477
|
+
```
|
|
478
|
+
|
|
479
|
+
---
|
|
480
|
+
|
|
481
|
+
### Post Hoc Power Analysis (After Study)
|
|
482
|
+
|
|
483
|
+
**⚠️ CAUTION**: Post hoc power is controversial and often not recommended
|
|
484
|
+
|
|
485
|
+
**Why it's problematic**:
|
|
486
|
+
- Observed power is a direct function of p-value
|
|
487
|
+
- If p > 0.05, power is always low
|
|
488
|
+
- Provides no additional information beyond p-value
|
|
489
|
+
- Can be misleading
|
|
490
|
+
|
|
491
|
+
**When it might be acceptable**:
|
|
492
|
+
- Study planning for future research
|
|
493
|
+
- Using effect size from multiple studies (not just your own)
|
|
494
|
+
- Explicit goal is sample size for replication
|
|
495
|
+
|
|
496
|
+
**Better alternatives**:
|
|
497
|
+
- Report confidence intervals for effect sizes
|
|
498
|
+
- Conduct sensitivity analysis
|
|
499
|
+
- Report minimum detectable effect size
|
|
500
|
+
|
|
501
|
+
---
|
|
502
|
+
|
|
503
|
+
### Sensitivity Analysis
|
|
504
|
+
|
|
505
|
+
**Purpose**: Determine minimum detectable effect size given study parameters
|
|
506
|
+
|
|
507
|
+
**When to use**: After study is complete, to understand study's capability
|
|
508
|
+
|
|
509
|
+
**Python implementation**:
|
|
510
|
+
```python
|
|
511
|
+
# What effect size could we detect with n=50 per group?
|
|
512
|
+
detectable_effect = tt_ind_solve_power(
|
|
513
|
+
effect_size=None, # Solve for this
|
|
514
|
+
nobs1=50,
|
|
515
|
+
alpha=0.05,
|
|
516
|
+
power=0.80,
|
|
517
|
+
ratio=1.0,
|
|
518
|
+
alternative='two-sided'
|
|
519
|
+
)
|
|
520
|
+
|
|
521
|
+
print(f"With n=50 per group, we could detect d ≥ {detectable_effect:.2f}")
|
|
522
|
+
```
|
|
523
|
+
|
|
524
|
+
---
|
|
525
|
+
|
|
526
|
+
## Reporting Effect Sizes
|
|
527
|
+
|
|
528
|
+
### APA Style Guidelines
|
|
529
|
+
|
|
530
|
+
**T-test example**:
|
|
531
|
+
> "Group A (M = 75.2, SD = 8.5) scored significantly higher than Group B (M = 68.3, SD = 9.2), t(98) = 3.82, p < .001, d = 0.77, 95% CI [0.36, 1.18]."
|
|
532
|
+
|
|
533
|
+
**ANOVA example**:
|
|
534
|
+
> "There was a significant main effect of treatment condition on test scores, F(2, 87) = 8.45, p < .001, η²p = .16. Post hoc comparisons using Tukey's HSD revealed..."
|
|
535
|
+
|
|
536
|
+
**Correlation example**:
|
|
537
|
+
> "There was a moderate positive correlation between study time and exam scores, r(148) = .42, p < .001, 95% CI [.27, .55]."
|
|
538
|
+
|
|
539
|
+
**Regression example**:
|
|
540
|
+
> "The regression model significantly predicted exam scores, F(3, 146) = 45.2, p < .001, R² = .48. Study hours (β = .52, p < .001) and prior GPA (β = .31, p < .001) were significant predictors."
|
|
541
|
+
|
|
542
|
+
**Bayesian example**:
|
|
543
|
+
> "A Bayesian independent samples t-test provided strong evidence for a difference between groups, BF₁₀ = 23.5, indicating the data are 23.5 times more likely under H₁ than H₀."
|
|
544
|
+
|
|
545
|
+
---
|
|
546
|
+
|
|
547
|
+
## Effect Size Pitfalls
|
|
548
|
+
|
|
549
|
+
1. **Don't only rely on benchmarks**: Context matters; small effects can be meaningful
|
|
550
|
+
2. **Report confidence intervals**: CIs show precision of effect size estimate
|
|
551
|
+
3. **Distinguish statistical vs. practical significance**: Large n can make trivial effects "significant"
|
|
552
|
+
4. **Consider cost-benefit**: Even small effects may be valuable if intervention is low-cost
|
|
553
|
+
5. **Multiple outcomes**: Effect sizes vary across outcomes; report all
|
|
554
|
+
6. **Don't cherry-pick**: Report effects for all planned analyses
|
|
555
|
+
7. **Publication bias**: Published effects are often overestimated
|
|
556
|
+
|
|
557
|
+
---
|
|
558
|
+
|
|
559
|
+
## Quick Reference Table
|
|
560
|
+
|
|
561
|
+
| Analysis | Effect Size | Small | Medium | Large |
|
|
562
|
+
|----------|-------------|-------|--------|-------|
|
|
563
|
+
| T-test | Cohen's d | 0.20 | 0.50 | 0.80 |
|
|
564
|
+
| ANOVA | η², ω² | 0.01 | 0.06 | 0.14 |
|
|
565
|
+
| ANOVA | Cohen's f | 0.10 | 0.25 | 0.40 |
|
|
566
|
+
| Correlation | r, ρ | 0.10 | 0.30 | 0.50 |
|
|
567
|
+
| Regression | R² | 0.02 | 0.13 | 0.26 |
|
|
568
|
+
| Regression | f² | 0.02 | 0.15 | 0.35 |
|
|
569
|
+
| Chi-square | Cramér's V | 0.07 | 0.21 | 0.35 |
|
|
570
|
+
| Chi-square (2×2) | φ | 0.10 | 0.30 | 0.50 |
|
|
571
|
+
|
|
572
|
+
---
|
|
573
|
+
|
|
574
|
+
## Resources
|
|
575
|
+
|
|
576
|
+
- Cohen, J. (1988). *Statistical Power Analysis for the Behavioral Sciences* (2nd ed.)
|
|
577
|
+
- Lakens, D. (2013). Calculating and reporting effect sizes
|
|
578
|
+
- Ellis, P. D. (2010). *The Essential Guide to Effect Sizes*
|