clearai-dsh 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (154) hide show
  1. package/CHANGELOG.md +26 -0
  2. package/LICENSE +201 -0
  3. package/README.md +138 -0
  4. package/README.zh-CN.md +138 -0
  5. package/bin/clearai.mjs +224 -0
  6. package/brand/README.md +41 -0
  7. package/brand/logo-512-dark.png +0 -0
  8. package/brand/logo-512.png +0 -0
  9. package/brand/logo-lockup-dark.png +0 -0
  10. package/brand/logo-lockup.png +0 -0
  11. package/brand/logo-lockup.svg +12 -0
  12. package/brand/logo-wordmark.svg +6 -0
  13. package/brand/logo.svg +19 -0
  14. package/cordis.patch.yml +39 -0
  15. package/lib/client.js +3071 -0
  16. package/lib/fold.js +1576 -0
  17. package/lib/host.js +605 -0
  18. package/package.json +65 -0
  19. package/presets/clearai/agent.cordis.yml +226 -0
  20. package/presets/clearai/plugins/brain.js +547 -0
  21. package/presets/clearai/plugins/clearai-kernel.js +5485 -0
  22. package/presets/clearai/plugins/ontology.js +306 -0
  23. package/presets/clearai/plugins/prompts.js +312 -0
  24. package/presets/clearai/preset.yml +5 -0
  25. package/presets/clearai/skills/clearai-loop/SKILL.md +89 -0
  26. package/presets/clearai/template/knowledge/README.md +25 -0
  27. package/presets/clearai/template/memory/README.md +34 -0
  28. package/presets/clearai/template/project.md +49 -0
  29. package/presets/clearai/template/skills/README.md +37 -0
  30. package/presets/clearai/template/skills/chart-diagram-qa/SKILL.md +43 -0
  31. package/presets/clearai/template/skills/citation-management/SKILL.md +73 -0
  32. package/presets/clearai/template/skills/citation-management/references/bibtex_formatting.md +908 -0
  33. package/presets/clearai/template/skills/citation-management/references/citation_validation.md +794 -0
  34. package/presets/clearai/template/skills/citation-management/references/google_scholar_search.md +725 -0
  35. package/presets/clearai/template/skills/citation-management/references/metadata_extraction.md +870 -0
  36. package/presets/clearai/template/skills/citation-management/references/pubmed_search.md +839 -0
  37. package/presets/clearai/template/skills/citation-management/scripts/doi_to_bibtex.py +204 -0
  38. package/presets/clearai/template/skills/citation-management/scripts/extract_metadata.py +569 -0
  39. package/presets/clearai/template/skills/citation-management/scripts/format_bibtex.py +349 -0
  40. package/presets/clearai/template/skills/citation-management/scripts/generate_schematic.py +139 -0
  41. package/presets/clearai/template/skills/citation-management/scripts/generate_schematic_ai.py +817 -0
  42. package/presets/clearai/template/skills/citation-management/scripts/search_google_scholar.py +282 -0
  43. package/presets/clearai/template/skills/citation-management/scripts/search_pubmed.py +398 -0
  44. package/presets/clearai/template/skills/citation-management/scripts/validate_citations.py +497 -0
  45. package/presets/clearai/template/skills/data-analysis/SKILL.md +92 -0
  46. package/presets/clearai/template/skills/data-analysis/checklists/readiness_check.md +23 -0
  47. package/presets/clearai/template/skills/data-analysis/templates/analysis_report.md.tpl +63 -0
  48. package/presets/clearai/template/skills/data-analysis/templates/cleaning_rules_draft.yaml.tpl +32 -0
  49. package/presets/clearai/template/skills/data-analysis/templates/data_dictionary.md.tpl +12 -0
  50. package/presets/clearai/template/skills/data-analysis/templates/domain_knowledge_template.md.tpl +316 -0
  51. package/presets/clearai/template/skills/data-analysis/templates/feature_candidates.json.tpl +20 -0
  52. package/presets/clearai/template/skills/data-analysis/templates/quality_scorecard.md.tpl +30 -0
  53. package/presets/clearai/template/skills/data-analysis/workflows/01-data-profiling.md +42 -0
  54. package/presets/clearai/template/skills/data-analysis/workflows/02-quality-audit.md +36 -0
  55. package/presets/clearai/template/skills/data-analysis/workflows/03-physical-correlation.md +25 -0
  56. package/presets/clearai/template/skills/data-analysis/workflows/04-unstructured-mining.md +26 -0
  57. package/presets/clearai/template/skills/data-qa-analysis/SKILL.md +102 -0
  58. package/presets/clearai/template/skills/data-qa-analysis/checklists/readiness_check.md +62 -0
  59. package/presets/clearai/template/skills/data-qa-analysis/templates/best_in_class_report.md.tpl +56 -0
  60. package/presets/clearai/template/skills/data-qa-analysis/templates/cleaning_rules_draft.yaml.tpl +56 -0
  61. package/presets/clearai/template/skills/data-qa-analysis/templates/data_dictionary.md.tpl +13 -0
  62. package/presets/clearai/template/skills/data-qa-analysis/templates/data_source_inventory_and_lineage.md.tpl +146 -0
  63. package/presets/clearai/template/skills/data-qa-analysis/templates/data_status_report.md.tpl +60 -0
  64. package/presets/clearai/template/skills/data-qa-analysis/templates/steady_state_rules.yaml.tpl +41 -0
  65. package/presets/clearai/template/skills/data-qa-analysis/templates/subsystem_registry.md.tpl +101 -0
  66. package/presets/clearai/template/skills/data-qa-analysis/templates/unified_execution_plan.md.tpl +100 -0
  67. package/presets/clearai/template/skills/data-qa-analysis/workflows/01-data-source-inventory-and-lineage.md +194 -0
  68. package/presets/clearai/template/skills/data-qa-analysis/workflows/02-data-alignment-and-tag-semantics.md +122 -0
  69. package/presets/clearai/template/skills/data-qa-analysis/workflows/03-steady-state-identification.md +126 -0
  70. package/presets/clearai/template/skills/data-qa-analysis/workflows/04-consumption-analysis.md +152 -0
  71. package/presets/clearai/template/skills/data-qa-analysis/workflows/05-best-in-class-and-optimization-space.md +78 -0
  72. package/presets/clearai/template/skills/domain-presearch/SKILL.md +131 -0
  73. package/presets/clearai/template/skills/domain-presearch/checklists/domain_checklist.md +24 -0
  74. package/presets/clearai/template/skills/domain-presearch/references/figure_code.md +78 -0
  75. package/presets/clearai/template/skills/domain-presearch/references/strategic_frameworks.md +38 -0
  76. package/presets/clearai/template/skills/exploration-loop/SKILL.md +81 -0
  77. package/presets/clearai/template/skills/exploratory-data-analysis/SKILL.md +77 -0
  78. package/presets/clearai/template/skills/exploratory-data-analysis/references/bioinformatics_genomics_formats.md +664 -0
  79. package/presets/clearai/template/skills/exploratory-data-analysis/references/chemistry_molecular_formats.md +664 -0
  80. package/presets/clearai/template/skills/exploratory-data-analysis/references/general_scientific_formats.md +518 -0
  81. package/presets/clearai/template/skills/exploratory-data-analysis/references/microscopy_imaging_formats.md +620 -0
  82. package/presets/clearai/template/skills/exploratory-data-analysis/references/proteomics_metabolomics_formats.md +517 -0
  83. package/presets/clearai/template/skills/exploratory-data-analysis/references/spectroscopy_analytical_formats.md +633 -0
  84. package/presets/clearai/template/skills/exploratory-data-analysis/scripts/eda_analyzer.py +547 -0
  85. package/presets/clearai/template/skills/hypothesis-generation/SKILL.md +73 -0
  86. package/presets/clearai/template/skills/hypothesis-generation/references/experimental_design_patterns.md +329 -0
  87. package/presets/clearai/template/skills/hypothesis-generation/references/hypothesis_quality_criteria.md +198 -0
  88. package/presets/clearai/template/skills/hypothesis-generation/references/literature_search_strategies.md +622 -0
  89. package/presets/clearai/template/skills/hypothesis-generation/scripts/generate_schematic.py +139 -0
  90. package/presets/clearai/template/skills/hypothesis-generation/scripts/generate_schematic_ai.py +817 -0
  91. package/presets/clearai/template/skills/literature-review/SKILL.md +72 -0
  92. package/presets/clearai/template/skills/literature-review/references/citation_styles.md +166 -0
  93. package/presets/clearai/template/skills/literature-review/references/database_strategies.md +455 -0
  94. package/presets/clearai/template/skills/literature-review/scripts/generate_pdf.py +176 -0
  95. package/presets/clearai/template/skills/literature-review/scripts/generate_schematic.py +139 -0
  96. package/presets/clearai/template/skills/literature-review/scripts/generate_schematic_ai.py +817 -0
  97. package/presets/clearai/template/skills/literature-review/scripts/search_databases.py +303 -0
  98. package/presets/clearai/template/skills/literature-review/scripts/verify_citations.py +221 -0
  99. package/presets/clearai/template/skills/paper-lookup/SKILL.md +59 -0
  100. package/presets/clearai/template/skills/paper-lookup/references/arxiv.md +161 -0
  101. package/presets/clearai/template/skills/paper-lookup/references/biorxiv.md +118 -0
  102. package/presets/clearai/template/skills/paper-lookup/references/core.md +150 -0
  103. package/presets/clearai/template/skills/paper-lookup/references/crossref.md +181 -0
  104. package/presets/clearai/template/skills/paper-lookup/references/medrxiv.md +104 -0
  105. package/presets/clearai/template/skills/paper-lookup/references/openalex.md +174 -0
  106. package/presets/clearai/template/skills/paper-lookup/references/pmc.md +152 -0
  107. package/presets/clearai/template/skills/paper-lookup/references/pubmed.md +124 -0
  108. package/presets/clearai/template/skills/paper-lookup/references/semantic-scholar.md +203 -0
  109. package/presets/clearai/template/skills/paper-lookup/references/unpaywall.md +127 -0
  110. package/presets/clearai/template/skills/process-presearch/SKILL.md +196 -0
  111. package/presets/clearai/template/skills/process-presearch/checklists/process_checklist.md +18 -0
  112. package/presets/clearai/template/skills/process-presearch/references/figure_code.md +107 -0
  113. package/presets/clearai/template/skills/process-presearch/references/source_attribution_example.md +22 -0
  114. package/presets/clearai/template/skills/process-understanding-extraction/SKILL.md +69 -0
  115. package/presets/clearai/template/skills/process-understanding-extraction/checklists/readiness_check.md +34 -0
  116. package/presets/clearai/template/skills/process-understanding-extraction/templates/docx_raw_dump_extractor.py.tpl +132 -0
  117. package/presets/clearai/template/skills/process-understanding-extraction/templates/entity_map_unit_topology.json.tpl +86 -0
  118. package/presets/clearai/template/skills/process-understanding-extraction/templates/process_brief.md.tpl +89 -0
  119. package/presets/clearai/template/skills/process-understanding-extraction/templates/process_brief_builder_from_raw_dump.py.tpl +203 -0
  120. package/presets/clearai/template/skills/process-understanding-extraction/templates/process_flow_mermaid.md.tpl +41 -0
  121. package/presets/clearai/template/skills/process-understanding-extraction/templates/unified_execution_plan.md.tpl +53 -0
  122. package/presets/clearai/template/skills/process-understanding-extraction/workflows/01-process-doc-discovery.md +173 -0
  123. package/presets/clearai/template/skills/process-understanding-extraction/workflows/02-process-understanding-and-diagramming.md +106 -0
  124. package/presets/clearai/template/skills/scientific-brainstorming/SKILL.md +64 -0
  125. package/presets/clearai/template/skills/scientific-brainstorming/references/brainstorming_methods.md +326 -0
  126. package/presets/clearai/template/skills/scientific-critical-thinking/SKILL.md +72 -0
  127. package/presets/clearai/template/skills/scientific-critical-thinking/references/common_biases.md +364 -0
  128. package/presets/clearai/template/skills/scientific-critical-thinking/references/evidence_hierarchy.md +485 -0
  129. package/presets/clearai/template/skills/scientific-critical-thinking/references/experimental_design.md +496 -0
  130. package/presets/clearai/template/skills/scientific-critical-thinking/references/logical_fallacies.md +478 -0
  131. package/presets/clearai/template/skills/scientific-critical-thinking/references/scientific_method.md +169 -0
  132. package/presets/clearai/template/skills/scientific-critical-thinking/references/statistical_pitfalls.md +506 -0
  133. package/presets/clearai/template/skills/skill-creator/SKILL.md +109 -0
  134. package/presets/clearai/template/skills/skill-creator/references/authoring-guide.md +89 -0
  135. package/presets/clearai/template/skills/statistical-analysis/SKILL.md +79 -0
  136. package/presets/clearai/template/skills/statistical-analysis/references/assumptions_and_diagnostics.md +369 -0
  137. package/presets/clearai/template/skills/statistical-analysis/references/bayesian_statistics.md +653 -0
  138. package/presets/clearai/template/skills/statistical-analysis/references/effect_sizes_and_power.md +578 -0
  139. package/presets/clearai/template/skills/statistical-analysis/references/reporting_standards.md +469 -0
  140. package/presets/clearai/template/skills/statistical-analysis/references/test_selection_guide.md +129 -0
  141. package/presets/clearai/template/skills/statistical-analysis/scripts/assumption_checks.py +538 -0
  142. package/presets/clearai/template/skills/web-artifact/SKILL.md +165 -0
  143. package/presets/clearai/template/skills/web-artifact/assets/renderer/renderer.css +229 -0
  144. package/presets/clearai/template/skills/web-artifact/assets/renderer/renderer.js +373 -0
  145. package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/LICENSE +263 -0
  146. package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/UPSTREAM.md +26 -0
  147. package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/elk.bundled.js +6605 -0
  148. package/presets/clearai/template/skills/web-artifact/references/when-drawing-a-topology.md +150 -0
  149. package/presets/clearai/template/skills/web-artifact/references/when-the-page-must-work-offline.md +62 -0
  150. package/presets/clearai/template/skills/web-artifact/scripts/check_artifact.py +167 -0
  151. package/presets/clearai/template/skills/web-artifact/scripts/render_topology.js +272 -0
  152. package/presets/clearai/template/skills/what-if-oracle/LICENSE.txt +5 -0
  153. package/presets/clearai/template/skills/what-if-oracle/SKILL.md +72 -0
  154. package/presets/clearai/template/skills/what-if-oracle/references/scenario-templates.md +154 -0
@@ -0,0 +1,79 @@
1
+ ---
2
+ name: statistical-analysis
3
+ description: |
4
+ 【假设验证·统计检验】选择合适检验方法、检查前提、执行分析并产出可审计结论。适用:验证两组条件差异、相关性、回归关系、组间对比。不适用:首触未知数据文件(用 exploratory-data-analysis);深度关联分析与诊断(用 data-analysis)。
5
+ license: MIT license
6
+ metadata:
7
+ version: 1.0-clearai
8
+ skill-author: K-Dense Inc. (adapted for ClearAI)
9
+ tier: system
10
+ origin: template
11
+ created_at: '2026-06-12T02:47:05.410760+00:00'
12
+ ---
13
+
14
+ # 统计分析 Skill(ClearAI 版)
15
+
16
+ ## 使用边界
17
+
18
+ - **适用**:已有明确假设与干净表格数据,需统计检验支撑结论。
19
+ - **不适用**:数据尚未探查 → `exploratory-data-analysis`;需深度关联与根因分析 → `data-analysis`。
20
+
21
+ ## ClearAI 工具与路径映射
22
+
23
+ - 数据 → `read` / `bash`(pandas 读 `input/` 或 `lab/`)
24
+ - 分析脚本 → `lab/scripts/stats_*.py`(可用已复制 `scripts/` 模板)
25
+ - 结果表 → `lab/extracted/stats_results.md` 或 `.csv`
26
+ - 图表 → `lab/diagrams/`
27
+ - 交付摘要 → `products/reports/`(若用户要求正式报告)
28
+ - 经验回写 → `clear/memory/statistical_analysis_lessons.md`
29
+
30
+ ## 环境契约
31
+
32
+ ```bash
33
+ python -c "import pandas, scipy, statsmodels, pingouin; print('statistics capability ready')"
34
+ ```
35
+
36
+ 上述统计栈属于 ClearAI 基础安装能力。若 import 验证失败,记录缺失包与当前 Python
37
+ 路径并停止,提示用户修复部署;**禁止**在 Agent 运行中执行 `pip install` 或 `uv`。
38
+
39
+ ## 工作流
40
+
41
+ ### 1. 检验选型
42
+
43
+ | 问题类型 | 典型检验 |
44
+ |----------|----------|
45
+ | 两组均值 | t 检验 / Mann-Whitney |
46
+ | 多组均值 | ANOVA / Kruskal-Wallis |
47
+ | 分类关联 | 卡方 |
48
+ | 两变量关系 | Pearson/Spearman 相关、OLS 回归 |
49
+ | 时序对比 | 配对检验、变化点(简述) |
50
+
51
+ 详见 `references/test_selection.md`、`references/assumptions_and_diagnostics.md`。
52
+
53
+ ### 2. 前提检查
54
+
55
+ - 正态性、方差齐性、样本量、独立性
56
+ - 不满足时选用非参数或明确声明局限
57
+
58
+ ### 3. 执行与报告(工程口径,非 APA)
59
+
60
+ 每条结论须含:
61
+ - 检验名称、统计量、p 值/置信区间、效应量
62
+ - **物理证据**:基于哪份文件、哪几列、样本量 n=?
63
+ - 工程意义:差异是否在业务上显著?
64
+
65
+ ### 4. 证据闭环
66
+
67
+ - 无统计输出文件/图表路径 → 不得宣称「已验证」
68
+ - 与 `hypothesis-generation` 产出的 H1/H2 逐条对应
69
+
70
+ ## 报告模板(`lab/extracted/stats_results.md`)
71
+
72
+ ```markdown
73
+ ## 假设 H1: ...
74
+ - 方法:...
75
+ - 结果:stat=..., p=..., effect_size=...
76
+ - 证据:lab/diagrams/xxx.png, 输入=input/yyy.csv
77
+ - 结论:支持/拒绝/ inconclusive
78
+ - 局限:...
79
+ ```
@@ -0,0 +1,369 @@
1
+ # Statistical Assumptions and Diagnostic Procedures
2
+
3
+ This document provides comprehensive guidance on checking and validating statistical assumptions for various analyses.
4
+
5
+ ## General Principles
6
+
7
+ 1. **Always check assumptions before interpreting test results**
8
+ 2. **Use multiple diagnostic methods** (visual + formal tests)
9
+ 3. **Consider robustness**: Some tests are robust to violations under certain conditions
10
+ 4. **Document all assumption checks** in analysis reports
11
+ 5. **Report violations and remedial actions taken**
12
+
13
+ ## Common Assumptions Across Tests
14
+
15
+ ### 1. Independence of Observations
16
+
17
+ **What it means**: Each observation is independent; measurements on one subject do not influence measurements on another.
18
+
19
+ **How to check**:
20
+ - Review study design and data collection procedures
21
+ - For time series: Check autocorrelation (ACF/PACF plots, Durbin-Watson test)
22
+ - For clustered data: Consider intraclass correlation (ICC)
23
+
24
+ **What to do if violated**:
25
+ - Use mixed-effects models for clustered/hierarchical data
26
+ - Use time series methods for temporally dependent data
27
+ - Use generalized estimating equations (GEE) for correlated data
28
+
29
+ **Critical severity**: HIGH - violations can severely inflate Type I error
30
+
31
+ ---
32
+
33
+ ### 2. Normality
34
+
35
+ **What it means**: Data or residuals follow a normal (Gaussian) distribution.
36
+
37
+ **When required**:
38
+ - t-tests (for small samples; robust for n > 30 per group)
39
+ - ANOVA (for small samples; robust for n > 30 per group)
40
+ - Linear regression (for residuals)
41
+ - Some correlation tests (Pearson)
42
+
43
+ **How to check**:
44
+
45
+ **Visual methods** (primary):
46
+ - Q-Q (quantile-quantile) plot: Points should fall on diagonal line
47
+ - Histogram with normal curve overlay
48
+ - Kernel density plot
49
+
50
+ **Formal tests** (secondary):
51
+ - Shapiro-Wilk test (recommended for n < 50)
52
+ - Kolmogorov-Smirnov test
53
+ - Anderson-Darling test
54
+
55
+ **Python implementation**:
56
+ ```python
57
+ from scipy import stats
58
+ import matplotlib.pyplot as plt
59
+
60
+ # Shapiro-Wilk test
61
+ statistic, p_value = stats.shapiro(data)
62
+
63
+ # Q-Q plot
64
+ stats.probplot(data, dist="norm", plot=plt)
65
+ ```
66
+
67
+ **Interpretation guidance**:
68
+ - For n < 30: Both visual and formal tests important
69
+ - For 30 ≤ n < 100: Visual inspection primary, formal tests secondary
70
+ - For n ≥ 100: Formal tests overly sensitive; rely on visual inspection
71
+ - Look for severe skewness, outliers, or bimodality
72
+
73
+ **What to do if violated**:
74
+ - **Mild violations** (slight skewness): Proceed if n > 30 per group
75
+ - **Moderate violations**: Use non-parametric alternatives (Mann-Whitney, Kruskal-Wallis, Wilcoxon)
76
+ - **Severe violations**:
77
+ - Transform data (log, square root, Box-Cox)
78
+ - Use non-parametric methods
79
+ - Use robust regression methods
80
+ - Consider bootstrapping
81
+
82
+ **Critical severity**: MEDIUM - parametric tests are often robust to mild violations with adequate sample size
83
+
84
+ ---
85
+
86
+ ### 3. Homogeneity of Variance (Homoscedasticity)
87
+
88
+ **What it means**: Variances are equal across groups or across the range of predictors.
89
+
90
+ **When required**:
91
+ - Independent samples t-test
92
+ - ANOVA
93
+ - Linear regression (constant variance of residuals)
94
+
95
+ **How to check**:
96
+
97
+ **Visual methods** (primary):
98
+ - Box plots by group (for t-test/ANOVA)
99
+ - Residuals vs. fitted values plot (for regression) - should show random scatter
100
+ - Scale-location plot (square root of standardized residuals vs. fitted)
101
+
102
+ **Formal tests** (secondary):
103
+ - Levene's test (robust to non-normality)
104
+ - Bartlett's test (sensitive to non-normality, not recommended)
105
+ - Brown-Forsythe test (median-based version of Levene's)
106
+ - Breusch-Pagan test (for regression)
107
+
108
+ **Python implementation**:
109
+ ```python
110
+ from scipy import stats
111
+ import pingouin as pg
112
+
113
+ # Levene's test
114
+ statistic, p_value = stats.levene(group1, group2, group3)
115
+
116
+ # For regression
117
+ # Breusch-Pagan test
118
+ from statsmodels.stats.diagnostic import het_breuschpagan
119
+ _, p_value, _, _ = het_breuschpagan(residuals, exog)
120
+ ```
121
+
122
+ **Interpretation guidance**:
123
+ - Variance ratio (max/min) < 2-3: Generally acceptable
124
+ - For ANOVA: Test is robust if groups have equal sizes
125
+ - For regression: Look for funnel patterns in residual plots
126
+
127
+ **What to do if violated**:
128
+ - **t-test**: Use Welch's t-test (does not assume equal variances)
129
+ - **ANOVA**: Use Welch's ANOVA or Brown-Forsythe ANOVA
130
+ - **Regression**:
131
+ - Transform dependent variable (log, square root)
132
+ - Use weighted least squares (WLS)
133
+ - Use robust standard errors (HC3)
134
+ - Use generalized linear models (GLM) with appropriate variance function
135
+
136
+ **Critical severity**: MEDIUM - tests can be robust with equal sample sizes
137
+
138
+ ---
139
+
140
+ ## Test-Specific Assumptions
141
+
142
+ ### T-Tests
143
+
144
+ **Assumptions**:
145
+ 1. Independence of observations
146
+ 2. Normality (each group for independent t-test; differences for paired t-test)
147
+ 3. Homogeneity of variance (independent t-test only)
148
+
149
+ **Diagnostic workflow**:
150
+ ```python
151
+ import scipy.stats as stats
152
+ import pingouin as pg
153
+
154
+ # Check normality for each group
155
+ stats.shapiro(group1)
156
+ stats.shapiro(group2)
157
+
158
+ # Check homogeneity of variance
159
+ stats.levene(group1, group2)
160
+
161
+ # If assumptions violated:
162
+ # Option 1: Welch's t-test (unequal variances)
163
+ pg.ttest(group1, group2, correction=False) # Welch's
164
+
165
+ # Option 2: Non-parametric alternative
166
+ pg.mwu(group1, group2) # Mann-Whitney U
167
+ ```
168
+
169
+ ---
170
+
171
+ ### ANOVA
172
+
173
+ **Assumptions**:
174
+ 1. Independence of observations within and between groups
175
+ 2. Normality in each group
176
+ 3. Homogeneity of variance across groups
177
+
178
+ **Additional considerations**:
179
+ - For repeated measures ANOVA: Sphericity assumption (Mauchly's test)
180
+
181
+ **Diagnostic workflow**:
182
+ ```python
183
+ import pingouin as pg
184
+
185
+ # Check normality per group
186
+ for group in df['group'].unique():
187
+ data = df[df['group'] == group]['value']
188
+ stats.shapiro(data)
189
+
190
+ # Check homogeneity of variance
191
+ pg.homoscedasticity(df, dv='value', group='group')
192
+
193
+ # For repeated measures: Check sphericity
194
+ # Automatically tested in pingouin's rm_anova
195
+ ```
196
+
197
+ **What to do if sphericity violated** (repeated measures):
198
+ - Greenhouse-Geisser correction (ε < 0.75)
199
+ - Huynh-Feldt correction (ε > 0.75)
200
+ - Use multivariate approach (MANOVA)
201
+
202
+ ---
203
+
204
+ ### Linear Regression
205
+
206
+ **Assumptions**:
207
+ 1. **Linearity**: Relationship between X and Y is linear
208
+ 2. **Independence**: Residuals are independent
209
+ 3. **Homoscedasticity**: Constant variance of residuals
210
+ 4. **Normality**: Residuals are normally distributed
211
+ 5. **No multicollinearity**: Predictors are not highly correlated (multiple regression)
212
+
213
+ **Diagnostic workflow**:
214
+
215
+ **1. Linearity**:
216
+ ```python
217
+ import matplotlib.pyplot as plt
218
+ import seaborn as sns
219
+
220
+ # Scatter plots of Y vs each X
221
+ # Residuals vs. fitted values (should be randomly scattered)
222
+ plt.scatter(fitted_values, residuals)
223
+ plt.axhline(y=0, color='r', linestyle='--')
224
+ ```
225
+
226
+ **2. Independence**:
227
+ ```python
228
+ from statsmodels.stats.stattools import durbin_watson
229
+
230
+ # Durbin-Watson test (for time series)
231
+ dw_statistic = durbin_watson(residuals)
232
+ # Values between 1.5-2.5 suggest independence
233
+ ```
234
+
235
+ **3. Homoscedasticity**:
236
+ ```python
237
+ # Breusch-Pagan test
238
+ from statsmodels.stats.diagnostic import het_breuschpagan
239
+ _, p_value, _, _ = het_breuschpagan(residuals, exog)
240
+
241
+ # Visual: Scale-location plot
242
+ plt.scatter(fitted_values, np.sqrt(np.abs(std_residuals)))
243
+ ```
244
+
245
+ **4. Normality of residuals**:
246
+ ```python
247
+ # Q-Q plot of residuals
248
+ stats.probplot(residuals, dist="norm", plot=plt)
249
+
250
+ # Shapiro-Wilk test
251
+ stats.shapiro(residuals)
252
+ ```
253
+
254
+ **5. Multicollinearity**:
255
+ ```python
256
+ from statsmodels.stats.outliers_influence import variance_inflation_factor
257
+
258
+ # Calculate VIF for each predictor
259
+ vif_data = pd.DataFrame()
260
+ vif_data["feature"] = X.columns
261
+ vif_data["VIF"] = [variance_inflation_factor(X.values, i) for i in range(len(X.columns))]
262
+
263
+ # VIF > 10 indicates severe multicollinearity
264
+ # VIF > 5 indicates moderate multicollinearity
265
+ ```
266
+
267
+ **What to do if violated**:
268
+ - **Non-linearity**: Add polynomial terms, use GAM, or transform variables
269
+ - **Heteroscedasticity**: Transform Y, use WLS, use robust SE
270
+ - **Non-normal residuals**: Transform Y, use robust methods, check for outliers
271
+ - **Multicollinearity**: Remove correlated predictors, use PCA, ridge regression
272
+
273
+ ---
274
+
275
+ ### Logistic Regression
276
+
277
+ **Assumptions**:
278
+ 1. **Independence**: Observations are independent
279
+ 2. **Linearity**: Linear relationship between log-odds and continuous predictors
280
+ 3. **No perfect multicollinearity**: Predictors not perfectly correlated
281
+ 4. **Large sample size**: At least 10-20 events per predictor
282
+
283
+ **Diagnostic workflow**:
284
+
285
+ **1. Linearity of logit**:
286
+ ```python
287
+ # Box-Tidwell test: Add interaction with log of continuous predictor
288
+ # If interaction is significant, linearity violated
289
+ ```
290
+
291
+ **2. Multicollinearity**:
292
+ ```python
293
+ # Use VIF as in linear regression
294
+ ```
295
+
296
+ **3. Influential observations**:
297
+ ```python
298
+ # Cook's distance, DFBetas, leverage
299
+ from statsmodels.stats.outliers_influence import OLSInfluence
300
+
301
+ influence = OLSInfluence(model)
302
+ cooks_d = influence.cooks_distance
303
+ ```
304
+
305
+ **4. Model fit**:
306
+ ```python
307
+ # Hosmer-Lemeshow test
308
+ # Pseudo R-squared
309
+ # Classification metrics (accuracy, AUC-ROC)
310
+ ```
311
+
312
+ ---
313
+
314
+ ## Outlier Detection
315
+
316
+ **Methods**:
317
+ 1. **Visual**: Box plots, scatter plots
318
+ 2. **Statistical**:
319
+ - Z-scores: |z| > 3 suggests outlier
320
+ - IQR method: Values < Q1 - 1.5×IQR or > Q3 + 1.5×IQR
321
+ - Modified Z-score using median absolute deviation (robust to outliers)
322
+
323
+ **For regression**:
324
+ - **Leverage**: High leverage points (hat values)
325
+ - **Influence**: Cook's distance > 4/n suggests influential point
326
+ - **Outliers**: Studentized residuals > ±3
327
+
328
+ **What to do**:
329
+ 1. Investigate data entry errors
330
+ 2. Consider if outliers are valid observations
331
+ 3. Report sensitivity analysis (results with and without outliers)
332
+ 4. Use robust methods if outliers are legitimate
333
+
334
+ ---
335
+
336
+ ## Sample Size Considerations
337
+
338
+ ### Minimum Sample Sizes (Rules of Thumb)
339
+
340
+ - **T-test**: n ≥ 30 per group for robustness to non-normality
341
+ - **ANOVA**: n ≥ 30 per group
342
+ - **Correlation**: n ≥ 30 for adequate power
343
+ - **Simple regression**: n ≥ 50
344
+ - **Multiple regression**: n ≥ 10-20 per predictor (minimum 10 + k predictors)
345
+ - **Logistic regression**: n ≥ 10-20 events per predictor
346
+
347
+ ### Small Sample Considerations
348
+
349
+ For small samples:
350
+ - Assumptions become more critical
351
+ - Use exact tests when available (Fisher's exact, exact logistic regression)
352
+ - Consider non-parametric alternatives
353
+ - Use permutation tests or bootstrap methods
354
+ - Be conservative with interpretation
355
+
356
+ ---
357
+
358
+ ## Reporting Assumption Checks
359
+
360
+ When reporting analyses, include:
361
+
362
+ 1. **Statement of assumptions checked**: List all assumptions tested
363
+ 2. **Methods used**: Describe visual and formal tests employed
364
+ 3. **Results of diagnostic tests**: Report test statistics and p-values
365
+ 4. **Assessment**: State whether assumptions were met or violated
366
+ 5. **Actions taken**: If violated, describe remedial actions (transformations, alternative tests, robust methods)
367
+
368
+ **Example reporting statement**:
369
+ > "Normality was assessed using Shapiro-Wilk tests and Q-Q plots. Data for Group A (W = 0.97, p = .18) and Group B (W = 0.96, p = .12) showed no significant departure from normality. Homogeneity of variance was assessed using Levene's test, which was non-significant (F(1, 58) = 1.23, p = .27), indicating equal variances across groups. Therefore, assumptions for the independent samples t-test were satisfied."