clearai-dsh 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +26 -0
- package/LICENSE +201 -0
- package/README.md +138 -0
- package/README.zh-CN.md +138 -0
- package/bin/clearai.mjs +224 -0
- package/brand/README.md +41 -0
- package/brand/logo-512-dark.png +0 -0
- package/brand/logo-512.png +0 -0
- package/brand/logo-lockup-dark.png +0 -0
- package/brand/logo-lockup.png +0 -0
- package/brand/logo-lockup.svg +12 -0
- package/brand/logo-wordmark.svg +6 -0
- package/brand/logo.svg +19 -0
- package/cordis.patch.yml +39 -0
- package/lib/client.js +3071 -0
- package/lib/fold.js +1576 -0
- package/lib/host.js +605 -0
- package/package.json +65 -0
- package/presets/clearai/agent.cordis.yml +226 -0
- package/presets/clearai/plugins/brain.js +547 -0
- package/presets/clearai/plugins/clearai-kernel.js +5485 -0
- package/presets/clearai/plugins/ontology.js +306 -0
- package/presets/clearai/plugins/prompts.js +312 -0
- package/presets/clearai/preset.yml +5 -0
- package/presets/clearai/skills/clearai-loop/SKILL.md +89 -0
- package/presets/clearai/template/knowledge/README.md +25 -0
- package/presets/clearai/template/memory/README.md +34 -0
- package/presets/clearai/template/project.md +49 -0
- package/presets/clearai/template/skills/README.md +37 -0
- package/presets/clearai/template/skills/chart-diagram-qa/SKILL.md +43 -0
- package/presets/clearai/template/skills/citation-management/SKILL.md +73 -0
- package/presets/clearai/template/skills/citation-management/references/bibtex_formatting.md +908 -0
- package/presets/clearai/template/skills/citation-management/references/citation_validation.md +794 -0
- package/presets/clearai/template/skills/citation-management/references/google_scholar_search.md +725 -0
- package/presets/clearai/template/skills/citation-management/references/metadata_extraction.md +870 -0
- package/presets/clearai/template/skills/citation-management/references/pubmed_search.md +839 -0
- package/presets/clearai/template/skills/citation-management/scripts/doi_to_bibtex.py +204 -0
- package/presets/clearai/template/skills/citation-management/scripts/extract_metadata.py +569 -0
- package/presets/clearai/template/skills/citation-management/scripts/format_bibtex.py +349 -0
- package/presets/clearai/template/skills/citation-management/scripts/generate_schematic.py +139 -0
- package/presets/clearai/template/skills/citation-management/scripts/generate_schematic_ai.py +817 -0
- package/presets/clearai/template/skills/citation-management/scripts/search_google_scholar.py +282 -0
- package/presets/clearai/template/skills/citation-management/scripts/search_pubmed.py +398 -0
- package/presets/clearai/template/skills/citation-management/scripts/validate_citations.py +497 -0
- package/presets/clearai/template/skills/data-analysis/SKILL.md +92 -0
- package/presets/clearai/template/skills/data-analysis/checklists/readiness_check.md +23 -0
- package/presets/clearai/template/skills/data-analysis/templates/analysis_report.md.tpl +63 -0
- package/presets/clearai/template/skills/data-analysis/templates/cleaning_rules_draft.yaml.tpl +32 -0
- package/presets/clearai/template/skills/data-analysis/templates/data_dictionary.md.tpl +12 -0
- package/presets/clearai/template/skills/data-analysis/templates/domain_knowledge_template.md.tpl +316 -0
- package/presets/clearai/template/skills/data-analysis/templates/feature_candidates.json.tpl +20 -0
- package/presets/clearai/template/skills/data-analysis/templates/quality_scorecard.md.tpl +30 -0
- package/presets/clearai/template/skills/data-analysis/workflows/01-data-profiling.md +42 -0
- package/presets/clearai/template/skills/data-analysis/workflows/02-quality-audit.md +36 -0
- package/presets/clearai/template/skills/data-analysis/workflows/03-physical-correlation.md +25 -0
- package/presets/clearai/template/skills/data-analysis/workflows/04-unstructured-mining.md +26 -0
- package/presets/clearai/template/skills/data-qa-analysis/SKILL.md +102 -0
- package/presets/clearai/template/skills/data-qa-analysis/checklists/readiness_check.md +62 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/best_in_class_report.md.tpl +56 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/cleaning_rules_draft.yaml.tpl +56 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/data_dictionary.md.tpl +13 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/data_source_inventory_and_lineage.md.tpl +146 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/data_status_report.md.tpl +60 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/steady_state_rules.yaml.tpl +41 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/subsystem_registry.md.tpl +101 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/unified_execution_plan.md.tpl +100 -0
- package/presets/clearai/template/skills/data-qa-analysis/workflows/01-data-source-inventory-and-lineage.md +194 -0
- package/presets/clearai/template/skills/data-qa-analysis/workflows/02-data-alignment-and-tag-semantics.md +122 -0
- package/presets/clearai/template/skills/data-qa-analysis/workflows/03-steady-state-identification.md +126 -0
- package/presets/clearai/template/skills/data-qa-analysis/workflows/04-consumption-analysis.md +152 -0
- package/presets/clearai/template/skills/data-qa-analysis/workflows/05-best-in-class-and-optimization-space.md +78 -0
- package/presets/clearai/template/skills/domain-presearch/SKILL.md +131 -0
- package/presets/clearai/template/skills/domain-presearch/checklists/domain_checklist.md +24 -0
- package/presets/clearai/template/skills/domain-presearch/references/figure_code.md +78 -0
- package/presets/clearai/template/skills/domain-presearch/references/strategic_frameworks.md +38 -0
- package/presets/clearai/template/skills/exploration-loop/SKILL.md +81 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/SKILL.md +77 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/references/bioinformatics_genomics_formats.md +664 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/references/chemistry_molecular_formats.md +664 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/references/general_scientific_formats.md +518 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/references/microscopy_imaging_formats.md +620 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/references/proteomics_metabolomics_formats.md +517 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/references/spectroscopy_analytical_formats.md +633 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/scripts/eda_analyzer.py +547 -0
- package/presets/clearai/template/skills/hypothesis-generation/SKILL.md +73 -0
- package/presets/clearai/template/skills/hypothesis-generation/references/experimental_design_patterns.md +329 -0
- package/presets/clearai/template/skills/hypothesis-generation/references/hypothesis_quality_criteria.md +198 -0
- package/presets/clearai/template/skills/hypothesis-generation/references/literature_search_strategies.md +622 -0
- package/presets/clearai/template/skills/hypothesis-generation/scripts/generate_schematic.py +139 -0
- package/presets/clearai/template/skills/hypothesis-generation/scripts/generate_schematic_ai.py +817 -0
- package/presets/clearai/template/skills/literature-review/SKILL.md +72 -0
- package/presets/clearai/template/skills/literature-review/references/citation_styles.md +166 -0
- package/presets/clearai/template/skills/literature-review/references/database_strategies.md +455 -0
- package/presets/clearai/template/skills/literature-review/scripts/generate_pdf.py +176 -0
- package/presets/clearai/template/skills/literature-review/scripts/generate_schematic.py +139 -0
- package/presets/clearai/template/skills/literature-review/scripts/generate_schematic_ai.py +817 -0
- package/presets/clearai/template/skills/literature-review/scripts/search_databases.py +303 -0
- package/presets/clearai/template/skills/literature-review/scripts/verify_citations.py +221 -0
- package/presets/clearai/template/skills/paper-lookup/SKILL.md +59 -0
- package/presets/clearai/template/skills/paper-lookup/references/arxiv.md +161 -0
- package/presets/clearai/template/skills/paper-lookup/references/biorxiv.md +118 -0
- package/presets/clearai/template/skills/paper-lookup/references/core.md +150 -0
- package/presets/clearai/template/skills/paper-lookup/references/crossref.md +181 -0
- package/presets/clearai/template/skills/paper-lookup/references/medrxiv.md +104 -0
- package/presets/clearai/template/skills/paper-lookup/references/openalex.md +174 -0
- package/presets/clearai/template/skills/paper-lookup/references/pmc.md +152 -0
- package/presets/clearai/template/skills/paper-lookup/references/pubmed.md +124 -0
- package/presets/clearai/template/skills/paper-lookup/references/semantic-scholar.md +203 -0
- package/presets/clearai/template/skills/paper-lookup/references/unpaywall.md +127 -0
- package/presets/clearai/template/skills/process-presearch/SKILL.md +196 -0
- package/presets/clearai/template/skills/process-presearch/checklists/process_checklist.md +18 -0
- package/presets/clearai/template/skills/process-presearch/references/figure_code.md +107 -0
- package/presets/clearai/template/skills/process-presearch/references/source_attribution_example.md +22 -0
- package/presets/clearai/template/skills/process-understanding-extraction/SKILL.md +69 -0
- package/presets/clearai/template/skills/process-understanding-extraction/checklists/readiness_check.md +34 -0
- package/presets/clearai/template/skills/process-understanding-extraction/templates/docx_raw_dump_extractor.py.tpl +132 -0
- package/presets/clearai/template/skills/process-understanding-extraction/templates/entity_map_unit_topology.json.tpl +86 -0
- package/presets/clearai/template/skills/process-understanding-extraction/templates/process_brief.md.tpl +89 -0
- package/presets/clearai/template/skills/process-understanding-extraction/templates/process_brief_builder_from_raw_dump.py.tpl +203 -0
- package/presets/clearai/template/skills/process-understanding-extraction/templates/process_flow_mermaid.md.tpl +41 -0
- package/presets/clearai/template/skills/process-understanding-extraction/templates/unified_execution_plan.md.tpl +53 -0
- package/presets/clearai/template/skills/process-understanding-extraction/workflows/01-process-doc-discovery.md +173 -0
- package/presets/clearai/template/skills/process-understanding-extraction/workflows/02-process-understanding-and-diagramming.md +106 -0
- package/presets/clearai/template/skills/scientific-brainstorming/SKILL.md +64 -0
- package/presets/clearai/template/skills/scientific-brainstorming/references/brainstorming_methods.md +326 -0
- package/presets/clearai/template/skills/scientific-critical-thinking/SKILL.md +72 -0
- package/presets/clearai/template/skills/scientific-critical-thinking/references/common_biases.md +364 -0
- package/presets/clearai/template/skills/scientific-critical-thinking/references/evidence_hierarchy.md +485 -0
- package/presets/clearai/template/skills/scientific-critical-thinking/references/experimental_design.md +496 -0
- package/presets/clearai/template/skills/scientific-critical-thinking/references/logical_fallacies.md +478 -0
- package/presets/clearai/template/skills/scientific-critical-thinking/references/scientific_method.md +169 -0
- package/presets/clearai/template/skills/scientific-critical-thinking/references/statistical_pitfalls.md +506 -0
- package/presets/clearai/template/skills/skill-creator/SKILL.md +109 -0
- package/presets/clearai/template/skills/skill-creator/references/authoring-guide.md +89 -0
- package/presets/clearai/template/skills/statistical-analysis/SKILL.md +79 -0
- package/presets/clearai/template/skills/statistical-analysis/references/assumptions_and_diagnostics.md +369 -0
- package/presets/clearai/template/skills/statistical-analysis/references/bayesian_statistics.md +653 -0
- package/presets/clearai/template/skills/statistical-analysis/references/effect_sizes_and_power.md +578 -0
- package/presets/clearai/template/skills/statistical-analysis/references/reporting_standards.md +469 -0
- package/presets/clearai/template/skills/statistical-analysis/references/test_selection_guide.md +129 -0
- package/presets/clearai/template/skills/statistical-analysis/scripts/assumption_checks.py +538 -0
- package/presets/clearai/template/skills/web-artifact/SKILL.md +165 -0
- package/presets/clearai/template/skills/web-artifact/assets/renderer/renderer.css +229 -0
- package/presets/clearai/template/skills/web-artifact/assets/renderer/renderer.js +373 -0
- package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/LICENSE +263 -0
- package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/UPSTREAM.md +26 -0
- package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/elk.bundled.js +6605 -0
- package/presets/clearai/template/skills/web-artifact/references/when-drawing-a-topology.md +150 -0
- package/presets/clearai/template/skills/web-artifact/references/when-the-page-must-work-offline.md +62 -0
- package/presets/clearai/template/skills/web-artifact/scripts/check_artifact.py +167 -0
- package/presets/clearai/template/skills/web-artifact/scripts/render_topology.js +272 -0
- package/presets/clearai/template/skills/what-if-oracle/LICENSE.txt +5 -0
- package/presets/clearai/template/skills/what-if-oracle/SKILL.md +72 -0
- package/presets/clearai/template/skills/what-if-oracle/references/scenario-templates.md +154 -0
|
@@ -0,0 +1,79 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: statistical-analysis
|
|
3
|
+
description: |
|
|
4
|
+
【假设验证·统计检验】选择合适检验方法、检查前提、执行分析并产出可审计结论。适用:验证两组条件差异、相关性、回归关系、组间对比。不适用:首触未知数据文件(用 exploratory-data-analysis);深度关联分析与诊断(用 data-analysis)。
|
|
5
|
+
license: MIT license
|
|
6
|
+
metadata:
|
|
7
|
+
version: 1.0-clearai
|
|
8
|
+
skill-author: K-Dense Inc. (adapted for ClearAI)
|
|
9
|
+
tier: system
|
|
10
|
+
origin: template
|
|
11
|
+
created_at: '2026-06-12T02:47:05.410760+00:00'
|
|
12
|
+
---
|
|
13
|
+
|
|
14
|
+
# 统计分析 Skill(ClearAI 版)
|
|
15
|
+
|
|
16
|
+
## 使用边界
|
|
17
|
+
|
|
18
|
+
- **适用**:已有明确假设与干净表格数据,需统计检验支撑结论。
|
|
19
|
+
- **不适用**:数据尚未探查 → `exploratory-data-analysis`;需深度关联与根因分析 → `data-analysis`。
|
|
20
|
+
|
|
21
|
+
## ClearAI 工具与路径映射
|
|
22
|
+
|
|
23
|
+
- 数据 → `read` / `bash`(pandas 读 `input/` 或 `lab/`)
|
|
24
|
+
- 分析脚本 → `lab/scripts/stats_*.py`(可用已复制 `scripts/` 模板)
|
|
25
|
+
- 结果表 → `lab/extracted/stats_results.md` 或 `.csv`
|
|
26
|
+
- 图表 → `lab/diagrams/`
|
|
27
|
+
- 交付摘要 → `products/reports/`(若用户要求正式报告)
|
|
28
|
+
- 经验回写 → `clear/memory/statistical_analysis_lessons.md`
|
|
29
|
+
|
|
30
|
+
## 环境契约
|
|
31
|
+
|
|
32
|
+
```bash
|
|
33
|
+
python -c "import pandas, scipy, statsmodels, pingouin; print('statistics capability ready')"
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
上述统计栈属于 ClearAI 基础安装能力。若 import 验证失败,记录缺失包与当前 Python
|
|
37
|
+
路径并停止,提示用户修复部署;**禁止**在 Agent 运行中执行 `pip install` 或 `uv`。
|
|
38
|
+
|
|
39
|
+
## 工作流
|
|
40
|
+
|
|
41
|
+
### 1. 检验选型
|
|
42
|
+
|
|
43
|
+
| 问题类型 | 典型检验 |
|
|
44
|
+
|----------|----------|
|
|
45
|
+
| 两组均值 | t 检验 / Mann-Whitney |
|
|
46
|
+
| 多组均值 | ANOVA / Kruskal-Wallis |
|
|
47
|
+
| 分类关联 | 卡方 |
|
|
48
|
+
| 两变量关系 | Pearson/Spearman 相关、OLS 回归 |
|
|
49
|
+
| 时序对比 | 配对检验、变化点(简述) |
|
|
50
|
+
|
|
51
|
+
详见 `references/test_selection.md`、`references/assumptions_and_diagnostics.md`。
|
|
52
|
+
|
|
53
|
+
### 2. 前提检查
|
|
54
|
+
|
|
55
|
+
- 正态性、方差齐性、样本量、独立性
|
|
56
|
+
- 不满足时选用非参数或明确声明局限
|
|
57
|
+
|
|
58
|
+
### 3. 执行与报告(工程口径,非 APA)
|
|
59
|
+
|
|
60
|
+
每条结论须含:
|
|
61
|
+
- 检验名称、统计量、p 值/置信区间、效应量
|
|
62
|
+
- **物理证据**:基于哪份文件、哪几列、样本量 n=?
|
|
63
|
+
- 工程意义:差异是否在业务上显著?
|
|
64
|
+
|
|
65
|
+
### 4. 证据闭环
|
|
66
|
+
|
|
67
|
+
- 无统计输出文件/图表路径 → 不得宣称「已验证」
|
|
68
|
+
- 与 `hypothesis-generation` 产出的 H1/H2 逐条对应
|
|
69
|
+
|
|
70
|
+
## 报告模板(`lab/extracted/stats_results.md`)
|
|
71
|
+
|
|
72
|
+
```markdown
|
|
73
|
+
## 假设 H1: ...
|
|
74
|
+
- 方法:...
|
|
75
|
+
- 结果:stat=..., p=..., effect_size=...
|
|
76
|
+
- 证据:lab/diagrams/xxx.png, 输入=input/yyy.csv
|
|
77
|
+
- 结论:支持/拒绝/ inconclusive
|
|
78
|
+
- 局限:...
|
|
79
|
+
```
|
|
@@ -0,0 +1,369 @@
|
|
|
1
|
+
# Statistical Assumptions and Diagnostic Procedures
|
|
2
|
+
|
|
3
|
+
This document provides comprehensive guidance on checking and validating statistical assumptions for various analyses.
|
|
4
|
+
|
|
5
|
+
## General Principles
|
|
6
|
+
|
|
7
|
+
1. **Always check assumptions before interpreting test results**
|
|
8
|
+
2. **Use multiple diagnostic methods** (visual + formal tests)
|
|
9
|
+
3. **Consider robustness**: Some tests are robust to violations under certain conditions
|
|
10
|
+
4. **Document all assumption checks** in analysis reports
|
|
11
|
+
5. **Report violations and remedial actions taken**
|
|
12
|
+
|
|
13
|
+
## Common Assumptions Across Tests
|
|
14
|
+
|
|
15
|
+
### 1. Independence of Observations
|
|
16
|
+
|
|
17
|
+
**What it means**: Each observation is independent; measurements on one subject do not influence measurements on another.
|
|
18
|
+
|
|
19
|
+
**How to check**:
|
|
20
|
+
- Review study design and data collection procedures
|
|
21
|
+
- For time series: Check autocorrelation (ACF/PACF plots, Durbin-Watson test)
|
|
22
|
+
- For clustered data: Consider intraclass correlation (ICC)
|
|
23
|
+
|
|
24
|
+
**What to do if violated**:
|
|
25
|
+
- Use mixed-effects models for clustered/hierarchical data
|
|
26
|
+
- Use time series methods for temporally dependent data
|
|
27
|
+
- Use generalized estimating equations (GEE) for correlated data
|
|
28
|
+
|
|
29
|
+
**Critical severity**: HIGH - violations can severely inflate Type I error
|
|
30
|
+
|
|
31
|
+
---
|
|
32
|
+
|
|
33
|
+
### 2. Normality
|
|
34
|
+
|
|
35
|
+
**What it means**: Data or residuals follow a normal (Gaussian) distribution.
|
|
36
|
+
|
|
37
|
+
**When required**:
|
|
38
|
+
- t-tests (for small samples; robust for n > 30 per group)
|
|
39
|
+
- ANOVA (for small samples; robust for n > 30 per group)
|
|
40
|
+
- Linear regression (for residuals)
|
|
41
|
+
- Some correlation tests (Pearson)
|
|
42
|
+
|
|
43
|
+
**How to check**:
|
|
44
|
+
|
|
45
|
+
**Visual methods** (primary):
|
|
46
|
+
- Q-Q (quantile-quantile) plot: Points should fall on diagonal line
|
|
47
|
+
- Histogram with normal curve overlay
|
|
48
|
+
- Kernel density plot
|
|
49
|
+
|
|
50
|
+
**Formal tests** (secondary):
|
|
51
|
+
- Shapiro-Wilk test (recommended for n < 50)
|
|
52
|
+
- Kolmogorov-Smirnov test
|
|
53
|
+
- Anderson-Darling test
|
|
54
|
+
|
|
55
|
+
**Python implementation**:
|
|
56
|
+
```python
|
|
57
|
+
from scipy import stats
|
|
58
|
+
import matplotlib.pyplot as plt
|
|
59
|
+
|
|
60
|
+
# Shapiro-Wilk test
|
|
61
|
+
statistic, p_value = stats.shapiro(data)
|
|
62
|
+
|
|
63
|
+
# Q-Q plot
|
|
64
|
+
stats.probplot(data, dist="norm", plot=plt)
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
**Interpretation guidance**:
|
|
68
|
+
- For n < 30: Both visual and formal tests important
|
|
69
|
+
- For 30 ≤ n < 100: Visual inspection primary, formal tests secondary
|
|
70
|
+
- For n ≥ 100: Formal tests overly sensitive; rely on visual inspection
|
|
71
|
+
- Look for severe skewness, outliers, or bimodality
|
|
72
|
+
|
|
73
|
+
**What to do if violated**:
|
|
74
|
+
- **Mild violations** (slight skewness): Proceed if n > 30 per group
|
|
75
|
+
- **Moderate violations**: Use non-parametric alternatives (Mann-Whitney, Kruskal-Wallis, Wilcoxon)
|
|
76
|
+
- **Severe violations**:
|
|
77
|
+
- Transform data (log, square root, Box-Cox)
|
|
78
|
+
- Use non-parametric methods
|
|
79
|
+
- Use robust regression methods
|
|
80
|
+
- Consider bootstrapping
|
|
81
|
+
|
|
82
|
+
**Critical severity**: MEDIUM - parametric tests are often robust to mild violations with adequate sample size
|
|
83
|
+
|
|
84
|
+
---
|
|
85
|
+
|
|
86
|
+
### 3. Homogeneity of Variance (Homoscedasticity)
|
|
87
|
+
|
|
88
|
+
**What it means**: Variances are equal across groups or across the range of predictors.
|
|
89
|
+
|
|
90
|
+
**When required**:
|
|
91
|
+
- Independent samples t-test
|
|
92
|
+
- ANOVA
|
|
93
|
+
- Linear regression (constant variance of residuals)
|
|
94
|
+
|
|
95
|
+
**How to check**:
|
|
96
|
+
|
|
97
|
+
**Visual methods** (primary):
|
|
98
|
+
- Box plots by group (for t-test/ANOVA)
|
|
99
|
+
- Residuals vs. fitted values plot (for regression) - should show random scatter
|
|
100
|
+
- Scale-location plot (square root of standardized residuals vs. fitted)
|
|
101
|
+
|
|
102
|
+
**Formal tests** (secondary):
|
|
103
|
+
- Levene's test (robust to non-normality)
|
|
104
|
+
- Bartlett's test (sensitive to non-normality, not recommended)
|
|
105
|
+
- Brown-Forsythe test (median-based version of Levene's)
|
|
106
|
+
- Breusch-Pagan test (for regression)
|
|
107
|
+
|
|
108
|
+
**Python implementation**:
|
|
109
|
+
```python
|
|
110
|
+
from scipy import stats
|
|
111
|
+
import pingouin as pg
|
|
112
|
+
|
|
113
|
+
# Levene's test
|
|
114
|
+
statistic, p_value = stats.levene(group1, group2, group3)
|
|
115
|
+
|
|
116
|
+
# For regression
|
|
117
|
+
# Breusch-Pagan test
|
|
118
|
+
from statsmodels.stats.diagnostic import het_breuschpagan
|
|
119
|
+
_, p_value, _, _ = het_breuschpagan(residuals, exog)
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
**Interpretation guidance**:
|
|
123
|
+
- Variance ratio (max/min) < 2-3: Generally acceptable
|
|
124
|
+
- For ANOVA: Test is robust if groups have equal sizes
|
|
125
|
+
- For regression: Look for funnel patterns in residual plots
|
|
126
|
+
|
|
127
|
+
**What to do if violated**:
|
|
128
|
+
- **t-test**: Use Welch's t-test (does not assume equal variances)
|
|
129
|
+
- **ANOVA**: Use Welch's ANOVA or Brown-Forsythe ANOVA
|
|
130
|
+
- **Regression**:
|
|
131
|
+
- Transform dependent variable (log, square root)
|
|
132
|
+
- Use weighted least squares (WLS)
|
|
133
|
+
- Use robust standard errors (HC3)
|
|
134
|
+
- Use generalized linear models (GLM) with appropriate variance function
|
|
135
|
+
|
|
136
|
+
**Critical severity**: MEDIUM - tests can be robust with equal sample sizes
|
|
137
|
+
|
|
138
|
+
---
|
|
139
|
+
|
|
140
|
+
## Test-Specific Assumptions
|
|
141
|
+
|
|
142
|
+
### T-Tests
|
|
143
|
+
|
|
144
|
+
**Assumptions**:
|
|
145
|
+
1. Independence of observations
|
|
146
|
+
2. Normality (each group for independent t-test; differences for paired t-test)
|
|
147
|
+
3. Homogeneity of variance (independent t-test only)
|
|
148
|
+
|
|
149
|
+
**Diagnostic workflow**:
|
|
150
|
+
```python
|
|
151
|
+
import scipy.stats as stats
|
|
152
|
+
import pingouin as pg
|
|
153
|
+
|
|
154
|
+
# Check normality for each group
|
|
155
|
+
stats.shapiro(group1)
|
|
156
|
+
stats.shapiro(group2)
|
|
157
|
+
|
|
158
|
+
# Check homogeneity of variance
|
|
159
|
+
stats.levene(group1, group2)
|
|
160
|
+
|
|
161
|
+
# If assumptions violated:
|
|
162
|
+
# Option 1: Welch's t-test (unequal variances)
|
|
163
|
+
pg.ttest(group1, group2, correction=False) # Welch's
|
|
164
|
+
|
|
165
|
+
# Option 2: Non-parametric alternative
|
|
166
|
+
pg.mwu(group1, group2) # Mann-Whitney U
|
|
167
|
+
```
|
|
168
|
+
|
|
169
|
+
---
|
|
170
|
+
|
|
171
|
+
### ANOVA
|
|
172
|
+
|
|
173
|
+
**Assumptions**:
|
|
174
|
+
1. Independence of observations within and between groups
|
|
175
|
+
2. Normality in each group
|
|
176
|
+
3. Homogeneity of variance across groups
|
|
177
|
+
|
|
178
|
+
**Additional considerations**:
|
|
179
|
+
- For repeated measures ANOVA: Sphericity assumption (Mauchly's test)
|
|
180
|
+
|
|
181
|
+
**Diagnostic workflow**:
|
|
182
|
+
```python
|
|
183
|
+
import pingouin as pg
|
|
184
|
+
|
|
185
|
+
# Check normality per group
|
|
186
|
+
for group in df['group'].unique():
|
|
187
|
+
data = df[df['group'] == group]['value']
|
|
188
|
+
stats.shapiro(data)
|
|
189
|
+
|
|
190
|
+
# Check homogeneity of variance
|
|
191
|
+
pg.homoscedasticity(df, dv='value', group='group')
|
|
192
|
+
|
|
193
|
+
# For repeated measures: Check sphericity
|
|
194
|
+
# Automatically tested in pingouin's rm_anova
|
|
195
|
+
```
|
|
196
|
+
|
|
197
|
+
**What to do if sphericity violated** (repeated measures):
|
|
198
|
+
- Greenhouse-Geisser correction (ε < 0.75)
|
|
199
|
+
- Huynh-Feldt correction (ε > 0.75)
|
|
200
|
+
- Use multivariate approach (MANOVA)
|
|
201
|
+
|
|
202
|
+
---
|
|
203
|
+
|
|
204
|
+
### Linear Regression
|
|
205
|
+
|
|
206
|
+
**Assumptions**:
|
|
207
|
+
1. **Linearity**: Relationship between X and Y is linear
|
|
208
|
+
2. **Independence**: Residuals are independent
|
|
209
|
+
3. **Homoscedasticity**: Constant variance of residuals
|
|
210
|
+
4. **Normality**: Residuals are normally distributed
|
|
211
|
+
5. **No multicollinearity**: Predictors are not highly correlated (multiple regression)
|
|
212
|
+
|
|
213
|
+
**Diagnostic workflow**:
|
|
214
|
+
|
|
215
|
+
**1. Linearity**:
|
|
216
|
+
```python
|
|
217
|
+
import matplotlib.pyplot as plt
|
|
218
|
+
import seaborn as sns
|
|
219
|
+
|
|
220
|
+
# Scatter plots of Y vs each X
|
|
221
|
+
# Residuals vs. fitted values (should be randomly scattered)
|
|
222
|
+
plt.scatter(fitted_values, residuals)
|
|
223
|
+
plt.axhline(y=0, color='r', linestyle='--')
|
|
224
|
+
```
|
|
225
|
+
|
|
226
|
+
**2. Independence**:
|
|
227
|
+
```python
|
|
228
|
+
from statsmodels.stats.stattools import durbin_watson
|
|
229
|
+
|
|
230
|
+
# Durbin-Watson test (for time series)
|
|
231
|
+
dw_statistic = durbin_watson(residuals)
|
|
232
|
+
# Values between 1.5-2.5 suggest independence
|
|
233
|
+
```
|
|
234
|
+
|
|
235
|
+
**3. Homoscedasticity**:
|
|
236
|
+
```python
|
|
237
|
+
# Breusch-Pagan test
|
|
238
|
+
from statsmodels.stats.diagnostic import het_breuschpagan
|
|
239
|
+
_, p_value, _, _ = het_breuschpagan(residuals, exog)
|
|
240
|
+
|
|
241
|
+
# Visual: Scale-location plot
|
|
242
|
+
plt.scatter(fitted_values, np.sqrt(np.abs(std_residuals)))
|
|
243
|
+
```
|
|
244
|
+
|
|
245
|
+
**4. Normality of residuals**:
|
|
246
|
+
```python
|
|
247
|
+
# Q-Q plot of residuals
|
|
248
|
+
stats.probplot(residuals, dist="norm", plot=plt)
|
|
249
|
+
|
|
250
|
+
# Shapiro-Wilk test
|
|
251
|
+
stats.shapiro(residuals)
|
|
252
|
+
```
|
|
253
|
+
|
|
254
|
+
**5. Multicollinearity**:
|
|
255
|
+
```python
|
|
256
|
+
from statsmodels.stats.outliers_influence import variance_inflation_factor
|
|
257
|
+
|
|
258
|
+
# Calculate VIF for each predictor
|
|
259
|
+
vif_data = pd.DataFrame()
|
|
260
|
+
vif_data["feature"] = X.columns
|
|
261
|
+
vif_data["VIF"] = [variance_inflation_factor(X.values, i) for i in range(len(X.columns))]
|
|
262
|
+
|
|
263
|
+
# VIF > 10 indicates severe multicollinearity
|
|
264
|
+
# VIF > 5 indicates moderate multicollinearity
|
|
265
|
+
```
|
|
266
|
+
|
|
267
|
+
**What to do if violated**:
|
|
268
|
+
- **Non-linearity**: Add polynomial terms, use GAM, or transform variables
|
|
269
|
+
- **Heteroscedasticity**: Transform Y, use WLS, use robust SE
|
|
270
|
+
- **Non-normal residuals**: Transform Y, use robust methods, check for outliers
|
|
271
|
+
- **Multicollinearity**: Remove correlated predictors, use PCA, ridge regression
|
|
272
|
+
|
|
273
|
+
---
|
|
274
|
+
|
|
275
|
+
### Logistic Regression
|
|
276
|
+
|
|
277
|
+
**Assumptions**:
|
|
278
|
+
1. **Independence**: Observations are independent
|
|
279
|
+
2. **Linearity**: Linear relationship between log-odds and continuous predictors
|
|
280
|
+
3. **No perfect multicollinearity**: Predictors not perfectly correlated
|
|
281
|
+
4. **Large sample size**: At least 10-20 events per predictor
|
|
282
|
+
|
|
283
|
+
**Diagnostic workflow**:
|
|
284
|
+
|
|
285
|
+
**1. Linearity of logit**:
|
|
286
|
+
```python
|
|
287
|
+
# Box-Tidwell test: Add interaction with log of continuous predictor
|
|
288
|
+
# If interaction is significant, linearity violated
|
|
289
|
+
```
|
|
290
|
+
|
|
291
|
+
**2. Multicollinearity**:
|
|
292
|
+
```python
|
|
293
|
+
# Use VIF as in linear regression
|
|
294
|
+
```
|
|
295
|
+
|
|
296
|
+
**3. Influential observations**:
|
|
297
|
+
```python
|
|
298
|
+
# Cook's distance, DFBetas, leverage
|
|
299
|
+
from statsmodels.stats.outliers_influence import OLSInfluence
|
|
300
|
+
|
|
301
|
+
influence = OLSInfluence(model)
|
|
302
|
+
cooks_d = influence.cooks_distance
|
|
303
|
+
```
|
|
304
|
+
|
|
305
|
+
**4. Model fit**:
|
|
306
|
+
```python
|
|
307
|
+
# Hosmer-Lemeshow test
|
|
308
|
+
# Pseudo R-squared
|
|
309
|
+
# Classification metrics (accuracy, AUC-ROC)
|
|
310
|
+
```
|
|
311
|
+
|
|
312
|
+
---
|
|
313
|
+
|
|
314
|
+
## Outlier Detection
|
|
315
|
+
|
|
316
|
+
**Methods**:
|
|
317
|
+
1. **Visual**: Box plots, scatter plots
|
|
318
|
+
2. **Statistical**:
|
|
319
|
+
- Z-scores: |z| > 3 suggests outlier
|
|
320
|
+
- IQR method: Values < Q1 - 1.5×IQR or > Q3 + 1.5×IQR
|
|
321
|
+
- Modified Z-score using median absolute deviation (robust to outliers)
|
|
322
|
+
|
|
323
|
+
**For regression**:
|
|
324
|
+
- **Leverage**: High leverage points (hat values)
|
|
325
|
+
- **Influence**: Cook's distance > 4/n suggests influential point
|
|
326
|
+
- **Outliers**: Studentized residuals > ±3
|
|
327
|
+
|
|
328
|
+
**What to do**:
|
|
329
|
+
1. Investigate data entry errors
|
|
330
|
+
2. Consider if outliers are valid observations
|
|
331
|
+
3. Report sensitivity analysis (results with and without outliers)
|
|
332
|
+
4. Use robust methods if outliers are legitimate
|
|
333
|
+
|
|
334
|
+
---
|
|
335
|
+
|
|
336
|
+
## Sample Size Considerations
|
|
337
|
+
|
|
338
|
+
### Minimum Sample Sizes (Rules of Thumb)
|
|
339
|
+
|
|
340
|
+
- **T-test**: n ≥ 30 per group for robustness to non-normality
|
|
341
|
+
- **ANOVA**: n ≥ 30 per group
|
|
342
|
+
- **Correlation**: n ≥ 30 for adequate power
|
|
343
|
+
- **Simple regression**: n ≥ 50
|
|
344
|
+
- **Multiple regression**: n ≥ 10-20 per predictor (minimum 10 + k predictors)
|
|
345
|
+
- **Logistic regression**: n ≥ 10-20 events per predictor
|
|
346
|
+
|
|
347
|
+
### Small Sample Considerations
|
|
348
|
+
|
|
349
|
+
For small samples:
|
|
350
|
+
- Assumptions become more critical
|
|
351
|
+
- Use exact tests when available (Fisher's exact, exact logistic regression)
|
|
352
|
+
- Consider non-parametric alternatives
|
|
353
|
+
- Use permutation tests or bootstrap methods
|
|
354
|
+
- Be conservative with interpretation
|
|
355
|
+
|
|
356
|
+
---
|
|
357
|
+
|
|
358
|
+
## Reporting Assumption Checks
|
|
359
|
+
|
|
360
|
+
When reporting analyses, include:
|
|
361
|
+
|
|
362
|
+
1. **Statement of assumptions checked**: List all assumptions tested
|
|
363
|
+
2. **Methods used**: Describe visual and formal tests employed
|
|
364
|
+
3. **Results of diagnostic tests**: Report test statistics and p-values
|
|
365
|
+
4. **Assessment**: State whether assumptions were met or violated
|
|
366
|
+
5. **Actions taken**: If violated, describe remedial actions (transformations, alternative tests, robust methods)
|
|
367
|
+
|
|
368
|
+
**Example reporting statement**:
|
|
369
|
+
> "Normality was assessed using Shapiro-Wilk tests and Q-Q plots. Data for Group A (W = 0.97, p = .18) and Group B (W = 0.96, p = .12) showed no significant departure from normality. Homogeneity of variance was assessed using Levene's test, which was non-significant (F(1, 58) = 1.23, p = .27), indicating equal variances across groups. Therefore, assumptions for the independent samples t-test were satisfied."
|