clearai-dsh 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (154) hide show
  1. package/CHANGELOG.md +26 -0
  2. package/LICENSE +201 -0
  3. package/README.md +138 -0
  4. package/README.zh-CN.md +138 -0
  5. package/bin/clearai.mjs +224 -0
  6. package/brand/README.md +41 -0
  7. package/brand/logo-512-dark.png +0 -0
  8. package/brand/logo-512.png +0 -0
  9. package/brand/logo-lockup-dark.png +0 -0
  10. package/brand/logo-lockup.png +0 -0
  11. package/brand/logo-lockup.svg +12 -0
  12. package/brand/logo-wordmark.svg +6 -0
  13. package/brand/logo.svg +19 -0
  14. package/cordis.patch.yml +39 -0
  15. package/lib/client.js +3071 -0
  16. package/lib/fold.js +1576 -0
  17. package/lib/host.js +605 -0
  18. package/package.json +65 -0
  19. package/presets/clearai/agent.cordis.yml +226 -0
  20. package/presets/clearai/plugins/brain.js +547 -0
  21. package/presets/clearai/plugins/clearai-kernel.js +5485 -0
  22. package/presets/clearai/plugins/ontology.js +306 -0
  23. package/presets/clearai/plugins/prompts.js +312 -0
  24. package/presets/clearai/preset.yml +5 -0
  25. package/presets/clearai/skills/clearai-loop/SKILL.md +89 -0
  26. package/presets/clearai/template/knowledge/README.md +25 -0
  27. package/presets/clearai/template/memory/README.md +34 -0
  28. package/presets/clearai/template/project.md +49 -0
  29. package/presets/clearai/template/skills/README.md +37 -0
  30. package/presets/clearai/template/skills/chart-diagram-qa/SKILL.md +43 -0
  31. package/presets/clearai/template/skills/citation-management/SKILL.md +73 -0
  32. package/presets/clearai/template/skills/citation-management/references/bibtex_formatting.md +908 -0
  33. package/presets/clearai/template/skills/citation-management/references/citation_validation.md +794 -0
  34. package/presets/clearai/template/skills/citation-management/references/google_scholar_search.md +725 -0
  35. package/presets/clearai/template/skills/citation-management/references/metadata_extraction.md +870 -0
  36. package/presets/clearai/template/skills/citation-management/references/pubmed_search.md +839 -0
  37. package/presets/clearai/template/skills/citation-management/scripts/doi_to_bibtex.py +204 -0
  38. package/presets/clearai/template/skills/citation-management/scripts/extract_metadata.py +569 -0
  39. package/presets/clearai/template/skills/citation-management/scripts/format_bibtex.py +349 -0
  40. package/presets/clearai/template/skills/citation-management/scripts/generate_schematic.py +139 -0
  41. package/presets/clearai/template/skills/citation-management/scripts/generate_schematic_ai.py +817 -0
  42. package/presets/clearai/template/skills/citation-management/scripts/search_google_scholar.py +282 -0
  43. package/presets/clearai/template/skills/citation-management/scripts/search_pubmed.py +398 -0
  44. package/presets/clearai/template/skills/citation-management/scripts/validate_citations.py +497 -0
  45. package/presets/clearai/template/skills/data-analysis/SKILL.md +92 -0
  46. package/presets/clearai/template/skills/data-analysis/checklists/readiness_check.md +23 -0
  47. package/presets/clearai/template/skills/data-analysis/templates/analysis_report.md.tpl +63 -0
  48. package/presets/clearai/template/skills/data-analysis/templates/cleaning_rules_draft.yaml.tpl +32 -0
  49. package/presets/clearai/template/skills/data-analysis/templates/data_dictionary.md.tpl +12 -0
  50. package/presets/clearai/template/skills/data-analysis/templates/domain_knowledge_template.md.tpl +316 -0
  51. package/presets/clearai/template/skills/data-analysis/templates/feature_candidates.json.tpl +20 -0
  52. package/presets/clearai/template/skills/data-analysis/templates/quality_scorecard.md.tpl +30 -0
  53. package/presets/clearai/template/skills/data-analysis/workflows/01-data-profiling.md +42 -0
  54. package/presets/clearai/template/skills/data-analysis/workflows/02-quality-audit.md +36 -0
  55. package/presets/clearai/template/skills/data-analysis/workflows/03-physical-correlation.md +25 -0
  56. package/presets/clearai/template/skills/data-analysis/workflows/04-unstructured-mining.md +26 -0
  57. package/presets/clearai/template/skills/data-qa-analysis/SKILL.md +102 -0
  58. package/presets/clearai/template/skills/data-qa-analysis/checklists/readiness_check.md +62 -0
  59. package/presets/clearai/template/skills/data-qa-analysis/templates/best_in_class_report.md.tpl +56 -0
  60. package/presets/clearai/template/skills/data-qa-analysis/templates/cleaning_rules_draft.yaml.tpl +56 -0
  61. package/presets/clearai/template/skills/data-qa-analysis/templates/data_dictionary.md.tpl +13 -0
  62. package/presets/clearai/template/skills/data-qa-analysis/templates/data_source_inventory_and_lineage.md.tpl +146 -0
  63. package/presets/clearai/template/skills/data-qa-analysis/templates/data_status_report.md.tpl +60 -0
  64. package/presets/clearai/template/skills/data-qa-analysis/templates/steady_state_rules.yaml.tpl +41 -0
  65. package/presets/clearai/template/skills/data-qa-analysis/templates/subsystem_registry.md.tpl +101 -0
  66. package/presets/clearai/template/skills/data-qa-analysis/templates/unified_execution_plan.md.tpl +100 -0
  67. package/presets/clearai/template/skills/data-qa-analysis/workflows/01-data-source-inventory-and-lineage.md +194 -0
  68. package/presets/clearai/template/skills/data-qa-analysis/workflows/02-data-alignment-and-tag-semantics.md +122 -0
  69. package/presets/clearai/template/skills/data-qa-analysis/workflows/03-steady-state-identification.md +126 -0
  70. package/presets/clearai/template/skills/data-qa-analysis/workflows/04-consumption-analysis.md +152 -0
  71. package/presets/clearai/template/skills/data-qa-analysis/workflows/05-best-in-class-and-optimization-space.md +78 -0
  72. package/presets/clearai/template/skills/domain-presearch/SKILL.md +131 -0
  73. package/presets/clearai/template/skills/domain-presearch/checklists/domain_checklist.md +24 -0
  74. package/presets/clearai/template/skills/domain-presearch/references/figure_code.md +78 -0
  75. package/presets/clearai/template/skills/domain-presearch/references/strategic_frameworks.md +38 -0
  76. package/presets/clearai/template/skills/exploration-loop/SKILL.md +81 -0
  77. package/presets/clearai/template/skills/exploratory-data-analysis/SKILL.md +77 -0
  78. package/presets/clearai/template/skills/exploratory-data-analysis/references/bioinformatics_genomics_formats.md +664 -0
  79. package/presets/clearai/template/skills/exploratory-data-analysis/references/chemistry_molecular_formats.md +664 -0
  80. package/presets/clearai/template/skills/exploratory-data-analysis/references/general_scientific_formats.md +518 -0
  81. package/presets/clearai/template/skills/exploratory-data-analysis/references/microscopy_imaging_formats.md +620 -0
  82. package/presets/clearai/template/skills/exploratory-data-analysis/references/proteomics_metabolomics_formats.md +517 -0
  83. package/presets/clearai/template/skills/exploratory-data-analysis/references/spectroscopy_analytical_formats.md +633 -0
  84. package/presets/clearai/template/skills/exploratory-data-analysis/scripts/eda_analyzer.py +547 -0
  85. package/presets/clearai/template/skills/hypothesis-generation/SKILL.md +73 -0
  86. package/presets/clearai/template/skills/hypothesis-generation/references/experimental_design_patterns.md +329 -0
  87. package/presets/clearai/template/skills/hypothesis-generation/references/hypothesis_quality_criteria.md +198 -0
  88. package/presets/clearai/template/skills/hypothesis-generation/references/literature_search_strategies.md +622 -0
  89. package/presets/clearai/template/skills/hypothesis-generation/scripts/generate_schematic.py +139 -0
  90. package/presets/clearai/template/skills/hypothesis-generation/scripts/generate_schematic_ai.py +817 -0
  91. package/presets/clearai/template/skills/literature-review/SKILL.md +72 -0
  92. package/presets/clearai/template/skills/literature-review/references/citation_styles.md +166 -0
  93. package/presets/clearai/template/skills/literature-review/references/database_strategies.md +455 -0
  94. package/presets/clearai/template/skills/literature-review/scripts/generate_pdf.py +176 -0
  95. package/presets/clearai/template/skills/literature-review/scripts/generate_schematic.py +139 -0
  96. package/presets/clearai/template/skills/literature-review/scripts/generate_schematic_ai.py +817 -0
  97. package/presets/clearai/template/skills/literature-review/scripts/search_databases.py +303 -0
  98. package/presets/clearai/template/skills/literature-review/scripts/verify_citations.py +221 -0
  99. package/presets/clearai/template/skills/paper-lookup/SKILL.md +59 -0
  100. package/presets/clearai/template/skills/paper-lookup/references/arxiv.md +161 -0
  101. package/presets/clearai/template/skills/paper-lookup/references/biorxiv.md +118 -0
  102. package/presets/clearai/template/skills/paper-lookup/references/core.md +150 -0
  103. package/presets/clearai/template/skills/paper-lookup/references/crossref.md +181 -0
  104. package/presets/clearai/template/skills/paper-lookup/references/medrxiv.md +104 -0
  105. package/presets/clearai/template/skills/paper-lookup/references/openalex.md +174 -0
  106. package/presets/clearai/template/skills/paper-lookup/references/pmc.md +152 -0
  107. package/presets/clearai/template/skills/paper-lookup/references/pubmed.md +124 -0
  108. package/presets/clearai/template/skills/paper-lookup/references/semantic-scholar.md +203 -0
  109. package/presets/clearai/template/skills/paper-lookup/references/unpaywall.md +127 -0
  110. package/presets/clearai/template/skills/process-presearch/SKILL.md +196 -0
  111. package/presets/clearai/template/skills/process-presearch/checklists/process_checklist.md +18 -0
  112. package/presets/clearai/template/skills/process-presearch/references/figure_code.md +107 -0
  113. package/presets/clearai/template/skills/process-presearch/references/source_attribution_example.md +22 -0
  114. package/presets/clearai/template/skills/process-understanding-extraction/SKILL.md +69 -0
  115. package/presets/clearai/template/skills/process-understanding-extraction/checklists/readiness_check.md +34 -0
  116. package/presets/clearai/template/skills/process-understanding-extraction/templates/docx_raw_dump_extractor.py.tpl +132 -0
  117. package/presets/clearai/template/skills/process-understanding-extraction/templates/entity_map_unit_topology.json.tpl +86 -0
  118. package/presets/clearai/template/skills/process-understanding-extraction/templates/process_brief.md.tpl +89 -0
  119. package/presets/clearai/template/skills/process-understanding-extraction/templates/process_brief_builder_from_raw_dump.py.tpl +203 -0
  120. package/presets/clearai/template/skills/process-understanding-extraction/templates/process_flow_mermaid.md.tpl +41 -0
  121. package/presets/clearai/template/skills/process-understanding-extraction/templates/unified_execution_plan.md.tpl +53 -0
  122. package/presets/clearai/template/skills/process-understanding-extraction/workflows/01-process-doc-discovery.md +173 -0
  123. package/presets/clearai/template/skills/process-understanding-extraction/workflows/02-process-understanding-and-diagramming.md +106 -0
  124. package/presets/clearai/template/skills/scientific-brainstorming/SKILL.md +64 -0
  125. package/presets/clearai/template/skills/scientific-brainstorming/references/brainstorming_methods.md +326 -0
  126. package/presets/clearai/template/skills/scientific-critical-thinking/SKILL.md +72 -0
  127. package/presets/clearai/template/skills/scientific-critical-thinking/references/common_biases.md +364 -0
  128. package/presets/clearai/template/skills/scientific-critical-thinking/references/evidence_hierarchy.md +485 -0
  129. package/presets/clearai/template/skills/scientific-critical-thinking/references/experimental_design.md +496 -0
  130. package/presets/clearai/template/skills/scientific-critical-thinking/references/logical_fallacies.md +478 -0
  131. package/presets/clearai/template/skills/scientific-critical-thinking/references/scientific_method.md +169 -0
  132. package/presets/clearai/template/skills/scientific-critical-thinking/references/statistical_pitfalls.md +506 -0
  133. package/presets/clearai/template/skills/skill-creator/SKILL.md +109 -0
  134. package/presets/clearai/template/skills/skill-creator/references/authoring-guide.md +89 -0
  135. package/presets/clearai/template/skills/statistical-analysis/SKILL.md +79 -0
  136. package/presets/clearai/template/skills/statistical-analysis/references/assumptions_and_diagnostics.md +369 -0
  137. package/presets/clearai/template/skills/statistical-analysis/references/bayesian_statistics.md +653 -0
  138. package/presets/clearai/template/skills/statistical-analysis/references/effect_sizes_and_power.md +578 -0
  139. package/presets/clearai/template/skills/statistical-analysis/references/reporting_standards.md +469 -0
  140. package/presets/clearai/template/skills/statistical-analysis/references/test_selection_guide.md +129 -0
  141. package/presets/clearai/template/skills/statistical-analysis/scripts/assumption_checks.py +538 -0
  142. package/presets/clearai/template/skills/web-artifact/SKILL.md +165 -0
  143. package/presets/clearai/template/skills/web-artifact/assets/renderer/renderer.css +229 -0
  144. package/presets/clearai/template/skills/web-artifact/assets/renderer/renderer.js +373 -0
  145. package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/LICENSE +263 -0
  146. package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/UPSTREAM.md +26 -0
  147. package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/elk.bundled.js +6605 -0
  148. package/presets/clearai/template/skills/web-artifact/references/when-drawing-a-topology.md +150 -0
  149. package/presets/clearai/template/skills/web-artifact/references/when-the-page-must-work-offline.md +62 -0
  150. package/presets/clearai/template/skills/web-artifact/scripts/check_artifact.py +167 -0
  151. package/presets/clearai/template/skills/web-artifact/scripts/render_topology.js +272 -0
  152. package/presets/clearai/template/skills/what-if-oracle/LICENSE.txt +5 -0
  153. package/presets/clearai/template/skills/what-if-oracle/SKILL.md +72 -0
  154. package/presets/clearai/template/skills/what-if-oracle/references/scenario-templates.md +154 -0
@@ -0,0 +1,653 @@
1
+ # Bayesian Statistical Analysis
2
+
3
+ This document provides guidance on conducting and interpreting Bayesian statistical analyses, which offer an alternative framework to frequentist (classical) statistics.
4
+
5
+ ## Bayesian vs. Frequentist Philosophy
6
+
7
+ ### Fundamental Differences
8
+
9
+ | Aspect | Frequentist | Bayesian |
10
+ |--------|-------------|----------|
11
+ | **Probability interpretation** | Long-run frequency of events | Degree of belief/uncertainty |
12
+ | **Parameters** | Fixed but unknown | Random variables with distributions |
13
+ | **Inference** | Based on sampling distributions | Based on posterior distributions |
14
+ | **Primary output** | p-values, confidence intervals | Posterior probabilities, credible intervals |
15
+ | **Prior information** | Not formally incorporated | Explicitly incorporated via priors |
16
+ | **Hypothesis testing** | Reject/fail to reject null | Probability of hypotheses given data |
17
+ | **Sample size** | Often requires minimum | Can work with any sample size |
18
+ | **Interpretation** | Indirect (probability of data given H₀) | Direct (probability of hypothesis given data) |
19
+
20
+ ### Key Question Difference
21
+
22
+ **Frequentist**: "If the null hypothesis is true, what is the probability of observing data this extreme or more extreme?"
23
+
24
+ **Bayesian**: "Given the observed data, what is the probability that the hypothesis is true?"
25
+
26
+ The Bayesian question is more intuitive and directly addresses what researchers want to know.
27
+
28
+ ---
29
+
30
+ ## Bayes' Theorem
31
+
32
+ **Formula**:
33
+ ```
34
+ P(θ|D) = P(D|θ) × P(θ) / P(D)
35
+ ```
36
+
37
+ **In words**:
38
+ ```
39
+ Posterior = Likelihood × Prior / Evidence
40
+ ```
41
+
42
+ Where:
43
+ - **θ (theta)**: Parameter of interest (e.g., mean difference, correlation)
44
+ - **D**: Observed data
45
+ - **P(θ|D)**: Posterior distribution (belief about θ after seeing data)
46
+ - **P(D|θ)**: Likelihood (probability of data given θ)
47
+ - **P(θ)**: Prior distribution (belief about θ before seeing data)
48
+ - **P(D)**: Marginal likelihood/evidence (normalizing constant)
49
+
50
+ ---
51
+
52
+ ## Prior Distributions
53
+
54
+ ### Types of Priors
55
+
56
+ #### 1. Informative Priors
57
+
58
+ **When to use**: When you have substantial prior knowledge from:
59
+ - Previous studies
60
+ - Expert knowledge
61
+ - Theory
62
+ - Pilot data
63
+
64
+ **Example**: Meta-analysis shows effect size d ≈ 0.40, SD = 0.15
65
+ - Prior: Normal(0.40, 0.15)
66
+
67
+ **Advantages**:
68
+ - Incorporates existing knowledge
69
+ - More efficient (smaller samples needed)
70
+ - Can stabilize estimates with small data
71
+
72
+ **Disadvantages**:
73
+ - Subjective (but subjectivity can be strength)
74
+ - Must be justified and transparent
75
+ - May be controversial if strong prior conflicts with data
76
+
77
+ ---
78
+
79
+ #### 2. Weakly Informative Priors
80
+
81
+ **When to use**: Default choice for most applications
82
+
83
+ **Characteristics**:
84
+ - Regularizes estimates (prevents extreme values)
85
+ - Has minimal influence on posterior with moderate data
86
+ - Prevents computational issues
87
+
88
+ **Example priors**:
89
+ - Effect size: Normal(0, 1) or Cauchy(0, 0.707)
90
+ - Variance: Half-Cauchy(0, 1)
91
+ - Correlation: Uniform(-1, 1) or Beta(2, 2)
92
+
93
+ **Advantages**:
94
+ - Balances objectivity and regularization
95
+ - Computationally stable
96
+ - Broadly acceptable
97
+
98
+ ---
99
+
100
+ #### 3. Non-Informative (Flat/Uniform) Priors
101
+
102
+ **When to use**: When attempting to be "objective"
103
+
104
+ **Example**: Uniform(-∞, ∞) for any value
105
+
106
+ **⚠️ Caution**:
107
+ - Can lead to improper posteriors
108
+ - May produce non-sensible results
109
+ - Not truly "non-informative" (still makes assumptions)
110
+ - Often not recommended in modern Bayesian practice
111
+
112
+ **Better alternative**: Use weakly informative priors
113
+
114
+ ---
115
+
116
+ ### Prior Sensitivity Analysis
117
+
118
+ **Always conduct**: Test how results change with different priors
119
+
120
+ **Process**:
121
+ 1. Fit model with default/planned prior
122
+ 2. Fit model with more diffuse prior
123
+ 3. Fit model with more concentrated prior
124
+ 4. Compare posterior distributions
125
+
126
+ **Reporting**:
127
+ - If results are similar: Evidence is robust
128
+ - If results differ substantially: Data are not strong enough to overwhelm prior
129
+
130
+ **Python example**:
131
+ ```python
132
+ import pymc as pm
133
+
134
+ prior_specs = [
135
+ ('weakly_informative', 0, 1),
136
+ ('diffuse', 0, 10),
137
+ ('informative', 0.5, 0.3),
138
+ ]
139
+
140
+ results = {}
141
+ for name, mu_prior, sigma_prior in prior_specs:
142
+ with pm.Model() as model:
143
+ effect = pm.Normal('effect', mu=mu_prior, sigma=sigma_prior)
144
+ # ... likelihood and observed data
145
+ trace = pm.sample(2000, tune=1000, return_inferencedata=True)
146
+ results[name] = trace
147
+ ```
148
+
149
+ ---
150
+
151
+ ## Bayesian Hypothesis Testing
152
+
153
+ ### Bayes Factor (BF)
154
+
155
+ **What it is**: Ratio of evidence for two competing hypotheses
156
+
157
+ **Formula**:
158
+ ```
159
+ BF₁₀ = P(D|H₁) / P(D|H₀)
160
+ ```
161
+
162
+ **Interpretation**:
163
+
164
+ | BF₁₀ | Evidence |
165
+ |------|----------|
166
+ | >100 | Decisive for H₁ |
167
+ | 30-100 | Very strong for H₁ |
168
+ | 10-30 | Strong for H₁ |
169
+ | 3-10 | Moderate for H₁ |
170
+ | 1-3 | Anecdotal for H₁ |
171
+ | 1 | No evidence |
172
+ | 1/3-1 | Anecdotal for H₀ |
173
+ | 1/10-1/3 | Moderate for H₀ |
174
+ | 1/30-1/10 | Strong for H₀ |
175
+ | 1/100-1/30 | Very strong for H₀ |
176
+ | <1/100 | Decisive for H₀ |
177
+
178
+ **Advantages over p-values**:
179
+ 1. Can provide evidence for null hypothesis
180
+ 2. Not dependent on sampling intentions (no "peeking" problem)
181
+ 3. Directly quantifies evidence
182
+ 4. Can be updated with more data
183
+
184
+ **Python calculation**:
185
+ ```python
186
+ # Pingouin 0.5+: BF10 for independent two-sided t-tests; one-sided BF removed.
187
+ import pingouin as pg
188
+
189
+ result = pg.ttest(group1, group2, correction=False)
190
+ bf10 = result['BF10'].values[0]
191
+
192
+ # Rigorous Bayes Factors: BayesFactor (R), JASP, or PyMC model comparison (see pymc skill)
193
+ ```
194
+
195
+ ---
196
+
197
+ ### Region of Practical Equivalence (ROPE)
198
+
199
+ **Purpose**: Define range of negligible effect sizes
200
+
201
+ **Process**:
202
+ 1. Define ROPE (e.g., d ∈ [-0.1, 0.1] for negligible effects)
203
+ 2. Calculate % of posterior inside ROPE
204
+ 3. Make decision:
205
+ - >95% in ROPE: Accept practical equivalence
206
+ - >95% outside ROPE: Reject equivalence
207
+ - Otherwise: Inconclusive
208
+
209
+ **Advantage**: Directly tests for practical significance
210
+
211
+ **Python example**:
212
+ ```python
213
+ # Define ROPE
214
+ rope_lower, rope_upper = -0.1, 0.1
215
+
216
+ # Calculate % of posterior in ROPE
217
+ in_rope = np.mean((posterior_samples > rope_lower) &
218
+ (posterior_samples < rope_upper))
219
+
220
+ print(f"{in_rope*100:.1f}% of posterior in ROPE")
221
+ ```
222
+
223
+ ---
224
+
225
+ ## Bayesian Estimation
226
+
227
+ ### Credible Intervals
228
+
229
+ **What it is**: Interval containing parameter with X% probability
230
+
231
+ **95% Credible Interval interpretation**:
232
+ > "There is a 95% probability that the true parameter lies in this interval."
233
+
234
+ **This is what people THINK confidence intervals mean** (but don't in frequentist framework)
235
+
236
+ **Types**:
237
+
238
+ #### Equal-Tailed Interval (ETI)
239
+ - 2.5th to 97.5th percentile
240
+ - Simple to calculate
241
+ - May not include mode for skewed distributions
242
+
243
+ #### Highest Density Interval (HDI)
244
+ - Narrowest interval containing 95% of distribution
245
+ - Always includes mode
246
+ - Better for skewed distributions
247
+
248
+ **Python calculation**:
249
+ ```python
250
+ import arviz as az
251
+
252
+ # Equal-tailed interval
253
+ eti = np.percentile(posterior_samples, [2.5, 97.5])
254
+
255
+ # HDI
256
+ hdi = az.hdi(posterior_samples, hdi_prob=0.95)
257
+ ```
258
+
259
+ ---
260
+
261
+ ### Posterior Distributions
262
+
263
+ **Interpreting posterior distributions**:
264
+
265
+ 1. **Central tendency**:
266
+ - Mean: Average posterior value
267
+ - Median: 50th percentile
268
+ - Mode: Most probable value (MAP - Maximum A Posteriori)
269
+
270
+ 2. **Uncertainty**:
271
+ - SD: Spread of posterior
272
+ - Credible intervals: Quantify uncertainty
273
+
274
+ 3. **Shape**:
275
+ - Symmetric: Similar to normal
276
+ - Skewed: Asymmetric uncertainty
277
+ - Multimodal: Multiple plausible values
278
+
279
+ **Visualization**:
280
+ ```python
281
+ import matplotlib.pyplot as plt
282
+ import arviz as az
283
+
284
+ # Posterior plot with HDI
285
+ az.plot_posterior(trace, hdi_prob=0.95)
286
+
287
+ # Trace plot (check convergence)
288
+ az.plot_trace(trace)
289
+
290
+ # Forest plot (multiple parameters)
291
+ az.plot_forest(trace)
292
+ ```
293
+
294
+ ---
295
+
296
+ ## Common Bayesian Analyses
297
+
298
+ ### Bayesian T-Test
299
+
300
+ **Purpose**: Compare two groups (Bayesian alternative to t-test)
301
+
302
+ **Outputs**:
303
+ 1. Posterior distribution of mean difference
304
+ 2. 95% credible interval
305
+ 3. Bayes Factor (BF₁₀)
306
+ 4. Probability of directional hypothesis (e.g., P(μ₁ > μ₂))
307
+
308
+ **Python implementation**:
309
+ ```python
310
+ import pymc as pm
311
+ import arviz as az
312
+
313
+ # Bayesian independent samples t-test
314
+ with pm.Model() as model:
315
+ # Priors for group means
316
+ mu1 = pm.Normal('mu1', mu=0, sigma=10)
317
+ mu2 = pm.Normal('mu2', mu=0, sigma=10)
318
+
319
+ # Prior for pooled standard deviation
320
+ sigma = pm.HalfNormal('sigma', sigma=10)
321
+
322
+ # Likelihood
323
+ y1 = pm.Normal('y1', mu=mu1, sigma=sigma, observed=group1)
324
+ y2 = pm.Normal('y2', mu=mu2, sigma=sigma, observed=group2)
325
+
326
+ # Derived quantity: mean difference
327
+ diff = pm.Deterministic('diff', mu1 - mu2)
328
+
329
+ # Sample posterior
330
+ trace = pm.sample(2000, tune=1000, return_inferencedata=True)
331
+
332
+ # Analyze results
333
+ print(az.summary(trace, var_names=['mu1', 'mu2', 'diff']))
334
+
335
+ # Probability that group1 > group2
336
+ prob_greater = np.mean(trace.posterior['diff'].values > 0)
337
+ print(f"P(μ₁ > μ₂) = {prob_greater:.3f}")
338
+
339
+ # Plot posterior
340
+ az.plot_posterior(trace, var_names=['diff'], ref_val=0)
341
+ ```
342
+
343
+ ---
344
+
345
+ ### Bayesian ANOVA
346
+
347
+ **Purpose**: Compare three or more groups
348
+
349
+ **Model**:
350
+ ```python
351
+ import pymc as pm
352
+
353
+ with pm.Model() as anova_model:
354
+ # Hyperpriors
355
+ mu_global = pm.Normal('mu_global', mu=0, sigma=10)
356
+ sigma_between = pm.HalfNormal('sigma_between', sigma=5)
357
+ sigma_within = pm.HalfNormal('sigma_within', sigma=5)
358
+
359
+ # Group means (hierarchical)
360
+ group_means = pm.Normal('group_means',
361
+ mu=mu_global,
362
+ sigma=sigma_between,
363
+ shape=n_groups)
364
+
365
+ # Likelihood
366
+ y = pm.Normal('y',
367
+ mu=group_means[group_idx],
368
+ sigma=sigma_within,
369
+ observed=data)
370
+
371
+ trace = pm.sample(2000, tune=1000, return_inferencedata=True)
372
+
373
+ # Posterior contrasts
374
+ contrast_1_2 = trace.posterior['group_means'][:,:,0] - trace.posterior['group_means'][:,:,1]
375
+ ```
376
+
377
+ ---
378
+
379
+ ### Bayesian Correlation
380
+
381
+ **Purpose**: Estimate correlation between two variables
382
+
383
+ **Advantage**: Provides distribution of correlation values
384
+
385
+ **Python implementation**:
386
+ ```python
387
+ import pymc as pm
388
+
389
+ with pm.Model() as corr_model:
390
+ # Prior on correlation
391
+ rho = pm.Uniform('rho', lower=-1, upper=1)
392
+
393
+ # Convert to covariance matrix
394
+ cov_matrix = pm.math.stack([[1, rho],
395
+ [rho, 1]])
396
+
397
+ # Likelihood (bivariate normal)
398
+ obs = pm.MvNormal('obs',
399
+ mu=[0, 0],
400
+ cov=cov_matrix,
401
+ observed=np.column_stack([x, y]))
402
+
403
+ trace = pm.sample(2000, tune=1000, return_inferencedata=True)
404
+
405
+ # Summarize correlation
406
+ print(az.summary(trace, var_names=['rho']))
407
+
408
+ # Probability that correlation is positive
409
+ prob_positive = np.mean(trace.posterior['rho'].values > 0)
410
+ ```
411
+
412
+ ---
413
+
414
+ ### Bayesian Linear Regression
415
+
416
+ **Purpose**: Model relationship between predictors and outcome
417
+
418
+ **Advantages**:
419
+ - Uncertainty in all parameters
420
+ - Natural regularization (via priors)
421
+ - Can incorporate prior knowledge
422
+ - Credible intervals for predictions
423
+
424
+ **Python implementation**:
425
+ ```python
426
+ import pymc as pm
427
+
428
+ with pm.Model() as regression_model:
429
+ # Priors for coefficients
430
+ alpha = pm.Normal('alpha', mu=0, sigma=10) # Intercept
431
+ beta = pm.Normal('beta', mu=0, sigma=10, shape=n_predictors)
432
+ sigma = pm.HalfNormal('sigma', sigma=10)
433
+
434
+ # Expected value
435
+ mu = alpha + pm.math.dot(X, beta)
436
+
437
+ # Likelihood
438
+ y_obs = pm.Normal('y_obs', mu=mu, sigma=sigma, observed=y)
439
+
440
+ trace = pm.sample(2000, tune=1000, return_inferencedata=True)
441
+
442
+ # Posterior predictive checks
443
+ with regression_model:
444
+ ppc = pm.sample_posterior_predictive(trace)
445
+
446
+ az.plot_ppc(ppc)
447
+
448
+ # Predictions with uncertainty
449
+ with regression_model:
450
+ pm.set_data({'X': X_new})
451
+ posterior_pred = pm.sample_posterior_predictive(trace)
452
+ ```
453
+
454
+ ---
455
+
456
+ ## Hierarchical (Multilevel) Models
457
+
458
+ **When to use**:
459
+ - Nested/clustered data (students within schools)
460
+ - Repeated measures
461
+ - Meta-analysis
462
+ - Varying effects across groups
463
+
464
+ **Key concept**: Partial pooling
465
+ - Complete pooling: Ignore groups (biased)
466
+ - No pooling: Analyze groups separately (high variance)
467
+ - Partial pooling: Borrow strength across groups (Bayesian)
468
+
469
+ **Example: Varying intercepts**:
470
+ ```python
471
+ with pm.Model() as hierarchical_model:
472
+ # Hyperpriors
473
+ mu_global = pm.Normal('mu_global', mu=0, sigma=10)
474
+ sigma_between = pm.HalfNormal('sigma_between', sigma=5)
475
+ sigma_within = pm.HalfNormal('sigma_within', sigma=5)
476
+
477
+ # Group-level intercepts
478
+ alpha = pm.Normal('alpha',
479
+ mu=mu_global,
480
+ sigma=sigma_between,
481
+ shape=n_groups)
482
+
483
+ # Likelihood
484
+ y_obs = pm.Normal('y_obs',
485
+ mu=alpha[group_idx],
486
+ sigma=sigma_within,
487
+ observed=y)
488
+
489
+ trace = pm.sample()
490
+ ```
491
+
492
+ ---
493
+
494
+ ## Model Comparison
495
+
496
+ ### Methods
497
+
498
+ #### 1. Bayes Factor
499
+ - Directly compares model evidence
500
+ - Sensitive to prior specification
501
+ - Can be computationally intensive
502
+
503
+ #### 2. Information Criteria
504
+
505
+ **WAIC (Widely Applicable Information Criterion)**:
506
+ - Bayesian analog of AIC
507
+ - Lower is better
508
+ - Accounts for effective number of parameters
509
+
510
+ **LOO (Leave-One-Out Cross-Validation)**:
511
+ - Estimates out-of-sample prediction error
512
+ - Lower is better
513
+ - More robust than WAIC
514
+
515
+ **Python calculation**:
516
+ ```python
517
+ import arviz as az
518
+
519
+ # Calculate WAIC and LOO
520
+ waic = az.waic(trace)
521
+ loo = az.loo(trace)
522
+
523
+ print(f"WAIC: {waic.elpd_waic:.2f}")
524
+ print(f"LOO: {loo.elpd_loo:.2f}")
525
+
526
+ # Compare multiple models
527
+ comparison = az.compare({
528
+ 'model1': trace1,
529
+ 'model2': trace2,
530
+ 'model3': trace3
531
+ })
532
+ print(comparison)
533
+ ```
534
+
535
+ ---
536
+
537
+ ## Checking Bayesian Models
538
+
539
+ ### 1. Convergence Diagnostics
540
+
541
+ **R-hat (Gelman-Rubin statistic)**:
542
+ - Compares within-chain and between-chain variance
543
+ - Values close to 1.0 indicate convergence
544
+ - R-hat < 1.01: Good
545
+ - R-hat > 1.05: Poor convergence
546
+
547
+ **Effective Sample Size (ESS)**:
548
+ - Number of independent samples
549
+ - Higher is better
550
+ - ESS > 400 per chain recommended
551
+
552
+ **Trace plots**:
553
+ - Should look like "fuzzy caterpillar"
554
+ - No trends, no stuck chains
555
+
556
+ **Python checking**:
557
+ ```python
558
+ # Automatic summary with diagnostics
559
+ print(az.summary(trace, var_names=['parameter']))
560
+
561
+ # Visual diagnostics
562
+ az.plot_trace(trace)
563
+ az.plot_rank(trace) # Rank plots
564
+ ```
565
+
566
+ ---
567
+
568
+ ### 2. Posterior Predictive Checks
569
+
570
+ **Purpose**: Does model generate data similar to observed data?
571
+
572
+ **Process**:
573
+ 1. Generate predictions from posterior
574
+ 2. Compare to actual data
575
+ 3. Look for systematic discrepancies
576
+
577
+ **Python implementation**:
578
+ ```python
579
+ with model:
580
+ ppc = pm.sample_posterior_predictive(trace)
581
+
582
+ # Visual check
583
+ az.plot_ppc(ppc, num_pp_samples=100)
584
+
585
+ # Quantitative checks
586
+ obs_mean = np.mean(observed_data)
587
+ pred_means = [np.mean(sample) for sample in ppc.posterior_predictive['y_obs']]
588
+ p_value = np.mean(pred_means >= obs_mean) # Bayesian p-value
589
+ ```
590
+
591
+ ---
592
+
593
+ ## Reporting Bayesian Results
594
+
595
+ ### Example T-Test Report
596
+
597
+ > "A Bayesian independent samples t-test was conducted to compare groups A and B. Weakly informative priors were used: Normal(0, 1) for the mean difference and Half-Cauchy(0, 1) for the pooled standard deviation. The posterior distribution of the mean difference had a mean of 5.2 (95% CI [2.3, 8.1]), indicating that Group A scored higher than Group B. The Bayes Factor BF₁₀ = 23.5 provided strong evidence for a difference between groups, and there was a 99.7% probability that Group A's mean exceeded Group B's mean."
598
+
599
+ ### Example Regression Report
600
+
601
+ > "A Bayesian linear regression was fitted with weakly informative priors (Normal(0, 10) for coefficients, Half-Cauchy(0, 5) for residual SD). The model explained substantial variance (R² = 0.47, 95% CI [0.38, 0.55]). Study hours (β = 0.52, 95% CI [0.38, 0.66]) and prior GPA (β = 0.31, 95% CI [0.17, 0.45]) were credible predictors (95% CIs excluded zero). Posterior predictive checks showed good model fit. Convergence diagnostics were satisfactory (all R-hat < 1.01, ESS > 1000)."
602
+
603
+ ---
604
+
605
+ ## Advantages and Limitations
606
+
607
+ ### Advantages
608
+
609
+ 1. **Intuitive interpretation**: Direct probability statements about parameters
610
+ 2. **Incorporates prior knowledge**: Uses all available information
611
+ 3. **Flexible**: Handles complex models easily
612
+ 4. **No p-hacking**: Can look at data as it arrives
613
+ 5. **Quantifies uncertainty**: Full posterior distribution
614
+ 6. **Small samples**: Works with any sample size
615
+
616
+ ### Limitations
617
+
618
+ 1. **Computational**: Requires MCMC sampling (can be slow)
619
+ 2. **Prior specification**: Requires thought and justification
620
+ 3. **Complexity**: Steeper learning curve
621
+ 4. **Software**: Fewer tools than frequentist methods
622
+ 5. **Communication**: May need to educate reviewers/readers
623
+
624
+ ---
625
+
626
+ ## Key Python Packages
627
+
628
+ ArviZ/Bambi 属于可选 Bayesian 能力,不在基础统计栈的保证范围。使用前先做 import
629
+ 探测;若当前部署缺失,记录能力缺口并停止该分支,禁止在 Agent 运行中安装。ArviZ
630
+ 0.23+ requires Python 3.12+.
631
+
632
+ - **PyMC** (`pymc>=5`): Full Bayesian modeling framework
633
+ - **ArviZ** (`arviz>=0.17`): Visualization and diagnostics ([docs](https://python.arviz.org))
634
+ - **Bambi**: High-level interface for regression models(仅在部署已提供该能力时使用)
635
+ - **PyStan**: Python interface to Stan
636
+ - **TensorFlow Probability**: Bayesian inference with TensorFlow
637
+
638
+ ---
639
+
640
+ ## When to Use Bayesian Methods
641
+
642
+ **Use Bayesian when**:
643
+ - You have prior information to incorporate
644
+ - You want direct probability statements
645
+ - Sample size is small
646
+ - Model is complex (hierarchical, missing data, etc.)
647
+ - You want to update analysis as data arrives
648
+
649
+ **Frequentist may be sufficient when**:
650
+ - Standard analysis with large sample
651
+ - No prior information
652
+ - Computational resources limited
653
+ - Reviewers unfamiliar with Bayesian methods