clearai-dsh 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (154) hide show
  1. package/CHANGELOG.md +26 -0
  2. package/LICENSE +201 -0
  3. package/README.md +138 -0
  4. package/README.zh-CN.md +138 -0
  5. package/bin/clearai.mjs +224 -0
  6. package/brand/README.md +41 -0
  7. package/brand/logo-512-dark.png +0 -0
  8. package/brand/logo-512.png +0 -0
  9. package/brand/logo-lockup-dark.png +0 -0
  10. package/brand/logo-lockup.png +0 -0
  11. package/brand/logo-lockup.svg +12 -0
  12. package/brand/logo-wordmark.svg +6 -0
  13. package/brand/logo.svg +19 -0
  14. package/cordis.patch.yml +39 -0
  15. package/lib/client.js +3071 -0
  16. package/lib/fold.js +1576 -0
  17. package/lib/host.js +605 -0
  18. package/package.json +65 -0
  19. package/presets/clearai/agent.cordis.yml +226 -0
  20. package/presets/clearai/plugins/brain.js +547 -0
  21. package/presets/clearai/plugins/clearai-kernel.js +5485 -0
  22. package/presets/clearai/plugins/ontology.js +306 -0
  23. package/presets/clearai/plugins/prompts.js +312 -0
  24. package/presets/clearai/preset.yml +5 -0
  25. package/presets/clearai/skills/clearai-loop/SKILL.md +89 -0
  26. package/presets/clearai/template/knowledge/README.md +25 -0
  27. package/presets/clearai/template/memory/README.md +34 -0
  28. package/presets/clearai/template/project.md +49 -0
  29. package/presets/clearai/template/skills/README.md +37 -0
  30. package/presets/clearai/template/skills/chart-diagram-qa/SKILL.md +43 -0
  31. package/presets/clearai/template/skills/citation-management/SKILL.md +73 -0
  32. package/presets/clearai/template/skills/citation-management/references/bibtex_formatting.md +908 -0
  33. package/presets/clearai/template/skills/citation-management/references/citation_validation.md +794 -0
  34. package/presets/clearai/template/skills/citation-management/references/google_scholar_search.md +725 -0
  35. package/presets/clearai/template/skills/citation-management/references/metadata_extraction.md +870 -0
  36. package/presets/clearai/template/skills/citation-management/references/pubmed_search.md +839 -0
  37. package/presets/clearai/template/skills/citation-management/scripts/doi_to_bibtex.py +204 -0
  38. package/presets/clearai/template/skills/citation-management/scripts/extract_metadata.py +569 -0
  39. package/presets/clearai/template/skills/citation-management/scripts/format_bibtex.py +349 -0
  40. package/presets/clearai/template/skills/citation-management/scripts/generate_schematic.py +139 -0
  41. package/presets/clearai/template/skills/citation-management/scripts/generate_schematic_ai.py +817 -0
  42. package/presets/clearai/template/skills/citation-management/scripts/search_google_scholar.py +282 -0
  43. package/presets/clearai/template/skills/citation-management/scripts/search_pubmed.py +398 -0
  44. package/presets/clearai/template/skills/citation-management/scripts/validate_citations.py +497 -0
  45. package/presets/clearai/template/skills/data-analysis/SKILL.md +92 -0
  46. package/presets/clearai/template/skills/data-analysis/checklists/readiness_check.md +23 -0
  47. package/presets/clearai/template/skills/data-analysis/templates/analysis_report.md.tpl +63 -0
  48. package/presets/clearai/template/skills/data-analysis/templates/cleaning_rules_draft.yaml.tpl +32 -0
  49. package/presets/clearai/template/skills/data-analysis/templates/data_dictionary.md.tpl +12 -0
  50. package/presets/clearai/template/skills/data-analysis/templates/domain_knowledge_template.md.tpl +316 -0
  51. package/presets/clearai/template/skills/data-analysis/templates/feature_candidates.json.tpl +20 -0
  52. package/presets/clearai/template/skills/data-analysis/templates/quality_scorecard.md.tpl +30 -0
  53. package/presets/clearai/template/skills/data-analysis/workflows/01-data-profiling.md +42 -0
  54. package/presets/clearai/template/skills/data-analysis/workflows/02-quality-audit.md +36 -0
  55. package/presets/clearai/template/skills/data-analysis/workflows/03-physical-correlation.md +25 -0
  56. package/presets/clearai/template/skills/data-analysis/workflows/04-unstructured-mining.md +26 -0
  57. package/presets/clearai/template/skills/data-qa-analysis/SKILL.md +102 -0
  58. package/presets/clearai/template/skills/data-qa-analysis/checklists/readiness_check.md +62 -0
  59. package/presets/clearai/template/skills/data-qa-analysis/templates/best_in_class_report.md.tpl +56 -0
  60. package/presets/clearai/template/skills/data-qa-analysis/templates/cleaning_rules_draft.yaml.tpl +56 -0
  61. package/presets/clearai/template/skills/data-qa-analysis/templates/data_dictionary.md.tpl +13 -0
  62. package/presets/clearai/template/skills/data-qa-analysis/templates/data_source_inventory_and_lineage.md.tpl +146 -0
  63. package/presets/clearai/template/skills/data-qa-analysis/templates/data_status_report.md.tpl +60 -0
  64. package/presets/clearai/template/skills/data-qa-analysis/templates/steady_state_rules.yaml.tpl +41 -0
  65. package/presets/clearai/template/skills/data-qa-analysis/templates/subsystem_registry.md.tpl +101 -0
  66. package/presets/clearai/template/skills/data-qa-analysis/templates/unified_execution_plan.md.tpl +100 -0
  67. package/presets/clearai/template/skills/data-qa-analysis/workflows/01-data-source-inventory-and-lineage.md +194 -0
  68. package/presets/clearai/template/skills/data-qa-analysis/workflows/02-data-alignment-and-tag-semantics.md +122 -0
  69. package/presets/clearai/template/skills/data-qa-analysis/workflows/03-steady-state-identification.md +126 -0
  70. package/presets/clearai/template/skills/data-qa-analysis/workflows/04-consumption-analysis.md +152 -0
  71. package/presets/clearai/template/skills/data-qa-analysis/workflows/05-best-in-class-and-optimization-space.md +78 -0
  72. package/presets/clearai/template/skills/domain-presearch/SKILL.md +131 -0
  73. package/presets/clearai/template/skills/domain-presearch/checklists/domain_checklist.md +24 -0
  74. package/presets/clearai/template/skills/domain-presearch/references/figure_code.md +78 -0
  75. package/presets/clearai/template/skills/domain-presearch/references/strategic_frameworks.md +38 -0
  76. package/presets/clearai/template/skills/exploration-loop/SKILL.md +81 -0
  77. package/presets/clearai/template/skills/exploratory-data-analysis/SKILL.md +77 -0
  78. package/presets/clearai/template/skills/exploratory-data-analysis/references/bioinformatics_genomics_formats.md +664 -0
  79. package/presets/clearai/template/skills/exploratory-data-analysis/references/chemistry_molecular_formats.md +664 -0
  80. package/presets/clearai/template/skills/exploratory-data-analysis/references/general_scientific_formats.md +518 -0
  81. package/presets/clearai/template/skills/exploratory-data-analysis/references/microscopy_imaging_formats.md +620 -0
  82. package/presets/clearai/template/skills/exploratory-data-analysis/references/proteomics_metabolomics_formats.md +517 -0
  83. package/presets/clearai/template/skills/exploratory-data-analysis/references/spectroscopy_analytical_formats.md +633 -0
  84. package/presets/clearai/template/skills/exploratory-data-analysis/scripts/eda_analyzer.py +547 -0
  85. package/presets/clearai/template/skills/hypothesis-generation/SKILL.md +73 -0
  86. package/presets/clearai/template/skills/hypothesis-generation/references/experimental_design_patterns.md +329 -0
  87. package/presets/clearai/template/skills/hypothesis-generation/references/hypothesis_quality_criteria.md +198 -0
  88. package/presets/clearai/template/skills/hypothesis-generation/references/literature_search_strategies.md +622 -0
  89. package/presets/clearai/template/skills/hypothesis-generation/scripts/generate_schematic.py +139 -0
  90. package/presets/clearai/template/skills/hypothesis-generation/scripts/generate_schematic_ai.py +817 -0
  91. package/presets/clearai/template/skills/literature-review/SKILL.md +72 -0
  92. package/presets/clearai/template/skills/literature-review/references/citation_styles.md +166 -0
  93. package/presets/clearai/template/skills/literature-review/references/database_strategies.md +455 -0
  94. package/presets/clearai/template/skills/literature-review/scripts/generate_pdf.py +176 -0
  95. package/presets/clearai/template/skills/literature-review/scripts/generate_schematic.py +139 -0
  96. package/presets/clearai/template/skills/literature-review/scripts/generate_schematic_ai.py +817 -0
  97. package/presets/clearai/template/skills/literature-review/scripts/search_databases.py +303 -0
  98. package/presets/clearai/template/skills/literature-review/scripts/verify_citations.py +221 -0
  99. package/presets/clearai/template/skills/paper-lookup/SKILL.md +59 -0
  100. package/presets/clearai/template/skills/paper-lookup/references/arxiv.md +161 -0
  101. package/presets/clearai/template/skills/paper-lookup/references/biorxiv.md +118 -0
  102. package/presets/clearai/template/skills/paper-lookup/references/core.md +150 -0
  103. package/presets/clearai/template/skills/paper-lookup/references/crossref.md +181 -0
  104. package/presets/clearai/template/skills/paper-lookup/references/medrxiv.md +104 -0
  105. package/presets/clearai/template/skills/paper-lookup/references/openalex.md +174 -0
  106. package/presets/clearai/template/skills/paper-lookup/references/pmc.md +152 -0
  107. package/presets/clearai/template/skills/paper-lookup/references/pubmed.md +124 -0
  108. package/presets/clearai/template/skills/paper-lookup/references/semantic-scholar.md +203 -0
  109. package/presets/clearai/template/skills/paper-lookup/references/unpaywall.md +127 -0
  110. package/presets/clearai/template/skills/process-presearch/SKILL.md +196 -0
  111. package/presets/clearai/template/skills/process-presearch/checklists/process_checklist.md +18 -0
  112. package/presets/clearai/template/skills/process-presearch/references/figure_code.md +107 -0
  113. package/presets/clearai/template/skills/process-presearch/references/source_attribution_example.md +22 -0
  114. package/presets/clearai/template/skills/process-understanding-extraction/SKILL.md +69 -0
  115. package/presets/clearai/template/skills/process-understanding-extraction/checklists/readiness_check.md +34 -0
  116. package/presets/clearai/template/skills/process-understanding-extraction/templates/docx_raw_dump_extractor.py.tpl +132 -0
  117. package/presets/clearai/template/skills/process-understanding-extraction/templates/entity_map_unit_topology.json.tpl +86 -0
  118. package/presets/clearai/template/skills/process-understanding-extraction/templates/process_brief.md.tpl +89 -0
  119. package/presets/clearai/template/skills/process-understanding-extraction/templates/process_brief_builder_from_raw_dump.py.tpl +203 -0
  120. package/presets/clearai/template/skills/process-understanding-extraction/templates/process_flow_mermaid.md.tpl +41 -0
  121. package/presets/clearai/template/skills/process-understanding-extraction/templates/unified_execution_plan.md.tpl +53 -0
  122. package/presets/clearai/template/skills/process-understanding-extraction/workflows/01-process-doc-discovery.md +173 -0
  123. package/presets/clearai/template/skills/process-understanding-extraction/workflows/02-process-understanding-and-diagramming.md +106 -0
  124. package/presets/clearai/template/skills/scientific-brainstorming/SKILL.md +64 -0
  125. package/presets/clearai/template/skills/scientific-brainstorming/references/brainstorming_methods.md +326 -0
  126. package/presets/clearai/template/skills/scientific-critical-thinking/SKILL.md +72 -0
  127. package/presets/clearai/template/skills/scientific-critical-thinking/references/common_biases.md +364 -0
  128. package/presets/clearai/template/skills/scientific-critical-thinking/references/evidence_hierarchy.md +485 -0
  129. package/presets/clearai/template/skills/scientific-critical-thinking/references/experimental_design.md +496 -0
  130. package/presets/clearai/template/skills/scientific-critical-thinking/references/logical_fallacies.md +478 -0
  131. package/presets/clearai/template/skills/scientific-critical-thinking/references/scientific_method.md +169 -0
  132. package/presets/clearai/template/skills/scientific-critical-thinking/references/statistical_pitfalls.md +506 -0
  133. package/presets/clearai/template/skills/skill-creator/SKILL.md +109 -0
  134. package/presets/clearai/template/skills/skill-creator/references/authoring-guide.md +89 -0
  135. package/presets/clearai/template/skills/statistical-analysis/SKILL.md +79 -0
  136. package/presets/clearai/template/skills/statistical-analysis/references/assumptions_and_diagnostics.md +369 -0
  137. package/presets/clearai/template/skills/statistical-analysis/references/bayesian_statistics.md +653 -0
  138. package/presets/clearai/template/skills/statistical-analysis/references/effect_sizes_and_power.md +578 -0
  139. package/presets/clearai/template/skills/statistical-analysis/references/reporting_standards.md +469 -0
  140. package/presets/clearai/template/skills/statistical-analysis/references/test_selection_guide.md +129 -0
  141. package/presets/clearai/template/skills/statistical-analysis/scripts/assumption_checks.py +538 -0
  142. package/presets/clearai/template/skills/web-artifact/SKILL.md +165 -0
  143. package/presets/clearai/template/skills/web-artifact/assets/renderer/renderer.css +229 -0
  144. package/presets/clearai/template/skills/web-artifact/assets/renderer/renderer.js +373 -0
  145. package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/LICENSE +263 -0
  146. package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/UPSTREAM.md +26 -0
  147. package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/elk.bundled.js +6605 -0
  148. package/presets/clearai/template/skills/web-artifact/references/when-drawing-a-topology.md +150 -0
  149. package/presets/clearai/template/skills/web-artifact/references/when-the-page-must-work-offline.md +62 -0
  150. package/presets/clearai/template/skills/web-artifact/scripts/check_artifact.py +167 -0
  151. package/presets/clearai/template/skills/web-artifact/scripts/render_topology.js +272 -0
  152. package/presets/clearai/template/skills/what-if-oracle/LICENSE.txt +5 -0
  153. package/presets/clearai/template/skills/what-if-oracle/SKILL.md +72 -0
  154. package/presets/clearai/template/skills/what-if-oracle/references/scenario-templates.md +154 -0
@@ -0,0 +1,506 @@
1
+ # Common Statistical Pitfalls
2
+
3
+ ## P-Value Misinterpretations
4
+
5
+ ### Pitfall 1: P-Value = Probability Hypothesis is True
6
+ **Misconception:** p = .05 means 5% chance the null hypothesis is true.
7
+
8
+ **Reality:** P-value is the probability of observing data this extreme (or more) *if* the null hypothesis is true. It says nothing about the probability the hypothesis is true.
9
+
10
+ **Correct interpretation:** "If there were truly no effect, we would observe data this extreme only 5% of the time."
11
+
12
+ ### Pitfall 2: Non-Significant = No Effect
13
+ **Misconception:** p > .05 proves there's no effect.
14
+
15
+ **Reality:** Absence of evidence ≠ evidence of absence. Non-significant results may indicate:
16
+ - Insufficient statistical power
17
+ - True effect too small to detect
18
+ - High variability
19
+ - Small sample size
20
+
21
+ **Better approach:**
22
+ - Report confidence intervals
23
+ - Conduct power analysis
24
+ - Consider equivalence testing
25
+
26
+ ### Pitfall 3: Significant = Important
27
+ **Misconception:** Statistical significance means practical importance.
28
+
29
+ **Reality:** With large samples, trivial effects become "significant." A statistically significant 0.1 IQ point difference is meaningless in practice.
30
+
31
+ **Better approach:**
32
+ - Report effect sizes
33
+ - Consider practical significance
34
+ - Use confidence intervals
35
+
36
+ ### Pitfall 4: P = .049 vs. P = .051
37
+ **Misconception:** These are meaningfully different because one crosses the .05 threshold.
38
+
39
+ **Reality:** These represent nearly identical evidence. The .05 threshold is arbitrary.
40
+
41
+ **Better approach:**
42
+ - Treat p-values as continuous measures of evidence
43
+ - Report exact p-values
44
+ - Consider context and prior evidence
45
+
46
+ ### Pitfall 5: One-Tailed Tests Without Justification
47
+ **Misconception:** One-tailed tests are free extra power.
48
+
49
+ **Reality:** One-tailed tests assume effects can only go one direction, which is rarely true. They're often used to artificially boost significance.
50
+
51
+ **When appropriate:** Only when effects in one direction are theoretically impossible or equivalent to null.
52
+
53
+ ## Multiple Comparisons Problems
54
+
55
+ ### Pitfall 6: Multiple Testing Without Correction
56
+ **Problem:** Testing 20 hypotheses at p < .05 gives ~65% chance of at least one false positive.
57
+
58
+ **Examples:**
59
+ - Testing many outcomes
60
+ - Testing many subgroups
61
+ - Conducting multiple interim analyses
62
+ - Testing at multiple time points
63
+
64
+ **Solutions:**
65
+ - Bonferroni correction (divide α by number of tests)
66
+ - False Discovery Rate (FDR) control
67
+ - Prespecify primary outcome
68
+ - Treat exploratory analyses as hypothesis-generating
69
+
70
+ ### Pitfall 7: Subgroup Analysis Fishing
71
+ **Problem:** Testing many subgroups until finding significance.
72
+
73
+ **Why problematic:**
74
+ - Inflates false positive rate
75
+ - Often reported without disclosure
76
+ - "Interaction was significant in women" may be random
77
+
78
+ **Solutions:**
79
+ - Prespecify subgroups
80
+ - Use interaction tests, not separate tests
81
+ - Require replication
82
+ - Correct for multiple comparisons
83
+
84
+ ### Pitfall 8: Outcome Switching
85
+ **Problem:** Analyzing many outcomes, reporting only significant ones.
86
+
87
+ **Detection signs:**
88
+ - Secondary outcomes emphasized
89
+ - Incomplete outcome reporting
90
+ - Discrepancy between registration and publication
91
+
92
+ **Solutions:**
93
+ - Preregister all outcomes
94
+ - Report all planned outcomes
95
+ - Distinguish primary from secondary
96
+
97
+ ## Sample Size and Power Issues
98
+
99
+ ### Pitfall 9: Underpowered Studies
100
+ **Problem:** Small samples have low probability of detecting true effects.
101
+
102
+ **Consequences:**
103
+ - High false negative rate
104
+ - Significant results more likely to be false positives
105
+ - Overestimated effect sizes (when significant)
106
+
107
+ **Solutions:**
108
+ - Conduct a priori power analysis
109
+ - Aim for 80-90% power
110
+ - Consider effect size from prior research
111
+
112
+ ### Pitfall 10: Post-Hoc Power Analysis
113
+ **Problem:** Calculating power after seeing results is circular and uninformative.
114
+
115
+ **Why useless:**
116
+ - Non-significant results always have low "post-hoc power"
117
+ - It recapitulates the p-value without new information
118
+
119
+ **Better approach:**
120
+ - Calculate confidence intervals
121
+ - Plan replication with adequate sample
122
+ - Conduct prospective power analysis for future studies
123
+
124
+ ### Pitfall 11: Small Sample Fallacy
125
+ **Problem:** Trusting results from very small samples.
126
+
127
+ **Issues:**
128
+ - High sampling variability
129
+ - Outliers have large influence
130
+ - Assumptions of tests violated
131
+ - Confidence intervals very wide
132
+
133
+ **Guidelines:**
134
+ - Be skeptical of n < 30
135
+ - Check assumptions carefully
136
+ - Consider non-parametric tests
137
+ - Replicate findings
138
+
139
+ ## Effect Size Misunderstandings
140
+
141
+ ### Pitfall 12: Ignoring Effect Size
142
+ **Problem:** Focusing only on significance, not magnitude.
143
+
144
+ **Why problematic:**
145
+ - Significance ≠ importance
146
+ - Can't compare across studies
147
+ - Doesn't inform practical decisions
148
+
149
+ **Solutions:**
150
+ - Always report effect sizes
151
+ - Use standardized measures (Cohen's d, r, η²)
152
+ - Interpret using field conventions
153
+ - Consider minimum clinically important difference
154
+
155
+ ### Pitfall 13: Misinterpreting Standardized Effect Sizes
156
+ **Problem:** Treating Cohen's d = 0.5 as "medium" without context.
157
+
158
+ **Reality:**
159
+ - Field-specific norms vary
160
+ - Some fields have larger typical effects
161
+ - Real-world importance depends on context
162
+
163
+ **Better approach:**
164
+ - Compare to effects in same domain
165
+ - Consider practical implications
166
+ - Look at raw effect sizes too
167
+
168
+ ### Pitfall 14: Confusing Explained Variance with Importance
169
+ **Problem:** "Only explains 5% of variance" = unimportant.
170
+
171
+ **Reality:**
172
+ - Height explains ~5% of variation in NBA player salary but is crucial
173
+ - Complex phenomena have many small contributors
174
+ - Predictive accuracy ≠ causal importance
175
+
176
+ **Consideration:** Context matters more than percentage alone.
177
+
178
+ ## Correlation and Causation
179
+
180
+ ### Pitfall 15: Correlation Implies Causation
181
+ **Problem:** Inferring causation from correlation.
182
+
183
+ **Alternative explanations:**
184
+ - Reverse causation (B causes A, not A causes B)
185
+ - Confounding (C causes both A and B)
186
+ - Coincidence
187
+ - Selection bias
188
+
189
+ **Criteria for causation:**
190
+ - Temporal precedence
191
+ - Covariation
192
+ - No plausible alternatives
193
+ - Ideally: experimental manipulation
194
+
195
+ ### Pitfall 16: Ecological Fallacy
196
+ **Problem:** Inferring individual-level relationships from group-level data.
197
+
198
+ **Example:** Countries with more chocolate consumption have more Nobel laureates doesn't mean eating chocolate makes you win Nobels.
199
+
200
+ **Why problematic:** Group-level correlations may not hold at individual level.
201
+
202
+ ### Pitfall 17: Simpson's Paradox
203
+ **Problem:** Trend appears in groups but reverses when combined (or vice versa).
204
+
205
+ **Example:** Treatment appears worse overall but better in every subgroup.
206
+
207
+ **Cause:** Confounding variable distributed differently across groups.
208
+
209
+ **Solution:** Consider confounders and look at appropriate level of analysis.
210
+
211
+ ## Regression and Modeling Pitfalls
212
+
213
+ ### Pitfall 18: Overfitting
214
+ **Problem:** Model fits sample data well but doesn't generalize.
215
+
216
+ **Causes:**
217
+ - Too many predictors relative to sample size
218
+ - Fitting noise rather than signal
219
+ - No cross-validation
220
+
221
+ **Solutions:**
222
+ - Use cross-validation
223
+ - Penalized regression (LASSO, ridge)
224
+ - Independent test set
225
+ - Simpler models
226
+
227
+ ### Pitfall 19: Extrapolation Beyond Data Range
228
+ **Problem:** Predicting outside the range of observed data.
229
+
230
+ **Why dangerous:**
231
+ - Relationships may not hold outside observed range
232
+ - Increased uncertainty not reflected in predictions
233
+
234
+ **Solution:** Only interpolate; avoid extrapolation.
235
+
236
+ ### Pitfall 20: Ignoring Model Assumptions
237
+ **Problem:** Using statistical tests without checking assumptions.
238
+
239
+ **Common violations:**
240
+ - Non-normality (for parametric tests)
241
+ - Heteroscedasticity (unequal variances)
242
+ - Non-independence
243
+ - Linearity
244
+ - No multicollinearity
245
+
246
+ **Solutions:**
247
+ - Check assumptions with diagnostics
248
+ - Use robust methods
249
+ - Transform data
250
+ - Use appropriate non-parametric alternatives
251
+
252
+ ### Pitfall 21: Treating Non-Significant Covariates as Eliminating Confounding
253
+ **Problem:** "We controlled for X and it wasn't significant, so it's not a confounder."
254
+
255
+ **Reality:** Non-significant covariates can still be important confounders. Significance ≠ confounding.
256
+
257
+ **Solution:** Include theoretically important covariates regardless of significance.
258
+
259
+ ### Pitfall 22: Collinearity Masking Effects
260
+ **Problem:** When predictors are highly correlated, true effects may appear non-significant.
261
+
262
+ **Manifestations:**
263
+ - Large standard errors
264
+ - Unstable coefficients
265
+ - Sign changes when adding/removing variables
266
+
267
+ **Detection:**
268
+ - Variance Inflation Factors (VIF)
269
+ - Correlation matrices
270
+
271
+ **Solutions:**
272
+ - Remove redundant predictors
273
+ - Combine correlated variables
274
+ - Use regularization methods
275
+
276
+ ## Specific Test Misuses
277
+
278
+ ### Pitfall 23: T-Test for Multiple Groups
279
+ **Problem:** Conducting multiple t-tests instead of ANOVA.
280
+
281
+ **Why wrong:** Inflates Type I error rate dramatically.
282
+
283
+ **Correct approach:**
284
+ - Use ANOVA first
285
+ - Follow with planned comparisons or post-hoc tests with correction
286
+
287
+ ### Pitfall 24: Pearson Correlation for Non-Linear Relationships
288
+ **Problem:** Using Pearson's r for curved relationships.
289
+
290
+ **Why misleading:** r measures linear relationships only.
291
+
292
+ **Solutions:**
293
+ - Check scatterplots first
294
+ - Use Spearman's ρ for monotonic relationships
295
+ - Consider polynomial or non-linear models
296
+
297
+ ### Pitfall 25: Chi-Square with Small Expected Frequencies
298
+ **Problem:** Chi-square test with expected cell counts < 5.
299
+
300
+ **Why wrong:** Violates test assumptions, p-values inaccurate.
301
+
302
+ **Solutions:**
303
+ - Fisher's exact test
304
+ - Combine categories
305
+ - Increase sample size
306
+
307
+ ### Pitfall 26: Paired vs. Independent Tests
308
+ **Problem:** Using independent samples test for paired data (or vice versa).
309
+
310
+ **Why wrong:**
311
+ - Wastes power (paired data analyzed as independent)
312
+ - Violates independence assumption (independent data analyzed as paired)
313
+
314
+ **Solution:** Match test to design.
315
+
316
+ ## Confidence Interval Misinterpretations
317
+
318
+ ### Pitfall 27: 95% CI = 95% Probability True Value Inside
319
+ **Misconception:** "95% chance the true value is in this interval."
320
+
321
+ **Reality:** The true value either is or isn't in this specific interval. If we repeated the study many times, 95% of resulting intervals would contain the true value.
322
+
323
+ **Better interpretation:** "We're 95% confident this interval contains the true value."
324
+
325
+ ### Pitfall 28: Overlapping CIs = No Difference
326
+ **Problem:** Assuming overlapping confidence intervals mean no significant difference.
327
+
328
+ **Reality:** Overlapping CIs are less stringent than difference tests. Two CIs can overlap while the difference between groups is significant.
329
+
330
+ **Guideline:** Overlap of point estimate with other CI is more relevant than overlap of intervals.
331
+
332
+ ### Pitfall 29: Ignoring CI Width
333
+ **Problem:** Focusing only on whether CI includes zero, not precision.
334
+
335
+ **Why important:** Wide CIs indicate high uncertainty. "Significant" effects with huge CIs are less convincing.
336
+
337
+ **Consider:** Both significance and precision.
338
+
339
+ ## Bayesian vs. Frequentist Confusions
340
+
341
+ ### Pitfall 30: Mixing Bayesian and Frequentist Interpretations
342
+ **Problem:** Making Bayesian statements from frequentist analyses.
343
+
344
+ **Examples:**
345
+ - "Probability hypothesis is true" (Bayesian) from p-value (frequentist)
346
+ - "Evidence for null" from non-significant result (frequentist can't support null)
347
+
348
+ **Solution:**
349
+ - Be clear about framework
350
+ - Use Bayesian methods for Bayesian questions
351
+ - Use Bayes factors to compare hypotheses
352
+
353
+ ### Pitfall 31: Ignoring Prior Probability
354
+ **Problem:** Treating all hypotheses as equally likely initially.
355
+
356
+ **Reality:** Extraordinary claims need extraordinary evidence. Prior plausibility matters.
357
+
358
+ **Consider:**
359
+ - Plausibility given existing knowledge
360
+ - Mechanism plausibility
361
+ - Base rates
362
+
363
+ ## Data Transformation Issues
364
+
365
+ ### Pitfall 32: Dichotomizing Continuous Variables
366
+ **Problem:** Splitting continuous variables at arbitrary cutoffs.
367
+
368
+ **Consequences:**
369
+ - Loss of information and power
370
+ - Arbitrary distinctions
371
+ - Discarding individual differences
372
+
373
+ **Exceptions:** Clinically meaningful cutoffs with strong justification.
374
+
375
+ **Better:** Keep continuous or use multiple categories.
376
+
377
+ ### Pitfall 33: Trying Multiple Transformations
378
+ **Problem:** Testing many transformations until finding significance.
379
+
380
+ **Why problematic:** Inflates Type I error, is a form of p-hacking.
381
+
382
+ **Better approach:**
383
+ - Prespecify transformations
384
+ - Use theory-driven transformations
385
+ - Correct for multiple testing if exploring
386
+
387
+ ## Missing Data Problems
388
+
389
+ ### Pitfall 34: Listwise Deletion by Default
390
+ **Problem:** Automatically deleting all cases with any missing data.
391
+
392
+ **Consequences:**
393
+ - Reduced power
394
+ - Potential bias if data not missing completely at random (MCAR)
395
+
396
+ **Better approaches:**
397
+ - Multiple imputation
398
+ - Maximum likelihood methods
399
+ - Analyze missingness patterns
400
+
401
+ ### Pitfall 35: Ignoring Missing Data Mechanisms
402
+ **Problem:** Not considering why data are missing.
403
+
404
+ **Types:**
405
+ - MCAR (Missing Completely at Random): Safe to delete
406
+ - MAR (Missing at Random): Can impute
407
+ - MNAR (Missing Not at Random): May bias results
408
+
409
+ **Solution:** Analyze patterns, use appropriate methods, consider sensitivity analyses.
410
+
411
+ ## Publication and Reporting Issues
412
+
413
+ ### Pitfall 36: Selective Reporting
414
+ **Problem:** Only reporting significant results or favorable analyses.
415
+
416
+ **Consequences:**
417
+ - Literature appears more consistent than reality
418
+ - Meta-analyses biased
419
+ - Wasted research effort
420
+
421
+ **Solutions:**
422
+ - Preregistration
423
+ - Report all analyses
424
+ - Use reporting guidelines (CONSORT, PRISMA, etc.)
425
+
426
+ ### Pitfall 37: Rounding to p < .05
427
+ **Problem:** Reporting exact p-values selectively (e.g., p = .049 but p < .05 for .051).
428
+
429
+ **Why problematic:** Obscures values near threshold, enables p-hacking detection evasion.
430
+
431
+ **Better:** Always report exact p-values.
432
+
433
+ ### Pitfall 38: No Data Sharing
434
+ **Problem:** Not making data available for verification or reanalysis.
435
+
436
+ **Consequences:**
437
+ - Can't verify results
438
+ - Can't include in meta-analyses
439
+ - Hinders scientific progress
440
+
441
+ **Best practice:** Share data unless privacy concerns prohibit.
442
+
443
+ ## Cross-Validation and Generalization
444
+
445
+ ### Pitfall 39: No Cross-Validation
446
+ **Problem:** Testing model on same data used to build it.
447
+
448
+ **Consequence:** Overly optimistic performance estimates.
449
+
450
+ **Solutions:**
451
+ - Split data (train/test)
452
+ - K-fold cross-validation
453
+ - Independent validation sample
454
+
455
+ ### Pitfall 40: Data Leakage
456
+ **Problem:** Information from test set leaking into training.
457
+
458
+ **Examples:**
459
+ - Normalizing before splitting
460
+ - Feature selection on full dataset
461
+ - Including temporal information
462
+
463
+ **Consequence:** Inflated performance metrics.
464
+
465
+ **Prevention:** All preprocessing decisions made using only training data.
466
+
467
+ ## Meta-Analysis Pitfalls
468
+
469
+ ### Pitfall 41: Apples and Oranges
470
+ **Problem:** Combining studies with different designs, populations, or measures.
471
+
472
+ **Balance:** Need homogeneity but also comprehensiveness.
473
+
474
+ **Solutions:**
475
+ - Clear inclusion criteria
476
+ - Subgroup analyses
477
+ - Meta-regression for moderators
478
+
479
+ ### Pitfall 42: Ignoring Publication Bias
480
+ **Problem:** Published studies overrepresent significant results.
481
+
482
+ **Consequences:** Overestimated effects in meta-analyses.
483
+
484
+ **Detection:**
485
+ - Funnel plots
486
+ - Trim-and-fill
487
+ - PET-PEESE
488
+ - P-curve analysis
489
+
490
+ **Solutions:**
491
+ - Include unpublished studies
492
+ - Register reviews
493
+ - Use bias-correction methods
494
+
495
+ ## General Best Practices
496
+
497
+ 1. **Preregister studies** - Distinguish confirmatory from exploratory
498
+ 2. **Report transparently** - All analyses, not just significant ones
499
+ 3. **Check assumptions** - Don't blindly apply tests
500
+ 4. **Use appropriate tests** - Match test to data and design
501
+ 5. **Report effect sizes** - Not just p-values
502
+ 6. **Consider practical significance** - Not just statistical
503
+ 7. **Replicate findings** - One study is rarely definitive
504
+ 8. **Share data and code** - Enable verification
505
+ 9. **Use confidence intervals** - Show uncertainty
506
+ 10. **Think causally carefully** - Most research is correlational
@@ -0,0 +1,109 @@
1
+ ---
2
+ name: skill-creator
3
+ description: |
4
+ 【元技能·创作工艺·触发词: 创建技能, 新建skill, 写skill, 改进技能, 优化skill, 沉淀SOP, 把流程做成技能, 让我以后能复用, 技能描述触发不准】把可复用经验结晶成一个好用的 Skill 的工艺指南。无论你是要从零创建新 Skill、改进已有 Skill、还是优化它的触发描述——动手写 clear/skills 下任何 SKILL.md 之前,都先加载本技能取工艺,确保描述是好触发器、结构精简、解释了 why。
5
+ 不适用:写普通文档、README 或代码注释(那是内容工作,不是技能沉淀);
6
+ 一次性经验先落 memory,够通用了再回来沉淀成技能。
7
+ license: MIT
8
+ metadata:
9
+ tier: system
10
+ origin: template
11
+ version: "1.0-clearai"
12
+ ---
13
+
14
+ # Skill Creator(ClearAI 版)
15
+
16
+ 把"这次怎么做成的"结晶为一个**未来千百次可复用**的 Skill。本技能是动态 Skill 进化 loop 的工艺底座:agent 沉淀/改进 Skill 前先读它,保证产出质量。
17
+
18
+ ## 何时用 · 三种入口
19
+
20
+ 1. **创建新 Skill**:手上有一段值得复用的流程(常来自 WriteMemory 的 Trigger-Action-Validation 三元组)。
21
+ 2. **改进已有 Skill**:某 Skill 用下来有缺漏,要补 checklist / 修步骤 / 改描述。
22
+ 3. **优化触发描述**:Skill 该触发时没触发(欠触发),调 description。
23
+
24
+ 判断走哪条:先扫外脑索引 `<skills>`/`<candidate_skills>`,有匹配就改、没有就建、都不够通用就只留 memory lesson。
25
+
26
+ ## 在 ClearAI 里怎么落地(工具映射)
27
+
28
+ 写 skill 的**唯一**方式是 `SaveSkill`(勿用 write/edit 改 skill——会被拦):
29
+ - **新建** → `SaveSkill(skill="<name>", content=<SKILL.md 全文>)`。系统自动盖 `tier=candidate`、校验 frontmatter,落入「技能收件箱」等采纳。
30
+ - **改进已有** → `SaveSkill(skill="<已有名>", content=<改进后全文>)`。若该 skill 是 trusted/system,系统自动生成「候选改版」、**活版不动**,采纳时才替换。
31
+ - **子资源** → `SaveSkill(skill="<name>", path="scripts/x.py"|"references/y.md", content=...)`,一次写一个文件(全文)。
32
+ - SaveSkill 不弹写时审批:所有产物都进收件箱,**用户采纳是唯一的生效闸**。
33
+ - **不要**在正文引用子代理、`claude -p`、浏览器评测器等 ClearAI 没有的能力。
34
+
35
+ ## frontmatter 契约
36
+
37
+ ```yaml
38
+ name: <kebab-case> # 仅小写字母/数字/连字符
39
+ description: |
40
+ 【阶段·触发词: kw1, kw2, kw3】一句话能力。适用:...;不适用:指向其他 skill。
41
+ metadata:
42
+ tier: candidate # 系统会自动盖,无需手填
43
+ ```
44
+
45
+ **description 是触发的唯一机制,也是最该用心的地方**:
46
+ - 首行用 `【阶段·触发词: ...】` 列 3-6 个用户任务里会出现、能命中本 skill 的关键词。
47
+ - 要**"pushy"对抗欠触发**:明确写"当用户提到 X、Y、Z 时就用本技能,即使没明说要用"。当前模型倾向于该用 skill 时不用,描述写得主动一点能纠偏。
48
+ - 写清**适用 / 不适用**,不适用处指向更合适的 skill 名,减少误触发。
49
+
50
+ ## 结构与渐进式披露
51
+
52
+ Skill 三级加载,写时按此分层、保持每层精简:
53
+ 1. **name + description**(始终在 context)——触发判断只看这层。
54
+ 2. **SKILL.md 正文**(触发后加载,目标 <500 行)——SOP 主干。
55
+ 3. **子资源**(按需读/执行)——`scripts/` 确定性代码、`references/` 详细文档、`assets/` 产物模板。
56
+
57
+ 正文超长就加一层层级:SKILL.md 留主干 + 指针,细节挪进 `references/`,并在指针处说明"何时去读"。多领域/多框架按变体拆 `references/<variant>.md`,正文只做选择与编排。
58
+
59
+ ## 先分清这是谁的活
60
+
61
+ harness 执行不变量(与具体任务无关、没商量余地),skill 提供工艺(这类事怎么做得好),
62
+ agent 做接合(把工艺用到具体情境)。动笔前先问这条内容属于哪一层——放错层的内容不会报错,
63
+ 只会慢慢腐坏。
64
+
65
+ **平台已经强制的,写「为什么」而不是重抄规则。** 读者撞上拒绝时需要理由才能改对;一份和
66
+ 平台重复的规则表既帮不上忙,还会先于平台过期。
67
+
68
+ **平台强制不了的,才是 skill 的本职。** 本仓的例子:`verify` 阶段隔离的是数据目录与 HOME,
69
+ 不隔离网络,所以「运行时不引 CDN」在验证时看不出来,产物拿到没网的地方才炸——这条只能靠
70
+ 工艺纪律加技能自带的检查脚本兜住。这类内容写得越具体越好。
71
+
72
+ **数据的形状归上游技能的契约或 harness 的校验**,别焊进呈现/渲染类技能。调用时自然会带
73
+ 进来;写死了反而让技能在别的数据上用不了。
74
+
75
+ ### 别复述平台行为,那会漂移
76
+
77
+ 写「平台会怎样」的句子迟早和平台对不上,而读者拿它当真。本仓踩过:某技能写「手建目录不会
78
+ 被识别,且这个失败是静默的」,而平台早已给出带修复建议的显式错误。这句不只是过时——它暗示
79
+ agent 不必去读错误信息,可答案恰恰就在那里。
80
+
81
+ 写「你该怎么做、以及为什么」,把「平台实际怎么反应」留给它自己的错误信息去说,那永远是最新的。
82
+
83
+ ### 平台常量分两种
84
+
85
+ 判据只有一个:**这个数字会改变 agent 的下一步动作吗?**
86
+
87
+ 会就写。一个外部调用的超时秒数是「这个操作要不要拆成后台作业」的直接输入,不写他就没法判断。
88
+
89
+ 不会就别抄。state 的字节上限对 agent 没有决策价值,写「有上限、大数组换成摘要与 state_ref」
90
+ 就够了,具体数字让平台的拒绝去说。抄下来的常量没有任何机制会随平台更新,而它看起来永远
91
+ 像是准确的。
92
+
93
+ ## 写作风格(决定 skill 好不好用)
94
+
95
+ - **解释 why,而非堆大写 MUST**。模型很聪明、有 theory of mind,给它理由比给铁律更有效、更通用。写到 ALWAYS/NEVER 全大写或极刚性结构时,是黄旗——回头重述成"为什么这样做"。
96
+ - 用**祈使句**,面向通用场景而非过拟合到具体例子。
97
+ - 先写草稿,再以新眼光重看一遍精简——删掉不拉动效果的部分。
98
+
99
+ ## 评测 = 复用 ClearAI 已有回路(不要另造)
100
+
101
+ skill-creator 原版有一套子代理 benchmark + HTML viewer。在 ClearAI 里**用我们已有的两条回路顶替**:
102
+
103
+ - **人评** = 把 candidate 交给用户在**技能收件箱**采纳/驳回。采纳即"通过评审"。
104
+ - **量化** = 看该 skill 的 **uses/wins 遥测**(`clear/audit/skill_usage.jsonl` 聚合,索引按 wins/uses 排序)。被反复加载且关联成功计划多,说明它好用,自然上浮。
105
+ - **重复脚本信号** = 若多次任务里你都手写了同一个脚本,把它写一次放进该 skill 的 `scripts/` 并在正文引用,省掉未来每次重造轮子。
106
+
107
+ ## 更细的写作模式 / 示例 / 反例
108
+
109
+ 需要 description 优化细则、好/坏示例、近义反例设计、领域拆分模板时,读 `references/authoring-guide.md`(按需加载,不必一上来就读)。
@@ -0,0 +1,89 @@
1
+ # 创作工艺细则(按需加载)
2
+
3
+ SKILL.md 已讲核心。本文是第二层:description 优化、好坏示例、反例设计、领域拆分。需要时再读对应小节。
4
+
5
+ ## 目录
6
+ - 1. description 写法与"pushy"度
7
+ - 2. 触发词怎么选
8
+ - 3. 好例 / 坏例
9
+ - 4. 适用 / 不适用 的写法
10
+ - 5. 输出格式与示例模式
11
+ - 6. 何时该 bundle 脚本
12
+ - 7. 多领域拆分
13
+ - 8. 自检清单
14
+
15
+ ## 1. description 写法与"pushy"度
16
+
17
+ description 是 Skill 触发的唯一依据。模型只看 name+description 决定要不要加载本 skill。两个常见失败:
18
+ - **欠触发**(更常见):该用时不用。对策是把描述写得主动——显式列触发场景,并加一句"即使用户没明说要用本技能,只要涉及 X/Y/Z 就用"。
19
+ - **过触发**:不该用时也用。对策是写清"不适用"并指向更合适的 skill。
20
+
21
+ 把握度:宁可稍微 pushy 一点纠正欠触发,但用"不适用"兜住边界。
22
+
23
+ ## 2. 触发词怎么选
24
+
25
+ 放进 `【阶段·触发词: ...】` 的词,应是**用户真实任务文本里会出现**的词,而非你对该领域的术语。想象用户会怎么口语化地描述这个需求:
26
+ - 包含同义/近义说法(正式的、口语的)。
27
+ - 包含用户**不点名**本 skill、但显然需要它的场景词。
28
+ - 3-6 个即可,过多稀释。
29
+
30
+ ## 3. 好例 / 坏例
31
+
32
+ **description 坏例**(太被动、无触发场景):
33
+ > 引用管理:检索论文元数据,DOI 转 BibTeX。
34
+
35
+ **好例**(主动 + 触发场景 + 边界):
36
+ > 【引用管理·触发词: 引用, 参考文献, BibTeX, DOI, 文末引用, 去重】检索论文元数据、DOI→BibTeX、校验去重。适用:预研/综述报告文末引用统一、交付前引用准确性检查。不适用:全文综述撰写(用 literature-review)。
37
+
38
+ ## 4. 适用 / 不适用 的写法
39
+
40
+ - **适用**:列 2-4 个该 skill 最该上场的具体情境。
41
+ - **不适用**:列易混淆的邻近情境,并**指向应该用的 skill 名**。这既减少误触发,也帮模型在多 skill 竞争时选对。
42
+
43
+ ## 5. 输出格式与示例模式
44
+
45
+ 要求固定产出格式时,直接在正文给模板:
46
+ ```markdown
47
+ ## 报告结构(严格按此)
48
+ # 标题
49
+ ## 摘要
50
+ ## 关键发现
51
+ ## 建议
52
+ ```
53
+
54
+ 给示例帮助很大,用 输入→输出 对照:
55
+ ```markdown
56
+ **示例 1**
57
+ 输入:为用户认证加了 JWT
58
+ 输出:feat(auth): 实现基于 JWT 的认证
59
+ ```
60
+
61
+ ## 6. 何时该 bundle 脚本
62
+
63
+ 读你自己几次任务的过程记录:如果**多次**都独立写了类似的 `create_docx.py` / `build_chart.py` / 同样的多步处理,这是强信号——把它写一次放进该 skill 的 `scripts/`,正文里指明"用 scripts/xxx.py 而非每次重写"。每个未来调用都省掉重造轮子。
64
+
65
+ ClearAI 里脚本由 bash 工具执行,写类动作由系统策略治理。脚本要自带用法注释、参数清晰、可独立运行。
66
+
67
+ ## 7. 多领域拆分
68
+
69
+ 一个 skill 覆盖多领域/框架时,按变体拆 `references/`,正文只做选择与编排:
70
+ ```
71
+ cloud-deploy/
72
+ ├── SKILL.md # 工作流 + 如何选变体
73
+ └── references/
74
+ ├── aws.md
75
+ ├── gcp.md
76
+ └── azure.md
77
+ ```
78
+ 模型只读相关那一个变体文件,省 context。大参考文件(>300 行)开头放目录。
79
+
80
+ ## 8. 自检清单(写完回看)
81
+
82
+ - [ ] description 首行有 `【阶段·触发词: ...】`,词是用户会说的话。
83
+ - [ ] 写了"适用 / 不适用",不适用指向了别的 skill。
84
+ - [ ] 主动度够(不会欠触发),但有边界(不会过触发)。
85
+ - [ ] 正文是祈使句、解释了 why、没堆大写 MUST。
86
+ - [ ] SKILL.md 主干精简(<500 行),长内容挪进 references/ 并留指针。
87
+ - [ ] 重复出现的脚本已 bundle 进 scripts/。
88
+ - [ ] name 是 kebab-case;frontmatter 只含允许的 key。
89
+ - [ ] 评测交给收件箱(人评)+ uses/wins(量化),没有另造评测机器。