clearai-dsh 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (154) hide show
  1. package/CHANGELOG.md +26 -0
  2. package/LICENSE +201 -0
  3. package/README.md +138 -0
  4. package/README.zh-CN.md +138 -0
  5. package/bin/clearai.mjs +224 -0
  6. package/brand/README.md +41 -0
  7. package/brand/logo-512-dark.png +0 -0
  8. package/brand/logo-512.png +0 -0
  9. package/brand/logo-lockup-dark.png +0 -0
  10. package/brand/logo-lockup.png +0 -0
  11. package/brand/logo-lockup.svg +12 -0
  12. package/brand/logo-wordmark.svg +6 -0
  13. package/brand/logo.svg +19 -0
  14. package/cordis.patch.yml +39 -0
  15. package/lib/client.js +3071 -0
  16. package/lib/fold.js +1576 -0
  17. package/lib/host.js +605 -0
  18. package/package.json +65 -0
  19. package/presets/clearai/agent.cordis.yml +226 -0
  20. package/presets/clearai/plugins/brain.js +547 -0
  21. package/presets/clearai/plugins/clearai-kernel.js +5485 -0
  22. package/presets/clearai/plugins/ontology.js +306 -0
  23. package/presets/clearai/plugins/prompts.js +312 -0
  24. package/presets/clearai/preset.yml +5 -0
  25. package/presets/clearai/skills/clearai-loop/SKILL.md +89 -0
  26. package/presets/clearai/template/knowledge/README.md +25 -0
  27. package/presets/clearai/template/memory/README.md +34 -0
  28. package/presets/clearai/template/project.md +49 -0
  29. package/presets/clearai/template/skills/README.md +37 -0
  30. package/presets/clearai/template/skills/chart-diagram-qa/SKILL.md +43 -0
  31. package/presets/clearai/template/skills/citation-management/SKILL.md +73 -0
  32. package/presets/clearai/template/skills/citation-management/references/bibtex_formatting.md +908 -0
  33. package/presets/clearai/template/skills/citation-management/references/citation_validation.md +794 -0
  34. package/presets/clearai/template/skills/citation-management/references/google_scholar_search.md +725 -0
  35. package/presets/clearai/template/skills/citation-management/references/metadata_extraction.md +870 -0
  36. package/presets/clearai/template/skills/citation-management/references/pubmed_search.md +839 -0
  37. package/presets/clearai/template/skills/citation-management/scripts/doi_to_bibtex.py +204 -0
  38. package/presets/clearai/template/skills/citation-management/scripts/extract_metadata.py +569 -0
  39. package/presets/clearai/template/skills/citation-management/scripts/format_bibtex.py +349 -0
  40. package/presets/clearai/template/skills/citation-management/scripts/generate_schematic.py +139 -0
  41. package/presets/clearai/template/skills/citation-management/scripts/generate_schematic_ai.py +817 -0
  42. package/presets/clearai/template/skills/citation-management/scripts/search_google_scholar.py +282 -0
  43. package/presets/clearai/template/skills/citation-management/scripts/search_pubmed.py +398 -0
  44. package/presets/clearai/template/skills/citation-management/scripts/validate_citations.py +497 -0
  45. package/presets/clearai/template/skills/data-analysis/SKILL.md +92 -0
  46. package/presets/clearai/template/skills/data-analysis/checklists/readiness_check.md +23 -0
  47. package/presets/clearai/template/skills/data-analysis/templates/analysis_report.md.tpl +63 -0
  48. package/presets/clearai/template/skills/data-analysis/templates/cleaning_rules_draft.yaml.tpl +32 -0
  49. package/presets/clearai/template/skills/data-analysis/templates/data_dictionary.md.tpl +12 -0
  50. package/presets/clearai/template/skills/data-analysis/templates/domain_knowledge_template.md.tpl +316 -0
  51. package/presets/clearai/template/skills/data-analysis/templates/feature_candidates.json.tpl +20 -0
  52. package/presets/clearai/template/skills/data-analysis/templates/quality_scorecard.md.tpl +30 -0
  53. package/presets/clearai/template/skills/data-analysis/workflows/01-data-profiling.md +42 -0
  54. package/presets/clearai/template/skills/data-analysis/workflows/02-quality-audit.md +36 -0
  55. package/presets/clearai/template/skills/data-analysis/workflows/03-physical-correlation.md +25 -0
  56. package/presets/clearai/template/skills/data-analysis/workflows/04-unstructured-mining.md +26 -0
  57. package/presets/clearai/template/skills/data-qa-analysis/SKILL.md +102 -0
  58. package/presets/clearai/template/skills/data-qa-analysis/checklists/readiness_check.md +62 -0
  59. package/presets/clearai/template/skills/data-qa-analysis/templates/best_in_class_report.md.tpl +56 -0
  60. package/presets/clearai/template/skills/data-qa-analysis/templates/cleaning_rules_draft.yaml.tpl +56 -0
  61. package/presets/clearai/template/skills/data-qa-analysis/templates/data_dictionary.md.tpl +13 -0
  62. package/presets/clearai/template/skills/data-qa-analysis/templates/data_source_inventory_and_lineage.md.tpl +146 -0
  63. package/presets/clearai/template/skills/data-qa-analysis/templates/data_status_report.md.tpl +60 -0
  64. package/presets/clearai/template/skills/data-qa-analysis/templates/steady_state_rules.yaml.tpl +41 -0
  65. package/presets/clearai/template/skills/data-qa-analysis/templates/subsystem_registry.md.tpl +101 -0
  66. package/presets/clearai/template/skills/data-qa-analysis/templates/unified_execution_plan.md.tpl +100 -0
  67. package/presets/clearai/template/skills/data-qa-analysis/workflows/01-data-source-inventory-and-lineage.md +194 -0
  68. package/presets/clearai/template/skills/data-qa-analysis/workflows/02-data-alignment-and-tag-semantics.md +122 -0
  69. package/presets/clearai/template/skills/data-qa-analysis/workflows/03-steady-state-identification.md +126 -0
  70. package/presets/clearai/template/skills/data-qa-analysis/workflows/04-consumption-analysis.md +152 -0
  71. package/presets/clearai/template/skills/data-qa-analysis/workflows/05-best-in-class-and-optimization-space.md +78 -0
  72. package/presets/clearai/template/skills/domain-presearch/SKILL.md +131 -0
  73. package/presets/clearai/template/skills/domain-presearch/checklists/domain_checklist.md +24 -0
  74. package/presets/clearai/template/skills/domain-presearch/references/figure_code.md +78 -0
  75. package/presets/clearai/template/skills/domain-presearch/references/strategic_frameworks.md +38 -0
  76. package/presets/clearai/template/skills/exploration-loop/SKILL.md +81 -0
  77. package/presets/clearai/template/skills/exploratory-data-analysis/SKILL.md +77 -0
  78. package/presets/clearai/template/skills/exploratory-data-analysis/references/bioinformatics_genomics_formats.md +664 -0
  79. package/presets/clearai/template/skills/exploratory-data-analysis/references/chemistry_molecular_formats.md +664 -0
  80. package/presets/clearai/template/skills/exploratory-data-analysis/references/general_scientific_formats.md +518 -0
  81. package/presets/clearai/template/skills/exploratory-data-analysis/references/microscopy_imaging_formats.md +620 -0
  82. package/presets/clearai/template/skills/exploratory-data-analysis/references/proteomics_metabolomics_formats.md +517 -0
  83. package/presets/clearai/template/skills/exploratory-data-analysis/references/spectroscopy_analytical_formats.md +633 -0
  84. package/presets/clearai/template/skills/exploratory-data-analysis/scripts/eda_analyzer.py +547 -0
  85. package/presets/clearai/template/skills/hypothesis-generation/SKILL.md +73 -0
  86. package/presets/clearai/template/skills/hypothesis-generation/references/experimental_design_patterns.md +329 -0
  87. package/presets/clearai/template/skills/hypothesis-generation/references/hypothesis_quality_criteria.md +198 -0
  88. package/presets/clearai/template/skills/hypothesis-generation/references/literature_search_strategies.md +622 -0
  89. package/presets/clearai/template/skills/hypothesis-generation/scripts/generate_schematic.py +139 -0
  90. package/presets/clearai/template/skills/hypothesis-generation/scripts/generate_schematic_ai.py +817 -0
  91. package/presets/clearai/template/skills/literature-review/SKILL.md +72 -0
  92. package/presets/clearai/template/skills/literature-review/references/citation_styles.md +166 -0
  93. package/presets/clearai/template/skills/literature-review/references/database_strategies.md +455 -0
  94. package/presets/clearai/template/skills/literature-review/scripts/generate_pdf.py +176 -0
  95. package/presets/clearai/template/skills/literature-review/scripts/generate_schematic.py +139 -0
  96. package/presets/clearai/template/skills/literature-review/scripts/generate_schematic_ai.py +817 -0
  97. package/presets/clearai/template/skills/literature-review/scripts/search_databases.py +303 -0
  98. package/presets/clearai/template/skills/literature-review/scripts/verify_citations.py +221 -0
  99. package/presets/clearai/template/skills/paper-lookup/SKILL.md +59 -0
  100. package/presets/clearai/template/skills/paper-lookup/references/arxiv.md +161 -0
  101. package/presets/clearai/template/skills/paper-lookup/references/biorxiv.md +118 -0
  102. package/presets/clearai/template/skills/paper-lookup/references/core.md +150 -0
  103. package/presets/clearai/template/skills/paper-lookup/references/crossref.md +181 -0
  104. package/presets/clearai/template/skills/paper-lookup/references/medrxiv.md +104 -0
  105. package/presets/clearai/template/skills/paper-lookup/references/openalex.md +174 -0
  106. package/presets/clearai/template/skills/paper-lookup/references/pmc.md +152 -0
  107. package/presets/clearai/template/skills/paper-lookup/references/pubmed.md +124 -0
  108. package/presets/clearai/template/skills/paper-lookup/references/semantic-scholar.md +203 -0
  109. package/presets/clearai/template/skills/paper-lookup/references/unpaywall.md +127 -0
  110. package/presets/clearai/template/skills/process-presearch/SKILL.md +196 -0
  111. package/presets/clearai/template/skills/process-presearch/checklists/process_checklist.md +18 -0
  112. package/presets/clearai/template/skills/process-presearch/references/figure_code.md +107 -0
  113. package/presets/clearai/template/skills/process-presearch/references/source_attribution_example.md +22 -0
  114. package/presets/clearai/template/skills/process-understanding-extraction/SKILL.md +69 -0
  115. package/presets/clearai/template/skills/process-understanding-extraction/checklists/readiness_check.md +34 -0
  116. package/presets/clearai/template/skills/process-understanding-extraction/templates/docx_raw_dump_extractor.py.tpl +132 -0
  117. package/presets/clearai/template/skills/process-understanding-extraction/templates/entity_map_unit_topology.json.tpl +86 -0
  118. package/presets/clearai/template/skills/process-understanding-extraction/templates/process_brief.md.tpl +89 -0
  119. package/presets/clearai/template/skills/process-understanding-extraction/templates/process_brief_builder_from_raw_dump.py.tpl +203 -0
  120. package/presets/clearai/template/skills/process-understanding-extraction/templates/process_flow_mermaid.md.tpl +41 -0
  121. package/presets/clearai/template/skills/process-understanding-extraction/templates/unified_execution_plan.md.tpl +53 -0
  122. package/presets/clearai/template/skills/process-understanding-extraction/workflows/01-process-doc-discovery.md +173 -0
  123. package/presets/clearai/template/skills/process-understanding-extraction/workflows/02-process-understanding-and-diagramming.md +106 -0
  124. package/presets/clearai/template/skills/scientific-brainstorming/SKILL.md +64 -0
  125. package/presets/clearai/template/skills/scientific-brainstorming/references/brainstorming_methods.md +326 -0
  126. package/presets/clearai/template/skills/scientific-critical-thinking/SKILL.md +72 -0
  127. package/presets/clearai/template/skills/scientific-critical-thinking/references/common_biases.md +364 -0
  128. package/presets/clearai/template/skills/scientific-critical-thinking/references/evidence_hierarchy.md +485 -0
  129. package/presets/clearai/template/skills/scientific-critical-thinking/references/experimental_design.md +496 -0
  130. package/presets/clearai/template/skills/scientific-critical-thinking/references/logical_fallacies.md +478 -0
  131. package/presets/clearai/template/skills/scientific-critical-thinking/references/scientific_method.md +169 -0
  132. package/presets/clearai/template/skills/scientific-critical-thinking/references/statistical_pitfalls.md +506 -0
  133. package/presets/clearai/template/skills/skill-creator/SKILL.md +109 -0
  134. package/presets/clearai/template/skills/skill-creator/references/authoring-guide.md +89 -0
  135. package/presets/clearai/template/skills/statistical-analysis/SKILL.md +79 -0
  136. package/presets/clearai/template/skills/statistical-analysis/references/assumptions_and_diagnostics.md +369 -0
  137. package/presets/clearai/template/skills/statistical-analysis/references/bayesian_statistics.md +653 -0
  138. package/presets/clearai/template/skills/statistical-analysis/references/effect_sizes_and_power.md +578 -0
  139. package/presets/clearai/template/skills/statistical-analysis/references/reporting_standards.md +469 -0
  140. package/presets/clearai/template/skills/statistical-analysis/references/test_selection_guide.md +129 -0
  141. package/presets/clearai/template/skills/statistical-analysis/scripts/assumption_checks.py +538 -0
  142. package/presets/clearai/template/skills/web-artifact/SKILL.md +165 -0
  143. package/presets/clearai/template/skills/web-artifact/assets/renderer/renderer.css +229 -0
  144. package/presets/clearai/template/skills/web-artifact/assets/renderer/renderer.js +373 -0
  145. package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/LICENSE +263 -0
  146. package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/UPSTREAM.md +26 -0
  147. package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/elk.bundled.js +6605 -0
  148. package/presets/clearai/template/skills/web-artifact/references/when-drawing-a-topology.md +150 -0
  149. package/presets/clearai/template/skills/web-artifact/references/when-the-page-must-work-offline.md +62 -0
  150. package/presets/clearai/template/skills/web-artifact/scripts/check_artifact.py +167 -0
  151. package/presets/clearai/template/skills/web-artifact/scripts/render_topology.js +272 -0
  152. package/presets/clearai/template/skills/what-if-oracle/LICENSE.txt +5 -0
  153. package/presets/clearai/template/skills/what-if-oracle/SKILL.md +72 -0
  154. package/presets/clearai/template/skills/what-if-oracle/references/scenario-templates.md +154 -0
@@ -0,0 +1,538 @@
1
+ """
2
+ Comprehensive statistical assumption checking utilities.
3
+
4
+ This module provides functions to check common statistical assumptions:
5
+ - Normality
6
+ - Homogeneity of variance
7
+ - Independence
8
+ - Linearity
9
+ - Outliers
10
+ """
11
+
12
+ import numpy as np
13
+ import pandas as pd
14
+ from scipy import stats
15
+ import matplotlib.pyplot as plt
16
+ from typing import Dict, List, Optional, Union
17
+
18
+
19
+ def check_normality(
20
+ data: Union[np.ndarray, pd.Series, List],
21
+ name: str = "data",
22
+ alpha: float = 0.05,
23
+ plot: bool = True
24
+ ) -> Dict:
25
+ """
26
+ Check normality assumption using Shapiro-Wilk test and visualizations.
27
+
28
+ Parameters
29
+ ----------
30
+ data : array-like
31
+ Data to check for normality
32
+ name : str
33
+ Name of the variable (for labeling)
34
+ alpha : float
35
+ Significance level for Shapiro-Wilk test
36
+ plot : bool
37
+ Whether to create Q-Q plot and histogram
38
+
39
+ Returns
40
+ -------
41
+ dict
42
+ Results including test statistic, p-value, and interpretation
43
+ """
44
+ data = np.asarray(data)
45
+ data_clean = data[~np.isnan(data)]
46
+
47
+ # Shapiro-Wilk test
48
+ statistic, p_value = stats.shapiro(data_clean)
49
+
50
+ # Interpretation
51
+ is_normal = p_value > alpha
52
+ interpretation = (
53
+ f"Data {'appear' if is_normal else 'do not appear'} normally distributed "
54
+ f"(W = {statistic:.3f}, p = {p_value:.3f})"
55
+ )
56
+
57
+ # Visual checks
58
+ if plot:
59
+ fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(12, 4))
60
+
61
+ # Q-Q plot
62
+ stats.probplot(data_clean, dist="norm", plot=ax1)
63
+ ax1.set_title(f"Q-Q Plot: {name}")
64
+ ax1.grid(alpha=0.3)
65
+
66
+ # Histogram with normal curve
67
+ ax2.hist(data_clean, bins='auto', density=True, alpha=0.7, color='steelblue', edgecolor='black')
68
+ mu, sigma = data_clean.mean(), data_clean.std()
69
+ x = np.linspace(data_clean.min(), data_clean.max(), 100)
70
+ ax2.plot(x, stats.norm.pdf(x, mu, sigma), 'r-', linewidth=2, label='Normal curve')
71
+ ax2.set_xlabel('Value')
72
+ ax2.set_ylabel('Density')
73
+ ax2.set_title(f'Histogram: {name}')
74
+ ax2.legend()
75
+ ax2.grid(alpha=0.3)
76
+
77
+ plt.tight_layout()
78
+ plt.show()
79
+
80
+ return {
81
+ 'test': 'Shapiro-Wilk',
82
+ 'statistic': statistic,
83
+ 'p_value': p_value,
84
+ 'is_normal': is_normal,
85
+ 'interpretation': interpretation,
86
+ 'n': len(data_clean),
87
+ 'recommendation': (
88
+ "Proceed with parametric test" if is_normal
89
+ else "Consider non-parametric alternative or transformation"
90
+ )
91
+ }
92
+
93
+
94
+ def check_normality_per_group(
95
+ data: pd.DataFrame,
96
+ value_col: str,
97
+ group_col: str,
98
+ alpha: float = 0.05,
99
+ plot: bool = True
100
+ ) -> pd.DataFrame:
101
+ """
102
+ Check normality assumption for each group separately.
103
+
104
+ Parameters
105
+ ----------
106
+ data : pd.DataFrame
107
+ Data containing values and group labels
108
+ value_col : str
109
+ Column name for values to check
110
+ group_col : str
111
+ Column name for group labels
112
+ alpha : float
113
+ Significance level
114
+ plot : bool
115
+ Whether to create Q-Q plots for each group
116
+
117
+ Returns
118
+ -------
119
+ pd.DataFrame
120
+ Results for each group
121
+ """
122
+ groups = data[group_col].unique()
123
+ results = []
124
+
125
+ if plot:
126
+ n_groups = len(groups)
127
+ fig, axes = plt.subplots(1, n_groups, figsize=(5 * n_groups, 4))
128
+ if n_groups == 1:
129
+ axes = [axes]
130
+
131
+ for idx, group in enumerate(groups):
132
+ group_data = data[data[group_col] == group][value_col].dropna()
133
+ stat, p = stats.shapiro(group_data)
134
+
135
+ results.append({
136
+ 'Group': group,
137
+ 'N': len(group_data),
138
+ 'W': stat,
139
+ 'p-value': p,
140
+ 'Normal': 'Yes' if p > alpha else 'No'
141
+ })
142
+
143
+ if plot:
144
+ stats.probplot(group_data, dist="norm", plot=axes[idx])
145
+ axes[idx].set_title(f"Q-Q Plot: {group}")
146
+ axes[idx].grid(alpha=0.3)
147
+
148
+ if plot:
149
+ plt.tight_layout()
150
+ plt.show()
151
+
152
+ return pd.DataFrame(results)
153
+
154
+
155
+ def check_homogeneity_of_variance(
156
+ data: pd.DataFrame,
157
+ value_col: str,
158
+ group_col: str,
159
+ alpha: float = 0.05,
160
+ plot: bool = True
161
+ ) -> Dict:
162
+ """
163
+ Check homogeneity of variance using Levene's test.
164
+
165
+ Parameters
166
+ ----------
167
+ data : pd.DataFrame
168
+ Data containing values and group labels
169
+ value_col : str
170
+ Column name for values
171
+ group_col : str
172
+ Column name for group labels
173
+ alpha : float
174
+ Significance level
175
+ plot : bool
176
+ Whether to create box plots
177
+
178
+ Returns
179
+ -------
180
+ dict
181
+ Results including test statistic, p-value, and interpretation
182
+ """
183
+ groups = [group[value_col].values for name, group in data.groupby(group_col)]
184
+
185
+ # Levene's test (robust to non-normality)
186
+ statistic, p_value = stats.levene(*groups)
187
+
188
+ # Variance ratio (max/min)
189
+ variances = [np.var(g, ddof=1) for g in groups]
190
+ var_ratio = max(variances) / min(variances)
191
+
192
+ is_homogeneous = p_value > alpha
193
+ interpretation = (
194
+ f"Variances {'appear' if is_homogeneous else 'do not appear'} homogeneous "
195
+ f"(F = {statistic:.3f}, p = {p_value:.3f}, variance ratio = {var_ratio:.2f})"
196
+ )
197
+
198
+ if plot:
199
+ fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(12, 4))
200
+
201
+ # Box plot
202
+ data.boxplot(column=value_col, by=group_col, ax=ax1)
203
+ ax1.set_title('Box Plots by Group')
204
+ ax1.set_xlabel(group_col)
205
+ ax1.set_ylabel(value_col)
206
+ plt.sca(ax1)
207
+ plt.xticks(rotation=45)
208
+
209
+ # Variance plot
210
+ group_names = data[group_col].unique()
211
+ ax2.bar(range(len(variances)), variances, color='steelblue', edgecolor='black')
212
+ ax2.set_xticks(range(len(variances)))
213
+ ax2.set_xticklabels(group_names, rotation=45)
214
+ ax2.set_ylabel('Variance')
215
+ ax2.set_title('Variance by Group')
216
+ ax2.grid(alpha=0.3, axis='y')
217
+
218
+ plt.tight_layout()
219
+ plt.show()
220
+
221
+ return {
222
+ 'test': 'Levene',
223
+ 'statistic': statistic,
224
+ 'p_value': p_value,
225
+ 'is_homogeneous': is_homogeneous,
226
+ 'variance_ratio': var_ratio,
227
+ 'interpretation': interpretation,
228
+ 'recommendation': (
229
+ "Proceed with standard test" if is_homogeneous
230
+ else "Consider Welch's correction or transformation"
231
+ )
232
+ }
233
+
234
+
235
+ def check_linearity(
236
+ x: Union[np.ndarray, pd.Series],
237
+ y: Union[np.ndarray, pd.Series],
238
+ x_name: str = "X",
239
+ y_name: str = "Y"
240
+ ) -> Dict:
241
+ """
242
+ Check linearity assumption for regression.
243
+
244
+ Parameters
245
+ ----------
246
+ x : array-like
247
+ Predictor variable
248
+ y : array-like
249
+ Outcome variable
250
+ x_name : str
251
+ Name of predictor
252
+ y_name : str
253
+ Name of outcome
254
+
255
+ Returns
256
+ -------
257
+ dict
258
+ Visualization and recommendations
259
+ """
260
+ x = np.asarray(x)
261
+ y = np.asarray(y)
262
+
263
+ # Fit linear regression
264
+ slope, intercept, r_value, p_value, std_err = stats.linregress(x, y)
265
+ y_pred = intercept + slope * x
266
+
267
+ # Calculate residuals
268
+ residuals = y - y_pred
269
+
270
+ # Visualization
271
+ fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(12, 4))
272
+
273
+ # Scatter plot with regression line
274
+ ax1.scatter(x, y, alpha=0.6, s=50, edgecolors='black', linewidths=0.5)
275
+ ax1.plot(x, y_pred, 'r-', linewidth=2, label=f'y = {intercept:.2f} + {slope:.2f}x')
276
+ ax1.set_xlabel(x_name)
277
+ ax1.set_ylabel(y_name)
278
+ ax1.set_title('Scatter Plot with Regression Line')
279
+ ax1.legend()
280
+ ax1.grid(alpha=0.3)
281
+
282
+ # Residuals vs fitted
283
+ ax2.scatter(y_pred, residuals, alpha=0.6, s=50, edgecolors='black', linewidths=0.5)
284
+ ax2.axhline(y=0, color='r', linestyle='--', linewidth=2)
285
+ ax2.set_xlabel('Fitted values')
286
+ ax2.set_ylabel('Residuals')
287
+ ax2.set_title('Residuals vs Fitted Values')
288
+ ax2.grid(alpha=0.3)
289
+
290
+ plt.tight_layout()
291
+ plt.show()
292
+
293
+ return {
294
+ 'r': r_value,
295
+ 'r_squared': r_value ** 2,
296
+ 'interpretation': (
297
+ "Examine residual plot. Points should be randomly scattered around zero. "
298
+ "Patterns (curves, funnels) suggest non-linearity or heteroscedasticity."
299
+ ),
300
+ 'recommendation': (
301
+ "If non-linear pattern detected: Consider polynomial terms, "
302
+ "transformations, or non-linear models"
303
+ )
304
+ }
305
+
306
+
307
+ def detect_outliers(
308
+ data: Union[np.ndarray, pd.Series, List],
309
+ name: str = "data",
310
+ method: str = "iqr",
311
+ threshold: float = 1.5,
312
+ plot: bool = True
313
+ ) -> Dict:
314
+ """
315
+ Detect outliers using IQR method or z-score method.
316
+
317
+ Parameters
318
+ ----------
319
+ data : array-like
320
+ Data to check for outliers
321
+ name : str
322
+ Name of variable
323
+ method : str
324
+ Method to use: 'iqr' or 'zscore'
325
+ threshold : float
326
+ Threshold for outlier detection
327
+ For IQR: typically 1.5 (mild) or 3 (extreme)
328
+ For z-score: typically 3
329
+ plot : bool
330
+ Whether to create visualizations
331
+
332
+ Returns
333
+ -------
334
+ dict
335
+ Outlier indices, values, and visualizations
336
+ """
337
+ data = np.asarray(data)
338
+ data_clean = data[~np.isnan(data)]
339
+
340
+ if method == "iqr":
341
+ q1 = np.percentile(data_clean, 25)
342
+ q3 = np.percentile(data_clean, 75)
343
+ iqr = q3 - q1
344
+ lower_bound = q1 - threshold * iqr
345
+ upper_bound = q3 + threshold * iqr
346
+ outlier_mask = (data_clean < lower_bound) | (data_clean > upper_bound)
347
+
348
+ elif method == "zscore":
349
+ z_scores = np.abs(stats.zscore(data_clean))
350
+ outlier_mask = z_scores > threshold
351
+ lower_bound = data_clean.mean() - threshold * data_clean.std()
352
+ upper_bound = data_clean.mean() + threshold * data_clean.std()
353
+
354
+ else:
355
+ raise ValueError("method must be 'iqr' or 'zscore'")
356
+
357
+ outlier_indices = np.where(outlier_mask)[0]
358
+ outlier_values = data_clean[outlier_mask]
359
+ n_outliers = len(outlier_indices)
360
+ pct_outliers = (n_outliers / len(data_clean)) * 100
361
+
362
+ if plot:
363
+ fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(12, 4))
364
+
365
+ # Box plot
366
+ bp = ax1.boxplot(data_clean, vert=True, patch_artist=True)
367
+ bp['boxes'][0].set_facecolor('steelblue')
368
+ ax1.set_ylabel('Value')
369
+ ax1.set_title(f'Box Plot: {name}')
370
+ ax1.grid(alpha=0.3, axis='y')
371
+
372
+ # Scatter plot highlighting outliers
373
+ x_coords = np.arange(len(data_clean))
374
+ ax2.scatter(x_coords[~outlier_mask], data_clean[~outlier_mask],
375
+ alpha=0.6, s=50, color='steelblue', label='Normal', edgecolors='black', linewidths=0.5)
376
+ if n_outliers > 0:
377
+ ax2.scatter(x_coords[outlier_mask], data_clean[outlier_mask],
378
+ alpha=0.8, s=100, color='red', label='Outliers', marker='D', edgecolors='black', linewidths=0.5)
379
+ ax2.axhline(y=lower_bound, color='orange', linestyle='--', linewidth=1.5, label='Bounds')
380
+ ax2.axhline(y=upper_bound, color='orange', linestyle='--', linewidth=1.5)
381
+ ax2.set_xlabel('Index')
382
+ ax2.set_ylabel('Value')
383
+ ax2.set_title(f'Outlier Detection: {name}')
384
+ ax2.legend()
385
+ ax2.grid(alpha=0.3)
386
+
387
+ plt.tight_layout()
388
+ plt.show()
389
+
390
+ return {
391
+ 'method': method,
392
+ 'threshold': threshold,
393
+ 'n_outliers': n_outliers,
394
+ 'pct_outliers': pct_outliers,
395
+ 'outlier_indices': outlier_indices,
396
+ 'outlier_values': outlier_values,
397
+ 'lower_bound': lower_bound,
398
+ 'upper_bound': upper_bound,
399
+ 'interpretation': f"Found {n_outliers} outliers ({pct_outliers:.1f}% of data)",
400
+ 'recommendation': (
401
+ "Investigate outliers for data entry errors. "
402
+ "Consider: (1) removing if errors, (2) winsorizing, "
403
+ "(3) keeping if legitimate, (4) using robust methods"
404
+ )
405
+ }
406
+
407
+
408
+ def comprehensive_assumption_check(
409
+ data: pd.DataFrame,
410
+ value_col: str,
411
+ group_col: Optional[str] = None,
412
+ alpha: float = 0.05
413
+ ) -> Dict:
414
+ """
415
+ Perform comprehensive assumption checking for common statistical tests.
416
+
417
+ Parameters
418
+ ----------
419
+ data : pd.DataFrame
420
+ Data to check
421
+ value_col : str
422
+ Column name for dependent variable
423
+ group_col : str, optional
424
+ Column name for grouping variable (if applicable)
425
+ alpha : float
426
+ Significance level
427
+
428
+ Returns
429
+ -------
430
+ dict
431
+ Summary of all assumption checks
432
+ """
433
+ print("=" * 70)
434
+ print("COMPREHENSIVE ASSUMPTION CHECK")
435
+ print("=" * 70)
436
+
437
+ results = {}
438
+
439
+ # Outlier detection
440
+ print("\n1. OUTLIER DETECTION")
441
+ print("-" * 70)
442
+ outlier_results = detect_outliers(
443
+ data[value_col].dropna(),
444
+ name=value_col,
445
+ method='iqr',
446
+ plot=True
447
+ )
448
+ results['outliers'] = outlier_results
449
+ print(f" {outlier_results['interpretation']}")
450
+ print(f" {outlier_results['recommendation']}")
451
+
452
+ # Check if grouped data
453
+ if group_col is not None:
454
+ # Normality per group
455
+ print(f"\n2. NORMALITY CHECK (by {group_col})")
456
+ print("-" * 70)
457
+ normality_results = check_normality_per_group(
458
+ data, value_col, group_col, alpha=alpha, plot=True
459
+ )
460
+ results['normality_per_group'] = normality_results
461
+ print(normality_results.to_string(index=False))
462
+
463
+ all_normal = normality_results['Normal'].eq('Yes').all()
464
+ print(f"\n All groups normal: {'Yes' if all_normal else 'No'}")
465
+ if not all_normal:
466
+ print(" → Consider non-parametric alternative (Mann-Whitney, Kruskal-Wallis)")
467
+
468
+ # Homogeneity of variance
469
+ print(f"\n3. HOMOGENEITY OF VARIANCE")
470
+ print("-" * 70)
471
+ homogeneity_results = check_homogeneity_of_variance(
472
+ data, value_col, group_col, alpha=alpha, plot=True
473
+ )
474
+ results['homogeneity'] = homogeneity_results
475
+ print(f" {homogeneity_results['interpretation']}")
476
+ print(f" {homogeneity_results['recommendation']}")
477
+
478
+ else:
479
+ # Overall normality
480
+ print(f"\n2. NORMALITY CHECK")
481
+ print("-" * 70)
482
+ normality_results = check_normality(
483
+ data[value_col].dropna(),
484
+ name=value_col,
485
+ alpha=alpha,
486
+ plot=True
487
+ )
488
+ results['normality'] = normality_results
489
+ print(f" {normality_results['interpretation']}")
490
+ print(f" {normality_results['recommendation']}")
491
+
492
+ # Summary
493
+ print("\n" + "=" * 70)
494
+ print("SUMMARY")
495
+ print("=" * 70)
496
+
497
+ if group_col is not None:
498
+ all_normal = results.get('normality_per_group', pd.DataFrame()).get('Normal', pd.Series()).eq('Yes').all()
499
+ is_homogeneous = results.get('homogeneity', {}).get('is_homogeneous', False)
500
+
501
+ if all_normal and is_homogeneous:
502
+ print("✓ All assumptions met. Proceed with parametric test (t-test, ANOVA).")
503
+ elif not all_normal:
504
+ print("✗ Normality violated. Use non-parametric alternative.")
505
+ elif not is_homogeneous:
506
+ print("✗ Homogeneity violated. Use Welch's correction or transformation.")
507
+ else:
508
+ is_normal = results.get('normality', {}).get('is_normal', False)
509
+ if is_normal:
510
+ print("✓ Normality assumption met.")
511
+ else:
512
+ print("✗ Normality violated. Consider transformation or non-parametric method.")
513
+
514
+ print("=" * 70)
515
+
516
+ return results
517
+
518
+
519
+ if __name__ == "__main__":
520
+ # Example usage
521
+ np.random.seed(42)
522
+
523
+ # Simulate data
524
+ group_a = np.random.normal(75, 8, 50)
525
+ group_b = np.random.normal(68, 10, 50)
526
+
527
+ df = pd.DataFrame({
528
+ 'score': np.concatenate([group_a, group_b]),
529
+ 'group': ['A'] * 50 + ['B'] * 50
530
+ })
531
+
532
+ # Run comprehensive check
533
+ results = comprehensive_assumption_check(
534
+ df,
535
+ value_col='score',
536
+ group_col='group',
537
+ alpha=0.05
538
+ )
@@ -0,0 +1,165 @@
1
+ ---
2
+ name: web-artifact
3
+ description: |
4
+ 【单文件网页制品】把数据渲染成一个自包含、能双击打开、可标注回传的网页,交给懂行的人挑错。
5
+ 触发词:拓扑图、流程图、检查件、评审件、预览件、发给专家看、离线单文件、双击打开的 HTML、
6
+ 可交互图表、annotations 反馈。
7
+ 适用:内容由数据生成而非手写、要发给不在 Workbench 里的人、且没有活的可被驱动的状态。
8
+ 典型是流程/空间拓扑确认、结构检查、参数对照、结论页。
9
+ 不适用:有活状态就建 App(用 app-authoring 判断);营销页与教学页;
10
+ 从 PDF/会议记录做提取本身(那是 process-understanding-extraction 与 semantica 的事,
11
+ 本技能只渲染提取结果)。
12
+ version: '1.0'
13
+ metadata:
14
+ tier: system
15
+ origin: template
16
+ artifacts: [products/artifacts/**, lab/**]
17
+ vendored: [elkjs 0.12.0 (EPL-2.0)]
18
+ ---
19
+
20
+ # 单文件网页制品
21
+
22
+ 做一个自包含、能双击打开、随数据重新生成的网页。它的读者是懂行但不在 Workbench 里的人,
23
+ 他的任务是挑错——所以这份东西要经得起离线打开、要说清自己从哪来、要让意见能带回来。
24
+
25
+ ## 什么时候用
26
+
27
+ 四条同时成立:产物是一个网页 · 内容由数据生成而非手写 · 要发给不在 Workbench 里的人 ·
28
+ 没有活的可被驱动的状态。最后一条不成立就不是本技能的活——有活状态的目标需要一个可被
29
+ 驱动的 App,本技能只负责静态网页这一条腿。
30
+
31
+ | 相邻职责 | 分工 |
32
+ |---|---|
33
+ | 数据契约与提取工艺 | 它们拥有数据契约。**本技能只渲染数据,不规定数据里必须有什么** |
34
+ | 工程几何 | 它管几何与基准对不对;本技能管把关系画得能让人看懂 |
35
+
36
+ **本技能不规定数据形状。** 节点与边上除渲染必需字段之外的属性一律原样透传到侧栏——
37
+ 有位号就显示位号,有出处就显示出处,什么都没有也能正常渲染。数据该长什么样,由上游技能
38
+ 的契约或 harness 的校验决定,调用时自然带进来。
39
+
40
+ ## 先校准处理档位
41
+
42
+ 校准的是处理规格,而非要不要设计。三档:
43
+
44
+ - **内部核对**——自己或同事看一眼就丢。够读就行,别花时间。
45
+ - **专家评审**(默认档)——外人要在上面挑错。可标注、能离线、有来历,排版清楚。
46
+ - **对外交付**——会被转发、会被存档。加完整的视觉身份。
47
+
48
+ 多数任务落在中间档。拿不准时按中间档做:一个排版干净的页面从不会是错的,一个过度设计的
49
+ 页面有时是。
50
+
51
+ ## 自包含是硬约束
52
+
53
+ 依赖全部内联,运行时零网络请求。构建期联网随意(本技能自己就在构建期跑 ELK)。
54
+
55
+ 这条守的是「没网的会议室里也能打开」,而不是「不许用 CDN」这句字面禁令。理解了本意,
56
+ 遇到新情况就能自己判断:Google Fonts 也是外部请求,所以字体只用系统栈;图片走 data URI;
57
+ 连 `<link rel="preconnect">` 都不必留。
58
+
59
+ 被移植进来的上游方案正是栽在这里——它从 esm.sh 与 jsdelivr 取 React 与 React Flow,
60
+ 而「拿给不在 Workbench 里的专家看」是它唯一的用途。
61
+
62
+ ## 从数据生成,不手工摆位
63
+
64
+ 几何由布局器算,渲染器只画。手改坐标、手挪标签、在布局器之外再跑一次路由,都会让制品与
65
+ 数据脱节——下次重新生成时那些手工调整全部丢失,而你多半已经忘了改过什么。
66
+
67
+ 需要更紧凑或更松散时,调布局器的参数,不要动它的输出。
68
+
69
+ ## 为什么不用框架
70
+
71
+ 默认无框架:布局在构建期算完,浏览器端用原生 SVG/DOM 加少量 JS 渲染。产物含数据约 40 KB,
72
+ 脚本直出,没有构建步骤。
73
+
74
+ 两条根本理由。**制品是生成物**,随数据重新生成,框架最大的卖点——可维护性与组件复用——
75
+ 在这里不成立。**运行时状态规模小且耦合低**,检查类制品通常只有选中项、过滤集、视图变换、
76
+ 标注列表几个状态,每个的影响都是局部且直接的,没有一处需要虚拟 DOM 去 diff。
77
+
78
+ 具体到量级:zoom/pan 约 30 行,点选高亮约 20 行,侧栏约 30 行,过滤约 40 行,标注导出
79
+ 约 60 行,minimap 约 40 行。自己写是实事求是,不是逞强。作为对照,内联 React 与 React Flow
80
+ 约 450 KB,换来的只是画盒子、画路径、放标签、pan/zoom 四件事。
81
+
82
+ ## 想引框架时,先回头看那道岔路
83
+
84
+ 需要虚拟 DOM 的信号是 UI 要随状态重算大量节点:虚拟滚动、多视图联动、表单编辑、undo/redo。
85
+ 出现这些时先回答另一个问题——
86
+
87
+ > 一个「制品」复杂到需要框架与构建步骤,多半已经跨过了「要不要 App」的分水岭。
88
+
89
+ 需要复杂状态管理,通常意味着有活的、可被驱动的状态,那正是建 App 的判据。真到那一步,走
90
+ `web-app-v1` scaffold(vite 构建,任何 npm 包随便用,vite 负责打成自包含产物),而不是在
91
+ 这里手搓半个框架。
92
+
93
+ ## 让读者知道从哪里开始看
94
+
95
+ 专家的时间很贵。让他从头扫一遍图,和让他先看那 12 处存疑,效率差一个量级。
96
+
97
+ 所以摘要先于细节,状态用形式编码而非只给数字。数据里带了置信度、告警、缺口、未覆盖区域时,
98
+ 把它们提到顶部做成入口,点一下就跳到现场。**具体报什么由数据决定**——本技能只要求有入口,
99
+ 没有可报的就不占地方。
100
+
101
+ ## 制品要自陈来历
102
+
103
+ 页脚固定交代:从什么数据生成、什么时候、生成器什么版本,以及**这份东西证明了什么、不证明
104
+ 什么**。最后一项最容易被略过,也最容易避免误读——一张拓扑图证明设备间的连接关系,不证明
105
+ 阀门、控制回路与管径。
106
+
107
+ 只声称观察到的设置支持的结论。
108
+
109
+ ## 反馈要能回来
110
+
111
+ 读者能在制品上标注,导出结构化反馈,且反馈带得回源数据的版本(稳定 id 加源数据摘要)。
112
+ 口头反馈会丢失、会走样、说不清针对哪一版;一份带 id 的 JSON 可以直接被下一轮生成消费。
113
+
114
+ ## 交付前
115
+
116
+ 先跑确定性检查,它比人眼可靠:
117
+
118
+ ```bash
119
+ python3 scripts/check_artifact.py out/topology.html --source data.json --json out/check.json
120
+ ```
121
+
122
+ 它查外部引用、内嵌数据合法性、id 唯一性、悬挂边、来历字段、主题三态。全 PASS 才算可交付。
123
+
124
+ `--json` 写出的那份结果是**证据门的实体供给**:`done_criteria` 直接引用它就行(例如
125
+ 「`out/check.json` 的 `ok` 为 true」),Evaluator跑一条命令便有定论,不必读散文。本技能是加载
126
+ 进上下文的文本,**没有执行力**——能规则化的判据交给脚本与证据门,别交给自觉。
127
+
128
+ 然后在浏览器里看一眼:页面不空白、控制台无错误、计数与源数据一致、缩放平移只影响该影响的
129
+ 东西。**先证明渲染在跑**——`file://` 下某些宿主根本不跑 requestAnimationFrame,所以要用
130
+ http origin 验证,不要只看双击打开的那一份。
131
+
132
+ ## 设计基本功
133
+
134
+ **双主题三态**。读者的浏览器有三种状态:显式浅色、显式深色、以及什么都没标的系统态——最后
135
+ 一种最常见也最常被漏掉。所以在裸 `:root` 里定义完整的浅色令牌,在
136
+ `@media (prefers-color-scheme: dark)` 里以 `:root:not([data-theme="light"])` 为守卫重定义,
137
+ 再在 `:root[data-theme="dark"]` 里重定义一次。组件只用令牌,任何颜色都不要只定义在媒体查询
138
+ 或属性选择器里面。`body` 必须显式上底色,否则会借用宿主的底,深浅一串就露馅。
139
+
140
+ **排版与布局**。正文宽度收在约 65 字符;定一套字号阶梯并守住;数字列用
141
+ `font-variant-numeric: tabular-nums`;宽内容(表格、代码、图)各自 `overflow-x: auto`,
142
+ 别让页面横向滚动。业务字符串用 `textContent` 渲染,不要拼进 `innerHTML`。
143
+
144
+ **避开一眼假的默认**。当前 AI 生成的页面高度雷同:暖奶油底配衬线标题与陶土色点缀、近黑底
145
+ 配一抹荧光绿、紫蓝渐变大标题、Inter 或 Space Grotesk 当万能字体、emoji 当章节标记、什么都
146
+ 居中、到处圆角卡片。用户明确要求某种风格时照做;没要求时,别把这份自由花在这些默认上。
147
+ 从主题本身找线索——工业流程图的调性在控制室与蓝图里,不在营销落地页里。
148
+
149
+ ## 门类
150
+
151
+ - [when-drawing-a-topology.md](references/when-drawing-a-topology.md) — 流程与空间拓扑图:
152
+ 布局参数基线、关系分类、排序与密度、节点内容克制。配 `scripts/render_topology.js`。
153
+ - [when-the-page-must-work-offline.md](references/when-the-page-must-work-offline.md) —
154
+ 内联的具体做法、体积预算、与 App 侧 freeze 的区别。
155
+
156
+ ## 已知盲区
157
+
158
+ 由失败复盘回填。写进来的条件是:作者读过本技能,仍然在同一个地方栽了。
159
+
160
+ - **「不引 CDN」被当成字面禁令,于是在「要双击即开」的需求下被判定为不适用**。被移植的上游
161
+ 方案与本仓的一次微体素沙盘开发都栽在这里。补上本意(运行时可离线)之后,正确解(依赖内联)
162
+ 是自明的。本技能因此不写禁令,写理由。
163
+ - **制品可能在 0 尺寸的容器里初始化**(面板折叠、iframe 未显示、打印预览)。此时
164
+ `getBoundingClientRect()` 返回 0×0,据此算出的取景位移会把内容推到画布外,表现为「打开就
165
+ 是空白,手动点一下适应窗口才好」。取景要等到容器真有尺寸,用 ResizeObserver 兜底。