clearai-dsh 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (154) hide show
  1. package/CHANGELOG.md +26 -0
  2. package/LICENSE +201 -0
  3. package/README.md +138 -0
  4. package/README.zh-CN.md +138 -0
  5. package/bin/clearai.mjs +224 -0
  6. package/brand/README.md +41 -0
  7. package/brand/logo-512-dark.png +0 -0
  8. package/brand/logo-512.png +0 -0
  9. package/brand/logo-lockup-dark.png +0 -0
  10. package/brand/logo-lockup.png +0 -0
  11. package/brand/logo-lockup.svg +12 -0
  12. package/brand/logo-wordmark.svg +6 -0
  13. package/brand/logo.svg +19 -0
  14. package/cordis.patch.yml +39 -0
  15. package/lib/client.js +3071 -0
  16. package/lib/fold.js +1576 -0
  17. package/lib/host.js +605 -0
  18. package/package.json +65 -0
  19. package/presets/clearai/agent.cordis.yml +226 -0
  20. package/presets/clearai/plugins/brain.js +547 -0
  21. package/presets/clearai/plugins/clearai-kernel.js +5485 -0
  22. package/presets/clearai/plugins/ontology.js +306 -0
  23. package/presets/clearai/plugins/prompts.js +312 -0
  24. package/presets/clearai/preset.yml +5 -0
  25. package/presets/clearai/skills/clearai-loop/SKILL.md +89 -0
  26. package/presets/clearai/template/knowledge/README.md +25 -0
  27. package/presets/clearai/template/memory/README.md +34 -0
  28. package/presets/clearai/template/project.md +49 -0
  29. package/presets/clearai/template/skills/README.md +37 -0
  30. package/presets/clearai/template/skills/chart-diagram-qa/SKILL.md +43 -0
  31. package/presets/clearai/template/skills/citation-management/SKILL.md +73 -0
  32. package/presets/clearai/template/skills/citation-management/references/bibtex_formatting.md +908 -0
  33. package/presets/clearai/template/skills/citation-management/references/citation_validation.md +794 -0
  34. package/presets/clearai/template/skills/citation-management/references/google_scholar_search.md +725 -0
  35. package/presets/clearai/template/skills/citation-management/references/metadata_extraction.md +870 -0
  36. package/presets/clearai/template/skills/citation-management/references/pubmed_search.md +839 -0
  37. package/presets/clearai/template/skills/citation-management/scripts/doi_to_bibtex.py +204 -0
  38. package/presets/clearai/template/skills/citation-management/scripts/extract_metadata.py +569 -0
  39. package/presets/clearai/template/skills/citation-management/scripts/format_bibtex.py +349 -0
  40. package/presets/clearai/template/skills/citation-management/scripts/generate_schematic.py +139 -0
  41. package/presets/clearai/template/skills/citation-management/scripts/generate_schematic_ai.py +817 -0
  42. package/presets/clearai/template/skills/citation-management/scripts/search_google_scholar.py +282 -0
  43. package/presets/clearai/template/skills/citation-management/scripts/search_pubmed.py +398 -0
  44. package/presets/clearai/template/skills/citation-management/scripts/validate_citations.py +497 -0
  45. package/presets/clearai/template/skills/data-analysis/SKILL.md +92 -0
  46. package/presets/clearai/template/skills/data-analysis/checklists/readiness_check.md +23 -0
  47. package/presets/clearai/template/skills/data-analysis/templates/analysis_report.md.tpl +63 -0
  48. package/presets/clearai/template/skills/data-analysis/templates/cleaning_rules_draft.yaml.tpl +32 -0
  49. package/presets/clearai/template/skills/data-analysis/templates/data_dictionary.md.tpl +12 -0
  50. package/presets/clearai/template/skills/data-analysis/templates/domain_knowledge_template.md.tpl +316 -0
  51. package/presets/clearai/template/skills/data-analysis/templates/feature_candidates.json.tpl +20 -0
  52. package/presets/clearai/template/skills/data-analysis/templates/quality_scorecard.md.tpl +30 -0
  53. package/presets/clearai/template/skills/data-analysis/workflows/01-data-profiling.md +42 -0
  54. package/presets/clearai/template/skills/data-analysis/workflows/02-quality-audit.md +36 -0
  55. package/presets/clearai/template/skills/data-analysis/workflows/03-physical-correlation.md +25 -0
  56. package/presets/clearai/template/skills/data-analysis/workflows/04-unstructured-mining.md +26 -0
  57. package/presets/clearai/template/skills/data-qa-analysis/SKILL.md +102 -0
  58. package/presets/clearai/template/skills/data-qa-analysis/checklists/readiness_check.md +62 -0
  59. package/presets/clearai/template/skills/data-qa-analysis/templates/best_in_class_report.md.tpl +56 -0
  60. package/presets/clearai/template/skills/data-qa-analysis/templates/cleaning_rules_draft.yaml.tpl +56 -0
  61. package/presets/clearai/template/skills/data-qa-analysis/templates/data_dictionary.md.tpl +13 -0
  62. package/presets/clearai/template/skills/data-qa-analysis/templates/data_source_inventory_and_lineage.md.tpl +146 -0
  63. package/presets/clearai/template/skills/data-qa-analysis/templates/data_status_report.md.tpl +60 -0
  64. package/presets/clearai/template/skills/data-qa-analysis/templates/steady_state_rules.yaml.tpl +41 -0
  65. package/presets/clearai/template/skills/data-qa-analysis/templates/subsystem_registry.md.tpl +101 -0
  66. package/presets/clearai/template/skills/data-qa-analysis/templates/unified_execution_plan.md.tpl +100 -0
  67. package/presets/clearai/template/skills/data-qa-analysis/workflows/01-data-source-inventory-and-lineage.md +194 -0
  68. package/presets/clearai/template/skills/data-qa-analysis/workflows/02-data-alignment-and-tag-semantics.md +122 -0
  69. package/presets/clearai/template/skills/data-qa-analysis/workflows/03-steady-state-identification.md +126 -0
  70. package/presets/clearai/template/skills/data-qa-analysis/workflows/04-consumption-analysis.md +152 -0
  71. package/presets/clearai/template/skills/data-qa-analysis/workflows/05-best-in-class-and-optimization-space.md +78 -0
  72. package/presets/clearai/template/skills/domain-presearch/SKILL.md +131 -0
  73. package/presets/clearai/template/skills/domain-presearch/checklists/domain_checklist.md +24 -0
  74. package/presets/clearai/template/skills/domain-presearch/references/figure_code.md +78 -0
  75. package/presets/clearai/template/skills/domain-presearch/references/strategic_frameworks.md +38 -0
  76. package/presets/clearai/template/skills/exploration-loop/SKILL.md +81 -0
  77. package/presets/clearai/template/skills/exploratory-data-analysis/SKILL.md +77 -0
  78. package/presets/clearai/template/skills/exploratory-data-analysis/references/bioinformatics_genomics_formats.md +664 -0
  79. package/presets/clearai/template/skills/exploratory-data-analysis/references/chemistry_molecular_formats.md +664 -0
  80. package/presets/clearai/template/skills/exploratory-data-analysis/references/general_scientific_formats.md +518 -0
  81. package/presets/clearai/template/skills/exploratory-data-analysis/references/microscopy_imaging_formats.md +620 -0
  82. package/presets/clearai/template/skills/exploratory-data-analysis/references/proteomics_metabolomics_formats.md +517 -0
  83. package/presets/clearai/template/skills/exploratory-data-analysis/references/spectroscopy_analytical_formats.md +633 -0
  84. package/presets/clearai/template/skills/exploratory-data-analysis/scripts/eda_analyzer.py +547 -0
  85. package/presets/clearai/template/skills/hypothesis-generation/SKILL.md +73 -0
  86. package/presets/clearai/template/skills/hypothesis-generation/references/experimental_design_patterns.md +329 -0
  87. package/presets/clearai/template/skills/hypothesis-generation/references/hypothesis_quality_criteria.md +198 -0
  88. package/presets/clearai/template/skills/hypothesis-generation/references/literature_search_strategies.md +622 -0
  89. package/presets/clearai/template/skills/hypothesis-generation/scripts/generate_schematic.py +139 -0
  90. package/presets/clearai/template/skills/hypothesis-generation/scripts/generate_schematic_ai.py +817 -0
  91. package/presets/clearai/template/skills/literature-review/SKILL.md +72 -0
  92. package/presets/clearai/template/skills/literature-review/references/citation_styles.md +166 -0
  93. package/presets/clearai/template/skills/literature-review/references/database_strategies.md +455 -0
  94. package/presets/clearai/template/skills/literature-review/scripts/generate_pdf.py +176 -0
  95. package/presets/clearai/template/skills/literature-review/scripts/generate_schematic.py +139 -0
  96. package/presets/clearai/template/skills/literature-review/scripts/generate_schematic_ai.py +817 -0
  97. package/presets/clearai/template/skills/literature-review/scripts/search_databases.py +303 -0
  98. package/presets/clearai/template/skills/literature-review/scripts/verify_citations.py +221 -0
  99. package/presets/clearai/template/skills/paper-lookup/SKILL.md +59 -0
  100. package/presets/clearai/template/skills/paper-lookup/references/arxiv.md +161 -0
  101. package/presets/clearai/template/skills/paper-lookup/references/biorxiv.md +118 -0
  102. package/presets/clearai/template/skills/paper-lookup/references/core.md +150 -0
  103. package/presets/clearai/template/skills/paper-lookup/references/crossref.md +181 -0
  104. package/presets/clearai/template/skills/paper-lookup/references/medrxiv.md +104 -0
  105. package/presets/clearai/template/skills/paper-lookup/references/openalex.md +174 -0
  106. package/presets/clearai/template/skills/paper-lookup/references/pmc.md +152 -0
  107. package/presets/clearai/template/skills/paper-lookup/references/pubmed.md +124 -0
  108. package/presets/clearai/template/skills/paper-lookup/references/semantic-scholar.md +203 -0
  109. package/presets/clearai/template/skills/paper-lookup/references/unpaywall.md +127 -0
  110. package/presets/clearai/template/skills/process-presearch/SKILL.md +196 -0
  111. package/presets/clearai/template/skills/process-presearch/checklists/process_checklist.md +18 -0
  112. package/presets/clearai/template/skills/process-presearch/references/figure_code.md +107 -0
  113. package/presets/clearai/template/skills/process-presearch/references/source_attribution_example.md +22 -0
  114. package/presets/clearai/template/skills/process-understanding-extraction/SKILL.md +69 -0
  115. package/presets/clearai/template/skills/process-understanding-extraction/checklists/readiness_check.md +34 -0
  116. package/presets/clearai/template/skills/process-understanding-extraction/templates/docx_raw_dump_extractor.py.tpl +132 -0
  117. package/presets/clearai/template/skills/process-understanding-extraction/templates/entity_map_unit_topology.json.tpl +86 -0
  118. package/presets/clearai/template/skills/process-understanding-extraction/templates/process_brief.md.tpl +89 -0
  119. package/presets/clearai/template/skills/process-understanding-extraction/templates/process_brief_builder_from_raw_dump.py.tpl +203 -0
  120. package/presets/clearai/template/skills/process-understanding-extraction/templates/process_flow_mermaid.md.tpl +41 -0
  121. package/presets/clearai/template/skills/process-understanding-extraction/templates/unified_execution_plan.md.tpl +53 -0
  122. package/presets/clearai/template/skills/process-understanding-extraction/workflows/01-process-doc-discovery.md +173 -0
  123. package/presets/clearai/template/skills/process-understanding-extraction/workflows/02-process-understanding-and-diagramming.md +106 -0
  124. package/presets/clearai/template/skills/scientific-brainstorming/SKILL.md +64 -0
  125. package/presets/clearai/template/skills/scientific-brainstorming/references/brainstorming_methods.md +326 -0
  126. package/presets/clearai/template/skills/scientific-critical-thinking/SKILL.md +72 -0
  127. package/presets/clearai/template/skills/scientific-critical-thinking/references/common_biases.md +364 -0
  128. package/presets/clearai/template/skills/scientific-critical-thinking/references/evidence_hierarchy.md +485 -0
  129. package/presets/clearai/template/skills/scientific-critical-thinking/references/experimental_design.md +496 -0
  130. package/presets/clearai/template/skills/scientific-critical-thinking/references/logical_fallacies.md +478 -0
  131. package/presets/clearai/template/skills/scientific-critical-thinking/references/scientific_method.md +169 -0
  132. package/presets/clearai/template/skills/scientific-critical-thinking/references/statistical_pitfalls.md +506 -0
  133. package/presets/clearai/template/skills/skill-creator/SKILL.md +109 -0
  134. package/presets/clearai/template/skills/skill-creator/references/authoring-guide.md +89 -0
  135. package/presets/clearai/template/skills/statistical-analysis/SKILL.md +79 -0
  136. package/presets/clearai/template/skills/statistical-analysis/references/assumptions_and_diagnostics.md +369 -0
  137. package/presets/clearai/template/skills/statistical-analysis/references/bayesian_statistics.md +653 -0
  138. package/presets/clearai/template/skills/statistical-analysis/references/effect_sizes_and_power.md +578 -0
  139. package/presets/clearai/template/skills/statistical-analysis/references/reporting_standards.md +469 -0
  140. package/presets/clearai/template/skills/statistical-analysis/references/test_selection_guide.md +129 -0
  141. package/presets/clearai/template/skills/statistical-analysis/scripts/assumption_checks.py +538 -0
  142. package/presets/clearai/template/skills/web-artifact/SKILL.md +165 -0
  143. package/presets/clearai/template/skills/web-artifact/assets/renderer/renderer.css +229 -0
  144. package/presets/clearai/template/skills/web-artifact/assets/renderer/renderer.js +373 -0
  145. package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/LICENSE +263 -0
  146. package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/UPSTREAM.md +26 -0
  147. package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/elk.bundled.js +6605 -0
  148. package/presets/clearai/template/skills/web-artifact/references/when-drawing-a-topology.md +150 -0
  149. package/presets/clearai/template/skills/web-artifact/references/when-the-page-must-work-offline.md +62 -0
  150. package/presets/clearai/template/skills/web-artifact/scripts/check_artifact.py +167 -0
  151. package/presets/clearai/template/skills/web-artifact/scripts/render_topology.js +272 -0
  152. package/presets/clearai/template/skills/what-if-oracle/LICENSE.txt +5 -0
  153. package/presets/clearai/template/skills/what-if-oracle/SKILL.md +72 -0
  154. package/presets/clearai/template/skills/what-if-oracle/references/scenario-templates.md +154 -0
@@ -0,0 +1,547 @@
1
+ #!/usr/bin/env python3
2
+ """
3
+ Exploratory Data Analysis Analyzer
4
+ Analyzes scientific data files and generates comprehensive markdown reports
5
+ """
6
+
7
+ import os
8
+ import sys
9
+ from pathlib import Path
10
+ from datetime import datetime
11
+ import json
12
+
13
+
14
+ def detect_file_type(filepath):
15
+ """
16
+ Detect the file type based on extension and content.
17
+
18
+ Returns:
19
+ tuple: (extension, file_category, reference_file)
20
+ """
21
+ file_path = Path(filepath)
22
+ extension = file_path.suffix.lower()
23
+ name = file_path.name.lower()
24
+
25
+ # Map extensions to categories and reference files
26
+ extension_map = {
27
+ # Chemistry/Molecular
28
+ 'pdb': ('chemistry_molecular', 'Protein Data Bank'),
29
+ 'cif': ('chemistry_molecular', 'Crystallographic Information File'),
30
+ 'mol': ('chemistry_molecular', 'MDL Molfile'),
31
+ 'mol2': ('chemistry_molecular', 'Tripos Mol2'),
32
+ 'sdf': ('chemistry_molecular', 'Structure Data File'),
33
+ 'xyz': ('chemistry_molecular', 'XYZ Coordinates'),
34
+ 'smi': ('chemistry_molecular', 'SMILES String'),
35
+ 'smiles': ('chemistry_molecular', 'SMILES String'),
36
+ 'pdbqt': ('chemistry_molecular', 'AutoDock PDBQT'),
37
+ 'mae': ('chemistry_molecular', 'Maestro Format'),
38
+ 'gro': ('chemistry_molecular', 'GROMACS Coordinate File'),
39
+ 'log': ('chemistry_molecular', 'Gaussian Log File'),
40
+ 'out': ('chemistry_molecular', 'Quantum Chemistry Output'),
41
+ 'wfn': ('chemistry_molecular', 'Wavefunction Files'),
42
+ 'wfx': ('chemistry_molecular', 'Wavefunction Files'),
43
+ 'fchk': ('chemistry_molecular', 'Gaussian Formatted Checkpoint'),
44
+ 'cube': ('chemistry_molecular', 'Gaussian Cube File'),
45
+ 'dcd': ('chemistry_molecular', 'Binary Trajectory'),
46
+ 'xtc': ('chemistry_molecular', 'Compressed Trajectory'),
47
+ 'trr': ('chemistry_molecular', 'GROMACS Trajectory'),
48
+ 'nc': ('chemistry_molecular', 'Amber NetCDF Trajectory'),
49
+ 'netcdf': ('chemistry_molecular', 'Amber NetCDF Trajectory'),
50
+
51
+ # Bioinformatics/Genomics
52
+ 'fasta': ('bioinformatics_genomics', 'FASTA Format'),
53
+ 'fa': ('bioinformatics_genomics', 'FASTA Format'),
54
+ 'fna': ('bioinformatics_genomics', 'FASTA Format'),
55
+ 'fastq': ('bioinformatics_genomics', 'FASTQ Format'),
56
+ 'fq': ('bioinformatics_genomics', 'FASTQ Format'),
57
+ 'sam': ('bioinformatics_genomics', 'Sequence Alignment/Map'),
58
+ 'bam': ('bioinformatics_genomics', 'Binary Alignment/Map'),
59
+ 'cram': ('bioinformatics_genomics', 'CRAM Format'),
60
+ 'bed': ('bioinformatics_genomics', 'Browser Extensible Data'),
61
+ 'bedgraph': ('bioinformatics_genomics', 'BED with Graph Data'),
62
+ 'bigwig': ('bioinformatics_genomics', 'Binary BigWig'),
63
+ 'bw': ('bioinformatics_genomics', 'Binary BigWig'),
64
+ 'bigbed': ('bioinformatics_genomics', 'Binary BigBed'),
65
+ 'bb': ('bioinformatics_genomics', 'Binary BigBed'),
66
+ 'gff': ('bioinformatics_genomics', 'General Feature Format'),
67
+ 'gff3': ('bioinformatics_genomics', 'General Feature Format'),
68
+ 'gtf': ('bioinformatics_genomics', 'Gene Transfer Format'),
69
+ 'vcf': ('bioinformatics_genomics', 'Variant Call Format'),
70
+ 'bcf': ('bioinformatics_genomics', 'Binary VCF'),
71
+ 'gvcf': ('bioinformatics_genomics', 'Genomic VCF'),
72
+
73
+ # Microscopy/Imaging
74
+ 'tif': ('microscopy_imaging', 'Tagged Image File Format'),
75
+ 'tiff': ('microscopy_imaging', 'Tagged Image File Format'),
76
+ 'nd2': ('microscopy_imaging', 'Nikon NIS-Elements'),
77
+ 'lif': ('microscopy_imaging', 'Leica Image Format'),
78
+ 'czi': ('microscopy_imaging', 'Carl Zeiss Image'),
79
+ 'oib': ('microscopy_imaging', 'Olympus Image Format'),
80
+ 'oif': ('microscopy_imaging', 'Olympus Image Format'),
81
+ 'vsi': ('microscopy_imaging', 'Olympus VSI'),
82
+ 'ims': ('microscopy_imaging', 'Imaris Format'),
83
+ 'lsm': ('microscopy_imaging', 'Zeiss LSM'),
84
+ 'stk': ('microscopy_imaging', 'MetaMorph Stack'),
85
+ 'dv': ('microscopy_imaging', 'DeltaVision'),
86
+ 'mrc': ('microscopy_imaging', 'Medical Research Council'),
87
+ 'dm3': ('microscopy_imaging', 'Gatan Digital Micrograph'),
88
+ 'dm4': ('microscopy_imaging', 'Gatan Digital Micrograph'),
89
+ 'dcm': ('microscopy_imaging', 'DICOM'),
90
+ 'nii': ('microscopy_imaging', 'NIfTI'),
91
+ 'nrrd': ('microscopy_imaging', 'Nearly Raw Raster Data'),
92
+
93
+ # Spectroscopy/Analytical
94
+ 'fid': ('spectroscopy_analytical', 'NMR Free Induction Decay'),
95
+ 'mzml': ('spectroscopy_analytical', 'Mass Spectrometry Markup Language'),
96
+ 'mzxml': ('spectroscopy_analytical', 'Mass Spectrometry XML'),
97
+ 'raw': ('spectroscopy_analytical', 'Vendor Raw Files'),
98
+ 'd': ('spectroscopy_analytical', 'Agilent Data Directory'),
99
+ 'mgf': ('spectroscopy_analytical', 'Mascot Generic Format'),
100
+ 'spc': ('spectroscopy_analytical', 'Galactic SPC'),
101
+ 'jdx': ('spectroscopy_analytical', 'JCAMP-DX'),
102
+ 'jcamp': ('spectroscopy_analytical', 'JCAMP-DX'),
103
+
104
+ # Proteomics/Metabolomics
105
+ 'pepxml': ('proteomics_metabolomics', 'Trans-Proteomic Pipeline Peptide XML'),
106
+ 'protxml': ('proteomics_metabolomics', 'Protein Inference Results'),
107
+ 'mzid': ('proteomics_metabolomics', 'Peptide Identification Format'),
108
+ 'mztab': ('proteomics_metabolomics', 'Proteomics/Metabolomics Tabular Format'),
109
+
110
+ # General Scientific
111
+ 'npy': ('general_scientific', 'NumPy Array'),
112
+ 'npz': ('general_scientific', 'Compressed NumPy Archive'),
113
+ 'csv': ('general_scientific', 'Comma-Separated Values'),
114
+ 'tsv': ('general_scientific', 'Tab-Separated Values'),
115
+ 'xlsx': ('general_scientific', 'Excel Spreadsheets'),
116
+ 'xls': ('general_scientific', 'Excel Spreadsheets'),
117
+ 'json': ('general_scientific', 'JavaScript Object Notation'),
118
+ 'xml': ('general_scientific', 'Extensible Markup Language'),
119
+ 'hdf5': ('general_scientific', 'Hierarchical Data Format 5'),
120
+ 'h5': ('general_scientific', 'Hierarchical Data Format 5'),
121
+ 'h5ad': ('bioinformatics_genomics', 'Anndata Format'),
122
+ 'zarr': ('general_scientific', 'Chunked Array Storage'),
123
+ 'parquet': ('general_scientific', 'Apache Parquet'),
124
+ 'mat': ('general_scientific', 'MATLAB Data'),
125
+ 'fits': ('general_scientific', 'Flexible Image Transport System'),
126
+ }
127
+
128
+ ext_clean = extension.lstrip('.')
129
+ if ext_clean in extension_map:
130
+ category, description = extension_map[ext_clean]
131
+ return ext_clean, category, description
132
+
133
+ return ext_clean, 'unknown', 'Unknown Format'
134
+
135
+
136
+ def get_file_basic_info(filepath):
137
+ """Get basic file information."""
138
+ file_path = Path(filepath)
139
+ stat = file_path.stat()
140
+
141
+ return {
142
+ 'filename': file_path.name,
143
+ 'path': str(file_path.absolute()),
144
+ 'size_bytes': stat.st_size,
145
+ 'size_human': format_bytes(stat.st_size),
146
+ 'modified': datetime.fromtimestamp(stat.st_mtime).isoformat(),
147
+ 'extension': file_path.suffix.lower(),
148
+ }
149
+
150
+
151
+ def format_bytes(size):
152
+ """Convert bytes to human-readable format."""
153
+ for unit in ['B', 'KB', 'MB', 'GB', 'TB']:
154
+ if size < 1024.0:
155
+ return f"{size:.2f} {unit}"
156
+ size /= 1024.0
157
+ return f"{size:.2f} PB"
158
+
159
+
160
+ def load_reference_info(category, extension):
161
+ """
162
+ Load reference information for the file type.
163
+
164
+ Args:
165
+ category: File category (e.g., 'chemistry_molecular')
166
+ extension: File extension
167
+
168
+ Returns:
169
+ dict: Reference information
170
+ """
171
+ # Map categories to reference files
172
+ category_files = {
173
+ 'chemistry_molecular': 'chemistry_molecular_formats.md',
174
+ 'bioinformatics_genomics': 'bioinformatics_genomics_formats.md',
175
+ 'microscopy_imaging': 'microscopy_imaging_formats.md',
176
+ 'spectroscopy_analytical': 'spectroscopy_analytical_formats.md',
177
+ 'proteomics_metabolomics': 'proteomics_metabolomics_formats.md',
178
+ 'general_scientific': 'general_scientific_formats.md',
179
+ }
180
+
181
+ if category not in category_files:
182
+ return None
183
+
184
+ # Get the reference file path
185
+ script_dir = Path(__file__).parent
186
+ ref_file = script_dir.parent / 'references' / category_files[category]
187
+
188
+ if not ref_file.exists():
189
+ return None
190
+
191
+ # Parse the reference file for the specific extension
192
+ # This is a simplified parser - could be more sophisticated
193
+ try:
194
+ with open(ref_file, 'r') as f:
195
+ content = f.read()
196
+
197
+ # Extract section for this file type
198
+ # Look for the extension heading
199
+ import re
200
+ pattern = rf'### \.{extension}[^#]*?(?=###|\Z)'
201
+ match = re.search(pattern, content, re.IGNORECASE | re.DOTALL)
202
+
203
+ if match:
204
+ section = match.group(0)
205
+ return {
206
+ 'raw_section': section,
207
+ 'reference_file': category_files[category]
208
+ }
209
+ except Exception as e:
210
+ print(f"Error loading reference: {e}", file=sys.stderr)
211
+
212
+ return None
213
+
214
+
215
+ def analyze_file(filepath):
216
+ """
217
+ Main analysis function that routes to specific analyzers.
218
+
219
+ Returns:
220
+ dict: Analysis results
221
+ """
222
+ basic_info = get_file_basic_info(filepath)
223
+ extension, category, description = detect_file_type(filepath)
224
+
225
+ analysis = {
226
+ 'basic_info': basic_info,
227
+ 'file_type': {
228
+ 'extension': extension,
229
+ 'category': category,
230
+ 'description': description
231
+ },
232
+ 'reference_info': load_reference_info(category, extension),
233
+ 'data_analysis': {}
234
+ }
235
+
236
+ # Try to perform data-specific analysis based on file type
237
+ try:
238
+ if category == 'general_scientific':
239
+ analysis['data_analysis'] = analyze_general_scientific(filepath, extension)
240
+ elif category == 'bioinformatics_genomics':
241
+ analysis['data_analysis'] = analyze_bioinformatics(filepath, extension)
242
+ elif category == 'microscopy_imaging':
243
+ analysis['data_analysis'] = analyze_imaging(filepath, extension)
244
+ # Add more specific analyzers as needed
245
+ except Exception as e:
246
+ analysis['data_analysis']['error'] = str(e)
247
+
248
+ return analysis
249
+
250
+
251
+ def analyze_general_scientific(filepath, extension):
252
+ """Analyze general scientific data formats."""
253
+ results = {}
254
+
255
+ try:
256
+ if extension in ['npy']:
257
+ import numpy as np
258
+ data = np.load(filepath)
259
+ results = {
260
+ 'shape': data.shape,
261
+ 'dtype': str(data.dtype),
262
+ 'size': data.size,
263
+ 'ndim': data.ndim,
264
+ 'statistics': {
265
+ 'min': float(np.min(data)) if np.issubdtype(data.dtype, np.number) else None,
266
+ 'max': float(np.max(data)) if np.issubdtype(data.dtype, np.number) else None,
267
+ 'mean': float(np.mean(data)) if np.issubdtype(data.dtype, np.number) else None,
268
+ 'std': float(np.std(data)) if np.issubdtype(data.dtype, np.number) else None,
269
+ }
270
+ }
271
+
272
+ elif extension in ['npz']:
273
+ import numpy as np
274
+ data = np.load(filepath)
275
+ results = {
276
+ 'arrays': list(data.files),
277
+ 'array_count': len(data.files),
278
+ 'array_shapes': {name: data[name].shape for name in data.files}
279
+ }
280
+
281
+ elif extension in ['csv', 'tsv']:
282
+ import pandas as pd
283
+ sep = '\t' if extension == 'tsv' else ','
284
+ df = pd.read_csv(filepath, sep=sep, nrows=10000) # Sample first 10k rows
285
+
286
+ results = {
287
+ 'shape': df.shape,
288
+ 'columns': list(df.columns),
289
+ 'dtypes': {col: str(dtype) for col, dtype in df.dtypes.items()},
290
+ 'missing_values': df.isnull().sum().to_dict(),
291
+ 'summary_statistics': df.describe().to_dict() if len(df.select_dtypes(include='number').columns) > 0 else {}
292
+ }
293
+
294
+ elif extension in ['json']:
295
+ with open(filepath, 'r') as f:
296
+ data = json.load(f)
297
+
298
+ results = {
299
+ 'type': type(data).__name__,
300
+ 'keys': list(data.keys()) if isinstance(data, dict) else None,
301
+ 'length': len(data) if isinstance(data, (list, dict)) else None
302
+ }
303
+
304
+ elif extension in ['h5', 'hdf5']:
305
+ import h5py
306
+ with h5py.File(filepath, 'r') as f:
307
+ def get_structure(group, prefix=''):
308
+ items = {}
309
+ for key in group.keys():
310
+ path = f"{prefix}/{key}"
311
+ if isinstance(group[key], h5py.Dataset):
312
+ items[path] = {
313
+ 'type': 'dataset',
314
+ 'shape': group[key].shape,
315
+ 'dtype': str(group[key].dtype)
316
+ }
317
+ elif isinstance(group[key], h5py.Group):
318
+ items[path] = {'type': 'group'}
319
+ items.update(get_structure(group[key], path))
320
+ return items
321
+
322
+ results = {
323
+ 'structure': get_structure(f),
324
+ 'attributes': dict(f.attrs)
325
+ }
326
+
327
+ except ImportError as e:
328
+ results['error'] = f"Required library not installed: {e}"
329
+ except Exception as e:
330
+ results['error'] = f"Analysis error: {e}"
331
+
332
+ return results
333
+
334
+
335
+ def analyze_bioinformatics(filepath, extension):
336
+ """Analyze bioinformatics/genomics formats."""
337
+ results = {}
338
+
339
+ try:
340
+ if extension in ['fasta', 'fa', 'fna']:
341
+ from Bio import SeqIO
342
+ sequences = list(SeqIO.parse(filepath, 'fasta'))
343
+ lengths = [len(seq) for seq in sequences]
344
+
345
+ results = {
346
+ 'sequence_count': len(sequences),
347
+ 'total_length': sum(lengths),
348
+ 'mean_length': sum(lengths) / len(lengths) if lengths else 0,
349
+ 'min_length': min(lengths) if lengths else 0,
350
+ 'max_length': max(lengths) if lengths else 0,
351
+ 'sequence_ids': [seq.id for seq in sequences[:10]] # First 10
352
+ }
353
+
354
+ elif extension in ['fastq', 'fq']:
355
+ from Bio import SeqIO
356
+ sequences = []
357
+ for i, seq in enumerate(SeqIO.parse(filepath, 'fastq')):
358
+ sequences.append(seq)
359
+ if i >= 9999: # Sample first 10k
360
+ break
361
+
362
+ lengths = [len(seq) for seq in sequences]
363
+ qualities = [sum(seq.letter_annotations['phred_quality']) / len(seq) for seq in sequences]
364
+
365
+ results = {
366
+ 'read_count_sampled': len(sequences),
367
+ 'mean_length': sum(lengths) / len(lengths) if lengths else 0,
368
+ 'mean_quality': sum(qualities) / len(qualities) if qualities else 0,
369
+ 'min_length': min(lengths) if lengths else 0,
370
+ 'max_length': max(lengths) if lengths else 0,
371
+ }
372
+
373
+ except ImportError as e:
374
+ results['error'] = f"Required library not installed (try: pip install biopython): {e}"
375
+ except Exception as e:
376
+ results['error'] = f"Analysis error: {e}"
377
+
378
+ return results
379
+
380
+
381
+ def analyze_imaging(filepath, extension):
382
+ """Analyze microscopy/imaging formats."""
383
+ results = {}
384
+
385
+ try:
386
+ if extension in ['tif', 'tiff', 'png', 'jpg', 'jpeg']:
387
+ from PIL import Image
388
+ import numpy as np
389
+
390
+ img = Image.open(filepath)
391
+ img_array = np.array(img)
392
+
393
+ results = {
394
+ 'size': img.size,
395
+ 'mode': img.mode,
396
+ 'format': img.format,
397
+ 'shape': img_array.shape,
398
+ 'dtype': str(img_array.dtype),
399
+ 'value_range': [int(img_array.min()), int(img_array.max())],
400
+ 'mean_intensity': float(img_array.mean()),
401
+ }
402
+
403
+ # Check for multi-page TIFF
404
+ if extension in ['tif', 'tiff']:
405
+ try:
406
+ frame_count = 0
407
+ while True:
408
+ img.seek(frame_count)
409
+ frame_count += 1
410
+ except EOFError:
411
+ results['page_count'] = frame_count
412
+
413
+ except ImportError as e:
414
+ results['error'] = f"Required library not installed (try: pip install pillow): {e}"
415
+ except Exception as e:
416
+ results['error'] = f"Analysis error: {e}"
417
+
418
+ return results
419
+
420
+
421
+ def generate_markdown_report(analysis, output_path=None):
422
+ """
423
+ Generate a comprehensive markdown report from analysis results.
424
+
425
+ Args:
426
+ analysis: Analysis results dictionary
427
+ output_path: Path to save the report (if None, prints to stdout)
428
+ """
429
+ lines = []
430
+
431
+ # Title
432
+ filename = analysis['basic_info']['filename']
433
+ lines.append(f"# Exploratory Data Analysis Report: {filename}\n")
434
+ lines.append(f"**Generated:** {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}\n")
435
+ lines.append("---\n")
436
+
437
+ # Basic Information
438
+ lines.append("## Basic Information\n")
439
+ basic = analysis['basic_info']
440
+ lines.append(f"- **Filename:** `{basic['filename']}`")
441
+ lines.append(f"- **Full Path:** `{basic['path']}`")
442
+ lines.append(f"- **File Size:** {basic['size_human']} ({basic['size_bytes']:,} bytes)")
443
+ lines.append(f"- **Last Modified:** {basic['modified']}")
444
+ lines.append(f"- **Extension:** `.{analysis['file_type']['extension']}`\n")
445
+
446
+ # File Type Information
447
+ lines.append("## File Type\n")
448
+ ft = analysis['file_type']
449
+ lines.append(f"- **Category:** {ft['category'].replace('_', ' ').title()}")
450
+ lines.append(f"- **Description:** {ft['description']}\n")
451
+
452
+ # Reference Information
453
+ if analysis.get('reference_info'):
454
+ lines.append("## Format Reference\n")
455
+ ref = analysis['reference_info']
456
+ if 'raw_section' in ref:
457
+ lines.append(ref['raw_section'])
458
+ lines.append(f"\n*Reference: {ref['reference_file']}*\n")
459
+
460
+ # Data Analysis
461
+ if analysis.get('data_analysis'):
462
+ lines.append("## Data Analysis\n")
463
+ data = analysis['data_analysis']
464
+
465
+ if 'error' in data:
466
+ lines.append(f"⚠️ **Analysis Error:** {data['error']}\n")
467
+ else:
468
+ # Format the data analysis based on what's present
469
+ lines.append("### Summary Statistics\n")
470
+ lines.append("```json")
471
+ lines.append(json.dumps(data, indent=2, default=str))
472
+ lines.append("```\n")
473
+
474
+ # Recommendations
475
+ lines.append("## Recommendations for Further Analysis\n")
476
+ lines.append(f"Based on the file type (`.{analysis['file_type']['extension']}`), consider the following analyses:\n")
477
+
478
+ # Add specific recommendations based on category
479
+ category = analysis['file_type']['category']
480
+ if category == 'general_scientific':
481
+ lines.append("- Statistical distribution analysis")
482
+ lines.append("- Missing value imputation strategies")
483
+ lines.append("- Correlation analysis between variables")
484
+ lines.append("- Outlier detection and handling")
485
+ lines.append("- Dimensionality reduction (PCA, t-SNE)")
486
+ elif category == 'bioinformatics_genomics':
487
+ lines.append("- Sequence quality control and filtering")
488
+ lines.append("- GC content analysis")
489
+ lines.append("- Read alignment and mapping statistics")
490
+ lines.append("- Variant calling and annotation")
491
+ lines.append("- Differential expression analysis")
492
+ elif category == 'microscopy_imaging':
493
+ lines.append("- Image quality assessment")
494
+ lines.append("- Background correction and normalization")
495
+ lines.append("- Segmentation and object detection")
496
+ lines.append("- Colocalization analysis")
497
+ lines.append("- Intensity measurements and quantification")
498
+
499
+ lines.append("")
500
+
501
+ # Footer
502
+ lines.append("---")
503
+ lines.append("*This report was generated by the exploratory-data-analysis skill.*")
504
+
505
+ report = '\n'.join(lines)
506
+
507
+ if output_path:
508
+ with open(output_path, 'w') as f:
509
+ f.write(report)
510
+ print(f"Report saved to: {output_path}")
511
+ else:
512
+ print(report)
513
+
514
+ return report
515
+
516
+
517
+ def main():
518
+ """Main CLI interface."""
519
+ if len(sys.argv) < 2:
520
+ print("Usage: python eda_analyzer.py <filepath> [output.md]")
521
+ print(" filepath: Path to the data file to analyze")
522
+ print(" output.md: Optional output path for markdown report")
523
+ sys.exit(1)
524
+
525
+ filepath = sys.argv[1]
526
+ output_path = sys.argv[2] if len(sys.argv) > 2 else None
527
+
528
+ if not os.path.exists(filepath):
529
+ print(f"Error: File not found: {filepath}")
530
+ sys.exit(1)
531
+
532
+ # If no output path specified, use the input filename
533
+ if output_path is None:
534
+ input_path = Path(filepath)
535
+ output_path = input_path.parent / f"{input_path.stem}_eda_report.md"
536
+
537
+ print(f"Analyzing: {filepath}")
538
+ analysis = analyze_file(filepath)
539
+
540
+ print(f"\nGenerating report...")
541
+ generate_markdown_report(analysis, output_path)
542
+
543
+ print(f"\n✓ Analysis complete!")
544
+
545
+
546
+ if __name__ == '__main__':
547
+ main()
@@ -0,0 +1,73 @@
1
+ ---
2
+ name: hypothesis-generation
3
+ description: |
4
+ 【假设形成·可证伪】从观察/文献/预研结论形成竞争假设与验证计划。适用:预研后需明确「待验证什么」、CM 分析中的可验证假设、根因候选排序。不适用:纯发散脑暴(用 scientific-brainstorming);已有数据直接做统计检验(用 statistical-analysis)。
5
+ license: MIT license
6
+ metadata:
7
+ version: 1.0-clearai
8
+ skill-author: K-Dense Inc. (adapted for ClearAI)
9
+ tier: system
10
+ origin: template
11
+ created_at: '2026-06-12T02:47:05.408447+00:00'
12
+ ---
13
+
14
+ # 科学假设形成 Skill(ClearAI 版)
15
+
16
+ ## 使用边界
17
+
18
+ - **适用**:有初步观察、预研结论或数据模式,需形成 3–5 个可区分、可证伪的假设。
19
+ - **不适用**:尚无观察的早期探索 → `scientific-brainstorming`;自动化批量假设挖掘 → P1 `hypogenic`。
20
+
21
+ ## ClearAI 工具与路径映射
22
+
23
+ - 文献 → `web_search` + `paper-lookup`(学术补充)
24
+ - 数据探查 → `read` / `bash`(pandas 摘要)
25
+ - 假设文档 → `write` 到 `lab/knowledge/hypotheses.md`
26
+ - 验证计划中的脚本 → `lab/scripts/`
27
+ - 经验回写 → `clear/memory/hypothesis_generation_lessons.md`
28
+
29
+ ## 工作流
30
+
31
+ ### 1. 澄清现象
32
+
33
+ - 核心观察是什么?范围与约束?已知 vs 未知?
34
+
35
+ ### 2. 文献与事实检索
36
+
37
+ - `web_search` 获取行业实践与学术线索
38
+ - 需要 DOI/论文细节时加载 `paper-lookup`
39
+
40
+ 检索策略见 `references/literature_search_strategies.md`。
41
+
42
+ ### 3. 综合证据
43
+
44
+ - 当前共识、冲突证据、空白点
45
+
46
+ ### 4. 生成竞争假设(3–5 个)
47
+
48
+ 每条假设须含:
49
+ - **机制解释**(非仅描述)
50
+ - **可观测预测**
51
+ - **证伪条件**(什么结果可推翻它)
52
+
53
+ ### 5. 质量评估
54
+
55
+ 用 `references/hypothesis_quality_criteria.md` 评估:可检验性、可证伪性、简约性、解释力。
56
+
57
+ ### 6. 设计验证步骤
58
+
59
+ - 需要什么数据/实验/统计检验?
60
+ - 下一步 Skill:`exploratory-data-analysis` → `statistical-analysis` 或 `data-analysis`
61
+
62
+ ## 交付物 `lab/knowledge/hypotheses.md`
63
+
64
+ ```markdown
65
+ ## H1: ...
66
+ - 机制:...
67
+ - 预测:...
68
+ - 证伪:...
69
+ - 验证动作:...
70
+ - 优先级:高/中/低
71
+ ```
72
+
73
+ 与 NEXT「可验证假设」口径对齐:CM 分析、diagnosis 根因分析可直接引用本文档。