clearai-dsh 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +26 -0
- package/LICENSE +201 -0
- package/README.md +138 -0
- package/README.zh-CN.md +138 -0
- package/bin/clearai.mjs +224 -0
- package/brand/README.md +41 -0
- package/brand/logo-512-dark.png +0 -0
- package/brand/logo-512.png +0 -0
- package/brand/logo-lockup-dark.png +0 -0
- package/brand/logo-lockup.png +0 -0
- package/brand/logo-lockup.svg +12 -0
- package/brand/logo-wordmark.svg +6 -0
- package/brand/logo.svg +19 -0
- package/cordis.patch.yml +39 -0
- package/lib/client.js +3071 -0
- package/lib/fold.js +1576 -0
- package/lib/host.js +605 -0
- package/package.json +65 -0
- package/presets/clearai/agent.cordis.yml +226 -0
- package/presets/clearai/plugins/brain.js +547 -0
- package/presets/clearai/plugins/clearai-kernel.js +5485 -0
- package/presets/clearai/plugins/ontology.js +306 -0
- package/presets/clearai/plugins/prompts.js +312 -0
- package/presets/clearai/preset.yml +5 -0
- package/presets/clearai/skills/clearai-loop/SKILL.md +89 -0
- package/presets/clearai/template/knowledge/README.md +25 -0
- package/presets/clearai/template/memory/README.md +34 -0
- package/presets/clearai/template/project.md +49 -0
- package/presets/clearai/template/skills/README.md +37 -0
- package/presets/clearai/template/skills/chart-diagram-qa/SKILL.md +43 -0
- package/presets/clearai/template/skills/citation-management/SKILL.md +73 -0
- package/presets/clearai/template/skills/citation-management/references/bibtex_formatting.md +908 -0
- package/presets/clearai/template/skills/citation-management/references/citation_validation.md +794 -0
- package/presets/clearai/template/skills/citation-management/references/google_scholar_search.md +725 -0
- package/presets/clearai/template/skills/citation-management/references/metadata_extraction.md +870 -0
- package/presets/clearai/template/skills/citation-management/references/pubmed_search.md +839 -0
- package/presets/clearai/template/skills/citation-management/scripts/doi_to_bibtex.py +204 -0
- package/presets/clearai/template/skills/citation-management/scripts/extract_metadata.py +569 -0
- package/presets/clearai/template/skills/citation-management/scripts/format_bibtex.py +349 -0
- package/presets/clearai/template/skills/citation-management/scripts/generate_schematic.py +139 -0
- package/presets/clearai/template/skills/citation-management/scripts/generate_schematic_ai.py +817 -0
- package/presets/clearai/template/skills/citation-management/scripts/search_google_scholar.py +282 -0
- package/presets/clearai/template/skills/citation-management/scripts/search_pubmed.py +398 -0
- package/presets/clearai/template/skills/citation-management/scripts/validate_citations.py +497 -0
- package/presets/clearai/template/skills/data-analysis/SKILL.md +92 -0
- package/presets/clearai/template/skills/data-analysis/checklists/readiness_check.md +23 -0
- package/presets/clearai/template/skills/data-analysis/templates/analysis_report.md.tpl +63 -0
- package/presets/clearai/template/skills/data-analysis/templates/cleaning_rules_draft.yaml.tpl +32 -0
- package/presets/clearai/template/skills/data-analysis/templates/data_dictionary.md.tpl +12 -0
- package/presets/clearai/template/skills/data-analysis/templates/domain_knowledge_template.md.tpl +316 -0
- package/presets/clearai/template/skills/data-analysis/templates/feature_candidates.json.tpl +20 -0
- package/presets/clearai/template/skills/data-analysis/templates/quality_scorecard.md.tpl +30 -0
- package/presets/clearai/template/skills/data-analysis/workflows/01-data-profiling.md +42 -0
- package/presets/clearai/template/skills/data-analysis/workflows/02-quality-audit.md +36 -0
- package/presets/clearai/template/skills/data-analysis/workflows/03-physical-correlation.md +25 -0
- package/presets/clearai/template/skills/data-analysis/workflows/04-unstructured-mining.md +26 -0
- package/presets/clearai/template/skills/data-qa-analysis/SKILL.md +102 -0
- package/presets/clearai/template/skills/data-qa-analysis/checklists/readiness_check.md +62 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/best_in_class_report.md.tpl +56 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/cleaning_rules_draft.yaml.tpl +56 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/data_dictionary.md.tpl +13 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/data_source_inventory_and_lineage.md.tpl +146 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/data_status_report.md.tpl +60 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/steady_state_rules.yaml.tpl +41 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/subsystem_registry.md.tpl +101 -0
- package/presets/clearai/template/skills/data-qa-analysis/templates/unified_execution_plan.md.tpl +100 -0
- package/presets/clearai/template/skills/data-qa-analysis/workflows/01-data-source-inventory-and-lineage.md +194 -0
- package/presets/clearai/template/skills/data-qa-analysis/workflows/02-data-alignment-and-tag-semantics.md +122 -0
- package/presets/clearai/template/skills/data-qa-analysis/workflows/03-steady-state-identification.md +126 -0
- package/presets/clearai/template/skills/data-qa-analysis/workflows/04-consumption-analysis.md +152 -0
- package/presets/clearai/template/skills/data-qa-analysis/workflows/05-best-in-class-and-optimization-space.md +78 -0
- package/presets/clearai/template/skills/domain-presearch/SKILL.md +131 -0
- package/presets/clearai/template/skills/domain-presearch/checklists/domain_checklist.md +24 -0
- package/presets/clearai/template/skills/domain-presearch/references/figure_code.md +78 -0
- package/presets/clearai/template/skills/domain-presearch/references/strategic_frameworks.md +38 -0
- package/presets/clearai/template/skills/exploration-loop/SKILL.md +81 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/SKILL.md +77 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/references/bioinformatics_genomics_formats.md +664 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/references/chemistry_molecular_formats.md +664 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/references/general_scientific_formats.md +518 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/references/microscopy_imaging_formats.md +620 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/references/proteomics_metabolomics_formats.md +517 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/references/spectroscopy_analytical_formats.md +633 -0
- package/presets/clearai/template/skills/exploratory-data-analysis/scripts/eda_analyzer.py +547 -0
- package/presets/clearai/template/skills/hypothesis-generation/SKILL.md +73 -0
- package/presets/clearai/template/skills/hypothesis-generation/references/experimental_design_patterns.md +329 -0
- package/presets/clearai/template/skills/hypothesis-generation/references/hypothesis_quality_criteria.md +198 -0
- package/presets/clearai/template/skills/hypothesis-generation/references/literature_search_strategies.md +622 -0
- package/presets/clearai/template/skills/hypothesis-generation/scripts/generate_schematic.py +139 -0
- package/presets/clearai/template/skills/hypothesis-generation/scripts/generate_schematic_ai.py +817 -0
- package/presets/clearai/template/skills/literature-review/SKILL.md +72 -0
- package/presets/clearai/template/skills/literature-review/references/citation_styles.md +166 -0
- package/presets/clearai/template/skills/literature-review/references/database_strategies.md +455 -0
- package/presets/clearai/template/skills/literature-review/scripts/generate_pdf.py +176 -0
- package/presets/clearai/template/skills/literature-review/scripts/generate_schematic.py +139 -0
- package/presets/clearai/template/skills/literature-review/scripts/generate_schematic_ai.py +817 -0
- package/presets/clearai/template/skills/literature-review/scripts/search_databases.py +303 -0
- package/presets/clearai/template/skills/literature-review/scripts/verify_citations.py +221 -0
- package/presets/clearai/template/skills/paper-lookup/SKILL.md +59 -0
- package/presets/clearai/template/skills/paper-lookup/references/arxiv.md +161 -0
- package/presets/clearai/template/skills/paper-lookup/references/biorxiv.md +118 -0
- package/presets/clearai/template/skills/paper-lookup/references/core.md +150 -0
- package/presets/clearai/template/skills/paper-lookup/references/crossref.md +181 -0
- package/presets/clearai/template/skills/paper-lookup/references/medrxiv.md +104 -0
- package/presets/clearai/template/skills/paper-lookup/references/openalex.md +174 -0
- package/presets/clearai/template/skills/paper-lookup/references/pmc.md +152 -0
- package/presets/clearai/template/skills/paper-lookup/references/pubmed.md +124 -0
- package/presets/clearai/template/skills/paper-lookup/references/semantic-scholar.md +203 -0
- package/presets/clearai/template/skills/paper-lookup/references/unpaywall.md +127 -0
- package/presets/clearai/template/skills/process-presearch/SKILL.md +196 -0
- package/presets/clearai/template/skills/process-presearch/checklists/process_checklist.md +18 -0
- package/presets/clearai/template/skills/process-presearch/references/figure_code.md +107 -0
- package/presets/clearai/template/skills/process-presearch/references/source_attribution_example.md +22 -0
- package/presets/clearai/template/skills/process-understanding-extraction/SKILL.md +69 -0
- package/presets/clearai/template/skills/process-understanding-extraction/checklists/readiness_check.md +34 -0
- package/presets/clearai/template/skills/process-understanding-extraction/templates/docx_raw_dump_extractor.py.tpl +132 -0
- package/presets/clearai/template/skills/process-understanding-extraction/templates/entity_map_unit_topology.json.tpl +86 -0
- package/presets/clearai/template/skills/process-understanding-extraction/templates/process_brief.md.tpl +89 -0
- package/presets/clearai/template/skills/process-understanding-extraction/templates/process_brief_builder_from_raw_dump.py.tpl +203 -0
- package/presets/clearai/template/skills/process-understanding-extraction/templates/process_flow_mermaid.md.tpl +41 -0
- package/presets/clearai/template/skills/process-understanding-extraction/templates/unified_execution_plan.md.tpl +53 -0
- package/presets/clearai/template/skills/process-understanding-extraction/workflows/01-process-doc-discovery.md +173 -0
- package/presets/clearai/template/skills/process-understanding-extraction/workflows/02-process-understanding-and-diagramming.md +106 -0
- package/presets/clearai/template/skills/scientific-brainstorming/SKILL.md +64 -0
- package/presets/clearai/template/skills/scientific-brainstorming/references/brainstorming_methods.md +326 -0
- package/presets/clearai/template/skills/scientific-critical-thinking/SKILL.md +72 -0
- package/presets/clearai/template/skills/scientific-critical-thinking/references/common_biases.md +364 -0
- package/presets/clearai/template/skills/scientific-critical-thinking/references/evidence_hierarchy.md +485 -0
- package/presets/clearai/template/skills/scientific-critical-thinking/references/experimental_design.md +496 -0
- package/presets/clearai/template/skills/scientific-critical-thinking/references/logical_fallacies.md +478 -0
- package/presets/clearai/template/skills/scientific-critical-thinking/references/scientific_method.md +169 -0
- package/presets/clearai/template/skills/scientific-critical-thinking/references/statistical_pitfalls.md +506 -0
- package/presets/clearai/template/skills/skill-creator/SKILL.md +109 -0
- package/presets/clearai/template/skills/skill-creator/references/authoring-guide.md +89 -0
- package/presets/clearai/template/skills/statistical-analysis/SKILL.md +79 -0
- package/presets/clearai/template/skills/statistical-analysis/references/assumptions_and_diagnostics.md +369 -0
- package/presets/clearai/template/skills/statistical-analysis/references/bayesian_statistics.md +653 -0
- package/presets/clearai/template/skills/statistical-analysis/references/effect_sizes_and_power.md +578 -0
- package/presets/clearai/template/skills/statistical-analysis/references/reporting_standards.md +469 -0
- package/presets/clearai/template/skills/statistical-analysis/references/test_selection_guide.md +129 -0
- package/presets/clearai/template/skills/statistical-analysis/scripts/assumption_checks.py +538 -0
- package/presets/clearai/template/skills/web-artifact/SKILL.md +165 -0
- package/presets/clearai/template/skills/web-artifact/assets/renderer/renderer.css +229 -0
- package/presets/clearai/template/skills/web-artifact/assets/renderer/renderer.js +373 -0
- package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/LICENSE +263 -0
- package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/UPSTREAM.md +26 -0
- package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/elk.bundled.js +6605 -0
- package/presets/clearai/template/skills/web-artifact/references/when-drawing-a-topology.md +150 -0
- package/presets/clearai/template/skills/web-artifact/references/when-the-page-must-work-offline.md +62 -0
- package/presets/clearai/template/skills/web-artifact/scripts/check_artifact.py +167 -0
- package/presets/clearai/template/skills/web-artifact/scripts/render_topology.js +272 -0
- package/presets/clearai/template/skills/what-if-oracle/LICENSE.txt +5 -0
- package/presets/clearai/template/skills/what-if-oracle/SKILL.md +72 -0
- package/presets/clearai/template/skills/what-if-oracle/references/scenario-templates.md +154 -0
|
@@ -0,0 +1,497 @@
|
|
|
1
|
+
#!/usr/bin/env python3
|
|
2
|
+
"""
|
|
3
|
+
Citation Validation Tool
|
|
4
|
+
Validate BibTeX files for accuracy, completeness, and format compliance.
|
|
5
|
+
"""
|
|
6
|
+
|
|
7
|
+
import sys
|
|
8
|
+
import re
|
|
9
|
+
import requests
|
|
10
|
+
import argparse
|
|
11
|
+
import json
|
|
12
|
+
from typing import Dict, List, Tuple, Optional
|
|
13
|
+
from collections import defaultdict
|
|
14
|
+
|
|
15
|
+
class CitationValidator:
|
|
16
|
+
"""Validate BibTeX entries for errors and inconsistencies."""
|
|
17
|
+
|
|
18
|
+
def __init__(self):
|
|
19
|
+
self.session = requests.Session()
|
|
20
|
+
self.session.headers.update({
|
|
21
|
+
'User-Agent': 'CitationValidator/1.0 (Citation Management Tool)'
|
|
22
|
+
})
|
|
23
|
+
|
|
24
|
+
# Required fields by entry type
|
|
25
|
+
self.required_fields = {
|
|
26
|
+
'article': ['author', 'title', 'journal', 'year'],
|
|
27
|
+
'book': ['title', 'publisher', 'year'], # author OR editor
|
|
28
|
+
'inproceedings': ['author', 'title', 'booktitle', 'year'],
|
|
29
|
+
'incollection': ['author', 'title', 'booktitle', 'publisher', 'year'],
|
|
30
|
+
'phdthesis': ['author', 'title', 'school', 'year'],
|
|
31
|
+
'mastersthesis': ['author', 'title', 'school', 'year'],
|
|
32
|
+
'techreport': ['author', 'title', 'institution', 'year'],
|
|
33
|
+
'misc': ['title', 'year']
|
|
34
|
+
}
|
|
35
|
+
|
|
36
|
+
# Recommended fields
|
|
37
|
+
self.recommended_fields = {
|
|
38
|
+
'article': ['volume', 'pages', 'doi'],
|
|
39
|
+
'book': ['isbn'],
|
|
40
|
+
'inproceedings': ['pages'],
|
|
41
|
+
}
|
|
42
|
+
|
|
43
|
+
def parse_bibtex_file(self, filepath: str) -> List[Dict]:
|
|
44
|
+
"""
|
|
45
|
+
Parse BibTeX file and extract entries.
|
|
46
|
+
|
|
47
|
+
Args:
|
|
48
|
+
filepath: Path to BibTeX file
|
|
49
|
+
|
|
50
|
+
Returns:
|
|
51
|
+
List of entry dictionaries
|
|
52
|
+
"""
|
|
53
|
+
try:
|
|
54
|
+
with open(filepath, 'r', encoding='utf-8') as f:
|
|
55
|
+
content = f.read()
|
|
56
|
+
except Exception as e:
|
|
57
|
+
print(f'Error reading file: {e}', file=sys.stderr)
|
|
58
|
+
return []
|
|
59
|
+
|
|
60
|
+
entries = []
|
|
61
|
+
|
|
62
|
+
# Match BibTeX entries
|
|
63
|
+
pattern = r'@(\w+)\s*\{\s*([^,\s]+)\s*,(.*?)\n\}'
|
|
64
|
+
matches = re.finditer(pattern, content, re.DOTALL | re.IGNORECASE)
|
|
65
|
+
|
|
66
|
+
for match in matches:
|
|
67
|
+
entry_type = match.group(1).lower()
|
|
68
|
+
citation_key = match.group(2).strip()
|
|
69
|
+
fields_text = match.group(3)
|
|
70
|
+
|
|
71
|
+
# Parse fields
|
|
72
|
+
fields = {}
|
|
73
|
+
field_pattern = r'(\w+)\s*=\s*\{([^}]*)\}|(\w+)\s*=\s*"([^"]*)"'
|
|
74
|
+
field_matches = re.finditer(field_pattern, fields_text)
|
|
75
|
+
|
|
76
|
+
for field_match in field_matches:
|
|
77
|
+
if field_match.group(1):
|
|
78
|
+
field_name = field_match.group(1).lower()
|
|
79
|
+
field_value = field_match.group(2)
|
|
80
|
+
else:
|
|
81
|
+
field_name = field_match.group(3).lower()
|
|
82
|
+
field_value = field_match.group(4)
|
|
83
|
+
|
|
84
|
+
fields[field_name] = field_value.strip()
|
|
85
|
+
|
|
86
|
+
entries.append({
|
|
87
|
+
'type': entry_type,
|
|
88
|
+
'key': citation_key,
|
|
89
|
+
'fields': fields,
|
|
90
|
+
'raw': match.group(0)
|
|
91
|
+
})
|
|
92
|
+
|
|
93
|
+
return entries
|
|
94
|
+
|
|
95
|
+
def validate_entry(self, entry: Dict) -> Tuple[List[Dict], List[Dict]]:
|
|
96
|
+
"""
|
|
97
|
+
Validate a single BibTeX entry.
|
|
98
|
+
|
|
99
|
+
Args:
|
|
100
|
+
entry: Entry dictionary
|
|
101
|
+
|
|
102
|
+
Returns:
|
|
103
|
+
Tuple of (errors, warnings)
|
|
104
|
+
"""
|
|
105
|
+
errors = []
|
|
106
|
+
warnings = []
|
|
107
|
+
|
|
108
|
+
entry_type = entry['type']
|
|
109
|
+
key = entry['key']
|
|
110
|
+
fields = entry['fields']
|
|
111
|
+
|
|
112
|
+
# Check required fields
|
|
113
|
+
if entry_type in self.required_fields:
|
|
114
|
+
for req_field in self.required_fields[entry_type]:
|
|
115
|
+
if req_field not in fields or not fields[req_field]:
|
|
116
|
+
# Special case: book can have author OR editor
|
|
117
|
+
if entry_type == 'book' and req_field == 'author':
|
|
118
|
+
if 'editor' not in fields or not fields['editor']:
|
|
119
|
+
errors.append({
|
|
120
|
+
'type': 'missing_required_field',
|
|
121
|
+
'field': 'author or editor',
|
|
122
|
+
'severity': 'high',
|
|
123
|
+
'message': f'Entry {key}: Missing required field "author" or "editor"'
|
|
124
|
+
})
|
|
125
|
+
else:
|
|
126
|
+
errors.append({
|
|
127
|
+
'type': 'missing_required_field',
|
|
128
|
+
'field': req_field,
|
|
129
|
+
'severity': 'high',
|
|
130
|
+
'message': f'Entry {key}: Missing required field "{req_field}"'
|
|
131
|
+
})
|
|
132
|
+
|
|
133
|
+
# Check recommended fields
|
|
134
|
+
if entry_type in self.recommended_fields:
|
|
135
|
+
for rec_field in self.recommended_fields[entry_type]:
|
|
136
|
+
if rec_field not in fields or not fields[rec_field]:
|
|
137
|
+
warnings.append({
|
|
138
|
+
'type': 'missing_recommended_field',
|
|
139
|
+
'field': rec_field,
|
|
140
|
+
'severity': 'medium',
|
|
141
|
+
'message': f'Entry {key}: Missing recommended field "{rec_field}"'
|
|
142
|
+
})
|
|
143
|
+
|
|
144
|
+
# Validate year
|
|
145
|
+
if 'year' in fields:
|
|
146
|
+
year = fields['year']
|
|
147
|
+
if not re.match(r'^\d{4}$', year):
|
|
148
|
+
errors.append({
|
|
149
|
+
'type': 'invalid_year',
|
|
150
|
+
'field': 'year',
|
|
151
|
+
'value': year,
|
|
152
|
+
'severity': 'high',
|
|
153
|
+
'message': f'Entry {key}: Invalid year format "{year}" (should be 4 digits)'
|
|
154
|
+
})
|
|
155
|
+
elif int(year) < 1600 or int(year) > 2030:
|
|
156
|
+
warnings.append({
|
|
157
|
+
'type': 'suspicious_year',
|
|
158
|
+
'field': 'year',
|
|
159
|
+
'value': year,
|
|
160
|
+
'severity': 'medium',
|
|
161
|
+
'message': f'Entry {key}: Suspicious year "{year}" (outside reasonable range)'
|
|
162
|
+
})
|
|
163
|
+
|
|
164
|
+
# Validate DOI format
|
|
165
|
+
if 'doi' in fields:
|
|
166
|
+
doi = fields['doi']
|
|
167
|
+
if not re.match(r'^10\.\d{4,}/[^\s]+$', doi):
|
|
168
|
+
warnings.append({
|
|
169
|
+
'type': 'invalid_doi_format',
|
|
170
|
+
'field': 'doi',
|
|
171
|
+
'value': doi,
|
|
172
|
+
'severity': 'medium',
|
|
173
|
+
'message': f'Entry {key}: Invalid DOI format "{doi}"'
|
|
174
|
+
})
|
|
175
|
+
|
|
176
|
+
# Check for single hyphen in pages (should be --)
|
|
177
|
+
if 'pages' in fields:
|
|
178
|
+
pages = fields['pages']
|
|
179
|
+
if re.search(r'\d-\d', pages) and '--' not in pages:
|
|
180
|
+
warnings.append({
|
|
181
|
+
'type': 'page_range_format',
|
|
182
|
+
'field': 'pages',
|
|
183
|
+
'value': pages,
|
|
184
|
+
'severity': 'low',
|
|
185
|
+
'message': f'Entry {key}: Page range uses single hyphen, should use -- (en-dash)'
|
|
186
|
+
})
|
|
187
|
+
|
|
188
|
+
# Check author format
|
|
189
|
+
if 'author' in fields:
|
|
190
|
+
author = fields['author']
|
|
191
|
+
if ';' in author or '&' in author:
|
|
192
|
+
errors.append({
|
|
193
|
+
'type': 'invalid_author_format',
|
|
194
|
+
'field': 'author',
|
|
195
|
+
'severity': 'high',
|
|
196
|
+
'message': f'Entry {key}: Authors should be separated by " and ", not ";" or "&"'
|
|
197
|
+
})
|
|
198
|
+
|
|
199
|
+
return errors, warnings
|
|
200
|
+
|
|
201
|
+
def verify_doi(self, doi: str) -> Tuple[bool, Optional[Dict]]:
|
|
202
|
+
"""
|
|
203
|
+
Verify DOI resolves correctly and get metadata.
|
|
204
|
+
|
|
205
|
+
Args:
|
|
206
|
+
doi: Digital Object Identifier
|
|
207
|
+
|
|
208
|
+
Returns:
|
|
209
|
+
Tuple of (is_valid, metadata)
|
|
210
|
+
"""
|
|
211
|
+
try:
|
|
212
|
+
url = f'https://doi.org/{doi}'
|
|
213
|
+
response = self.session.head(url, timeout=10, allow_redirects=True)
|
|
214
|
+
|
|
215
|
+
if response.status_code < 400:
|
|
216
|
+
# DOI resolves, now get metadata from CrossRef
|
|
217
|
+
crossref_url = f'https://api.crossref.org/works/{doi}'
|
|
218
|
+
metadata_response = self.session.get(crossref_url, timeout=10)
|
|
219
|
+
|
|
220
|
+
if metadata_response.status_code == 200:
|
|
221
|
+
data = metadata_response.json()
|
|
222
|
+
message = data.get('message', {})
|
|
223
|
+
|
|
224
|
+
# Extract key metadata
|
|
225
|
+
metadata = {
|
|
226
|
+
'title': message.get('title', [''])[0],
|
|
227
|
+
'year': self._extract_year_crossref(message),
|
|
228
|
+
'authors': self._format_authors_crossref(message.get('author', [])),
|
|
229
|
+
}
|
|
230
|
+
return True, metadata
|
|
231
|
+
else:
|
|
232
|
+
return True, None # DOI resolves but no CrossRef metadata
|
|
233
|
+
else:
|
|
234
|
+
return False, None
|
|
235
|
+
|
|
236
|
+
except Exception:
|
|
237
|
+
return False, None
|
|
238
|
+
|
|
239
|
+
def detect_duplicates(self, entries: List[Dict]) -> List[Dict]:
|
|
240
|
+
"""
|
|
241
|
+
Detect duplicate entries.
|
|
242
|
+
|
|
243
|
+
Args:
|
|
244
|
+
entries: List of entry dictionaries
|
|
245
|
+
|
|
246
|
+
Returns:
|
|
247
|
+
List of duplicate groups
|
|
248
|
+
"""
|
|
249
|
+
duplicates = []
|
|
250
|
+
|
|
251
|
+
# Check for duplicate DOIs
|
|
252
|
+
doi_map = defaultdict(list)
|
|
253
|
+
for entry in entries:
|
|
254
|
+
doi = entry['fields'].get('doi', '').strip()
|
|
255
|
+
if doi:
|
|
256
|
+
doi_map[doi].append(entry['key'])
|
|
257
|
+
|
|
258
|
+
for doi, keys in doi_map.items():
|
|
259
|
+
if len(keys) > 1:
|
|
260
|
+
duplicates.append({
|
|
261
|
+
'type': 'duplicate_doi',
|
|
262
|
+
'doi': doi,
|
|
263
|
+
'entries': keys,
|
|
264
|
+
'severity': 'high',
|
|
265
|
+
'message': f'Duplicate DOI {doi} found in entries: {", ".join(keys)}'
|
|
266
|
+
})
|
|
267
|
+
|
|
268
|
+
# Check for duplicate citation keys
|
|
269
|
+
key_counts = defaultdict(int)
|
|
270
|
+
for entry in entries:
|
|
271
|
+
key_counts[entry['key']] += 1
|
|
272
|
+
|
|
273
|
+
for key, count in key_counts.items():
|
|
274
|
+
if count > 1:
|
|
275
|
+
duplicates.append({
|
|
276
|
+
'type': 'duplicate_key',
|
|
277
|
+
'key': key,
|
|
278
|
+
'count': count,
|
|
279
|
+
'severity': 'high',
|
|
280
|
+
'message': f'Citation key "{key}" appears {count} times'
|
|
281
|
+
})
|
|
282
|
+
|
|
283
|
+
# Check for similar titles (possible duplicates)
|
|
284
|
+
titles = {}
|
|
285
|
+
for entry in entries:
|
|
286
|
+
title = entry['fields'].get('title', '').lower()
|
|
287
|
+
title = re.sub(r'[^\w\s]', '', title) # Remove punctuation
|
|
288
|
+
title = ' '.join(title.split()) # Normalize whitespace
|
|
289
|
+
|
|
290
|
+
if title:
|
|
291
|
+
if title in titles:
|
|
292
|
+
duplicates.append({
|
|
293
|
+
'type': 'similar_title',
|
|
294
|
+
'entries': [titles[title], entry['key']],
|
|
295
|
+
'severity': 'medium',
|
|
296
|
+
'message': f'Possible duplicate: "{titles[title]}" and "{entry["key"]}" have identical titles'
|
|
297
|
+
})
|
|
298
|
+
else:
|
|
299
|
+
titles[title] = entry['key']
|
|
300
|
+
|
|
301
|
+
return duplicates
|
|
302
|
+
|
|
303
|
+
def validate_file(self, filepath: str, check_dois: bool = False) -> Dict:
|
|
304
|
+
"""
|
|
305
|
+
Validate entire BibTeX file.
|
|
306
|
+
|
|
307
|
+
Args:
|
|
308
|
+
filepath: Path to BibTeX file
|
|
309
|
+
check_dois: Whether to verify DOIs (slow)
|
|
310
|
+
|
|
311
|
+
Returns:
|
|
312
|
+
Validation report dictionary
|
|
313
|
+
"""
|
|
314
|
+
print(f'Parsing {filepath}...', file=sys.stderr)
|
|
315
|
+
entries = self.parse_bibtex_file(filepath)
|
|
316
|
+
|
|
317
|
+
if not entries:
|
|
318
|
+
return {
|
|
319
|
+
'total_entries': 0,
|
|
320
|
+
'errors': [],
|
|
321
|
+
'warnings': [],
|
|
322
|
+
'duplicates': []
|
|
323
|
+
}
|
|
324
|
+
|
|
325
|
+
print(f'Found {len(entries)} entries', file=sys.stderr)
|
|
326
|
+
|
|
327
|
+
all_errors = []
|
|
328
|
+
all_warnings = []
|
|
329
|
+
|
|
330
|
+
# Validate each entry
|
|
331
|
+
for i, entry in enumerate(entries):
|
|
332
|
+
print(f'Validating entry {i+1}/{len(entries)}: {entry["key"]}', file=sys.stderr)
|
|
333
|
+
errors, warnings = self.validate_entry(entry)
|
|
334
|
+
|
|
335
|
+
for error in errors:
|
|
336
|
+
error['entry'] = entry['key']
|
|
337
|
+
all_errors.append(error)
|
|
338
|
+
|
|
339
|
+
for warning in warnings:
|
|
340
|
+
warning['entry'] = entry['key']
|
|
341
|
+
all_warnings.append(warning)
|
|
342
|
+
|
|
343
|
+
# Check for duplicates
|
|
344
|
+
print('Checking for duplicates...', file=sys.stderr)
|
|
345
|
+
duplicates = self.detect_duplicates(entries)
|
|
346
|
+
|
|
347
|
+
# Verify DOIs if requested
|
|
348
|
+
doi_errors = []
|
|
349
|
+
if check_dois:
|
|
350
|
+
print('Verifying DOIs...', file=sys.stderr)
|
|
351
|
+
for i, entry in enumerate(entries):
|
|
352
|
+
doi = entry['fields'].get('doi', '')
|
|
353
|
+
if doi:
|
|
354
|
+
print(f'Verifying DOI {i+1}: {doi}', file=sys.stderr)
|
|
355
|
+
is_valid, metadata = self.verify_doi(doi)
|
|
356
|
+
|
|
357
|
+
if not is_valid:
|
|
358
|
+
doi_errors.append({
|
|
359
|
+
'type': 'invalid_doi',
|
|
360
|
+
'entry': entry['key'],
|
|
361
|
+
'doi': doi,
|
|
362
|
+
'severity': 'high',
|
|
363
|
+
'message': f'Entry {entry["key"]}: DOI does not resolve: {doi}'
|
|
364
|
+
})
|
|
365
|
+
|
|
366
|
+
all_errors.extend(doi_errors)
|
|
367
|
+
|
|
368
|
+
return {
|
|
369
|
+
'filepath': filepath,
|
|
370
|
+
'total_entries': len(entries),
|
|
371
|
+
'valid_entries': len(entries) - len([e for e in all_errors if e['severity'] == 'high']),
|
|
372
|
+
'errors': all_errors,
|
|
373
|
+
'warnings': all_warnings,
|
|
374
|
+
'duplicates': duplicates
|
|
375
|
+
}
|
|
376
|
+
|
|
377
|
+
def _extract_year_crossref(self, message: Dict) -> str:
|
|
378
|
+
"""Extract year from CrossRef message."""
|
|
379
|
+
date_parts = message.get('published-print', {}).get('date-parts', [[]])
|
|
380
|
+
if not date_parts or not date_parts[0]:
|
|
381
|
+
date_parts = message.get('published-online', {}).get('date-parts', [[]])
|
|
382
|
+
|
|
383
|
+
if date_parts and date_parts[0]:
|
|
384
|
+
return str(date_parts[0][0])
|
|
385
|
+
return ''
|
|
386
|
+
|
|
387
|
+
def _format_authors_crossref(self, authors: List[Dict]) -> str:
|
|
388
|
+
"""Format author list from CrossRef."""
|
|
389
|
+
if not authors:
|
|
390
|
+
return ''
|
|
391
|
+
|
|
392
|
+
formatted = []
|
|
393
|
+
for author in authors[:3]: # First 3 authors
|
|
394
|
+
given = author.get('given', '')
|
|
395
|
+
family = author.get('family', '')
|
|
396
|
+
if family:
|
|
397
|
+
formatted.append(f'{family}, {given}' if given else family)
|
|
398
|
+
|
|
399
|
+
if len(authors) > 3:
|
|
400
|
+
formatted.append('et al.')
|
|
401
|
+
|
|
402
|
+
return ', '.join(formatted)
|
|
403
|
+
|
|
404
|
+
|
|
405
|
+
def main():
|
|
406
|
+
"""Command-line interface."""
|
|
407
|
+
parser = argparse.ArgumentParser(
|
|
408
|
+
description='Validate BibTeX files for errors and inconsistencies',
|
|
409
|
+
epilog='Example: python validate_citations.py references.bib'
|
|
410
|
+
)
|
|
411
|
+
|
|
412
|
+
parser.add_argument(
|
|
413
|
+
'file',
|
|
414
|
+
help='BibTeX file to validate'
|
|
415
|
+
)
|
|
416
|
+
|
|
417
|
+
parser.add_argument(
|
|
418
|
+
'--check-dois',
|
|
419
|
+
action='store_true',
|
|
420
|
+
help='Verify DOIs resolve correctly (slow)'
|
|
421
|
+
)
|
|
422
|
+
|
|
423
|
+
parser.add_argument(
|
|
424
|
+
'--auto-fix',
|
|
425
|
+
action='store_true',
|
|
426
|
+
help='Attempt to auto-fix common issues (not implemented yet)'
|
|
427
|
+
)
|
|
428
|
+
|
|
429
|
+
parser.add_argument(
|
|
430
|
+
'--report',
|
|
431
|
+
help='Output file for JSON validation report'
|
|
432
|
+
)
|
|
433
|
+
|
|
434
|
+
parser.add_argument(
|
|
435
|
+
'--verbose',
|
|
436
|
+
action='store_true',
|
|
437
|
+
help='Show detailed output'
|
|
438
|
+
)
|
|
439
|
+
|
|
440
|
+
args = parser.parse_args()
|
|
441
|
+
|
|
442
|
+
# Validate file
|
|
443
|
+
validator = CitationValidator()
|
|
444
|
+
report = validator.validate_file(args.file, check_dois=args.check_dois)
|
|
445
|
+
|
|
446
|
+
# Print summary
|
|
447
|
+
print('\n' + '='*60)
|
|
448
|
+
print('CITATION VALIDATION REPORT')
|
|
449
|
+
print('='*60)
|
|
450
|
+
print(f'\nFile: {args.file}')
|
|
451
|
+
print(f'Total entries: {report["total_entries"]}')
|
|
452
|
+
print(f'Valid entries: {report["valid_entries"]}')
|
|
453
|
+
print(f'Errors: {len(report["errors"])}')
|
|
454
|
+
print(f'Warnings: {len(report["warnings"])}')
|
|
455
|
+
print(f'Duplicates: {len(report["duplicates"])}')
|
|
456
|
+
|
|
457
|
+
# Print errors
|
|
458
|
+
if report['errors']:
|
|
459
|
+
print('\n' + '-'*60)
|
|
460
|
+
print('ERRORS (must fix):')
|
|
461
|
+
print('-'*60)
|
|
462
|
+
for error in report['errors']:
|
|
463
|
+
print(f'\n{error["message"]}')
|
|
464
|
+
if args.verbose:
|
|
465
|
+
print(f' Type: {error["type"]}')
|
|
466
|
+
print(f' Severity: {error["severity"]}')
|
|
467
|
+
|
|
468
|
+
# Print warnings
|
|
469
|
+
if report['warnings'] and args.verbose:
|
|
470
|
+
print('\n' + '-'*60)
|
|
471
|
+
print('WARNINGS (should fix):')
|
|
472
|
+
print('-'*60)
|
|
473
|
+
for warning in report['warnings']:
|
|
474
|
+
print(f'\n{warning["message"]}')
|
|
475
|
+
|
|
476
|
+
# Print duplicates
|
|
477
|
+
if report['duplicates']:
|
|
478
|
+
print('\n' + '-'*60)
|
|
479
|
+
print('DUPLICATES:')
|
|
480
|
+
print('-'*60)
|
|
481
|
+
for dup in report['duplicates']:
|
|
482
|
+
print(f'\n{dup["message"]}')
|
|
483
|
+
|
|
484
|
+
# Save report
|
|
485
|
+
if args.report:
|
|
486
|
+
with open(args.report, 'w', encoding='utf-8') as f:
|
|
487
|
+
json.dump(report, f, indent=2)
|
|
488
|
+
print(f'\nDetailed report saved to: {args.report}')
|
|
489
|
+
|
|
490
|
+
# Exit with error code if there are errors
|
|
491
|
+
if report['errors']:
|
|
492
|
+
sys.exit(1)
|
|
493
|
+
|
|
494
|
+
|
|
495
|
+
if __name__ == '__main__':
|
|
496
|
+
main()
|
|
497
|
+
|
|
@@ -0,0 +1,92 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: data-analysis
|
|
3
|
+
description: |
|
|
4
|
+
【数据·深度分析】数据生成过程还原、机理关联、领域知识提取与质量审计。适用:已了解数据结构后的深度分析。不适用:首次拿到未知数据文件(用 exploratory-data-analysis);流程型系统专项 QA(用 data-qa-analysis)。
|
|
5
|
+
version: '1.0'
|
|
6
|
+
metadata:
|
|
7
|
+
tier: system
|
|
8
|
+
origin: template
|
|
9
|
+
created_at: '2026-06-12T02:47:05.407046+00:00'
|
|
10
|
+
---
|
|
11
|
+
|
|
12
|
+
# 数据深度分析 Skill (SOP)
|
|
13
|
+
|
|
14
|
+
> 本 Skill 旨在将数据分析提升至领域专家水平。你不再是一个简单的统计员,而是**资深数据分析专家**。
|
|
15
|
+
|
|
16
|
+
## 使用边界
|
|
17
|
+
|
|
18
|
+
- **适用**:数据结构已初步了解后的深度分析(生成过程还原、机理关联、领域知识)。
|
|
19
|
+
- **不适用**:首次未知数据文件 → `exploratory-data-analysis`;流程型系统专项 → `data-qa-analysis`;纯统计检验 → `statistical-analysis`。
|
|
20
|
+
|
|
21
|
+
## 落盘约定
|
|
22
|
+
|
|
23
|
+
- **过程脚本**:`lab/scripts/`(探索性分析、绘图)
|
|
24
|
+
- **交付目录**:`products/extracted/`(结构化产物与报告)
|
|
25
|
+
- `products/extracted/tag_entity_map.json`(变量→实体映射;**禁止**写入 `entity_map.json`,该路径保留给流程拓扑 skill)
|
|
26
|
+
- `products/extracted/data_dictionary.md`
|
|
27
|
+
- `products/extracted/feature_candidates.json`
|
|
28
|
+
- `products/extracted/expert_rules.json`
|
|
29
|
+
- `products/extracted/cleaning_rules.yaml`
|
|
30
|
+
- `products/extracted/data_quality_report.md`
|
|
31
|
+
- `products/extracted/domain_knowledge.md`
|
|
32
|
+
- `products/extracted/analysis_report.md`
|
|
33
|
+
- 填写 `templates/domain_knowledge_template.md.tpl` 时,内容须汇总写入 `products/extracted/domain_knowledge.md`
|
|
34
|
+
|
|
35
|
+
## 1. 核心方法论 (The Four Lenses)
|
|
36
|
+
你的分析必须贯穿以下四个维度(Meta-Knowledge):
|
|
37
|
+
- **Process (过程)**: 还原数据生成过程,识别运行/实验模式,评估过程稳定性。
|
|
38
|
+
- **Quality (质量)**: 审计数据质量,校验机理与单位一致性,关联异常与缺陷。
|
|
39
|
+
- **Cost (代价)**: 识别隐性的信息损失与采集代价,量化异常的影响。
|
|
40
|
+
- **Continuity (连续性)**: 分析时间序列的连续性,识别中断与效率瓶颈。
|
|
41
|
+
|
|
42
|
+
## 2. 决策树 (Decision Tree)
|
|
43
|
+
根据用户意图选择子流程:
|
|
44
|
+
1. **"理解这份数据 / 还原数据生成过程"** -> `workflows/01-data-profiling.md` (深度画像与过程还原)
|
|
45
|
+
2. **"数据质量如何 / 能否物理自洽"** -> `workflows/02-quality-audit.md` (时序对齐与机理校验)
|
|
46
|
+
3. **"寻找因果关系 / 挖掘特征"** -> `workflows/03-physical-correlation.md` (滞后分析与因果发现)
|
|
47
|
+
4. **"分析日志 / 提取专家经验"** -> `workflows/04-unstructured-mining.md` (非结构化挖掘)
|
|
48
|
+
5. **"全面体检"** -> 按顺序执行 1 -> 2 -> 3 -> 4,最后基于 `templates/analysis_report.md.tpl` 生成 `products/extracted/analysis_report.md`。
|
|
49
|
+
|
|
50
|
+
## 3. 核心指令 (Core Instructions)
|
|
51
|
+
|
|
52
|
+
<instruction>
|
|
53
|
+
<role>
|
|
54
|
+
你是顶尖研究机构级别的数据分析专家。你不仅仅看数字,你看到的是数字背后真实运转的系统:流动的物质、变化的状态和演化的过程。你对时间戳极其敏感,对量纲与单位一丝不苟。
|
|
55
|
+
</role>
|
|
56
|
+
|
|
57
|
+
<rule>
|
|
58
|
+
1. **机理优先 (Mechanism-First)**:
|
|
59
|
+
- 严禁在未确认单位的情况下计算统计量。
|
|
60
|
+
- 必须检查机理一致性(如:守恒量是否平衡?比例是否越界?)。
|
|
61
|
+
2. **过程还原 (Process Reconstruction)**:
|
|
62
|
+
- 不要把数据看作静态表格,而要看作动态过程的快照。
|
|
63
|
+
- 尝试通过数据还原出 "输入 -> 处理 -> 输出" 的时序逻辑。
|
|
64
|
+
3. **知识资产化**:
|
|
65
|
+
- 所有发现须汇总到 `products/extracted/domain_knowledge.md`(结构参考 `templates/domain_knowledge_template.md.tpl`)。
|
|
66
|
+
- 必须区分 "硬性约束" (Mechanism) 和 "软性规则" (Heuristics)。
|
|
67
|
+
4. **先读后算**:统计缺失率、采样间隔、相关系数、稳定性指标前,必须 read 或 bash 读取 `input/` 下实际文件;未取证前不得写出具体数字。
|
|
68
|
+
5. **变量必须存在**:`products/extracted/tag_entity_map.json` 中的变量/列名必须来自已读取数据的表头或字典,禁止凭命名惯例虚构变量。
|
|
69
|
+
6. **报告数值带来源**:`products/extracted/data_quality_report.md` / `products/extracted/analysis_report.md` 中每个量化结论附 `来源: <path>`。
|
|
70
|
+
</rule>
|
|
71
|
+
|
|
72
|
+
<thinking>
|
|
73
|
+
在执行每一步前,强制进行 Chain-of-Thought:
|
|
74
|
+
1. 这个变量在真实世界中对应什么实体?(观测通道?控制输入?)
|
|
75
|
+
2. 这段数据的变化趋势是否符合领域常识?
|
|
76
|
+
3. 如果我不理解这个异常,是否可以通过时间轴关联到其他变量?
|
|
77
|
+
4. 我提取的这条规则,在实践中是否可执行、可验证?
|
|
78
|
+
</thinking>
|
|
79
|
+
</instruction>
|
|
80
|
+
|
|
81
|
+
## 4. 资源索引 (Resource Index)
|
|
82
|
+
- **Workflows**:
|
|
83
|
+
- `workflows/01-data-profiling.md`
|
|
84
|
+
- `workflows/02-quality-audit.md`
|
|
85
|
+
- `workflows/03-physical-correlation.md`
|
|
86
|
+
- `workflows/04-unstructured-mining.md`
|
|
87
|
+
- **Templates**:
|
|
88
|
+
- `templates/domain_knowledge_template.md.tpl` (核心)
|
|
89
|
+
- `templates/analysis_report.md.tpl`
|
|
90
|
+
- `templates/data_dictionary.md.tpl`
|
|
91
|
+
- **Checklists**:
|
|
92
|
+
- `checklists/readiness_check.md`
|
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
# 建模准入自检清单 (Readiness Checklist)
|
|
2
|
+
|
|
3
|
+
> 在宣布分析完成并建议进入 Modeling 阶段前,必须通过以下所有检查。
|
|
4
|
+
|
|
5
|
+
## 1. 物理语义自检
|
|
6
|
+
- [ ] 所有字段的 **物理单位** 已确认且合理(如:没有出现“负的绝对温度”)。
|
|
7
|
+
- [ ] 已识别出 **Business Keys**(如批次号、样本编号、会话 ID),并验证了其唯一性。
|
|
8
|
+
- [ ] 已确认数据的 **采样频率**,并评估其是否足以捕捉目标过程(奈奎斯特定理)。
|
|
9
|
+
|
|
10
|
+
## 2. 数据质量自检
|
|
11
|
+
- [ ] **Data Dictionary** 已生成且完整。
|
|
12
|
+
- [ ] `products/extracted/data_quality_report.md` 已生成,且无致命数据质量缺陷。
|
|
13
|
+
- [ ] 所有识别出的异常模式(卡死、丢包)都有对应的 **Cleaning Rule** 建议。
|
|
14
|
+
|
|
15
|
+
## 3. 建模可行性自检
|
|
16
|
+
- [ ] 已识别出至少 1 个 **Target** (预测目标) 和 3 个以上的强相关 **Features**。
|
|
17
|
+
- [ ] 已评估数据量是否满足建模需求(如:复杂模型至少需要 1000+ 样本)。
|
|
18
|
+
- [ ] 已明确区分 **训练集** 和 **测试集** 的划分策略(如:按时间切分,而非随机 Shuffle)。
|
|
19
|
+
|
|
20
|
+
## 4. 交付物完整性
|
|
21
|
+
- [ ] `products/extracted/domain_knowledge.md` 已汇总各 workflow 章节
|
|
22
|
+
- [ ] `products/extracted/analysis_report.md` 包含四大维度视角的分析结论。
|
|
23
|
+
- [ ] `products/extracted/feature_candidates.json` 包含机理特征建议。
|
|
@@ -0,0 +1,63 @@
|
|
|
1
|
+
# 数据深度分析报告 (Deep Data Analysis Report)
|
|
2
|
+
|
|
3
|
+
**分析对象**: `{{ file_name }}`
|
|
4
|
+
**分析时间**: `{{ analysis_time }}`
|
|
5
|
+
**分析师**: ClearAI
|
|
6
|
+
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## 1. 核心结论 (Executive Summary)
|
|
10
|
+
> **一句话总结**: 数据质量评级为 **{{ quality_grade }}**,主要问题集中在 [XX] 方面,建议 [Action]。
|
|
11
|
+
|
|
12
|
+
- **数据可用性**: [High/Medium/Low]
|
|
13
|
+
- **关键发现**:
|
|
14
|
+
1. ...
|
|
15
|
+
2. ...
|
|
16
|
+
|
|
17
|
+
---
|
|
18
|
+
|
|
19
|
+
## 2. 四大维度视角分析 (Four-Lens Analysis)
|
|
20
|
+
|
|
21
|
+
### 2.1 过程视角 (Process) - 稳定性与能力
|
|
22
|
+
- **稳定性分析**: 关键参数 `{{ key_param }}` 的过程稳定性指标为 `{{ stability_value }}`(如 Cpk,目标: >1.33)。
|
|
23
|
+
- **过程状态**: 是否处于受控/稳定状态?[Yes/No]
|
|
24
|
+
- **过程异常**: 发现 `{{ anomaly_count }}` 次关键参数越限。
|
|
25
|
+
|
|
26
|
+
### 2.2 质量视角 (Quality) - 缺陷与关联
|
|
27
|
+
- **缺陷分布**: 主要缺陷类型为 `{{ top_defect }}`。
|
|
28
|
+
- **关联分析**: `{{ key_param }}` 的波动与 `{{ defect_type }}` 呈现 [正/负] 相关。
|
|
29
|
+
- **潜在风险**: ...
|
|
30
|
+
|
|
31
|
+
### 2.3 代价视角 (Cost) - 损失量化
|
|
32
|
+
- **质量损失估算**: 约 `{{ quality_loss_cost }}`(按项目约定的资源/成本单位)。
|
|
33
|
+
- **优化潜力**: 若将过程稳定性提升至目标水平,预计节约 `{{ potential_saving }}`。
|
|
34
|
+
|
|
35
|
+
### 2.4 连续性视角 (Continuity) - 覆盖与可用率
|
|
36
|
+
- **有效时长**: 数据覆盖率为 `{{ coverage_rate }}%`。
|
|
37
|
+
- **中断分析**: 识别出 `{{ downtime_count }}` 次异常中断/断数。
|
|
38
|
+
|
|
39
|
+
---
|
|
40
|
+
|
|
41
|
+
## 3. 详细画像 (Detailed Profiling)
|
|
42
|
+
|
|
43
|
+
### 3.1 基础统计
|
|
44
|
+
| 字段 | Min | Max | Mean | Std | Skew | Missing% |
|
|
45
|
+
| :--- | :--- | :--- | :--- | :--- | :--- | :--- |
|
|
46
|
+
| ... | ... | ... | ... | ... | ... | ... |
|
|
47
|
+
|
|
48
|
+
### 3.2 物理/业务键分析
|
|
49
|
+
- **Business Keys**: 识别出的主键为 `{{ primary_key }}`。
|
|
50
|
+
- **实体关联**: 包含 `{{ entity_count }}` 个唯一实体(如批次、样本、会话)。
|
|
51
|
+
|
|
52
|
+
---
|
|
53
|
+
|
|
54
|
+
## 4. 建模建议 (Modeling Recommendations)
|
|
55
|
+
1. **预处理建议**: ...
|
|
56
|
+
2. **特征工程**: 建议构建特征见 `products/extracted/feature_candidates.json`。
|
|
57
|
+
3. **模型选择**: 适合 [简单机理/复杂数据驱动] 建模。
|
|
58
|
+
|
|
59
|
+
---
|
|
60
|
+
|
|
61
|
+
## 5. 附录
|
|
62
|
+
- [数据字典](products/extracted/data_dictionary.md)
|
|
63
|
+
- [数据质量报告](products/extracted/data_quality_report.md)
|