clearai-dsh 0.3.0 → 0.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +147 -0
- package/README.md +28 -28
- package/README.zh-CN.md +28 -28
- package/lib/client.js +861 -2555
- package/lib/domain-language.js +552 -48
- package/lib/fold.js +351 -1015
- package/lib/host.js +47 -557
- package/lib/invariant.js +9 -12
- package/lib/knowledge-view.js +363 -226
- package/lib/lang.js +81 -0
- package/package.json +1 -1
- package/presets/clearai/agent.cordis.yml +54 -82
- package/presets/clearai/clearai.patch.yml +54 -82
- package/presets/clearai/plugins/clearai-kernel.js +1590 -5437
- package/presets/clearai/plugins/ontology.js +56 -14
- package/presets/clearai/plugins/prompts.js +73 -363
- package/presets/clearai/skills/clearai-loop/SKILL.md +73 -59
- package/presets/clearai/plugins/brain.js +0 -547
- package/presets/clearai/plugins/commands.js +0 -199
- package/presets/clearai/template/knowledge/README.md +0 -25
- package/presets/clearai/template/memory/README.md +0 -34
- package/presets/clearai/template/project.md +0 -49
- package/presets/clearai/template/skills/README.md +0 -37
- package/presets/clearai/template/skills/chart-diagram-qa/SKILL.md +0 -43
- package/presets/clearai/template/skills/citation-management/SKILL.md +0 -73
- package/presets/clearai/template/skills/citation-management/references/bibtex_formatting.md +0 -908
- package/presets/clearai/template/skills/citation-management/references/citation_validation.md +0 -794
- package/presets/clearai/template/skills/citation-management/references/google_scholar_search.md +0 -725
- package/presets/clearai/template/skills/citation-management/references/metadata_extraction.md +0 -870
- package/presets/clearai/template/skills/citation-management/references/pubmed_search.md +0 -839
- package/presets/clearai/template/skills/citation-management/scripts/doi_to_bibtex.py +0 -204
- package/presets/clearai/template/skills/citation-management/scripts/extract_metadata.py +0 -569
- package/presets/clearai/template/skills/citation-management/scripts/format_bibtex.py +0 -349
- package/presets/clearai/template/skills/citation-management/scripts/generate_schematic.py +0 -139
- package/presets/clearai/template/skills/citation-management/scripts/generate_schematic_ai.py +0 -817
- package/presets/clearai/template/skills/citation-management/scripts/search_google_scholar.py +0 -282
- package/presets/clearai/template/skills/citation-management/scripts/search_pubmed.py +0 -398
- package/presets/clearai/template/skills/citation-management/scripts/validate_citations.py +0 -497
- package/presets/clearai/template/skills/data-analysis/SKILL.md +0 -92
- package/presets/clearai/template/skills/data-analysis/checklists/readiness_check.md +0 -23
- package/presets/clearai/template/skills/data-analysis/templates/analysis_report.md.tpl +0 -63
- package/presets/clearai/template/skills/data-analysis/templates/cleaning_rules_draft.yaml.tpl +0 -32
- package/presets/clearai/template/skills/data-analysis/templates/data_dictionary.md.tpl +0 -12
- package/presets/clearai/template/skills/data-analysis/templates/domain_knowledge_template.md.tpl +0 -316
- package/presets/clearai/template/skills/data-analysis/templates/feature_candidates.json.tpl +0 -20
- package/presets/clearai/template/skills/data-analysis/templates/quality_scorecard.md.tpl +0 -30
- package/presets/clearai/template/skills/data-analysis/workflows/01-data-profiling.md +0 -42
- package/presets/clearai/template/skills/data-analysis/workflows/02-quality-audit.md +0 -36
- package/presets/clearai/template/skills/data-analysis/workflows/03-physical-correlation.md +0 -25
- package/presets/clearai/template/skills/data-analysis/workflows/04-unstructured-mining.md +0 -26
- package/presets/clearai/template/skills/data-qa-analysis/SKILL.md +0 -102
- package/presets/clearai/template/skills/data-qa-analysis/checklists/readiness_check.md +0 -62
- package/presets/clearai/template/skills/data-qa-analysis/templates/best_in_class_report.md.tpl +0 -56
- package/presets/clearai/template/skills/data-qa-analysis/templates/cleaning_rules_draft.yaml.tpl +0 -56
- package/presets/clearai/template/skills/data-qa-analysis/templates/data_dictionary.md.tpl +0 -13
- package/presets/clearai/template/skills/data-qa-analysis/templates/data_source_inventory_and_lineage.md.tpl +0 -146
- package/presets/clearai/template/skills/data-qa-analysis/templates/data_status_report.md.tpl +0 -60
- package/presets/clearai/template/skills/data-qa-analysis/templates/steady_state_rules.yaml.tpl +0 -41
- package/presets/clearai/template/skills/data-qa-analysis/templates/subsystem_registry.md.tpl +0 -101
- package/presets/clearai/template/skills/data-qa-analysis/templates/unified_execution_plan.md.tpl +0 -100
- package/presets/clearai/template/skills/data-qa-analysis/workflows/01-data-source-inventory-and-lineage.md +0 -194
- package/presets/clearai/template/skills/data-qa-analysis/workflows/02-data-alignment-and-tag-semantics.md +0 -122
- package/presets/clearai/template/skills/data-qa-analysis/workflows/03-steady-state-identification.md +0 -126
- package/presets/clearai/template/skills/data-qa-analysis/workflows/04-consumption-analysis.md +0 -152
- package/presets/clearai/template/skills/data-qa-analysis/workflows/05-best-in-class-and-optimization-space.md +0 -78
- package/presets/clearai/template/skills/domain-presearch/SKILL.md +0 -131
- package/presets/clearai/template/skills/domain-presearch/checklists/domain_checklist.md +0 -24
- package/presets/clearai/template/skills/domain-presearch/references/figure_code.md +0 -78
- package/presets/clearai/template/skills/domain-presearch/references/strategic_frameworks.md +0 -38
- package/presets/clearai/template/skills/exploration-loop/SKILL.md +0 -81
- package/presets/clearai/template/skills/exploratory-data-analysis/SKILL.md +0 -77
- package/presets/clearai/template/skills/exploratory-data-analysis/references/bioinformatics_genomics_formats.md +0 -664
- package/presets/clearai/template/skills/exploratory-data-analysis/references/chemistry_molecular_formats.md +0 -664
- package/presets/clearai/template/skills/exploratory-data-analysis/references/general_scientific_formats.md +0 -518
- package/presets/clearai/template/skills/exploratory-data-analysis/references/microscopy_imaging_formats.md +0 -620
- package/presets/clearai/template/skills/exploratory-data-analysis/references/proteomics_metabolomics_formats.md +0 -517
- package/presets/clearai/template/skills/exploratory-data-analysis/references/spectroscopy_analytical_formats.md +0 -633
- package/presets/clearai/template/skills/exploratory-data-analysis/scripts/eda_analyzer.py +0 -547
- package/presets/clearai/template/skills/hypothesis-generation/SKILL.md +0 -73
- package/presets/clearai/template/skills/hypothesis-generation/references/experimental_design_patterns.md +0 -329
- package/presets/clearai/template/skills/hypothesis-generation/references/hypothesis_quality_criteria.md +0 -198
- package/presets/clearai/template/skills/hypothesis-generation/references/literature_search_strategies.md +0 -622
- package/presets/clearai/template/skills/hypothesis-generation/scripts/generate_schematic.py +0 -139
- package/presets/clearai/template/skills/hypothesis-generation/scripts/generate_schematic_ai.py +0 -817
- package/presets/clearai/template/skills/literature-review/SKILL.md +0 -72
- package/presets/clearai/template/skills/literature-review/references/citation_styles.md +0 -166
- package/presets/clearai/template/skills/literature-review/references/database_strategies.md +0 -455
- package/presets/clearai/template/skills/literature-review/scripts/generate_pdf.py +0 -176
- package/presets/clearai/template/skills/literature-review/scripts/generate_schematic.py +0 -139
- package/presets/clearai/template/skills/literature-review/scripts/generate_schematic_ai.py +0 -817
- package/presets/clearai/template/skills/literature-review/scripts/search_databases.py +0 -303
- package/presets/clearai/template/skills/literature-review/scripts/verify_citations.py +0 -221
- package/presets/clearai/template/skills/paper-lookup/SKILL.md +0 -59
- package/presets/clearai/template/skills/paper-lookup/references/arxiv.md +0 -161
- package/presets/clearai/template/skills/paper-lookup/references/biorxiv.md +0 -118
- package/presets/clearai/template/skills/paper-lookup/references/core.md +0 -150
- package/presets/clearai/template/skills/paper-lookup/references/crossref.md +0 -181
- package/presets/clearai/template/skills/paper-lookup/references/medrxiv.md +0 -104
- package/presets/clearai/template/skills/paper-lookup/references/openalex.md +0 -174
- package/presets/clearai/template/skills/paper-lookup/references/pmc.md +0 -152
- package/presets/clearai/template/skills/paper-lookup/references/pubmed.md +0 -124
- package/presets/clearai/template/skills/paper-lookup/references/semantic-scholar.md +0 -203
- package/presets/clearai/template/skills/paper-lookup/references/unpaywall.md +0 -127
- package/presets/clearai/template/skills/process-presearch/SKILL.md +0 -196
- package/presets/clearai/template/skills/process-presearch/checklists/process_checklist.md +0 -18
- package/presets/clearai/template/skills/process-presearch/references/figure_code.md +0 -107
- package/presets/clearai/template/skills/process-presearch/references/source_attribution_example.md +0 -22
- package/presets/clearai/template/skills/process-understanding-extraction/SKILL.md +0 -69
- package/presets/clearai/template/skills/process-understanding-extraction/checklists/readiness_check.md +0 -34
- package/presets/clearai/template/skills/process-understanding-extraction/templates/docx_raw_dump_extractor.py.tpl +0 -132
- package/presets/clearai/template/skills/process-understanding-extraction/templates/entity_map_unit_topology.json.tpl +0 -86
- package/presets/clearai/template/skills/process-understanding-extraction/templates/process_brief.md.tpl +0 -89
- package/presets/clearai/template/skills/process-understanding-extraction/templates/process_brief_builder_from_raw_dump.py.tpl +0 -203
- package/presets/clearai/template/skills/process-understanding-extraction/templates/process_flow_mermaid.md.tpl +0 -41
- package/presets/clearai/template/skills/process-understanding-extraction/templates/unified_execution_plan.md.tpl +0 -53
- package/presets/clearai/template/skills/process-understanding-extraction/workflows/01-process-doc-discovery.md +0 -173
- package/presets/clearai/template/skills/process-understanding-extraction/workflows/02-process-understanding-and-diagramming.md +0 -106
- package/presets/clearai/template/skills/scientific-brainstorming/SKILL.md +0 -64
- package/presets/clearai/template/skills/scientific-brainstorming/references/brainstorming_methods.md +0 -326
- package/presets/clearai/template/skills/scientific-critical-thinking/SKILL.md +0 -72
- package/presets/clearai/template/skills/scientific-critical-thinking/references/common_biases.md +0 -364
- package/presets/clearai/template/skills/scientific-critical-thinking/references/evidence_hierarchy.md +0 -485
- package/presets/clearai/template/skills/scientific-critical-thinking/references/experimental_design.md +0 -496
- package/presets/clearai/template/skills/scientific-critical-thinking/references/logical_fallacies.md +0 -478
- package/presets/clearai/template/skills/scientific-critical-thinking/references/scientific_method.md +0 -169
- package/presets/clearai/template/skills/scientific-critical-thinking/references/statistical_pitfalls.md +0 -506
- package/presets/clearai/template/skills/skill-creator/SKILL.md +0 -109
- package/presets/clearai/template/skills/skill-creator/references/authoring-guide.md +0 -89
- package/presets/clearai/template/skills/statistical-analysis/SKILL.md +0 -79
- package/presets/clearai/template/skills/statistical-analysis/references/assumptions_and_diagnostics.md +0 -369
- package/presets/clearai/template/skills/statistical-analysis/references/bayesian_statistics.md +0 -653
- package/presets/clearai/template/skills/statistical-analysis/references/effect_sizes_and_power.md +0 -578
- package/presets/clearai/template/skills/statistical-analysis/references/reporting_standards.md +0 -469
- package/presets/clearai/template/skills/statistical-analysis/references/test_selection_guide.md +0 -129
- package/presets/clearai/template/skills/statistical-analysis/scripts/assumption_checks.py +0 -538
- package/presets/clearai/template/skills/web-artifact/SKILL.md +0 -165
- package/presets/clearai/template/skills/web-artifact/assets/renderer/renderer.css +0 -229
- package/presets/clearai/template/skills/web-artifact/assets/renderer/renderer.js +0 -373
- package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/LICENSE +0 -263
- package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/UPSTREAM.md +0 -26
- package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/elk.bundled.js +0 -6605
- package/presets/clearai/template/skills/web-artifact/references/when-drawing-a-topology.md +0 -150
- package/presets/clearai/template/skills/web-artifact/references/when-the-page-must-work-offline.md +0 -62
- package/presets/clearai/template/skills/web-artifact/scripts/check_artifact.py +0 -167
- package/presets/clearai/template/skills/web-artifact/scripts/render_topology.js +0 -272
- package/presets/clearai/template/skills/what-if-oracle/LICENSE.txt +0 -5
- package/presets/clearai/template/skills/what-if-oracle/SKILL.md +0 -72
- package/presets/clearai/template/skills/what-if-oracle/references/scenario-templates.md +0 -154
|
@@ -1,303 +0,0 @@
|
|
|
1
|
-
#!/usr/bin/env python3
|
|
2
|
-
"""
|
|
3
|
-
Literature Database Search Script
|
|
4
|
-
Searches multiple literature databases and aggregates results.
|
|
5
|
-
"""
|
|
6
|
-
|
|
7
|
-
import json
|
|
8
|
-
import sys
|
|
9
|
-
from typing import Dict, List
|
|
10
|
-
from datetime import datetime
|
|
11
|
-
|
|
12
|
-
def format_search_results(results: List[Dict], output_format: str = 'json') -> str:
|
|
13
|
-
"""
|
|
14
|
-
Format search results for output.
|
|
15
|
-
|
|
16
|
-
Args:
|
|
17
|
-
results: List of search results
|
|
18
|
-
output_format: Format (json, markdown, or bibtex)
|
|
19
|
-
|
|
20
|
-
Returns:
|
|
21
|
-
Formatted string
|
|
22
|
-
"""
|
|
23
|
-
if output_format == 'json':
|
|
24
|
-
return json.dumps(results, indent=2)
|
|
25
|
-
|
|
26
|
-
elif output_format == 'markdown':
|
|
27
|
-
md = f"# Literature Search Results\n\n"
|
|
28
|
-
md += f"**Search Date**: {datetime.now().strftime('%Y-%m-%d %H:%M')}\n"
|
|
29
|
-
md += f"**Total Results**: {len(results)}\n\n"
|
|
30
|
-
|
|
31
|
-
for i, result in enumerate(results, 1):
|
|
32
|
-
md += f"## {i}. {result.get('title', 'Untitled')}\n\n"
|
|
33
|
-
md += f"**Authors**: {result.get('authors', 'Unknown')}\n\n"
|
|
34
|
-
md += f"**Year**: {result.get('year', 'N/A')}\n\n"
|
|
35
|
-
md += f"**Source**: {result.get('source', 'Unknown')}\n\n"
|
|
36
|
-
|
|
37
|
-
if result.get('abstract'):
|
|
38
|
-
md += f"**Abstract**: {result['abstract']}\n\n"
|
|
39
|
-
|
|
40
|
-
if result.get('doi'):
|
|
41
|
-
md += f"**DOI**: [{result['doi']}](https://doi.org/{result['doi']})\n\n"
|
|
42
|
-
|
|
43
|
-
if result.get('url'):
|
|
44
|
-
md += f"**URL**: {result['url']}\n\n"
|
|
45
|
-
|
|
46
|
-
if result.get('citations'):
|
|
47
|
-
md += f"**Citations**: {result['citations']}\n\n"
|
|
48
|
-
|
|
49
|
-
md += "---\n\n"
|
|
50
|
-
|
|
51
|
-
return md
|
|
52
|
-
|
|
53
|
-
elif output_format == 'bibtex':
|
|
54
|
-
bibtex = ""
|
|
55
|
-
for i, result in enumerate(results, 1):
|
|
56
|
-
entry_type = result.get('type', 'article')
|
|
57
|
-
cite_key = f"{result.get('first_author', 'unknown')}{result.get('year', '0000')}"
|
|
58
|
-
|
|
59
|
-
bibtex += f"@{entry_type}{{{cite_key},\n"
|
|
60
|
-
bibtex += f" title = {{{result.get('title', '')}}},\n"
|
|
61
|
-
bibtex += f" author = {{{result.get('authors', '')}}},\n"
|
|
62
|
-
bibtex += f" year = {{{result.get('year', '')}}},\n"
|
|
63
|
-
|
|
64
|
-
if result.get('journal'):
|
|
65
|
-
bibtex += f" journal = {{{result['journal']}}},\n"
|
|
66
|
-
|
|
67
|
-
if result.get('volume'):
|
|
68
|
-
bibtex += f" volume = {{{result['volume']}}},\n"
|
|
69
|
-
|
|
70
|
-
if result.get('pages'):
|
|
71
|
-
bibtex += f" pages = {{{result['pages']}}},\n"
|
|
72
|
-
|
|
73
|
-
if result.get('doi'):
|
|
74
|
-
bibtex += f" doi = {{{result['doi']}}},\n"
|
|
75
|
-
|
|
76
|
-
bibtex += "}\n\n"
|
|
77
|
-
|
|
78
|
-
return bibtex
|
|
79
|
-
|
|
80
|
-
else:
|
|
81
|
-
raise ValueError(f"Unknown format: {output_format}")
|
|
82
|
-
|
|
83
|
-
def deduplicate_results(results: List[Dict]) -> List[Dict]:
|
|
84
|
-
"""
|
|
85
|
-
Remove duplicate results based on DOI or title.
|
|
86
|
-
|
|
87
|
-
Args:
|
|
88
|
-
results: List of search results
|
|
89
|
-
|
|
90
|
-
Returns:
|
|
91
|
-
Deduplicated list
|
|
92
|
-
"""
|
|
93
|
-
seen_dois = set()
|
|
94
|
-
seen_titles = set()
|
|
95
|
-
unique_results = []
|
|
96
|
-
|
|
97
|
-
for result in results:
|
|
98
|
-
doi = result.get('doi', '').lower().strip()
|
|
99
|
-
title = result.get('title', '').lower().strip()
|
|
100
|
-
|
|
101
|
-
# Check DOI first (more reliable)
|
|
102
|
-
if doi and doi in seen_dois:
|
|
103
|
-
continue
|
|
104
|
-
|
|
105
|
-
# Check title as fallback
|
|
106
|
-
if not doi and title in seen_titles:
|
|
107
|
-
continue
|
|
108
|
-
|
|
109
|
-
# Add to results
|
|
110
|
-
if doi:
|
|
111
|
-
seen_dois.add(doi)
|
|
112
|
-
if title:
|
|
113
|
-
seen_titles.add(title)
|
|
114
|
-
|
|
115
|
-
unique_results.append(result)
|
|
116
|
-
|
|
117
|
-
return unique_results
|
|
118
|
-
|
|
119
|
-
def rank_results(results: List[Dict], criteria: str = 'citations') -> List[Dict]:
|
|
120
|
-
"""
|
|
121
|
-
Rank results by specified criteria.
|
|
122
|
-
|
|
123
|
-
Args:
|
|
124
|
-
results: List of search results
|
|
125
|
-
criteria: Ranking criteria (citations, year, relevance)
|
|
126
|
-
|
|
127
|
-
Returns:
|
|
128
|
-
Ranked list
|
|
129
|
-
"""
|
|
130
|
-
if criteria == 'citations':
|
|
131
|
-
return sorted(results, key=lambda x: x.get('citations', 0), reverse=True)
|
|
132
|
-
elif criteria == 'year':
|
|
133
|
-
return sorted(results, key=lambda x: x.get('year', '0'), reverse=True)
|
|
134
|
-
elif criteria == 'relevance':
|
|
135
|
-
return sorted(results, key=lambda x: x.get('relevance_score', 0), reverse=True)
|
|
136
|
-
else:
|
|
137
|
-
return results
|
|
138
|
-
|
|
139
|
-
def filter_by_year(results: List[Dict], start_year: int = None, end_year: int = None) -> List[Dict]:
|
|
140
|
-
"""
|
|
141
|
-
Filter results by publication year range.
|
|
142
|
-
|
|
143
|
-
Args:
|
|
144
|
-
results: List of search results
|
|
145
|
-
start_year: Minimum year (inclusive)
|
|
146
|
-
end_year: Maximum year (inclusive)
|
|
147
|
-
|
|
148
|
-
Returns:
|
|
149
|
-
Filtered list
|
|
150
|
-
"""
|
|
151
|
-
filtered = []
|
|
152
|
-
|
|
153
|
-
for result in results:
|
|
154
|
-
try:
|
|
155
|
-
year = int(result.get('year', 0))
|
|
156
|
-
if start_year and year < start_year:
|
|
157
|
-
continue
|
|
158
|
-
if end_year and year > end_year:
|
|
159
|
-
continue
|
|
160
|
-
filtered.append(result)
|
|
161
|
-
except (ValueError, TypeError):
|
|
162
|
-
# Include if year parsing fails
|
|
163
|
-
filtered.append(result)
|
|
164
|
-
|
|
165
|
-
return filtered
|
|
166
|
-
|
|
167
|
-
def generate_search_summary(results: List[Dict]) -> Dict:
|
|
168
|
-
"""
|
|
169
|
-
Generate summary statistics for search results.
|
|
170
|
-
|
|
171
|
-
Args:
|
|
172
|
-
results: List of search results
|
|
173
|
-
|
|
174
|
-
Returns:
|
|
175
|
-
Summary dictionary
|
|
176
|
-
"""
|
|
177
|
-
summary = {
|
|
178
|
-
'total_results': len(results),
|
|
179
|
-
'sources': {},
|
|
180
|
-
'year_distribution': {},
|
|
181
|
-
'avg_citations': 0,
|
|
182
|
-
'total_citations': 0
|
|
183
|
-
}
|
|
184
|
-
|
|
185
|
-
citations = []
|
|
186
|
-
|
|
187
|
-
for result in results:
|
|
188
|
-
# Count by source
|
|
189
|
-
source = result.get('source', 'Unknown')
|
|
190
|
-
summary['sources'][source] = summary['sources'].get(source, 0) + 1
|
|
191
|
-
|
|
192
|
-
# Count by year
|
|
193
|
-
year = result.get('year', 'Unknown')
|
|
194
|
-
summary['year_distribution'][year] = summary['year_distribution'].get(year, 0) + 1
|
|
195
|
-
|
|
196
|
-
# Collect citations
|
|
197
|
-
if result.get('citations'):
|
|
198
|
-
try:
|
|
199
|
-
citations.append(int(result['citations']))
|
|
200
|
-
except (ValueError, TypeError):
|
|
201
|
-
pass
|
|
202
|
-
|
|
203
|
-
if citations:
|
|
204
|
-
summary['avg_citations'] = sum(citations) / len(citations)
|
|
205
|
-
summary['total_citations'] = sum(citations)
|
|
206
|
-
|
|
207
|
-
return summary
|
|
208
|
-
|
|
209
|
-
def main():
|
|
210
|
-
"""Command-line interface for search result processing."""
|
|
211
|
-
if len(sys.argv) < 2:
|
|
212
|
-
print("Usage: python search_databases.py <results.json> [options]")
|
|
213
|
-
print("\nOptions:")
|
|
214
|
-
print(" --format FORMAT Output format (json, markdown, bibtex)")
|
|
215
|
-
print(" --output FILE Output file (default: stdout)")
|
|
216
|
-
print(" --rank CRITERIA Rank by (citations, year, relevance)")
|
|
217
|
-
print(" --year-start YEAR Filter by start year")
|
|
218
|
-
print(" --year-end YEAR Filter by end year")
|
|
219
|
-
print(" --deduplicate Remove duplicates")
|
|
220
|
-
print(" --summary Show summary statistics")
|
|
221
|
-
sys.exit(1)
|
|
222
|
-
|
|
223
|
-
# Load results
|
|
224
|
-
results_file = sys.argv[1]
|
|
225
|
-
try:
|
|
226
|
-
with open(results_file, 'r', encoding='utf-8') as f:
|
|
227
|
-
results = json.load(f)
|
|
228
|
-
except Exception as e:
|
|
229
|
-
print(f"Error loading results: {e}")
|
|
230
|
-
sys.exit(1)
|
|
231
|
-
|
|
232
|
-
# Parse options
|
|
233
|
-
output_format = 'markdown'
|
|
234
|
-
output_file = None
|
|
235
|
-
rank_criteria = None
|
|
236
|
-
year_start = None
|
|
237
|
-
year_end = None
|
|
238
|
-
do_dedup = False
|
|
239
|
-
show_summary = False
|
|
240
|
-
|
|
241
|
-
i = 2
|
|
242
|
-
while i < len(sys.argv):
|
|
243
|
-
arg = sys.argv[i]
|
|
244
|
-
|
|
245
|
-
if arg == '--format' and i + 1 < len(sys.argv):
|
|
246
|
-
output_format = sys.argv[i + 1]
|
|
247
|
-
i += 2
|
|
248
|
-
elif arg == '--output' and i + 1 < len(sys.argv):
|
|
249
|
-
output_file = sys.argv[i + 1]
|
|
250
|
-
i += 2
|
|
251
|
-
elif arg == '--rank' and i + 1 < len(sys.argv):
|
|
252
|
-
rank_criteria = sys.argv[i + 1]
|
|
253
|
-
i += 2
|
|
254
|
-
elif arg == '--year-start' and i + 1 < len(sys.argv):
|
|
255
|
-
year_start = int(sys.argv[i + 1])
|
|
256
|
-
i += 2
|
|
257
|
-
elif arg == '--year-end' and i + 1 < len(sys.argv):
|
|
258
|
-
year_end = int(sys.argv[i + 1])
|
|
259
|
-
i += 2
|
|
260
|
-
elif arg == '--deduplicate':
|
|
261
|
-
do_dedup = True
|
|
262
|
-
i += 1
|
|
263
|
-
elif arg == '--summary':
|
|
264
|
-
show_summary = True
|
|
265
|
-
i += 1
|
|
266
|
-
else:
|
|
267
|
-
i += 1
|
|
268
|
-
|
|
269
|
-
# Process results
|
|
270
|
-
if do_dedup:
|
|
271
|
-
results = deduplicate_results(results)
|
|
272
|
-
print(f"After deduplication: {len(results)} results")
|
|
273
|
-
|
|
274
|
-
if year_start or year_end:
|
|
275
|
-
results = filter_by_year(results, year_start, year_end)
|
|
276
|
-
print(f"After year filter: {len(results)} results")
|
|
277
|
-
|
|
278
|
-
if rank_criteria:
|
|
279
|
-
results = rank_results(results, rank_criteria)
|
|
280
|
-
print(f"Ranked by: {rank_criteria}")
|
|
281
|
-
|
|
282
|
-
# Show summary
|
|
283
|
-
if show_summary:
|
|
284
|
-
summary = generate_search_summary(results)
|
|
285
|
-
print("\n" + "="*60)
|
|
286
|
-
print("SEARCH SUMMARY")
|
|
287
|
-
print("="*60)
|
|
288
|
-
print(json.dumps(summary, indent=2))
|
|
289
|
-
print()
|
|
290
|
-
|
|
291
|
-
# Format output
|
|
292
|
-
output = format_search_results(results, output_format)
|
|
293
|
-
|
|
294
|
-
# Write output
|
|
295
|
-
if output_file:
|
|
296
|
-
with open(output_file, 'w', encoding='utf-8') as f:
|
|
297
|
-
f.write(output)
|
|
298
|
-
print(f"✓ Results saved to: {output_file}")
|
|
299
|
-
else:
|
|
300
|
-
print(output)
|
|
301
|
-
|
|
302
|
-
if __name__ == "__main__":
|
|
303
|
-
main()
|
|
@@ -1,221 +0,0 @@
|
|
|
1
|
-
#!/usr/bin/env python3
|
|
2
|
-
"""
|
|
3
|
-
Citation Verification Script
|
|
4
|
-
Verifies DOIs, URLs, and citation metadata for accuracy.
|
|
5
|
-
"""
|
|
6
|
-
|
|
7
|
-
import re
|
|
8
|
-
import requests
|
|
9
|
-
import json
|
|
10
|
-
from typing import Dict, List, Tuple
|
|
11
|
-
import time
|
|
12
|
-
|
|
13
|
-
class CitationVerifier:
|
|
14
|
-
def __init__(self):
|
|
15
|
-
self.session = requests.Session()
|
|
16
|
-
self.session.headers.update({
|
|
17
|
-
'User-Agent': 'CitationVerifier/1.0 (Literature Review Tool)'
|
|
18
|
-
})
|
|
19
|
-
|
|
20
|
-
def extract_dois(self, text: str) -> List[str]:
|
|
21
|
-
"""Extract all DOIs from text."""
|
|
22
|
-
doi_pattern = r'10\.\d{4,}/[^\s\]\)"]+'
|
|
23
|
-
return re.findall(doi_pattern, text)
|
|
24
|
-
|
|
25
|
-
def verify_doi(self, doi: str) -> Tuple[bool, Dict]:
|
|
26
|
-
"""
|
|
27
|
-
Verify a DOI and retrieve metadata.
|
|
28
|
-
Returns (is_valid, metadata)
|
|
29
|
-
"""
|
|
30
|
-
try:
|
|
31
|
-
url = f"https://doi.org/api/handles/{doi}"
|
|
32
|
-
response = self.session.get(url, timeout=10)
|
|
33
|
-
|
|
34
|
-
if response.status_code == 200:
|
|
35
|
-
# DOI exists, now get metadata from CrossRef
|
|
36
|
-
metadata = self._get_crossref_metadata(doi)
|
|
37
|
-
return True, metadata
|
|
38
|
-
else:
|
|
39
|
-
return False, {}
|
|
40
|
-
except Exception as e:
|
|
41
|
-
return False, {"error": str(e)}
|
|
42
|
-
|
|
43
|
-
def _get_crossref_metadata(self, doi: str) -> Dict:
|
|
44
|
-
"""Get metadata from CrossRef API."""
|
|
45
|
-
try:
|
|
46
|
-
url = f"https://api.crossref.org/works/{doi}"
|
|
47
|
-
response = self.session.get(url, timeout=10)
|
|
48
|
-
|
|
49
|
-
if response.status_code == 200:
|
|
50
|
-
data = response.json()
|
|
51
|
-
message = data.get('message', {})
|
|
52
|
-
|
|
53
|
-
# Extract key metadata
|
|
54
|
-
metadata = {
|
|
55
|
-
'title': message.get('title', [''])[0],
|
|
56
|
-
'authors': self._format_authors(message.get('author', [])),
|
|
57
|
-
'year': self._extract_year(message),
|
|
58
|
-
'journal': message.get('container-title', [''])[0],
|
|
59
|
-
'volume': message.get('volume', ''),
|
|
60
|
-
'pages': message.get('page', ''),
|
|
61
|
-
'doi': doi
|
|
62
|
-
}
|
|
63
|
-
return metadata
|
|
64
|
-
return {}
|
|
65
|
-
except Exception as e:
|
|
66
|
-
return {"error": str(e)}
|
|
67
|
-
|
|
68
|
-
def _format_authors(self, authors: List[Dict]) -> str:
|
|
69
|
-
"""Format author list."""
|
|
70
|
-
if not authors:
|
|
71
|
-
return ""
|
|
72
|
-
|
|
73
|
-
formatted = []
|
|
74
|
-
for author in authors[:3]: # First 3 authors
|
|
75
|
-
given = author.get('given', '')
|
|
76
|
-
family = author.get('family', '')
|
|
77
|
-
if family:
|
|
78
|
-
formatted.append(f"{family}, {given[0]}." if given else family)
|
|
79
|
-
|
|
80
|
-
if len(authors) > 3:
|
|
81
|
-
formatted.append("et al.")
|
|
82
|
-
|
|
83
|
-
return ", ".join(formatted)
|
|
84
|
-
|
|
85
|
-
def _extract_year(self, message: Dict) -> str:
|
|
86
|
-
"""Extract publication year."""
|
|
87
|
-
date_parts = message.get('published-print', {}).get('date-parts', [[]])
|
|
88
|
-
if not date_parts or not date_parts[0]:
|
|
89
|
-
date_parts = message.get('published-online', {}).get('date-parts', [[]])
|
|
90
|
-
|
|
91
|
-
if date_parts and date_parts[0]:
|
|
92
|
-
return str(date_parts[0][0])
|
|
93
|
-
return ""
|
|
94
|
-
|
|
95
|
-
def verify_url(self, url: str) -> Tuple[bool, int]:
|
|
96
|
-
"""
|
|
97
|
-
Verify a URL is accessible.
|
|
98
|
-
Returns (is_accessible, status_code)
|
|
99
|
-
"""
|
|
100
|
-
try:
|
|
101
|
-
response = self.session.head(url, timeout=10, allow_redirects=True)
|
|
102
|
-
is_accessible = response.status_code < 400
|
|
103
|
-
return is_accessible, response.status_code
|
|
104
|
-
except Exception:
|
|
105
|
-
return False, 0
|
|
106
|
-
|
|
107
|
-
def verify_citations_in_file(self, filepath: str) -> Dict:
|
|
108
|
-
"""
|
|
109
|
-
Verify all citations in a markdown file.
|
|
110
|
-
Returns a report of verification results.
|
|
111
|
-
"""
|
|
112
|
-
with open(filepath, 'r', encoding='utf-8') as f:
|
|
113
|
-
content = f.read()
|
|
114
|
-
|
|
115
|
-
dois = self.extract_dois(content)
|
|
116
|
-
|
|
117
|
-
report = {
|
|
118
|
-
'total_dois': len(dois),
|
|
119
|
-
'verified': [],
|
|
120
|
-
'failed': [],
|
|
121
|
-
'metadata': {}
|
|
122
|
-
}
|
|
123
|
-
|
|
124
|
-
for doi in dois:
|
|
125
|
-
print(f"Verifying DOI: {doi}")
|
|
126
|
-
is_valid, metadata = self.verify_doi(doi)
|
|
127
|
-
|
|
128
|
-
if is_valid:
|
|
129
|
-
report['verified'].append(doi)
|
|
130
|
-
report['metadata'][doi] = metadata
|
|
131
|
-
else:
|
|
132
|
-
report['failed'].append(doi)
|
|
133
|
-
|
|
134
|
-
time.sleep(0.5) # Rate limiting
|
|
135
|
-
|
|
136
|
-
return report
|
|
137
|
-
|
|
138
|
-
def format_citation_apa(self, metadata: Dict) -> str:
|
|
139
|
-
"""Format citation in APA style."""
|
|
140
|
-
authors = metadata.get('authors', '')
|
|
141
|
-
year = metadata.get('year', 'n.d.')
|
|
142
|
-
title = metadata.get('title', '')
|
|
143
|
-
journal = metadata.get('journal', '')
|
|
144
|
-
volume = metadata.get('volume', '')
|
|
145
|
-
pages = metadata.get('pages', '')
|
|
146
|
-
doi = metadata.get('doi', '')
|
|
147
|
-
|
|
148
|
-
citation = f"{authors} ({year}). {title}. "
|
|
149
|
-
if journal:
|
|
150
|
-
citation += f"*{journal}*"
|
|
151
|
-
if volume:
|
|
152
|
-
citation += f", *{volume}*"
|
|
153
|
-
if pages:
|
|
154
|
-
citation += f", {pages}"
|
|
155
|
-
if doi:
|
|
156
|
-
citation += f". https://doi.org/{doi}"
|
|
157
|
-
|
|
158
|
-
return citation
|
|
159
|
-
|
|
160
|
-
def format_citation_nature(self, metadata: Dict) -> str:
|
|
161
|
-
"""Format citation in Nature style."""
|
|
162
|
-
authors = metadata.get('authors', '')
|
|
163
|
-
title = metadata.get('title', '')
|
|
164
|
-
journal = metadata.get('journal', '')
|
|
165
|
-
volume = metadata.get('volume', '')
|
|
166
|
-
pages = metadata.get('pages', '')
|
|
167
|
-
year = metadata.get('year', '')
|
|
168
|
-
|
|
169
|
-
citation = f"{authors} {title}. "
|
|
170
|
-
if journal:
|
|
171
|
-
citation += f"*{journal}* "
|
|
172
|
-
if volume:
|
|
173
|
-
citation += f"**{volume}**, "
|
|
174
|
-
if pages:
|
|
175
|
-
citation += f"{pages} "
|
|
176
|
-
if year:
|
|
177
|
-
citation += f"({year})"
|
|
178
|
-
|
|
179
|
-
return citation
|
|
180
|
-
|
|
181
|
-
def main():
|
|
182
|
-
"""Example usage."""
|
|
183
|
-
import sys
|
|
184
|
-
|
|
185
|
-
if len(sys.argv) < 2:
|
|
186
|
-
print("Usage: python verify_citations.py <markdown_file>")
|
|
187
|
-
sys.exit(1)
|
|
188
|
-
|
|
189
|
-
filepath = sys.argv[1]
|
|
190
|
-
verifier = CitationVerifier()
|
|
191
|
-
|
|
192
|
-
print(f"Verifying citations in: {filepath}")
|
|
193
|
-
report = verifier.verify_citations_in_file(filepath)
|
|
194
|
-
|
|
195
|
-
print("\n" + "="*60)
|
|
196
|
-
print("CITATION VERIFICATION REPORT")
|
|
197
|
-
print("="*60)
|
|
198
|
-
print(f"\nTotal DOIs found: {report['total_dois']}")
|
|
199
|
-
print(f"Verified: {len(report['verified'])}")
|
|
200
|
-
print(f"Failed: {len(report['failed'])}")
|
|
201
|
-
|
|
202
|
-
if report['failed']:
|
|
203
|
-
print("\nFailed DOIs:")
|
|
204
|
-
for doi in report['failed']:
|
|
205
|
-
print(f" - {doi}")
|
|
206
|
-
|
|
207
|
-
if report['metadata']:
|
|
208
|
-
print("\n\nVerified Citations (APA format):")
|
|
209
|
-
for doi, metadata in report['metadata'].items():
|
|
210
|
-
citation = verifier.format_citation_apa(metadata)
|
|
211
|
-
print(f"\n{citation}")
|
|
212
|
-
|
|
213
|
-
# Save detailed report
|
|
214
|
-
output_file = filepath.replace('.md', '_citation_report.json')
|
|
215
|
-
with open(output_file, 'w', encoding='utf-8') as f:
|
|
216
|
-
json.dump(report, f, indent=2)
|
|
217
|
-
|
|
218
|
-
print(f"\n\nDetailed report saved to: {output_file}")
|
|
219
|
-
|
|
220
|
-
if __name__ == "__main__":
|
|
221
|
-
main()
|
|
@@ -1,59 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: paper-lookup
|
|
3
|
-
description: |
|
|
4
|
-
【论文检索·学术 API】通过 REST API 查 PubMed/arXiv/OpenAlex/Crossref 等学术库。适用:DOI/PMID 解析、论文元数据、引用关系、开放获取链接。不适用:行业快研与商业信息(用 web_search + domain-presearch);完整综述流程(用 literature-review)。
|
|
5
|
-
license: MIT license
|
|
6
|
-
metadata:
|
|
7
|
-
version: 1.0-clearai
|
|
8
|
-
skill-author: K-Dense Inc. (adapted for ClearAI)
|
|
9
|
-
tier: system
|
|
10
|
-
origin: template
|
|
11
|
-
created_at: '2026-06-12T02:47:05.409595+00:00'
|
|
12
|
-
---
|
|
13
|
-
|
|
14
|
-
# Paper Lookup Skill(ClearAI 版)
|
|
15
|
-
|
|
16
|
-
## 使用边界
|
|
17
|
-
|
|
18
|
-
- **适用**:需要学术论文元数据、DOI 互转、预印本、引用图;补充 `web_search` 的学术精度。
|
|
19
|
-
- **不适用**:替代 `web_search` 做行业新闻/企业情报;不单独产出完整综述。
|
|
20
|
-
|
|
21
|
-
## ClearAI 工具与路径映射
|
|
22
|
-
|
|
23
|
-
- API 调用 → `bash`(`curl` / Python `requests`),脚本可放 `lab/scripts/`
|
|
24
|
-
- 结果落盘 → `lab/knowledge/paper_lookup_results.json`
|
|
25
|
-
- 与综述衔接 → 供 `literature-review` / `citation-management` 消费
|
|
26
|
-
- 经验回写 → `clear/memory/paper_lookup_lessons.md`
|
|
27
|
-
|
|
28
|
-
## 核心工作流
|
|
29
|
-
|
|
30
|
-
1. **理解查询**:主题检索 / 特定 DOI / 作者 / OA PDF?
|
|
31
|
-
2. **选库**:见下表;详 endpoint 见 `references/` 各库文件
|
|
32
|
-
3. **并行查询**:独立库可用同轮多个 `bash`
|
|
33
|
-
4. **返回**:原始 JSON + 已查库列表 + 无结果须显式说明
|
|
34
|
-
|
|
35
|
-
## 数据库选择(摘要)
|
|
36
|
-
|
|
37
|
-
| 意图 | 主库 | 补充 |
|
|
38
|
-
|------|------|------|
|
|
39
|
-
| 生物医学主题 | PubMed | Semantic Scholar, OpenAlex |
|
|
40
|
-
| 物理/数学/CS 预印本 | arXiv | OpenAlex |
|
|
41
|
-
| 跨学科 | OpenAlex | Crossref, Semantic Scholar |
|
|
42
|
-
| DOI 元数据 | Crossref | Unpaywall |
|
|
43
|
-
| 开放获取链接 | Unpaywall | PMC, CORE |
|
|
44
|
-
| 引用关系 | Semantic Scholar | OpenAlex |
|
|
45
|
-
|
|
46
|
-
## 标识符
|
|
47
|
-
|
|
48
|
-
| 类型 | 示例 |
|
|
49
|
-
|------|------|
|
|
50
|
-
| DOI | `10.1038/nature12373` |
|
|
51
|
-
| PMID | `34567890` |
|
|
52
|
-
| arXiv | `2103.15348` |
|
|
53
|
-
|
|
54
|
-
## 与 web_search 分工
|
|
55
|
-
|
|
56
|
-
- 行业案例、企业动态、中文工艺资讯 → `web_search`
|
|
57
|
-
- 论文题名、摘要、DOI、引用链 → 本 Skill
|
|
58
|
-
|
|
59
|
-
详细 API 示例见 `references/pubmed.md`、`references/crossref.md` 等(已从上游复制)。
|