clearai-dsh 0.3.0 → 0.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +147 -0
- package/README.md +28 -28
- package/README.zh-CN.md +28 -28
- package/lib/client.js +861 -2555
- package/lib/domain-language.js +552 -48
- package/lib/fold.js +351 -1015
- package/lib/host.js +47 -557
- package/lib/invariant.js +9 -12
- package/lib/knowledge-view.js +363 -226
- package/lib/lang.js +81 -0
- package/package.json +1 -1
- package/presets/clearai/agent.cordis.yml +54 -82
- package/presets/clearai/clearai.patch.yml +54 -82
- package/presets/clearai/plugins/clearai-kernel.js +1590 -5437
- package/presets/clearai/plugins/ontology.js +56 -14
- package/presets/clearai/plugins/prompts.js +73 -363
- package/presets/clearai/skills/clearai-loop/SKILL.md +73 -59
- package/presets/clearai/plugins/brain.js +0 -547
- package/presets/clearai/plugins/commands.js +0 -199
- package/presets/clearai/template/knowledge/README.md +0 -25
- package/presets/clearai/template/memory/README.md +0 -34
- package/presets/clearai/template/project.md +0 -49
- package/presets/clearai/template/skills/README.md +0 -37
- package/presets/clearai/template/skills/chart-diagram-qa/SKILL.md +0 -43
- package/presets/clearai/template/skills/citation-management/SKILL.md +0 -73
- package/presets/clearai/template/skills/citation-management/references/bibtex_formatting.md +0 -908
- package/presets/clearai/template/skills/citation-management/references/citation_validation.md +0 -794
- package/presets/clearai/template/skills/citation-management/references/google_scholar_search.md +0 -725
- package/presets/clearai/template/skills/citation-management/references/metadata_extraction.md +0 -870
- package/presets/clearai/template/skills/citation-management/references/pubmed_search.md +0 -839
- package/presets/clearai/template/skills/citation-management/scripts/doi_to_bibtex.py +0 -204
- package/presets/clearai/template/skills/citation-management/scripts/extract_metadata.py +0 -569
- package/presets/clearai/template/skills/citation-management/scripts/format_bibtex.py +0 -349
- package/presets/clearai/template/skills/citation-management/scripts/generate_schematic.py +0 -139
- package/presets/clearai/template/skills/citation-management/scripts/generate_schematic_ai.py +0 -817
- package/presets/clearai/template/skills/citation-management/scripts/search_google_scholar.py +0 -282
- package/presets/clearai/template/skills/citation-management/scripts/search_pubmed.py +0 -398
- package/presets/clearai/template/skills/citation-management/scripts/validate_citations.py +0 -497
- package/presets/clearai/template/skills/data-analysis/SKILL.md +0 -92
- package/presets/clearai/template/skills/data-analysis/checklists/readiness_check.md +0 -23
- package/presets/clearai/template/skills/data-analysis/templates/analysis_report.md.tpl +0 -63
- package/presets/clearai/template/skills/data-analysis/templates/cleaning_rules_draft.yaml.tpl +0 -32
- package/presets/clearai/template/skills/data-analysis/templates/data_dictionary.md.tpl +0 -12
- package/presets/clearai/template/skills/data-analysis/templates/domain_knowledge_template.md.tpl +0 -316
- package/presets/clearai/template/skills/data-analysis/templates/feature_candidates.json.tpl +0 -20
- package/presets/clearai/template/skills/data-analysis/templates/quality_scorecard.md.tpl +0 -30
- package/presets/clearai/template/skills/data-analysis/workflows/01-data-profiling.md +0 -42
- package/presets/clearai/template/skills/data-analysis/workflows/02-quality-audit.md +0 -36
- package/presets/clearai/template/skills/data-analysis/workflows/03-physical-correlation.md +0 -25
- package/presets/clearai/template/skills/data-analysis/workflows/04-unstructured-mining.md +0 -26
- package/presets/clearai/template/skills/data-qa-analysis/SKILL.md +0 -102
- package/presets/clearai/template/skills/data-qa-analysis/checklists/readiness_check.md +0 -62
- package/presets/clearai/template/skills/data-qa-analysis/templates/best_in_class_report.md.tpl +0 -56
- package/presets/clearai/template/skills/data-qa-analysis/templates/cleaning_rules_draft.yaml.tpl +0 -56
- package/presets/clearai/template/skills/data-qa-analysis/templates/data_dictionary.md.tpl +0 -13
- package/presets/clearai/template/skills/data-qa-analysis/templates/data_source_inventory_and_lineage.md.tpl +0 -146
- package/presets/clearai/template/skills/data-qa-analysis/templates/data_status_report.md.tpl +0 -60
- package/presets/clearai/template/skills/data-qa-analysis/templates/steady_state_rules.yaml.tpl +0 -41
- package/presets/clearai/template/skills/data-qa-analysis/templates/subsystem_registry.md.tpl +0 -101
- package/presets/clearai/template/skills/data-qa-analysis/templates/unified_execution_plan.md.tpl +0 -100
- package/presets/clearai/template/skills/data-qa-analysis/workflows/01-data-source-inventory-and-lineage.md +0 -194
- package/presets/clearai/template/skills/data-qa-analysis/workflows/02-data-alignment-and-tag-semantics.md +0 -122
- package/presets/clearai/template/skills/data-qa-analysis/workflows/03-steady-state-identification.md +0 -126
- package/presets/clearai/template/skills/data-qa-analysis/workflows/04-consumption-analysis.md +0 -152
- package/presets/clearai/template/skills/data-qa-analysis/workflows/05-best-in-class-and-optimization-space.md +0 -78
- package/presets/clearai/template/skills/domain-presearch/SKILL.md +0 -131
- package/presets/clearai/template/skills/domain-presearch/checklists/domain_checklist.md +0 -24
- package/presets/clearai/template/skills/domain-presearch/references/figure_code.md +0 -78
- package/presets/clearai/template/skills/domain-presearch/references/strategic_frameworks.md +0 -38
- package/presets/clearai/template/skills/exploration-loop/SKILL.md +0 -81
- package/presets/clearai/template/skills/exploratory-data-analysis/SKILL.md +0 -77
- package/presets/clearai/template/skills/exploratory-data-analysis/references/bioinformatics_genomics_formats.md +0 -664
- package/presets/clearai/template/skills/exploratory-data-analysis/references/chemistry_molecular_formats.md +0 -664
- package/presets/clearai/template/skills/exploratory-data-analysis/references/general_scientific_formats.md +0 -518
- package/presets/clearai/template/skills/exploratory-data-analysis/references/microscopy_imaging_formats.md +0 -620
- package/presets/clearai/template/skills/exploratory-data-analysis/references/proteomics_metabolomics_formats.md +0 -517
- package/presets/clearai/template/skills/exploratory-data-analysis/references/spectroscopy_analytical_formats.md +0 -633
- package/presets/clearai/template/skills/exploratory-data-analysis/scripts/eda_analyzer.py +0 -547
- package/presets/clearai/template/skills/hypothesis-generation/SKILL.md +0 -73
- package/presets/clearai/template/skills/hypothesis-generation/references/experimental_design_patterns.md +0 -329
- package/presets/clearai/template/skills/hypothesis-generation/references/hypothesis_quality_criteria.md +0 -198
- package/presets/clearai/template/skills/hypothesis-generation/references/literature_search_strategies.md +0 -622
- package/presets/clearai/template/skills/hypothesis-generation/scripts/generate_schematic.py +0 -139
- package/presets/clearai/template/skills/hypothesis-generation/scripts/generate_schematic_ai.py +0 -817
- package/presets/clearai/template/skills/literature-review/SKILL.md +0 -72
- package/presets/clearai/template/skills/literature-review/references/citation_styles.md +0 -166
- package/presets/clearai/template/skills/literature-review/references/database_strategies.md +0 -455
- package/presets/clearai/template/skills/literature-review/scripts/generate_pdf.py +0 -176
- package/presets/clearai/template/skills/literature-review/scripts/generate_schematic.py +0 -139
- package/presets/clearai/template/skills/literature-review/scripts/generate_schematic_ai.py +0 -817
- package/presets/clearai/template/skills/literature-review/scripts/search_databases.py +0 -303
- package/presets/clearai/template/skills/literature-review/scripts/verify_citations.py +0 -221
- package/presets/clearai/template/skills/paper-lookup/SKILL.md +0 -59
- package/presets/clearai/template/skills/paper-lookup/references/arxiv.md +0 -161
- package/presets/clearai/template/skills/paper-lookup/references/biorxiv.md +0 -118
- package/presets/clearai/template/skills/paper-lookup/references/core.md +0 -150
- package/presets/clearai/template/skills/paper-lookup/references/crossref.md +0 -181
- package/presets/clearai/template/skills/paper-lookup/references/medrxiv.md +0 -104
- package/presets/clearai/template/skills/paper-lookup/references/openalex.md +0 -174
- package/presets/clearai/template/skills/paper-lookup/references/pmc.md +0 -152
- package/presets/clearai/template/skills/paper-lookup/references/pubmed.md +0 -124
- package/presets/clearai/template/skills/paper-lookup/references/semantic-scholar.md +0 -203
- package/presets/clearai/template/skills/paper-lookup/references/unpaywall.md +0 -127
- package/presets/clearai/template/skills/process-presearch/SKILL.md +0 -196
- package/presets/clearai/template/skills/process-presearch/checklists/process_checklist.md +0 -18
- package/presets/clearai/template/skills/process-presearch/references/figure_code.md +0 -107
- package/presets/clearai/template/skills/process-presearch/references/source_attribution_example.md +0 -22
- package/presets/clearai/template/skills/process-understanding-extraction/SKILL.md +0 -69
- package/presets/clearai/template/skills/process-understanding-extraction/checklists/readiness_check.md +0 -34
- package/presets/clearai/template/skills/process-understanding-extraction/templates/docx_raw_dump_extractor.py.tpl +0 -132
- package/presets/clearai/template/skills/process-understanding-extraction/templates/entity_map_unit_topology.json.tpl +0 -86
- package/presets/clearai/template/skills/process-understanding-extraction/templates/process_brief.md.tpl +0 -89
- package/presets/clearai/template/skills/process-understanding-extraction/templates/process_brief_builder_from_raw_dump.py.tpl +0 -203
- package/presets/clearai/template/skills/process-understanding-extraction/templates/process_flow_mermaid.md.tpl +0 -41
- package/presets/clearai/template/skills/process-understanding-extraction/templates/unified_execution_plan.md.tpl +0 -53
- package/presets/clearai/template/skills/process-understanding-extraction/workflows/01-process-doc-discovery.md +0 -173
- package/presets/clearai/template/skills/process-understanding-extraction/workflows/02-process-understanding-and-diagramming.md +0 -106
- package/presets/clearai/template/skills/scientific-brainstorming/SKILL.md +0 -64
- package/presets/clearai/template/skills/scientific-brainstorming/references/brainstorming_methods.md +0 -326
- package/presets/clearai/template/skills/scientific-critical-thinking/SKILL.md +0 -72
- package/presets/clearai/template/skills/scientific-critical-thinking/references/common_biases.md +0 -364
- package/presets/clearai/template/skills/scientific-critical-thinking/references/evidence_hierarchy.md +0 -485
- package/presets/clearai/template/skills/scientific-critical-thinking/references/experimental_design.md +0 -496
- package/presets/clearai/template/skills/scientific-critical-thinking/references/logical_fallacies.md +0 -478
- package/presets/clearai/template/skills/scientific-critical-thinking/references/scientific_method.md +0 -169
- package/presets/clearai/template/skills/scientific-critical-thinking/references/statistical_pitfalls.md +0 -506
- package/presets/clearai/template/skills/skill-creator/SKILL.md +0 -109
- package/presets/clearai/template/skills/skill-creator/references/authoring-guide.md +0 -89
- package/presets/clearai/template/skills/statistical-analysis/SKILL.md +0 -79
- package/presets/clearai/template/skills/statistical-analysis/references/assumptions_and_diagnostics.md +0 -369
- package/presets/clearai/template/skills/statistical-analysis/references/bayesian_statistics.md +0 -653
- package/presets/clearai/template/skills/statistical-analysis/references/effect_sizes_and_power.md +0 -578
- package/presets/clearai/template/skills/statistical-analysis/references/reporting_standards.md +0 -469
- package/presets/clearai/template/skills/statistical-analysis/references/test_selection_guide.md +0 -129
- package/presets/clearai/template/skills/statistical-analysis/scripts/assumption_checks.py +0 -538
- package/presets/clearai/template/skills/web-artifact/SKILL.md +0 -165
- package/presets/clearai/template/skills/web-artifact/assets/renderer/renderer.css +0 -229
- package/presets/clearai/template/skills/web-artifact/assets/renderer/renderer.js +0 -373
- package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/LICENSE +0 -263
- package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/UPSTREAM.md +0 -26
- package/presets/clearai/template/skills/web-artifact/assets/vendor/elkjs/elk.bundled.js +0 -6605
- package/presets/clearai/template/skills/web-artifact/references/when-drawing-a-topology.md +0 -150
- package/presets/clearai/template/skills/web-artifact/references/when-the-page-must-work-offline.md +0 -62
- package/presets/clearai/template/skills/web-artifact/scripts/check_artifact.py +0 -167
- package/presets/clearai/template/skills/web-artifact/scripts/render_topology.js +0 -272
- package/presets/clearai/template/skills/what-if-oracle/LICENSE.txt +0 -5
- package/presets/clearai/template/skills/what-if-oracle/SKILL.md +0 -72
- package/presets/clearai/template/skills/what-if-oracle/references/scenario-templates.md +0 -154
|
@@ -1,104 +0,0 @@
|
|
|
1
|
-
# medRxiv API
|
|
2
|
-
|
|
3
|
-
medRxiv is a preprint server for health sciences. The API is identical to bioRxiv's API -- same endpoints, same response format -- just use `medrxiv` as the server parameter.
|
|
4
|
-
|
|
5
|
-
**Important:** Like bioRxiv, there is **no keyword search**. Use Semantic Scholar, OpenAlex, or PubMed for keyword searches of medRxiv content.
|
|
6
|
-
|
|
7
|
-
## Base URL
|
|
8
|
-
|
|
9
|
-
```
|
|
10
|
-
https://api.biorxiv.org
|
|
11
|
-
```
|
|
12
|
-
|
|
13
|
-
(Same base URL as bioRxiv -- the server is specified in the path.)
|
|
14
|
-
|
|
15
|
-
## Authentication
|
|
16
|
-
|
|
17
|
-
None required. Fully public API.
|
|
18
|
-
|
|
19
|
-
## Key Endpoints
|
|
20
|
-
|
|
21
|
-
### 1. Content Detail -- Browse by date range
|
|
22
|
-
|
|
23
|
-
```
|
|
24
|
-
GET /details/medrxiv/{interval}/{cursor}/{format}
|
|
25
|
-
```
|
|
26
|
-
|
|
27
|
-
| Parameter | Values | Description |
|
|
28
|
-
|-----------|--------|-------------|
|
|
29
|
-
| `interval` | `YYYY-MM-DD/YYYY-MM-DD` | Date range (inclusive) |
|
|
30
|
-
| | `N` (integer) | N most recent preprints |
|
|
31
|
-
| | `Nd` (integer + "d") | Last N days |
|
|
32
|
-
| `cursor` | Integer (default `0`) | Pagination offset (100 per page) |
|
|
33
|
-
| `format` | `json` (default), `xml` | Response format |
|
|
34
|
-
|
|
35
|
-
Optional: `?category=cardiovascular%20medicine` (use URL-encoding for spaces)
|
|
36
|
-
|
|
37
|
-
**Examples:**
|
|
38
|
-
```
|
|
39
|
-
https://api.biorxiv.org/details/medrxiv/2024-01-01/2024-01-31/0
|
|
40
|
-
https://api.biorxiv.org/details/medrxiv/5
|
|
41
|
-
https://api.biorxiv.org/details/medrxiv/10d
|
|
42
|
-
```
|
|
43
|
-
|
|
44
|
-
### 2. Content Detail -- DOI lookup
|
|
45
|
-
|
|
46
|
-
```
|
|
47
|
-
GET /details/medrxiv/{doi}/na/{format}
|
|
48
|
-
```
|
|
49
|
-
|
|
50
|
-
**Example:**
|
|
51
|
-
```
|
|
52
|
-
https://api.biorxiv.org/details/medrxiv/10.1101/2021.04.29.21256344/na/json
|
|
53
|
-
```
|
|
54
|
-
|
|
55
|
-
### 3. Published Article Links
|
|
56
|
-
|
|
57
|
-
```
|
|
58
|
-
GET /pubs/medrxiv/{interval}/{cursor}
|
|
59
|
-
GET /pubs/medrxiv/{doi}/na
|
|
60
|
-
```
|
|
61
|
-
|
|
62
|
-
Links preprints to their published journal versions. Accepts both preprint DOI and published DOI.
|
|
63
|
-
|
|
64
|
-
## Response Format
|
|
65
|
-
|
|
66
|
-
Same as bioRxiv:
|
|
67
|
-
|
|
68
|
-
```json
|
|
69
|
-
{
|
|
70
|
-
"messages": [{
|
|
71
|
-
"status": "ok",
|
|
72
|
-
"count": 100,
|
|
73
|
-
"total": "502",
|
|
74
|
-
"cursor": 0
|
|
75
|
-
}],
|
|
76
|
-
"collection": [{
|
|
77
|
-
"title": "Paper title...",
|
|
78
|
-
"authors": "Surname, A.; Surname, B.",
|
|
79
|
-
"author_corresponding": "Full Name",
|
|
80
|
-
"author_corresponding_institution": "Institution",
|
|
81
|
-
"doi": "10.1101/2021.04.29.21256344",
|
|
82
|
-
"date": "2021-05-03",
|
|
83
|
-
"version": "1",
|
|
84
|
-
"type": "PUBLISHAHEADOFPRINT",
|
|
85
|
-
"license": "cc_by_nc_nd",
|
|
86
|
-
"category": "cardiovascular medicine",
|
|
87
|
-
"abstract": "Full abstract text...",
|
|
88
|
-
"published": "10.1371/journal.pone.0256482",
|
|
89
|
-
"server": "medRxiv"
|
|
90
|
-
}]
|
|
91
|
-
}
|
|
92
|
-
```
|
|
93
|
-
|
|
94
|
-
## Pagination
|
|
95
|
-
|
|
96
|
-
100 results per page. Use `cursor` parameter to paginate.
|
|
97
|
-
|
|
98
|
-
## Rate Limits
|
|
99
|
-
|
|
100
|
-
No documented rate limits. No authentication required.
|
|
101
|
-
|
|
102
|
-
## Categories
|
|
103
|
-
|
|
104
|
-
`addiction-medicine`, `allergy-and-immunology`, `anesthesia`, `cardiovascular-medicine`, `dentistry-and-oral-medicine`, `dermatology`, `emergency-medicine`, `endocrinology`, `epidemiology`, `forensic-medicine`, `gastroenterology`, `genetic-and-genomic-medicine`, `geriatric-medicine`, `health-economics`, `health-informatics`, `health-policy`, `health-systems-and-quality-improvement`, `hematology`, `hiv-aids`, `infectious-diseases`, `intensive-care-and-critical-care-medicine`, `medical-education`, `medical-ethics`, `nephrology`, `neurology`, `nursing`, `nutrition`, `obstetrics-and-gynecology`, `occupational-and-environmental-health`, `oncology`, `ophthalmology`, `orthopedics`, `otolaryngology`, `pain-medicine`, `palliative-medicine`, `pathology`, `pediatrics`, `pharmacology-and-therapeutics`, `primary-care-research`, `psychiatry-and-clinical-psychology`, `public-and-global-health`, `radiology-and-imaging`, `rehabilitation-medicine-and-physical-therapy`, `respiratory-medicine`, `rheumatology`, `sexual-and-reproductive-health`, `sports-medicine`, `surgery`, `toxicology`, `transplantation`, `urology`
|
|
@@ -1,174 +0,0 @@
|
|
|
1
|
-
# OpenAlex API
|
|
2
|
-
|
|
3
|
-
OpenAlex is a comprehensive index of 250M+ scholarly works, authors, institutions, sources, and topics. It's the broadest multidisciplinary database in this skill.
|
|
4
|
-
|
|
5
|
-
## Base URL
|
|
6
|
-
|
|
7
|
-
```
|
|
8
|
-
https://api.openalex.org
|
|
9
|
-
```
|
|
10
|
-
|
|
11
|
-
## Authentication
|
|
12
|
-
|
|
13
|
-
- **API key recommended** (free). Get one at https://openalex.org/settings/api
|
|
14
|
-
- Pass as: `?api_key=YOUR_KEY`
|
|
15
|
-
- Legacy polite pool still works: add `?mailto=you@example.com` for better rate limits
|
|
16
|
-
|
|
17
|
-
## Rate Limits
|
|
18
|
-
|
|
19
|
-
- **100 requests/second** max
|
|
20
|
-
- Usage-based pricing with $1/day free allowance
|
|
21
|
-
- Single entity lookups by ID/DOI are free (unlimited)
|
|
22
|
-
- List + filter queries: ~$0.0001 each (~10,000/day free)
|
|
23
|
-
- Search queries: ~$0.001 each (~1,000/day free)
|
|
24
|
-
|
|
25
|
-
## Key Endpoints
|
|
26
|
-
|
|
27
|
-
### 1. Get a single work
|
|
28
|
-
|
|
29
|
-
```
|
|
30
|
-
GET /works/{id}
|
|
31
|
-
```
|
|
32
|
-
|
|
33
|
-
Accepts multiple ID formats:
|
|
34
|
-
```
|
|
35
|
-
/works/W2741809807 (OpenAlex ID)
|
|
36
|
-
/works/doi:10.7717/peerj.4375 (DOI)
|
|
37
|
-
/works/pmid:29456894 (PMID)
|
|
38
|
-
/works/https://doi.org/10.7717/peerj.4375 (full DOI URL)
|
|
39
|
-
```
|
|
40
|
-
|
|
41
|
-
### 2. Search works
|
|
42
|
-
|
|
43
|
-
```
|
|
44
|
-
GET /works?search={query}&per_page={n}&page={n}
|
|
45
|
-
```
|
|
46
|
-
|
|
47
|
-
| Parameter | Default | Description |
|
|
48
|
-
|-----------|---------|-------------|
|
|
49
|
-
| `search` | -- | Full-text search (title, abstract, fulltext). Supports boolean: `AND`, `OR`, `NOT` (uppercase) |
|
|
50
|
-
| `search.exact` | -- | No stemming |
|
|
51
|
-
| `search.semantic` | -- | AI embedding search (beta, 1 req/s, max 50 results) |
|
|
52
|
-
| `filter` | -- | Comma-separated `field:value` pairs |
|
|
53
|
-
| `sort` | relevance | `cited_by_count:desc`, `publication_date:desc`, `relevance_score:desc` |
|
|
54
|
-
| `per_page` | 25 | Results per page (max 100) |
|
|
55
|
-
| `page` | 1 | Page number (max `page * per_page` = 10,000) |
|
|
56
|
-
| `cursor` | -- | Use `*` for first page of deep pagination |
|
|
57
|
-
| `select` | -- | Comma-separated fields to return |
|
|
58
|
-
| `group_by` | -- | Aggregate by field |
|
|
59
|
-
|
|
60
|
-
**Advanced search:** Supports wildcards (`machin*`), fuzzy (`machin~1`), proximity (`"climate change"~5`), boolean grouping.
|
|
61
|
-
|
|
62
|
-
**Example:**
|
|
63
|
-
```
|
|
64
|
-
https://api.openalex.org/works?search=CRISPR+gene+therapy&filter=from_publication_date:2023-01-01&sort=cited_by_count:desc&per_page=10
|
|
65
|
-
```
|
|
66
|
-
|
|
67
|
-
### 3. Filter works
|
|
68
|
-
|
|
69
|
-
```
|
|
70
|
-
GET /works?filter={filters}
|
|
71
|
-
```
|
|
72
|
-
|
|
73
|
-
Key filter fields:
|
|
74
|
-
| Filter | Example | Description |
|
|
75
|
-
|--------|---------|-------------|
|
|
76
|
-
| `from_publication_date` | `2023-01-01` | Published after date |
|
|
77
|
-
| `to_publication_date` | `2024-12-31` | Published before date |
|
|
78
|
-
| `publication_year` | `2024` | Exact year |
|
|
79
|
-
| `type` | `article` | Work type |
|
|
80
|
-
| `cited_by_count` | `>100` | Citation threshold |
|
|
81
|
-
| `is_oa` | `true` | Open access only |
|
|
82
|
-
| `has_abstract` | `true` | Has abstract |
|
|
83
|
-
| `authorships.author.id` | `A5048491430` | By author ID |
|
|
84
|
-
| `primary_location.source.id` | `S137773608` | By journal/source |
|
|
85
|
-
| `institutions.country_code` | `us` | By country |
|
|
86
|
-
| `concepts.id` | `C41008148` | By concept/topic |
|
|
87
|
-
| `doi` | `10.1038/nature12373` | By DOI |
|
|
88
|
-
|
|
89
|
-
**Operators:** `>`, `<`, `!` (negation), `|` (OR within filter)
|
|
90
|
-
|
|
91
|
-
**Example:**
|
|
92
|
-
```
|
|
93
|
-
https://api.openalex.org/works?filter=from_publication_date:2024-01-01,type:article,is_oa:true,cited_by_count:>50
|
|
94
|
-
```
|
|
95
|
-
|
|
96
|
-
### 4. Other entities
|
|
97
|
-
|
|
98
|
-
```
|
|
99
|
-
GET /authors?search={name}
|
|
100
|
-
GET /authors/{id}
|
|
101
|
-
GET /sources?search={name} (journals, repositories)
|
|
102
|
-
GET /sources/{id}
|
|
103
|
-
GET /institutions?search={name}
|
|
104
|
-
GET /institutions/{id}
|
|
105
|
-
GET /topics/{id}
|
|
106
|
-
```
|
|
107
|
-
|
|
108
|
-
Authors and institutions accept similar filter/sort/pagination parameters.
|
|
109
|
-
|
|
110
|
-
### 5. Cursor pagination (for >10,000 results)
|
|
111
|
-
|
|
112
|
-
```
|
|
113
|
-
GET /works?filter=publication_year:2024&cursor=*&per_page=100
|
|
114
|
-
```
|
|
115
|
-
|
|
116
|
-
Response includes `meta.next_cursor`. Pass it as `cursor={value}` in the next request. Stop when `next_cursor` is null.
|
|
117
|
-
|
|
118
|
-
## Response Format
|
|
119
|
-
|
|
120
|
-
### Work object (key fields)
|
|
121
|
-
|
|
122
|
-
```json
|
|
123
|
-
{
|
|
124
|
-
"id": "https://openalex.org/W2741809807",
|
|
125
|
-
"doi": "https://doi.org/10.7717/peerj.4375",
|
|
126
|
-
"title": "The state of OA",
|
|
127
|
-
"publication_year": 2018,
|
|
128
|
-
"publication_date": "2018-02-13",
|
|
129
|
-
"type": "article",
|
|
130
|
-
"language": "en",
|
|
131
|
-
"is_retracted": false,
|
|
132
|
-
"cited_by_count": 1169,
|
|
133
|
-
"open_access": {
|
|
134
|
-
"is_oa": true,
|
|
135
|
-
"oa_status": "gold",
|
|
136
|
-
"oa_url": "https://doi.org/10.7717/peerj.4375"
|
|
137
|
-
},
|
|
138
|
-
"authorships": [{
|
|
139
|
-
"author": {"id": "https://openalex.org/A5048491430", "display_name": "Heather Piwowar"},
|
|
140
|
-
"institutions": [{"display_name": "Impactstory"}]
|
|
141
|
-
}],
|
|
142
|
-
"primary_location": {
|
|
143
|
-
"source": {"display_name": "PeerJ", "issn_l": "2167-8359"}
|
|
144
|
-
},
|
|
145
|
-
"abstract_inverted_index": {"Despite": [0], "growing": [1], "interest": [2], ...},
|
|
146
|
-
"referenced_works": ["https://openalex.org/W123...", ...],
|
|
147
|
-
"ids": {"openalex": "...", "doi": "...", "pmid": "..."}
|
|
148
|
-
}
|
|
149
|
-
```
|
|
150
|
-
|
|
151
|
-
### Abstract inverted index
|
|
152
|
-
|
|
153
|
-
Abstracts are stored as `{word: [positions]}`. To reconstruct:
|
|
154
|
-
```python
|
|
155
|
-
def reconstruct(inverted_index):
|
|
156
|
-
positions = {}
|
|
157
|
-
for word, indices in inverted_index.items():
|
|
158
|
-
for idx in indices:
|
|
159
|
-
positions[idx] = word
|
|
160
|
-
return ' '.join(positions[i] for i in sorted(positions.keys()))
|
|
161
|
-
```
|
|
162
|
-
|
|
163
|
-
### List response
|
|
164
|
-
|
|
165
|
-
```json
|
|
166
|
-
{
|
|
167
|
-
"meta": {"count": 3771834, "page": 1, "per_page": 10},
|
|
168
|
-
"results": [...]
|
|
169
|
-
}
|
|
170
|
-
```
|
|
171
|
-
|
|
172
|
-
## Error Format
|
|
173
|
-
|
|
174
|
-
HTTP 403 for invalid API key, 429 for rate limit exceeded. Error responses include a message field.
|
|
@@ -1,152 +0,0 @@
|
|
|
1
|
-
# PMC (PubMed Central)
|
|
2
|
-
|
|
3
|
-
PMC is a **full-text archive** of biomedical and life sciences articles. It is separate from PubMed -- PubMed has citations/abstracts, PMC has full text. Not all PubMed articles are in PMC, and vice versa.
|
|
4
|
-
|
|
5
|
-
## E-utilities for PMC
|
|
6
|
-
|
|
7
|
-
### Base URL
|
|
8
|
-
|
|
9
|
-
```
|
|
10
|
-
https://eutils.ncbi.nlm.nih.gov/entrez/eutils/
|
|
11
|
-
```
|
|
12
|
-
|
|
13
|
-
Same E-utilities as PubMed, but with `db=pmc`.
|
|
14
|
-
|
|
15
|
-
### eSearch -- Search PMC
|
|
16
|
-
|
|
17
|
-
```
|
|
18
|
-
GET /esearch.fcgi?db=pmc&term={query}&retmode=json
|
|
19
|
-
```
|
|
20
|
-
|
|
21
|
-
Same parameters as PubMed eSearch. Returns PMC UIDs (numeric, e.g., `13033346`). You need to prepend "PMC" to get a PMCID (e.g., `PMC13033346`).
|
|
22
|
-
|
|
23
|
-
### eFetch -- Get Full Text XML
|
|
24
|
-
|
|
25
|
-
```
|
|
26
|
-
GET /efetch.fcgi?db=pmc&id={pmcid}&retmode=xml
|
|
27
|
-
```
|
|
28
|
-
|
|
29
|
-
| rettype | retmode | Returns |
|
|
30
|
-
|---------|---------|---------|
|
|
31
|
-
| *(omit)* | `xml` | **Full text JATS XML** (body, figures, references) |
|
|
32
|
-
| `medline` | `text` | MEDLINE format |
|
|
33
|
-
|
|
34
|
-
**Example:**
|
|
35
|
-
```
|
|
36
|
-
https://eutils.ncbi.nlm.nih.gov/entrez/eutils/efetch.fcgi?db=pmc&id=7029759&retmode=xml
|
|
37
|
-
```
|
|
38
|
-
|
|
39
|
-
The XML uses JATS (Journal Article Tag Suite) format:
|
|
40
|
-
- `<front>` -- journal metadata, article metadata, author info
|
|
41
|
-
- `<body>` -- full article text with `<sec>` sections, `<p>` paragraphs, `<fig>` figures
|
|
42
|
-
- `<back>` -- `<ref-list>` with all references
|
|
43
|
-
|
|
44
|
-
Pass numeric IDs only (not "PMC7029759", just "7029759").
|
|
45
|
-
|
|
46
|
-
## BioC API -- Structured Full Text
|
|
47
|
-
|
|
48
|
-
An alternative way to get full text in a structured passage format.
|
|
49
|
-
|
|
50
|
-
### Base URL
|
|
51
|
-
|
|
52
|
-
```
|
|
53
|
-
https://www.ncbi.nlm.nih.gov/research/bionlp/RESTful/pmcoa.cgi/
|
|
54
|
-
```
|
|
55
|
-
|
|
56
|
-
### Endpoint
|
|
57
|
-
|
|
58
|
-
```
|
|
59
|
-
GET /BioC_{format}/{id}/{encoding}
|
|
60
|
-
```
|
|
61
|
-
|
|
62
|
-
| Parameter | Values |
|
|
63
|
-
|-----------|--------|
|
|
64
|
-
| `format` | `json` or `xml` |
|
|
65
|
-
| `id` | PMID (e.g., `17299597`) or PMCID (e.g., `PMC7029759`) |
|
|
66
|
-
| `encoding` | `unicode` or `ascii` |
|
|
67
|
-
|
|
68
|
-
**Example:**
|
|
69
|
-
```
|
|
70
|
-
https://www.ncbi.nlm.nih.gov/research/bionlp/RESTful/pmcoa.cgi/BioC_json/PMC7029759/unicode
|
|
71
|
-
```
|
|
72
|
-
|
|
73
|
-
**Response structure (JSON):**
|
|
74
|
-
```json
|
|
75
|
-
{
|
|
76
|
-
"source": "PMC",
|
|
77
|
-
"documents": [{
|
|
78
|
-
"id": "PMC7029759",
|
|
79
|
-
"infons": {"license": "...", "doi": "..."},
|
|
80
|
-
"passages": [
|
|
81
|
-
{
|
|
82
|
-
"offset": 0,
|
|
83
|
-
"infons": {"section_type": "TITLE"},
|
|
84
|
-
"text": "Article title..."
|
|
85
|
-
},
|
|
86
|
-
{
|
|
87
|
-
"offset": 42,
|
|
88
|
-
"infons": {"section_type": "ABSTRACT"},
|
|
89
|
-
"text": "Abstract text..."
|
|
90
|
-
},
|
|
91
|
-
{
|
|
92
|
-
"offset": 500,
|
|
93
|
-
"infons": {"section_type": "INTRO"},
|
|
94
|
-
"text": "Introduction text..."
|
|
95
|
-
}
|
|
96
|
-
]
|
|
97
|
-
}]
|
|
98
|
-
}
|
|
99
|
-
```
|
|
100
|
-
|
|
101
|
-
Section types: `TITLE`, `ABSTRACT`, `INTRO`, `METHODS`, `RESULTS`, `DISCUSS`, `CONCL`, `REF`, `SUPPL`, `FIG`, `TABLE`
|
|
102
|
-
|
|
103
|
-
**Coverage:** ~3 million articles from the PMC Open Access Subset.
|
|
104
|
-
|
|
105
|
-
## PMC ID Converter API
|
|
106
|
-
|
|
107
|
-
Converts between PMID, PMCID, DOI, and Manuscript ID.
|
|
108
|
-
|
|
109
|
-
### Base URL
|
|
110
|
-
|
|
111
|
-
```
|
|
112
|
-
https://pmc.ncbi.nlm.nih.gov/tools/idconv/api/v1/articles/
|
|
113
|
-
```
|
|
114
|
-
|
|
115
|
-
### Parameters
|
|
116
|
-
|
|
117
|
-
| Parameter | Required | Description |
|
|
118
|
-
|-----------|----------|-------------|
|
|
119
|
-
| `ids` | Yes | Up to 200 comma-separated IDs |
|
|
120
|
-
| `idtype` | No | `pmcid`, `pmid`, `mid`, `doi` (default: auto-detect) |
|
|
121
|
-
| `format` | No | `json`, `xml`, `csv` (default: xml) |
|
|
122
|
-
| `tool` | Recommended | Your application name |
|
|
123
|
-
| `email` | Recommended | Your contact email |
|
|
124
|
-
|
|
125
|
-
**Example:**
|
|
126
|
-
```
|
|
127
|
-
https://pmc.ncbi.nlm.nih.gov/tools/idconv/api/v1/articles/?ids=PMC7029759&format=json
|
|
128
|
-
```
|
|
129
|
-
|
|
130
|
-
**Response:**
|
|
131
|
-
```json
|
|
132
|
-
{
|
|
133
|
-
"status": "ok",
|
|
134
|
-
"records": [{
|
|
135
|
-
"pmcid": "PMC7029759",
|
|
136
|
-
"pmid": "32117569",
|
|
137
|
-
"doi": "10.12688/f1000research.22211.2"
|
|
138
|
-
}]
|
|
139
|
-
}
|
|
140
|
-
```
|
|
141
|
-
|
|
142
|
-
Only returns results for articles that are in PMC. If an article is in PubMed but not PMC, no PMCID will be returned.
|
|
143
|
-
|
|
144
|
-
## Rate Limits
|
|
145
|
-
|
|
146
|
-
| Service | Limit |
|
|
147
|
-
|---------|-------|
|
|
148
|
-
| E-utilities (`db=pmc`) | 3/sec without key, 10/sec with key |
|
|
149
|
-
| BioC API | Follow general NCBI policy (3/sec without key) |
|
|
150
|
-
| ID Converter | Follow general NCBI policy |
|
|
151
|
-
|
|
152
|
-
Include `tool` and `email` parameters on E-utility requests. Large batch jobs should run outside peak hours (Mon-Fri 5AM-9PM ET).
|
|
@@ -1,124 +0,0 @@
|
|
|
1
|
-
# PubMed (NCBI E-utilities)
|
|
2
|
-
|
|
3
|
-
PubMed provides citations, abstracts, and metadata for 37M+ biomedical and life science articles. It does NOT contain full text -- for that, use PMC.
|
|
4
|
-
|
|
5
|
-
## Base URL
|
|
6
|
-
|
|
7
|
-
```
|
|
8
|
-
https://eutils.ncbi.nlm.nih.gov/entrez/eutils/
|
|
9
|
-
```
|
|
10
|
-
|
|
11
|
-
## Authentication
|
|
12
|
-
|
|
13
|
-
- **API key optional** but recommended. Without: 3 req/sec. With: 10 req/sec.
|
|
14
|
-
- Pass as: `&api_key=YOUR_KEY`
|
|
15
|
-
- Also include `&tool=your_app_name&email=your@email.com` on all requests.
|
|
16
|
-
|
|
17
|
-
## Key Endpoints
|
|
18
|
-
|
|
19
|
-
### 1. eSearch -- Search and get PMIDs
|
|
20
|
-
|
|
21
|
-
```
|
|
22
|
-
GET /esearch.fcgi?db=pubmed&term={query}&retmode=json
|
|
23
|
-
```
|
|
24
|
-
|
|
25
|
-
| Parameter | Required | Default | Description |
|
|
26
|
-
|-----------|----------|---------|-------------|
|
|
27
|
-
| `db` | Yes | -- | `pubmed` |
|
|
28
|
-
| `term` | Yes | -- | Search query. Supports PubMed syntax: field tags `[AU]`, `[TI]`, `[TA]`, `[MH]` (MeSH), boolean AND/OR/NOT |
|
|
29
|
-
| `retmax` | No | 20 | Max PMIDs returned (max 10,000) |
|
|
30
|
-
| `retstart` | No | 0 | Pagination offset |
|
|
31
|
-
| `retmode` | No | `xml` | `json` or `xml` |
|
|
32
|
-
| `rettype` | No | `uilist` | `uilist` (IDs) or `count` (count only) |
|
|
33
|
-
| `sort` | No | `relevance` | `relevance`, `pub_date`, `Author`, `JournalName` |
|
|
34
|
-
| `datetype` | No | -- | `pdat` (publication), `mdat` (modification), `edat` (entrez) |
|
|
35
|
-
| `mindate` / `maxdate` | No | -- | Date range `YYYY/MM/DD` |
|
|
36
|
-
| `reldate` | No | -- | Items from last N days |
|
|
37
|
-
| `usehistory` | No | -- | `y` to store on History Server for large result sets |
|
|
38
|
-
|
|
39
|
-
**Example:**
|
|
40
|
-
```
|
|
41
|
-
https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esearch.fcgi?db=pubmed&term=CRISPR+gene+therapy&retmode=json&retmax=5&sort=pub_date
|
|
42
|
-
```
|
|
43
|
-
|
|
44
|
-
**Response:**
|
|
45
|
-
```json
|
|
46
|
-
{
|
|
47
|
-
"esearchresult": {
|
|
48
|
-
"count": "224107",
|
|
49
|
-
"retmax": "5",
|
|
50
|
-
"retstart": "0",
|
|
51
|
-
"idlist": ["39984857", "39984678", "39984543", "39984210", "39983901"]
|
|
52
|
-
}
|
|
53
|
-
}
|
|
54
|
-
```
|
|
55
|
-
|
|
56
|
-
### 2. eSummary -- Get document summaries
|
|
57
|
-
|
|
58
|
-
```
|
|
59
|
-
GET /esummary.fcgi?db=pubmed&id={pmids}&retmode=json
|
|
60
|
-
```
|
|
61
|
-
|
|
62
|
-
| Parameter | Required | Description |
|
|
63
|
-
|-----------|----------|-------------|
|
|
64
|
-
| `db` | Yes | `pubmed` |
|
|
65
|
-
| `id` | Yes | Comma-separated PMIDs (max 10,000) |
|
|
66
|
-
| `retmode` | No | `json` or `xml` |
|
|
67
|
-
|
|
68
|
-
**Example:**
|
|
69
|
-
```
|
|
70
|
-
https://eutils.ncbi.nlm.nih.gov/entrez/eutils/esummary.fcgi?db=pubmed&id=39984857,39984678&retmode=json
|
|
71
|
-
```
|
|
72
|
-
|
|
73
|
-
**Response fields:** `uid`, `pubdate`, `source` (journal), `authors`, `title`, `volume`, `issue`, `pages`, `fulljournalname`, `elocationid` (DOI), `articleids` (PMC, DOI, etc.), `pubtype`, `pmcrefcount`
|
|
74
|
-
|
|
75
|
-
### 3. eFetch -- Retrieve full records (abstracts, MEDLINE)
|
|
76
|
-
|
|
77
|
-
```
|
|
78
|
-
GET /efetch.fcgi?db=pubmed&id={pmids}&rettype={type}&retmode={mode}
|
|
79
|
-
```
|
|
80
|
-
|
|
81
|
-
| rettype | retmode | Returns |
|
|
82
|
-
|---------|---------|---------|
|
|
83
|
-
| *(omit)* | `xml` | Full PubMed XML (citation + abstract) |
|
|
84
|
-
| `medline` | `text` | MEDLINE format |
|
|
85
|
-
| `abstract` | `text` | Plain text abstract |
|
|
86
|
-
| `uilist` | `text` | PMID list |
|
|
87
|
-
|
|
88
|
-
**Example -- get abstracts as XML:**
|
|
89
|
-
```
|
|
90
|
-
https://eutils.ncbi.nlm.nih.gov/entrez/eutils/efetch.fcgi?db=pubmed&id=39984857&retmode=xml
|
|
91
|
-
```
|
|
92
|
-
|
|
93
|
-
The XML contains `<PubmedArticle>` with `<MedlineCitation>` (title, abstract, MeSH terms, authors) and `<PubmedData>` (article IDs, publication history).
|
|
94
|
-
|
|
95
|
-
### 4. eLink -- Find related articles
|
|
96
|
-
|
|
97
|
-
```
|
|
98
|
-
GET /elink.fcgi?dbfrom=pubmed&db=pubmed&id={pmid}&cmd=neighbor_score&retmode=json
|
|
99
|
-
```
|
|
100
|
-
|
|
101
|
-
Returns related PMIDs with relevance scores.
|
|
102
|
-
|
|
103
|
-
## Search Syntax Tips
|
|
104
|
-
|
|
105
|
-
- **Field tags:** `aspirin[TI]` (title), `Smith J[AU]` (author), `Nature[TA]` (journal), `neoplasms[MH]` (MeSH heading)
|
|
106
|
-
- **Boolean:** `CRISPR AND (therapy OR treatment)`
|
|
107
|
-
- **Date range:** `2020/01/01:2024/12/31[PDAT]`
|
|
108
|
-
- **Publication type:** `review[PT]`, `clinical trial[PT]`
|
|
109
|
-
- **Organism:** `humans[MH]`, `mice[MH]`
|
|
110
|
-
|
|
111
|
-
## Rate Limits
|
|
112
|
-
|
|
113
|
-
- **3 requests/second** without API key
|
|
114
|
-
- **10 requests/second** with API key
|
|
115
|
-
- Include `tool` and `email` parameters on every request
|
|
116
|
-
- Large batch jobs should run outside peak hours (Mon-Fri 5AM-9PM ET)
|
|
117
|
-
|
|
118
|
-
## Error Format
|
|
119
|
-
|
|
120
|
-
```json
|
|
121
|
-
{"error": "API rate limit exceeded", "count": "11"}
|
|
122
|
-
```
|
|
123
|
-
|
|
124
|
-
HTTP 400 for bad requests, 429 for rate limiting.
|