@pikaa-ai/pikaa 0.3.22 → 0.3.24
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/assets/brand/orbit-logo-option4-whale.jpg +0 -0
- package/assets/brand/orbit-logo.jpg +0 -0
- package/assets/brand/orbit-logo.png +0 -0
- package/assets/brand/orbit-logo.svg +3 -0
- package/dist/cli.js +448 -181
- package/dist/index.js +22 -2
- package/package.json +1 -2
- package/skills/adaptyv/SKILL.md +0 -240
- package/skills/aeon/SKILL.md +0 -402
- package/skills/analytical-method-validation/SKILL.md +0 -299
- package/skills/anndata/SKILL.md +0 -431
- package/skills/arbor/SKILL.md +0 -152
- package/skills/arboreto/SKILL.md +0 -267
- package/skills/astropy/SKILL.md +0 -353
- package/skills/autoskill/SKILL.md +0 -233
- package/skills/benchling-integration/SKILL.md +0 -229
- package/skills/bgpt-paper-search/SKILL.md +0 -75
- package/skills/bids/SKILL.md +0 -237
- package/skills/biopython/SKILL.md +0 -472
- package/skills/bioservices/SKILL.md +0 -399
- package/skills/bulk-rnaseq/SKILL.md +0 -198
- package/skills/cellxgene-census/SKILL.md +0 -283
- package/skills/cirq/SKILL.md +0 -370
- package/skills/citation-management/SKILL.md +0 -329
- package/skills/clinical-decision-support/SKILL.md +0 -238
- package/skills/clinical-decision-support/references/README.md +0 -62
- package/skills/clinical-reports/SKILL.md +0 -248
- package/skills/clinical-reports/references/README.md +0 -34
- package/skills/cobrapy/SKILL.md +0 -496
- package/skills/consciousness-council/SKILL.md +0 -151
- package/skills/dask/SKILL.md +0 -482
- package/skills/database-lookup/SKILL.md +0 -386
- package/skills/datamol/SKILL.md +0 -200
- package/skills/deepchem/SKILL.md +0 -244
- package/skills/deepspot-m/SKILL.md +0 -175
- package/skills/deeptools/SKILL.md +0 -412
- package/skills/depmap/SKILL.md +0 -301
- package/skills/dhdna-profiler/SKILL.md +0 -184
- package/skills/diffdock/SKILL.md +0 -488
- package/skills/dnanexus-integration/SKILL.md +0 -325
- package/skills/docx/SKILL.md +0 -99
- package/skills/esm/SKILL.md +0 -334
- package/skills/etetoolkit/SKILL.md +0 -327
- package/skills/exa-search/SKILL.md +0 -102
- package/skills/executing-plans/SKILL.md +0 -14
- package/skills/experimental-design/SKILL.md +0 -234
- package/skills/exploratory-data-analysis/SKILL.md +0 -280
- package/skills/flowio/SKILL.md +0 -310
- package/skills/fluidsim/SKILL.md +0 -279
- package/skills/frontend-design/SKILL.md +0 -100
- package/skills/generate-image/SKILL.md +0 -304
- package/skills/geniml/SKILL.md +0 -310
- package/skills/genomic-coordinates/SKILL.md +0 -189
- package/skills/genomic-intelligence/SKILL.md +0 -243
- package/skills/geomaster/README.md +0 -105
- package/skills/geomaster/SKILL.md +0 -366
- package/skills/geopandas/SKILL.md +0 -250
- package/skills/get-available-resources/SKILL.md +0 -260
- package/skills/gget/SKILL.md +0 -153
- package/skills/ginkgo-cloud-lab/SKILL.md +0 -106
- package/skills/glycoengineering/SKILL.md +0 -339
- package/skills/gtars/SKILL.md +0 -282
- package/skills/guardian-rails/SKILL.md +0 -54
- package/skills/histolab/SKILL.md +0 -243
- package/skills/hugging-science/SKILL.md +0 -132
- package/skills/hypogenic/SKILL.md +0 -290
- package/skills/hypothesis-generation/SKILL.md +0 -264
- package/skills/imaging-data-commons/SKILL.md +0 -496
- package/skills/infographics/SKILL.md +0 -315
- package/skills/iso-standards-readiness/SKILL.md +0 -352
- package/skills/lab-hardware-cad/SKILL.md +0 -372
- package/skills/labarchive-integration/SKILL.md +0 -216
- package/skills/lamindb/SKILL.md +0 -408
- package/skills/latchbio-integration/SKILL.md +0 -227
- package/skills/latex-posters/SKILL.md +0 -369
- package/skills/latex-posters/references/README.md +0 -439
- package/skills/liteparse/SKILL.md +0 -295
- package/skills/literature-review/SKILL.md +0 -263
- package/skills/markdown-mermaid-writing/SKILL.md +0 -322
- package/skills/market-research-reports/SKILL.md +0 -337
- package/skills/markitdown/SKILL.md +0 -264
- package/skills/matchms/SKILL.md +0 -276
- package/skills/matlab/SKILL.md +0 -274
- package/skills/matplotlib/SKILL.md +0 -378
- package/skills/medchem/SKILL.md +0 -321
- package/skills/modal/SKILL.md +0 -468
- package/skills/molecular-dynamics/SKILL.md +0 -458
- package/skills/molfeat/SKILL.md +0 -348
- package/skills/ncats-arax/SKILL.md +0 -178
- package/skills/networkx/SKILL.md +0 -440
- package/skills/neurokit2/SKILL.md +0 -323
- package/skills/neuropixels-analysis/SKILL.md +0 -412
- package/skills/nextflow/SKILL.md +0 -195
- package/skills/omero-integration/SKILL.md +0 -222
- package/skills/onekgpd/SKILL.md +0 -371
- package/skills/ontology-term-resolution/SKILL.md +0 -147
- package/skills/open-notebook/SKILL.md +0 -297
- package/skills/openpiv/SKILL.md +0 -469
- package/skills/opentrons-integration/SKILL.md +0 -322
- package/skills/optimize-for-gpu/SKILL.md +0 -176
- package/skills/owasp-top10/SKILL.md +0 -48
- package/skills/pacsomatic/LICENSE +0 -21
- package/skills/pacsomatic/SKILL.md +0 -150
- package/skills/paper-lookup/SKILL.md +0 -263
- package/skills/paperclip/SKILL.md +0 -413
- package/skills/paperzilla/SKILL.md +0 -159
- package/skills/parallel-web/SKILL.md +0 -128
- package/skills/pathml/SKILL.md +0 -222
- package/skills/pathogen-variant-surveillance/SKILL.md +0 -208
- package/skills/pathway-enrichment/SKILL.md +0 -194
- package/skills/pdf/SKILL.md +0 -322
- package/skills/peer-review/SKILL.md +0 -288
- package/skills/penetration-testing/SKILL.md +0 -31
- package/skills/pennylane/SKILL.md +0 -240
- package/skills/phylogenetics/SKILL.md +0 -409
- package/skills/pi-agent/SKILL.md +0 -83
- package/skills/pkpd-modeling/SKILL.md +0 -381
- package/skills/polars/SKILL.md +0 -393
- package/skills/polars-bio/SKILL.md +0 -379
- package/skills/ponytail/SKILL.md +0 -31
- package/skills/ponytail-audit/SKILL.md +0 -18
- package/skills/pptx/SKILL.md +0 -246
- package/skills/pptx-posters/SKILL.md +0 -258
- package/skills/primekg/SKILL.md +0 -99
- package/skills/protocolsio-integration/SKILL.md +0 -236
- package/skills/pufferlib/SKILL.md +0 -328
- package/skills/pydeseq2/SKILL.md +0 -369
- package/skills/pydicom/SKILL.md +0 -381
- package/skills/pyhealth/SKILL.md +0 -124
- package/skills/pylabrobot/SKILL.md +0 -216
- package/skills/pymatgen/SKILL.md +0 -404
- package/skills/pymc/SKILL.md +0 -310
- package/skills/pymoo/SKILL.md +0 -276
- package/skills/pyopenms/SKILL.md +0 -179
- package/skills/pysam/SKILL.md +0 -330
- package/skills/pytdc/SKILL.md +0 -297
- package/skills/pytorch-lightning/SKILL.md +0 -191
- package/skills/pyzotero/SKILL.md +0 -137
- package/skills/qiskit/SKILL.md +0 -259
- package/skills/qutip/SKILL.md +0 -317
- package/skills/rdkit/SKILL.md +0 -94
- package/skills/relsa-severity-assessment/SKILL.md +0 -354
- package/skills/research-grants/SKILL.md +0 -296
- package/skills/research-grants/references/README.md +0 -287
- package/skills/research-lookup/README.md +0 -106
- package/skills/research-lookup/SKILL.md +0 -338
- package/skills/rowan/SKILL.md +0 -398
- package/skills/scanpy/SKILL.md +0 -303
- package/skills/scholar-evaluation/SKILL.md +0 -296
- package/skills/scientific-brainstorming/SKILL.md +0 -282
- package/skills/scientific-critical-thinking/SKILL.md +0 -180
- package/skills/scientific-schematics/SKILL.md +0 -370
- package/skills/scientific-slides/SKILL.md +0 -379
- package/skills/scientific-visualization/SKILL.md +0 -285
- package/skills/scientific-writing/SKILL.md +0 -356
- package/skills/scikit-bio/SKILL.md +0 -470
- package/skills/scikit-learn/SKILL.md +0 -324
- package/skills/scikit-survival/SKILL.md +0 -313
- package/skills/scvelo/SKILL.md +0 -328
- package/skills/scvi-tools/SKILL.md +0 -201
- package/skills/seaborn/SKILL.md +0 -254
- package/skills/security-auditor/SKILL.md +0 -37
- package/skills/shap/SKILL.md +0 -282
- package/skills/simpy/SKILL.md +0 -283
- package/skills/stable-baselines3/SKILL.md +0 -325
- package/skills/statistical-analysis/SKILL.md +0 -446
- package/skills/statistical-power/SKILL.md +0 -200
- package/skills/statsmodels/SKILL.md +0 -238
- package/skills/sympy/SKILL.md +0 -354
- package/skills/systematic-debugging/SKILL.md +0 -35
- package/skills/tamarind/SKILL.md +0 -285
- package/skills/tdd/SKILL.md +0 -26
- package/skills/tiledbvcf/SKILL.md +0 -456
- package/skills/timesfm-forecasting/SKILL.md +0 -408
- package/skills/timesfm-forecasting/examples/global-temperature/README.md +0 -178
- package/skills/torch-geometric/SKILL.md +0 -458
- package/skills/torchdrug/SKILL.md +0 -241
- package/skills/transformers/SKILL.md +0 -195
- package/skills/treatment-plans/SKILL.md +0 -174
- package/skills/treatment-plans/references/README.md +0 -19
- package/skills/umap-learn/SKILL.md +0 -488
- package/skills/uncertainty-and-units/SKILL.md +0 -384
- package/skills/usfiscaldata/SKILL.md +0 -171
- package/skills/vaex/SKILL.md +0 -204
- package/skills/venue-templates/SKILL.md +0 -269
- package/skills/verification-before-completion/SKILL.md +0 -22
- package/skills/waypoint-bio/SKILL.md +0 -273
- package/skills/what-if-oracle/SKILL.md +0 -184
- package/skills/writing-plans/SKILL.md +0 -15
- package/skills/xlsx/SKILL.md +0 -110
- package/skills/zarr-python/SKILL.md +0 -241
|
@@ -1,295 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: liteparse
|
|
3
|
-
description: Local document and PDF parsing that returns spatial text with bounding boxes. Use for extracting text from PDFs, DOCX, Office files, and images; running OCR on scans; producing layout-preserved JSON for RAG; batch-ingesting folders of papers; or rendering pages to PNG for multimodal agents. Distinguishing capabilities are per-token bounding boxes, page raster output, and fully local processing with no cloud API.
|
|
4
|
-
license: Apache-2.0
|
|
5
|
-
allowed-tools: Read Write Edit Bash
|
|
6
|
-
compatibility: Python 3.10+. Optional LibreOffice (Office formats) and ImageMagick (images). Bundled Tesseract for OCR. All processing is local — no cloud API required.
|
|
7
|
-
metadata:
|
|
8
|
-
version: "1.1"
|
|
9
|
-
skill-author: K-Dense Inc.
|
|
10
|
-
---
|
|
11
|
-
|
|
12
|
-
# LiteParse — Local Document Parsing
|
|
13
|
-
|
|
14
|
-
## Overview
|
|
15
|
-
|
|
16
|
-
LiteParse is a fast, open-source document parser (Rust core, Python/Node bindings) focused on **local, layout-aware text extraction** with bounding boxes. It does not produce Markdown and does not call cloud LLMs. Outputs are **plain text** (layout-preserved) or **structured JSON** with per-page `text_items` (position, font metadata, optional confidence).
|
|
17
|
-
|
|
18
|
-
**Version note:** Examples target **liteparse 2.0.0** (PyPI, May 2026). The upstream V1 branch is legacy; this skill documents **V2 / main** only.
|
|
19
|
-
|
|
20
|
-
For parser selection vs MarkItDown, the `pdf` skill, or LlamaParse, see `references/choosing_a_parser.md`.
|
|
21
|
-
|
|
22
|
-
## When to Use This Skill
|
|
23
|
-
|
|
24
|
-
Use LiteParse when you need:
|
|
25
|
-
|
|
26
|
-
- **Fast local parsing** of PDFs or converted Office/image files without cloud dependencies
|
|
27
|
-
- **Spatial text** with bounding boxes for layout-aware RAG, citation grounding, or figure/table region logic
|
|
28
|
-
- **OCR** on scanned PDFs or images (bundled Tesseract, or a user-run HTTP OCR server)
|
|
29
|
-
- **Page screenshots** (PNG) for multimodal agents that must see charts, figures, or handwriting
|
|
30
|
-
- **Batch ingestion** of literature folders, supplementary PDFs, or protocol libraries
|
|
31
|
-
- **Page subsets** or **password-protected** PDFs
|
|
32
|
-
|
|
33
|
-
## When Not to Use
|
|
34
|
-
|
|
35
|
-
| Task | Use instead |
|
|
36
|
-
|------|-------------|
|
|
37
|
-
| Markdown for LLM ingestion (EPUB, audio, YouTube, HTML) | `markitdown` skill |
|
|
38
|
-
| Merge/split PDFs, forms, watermarks, rotation | `pdf` skill |
|
|
39
|
-
| Dense tables, handwriting, production cloud pipelines | [LlamaParse](https://docs.cloud.llamaindex.ai/llamaparse/overview) (cloud; sign up separately) |
|
|
40
|
-
|
|
41
|
-
## Installation
|
|
42
|
-
|
|
43
|
-
```bash
|
|
44
|
-
uv pip install "liteparse==2.0.0"
|
|
45
|
-
```
|
|
46
|
-
|
|
47
|
-
This installs the Python bindings and the **`lit`** CLI. Verify:
|
|
48
|
-
|
|
49
|
-
```bash
|
|
50
|
-
lit --help
|
|
51
|
-
python -c "import liteparse; print(liteparse.__version__)"
|
|
52
|
-
```
|
|
53
|
-
|
|
54
|
-
**Optional system tools** (for non-PDF inputs):
|
|
55
|
-
|
|
56
|
-
- **LibreOffice** — Word, Excel, PowerPoint, OpenDocument, CSV/TSV
|
|
57
|
-
- **ImageMagick** — PNG, JPEG, TIFF, WebP, SVG, etc.
|
|
58
|
-
|
|
59
|
-
Install commands are in `references/ocr_and_formats.md`.
|
|
60
|
-
|
|
61
|
-
**Node.js / TypeScript** (optional): `npm i @llamaindex/liteparse` — see `references/api_reference.md`.
|
|
62
|
-
|
|
63
|
-
---
|
|
64
|
-
|
|
65
|
-
## Quick Start
|
|
66
|
-
|
|
67
|
-
### Python
|
|
68
|
-
|
|
69
|
-
```python
|
|
70
|
-
from liteparse import LiteParse
|
|
71
|
-
|
|
72
|
-
parser = LiteParse(quiet=True)
|
|
73
|
-
result = parser.parse("paper.pdf")
|
|
74
|
-
print(result.text)
|
|
75
|
-
|
|
76
|
-
for page in result.pages:
|
|
77
|
-
print(f"Page {page.page_num}: {len(page.text_items)} items")
|
|
78
|
-
```
|
|
79
|
-
|
|
80
|
-
### CLI
|
|
81
|
-
|
|
82
|
-
```bash
|
|
83
|
-
# Layout-preserved text (default)
|
|
84
|
-
lit parse paper.pdf
|
|
85
|
-
|
|
86
|
-
# Structured JSON with bounding boxes
|
|
87
|
-
lit parse paper.pdf --format json -o paper.json
|
|
88
|
-
|
|
89
|
-
# Disable OCR on text-native PDFs (faster)
|
|
90
|
-
lit parse paper.pdf --no-ocr
|
|
91
|
-
```
|
|
92
|
-
|
|
93
|
-
---
|
|
94
|
-
|
|
95
|
-
## Core Workflows
|
|
96
|
-
|
|
97
|
-
### 1. Parse to layout-preserved text
|
|
98
|
-
|
|
99
|
-
Best for quick full-document text or feeding chunkers that do not need coordinates.
|
|
100
|
-
|
|
101
|
-
```python
|
|
102
|
-
parser = LiteParse(ocr_enabled=True, quiet=True)
|
|
103
|
-
result = parser.parse("document.pdf")
|
|
104
|
-
full_text = result.text
|
|
105
|
-
```
|
|
106
|
-
|
|
107
|
-
```bash
|
|
108
|
-
lit parse document.pdf -o output.txt
|
|
109
|
-
```
|
|
110
|
-
|
|
111
|
-
### 2. Parse to structured JSON (bounding boxes)
|
|
112
|
-
|
|
113
|
-
Use when building layout-aware RAG, highlighting source regions, or joining text with screenshots.
|
|
114
|
-
|
|
115
|
-
```python
|
|
116
|
-
import json
|
|
117
|
-
from liteparse import LiteParse
|
|
118
|
-
|
|
119
|
-
parser = LiteParse(output_format="json", quiet=True)
|
|
120
|
-
result = parser.parse("document.pdf")
|
|
121
|
-
|
|
122
|
-
# Programmatic access
|
|
123
|
-
for page in result.pages:
|
|
124
|
-
for item in page.text_items:
|
|
125
|
-
bbox = (item.x, item.y, item.width, item.height)
|
|
126
|
-
# item.text, item.confidence, item.font_name, item.font_size
|
|
127
|
-
```
|
|
128
|
-
|
|
129
|
-
```bash
|
|
130
|
-
lit parse document.pdf --format json -o document.json
|
|
131
|
-
```
|
|
132
|
-
|
|
133
|
-
JSON field layout: `references/output_formats.md`.
|
|
134
|
-
|
|
135
|
-
### 3. Parse specific pages
|
|
136
|
-
|
|
137
|
-
```python
|
|
138
|
-
parser = LiteParse(target_pages="1-5,10,15-20", quiet=True)
|
|
139
|
-
result = parser.parse("long_paper.pdf")
|
|
140
|
-
```
|
|
141
|
-
|
|
142
|
-
```bash
|
|
143
|
-
lit parse long_paper.pdf --target-pages "1-5,10"
|
|
144
|
-
```
|
|
145
|
-
|
|
146
|
-
### 4. Parse from bytes or stdin
|
|
147
|
-
|
|
148
|
-
Useful for uploads, S3 downloads, or piping remote PDFs.
|
|
149
|
-
|
|
150
|
-
```python
|
|
151
|
-
with open("document.pdf", "rb") as f:
|
|
152
|
-
result = parser.parse(f.read())
|
|
153
|
-
```
|
|
154
|
-
|
|
155
|
-
```bash
|
|
156
|
-
curl -sL https://example.com/report.pdf | lit parse -
|
|
157
|
-
```
|
|
158
|
-
|
|
159
|
-
### 5. Page screenshots for multimodal agents
|
|
160
|
-
|
|
161
|
-
Screenshots capture visual content that text extraction alone misses (figures, complex tables, handwriting).
|
|
162
|
-
|
|
163
|
-
```python
|
|
164
|
-
from pathlib import Path
|
|
165
|
-
|
|
166
|
-
parser = LiteParse(dpi=150, quiet=True)
|
|
167
|
-
shots = parser.screenshot("document.pdf", page_numbers=[1, 2, 3])
|
|
168
|
-
out = Path("screenshots")
|
|
169
|
-
out.mkdir(exist_ok=True)
|
|
170
|
-
for s in shots:
|
|
171
|
-
(out / f"page_{s.page_num}.png").write_bytes(s.image_bytes)
|
|
172
|
-
```
|
|
173
|
-
|
|
174
|
-
```bash
|
|
175
|
-
lit screenshot document.pdf --target-pages "1,3,5" -o ./screenshots
|
|
176
|
-
lit screenshot document.pdf --dpi 300 -o ./screenshots
|
|
177
|
-
```
|
|
178
|
-
|
|
179
|
-
Combine **JSON parse + screenshots** when an agent needs both coordinates and pixels for the same pages.
|
|
180
|
-
|
|
181
|
-
### 6. Batch-parse a directory
|
|
182
|
-
|
|
183
|
-
For large corpora, prefer the CLI (parallel OCR workers) or the bundled script.
|
|
184
|
-
|
|
185
|
-
```bash
|
|
186
|
-
lit batch-parse ./papers ./parsed --format json --recursive
|
|
187
|
-
lit batch-parse ./papers ./parsed --extension .pdf --no-ocr
|
|
188
|
-
```
|
|
189
|
-
|
|
190
|
-
```bash
|
|
191
|
-
python scripts/batch_parse_dir.py ./papers ./parsed --format json --recursive
|
|
192
|
-
```
|
|
193
|
-
|
|
194
|
-
See `scripts/batch_parse_dir.py` for a Python batch wrapper without network calls.
|
|
195
|
-
|
|
196
|
-
### 7. OCR configuration
|
|
197
|
-
|
|
198
|
-
OCR is **on by default**. Tesseract is bundled; no extra install for basic English OCR.
|
|
199
|
-
|
|
200
|
-
```python
|
|
201
|
-
parser = LiteParse(
|
|
202
|
-
ocr_enabled=True,
|
|
203
|
-
ocr_language="eng", # Tesseract codes: fra, deu, etc.
|
|
204
|
-
num_workers=4, # parallel OCR (default: CPU cores - 1)
|
|
205
|
-
dpi=150, # higher DPI → better OCR, slower
|
|
206
|
-
)
|
|
207
|
-
```
|
|
208
|
-
|
|
209
|
-
```bash
|
|
210
|
-
lit parse scan.pdf --ocr-language fra
|
|
211
|
-
lit parse scan.pdf --no-ocr
|
|
212
|
-
lit parse scan.pdf --ocr-server-url http://localhost:8080/ocr
|
|
213
|
-
```
|
|
214
|
-
|
|
215
|
-
**Offline / air-gapped:** set `TESSDATA_PREFIX` to a directory of `.traineddata` files, or pass `--tessdata-path`. Details: `references/ocr_and_formats.md`.
|
|
216
|
-
|
|
217
|
-
### 8. Encrypted PDFs
|
|
218
|
-
|
|
219
|
-
```python
|
|
220
|
-
parser = LiteParse(password="secret", quiet=True)
|
|
221
|
-
result = parser.parse("protected.pdf")
|
|
222
|
-
```
|
|
223
|
-
|
|
224
|
-
```bash
|
|
225
|
-
lit parse protected.pdf --password secret
|
|
226
|
-
```
|
|
227
|
-
|
|
228
|
-
### 9. Search text items by phrase
|
|
229
|
-
|
|
230
|
-
Merge adjacent items and return combined bounding boxes for a phrase (e.g. section titles).
|
|
231
|
-
|
|
232
|
-
```python
|
|
233
|
-
from liteparse import search_items
|
|
234
|
-
|
|
235
|
-
page = result.get_page(1)
|
|
236
|
-
matches = search_items(page.text_items, "Materials and Methods", case_sensitive=False)
|
|
237
|
-
```
|
|
238
|
-
|
|
239
|
-
---
|
|
240
|
-
|
|
241
|
-
## Multi-Format Inputs
|
|
242
|
-
|
|
243
|
-
| Category | Extensions (examples) | Requirement |
|
|
244
|
-
|----------|----------------------|-------------|
|
|
245
|
-
| PDF | `.pdf` | Native |
|
|
246
|
-
| Office | `.docx`, `.xlsx`, `.pptx`, `.doc`, `.odt`, … | LibreOffice |
|
|
247
|
-
| Images | `.png`, `.jpg`, `.tiff`, `.webp`, `.svg`, … | ImageMagick |
|
|
248
|
-
|
|
249
|
-
Files are converted to PDF internally, then parsed. If conversion tools are missing, parsing fails with an actionable error — install the dependency and retry.
|
|
250
|
-
|
|
251
|
-
---
|
|
252
|
-
|
|
253
|
-
## Performance Tips
|
|
254
|
-
|
|
255
|
-
- **`--no-ocr`** on born-digital PDFs — largest speedup
|
|
256
|
-
- **`target_pages`** — parse only methods/supplement sections
|
|
257
|
-
- **`num_workers`** — scale OCR across CPU cores
|
|
258
|
-
- **`max_pages`** — cap very large files (default 1000)
|
|
259
|
-
- **`lit batch-parse`** — directory-scale jobs with `--recursive` and `--extension`
|
|
260
|
-
- Lower **`dpi`** (e.g. 100) when OCR quality is already sufficient
|
|
261
|
-
|
|
262
|
-
---
|
|
263
|
-
|
|
264
|
-
## Reference Files
|
|
265
|
-
|
|
266
|
-
| File | Read when |
|
|
267
|
-
|------|-----------|
|
|
268
|
-
| `references/choosing_a_parser.md` | Unsure whether to use LiteParse, MarkItDown, pdf, or LlamaParse |
|
|
269
|
-
| `references/api_reference.md` | Python/TypeScript API, types, `search_items` |
|
|
270
|
-
| `references/cli_reference.md` | Full `lit` command flags |
|
|
271
|
-
| `references/output_formats.md` | JSON schema, bboxes, confidence scores |
|
|
272
|
-
| `references/ocr_and_formats.md` | Tesseract, HTTP OCR, LibreOffice, ImageMagick |
|
|
273
|
-
|
|
274
|
-
---
|
|
275
|
-
|
|
276
|
-
## Troubleshooting
|
|
277
|
-
|
|
278
|
-
| Issue | Fix |
|
|
279
|
-
|-------|-----|
|
|
280
|
-
| Office file fails | Install LibreOffice; ensure `soffice` is on PATH (Windows: add LibreOffice `program` dir) |
|
|
281
|
-
| Image fails | Install ImageMagick; verify `convert` or `magick` works |
|
|
282
|
-
| OCR poor quality | Increase `--dpi`; try `--ocr-language`; or HTTP OCR server |
|
|
283
|
-
| OCR slow | `--no-ocr` if not needed; reduce pages; increase `num_workers` |
|
|
284
|
-
| Air-gapped OCR | `export TESSDATA_PREFIX=/path/to/tessdata` or `--tessdata-path` |
|
|
285
|
-
| `ParseError` on bytes | Ensure input is valid PDF bytes (Office bytes need a file path + conversion) |
|
|
286
|
-
|
|
287
|
-
---
|
|
288
|
-
|
|
289
|
-
## Resources
|
|
290
|
-
|
|
291
|
-
- **GitHub**: https://github.com/run-llama/liteparse
|
|
292
|
-
- **Docs**: https://developers.llamaindex.ai/liteparse/
|
|
293
|
-
- **PyPI**: https://pypi.org/project/liteparse/2.0.0/
|
|
294
|
-
- **npm**: https://www.npmjs.com/package/@llamaindex/liteparse
|
|
295
|
-
- **OCR API spec**: https://github.com/run-llama/liteparse/blob/main/OCR_API_SPEC.md
|
|
@@ -1,263 +0,0 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: literature-review
|
|
3
|
-
description: Conduct comprehensive, systematic literature reviews using multiple academic databases (PubMed, arXiv, bioRxiv, Semantic Scholar, etc.). This skill should be used when conducting systematic literature reviews, meta-analyses, research synthesis, or comprehensive literature searches across biomedical, scientific, and technical domains. Creates professionally formatted markdown documents and PDFs with verified citations in multiple citation styles (APA, Nature, Vancouver, etc.).
|
|
4
|
-
allowed-tools: Read Write Edit Bash
|
|
5
|
-
license: MIT license
|
|
6
|
-
metadata:
|
|
7
|
-
version: "1.7"
|
|
8
|
-
skill-author: K-Dense Inc.
|
|
9
|
-
openclaw:
|
|
10
|
-
primaryEnv: OPENROUTER_API_KEY
|
|
11
|
-
envVars:
|
|
12
|
-
- name: OPENROUTER_API_KEY
|
|
13
|
-
required: false
|
|
14
|
-
description: OpenRouter API key for the skill's LLM-powered steps.
|
|
15
|
-
---
|
|
16
|
-
|
|
17
|
-
# Literature Review
|
|
18
|
-
|
|
19
|
-
## Overview
|
|
20
|
-
|
|
21
|
-
Conduct systematic, comprehensive literature reviews following rigorous academic methodology. Search multiple literature databases, synthesize findings thematically, verify all citations for accuracy, and generate professional output documents in markdown and PDF formats.
|
|
22
|
-
|
|
23
|
-
This skill uses the **parallel-web skill** (`parallel-cli search`) as the primary web search tool for broad academic literature discovery, supplemented by specialized database access skills (gget, bioservices, datacommons-client). It provides specialized tools for citation verification, result aggregation, and document generation.
|
|
24
|
-
|
|
25
|
-
## When to Use This Skill
|
|
26
|
-
|
|
27
|
-
Use this skill when:
|
|
28
|
-
- Conducting a systematic literature review for research or publication
|
|
29
|
-
- Synthesizing current knowledge on a specific topic across multiple sources
|
|
30
|
-
- Performing meta-analysis or scoping reviews
|
|
31
|
-
- Writing the literature review section of a research paper or thesis
|
|
32
|
-
- Investigating the state of the art in a research domain
|
|
33
|
-
- Identifying research gaps and future directions
|
|
34
|
-
- Requiring verified citations and professional formatting
|
|
35
|
-
|
|
36
|
-
## Visual Enhancement with Scientific Schematics
|
|
37
|
-
|
|
38
|
-
**⚠️ MANDATORY: Every literature review MUST include at least 1-2 AI-generated figures using the scientific-schematics skill.**
|
|
39
|
-
|
|
40
|
-
This is not optional. Literature reviews without visual elements are incomplete. Before finalizing any document:
|
|
41
|
-
1. Generate at minimum ONE schematic or diagram (e.g., PRISMA flow diagram for systematic reviews)
|
|
42
|
-
2. Prefer 2-3 figures for comprehensive reviews (search strategy flowchart, thematic synthesis diagram, conceptual framework)
|
|
43
|
-
|
|
44
|
-
**How to generate figures:**
|
|
45
|
-
- Use the **scientific-schematics** skill to generate AI-powered publication-quality diagrams
|
|
46
|
-
- Simply describe your desired diagram in natural language
|
|
47
|
-
- Nano Banana Pro will automatically generate, review, and refine the schematic
|
|
48
|
-
|
|
49
|
-
**How to generate schematics:**
|
|
50
|
-
```bash
|
|
51
|
-
python scripts/generate_schematic.py "your diagram description" -o figures/output.png
|
|
52
|
-
```
|
|
53
|
-
|
|
54
|
-
The AI will automatically:
|
|
55
|
-
- Create publication-quality images with proper formatting
|
|
56
|
-
- Review and refine through multiple iterations
|
|
57
|
-
- Ensure accessibility (colorblind-friendly, high contrast)
|
|
58
|
-
- Save outputs in the figures/ directory
|
|
59
|
-
|
|
60
|
-
**When to add schematics:**
|
|
61
|
-
- PRISMA flow diagrams for systematic reviews
|
|
62
|
-
- Literature search strategy flowcharts
|
|
63
|
-
- Thematic synthesis diagrams
|
|
64
|
-
- Research gap visualization maps
|
|
65
|
-
- Citation network diagrams
|
|
66
|
-
- Conceptual framework illustrations
|
|
67
|
-
- Any complex concept that benefits from visualization
|
|
68
|
-
|
|
69
|
-
For detailed guidance on creating schematics, refer to the scientific-schematics skill documentation.
|
|
70
|
-
|
|
71
|
-
---
|
|
72
|
-
|
|
73
|
-
## Core Workflow
|
|
74
|
-
|
|
75
|
-
A literature review runs in seven phases, documented in full with commands and templates
|
|
76
|
-
in [references/core_workflow.md](references/core_workflow.md):
|
|
77
|
-
|
|
78
|
-
1. **Planning and scoping** — the question, inclusion and exclusion criteria, and scope.
|
|
79
|
-
2. **Systematic literature search** — multi-database searching with recorded queries.
|
|
80
|
-
3. **Screening and selection** — title/abstract then full-text screening with counts kept
|
|
81
|
-
for the PRISMA flow.
|
|
82
|
-
4. **Data extraction and quality assessment** — structured extraction and risk-of-bias
|
|
83
|
-
or quality appraisal.
|
|
84
|
-
5. **Synthesis and analysis** — thematic or quantitative synthesis across studies.
|
|
85
|
-
6. **Citation verification** — every citation checked against the actual source.
|
|
86
|
-
7. **Document generation** — assembling the review with a complete bibliography.
|
|
87
|
-
|
|
88
|
-
Record every search string and date as you go: a review that cannot reproduce its own
|
|
89
|
-
search is not systematic. Per-database search guidance and citation styles are in
|
|
90
|
-
[references/search_and_citation.md](references/search_and_citation.md), and a full worked
|
|
91
|
-
review is in [references/example_workflow.md](references/example_workflow.md).
|
|
92
|
-
|
|
93
|
-
## Best Practices
|
|
94
|
-
|
|
95
|
-
### Search Strategy
|
|
96
|
-
1. **Start with parallel-web**: Use `parallel-cli search` with academic domains for initial broad coverage before querying specialized databases
|
|
97
|
-
2. **Use multiple databases** (minimum 3): Ensures comprehensive coverage — parallel-web counts as one source
|
|
98
|
-
3. **Include preprint servers**: Captures latest unpublished findings
|
|
99
|
-
4. **Document everything**: Search strings, dates, result counts for reproducibility — save all parallel-cli output to `sources/`
|
|
100
|
-
5. **Test and refine**: Run pilot searches, review results, adjust search terms
|
|
101
|
-
6. **Sort by citations**: When available, sort search results by citation count to surface influential work first
|
|
102
|
-
7. **Use parallel-cli extract**: Fetch full content from promising URLs found during search to verify relevance before full-text screening
|
|
103
|
-
|
|
104
|
-
### Screening and Selection
|
|
105
|
-
1. **Use multiple databases** (minimum 3): Ensures comprehensive coverage
|
|
106
|
-
2. **Include preprint servers**: Captures latest unpublished findings
|
|
107
|
-
3. **Document everything**: Search strings, dates, result counts for reproducibility
|
|
108
|
-
4. **Test and refine**: Run pilot searches, review results, adjust search terms
|
|
109
|
-
|
|
110
|
-
### Screening and Selection
|
|
111
|
-
1. **Use clear criteria**: Document inclusion/exclusion criteria before screening
|
|
112
|
-
2. **Screen systematically**: Title → Abstract → Full text
|
|
113
|
-
3. **Document exclusions**: Record reasons for excluding studies
|
|
114
|
-
4. **Consider dual screening**: For systematic reviews, have two reviewers screen independently
|
|
115
|
-
|
|
116
|
-
### Synthesis
|
|
117
|
-
1. **Organize thematically**: Group by themes, NOT by individual studies
|
|
118
|
-
2. **Synthesize across studies**: Compare, contrast, identify patterns
|
|
119
|
-
3. **Be critical**: Evaluate quality and consistency of evidence
|
|
120
|
-
4. **Identify gaps**: Note what's missing or understudied
|
|
121
|
-
|
|
122
|
-
### Quality and Reproducibility
|
|
123
|
-
1. **Assess study quality**: Use appropriate quality assessment tools
|
|
124
|
-
2. **Verify all citations**: Run verify_citations.py script
|
|
125
|
-
3. **Document methodology**: Provide enough detail for others to reproduce
|
|
126
|
-
4. **Follow guidelines**: Use PRISMA for systematic reviews
|
|
127
|
-
|
|
128
|
-
### Writing
|
|
129
|
-
1. **Be objective**: Present evidence fairly, acknowledge limitations
|
|
130
|
-
2. **Be systematic**: Follow structured template
|
|
131
|
-
3. **Be specific**: Include numbers, statistics, effect sizes where available
|
|
132
|
-
4. **Be clear**: Use clear headings, logical flow, thematic organization
|
|
133
|
-
|
|
134
|
-
## Common Pitfalls to Avoid
|
|
135
|
-
|
|
136
|
-
1. **Single database search**: Misses relevant papers; always search multiple databases
|
|
137
|
-
2. **No search documentation**: Makes review irreproducible; document all searches
|
|
138
|
-
3. **Study-by-study summary**: Lacks synthesis; organize thematically instead
|
|
139
|
-
4. **Unverified citations**: Leads to errors; always run verify_citations.py
|
|
140
|
-
5. **Too broad search**: Yields thousands of irrelevant results; refine with specific terms
|
|
141
|
-
6. **Too narrow search**: Misses relevant papers; include synonyms and related terms
|
|
142
|
-
7. **Ignoring preprints**: Misses latest findings; include bioRxiv, medRxiv, arXiv
|
|
143
|
-
8. **No quality assessment**: Treats all evidence equally; assess and report quality
|
|
144
|
-
9. **Publication bias**: Only positive results published; note potential bias
|
|
145
|
-
10. **Outdated search**: Field evolves rapidly; clearly state search date
|
|
146
|
-
|
|
147
|
-
## Integration with Other Skills
|
|
148
|
-
|
|
149
|
-
This skill works seamlessly with other scientific skills:
|
|
150
|
-
|
|
151
|
-
### Web Search & Extraction (parallel-web skill — PRIMARY)
|
|
152
|
-
- **parallel-cli search**: Broad academic and general web search with domain filtering — use for initial scoping, finding papers, citation chaining, and supplementary searches
|
|
153
|
-
- **parallel-cli extract**: Fetch full content from paper URLs, journal websites, and preprint servers — use for reading abstracts, extracting reference lists, and verifying paper details
|
|
154
|
-
- **parallel-cli search --include-domains**: Academic-focused search across scholarly domains (arxiv.org, pubmed, nature.com, etc.)
|
|
155
|
-
|
|
156
|
-
### Database Access Skills
|
|
157
|
-
- **gget**: PubMed, bioRxiv, COSMIC, AlphaFold, Ensembl, UniProt
|
|
158
|
-
- **bioservices**: ChEMBL, KEGG, Reactome, UniProt, PubChem
|
|
159
|
-
- **datacommons-client**: Demographics, economics, health statistics
|
|
160
|
-
|
|
161
|
-
### Analysis Skills
|
|
162
|
-
- **pydeseq2**: RNA-seq differential expression (for methods sections)
|
|
163
|
-
- **scanpy**: Single-cell analysis (for methods sections)
|
|
164
|
-
- **anndata**: Single-cell data (for methods sections)
|
|
165
|
-
- **biopython**: Sequence analysis (for background sections)
|
|
166
|
-
|
|
167
|
-
### Visualization Skills
|
|
168
|
-
- **matplotlib**: Generate figures and plots for review
|
|
169
|
-
- **seaborn**: Statistical visualizations
|
|
170
|
-
|
|
171
|
-
### Writing Skills
|
|
172
|
-
- **brand-guidelines**: Apply institutional branding to PDF
|
|
173
|
-
- **internal-comms**: Adapt review for different audiences
|
|
174
|
-
- **venue-templates**: Access venue-specific writing style guides when preparing reviews for publication
|
|
175
|
-
|
|
176
|
-
### Venue-Specific Writing Styles
|
|
177
|
-
|
|
178
|
-
When preparing a literature review for a specific journal, consult the **venue-templates** skill for writing style guidance:
|
|
179
|
-
- `venue_writing_styles.md`: Master style comparison across venues
|
|
180
|
-
- `nature_science_style.md`: Nature/Science flowing abstract style, story-driven structure
|
|
181
|
-
- `cell_press_style.md`: Cell Press graphical abstracts, Highlights format
|
|
182
|
-
- `medical_journal_styles.md`: NEJM/Lancet/JAMA structured abstracts, PRISMA compliance
|
|
183
|
-
|
|
184
|
-
These guides help adapt your review's tone, abstract format, and structure to match the target venue's expectations.
|
|
185
|
-
|
|
186
|
-
## Resources
|
|
187
|
-
|
|
188
|
-
### Bundled Resources
|
|
189
|
-
|
|
190
|
-
**Scripts:**
|
|
191
|
-
- `scripts/verify_citations.py`: Verify DOIs and generate formatted citations
|
|
192
|
-
- `scripts/generate_pdf.py`: Convert markdown to professional PDF
|
|
193
|
-
- `scripts/search_databases.py`: Process, deduplicate, and format search results
|
|
194
|
-
|
|
195
|
-
**References:**
|
|
196
|
-
- `references/citation_styles.md`: Detailed citation formatting guide (APA, Nature, Vancouver, Chicago, IEEE)
|
|
197
|
-
- `references/database_strategies.md`: Comprehensive database search strategies
|
|
198
|
-
|
|
199
|
-
**Assets:**
|
|
200
|
-
- `assets/review_template.md`: Complete literature review template with all sections
|
|
201
|
-
|
|
202
|
-
### External Resources
|
|
203
|
-
|
|
204
|
-
**Guidelines:**
|
|
205
|
-
- PRISMA (Systematic Reviews): http://www.prisma-statement.org/
|
|
206
|
-
- Cochrane Handbook: https://training.cochrane.org/handbook
|
|
207
|
-
- AMSTAR 2 (Review Quality): https://amstar.ca/
|
|
208
|
-
|
|
209
|
-
**Tools:**
|
|
210
|
-
- MeSH Browser: https://meshb.nlm.nih.gov/search
|
|
211
|
-
- PubMed Advanced Search: https://pubmed.ncbi.nlm.nih.gov/advanced/
|
|
212
|
-
- Boolean Search Guide: https://www.ncbi.nlm.nih.gov/books/NBK3827/
|
|
213
|
-
|
|
214
|
-
**Citation Styles:**
|
|
215
|
-
- APA Style: https://apastyle.apa.org/
|
|
216
|
-
- Nature Portfolio: https://www.nature.com/nature-portfolio/editorial-policies/reporting-standards
|
|
217
|
-
- NLM/Vancouver: https://www.nlm.nih.gov/bsd/uniform_requirements.html
|
|
218
|
-
|
|
219
|
-
## Dependencies
|
|
220
|
-
|
|
221
|
-
### Required CLI Tools
|
|
222
|
-
```bash
|
|
223
|
-
# parallel-cli (PRIMARY — for web search and URL extraction)
|
|
224
|
-
curl -fsSL https://parallel.ai/install.sh | bash
|
|
225
|
-
# Or: uv tool install "parallel-web-tools[cli]"
|
|
226
|
-
# Authenticate: parallel-cli auth
|
|
227
|
-
```
|
|
228
|
-
|
|
229
|
-
### Required Python Packages
|
|
230
|
-
```bash
|
|
231
|
-
uv pip install requests # For citation verification
|
|
232
|
-
```
|
|
233
|
-
|
|
234
|
-
### Required System Tools
|
|
235
|
-
```bash
|
|
236
|
-
# For PDF generation
|
|
237
|
-
brew install pandoc # macOS
|
|
238
|
-
apt-get install pandoc # Linux
|
|
239
|
-
|
|
240
|
-
# For LaTeX (PDF generation)
|
|
241
|
-
brew install --cask mactex # macOS
|
|
242
|
-
apt-get install texlive-xetex # Linux
|
|
243
|
-
```
|
|
244
|
-
|
|
245
|
-
Check dependencies:
|
|
246
|
-
```bash
|
|
247
|
-
python scripts/generate_pdf.py --check-deps
|
|
248
|
-
```
|
|
249
|
-
|
|
250
|
-
## Summary
|
|
251
|
-
|
|
252
|
-
This literature-review skill provides:
|
|
253
|
-
|
|
254
|
-
1. **Systematic methodology** following academic best practices
|
|
255
|
-
2. **Parallel-web powered search** using `parallel-cli search` for fast, broad academic literature discovery with scholarly domain filtering
|
|
256
|
-
3. **Multi-database integration** via existing scientific skills (gget, bioservices, datacommons-client)
|
|
257
|
-
4. **Citation verification** ensuring accuracy and credibility
|
|
258
|
-
5. **Professional output** in markdown and PDF formats
|
|
259
|
-
6. **Comprehensive guidance** covering the entire review process
|
|
260
|
-
7. **Quality assurance** with verification and validation tools
|
|
261
|
-
8. **Reproducibility** through detailed documentation requirements
|
|
262
|
-
|
|
263
|
-
Conduct thorough, rigorous literature reviews that meet academic standards and provide comprehensive synthesis of current knowledge in any domain.
|