litsurvey 1.0.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- litsurvey-1.0.0/CHANGELOG.md +37 -0
- litsurvey-1.0.0/CITATION.cff +38 -0
- litsurvey-1.0.0/CONTRIBUTING.md +63 -0
- litsurvey-1.0.0/LICENSE +21 -0
- litsurvey-1.0.0/MANIFEST.in +4 -0
- litsurvey-1.0.0/PKG-INFO +203 -0
- litsurvey-1.0.0/README.md +178 -0
- litsurvey-1.0.0/docs/api-keys.md +71 -0
- litsurvey-1.0.0/docs/cli.md +228 -0
- litsurvey-1.0.0/docs/confidentiality.md +95 -0
- litsurvey-1.0.0/docs/llm-integration.md +208 -0
- litsurvey-1.0.0/docs/use-cases.md +158 -0
- litsurvey-1.0.0/docs/web.md +69 -0
- litsurvey-1.0.0/integrations/claude-code/litsurvey/SKILL.md +49 -0
- litsurvey-1.0.0/integrations/codex/AGENTS.md +22 -0
- litsurvey-1.0.0/litsurvey/__init__.py +4 -0
- litsurvey-1.0.0/litsurvey/__main__.py +4 -0
- litsurvey-1.0.0/litsurvey/agent.py +223 -0
- litsurvey-1.0.0/litsurvey/backends.py +291 -0
- litsurvey-1.0.0/litsurvey/cli.py +345 -0
- litsurvey-1.0.0/litsurvey/config.py +99 -0
- litsurvey-1.0.0/litsurvey/export.py +132 -0
- litsurvey-1.0.0/litsurvey/history.py +60 -0
- litsurvey-1.0.0/litsurvey/http.py +160 -0
- litsurvey-1.0.0/litsurvey/ops.py +217 -0
- litsurvey-1.0.0/litsurvey/papers.py +140 -0
- litsurvey-1.0.0/litsurvey/sources/__init__.py +41 -0
- litsurvey-1.0.0/litsurvey/sources/arxiv.py +57 -0
- litsurvey-1.0.0/litsurvey/sources/crossref.py +55 -0
- litsurvey-1.0.0/litsurvey/sources/iacr.py +45 -0
- litsurvey-1.0.0/litsurvey/sources/openalex.py +108 -0
- litsurvey-1.0.0/litsurvey/sources/semanticscholar.py +65 -0
- litsurvey-1.0.0/litsurvey/sources/unpaywall.py +20 -0
- litsurvey-1.0.0/litsurvey/web/index.html +387 -0
- litsurvey-1.0.0/litsurvey/web.py +181 -0
- litsurvey-1.0.0/litsurvey.egg-info/PKG-INFO +203 -0
- litsurvey-1.0.0/litsurvey.egg-info/SOURCES.txt +46 -0
- litsurvey-1.0.0/litsurvey.egg-info/dependency_links.txt +1 -0
- litsurvey-1.0.0/litsurvey.egg-info/entry_points.txt +2 -0
- litsurvey-1.0.0/litsurvey.egg-info/requires.txt +3 -0
- litsurvey-1.0.0/litsurvey.egg-info/top_level.txt +1 -0
- litsurvey-1.0.0/pyproject.toml +44 -0
- litsurvey-1.0.0/setup.cfg +4 -0
- litsurvey-1.0.0/tests/test_backends_agent.py +186 -0
- litsurvey-1.0.0/tests/test_cli_history.py +150 -0
- litsurvey-1.0.0/tests/test_export.py +49 -0
- litsurvey-1.0.0/tests/test_papers.py +121 -0
- litsurvey-1.0.0/tests/test_sources.py +164 -0
|
@@ -0,0 +1,37 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
All notable changes to this project are documented here. The format follows
|
|
4
|
+
[Keep a Changelog](https://keepachangelog.com/) and the project uses
|
|
5
|
+
[Semantic Versioning](https://semver.org/).
|
|
6
|
+
|
|
7
|
+
## [Unreleased]
|
|
8
|
+
|
|
9
|
+
## [1.0.0] - 2026-09-11
|
|
10
|
+
|
|
11
|
+
First public release.
|
|
12
|
+
|
|
13
|
+
### Added
|
|
14
|
+
- `search`: fused, de-duplicated keyword search over OpenAlex, Semantic Scholar,
|
|
15
|
+
arXiv, TechRxiv, Research Square (via Crossref) and the IACR ePrint archive,
|
|
16
|
+
with `--sort relevance|citations|year`.
|
|
17
|
+
- `paper`, `cites`, `refs`, `related`: one paper and its citation neighbourhood;
|
|
18
|
+
a title is accepted in place of an id, with a candidate list and `--pick N`.
|
|
19
|
+
`cites`/`refs` rank by citation count via OpenAlex, with Semantic Scholar as
|
|
20
|
+
fallback; `related` and `paper` fall back to OpenAlex when Semantic Scholar
|
|
21
|
+
is unavailable.
|
|
22
|
+
- `oa`: legal open-access copies via Unpaywall.
|
|
23
|
+
- `novelty` and `research`: LLM-driven agents with pluggable backends and a
|
|
24
|
+
reproducible search log. Backends: `cli` (delegates the task to a signed-in
|
|
25
|
+
Claude Code, Codex CLI or Gemini CLI under the user's subscription),
|
|
26
|
+
`ollama` (local), `openai` (any OpenAI-compatible server, `--base-url` for
|
|
27
|
+
OpenRouter, LM Studio and others), `anthropic`.
|
|
28
|
+
- Export to BibTeX, RIS, CSV, JSON and Markdown from any result list.
|
|
29
|
+
- Run history (`litsurvey history`) shared by the CLI and the web page.
|
|
30
|
+
- `litsurvey web`: local browser interface with a History tab, a photo
|
|
31
|
+
banner, a three-way choice of where the LLM runs (subscription CLI, local
|
|
32
|
+
model, cloud API key) with a privacy banner, BibTeX/RIS/CSV downloads,
|
|
33
|
+
URL detection with DOI/arXiv conversion, "waiting for <source>" status
|
|
34
|
+
while a run is in progress, and a report-a-bug link.
|
|
35
|
+
- `init` and `doctor` commands.
|
|
36
|
+
- Per-host rate limiting (Semantic Scholar 1 req/s, arXiv 1 req/3 s) and
|
|
37
|
+
`Retry-After` handling.
|
|
@@ -0,0 +1,38 @@
|
|
|
1
|
+
cff-version: 1.2.0
|
|
2
|
+
title: "litsurvey: multi-source scientific literature search with LLM-assisted novelty assessment"
|
|
3
|
+
message: >-
|
|
4
|
+
If you use this software in your research, please cite it using the
|
|
5
|
+
metadata below.
|
|
6
|
+
type: software
|
|
7
|
+
authors:
|
|
8
|
+
- family-names: Khankhoje
|
|
9
|
+
given-names: Uday
|
|
10
|
+
email: uday@ee.iitm.ac.in
|
|
11
|
+
affiliation: >-
|
|
12
|
+
Department of Electrical Engineering, Indian Institute of Technology
|
|
13
|
+
Madras
|
|
14
|
+
orcid: 'https://orcid.org/0000-0002-9629-3922'
|
|
15
|
+
repository-code: 'https://github.com/udaykdk/litsurvey'
|
|
16
|
+
abstract: >-
|
|
17
|
+
A zero-dependency Python tool for literature surveys: it searches OpenAlex,
|
|
18
|
+
Semantic Scholar, arXiv, TechRxiv, Research Square and the IACR ePrint
|
|
19
|
+
archive, fuses and de-duplicates the results, walks the citation graph,
|
|
20
|
+
finds open-access copies, and exports BibTeX/RIS/CSV. An LLM (a local
|
|
21
|
+
model, the user's Claude Code / Codex / Gemini subscription CLI, or a cloud
|
|
22
|
+
API) drives multi-step novelty assessment and cited state-of-the-art
|
|
23
|
+
surveys, with a reproducible search log. A local browser interface with
|
|
24
|
+
run history is included. Designed so that confidential manuscript content
|
|
25
|
+
never has to leave the user's machine.
|
|
26
|
+
keywords:
|
|
27
|
+
- literature search
|
|
28
|
+
- peer review
|
|
29
|
+
- systematic review
|
|
30
|
+
- novelty assessment
|
|
31
|
+
- Semantic Scholar
|
|
32
|
+
- OpenAlex
|
|
33
|
+
- arXiv
|
|
34
|
+
- BibTeX
|
|
35
|
+
- large language models
|
|
36
|
+
license: MIT
|
|
37
|
+
version: 1.0.0
|
|
38
|
+
date-released: '2026-09-11'
|
|
@@ -0,0 +1,63 @@
|
|
|
1
|
+
# Contributing
|
|
2
|
+
|
|
3
|
+
## Reporting a bug
|
|
4
|
+
|
|
5
|
+
Open an issue with the bug template. Include the command, the output with
|
|
6
|
+
`--debug`, and the output of `litsurvey --version` and `litsurvey doctor`.
|
|
7
|
+
Remove your API key if it appears anywhere. Most problems are API changes on
|
|
8
|
+
the provider side, and the exact URL from `--debug` is what makes them
|
|
9
|
+
fixable.
|
|
10
|
+
|
|
11
|
+
## Development setup
|
|
12
|
+
|
|
13
|
+
```bash
|
|
14
|
+
git clone https://github.com/udaykdk/litsurvey
|
|
15
|
+
cd litsurvey
|
|
16
|
+
python -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
|
|
17
|
+
pip install -e ".[dev]"
|
|
18
|
+
python -m pytest
|
|
19
|
+
```
|
|
20
|
+
|
|
21
|
+
The tests are offline: every HTTP call is mocked. Please keep it that way, so
|
|
22
|
+
CI does not depend on API availability. If you change how a source is parsed,
|
|
23
|
+
add a fixture for the new response shape in `tests/test_sources.py`.
|
|
24
|
+
|
|
25
|
+
## Design rules
|
|
26
|
+
|
|
27
|
+
- **Zero runtime dependencies.** Standard library only. This keeps
|
|
28
|
+
installation trivial on university machines and inside agents. If a
|
|
29
|
+
feature needs a dependency, make it optional and import it lazily.
|
|
30
|
+
- **Only public, official APIs, with one documented exception.** No
|
|
31
|
+
scraping of sites whose terms forbid it (Google Scholar in particular).
|
|
32
|
+
New sources should have a documented API and a rate limit we can honour in
|
|
33
|
+
`http.HOST_INTERVAL`. The IACR ePrint source reads a server-rendered
|
|
34
|
+
search page because the archive offers no search API and does not
|
|
35
|
+
prohibit it; it is rate-limited to one request per two seconds and is
|
|
36
|
+
isolated so that a markup change only silences that source.
|
|
37
|
+
- **Nothing leaves the machine that the docs do not say leaves.** Any change
|
|
38
|
+
to what is sent where must be reflected in `docs/confidentiality.md`.
|
|
39
|
+
- **Every successfully completed run is recorded** in the history store
|
|
40
|
+
(search runs with per-source hit counts), so results are reproducible.
|
|
41
|
+
|
|
42
|
+
## Adding a source
|
|
43
|
+
|
|
44
|
+
1. Create `litsurvey/sources/<name>.py` with a `search(query, limit,
|
|
45
|
+
year_from)` function that returns a list of `papers.make(...)` records
|
|
46
|
+
with `sources=["<name>"]`.
|
|
47
|
+
2. Register it in `litsurvey/sources/__init__.py` (`SOURCES`).
|
|
48
|
+
3. Add the host's rate limit to `http.HOST_INTERVAL`.
|
|
49
|
+
4. Add a parsing test with a fixture, and a line in `docs/cli.md`.
|
|
50
|
+
|
|
51
|
+
## Pull requests
|
|
52
|
+
|
|
53
|
+
Small, focused pull requests are easier to review. Add a line under
|
|
54
|
+
*Unreleased* in `CHANGELOG.md`. CI runs the tests on Linux, macOS and
|
|
55
|
+
Windows with Python 3.10, 3.12 and 3.13.
|
|
56
|
+
|
|
57
|
+
## Releasing (maintainers)
|
|
58
|
+
|
|
59
|
+
1. Bump `version` in `pyproject.toml`, `litsurvey/__init__.py` and
|
|
60
|
+
`CITATION.cff`; move the changelog entries under the new version.
|
|
61
|
+
2. `git tag -a vX.Y.Z -m "vX.Y.Z"` and push the tag.
|
|
62
|
+
3. Create a GitHub release from the tag; the publish workflow uploads to
|
|
63
|
+
PyPI via trusted publishing.
|
litsurvey-1.0.0/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Uday K Khankhoje
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
litsurvey-1.0.0/PKG-INFO
ADDED
|
@@ -0,0 +1,203 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: litsurvey
|
|
3
|
+
Version: 1.0.0
|
|
4
|
+
Summary: Multi-source scientific literature search with optional LLM-driven novelty assessment and deep research. Zero dependencies.
|
|
5
|
+
Author-email: Uday Khankhoje <uday@ee.iitm.ac.in>
|
|
6
|
+
License: MIT
|
|
7
|
+
Project-URL: Homepage, https://github.com/udaykdk/litsurvey
|
|
8
|
+
Project-URL: Issues, https://github.com/udaykdk/litsurvey/issues
|
|
9
|
+
Project-URL: Changelog, https://github.com/udaykdk/litsurvey/blob/main/CHANGELOG.md
|
|
10
|
+
Keywords: literature search,literature survey,semantic scholar,openalex,arxiv,crossref,peer review,novelty,systematic review,bibtex,llm,ollama,claude code
|
|
11
|
+
Classifier: Development Status :: 4 - Beta
|
|
12
|
+
Classifier: Environment :: Console
|
|
13
|
+
Classifier: Intended Audience :: Science/Research
|
|
14
|
+
Classifier: License :: OSI Approved :: MIT License
|
|
15
|
+
Classifier: Operating System :: OS Independent
|
|
16
|
+
Classifier: Programming Language :: Python :: 3
|
|
17
|
+
Classifier: Programming Language :: Python :: 3 :: Only
|
|
18
|
+
Classifier: Topic :: Scientific/Engineering
|
|
19
|
+
Requires-Python: >=3.10
|
|
20
|
+
Description-Content-Type: text/markdown
|
|
21
|
+
License-File: LICENSE
|
|
22
|
+
Provides-Extra: dev
|
|
23
|
+
Requires-Dist: pytest>=7; extra == "dev"
|
|
24
|
+
Dynamic: license-file
|
|
25
|
+
|
|
26
|
+
# litsurvey
|
|
27
|
+
|
|
28
|
+
A research assistant for the literature: ask a question, get a cited
|
|
29
|
+
state-of-the-art survey; state a claim, get a prior-art verdict; give a
|
|
30
|
+
paper title, get everything that cites it. It runs on your laptop, from the
|
|
31
|
+
command line or a small browser page, and uses the LLM you already have,
|
|
32
|
+
whether that is a Claude, ChatGPT or Gemini subscription, a local model, or
|
|
33
|
+
an API key. One Python package, no runtime dependencies, Python 3.10 or
|
|
34
|
+
newer.
|
|
35
|
+
|
|
36
|
+
## What it does
|
|
37
|
+
|
|
38
|
+
**LLM-driven (the main draw)**
|
|
39
|
+
|
|
40
|
+
- `research "question"`: the model breaks the question into sub-questions,
|
|
41
|
+
runs searches across six scholarly indexes, walks the citation graph of
|
|
42
|
+
the central papers, reads one or two arXiv papers in full, and writes a
|
|
43
|
+
survey organised by theme with numbered citations, open problems, and an
|
|
44
|
+
honest coverage caveat. Every citation comes from a real search result,
|
|
45
|
+
and the report ends with a log of every query that was run.
|
|
46
|
+
- `novelty "claim"`: the same machinery aimed at one claim: closest prior
|
|
47
|
+
work, what is new versus known, a verdict (clearly novel / incremental /
|
|
48
|
+
substantially anticipated / cannot determine) with confidence, and
|
|
49
|
+
citations a reviewer can use.
|
|
50
|
+
- Runs on any of: **your subscription's command-line agent** (Claude Code,
|
|
51
|
+
Codex CLI, Gemini CLI; no API key), **a local model** through Ollama
|
|
52
|
+
(nothing leaves the machine), or **a cloud API key** (OpenAI, OpenRouter,
|
|
53
|
+
Anthropic, any OpenAI-compatible server).
|
|
54
|
+
|
|
55
|
+
**Search and lookup (no LLM)**
|
|
56
|
+
|
|
57
|
+
- Fused search over OpenAlex, Semantic Scholar, arXiv, TechRxiv, Research
|
|
58
|
+
Square and the IACR ePrint archive, de-duplicated, sortable by relevance,
|
|
59
|
+
citations or year.
|
|
60
|
+
- For any paper, by title or id: who cites it (most-cited first), what it
|
|
61
|
+
cites, similar papers from a recommender, its full record and BibTeX, and
|
|
62
|
+
a legal open-access copy.
|
|
63
|
+
- Export to BibTeX, RIS, CSV, JSON or Markdown for Zotero, Overleaf, Mendeley
|
|
64
|
+
or a spreadsheet.
|
|
65
|
+
- A history of every completed run, shared by the command line and the
|
|
66
|
+
browser page, with the same downloads.
|
|
67
|
+
|
|
68
|
+
## Install
|
|
69
|
+
|
|
70
|
+
From GitHub (works today):
|
|
71
|
+
|
|
72
|
+
```bash
|
|
73
|
+
pipx install git+https://github.com/udaykdk/litsurvey
|
|
74
|
+
# or
|
|
75
|
+
pip install git+https://github.com/udaykdk/litsurvey
|
|
76
|
+
```
|
|
77
|
+
|
|
78
|
+
From PyPI, after the first release is published:
|
|
79
|
+
|
|
80
|
+
```bash
|
|
81
|
+
pipx install litsurvey
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
Or clone and run without installing: `python3 -m litsurvey --help`.
|
|
85
|
+
|
|
86
|
+
Then, optional but recommended: get a free Semantic Scholar API key
|
|
87
|
+
(a one-minute form, see [docs/api-keys.md](https://github.com/udaykdk/litsurvey/blob/main/docs/api-keys.md))
|
|
88
|
+
and store it with `litsurvey init`. Without it everything works, only
|
|
89
|
+
slower. `litsurvey doctor` checks the setup and prints a `RESULT:` verdict.
|
|
90
|
+
|
|
91
|
+
## Examples
|
|
92
|
+
|
|
93
|
+
A state-of-the-art survey, written by the model from live search results:
|
|
94
|
+
|
|
95
|
+
```bash
|
|
96
|
+
litsurvey research "uncertainty quantification methods for physics-informed neural networks" --out uq-survey.md
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
A prior-art check on a claim, phrased as a generic topic (never a sentence
|
|
100
|
+
from a confidential manuscript):
|
|
101
|
+
|
|
102
|
+
```bash
|
|
103
|
+
litsurvey novelty "adaptive sampling of collocation points in physics-informed neural networks" --out claim.md
|
|
104
|
+
```
|
|
105
|
+
|
|
106
|
+
Choosing where the model runs:
|
|
107
|
+
|
|
108
|
+
```bash
|
|
109
|
+
litsurvey research "..." --backend cli --model claude # Claude Code, under your Claude Pro/Max plan
|
|
110
|
+
litsurvey research "..." --backend ollama --model qwen3:30b # local; nothing leaves the machine
|
|
111
|
+
litsurvey research "..." --backend openai --base-url https://openrouter.ai/api --model anthropic/claude-sonnet-4.5
|
|
112
|
+
```
|
|
113
|
+
|
|
114
|
+
Search and lookup, with titles instead of ids:
|
|
115
|
+
|
|
116
|
+
```bash
|
|
117
|
+
litsurvey search "physics informed neural networks inverse problems" --year-from 2022 --sort citations
|
|
118
|
+
litsurvey cites "Superior thermal conductivity of single-layer graphene" # picks the paper, lists the most-cited citers
|
|
119
|
+
litsurvey related "Embedding deep learning in inverse scattering problems"
|
|
120
|
+
litsurvey paper "KAN: Kolmogorov-Arnold Networks" --out ref.bib # BibTeX for a paper you know by title
|
|
121
|
+
litsurvey oa "10.1016/j.cma.2022.114823" # legal free PDF
|
|
122
|
+
litsurvey web # the same, in your browser
|
|
123
|
+
```
|
|
124
|
+
|
|
125
|
+
When a title matches several papers, litsurvey shows the candidates and
|
|
126
|
+
asks; scripts and agents get the list and use `--pick N`.
|
|
127
|
+
|
|
128
|
+
Typical times: a search takes a few seconds; `novelty` and `research` take
|
|
129
|
+
two to fifteen minutes depending on the model.
|
|
130
|
+
|
|
131
|
+
## What needs an LLM, and what leaves your machine
|
|
132
|
+
|
|
133
|
+
| Command | LLM | What is sent out, and to whom |
|
|
134
|
+
|---|---|---|
|
|
135
|
+
| `search`, `paper`, `cites`, `refs`, `related` | no | your query, or a paper id or title, to the scholarly APIs (OpenAlex, Semantic Scholar, arXiv, Crossref, eprint.iacr.org) |
|
|
136
|
+
| `oa` | no | one DOI to Unpaywall (a title is first resolved through the search APIs) |
|
|
137
|
+
| `history`, `init`, `doctor`, `web` | no | nothing (`doctor` makes one test query per source; actions on the web page follow the rows above) |
|
|
138
|
+
| `novelty`, `research`, local model (Ollama, or an OpenAI-compatible server on localhost) | yes, local | the keyword queries the model composes and the paper ids it looks up, to the scholarly APIs; your text stays on the machine |
|
|
139
|
+
| `novelty`, `research`, subscription CLI (Claude Code, Codex, Gemini) | yes, vendor cloud | your claim or question and everything the agent reads, to that vendor under your subscription |
|
|
140
|
+
| `novelty`, `research`, cloud API key | yes, cloud | your claim or question and every search result, to that provider |
|
|
141
|
+
|
|
142
|
+
For confidential work (a manuscript under review, an unpublished result)
|
|
143
|
+
use a local model and phrase the input as generic topic terms. Completed
|
|
144
|
+
runs, including inputs and results, are stored under `~/.litsurvey/` on
|
|
145
|
+
your disk. LLM reports are aids, not verified reviews: check the cited
|
|
146
|
+
paper and DOI before using any claim. Details in
|
|
147
|
+
[docs/confidentiality.md](https://github.com/udaykdk/litsurvey/blob/main/docs/confidentiality.md).
|
|
148
|
+
|
|
149
|
+
## Documentation
|
|
150
|
+
|
|
151
|
+
- [docs/cli.md](https://github.com/udaykdk/litsurvey/blob/main/docs/cli.md): every command, option and output format, with sample output
|
|
152
|
+
- [docs/web.md](https://github.com/udaykdk/litsurvey/blob/main/docs/web.md): the browser interface and the History tab
|
|
153
|
+
- [docs/use-cases.md](https://github.com/udaykdk/litsurvey/blob/main/docs/use-cases.md): seven scenarios with the exact commands, from student discovery to confidential peer review
|
|
154
|
+
- [docs/llm-integration.md](https://github.com/udaykdk/litsurvey/blob/main/docs/llm-integration.md): LLM backends and models; using litsurvey as a tool from Claude Code, Codex and other agents
|
|
155
|
+
- [docs/confidentiality.md](https://github.com/udaykdk/litsurvey/blob/main/docs/confidentiality.md): what leaves the machine, per mode and backend
|
|
156
|
+
- [docs/api-keys.md](https://github.com/udaykdk/litsurvey/blob/main/docs/api-keys.md): getting and storing the Semantic Scholar key; running without one
|
|
157
|
+
- [CONTRIBUTING.md](https://github.com/udaykdk/litsurvey/blob/main/CONTRIBUTING.md): reporting bugs, running tests, adding a source
|
|
158
|
+
|
|
159
|
+
## Coverage, and Google Scholar
|
|
160
|
+
|
|
161
|
+
Six sources are searched by default: OpenAlex (about 250 million works),
|
|
162
|
+
Semantic Scholar (about 220 million), arXiv, and three preprint portals:
|
|
163
|
+
TechRxiv and Research Square (through Crossref, by their DOI prefixes) and
|
|
164
|
+
the IACR Cryptology ePrint Archive (which has no search API, so litsurvey
|
|
165
|
+
reads its search page; if that page changes, only that source goes quiet).
|
|
166
|
+
`--sources` restricts the set; `crossref` (all DOI-registered works) can be
|
|
167
|
+
added but overlaps OpenAlex. Together they cover most of what Google
|
|
168
|
+
Scholar shows, except some grey literature (technical reports, theses,
|
|
169
|
+
standards). Google Scholar has no API and its terms of service forbid
|
|
170
|
+
automated access, so litsurvey does not scrape it; `litsurvey search
|
|
171
|
+
--scholar` prints the matching Google Scholar URL for a manual comparison,
|
|
172
|
+
and the web page has the same link.
|
|
173
|
+
|
|
174
|
+
## Data attribution
|
|
175
|
+
|
|
176
|
+
Results come from [OpenAlex](https://openalex.org) (CC0),
|
|
177
|
+
[Semantic Scholar](https://www.semanticscholar.org) (Semantic Scholar Open
|
|
178
|
+
Data Platform, Allen Institute for AI), [arXiv](https://arxiv.org) (thank you
|
|
179
|
+
to arXiv for use of its open access interoperability),
|
|
180
|
+
[Crossref](https://www.crossref.org), the
|
|
181
|
+
[IACR Cryptology ePrint Archive](https://eprint.iacr.org) and
|
|
182
|
+
[Unpaywall](https://unpaywall.org). If you publish work that used these
|
|
183
|
+
results, please credit them.
|
|
184
|
+
|
|
185
|
+
## Author
|
|
186
|
+
|
|
187
|
+
Designed by Uday Khankhoje, Department of Electrical Engineering, IIT
|
|
188
|
+
Madras; implemented with Anthropic Claude. Questions and bug reports: the
|
|
189
|
+
[issue tracker](https://github.com/udaykdk/litsurvey/issues).
|
|
190
|
+
|
|
191
|
+
## Credits
|
|
192
|
+
|
|
193
|
+
The banner of the web page uses slivers of an aerial beach photograph by
|
|
194
|
+
[Lance Asper on Unsplash](https://unsplash.com/@lance_asper).
|
|
195
|
+
|
|
196
|
+
## Citing
|
|
197
|
+
|
|
198
|
+
See [CITATION.cff](https://github.com/udaykdk/litsurvey/blob/main/CITATION.cff).
|
|
199
|
+
GitHub shows a "Cite this repository" button on the project page.
|
|
200
|
+
|
|
201
|
+
## License
|
|
202
|
+
|
|
203
|
+
MIT. See [LICENSE](https://github.com/udaykdk/litsurvey/blob/main/LICENSE).
|
|
@@ -0,0 +1,178 @@
|
|
|
1
|
+
# litsurvey
|
|
2
|
+
|
|
3
|
+
A research assistant for the literature: ask a question, get a cited
|
|
4
|
+
state-of-the-art survey; state a claim, get a prior-art verdict; give a
|
|
5
|
+
paper title, get everything that cites it. It runs on your laptop, from the
|
|
6
|
+
command line or a small browser page, and uses the LLM you already have,
|
|
7
|
+
whether that is a Claude, ChatGPT or Gemini subscription, a local model, or
|
|
8
|
+
an API key. One Python package, no runtime dependencies, Python 3.10 or
|
|
9
|
+
newer.
|
|
10
|
+
|
|
11
|
+
## What it does
|
|
12
|
+
|
|
13
|
+
**LLM-driven (the main draw)**
|
|
14
|
+
|
|
15
|
+
- `research "question"`: the model breaks the question into sub-questions,
|
|
16
|
+
runs searches across six scholarly indexes, walks the citation graph of
|
|
17
|
+
the central papers, reads one or two arXiv papers in full, and writes a
|
|
18
|
+
survey organised by theme with numbered citations, open problems, and an
|
|
19
|
+
honest coverage caveat. Every citation comes from a real search result,
|
|
20
|
+
and the report ends with a log of every query that was run.
|
|
21
|
+
- `novelty "claim"`: the same machinery aimed at one claim: closest prior
|
|
22
|
+
work, what is new versus known, a verdict (clearly novel / incremental /
|
|
23
|
+
substantially anticipated / cannot determine) with confidence, and
|
|
24
|
+
citations a reviewer can use.
|
|
25
|
+
- Runs on any of: **your subscription's command-line agent** (Claude Code,
|
|
26
|
+
Codex CLI, Gemini CLI; no API key), **a local model** through Ollama
|
|
27
|
+
(nothing leaves the machine), or **a cloud API key** (OpenAI, OpenRouter,
|
|
28
|
+
Anthropic, any OpenAI-compatible server).
|
|
29
|
+
|
|
30
|
+
**Search and lookup (no LLM)**
|
|
31
|
+
|
|
32
|
+
- Fused search over OpenAlex, Semantic Scholar, arXiv, TechRxiv, Research
|
|
33
|
+
Square and the IACR ePrint archive, de-duplicated, sortable by relevance,
|
|
34
|
+
citations or year.
|
|
35
|
+
- For any paper, by title or id: who cites it (most-cited first), what it
|
|
36
|
+
cites, similar papers from a recommender, its full record and BibTeX, and
|
|
37
|
+
a legal open-access copy.
|
|
38
|
+
- Export to BibTeX, RIS, CSV, JSON or Markdown for Zotero, Overleaf, Mendeley
|
|
39
|
+
or a spreadsheet.
|
|
40
|
+
- A history of every completed run, shared by the command line and the
|
|
41
|
+
browser page, with the same downloads.
|
|
42
|
+
|
|
43
|
+
## Install
|
|
44
|
+
|
|
45
|
+
From GitHub (works today):
|
|
46
|
+
|
|
47
|
+
```bash
|
|
48
|
+
pipx install git+https://github.com/udaykdk/litsurvey
|
|
49
|
+
# or
|
|
50
|
+
pip install git+https://github.com/udaykdk/litsurvey
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
From PyPI, after the first release is published:
|
|
54
|
+
|
|
55
|
+
```bash
|
|
56
|
+
pipx install litsurvey
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
Or clone and run without installing: `python3 -m litsurvey --help`.
|
|
60
|
+
|
|
61
|
+
Then, optional but recommended: get a free Semantic Scholar API key
|
|
62
|
+
(a one-minute form, see [docs/api-keys.md](https://github.com/udaykdk/litsurvey/blob/main/docs/api-keys.md))
|
|
63
|
+
and store it with `litsurvey init`. Without it everything works, only
|
|
64
|
+
slower. `litsurvey doctor` checks the setup and prints a `RESULT:` verdict.
|
|
65
|
+
|
|
66
|
+
## Examples
|
|
67
|
+
|
|
68
|
+
A state-of-the-art survey, written by the model from live search results:
|
|
69
|
+
|
|
70
|
+
```bash
|
|
71
|
+
litsurvey research "uncertainty quantification methods for physics-informed neural networks" --out uq-survey.md
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
A prior-art check on a claim, phrased as a generic topic (never a sentence
|
|
75
|
+
from a confidential manuscript):
|
|
76
|
+
|
|
77
|
+
```bash
|
|
78
|
+
litsurvey novelty "adaptive sampling of collocation points in physics-informed neural networks" --out claim.md
|
|
79
|
+
```
|
|
80
|
+
|
|
81
|
+
Choosing where the model runs:
|
|
82
|
+
|
|
83
|
+
```bash
|
|
84
|
+
litsurvey research "..." --backend cli --model claude # Claude Code, under your Claude Pro/Max plan
|
|
85
|
+
litsurvey research "..." --backend ollama --model qwen3:30b # local; nothing leaves the machine
|
|
86
|
+
litsurvey research "..." --backend openai --base-url https://openrouter.ai/api --model anthropic/claude-sonnet-4.5
|
|
87
|
+
```
|
|
88
|
+
|
|
89
|
+
Search and lookup, with titles instead of ids:
|
|
90
|
+
|
|
91
|
+
```bash
|
|
92
|
+
litsurvey search "physics informed neural networks inverse problems" --year-from 2022 --sort citations
|
|
93
|
+
litsurvey cites "Superior thermal conductivity of single-layer graphene" # picks the paper, lists the most-cited citers
|
|
94
|
+
litsurvey related "Embedding deep learning in inverse scattering problems"
|
|
95
|
+
litsurvey paper "KAN: Kolmogorov-Arnold Networks" --out ref.bib # BibTeX for a paper you know by title
|
|
96
|
+
litsurvey oa "10.1016/j.cma.2022.114823" # legal free PDF
|
|
97
|
+
litsurvey web # the same, in your browser
|
|
98
|
+
```
|
|
99
|
+
|
|
100
|
+
When a title matches several papers, litsurvey shows the candidates and
|
|
101
|
+
asks; scripts and agents get the list and use `--pick N`.
|
|
102
|
+
|
|
103
|
+
Typical times: a search takes a few seconds; `novelty` and `research` take
|
|
104
|
+
two to fifteen minutes depending on the model.
|
|
105
|
+
|
|
106
|
+
## What needs an LLM, and what leaves your machine
|
|
107
|
+
|
|
108
|
+
| Command | LLM | What is sent out, and to whom |
|
|
109
|
+
|---|---|---|
|
|
110
|
+
| `search`, `paper`, `cites`, `refs`, `related` | no | your query, or a paper id or title, to the scholarly APIs (OpenAlex, Semantic Scholar, arXiv, Crossref, eprint.iacr.org) |
|
|
111
|
+
| `oa` | no | one DOI to Unpaywall (a title is first resolved through the search APIs) |
|
|
112
|
+
| `history`, `init`, `doctor`, `web` | no | nothing (`doctor` makes one test query per source; actions on the web page follow the rows above) |
|
|
113
|
+
| `novelty`, `research`, local model (Ollama, or an OpenAI-compatible server on localhost) | yes, local | the keyword queries the model composes and the paper ids it looks up, to the scholarly APIs; your text stays on the machine |
|
|
114
|
+
| `novelty`, `research`, subscription CLI (Claude Code, Codex, Gemini) | yes, vendor cloud | your claim or question and everything the agent reads, to that vendor under your subscription |
|
|
115
|
+
| `novelty`, `research`, cloud API key | yes, cloud | your claim or question and every search result, to that provider |
|
|
116
|
+
|
|
117
|
+
For confidential work (a manuscript under review, an unpublished result)
|
|
118
|
+
use a local model and phrase the input as generic topic terms. Completed
|
|
119
|
+
runs, including inputs and results, are stored under `~/.litsurvey/` on
|
|
120
|
+
your disk. LLM reports are aids, not verified reviews: check the cited
|
|
121
|
+
paper and DOI before using any claim. Details in
|
|
122
|
+
[docs/confidentiality.md](https://github.com/udaykdk/litsurvey/blob/main/docs/confidentiality.md).
|
|
123
|
+
|
|
124
|
+
## Documentation
|
|
125
|
+
|
|
126
|
+
- [docs/cli.md](https://github.com/udaykdk/litsurvey/blob/main/docs/cli.md): every command, option and output format, with sample output
|
|
127
|
+
- [docs/web.md](https://github.com/udaykdk/litsurvey/blob/main/docs/web.md): the browser interface and the History tab
|
|
128
|
+
- [docs/use-cases.md](https://github.com/udaykdk/litsurvey/blob/main/docs/use-cases.md): seven scenarios with the exact commands, from student discovery to confidential peer review
|
|
129
|
+
- [docs/llm-integration.md](https://github.com/udaykdk/litsurvey/blob/main/docs/llm-integration.md): LLM backends and models; using litsurvey as a tool from Claude Code, Codex and other agents
|
|
130
|
+
- [docs/confidentiality.md](https://github.com/udaykdk/litsurvey/blob/main/docs/confidentiality.md): what leaves the machine, per mode and backend
|
|
131
|
+
- [docs/api-keys.md](https://github.com/udaykdk/litsurvey/blob/main/docs/api-keys.md): getting and storing the Semantic Scholar key; running without one
|
|
132
|
+
- [CONTRIBUTING.md](https://github.com/udaykdk/litsurvey/blob/main/CONTRIBUTING.md): reporting bugs, running tests, adding a source
|
|
133
|
+
|
|
134
|
+
## Coverage, and Google Scholar
|
|
135
|
+
|
|
136
|
+
Six sources are searched by default: OpenAlex (about 250 million works),
|
|
137
|
+
Semantic Scholar (about 220 million), arXiv, and three preprint portals:
|
|
138
|
+
TechRxiv and Research Square (through Crossref, by their DOI prefixes) and
|
|
139
|
+
the IACR Cryptology ePrint Archive (which has no search API, so litsurvey
|
|
140
|
+
reads its search page; if that page changes, only that source goes quiet).
|
|
141
|
+
`--sources` restricts the set; `crossref` (all DOI-registered works) can be
|
|
142
|
+
added but overlaps OpenAlex. Together they cover most of what Google
|
|
143
|
+
Scholar shows, except some grey literature (technical reports, theses,
|
|
144
|
+
standards). Google Scholar has no API and its terms of service forbid
|
|
145
|
+
automated access, so litsurvey does not scrape it; `litsurvey search
|
|
146
|
+
--scholar` prints the matching Google Scholar URL for a manual comparison,
|
|
147
|
+
and the web page has the same link.
|
|
148
|
+
|
|
149
|
+
## Data attribution
|
|
150
|
+
|
|
151
|
+
Results come from [OpenAlex](https://openalex.org) (CC0),
|
|
152
|
+
[Semantic Scholar](https://www.semanticscholar.org) (Semantic Scholar Open
|
|
153
|
+
Data Platform, Allen Institute for AI), [arXiv](https://arxiv.org) (thank you
|
|
154
|
+
to arXiv for use of its open access interoperability),
|
|
155
|
+
[Crossref](https://www.crossref.org), the
|
|
156
|
+
[IACR Cryptology ePrint Archive](https://eprint.iacr.org) and
|
|
157
|
+
[Unpaywall](https://unpaywall.org). If you publish work that used these
|
|
158
|
+
results, please credit them.
|
|
159
|
+
|
|
160
|
+
## Author
|
|
161
|
+
|
|
162
|
+
Designed by Uday Khankhoje, Department of Electrical Engineering, IIT
|
|
163
|
+
Madras; implemented with Anthropic Claude. Questions and bug reports: the
|
|
164
|
+
[issue tracker](https://github.com/udaykdk/litsurvey/issues).
|
|
165
|
+
|
|
166
|
+
## Credits
|
|
167
|
+
|
|
168
|
+
The banner of the web page uses slivers of an aerial beach photograph by
|
|
169
|
+
[Lance Asper on Unsplash](https://unsplash.com/@lance_asper).
|
|
170
|
+
|
|
171
|
+
## Citing
|
|
172
|
+
|
|
173
|
+
See [CITATION.cff](https://github.com/udaykdk/litsurvey/blob/main/CITATION.cff).
|
|
174
|
+
GitHub shows a "Cite this repository" button on the project page.
|
|
175
|
+
|
|
176
|
+
## License
|
|
177
|
+
|
|
178
|
+
MIT. See [LICENSE](https://github.com/udaykdk/litsurvey/blob/main/LICENSE).
|
|
@@ -0,0 +1,71 @@
|
|
|
1
|
+
# API keys and running without them
|
|
2
|
+
|
|
3
|
+
litsurvey uses six public services. None requires payment. Only one offers
|
|
4
|
+
a key, and it is optional.
|
|
5
|
+
|
|
6
|
+
| Service | Used for | Needs | Without it |
|
|
7
|
+
|---|---|---|---|
|
|
8
|
+
| OpenAlex | search, citation ranking, fallbacks | nothing; an email address gets the faster "polite pool" | works, slightly slower |
|
|
9
|
+
| Semantic Scholar | search, citation graph, recommendations | nothing; a free API key gives a dedicated 1 request/second | works on a shared pool; frequent rate-limit retries, `related` most affected |
|
|
10
|
+
| arXiv | search, full text for the agent | nothing | n/a |
|
|
11
|
+
| Crossref | TechRxiv and Research Square search | nothing; the same email gets its polite pool | works, slightly slower |
|
|
12
|
+
| IACR ePrint | cryptography preprint search | nothing | n/a |
|
|
13
|
+
| Unpaywall | open-access lookup | an email address | works with a placeholder address |
|
|
14
|
+
|
|
15
|
+
## Getting a Semantic Scholar key
|
|
16
|
+
|
|
17
|
+
1. Open https://www.semanticscholar.org/product/api and find the API key
|
|
18
|
+
request form.
|
|
19
|
+
2. Fill in your name, email, and a one-line purpose such as "literature
|
|
20
|
+
search for academic research". Institutional email helps.
|
|
21
|
+
3. Approval arrives by email, usually within a day or two. The key looks like
|
|
22
|
+
`s2k-…` and the email states the limit: 1 request per second, cumulative
|
|
23
|
+
across all endpoints.
|
|
24
|
+
|
|
25
|
+
litsurvey enforces that limit itself, across the CLI and the web page at the
|
|
26
|
+
same time, so you will not be blocked for exceeding it.
|
|
27
|
+
|
|
28
|
+
## Storing the key
|
|
29
|
+
|
|
30
|
+
Interactive, recommended:
|
|
31
|
+
|
|
32
|
+
```bash
|
|
33
|
+
litsurvey init
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
This writes `~/.litsurvey/config.json` with file permissions 600 and also
|
|
37
|
+
asks for your email and the default LLM backend. Or write the file yourself:
|
|
38
|
+
|
|
39
|
+
```json
|
|
40
|
+
{
|
|
41
|
+
"s2_api_key": "s2k-…",
|
|
42
|
+
"openalex_mailto": "you@university.edu"
|
|
43
|
+
}
|
|
44
|
+
```
|
|
45
|
+
|
|
46
|
+
Or use environment variables, which override the file: `S2_API_KEY`,
|
|
47
|
+
`OPENALEX_MAILTO`.
|
|
48
|
+
|
|
49
|
+
Check with `litsurvey doctor`; look for the line
|
|
50
|
+
`S2 API key : set (s2k-…)`.
|
|
51
|
+
|
|
52
|
+
Never paste the key into an issue report; `--debug` output does not print it.
|
|
53
|
+
|
|
54
|
+
## What the key changes
|
|
55
|
+
|
|
56
|
+
With a key, searches return in a few seconds and the citation-graph and
|
|
57
|
+
recommendation commands are reliable. Without a key, Semantic Scholar's
|
|
58
|
+
shared pool often answers 429; litsurvey waits and retries (2, 4, 8, 16
|
|
59
|
+
seconds), so a search may take half a minute at busy times, and a long agent
|
|
60
|
+
run will be slow. OpenAlex and arXiv are unaffected, so `search` still
|
|
61
|
+
returns results even when Semantic Scholar is refusing.
|
|
62
|
+
|
|
63
|
+
## Other keys
|
|
64
|
+
|
|
65
|
+
- `OPENAI_API_KEY` and `OPENAI_BASE_URL`: for the `openai` backend. Local
|
|
66
|
+
servers such as LM Studio need no key; set the base URL to
|
|
67
|
+
`http://localhost:1234`.
|
|
68
|
+
- `ANTHROPIC_API_KEY`: for the `anthropic` backend.
|
|
69
|
+
|
|
70
|
+
These are only used by `novelty` and `research`. See
|
|
71
|
+
[llm-integration.md](llm-integration.md).
|