litsurvey 1.0.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (48) hide show
  1. litsurvey-1.0.0/CHANGELOG.md +37 -0
  2. litsurvey-1.0.0/CITATION.cff +38 -0
  3. litsurvey-1.0.0/CONTRIBUTING.md +63 -0
  4. litsurvey-1.0.0/LICENSE +21 -0
  5. litsurvey-1.0.0/MANIFEST.in +4 -0
  6. litsurvey-1.0.0/PKG-INFO +203 -0
  7. litsurvey-1.0.0/README.md +178 -0
  8. litsurvey-1.0.0/docs/api-keys.md +71 -0
  9. litsurvey-1.0.0/docs/cli.md +228 -0
  10. litsurvey-1.0.0/docs/confidentiality.md +95 -0
  11. litsurvey-1.0.0/docs/llm-integration.md +208 -0
  12. litsurvey-1.0.0/docs/use-cases.md +158 -0
  13. litsurvey-1.0.0/docs/web.md +69 -0
  14. litsurvey-1.0.0/integrations/claude-code/litsurvey/SKILL.md +49 -0
  15. litsurvey-1.0.0/integrations/codex/AGENTS.md +22 -0
  16. litsurvey-1.0.0/litsurvey/__init__.py +4 -0
  17. litsurvey-1.0.0/litsurvey/__main__.py +4 -0
  18. litsurvey-1.0.0/litsurvey/agent.py +223 -0
  19. litsurvey-1.0.0/litsurvey/backends.py +291 -0
  20. litsurvey-1.0.0/litsurvey/cli.py +345 -0
  21. litsurvey-1.0.0/litsurvey/config.py +99 -0
  22. litsurvey-1.0.0/litsurvey/export.py +132 -0
  23. litsurvey-1.0.0/litsurvey/history.py +60 -0
  24. litsurvey-1.0.0/litsurvey/http.py +160 -0
  25. litsurvey-1.0.0/litsurvey/ops.py +217 -0
  26. litsurvey-1.0.0/litsurvey/papers.py +140 -0
  27. litsurvey-1.0.0/litsurvey/sources/__init__.py +41 -0
  28. litsurvey-1.0.0/litsurvey/sources/arxiv.py +57 -0
  29. litsurvey-1.0.0/litsurvey/sources/crossref.py +55 -0
  30. litsurvey-1.0.0/litsurvey/sources/iacr.py +45 -0
  31. litsurvey-1.0.0/litsurvey/sources/openalex.py +108 -0
  32. litsurvey-1.0.0/litsurvey/sources/semanticscholar.py +65 -0
  33. litsurvey-1.0.0/litsurvey/sources/unpaywall.py +20 -0
  34. litsurvey-1.0.0/litsurvey/web/index.html +387 -0
  35. litsurvey-1.0.0/litsurvey/web.py +181 -0
  36. litsurvey-1.0.0/litsurvey.egg-info/PKG-INFO +203 -0
  37. litsurvey-1.0.0/litsurvey.egg-info/SOURCES.txt +46 -0
  38. litsurvey-1.0.0/litsurvey.egg-info/dependency_links.txt +1 -0
  39. litsurvey-1.0.0/litsurvey.egg-info/entry_points.txt +2 -0
  40. litsurvey-1.0.0/litsurvey.egg-info/requires.txt +3 -0
  41. litsurvey-1.0.0/litsurvey.egg-info/top_level.txt +1 -0
  42. litsurvey-1.0.0/pyproject.toml +44 -0
  43. litsurvey-1.0.0/setup.cfg +4 -0
  44. litsurvey-1.0.0/tests/test_backends_agent.py +186 -0
  45. litsurvey-1.0.0/tests/test_cli_history.py +150 -0
  46. litsurvey-1.0.0/tests/test_export.py +49 -0
  47. litsurvey-1.0.0/tests/test_papers.py +121 -0
  48. litsurvey-1.0.0/tests/test_sources.py +164 -0
@@ -0,0 +1,37 @@
1
+ # Changelog
2
+
3
+ All notable changes to this project are documented here. The format follows
4
+ [Keep a Changelog](https://keepachangelog.com/) and the project uses
5
+ [Semantic Versioning](https://semver.org/).
6
+
7
+ ## [Unreleased]
8
+
9
+ ## [1.0.0] - 2026-09-11
10
+
11
+ First public release.
12
+
13
+ ### Added
14
+ - `search`: fused, de-duplicated keyword search over OpenAlex, Semantic Scholar,
15
+ arXiv, TechRxiv, Research Square (via Crossref) and the IACR ePrint archive,
16
+ with `--sort relevance|citations|year`.
17
+ - `paper`, `cites`, `refs`, `related`: one paper and its citation neighbourhood;
18
+ a title is accepted in place of an id, with a candidate list and `--pick N`.
19
+ `cites`/`refs` rank by citation count via OpenAlex, with Semantic Scholar as
20
+ fallback; `related` and `paper` fall back to OpenAlex when Semantic Scholar
21
+ is unavailable.
22
+ - `oa`: legal open-access copies via Unpaywall.
23
+ - `novelty` and `research`: LLM-driven agents with pluggable backends and a
24
+ reproducible search log. Backends: `cli` (delegates the task to a signed-in
25
+ Claude Code, Codex CLI or Gemini CLI under the user's subscription),
26
+ `ollama` (local), `openai` (any OpenAI-compatible server, `--base-url` for
27
+ OpenRouter, LM Studio and others), `anthropic`.
28
+ - Export to BibTeX, RIS, CSV, JSON and Markdown from any result list.
29
+ - Run history (`litsurvey history`) shared by the CLI and the web page.
30
+ - `litsurvey web`: local browser interface with a History tab, a photo
31
+ banner, a three-way choice of where the LLM runs (subscription CLI, local
32
+ model, cloud API key) with a privacy banner, BibTeX/RIS/CSV downloads,
33
+ URL detection with DOI/arXiv conversion, "waiting for <source>" status
34
+ while a run is in progress, and a report-a-bug link.
35
+ - `init` and `doctor` commands.
36
+ - Per-host rate limiting (Semantic Scholar 1 req/s, arXiv 1 req/3 s) and
37
+ `Retry-After` handling.
@@ -0,0 +1,38 @@
1
+ cff-version: 1.2.0
2
+ title: "litsurvey: multi-source scientific literature search with LLM-assisted novelty assessment"
3
+ message: >-
4
+ If you use this software in your research, please cite it using the
5
+ metadata below.
6
+ type: software
7
+ authors:
8
+ - family-names: Khankhoje
9
+ given-names: Uday
10
+ email: uday@ee.iitm.ac.in
11
+ affiliation: >-
12
+ Department of Electrical Engineering, Indian Institute of Technology
13
+ Madras
14
+ orcid: 'https://orcid.org/0000-0002-9629-3922'
15
+ repository-code: 'https://github.com/udaykdk/litsurvey'
16
+ abstract: >-
17
+ A zero-dependency Python tool for literature surveys: it searches OpenAlex,
18
+ Semantic Scholar, arXiv, TechRxiv, Research Square and the IACR ePrint
19
+ archive, fuses and de-duplicates the results, walks the citation graph,
20
+ finds open-access copies, and exports BibTeX/RIS/CSV. An LLM (a local
21
+ model, the user's Claude Code / Codex / Gemini subscription CLI, or a cloud
22
+ API) drives multi-step novelty assessment and cited state-of-the-art
23
+ surveys, with a reproducible search log. A local browser interface with
24
+ run history is included. Designed so that confidential manuscript content
25
+ never has to leave the user's machine.
26
+ keywords:
27
+ - literature search
28
+ - peer review
29
+ - systematic review
30
+ - novelty assessment
31
+ - Semantic Scholar
32
+ - OpenAlex
33
+ - arXiv
34
+ - BibTeX
35
+ - large language models
36
+ license: MIT
37
+ version: 1.0.0
38
+ date-released: '2026-09-11'
@@ -0,0 +1,63 @@
1
+ # Contributing
2
+
3
+ ## Reporting a bug
4
+
5
+ Open an issue with the bug template. Include the command, the output with
6
+ `--debug`, and the output of `litsurvey --version` and `litsurvey doctor`.
7
+ Remove your API key if it appears anywhere. Most problems are API changes on
8
+ the provider side, and the exact URL from `--debug` is what makes them
9
+ fixable.
10
+
11
+ ## Development setup
12
+
13
+ ```bash
14
+ git clone https://github.com/udaykdk/litsurvey
15
+ cd litsurvey
16
+ python -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
17
+ pip install -e ".[dev]"
18
+ python -m pytest
19
+ ```
20
+
21
+ The tests are offline: every HTTP call is mocked. Please keep it that way, so
22
+ CI does not depend on API availability. If you change how a source is parsed,
23
+ add a fixture for the new response shape in `tests/test_sources.py`.
24
+
25
+ ## Design rules
26
+
27
+ - **Zero runtime dependencies.** Standard library only. This keeps
28
+ installation trivial on university machines and inside agents. If a
29
+ feature needs a dependency, make it optional and import it lazily.
30
+ - **Only public, official APIs, with one documented exception.** No
31
+ scraping of sites whose terms forbid it (Google Scholar in particular).
32
+ New sources should have a documented API and a rate limit we can honour in
33
+ `http.HOST_INTERVAL`. The IACR ePrint source reads a server-rendered
34
+ search page because the archive offers no search API and does not
35
+ prohibit it; it is rate-limited to one request per two seconds and is
36
+ isolated so that a markup change only silences that source.
37
+ - **Nothing leaves the machine that the docs do not say leaves.** Any change
38
+ to what is sent where must be reflected in `docs/confidentiality.md`.
39
+ - **Every successfully completed run is recorded** in the history store
40
+ (search runs with per-source hit counts), so results are reproducible.
41
+
42
+ ## Adding a source
43
+
44
+ 1. Create `litsurvey/sources/<name>.py` with a `search(query, limit,
45
+ year_from)` function that returns a list of `papers.make(...)` records
46
+ with `sources=["<name>"]`.
47
+ 2. Register it in `litsurvey/sources/__init__.py` (`SOURCES`).
48
+ 3. Add the host's rate limit to `http.HOST_INTERVAL`.
49
+ 4. Add a parsing test with a fixture, and a line in `docs/cli.md`.
50
+
51
+ ## Pull requests
52
+
53
+ Small, focused pull requests are easier to review. Add a line under
54
+ *Unreleased* in `CHANGELOG.md`. CI runs the tests on Linux, macOS and
55
+ Windows with Python 3.10, 3.12 and 3.13.
56
+
57
+ ## Releasing (maintainers)
58
+
59
+ 1. Bump `version` in `pyproject.toml`, `litsurvey/__init__.py` and
60
+ `CITATION.cff`; move the changelog entries under the new version.
61
+ 2. `git tag -a vX.Y.Z -m "vX.Y.Z"` and push the tag.
62
+ 3. Create a GitHub release from the tag; the publish workflow uploads to
63
+ PyPI via trusted publishing.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Uday K Khankhoje
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,4 @@
1
+ include LICENSE README.md CHANGELOG.md CONTRIBUTING.md CITATION.cff
2
+ recursive-include docs *.md
3
+ recursive-include integrations *.md
4
+ recursive-include litsurvey/web *.html
@@ -0,0 +1,203 @@
1
+ Metadata-Version: 2.4
2
+ Name: litsurvey
3
+ Version: 1.0.0
4
+ Summary: Multi-source scientific literature search with optional LLM-driven novelty assessment and deep research. Zero dependencies.
5
+ Author-email: Uday Khankhoje <uday@ee.iitm.ac.in>
6
+ License: MIT
7
+ Project-URL: Homepage, https://github.com/udaykdk/litsurvey
8
+ Project-URL: Issues, https://github.com/udaykdk/litsurvey/issues
9
+ Project-URL: Changelog, https://github.com/udaykdk/litsurvey/blob/main/CHANGELOG.md
10
+ Keywords: literature search,literature survey,semantic scholar,openalex,arxiv,crossref,peer review,novelty,systematic review,bibtex,llm,ollama,claude code
11
+ Classifier: Development Status :: 4 - Beta
12
+ Classifier: Environment :: Console
13
+ Classifier: Intended Audience :: Science/Research
14
+ Classifier: License :: OSI Approved :: MIT License
15
+ Classifier: Operating System :: OS Independent
16
+ Classifier: Programming Language :: Python :: 3
17
+ Classifier: Programming Language :: Python :: 3 :: Only
18
+ Classifier: Topic :: Scientific/Engineering
19
+ Requires-Python: >=3.10
20
+ Description-Content-Type: text/markdown
21
+ License-File: LICENSE
22
+ Provides-Extra: dev
23
+ Requires-Dist: pytest>=7; extra == "dev"
24
+ Dynamic: license-file
25
+
26
+ # litsurvey
27
+
28
+ A research assistant for the literature: ask a question, get a cited
29
+ state-of-the-art survey; state a claim, get a prior-art verdict; give a
30
+ paper title, get everything that cites it. It runs on your laptop, from the
31
+ command line or a small browser page, and uses the LLM you already have,
32
+ whether that is a Claude, ChatGPT or Gemini subscription, a local model, or
33
+ an API key. One Python package, no runtime dependencies, Python 3.10 or
34
+ newer.
35
+
36
+ ## What it does
37
+
38
+ **LLM-driven (the main draw)**
39
+
40
+ - `research "question"`: the model breaks the question into sub-questions,
41
+ runs searches across six scholarly indexes, walks the citation graph of
42
+ the central papers, reads one or two arXiv papers in full, and writes a
43
+ survey organised by theme with numbered citations, open problems, and an
44
+ honest coverage caveat. Every citation comes from a real search result,
45
+ and the report ends with a log of every query that was run.
46
+ - `novelty "claim"`: the same machinery aimed at one claim: closest prior
47
+ work, what is new versus known, a verdict (clearly novel / incremental /
48
+ substantially anticipated / cannot determine) with confidence, and
49
+ citations a reviewer can use.
50
+ - Runs on any of: **your subscription's command-line agent** (Claude Code,
51
+ Codex CLI, Gemini CLI; no API key), **a local model** through Ollama
52
+ (nothing leaves the machine), or **a cloud API key** (OpenAI, OpenRouter,
53
+ Anthropic, any OpenAI-compatible server).
54
+
55
+ **Search and lookup (no LLM)**
56
+
57
+ - Fused search over OpenAlex, Semantic Scholar, arXiv, TechRxiv, Research
58
+ Square and the IACR ePrint archive, de-duplicated, sortable by relevance,
59
+ citations or year.
60
+ - For any paper, by title or id: who cites it (most-cited first), what it
61
+ cites, similar papers from a recommender, its full record and BibTeX, and
62
+ a legal open-access copy.
63
+ - Export to BibTeX, RIS, CSV, JSON or Markdown for Zotero, Overleaf, Mendeley
64
+ or a spreadsheet.
65
+ - A history of every completed run, shared by the command line and the
66
+ browser page, with the same downloads.
67
+
68
+ ## Install
69
+
70
+ From GitHub (works today):
71
+
72
+ ```bash
73
+ pipx install git+https://github.com/udaykdk/litsurvey
74
+ # or
75
+ pip install git+https://github.com/udaykdk/litsurvey
76
+ ```
77
+
78
+ From PyPI, after the first release is published:
79
+
80
+ ```bash
81
+ pipx install litsurvey
82
+ ```
83
+
84
+ Or clone and run without installing: `python3 -m litsurvey --help`.
85
+
86
+ Then, optional but recommended: get a free Semantic Scholar API key
87
+ (a one-minute form, see [docs/api-keys.md](https://github.com/udaykdk/litsurvey/blob/main/docs/api-keys.md))
88
+ and store it with `litsurvey init`. Without it everything works, only
89
+ slower. `litsurvey doctor` checks the setup and prints a `RESULT:` verdict.
90
+
91
+ ## Examples
92
+
93
+ A state-of-the-art survey, written by the model from live search results:
94
+
95
+ ```bash
96
+ litsurvey research "uncertainty quantification methods for physics-informed neural networks" --out uq-survey.md
97
+ ```
98
+
99
+ A prior-art check on a claim, phrased as a generic topic (never a sentence
100
+ from a confidential manuscript):
101
+
102
+ ```bash
103
+ litsurvey novelty "adaptive sampling of collocation points in physics-informed neural networks" --out claim.md
104
+ ```
105
+
106
+ Choosing where the model runs:
107
+
108
+ ```bash
109
+ litsurvey research "..." --backend cli --model claude # Claude Code, under your Claude Pro/Max plan
110
+ litsurvey research "..." --backend ollama --model qwen3:30b # local; nothing leaves the machine
111
+ litsurvey research "..." --backend openai --base-url https://openrouter.ai/api --model anthropic/claude-sonnet-4.5
112
+ ```
113
+
114
+ Search and lookup, with titles instead of ids:
115
+
116
+ ```bash
117
+ litsurvey search "physics informed neural networks inverse problems" --year-from 2022 --sort citations
118
+ litsurvey cites "Superior thermal conductivity of single-layer graphene" # picks the paper, lists the most-cited citers
119
+ litsurvey related "Embedding deep learning in inverse scattering problems"
120
+ litsurvey paper "KAN: Kolmogorov-Arnold Networks" --out ref.bib # BibTeX for a paper you know by title
121
+ litsurvey oa "10.1016/j.cma.2022.114823" # legal free PDF
122
+ litsurvey web # the same, in your browser
123
+ ```
124
+
125
+ When a title matches several papers, litsurvey shows the candidates and
126
+ asks; scripts and agents get the list and use `--pick N`.
127
+
128
+ Typical times: a search takes a few seconds; `novelty` and `research` take
129
+ two to fifteen minutes depending on the model.
130
+
131
+ ## What needs an LLM, and what leaves your machine
132
+
133
+ | Command | LLM | What is sent out, and to whom |
134
+ |---|---|---|
135
+ | `search`, `paper`, `cites`, `refs`, `related` | no | your query, or a paper id or title, to the scholarly APIs (OpenAlex, Semantic Scholar, arXiv, Crossref, eprint.iacr.org) |
136
+ | `oa` | no | one DOI to Unpaywall (a title is first resolved through the search APIs) |
137
+ | `history`, `init`, `doctor`, `web` | no | nothing (`doctor` makes one test query per source; actions on the web page follow the rows above) |
138
+ | `novelty`, `research`, local model (Ollama, or an OpenAI-compatible server on localhost) | yes, local | the keyword queries the model composes and the paper ids it looks up, to the scholarly APIs; your text stays on the machine |
139
+ | `novelty`, `research`, subscription CLI (Claude Code, Codex, Gemini) | yes, vendor cloud | your claim or question and everything the agent reads, to that vendor under your subscription |
140
+ | `novelty`, `research`, cloud API key | yes, cloud | your claim or question and every search result, to that provider |
141
+
142
+ For confidential work (a manuscript under review, an unpublished result)
143
+ use a local model and phrase the input as generic topic terms. Completed
144
+ runs, including inputs and results, are stored under `~/.litsurvey/` on
145
+ your disk. LLM reports are aids, not verified reviews: check the cited
146
+ paper and DOI before using any claim. Details in
147
+ [docs/confidentiality.md](https://github.com/udaykdk/litsurvey/blob/main/docs/confidentiality.md).
148
+
149
+ ## Documentation
150
+
151
+ - [docs/cli.md](https://github.com/udaykdk/litsurvey/blob/main/docs/cli.md): every command, option and output format, with sample output
152
+ - [docs/web.md](https://github.com/udaykdk/litsurvey/blob/main/docs/web.md): the browser interface and the History tab
153
+ - [docs/use-cases.md](https://github.com/udaykdk/litsurvey/blob/main/docs/use-cases.md): seven scenarios with the exact commands, from student discovery to confidential peer review
154
+ - [docs/llm-integration.md](https://github.com/udaykdk/litsurvey/blob/main/docs/llm-integration.md): LLM backends and models; using litsurvey as a tool from Claude Code, Codex and other agents
155
+ - [docs/confidentiality.md](https://github.com/udaykdk/litsurvey/blob/main/docs/confidentiality.md): what leaves the machine, per mode and backend
156
+ - [docs/api-keys.md](https://github.com/udaykdk/litsurvey/blob/main/docs/api-keys.md): getting and storing the Semantic Scholar key; running without one
157
+ - [CONTRIBUTING.md](https://github.com/udaykdk/litsurvey/blob/main/CONTRIBUTING.md): reporting bugs, running tests, adding a source
158
+
159
+ ## Coverage, and Google Scholar
160
+
161
+ Six sources are searched by default: OpenAlex (about 250 million works),
162
+ Semantic Scholar (about 220 million), arXiv, and three preprint portals:
163
+ TechRxiv and Research Square (through Crossref, by their DOI prefixes) and
164
+ the IACR Cryptology ePrint Archive (which has no search API, so litsurvey
165
+ reads its search page; if that page changes, only that source goes quiet).
166
+ `--sources` restricts the set; `crossref` (all DOI-registered works) can be
167
+ added but overlaps OpenAlex. Together they cover most of what Google
168
+ Scholar shows, except some grey literature (technical reports, theses,
169
+ standards). Google Scholar has no API and its terms of service forbid
170
+ automated access, so litsurvey does not scrape it; `litsurvey search
171
+ --scholar` prints the matching Google Scholar URL for a manual comparison,
172
+ and the web page has the same link.
173
+
174
+ ## Data attribution
175
+
176
+ Results come from [OpenAlex](https://openalex.org) (CC0),
177
+ [Semantic Scholar](https://www.semanticscholar.org) (Semantic Scholar Open
178
+ Data Platform, Allen Institute for AI), [arXiv](https://arxiv.org) (thank you
179
+ to arXiv for use of its open access interoperability),
180
+ [Crossref](https://www.crossref.org), the
181
+ [IACR Cryptology ePrint Archive](https://eprint.iacr.org) and
182
+ [Unpaywall](https://unpaywall.org). If you publish work that used these
183
+ results, please credit them.
184
+
185
+ ## Author
186
+
187
+ Designed by Uday Khankhoje, Department of Electrical Engineering, IIT
188
+ Madras; implemented with Anthropic Claude. Questions and bug reports: the
189
+ [issue tracker](https://github.com/udaykdk/litsurvey/issues).
190
+
191
+ ## Credits
192
+
193
+ The banner of the web page uses slivers of an aerial beach photograph by
194
+ [Lance Asper on Unsplash](https://unsplash.com/@lance_asper).
195
+
196
+ ## Citing
197
+
198
+ See [CITATION.cff](https://github.com/udaykdk/litsurvey/blob/main/CITATION.cff).
199
+ GitHub shows a "Cite this repository" button on the project page.
200
+
201
+ ## License
202
+
203
+ MIT. See [LICENSE](https://github.com/udaykdk/litsurvey/blob/main/LICENSE).
@@ -0,0 +1,178 @@
1
+ # litsurvey
2
+
3
+ A research assistant for the literature: ask a question, get a cited
4
+ state-of-the-art survey; state a claim, get a prior-art verdict; give a
5
+ paper title, get everything that cites it. It runs on your laptop, from the
6
+ command line or a small browser page, and uses the LLM you already have,
7
+ whether that is a Claude, ChatGPT or Gemini subscription, a local model, or
8
+ an API key. One Python package, no runtime dependencies, Python 3.10 or
9
+ newer.
10
+
11
+ ## What it does
12
+
13
+ **LLM-driven (the main draw)**
14
+
15
+ - `research "question"`: the model breaks the question into sub-questions,
16
+ runs searches across six scholarly indexes, walks the citation graph of
17
+ the central papers, reads one or two arXiv papers in full, and writes a
18
+ survey organised by theme with numbered citations, open problems, and an
19
+ honest coverage caveat. Every citation comes from a real search result,
20
+ and the report ends with a log of every query that was run.
21
+ - `novelty "claim"`: the same machinery aimed at one claim: closest prior
22
+ work, what is new versus known, a verdict (clearly novel / incremental /
23
+ substantially anticipated / cannot determine) with confidence, and
24
+ citations a reviewer can use.
25
+ - Runs on any of: **your subscription's command-line agent** (Claude Code,
26
+ Codex CLI, Gemini CLI; no API key), **a local model** through Ollama
27
+ (nothing leaves the machine), or **a cloud API key** (OpenAI, OpenRouter,
28
+ Anthropic, any OpenAI-compatible server).
29
+
30
+ **Search and lookup (no LLM)**
31
+
32
+ - Fused search over OpenAlex, Semantic Scholar, arXiv, TechRxiv, Research
33
+ Square and the IACR ePrint archive, de-duplicated, sortable by relevance,
34
+ citations or year.
35
+ - For any paper, by title or id: who cites it (most-cited first), what it
36
+ cites, similar papers from a recommender, its full record and BibTeX, and
37
+ a legal open-access copy.
38
+ - Export to BibTeX, RIS, CSV, JSON or Markdown for Zotero, Overleaf, Mendeley
39
+ or a spreadsheet.
40
+ - A history of every completed run, shared by the command line and the
41
+ browser page, with the same downloads.
42
+
43
+ ## Install
44
+
45
+ From GitHub (works today):
46
+
47
+ ```bash
48
+ pipx install git+https://github.com/udaykdk/litsurvey
49
+ # or
50
+ pip install git+https://github.com/udaykdk/litsurvey
51
+ ```
52
+
53
+ From PyPI, after the first release is published:
54
+
55
+ ```bash
56
+ pipx install litsurvey
57
+ ```
58
+
59
+ Or clone and run without installing: `python3 -m litsurvey --help`.
60
+
61
+ Then, optional but recommended: get a free Semantic Scholar API key
62
+ (a one-minute form, see [docs/api-keys.md](https://github.com/udaykdk/litsurvey/blob/main/docs/api-keys.md))
63
+ and store it with `litsurvey init`. Without it everything works, only
64
+ slower. `litsurvey doctor` checks the setup and prints a `RESULT:` verdict.
65
+
66
+ ## Examples
67
+
68
+ A state-of-the-art survey, written by the model from live search results:
69
+
70
+ ```bash
71
+ litsurvey research "uncertainty quantification methods for physics-informed neural networks" --out uq-survey.md
72
+ ```
73
+
74
+ A prior-art check on a claim, phrased as a generic topic (never a sentence
75
+ from a confidential manuscript):
76
+
77
+ ```bash
78
+ litsurvey novelty "adaptive sampling of collocation points in physics-informed neural networks" --out claim.md
79
+ ```
80
+
81
+ Choosing where the model runs:
82
+
83
+ ```bash
84
+ litsurvey research "..." --backend cli --model claude # Claude Code, under your Claude Pro/Max plan
85
+ litsurvey research "..." --backend ollama --model qwen3:30b # local; nothing leaves the machine
86
+ litsurvey research "..." --backend openai --base-url https://openrouter.ai/api --model anthropic/claude-sonnet-4.5
87
+ ```
88
+
89
+ Search and lookup, with titles instead of ids:
90
+
91
+ ```bash
92
+ litsurvey search "physics informed neural networks inverse problems" --year-from 2022 --sort citations
93
+ litsurvey cites "Superior thermal conductivity of single-layer graphene" # picks the paper, lists the most-cited citers
94
+ litsurvey related "Embedding deep learning in inverse scattering problems"
95
+ litsurvey paper "KAN: Kolmogorov-Arnold Networks" --out ref.bib # BibTeX for a paper you know by title
96
+ litsurvey oa "10.1016/j.cma.2022.114823" # legal free PDF
97
+ litsurvey web # the same, in your browser
98
+ ```
99
+
100
+ When a title matches several papers, litsurvey shows the candidates and
101
+ asks; scripts and agents get the list and use `--pick N`.
102
+
103
+ Typical times: a search takes a few seconds; `novelty` and `research` take
104
+ two to fifteen minutes depending on the model.
105
+
106
+ ## What needs an LLM, and what leaves your machine
107
+
108
+ | Command | LLM | What is sent out, and to whom |
109
+ |---|---|---|
110
+ | `search`, `paper`, `cites`, `refs`, `related` | no | your query, or a paper id or title, to the scholarly APIs (OpenAlex, Semantic Scholar, arXiv, Crossref, eprint.iacr.org) |
111
+ | `oa` | no | one DOI to Unpaywall (a title is first resolved through the search APIs) |
112
+ | `history`, `init`, `doctor`, `web` | no | nothing (`doctor` makes one test query per source; actions on the web page follow the rows above) |
113
+ | `novelty`, `research`, local model (Ollama, or an OpenAI-compatible server on localhost) | yes, local | the keyword queries the model composes and the paper ids it looks up, to the scholarly APIs; your text stays on the machine |
114
+ | `novelty`, `research`, subscription CLI (Claude Code, Codex, Gemini) | yes, vendor cloud | your claim or question and everything the agent reads, to that vendor under your subscription |
115
+ | `novelty`, `research`, cloud API key | yes, cloud | your claim or question and every search result, to that provider |
116
+
117
+ For confidential work (a manuscript under review, an unpublished result)
118
+ use a local model and phrase the input as generic topic terms. Completed
119
+ runs, including inputs and results, are stored under `~/.litsurvey/` on
120
+ your disk. LLM reports are aids, not verified reviews: check the cited
121
+ paper and DOI before using any claim. Details in
122
+ [docs/confidentiality.md](https://github.com/udaykdk/litsurvey/blob/main/docs/confidentiality.md).
123
+
124
+ ## Documentation
125
+
126
+ - [docs/cli.md](https://github.com/udaykdk/litsurvey/blob/main/docs/cli.md): every command, option and output format, with sample output
127
+ - [docs/web.md](https://github.com/udaykdk/litsurvey/blob/main/docs/web.md): the browser interface and the History tab
128
+ - [docs/use-cases.md](https://github.com/udaykdk/litsurvey/blob/main/docs/use-cases.md): seven scenarios with the exact commands, from student discovery to confidential peer review
129
+ - [docs/llm-integration.md](https://github.com/udaykdk/litsurvey/blob/main/docs/llm-integration.md): LLM backends and models; using litsurvey as a tool from Claude Code, Codex and other agents
130
+ - [docs/confidentiality.md](https://github.com/udaykdk/litsurvey/blob/main/docs/confidentiality.md): what leaves the machine, per mode and backend
131
+ - [docs/api-keys.md](https://github.com/udaykdk/litsurvey/blob/main/docs/api-keys.md): getting and storing the Semantic Scholar key; running without one
132
+ - [CONTRIBUTING.md](https://github.com/udaykdk/litsurvey/blob/main/CONTRIBUTING.md): reporting bugs, running tests, adding a source
133
+
134
+ ## Coverage, and Google Scholar
135
+
136
+ Six sources are searched by default: OpenAlex (about 250 million works),
137
+ Semantic Scholar (about 220 million), arXiv, and three preprint portals:
138
+ TechRxiv and Research Square (through Crossref, by their DOI prefixes) and
139
+ the IACR Cryptology ePrint Archive (which has no search API, so litsurvey
140
+ reads its search page; if that page changes, only that source goes quiet).
141
+ `--sources` restricts the set; `crossref` (all DOI-registered works) can be
142
+ added but overlaps OpenAlex. Together they cover most of what Google
143
+ Scholar shows, except some grey literature (technical reports, theses,
144
+ standards). Google Scholar has no API and its terms of service forbid
145
+ automated access, so litsurvey does not scrape it; `litsurvey search
146
+ --scholar` prints the matching Google Scholar URL for a manual comparison,
147
+ and the web page has the same link.
148
+
149
+ ## Data attribution
150
+
151
+ Results come from [OpenAlex](https://openalex.org) (CC0),
152
+ [Semantic Scholar](https://www.semanticscholar.org) (Semantic Scholar Open
153
+ Data Platform, Allen Institute for AI), [arXiv](https://arxiv.org) (thank you
154
+ to arXiv for use of its open access interoperability),
155
+ [Crossref](https://www.crossref.org), the
156
+ [IACR Cryptology ePrint Archive](https://eprint.iacr.org) and
157
+ [Unpaywall](https://unpaywall.org). If you publish work that used these
158
+ results, please credit them.
159
+
160
+ ## Author
161
+
162
+ Designed by Uday Khankhoje, Department of Electrical Engineering, IIT
163
+ Madras; implemented with Anthropic Claude. Questions and bug reports: the
164
+ [issue tracker](https://github.com/udaykdk/litsurvey/issues).
165
+
166
+ ## Credits
167
+
168
+ The banner of the web page uses slivers of an aerial beach photograph by
169
+ [Lance Asper on Unsplash](https://unsplash.com/@lance_asper).
170
+
171
+ ## Citing
172
+
173
+ See [CITATION.cff](https://github.com/udaykdk/litsurvey/blob/main/CITATION.cff).
174
+ GitHub shows a "Cite this repository" button on the project page.
175
+
176
+ ## License
177
+
178
+ MIT. See [LICENSE](https://github.com/udaykdk/litsurvey/blob/main/LICENSE).
@@ -0,0 +1,71 @@
1
+ # API keys and running without them
2
+
3
+ litsurvey uses six public services. None requires payment. Only one offers
4
+ a key, and it is optional.
5
+
6
+ | Service | Used for | Needs | Without it |
7
+ |---|---|---|---|
8
+ | OpenAlex | search, citation ranking, fallbacks | nothing; an email address gets the faster "polite pool" | works, slightly slower |
9
+ | Semantic Scholar | search, citation graph, recommendations | nothing; a free API key gives a dedicated 1 request/second | works on a shared pool; frequent rate-limit retries, `related` most affected |
10
+ | arXiv | search, full text for the agent | nothing | n/a |
11
+ | Crossref | TechRxiv and Research Square search | nothing; the same email gets its polite pool | works, slightly slower |
12
+ | IACR ePrint | cryptography preprint search | nothing | n/a |
13
+ | Unpaywall | open-access lookup | an email address | works with a placeholder address |
14
+
15
+ ## Getting a Semantic Scholar key
16
+
17
+ 1. Open https://www.semanticscholar.org/product/api and find the API key
18
+ request form.
19
+ 2. Fill in your name, email, and a one-line purpose such as "literature
20
+ search for academic research". Institutional email helps.
21
+ 3. Approval arrives by email, usually within a day or two. The key looks like
22
+ `s2k-…` and the email states the limit: 1 request per second, cumulative
23
+ across all endpoints.
24
+
25
+ litsurvey enforces that limit itself, across the CLI and the web page at the
26
+ same time, so you will not be blocked for exceeding it.
27
+
28
+ ## Storing the key
29
+
30
+ Interactive, recommended:
31
+
32
+ ```bash
33
+ litsurvey init
34
+ ```
35
+
36
+ This writes `~/.litsurvey/config.json` with file permissions 600 and also
37
+ asks for your email and the default LLM backend. Or write the file yourself:
38
+
39
+ ```json
40
+ {
41
+ "s2_api_key": "s2k-…",
42
+ "openalex_mailto": "you@university.edu"
43
+ }
44
+ ```
45
+
46
+ Or use environment variables, which override the file: `S2_API_KEY`,
47
+ `OPENALEX_MAILTO`.
48
+
49
+ Check with `litsurvey doctor`; look for the line
50
+ `S2 API key : set (s2k-…)`.
51
+
52
+ Never paste the key into an issue report; `--debug` output does not print it.
53
+
54
+ ## What the key changes
55
+
56
+ With a key, searches return in a few seconds and the citation-graph and
57
+ recommendation commands are reliable. Without a key, Semantic Scholar's
58
+ shared pool often answers 429; litsurvey waits and retries (2, 4, 8, 16
59
+ seconds), so a search may take half a minute at busy times, and a long agent
60
+ run will be slow. OpenAlex and arXiv are unaffected, so `search` still
61
+ returns results even when Semantic Scholar is refusing.
62
+
63
+ ## Other keys
64
+
65
+ - `OPENAI_API_KEY` and `OPENAI_BASE_URL`: for the `openai` backend. Local
66
+ servers such as LM Studio need no key; set the base URL to
67
+ `http://localhost:1234`.
68
+ - `ANTHROPIC_API_KEY`: for the `anthropic` backend.
69
+
70
+ These are only used by `novelty` and `research`. See
71
+ [llm-integration.md](llm-integration.md).