apollodorus-client 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,2 @@
1
+ [flake8]
2
+ max-line-length = 120
@@ -0,0 +1,4 @@
1
+ .env
2
+ .env.*
3
+ dist/
4
+ .venv/
@@ -0,0 +1,48 @@
1
+ # Changelog
2
+
3
+ All notable changes to `apollodorus-client` are recorded here. The format follows
4
+ [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and versions follow
5
+ [Semantic Versioning](https://semver.org/) under the 0.x rules in
6
+ [`CLAUDE.md`](CLAUDE.md#choosing-the-version).
7
+
8
+ Every release needs its own `## [X.Y.Z] - YYYY-MM-DD` section, because the release
9
+ workflow refuses a version without one. Add changes under `## [Unreleased]` as you make
10
+ them. When you release, rename that heading to the new version, and start a fresh
11
+ `## [Unreleased]` above it.
12
+
13
+ ## [Unreleased]
14
+
15
+ ## [0.1.0] - 2026-10-08
16
+
17
+ The first release.
18
+
19
+ ### Added
20
+
21
+ - `Apollodorus.search(query, effort=...)` runs effort-level searches (`low`, `medium`,
22
+ `high`, `extra`, `max`):
23
+ - it submits a search job and long-polls it to the end;
24
+ - it resubmits after a server restart or a stalled job, up to 3 times;
25
+ - its timeout counts only running time: time spent queued doesn't count, and the
26
+ clock restarts after a resubmit;
27
+ - it takes the filters `limit`, `since`, `until`, `field`, and `top_n`, plus an
28
+ `on_progress` callback.
29
+ - `paper`, `text`, `fetch`, `bookmark` (with an optional note), `unbookmark`, `bookmarks`,
30
+ `stats`, and `health`.
31
+ - `Apollodorus.from_env()` reads `APOLLODORUS_URL`, `APOLLODORUS_API_KEY`, and
32
+ `APOLLODORUS_BOOKMARK_OWNER` from the environment. With the `[dotenv]` extra it also
33
+ reads them from a `.env` file.
34
+ - Retries:
35
+ - `search` and the read-only calls retry transport errors and 429/502/503/504 up to
36
+ 5 times, with a jittered backoff;
37
+ - a `Retry-After` header is honored, capped at 30 s.
38
+ - Typed, immutable results: `SearchResult`, a `SearchResults` list that carries the job's
39
+ `partial` flag and its `sources`, and `JobStatus`. The package ships `py.typed`.
40
+ - Errors:
41
+ - `ApollodorusError` carries `status` and `detail`, and can be pickled;
42
+ - a search that runs out of time raises `SearchTimeout`;
43
+ - a wrong `base_url` gets a clear hint;
44
+ - a key or owner that can't be sent as a header is refused, and the error text never
45
+ includes it.
46
+ - Safety:
47
+ - the client never follows a redirect, so the key never leaves the API host;
48
+ - an `http://` URL to a non-loopback host warns once.
@@ -0,0 +1,104 @@
1
+ # client/ — apollodorus-client, a public PyPI package
2
+
3
+ This is the pip-installable Python client for the Apollodorus HTTP API (`docs/api.md`):
4
+
5
+ ```bash
6
+ pip install apollodorus-client
7
+ ```
8
+
9
+ Anyone can read it on PyPI, so it must hold no secrets and no internal hostnames. Write
10
+ examples with `https://<your-host>/apollodorus/api`. Its only dependency is `httpx`;
11
+ `python-dotenv` comes with the `[dotenv]` extra.
12
+
13
+ ## Layout
14
+
15
+ - `src/apollodorus_client/__init__.py` holds the public exports (`__all__`) and
16
+ `__version__`, the single source of the version (hatch reads it, see `pyproject.toml`).
17
+ - `src/apollodorus_client/_client.py` holds all the code.
18
+ - `tests/` runs offline against `httpx.MockTransport`.
19
+ - The root suite's `tests/unit/test_client_e2e.py` drives the client against the real app,
20
+ and pins the client's deadlines to the server's.
21
+ - `README.md` is the package's PyPI page. `CHANGELOG.md` has one section per release.
22
+
23
+ Check from `client/` (CI's `client` job runs the same):
24
+
25
+ ```bash
26
+ pytest -q && flake8 src tests && black --check src tests && isort --check-only src tests && mypy
27
+ ```
28
+
29
+ ## Releasing (automatic)
30
+
31
+ `.github/workflows/publish-client.yml` runs on every push to `main` that touches
32
+ `client/`. It reads `__version__` and compares it with what PyPI already has:
33
+
34
+ - **The version is already on PyPI:** nothing is published. A test, CI, or comment change
35
+ can merge without a version bump.
36
+ - **The version is new:**
37
+ 1. The `check` job validates it: valid PEP 440, `MAJOR.MINOR.PATCH` with an optional
38
+ `aN`/`bN`/`rcN` suffix, higher than every version on PyPI, and with a
39
+ `## [X.Y.Z]` section in `CHANGELOG.md`. Any failure fails the run and names the fix.
40
+ 2. The workflow then tests, builds, checks the wheel on Python 3.10 and with
41
+ `twine check --strict`, and publishes to PyPI.
42
+ 3. Finally it tags `client-vX.Y.Z` and opens a GitHub release whose notes are that
43
+ CHANGELOG section.
44
+
45
+ So **a release is one merge**. Its pull request carries three changes together:
46
+
47
+ 1. the code change and its tests;
48
+ 2. the `CHANGELOG.md` section: move the `[Unreleased]` notes under
49
+ `## [X.Y.Z] - YYYY-MM-DD`, and leave an empty `## [Unreleased]` above it;
50
+ 3. the bumped `__version__`.
51
+
52
+ Do not create or push `client-v*` tags by hand: the workflow makes them, and nothing
53
+ triggers on them. To retry a failed run, use "Run workflow" on `main`. A re-run is safe:
54
+ files PyPI already has are skipped, and so is a tag or release that already exists.
55
+
56
+ Anything a user would notice needs a release, including a README fix, because the README
57
+ is the PyPI page and only a new version updates it. Changes outside `client/` (the server,
58
+ the UI, deploys) never need a client release, unless the client has to change with them.
59
+
60
+ ## Choosing the version
61
+
62
+ Use [SemVer](https://semver.org/) with 0.x rules: the package isn't 1.0 yet. The public
63
+ API is:
64
+
65
+ - everything in `__all__`;
66
+ - the methods of `Apollodorus`: their names, parameters, defaults, and return types;
67
+ - the fields of `SearchResult`, `SearchResults`, and `JobStatus`;
68
+ - the exception types and their `status` and `detail`;
69
+ - the environment variables `APOLLODORUS_URL`, `APOLLODORUS_API_KEY`, and
70
+ `APOLLODORUS_BOOKMARK_OWNER`;
71
+ - the retry, timeout, and resubmit behavior the README documents.
72
+
73
+ | Bump | When | Example |
74
+ |---|---|---|
75
+ | **PATCH** (0.1.0 → 0.1.1) | A fix that keeps the documented behavior: bugs, wrong error text, README or docstring fixes, internal refactors, and loosening a dependency bound | a poll that spun is now bounded |
76
+ | **MINOR** (0.1.x → 0.2.0) | Everything else, breaking changes included while 0.x: new methods, parameters, or result fields; changed defaults or retry/timeout behavior; a call to a route older servers lack; a higher minimum Python or `httpx`; anything renamed or removed | `search(..., sort=)`; a renamed exception |
77
+ | **MAJOR** (→ 1.0.0) | Only when the owner declares the API stable. Normal SemVer applies after that: breaking means MAJOR. | — |
78
+
79
+ Rules:
80
+
81
+ - **When in doubt between PATCH and MINOR, choose MINOR.**
82
+ - **List every breaking change** under `### Breaking` in that version's CHANGELOG section,
83
+ and say how to migrate.
84
+ - **Pre-releases** (`0.2.0rc1`) are for trying a risky release first. pip ignores them
85
+ unless asked (`pip install --pre apollodorus-client`). The final `0.2.0` then
86
+ supersedes them.
87
+ - **Versions only go up.** PyPI never accepts a version twice, even one deleted or
88
+ yanked, and the workflow refuses a version below the highest on PyPI. That means there
89
+ are no backport releases for an old minor line: fix forward.
90
+ - **If a bad release got out,** yank it on PyPI (it stays installable when pinned) and
91
+ release the fix as the next PATCH.
92
+
93
+ ## Compatibility with the server
94
+
95
+ The client follows the HTTP API in `docs/api.md`:
96
+
97
+ - **A new response field:** old clients ignore it, so no release is needed until the
98
+ client exposes it (MINOR).
99
+ - **A server change that would break released clients:** keep the old behavior on the
100
+ server for a while, release a client that handles both, and only then drop the old one.
101
+ - **Pinned values:** `DEADLINES`, `STALE_HEARTBEAT_SECONDS`, and `POLL_WAIT_SECONDS` must
102
+ equal the server's (`effort.DEADLINES`, `jobs.STALE_SECONDS`, the long poll's cap), and
103
+ `tests/unit/test_client_e2e.py` fails when they drift. A server change to any of them
104
+ needs a client release, usually MINOR because the timeout behavior changes.
@@ -0,0 +1,266 @@
1
+ Metadata-Version: 2.5
2
+ Name: apollodorus-client
3
+ Version: 0.1.0
4
+ Summary: Python client for the Apollodorus paper-search API
5
+ Classifier: Intended Audience :: Developers
6
+ Classifier: Operating System :: OS Independent
7
+ Classifier: Programming Language :: Python :: 3
8
+ Classifier: Programming Language :: Python :: 3 :: Only
9
+ Classifier: Typing :: Typed
10
+ Requires-Python: >=3.10
11
+ Requires-Dist: httpx>=0.27
12
+ Provides-Extra: dev
13
+ Requires-Dist: black==26.5.1; extra == 'dev'
14
+ Requires-Dist: flake8==7.3.0; extra == 'dev'
15
+ Requires-Dist: isort==8.0.1; extra == 'dev'
16
+ Requires-Dist: mypy==2.1.0; extra == 'dev'
17
+ Requires-Dist: pytest==9.0.3; extra == 'dev'
18
+ Requires-Dist: python-dotenv==1.2.2; extra == 'dev'
19
+ Provides-Extra: dotenv
20
+ Requires-Dist: python-dotenv>=1.0; extra == 'dotenv'
21
+ Description-Content-Type: text/markdown
22
+
23
+ # apollodorus-client
24
+
25
+ Python client for the Apollodorus paper-search API: FluxGate's cache of top research
26
+ papers (quant finance, CS, statistics, economics), with search that reaches beyond it.
27
+
28
+ ```bash
29
+ pip install apollodorus-client # add [dotenv] to read a .env file
30
+ ```
31
+
32
+ ```python
33
+ from apollodorus_client import Apollodorus
34
+
35
+ with Apollodorus.from_env() as ap: # APOLLODORUS_URL + APOLLODORUS_API_KEY
36
+ hits = ap.search("momentum crash", effort="medium")
37
+ for hit in hits:
38
+ print(f"{hit.score:8.3f} {hit.id} cached={hit.cached} {hit.title}")
39
+ print(hits.sources, "partial" if hits.partial else "complete")
40
+ ```
41
+
42
+ Python 3.10 or later; the only dependency is `httpx`. The server also serves a guide for AI
43
+ agents at `{APOLLODORUS_URL}/llms.txt` and its OpenAPI schema at
44
+ `{APOLLODORUS_URL}/openapi.json`.
45
+
46
+ ## Configuration
47
+
48
+ `Apollodorus.from_env(*, owner=None, timeout=60.0)` reads three variables:
49
+
50
+ | Variable | Meaning |
51
+ |---|---|
52
+ | `APOLLODORUS_URL` | the API root, e.g. `https://<your-host>/apollodorus/api` (required) |
53
+ | `APOLLODORUS_API_KEY` | the API key, sent as the `X-API-Key` header (required) |
54
+ | `APOLLODORUS_BOOKMARK_OWNER` | optional: the name your bookmarks are kept under (`owner` overrides it) |
55
+
56
+ With the `[dotenv]` extra installed, a `.env` file in the working directory, or in a
57
+ directory above it, supplies what the environment lacks; a real environment variable wins
58
+ over the file. Only those three names are read, from either. `from_env()` raises
59
+ `RuntimeError` when the URL or the key is missing.
60
+
61
+ Or build the client yourself:
62
+
63
+ ```python
64
+ import os
65
+
66
+ from apollodorus_client import Apollodorus
67
+
68
+ ap = Apollodorus(
69
+ "https://<your-host>/apollodorus/api",
70
+ os.environ["APOLLODORUS_API_KEY"],
71
+ owner="my-project", # optional: sent as X-Bookmark-Owner
72
+ timeout=60.0, # seconds per request; see "Papers, bookmarks, and the corpus" for the exceptions
73
+ )
74
+ ```
75
+
76
+ - `base_url` must be an absolute `http://` or `https://` URL with a host.
77
+ - `api_key` must be visible ASCII: no spaces, control characters, or curly quotes.
78
+ Whitespace around it, such as a trailing newline from a secrets file, is dropped. A
79
+ refused key is never repeated in the error.
80
+ - `owner` is 1 to 64 lowercase letters, digits, `.`, `_`, or `-`, starting with a letter or
81
+ digit. Without one, the server files your bookmarks under the shared owner `api-key`.
82
+ - `http_client` (optional) is an `httpx.Client` to send through. It is used as it is: never
83
+ modified, never closed, and never given the key or the base URL, which go only on this
84
+ client's own requests. Its own timeout then applies in place of `timeout`.
85
+
86
+ Keep the key in the environment or a `.env` file, never in code or URLs, and use an
87
+ `https://` URL. A plain `http://` `base_url` to any host but `localhost`, `127.0.0.0/8`, or
88
+ `::1` makes the constructor issue a `UserWarning`, because the key would travel in
89
+ cleartext.
90
+
91
+ ## Effort levels
92
+
93
+ | `effort` | Matches by | Searches | Server time budget |
94
+ |---|---|---|---|
95
+ | `low` (default) | keywords | the cache: titles, abstracts, full text | 10 s |
96
+ | `medium` | keywords | the cache + Semantic Scholar, OpenAlex, CORE, arXiv | 30 s |
97
+ | `high` | meaning | the cache + Semantic Scholar | 60 s |
98
+ | `extra` | meaning | the cache + Semantic Scholar, OpenAlex, CORE, arXiv | 120 s |
99
+ | `max` | meaning, over full text | `extra`, then reads the top `top_n` papers in full | none (minutes) |
100
+
101
+ Keyword levels need every word, in any order; `"quoted words"` must appear as a phrase.
102
+ Meaning levels find papers about the query even when they share none of its words. A level
103
+ that runs out of time answers with what it found, and `hits.partial` is `True`. Start at
104
+ `low`, and go up when the results fall short. `EFFORTS` and `DEADLINES` (seconds, `None` for
105
+ `max`) are importable from the package.
106
+
107
+ ## Searching
108
+
109
+ ```python
110
+ from apollodorus_client import Apollodorus, JobStatus
111
+
112
+
113
+ def show(status: JobStatus) -> None: # called with each poll's answer
114
+ print(f"{status.status:<8} {status.stage} {status.elapsed_s:.0f} s")
115
+
116
+
117
+ with Apollodorus.from_env() as ap:
118
+ hits = ap.search(
119
+ "regime switching asset allocation",
120
+ effort="high",
121
+ limit=50, # 1-100 results, default 20
122
+ since=2015, until=2024, # publication years
123
+ field="q-fin", # a coarse field group: q-fin, econ, cs, stat, ...
124
+ timeout=120, # seconds the search may run; None waits as long as it takes
125
+ on_progress=show,
126
+ )
127
+ print(hits.job_id, hits.partial, hits.sources, hits.elapsed_s)
128
+
129
+ deep = ap.search("intraday momentum and liquidity", effort="max", top_n=10) # minutes
130
+ for hit in deep[:10]:
131
+ print(hit.fulltext, hit.match, (hit.passage or "")[:200])
132
+ ```
133
+
134
+ `search(query, *, effort="low", limit=20, since=None, until=None, field=None, top_n=20,
135
+ timeout=..., on_progress=None)` submits the search as a job on the server and long-polls it
136
+ until it is done. It returns `SearchResults`: a list of `SearchResult`, best first, that also
137
+ carries `.partial`, `.sources` (each source's `ok`, `failed`, `skipped`, or `timeout`),
138
+ `.job_id`, and `.elapsed_s`. `top_n` (1–50) is read at `max` only: the papers whose full text
139
+ is read. Out-of-range values (a `limit` of 500, say) are refused by the server, as an
140
+ `ApollodorusError` with status 422. `on_progress` receives a `JobStatus` (`id`, `status`,
141
+ `stage`, `elapsed_s`, `heartbeat_age_s`, `partial`, `sources`) for every answer the server
142
+ gives.
143
+
144
+ Each `SearchResult` has:
145
+
146
+ | Field | Meaning |
147
+ |---|---|
148
+ | `id` | what `fetch()` takes: a paper id (`2301.12345`, `ssrn-…`, `journal-W…`), `doi:…` for a paper known only by its DOI, or `None` when no source we can fetch from knows it |
149
+ | `title`, `authors`, `year`, `source`, `citation_count`, `url` | the paper |
150
+ | `cached` | whether the cache holds it |
151
+ | `score` | BM25 at the keyword levels, cosine similarity otherwise: compare it only within one search |
152
+ | `match` | `keyword`, `meaning`, or `fulltext` |
153
+ | `exact` | every query word (and quoted phrase) occurs in its text |
154
+ | `passage` | the matching snippet (at `max`, the best full-text chunk), or `None` |
155
+ | `fulltext` | at `max`: whether its full text was read; `None` otherwise |
156
+ | `found_by` | the searches that found it: `index` (the cache), `s2`, `openalex`, `core`, `arxiv` |
157
+
158
+ A `SearchResult` is immutable: `authors` and `found_by` are tuples, so a result can go in a
159
+ set or key a dict. A result only CORE found has `found_by == ("core",)`.
160
+
161
+ ### How long a search may run
162
+
163
+ - By default `timeout` is the level's server budget plus 30 s: 40 s at `low`, 60 s at
164
+ `medium`, 90 s at `high`, and 150 s at `extra`. `max` has no limit. Pass a number of
165
+ seconds to set one, or `None` for none.
166
+ - The clock starts once the server has taken the search. Time spent queued behind other
167
+ searches doesn't count: the clock restarts on every poll that finds the search still
168
+ queued.
169
+ - A resubmitted search (see Retries) gets the whole `timeout` again, and the pause before
170
+ it isn't charged to it.
171
+ - Each poll waits on the server for up to 25 s, and never past the time left.
172
+ - Past the limit, `search()` raises `SearchTimeout`.
173
+
174
+ ### Results the cache doesn't hold
175
+
176
+ Results with `cached=False` can still be read: when `hit.id` isn't `None`, `fetch(hit.id)`
177
+ caches the paper now. Do it only for papers whose text you need now. The server caches
178
+ uncached results that clear its citation bar in the background, when its auto-cache is on,
179
+ unless only CORE found them: CORE isn't limited to the project's fields, so those wait for
180
+ a `fetch`.
181
+
182
+ ## Papers, bookmarks, and the corpus
183
+
184
+ ```python
185
+ from apollodorus_client import Apollodorus
186
+
187
+ with Apollodorus.from_env(owner="my-project") as ap:
188
+ hits = ap.search("volatility targeting", limit=5)
189
+ meta = ap.paper(hits[0].id) # dict: title, abstract, authors, status, text_source, citation_count, ...
190
+ body = ap.text(hits[0].id) # str: the full text (just the abstract for abstract-only papers)
191
+
192
+ uncached = next((hit for hit in hits if not hit.cached and hit.id), None)
193
+ if uncached is not None: # only when you need its text now
194
+ meta = ap.fetch(uncached.id) # slow (tens of seconds); the id may be "doi:..."
195
+ body = ap.text(meta["arxiv_id"]) # read it under the id the fetch returns
196
+
197
+ paper_id = meta["arxiv_id"]
198
+ ap.bookmark(paper_id, note="hedging section") # pins the paper against eviction
199
+ ap.bookmark(paper_id) # no note given: the note stays
200
+ ap.bookmark(paper_id, note=None) # clears the note
201
+ page = ap.bookmarks(limit=20, offset=0) # yours, newest first: {"results", "total", "limit", "offset"}
202
+ ap.unbookmark(paper_id) # {"removed": True, "published": ...}
203
+
204
+ print(ap.stats()) # corpus counts and footprint, from the nightly snapshot
205
+ print(ap.health()) # {"status": "ok"}
206
+ ```
207
+
208
+ - `paper(id)`: one paper's metadata, whatever its `status` (`active`, `evicted`, or
209
+ `parse_failed`), including whether you bookmarked it.
210
+ - `text(id)`: a cached paper's full text.
211
+ - `fetch(id)`: caches a paper now, from any source, and returns its metadata. The id may be
212
+ a search result's `doi:…`. Use the `arxiv_id` it returns from then on: it can differ from
213
+ the id you sent (a version suffix dropped, a merged work, the same paper under another
214
+ source). Fetching a paper the cache already holds returns it as it is.
215
+ - `bookmark(id, note=...)`: bookmarks a paper for your owner, fetching it first if it isn't
216
+ cached. The answer's `published` is `False` while the change is still on its way to the
217
+ shared index: in flight, never lost.
218
+ - `unbookmark(id)`: removes it. Removing a bookmark that isn't there is not an error.
219
+ - `bookmarks(*, limit=20, offset=0)`: one page of your bookmarks.
220
+ - `stats()` and `health()`: the corpus's counts and footprint, and a liveness check.
221
+
222
+ `fetch` and `bookmark` allow 300 s, since they can download and parse a paper, and each of
223
+ `search()`'s polls allows its wait plus 30 s. Every other request uses the client's
224
+ `timeout`.
225
+
226
+ ## Retries
227
+
228
+ - `search()` and the reads (`paper`, `text`, `bookmarks`, `stats`, `health`) retry a 429, a
229
+ 502, 503, or 504, and a dropped connection: 5 attempts in all, about 1, 2, 4, then 8 s
230
+ apart. Each pause varies by up to 25% either way, so clients that failed together don't
231
+ come back together.
232
+ - A `Retry-After` header on a 429 or 503, in seconds or as an HTTP date, is waited for when
233
+ it asks for longer than that pause, for at most 30 s. The server sends `Retry-After: 5`
234
+ when its search queue is full or it is restarting.
235
+ - A search the server lost (a poll answered 404, after a restart, say) is submitted again at
236
+ once. One that failed, or that stalled (running, but silent for over 60 s), is submitted
237
+ again after a pause of about 1, 2, then 4 s. A search is resubmitted 3 times at most,
238
+ whatever the causes.
239
+ - The writes (`fetch`, `bookmark`, `unbookmark`) make one attempt, since a write whose answer
240
+ was lost may still have landed.
241
+ - A request this client built wrong (`httpx.LocalProtocolError`) is never retried.
242
+
243
+ ## Errors
244
+
245
+ - `ApollodorusError` (`.status`, `.detail`): the API answered with an error, after any
246
+ retries. It is also raised when a search failed or stalled through 3 resubmits, or
247
+ reported a status this client doesn't know, and when the answer can't be the API's:
248
+ - A redirect is never followed, so the key can't be forwarded to another host. Use the
249
+ final `https://` address.
250
+ - A web page means `base_url` is likely the web UI's address: the API is under it at
251
+ `/api`.
252
+ - `SearchTimeout` (a `TimeoutError`): the search ran past `search()`'s `timeout`.
253
+ - `httpx.TransportError`: the server stayed unreachable through every attempt.
254
+ - `ValueError`: a bad argument, such as an unknown `effort`, a malformed `owner` or
255
+ `api_key`, or a `base_url` that isn't an absolute `http(s)://` URL with a host.
256
+ - `RuntimeError`: `from_env()` found no URL or no key.
257
+
258
+ Both exception types pickle and copy cleanly, so they cross process boundaries.
259
+
260
+ ## Threads and lifetime
261
+
262
+ One client per process is enough: it pools connections. It keeps no per-call state, and the
263
+ `httpx.Client` it wraps can be shared between threads, so every thread can use the same
264
+ client. `search()` blocks the calling thread until the results arrive. Use the client as a
265
+ context manager, or call `close()` when you're done with it. Closing never closes an
266
+ `http_client` you passed in.
@@ -0,0 +1,244 @@
1
+ # apollodorus-client
2
+
3
+ Python client for the Apollodorus paper-search API: FluxGate's cache of top research
4
+ papers (quant finance, CS, statistics, economics), with search that reaches beyond it.
5
+
6
+ ```bash
7
+ pip install apollodorus-client # add [dotenv] to read a .env file
8
+ ```
9
+
10
+ ```python
11
+ from apollodorus_client import Apollodorus
12
+
13
+ with Apollodorus.from_env() as ap: # APOLLODORUS_URL + APOLLODORUS_API_KEY
14
+ hits = ap.search("momentum crash", effort="medium")
15
+ for hit in hits:
16
+ print(f"{hit.score:8.3f} {hit.id} cached={hit.cached} {hit.title}")
17
+ print(hits.sources, "partial" if hits.partial else "complete")
18
+ ```
19
+
20
+ Python 3.10 or later; the only dependency is `httpx`. The server also serves a guide for AI
21
+ agents at `{APOLLODORUS_URL}/llms.txt` and its OpenAPI schema at
22
+ `{APOLLODORUS_URL}/openapi.json`.
23
+
24
+ ## Configuration
25
+
26
+ `Apollodorus.from_env(*, owner=None, timeout=60.0)` reads three variables:
27
+
28
+ | Variable | Meaning |
29
+ |---|---|
30
+ | `APOLLODORUS_URL` | the API root, e.g. `https://<your-host>/apollodorus/api` (required) |
31
+ | `APOLLODORUS_API_KEY` | the API key, sent as the `X-API-Key` header (required) |
32
+ | `APOLLODORUS_BOOKMARK_OWNER` | optional: the name your bookmarks are kept under (`owner` overrides it) |
33
+
34
+ With the `[dotenv]` extra installed, a `.env` file in the working directory, or in a
35
+ directory above it, supplies what the environment lacks; a real environment variable wins
36
+ over the file. Only those three names are read, from either. `from_env()` raises
37
+ `RuntimeError` when the URL or the key is missing.
38
+
39
+ Or build the client yourself:
40
+
41
+ ```python
42
+ import os
43
+
44
+ from apollodorus_client import Apollodorus
45
+
46
+ ap = Apollodorus(
47
+ "https://<your-host>/apollodorus/api",
48
+ os.environ["APOLLODORUS_API_KEY"],
49
+ owner="my-project", # optional: sent as X-Bookmark-Owner
50
+ timeout=60.0, # seconds per request; see "Papers, bookmarks, and the corpus" for the exceptions
51
+ )
52
+ ```
53
+
54
+ - `base_url` must be an absolute `http://` or `https://` URL with a host.
55
+ - `api_key` must be visible ASCII: no spaces, control characters, or curly quotes.
56
+ Whitespace around it, such as a trailing newline from a secrets file, is dropped. A
57
+ refused key is never repeated in the error.
58
+ - `owner` is 1 to 64 lowercase letters, digits, `.`, `_`, or `-`, starting with a letter or
59
+ digit. Without one, the server files your bookmarks under the shared owner `api-key`.
60
+ - `http_client` (optional) is an `httpx.Client` to send through. It is used as it is: never
61
+ modified, never closed, and never given the key or the base URL, which go only on this
62
+ client's own requests. Its own timeout then applies in place of `timeout`.
63
+
64
+ Keep the key in the environment or a `.env` file, never in code or URLs, and use an
65
+ `https://` URL. A plain `http://` `base_url` to any host but `localhost`, `127.0.0.0/8`, or
66
+ `::1` makes the constructor issue a `UserWarning`, because the key would travel in
67
+ cleartext.
68
+
69
+ ## Effort levels
70
+
71
+ | `effort` | Matches by | Searches | Server time budget |
72
+ |---|---|---|---|
73
+ | `low` (default) | keywords | the cache: titles, abstracts, full text | 10 s |
74
+ | `medium` | keywords | the cache + Semantic Scholar, OpenAlex, CORE, arXiv | 30 s |
75
+ | `high` | meaning | the cache + Semantic Scholar | 60 s |
76
+ | `extra` | meaning | the cache + Semantic Scholar, OpenAlex, CORE, arXiv | 120 s |
77
+ | `max` | meaning, over full text | `extra`, then reads the top `top_n` papers in full | none (minutes) |
78
+
79
+ Keyword levels need every word, in any order; `"quoted words"` must appear as a phrase.
80
+ Meaning levels find papers about the query even when they share none of its words. A level
81
+ that runs out of time answers with what it found, and `hits.partial` is `True`. Start at
82
+ `low`, and go up when the results fall short. `EFFORTS` and `DEADLINES` (seconds, `None` for
83
+ `max`) are importable from the package.
84
+
85
+ ## Searching
86
+
87
+ ```python
88
+ from apollodorus_client import Apollodorus, JobStatus
89
+
90
+
91
+ def show(status: JobStatus) -> None: # called with each poll's answer
92
+ print(f"{status.status:<8} {status.stage} {status.elapsed_s:.0f} s")
93
+
94
+
95
+ with Apollodorus.from_env() as ap:
96
+ hits = ap.search(
97
+ "regime switching asset allocation",
98
+ effort="high",
99
+ limit=50, # 1-100 results, default 20
100
+ since=2015, until=2024, # publication years
101
+ field="q-fin", # a coarse field group: q-fin, econ, cs, stat, ...
102
+ timeout=120, # seconds the search may run; None waits as long as it takes
103
+ on_progress=show,
104
+ )
105
+ print(hits.job_id, hits.partial, hits.sources, hits.elapsed_s)
106
+
107
+ deep = ap.search("intraday momentum and liquidity", effort="max", top_n=10) # minutes
108
+ for hit in deep[:10]:
109
+ print(hit.fulltext, hit.match, (hit.passage or "")[:200])
110
+ ```
111
+
112
+ `search(query, *, effort="low", limit=20, since=None, until=None, field=None, top_n=20,
113
+ timeout=..., on_progress=None)` submits the search as a job on the server and long-polls it
114
+ until it is done. It returns `SearchResults`: a list of `SearchResult`, best first, that also
115
+ carries `.partial`, `.sources` (each source's `ok`, `failed`, `skipped`, or `timeout`),
116
+ `.job_id`, and `.elapsed_s`. `top_n` (1–50) is read at `max` only: the papers whose full text
117
+ is read. Out-of-range values (a `limit` of 500, say) are refused by the server, as an
118
+ `ApollodorusError` with status 422. `on_progress` receives a `JobStatus` (`id`, `status`,
119
+ `stage`, `elapsed_s`, `heartbeat_age_s`, `partial`, `sources`) for every answer the server
120
+ gives.
121
+
122
+ Each `SearchResult` has:
123
+
124
+ | Field | Meaning |
125
+ |---|---|
126
+ | `id` | what `fetch()` takes: a paper id (`2301.12345`, `ssrn-…`, `journal-W…`), `doi:…` for a paper known only by its DOI, or `None` when no source we can fetch from knows it |
127
+ | `title`, `authors`, `year`, `source`, `citation_count`, `url` | the paper |
128
+ | `cached` | whether the cache holds it |
129
+ | `score` | BM25 at the keyword levels, cosine similarity otherwise: compare it only within one search |
130
+ | `match` | `keyword`, `meaning`, or `fulltext` |
131
+ | `exact` | every query word (and quoted phrase) occurs in its text |
132
+ | `passage` | the matching snippet (at `max`, the best full-text chunk), or `None` |
133
+ | `fulltext` | at `max`: whether its full text was read; `None` otherwise |
134
+ | `found_by` | the searches that found it: `index` (the cache), `s2`, `openalex`, `core`, `arxiv` |
135
+
136
+ A `SearchResult` is immutable: `authors` and `found_by` are tuples, so a result can go in a
137
+ set or key a dict. A result only CORE found has `found_by == ("core",)`.
138
+
139
+ ### How long a search may run
140
+
141
+ - By default `timeout` is the level's server budget plus 30 s: 40 s at `low`, 60 s at
142
+ `medium`, 90 s at `high`, and 150 s at `extra`. `max` has no limit. Pass a number of
143
+ seconds to set one, or `None` for none.
144
+ - The clock starts once the server has taken the search. Time spent queued behind other
145
+ searches doesn't count: the clock restarts on every poll that finds the search still
146
+ queued.
147
+ - A resubmitted search (see Retries) gets the whole `timeout` again, and the pause before
148
+ it isn't charged to it.
149
+ - Each poll waits on the server for up to 25 s, and never past the time left.
150
+ - Past the limit, `search()` raises `SearchTimeout`.
151
+
152
+ ### Results the cache doesn't hold
153
+
154
+ Results with `cached=False` can still be read: when `hit.id` isn't `None`, `fetch(hit.id)`
155
+ caches the paper now. Do it only for papers whose text you need now. The server caches
156
+ uncached results that clear its citation bar in the background, when its auto-cache is on,
157
+ unless only CORE found them: CORE isn't limited to the project's fields, so those wait for
158
+ a `fetch`.
159
+
160
+ ## Papers, bookmarks, and the corpus
161
+
162
+ ```python
163
+ from apollodorus_client import Apollodorus
164
+
165
+ with Apollodorus.from_env(owner="my-project") as ap:
166
+ hits = ap.search("volatility targeting", limit=5)
167
+ meta = ap.paper(hits[0].id) # dict: title, abstract, authors, status, text_source, citation_count, ...
168
+ body = ap.text(hits[0].id) # str: the full text (just the abstract for abstract-only papers)
169
+
170
+ uncached = next((hit for hit in hits if not hit.cached and hit.id), None)
171
+ if uncached is not None: # only when you need its text now
172
+ meta = ap.fetch(uncached.id) # slow (tens of seconds); the id may be "doi:..."
173
+ body = ap.text(meta["arxiv_id"]) # read it under the id the fetch returns
174
+
175
+ paper_id = meta["arxiv_id"]
176
+ ap.bookmark(paper_id, note="hedging section") # pins the paper against eviction
177
+ ap.bookmark(paper_id) # no note given: the note stays
178
+ ap.bookmark(paper_id, note=None) # clears the note
179
+ page = ap.bookmarks(limit=20, offset=0) # yours, newest first: {"results", "total", "limit", "offset"}
180
+ ap.unbookmark(paper_id) # {"removed": True, "published": ...}
181
+
182
+ print(ap.stats()) # corpus counts and footprint, from the nightly snapshot
183
+ print(ap.health()) # {"status": "ok"}
184
+ ```
185
+
186
+ - `paper(id)`: one paper's metadata, whatever its `status` (`active`, `evicted`, or
187
+ `parse_failed`), including whether you bookmarked it.
188
+ - `text(id)`: a cached paper's full text.
189
+ - `fetch(id)`: caches a paper now, from any source, and returns its metadata. The id may be
190
+ a search result's `doi:…`. Use the `arxiv_id` it returns from then on: it can differ from
191
+ the id you sent (a version suffix dropped, a merged work, the same paper under another
192
+ source). Fetching a paper the cache already holds returns it as it is.
193
+ - `bookmark(id, note=...)`: bookmarks a paper for your owner, fetching it first if it isn't
194
+ cached. The answer's `published` is `False` while the change is still on its way to the
195
+ shared index: in flight, never lost.
196
+ - `unbookmark(id)`: removes it. Removing a bookmark that isn't there is not an error.
197
+ - `bookmarks(*, limit=20, offset=0)`: one page of your bookmarks.
198
+ - `stats()` and `health()`: the corpus's counts and footprint, and a liveness check.
199
+
200
+ `fetch` and `bookmark` allow 300 s, since they can download and parse a paper, and each of
201
+ `search()`'s polls allows its wait plus 30 s. Every other request uses the client's
202
+ `timeout`.
203
+
204
+ ## Retries
205
+
206
+ - `search()` and the reads (`paper`, `text`, `bookmarks`, `stats`, `health`) retry a 429, a
207
+ 502, 503, or 504, and a dropped connection: 5 attempts in all, about 1, 2, 4, then 8 s
208
+ apart. Each pause varies by up to 25% either way, so clients that failed together don't
209
+ come back together.
210
+ - A `Retry-After` header on a 429 or 503, in seconds or as an HTTP date, is waited for when
211
+ it asks for longer than that pause, for at most 30 s. The server sends `Retry-After: 5`
212
+ when its search queue is full or it is restarting.
213
+ - A search the server lost (a poll answered 404, after a restart, say) is submitted again at
214
+ once. One that failed, or that stalled (running, but silent for over 60 s), is submitted
215
+ again after a pause of about 1, 2, then 4 s. A search is resubmitted 3 times at most,
216
+ whatever the causes.
217
+ - The writes (`fetch`, `bookmark`, `unbookmark`) make one attempt, since a write whose answer
218
+ was lost may still have landed.
219
+ - A request this client built wrong (`httpx.LocalProtocolError`) is never retried.
220
+
221
+ ## Errors
222
+
223
+ - `ApollodorusError` (`.status`, `.detail`): the API answered with an error, after any
224
+ retries. It is also raised when a search failed or stalled through 3 resubmits, or
225
+ reported a status this client doesn't know, and when the answer can't be the API's:
226
+ - A redirect is never followed, so the key can't be forwarded to another host. Use the
227
+ final `https://` address.
228
+ - A web page means `base_url` is likely the web UI's address: the API is under it at
229
+ `/api`.
230
+ - `SearchTimeout` (a `TimeoutError`): the search ran past `search()`'s `timeout`.
231
+ - `httpx.TransportError`: the server stayed unreachable through every attempt.
232
+ - `ValueError`: a bad argument, such as an unknown `effort`, a malformed `owner` or
233
+ `api_key`, or a `base_url` that isn't an absolute `http(s)://` URL with a host.
234
+ - `RuntimeError`: `from_env()` found no URL or no key.
235
+
236
+ Both exception types pickle and copy cleanly, so they cross process boundaries.
237
+
238
+ ## Threads and lifetime
239
+
240
+ One client per process is enough: it pools connections. It keeps no per-call state, and the
241
+ `httpx.Client` it wraps can be shared between threads, so every thread can use the same
242
+ client. `search()` blocks the calling thread until the results arrive. Use the client as a
243
+ context manager, or call `close()` when you're done with it. Closing never closes an
244
+ `http_client` you passed in.