apollodorus-client 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- apollodorus_client-0.1.0/.flake8 +2 -0
- apollodorus_client-0.1.0/.gitignore +4 -0
- apollodorus_client-0.1.0/CHANGELOG.md +48 -0
- apollodorus_client-0.1.0/CLAUDE.md +104 -0
- apollodorus_client-0.1.0/PKG-INFO +266 -0
- apollodorus_client-0.1.0/README.md +244 -0
- apollodorus_client-0.1.0/pyproject.toml +44 -0
- apollodorus_client-0.1.0/src/apollodorus_client/__init__.py +33 -0
- apollodorus_client-0.1.0/src/apollodorus_client/_client.py +493 -0
- apollodorus_client-0.1.0/src/apollodorus_client/py.typed +0 -0
- apollodorus_client-0.1.0/tests/test_client.py +1023 -0
|
@@ -0,0 +1,48 @@
|
|
|
1
|
+
# Changelog
|
|
2
|
+
|
|
3
|
+
All notable changes to `apollodorus-client` are recorded here. The format follows
|
|
4
|
+
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and versions follow
|
|
5
|
+
[Semantic Versioning](https://semver.org/) under the 0.x rules in
|
|
6
|
+
[`CLAUDE.md`](CLAUDE.md#choosing-the-version).
|
|
7
|
+
|
|
8
|
+
Every release needs its own `## [X.Y.Z] - YYYY-MM-DD` section, because the release
|
|
9
|
+
workflow refuses a version without one. Add changes under `## [Unreleased]` as you make
|
|
10
|
+
them. When you release, rename that heading to the new version, and start a fresh
|
|
11
|
+
`## [Unreleased]` above it.
|
|
12
|
+
|
|
13
|
+
## [Unreleased]
|
|
14
|
+
|
|
15
|
+
## [0.1.0] - 2026-10-08
|
|
16
|
+
|
|
17
|
+
The first release.
|
|
18
|
+
|
|
19
|
+
### Added
|
|
20
|
+
|
|
21
|
+
- `Apollodorus.search(query, effort=...)` runs effort-level searches (`low`, `medium`,
|
|
22
|
+
`high`, `extra`, `max`):
|
|
23
|
+
- it submits a search job and long-polls it to the end;
|
|
24
|
+
- it resubmits after a server restart or a stalled job, up to 3 times;
|
|
25
|
+
- its timeout counts only running time: time spent queued doesn't count, and the
|
|
26
|
+
clock restarts after a resubmit;
|
|
27
|
+
- it takes the filters `limit`, `since`, `until`, `field`, and `top_n`, plus an
|
|
28
|
+
`on_progress` callback.
|
|
29
|
+
- `paper`, `text`, `fetch`, `bookmark` (with an optional note), `unbookmark`, `bookmarks`,
|
|
30
|
+
`stats`, and `health`.
|
|
31
|
+
- `Apollodorus.from_env()` reads `APOLLODORUS_URL`, `APOLLODORUS_API_KEY`, and
|
|
32
|
+
`APOLLODORUS_BOOKMARK_OWNER` from the environment. With the `[dotenv]` extra it also
|
|
33
|
+
reads them from a `.env` file.
|
|
34
|
+
- Retries:
|
|
35
|
+
- `search` and the read-only calls retry transport errors and 429/502/503/504 up to
|
|
36
|
+
5 times, with a jittered backoff;
|
|
37
|
+
- a `Retry-After` header is honored, capped at 30 s.
|
|
38
|
+
- Typed, immutable results: `SearchResult`, a `SearchResults` list that carries the job's
|
|
39
|
+
`partial` flag and its `sources`, and `JobStatus`. The package ships `py.typed`.
|
|
40
|
+
- Errors:
|
|
41
|
+
- `ApollodorusError` carries `status` and `detail`, and can be pickled;
|
|
42
|
+
- a search that runs out of time raises `SearchTimeout`;
|
|
43
|
+
- a wrong `base_url` gets a clear hint;
|
|
44
|
+
- a key or owner that can't be sent as a header is refused, and the error text never
|
|
45
|
+
includes it.
|
|
46
|
+
- Safety:
|
|
47
|
+
- the client never follows a redirect, so the key never leaves the API host;
|
|
48
|
+
- an `http://` URL to a non-loopback host warns once.
|
|
@@ -0,0 +1,104 @@
|
|
|
1
|
+
# client/ — apollodorus-client, a public PyPI package
|
|
2
|
+
|
|
3
|
+
This is the pip-installable Python client for the Apollodorus HTTP API (`docs/api.md`):
|
|
4
|
+
|
|
5
|
+
```bash
|
|
6
|
+
pip install apollodorus-client
|
|
7
|
+
```
|
|
8
|
+
|
|
9
|
+
Anyone can read it on PyPI, so it must hold no secrets and no internal hostnames. Write
|
|
10
|
+
examples with `https://<your-host>/apollodorus/api`. Its only dependency is `httpx`;
|
|
11
|
+
`python-dotenv` comes with the `[dotenv]` extra.
|
|
12
|
+
|
|
13
|
+
## Layout
|
|
14
|
+
|
|
15
|
+
- `src/apollodorus_client/__init__.py` holds the public exports (`__all__`) and
|
|
16
|
+
`__version__`, the single source of the version (hatch reads it, see `pyproject.toml`).
|
|
17
|
+
- `src/apollodorus_client/_client.py` holds all the code.
|
|
18
|
+
- `tests/` runs offline against `httpx.MockTransport`.
|
|
19
|
+
- The root suite's `tests/unit/test_client_e2e.py` drives the client against the real app,
|
|
20
|
+
and pins the client's deadlines to the server's.
|
|
21
|
+
- `README.md` is the package's PyPI page. `CHANGELOG.md` has one section per release.
|
|
22
|
+
|
|
23
|
+
Check from `client/` (CI's `client` job runs the same):
|
|
24
|
+
|
|
25
|
+
```bash
|
|
26
|
+
pytest -q && flake8 src tests && black --check src tests && isort --check-only src tests && mypy
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
## Releasing (automatic)
|
|
30
|
+
|
|
31
|
+
`.github/workflows/publish-client.yml` runs on every push to `main` that touches
|
|
32
|
+
`client/`. It reads `__version__` and compares it with what PyPI already has:
|
|
33
|
+
|
|
34
|
+
- **The version is already on PyPI:** nothing is published. A test, CI, or comment change
|
|
35
|
+
can merge without a version bump.
|
|
36
|
+
- **The version is new:**
|
|
37
|
+
1. The `check` job validates it: valid PEP 440, `MAJOR.MINOR.PATCH` with an optional
|
|
38
|
+
`aN`/`bN`/`rcN` suffix, higher than every version on PyPI, and with a
|
|
39
|
+
`## [X.Y.Z]` section in `CHANGELOG.md`. Any failure fails the run and names the fix.
|
|
40
|
+
2. The workflow then tests, builds, checks the wheel on Python 3.10 and with
|
|
41
|
+
`twine check --strict`, and publishes to PyPI.
|
|
42
|
+
3. Finally it tags `client-vX.Y.Z` and opens a GitHub release whose notes are that
|
|
43
|
+
CHANGELOG section.
|
|
44
|
+
|
|
45
|
+
So **a release is one merge**. Its pull request carries three changes together:
|
|
46
|
+
|
|
47
|
+
1. the code change and its tests;
|
|
48
|
+
2. the `CHANGELOG.md` section: move the `[Unreleased]` notes under
|
|
49
|
+
`## [X.Y.Z] - YYYY-MM-DD`, and leave an empty `## [Unreleased]` above it;
|
|
50
|
+
3. the bumped `__version__`.
|
|
51
|
+
|
|
52
|
+
Do not create or push `client-v*` tags by hand: the workflow makes them, and nothing
|
|
53
|
+
triggers on them. To retry a failed run, use "Run workflow" on `main`. A re-run is safe:
|
|
54
|
+
files PyPI already has are skipped, and so is a tag or release that already exists.
|
|
55
|
+
|
|
56
|
+
Anything a user would notice needs a release, including a README fix, because the README
|
|
57
|
+
is the PyPI page and only a new version updates it. Changes outside `client/` (the server,
|
|
58
|
+
the UI, deploys) never need a client release, unless the client has to change with them.
|
|
59
|
+
|
|
60
|
+
## Choosing the version
|
|
61
|
+
|
|
62
|
+
Use [SemVer](https://semver.org/) with 0.x rules: the package isn't 1.0 yet. The public
|
|
63
|
+
API is:
|
|
64
|
+
|
|
65
|
+
- everything in `__all__`;
|
|
66
|
+
- the methods of `Apollodorus`: their names, parameters, defaults, and return types;
|
|
67
|
+
- the fields of `SearchResult`, `SearchResults`, and `JobStatus`;
|
|
68
|
+
- the exception types and their `status` and `detail`;
|
|
69
|
+
- the environment variables `APOLLODORUS_URL`, `APOLLODORUS_API_KEY`, and
|
|
70
|
+
`APOLLODORUS_BOOKMARK_OWNER`;
|
|
71
|
+
- the retry, timeout, and resubmit behavior the README documents.
|
|
72
|
+
|
|
73
|
+
| Bump | When | Example |
|
|
74
|
+
|---|---|---|
|
|
75
|
+
| **PATCH** (0.1.0 → 0.1.1) | A fix that keeps the documented behavior: bugs, wrong error text, README or docstring fixes, internal refactors, and loosening a dependency bound | a poll that spun is now bounded |
|
|
76
|
+
| **MINOR** (0.1.x → 0.2.0) | Everything else, breaking changes included while 0.x: new methods, parameters, or result fields; changed defaults or retry/timeout behavior; a call to a route older servers lack; a higher minimum Python or `httpx`; anything renamed or removed | `search(..., sort=)`; a renamed exception |
|
|
77
|
+
| **MAJOR** (→ 1.0.0) | Only when the owner declares the API stable. Normal SemVer applies after that: breaking means MAJOR. | — |
|
|
78
|
+
|
|
79
|
+
Rules:
|
|
80
|
+
|
|
81
|
+
- **When in doubt between PATCH and MINOR, choose MINOR.**
|
|
82
|
+
- **List every breaking change** under `### Breaking` in that version's CHANGELOG section,
|
|
83
|
+
and say how to migrate.
|
|
84
|
+
- **Pre-releases** (`0.2.0rc1`) are for trying a risky release first. pip ignores them
|
|
85
|
+
unless asked (`pip install --pre apollodorus-client`). The final `0.2.0` then
|
|
86
|
+
supersedes them.
|
|
87
|
+
- **Versions only go up.** PyPI never accepts a version twice, even one deleted or
|
|
88
|
+
yanked, and the workflow refuses a version below the highest on PyPI. That means there
|
|
89
|
+
are no backport releases for an old minor line: fix forward.
|
|
90
|
+
- **If a bad release got out,** yank it on PyPI (it stays installable when pinned) and
|
|
91
|
+
release the fix as the next PATCH.
|
|
92
|
+
|
|
93
|
+
## Compatibility with the server
|
|
94
|
+
|
|
95
|
+
The client follows the HTTP API in `docs/api.md`:
|
|
96
|
+
|
|
97
|
+
- **A new response field:** old clients ignore it, so no release is needed until the
|
|
98
|
+
client exposes it (MINOR).
|
|
99
|
+
- **A server change that would break released clients:** keep the old behavior on the
|
|
100
|
+
server for a while, release a client that handles both, and only then drop the old one.
|
|
101
|
+
- **Pinned values:** `DEADLINES`, `STALE_HEARTBEAT_SECONDS`, and `POLL_WAIT_SECONDS` must
|
|
102
|
+
equal the server's (`effort.DEADLINES`, `jobs.STALE_SECONDS`, the long poll's cap), and
|
|
103
|
+
`tests/unit/test_client_e2e.py` fails when they drift. A server change to any of them
|
|
104
|
+
needs a client release, usually MINOR because the timeout behavior changes.
|
|
@@ -0,0 +1,266 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: apollodorus-client
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Python client for the Apollodorus paper-search API
|
|
5
|
+
Classifier: Intended Audience :: Developers
|
|
6
|
+
Classifier: Operating System :: OS Independent
|
|
7
|
+
Classifier: Programming Language :: Python :: 3
|
|
8
|
+
Classifier: Programming Language :: Python :: 3 :: Only
|
|
9
|
+
Classifier: Typing :: Typed
|
|
10
|
+
Requires-Python: >=3.10
|
|
11
|
+
Requires-Dist: httpx>=0.27
|
|
12
|
+
Provides-Extra: dev
|
|
13
|
+
Requires-Dist: black==26.5.1; extra == 'dev'
|
|
14
|
+
Requires-Dist: flake8==7.3.0; extra == 'dev'
|
|
15
|
+
Requires-Dist: isort==8.0.1; extra == 'dev'
|
|
16
|
+
Requires-Dist: mypy==2.1.0; extra == 'dev'
|
|
17
|
+
Requires-Dist: pytest==9.0.3; extra == 'dev'
|
|
18
|
+
Requires-Dist: python-dotenv==1.2.2; extra == 'dev'
|
|
19
|
+
Provides-Extra: dotenv
|
|
20
|
+
Requires-Dist: python-dotenv>=1.0; extra == 'dotenv'
|
|
21
|
+
Description-Content-Type: text/markdown
|
|
22
|
+
|
|
23
|
+
# apollodorus-client
|
|
24
|
+
|
|
25
|
+
Python client for the Apollodorus paper-search API: FluxGate's cache of top research
|
|
26
|
+
papers (quant finance, CS, statistics, economics), with search that reaches beyond it.
|
|
27
|
+
|
|
28
|
+
```bash
|
|
29
|
+
pip install apollodorus-client # add [dotenv] to read a .env file
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
```python
|
|
33
|
+
from apollodorus_client import Apollodorus
|
|
34
|
+
|
|
35
|
+
with Apollodorus.from_env() as ap: # APOLLODORUS_URL + APOLLODORUS_API_KEY
|
|
36
|
+
hits = ap.search("momentum crash", effort="medium")
|
|
37
|
+
for hit in hits:
|
|
38
|
+
print(f"{hit.score:8.3f} {hit.id} cached={hit.cached} {hit.title}")
|
|
39
|
+
print(hits.sources, "partial" if hits.partial else "complete")
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
Python 3.10 or later; the only dependency is `httpx`. The server also serves a guide for AI
|
|
43
|
+
agents at `{APOLLODORUS_URL}/llms.txt` and its OpenAPI schema at
|
|
44
|
+
`{APOLLODORUS_URL}/openapi.json`.
|
|
45
|
+
|
|
46
|
+
## Configuration
|
|
47
|
+
|
|
48
|
+
`Apollodorus.from_env(*, owner=None, timeout=60.0)` reads three variables:
|
|
49
|
+
|
|
50
|
+
| Variable | Meaning |
|
|
51
|
+
|---|---|
|
|
52
|
+
| `APOLLODORUS_URL` | the API root, e.g. `https://<your-host>/apollodorus/api` (required) |
|
|
53
|
+
| `APOLLODORUS_API_KEY` | the API key, sent as the `X-API-Key` header (required) |
|
|
54
|
+
| `APOLLODORUS_BOOKMARK_OWNER` | optional: the name your bookmarks are kept under (`owner` overrides it) |
|
|
55
|
+
|
|
56
|
+
With the `[dotenv]` extra installed, a `.env` file in the working directory, or in a
|
|
57
|
+
directory above it, supplies what the environment lacks; a real environment variable wins
|
|
58
|
+
over the file. Only those three names are read, from either. `from_env()` raises
|
|
59
|
+
`RuntimeError` when the URL or the key is missing.
|
|
60
|
+
|
|
61
|
+
Or build the client yourself:
|
|
62
|
+
|
|
63
|
+
```python
|
|
64
|
+
import os
|
|
65
|
+
|
|
66
|
+
from apollodorus_client import Apollodorus
|
|
67
|
+
|
|
68
|
+
ap = Apollodorus(
|
|
69
|
+
"https://<your-host>/apollodorus/api",
|
|
70
|
+
os.environ["APOLLODORUS_API_KEY"],
|
|
71
|
+
owner="my-project", # optional: sent as X-Bookmark-Owner
|
|
72
|
+
timeout=60.0, # seconds per request; see "Papers, bookmarks, and the corpus" for the exceptions
|
|
73
|
+
)
|
|
74
|
+
```
|
|
75
|
+
|
|
76
|
+
- `base_url` must be an absolute `http://` or `https://` URL with a host.
|
|
77
|
+
- `api_key` must be visible ASCII: no spaces, control characters, or curly quotes.
|
|
78
|
+
Whitespace around it, such as a trailing newline from a secrets file, is dropped. A
|
|
79
|
+
refused key is never repeated in the error.
|
|
80
|
+
- `owner` is 1 to 64 lowercase letters, digits, `.`, `_`, or `-`, starting with a letter or
|
|
81
|
+
digit. Without one, the server files your bookmarks under the shared owner `api-key`.
|
|
82
|
+
- `http_client` (optional) is an `httpx.Client` to send through. It is used as it is: never
|
|
83
|
+
modified, never closed, and never given the key or the base URL, which go only on this
|
|
84
|
+
client's own requests. Its own timeout then applies in place of `timeout`.
|
|
85
|
+
|
|
86
|
+
Keep the key in the environment or a `.env` file, never in code or URLs, and use an
|
|
87
|
+
`https://` URL. A plain `http://` `base_url` to any host but `localhost`, `127.0.0.0/8`, or
|
|
88
|
+
`::1` makes the constructor issue a `UserWarning`, because the key would travel in
|
|
89
|
+
cleartext.
|
|
90
|
+
|
|
91
|
+
## Effort levels
|
|
92
|
+
|
|
93
|
+
| `effort` | Matches by | Searches | Server time budget |
|
|
94
|
+
|---|---|---|---|
|
|
95
|
+
| `low` (default) | keywords | the cache: titles, abstracts, full text | 10 s |
|
|
96
|
+
| `medium` | keywords | the cache + Semantic Scholar, OpenAlex, CORE, arXiv | 30 s |
|
|
97
|
+
| `high` | meaning | the cache + Semantic Scholar | 60 s |
|
|
98
|
+
| `extra` | meaning | the cache + Semantic Scholar, OpenAlex, CORE, arXiv | 120 s |
|
|
99
|
+
| `max` | meaning, over full text | `extra`, then reads the top `top_n` papers in full | none (minutes) |
|
|
100
|
+
|
|
101
|
+
Keyword levels need every word, in any order; `"quoted words"` must appear as a phrase.
|
|
102
|
+
Meaning levels find papers about the query even when they share none of its words. A level
|
|
103
|
+
that runs out of time answers with what it found, and `hits.partial` is `True`. Start at
|
|
104
|
+
`low`, and go up when the results fall short. `EFFORTS` and `DEADLINES` (seconds, `None` for
|
|
105
|
+
`max`) are importable from the package.
|
|
106
|
+
|
|
107
|
+
## Searching
|
|
108
|
+
|
|
109
|
+
```python
|
|
110
|
+
from apollodorus_client import Apollodorus, JobStatus
|
|
111
|
+
|
|
112
|
+
|
|
113
|
+
def show(status: JobStatus) -> None: # called with each poll's answer
|
|
114
|
+
print(f"{status.status:<8} {status.stage} {status.elapsed_s:.0f} s")
|
|
115
|
+
|
|
116
|
+
|
|
117
|
+
with Apollodorus.from_env() as ap:
|
|
118
|
+
hits = ap.search(
|
|
119
|
+
"regime switching asset allocation",
|
|
120
|
+
effort="high",
|
|
121
|
+
limit=50, # 1-100 results, default 20
|
|
122
|
+
since=2015, until=2024, # publication years
|
|
123
|
+
field="q-fin", # a coarse field group: q-fin, econ, cs, stat, ...
|
|
124
|
+
timeout=120, # seconds the search may run; None waits as long as it takes
|
|
125
|
+
on_progress=show,
|
|
126
|
+
)
|
|
127
|
+
print(hits.job_id, hits.partial, hits.sources, hits.elapsed_s)
|
|
128
|
+
|
|
129
|
+
deep = ap.search("intraday momentum and liquidity", effort="max", top_n=10) # minutes
|
|
130
|
+
for hit in deep[:10]:
|
|
131
|
+
print(hit.fulltext, hit.match, (hit.passage or "")[:200])
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
`search(query, *, effort="low", limit=20, since=None, until=None, field=None, top_n=20,
|
|
135
|
+
timeout=..., on_progress=None)` submits the search as a job on the server and long-polls it
|
|
136
|
+
until it is done. It returns `SearchResults`: a list of `SearchResult`, best first, that also
|
|
137
|
+
carries `.partial`, `.sources` (each source's `ok`, `failed`, `skipped`, or `timeout`),
|
|
138
|
+
`.job_id`, and `.elapsed_s`. `top_n` (1–50) is read at `max` only: the papers whose full text
|
|
139
|
+
is read. Out-of-range values (a `limit` of 500, say) are refused by the server, as an
|
|
140
|
+
`ApollodorusError` with status 422. `on_progress` receives a `JobStatus` (`id`, `status`,
|
|
141
|
+
`stage`, `elapsed_s`, `heartbeat_age_s`, `partial`, `sources`) for every answer the server
|
|
142
|
+
gives.
|
|
143
|
+
|
|
144
|
+
Each `SearchResult` has:
|
|
145
|
+
|
|
146
|
+
| Field | Meaning |
|
|
147
|
+
|---|---|
|
|
148
|
+
| `id` | what `fetch()` takes: a paper id (`2301.12345`, `ssrn-…`, `journal-W…`), `doi:…` for a paper known only by its DOI, or `None` when no source we can fetch from knows it |
|
|
149
|
+
| `title`, `authors`, `year`, `source`, `citation_count`, `url` | the paper |
|
|
150
|
+
| `cached` | whether the cache holds it |
|
|
151
|
+
| `score` | BM25 at the keyword levels, cosine similarity otherwise: compare it only within one search |
|
|
152
|
+
| `match` | `keyword`, `meaning`, or `fulltext` |
|
|
153
|
+
| `exact` | every query word (and quoted phrase) occurs in its text |
|
|
154
|
+
| `passage` | the matching snippet (at `max`, the best full-text chunk), or `None` |
|
|
155
|
+
| `fulltext` | at `max`: whether its full text was read; `None` otherwise |
|
|
156
|
+
| `found_by` | the searches that found it: `index` (the cache), `s2`, `openalex`, `core`, `arxiv` |
|
|
157
|
+
|
|
158
|
+
A `SearchResult` is immutable: `authors` and `found_by` are tuples, so a result can go in a
|
|
159
|
+
set or key a dict. A result only CORE found has `found_by == ("core",)`.
|
|
160
|
+
|
|
161
|
+
### How long a search may run
|
|
162
|
+
|
|
163
|
+
- By default `timeout` is the level's server budget plus 30 s: 40 s at `low`, 60 s at
|
|
164
|
+
`medium`, 90 s at `high`, and 150 s at `extra`. `max` has no limit. Pass a number of
|
|
165
|
+
seconds to set one, or `None` for none.
|
|
166
|
+
- The clock starts once the server has taken the search. Time spent queued behind other
|
|
167
|
+
searches doesn't count: the clock restarts on every poll that finds the search still
|
|
168
|
+
queued.
|
|
169
|
+
- A resubmitted search (see Retries) gets the whole `timeout` again, and the pause before
|
|
170
|
+
it isn't charged to it.
|
|
171
|
+
- Each poll waits on the server for up to 25 s, and never past the time left.
|
|
172
|
+
- Past the limit, `search()` raises `SearchTimeout`.
|
|
173
|
+
|
|
174
|
+
### Results the cache doesn't hold
|
|
175
|
+
|
|
176
|
+
Results with `cached=False` can still be read: when `hit.id` isn't `None`, `fetch(hit.id)`
|
|
177
|
+
caches the paper now. Do it only for papers whose text you need now. The server caches
|
|
178
|
+
uncached results that clear its citation bar in the background, when its auto-cache is on,
|
|
179
|
+
unless only CORE found them: CORE isn't limited to the project's fields, so those wait for
|
|
180
|
+
a `fetch`.
|
|
181
|
+
|
|
182
|
+
## Papers, bookmarks, and the corpus
|
|
183
|
+
|
|
184
|
+
```python
|
|
185
|
+
from apollodorus_client import Apollodorus
|
|
186
|
+
|
|
187
|
+
with Apollodorus.from_env(owner="my-project") as ap:
|
|
188
|
+
hits = ap.search("volatility targeting", limit=5)
|
|
189
|
+
meta = ap.paper(hits[0].id) # dict: title, abstract, authors, status, text_source, citation_count, ...
|
|
190
|
+
body = ap.text(hits[0].id) # str: the full text (just the abstract for abstract-only papers)
|
|
191
|
+
|
|
192
|
+
uncached = next((hit for hit in hits if not hit.cached and hit.id), None)
|
|
193
|
+
if uncached is not None: # only when you need its text now
|
|
194
|
+
meta = ap.fetch(uncached.id) # slow (tens of seconds); the id may be "doi:..."
|
|
195
|
+
body = ap.text(meta["arxiv_id"]) # read it under the id the fetch returns
|
|
196
|
+
|
|
197
|
+
paper_id = meta["arxiv_id"]
|
|
198
|
+
ap.bookmark(paper_id, note="hedging section") # pins the paper against eviction
|
|
199
|
+
ap.bookmark(paper_id) # no note given: the note stays
|
|
200
|
+
ap.bookmark(paper_id, note=None) # clears the note
|
|
201
|
+
page = ap.bookmarks(limit=20, offset=0) # yours, newest first: {"results", "total", "limit", "offset"}
|
|
202
|
+
ap.unbookmark(paper_id) # {"removed": True, "published": ...}
|
|
203
|
+
|
|
204
|
+
print(ap.stats()) # corpus counts and footprint, from the nightly snapshot
|
|
205
|
+
print(ap.health()) # {"status": "ok"}
|
|
206
|
+
```
|
|
207
|
+
|
|
208
|
+
- `paper(id)`: one paper's metadata, whatever its `status` (`active`, `evicted`, or
|
|
209
|
+
`parse_failed`), including whether you bookmarked it.
|
|
210
|
+
- `text(id)`: a cached paper's full text.
|
|
211
|
+
- `fetch(id)`: caches a paper now, from any source, and returns its metadata. The id may be
|
|
212
|
+
a search result's `doi:…`. Use the `arxiv_id` it returns from then on: it can differ from
|
|
213
|
+
the id you sent (a version suffix dropped, a merged work, the same paper under another
|
|
214
|
+
source). Fetching a paper the cache already holds returns it as it is.
|
|
215
|
+
- `bookmark(id, note=...)`: bookmarks a paper for your owner, fetching it first if it isn't
|
|
216
|
+
cached. The answer's `published` is `False` while the change is still on its way to the
|
|
217
|
+
shared index: in flight, never lost.
|
|
218
|
+
- `unbookmark(id)`: removes it. Removing a bookmark that isn't there is not an error.
|
|
219
|
+
- `bookmarks(*, limit=20, offset=0)`: one page of your bookmarks.
|
|
220
|
+
- `stats()` and `health()`: the corpus's counts and footprint, and a liveness check.
|
|
221
|
+
|
|
222
|
+
`fetch` and `bookmark` allow 300 s, since they can download and parse a paper, and each of
|
|
223
|
+
`search()`'s polls allows its wait plus 30 s. Every other request uses the client's
|
|
224
|
+
`timeout`.
|
|
225
|
+
|
|
226
|
+
## Retries
|
|
227
|
+
|
|
228
|
+
- `search()` and the reads (`paper`, `text`, `bookmarks`, `stats`, `health`) retry a 429, a
|
|
229
|
+
502, 503, or 504, and a dropped connection: 5 attempts in all, about 1, 2, 4, then 8 s
|
|
230
|
+
apart. Each pause varies by up to 25% either way, so clients that failed together don't
|
|
231
|
+
come back together.
|
|
232
|
+
- A `Retry-After` header on a 429 or 503, in seconds or as an HTTP date, is waited for when
|
|
233
|
+
it asks for longer than that pause, for at most 30 s. The server sends `Retry-After: 5`
|
|
234
|
+
when its search queue is full or it is restarting.
|
|
235
|
+
- A search the server lost (a poll answered 404, after a restart, say) is submitted again at
|
|
236
|
+
once. One that failed, or that stalled (running, but silent for over 60 s), is submitted
|
|
237
|
+
again after a pause of about 1, 2, then 4 s. A search is resubmitted 3 times at most,
|
|
238
|
+
whatever the causes.
|
|
239
|
+
- The writes (`fetch`, `bookmark`, `unbookmark`) make one attempt, since a write whose answer
|
|
240
|
+
was lost may still have landed.
|
|
241
|
+
- A request this client built wrong (`httpx.LocalProtocolError`) is never retried.
|
|
242
|
+
|
|
243
|
+
## Errors
|
|
244
|
+
|
|
245
|
+
- `ApollodorusError` (`.status`, `.detail`): the API answered with an error, after any
|
|
246
|
+
retries. It is also raised when a search failed or stalled through 3 resubmits, or
|
|
247
|
+
reported a status this client doesn't know, and when the answer can't be the API's:
|
|
248
|
+
- A redirect is never followed, so the key can't be forwarded to another host. Use the
|
|
249
|
+
final `https://` address.
|
|
250
|
+
- A web page means `base_url` is likely the web UI's address: the API is under it at
|
|
251
|
+
`/api`.
|
|
252
|
+
- `SearchTimeout` (a `TimeoutError`): the search ran past `search()`'s `timeout`.
|
|
253
|
+
- `httpx.TransportError`: the server stayed unreachable through every attempt.
|
|
254
|
+
- `ValueError`: a bad argument, such as an unknown `effort`, a malformed `owner` or
|
|
255
|
+
`api_key`, or a `base_url` that isn't an absolute `http(s)://` URL with a host.
|
|
256
|
+
- `RuntimeError`: `from_env()` found no URL or no key.
|
|
257
|
+
|
|
258
|
+
Both exception types pickle and copy cleanly, so they cross process boundaries.
|
|
259
|
+
|
|
260
|
+
## Threads and lifetime
|
|
261
|
+
|
|
262
|
+
One client per process is enough: it pools connections. It keeps no per-call state, and the
|
|
263
|
+
`httpx.Client` it wraps can be shared between threads, so every thread can use the same
|
|
264
|
+
client. `search()` blocks the calling thread until the results arrive. Use the client as a
|
|
265
|
+
context manager, or call `close()` when you're done with it. Closing never closes an
|
|
266
|
+
`http_client` you passed in.
|
|
@@ -0,0 +1,244 @@
|
|
|
1
|
+
# apollodorus-client
|
|
2
|
+
|
|
3
|
+
Python client for the Apollodorus paper-search API: FluxGate's cache of top research
|
|
4
|
+
papers (quant finance, CS, statistics, economics), with search that reaches beyond it.
|
|
5
|
+
|
|
6
|
+
```bash
|
|
7
|
+
pip install apollodorus-client # add [dotenv] to read a .env file
|
|
8
|
+
```
|
|
9
|
+
|
|
10
|
+
```python
|
|
11
|
+
from apollodorus_client import Apollodorus
|
|
12
|
+
|
|
13
|
+
with Apollodorus.from_env() as ap: # APOLLODORUS_URL + APOLLODORUS_API_KEY
|
|
14
|
+
hits = ap.search("momentum crash", effort="medium")
|
|
15
|
+
for hit in hits:
|
|
16
|
+
print(f"{hit.score:8.3f} {hit.id} cached={hit.cached} {hit.title}")
|
|
17
|
+
print(hits.sources, "partial" if hits.partial else "complete")
|
|
18
|
+
```
|
|
19
|
+
|
|
20
|
+
Python 3.10 or later; the only dependency is `httpx`. The server also serves a guide for AI
|
|
21
|
+
agents at `{APOLLODORUS_URL}/llms.txt` and its OpenAPI schema at
|
|
22
|
+
`{APOLLODORUS_URL}/openapi.json`.
|
|
23
|
+
|
|
24
|
+
## Configuration
|
|
25
|
+
|
|
26
|
+
`Apollodorus.from_env(*, owner=None, timeout=60.0)` reads three variables:
|
|
27
|
+
|
|
28
|
+
| Variable | Meaning |
|
|
29
|
+
|---|---|
|
|
30
|
+
| `APOLLODORUS_URL` | the API root, e.g. `https://<your-host>/apollodorus/api` (required) |
|
|
31
|
+
| `APOLLODORUS_API_KEY` | the API key, sent as the `X-API-Key` header (required) |
|
|
32
|
+
| `APOLLODORUS_BOOKMARK_OWNER` | optional: the name your bookmarks are kept under (`owner` overrides it) |
|
|
33
|
+
|
|
34
|
+
With the `[dotenv]` extra installed, a `.env` file in the working directory, or in a
|
|
35
|
+
directory above it, supplies what the environment lacks; a real environment variable wins
|
|
36
|
+
over the file. Only those three names are read, from either. `from_env()` raises
|
|
37
|
+
`RuntimeError` when the URL or the key is missing.
|
|
38
|
+
|
|
39
|
+
Or build the client yourself:
|
|
40
|
+
|
|
41
|
+
```python
|
|
42
|
+
import os
|
|
43
|
+
|
|
44
|
+
from apollodorus_client import Apollodorus
|
|
45
|
+
|
|
46
|
+
ap = Apollodorus(
|
|
47
|
+
"https://<your-host>/apollodorus/api",
|
|
48
|
+
os.environ["APOLLODORUS_API_KEY"],
|
|
49
|
+
owner="my-project", # optional: sent as X-Bookmark-Owner
|
|
50
|
+
timeout=60.0, # seconds per request; see "Papers, bookmarks, and the corpus" for the exceptions
|
|
51
|
+
)
|
|
52
|
+
```
|
|
53
|
+
|
|
54
|
+
- `base_url` must be an absolute `http://` or `https://` URL with a host.
|
|
55
|
+
- `api_key` must be visible ASCII: no spaces, control characters, or curly quotes.
|
|
56
|
+
Whitespace around it, such as a trailing newline from a secrets file, is dropped. A
|
|
57
|
+
refused key is never repeated in the error.
|
|
58
|
+
- `owner` is 1 to 64 lowercase letters, digits, `.`, `_`, or `-`, starting with a letter or
|
|
59
|
+
digit. Without one, the server files your bookmarks under the shared owner `api-key`.
|
|
60
|
+
- `http_client` (optional) is an `httpx.Client` to send through. It is used as it is: never
|
|
61
|
+
modified, never closed, and never given the key or the base URL, which go only on this
|
|
62
|
+
client's own requests. Its own timeout then applies in place of `timeout`.
|
|
63
|
+
|
|
64
|
+
Keep the key in the environment or a `.env` file, never in code or URLs, and use an
|
|
65
|
+
`https://` URL. A plain `http://` `base_url` to any host but `localhost`, `127.0.0.0/8`, or
|
|
66
|
+
`::1` makes the constructor issue a `UserWarning`, because the key would travel in
|
|
67
|
+
cleartext.
|
|
68
|
+
|
|
69
|
+
## Effort levels
|
|
70
|
+
|
|
71
|
+
| `effort` | Matches by | Searches | Server time budget |
|
|
72
|
+
|---|---|---|---|
|
|
73
|
+
| `low` (default) | keywords | the cache: titles, abstracts, full text | 10 s |
|
|
74
|
+
| `medium` | keywords | the cache + Semantic Scholar, OpenAlex, CORE, arXiv | 30 s |
|
|
75
|
+
| `high` | meaning | the cache + Semantic Scholar | 60 s |
|
|
76
|
+
| `extra` | meaning | the cache + Semantic Scholar, OpenAlex, CORE, arXiv | 120 s |
|
|
77
|
+
| `max` | meaning, over full text | `extra`, then reads the top `top_n` papers in full | none (minutes) |
|
|
78
|
+
|
|
79
|
+
Keyword levels need every word, in any order; `"quoted words"` must appear as a phrase.
|
|
80
|
+
Meaning levels find papers about the query even when they share none of its words. A level
|
|
81
|
+
that runs out of time answers with what it found, and `hits.partial` is `True`. Start at
|
|
82
|
+
`low`, and go up when the results fall short. `EFFORTS` and `DEADLINES` (seconds, `None` for
|
|
83
|
+
`max`) are importable from the package.
|
|
84
|
+
|
|
85
|
+
## Searching
|
|
86
|
+
|
|
87
|
+
```python
|
|
88
|
+
from apollodorus_client import Apollodorus, JobStatus
|
|
89
|
+
|
|
90
|
+
|
|
91
|
+
def show(status: JobStatus) -> None: # called with each poll's answer
|
|
92
|
+
print(f"{status.status:<8} {status.stage} {status.elapsed_s:.0f} s")
|
|
93
|
+
|
|
94
|
+
|
|
95
|
+
with Apollodorus.from_env() as ap:
|
|
96
|
+
hits = ap.search(
|
|
97
|
+
"regime switching asset allocation",
|
|
98
|
+
effort="high",
|
|
99
|
+
limit=50, # 1-100 results, default 20
|
|
100
|
+
since=2015, until=2024, # publication years
|
|
101
|
+
field="q-fin", # a coarse field group: q-fin, econ, cs, stat, ...
|
|
102
|
+
timeout=120, # seconds the search may run; None waits as long as it takes
|
|
103
|
+
on_progress=show,
|
|
104
|
+
)
|
|
105
|
+
print(hits.job_id, hits.partial, hits.sources, hits.elapsed_s)
|
|
106
|
+
|
|
107
|
+
deep = ap.search("intraday momentum and liquidity", effort="max", top_n=10) # minutes
|
|
108
|
+
for hit in deep[:10]:
|
|
109
|
+
print(hit.fulltext, hit.match, (hit.passage or "")[:200])
|
|
110
|
+
```
|
|
111
|
+
|
|
112
|
+
`search(query, *, effort="low", limit=20, since=None, until=None, field=None, top_n=20,
|
|
113
|
+
timeout=..., on_progress=None)` submits the search as a job on the server and long-polls it
|
|
114
|
+
until it is done. It returns `SearchResults`: a list of `SearchResult`, best first, that also
|
|
115
|
+
carries `.partial`, `.sources` (each source's `ok`, `failed`, `skipped`, or `timeout`),
|
|
116
|
+
`.job_id`, and `.elapsed_s`. `top_n` (1–50) is read at `max` only: the papers whose full text
|
|
117
|
+
is read. Out-of-range values (a `limit` of 500, say) are refused by the server, as an
|
|
118
|
+
`ApollodorusError` with status 422. `on_progress` receives a `JobStatus` (`id`, `status`,
|
|
119
|
+
`stage`, `elapsed_s`, `heartbeat_age_s`, `partial`, `sources`) for every answer the server
|
|
120
|
+
gives.
|
|
121
|
+
|
|
122
|
+
Each `SearchResult` has:
|
|
123
|
+
|
|
124
|
+
| Field | Meaning |
|
|
125
|
+
|---|---|
|
|
126
|
+
| `id` | what `fetch()` takes: a paper id (`2301.12345`, `ssrn-…`, `journal-W…`), `doi:…` for a paper known only by its DOI, or `None` when no source we can fetch from knows it |
|
|
127
|
+
| `title`, `authors`, `year`, `source`, `citation_count`, `url` | the paper |
|
|
128
|
+
| `cached` | whether the cache holds it |
|
|
129
|
+
| `score` | BM25 at the keyword levels, cosine similarity otherwise: compare it only within one search |
|
|
130
|
+
| `match` | `keyword`, `meaning`, or `fulltext` |
|
|
131
|
+
| `exact` | every query word (and quoted phrase) occurs in its text |
|
|
132
|
+
| `passage` | the matching snippet (at `max`, the best full-text chunk), or `None` |
|
|
133
|
+
| `fulltext` | at `max`: whether its full text was read; `None` otherwise |
|
|
134
|
+
| `found_by` | the searches that found it: `index` (the cache), `s2`, `openalex`, `core`, `arxiv` |
|
|
135
|
+
|
|
136
|
+
A `SearchResult` is immutable: `authors` and `found_by` are tuples, so a result can go in a
|
|
137
|
+
set or key a dict. A result only CORE found has `found_by == ("core",)`.
|
|
138
|
+
|
|
139
|
+
### How long a search may run
|
|
140
|
+
|
|
141
|
+
- By default `timeout` is the level's server budget plus 30 s: 40 s at `low`, 60 s at
|
|
142
|
+
`medium`, 90 s at `high`, and 150 s at `extra`. `max` has no limit. Pass a number of
|
|
143
|
+
seconds to set one, or `None` for none.
|
|
144
|
+
- The clock starts once the server has taken the search. Time spent queued behind other
|
|
145
|
+
searches doesn't count: the clock restarts on every poll that finds the search still
|
|
146
|
+
queued.
|
|
147
|
+
- A resubmitted search (see Retries) gets the whole `timeout` again, and the pause before
|
|
148
|
+
it isn't charged to it.
|
|
149
|
+
- Each poll waits on the server for up to 25 s, and never past the time left.
|
|
150
|
+
- Past the limit, `search()` raises `SearchTimeout`.
|
|
151
|
+
|
|
152
|
+
### Results the cache doesn't hold
|
|
153
|
+
|
|
154
|
+
Results with `cached=False` can still be read: when `hit.id` isn't `None`, `fetch(hit.id)`
|
|
155
|
+
caches the paper now. Do it only for papers whose text you need now. The server caches
|
|
156
|
+
uncached results that clear its citation bar in the background, when its auto-cache is on,
|
|
157
|
+
unless only CORE found them: CORE isn't limited to the project's fields, so those wait for
|
|
158
|
+
a `fetch`.
|
|
159
|
+
|
|
160
|
+
## Papers, bookmarks, and the corpus
|
|
161
|
+
|
|
162
|
+
```python
|
|
163
|
+
from apollodorus_client import Apollodorus
|
|
164
|
+
|
|
165
|
+
with Apollodorus.from_env(owner="my-project") as ap:
|
|
166
|
+
hits = ap.search("volatility targeting", limit=5)
|
|
167
|
+
meta = ap.paper(hits[0].id) # dict: title, abstract, authors, status, text_source, citation_count, ...
|
|
168
|
+
body = ap.text(hits[0].id) # str: the full text (just the abstract for abstract-only papers)
|
|
169
|
+
|
|
170
|
+
uncached = next((hit for hit in hits if not hit.cached and hit.id), None)
|
|
171
|
+
if uncached is not None: # only when you need its text now
|
|
172
|
+
meta = ap.fetch(uncached.id) # slow (tens of seconds); the id may be "doi:..."
|
|
173
|
+
body = ap.text(meta["arxiv_id"]) # read it under the id the fetch returns
|
|
174
|
+
|
|
175
|
+
paper_id = meta["arxiv_id"]
|
|
176
|
+
ap.bookmark(paper_id, note="hedging section") # pins the paper against eviction
|
|
177
|
+
ap.bookmark(paper_id) # no note given: the note stays
|
|
178
|
+
ap.bookmark(paper_id, note=None) # clears the note
|
|
179
|
+
page = ap.bookmarks(limit=20, offset=0) # yours, newest first: {"results", "total", "limit", "offset"}
|
|
180
|
+
ap.unbookmark(paper_id) # {"removed": True, "published": ...}
|
|
181
|
+
|
|
182
|
+
print(ap.stats()) # corpus counts and footprint, from the nightly snapshot
|
|
183
|
+
print(ap.health()) # {"status": "ok"}
|
|
184
|
+
```
|
|
185
|
+
|
|
186
|
+
- `paper(id)`: one paper's metadata, whatever its `status` (`active`, `evicted`, or
|
|
187
|
+
`parse_failed`), including whether you bookmarked it.
|
|
188
|
+
- `text(id)`: a cached paper's full text.
|
|
189
|
+
- `fetch(id)`: caches a paper now, from any source, and returns its metadata. The id may be
|
|
190
|
+
a search result's `doi:…`. Use the `arxiv_id` it returns from then on: it can differ from
|
|
191
|
+
the id you sent (a version suffix dropped, a merged work, the same paper under another
|
|
192
|
+
source). Fetching a paper the cache already holds returns it as it is.
|
|
193
|
+
- `bookmark(id, note=...)`: bookmarks a paper for your owner, fetching it first if it isn't
|
|
194
|
+
cached. The answer's `published` is `False` while the change is still on its way to the
|
|
195
|
+
shared index: in flight, never lost.
|
|
196
|
+
- `unbookmark(id)`: removes it. Removing a bookmark that isn't there is not an error.
|
|
197
|
+
- `bookmarks(*, limit=20, offset=0)`: one page of your bookmarks.
|
|
198
|
+
- `stats()` and `health()`: the corpus's counts and footprint, and a liveness check.
|
|
199
|
+
|
|
200
|
+
`fetch` and `bookmark` allow 300 s, since they can download and parse a paper, and each of
|
|
201
|
+
`search()`'s polls allows its wait plus 30 s. Every other request uses the client's
|
|
202
|
+
`timeout`.
|
|
203
|
+
|
|
204
|
+
## Retries
|
|
205
|
+
|
|
206
|
+
- `search()` and the reads (`paper`, `text`, `bookmarks`, `stats`, `health`) retry a 429, a
|
|
207
|
+
502, 503, or 504, and a dropped connection: 5 attempts in all, about 1, 2, 4, then 8 s
|
|
208
|
+
apart. Each pause varies by up to 25% either way, so clients that failed together don't
|
|
209
|
+
come back together.
|
|
210
|
+
- A `Retry-After` header on a 429 or 503, in seconds or as an HTTP date, is waited for when
|
|
211
|
+
it asks for longer than that pause, for at most 30 s. The server sends `Retry-After: 5`
|
|
212
|
+
when its search queue is full or it is restarting.
|
|
213
|
+
- A search the server lost (a poll answered 404, after a restart, say) is submitted again at
|
|
214
|
+
once. One that failed, or that stalled (running, but silent for over 60 s), is submitted
|
|
215
|
+
again after a pause of about 1, 2, then 4 s. A search is resubmitted 3 times at most,
|
|
216
|
+
whatever the causes.
|
|
217
|
+
- The writes (`fetch`, `bookmark`, `unbookmark`) make one attempt, since a write whose answer
|
|
218
|
+
was lost may still have landed.
|
|
219
|
+
- A request this client built wrong (`httpx.LocalProtocolError`) is never retried.
|
|
220
|
+
|
|
221
|
+
## Errors
|
|
222
|
+
|
|
223
|
+
- `ApollodorusError` (`.status`, `.detail`): the API answered with an error, after any
|
|
224
|
+
retries. It is also raised when a search failed or stalled through 3 resubmits, or
|
|
225
|
+
reported a status this client doesn't know, and when the answer can't be the API's:
|
|
226
|
+
- A redirect is never followed, so the key can't be forwarded to another host. Use the
|
|
227
|
+
final `https://` address.
|
|
228
|
+
- A web page means `base_url` is likely the web UI's address: the API is under it at
|
|
229
|
+
`/api`.
|
|
230
|
+
- `SearchTimeout` (a `TimeoutError`): the search ran past `search()`'s `timeout`.
|
|
231
|
+
- `httpx.TransportError`: the server stayed unreachable through every attempt.
|
|
232
|
+
- `ValueError`: a bad argument, such as an unknown `effort`, a malformed `owner` or
|
|
233
|
+
`api_key`, or a `base_url` that isn't an absolute `http(s)://` URL with a host.
|
|
234
|
+
- `RuntimeError`: `from_env()` found no URL or no key.
|
|
235
|
+
|
|
236
|
+
Both exception types pickle and copy cleanly, so they cross process boundaries.
|
|
237
|
+
|
|
238
|
+
## Threads and lifetime
|
|
239
|
+
|
|
240
|
+
One client per process is enough: it pools connections. It keeps no per-call state, and the
|
|
241
|
+
`httpx.Client` it wraps can be shared between threads, so every thread can use the same
|
|
242
|
+
client. `search()` blocks the calling thread until the results arrive. Use the client as a
|
|
243
|
+
context manager, or call `close()` when you're done with it. Closing never closes an
|
|
244
|
+
`http_client` you passed in.
|