arxiv-dl 0.1.1 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +32 -0
- data/CODE_OF_CONDUCT.md +80 -7
- data/README.md +46 -23
- data/lib/arxiv/downloader/abstract_page.rb +1 -1
- data/lib/arxiv/downloader/archive.rb +31 -13
- data/lib/arxiv/downloader/assets_cache.rb +2 -1
- data/lib/arxiv/downloader/author.rb +11 -0
- data/lib/arxiv/downloader/bibtex.rb +5 -6
- data/lib/arxiv/downloader/categories.rb +7 -2
- data/lib/arxiv/downloader/cli.rb +32 -11
- data/lib/arxiv/downloader/client.rb +41 -2
- data/lib/arxiv/downloader/feed_parser/atom_entry.rb +14 -0
- data/lib/arxiv/downloader/feed_parser/atom_feed.rb +16 -0
- data/lib/arxiv/downloader/feed_parser/author_element.rb +12 -0
- data/lib/arxiv/downloader/feed_parser.rb +19 -30
- data/lib/arxiv/downloader/html_archive.rb +19 -4
- data/lib/arxiv/downloader/http_error.rb +14 -0
- data/lib/arxiv/downloader/identifier.rb +6 -0
- data/lib/arxiv/downloader/metadata/json.rb +1 -0
- data/lib/arxiv/downloader/metadata/markdown.rb +8 -1
- data/lib/arxiv/downloader/metadata/yaml.rb +1 -0
- data/lib/arxiv/downloader/metadata.rb +1 -0
- data/lib/arxiv/downloader/paper_not_found.rb +9 -0
- data/lib/arxiv/downloader/pdf.rb +1 -1
- data/lib/arxiv/downloader/source_archive.rb +41 -8
- data/lib/arxiv/downloader/version.rb +1 -1
- data/lib/arxiv/downloader.rb +18 -15
- metadata +10 -3
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 36af277147beae3a55221ba50443204db5b9e4bdae1613a3d9e923a465d4bc31
|
|
4
|
+
data.tar.gz: 04d405743856249d2f78d4f9e711e73b716ddb27f1b6633ad92a96e6385677be
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 92527994ad88aec89a64f3f678f60ff8a9d286d3189d06a6eafa4a9ba4f5a02fa815e9d23c0e429c0ea44ad03d3d58ae870c79a001c0705a682fd22818210412
|
|
7
|
+
data.tar.gz: 5058f511ec0e2707275605bee72d29d44c190e563e9b83da65e8b8619199186345c78231b35aa685fe183d6c0dd6f6996115059fd72c991c8d7babb39dfe7be8
|
data/CHANGELOG.md
CHANGED
|
@@ -1,3 +1,35 @@
|
|
|
1
|
+
## [0.2.0]
|
|
2
|
+
|
|
3
|
+
Versioned archives, author affiliations, and a round of robustness fixes. Two breaking changes to the output layout and metadata; see below.
|
|
4
|
+
|
|
5
|
+
### Breaking
|
|
6
|
+
|
|
7
|
+
- Versioned layout. Each archived version goes in its own `v<N>/` folder under the paper folder, with versioned filenames (`2508.16190v1.pdf`, `2508.16190v1-abstract.html`, `html/2508.16190v1.html`). A versioned ID (`2508.16190v1`) now archives that version instead of silently downloading the latest; an unversioned ID archives the latest.
|
|
8
|
+
- Author affiliations. `Metadata#authors` is now a list of `Arxiv::Downloader::Author` (`name`, `affiliations`) instead of name strings. `metadata.yaml`, `metadata.json`, and the `metadata.md` frontmatter write `authors` as a list of `{ name, affiliations }`.
|
|
9
|
+
- Minimum Ruby version is now 4.0.7.
|
|
10
|
+
|
|
11
|
+
### Added
|
|
12
|
+
|
|
13
|
+
- Author affiliations (`arxiv:affiliation`) are parsed from the Atom feed. The `metadata.md` body lists them after each author name: `- Jon S. Lawrence (Australian Astronomical Observatory; Macquarie University)`. BibTeX output is unchanged: names only.
|
|
14
|
+
- `Metadata#version` records which paper version was archived; the sidecars include it as `version`.
|
|
15
|
+
- Re-running skips versions already archived. Each version downloads into `v<N>.partial/` and is renamed to `v<N>/` only when complete, so a failed run never leaves a folder that looks finished.
|
|
16
|
+
- CLI: `-i FILE` / `--input FILE` reads targets one per line (`-` for stdin; blank lines and `#` comments skipped), combined with any argument targets.
|
|
17
|
+
- 429 and 503 responses are retried up to 3 times, waiting for `Retry-After` seconds when arxiv sends it, otherwise backing off 10s, 20s, 40s. Retries are logged with `-v`.
|
|
18
|
+
- HTTP requests time out (10s connect, 10s write, 60s per read) instead of hanging forever on a stalled connection.
|
|
19
|
+
|
|
20
|
+
### Fixed
|
|
21
|
+
|
|
22
|
+
- `Client#get` raises `Arxiv::Downloader::HTTPError` (with `status` and `url`) on non-success responses, instead of returning error bodies that were written to disk as PDFs/HTML or crashed the Atom parser.
|
|
23
|
+
- CLI: a failing target is reported on stderr as `<target>: <message>` and the remaining targets still download. Exit status is `1` if any target failed. Previously the first failure aborted the batch with a stack trace.
|
|
24
|
+
- An ID arxiv has no paper for raises `Arxiv::Downloader::PaperNotFound` instead of `NoMethodError`.
|
|
25
|
+
- Papers with no HTML version (404) skip the `html/` archive instead of saving arxiv's 404 page. A missing HTML asset (404) leaves its reference untouched instead of failing the paper.
|
|
26
|
+
- Source downloads handle more than gzipped tarballs: a single gzipped file is written under its original name, a PDF-only submission's source is skipped (the PDF is already archived), and unrecognized formats are kept as raw bytes in `src/<id>`.
|
|
27
|
+
- Legacy IDs (`cs/0002001`) no longer fail to write files: the `/` is replaced with `-` in file names, as it already was in the paper folder name.
|
|
28
|
+
|
|
29
|
+
### Security
|
|
30
|
+
|
|
31
|
+
- Source tarball entries that resolve outside `src/` (`../` or absolute paths) are skipped instead of written.
|
|
32
|
+
|
|
1
33
|
## [0.1.1]
|
|
2
34
|
|
|
3
35
|
Bug fix: the HTML archive step crashed with `NoMethodError` on papers whose HTML embeds `data:` URI images (e.g. arxiv's feedback-overlay mascot), and mis-fetched root-relative `/static/...` asset references from a wrong page-relative URL.
|
data/CODE_OF_CONDUCT.md
CHANGED
|
@@ -1,10 +1,83 @@
|
|
|
1
|
-
Code of Conduct
|
|
1
|
+
# Contributor Covenant 3.0 Code of Conduct
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
## Our Pledge
|
|
4
4
|
|
|
5
|
-
|
|
6
|
-
- Participants must ensure that their language and actions are free of personal attacks and disparaging personal remarks.
|
|
7
|
-
- When interpreting the words and actions of others, participants should always assume good intentions.
|
|
8
|
-
- Behaviour which can be reasonably considered harassment will not be tolerated.
|
|
5
|
+
We pledge to make our community welcoming, safe, and equitable for all.
|
|
9
6
|
|
|
10
|
-
|
|
7
|
+
We are committed to fostering an environment that respects and promotes the dignity, rights, and contributions of all individuals, regardless of characteristics including race, ethnicity, caste, color, age, physical characteristics, neurodiversity, disability, sex or gender, gender identity or expression, sexual orientation, language, philosophy or religion, national or social origin, socio-economic position, level of education, or other status. The same privileges of participation are extended to everyone who participates in good faith and in accordance with this Covenant.
|
|
8
|
+
|
|
9
|
+
## Encouraged Behaviors
|
|
10
|
+
|
|
11
|
+
While acknowledging differences in social norms, we all strive to meet our community's expectations for positive behavior. We also understand that our words and actions may be interpreted differently than we intend based on culture, background, or native language.
|
|
12
|
+
|
|
13
|
+
With these considerations in mind, we agree to behave mindfully toward each other and act in ways that center our shared values, including:
|
|
14
|
+
|
|
15
|
+
1. Respecting the **purpose of our community**, our activities, and our ways of gathering.
|
|
16
|
+
2. Engaging **kindly and honestly** with others.
|
|
17
|
+
3. Respecting **different viewpoints** and experiences.
|
|
18
|
+
4. **Taking responsibility** for our actions and contributions.
|
|
19
|
+
5. Gracefully giving and accepting **constructive feedback**.
|
|
20
|
+
6. Committing to **repairing harm** when it occurs.
|
|
21
|
+
7. Behaving in other ways that promote and sustain the **well-being of our community**.
|
|
22
|
+
|
|
23
|
+
## Restricted Behaviors
|
|
24
|
+
|
|
25
|
+
We agree to restrict the following behaviors in our community. Instances, threats, and promotion of these behaviors are violations of this Code of Conduct.
|
|
26
|
+
|
|
27
|
+
1. **Harassment.** Violating explicitly expressed boundaries or engaging in unnecessary personal attention after any clear request to stop.
|
|
28
|
+
2. **Character attacks.** Making insulting, demeaning, or pejorative comments directed at a community member or group of people.
|
|
29
|
+
3. **Stereotyping or discrimination.** Characterizing anyone’s personality or behavior on the basis of immutable identities or traits.
|
|
30
|
+
4. **Sexualization.** Behaving in a way that would generally be considered inappropriately intimate in the context or purpose of the community.
|
|
31
|
+
5. **Violating confidentiality**. Sharing or acting on someone's personal or private information without their permission.
|
|
32
|
+
6. **Endangerment.** Causing, encouraging, or threatening violence or other harm toward any person or group.
|
|
33
|
+
7. Behaving in other ways that **threaten the well-being** of our community.
|
|
34
|
+
|
|
35
|
+
### Other Restrictions
|
|
36
|
+
|
|
37
|
+
1. **Misleading identity.** Impersonating someone else for any reason, or pretending to be someone else to evade enforcement actions.
|
|
38
|
+
2. **Failing to credit sources.** Not properly crediting the sources of content you contribute.
|
|
39
|
+
3. **Promotional materials**. Sharing marketing or other commercial content in a way that is outside the norms of the community.
|
|
40
|
+
4. **Irresponsible communication.** Failing to responsibly present content which includes, links or describes any other restricted behaviors.
|
|
41
|
+
|
|
42
|
+
## Reporting an Issue
|
|
43
|
+
|
|
44
|
+
Tensions can occur between community members even when they are trying their best to collaborate. Not every conflict represents a code of conduct violation, and this Code of Conduct reinforces encouraged behaviors and norms that can help avoid conflicts and minimize harm.
|
|
45
|
+
|
|
46
|
+
When an incident does occur, it is important to report it promptly. To report a possible violation, **email the maintainer at [veganstraightedge@gmail.com](mailto:veganstraightedge@gmail.com).**
|
|
47
|
+
|
|
48
|
+
Community Moderators take reports of violations seriously and will make every effort to respond in a timely manner. They will investigate all reports of code of conduct violations, reviewing messages, logs, and recordings, or interviewing witnesses and other participants. Community Moderators will keep investigation and enforcement actions as transparent as possible while prioritizing safety and confidentiality. In order to honor these values, enforcement actions are carried out in private with the involved parties, but communicating to the whole community may be part of a mutually agreed upon resolution.
|
|
49
|
+
|
|
50
|
+
## Addressing and Repairing Harm
|
|
51
|
+
|
|
52
|
+
If an investigation by the Community Moderators finds that this Code of Conduct has been violated, the following enforcement ladder may be used to determine how best to repair harm, based on the incident's impact on the individuals involved and the community as a whole. Depending on the severity of a violation, lower rungs on the ladder may be skipped.
|
|
53
|
+
|
|
54
|
+
1. Warning
|
|
55
|
+
1. Event: A violation involving a single incident or series of incidents.
|
|
56
|
+
2. Consequence: A private, written warning from the Community Moderators.
|
|
57
|
+
3. Repair: Examples of repair include a private written apology, acknowledgement of responsibility, and seeking clarification on expectations.
|
|
58
|
+
2. Temporarily Limited Activities
|
|
59
|
+
1. Event: A repeated incidence of a violation that previously resulted in a warning, or the first incidence of a more serious violation.
|
|
60
|
+
2. Consequence: A private, written warning with a time-limited cooldown period designed to underscore the seriousness of the situation and give the community members involved time to process the incident. The cooldown period may be limited to particular communication channels or interactions with particular community members.
|
|
61
|
+
3. Repair: Examples of repair may include making an apology, using the cooldown period to reflect on actions and impact, and being thoughtful about re-entering community spaces after the period is over.
|
|
62
|
+
3. Temporary Suspension
|
|
63
|
+
1. Event: A pattern of repeated violation which the Community Moderators have tried to address with warnings, or a single serious violation.
|
|
64
|
+
2. Consequence: A private written warning with conditions for return from suspension. In general, temporary suspensions give the person being suspended time to reflect upon their behavior and possible corrective actions.
|
|
65
|
+
3. Repair: Examples of repair include respecting the spirit of the suspension, meeting the specified conditions for return, and being thoughtful about how to reintegrate with the community when the suspension is lifted.
|
|
66
|
+
4. Permanent Ban
|
|
67
|
+
1. Event: A pattern of repeated code of conduct violations that other steps on the ladder have failed to resolve, or a violation so serious that the Community Moderators determine there is no way to keep the community safe with this person as a member.
|
|
68
|
+
2. Consequence: Access to all community spaces, tools, and communication channels is removed. In general, permanent bans should be rarely used, should have strong reasoning behind them, and should only be resorted to if working through other remedies has failed to change the behavior.
|
|
69
|
+
3. Repair: There is no possible repair in cases of this severity.
|
|
70
|
+
|
|
71
|
+
This enforcement ladder is intended as a guideline. It does not limit the ability of Community Managers to use their discretion and judgment, in keeping with the best interests of our community.
|
|
72
|
+
|
|
73
|
+
## Scope
|
|
74
|
+
|
|
75
|
+
This Code of Conduct applies within all community spaces, and also applies when an individual is officially representing the community in public or other spaces. Examples of representing our community include using an official email address, posting via an official social media account, or acting as an appointed representative at an online or offline event.
|
|
76
|
+
|
|
77
|
+
## Attribution
|
|
78
|
+
|
|
79
|
+
This Code of Conduct is adapted from the Contributor Covenant, version 3.0, permanently available at [https://www.contributor-covenant.org/version/3/0/](https://www.contributor-covenant.org/version/3/0/).
|
|
80
|
+
|
|
81
|
+
Contributor Covenant is stewarded by the Organization for Ethical Source and licensed under CC BY-SA 4.0. To view a copy of this license, visit [https://creativecommons.org/licenses/by-sa/4.0/](https://creativecommons.org/licenses/by-sa/4.0/)
|
|
82
|
+
|
|
83
|
+
For answers to common questions about Contributor Covenant, see the FAQ at [https://www.contributor-covenant.org/faq](https://www.contributor-covenant.org/faq). Translations are provided at [https://www.contributor-covenant.org/translations](https://www.contributor-covenant.org/translations). Additional enforcement and community guideline resources can be found at [https://www.contributor-covenant.org/resources](https://www.contributor-covenant.org/resources). The enforcement ladder was inspired by the work of [Mozilla’s code of conduct team](https://github.com/mozilla/inclusion).
|
data/README.md
CHANGED
|
@@ -42,14 +42,15 @@ Accepted input forms:
|
|
|
42
42
|
|
|
43
43
|
### Flags
|
|
44
44
|
|
|
45
|
-
| Flag
|
|
46
|
-
|
|
|
47
|
-
| `-
|
|
48
|
-
| `--
|
|
49
|
-
|
|
|
50
|
-
| `-
|
|
51
|
-
| `--
|
|
52
|
-
|
|
|
45
|
+
| Flag | Description |
|
|
46
|
+
| ------------------------- | ----------------------------------------------------------------------------- |
|
|
47
|
+
| `-i FILE`, `--input FILE` | Read IDs/URLs from FILE, one per line (`-` for stdin. blanks and `#` skipped) |
|
|
48
|
+
| `-p PATH`, `--path PATH` | Root download directory |
|
|
49
|
+
| `--rate-limit SECONDS` | Seconds between HTTP requests. `0` disables throttling |
|
|
50
|
+
| `-v`, `--verbose` | Print step lines and per-request URL/byte logs to stdout |
|
|
51
|
+
| `-q`, `--quiet` | Print nothing to stdout. errors still go to stderr |
|
|
52
|
+
| `--version` | Print the gem version and exit |
|
|
53
|
+
| `-h`, `--help` | Print help and exit |
|
|
53
54
|
|
|
54
55
|
`-v` and `-q` are mutually exclusive.
|
|
55
56
|
|
|
@@ -58,10 +59,14 @@ Accepted input forms:
|
|
|
58
59
|
| Variable | Effect |
|
|
59
60
|
| --------------------- | ------------------------------------------------------------------------------- |
|
|
60
61
|
| `ARXIV_DOWNLOAD_PATH` | Root download directory (default: `$HOME/Downloads/ArXiv_Papers`) |
|
|
61
|
-
| `ARXIV_RATE_LIMIT` | Seconds between HTTP requests (default: `3`, per arxiv etiquette
|
|
62
|
+
| `ARXIV_RATE_LIMIT` | Seconds between HTTP requests (default: `3`, per arxiv etiquette. `0` disables) |
|
|
62
63
|
|
|
63
64
|
Precedence: CLI flag > ENV var > default.
|
|
64
65
|
|
|
66
|
+
### Errors and exit status
|
|
67
|
+
|
|
68
|
+
A target that fails (unrecognized ID, no such paper, HTTP error, network failure) is reported on stderr as `<target>: <message>`, and the remaining targets still download. Exit status is `0` when every target succeeds and `1` when any fails.
|
|
69
|
+
|
|
65
70
|
### Examples
|
|
66
71
|
|
|
67
72
|
Download a single paper to the default location:
|
|
@@ -82,6 +87,13 @@ Download multiple papers, verbose:
|
|
|
82
87
|
arxiv-dl -v 2508.16190 1207.7214 cs/0002001
|
|
83
88
|
```
|
|
84
89
|
|
|
90
|
+
Download every paper listed in a file, or piped in:
|
|
91
|
+
|
|
92
|
+
```sh
|
|
93
|
+
arxiv-dl --input reading-list.txt
|
|
94
|
+
cat reading-list.txt | arxiv-dl --input -
|
|
95
|
+
```
|
|
96
|
+
|
|
85
97
|
Disable rate limiting (when running against a local mirror, etc):
|
|
86
98
|
|
|
87
99
|
```sh
|
|
@@ -96,22 +108,27 @@ $ARXIV_DOWNLOAD_PATH/ # default: $HOME/Downloads/ArXiv_Papers
|
|
|
96
108
|
arxiv.org/static/...
|
|
97
109
|
cdn.jsdelivr.net/...
|
|
98
110
|
YYYY/MM/DD/<primary_category>/<arxiv-id>-<slug>/
|
|
99
|
-
<
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
111
|
+
v<N>/ # one folder per archived version
|
|
112
|
+
<arxiv-id>v<N>.pdf
|
|
113
|
+
<arxiv-id>v<N>-abstract.html
|
|
114
|
+
metadata.md # YAML frontmatter + Markdown body
|
|
115
|
+
metadata.yaml
|
|
116
|
+
metadata.json
|
|
117
|
+
metadata.bib # upstream BibTeX, falls back to synthesized
|
|
118
|
+
html/ # absent when the paper has no HTML version
|
|
119
|
+
<arxiv-id>v<N>.html # path-rewritten to local assets
|
|
120
|
+
x1.png, x2.png, ... # paper-specific images
|
|
121
|
+
src/ # absent for PDF-only submissions
|
|
122
|
+
*.tex, *.bbl, ... # extracted from /src/<id>v<N>
|
|
110
123
|
```
|
|
111
124
|
|
|
125
|
+
An unversioned ID (`2508.16190`) archives the latest version. A versioned ID (`2508.16190v1`) archives that version. Different versions of the same paper sit side by side under the same paper folder.
|
|
126
|
+
|
|
127
|
+
Each version downloads into `v<N>.partial/` and is renamed to `v<N>/` only when every file succeeded. Re-running skips versions whose `v<N>/` already exists and retries interrupted ones from scratch.
|
|
128
|
+
|
|
112
129
|
`YYYY/MM/DD` is the original submission date. `<primary_category>` is from the paper's metadata (`cs.CL`, `math.NT`, etc). `<slug>` is derived from the paper title (Unicode → ASCII, hyphenated, truncated to 80 chars at a word boundary).
|
|
113
130
|
|
|
114
|
-
For legacy IDs containing `/` (e.g. `cs/0002001`), the slash is replaced with `-` in the directory name (`cs-0002001
|
|
131
|
+
For legacy IDs containing `/` (e.g. `cs/0002001`), the slash is replaced with `-` in the directory name and file names (`cs-0002001-.../v1/cs-0002001v1.pdf`).
|
|
115
132
|
|
|
116
133
|
## Library usage
|
|
117
134
|
|
|
@@ -121,7 +138,7 @@ require 'arxiv/downloader'
|
|
|
121
138
|
identifier = Arxiv::Downloader::Identifier.new '2508.16190'
|
|
122
139
|
client = Arxiv::Downloader::Client.new # 3-second rate limit by default
|
|
123
140
|
path = Arxiv::Downloader::Archive.new(identifier, root: '/tmp/papers', client: client).run
|
|
124
|
-
# => "/tmp/papers/2025/08/22/cs.CL/2508.16190-comicscene154-a-scene-dataset-for-comic-analysis"
|
|
141
|
+
# => "/tmp/papers/2025/08/22/cs.CL/2508.16190-comicscene154-a-scene-dataset-for-comic-analysis/v1"
|
|
125
142
|
```
|
|
126
143
|
|
|
127
144
|
## Development
|
|
@@ -132,10 +149,16 @@ script/test # run specs and rubocop
|
|
|
132
149
|
script/console # interactive prompt
|
|
133
150
|
```
|
|
134
151
|
|
|
152
|
+
Specs run offline against recorded fixtures in `spec/fixtures/http/`. To check those fixtures against the live arxiv API, run:
|
|
153
|
+
|
|
154
|
+
```sh
|
|
155
|
+
ARXIV_LIVE=1 script/test
|
|
156
|
+
```
|
|
157
|
+
|
|
135
158
|
## License
|
|
136
159
|
|
|
137
160
|
MIT — see [LICENSE.md](LICENSE.md).
|
|
138
161
|
|
|
139
162
|
## Code of Conduct
|
|
140
163
|
|
|
141
|
-
|
|
164
|
+
This project follows the [Contributor Covenant](https://www.contributor-covenant.org/version/3/0/) 3.0 — see [CODE_OF_CONDUCT.md](CODE_OF_CONDUCT.md).
|
|
@@ -9,8 +9,14 @@ module Arxiv
|
|
|
9
9
|
@client = client
|
|
10
10
|
end
|
|
11
11
|
|
|
12
|
+
# Downloads into v<N>.partial/ and
|
|
13
|
+
# renames it to v<N>/ only once everything succeeded.
|
|
14
|
+
# An existing v<N>/ is always complete and is skipped.
|
|
12
15
|
def run
|
|
13
|
-
|
|
16
|
+
return paper_dir if Dir.exist? paper_dir
|
|
17
|
+
|
|
18
|
+
FileUtils.rm_rf staging_dir
|
|
19
|
+
FileUtils.mkdir_p staging_dir
|
|
14
20
|
|
|
15
21
|
download_pdf
|
|
16
22
|
download_abstract
|
|
@@ -18,6 +24,7 @@ module Arxiv
|
|
|
18
24
|
download_source_archive
|
|
19
25
|
write_sidecars
|
|
20
26
|
|
|
27
|
+
File.rename staging_dir, paper_dir
|
|
21
28
|
paper_dir
|
|
22
29
|
end
|
|
23
30
|
|
|
@@ -27,30 +34,41 @@ module Arxiv
|
|
|
27
34
|
@metadata ||= FeedParser.new(@client.get(atom_url).to_s).metadata
|
|
28
35
|
end
|
|
29
36
|
|
|
37
|
+
# the requested version, or the latest when none was requested
|
|
30
38
|
def atom_url
|
|
31
|
-
"https://export.arxiv.org/api/query?id_list=#{@identifier
|
|
39
|
+
"https://export.arxiv.org/api/query?id_list=#{@identifier}"
|
|
40
|
+
end
|
|
41
|
+
|
|
42
|
+
# the version the API actually returned, so every download matches the metadata
|
|
43
|
+
def archived
|
|
44
|
+
@archived ||= Identifier.new "#{metadata.arxiv_id}v#{metadata.version}"
|
|
32
45
|
end
|
|
33
46
|
|
|
34
47
|
def paper_dir
|
|
35
|
-
@paper_dir ||= File.join @root, Path.new(metadata).to_s
|
|
48
|
+
@paper_dir ||= File.join @root, Path.new(metadata).to_s, "v#{metadata.version}"
|
|
49
|
+
end
|
|
50
|
+
|
|
51
|
+
# a sibling of paper_dir, so relative ../_shared/ refs survive the rename
|
|
52
|
+
def staging_dir
|
|
53
|
+
"#{paper_dir}.partial"
|
|
36
54
|
end
|
|
37
55
|
|
|
38
56
|
def download_pdf
|
|
39
|
-
PDF.new(
|
|
57
|
+
PDF.new(archived, client: @client).download to: File.join(staging_dir, "#{archived.file_stem}.pdf")
|
|
40
58
|
end
|
|
41
59
|
|
|
42
60
|
def download_abstract
|
|
43
|
-
AbstractPage.new(
|
|
44
|
-
.download to: File.join(
|
|
61
|
+
AbstractPage.new(archived, client: @client)
|
|
62
|
+
.download to: File.join(staging_dir, "#{archived.file_stem}-abstract.html")
|
|
45
63
|
end
|
|
46
64
|
|
|
47
65
|
def download_html_archive
|
|
48
|
-
HTMLArchive.new(
|
|
49
|
-
.download to: File.join(
|
|
66
|
+
HTMLArchive.new(archived, client: @client, assets_cache: assets_cache)
|
|
67
|
+
.download to: File.join(staging_dir, 'html')
|
|
50
68
|
end
|
|
51
69
|
|
|
52
70
|
def download_source_archive
|
|
53
|
-
SourceArchive.new(
|
|
71
|
+
SourceArchive.new(archived, client: @client).download to: File.join(staging_dir, 'src')
|
|
54
72
|
end
|
|
55
73
|
|
|
56
74
|
def assets_cache
|
|
@@ -58,10 +76,10 @@ module Arxiv
|
|
|
58
76
|
end
|
|
59
77
|
|
|
60
78
|
def write_sidecars
|
|
61
|
-
Metadata::Markdown.new(metadata).write to:
|
|
62
|
-
Metadata::YAML.new(metadata).write to:
|
|
63
|
-
Metadata::JSON.new(metadata).write to:
|
|
64
|
-
Metadata::Bibtex.new(metadata, client: @client).write to:
|
|
79
|
+
Metadata::Markdown.new(metadata).write to: staging_dir
|
|
80
|
+
Metadata::YAML.new(metadata).write to: staging_dir
|
|
81
|
+
Metadata::JSON.new(metadata).write to: staging_dir
|
|
82
|
+
Metadata::Bibtex.new(metadata, client: @client).write to: staging_dir
|
|
65
83
|
end
|
|
66
84
|
end
|
|
67
85
|
end
|
|
@@ -23,10 +23,9 @@ module Arxiv
|
|
|
23
23
|
def fetch
|
|
24
24
|
return nil if @client.nil?
|
|
25
25
|
|
|
26
|
-
|
|
27
|
-
|
|
28
|
-
|
|
29
|
-
response.to_s
|
|
26
|
+
@client.get(url).to_s
|
|
27
|
+
rescue HTTPError
|
|
28
|
+
nil
|
|
30
29
|
end
|
|
31
30
|
|
|
32
31
|
def to_s
|
|
@@ -34,7 +33,7 @@ module Arxiv
|
|
|
34
33
|
end
|
|
35
34
|
|
|
36
35
|
def key
|
|
37
|
-
last_name = @metadata.authors.first.split.last.downcase.gsub(/[^a-z]/, '')
|
|
36
|
+
last_name = @metadata.authors.first.name.split.last.downcase.gsub(/[^a-z]/, '')
|
|
38
37
|
year = @metadata.published.year
|
|
39
38
|
title_word = Slug.new(@metadata.title).to_s.split('-').first
|
|
40
39
|
|
|
@@ -48,7 +47,7 @@ module Arxiv
|
|
|
48
47
|
end
|
|
49
48
|
|
|
50
49
|
def authors
|
|
51
|
-
@metadata.authors.join ' and '
|
|
50
|
+
@metadata.authors.map(&:name).join ' and '
|
|
52
51
|
end
|
|
53
52
|
end
|
|
54
53
|
end
|
|
@@ -6,14 +6,19 @@ module Arxiv
|
|
|
6
6
|
DATA_PATH = File.expand_path('categories.yaml', __dir__).freeze
|
|
7
7
|
|
|
8
8
|
def initialize
|
|
9
|
-
@data = YAML.load_file
|
|
9
|
+
@data = YAML.load_file DATA_PATH
|
|
10
10
|
end
|
|
11
11
|
|
|
12
12
|
def lookup id
|
|
13
13
|
entry = @data[id]
|
|
14
14
|
return nil if entry.nil?
|
|
15
15
|
|
|
16
|
-
{
|
|
16
|
+
{
|
|
17
|
+
id: id,
|
|
18
|
+
name: entry['name'],
|
|
19
|
+
group: entry['group'],
|
|
20
|
+
description: entry['description']
|
|
21
|
+
}
|
|
17
22
|
end
|
|
18
23
|
end
|
|
19
24
|
end
|
data/lib/arxiv/downloader/cli.rb
CHANGED
|
@@ -6,10 +6,11 @@ module Arxiv
|
|
|
6
6
|
DEFAULT_DOWNLOAD_PATH = File.join Dir.home, 'Downloads', 'ArXiv_Papers'
|
|
7
7
|
USAGE = 'Usage: arxiv-dl [options] <ARXIV_ID_OR_URL> [<ARXIV_ID_OR_URL>...]'.freeze
|
|
8
8
|
|
|
9
|
-
def initialize argv,
|
|
9
|
+
def initialize argv, stderr: $stderr, stdin: $stdin, stdout: $stdout
|
|
10
10
|
@argv = argv
|
|
11
|
-
@stdout = stdout
|
|
12
11
|
@stderr = stderr
|
|
12
|
+
@stdin = stdin
|
|
13
|
+
@stdout = stdout
|
|
13
14
|
end
|
|
14
15
|
|
|
15
16
|
def run
|
|
@@ -19,8 +20,8 @@ module Arxiv
|
|
|
19
20
|
return error_with USAGE if options[:targets].empty?
|
|
20
21
|
return error_with conflict if options[:verbose] && options[:quiet]
|
|
21
22
|
|
|
22
|
-
download_each options
|
|
23
|
-
0
|
|
23
|
+
failures = download_each options
|
|
24
|
+
failures.zero? ? 0 : 1
|
|
24
25
|
end
|
|
25
26
|
|
|
26
27
|
private
|
|
@@ -35,12 +36,12 @@ module Arxiv
|
|
|
35
36
|
|
|
36
37
|
begin
|
|
37
38
|
parser.parse! @argv
|
|
38
|
-
|
|
39
|
+
options[:targets] = @argv + input_targets(options[:input])
|
|
40
|
+
rescue OptionParser::ParseError, SystemCallError => e
|
|
39
41
|
@stderr.puts e.message
|
|
40
42
|
return { exit_status: 1 }
|
|
41
43
|
end
|
|
42
44
|
|
|
43
|
-
options[:targets] = @argv
|
|
44
45
|
options[:path] ||= ENV['ARXIV_DOWNLOAD_PATH'] || DEFAULT_DOWNLOAD_PATH
|
|
45
46
|
options[:rate_limit] ||= (ENV['ARXIV_RATE_LIMIT'] || Client::DEFAULT_RATE_LIMIT).to_i
|
|
46
47
|
options
|
|
@@ -49,6 +50,7 @@ module Arxiv
|
|
|
49
50
|
def build_parser options
|
|
50
51
|
OptionParser.new do |parser|
|
|
51
52
|
parser.banner = USAGE
|
|
53
|
+
parser.on('-i FILE', '--input FILE') { |value| options[:input] = value }
|
|
52
54
|
parser.on('-p PATH', '--path PATH') { |value| options[:path] = value }
|
|
53
55
|
parser.on('--rate-limit SECONDS', Integer) { |value| options[:rate_limit] = value }
|
|
54
56
|
parser.on('-v', '--verbose') { options[:verbose] = true }
|
|
@@ -64,6 +66,14 @@ module Arxiv
|
|
|
64
66
|
end
|
|
65
67
|
end
|
|
66
68
|
|
|
69
|
+
# one target per line from FILE, or stdin for "-"; blank lines and # comments skipped
|
|
70
|
+
def input_targets input
|
|
71
|
+
return [] if input.nil?
|
|
72
|
+
|
|
73
|
+
text = input == '-' ? @stdin.read : File.read(input)
|
|
74
|
+
text.lines.map(&:strip).reject { it.empty? || it.start_with?('#') }
|
|
75
|
+
end
|
|
76
|
+
|
|
67
77
|
def error_with message
|
|
68
78
|
@stderr.puts message
|
|
69
79
|
1
|
|
@@ -72,13 +82,24 @@ module Arxiv
|
|
|
72
82
|
def download_each options
|
|
73
83
|
client = Client.new rate_limit: options[:rate_limit], log: (options[:verbose] ? @stdout : nil)
|
|
74
84
|
|
|
85
|
+
failures = 0
|
|
75
86
|
options[:targets].each do |target|
|
|
76
|
-
|
|
77
|
-
@stdout.puts "==> Downloading #{identifier.id}" if options[:verbose]
|
|
78
|
-
|
|
79
|
-
path = Archive.new(identifier, root: options[:path], client: client).run
|
|
80
|
-
@stdout.puts path unless options[:quiet]
|
|
87
|
+
failures += 1 unless download_one(target, client:, options:)
|
|
81
88
|
end
|
|
89
|
+
failures
|
|
90
|
+
end
|
|
91
|
+
|
|
92
|
+
# true on success; reports the failure and returns false otherwise
|
|
93
|
+
def download_one target, client:, options:
|
|
94
|
+
identifier = Identifier.new target
|
|
95
|
+
@stdout.puts "==> Downloading #{identifier.id}" if options[:verbose]
|
|
96
|
+
|
|
97
|
+
path = Archive.new(identifier, root: options[:path], client: client).run
|
|
98
|
+
@stdout.puts path unless options[:quiet]
|
|
99
|
+
true
|
|
100
|
+
rescue Error, HTTP::Error => e
|
|
101
|
+
@stderr.puts "#{target}: #{e.message}"
|
|
102
|
+
false
|
|
82
103
|
end
|
|
83
104
|
end
|
|
84
105
|
end
|
|
@@ -5,6 +5,10 @@ module Arxiv
|
|
|
5
5
|
class Client
|
|
6
6
|
SOURCE_URL = 'https://github.com/xoengineering/arxiv-dl'.freeze
|
|
7
7
|
DEFAULT_RATE_LIMIT = 3
|
|
8
|
+
TIMEOUTS = { connect: 10, read: 60, write: 10 }.freeze # seconds, per operation
|
|
9
|
+
MAX_RETRIES = 3
|
|
10
|
+
RETRY_BACKOFF = 10 # seconds before the first retry. doubles on each retry.
|
|
11
|
+
RETRYABLE_STATUSES = [429, 503].freeze
|
|
8
12
|
|
|
9
13
|
attr_reader :rate_limit
|
|
10
14
|
|
|
@@ -18,14 +22,49 @@ module Arxiv
|
|
|
18
22
|
end
|
|
19
23
|
|
|
20
24
|
def get url
|
|
25
|
+
retries = 0
|
|
26
|
+
|
|
27
|
+
loop do
|
|
28
|
+
response = request url
|
|
29
|
+
return response if response.status.success?
|
|
30
|
+
raise http_error(url, response) unless retryable? response, retries
|
|
31
|
+
|
|
32
|
+
retries += 1
|
|
33
|
+
wait_before_retry response, retries
|
|
34
|
+
end
|
|
35
|
+
end
|
|
36
|
+
|
|
37
|
+
private
|
|
38
|
+
|
|
39
|
+
def request url
|
|
21
40
|
throttle
|
|
22
|
-
response = HTTP.headers('User-Agent' => user_agent).follow.get(url)
|
|
41
|
+
response = HTTP.timeout(TIMEOUTS).headers('User-Agent' => user_agent).follow.get(url)
|
|
23
42
|
@last_request_at = Time.now
|
|
24
43
|
log_request url, response
|
|
25
44
|
response
|
|
26
45
|
end
|
|
27
46
|
|
|
28
|
-
|
|
47
|
+
def http_error url, response
|
|
48
|
+
HTTPError.new status: response.status.code, url: url, reason: response.status.reason
|
|
49
|
+
end
|
|
50
|
+
|
|
51
|
+
def retryable? response, retries
|
|
52
|
+
RETRYABLE_STATUSES.include?(response.status.code) && retries < MAX_RETRIES
|
|
53
|
+
end
|
|
54
|
+
|
|
55
|
+
def wait_before_retry response, retries
|
|
56
|
+
seconds = retry_after(response) || (RETRY_BACKOFF * (2**(retries - 1)))
|
|
57
|
+
@log&.puts "==> #{response.status}. Retrying in #{seconds}s"
|
|
58
|
+
sleep seconds
|
|
59
|
+
end
|
|
60
|
+
|
|
61
|
+
# Retry-After in delay-seconds form. The HTTP-date form falls back to backoff.
|
|
62
|
+
def retry_after response
|
|
63
|
+
value = response.headers['Retry-After']
|
|
64
|
+
return if value.nil?
|
|
65
|
+
|
|
66
|
+
Integer(value, exception: false)
|
|
67
|
+
end
|
|
29
68
|
|
|
30
69
|
def log_request url, response
|
|
31
70
|
return if @log.nil?
|
|
@@ -0,0 +1,14 @@
|
|
|
1
|
+
module Arxiv
|
|
2
|
+
module Downloader
|
|
3
|
+
class FeedParser
|
|
4
|
+
class AtomEntry < Feedjira::Parser::AtomEntry
|
|
5
|
+
elements :author, as: :authors, class: AuthorElement
|
|
6
|
+
|
|
7
|
+
element 'arxiv:primary_category', as: :primary_category_id, value: :term
|
|
8
|
+
element 'arxiv:comment', as: :comment
|
|
9
|
+
element 'arxiv:doi', as: :doi
|
|
10
|
+
element 'arxiv:journal_ref', as: :journal_ref
|
|
11
|
+
end
|
|
12
|
+
end
|
|
13
|
+
end
|
|
14
|
+
end
|
|
@@ -0,0 +1,16 @@
|
|
|
1
|
+
module Arxiv
|
|
2
|
+
module Downloader
|
|
3
|
+
class FeedParser
|
|
4
|
+
class AtomFeed
|
|
5
|
+
include SAXMachine
|
|
6
|
+
include Feedjira::FeedUtilities
|
|
7
|
+
|
|
8
|
+
elements :entry, as: :entries, class: AtomEntry
|
|
9
|
+
|
|
10
|
+
def self.able_to_parse? xml
|
|
11
|
+
xml.include? 'http://www.w3.org/2005/Atom'
|
|
12
|
+
end
|
|
13
|
+
end
|
|
14
|
+
end
|
|
15
|
+
end
|
|
16
|
+
end
|
|
@@ -1,34 +1,11 @@
|
|
|
1
|
-
require 'feedjira'
|
|
1
|
+
require 'feedjira' # before all: SAXMachine, Feedjira::Parser::AtomEntry
|
|
2
|
+
require_relative 'feed_parser/author_element' # before atom_entry: AtomEntry uses class: AuthorElement
|
|
3
|
+
require_relative 'feed_parser/atom_entry' # before atom_feed: AtomFeed uses class: AtomEntry
|
|
4
|
+
require_relative 'feed_parser/atom_feed' # after atom_entry
|
|
2
5
|
|
|
3
6
|
module Arxiv
|
|
4
7
|
module Downloader
|
|
5
8
|
class FeedParser
|
|
6
|
-
class Author
|
|
7
|
-
include SAXMachine
|
|
8
|
-
|
|
9
|
-
element :name
|
|
10
|
-
end
|
|
11
|
-
|
|
12
|
-
class AtomEntry < Feedjira::Parser::AtomEntry
|
|
13
|
-
elements :author, as: :authors, class: Author
|
|
14
|
-
|
|
15
|
-
element 'arxiv:primary_category', as: :primary_category_id, value: :term
|
|
16
|
-
element 'arxiv:comment', as: :comment
|
|
17
|
-
element 'arxiv:doi', as: :doi
|
|
18
|
-
element 'arxiv:journal_ref', as: :journal_ref
|
|
19
|
-
end
|
|
20
|
-
|
|
21
|
-
class AtomFeed
|
|
22
|
-
include SAXMachine
|
|
23
|
-
include Feedjira::FeedUtilities
|
|
24
|
-
|
|
25
|
-
elements :entry, as: :entries, class: AtomEntry
|
|
26
|
-
|
|
27
|
-
def self.able_to_parse? xml
|
|
28
|
-
xml.include? 'http://www.w3.org/2005/Atom'
|
|
29
|
-
end
|
|
30
|
-
end
|
|
31
|
-
|
|
32
9
|
Feedjira.configure { |config| config.parsers = [AtomFeed] + config.parsers }
|
|
33
10
|
|
|
34
11
|
def initialize xml
|
|
@@ -37,15 +14,19 @@ module Arxiv
|
|
|
37
14
|
end
|
|
38
15
|
|
|
39
16
|
def metadata
|
|
40
|
-
entry
|
|
41
|
-
|
|
17
|
+
entry = @feed.entries.first
|
|
18
|
+
raise PaperNotFound if entry.nil?
|
|
19
|
+
|
|
20
|
+
identifier = Identifier.new entry.entry_id
|
|
21
|
+
arxiv_id = identifier.id
|
|
42
22
|
|
|
43
23
|
Metadata.new(
|
|
44
24
|
arxiv_id: arxiv_id,
|
|
25
|
+
version: identifier.version,
|
|
45
26
|
arxiv_url: "https://arxiv.org/abs/#{arxiv_id}",
|
|
46
27
|
pdf_url: "https://arxiv.org/pdf/#{arxiv_id}.pdf",
|
|
47
28
|
title: entry.title,
|
|
48
|
-
authors: entry
|
|
29
|
+
authors: authors_of(entry),
|
|
49
30
|
abstract: entry.summary.strip,
|
|
50
31
|
published: entry.published.to_date,
|
|
51
32
|
updated: entry.updated.to_date,
|
|
@@ -56,6 +37,14 @@ module Arxiv
|
|
|
56
37
|
journal_ref: entry.journal_ref
|
|
57
38
|
)
|
|
58
39
|
end
|
|
40
|
+
|
|
41
|
+
private
|
|
42
|
+
|
|
43
|
+
def authors_of entry
|
|
44
|
+
entry.authors.map do |author|
|
|
45
|
+
Author.new name: author.name, affiliations: author.affiliations
|
|
46
|
+
end
|
|
47
|
+
end
|
|
59
48
|
end
|
|
60
49
|
end
|
|
61
50
|
end
|
|
@@ -18,20 +18,32 @@ module Arxiv
|
|
|
18
18
|
end
|
|
19
19
|
|
|
20
20
|
def download to:
|
|
21
|
+
html = page
|
|
22
|
+
return if html.nil?
|
|
23
|
+
|
|
21
24
|
FileUtils.mkdir_p to
|
|
22
25
|
|
|
23
|
-
document = Nokogiri::HTML
|
|
26
|
+
document = Nokogiri::HTML html
|
|
24
27
|
ASSET_SELECTORS.each do |selector, attribute|
|
|
25
28
|
document.css(selector).each { |node| process node, attribute, to }
|
|
26
29
|
end
|
|
27
30
|
|
|
28
|
-
File.write File.join(to, "#{@identifier.
|
|
31
|
+
File.write File.join(to, "#{@identifier.file_stem}.html"), document.to_html
|
|
29
32
|
end
|
|
30
33
|
|
|
31
34
|
private
|
|
32
35
|
|
|
36
|
+
# nil when the paper has no HTML version (older papers, or conversion failed)
|
|
37
|
+
def page
|
|
38
|
+
@client.get(html_url).to_s
|
|
39
|
+
rescue HTTPError => e
|
|
40
|
+
raise unless e.status == 404
|
|
41
|
+
|
|
42
|
+
nil
|
|
43
|
+
end
|
|
44
|
+
|
|
33
45
|
def html_url
|
|
34
|
-
"https://arxiv.org/html/#{@identifier
|
|
46
|
+
"https://arxiv.org/html/#{@identifier}"
|
|
35
47
|
end
|
|
36
48
|
|
|
37
49
|
def process node, attribute, html_dir
|
|
@@ -43,6 +55,9 @@ module Arxiv
|
|
|
43
55
|
when :remote then cache_remote node, attribute, reference, html_dir
|
|
44
56
|
when :root_relative then cache_remote node, attribute, "https://arxiv.org#{reference}", html_dir
|
|
45
57
|
end
|
|
58
|
+
rescue HTTPError => e
|
|
59
|
+
# a missing asset leaves its reference untouched rather than failing the paper
|
|
60
|
+
raise unless e.status == 404
|
|
46
61
|
end
|
|
47
62
|
|
|
48
63
|
# :skip covers refs that can't or shouldn't be fetched: data:/javascript:/
|
|
@@ -64,7 +79,7 @@ module Arxiv
|
|
|
64
79
|
end
|
|
65
80
|
|
|
66
81
|
def absolute_for reference
|
|
67
|
-
"https://arxiv.org/html/#{@identifier
|
|
82
|
+
"https://arxiv.org/html/#{@identifier}/#{reference}"
|
|
68
83
|
end
|
|
69
84
|
|
|
70
85
|
def cache_remote node, attribute, reference, html_dir
|
|
@@ -0,0 +1,14 @@
|
|
|
1
|
+
module Arxiv
|
|
2
|
+
module Downloader
|
|
3
|
+
class HTTPError < Error
|
|
4
|
+
attr_reader :status, :url
|
|
5
|
+
|
|
6
|
+
def initialize status:, url:, reason: nil
|
|
7
|
+
@status = status
|
|
8
|
+
@url = url
|
|
9
|
+
|
|
10
|
+
super("GET #{url} failed: #{[status, reason].compact.join ' '}")
|
|
11
|
+
end
|
|
12
|
+
end
|
|
13
|
+
end
|
|
14
|
+
end
|
|
@@ -13,6 +13,12 @@ module Arxiv
|
|
|
13
13
|
set_id_and_version
|
|
14
14
|
end
|
|
15
15
|
|
|
16
|
+
# the API/URL form: 2508.16190, or 2508.16190v2 when a version is known
|
|
17
|
+
def to_s = version.nil? ? id : "#{id}v#{version}"
|
|
18
|
+
|
|
19
|
+
# to_s as a single path segment: legacy IDs like cs/0002001 become cs-0002001
|
|
20
|
+
def file_stem = to_s.tr('/', '-')
|
|
21
|
+
|
|
16
22
|
private
|
|
17
23
|
|
|
18
24
|
# validations
|
|
@@ -34,6 +34,7 @@ module Arxiv
|
|
|
34
34
|
case object
|
|
35
35
|
when Hash then object.to_h { |key, value| [key.to_s, stringify(value)] }
|
|
36
36
|
when Array then object.map { |item| stringify item }
|
|
37
|
+
when Data then stringify object.to_h
|
|
37
38
|
else object
|
|
38
39
|
end
|
|
39
40
|
end
|
|
@@ -57,7 +58,13 @@ module Arxiv
|
|
|
57
58
|
end
|
|
58
59
|
|
|
59
60
|
def authors_list
|
|
60
|
-
@metadata.authors.map { |author| "- #{author}" }.join "\n"
|
|
61
|
+
@metadata.authors.map { |author| "- #{author_line author}" }.join "\n"
|
|
62
|
+
end
|
|
63
|
+
|
|
64
|
+
def author_line author
|
|
65
|
+
return author.name if author.affiliations.empty?
|
|
66
|
+
|
|
67
|
+
"#{author.name} (#{author.affiliations.join '; '})"
|
|
61
68
|
end
|
|
62
69
|
end
|
|
63
70
|
end
|
data/lib/arxiv/downloader/pdf.rb
CHANGED
|
@@ -6,30 +6,58 @@ require 'zlib'
|
|
|
6
6
|
module Arxiv
|
|
7
7
|
module Downloader
|
|
8
8
|
class SourceArchive
|
|
9
|
+
GZIP_MAGIC = "\x1F\x8B".b.freeze
|
|
10
|
+
PDF_MAGIC = '%PDF'.b.freeze
|
|
11
|
+
TAR_MAGIC = 'ustar'.b.freeze
|
|
12
|
+
|
|
9
13
|
def initialize identifier, client:
|
|
10
14
|
@identifier = identifier
|
|
11
15
|
@client = client
|
|
12
16
|
end
|
|
13
17
|
|
|
14
18
|
def download to:
|
|
19
|
+
body = @client.get(url).to_s.b
|
|
20
|
+
|
|
21
|
+
# PDF-only submission: the PDF is already archived alongside src/
|
|
22
|
+
return if body.start_with? PDF_MAGIC
|
|
23
|
+
|
|
15
24
|
FileUtils.mkdir_p to
|
|
25
|
+
return extract_gzip body, to if body.start_with? GZIP_MAGIC
|
|
16
26
|
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
Gem::Package::TarReader.new(gz) do |tar|
|
|
20
|
-
tar.each { |entry| extract entry, to }
|
|
21
|
-
end
|
|
22
|
-
end
|
|
27
|
+
# unrecognized format: keep the raw bytes rather than lose them
|
|
28
|
+
File.binwrite File.join(to, @identifier.file_stem), body
|
|
23
29
|
end
|
|
24
30
|
|
|
25
31
|
private
|
|
26
32
|
|
|
33
|
+
# arxiv serves either a gzipped tarball or a single gzipped file
|
|
34
|
+
def extract_gzip body, to
|
|
35
|
+
Zlib::GzipReader.wrap StringIO.new(body) do |gz|
|
|
36
|
+
contents = gz.read
|
|
37
|
+
next extract_tar contents, to if tar? contents
|
|
38
|
+
|
|
39
|
+
filename = File.basename(gz.orig_name || @identifier.file_stem)
|
|
40
|
+
File.binwrite File.join(to, filename), contents
|
|
41
|
+
end
|
|
42
|
+
end
|
|
43
|
+
|
|
27
44
|
def url
|
|
28
|
-
"https://arxiv.org/src/#{@identifier
|
|
45
|
+
"https://arxiv.org/src/#{@identifier}"
|
|
46
|
+
end
|
|
47
|
+
|
|
48
|
+
def tar? contents
|
|
49
|
+
contents.byteslice(257, TAR_MAGIC.bytesize) == TAR_MAGIC
|
|
50
|
+
end
|
|
51
|
+
|
|
52
|
+
def extract_tar contents, to
|
|
53
|
+
Gem::Package::TarReader.new StringIO.new(contents) do |tar|
|
|
54
|
+
tar.each { |entry| extract entry, to }
|
|
55
|
+
end
|
|
29
56
|
end
|
|
30
57
|
|
|
31
58
|
def extract entry, root
|
|
32
|
-
path = File.
|
|
59
|
+
path = File.expand_path entry.full_name, root
|
|
60
|
+
return unless inside? path, root
|
|
33
61
|
|
|
34
62
|
if entry.directory?
|
|
35
63
|
FileUtils.mkdir_p path
|
|
@@ -38,6 +66,11 @@ module Arxiv
|
|
|
38
66
|
File.binwrite path, entry.read
|
|
39
67
|
end
|
|
40
68
|
end
|
|
69
|
+
|
|
70
|
+
# guards against tar entries like ../escaped.txt or /etc/passwd
|
|
71
|
+
def inside? path, root
|
|
72
|
+
path.start_with? File.join(File.expand_path(root), '')
|
|
73
|
+
end
|
|
41
74
|
end
|
|
42
75
|
end
|
|
43
76
|
end
|
data/lib/arxiv/downloader.rb
CHANGED
|
@@ -1,24 +1,27 @@
|
|
|
1
|
-
require_relative 'downloader/
|
|
2
|
-
require_relative 'downloader/
|
|
3
|
-
require_relative 'downloader/
|
|
4
|
-
require_relative 'downloader/
|
|
1
|
+
require_relative 'downloader/abstract_page'
|
|
2
|
+
require_relative 'downloader/archive'
|
|
3
|
+
require_relative 'downloader/assets_cache'
|
|
4
|
+
require_relative 'downloader/author'
|
|
5
|
+
require_relative 'downloader/bibtex'
|
|
5
6
|
require_relative 'downloader/categories'
|
|
7
|
+
require_relative 'downloader/cli'
|
|
6
8
|
require_relative 'downloader/client'
|
|
7
|
-
require_relative 'downloader/
|
|
9
|
+
require_relative 'downloader/error' # before errors below that subclass Error
|
|
8
10
|
require_relative 'downloader/feed_parser'
|
|
11
|
+
require_relative 'downloader/html_archive'
|
|
12
|
+
require_relative 'downloader/http_error' # after error
|
|
13
|
+
require_relative 'downloader/identifier' # after error
|
|
14
|
+
require_relative 'downloader/metadata' # before metadata/*: they reopen class Metadata
|
|
15
|
+
require_relative 'downloader/metadata/bibtex' # after metadata
|
|
16
|
+
require_relative 'downloader/metadata/json' # after metadata
|
|
17
|
+
require_relative 'downloader/metadata/markdown' # after metadata
|
|
18
|
+
require_relative 'downloader/metadata/yaml' # after metadata
|
|
19
|
+
require_relative 'downloader/paper_not_found' # after error
|
|
9
20
|
require_relative 'downloader/path'
|
|
10
21
|
require_relative 'downloader/pdf'
|
|
11
|
-
require_relative 'downloader/
|
|
12
|
-
require_relative 'downloader/abstract_page'
|
|
22
|
+
require_relative 'downloader/slug'
|
|
13
23
|
require_relative 'downloader/source_archive'
|
|
14
|
-
require_relative 'downloader/
|
|
15
|
-
require_relative 'downloader/html_archive'
|
|
16
|
-
require_relative 'downloader/metadata/yaml'
|
|
17
|
-
require_relative 'downloader/metadata/json'
|
|
18
|
-
require_relative 'downloader/metadata/bibtex'
|
|
19
|
-
require_relative 'downloader/metadata/markdown'
|
|
20
|
-
require_relative 'downloader/archive'
|
|
21
|
-
require_relative 'downloader/cli'
|
|
24
|
+
require_relative 'downloader/version'
|
|
22
25
|
|
|
23
26
|
module Arxiv
|
|
24
27
|
module Downloader
|
metadata
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: arxiv-dl
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 0.
|
|
4
|
+
version: 0.2.0
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Shane Becker
|
|
@@ -97,6 +97,7 @@ files:
|
|
|
97
97
|
- lib/arxiv/downloader/abstract_page.rb
|
|
98
98
|
- lib/arxiv/downloader/archive.rb
|
|
99
99
|
- lib/arxiv/downloader/assets_cache.rb
|
|
100
|
+
- lib/arxiv/downloader/author.rb
|
|
100
101
|
- lib/arxiv/downloader/bibtex.rb
|
|
101
102
|
- lib/arxiv/downloader/categories.rb
|
|
102
103
|
- lib/arxiv/downloader/categories.yaml
|
|
@@ -104,13 +105,18 @@ files:
|
|
|
104
105
|
- lib/arxiv/downloader/client.rb
|
|
105
106
|
- lib/arxiv/downloader/error.rb
|
|
106
107
|
- lib/arxiv/downloader/feed_parser.rb
|
|
108
|
+
- lib/arxiv/downloader/feed_parser/atom_entry.rb
|
|
109
|
+
- lib/arxiv/downloader/feed_parser/atom_feed.rb
|
|
110
|
+
- lib/arxiv/downloader/feed_parser/author_element.rb
|
|
107
111
|
- lib/arxiv/downloader/html_archive.rb
|
|
112
|
+
- lib/arxiv/downloader/http_error.rb
|
|
108
113
|
- lib/arxiv/downloader/identifier.rb
|
|
109
114
|
- lib/arxiv/downloader/metadata.rb
|
|
110
115
|
- lib/arxiv/downloader/metadata/bibtex.rb
|
|
111
116
|
- lib/arxiv/downloader/metadata/json.rb
|
|
112
117
|
- lib/arxiv/downloader/metadata/markdown.rb
|
|
113
118
|
- lib/arxiv/downloader/metadata/yaml.rb
|
|
119
|
+
- lib/arxiv/downloader/paper_not_found.rb
|
|
114
120
|
- lib/arxiv/downloader/path.rb
|
|
115
121
|
- lib/arxiv/downloader/pdf.rb
|
|
116
122
|
- lib/arxiv/downloader/slug.rb
|
|
@@ -124,6 +130,7 @@ metadata:
|
|
|
124
130
|
allowed_push_host: https://rubygems.org
|
|
125
131
|
homepage_uri: https://github.com/xoengineering/arxiv-dl
|
|
126
132
|
source_code_uri: https://github.com/xoengineering/arxiv-dl
|
|
133
|
+
bug_tracker_uri: https://github.com/xoengineering/arxiv-dl/issues
|
|
127
134
|
changelog_uri: https://github.com/xoengineering/arxiv-dl/blob/main/CHANGELOG.md
|
|
128
135
|
rubygems_mfa_required: 'true'
|
|
129
136
|
rdoc_options: []
|
|
@@ -133,14 +140,14 @@ required_ruby_version: !ruby/object:Gem::Requirement
|
|
|
133
140
|
requirements:
|
|
134
141
|
- - ">="
|
|
135
142
|
- !ruby/object:Gem::Version
|
|
136
|
-
version: 4.0.
|
|
143
|
+
version: 4.0.7
|
|
137
144
|
required_rubygems_version: !ruby/object:Gem::Requirement
|
|
138
145
|
requirements:
|
|
139
146
|
- - ">="
|
|
140
147
|
- !ruby/object:Gem::Version
|
|
141
148
|
version: '0'
|
|
142
149
|
requirements: []
|
|
143
|
-
rubygems_version: 4.0.
|
|
150
|
+
rubygems_version: 4.0.21
|
|
144
151
|
specification_version: 4
|
|
145
152
|
summary: Download papers and metadata from arxiv.org for offline archives.
|
|
146
153
|
test_files: []
|