jstor-dl 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +7 -0
- data/CHANGELOG.md +9 -0
- data/CODE_OF_CONDUCT.md +83 -0
- data/LICENSE.md +21 -0
- data/README.md +121 -0
- data/exe/jstor-dl +5 -0
- data/lib/jstor/downloader/archive.rb +78 -0
- data/lib/jstor/downloader/author.rb +11 -0
- data/lib/jstor/downloader/bibtex.rb +57 -0
- data/lib/jstor/downloader/cli.rb +106 -0
- data/lib/jstor/downloader/client.rb +86 -0
- data/lib/jstor/downloader/error.rb +5 -0
- data/lib/jstor/downloader/http_error.rb +14 -0
- data/lib/jstor/downloader/identifier.rb +66 -0
- data/lib/jstor/downloader/item_file.rb +23 -0
- data/lib/jstor/downloader/item_not_found.rb +13 -0
- data/lib/jstor/downloader/metadata/bibtex.rb +20 -0
- data/lib/jstor/downloader/metadata/json.rb +33 -0
- data/lib/jstor/downloader/metadata/markdown.rb +69 -0
- data/lib/jstor/downloader/metadata/yaml.rb +32 -0
- data/lib/jstor/downloader/metadata.rb +21 -0
- data/lib/jstor/downloader/metadata_parser.rb +43 -0
- data/lib/jstor/downloader/path.rb +28 -0
- data/lib/jstor/downloader/slug.rb +44 -0
- data/lib/jstor/downloader/version.rb +5 -0
- data/lib/jstor/downloader.rb +24 -0
- data/sig/jstor/downloader.rbs +5 -0
- metadata +119 -0
checksums.yaml
ADDED
|
@@ -0,0 +1,7 @@
|
|
|
1
|
+
---
|
|
2
|
+
SHA256:
|
|
3
|
+
metadata.gz: cd0503a0a9da1a8298c012fab0dd19f866e907a87a58537f6f1b26174b3bb1ab
|
|
4
|
+
data.tar.gz: 20779888e79ac989cc10b14ac96877b9200641f7e4cc3fe679f252b4602b1570
|
|
5
|
+
SHA512:
|
|
6
|
+
metadata.gz: 6839ae6451e76bfe92b282b4829a7dcea3b117c103bf68d345ad8e91d7a922d56d30bd27007a3be8c0920ac219c67c0b04ef98afa0e0206e0e9e7e8b3ebdf758
|
|
7
|
+
data.tar.gz: d0a458946ce663a24638a25b4fbc23d8579352c101105cb371e764b086524fef22723080e93314512cdfcdd0a8fd061cfb929edaac067f1c06087e24cef0d392
|
data/CHANGELOG.md
ADDED
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
## [0.1.0]
|
|
2
|
+
|
|
3
|
+
First version. Per-article offline archive of JSTOR's public-domain Early Journal Content, fetched only from the Internet Archive's copy, never from jstor.org.
|
|
4
|
+
|
|
5
|
+
- Accepts JSTOR stable IDs, jstor.org URLs, `10.2307/…` DOIs and doi.org URLs, and archive.org `jstor-…` items and URLs.
|
|
6
|
+
- Saves the scanned PDF, the OCR plaintext, JSTOR's article metadata XML verbatim, and four sidecar metadata files (`metadata.md`, `metadata.yaml`, `metadata.json`, `metadata.bib`).
|
|
7
|
+
- Layout: `YYYY/MM/DD/<journal>/<jstor-id>-<slug>/`.
|
|
8
|
+
- Articles not in the Early Journal Content on archive.org raise `Jstor::Downloader::ItemNotFound`.
|
|
9
|
+
- Rate-limited HTTP client (3s default) with timeouts and retries on 429/503, staged downloads that skip already-archived articles, and a CLI with `--input FILE|-`, per-target error reporting, and exit status 1 on any failure. All copied from arxiv-dl.
|
data/CODE_OF_CONDUCT.md
ADDED
|
@@ -0,0 +1,83 @@
|
|
|
1
|
+
# Contributor Covenant 3.0 Code of Conduct
|
|
2
|
+
|
|
3
|
+
## Our Pledge
|
|
4
|
+
|
|
5
|
+
We pledge to make our community welcoming, safe, and equitable for all.
|
|
6
|
+
|
|
7
|
+
We are committed to fostering an environment that respects and promotes the dignity, rights, and contributions of all individuals, regardless of characteristics including race, ethnicity, caste, color, age, physical characteristics, neurodiversity, disability, sex or gender, gender identity or expression, sexual orientation, language, philosophy or religion, national or social origin, socio-economic position, level of education, or other status. The same privileges of participation are extended to everyone who participates in good faith and in accordance with this Covenant.
|
|
8
|
+
|
|
9
|
+
## Encouraged Behaviors
|
|
10
|
+
|
|
11
|
+
While acknowledging differences in social norms, we all strive to meet our community's expectations for positive behavior. We also understand that our words and actions may be interpreted differently than we intend based on culture, background, or native language.
|
|
12
|
+
|
|
13
|
+
With these considerations in mind, we agree to behave mindfully toward each other and act in ways that center our shared values, including:
|
|
14
|
+
|
|
15
|
+
1. Respecting the **purpose of our community**, our activities, and our ways of gathering.
|
|
16
|
+
2. Engaging **kindly and honestly** with others.
|
|
17
|
+
3. Respecting **different viewpoints** and experiences.
|
|
18
|
+
4. **Taking responsibility** for our actions and contributions.
|
|
19
|
+
5. Gracefully giving and accepting **constructive feedback**.
|
|
20
|
+
6. Committing to **repairing harm** when it occurs.
|
|
21
|
+
7. Behaving in other ways that promote and sustain the **well-being of our community**.
|
|
22
|
+
|
|
23
|
+
## Restricted Behaviors
|
|
24
|
+
|
|
25
|
+
We agree to restrict the following behaviors in our community. Instances, threats, and promotion of these behaviors are violations of this Code of Conduct.
|
|
26
|
+
|
|
27
|
+
1. **Harassment.** Violating explicitly expressed boundaries or engaging in unnecessary personal attention after any clear request to stop.
|
|
28
|
+
2. **Character attacks.** Making insulting, demeaning, or pejorative comments directed at a community member or group of people.
|
|
29
|
+
3. **Stereotyping or discrimination.** Characterizing anyone’s personality or behavior on the basis of immutable identities or traits.
|
|
30
|
+
4. **Sexualization.** Behaving in a way that would generally be considered inappropriately intimate in the context or purpose of the community.
|
|
31
|
+
5. **Violating confidentiality**. Sharing or acting on someone's personal or private information without their permission.
|
|
32
|
+
6. **Endangerment.** Causing, encouraging, or threatening violence or other harm toward any person or group.
|
|
33
|
+
7. Behaving in other ways that **threaten the well-being** of our community.
|
|
34
|
+
|
|
35
|
+
### Other Restrictions
|
|
36
|
+
|
|
37
|
+
1. **Misleading identity.** Impersonating someone else for any reason, or pretending to be someone else to evade enforcement actions.
|
|
38
|
+
2. **Failing to credit sources.** Not properly crediting the sources of content you contribute.
|
|
39
|
+
3. **Promotional materials**. Sharing marketing or other commercial content in a way that is outside the norms of the community.
|
|
40
|
+
4. **Irresponsible communication.** Failing to responsibly present content which includes, links or describes any other restricted behaviors.
|
|
41
|
+
|
|
42
|
+
## Reporting an Issue
|
|
43
|
+
|
|
44
|
+
Tensions can occur between community members even when they are trying their best to collaborate. Not every conflict represents a code of conduct violation, and this Code of Conduct reinforces encouraged behaviors and norms that can help avoid conflicts and minimize harm.
|
|
45
|
+
|
|
46
|
+
When an incident does occur, it is important to report it promptly. To report a possible violation, **email the maintainer at [veganstraightedge@gmail.com](mailto:veganstraightedge@gmail.com).**
|
|
47
|
+
|
|
48
|
+
Community Moderators take reports of violations seriously and will make every effort to respond in a timely manner. They will investigate all reports of code of conduct violations, reviewing messages, logs, and recordings, or interviewing witnesses and other participants. Community Moderators will keep investigation and enforcement actions as transparent as possible while prioritizing safety and confidentiality. In order to honor these values, enforcement actions are carried out in private with the involved parties, but communicating to the whole community may be part of a mutually agreed upon resolution.
|
|
49
|
+
|
|
50
|
+
## Addressing and Repairing Harm
|
|
51
|
+
|
|
52
|
+
If an investigation by the Community Moderators finds that this Code of Conduct has been violated, the following enforcement ladder may be used to determine how best to repair harm, based on the incident's impact on the individuals involved and the community as a whole. Depending on the severity of a violation, lower rungs on the ladder may be skipped.
|
|
53
|
+
|
|
54
|
+
1. Warning
|
|
55
|
+
1. Event: A violation involving a single incident or series of incidents.
|
|
56
|
+
2. Consequence: A private, written warning from the Community Moderators.
|
|
57
|
+
3. Repair: Examples of repair include a private written apology, acknowledgement of responsibility, and seeking clarification on expectations.
|
|
58
|
+
2. Temporarily Limited Activities
|
|
59
|
+
1. Event: A repeated incidence of a violation that previously resulted in a warning, or the first incidence of a more serious violation.
|
|
60
|
+
2. Consequence: A private, written warning with a time-limited cooldown period designed to underscore the seriousness of the situation and give the community members involved time to process the incident. The cooldown period may be limited to particular communication channels or interactions with particular community members.
|
|
61
|
+
3. Repair: Examples of repair may include making an apology, using the cooldown period to reflect on actions and impact, and being thoughtful about re-entering community spaces after the period is over.
|
|
62
|
+
3. Temporary Suspension
|
|
63
|
+
1. Event: A pattern of repeated violation which the Community Moderators have tried to address with warnings, or a single serious violation.
|
|
64
|
+
2. Consequence: A private written warning with conditions for return from suspension. In general, temporary suspensions give the person being suspended time to reflect upon their behavior and possible corrective actions.
|
|
65
|
+
3. Repair: Examples of repair include respecting the spirit of the suspension, meeting the specified conditions for return, and being thoughtful about how to reintegrate with the community when the suspension is lifted.
|
|
66
|
+
4. Permanent Ban
|
|
67
|
+
1. Event: A pattern of repeated code of conduct violations that other steps on the ladder have failed to resolve, or a violation so serious that the Community Moderators determine there is no way to keep the community safe with this person as a member.
|
|
68
|
+
2. Consequence: Access to all community spaces, tools, and communication channels is removed. In general, permanent bans should be rarely used, should have strong reasoning behind them, and should only be resorted to if working through other remedies has failed to change the behavior.
|
|
69
|
+
3. Repair: There is no possible repair in cases of this severity.
|
|
70
|
+
|
|
71
|
+
This enforcement ladder is intended as a guideline. It does not limit the ability of Community Managers to use their discretion and judgment, in keeping with the best interests of our community.
|
|
72
|
+
|
|
73
|
+
## Scope
|
|
74
|
+
|
|
75
|
+
This Code of Conduct applies within all community spaces, and also applies when an individual is officially representing the community in public or other spaces. Examples of representing our community include using an official email address, posting via an official social media account, or acting as an appointed representative at an online or offline event.
|
|
76
|
+
|
|
77
|
+
## Attribution
|
|
78
|
+
|
|
79
|
+
This Code of Conduct is adapted from the Contributor Covenant, version 3.0, permanently available at [https://www.contributor-covenant.org/version/3/0/](https://www.contributor-covenant.org/version/3/0/).
|
|
80
|
+
|
|
81
|
+
Contributor Covenant is stewarded by the Organization for Ethical Source and licensed under CC BY-SA 4.0. To view a copy of this license, visit [https://creativecommons.org/licenses/by-sa/4.0/](https://creativecommons.org/licenses/by-sa/4.0/)
|
|
82
|
+
|
|
83
|
+
For answers to common questions about Contributor Covenant, see the FAQ at [https://www.contributor-covenant.org/faq](https://www.contributor-covenant.org/faq). Translations are provided at [https://www.contributor-covenant.org/translations](https://www.contributor-covenant.org/translations). Additional enforcement and community guideline resources can be found at [https://www.contributor-covenant.org/resources](https://www.contributor-covenant.org/resources). The enforcement ladder was inspired by the work of [Mozilla’s code of conduct team](https://github.com/mozilla/inclusion).
|
data/LICENSE.md
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
The MIT License (MIT)
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Shane Becker
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in
|
|
13
|
+
all copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN
|
|
21
|
+
THE SOFTWARE.
|
data/README.md
ADDED
|
@@ -0,0 +1,121 @@
|
|
|
1
|
+
# jstor-dl
|
|
2
|
+
|
|
3
|
+
Download articles from JSTOR's public-domain Early Journal Content for offline archives.
|
|
4
|
+
|
|
5
|
+
For each article, `jstor-dl` saves:
|
|
6
|
+
|
|
7
|
+
- The scanned PDF
|
|
8
|
+
- The OCR plaintext
|
|
9
|
+
- JSTOR's own article metadata (XML), verbatim
|
|
10
|
+
- Four sidecar metadata files: `metadata.md`, `metadata.yaml`, `metadata.json`, `metadata.bib`
|
|
11
|
+
|
|
12
|
+
## Where the content comes from
|
|
13
|
+
|
|
14
|
+
`jstor-dl` fetches only from the [Internet Archive's copy](https://archive.org/details/jstor_ejc) of JSTOR's Early Journal Content, never from jstor.org. [JSTOR's terms](https://about.jstor.org/terms) forbid any tool from downloading from jstor.org, even a single article; they allow only manual downloading.
|
|
15
|
+
|
|
16
|
+
The Early Journal Content is nearly 500,000 public-domain articles from 200+ journals (published before 1923 in the US, before 1870 elsewhere), released by JSTOR for free non-commercial use with acknowledgement, and uploaded to the Internet Archive in 2013 for bulk harvesting. Articles JSTOR added to its Early Journal Content after 2013 may be missing from the Internet Archive's copy.
|
|
17
|
+
|
|
18
|
+
## Installation
|
|
19
|
+
|
|
20
|
+
```sh
|
|
21
|
+
gem install jstor-dl
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
## CLI usage
|
|
25
|
+
|
|
26
|
+
```sh
|
|
27
|
+
jstor-dl <JSTOR_ID_OR_URL> [<JSTOR_ID_OR_URL>...]
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
Accepted input forms:
|
|
31
|
+
|
|
32
|
+
| Form | Example |
|
|
33
|
+
| ----------------------- | -------------------------------------------------------- |
|
|
34
|
+
| Stable ID | `4385670` |
|
|
35
|
+
| JSTOR URL | `https://www.jstor.org/stable/4385670` |
|
|
36
|
+
| JSTOR URL, DOI form | `https://www.jstor.org/stable/10.2307/4385670` |
|
|
37
|
+
| JSTOR PDF URL | `https://www.jstor.org/stable/pdf/4385670.pdf` |
|
|
38
|
+
| DOI | `10.2307/4385670`, `doi:10.2307/4385670` |
|
|
39
|
+
| DOI URL | `https://doi.org/10.2307/4385670` |
|
|
40
|
+
| Internet Archive item | `jstor-4385670` |
|
|
41
|
+
| Internet Archive URL | `https://archive.org/details/jstor-4385670` |
|
|
42
|
+
|
|
43
|
+
JSTOR URLs are only read for the article's ID; nothing is requested from jstor.org.
|
|
44
|
+
|
|
45
|
+
### Flags
|
|
46
|
+
|
|
47
|
+
| Flag | Description |
|
|
48
|
+
| ------------------------- | ----------------------------------------------------------------------------- |
|
|
49
|
+
| `-i FILE`, `--input FILE` | Read IDs/URLs from FILE, one per line (`-` for stdin; blanks and `#` skipped) |
|
|
50
|
+
| `-p PATH`, `--path PATH` | Root download directory |
|
|
51
|
+
| `--rate-limit SECONDS` | Seconds between HTTP requests; `0` disables throttling |
|
|
52
|
+
| `-v`, `--verbose` | Print step lines and per-request URL/byte logs to stdout |
|
|
53
|
+
| `-q`, `--quiet` | Print nothing to stdout; errors still go to stderr |
|
|
54
|
+
| `--version` | Print the gem version and exit |
|
|
55
|
+
| `-h`, `--help` | Print help and exit |
|
|
56
|
+
|
|
57
|
+
`-v` and `-q` are mutually exclusive.
|
|
58
|
+
|
|
59
|
+
### Environment variables
|
|
60
|
+
|
|
61
|
+
| Variable | Effect |
|
|
62
|
+
| --------------------- | ----------------------------------------------------------------- |
|
|
63
|
+
| `JSTOR_DOWNLOAD_PATH` | Root download directory (default: `$HOME/Downloads/JSTOR_Papers`) |
|
|
64
|
+
| `JSTOR_RATE_LIMIT` | Seconds between HTTP requests (default: `3`; `0` disables) |
|
|
65
|
+
|
|
66
|
+
Precedence: CLI flag > ENV var > default.
|
|
67
|
+
|
|
68
|
+
### Errors and exit status
|
|
69
|
+
|
|
70
|
+
A target that fails (unrecognized ID, not in the Early Journal Content on archive.org, HTTP error, network failure) is reported on stderr as `<target>: <message>`, and the remaining targets still download. Exit status is `0` when every target succeeds and `1` when any fails.
|
|
71
|
+
|
|
72
|
+
## Output layout
|
|
73
|
+
|
|
74
|
+
```txt
|
|
75
|
+
$JSTOR_DOWNLOAD_PATH/ # default: $HOME/Downloads/JSTOR_Papers
|
|
76
|
+
YYYY/MM/DD/<journal>/<jstor-id>-<slug>/
|
|
77
|
+
<jstor-id>.pdf # scanned article
|
|
78
|
+
<jstor-id>.txt # OCR plaintext
|
|
79
|
+
jstor.xml # JSTOR's article metadata, verbatim
|
|
80
|
+
metadata.md # YAML frontmatter + Markdown body
|
|
81
|
+
metadata.yaml
|
|
82
|
+
metadata.json
|
|
83
|
+
metadata.bib # synthesized from the metadata
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
`YYYY/MM/DD` is the publication date (shorter when only the year or month is known). `<journal>` is JSTOR's journal abbreviation (`clasweek` for The Classical Weekly). `<slug>` is derived from the article title.
|
|
87
|
+
|
|
88
|
+
Each article downloads into a sibling `.partial` folder and is renamed into place only when every file succeeded. Re-running skips articles already archived.
|
|
89
|
+
|
|
90
|
+
## Library usage
|
|
91
|
+
|
|
92
|
+
```ruby
|
|
93
|
+
require 'jstor/downloader'
|
|
94
|
+
|
|
95
|
+
identifier = Jstor::Downloader::Identifier.new 'https://www.jstor.org/stable/4385670'
|
|
96
|
+
client = Jstor::Downloader::Client.new # 3-second rate limit by default
|
|
97
|
+
path = Jstor::Downloader::Archive.new(identifier, root: '/tmp/papers', client: client).run
|
|
98
|
+
# => "/tmp/papers/1907/10/05/clasweek/4385670-the-elements-of-the-translation-of-latin"
|
|
99
|
+
```
|
|
100
|
+
|
|
101
|
+
## Development
|
|
102
|
+
|
|
103
|
+
```sh
|
|
104
|
+
script/setup # install dependencies
|
|
105
|
+
script/test # run specs and rubocop
|
|
106
|
+
script/console # interactive prompt
|
|
107
|
+
```
|
|
108
|
+
|
|
109
|
+
Specs run offline against recorded fixtures in `spec/fixtures/http/`. To check those fixtures against the live archive.org API, run:
|
|
110
|
+
|
|
111
|
+
```sh
|
|
112
|
+
ARCHIVE_LIVE=1 script/test
|
|
113
|
+
```
|
|
114
|
+
|
|
115
|
+
## License
|
|
116
|
+
|
|
117
|
+
MIT — see [LICENSE.md](LICENSE.md).
|
|
118
|
+
|
|
119
|
+
## Code of Conduct
|
|
120
|
+
|
|
121
|
+
This project follows the [Contributor Covenant](https://www.contributor-covenant.org/version/3/0/) 3.0 — see [CODE_OF_CONDUCT.md](CODE_OF_CONDUCT.md).
|
data/exe/jstor-dl
ADDED
|
@@ -0,0 +1,78 @@
|
|
|
1
|
+
require 'fileutils'
|
|
2
|
+
|
|
3
|
+
module Jstor
|
|
4
|
+
module Downloader
|
|
5
|
+
class Archive
|
|
6
|
+
def initialize identifier, root:, client: Client.new
|
|
7
|
+
@identifier = identifier
|
|
8
|
+
@root = root
|
|
9
|
+
@client = client
|
|
10
|
+
end
|
|
11
|
+
|
|
12
|
+
# Downloads into <article>.partial/ and
|
|
13
|
+
# renames it into place only once everything succeeded.
|
|
14
|
+
# An existing article folder is always complete and is skipped.
|
|
15
|
+
def run
|
|
16
|
+
return article_dir if Dir.exist? article_dir
|
|
17
|
+
|
|
18
|
+
FileUtils.rm_rf staging_dir
|
|
19
|
+
FileUtils.mkdir_p staging_dir
|
|
20
|
+
|
|
21
|
+
download_pdf
|
|
22
|
+
download_text
|
|
23
|
+
download_jstor_xml
|
|
24
|
+
write_sidecars
|
|
25
|
+
|
|
26
|
+
File.rename staging_dir, article_dir
|
|
27
|
+
article_dir
|
|
28
|
+
end
|
|
29
|
+
|
|
30
|
+
private
|
|
31
|
+
|
|
32
|
+
def metadata
|
|
33
|
+
@metadata ||= MetadataParser.new(@client.get(metadata_url).to_s).metadata
|
|
34
|
+
end
|
|
35
|
+
|
|
36
|
+
def metadata_url
|
|
37
|
+
"https://archive.org/metadata/#{@identifier.item}"
|
|
38
|
+
end
|
|
39
|
+
|
|
40
|
+
def article_dir
|
|
41
|
+
@article_dir ||= File.join @root, Path.new(metadata).to_s
|
|
42
|
+
end
|
|
43
|
+
|
|
44
|
+
# a sibling of article_dir, so the rename is a single atomic step
|
|
45
|
+
def staging_dir
|
|
46
|
+
"#{article_dir}.partial"
|
|
47
|
+
end
|
|
48
|
+
|
|
49
|
+
def download_pdf
|
|
50
|
+
ItemFile.new(@identifier, "#{@identifier}.pdf", client: @client)
|
|
51
|
+
.download to: File.join(staging_dir, "#{@identifier}.pdf")
|
|
52
|
+
end
|
|
53
|
+
|
|
54
|
+
# the OCR plaintext, when the item has it
|
|
55
|
+
def download_text
|
|
56
|
+
download_optional "#{@identifier}_djvu.txt", to: File.join(staging_dir, "#{@identifier}.txt")
|
|
57
|
+
end
|
|
58
|
+
|
|
59
|
+
# JSTOR's own article metadata, verbatim, when the item has it
|
|
60
|
+
def download_jstor_xml
|
|
61
|
+
download_optional "10.2307_#{@identifier}.xml", to: File.join(staging_dir, 'jstor.xml')
|
|
62
|
+
end
|
|
63
|
+
|
|
64
|
+
def download_optional name, to:
|
|
65
|
+
ItemFile.new(@identifier, name, client: @client).download to: to
|
|
66
|
+
rescue HTTPError => e
|
|
67
|
+
raise unless e.status == 404
|
|
68
|
+
end
|
|
69
|
+
|
|
70
|
+
def write_sidecars
|
|
71
|
+
Metadata::Markdown.new(metadata).write to: staging_dir
|
|
72
|
+
Metadata::YAML.new(metadata).write to: staging_dir
|
|
73
|
+
Metadata::JSON.new(metadata).write to: staging_dir
|
|
74
|
+
Metadata::Bibtex.new(metadata).write to: staging_dir
|
|
75
|
+
end
|
|
76
|
+
end
|
|
77
|
+
end
|
|
78
|
+
end
|
|
@@ -0,0 +1,57 @@
|
|
|
1
|
+
module Jstor
|
|
2
|
+
module Downloader
|
|
3
|
+
# Synthesized from the metadata only: fetching JSTOR's own citation
|
|
4
|
+
# export would be an automated download from jstor.org.
|
|
5
|
+
class Bibtex
|
|
6
|
+
def initialize metadata
|
|
7
|
+
@metadata = metadata
|
|
8
|
+
end
|
|
9
|
+
|
|
10
|
+
def to_s
|
|
11
|
+
<<~BIBTEX
|
|
12
|
+
@article{#{key},
|
|
13
|
+
title={#{@metadata.title}},
|
|
14
|
+
author={#{authors}},
|
|
15
|
+
journal={#{@metadata.journal[:name]}},
|
|
16
|
+
year={#{year}},
|
|
17
|
+
volume={#{@metadata.volume}},
|
|
18
|
+
pages={#{pages}},
|
|
19
|
+
doi={#{@metadata.doi}},
|
|
20
|
+
url={#{@metadata.jstor_url}},
|
|
21
|
+
}
|
|
22
|
+
BIBTEX
|
|
23
|
+
end
|
|
24
|
+
|
|
25
|
+
def key
|
|
26
|
+
title_word = Slug.new(@metadata.title).to_s.split('-').first
|
|
27
|
+
|
|
28
|
+
"#{surname}#{year}#{title_word}"
|
|
29
|
+
end
|
|
30
|
+
|
|
31
|
+
private
|
|
32
|
+
|
|
33
|
+
def authors
|
|
34
|
+
@metadata.authors.map(&:name).join ' and '
|
|
35
|
+
end
|
|
36
|
+
|
|
37
|
+
# "Stebbins, Joel" and "Ella Catherine Greene" both give a lowercase surname
|
|
38
|
+
def surname
|
|
39
|
+
first_author = @metadata.authors.first
|
|
40
|
+
return 'anonymous' if first_author.nil?
|
|
41
|
+
|
|
42
|
+
name = first_author.name
|
|
43
|
+
last = name.include?(',') ? name.split(',').first : name.split.last
|
|
44
|
+
last.downcase.gsub(/[^a-z]/, '')
|
|
45
|
+
end
|
|
46
|
+
|
|
47
|
+
def year
|
|
48
|
+
@metadata.published.to_s[0, 4]
|
|
49
|
+
end
|
|
50
|
+
|
|
51
|
+
# BibTeX page ranges use an en dash: 2--5
|
|
52
|
+
def pages
|
|
53
|
+
@metadata.pages.to_s.sub '-', '--'
|
|
54
|
+
end
|
|
55
|
+
end
|
|
56
|
+
end
|
|
57
|
+
end
|
|
@@ -0,0 +1,106 @@
|
|
|
1
|
+
require 'optparse'
|
|
2
|
+
|
|
3
|
+
module Jstor
|
|
4
|
+
module Downloader
|
|
5
|
+
class CLI
|
|
6
|
+
DEFAULT_DOWNLOAD_PATH = File.join Dir.home, 'Downloads', 'JSTOR_Papers'
|
|
7
|
+
USAGE = 'Usage: jstor-dl [options] <JSTOR_ID_OR_URL> [<JSTOR_ID_OR_URL>...]'.freeze
|
|
8
|
+
|
|
9
|
+
def initialize argv, stderr: $stderr, stdin: $stdin, stdout: $stdout
|
|
10
|
+
@argv = argv
|
|
11
|
+
@stderr = stderr
|
|
12
|
+
@stdin = stdin
|
|
13
|
+
@stdout = stdout
|
|
14
|
+
end
|
|
15
|
+
|
|
16
|
+
def run
|
|
17
|
+
options = parse
|
|
18
|
+
return options.fetch(:exit_status) if options.key? :exit_status
|
|
19
|
+
|
|
20
|
+
return error_with USAGE if options[:targets].empty?
|
|
21
|
+
return error_with conflict if options[:verbose] && options[:quiet]
|
|
22
|
+
|
|
23
|
+
failures = download_each options
|
|
24
|
+
failures.zero? ? 0 : 1
|
|
25
|
+
end
|
|
26
|
+
|
|
27
|
+
private
|
|
28
|
+
|
|
29
|
+
def conflict
|
|
30
|
+
'-v and -q are mutually exclusive'
|
|
31
|
+
end
|
|
32
|
+
|
|
33
|
+
def parse
|
|
34
|
+
options = { targets: [], verbose: false, quiet: false }
|
|
35
|
+
parser = build_parser options
|
|
36
|
+
|
|
37
|
+
begin
|
|
38
|
+
parser.parse! @argv
|
|
39
|
+
options[:targets] = @argv + input_targets(options[:input])
|
|
40
|
+
rescue OptionParser::ParseError, SystemCallError => e
|
|
41
|
+
@stderr.puts e.message
|
|
42
|
+
return { exit_status: 1 }
|
|
43
|
+
end
|
|
44
|
+
|
|
45
|
+
options[:path] ||= ENV['JSTOR_DOWNLOAD_PATH'] || DEFAULT_DOWNLOAD_PATH
|
|
46
|
+
options[:rate_limit] ||= (ENV['JSTOR_RATE_LIMIT'] || Client::DEFAULT_RATE_LIMIT).to_i
|
|
47
|
+
options
|
|
48
|
+
end
|
|
49
|
+
|
|
50
|
+
def build_parser options
|
|
51
|
+
OptionParser.new do |parser|
|
|
52
|
+
parser.banner = USAGE
|
|
53
|
+
parser.on('-i FILE', '--input FILE') { |value| options[:input] = value }
|
|
54
|
+
parser.on('-p PATH', '--path PATH') { |value| options[:path] = value }
|
|
55
|
+
parser.on('--rate-limit SECONDS', Integer) { |value| options[:rate_limit] = value }
|
|
56
|
+
parser.on('-v', '--verbose') { options[:verbose] = true }
|
|
57
|
+
parser.on('-q', '--quiet') { options[:quiet] = true }
|
|
58
|
+
parser.on('--version') do
|
|
59
|
+
@stdout.puts VERSION
|
|
60
|
+
options[:exit_status] = 0
|
|
61
|
+
end
|
|
62
|
+
parser.on('-h', '--help') do
|
|
63
|
+
@stdout.puts parser.help
|
|
64
|
+
options[:exit_status] = 0
|
|
65
|
+
end
|
|
66
|
+
end
|
|
67
|
+
end
|
|
68
|
+
|
|
69
|
+
# one target per line from FILE, or stdin for "-"; blank lines and # comments skipped
|
|
70
|
+
def input_targets input
|
|
71
|
+
return [] if input.nil?
|
|
72
|
+
|
|
73
|
+
text = input == '-' ? @stdin.read : File.read(input)
|
|
74
|
+
text.lines.map(&:strip).reject { it.empty? || it.start_with?('#') }
|
|
75
|
+
end
|
|
76
|
+
|
|
77
|
+
def error_with message
|
|
78
|
+
@stderr.puts message
|
|
79
|
+
1
|
|
80
|
+
end
|
|
81
|
+
|
|
82
|
+
def download_each options
|
|
83
|
+
client = Client.new rate_limit: options[:rate_limit], log: (options[:verbose] ? @stdout : nil)
|
|
84
|
+
|
|
85
|
+
failures = 0
|
|
86
|
+
options[:targets].each do |target|
|
|
87
|
+
failures += 1 unless download_one(target, client:, options:)
|
|
88
|
+
end
|
|
89
|
+
failures
|
|
90
|
+
end
|
|
91
|
+
|
|
92
|
+
# true on success; reports the failure and returns false otherwise
|
|
93
|
+
def download_one target, client:, options:
|
|
94
|
+
identifier = Identifier.new target
|
|
95
|
+
@stdout.puts "==> Downloading #{identifier.id}" if options[:verbose]
|
|
96
|
+
|
|
97
|
+
path = Archive.new(identifier, root: options[:path], client: client).run
|
|
98
|
+
@stdout.puts path unless options[:quiet]
|
|
99
|
+
true
|
|
100
|
+
rescue Error, HTTP::Error => e
|
|
101
|
+
@stderr.puts "#{target}: #{e.message}"
|
|
102
|
+
false
|
|
103
|
+
end
|
|
104
|
+
end
|
|
105
|
+
end
|
|
106
|
+
end
|
|
@@ -0,0 +1,86 @@
|
|
|
1
|
+
require 'http'
|
|
2
|
+
|
|
3
|
+
module Jstor
|
|
4
|
+
module Downloader
|
|
5
|
+
class Client
|
|
6
|
+
SOURCE_URL = 'https://github.com/xoengineering/jstor-dl'.freeze
|
|
7
|
+
DEFAULT_RATE_LIMIT = 3
|
|
8
|
+
TIMEOUTS = { connect: 10, read: 60, write: 10 }.freeze # seconds, per operation
|
|
9
|
+
MAX_RETRIES = 3
|
|
10
|
+
RETRY_BACKOFF = 10 # seconds before the first retry. doubles on each retry.
|
|
11
|
+
RETRYABLE_STATUSES = [429, 503].freeze
|
|
12
|
+
|
|
13
|
+
attr_reader :rate_limit
|
|
14
|
+
|
|
15
|
+
def initialize rate_limit: DEFAULT_RATE_LIMIT, log: nil
|
|
16
|
+
@rate_limit = rate_limit
|
|
17
|
+
@log = log
|
|
18
|
+
end
|
|
19
|
+
|
|
20
|
+
def user_agent
|
|
21
|
+
"jstor-dl/#{VERSION} (+#{SOURCE_URL})"
|
|
22
|
+
end
|
|
23
|
+
|
|
24
|
+
def get url
|
|
25
|
+
retries = 0
|
|
26
|
+
|
|
27
|
+
loop do
|
|
28
|
+
response = request url
|
|
29
|
+
return response if response.status.success?
|
|
30
|
+
raise http_error(url, response) unless retryable? response, retries
|
|
31
|
+
|
|
32
|
+
retries += 1
|
|
33
|
+
wait_before_retry response, retries
|
|
34
|
+
end
|
|
35
|
+
end
|
|
36
|
+
|
|
37
|
+
private
|
|
38
|
+
|
|
39
|
+
def request url
|
|
40
|
+
throttle
|
|
41
|
+
response = HTTP.timeout(TIMEOUTS).headers('User-Agent' => user_agent).follow.get(url)
|
|
42
|
+
@last_request_at = Time.now
|
|
43
|
+
log_request url, response
|
|
44
|
+
response
|
|
45
|
+
end
|
|
46
|
+
|
|
47
|
+
def http_error url, response
|
|
48
|
+
HTTPError.new status: response.status.code, url: url, reason: response.status.reason
|
|
49
|
+
end
|
|
50
|
+
|
|
51
|
+
def retryable? response, retries
|
|
52
|
+
RETRYABLE_STATUSES.include?(response.status.code) && retries < MAX_RETRIES
|
|
53
|
+
end
|
|
54
|
+
|
|
55
|
+
def wait_before_retry response, retries
|
|
56
|
+
seconds = retry_after(response) || (RETRY_BACKOFF * (2**(retries - 1)))
|
|
57
|
+
@log&.puts "==> #{response.status}. Retrying in #{seconds}s"
|
|
58
|
+
sleep seconds
|
|
59
|
+
end
|
|
60
|
+
|
|
61
|
+
# Retry-After in delay-seconds form. The HTTP-date form falls back to backoff.
|
|
62
|
+
def retry_after response
|
|
63
|
+
value = response.headers['Retry-After']
|
|
64
|
+
return if value.nil?
|
|
65
|
+
|
|
66
|
+
Integer(value, exception: false)
|
|
67
|
+
end
|
|
68
|
+
|
|
69
|
+
def log_request url, response
|
|
70
|
+
return if @log.nil?
|
|
71
|
+
|
|
72
|
+
@log.puts "==> GET #{url} (#{response.body.to_s.bytesize} bytes)"
|
|
73
|
+
end
|
|
74
|
+
|
|
75
|
+
def throttle
|
|
76
|
+
return if @rate_limit.zero?
|
|
77
|
+
return if @last_request_at.nil?
|
|
78
|
+
|
|
79
|
+
elapsed = Time.now - @last_request_at
|
|
80
|
+
return if elapsed >= @rate_limit
|
|
81
|
+
|
|
82
|
+
sleep(@rate_limit - elapsed)
|
|
83
|
+
end
|
|
84
|
+
end
|
|
85
|
+
end
|
|
86
|
+
end
|
|
@@ -0,0 +1,14 @@
|
|
|
1
|
+
module Jstor
|
|
2
|
+
module Downloader
|
|
3
|
+
class HTTPError < Error
|
|
4
|
+
attr_reader :status, :url
|
|
5
|
+
|
|
6
|
+
def initialize status:, url:, reason: nil
|
|
7
|
+
@status = status
|
|
8
|
+
@url = url
|
|
9
|
+
|
|
10
|
+
super("GET #{url} failed: #{[status, reason].compact.join ' '}")
|
|
11
|
+
end
|
|
12
|
+
end
|
|
13
|
+
end
|
|
14
|
+
end
|
|
@@ -0,0 +1,66 @@
|
|
|
1
|
+
require 'uri'
|
|
2
|
+
|
|
3
|
+
module Jstor
|
|
4
|
+
module Downloader
|
|
5
|
+
# A JSTOR stable ID for an Early Journal Content article. EJC stable IDs are
|
|
6
|
+
# numeric; JSTOR's newer IDs (j.ctt…, resrep…) are never in the EJC.
|
|
7
|
+
class Identifier
|
|
8
|
+
class Invalid < Error; end
|
|
9
|
+
|
|
10
|
+
STABLE_ID = /\A\d+\z/
|
|
11
|
+
DOI = %r{\A(?:doi:)?10\.2307/(\d+)\z}i
|
|
12
|
+
ITEM = /\Ajstor-(\d+)\z/
|
|
13
|
+
|
|
14
|
+
JSTOR_PATHS = [%r{\A/stable/(?:10\.2307/)?(\d+)\z}, %r{\A/stable/pdf/(\d+)\.pdf\z}].freeze
|
|
15
|
+
DOI_PATH = %r{\A/10\.2307/(\d+)\z}
|
|
16
|
+
ARCHIVE_PATH = %r{\A/(?:details|download)/jstor-(\d+)(?:/|\z)}
|
|
17
|
+
JSTOR_HOSTS = %w[jstor.org www.jstor.org].freeze
|
|
18
|
+
DOI_HOSTS = %w[doi.org dx.doi.org].freeze
|
|
19
|
+
ARCHIVE_HOSTS = %w[archive.org www.archive.org].freeze
|
|
20
|
+
|
|
21
|
+
attr_reader :id, :input
|
|
22
|
+
|
|
23
|
+
def initialize input
|
|
24
|
+
@input = input
|
|
25
|
+
@id = parse input.to_s.strip
|
|
26
|
+
raise Invalid, "not a JSTOR Early Journal Content identifier: #{input}" if @id.nil?
|
|
27
|
+
end
|
|
28
|
+
|
|
29
|
+
# the archive.org item holding this article
|
|
30
|
+
def item = "jstor-#{id}"
|
|
31
|
+
|
|
32
|
+
def doi = "10.2307/#{id}"
|
|
33
|
+
|
|
34
|
+
def to_s = id
|
|
35
|
+
|
|
36
|
+
private
|
|
37
|
+
|
|
38
|
+
def parse text
|
|
39
|
+
return text if STABLE_ID.match? text
|
|
40
|
+
return Regexp.last_match(1) if DOI.match(text) || ITEM.match(text)
|
|
41
|
+
|
|
42
|
+
parse_url text
|
|
43
|
+
end
|
|
44
|
+
|
|
45
|
+
def parse_url text
|
|
46
|
+
uri = URI.parse(text.include?('://') ? text : "https://#{text}")
|
|
47
|
+
path = uri.path.to_s
|
|
48
|
+
|
|
49
|
+
return first_capture(JSTOR_PATHS, path) if JSTOR_HOSTS.include? uri.host
|
|
50
|
+
return first_capture([DOI_PATH], path) if DOI_HOSTS.include? uri.host
|
|
51
|
+
|
|
52
|
+
first_capture([ARCHIVE_PATH], path) if ARCHIVE_HOSTS.include? uri.host
|
|
53
|
+
rescue URI::InvalidURIError
|
|
54
|
+
nil
|
|
55
|
+
end
|
|
56
|
+
|
|
57
|
+
def first_capture patterns, path
|
|
58
|
+
patterns.each do |pattern|
|
|
59
|
+
match = pattern.match path
|
|
60
|
+
return match[1] if match
|
|
61
|
+
end
|
|
62
|
+
nil
|
|
63
|
+
end
|
|
64
|
+
end
|
|
65
|
+
end
|
|
66
|
+
end
|
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
module Jstor
|
|
2
|
+
module Downloader
|
|
3
|
+
# One file from the article's archive.org item. /download/ redirects to a
|
|
4
|
+
# storage node; Client follows redirects across hosts.
|
|
5
|
+
class ItemFile
|
|
6
|
+
def initialize identifier, name, client:
|
|
7
|
+
@identifier = identifier
|
|
8
|
+
@name = name
|
|
9
|
+
@client = client
|
|
10
|
+
end
|
|
11
|
+
|
|
12
|
+
def download to:
|
|
13
|
+
File.binwrite to, @client.get(url).to_s
|
|
14
|
+
end
|
|
15
|
+
|
|
16
|
+
private
|
|
17
|
+
|
|
18
|
+
def url
|
|
19
|
+
"https://archive.org/download/#{@identifier.item}/#{@name}"
|
|
20
|
+
end
|
|
21
|
+
end
|
|
22
|
+
end
|
|
23
|
+
end
|
|
@@ -0,0 +1,13 @@
|
|
|
1
|
+
module Jstor
|
|
2
|
+
module Downloader
|
|
3
|
+
class ItemNotFound < Error
|
|
4
|
+
MESSAGE = <<~MESSAGE.chomp
|
|
5
|
+
not in the Early Journal Content on archive.org (JSTOR's terms allow only manual download from jstor.org)
|
|
6
|
+
MESSAGE
|
|
7
|
+
|
|
8
|
+
def initialize message = MESSAGE
|
|
9
|
+
super
|
|
10
|
+
end
|
|
11
|
+
end
|
|
12
|
+
end
|
|
13
|
+
end
|
|
@@ -0,0 +1,20 @@
|
|
|
1
|
+
require 'fileutils'
|
|
2
|
+
|
|
3
|
+
module Jstor
|
|
4
|
+
module Downloader
|
|
5
|
+
class Metadata
|
|
6
|
+
class Bibtex
|
|
7
|
+
FILENAME = 'metadata.bib'.freeze
|
|
8
|
+
|
|
9
|
+
def initialize metadata
|
|
10
|
+
@metadata = metadata
|
|
11
|
+
end
|
|
12
|
+
|
|
13
|
+
def write to:
|
|
14
|
+
FileUtils.mkdir_p to
|
|
15
|
+
File.write File.join(to, FILENAME), Downloader::Bibtex.new(@metadata).to_s
|
|
16
|
+
end
|
|
17
|
+
end
|
|
18
|
+
end
|
|
19
|
+
end
|
|
20
|
+
end
|
|
@@ -0,0 +1,33 @@
|
|
|
1
|
+
require 'fileutils'
|
|
2
|
+
require 'json'
|
|
3
|
+
|
|
4
|
+
module Jstor
|
|
5
|
+
module Downloader
|
|
6
|
+
class Metadata
|
|
7
|
+
class JSON
|
|
8
|
+
FILENAME = 'metadata.json'.freeze
|
|
9
|
+
|
|
10
|
+
def initialize metadata
|
|
11
|
+
@metadata = metadata
|
|
12
|
+
end
|
|
13
|
+
|
|
14
|
+
def write to:
|
|
15
|
+
FileUtils.mkdir_p to
|
|
16
|
+
File.write File.join(to, FILENAME), "#{::JSON.pretty_generate(serialize(@metadata.to_h))}\n"
|
|
17
|
+
end
|
|
18
|
+
|
|
19
|
+
private
|
|
20
|
+
|
|
21
|
+
def serialize object
|
|
22
|
+
case object
|
|
23
|
+
when Hash then object.to_h { |key, value| [key.to_s, serialize(value)] }
|
|
24
|
+
when Array then object.map { |item| serialize item }
|
|
25
|
+
when Data then serialize object.to_h
|
|
26
|
+
when Date, Time then object.iso8601
|
|
27
|
+
else object
|
|
28
|
+
end
|
|
29
|
+
end
|
|
30
|
+
end
|
|
31
|
+
end
|
|
32
|
+
end
|
|
33
|
+
end
|
|
@@ -0,0 +1,69 @@
|
|
|
1
|
+
require 'fileutils'
|
|
2
|
+
require 'yaml'
|
|
3
|
+
|
|
4
|
+
module Jstor
|
|
5
|
+
module Downloader
|
|
6
|
+
class Metadata
|
|
7
|
+
class Markdown
|
|
8
|
+
FILENAME = 'metadata.md'.freeze
|
|
9
|
+
|
|
10
|
+
def initialize metadata
|
|
11
|
+
@metadata = metadata
|
|
12
|
+
end
|
|
13
|
+
|
|
14
|
+
def write to:
|
|
15
|
+
FileUtils.mkdir_p to
|
|
16
|
+
File.write File.join(to, FILENAME), "#{frontmatter}\n#{body}"
|
|
17
|
+
end
|
|
18
|
+
|
|
19
|
+
private
|
|
20
|
+
|
|
21
|
+
def frontmatter
|
|
22
|
+
"---\n#{::YAML.dump(stringify(frontmatter_hash)).delete_prefix("---\n")}---"
|
|
23
|
+
end
|
|
24
|
+
|
|
25
|
+
def frontmatter_hash
|
|
26
|
+
@metadata.to_h.merge bibtex_key: bibtex_key
|
|
27
|
+
end
|
|
28
|
+
|
|
29
|
+
def bibtex_key
|
|
30
|
+
Downloader::Bibtex.new(@metadata).key
|
|
31
|
+
end
|
|
32
|
+
|
|
33
|
+
def stringify object
|
|
34
|
+
case object
|
|
35
|
+
when Hash then object.to_h { |key, value| [key.to_s, stringify(value)] }
|
|
36
|
+
when Array then object.map { |item| stringify item }
|
|
37
|
+
when Data then stringify object.to_h
|
|
38
|
+
else object
|
|
39
|
+
end
|
|
40
|
+
end
|
|
41
|
+
|
|
42
|
+
# JSTOR asks for acknowledgement as the source of Early Journal Content
|
|
43
|
+
def body
|
|
44
|
+
<<~MARKDOWN
|
|
45
|
+
|
|
46
|
+
# #{@metadata.title}
|
|
47
|
+
|
|
48
|
+
#{authors_list}
|
|
49
|
+
|
|
50
|
+
- Published: #{@metadata.published}
|
|
51
|
+
- #{citation}
|
|
52
|
+
- JSTOR: [#{@metadata.doi}](#{@metadata.jstor_url})
|
|
53
|
+
- Internet Archive: [jstor-#{@metadata.jstor_id}](#{@metadata.archive_url})
|
|
54
|
+
|
|
55
|
+
From JSTOR Early Journal Content, via the Internet Archive.
|
|
56
|
+
MARKDOWN
|
|
57
|
+
end
|
|
58
|
+
|
|
59
|
+
def authors_list
|
|
60
|
+
@metadata.authors.map { |author| "- #{author.name}" }.join "\n"
|
|
61
|
+
end
|
|
62
|
+
|
|
63
|
+
def citation
|
|
64
|
+
"#{@metadata.journal[:name]}, volume #{@metadata.volume}, pages #{@metadata.pages}"
|
|
65
|
+
end
|
|
66
|
+
end
|
|
67
|
+
end
|
|
68
|
+
end
|
|
69
|
+
end
|
|
@@ -0,0 +1,32 @@
|
|
|
1
|
+
require 'fileutils'
|
|
2
|
+
require 'yaml'
|
|
3
|
+
|
|
4
|
+
module Jstor
|
|
5
|
+
module Downloader
|
|
6
|
+
class Metadata
|
|
7
|
+
class YAML
|
|
8
|
+
FILENAME = 'metadata.yaml'.freeze
|
|
9
|
+
|
|
10
|
+
def initialize metadata
|
|
11
|
+
@metadata = metadata
|
|
12
|
+
end
|
|
13
|
+
|
|
14
|
+
def write to:
|
|
15
|
+
FileUtils.mkdir_p to
|
|
16
|
+
File.write File.join(to, FILENAME), ::YAML.dump(stringify(@metadata.to_h))
|
|
17
|
+
end
|
|
18
|
+
|
|
19
|
+
private
|
|
20
|
+
|
|
21
|
+
def stringify object
|
|
22
|
+
case object
|
|
23
|
+
when Hash then object.to_h { |key, value| [key.to_s, stringify(value)] }
|
|
24
|
+
when Array then object.map { |item| stringify item }
|
|
25
|
+
when Data then stringify object.to_h
|
|
26
|
+
else object
|
|
27
|
+
end
|
|
28
|
+
end
|
|
29
|
+
end
|
|
30
|
+
end
|
|
31
|
+
end
|
|
32
|
+
end
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
module Jstor
|
|
2
|
+
module Downloader
|
|
3
|
+
# Field order is the reading order of the metadata sidecars.
|
|
4
|
+
Metadata = Data.define(
|
|
5
|
+
:jstor_id,
|
|
6
|
+
:doi,
|
|
7
|
+
:jstor_url,
|
|
8
|
+
:archive_url,
|
|
9
|
+
:title,
|
|
10
|
+
:authors,
|
|
11
|
+
:published,
|
|
12
|
+
:journal,
|
|
13
|
+
:volume,
|
|
14
|
+
:pages,
|
|
15
|
+
:issn,
|
|
16
|
+
:language,
|
|
17
|
+
:publisher,
|
|
18
|
+
:article_type
|
|
19
|
+
)
|
|
20
|
+
end
|
|
21
|
+
end
|
|
@@ -0,0 +1,43 @@
|
|
|
1
|
+
require 'json'
|
|
2
|
+
|
|
3
|
+
module Jstor
|
|
4
|
+
module Downloader
|
|
5
|
+
# Parses an archive.org item metadata response (https://archive.org/metadata/jstor-<id>)
|
|
6
|
+
class MetadataParser
|
|
7
|
+
def initialize json
|
|
8
|
+
@json = JSON.parse json
|
|
9
|
+
end
|
|
10
|
+
|
|
11
|
+
def metadata
|
|
12
|
+
fields = @json['metadata']
|
|
13
|
+
raise ItemNotFound if fields.nil?
|
|
14
|
+
|
|
15
|
+
identifier = Identifier.new fields.fetch('identifier')
|
|
16
|
+
|
|
17
|
+
Metadata.new(
|
|
18
|
+
jstor_id: identifier.id,
|
|
19
|
+
doi: identifier.doi,
|
|
20
|
+
jstor_url: "https://www.jstor.org/stable/#{identifier}",
|
|
21
|
+
archive_url: "https://archive.org/details/#{identifier.item}",
|
|
22
|
+
title: fields['title'],
|
|
23
|
+
authors: authors_from(fields['creator']),
|
|
24
|
+
published: fields['date'],
|
|
25
|
+
journal: { id: fields['journalabbrv'], name: fields['journaltitle'] },
|
|
26
|
+
volume: fields['volume'],
|
|
27
|
+
pages: fields['pagerange'],
|
|
28
|
+
issn: fields['issn'],
|
|
29
|
+
language: fields['language'],
|
|
30
|
+
publisher: fields['publisher'],
|
|
31
|
+
article_type: fields['article-type']
|
|
32
|
+
)
|
|
33
|
+
end
|
|
34
|
+
|
|
35
|
+
private
|
|
36
|
+
|
|
37
|
+
# archive.org gives a string for one creator and a list for several
|
|
38
|
+
def authors_from creator
|
|
39
|
+
Array(creator).map { Author.new name: it }
|
|
40
|
+
end
|
|
41
|
+
end
|
|
42
|
+
end
|
|
43
|
+
end
|
|
@@ -0,0 +1,28 @@
|
|
|
1
|
+
module Jstor
|
|
2
|
+
module Downloader
|
|
3
|
+
class Path
|
|
4
|
+
def initialize metadata
|
|
5
|
+
@metadata = metadata
|
|
6
|
+
end
|
|
7
|
+
|
|
8
|
+
def to_s
|
|
9
|
+
[date_dir, journal, "#{@metadata.jstor_id}-#{slug}"].join '/'
|
|
10
|
+
end
|
|
11
|
+
|
|
12
|
+
private
|
|
13
|
+
|
|
14
|
+
# 1907-10-05 becomes 1907/10/05; a year-only or year-month date gives a shorter path
|
|
15
|
+
def date_dir
|
|
16
|
+
@metadata.published.to_s.split('-').join '/'
|
|
17
|
+
end
|
|
18
|
+
|
|
19
|
+
def journal
|
|
20
|
+
@metadata.journal[:id]
|
|
21
|
+
end
|
|
22
|
+
|
|
23
|
+
def slug
|
|
24
|
+
Slug.new(@metadata.title).to_s
|
|
25
|
+
end
|
|
26
|
+
end
|
|
27
|
+
end
|
|
28
|
+
end
|
|
@@ -0,0 +1,44 @@
|
|
|
1
|
+
require 'stringex'
|
|
2
|
+
|
|
3
|
+
module Jstor
|
|
4
|
+
module Downloader
|
|
5
|
+
class Slug
|
|
6
|
+
MAX_LENGTH = 80
|
|
7
|
+
|
|
8
|
+
TEX_INLINE_MATH = /\$[^$]*\$/ # $...$
|
|
9
|
+
TEX_DISPLAY_MATH = /\\\(.*?\\\)|\\\[.*?\\\]/m # \(...\) or \[...\]
|
|
10
|
+
TEX_COMMAND = /\\[a-zA-Z]+\*?/ # \emph, \alpha, etc.
|
|
11
|
+
|
|
12
|
+
def initialize title
|
|
13
|
+
@title = title
|
|
14
|
+
end
|
|
15
|
+
|
|
16
|
+
def to_s
|
|
17
|
+
truncate strip_tex(@title).to_url
|
|
18
|
+
end
|
|
19
|
+
|
|
20
|
+
private
|
|
21
|
+
|
|
22
|
+
def strip_tex string
|
|
23
|
+
string
|
|
24
|
+
.gsub(TEX_INLINE_MATH, ' ')
|
|
25
|
+
.gsub(TEX_DISPLAY_MATH, ' ')
|
|
26
|
+
.gsub(TEX_COMMAND, ' ')
|
|
27
|
+
end
|
|
28
|
+
|
|
29
|
+
def truncate slug
|
|
30
|
+
return slug if slug.length <= MAX_LENGTH
|
|
31
|
+
|
|
32
|
+
words = slug.split '-'
|
|
33
|
+
result = +''
|
|
34
|
+
words.each do |word|
|
|
35
|
+
break if result.length + 1 + word.length > MAX_LENGTH
|
|
36
|
+
|
|
37
|
+
result << '-' unless result.empty?
|
|
38
|
+
result << word
|
|
39
|
+
end
|
|
40
|
+
result
|
|
41
|
+
end
|
|
42
|
+
end
|
|
43
|
+
end
|
|
44
|
+
end
|
|
@@ -0,0 +1,24 @@
|
|
|
1
|
+
require_relative 'downloader/archive'
|
|
2
|
+
require_relative 'downloader/author'
|
|
3
|
+
require_relative 'downloader/bibtex'
|
|
4
|
+
require_relative 'downloader/cli'
|
|
5
|
+
require_relative 'downloader/client'
|
|
6
|
+
require_relative 'downloader/error' # before errors below that subclass Error
|
|
7
|
+
require_relative 'downloader/http_error' # after error
|
|
8
|
+
require_relative 'downloader/identifier' # after error
|
|
9
|
+
require_relative 'downloader/item_file'
|
|
10
|
+
require_relative 'downloader/item_not_found' # after error
|
|
11
|
+
require_relative 'downloader/metadata' # before metadata/*: they reopen class Metadata
|
|
12
|
+
require_relative 'downloader/metadata/bibtex' # after metadata
|
|
13
|
+
require_relative 'downloader/metadata/json' # after metadata
|
|
14
|
+
require_relative 'downloader/metadata/markdown' # after metadata
|
|
15
|
+
require_relative 'downloader/metadata/yaml' # after metadata
|
|
16
|
+
require_relative 'downloader/metadata_parser'
|
|
17
|
+
require_relative 'downloader/path'
|
|
18
|
+
require_relative 'downloader/slug'
|
|
19
|
+
require_relative 'downloader/version'
|
|
20
|
+
|
|
21
|
+
module Jstor
|
|
22
|
+
module Downloader
|
|
23
|
+
end
|
|
24
|
+
end
|
metadata
ADDED
|
@@ -0,0 +1,119 @@
|
|
|
1
|
+
--- !ruby/object:Gem::Specification
|
|
2
|
+
name: jstor-dl
|
|
3
|
+
version: !ruby/object:Gem::Version
|
|
4
|
+
version: 0.1.0
|
|
5
|
+
platform: ruby
|
|
6
|
+
authors:
|
|
7
|
+
- Shane Becker
|
|
8
|
+
bindir: exe
|
|
9
|
+
cert_chain: []
|
|
10
|
+
date: 1980-01-02 00:00:00.000000000 Z
|
|
11
|
+
dependencies:
|
|
12
|
+
- !ruby/object:Gem::Dependency
|
|
13
|
+
name: http
|
|
14
|
+
requirement: !ruby/object:Gem::Requirement
|
|
15
|
+
requirements:
|
|
16
|
+
- - "~>"
|
|
17
|
+
- !ruby/object:Gem::Version
|
|
18
|
+
version: '6.0'
|
|
19
|
+
type: :runtime
|
|
20
|
+
prerelease: false
|
|
21
|
+
version_requirements: !ruby/object:Gem::Requirement
|
|
22
|
+
requirements:
|
|
23
|
+
- - "~>"
|
|
24
|
+
- !ruby/object:Gem::Version
|
|
25
|
+
version: '6.0'
|
|
26
|
+
- !ruby/object:Gem::Dependency
|
|
27
|
+
name: ostruct
|
|
28
|
+
requirement: !ruby/object:Gem::Requirement
|
|
29
|
+
requirements:
|
|
30
|
+
- - "~>"
|
|
31
|
+
- !ruby/object:Gem::Version
|
|
32
|
+
version: '0.6'
|
|
33
|
+
type: :runtime
|
|
34
|
+
prerelease: false
|
|
35
|
+
version_requirements: !ruby/object:Gem::Requirement
|
|
36
|
+
requirements:
|
|
37
|
+
- - "~>"
|
|
38
|
+
- !ruby/object:Gem::Version
|
|
39
|
+
version: '0.6'
|
|
40
|
+
- !ruby/object:Gem::Dependency
|
|
41
|
+
name: stringex
|
|
42
|
+
requirement: !ruby/object:Gem::Requirement
|
|
43
|
+
requirements:
|
|
44
|
+
- - "~>"
|
|
45
|
+
- !ruby/object:Gem::Version
|
|
46
|
+
version: '2.8'
|
|
47
|
+
type: :runtime
|
|
48
|
+
prerelease: false
|
|
49
|
+
version_requirements: !ruby/object:Gem::Requirement
|
|
50
|
+
requirements:
|
|
51
|
+
- - "~>"
|
|
52
|
+
- !ruby/object:Gem::Version
|
|
53
|
+
version: '2.8'
|
|
54
|
+
description: |
|
|
55
|
+
Command line tool and Ruby library for archiving JSTOR Early Journal Content articles
|
|
56
|
+
as PDFs and OCR plaintext with sidecar metadata. Fetches only from the Internet Archive's copy,
|
|
57
|
+
never from jstor.org.
|
|
58
|
+
email:
|
|
59
|
+
- veganstraightedge@gmail.com
|
|
60
|
+
executables:
|
|
61
|
+
- jstor-dl
|
|
62
|
+
extensions: []
|
|
63
|
+
extra_rdoc_files: []
|
|
64
|
+
files:
|
|
65
|
+
- CHANGELOG.md
|
|
66
|
+
- CODE_OF_CONDUCT.md
|
|
67
|
+
- LICENSE.md
|
|
68
|
+
- README.md
|
|
69
|
+
- exe/jstor-dl
|
|
70
|
+
- lib/jstor/downloader.rb
|
|
71
|
+
- lib/jstor/downloader/archive.rb
|
|
72
|
+
- lib/jstor/downloader/author.rb
|
|
73
|
+
- lib/jstor/downloader/bibtex.rb
|
|
74
|
+
- lib/jstor/downloader/cli.rb
|
|
75
|
+
- lib/jstor/downloader/client.rb
|
|
76
|
+
- lib/jstor/downloader/error.rb
|
|
77
|
+
- lib/jstor/downloader/http_error.rb
|
|
78
|
+
- lib/jstor/downloader/identifier.rb
|
|
79
|
+
- lib/jstor/downloader/item_file.rb
|
|
80
|
+
- lib/jstor/downloader/item_not_found.rb
|
|
81
|
+
- lib/jstor/downloader/metadata.rb
|
|
82
|
+
- lib/jstor/downloader/metadata/bibtex.rb
|
|
83
|
+
- lib/jstor/downloader/metadata/json.rb
|
|
84
|
+
- lib/jstor/downloader/metadata/markdown.rb
|
|
85
|
+
- lib/jstor/downloader/metadata/yaml.rb
|
|
86
|
+
- lib/jstor/downloader/metadata_parser.rb
|
|
87
|
+
- lib/jstor/downloader/path.rb
|
|
88
|
+
- lib/jstor/downloader/slug.rb
|
|
89
|
+
- lib/jstor/downloader/version.rb
|
|
90
|
+
- sig/jstor/downloader.rbs
|
|
91
|
+
homepage: https://github.com/xoengineering/jstor-dl
|
|
92
|
+
licenses:
|
|
93
|
+
- MIT
|
|
94
|
+
metadata:
|
|
95
|
+
allowed_push_host: https://rubygems.org
|
|
96
|
+
homepage_uri: https://github.com/xoengineering/jstor-dl
|
|
97
|
+
source_code_uri: https://github.com/xoengineering/jstor-dl
|
|
98
|
+
bug_tracker_uri: https://github.com/xoengineering/jstor-dl/issues
|
|
99
|
+
changelog_uri: https://github.com/xoengineering/jstor-dl/blob/main/CHANGELOG.md
|
|
100
|
+
rubygems_mfa_required: 'true'
|
|
101
|
+
rdoc_options: []
|
|
102
|
+
require_paths:
|
|
103
|
+
- lib
|
|
104
|
+
required_ruby_version: !ruby/object:Gem::Requirement
|
|
105
|
+
requirements:
|
|
106
|
+
- - ">="
|
|
107
|
+
- !ruby/object:Gem::Version
|
|
108
|
+
version: 4.0.7
|
|
109
|
+
required_rubygems_version: !ruby/object:Gem::Requirement
|
|
110
|
+
requirements:
|
|
111
|
+
- - ">="
|
|
112
|
+
- !ruby/object:Gem::Version
|
|
113
|
+
version: '0'
|
|
114
|
+
requirements: []
|
|
115
|
+
rubygems_version: 4.0.21
|
|
116
|
+
specification_version: 4
|
|
117
|
+
summary: Download JSTOR's public-domain Early Journal Content from archive.org for
|
|
118
|
+
offline archives.
|
|
119
|
+
test_files: []
|