jstor-dl 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml ADDED
@@ -0,0 +1,7 @@
1
+ ---
2
+ SHA256:
3
+ metadata.gz: cd0503a0a9da1a8298c012fab0dd19f866e907a87a58537f6f1b26174b3bb1ab
4
+ data.tar.gz: 20779888e79ac989cc10b14ac96877b9200641f7e4cc3fe679f252b4602b1570
5
+ SHA512:
6
+ metadata.gz: 6839ae6451e76bfe92b282b4829a7dcea3b117c103bf68d345ad8e91d7a922d56d30bd27007a3be8c0920ac219c67c0b04ef98afa0e0206e0e9e7e8b3ebdf758
7
+ data.tar.gz: d0a458946ce663a24638a25b4fbc23d8579352c101105cb371e764b086524fef22723080e93314512cdfcdd0a8fd061cfb929edaac067f1c06087e24cef0d392
data/CHANGELOG.md ADDED
@@ -0,0 +1,9 @@
1
+ ## [0.1.0]
2
+
3
+ First version. Per-article offline archive of JSTOR's public-domain Early Journal Content, fetched only from the Internet Archive's copy, never from jstor.org.
4
+
5
+ - Accepts JSTOR stable IDs, jstor.org URLs, `10.2307/…` DOIs and doi.org URLs, and archive.org `jstor-…` items and URLs.
6
+ - Saves the scanned PDF, the OCR plaintext, JSTOR's article metadata XML verbatim, and four sidecar metadata files (`metadata.md`, `metadata.yaml`, `metadata.json`, `metadata.bib`).
7
+ - Layout: `YYYY/MM/DD/<journal>/<jstor-id>-<slug>/`.
8
+ - Articles not in the Early Journal Content on archive.org raise `Jstor::Downloader::ItemNotFound`.
9
+ - Rate-limited HTTP client (3s default) with timeouts and retries on 429/503, staged downloads that skip already-archived articles, and a CLI with `--input FILE|-`, per-target error reporting, and exit status 1 on any failure. All copied from arxiv-dl.
@@ -0,0 +1,83 @@
1
+ # Contributor Covenant 3.0 Code of Conduct
2
+
3
+ ## Our Pledge
4
+
5
+ We pledge to make our community welcoming, safe, and equitable for all.
6
+
7
+ We are committed to fostering an environment that respects and promotes the dignity, rights, and contributions of all individuals, regardless of characteristics including race, ethnicity, caste, color, age, physical characteristics, neurodiversity, disability, sex or gender, gender identity or expression, sexual orientation, language, philosophy or religion, national or social origin, socio-economic position, level of education, or other status. The same privileges of participation are extended to everyone who participates in good faith and in accordance with this Covenant.
8
+
9
+ ## Encouraged Behaviors
10
+
11
+ While acknowledging differences in social norms, we all strive to meet our community's expectations for positive behavior. We also understand that our words and actions may be interpreted differently than we intend based on culture, background, or native language.
12
+
13
+ With these considerations in mind, we agree to behave mindfully toward each other and act in ways that center our shared values, including:
14
+
15
+ 1. Respecting the **purpose of our community**, our activities, and our ways of gathering.
16
+ 2. Engaging **kindly and honestly** with others.
17
+ 3. Respecting **different viewpoints** and experiences.
18
+ 4. **Taking responsibility** for our actions and contributions.
19
+ 5. Gracefully giving and accepting **constructive feedback**.
20
+ 6. Committing to **repairing harm** when it occurs.
21
+ 7. Behaving in other ways that promote and sustain the **well-being of our community**.
22
+
23
+ ## Restricted Behaviors
24
+
25
+ We agree to restrict the following behaviors in our community. Instances, threats, and promotion of these behaviors are violations of this Code of Conduct.
26
+
27
+ 1. **Harassment.** Violating explicitly expressed boundaries or engaging in unnecessary personal attention after any clear request to stop.
28
+ 2. **Character attacks.** Making insulting, demeaning, or pejorative comments directed at a community member or group of people.
29
+ 3. **Stereotyping or discrimination.** Characterizing anyone’s personality or behavior on the basis of immutable identities or traits.
30
+ 4. **Sexualization.** Behaving in a way that would generally be considered inappropriately intimate in the context or purpose of the community.
31
+ 5. **Violating confidentiality**. Sharing or acting on someone's personal or private information without their permission.
32
+ 6. **Endangerment.** Causing, encouraging, or threatening violence or other harm toward any person or group.
33
+ 7. Behaving in other ways that **threaten the well-being** of our community.
34
+
35
+ ### Other Restrictions
36
+
37
+ 1. **Misleading identity.** Impersonating someone else for any reason, or pretending to be someone else to evade enforcement actions.
38
+ 2. **Failing to credit sources.** Not properly crediting the sources of content you contribute.
39
+ 3. **Promotional materials**. Sharing marketing or other commercial content in a way that is outside the norms of the community.
40
+ 4. **Irresponsible communication.** Failing to responsibly present content which includes, links or describes any other restricted behaviors.
41
+
42
+ ## Reporting an Issue
43
+
44
+ Tensions can occur between community members even when they are trying their best to collaborate. Not every conflict represents a code of conduct violation, and this Code of Conduct reinforces encouraged behaviors and norms that can help avoid conflicts and minimize harm.
45
+
46
+ When an incident does occur, it is important to report it promptly. To report a possible violation, **email the maintainer at [veganstraightedge@gmail.com](mailto:veganstraightedge@gmail.com).**
47
+
48
+ Community Moderators take reports of violations seriously and will make every effort to respond in a timely manner. They will investigate all reports of code of conduct violations, reviewing messages, logs, and recordings, or interviewing witnesses and other participants. Community Moderators will keep investigation and enforcement actions as transparent as possible while prioritizing safety and confidentiality. In order to honor these values, enforcement actions are carried out in private with the involved parties, but communicating to the whole community may be part of a mutually agreed upon resolution.
49
+
50
+ ## Addressing and Repairing Harm
51
+
52
+ If an investigation by the Community Moderators finds that this Code of Conduct has been violated, the following enforcement ladder may be used to determine how best to repair harm, based on the incident's impact on the individuals involved and the community as a whole. Depending on the severity of a violation, lower rungs on the ladder may be skipped.
53
+
54
+ 1. Warning
55
+ 1. Event: A violation involving a single incident or series of incidents.
56
+ 2. Consequence: A private, written warning from the Community Moderators.
57
+ 3. Repair: Examples of repair include a private written apology, acknowledgement of responsibility, and seeking clarification on expectations.
58
+ 2. Temporarily Limited Activities
59
+ 1. Event: A repeated incidence of a violation that previously resulted in a warning, or the first incidence of a more serious violation.
60
+ 2. Consequence: A private, written warning with a time-limited cooldown period designed to underscore the seriousness of the situation and give the community members involved time to process the incident. The cooldown period may be limited to particular communication channels or interactions with particular community members.
61
+ 3. Repair: Examples of repair may include making an apology, using the cooldown period to reflect on actions and impact, and being thoughtful about re-entering community spaces after the period is over.
62
+ 3. Temporary Suspension
63
+ 1. Event: A pattern of repeated violation which the Community Moderators have tried to address with warnings, or a single serious violation.
64
+ 2. Consequence: A private written warning with conditions for return from suspension. In general, temporary suspensions give the person being suspended time to reflect upon their behavior and possible corrective actions.
65
+ 3. Repair: Examples of repair include respecting the spirit of the suspension, meeting the specified conditions for return, and being thoughtful about how to reintegrate with the community when the suspension is lifted.
66
+ 4. Permanent Ban
67
+ 1. Event: A pattern of repeated code of conduct violations that other steps on the ladder have failed to resolve, or a violation so serious that the Community Moderators determine there is no way to keep the community safe with this person as a member.
68
+ 2. Consequence: Access to all community spaces, tools, and communication channels is removed. In general, permanent bans should be rarely used, should have strong reasoning behind them, and should only be resorted to if working through other remedies has failed to change the behavior.
69
+ 3. Repair: There is no possible repair in cases of this severity.
70
+
71
+ This enforcement ladder is intended as a guideline. It does not limit the ability of Community Managers to use their discretion and judgment, in keeping with the best interests of our community.
72
+
73
+ ## Scope
74
+
75
+ This Code of Conduct applies within all community spaces, and also applies when an individual is officially representing the community in public or other spaces. Examples of representing our community include using an official email address, posting via an official social media account, or acting as an appointed representative at an online or offline event.
76
+
77
+ ## Attribution
78
+
79
+ This Code of Conduct is adapted from the Contributor Covenant, version 3.0, permanently available at [https://www.contributor-covenant.org/version/3/0/](https://www.contributor-covenant.org/version/3/0/).
80
+
81
+ Contributor Covenant is stewarded by the Organization for Ethical Source and licensed under CC BY-SA 4.0. To view a copy of this license, visit [https://creativecommons.org/licenses/by-sa/4.0/](https://creativecommons.org/licenses/by-sa/4.0/)
82
+
83
+ For answers to common questions about Contributor Covenant, see the FAQ at [https://www.contributor-covenant.org/faq](https://www.contributor-covenant.org/faq). Translations are provided at [https://www.contributor-covenant.org/translations](https://www.contributor-covenant.org/translations). Additional enforcement and community guideline resources can be found at [https://www.contributor-covenant.org/resources](https://www.contributor-covenant.org/resources). The enforcement ladder was inspired by the work of [Mozilla’s code of conduct team](https://github.com/mozilla/inclusion).
data/LICENSE.md ADDED
@@ -0,0 +1,21 @@
1
+ The MIT License (MIT)
2
+
3
+ Copyright (c) 2026 Shane Becker
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in
13
+ all copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN
21
+ THE SOFTWARE.
data/README.md ADDED
@@ -0,0 +1,121 @@
1
+ # jstor-dl
2
+
3
+ Download articles from JSTOR's public-domain Early Journal Content for offline archives.
4
+
5
+ For each article, `jstor-dl` saves:
6
+
7
+ - The scanned PDF
8
+ - The OCR plaintext
9
+ - JSTOR's own article metadata (XML), verbatim
10
+ - Four sidecar metadata files: `metadata.md`, `metadata.yaml`, `metadata.json`, `metadata.bib`
11
+
12
+ ## Where the content comes from
13
+
14
+ `jstor-dl` fetches only from the [Internet Archive's copy](https://archive.org/details/jstor_ejc) of JSTOR's Early Journal Content, never from jstor.org. [JSTOR's terms](https://about.jstor.org/terms) forbid any tool from downloading from jstor.org, even a single article; they allow only manual downloading.
15
+
16
+ The Early Journal Content is nearly 500,000 public-domain articles from 200+ journals (published before 1923 in the US, before 1870 elsewhere), released by JSTOR for free non-commercial use with acknowledgement, and uploaded to the Internet Archive in 2013 for bulk harvesting. Articles JSTOR added to its Early Journal Content after 2013 may be missing from the Internet Archive's copy.
17
+
18
+ ## Installation
19
+
20
+ ```sh
21
+ gem install jstor-dl
22
+ ```
23
+
24
+ ## CLI usage
25
+
26
+ ```sh
27
+ jstor-dl <JSTOR_ID_OR_URL> [<JSTOR_ID_OR_URL>...]
28
+ ```
29
+
30
+ Accepted input forms:
31
+
32
+ | Form | Example |
33
+ | ----------------------- | -------------------------------------------------------- |
34
+ | Stable ID | `4385670` |
35
+ | JSTOR URL | `https://www.jstor.org/stable/4385670` |
36
+ | JSTOR URL, DOI form | `https://www.jstor.org/stable/10.2307/4385670` |
37
+ | JSTOR PDF URL | `https://www.jstor.org/stable/pdf/4385670.pdf` |
38
+ | DOI | `10.2307/4385670`, `doi:10.2307/4385670` |
39
+ | DOI URL | `https://doi.org/10.2307/4385670` |
40
+ | Internet Archive item | `jstor-4385670` |
41
+ | Internet Archive URL | `https://archive.org/details/jstor-4385670` |
42
+
43
+ JSTOR URLs are only read for the article's ID; nothing is requested from jstor.org.
44
+
45
+ ### Flags
46
+
47
+ | Flag | Description |
48
+ | ------------------------- | ----------------------------------------------------------------------------- |
49
+ | `-i FILE`, `--input FILE` | Read IDs/URLs from FILE, one per line (`-` for stdin; blanks and `#` skipped) |
50
+ | `-p PATH`, `--path PATH` | Root download directory |
51
+ | `--rate-limit SECONDS` | Seconds between HTTP requests; `0` disables throttling |
52
+ | `-v`, `--verbose` | Print step lines and per-request URL/byte logs to stdout |
53
+ | `-q`, `--quiet` | Print nothing to stdout; errors still go to stderr |
54
+ | `--version` | Print the gem version and exit |
55
+ | `-h`, `--help` | Print help and exit |
56
+
57
+ `-v` and `-q` are mutually exclusive.
58
+
59
+ ### Environment variables
60
+
61
+ | Variable | Effect |
62
+ | --------------------- | ----------------------------------------------------------------- |
63
+ | `JSTOR_DOWNLOAD_PATH` | Root download directory (default: `$HOME/Downloads/JSTOR_Papers`) |
64
+ | `JSTOR_RATE_LIMIT` | Seconds between HTTP requests (default: `3`; `0` disables) |
65
+
66
+ Precedence: CLI flag > ENV var > default.
67
+
68
+ ### Errors and exit status
69
+
70
+ A target that fails (unrecognized ID, not in the Early Journal Content on archive.org, HTTP error, network failure) is reported on stderr as `<target>: <message>`, and the remaining targets still download. Exit status is `0` when every target succeeds and `1` when any fails.
71
+
72
+ ## Output layout
73
+
74
+ ```txt
75
+ $JSTOR_DOWNLOAD_PATH/ # default: $HOME/Downloads/JSTOR_Papers
76
+ YYYY/MM/DD/<journal>/<jstor-id>-<slug>/
77
+ <jstor-id>.pdf # scanned article
78
+ <jstor-id>.txt # OCR plaintext
79
+ jstor.xml # JSTOR's article metadata, verbatim
80
+ metadata.md # YAML frontmatter + Markdown body
81
+ metadata.yaml
82
+ metadata.json
83
+ metadata.bib # synthesized from the metadata
84
+ ```
85
+
86
+ `YYYY/MM/DD` is the publication date (shorter when only the year or month is known). `<journal>` is JSTOR's journal abbreviation (`clasweek` for The Classical Weekly). `<slug>` is derived from the article title.
87
+
88
+ Each article downloads into a sibling `.partial` folder and is renamed into place only when every file succeeded. Re-running skips articles already archived.
89
+
90
+ ## Library usage
91
+
92
+ ```ruby
93
+ require 'jstor/downloader'
94
+
95
+ identifier = Jstor::Downloader::Identifier.new 'https://www.jstor.org/stable/4385670'
96
+ client = Jstor::Downloader::Client.new # 3-second rate limit by default
97
+ path = Jstor::Downloader::Archive.new(identifier, root: '/tmp/papers', client: client).run
98
+ # => "/tmp/papers/1907/10/05/clasweek/4385670-the-elements-of-the-translation-of-latin"
99
+ ```
100
+
101
+ ## Development
102
+
103
+ ```sh
104
+ script/setup # install dependencies
105
+ script/test # run specs and rubocop
106
+ script/console # interactive prompt
107
+ ```
108
+
109
+ Specs run offline against recorded fixtures in `spec/fixtures/http/`. To check those fixtures against the live archive.org API, run:
110
+
111
+ ```sh
112
+ ARCHIVE_LIVE=1 script/test
113
+ ```
114
+
115
+ ## License
116
+
117
+ MIT — see [LICENSE.md](LICENSE.md).
118
+
119
+ ## Code of Conduct
120
+
121
+ This project follows the [Contributor Covenant](https://www.contributor-covenant.org/version/3/0/) 3.0 — see [CODE_OF_CONDUCT.md](CODE_OF_CONDUCT.md).
data/exe/jstor-dl ADDED
@@ -0,0 +1,5 @@
1
+ #!/usr/bin/env ruby
2
+
3
+ require 'jstor/downloader'
4
+
5
+ exit Jstor::Downloader::CLI.new(ARGV).run
@@ -0,0 +1,78 @@
1
+ require 'fileutils'
2
+
3
+ module Jstor
4
+ module Downloader
5
+ class Archive
6
+ def initialize identifier, root:, client: Client.new
7
+ @identifier = identifier
8
+ @root = root
9
+ @client = client
10
+ end
11
+
12
+ # Downloads into <article>.partial/ and
13
+ # renames it into place only once everything succeeded.
14
+ # An existing article folder is always complete and is skipped.
15
+ def run
16
+ return article_dir if Dir.exist? article_dir
17
+
18
+ FileUtils.rm_rf staging_dir
19
+ FileUtils.mkdir_p staging_dir
20
+
21
+ download_pdf
22
+ download_text
23
+ download_jstor_xml
24
+ write_sidecars
25
+
26
+ File.rename staging_dir, article_dir
27
+ article_dir
28
+ end
29
+
30
+ private
31
+
32
+ def metadata
33
+ @metadata ||= MetadataParser.new(@client.get(metadata_url).to_s).metadata
34
+ end
35
+
36
+ def metadata_url
37
+ "https://archive.org/metadata/#{@identifier.item}"
38
+ end
39
+
40
+ def article_dir
41
+ @article_dir ||= File.join @root, Path.new(metadata).to_s
42
+ end
43
+
44
+ # a sibling of article_dir, so the rename is a single atomic step
45
+ def staging_dir
46
+ "#{article_dir}.partial"
47
+ end
48
+
49
+ def download_pdf
50
+ ItemFile.new(@identifier, "#{@identifier}.pdf", client: @client)
51
+ .download to: File.join(staging_dir, "#{@identifier}.pdf")
52
+ end
53
+
54
+ # the OCR plaintext, when the item has it
55
+ def download_text
56
+ download_optional "#{@identifier}_djvu.txt", to: File.join(staging_dir, "#{@identifier}.txt")
57
+ end
58
+
59
+ # JSTOR's own article metadata, verbatim, when the item has it
60
+ def download_jstor_xml
61
+ download_optional "10.2307_#{@identifier}.xml", to: File.join(staging_dir, 'jstor.xml')
62
+ end
63
+
64
+ def download_optional name, to:
65
+ ItemFile.new(@identifier, name, client: @client).download to: to
66
+ rescue HTTPError => e
67
+ raise unless e.status == 404
68
+ end
69
+
70
+ def write_sidecars
71
+ Metadata::Markdown.new(metadata).write to: staging_dir
72
+ Metadata::YAML.new(metadata).write to: staging_dir
73
+ Metadata::JSON.new(metadata).write to: staging_dir
74
+ Metadata::Bibtex.new(metadata).write to: staging_dir
75
+ end
76
+ end
77
+ end
78
+ end
@@ -0,0 +1,11 @@
1
+ module Jstor
2
+ module Downloader
3
+ Author = Data.define :name, :affiliations do
4
+ def initialize name:, affiliations: []
5
+ super
6
+ end
7
+
8
+ def to_s = name
9
+ end
10
+ end
11
+ end
@@ -0,0 +1,57 @@
1
+ module Jstor
2
+ module Downloader
3
+ # Synthesized from the metadata only: fetching JSTOR's own citation
4
+ # export would be an automated download from jstor.org.
5
+ class Bibtex
6
+ def initialize metadata
7
+ @metadata = metadata
8
+ end
9
+
10
+ def to_s
11
+ <<~BIBTEX
12
+ @article{#{key},
13
+ title={#{@metadata.title}},
14
+ author={#{authors}},
15
+ journal={#{@metadata.journal[:name]}},
16
+ year={#{year}},
17
+ volume={#{@metadata.volume}},
18
+ pages={#{pages}},
19
+ doi={#{@metadata.doi}},
20
+ url={#{@metadata.jstor_url}},
21
+ }
22
+ BIBTEX
23
+ end
24
+
25
+ def key
26
+ title_word = Slug.new(@metadata.title).to_s.split('-').first
27
+
28
+ "#{surname}#{year}#{title_word}"
29
+ end
30
+
31
+ private
32
+
33
+ def authors
34
+ @metadata.authors.map(&:name).join ' and '
35
+ end
36
+
37
+ # "Stebbins, Joel" and "Ella Catherine Greene" both give a lowercase surname
38
+ def surname
39
+ first_author = @metadata.authors.first
40
+ return 'anonymous' if first_author.nil?
41
+
42
+ name = first_author.name
43
+ last = name.include?(',') ? name.split(',').first : name.split.last
44
+ last.downcase.gsub(/[^a-z]/, '')
45
+ end
46
+
47
+ def year
48
+ @metadata.published.to_s[0, 4]
49
+ end
50
+
51
+ # BibTeX page ranges use an en dash: 2--5
52
+ def pages
53
+ @metadata.pages.to_s.sub '-', '--'
54
+ end
55
+ end
56
+ end
57
+ end
@@ -0,0 +1,106 @@
1
+ require 'optparse'
2
+
3
+ module Jstor
4
+ module Downloader
5
+ class CLI
6
+ DEFAULT_DOWNLOAD_PATH = File.join Dir.home, 'Downloads', 'JSTOR_Papers'
7
+ USAGE = 'Usage: jstor-dl [options] <JSTOR_ID_OR_URL> [<JSTOR_ID_OR_URL>...]'.freeze
8
+
9
+ def initialize argv, stderr: $stderr, stdin: $stdin, stdout: $stdout
10
+ @argv = argv
11
+ @stderr = stderr
12
+ @stdin = stdin
13
+ @stdout = stdout
14
+ end
15
+
16
+ def run
17
+ options = parse
18
+ return options.fetch(:exit_status) if options.key? :exit_status
19
+
20
+ return error_with USAGE if options[:targets].empty?
21
+ return error_with conflict if options[:verbose] && options[:quiet]
22
+
23
+ failures = download_each options
24
+ failures.zero? ? 0 : 1
25
+ end
26
+
27
+ private
28
+
29
+ def conflict
30
+ '-v and -q are mutually exclusive'
31
+ end
32
+
33
+ def parse
34
+ options = { targets: [], verbose: false, quiet: false }
35
+ parser = build_parser options
36
+
37
+ begin
38
+ parser.parse! @argv
39
+ options[:targets] = @argv + input_targets(options[:input])
40
+ rescue OptionParser::ParseError, SystemCallError => e
41
+ @stderr.puts e.message
42
+ return { exit_status: 1 }
43
+ end
44
+
45
+ options[:path] ||= ENV['JSTOR_DOWNLOAD_PATH'] || DEFAULT_DOWNLOAD_PATH
46
+ options[:rate_limit] ||= (ENV['JSTOR_RATE_LIMIT'] || Client::DEFAULT_RATE_LIMIT).to_i
47
+ options
48
+ end
49
+
50
+ def build_parser options
51
+ OptionParser.new do |parser|
52
+ parser.banner = USAGE
53
+ parser.on('-i FILE', '--input FILE') { |value| options[:input] = value }
54
+ parser.on('-p PATH', '--path PATH') { |value| options[:path] = value }
55
+ parser.on('--rate-limit SECONDS', Integer) { |value| options[:rate_limit] = value }
56
+ parser.on('-v', '--verbose') { options[:verbose] = true }
57
+ parser.on('-q', '--quiet') { options[:quiet] = true }
58
+ parser.on('--version') do
59
+ @stdout.puts VERSION
60
+ options[:exit_status] = 0
61
+ end
62
+ parser.on('-h', '--help') do
63
+ @stdout.puts parser.help
64
+ options[:exit_status] = 0
65
+ end
66
+ end
67
+ end
68
+
69
+ # one target per line from FILE, or stdin for "-"; blank lines and # comments skipped
70
+ def input_targets input
71
+ return [] if input.nil?
72
+
73
+ text = input == '-' ? @stdin.read : File.read(input)
74
+ text.lines.map(&:strip).reject { it.empty? || it.start_with?('#') }
75
+ end
76
+
77
+ def error_with message
78
+ @stderr.puts message
79
+ 1
80
+ end
81
+
82
+ def download_each options
83
+ client = Client.new rate_limit: options[:rate_limit], log: (options[:verbose] ? @stdout : nil)
84
+
85
+ failures = 0
86
+ options[:targets].each do |target|
87
+ failures += 1 unless download_one(target, client:, options:)
88
+ end
89
+ failures
90
+ end
91
+
92
+ # true on success; reports the failure and returns false otherwise
93
+ def download_one target, client:, options:
94
+ identifier = Identifier.new target
95
+ @stdout.puts "==> Downloading #{identifier.id}" if options[:verbose]
96
+
97
+ path = Archive.new(identifier, root: options[:path], client: client).run
98
+ @stdout.puts path unless options[:quiet]
99
+ true
100
+ rescue Error, HTTP::Error => e
101
+ @stderr.puts "#{target}: #{e.message}"
102
+ false
103
+ end
104
+ end
105
+ end
106
+ end
@@ -0,0 +1,86 @@
1
+ require 'http'
2
+
3
+ module Jstor
4
+ module Downloader
5
+ class Client
6
+ SOURCE_URL = 'https://github.com/xoengineering/jstor-dl'.freeze
7
+ DEFAULT_RATE_LIMIT = 3
8
+ TIMEOUTS = { connect: 10, read: 60, write: 10 }.freeze # seconds, per operation
9
+ MAX_RETRIES = 3
10
+ RETRY_BACKOFF = 10 # seconds before the first retry. doubles on each retry.
11
+ RETRYABLE_STATUSES = [429, 503].freeze
12
+
13
+ attr_reader :rate_limit
14
+
15
+ def initialize rate_limit: DEFAULT_RATE_LIMIT, log: nil
16
+ @rate_limit = rate_limit
17
+ @log = log
18
+ end
19
+
20
+ def user_agent
21
+ "jstor-dl/#{VERSION} (+#{SOURCE_URL})"
22
+ end
23
+
24
+ def get url
25
+ retries = 0
26
+
27
+ loop do
28
+ response = request url
29
+ return response if response.status.success?
30
+ raise http_error(url, response) unless retryable? response, retries
31
+
32
+ retries += 1
33
+ wait_before_retry response, retries
34
+ end
35
+ end
36
+
37
+ private
38
+
39
+ def request url
40
+ throttle
41
+ response = HTTP.timeout(TIMEOUTS).headers('User-Agent' => user_agent).follow.get(url)
42
+ @last_request_at = Time.now
43
+ log_request url, response
44
+ response
45
+ end
46
+
47
+ def http_error url, response
48
+ HTTPError.new status: response.status.code, url: url, reason: response.status.reason
49
+ end
50
+
51
+ def retryable? response, retries
52
+ RETRYABLE_STATUSES.include?(response.status.code) && retries < MAX_RETRIES
53
+ end
54
+
55
+ def wait_before_retry response, retries
56
+ seconds = retry_after(response) || (RETRY_BACKOFF * (2**(retries - 1)))
57
+ @log&.puts "==> #{response.status}. Retrying in #{seconds}s"
58
+ sleep seconds
59
+ end
60
+
61
+ # Retry-After in delay-seconds form. The HTTP-date form falls back to backoff.
62
+ def retry_after response
63
+ value = response.headers['Retry-After']
64
+ return if value.nil?
65
+
66
+ Integer(value, exception: false)
67
+ end
68
+
69
+ def log_request url, response
70
+ return if @log.nil?
71
+
72
+ @log.puts "==> GET #{url} (#{response.body.to_s.bytesize} bytes)"
73
+ end
74
+
75
+ def throttle
76
+ return if @rate_limit.zero?
77
+ return if @last_request_at.nil?
78
+
79
+ elapsed = Time.now - @last_request_at
80
+ return if elapsed >= @rate_limit
81
+
82
+ sleep(@rate_limit - elapsed)
83
+ end
84
+ end
85
+ end
86
+ end
@@ -0,0 +1,5 @@
1
+ module Jstor
2
+ module Downloader
3
+ class Error < StandardError; end
4
+ end
5
+ end
@@ -0,0 +1,14 @@
1
+ module Jstor
2
+ module Downloader
3
+ class HTTPError < Error
4
+ attr_reader :status, :url
5
+
6
+ def initialize status:, url:, reason: nil
7
+ @status = status
8
+ @url = url
9
+
10
+ super("GET #{url} failed: #{[status, reason].compact.join ' '}")
11
+ end
12
+ end
13
+ end
14
+ end
@@ -0,0 +1,66 @@
1
+ require 'uri'
2
+
3
+ module Jstor
4
+ module Downloader
5
+ # A JSTOR stable ID for an Early Journal Content article. EJC stable IDs are
6
+ # numeric; JSTOR's newer IDs (j.ctt…, resrep…) are never in the EJC.
7
+ class Identifier
8
+ class Invalid < Error; end
9
+
10
+ STABLE_ID = /\A\d+\z/
11
+ DOI = %r{\A(?:doi:)?10\.2307/(\d+)\z}i
12
+ ITEM = /\Ajstor-(\d+)\z/
13
+
14
+ JSTOR_PATHS = [%r{\A/stable/(?:10\.2307/)?(\d+)\z}, %r{\A/stable/pdf/(\d+)\.pdf\z}].freeze
15
+ DOI_PATH = %r{\A/10\.2307/(\d+)\z}
16
+ ARCHIVE_PATH = %r{\A/(?:details|download)/jstor-(\d+)(?:/|\z)}
17
+ JSTOR_HOSTS = %w[jstor.org www.jstor.org].freeze
18
+ DOI_HOSTS = %w[doi.org dx.doi.org].freeze
19
+ ARCHIVE_HOSTS = %w[archive.org www.archive.org].freeze
20
+
21
+ attr_reader :id, :input
22
+
23
+ def initialize input
24
+ @input = input
25
+ @id = parse input.to_s.strip
26
+ raise Invalid, "not a JSTOR Early Journal Content identifier: #{input}" if @id.nil?
27
+ end
28
+
29
+ # the archive.org item holding this article
30
+ def item = "jstor-#{id}"
31
+
32
+ def doi = "10.2307/#{id}"
33
+
34
+ def to_s = id
35
+
36
+ private
37
+
38
+ def parse text
39
+ return text if STABLE_ID.match? text
40
+ return Regexp.last_match(1) if DOI.match(text) || ITEM.match(text)
41
+
42
+ parse_url text
43
+ end
44
+
45
+ def parse_url text
46
+ uri = URI.parse(text.include?('://') ? text : "https://#{text}")
47
+ path = uri.path.to_s
48
+
49
+ return first_capture(JSTOR_PATHS, path) if JSTOR_HOSTS.include? uri.host
50
+ return first_capture([DOI_PATH], path) if DOI_HOSTS.include? uri.host
51
+
52
+ first_capture([ARCHIVE_PATH], path) if ARCHIVE_HOSTS.include? uri.host
53
+ rescue URI::InvalidURIError
54
+ nil
55
+ end
56
+
57
+ def first_capture patterns, path
58
+ patterns.each do |pattern|
59
+ match = pattern.match path
60
+ return match[1] if match
61
+ end
62
+ nil
63
+ end
64
+ end
65
+ end
66
+ end
@@ -0,0 +1,23 @@
1
+ module Jstor
2
+ module Downloader
3
+ # One file from the article's archive.org item. /download/ redirects to a
4
+ # storage node; Client follows redirects across hosts.
5
+ class ItemFile
6
+ def initialize identifier, name, client:
7
+ @identifier = identifier
8
+ @name = name
9
+ @client = client
10
+ end
11
+
12
+ def download to:
13
+ File.binwrite to, @client.get(url).to_s
14
+ end
15
+
16
+ private
17
+
18
+ def url
19
+ "https://archive.org/download/#{@identifier.item}/#{@name}"
20
+ end
21
+ end
22
+ end
23
+ end
@@ -0,0 +1,13 @@
1
+ module Jstor
2
+ module Downloader
3
+ class ItemNotFound < Error
4
+ MESSAGE = <<~MESSAGE.chomp
5
+ not in the Early Journal Content on archive.org (JSTOR's terms allow only manual download from jstor.org)
6
+ MESSAGE
7
+
8
+ def initialize message = MESSAGE
9
+ super
10
+ end
11
+ end
12
+ end
13
+ end
@@ -0,0 +1,20 @@
1
+ require 'fileutils'
2
+
3
+ module Jstor
4
+ module Downloader
5
+ class Metadata
6
+ class Bibtex
7
+ FILENAME = 'metadata.bib'.freeze
8
+
9
+ def initialize metadata
10
+ @metadata = metadata
11
+ end
12
+
13
+ def write to:
14
+ FileUtils.mkdir_p to
15
+ File.write File.join(to, FILENAME), Downloader::Bibtex.new(@metadata).to_s
16
+ end
17
+ end
18
+ end
19
+ end
20
+ end
@@ -0,0 +1,33 @@
1
+ require 'fileutils'
2
+ require 'json'
3
+
4
+ module Jstor
5
+ module Downloader
6
+ class Metadata
7
+ class JSON
8
+ FILENAME = 'metadata.json'.freeze
9
+
10
+ def initialize metadata
11
+ @metadata = metadata
12
+ end
13
+
14
+ def write to:
15
+ FileUtils.mkdir_p to
16
+ File.write File.join(to, FILENAME), "#{::JSON.pretty_generate(serialize(@metadata.to_h))}\n"
17
+ end
18
+
19
+ private
20
+
21
+ def serialize object
22
+ case object
23
+ when Hash then object.to_h { |key, value| [key.to_s, serialize(value)] }
24
+ when Array then object.map { |item| serialize item }
25
+ when Data then serialize object.to_h
26
+ when Date, Time then object.iso8601
27
+ else object
28
+ end
29
+ end
30
+ end
31
+ end
32
+ end
33
+ end
@@ -0,0 +1,69 @@
1
+ require 'fileutils'
2
+ require 'yaml'
3
+
4
+ module Jstor
5
+ module Downloader
6
+ class Metadata
7
+ class Markdown
8
+ FILENAME = 'metadata.md'.freeze
9
+
10
+ def initialize metadata
11
+ @metadata = metadata
12
+ end
13
+
14
+ def write to:
15
+ FileUtils.mkdir_p to
16
+ File.write File.join(to, FILENAME), "#{frontmatter}\n#{body}"
17
+ end
18
+
19
+ private
20
+
21
+ def frontmatter
22
+ "---\n#{::YAML.dump(stringify(frontmatter_hash)).delete_prefix("---\n")}---"
23
+ end
24
+
25
+ def frontmatter_hash
26
+ @metadata.to_h.merge bibtex_key: bibtex_key
27
+ end
28
+
29
+ def bibtex_key
30
+ Downloader::Bibtex.new(@metadata).key
31
+ end
32
+
33
+ def stringify object
34
+ case object
35
+ when Hash then object.to_h { |key, value| [key.to_s, stringify(value)] }
36
+ when Array then object.map { |item| stringify item }
37
+ when Data then stringify object.to_h
38
+ else object
39
+ end
40
+ end
41
+
42
+ # JSTOR asks for acknowledgement as the source of Early Journal Content
43
+ def body
44
+ <<~MARKDOWN
45
+
46
+ # #{@metadata.title}
47
+
48
+ #{authors_list}
49
+
50
+ - Published: #{@metadata.published}
51
+ - #{citation}
52
+ - JSTOR: [#{@metadata.doi}](#{@metadata.jstor_url})
53
+ - Internet Archive: [jstor-#{@metadata.jstor_id}](#{@metadata.archive_url})
54
+
55
+ From JSTOR Early Journal Content, via the Internet Archive.
56
+ MARKDOWN
57
+ end
58
+
59
+ def authors_list
60
+ @metadata.authors.map { |author| "- #{author.name}" }.join "\n"
61
+ end
62
+
63
+ def citation
64
+ "#{@metadata.journal[:name]}, volume #{@metadata.volume}, pages #{@metadata.pages}"
65
+ end
66
+ end
67
+ end
68
+ end
69
+ end
@@ -0,0 +1,32 @@
1
+ require 'fileutils'
2
+ require 'yaml'
3
+
4
+ module Jstor
5
+ module Downloader
6
+ class Metadata
7
+ class YAML
8
+ FILENAME = 'metadata.yaml'.freeze
9
+
10
+ def initialize metadata
11
+ @metadata = metadata
12
+ end
13
+
14
+ def write to:
15
+ FileUtils.mkdir_p to
16
+ File.write File.join(to, FILENAME), ::YAML.dump(stringify(@metadata.to_h))
17
+ end
18
+
19
+ private
20
+
21
+ def stringify object
22
+ case object
23
+ when Hash then object.to_h { |key, value| [key.to_s, stringify(value)] }
24
+ when Array then object.map { |item| stringify item }
25
+ when Data then stringify object.to_h
26
+ else object
27
+ end
28
+ end
29
+ end
30
+ end
31
+ end
32
+ end
@@ -0,0 +1,21 @@
1
+ module Jstor
2
+ module Downloader
3
+ # Field order is the reading order of the metadata sidecars.
4
+ Metadata = Data.define(
5
+ :jstor_id,
6
+ :doi,
7
+ :jstor_url,
8
+ :archive_url,
9
+ :title,
10
+ :authors,
11
+ :published,
12
+ :journal,
13
+ :volume,
14
+ :pages,
15
+ :issn,
16
+ :language,
17
+ :publisher,
18
+ :article_type
19
+ )
20
+ end
21
+ end
@@ -0,0 +1,43 @@
1
+ require 'json'
2
+
3
+ module Jstor
4
+ module Downloader
5
+ # Parses an archive.org item metadata response (https://archive.org/metadata/jstor-<id>)
6
+ class MetadataParser
7
+ def initialize json
8
+ @json = JSON.parse json
9
+ end
10
+
11
+ def metadata
12
+ fields = @json['metadata']
13
+ raise ItemNotFound if fields.nil?
14
+
15
+ identifier = Identifier.new fields.fetch('identifier')
16
+
17
+ Metadata.new(
18
+ jstor_id: identifier.id,
19
+ doi: identifier.doi,
20
+ jstor_url: "https://www.jstor.org/stable/#{identifier}",
21
+ archive_url: "https://archive.org/details/#{identifier.item}",
22
+ title: fields['title'],
23
+ authors: authors_from(fields['creator']),
24
+ published: fields['date'],
25
+ journal: { id: fields['journalabbrv'], name: fields['journaltitle'] },
26
+ volume: fields['volume'],
27
+ pages: fields['pagerange'],
28
+ issn: fields['issn'],
29
+ language: fields['language'],
30
+ publisher: fields['publisher'],
31
+ article_type: fields['article-type']
32
+ )
33
+ end
34
+
35
+ private
36
+
37
+ # archive.org gives a string for one creator and a list for several
38
+ def authors_from creator
39
+ Array(creator).map { Author.new name: it }
40
+ end
41
+ end
42
+ end
43
+ end
@@ -0,0 +1,28 @@
1
+ module Jstor
2
+ module Downloader
3
+ class Path
4
+ def initialize metadata
5
+ @metadata = metadata
6
+ end
7
+
8
+ def to_s
9
+ [date_dir, journal, "#{@metadata.jstor_id}-#{slug}"].join '/'
10
+ end
11
+
12
+ private
13
+
14
+ # 1907-10-05 becomes 1907/10/05; a year-only or year-month date gives a shorter path
15
+ def date_dir
16
+ @metadata.published.to_s.split('-').join '/'
17
+ end
18
+
19
+ def journal
20
+ @metadata.journal[:id]
21
+ end
22
+
23
+ def slug
24
+ Slug.new(@metadata.title).to_s
25
+ end
26
+ end
27
+ end
28
+ end
@@ -0,0 +1,44 @@
1
+ require 'stringex'
2
+
3
+ module Jstor
4
+ module Downloader
5
+ class Slug
6
+ MAX_LENGTH = 80
7
+
8
+ TEX_INLINE_MATH = /\$[^$]*\$/ # $...$
9
+ TEX_DISPLAY_MATH = /\\\(.*?\\\)|\\\[.*?\\\]/m # \(...\) or \[...\]
10
+ TEX_COMMAND = /\\[a-zA-Z]+\*?/ # \emph, \alpha, etc.
11
+
12
+ def initialize title
13
+ @title = title
14
+ end
15
+
16
+ def to_s
17
+ truncate strip_tex(@title).to_url
18
+ end
19
+
20
+ private
21
+
22
+ def strip_tex string
23
+ string
24
+ .gsub(TEX_INLINE_MATH, ' ')
25
+ .gsub(TEX_DISPLAY_MATH, ' ')
26
+ .gsub(TEX_COMMAND, ' ')
27
+ end
28
+
29
+ def truncate slug
30
+ return slug if slug.length <= MAX_LENGTH
31
+
32
+ words = slug.split '-'
33
+ result = +''
34
+ words.each do |word|
35
+ break if result.length + 1 + word.length > MAX_LENGTH
36
+
37
+ result << '-' unless result.empty?
38
+ result << word
39
+ end
40
+ result
41
+ end
42
+ end
43
+ end
44
+ end
@@ -0,0 +1,5 @@
1
+ module Jstor
2
+ module Downloader
3
+ VERSION = '0.1.0'.freeze
4
+ end
5
+ end
@@ -0,0 +1,24 @@
1
+ require_relative 'downloader/archive'
2
+ require_relative 'downloader/author'
3
+ require_relative 'downloader/bibtex'
4
+ require_relative 'downloader/cli'
5
+ require_relative 'downloader/client'
6
+ require_relative 'downloader/error' # before errors below that subclass Error
7
+ require_relative 'downloader/http_error' # after error
8
+ require_relative 'downloader/identifier' # after error
9
+ require_relative 'downloader/item_file'
10
+ require_relative 'downloader/item_not_found' # after error
11
+ require_relative 'downloader/metadata' # before metadata/*: they reopen class Metadata
12
+ require_relative 'downloader/metadata/bibtex' # after metadata
13
+ require_relative 'downloader/metadata/json' # after metadata
14
+ require_relative 'downloader/metadata/markdown' # after metadata
15
+ require_relative 'downloader/metadata/yaml' # after metadata
16
+ require_relative 'downloader/metadata_parser'
17
+ require_relative 'downloader/path'
18
+ require_relative 'downloader/slug'
19
+ require_relative 'downloader/version'
20
+
21
+ module Jstor
22
+ module Downloader
23
+ end
24
+ end
@@ -0,0 +1,5 @@
1
+ module Jstor
2
+ module Downloader
3
+ VERSION: String
4
+ end
5
+ end
metadata ADDED
@@ -0,0 +1,119 @@
1
+ --- !ruby/object:Gem::Specification
2
+ name: jstor-dl
3
+ version: !ruby/object:Gem::Version
4
+ version: 0.1.0
5
+ platform: ruby
6
+ authors:
7
+ - Shane Becker
8
+ bindir: exe
9
+ cert_chain: []
10
+ date: 1980-01-02 00:00:00.000000000 Z
11
+ dependencies:
12
+ - !ruby/object:Gem::Dependency
13
+ name: http
14
+ requirement: !ruby/object:Gem::Requirement
15
+ requirements:
16
+ - - "~>"
17
+ - !ruby/object:Gem::Version
18
+ version: '6.0'
19
+ type: :runtime
20
+ prerelease: false
21
+ version_requirements: !ruby/object:Gem::Requirement
22
+ requirements:
23
+ - - "~>"
24
+ - !ruby/object:Gem::Version
25
+ version: '6.0'
26
+ - !ruby/object:Gem::Dependency
27
+ name: ostruct
28
+ requirement: !ruby/object:Gem::Requirement
29
+ requirements:
30
+ - - "~>"
31
+ - !ruby/object:Gem::Version
32
+ version: '0.6'
33
+ type: :runtime
34
+ prerelease: false
35
+ version_requirements: !ruby/object:Gem::Requirement
36
+ requirements:
37
+ - - "~>"
38
+ - !ruby/object:Gem::Version
39
+ version: '0.6'
40
+ - !ruby/object:Gem::Dependency
41
+ name: stringex
42
+ requirement: !ruby/object:Gem::Requirement
43
+ requirements:
44
+ - - "~>"
45
+ - !ruby/object:Gem::Version
46
+ version: '2.8'
47
+ type: :runtime
48
+ prerelease: false
49
+ version_requirements: !ruby/object:Gem::Requirement
50
+ requirements:
51
+ - - "~>"
52
+ - !ruby/object:Gem::Version
53
+ version: '2.8'
54
+ description: |
55
+ Command line tool and Ruby library for archiving JSTOR Early Journal Content articles
56
+ as PDFs and OCR plaintext with sidecar metadata. Fetches only from the Internet Archive's copy,
57
+ never from jstor.org.
58
+ email:
59
+ - veganstraightedge@gmail.com
60
+ executables:
61
+ - jstor-dl
62
+ extensions: []
63
+ extra_rdoc_files: []
64
+ files:
65
+ - CHANGELOG.md
66
+ - CODE_OF_CONDUCT.md
67
+ - LICENSE.md
68
+ - README.md
69
+ - exe/jstor-dl
70
+ - lib/jstor/downloader.rb
71
+ - lib/jstor/downloader/archive.rb
72
+ - lib/jstor/downloader/author.rb
73
+ - lib/jstor/downloader/bibtex.rb
74
+ - lib/jstor/downloader/cli.rb
75
+ - lib/jstor/downloader/client.rb
76
+ - lib/jstor/downloader/error.rb
77
+ - lib/jstor/downloader/http_error.rb
78
+ - lib/jstor/downloader/identifier.rb
79
+ - lib/jstor/downloader/item_file.rb
80
+ - lib/jstor/downloader/item_not_found.rb
81
+ - lib/jstor/downloader/metadata.rb
82
+ - lib/jstor/downloader/metadata/bibtex.rb
83
+ - lib/jstor/downloader/metadata/json.rb
84
+ - lib/jstor/downloader/metadata/markdown.rb
85
+ - lib/jstor/downloader/metadata/yaml.rb
86
+ - lib/jstor/downloader/metadata_parser.rb
87
+ - lib/jstor/downloader/path.rb
88
+ - lib/jstor/downloader/slug.rb
89
+ - lib/jstor/downloader/version.rb
90
+ - sig/jstor/downloader.rbs
91
+ homepage: https://github.com/xoengineering/jstor-dl
92
+ licenses:
93
+ - MIT
94
+ metadata:
95
+ allowed_push_host: https://rubygems.org
96
+ homepage_uri: https://github.com/xoengineering/jstor-dl
97
+ source_code_uri: https://github.com/xoengineering/jstor-dl
98
+ bug_tracker_uri: https://github.com/xoengineering/jstor-dl/issues
99
+ changelog_uri: https://github.com/xoengineering/jstor-dl/blob/main/CHANGELOG.md
100
+ rubygems_mfa_required: 'true'
101
+ rdoc_options: []
102
+ require_paths:
103
+ - lib
104
+ required_ruby_version: !ruby/object:Gem::Requirement
105
+ requirements:
106
+ - - ">="
107
+ - !ruby/object:Gem::Version
108
+ version: 4.0.7
109
+ required_rubygems_version: !ruby/object:Gem::Requirement
110
+ requirements:
111
+ - - ">="
112
+ - !ruby/object:Gem::Version
113
+ version: '0'
114
+ requirements: []
115
+ rubygems_version: 4.0.21
116
+ specification_version: 4
117
+ summary: Download JSTOR's public-domain Early Journal Content from archive.org for
118
+ offline archives.
119
+ test_files: []