arxiv-dl 0.2.0 → 0.3.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 36af277147beae3a55221ba50443204db5b9e4bdae1613a3d9e923a465d4bc31
4
- data.tar.gz: 04d405743856249d2f78d4f9e711e73b716ddb27f1b6633ad92a96e6385677be
3
+ metadata.gz: 14d23483457b1dbb20c8aa9fa1384e52caf7144cab5ed357450f235dbcd29b80
4
+ data.tar.gz: 4d7255130c208f959bb0bb7249a33b3a076a362da7894dc370e480f2511109b7
5
5
  SHA512:
6
- metadata.gz: 92527994ad88aec89a64f3f678f60ff8a9d286d3189d06a6eafa4a9ba4f5a02fa815e9d23c0e429c0ea44ad03d3d58ae870c79a001c0705a682fd22818210412
7
- data.tar.gz: 5058f511ec0e2707275605bee72d29d44c190e563e9b83da65e8b8619199186345c78231b35aa685fe183d6c0dd6f6996115059fd72c991c8d7babb39dfe7be8
6
+ metadata.gz: '0572297cdb2a74bb5cb46cdcd8449e2e41611a9649493e8ba5bc22b2fcb84d264e5982f114d2cff640cfd76e1761de592c8a472815ee6a8116b70f35807a11a0'
7
+ data.tar.gz: 8fd1844b84d8fc2044b8d92078ad2dae136713c93768269b64fe9a83b169a9da37466661dbaf61bb12f81d2ba827276f9c56e00291b978e8bdf7524de4a858ad
data/CHANGELOG.md CHANGED
@@ -1,3 +1,16 @@
1
+ ## [0.3.1]
2
+
3
+ - Built on [dl-core](https://github.com/xoengineering/dl-core) 0.1: the HTTP client, errors, `Author`, `Slug`, sidecar writers, and CLI now come from it instead of copies. No change in behavior or output. `Arxiv::Downloader::Client`, `HTTPError`, `Author`, `Slug`, and `Metadata::YAML`/`JSON` still work; they are now `dl-core`'s classes. `Arxiv::Downloader::Error` now subclasses `DL::Core::Error`.
4
+ - Direct dependencies on `http`, `stringex`, and `ostruct` replaced by `dl-core`.
5
+
6
+ ## [0.3.0]
7
+
8
+ Less folder nesting for the common case: most papers only ever have one version.
9
+
10
+ ### Breaking
11
+
12
+ - Single-version papers are archived flat again: a paper whose only archived version is v1 keeps its files directly in the paper folder, without a `v1/` level. `v<N>/` folders are used only when a paper has more than one version. Archiving a second version of a flat paper moves the existing files into `v<N>/` first, rewriting `html/` links into `_shared/` for the new depth. A paper whose latest version is v2 or later starts out in `v<N>/` folders. Existing 0.2.0 archives keep their `v1/` folders.
13
+
1
14
  ## [0.2.0]
2
15
 
3
16
  Versioned archives, author affiliations, and a round of robustness fixes. Two breaking changes to the output layout and metadata; see below.
data/README.md CHANGED
@@ -108,27 +108,40 @@ $ARXIV_DOWNLOAD_PATH/ # default: $HOME/Downloads/ArXiv_Papers
108
108
  arxiv.org/static/...
109
109
  cdn.jsdelivr.net/...
110
110
  YYYY/MM/DD/<primary_category>/<arxiv-id>-<slug>/
111
- v<N>/ # one folder per archived version
112
- <arxiv-id>v<N>.pdf
113
- <arxiv-id>v<N>-abstract.html
114
- metadata.md # YAML frontmatter + Markdown body
115
- metadata.yaml
116
- metadata.json
117
- metadata.bib # upstream BibTeX, falls back to synthesized
118
- html/ # absent when the paper has no HTML version
119
- <arxiv-id>v<N>.html # path-rewritten to local assets
120
- x1.png, x2.png, ... # paper-specific images
121
- src/ # absent for PDF-only submissions
122
- *.tex, *.bbl, ... # extracted from /src/<id>v<N>
111
+ <arxiv-id>v<N>.pdf
112
+ <arxiv-id>v<N>-abstract.html
113
+ metadata.md # YAML frontmatter + Markdown body
114
+ metadata.yaml
115
+ metadata.json
116
+ metadata.bib # upstream BibTeX, falls back to synthesized
117
+ html/ # absent when the paper has no HTML version
118
+ <arxiv-id>v<N>.html # path-rewritten to local assets
119
+ x1.png, x2.png, ... # paper-specific images
120
+ src/ # absent for PDF-only submissions
121
+ *.tex, *.bbl, ... # extracted from /src/<id>v<N>
123
122
  ```
124
123
 
125
- An unversioned ID (`2508.16190`) archives the latest version. A versioned ID (`2508.16190v1`) archives that version. Different versions of the same paper sit side by side under the same paper folder.
124
+ An unversioned ID (`2508.16190`) archives the latest version. A versioned ID (`2508.16190v1`) archives that version.
126
125
 
127
- Each version downloads into `v<N>.partial/` and is renamed to `v<N>/` only when every file succeeded. Re-running skips versions whose `v<N>/` already exists and retries interrupted ones from scratch.
126
+ A paper with a single archived version v1 is kept flat, as above. When a paper has more than one version, each version gets its own `v<N>/` folder with the same contents:
127
+
128
+ ```txt
129
+ YYYY/MM/DD/<primary_category>/<arxiv-id>-<slug>/
130
+ v1/
131
+ <arxiv-id>v1.pdf
132
+ ...
133
+ v2/
134
+ <arxiv-id>v2.pdf
135
+ ...
136
+ ```
137
+
138
+ Archiving a second version of a flat paper first moves the existing files into `v<N>/` (rewriting `html/` links into `_shared/` for the new depth). A paper whose latest version is v2 or later starts out in `v<N>/` folders.
139
+
140
+ Each version downloads into a sibling `.partial` folder and is renamed into place only when every file succeeded. Re-running skips versions already archived and retries interrupted ones from scratch.
128
141
 
129
142
  `YYYY/MM/DD` is the original submission date. `<primary_category>` is from the paper's metadata (`cs.CL`, `math.NT`, etc). `<slug>` is derived from the paper title (Unicode → ASCII, hyphenated, truncated to 80 chars at a word boundary).
130
143
 
131
- For legacy IDs containing `/` (e.g. `cs/0002001`), the slash is replaced with `-` in the directory name and file names (`cs-0002001-.../v1/cs-0002001v1.pdf`).
144
+ For legacy IDs containing `/` (e.g. `cs/0002001`), the slash is replaced with `-` in the directory name and file names (`cs-0002001-.../cs-0002001v1.pdf`).
132
145
 
133
146
  ## Library usage
134
147
 
@@ -138,7 +151,7 @@ require 'arxiv/downloader'
138
151
  identifier = Arxiv::Downloader::Identifier.new '2508.16190'
139
152
  client = Arxiv::Downloader::Client.new # 3-second rate limit by default
140
153
  path = Arxiv::Downloader::Archive.new(identifier, root: '/tmp/papers', client: client).run
141
- # => "/tmp/papers/2025/08/22/cs.CL/2508.16190-comicscene154-a-scene-dataset-for-comic-analysis/v1"
154
+ # => "/tmp/papers/2025/08/22/cs.CL/2508.16190-comicscene154-a-scene-dataset-for-comic-analysis"
142
155
  ```
143
156
 
144
157
  ## Development
@@ -9,12 +9,14 @@ module Arxiv
9
9
  @client = client
10
10
  end
11
11
 
12
- # Downloads into v<N>.partial/ and
13
- # renames it to v<N>/ only once everything succeeded.
14
- # An existing v<N>/ is always complete and is skipped.
12
+ # Downloads into <destination>.partial/ and
13
+ # renames it into place only once everything succeeded.
14
+ # An archived version is always complete and is skipped.
15
+ # PaperFolder decides flat (single version) vs v<N>/ (several versions).
15
16
  def run
16
- return paper_dir if Dir.exist? paper_dir
17
+ return paper_folder.location_of(metadata.version) if paper_folder.archived? metadata.version
17
18
 
19
+ paper_folder.unflatten!
18
20
  FileUtils.rm_rf staging_dir
19
21
  FileUtils.mkdir_p staging_dir
20
22
 
@@ -24,8 +26,8 @@ module Arxiv
24
26
  download_source_archive
25
27
  write_sidecars
26
28
 
27
- File.rename staging_dir, paper_dir
28
- paper_dir
29
+ File.rename staging_dir, destination
30
+ destination
29
31
  end
30
32
 
31
33
  private
@@ -44,13 +46,18 @@ module Arxiv
44
46
  @archived ||= Identifier.new "#{metadata.arxiv_id}v#{metadata.version}"
45
47
  end
46
48
 
47
- def paper_dir
48
- @paper_dir ||= File.join @root, Path.new(metadata).to_s, "v#{metadata.version}"
49
+ def paper_folder
50
+ @paper_folder ||= PaperFolder.new File.join(@root, Path.new(metadata).to_s)
49
51
  end
50
52
 
51
- # a sibling of paper_dir, so relative ../_shared/ refs survive the rename
53
+ # evaluated after unflatten!, which can turn a flat folder into v<N>/ folders
54
+ def destination
55
+ @destination ||= paper_folder.destination_for metadata.version
56
+ end
57
+
58
+ # a sibling of destination, so relative ../_shared/ refs survive the rename
52
59
  def staging_dir
53
- "#{paper_dir}.partial"
60
+ "#{destination}.partial"
54
61
  end
55
62
 
56
63
  def download_pdf
@@ -1,11 +1,5 @@
1
1
  module Arxiv
2
2
  module Downloader
3
- Author = Data.define :name, :affiliations do
4
- def initialize name:, affiliations: []
5
- super
6
- end
7
-
8
- def to_s = name
9
- end
3
+ Author = DL::Core::Author
10
4
  end
11
5
  end
@@ -1,106 +1,24 @@
1
- require 'optparse'
2
-
3
1
  module Arxiv
4
2
  module Downloader
5
- class CLI
6
- DEFAULT_DOWNLOAD_PATH = File.join Dir.home, 'Downloads', 'ArXiv_Papers'
7
- USAGE = 'Usage: arxiv-dl [options] <ARXIV_ID_OR_URL> [<ARXIV_ID_OR_URL>...]'.freeze
8
-
9
- def initialize argv, stderr: $stderr, stdin: $stdin, stdout: $stdout
10
- @argv = argv
11
- @stderr = stderr
12
- @stdin = stdin
13
- @stdout = stdout
14
- end
15
-
16
- def run
17
- options = parse
18
- return options.fetch(:exit_status) if options.key? :exit_status
19
-
20
- return error_with USAGE if options[:targets].empty?
21
- return error_with conflict if options[:verbose] && options[:quiet]
22
-
23
- failures = download_each options
24
- failures.zero? ? 0 : 1
25
- end
26
-
27
- private
28
-
29
- def conflict
30
- '-v and -q are mutually exclusive'
31
- end
32
-
33
- def parse
34
- options = { targets: [], verbose: false, quiet: false }
35
- parser = build_parser options
36
-
37
- begin
38
- parser.parse! @argv
39
- options[:targets] = @argv + input_targets(options[:input])
40
- rescue OptionParser::ParseError, SystemCallError => e
41
- @stderr.puts e.message
42
- return { exit_status: 1 }
43
- end
44
-
45
- options[:path] ||= ENV['ARXIV_DOWNLOAD_PATH'] || DEFAULT_DOWNLOAD_PATH
46
- options[:rate_limit] ||= (ENV['ARXIV_RATE_LIMIT'] || Client::DEFAULT_RATE_LIMIT).to_i
47
- options
48
- end
3
+ class CLI < DL::Core::CLI
4
+ def program = 'arxiv-dl'
49
5
 
50
- def build_parser options
51
- OptionParser.new do |parser|
52
- parser.banner = USAGE
53
- parser.on('-i FILE', '--input FILE') { |value| options[:input] = value }
54
- parser.on('-p PATH', '--path PATH') { |value| options[:path] = value }
55
- parser.on('--rate-limit SECONDS', Integer) { |value| options[:rate_limit] = value }
56
- parser.on('-v', '--verbose') { options[:verbose] = true }
57
- parser.on('-q', '--quiet') { options[:quiet] = true }
58
- parser.on('--version') do
59
- @stdout.puts VERSION
60
- options[:exit_status] = 0
61
- end
62
- parser.on('-h', '--help') do
63
- @stdout.puts parser.help
64
- options[:exit_status] = 0
65
- end
66
- end
67
- end
6
+ def target_name = 'ARXIV_ID_OR_URL'
68
7
 
69
- # one target per line from FILE, or stdin for "-"; blank lines and # comments skipped
70
- def input_targets input
71
- return [] if input.nil?
8
+ # ARXIV_DOWNLOAD_PATH, ARXIV_RATE_LIMIT
9
+ def env_prefix = 'ARXIV'
72
10
 
73
- text = input == '-' ? @stdin.read : File.read(input)
74
- text.lines.map(&:strip).reject { it.empty? || it.start_with?('#') }
75
- end
11
+ def default_path = File.join(Dir.home, 'Downloads', 'ArXiv_Papers')
76
12
 
77
- def error_with message
78
- @stderr.puts message
79
- 1
80
- end
13
+ def version = VERSION
81
14
 
82
- def download_each options
83
- client = Client.new rate_limit: options[:rate_limit], log: (options[:verbose] ? @stdout : nil)
15
+ def user_agent = Client::USER_AGENT
84
16
 
85
- failures = 0
86
- options[:targets].each do |target|
87
- failures += 1 unless download_one(target, client:, options:)
88
- end
89
- failures
90
- end
17
+ def client_for(rate_limit:, log:) = Client.new(rate_limit:, log:)
91
18
 
92
- # true on success; reports the failure and returns false otherwise
93
- def download_one target, client:, options:
94
- identifier = Identifier.new target
95
- @stdout.puts "==> Downloading #{identifier.id}" if options[:verbose]
19
+ def identifier_for(target) = Identifier.new(target)
96
20
 
97
- path = Archive.new(identifier, root: options[:path], client: client).run
98
- @stdout.puts path unless options[:quiet]
99
- true
100
- rescue Error, HTTP::Error => e
101
- @stderr.puts "#{target}: #{e.message}"
102
- false
103
- end
21
+ def archive_for(identifier, root:, client:) = Archive.new(identifier, root:, client:)
104
22
  end
105
23
  end
106
24
  end
@@ -1,85 +1,12 @@
1
- require 'http'
2
-
3
1
  module Arxiv
4
2
  module Downloader
5
- class Client
6
- SOURCE_URL = 'https://github.com/xoengineering/arxiv-dl'.freeze
7
- DEFAULT_RATE_LIMIT = 3
8
- TIMEOUTS = { connect: 10, read: 60, write: 10 }.freeze # seconds, per operation
9
- MAX_RETRIES = 3
10
- RETRY_BACKOFF = 10 # seconds before the first retry. doubles on each retry.
11
- RETRYABLE_STATUSES = [429, 503].freeze
12
-
13
- attr_reader :rate_limit
14
-
15
- def initialize rate_limit: DEFAULT_RATE_LIMIT, log: nil
16
- @rate_limit = rate_limit
17
- @log = log
18
- end
19
-
20
- def user_agent
21
- "arxiv-dl/#{VERSION} (+#{SOURCE_URL})"
22
- end
23
-
24
- def get url
25
- retries = 0
26
-
27
- loop do
28
- response = request url
29
- return response if response.status.success?
30
- raise http_error(url, response) unless retryable? response, retries
31
-
32
- retries += 1
33
- wait_before_retry response, retries
34
- end
35
- end
36
-
37
- private
38
-
39
- def request url
40
- throttle
41
- response = HTTP.timeout(TIMEOUTS).headers('User-Agent' => user_agent).follow.get(url)
42
- @last_request_at = Time.now
43
- log_request url, response
44
- response
45
- end
46
-
47
- def http_error url, response
48
- HTTPError.new status: response.status.code, url: url, reason: response.status.reason
49
- end
50
-
51
- def retryable? response, retries
52
- RETRYABLE_STATUSES.include?(response.status.code) && retries < MAX_RETRIES
53
- end
54
-
55
- def wait_before_retry response, retries
56
- seconds = retry_after(response) || (RETRY_BACKOFF * (2**(retries - 1)))
57
- @log&.puts "==> #{response.status}. Retrying in #{seconds}s"
58
- sleep seconds
59
- end
60
-
61
- # Retry-After in delay-seconds form. The HTTP-date form falls back to backoff.
62
- def retry_after response
63
- value = response.headers['Retry-After']
64
- return if value.nil?
65
-
66
- Integer(value, exception: false)
67
- end
68
-
69
- def log_request url, response
70
- return if @log.nil?
71
-
72
- @log.puts "==> GET #{url} (#{response.body.to_s.bytesize} bytes)"
73
- end
74
-
75
- def throttle
76
- return if @rate_limit.zero?
77
- return if @last_request_at.nil?
78
-
79
- elapsed = Time.now - @last_request_at
80
- return if elapsed >= @rate_limit
3
+ # DL::Core::Client with arxiv-dl's User-Agent filled in
4
+ class Client < DL::Core::Client
5
+ SOURCE_URL = 'https://github.com/xoengineering/arxiv-dl'.freeze
6
+ USER_AGENT = "arxiv-dl/#{VERSION} (+#{SOURCE_URL})".freeze
81
7
 
82
- sleep(@rate_limit - elapsed)
8
+ def initialize log: nil, rate_limit: DEFAULT_RATE_LIMIT
9
+ super(user_agent: USER_AGENT, log:, rate_limit:)
83
10
  end
84
11
  end
85
12
  end
@@ -1,5 +1,5 @@
1
1
  module Arxiv
2
2
  module Downloader
3
- class Error < StandardError; end
3
+ class Error < DL::Core::Error; end
4
4
  end
5
5
  end
@@ -1,14 +1,5 @@
1
1
  module Arxiv
2
2
  module Downloader
3
- class HTTPError < Error
4
- attr_reader :status, :url
5
-
6
- def initialize status:, url:, reason: nil
7
- @status = status
8
- @url = url
9
-
10
- super("GET #{url} failed: #{[status, reason].compact.join ' '}")
11
- end
12
- end
3
+ HTTPError = DL::Core::HTTPError
13
4
  end
14
5
  end
@@ -1,33 +1,7 @@
1
- require 'fileutils'
2
- require 'json'
3
-
4
1
  module Arxiv
5
2
  module Downloader
6
3
  class Metadata
7
- class JSON
8
- FILENAME = 'metadata.json'.freeze
9
-
10
- def initialize metadata
11
- @metadata = metadata
12
- end
13
-
14
- def write to:
15
- FileUtils.mkdir_p to
16
- File.write File.join(to, FILENAME), "#{::JSON.pretty_generate(serialize(@metadata.to_h))}\n"
17
- end
18
-
19
- private
20
-
21
- def serialize object
22
- case object
23
- when Hash then object.to_h { |key, value| [key.to_s, serialize(value)] }
24
- when Array then object.map { |item| serialize item }
25
- when Data then serialize object.to_h
26
- when Date, Time then object.iso8601
27
- else object
28
- end
29
- end
30
- end
4
+ JSON = DL::Core::Sidecar::JSON
31
5
  end
32
6
  end
33
7
  end
@@ -1,44 +1,22 @@
1
- require 'fileutils'
2
- require 'yaml'
3
-
4
1
  module Arxiv
5
2
  module Downloader
6
3
  class Metadata
4
+ # metadata.md: dl-core writes the frontmatter; the body is arxiv-specific
7
5
  class Markdown
8
- FILENAME = 'metadata.md'.freeze
9
-
10
6
  def initialize metadata
11
7
  @metadata = metadata
12
8
  end
13
9
 
14
10
  def write to:
15
- FileUtils.mkdir_p to
16
- File.write File.join(to, FILENAME), "#{frontmatter}\n#{body}"
11
+ DL::Core::Sidecar::Markdown.new(@metadata, body:, extras: { bibtex_key: }).write to:
17
12
  end
18
13
 
19
14
  private
20
15
 
21
- def frontmatter
22
- "---\n#{::YAML.dump(stringify(frontmatter_hash)).delete_prefix("---\n")}---"
23
- end
24
-
25
- def frontmatter_hash
26
- @metadata.to_h.merge bibtex_key: bibtex_key
27
- end
28
-
29
16
  def bibtex_key
30
17
  Downloader::Bibtex.new(@metadata).key
31
18
  end
32
19
 
33
- def stringify object
34
- case object
35
- when Hash then object.to_h { |key, value| [key.to_s, stringify(value)] }
36
- when Array then object.map { |item| stringify item }
37
- when Data then stringify object.to_h
38
- else object
39
- end
40
- end
41
-
42
20
  def body
43
21
  <<~MARKDOWN
44
22
 
@@ -1,32 +1,7 @@
1
- require 'fileutils'
2
- require 'yaml'
3
-
4
1
  module Arxiv
5
2
  module Downloader
6
3
  class Metadata
7
- class YAML
8
- FILENAME = 'metadata.yaml'.freeze
9
-
10
- def initialize metadata
11
- @metadata = metadata
12
- end
13
-
14
- def write to:
15
- FileUtils.mkdir_p to
16
- File.write File.join(to, FILENAME), ::YAML.dump(stringify(@metadata.to_h))
17
- end
18
-
19
- private
20
-
21
- def stringify object
22
- case object
23
- when Hash then object.to_h { |key, value| [key.to_s, stringify(value)] }
24
- when Array then object.map { |item| stringify item }
25
- when Data then stringify object.to_h
26
- else object
27
- end
28
- end
29
- end
4
+ YAML = DL::Core::Sidecar::YAML
30
5
  end
31
6
  end
32
7
  end
@@ -0,0 +1,85 @@
1
+ require 'fileutils'
2
+ require 'nokogiri'
3
+ require 'yaml'
4
+
5
+ module Arxiv
6
+ module Downloader
7
+ # One paper's folder. A paper with a single archived version is kept flat
8
+ # (files directly in the folder); once a second version arrives, each
9
+ # version gets its own v<N>/ folder.
10
+ class PaperFolder
11
+ VERSION_FOLDER = /\Av\d+(\.partial)?\z/
12
+
13
+ attr_reader :path
14
+
15
+ def initialize path
16
+ @path = path
17
+ end
18
+
19
+ def archived? version
20
+ Dir.exist? location_of(version)
21
+ end
22
+
23
+ def location_of version
24
+ return path if flat_version == version
25
+
26
+ version_path version
27
+ end
28
+
29
+ # where a new download of `version` goes; call unflatten! first if the folder is flat
30
+ def destination_for version
31
+ return path if version == 1 && !versioned?
32
+
33
+ version_path version
34
+ end
35
+
36
+ # moves the flat version into v<N>/ so another version can sit beside it
37
+ def unflatten!
38
+ return if flat_version.nil?
39
+
40
+ destination = version_path flat_version
41
+ FileUtils.mkdir_p destination
42
+
43
+ # metadata.yaml marks the folder as flat, so it moves last
44
+ children = Dir.children(path) - [File.basename(destination), metadata_filename]
45
+ children.each { FileUtils.mv File.join(path, it), destination }
46
+ deepen_shared_links File.join(destination, 'html')
47
+ FileUtils.mv File.join(path, metadata_filename), destination
48
+ end
49
+
50
+ private
51
+
52
+ def flat_version
53
+ metadata_path = File.join path, metadata_filename
54
+ return unless File.exist? metadata_path
55
+
56
+ ::YAML.safe_load_file(metadata_path, permitted_classes: [Date]).fetch 'version'
57
+ end
58
+
59
+ def versioned?
60
+ Dir.exist?(path) && Dir.children(path).any? { VERSION_FOLDER.match? it }
61
+ end
62
+
63
+ def version_path version
64
+ File.join path, "v#{version}"
65
+ end
66
+
67
+ def metadata_filename
68
+ Metadata::YAML::FILENAME
69
+ end
70
+
71
+ # html/ moved one level deeper, so relative links up into _shared/ need one more ../
72
+ def deepen_shared_links html_dir
73
+ Dir.glob(File.join(html_dir, '*.html')).each do |html_path|
74
+ document = Nokogiri::HTML File.read(html_path)
75
+ HTMLArchive::ASSET_SELECTORS.each do |selector, attribute|
76
+ document.css(selector).each do |node|
77
+ node[attribute] = "../#{node[attribute]}" if node[attribute].start_with? '../'
78
+ end
79
+ end
80
+ File.write html_path, document.to_html
81
+ end
82
+ end
83
+ end
84
+ end
85
+ end
@@ -1,44 +1,5 @@
1
- require 'stringex'
2
-
3
1
  module Arxiv
4
2
  module Downloader
5
- class Slug
6
- MAX_LENGTH = 80
7
-
8
- TEX_INLINE_MATH = /\$[^$]*\$/ # $...$
9
- TEX_DISPLAY_MATH = /\\\(.*?\\\)|\\\[.*?\\\]/m # \(...\) or \[...\]
10
- TEX_COMMAND = /\\[a-zA-Z]+\*?/ # \emph, \alpha, etc.
11
-
12
- def initialize title
13
- @title = title
14
- end
15
-
16
- def to_s
17
- truncate strip_tex(@title).to_url
18
- end
19
-
20
- private
21
-
22
- def strip_tex string
23
- string
24
- .gsub(TEX_INLINE_MATH, ' ')
25
- .gsub(TEX_DISPLAY_MATH, ' ')
26
- .gsub(TEX_COMMAND, ' ')
27
- end
28
-
29
- def truncate slug
30
- return slug if slug.length <= MAX_LENGTH
31
-
32
- words = slug.split '-'
33
- result = +''
34
- words.each do |word|
35
- break if result.length + 1 + word.length > MAX_LENGTH
36
-
37
- result << '-' unless result.empty?
38
- result << word
39
- end
40
- result
41
- end
42
- end
3
+ Slug = DL::Core::Slug
43
4
  end
44
5
  end
@@ -1,5 +1,5 @@
1
1
  module Arxiv
2
2
  module Downloader
3
- VERSION = '0.2.0'.freeze
3
+ VERSION = '0.3.1'.freeze
4
4
  end
5
5
  end
@@ -1,3 +1,7 @@
1
+ require 'dl/core'
2
+
3
+ require_relative 'downloader/version' # before client: Client::USER_AGENT uses VERSION
4
+
1
5
  require_relative 'downloader/abstract_page'
2
6
  require_relative 'downloader/archive'
3
7
  require_relative 'downloader/assets_cache'
@@ -5,23 +9,23 @@ require_relative 'downloader/author'
5
9
  require_relative 'downloader/bibtex'
6
10
  require_relative 'downloader/categories'
7
11
  require_relative 'downloader/cli'
8
- require_relative 'downloader/client'
12
+ require_relative 'downloader/client' # after version
9
13
  require_relative 'downloader/error' # before errors below that subclass Error
10
14
  require_relative 'downloader/feed_parser'
11
15
  require_relative 'downloader/html_archive'
12
- require_relative 'downloader/http_error' # after error
16
+ require_relative 'downloader/http_error'
13
17
  require_relative 'downloader/identifier' # after error
14
18
  require_relative 'downloader/metadata' # before metadata/*: they reopen class Metadata
15
19
  require_relative 'downloader/metadata/bibtex' # after metadata
16
20
  require_relative 'downloader/metadata/json' # after metadata
17
21
  require_relative 'downloader/metadata/markdown' # after metadata
18
22
  require_relative 'downloader/metadata/yaml' # after metadata
23
+ require_relative 'downloader/paper_folder'
19
24
  require_relative 'downloader/paper_not_found' # after error
20
25
  require_relative 'downloader/path'
21
26
  require_relative 'downloader/pdf'
22
27
  require_relative 'downloader/slug'
23
28
  require_relative 'downloader/source_archive'
24
- require_relative 'downloader/version'
25
29
 
26
30
  module Arxiv
27
31
  module Downloader
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: arxiv-dl
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.2.0
4
+ version: 0.3.1
5
5
  platform: ruby
6
6
  authors:
7
7
  - Shane Becker
@@ -10,33 +10,33 @@ cert_chain: []
10
10
  date: 1980-01-02 00:00:00.000000000 Z
11
11
  dependencies:
12
12
  - !ruby/object:Gem::Dependency
13
- name: feedjira
13
+ name: dl-core
14
14
  requirement: !ruby/object:Gem::Requirement
15
15
  requirements:
16
16
  - - "~>"
17
17
  - !ruby/object:Gem::Version
18
- version: '4.0'
18
+ version: '0.1'
19
19
  type: :runtime
20
20
  prerelease: false
21
21
  version_requirements: !ruby/object:Gem::Requirement
22
22
  requirements:
23
23
  - - "~>"
24
24
  - !ruby/object:Gem::Version
25
- version: '4.0'
25
+ version: '0.1'
26
26
  - !ruby/object:Gem::Dependency
27
- name: http
27
+ name: feedjira
28
28
  requirement: !ruby/object:Gem::Requirement
29
29
  requirements:
30
30
  - - "~>"
31
31
  - !ruby/object:Gem::Version
32
- version: '6.0'
32
+ version: '4.0'
33
33
  type: :runtime
34
34
  prerelease: false
35
35
  version_requirements: !ruby/object:Gem::Requirement
36
36
  requirements:
37
37
  - - "~>"
38
38
  - !ruby/object:Gem::Version
39
- version: '6.0'
39
+ version: '4.0'
40
40
  - !ruby/object:Gem::Dependency
41
41
  name: nokogiri
42
42
  requirement: !ruby/object:Gem::Requirement
@@ -51,34 +51,6 @@ dependencies:
51
51
  - - "~>"
52
52
  - !ruby/object:Gem::Version
53
53
  version: '1.19'
54
- - !ruby/object:Gem::Dependency
55
- name: ostruct
56
- requirement: !ruby/object:Gem::Requirement
57
- requirements:
58
- - - "~>"
59
- - !ruby/object:Gem::Version
60
- version: '0.6'
61
- type: :runtime
62
- prerelease: false
63
- version_requirements: !ruby/object:Gem::Requirement
64
- requirements:
65
- - - "~>"
66
- - !ruby/object:Gem::Version
67
- version: '0.6'
68
- - !ruby/object:Gem::Dependency
69
- name: stringex
70
- requirement: !ruby/object:Gem::Requirement
71
- requirements:
72
- - - "~>"
73
- - !ruby/object:Gem::Version
74
- version: '2.8'
75
- type: :runtime
76
- prerelease: false
77
- version_requirements: !ruby/object:Gem::Requirement
78
- requirements:
79
- - - "~>"
80
- - !ruby/object:Gem::Version
81
- version: '2.8'
82
54
  description: Command line tool and Ruby library for archiving arxiv.org papers as
83
55
  PDFs with sidecar metadata.
84
56
  email:
@@ -116,6 +88,7 @@ files:
116
88
  - lib/arxiv/downloader/metadata/json.rb
117
89
  - lib/arxiv/downloader/metadata/markdown.rb
118
90
  - lib/arxiv/downloader/metadata/yaml.rb
91
+ - lib/arxiv/downloader/paper_folder.rb
119
92
  - lib/arxiv/downloader/paper_not_found.rb
120
93
  - lib/arxiv/downloader/path.rb
121
94
  - lib/arxiv/downloader/pdf.rb