word-to-markdown 1.1.8 → 1.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.md +31 -10
- data/bin/w2m +26 -8
- data/lib/cliver/dependency_ext.rb +4 -3
- data/lib/word-to-markdown/converter.rb +35 -10
- data/lib/word-to-markdown/document.rb +29 -10
- data/lib/word-to-markdown/input_format.rb +86 -0
- data/lib/word-to-markdown/url_scrubber.rb +110 -0
- data/lib/word-to-markdown/version.rb +1 -1
- data/lib/word-to-markdown.rb +69 -8
- metadata +74 -40
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: d5fff6778e90bf76409d47c5177fd2400a14308cf40be3ea2e3ee118917945b0
|
|
4
|
+
data.tar.gz: 18062303949a156da4c940daeb79854c820036fe623cdb9c4c07b1bce9dca242
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: f041232dbf6ae3aee4738d9562fab3f820a492dfd8d71fdd257397b278ca03550a27a59a0d0b5c3514254bab49367ba067970f7551ea11c3a7b90b8eda434499
|
|
7
|
+
data.tar.gz: e8c267ccd410ccfc978ef873805696299b1a658d1894dccae2f63d39f8afd12595ced23fbce6c8f4173188cc0b1e6fa44957d37c29c32b8ea18c08fb9a8745e1
|
data/README.md
CHANGED
|
@@ -1,24 +1,29 @@
|
|
|
1
1
|
# Word to Markdown converter
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
> [!IMPORTANT]
|
|
4
|
+
> **Looking for the latest and greatest?** Check out [**word-to-markdown-js**](https://github.com/benbalter/word-to-markdown-js), the newer, better successor to this project. It's a modern, actively maintained rewrite and is recommended for new projects. This Ruby gem remains available for existing users.
|
|
4
5
|
|
|
5
|
-
|
|
6
|
+
A Ruby gem to liberate content from [the jail that is Word documents](https://ben.balter.com/2012/10/19/we-ve-been-trained-to-make-paper/#jailbreaking-content)
|
|
7
|
+
|
|
8
|
+
[](https://github.com/benbalter/word-to-markdown/actions/workflows/ci.yml) [](https://rubygems.org/gems/word-to-markdown)
|
|
6
9
|
|
|
7
10
|
## The problem
|
|
8
11
|
|
|
9
|
-
> Our default content publishing workflow is terribly broken. [We've all been trained to make paper](
|
|
12
|
+
> Our default content publishing workflow is terribly broken. [We've all been trained to make paper](https://ben.balter.com/2012/10/19/we-ve-been-trained-to-make-paper/), yet today, content authored once is more commonly consumed in multiple formats, and rarely, if ever, does it embody physical form. Put another way, our go-to content authoring workflow remains relatively unchanged since it was conceived in the early 80s.
|
|
10
13
|
>
|
|
11
|
-
> I'm asked regularly by government employees — knowledge workers who fire up a desktop word processor as the first step to any project — for an automated pipeline to convert Microsoft Word documents to [Markdown](
|
|
14
|
+
> I'm asked regularly by government employees — knowledge workers who fire up a desktop word processor as the first step to any project — for an automated pipeline to convert Microsoft Word documents to [Markdown](https://docs.github.com/en/get-started/writing-on-github/getting-started-with-writing-and-formatting-on-github/basic-writing-and-formatting-syntax), the *lingua franca* of the internet, but as my recent foray into building [just such a converter](https://word2md.com/) proves, it's not that simple.
|
|
12
15
|
>
|
|
13
16
|
> Markdown isn't just an alternative format. Markdown forces you to write for the web.
|
|
14
17
|
|
|
15
|
-
**[Read more](
|
|
18
|
+
**[Read more](https://ben.balter.com/2014/03/31/word-versus-markdown-more-than-mere-semantics/)**
|
|
19
|
+
|
|
20
|
+
## Just want to convert a Microsoft Word (or Google) document to Markdown?
|
|
16
21
|
|
|
17
|
-
**[
|
|
22
|
+
You can use this **[hosted service](https://word2md.com/)** (or check out [its source](https://github.com/benbalter/word-to-markdown-server)).
|
|
18
23
|
|
|
19
24
|
## Install
|
|
20
25
|
|
|
21
|
-
You'll need to install [LibreOffice](
|
|
26
|
+
You'll need to install [LibreOffice](https://www.libreoffice.org/). Then:
|
|
22
27
|
|
|
23
28
|
```bash
|
|
24
29
|
gem install word-to-markdown
|
|
@@ -65,14 +70,30 @@ $ w2m path/to/document.docx
|
|
|
65
70
|
|
|
66
71
|
Word-to-markdown requires `soffice` a command line interface to LibreOffice that works on Linux, Mac, and Windows. To install soffice, see [the LibreOffice documentation](https://www.libreoffice.org/get-help/install-howto/).
|
|
67
72
|
|
|
73
|
+
Word-to-markdown only accepts Word documents (`.docx` and `.doc`), identified by their contents rather than their file extension. Other files, including those LibreOffice could otherwise open, such as HTML, ODT, or RTF, raise `WordToMarkdown::Document::UnsupportedFormatError`.
|
|
74
|
+
|
|
75
|
+
LibreOffice is killed if a conversion takes longer than 60 seconds, raising `WordToMarkdown::TimeoutError`. To change the limit, set `WordToMarkdown.timeout = 120` or the `WORD_TO_MARKDOWN_TIMEOUT` environment variable.
|
|
76
|
+
|
|
77
|
+
### Converting untrusted documents
|
|
78
|
+
|
|
79
|
+
Word documents can link to remote resources, such as images, which LibreOffice fetches while converting the document. If you convert documents from untrusted sources (for example, files uploaded to a web service), run the conversion without network access, such as in a container or sandbox with no outbound network, so that a document can't make requests to internal services or other hosts on your behalf.
|
|
80
|
+
|
|
68
81
|
## Testing
|
|
69
82
|
|
|
70
83
|
```
|
|
71
84
|
script/cibuild
|
|
72
85
|
```
|
|
73
86
|
|
|
74
|
-
##
|
|
87
|
+
## Docker
|
|
88
|
+
|
|
89
|
+
Everything you need to run the executable locally:
|
|
90
|
+
|
|
91
|
+
```
|
|
92
|
+
docker compose build
|
|
93
|
+
docker compose run --rm app bundle exec w2m --help
|
|
94
|
+
docker compose run --rm app bundle exec w2m test/fixtures/em.docx
|
|
95
|
+
```
|
|
75
96
|
|
|
76
|
-
|
|
97
|
+
## Hosted service
|
|
77
98
|
|
|
78
|
-
A live version runs at [
|
|
99
|
+
[Word-to-markdown-server](https://github.com/benbalter/word-to-markdown-server) contains a lightweight server for converting Word Documents as a service. A live version runs at [word2md.com](https://word2md.com).
|
data/bin/w2m
CHANGED
|
@@ -1,17 +1,35 @@
|
|
|
1
1
|
#!/usr/bin/env ruby
|
|
2
2
|
# frozen_string_literal: true
|
|
3
3
|
|
|
4
|
+
require 'optparse'
|
|
4
5
|
require 'word-to-markdown'
|
|
5
6
|
|
|
6
|
-
|
|
7
|
-
|
|
7
|
+
parser = OptionParser.new do |opts|
|
|
8
|
+
opts.banner = 'Usage: w2m path/to/document.docx'
|
|
9
|
+
|
|
10
|
+
opts.on('-h', '--help', 'Show this help') do
|
|
11
|
+
puts opts
|
|
12
|
+
exit
|
|
13
|
+
end
|
|
14
|
+
|
|
15
|
+
opts.on('-v', '--version', 'Show the WordToMarkdown and LibreOffice versions') do
|
|
16
|
+
puts "WordToMarkdown v#{WordToMarkdown::VERSION}"
|
|
17
|
+
puts "LibreOffice v#{WordToMarkdown.soffice.version}" unless Gem.win_platform?
|
|
18
|
+
exit
|
|
19
|
+
end
|
|
20
|
+
end
|
|
21
|
+
|
|
22
|
+
begin
|
|
23
|
+
parser.parse!
|
|
24
|
+
rescue OptionParser::ParseError => e
|
|
25
|
+
warn e.message
|
|
26
|
+
warn parser
|
|
8
27
|
exit 1
|
|
9
28
|
end
|
|
10
29
|
|
|
11
|
-
if ARGV
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
else
|
|
15
|
-
doc = WordToMarkdown.new ARGV[0]
|
|
16
|
-
puts doc.to_s
|
|
30
|
+
if ARGV.size != 1
|
|
31
|
+
warn parser
|
|
32
|
+
exit 1
|
|
17
33
|
end
|
|
34
|
+
|
|
35
|
+
puts WordToMarkdown.new(ARGV[0])
|
|
@@ -24,14 +24,15 @@ module Cliver
|
|
|
24
24
|
|
|
25
25
|
# Returns the version of the resolved dependency
|
|
26
26
|
def version
|
|
27
|
-
return @
|
|
27
|
+
return @version if defined? @version
|
|
28
28
|
return if Gem.win_platform?
|
|
29
|
+
|
|
29
30
|
version = installed_versions.find { |p, _v| p == path }
|
|
30
|
-
@
|
|
31
|
+
@version = version.nil? ? nil : version[1]
|
|
31
32
|
end
|
|
32
33
|
|
|
33
34
|
def major_version
|
|
34
|
-
version
|
|
35
|
+
version&.split('.')&.first
|
|
35
36
|
end
|
|
36
37
|
end
|
|
37
38
|
end
|
|
@@ -7,14 +7,22 @@ class WordToMarkdown
|
|
|
7
7
|
# Number of headings to guess, e.g., h6
|
|
8
8
|
HEADING_DEPTH = 6
|
|
9
9
|
|
|
10
|
-
# Percentile step for
|
|
10
|
+
# Percentile step for each heading
|
|
11
11
|
HEADING_STEP = 100 / HEADING_DEPTH
|
|
12
12
|
|
|
13
13
|
# Minimum heading size
|
|
14
14
|
MIN_HEADING_SIZE = 20
|
|
15
15
|
|
|
16
16
|
# Unicode bullets to strip when processing
|
|
17
|
-
UNICODE_BULLETS = ['○', '
|
|
17
|
+
UNICODE_BULLETS = ['○', '●', "\u2022", '\\p{C}'].freeze
|
|
18
|
+
|
|
19
|
+
# Leading bullets to strip from list items. A plain "o" only counts as a
|
|
20
|
+
# bullet when followed by whitespace, so words like "orange" survive.
|
|
21
|
+
BULLET_REGEX = /\A(?:[#{UNICODE_BULLETS.join}]|o(?=[[:space:]]))+[[:space:]]*/
|
|
22
|
+
|
|
23
|
+
# Leading list numbering to strip, e.g., "1.", "a.", or "iv.", along with
|
|
24
|
+
# any whitespace (including non-breaking spaces) that follows it
|
|
25
|
+
NUMBERING_REGEX = /\A(?:\d+|[a-zA-Z]|[ivxlcdm]+|[IVXLCDM]+)\.(?:[[:space:]]+|\z)/
|
|
18
26
|
|
|
19
27
|
# @param document [WordToMarkdown::Document] The document to convert
|
|
20
28
|
def initialize(document)
|
|
@@ -64,10 +72,11 @@ class WordToMarkdown
|
|
|
64
72
|
|
|
65
73
|
# Given a Nokogiri node, guess what heading it represents, if any
|
|
66
74
|
#
|
|
67
|
-
# @param node [
|
|
75
|
+
# @param node [Nokogiri::Node] the nokogiri node
|
|
68
76
|
# @return [String, nil] the heading tag (e.g., H1), or nil
|
|
69
77
|
def guess_heading(node)
|
|
70
78
|
return nil if node.font_size.nil?
|
|
79
|
+
|
|
71
80
|
[*1...HEADING_DEPTH].each do |heading|
|
|
72
81
|
return "h#{heading}" if node.font_size >= h(heading)
|
|
73
82
|
end
|
|
@@ -81,7 +90,23 @@ class WordToMarkdown
|
|
|
81
90
|
#
|
|
82
91
|
# @return [Integer] the minimum font size
|
|
83
92
|
def h(num)
|
|
84
|
-
|
|
93
|
+
self.class.percentile(font_sizes, ((HEADING_DEPTH - 1) - num) * HEADING_STEP)
|
|
94
|
+
end
|
|
95
|
+
|
|
96
|
+
# Linearly interpolated percentile, matching the algorithm used by the
|
|
97
|
+
# descriptive_statistics gem (v2.5.1) that this replaces
|
|
98
|
+
#
|
|
99
|
+
# @param values [Array<Numeric>] the values
|
|
100
|
+
# @param pct [Numeric] the percentile, from 0 to 100
|
|
101
|
+
#
|
|
102
|
+
# @return [Float, nil] the percentile, or nil if values is empty
|
|
103
|
+
def self.percentile(values, pct)
|
|
104
|
+
sorted = values.map(&:to_f).sort
|
|
105
|
+
rank = pct / 100.0 * (sorted.size - 1)
|
|
106
|
+
lower, upper = sorted[rank.floor, 2]
|
|
107
|
+
return lower if upper.nil? # empty or single-element collection, or the 100th percentile
|
|
108
|
+
|
|
109
|
+
lower + ((upper - lower) * (rank - rank.floor))
|
|
85
110
|
end
|
|
86
111
|
|
|
87
112
|
# Convert span-based font styles to `strong`s and `em`s
|
|
@@ -109,7 +134,7 @@ class WordToMarkdown
|
|
|
109
134
|
def remove_unicode_bullets_from_list_items!
|
|
110
135
|
path = WordToMarkdown.soffice.major_version == '5' ? 'li span span' : 'li span'
|
|
111
136
|
@document.tree.search(path).each do |span|
|
|
112
|
-
span.inner_html = span.inner_html.
|
|
137
|
+
span.inner_html = span.inner_html.sub(BULLET_REGEX, '')
|
|
113
138
|
end
|
|
114
139
|
end
|
|
115
140
|
|
|
@@ -117,21 +142,21 @@ class WordToMarkdown
|
|
|
117
142
|
def remove_numbering_from_list_items!
|
|
118
143
|
path = WordToMarkdown.soffice.major_version == '5' ? 'li span span' : 'li span'
|
|
119
144
|
@document.tree.search(path).each do |span|
|
|
120
|
-
span.inner_html = span.inner_html.
|
|
145
|
+
span.inner_html = span.inner_html.sub(NUMBERING_REGEX, '')
|
|
121
146
|
end
|
|
122
147
|
end
|
|
123
148
|
|
|
124
|
-
#
|
|
149
|
+
# Remove whitespace from list items
|
|
125
150
|
def remove_whitespace_from_list_items!
|
|
126
|
-
@document.tree.search('li span').each { |span| span.inner_html.strip
|
|
151
|
+
@document.tree.search('li span').each { |span| span.inner_html = span.inner_html.strip }
|
|
127
152
|
end
|
|
128
153
|
|
|
129
|
-
# Convert table headers to `th`
|
|
154
|
+
# Convert table headers to `th`s
|
|
130
155
|
def semanticize_table_headers!
|
|
131
156
|
@document.tree.search('table tr:first td').each { |node| node.node_name = 'th' }
|
|
132
157
|
end
|
|
133
158
|
|
|
134
|
-
# Try to guess heading where implicit
|
|
159
|
+
# Try to guess heading where implicit based on font size
|
|
135
160
|
def semanticize_headings!
|
|
136
161
|
implicit_headings.each do |element|
|
|
137
162
|
heading = guess_heading element
|
|
@@ -3,16 +3,24 @@
|
|
|
3
3
|
class WordToMarkdown
|
|
4
4
|
class Document
|
|
5
5
|
class NotFoundError < StandardError; end
|
|
6
|
+
|
|
6
7
|
class ConversionError < StandardError; end
|
|
7
8
|
|
|
8
|
-
|
|
9
|
+
class UnsupportedFormatError < StandardError; end
|
|
10
|
+
|
|
11
|
+
attr_reader :path, :tmpdir, :import_filter
|
|
9
12
|
|
|
10
13
|
# @param path [string] Path to the Word document
|
|
11
14
|
# @param tmpdir [string] Path to a working directory to use
|
|
12
15
|
def initialize(path, tmpdir = nil)
|
|
13
16
|
@path = File.expand_path path, Dir.pwd
|
|
14
|
-
@tmpdir = tmpdir || Dir.mktmpdir
|
|
15
17
|
raise NotFoundError, "File #{@path} does not exist" unless File.exist?(@path)
|
|
18
|
+
|
|
19
|
+
@import_filter = InputFormat.filter_for(@path)
|
|
20
|
+
raise UnsupportedFormatError, "File #{@path} is not a Word document (.docx or .doc)" if @import_filter.nil?
|
|
21
|
+
|
|
22
|
+
@own_tmpdir = tmpdir.nil?
|
|
23
|
+
@tmpdir = tmpdir || Dir.mktmpdir
|
|
16
24
|
end
|
|
17
25
|
|
|
18
26
|
# @return [String] the document's extension
|
|
@@ -20,7 +28,7 @@ class WordToMarkdown
|
|
|
20
28
|
File.extname path
|
|
21
29
|
end
|
|
22
30
|
|
|
23
|
-
# @return [
|
|
31
|
+
# @return [Nokogiri::Document]
|
|
24
32
|
def tree
|
|
25
33
|
@tree ||= begin
|
|
26
34
|
tree = Nokogiri::HTML(normalized_html)
|
|
@@ -31,6 +39,7 @@ class WordToMarkdown
|
|
|
31
39
|
|
|
32
40
|
# @return [String] the html representation of the document
|
|
33
41
|
def html
|
|
42
|
+
UrlScrubber.scrub!(tree)
|
|
34
43
|
tree.to_html.gsub("</li>\n", '</li>')
|
|
35
44
|
end
|
|
36
45
|
|
|
@@ -44,7 +53,7 @@ class WordToMarkdown
|
|
|
44
53
|
#
|
|
45
54
|
# @return [String] the encoding, defaulting to "UTF-8"
|
|
46
55
|
def encoding
|
|
47
|
-
match = raw_html.encode('UTF-8', invalid: :replace, replace: '').match(/charset=([
|
|
56
|
+
match = raw_html.encode('UTF-8', invalid: :replace, replace: '').match(/charset=([^"]+)/)
|
|
48
57
|
if match
|
|
49
58
|
match[1].sub('macintosh', 'MacRoman')
|
|
50
59
|
else
|
|
@@ -78,30 +87,40 @@ class WordToMarkdown
|
|
|
78
87
|
string.gsub!(' ', ' ') # HTML encoded spaces
|
|
79
88
|
string.sub!(/\A[[:space:]]+/, '') # document leading whitespace
|
|
80
89
|
string.sub!(/[[:space:]]+\z/, '') # document trailing whitespace
|
|
81
|
-
string.gsub!(/(
|
|
82
|
-
string.gsub!(
|
|
90
|
+
string.gsub!(/( +)$/, '') # line trailing whitespace
|
|
91
|
+
string.gsub!("\n\n\n\n", "\n\n") # Quadruple line breaks
|
|
83
92
|
string.delete!(' ') # Unicode non-breaking spaces, injected as tabs
|
|
93
|
+
string.gsub!(/\*\*\ +(?!\*|_)([[:punct:]])/, '**\1') # Remove extra space after bold
|
|
84
94
|
string
|
|
85
95
|
end
|
|
86
96
|
|
|
87
97
|
# @return [String] the path to the intermediary HTML document
|
|
88
98
|
def dest_path
|
|
89
|
-
|
|
90
|
-
File.expand_path(dest_filename, tmpdir)
|
|
99
|
+
File.expand_path("#{File.basename(path, '.*')}.html", tmpdir)
|
|
91
100
|
end
|
|
92
101
|
|
|
93
102
|
# @return [String] the unnormalized HTML representation
|
|
94
103
|
def raw_html
|
|
95
104
|
@raw_html ||= begin
|
|
96
|
-
WordToMarkdown.run_command '--headless', '--convert-to', filter, path, '--outdir', tmpdir
|
|
105
|
+
WordToMarkdown.run_command '--headless', "--infilter=#{import_filter}", '--convert-to', filter, path, '--outdir', tmpdir
|
|
97
106
|
raise ConversionError, "Failed to convert #{path}" unless File.exist?(dest_path)
|
|
107
|
+
|
|
98
108
|
html = File.read dest_path
|
|
99
109
|
File.delete dest_path
|
|
100
110
|
html
|
|
111
|
+
ensure
|
|
112
|
+
remove_tmpdir
|
|
101
113
|
end
|
|
102
114
|
end
|
|
103
115
|
|
|
104
|
-
#
|
|
116
|
+
# Remove the working directory if we created it and nothing else is in it.
|
|
117
|
+
# Non-empty directories are kept, since LibreOffice may have written
|
|
118
|
+
# images there that the markdown references.
|
|
119
|
+
def remove_tmpdir
|
|
120
|
+
Dir.rmdir(tmpdir) if @own_tmpdir && Dir.empty?(tmpdir)
|
|
121
|
+
end
|
|
122
|
+
|
|
123
|
+
# @return [String] the LibreOffice filter to use for export
|
|
105
124
|
def filter
|
|
106
125
|
if WordToMarkdown.soffice.major_version == '5'
|
|
107
126
|
'html:XHTML Writer File:UTF8'
|
|
@@ -0,0 +1,86 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
class WordToMarkdown
|
|
4
|
+
# Identifies Word documents by their content, rather than their extension,
|
|
5
|
+
# so that other formats LibreOffice would otherwise sniff and import (e.g.,
|
|
6
|
+
# HTML, which can reference remote resources) are never passed to it
|
|
7
|
+
module InputFormat
|
|
8
|
+
# Signature of an OLE compound file, used by .doc files
|
|
9
|
+
OLE_SIGNATURE = "\xD0\xCF\x11\xE0\xA1\xB1\x1A\xE1".b.freeze
|
|
10
|
+
|
|
11
|
+
# Signature of a ZIP local file header, used by .docx files
|
|
12
|
+
ZIP_SIGNATURE = "PK\x03\x04".b.freeze
|
|
13
|
+
|
|
14
|
+
# Signature of a ZIP end of central directory record
|
|
15
|
+
ZIP_EOCD_SIGNATURE = "PK\x05\x06".b.freeze
|
|
16
|
+
|
|
17
|
+
# Size of the end of central directory record, excluding the comment
|
|
18
|
+
ZIP_EOCD_SIZE = 22
|
|
19
|
+
|
|
20
|
+
# Maximum length of a ZIP archive comment
|
|
21
|
+
ZIP_MAX_COMMENT_SIZE = 0xFFFF
|
|
22
|
+
|
|
23
|
+
# The main document part every Word OOXML package contains
|
|
24
|
+
OOXML_DOCUMENT_PART = 'word/document.xml'.b.freeze
|
|
25
|
+
|
|
26
|
+
# LibreOffice import filters, by format
|
|
27
|
+
FILTERS = {
|
|
28
|
+
ooxml: 'MS Word 2007 XML',
|
|
29
|
+
ole: 'MS Word 97'
|
|
30
|
+
}.freeze
|
|
31
|
+
|
|
32
|
+
class << self
|
|
33
|
+
# @param path [String] path to the file
|
|
34
|
+
# @return [String, nil] the LibreOffice import filter for the file, or nil if it isn't a Word document
|
|
35
|
+
def filter_for(path)
|
|
36
|
+
FILTERS[detect(path)]
|
|
37
|
+
end
|
|
38
|
+
|
|
39
|
+
# @param path [String] path to the file
|
|
40
|
+
# @return [Symbol, nil] :ooxml for .docx, :ole for .doc, or nil if it isn't a Word document
|
|
41
|
+
def detect(path)
|
|
42
|
+
File.open(path, 'rb') do |file|
|
|
43
|
+
signature = file.read(OLE_SIGNATURE.bytesize).to_s
|
|
44
|
+
if signature == OLE_SIGNATURE
|
|
45
|
+
:ole
|
|
46
|
+
elsif signature.start_with?(ZIP_SIGNATURE) && ooxml_document?(file)
|
|
47
|
+
:ooxml
|
|
48
|
+
end
|
|
49
|
+
end
|
|
50
|
+
end
|
|
51
|
+
|
|
52
|
+
private
|
|
53
|
+
|
|
54
|
+
# @param file [File] an open ZIP file
|
|
55
|
+
# @return [Boolean] true if the ZIP's central directory lists the Word main document part
|
|
56
|
+
def ooxml_document?(file)
|
|
57
|
+
directory = central_directory(file)
|
|
58
|
+
!directory.nil? && directory.include?(OOXML_DOCUMENT_PART)
|
|
59
|
+
end
|
|
60
|
+
|
|
61
|
+
# Read the ZIP central directory, whose entries include each file name, uncompressed
|
|
62
|
+
#
|
|
63
|
+
# @param file [File] an open ZIP file
|
|
64
|
+
# @return [String, nil] the raw central directory, or nil if it can't be found
|
|
65
|
+
def central_directory(file)
|
|
66
|
+
size, offset = central_directory_location(file)
|
|
67
|
+
return if size.nil? || offset + size > file.size
|
|
68
|
+
|
|
69
|
+
file.seek(offset)
|
|
70
|
+
file.read(size)
|
|
71
|
+
end
|
|
72
|
+
|
|
73
|
+
# @param file [File] an open ZIP file
|
|
74
|
+
# @return [Array<Integer>, nil] the central directory's size and offset, or nil if not found
|
|
75
|
+
def central_directory_location(file)
|
|
76
|
+
tail_size = [file.size, ZIP_EOCD_SIZE + ZIP_MAX_COMMENT_SIZE].min
|
|
77
|
+
file.seek(-tail_size, IO::SEEK_END)
|
|
78
|
+
tail = file.read(tail_size)
|
|
79
|
+
index = tail.rindex(ZIP_EOCD_SIGNATURE)
|
|
80
|
+
return if index.nil? || tail.bytesize - index < ZIP_EOCD_SIZE
|
|
81
|
+
|
|
82
|
+
tail.byteslice(index + 12, 8).unpack('VV')
|
|
83
|
+
end
|
|
84
|
+
end
|
|
85
|
+
end
|
|
86
|
+
end
|
|
@@ -0,0 +1,110 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
class WordToMarkdown
|
|
4
|
+
# Removes links and images with unsafe URL schemes (e.g., javascript: or
|
|
5
|
+
# vbscript:) so that they don't survive into the markdown output
|
|
6
|
+
module UrlScrubber
|
|
7
|
+
# URL schemes permitted in link targets. Links with any other scheme are
|
|
8
|
+
# unwrapped to their text. Relative URLs and fragments are always permitted.
|
|
9
|
+
SAFE_LINK_SCHEMES = %w[http https mailto].freeze
|
|
10
|
+
|
|
11
|
+
# URL schemes permitted in image sources, in addition to data:image/ URIs.
|
|
12
|
+
# Images with any other scheme are removed.
|
|
13
|
+
SAFE_IMAGE_SCHEMES = %w[http https].freeze
|
|
14
|
+
|
|
15
|
+
# Browsers remove ASCII tabs and newlines anywhere in a URL, and C0
|
|
16
|
+
# control characters and spaces at either end, before parsing the scheme
|
|
17
|
+
STRIPPED_CHARS = /[\t\n\r]/
|
|
18
|
+
EDGE_CHARS = /\A[\x00-\x20]+|[\x00-\x20]+\z/
|
|
19
|
+
|
|
20
|
+
# Matches a URL scheme, e.g., "https:"
|
|
21
|
+
SCHEME_REGEX = /\A([a-z][a-z0-9+.-]*):/i
|
|
22
|
+
|
|
23
|
+
# Matches a URL without a recognizable scheme whose first segment contains
|
|
24
|
+
# characters a Markdown renderer or browser may decode into a scheme
|
|
25
|
+
# separator, e.g., "javascript:", "javascript:", or "javascript\:"
|
|
26
|
+
AMBIGUOUS_REGEX = %r{\A[^/?#]*[:&\\]}
|
|
27
|
+
|
|
28
|
+
# Characters that could end or alter a Markdown link destination, which
|
|
29
|
+
# are percent-encoded in the URLs that are kept
|
|
30
|
+
DESTINATION_UNSAFE_CHARS = /[\x00-\x20\x7F()<>\[\]\\`"]/
|
|
31
|
+
|
|
32
|
+
# Matches a data URI for an image
|
|
33
|
+
DATA_IMAGE_REGEX = %r{\Adata:image/}i
|
|
34
|
+
|
|
35
|
+
class << self
|
|
36
|
+
# Unwrap links and remove images whose URL scheme isn't permitted
|
|
37
|
+
#
|
|
38
|
+
# @param tree [Nokogiri::HTML::Document] the document to scrub, in place
|
|
39
|
+
# @return [Nokogiri::HTML::Document] the scrubbed document
|
|
40
|
+
def scrub!(tree)
|
|
41
|
+
tree.css('a[href]').each do |node|
|
|
42
|
+
safe_link?(node['href']) ? escape_attributes!(node, 'href') : node.replace(node.children)
|
|
43
|
+
end
|
|
44
|
+
|
|
45
|
+
tree.css('img[src]').each do |node|
|
|
46
|
+
safe_image?(node['src']) ? escape_attributes!(node, 'src') : node.remove
|
|
47
|
+
end
|
|
48
|
+
|
|
49
|
+
tree
|
|
50
|
+
end
|
|
51
|
+
|
|
52
|
+
# Percent-encode characters in a URL that could end or alter a Markdown link destination
|
|
53
|
+
#
|
|
54
|
+
# @param url [String] the URL
|
|
55
|
+
# @return [String] the escaped URL
|
|
56
|
+
def escape_destination(url)
|
|
57
|
+
url.gsub(DESTINATION_UNSAFE_CHARS) { |char| format('%%%02X', char.ord) }
|
|
58
|
+
end
|
|
59
|
+
|
|
60
|
+
# @param url [String] a link target
|
|
61
|
+
# @return [Boolean] true if the URL is relative, a fragment, or has a permitted scheme
|
|
62
|
+
def safe_link?(url)
|
|
63
|
+
scheme = scheme(url)
|
|
64
|
+
scheme.nil? ? relative?(url) : SAFE_LINK_SCHEMES.include?(scheme)
|
|
65
|
+
end
|
|
66
|
+
|
|
67
|
+
# @param url [String] an image source
|
|
68
|
+
# @return [Boolean] true if the URL is relative, a data:image/ URI, or has a permitted scheme
|
|
69
|
+
def safe_image?(url)
|
|
70
|
+
scheme = scheme(url)
|
|
71
|
+
return relative?(url) if scheme.nil?
|
|
72
|
+
|
|
73
|
+
SAFE_IMAGE_SCHEMES.include?(scheme) || normalize(url).match?(DATA_IMAGE_REGEX)
|
|
74
|
+
end
|
|
75
|
+
|
|
76
|
+
# @param url [String] the URL
|
|
77
|
+
# @return [String, nil] the URL's lowercased scheme, or nil if it has none
|
|
78
|
+
def scheme(url)
|
|
79
|
+
match = normalize(url).match(SCHEME_REGEX)
|
|
80
|
+
match && match[1].downcase
|
|
81
|
+
end
|
|
82
|
+
|
|
83
|
+
private
|
|
84
|
+
|
|
85
|
+
# Escape a kept link or image's URL and title so they can't break out
|
|
86
|
+
# of the Markdown link that ReverseMarkdown writes for them
|
|
87
|
+
#
|
|
88
|
+
# @param node [Nokogiri::XML::Element] the link or image
|
|
89
|
+
# @param attribute [String] the name of the URL attribute
|
|
90
|
+
def escape_attributes!(node, attribute)
|
|
91
|
+
node[attribute] = escape_destination(node[attribute])
|
|
92
|
+
node['title'] = node['title'].tr('"', "'") if node['title']
|
|
93
|
+
end
|
|
94
|
+
|
|
95
|
+
# @param url [String] a URL without a scheme
|
|
96
|
+
# @return [Boolean] true if the URL is a relative path or fragment that can't be decoded into one with a scheme
|
|
97
|
+
def relative?(url)
|
|
98
|
+
!normalize(url).match?(AMBIGUOUS_REGEX)
|
|
99
|
+
end
|
|
100
|
+
|
|
101
|
+
# Normalize a URL the way a browser does before parsing its scheme
|
|
102
|
+
#
|
|
103
|
+
# @param url [String] the URL
|
|
104
|
+
# @return [String] the normalized URL
|
|
105
|
+
def normalize(url)
|
|
106
|
+
url.to_s.gsub(STRIPPED_CHARS, '').gsub(EDGE_CHARS, '')
|
|
107
|
+
end
|
|
108
|
+
end
|
|
109
|
+
end
|
|
110
|
+
end
|
data/lib/word-to-markdown.rb
CHANGED
|
@@ -1,6 +1,5 @@
|
|
|
1
1
|
# frozen_string_literal: true
|
|
2
2
|
|
|
3
|
-
require 'descriptive_statistics'
|
|
4
3
|
require 'reverse_markdown'
|
|
5
4
|
require 'nokogiri-styles'
|
|
6
5
|
require 'premailer'
|
|
@@ -12,28 +11,38 @@ require 'cliver'
|
|
|
12
11
|
require 'open3'
|
|
13
12
|
|
|
14
13
|
require_relative 'word-to-markdown/version'
|
|
14
|
+
require_relative 'word-to-markdown/input_format'
|
|
15
|
+
require_relative 'word-to-markdown/url_scrubber'
|
|
15
16
|
require_relative 'word-to-markdown/document'
|
|
16
17
|
require_relative 'word-to-markdown/converter'
|
|
17
18
|
require_relative 'nokogiri/xml/element'
|
|
18
19
|
require_relative 'cliver/dependency_ext'
|
|
19
20
|
|
|
20
21
|
class WordToMarkdown
|
|
22
|
+
class TimeoutError < StandardError; end
|
|
23
|
+
|
|
21
24
|
attr_reader :document, :converter
|
|
22
25
|
|
|
23
26
|
# Options to be passed to Reverse Markdown
|
|
24
27
|
REVERSE_MARKDOWN_OPTIONS = {
|
|
25
|
-
unknown_tags:
|
|
28
|
+
unknown_tags: :bypass,
|
|
26
29
|
github_flavored: true
|
|
27
30
|
}.freeze
|
|
28
31
|
|
|
32
|
+
# Default number of seconds to wait for LibreOffice to convert a document
|
|
33
|
+
# before giving up. Can be overridden with the WORD_TO_MARKDOWN_TIMEOUT
|
|
34
|
+
# environment variable, or by setting WordToMarkdown.timeout
|
|
35
|
+
DEFAULT_TIMEOUT = 60
|
|
36
|
+
|
|
29
37
|
# Minimum version of LibreOffice Required
|
|
30
|
-
SOFFICE_VERSION_REQUIREMENT = '> 4.0'
|
|
38
|
+
SOFFICE_VERSION_REQUIREMENT = '> 4.0'
|
|
31
39
|
|
|
32
40
|
# Paths to look for LibreOffice, in order of preference
|
|
33
41
|
PATHS = [
|
|
34
42
|
'*', # Sub'd for ENV["PATH"]
|
|
35
43
|
'~/Applications/LibreOffice.app/Contents/MacOS',
|
|
36
44
|
'/Applications/LibreOffice.app/Contents/MacOS',
|
|
45
|
+
'/Program Files/LibreOffice/program',
|
|
37
46
|
'/Program Files/LibreOffice 5/program',
|
|
38
47
|
'/Program Files (x86)/LibreOffice 4/program'
|
|
39
48
|
].freeze
|
|
@@ -56,19 +65,46 @@ class WordToMarkdown
|
|
|
56
65
|
end
|
|
57
66
|
|
|
58
67
|
class << self
|
|
68
|
+
attr_writer :timeout
|
|
69
|
+
|
|
70
|
+
# @return [Numeric] seconds to wait for LibreOffice before giving up
|
|
71
|
+
def timeout
|
|
72
|
+
@timeout ||= Float(ENV.fetch('WORD_TO_MARKDOWN_TIMEOUT', DEFAULT_TIMEOUT))
|
|
73
|
+
end
|
|
74
|
+
|
|
59
75
|
# Run an soffice command
|
|
60
76
|
#
|
|
61
|
-
# @param args [string] one or more arguments to pass to the
|
|
77
|
+
# @param args [string] one or more arguments to pass to the soffice command
|
|
62
78
|
# @return [string] the command output
|
|
63
79
|
def run_command(*args)
|
|
64
80
|
raise 'LibreOffice already running' if soffice.open?
|
|
65
81
|
|
|
66
|
-
output, status =
|
|
82
|
+
output, status = capture_with_timeout(soffice.path, *args, timeout: timeout)
|
|
67
83
|
logger.debug output
|
|
68
84
|
raise "Command `#{soffice.path} #{args.join(' ')}` failed: #{output}" if status.exitstatus != 0
|
|
85
|
+
|
|
69
86
|
output
|
|
70
87
|
end
|
|
71
88
|
|
|
89
|
+
# Run a command, capturing its combined output, and kill it (along with
|
|
90
|
+
# any child processes) if it runs longer than the timeout
|
|
91
|
+
#
|
|
92
|
+
# @param command [Array<String>] the command and its arguments
|
|
93
|
+
# @param timeout [Numeric] seconds to wait before killing the command
|
|
94
|
+
# @return [Array(String, Process::Status)] the command output and status
|
|
95
|
+
def capture_with_timeout(*command, timeout:)
|
|
96
|
+
Open3.popen2e(*command, process_group_option) do |stdin, output, waiter|
|
|
97
|
+
stdin.close
|
|
98
|
+
reader = read_in_background(output)
|
|
99
|
+
unless waiter.join(timeout)
|
|
100
|
+
kill_process_group(waiter.pid)
|
|
101
|
+
raise TimeoutError, "Command `#{command.join(' ')}` timed out after #{timeout} seconds"
|
|
102
|
+
end
|
|
103
|
+
|
|
104
|
+
[reader.value, waiter.value]
|
|
105
|
+
end
|
|
106
|
+
end
|
|
107
|
+
|
|
72
108
|
# Returns a Cliver::Dependency object representing our soffice dependency
|
|
73
109
|
#
|
|
74
110
|
# Attempts to resolve by looking at PATH followed by paths in the PATHS constant
|
|
@@ -85,7 +121,7 @@ class WordToMarkdown
|
|
|
85
121
|
# @return Logger instance
|
|
86
122
|
def logger
|
|
87
123
|
@logger ||= begin
|
|
88
|
-
logger = Logger.new(
|
|
124
|
+
logger = Logger.new($stdout)
|
|
89
125
|
logger.level = Logger::ERROR unless ENV['DEBUG']
|
|
90
126
|
logger
|
|
91
127
|
end
|
|
@@ -93,13 +129,38 @@ class WordToMarkdown
|
|
|
93
129
|
|
|
94
130
|
private
|
|
95
131
|
|
|
132
|
+
# Start commands in their own process group, so that on timeout any
|
|
133
|
+
# processes they spawn (e.g., soffice.bin) are killed too
|
|
134
|
+
def process_group_option
|
|
135
|
+
Gem.win_platform? ? { new_pgroup: true } : { pgroup: true }
|
|
136
|
+
end
|
|
137
|
+
|
|
138
|
+
# Read a stream in a separate thread, so the command can't block on a full pipe
|
|
139
|
+
#
|
|
140
|
+
# @param io [IO] the stream to read
|
|
141
|
+
# @return [Thread] a thread whose value is the stream's contents
|
|
142
|
+
def read_in_background(io)
|
|
143
|
+
Thread.new do
|
|
144
|
+
io.read
|
|
145
|
+
rescue IOError
|
|
146
|
+
'' # The stream was closed after the command timed out
|
|
147
|
+
end
|
|
148
|
+
end
|
|
149
|
+
|
|
150
|
+
# @param pid [Integer] the pid of a process group leader
|
|
151
|
+
def kill_process_group(pid)
|
|
152
|
+
Process.kill('KILL', Gem.win_platform? ? pid : -pid)
|
|
153
|
+
rescue Errno::ESRCH
|
|
154
|
+
nil
|
|
155
|
+
end
|
|
156
|
+
|
|
96
157
|
# Workaround for two upstream bugs:
|
|
97
|
-
# 1. `soffice.exe --version` on windows opens a popup and
|
|
158
|
+
# 1. `soffice.exe --version` on windows opens a popup and returns a null string when manually closed
|
|
98
159
|
# 2. Even if the second argument to Cliver is nil, Cliver thinks there's a requirement
|
|
99
160
|
# and will shell out to `soffice.exe --version`
|
|
100
161
|
# In order to support Windows, don't pass *any* version requirement to Cliver
|
|
101
162
|
def soffice_dependency_args
|
|
102
|
-
args = [path: PATHS.join(File::PATH_SEPARATOR)]
|
|
163
|
+
args = [{ path: PATHS.join(File::PATH_SEPARATOR) }]
|
|
103
164
|
if Gem.win_platform?
|
|
104
165
|
args
|
|
105
166
|
else
|
metadata
CHANGED
|
@@ -1,14 +1,13 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: word-to-markdown
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 1.
|
|
4
|
+
version: 1.2.0
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Ben Balter
|
|
8
|
-
autorequire:
|
|
9
8
|
bindir: bin
|
|
10
9
|
cert_chain: []
|
|
11
|
-
date:
|
|
10
|
+
date: 1980-01-02 00:00:00.000000000 Z
|
|
12
11
|
dependencies:
|
|
13
12
|
- !ruby/object:Gem::Dependency
|
|
14
13
|
name: cliver
|
|
@@ -25,19 +24,19 @@ dependencies:
|
|
|
25
24
|
- !ruby/object:Gem::Version
|
|
26
25
|
version: '0.3'
|
|
27
26
|
- !ruby/object:Gem::Dependency
|
|
28
|
-
name:
|
|
27
|
+
name: logger
|
|
29
28
|
requirement: !ruby/object:Gem::Requirement
|
|
30
29
|
requirements:
|
|
31
30
|
- - "~>"
|
|
32
31
|
- !ruby/object:Gem::Version
|
|
33
|
-
version: '
|
|
32
|
+
version: '1.4'
|
|
34
33
|
type: :runtime
|
|
35
34
|
prerelease: false
|
|
36
35
|
version_requirements: !ruby/object:Gem::Requirement
|
|
37
36
|
requirements:
|
|
38
37
|
- - "~>"
|
|
39
38
|
- !ruby/object:Gem::Version
|
|
40
|
-
version: '
|
|
39
|
+
version: '1.4'
|
|
41
40
|
- !ruby/object:Gem::Dependency
|
|
42
41
|
name: nokogiri-styles
|
|
43
42
|
requirement: !ruby/object:Gem::Requirement
|
|
@@ -70,128 +69,160 @@ dependencies:
|
|
|
70
69
|
name: reverse_markdown
|
|
71
70
|
requirement: !ruby/object:Gem::Requirement
|
|
72
71
|
requirements:
|
|
73
|
-
- - "
|
|
72
|
+
- - ">="
|
|
74
73
|
- !ruby/object:Gem::Version
|
|
75
|
-
version: '1
|
|
74
|
+
version: '1'
|
|
75
|
+
- - "<"
|
|
76
|
+
- !ruby/object:Gem::Version
|
|
77
|
+
version: '3'
|
|
76
78
|
type: :runtime
|
|
77
79
|
prerelease: false
|
|
78
80
|
version_requirements: !ruby/object:Gem::Requirement
|
|
79
81
|
requirements:
|
|
80
|
-
- - "
|
|
82
|
+
- - ">="
|
|
81
83
|
- !ruby/object:Gem::Version
|
|
82
|
-
version: '1
|
|
84
|
+
version: '1'
|
|
85
|
+
- - "<"
|
|
86
|
+
- !ruby/object:Gem::Version
|
|
87
|
+
version: '3'
|
|
83
88
|
- !ruby/object:Gem::Dependency
|
|
84
89
|
name: sys-proctable
|
|
85
90
|
requirement: !ruby/object:Gem::Requirement
|
|
86
91
|
requirements:
|
|
87
92
|
- - "~>"
|
|
88
93
|
- !ruby/object:Gem::Version
|
|
89
|
-
version: '1.
|
|
94
|
+
version: '1.3'
|
|
90
95
|
type: :runtime
|
|
91
96
|
prerelease: false
|
|
92
97
|
version_requirements: !ruby/object:Gem::Requirement
|
|
93
98
|
requirements:
|
|
94
99
|
- - "~>"
|
|
95
100
|
- !ruby/object:Gem::Version
|
|
96
|
-
version: '1.
|
|
101
|
+
version: '1.3'
|
|
97
102
|
- !ruby/object:Gem::Dependency
|
|
98
|
-
name:
|
|
103
|
+
name: minitest
|
|
99
104
|
requirement: !ruby/object:Gem::Requirement
|
|
100
105
|
requirements:
|
|
101
|
-
- - "
|
|
106
|
+
- - ">="
|
|
107
|
+
- !ruby/object:Gem::Version
|
|
108
|
+
version: '5'
|
|
109
|
+
- - "<"
|
|
102
110
|
- !ruby/object:Gem::Version
|
|
103
|
-
version: '
|
|
111
|
+
version: '7'
|
|
104
112
|
type: :development
|
|
105
113
|
prerelease: false
|
|
106
114
|
version_requirements: !ruby/object:Gem::Requirement
|
|
107
115
|
requirements:
|
|
108
|
-
- - "
|
|
116
|
+
- - ">="
|
|
117
|
+
- !ruby/object:Gem::Version
|
|
118
|
+
version: '5'
|
|
119
|
+
- - "<"
|
|
109
120
|
- !ruby/object:Gem::Version
|
|
110
|
-
version: '
|
|
121
|
+
version: '7'
|
|
111
122
|
- !ruby/object:Gem::Dependency
|
|
112
|
-
name:
|
|
123
|
+
name: mocha
|
|
124
|
+
requirement: !ruby/object:Gem::Requirement
|
|
125
|
+
requirements:
|
|
126
|
+
- - ">="
|
|
127
|
+
- !ruby/object:Gem::Version
|
|
128
|
+
version: '2'
|
|
129
|
+
- - "<"
|
|
130
|
+
- !ruby/object:Gem::Version
|
|
131
|
+
version: '4'
|
|
132
|
+
type: :development
|
|
133
|
+
prerelease: false
|
|
134
|
+
version_requirements: !ruby/object:Gem::Requirement
|
|
135
|
+
requirements:
|
|
136
|
+
- - ">="
|
|
137
|
+
- !ruby/object:Gem::Version
|
|
138
|
+
version: '2'
|
|
139
|
+
- - "<"
|
|
140
|
+
- !ruby/object:Gem::Version
|
|
141
|
+
version: '4'
|
|
142
|
+
- !ruby/object:Gem::Dependency
|
|
143
|
+
name: pry
|
|
113
144
|
requirement: !ruby/object:Gem::Requirement
|
|
114
145
|
requirements:
|
|
115
146
|
- - "~>"
|
|
116
147
|
- !ruby/object:Gem::Version
|
|
117
|
-
version: '
|
|
148
|
+
version: '0.10'
|
|
118
149
|
type: :development
|
|
119
150
|
prerelease: false
|
|
120
151
|
version_requirements: !ruby/object:Gem::Requirement
|
|
121
152
|
requirements:
|
|
122
153
|
- - "~>"
|
|
123
154
|
- !ruby/object:Gem::Version
|
|
124
|
-
version: '
|
|
155
|
+
version: '0.10'
|
|
125
156
|
- !ruby/object:Gem::Dependency
|
|
126
|
-
name:
|
|
157
|
+
name: rake
|
|
127
158
|
requirement: !ruby/object:Gem::Requirement
|
|
128
159
|
requirements:
|
|
129
160
|
- - "~>"
|
|
130
161
|
- !ruby/object:Gem::Version
|
|
131
|
-
version: '
|
|
162
|
+
version: '13.0'
|
|
132
163
|
type: :development
|
|
133
164
|
prerelease: false
|
|
134
165
|
version_requirements: !ruby/object:Gem::Requirement
|
|
135
166
|
requirements:
|
|
136
167
|
- - "~>"
|
|
137
168
|
- !ruby/object:Gem::Version
|
|
138
|
-
version: '
|
|
169
|
+
version: '13.0'
|
|
139
170
|
- !ruby/object:Gem::Dependency
|
|
140
|
-
name:
|
|
171
|
+
name: rubocop
|
|
141
172
|
requirement: !ruby/object:Gem::Requirement
|
|
142
173
|
requirements:
|
|
143
174
|
- - "~>"
|
|
144
175
|
- !ruby/object:Gem::Version
|
|
145
|
-
version: '0
|
|
176
|
+
version: '1.0'
|
|
146
177
|
type: :development
|
|
147
178
|
prerelease: false
|
|
148
179
|
version_requirements: !ruby/object:Gem::Requirement
|
|
149
180
|
requirements:
|
|
150
181
|
- - "~>"
|
|
151
182
|
- !ruby/object:Gem::Version
|
|
152
|
-
version: '0
|
|
183
|
+
version: '1.0'
|
|
153
184
|
- !ruby/object:Gem::Dependency
|
|
154
|
-
name:
|
|
185
|
+
name: rubocop-minitest
|
|
155
186
|
requirement: !ruby/object:Gem::Requirement
|
|
156
187
|
requirements:
|
|
157
188
|
- - "~>"
|
|
158
189
|
- !ruby/object:Gem::Version
|
|
159
|
-
version: '
|
|
190
|
+
version: '0.3'
|
|
160
191
|
type: :development
|
|
161
192
|
prerelease: false
|
|
162
193
|
version_requirements: !ruby/object:Gem::Requirement
|
|
163
194
|
requirements:
|
|
164
195
|
- - "~>"
|
|
165
196
|
- !ruby/object:Gem::Version
|
|
166
|
-
version: '
|
|
197
|
+
version: '0.3'
|
|
167
198
|
- !ruby/object:Gem::Dependency
|
|
168
|
-
name: rubocop
|
|
199
|
+
name: rubocop-performance
|
|
169
200
|
requirement: !ruby/object:Gem::Requirement
|
|
170
201
|
requirements:
|
|
171
202
|
- - "~>"
|
|
172
203
|
- !ruby/object:Gem::Version
|
|
173
|
-
version: '
|
|
204
|
+
version: '1.5'
|
|
174
205
|
type: :development
|
|
175
206
|
prerelease: false
|
|
176
207
|
version_requirements: !ruby/object:Gem::Requirement
|
|
177
208
|
requirements:
|
|
178
209
|
- - "~>"
|
|
179
210
|
- !ruby/object:Gem::Version
|
|
180
|
-
version: '
|
|
211
|
+
version: '1.5'
|
|
181
212
|
- !ruby/object:Gem::Dependency
|
|
182
213
|
name: shoulda
|
|
183
214
|
requirement: !ruby/object:Gem::Requirement
|
|
184
215
|
requirements:
|
|
185
216
|
- - "~>"
|
|
186
217
|
- !ruby/object:Gem::Version
|
|
187
|
-
version: '
|
|
218
|
+
version: '4.0'
|
|
188
219
|
type: :development
|
|
189
220
|
prerelease: false
|
|
190
221
|
version_requirements: !ruby/object:Gem::Requirement
|
|
191
222
|
requirements:
|
|
192
223
|
- - "~>"
|
|
193
224
|
- !ruby/object:Gem::Version
|
|
194
|
-
version: '
|
|
225
|
+
version: '4.0'
|
|
195
226
|
description: Ruby Gem to convert Word documents to markdown.
|
|
196
227
|
email: ben.balter@github.com
|
|
197
228
|
executables:
|
|
@@ -207,12 +238,17 @@ files:
|
|
|
207
238
|
- lib/word-to-markdown.rb
|
|
208
239
|
- lib/word-to-markdown/converter.rb
|
|
209
240
|
- lib/word-to-markdown/document.rb
|
|
241
|
+
- lib/word-to-markdown/input_format.rb
|
|
242
|
+
- lib/word-to-markdown/url_scrubber.rb
|
|
210
243
|
- lib/word-to-markdown/version.rb
|
|
211
244
|
homepage: https://github.com/benbalter/word-to-markdown
|
|
212
245
|
licenses:
|
|
213
246
|
- MIT
|
|
214
|
-
metadata:
|
|
215
|
-
|
|
247
|
+
metadata:
|
|
248
|
+
rubygems_mfa_required: 'true'
|
|
249
|
+
source_code_uri: https://github.com/benbalter/word-to-markdown
|
|
250
|
+
bug_tracker_uri: https://github.com/benbalter/word-to-markdown/issues
|
|
251
|
+
changelog_uri: https://github.com/benbalter/word-to-markdown/releases
|
|
216
252
|
rdoc_options: []
|
|
217
253
|
require_paths:
|
|
218
254
|
- lib
|
|
@@ -220,16 +256,14 @@ required_ruby_version: !ruby/object:Gem::Requirement
|
|
|
220
256
|
requirements:
|
|
221
257
|
- - ">="
|
|
222
258
|
- !ruby/object:Gem::Version
|
|
223
|
-
version: '
|
|
259
|
+
version: '3.2'
|
|
224
260
|
required_rubygems_version: !ruby/object:Gem::Requirement
|
|
225
261
|
requirements:
|
|
226
262
|
- - ">="
|
|
227
263
|
- !ruby/object:Gem::Version
|
|
228
264
|
version: '0'
|
|
229
265
|
requirements: []
|
|
230
|
-
|
|
231
|
-
rubygems_version: 2.7.6
|
|
232
|
-
signing_key:
|
|
266
|
+
rubygems_version: 3.6.9
|
|
233
267
|
specification_version: 4
|
|
234
268
|
summary: Ruby Gem to convert Word documents to markdown
|
|
235
269
|
test_files: []
|