word-to-markdown 1.1.9 → 1.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.md +20 -15
- data/bin/w2m +26 -8
- data/lib/word-to-markdown/converter.rb +35 -11
- data/lib/word-to-markdown/document.rb +25 -9
- data/lib/word-to-markdown/input_format.rb +86 -0
- data/lib/word-to-markdown/url_scrubber.rb +110 -0
- data/lib/word-to-markdown/version.rb +1 -1
- data/lib/word-to-markdown.rb +65 -5
- metadata +34 -20
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: d5fff6778e90bf76409d47c5177fd2400a14308cf40be3ea2e3ee118917945b0
|
|
4
|
+
data.tar.gz: 18062303949a156da4c940daeb79854c820036fe623cdb9c4c07b1bce9dca242
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: f041232dbf6ae3aee4738d9562fab3f820a492dfd8d71fdd257397b278ca03550a27a59a0d0b5c3514254bab49367ba067970f7551ea11c3a7b90b8eda434499
|
|
7
|
+
data.tar.gz: e8c267ccd410ccfc978ef873805696299b1a658d1894dccae2f63d39f8afd12595ced23fbce6c8f4173188cc0b1e6fa44957d37c29c32b8ea18c08fb9a8745e1
|
data/README.md
CHANGED
|
@@ -1,18 +1,21 @@
|
|
|
1
1
|
# Word to Markdown converter
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
> [!IMPORTANT]
|
|
4
|
+
> **Looking for the latest and greatest?** Check out [**word-to-markdown-js**](https://github.com/benbalter/word-to-markdown-js), the newer, better successor to this project. It's a modern, actively maintained rewrite and is recommended for new projects. This Ruby gem remains available for existing users.
|
|
4
5
|
|
|
5
|
-
|
|
6
|
+
A Ruby gem to liberate content from [the jail that is Word documents](https://ben.balter.com/2012/10/19/we-ve-been-trained-to-make-paper/#jailbreaking-content)
|
|
7
|
+
|
|
8
|
+
[](https://github.com/benbalter/word-to-markdown/actions/workflows/ci.yml) [](https://rubygems.org/gems/word-to-markdown)
|
|
6
9
|
|
|
7
10
|
## The problem
|
|
8
11
|
|
|
9
|
-
> Our default content publishing workflow is terribly broken. [We've all been trained to make paper](
|
|
12
|
+
> Our default content publishing workflow is terribly broken. [We've all been trained to make paper](https://ben.balter.com/2012/10/19/we-ve-been-trained-to-make-paper/), yet today, content authored once is more commonly consumed in multiple formats, and rarely, if ever, does it embody physical form. Put another way, our go-to content authoring workflow remains relatively unchanged since it was conceived in the early 80s.
|
|
10
13
|
>
|
|
11
|
-
> I'm asked regularly by government employees — knowledge workers who fire up a desktop word processor as the first step to any project — for an automated pipeline to convert Microsoft Word documents to [Markdown](
|
|
14
|
+
> I'm asked regularly by government employees — knowledge workers who fire up a desktop word processor as the first step to any project — for an automated pipeline to convert Microsoft Word documents to [Markdown](https://docs.github.com/en/get-started/writing-on-github/getting-started-with-writing-and-formatting-on-github/basic-writing-and-formatting-syntax), the *lingua franca* of the internet, but as my recent foray into building [just such a converter](https://word2md.com/) proves, it's not that simple.
|
|
12
15
|
>
|
|
13
16
|
> Markdown isn't just an alternative format. Markdown forces you to write for the web.
|
|
14
17
|
|
|
15
|
-
**[Read more](
|
|
18
|
+
**[Read more](https://ben.balter.com/2014/03/31/word-versus-markdown-more-than-mere-semantics/)**
|
|
16
19
|
|
|
17
20
|
## Just want to convert a Microsoft Word (or Google) document to Markdown?
|
|
18
21
|
|
|
@@ -20,7 +23,7 @@ You can use this **[hosted service](https://word2md.com/)** (or check out [its s
|
|
|
20
23
|
|
|
21
24
|
## Install
|
|
22
25
|
|
|
23
|
-
You'll need to install [LibreOffice](
|
|
26
|
+
You'll need to install [LibreOffice](https://www.libreoffice.org/). Then:
|
|
24
27
|
|
|
25
28
|
```bash
|
|
26
29
|
gem install word-to-markdown
|
|
@@ -67,6 +70,14 @@ $ w2m path/to/document.docx
|
|
|
67
70
|
|
|
68
71
|
Word-to-markdown requires `soffice` a command line interface to LibreOffice that works on Linux, Mac, and Windows. To install soffice, see [the LibreOffice documentation](https://www.libreoffice.org/get-help/install-howto/).
|
|
69
72
|
|
|
73
|
+
Word-to-markdown only accepts Word documents (`.docx` and `.doc`), identified by their contents rather than their file extension. Other files, including those LibreOffice could otherwise open, such as HTML, ODT, or RTF, raise `WordToMarkdown::Document::UnsupportedFormatError`.
|
|
74
|
+
|
|
75
|
+
LibreOffice is killed if a conversion takes longer than 60 seconds, raising `WordToMarkdown::TimeoutError`. To change the limit, set `WordToMarkdown.timeout = 120` or the `WORD_TO_MARKDOWN_TIMEOUT` environment variable.
|
|
76
|
+
|
|
77
|
+
### Converting untrusted documents
|
|
78
|
+
|
|
79
|
+
Word documents can link to remote resources, such as images, which LibreOffice fetches while converting the document. If you convert documents from untrusted sources (for example, files uploaded to a web service), run the conversion without network access, such as in a container or sandbox with no outbound network, so that a document can't make requests to internal services or other hosts on your behalf.
|
|
80
|
+
|
|
70
81
|
## Testing
|
|
71
82
|
|
|
72
83
|
```
|
|
@@ -75,18 +86,12 @@ script/cibuild
|
|
|
75
86
|
|
|
76
87
|
## Docker
|
|
77
88
|
|
|
78
|
-
First, create the `Gemfile.lock` by installing the dependencies:
|
|
79
|
-
|
|
80
|
-
```
|
|
81
|
-
bundle install
|
|
82
|
-
```
|
|
83
|
-
|
|
84
89
|
Everything you need to run the executable locally:
|
|
85
90
|
|
|
86
91
|
```
|
|
87
|
-
docker
|
|
88
|
-
docker
|
|
89
|
-
docker
|
|
92
|
+
docker compose build
|
|
93
|
+
docker compose run --rm app bundle exec w2m --help
|
|
94
|
+
docker compose run --rm app bundle exec w2m test/fixtures/em.docx
|
|
90
95
|
```
|
|
91
96
|
|
|
92
97
|
## Hosted service
|
data/bin/w2m
CHANGED
|
@@ -1,17 +1,35 @@
|
|
|
1
1
|
#!/usr/bin/env ruby
|
|
2
2
|
# frozen_string_literal: true
|
|
3
3
|
|
|
4
|
+
require 'optparse'
|
|
4
5
|
require 'word-to-markdown'
|
|
5
6
|
|
|
6
|
-
|
|
7
|
-
|
|
7
|
+
parser = OptionParser.new do |opts|
|
|
8
|
+
opts.banner = 'Usage: w2m path/to/document.docx'
|
|
9
|
+
|
|
10
|
+
opts.on('-h', '--help', 'Show this help') do
|
|
11
|
+
puts opts
|
|
12
|
+
exit
|
|
13
|
+
end
|
|
14
|
+
|
|
15
|
+
opts.on('-v', '--version', 'Show the WordToMarkdown and LibreOffice versions') do
|
|
16
|
+
puts "WordToMarkdown v#{WordToMarkdown::VERSION}"
|
|
17
|
+
puts "LibreOffice v#{WordToMarkdown.soffice.version}" unless Gem.win_platform?
|
|
18
|
+
exit
|
|
19
|
+
end
|
|
20
|
+
end
|
|
21
|
+
|
|
22
|
+
begin
|
|
23
|
+
parser.parse!
|
|
24
|
+
rescue OptionParser::ParseError => e
|
|
25
|
+
warn e.message
|
|
26
|
+
warn parser
|
|
8
27
|
exit 1
|
|
9
28
|
end
|
|
10
29
|
|
|
11
|
-
if ARGV
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
else
|
|
15
|
-
doc = WordToMarkdown.new ARGV[0]
|
|
16
|
-
puts doc.to_s
|
|
30
|
+
if ARGV.size != 1
|
|
31
|
+
warn parser
|
|
32
|
+
exit 1
|
|
17
33
|
end
|
|
34
|
+
|
|
35
|
+
puts WordToMarkdown.new(ARGV[0])
|
|
@@ -7,14 +7,22 @@ class WordToMarkdown
|
|
|
7
7
|
# Number of headings to guess, e.g., h6
|
|
8
8
|
HEADING_DEPTH = 6
|
|
9
9
|
|
|
10
|
-
# Percentile step for
|
|
10
|
+
# Percentile step for each heading
|
|
11
11
|
HEADING_STEP = 100 / HEADING_DEPTH
|
|
12
12
|
|
|
13
13
|
# Minimum heading size
|
|
14
14
|
MIN_HEADING_SIZE = 20
|
|
15
15
|
|
|
16
16
|
# Unicode bullets to strip when processing
|
|
17
|
-
UNICODE_BULLETS = ['○', '
|
|
17
|
+
UNICODE_BULLETS = ['○', '●', "\u2022", '\\p{C}'].freeze
|
|
18
|
+
|
|
19
|
+
# Leading bullets to strip from list items. A plain "o" only counts as a
|
|
20
|
+
# bullet when followed by whitespace, so words like "orange" survive.
|
|
21
|
+
BULLET_REGEX = /\A(?:[#{UNICODE_BULLETS.join}]|o(?=[[:space:]]))+[[:space:]]*/
|
|
22
|
+
|
|
23
|
+
# Leading list numbering to strip, e.g., "1.", "a.", or "iv.", along with
|
|
24
|
+
# any whitespace (including non-breaking spaces) that follows it
|
|
25
|
+
NUMBERING_REGEX = /\A(?:\d+|[a-zA-Z]|[ivxlcdm]+|[IVXLCDM]+)\.(?:[[:space:]]+|\z)/
|
|
18
26
|
|
|
19
27
|
# @param document [WordToMarkdown::Document] The document to convert
|
|
20
28
|
def initialize(document)
|
|
@@ -58,13 +66,13 @@ class WordToMarkdown
|
|
|
58
66
|
@document.tree.css('[style]').each do |element|
|
|
59
67
|
sizes.push element.font_size.round(-1) unless element.font_size.nil?
|
|
60
68
|
end
|
|
61
|
-
sizes.uniq.sort
|
|
69
|
+
sizes.uniq.sort
|
|
62
70
|
end
|
|
63
71
|
end
|
|
64
72
|
|
|
65
73
|
# Given a Nokogiri node, guess what heading it represents, if any
|
|
66
74
|
#
|
|
67
|
-
# @param node [
|
|
75
|
+
# @param node [Nokogiri::Node] the nokogiri node
|
|
68
76
|
# @return [String, nil] the heading tag (e.g., H1), or nil
|
|
69
77
|
def guess_heading(node)
|
|
70
78
|
return nil if node.font_size.nil?
|
|
@@ -82,7 +90,23 @@ class WordToMarkdown
|
|
|
82
90
|
#
|
|
83
91
|
# @return [Integer] the minimum font size
|
|
84
92
|
def h(num)
|
|
85
|
-
|
|
93
|
+
self.class.percentile(font_sizes, ((HEADING_DEPTH - 1) - num) * HEADING_STEP)
|
|
94
|
+
end
|
|
95
|
+
|
|
96
|
+
# Linearly interpolated percentile, matching the algorithm used by the
|
|
97
|
+
# descriptive_statistics gem (v2.5.1) that this replaces
|
|
98
|
+
#
|
|
99
|
+
# @param values [Array<Numeric>] the values
|
|
100
|
+
# @param pct [Numeric] the percentile, from 0 to 100
|
|
101
|
+
#
|
|
102
|
+
# @return [Float, nil] the percentile, or nil if values is empty
|
|
103
|
+
def self.percentile(values, pct)
|
|
104
|
+
sorted = values.map(&:to_f).sort
|
|
105
|
+
rank = pct / 100.0 * (sorted.size - 1)
|
|
106
|
+
lower, upper = sorted[rank.floor, 2]
|
|
107
|
+
return lower if upper.nil? # empty or single-element collection, or the 100th percentile
|
|
108
|
+
|
|
109
|
+
lower + ((upper - lower) * (rank - rank.floor))
|
|
86
110
|
end
|
|
87
111
|
|
|
88
112
|
# Convert span-based font styles to `strong`s and `em`s
|
|
@@ -110,7 +134,7 @@ class WordToMarkdown
|
|
|
110
134
|
def remove_unicode_bullets_from_list_items!
|
|
111
135
|
path = WordToMarkdown.soffice.major_version == '5' ? 'li span span' : 'li span'
|
|
112
136
|
@document.tree.search(path).each do |span|
|
|
113
|
-
span.inner_html = span.inner_html.
|
|
137
|
+
span.inner_html = span.inner_html.sub(BULLET_REGEX, '')
|
|
114
138
|
end
|
|
115
139
|
end
|
|
116
140
|
|
|
@@ -118,21 +142,21 @@ class WordToMarkdown
|
|
|
118
142
|
def remove_numbering_from_list_items!
|
|
119
143
|
path = WordToMarkdown.soffice.major_version == '5' ? 'li span span' : 'li span'
|
|
120
144
|
@document.tree.search(path).each do |span|
|
|
121
|
-
span.inner_html = span.inner_html.
|
|
145
|
+
span.inner_html = span.inner_html.sub(NUMBERING_REGEX, '')
|
|
122
146
|
end
|
|
123
147
|
end
|
|
124
148
|
|
|
125
|
-
#
|
|
149
|
+
# Remove whitespace from list items
|
|
126
150
|
def remove_whitespace_from_list_items!
|
|
127
|
-
@document.tree.search('li span').each { |span| span.inner_html.strip
|
|
151
|
+
@document.tree.search('li span').each { |span| span.inner_html = span.inner_html.strip }
|
|
128
152
|
end
|
|
129
153
|
|
|
130
|
-
# Convert table headers to `th`
|
|
154
|
+
# Convert table headers to `th`s
|
|
131
155
|
def semanticize_table_headers!
|
|
132
156
|
@document.tree.search('table tr:first td').each { |node| node.node_name = 'th' }
|
|
133
157
|
end
|
|
134
158
|
|
|
135
|
-
# Try to guess heading where implicit
|
|
159
|
+
# Try to guess heading where implicit based on font size
|
|
136
160
|
def semanticize_headings!
|
|
137
161
|
implicit_headings.each do |element|
|
|
138
162
|
heading = guess_heading element
|
|
@@ -6,14 +6,21 @@ class WordToMarkdown
|
|
|
6
6
|
|
|
7
7
|
class ConversionError < StandardError; end
|
|
8
8
|
|
|
9
|
-
|
|
9
|
+
class UnsupportedFormatError < StandardError; end
|
|
10
|
+
|
|
11
|
+
attr_reader :path, :tmpdir, :import_filter
|
|
10
12
|
|
|
11
13
|
# @param path [string] Path to the Word document
|
|
12
14
|
# @param tmpdir [string] Path to a working directory to use
|
|
13
15
|
def initialize(path, tmpdir = nil)
|
|
14
16
|
@path = File.expand_path path, Dir.pwd
|
|
15
|
-
@tmpdir = tmpdir || Dir.mktmpdir
|
|
16
17
|
raise NotFoundError, "File #{@path} does not exist" unless File.exist?(@path)
|
|
18
|
+
|
|
19
|
+
@import_filter = InputFormat.filter_for(@path)
|
|
20
|
+
raise UnsupportedFormatError, "File #{@path} is not a Word document (.docx or .doc)" if @import_filter.nil?
|
|
21
|
+
|
|
22
|
+
@own_tmpdir = tmpdir.nil?
|
|
23
|
+
@tmpdir = tmpdir || Dir.mktmpdir
|
|
17
24
|
end
|
|
18
25
|
|
|
19
26
|
# @return [String] the document's extension
|
|
@@ -21,7 +28,7 @@ class WordToMarkdown
|
|
|
21
28
|
File.extname path
|
|
22
29
|
end
|
|
23
30
|
|
|
24
|
-
# @return [
|
|
31
|
+
# @return [Nokogiri::Document]
|
|
25
32
|
def tree
|
|
26
33
|
@tree ||= begin
|
|
27
34
|
tree = Nokogiri::HTML(normalized_html)
|
|
@@ -32,6 +39,7 @@ class WordToMarkdown
|
|
|
32
39
|
|
|
33
40
|
# @return [String] the html representation of the document
|
|
34
41
|
def html
|
|
42
|
+
UrlScrubber.scrub!(tree)
|
|
35
43
|
tree.to_html.gsub("</li>\n", '</li>')
|
|
36
44
|
end
|
|
37
45
|
|
|
@@ -79,8 +87,8 @@ class WordToMarkdown
|
|
|
79
87
|
string.gsub!(' ', ' ') # HTML encoded spaces
|
|
80
88
|
string.sub!(/\A[[:space:]]+/, '') # document leading whitespace
|
|
81
89
|
string.sub!(/[[:space:]]+\z/, '') # document trailing whitespace
|
|
82
|
-
string.gsub!(/(
|
|
83
|
-
string.gsub!(
|
|
90
|
+
string.gsub!(/( +)$/, '') # line trailing whitespace
|
|
91
|
+
string.gsub!("\n\n\n\n", "\n\n") # Quadruple line breaks
|
|
84
92
|
string.delete!(' ') # Unicode non-breaking spaces, injected as tabs
|
|
85
93
|
string.gsub!(/\*\*\ +(?!\*|_)([[:punct:]])/, '**\1') # Remove extra space after bold
|
|
86
94
|
string
|
|
@@ -88,23 +96,31 @@ class WordToMarkdown
|
|
|
88
96
|
|
|
89
97
|
# @return [String] the path to the intermediary HTML document
|
|
90
98
|
def dest_path
|
|
91
|
-
|
|
92
|
-
File.expand_path(dest_filename, tmpdir)
|
|
99
|
+
File.expand_path("#{File.basename(path, '.*')}.html", tmpdir)
|
|
93
100
|
end
|
|
94
101
|
|
|
95
102
|
# @return [String] the unnormalized HTML representation
|
|
96
103
|
def raw_html
|
|
97
104
|
@raw_html ||= begin
|
|
98
|
-
WordToMarkdown.run_command '--headless', '--convert-to', filter, path, '--outdir', tmpdir
|
|
105
|
+
WordToMarkdown.run_command '--headless', "--infilter=#{import_filter}", '--convert-to', filter, path, '--outdir', tmpdir
|
|
99
106
|
raise ConversionError, "Failed to convert #{path}" unless File.exist?(dest_path)
|
|
100
107
|
|
|
101
108
|
html = File.read dest_path
|
|
102
109
|
File.delete dest_path
|
|
103
110
|
html
|
|
111
|
+
ensure
|
|
112
|
+
remove_tmpdir
|
|
104
113
|
end
|
|
105
114
|
end
|
|
106
115
|
|
|
107
|
-
#
|
|
116
|
+
# Remove the working directory if we created it and nothing else is in it.
|
|
117
|
+
# Non-empty directories are kept, since LibreOffice may have written
|
|
118
|
+
# images there that the markdown references.
|
|
119
|
+
def remove_tmpdir
|
|
120
|
+
Dir.rmdir(tmpdir) if @own_tmpdir && Dir.empty?(tmpdir)
|
|
121
|
+
end
|
|
122
|
+
|
|
123
|
+
# @return [String] the LibreOffice filter to use for export
|
|
108
124
|
def filter
|
|
109
125
|
if WordToMarkdown.soffice.major_version == '5'
|
|
110
126
|
'html:XHTML Writer File:UTF8'
|
|
@@ -0,0 +1,86 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
class WordToMarkdown
|
|
4
|
+
# Identifies Word documents by their content, rather than their extension,
|
|
5
|
+
# so that other formats LibreOffice would otherwise sniff and import (e.g.,
|
|
6
|
+
# HTML, which can reference remote resources) are never passed to it
|
|
7
|
+
module InputFormat
|
|
8
|
+
# Signature of an OLE compound file, used by .doc files
|
|
9
|
+
OLE_SIGNATURE = "\xD0\xCF\x11\xE0\xA1\xB1\x1A\xE1".b.freeze
|
|
10
|
+
|
|
11
|
+
# Signature of a ZIP local file header, used by .docx files
|
|
12
|
+
ZIP_SIGNATURE = "PK\x03\x04".b.freeze
|
|
13
|
+
|
|
14
|
+
# Signature of a ZIP end of central directory record
|
|
15
|
+
ZIP_EOCD_SIGNATURE = "PK\x05\x06".b.freeze
|
|
16
|
+
|
|
17
|
+
# Size of the end of central directory record, excluding the comment
|
|
18
|
+
ZIP_EOCD_SIZE = 22
|
|
19
|
+
|
|
20
|
+
# Maximum length of a ZIP archive comment
|
|
21
|
+
ZIP_MAX_COMMENT_SIZE = 0xFFFF
|
|
22
|
+
|
|
23
|
+
# The main document part every Word OOXML package contains
|
|
24
|
+
OOXML_DOCUMENT_PART = 'word/document.xml'.b.freeze
|
|
25
|
+
|
|
26
|
+
# LibreOffice import filters, by format
|
|
27
|
+
FILTERS = {
|
|
28
|
+
ooxml: 'MS Word 2007 XML',
|
|
29
|
+
ole: 'MS Word 97'
|
|
30
|
+
}.freeze
|
|
31
|
+
|
|
32
|
+
class << self
|
|
33
|
+
# @param path [String] path to the file
|
|
34
|
+
# @return [String, nil] the LibreOffice import filter for the file, or nil if it isn't a Word document
|
|
35
|
+
def filter_for(path)
|
|
36
|
+
FILTERS[detect(path)]
|
|
37
|
+
end
|
|
38
|
+
|
|
39
|
+
# @param path [String] path to the file
|
|
40
|
+
# @return [Symbol, nil] :ooxml for .docx, :ole for .doc, or nil if it isn't a Word document
|
|
41
|
+
def detect(path)
|
|
42
|
+
File.open(path, 'rb') do |file|
|
|
43
|
+
signature = file.read(OLE_SIGNATURE.bytesize).to_s
|
|
44
|
+
if signature == OLE_SIGNATURE
|
|
45
|
+
:ole
|
|
46
|
+
elsif signature.start_with?(ZIP_SIGNATURE) && ooxml_document?(file)
|
|
47
|
+
:ooxml
|
|
48
|
+
end
|
|
49
|
+
end
|
|
50
|
+
end
|
|
51
|
+
|
|
52
|
+
private
|
|
53
|
+
|
|
54
|
+
# @param file [File] an open ZIP file
|
|
55
|
+
# @return [Boolean] true if the ZIP's central directory lists the Word main document part
|
|
56
|
+
def ooxml_document?(file)
|
|
57
|
+
directory = central_directory(file)
|
|
58
|
+
!directory.nil? && directory.include?(OOXML_DOCUMENT_PART)
|
|
59
|
+
end
|
|
60
|
+
|
|
61
|
+
# Read the ZIP central directory, whose entries include each file name, uncompressed
|
|
62
|
+
#
|
|
63
|
+
# @param file [File] an open ZIP file
|
|
64
|
+
# @return [String, nil] the raw central directory, or nil if it can't be found
|
|
65
|
+
def central_directory(file)
|
|
66
|
+
size, offset = central_directory_location(file)
|
|
67
|
+
return if size.nil? || offset + size > file.size
|
|
68
|
+
|
|
69
|
+
file.seek(offset)
|
|
70
|
+
file.read(size)
|
|
71
|
+
end
|
|
72
|
+
|
|
73
|
+
# @param file [File] an open ZIP file
|
|
74
|
+
# @return [Array<Integer>, nil] the central directory's size and offset, or nil if not found
|
|
75
|
+
def central_directory_location(file)
|
|
76
|
+
tail_size = [file.size, ZIP_EOCD_SIZE + ZIP_MAX_COMMENT_SIZE].min
|
|
77
|
+
file.seek(-tail_size, IO::SEEK_END)
|
|
78
|
+
tail = file.read(tail_size)
|
|
79
|
+
index = tail.rindex(ZIP_EOCD_SIGNATURE)
|
|
80
|
+
return if index.nil? || tail.bytesize - index < ZIP_EOCD_SIZE
|
|
81
|
+
|
|
82
|
+
tail.byteslice(index + 12, 8).unpack('VV')
|
|
83
|
+
end
|
|
84
|
+
end
|
|
85
|
+
end
|
|
86
|
+
end
|
|
@@ -0,0 +1,110 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
class WordToMarkdown
|
|
4
|
+
# Removes links and images with unsafe URL schemes (e.g., javascript: or
|
|
5
|
+
# vbscript:) so that they don't survive into the markdown output
|
|
6
|
+
module UrlScrubber
|
|
7
|
+
# URL schemes permitted in link targets. Links with any other scheme are
|
|
8
|
+
# unwrapped to their text. Relative URLs and fragments are always permitted.
|
|
9
|
+
SAFE_LINK_SCHEMES = %w[http https mailto].freeze
|
|
10
|
+
|
|
11
|
+
# URL schemes permitted in image sources, in addition to data:image/ URIs.
|
|
12
|
+
# Images with any other scheme are removed.
|
|
13
|
+
SAFE_IMAGE_SCHEMES = %w[http https].freeze
|
|
14
|
+
|
|
15
|
+
# Browsers remove ASCII tabs and newlines anywhere in a URL, and C0
|
|
16
|
+
# control characters and spaces at either end, before parsing the scheme
|
|
17
|
+
STRIPPED_CHARS = /[\t\n\r]/
|
|
18
|
+
EDGE_CHARS = /\A[\x00-\x20]+|[\x00-\x20]+\z/
|
|
19
|
+
|
|
20
|
+
# Matches a URL scheme, e.g., "https:"
|
|
21
|
+
SCHEME_REGEX = /\A([a-z][a-z0-9+.-]*):/i
|
|
22
|
+
|
|
23
|
+
# Matches a URL without a recognizable scheme whose first segment contains
|
|
24
|
+
# characters a Markdown renderer or browser may decode into a scheme
|
|
25
|
+
# separator, e.g., "javascript:", "javascript:", or "javascript\:"
|
|
26
|
+
AMBIGUOUS_REGEX = %r{\A[^/?#]*[:&\\]}
|
|
27
|
+
|
|
28
|
+
# Characters that could end or alter a Markdown link destination, which
|
|
29
|
+
# are percent-encoded in the URLs that are kept
|
|
30
|
+
DESTINATION_UNSAFE_CHARS = /[\x00-\x20\x7F()<>\[\]\\`"]/
|
|
31
|
+
|
|
32
|
+
# Matches a data URI for an image
|
|
33
|
+
DATA_IMAGE_REGEX = %r{\Adata:image/}i
|
|
34
|
+
|
|
35
|
+
class << self
|
|
36
|
+
# Unwrap links and remove images whose URL scheme isn't permitted
|
|
37
|
+
#
|
|
38
|
+
# @param tree [Nokogiri::HTML::Document] the document to scrub, in place
|
|
39
|
+
# @return [Nokogiri::HTML::Document] the scrubbed document
|
|
40
|
+
def scrub!(tree)
|
|
41
|
+
tree.css('a[href]').each do |node|
|
|
42
|
+
safe_link?(node['href']) ? escape_attributes!(node, 'href') : node.replace(node.children)
|
|
43
|
+
end
|
|
44
|
+
|
|
45
|
+
tree.css('img[src]').each do |node|
|
|
46
|
+
safe_image?(node['src']) ? escape_attributes!(node, 'src') : node.remove
|
|
47
|
+
end
|
|
48
|
+
|
|
49
|
+
tree
|
|
50
|
+
end
|
|
51
|
+
|
|
52
|
+
# Percent-encode characters in a URL that could end or alter a Markdown link destination
|
|
53
|
+
#
|
|
54
|
+
# @param url [String] the URL
|
|
55
|
+
# @return [String] the escaped URL
|
|
56
|
+
def escape_destination(url)
|
|
57
|
+
url.gsub(DESTINATION_UNSAFE_CHARS) { |char| format('%%%02X', char.ord) }
|
|
58
|
+
end
|
|
59
|
+
|
|
60
|
+
# @param url [String] a link target
|
|
61
|
+
# @return [Boolean] true if the URL is relative, a fragment, or has a permitted scheme
|
|
62
|
+
def safe_link?(url)
|
|
63
|
+
scheme = scheme(url)
|
|
64
|
+
scheme.nil? ? relative?(url) : SAFE_LINK_SCHEMES.include?(scheme)
|
|
65
|
+
end
|
|
66
|
+
|
|
67
|
+
# @param url [String] an image source
|
|
68
|
+
# @return [Boolean] true if the URL is relative, a data:image/ URI, or has a permitted scheme
|
|
69
|
+
def safe_image?(url)
|
|
70
|
+
scheme = scheme(url)
|
|
71
|
+
return relative?(url) if scheme.nil?
|
|
72
|
+
|
|
73
|
+
SAFE_IMAGE_SCHEMES.include?(scheme) || normalize(url).match?(DATA_IMAGE_REGEX)
|
|
74
|
+
end
|
|
75
|
+
|
|
76
|
+
# @param url [String] the URL
|
|
77
|
+
# @return [String, nil] the URL's lowercased scheme, or nil if it has none
|
|
78
|
+
def scheme(url)
|
|
79
|
+
match = normalize(url).match(SCHEME_REGEX)
|
|
80
|
+
match && match[1].downcase
|
|
81
|
+
end
|
|
82
|
+
|
|
83
|
+
private
|
|
84
|
+
|
|
85
|
+
# Escape a kept link or image's URL and title so they can't break out
|
|
86
|
+
# of the Markdown link that ReverseMarkdown writes for them
|
|
87
|
+
#
|
|
88
|
+
# @param node [Nokogiri::XML::Element] the link or image
|
|
89
|
+
# @param attribute [String] the name of the URL attribute
|
|
90
|
+
def escape_attributes!(node, attribute)
|
|
91
|
+
node[attribute] = escape_destination(node[attribute])
|
|
92
|
+
node['title'] = node['title'].tr('"', "'") if node['title']
|
|
93
|
+
end
|
|
94
|
+
|
|
95
|
+
# @param url [String] a URL without a scheme
|
|
96
|
+
# @return [Boolean] true if the URL is a relative path or fragment that can't be decoded into one with a scheme
|
|
97
|
+
def relative?(url)
|
|
98
|
+
!normalize(url).match?(AMBIGUOUS_REGEX)
|
|
99
|
+
end
|
|
100
|
+
|
|
101
|
+
# Normalize a URL the way a browser does before parsing its scheme
|
|
102
|
+
#
|
|
103
|
+
# @param url [String] the URL
|
|
104
|
+
# @return [String] the normalized URL
|
|
105
|
+
def normalize(url)
|
|
106
|
+
url.to_s.gsub(STRIPPED_CHARS, '').gsub(EDGE_CHARS, '')
|
|
107
|
+
end
|
|
108
|
+
end
|
|
109
|
+
end
|
|
110
|
+
end
|
data/lib/word-to-markdown.rb
CHANGED
|
@@ -1,6 +1,5 @@
|
|
|
1
1
|
# frozen_string_literal: true
|
|
2
2
|
|
|
3
|
-
require 'descriptive_statistics/safe'
|
|
4
3
|
require 'reverse_markdown'
|
|
5
4
|
require 'nokogiri-styles'
|
|
6
5
|
require 'premailer'
|
|
@@ -12,12 +11,16 @@ require 'cliver'
|
|
|
12
11
|
require 'open3'
|
|
13
12
|
|
|
14
13
|
require_relative 'word-to-markdown/version'
|
|
14
|
+
require_relative 'word-to-markdown/input_format'
|
|
15
|
+
require_relative 'word-to-markdown/url_scrubber'
|
|
15
16
|
require_relative 'word-to-markdown/document'
|
|
16
17
|
require_relative 'word-to-markdown/converter'
|
|
17
18
|
require_relative 'nokogiri/xml/element'
|
|
18
19
|
require_relative 'cliver/dependency_ext'
|
|
19
20
|
|
|
20
21
|
class WordToMarkdown
|
|
22
|
+
class TimeoutError < StandardError; end
|
|
23
|
+
|
|
21
24
|
attr_reader :document, :converter
|
|
22
25
|
|
|
23
26
|
# Options to be passed to Reverse Markdown
|
|
@@ -26,6 +29,11 @@ class WordToMarkdown
|
|
|
26
29
|
github_flavored: true
|
|
27
30
|
}.freeze
|
|
28
31
|
|
|
32
|
+
# Default number of seconds to wait for LibreOffice to convert a document
|
|
33
|
+
# before giving up. Can be overridden with the WORD_TO_MARKDOWN_TIMEOUT
|
|
34
|
+
# environment variable, or by setting WordToMarkdown.timeout
|
|
35
|
+
DEFAULT_TIMEOUT = 60
|
|
36
|
+
|
|
29
37
|
# Minimum version of LibreOffice Required
|
|
30
38
|
SOFFICE_VERSION_REQUIREMENT = '> 4.0'
|
|
31
39
|
|
|
@@ -34,6 +42,7 @@ class WordToMarkdown
|
|
|
34
42
|
'*', # Sub'd for ENV["PATH"]
|
|
35
43
|
'~/Applications/LibreOffice.app/Contents/MacOS',
|
|
36
44
|
'/Applications/LibreOffice.app/Contents/MacOS',
|
|
45
|
+
'/Program Files/LibreOffice/program',
|
|
37
46
|
'/Program Files/LibreOffice 5/program',
|
|
38
47
|
'/Program Files (x86)/LibreOffice 4/program'
|
|
39
48
|
].freeze
|
|
@@ -56,20 +65,46 @@ class WordToMarkdown
|
|
|
56
65
|
end
|
|
57
66
|
|
|
58
67
|
class << self
|
|
68
|
+
attr_writer :timeout
|
|
69
|
+
|
|
70
|
+
# @return [Numeric] seconds to wait for LibreOffice before giving up
|
|
71
|
+
def timeout
|
|
72
|
+
@timeout ||= Float(ENV.fetch('WORD_TO_MARKDOWN_TIMEOUT', DEFAULT_TIMEOUT))
|
|
73
|
+
end
|
|
74
|
+
|
|
59
75
|
# Run an soffice command
|
|
60
76
|
#
|
|
61
|
-
# @param args [string] one or more arguments to pass to the
|
|
77
|
+
# @param args [string] one or more arguments to pass to the soffice command
|
|
62
78
|
# @return [string] the command output
|
|
63
79
|
def run_command(*args)
|
|
64
80
|
raise 'LibreOffice already running' if soffice.open?
|
|
65
81
|
|
|
66
|
-
output, status =
|
|
82
|
+
output, status = capture_with_timeout(soffice.path, *args, timeout: timeout)
|
|
67
83
|
logger.debug output
|
|
68
84
|
raise "Command `#{soffice.path} #{args.join(' ')}` failed: #{output}" if status.exitstatus != 0
|
|
69
85
|
|
|
70
86
|
output
|
|
71
87
|
end
|
|
72
88
|
|
|
89
|
+
# Run a command, capturing its combined output, and kill it (along with
|
|
90
|
+
# any child processes) if it runs longer than the timeout
|
|
91
|
+
#
|
|
92
|
+
# @param command [Array<String>] the command and its arguments
|
|
93
|
+
# @param timeout [Numeric] seconds to wait before killing the command
|
|
94
|
+
# @return [Array(String, Process::Status)] the command output and status
|
|
95
|
+
def capture_with_timeout(*command, timeout:)
|
|
96
|
+
Open3.popen2e(*command, process_group_option) do |stdin, output, waiter|
|
|
97
|
+
stdin.close
|
|
98
|
+
reader = read_in_background(output)
|
|
99
|
+
unless waiter.join(timeout)
|
|
100
|
+
kill_process_group(waiter.pid)
|
|
101
|
+
raise TimeoutError, "Command `#{command.join(' ')}` timed out after #{timeout} seconds"
|
|
102
|
+
end
|
|
103
|
+
|
|
104
|
+
[reader.value, waiter.value]
|
|
105
|
+
end
|
|
106
|
+
end
|
|
107
|
+
|
|
73
108
|
# Returns a Cliver::Dependency object representing our soffice dependency
|
|
74
109
|
#
|
|
75
110
|
# Attempts to resolve by looking at PATH followed by paths in the PATHS constant
|
|
@@ -94,13 +129,38 @@ class WordToMarkdown
|
|
|
94
129
|
|
|
95
130
|
private
|
|
96
131
|
|
|
132
|
+
# Start commands in their own process group, so that on timeout any
|
|
133
|
+
# processes they spawn (e.g., soffice.bin) are killed too
|
|
134
|
+
def process_group_option
|
|
135
|
+
Gem.win_platform? ? { new_pgroup: true } : { pgroup: true }
|
|
136
|
+
end
|
|
137
|
+
|
|
138
|
+
# Read a stream in a separate thread, so the command can't block on a full pipe
|
|
139
|
+
#
|
|
140
|
+
# @param io [IO] the stream to read
|
|
141
|
+
# @return [Thread] a thread whose value is the stream's contents
|
|
142
|
+
def read_in_background(io)
|
|
143
|
+
Thread.new do
|
|
144
|
+
io.read
|
|
145
|
+
rescue IOError
|
|
146
|
+
'' # The stream was closed after the command timed out
|
|
147
|
+
end
|
|
148
|
+
end
|
|
149
|
+
|
|
150
|
+
# @param pid [Integer] the pid of a process group leader
|
|
151
|
+
def kill_process_group(pid)
|
|
152
|
+
Process.kill('KILL', Gem.win_platform? ? pid : -pid)
|
|
153
|
+
rescue Errno::ESRCH
|
|
154
|
+
nil
|
|
155
|
+
end
|
|
156
|
+
|
|
97
157
|
# Workaround for two upstream bugs:
|
|
98
|
-
# 1. `soffice.exe --version` on windows opens a popup and
|
|
158
|
+
# 1. `soffice.exe --version` on windows opens a popup and returns a null string when manually closed
|
|
99
159
|
# 2. Even if the second argument to Cliver is nil, Cliver thinks there's a requirement
|
|
100
160
|
# and will shell out to `soffice.exe --version`
|
|
101
161
|
# In order to support Windows, don't pass *any* version requirement to Cliver
|
|
102
162
|
def soffice_dependency_args
|
|
103
|
-
args = [path: PATHS.join(File::PATH_SEPARATOR)]
|
|
163
|
+
args = [{ path: PATHS.join(File::PATH_SEPARATOR) }]
|
|
104
164
|
if Gem.win_platform?
|
|
105
165
|
args
|
|
106
166
|
else
|
metadata
CHANGED
|
@@ -1,14 +1,13 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: word-to-markdown
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 1.
|
|
4
|
+
version: 1.2.0
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Ben Balter
|
|
8
|
-
autorequire:
|
|
9
8
|
bindir: bin
|
|
10
9
|
cert_chain: []
|
|
11
|
-
date:
|
|
10
|
+
date: 1980-01-02 00:00:00.000000000 Z
|
|
12
11
|
dependencies:
|
|
13
12
|
- !ruby/object:Gem::Dependency
|
|
14
13
|
name: cliver
|
|
@@ -25,19 +24,19 @@ dependencies:
|
|
|
25
24
|
- !ruby/object:Gem::Version
|
|
26
25
|
version: '0.3'
|
|
27
26
|
- !ruby/object:Gem::Dependency
|
|
28
|
-
name:
|
|
27
|
+
name: logger
|
|
29
28
|
requirement: !ruby/object:Gem::Requirement
|
|
30
29
|
requirements:
|
|
31
30
|
- - "~>"
|
|
32
31
|
- !ruby/object:Gem::Version
|
|
33
|
-
version: '
|
|
32
|
+
version: '1.4'
|
|
34
33
|
type: :runtime
|
|
35
34
|
prerelease: false
|
|
36
35
|
version_requirements: !ruby/object:Gem::Requirement
|
|
37
36
|
requirements:
|
|
38
37
|
- - "~>"
|
|
39
38
|
- !ruby/object:Gem::Version
|
|
40
|
-
version: '
|
|
39
|
+
version: '1.4'
|
|
41
40
|
- !ruby/object:Gem::Dependency
|
|
42
41
|
name: nokogiri-styles
|
|
43
42
|
requirement: !ruby/object:Gem::Requirement
|
|
@@ -92,42 +91,54 @@ dependencies:
|
|
|
92
91
|
requirements:
|
|
93
92
|
- - "~>"
|
|
94
93
|
- !ruby/object:Gem::Version
|
|
95
|
-
version: '1.
|
|
94
|
+
version: '1.3'
|
|
96
95
|
type: :runtime
|
|
97
96
|
prerelease: false
|
|
98
97
|
version_requirements: !ruby/object:Gem::Requirement
|
|
99
98
|
requirements:
|
|
100
99
|
- - "~>"
|
|
101
100
|
- !ruby/object:Gem::Version
|
|
102
|
-
version: '1.
|
|
101
|
+
version: '1.3'
|
|
103
102
|
- !ruby/object:Gem::Dependency
|
|
104
103
|
name: minitest
|
|
105
104
|
requirement: !ruby/object:Gem::Requirement
|
|
106
105
|
requirements:
|
|
107
|
-
- - "
|
|
106
|
+
- - ">="
|
|
107
|
+
- !ruby/object:Gem::Version
|
|
108
|
+
version: '5'
|
|
109
|
+
- - "<"
|
|
108
110
|
- !ruby/object:Gem::Version
|
|
109
|
-
version: '
|
|
111
|
+
version: '7'
|
|
110
112
|
type: :development
|
|
111
113
|
prerelease: false
|
|
112
114
|
version_requirements: !ruby/object:Gem::Requirement
|
|
113
115
|
requirements:
|
|
114
|
-
- - "
|
|
116
|
+
- - ">="
|
|
117
|
+
- !ruby/object:Gem::Version
|
|
118
|
+
version: '5'
|
|
119
|
+
- - "<"
|
|
115
120
|
- !ruby/object:Gem::Version
|
|
116
|
-
version: '
|
|
121
|
+
version: '7'
|
|
117
122
|
- !ruby/object:Gem::Dependency
|
|
118
123
|
name: mocha
|
|
119
124
|
requirement: !ruby/object:Gem::Requirement
|
|
120
125
|
requirements:
|
|
121
|
-
- - "
|
|
126
|
+
- - ">="
|
|
122
127
|
- !ruby/object:Gem::Version
|
|
123
|
-
version: '
|
|
128
|
+
version: '2'
|
|
129
|
+
- - "<"
|
|
130
|
+
- !ruby/object:Gem::Version
|
|
131
|
+
version: '4'
|
|
124
132
|
type: :development
|
|
125
133
|
prerelease: false
|
|
126
134
|
version_requirements: !ruby/object:Gem::Requirement
|
|
127
135
|
requirements:
|
|
128
|
-
- - "
|
|
136
|
+
- - ">="
|
|
137
|
+
- !ruby/object:Gem::Version
|
|
138
|
+
version: '2'
|
|
139
|
+
- - "<"
|
|
129
140
|
- !ruby/object:Gem::Version
|
|
130
|
-
version: '
|
|
141
|
+
version: '4'
|
|
131
142
|
- !ruby/object:Gem::Dependency
|
|
132
143
|
name: pry
|
|
133
144
|
requirement: !ruby/object:Gem::Requirement
|
|
@@ -227,13 +238,17 @@ files:
|
|
|
227
238
|
- lib/word-to-markdown.rb
|
|
228
239
|
- lib/word-to-markdown/converter.rb
|
|
229
240
|
- lib/word-to-markdown/document.rb
|
|
241
|
+
- lib/word-to-markdown/input_format.rb
|
|
242
|
+
- lib/word-to-markdown/url_scrubber.rb
|
|
230
243
|
- lib/word-to-markdown/version.rb
|
|
231
244
|
homepage: https://github.com/benbalter/word-to-markdown
|
|
232
245
|
licenses:
|
|
233
246
|
- MIT
|
|
234
247
|
metadata:
|
|
235
248
|
rubygems_mfa_required: 'true'
|
|
236
|
-
|
|
249
|
+
source_code_uri: https://github.com/benbalter/word-to-markdown
|
|
250
|
+
bug_tracker_uri: https://github.com/benbalter/word-to-markdown/issues
|
|
251
|
+
changelog_uri: https://github.com/benbalter/word-to-markdown/releases
|
|
237
252
|
rdoc_options: []
|
|
238
253
|
require_paths:
|
|
239
254
|
- lib
|
|
@@ -241,15 +256,14 @@ required_ruby_version: !ruby/object:Gem::Requirement
|
|
|
241
256
|
requirements:
|
|
242
257
|
- - ">="
|
|
243
258
|
- !ruby/object:Gem::Version
|
|
244
|
-
version: '
|
|
259
|
+
version: '3.2'
|
|
245
260
|
required_rubygems_version: !ruby/object:Gem::Requirement
|
|
246
261
|
requirements:
|
|
247
262
|
- - ">="
|
|
248
263
|
- !ruby/object:Gem::Version
|
|
249
264
|
version: '0'
|
|
250
265
|
requirements: []
|
|
251
|
-
rubygems_version: 3.
|
|
252
|
-
signing_key:
|
|
266
|
+
rubygems_version: 3.6.9
|
|
253
267
|
specification_version: 4
|
|
254
268
|
summary: Ruby Gem to convert Word documents to markdown
|
|
255
269
|
test_files: []
|