word-to-markdown 1.1.8 → 1.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 3febb4398acdc4eacedcc62e09f4beeaee625858043a27e6df7e597fee1e0d17
4
- data.tar.gz: 7a76057aeca2db8f321282bc309a835f282ea921343b28c8bce6d83cd0fc4582
3
+ metadata.gz: d5fff6778e90bf76409d47c5177fd2400a14308cf40be3ea2e3ee118917945b0
4
+ data.tar.gz: 18062303949a156da4c940daeb79854c820036fe623cdb9c4c07b1bce9dca242
5
5
  SHA512:
6
- metadata.gz: ee2340688c2d5f3f21c7e47f85220bebc88201e28200a682146041fa7f5a47e89c56d687b4f6565a95d7d3e7f4c70fb1551894ad297620e7bd49cef520975e18
7
- data.tar.gz: 44d405387990ee9a09cb33572a1d7843461c0817789c92a08df6d3542082ffc837bbd39d544a2e6a734a0772c67da3afed77049dc96b0ace0d562473515816e1
6
+ metadata.gz: f041232dbf6ae3aee4738d9562fab3f820a492dfd8d71fdd257397b278ca03550a27a59a0d0b5c3514254bab49367ba067970f7551ea11c3a7b90b8eda434499
7
+ data.tar.gz: e8c267ccd410ccfc978ef873805696299b1a658d1894dccae2f63d39f8afd12595ced23fbce6c8f4173188cc0b1e6fa44957d37c29c32b8ea18c08fb9a8745e1
data/README.md CHANGED
@@ -1,24 +1,29 @@
1
1
  # Word to Markdown converter
2
2
 
3
- A Ruby gem to liberate content from [the jail that is Word documents](http://ben.balter.com/2012/10/19/we-ve-been-trained-to-make-paper/#jailbreaking-content)
3
+ > [!IMPORTANT]
4
+ > **Looking for the latest and greatest?** Check out [**word-to-markdown-js**](https://github.com/benbalter/word-to-markdown-js), the newer, better successor to this project. It's a modern, actively maintained rewrite and is recommended for new projects. This Ruby gem remains available for existing users.
4
5
 
5
- [![Build Status](https://travis-ci.org/benbalter/word-to-markdown.svg?branch=master)](https://travis-ci.org/benbalter/word-to-markdown) [![Gem Version](https://badge.fury.io/rb/word-to-markdown.png)](http://badge.fury.io/rb/word-to-markdown) [![Inline docs](http://inch-ci.org/github/benbalter/word-to-markdown.png)](http://inch-ci.org/github/benbalter/word-to-markdown) [![Build status](https://ci.appveyor.com/api/projects/status/x2gnsfvli3q47a2e/branch/master?svg=true)](https://ci.appveyor.com/project/benbalter/word-to-markdown/branch/master)
6
+ A Ruby gem to liberate content from [the jail that is Word documents](https://ben.balter.com/2012/10/19/we-ve-been-trained-to-make-paper/#jailbreaking-content)
7
+
8
+ [![CI](https://github.com/benbalter/word-to-markdown/actions/workflows/ci.yml/badge.svg)](https://github.com/benbalter/word-to-markdown/actions/workflows/ci.yml) [![Gem Version](https://img.shields.io/gem/v/word-to-markdown)](https://rubygems.org/gems/word-to-markdown)
6
9
 
7
10
  ## The problem
8
11
 
9
- > Our default content publishing workflow is terribly broken. [We've all been trained to make paper](http://ben.balter.com/2012/10/19/we-ve-been-trained-to-make-paper/), yet today, content authored once is more commonly consumed in multiple formats, and rarely, if ever, does it embody physical form. Put another way, our go-to content authoring workflow remains relatively unchanged since it was conceived in the early 80s.
12
+ > Our default content publishing workflow is terribly broken. [We've all been trained to make paper](https://ben.balter.com/2012/10/19/we-ve-been-trained-to-make-paper/), yet today, content authored once is more commonly consumed in multiple formats, and rarely, if ever, does it embody physical form. Put another way, our go-to content authoring workflow remains relatively unchanged since it was conceived in the early 80s.
10
13
  >
11
- > I'm asked regularly by government employees — knowledge workers who fire up a desktop word processor as the first step to any project — for an automated pipeline to convert Microsoft Word documents to [Markdown](http://guides.github.com/overviews/mastering-markdown/), the *lingua franca* of the internet, but as my recent foray into building [just such a converter](http://word-to-markdown.herokuapp.com/) proves, it's not that simple.
14
+ > I'm asked regularly by government employees — knowledge workers who fire up a desktop word processor as the first step to any project — for an automated pipeline to convert Microsoft Word documents to [Markdown](https://docs.github.com/en/get-started/writing-on-github/getting-started-with-writing-and-formatting-on-github/basic-writing-and-formatting-syntax), the *lingua franca* of the internet, but as my recent foray into building [just such a converter](https://word2md.com/) proves, it's not that simple.
12
15
  >
13
16
  > Markdown isn't just an alternative format. Markdown forces you to write for the web.
14
17
 
15
- **[Read more](http://ben.balter.com/2014/03/31/word-versus-markdown-more-than-mere-semantics/)**
18
+ **[Read more](https://ben.balter.com/2014/03/31/word-versus-markdown-more-than-mere-semantics/)**
19
+
20
+ ## Just want to convert a Microsoft Word (or Google) document to Markdown?
16
21
 
17
- **[Demo](http://word-to-markdown.herokuapp.com/)**
22
+ You can use this **[hosted service](https://word2md.com/)** (or check out [its source](https://github.com/benbalter/word-to-markdown-server)).
18
23
 
19
24
  ## Install
20
25
 
21
- You'll need to install [LibreOffice](http://www.libreoffice.org/). Then:
26
+ You'll need to install [LibreOffice](https://www.libreoffice.org/). Then:
22
27
 
23
28
  ```bash
24
29
  gem install word-to-markdown
@@ -65,14 +70,30 @@ $ w2m path/to/document.docx
65
70
 
66
71
  Word-to-markdown requires `soffice` a command line interface to LibreOffice that works on Linux, Mac, and Windows. To install soffice, see [the LibreOffice documentation](https://www.libreoffice.org/get-help/install-howto/).
67
72
 
73
+ Word-to-markdown only accepts Word documents (`.docx` and `.doc`), identified by their contents rather than their file extension. Other files, including those LibreOffice could otherwise open, such as HTML, ODT, or RTF, raise `WordToMarkdown::Document::UnsupportedFormatError`.
74
+
75
+ LibreOffice is killed if a conversion takes longer than 60 seconds, raising `WordToMarkdown::TimeoutError`. To change the limit, set `WordToMarkdown.timeout = 120` or the `WORD_TO_MARKDOWN_TIMEOUT` environment variable.
76
+
77
+ ### Converting untrusted documents
78
+
79
+ Word documents can link to remote resources, such as images, which LibreOffice fetches while converting the document. If you convert documents from untrusted sources (for example, files uploaded to a web service), run the conversion without network access, such as in a container or sandbox with no outbound network, so that a document can't make requests to internal services or other hosts on your behalf.
80
+
68
81
  ## Testing
69
82
 
70
83
  ```
71
84
  script/cibuild
72
85
  ```
73
86
 
74
- ## Server
87
+ ## Docker
88
+
89
+ Everything you need to run the executable locally:
90
+
91
+ ```
92
+ docker compose build
93
+ docker compose run --rm app bundle exec w2m --help
94
+ docker compose run --rm app bundle exec w2m test/fixtures/em.docx
95
+ ```
75
96
 
76
- [Word-to-markdown-demo](https://github.com/benbalter/word-to-markdown-demo) contains a lightweight server for converting Word Documents as a service.
97
+ ## Hosted service
77
98
 
78
- A live version runs at [word-to-markdown.herokuapp.com](http://word-to-markdown.herokuapp.com).
99
+ [Word-to-markdown-server](https://github.com/benbalter/word-to-markdown-server) contains a lightweight server for converting Word Documents as a service. A live version runs at [word2md.com](https://word2md.com).
data/bin/w2m CHANGED
@@ -1,17 +1,35 @@
1
1
  #!/usr/bin/env ruby
2
2
  # frozen_string_literal: true
3
3
 
4
+ require 'optparse'
4
5
  require 'word-to-markdown'
5
6
 
6
- if ARGV.size != 1 || ARGV[0] == '--help'
7
- puts 'Usage: bundle exec w2m path/to/document.docx'
7
+ parser = OptionParser.new do |opts|
8
+ opts.banner = 'Usage: w2m path/to/document.docx'
9
+
10
+ opts.on('-h', '--help', 'Show this help') do
11
+ puts opts
12
+ exit
13
+ end
14
+
15
+ opts.on('-v', '--version', 'Show the WordToMarkdown and LibreOffice versions') do
16
+ puts "WordToMarkdown v#{WordToMarkdown::VERSION}"
17
+ puts "LibreOffice v#{WordToMarkdown.soffice.version}" unless Gem.win_platform?
18
+ exit
19
+ end
20
+ end
21
+
22
+ begin
23
+ parser.parse!
24
+ rescue OptionParser::ParseError => e
25
+ warn e.message
26
+ warn parser
8
27
  exit 1
9
28
  end
10
29
 
11
- if ARGV[0] == '--version'
12
- puts "WordToMarkdown v#{WordToMarkdown::VERSION}"
13
- puts "LibreOffice v#{WordToMarkdown.soffice.version}" unless Gem.win_platform?
14
- else
15
- doc = WordToMarkdown.new ARGV[0]
16
- puts doc.to_s
30
+ if ARGV.size != 1
31
+ warn parser
32
+ exit 1
17
33
  end
34
+
35
+ puts WordToMarkdown.new(ARGV[0])
@@ -24,14 +24,15 @@ module Cliver
24
24
 
25
25
  # Returns the version of the resolved dependency
26
26
  def version
27
- return @detected_version if defined? @detected_version
27
+ return @version if defined? @version
28
28
  return if Gem.win_platform?
29
+
29
30
  version = installed_versions.find { |p, _v| p == path }
30
- @detected_version = version.nil? ? nil : version[1]
31
+ @version = version.nil? ? nil : version[1]
31
32
  end
32
33
 
33
34
  def major_version
34
- version.split('.').first if version
35
+ version&.split('.')&.first
35
36
  end
36
37
  end
37
38
  end
@@ -7,14 +7,22 @@ class WordToMarkdown
7
7
  # Number of headings to guess, e.g., h6
8
8
  HEADING_DEPTH = 6
9
9
 
10
- # Percentile step for eaceh eheading
10
+ # Percentile step for each heading
11
11
  HEADING_STEP = 100 / HEADING_DEPTH
12
12
 
13
13
  # Minimum heading size
14
14
  MIN_HEADING_SIZE = 20
15
15
 
16
16
  # Unicode bullets to strip when processing
17
- UNICODE_BULLETS = ['○', 'o', '●', "\u2022", '\\p{C}'].freeze
17
+ UNICODE_BULLETS = ['○', '●', "\u2022", '\\p{C}'].freeze
18
+
19
+ # Leading bullets to strip from list items. A plain "o" only counts as a
20
+ # bullet when followed by whitespace, so words like "orange" survive.
21
+ BULLET_REGEX = /\A(?:[#{UNICODE_BULLETS.join}]|o(?=[[:space:]]))+[[:space:]]*/
22
+
23
+ # Leading list numbering to strip, e.g., "1.", "a.", or "iv.", along with
24
+ # any whitespace (including non-breaking spaces) that follows it
25
+ NUMBERING_REGEX = /\A(?:\d+|[a-zA-Z]|[ivxlcdm]+|[IVXLCDM]+)\.(?:[[:space:]]+|\z)/
18
26
 
19
27
  # @param document [WordToMarkdown::Document] The document to convert
20
28
  def initialize(document)
@@ -64,10 +72,11 @@ class WordToMarkdown
64
72
 
65
73
  # Given a Nokogiri node, guess what heading it represents, if any
66
74
  #
67
- # @param node [Nokigiri::Node] the nokigiri node
75
+ # @param node [Nokogiri::Node] the nokogiri node
68
76
  # @return [String, nil] the heading tag (e.g., H1), or nil
69
77
  def guess_heading(node)
70
78
  return nil if node.font_size.nil?
79
+
71
80
  [*1...HEADING_DEPTH].each do |heading|
72
81
  return "h#{heading}" if node.font_size >= h(heading)
73
82
  end
@@ -81,7 +90,23 @@ class WordToMarkdown
81
90
  #
82
91
  # @return [Integer] the minimum font size
83
92
  def h(num)
84
- font_sizes.percentile(((HEADING_DEPTH - 1) - num) * HEADING_STEP)
93
+ self.class.percentile(font_sizes, ((HEADING_DEPTH - 1) - num) * HEADING_STEP)
94
+ end
95
+
96
+ # Linearly interpolated percentile, matching the algorithm used by the
97
+ # descriptive_statistics gem (v2.5.1) that this replaces
98
+ #
99
+ # @param values [Array<Numeric>] the values
100
+ # @param pct [Numeric] the percentile, from 0 to 100
101
+ #
102
+ # @return [Float, nil] the percentile, or nil if values is empty
103
+ def self.percentile(values, pct)
104
+ sorted = values.map(&:to_f).sort
105
+ rank = pct / 100.0 * (sorted.size - 1)
106
+ lower, upper = sorted[rank.floor, 2]
107
+ return lower if upper.nil? # empty or single-element collection, or the 100th percentile
108
+
109
+ lower + ((upper - lower) * (rank - rank.floor))
85
110
  end
86
111
 
87
112
  # Convert span-based font styles to `strong`s and `em`s
@@ -109,7 +134,7 @@ class WordToMarkdown
109
134
  def remove_unicode_bullets_from_list_items!
110
135
  path = WordToMarkdown.soffice.major_version == '5' ? 'li span span' : 'li span'
111
136
  @document.tree.search(path).each do |span|
112
- span.inner_html = span.inner_html.gsub(/^([#{UNICODE_BULLETS.join("")}]+)/, '')
137
+ span.inner_html = span.inner_html.sub(BULLET_REGEX, '')
113
138
  end
114
139
  end
115
140
 
@@ -117,21 +142,21 @@ class WordToMarkdown
117
142
  def remove_numbering_from_list_items!
118
143
  path = WordToMarkdown.soffice.major_version == '5' ? 'li span span' : 'li span'
119
144
  @document.tree.search(path).each do |span|
120
- span.inner_html = span.inner_html.gsub(/^[a-zA-Z0-9]+\./m, '')
145
+ span.inner_html = span.inner_html.sub(NUMBERING_REGEX, '')
121
146
  end
122
147
  end
123
148
 
124
- # Remvoe whitespace from list items
149
+ # Remove whitespace from list items
125
150
  def remove_whitespace_from_list_items!
126
- @document.tree.search('li span').each { |span| span.inner_html.strip! }
151
+ @document.tree.search('li span').each { |span| span.inner_html = span.inner_html.strip }
127
152
  end
128
153
 
129
- # Convert table headers to `th`s2
154
+ # Convert table headers to `th`s
130
155
  def semanticize_table_headers!
131
156
  @document.tree.search('table tr:first td').each { |node| node.node_name = 'th' }
132
157
  end
133
158
 
134
- # Try to guess heading where implicit bassed on font size
159
+ # Try to guess heading where implicit based on font size
135
160
  def semanticize_headings!
136
161
  implicit_headings.each do |element|
137
162
  heading = guess_heading element
@@ -3,16 +3,24 @@
3
3
  class WordToMarkdown
4
4
  class Document
5
5
  class NotFoundError < StandardError; end
6
+
6
7
  class ConversionError < StandardError; end
7
8
 
8
- attr_reader :path, :tmpdir
9
+ class UnsupportedFormatError < StandardError; end
10
+
11
+ attr_reader :path, :tmpdir, :import_filter
9
12
 
10
13
  # @param path [string] Path to the Word document
11
14
  # @param tmpdir [string] Path to a working directory to use
12
15
  def initialize(path, tmpdir = nil)
13
16
  @path = File.expand_path path, Dir.pwd
14
- @tmpdir = tmpdir || Dir.mktmpdir
15
17
  raise NotFoundError, "File #{@path} does not exist" unless File.exist?(@path)
18
+
19
+ @import_filter = InputFormat.filter_for(@path)
20
+ raise UnsupportedFormatError, "File #{@path} is not a Word document (.docx or .doc)" if @import_filter.nil?
21
+
22
+ @own_tmpdir = tmpdir.nil?
23
+ @tmpdir = tmpdir || Dir.mktmpdir
16
24
  end
17
25
 
18
26
  # @return [String] the document's extension
@@ -20,7 +28,7 @@ class WordToMarkdown
20
28
  File.extname path
21
29
  end
22
30
 
23
- # @return [Nokigiri::Document]
31
+ # @return [Nokogiri::Document]
24
32
  def tree
25
33
  @tree ||= begin
26
34
  tree = Nokogiri::HTML(normalized_html)
@@ -31,6 +39,7 @@ class WordToMarkdown
31
39
 
32
40
  # @return [String] the html representation of the document
33
41
  def html
42
+ UrlScrubber.scrub!(tree)
34
43
  tree.to_html.gsub("</li>\n", '</li>')
35
44
  end
36
45
 
@@ -44,7 +53,7 @@ class WordToMarkdown
44
53
  #
45
54
  # @return [String] the encoding, defaulting to "UTF-8"
46
55
  def encoding
47
- match = raw_html.encode('UTF-8', invalid: :replace, replace: '').match(/charset=([^\"]+)/)
56
+ match = raw_html.encode('UTF-8', invalid: :replace, replace: '').match(/charset=([^"]+)/)
48
57
  if match
49
58
  match[1].sub('macintosh', 'MacRoman')
50
59
  else
@@ -78,30 +87,40 @@ class WordToMarkdown
78
87
  string.gsub!('&nbsp;', ' ') # HTML encoded spaces
79
88
  string.sub!(/\A[[:space:]]+/, '') # document leading whitespace
80
89
  string.sub!(/[[:space:]]+\z/, '') # document trailing whitespace
81
- string.gsub!(/([ ]+)$/, '') # line trailing whitespace
82
- string.gsub!(/\n\n\n\n/, "\n\n") # Quadruple line breaks
90
+ string.gsub!(/( +)$/, '') # line trailing whitespace
91
+ string.gsub!("\n\n\n\n", "\n\n") # Quadruple line breaks
83
92
  string.delete!(' ') # Unicode non-breaking spaces, injected as tabs
93
+ string.gsub!(/\*\*\ +(?!\*|_)([[:punct:]])/, '**\1') # Remove extra space after bold
84
94
  string
85
95
  end
86
96
 
87
97
  # @return [String] the path to the intermediary HTML document
88
98
  def dest_path
89
- dest_filename = File.basename(path).gsub(/#{Regexp.escape(extension)}$/, '.html')
90
- File.expand_path(dest_filename, tmpdir)
99
+ File.expand_path("#{File.basename(path, '.*')}.html", tmpdir)
91
100
  end
92
101
 
93
102
  # @return [String] the unnormalized HTML representation
94
103
  def raw_html
95
104
  @raw_html ||= begin
96
- WordToMarkdown.run_command '--headless', '--convert-to', filter, path, '--outdir', tmpdir
105
+ WordToMarkdown.run_command '--headless', "--infilter=#{import_filter}", '--convert-to', filter, path, '--outdir', tmpdir
97
106
  raise ConversionError, "Failed to convert #{path}" unless File.exist?(dest_path)
107
+
98
108
  html = File.read dest_path
99
109
  File.delete dest_path
100
110
  html
111
+ ensure
112
+ remove_tmpdir
101
113
  end
102
114
  end
103
115
 
104
- # @return [String] the LibreOffice filter to use for conversion
116
+ # Remove the working directory if we created it and nothing else is in it.
117
+ # Non-empty directories are kept, since LibreOffice may have written
118
+ # images there that the markdown references.
119
+ def remove_tmpdir
120
+ Dir.rmdir(tmpdir) if @own_tmpdir && Dir.empty?(tmpdir)
121
+ end
122
+
123
+ # @return [String] the LibreOffice filter to use for export
105
124
  def filter
106
125
  if WordToMarkdown.soffice.major_version == '5'
107
126
  'html:XHTML Writer File:UTF8'
@@ -0,0 +1,86 @@
1
+ # frozen_string_literal: true
2
+
3
+ class WordToMarkdown
4
+ # Identifies Word documents by their content, rather than their extension,
5
+ # so that other formats LibreOffice would otherwise sniff and import (e.g.,
6
+ # HTML, which can reference remote resources) are never passed to it
7
+ module InputFormat
8
+ # Signature of an OLE compound file, used by .doc files
9
+ OLE_SIGNATURE = "\xD0\xCF\x11\xE0\xA1\xB1\x1A\xE1".b.freeze
10
+
11
+ # Signature of a ZIP local file header, used by .docx files
12
+ ZIP_SIGNATURE = "PK\x03\x04".b.freeze
13
+
14
+ # Signature of a ZIP end of central directory record
15
+ ZIP_EOCD_SIGNATURE = "PK\x05\x06".b.freeze
16
+
17
+ # Size of the end of central directory record, excluding the comment
18
+ ZIP_EOCD_SIZE = 22
19
+
20
+ # Maximum length of a ZIP archive comment
21
+ ZIP_MAX_COMMENT_SIZE = 0xFFFF
22
+
23
+ # The main document part every Word OOXML package contains
24
+ OOXML_DOCUMENT_PART = 'word/document.xml'.b.freeze
25
+
26
+ # LibreOffice import filters, by format
27
+ FILTERS = {
28
+ ooxml: 'MS Word 2007 XML',
29
+ ole: 'MS Word 97'
30
+ }.freeze
31
+
32
+ class << self
33
+ # @param path [String] path to the file
34
+ # @return [String, nil] the LibreOffice import filter for the file, or nil if it isn't a Word document
35
+ def filter_for(path)
36
+ FILTERS[detect(path)]
37
+ end
38
+
39
+ # @param path [String] path to the file
40
+ # @return [Symbol, nil] :ooxml for .docx, :ole for .doc, or nil if it isn't a Word document
41
+ def detect(path)
42
+ File.open(path, 'rb') do |file|
43
+ signature = file.read(OLE_SIGNATURE.bytesize).to_s
44
+ if signature == OLE_SIGNATURE
45
+ :ole
46
+ elsif signature.start_with?(ZIP_SIGNATURE) && ooxml_document?(file)
47
+ :ooxml
48
+ end
49
+ end
50
+ end
51
+
52
+ private
53
+
54
+ # @param file [File] an open ZIP file
55
+ # @return [Boolean] true if the ZIP's central directory lists the Word main document part
56
+ def ooxml_document?(file)
57
+ directory = central_directory(file)
58
+ !directory.nil? && directory.include?(OOXML_DOCUMENT_PART)
59
+ end
60
+
61
+ # Read the ZIP central directory, whose entries include each file name, uncompressed
62
+ #
63
+ # @param file [File] an open ZIP file
64
+ # @return [String, nil] the raw central directory, or nil if it can't be found
65
+ def central_directory(file)
66
+ size, offset = central_directory_location(file)
67
+ return if size.nil? || offset + size > file.size
68
+
69
+ file.seek(offset)
70
+ file.read(size)
71
+ end
72
+
73
+ # @param file [File] an open ZIP file
74
+ # @return [Array<Integer>, nil] the central directory's size and offset, or nil if not found
75
+ def central_directory_location(file)
76
+ tail_size = [file.size, ZIP_EOCD_SIZE + ZIP_MAX_COMMENT_SIZE].min
77
+ file.seek(-tail_size, IO::SEEK_END)
78
+ tail = file.read(tail_size)
79
+ index = tail.rindex(ZIP_EOCD_SIGNATURE)
80
+ return if index.nil? || tail.bytesize - index < ZIP_EOCD_SIZE
81
+
82
+ tail.byteslice(index + 12, 8).unpack('VV')
83
+ end
84
+ end
85
+ end
86
+ end
@@ -0,0 +1,110 @@
1
+ # frozen_string_literal: true
2
+
3
+ class WordToMarkdown
4
+ # Removes links and images with unsafe URL schemes (e.g., javascript: or
5
+ # vbscript:) so that they don't survive into the markdown output
6
+ module UrlScrubber
7
+ # URL schemes permitted in link targets. Links with any other scheme are
8
+ # unwrapped to their text. Relative URLs and fragments are always permitted.
9
+ SAFE_LINK_SCHEMES = %w[http https mailto].freeze
10
+
11
+ # URL schemes permitted in image sources, in addition to data:image/ URIs.
12
+ # Images with any other scheme are removed.
13
+ SAFE_IMAGE_SCHEMES = %w[http https].freeze
14
+
15
+ # Browsers remove ASCII tabs and newlines anywhere in a URL, and C0
16
+ # control characters and spaces at either end, before parsing the scheme
17
+ STRIPPED_CHARS = /[\t\n\r]/
18
+ EDGE_CHARS = /\A[\x00-\x20]+|[\x00-\x20]+\z/
19
+
20
+ # Matches a URL scheme, e.g., "https:"
21
+ SCHEME_REGEX = /\A([a-z][a-z0-9+.-]*):/i
22
+
23
+ # Matches a URL without a recognizable scheme whose first segment contains
24
+ # characters a Markdown renderer or browser may decode into a scheme
25
+ # separator, e.g., "javascript&colon;", "javascript&#58;", or "javascript\:"
26
+ AMBIGUOUS_REGEX = %r{\A[^/?#]*[:&\\]}
27
+
28
+ # Characters that could end or alter a Markdown link destination, which
29
+ # are percent-encoded in the URLs that are kept
30
+ DESTINATION_UNSAFE_CHARS = /[\x00-\x20\x7F()<>\[\]\\`"]/
31
+
32
+ # Matches a data URI for an image
33
+ DATA_IMAGE_REGEX = %r{\Adata:image/}i
34
+
35
+ class << self
36
+ # Unwrap links and remove images whose URL scheme isn't permitted
37
+ #
38
+ # @param tree [Nokogiri::HTML::Document] the document to scrub, in place
39
+ # @return [Nokogiri::HTML::Document] the scrubbed document
40
+ def scrub!(tree)
41
+ tree.css('a[href]').each do |node|
42
+ safe_link?(node['href']) ? escape_attributes!(node, 'href') : node.replace(node.children)
43
+ end
44
+
45
+ tree.css('img[src]').each do |node|
46
+ safe_image?(node['src']) ? escape_attributes!(node, 'src') : node.remove
47
+ end
48
+
49
+ tree
50
+ end
51
+
52
+ # Percent-encode characters in a URL that could end or alter a Markdown link destination
53
+ #
54
+ # @param url [String] the URL
55
+ # @return [String] the escaped URL
56
+ def escape_destination(url)
57
+ url.gsub(DESTINATION_UNSAFE_CHARS) { |char| format('%%%02X', char.ord) }
58
+ end
59
+
60
+ # @param url [String] a link target
61
+ # @return [Boolean] true if the URL is relative, a fragment, or has a permitted scheme
62
+ def safe_link?(url)
63
+ scheme = scheme(url)
64
+ scheme.nil? ? relative?(url) : SAFE_LINK_SCHEMES.include?(scheme)
65
+ end
66
+
67
+ # @param url [String] an image source
68
+ # @return [Boolean] true if the URL is relative, a data:image/ URI, or has a permitted scheme
69
+ def safe_image?(url)
70
+ scheme = scheme(url)
71
+ return relative?(url) if scheme.nil?
72
+
73
+ SAFE_IMAGE_SCHEMES.include?(scheme) || normalize(url).match?(DATA_IMAGE_REGEX)
74
+ end
75
+
76
+ # @param url [String] the URL
77
+ # @return [String, nil] the URL's lowercased scheme, or nil if it has none
78
+ def scheme(url)
79
+ match = normalize(url).match(SCHEME_REGEX)
80
+ match && match[1].downcase
81
+ end
82
+
83
+ private
84
+
85
+ # Escape a kept link or image's URL and title so they can't break out
86
+ # of the Markdown link that ReverseMarkdown writes for them
87
+ #
88
+ # @param node [Nokogiri::XML::Element] the link or image
89
+ # @param attribute [String] the name of the URL attribute
90
+ def escape_attributes!(node, attribute)
91
+ node[attribute] = escape_destination(node[attribute])
92
+ node['title'] = node['title'].tr('"', "'") if node['title']
93
+ end
94
+
95
+ # @param url [String] a URL without a scheme
96
+ # @return [Boolean] true if the URL is a relative path or fragment that can't be decoded into one with a scheme
97
+ def relative?(url)
98
+ !normalize(url).match?(AMBIGUOUS_REGEX)
99
+ end
100
+
101
+ # Normalize a URL the way a browser does before parsing its scheme
102
+ #
103
+ # @param url [String] the URL
104
+ # @return [String] the normalized URL
105
+ def normalize(url)
106
+ url.to_s.gsub(STRIPPED_CHARS, '').gsub(EDGE_CHARS, '')
107
+ end
108
+ end
109
+ end
110
+ end
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  class WordToMarkdown
4
- VERSION = '1.1.8'.freeze
4
+ VERSION = '1.2.0'
5
5
  end
@@ -1,6 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
- require 'descriptive_statistics'
4
3
  require 'reverse_markdown'
5
4
  require 'nokogiri-styles'
6
5
  require 'premailer'
@@ -12,28 +11,38 @@ require 'cliver'
12
11
  require 'open3'
13
12
 
14
13
  require_relative 'word-to-markdown/version'
14
+ require_relative 'word-to-markdown/input_format'
15
+ require_relative 'word-to-markdown/url_scrubber'
15
16
  require_relative 'word-to-markdown/document'
16
17
  require_relative 'word-to-markdown/converter'
17
18
  require_relative 'nokogiri/xml/element'
18
19
  require_relative 'cliver/dependency_ext'
19
20
 
20
21
  class WordToMarkdown
22
+ class TimeoutError < StandardError; end
23
+
21
24
  attr_reader :document, :converter
22
25
 
23
26
  # Options to be passed to Reverse Markdown
24
27
  REVERSE_MARKDOWN_OPTIONS = {
25
- unknown_tags: :bypass,
28
+ unknown_tags: :bypass,
26
29
  github_flavored: true
27
30
  }.freeze
28
31
 
32
+ # Default number of seconds to wait for LibreOffice to convert a document
33
+ # before giving up. Can be overridden with the WORD_TO_MARKDOWN_TIMEOUT
34
+ # environment variable, or by setting WordToMarkdown.timeout
35
+ DEFAULT_TIMEOUT = 60
36
+
29
37
  # Minimum version of LibreOffice Required
30
- SOFFICE_VERSION_REQUIREMENT = '> 4.0'.freeze
38
+ SOFFICE_VERSION_REQUIREMENT = '> 4.0'
31
39
 
32
40
  # Paths to look for LibreOffice, in order of preference
33
41
  PATHS = [
34
42
  '*', # Sub'd for ENV["PATH"]
35
43
  '~/Applications/LibreOffice.app/Contents/MacOS',
36
44
  '/Applications/LibreOffice.app/Contents/MacOS',
45
+ '/Program Files/LibreOffice/program',
37
46
  '/Program Files/LibreOffice 5/program',
38
47
  '/Program Files (x86)/LibreOffice 4/program'
39
48
  ].freeze
@@ -56,19 +65,46 @@ class WordToMarkdown
56
65
  end
57
66
 
58
67
  class << self
68
+ attr_writer :timeout
69
+
70
+ # @return [Numeric] seconds to wait for LibreOffice before giving up
71
+ def timeout
72
+ @timeout ||= Float(ENV.fetch('WORD_TO_MARKDOWN_TIMEOUT', DEFAULT_TIMEOUT))
73
+ end
74
+
59
75
  # Run an soffice command
60
76
  #
61
- # @param args [string] one or more arguments to pass to the sofice command
77
+ # @param args [string] one or more arguments to pass to the soffice command
62
78
  # @return [string] the command output
63
79
  def run_command(*args)
64
80
  raise 'LibreOffice already running' if soffice.open?
65
81
 
66
- output, status = Open3.capture2e(soffice.path, *args)
82
+ output, status = capture_with_timeout(soffice.path, *args, timeout: timeout)
67
83
  logger.debug output
68
84
  raise "Command `#{soffice.path} #{args.join(' ')}` failed: #{output}" if status.exitstatus != 0
85
+
69
86
  output
70
87
  end
71
88
 
89
+ # Run a command, capturing its combined output, and kill it (along with
90
+ # any child processes) if it runs longer than the timeout
91
+ #
92
+ # @param command [Array<String>] the command and its arguments
93
+ # @param timeout [Numeric] seconds to wait before killing the command
94
+ # @return [Array(String, Process::Status)] the command output and status
95
+ def capture_with_timeout(*command, timeout:)
96
+ Open3.popen2e(*command, process_group_option) do |stdin, output, waiter|
97
+ stdin.close
98
+ reader = read_in_background(output)
99
+ unless waiter.join(timeout)
100
+ kill_process_group(waiter.pid)
101
+ raise TimeoutError, "Command `#{command.join(' ')}` timed out after #{timeout} seconds"
102
+ end
103
+
104
+ [reader.value, waiter.value]
105
+ end
106
+ end
107
+
72
108
  # Returns a Cliver::Dependency object representing our soffice dependency
73
109
  #
74
110
  # Attempts to resolve by looking at PATH followed by paths in the PATHS constant
@@ -85,7 +121,7 @@ class WordToMarkdown
85
121
  # @return Logger instance
86
122
  def logger
87
123
  @logger ||= begin
88
- logger = Logger.new(STDOUT)
124
+ logger = Logger.new($stdout)
89
125
  logger.level = Logger::ERROR unless ENV['DEBUG']
90
126
  logger
91
127
  end
@@ -93,13 +129,38 @@ class WordToMarkdown
93
129
 
94
130
  private
95
131
 
132
+ # Start commands in their own process group, so that on timeout any
133
+ # processes they spawn (e.g., soffice.bin) are killed too
134
+ def process_group_option
135
+ Gem.win_platform? ? { new_pgroup: true } : { pgroup: true }
136
+ end
137
+
138
+ # Read a stream in a separate thread, so the command can't block on a full pipe
139
+ #
140
+ # @param io [IO] the stream to read
141
+ # @return [Thread] a thread whose value is the stream's contents
142
+ def read_in_background(io)
143
+ Thread.new do
144
+ io.read
145
+ rescue IOError
146
+ '' # The stream was closed after the command timed out
147
+ end
148
+ end
149
+
150
+ # @param pid [Integer] the pid of a process group leader
151
+ def kill_process_group(pid)
152
+ Process.kill('KILL', Gem.win_platform? ? pid : -pid)
153
+ rescue Errno::ESRCH
154
+ nil
155
+ end
156
+
96
157
  # Workaround for two upstream bugs:
97
- # 1. `soffice.exe --version` on windows opens a popup and retuns a null string when manually closed
158
+ # 1. `soffice.exe --version` on windows opens a popup and returns a null string when manually closed
98
159
  # 2. Even if the second argument to Cliver is nil, Cliver thinks there's a requirement
99
160
  # and will shell out to `soffice.exe --version`
100
161
  # In order to support Windows, don't pass *any* version requirement to Cliver
101
162
  def soffice_dependency_args
102
- args = [path: PATHS.join(File::PATH_SEPARATOR)]
163
+ args = [{ path: PATHS.join(File::PATH_SEPARATOR) }]
103
164
  if Gem.win_platform?
104
165
  args
105
166
  else
metadata CHANGED
@@ -1,14 +1,13 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: word-to-markdown
3
3
  version: !ruby/object:Gem::Version
4
- version: 1.1.8
4
+ version: 1.2.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Ben Balter
8
- autorequire:
9
8
  bindir: bin
10
9
  cert_chain: []
11
- date: 2018-08-01 00:00:00.000000000 Z
10
+ date: 1980-01-02 00:00:00.000000000 Z
12
11
  dependencies:
13
12
  - !ruby/object:Gem::Dependency
14
13
  name: cliver
@@ -25,19 +24,19 @@ dependencies:
25
24
  - !ruby/object:Gem::Version
26
25
  version: '0.3'
27
26
  - !ruby/object:Gem::Dependency
28
- name: descriptive_statistics
27
+ name: logger
29
28
  requirement: !ruby/object:Gem::Requirement
30
29
  requirements:
31
30
  - - "~>"
32
31
  - !ruby/object:Gem::Version
33
- version: '2.5'
32
+ version: '1.4'
34
33
  type: :runtime
35
34
  prerelease: false
36
35
  version_requirements: !ruby/object:Gem::Requirement
37
36
  requirements:
38
37
  - - "~>"
39
38
  - !ruby/object:Gem::Version
40
- version: '2.5'
39
+ version: '1.4'
41
40
  - !ruby/object:Gem::Dependency
42
41
  name: nokogiri-styles
43
42
  requirement: !ruby/object:Gem::Requirement
@@ -70,128 +69,160 @@ dependencies:
70
69
  name: reverse_markdown
71
70
  requirement: !ruby/object:Gem::Requirement
72
71
  requirements:
73
- - - "~>"
72
+ - - ">="
74
73
  - !ruby/object:Gem::Version
75
- version: '1.0'
74
+ version: '1'
75
+ - - "<"
76
+ - !ruby/object:Gem::Version
77
+ version: '3'
76
78
  type: :runtime
77
79
  prerelease: false
78
80
  version_requirements: !ruby/object:Gem::Requirement
79
81
  requirements:
80
- - - "~>"
82
+ - - ">="
81
83
  - !ruby/object:Gem::Version
82
- version: '1.0'
84
+ version: '1'
85
+ - - "<"
86
+ - !ruby/object:Gem::Version
87
+ version: '3'
83
88
  - !ruby/object:Gem::Dependency
84
89
  name: sys-proctable
85
90
  requirement: !ruby/object:Gem::Requirement
86
91
  requirements:
87
92
  - - "~>"
88
93
  - !ruby/object:Gem::Version
89
- version: '1.0'
94
+ version: '1.3'
90
95
  type: :runtime
91
96
  prerelease: false
92
97
  version_requirements: !ruby/object:Gem::Requirement
93
98
  requirements:
94
99
  - - "~>"
95
100
  - !ruby/object:Gem::Version
96
- version: '1.0'
101
+ version: '1.3'
97
102
  - !ruby/object:Gem::Dependency
98
- name: bundler
103
+ name: minitest
99
104
  requirement: !ruby/object:Gem::Requirement
100
105
  requirements:
101
- - - "~>"
106
+ - - ">="
107
+ - !ruby/object:Gem::Version
108
+ version: '5'
109
+ - - "<"
102
110
  - !ruby/object:Gem::Version
103
- version: '1.6'
111
+ version: '7'
104
112
  type: :development
105
113
  prerelease: false
106
114
  version_requirements: !ruby/object:Gem::Requirement
107
115
  requirements:
108
- - - "~>"
116
+ - - ">="
117
+ - !ruby/object:Gem::Version
118
+ version: '5'
119
+ - - "<"
109
120
  - !ruby/object:Gem::Version
110
- version: '1.6'
121
+ version: '7'
111
122
  - !ruby/object:Gem::Dependency
112
- name: minitest
123
+ name: mocha
124
+ requirement: !ruby/object:Gem::Requirement
125
+ requirements:
126
+ - - ">="
127
+ - !ruby/object:Gem::Version
128
+ version: '2'
129
+ - - "<"
130
+ - !ruby/object:Gem::Version
131
+ version: '4'
132
+ type: :development
133
+ prerelease: false
134
+ version_requirements: !ruby/object:Gem::Requirement
135
+ requirements:
136
+ - - ">="
137
+ - !ruby/object:Gem::Version
138
+ version: '2'
139
+ - - "<"
140
+ - !ruby/object:Gem::Version
141
+ version: '4'
142
+ - !ruby/object:Gem::Dependency
143
+ name: pry
113
144
  requirement: !ruby/object:Gem::Requirement
114
145
  requirements:
115
146
  - - "~>"
116
147
  - !ruby/object:Gem::Version
117
- version: '5.0'
148
+ version: '0.10'
118
149
  type: :development
119
150
  prerelease: false
120
151
  version_requirements: !ruby/object:Gem::Requirement
121
152
  requirements:
122
153
  - - "~>"
123
154
  - !ruby/object:Gem::Version
124
- version: '5.0'
155
+ version: '0.10'
125
156
  - !ruby/object:Gem::Dependency
126
- name: mocha
157
+ name: rake
127
158
  requirement: !ruby/object:Gem::Requirement
128
159
  requirements:
129
160
  - - "~>"
130
161
  - !ruby/object:Gem::Version
131
- version: '1.1'
162
+ version: '13.0'
132
163
  type: :development
133
164
  prerelease: false
134
165
  version_requirements: !ruby/object:Gem::Requirement
135
166
  requirements:
136
167
  - - "~>"
137
168
  - !ruby/object:Gem::Version
138
- version: '1.1'
169
+ version: '13.0'
139
170
  - !ruby/object:Gem::Dependency
140
- name: pry
171
+ name: rubocop
141
172
  requirement: !ruby/object:Gem::Requirement
142
173
  requirements:
143
174
  - - "~>"
144
175
  - !ruby/object:Gem::Version
145
- version: '0.10'
176
+ version: '1.0'
146
177
  type: :development
147
178
  prerelease: false
148
179
  version_requirements: !ruby/object:Gem::Requirement
149
180
  requirements:
150
181
  - - "~>"
151
182
  - !ruby/object:Gem::Version
152
- version: '0.10'
183
+ version: '1.0'
153
184
  - !ruby/object:Gem::Dependency
154
- name: rake
185
+ name: rubocop-minitest
155
186
  requirement: !ruby/object:Gem::Requirement
156
187
  requirements:
157
188
  - - "~>"
158
189
  - !ruby/object:Gem::Version
159
- version: '10.4'
190
+ version: '0.3'
160
191
  type: :development
161
192
  prerelease: false
162
193
  version_requirements: !ruby/object:Gem::Requirement
163
194
  requirements:
164
195
  - - "~>"
165
196
  - !ruby/object:Gem::Version
166
- version: '10.4'
197
+ version: '0.3'
167
198
  - !ruby/object:Gem::Dependency
168
- name: rubocop
199
+ name: rubocop-performance
169
200
  requirement: !ruby/object:Gem::Requirement
170
201
  requirements:
171
202
  - - "~>"
172
203
  - !ruby/object:Gem::Version
173
- version: '0.49'
204
+ version: '1.5'
174
205
  type: :development
175
206
  prerelease: false
176
207
  version_requirements: !ruby/object:Gem::Requirement
177
208
  requirements:
178
209
  - - "~>"
179
210
  - !ruby/object:Gem::Version
180
- version: '0.49'
211
+ version: '1.5'
181
212
  - !ruby/object:Gem::Dependency
182
213
  name: shoulda
183
214
  requirement: !ruby/object:Gem::Requirement
184
215
  requirements:
185
216
  - - "~>"
186
217
  - !ruby/object:Gem::Version
187
- version: '3.5'
218
+ version: '4.0'
188
219
  type: :development
189
220
  prerelease: false
190
221
  version_requirements: !ruby/object:Gem::Requirement
191
222
  requirements:
192
223
  - - "~>"
193
224
  - !ruby/object:Gem::Version
194
- version: '3.5'
225
+ version: '4.0'
195
226
  description: Ruby Gem to convert Word documents to markdown.
196
227
  email: ben.balter@github.com
197
228
  executables:
@@ -207,12 +238,17 @@ files:
207
238
  - lib/word-to-markdown.rb
208
239
  - lib/word-to-markdown/converter.rb
209
240
  - lib/word-to-markdown/document.rb
241
+ - lib/word-to-markdown/input_format.rb
242
+ - lib/word-to-markdown/url_scrubber.rb
210
243
  - lib/word-to-markdown/version.rb
211
244
  homepage: https://github.com/benbalter/word-to-markdown
212
245
  licenses:
213
246
  - MIT
214
- metadata: {}
215
- post_install_message:
247
+ metadata:
248
+ rubygems_mfa_required: 'true'
249
+ source_code_uri: https://github.com/benbalter/word-to-markdown
250
+ bug_tracker_uri: https://github.com/benbalter/word-to-markdown/issues
251
+ changelog_uri: https://github.com/benbalter/word-to-markdown/releases
216
252
  rdoc_options: []
217
253
  require_paths:
218
254
  - lib
@@ -220,16 +256,14 @@ required_ruby_version: !ruby/object:Gem::Requirement
220
256
  requirements:
221
257
  - - ">="
222
258
  - !ruby/object:Gem::Version
223
- version: '0'
259
+ version: '3.2'
224
260
  required_rubygems_version: !ruby/object:Gem::Requirement
225
261
  requirements:
226
262
  - - ">="
227
263
  - !ruby/object:Gem::Version
228
264
  version: '0'
229
265
  requirements: []
230
- rubyforge_project:
231
- rubygems_version: 2.7.6
232
- signing_key:
266
+ rubygems_version: 3.6.9
233
267
  specification_version: 4
234
268
  summary: Ruby Gem to convert Word documents to markdown
235
269
  test_files: []