word-to-markdown 1.1.9 → 1.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: f2b816a2ad9402eb1c45f74f806482502608337ac9127d072647bd4f1f97cd39
4
- data.tar.gz: ff35f5c9f2e89c0e781ea864552f6f20bafd23083301c4d72c5c95c7aae4f38c
3
+ metadata.gz: d5fff6778e90bf76409d47c5177fd2400a14308cf40be3ea2e3ee118917945b0
4
+ data.tar.gz: 18062303949a156da4c940daeb79854c820036fe623cdb9c4c07b1bce9dca242
5
5
  SHA512:
6
- metadata.gz: e74c055913709cd0fa871ba95cf22b22a86089bb18fd9e88cd27d1dec3ec3b4927708fc9a591ba278b3e5cdc98150178cd17547ba76579702893645554370c83
7
- data.tar.gz: 9b01d816e8f95d43fb19dc828213e3db4f7b0897957c6e3eba42b4f4fe888d3564db3b3730b416f9eb10837ca4911fd7083c73b3556c3313dd12cefe2f8ae597
6
+ metadata.gz: f041232dbf6ae3aee4738d9562fab3f820a492dfd8d71fdd257397b278ca03550a27a59a0d0b5c3514254bab49367ba067970f7551ea11c3a7b90b8eda434499
7
+ data.tar.gz: e8c267ccd410ccfc978ef873805696299b1a658d1894dccae2f63d39f8afd12595ced23fbce6c8f4173188cc0b1e6fa44957d37c29c32b8ea18c08fb9a8745e1
data/README.md CHANGED
@@ -1,18 +1,21 @@
1
1
  # Word to Markdown converter
2
2
 
3
- A Ruby gem to liberate content from [the jail that is Word documents](http://ben.balter.com/2012/10/19/we-ve-been-trained-to-make-paper/#jailbreaking-content)
3
+ > [!IMPORTANT]
4
+ > **Looking for the latest and greatest?** Check out [**word-to-markdown-js**](https://github.com/benbalter/word-to-markdown-js), the newer, better successor to this project. It's a modern, actively maintained rewrite and is recommended for new projects. This Ruby gem remains available for existing users.
4
5
 
5
- [![CI](https://github.com/benbalter/word-to-markdown/actions/workflows/ci.yml/badge.svg)](https://github.com/benbalter/word-to-markdown/actions/workflows/ci.yml) [![Gem Version](https://badge.fury.io/rb/word-to-markdown.png)](http://badge.fury.io/rb/word-to-markdown) [![Inline docs](http://inch-ci.org/github/benbalter/word-to-markdown.png)](http://inch-ci.org/github/benbalter/word-to-markdown) [![Build status](https://ci.appveyor.com/api/projects/status/x2gnsfvli3q47a2e/branch/master?svg=true)](https://ci.appveyor.com/project/benbalter/word-to-markdown/branch/master) [![Maintainability](https://api.codeclimate.com/v1/badges/aae0d67ea7db185f1595/maintainability)](https://codeclimate.com/github/benbalter/word-to-markdown/maintainability) [![Test Coverage](https://api.codeclimate.com/v1/badges/aae0d67ea7db185f1595/test_coverage)](https://codeclimate.com/github/benbalter/word-to-markdown/test_coverage)
6
+ A Ruby gem to liberate content from [the jail that is Word documents](https://ben.balter.com/2012/10/19/we-ve-been-trained-to-make-paper/#jailbreaking-content)
7
+
8
+ [![CI](https://github.com/benbalter/word-to-markdown/actions/workflows/ci.yml/badge.svg)](https://github.com/benbalter/word-to-markdown/actions/workflows/ci.yml) [![Gem Version](https://img.shields.io/gem/v/word-to-markdown)](https://rubygems.org/gems/word-to-markdown)
6
9
 
7
10
  ## The problem
8
11
 
9
- > Our default content publishing workflow is terribly broken. [We've all been trained to make paper](http://ben.balter.com/2012/10/19/we-ve-been-trained-to-make-paper/), yet today, content authored once is more commonly consumed in multiple formats, and rarely, if ever, does it embody physical form. Put another way, our go-to content authoring workflow remains relatively unchanged since it was conceived in the early 80s.
12
+ > Our default content publishing workflow is terribly broken. [We've all been trained to make paper](https://ben.balter.com/2012/10/19/we-ve-been-trained-to-make-paper/), yet today, content authored once is more commonly consumed in multiple formats, and rarely, if ever, does it embody physical form. Put another way, our go-to content authoring workflow remains relatively unchanged since it was conceived in the early 80s.
10
13
  >
11
- > I'm asked regularly by government employees — knowledge workers who fire up a desktop word processor as the first step to any project — for an automated pipeline to convert Microsoft Word documents to [Markdown](http://guides.github.com/overviews/mastering-markdown/), the *lingua franca* of the internet, but as my recent foray into building [just such a converter](http://word-to-markdown.herokuapp.com/) proves, it's not that simple.
14
+ > I'm asked regularly by government employees — knowledge workers who fire up a desktop word processor as the first step to any project — for an automated pipeline to convert Microsoft Word documents to [Markdown](https://docs.github.com/en/get-started/writing-on-github/getting-started-with-writing-and-formatting-on-github/basic-writing-and-formatting-syntax), the *lingua franca* of the internet, but as my recent foray into building [just such a converter](https://word2md.com/) proves, it's not that simple.
12
15
  >
13
16
  > Markdown isn't just an alternative format. Markdown forces you to write for the web.
14
17
 
15
- **[Read more](http://ben.balter.com/2014/03/31/word-versus-markdown-more-than-mere-semantics/)**
18
+ **[Read more](https://ben.balter.com/2014/03/31/word-versus-markdown-more-than-mere-semantics/)**
16
19
 
17
20
  ## Just want to convert a Microsoft Word (or Google) document to Markdown?
18
21
 
@@ -20,7 +23,7 @@ You can use this **[hosted service](https://word2md.com/)** (or check out [its s
20
23
 
21
24
  ## Install
22
25
 
23
- You'll need to install [LibreOffice](http://www.libreoffice.org/). Then:
26
+ You'll need to install [LibreOffice](https://www.libreoffice.org/). Then:
24
27
 
25
28
  ```bash
26
29
  gem install word-to-markdown
@@ -67,6 +70,14 @@ $ w2m path/to/document.docx
67
70
 
68
71
  Word-to-markdown requires `soffice` a command line interface to LibreOffice that works on Linux, Mac, and Windows. To install soffice, see [the LibreOffice documentation](https://www.libreoffice.org/get-help/install-howto/).
69
72
 
73
+ Word-to-markdown only accepts Word documents (`.docx` and `.doc`), identified by their contents rather than their file extension. Other files, including those LibreOffice could otherwise open, such as HTML, ODT, or RTF, raise `WordToMarkdown::Document::UnsupportedFormatError`.
74
+
75
+ LibreOffice is killed if a conversion takes longer than 60 seconds, raising `WordToMarkdown::TimeoutError`. To change the limit, set `WordToMarkdown.timeout = 120` or the `WORD_TO_MARKDOWN_TIMEOUT` environment variable.
76
+
77
+ ### Converting untrusted documents
78
+
79
+ Word documents can link to remote resources, such as images, which LibreOffice fetches while converting the document. If you convert documents from untrusted sources (for example, files uploaded to a web service), run the conversion without network access, such as in a container or sandbox with no outbound network, so that a document can't make requests to internal services or other hosts on your behalf.
80
+
70
81
  ## Testing
71
82
 
72
83
  ```
@@ -75,18 +86,12 @@ script/cibuild
75
86
 
76
87
  ## Docker
77
88
 
78
- First, create the `Gemfile.lock` by installing the dependencies:
79
-
80
- ```
81
- bundle install
82
- ```
83
-
84
89
  Everything you need to run the executable locally:
85
90
 
86
91
  ```
87
- docker-compose build
88
- docker-compose run --rm app bundle exec w2m --help
89
- docker-compose run --rm app bundle exec w2m test/fixtures/em.docx
92
+ docker compose build
93
+ docker compose run --rm app bundle exec w2m --help
94
+ docker compose run --rm app bundle exec w2m test/fixtures/em.docx
90
95
  ```
91
96
 
92
97
  ## Hosted service
data/bin/w2m CHANGED
@@ -1,17 +1,35 @@
1
1
  #!/usr/bin/env ruby
2
2
  # frozen_string_literal: true
3
3
 
4
+ require 'optparse'
4
5
  require 'word-to-markdown'
5
6
 
6
- if ARGV.size != 1 || ARGV[0] == '--help'
7
- puts 'Usage: bundle exec w2m path/to/document.docx'
7
+ parser = OptionParser.new do |opts|
8
+ opts.banner = 'Usage: w2m path/to/document.docx'
9
+
10
+ opts.on('-h', '--help', 'Show this help') do
11
+ puts opts
12
+ exit
13
+ end
14
+
15
+ opts.on('-v', '--version', 'Show the WordToMarkdown and LibreOffice versions') do
16
+ puts "WordToMarkdown v#{WordToMarkdown::VERSION}"
17
+ puts "LibreOffice v#{WordToMarkdown.soffice.version}" unless Gem.win_platform?
18
+ exit
19
+ end
20
+ end
21
+
22
+ begin
23
+ parser.parse!
24
+ rescue OptionParser::ParseError => e
25
+ warn e.message
26
+ warn parser
8
27
  exit 1
9
28
  end
10
29
 
11
- if ARGV[0] == '--version'
12
- puts "WordToMarkdown v#{WordToMarkdown::VERSION}"
13
- puts "LibreOffice v#{WordToMarkdown.soffice.version}" unless Gem.win_platform?
14
- else
15
- doc = WordToMarkdown.new ARGV[0]
16
- puts doc.to_s
30
+ if ARGV.size != 1
31
+ warn parser
32
+ exit 1
17
33
  end
34
+
35
+ puts WordToMarkdown.new(ARGV[0])
@@ -7,14 +7,22 @@ class WordToMarkdown
7
7
  # Number of headings to guess, e.g., h6
8
8
  HEADING_DEPTH = 6
9
9
 
10
- # Percentile step for eaceh eheading
10
+ # Percentile step for each heading
11
11
  HEADING_STEP = 100 / HEADING_DEPTH
12
12
 
13
13
  # Minimum heading size
14
14
  MIN_HEADING_SIZE = 20
15
15
 
16
16
  # Unicode bullets to strip when processing
17
- UNICODE_BULLETS = ['○', 'o', '●', "\u2022", '\\p{C}'].freeze
17
+ UNICODE_BULLETS = ['○', '●', "\u2022", '\\p{C}'].freeze
18
+
19
+ # Leading bullets to strip from list items. A plain "o" only counts as a
20
+ # bullet when followed by whitespace, so words like "orange" survive.
21
+ BULLET_REGEX = /\A(?:[#{UNICODE_BULLETS.join}]|o(?=[[:space:]]))+[[:space:]]*/
22
+
23
+ # Leading list numbering to strip, e.g., "1.", "a.", or "iv.", along with
24
+ # any whitespace (including non-breaking spaces) that follows it
25
+ NUMBERING_REGEX = /\A(?:\d+|[a-zA-Z]|[ivxlcdm]+|[IVXLCDM]+)\.(?:[[:space:]]+|\z)/
18
26
 
19
27
  # @param document [WordToMarkdown::Document] The document to convert
20
28
  def initialize(document)
@@ -58,13 +66,13 @@ class WordToMarkdown
58
66
  @document.tree.css('[style]').each do |element|
59
67
  sizes.push element.font_size.round(-1) unless element.font_size.nil?
60
68
  end
61
- sizes.uniq.sort.extend(DescriptiveStatistics)
69
+ sizes.uniq.sort
62
70
  end
63
71
  end
64
72
 
65
73
  # Given a Nokogiri node, guess what heading it represents, if any
66
74
  #
67
- # @param node [Nokigiri::Node] the nokigiri node
75
+ # @param node [Nokogiri::Node] the nokogiri node
68
76
  # @return [String, nil] the heading tag (e.g., H1), or nil
69
77
  def guess_heading(node)
70
78
  return nil if node.font_size.nil?
@@ -82,7 +90,23 @@ class WordToMarkdown
82
90
  #
83
91
  # @return [Integer] the minimum font size
84
92
  def h(num)
85
- font_sizes.percentile(((HEADING_DEPTH - 1) - num) * HEADING_STEP)
93
+ self.class.percentile(font_sizes, ((HEADING_DEPTH - 1) - num) * HEADING_STEP)
94
+ end
95
+
96
+ # Linearly interpolated percentile, matching the algorithm used by the
97
+ # descriptive_statistics gem (v2.5.1) that this replaces
98
+ #
99
+ # @param values [Array<Numeric>] the values
100
+ # @param pct [Numeric] the percentile, from 0 to 100
101
+ #
102
+ # @return [Float, nil] the percentile, or nil if values is empty
103
+ def self.percentile(values, pct)
104
+ sorted = values.map(&:to_f).sort
105
+ rank = pct / 100.0 * (sorted.size - 1)
106
+ lower, upper = sorted[rank.floor, 2]
107
+ return lower if upper.nil? # empty or single-element collection, or the 100th percentile
108
+
109
+ lower + ((upper - lower) * (rank - rank.floor))
86
110
  end
87
111
 
88
112
  # Convert span-based font styles to `strong`s and `em`s
@@ -110,7 +134,7 @@ class WordToMarkdown
110
134
  def remove_unicode_bullets_from_list_items!
111
135
  path = WordToMarkdown.soffice.major_version == '5' ? 'li span span' : 'li span'
112
136
  @document.tree.search(path).each do |span|
113
- span.inner_html = span.inner_html.gsub(/^([#{UNICODE_BULLETS.join}]+)/, '')
137
+ span.inner_html = span.inner_html.sub(BULLET_REGEX, '')
114
138
  end
115
139
  end
116
140
 
@@ -118,21 +142,21 @@ class WordToMarkdown
118
142
  def remove_numbering_from_list_items!
119
143
  path = WordToMarkdown.soffice.major_version == '5' ? 'li span span' : 'li span'
120
144
  @document.tree.search(path).each do |span|
121
- span.inner_html = span.inner_html.gsub(/^[a-zA-Z0-9]+\./m, '')
145
+ span.inner_html = span.inner_html.sub(NUMBERING_REGEX, '')
122
146
  end
123
147
  end
124
148
 
125
- # Remvoe whitespace from list items
149
+ # Remove whitespace from list items
126
150
  def remove_whitespace_from_list_items!
127
- @document.tree.search('li span').each { |span| span.inner_html.strip! }
151
+ @document.tree.search('li span').each { |span| span.inner_html = span.inner_html.strip }
128
152
  end
129
153
 
130
- # Convert table headers to `th`s2
154
+ # Convert table headers to `th`s
131
155
  def semanticize_table_headers!
132
156
  @document.tree.search('table tr:first td').each { |node| node.node_name = 'th' }
133
157
  end
134
158
 
135
- # Try to guess heading where implicit bassed on font size
159
+ # Try to guess heading where implicit based on font size
136
160
  def semanticize_headings!
137
161
  implicit_headings.each do |element|
138
162
  heading = guess_heading element
@@ -6,14 +6,21 @@ class WordToMarkdown
6
6
 
7
7
  class ConversionError < StandardError; end
8
8
 
9
- attr_reader :path, :tmpdir
9
+ class UnsupportedFormatError < StandardError; end
10
+
11
+ attr_reader :path, :tmpdir, :import_filter
10
12
 
11
13
  # @param path [string] Path to the Word document
12
14
  # @param tmpdir [string] Path to a working directory to use
13
15
  def initialize(path, tmpdir = nil)
14
16
  @path = File.expand_path path, Dir.pwd
15
- @tmpdir = tmpdir || Dir.mktmpdir
16
17
  raise NotFoundError, "File #{@path} does not exist" unless File.exist?(@path)
18
+
19
+ @import_filter = InputFormat.filter_for(@path)
20
+ raise UnsupportedFormatError, "File #{@path} is not a Word document (.docx or .doc)" if @import_filter.nil?
21
+
22
+ @own_tmpdir = tmpdir.nil?
23
+ @tmpdir = tmpdir || Dir.mktmpdir
17
24
  end
18
25
 
19
26
  # @return [String] the document's extension
@@ -21,7 +28,7 @@ class WordToMarkdown
21
28
  File.extname path
22
29
  end
23
30
 
24
- # @return [Nokigiri::Document]
31
+ # @return [Nokogiri::Document]
25
32
  def tree
26
33
  @tree ||= begin
27
34
  tree = Nokogiri::HTML(normalized_html)
@@ -32,6 +39,7 @@ class WordToMarkdown
32
39
 
33
40
  # @return [String] the html representation of the document
34
41
  def html
42
+ UrlScrubber.scrub!(tree)
35
43
  tree.to_html.gsub("</li>\n", '</li>')
36
44
  end
37
45
 
@@ -79,8 +87,8 @@ class WordToMarkdown
79
87
  string.gsub!('&nbsp;', ' ') # HTML encoded spaces
80
88
  string.sub!(/\A[[:space:]]+/, '') # document leading whitespace
81
89
  string.sub!(/[[:space:]]+\z/, '') # document trailing whitespace
82
- string.gsub!(/([ ]+)$/, '') # line trailing whitespace
83
- string.gsub!(/\n\n\n\n/, "\n\n") # Quadruple line breaks
90
+ string.gsub!(/( +)$/, '') # line trailing whitespace
91
+ string.gsub!("\n\n\n\n", "\n\n") # Quadruple line breaks
84
92
  string.delete!(' ') # Unicode non-breaking spaces, injected as tabs
85
93
  string.gsub!(/\*\*\ +(?!\*|_)([[:punct:]])/, '**\1') # Remove extra space after bold
86
94
  string
@@ -88,23 +96,31 @@ class WordToMarkdown
88
96
 
89
97
  # @return [String] the path to the intermediary HTML document
90
98
  def dest_path
91
- dest_filename = File.basename(path).gsub(/#{Regexp.escape(extension)}$/, '.html')
92
- File.expand_path(dest_filename, tmpdir)
99
+ File.expand_path("#{File.basename(path, '.*')}.html", tmpdir)
93
100
  end
94
101
 
95
102
  # @return [String] the unnormalized HTML representation
96
103
  def raw_html
97
104
  @raw_html ||= begin
98
- WordToMarkdown.run_command '--headless', '--convert-to', filter, path, '--outdir', tmpdir
105
+ WordToMarkdown.run_command '--headless', "--infilter=#{import_filter}", '--convert-to', filter, path, '--outdir', tmpdir
99
106
  raise ConversionError, "Failed to convert #{path}" unless File.exist?(dest_path)
100
107
 
101
108
  html = File.read dest_path
102
109
  File.delete dest_path
103
110
  html
111
+ ensure
112
+ remove_tmpdir
104
113
  end
105
114
  end
106
115
 
107
- # @return [String] the LibreOffice filter to use for conversion
116
+ # Remove the working directory if we created it and nothing else is in it.
117
+ # Non-empty directories are kept, since LibreOffice may have written
118
+ # images there that the markdown references.
119
+ def remove_tmpdir
120
+ Dir.rmdir(tmpdir) if @own_tmpdir && Dir.empty?(tmpdir)
121
+ end
122
+
123
+ # @return [String] the LibreOffice filter to use for export
108
124
  def filter
109
125
  if WordToMarkdown.soffice.major_version == '5'
110
126
  'html:XHTML Writer File:UTF8'
@@ -0,0 +1,86 @@
1
+ # frozen_string_literal: true
2
+
3
+ class WordToMarkdown
4
+ # Identifies Word documents by their content, rather than their extension,
5
+ # so that other formats LibreOffice would otherwise sniff and import (e.g.,
6
+ # HTML, which can reference remote resources) are never passed to it
7
+ module InputFormat
8
+ # Signature of an OLE compound file, used by .doc files
9
+ OLE_SIGNATURE = "\xD0\xCF\x11\xE0\xA1\xB1\x1A\xE1".b.freeze
10
+
11
+ # Signature of a ZIP local file header, used by .docx files
12
+ ZIP_SIGNATURE = "PK\x03\x04".b.freeze
13
+
14
+ # Signature of a ZIP end of central directory record
15
+ ZIP_EOCD_SIGNATURE = "PK\x05\x06".b.freeze
16
+
17
+ # Size of the end of central directory record, excluding the comment
18
+ ZIP_EOCD_SIZE = 22
19
+
20
+ # Maximum length of a ZIP archive comment
21
+ ZIP_MAX_COMMENT_SIZE = 0xFFFF
22
+
23
+ # The main document part every Word OOXML package contains
24
+ OOXML_DOCUMENT_PART = 'word/document.xml'.b.freeze
25
+
26
+ # LibreOffice import filters, by format
27
+ FILTERS = {
28
+ ooxml: 'MS Word 2007 XML',
29
+ ole: 'MS Word 97'
30
+ }.freeze
31
+
32
+ class << self
33
+ # @param path [String] path to the file
34
+ # @return [String, nil] the LibreOffice import filter for the file, or nil if it isn't a Word document
35
+ def filter_for(path)
36
+ FILTERS[detect(path)]
37
+ end
38
+
39
+ # @param path [String] path to the file
40
+ # @return [Symbol, nil] :ooxml for .docx, :ole for .doc, or nil if it isn't a Word document
41
+ def detect(path)
42
+ File.open(path, 'rb') do |file|
43
+ signature = file.read(OLE_SIGNATURE.bytesize).to_s
44
+ if signature == OLE_SIGNATURE
45
+ :ole
46
+ elsif signature.start_with?(ZIP_SIGNATURE) && ooxml_document?(file)
47
+ :ooxml
48
+ end
49
+ end
50
+ end
51
+
52
+ private
53
+
54
+ # @param file [File] an open ZIP file
55
+ # @return [Boolean] true if the ZIP's central directory lists the Word main document part
56
+ def ooxml_document?(file)
57
+ directory = central_directory(file)
58
+ !directory.nil? && directory.include?(OOXML_DOCUMENT_PART)
59
+ end
60
+
61
+ # Read the ZIP central directory, whose entries include each file name, uncompressed
62
+ #
63
+ # @param file [File] an open ZIP file
64
+ # @return [String, nil] the raw central directory, or nil if it can't be found
65
+ def central_directory(file)
66
+ size, offset = central_directory_location(file)
67
+ return if size.nil? || offset + size > file.size
68
+
69
+ file.seek(offset)
70
+ file.read(size)
71
+ end
72
+
73
+ # @param file [File] an open ZIP file
74
+ # @return [Array<Integer>, nil] the central directory's size and offset, or nil if not found
75
+ def central_directory_location(file)
76
+ tail_size = [file.size, ZIP_EOCD_SIZE + ZIP_MAX_COMMENT_SIZE].min
77
+ file.seek(-tail_size, IO::SEEK_END)
78
+ tail = file.read(tail_size)
79
+ index = tail.rindex(ZIP_EOCD_SIGNATURE)
80
+ return if index.nil? || tail.bytesize - index < ZIP_EOCD_SIZE
81
+
82
+ tail.byteslice(index + 12, 8).unpack('VV')
83
+ end
84
+ end
85
+ end
86
+ end
@@ -0,0 +1,110 @@
1
+ # frozen_string_literal: true
2
+
3
+ class WordToMarkdown
4
+ # Removes links and images with unsafe URL schemes (e.g., javascript: or
5
+ # vbscript:) so that they don't survive into the markdown output
6
+ module UrlScrubber
7
+ # URL schemes permitted in link targets. Links with any other scheme are
8
+ # unwrapped to their text. Relative URLs and fragments are always permitted.
9
+ SAFE_LINK_SCHEMES = %w[http https mailto].freeze
10
+
11
+ # URL schemes permitted in image sources, in addition to data:image/ URIs.
12
+ # Images with any other scheme are removed.
13
+ SAFE_IMAGE_SCHEMES = %w[http https].freeze
14
+
15
+ # Browsers remove ASCII tabs and newlines anywhere in a URL, and C0
16
+ # control characters and spaces at either end, before parsing the scheme
17
+ STRIPPED_CHARS = /[\t\n\r]/
18
+ EDGE_CHARS = /\A[\x00-\x20]+|[\x00-\x20]+\z/
19
+
20
+ # Matches a URL scheme, e.g., "https:"
21
+ SCHEME_REGEX = /\A([a-z][a-z0-9+.-]*):/i
22
+
23
+ # Matches a URL without a recognizable scheme whose first segment contains
24
+ # characters a Markdown renderer or browser may decode into a scheme
25
+ # separator, e.g., "javascript&colon;", "javascript&#58;", or "javascript\:"
26
+ AMBIGUOUS_REGEX = %r{\A[^/?#]*[:&\\]}
27
+
28
+ # Characters that could end or alter a Markdown link destination, which
29
+ # are percent-encoded in the URLs that are kept
30
+ DESTINATION_UNSAFE_CHARS = /[\x00-\x20\x7F()<>\[\]\\`"]/
31
+
32
+ # Matches a data URI for an image
33
+ DATA_IMAGE_REGEX = %r{\Adata:image/}i
34
+
35
+ class << self
36
+ # Unwrap links and remove images whose URL scheme isn't permitted
37
+ #
38
+ # @param tree [Nokogiri::HTML::Document] the document to scrub, in place
39
+ # @return [Nokogiri::HTML::Document] the scrubbed document
40
+ def scrub!(tree)
41
+ tree.css('a[href]').each do |node|
42
+ safe_link?(node['href']) ? escape_attributes!(node, 'href') : node.replace(node.children)
43
+ end
44
+
45
+ tree.css('img[src]').each do |node|
46
+ safe_image?(node['src']) ? escape_attributes!(node, 'src') : node.remove
47
+ end
48
+
49
+ tree
50
+ end
51
+
52
+ # Percent-encode characters in a URL that could end or alter a Markdown link destination
53
+ #
54
+ # @param url [String] the URL
55
+ # @return [String] the escaped URL
56
+ def escape_destination(url)
57
+ url.gsub(DESTINATION_UNSAFE_CHARS) { |char| format('%%%02X', char.ord) }
58
+ end
59
+
60
+ # @param url [String] a link target
61
+ # @return [Boolean] true if the URL is relative, a fragment, or has a permitted scheme
62
+ def safe_link?(url)
63
+ scheme = scheme(url)
64
+ scheme.nil? ? relative?(url) : SAFE_LINK_SCHEMES.include?(scheme)
65
+ end
66
+
67
+ # @param url [String] an image source
68
+ # @return [Boolean] true if the URL is relative, a data:image/ URI, or has a permitted scheme
69
+ def safe_image?(url)
70
+ scheme = scheme(url)
71
+ return relative?(url) if scheme.nil?
72
+
73
+ SAFE_IMAGE_SCHEMES.include?(scheme) || normalize(url).match?(DATA_IMAGE_REGEX)
74
+ end
75
+
76
+ # @param url [String] the URL
77
+ # @return [String, nil] the URL's lowercased scheme, or nil if it has none
78
+ def scheme(url)
79
+ match = normalize(url).match(SCHEME_REGEX)
80
+ match && match[1].downcase
81
+ end
82
+
83
+ private
84
+
85
+ # Escape a kept link or image's URL and title so they can't break out
86
+ # of the Markdown link that ReverseMarkdown writes for them
87
+ #
88
+ # @param node [Nokogiri::XML::Element] the link or image
89
+ # @param attribute [String] the name of the URL attribute
90
+ def escape_attributes!(node, attribute)
91
+ node[attribute] = escape_destination(node[attribute])
92
+ node['title'] = node['title'].tr('"', "'") if node['title']
93
+ end
94
+
95
+ # @param url [String] a URL without a scheme
96
+ # @return [Boolean] true if the URL is a relative path or fragment that can't be decoded into one with a scheme
97
+ def relative?(url)
98
+ !normalize(url).match?(AMBIGUOUS_REGEX)
99
+ end
100
+
101
+ # Normalize a URL the way a browser does before parsing its scheme
102
+ #
103
+ # @param url [String] the URL
104
+ # @return [String] the normalized URL
105
+ def normalize(url)
106
+ url.to_s.gsub(STRIPPED_CHARS, '').gsub(EDGE_CHARS, '')
107
+ end
108
+ end
109
+ end
110
+ end
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  class WordToMarkdown
4
- VERSION = '1.1.9'
4
+ VERSION = '1.2.0'
5
5
  end
@@ -1,6 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
- require 'descriptive_statistics/safe'
4
3
  require 'reverse_markdown'
5
4
  require 'nokogiri-styles'
6
5
  require 'premailer'
@@ -12,12 +11,16 @@ require 'cliver'
12
11
  require 'open3'
13
12
 
14
13
  require_relative 'word-to-markdown/version'
14
+ require_relative 'word-to-markdown/input_format'
15
+ require_relative 'word-to-markdown/url_scrubber'
15
16
  require_relative 'word-to-markdown/document'
16
17
  require_relative 'word-to-markdown/converter'
17
18
  require_relative 'nokogiri/xml/element'
18
19
  require_relative 'cliver/dependency_ext'
19
20
 
20
21
  class WordToMarkdown
22
+ class TimeoutError < StandardError; end
23
+
21
24
  attr_reader :document, :converter
22
25
 
23
26
  # Options to be passed to Reverse Markdown
@@ -26,6 +29,11 @@ class WordToMarkdown
26
29
  github_flavored: true
27
30
  }.freeze
28
31
 
32
+ # Default number of seconds to wait for LibreOffice to convert a document
33
+ # before giving up. Can be overridden with the WORD_TO_MARKDOWN_TIMEOUT
34
+ # environment variable, or by setting WordToMarkdown.timeout
35
+ DEFAULT_TIMEOUT = 60
36
+
29
37
  # Minimum version of LibreOffice Required
30
38
  SOFFICE_VERSION_REQUIREMENT = '> 4.0'
31
39
 
@@ -34,6 +42,7 @@ class WordToMarkdown
34
42
  '*', # Sub'd for ENV["PATH"]
35
43
  '~/Applications/LibreOffice.app/Contents/MacOS',
36
44
  '/Applications/LibreOffice.app/Contents/MacOS',
45
+ '/Program Files/LibreOffice/program',
37
46
  '/Program Files/LibreOffice 5/program',
38
47
  '/Program Files (x86)/LibreOffice 4/program'
39
48
  ].freeze
@@ -56,20 +65,46 @@ class WordToMarkdown
56
65
  end
57
66
 
58
67
  class << self
68
+ attr_writer :timeout
69
+
70
+ # @return [Numeric] seconds to wait for LibreOffice before giving up
71
+ def timeout
72
+ @timeout ||= Float(ENV.fetch('WORD_TO_MARKDOWN_TIMEOUT', DEFAULT_TIMEOUT))
73
+ end
74
+
59
75
  # Run an soffice command
60
76
  #
61
- # @param args [string] one or more arguments to pass to the sofice command
77
+ # @param args [string] one or more arguments to pass to the soffice command
62
78
  # @return [string] the command output
63
79
  def run_command(*args)
64
80
  raise 'LibreOffice already running' if soffice.open?
65
81
 
66
- output, status = Open3.capture2e(soffice.path, *args)
82
+ output, status = capture_with_timeout(soffice.path, *args, timeout: timeout)
67
83
  logger.debug output
68
84
  raise "Command `#{soffice.path} #{args.join(' ')}` failed: #{output}" if status.exitstatus != 0
69
85
 
70
86
  output
71
87
  end
72
88
 
89
+ # Run a command, capturing its combined output, and kill it (along with
90
+ # any child processes) if it runs longer than the timeout
91
+ #
92
+ # @param command [Array<String>] the command and its arguments
93
+ # @param timeout [Numeric] seconds to wait before killing the command
94
+ # @return [Array(String, Process::Status)] the command output and status
95
+ def capture_with_timeout(*command, timeout:)
96
+ Open3.popen2e(*command, process_group_option) do |stdin, output, waiter|
97
+ stdin.close
98
+ reader = read_in_background(output)
99
+ unless waiter.join(timeout)
100
+ kill_process_group(waiter.pid)
101
+ raise TimeoutError, "Command `#{command.join(' ')}` timed out after #{timeout} seconds"
102
+ end
103
+
104
+ [reader.value, waiter.value]
105
+ end
106
+ end
107
+
73
108
  # Returns a Cliver::Dependency object representing our soffice dependency
74
109
  #
75
110
  # Attempts to resolve by looking at PATH followed by paths in the PATHS constant
@@ -94,13 +129,38 @@ class WordToMarkdown
94
129
 
95
130
  private
96
131
 
132
+ # Start commands in their own process group, so that on timeout any
133
+ # processes they spawn (e.g., soffice.bin) are killed too
134
+ def process_group_option
135
+ Gem.win_platform? ? { new_pgroup: true } : { pgroup: true }
136
+ end
137
+
138
+ # Read a stream in a separate thread, so the command can't block on a full pipe
139
+ #
140
+ # @param io [IO] the stream to read
141
+ # @return [Thread] a thread whose value is the stream's contents
142
+ def read_in_background(io)
143
+ Thread.new do
144
+ io.read
145
+ rescue IOError
146
+ '' # The stream was closed after the command timed out
147
+ end
148
+ end
149
+
150
+ # @param pid [Integer] the pid of a process group leader
151
+ def kill_process_group(pid)
152
+ Process.kill('KILL', Gem.win_platform? ? pid : -pid)
153
+ rescue Errno::ESRCH
154
+ nil
155
+ end
156
+
97
157
  # Workaround for two upstream bugs:
98
- # 1. `soffice.exe --version` on windows opens a popup and retuns a null string when manually closed
158
+ # 1. `soffice.exe --version` on windows opens a popup and returns a null string when manually closed
99
159
  # 2. Even if the second argument to Cliver is nil, Cliver thinks there's a requirement
100
160
  # and will shell out to `soffice.exe --version`
101
161
  # In order to support Windows, don't pass *any* version requirement to Cliver
102
162
  def soffice_dependency_args
103
- args = [path: PATHS.join(File::PATH_SEPARATOR)]
163
+ args = [{ path: PATHS.join(File::PATH_SEPARATOR) }]
104
164
  if Gem.win_platform?
105
165
  args
106
166
  else
metadata CHANGED
@@ -1,14 +1,13 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: word-to-markdown
3
3
  version: !ruby/object:Gem::Version
4
- version: 1.1.9
4
+ version: 1.2.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Ben Balter
8
- autorequire:
9
8
  bindir: bin
10
9
  cert_chain: []
11
- date: 2025-01-08 00:00:00.000000000 Z
10
+ date: 1980-01-02 00:00:00.000000000 Z
12
11
  dependencies:
13
12
  - !ruby/object:Gem::Dependency
14
13
  name: cliver
@@ -25,19 +24,19 @@ dependencies:
25
24
  - !ruby/object:Gem::Version
26
25
  version: '0.3'
27
26
  - !ruby/object:Gem::Dependency
28
- name: descriptive_statistics
27
+ name: logger
29
28
  requirement: !ruby/object:Gem::Requirement
30
29
  requirements:
31
30
  - - "~>"
32
31
  - !ruby/object:Gem::Version
33
- version: '2.5'
32
+ version: '1.4'
34
33
  type: :runtime
35
34
  prerelease: false
36
35
  version_requirements: !ruby/object:Gem::Requirement
37
36
  requirements:
38
37
  - - "~>"
39
38
  - !ruby/object:Gem::Version
40
- version: '2.5'
39
+ version: '1.4'
41
40
  - !ruby/object:Gem::Dependency
42
41
  name: nokogiri-styles
43
42
  requirement: !ruby/object:Gem::Requirement
@@ -92,42 +91,54 @@ dependencies:
92
91
  requirements:
93
92
  - - "~>"
94
93
  - !ruby/object:Gem::Version
95
- version: '1.0'
94
+ version: '1.3'
96
95
  type: :runtime
97
96
  prerelease: false
98
97
  version_requirements: !ruby/object:Gem::Requirement
99
98
  requirements:
100
99
  - - "~>"
101
100
  - !ruby/object:Gem::Version
102
- version: '1.0'
101
+ version: '1.3'
103
102
  - !ruby/object:Gem::Dependency
104
103
  name: minitest
105
104
  requirement: !ruby/object:Gem::Requirement
106
105
  requirements:
107
- - - "~>"
106
+ - - ">="
107
+ - !ruby/object:Gem::Version
108
+ version: '5'
109
+ - - "<"
108
110
  - !ruby/object:Gem::Version
109
- version: '5.0'
111
+ version: '7'
110
112
  type: :development
111
113
  prerelease: false
112
114
  version_requirements: !ruby/object:Gem::Requirement
113
115
  requirements:
114
- - - "~>"
116
+ - - ">="
117
+ - !ruby/object:Gem::Version
118
+ version: '5'
119
+ - - "<"
115
120
  - !ruby/object:Gem::Version
116
- version: '5.0'
121
+ version: '7'
117
122
  - !ruby/object:Gem::Dependency
118
123
  name: mocha
119
124
  requirement: !ruby/object:Gem::Requirement
120
125
  requirements:
121
- - - "~>"
126
+ - - ">="
122
127
  - !ruby/object:Gem::Version
123
- version: '1.1'
128
+ version: '2'
129
+ - - "<"
130
+ - !ruby/object:Gem::Version
131
+ version: '4'
124
132
  type: :development
125
133
  prerelease: false
126
134
  version_requirements: !ruby/object:Gem::Requirement
127
135
  requirements:
128
- - - "~>"
136
+ - - ">="
137
+ - !ruby/object:Gem::Version
138
+ version: '2'
139
+ - - "<"
129
140
  - !ruby/object:Gem::Version
130
- version: '1.1'
141
+ version: '4'
131
142
  - !ruby/object:Gem::Dependency
132
143
  name: pry
133
144
  requirement: !ruby/object:Gem::Requirement
@@ -227,13 +238,17 @@ files:
227
238
  - lib/word-to-markdown.rb
228
239
  - lib/word-to-markdown/converter.rb
229
240
  - lib/word-to-markdown/document.rb
241
+ - lib/word-to-markdown/input_format.rb
242
+ - lib/word-to-markdown/url_scrubber.rb
230
243
  - lib/word-to-markdown/version.rb
231
244
  homepage: https://github.com/benbalter/word-to-markdown
232
245
  licenses:
233
246
  - MIT
234
247
  metadata:
235
248
  rubygems_mfa_required: 'true'
236
- post_install_message:
249
+ source_code_uri: https://github.com/benbalter/word-to-markdown
250
+ bug_tracker_uri: https://github.com/benbalter/word-to-markdown/issues
251
+ changelog_uri: https://github.com/benbalter/word-to-markdown/releases
237
252
  rdoc_options: []
238
253
  require_paths:
239
254
  - lib
@@ -241,15 +256,14 @@ required_ruby_version: !ruby/object:Gem::Requirement
241
256
  requirements:
242
257
  - - ">="
243
258
  - !ruby/object:Gem::Version
244
- version: '0'
259
+ version: '3.2'
245
260
  required_rubygems_version: !ruby/object:Gem::Requirement
246
261
  requirements:
247
262
  - - ">="
248
263
  - !ruby/object:Gem::Version
249
264
  version: '0'
250
265
  requirements: []
251
- rubygems_version: 3.5.16
252
- signing_key:
266
+ rubygems_version: 3.6.9
253
267
  specification_version: 4
254
268
  summary: Ruby Gem to convert Word documents to markdown
255
269
  test_files: []