lesath 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml ADDED
@@ -0,0 +1,7 @@
1
+ ---
2
+ SHA256:
3
+ metadata.gz: '08725b38f6494e679c2239f228fdee59dc79febcb8f9d14daadb44fdcec5848d'
4
+ data.tar.gz: 44d47d31ad5fbdf6a61bfeb5f6e1202f7aab7b508670468360dbe6e5ccccfabd
5
+ SHA512:
6
+ metadata.gz: '09e88095d742b34f4e22bfb6a58c3e9ec86234ca40c5ed0e70d41787f5a59b843b3ccf2a34bdc1dc1e16334a03cf82e622646b0cca90cd152c63dc803b429735'
7
+ data.tar.gz: c47989d1574c9decd0f806827ff002f741553fab9078724fa92ab601d08a6e716f0997d2eaa7b0ae14a713af44ac0fe11d42ea2ae34789ff0dbf610db35a0685
data/CHANGELOG.md ADDED
@@ -0,0 +1,5 @@
1
+ # Changelog
2
+
3
+ ## [0.1.0] - 2026-09-24
4
+
5
+ - Initial release.
data/LICENSE.txt ADDED
@@ -0,0 +1,21 @@
1
+ The MIT License (MIT)
2
+
3
+ Copyright (c) 2026 Yudai Takada
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in
13
+ all copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN
21
+ THE SOFTWARE.
data/README.md ADDED
@@ -0,0 +1,53 @@
1
+ # Lesath
2
+
3
+ Lesath (υ Scorpii) is a small library for strict spreadsheet interchange.
4
+ The star's name comes from Arabic *lasʿa*, “sting.” It reads and writes a
5
+ documented subset of `.xlsx` and `.ods` with no UI or calculation engine.
6
+
7
+ ```ruby
8
+ require "lesath"
9
+
10
+ book = Lesath::Workbook.new
11
+ book.add_sheet("Sales").add_sheet("Summary")
12
+ book.set("Sales", 1, 1, "Item")
13
+ book.set("Sales", 2, 1, "Tea")
14
+ book.set("Sales", 2, 2, 12.5)
15
+ Lesath.write(book, "sales.xlsx")
16
+ copy = Lesath.read("sales.xlsx")
17
+ copy.cell("Sales", 2, 2).value # => 12.5
18
+ ```
19
+
20
+ Coordinates are one-based. `Lesath.read` / `Lesath.write` infer the format from
21
+ the filename or accept `format: :xlsx` / `format: :ods`. Writing refuses to
22
+ replace an existing target. The API stores formula text and its cached scalar;
23
+ it does not calculate or translate formulas. XLSX formulas begin with `=`, ODS
24
+ formulas with `of:=`. A workbook containing formulas cannot be written to the
25
+ other format. ODS formulas need a cached value. The caller must recalculate
26
+ before exporting changed formula inputs.
27
+
28
+ Supported cells are strings, finite integers/floats, booleans and empty cells,
29
+ with up to 100,000 populated cells, 200 sheets, 1,048,576 rows, 16,384
30
+ columns, 15-digit integers, 32,767 UTF-16 units per string and 31 characters per sheet name. XLSX inline strings and simple scalar cached
31
+ formulas are supported. ODS unstyled scalar cells and simple repeated blank
32
+ rows/cells are supported. Multiple sheets retain their order and names.
33
+ ODS strings needing `<text:s>`, tabs or line breaks are rejected rather than
34
+ exported with altered display text. Outputs are checked against the same ZIP
35
+ limits as imports before being moved to the target path.
36
+
37
+ Imported files containing cell styles, rich text, dates/times, merged cells,
38
+ comments, charts, images, macros, validations, hyperlinks, named ranges,
39
+ external links, other unsupported package parts or XML elements raise
40
+ `Lesath::UnsupportedFeature`. ZIP and XML limits raise
41
+ `Lesath::InvalidPackage`. This is deliberately not a general Excel/LibreOffice
42
+ round-trip editor. See [ADR 001](docs/adr/001-lossless-subset.md) for the exact
43
+ boundary and source standards.
44
+
45
+ `Rukbat::Workbook` is not directly accepted: it has richer formatting and
46
+ formula semantics, so an adapter must explicitly choose a loss policy. CSV
47
+ remains the appropriate broad interchange path until that adapter exists.
48
+
49
+ ```sh
50
+ bundle install
51
+ bundle exec rake spec
52
+ gem build --strict lesath.gemspec
53
+ ```
@@ -0,0 +1,60 @@
1
+ # ADR 001: Reject unsupported spreadsheet features on import
2
+
3
+ - Status: Accepted
4
+ - Date: 2026-09-24
5
+ - Author: Yudai Takada
6
+
7
+ ## Context
8
+
9
+ Q9-a asks for xlsx/ODS interchange. Rukbat's workbook holds sparse cell input,
10
+ formulas, formats, comments, ranges and print metadata, while its CSV route
11
+ exports a single sheet with explicit input/calculated-value choice. XLSX and
12
+ ODS are package formats with many more features. Reading only visible values
13
+ then saving would quietly destroy unsupported content.
14
+
15
+ ## Decision
16
+
17
+ Lesath has its own small workbook model. Import validates every ZIP part and
18
+ supported XML element/attribute. Unknown or unsupported content raises an
19
+ error before returning a workbook. Export creates a new package and never
20
+ overwrites an existing path. The initial subset contains ordered sheets,
21
+ one-based sparse cells, strings, finite numbers, booleans and same-format
22
+ formula text with a cached scalar. It does not evaluate formulas. Cross-format
23
+ formula export is rejected because SpreadsheetML and OpenFormula use different
24
+ syntax and function semantics. Input values are stored separately from
25
+ formula text. Formula caches may be stale; callers must recalculate them.
26
+
27
+ Cell formatting, date/time serial interpretation, rich text, shared formulas,
28
+ styles, merged cells, validation, protection, named ranges, drawing/chart
29
+ parts, embedded objects, signatures, encryption, macros and external
30
+ relationships are outside the subset and rejected. XLSX's default numeric
31
+ cell has no date inference; non-default style references are rejected. ODS
32
+ supports only `string`, `float`, `boolean` value types. These are explicit
33
+ compatibility limits, not an implicit lossy conversion option.
34
+
35
+ ZIP input is restricted to 256 entries, 20 MiB per part, 50 MiB total and a
36
+ 1,000:1 advertised compression ratio; paths, duplicate names and encrypted
37
+ entries are rejected. XML DTD/entity declarations are rejected before parsing.
38
+ Package part allowlists and XML element/attribute allowlists prevent silent
39
+ loss of unknown features. This limits supported real-world files, including
40
+ many office-generated workbooks with default style or metadata parts. Expand
41
+ the subset only with a fixture proving that each added feature survives a
42
+ read/write round trip or produces a clear error.
43
+
44
+ ODS whitespace that requires `text:s`, tabs or line-break markup is rejected
45
+ until those elements are supported. XLSX inline strings use `xml:space` when
46
+ needed. Public cell text is immutable, and generated packages must pass the
47
+ same size and compression checks as imported packages before publication.
48
+
49
+ The model is intentionally independent of Rukbat. A future adapter should
50
+ map only representable values and explicitly reject Rukbat formatting,
51
+ comments, named ranges and metadata until those are implemented here. It
52
+ must not silently discard them.
53
+
54
+ ## Standards
55
+
56
+ - [ECMA-376, Office Open XML](https://ecma-international.org/publications-and-standards/standards/ecma-376/) — SpreadsheetML and Open Packaging Conventions.
57
+ - [Microsoft Open XML SpreadsheetML structure](https://learn.microsoft.com/en-us/office/open-xml/spreadsheet/structure-of-a-spreadsheetml-document) — workbook, worksheet, relationship and cell packaging examples.
58
+ - [OASIS OpenDocument 1.3 Part 2: Packages](https://docs.oasis-open.org/office/OpenDocument/v1.3/OpenDocument-v1.3-part2-packages.html) — manifest and uncompressed first `mimetype` entry.
59
+ - [OASIS OpenDocument 1.3 Part 3: Schema](https://docs.oasis-open.org/office/OpenDocument/v1.3/OpenDocument-v1.3-part3-schema.html) — spreadsheet tables, value types and repeated cells.
60
+ - [OASIS OpenDocument 1.3 Part 4: OpenFormula](https://docs.oasis-open.org/office/OpenDocument/v1.3/OpenDocument-v1.3-part4-formula.html) — ODS formula syntax.
data/lib/lesath/ods.rb ADDED
@@ -0,0 +1,169 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Lesath
4
+ module ODS
5
+ OFFICE = "urn:oasis:names:tc:opendocument:xmlns:office:1.0"
6
+ TABLE = "urn:oasis:names:tc:opendocument:xmlns:table:1.0"
7
+ TEXT = "urn:oasis:names:tc:opendocument:xmlns:text:1.0"
8
+ OF = "urn:oasis:names:tc:opendocument:xmlns:of:1.2"
9
+ MANIFEST = "urn:oasis:names:tc:opendocument:xmlns:manifest:1.0"
10
+ MIME = "application/vnd.oasis.opendocument.spreadsheet"
11
+ module_function
12
+
13
+ def read(path)
14
+ parts = Package.read(path)
15
+ raise UnsupportedFeature, "unsupported ODS package parts" unless parts.keys.sort == ["META-INF/manifest.xml", "content.xml", "mimetype"].sort
16
+ raise InvalidPackage, "not an ODS spreadsheet" unless parts["mimetype"] == MIME
17
+ check_manifest(parts.fetch("META-INF/manifest.xml"))
18
+ root = Package.xml(parts.fetch("content.xml"))
19
+ Package.check(root, OFFICE, "document-content", attributes: [[OFFICE, "version"]], children: [[OFFICE, "body"]])
20
+ raise UnsupportedFeature, "unsupported ODS version" unless root.attributes["office:version"] == "1.3"
21
+ body = root.elements.to_a
22
+ raise InvalidPackage, "missing ODS body" unless body.length == 1
23
+ Package.check(body.first, OFFICE, "body", children: [[OFFICE, "spreadsheet"]])
24
+ spreadsheet = body.first.elements.to_a
25
+ raise InvalidPackage, "missing ODS spreadsheet" unless spreadsheet.length == 1
26
+ Package.check(spreadsheet.first, OFFICE, "spreadsheet", children: [[TABLE, "table"]])
27
+ book = Workbook.new(source_format: :ods)
28
+ spreadsheet.first.elements.each { |table| parse_table(table, book) }
29
+ raise InvalidPackage, "ODS has no sheets" if book.sheet_names.empty?
30
+ book
31
+ end
32
+
33
+ def write(book, path)
34
+ raise Error, "workbook has no sheets" if book.sheet_names.empty?
35
+ raise UnsupportedFeature, "cannot translate xlsx formulas to ODS" if book.source_format == :xlsx && book.sheet_names.any? { |name| book.each_cell(name).any? { |_r, _c, cell| cell.formula } }
36
+ parts = {
37
+ "mimetype" => MIME,
38
+ "META-INF/manifest.xml" => manifest_xml,
39
+ "content.xml" => content_xml(book)
40
+ }
41
+ Package.write(parts, path, first_stored: "mimetype")
42
+ end
43
+
44
+ def check_manifest(bytes)
45
+ root = Package.xml(bytes)
46
+ Package.check(root, MANIFEST, "manifest", attributes: [[MANIFEST, "version"]], children: [[MANIFEST, "file-entry"]])
47
+ entries = root.elements.to_a.to_h do |entry|
48
+ Package.check(entry, MANIFEST, "file-entry", attributes: [[MANIFEST, "full-path"], [MANIFEST, "media-type"]])
49
+ [entry.attributes["manifest:full-path"], entry.attributes["manifest:media-type"]]
50
+ end
51
+ raise InvalidPackage, "duplicate ODS manifest entry" if entries.length != root.elements.to_a.length
52
+ raise UnsupportedFeature, "unsupported ODS manifest" unless entries == {"/" => MIME, "content.xml" => "text/xml"}
53
+ end
54
+
55
+ def parse_table(element, book)
56
+ Package.check(element, TABLE, "table", attributes: [[TABLE, "name"]], children: [[TABLE, "table-row"]])
57
+ name = element.attributes["table:name"]
58
+ raise InvalidPackage, "missing sheet name" unless name
59
+ book.add_sheet(name)
60
+ row_number = 1
61
+ element.elements.each do |row|
62
+ Package.check(row, TABLE, "table-row", attributes: [[TABLE, "number-rows-repeated"]], children: [[TABLE, "table-cell"]])
63
+ repeat = count(row.attributes["table:number-rows-repeated"])
64
+ raise InvalidPackage, "row out of range" if row_number + repeat - 1 > Workbook::MAX_ROWS
65
+ column_number = 1
66
+ row.elements.each do |cell|
67
+ Package.check(cell, TABLE, "table-cell", attributes: [[TABLE, "number-columns-repeated"], [TABLE, "formula"], [OFFICE, "value-type"], [OFFICE, "value"], [OFFICE, "boolean-value"], [OFFICE, "string-value"]], children: [[TEXT, "p"]])
68
+ columns = count(cell.attributes["table:number-columns-repeated"])
69
+ raise InvalidPackage, "column out of range" if column_number + columns - 1 > Workbook::MAX_COLUMNS
70
+ formula = cell.attributes["table:formula"]
71
+ raise UnsupportedFeature, "unsupported ODS formula syntax" if formula && !formula.start_with?("of:=")
72
+ raise InvalidPackage, "OpenFormula namespace is not bound" if formula && cell.namespace("of") != OF
73
+ value = parse_value(cell)
74
+ raise UnsupportedFeature, "ODS formula requires a cached value" if formula && value.nil?
75
+ populated = !value.nil? || formula
76
+ raise UnsupportedFeature, "repeated populated rows are not supported" if repeat > 1 && populated
77
+ raise UnsupportedFeature, "repeated populated cells are not supported" if columns > 1 && populated
78
+ book.set(name, row_number, column_number, value, formula:) if populated
79
+ column_number += columns
80
+ end
81
+ row_number += repeat
82
+ end
83
+ end
84
+
85
+ def parse_value(cell)
86
+ paragraphs = cell.elements.to_a
87
+ paragraphs.each { |paragraph| Package.check(paragraph, TEXT, "p", text: true) }
88
+ raise UnsupportedFeature, "multiline or rich text is not supported" if paragraphs.length > 1 || paragraphs.any? { |p| !p.elements.to_a.empty? }
89
+ type = cell.attributes["office:value-type"]
90
+ values = %w[value boolean-value string-value].select { |key| cell.attributes["office:#{key}"] }
91
+ allowed = {nil => [], "string" => ["string-value"], "float" => ["value"], "boolean" => ["boolean-value"]}[type]
92
+ raise UnsupportedFeature, "unsupported ODS value type: #{type}" unless allowed
93
+ raise UnsupportedFeature, "conflicting ODS value attributes" unless (values - allowed).empty?
94
+ case type
95
+ when nil
96
+ raise InvalidPackage, "text without a value type" unless paragraphs.empty?
97
+ nil
98
+ when "string"
99
+ text = cell.attributes["office:string-value"] || paragraphs.first&.text.to_s
100
+ raise UnsupportedFeature, "ODS string display text is required" if paragraphs.empty?
101
+ safe_text!(text)
102
+ raise InvalidPackage, "string cache differs from display text" if paragraphs.any? && paragraphs.first.text.to_s != text
103
+ text
104
+ when "float"
105
+ raise UnsupportedFeature, "formatted numeric display text" unless paragraphs.empty?
106
+ XLSX.number(cell.attributes["office:value"])
107
+ when "boolean"
108
+ raise UnsupportedFeature, "formatted boolean display text" unless paragraphs.empty?
109
+ {"true" => true, "false" => false}.fetch(cell.attributes["office:boolean-value"]) { raise InvalidPackage, "invalid ODS boolean" }
110
+ end
111
+ end
112
+
113
+ def count(value)
114
+ return 1 unless value
115
+ raise InvalidPackage, "invalid repeat count" unless value.length <= 7 && value.match?(/\A[1-9]\d*\z/)
116
+ Integer(value, 10)
117
+ end
118
+
119
+ def manifest_xml
120
+ root = REXML::Element.new("manifest:manifest")
121
+ root.add_attributes("xmlns:manifest" => MANIFEST, "manifest:version" => "1.3")
122
+ root.add_element("manifest:file-entry").add_attributes("manifest:full-path" => "/", "manifest:media-type" => MIME)
123
+ root.add_element("manifest:file-entry").add_attributes("manifest:full-path" => "content.xml", "manifest:media-type" => "text/xml")
124
+ Package.render(root)
125
+ end
126
+
127
+ def content_xml(book)
128
+ root = REXML::Element.new("office:document-content")
129
+ root.add_attributes("xmlns:office" => OFFICE, "xmlns:table" => TABLE, "xmlns:text" => TEXT, "xmlns:of" => OF, "office:version" => "1.3")
130
+ spreadsheet = root.add_element("office:body").add_element("office:spreadsheet")
131
+ book.sheet_names.each do |name|
132
+ table = spreadsheet.add_element("table:table", {"table:name" => name})
133
+ previous_row = 0
134
+ book.each_cell(name).group_by(&:first).each do |row, cells|
135
+ gap = row - previous_row - 1
136
+ table.add_element("table:table-row", {"table:number-rows-repeated" => gap.to_s}) if gap.positive?
137
+ row_node = table.add_element("table:table-row")
138
+ previous_column = 0
139
+ cells.each do |_row, column, cell|
140
+ raise UnsupportedFeature, "ODS formula must start with of:=" if cell.formula && !cell.formula.start_with?("of:=")
141
+ gap = column - previous_column - 1
142
+ row_node.add_element("table:table-cell", {"table:number-columns-repeated" => gap.to_s}) if gap.positive?
143
+ node = row_node.add_element("table:table-cell")
144
+ node.add_attribute("table:formula", cell.formula) if cell.formula
145
+ case cell.value
146
+ when String
147
+ safe_text!(cell.value)
148
+ node.add_attributes("office:value-type" => "string", "office:string-value" => cell.value)
149
+ node.add_element("text:p").text = cell.value
150
+ when TrueClass, FalseClass then node.add_attributes("office:value-type" => "boolean", "office:boolean-value" => cell.value.to_s)
151
+ when Numeric then node.add_attributes("office:value-type" => "float", "office:value" => cell.value.to_s)
152
+ when NilClass
153
+ raise UnsupportedFeature, "ODS formula requires a cached value" if cell.formula
154
+ end
155
+ previous_column = column
156
+ end
157
+ previous_row = row
158
+ end
159
+ end
160
+ Package.render(root)
161
+ end
162
+
163
+ def safe_text!(text)
164
+ if text.match?(/\A | \z| {2}|[\t\r\n]/)
165
+ raise UnsupportedFeature, "ODS string whitespace needs text:s, tab, or line-break support"
166
+ end
167
+ end
168
+ end
169
+ end
@@ -0,0 +1,120 @@
1
+ # frozen_string_literal: true
2
+
3
+ require "tempfile"
4
+
5
+ module Lesath
6
+ module Package
7
+ MAX_PART = 20 * 1024 * 1024
8
+ MAX_TOTAL = 50 * 1024 * 1024
9
+ MAX_ENTRIES = 256
10
+
11
+ def self.read(path)
12
+ Zip::File.open(path) do |zip|
13
+ raise InvalidPackage, "too many ZIP entries" if zip.entries.length > MAX_ENTRIES
14
+
15
+ parts = {}
16
+ total = 0
17
+ zip.entries.each do |entry|
18
+ name = entry.name
19
+ raise InvalidPackage, "unsafe ZIP path" if name.start_with?("/") || name.include?("\\") || name.split("/").include?("..") || name.include?("\0")
20
+ raise InvalidPackage, "duplicate ZIP part: #{name}" if parts.key?(name)
21
+ raise InvalidPackage, "encrypted ZIP entry" if (entry.gp_flags & 1) != 0
22
+ raise InvalidPackage, "ZIP part too large" if entry.size > MAX_PART || total + entry.size > MAX_TOTAL
23
+ raise InvalidPackage, "suspicious ZIP compression" if entry.compressed_size.positive? && entry.size > entry.compressed_size * 1_000
24
+
25
+ bytes = entry.get_input_stream { |input| input.read(MAX_PART + 1) }
26
+ raise InvalidPackage, "ZIP part too large" if bytes.bytesize > MAX_PART
27
+ total += bytes.bytesize
28
+ raise InvalidPackage, "ZIP package too large" if total > MAX_TOTAL
29
+ parts[name] = bytes
30
+ end
31
+ parts
32
+ end
33
+ rescue Zip::Error, SystemCallError => error
34
+ raise InvalidPackage, "cannot read package: #{error.message}"
35
+ end
36
+
37
+ def self.write(parts, path, first_stored: nil)
38
+ raise Error, "target already exists: #{path}" if File.exist?(path)
39
+
40
+ target = File.expand_path(path)
41
+ Tempfile.create([".lesath-", ".tmp"], File.dirname(target)) do |file|
42
+ file.close
43
+ Zip::OutputStream.open(file.path) do |zip|
44
+ parts.each do |name, bytes|
45
+ zip.put_next_entry(name, nil, nil, first_stored == name ? Zip::Entry::STORED : Zip::Entry::DEFLATED)
46
+ zip.write(bytes)
47
+ end
48
+ end
49
+ read(file.path) # An output must satisfy the same ZIP limits as an input.
50
+ if Gem.win_platform?
51
+ created = completed = false
52
+ begin
53
+ File.open(target, File::WRONLY | File::CREAT | File::EXCL) do |output|
54
+ created = true
55
+ IO.copy_stream(file.path, output)
56
+ output.flush
57
+ output.fsync
58
+ end
59
+ completed = true
60
+ ensure
61
+ File.unlink(target) if created && !completed && File.file?(target)
62
+ end
63
+ else
64
+ File.link(file.path, target)
65
+ end
66
+ end
67
+ path
68
+ rescue Errno::EEXIST
69
+ raise Error, "target already exists: #{path}"
70
+ end
71
+
72
+ def self.xml(bytes)
73
+ raise InvalidPackage, "XML DTD is not supported" if bytes.match?(/<!\s*(?:DOCTYPE|ENTITY)/i)
74
+ document = REXML::Document.new(bytes)
75
+ raise InvalidPackage, "empty XML part" unless document.root
76
+ document.children.each do |child|
77
+ next if child.is_a?(REXML::XMLDecl) || child.equal?(document.root)
78
+ next if child.is_a?(REXML::Text) && child.value.strip.empty?
79
+
80
+ raise UnsupportedFeature, "unsupported XML prolog content"
81
+ end
82
+
83
+ document.root
84
+ rescue REXML::ParseException => error
85
+ raise InvalidPackage, "invalid XML: #{error.message}"
86
+ end
87
+
88
+ def self.check(element, namespace, name, attributes: [], children: [], text: false)
89
+ raise UnsupportedFeature, "unexpected XML element: #{element.expanded_name}" unless element.namespace == namespace && element.name == name
90
+
91
+ element.attributes.each_attribute do |attribute|
92
+ next if attribute.expanded_name == "xmlns" || attribute.expanded_name.start_with?("xmlns:")
93
+ namespace = attribute.expanded_name == "xml:space" ? "http://www.w3.org/XML/1998/namespace" : attribute.namespace
94
+ raise UnsupportedFeature, "unsupported attribute: #{attribute.expanded_name}" unless attributes.include?([namespace, attribute.name])
95
+ end
96
+ element.elements.each do |child|
97
+ raise UnsupportedFeature, "unsupported XML element: #{child.expanded_name}" unless children.include?([child.namespace, child.name])
98
+ end
99
+ text_nodes = element.children.select { |child| child.is_a?(REXML::Text) }
100
+ if text_nodes.any? { |child| child.is_a?(REXML::CData) } || (text && text_nodes.length > 1)
101
+ raise UnsupportedFeature, "CDATA or split XML text is not supported"
102
+ end
103
+ element.children.each do |child|
104
+ if child.is_a?(REXML::Text)
105
+ raise UnsupportedFeature, "unexpected XML text" if !text && !child.value.strip.empty?
106
+ elsif !child.is_a?(REXML::Element)
107
+ raise UnsupportedFeature, "unsupported XML content"
108
+ end
109
+ end
110
+ element
111
+ end
112
+
113
+ def self.render(root)
114
+ document = REXML::Document.new
115
+ document << REXML::XMLDecl.new("1.0", "UTF-8")
116
+ document.add_element(root)
117
+ document.to_s
118
+ end
119
+ end
120
+ end
@@ -0,0 +1,5 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Lesath
4
+ VERSION = "0.1.0"
5
+ end
@@ -0,0 +1,75 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Lesath
4
+ class Error < StandardError; end
5
+ class UnsupportedFeature < Error; end
6
+ class InvalidPackage < Error; end
7
+
8
+ Cell = Data.define(:value, :formula)
9
+
10
+ class Workbook
11
+ MAX_ROWS = 1_048_576
12
+ MAX_COLUMNS = 16_384
13
+ MAX_CELLS = 100_000
14
+ MAX_SHEETS = 200
15
+
16
+ attr_reader :source_format
17
+
18
+ def initialize(source_format: nil)
19
+ @source_format = source_format
20
+ @sheets = {}
21
+ end
22
+
23
+ def sheet_names = @sheets.keys.freeze
24
+
25
+ def add_sheet(name)
26
+ name = xml_text(name.to_s)
27
+ raise Error, "invalid sheet name" if name.empty? || name.length > 31 || name.match?(/[\[\]:*?\\\/]/) || name.start_with?("'") || name.end_with?("'")
28
+ raise Error, "duplicate sheet name: #{name}" if @sheets.keys.any? { |key| key.casecmp?(name) }
29
+ raise Error, "too many sheets" if @sheets.length >= MAX_SHEETS
30
+
31
+ @sheets[name.freeze] = {}
32
+ self
33
+ end
34
+
35
+ def cell(sheet, row, column)
36
+ @sheets.fetch(sheet) { raise Error, "unknown sheet: #{sheet}" }[[row, column]]
37
+ end
38
+
39
+ def set(sheet, row, column, value = nil, formula: nil)
40
+ cells = @sheets.fetch(sheet) { raise Error, "unknown sheet: #{sheet}" }
41
+ raise Error, "cell coordinates out of range" unless row.is_a?(Integer) && row.between?(1, MAX_ROWS) && column.is_a?(Integer) && column.between?(1, MAX_COLUMNS)
42
+ raise Error, "unsupported cell value" unless value.nil? || value.is_a?(String) || value == true || value == false || value.is_a?(Integer) || (value.is_a?(Float) && value.finite?)
43
+ raise Error, "integer exceeds spreadsheet precision" if value.is_a?(Integer) && value.abs > 999_999_999_999_999
44
+ raise Error, "invalid formula" if formula && (!formula.is_a?(String) || formula.empty?)
45
+ value = xml_text(value) if value.is_a?(String)
46
+ formula = xml_text(formula) if formula
47
+ raise Error, "cell text exceeds 32767 UTF-16 units" if value.is_a?(String) && value.encode("UTF-16LE").bytesize / 2 > 32_767
48
+ raise Error, "too many cells" if !cells.key?([row, column]) && cell_count >= MAX_CELLS
49
+
50
+ value.nil? && formula.nil? ? cells.delete([row, column]) : cells[[row, column]] = Cell.new(value:, formula:)
51
+ self
52
+ end
53
+
54
+ def each_cell(sheet, &block)
55
+ return enum_for(__method__, sheet) unless block
56
+
57
+ @sheets.fetch(sheet) { raise Error, "unknown sheet: #{sheet}" }.sort.each do |(row, column), cell|
58
+ block.call(row, column, cell)
59
+ end
60
+ self
61
+ end
62
+
63
+ def cell_count = @sheets.values.sum(&:size)
64
+
65
+ private
66
+
67
+ def xml_text(value)
68
+ text = value.encode(Encoding::UTF_8)
69
+ raise Error, "invalid XML text" unless text.valid_encoding? && !text.match?(/[\u0000-\u0008\u000B\u000C\u000E-\u001F]/)
70
+ text.freeze
71
+ rescue EncodingError
72
+ raise Error, "invalid XML text"
73
+ end
74
+ end
75
+ end
@@ -0,0 +1,240 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Lesath
4
+ module XLSX
5
+ MAIN = "http://schemas.openxmlformats.org/spreadsheetml/2006/main"
6
+ DOC_REL = "http://schemas.openxmlformats.org/officeDocument/2006/relationships"
7
+ PKG_REL = "http://schemas.openxmlformats.org/package/2006/relationships"
8
+ TYPES = "http://schemas.openxmlformats.org/package/2006/content-types"
9
+ WORKBOOK_REL = "http://schemas.openxmlformats.org/officeDocument/2006/relationships/officeDocument"
10
+ SHEET_REL = "http://schemas.openxmlformats.org/officeDocument/2006/relationships/worksheet"
11
+ module_function
12
+
13
+ def read(path)
14
+ parts = Package.read(path)
15
+ root_rels(parts.fetch("_rels/.rels") { raise InvalidPackage, "missing root relationships" })
16
+ relationships = workbook_rels(parts.fetch("xl/_rels/workbook.xml.rels") { raise InvalidPackage, "missing workbook relationships" })
17
+ workbook_xml = parts.fetch("xl/workbook.xml") { raise InvalidPackage, "missing workbook" }
18
+ names = workbook_sheets(workbook_xml)
19
+ expected = ["[Content_Types].xml", "_rels/.rels", "xl/workbook.xml", "xl/_rels/workbook.xml.rels"]
20
+ book = Workbook.new(source_format: :xlsx)
21
+ names.each do |name, id|
22
+ target = relationships.delete(id) { raise InvalidPackage, "sheet relationship missing: #{id}" }
23
+ expected << target
24
+ book.add_sheet(name)
25
+ parse_sheet(parts.fetch(target) { raise InvalidPackage, "sheet part missing: #{target}" }, book, name)
26
+ end
27
+ raise UnsupportedFeature, "unused workbook relationship" unless relationships.empty?
28
+ raise UnsupportedFeature, "unsupported package parts: #{(parts.keys - expected).join(', ')}" unless (parts.keys - expected).empty?
29
+ content_types(parts.fetch("[Content_Types].xml"), expected)
30
+ book
31
+ end
32
+
33
+ def write(book, path)
34
+ raise Error, "workbook has no sheets" if book.sheet_names.empty?
35
+ raise UnsupportedFeature, "cannot translate ODS formulas to xlsx" if book.source_format == :ods && book.sheet_names.any? { |name| book.each_cell(name).any? { |_r, _c, cell| cell.formula } }
36
+
37
+ parts = {}
38
+ parts["[Content_Types].xml"] = types_xml(book.sheet_names.length)
39
+ parts["_rels/.rels"] = relationships_xml([["rId1", WORKBOOK_REL, "xl/workbook.xml"]])
40
+ parts["xl/workbook.xml"] = workbook_xml(book.sheet_names)
41
+ parts["xl/_rels/workbook.xml.rels"] = relationships_xml(book.sheet_names.each_index.map { |index| ["rId#{index + 1}", SHEET_REL, "worksheets/sheet#{index + 1}.xml"] })
42
+ book.sheet_names.each_with_index { |name, index| parts["xl/worksheets/sheet#{index + 1}.xml"] = sheet_xml(book, name) }
43
+ Package.write(parts, path)
44
+ end
45
+
46
+ def root_rels(bytes)
47
+ root = Package.xml(bytes)
48
+ Package.check(root, PKG_REL, "Relationships", children: [[PKG_REL, "Relationship"]])
49
+ rels = root.elements.to_a
50
+ raise InvalidPackage, "invalid root relationships" unless rels.length == 1
51
+ rel = Package.check(rels.first, PKG_REL, "Relationship", attributes: [["", "Id"], ["", "Type"], ["", "Target"]])
52
+ raise UnsupportedFeature, "unsupported root relationship" unless rel.attributes["Type"] == WORKBOOK_REL && rel.attributes["Target"] == "xl/workbook.xml"
53
+ end
54
+
55
+ def workbook_rels(bytes)
56
+ root = Package.xml(bytes)
57
+ Package.check(root, PKG_REL, "Relationships", children: [[PKG_REL, "Relationship"]])
58
+ relationships = root.elements.to_a.to_h do |rel|
59
+ Package.check(rel, PKG_REL, "Relationship", attributes: [["", "Id"], ["", "Type"], ["", "Target"]])
60
+ raise UnsupportedFeature, "unsupported workbook relationship" unless rel.attributes["Type"] == SHEET_REL
61
+ target = rel.attributes["Target"]
62
+ raise UnsupportedFeature, "unsafe worksheet relationship" unless target&.match?(/\Aworksheets\/sheet\d+\.xml\z/)
63
+ [rel.attributes["Id"], "xl/#{target}"]
64
+ end
65
+ raise InvalidPackage, "duplicate workbook relationship" if relationships.length != root.elements.to_a.length
66
+ relationships
67
+ end
68
+
69
+ def workbook_sheets(bytes)
70
+ root = Package.xml(bytes)
71
+ Package.check(root, MAIN, "workbook", children: [[MAIN, "sheets"], [MAIN, "calcPr"]])
72
+ sheets = root.elements.to_a.select { |element| element.name == "sheets" }
73
+ raise InvalidPackage, "missing sheets" unless sheets.length == 1
74
+ root.elements.to_a.select { |element| element.name == "calcPr" }.each do |element|
75
+ Package.check(element, MAIN, "calcPr", attributes: [["", "fullCalcOnLoad"]])
76
+ end
77
+ Package.check(sheets.first, MAIN, "sheets", children: [[MAIN, "sheet"]])
78
+ names = sheets.first.elements.to_a.map do |sheet|
79
+ Package.check(sheet, MAIN, "sheet", attributes: [["", "name"], ["", "sheetId"], [DOC_REL, "id"]])
80
+ [sheet.attributes["name"], sheet.attributes["r:id"]]
81
+ end
82
+ raise InvalidPackage, "workbook has no sheets" if names.empty? || names.any? { |name, id| !name || !id }
83
+ names
84
+ end
85
+
86
+ def content_types(bytes, expected)
87
+ root = Package.xml(bytes)
88
+ Package.check(root, TYPES, "Types", children: [[TYPES, "Default"], [TYPES, "Override"]])
89
+ overrides = root.elements.to_a.filter_map do |element|
90
+ if element.name == "Default"
91
+ Package.check(element, TYPES, "Default", attributes: [["", "Extension"], ["", "ContentType"]])
92
+ raise UnsupportedFeature, "unsupported content type" unless element.attributes["Extension"] == "rels" && element.attributes["ContentType"] == "application/vnd.openxmlformats-package.relationships+xml"
93
+ nil
94
+ else
95
+ Package.check(element, TYPES, "Override", attributes: [["", "PartName"], ["", "ContentType"]])
96
+ part = element.attributes["PartName"].to_s.delete_prefix("/")
97
+ expected_type = part == "xl/workbook.xml" ? "application/vnd.openxmlformats-officedocument.spreadsheetml.sheet.main+xml" : "application/vnd.openxmlformats-officedocument.spreadsheetml.worksheet+xml"
98
+ raise UnsupportedFeature, "unsupported part content type" unless element.attributes["ContentType"] == expected_type
99
+ part
100
+ end
101
+ end
102
+ raise UnsupportedFeature, "content types do not match workbook parts" unless overrides.sort == (expected - ["[Content_Types].xml", "_rels/.rels", "xl/_rels/workbook.xml.rels"]).sort
103
+ end
104
+
105
+ def parse_sheet(bytes, book, name)
106
+ root = Package.xml(bytes)
107
+ Package.check(root, MAIN, "worksheet", children: [[MAIN, "sheetData"]])
108
+ data = root.elements.to_a
109
+ raise InvalidPackage, "missing sheet data" unless data.length == 1
110
+ Package.check(data.first, MAIN, "sheetData", children: [[MAIN, "row"]])
111
+ data.first.elements.each do |row|
112
+ Package.check(row, MAIN, "row", attributes: [["", "r"]], children: [[MAIN, "c"]])
113
+ row.elements.each { |cell| parse_cell(cell, book, name) }
114
+ end
115
+ end
116
+
117
+ def parse_cell(element, book, name)
118
+ Package.check(element, MAIN, "c", attributes: [["", "r"], ["", "t"], ["", "s"]], children: [[MAIN, "f"], [MAIN, "v"], [MAIN, "is"]])
119
+ raise InvalidPackage, "duplicate cell child" if element.elements.to_a.map(&:name).tally.values.any? { |count| count > 1 }
120
+ raise UnsupportedFeature, "cell style is not supported" if element.attributes["s"] && element.attributes["s"] != "0"
121
+ reference = element.attributes["r"]
122
+ raise InvalidPackage, "invalid cell reference" unless reference&.match?(/\A[A-Z]+[1-9]\d*\z/)
123
+ letters, row = reference.match(/\A([A-Z]+)(\d+)\z/).captures
124
+ column = letters.bytes.reduce(0) { |value, byte| value * 26 + byte - 64 }
125
+ row = row.to_i
126
+ raise InvalidPackage, "cell outside row" unless element.parent.attributes["r"].to_i == row
127
+ raise InvalidPackage, "duplicate cell: #{reference}" if book.cell(name, row, column)
128
+ formula = element.elements["f"]
129
+ value = element.elements["v"]
130
+ inline = element.elements["is"]
131
+ Package.check(formula, MAIN, "f", text: true) if formula
132
+ Package.check(value, MAIN, "v", text: true) if value
133
+ if inline
134
+ Package.check(inline, MAIN, "is", children: [[MAIN, "t"]])
135
+ raise UnsupportedFeature, "rich text is not supported" unless inline.elements.to_a.length == 1
136
+ Package.check(inline.elements[1], MAIN, "t", attributes: [["http://www.w3.org/XML/1998/namespace", "space"]], text: true)
137
+ if whitespace_sensitive?(inline.elements[1].text.to_s) && inline.elements[1].attributes["xml:space"] != "preserve"
138
+ raise UnsupportedFeature, "inline string needs xml:space=preserve"
139
+ end
140
+ end
141
+ raise InvalidPackage, "conflicting cell values" if inline && value
142
+ type = element.attributes["t"] || "n"
143
+ parsed = case type
144
+ when "inlineStr" then inline&.elements&.[]("t")&.text.to_s
145
+ when "str" then value&.text.to_s
146
+ when "b" then {"1" => true, "0" => false}.fetch(value&.text) { raise InvalidPackage, "invalid boolean" }
147
+ when "n" then value ? number(value.text) : nil
148
+ else raise UnsupportedFeature, "unsupported cell type: #{type}"
149
+ end
150
+ raise InvalidPackage, "missing inline string" if type == "inlineStr" && !inline
151
+ raise InvalidPackage, "unexpected inline string" if type != "inlineStr" && inline
152
+ raise UnsupportedFeature, "inline formula is not supported" if formula && inline
153
+ raise InvalidPackage, "empty formula" if formula && formula.text.to_s.empty?
154
+ raise UnsupportedFeature, "formula requires a cached value" if formula && !value
155
+ raise UnsupportedFeature, "formula requires a cached value" if formula && parsed.nil?
156
+ book.set(name, row, column, parsed, formula: formula && "=#{formula.text}")
157
+ end
158
+
159
+ def number(text)
160
+ raise InvalidPackage, "invalid number" unless text&.match?(/\A[+-]?(?:\d+\.?\d*|\.\d+)(?:[eE][+-]?\d+)?\z/)
161
+ value = text.match?(/\A[+-]?\d+\z/) ? Integer(text, 10) : Float(text)
162
+ raise InvalidPackage, "non-finite number" unless value.finite?
163
+ value
164
+ end
165
+
166
+ def relationships_xml(entries)
167
+ root = REXML::Element.new("Relationships")
168
+ root.add_attribute("xmlns", PKG_REL)
169
+ entries.each do |id, type, target|
170
+ rel = root.add_element("Relationship")
171
+ rel.add_attributes("Id" => id, "Type" => type, "Target" => target)
172
+ end
173
+ Package.render(root)
174
+ end
175
+
176
+ def types_xml(count)
177
+ root = REXML::Element.new("Types")
178
+ root.add_attribute("xmlns", TYPES)
179
+ root.add_element("Default").add_attributes("Extension" => "rels", "ContentType" => "application/vnd.openxmlformats-package.relationships+xml")
180
+ root.add_element("Override").add_attributes("PartName" => "/xl/workbook.xml", "ContentType" => "application/vnd.openxmlformats-officedocument.spreadsheetml.sheet.main+xml")
181
+ count.times do |index|
182
+ root.add_element("Override").add_attributes("PartName" => "/xl/worksheets/sheet#{index + 1}.xml", "ContentType" => "application/vnd.openxmlformats-officedocument.spreadsheetml.worksheet+xml")
183
+ end
184
+ Package.render(root)
185
+ end
186
+
187
+ def workbook_xml(names)
188
+ root = REXML::Element.new("workbook")
189
+ root.add_attributes("xmlns" => MAIN, "xmlns:r" => DOC_REL)
190
+ sheets = root.add_element("sheets")
191
+ names.each_with_index { |name, index| sheets.add_element("sheet").add_attributes("name" => name, "sheetId" => (index + 1).to_s, "r:id" => "rId#{index + 1}") }
192
+ root.add_element("calcPr").add_attribute("fullCalcOnLoad", "1")
193
+ Package.render(root)
194
+ end
195
+
196
+ def sheet_xml(book, name)
197
+ root = REXML::Element.new("worksheet")
198
+ root.add_attribute("xmlns", MAIN)
199
+ data = root.add_element("sheetData")
200
+ previous = nil
201
+ book.each_cell(name) do |row, column, cell|
202
+ raise UnsupportedFeature, "xlsx formula must start with =" if cell.formula && !cell.formula.start_with?("=")
203
+ current = row == previous ? data.elements.to_a.last : data.add_element("row", {"r" => row.to_s})
204
+ previous = row
205
+ node = current.add_element("c", {"r" => reference(row, column)})
206
+ node.add_element("f").text = cell.formula.delete_prefix("=") if cell.formula
207
+ case cell.value
208
+ when String
209
+ if cell.formula
210
+ node.add_attribute("t", "str")
211
+ node.add_element("v").text = cell.value
212
+ else
213
+ node.add_attribute("t", "inlineStr")
214
+ text = node.add_element("is").add_element("t")
215
+ text.add_attribute("xml:space", "preserve") if whitespace_sensitive?(cell.value)
216
+ text.text = cell.value
217
+ end
218
+ when TrueClass, FalseClass
219
+ node.add_attribute("t", "b")
220
+ node.add_element("v").text = cell.value ? "1" : "0"
221
+ when Numeric then node.add_element("v").text = cell.value.to_s
222
+ end
223
+ end
224
+ Package.render(root)
225
+ end
226
+
227
+ def reference(row, column)
228
+ letters = +""
229
+ while column.positive?
230
+ column, remainder = (column - 1).divmod(26)
231
+ letters.prepend((65 + remainder).chr)
232
+ end
233
+ "#{letters}#{row}"
234
+ end
235
+
236
+ def whitespace_sensitive?(text)
237
+ text.match?(/\A\s|\s\z|[ \t]{2}|[\r\n\t]/)
238
+ end
239
+ end
240
+ end
data/lib/lesath.rb ADDED
@@ -0,0 +1,26 @@
1
+ # frozen_string_literal: true
2
+
3
+ require "rexml/document"
4
+ require "zip"
5
+ require_relative "lesath/version"
6
+ require_relative "lesath/workbook"
7
+ require_relative "lesath/package"
8
+ require_relative "lesath/xlsx"
9
+ require_relative "lesath/ods"
10
+
11
+ module Lesath
12
+ def self.read(path, format: nil)
13
+ format ||= File.extname(path).downcase.delete_prefix(".").to_sym
14
+ codec(format).read(path)
15
+ end
16
+
17
+ def self.write(workbook, path, format: nil)
18
+ format ||= File.extname(path).downcase.delete_prefix(".").to_sym
19
+ codec(format).write(workbook, path)
20
+ end
21
+
22
+ def self.codec(format)
23
+ {xlsx: XLSX, ods: ODS}.fetch(format) { raise ArgumentError, "format must be :xlsx or :ods" }
24
+ end
25
+ private_class_method :codec
26
+ end
metadata ADDED
@@ -0,0 +1,83 @@
1
+ --- !ruby/object:Gem::Specification
2
+ name: lesath
3
+ version: !ruby/object:Gem::Version
4
+ version: 0.1.0
5
+ platform: ruby
6
+ authors:
7
+ - Yudai Takada
8
+ bindir: bin
9
+ cert_chain: []
10
+ date: 1980-01-02 00:00:00.000000000 Z
11
+ dependencies:
12
+ - !ruby/object:Gem::Dependency
13
+ name: rexml
14
+ requirement: !ruby/object:Gem::Requirement
15
+ requirements:
16
+ - - "~>"
17
+ - !ruby/object:Gem::Version
18
+ version: '3.3'
19
+ type: :runtime
20
+ prerelease: false
21
+ version_requirements: !ruby/object:Gem::Requirement
22
+ requirements:
23
+ - - "~>"
24
+ - !ruby/object:Gem::Version
25
+ version: '3.3'
26
+ - !ruby/object:Gem::Dependency
27
+ name: rubyzip
28
+ requirement: !ruby/object:Gem::Requirement
29
+ requirements:
30
+ - - "~>"
31
+ - !ruby/object:Gem::Version
32
+ version: '3.0'
33
+ type: :runtime
34
+ prerelease: false
35
+ version_requirements: !ruby/object:Gem::Requirement
36
+ requirements:
37
+ - - "~>"
38
+ - !ruby/object:Gem::Version
39
+ version: '3.0'
40
+ description: Reads and writes a documented, lossless subset of xlsx and ODS workbooks.
41
+ email:
42
+ - t.yudai92@gmail.com
43
+ executables: []
44
+ extensions: []
45
+ extra_rdoc_files: []
46
+ files:
47
+ - CHANGELOG.md
48
+ - LICENSE.txt
49
+ - README.md
50
+ - docs/adr/001-lossless-subset.md
51
+ - lib/lesath.rb
52
+ - lib/lesath/ods.rb
53
+ - lib/lesath/package.rb
54
+ - lib/lesath/version.rb
55
+ - lib/lesath/workbook.rb
56
+ - lib/lesath/xlsx.rb
57
+ homepage: https://github.com/noxdea/lesath
58
+ licenses:
59
+ - MIT
60
+ metadata:
61
+ allowed_push_host: https://rubygems.org
62
+ homepage_uri: https://github.com/noxdea/lesath
63
+ source_code_uri: https://github.com/noxdea/lesath/tree/main
64
+ changelog_uri: https://github.com/noxdea/lesath/blob/main/CHANGELOG.md
65
+ rubygems_mfa_required: 'true'
66
+ rdoc_options: []
67
+ require_paths:
68
+ - lib
69
+ required_ruby_version: !ruby/object:Gem::Requirement
70
+ requirements:
71
+ - - ">="
72
+ - !ruby/object:Gem::Version
73
+ version: 3.2.0
74
+ required_rubygems_version: !ruby/object:Gem::Requirement
75
+ requirements:
76
+ - - ">="
77
+ - !ruby/object:Gem::Version
78
+ version: '0'
79
+ requirements: []
80
+ rubygems_version: 4.0.16
81
+ specification_version: 4
82
+ summary: Strict xlsx and ODS spreadsheet interchange
83
+ test_files: []