iostreams 1.11.0 → 2.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.md +14 -13
- data/Rakefile +52 -0
- data/docs/CLAUDE.md +9 -0
- data/docs/config.md +157 -0
- data/docs/copy_files.md +75 -0
- data/docs/extensions.md +111 -0
- data/docs/formats.md +188 -0
- data/docs/index.md +388 -0
- data/docs/path.md +652 -0
- data/docs/pgp.md +436 -0
- data/docs/streams.md +337 -0
- data/docs/tutorial.md +483 -0
- data/docs/upgrading.md +217 -0
- data/lib/io_streams/builder.rb +71 -11
- data/lib/io_streams/bzip2/reader.rb +25 -2
- data/lib/io_streams/bzip2/writer.rb +26 -2
- data/lib/io_streams/encode/reader.rb +6 -2
- data/lib/io_streams/encode/writer.rb +9 -5
- data/lib/io_streams/errors.rb +4 -0
- data/lib/io_streams/gzip/reader.rb +5 -1
- data/lib/io_streams/gzip/writer.rb +11 -2
- data/lib/io_streams/io_streams.rb +156 -20
- data/lib/io_streams/line/reader.rb +9 -4
- data/lib/io_streams/line/writer.rb +1 -1
- data/lib/io_streams/path.rb +117 -8
- data/lib/io_streams/paths/file.rb +57 -11
- data/lib/io_streams/paths/http.rb +123 -9
- data/lib/io_streams/paths/matcher.rb +3 -3
- data/lib/io_streams/paths/s3.rb +69 -18
- data/lib/io_streams/paths/sftp/net_ssh.rb +104 -0
- data/lib/io_streams/paths/sftp.rb +103 -64
- data/lib/io_streams/pgp/reader.rb +63 -10
- data/lib/io_streams/pgp/writer.rb +111 -30
- data/lib/io_streams/pgp.rb +256 -71
- data/lib/io_streams/reader.rb +14 -5
- data/lib/io_streams/record/reader.rb +75 -6
- data/lib/io_streams/record/writer.rb +3 -4
- data/lib/io_streams/row/reader.rb +1 -1
- data/lib/io_streams/row/writer.rb +1 -1
- data/lib/io_streams/stream.rb +48 -37
- data/lib/io_streams/symmetric_encryption/reader.rb +6 -2
- data/lib/io_streams/symmetric_encryption/writer.rb +8 -4
- data/lib/io_streams/tabular/header.rb +49 -10
- data/lib/io_streams/tabular/parser/array.rb +0 -10
- data/lib/io_streams/tabular/parser/base.rb +10 -0
- data/lib/io_streams/tabular/parser/csv.rb +9 -36
- data/lib/io_streams/tabular/parser/fixed.rb +8 -6
- data/lib/io_streams/tabular/parser/psv.rb +6 -14
- data/lib/io_streams/tabular.rb +5 -10
- data/lib/io_streams/utils.rb +34 -2
- data/lib/io_streams/version.rb +1 -1
- data/lib/io_streams/writer.rb +16 -7
- data/lib/io_streams/xlsx/reader.rb +6 -2
- data/lib/io_streams/zip/reader.rb +4 -0
- data/lib/io_streams/zip/writer.rb +26 -10
- data/lib/iostreams.rb +0 -1
- metadata +46 -112
- data/lib/io_streams/deprecated.rb +0 -216
- data/lib/io_streams/tabular/utility/csv_row.rb +0 -105
- data/test/builder_test.rb +0 -311
- data/test/bzip2_reader_test.rb +0 -27
- data/test/bzip2_writer_test.rb +0 -56
- data/test/deprecated_test.rb +0 -121
- data/test/encode_reader_test.rb +0 -51
- data/test/encode_writer_test.rb +0 -90
- data/test/files/embedded_lines_test.csv +0 -7
- data/test/files/multiple_files.zip +0 -0
- data/test/files/spreadsheet.xlsx +0 -0
- data/test/files/test.csv +0 -4
- data/test/files/test.json +0 -3
- data/test/files/test.psv +0 -4
- data/test/files/text file.txt +0 -3
- data/test/files/text.txt +0 -3
- data/test/files/text.txt.bz2 +0 -0
- data/test/files/text.txt.gz +0 -0
- data/test/files/text.txt.gz.zip +0 -0
- data/test/files/text.zip +0 -0
- data/test/files/text.zip.gz +0 -0
- data/test/files/unclosed_quote_large_test.csv +0 -1658
- data/test/files/unclosed_quote_test.csv +0 -4
- data/test/files/unclosed_quote_test2.csv +0 -3
- data/test/gzip_reader_test.rb +0 -27
- data/test/gzip_writer_test.rb +0 -52
- data/test/io_streams_test.rb +0 -132
- data/test/line_reader_test.rb +0 -325
- data/test/line_writer_test.rb +0 -59
- data/test/minimal_file_reader.rb +0 -25
- data/test/path_test.rb +0 -55
- data/test/paths/file_test.rb +0 -213
- data/test/paths/http_test.rb +0 -34
- data/test/paths/matcher_test.rb +0 -120
- data/test/paths/s3_test.rb +0 -220
- data/test/paths/sftp_test.rb +0 -106
- data/test/pgp_reader_test.rb +0 -46
- data/test/pgp_test.rb +0 -267
- data/test/pgp_writer_test.rb +0 -130
- data/test/record_reader_test.rb +0 -60
- data/test/record_writer_test.rb +0 -82
- data/test/row_reader_test.rb +0 -35
- data/test/row_writer_test.rb +0 -56
- data/test/stream_test.rb +0 -577
- data/test/tabular_test.rb +0 -338
- data/test/test_helper.rb +0 -40
- data/test/utils_test.rb +0 -20
- data/test/xlsx_reader_test.rb +0 -37
- data/test/zip_reader_test.rb +0 -53
- data/test/zip_writer_test.rb +0 -48
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: d74053d56262767ee38729e8fc3810416b340ec7e1a318c37a95579f392a92d6
|
|
4
|
+
data.tar.gz: 7e6e2fa0088656a52a8fd9c45c78b96ed8b2c324e9f9379a091fcfdd3939dde9
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: a492f9f89787eae3675fcf0aef1a1fbc9f67e461a0c1d0c9422888fa4cc66277c55d39248041b79a50de1b0d321a51a013445379de93404b0f85105d48914962
|
|
7
|
+
data.tar.gz: 2e6303a89a55bbbda4a9c73a19bdd6780823633afa4391d2fe3071ba8ca85e360a08214e76ed2e869182d8ed94388ed0454769b5d6dc41e540abf704d1c8a794
|
data/README.md
CHANGED
|
@@ -1,8 +1,9 @@
|
|
|
1
1
|
# IOStreams
|
|
2
2
|
[](https://rubygems.org/gems/iostreams) [](https://rubygems.org/gems/iostreams) [](http://opensource.org/licenses/Apache-2.0) 
|
|
3
3
|
|
|
4
|
-
IOStreams is
|
|
5
|
-
|
|
4
|
+
IOStreams is a streaming library for Ruby that makes compression, encryption, file format, and storage
|
|
5
|
+
location transparent to your code. Read and write files of any size, one block at a time, whether they
|
|
6
|
+
are gzip, zip, or PGP encrypted, and whether they live on local disk, AWS S3, SFTP, or are fetched over HTTP.
|
|
6
7
|
|
|
7
8
|
## Project Status
|
|
8
9
|
|
|
@@ -10,26 +11,26 @@ Production Ready, heavily used in production environments, many as part of Rocke
|
|
|
10
11
|
|
|
11
12
|
## Documentation
|
|
12
13
|
|
|
13
|
-
Start with the [IOStreams tutorial](https://iostreams.
|
|
14
|
+
Start with the [IOStreams tutorial](https://iostreams.reidmorrison.com/tutorial) to get a great introduction to IOStreams.
|
|
14
15
|
|
|
15
|
-
Next, checkout the remaining [IOStreams documentation](https://iostreams.
|
|
16
|
+
Next, checkout the remaining [IOStreams documentation](https://iostreams.reidmorrison.com/)
|
|
16
17
|
|
|
17
|
-
|
|
18
|
+
See the [CHANGELOG](CHANGELOG.md) for the release history and notable changes.
|
|
18
19
|
|
|
19
|
-
|
|
20
|
-
the following line to your code:
|
|
20
|
+
## Upgrading
|
|
21
21
|
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
~~~
|
|
25
|
-
|
|
26
|
-
It is important to move any of the old deprecated apis over to the new api, since they will be removed in a future
|
|
27
|
-
release.
|
|
22
|
+
See [Upgrading IOStreams](https://iostreams.reidmorrison.com/upgrading) for the changes that may need
|
|
23
|
+
updates to your application, and the security settings to review, when upgrading.
|
|
28
24
|
|
|
29
25
|
## Versioning
|
|
30
26
|
|
|
31
27
|
This project adheres to [Semantic Versioning](http://semver.org/).
|
|
32
28
|
|
|
29
|
+
## Contributing
|
|
30
|
+
|
|
31
|
+
Contributions are welcome. See [CONTRIBUTING.md](CONTRIBUTING.md) for guidelines on documentation
|
|
32
|
+
updates, code changes, the project architecture, and the code of conduct.
|
|
33
|
+
|
|
33
34
|
## Author
|
|
34
35
|
|
|
35
36
|
[Reid Morrison](https://github.com/reidmorrison)
|
data/Rakefile
CHANGED
|
@@ -1,10 +1,12 @@
|
|
|
1
1
|
require "rake/testtask"
|
|
2
2
|
require_relative "lib/io_streams/version"
|
|
3
3
|
|
|
4
|
+
desc "Build the iostreams gem"
|
|
4
5
|
task :gem do
|
|
5
6
|
system "gem build iostreams.gemspec"
|
|
6
7
|
end
|
|
7
8
|
|
|
9
|
+
desc "Build and publish the iostreams gem, then tag and push the release"
|
|
8
10
|
task publish: :gem do
|
|
9
11
|
system "git tag -a v#{IOStreams::VERSION} -m 'Tagging #{IOStreams::VERSION}'"
|
|
10
12
|
system "git push --tags"
|
|
@@ -12,6 +14,56 @@ task publish: :gem do
|
|
|
12
14
|
system "rm iostreams-#{IOStreams::VERSION}.gem"
|
|
13
15
|
end
|
|
14
16
|
|
|
17
|
+
desc "Start an IRB console with the gem loaded"
|
|
18
|
+
task :console do
|
|
19
|
+
exec "irb -I lib -r iostreams"
|
|
20
|
+
end
|
|
21
|
+
|
|
22
|
+
desc "Generate docs/llms-full.txt from docs/llms.txt and the doc pages it links to"
|
|
23
|
+
task :llms_full do
|
|
24
|
+
require "uri"
|
|
25
|
+
|
|
26
|
+
docs_dir = File.join(__dir__, "docs")
|
|
27
|
+
llms_path = File.join(docs_dir, "llms.txt")
|
|
28
|
+
out_path = File.join(docs_dir, "llms-full.txt")
|
|
29
|
+
|
|
30
|
+
llms_txt = File.read(llms_path)
|
|
31
|
+
header = llms_txt[/\A.*?(?=\n## Docs)/m]&.strip
|
|
32
|
+
raise "Could not find intro text before '## Docs' in #{llms_path}" unless header
|
|
33
|
+
|
|
34
|
+
docs_section = llms_txt[/^## Docs\n(.*?)(?=\n## |\z)/m, 1]
|
|
35
|
+
raise "Could not find '## Docs' section in #{llms_path}" unless docs_section
|
|
36
|
+
|
|
37
|
+
sections = docs_section.each_line.filter_map do |line|
|
|
38
|
+
next unless line =~ /^- \[(?<title>[^\]]+)\]\((?<url>[^)]+)\)/
|
|
39
|
+
|
|
40
|
+
title = $~[:title]
|
|
41
|
+
path = URI.parse($~[:url]).path.sub(%r{\A/}, "")
|
|
42
|
+
path = "index" if path.empty?
|
|
43
|
+
file = File.join(docs_dir, "#{path}.md")
|
|
44
|
+
raise "Missing doc file for #{$~[:url]}: #{file}" unless File.exist?(file)
|
|
45
|
+
|
|
46
|
+
raw = File.read(file)
|
|
47
|
+
|
|
48
|
+
# The page heading lives in front matter, which the shared docs theme
|
|
49
|
+
# renders as the page's h1. Front matter is stripped below, so lift the
|
|
50
|
+
# heading back out and re-emit it. `heading` wins over `title` where a page
|
|
51
|
+
# sets both, which is the same precedence the theme uses. Falls back to the
|
|
52
|
+
# link text in llms.txt for a page that sets neither.
|
|
53
|
+
front_matter = raw[/\A---\n(.*?)\n---\n/m, 1].to_s
|
|
54
|
+
heading = %w[heading title].
|
|
55
|
+
filter_map { |key| front_matter[/^#{key}:[ \t]*(.+)$/, 1] }.
|
|
56
|
+
first.to_s.strip.delete_prefix('"').delete_suffix('"')
|
|
57
|
+
heading = title if heading.empty?
|
|
58
|
+
|
|
59
|
+
body = raw.sub(/\A---\n.*?\n---\n/m, "").strip
|
|
60
|
+
"# #{heading}\n\n#{body}"
|
|
61
|
+
end
|
|
62
|
+
|
|
63
|
+
File.write(out_path, "#{header}\n\n#{sections.join("\n\n---\n\n")}\n")
|
|
64
|
+
puts "Wrote #{out_path}"
|
|
65
|
+
end
|
|
66
|
+
|
|
15
67
|
Rake::TestTask.new(:test) do |t|
|
|
16
68
|
t.pattern = "test/**/*_test.rb"
|
|
17
69
|
t.verbose = true
|
data/docs/CLAUDE.md
ADDED
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# Documentation site
|
|
2
|
+
|
|
3
|
+
User-facing documentation is a Jekyll site under [docs/](docs/), published to iostreams.reidmorrison.com. The markdown pages are what matter when reading or updating documentation.
|
|
4
|
+
|
|
5
|
+
**The look and feel is not in this repo.** `docs/_config.yml` sets `remote_theme: reidmorrison/rm-docs-theme@v1`, and the layout, stylesheet, sidebar and syntax highlighting all come from there. This repo holds only its content: the markdown pages and `docs/images`. **Do not add a `docs/_layouts`, `docs/stylesheets` or `docs/javascripts` directory**; they were deleted deliberately, because six gem repos each carried a near-identical copy of the same theme and the copies had drifted. A styling change belongs in `rm-docs-theme`, where it reaches every doc site at once. `v1` is a moving major tag, so theme fixes arrive on the next build; breaking changes go to `v2` and are opted into by editing the pin. `jekyll-remote-theme` must stay in `plugins`: GitHub Pages enables it on its own, but a local build does not, and without it every page silently renders with no layout. Preview against a local theme checkout with `~/src/rm-docs-theme/bin/preview ~/src/iostreams/docs`.
|
|
6
|
+
|
|
7
|
+
**A page's title lives in its front matter**, not in a `#` heading at the top of the markdown; the theme renders it as the page's `h1`. `title` is the browser title and the default heading, `heading` overrides the h1 where the two should differ, and `description` is the page's meta description. `index.md` sets `heading` only, so the home page keeps the tuned SEO `<title>` from `_config.yml`. Adding or renaming a page means editing the `nav` block in `docs/_config.yml` and `docs/llms.txt`.
|
|
8
|
+
|
|
9
|
+
The site also serves two files for AI assistants: [docs/llms.txt](docs/llms.txt), a hand-maintained index of the docs pages, and `docs/llms-full.txt`, all pages concatenated, regenerated with `bundle exec rake llms_full`. **After editing any `docs/*.md` page, re-run `bundle exec rake llms_full`** and commit the result; never edit `llms-full.txt` by hand. That task reads the page heading out of the front matter, so a page that sets neither `heading` nor `title` falls back to its link text in `llms.txt`.
|
data/docs/config.md
ADDED
|
@@ -0,0 +1,157 @@
|
|
|
1
|
+
---
|
|
2
|
+
layout: default
|
|
3
|
+
title: Configuring IOStreams
|
|
4
|
+
description: >-
|
|
5
|
+
Named roots via IOStreams.add_root so the same code targets different storage
|
|
6
|
+
per environment, plus the temp directory and logger settings.
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## add_root
|
|
10
|
+
|
|
11
|
+
Roots allow paths to reference a particular root directory, so that all path names are appended to that root.
|
|
12
|
+
Their primary purpose is to allow the exact same code to run in production and development, yet use completely
|
|
13
|
+
different data sources in each. For example, in production a root can point to an S3 bucket, while in
|
|
14
|
+
development it points to the local file system.
|
|
15
|
+
|
|
16
|
+
Roots are configured via an initializer at startup. `IOStreams.join` then joins the supplied path
|
|
17
|
+
elements onto the named root, using the `:default` root whenever a root is not supplied.
|
|
18
|
+
|
|
19
|
+
Set the default root for this environment in an initializer:
|
|
20
|
+
~~~ruby
|
|
21
|
+
IOStreams.add_root(:default, "/var/my_app/files")
|
|
22
|
+
~~~
|
|
23
|
+
|
|
24
|
+
Now the default root path is available:
|
|
25
|
+
~~~ruby
|
|
26
|
+
IOStreams.root
|
|
27
|
+
# => #<IOStreams::Paths::File:/var/my_app/files pipeline={}>
|
|
28
|
+
|
|
29
|
+
IOStreams.root.to_s
|
|
30
|
+
# => "/var/my_app/files"
|
|
31
|
+
~~~
|
|
32
|
+
|
|
33
|
+
Comparing the final path using `path` and then `join` that uses a root path:
|
|
34
|
+
~~~ruby
|
|
35
|
+
IOStreams.path("/var/my_app/files", "my_test_file.txt").to_s
|
|
36
|
+
# => "/var/my_app/files/my_test_file.txt"
|
|
37
|
+
|
|
38
|
+
IOStreams.join("my_test_file.txt").to_s
|
|
39
|
+
# => "/var/my_app/files/my_test_file.txt"
|
|
40
|
+
~~~
|
|
41
|
+
|
|
42
|
+
|
|
43
|
+
Using `path`:
|
|
44
|
+
~~~ruby
|
|
45
|
+
IOStreams.path("/var/my_app/files", "my_test_file.txt").write("Hello World")
|
|
46
|
+
~~~
|
|
47
|
+
|
|
48
|
+
With the default root path configured the above code can be simplified by using `join` since it resolves to the same path.
|
|
49
|
+
~~~ruby
|
|
50
|
+
IOStreams.join("my_test_file.txt").write("Hello World")
|
|
51
|
+
~~~
|
|
52
|
+
|
|
53
|
+
Multiple roots can be setup, for example one for input files, another for output files, another for
|
|
54
|
+
reports, etc. During development the roots can all point to a common location, while in production
|
|
55
|
+
they could be completely different S3 buckets.
|
|
56
|
+
|
|
57
|
+
For example add special paths for `downloads` and `uploads`.
|
|
58
|
+
~~~ruby
|
|
59
|
+
IOStreams.add_root(:downloads, "/var/my_app/downloads")
|
|
60
|
+
IOStreams.add_root(:uploads, "/var/my_app/uploads")
|
|
61
|
+
~~~
|
|
62
|
+
|
|
63
|
+
An example that writes a file into the `/var/my_app/downloads` directory:
|
|
64
|
+
~~~ruby
|
|
65
|
+
IOStreams.join("my_test_file.txt", root: :downloads).write("Hello World")
|
|
66
|
+
~~~
|
|
67
|
+
|
|
68
|
+
The other benefit is that the root paths used in an application are externalized from the code base. That way the
|
|
69
|
+
roots can be changed to different locations depending on the environment.
|
|
70
|
+
|
|
71
|
+
We can also change the storage mechanism by changing the root:
|
|
72
|
+
~~~ruby
|
|
73
|
+
IOStreams.add_root(:downloads, "s3://my-app-bucket-name/downloads")
|
|
74
|
+
IOStreams.add_root(:uploads, "s3://my-app-bucket-name/uploads")
|
|
75
|
+
~~~
|
|
76
|
+
|
|
77
|
+
Now the application will write to S3 and the code does not change at all.
|
|
78
|
+
~~~ruby
|
|
79
|
+
IOStreams.join("my_test_file.txt", root: :downloads).write("Hello World")
|
|
80
|
+
~~~
|
|
81
|
+
|
|
82
|
+
To use or query a configured root path:
|
|
83
|
+
~~~ruby
|
|
84
|
+
IOStreams.root(:downloads).to_s
|
|
85
|
+
# => "s3://my-app-bucket-name/downloads"
|
|
86
|
+
~~~
|
|
87
|
+
|
|
88
|
+
## temp_dir
|
|
89
|
+
|
|
90
|
+
When working with large files the standard temp file system location can be too small to handle downloading large
|
|
91
|
+
files. For example to decrypt a pgp file from S3, because GnuPG is not streaming capable and only operates on local files.
|
|
92
|
+
|
|
93
|
+
By default IOStreams looks up the location to store temp files in the following order:
|
|
94
|
+
* `ENV['TMPDIR']`
|
|
95
|
+
* `ENV['TMP']`
|
|
96
|
+
* `ENV['TEMP']`
|
|
97
|
+
* `Etc.systmpdir`
|
|
98
|
+
* `/tmp` (if it exists)
|
|
99
|
+
* Otherwise `.`
|
|
100
|
+
|
|
101
|
+
To explicity set the temp file location the following config option can be used:
|
|
102
|
+
|
|
103
|
+
~~~ruby
|
|
104
|
+
IOStreams.temp_dir = "/var/really_big_temp"
|
|
105
|
+
~~~
|
|
106
|
+
|
|
107
|
+
### temp_file
|
|
108
|
+
|
|
109
|
+
To work with a temporary file directly, `IOStreams.temp_file` yields a path inside `temp_dir`
|
|
110
|
+
and deletes the file when the block completes:
|
|
111
|
+
|
|
112
|
+
~~~ruby
|
|
113
|
+
IOStreams.temp_file("export", ".csv") do |path|
|
|
114
|
+
path.write("Hello World")
|
|
115
|
+
# ... use the temp file ...
|
|
116
|
+
end
|
|
117
|
+
# The temp file has been deleted.
|
|
118
|
+
~~~
|
|
119
|
+
|
|
120
|
+
The first argument is a base file name to include in the generated temp file name, and the
|
|
121
|
+
optional second argument is the file extension.
|
|
122
|
+
|
|
123
|
+
## logger
|
|
124
|
+
|
|
125
|
+
IOStreams can log debug information, such as the external commands it runs for PGP and SFTP.
|
|
126
|
+
|
|
127
|
+
When [Semantic Logger](https://logger.reidmorrison.com) is loaded it is detected automatically, and IOStreams
|
|
128
|
+
logs to it without any additional configuration.
|
|
129
|
+
|
|
130
|
+
To use a different logger, or to log when Semantic Logger is not present, assign any logger that
|
|
131
|
+
responds to the standard logging methods:
|
|
132
|
+
|
|
133
|
+
~~~ruby
|
|
134
|
+
require "logger"
|
|
135
|
+
IOStreams.logger = Logger.new($stdout)
|
|
136
|
+
~~~
|
|
137
|
+
|
|
138
|
+
To disable logging entirely, set the logger to `nil`:
|
|
139
|
+
|
|
140
|
+
~~~ruby
|
|
141
|
+
IOStreams.logger = nil
|
|
142
|
+
~~~
|
|
143
|
+
|
|
144
|
+
IOStreams also logs warnings for behavior that will change in the next major version. See
|
|
145
|
+
[Coming in v3.0](upgrading#coming-in-v30).
|
|
146
|
+
|
|
147
|
+
## enforce_column_restrictions
|
|
148
|
+
|
|
149
|
+
Applies `allowed_columns`, `required_columns` and `skip_unknown` to every input when reading
|
|
150
|
+
records, including JSON records, so that renaming an uploaded file from `.csv` to `.json` cannot
|
|
151
|
+
bypass them. Set it in an initializer:
|
|
152
|
+
|
|
153
|
+
~~~ruby
|
|
154
|
+
IOStreams.enforce_column_restrictions = true
|
|
155
|
+
~~~
|
|
156
|
+
|
|
157
|
+
Default: false. It will default to true in v3.0. See [Header options](formats#header-options).
|
data/docs/copy_files.md
ADDED
|
@@ -0,0 +1,75 @@
|
|
|
1
|
+
---
|
|
2
|
+
layout: default
|
|
3
|
+
title: Copying Between Files
|
|
4
|
+
description: >-
|
|
5
|
+
Using IOStreams.copy to move a file between storage locations and to compress,
|
|
6
|
+
decompress, encrypt or decrypt it on the way.
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
File copying can be used to:
|
|
10
|
+
* copy from one storage location to another.
|
|
11
|
+
* create a decrypted / encrypted copy of an existing file.
|
|
12
|
+
* create a decompressed / compressed copy of an existing file.
|
|
13
|
+
|
|
14
|
+
## Examples
|
|
15
|
+
|
|
16
|
+
Decompress `example.csv.gz` into `example.csv`:
|
|
17
|
+
|
|
18
|
+
~~~ruby
|
|
19
|
+
source = IOStreams.path("example.csv.gz")
|
|
20
|
+
IOStreams.path("example.csv").copy_from(source)
|
|
21
|
+
~~~
|
|
22
|
+
|
|
23
|
+
Decrypt a file encrypted with Symmetric Encryption:
|
|
24
|
+
|
|
25
|
+
~~~ruby
|
|
26
|
+
source = IOStreams.path("example.csv.enc")
|
|
27
|
+
IOStreams.path("example.csv").copy_from(source)
|
|
28
|
+
~~~
|
|
29
|
+
|
|
30
|
+
Encrypt a file using PGP encryption so that it can only be read by `receiver@example.org`.
|
|
31
|
+
|
|
32
|
+
~~~ruby
|
|
33
|
+
source = IOStreams.path("example.csv")
|
|
34
|
+
target = IOStreams.path("example.csv.pgp")
|
|
35
|
+
target.option(:pgp, recipient: "receiver@example.org")
|
|
36
|
+
target.copy_from(source)
|
|
37
|
+
~~~
|
|
38
|
+
|
|
39
|
+
When the file name does not have file extensions that would allow IOStreams to infer what streams to apply,
|
|
40
|
+
the streams can be explicitly set using `stream`.
|
|
41
|
+
|
|
42
|
+
In this example, the file `CUSTOMER_DATA` has no extensions, so `stream(:enc)` tells IOStreams
|
|
43
|
+
that its contents were encrypted with Symmetric Encryption. The decrypted contents are then
|
|
44
|
+
PGP encrypted and written to `xyz.csv.pgp` using the pgp key for `receiver@example.org`.
|
|
45
|
+
|
|
46
|
+
~~~ruby
|
|
47
|
+
input = IOStreams.path("CUSTOMER_DATA").stream(:enc)
|
|
48
|
+
IOStreams.path("xyz.csv.pgp").option(:pgp, recipient: "receiver@example.org").copy_from(input)
|
|
49
|
+
~~~
|
|
50
|
+
|
|
51
|
+
To copy a file _without_ performing any conversions (ignore file extensions), set `convert` to `false`:
|
|
52
|
+
|
|
53
|
+
~~~ruby
|
|
54
|
+
input = IOStreams.path("sample.json.zip")
|
|
55
|
+
IOStreams.path("sample.copy").copy_from(input, convert: false)
|
|
56
|
+
~~~
|
|
57
|
+
|
|
58
|
+
Custom stream conversions can be applied to both the source and the target in a single copy.
|
|
59
|
+
Here the source is read as binary and the target is PGP encrypted:
|
|
60
|
+
|
|
61
|
+
~~~ruby
|
|
62
|
+
source = IOStreams.path("source_file").stream(:encode, encoding: "BINARY")
|
|
63
|
+
IOStreams.path("target_file.pgp").option(:pgp, passphrase: "hello").copy_from(source)
|
|
64
|
+
~~~
|
|
65
|
+
|
|
66
|
+
To convert the contents row by row, or record by record, during the copy, supply `mode`.
|
|
67
|
+
For example, copy a CSV file into JSON, parsing and rendering each record:
|
|
68
|
+
|
|
69
|
+
~~~ruby
|
|
70
|
+
source = IOStreams.path("source_file.csv")
|
|
71
|
+
IOStreams.path("target_file.json").copy_from(source, mode: :hash)
|
|
72
|
+
~~~
|
|
73
|
+
|
|
74
|
+
Notes:
|
|
75
|
+
* `mode` accepts `:line`, `:array`, or `:hash`, and only applies when `convert` is `true`.
|
data/docs/extensions.md
ADDED
|
@@ -0,0 +1,111 @@
|
|
|
1
|
+
---
|
|
2
|
+
layout: default
|
|
3
|
+
title: File Extensions
|
|
4
|
+
description: >-
|
|
5
|
+
How the extensions in a file name, such as .csv.gz.pgp, decide which streams
|
|
6
|
+
IOStreams applies when reading or writing, and how to register your own.
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
IOStreams uses the extensions in the file name to determine which streams to apply when
|
|
10
|
+
reading or writing a file. Multiple extensions are applied in order, so `sample.csv.gz.pgp`
|
|
11
|
+
is first decrypted with PGP and then decompressed with GZip when read.
|
|
12
|
+
|
|
13
|
+
Supported extensions:
|
|
14
|
+
|
|
15
|
+
| Extension | Stream | Read | Write | Required gem / program |
|
|
16
|
+
|:-----------------|:---------------------|:-----|:------|:----------------------------------|
|
|
17
|
+
| `.bz2` | BZip2 | Yes | Yes | `bzip2-ffi` |
|
|
18
|
+
| `.enc` | Symmetric Encryption | Yes | Yes | `symmetric-encryption` |
|
|
19
|
+
| `.gz`, `.gzip` | GZip | Yes | Yes | None (Ruby standard library) |
|
|
20
|
+
| `.zip` | Zip | Yes | Yes | `rubyzip` (read), `zip_kit` (write). On JRuby the built-in Java zip support is used for reading. |
|
|
21
|
+
| `.pgp`, `.gpg` | PGP | Yes | Yes | GnuPG command line program (`gpg`) |
|
|
22
|
+
| `.xlsx`, `.xlsm` | Excel Spreadsheet | Yes | No | `creek` |
|
|
23
|
+
|
|
24
|
+
The gems above are soft dependencies: IOStreams does not require them for installation,
|
|
25
|
+
they only need to be added to the `Gemfile` when the corresponding extension is used.
|
|
26
|
+
|
|
27
|
+
## Compression options
|
|
28
|
+
|
|
29
|
+
GZip accepts a compression `level` when writing, from `0` (no compression) to `9` (best compression):
|
|
30
|
+
|
|
31
|
+
~~~ruby
|
|
32
|
+
IOStreams.path("sample.csv.gz").option(:gz, level: 9).write(data)
|
|
33
|
+
~~~
|
|
34
|
+
|
|
35
|
+
BZip2 passes its options through to `bzip2-ffi`: `block_size` (`1` to `9`) and `work_factor` (`0` to `250`)
|
|
36
|
+
when writing, and `small` and `first_only` when reading:
|
|
37
|
+
|
|
38
|
+
~~~ruby
|
|
39
|
+
IOStreams.path("sample.csv.bz2").option(:bz2, block_size: 9).write(data)
|
|
40
|
+
IOStreams.path("sample.csv.bz2").option(:bz2, small: true).read
|
|
41
|
+
~~~
|
|
42
|
+
|
|
43
|
+
Options are strict, so an option that a stream does not accept raises an `ArgumentError`.
|
|
44
|
+
See [Streams](streams#reading-and-writing-need-separate-options). The exception is BZip2, which
|
|
45
|
+
ignores any other option and logs a warning. In v3.0 it will raise an `ArgumentError` too.
|
|
46
|
+
|
|
47
|
+
## Reading an Excel Spreadsheet
|
|
48
|
+
|
|
49
|
+
Each row in the spreadsheet is converted into a CSV line, so the regular `:line`, `:array`,
|
|
50
|
+
and `:hash` modes apply:
|
|
51
|
+
|
|
52
|
+
~~~ruby
|
|
53
|
+
IOStreams.path("spreadsheet.xlsx").each(:hash) do |record|
|
|
54
|
+
p record
|
|
55
|
+
end
|
|
56
|
+
~~~
|
|
57
|
+
|
|
58
|
+
Notes:
|
|
59
|
+
* Since the underlying `creek` gem operates on files, when reading from a stream (for example S3 or HTTP)
|
|
60
|
+
the contents are first downloaded into a temp file.
|
|
61
|
+
* Writing xlsx files is not supported.
|
|
62
|
+
|
|
63
|
+
## Character encoding
|
|
64
|
+
|
|
65
|
+
The special `:encode` stream converts the character encoding of the data being read or written.
|
|
66
|
+
It is applied with `option` or `stream` rather than a file name extension:
|
|
67
|
+
|
|
68
|
+
~~~ruby
|
|
69
|
+
IOStreams.path("sample.csv.gz").
|
|
70
|
+
option(:encode, encoding: "UTF-8", cleaner: :printable, replace: "").
|
|
71
|
+
each do |line|
|
|
72
|
+
puts line
|
|
73
|
+
end
|
|
74
|
+
~~~
|
|
75
|
+
|
|
76
|
+
Options:
|
|
77
|
+
|
|
78
|
+
* `encoding: [String|Encoding]`
|
|
79
|
+
The target encoding, for example `"UTF-8"`, `"US-ASCII"`, or `"ASCII-8BIT"`.
|
|
80
|
+
Default: `"UTF-8"`
|
|
81
|
+
|
|
82
|
+
* `replace: [String]`
|
|
83
|
+
The character to replace with when a character cannot be converted to the target encoding.
|
|
84
|
+
Default: nil (raise `Encoding::UndefinedConversionError` on invalid characters)
|
|
85
|
+
|
|
86
|
+
* `cleaner: [nil|Symbol|Proc]`
|
|
87
|
+
Cleanse the data. Built-in rules:
|
|
88
|
+
* `:printable` removes all non-printable characters except `\r` and `\n`.
|
|
89
|
+
* `:replace_non_printable` replaces all non-printable characters except `\r` and `\n`
|
|
90
|
+
with the `replace` value, or an empty string when `replace` is nil.
|
|
91
|
+
A Proc can also be supplied to perform custom cleansing; it is called with the data
|
|
92
|
+
and the `replace` value after every read or write.
|
|
93
|
+
Default: nil
|
|
94
|
+
|
|
95
|
+
## Registering a custom extension
|
|
96
|
+
|
|
97
|
+
To add a new extension, supply its reader and writer classes. Both must implement `.open`
|
|
98
|
+
that yields a stream implementing `#read` or `#write` respectively. See any of the streams
|
|
99
|
+
under `lib/io_streams` for examples.
|
|
100
|
+
|
|
101
|
+
~~~ruby
|
|
102
|
+
IOStreams.register_extension(:xls, MyXls::Reader, MyXls::Writer)
|
|
103
|
+
~~~
|
|
104
|
+
|
|
105
|
+
Similarly, to support a new storage location, supply a Path class for its URI scheme.
|
|
106
|
+
See [IOStreams::Paths::S3](https://github.com/reidmorrison/iostreams/blob/main/lib/io_streams/paths/s3.rb)
|
|
107
|
+
for an example of what is required.
|
|
108
|
+
|
|
109
|
+
~~~ruby
|
|
110
|
+
IOStreams.register_scheme(:gcs, MyGoogleCloudStoragePath)
|
|
111
|
+
~~~
|
data/docs/formats.md
ADDED
|
@@ -0,0 +1,188 @@
|
|
|
1
|
+
---
|
|
2
|
+
layout: default
|
|
3
|
+
title: File Formats
|
|
4
|
+
description: >-
|
|
5
|
+
Converting rows and records to and from CSV, PSV, JSON and fixed width files,
|
|
6
|
+
including format inference, format options and header handling.
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
When reading or writing rows (`:array`) or records (`:hash`), IOStreams converts each line
|
|
10
|
+
to or from the file's tabular format. The following formats are supported:
|
|
11
|
+
|
|
12
|
+
* `:csv` Comma Separated Values
|
|
13
|
+
* `:psv` Pipe Separated Values
|
|
14
|
+
* `:json` One JSON document per line
|
|
15
|
+
* `:fixed` Fixed width columns
|
|
16
|
+
* `:array` Each line is already an array of values
|
|
17
|
+
* `:hash` Each line is already a hash
|
|
18
|
+
|
|
19
|
+
PSV has no way to escape values, so when writing PSV, a `|` within a value is replaced with `:`
|
|
20
|
+
and a line break with a space, so that a value cannot add columns or records.
|
|
21
|
+
|
|
22
|
+
## Format inference
|
|
23
|
+
|
|
24
|
+
The format is inferred from the file name when it contains a recognized extension:
|
|
25
|
+
|
|
26
|
+
~~~ruby
|
|
27
|
+
IOStreams.path("sample.csv").each(:hash) { |record| p record }
|
|
28
|
+
IOStreams.path("sample.json").each(:hash) { |record| p record }
|
|
29
|
+
IOStreams.path("sample.psv").each(:hash) { |record| p record }
|
|
30
|
+
~~~
|
|
31
|
+
|
|
32
|
+
The format extension can appear anywhere in the file name, so `sample.csv.gz` and
|
|
33
|
+
`sample.json.pgp` are recognized as CSV and JSON respectively.
|
|
34
|
+
|
|
35
|
+
When the file name does not contain a recognized format extension, the format defaults to `:csv`.
|
|
36
|
+
|
|
37
|
+
## Specifying the format
|
|
38
|
+
|
|
39
|
+
When the file name cannot be used to infer the format, set it explicitly with `format`:
|
|
40
|
+
|
|
41
|
+
~~~ruby
|
|
42
|
+
path = IOStreams.path("sample_data")
|
|
43
|
+
path.format(:json)
|
|
44
|
+
path.each(:hash) { |record| p record }
|
|
45
|
+
~~~
|
|
46
|
+
|
|
47
|
+
`format` can be chained with the other path methods:
|
|
48
|
+
|
|
49
|
+
~~~ruby
|
|
50
|
+
IOStreams.path("sample_data").format(:json).each(:hash) { |record| p record }
|
|
51
|
+
~~~
|
|
52
|
+
|
|
53
|
+
## Format options
|
|
54
|
+
|
|
55
|
+
Format specific options are supplied with `format_options`. They are passed to the parser
|
|
56
|
+
for the chosen format. The `:fixed` format requires its file layout to be supplied this way,
|
|
57
|
+
as shown in the next section. The other formats do not currently take any options.
|
|
58
|
+
|
|
59
|
+
## Fixed width files
|
|
60
|
+
|
|
61
|
+
Fixed width files have no delimiters; each column is identified by its position within the line.
|
|
62
|
+
Since the layout cannot be inferred from the file, supply it using `format_options`:
|
|
63
|
+
|
|
64
|
+
~~~ruby
|
|
65
|
+
path = IOStreams.path("sample_data")
|
|
66
|
+
path.format(:fixed)
|
|
67
|
+
path.format_options(
|
|
68
|
+
layout: [
|
|
69
|
+
{size: 23, key: "name"},
|
|
70
|
+
{size: 40, key: "address"},
|
|
71
|
+
{size: 5, key: "zip"}
|
|
72
|
+
]
|
|
73
|
+
)
|
|
74
|
+
path.each(:hash) { |record| p record }
|
|
75
|
+
~~~
|
|
76
|
+
|
|
77
|
+
Writing a fixed width file uses the same layout to render each record:
|
|
78
|
+
|
|
79
|
+
~~~ruby
|
|
80
|
+
path = IOStreams.path("sample_data")
|
|
81
|
+
path.format(:fixed)
|
|
82
|
+
path.format_options(
|
|
83
|
+
layout: [
|
|
84
|
+
{size: 23, key: "name"},
|
|
85
|
+
{size: 40, key: "address"},
|
|
86
|
+
{size: 5, key: "zip"}
|
|
87
|
+
]
|
|
88
|
+
)
|
|
89
|
+
path.writer(:hash) do |io|
|
|
90
|
+
io << {"name" => "Jack Jones", "address" => "Somewhere", "zip" => 12345}
|
|
91
|
+
end
|
|
92
|
+
~~~
|
|
93
|
+
|
|
94
|
+
Note: The keys in the hashes being written must match the layout `:key` values exactly,
|
|
95
|
+
including whether they are strings or symbols.
|
|
96
|
+
|
|
97
|
+
Layout column definitions:
|
|
98
|
+
|
|
99
|
+
* `:size` The number of characters this column occupies.
|
|
100
|
+
The last column may use a size of `:remainder` to take the rest of the line as its value.
|
|
101
|
+
* `:key` The name for this column. Leave out the key to ignore the column during parsing,
|
|
102
|
+
and to space fill when rendering.
|
|
103
|
+
* `:type` `:string` (default), `:integer`, or `:float`.
|
|
104
|
+
Strings are left justified and space padded, numbers are right justified and zero padded.
|
|
105
|
+
When writing, line breaks within a string are replaced with a space so that a value cannot
|
|
106
|
+
add records.
|
|
107
|
+
Raises `IOStreams::Errors::ValueTooLong` when an `:integer` or `:float` value cannot be
|
|
108
|
+
rendered in `size` characters.
|
|
109
|
+
* `:decimals` For `:float` columns, the number of decimal places to render.
|
|
110
|
+
Default: 2
|
|
111
|
+
|
|
112
|
+
In addition to `layout`, the `:fixed` format takes one more option:
|
|
113
|
+
|
|
114
|
+
* `truncate: [true|false]`
|
|
115
|
+
Whether to truncate string values that are longer than their column `:size` when writing.
|
|
116
|
+
When false, a string value that is too long raises `IOStreams::Errors::ValueTooLong`
|
|
117
|
+
instead of being truncated. Numeric values are never truncated.
|
|
118
|
+
Default: true
|
|
119
|
+
|
|
120
|
+
## Header options
|
|
121
|
+
|
|
122
|
+
When reading or writing records (`:hash`), the following options control the header row:
|
|
123
|
+
|
|
124
|
+
* `columns: [Array<String>]`
|
|
125
|
+
When reading, supplies the header columns for files that do not include a header row.
|
|
126
|
+
When writing, sets the columns to write, including their order. Keys not listed in
|
|
127
|
+
`columns` are ignored during writes.
|
|
128
|
+
|
|
129
|
+
* `cleanse_header: [true|false]`
|
|
130
|
+
Whether to cleanse the column names read from the header row.
|
|
131
|
+
Column names are stripped of leading and trailing whitespace, lowercased, and spaces
|
|
132
|
+
and dashes are converted to underscores, so the header `" First Name "` becomes `"first_name"`.
|
|
133
|
+
Default: true
|
|
134
|
+
|
|
135
|
+
* `allowed_columns: [Array<String>]`
|
|
136
|
+
List of columns to allow. Any other columns are ignored when `skip_unknown` is true,
|
|
137
|
+
otherwise an `IOStreams::Errors::InvalidHeader` exception is raised.
|
|
138
|
+
Default: nil (allow all columns)
|
|
139
|
+
|
|
140
|
+
* `required_columns: [Array<String>]`
|
|
141
|
+
List of columns that must be present, otherwise an exception is raised.
|
|
142
|
+
|
|
143
|
+
* `skip_unknown: [true|false]`
|
|
144
|
+
When true, any columns not present in `allowed_columns` are skipped entirely as if they
|
|
145
|
+
were not in the file at all. When false, an unknown column raises
|
|
146
|
+
`IOStreams::Errors::InvalidHeader`.
|
|
147
|
+
Default: true
|
|
148
|
+
|
|
149
|
+
By default, when reading records, `allowed_columns`, `required_columns` and `skip_unknown` only
|
|
150
|
+
apply to a header row read from the file, and only when `cleanse_header` is true. They are ignored
|
|
151
|
+
for JSON and `:hash` input, when `columns` are supplied, and with `cleanse_header: false`, and a
|
|
152
|
+
warning is logged when applying them would change the records read.
|
|
153
|
+
|
|
154
|
+
Since the format is usually inferred from the file name, renaming an uploaded file from `.csv` to
|
|
155
|
+
`.json` bypasses them. When they restrict which columns an upload can set, apply them to every input
|
|
156
|
+
in an initializer:
|
|
157
|
+
|
|
158
|
+
~~~ruby
|
|
159
|
+
IOStreams.enforce_column_restrictions = true
|
|
160
|
+
~~~
|
|
161
|
+
|
|
162
|
+
They then also apply to the supplied `columns`, to a header row read with `cleanse_header: false`,
|
|
163
|
+
and, for formats without a header row such as JSON, to the keys of each record. When either
|
|
164
|
+
`allowed_columns` or `required_columns` is set, JSON keys are cleansed the same way as a header row,
|
|
165
|
+
unless `cleanse_header: false` is supplied. This will be the default in v3.0.
|
|
166
|
+
|
|
167
|
+
Example, reading a headerless CSV file:
|
|
168
|
+
|
|
169
|
+
~~~ruby
|
|
170
|
+
path = IOStreams.path("no_header.csv")
|
|
171
|
+
path.each(:hash, columns: ["name", "address", "zip"]) do |record|
|
|
172
|
+
p record
|
|
173
|
+
end
|
|
174
|
+
~~~
|
|
175
|
+
|
|
176
|
+
Example, writing only specific columns in a fixed order:
|
|
177
|
+
|
|
178
|
+
~~~ruby
|
|
179
|
+
path = IOStreams.path("sample.csv")
|
|
180
|
+
path.writer(:hash, columns: ["name", "zip"]) do |io|
|
|
181
|
+
io << {"name" => "Jack Jones", "address" => "Somewhere", "zip" => 12345}
|
|
182
|
+
end
|
|
183
|
+
path.read
|
|
184
|
+
# => "name,zip\nJack Jones,12345\n"
|
|
185
|
+
~~~
|
|
186
|
+
|
|
187
|
+
Note: Column names are converted to strings, and the keys in the hashes being written may
|
|
188
|
+
be strings or symbols.
|