iostreams 1.11.0 → 2.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (108) hide show
  1. checksums.yaml +4 -4
  2. data/README.md +14 -13
  3. data/Rakefile +52 -0
  4. data/docs/CLAUDE.md +9 -0
  5. data/docs/config.md +157 -0
  6. data/docs/copy_files.md +75 -0
  7. data/docs/extensions.md +111 -0
  8. data/docs/formats.md +188 -0
  9. data/docs/index.md +388 -0
  10. data/docs/path.md +652 -0
  11. data/docs/pgp.md +436 -0
  12. data/docs/streams.md +337 -0
  13. data/docs/tutorial.md +483 -0
  14. data/docs/upgrading.md +217 -0
  15. data/lib/io_streams/builder.rb +71 -11
  16. data/lib/io_streams/bzip2/reader.rb +25 -2
  17. data/lib/io_streams/bzip2/writer.rb +26 -2
  18. data/lib/io_streams/encode/reader.rb +6 -2
  19. data/lib/io_streams/encode/writer.rb +9 -5
  20. data/lib/io_streams/errors.rb +4 -0
  21. data/lib/io_streams/gzip/reader.rb +5 -1
  22. data/lib/io_streams/gzip/writer.rb +11 -2
  23. data/lib/io_streams/io_streams.rb +156 -20
  24. data/lib/io_streams/line/reader.rb +9 -4
  25. data/lib/io_streams/line/writer.rb +1 -1
  26. data/lib/io_streams/path.rb +117 -8
  27. data/lib/io_streams/paths/file.rb +57 -11
  28. data/lib/io_streams/paths/http.rb +123 -9
  29. data/lib/io_streams/paths/matcher.rb +3 -3
  30. data/lib/io_streams/paths/s3.rb +69 -18
  31. data/lib/io_streams/paths/sftp/net_ssh.rb +104 -0
  32. data/lib/io_streams/paths/sftp.rb +103 -64
  33. data/lib/io_streams/pgp/reader.rb +63 -10
  34. data/lib/io_streams/pgp/writer.rb +111 -30
  35. data/lib/io_streams/pgp.rb +256 -71
  36. data/lib/io_streams/reader.rb +14 -5
  37. data/lib/io_streams/record/reader.rb +75 -6
  38. data/lib/io_streams/record/writer.rb +3 -4
  39. data/lib/io_streams/row/reader.rb +1 -1
  40. data/lib/io_streams/row/writer.rb +1 -1
  41. data/lib/io_streams/stream.rb +48 -37
  42. data/lib/io_streams/symmetric_encryption/reader.rb +6 -2
  43. data/lib/io_streams/symmetric_encryption/writer.rb +8 -4
  44. data/lib/io_streams/tabular/header.rb +49 -10
  45. data/lib/io_streams/tabular/parser/array.rb +0 -10
  46. data/lib/io_streams/tabular/parser/base.rb +10 -0
  47. data/lib/io_streams/tabular/parser/csv.rb +9 -36
  48. data/lib/io_streams/tabular/parser/fixed.rb +8 -6
  49. data/lib/io_streams/tabular/parser/psv.rb +6 -14
  50. data/lib/io_streams/tabular.rb +5 -10
  51. data/lib/io_streams/utils.rb +34 -2
  52. data/lib/io_streams/version.rb +1 -1
  53. data/lib/io_streams/writer.rb +16 -7
  54. data/lib/io_streams/xlsx/reader.rb +6 -2
  55. data/lib/io_streams/zip/reader.rb +4 -0
  56. data/lib/io_streams/zip/writer.rb +26 -10
  57. data/lib/iostreams.rb +0 -1
  58. metadata +46 -112
  59. data/lib/io_streams/deprecated.rb +0 -216
  60. data/lib/io_streams/tabular/utility/csv_row.rb +0 -105
  61. data/test/builder_test.rb +0 -311
  62. data/test/bzip2_reader_test.rb +0 -27
  63. data/test/bzip2_writer_test.rb +0 -56
  64. data/test/deprecated_test.rb +0 -121
  65. data/test/encode_reader_test.rb +0 -51
  66. data/test/encode_writer_test.rb +0 -90
  67. data/test/files/embedded_lines_test.csv +0 -7
  68. data/test/files/multiple_files.zip +0 -0
  69. data/test/files/spreadsheet.xlsx +0 -0
  70. data/test/files/test.csv +0 -4
  71. data/test/files/test.json +0 -3
  72. data/test/files/test.psv +0 -4
  73. data/test/files/text file.txt +0 -3
  74. data/test/files/text.txt +0 -3
  75. data/test/files/text.txt.bz2 +0 -0
  76. data/test/files/text.txt.gz +0 -0
  77. data/test/files/text.txt.gz.zip +0 -0
  78. data/test/files/text.zip +0 -0
  79. data/test/files/text.zip.gz +0 -0
  80. data/test/files/unclosed_quote_large_test.csv +0 -1658
  81. data/test/files/unclosed_quote_test.csv +0 -4
  82. data/test/files/unclosed_quote_test2.csv +0 -3
  83. data/test/gzip_reader_test.rb +0 -27
  84. data/test/gzip_writer_test.rb +0 -52
  85. data/test/io_streams_test.rb +0 -132
  86. data/test/line_reader_test.rb +0 -325
  87. data/test/line_writer_test.rb +0 -59
  88. data/test/minimal_file_reader.rb +0 -25
  89. data/test/path_test.rb +0 -55
  90. data/test/paths/file_test.rb +0 -213
  91. data/test/paths/http_test.rb +0 -34
  92. data/test/paths/matcher_test.rb +0 -120
  93. data/test/paths/s3_test.rb +0 -220
  94. data/test/paths/sftp_test.rb +0 -106
  95. data/test/pgp_reader_test.rb +0 -46
  96. data/test/pgp_test.rb +0 -267
  97. data/test/pgp_writer_test.rb +0 -130
  98. data/test/record_reader_test.rb +0 -60
  99. data/test/record_writer_test.rb +0 -82
  100. data/test/row_reader_test.rb +0 -35
  101. data/test/row_writer_test.rb +0 -56
  102. data/test/stream_test.rb +0 -577
  103. data/test/tabular_test.rb +0 -338
  104. data/test/test_helper.rb +0 -40
  105. data/test/utils_test.rb +0 -20
  106. data/test/xlsx_reader_test.rb +0 -37
  107. data/test/zip_reader_test.rb +0 -53
  108. data/test/zip_writer_test.rb +0 -48
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: dfb827d5403c211fcdcbefbdb79e8d28f0a29f8d5982cc5de0eb832cf72d60bc
4
- data.tar.gz: 710894a9919d7d3935867f67dd7a7bdfae041c29c3c5700dd80e0cccd7b7777c
3
+ metadata.gz: d74053d56262767ee38729e8fc3810416b340ec7e1a318c37a95579f392a92d6
4
+ data.tar.gz: 7e6e2fa0088656a52a8fd9c45c78b96ed8b2c324e9f9379a091fcfdd3939dde9
5
5
  SHA512:
6
- metadata.gz: d246b2d041bbedf6de44340c6db298cfabf71ad14180c7a4743c83c3a60deca81b4b1f996ce6cf27fc8674e708f6276432929870ab7b4160669d96b4aeae1f53
7
- data.tar.gz: 3abfbfba9845a6fe135612a5048f8f3f2cf961907898f1ae4759cb2fd5093460cff0093536764f9b5c7600c313dc1b2043c7795a6fde660be093daa38d9f1d86
6
+ metadata.gz: a492f9f89787eae3675fcf0aef1a1fbc9f67e461a0c1d0c9422888fa4cc66277c55d39248041b79a50de1b0d321a51a013445379de93404b0f85105d48914962
7
+ data.tar.gz: 2e6303a89a55bbbda4a9c73a19bdd6780823633afa4391d2fe3071ba8ca85e360a08214e76ed2e869182d8ed94388ed0454769b5d6dc41e540abf704d1c8a794
data/README.md CHANGED
@@ -1,8 +1,9 @@
1
1
  # IOStreams
2
2
  [![Gem Version](https://img.shields.io/gem/v/iostreams.svg)](https://rubygems.org/gems/iostreams) [![Downloads](https://img.shields.io/gem/dt/iostreams.svg)](https://rubygems.org/gems/iostreams) [![License](https://img.shields.io/badge/license-Apache%202.0-brightgreen.svg)](http://opensource.org/licenses/Apache-2.0) ![](https://img.shields.io/badge/status-Production%20Ready-blue.svg)
3
3
 
4
- IOStreams is an incredibly powerful streaming library that makes changes to file formats, compression, encryption,
5
- or storage mechanism transparent to the application.
4
+ IOStreams is a streaming library for Ruby that makes compression, encryption, file format, and storage
5
+ location transparent to your code. Read and write files of any size, one block at a time, whether they
6
+ are gzip, zip, or PGP encrypted, and whether they live on local disk, AWS S3, SFTP, or are fetched over HTTP.
6
7
 
7
8
  ## Project Status
8
9
 
@@ -10,26 +11,26 @@ Production Ready, heavily used in production environments, many as part of Rocke
10
11
 
11
12
  ## Documentation
12
13
 
13
- Start with the [IOStreams tutorial](https://iostreams.rocketjob.io/tutorial) to get a great introduction to IOStreams.
14
+ Start with the [IOStreams tutorial](https://iostreams.reidmorrison.com/tutorial) to get a great introduction to IOStreams.
14
15
 
15
- Next, checkout the remaining [IOStreams documentation](https://iostreams.rocketjob.io/)
16
+ Next, checkout the remaining [IOStreams documentation](https://iostreams.reidmorrison.com/)
16
17
 
17
- ## Upgrading to v1.6
18
+ See the [CHANGELOG](CHANGELOG.md) for the release history and notable changes.
18
19
 
19
- The old, deprecated api's are no longer loaded by default with v1.6. To add back the deprecated api support, add
20
- the following line to your code:
20
+ ## Upgrading
21
21
 
22
- ~~~ruby
23
- IOStreams.include(IOStreams::Deprecated)
24
- ~~~
25
-
26
- It is important to move any of the old deprecated apis over to the new api, since they will be removed in a future
27
- release.
22
+ See [Upgrading IOStreams](https://iostreams.reidmorrison.com/upgrading) for the changes that may need
23
+ updates to your application, and the security settings to review, when upgrading.
28
24
 
29
25
  ## Versioning
30
26
 
31
27
  This project adheres to [Semantic Versioning](http://semver.org/).
32
28
 
29
+ ## Contributing
30
+
31
+ Contributions are welcome. See [CONTRIBUTING.md](CONTRIBUTING.md) for guidelines on documentation
32
+ updates, code changes, the project architecture, and the code of conduct.
33
+
33
34
  ## Author
34
35
 
35
36
  [Reid Morrison](https://github.com/reidmorrison)
data/Rakefile CHANGED
@@ -1,10 +1,12 @@
1
1
  require "rake/testtask"
2
2
  require_relative "lib/io_streams/version"
3
3
 
4
+ desc "Build the iostreams gem"
4
5
  task :gem do
5
6
  system "gem build iostreams.gemspec"
6
7
  end
7
8
 
9
+ desc "Build and publish the iostreams gem, then tag and push the release"
8
10
  task publish: :gem do
9
11
  system "git tag -a v#{IOStreams::VERSION} -m 'Tagging #{IOStreams::VERSION}'"
10
12
  system "git push --tags"
@@ -12,6 +14,56 @@ task publish: :gem do
12
14
  system "rm iostreams-#{IOStreams::VERSION}.gem"
13
15
  end
14
16
 
17
+ desc "Start an IRB console with the gem loaded"
18
+ task :console do
19
+ exec "irb -I lib -r iostreams"
20
+ end
21
+
22
+ desc "Generate docs/llms-full.txt from docs/llms.txt and the doc pages it links to"
23
+ task :llms_full do
24
+ require "uri"
25
+
26
+ docs_dir = File.join(__dir__, "docs")
27
+ llms_path = File.join(docs_dir, "llms.txt")
28
+ out_path = File.join(docs_dir, "llms-full.txt")
29
+
30
+ llms_txt = File.read(llms_path)
31
+ header = llms_txt[/\A.*?(?=\n## Docs)/m]&.strip
32
+ raise "Could not find intro text before '## Docs' in #{llms_path}" unless header
33
+
34
+ docs_section = llms_txt[/^## Docs\n(.*?)(?=\n## |\z)/m, 1]
35
+ raise "Could not find '## Docs' section in #{llms_path}" unless docs_section
36
+
37
+ sections = docs_section.each_line.filter_map do |line|
38
+ next unless line =~ /^- \[(?<title>[^\]]+)\]\((?<url>[^)]+)\)/
39
+
40
+ title = $~[:title]
41
+ path = URI.parse($~[:url]).path.sub(%r{\A/}, "")
42
+ path = "index" if path.empty?
43
+ file = File.join(docs_dir, "#{path}.md")
44
+ raise "Missing doc file for #{$~[:url]}: #{file}" unless File.exist?(file)
45
+
46
+ raw = File.read(file)
47
+
48
+ # The page heading lives in front matter, which the shared docs theme
49
+ # renders as the page's h1. Front matter is stripped below, so lift the
50
+ # heading back out and re-emit it. `heading` wins over `title` where a page
51
+ # sets both, which is the same precedence the theme uses. Falls back to the
52
+ # link text in llms.txt for a page that sets neither.
53
+ front_matter = raw[/\A---\n(.*?)\n---\n/m, 1].to_s
54
+ heading = %w[heading title].
55
+ filter_map { |key| front_matter[/^#{key}:[ \t]*(.+)$/, 1] }.
56
+ first.to_s.strip.delete_prefix('"').delete_suffix('"')
57
+ heading = title if heading.empty?
58
+
59
+ body = raw.sub(/\A---\n.*?\n---\n/m, "").strip
60
+ "# #{heading}\n\n#{body}"
61
+ end
62
+
63
+ File.write(out_path, "#{header}\n\n#{sections.join("\n\n---\n\n")}\n")
64
+ puts "Wrote #{out_path}"
65
+ end
66
+
15
67
  Rake::TestTask.new(:test) do |t|
16
68
  t.pattern = "test/**/*_test.rb"
17
69
  t.verbose = true
data/docs/CLAUDE.md ADDED
@@ -0,0 +1,9 @@
1
+ # Documentation site
2
+
3
+ User-facing documentation is a Jekyll site under [docs/](docs/), published to iostreams.reidmorrison.com. The markdown pages are what matter when reading or updating documentation.
4
+
5
+ **The look and feel is not in this repo.** `docs/_config.yml` sets `remote_theme: reidmorrison/rm-docs-theme@v1`, and the layout, stylesheet, sidebar and syntax highlighting all come from there. This repo holds only its content: the markdown pages and `docs/images`. **Do not add a `docs/_layouts`, `docs/stylesheets` or `docs/javascripts` directory**; they were deleted deliberately, because six gem repos each carried a near-identical copy of the same theme and the copies had drifted. A styling change belongs in `rm-docs-theme`, where it reaches every doc site at once. `v1` is a moving major tag, so theme fixes arrive on the next build; breaking changes go to `v2` and are opted into by editing the pin. `jekyll-remote-theme` must stay in `plugins`: GitHub Pages enables it on its own, but a local build does not, and without it every page silently renders with no layout. Preview against a local theme checkout with `~/src/rm-docs-theme/bin/preview ~/src/iostreams/docs`.
6
+
7
+ **A page's title lives in its front matter**, not in a `#` heading at the top of the markdown; the theme renders it as the page's `h1`. `title` is the browser title and the default heading, `heading` overrides the h1 where the two should differ, and `description` is the page's meta description. `index.md` sets `heading` only, so the home page keeps the tuned SEO `<title>` from `_config.yml`. Adding or renaming a page means editing the `nav` block in `docs/_config.yml` and `docs/llms.txt`.
8
+
9
+ The site also serves two files for AI assistants: [docs/llms.txt](docs/llms.txt), a hand-maintained index of the docs pages, and `docs/llms-full.txt`, all pages concatenated, regenerated with `bundle exec rake llms_full`. **After editing any `docs/*.md` page, re-run `bundle exec rake llms_full`** and commit the result; never edit `llms-full.txt` by hand. That task reads the page heading out of the front matter, so a page that sets neither `heading` nor `title` falls back to its link text in `llms.txt`.
data/docs/config.md ADDED
@@ -0,0 +1,157 @@
1
+ ---
2
+ layout: default
3
+ title: Configuring IOStreams
4
+ description: >-
5
+ Named roots via IOStreams.add_root so the same code targets different storage
6
+ per environment, plus the temp directory and logger settings.
7
+ ---
8
+
9
+ ## add_root
10
+
11
+ Roots allow paths to reference a particular root directory, so that all path names are appended to that root.
12
+ Their primary purpose is to allow the exact same code to run in production and development, yet use completely
13
+ different data sources in each. For example, in production a root can point to an S3 bucket, while in
14
+ development it points to the local file system.
15
+
16
+ Roots are configured via an initializer at startup. `IOStreams.join` then joins the supplied path
17
+ elements onto the named root, using the `:default` root whenever a root is not supplied.
18
+
19
+ Set the default root for this environment in an initializer:
20
+ ~~~ruby
21
+ IOStreams.add_root(:default, "/var/my_app/files")
22
+ ~~~
23
+
24
+ Now the default root path is available:
25
+ ~~~ruby
26
+ IOStreams.root
27
+ # => #<IOStreams::Paths::File:/var/my_app/files pipeline={}>
28
+
29
+ IOStreams.root.to_s
30
+ # => "/var/my_app/files"
31
+ ~~~
32
+
33
+ Comparing the final path using `path` and then `join` that uses a root path:
34
+ ~~~ruby
35
+ IOStreams.path("/var/my_app/files", "my_test_file.txt").to_s
36
+ # => "/var/my_app/files/my_test_file.txt"
37
+
38
+ IOStreams.join("my_test_file.txt").to_s
39
+ # => "/var/my_app/files/my_test_file.txt"
40
+ ~~~
41
+
42
+
43
+ Using `path`:
44
+ ~~~ruby
45
+ IOStreams.path("/var/my_app/files", "my_test_file.txt").write("Hello World")
46
+ ~~~
47
+
48
+ With the default root path configured the above code can be simplified by using `join` since it resolves to the same path.
49
+ ~~~ruby
50
+ IOStreams.join("my_test_file.txt").write("Hello World")
51
+ ~~~
52
+
53
+ Multiple roots can be setup, for example one for input files, another for output files, another for
54
+ reports, etc. During development the roots can all point to a common location, while in production
55
+ they could be completely different S3 buckets.
56
+
57
+ For example add special paths for `downloads` and `uploads`.
58
+ ~~~ruby
59
+ IOStreams.add_root(:downloads, "/var/my_app/downloads")
60
+ IOStreams.add_root(:uploads, "/var/my_app/uploads")
61
+ ~~~
62
+
63
+ An example that writes a file into the `/var/my_app/downloads` directory:
64
+ ~~~ruby
65
+ IOStreams.join("my_test_file.txt", root: :downloads).write("Hello World")
66
+ ~~~
67
+
68
+ The other benefit is that the root paths used in an application are externalized from the code base. That way the
69
+ roots can be changed to different locations depending on the environment.
70
+
71
+ We can also change the storage mechanism by changing the root:
72
+ ~~~ruby
73
+ IOStreams.add_root(:downloads, "s3://my-app-bucket-name/downloads")
74
+ IOStreams.add_root(:uploads, "s3://my-app-bucket-name/uploads")
75
+ ~~~
76
+
77
+ Now the application will write to S3 and the code does not change at all.
78
+ ~~~ruby
79
+ IOStreams.join("my_test_file.txt", root: :downloads).write("Hello World")
80
+ ~~~
81
+
82
+ To use or query a configured root path:
83
+ ~~~ruby
84
+ IOStreams.root(:downloads).to_s
85
+ # => "s3://my-app-bucket-name/downloads"
86
+ ~~~
87
+
88
+ ## temp_dir
89
+
90
+ When working with large files the standard temp file system location can be too small to handle downloading large
91
+ files. For example to decrypt a pgp file from S3, because GnuPG is not streaming capable and only operates on local files.
92
+
93
+ By default IOStreams looks up the location to store temp files in the following order:
94
+ * `ENV['TMPDIR']`
95
+ * `ENV['TMP']`
96
+ * `ENV['TEMP']`
97
+ * `Etc.systmpdir`
98
+ * `/tmp` (if it exists)
99
+ * Otherwise `.`
100
+
101
+ To explicity set the temp file location the following config option can be used:
102
+
103
+ ~~~ruby
104
+ IOStreams.temp_dir = "/var/really_big_temp"
105
+ ~~~
106
+
107
+ ### temp_file
108
+
109
+ To work with a temporary file directly, `IOStreams.temp_file` yields a path inside `temp_dir`
110
+ and deletes the file when the block completes:
111
+
112
+ ~~~ruby
113
+ IOStreams.temp_file("export", ".csv") do |path|
114
+ path.write("Hello World")
115
+ # ... use the temp file ...
116
+ end
117
+ # The temp file has been deleted.
118
+ ~~~
119
+
120
+ The first argument is a base file name to include in the generated temp file name, and the
121
+ optional second argument is the file extension.
122
+
123
+ ## logger
124
+
125
+ IOStreams can log debug information, such as the external commands it runs for PGP and SFTP.
126
+
127
+ When [Semantic Logger](https://logger.reidmorrison.com) is loaded it is detected automatically, and IOStreams
128
+ logs to it without any additional configuration.
129
+
130
+ To use a different logger, or to log when Semantic Logger is not present, assign any logger that
131
+ responds to the standard logging methods:
132
+
133
+ ~~~ruby
134
+ require "logger"
135
+ IOStreams.logger = Logger.new($stdout)
136
+ ~~~
137
+
138
+ To disable logging entirely, set the logger to `nil`:
139
+
140
+ ~~~ruby
141
+ IOStreams.logger = nil
142
+ ~~~
143
+
144
+ IOStreams also logs warnings for behavior that will change in the next major version. See
145
+ [Coming in v3.0](upgrading#coming-in-v30).
146
+
147
+ ## enforce_column_restrictions
148
+
149
+ Applies `allowed_columns`, `required_columns` and `skip_unknown` to every input when reading
150
+ records, including JSON records, so that renaming an uploaded file from `.csv` to `.json` cannot
151
+ bypass them. Set it in an initializer:
152
+
153
+ ~~~ruby
154
+ IOStreams.enforce_column_restrictions = true
155
+ ~~~
156
+
157
+ Default: false. It will default to true in v3.0. See [Header options](formats#header-options).
@@ -0,0 +1,75 @@
1
+ ---
2
+ layout: default
3
+ title: Copying Between Files
4
+ description: >-
5
+ Using IOStreams.copy to move a file between storage locations and to compress,
6
+ decompress, encrypt or decrypt it on the way.
7
+ ---
8
+
9
+ File copying can be used to:
10
+ * copy from one storage location to another.
11
+ * create a decrypted / encrypted copy of an existing file.
12
+ * create a decompressed / compressed copy of an existing file.
13
+
14
+ ## Examples
15
+
16
+ Decompress `example.csv.gz` into `example.csv`:
17
+
18
+ ~~~ruby
19
+ source = IOStreams.path("example.csv.gz")
20
+ IOStreams.path("example.csv").copy_from(source)
21
+ ~~~
22
+
23
+ Decrypt a file encrypted with Symmetric Encryption:
24
+
25
+ ~~~ruby
26
+ source = IOStreams.path("example.csv.enc")
27
+ IOStreams.path("example.csv").copy_from(source)
28
+ ~~~
29
+
30
+ Encrypt a file using PGP encryption so that it can only be read by `receiver@example.org`.
31
+
32
+ ~~~ruby
33
+ source = IOStreams.path("example.csv")
34
+ target = IOStreams.path("example.csv.pgp")
35
+ target.option(:pgp, recipient: "receiver@example.org")
36
+ target.copy_from(source)
37
+ ~~~
38
+
39
+ When the file name does not have file extensions that would allow IOStreams to infer what streams to apply,
40
+ the streams can be explicitly set using `stream`.
41
+
42
+ In this example, the file `CUSTOMER_DATA` has no extensions, so `stream(:enc)` tells IOStreams
43
+ that its contents were encrypted with Symmetric Encryption. The decrypted contents are then
44
+ PGP encrypted and written to `xyz.csv.pgp` using the pgp key for `receiver@example.org`.
45
+
46
+ ~~~ruby
47
+ input = IOStreams.path("CUSTOMER_DATA").stream(:enc)
48
+ IOStreams.path("xyz.csv.pgp").option(:pgp, recipient: "receiver@example.org").copy_from(input)
49
+ ~~~
50
+
51
+ To copy a file _without_ performing any conversions (ignore file extensions), set `convert` to `false`:
52
+
53
+ ~~~ruby
54
+ input = IOStreams.path("sample.json.zip")
55
+ IOStreams.path("sample.copy").copy_from(input, convert: false)
56
+ ~~~
57
+
58
+ Custom stream conversions can be applied to both the source and the target in a single copy.
59
+ Here the source is read as binary and the target is PGP encrypted:
60
+
61
+ ~~~ruby
62
+ source = IOStreams.path("source_file").stream(:encode, encoding: "BINARY")
63
+ IOStreams.path("target_file.pgp").option(:pgp, passphrase: "hello").copy_from(source)
64
+ ~~~
65
+
66
+ To convert the contents row by row, or record by record, during the copy, supply `mode`.
67
+ For example, copy a CSV file into JSON, parsing and rendering each record:
68
+
69
+ ~~~ruby
70
+ source = IOStreams.path("source_file.csv")
71
+ IOStreams.path("target_file.json").copy_from(source, mode: :hash)
72
+ ~~~
73
+
74
+ Notes:
75
+ * `mode` accepts `:line`, `:array`, or `:hash`, and only applies when `convert` is `true`.
@@ -0,0 +1,111 @@
1
+ ---
2
+ layout: default
3
+ title: File Extensions
4
+ description: >-
5
+ How the extensions in a file name, such as .csv.gz.pgp, decide which streams
6
+ IOStreams applies when reading or writing, and how to register your own.
7
+ ---
8
+
9
+ IOStreams uses the extensions in the file name to determine which streams to apply when
10
+ reading or writing a file. Multiple extensions are applied in order, so `sample.csv.gz.pgp`
11
+ is first decrypted with PGP and then decompressed with GZip when read.
12
+
13
+ Supported extensions:
14
+
15
+ | Extension | Stream | Read | Write | Required gem / program |
16
+ |:-----------------|:---------------------|:-----|:------|:----------------------------------|
17
+ | `.bz2` | BZip2 | Yes | Yes | `bzip2-ffi` |
18
+ | `.enc` | Symmetric Encryption | Yes | Yes | `symmetric-encryption` |
19
+ | `.gz`, `.gzip` | GZip | Yes | Yes | None (Ruby standard library) |
20
+ | `.zip` | Zip | Yes | Yes | `rubyzip` (read), `zip_kit` (write). On JRuby the built-in Java zip support is used for reading. |
21
+ | `.pgp`, `.gpg` | PGP | Yes | Yes | GnuPG command line program (`gpg`) |
22
+ | `.xlsx`, `.xlsm` | Excel Spreadsheet | Yes | No | `creek` |
23
+
24
+ The gems above are soft dependencies: IOStreams does not require them for installation,
25
+ they only need to be added to the `Gemfile` when the corresponding extension is used.
26
+
27
+ ## Compression options
28
+
29
+ GZip accepts a compression `level` when writing, from `0` (no compression) to `9` (best compression):
30
+
31
+ ~~~ruby
32
+ IOStreams.path("sample.csv.gz").option(:gz, level: 9).write(data)
33
+ ~~~
34
+
35
+ BZip2 passes its options through to `bzip2-ffi`: `block_size` (`1` to `9`) and `work_factor` (`0` to `250`)
36
+ when writing, and `small` and `first_only` when reading:
37
+
38
+ ~~~ruby
39
+ IOStreams.path("sample.csv.bz2").option(:bz2, block_size: 9).write(data)
40
+ IOStreams.path("sample.csv.bz2").option(:bz2, small: true).read
41
+ ~~~
42
+
43
+ Options are strict, so an option that a stream does not accept raises an `ArgumentError`.
44
+ See [Streams](streams#reading-and-writing-need-separate-options). The exception is BZip2, which
45
+ ignores any other option and logs a warning. In v3.0 it will raise an `ArgumentError` too.
46
+
47
+ ## Reading an Excel Spreadsheet
48
+
49
+ Each row in the spreadsheet is converted into a CSV line, so the regular `:line`, `:array`,
50
+ and `:hash` modes apply:
51
+
52
+ ~~~ruby
53
+ IOStreams.path("spreadsheet.xlsx").each(:hash) do |record|
54
+ p record
55
+ end
56
+ ~~~
57
+
58
+ Notes:
59
+ * Since the underlying `creek` gem operates on files, when reading from a stream (for example S3 or HTTP)
60
+ the contents are first downloaded into a temp file.
61
+ * Writing xlsx files is not supported.
62
+
63
+ ## Character encoding
64
+
65
+ The special `:encode` stream converts the character encoding of the data being read or written.
66
+ It is applied with `option` or `stream` rather than a file name extension:
67
+
68
+ ~~~ruby
69
+ IOStreams.path("sample.csv.gz").
70
+ option(:encode, encoding: "UTF-8", cleaner: :printable, replace: "").
71
+ each do |line|
72
+ puts line
73
+ end
74
+ ~~~
75
+
76
+ Options:
77
+
78
+ * `encoding: [String|Encoding]`
79
+ The target encoding, for example `"UTF-8"`, `"US-ASCII"`, or `"ASCII-8BIT"`.
80
+ Default: `"UTF-8"`
81
+
82
+ * `replace: [String]`
83
+ The character to replace with when a character cannot be converted to the target encoding.
84
+ Default: nil (raise `Encoding::UndefinedConversionError` on invalid characters)
85
+
86
+ * `cleaner: [nil|Symbol|Proc]`
87
+ Cleanse the data. Built-in rules:
88
+ * `:printable` removes all non-printable characters except `\r` and `\n`.
89
+ * `:replace_non_printable` replaces all non-printable characters except `\r` and `\n`
90
+ with the `replace` value, or an empty string when `replace` is nil.
91
+ A Proc can also be supplied to perform custom cleansing; it is called with the data
92
+ and the `replace` value after every read or write.
93
+ Default: nil
94
+
95
+ ## Registering a custom extension
96
+
97
+ To add a new extension, supply its reader and writer classes. Both must implement `.open`
98
+ that yields a stream implementing `#read` or `#write` respectively. See any of the streams
99
+ under `lib/io_streams` for examples.
100
+
101
+ ~~~ruby
102
+ IOStreams.register_extension(:xls, MyXls::Reader, MyXls::Writer)
103
+ ~~~
104
+
105
+ Similarly, to support a new storage location, supply a Path class for its URI scheme.
106
+ See [IOStreams::Paths::S3](https://github.com/reidmorrison/iostreams/blob/main/lib/io_streams/paths/s3.rb)
107
+ for an example of what is required.
108
+
109
+ ~~~ruby
110
+ IOStreams.register_scheme(:gcs, MyGoogleCloudStoragePath)
111
+ ~~~
data/docs/formats.md ADDED
@@ -0,0 +1,188 @@
1
+ ---
2
+ layout: default
3
+ title: File Formats
4
+ description: >-
5
+ Converting rows and records to and from CSV, PSV, JSON and fixed width files,
6
+ including format inference, format options and header handling.
7
+ ---
8
+
9
+ When reading or writing rows (`:array`) or records (`:hash`), IOStreams converts each line
10
+ to or from the file's tabular format. The following formats are supported:
11
+
12
+ * `:csv` Comma Separated Values
13
+ * `:psv` Pipe Separated Values
14
+ * `:json` One JSON document per line
15
+ * `:fixed` Fixed width columns
16
+ * `:array` Each line is already an array of values
17
+ * `:hash` Each line is already a hash
18
+
19
+ PSV has no way to escape values, so when writing PSV, a `|` within a value is replaced with `:`
20
+ and a line break with a space, so that a value cannot add columns or records.
21
+
22
+ ## Format inference
23
+
24
+ The format is inferred from the file name when it contains a recognized extension:
25
+
26
+ ~~~ruby
27
+ IOStreams.path("sample.csv").each(:hash) { |record| p record }
28
+ IOStreams.path("sample.json").each(:hash) { |record| p record }
29
+ IOStreams.path("sample.psv").each(:hash) { |record| p record }
30
+ ~~~
31
+
32
+ The format extension can appear anywhere in the file name, so `sample.csv.gz` and
33
+ `sample.json.pgp` are recognized as CSV and JSON respectively.
34
+
35
+ When the file name does not contain a recognized format extension, the format defaults to `:csv`.
36
+
37
+ ## Specifying the format
38
+
39
+ When the file name cannot be used to infer the format, set it explicitly with `format`:
40
+
41
+ ~~~ruby
42
+ path = IOStreams.path("sample_data")
43
+ path.format(:json)
44
+ path.each(:hash) { |record| p record }
45
+ ~~~
46
+
47
+ `format` can be chained with the other path methods:
48
+
49
+ ~~~ruby
50
+ IOStreams.path("sample_data").format(:json).each(:hash) { |record| p record }
51
+ ~~~
52
+
53
+ ## Format options
54
+
55
+ Format specific options are supplied with `format_options`. They are passed to the parser
56
+ for the chosen format. The `:fixed` format requires its file layout to be supplied this way,
57
+ as shown in the next section. The other formats do not currently take any options.
58
+
59
+ ## Fixed width files
60
+
61
+ Fixed width files have no delimiters; each column is identified by its position within the line.
62
+ Since the layout cannot be inferred from the file, supply it using `format_options`:
63
+
64
+ ~~~ruby
65
+ path = IOStreams.path("sample_data")
66
+ path.format(:fixed)
67
+ path.format_options(
68
+ layout: [
69
+ {size: 23, key: "name"},
70
+ {size: 40, key: "address"},
71
+ {size: 5, key: "zip"}
72
+ ]
73
+ )
74
+ path.each(:hash) { |record| p record }
75
+ ~~~
76
+
77
+ Writing a fixed width file uses the same layout to render each record:
78
+
79
+ ~~~ruby
80
+ path = IOStreams.path("sample_data")
81
+ path.format(:fixed)
82
+ path.format_options(
83
+ layout: [
84
+ {size: 23, key: "name"},
85
+ {size: 40, key: "address"},
86
+ {size: 5, key: "zip"}
87
+ ]
88
+ )
89
+ path.writer(:hash) do |io|
90
+ io << {"name" => "Jack Jones", "address" => "Somewhere", "zip" => 12345}
91
+ end
92
+ ~~~
93
+
94
+ Note: The keys in the hashes being written must match the layout `:key` values exactly,
95
+ including whether they are strings or symbols.
96
+
97
+ Layout column definitions:
98
+
99
+ * `:size` The number of characters this column occupies.
100
+ The last column may use a size of `:remainder` to take the rest of the line as its value.
101
+ * `:key` The name for this column. Leave out the key to ignore the column during parsing,
102
+ and to space fill when rendering.
103
+ * `:type` `:string` (default), `:integer`, or `:float`.
104
+ Strings are left justified and space padded, numbers are right justified and zero padded.
105
+ When writing, line breaks within a string are replaced with a space so that a value cannot
106
+ add records.
107
+ Raises `IOStreams::Errors::ValueTooLong` when an `:integer` or `:float` value cannot be
108
+ rendered in `size` characters.
109
+ * `:decimals` For `:float` columns, the number of decimal places to render.
110
+ Default: 2
111
+
112
+ In addition to `layout`, the `:fixed` format takes one more option:
113
+
114
+ * `truncate: [true|false]`
115
+ Whether to truncate string values that are longer than their column `:size` when writing.
116
+ When false, a string value that is too long raises `IOStreams::Errors::ValueTooLong`
117
+ instead of being truncated. Numeric values are never truncated.
118
+ Default: true
119
+
120
+ ## Header options
121
+
122
+ When reading or writing records (`:hash`), the following options control the header row:
123
+
124
+ * `columns: [Array<String>]`
125
+ When reading, supplies the header columns for files that do not include a header row.
126
+ When writing, sets the columns to write, including their order. Keys not listed in
127
+ `columns` are ignored during writes.
128
+
129
+ * `cleanse_header: [true|false]`
130
+ Whether to cleanse the column names read from the header row.
131
+ Column names are stripped of leading and trailing whitespace, lowercased, and spaces
132
+ and dashes are converted to underscores, so the header `" First Name "` becomes `"first_name"`.
133
+ Default: true
134
+
135
+ * `allowed_columns: [Array<String>]`
136
+ List of columns to allow. Any other columns are ignored when `skip_unknown` is true,
137
+ otherwise an `IOStreams::Errors::InvalidHeader` exception is raised.
138
+ Default: nil (allow all columns)
139
+
140
+ * `required_columns: [Array<String>]`
141
+ List of columns that must be present, otherwise an exception is raised.
142
+
143
+ * `skip_unknown: [true|false]`
144
+ When true, any columns not present in `allowed_columns` are skipped entirely as if they
145
+ were not in the file at all. When false, an unknown column raises
146
+ `IOStreams::Errors::InvalidHeader`.
147
+ Default: true
148
+
149
+ By default, when reading records, `allowed_columns`, `required_columns` and `skip_unknown` only
150
+ apply to a header row read from the file, and only when `cleanse_header` is true. They are ignored
151
+ for JSON and `:hash` input, when `columns` are supplied, and with `cleanse_header: false`, and a
152
+ warning is logged when applying them would change the records read.
153
+
154
+ Since the format is usually inferred from the file name, renaming an uploaded file from `.csv` to
155
+ `.json` bypasses them. When they restrict which columns an upload can set, apply them to every input
156
+ in an initializer:
157
+
158
+ ~~~ruby
159
+ IOStreams.enforce_column_restrictions = true
160
+ ~~~
161
+
162
+ They then also apply to the supplied `columns`, to a header row read with `cleanse_header: false`,
163
+ and, for formats without a header row such as JSON, to the keys of each record. When either
164
+ `allowed_columns` or `required_columns` is set, JSON keys are cleansed the same way as a header row,
165
+ unless `cleanse_header: false` is supplied. This will be the default in v3.0.
166
+
167
+ Example, reading a headerless CSV file:
168
+
169
+ ~~~ruby
170
+ path = IOStreams.path("no_header.csv")
171
+ path.each(:hash, columns: ["name", "address", "zip"]) do |record|
172
+ p record
173
+ end
174
+ ~~~
175
+
176
+ Example, writing only specific columns in a fixed order:
177
+
178
+ ~~~ruby
179
+ path = IOStreams.path("sample.csv")
180
+ path.writer(:hash, columns: ["name", "zip"]) do |io|
181
+ io << {"name" => "Jack Jones", "address" => "Somewhere", "zip" => 12345}
182
+ end
183
+ path.read
184
+ # => "name,zip\nJack Jones,12345\n"
185
+ ~~~
186
+
187
+ Note: Column names are converted to strings, and the keys in the hashes being written may
188
+ be strings or symbols.