iostreams 1.11.0 → 2.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (108) hide show
  1. checksums.yaml +4 -4
  2. data/README.md +14 -13
  3. data/Rakefile +52 -0
  4. data/docs/CLAUDE.md +9 -0
  5. data/docs/config.md +157 -0
  6. data/docs/copy_files.md +75 -0
  7. data/docs/extensions.md +111 -0
  8. data/docs/formats.md +188 -0
  9. data/docs/index.md +388 -0
  10. data/docs/path.md +652 -0
  11. data/docs/pgp.md +436 -0
  12. data/docs/streams.md +337 -0
  13. data/docs/tutorial.md +483 -0
  14. data/docs/upgrading.md +217 -0
  15. data/lib/io_streams/builder.rb +71 -11
  16. data/lib/io_streams/bzip2/reader.rb +25 -2
  17. data/lib/io_streams/bzip2/writer.rb +26 -2
  18. data/lib/io_streams/encode/reader.rb +6 -2
  19. data/lib/io_streams/encode/writer.rb +9 -5
  20. data/lib/io_streams/errors.rb +4 -0
  21. data/lib/io_streams/gzip/reader.rb +5 -1
  22. data/lib/io_streams/gzip/writer.rb +11 -2
  23. data/lib/io_streams/io_streams.rb +156 -20
  24. data/lib/io_streams/line/reader.rb +9 -4
  25. data/lib/io_streams/line/writer.rb +1 -1
  26. data/lib/io_streams/path.rb +117 -8
  27. data/lib/io_streams/paths/file.rb +57 -11
  28. data/lib/io_streams/paths/http.rb +123 -9
  29. data/lib/io_streams/paths/matcher.rb +3 -3
  30. data/lib/io_streams/paths/s3.rb +69 -18
  31. data/lib/io_streams/paths/sftp/net_ssh.rb +104 -0
  32. data/lib/io_streams/paths/sftp.rb +103 -64
  33. data/lib/io_streams/pgp/reader.rb +63 -10
  34. data/lib/io_streams/pgp/writer.rb +111 -30
  35. data/lib/io_streams/pgp.rb +256 -71
  36. data/lib/io_streams/reader.rb +14 -5
  37. data/lib/io_streams/record/reader.rb +75 -6
  38. data/lib/io_streams/record/writer.rb +3 -4
  39. data/lib/io_streams/row/reader.rb +1 -1
  40. data/lib/io_streams/row/writer.rb +1 -1
  41. data/lib/io_streams/stream.rb +48 -37
  42. data/lib/io_streams/symmetric_encryption/reader.rb +6 -2
  43. data/lib/io_streams/symmetric_encryption/writer.rb +8 -4
  44. data/lib/io_streams/tabular/header.rb +49 -10
  45. data/lib/io_streams/tabular/parser/array.rb +0 -10
  46. data/lib/io_streams/tabular/parser/base.rb +10 -0
  47. data/lib/io_streams/tabular/parser/csv.rb +9 -36
  48. data/lib/io_streams/tabular/parser/fixed.rb +8 -6
  49. data/lib/io_streams/tabular/parser/psv.rb +6 -14
  50. data/lib/io_streams/tabular.rb +5 -10
  51. data/lib/io_streams/utils.rb +34 -2
  52. data/lib/io_streams/version.rb +1 -1
  53. data/lib/io_streams/writer.rb +16 -7
  54. data/lib/io_streams/xlsx/reader.rb +6 -2
  55. data/lib/io_streams/zip/reader.rb +4 -0
  56. data/lib/io_streams/zip/writer.rb +26 -10
  57. data/lib/iostreams.rb +0 -1
  58. metadata +46 -112
  59. data/lib/io_streams/deprecated.rb +0 -216
  60. data/lib/io_streams/tabular/utility/csv_row.rb +0 -105
  61. data/test/builder_test.rb +0 -311
  62. data/test/bzip2_reader_test.rb +0 -27
  63. data/test/bzip2_writer_test.rb +0 -56
  64. data/test/deprecated_test.rb +0 -121
  65. data/test/encode_reader_test.rb +0 -51
  66. data/test/encode_writer_test.rb +0 -90
  67. data/test/files/embedded_lines_test.csv +0 -7
  68. data/test/files/multiple_files.zip +0 -0
  69. data/test/files/spreadsheet.xlsx +0 -0
  70. data/test/files/test.csv +0 -4
  71. data/test/files/test.json +0 -3
  72. data/test/files/test.psv +0 -4
  73. data/test/files/text file.txt +0 -3
  74. data/test/files/text.txt +0 -3
  75. data/test/files/text.txt.bz2 +0 -0
  76. data/test/files/text.txt.gz +0 -0
  77. data/test/files/text.txt.gz.zip +0 -0
  78. data/test/files/text.zip +0 -0
  79. data/test/files/text.zip.gz +0 -0
  80. data/test/files/unclosed_quote_large_test.csv +0 -1658
  81. data/test/files/unclosed_quote_test.csv +0 -4
  82. data/test/files/unclosed_quote_test2.csv +0 -3
  83. data/test/gzip_reader_test.rb +0 -27
  84. data/test/gzip_writer_test.rb +0 -52
  85. data/test/io_streams_test.rb +0 -132
  86. data/test/line_reader_test.rb +0 -325
  87. data/test/line_writer_test.rb +0 -59
  88. data/test/minimal_file_reader.rb +0 -25
  89. data/test/path_test.rb +0 -55
  90. data/test/paths/file_test.rb +0 -213
  91. data/test/paths/http_test.rb +0 -34
  92. data/test/paths/matcher_test.rb +0 -120
  93. data/test/paths/s3_test.rb +0 -220
  94. data/test/paths/sftp_test.rb +0 -106
  95. data/test/pgp_reader_test.rb +0 -46
  96. data/test/pgp_test.rb +0 -267
  97. data/test/pgp_writer_test.rb +0 -130
  98. data/test/record_reader_test.rb +0 -60
  99. data/test/record_writer_test.rb +0 -82
  100. data/test/row_reader_test.rb +0 -35
  101. data/test/row_writer_test.rb +0 -56
  102. data/test/stream_test.rb +0 -577
  103. data/test/tabular_test.rb +0 -338
  104. data/test/test_helper.rb +0 -40
  105. data/test/utils_test.rb +0 -20
  106. data/test/xlsx_reader_test.rb +0 -37
  107. data/test/zip_reader_test.rb +0 -53
  108. data/test/zip_writer_test.rb +0 -48
data/docs/streams.md ADDED
@@ -0,0 +1,337 @@
1
+ ---
2
+ layout: default
3
+ title: Streams
4
+ description: >-
5
+ Reading and writing a path a block, line, row or record at a time, and the
6
+ pipeline of compression, encryption and format streams applied to it.
7
+ ---
8
+
9
+ Once you have a [path](path), you read from and write to it with a small, consistent set of methods.
10
+ Reading and writing always happen a block, line, or record at a time, so memory use stays low no
11
+ matter how large the file is. Choose how each chunk is delivered by passing a mode: the default
12
+ streams raw data, `:line` yields one line at a time, `:array` yields each row as an array, and
13
+ `:hash` yields each record as a hash keyed by the header row.
14
+
15
+ Read 128 characters at a time from the file:
16
+ ~~~ruby
17
+ IOStreams.path("example.csv").reader do |io|
18
+ while (data = io.read(128))
19
+ p data
20
+ end
21
+ end
22
+ ~~~
23
+
24
+ Read one line at a time from the file:
25
+ ~~~ruby
26
+ IOStreams.path("example.csv").each do |line|
27
+ puts line
28
+ end
29
+ ~~~
30
+
31
+ By default the line delimiter is auto-detected from the file, handling both Windows (`\r\n`)
32
+ and Linux (`\n`) line endings. To break the file up by something other than its line endings,
33
+ supply `delimiter`:
34
+ ~~~ruby
35
+ IOStreams.path("example.txt").each(:line, delimiter: "|") do |line|
36
+ puts line
37
+ end
38
+ ~~~
39
+
40
+ When a line can contain embedded newlines, such as a CSV field wrapped in double quotes that
41
+ spans multiple lines, supply `embedded_within` so those newlines are not treated as line endings:
42
+ ~~~ruby
43
+ IOStreams.path("example.csv").each(:line, embedded_within: '"') do |line|
44
+ puts line
45
+ end
46
+ ~~~
47
+
48
+ Notes:
49
+ * Newlines embedded within quoted fields are kept on the same line automatically when the resolved
50
+ tabular format quotes its fields, such as CSV (detected from a `.csv` file name or set explicitly via
51
+ `.format(:csv)`). `embedded_within` only needs to be supplied for quoted formats that are not
52
+ auto-detected, or to override the quote character.
53
+ * A file that is named `.csv` but is actually pipe-delimited can avoid quote parsing by declaring its
54
+ real format with `.format(:psv)`, or by passing `embedded_within: nil` to disable it explicitly:
55
+ ~~~ruby
56
+ IOStreams.path("pipe_delimited.csv").format(:psv).each(:line) do |line|
57
+ puts line
58
+ end
59
+ ~~~
60
+
61
+ Display each row from the csv file as an array:
62
+ ~~~ruby
63
+ IOStreams.path("example.csv").each(:array) do |array|
64
+ p array
65
+ end
66
+ ~~~
67
+
68
+ Display each row from the csv file as a hash, where the first line in the CSV file is the header:
69
+ ~~~ruby
70
+ IOStreams.path("example.csv").each(:hash) do |hash|
71
+ p hash
72
+ end
73
+ ~~~
74
+
75
+ Write data to the file.
76
+ ~~~ruby
77
+ IOStreams.path("abc.txt").writer do |io|
78
+ io << "This"
79
+ io << " is "
80
+ io << " one line\n"
81
+ end
82
+ ~~~
83
+
84
+ Write lines to the file. By adding `:line` to `writer`, each write appends a new line character.
85
+ ~~~ruby
86
+ IOStreams.path("example.csv").writer(:line) do |file|
87
+ file << "these"
88
+ file << "are"
89
+ file << "all"
90
+ file << "separate"
91
+ file << "lines"
92
+ end
93
+ ~~~
94
+
95
+ Write an array (row) at a time to the file.
96
+ Each array is converted to csv before being compressed with zip.
97
+
98
+ ~~~ruby
99
+ IOStreams.path("example.csv").writer(:array) do |io|
100
+ io << ["name", "address", "zip_code"]
101
+ io << ["Jack", "There", "1234"]
102
+ io << ["Joe", "Over There somewhere", 1234]
103
+ end
104
+ ~~~
105
+
106
+ Write a hash (record) at a time to the file.
107
+ Each hash is converted to csv before being compressed with zip.
108
+ The header row is extracted from the first hash write that is performed.
109
+
110
+ ~~~ruby
111
+ IOStreams.path("example.csv").writer(:hash) do |stream|
112
+ stream << {name: "Jack", address: "There", zip_code: 1234}
113
+ stream << {zip_code: 1234, address: "Over There somewhere", name: "Joe"}
114
+ end
115
+ ~~~
116
+
117
+ Notes
118
+ * Any additional keys supplied during subsequent write operations will be ignored
119
+ since the header row has already been written to the file.
120
+ * The order of the header and values is determined by the order of the keys supplied
121
+ during the first write.
122
+ * The order of keys in the subsequent writes does not matter.
123
+
124
+ Stream into an in-memory buffer, useful for testing.
125
+ The original filename still needs to be supplied so that the streaming pipeline can still be inferred.
126
+
127
+ ~~~ruby
128
+ io = StringIO.new
129
+ IOStreams.stream(io).file_name("example.csv.gz").writer(:hash) do |stream|
130
+ stream << {name: "Jack", address: "There", zip_code: 1234}
131
+ stream << {name: "Joe", zip_code: 1234, address: "Over There somewhere"}
132
+ end
133
+ puts io.string
134
+ ~~~
135
+
136
+ Read a CSV file and write the output to an encrypted file in JSON format.
137
+
138
+ ~~~ruby
139
+ IOStreams.path("sample.json.enc").writer(:hash) do |output|
140
+ IOStreams.path("sample.csv").each(:hash) do |record|
141
+ output << record
142
+ end
143
+ end
144
+ ~~~
145
+
146
+ Read a zip file hosted on a HTTP Web Server, returning each row as a hash:
147
+ ~~~ruby
148
+ IOStreams.
149
+ path("https://www5.fdic.gov/idasp/Offices2.zip").
150
+ option(:zip, entry_file_name: "OFFICES2_ALL.CSV").
151
+ each(:hash) do |row|
152
+ p row
153
+ end
154
+ ~~~
155
+
156
+ Notes:
157
+ * By default IOStreams will read the first file in the zip file.
158
+ * To choose a specific file name within the zip file, supply: `entry_file_name`
159
+
160
+ ## Notes
161
+
162
+ * Reading a Zip file requires the entire file to be available locally, so reading from a
163
+ stream (for example S3 or HTTP) downloads it into a temp file first.
164
+ Writing Zip is fully streamed, no temp file is required.
165
+ * When writing, `entry_file_name` sets the name of the file entry within the zip file.
166
+ It defaults to the file name without the `.zip` extension, so writing to
167
+ `example.csv.zip` creates an entry named `example.csv`.
168
+ * Gzip is still recommended over Zip for very large files, since Zip files can only
169
+ be read via a local file.
170
+
171
+ ## Pipeline
172
+
173
+ If the file is compressed, the pipeline will infer the necessary streams that need to be applied to it:
174
+
175
+ ~~~ruby
176
+ path = IOStreams.path("somewhere/example.csv.gz")
177
+ # => #<IOStreams::Paths::File:somewhere/example.csv.gz pipeline={:gz=>{}}>
178
+
179
+ path.pipeline
180
+ # => {:gz=>{}}
181
+ ~~~
182
+
183
+ The `pipeline` above includes `:gz` to indicate that the file should compressed / decompressed with GZip.
184
+
185
+ #### Option
186
+
187
+ Each path supports several options which can be supplied using the `option` method.
188
+
189
+ Set the options for a stream in the pipeline for this file. Each stream can only be applied once and is uniquely
190
+ identified by its symbolic name.
191
+
192
+ To see the pipeline of streams that IOStreams would infer:
193
+ ~~~ruby
194
+ IOStreams.path("example.pgp").pipeline
195
+ # => {:pgp=>{}}
196
+
197
+ IOStreams.path("example.gz").pipeline
198
+ # => {:gz=>{}}
199
+
200
+ IOStreams.path("example.gz.pgp").pipeline
201
+ # => {:gz=>{}, :pgp=>{}}
202
+ ~~~
203
+
204
+ If the relevant stream is not found for this file it is ignored.
205
+ For example, if the file does not have a pgp extension then the pgp option is ignored.
206
+ ~~~ruby
207
+ IOStreams.path("example.csv.gz").
208
+ option(:pgp, passphrase: "receiver_passphrase").
209
+ read
210
+ ~~~
211
+
212
+ This is great way to pass in stream specific options for when they are required, and to still support
213
+ paths that do not use that stream. For example, the same code can support pgp encrypted, Symmetric Encryption encrypted,
214
+ and plain text files.
215
+ ~~~ruby
216
+ IOStreams.path("example.csv.enc").
217
+ option(:pgp, passphrase: "receiver_passphrase").
218
+ read
219
+ ~~~
220
+
221
+ To see the what value was previously set for a particular option:
222
+ ~~~ruby
223
+ path = IOStreams.path("example.pgp")
224
+ path.option(:pgp, passphrase: "receiver_passphrase")
225
+
226
+ path.setting(:pgp)
227
+ # => {:passphrase=>"receiver_passphrase"}
228
+ ~~~
229
+
230
+ #### Reading and writing need separate options
231
+
232
+ The options for a stream are passed to its reader when reading and to its writer when writing,
233
+ and the two directions usually accept different options. For example, the PGP writer needs
234
+ the `recipient` to encrypt for, while the PGP reader needs the `passphrase` for the private key.
235
+
236
+ Options are strict: an option that does not apply to the direction being used raises an
237
+ `ArgumentError` that names the direction it belongs to, rather than being silently ignored.
238
+ So configure one path for writing and a separate path for reading:
239
+
240
+ ~~~ruby
241
+ IOStreams.path("example.csv.pgp").
242
+ option(:pgp, recipient: "receiver@example.org").
243
+ write("name,login\nJack Jones,jjones\n")
244
+
245
+ IOStreams.path("example.csv.pgp").
246
+ option(:pgp, passphrase: "receiver_passphrase").
247
+ read
248
+ ~~~
249
+
250
+ Reading with the writer's options fails before any data is read:
251
+ ~~~ruby
252
+ IOStreams.path("example.csv.pgp").
253
+ option(:pgp, recipient: "receiver@example.org").
254
+ read
255
+ # ArgumentError: :recipient only applies when writing a :pgp stream and cannot be used when reading.
256
+ # Configure a separate path or stream without it for reading.
257
+ ~~~
258
+
259
+ An option that neither direction accepts, such as a misspelled one, raises an `ArgumentError`
260
+ that lists the valid options. The exception is BZip2, which ignores options it does not accept and
261
+ logs a warning instead. In v3.0 it will raise an `ArgumentError` too.
262
+
263
+ Options for a stream that is not in the pipeline are still ignored, as described above,
264
+ since they are not passed to any reader or writer.
265
+
266
+ #### Stream
267
+
268
+ The `stream` method stops IOStreams from inferring the streams for this path and only uses the specified streams.
269
+
270
+ For example when using a filename that does not have the necessary file extensions.
271
+ In this case the file was compressed with Zip, so tell IOStreams to unzip it:
272
+ ~~~ruby
273
+ path = IOStreams.path("tempfile2527")
274
+ path.stream(:zip)
275
+ path.read
276
+ ~~~
277
+
278
+ The above example could also be written as:
279
+ ~~~ruby
280
+ IOStreams.path("tempfile2527").
281
+ stream(:zip).
282
+ read
283
+ ~~~
284
+
285
+ Multiple streams can also be specified:
286
+
287
+ ~~~ruby
288
+ IOStreams.path("tempfile2527").
289
+ stream(:zip).
290
+ stream(:pgp, passphrase: "receiver_passphrase").
291
+ read
292
+ ~~~
293
+
294
+ Now that IOStreams is not inferring the pipeline from the filename, we can still see the above streams:
295
+ ~~~ruby
296
+ IOStreams.path("tempfile2527").
297
+ stream(:zip).
298
+ stream(:pgp, passphrase: "receiver_passphrase").
299
+ pipeline
300
+ # => {:zip=>{}, :pgp=>{:passphrase=>"receiver_passphrase"}}
301
+ ~~~
302
+
303
+
304
+ In this example the file contains JSON data that was compressed with Zip, and since we want to read each row as a hash:
305
+ ~~~ruby
306
+ IOStreams.path("tempfile2527").
307
+ stream(:zip).
308
+ each(:hash, format: :json) do |row|
309
+ p row
310
+ end
311
+ ~~~
312
+
313
+ Alternatively if the original file name is available it can also be supplied allowing IOStreams to infer the above streams:
314
+ ~~~ruby
315
+ IOStreams.path("tempfile2527").
316
+ file_name("file.json.zip").
317
+ each(:hash) do |row|
318
+ p row
319
+ end
320
+ ~~~
321
+
322
+ To see the what value was previously set for a particular stream:
323
+ ~~~ruby
324
+ path = IOStreams.path("tempfile2527")
325
+ path.stream(:zip)
326
+ path.stream(:pgp, passphrase: "receiver_passphrase")
327
+
328
+ path.setting(:pgp)
329
+ # => {:passphrase=>"receiver_passphrase"}
330
+ ~~~
331
+
332
+ To ensure no streams are inferred or applied use stream `:none`
333
+ ~~~ruby
334
+ path = IOStreams.path("file.zip")
335
+ path.stream(:none)
336
+ path.read
337
+ ~~~