iostreams 1.11.0 → 2.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.md +14 -13
- data/Rakefile +52 -0
- data/docs/CLAUDE.md +9 -0
- data/docs/config.md +157 -0
- data/docs/copy_files.md +75 -0
- data/docs/extensions.md +111 -0
- data/docs/formats.md +188 -0
- data/docs/index.md +388 -0
- data/docs/path.md +652 -0
- data/docs/pgp.md +436 -0
- data/docs/streams.md +337 -0
- data/docs/tutorial.md +483 -0
- data/docs/upgrading.md +217 -0
- data/lib/io_streams/builder.rb +71 -11
- data/lib/io_streams/bzip2/reader.rb +25 -2
- data/lib/io_streams/bzip2/writer.rb +26 -2
- data/lib/io_streams/encode/reader.rb +6 -2
- data/lib/io_streams/encode/writer.rb +9 -5
- data/lib/io_streams/errors.rb +4 -0
- data/lib/io_streams/gzip/reader.rb +5 -1
- data/lib/io_streams/gzip/writer.rb +11 -2
- data/lib/io_streams/io_streams.rb +156 -20
- data/lib/io_streams/line/reader.rb +9 -4
- data/lib/io_streams/line/writer.rb +1 -1
- data/lib/io_streams/path.rb +117 -8
- data/lib/io_streams/paths/file.rb +57 -11
- data/lib/io_streams/paths/http.rb +123 -9
- data/lib/io_streams/paths/matcher.rb +3 -3
- data/lib/io_streams/paths/s3.rb +69 -18
- data/lib/io_streams/paths/sftp/net_ssh.rb +104 -0
- data/lib/io_streams/paths/sftp.rb +103 -64
- data/lib/io_streams/pgp/reader.rb +63 -10
- data/lib/io_streams/pgp/writer.rb +111 -30
- data/lib/io_streams/pgp.rb +256 -71
- data/lib/io_streams/reader.rb +14 -5
- data/lib/io_streams/record/reader.rb +75 -6
- data/lib/io_streams/record/writer.rb +3 -4
- data/lib/io_streams/row/reader.rb +1 -1
- data/lib/io_streams/row/writer.rb +1 -1
- data/lib/io_streams/stream.rb +48 -37
- data/lib/io_streams/symmetric_encryption/reader.rb +6 -2
- data/lib/io_streams/symmetric_encryption/writer.rb +8 -4
- data/lib/io_streams/tabular/header.rb +49 -10
- data/lib/io_streams/tabular/parser/array.rb +0 -10
- data/lib/io_streams/tabular/parser/base.rb +10 -0
- data/lib/io_streams/tabular/parser/csv.rb +9 -36
- data/lib/io_streams/tabular/parser/fixed.rb +8 -6
- data/lib/io_streams/tabular/parser/psv.rb +6 -14
- data/lib/io_streams/tabular.rb +5 -10
- data/lib/io_streams/utils.rb +34 -2
- data/lib/io_streams/version.rb +1 -1
- data/lib/io_streams/writer.rb +16 -7
- data/lib/io_streams/xlsx/reader.rb +6 -2
- data/lib/io_streams/zip/reader.rb +4 -0
- data/lib/io_streams/zip/writer.rb +26 -10
- data/lib/iostreams.rb +0 -1
- metadata +46 -112
- data/lib/io_streams/deprecated.rb +0 -216
- data/lib/io_streams/tabular/utility/csv_row.rb +0 -105
- data/test/builder_test.rb +0 -311
- data/test/bzip2_reader_test.rb +0 -27
- data/test/bzip2_writer_test.rb +0 -56
- data/test/deprecated_test.rb +0 -121
- data/test/encode_reader_test.rb +0 -51
- data/test/encode_writer_test.rb +0 -90
- data/test/files/embedded_lines_test.csv +0 -7
- data/test/files/multiple_files.zip +0 -0
- data/test/files/spreadsheet.xlsx +0 -0
- data/test/files/test.csv +0 -4
- data/test/files/test.json +0 -3
- data/test/files/test.psv +0 -4
- data/test/files/text file.txt +0 -3
- data/test/files/text.txt +0 -3
- data/test/files/text.txt.bz2 +0 -0
- data/test/files/text.txt.gz +0 -0
- data/test/files/text.txt.gz.zip +0 -0
- data/test/files/text.zip +0 -0
- data/test/files/text.zip.gz +0 -0
- data/test/files/unclosed_quote_large_test.csv +0 -1658
- data/test/files/unclosed_quote_test.csv +0 -4
- data/test/files/unclosed_quote_test2.csv +0 -3
- data/test/gzip_reader_test.rb +0 -27
- data/test/gzip_writer_test.rb +0 -52
- data/test/io_streams_test.rb +0 -132
- data/test/line_reader_test.rb +0 -325
- data/test/line_writer_test.rb +0 -59
- data/test/minimal_file_reader.rb +0 -25
- data/test/path_test.rb +0 -55
- data/test/paths/file_test.rb +0 -213
- data/test/paths/http_test.rb +0 -34
- data/test/paths/matcher_test.rb +0 -120
- data/test/paths/s3_test.rb +0 -220
- data/test/paths/sftp_test.rb +0 -106
- data/test/pgp_reader_test.rb +0 -46
- data/test/pgp_test.rb +0 -267
- data/test/pgp_writer_test.rb +0 -130
- data/test/record_reader_test.rb +0 -60
- data/test/record_writer_test.rb +0 -82
- data/test/row_reader_test.rb +0 -35
- data/test/row_writer_test.rb +0 -56
- data/test/stream_test.rb +0 -577
- data/test/tabular_test.rb +0 -338
- data/test/test_helper.rb +0 -40
- data/test/utils_test.rb +0 -20
- data/test/xlsx_reader_test.rb +0 -37
- data/test/zip_reader_test.rb +0 -53
- data/test/zip_writer_test.rb +0 -48
data/docs/streams.md
ADDED
|
@@ -0,0 +1,337 @@
|
|
|
1
|
+
---
|
|
2
|
+
layout: default
|
|
3
|
+
title: Streams
|
|
4
|
+
description: >-
|
|
5
|
+
Reading and writing a path a block, line, row or record at a time, and the
|
|
6
|
+
pipeline of compression, encryption and format streams applied to it.
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
Once you have a [path](path), you read from and write to it with a small, consistent set of methods.
|
|
10
|
+
Reading and writing always happen a block, line, or record at a time, so memory use stays low no
|
|
11
|
+
matter how large the file is. Choose how each chunk is delivered by passing a mode: the default
|
|
12
|
+
streams raw data, `:line` yields one line at a time, `:array` yields each row as an array, and
|
|
13
|
+
`:hash` yields each record as a hash keyed by the header row.
|
|
14
|
+
|
|
15
|
+
Read 128 characters at a time from the file:
|
|
16
|
+
~~~ruby
|
|
17
|
+
IOStreams.path("example.csv").reader do |io|
|
|
18
|
+
while (data = io.read(128))
|
|
19
|
+
p data
|
|
20
|
+
end
|
|
21
|
+
end
|
|
22
|
+
~~~
|
|
23
|
+
|
|
24
|
+
Read one line at a time from the file:
|
|
25
|
+
~~~ruby
|
|
26
|
+
IOStreams.path("example.csv").each do |line|
|
|
27
|
+
puts line
|
|
28
|
+
end
|
|
29
|
+
~~~
|
|
30
|
+
|
|
31
|
+
By default the line delimiter is auto-detected from the file, handling both Windows (`\r\n`)
|
|
32
|
+
and Linux (`\n`) line endings. To break the file up by something other than its line endings,
|
|
33
|
+
supply `delimiter`:
|
|
34
|
+
~~~ruby
|
|
35
|
+
IOStreams.path("example.txt").each(:line, delimiter: "|") do |line|
|
|
36
|
+
puts line
|
|
37
|
+
end
|
|
38
|
+
~~~
|
|
39
|
+
|
|
40
|
+
When a line can contain embedded newlines, such as a CSV field wrapped in double quotes that
|
|
41
|
+
spans multiple lines, supply `embedded_within` so those newlines are not treated as line endings:
|
|
42
|
+
~~~ruby
|
|
43
|
+
IOStreams.path("example.csv").each(:line, embedded_within: '"') do |line|
|
|
44
|
+
puts line
|
|
45
|
+
end
|
|
46
|
+
~~~
|
|
47
|
+
|
|
48
|
+
Notes:
|
|
49
|
+
* Newlines embedded within quoted fields are kept on the same line automatically when the resolved
|
|
50
|
+
tabular format quotes its fields, such as CSV (detected from a `.csv` file name or set explicitly via
|
|
51
|
+
`.format(:csv)`). `embedded_within` only needs to be supplied for quoted formats that are not
|
|
52
|
+
auto-detected, or to override the quote character.
|
|
53
|
+
* A file that is named `.csv` but is actually pipe-delimited can avoid quote parsing by declaring its
|
|
54
|
+
real format with `.format(:psv)`, or by passing `embedded_within: nil` to disable it explicitly:
|
|
55
|
+
~~~ruby
|
|
56
|
+
IOStreams.path("pipe_delimited.csv").format(:psv).each(:line) do |line|
|
|
57
|
+
puts line
|
|
58
|
+
end
|
|
59
|
+
~~~
|
|
60
|
+
|
|
61
|
+
Display each row from the csv file as an array:
|
|
62
|
+
~~~ruby
|
|
63
|
+
IOStreams.path("example.csv").each(:array) do |array|
|
|
64
|
+
p array
|
|
65
|
+
end
|
|
66
|
+
~~~
|
|
67
|
+
|
|
68
|
+
Display each row from the csv file as a hash, where the first line in the CSV file is the header:
|
|
69
|
+
~~~ruby
|
|
70
|
+
IOStreams.path("example.csv").each(:hash) do |hash|
|
|
71
|
+
p hash
|
|
72
|
+
end
|
|
73
|
+
~~~
|
|
74
|
+
|
|
75
|
+
Write data to the file.
|
|
76
|
+
~~~ruby
|
|
77
|
+
IOStreams.path("abc.txt").writer do |io|
|
|
78
|
+
io << "This"
|
|
79
|
+
io << " is "
|
|
80
|
+
io << " one line\n"
|
|
81
|
+
end
|
|
82
|
+
~~~
|
|
83
|
+
|
|
84
|
+
Write lines to the file. By adding `:line` to `writer`, each write appends a new line character.
|
|
85
|
+
~~~ruby
|
|
86
|
+
IOStreams.path("example.csv").writer(:line) do |file|
|
|
87
|
+
file << "these"
|
|
88
|
+
file << "are"
|
|
89
|
+
file << "all"
|
|
90
|
+
file << "separate"
|
|
91
|
+
file << "lines"
|
|
92
|
+
end
|
|
93
|
+
~~~
|
|
94
|
+
|
|
95
|
+
Write an array (row) at a time to the file.
|
|
96
|
+
Each array is converted to csv before being compressed with zip.
|
|
97
|
+
|
|
98
|
+
~~~ruby
|
|
99
|
+
IOStreams.path("example.csv").writer(:array) do |io|
|
|
100
|
+
io << ["name", "address", "zip_code"]
|
|
101
|
+
io << ["Jack", "There", "1234"]
|
|
102
|
+
io << ["Joe", "Over There somewhere", 1234]
|
|
103
|
+
end
|
|
104
|
+
~~~
|
|
105
|
+
|
|
106
|
+
Write a hash (record) at a time to the file.
|
|
107
|
+
Each hash is converted to csv before being compressed with zip.
|
|
108
|
+
The header row is extracted from the first hash write that is performed.
|
|
109
|
+
|
|
110
|
+
~~~ruby
|
|
111
|
+
IOStreams.path("example.csv").writer(:hash) do |stream|
|
|
112
|
+
stream << {name: "Jack", address: "There", zip_code: 1234}
|
|
113
|
+
stream << {zip_code: 1234, address: "Over There somewhere", name: "Joe"}
|
|
114
|
+
end
|
|
115
|
+
~~~
|
|
116
|
+
|
|
117
|
+
Notes
|
|
118
|
+
* Any additional keys supplied during subsequent write operations will be ignored
|
|
119
|
+
since the header row has already been written to the file.
|
|
120
|
+
* The order of the header and values is determined by the order of the keys supplied
|
|
121
|
+
during the first write.
|
|
122
|
+
* The order of keys in the subsequent writes does not matter.
|
|
123
|
+
|
|
124
|
+
Stream into an in-memory buffer, useful for testing.
|
|
125
|
+
The original filename still needs to be supplied so that the streaming pipeline can still be inferred.
|
|
126
|
+
|
|
127
|
+
~~~ruby
|
|
128
|
+
io = StringIO.new
|
|
129
|
+
IOStreams.stream(io).file_name("example.csv.gz").writer(:hash) do |stream|
|
|
130
|
+
stream << {name: "Jack", address: "There", zip_code: 1234}
|
|
131
|
+
stream << {name: "Joe", zip_code: 1234, address: "Over There somewhere"}
|
|
132
|
+
end
|
|
133
|
+
puts io.string
|
|
134
|
+
~~~
|
|
135
|
+
|
|
136
|
+
Read a CSV file and write the output to an encrypted file in JSON format.
|
|
137
|
+
|
|
138
|
+
~~~ruby
|
|
139
|
+
IOStreams.path("sample.json.enc").writer(:hash) do |output|
|
|
140
|
+
IOStreams.path("sample.csv").each(:hash) do |record|
|
|
141
|
+
output << record
|
|
142
|
+
end
|
|
143
|
+
end
|
|
144
|
+
~~~
|
|
145
|
+
|
|
146
|
+
Read a zip file hosted on a HTTP Web Server, returning each row as a hash:
|
|
147
|
+
~~~ruby
|
|
148
|
+
IOStreams.
|
|
149
|
+
path("https://www5.fdic.gov/idasp/Offices2.zip").
|
|
150
|
+
option(:zip, entry_file_name: "OFFICES2_ALL.CSV").
|
|
151
|
+
each(:hash) do |row|
|
|
152
|
+
p row
|
|
153
|
+
end
|
|
154
|
+
~~~
|
|
155
|
+
|
|
156
|
+
Notes:
|
|
157
|
+
* By default IOStreams will read the first file in the zip file.
|
|
158
|
+
* To choose a specific file name within the zip file, supply: `entry_file_name`
|
|
159
|
+
|
|
160
|
+
## Notes
|
|
161
|
+
|
|
162
|
+
* Reading a Zip file requires the entire file to be available locally, so reading from a
|
|
163
|
+
stream (for example S3 or HTTP) downloads it into a temp file first.
|
|
164
|
+
Writing Zip is fully streamed, no temp file is required.
|
|
165
|
+
* When writing, `entry_file_name` sets the name of the file entry within the zip file.
|
|
166
|
+
It defaults to the file name without the `.zip` extension, so writing to
|
|
167
|
+
`example.csv.zip` creates an entry named `example.csv`.
|
|
168
|
+
* Gzip is still recommended over Zip for very large files, since Zip files can only
|
|
169
|
+
be read via a local file.
|
|
170
|
+
|
|
171
|
+
## Pipeline
|
|
172
|
+
|
|
173
|
+
If the file is compressed, the pipeline will infer the necessary streams that need to be applied to it:
|
|
174
|
+
|
|
175
|
+
~~~ruby
|
|
176
|
+
path = IOStreams.path("somewhere/example.csv.gz")
|
|
177
|
+
# => #<IOStreams::Paths::File:somewhere/example.csv.gz pipeline={:gz=>{}}>
|
|
178
|
+
|
|
179
|
+
path.pipeline
|
|
180
|
+
# => {:gz=>{}}
|
|
181
|
+
~~~
|
|
182
|
+
|
|
183
|
+
The `pipeline` above includes `:gz` to indicate that the file should compressed / decompressed with GZip.
|
|
184
|
+
|
|
185
|
+
#### Option
|
|
186
|
+
|
|
187
|
+
Each path supports several options which can be supplied using the `option` method.
|
|
188
|
+
|
|
189
|
+
Set the options for a stream in the pipeline for this file. Each stream can only be applied once and is uniquely
|
|
190
|
+
identified by its symbolic name.
|
|
191
|
+
|
|
192
|
+
To see the pipeline of streams that IOStreams would infer:
|
|
193
|
+
~~~ruby
|
|
194
|
+
IOStreams.path("example.pgp").pipeline
|
|
195
|
+
# => {:pgp=>{}}
|
|
196
|
+
|
|
197
|
+
IOStreams.path("example.gz").pipeline
|
|
198
|
+
# => {:gz=>{}}
|
|
199
|
+
|
|
200
|
+
IOStreams.path("example.gz.pgp").pipeline
|
|
201
|
+
# => {:gz=>{}, :pgp=>{}}
|
|
202
|
+
~~~
|
|
203
|
+
|
|
204
|
+
If the relevant stream is not found for this file it is ignored.
|
|
205
|
+
For example, if the file does not have a pgp extension then the pgp option is ignored.
|
|
206
|
+
~~~ruby
|
|
207
|
+
IOStreams.path("example.csv.gz").
|
|
208
|
+
option(:pgp, passphrase: "receiver_passphrase").
|
|
209
|
+
read
|
|
210
|
+
~~~
|
|
211
|
+
|
|
212
|
+
This is great way to pass in stream specific options for when they are required, and to still support
|
|
213
|
+
paths that do not use that stream. For example, the same code can support pgp encrypted, Symmetric Encryption encrypted,
|
|
214
|
+
and plain text files.
|
|
215
|
+
~~~ruby
|
|
216
|
+
IOStreams.path("example.csv.enc").
|
|
217
|
+
option(:pgp, passphrase: "receiver_passphrase").
|
|
218
|
+
read
|
|
219
|
+
~~~
|
|
220
|
+
|
|
221
|
+
To see the what value was previously set for a particular option:
|
|
222
|
+
~~~ruby
|
|
223
|
+
path = IOStreams.path("example.pgp")
|
|
224
|
+
path.option(:pgp, passphrase: "receiver_passphrase")
|
|
225
|
+
|
|
226
|
+
path.setting(:pgp)
|
|
227
|
+
# => {:passphrase=>"receiver_passphrase"}
|
|
228
|
+
~~~
|
|
229
|
+
|
|
230
|
+
#### Reading and writing need separate options
|
|
231
|
+
|
|
232
|
+
The options for a stream are passed to its reader when reading and to its writer when writing,
|
|
233
|
+
and the two directions usually accept different options. For example, the PGP writer needs
|
|
234
|
+
the `recipient` to encrypt for, while the PGP reader needs the `passphrase` for the private key.
|
|
235
|
+
|
|
236
|
+
Options are strict: an option that does not apply to the direction being used raises an
|
|
237
|
+
`ArgumentError` that names the direction it belongs to, rather than being silently ignored.
|
|
238
|
+
So configure one path for writing and a separate path for reading:
|
|
239
|
+
|
|
240
|
+
~~~ruby
|
|
241
|
+
IOStreams.path("example.csv.pgp").
|
|
242
|
+
option(:pgp, recipient: "receiver@example.org").
|
|
243
|
+
write("name,login\nJack Jones,jjones\n")
|
|
244
|
+
|
|
245
|
+
IOStreams.path("example.csv.pgp").
|
|
246
|
+
option(:pgp, passphrase: "receiver_passphrase").
|
|
247
|
+
read
|
|
248
|
+
~~~
|
|
249
|
+
|
|
250
|
+
Reading with the writer's options fails before any data is read:
|
|
251
|
+
~~~ruby
|
|
252
|
+
IOStreams.path("example.csv.pgp").
|
|
253
|
+
option(:pgp, recipient: "receiver@example.org").
|
|
254
|
+
read
|
|
255
|
+
# ArgumentError: :recipient only applies when writing a :pgp stream and cannot be used when reading.
|
|
256
|
+
# Configure a separate path or stream without it for reading.
|
|
257
|
+
~~~
|
|
258
|
+
|
|
259
|
+
An option that neither direction accepts, such as a misspelled one, raises an `ArgumentError`
|
|
260
|
+
that lists the valid options. The exception is BZip2, which ignores options it does not accept and
|
|
261
|
+
logs a warning instead. In v3.0 it will raise an `ArgumentError` too.
|
|
262
|
+
|
|
263
|
+
Options for a stream that is not in the pipeline are still ignored, as described above,
|
|
264
|
+
since they are not passed to any reader or writer.
|
|
265
|
+
|
|
266
|
+
#### Stream
|
|
267
|
+
|
|
268
|
+
The `stream` method stops IOStreams from inferring the streams for this path and only uses the specified streams.
|
|
269
|
+
|
|
270
|
+
For example when using a filename that does not have the necessary file extensions.
|
|
271
|
+
In this case the file was compressed with Zip, so tell IOStreams to unzip it:
|
|
272
|
+
~~~ruby
|
|
273
|
+
path = IOStreams.path("tempfile2527")
|
|
274
|
+
path.stream(:zip)
|
|
275
|
+
path.read
|
|
276
|
+
~~~
|
|
277
|
+
|
|
278
|
+
The above example could also be written as:
|
|
279
|
+
~~~ruby
|
|
280
|
+
IOStreams.path("tempfile2527").
|
|
281
|
+
stream(:zip).
|
|
282
|
+
read
|
|
283
|
+
~~~
|
|
284
|
+
|
|
285
|
+
Multiple streams can also be specified:
|
|
286
|
+
|
|
287
|
+
~~~ruby
|
|
288
|
+
IOStreams.path("tempfile2527").
|
|
289
|
+
stream(:zip).
|
|
290
|
+
stream(:pgp, passphrase: "receiver_passphrase").
|
|
291
|
+
read
|
|
292
|
+
~~~
|
|
293
|
+
|
|
294
|
+
Now that IOStreams is not inferring the pipeline from the filename, we can still see the above streams:
|
|
295
|
+
~~~ruby
|
|
296
|
+
IOStreams.path("tempfile2527").
|
|
297
|
+
stream(:zip).
|
|
298
|
+
stream(:pgp, passphrase: "receiver_passphrase").
|
|
299
|
+
pipeline
|
|
300
|
+
# => {:zip=>{}, :pgp=>{:passphrase=>"receiver_passphrase"}}
|
|
301
|
+
~~~
|
|
302
|
+
|
|
303
|
+
|
|
304
|
+
In this example the file contains JSON data that was compressed with Zip, and since we want to read each row as a hash:
|
|
305
|
+
~~~ruby
|
|
306
|
+
IOStreams.path("tempfile2527").
|
|
307
|
+
stream(:zip).
|
|
308
|
+
each(:hash, format: :json) do |row|
|
|
309
|
+
p row
|
|
310
|
+
end
|
|
311
|
+
~~~
|
|
312
|
+
|
|
313
|
+
Alternatively if the original file name is available it can also be supplied allowing IOStreams to infer the above streams:
|
|
314
|
+
~~~ruby
|
|
315
|
+
IOStreams.path("tempfile2527").
|
|
316
|
+
file_name("file.json.zip").
|
|
317
|
+
each(:hash) do |row|
|
|
318
|
+
p row
|
|
319
|
+
end
|
|
320
|
+
~~~
|
|
321
|
+
|
|
322
|
+
To see the what value was previously set for a particular stream:
|
|
323
|
+
~~~ruby
|
|
324
|
+
path = IOStreams.path("tempfile2527")
|
|
325
|
+
path.stream(:zip)
|
|
326
|
+
path.stream(:pgp, passphrase: "receiver_passphrase")
|
|
327
|
+
|
|
328
|
+
path.setting(:pgp)
|
|
329
|
+
# => {:passphrase=>"receiver_passphrase"}
|
|
330
|
+
~~~
|
|
331
|
+
|
|
332
|
+
To ensure no streams are inferred or applied use stream `:none`
|
|
333
|
+
~~~ruby
|
|
334
|
+
path = IOStreams.path("file.zip")
|
|
335
|
+
path.stream(:none)
|
|
336
|
+
path.read
|
|
337
|
+
~~~
|