iostreams 2.0.0 → 2.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.md +5 -22
- data/Rakefile +45 -0
- data/docs/CLAUDE.md +9 -0
- data/docs/config.md +157 -0
- data/docs/copy_files.md +75 -0
- data/docs/extensions.md +111 -0
- data/docs/formats.md +188 -0
- data/docs/index.md +388 -0
- data/docs/path.md +652 -0
- data/docs/pgp.md +436 -0
- data/docs/streams.md +337 -0
- data/docs/tutorial.md +483 -0
- data/docs/upgrading.md +217 -0
- data/lib/io_streams/builder.rb +62 -2
- data/lib/io_streams/bzip2/reader.rb +25 -2
- data/lib/io_streams/bzip2/writer.rb +26 -2
- data/lib/io_streams/encode/reader.rb +4 -0
- data/lib/io_streams/encode/writer.rb +4 -0
- data/lib/io_streams/errors.rb +4 -0
- data/lib/io_streams/gzip/reader.rb +4 -0
- data/lib/io_streams/gzip/writer.rb +11 -2
- data/lib/io_streams/io_streams.rb +111 -1
- data/lib/io_streams/line/reader.rb +7 -2
- data/lib/io_streams/path.rb +115 -6
- data/lib/io_streams/paths/file.rb +47 -1
- data/lib/io_streams/paths/http.rb +45 -4
- data/lib/io_streams/paths/s3.rb +66 -15
- data/lib/io_streams/paths/sftp/net_ssh.rb +104 -0
- data/lib/io_streams/paths/sftp.rb +97 -57
- data/lib/io_streams/pgp/reader.rb +42 -2
- data/lib/io_streams/pgp/writer.rb +26 -6
- data/lib/io_streams/pgp.rb +78 -21
- data/lib/io_streams/reader.rb +10 -1
- data/lib/io_streams/record/reader.rb +72 -2
- data/lib/io_streams/stream.rb +12 -7
- data/lib/io_streams/symmetric_encryption/reader.rb +4 -0
- data/lib/io_streams/symmetric_encryption/writer.rb +4 -0
- data/lib/io_streams/tabular/header.rb +31 -4
- data/lib/io_streams/tabular/parser/base.rb +10 -0
- data/lib/io_streams/tabular/parser/csv.rb +5 -0
- data/lib/io_streams/tabular/parser/fixed.rb +3 -1
- data/lib/io_streams/tabular/parser/psv.rb +6 -2
- data/lib/io_streams/utils.rb +31 -0
- data/lib/io_streams/version.rb +1 -1
- data/lib/io_streams/writer.rb +10 -1
- data/lib/io_streams/xlsx/reader.rb +5 -1
- data/lib/io_streams/zip/reader.rb +4 -0
- data/lib/io_streams/zip/writer.rb +4 -0
- metadata +24 -7
data/docs/index.md
ADDED
|
@@ -0,0 +1,388 @@
|
|
|
1
|
+
---
|
|
2
|
+
layout: default
|
|
3
|
+
heading: What is IOStreams?
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
IOStreams is a streaming library for Ruby that makes compression, encryption, file format, and
|
|
7
|
+
storage location transparent to your code. Read and write files as if they were plain, local text,
|
|
8
|
+
whether they are gzip, zip, or PGP encrypted, and whether they live on local disk, AWS S3, SFTP,
|
|
9
|
+
or are fetched over HTTP.
|
|
10
|
+
|
|
11
|
+
## Why IOStreams?
|
|
12
|
+
|
|
13
|
+
Processing files in Ruby usually means writing different code for every variation. One customer
|
|
14
|
+
sends a gzip compressed CSV, another sends a PGP encrypted file, a third uploads an Excel
|
|
15
|
+
spreadsheet. The files sit on local disk in development, but in AWS S3 in production. Each
|
|
16
|
+
combination needs its own handling, and reading a large file into memory all at once risks
|
|
17
|
+
exhausting it.
|
|
18
|
+
|
|
19
|
+
Consider reading a gzip compressed CSV file from S3, one record at a time, _without_ IOStreams:
|
|
20
|
+
|
|
21
|
+
~~~ruby
|
|
22
|
+
require "aws-sdk-s3"
|
|
23
|
+
require "zlib"
|
|
24
|
+
require "csv"
|
|
25
|
+
|
|
26
|
+
response = Aws::S3::Client.new.get_object(bucket: "my-bucket", key: "data.csv.gz")
|
|
27
|
+
gz = Zlib::GzipReader.new(response.body)
|
|
28
|
+
headers = nil
|
|
29
|
+
gz.each_line do |line|
|
|
30
|
+
row = CSV.parse_line(line)
|
|
31
|
+
if headers.nil?
|
|
32
|
+
headers = row
|
|
33
|
+
else
|
|
34
|
+
record = headers.zip(row).to_h
|
|
35
|
+
# ... process record ...
|
|
36
|
+
end
|
|
37
|
+
end
|
|
38
|
+
~~~
|
|
39
|
+
|
|
40
|
+
Switch that file to plain CSV, to PGP encrypted, or move it back to local disk, and this code has to
|
|
41
|
+
change every time. With IOStreams the same single line handles all of them:
|
|
42
|
+
|
|
43
|
+
~~~ruby
|
|
44
|
+
IOStreams.path("s3://my-bucket/data.csv.gz").each(:hash) do |record|
|
|
45
|
+
# ... process record ...
|
|
46
|
+
end
|
|
47
|
+
~~~
|
|
48
|
+
|
|
49
|
+
IOStreams reads the file name, `data.csv.gz`, infers that it is a gzip compressed CSV, and assembles
|
|
50
|
+
the streaming pipeline to fetch, decompress, and parse it. Point the same code at `data.csv`,
|
|
51
|
+
`data.csv.pgp`, or a local path instead, and nothing else changes.
|
|
52
|
+
|
|
53
|
+
### One API, whatever the format, compression, or encryption
|
|
54
|
+
|
|
55
|
+
IOStreams detects the file type from its extensions and applies the matching streams in order, so
|
|
56
|
+
`sample.csv.gz.pgp` is decrypted, then decompressed, then parsed as CSV without a line of special
|
|
57
|
+
handling. The same code that reads a plain CSV also reads an Excel spreadsheet, a PGP encrypted JSON
|
|
58
|
+
file, or a pipe separated file. A single background job can ingest files from many senders, in many
|
|
59
|
+
formats, with no per-format code.
|
|
60
|
+
|
|
61
|
+
### One API, wherever the file is stored
|
|
62
|
+
|
|
63
|
+
Local disk, AWS S3, Google Cloud Storage, SFTP, and HTTP(S) all share the same `IOStreams.path`
|
|
64
|
+
interface. The only thing that changes between them is the file name, or nothing at all when you
|
|
65
|
+
[configure roots](config).
|
|
66
|
+
|
|
67
|
+
### Constant memory, even for huge files
|
|
68
|
+
|
|
69
|
+
Everything is streamed a block at a time, so a 10 GB compressed file uses about as much memory as a
|
|
70
|
+
10 KB one. Scaling up to large files is trivial: the code that processes ten rows processes ten
|
|
71
|
+
million without changes.
|
|
72
|
+
|
|
73
|
+
### Configure storage once, switch environments with no code change
|
|
74
|
+
|
|
75
|
+
With [roots](config), the storage location lives in a startup initializer instead of being scattered
|
|
76
|
+
through the code. Point the `:default` root at the local file system in development and at an S3
|
|
77
|
+
bucket in production, and the exact same application code runs in both.
|
|
78
|
+
|
|
79
|
+
### Capabilities
|
|
80
|
+
|
|
81
|
+
* Low memory utilization, even when processing very large files.
|
|
82
|
+
* Parse JSON, CSV, PSV, or fixed width data on the fly.
|
|
83
|
+
* Encrypt / Decrypt data on the fly.
|
|
84
|
+
* Compress / Decompress data on the fly.
|
|
85
|
+
* Change storage location / mechanism transparently without any code changes.
|
|
86
|
+
|
|
87
|
+
#### File Extensions
|
|
88
|
+
* Zip
|
|
89
|
+
* Gzip
|
|
90
|
+
* BZip2
|
|
91
|
+
* PGP (Requires GnuPG)
|
|
92
|
+
* Xlsx (Reading)
|
|
93
|
+
* Encryption using [Symmetric Encryption](https://encryption.reidmorrison.com/)
|
|
94
|
+
|
|
95
|
+
#### File Storage
|
|
96
|
+
* File
|
|
97
|
+
* AWS S3
|
|
98
|
+
* Google Cloud Storage (Using the AWS S3 Client)
|
|
99
|
+
* SFTP
|
|
100
|
+
* HTTP(S) (Read only)
|
|
101
|
+
|
|
102
|
+
#### File formats
|
|
103
|
+
* CSV
|
|
104
|
+
* Fixed width formats
|
|
105
|
+
* JSON
|
|
106
|
+
* PSV
|
|
107
|
+
|
|
108
|
+
## Example usages
|
|
109
|
+
|
|
110
|
+
### Creating files
|
|
111
|
+
|
|
112
|
+
Write a string to a local file called `sample.txt`:
|
|
113
|
+
|
|
114
|
+
~~~ruby
|
|
115
|
+
path = IOStreams.path("sample.txt")
|
|
116
|
+
path.write("Hello World")
|
|
117
|
+
~~~
|
|
118
|
+
|
|
119
|
+
Write a string to AWS S3, storing in the S3 bucket `sample-bucket`, under the path `demo` with a file name of `sample.txt`.
|
|
120
|
+
|
|
121
|
+
~~~ruby
|
|
122
|
+
path = IOStreams.path("s3://sample-bucket/demo/sample.txt")
|
|
123
|
+
path.write("Hello World")
|
|
124
|
+
~~~
|
|
125
|
+
|
|
126
|
+
Write a string into a compressed file by adding the `.gz` extension to the file name:
|
|
127
|
+
|
|
128
|
+
~~~ruby
|
|
129
|
+
path = IOStreams.path("sample.txt.gz")
|
|
130
|
+
path.write("Hello World")
|
|
131
|
+
~~~
|
|
132
|
+
|
|
133
|
+
Compress and encrypt the data into a PGP encrypted file, called `sample.txt.pgp`:
|
|
134
|
+
|
|
135
|
+
~~~ruby
|
|
136
|
+
path = IOStreams.path("sample.txt.pgp")
|
|
137
|
+
# Recipient that can decrypt this file:
|
|
138
|
+
path.option(:pgp, recipient: "receiver@example.org")
|
|
139
|
+
path.write("Hello World")
|
|
140
|
+
~~~
|
|
141
|
+
|
|
142
|
+
Note: GnuPG needs to be installed locally for the above PGP example to work.
|
|
143
|
+
|
|
144
|
+
Write a string to a SFTP server, with a host name of `example.org`, under the path `demo`,
|
|
145
|
+
with a file name of `sample.txt`. Adds the optional `username` and `password`.
|
|
146
|
+
|
|
147
|
+
~~~ruby
|
|
148
|
+
path = IOStreams.path("sftp://example.org/demo/sample.txt",
|
|
149
|
+
username: "example",
|
|
150
|
+
password: "topsecret")
|
|
151
|
+
path.write("Hello World")
|
|
152
|
+
~~~
|
|
153
|
+
|
|
154
|
+
Write a string to AWS S3, storing in the S3 bucket `sample-bucket`, under the path `demo`,
|
|
155
|
+
encrypted with pgp, with a file name of `sample.txt.pgp`.
|
|
156
|
+
|
|
157
|
+
~~~ruby
|
|
158
|
+
path = IOStreams.path("s3://sample-bucket/demo/sample.txt.pgp")
|
|
159
|
+
path.option(:pgp, recipient: "receiver@example.org")
|
|
160
|
+
path.write("Hello World")
|
|
161
|
+
~~~
|
|
162
|
+
|
|
163
|
+
### Reading files
|
|
164
|
+
|
|
165
|
+
Read an entire local file called `sample.txt`, into a string:
|
|
166
|
+
|
|
167
|
+
~~~ruby
|
|
168
|
+
path = IOStreams.path("sample.txt")
|
|
169
|
+
path.read
|
|
170
|
+
# => "Hello World"
|
|
171
|
+
~~~
|
|
172
|
+
|
|
173
|
+
Read an entire file called `sample.txt`, into a string, from the S3 bucket `sample-bucket`, under the path `demo`:
|
|
174
|
+
|
|
175
|
+
~~~ruby
|
|
176
|
+
path = IOStreams.path("s3://sample-bucket/demo/sample.txt")
|
|
177
|
+
path.read
|
|
178
|
+
# => "Hello World"
|
|
179
|
+
~~~
|
|
180
|
+
|
|
181
|
+
Read an entire local file called `sample.txt.gz`, and decompress the contents into a string:
|
|
182
|
+
|
|
183
|
+
~~~ruby
|
|
184
|
+
path = IOStreams.path("sample.txt.gz")
|
|
185
|
+
path.read
|
|
186
|
+
# => "Hello World"
|
|
187
|
+
~~~
|
|
188
|
+
|
|
189
|
+
Read an entire local file called `sample.txt.pgp`, decompress, and decrypt the contents into a string:
|
|
190
|
+
|
|
191
|
+
~~~ruby
|
|
192
|
+
path = IOStreams.path("sample.txt.pgp")
|
|
193
|
+
path.read
|
|
194
|
+
# => "Hello World"
|
|
195
|
+
~~~
|
|
196
|
+
|
|
197
|
+
Notes:
|
|
198
|
+
* GnuPG needs to be installed locally for the above PGP example to work.
|
|
199
|
+
|
|
200
|
+
## Streaming Examples
|
|
201
|
+
|
|
202
|
+
When dealing with large files it is important _not_ to load the entire file into memory.
|
|
203
|
+
Efficiently read the files data in chunks / lines / records.
|
|
204
|
+
|
|
205
|
+
Read 128 characters at a time from a file:
|
|
206
|
+
~~~ruby
|
|
207
|
+
path = IOStreams.path("sample.txt")
|
|
208
|
+
path.reader do |io|
|
|
209
|
+
while (data = io.read(128))
|
|
210
|
+
p data
|
|
211
|
+
end
|
|
212
|
+
end
|
|
213
|
+
~~~
|
|
214
|
+
|
|
215
|
+
Read one line at a time from the file:
|
|
216
|
+
~~~ruby
|
|
217
|
+
path = IOStreams.path("sample.txt")
|
|
218
|
+
path.each do |line|
|
|
219
|
+
puts line
|
|
220
|
+
end
|
|
221
|
+
~~~
|
|
222
|
+
|
|
223
|
+
Write data to the file.
|
|
224
|
+
~~~ruby
|
|
225
|
+
path = IOStreams.path("sample.txt")
|
|
226
|
+
path.writer do |io|
|
|
227
|
+
io << "This"
|
|
228
|
+
io << " is "
|
|
229
|
+
io << " one line\n"
|
|
230
|
+
end
|
|
231
|
+
~~~
|
|
232
|
+
|
|
233
|
+
Write lines to the file. By adding `:line` to `writer`, each write appends a new line character.
|
|
234
|
+
~~~ruby
|
|
235
|
+
path = IOStreams.path("sample.txt")
|
|
236
|
+
path.writer(:line) do |file|
|
|
237
|
+
file << "these"
|
|
238
|
+
file << "are"
|
|
239
|
+
file << "all"
|
|
240
|
+
file << "separate"
|
|
241
|
+
file << "lines"
|
|
242
|
+
end
|
|
243
|
+
~~~
|
|
244
|
+
|
|
245
|
+
### Reading CSV Files
|
|
246
|
+
|
|
247
|
+
Example CSV file, `example.csv`:
|
|
248
|
+
|
|
249
|
+
~~~csv
|
|
250
|
+
name,address,zip_code
|
|
251
|
+
Jack,There,1234
|
|
252
|
+
Joe,Over There somewhere,1234
|
|
253
|
+
~~~
|
|
254
|
+
|
|
255
|
+
Read each line from the CSV file as lines of strings:
|
|
256
|
+
~~~ruby
|
|
257
|
+
path = IOStreams.path("example.csv")
|
|
258
|
+
path.each do |line|
|
|
259
|
+
p line
|
|
260
|
+
end
|
|
261
|
+
~~~
|
|
262
|
+
|
|
263
|
+
Output:
|
|
264
|
+
~~~ruby
|
|
265
|
+
"name,address,zip_code"
|
|
266
|
+
"Jack,There,1234"
|
|
267
|
+
"Joe,Over There somewhere,1234"
|
|
268
|
+
~~~
|
|
269
|
+
|
|
270
|
+
Read each row from the CSV file as arrays:
|
|
271
|
+
~~~ruby
|
|
272
|
+
path = IOStreams.path("example.csv")
|
|
273
|
+
path.each(:array) do |array|
|
|
274
|
+
p array
|
|
275
|
+
end
|
|
276
|
+
~~~
|
|
277
|
+
|
|
278
|
+
Output:
|
|
279
|
+
~~~ruby
|
|
280
|
+
["name", "address", "zip_code"]
|
|
281
|
+
["Jack", "There", "1234"]
|
|
282
|
+
["Joe", "Over There somewhere", "1234"]
|
|
283
|
+
~~~
|
|
284
|
+
|
|
285
|
+
Read each row from a csv file as key-value pairs, where the key is the CSV column header, and the value is the value for that row.
|
|
286
|
+
~~~ruby
|
|
287
|
+
path = IOStreams.path("example.csv")
|
|
288
|
+
path.each(:hash) do |record|
|
|
289
|
+
p record
|
|
290
|
+
end
|
|
291
|
+
~~~
|
|
292
|
+
|
|
293
|
+
Output:
|
|
294
|
+
|
|
295
|
+
~~~ruby
|
|
296
|
+
{"name"=>"Jack", "address"=>"There", "zip_code"=>"1234"}
|
|
297
|
+
{"name"=>"Joe", "address"=>"Over There somewhere", "zip_code"=>"1234"}
|
|
298
|
+
~~~
|
|
299
|
+
|
|
300
|
+
### Writing CSV Files
|
|
301
|
+
|
|
302
|
+
Write an array (row) at a time to the file.
|
|
303
|
+
Each array is converted to csv before being written to the file.
|
|
304
|
+
|
|
305
|
+
~~~ruby
|
|
306
|
+
IOStreams.path("example.csv").writer(:array) do |io|
|
|
307
|
+
io << ["name", "address", "zip_code"]
|
|
308
|
+
io << ["Jack", "There", "1234"]
|
|
309
|
+
io << ["Joe", "Over There somewhere", 1234]
|
|
310
|
+
end
|
|
311
|
+
~~~
|
|
312
|
+
|
|
313
|
+
Write a hash (record) at a time to the file.
|
|
314
|
+
Each hash is converted to csv before being written to the file.
|
|
315
|
+
The header row is extracted from the first hash write that is performed.
|
|
316
|
+
|
|
317
|
+
~~~ruby
|
|
318
|
+
path = IOStreams.path("example.csv")
|
|
319
|
+
path.writer(:hash) do |stream|
|
|
320
|
+
stream << {name: "Jack", address: "There", zip_code: 1234}
|
|
321
|
+
stream << {zip_code: 1234, address: "Over There somewhere", name: "Joe"}
|
|
322
|
+
end
|
|
323
|
+
~~~
|
|
324
|
+
|
|
325
|
+
This time write the CSV data to a compressed zip file, by adding `.zip` to the file name.
|
|
326
|
+
|
|
327
|
+
~~~ruby
|
|
328
|
+
path = IOStreams.path("example.csv.zip")
|
|
329
|
+
path.writer(:hash) do |stream|
|
|
330
|
+
stream << {name: "Jack", address: "There", zip_code: 1234}
|
|
331
|
+
stream << {zip_code: 1234, address: "Over There somewhere", name: "Joe"}
|
|
332
|
+
end
|
|
333
|
+
~~~
|
|
334
|
+
|
|
335
|
+
Changing the file name to change its compression, encryption, or even whether it is local or remote
|
|
336
|
+
has no effect on the code reading from or writing to the path.
|
|
337
|
+
|
|
338
|
+
## PSV Files
|
|
339
|
+
|
|
340
|
+
PSV files are faster than CSV files, since CSV files have complex rules for dealing with embedded quotes and newlines.
|
|
341
|
+
|
|
342
|
+
PSV files in IOStreams follow the following simple rules:
|
|
343
|
+
* Values are delimited using `|`.
|
|
344
|
+
* Rows are delimeted with new lines.
|
|
345
|
+
* Values may _not_ contain `|`, or new lines.
|
|
346
|
+
|
|
347
|
+
Example PSV file, `example.psv`:
|
|
348
|
+
|
|
349
|
+
~~~csv
|
|
350
|
+
name|address|zip_code
|
|
351
|
+
Jack|There|1234
|
|
352
|
+
Joe|Over There somewhere|1234
|
|
353
|
+
~~~
|
|
354
|
+
|
|
355
|
+
### Reading PSV Files
|
|
356
|
+
|
|
357
|
+
Read each row from a psv file as key-value pairs, where the key is the PSV column header, and the value is the value for that row.
|
|
358
|
+
~~~ruby
|
|
359
|
+
path = IOStreams.path("example.psv")
|
|
360
|
+
path.each(:hash) do |record|
|
|
361
|
+
p record
|
|
362
|
+
end
|
|
363
|
+
~~~
|
|
364
|
+
|
|
365
|
+
Output:
|
|
366
|
+
|
|
367
|
+
~~~ruby
|
|
368
|
+
{"name"=>"Jack", "address"=>"There", "zip_code"=>"1234"}
|
|
369
|
+
{"name"=>"Joe", "address"=>"Over There somewhere", "zip_code"=>"1234"}
|
|
370
|
+
~~~
|
|
371
|
+
|
|
372
|
+
### Writing PSV Files
|
|
373
|
+
|
|
374
|
+
Write a hash (record) at a time to the file.
|
|
375
|
+
Each hash is converted to psv before being written to the file.
|
|
376
|
+
The header row is extracted from the first hash write that is performed.
|
|
377
|
+
|
|
378
|
+
~~~ruby
|
|
379
|
+
path = IOStreams.path("example.psv")
|
|
380
|
+
path.writer(:hash) do |stream|
|
|
381
|
+
stream << {name: "Jack", address: "There", zip_code: 1234}
|
|
382
|
+
stream << {zip_code: 1234, address: "Over There somewhere", name: "Joe"}
|
|
383
|
+
end
|
|
384
|
+
~~~
|
|
385
|
+
|
|
386
|
+
## Getting Started
|
|
387
|
+
|
|
388
|
+
Start with the [IOStreams tutorial](tutorial) for a great introduction to IOStreams.
|