iostreams 2.0.0 → 2.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (50) hide show
  1. checksums.yaml +4 -4
  2. data/README.md +5 -22
  3. data/Rakefile +45 -0
  4. data/docs/CLAUDE.md +9 -0
  5. data/docs/config.md +157 -0
  6. data/docs/copy_files.md +75 -0
  7. data/docs/extensions.md +111 -0
  8. data/docs/formats.md +188 -0
  9. data/docs/index.md +388 -0
  10. data/docs/path.md +652 -0
  11. data/docs/pgp.md +436 -0
  12. data/docs/streams.md +337 -0
  13. data/docs/tutorial.md +483 -0
  14. data/docs/upgrading.md +217 -0
  15. data/lib/io_streams/builder.rb +62 -2
  16. data/lib/io_streams/bzip2/reader.rb +25 -2
  17. data/lib/io_streams/bzip2/writer.rb +26 -2
  18. data/lib/io_streams/encode/reader.rb +4 -0
  19. data/lib/io_streams/encode/writer.rb +4 -0
  20. data/lib/io_streams/errors.rb +4 -0
  21. data/lib/io_streams/gzip/reader.rb +4 -0
  22. data/lib/io_streams/gzip/writer.rb +11 -2
  23. data/lib/io_streams/io_streams.rb +111 -1
  24. data/lib/io_streams/line/reader.rb +7 -2
  25. data/lib/io_streams/path.rb +115 -6
  26. data/lib/io_streams/paths/file.rb +47 -1
  27. data/lib/io_streams/paths/http.rb +45 -4
  28. data/lib/io_streams/paths/s3.rb +66 -15
  29. data/lib/io_streams/paths/sftp/net_ssh.rb +104 -0
  30. data/lib/io_streams/paths/sftp.rb +97 -57
  31. data/lib/io_streams/pgp/reader.rb +42 -2
  32. data/lib/io_streams/pgp/writer.rb +26 -6
  33. data/lib/io_streams/pgp.rb +78 -21
  34. data/lib/io_streams/reader.rb +10 -1
  35. data/lib/io_streams/record/reader.rb +72 -2
  36. data/lib/io_streams/stream.rb +12 -7
  37. data/lib/io_streams/symmetric_encryption/reader.rb +4 -0
  38. data/lib/io_streams/symmetric_encryption/writer.rb +4 -0
  39. data/lib/io_streams/tabular/header.rb +31 -4
  40. data/lib/io_streams/tabular/parser/base.rb +10 -0
  41. data/lib/io_streams/tabular/parser/csv.rb +5 -0
  42. data/lib/io_streams/tabular/parser/fixed.rb +3 -1
  43. data/lib/io_streams/tabular/parser/psv.rb +6 -2
  44. data/lib/io_streams/utils.rb +31 -0
  45. data/lib/io_streams/version.rb +1 -1
  46. data/lib/io_streams/writer.rb +10 -1
  47. data/lib/io_streams/xlsx/reader.rb +5 -1
  48. data/lib/io_streams/zip/reader.rb +4 -0
  49. data/lib/io_streams/zip/writer.rb +4 -0
  50. metadata +24 -7
data/docs/index.md ADDED
@@ -0,0 +1,388 @@
1
+ ---
2
+ layout: default
3
+ heading: What is IOStreams?
4
+ ---
5
+
6
+ IOStreams is a streaming library for Ruby that makes compression, encryption, file format, and
7
+ storage location transparent to your code. Read and write files as if they were plain, local text,
8
+ whether they are gzip, zip, or PGP encrypted, and whether they live on local disk, AWS S3, SFTP,
9
+ or are fetched over HTTP.
10
+
11
+ ## Why IOStreams?
12
+
13
+ Processing files in Ruby usually means writing different code for every variation. One customer
14
+ sends a gzip compressed CSV, another sends a PGP encrypted file, a third uploads an Excel
15
+ spreadsheet. The files sit on local disk in development, but in AWS S3 in production. Each
16
+ combination needs its own handling, and reading a large file into memory all at once risks
17
+ exhausting it.
18
+
19
+ Consider reading a gzip compressed CSV file from S3, one record at a time, _without_ IOStreams:
20
+
21
+ ~~~ruby
22
+ require "aws-sdk-s3"
23
+ require "zlib"
24
+ require "csv"
25
+
26
+ response = Aws::S3::Client.new.get_object(bucket: "my-bucket", key: "data.csv.gz")
27
+ gz = Zlib::GzipReader.new(response.body)
28
+ headers = nil
29
+ gz.each_line do |line|
30
+ row = CSV.parse_line(line)
31
+ if headers.nil?
32
+ headers = row
33
+ else
34
+ record = headers.zip(row).to_h
35
+ # ... process record ...
36
+ end
37
+ end
38
+ ~~~
39
+
40
+ Switch that file to plain CSV, to PGP encrypted, or move it back to local disk, and this code has to
41
+ change every time. With IOStreams the same single line handles all of them:
42
+
43
+ ~~~ruby
44
+ IOStreams.path("s3://my-bucket/data.csv.gz").each(:hash) do |record|
45
+ # ... process record ...
46
+ end
47
+ ~~~
48
+
49
+ IOStreams reads the file name, `data.csv.gz`, infers that it is a gzip compressed CSV, and assembles
50
+ the streaming pipeline to fetch, decompress, and parse it. Point the same code at `data.csv`,
51
+ `data.csv.pgp`, or a local path instead, and nothing else changes.
52
+
53
+ ### One API, whatever the format, compression, or encryption
54
+
55
+ IOStreams detects the file type from its extensions and applies the matching streams in order, so
56
+ `sample.csv.gz.pgp` is decrypted, then decompressed, then parsed as CSV without a line of special
57
+ handling. The same code that reads a plain CSV also reads an Excel spreadsheet, a PGP encrypted JSON
58
+ file, or a pipe separated file. A single background job can ingest files from many senders, in many
59
+ formats, with no per-format code.
60
+
61
+ ### One API, wherever the file is stored
62
+
63
+ Local disk, AWS S3, Google Cloud Storage, SFTP, and HTTP(S) all share the same `IOStreams.path`
64
+ interface. The only thing that changes between them is the file name, or nothing at all when you
65
+ [configure roots](config).
66
+
67
+ ### Constant memory, even for huge files
68
+
69
+ Everything is streamed a block at a time, so a 10 GB compressed file uses about as much memory as a
70
+ 10 KB one. Scaling up to large files is trivial: the code that processes ten rows processes ten
71
+ million without changes.
72
+
73
+ ### Configure storage once, switch environments with no code change
74
+
75
+ With [roots](config), the storage location lives in a startup initializer instead of being scattered
76
+ through the code. Point the `:default` root at the local file system in development and at an S3
77
+ bucket in production, and the exact same application code runs in both.
78
+
79
+ ### Capabilities
80
+
81
+ * Low memory utilization, even when processing very large files.
82
+ * Parse JSON, CSV, PSV, or fixed width data on the fly.
83
+ * Encrypt / Decrypt data on the fly.
84
+ * Compress / Decompress data on the fly.
85
+ * Change storage location / mechanism transparently without any code changes.
86
+
87
+ #### File Extensions
88
+ * Zip
89
+ * Gzip
90
+ * BZip2
91
+ * PGP (Requires GnuPG)
92
+ * Xlsx (Reading)
93
+ * Encryption using [Symmetric Encryption](https://encryption.reidmorrison.com/)
94
+
95
+ #### File Storage
96
+ * File
97
+ * AWS S3
98
+ * Google Cloud Storage (Using the AWS S3 Client)
99
+ * SFTP
100
+ * HTTP(S) (Read only)
101
+
102
+ #### File formats
103
+ * CSV
104
+ * Fixed width formats
105
+ * JSON
106
+ * PSV
107
+
108
+ ## Example usages
109
+
110
+ ### Creating files
111
+
112
+ Write a string to a local file called `sample.txt`:
113
+
114
+ ~~~ruby
115
+ path = IOStreams.path("sample.txt")
116
+ path.write("Hello World")
117
+ ~~~
118
+
119
+ Write a string to AWS S3, storing in the S3 bucket `sample-bucket`, under the path `demo` with a file name of `sample.txt`.
120
+
121
+ ~~~ruby
122
+ path = IOStreams.path("s3://sample-bucket/demo/sample.txt")
123
+ path.write("Hello World")
124
+ ~~~
125
+
126
+ Write a string into a compressed file by adding the `.gz` extension to the file name:
127
+
128
+ ~~~ruby
129
+ path = IOStreams.path("sample.txt.gz")
130
+ path.write("Hello World")
131
+ ~~~
132
+
133
+ Compress and encrypt the data into a PGP encrypted file, called `sample.txt.pgp`:
134
+
135
+ ~~~ruby
136
+ path = IOStreams.path("sample.txt.pgp")
137
+ # Recipient that can decrypt this file:
138
+ path.option(:pgp, recipient: "receiver@example.org")
139
+ path.write("Hello World")
140
+ ~~~
141
+
142
+ Note: GnuPG needs to be installed locally for the above PGP example to work.
143
+
144
+ Write a string to a SFTP server, with a host name of `example.org`, under the path `demo`,
145
+ with a file name of `sample.txt`. Adds the optional `username` and `password`.
146
+
147
+ ~~~ruby
148
+ path = IOStreams.path("sftp://example.org/demo/sample.txt",
149
+ username: "example",
150
+ password: "topsecret")
151
+ path.write("Hello World")
152
+ ~~~
153
+
154
+ Write a string to AWS S3, storing in the S3 bucket `sample-bucket`, under the path `demo`,
155
+ encrypted with pgp, with a file name of `sample.txt.pgp`.
156
+
157
+ ~~~ruby
158
+ path = IOStreams.path("s3://sample-bucket/demo/sample.txt.pgp")
159
+ path.option(:pgp, recipient: "receiver@example.org")
160
+ path.write("Hello World")
161
+ ~~~
162
+
163
+ ### Reading files
164
+
165
+ Read an entire local file called `sample.txt`, into a string:
166
+
167
+ ~~~ruby
168
+ path = IOStreams.path("sample.txt")
169
+ path.read
170
+ # => "Hello World"
171
+ ~~~
172
+
173
+ Read an entire file called `sample.txt`, into a string, from the S3 bucket `sample-bucket`, under the path `demo`:
174
+
175
+ ~~~ruby
176
+ path = IOStreams.path("s3://sample-bucket/demo/sample.txt")
177
+ path.read
178
+ # => "Hello World"
179
+ ~~~
180
+
181
+ Read an entire local file called `sample.txt.gz`, and decompress the contents into a string:
182
+
183
+ ~~~ruby
184
+ path = IOStreams.path("sample.txt.gz")
185
+ path.read
186
+ # => "Hello World"
187
+ ~~~
188
+
189
+ Read an entire local file called `sample.txt.pgp`, decompress, and decrypt the contents into a string:
190
+
191
+ ~~~ruby
192
+ path = IOStreams.path("sample.txt.pgp")
193
+ path.read
194
+ # => "Hello World"
195
+ ~~~
196
+
197
+ Notes:
198
+ * GnuPG needs to be installed locally for the above PGP example to work.
199
+
200
+ ## Streaming Examples
201
+
202
+ When dealing with large files it is important _not_ to load the entire file into memory.
203
+ Efficiently read the files data in chunks / lines / records.
204
+
205
+ Read 128 characters at a time from a file:
206
+ ~~~ruby
207
+ path = IOStreams.path("sample.txt")
208
+ path.reader do |io|
209
+ while (data = io.read(128))
210
+ p data
211
+ end
212
+ end
213
+ ~~~
214
+
215
+ Read one line at a time from the file:
216
+ ~~~ruby
217
+ path = IOStreams.path("sample.txt")
218
+ path.each do |line|
219
+ puts line
220
+ end
221
+ ~~~
222
+
223
+ Write data to the file.
224
+ ~~~ruby
225
+ path = IOStreams.path("sample.txt")
226
+ path.writer do |io|
227
+ io << "This"
228
+ io << " is "
229
+ io << " one line\n"
230
+ end
231
+ ~~~
232
+
233
+ Write lines to the file. By adding `:line` to `writer`, each write appends a new line character.
234
+ ~~~ruby
235
+ path = IOStreams.path("sample.txt")
236
+ path.writer(:line) do |file|
237
+ file << "these"
238
+ file << "are"
239
+ file << "all"
240
+ file << "separate"
241
+ file << "lines"
242
+ end
243
+ ~~~
244
+
245
+ ### Reading CSV Files
246
+
247
+ Example CSV file, `example.csv`:
248
+
249
+ ~~~csv
250
+ name,address,zip_code
251
+ Jack,There,1234
252
+ Joe,Over There somewhere,1234
253
+ ~~~
254
+
255
+ Read each line from the CSV file as lines of strings:
256
+ ~~~ruby
257
+ path = IOStreams.path("example.csv")
258
+ path.each do |line|
259
+ p line
260
+ end
261
+ ~~~
262
+
263
+ Output:
264
+ ~~~ruby
265
+ "name,address,zip_code"
266
+ "Jack,There,1234"
267
+ "Joe,Over There somewhere,1234"
268
+ ~~~
269
+
270
+ Read each row from the CSV file as arrays:
271
+ ~~~ruby
272
+ path = IOStreams.path("example.csv")
273
+ path.each(:array) do |array|
274
+ p array
275
+ end
276
+ ~~~
277
+
278
+ Output:
279
+ ~~~ruby
280
+ ["name", "address", "zip_code"]
281
+ ["Jack", "There", "1234"]
282
+ ["Joe", "Over There somewhere", "1234"]
283
+ ~~~
284
+
285
+ Read each row from a csv file as key-value pairs, where the key is the CSV column header, and the value is the value for that row.
286
+ ~~~ruby
287
+ path = IOStreams.path("example.csv")
288
+ path.each(:hash) do |record|
289
+ p record
290
+ end
291
+ ~~~
292
+
293
+ Output:
294
+
295
+ ~~~ruby
296
+ {"name"=>"Jack", "address"=>"There", "zip_code"=>"1234"}
297
+ {"name"=>"Joe", "address"=>"Over There somewhere", "zip_code"=>"1234"}
298
+ ~~~
299
+
300
+ ### Writing CSV Files
301
+
302
+ Write an array (row) at a time to the file.
303
+ Each array is converted to csv before being written to the file.
304
+
305
+ ~~~ruby
306
+ IOStreams.path("example.csv").writer(:array) do |io|
307
+ io << ["name", "address", "zip_code"]
308
+ io << ["Jack", "There", "1234"]
309
+ io << ["Joe", "Over There somewhere", 1234]
310
+ end
311
+ ~~~
312
+
313
+ Write a hash (record) at a time to the file.
314
+ Each hash is converted to csv before being written to the file.
315
+ The header row is extracted from the first hash write that is performed.
316
+
317
+ ~~~ruby
318
+ path = IOStreams.path("example.csv")
319
+ path.writer(:hash) do |stream|
320
+ stream << {name: "Jack", address: "There", zip_code: 1234}
321
+ stream << {zip_code: 1234, address: "Over There somewhere", name: "Joe"}
322
+ end
323
+ ~~~
324
+
325
+ This time write the CSV data to a compressed zip file, by adding `.zip` to the file name.
326
+
327
+ ~~~ruby
328
+ path = IOStreams.path("example.csv.zip")
329
+ path.writer(:hash) do |stream|
330
+ stream << {name: "Jack", address: "There", zip_code: 1234}
331
+ stream << {zip_code: 1234, address: "Over There somewhere", name: "Joe"}
332
+ end
333
+ ~~~
334
+
335
+ Changing the file name to change its compression, encryption, or even whether it is local or remote
336
+ has no effect on the code reading from or writing to the path.
337
+
338
+ ## PSV Files
339
+
340
+ PSV files are faster than CSV files, since CSV files have complex rules for dealing with embedded quotes and newlines.
341
+
342
+ PSV files in IOStreams follow the following simple rules:
343
+ * Values are delimited using `|`.
344
+ * Rows are delimeted with new lines.
345
+ * Values may _not_ contain `|`, or new lines.
346
+
347
+ Example PSV file, `example.psv`:
348
+
349
+ ~~~csv
350
+ name|address|zip_code
351
+ Jack|There|1234
352
+ Joe|Over There somewhere|1234
353
+ ~~~
354
+
355
+ ### Reading PSV Files
356
+
357
+ Read each row from a psv file as key-value pairs, where the key is the PSV column header, and the value is the value for that row.
358
+ ~~~ruby
359
+ path = IOStreams.path("example.psv")
360
+ path.each(:hash) do |record|
361
+ p record
362
+ end
363
+ ~~~
364
+
365
+ Output:
366
+
367
+ ~~~ruby
368
+ {"name"=>"Jack", "address"=>"There", "zip_code"=>"1234"}
369
+ {"name"=>"Joe", "address"=>"Over There somewhere", "zip_code"=>"1234"}
370
+ ~~~
371
+
372
+ ### Writing PSV Files
373
+
374
+ Write a hash (record) at a time to the file.
375
+ Each hash is converted to psv before being written to the file.
376
+ The header row is extracted from the first hash write that is performed.
377
+
378
+ ~~~ruby
379
+ path = IOStreams.path("example.psv")
380
+ path.writer(:hash) do |stream|
381
+ stream << {name: "Jack", address: "There", zip_code: 1234}
382
+ stream << {zip_code: 1234, address: "Over There somewhere", name: "Joe"}
383
+ end
384
+ ~~~
385
+
386
+ ## Getting Started
387
+
388
+ Start with the [IOStreams tutorial](tutorial) for a great introduction to IOStreams.