iostreams 1.11.0 → 2.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (108) hide show
  1. checksums.yaml +4 -4
  2. data/README.md +14 -13
  3. data/Rakefile +52 -0
  4. data/docs/CLAUDE.md +9 -0
  5. data/docs/config.md +157 -0
  6. data/docs/copy_files.md +75 -0
  7. data/docs/extensions.md +111 -0
  8. data/docs/formats.md +188 -0
  9. data/docs/index.md +388 -0
  10. data/docs/path.md +652 -0
  11. data/docs/pgp.md +436 -0
  12. data/docs/streams.md +337 -0
  13. data/docs/tutorial.md +483 -0
  14. data/docs/upgrading.md +217 -0
  15. data/lib/io_streams/builder.rb +71 -11
  16. data/lib/io_streams/bzip2/reader.rb +25 -2
  17. data/lib/io_streams/bzip2/writer.rb +26 -2
  18. data/lib/io_streams/encode/reader.rb +6 -2
  19. data/lib/io_streams/encode/writer.rb +9 -5
  20. data/lib/io_streams/errors.rb +4 -0
  21. data/lib/io_streams/gzip/reader.rb +5 -1
  22. data/lib/io_streams/gzip/writer.rb +11 -2
  23. data/lib/io_streams/io_streams.rb +156 -20
  24. data/lib/io_streams/line/reader.rb +9 -4
  25. data/lib/io_streams/line/writer.rb +1 -1
  26. data/lib/io_streams/path.rb +117 -8
  27. data/lib/io_streams/paths/file.rb +57 -11
  28. data/lib/io_streams/paths/http.rb +123 -9
  29. data/lib/io_streams/paths/matcher.rb +3 -3
  30. data/lib/io_streams/paths/s3.rb +69 -18
  31. data/lib/io_streams/paths/sftp/net_ssh.rb +104 -0
  32. data/lib/io_streams/paths/sftp.rb +103 -64
  33. data/lib/io_streams/pgp/reader.rb +63 -10
  34. data/lib/io_streams/pgp/writer.rb +111 -30
  35. data/lib/io_streams/pgp.rb +256 -71
  36. data/lib/io_streams/reader.rb +14 -5
  37. data/lib/io_streams/record/reader.rb +75 -6
  38. data/lib/io_streams/record/writer.rb +3 -4
  39. data/lib/io_streams/row/reader.rb +1 -1
  40. data/lib/io_streams/row/writer.rb +1 -1
  41. data/lib/io_streams/stream.rb +48 -37
  42. data/lib/io_streams/symmetric_encryption/reader.rb +6 -2
  43. data/lib/io_streams/symmetric_encryption/writer.rb +8 -4
  44. data/lib/io_streams/tabular/header.rb +49 -10
  45. data/lib/io_streams/tabular/parser/array.rb +0 -10
  46. data/lib/io_streams/tabular/parser/base.rb +10 -0
  47. data/lib/io_streams/tabular/parser/csv.rb +9 -36
  48. data/lib/io_streams/tabular/parser/fixed.rb +8 -6
  49. data/lib/io_streams/tabular/parser/psv.rb +6 -14
  50. data/lib/io_streams/tabular.rb +5 -10
  51. data/lib/io_streams/utils.rb +34 -2
  52. data/lib/io_streams/version.rb +1 -1
  53. data/lib/io_streams/writer.rb +16 -7
  54. data/lib/io_streams/xlsx/reader.rb +6 -2
  55. data/lib/io_streams/zip/reader.rb +4 -0
  56. data/lib/io_streams/zip/writer.rb +26 -10
  57. data/lib/iostreams.rb +0 -1
  58. metadata +46 -112
  59. data/lib/io_streams/deprecated.rb +0 -216
  60. data/lib/io_streams/tabular/utility/csv_row.rb +0 -105
  61. data/test/builder_test.rb +0 -311
  62. data/test/bzip2_reader_test.rb +0 -27
  63. data/test/bzip2_writer_test.rb +0 -56
  64. data/test/deprecated_test.rb +0 -121
  65. data/test/encode_reader_test.rb +0 -51
  66. data/test/encode_writer_test.rb +0 -90
  67. data/test/files/embedded_lines_test.csv +0 -7
  68. data/test/files/multiple_files.zip +0 -0
  69. data/test/files/spreadsheet.xlsx +0 -0
  70. data/test/files/test.csv +0 -4
  71. data/test/files/test.json +0 -3
  72. data/test/files/test.psv +0 -4
  73. data/test/files/text file.txt +0 -3
  74. data/test/files/text.txt +0 -3
  75. data/test/files/text.txt.bz2 +0 -0
  76. data/test/files/text.txt.gz +0 -0
  77. data/test/files/text.txt.gz.zip +0 -0
  78. data/test/files/text.zip +0 -0
  79. data/test/files/text.zip.gz +0 -0
  80. data/test/files/unclosed_quote_large_test.csv +0 -1658
  81. data/test/files/unclosed_quote_test.csv +0 -4
  82. data/test/files/unclosed_quote_test2.csv +0 -3
  83. data/test/gzip_reader_test.rb +0 -27
  84. data/test/gzip_writer_test.rb +0 -52
  85. data/test/io_streams_test.rb +0 -132
  86. data/test/line_reader_test.rb +0 -325
  87. data/test/line_writer_test.rb +0 -59
  88. data/test/minimal_file_reader.rb +0 -25
  89. data/test/path_test.rb +0 -55
  90. data/test/paths/file_test.rb +0 -213
  91. data/test/paths/http_test.rb +0 -34
  92. data/test/paths/matcher_test.rb +0 -120
  93. data/test/paths/s3_test.rb +0 -220
  94. data/test/paths/sftp_test.rb +0 -106
  95. data/test/pgp_reader_test.rb +0 -46
  96. data/test/pgp_test.rb +0 -267
  97. data/test/pgp_writer_test.rb +0 -130
  98. data/test/record_reader_test.rb +0 -60
  99. data/test/record_writer_test.rb +0 -82
  100. data/test/row_reader_test.rb +0 -35
  101. data/test/row_writer_test.rb +0 -56
  102. data/test/stream_test.rb +0 -577
  103. data/test/tabular_test.rb +0 -338
  104. data/test/test_helper.rb +0 -40
  105. data/test/utils_test.rb +0 -20
  106. data/test/xlsx_reader_test.rb +0 -37
  107. data/test/zip_reader_test.rb +0 -53
  108. data/test/zip_writer_test.rb +0 -48
data/docs/tutorial.md ADDED
@@ -0,0 +1,483 @@
1
+ ---
2
+ layout: default
3
+ title: Tutorial
4
+ heading: File / Data Streaming with Ruby
5
+ description: >-
6
+ A step by step introduction to streaming files with IOStreams, from counting
7
+ the lines in a gzip file to reading tabular data out of S3.
8
+ ---
9
+
10
+ If all files were small, they could just be loaded into memory in their entirety.
11
+ However, multi Gigabytes, or even Terabytes in size, loading them into memory is not feasible.
12
+
13
+ In linux it is common to use pipes to stream data between processes.
14
+ For example:
15
+
16
+ ~~~
17
+ # Count the number of lines in a file that has been compressed with gzip
18
+ cat abc.gz | gunzip -c | wc -l
19
+ ~~~
20
+
21
+ For large files it is critical to be able to read and write these files as streams. Ruby has support
22
+ for reading and writing files using streams, but has no built-in way of passing one stream through
23
+ another to support for example compressing the data, encrypting it and then finally writing the result
24
+ to a file. Several streaming implementations exist for languages such as `C++` and `Java` to chain
25
+ together several streams, `IOStreams` offers similar features for Ruby.
26
+
27
+ ~~~ruby
28
+ # Read the first 1024 characters from a compressed file:
29
+ path = IOStreams.path("hello.gz")
30
+ path.reader do |io|
31
+ data = io.read(1024)
32
+ puts "Read: #{data}"
33
+ end
34
+ ~~~
35
+
36
+ The true power of streams is shown when many streams are chained together to achieve the end
37
+ result, without holding the entire file in memory, or ideally without needing to create
38
+ any temporary files to process the stream.
39
+
40
+ ~~~ruby
41
+ # Create a file that is compressed with GZip and then encrypted with Symmetric Encryption:
42
+ path = IOStreams.path("hello.gz.enc")
43
+ path.writer do |io|
44
+ io << "Hello World"
45
+ io << "and some more"
46
+ end
47
+ ~~~
48
+
49
+ The power of the above example applies when the data being written starts to exceed hundreds of megabytes,
50
+ or even gigabytes.
51
+
52
+ By looking at the file name supplied above, IOStreams is able to determine which streams to apply
53
+ to the data being read or written. For example:
54
+ * `hello.zip` => Compressed using Zip
55
+ * `hello.zip.enc` => Compressed using Zip and then encrypted using Symmetric Encryption
56
+ * `hello.gz.enc` => Compressed using GZip and then encrypted using Symmetric Encryption
57
+
58
+ The objective is that all of these streaming processes are performed used streaming
59
+ so that only the current portion of the file is loaded into memory as it moves
60
+ through the entire file.
61
+
62
+ ## Step by Step
63
+
64
+ Install IOStreams gem:
65
+ ~~~
66
+ gem install iostreams --no-doc
67
+ ~~~
68
+
69
+ If you want to follow the AWS S3 examples below install the AWS S3 gem:
70
+ ~~~
71
+ gem install aws-sdk-s3 --no-doc
72
+ ~~~
73
+
74
+ See [Configuring the AWS SDK for Ruby](https://docs.aws.amazon.com/sdk-for-ruby/v3/developer-guide/setup-config.html)
75
+ to configure the Ruby AWS library.
76
+
77
+ Open a ruby interactive console:
78
+
79
+ ~~~
80
+ irb
81
+ ~~~
82
+
83
+ Load iostreams:
84
+
85
+ ~~~ruby
86
+ require "iostreams"
87
+ ~~~
88
+
89
+ Reference a file path to hold CSV data and then for fun lets also compress it with GZip:
90
+ ~~~ruby
91
+ path = IOStreams.path("sample/example.csv.gz")
92
+ # => #<IOStreams::Paths::File:sample/example.csv.gz pipeline={:gz=>{}}>
93
+ ~~~
94
+
95
+ The path and file name does not exist yet:
96
+ ~~~ruby
97
+ path.exist?
98
+ # => false
99
+ ~~~
100
+
101
+ If the path `sample` does not exist, it is created automatically during the first write.
102
+ Write CSV data to the file, compressing to GZip as we go.
103
+ ~~~ruby
104
+ path.writer do |io|
105
+ io << "name,login\n"
106
+ io << "Jack Jones,jjones\n"
107
+ io << "Jill Smith,jsmith\n"
108
+ end
109
+ ~~~
110
+
111
+ To verify the data written above, read the entire file:
112
+ ~~~ruby
113
+ path.read
114
+ # => "name,login\nJack Jones,jjones\nJill Smith,jsmith\n"
115
+ ~~~
116
+
117
+ It would be much easier if we could write the CSV data as hashes and let IOStreams deal
118
+ with all the details on how to create properly formatted CSV data:
119
+ ~~~ruby
120
+ path.writer(:hash) do |io|
121
+ io << {name: "Jack Jones", login: "jjones"}
122
+ io << {name: "Jill Smith", login: "jsmith"}
123
+ end
124
+ ~~~
125
+
126
+ Verify the data written by reading the entire file:
127
+ ~~~ruby
128
+ path.read
129
+ # => "name,login\nJack Jones,jjones\nJill Smith,jsmith\n"
130
+ ~~~
131
+
132
+ Now lets read the file one line at a time:
133
+ ~~~ruby
134
+ path.each do |line|
135
+ puts line
136
+ end
137
+ ~~~
138
+ Output:
139
+ ~~~
140
+ name,login
141
+ Jack Jones,jjones
142
+ Jill Smith,jsmith
143
+ ~~~
144
+
145
+ But who wants to do CSV parsing by hand, lets get IOStreams to do that for us by passing `:array` to `each`:
146
+ ~~~ruby
147
+ path.each(:array) do |array|
148
+ p array
149
+ end
150
+ ~~~
151
+ Output:
152
+ ~~~ruby
153
+ ["name", "login"]
154
+ ["Jack Jones", "jjones"]
155
+ ["Jill Smith", "jsmith"]
156
+ ~~~
157
+
158
+ That was better, but we really want a hash back where IOStreams takes care of the CSV header:
159
+ ~~~ruby
160
+ path.each(:hash) do |hash|
161
+ p hash
162
+ end
163
+ ~~~
164
+ Output:
165
+ ~~~ruby
166
+ {"name"=>"Jack Jones", "login"=>"jjones"}
167
+ {"name"=>"Jill Smith", "login"=>"jsmith"}
168
+ ~~~
169
+
170
+ As the file gets larger and we reach millions of rows the above code does not have to change at all.
171
+ And memory utilization stays about the same since each block is read in, decompressed, and parsed
172
+ from CSV one block at a time. The garbage collector can then free the released blocks from memory.
173
+
174
+ Now lets read a zip file hosted on an HTTP Web Server, displaying the first row as a hash:
175
+
176
+ ~~~ruby
177
+ IOStreams.
178
+ path("https://www5.fdic.gov/idasp/Offices2.zip").
179
+ option(:zip, entry_file_name: "OFFICES2_ALL.CSV").
180
+ each(:hash) do |row|
181
+ p row
182
+ # Just show the first line for this tutorial
183
+ break
184
+ end
185
+ ~~~
186
+ Output:
187
+ ~~~ruby
188
+ {"address"=>"1 Lincoln St. Fl 1", "bkclass"=>"SM", "cbsa"=>"Boston-Cambridge-Newton, MA-NH", "cbsa_div"=>"Boston, MA", "cbsa_div_flg"=>"1", "cbsa_div_no"=>"14454", "cbsa_metro"=>"14460", "cbsa_metro_flg"=>"1", "cbsa_metro_name"=>"Boston-Cambridge-Newton, MA-NH", "cbsa_micro_flg"=>"0", "cbsa_no"=>"14460", "cert"=>"14", "city"=>"Boston", "county"=>"Suffolk", "csa"=>"Boston-Worcester-Providence, MA-RI-NH-CT", "csa_flg"=>"1", "csa_no"=>"148", "estymd"=>"1792-01-01", "fi_uninum"=>"6", "mainoff"=>"1", "name"=>"State Street Bank And Trust Company", "offname"=>"State Street Bank And Trust Company", "offnum"=>nil, "rundate"=>"2020-05-14", "servtype"=>"11", "stalp"=>"MA", "stcnty"=>"25025", "stname"=>"Massachusetts", "uninum"=>"6", "zip"=>"2111"}
189
+ ~~~
190
+
191
+ Noticed that it took a while to return the first line?
192
+
193
+ That is because `zip` requires the entire file to be downloaded before it can decompress anything
194
+ in the file. And HTTP uses a push protocol when reading files, so it is downloaded automatically
195
+ into a temp file behind the scenes so that we can read it as if it was a local file.
196
+
197
+ ## Same Code - Varying File Types
198
+
199
+ Lets define a method to write data to a file.
200
+ ~~~ruby
201
+ def write_lines(file_name)
202
+ path = IOStreams.path(file_name)
203
+ path.writer do |io|
204
+ io << "name,login\n"
205
+ io << "Jack Jones,jjones\n"
206
+ io << "Jill Smith,jsmith\n"
207
+ end
208
+ end
209
+ ~~~
210
+
211
+ Create some sample files to work with
212
+ ~~~ruby
213
+ write_lines("sample/example.csv")
214
+ write_lines("sample/example.csv.gz")
215
+ ~~~
216
+
217
+ For PGP files we also need to specify the recipient that can decrypt the file.
218
+ ~~~ruby
219
+ path = IOStreams.path("sample/example.csv.pgp")
220
+ path.option(:pgp, recipient: "receiver@example.org")
221
+ write_lines(path)
222
+ ~~~
223
+
224
+ `IOStreams.path` takes a string as its argument, it can also accept an existing instance of `IOStreams`.
225
+ That allows the same method to accept the pgp recipient without having to pass the pgp specific recipient information
226
+ as an argument to the method.
227
+
228
+ Consider a simple method to display the contents of a file a line at a time
229
+ prefixed with the line number within the file:
230
+ ~~~ruby
231
+ def show_lines(file_name)
232
+ line_number = 1
233
+ path = IOStreams.path(file_name)
234
+ path.each(:line) do |line|
235
+ puts "[#{line_number}] #{line}"
236
+ line_number += 1
237
+ end
238
+ end
239
+ ~~~
240
+
241
+ Lets read all of the files created above with the new `show_lines` method:
242
+
243
+ ~~~ruby
244
+ show_lines("sample/example.csv")
245
+ show_lines("sample/example.csv.gz")
246
+ show_lines("sample/example.csv.pgp")
247
+ ~~~
248
+
249
+ Noticed how they all returned the exact same output, even though the first file was plain text, the second was
250
+ compressed with Gzip and the third was encrypted with PGP. They all returned:
251
+ ~~~
252
+ [1] name,login
253
+ [2] Jack Jones,jjones
254
+ [3] Jill Smith,jsmith
255
+ ~~~
256
+
257
+ Now a program can be developed using IOStreams and then without any code changes is able read plain text, compressed,
258
+ or encrypted files.
259
+
260
+ ## Same Code - Any File Storage
261
+
262
+ Using the unchanged `write_lines` and `show_lines` methods above, lets use them to read and write from S3.
263
+
264
+ But how is that possible since our program / methods above were only tested against local files?
265
+
266
+ Create the same sample files to work with, but this time on AWS S3 in a bucket name `my-iostreams-bucket`
267
+ ~~~ruby
268
+ write_lines("s3://my-iostreams-bucket/sample/example.csv")
269
+ write_lines("s3://my-iostreams-bucket/sample/example.csv.gz")
270
+ ~~~
271
+
272
+ For PGP files we also need to specify the recipient that can decrypt the file.
273
+ ~~~ruby
274
+ path = IOStreams.path("s3://my-iostreams-bucket/sample/example.csv.pgp")
275
+ path.option(:pgp, recipient: "receiver@example.org")
276
+ write_lines(path)
277
+ ~~~
278
+
279
+ The only change to switch to S3 storage was to prefix the file name passed in with `s3://my-iostreams-bucket/`.
280
+
281
+ Lets read all of the files created above with the new `show_lines` method:
282
+
283
+ ~~~ruby
284
+ show_lines("s3://my-iostreams-bucket/sample/example.csv")
285
+ show_lines("s3://my-iostreams-bucket/sample/example.csv.gz")
286
+ show_lines("s3://my-iostreams-bucket/sample/example.csv.pgp")
287
+ ~~~
288
+
289
+ Noticed how they all returned the exact same output, even though the first file was plain text, the second was
290
+ compressed with Gzip and the third was encrypted with PGP. They all returned:
291
+ ~~~
292
+ [1] name,login
293
+ [2] Jack Jones,jjones
294
+ [3] Jill Smith,jsmith
295
+ ~~~
296
+
297
+ Now a program can be developed using IOStreams and then without any code changes is able to read and write across
298
+ multiple storage locations.
299
+
300
+ ### Configure storage once, not in every call
301
+
302
+ In the examples above we still wrote the `s3://my-iostreams-bucket/` prefix into every call, which
303
+ means the code itself decides where files are stored. The real goal is for the application code to
304
+ say nothing at all about storage, so the same code can run unchanged against local files in
305
+ development and S3 in production.
306
+
307
+ Roots make that possible. Configure the storage location once in a startup initializer:
308
+ ~~~ruby
309
+ # config/initializers/iostreams.rb (development)
310
+ IOStreams.add_root(:default, "tmp/files")
311
+ ~~~
312
+
313
+ Then use `IOStreams.join` instead of `IOStreams.path`, and never mention the storage location again:
314
+ ~~~ruby
315
+ def write_lines(file_name)
316
+ path = IOStreams.join(file_name)
317
+ path.writer do |io|
318
+ io << "name,login\n"
319
+ io << "Jack Jones,jjones\n"
320
+ io << "Jill Smith,jsmith\n"
321
+ end
322
+ end
323
+
324
+ write_lines("sample/example.csv")
325
+ write_lines("sample/example.csv.gz")
326
+ ~~~
327
+
328
+ To run the exact same code against S3 in production, change only the initializer:
329
+ ~~~ruby
330
+ # config/initializers/iostreams.rb (production)
331
+ IOStreams.add_root(:default, "s3://my-iostreams-bucket/files")
332
+ ~~~
333
+
334
+ The application code does not change at all. See [Configuring IOStreams](config) for multiple roots,
335
+ for example separate locations for downloads, uploads, and reports.
336
+
337
+ ## Tabular Files
338
+
339
+ Tabular files are any files that start with a header row and then follows with rows of data with each row
340
+ on a separate line.
341
+
342
+ For example "example.csv"
343
+ ~~~
344
+ name,login
345
+ Jack Jones,jjones
346
+ Jill Smith,jsmith
347
+ ~~~
348
+
349
+ The first line contains the header: `name,login`
350
+ Each subsequent line contains the data delimited by a special character such as `,` in the same order as the header.
351
+
352
+ Another example is PSV (Pipe Separated Files)
353
+ ~~~
354
+ name|login
355
+ Jack Jones|jjones
356
+ Jill Smith|jsmith
357
+ ~~~
358
+
359
+ Of course these are simple examples and there are lots of rules on how to embed or escape the row or column delimiters.
360
+
361
+ ### Reading Tabular Files
362
+
363
+ When reading these files, IOStreams can handle the complexity of the files format and always return the data as a
364
+ `hash`, or `array`.
365
+
366
+ Lets create another method along the lines of `show_lines` above:
367
+ ~~~ruby
368
+ def show_rows(file_name)
369
+ line_number = 1
370
+ path = IOStreams.path(file_name)
371
+ path.each(:hash) do |row|
372
+ puts "[#{line_number}] #{row.inspect}"
373
+ line_number += 1
374
+ end
375
+ end
376
+ ~~~
377
+
378
+ The key difference is that `:hash` is being passed into `each` instead of `:line`.
379
+
380
+ Using the sample files created above:
381
+ ~~~ruby
382
+ show_rows("sample/example.csv")
383
+ ~~~
384
+
385
+ Outputs:
386
+ ~~~
387
+ [1] {"name"=>"Jack Jones", "login"=>"jjones"}
388
+ [2] {"name"=>"Jill Smith", "login"=>"jsmith"}
389
+ ~~~
390
+
391
+ Notice how only 2 rows are returned, since the header row is not actual data, it is just the definition of the
392
+ rows that follow.
393
+
394
+ The same method works without changes regardless of where the file was stored, or whether it was encrypted or
395
+ compressed.
396
+ ~~~ruby
397
+ show_rows("s3://my-iostreams-bucket/sample/example.csv")
398
+ show_rows("s3://my-iostreams-bucket/sample/example.csv.gz")
399
+ show_rows("s3://my-iostreams-bucket/sample/example.csv.pgp")
400
+ ~~~
401
+
402
+ ### Writing Tabular Files
403
+
404
+ Lets define a new method that uses a tabular api to write the data.
405
+ ~~~ruby
406
+ def write_tabular(file_name)
407
+ path = IOStreams.path(file_name)
408
+ path.writer(:hash) do |io|
409
+ io << {"name"=>"Jack Jones", "login"=>"jjones"}
410
+ io << {"name"=>"Jill Smith", "login"=>"jsmith"}
411
+ end
412
+ end
413
+ ~~~
414
+
415
+ The key difference is that `:hash` is being passed into `writer` to indicate that it will receive hashes instead of
416
+ raw data.
417
+
418
+ Lets create a sample file, and then read it to compare its contents to the raw writer above.
419
+ ~~~ruby
420
+ write_tabular("sample/example.csv")
421
+ IOStreams.path("sample/example.csv").read
422
+ # => "name,login\nJack Jones,jjones\nJill Smith,jsmith\n"
423
+ ~~~
424
+
425
+ Note how the output file is identical to the one created above.
426
+ Using `writer(:hash)` makes it easier to develop the application without regard for:
427
+ - The order of columns
428
+ - Missing columns
429
+ - Specialized escaping of values to handle row or column delimiters
430
+
431
+ Note: The first row written determines the column names as well as the order of the elements to be written.
432
+ To supply the header columns up front, to set the order or to filter out which columns should be written to the
433
+ target file, pass them to the writer: `path.writer(:hash, columns: ["name", "login"])`.
434
+
435
+ Now lets write the same data into a JSON file, then read it to see what it looks like:
436
+ ~~~ruby
437
+ write_tabular("sample/example.json")
438
+ puts IOStreams.path("sample/example.json").read
439
+ # => "{\"name\":\"Jack Jones\",\"login\":\"jjones\"}\n{\"name\":\"Jill Smith\",\"login\":\"jsmith\"}\n"
440
+ ~~~
441
+
442
+ Using the same `show_rows` method above to display the file line by line
443
+ ~~~ruby
444
+ show_rows("sample/example.json")
445
+ ~~~
446
+
447
+ Outputs the same data even though the file is now json instead of the previous file that was csv:
448
+ ~~~
449
+ [1] {"name"=>"Jack Jones", "login"=>"jjones"}
450
+ [2] {"name"=>"Jill Smith", "login"=>"jsmith"}
451
+ ~~~
452
+
453
+ The same method works without changes regardless of where the file was stored, or whether it was encrypted or
454
+ compressed, or whether the format was csv or json.
455
+ ~~~ruby
456
+ show_rows("sample/example.csv")
457
+ show_rows("sample/example.json")
458
+ show_rows("s3://my-iostreams-bucket/sample/example.csv")
459
+ show_rows("s3://my-iostreams-bucket/sample/example.json.gz")
460
+ show_rows("s3://my-iostreams-bucket/sample/example.json.pgp")
461
+ ~~~
462
+
463
+ ## Conclusion
464
+
465
+ IOStreams makes it possible to write an application to a common api so that
466
+ * the file can be accessed anywhere ( at least a local file, AWS S3, HTTP(S) and SFTP for now).
467
+ * the application does not care if or how the file was compressed.
468
+ * the application does not care if or how the file was encrypted.
469
+ * the actual file storage mechanism can be determined at runtime, or per environment.
470
+ * it is transparent whether the application receives an Excel Spreadsheet, CSV, or PSV formatted file.
471
+ It just works with hashes when desired.
472
+
473
+ IOStreams is an incredibly powerful streaming library that makes changes to file formats, compression, encryption,
474
+ or storage mechanism transparent to the application.
475
+
476
+ ## Next Steps
477
+
478
+ Move onto the [IOStreams Path Documentation](path) to see how to create paths that support AWS S3, HTTP(S), or SFTP.
479
+
480
+ Or jump straight into the [IOStreams Streams Documentation](streams) for detailed information on working with streams.
481
+
482
+ Read the [IOStreams PGP Documentation](pgp) for a tutorial on how to work with PGP Encrypted files.
483
+
data/docs/upgrading.md ADDED
@@ -0,0 +1,217 @@
1
+ ---
2
+ layout: default
3
+ title: Upgrading
4
+ heading: Upgrading IOStreams
5
+ description: >-
6
+ What changes when upgrading IOStreams to v2.1 or v2.0, and the security
7
+ settings to review in your application when upgrading.
8
+ ---
9
+
10
+ This page covers the changes that may need updates to your application when upgrading IOStreams,
11
+ and the security issues to check. For every change in each release, see the
12
+ [CHANGELOG](https://github.com/reidmorrison/iostreams/blob/main/CHANGELOG.md).
13
+
14
+ ## Upgrading to v2.1
15
+
16
+ v2.1 is a security release, and is backward compatible except for `IOStreams::Pgp.delete_keys`,
17
+ described below. Most applications upgrade without any code changes. The changes that would break
18
+ existing code are postponed to v3.0, and log a warning in v2.1. See [Coming in v3.0](#coming-in-v30).
19
+
20
+ ### Changes that may need code changes
21
+
22
+ #### PGP `delete_keys` requires an email or key id
23
+
24
+ `IOStreams::Pgp.delete_keys` raises `ArgumentError` unless an `email:` or `key_id:` is supplied.
25
+ Previously, on GnuPG 2.1 and later, calling it without either deleted every key in the keyring, and
26
+ with `private: true`, every secret key. This is the only breaking change in v2.1, since that could
27
+ not be undone. To delete several keys, call it once for each key, for example for each key returned
28
+ by `IOStreams::Pgp.list_keys`.
29
+
30
+ #### SFTP `#each_child` requires a known host key
31
+
32
+ `#each_child` on an SFTP path now verifies the server's host key, the same way reading and writing
33
+ already did. Previously it trusted a host key the first time it was seen. If listing files now
34
+ fails, supply the server's host key with the `HostKey` ssh option, or add it to the `known_hosts`
35
+ file of the user running the application. See [SFTP](path#sftp-sftp).
36
+
37
+ `#each_child` now also supports the `HostKey`, `IdentityKey` and `IdentityFile` ssh options, and
38
+ some others. Previously any ssh option made it raise `ArgumentError`. Any ssh option it does not
39
+ support raises `ArgumentError`.
40
+
41
+ #### Other changes in behavior
42
+
43
+ - SFTP `#each_child` lists the files in the path's directory. Previously it ignored the path and
44
+ listed the login directory, so `IOStreams.path("sftp://host/data/in").each_child("*.csv")`
45
+ listed `~/*.csv`, returning paths such as `/a.csv`. A url without a path, such as `sftp://host`,
46
+ still lists the login directory.
47
+ - `Path#join` and `IOStreams.join` only return an element unchanged when it is the path itself or
48
+ is inside it. Previously any element that started with the same characters was returned as-is,
49
+ so `IOStreams.path("s3://bucket/reports").join("reports_2024.csv")` gave
50
+ `s3://bucket/reports_2024.csv`, instead of `s3://bucket/reports/reports_2024.csv`.
51
+ - When writing PSV or fixed width files, a line break within a value is replaced with a space, so
52
+ that a value can no longer start a separate record.
53
+ - Internal temp files are created so that only the current user can read them.
54
+ - When reading lines, newlines within quoted values are kept within the line when the tabular format
55
+ quotes its values, such as CSV. Previously this depended on whether the file name contained `.csv`,
56
+ so `.format(:psv)` on a `.csv` file still joined quoted lines, and `.format(:csv)` on a stream
57
+ without a file name did not.
58
+ - An `ArgumentError` for an option that a stream does not accept now says which direction the option
59
+ belongs to, or lists the valid options. Previously it was Ruby's `unknown keyword` error.
60
+
61
+ ### Security checklist
62
+
63
+ When upgrading, review how your application uses IOStreams for the following documented security
64
+ issues. Each one depends on how the application is written, so upgrading alone does not fix it.
65
+
66
+ #### File names from untrusted input
67
+
68
+ If a user, a job, or any other untrusted input can supply a file name, it can also supply `..`
69
+ or an absolute path, and read or overwrite any file that the process can access. `IOStreams.join`
70
+ and root paths do not prevent this.
71
+
72
+ New in v2.1, add **allowed paths** in an initializer, so that accessing any other path raises
73
+ `IOStreams::Errors::AccessDenied`:
74
+
75
+ ~~~ruby
76
+ IOStreams.add_allowed_path("/var/my_app/uploads")
77
+ IOStreams.add_allowed_path("s3://my-app-bucket-name/export")
78
+ ~~~
79
+
80
+ Once any allowed path is added, every path IOStreams accesses in the process must be within one of
81
+ them, so add every location that your application reads or writes. See
82
+ [Restricting access with allowed paths](path#restricting-access-with-allowed-paths).
83
+
84
+ #### Untrusted names in S3 urls
85
+
86
+ Any query string in an S3 url is added to the S3 request parameters. Do not interpolate an
87
+ untrusted file name into the url, since a name such as `file.csv?acl=public-read` would set
88
+ request parameters. Join it onto the path instead. See [AWS S3](path#aws-s3-s3).
89
+
90
+ ~~~ruby
91
+ # Unsafe
92
+ IOStreams.path("s3://my-bucket-name/uploads/#{untrusted_name}")
93
+
94
+ # Safe
95
+ IOStreams.path("s3://my-bucket-name/uploads").join(untrusted_name)
96
+ ~~~
97
+
98
+ #### Credentials in HTTP and SFTP urls
99
+
100
+ A username and password supplied in an HTTP or SFTP url remain part of it, so `#to_s` returns
101
+ them, as does any log or error message that includes the path. Supply them with the `username:`
102
+ and `password:` arguments instead:
103
+
104
+ ~~~ruby
105
+ # Returns the password from #to_s
106
+ IOStreams.path("sftp://jack:secret@sftp.example.org/file.csv")
107
+
108
+ # Does not
109
+ IOStreams.path("sftp://sftp.example.org/file.csv", username: "jack", password: "secret")
110
+ ~~~
111
+
112
+ #### Untrusted HTTP urls
113
+
114
+ When any part of an HTTP url comes from untrusted input, an attacker can point it at internal
115
+ services or cloud metadata endpoints (Server Side Request Forgery), either directly or with a
116
+ redirect. Use `allow_hosts`, `http_redirect_count` and `maximum_file_size`, or allowed paths, which
117
+ also check every redirect. See [Security: untrusted URLs (SSRF)](path#security-untrusted-urls-ssrf).
118
+
119
+ #### Processing PGP files before they are verified
120
+
121
+ By default IOStreams passes the decrypted contents of a PGP file to your block as gpg decrypts it.
122
+ gpg can only check the file's integrity and signature once it has read the whole file, so a
123
+ tampered file is only rejected after your block has processed its contents. Either do not commit
124
+ any side effects until the block returns, for example by using a database transaction, or supply
125
+ the `verify_first` option, new in v2.1. See
126
+ [Verifying a file before processing it](pgp#verifying-a-file-before-processing-it).
127
+
128
+ ~~~ruby
129
+ path.option(:pgp, passphrase: "receiver_passphrase", verify_first: true)
130
+ ~~~
131
+
132
+ #### Importing and trusting PGP keys
133
+
134
+ `import_and_trust` and the writer's `import_and_trust_key` option trust a key at the Ultimate
135
+ level by default, so only use them with keys received from a verified, trusted source. When a key
136
+ cannot be fully verified, supply a lower `trust_level:` to `import_and_trust`, or a lower
137
+ `import_and_trust_level` option to the writer. See
138
+ [Trust level](pgp#trust-level).
139
+
140
+ In v2.1, `import_and_trust_key` encrypts to the imported key's fingerprint on GnuPG 2.1 and later.
141
+ Previously it encrypted to the key's email address, so if your keyring holds another key with the
142
+ same email address, that key could have been used instead. Check your keyring for duplicate email
143
+ addresses.
144
+
145
+ #### PGP files without integrity protection
146
+
147
+ Only enable `ignore_mdc_error` for files from a trusted source, since without MDC the decrypted
148
+ contents are not protected against tampering. See
149
+ [Reading legacy files without MDC integrity protection](pgp#reading-legacy-files-without-mdc-integrity-protection).
150
+
151
+ #### Column restrictions on uploaded files
152
+
153
+ By default, `allowed_columns` and `required_columns` are ignored for JSON input. Since the format
154
+ is usually inferred from the file name, renaming an upload from `.csv` to `.json` bypasses them.
155
+ If your application relies on them to restrict which columns an upload can set, apply them to every
156
+ input in an initializer, new in v2.1, and check whether any uploads were affected:
157
+
158
+ ~~~ruby
159
+ IOStreams.enforce_column_restrictions = true
160
+ ~~~
161
+
162
+ See [Column restrictions apply to every input](#column-restrictions-apply-to-every-input) for what
163
+ it changes.
164
+
165
+ ## Coming in v3.0
166
+
167
+ These changes are postponed to v3.0 since they could break existing code. Each one logs a warning
168
+ in v2.1 via `IOStreams.logger` when it would change the result, so check your logs for them.
169
+
170
+ ### Column restrictions apply to every input
171
+
172
+ When reading records, `allowed_columns`, `required_columns` and `skip_unknown` will apply to every
173
+ input. In v2.1 they are ignored for JSON and `:hash` input, when `columns:` is supplied, and with
174
+ `cleanse_header: false`.
175
+
176
+ When either `allowed_columns` or `required_columns` is set, JSON keys will be cleansed the same way
177
+ as a header row, for example `"Name"` becomes `"name"`. Unknown keys will be skipped, or raise
178
+ `IOStreams::Errors::InvalidHeader` when `skip_unknown: false`, and a record that is missing a
179
+ required column will raise `IOStreams::Errors::InvalidHeader`.
180
+
181
+ To prepare: set `IOStreams.enforce_column_restrictions = true`, and check that the JSON files you
182
+ read with these options have the expected keys. See [Header options](formats#header-options).
183
+
184
+ ### BZip2 options are strict
185
+
186
+ The BZip2 reader and writer ignore any option they do not accept, and log a warning. In v3.0 it will
187
+ raise `ArgumentError`, like every other stream. They accept `autoclose`, `first_only` and `small` when
188
+ reading, and `autoclose`, `block_size` and `work_factor` when writing.
189
+
190
+ To prepare: remove the option, or use a separate path for reading and for writing. See
191
+ [Reading and writing need separate options](streams#reading-and-writing-need-separate-options).
192
+
193
+ ### PGP `export` without an email or key id
194
+
195
+ `IOStreams::Pgp.export(email: nil)` without a `key_id:` raises `IOStreams::Pgp::Failure`, as it did
196
+ before v2.1. In v3.0 it will raise `ArgumentError`. Calling it without either argument already raises
197
+ `ArgumentError`.
198
+
199
+ ## Upgrading to v2.0
200
+
201
+ v2.0 is a major release with breaking changes:
202
+
203
+ - **Ruby 3.2 or later is required.**
204
+ - **Writing Zip files requires the `zip_kit` gem.** The retired `zip_tricks` gem has been replaced
205
+ by its successor, `zip_kit`. If your application writes Zip files, replace `gem "zip_tricks"` with
206
+ `gem "zip_kit"` in your Gemfile. Reading Zip files is unaffected.
207
+ - **The deprecated pre-v1.6 API has been removed**, including the `IOStreams::Deprecated` mix-in.
208
+ Move any code still using it to the `IOStreams.path` and `IOStreams.stream` API.
209
+ - **The deprecated PGP writer `compression:` option has been removed.** Use `compress:` instead,
210
+ available since v1.11.0.
211
+ - **`IOStreams::Pgp.logger` and `IOStreams::Pgp.logger=` have been removed.** Logging is configured
212
+ for the whole library with `IOStreams.logger=`. See [logger](config#logger).
213
+
214
+ ## Upgrading from before v1.6
215
+
216
+ The pre-v1.6 API was deprecated in v1.6 and removed in v2.0. Upgrade to v1.11 first, move all code to
217
+ the `IOStreams.path` and `IOStreams.stream` API, and then upgrade to v2.