iostreams 2.0.0 → 2.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (50) hide show
  1. checksums.yaml +4 -4
  2. data/README.md +5 -22
  3. data/Rakefile +45 -0
  4. data/docs/CLAUDE.md +9 -0
  5. data/docs/config.md +157 -0
  6. data/docs/copy_files.md +75 -0
  7. data/docs/extensions.md +111 -0
  8. data/docs/formats.md +188 -0
  9. data/docs/index.md +388 -0
  10. data/docs/path.md +652 -0
  11. data/docs/pgp.md +436 -0
  12. data/docs/streams.md +337 -0
  13. data/docs/tutorial.md +483 -0
  14. data/docs/upgrading.md +217 -0
  15. data/lib/io_streams/builder.rb +62 -2
  16. data/lib/io_streams/bzip2/reader.rb +25 -2
  17. data/lib/io_streams/bzip2/writer.rb +26 -2
  18. data/lib/io_streams/encode/reader.rb +4 -0
  19. data/lib/io_streams/encode/writer.rb +4 -0
  20. data/lib/io_streams/errors.rb +4 -0
  21. data/lib/io_streams/gzip/reader.rb +4 -0
  22. data/lib/io_streams/gzip/writer.rb +11 -2
  23. data/lib/io_streams/io_streams.rb +111 -1
  24. data/lib/io_streams/line/reader.rb +7 -2
  25. data/lib/io_streams/path.rb +115 -6
  26. data/lib/io_streams/paths/file.rb +47 -1
  27. data/lib/io_streams/paths/http.rb +45 -4
  28. data/lib/io_streams/paths/s3.rb +66 -15
  29. data/lib/io_streams/paths/sftp/net_ssh.rb +104 -0
  30. data/lib/io_streams/paths/sftp.rb +97 -57
  31. data/lib/io_streams/pgp/reader.rb +42 -2
  32. data/lib/io_streams/pgp/writer.rb +26 -6
  33. data/lib/io_streams/pgp.rb +78 -21
  34. data/lib/io_streams/reader.rb +10 -1
  35. data/lib/io_streams/record/reader.rb +72 -2
  36. data/lib/io_streams/stream.rb +12 -7
  37. data/lib/io_streams/symmetric_encryption/reader.rb +4 -0
  38. data/lib/io_streams/symmetric_encryption/writer.rb +4 -0
  39. data/lib/io_streams/tabular/header.rb +31 -4
  40. data/lib/io_streams/tabular/parser/base.rb +10 -0
  41. data/lib/io_streams/tabular/parser/csv.rb +5 -0
  42. data/lib/io_streams/tabular/parser/fixed.rb +3 -1
  43. data/lib/io_streams/tabular/parser/psv.rb +6 -2
  44. data/lib/io_streams/utils.rb +31 -0
  45. data/lib/io_streams/version.rb +1 -1
  46. data/lib/io_streams/writer.rb +10 -1
  47. data/lib/io_streams/xlsx/reader.rb +5 -1
  48. data/lib/io_streams/zip/reader.rb +4 -0
  49. data/lib/io_streams/zip/writer.rb +4 -0
  50. metadata +24 -7
data/docs/path.md ADDED
@@ -0,0 +1,652 @@
1
+ ---
2
+ layout: default
3
+ title: Path
4
+ description: >-
5
+ How a path identifies where a file is stored and how to reach it, with the
6
+ arguments for each location: local disk, AWS S3, SFTP and HTTP.
7
+ ---
8
+
9
+ A path identifies _where_ a file is stored and how to reach it, so that the streaming pipeline knows
10
+ where to read the data from or write it to.
11
+
12
+ Create a path with `IOStreams.path`, passing the file name, which may also be a URI, followed by any
13
+ arguments specific to that storage location. IOStreams infers the storage mechanism from the URI
14
+ scheme, so the same call returns a local file path, an S3 path, an SFTP path, and so on, all sharing
15
+ the identical interface.
16
+
17
+ IOStreams supports accessing files in the following places:
18
+
19
+ * File
20
+ * AWS S3
21
+ * Google Cloud Storage (Using the AWS S3 Client)
22
+ * SFTP
23
+ * HTTP(S) (Read only)
24
+
25
+ Are you using another cloud provider and want to add support for your favorite?
26
+ Checkout the supplied [IOStreams S3 path provider](https://github.com/reidmorrison/iostreams/blob/main/lib/io_streams/paths/s3.rb)
27
+ for an example of what is required. Pull requests welcome.
28
+
29
+ ### File
30
+
31
+ The simplest case is a file on the local disk:
32
+
33
+ ~~~ruby
34
+ path = IOStreams.path("somewhere/example.csv")
35
+ ~~~
36
+
37
+ #### Optional Arguments:
38
+
39
+ * `:create_path` set to false to stop IOStreams from automatically creating the output directories
40
+ if they do not exist.
41
+ Default: true
42
+ ~~~ruby
43
+ path = IOStreams.path("somewhere/example.csv.gz", create_path: false)
44
+ ~~~
45
+
46
+ ### AWS S3 (s3://)
47
+
48
+ If the supplied file name string includes a URI. For example if AWS is configured locally:
49
+
50
+ ~~~ruby
51
+ path = IOStreams.path("s3://bucket-name/path/example.csv")
52
+ ~~~
53
+
54
+ #### Required Arguments:
55
+
56
+ * url [String]
57
+
58
+ Prefix must be: `s3://`, followed by bucket name, followed by key.
59
+ Any query string in the url is added to the S3 request parameters, for example
60
+ `s3://my-bucket-name/file_name.csv?acl=bucket-owner-full-control`.
61
+ Examples:
62
+ s3://my-bucket-name/file_name.txt
63
+ s3://my-bucket-name/some_path/file_name.csv
64
+
65
+ Security warning: do not interpolate an untrusted file name into the url, since a name such as
66
+ `file.csv?acl=public-read` would set request parameters. Join it onto the path instead, which
67
+ does not parse it as a query:
68
+ `IOStreams.path("s3://my-bucket-name/uploads").join(untrusted_name)`
69
+
70
+ #### Optional Arguments:
71
+
72
+ * :access_key_id [String]
73
+
74
+ AWS Access Key Id to use to access this bucket.
75
+
76
+ * :secret_access_key [String]
77
+
78
+ AWS Secret Access Key Id to use to access this bucket.
79
+
80
+ * :region [String]
81
+
82
+ The AWS region to connect to.
83
+ Default: the region set in the environment variables or credential files.
84
+
85
+ * :client [Aws::S3::Client | Hash]
86
+
87
+ Supply the AWS S3 Client instance to use for this path.
88
+ Or, when a Hash, build a new client using the hash parameters.
89
+
90
+ ~~~ruby
91
+ client = Aws::S3::Client.new(endpoint: "https://s3.test.com")
92
+ path = IOStreams.path("s3://bucket/path/file_name.txt", client: client)
93
+
94
+ # Or, pass the client parameters directly:
95
+ path = IOStreams.path("s3://bucket/path/file_name.txt", client: {endpoint: "https://s3.test.com"})
96
+ ~~~
97
+
98
+ Writer specific options:
99
+
100
+ * :acl [String]
101
+
102
+ The canned ACL to apply to the object.
103
+
104
+ * :cache_control [String]
105
+
106
+ Specifies caching behavior along the request/reply chain.
107
+
108
+ * :content_disposition [String]
109
+
110
+ Specifies presentational information for the object.
111
+
112
+ * :content_encoding [String]
113
+
114
+ Specifies what content encodings have been applied to the object and
115
+ thus what decoding mechanisms must be applied to obtain the media-type
116
+ referenced by the Content-Type header field.
117
+
118
+ * :content_language [String]
119
+
120
+ The language the content is in.
121
+
122
+ * :content_length [Integer]
123
+
124
+ Size of the body in bytes. This parameter is useful when the size of
125
+ the body cannot be determined automatically.
126
+
127
+ * :content_md5 [String]
128
+
129
+ The base64-encoded 128-bit MD5 digest of the part data. This parameter
130
+ is auto-populated when using the command from the CLI. This parameted
131
+ is required if object lock parameters are specified.
132
+
133
+ * :content_type [String]
134
+
135
+ A standard MIME type describing the format of the object data.
136
+
137
+ * :expires [Time,DateTime,Date,Integer,String]
138
+
139
+ The date and time at which the object is no longer cacheable.
140
+
141
+ * :grant_full_control [String]
142
+
143
+ Gives the grantee READ, READ\_ACP, and WRITE\_ACP permissions on the
144
+ object.
145
+
146
+ * :grant_read [String]
147
+
148
+ Allows grantee to read the object data and its metadata.
149
+
150
+ * :grant_read_acp [String]
151
+
152
+ Allows grantee to read the object ACL.
153
+
154
+ * :grant_write_acp [String]
155
+
156
+ Allows grantee to write the ACL for the applicable object.
157
+
158
+ * :metadata [Hash<String,String>]
159
+
160
+ A map of metadata to store with the object in S3.
161
+
162
+ * :server_side_encryption [String]
163
+
164
+ The Server-side encryption algorithm used when storing this object in
165
+ S3 (e.g., AES256, aws:kms).
166
+
167
+ * :storage_class [String]
168
+
169
+ The type of storage to use for the object. Defaults to 'STANDARD'.
170
+
171
+ * :website_redirect_location [String]
172
+
173
+ If the bucket is configured as a website, redirects requests for this
174
+ object to another object in the same bucket or to an external URL.
175
+ Amazon S3 stores the value of this header in the object metadata.
176
+
177
+ * :sse_customer_algorithm [String]
178
+
179
+ Specifies the algorithm to use to when encrypting the object (e.g.,
180
+ AES256).
181
+
182
+ * :sse_customer_key [String]
183
+
184
+ Specifies the customer-provided encryption key for Amazon S3 to use in
185
+ encrypting data. This value is used to store the object and then it is
186
+ discarded; Amazon does not store the encryption key. The key must be
187
+ appropriate for use with the algorithm specified in the
188
+ x-amz-server-side​-encryption​-customer-algorithm header.
189
+
190
+ * :sse_customer_key_md5 [String]
191
+
192
+ Specifies the 128-bit MD5 digest of the encryption key according to
193
+ RFC 1321. Amazon S3 uses this header for a message integrity check to
194
+ ensure the encryption key was transmitted without error.
195
+
196
+ * :ssekms_key_id [String]
197
+
198
+ Specifies the AWS KMS key ID to use for object encryption. All GET and
199
+ PUT requests for an object protected by AWS KMS will fail if not made
200
+ via SSL or using SigV4. Documentation on configuring any of the
201
+ officially supported AWS SDKs and CLI can be found at
202
+ http://docs.aws.amazon.com/AmazonS3/latest/dev/UsingAWSSDK.html#specify-signature-version
203
+
204
+ * :ssekms_encryption_context [String]
205
+
206
+ Specifies the AWS KMS Encryption Context to use for object encryption.
207
+ The value of this header is a base64-encoded UTF-8 string holding JSON
208
+ with the encryption context key-value pairs.
209
+
210
+ * :request_payer [String]
211
+
212
+ Confirms that the requester knows that she or he will be charged for
213
+ the request. Bucket owners need not specify this parameter in their
214
+ requests. Documentation on downloading objects from requester pays
215
+ buckets can be found at
216
+ http://docs.aws.amazon.com/AmazonS3/latest/dev/ObjectsinRequesterPaysBuckets.html
217
+
218
+ * :tagging [String]
219
+
220
+ The tag-set for the object. The tag-set must be encoded as URL Query
221
+ parameters. (For example, "Key1=Value1")
222
+
223
+ * :object_lock_mode [String]
224
+
225
+ The object lock mode that you want to apply to this object.
226
+
227
+ * object_lock_retain_until_date: [Time,DateTime,Date,Integer,String]
228
+
229
+ The date and time when you want this object's object lock to expire.
230
+
231
+ * object_lock_legal_hold_status: [String]
232
+ The Legal Hold status that you want to apply to the specified object.
233
+
234
+ ### SFTP (sftp://)
235
+
236
+ If the supplied file name string includes the `sftp` URI.
237
+
238
+ ~~~ruby
239
+ path = IOStreams.path("sftp://hostname/path/example.csv")
240
+ ~~~
241
+
242
+ IOStreams reads and writes SFTP files by shelling out to the `sftp` command line program,
243
+ so it must be installed and on the `PATH`. When a password is supplied the `sshpass`
244
+ program is also required to pass the password to `sftp`. Additionally the `net-sftp` gem
245
+ must be added to the `Gemfile` to use `each_child`. `each_child` lists the files within the path's directory,
246
+ or within the login directory when the url has no path, for example `sftp://hostname`.
247
+
248
+ Read a file from a remote sftp server.
249
+ ~~~ruby
250
+ IOStreams.path("sftp://example.org/path/file.txt",
251
+ username: "jbloggs",
252
+ password: "secret").
253
+ reader do |input|
254
+ puts input.read
255
+ end
256
+ ~~~
257
+
258
+ Raises `IOStreams::Errors::CommunicationsFailure` when the file could not be read or written.
259
+
260
+ Write to a file on a remote sftp server.
261
+ ~~~ruby
262
+ IOStreams.path("sftp://example.org/path/file.txt",
263
+ username: "jbloggs",
264
+ password: "secret").
265
+ writer do |output|
266
+ output.write('Hello World')
267
+ end
268
+ ~~~
269
+
270
+ Display the contents of a remote file, supplying the username and password in the url.
271
+ Note that `#to_s` then includes the password, so prefer the `username:` and `password:` arguments above:
272
+ ~~~ruby
273
+ IOStreams.path("sftp://jack:OpenSesame@test.com:22/path/file_name.csv").reader do |io|
274
+ puts io.read
275
+ end
276
+ ~~~
277
+
278
+ Use an identity file instead of a password to authenticate:
279
+ ~~~ruby
280
+ path = IOStreams.path("sftp://test.com/path/file_name.csv",
281
+ username: "jack",
282
+ ssh_options: {IdentityFile: "~/.ssh/private_key"})
283
+ path.reader do |io|
284
+ puts io.read
285
+ end
286
+ ~~~
287
+
288
+ Pass in the IdentityKey itself instead of a password to authenticate.
289
+ For example, retrieve the identity key stored in Secret Config:
290
+ ~~~ruby
291
+ identity_key = SecretConfig.fetch("suppliers/sftp/identity_key")
292
+
293
+ path = IOStreams.path("sftp://test.com/path/file_name.csv",
294
+ username: "jack",
295
+ ssh_options: {IdentityKey: identity_key})
296
+ path.reader do |io|
297
+ puts io.read
298
+ end
299
+ ~~~
300
+
301
+ #### Required Arguments:
302
+
303
+ * url [String]
304
+
305
+ Prefix must be: `sftp://`, followed by host name, followed by file name.
306
+ Format:
307
+ "sftp://<host_name>/<file_name>"
308
+ "sftp://username:password@hostname:22/path/file_name"
309
+
310
+ A username and password supplied in the url remain part of it, so `#to_s` returns them,
311
+ as does any log or error message that includes the path. To keep them out of logs, supply them
312
+ with the `username:` and `password:` arguments instead.
313
+
314
+ #### Optional Arguments:
315
+
316
+ * username: [String]
317
+
318
+ Name of user to login with.
319
+
320
+ * password: [String]
321
+
322
+ Password for the user.
323
+
324
+ * ssh_options: [Hash]
325
+
326
+ * IdentityFile [String]
327
+
328
+ Path to the local identity (private key) file to authenticate with, instead of a password.
329
+
330
+ * IdentityKey [String]
331
+
332
+ The identity (private key) itself, supplied as a string.
333
+ Under the covers the key is written to a temp file and then passed as `IdentityFile`.
334
+
335
+ * HostKey [String]
336
+
337
+ The expected SSH host key presented by the remote host, instead of storing it in the
338
+ `known_hosts` file. It must contain the entire line that would be stored in `known_hosts`,
339
+ including the hostname, ip address, key type and key value. The easiest way to generate
340
+ the required value is with `ssh-keyscan hostname`.
341
+ Under the covers the value is written to a temp file and then passed as `UserKnownHostsFile`.
342
+
343
+ * Any other options supported by ssh_config.
344
+ `man ssh_config` to see all available options.
345
+
346
+ `each_child` lists files with the `net-sftp` gem instead of the `sftp` program, so it only supports
347
+ these ssh options: `HostKey`, `IdentityKey`, `IdentityFile`, `UserKnownHostsFile`,
348
+ `StrictHostKeyChecking`, `ConnectTimeout`, `ServerAliveInterval`, `ServerAliveCountMax` and
349
+ `LogLevel`. Any other option raises `ArgumentError`. Unlike the `sftp` program, `net-sftp` needs the
350
+ `ed25519` and `bcrypt_pbkdf` gems to use ed25519 host or identity keys.
351
+
352
+ Notes:
353
+ * Since the `sftp` program operates on local files, reading from or writing to an SFTP path
354
+ streams through a local temp file behind the scenes.
355
+
356
+ ### HTTP (http://, https://)
357
+
358
+ Read from a remote file over HTTP or HTTPS using an HTTP Get.
359
+
360
+ ~~~ruby
361
+ IOStreams.path('https://www5.fdic.gov/idasp/Offices2.zip').read
362
+ ~~~
363
+
364
+ Notes:
365
+ * Since Net::HTTP download only supports a push stream, the data is streamed into a tempfile first.
366
+ * Currently writing to an HTTP(S) server is not supported. Up to submitting a Pull Request with capability?
367
+
368
+ #### Required Arguments:
369
+
370
+ * url [String]
371
+
372
+ Prefix must be: `http://`, or `https://` followed by host name, followed by path and file name.
373
+ Also supports passing the username and password for basic authentication in the URI.
374
+
375
+ Format:
376
+ * http://hostname/path/file_name
377
+ * https://username:password@hostname/path/file_name
378
+
379
+ A username and password supplied in the url remain part of it, so `#to_s` returns them,
380
+ as does any log or error message that includes the path. To keep them out of logs, supply them
381
+ with the `username:` and `password:` arguments instead.
382
+
383
+ #### Optional Arguments:
384
+
385
+ * username: [String]
386
+
387
+ When supplied, basic authentication is used with the username and password.
388
+
389
+ * password: [String]
390
+
391
+ Password to use use with basic authentication when the username is supplied.
392
+
393
+ * parameters: [Hash]
394
+
395
+ Query parameters to append to the url as a query string.
396
+ For example, `parameters: {"type" => "csv"}` appends `?type=csv` to the url.
397
+
398
+ * http_redirect_count: [Integer]
399
+
400
+ Maximum number of http redirects to follow.
401
+ Set to `0` to disable following redirects entirely.
402
+ Default: `10`
403
+
404
+ * allow_hosts: [String | Array<String>]
405
+
406
+ Optional allow-list of host names that may be contacted. It is applied both to the
407
+ supplied url and to every redirect that is followed; a request to any other host raises
408
+ `IOStreams::Errors::CommunicationsFailure`.
409
+ Default: `nil` (any host is allowed).
410
+
411
+ * maximum_file_size: [Integer]
412
+
413
+ Optional maximum number of bytes to download. When the response body exceeds this size the
414
+ download is aborted with an `IOStreams::Errors::CommunicationsFailure`.
415
+ Default: `nil` (no limit).
416
+
417
+ ~~~ruby
418
+ path = IOStreams.path("http://hostname/path/example.csv")
419
+ ~~~
420
+
421
+ #### Security: untrusted URLs (SSRF)
422
+
423
+ Reading an HTTP(S) path causes the application to issue a request to the host named in the url.
424
+ When the url, or any part of it, can be influenced by untrusted input, an attacker can point it
425
+ at internal services or cloud metadata endpoints (Server Side Request Forgery).
426
+
427
+ Because redirect targets are chosen by the remote server, validating only the url that is passed
428
+ in is not sufficient: a trusted (or compromised) server can redirect the request to an internal
429
+ address. IOStreams provides a few controls to reduce this exposure:
430
+
431
+ * Restrict which hosts may be contacted, including across redirects:
432
+
433
+ ~~~ruby
434
+ IOStreams.path("https://supplier.example.com/report.csv", allow_hosts: ["supplier.example.com"]).read
435
+ ~~~
436
+
437
+ * Disable redirects entirely for untrusted urls:
438
+
439
+ ~~~ruby
440
+ IOStreams.path(untrusted_url, http_redirect_count: 0).read
441
+ ~~~
442
+
443
+ * Cap the download size to avoid unbounded (denial of service) responses:
444
+
445
+ ~~~ruby
446
+ IOStreams.path(untrusted_url, maximum_file_size: 50 * 1024 * 1024).read
447
+ ~~~
448
+
449
+ Basic authentication credentials are only ever sent to the original host. They are not resent
450
+ when a redirect points at a different scheme, host, or port, so a redirect cannot leak them to
451
+ another server. For stronger guarantees, route these downloads through an egress proxy or network
452
+ policy that blocks private, loopback, and link-local (cloud metadata) addresses.
453
+
454
+ Similarly when using https:
455
+
456
+ ~~~ruby
457
+ path = IOStreams.path("https://hostname/path/example.csv")
458
+ ~~~
459
+
460
+ This time IOStreams inferred that the file lives on an HTTP Server and returns `IOStreams::Paths::HTTP`.
461
+
462
+ ### Path Operations
463
+
464
+ Paths support common file operations, regardless of where the file is stored:
465
+
466
+ ~~~ruby
467
+ path = IOStreams.path("sample/example.csv")
468
+
469
+ # Does the file exist?
470
+ path.exist?
471
+ # => true
472
+
473
+ # Size of the file in bytes.
474
+ path.size
475
+ # => 64
476
+
477
+ # Delete the file.
478
+ path.delete
479
+
480
+ # Move the file to another path, returning the target path.
481
+ path.move_to("sample/moved.csv")
482
+
483
+ # Create the directory path, when it does not already exist.
484
+ IOStreams.path("sample/data").mkpath
485
+ ~~~
486
+
487
+ Inspect the components of a path's file name:
488
+
489
+ ~~~ruby
490
+ # The last component of the path.
491
+ IOStreams.path("/home/gumby/work/ruby.rb").basename
492
+ # => "ruby.rb"
493
+
494
+ # Remove a specific suffix from the file name.
495
+ IOStreams.path("/home/gumby/work/ruby.rb").basename(".rb")
496
+ # => "ruby"
497
+
498
+ # Remove any extension by supplying ".*".
499
+ IOStreams.path("/home/gumby/work/ruby.rb").basename(".*")
500
+ # => "ruby"
501
+
502
+ # The directory portion of the path.
503
+ IOStreams.path("a/b/d/test.rb").dirname
504
+ # => "a/b/d"
505
+
506
+ # The extension, including the leading period.
507
+ IOStreams.path("a/b/d/test.rb").extname
508
+ # => ".rb"
509
+
510
+ # The extension, without the leading period.
511
+ IOStreams.path("a/b/d/test.rb").extension
512
+ # => "rb"
513
+ ~~~
514
+
515
+ Notes:
516
+ * `basename`, `dirname`, `extname`, and `extension` return `nil` when no file name was set.
517
+ * A leading period on a dotfile is not treated as an extension, so `.profile` has no extension,
518
+ while `.profile.sh` has the extension `sh`.
519
+ * A file name ending in a period, such as `foo.`, returns an empty string for the extension.
520
+
521
+ Iterate over the files in a path using a wildcard pattern:
522
+
523
+ ~~~ruby
524
+ IOStreams.path("sample").each_child("*.csv") do |child|
525
+ puts child
526
+ end
527
+
528
+ # Recursively, including sub-directories:
529
+ IOStreams.path("sample").each_child("**/*.csv") do |child|
530
+ puts child
531
+ end
532
+ ~~~
533
+
534
+ `each_child` is also available directly on `IOStreams` when the pattern includes the full path:
535
+
536
+ ~~~ruby
537
+ IOStreams.each_child("sample/**/*.csv") { |child| puts child }
538
+ ~~~
539
+
540
+ Notes:
541
+ * These operations are supported by File and S3 paths. SFTP supports `each_child`,
542
+ and HTTP paths are read-only so they do not support any of them.
543
+ * By default `each_child` patterns are case-insensitive and hidden files are excluded.
544
+ Supply `case_sensitive: true` or `hidden: true` to change this behavior.
545
+
546
+ ### Using root paths
547
+
548
+ Roots allow paths to reference a particular root directory, so that all path names are appended to that root.
549
+ By using `IOStreams.join` instead of `IOStreams.path`, the storage location is no longer embedded in the
550
+ application code, it is configured once at startup.
551
+
552
+ The primary purpose of roots is to allow the exact same code to run in production and development,
553
+ yet use completely different data sources in each. For example, in production the root can point to an
554
+ S3 bucket, while in development it points to the local file system.
555
+
556
+ Roots are configured via an initializer at startup. Multiple roots can be setup, for example one for
557
+ input files, another for output files, another for reports, etc. During development the roots can all
558
+ point to a common location, while in production they could be completely different S3 buckets.
559
+
560
+ For example, inside an initializer:
561
+ ~~~ruby
562
+ IOStreams.add_root(:default, "tmp/export")
563
+ IOStreams.add_root(:ftp, "tmp/ftp")
564
+ ~~~
565
+
566
+ `:default` is used whenever a root is not supplied when calling `IOStreams.join`:
567
+ ~~~ruby
568
+ # Uses the :default root: "tmp/export/sample/example.csv"
569
+ path = IOStreams.join("sample", "example.csv")
570
+
571
+ # Uses the :ftp root: "tmp/ftp/sample/example.csv"
572
+ path = IOStreams.join("sample", "example.csv", root: :ftp)
573
+ ~~~
574
+
575
+ The following code:
576
+ ~~~ruby
577
+ path = IOStreams.path("tmp/export", "sample", "example.csv")
578
+ path.writer(:line) do |io|
579
+ io << "Welcome"
580
+ io << "To IOStreams"
581
+ end
582
+ ~~~
583
+
584
+ Can be reduced to:
585
+ ~~~ruby
586
+ path = IOStreams.join("sample", "example.csv")
587
+ path.writer(:line) do |io|
588
+ io << "Welcome"
589
+ io << "To IOStreams"
590
+ end
591
+ ~~~
592
+
593
+ Most importantly the root path information and storage mechanism are externalized from the application code.
594
+
595
+ For example, to make the above code write to S3 in production, change the initializer to:
596
+ ~~~ruby
597
+ IOStreams.add_root(:default, "s3://my-app-bucket-name/export")
598
+ IOStreams.add_root(:ftp, "s3://my-app-ftp-bucket-name/ftp")
599
+ ~~~
600
+
601
+ The code calling `IOStreams.join` does not change at all, see [Config](config) for more examples.
602
+
603
+ ### Restricting access with allowed paths
604
+
605
+ Roots make paths easy to build, but they do not stop a path from leaving the root, for example
606
+ `IOStreams.join("../../etc/passwd")`. When file names come from untrusted input, such as a user or a
607
+ job's configuration, add allowed paths in an initializer to restrict which paths IOStreams can access:
608
+
609
+ ~~~ruby
610
+ IOStreams.add_allowed_path("/var/my_app/uploads")
611
+ IOStreams.add_allowed_path("s3://my-app-bucket-name/export")
612
+ IOStreams.add_allowed_path("sftp://sftp.example.org/outbound")
613
+ IOStreams.add_allowed_path("https://reports.example.org/daily")
614
+ ~~~
615
+
616
+ Once any allowed path has been added, reading, writing, listing, deleting or otherwise accessing a
617
+ path that is not within one of them raises `IOStreams::Errors::AccessDenied`:
618
+
619
+ ~~~ruby
620
+ IOStreams.path("/var/my_app/uploads/file.csv").read
621
+ # => "..."
622
+
623
+ IOStreams.path("/var/my_app/uploads/../secrets.yml").read
624
+ # => IOStreams::Errors::AccessDenied
625
+
626
+ # Check without raising:
627
+ IOStreams.allowed_path?("/etc/passwd")
628
+ # => false
629
+ ~~~
630
+
631
+ Paths are normalized before they are compared:
632
+
633
+ - Local file names are resolved to their real path, so neither `..` nor a symbolic link can be used to
634
+ leave an allowed path. A relative allowed path is resolved against the current working directory
635
+ when it is added.
636
+ - S3 paths must be in the same bucket. Keys containing `.` or `..` segments are denied, since some
637
+ services that implement the S3 API resolve them.
638
+ - SFTP and HTTP paths must have the same host and port, and for HTTP the same scheme. `.` and `..`
639
+ are resolved the way the server resolves them. Every HTTP redirect is checked as well.
640
+
641
+ Notes:
642
+
643
+ - By default no allowed paths are added, and every path is accessible.
644
+ - `each_child` skips children that are not within the allowed paths, for example a symbolic link to a
645
+ file elsewhere.
646
+ - Temp files from `IOStreams.temp_file` are always accessible.
647
+ - `IOStreams.allowed_paths` returns the normalized allowed paths, and `IOStreams.delete_allowed_path`
648
+ removes one.
649
+ - Paths from a scheme added with `IOStreams.register_scheme` are denied, unless its path class
650
+ implements the private method `#allowed_location`.
651
+ - A local file could be replaced with a symbolic link after it is checked but before it is opened.
652
+ Do not allow paths where untrusted users can create files.