iostreams 1.11.0 → 2.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.md +14 -13
- data/Rakefile +52 -0
- data/docs/CLAUDE.md +9 -0
- data/docs/config.md +157 -0
- data/docs/copy_files.md +75 -0
- data/docs/extensions.md +111 -0
- data/docs/formats.md +188 -0
- data/docs/index.md +388 -0
- data/docs/path.md +652 -0
- data/docs/pgp.md +436 -0
- data/docs/streams.md +337 -0
- data/docs/tutorial.md +483 -0
- data/docs/upgrading.md +217 -0
- data/lib/io_streams/builder.rb +71 -11
- data/lib/io_streams/bzip2/reader.rb +25 -2
- data/lib/io_streams/bzip2/writer.rb +26 -2
- data/lib/io_streams/encode/reader.rb +6 -2
- data/lib/io_streams/encode/writer.rb +9 -5
- data/lib/io_streams/errors.rb +4 -0
- data/lib/io_streams/gzip/reader.rb +5 -1
- data/lib/io_streams/gzip/writer.rb +11 -2
- data/lib/io_streams/io_streams.rb +156 -20
- data/lib/io_streams/line/reader.rb +9 -4
- data/lib/io_streams/line/writer.rb +1 -1
- data/lib/io_streams/path.rb +117 -8
- data/lib/io_streams/paths/file.rb +57 -11
- data/lib/io_streams/paths/http.rb +123 -9
- data/lib/io_streams/paths/matcher.rb +3 -3
- data/lib/io_streams/paths/s3.rb +69 -18
- data/lib/io_streams/paths/sftp/net_ssh.rb +104 -0
- data/lib/io_streams/paths/sftp.rb +103 -64
- data/lib/io_streams/pgp/reader.rb +63 -10
- data/lib/io_streams/pgp/writer.rb +111 -30
- data/lib/io_streams/pgp.rb +256 -71
- data/lib/io_streams/reader.rb +14 -5
- data/lib/io_streams/record/reader.rb +75 -6
- data/lib/io_streams/record/writer.rb +3 -4
- data/lib/io_streams/row/reader.rb +1 -1
- data/lib/io_streams/row/writer.rb +1 -1
- data/lib/io_streams/stream.rb +48 -37
- data/lib/io_streams/symmetric_encryption/reader.rb +6 -2
- data/lib/io_streams/symmetric_encryption/writer.rb +8 -4
- data/lib/io_streams/tabular/header.rb +49 -10
- data/lib/io_streams/tabular/parser/array.rb +0 -10
- data/lib/io_streams/tabular/parser/base.rb +10 -0
- data/lib/io_streams/tabular/parser/csv.rb +9 -36
- data/lib/io_streams/tabular/parser/fixed.rb +8 -6
- data/lib/io_streams/tabular/parser/psv.rb +6 -14
- data/lib/io_streams/tabular.rb +5 -10
- data/lib/io_streams/utils.rb +34 -2
- data/lib/io_streams/version.rb +1 -1
- data/lib/io_streams/writer.rb +16 -7
- data/lib/io_streams/xlsx/reader.rb +6 -2
- data/lib/io_streams/zip/reader.rb +4 -0
- data/lib/io_streams/zip/writer.rb +26 -10
- data/lib/iostreams.rb +0 -1
- metadata +46 -112
- data/lib/io_streams/deprecated.rb +0 -216
- data/lib/io_streams/tabular/utility/csv_row.rb +0 -105
- data/test/builder_test.rb +0 -311
- data/test/bzip2_reader_test.rb +0 -27
- data/test/bzip2_writer_test.rb +0 -56
- data/test/deprecated_test.rb +0 -121
- data/test/encode_reader_test.rb +0 -51
- data/test/encode_writer_test.rb +0 -90
- data/test/files/embedded_lines_test.csv +0 -7
- data/test/files/multiple_files.zip +0 -0
- data/test/files/spreadsheet.xlsx +0 -0
- data/test/files/test.csv +0 -4
- data/test/files/test.json +0 -3
- data/test/files/test.psv +0 -4
- data/test/files/text file.txt +0 -3
- data/test/files/text.txt +0 -3
- data/test/files/text.txt.bz2 +0 -0
- data/test/files/text.txt.gz +0 -0
- data/test/files/text.txt.gz.zip +0 -0
- data/test/files/text.zip +0 -0
- data/test/files/text.zip.gz +0 -0
- data/test/files/unclosed_quote_large_test.csv +0 -1658
- data/test/files/unclosed_quote_test.csv +0 -4
- data/test/files/unclosed_quote_test2.csv +0 -3
- data/test/gzip_reader_test.rb +0 -27
- data/test/gzip_writer_test.rb +0 -52
- data/test/io_streams_test.rb +0 -132
- data/test/line_reader_test.rb +0 -325
- data/test/line_writer_test.rb +0 -59
- data/test/minimal_file_reader.rb +0 -25
- data/test/path_test.rb +0 -55
- data/test/paths/file_test.rb +0 -213
- data/test/paths/http_test.rb +0 -34
- data/test/paths/matcher_test.rb +0 -120
- data/test/paths/s3_test.rb +0 -220
- data/test/paths/sftp_test.rb +0 -106
- data/test/pgp_reader_test.rb +0 -46
- data/test/pgp_test.rb +0 -267
- data/test/pgp_writer_test.rb +0 -130
- data/test/record_reader_test.rb +0 -60
- data/test/record_writer_test.rb +0 -82
- data/test/row_reader_test.rb +0 -35
- data/test/row_writer_test.rb +0 -56
- data/test/stream_test.rb +0 -577
- data/test/tabular_test.rb +0 -338
- data/test/test_helper.rb +0 -40
- data/test/utils_test.rb +0 -20
- data/test/xlsx_reader_test.rb +0 -37
- data/test/zip_reader_test.rb +0 -53
- data/test/zip_writer_test.rb +0 -48
data/docs/path.md
ADDED
|
@@ -0,0 +1,652 @@
|
|
|
1
|
+
---
|
|
2
|
+
layout: default
|
|
3
|
+
title: Path
|
|
4
|
+
description: >-
|
|
5
|
+
How a path identifies where a file is stored and how to reach it, with the
|
|
6
|
+
arguments for each location: local disk, AWS S3, SFTP and HTTP.
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
A path identifies _where_ a file is stored and how to reach it, so that the streaming pipeline knows
|
|
10
|
+
where to read the data from or write it to.
|
|
11
|
+
|
|
12
|
+
Create a path with `IOStreams.path`, passing the file name, which may also be a URI, followed by any
|
|
13
|
+
arguments specific to that storage location. IOStreams infers the storage mechanism from the URI
|
|
14
|
+
scheme, so the same call returns a local file path, an S3 path, an SFTP path, and so on, all sharing
|
|
15
|
+
the identical interface.
|
|
16
|
+
|
|
17
|
+
IOStreams supports accessing files in the following places:
|
|
18
|
+
|
|
19
|
+
* File
|
|
20
|
+
* AWS S3
|
|
21
|
+
* Google Cloud Storage (Using the AWS S3 Client)
|
|
22
|
+
* SFTP
|
|
23
|
+
* HTTP(S) (Read only)
|
|
24
|
+
|
|
25
|
+
Are you using another cloud provider and want to add support for your favorite?
|
|
26
|
+
Checkout the supplied [IOStreams S3 path provider](https://github.com/reidmorrison/iostreams/blob/main/lib/io_streams/paths/s3.rb)
|
|
27
|
+
for an example of what is required. Pull requests welcome.
|
|
28
|
+
|
|
29
|
+
### File
|
|
30
|
+
|
|
31
|
+
The simplest case is a file on the local disk:
|
|
32
|
+
|
|
33
|
+
~~~ruby
|
|
34
|
+
path = IOStreams.path("somewhere/example.csv")
|
|
35
|
+
~~~
|
|
36
|
+
|
|
37
|
+
#### Optional Arguments:
|
|
38
|
+
|
|
39
|
+
* `:create_path` set to false to stop IOStreams from automatically creating the output directories
|
|
40
|
+
if they do not exist.
|
|
41
|
+
Default: true
|
|
42
|
+
~~~ruby
|
|
43
|
+
path = IOStreams.path("somewhere/example.csv.gz", create_path: false)
|
|
44
|
+
~~~
|
|
45
|
+
|
|
46
|
+
### AWS S3 (s3://)
|
|
47
|
+
|
|
48
|
+
If the supplied file name string includes a URI. For example if AWS is configured locally:
|
|
49
|
+
|
|
50
|
+
~~~ruby
|
|
51
|
+
path = IOStreams.path("s3://bucket-name/path/example.csv")
|
|
52
|
+
~~~
|
|
53
|
+
|
|
54
|
+
#### Required Arguments:
|
|
55
|
+
|
|
56
|
+
* url [String]
|
|
57
|
+
|
|
58
|
+
Prefix must be: `s3://`, followed by bucket name, followed by key.
|
|
59
|
+
Any query string in the url is added to the S3 request parameters, for example
|
|
60
|
+
`s3://my-bucket-name/file_name.csv?acl=bucket-owner-full-control`.
|
|
61
|
+
Examples:
|
|
62
|
+
s3://my-bucket-name/file_name.txt
|
|
63
|
+
s3://my-bucket-name/some_path/file_name.csv
|
|
64
|
+
|
|
65
|
+
Security warning: do not interpolate an untrusted file name into the url, since a name such as
|
|
66
|
+
`file.csv?acl=public-read` would set request parameters. Join it onto the path instead, which
|
|
67
|
+
does not parse it as a query:
|
|
68
|
+
`IOStreams.path("s3://my-bucket-name/uploads").join(untrusted_name)`
|
|
69
|
+
|
|
70
|
+
#### Optional Arguments:
|
|
71
|
+
|
|
72
|
+
* :access_key_id [String]
|
|
73
|
+
|
|
74
|
+
AWS Access Key Id to use to access this bucket.
|
|
75
|
+
|
|
76
|
+
* :secret_access_key [String]
|
|
77
|
+
|
|
78
|
+
AWS Secret Access Key Id to use to access this bucket.
|
|
79
|
+
|
|
80
|
+
* :region [String]
|
|
81
|
+
|
|
82
|
+
The AWS region to connect to.
|
|
83
|
+
Default: the region set in the environment variables or credential files.
|
|
84
|
+
|
|
85
|
+
* :client [Aws::S3::Client | Hash]
|
|
86
|
+
|
|
87
|
+
Supply the AWS S3 Client instance to use for this path.
|
|
88
|
+
Or, when a Hash, build a new client using the hash parameters.
|
|
89
|
+
|
|
90
|
+
~~~ruby
|
|
91
|
+
client = Aws::S3::Client.new(endpoint: "https://s3.test.com")
|
|
92
|
+
path = IOStreams.path("s3://bucket/path/file_name.txt", client: client)
|
|
93
|
+
|
|
94
|
+
# Or, pass the client parameters directly:
|
|
95
|
+
path = IOStreams.path("s3://bucket/path/file_name.txt", client: {endpoint: "https://s3.test.com"})
|
|
96
|
+
~~~
|
|
97
|
+
|
|
98
|
+
Writer specific options:
|
|
99
|
+
|
|
100
|
+
* :acl [String]
|
|
101
|
+
|
|
102
|
+
The canned ACL to apply to the object.
|
|
103
|
+
|
|
104
|
+
* :cache_control [String]
|
|
105
|
+
|
|
106
|
+
Specifies caching behavior along the request/reply chain.
|
|
107
|
+
|
|
108
|
+
* :content_disposition [String]
|
|
109
|
+
|
|
110
|
+
Specifies presentational information for the object.
|
|
111
|
+
|
|
112
|
+
* :content_encoding [String]
|
|
113
|
+
|
|
114
|
+
Specifies what content encodings have been applied to the object and
|
|
115
|
+
thus what decoding mechanisms must be applied to obtain the media-type
|
|
116
|
+
referenced by the Content-Type header field.
|
|
117
|
+
|
|
118
|
+
* :content_language [String]
|
|
119
|
+
|
|
120
|
+
The language the content is in.
|
|
121
|
+
|
|
122
|
+
* :content_length [Integer]
|
|
123
|
+
|
|
124
|
+
Size of the body in bytes. This parameter is useful when the size of
|
|
125
|
+
the body cannot be determined automatically.
|
|
126
|
+
|
|
127
|
+
* :content_md5 [String]
|
|
128
|
+
|
|
129
|
+
The base64-encoded 128-bit MD5 digest of the part data. This parameter
|
|
130
|
+
is auto-populated when using the command from the CLI. This parameted
|
|
131
|
+
is required if object lock parameters are specified.
|
|
132
|
+
|
|
133
|
+
* :content_type [String]
|
|
134
|
+
|
|
135
|
+
A standard MIME type describing the format of the object data.
|
|
136
|
+
|
|
137
|
+
* :expires [Time,DateTime,Date,Integer,String]
|
|
138
|
+
|
|
139
|
+
The date and time at which the object is no longer cacheable.
|
|
140
|
+
|
|
141
|
+
* :grant_full_control [String]
|
|
142
|
+
|
|
143
|
+
Gives the grantee READ, READ\_ACP, and WRITE\_ACP permissions on the
|
|
144
|
+
object.
|
|
145
|
+
|
|
146
|
+
* :grant_read [String]
|
|
147
|
+
|
|
148
|
+
Allows grantee to read the object data and its metadata.
|
|
149
|
+
|
|
150
|
+
* :grant_read_acp [String]
|
|
151
|
+
|
|
152
|
+
Allows grantee to read the object ACL.
|
|
153
|
+
|
|
154
|
+
* :grant_write_acp [String]
|
|
155
|
+
|
|
156
|
+
Allows grantee to write the ACL for the applicable object.
|
|
157
|
+
|
|
158
|
+
* :metadata [Hash<String,String>]
|
|
159
|
+
|
|
160
|
+
A map of metadata to store with the object in S3.
|
|
161
|
+
|
|
162
|
+
* :server_side_encryption [String]
|
|
163
|
+
|
|
164
|
+
The Server-side encryption algorithm used when storing this object in
|
|
165
|
+
S3 (e.g., AES256, aws:kms).
|
|
166
|
+
|
|
167
|
+
* :storage_class [String]
|
|
168
|
+
|
|
169
|
+
The type of storage to use for the object. Defaults to 'STANDARD'.
|
|
170
|
+
|
|
171
|
+
* :website_redirect_location [String]
|
|
172
|
+
|
|
173
|
+
If the bucket is configured as a website, redirects requests for this
|
|
174
|
+
object to another object in the same bucket or to an external URL.
|
|
175
|
+
Amazon S3 stores the value of this header in the object metadata.
|
|
176
|
+
|
|
177
|
+
* :sse_customer_algorithm [String]
|
|
178
|
+
|
|
179
|
+
Specifies the algorithm to use to when encrypting the object (e.g.,
|
|
180
|
+
AES256).
|
|
181
|
+
|
|
182
|
+
* :sse_customer_key [String]
|
|
183
|
+
|
|
184
|
+
Specifies the customer-provided encryption key for Amazon S3 to use in
|
|
185
|
+
encrypting data. This value is used to store the object and then it is
|
|
186
|
+
discarded; Amazon does not store the encryption key. The key must be
|
|
187
|
+
appropriate for use with the algorithm specified in the
|
|
188
|
+
x-amz-server-side-encryption-customer-algorithm header.
|
|
189
|
+
|
|
190
|
+
* :sse_customer_key_md5 [String]
|
|
191
|
+
|
|
192
|
+
Specifies the 128-bit MD5 digest of the encryption key according to
|
|
193
|
+
RFC 1321. Amazon S3 uses this header for a message integrity check to
|
|
194
|
+
ensure the encryption key was transmitted without error.
|
|
195
|
+
|
|
196
|
+
* :ssekms_key_id [String]
|
|
197
|
+
|
|
198
|
+
Specifies the AWS KMS key ID to use for object encryption. All GET and
|
|
199
|
+
PUT requests for an object protected by AWS KMS will fail if not made
|
|
200
|
+
via SSL or using SigV4. Documentation on configuring any of the
|
|
201
|
+
officially supported AWS SDKs and CLI can be found at
|
|
202
|
+
http://docs.aws.amazon.com/AmazonS3/latest/dev/UsingAWSSDK.html#specify-signature-version
|
|
203
|
+
|
|
204
|
+
* :ssekms_encryption_context [String]
|
|
205
|
+
|
|
206
|
+
Specifies the AWS KMS Encryption Context to use for object encryption.
|
|
207
|
+
The value of this header is a base64-encoded UTF-8 string holding JSON
|
|
208
|
+
with the encryption context key-value pairs.
|
|
209
|
+
|
|
210
|
+
* :request_payer [String]
|
|
211
|
+
|
|
212
|
+
Confirms that the requester knows that she or he will be charged for
|
|
213
|
+
the request. Bucket owners need not specify this parameter in their
|
|
214
|
+
requests. Documentation on downloading objects from requester pays
|
|
215
|
+
buckets can be found at
|
|
216
|
+
http://docs.aws.amazon.com/AmazonS3/latest/dev/ObjectsinRequesterPaysBuckets.html
|
|
217
|
+
|
|
218
|
+
* :tagging [String]
|
|
219
|
+
|
|
220
|
+
The tag-set for the object. The tag-set must be encoded as URL Query
|
|
221
|
+
parameters. (For example, "Key1=Value1")
|
|
222
|
+
|
|
223
|
+
* :object_lock_mode [String]
|
|
224
|
+
|
|
225
|
+
The object lock mode that you want to apply to this object.
|
|
226
|
+
|
|
227
|
+
* object_lock_retain_until_date: [Time,DateTime,Date,Integer,String]
|
|
228
|
+
|
|
229
|
+
The date and time when you want this object's object lock to expire.
|
|
230
|
+
|
|
231
|
+
* object_lock_legal_hold_status: [String]
|
|
232
|
+
The Legal Hold status that you want to apply to the specified object.
|
|
233
|
+
|
|
234
|
+
### SFTP (sftp://)
|
|
235
|
+
|
|
236
|
+
If the supplied file name string includes the `sftp` URI.
|
|
237
|
+
|
|
238
|
+
~~~ruby
|
|
239
|
+
path = IOStreams.path("sftp://hostname/path/example.csv")
|
|
240
|
+
~~~
|
|
241
|
+
|
|
242
|
+
IOStreams reads and writes SFTP files by shelling out to the `sftp` command line program,
|
|
243
|
+
so it must be installed and on the `PATH`. When a password is supplied the `sshpass`
|
|
244
|
+
program is also required to pass the password to `sftp`. Additionally the `net-sftp` gem
|
|
245
|
+
must be added to the `Gemfile` to use `each_child`. `each_child` lists the files within the path's directory,
|
|
246
|
+
or within the login directory when the url has no path, for example `sftp://hostname`.
|
|
247
|
+
|
|
248
|
+
Read a file from a remote sftp server.
|
|
249
|
+
~~~ruby
|
|
250
|
+
IOStreams.path("sftp://example.org/path/file.txt",
|
|
251
|
+
username: "jbloggs",
|
|
252
|
+
password: "secret").
|
|
253
|
+
reader do |input|
|
|
254
|
+
puts input.read
|
|
255
|
+
end
|
|
256
|
+
~~~
|
|
257
|
+
|
|
258
|
+
Raises `IOStreams::Errors::CommunicationsFailure` when the file could not be read or written.
|
|
259
|
+
|
|
260
|
+
Write to a file on a remote sftp server.
|
|
261
|
+
~~~ruby
|
|
262
|
+
IOStreams.path("sftp://example.org/path/file.txt",
|
|
263
|
+
username: "jbloggs",
|
|
264
|
+
password: "secret").
|
|
265
|
+
writer do |output|
|
|
266
|
+
output.write('Hello World')
|
|
267
|
+
end
|
|
268
|
+
~~~
|
|
269
|
+
|
|
270
|
+
Display the contents of a remote file, supplying the username and password in the url.
|
|
271
|
+
Note that `#to_s` then includes the password, so prefer the `username:` and `password:` arguments above:
|
|
272
|
+
~~~ruby
|
|
273
|
+
IOStreams.path("sftp://jack:OpenSesame@test.com:22/path/file_name.csv").reader do |io|
|
|
274
|
+
puts io.read
|
|
275
|
+
end
|
|
276
|
+
~~~
|
|
277
|
+
|
|
278
|
+
Use an identity file instead of a password to authenticate:
|
|
279
|
+
~~~ruby
|
|
280
|
+
path = IOStreams.path("sftp://test.com/path/file_name.csv",
|
|
281
|
+
username: "jack",
|
|
282
|
+
ssh_options: {IdentityFile: "~/.ssh/private_key"})
|
|
283
|
+
path.reader do |io|
|
|
284
|
+
puts io.read
|
|
285
|
+
end
|
|
286
|
+
~~~
|
|
287
|
+
|
|
288
|
+
Pass in the IdentityKey itself instead of a password to authenticate.
|
|
289
|
+
For example, retrieve the identity key stored in Secret Config:
|
|
290
|
+
~~~ruby
|
|
291
|
+
identity_key = SecretConfig.fetch("suppliers/sftp/identity_key")
|
|
292
|
+
|
|
293
|
+
path = IOStreams.path("sftp://test.com/path/file_name.csv",
|
|
294
|
+
username: "jack",
|
|
295
|
+
ssh_options: {IdentityKey: identity_key})
|
|
296
|
+
path.reader do |io|
|
|
297
|
+
puts io.read
|
|
298
|
+
end
|
|
299
|
+
~~~
|
|
300
|
+
|
|
301
|
+
#### Required Arguments:
|
|
302
|
+
|
|
303
|
+
* url [String]
|
|
304
|
+
|
|
305
|
+
Prefix must be: `sftp://`, followed by host name, followed by file name.
|
|
306
|
+
Format:
|
|
307
|
+
"sftp://<host_name>/<file_name>"
|
|
308
|
+
"sftp://username:password@hostname:22/path/file_name"
|
|
309
|
+
|
|
310
|
+
A username and password supplied in the url remain part of it, so `#to_s` returns them,
|
|
311
|
+
as does any log or error message that includes the path. To keep them out of logs, supply them
|
|
312
|
+
with the `username:` and `password:` arguments instead.
|
|
313
|
+
|
|
314
|
+
#### Optional Arguments:
|
|
315
|
+
|
|
316
|
+
* username: [String]
|
|
317
|
+
|
|
318
|
+
Name of user to login with.
|
|
319
|
+
|
|
320
|
+
* password: [String]
|
|
321
|
+
|
|
322
|
+
Password for the user.
|
|
323
|
+
|
|
324
|
+
* ssh_options: [Hash]
|
|
325
|
+
|
|
326
|
+
* IdentityFile [String]
|
|
327
|
+
|
|
328
|
+
Path to the local identity (private key) file to authenticate with, instead of a password.
|
|
329
|
+
|
|
330
|
+
* IdentityKey [String]
|
|
331
|
+
|
|
332
|
+
The identity (private key) itself, supplied as a string.
|
|
333
|
+
Under the covers the key is written to a temp file and then passed as `IdentityFile`.
|
|
334
|
+
|
|
335
|
+
* HostKey [String]
|
|
336
|
+
|
|
337
|
+
The expected SSH host key presented by the remote host, instead of storing it in the
|
|
338
|
+
`known_hosts` file. It must contain the entire line that would be stored in `known_hosts`,
|
|
339
|
+
including the hostname, ip address, key type and key value. The easiest way to generate
|
|
340
|
+
the required value is with `ssh-keyscan hostname`.
|
|
341
|
+
Under the covers the value is written to a temp file and then passed as `UserKnownHostsFile`.
|
|
342
|
+
|
|
343
|
+
* Any other options supported by ssh_config.
|
|
344
|
+
`man ssh_config` to see all available options.
|
|
345
|
+
|
|
346
|
+
`each_child` lists files with the `net-sftp` gem instead of the `sftp` program, so it only supports
|
|
347
|
+
these ssh options: `HostKey`, `IdentityKey`, `IdentityFile`, `UserKnownHostsFile`,
|
|
348
|
+
`StrictHostKeyChecking`, `ConnectTimeout`, `ServerAliveInterval`, `ServerAliveCountMax` and
|
|
349
|
+
`LogLevel`. Any other option raises `ArgumentError`. Unlike the `sftp` program, `net-sftp` needs the
|
|
350
|
+
`ed25519` and `bcrypt_pbkdf` gems to use ed25519 host or identity keys.
|
|
351
|
+
|
|
352
|
+
Notes:
|
|
353
|
+
* Since the `sftp` program operates on local files, reading from or writing to an SFTP path
|
|
354
|
+
streams through a local temp file behind the scenes.
|
|
355
|
+
|
|
356
|
+
### HTTP (http://, https://)
|
|
357
|
+
|
|
358
|
+
Read from a remote file over HTTP or HTTPS using an HTTP Get.
|
|
359
|
+
|
|
360
|
+
~~~ruby
|
|
361
|
+
IOStreams.path('https://www5.fdic.gov/idasp/Offices2.zip').read
|
|
362
|
+
~~~
|
|
363
|
+
|
|
364
|
+
Notes:
|
|
365
|
+
* Since Net::HTTP download only supports a push stream, the data is streamed into a tempfile first.
|
|
366
|
+
* Currently writing to an HTTP(S) server is not supported. Up to submitting a Pull Request with capability?
|
|
367
|
+
|
|
368
|
+
#### Required Arguments:
|
|
369
|
+
|
|
370
|
+
* url [String]
|
|
371
|
+
|
|
372
|
+
Prefix must be: `http://`, or `https://` followed by host name, followed by path and file name.
|
|
373
|
+
Also supports passing the username and password for basic authentication in the URI.
|
|
374
|
+
|
|
375
|
+
Format:
|
|
376
|
+
* http://hostname/path/file_name
|
|
377
|
+
* https://username:password@hostname/path/file_name
|
|
378
|
+
|
|
379
|
+
A username and password supplied in the url remain part of it, so `#to_s` returns them,
|
|
380
|
+
as does any log or error message that includes the path. To keep them out of logs, supply them
|
|
381
|
+
with the `username:` and `password:` arguments instead.
|
|
382
|
+
|
|
383
|
+
#### Optional Arguments:
|
|
384
|
+
|
|
385
|
+
* username: [String]
|
|
386
|
+
|
|
387
|
+
When supplied, basic authentication is used with the username and password.
|
|
388
|
+
|
|
389
|
+
* password: [String]
|
|
390
|
+
|
|
391
|
+
Password to use use with basic authentication when the username is supplied.
|
|
392
|
+
|
|
393
|
+
* parameters: [Hash]
|
|
394
|
+
|
|
395
|
+
Query parameters to append to the url as a query string.
|
|
396
|
+
For example, `parameters: {"type" => "csv"}` appends `?type=csv` to the url.
|
|
397
|
+
|
|
398
|
+
* http_redirect_count: [Integer]
|
|
399
|
+
|
|
400
|
+
Maximum number of http redirects to follow.
|
|
401
|
+
Set to `0` to disable following redirects entirely.
|
|
402
|
+
Default: `10`
|
|
403
|
+
|
|
404
|
+
* allow_hosts: [String | Array<String>]
|
|
405
|
+
|
|
406
|
+
Optional allow-list of host names that may be contacted. It is applied both to the
|
|
407
|
+
supplied url and to every redirect that is followed; a request to any other host raises
|
|
408
|
+
`IOStreams::Errors::CommunicationsFailure`.
|
|
409
|
+
Default: `nil` (any host is allowed).
|
|
410
|
+
|
|
411
|
+
* maximum_file_size: [Integer]
|
|
412
|
+
|
|
413
|
+
Optional maximum number of bytes to download. When the response body exceeds this size the
|
|
414
|
+
download is aborted with an `IOStreams::Errors::CommunicationsFailure`.
|
|
415
|
+
Default: `nil` (no limit).
|
|
416
|
+
|
|
417
|
+
~~~ruby
|
|
418
|
+
path = IOStreams.path("http://hostname/path/example.csv")
|
|
419
|
+
~~~
|
|
420
|
+
|
|
421
|
+
#### Security: untrusted URLs (SSRF)
|
|
422
|
+
|
|
423
|
+
Reading an HTTP(S) path causes the application to issue a request to the host named in the url.
|
|
424
|
+
When the url, or any part of it, can be influenced by untrusted input, an attacker can point it
|
|
425
|
+
at internal services or cloud metadata endpoints (Server Side Request Forgery).
|
|
426
|
+
|
|
427
|
+
Because redirect targets are chosen by the remote server, validating only the url that is passed
|
|
428
|
+
in is not sufficient: a trusted (or compromised) server can redirect the request to an internal
|
|
429
|
+
address. IOStreams provides a few controls to reduce this exposure:
|
|
430
|
+
|
|
431
|
+
* Restrict which hosts may be contacted, including across redirects:
|
|
432
|
+
|
|
433
|
+
~~~ruby
|
|
434
|
+
IOStreams.path("https://supplier.example.com/report.csv", allow_hosts: ["supplier.example.com"]).read
|
|
435
|
+
~~~
|
|
436
|
+
|
|
437
|
+
* Disable redirects entirely for untrusted urls:
|
|
438
|
+
|
|
439
|
+
~~~ruby
|
|
440
|
+
IOStreams.path(untrusted_url, http_redirect_count: 0).read
|
|
441
|
+
~~~
|
|
442
|
+
|
|
443
|
+
* Cap the download size to avoid unbounded (denial of service) responses:
|
|
444
|
+
|
|
445
|
+
~~~ruby
|
|
446
|
+
IOStreams.path(untrusted_url, maximum_file_size: 50 * 1024 * 1024).read
|
|
447
|
+
~~~
|
|
448
|
+
|
|
449
|
+
Basic authentication credentials are only ever sent to the original host. They are not resent
|
|
450
|
+
when a redirect points at a different scheme, host, or port, so a redirect cannot leak them to
|
|
451
|
+
another server. For stronger guarantees, route these downloads through an egress proxy or network
|
|
452
|
+
policy that blocks private, loopback, and link-local (cloud metadata) addresses.
|
|
453
|
+
|
|
454
|
+
Similarly when using https:
|
|
455
|
+
|
|
456
|
+
~~~ruby
|
|
457
|
+
path = IOStreams.path("https://hostname/path/example.csv")
|
|
458
|
+
~~~
|
|
459
|
+
|
|
460
|
+
This time IOStreams inferred that the file lives on an HTTP Server and returns `IOStreams::Paths::HTTP`.
|
|
461
|
+
|
|
462
|
+
### Path Operations
|
|
463
|
+
|
|
464
|
+
Paths support common file operations, regardless of where the file is stored:
|
|
465
|
+
|
|
466
|
+
~~~ruby
|
|
467
|
+
path = IOStreams.path("sample/example.csv")
|
|
468
|
+
|
|
469
|
+
# Does the file exist?
|
|
470
|
+
path.exist?
|
|
471
|
+
# => true
|
|
472
|
+
|
|
473
|
+
# Size of the file in bytes.
|
|
474
|
+
path.size
|
|
475
|
+
# => 64
|
|
476
|
+
|
|
477
|
+
# Delete the file.
|
|
478
|
+
path.delete
|
|
479
|
+
|
|
480
|
+
# Move the file to another path, returning the target path.
|
|
481
|
+
path.move_to("sample/moved.csv")
|
|
482
|
+
|
|
483
|
+
# Create the directory path, when it does not already exist.
|
|
484
|
+
IOStreams.path("sample/data").mkpath
|
|
485
|
+
~~~
|
|
486
|
+
|
|
487
|
+
Inspect the components of a path's file name:
|
|
488
|
+
|
|
489
|
+
~~~ruby
|
|
490
|
+
# The last component of the path.
|
|
491
|
+
IOStreams.path("/home/gumby/work/ruby.rb").basename
|
|
492
|
+
# => "ruby.rb"
|
|
493
|
+
|
|
494
|
+
# Remove a specific suffix from the file name.
|
|
495
|
+
IOStreams.path("/home/gumby/work/ruby.rb").basename(".rb")
|
|
496
|
+
# => "ruby"
|
|
497
|
+
|
|
498
|
+
# Remove any extension by supplying ".*".
|
|
499
|
+
IOStreams.path("/home/gumby/work/ruby.rb").basename(".*")
|
|
500
|
+
# => "ruby"
|
|
501
|
+
|
|
502
|
+
# The directory portion of the path.
|
|
503
|
+
IOStreams.path("a/b/d/test.rb").dirname
|
|
504
|
+
# => "a/b/d"
|
|
505
|
+
|
|
506
|
+
# The extension, including the leading period.
|
|
507
|
+
IOStreams.path("a/b/d/test.rb").extname
|
|
508
|
+
# => ".rb"
|
|
509
|
+
|
|
510
|
+
# The extension, without the leading period.
|
|
511
|
+
IOStreams.path("a/b/d/test.rb").extension
|
|
512
|
+
# => "rb"
|
|
513
|
+
~~~
|
|
514
|
+
|
|
515
|
+
Notes:
|
|
516
|
+
* `basename`, `dirname`, `extname`, and `extension` return `nil` when no file name was set.
|
|
517
|
+
* A leading period on a dotfile is not treated as an extension, so `.profile` has no extension,
|
|
518
|
+
while `.profile.sh` has the extension `sh`.
|
|
519
|
+
* A file name ending in a period, such as `foo.`, returns an empty string for the extension.
|
|
520
|
+
|
|
521
|
+
Iterate over the files in a path using a wildcard pattern:
|
|
522
|
+
|
|
523
|
+
~~~ruby
|
|
524
|
+
IOStreams.path("sample").each_child("*.csv") do |child|
|
|
525
|
+
puts child
|
|
526
|
+
end
|
|
527
|
+
|
|
528
|
+
# Recursively, including sub-directories:
|
|
529
|
+
IOStreams.path("sample").each_child("**/*.csv") do |child|
|
|
530
|
+
puts child
|
|
531
|
+
end
|
|
532
|
+
~~~
|
|
533
|
+
|
|
534
|
+
`each_child` is also available directly on `IOStreams` when the pattern includes the full path:
|
|
535
|
+
|
|
536
|
+
~~~ruby
|
|
537
|
+
IOStreams.each_child("sample/**/*.csv") { |child| puts child }
|
|
538
|
+
~~~
|
|
539
|
+
|
|
540
|
+
Notes:
|
|
541
|
+
* These operations are supported by File and S3 paths. SFTP supports `each_child`,
|
|
542
|
+
and HTTP paths are read-only so they do not support any of them.
|
|
543
|
+
* By default `each_child` patterns are case-insensitive and hidden files are excluded.
|
|
544
|
+
Supply `case_sensitive: true` or `hidden: true` to change this behavior.
|
|
545
|
+
|
|
546
|
+
### Using root paths
|
|
547
|
+
|
|
548
|
+
Roots allow paths to reference a particular root directory, so that all path names are appended to that root.
|
|
549
|
+
By using `IOStreams.join` instead of `IOStreams.path`, the storage location is no longer embedded in the
|
|
550
|
+
application code, it is configured once at startup.
|
|
551
|
+
|
|
552
|
+
The primary purpose of roots is to allow the exact same code to run in production and development,
|
|
553
|
+
yet use completely different data sources in each. For example, in production the root can point to an
|
|
554
|
+
S3 bucket, while in development it points to the local file system.
|
|
555
|
+
|
|
556
|
+
Roots are configured via an initializer at startup. Multiple roots can be setup, for example one for
|
|
557
|
+
input files, another for output files, another for reports, etc. During development the roots can all
|
|
558
|
+
point to a common location, while in production they could be completely different S3 buckets.
|
|
559
|
+
|
|
560
|
+
For example, inside an initializer:
|
|
561
|
+
~~~ruby
|
|
562
|
+
IOStreams.add_root(:default, "tmp/export")
|
|
563
|
+
IOStreams.add_root(:ftp, "tmp/ftp")
|
|
564
|
+
~~~
|
|
565
|
+
|
|
566
|
+
`:default` is used whenever a root is not supplied when calling `IOStreams.join`:
|
|
567
|
+
~~~ruby
|
|
568
|
+
# Uses the :default root: "tmp/export/sample/example.csv"
|
|
569
|
+
path = IOStreams.join("sample", "example.csv")
|
|
570
|
+
|
|
571
|
+
# Uses the :ftp root: "tmp/ftp/sample/example.csv"
|
|
572
|
+
path = IOStreams.join("sample", "example.csv", root: :ftp)
|
|
573
|
+
~~~
|
|
574
|
+
|
|
575
|
+
The following code:
|
|
576
|
+
~~~ruby
|
|
577
|
+
path = IOStreams.path("tmp/export", "sample", "example.csv")
|
|
578
|
+
path.writer(:line) do |io|
|
|
579
|
+
io << "Welcome"
|
|
580
|
+
io << "To IOStreams"
|
|
581
|
+
end
|
|
582
|
+
~~~
|
|
583
|
+
|
|
584
|
+
Can be reduced to:
|
|
585
|
+
~~~ruby
|
|
586
|
+
path = IOStreams.join("sample", "example.csv")
|
|
587
|
+
path.writer(:line) do |io|
|
|
588
|
+
io << "Welcome"
|
|
589
|
+
io << "To IOStreams"
|
|
590
|
+
end
|
|
591
|
+
~~~
|
|
592
|
+
|
|
593
|
+
Most importantly the root path information and storage mechanism are externalized from the application code.
|
|
594
|
+
|
|
595
|
+
For example, to make the above code write to S3 in production, change the initializer to:
|
|
596
|
+
~~~ruby
|
|
597
|
+
IOStreams.add_root(:default, "s3://my-app-bucket-name/export")
|
|
598
|
+
IOStreams.add_root(:ftp, "s3://my-app-ftp-bucket-name/ftp")
|
|
599
|
+
~~~
|
|
600
|
+
|
|
601
|
+
The code calling `IOStreams.join` does not change at all, see [Config](config) for more examples.
|
|
602
|
+
|
|
603
|
+
### Restricting access with allowed paths
|
|
604
|
+
|
|
605
|
+
Roots make paths easy to build, but they do not stop a path from leaving the root, for example
|
|
606
|
+
`IOStreams.join("../../etc/passwd")`. When file names come from untrusted input, such as a user or a
|
|
607
|
+
job's configuration, add allowed paths in an initializer to restrict which paths IOStreams can access:
|
|
608
|
+
|
|
609
|
+
~~~ruby
|
|
610
|
+
IOStreams.add_allowed_path("/var/my_app/uploads")
|
|
611
|
+
IOStreams.add_allowed_path("s3://my-app-bucket-name/export")
|
|
612
|
+
IOStreams.add_allowed_path("sftp://sftp.example.org/outbound")
|
|
613
|
+
IOStreams.add_allowed_path("https://reports.example.org/daily")
|
|
614
|
+
~~~
|
|
615
|
+
|
|
616
|
+
Once any allowed path has been added, reading, writing, listing, deleting or otherwise accessing a
|
|
617
|
+
path that is not within one of them raises `IOStreams::Errors::AccessDenied`:
|
|
618
|
+
|
|
619
|
+
~~~ruby
|
|
620
|
+
IOStreams.path("/var/my_app/uploads/file.csv").read
|
|
621
|
+
# => "..."
|
|
622
|
+
|
|
623
|
+
IOStreams.path("/var/my_app/uploads/../secrets.yml").read
|
|
624
|
+
# => IOStreams::Errors::AccessDenied
|
|
625
|
+
|
|
626
|
+
# Check without raising:
|
|
627
|
+
IOStreams.allowed_path?("/etc/passwd")
|
|
628
|
+
# => false
|
|
629
|
+
~~~
|
|
630
|
+
|
|
631
|
+
Paths are normalized before they are compared:
|
|
632
|
+
|
|
633
|
+
- Local file names are resolved to their real path, so neither `..` nor a symbolic link can be used to
|
|
634
|
+
leave an allowed path. A relative allowed path is resolved against the current working directory
|
|
635
|
+
when it is added.
|
|
636
|
+
- S3 paths must be in the same bucket. Keys containing `.` or `..` segments are denied, since some
|
|
637
|
+
services that implement the S3 API resolve them.
|
|
638
|
+
- SFTP and HTTP paths must have the same host and port, and for HTTP the same scheme. `.` and `..`
|
|
639
|
+
are resolved the way the server resolves them. Every HTTP redirect is checked as well.
|
|
640
|
+
|
|
641
|
+
Notes:
|
|
642
|
+
|
|
643
|
+
- By default no allowed paths are added, and every path is accessible.
|
|
644
|
+
- `each_child` skips children that are not within the allowed paths, for example a symbolic link to a
|
|
645
|
+
file elsewhere.
|
|
646
|
+
- Temp files from `IOStreams.temp_file` are always accessible.
|
|
647
|
+
- `IOStreams.allowed_paths` returns the normalized allowed paths, and `IOStreams.delete_allowed_path`
|
|
648
|
+
removes one.
|
|
649
|
+
- Paths from a scheme added with `IOStreams.register_scheme` are denied, unless its path class
|
|
650
|
+
implements the private method `#allowed_location`.
|
|
651
|
+
- A local file could be replaced with a symbolic link after it is checked but before it is opened.
|
|
652
|
+
Do not allow paths where untrusted users can create files.
|