herringbone 0.4.0 → 0.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 826898400d423a36ccd89740d408da3d367726cf84c6cf33c4067a3c54b8861d
4
- data.tar.gz: 0667170a3ede8376299aff04b4081259533c8877f9823ad8fcc9aee916a86650
3
+ metadata.gz: 4aca873b946a39da4bee947137c6914f0b321f95266f252f35487087e40ce3f4
4
+ data.tar.gz: 6d7f439f61d6d871cfbf8682e2af71bf2c502b5360355570aabeed373240d0d7
5
5
  SHA512:
6
- metadata.gz: 21984e5695bf75e086fa091a4c69456edd4fc7ef21ea2d73c89e0ad3de1316238d2f66cf9d8c86183c5e962a1ce965aaeac8036dc1a5c1a43aa6a3b85c983405
7
- data.tar.gz: 9377178bd66028014087cc55ed255c4abc2ed1c0217cdfbd86d5f9fe647a4ac0d741a973415f393c31fc7af83d653ce14df33d7a7629dee69aaed554e1fa5719
6
+ metadata.gz: 9cc7d70651be9426aec7608b1fa456545ea0c5318c4aaf193a67b14c7d2371a7173c62cca23a73f603d412dd2ddafa5d8f8e0833b08e6efc9c0632fe367e56d2
7
+ data.tar.gz: bb1b6ed8e32eb6bcbdb0c29ca1dbd773a74dd2f5d2587253a4a297d76a27ffddc7f0a508baabb9a3d45d0bcd40423f3f5ea531cdc4adf3b4564c3344f61f3fda
data/CHANGELOG.md CHANGED
@@ -2,6 +2,20 @@
2
2
 
3
3
  ## Unreleased
4
4
 
5
+ ## 0.5.0
6
+
7
+ - `Herringbone.redact` and `Herringbone::Redaction` rewrite a file with rows deleted, values
8
+ replaced or columns dropped, for GDPR erasure and pseudonymization, copying the row groups they
9
+ don't touch byte for byte.
10
+ - `Herringbone::Redaction#affects?` says whether a redaction would change a file, reading only
11
+ statistics, bloom filters and the `where` columns.
12
+ - `Herringbone::Writer#inspect` (and `SimpleWriter`, `ByteValues`) prints a one-line summary with
13
+ the writer's state and row count instead of dumping the buffered rows.
14
+ - **Breaking:** schema DSL blocks receive the builder as a parameter instead of running with
15
+ `instance_eval`: `Herringbone::Schema.define { |s| s.int64 :id }`, likewise for `Schema.infer`,
16
+ `Herringbone.write`, `SimpleWriter.new` and nested `struct`/`list`/`map` blocks. A block without
17
+ a parameter raises `ArgumentError`.
18
+
5
19
  ## 0.4.0
6
20
 
7
21
  - herringbone is now Ractor-enabled! Reading and writing work from any Ractor on Ruby 3.1+, and schemas can be passed to Ractors with
data/README.md CHANGED
@@ -148,19 +148,19 @@ with `each_batch` it can differ between batches.
148
148
  ## Writing
149
149
 
150
150
  ```ruby
151
- schema = Herringbone::Schema.define do
152
- int64 :id, null: false
153
- string :name
154
- enum :status, values: %w[pending paid shipped]
155
- list :tags, :string
156
- map :scores, :string, :double
157
- struct :address do
158
- string :city
159
- string :zip
151
+ schema = Herringbone::Schema.define do |s|
152
+ s.int64 :id, null: false
153
+ s.string :name
154
+ s.enum :status, values: %w[pending paid shipped]
155
+ s.list :tags, :string
156
+ s.map :scores, :string, :double
157
+ s.struct :address do |address|
158
+ address.string :city
159
+ address.string :zip
160
160
  end
161
- decimal :price, precision: 12, scale: 2
162
- json :payload
163
- timestamp :created_at
161
+ s.decimal :price, precision: 12, scale: 2
162
+ s.json :payload
163
+ s.timestamp :created_at
164
164
  end
165
165
 
166
166
  File.open("out.parquet", "wb") do |file|
@@ -177,7 +177,7 @@ end
177
177
 
178
178
  `Herringbone.write(io, rows)` writes an Enumerable of rows in one go, inferring the schema from the
179
179
  first 1000 rows unless `schema:` is given; fields declared in a block replace inferred ones:
180
- `Herringbone.write(io, rows) { json :payload }`. The rows are iterated once, holding back only
180
+ `Herringbone.write(io, rows) { |s| s.json :payload }`. The rows are iterated once, holding back only
181
181
  those first 1000, so lazy Enumerators and cursors that can't be rewound work. A later row that
182
182
  doesn't fit the inferred types raises `Herringbone::SchemaMismatch`, which explains what was
183
183
  inferred and how to declare the column, and leaves the file unfinished.
@@ -188,11 +188,15 @@ the `Writer.open` block raises (or `#abort` is called), no footer is written and
188
188
  so far is left for you to discard. To replace a file only once it is complete, write to a temporary
189
189
  file and rename it.
190
190
 
191
+ The block receives the schema builder, and the blocks of `struct`, `list` and `map` receive a
192
+ builder of their own, so the block can still call your methods and read your instance variables.
193
+ A block that takes no parameter raises `ArgumentError`.
194
+
191
195
  Column types: `boolean int8 int16 int32 int64 uint8 uint16 uint32 uint64 float double float16
192
196
  string binary json bson enum uuid date int96 time timestamp decimal fixed`, plus `struct`, `list`
193
- and `map`; `column :name, :int32` declares one by name. Fields are nullable unless `null: false`
197
+ and `map`; `s.column :name, :int32` declares one by name. Fields are nullable unless `null: false`
194
198
  is given; list elements unless `element_null: false`, map values unless `value_null: false`.
195
- Nested lists: `list :matrix do list :element, :double end`. `time` and `timestamp` take `unit:`
199
+ Nested lists: `s.list :matrix do |matrix| matrix.list :element, :double end`. `time` and `timestamp` take `unit:`
196
200
  (`:millis`, `:micros`, `:nanos`) and `utc:`.
197
201
 
198
202
  `enum` is a string column. `values:` restricts what may be written, and also takes a Rails-style
@@ -248,7 +252,7 @@ end
248
252
  ```
249
253
 
250
254
  Unlike CSV, headers are required. A column that holds more than one type can be declared up front:
251
- `Herringbone::SimpleWriter.new(io) { string :code }` (then call `close` when done).
255
+ `Herringbone::SimpleWriter.new(io) { |s| s.string :code }` (then call `close` when done).
252
256
 
253
257
  ### Statistics, page indexes and bloom filters
254
258
 
@@ -285,6 +289,136 @@ If the block raises, the SDK aborts the multipart upload and raises `Aws::S3::Mu
285
289
  so no partial object is left. S3 allows at most 10,000 parts, which with the default 5MB parts caps
286
290
  the file at about 48GB; raise `part_size:` for bigger files.
287
291
 
292
+ ## Redaction
293
+
294
+ `Herringbone.redact` rewrites a file with rows removed or values replaced, for GDPR "forget me"
295
+ requests and pseudonymization. One file goes in and one comes out, and the parts of the file
296
+ nothing touches are copied byte for byte.
297
+
298
+ ```ruby
299
+ # Forget me: remove the rows
300
+ Herringbone.redact(input, output) do |r|
301
+ r.where(user_id: 42).delete
302
+ end
303
+
304
+ # Forget me, but keep the row for accounting: blank the personal columns
305
+ Herringbone.redact(input, output) do |r|
306
+ r.where(user_id: 42).replace(email: nil, name: nil, address: nil)
307
+ end
308
+
309
+ # Forget me, but keep the row linkable: pseudonymize just that user's email
310
+ Herringbone.redact(input, output) do |r|
311
+ r.where(user_id: 42).replace(:email) { |email| OpenSSL::HMAC.hexdigest("SHA256", KEY, email) }
312
+ end
313
+
314
+ # Pseudonymize a column across the whole file, mask another, drop a third
315
+ Herringbone.redact(input, output) do |r|
316
+ r.replace(:email) { |email| OpenSSL::HMAC.hexdigest("SHA256", KEY, email.downcase) }
317
+ r.replace(:phone) { |phone| phone && "***#{phone[-3..]}" }
318
+ r.drop :ssn, :ip_address
319
+ end
320
+ ```
321
+
322
+ A `Herringbone::Redaction` is built once and applies itself to any number of files, which suits a
323
+ forget-me job going over a bucket:
324
+
325
+ ```ruby
326
+ forget = Herringbone::Redaction.new do |r|
327
+ r.where(user_id: 42).delete
328
+ r.where(email: "anna@example.com").delete # separate statements OR together
329
+ r.where(created_at: ..2.years.ago).delete # retention works the same way
330
+ end
331
+
332
+ if forget.affects?(input) # statistics and bloom filters first, then the where columns only
333
+ report = forget.apply(input, output)
334
+ report.rows_deleted # => 3
335
+ report.row_groups # => { copied: 61, rewritten: 1 }
336
+ end
337
+ ```
338
+
339
+ The block receives the redaction being built, like `Writer.open` passes the writer, so it can use
340
+ the methods and instance variables around it; a block without the `|r|` raises `ArgumentError`.
341
+ `where` takes what `read(where:)` takes (values, Arrays, Ranges, `nil`, callables and dotted
342
+ struct paths, but no columns inside lists or maps), and skips row groups and pages the same way.
343
+ `where(...).delete` removes the matching rows. `replace` sets constants,
344
+ `replace(email: nil, name: "[deleted]")`, or computes each value with a block,
345
+ `replace(:email, :phone) { |value| ... }`. A block that takes two parameters also gets the whole
346
+ row, as a Hash with String keys like `read` returns. Without `where`, `replace` applies to every
347
+ row. A column is a top-level field or a struct member by dotted path (`"address.city"`; a member
348
+ of a null struct is left alone). To change what is inside a list or map, replace the whole field
349
+ with a block that receives the Array or Hash. `drop :a, :b` removes columns from the schema.
350
+
351
+ Statements apply in declared order, row by row: a deleted row is gone for later statements, and
352
+ later statements see what earlier ones replaced, in their conditions as well as in their blocks.
353
+ Mistakes raise `ArgumentError` before anything is written: a `where` without a verb, a column that
354
+ doesn't exist, `nil` (or a value of the wrong type) for a `null: false` column. A block that
355
+ returns something its column can't store raises `Herringbone::EncodeError` when it gets there and
356
+ leaves the output unfinished.
357
+
358
+ `apply` and `Herringbone.redact` take IOs like `Reader` and `Writer` do: the input must be
359
+ seekable, the output only needs `#write`, and neither is closed. To redact in place, write to a
360
+ temporary file and rename it over the original; on S3, read the object and write the new one
361
+ through `upload_stream`. The `Redaction::Report` that `apply` returns has `rows_read`,
362
+ `rows_deleted`, `rows_changed` and `row_groups` (`{ copied:, rewritten: }`), which is the audit
363
+ trail an erasure needs.
364
+
365
+ Each row group is handled in the cheapest way that is still exact. If no statement can match it,
366
+ judging by statistics, bloom filters and the page index, or by reading only the `where` columns,
367
+ its column chunks are copied as they are and only their offsets are rebased. If rows match but
368
+ none is deleted and only leaf columns change, just those column chunks are encoded again. If rows
369
+ are deleted, or a nested field is replaced whole, the row group is rewritten, as one row group
370
+ (or none, when every row goes). Rewritten chunks keep the codec of the original chunk (LZO, which
371
+ Herringbone can't write, becomes Snappy) and get a new bloom filter if they had one. Writer
372
+ options (`compression:`, `bloom_filters:`, `page_rows:`...) apply to the rewritten chunks;
373
+ `row_group_bytes:` and `row_group_rows:` are refused, since row groups keep their boundaries. The
374
+ footer's key/value metadata is copied (without `ARROW:schema` and `pandas` when columns are
375
+ dropped, since those describe the columns) unless `metadata:` is given.
376
+
377
+ No deleted or replaced value survives in the output: not in data pages, dictionary pages, column
378
+ chunk min/max statistics, the page index or bloom filters. What Herringbone can't reach is up to
379
+ you: the original file, its S3 versions and backups still hold the data until you delete them,
380
+ only the columns you name are touched (an email that also sits in a free-text `notes` column stays
381
+ there), and other files with the same person in them need their own pass.
382
+
383
+ Some recipes. Keyed hashing keeps a column joinable without keeping the value; keep the key out
384
+ of the data, and rotate or destroy it to cut the link:
385
+
386
+ ```ruby
387
+ require "openssl"
388
+ KEY = ENV.fetch("PSEUDONYM_KEY")
389
+ Herringbone.redact(input, output) do |r|
390
+ r.replace(:email) { |email| email && OpenSSL::HMAC.hexdigest("SHA256", KEY, email.downcase) }
391
+ end
392
+ ```
393
+
394
+ Masking keeps enough to be recognizable to the person but not to anyone else:
395
+
396
+ ```ruby
397
+ Herringbone.redact(input, output) do |r|
398
+ r.replace(:card_number) { |number| number && number[-4..].rjust(number.size, "*") }
399
+ r.replace(:email) { |email| email&.sub(/\A(.).*@/, '\1***@') }
400
+ r.replace(:birth_date) { |date| date && Date.new(date.year, 1, 1) }
401
+ end
402
+ ```
403
+
404
+ Fake values (with the `faker` gem) make a copy for staging that looks real. Seed Faker from the
405
+ original value to get the same fake for the same person in every file:
406
+
407
+ ```ruby
408
+ require "faker"
409
+ require "zlib"
410
+ Herringbone.redact(input, output) do |r|
411
+ r.replace(:name) do |name|
412
+ next nil unless name
413
+ Faker::Config.random = Random.new(Zlib.crc32(name))
414
+ Faker::Name.name
415
+ end
416
+ r.replace(:address) do |address|
417
+ address && { "city" => Faker::Address.city, "zip" => Faker::Address.zip_code }
418
+ end
419
+ end
420
+ ```
421
+
288
422
  ## ActiveRecord
289
423
 
290
424
  `Herringbone.write` also takes a model or relation, which it reads with `find_each`, using a schema
@@ -92,7 +92,7 @@ module Herringbone
92
92
 
93
93
  if type == :hstore
94
94
  return builder.map(column.name, :string, :string, null: nullable) unless array
95
- return builder.list(column.name, null: nullable) { map :element, :string, :string }
95
+ return builder.list(column.name, null: nullable) { |list| list.map :element, :string, :string }
96
96
  end
97
97
 
98
98
  dsl_type, opts = scalar_type(column, type, sql_type, primary)
@@ -45,6 +45,15 @@ module Herringbone
45
45
  # @return [Boolean] true while still dictionary-encoding (not yet switched to raw bytes)
46
46
  def dictionary? = !@indices.nil?
47
47
 
48
+ # Short summary for the console, without the values
49
+ #
50
+ # @return [String] value count, the dictionary size (or "raw" after switching to raw bytes),
51
+ # the width for FIXED_LEN_BYTE_ARRAY and the estimated memory
52
+ def inspect
53
+ mode = dictionary? ? "dictionary=#{@dictionary.size}" : "raw"
54
+ "#<#{self.class.name} size=#{size} #{mode}#{" width=#{@width}" if @width} memory_bytes=#{memory_bytes}>"
55
+ end
56
+
48
57
  # Appends a value, switching to raw bytes when the dictionary limits are exceeded.
49
58
  # @param value [String] bytes of the value; for FIXED_LEN_BYTE_ARRAY it must be +width+ bytes
50
59
  # long (not checked here)
@@ -73,6 +73,13 @@ module Herringbone
73
73
  @writer&.abort
74
74
  end
75
75
 
76
+ # Short summary for the console, without the held-back rows
77
+ #
78
+ # @return [String] the number of rows held back, or the Writer's summary once it is open
79
+ def inspect
80
+ "#<#{self.class.name} #{@writer ? "writer=#{@writer.inspect}" : "held_back=#{@sample.size}"}>"
81
+ end
82
+
76
83
  private
77
84
 
78
85
  # Builds the schema from the held-back rows, opens the Writer and writes them