herringbone 0.5.0 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
data/MANUAL.md ADDED
@@ -0,0 +1,773 @@
1
+ # Herringbone manual
2
+
3
+ The full reference for Herringbone. For a quick start, see the [README](README.md).
4
+
5
+ - [Installation](#installation)
6
+ - [Optional gems](#optional-gems)
7
+ - [Ractors](#ractors)
8
+ - [Reading](#reading)
9
+ - [Reader options](#reader-options)
10
+ - [Selecting rows](#selecting-rows)
11
+ - [Numo arrays](#numo-arrays)
12
+ - [Writing](#writing)
13
+ - [Defining a schema](#defining-a-schema)
14
+ - [Writing with an inferred schema](#writing-with-an-inferred-schema)
15
+ - [Output IOs and failed writes](#output-ios-and-failed-writes)
16
+ - [Column types](#column-types)
17
+ - [Accepted values](#accepted-values)
18
+ - [Writer options](#writer-options)
19
+ - [SimpleWriter: coming from CSV](#simplewriter-coming-from-csv)
20
+ - [Statistics, page indexes and bloom filters](#statistics-page-indexes-and-bloom-filters)
21
+ - [Writing to S3](#writing-to-s3)
22
+ - [Encryption](#encryption)
23
+ - [Writing encrypted files](#writing-encrypted-files)
24
+ - [Which settings other tools read](#which-settings-other-tools-read)
25
+ - [Reading encrypted files](#reading-encrypted-files)
26
+ - [Key management](#key-management)
27
+ - [ActiveRecord](#activerecord)
28
+ - [Redaction](#redaction)
29
+ - [Statements](#statements)
30
+ - [Reusable redactions and reports](#reusable-redactions-and-reports)
31
+ - [Input and output](#input-and-output)
32
+ - [How row groups are rewritten](#how-row-groups-are-rewritten)
33
+ - [What gets erased, and what doesn't](#what-gets-erased-and-what-doesnt)
34
+ - [Encrypted files](#encrypted-files)
35
+ - [Recipes](#recipes)
36
+ - [Type mapping](#type-mapping)
37
+ - [Inspecting files](#inspecting-files)
38
+ - [Command line](#command-line)
39
+ - [Supported format features](#supported-format-features)
40
+ - [Development](#development)
41
+
42
+ ## Installation
43
+
44
+ ```ruby
45
+ gem "herringbone"
46
+ ```
47
+
48
+ ### Optional gems
49
+
50
+ Add these libraries to speed up and enable certain functionality:
51
+
52
+ ```ruby
53
+ gem "snappy" # native Snappy, 2-3x faster reads and writes of typical files
54
+ gem "zstd-ruby" # ZSTD
55
+ gem "brotli" # Brotli
56
+ gem "xxhash" # faster bloom filters
57
+ gem "numo-narray-alt" # read(as: :numo)
58
+ ```
59
+
60
+ Herringbone never requires them itself: it uses whichever are loaded. Bundler.require (as in Rails)
61
+ loads them, otherwise require them yourself:
62
+
63
+ ```ruby
64
+ require "herringbone"
65
+ require "zstd-ruby"
66
+ require "numo/narray"
67
+ ```
68
+
69
+ ### Ractors
70
+
71
+ On Ruby 3.1 and later, Herringbone works from any Ractor. Pass a schema to a Ractor with
72
+ `Ractor.make_shareable(schema)`. Rows made shareable the same way reach a writing Ractor by
73
+ reference, without being copied:
74
+
75
+ ```ruby
76
+ writer = Ractor.new("out.parquet", Ractor.make_shareable(schema)) do |path, schema|
77
+ File.open(path, "wb") do |f|
78
+ Herringbone::Writer.open(f, schema) do |w|
79
+ while (row = Ractor.receive) != :done
80
+ w << row
81
+ end
82
+ end
83
+ end
84
+ end
85
+ rows.each { |row| writer.send(Ractor.make_shareable(row)) }
86
+ writer.send(:done)
87
+ ```
88
+
89
+ Herringbone loads its parts on first use (`require "herringbone"` itself takes under a
90
+ millisecond). Before Ruby 3.4, Ractors other than the main one cannot load code, so call
91
+ `Herringbone.eager_load!` before starting them. It also moves all the loading to boot, for
92
+ example before a server forks.
93
+
94
+ ## Reading
95
+
96
+ Herringbone can read from any IO-ish object with random access (the IO should be seekable).
97
+
98
+ ```ruby
99
+ require "herringbone"
100
+
101
+ File.open("data.parquet", "rb") do |file|
102
+ reader = Herringbone::Reader.new(file)
103
+ reader.schema
104
+ reader.num_rows
105
+
106
+ reader.each_row { |row| p row } # Hashes with String keys, nested values as Hash/Array
107
+ reader.each_batch(1000) { |rows| ... } # Arrays of up to 1000 rows
108
+ reader.each_batch(10_000, as: :columns) { |batch| batch["amount"].sum } # { "id" => [...], ... }
109
+ reader.read # all rows at once; read(as: :columns) for column Arrays
110
+ end
111
+ ```
112
+
113
+ Reads stream: pages are read and decoded one at a time per column and rows are assembled in
114
+ batches, so memory depends on the batch and page sizes rather than on the row group size (1M
115
+ rows in a single row group peak at about 60–160 MB RSS growth instead of 660 MB). Column-order
116
+ batches (`as: :columns`) skip building a Hash per row and are 25–30% faster.
117
+
118
+ ### Reader options
119
+
120
+ `Reader.new` takes `keys: :symbol` for Symbol keys (in rows and in structs; map keys stay as
121
+ stored) and `time_zone:` to return timestamps in a zone instead of UTC: a UTC offset (`"+02:00"`,
122
+ or seconds), `"UTC"`, a `TZInfo::Timezone`, or anything responding to `#at` such as `Time.zone`
123
+ in Rails (which yields `ActiveSupport::TimeWithZone`). Zone names like `"Europe/Amsterdam"` work
124
+ when ActiveSupport or TZInfo is loaded. Timestamps stored with `isAdjustedToUTC=false` are
125
+ wall-clock values and stay as they are.
126
+
127
+ ### Selecting rows
128
+
129
+ `each_row`, `each_batch` and `read` take `columns:`, `where:`, `from:` and `limit:`:
130
+
131
+ ```ruby
132
+ reader.read(columns: %w[id name]) # projection
133
+ reader.read(where: { user_id: 42 })
134
+ reader.read(where: { status: %w[paid shipped], created_at: 1.week.ago.. }) # IN, Ranges
135
+ reader.read(where: { "address.city" => "Amsterdam", deleted_at: nil }) # struct members, IS NULL
136
+ reader.read(where: { amount: ->(v) { v && v > 100 } }) # any callable
137
+ reader.read(from: 1_000_000, limit: 100) # rows 1,000,000..1,000,099
138
+ ```
139
+
140
+ All conditions must hold; filtered columns need not be among `columns:`, and any column not inside
141
+ a list or map can be filtered on. Row groups are skipped using min/max statistics, null counts and
142
+ bloom filters, pages using the page index (written by Herringbone, parquet-mr, Arrow and others),
143
+ and `from:` jumps to its row through the offset index; every row read is then checked, so results
144
+ are exact. On a 1M-row file with 20k-row pages, looking up one `id` takes 0.1 s instead of 6.3 s.
145
+ Filters help most on columns the data is sorted or clustered by. `reader.scan_plan(where: ...)`
146
+ shows which row groups and row ranges a read would touch, without reading them.
147
+
148
+ ### Numo arrays
149
+
150
+ `as: :numo` returns (or yields, with `each_batch`) a Hash of column name => Numo array, and takes
151
+ the same `columns:`, `where:`, `from:` and `limit:`. Add `gem "numo-narray-alt"` (or
152
+ `numo-narray`) to your Gemfile and `require "numo/narray"`: without it, `as: :numo` raises
153
+ `Herringbone::UnsupportedError` naming the gem.
154
+
155
+ ```ruby
156
+ cols = reader.read(as: :numo, columns: %w[id amount]) # { "id" => Numo::Int64, "amount" => Numo::DFloat }
157
+ df = Rover::DataFrame.new(reader.read(as: :numo)) # no conversion needed
158
+ ```
159
+
160
+ Flat numeric and boolean columns are decoded from the page bytes without a Ruby object per value:
161
+ two numeric columns of 1M uncompressed rows read in about 25 ms, against 50 ms with `as: :columns`.
162
+
163
+ | Parquet | Numo |
164
+ |---|---|
165
+ | INT32, INT64 (also TIME) | `Int32`, `Int64` |
166
+ | INT(8/16) signed; INT(8/16/32/64) unsigned | `Int8`, `Int16`; `UInt8` … `UInt64` |
167
+ | FLOAT, FLOAT16, DOUBLE | `SFloat`, `SFloat`, `DFloat` (nulls are NaN) |
168
+ | integers with nulls | `DFloat` with NaN (exact up to 2**53) |
169
+ | BOOLEAN | `Bit`; with nulls `RObject` of true/false/nil |
170
+ | list of numbers, every row the same length and no nulls | 2-D `[rows, length]` (e.g. embeddings) |
171
+ | strings, binary, decimals, dates, timestamps, UUIDs, structs, maps, other lists | `RObject` of the values `read` returns |
172
+
173
+ Whether a column has nulls (or a list column is rectangular) is decided from the rows read, so
174
+ with `each_batch` it can differ between batches.
175
+
176
+ ## Writing
177
+
178
+ ### Defining a schema
179
+
180
+ ```ruby
181
+ schema = Herringbone::Schema.define do |s|
182
+ s.int64 :id, null: false
183
+ s.string :name
184
+ s.enum :status, values: %w[pending paid shipped]
185
+ s.list :tags, :string
186
+ s.map :scores, :string, :double
187
+ s.struct :address do |address|
188
+ address.string :city
189
+ address.string :zip
190
+ end
191
+ s.decimal :price, precision: 12, scale: 2
192
+ s.json :payload
193
+ s.timestamp :created_at
194
+ end
195
+
196
+ File.open("out.parquet", "wb") do |file|
197
+ Herringbone::Writer.open(file, schema) do |w|
198
+ w << { "id" => 1, "name" => "Anna", "status" => "paid", "tags" => ["a", "b"],
199
+ "scores" => { "x" => 1.5 }, "address" => { "city" => "Amsterdam" },
200
+ "price" => BigDecimal("9.99"), "payload" => { "any" => ["json"] }, "created_at" => Time.now }
201
+ w << { id: 2, status: :pending } # Symbol keys and values work; missing keys are nulls
202
+ w << [3, "Bo", nil, nil, nil, nil, nil, nil, nil, nil] # Arrays in schema order
203
+ w << order # anything with #attributes (ActiveRecord) or #to_h (Struct, Data)
204
+ end
205
+ end
206
+ ```
207
+
208
+ The block receives the schema builder, and the blocks of `struct`, `list` and `map` receive a
209
+ builder of their own, so the block can still call your methods and read your instance variables.
210
+ A block that takes no parameter raises `ArgumentError`.
211
+
212
+ ### Writing with an inferred schema
213
+
214
+ `Herringbone.write(io, rows)` writes an Enumerable of rows in one go, inferring the schema from the
215
+ first 1000 rows unless `schema:` is given; fields declared in a block replace inferred ones:
216
+ `Herringbone.write(io, rows) { |s| s.json :payload }`. The rows are iterated once, holding back only
217
+ those first 1000, so lazy Enumerators and cursors that can't be rewound work. A later row that
218
+ doesn't fit the inferred types raises `Herringbone::SchemaMismatch`, which explains what was
219
+ inferred and how to declare the column, and leaves the file unfinished.
220
+
221
+ ### Output IOs and failed writes
222
+
223
+ The writer writes to any IO that responds to `#write` (a `File`, `StringIO`, `Tempfile`, socket or
224
+ pipe), sequentially, and never seeks, rewinds or closes it (it does switch it to binary mode). If
225
+ the `Writer.open` block raises (or `#abort` is called), no footer is written and what was written
226
+ so far is left for you to discard. To replace a file only once it is complete, write to a temporary
227
+ file and rename it.
228
+
229
+ ### Column types
230
+
231
+ `boolean int8 int16 int32 int64 uint8 uint16 uint32 uint64 float double float16 string binary
232
+ json bson enum uuid date int96 time timestamp decimal fixed`, plus `struct`, `list` and `map`;
233
+ `s.column :name, :int32` declares one by name. Fields are nullable unless `null: false` is given;
234
+ list elements unless `element_null: false`, map values unless `value_null: false`. Nested lists:
235
+ `s.list :matrix do |matrix| matrix.list :element, :double end`. `time` and `timestamp` take `unit:`
236
+ (`:millis`, `:micros`, `:nanos`) and `utc:`.
237
+
238
+ `enum` is a string column. `values:` restricts what may be written, and also takes a Rails-style
239
+ Hash (`values: Order.statuses`), in which case both labels and stored integers are accepted and
240
+ the label is written. `parquet_enum: true` adds the Parquet ENUM annotation (pyarrow and pandas
241
+ read such columns as binary, which is why it is off by default).
242
+
243
+ ### Accepted values
244
+
245
+ Columns accept the values Ruby and Rails code usually has at hand:
246
+
247
+ | column | accepts |
248
+ |---|---|
249
+ | `date` | `Date`, `Time`/`DateTime` (their date), `"2024-05-01"` |
250
+ | `timestamp` | `Time`, `DateTime`, `ActiveSupport::TimeWithZone`, `Date` (midnight UTC), ISO-8601 strings, Integers in the column's unit |
251
+ | `time` | `Time` (its time of day, as Rails returns for `time` columns), `"13:45:30.25"`, Integers |
252
+ | `json` | Strings as-is, anything else through `JSON.generate` |
253
+ | `string`, `enum` | Strings, Symbols, anything with `to_s` |
254
+ | `boolean` | `true`/`false`, `1`/`0`, `"t"`/`"f"`, `"true"`/`"false"`, `"yes"`/`"no"` |
255
+ | integers | Integers, whole-number Floats/BigDecimals/Rationals, numeric Strings; out-of-range values raise |
256
+ | `decimal` | `BigDecimal`, Integer, Rational, Float, numeric Strings |
257
+ | `uuid` | Strings with or without dashes, or 16 raw bytes |
258
+
259
+ Values that don't fit raise `Herringbone::EncodeError` naming the row number and column path; the
260
+ failed row is discarded and the writer can carry on.
261
+
262
+ ### Writer options
263
+
264
+ | option | default | |
265
+ |---|---|---|
266
+ | `compression` | `:snappy` | `:none`, `:snappy`, `:gzip`, `:lz4` (LZ4_RAW), `:lz4_hadoop`, `:zstd`, `:brotli` |
267
+ | `compression_level` | codec default | level for `:zstd` (up to 22, default 3), `:gzip` (0-9) or `:brotli` (0-11) |
268
+ | `row_group_bytes` | 16MB | flush a row group once the buffered values take about this much memory, which bounds memory use (a 15-column table peaks around 290 MB RSS) |
269
+ | `row_group_rows` | none | also flush after this many rows |
270
+ | `page_bytes` | 1MB | approximate data page size |
271
+ | `page_rows` | `20_000` | at most this many rows per data page |
272
+ | `data_page_version` | `1` | `1` or `2` |
273
+ | `dictionary` | `true` | `false`, or an Array of column paths to dictionary-encode |
274
+ | `encodings` | `{}` | e.g. `{ "id" => :delta_binary_packed, "x" => :byte_stream_split }` |
275
+ | `metadata` | `{}` | footer key/value metadata, read back with `reader.metadata` |
276
+ | `bloom_filters` | none | `true`, an Array of column paths, or `{ "path" => { ndv:, fpp:, max_bytes: } }` |
277
+
278
+ ### SimpleWriter: coming from CSV
279
+
280
+ `SimpleWriter` writes like the CSV gem: name the columns, then append Arrays. The types are
281
+ inferred from the first 1000 rows, as with `Herringbone.write`.
282
+
283
+ ```ruby
284
+ File.open("people.parquet", "wb") do |file|
285
+ Herringbone::SimpleWriter.open(file) do |sw|
286
+ sw.headers!(:id, :name, :age)
287
+ sw << [123, "John", 12]
288
+ sw << { id: 124, name: "Jane" } # Hashes work too
289
+ end
290
+ end
291
+ ```
292
+
293
+ Unlike CSV, headers are required. A column that holds more than one type can be declared up front:
294
+ `Herringbone::SimpleWriter.new(io) { |s| s.string :code }` (then call `close` when done).
295
+
296
+ `encrypt!` encrypts the file with one key (see [Encryption](#encryption)): give it the key's hex
297
+ or a `Herringbone::Key` (`key: Herringbone::Key.generate` for a new one; keep it).
298
+
299
+ ```ruby
300
+ Herringbone::SimpleWriter.open(file) do |sw|
301
+ sw.encrypt!(key: ENV["PARQUET_KEY"])
302
+ sw.headers!(:id, :name)
303
+ sw << [1, "John"]
304
+ end
305
+ ```
306
+
307
+ ### Statistics, page indexes and bloom filters
308
+
309
+ Every column chunk gets min/max/null-count statistics and the Parquet page index (per-page min/max
310
+ and row offsets), which Herringbone, DuckDB, Spark, Trino, Arrow and DataFusion use to skip pages.
311
+ With 20,000 rows per page, `WHERE id BETWEEN ...` on a sorted column reads a handful of pages
312
+ instead of the whole row group, so sort rows by the columns you filter on.
313
+
314
+ Min/max cannot rule out a value inside a row group's range, the usual case for unsorted IDs,
315
+ emails or UUIDs; a bloom filter can. They are off by default. `bloom_filters: ["user_id", "email"]`
316
+ (or `true` for every non-boolean column) writes Parquet's split block bloom filters, sized from
317
+ each row group's distinct values at a false positive probability of 1% (`fpp:`, up to `max_bytes:`,
318
+ default 1MB) unless `ndv:` gives the number of distinct values. Nested leaves are named by their
319
+ dotted path (`"tags.list.element"`). Hashing is pure Ruby unless the `xxhash` gem is loaded:
320
+ a million rows take about 1.2 s longer to write with a filter on an INT64 column and 2.8 s longer
321
+ with one on a ~22-byte string column, and about 0.5 s longer with `xxhash`.
322
+
323
+ ### Writing to S3
324
+
325
+ Parquet keeps its metadata in a footer, so a file can be streamed into an S3 multipart upload
326
+ without a local copy, using `upload_stream` from `aws-sdk-s3`:
327
+
328
+ ```ruby
329
+ s3 = Aws::S3::TransferManager.new # aws-sdk-s3 1.197+; before that, Aws::S3::Object#upload_stream
330
+ s3.upload_stream(bucket: "exports", key: "events.parquet", part_size: 16 * 1024 * 1024) do |io|
331
+ Herringbone::Writer.open(io, schema) do |w|
332
+ events.each { |event| w << event }
333
+ end
334
+ end
335
+ s3.upload_stream(bucket: "exports", key: "orders.parquet") { |io| Herringbone.write(io, Order.all) }
336
+ ```
337
+
338
+ If the block raises, the SDK aborts the multipart upload and raises `Aws::S3::MultipartUploadError`,
339
+ so no partial object is left. S3 allows at most 10,000 parts, which with the default 5MB parts caps
340
+ the file at about 48GB; raise `part_size:` for bigger files.
341
+
342
+ ## Encryption
343
+
344
+ Herringbone reads and writes files encrypted with
345
+ [Parquet modular encryption](https://parquet.apache.org/docs/file-format/data-pages/encryption/),
346
+ the scheme parquet-mr (Spark), Arrow (pyarrow) and parquet-rs implement. Each column can be
347
+ encrypted with its own key or left in the clear, and the footer (schema, row counts, statistics)
348
+ is either encrypted too or left readable and signed. AES-GCM comes from OpenSSL, which Ruby ships
349
+ with; no other gem is needed.
350
+
351
+ ### Writing encrypted files
352
+
353
+ Most of the time, one key will do:
354
+
355
+ ```ruby
356
+ key = Herringbone::Key.generate # AES-256; keep it safe, without it the data is gone
357
+ key = Herringbone::Key.from_hex(ENV["PARQUET_KEY"]) # 32 or 64 hex digits; key.hex gives them back
358
+
359
+ Herringbone.write(io, rows, encryption: key)
360
+ Herringbone::Reader.new(io, decryption: key)
361
+ ```
362
+
363
+ That encrypts every column and the footer with the key (AES-GCM, no AAD prefix), the setup the
364
+ most other readers support, see below; they take the key itself (`key.bytes`). A key's hex (32 or
365
+ 64 digits) works in place of a `Key` too: `encryption: ENV["PARQUET_KEY"]`. Any other String is
366
+ refused, raw key bytes included, since those are easy to mangle and to mistake for text: wrap
367
+ them in `Herringbone::Key.new(bytes)`.
368
+
369
+ Each key has an id, stored in the clear in the files it encrypts: a fingerprint of the key (an
370
+ HMAC, which gives nothing away about the key), or one you choose with
371
+ `Herringbone::Key.new(bytes, id: "2026-10")`. A reader given several keys, a keyring, picks the
372
+ one the file names, so rotating keys is writing with a new one and reading with all of them:
373
+
374
+ ```ruby
375
+ Herringbone::Reader.new(io, decryption: [current_key, previous_key])
376
+ Herringbone::Reader.new(io, decryption: ->(key_id) { Rails.application.credentials.parquet_keys[key_id] })
377
+ ```
378
+
379
+ The block is your key management: it gets the id from the file and returns the key (a `Key` or
380
+ bytes), or nil. There is no envelope encryption (per-file data keys wrapped by a KMS), see
381
+ [Key management](#key-management). `Herringbone::EncryptionConfiguration.simple(key)` is what
382
+ `encryption: key` turns into.
383
+
384
+ For anything finer:
385
+
386
+ ```ruby
387
+ Herringbone::Writer.open(file, schema, encryption: {
388
+ footer_key: FOOTER_KEY, # 16, 24 or 32 bytes (AES-128/192/256)
389
+ footer_key_metadata: "orders-footer-v3", # stored in the file, to find the key by
390
+ columns: {
391
+ "ssn" => {key: SSN_KEY, key_metadata: "pii-v7"}, # its own key
392
+ "address" => ADDRESS_KEY, # a struct: all of its columns
393
+ "email" => :footer # the footer key
394
+ }
395
+ }) { |w| rows.each { |row| w << row } }
396
+
397
+ Herringbone.write(io, rows, encryption: {footer_key: FOOTER_KEY}) # every column, with the footer key
398
+ ```
399
+
400
+ The Hash is turned into a `Herringbone::EncryptionConfiguration`, which checks every setting
401
+ (keys, algorithm, column settings) when it is built; build one yourself to check settings up front
402
+ or to reuse them. It is frozen and its `#inspect` leaves the keys out. `decryption:` likewise takes
403
+ a Hash or a `Herringbone::DecryptionConfiguration`.
404
+
405
+ Columns not listed in `columns:` are written in the clear; without `columns:` every column is
406
+ encrypted with the footer key. Keys are binary Strings (`["00112233..."].pack("H*")` for hex).
407
+ Encrypted columns have their pages, page headers, statistics, page index and bloom filter
408
+ encrypted; statistics and the page index still drive `where:` for readers that have the key.
409
+
410
+ The other settings:
411
+
412
+ - `plaintext_footer: true` leaves the footer readable (and signs it), so readers without keys,
413
+ including those that don't know about encryption, can read the plaintext columns. The footer
414
+ then keeps no statistics of the encrypted columns.
415
+ - `algorithm: :aes_gcm_ctr` encrypts pages with AES-CTR instead of AES-GCM: a little faster, but
416
+ page contents are no longer authenticated (headers and metadata still are).
417
+ - `aad_prefix: "orders/2026-10-03/part-0"` binds the file to an identity, so it can't be passed
418
+ off as another file encrypted with the same keys. It is stored in the file, unless
419
+ `store_aad_prefix: false`, in which case readers have to supply it.
420
+
421
+ ### Which settings other tools read
422
+
423
+ What each Parquet implementation can decrypt, as of late 2026. Herringbone's own interop tests
424
+ cover pyarrow; the rest comes from their documentation and source code.
425
+
426
+ | Reader | Keys | Limits |
427
+ |---|---|---|
428
+ | pyarrow 25+ | one key: `pq.read_table(path, decryption_properties=pyarrow.parquet.encryption.create_decryption_properties(key))` | single-key API: uniform encryption only; anything else needs `CryptoFactory` and a KMS client reading key tools JSON (see Key management) |
429
+ | Arrow C++, arrow-go, ParquetSharp | footer key, column keys, key retriever | none |
430
+ | arrow-rs, DataFusion | footer key (`format.crypto.file_decryption.footer_key_as_hex`), column keys, key retriever | no AES-CTR, no 192-bit keys; the `datafusion` Python wheel is built without encryption |
431
+ | parquet-java, Spark 3.2+, Hive | key tools JSON through a KMS client class; plain keys need a small `DecryptionPropertiesFactory` returning the footer key | none |
432
+ | Trino 478+ (Hive connector) | footer and column keys from environment variables | read only |
433
+ | DuckDB 1.5 | one key (`PRAGMA add_parquet_key`) | encrypted footer, uniform, AES-GCM, no AAD prefix; 1.5.6 fails on column chunks with more than one page or with a bloom filter, also for files pyarrow writes |
434
+ | Polars, ClickHouse, Parquet.Net | none | can't read encrypted files |
435
+
436
+ `EncryptionConfiguration.simple` stays inside all those limits except DuckDB's page bug. Files
437
+ written by DuckDB don't follow the spec (no AAD, no column crypto metadata), and Herringbone can't
438
+ read them.
439
+
440
+ ### Reading encrypted files
441
+
442
+ ```ruby
443
+ reader = Herringbone::Reader.new(file, decryption: {
444
+ footer_key: FOOTER_KEY,
445
+ columns: {"ssn" => SSN_KEY, "address" => ADDRESS_KEY}
446
+ })
447
+ reader.read(where: {ssn: "123-45-6789"})
448
+ ```
449
+
450
+ Instead of (or besides) giving keys, `keys:` looks them up by the key metadata stored in the file:
451
+ a Hash, or anything responding to `#call`, returning nil for keys the caller has no access to.
452
+ Each key metadata is looked up once per Reader. A callable that takes two parameters also gets
453
+ what the key is for (`:footer` or the column's dotted path), and is called for keys the file
454
+ stores no metadata for, with nil.
455
+
456
+ ```ruby
457
+ Herringbone::Reader.new(file, decryption: {keys: ->(key_metadata) { kms.data_key(key_metadata) }})
458
+ ```
459
+
460
+ A file with a plaintext footer opens without keys; its plaintext columns read as usual, and the
461
+ footer signature is checked whenever the footer key is available. Reading or filtering on an
462
+ encrypted column whose key is missing raises `Herringbone::DecryptionError`, naming the column and
463
+ its key metadata, so `columns:` can leave it out. A wrong key, an AAD prefix that doesn't match,
464
+ or changed bytes also raise `DecryptionError`. `reader.encryption` describes how a file is
465
+ encrypted (algorithm, footer mode, key metadata, which columns are readable), nil when it isn't.
466
+
467
+ Files from parquet-mr 1.12+ and Arrow (both algorithms, both footer modes, AAD prefixes stored or
468
+ supplied) are read, and pyarrow reads the files Herringbone writes.
469
+
470
+ ### Key management
471
+
472
+ The file only stores the key metadata you give it; turning that back into a key is up to you.
473
+ Common schemes: the metadata names a key in a KMS or secret store; or it holds a data key
474
+ encrypted ("wrapped") with a master key, which the KMS unwraps. Herringbone doesn't implement
475
+ parquet-mr's and pyarrow's key tools format (JSON key material with wrapped data keys), but a
476
+ `keys:` resolver can read it: `JSON.parse(key_metadata)` gives the master key id and the wrapped
477
+ data key for your KMS to unwrap. Keys are per file and per column, not per row, so encryption
478
+ does not replace deleting a person's rows (see [Redaction](#redaction)). AES-GCM allows about 4
479
+ billion encryptions per key, two per page; the writer raises before going over.
480
+
481
+ ## ActiveRecord
482
+
483
+ `Herringbone.write` also takes a model or relation, which it reads with `find_each`, using a schema
484
+ built from the model's columns. It returns the number of rows written:
485
+
486
+ ```ruby
487
+ File.open("orders.parquet", "wb") do |file|
488
+ Herringbone.write(file, Order.where(created_at: 1.year.ago..), compression: :zstd)
489
+ end
490
+ schema = Herringbone::Schema.from_active_record(Order, only: %w[id status total])
491
+ Herringbone.write(io, Order, schema: schema)
492
+ ```
493
+
494
+ `from_active_record` takes `only:`, `except:` and `parquet_enum: true`. Rails is not a dependency:
495
+ it only calls `columns`, `primary_key` and `defined_enums` on the model. Enum attributes are
496
+ written as their labels, and the writer rejects values outside the enum.
497
+
498
+ | Column | Parquet |
499
+ |---|---|
500
+ | `integer` | `int16`/`int32`/`int64` by SQL type (`smallint`, `integer`, `bigint`...) or limit; `uintN` for `unsigned`; primary keys are always `int64` |
501
+ | `float` | `double` (`float` for Postgres `float4`) |
502
+ | `decimal` | `decimal(precision, scale)`; `decimal(38, 9)` when the column has no precision |
503
+ | `boolean`, `date`, `binary`, `uuid` | same |
504
+ | `json`, `jsonb` | `json` |
505
+ | `datetime`, `timestamp`, `timestamptz` | `timestamp` (microseconds, UTC) |
506
+ | `time` | `time` (microseconds) |
507
+ | `hstore` | `map` of `string` to `string` |
508
+ | enum attributes | `string` (or `enum`) |
509
+ | `string`, `text`, `citext`, anything else | `string` |
510
+ | Postgres array columns | `list` of the element type |
511
+
512
+ Columns declared `NOT NULL` (and primary keys) are required, all others nullable. Column order
513
+ follows `Model.columns`.
514
+
515
+ ## Redaction
516
+
517
+ `Herringbone.redact` rewrites a file with rows removed or values replaced, for GDPR "forget me"
518
+ requests and pseudonymization. One file goes in and one comes out, and the parts of the file
519
+ nothing touches are copied byte for byte.
520
+
521
+ ```ruby
522
+ # Forget me: remove the rows
523
+ Herringbone.redact(input, output) do |r|
524
+ r.where(user_id: 42).delete
525
+ end
526
+
527
+ # Forget me, but keep the row for accounting: blank the personal columns
528
+ Herringbone.redact(input, output) do |r|
529
+ r.where(user_id: 42).replace(email: nil, name: nil, address: nil)
530
+ end
531
+
532
+ # Forget me, but keep the row linkable: pseudonymize just that user's email
533
+ Herringbone.redact(input, output) do |r|
534
+ r.where(user_id: 42).replace(:email) { |email| OpenSSL::HMAC.hexdigest("SHA256", KEY, email) }
535
+ end
536
+
537
+ # Pseudonymize a column across the whole file, mask another, drop a third
538
+ Herringbone.redact(input, output) do |r|
539
+ r.replace(:email) { |email| OpenSSL::HMAC.hexdigest("SHA256", KEY, email.downcase) }
540
+ r.replace(:phone) { |phone| phone && "***#{phone[-3..]}" }
541
+ r.drop :ssn, :ip_address
542
+ end
543
+ ```
544
+
545
+ ### Statements
546
+
547
+ The block receives the redaction being built, like `Writer.open` passes the writer, so it can use
548
+ the methods and instance variables around it; a block without the `|r|` raises `ArgumentError`.
549
+ `where` takes what `read(where:)` takes (values, Arrays, Ranges, `nil`, callables and dotted
550
+ struct paths, but no columns inside lists or maps), and skips row groups and pages the same way.
551
+ `where(...).delete` removes the matching rows. `replace` sets constants,
552
+ `replace(email: nil, name: "[deleted]")`, or computes each value with a block,
553
+ `replace(:email, :phone) { |value| ... }`. A block that takes two parameters also gets the whole
554
+ row, as a Hash with String keys like `read` returns. Without `where`, `replace` applies to every
555
+ row. A column is a top-level field or a struct member by dotted path (`"address.city"`; a member
556
+ of a null struct is left alone). To change what is inside a list or map, replace the whole field
557
+ with a block that receives the Array or Hash. `drop :a, :b` removes columns from the schema.
558
+
559
+ Statements apply in declared order, row by row: a deleted row is gone for later statements, and
560
+ later statements see what earlier ones replaced, in their conditions as well as in their blocks.
561
+ Mistakes raise `ArgumentError` before anything is written: a `where` without a verb, a column that
562
+ doesn't exist, `nil` (or a value of the wrong type) for a `null: false` column. A block that
563
+ returns something its column can't store raises `Herringbone::EncodeError` when it gets there and
564
+ leaves the output unfinished.
565
+
566
+ ### Reusable redactions and reports
567
+
568
+ A `Herringbone::Redaction` is built once and applies itself to any number of files, which suits a
569
+ forget-me job going over a bucket:
570
+
571
+ ```ruby
572
+ forget = Herringbone::Redaction.new do |r|
573
+ r.where(user_id: 42).delete
574
+ r.where(email: "anna@example.com").delete # separate statements OR together
575
+ r.where(created_at: ..2.years.ago).delete # retention works the same way
576
+ end
577
+
578
+ if forget.affects?(input) # statistics and bloom filters first, then the where columns only
579
+ report = forget.apply(input, output)
580
+ report.rows_deleted # => 3
581
+ report.row_groups # => { copied: 61, rewritten: 1 }
582
+ end
583
+ ```
584
+
585
+ The `Redaction::Report` that `apply` returns has `rows_read`, `rows_deleted`, `rows_changed` and
586
+ `row_groups` (`{ copied:, rewritten: }`), which is the audit trail an erasure needs.
587
+
588
+ ### Input and output
589
+
590
+ `apply` and `Herringbone.redact` take IOs like `Reader` and `Writer` do: the input must be
591
+ seekable, the output only needs `#write`, and neither is closed. To redact in place, write to a
592
+ temporary file and rename it over the original; on S3, read the object and write the new one
593
+ through `upload_stream`.
594
+
595
+ ### How row groups are rewritten
596
+
597
+ Each row group is handled in the cheapest way that is still exact. If no statement can match it,
598
+ judging by statistics, bloom filters and the page index, or by reading only the `where` columns,
599
+ its column chunks are copied as they are and only their offsets are rebased. If rows match but
600
+ none is deleted and only leaf columns change, just those column chunks are encoded again. If rows
601
+ are deleted, or a nested field is replaced whole, the row group is rewritten, as one row group
602
+ (or none, when every row goes). Rewritten chunks keep the codec of the original chunk (LZO, which
603
+ Herringbone can't write, becomes Snappy) and get a new bloom filter if they had one. Writer
604
+ options (`compression:`, `bloom_filters:`, `page_rows:`...) apply to the rewritten chunks;
605
+ `row_group_bytes:` and `row_group_rows:` are refused, since row groups keep their boundaries. The
606
+ footer's key/value metadata is copied (without `ARROW:schema` and `pandas` when columns are
607
+ dropped, since those describe the columns) unless `metadata:` is given.
608
+
609
+ ### What gets erased, and what doesn't
610
+
611
+ No deleted or replaced value survives in the output: not in data pages, dictionary pages, column
612
+ chunk min/max statistics, the page index or bloom filters. What Herringbone can't reach is up to
613
+ you: the original file, its S3 versions and backups still hold the data until you delete them,
614
+ only the columns you name are touched (an email that also sits in a free-text `notes` column stays
615
+ there), and other files with the same person in them need their own pass.
616
+
617
+ ### Encrypted files
618
+
619
+ Pass the input's keys in `decryption:`. The output is then encrypted the same way: same algorithm,
620
+ footer mode, AAD prefix, keys and key metadata, for the columns that are kept. `encryption:` (as
621
+ for `Writer`) writes it with other settings, and `encryption: false` writes a plaintext file.
622
+ `Redaction#affects?` takes `decryption:` too.
623
+
624
+ ```ruby
625
+ Herringbone.redact(input, output, decryption: {keys: kms_lookup}) do |r|
626
+ r.where(user_id: 42).delete
627
+ end
628
+ ```
629
+
630
+ Encrypted column chunks can't be copied byte for byte, since their encryption is tied to the file
631
+ and to their position in it, so they are encoded again even in row groups no statement touches.
632
+ Every key of the columns being kept must be available.
633
+
634
+ ### Recipes
635
+
636
+ Keyed hashing keeps a column joinable without keeping the value; keep the key out of the data, and
637
+ rotate or destroy it to cut the link:
638
+
639
+ ```ruby
640
+ require "openssl"
641
+ KEY = ENV.fetch("PSEUDONYM_KEY")
642
+ Herringbone.redact(input, output) do |r|
643
+ r.replace(:email) { |email| email && OpenSSL::HMAC.hexdigest("SHA256", KEY, email.downcase) }
644
+ end
645
+ ```
646
+
647
+ Masking keeps enough to be recognizable to the person but not to anyone else:
648
+
649
+ ```ruby
650
+ Herringbone.redact(input, output) do |r|
651
+ r.replace(:card_number) { |number| number && number[-4..].rjust(number.size, "*") }
652
+ r.replace(:email) { |email| email&.sub(/\A(.).*@/, '\1***@') }
653
+ r.replace(:birth_date) { |date| date && Date.new(date.year, 1, 1) }
654
+ end
655
+ ```
656
+
657
+ Fake values (with the `faker` gem) make a copy for staging that looks real. Seed Faker from the
658
+ original value to get the same fake for the same person in every file:
659
+
660
+ ```ruby
661
+ require "faker"
662
+ require "zlib"
663
+ Herringbone.redact(input, output) do |r|
664
+ r.replace(:name) do |name|
665
+ next nil unless name
666
+ Faker::Config.random = Random.new(Zlib.crc32(name))
667
+ Faker::Name.name
668
+ end
669
+ r.replace(:address) do |address|
670
+ address && { "city" => Faker::Address.city, "zip" => Faker::Address.zip_code }
671
+ end
672
+ end
673
+ ```
674
+
675
+ ## Type mapping
676
+
677
+ | Parquet | Ruby |
678
+ |---|---|
679
+ | BOOLEAN | `true`/`false` |
680
+ | INT32/INT64 (incl. signed/unsigned INTEGER) | `Integer` |
681
+ | FLOAT, DOUBLE, FLOAT16 | `Float` |
682
+ | STRING, ENUM, JSON | `String` (UTF-8) |
683
+ | BYTE_ARRAY, FIXED_LEN_BYTE_ARRAY, BSON | `String` (binary) |
684
+ | DATE | `Date` (proleptic Gregorian) |
685
+ | TIMESTAMP, INT96 | `Time` (UTC) |
686
+ | TIME | `Integer` in the column's unit since midnight |
687
+ | DECIMAL | `BigDecimal` |
688
+ | UUID | `String` like `"0f1e2d3c-..."` |
689
+ | struct / list / map | `Hash` / `Array` / `Hash` |
690
+
691
+ ## Inspecting files
692
+
693
+ > The inspector and its HTML view are modelled on
694
+ > **[Parquet X-ray](https://huggingface.co/spaces/cfahlgren1/parquet-xray) by cfahlgren1** —
695
+ > the design and the idea are theirs. Go check it out. It is amazing!
696
+
697
+ `Herringbone::Inspector` examines a file using only its footer, page headers, page indexes and
698
+ bloom filter headers. Nothing is decompressed, so it is fast on big files and works for ZSTD and
699
+ Brotli files without the codec gems.
700
+
701
+ ```ruby
702
+ File.open("data.parquet", "rb") do |file|
703
+ inspector = Herringbone::Inspector.new(file)
704
+ inspector.summary # size, rows, row groups, codecs, page index / bloom filter presence...
705
+ inspector.row_groups[0].column("name").pages # page headers; also statistics, column_index...
706
+ puts inspector.report # the schema, every column chunk, key/value metadata (Arrow schema decoded)
707
+ inspector.to_h # all of it, JSON-serializable
708
+ inspector.to_html # one self-contained HTML page with a to-scale byte map of the file
709
+ end
710
+ ```
711
+
712
+ An encrypted file takes `decryption:` as for `Reader`. With a plaintext footer it opens without
713
+ keys, showing the encrypted chunks without their pages and page indexes; `summary` and `report`
714
+ say how the file is encrypted. Encrypted pages are checked against their CRCs as stored, without
715
+ decrypting them.
716
+
717
+ `inspector.verify_checksums` reads every page body (still without decompressing) to check the page
718
+ CRCs; the results then appear in `summary`, `report`, `to_h` and `to_html`.
719
+
720
+ ## Command line
721
+
722
+ ```
723
+ bin/herringbone cat FILE [N] # rows as JSON lines
724
+ bin/herringbone inspect FILE [--pages] # text report (--pages lists every page header)
725
+ bin/herringbone inspect FILE --format=json # everything as JSON
726
+ bin/herringbone inspect FILE --format=html # the HTML page, opened in your browser
727
+ bin/herringbone inspect FILE --format=html > out.html # the HTML page, saved
728
+ ```
729
+
730
+ Add `--verify-checksums` to any `inspect` form to check page CRCs.
731
+
732
+ Encrypted files take keys on the command line: as they are (picked by their fingerprint id, like
733
+ a keyring), or for the footer, a column path or the key id the file stores (which `inspect`
734
+ shows):
735
+
736
+ ```
737
+ bin/herringbone cat FILE --key=$PARQUET_KEY
738
+ bin/herringbone inspect FILE --footer-key=00112233445566778899aabbccddeeff
739
+ bin/herringbone cat FILE 10 --key=kc1=base64:MTIzNDU2Nzg5MDEyMzQ1MA== --column-key=ssn=raw:1234567890123450
740
+ bin/herringbone inspect FILE --aad-prefix=orders/part-0 --no-prompt
741
+ ```
742
+
743
+ Keys are hex, `base64:...` or `raw:...` (a 16, 24 or 32-character key that isn't valid hex is
744
+ taken as typed too). A key the file needs but wasn't given is asked for on stdin, without echo in
745
+ a terminal; an empty answer leaves that column unreadable. `--no-prompt` asks for nothing.
746
+
747
+ `--format` is `text` (the default), `json` or `html`; `--pages` only goes with `text`.
748
+ In a terminal, `--format=html` writes the page to a temp file, prints its path and opens it with
749
+ `open` on macOS, `start` on Windows or `xdg-open` elsewhere. When stdout is redirected or piped it
750
+ prints the page instead.
751
+
752
+ ## Supported format features
753
+
754
+ - Encodings (read and write): PLAIN, PLAIN_DICTIONARY/RLE_DICTIONARY, RLE, DELTA_BINARY_PACKED,
755
+ DELTA_LENGTH_BYTE_ARRAY, DELTA_BYTE_ARRAY, BYTE_STREAM_SPLIT; legacy BIT_PACKED levels (read)
756
+ - Data page v1 and v2, dictionary pages, page CRCs (written)
757
+ - Page indexes and split block bloom filters (read and written)
758
+ - Legacy list and map layouts per the Parquet backward-compatibility rules
759
+ - LZO-compressed files (read only)
760
+ - Modular encryption (read and write): AES_GCM_V1 and AES_GCM_CTR_V1, encrypted and plaintext
761
+ footers, footer and column keys, AAD prefixes
762
+ - Not supported: column chunks in external files
763
+
764
+ ## Development
765
+
766
+ ```
767
+ bundle install
768
+ bundle exec rake test
769
+ HERRINGBONE_PYTHON=/path/to/python-with-pyarrow bundle exec rake test # also run pyarrow interop tests
770
+ ```
771
+
772
+ `test/fixtures/parquet-testing` holds files from [apache/parquet-testing](https://github.com/apache/parquet-testing)
773
+ (Apache-2.0); expectations for them were generated with pyarrow, see `test/fixtures/generate_expectations.py`.