herringbone 0.5.0 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
data/README.md CHANGED
@@ -7,20 +7,13 @@ A pure-Ruby reader and writer for [Apache Parquet](https://parquet.apache.org/)
7
7
  LZO, from Hadoop-era files, can be read but not written
8
8
  - Full nesting support (structs, lists, maps, any depth) via Dremel record shredding/assembly
9
9
  - Reads files from parquet-mr, Arrow, Spark, Impala, DuckDB, Rust writers etc.
10
- - Optimized reads with batches and pages
10
+ - Optimized reads with batches and pages
11
+ - Parquet modular encryption (per column and of the footer), interoperable with parquet-mr and Arrow
11
12
  - Ruby 3.0+
12
13
 
13
- ## Diving in: dumping records in Rails
14
-
15
- Stream a relation straight into S3, no temp file needed:
16
-
17
- ```ruby
18
- s3 = Aws::S3::TransferManager.new
19
- s3.upload_stream(bucket: "exports", key: "payments.parquet") do |io|
20
- # Schema will be auto-inferred, find_each will be used automatically
21
- Herringbone.write(io, Payment.where(status: "settled", created_at: 1.month.ago..))
22
- end
23
- ```
14
+ This README covers the common cases. The [MANUAL](MANUAL.md) has everything else: reader and
15
+ writer options, schemas and column types, bloom filters, Numo arrays, encryption, the redaction
16
+ rules, type mapping and the inspector.
24
17
 
25
18
  ## Installation
26
19
 
@@ -28,272 +21,110 @@ end
28
21
  gem "herringbone"
29
22
  ```
30
23
 
31
- Add these libraries to speed up and enable certain functionality:
24
+ Optionally add `snappy` (faster), `zstd-ruby` and `brotli` (more codecs), `xxhash` (faster bloom
25
+ filters) or `numo-narray-alt` (`read(as: :numo)`). Herringbone uses whichever of them are loaded,
26
+ see [Optional gems](MANUAL.md#optional-gems).
32
27
 
33
- ```ruby
34
- gem "snappy" # native Snappy, 2-3x faster reads and writes of typical files
35
- gem "zstd-ruby" # ZSTD
36
- gem "brotli" # Brotli
37
- gem "xxhash" # faster bloom filters
38
- gem "numo-narray-alt" # read(as: :numo)
39
- ```
28
+ ## Exporting from Rails
40
29
 
41
- Herringbone never requires them itself: it uses whichever are loaded. Bundler.require (as in Rails)
42
- loads them, otherwise require them yourself:
30
+ Stream a relation straight into S3, no temp file needed:
43
31
 
44
32
  ```ruby
45
- require "herringbone"
46
- require "zstd-ruby"
47
- require "numo/narray"
33
+ s3 = Aws::S3::TransferManager.new
34
+ s3.upload_stream(bucket: "exports", key: "payments.parquet") do |io|
35
+ # Schema will be auto-inferred, find_each will be used automatically
36
+ Herringbone.write(io, Payment.where(status: "settled", created_at: 1.month.ago..))
37
+ end
48
38
  ```
49
39
 
50
- On Ruby 3.1 and later, Herringbone works from any Ractor. Pass a schema to a Ractor with
51
- `Ractor.make_shareable(schema)`. Rows made shareable the same way reach a writing Ractor by
52
- reference, without being copied:
40
+ Or to a local file:
53
41
 
54
42
  ```ruby
55
- writer = Ractor.new("out.parquet", Ractor.make_shareable(schema)) do |path, schema|
56
- File.open(path, "wb") do |f|
57
- Herringbone::Writer.open(f, schema) do |w|
58
- while (row = Ractor.receive) != :done
59
- w << row
60
- end
61
- end
62
- end
43
+ File.open("orders.parquet", "wb") do |file|
44
+ Herringbone.write(file, Order.where(created_at: 1.year.ago..), compression: :zstd)
63
45
  end
64
- rows.each { |row| writer.send(Ractor.make_shareable(row)) }
65
- writer.send(:done)
66
46
  ```
67
47
 
68
- ## Reading
48
+ Any Enumerable of Hashes, Structs or Arrays works too; the schema is inferred from the first 1000
49
+ rows. See [ActiveRecord](MANUAL.md#activerecord) and [Writing to S3](MANUAL.md#writing-to-s3).
69
50
 
70
- Herringbone can read from any IO-ish object with random access (the IO should be seekable).
51
+ ## Reading
71
52
 
72
53
  ```ruby
73
- require "herringbone"
74
-
75
54
  File.open("data.parquet", "rb") do |file|
76
55
  reader = Herringbone::Reader.new(file)
77
- reader.schema
78
56
  reader.num_rows
79
57
 
80
- reader.each_row { |row| p row } # Hashes with String keys, nested values as Hash/Array
81
- reader.each_batch(1000) { |rows| ... } # Arrays of up to 1000 rows
82
- reader.each_batch(10_000, as: :columns) { |batch| batch["amount"].sum } # { "id" => [...], ... }
83
- reader.read # all rows at once; read(as: :columns) for column Arrays
84
- end
85
- ```
86
-
87
- Reads stream: pages are read and decoded one at a time per column and rows are assembled in
88
- batches, so memory depends on the batch and page sizes rather than on the row group size (1M
89
- rows in a single row group peak at about 60–160 MB RSS growth instead of 660 MB). Column-order
90
- batches (`as: :columns`) skip building a Hash per row and are 25–30% faster.
58
+ reader.each_row { |row| p row } # Hashes with String keys, nested values as Hash/Array
59
+ reader.each_batch(1000) { |rows| ... } # Arrays of up to 1000 rows
60
+ reader.read # all rows at once
91
61
 
92
- `Reader.new` takes `keys: :symbol` for Symbol keys (in rows and in structs; map keys stay as
93
- stored) and `time_zone:` to return timestamps in a zone instead of UTC: a UTC offset (`"+02:00"`,
94
- or seconds), `"UTC"`, a `TZInfo::Timezone`, or anything responding to `#at` such as `Time.zone`
95
- in Rails (which yields `ActiveSupport::TimeWithZone`). Zone names like `"Europe/Amsterdam"` work
96
- when ActiveSupport or TZInfo is loaded. Timestamps stored with `isAdjustedToUTC=false` are
97
- wall-clock values and stay as they are.
98
-
99
- ### Selecting rows
100
-
101
- `each_row`, `each_batch` and `read` take `columns:`, `where:`, `from:` and `limit:`:
102
-
103
- ```ruby
104
- reader.read(columns: %w[id name]) # projection
105
- reader.read(where: { user_id: 42 })
106
- reader.read(where: { status: %w[paid shipped], created_at: 1.week.ago.. }) # IN, Ranges
107
- reader.read(where: { "address.city" => "Amsterdam", deleted_at: nil }) # struct members, IS NULL
108
- reader.read(where: { amount: ->(v) { v && v > 100 } }) # any callable
109
- reader.read(from: 1_000_000, limit: 100) # rows 1,000,000..1,000,099
62
+ reader.read(columns: %w[id name], where: { user_id: 42 })
63
+ reader.read(where: { status: %w[paid shipped], created_at: 1.week.ago.. })
64
+ end
110
65
  ```
111
66
 
112
- All conditions must hold; filtered columns need not be among `columns:`, and any column not inside
113
- a list or map can be filtered on. Row groups are skipped using min/max statistics, null counts and
114
- bloom filters, pages using the page index (written by Herringbone, parquet-mr, Arrow and others),
115
- and `from:` jumps to its row through the offset index; every row read is then checked, so results
116
- are exact. On a 1M-row file with 20k-row pages, looking up one `id` takes 0.1 s instead of 6.3 s.
117
- Filters help most on columns the data is sorted or clustered by. `reader.scan_plan(where: ...)`
118
- shows which row groups and row ranges a read would touch, without reading them.
67
+ Reads stream page by page, and `where:` skips row groups and pages it can rule out using
68
+ statistics, page indexes and bloom filters. See [Reading](MANUAL.md#reading).
119
69
 
120
- ### Numo arrays
70
+ ## Writing CSV-style
121
71
 
122
- `as: :numo` returns (or yields, with `each_batch`) a Hash of column name => Numo array, and takes
123
- the same `columns:`, `where:`, `from:` and `limit:`. Add `gem "numo-narray-alt"` (or
124
- `numo-narray`) to your Gemfile and `require "numo/narray"`: without it, `as: :numo` raises
125
- `Herringbone::UnsupportedError` naming the gem.
72
+ `SimpleWriter` writes like the CSV gem: name the columns, then append rows. The types are inferred.
126
73
 
127
74
  ```ruby
128
- cols = reader.read(as: :numo, columns: %w[id amount]) # { "id" => Numo::Int64, "amount" => Numo::DFloat }
129
- df = Rover::DataFrame.new(reader.read(as: :numo)) # no conversion needed
75
+ File.open("people.parquet", "wb") do |file|
76
+ Herringbone::SimpleWriter.open(file) do |sw|
77
+ sw.headers!(:id, :name, :age)
78
+ sw << [123, "John", 12]
79
+ sw << { id: 124, name: "Jane" } # Hashes work too
80
+ end
81
+ end
130
82
  ```
131
83
 
132
- Flat numeric and boolean columns are decoded from the page bytes without a Ruby object per value:
133
- two numeric columns of 1M uncompressed rows read in about 25 ms, against 50 ms with `as: :columns`.
134
-
135
- | Parquet | Numo |
136
- |---|---|
137
- | INT32, INT64 (also TIME) | `Int32`, `Int64` |
138
- | INT(8/16) signed; INT(8/16/32/64) unsigned | `Int8`, `Int16`; `UInt8` … `UInt64` |
139
- | FLOAT, FLOAT16, DOUBLE | `SFloat`, `SFloat`, `DFloat` (nulls are NaN) |
140
- | integers with nulls | `DFloat` with NaN (exact up to 2**53) |
141
- | BOOLEAN | `Bit`; with nulls `RObject` of true/false/nil |
142
- | list of numbers, every row the same length and no nulls | 2-D `[rows, length]` (e.g. embeddings) |
143
- | strings, binary, decimals, dates, timestamps, UUIDs, structs, maps, other lists | `RObject` of the values `read` returns |
84
+ To control the types, declare a schema and use `Herringbone::Writer`, see
85
+ [Writing](MANUAL.md#writing).
144
86
 
145
- Whether a column has nulls (or a list column is rectangular) is decided from the rows read, so
146
- with `each_batch` it can differ between batches.
87
+ ## Encrypting
147
88
 
148
- ## Writing
89
+ Parquet files tend to wander off: to S3 buckets, laptops, other teams. Encrypting them takes one
90
+ key and one line, no KMS, no Hadoop configuration. Make a key once and keep it with your other
91
+ secrets:
149
92
 
150
93
  ```ruby
151
- schema = Herringbone::Schema.define do |s|
152
- s.int64 :id, null: false
153
- s.string :name
154
- s.enum :status, values: %w[pending paid shipped]
155
- s.list :tags, :string
156
- s.map :scores, :string, :double
157
- s.struct :address do |address|
158
- address.string :city
159
- address.string :zip
160
- end
161
- s.decimal :price, precision: 12, scale: 2
162
- s.json :payload
163
- s.timestamp :created_at
164
- end
165
-
166
- File.open("out.parquet", "wb") do |file|
167
- Herringbone::Writer.open(file, schema) do |w|
168
- w << { "id" => 1, "name" => "Anna", "status" => "paid", "tags" => ["a", "b"],
169
- "scores" => { "x" => 1.5 }, "address" => { "city" => "Amsterdam" },
170
- "price" => BigDecimal("9.99"), "payload" => { "any" => ["json"] }, "created_at" => Time.now }
171
- w << { id: 2, status: :pending } # Symbol keys and values work; missing keys are nulls
172
- w << [3, "Bo", nil, nil, nil, nil, nil, nil, nil, nil] # Arrays in schema order
173
- w << order # anything with #attributes (ActiveRecord) or #to_h (Struct, Data)
174
- end
175
- end
94
+ Herringbone::Key.generate.hex # => "9f86d081884c7d65..." - put it in your credentials or ENV
176
95
  ```
177
96
 
178
- `Herringbone.write(io, rows)` writes an Enumerable of rows in one go, inferring the schema from the
179
- first 1000 rows unless `schema:` is given; fields declared in a block replace inferred ones:
180
- `Herringbone.write(io, rows) { |s| s.json :payload }`. The rows are iterated once, holding back only
181
- those first 1000, so lazy Enumerators and cursors that can't be rewound work. A later row that
182
- doesn't fit the inferred types raises `Herringbone::SchemaMismatch`, which explains what was
183
- inferred and how to declare the column, and leaves the file unfinished.
184
-
185
- The writer writes to any IO that responds to `#write` (a `File`, `StringIO`, `Tempfile`, socket or
186
- pipe), sequentially, and never seeks, rewinds or closes it (it does switch it to binary mode). If
187
- the `Writer.open` block raises (or `#abort` is called), no footer is written and what was written
188
- so far is left for you to discard. To replace a file only once it is complete, write to a temporary
189
- file and rename it.
190
-
191
- The block receives the schema builder, and the blocks of `struct`, `list` and `map` receive a
192
- builder of their own, so the block can still call your methods and read your instance variables.
193
- A block that takes no parameter raises `ArgumentError`.
194
-
195
- Column types: `boolean int8 int16 int32 int64 uint8 uint16 uint32 uint64 float double float16
196
- string binary json bson enum uuid date int96 time timestamp decimal fixed`, plus `struct`, `list`
197
- and `map`; `s.column :name, :int32` declares one by name. Fields are nullable unless `null: false`
198
- is given; list elements unless `element_null: false`, map values unless `value_null: false`.
199
- Nested lists: `s.list :matrix do |matrix| matrix.list :element, :double end`. `time` and `timestamp` take `unit:`
200
- (`:millis`, `:micros`, `:nanos`) and `utc:`.
201
-
202
- `enum` is a string column. `values:` restricts what may be written, and also takes a Rails-style
203
- Hash (`values: Order.statuses`), in which case both labels and stored integers are accepted and
204
- the label is written. `parquet_enum: true` adds the Parquet ENUM annotation (pyarrow and pandas
205
- read such columns as binary, which is why it is off by default).
206
-
207
- Columns accept the values Ruby and Rails code usually has at hand:
208
-
209
- | column | accepts |
210
- |---|---|
211
- | `date` | `Date`, `Time`/`DateTime` (their date), `"2024-05-01"` |
212
- | `timestamp` | `Time`, `DateTime`, `ActiveSupport::TimeWithZone`, `Date` (midnight UTC), ISO-8601 strings, Integers in the column's unit |
213
- | `time` | `Time` (its time of day, as Rails returns for `time` columns), `"13:45:30.25"`, Integers |
214
- | `json` | Strings as-is, anything else through `JSON.generate` |
215
- | `string`, `enum` | Strings, Symbols, anything with `to_s` |
216
- | `boolean` | `true`/`false`, `1`/`0`, `"t"`/`"f"`, `"true"`/`"false"`, `"yes"`/`"no"` |
217
- | integers | Integers, whole-number Floats/BigDecimals/Rationals, numeric Strings; out-of-range values raise |
218
- | `decimal` | `BigDecimal`, Integer, Rational, Float, numeric Strings |
219
- | `uuid` | Strings with or without dashes, or 16 raw bytes |
220
-
221
- Values that don't fit raise `Herringbone::EncodeError` naming the row number and column path; the
222
- failed row is discarded and the writer can carry on.
223
-
224
- Writer options:
225
-
226
- | option | default | |
227
- |---|---|---|
228
- | `compression` | `:snappy` | `:none`, `:snappy`, `:gzip`, `:lz4` (LZ4_RAW), `:lz4_hadoop`, `:zstd`, `:brotli` |
229
- | `row_group_bytes` | 16MB | flush a row group once the buffered values take about this much memory, which bounds memory use (a 15-column table peaks around 290 MB RSS) |
230
- | `row_group_rows` | none | also flush after this many rows |
231
- | `page_bytes` | 1MB | approximate data page size |
232
- | `page_rows` | `20_000` | at most this many rows per data page |
233
- | `data_page_version` | `1` | `1` or `2` |
234
- | `dictionary` | `true` | `false`, or an Array of column paths to dictionary-encode |
235
- | `encodings` | `{}` | e.g. `{ "id" => :delta_binary_packed, "x" => :byte_stream_split }` |
236
- | `metadata` | `{}` | footer key/value metadata, read back with `reader.metadata` |
237
- | `bloom_filters` | none | `true`, an Array of column paths, or `{ "path" => { ndv:, fpp:, max_bytes: } }` |
238
-
239
- ### Coming from CSV
240
-
241
- `SimpleWriter` writes like the CSV gem: name the columns, then append Arrays. The types are
242
- inferred from the first 1000 rows, as above.
97
+ Then write with it:
243
98
 
244
99
  ```ruby
245
100
  File.open("people.parquet", "wb") do |file|
246
101
  Herringbone::SimpleWriter.open(file) do |sw|
247
- sw.headers!(:id, :name, :age)
248
- sw << [123, "John", 12]
249
- sw << { id: 124, name: "Jane" } # Hashes work too
102
+ sw.encrypt!(key: ENV["PARQUET_KEY"])
103
+ sw.headers!(:id, :name, :email)
104
+ sw << [1, "John", "john@example.com"]
250
105
  end
251
106
  end
252
107
  ```
253
108
 
254
- Unlike CSV, headers are required. A column that holds more than one type can be declared up front:
255
- `Herringbone::SimpleWriter.new(io) { |s| s.string :code }` (then call `close` when done).
256
-
257
- ### Statistics, page indexes and bloom filters
258
-
259
- Every column chunk gets min/max/null-count statistics and the Parquet page index (per-page min/max
260
- and row offsets), which Herringbone, DuckDB, Spark, Trino, Arrow and DataFusion use to skip pages.
261
- With 20,000 rows per page, `WHERE id BETWEEN ...` on a sorted column reads a handful of pages
262
- instead of the whole row group, so sort rows by the columns you filter on.
263
-
264
- Min/max cannot rule out a value inside a row group's range, the usual case for unsorted IDs,
265
- emails or UUIDs; a bloom filter can. They are off by default. `bloom_filters: ["user_id", "email"]`
266
- (or `true` for every non-boolean column) writes Parquet's split block bloom filters, sized from
267
- each row group's distinct values at a false positive probability of 1% (`fpp:`, up to `max_bytes:`,
268
- default 1MB) unless `ndv:` gives the number of distinct values. Nested leaves are named by their
269
- dotted path (`"tags.list.element"`). Hashing is pure Ruby unless the `xxhash` gem is loaded:
270
- a million rows take about 1.2 s longer to write with a filter on an INT64 column and 2.8 s longer
271
- with one on a ~22-byte string column, and about 0.5 s longer with `xxhash`.
272
-
273
- ### Writing to S3
274
-
275
- Parquet keeps its metadata in a footer, so a file can be streamed into an S3 multipart upload
276
- without a local copy, using `upload_stream` from `aws-sdk-s3`:
109
+ and read with it:
277
110
 
278
111
  ```ruby
279
- s3 = Aws::S3::TransferManager.new # aws-sdk-s3 1.197+; before that, Aws::S3::Object#upload_stream
280
- s3.upload_stream(bucket: "exports", key: "events.parquet", part_size: 16 * 1024 * 1024) do |io|
281
- Herringbone::Writer.open(io, schema) do |w|
282
- events.each { |event| w << event }
283
- end
284
- end
285
- s3.upload_stream(bucket: "exports", key: "orders.parquet") { |io| Herringbone.write(io, Order.all) }
112
+ Herringbone::Reader.new(file, decryption: ENV["PARQUET_KEY"]).read
286
113
  ```
287
114
 
288
- If the block raises, the SDK aborts the multipart upload and raises `Aws::S3::MultipartUploadError`,
289
- so no partial object is left. S3 allows at most 10,000 parts, which with the default 5MB parts caps
290
- the file at about 48GB; raise `part_size:` for bigger files.
115
+ The schema, the values and the statistics are all encrypted (AES-256-GCM); without the key the
116
+ file is just noise. `encryption: key` does the same for `Herringbone.write` and
117
+ `Herringbone::Writer`. Your colleagues can open the file with the same key in pyarrow 25+
118
+ (`pq.read_table(path, decryption_properties=pyarrow.parquet.encryption.create_decryption_properties(bytes.fromhex(key)))`),
119
+ Arrow, DataFusion, Trino or Spark, and `herringbone inspect people.parquet` asks for the key.
120
+ When the time comes to rotate, read with all your keys, `decryption: [new_key, old_key]`, and
121
+ each file finds its own. See [Encryption](MANUAL.md#encryption) for per-column keys, plaintext
122
+ footers and which tools read what.
291
123
 
292
- ## Redaction
124
+ ## Redacting
293
125
 
294
126
  `Herringbone.redact` rewrites a file with rows removed or values replaced, for GDPR "forget me"
295
- requests and pseudonymization. One file goes in and one comes out, and the parts of the file
296
- nothing touches are copied byte for byte.
127
+ requests and pseudonymization. The parts of the file nothing touches are copied byte for byte.
297
128
 
298
129
  ```ruby
299
130
  # Forget me: remove the rows
@@ -306,11 +137,6 @@ Herringbone.redact(input, output) do |r|
306
137
  r.where(user_id: 42).replace(email: nil, name: nil, address: nil)
307
138
  end
308
139
 
309
- # Forget me, but keep the row linkable: pseudonymize just that user's email
310
- Herringbone.redact(input, output) do |r|
311
- r.where(user_id: 42).replace(:email) { |email| OpenSSL::HMAC.hexdigest("SHA256", KEY, email) }
312
- end
313
-
314
140
  # Pseudonymize a column across the whole file, mask another, drop a third
315
141
  Herringbone.redact(input, output) do |r|
316
142
  r.replace(:email) { |email| OpenSSL::HMAC.hexdigest("SHA256", KEY, email.downcase) }
@@ -319,213 +145,18 @@ Herringbone.redact(input, output) do |r|
319
145
  end
320
146
  ```
321
147
 
322
- A `Herringbone::Redaction` is built once and applies itself to any number of files, which suits a
323
- forget-me job going over a bucket:
148
+ See [Redaction](MANUAL.md#redaction) for reusable redactions, reports and recipes.
324
149
 
325
- ```ruby
326
- forget = Herringbone::Redaction.new do |r|
327
- r.where(user_id: 42).delete
328
- r.where(email: "anna@example.com").delete # separate statements OR together
329
- r.where(created_at: ..2.years.ago).delete # retention works the same way
330
- end
150
+ ## Looking inside a file
331
151
 
332
- if forget.affects?(input) # statistics and bloom filters first, then the where columns only
333
- report = forget.apply(input, output)
334
- report.rows_deleted # => 3
335
- report.row_groups # => { copied: 61, rewritten: 1 }
336
- end
337
152
  ```
338
-
339
- The block receives the redaction being built, like `Writer.open` passes the writer, so it can use
340
- the methods and instance variables around it; a block without the `|r|` raises `ArgumentError`.
341
- `where` takes what `read(where:)` takes (values, Arrays, Ranges, `nil`, callables and dotted
342
- struct paths, but no columns inside lists or maps), and skips row groups and pages the same way.
343
- `where(...).delete` removes the matching rows. `replace` sets constants,
344
- `replace(email: nil, name: "[deleted]")`, or computes each value with a block,
345
- `replace(:email, :phone) { |value| ... }`. A block that takes two parameters also gets the whole
346
- row, as a Hash with String keys like `read` returns. Without `where`, `replace` applies to every
347
- row. A column is a top-level field or a struct member by dotted path (`"address.city"`; a member
348
- of a null struct is left alone). To change what is inside a list or map, replace the whole field
349
- with a block that receives the Array or Hash. `drop :a, :b` removes columns from the schema.
350
-
351
- Statements apply in declared order, row by row: a deleted row is gone for later statements, and
352
- later statements see what earlier ones replaced, in their conditions as well as in their blocks.
353
- Mistakes raise `ArgumentError` before anything is written: a `where` without a verb, a column that
354
- doesn't exist, `nil` (or a value of the wrong type) for a `null: false` column. A block that
355
- returns something its column can't store raises `Herringbone::EncodeError` when it gets there and
356
- leaves the output unfinished.
357
-
358
- `apply` and `Herringbone.redact` take IOs like `Reader` and `Writer` do: the input must be
359
- seekable, the output only needs `#write`, and neither is closed. To redact in place, write to a
360
- temporary file and rename it over the original; on S3, read the object and write the new one
361
- through `upload_stream`. The `Redaction::Report` that `apply` returns has `rows_read`,
362
- `rows_deleted`, `rows_changed` and `row_groups` (`{ copied:, rewritten: }`), which is the audit
363
- trail an erasure needs.
364
-
365
- Each row group is handled in the cheapest way that is still exact. If no statement can match it,
366
- judging by statistics, bloom filters and the page index, or by reading only the `where` columns,
367
- its column chunks are copied as they are and only their offsets are rebased. If rows match but
368
- none is deleted and only leaf columns change, just those column chunks are encoded again. If rows
369
- are deleted, or a nested field is replaced whole, the row group is rewritten, as one row group
370
- (or none, when every row goes). Rewritten chunks keep the codec of the original chunk (LZO, which
371
- Herringbone can't write, becomes Snappy) and get a new bloom filter if they had one. Writer
372
- options (`compression:`, `bloom_filters:`, `page_rows:`...) apply to the rewritten chunks;
373
- `row_group_bytes:` and `row_group_rows:` are refused, since row groups keep their boundaries. The
374
- footer's key/value metadata is copied (without `ARROW:schema` and `pandas` when columns are
375
- dropped, since those describe the columns) unless `metadata:` is given.
376
-
377
- No deleted or replaced value survives in the output: not in data pages, dictionary pages, column
378
- chunk min/max statistics, the page index or bloom filters. What Herringbone can't reach is up to
379
- you: the original file, its S3 versions and backups still hold the data until you delete them,
380
- only the columns you name are touched (an email that also sits in a free-text `notes` column stays
381
- there), and other files with the same person in them need their own pass.
382
-
383
- Some recipes. Keyed hashing keeps a column joinable without keeping the value; keep the key out
384
- of the data, and rotate or destroy it to cut the link:
385
-
386
- ```ruby
387
- require "openssl"
388
- KEY = ENV.fetch("PSEUDONYM_KEY")
389
- Herringbone.redact(input, output) do |r|
390
- r.replace(:email) { |email| email && OpenSSL::HMAC.hexdigest("SHA256", KEY, email.downcase) }
391
- end
153
+ bin/herringbone cat FILE [N] # rows as JSON lines
154
+ bin/herringbone inspect FILE # schema, row groups, column chunks
155
+ bin/herringbone inspect FILE --format=html # a byte map of the file, opened in your browser
392
156
  ```
393
157
 
394
- Masking keeps enough to be recognizable to the person but not to anyone else:
395
-
396
- ```ruby
397
- Herringbone.redact(input, output) do |r|
398
- r.replace(:card_number) { |number| number && number[-4..].rjust(number.size, "*") }
399
- r.replace(:email) { |email| email&.sub(/\A(.).*@/, '\1***@') }
400
- r.replace(:birth_date) { |date| date && Date.new(date.year, 1, 1) }
401
- end
402
- ```
403
-
404
- Fake values (with the `faker` gem) make a copy for staging that looks real. Seed Faker from the
405
- original value to get the same fake for the same person in every file:
406
-
407
- ```ruby
408
- require "faker"
409
- require "zlib"
410
- Herringbone.redact(input, output) do |r|
411
- r.replace(:name) do |name|
412
- next nil unless name
413
- Faker::Config.random = Random.new(Zlib.crc32(name))
414
- Faker::Name.name
415
- end
416
- r.replace(:address) do |address|
417
- address && { "city" => Faker::Address.city, "zip" => Faker::Address.zip_code }
418
- end
419
- end
420
- ```
421
-
422
- ## ActiveRecord
423
-
424
- `Herringbone.write` also takes a model or relation, which it reads with `find_each`, using a schema
425
- built from the model's columns. It returns the number of rows written:
426
-
427
- ```ruby
428
- File.open("orders.parquet", "wb") do |file|
429
- Herringbone.write(file, Order.where(created_at: 1.year.ago..), compression: :zstd)
430
- end
431
- schema = Herringbone::Schema.from_active_record(Order, only: %w[id status total])
432
- Herringbone.write(io, Order, schema: schema)
433
- ```
434
-
435
- `from_active_record` takes `only:`, `except:` and `parquet_enum: true`. Rails is not a dependency:
436
- it only calls `columns`, `primary_key` and `defined_enums` on the model. Enum attributes are
437
- written as their labels, and the writer rejects values outside the enum.
438
-
439
- | Column | Parquet |
440
- |---|---|
441
- | `integer` | `int16`/`int32`/`int64` by SQL type (`smallint`, `integer`, `bigint`...) or limit; `uintN` for `unsigned`; primary keys are always `int64` |
442
- | `float` | `double` (`float` for Postgres `float4`) |
443
- | `decimal` | `decimal(precision, scale)`; `decimal(38, 9)` when the column has no precision |
444
- | `boolean`, `date`, `binary`, `uuid` | same |
445
- | `json`, `jsonb` | `json` |
446
- | `datetime`, `timestamp`, `timestamptz` | `timestamp` (microseconds, UTC) |
447
- | `time` | `time` (microseconds) |
448
- | `hstore` | `map` of `string` to `string` |
449
- | enum attributes | `string` (or `enum`) |
450
- | `string`, `text`, `citext`, anything else | `string` |
451
- | Postgres array columns | `list` of the element type |
452
-
453
- Columns declared `NOT NULL` (and primary keys) are required, all others nullable. Column order
454
- follows `Model.columns`.
455
-
456
- ## Type mapping
457
-
458
- | Parquet | Ruby |
459
- |---|---|
460
- | BOOLEAN | `true`/`false` |
461
- | INT32/INT64 (incl. signed/unsigned INTEGER) | `Integer` |
462
- | FLOAT, DOUBLE, FLOAT16 | `Float` |
463
- | STRING, ENUM, JSON | `String` (UTF-8) |
464
- | BYTE_ARRAY, FIXED_LEN_BYTE_ARRAY, BSON | `String` (binary) |
465
- | DATE | `Date` (proleptic Gregorian) |
466
- | TIMESTAMP, INT96 | `Time` (UTC) |
467
- | TIME | `Integer` in the column's unit since midnight |
468
- | DECIMAL | `BigDecimal` |
469
- | UUID | `String` like `"0f1e2d3c-..."` |
470
- | struct / list / map | `Hash` / `Array` / `Hash` |
471
-
472
- ## Inspecting files
473
-
474
- > The inspector and its HTML view are modelled on
475
- > **[Parquet X-ray](https://huggingface.co/spaces/cfahlgren1/parquet-xray) by cfahlgren1** —
476
- > the design and the idea are theirs. Go check it out. It is amazing!
477
-
478
- `Herringbone::Inspector` examines a file using only its footer, page headers, page indexes and
479
- bloom filter headers. Nothing is decompressed, so it is fast on big files and works for ZSTD and
480
- Brotli files without the codec gems.
481
-
482
- ```ruby
483
- File.open("data.parquet", "rb") do |file|
484
- inspector = Herringbone::Inspector.new(file)
485
- inspector.summary # size, rows, row groups, codecs, page index / bloom filter presence...
486
- inspector.row_groups[0].column("name").pages # page headers; also statistics, column_index...
487
- puts inspector.report # the schema, every column chunk, key/value metadata (Arrow schema decoded)
488
- inspector.to_h # all of it, JSON-serializable
489
- inspector.to_html # one self-contained HTML page with a to-scale byte map of the file
490
- end
491
- ```
492
-
493
- `inspector.verify_checksums` reads every page body (still without decompressing) to check the page
494
- CRCs; the results then appear in `summary`, `report`, `to_h` and `to_html`.
495
-
496
- ## Command line
497
-
498
- ```
499
- bin/herringbone cat FILE [N] # rows as JSON lines
500
- bin/herringbone inspect FILE [--pages] # text report (--pages lists every page header)
501
- bin/herringbone inspect FILE --format=json # everything as JSON
502
- bin/herringbone inspect FILE --format=html # the HTML page, opened in your browser
503
- bin/herringbone inspect FILE --format=html > out.html # the HTML page, saved
504
- ```
505
-
506
- Add `--verify-checksums` to any `inspect` form to check page CRCs.
507
-
508
- `--format` is `text` (the default), `json` or `html`; `--pages` only goes with `text`.
509
- In a terminal, `--format=html` writes the page to a temp file, prints its path and opens it with `open` on
510
- macOS, `start` on Windows or `xdg-open` elsewhere. When stdout is redirected or piped it prints the
511
- page instead.
512
-
513
- ## Supported format features
514
-
515
- - Encodings (read and write): PLAIN, PLAIN_DICTIONARY/RLE_DICTIONARY, RLE, DELTA_BINARY_PACKED,
516
- DELTA_LENGTH_BYTE_ARRAY, DELTA_BYTE_ARRAY, BYTE_STREAM_SPLIT; legacy BIT_PACKED levels (read)
517
- - Data page v1 and v2, dictionary pages, page CRCs (written)
518
- - Page indexes and split block bloom filters (read and written)
519
- - Legacy list and map layouts per the Parquet backward-compatibility rules
520
- - Not supported: encryption, column chunks in external files
158
+ See [Inspecting files](MANUAL.md#inspecting-files) and [Command line](MANUAL.md#command-line).
521
159
 
522
160
  ## Development
523
161
 
524
- ```
525
- bundle install
526
- bundle exec rake test
527
- HERRINGBONE_PYTHON=/path/to/python-with-pyarrow bundle exec rake test # also run pyarrow interop tests
528
- ```
529
-
530
- `test/fixtures/parquet-testing` holds files from [apache/parquet-testing](https://github.com/apache/parquet-testing)
531
- (Apache-2.0); expectations for them were generated with pyarrow, see `test/fixtures/generate_expectations.py`.
162
+ See [Development](MANUAL.md#development) and [CONTRIBUTING](CONTRIBUTING.md).