herringbone 0.5.0 → 0.6.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +28 -0
- data/MANUAL.md +773 -0
- data/README.md +68 -437
- data/bin/herringbone +86 -7
- data/lib/herringbone/bloom_filter.rb +0 -55
- data/lib/herringbone/compression.rb +27 -4
- data/lib/herringbone/encryption.rb +716 -0
- data/lib/herringbone/encryption_configuration.rb +304 -0
- data/lib/herringbone/format.rb +55 -3
- data/lib/herringbone/inferring_writer.rb +17 -2
- data/lib/herringbone/inspector.rb +158 -28
- data/lib/herringbone/key.rb +107 -0
- data/lib/herringbone/reader/bloom_filters.rb +91 -0
- data/lib/herringbone/reader/column_chunk_reader.rb +55 -1
- data/lib/herringbone/reader.rb +100 -15
- data/lib/herringbone/redaction/rewriter.rb +59 -7
- data/lib/herringbone/redaction.rb +20 -5
- data/lib/herringbone/schema.rb +2 -0
- data/lib/herringbone/simple_writer.rb +28 -0
- data/lib/herringbone/version.rb +1 -1
- data/lib/herringbone/writer.rb +102 -24
- data/lib/herringbone.rb +69 -25
- metadata +7 -2
data/README.md
CHANGED
|
@@ -7,20 +7,13 @@ A pure-Ruby reader and writer for [Apache Parquet](https://parquet.apache.org/)
|
|
|
7
7
|
LZO, from Hadoop-era files, can be read but not written
|
|
8
8
|
- Full nesting support (structs, lists, maps, any depth) via Dremel record shredding/assembly
|
|
9
9
|
- Reads files from parquet-mr, Arrow, Spark, Impala, DuckDB, Rust writers etc.
|
|
10
|
-
- Optimized reads with batches and pages
|
|
10
|
+
- Optimized reads with batches and pages
|
|
11
|
+
- Parquet modular encryption (per column and of the footer), interoperable with parquet-mr and Arrow
|
|
11
12
|
- Ruby 3.0+
|
|
12
13
|
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
```ruby
|
|
18
|
-
s3 = Aws::S3::TransferManager.new
|
|
19
|
-
s3.upload_stream(bucket: "exports", key: "payments.parquet") do |io|
|
|
20
|
-
# Schema will be auto-inferred, find_each will be used automatically
|
|
21
|
-
Herringbone.write(io, Payment.where(status: "settled", created_at: 1.month.ago..))
|
|
22
|
-
end
|
|
23
|
-
```
|
|
14
|
+
This README covers the common cases. The [MANUAL](MANUAL.md) has everything else: reader and
|
|
15
|
+
writer options, schemas and column types, bloom filters, Numo arrays, encryption, the redaction
|
|
16
|
+
rules, type mapping and the inspector.
|
|
24
17
|
|
|
25
18
|
## Installation
|
|
26
19
|
|
|
@@ -28,272 +21,110 @@ end
|
|
|
28
21
|
gem "herringbone"
|
|
29
22
|
```
|
|
30
23
|
|
|
31
|
-
|
|
24
|
+
Optionally add `snappy` (faster), `zstd-ruby` and `brotli` (more codecs), `xxhash` (faster bloom
|
|
25
|
+
filters) or `numo-narray-alt` (`read(as: :numo)`). Herringbone uses whichever of them are loaded,
|
|
26
|
+
see [Optional gems](MANUAL.md#optional-gems).
|
|
32
27
|
|
|
33
|
-
|
|
34
|
-
gem "snappy" # native Snappy, 2-3x faster reads and writes of typical files
|
|
35
|
-
gem "zstd-ruby" # ZSTD
|
|
36
|
-
gem "brotli" # Brotli
|
|
37
|
-
gem "xxhash" # faster bloom filters
|
|
38
|
-
gem "numo-narray-alt" # read(as: :numo)
|
|
39
|
-
```
|
|
28
|
+
## Exporting from Rails
|
|
40
29
|
|
|
41
|
-
|
|
42
|
-
loads them, otherwise require them yourself:
|
|
30
|
+
Stream a relation straight into S3, no temp file needed:
|
|
43
31
|
|
|
44
32
|
```ruby
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
33
|
+
s3 = Aws::S3::TransferManager.new
|
|
34
|
+
s3.upload_stream(bucket: "exports", key: "payments.parquet") do |io|
|
|
35
|
+
# Schema will be auto-inferred, find_each will be used automatically
|
|
36
|
+
Herringbone.write(io, Payment.where(status: "settled", created_at: 1.month.ago..))
|
|
37
|
+
end
|
|
48
38
|
```
|
|
49
39
|
|
|
50
|
-
|
|
51
|
-
`Ractor.make_shareable(schema)`. Rows made shareable the same way reach a writing Ractor by
|
|
52
|
-
reference, without being copied:
|
|
40
|
+
Or to a local file:
|
|
53
41
|
|
|
54
42
|
```ruby
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
Herringbone::Writer.open(f, schema) do |w|
|
|
58
|
-
while (row = Ractor.receive) != :done
|
|
59
|
-
w << row
|
|
60
|
-
end
|
|
61
|
-
end
|
|
62
|
-
end
|
|
43
|
+
File.open("orders.parquet", "wb") do |file|
|
|
44
|
+
Herringbone.write(file, Order.where(created_at: 1.year.ago..), compression: :zstd)
|
|
63
45
|
end
|
|
64
|
-
rows.each { |row| writer.send(Ractor.make_shareable(row)) }
|
|
65
|
-
writer.send(:done)
|
|
66
46
|
```
|
|
67
47
|
|
|
68
|
-
|
|
48
|
+
Any Enumerable of Hashes, Structs or Arrays works too; the schema is inferred from the first 1000
|
|
49
|
+
rows. See [ActiveRecord](MANUAL.md#activerecord) and [Writing to S3](MANUAL.md#writing-to-s3).
|
|
69
50
|
|
|
70
|
-
|
|
51
|
+
## Reading
|
|
71
52
|
|
|
72
53
|
```ruby
|
|
73
|
-
require "herringbone"
|
|
74
|
-
|
|
75
54
|
File.open("data.parquet", "rb") do |file|
|
|
76
55
|
reader = Herringbone::Reader.new(file)
|
|
77
|
-
reader.schema
|
|
78
56
|
reader.num_rows
|
|
79
57
|
|
|
80
|
-
reader.each_row { |row| p row }
|
|
81
|
-
reader.each_batch(1000) { |rows| ... }
|
|
82
|
-
reader.
|
|
83
|
-
reader.read # all rows at once; read(as: :columns) for column Arrays
|
|
84
|
-
end
|
|
85
|
-
```
|
|
86
|
-
|
|
87
|
-
Reads stream: pages are read and decoded one at a time per column and rows are assembled in
|
|
88
|
-
batches, so memory depends on the batch and page sizes rather than on the row group size (1M
|
|
89
|
-
rows in a single row group peak at about 60–160 MB RSS growth instead of 660 MB). Column-order
|
|
90
|
-
batches (`as: :columns`) skip building a Hash per row and are 25–30% faster.
|
|
58
|
+
reader.each_row { |row| p row } # Hashes with String keys, nested values as Hash/Array
|
|
59
|
+
reader.each_batch(1000) { |rows| ... } # Arrays of up to 1000 rows
|
|
60
|
+
reader.read # all rows at once
|
|
91
61
|
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
in Rails (which yields `ActiveSupport::TimeWithZone`). Zone names like `"Europe/Amsterdam"` work
|
|
96
|
-
when ActiveSupport or TZInfo is loaded. Timestamps stored with `isAdjustedToUTC=false` are
|
|
97
|
-
wall-clock values and stay as they are.
|
|
98
|
-
|
|
99
|
-
### Selecting rows
|
|
100
|
-
|
|
101
|
-
`each_row`, `each_batch` and `read` take `columns:`, `where:`, `from:` and `limit:`:
|
|
102
|
-
|
|
103
|
-
```ruby
|
|
104
|
-
reader.read(columns: %w[id name]) # projection
|
|
105
|
-
reader.read(where: { user_id: 42 })
|
|
106
|
-
reader.read(where: { status: %w[paid shipped], created_at: 1.week.ago.. }) # IN, Ranges
|
|
107
|
-
reader.read(where: { "address.city" => "Amsterdam", deleted_at: nil }) # struct members, IS NULL
|
|
108
|
-
reader.read(where: { amount: ->(v) { v && v > 100 } }) # any callable
|
|
109
|
-
reader.read(from: 1_000_000, limit: 100) # rows 1,000,000..1,000,099
|
|
62
|
+
reader.read(columns: %w[id name], where: { user_id: 42 })
|
|
63
|
+
reader.read(where: { status: %w[paid shipped], created_at: 1.week.ago.. })
|
|
64
|
+
end
|
|
110
65
|
```
|
|
111
66
|
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
bloom filters, pages using the page index (written by Herringbone, parquet-mr, Arrow and others),
|
|
115
|
-
and `from:` jumps to its row through the offset index; every row read is then checked, so results
|
|
116
|
-
are exact. On a 1M-row file with 20k-row pages, looking up one `id` takes 0.1 s instead of 6.3 s.
|
|
117
|
-
Filters help most on columns the data is sorted or clustered by. `reader.scan_plan(where: ...)`
|
|
118
|
-
shows which row groups and row ranges a read would touch, without reading them.
|
|
67
|
+
Reads stream page by page, and `where:` skips row groups and pages it can rule out using
|
|
68
|
+
statistics, page indexes and bloom filters. See [Reading](MANUAL.md#reading).
|
|
119
69
|
|
|
120
|
-
|
|
70
|
+
## Writing CSV-style
|
|
121
71
|
|
|
122
|
-
`
|
|
123
|
-
the same `columns:`, `where:`, `from:` and `limit:`. Add `gem "numo-narray-alt"` (or
|
|
124
|
-
`numo-narray`) to your Gemfile and `require "numo/narray"`: without it, `as: :numo` raises
|
|
125
|
-
`Herringbone::UnsupportedError` naming the gem.
|
|
72
|
+
`SimpleWriter` writes like the CSV gem: name the columns, then append rows. The types are inferred.
|
|
126
73
|
|
|
127
74
|
```ruby
|
|
128
|
-
|
|
129
|
-
|
|
75
|
+
File.open("people.parquet", "wb") do |file|
|
|
76
|
+
Herringbone::SimpleWriter.open(file) do |sw|
|
|
77
|
+
sw.headers!(:id, :name, :age)
|
|
78
|
+
sw << [123, "John", 12]
|
|
79
|
+
sw << { id: 124, name: "Jane" } # Hashes work too
|
|
80
|
+
end
|
|
81
|
+
end
|
|
130
82
|
```
|
|
131
83
|
|
|
132
|
-
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
| Parquet | Numo |
|
|
136
|
-
|---|---|
|
|
137
|
-
| INT32, INT64 (also TIME) | `Int32`, `Int64` |
|
|
138
|
-
| INT(8/16) signed; INT(8/16/32/64) unsigned | `Int8`, `Int16`; `UInt8` … `UInt64` |
|
|
139
|
-
| FLOAT, FLOAT16, DOUBLE | `SFloat`, `SFloat`, `DFloat` (nulls are NaN) |
|
|
140
|
-
| integers with nulls | `DFloat` with NaN (exact up to 2**53) |
|
|
141
|
-
| BOOLEAN | `Bit`; with nulls `RObject` of true/false/nil |
|
|
142
|
-
| list of numbers, every row the same length and no nulls | 2-D `[rows, length]` (e.g. embeddings) |
|
|
143
|
-
| strings, binary, decimals, dates, timestamps, UUIDs, structs, maps, other lists | `RObject` of the values `read` returns |
|
|
84
|
+
To control the types, declare a schema and use `Herringbone::Writer`, see
|
|
85
|
+
[Writing](MANUAL.md#writing).
|
|
144
86
|
|
|
145
|
-
|
|
146
|
-
with `each_batch` it can differ between batches.
|
|
87
|
+
## Encrypting
|
|
147
88
|
|
|
148
|
-
|
|
89
|
+
Parquet files tend to wander off: to S3 buckets, laptops, other teams. Encrypting them takes one
|
|
90
|
+
key and one line, no KMS, no Hadoop configuration. Make a key once and keep it with your other
|
|
91
|
+
secrets:
|
|
149
92
|
|
|
150
93
|
```ruby
|
|
151
|
-
|
|
152
|
-
s.int64 :id, null: false
|
|
153
|
-
s.string :name
|
|
154
|
-
s.enum :status, values: %w[pending paid shipped]
|
|
155
|
-
s.list :tags, :string
|
|
156
|
-
s.map :scores, :string, :double
|
|
157
|
-
s.struct :address do |address|
|
|
158
|
-
address.string :city
|
|
159
|
-
address.string :zip
|
|
160
|
-
end
|
|
161
|
-
s.decimal :price, precision: 12, scale: 2
|
|
162
|
-
s.json :payload
|
|
163
|
-
s.timestamp :created_at
|
|
164
|
-
end
|
|
165
|
-
|
|
166
|
-
File.open("out.parquet", "wb") do |file|
|
|
167
|
-
Herringbone::Writer.open(file, schema) do |w|
|
|
168
|
-
w << { "id" => 1, "name" => "Anna", "status" => "paid", "tags" => ["a", "b"],
|
|
169
|
-
"scores" => { "x" => 1.5 }, "address" => { "city" => "Amsterdam" },
|
|
170
|
-
"price" => BigDecimal("9.99"), "payload" => { "any" => ["json"] }, "created_at" => Time.now }
|
|
171
|
-
w << { id: 2, status: :pending } # Symbol keys and values work; missing keys are nulls
|
|
172
|
-
w << [3, "Bo", nil, nil, nil, nil, nil, nil, nil, nil] # Arrays in schema order
|
|
173
|
-
w << order # anything with #attributes (ActiveRecord) or #to_h (Struct, Data)
|
|
174
|
-
end
|
|
175
|
-
end
|
|
94
|
+
Herringbone::Key.generate.hex # => "9f86d081884c7d65..." - put it in your credentials or ENV
|
|
176
95
|
```
|
|
177
96
|
|
|
178
|
-
|
|
179
|
-
first 1000 rows unless `schema:` is given; fields declared in a block replace inferred ones:
|
|
180
|
-
`Herringbone.write(io, rows) { |s| s.json :payload }`. The rows are iterated once, holding back only
|
|
181
|
-
those first 1000, so lazy Enumerators and cursors that can't be rewound work. A later row that
|
|
182
|
-
doesn't fit the inferred types raises `Herringbone::SchemaMismatch`, which explains what was
|
|
183
|
-
inferred and how to declare the column, and leaves the file unfinished.
|
|
184
|
-
|
|
185
|
-
The writer writes to any IO that responds to `#write` (a `File`, `StringIO`, `Tempfile`, socket or
|
|
186
|
-
pipe), sequentially, and never seeks, rewinds or closes it (it does switch it to binary mode). If
|
|
187
|
-
the `Writer.open` block raises (or `#abort` is called), no footer is written and what was written
|
|
188
|
-
so far is left for you to discard. To replace a file only once it is complete, write to a temporary
|
|
189
|
-
file and rename it.
|
|
190
|
-
|
|
191
|
-
The block receives the schema builder, and the blocks of `struct`, `list` and `map` receive a
|
|
192
|
-
builder of their own, so the block can still call your methods and read your instance variables.
|
|
193
|
-
A block that takes no parameter raises `ArgumentError`.
|
|
194
|
-
|
|
195
|
-
Column types: `boolean int8 int16 int32 int64 uint8 uint16 uint32 uint64 float double float16
|
|
196
|
-
string binary json bson enum uuid date int96 time timestamp decimal fixed`, plus `struct`, `list`
|
|
197
|
-
and `map`; `s.column :name, :int32` declares one by name. Fields are nullable unless `null: false`
|
|
198
|
-
is given; list elements unless `element_null: false`, map values unless `value_null: false`.
|
|
199
|
-
Nested lists: `s.list :matrix do |matrix| matrix.list :element, :double end`. `time` and `timestamp` take `unit:`
|
|
200
|
-
(`:millis`, `:micros`, `:nanos`) and `utc:`.
|
|
201
|
-
|
|
202
|
-
`enum` is a string column. `values:` restricts what may be written, and also takes a Rails-style
|
|
203
|
-
Hash (`values: Order.statuses`), in which case both labels and stored integers are accepted and
|
|
204
|
-
the label is written. `parquet_enum: true` adds the Parquet ENUM annotation (pyarrow and pandas
|
|
205
|
-
read such columns as binary, which is why it is off by default).
|
|
206
|
-
|
|
207
|
-
Columns accept the values Ruby and Rails code usually has at hand:
|
|
208
|
-
|
|
209
|
-
| column | accepts |
|
|
210
|
-
|---|---|
|
|
211
|
-
| `date` | `Date`, `Time`/`DateTime` (their date), `"2024-05-01"` |
|
|
212
|
-
| `timestamp` | `Time`, `DateTime`, `ActiveSupport::TimeWithZone`, `Date` (midnight UTC), ISO-8601 strings, Integers in the column's unit |
|
|
213
|
-
| `time` | `Time` (its time of day, as Rails returns for `time` columns), `"13:45:30.25"`, Integers |
|
|
214
|
-
| `json` | Strings as-is, anything else through `JSON.generate` |
|
|
215
|
-
| `string`, `enum` | Strings, Symbols, anything with `to_s` |
|
|
216
|
-
| `boolean` | `true`/`false`, `1`/`0`, `"t"`/`"f"`, `"true"`/`"false"`, `"yes"`/`"no"` |
|
|
217
|
-
| integers | Integers, whole-number Floats/BigDecimals/Rationals, numeric Strings; out-of-range values raise |
|
|
218
|
-
| `decimal` | `BigDecimal`, Integer, Rational, Float, numeric Strings |
|
|
219
|
-
| `uuid` | Strings with or without dashes, or 16 raw bytes |
|
|
220
|
-
|
|
221
|
-
Values that don't fit raise `Herringbone::EncodeError` naming the row number and column path; the
|
|
222
|
-
failed row is discarded and the writer can carry on.
|
|
223
|
-
|
|
224
|
-
Writer options:
|
|
225
|
-
|
|
226
|
-
| option | default | |
|
|
227
|
-
|---|---|---|
|
|
228
|
-
| `compression` | `:snappy` | `:none`, `:snappy`, `:gzip`, `:lz4` (LZ4_RAW), `:lz4_hadoop`, `:zstd`, `:brotli` |
|
|
229
|
-
| `row_group_bytes` | 16MB | flush a row group once the buffered values take about this much memory, which bounds memory use (a 15-column table peaks around 290 MB RSS) |
|
|
230
|
-
| `row_group_rows` | none | also flush after this many rows |
|
|
231
|
-
| `page_bytes` | 1MB | approximate data page size |
|
|
232
|
-
| `page_rows` | `20_000` | at most this many rows per data page |
|
|
233
|
-
| `data_page_version` | `1` | `1` or `2` |
|
|
234
|
-
| `dictionary` | `true` | `false`, or an Array of column paths to dictionary-encode |
|
|
235
|
-
| `encodings` | `{}` | e.g. `{ "id" => :delta_binary_packed, "x" => :byte_stream_split }` |
|
|
236
|
-
| `metadata` | `{}` | footer key/value metadata, read back with `reader.metadata` |
|
|
237
|
-
| `bloom_filters` | none | `true`, an Array of column paths, or `{ "path" => { ndv:, fpp:, max_bytes: } }` |
|
|
238
|
-
|
|
239
|
-
### Coming from CSV
|
|
240
|
-
|
|
241
|
-
`SimpleWriter` writes like the CSV gem: name the columns, then append Arrays. The types are
|
|
242
|
-
inferred from the first 1000 rows, as above.
|
|
97
|
+
Then write with it:
|
|
243
98
|
|
|
244
99
|
```ruby
|
|
245
100
|
File.open("people.parquet", "wb") do |file|
|
|
246
101
|
Herringbone::SimpleWriter.open(file) do |sw|
|
|
247
|
-
sw.
|
|
248
|
-
sw
|
|
249
|
-
sw <<
|
|
102
|
+
sw.encrypt!(key: ENV["PARQUET_KEY"])
|
|
103
|
+
sw.headers!(:id, :name, :email)
|
|
104
|
+
sw << [1, "John", "john@example.com"]
|
|
250
105
|
end
|
|
251
106
|
end
|
|
252
107
|
```
|
|
253
108
|
|
|
254
|
-
|
|
255
|
-
`Herringbone::SimpleWriter.new(io) { |s| s.string :code }` (then call `close` when done).
|
|
256
|
-
|
|
257
|
-
### Statistics, page indexes and bloom filters
|
|
258
|
-
|
|
259
|
-
Every column chunk gets min/max/null-count statistics and the Parquet page index (per-page min/max
|
|
260
|
-
and row offsets), which Herringbone, DuckDB, Spark, Trino, Arrow and DataFusion use to skip pages.
|
|
261
|
-
With 20,000 rows per page, `WHERE id BETWEEN ...` on a sorted column reads a handful of pages
|
|
262
|
-
instead of the whole row group, so sort rows by the columns you filter on.
|
|
263
|
-
|
|
264
|
-
Min/max cannot rule out a value inside a row group's range, the usual case for unsorted IDs,
|
|
265
|
-
emails or UUIDs; a bloom filter can. They are off by default. `bloom_filters: ["user_id", "email"]`
|
|
266
|
-
(or `true` for every non-boolean column) writes Parquet's split block bloom filters, sized from
|
|
267
|
-
each row group's distinct values at a false positive probability of 1% (`fpp:`, up to `max_bytes:`,
|
|
268
|
-
default 1MB) unless `ndv:` gives the number of distinct values. Nested leaves are named by their
|
|
269
|
-
dotted path (`"tags.list.element"`). Hashing is pure Ruby unless the `xxhash` gem is loaded:
|
|
270
|
-
a million rows take about 1.2 s longer to write with a filter on an INT64 column and 2.8 s longer
|
|
271
|
-
with one on a ~22-byte string column, and about 0.5 s longer with `xxhash`.
|
|
272
|
-
|
|
273
|
-
### Writing to S3
|
|
274
|
-
|
|
275
|
-
Parquet keeps its metadata in a footer, so a file can be streamed into an S3 multipart upload
|
|
276
|
-
without a local copy, using `upload_stream` from `aws-sdk-s3`:
|
|
109
|
+
and read with it:
|
|
277
110
|
|
|
278
111
|
```ruby
|
|
279
|
-
|
|
280
|
-
s3.upload_stream(bucket: "exports", key: "events.parquet", part_size: 16 * 1024 * 1024) do |io|
|
|
281
|
-
Herringbone::Writer.open(io, schema) do |w|
|
|
282
|
-
events.each { |event| w << event }
|
|
283
|
-
end
|
|
284
|
-
end
|
|
285
|
-
s3.upload_stream(bucket: "exports", key: "orders.parquet") { |io| Herringbone.write(io, Order.all) }
|
|
112
|
+
Herringbone::Reader.new(file, decryption: ENV["PARQUET_KEY"]).read
|
|
286
113
|
```
|
|
287
114
|
|
|
288
|
-
|
|
289
|
-
|
|
290
|
-
the file
|
|
115
|
+
The schema, the values and the statistics are all encrypted (AES-256-GCM); without the key the
|
|
116
|
+
file is just noise. `encryption: key` does the same for `Herringbone.write` and
|
|
117
|
+
`Herringbone::Writer`. Your colleagues can open the file with the same key in pyarrow 25+
|
|
118
|
+
(`pq.read_table(path, decryption_properties=pyarrow.parquet.encryption.create_decryption_properties(bytes.fromhex(key)))`),
|
|
119
|
+
Arrow, DataFusion, Trino or Spark, and `herringbone inspect people.parquet` asks for the key.
|
|
120
|
+
When the time comes to rotate, read with all your keys, `decryption: [new_key, old_key]`, and
|
|
121
|
+
each file finds its own. See [Encryption](MANUAL.md#encryption) for per-column keys, plaintext
|
|
122
|
+
footers and which tools read what.
|
|
291
123
|
|
|
292
|
-
##
|
|
124
|
+
## Redacting
|
|
293
125
|
|
|
294
126
|
`Herringbone.redact` rewrites a file with rows removed or values replaced, for GDPR "forget me"
|
|
295
|
-
requests and pseudonymization.
|
|
296
|
-
nothing touches are copied byte for byte.
|
|
127
|
+
requests and pseudonymization. The parts of the file nothing touches are copied byte for byte.
|
|
297
128
|
|
|
298
129
|
```ruby
|
|
299
130
|
# Forget me: remove the rows
|
|
@@ -306,11 +137,6 @@ Herringbone.redact(input, output) do |r|
|
|
|
306
137
|
r.where(user_id: 42).replace(email: nil, name: nil, address: nil)
|
|
307
138
|
end
|
|
308
139
|
|
|
309
|
-
# Forget me, but keep the row linkable: pseudonymize just that user's email
|
|
310
|
-
Herringbone.redact(input, output) do |r|
|
|
311
|
-
r.where(user_id: 42).replace(:email) { |email| OpenSSL::HMAC.hexdigest("SHA256", KEY, email) }
|
|
312
|
-
end
|
|
313
|
-
|
|
314
140
|
# Pseudonymize a column across the whole file, mask another, drop a third
|
|
315
141
|
Herringbone.redact(input, output) do |r|
|
|
316
142
|
r.replace(:email) { |email| OpenSSL::HMAC.hexdigest("SHA256", KEY, email.downcase) }
|
|
@@ -319,213 +145,18 @@ Herringbone.redact(input, output) do |r|
|
|
|
319
145
|
end
|
|
320
146
|
```
|
|
321
147
|
|
|
322
|
-
|
|
323
|
-
forget-me job going over a bucket:
|
|
148
|
+
See [Redaction](MANUAL.md#redaction) for reusable redactions, reports and recipes.
|
|
324
149
|
|
|
325
|
-
|
|
326
|
-
forget = Herringbone::Redaction.new do |r|
|
|
327
|
-
r.where(user_id: 42).delete
|
|
328
|
-
r.where(email: "anna@example.com").delete # separate statements OR together
|
|
329
|
-
r.where(created_at: ..2.years.ago).delete # retention works the same way
|
|
330
|
-
end
|
|
150
|
+
## Looking inside a file
|
|
331
151
|
|
|
332
|
-
if forget.affects?(input) # statistics and bloom filters first, then the where columns only
|
|
333
|
-
report = forget.apply(input, output)
|
|
334
|
-
report.rows_deleted # => 3
|
|
335
|
-
report.row_groups # => { copied: 61, rewritten: 1 }
|
|
336
|
-
end
|
|
337
152
|
```
|
|
338
|
-
|
|
339
|
-
|
|
340
|
-
|
|
341
|
-
`where` takes what `read(where:)` takes (values, Arrays, Ranges, `nil`, callables and dotted
|
|
342
|
-
struct paths, but no columns inside lists or maps), and skips row groups and pages the same way.
|
|
343
|
-
`where(...).delete` removes the matching rows. `replace` sets constants,
|
|
344
|
-
`replace(email: nil, name: "[deleted]")`, or computes each value with a block,
|
|
345
|
-
`replace(:email, :phone) { |value| ... }`. A block that takes two parameters also gets the whole
|
|
346
|
-
row, as a Hash with String keys like `read` returns. Without `where`, `replace` applies to every
|
|
347
|
-
row. A column is a top-level field or a struct member by dotted path (`"address.city"`; a member
|
|
348
|
-
of a null struct is left alone). To change what is inside a list or map, replace the whole field
|
|
349
|
-
with a block that receives the Array or Hash. `drop :a, :b` removes columns from the schema.
|
|
350
|
-
|
|
351
|
-
Statements apply in declared order, row by row: a deleted row is gone for later statements, and
|
|
352
|
-
later statements see what earlier ones replaced, in their conditions as well as in their blocks.
|
|
353
|
-
Mistakes raise `ArgumentError` before anything is written: a `where` without a verb, a column that
|
|
354
|
-
doesn't exist, `nil` (or a value of the wrong type) for a `null: false` column. A block that
|
|
355
|
-
returns something its column can't store raises `Herringbone::EncodeError` when it gets there and
|
|
356
|
-
leaves the output unfinished.
|
|
357
|
-
|
|
358
|
-
`apply` and `Herringbone.redact` take IOs like `Reader` and `Writer` do: the input must be
|
|
359
|
-
seekable, the output only needs `#write`, and neither is closed. To redact in place, write to a
|
|
360
|
-
temporary file and rename it over the original; on S3, read the object and write the new one
|
|
361
|
-
through `upload_stream`. The `Redaction::Report` that `apply` returns has `rows_read`,
|
|
362
|
-
`rows_deleted`, `rows_changed` and `row_groups` (`{ copied:, rewritten: }`), which is the audit
|
|
363
|
-
trail an erasure needs.
|
|
364
|
-
|
|
365
|
-
Each row group is handled in the cheapest way that is still exact. If no statement can match it,
|
|
366
|
-
judging by statistics, bloom filters and the page index, or by reading only the `where` columns,
|
|
367
|
-
its column chunks are copied as they are and only their offsets are rebased. If rows match but
|
|
368
|
-
none is deleted and only leaf columns change, just those column chunks are encoded again. If rows
|
|
369
|
-
are deleted, or a nested field is replaced whole, the row group is rewritten, as one row group
|
|
370
|
-
(or none, when every row goes). Rewritten chunks keep the codec of the original chunk (LZO, which
|
|
371
|
-
Herringbone can't write, becomes Snappy) and get a new bloom filter if they had one. Writer
|
|
372
|
-
options (`compression:`, `bloom_filters:`, `page_rows:`...) apply to the rewritten chunks;
|
|
373
|
-
`row_group_bytes:` and `row_group_rows:` are refused, since row groups keep their boundaries. The
|
|
374
|
-
footer's key/value metadata is copied (without `ARROW:schema` and `pandas` when columns are
|
|
375
|
-
dropped, since those describe the columns) unless `metadata:` is given.
|
|
376
|
-
|
|
377
|
-
No deleted or replaced value survives in the output: not in data pages, dictionary pages, column
|
|
378
|
-
chunk min/max statistics, the page index or bloom filters. What Herringbone can't reach is up to
|
|
379
|
-
you: the original file, its S3 versions and backups still hold the data until you delete them,
|
|
380
|
-
only the columns you name are touched (an email that also sits in a free-text `notes` column stays
|
|
381
|
-
there), and other files with the same person in them need their own pass.
|
|
382
|
-
|
|
383
|
-
Some recipes. Keyed hashing keeps a column joinable without keeping the value; keep the key out
|
|
384
|
-
of the data, and rotate or destroy it to cut the link:
|
|
385
|
-
|
|
386
|
-
```ruby
|
|
387
|
-
require "openssl"
|
|
388
|
-
KEY = ENV.fetch("PSEUDONYM_KEY")
|
|
389
|
-
Herringbone.redact(input, output) do |r|
|
|
390
|
-
r.replace(:email) { |email| email && OpenSSL::HMAC.hexdigest("SHA256", KEY, email.downcase) }
|
|
391
|
-
end
|
|
153
|
+
bin/herringbone cat FILE [N] # rows as JSON lines
|
|
154
|
+
bin/herringbone inspect FILE # schema, row groups, column chunks
|
|
155
|
+
bin/herringbone inspect FILE --format=html # a byte map of the file, opened in your browser
|
|
392
156
|
```
|
|
393
157
|
|
|
394
|
-
|
|
395
|
-
|
|
396
|
-
```ruby
|
|
397
|
-
Herringbone.redact(input, output) do |r|
|
|
398
|
-
r.replace(:card_number) { |number| number && number[-4..].rjust(number.size, "*") }
|
|
399
|
-
r.replace(:email) { |email| email&.sub(/\A(.).*@/, '\1***@') }
|
|
400
|
-
r.replace(:birth_date) { |date| date && Date.new(date.year, 1, 1) }
|
|
401
|
-
end
|
|
402
|
-
```
|
|
403
|
-
|
|
404
|
-
Fake values (with the `faker` gem) make a copy for staging that looks real. Seed Faker from the
|
|
405
|
-
original value to get the same fake for the same person in every file:
|
|
406
|
-
|
|
407
|
-
```ruby
|
|
408
|
-
require "faker"
|
|
409
|
-
require "zlib"
|
|
410
|
-
Herringbone.redact(input, output) do |r|
|
|
411
|
-
r.replace(:name) do |name|
|
|
412
|
-
next nil unless name
|
|
413
|
-
Faker::Config.random = Random.new(Zlib.crc32(name))
|
|
414
|
-
Faker::Name.name
|
|
415
|
-
end
|
|
416
|
-
r.replace(:address) do |address|
|
|
417
|
-
address && { "city" => Faker::Address.city, "zip" => Faker::Address.zip_code }
|
|
418
|
-
end
|
|
419
|
-
end
|
|
420
|
-
```
|
|
421
|
-
|
|
422
|
-
## ActiveRecord
|
|
423
|
-
|
|
424
|
-
`Herringbone.write` also takes a model or relation, which it reads with `find_each`, using a schema
|
|
425
|
-
built from the model's columns. It returns the number of rows written:
|
|
426
|
-
|
|
427
|
-
```ruby
|
|
428
|
-
File.open("orders.parquet", "wb") do |file|
|
|
429
|
-
Herringbone.write(file, Order.where(created_at: 1.year.ago..), compression: :zstd)
|
|
430
|
-
end
|
|
431
|
-
schema = Herringbone::Schema.from_active_record(Order, only: %w[id status total])
|
|
432
|
-
Herringbone.write(io, Order, schema: schema)
|
|
433
|
-
```
|
|
434
|
-
|
|
435
|
-
`from_active_record` takes `only:`, `except:` and `parquet_enum: true`. Rails is not a dependency:
|
|
436
|
-
it only calls `columns`, `primary_key` and `defined_enums` on the model. Enum attributes are
|
|
437
|
-
written as their labels, and the writer rejects values outside the enum.
|
|
438
|
-
|
|
439
|
-
| Column | Parquet |
|
|
440
|
-
|---|---|
|
|
441
|
-
| `integer` | `int16`/`int32`/`int64` by SQL type (`smallint`, `integer`, `bigint`...) or limit; `uintN` for `unsigned`; primary keys are always `int64` |
|
|
442
|
-
| `float` | `double` (`float` for Postgres `float4`) |
|
|
443
|
-
| `decimal` | `decimal(precision, scale)`; `decimal(38, 9)` when the column has no precision |
|
|
444
|
-
| `boolean`, `date`, `binary`, `uuid` | same |
|
|
445
|
-
| `json`, `jsonb` | `json` |
|
|
446
|
-
| `datetime`, `timestamp`, `timestamptz` | `timestamp` (microseconds, UTC) |
|
|
447
|
-
| `time` | `time` (microseconds) |
|
|
448
|
-
| `hstore` | `map` of `string` to `string` |
|
|
449
|
-
| enum attributes | `string` (or `enum`) |
|
|
450
|
-
| `string`, `text`, `citext`, anything else | `string` |
|
|
451
|
-
| Postgres array columns | `list` of the element type |
|
|
452
|
-
|
|
453
|
-
Columns declared `NOT NULL` (and primary keys) are required, all others nullable. Column order
|
|
454
|
-
follows `Model.columns`.
|
|
455
|
-
|
|
456
|
-
## Type mapping
|
|
457
|
-
|
|
458
|
-
| Parquet | Ruby |
|
|
459
|
-
|---|---|
|
|
460
|
-
| BOOLEAN | `true`/`false` |
|
|
461
|
-
| INT32/INT64 (incl. signed/unsigned INTEGER) | `Integer` |
|
|
462
|
-
| FLOAT, DOUBLE, FLOAT16 | `Float` |
|
|
463
|
-
| STRING, ENUM, JSON | `String` (UTF-8) |
|
|
464
|
-
| BYTE_ARRAY, FIXED_LEN_BYTE_ARRAY, BSON | `String` (binary) |
|
|
465
|
-
| DATE | `Date` (proleptic Gregorian) |
|
|
466
|
-
| TIMESTAMP, INT96 | `Time` (UTC) |
|
|
467
|
-
| TIME | `Integer` in the column's unit since midnight |
|
|
468
|
-
| DECIMAL | `BigDecimal` |
|
|
469
|
-
| UUID | `String` like `"0f1e2d3c-..."` |
|
|
470
|
-
| struct / list / map | `Hash` / `Array` / `Hash` |
|
|
471
|
-
|
|
472
|
-
## Inspecting files
|
|
473
|
-
|
|
474
|
-
> The inspector and its HTML view are modelled on
|
|
475
|
-
> **[Parquet X-ray](https://huggingface.co/spaces/cfahlgren1/parquet-xray) by cfahlgren1** —
|
|
476
|
-
> the design and the idea are theirs. Go check it out. It is amazing!
|
|
477
|
-
|
|
478
|
-
`Herringbone::Inspector` examines a file using only its footer, page headers, page indexes and
|
|
479
|
-
bloom filter headers. Nothing is decompressed, so it is fast on big files and works for ZSTD and
|
|
480
|
-
Brotli files without the codec gems.
|
|
481
|
-
|
|
482
|
-
```ruby
|
|
483
|
-
File.open("data.parquet", "rb") do |file|
|
|
484
|
-
inspector = Herringbone::Inspector.new(file)
|
|
485
|
-
inspector.summary # size, rows, row groups, codecs, page index / bloom filter presence...
|
|
486
|
-
inspector.row_groups[0].column("name").pages # page headers; also statistics, column_index...
|
|
487
|
-
puts inspector.report # the schema, every column chunk, key/value metadata (Arrow schema decoded)
|
|
488
|
-
inspector.to_h # all of it, JSON-serializable
|
|
489
|
-
inspector.to_html # one self-contained HTML page with a to-scale byte map of the file
|
|
490
|
-
end
|
|
491
|
-
```
|
|
492
|
-
|
|
493
|
-
`inspector.verify_checksums` reads every page body (still without decompressing) to check the page
|
|
494
|
-
CRCs; the results then appear in `summary`, `report`, `to_h` and `to_html`.
|
|
495
|
-
|
|
496
|
-
## Command line
|
|
497
|
-
|
|
498
|
-
```
|
|
499
|
-
bin/herringbone cat FILE [N] # rows as JSON lines
|
|
500
|
-
bin/herringbone inspect FILE [--pages] # text report (--pages lists every page header)
|
|
501
|
-
bin/herringbone inspect FILE --format=json # everything as JSON
|
|
502
|
-
bin/herringbone inspect FILE --format=html # the HTML page, opened in your browser
|
|
503
|
-
bin/herringbone inspect FILE --format=html > out.html # the HTML page, saved
|
|
504
|
-
```
|
|
505
|
-
|
|
506
|
-
Add `--verify-checksums` to any `inspect` form to check page CRCs.
|
|
507
|
-
|
|
508
|
-
`--format` is `text` (the default), `json` or `html`; `--pages` only goes with `text`.
|
|
509
|
-
In a terminal, `--format=html` writes the page to a temp file, prints its path and opens it with `open` on
|
|
510
|
-
macOS, `start` on Windows or `xdg-open` elsewhere. When stdout is redirected or piped it prints the
|
|
511
|
-
page instead.
|
|
512
|
-
|
|
513
|
-
## Supported format features
|
|
514
|
-
|
|
515
|
-
- Encodings (read and write): PLAIN, PLAIN_DICTIONARY/RLE_DICTIONARY, RLE, DELTA_BINARY_PACKED,
|
|
516
|
-
DELTA_LENGTH_BYTE_ARRAY, DELTA_BYTE_ARRAY, BYTE_STREAM_SPLIT; legacy BIT_PACKED levels (read)
|
|
517
|
-
- Data page v1 and v2, dictionary pages, page CRCs (written)
|
|
518
|
-
- Page indexes and split block bloom filters (read and written)
|
|
519
|
-
- Legacy list and map layouts per the Parquet backward-compatibility rules
|
|
520
|
-
- Not supported: encryption, column chunks in external files
|
|
158
|
+
See [Inspecting files](MANUAL.md#inspecting-files) and [Command line](MANUAL.md#command-line).
|
|
521
159
|
|
|
522
160
|
## Development
|
|
523
161
|
|
|
524
|
-
|
|
525
|
-
bundle install
|
|
526
|
-
bundle exec rake test
|
|
527
|
-
HERRINGBONE_PYTHON=/path/to/python-with-pyarrow bundle exec rake test # also run pyarrow interop tests
|
|
528
|
-
```
|
|
529
|
-
|
|
530
|
-
`test/fixtures/parquet-testing` holds files from [apache/parquet-testing](https://github.com/apache/parquet-testing)
|
|
531
|
-
(Apache-2.0); expectations for them were generated with pyarrow, see `test/fixtures/generate_expectations.py`.
|
|
162
|
+
See [Development](MANUAL.md#development) and [CONTRIBUTING](CONTRIBUTING.md).
|