herringbone 0.2.0 → 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.md +56 -4
- data/lib/herringbone/active_record.rb +44 -12
- data/lib/herringbone/bloom_filter.rb +112 -12
- data/lib/herringbone/byte_values.rb +29 -0
- data/lib/herringbone/codecs/lz4.rb +66 -3
- data/lib/herringbone/codecs/snappy.rb +132 -3
- data/lib/herringbone/compression.rb +48 -0
- data/lib/herringbone/encodings/delta.rb +70 -13
- data/lib/herringbone/encodings/plain.rb +22 -0
- data/lib/herringbone/encodings/rle.rb +53 -1
- data/lib/herringbone/format.rb +150 -0
- data/lib/herringbone/inspector.rb +477 -81
- data/lib/herringbone/io_buffer_support.rb +3 -1
- data/lib/herringbone/reader/column_chunk_reader.rb +98 -6
- data/lib/herringbone/reader/column_cursor.rb +46 -2
- data/lib/herringbone/reader/numo.rb +365 -0
- data/lib/herringbone/reader/page_stream.rb +280 -4
- data/lib/herringbone/reader/scan.rb +103 -9
- data/lib/herringbone/reader.rb +308 -17
- data/lib/herringbone/schema.rb +306 -15
- data/lib/herringbone/thrift.rb +155 -4
- data/lib/herringbone/types.rb +176 -41
- data/lib/herringbone/version.rb +2 -1
- data/lib/herringbone/visualizer.rb +51 -11
- data/lib/herringbone/writer.rb +308 -31
- data/lib/herringbone/xxhash.rb +108 -2
- data/lib/herringbone.rb +28 -0
- metadata +3 -2
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: '08770d40f03a2320d9aa60e6fe63fbfe40ea5ddeddc627bff3be2bf46f431795'
|
|
4
|
+
data.tar.gz: b18fef0b86fd290576d254f338da456ad80dc2600c68eb0c903e6085a024ebe2
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 5c941e4b5f31f278ff819d25d464dab518c46e50d10435c360faf2da5994a9cff891754d00a73d8834aa08f2cd00be0387d9f6c8536459389797ba7daa1a82d0
|
|
7
|
+
data.tar.gz: 4532d47095f0832bf4cebb86ef370882dd7287625d80b9b94f6bc49fffb1ad3317075ecf8c45a1fef3fdb2fa19c3a641c0a6c2c4664bf24f5aa664e5b686e03e
|
data/README.md
CHANGED
|
@@ -12,12 +12,16 @@ A pure-Ruby reader and writer for [Apache Parquet](https://parquet.apache.org/)
|
|
|
12
12
|
|
|
13
13
|
```ruby
|
|
14
14
|
gem "herringbone"
|
|
15
|
+
gem "snappy" # optional: native Snappy, 2-3x faster reads and writes of typical files
|
|
15
16
|
gem "zstd-ruby" # optional: ZSTD (faster writes and smaller files than the default Snappy)
|
|
16
17
|
gem "brotli" # optional: Brotli
|
|
17
18
|
gem "xxhash" # optional: faster bloom filters
|
|
19
|
+
gem "numo-narray-alt" # optional: read(as: :numo)
|
|
18
20
|
```
|
|
19
21
|
|
|
20
|
-
The only dependency is `bigdecimal`. Snappy, LZ4 and GZIP always work.
|
|
22
|
+
The only dependency is `bigdecimal`. Snappy, LZ4 and GZIP always work. Snappy, Parquet's most
|
|
23
|
+
common codec, is pure Ruby unless the `snappy` gem is installed (it needs libsnappy or cmake to
|
|
24
|
+
build); Herringbone then uses it automatically. `Herringbone.codecs` lists
|
|
21
25
|
the codecs this process can use, e.g. `[:none, :snappy, :gzip, :lz4, :lz4_hadoop, :zstd]`. Using
|
|
22
26
|
a missing one raises `Herringbone::MissingCodecError` naming the gem to add: a writer raises it
|
|
23
27
|
before writing anything, a reader when it reaches the first such page (the schema and metadata
|
|
@@ -76,6 +80,34 @@ are exact. On a 1M-row file with 20k-row pages, looking up one `id` takes 0.1 s
|
|
|
76
80
|
Filters help most on columns the data is sorted or clustered by. `reader.scan_plan(where: ...)`
|
|
77
81
|
shows which row groups and row ranges a read would touch, without reading them.
|
|
78
82
|
|
|
83
|
+
### Numo arrays
|
|
84
|
+
|
|
85
|
+
`as: :numo` returns (or yields, with `each_batch`) a Hash of column name => Numo array, and takes
|
|
86
|
+
the same `columns:`, `where:`, `from:` and `limit:`. Add `gem "numo-narray-alt"` (or
|
|
87
|
+
`numo-narray`) to your Gemfile: Herringbone requires it on first use and raises
|
|
88
|
+
`Herringbone::UnsupportedError` naming the gem when it is missing.
|
|
89
|
+
|
|
90
|
+
```ruby
|
|
91
|
+
cols = reader.read(as: :numo, columns: %w[id amount]) # { "id" => Numo::Int64, "amount" => Numo::DFloat }
|
|
92
|
+
df = Rover::DataFrame.new(reader.read(as: :numo)) # no conversion needed
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
Flat numeric and boolean columns are decoded from the page bytes without a Ruby object per value:
|
|
96
|
+
two numeric columns of 1M uncompressed rows read in about 25 ms, against 50 ms with `as: :columns`.
|
|
97
|
+
|
|
98
|
+
| Parquet | Numo |
|
|
99
|
+
|---|---|
|
|
100
|
+
| INT32, INT64 (also TIME) | `Int32`, `Int64` |
|
|
101
|
+
| INT(8/16) signed; INT(8/16/32/64) unsigned | `Int8`, `Int16`; `UInt8` … `UInt64` |
|
|
102
|
+
| FLOAT, FLOAT16, DOUBLE | `SFloat`, `SFloat`, `DFloat` (nulls are NaN) |
|
|
103
|
+
| integers with nulls | `DFloat` with NaN (exact up to 2**53) |
|
|
104
|
+
| BOOLEAN | `Bit`; with nulls `RObject` of true/false/nil |
|
|
105
|
+
| list of numbers, every row the same length and no nulls | 2-D `[rows, length]` (e.g. embeddings) |
|
|
106
|
+
| strings, binary, decimals, dates, timestamps, UUIDs, structs, maps, other lists | `RObject` of the values `read` returns |
|
|
107
|
+
|
|
108
|
+
Whether a column has nulls (or a list column is rectangular) is decided from the rows read, so
|
|
109
|
+
with `each_batch` it can differ between batches.
|
|
110
|
+
|
|
79
111
|
## Writing
|
|
80
112
|
|
|
81
113
|
```ruby
|
|
@@ -111,9 +143,10 @@ first 1000 rows unless `schema:` is given; fields declared in a block replace in
|
|
|
111
143
|
`Herringbone.write(io, rows, schema: Herringbone::Schema.infer(rows) { json :payload })`.
|
|
112
144
|
|
|
113
145
|
The writer writes to any IO that responds to `#write` (a `File`, `StringIO`, `Tempfile`, socket or
|
|
114
|
-
pipe), sequentially, and never seeks, rewinds or closes it
|
|
115
|
-
`#abort` is called), no footer is written and what was written
|
|
116
|
-
To replace a file only once it is complete, write to a temporary
|
|
146
|
+
pipe), sequentially, and never seeks, rewinds or closes it (it does switch it to binary mode). If
|
|
147
|
+
the `Writer.open` block raises (or `#abort` is called), no footer is written and what was written
|
|
148
|
+
so far is left for you to discard. To replace a file only once it is complete, write to a temporary
|
|
149
|
+
file and rename it.
|
|
117
150
|
|
|
118
151
|
Column types: `boolean int8 int16 int32 int64 uint8 uint16 uint32 uint64 float double float16
|
|
119
152
|
string binary json bson enum uuid date int96 time timestamp decimal fixed`, plus `struct`, `list`
|
|
@@ -175,6 +208,25 @@ dotted path (`"tags.list.element"`). Hashing is pure Ruby unless the `xxhash` ge
|
|
|
175
208
|
a million rows take about 1.2 s longer to write with a filter on an INT64 column and 2.8 s longer
|
|
176
209
|
with one on a ~22-byte string column, and about 0.5 s longer with `xxhash`.
|
|
177
210
|
|
|
211
|
+
### Writing to S3
|
|
212
|
+
|
|
213
|
+
Parquet keeps its metadata in a footer, so a file can be streamed into an S3 multipart upload
|
|
214
|
+
without a local copy, using `upload_stream` from `aws-sdk-s3`:
|
|
215
|
+
|
|
216
|
+
```ruby
|
|
217
|
+
s3 = Aws::S3::TransferManager.new # aws-sdk-s3 1.197+; before that, Aws::S3::Object#upload_stream
|
|
218
|
+
s3.upload_stream(bucket: "exports", key: "events.parquet", part_size: 16 * 1024 * 1024) do |io|
|
|
219
|
+
Herringbone::Writer.open(io, schema) do |w|
|
|
220
|
+
events.each { |event| w << event }
|
|
221
|
+
end
|
|
222
|
+
end
|
|
223
|
+
s3.upload_stream(bucket: "exports", key: "orders.parquet") { |io| Herringbone.write(io, Order.all) }
|
|
224
|
+
```
|
|
225
|
+
|
|
226
|
+
If the block raises, the SDK aborts the multipart upload and raises `Aws::S3::MultipartUploadError`,
|
|
227
|
+
so no partial object is left. S3 allows at most 10,000 parts, which with the default 5MB parts caps
|
|
228
|
+
the file at about 48GB; raise `part_size:` for bigger files.
|
|
229
|
+
|
|
178
230
|
## ActiveRecord
|
|
179
231
|
|
|
180
232
|
`Herringbone.write` also takes a model or relation, which it reads with `find_each`, using a schema
|
|
@@ -11,13 +11,18 @@ module Herringbone
|
|
|
11
11
|
# ActiveRecord is not required: this only uses what a model class exposes
|
|
12
12
|
# (+columns+, +primary_key+ and, when present, +defined_enums+).
|
|
13
13
|
#
|
|
14
|
-
#
|
|
15
|
-
#
|
|
16
|
-
#
|
|
17
|
-
#
|
|
18
|
-
#
|
|
14
|
+
# Primary key columns are never nullable. Rails enum attributes are string columns holding the
|
|
15
|
+
# labels (the writer rejects values outside the enum; stored values such as 0/1 are written as
|
|
16
|
+
# their labels).
|
|
17
|
+
#
|
|
18
|
+
# @param model [Class] ActiveRecord model class, or anything exposing +columns+ the same way
|
|
19
|
+
# @param only [Array<String, Symbol>, String, Symbol, nil] attribute names to include
|
|
20
|
+
# @param except [Array<String, Symbol>, String, Symbol, nil] attribute names to leave out
|
|
21
|
+
# @param parquet_enum [Boolean] true adds the Parquet ENUM annotation to enum columns, see Builder#enum
|
|
22
|
+
# @return [Schema] schema with one top-level field per selected column, in model column order
|
|
23
|
+
# @raise [ArgumentError] when no columns are left after +only+ / +except+
|
|
19
24
|
def self.from_active_record(model, only: nil, except: nil, parquet_enum: false)
|
|
20
|
-
only
|
|
25
|
+
only &&= Array(only).map(&:to_s)
|
|
21
26
|
except = Array(except).map(&:to_s)
|
|
22
27
|
defined_enums = model.respond_to?(:defined_enums) ? model.defined_enums.to_h { |k, v| [k.to_s, v] } : {}
|
|
23
28
|
primary_keys = Array(model.respond_to?(:primary_key) ? model.primary_key : nil).map(&:to_s)
|
|
@@ -54,9 +59,14 @@ module Herringbone
|
|
|
54
59
|
module ActiveRecordMapping
|
|
55
60
|
module_function
|
|
56
61
|
|
|
62
|
+
# Precision for decimal columns that declare none (Postgres +numeric+ without arguments);
|
|
63
|
+
# 38 digits is the most a 16-byte FIXED_LEN_BYTE_ARRAY decimal holds
|
|
57
64
|
DEFAULT_DECIMAL_PRECISION = 38
|
|
65
|
+
# Scale used together with DEFAULT_DECIMAL_PRECISION when the column declares neither
|
|
58
66
|
DEFAULT_DECIMAL_SCALE = 9
|
|
59
67
|
|
|
68
|
+
# Bit width of named SQL integer types (MySQL, Postgres and SQLite spellings), keyed by the
|
|
69
|
+
# leading word of the lowercased +sql_type+
|
|
60
70
|
INTEGER_SQL_TYPES = {
|
|
61
71
|
"tinyint" => 8, "int1" => 8,
|
|
62
72
|
"smallint" => 16, "int2" => 16, "smallserial" => 16, "serial2" => 16,
|
|
@@ -64,6 +74,16 @@ module Herringbone
|
|
|
64
74
|
"bigint" => 64, "int8" => 64, "bigserial" => 64, "serial8" => 64
|
|
65
75
|
}.freeze
|
|
66
76
|
|
|
77
|
+
# Declares +column+ on +builder+: hstore as map<string, string>, array columns as a list of
|
|
78
|
+
# the element type, everything else as the scalar type picked by #scalar_type.
|
|
79
|
+
#
|
|
80
|
+
# @param builder [Builder] builder to add the field to
|
|
81
|
+
# @param column [ActiveRecord::ConnectionAdapters::Column] column metadata (+name+, +type+, and
|
|
82
|
+
# when available +sql_type+, +array+, +limit+, +precision+, +scale+)
|
|
83
|
+
# @param nullable [Boolean] whether the field may be null
|
|
84
|
+
# @param primary [Boolean] whether the column is (part of) the primary key
|
|
85
|
+
# @return [Node] the added schema node
|
|
86
|
+
# @raise [ArgumentError] when the builder rejects the field (e.g. a duplicate name)
|
|
67
87
|
def add_column(builder, column, nullable, primary)
|
|
68
88
|
sql_type = column.respond_to?(:sql_type) ? column.sql_type.to_s.downcase : ""
|
|
69
89
|
array = (column.respond_to?(:array) && column.array) || sql_type.end_with?("[]")
|
|
@@ -83,10 +103,18 @@ module Herringbone
|
|
|
83
103
|
end
|
|
84
104
|
end
|
|
85
105
|
|
|
106
|
+
# Picks the DSL type for a non-array, non-hstore column, following the table above.
|
|
107
|
+
#
|
|
108
|
+
# @param column [ActiveRecord::ConnectionAdapters::Column] column metadata, read for
|
|
109
|
+
# +limit+, +precision+ and +scale+
|
|
110
|
+
# @param type [Symbol, nil] ActiveRecord's abstract type (+column.type+)
|
|
111
|
+
# @param sql_type [String] lowercased database type with any array suffix removed
|
|
112
|
+
# @param primary [Boolean] whether the column is (part of) the primary key
|
|
113
|
+
# @return [Array(Symbol, Hash{Symbol => Object})] DSL type and its options for Builder#column
|
|
86
114
|
def scalar_type(column, type, sql_type, primary)
|
|
87
115
|
case type
|
|
88
116
|
when :integer, :bigint then [integer_type(column, sql_type, primary), {}]
|
|
89
|
-
when :float then [sql_type == "float4" ? :float : :double, {}]
|
|
117
|
+
when :float then [(sql_type == "float4") ? :float : :double, {}]
|
|
90
118
|
when :decimal, :money
|
|
91
119
|
precision = column.respond_to?(:precision) ? column.precision : nil
|
|
92
120
|
scale = column.respond_to?(:scale) ? column.scale : nil
|
|
@@ -94,12 +122,12 @@ module Herringbone
|
|
|
94
122
|
precision = DEFAULT_DECIMAL_PRECISION
|
|
95
123
|
scale ||= DEFAULT_DECIMAL_SCALE
|
|
96
124
|
end
|
|
97
|
-
[:decimal, {
|
|
125
|
+
[:decimal, {precision: precision, scale: scale || 0}]
|
|
98
126
|
when :boolean then [:boolean, {}]
|
|
99
127
|
when :binary then [:binary, {}]
|
|
100
128
|
when :date then [:date, {}]
|
|
101
|
-
when :datetime, :timestamp, :timestamptz then [:timestamp, {
|
|
102
|
-
when :time then [:time, {
|
|
129
|
+
when :datetime, :timestamp, :timestamptz then [:timestamp, {unit: :micros, utc: true}]
|
|
130
|
+
when :time then [:time, {unit: :micros}]
|
|
103
131
|
when :json, :jsonb then [:json, {}]
|
|
104
132
|
when :uuid then [:uuid, {}]
|
|
105
133
|
else [:string, {}]
|
|
@@ -108,6 +136,11 @@ module Herringbone
|
|
|
108
136
|
|
|
109
137
|
# Named SQL types (smallint, bigint, ...) decide the width; a generic "integer"/"int"
|
|
110
138
|
# uses the column limit in bytes, since e.g. SQLite reports `t.integer limit: 2` as "integer(2)".
|
|
139
|
+
#
|
|
140
|
+
# @param column [ActiveRecord::ConnectionAdapters::Column] column metadata, read for +limit+
|
|
141
|
+
# @param sql_type [String] lowercased database type with any array suffix removed
|
|
142
|
+
# @param primary [Boolean] true forces int64, since row ids can outgrow the declared width
|
|
143
|
+
# @return [Symbol] one of +:int8+ .. +:int64+ or +:uint8+ .. +:uint64+
|
|
111
144
|
def integer_type(column, sql_type, primary)
|
|
112
145
|
bits = INTEGER_SQL_TYPES[sql_type[/\A[a-z0-9]+/]]
|
|
113
146
|
bits ||= case (column.respond_to?(:limit) ? column.limit : nil)
|
|
@@ -120,9 +153,8 @@ module Herringbone
|
|
|
120
153
|
bits = 64 if primary
|
|
121
154
|
unsigned = sql_type.include?("unsigned")
|
|
122
155
|
return :"uint#{bits}" if unsigned
|
|
123
|
-
{
|
|
156
|
+
{8 => :int8, 16 => :int16, 32 => :int32, 64 => :int64}.fetch(bits)
|
|
124
157
|
end
|
|
125
158
|
end
|
|
126
159
|
end
|
|
127
160
|
end
|
|
128
|
-
|
|
@@ -13,19 +13,27 @@ module Herringbone
|
|
|
13
13
|
# String hashes the same bytes as the stored value. Nulls are never in a bloom filter.
|
|
14
14
|
# The writer builds them (bloom_filters: option) and reads with where: consult them.
|
|
15
15
|
class BloomFilter
|
|
16
|
+
# The spec's eight salt constants, one per word of a block
|
|
16
17
|
SALT = [0x47b6137b, 0x44974d91, 0x8824ad5b, 0xa2b7289d, 0x705495c7, 0x2df1424b, 0x9efc4947, 0x5c6bfb31].freeze
|
|
17
18
|
S0, S1, S2, S3, S4, S5, S6, S7 = SALT
|
|
18
19
|
# Low 16 bits of the salts: (lo * salt) mod 2**32 is computed as
|
|
19
20
|
# (lo & 0xFFFF) * salt + ((lo >> 16) * (salt & 0xFFFF) << 16), which stays a Fixnum
|
|
20
21
|
L0, L1, L2, L3, L4, L5, L6, L7 = SALT.map { |s| s & 0xFFFF }
|
|
22
|
+
# Size of a block: eight 32-bit words
|
|
21
23
|
BLOCK_BYTES = 32
|
|
24
|
+
# Smallest bitset: a single block
|
|
22
25
|
MIN_BYTES = 32
|
|
26
|
+
# Largest bitset accepted when reading or building (128 MiB, the cap parquet-mr uses)
|
|
23
27
|
MAX_BYTES = 128 * 1024 * 1024
|
|
24
28
|
# Default cap for filters sized by the writer (as in parquet-mr)
|
|
25
29
|
DEFAULT_MAX_BYTES = 1024 * 1024
|
|
30
|
+
# Default false positive probability for filters sized by the writer
|
|
26
31
|
DEFAULT_FPP = 0.01
|
|
32
|
+
# 32-bit mask
|
|
27
33
|
M32 = 0xFFFF_FFFF
|
|
34
|
+
# 64-bit mask
|
|
28
35
|
M64 = 0xFFFF_FFFF_FFFF_FFFF
|
|
36
|
+
# Shorthand for the physical type constants
|
|
29
37
|
T = Format::Type
|
|
30
38
|
# Physical types a bloom filter can be built for (the spec does not define BOOLEAN hashing)
|
|
31
39
|
TYPES = [T::INT32, T::INT64, T::INT96, T::FLOAT, T::DOUBLE, T::BYTE_ARRAY, T::FIXED_LEN_BYTE_ARRAY].freeze
|
|
@@ -33,10 +41,17 @@ module Herringbone
|
|
|
33
41
|
# Bitset size in bytes for +ndv+ distinct values at false positive probability +fpp+, per the
|
|
34
42
|
# spec's formula (m = -8 * ndv / ln(1 - fpp ** (1/8)) bits), rounded up to a power of two
|
|
35
43
|
# and clamped to MIN_BYTES..max_bytes
|
|
44
|
+
#
|
|
45
|
+
# @param ndv [Integer] expected number of distinct values (values below 1 count as 1)
|
|
46
|
+
# @param fpp [Float] false positive probability, strictly between 0 and 1
|
|
47
|
+
# @param max_bytes [Integer] upper bound, clamped to MIN_BYTES..MAX_BYTES and rounded down to a
|
|
48
|
+
# power of two
|
|
49
|
+
# @return [Integer] bitset size in bytes, a power of two
|
|
50
|
+
# @raise [ArgumentError] when +fpp+ is not between 0 and 1
|
|
36
51
|
def self.optimal_num_bytes(ndv, fpp = DEFAULT_FPP, max_bytes: DEFAULT_MAX_BYTES)
|
|
37
52
|
fpp = Float(fpp)
|
|
38
53
|
raise ArgumentError, "fpp must be between 0 and 1, got #{fpp}" unless fpp > 0 && fpp < 1
|
|
39
|
-
max_bytes =
|
|
54
|
+
max_bytes = Integer(max_bytes).clamp(MIN_BYTES, MAX_BYTES)
|
|
40
55
|
max_bytes = 1 << (max_bytes.bit_length - 1) # a power of two
|
|
41
56
|
ndv = [Integer(ndv), 1].max
|
|
42
57
|
bits = -8.0 * ndv / Math.log(1 - fpp**(1.0 / 8))
|
|
@@ -47,6 +62,12 @@ module Herringbone
|
|
|
47
62
|
end
|
|
48
63
|
|
|
49
64
|
# XXH64 of the PLAIN encoding of a physical value of +type+ (as the column encoders produce it)
|
|
65
|
+
#
|
|
66
|
+
# @param value [Integer, Float, String, Array<Integer>] the physical value; for INT96 the
|
|
67
|
+
# +[nanos_of_day, julian_day]+ pair the encoder produces
|
|
68
|
+
# @param type [Integer] Format::Type physical type
|
|
69
|
+
# @return [Integer] the unsigned 64-bit hash
|
|
70
|
+
# @raise [UnsupportedError] for BOOLEAN (or any type not in TYPES)
|
|
50
71
|
def self.hash_physical(value, type)
|
|
51
72
|
case type
|
|
52
73
|
when T::INT32 then XXHash.xxh64_u32(value)
|
|
@@ -62,6 +83,12 @@ module Herringbone
|
|
|
62
83
|
# Hashes of many physical values of +type+, converting them in bulk. With +distinct+, each
|
|
63
84
|
# distinct physical value is hashed once (floats are compared by their bytes, so -0.0 and 0.0,
|
|
64
85
|
# or NaNs with different payloads, stay apart as they hash differently).
|
|
86
|
+
#
|
|
87
|
+
# @param values [Array] physical values, as for .hash_physical
|
|
88
|
+
# @param type [Integer] Format::Type physical type
|
|
89
|
+
# @param distinct [Boolean] hash each distinct value once (the result is then shorter)
|
|
90
|
+
# @return [Array<Integer>] the unsigned 64-bit hashes
|
|
91
|
+
# @raise [UnsupportedError] for BOOLEAN (or any type not in TYPES)
|
|
65
92
|
def self.hash_physical_all(values, type, distinct: false)
|
|
66
93
|
case type
|
|
67
94
|
when T::INT32 then XXHash.xxh64_u32_all(distinct ? values.uniq : values)
|
|
@@ -80,6 +107,13 @@ module Herringbone
|
|
|
80
107
|
|
|
81
108
|
# Reads a filter (header and bitset) from +buf+ at +pos+. Returns nil for algorithms, hashes
|
|
82
109
|
# or compressions this implementation does not know.
|
|
110
|
+
#
|
|
111
|
+
# @param buf [String] bytes holding the BloomFilterHeader followed by the bitset
|
|
112
|
+
# @param pos [Integer] byte offset of the header in +buf+
|
|
113
|
+
# @param column [Schema::Column, nil] column the filter belongs to, for converting values
|
|
114
|
+
# @return [BloomFilter, nil] the filter, or nil when it is of an unsupported kind
|
|
115
|
+
# @raise [FormatError] when the bitset is shorter than the header announces
|
|
116
|
+
# @raise [Thrift::Error] when the header cannot be decoded
|
|
83
117
|
def self.decode(buf, pos = 0, column: nil)
|
|
84
118
|
reader = Thrift::Reader.new(buf, pos)
|
|
85
119
|
header = reader.read_struct(Format::BloomFilterHeader)
|
|
@@ -89,48 +123,78 @@ module Herringbone
|
|
|
89
123
|
new(bitset: bitset, column: column)
|
|
90
124
|
end
|
|
91
125
|
|
|
126
|
+
# Whether a header describes a filter this implementation can read: split block algorithm,
|
|
127
|
+
# XXH64, uncompressed, and a size that is a multiple of 32 bytes up to MAX_BYTES
|
|
128
|
+
#
|
|
129
|
+
# @param header [Format::BloomFilterHeader] decoded header
|
|
130
|
+
# @return [Boolean] true when supported
|
|
92
131
|
def self.supported_header?(header)
|
|
93
132
|
n = header.num_bytes
|
|
94
133
|
n.is_a?(Integer) && n >= BLOCK_BYTES && (n % BLOCK_BYTES).zero? && n <= MAX_BYTES &&
|
|
95
|
-
header.algorithm
|
|
96
|
-
(header.
|
|
134
|
+
!(header.algorithm && header.algorithm.block).nil? &&
|
|
135
|
+
!(header.hash_function && header.hash_function.xxhash).nil? &&
|
|
136
|
+
(header.compression.nil? || !header.compression.uncompressed.nil?)
|
|
97
137
|
end
|
|
98
138
|
|
|
139
|
+
# @return [Schema::Column, nil] column whose encoder converts values, nil for raw Strings
|
|
99
140
|
attr_reader :column
|
|
100
141
|
|
|
101
142
|
# A filter of +num_bytes+ (a multiple of 32, normally a power of two), or one using an existing
|
|
102
143
|
# +bitset+ String. With a +column+ (a Schema::Column), values are Ruby values converted with
|
|
103
144
|
# the column's encoder; without one, values must be Strings and their bytes are hashed.
|
|
145
|
+
#
|
|
146
|
+
# @param num_bytes [Integer, nil] bitset size for an empty filter (MIN_BYTES when nil); ignored
|
|
147
|
+
# with +bitset+
|
|
148
|
+
# @param bitset [String, nil] existing bitset (little-endian 32-bit words)
|
|
149
|
+
# @param column [Schema::Column, nil] column the filter is for
|
|
150
|
+
# @raise [ArgumentError] when the size (or +bitset+'s size) is not a multiple of 32 bytes or
|
|
151
|
+
# exceeds MAX_BYTES
|
|
152
|
+
# @raise [UnsupportedError] when the column's physical type cannot have a bloom filter
|
|
104
153
|
def initialize(num_bytes = nil, bitset: nil, column: nil)
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
num_bytes = Integer(num_bytes || MIN_BYTES)
|
|
109
|
-
unless num_bytes >= BLOCK_BYTES && (num_bytes % BLOCK_BYTES).zero? && num_bytes <= MAX_BYTES
|
|
110
|
-
raise ArgumentError, "Bloom filter size must be a multiple of 32 bytes up to 128MB, got #{num_bytes}"
|
|
111
|
-
end
|
|
112
|
-
@words = Array.new(num_bytes / 4, 0)
|
|
154
|
+
num_bytes = bitset ? bitset.bytesize : Integer(num_bytes || MIN_BYTES)
|
|
155
|
+
unless num_bytes >= BLOCK_BYTES && (num_bytes % BLOCK_BYTES).zero? && num_bytes <= MAX_BYTES
|
|
156
|
+
raise ArgumentError, "Bloom filter size must be a multiple of 32 bytes up to 128MB, got #{num_bytes}"
|
|
113
157
|
end
|
|
158
|
+
@words = bitset ? bitset.unpack("V*") : Array.new(num_bytes / 4, 0)
|
|
114
159
|
@num_blocks = @words.size / 8
|
|
115
|
-
raise ArgumentError, "Bloom filter bitset must be a multiple of 32 bytes" if @num_blocks.zero? || @words.size % 8 != 0
|
|
116
160
|
@column = column
|
|
117
161
|
if column && !TYPES.include?(column.type)
|
|
118
162
|
raise UnsupportedError, "Bloom filters are not supported for #{T::NAMES[column.type]} column #{column.dotted_path}"
|
|
119
163
|
end
|
|
120
164
|
end
|
|
121
165
|
|
|
166
|
+
# @return [Integer] bitset size in bytes
|
|
122
167
|
def num_bytes = @words.size * 4
|
|
123
168
|
|
|
169
|
+
# Adds a value to the filter
|
|
170
|
+
#
|
|
171
|
+
# @param value [Object] Ruby value, converted with the column's encoder (a String without a
|
|
172
|
+
# column)
|
|
173
|
+
# @return [BloomFilter] self
|
|
174
|
+
# @raise [ArgumentError] for nil or a value the column cannot encode
|
|
124
175
|
def insert(value)
|
|
125
176
|
insert_hash(hash_of(value))
|
|
126
177
|
self
|
|
127
178
|
end
|
|
128
179
|
|
|
180
|
+
# Whether the value may be in the filter. Never false for an inserted value; true for a
|
|
181
|
+
# value that was not inserted with about the false positive probability the filter was sized for.
|
|
182
|
+
#
|
|
183
|
+
# @param value [Object] Ruby value, converted with the column's encoder (a String without a
|
|
184
|
+
# column)
|
|
185
|
+
# @return [Boolean] false when the value is definitely absent
|
|
186
|
+
# @raise [ArgumentError] for nil or a value the column cannot encode
|
|
129
187
|
def might_contain?(value)
|
|
130
188
|
might_contain_hash?(hash_of(value))
|
|
131
189
|
end
|
|
132
190
|
|
|
133
191
|
# XXH64 hash the filter uses for a Ruby +value+
|
|
192
|
+
#
|
|
193
|
+
# @param value [Object] Ruby value, converted with the column's encoder (a String without a
|
|
194
|
+
# column)
|
|
195
|
+
# @return [Integer] the unsigned 64-bit hash
|
|
196
|
+
# @raise [ArgumentError] for nil, a value the column cannot encode, or (without a column) a
|
|
197
|
+
# non-String
|
|
134
198
|
def hash_of(value)
|
|
135
199
|
raise ArgumentError, "Nulls are not recorded in bloom filters" if value.nil?
|
|
136
200
|
if @column
|
|
@@ -146,6 +210,11 @@ module Herringbone
|
|
|
146
210
|
end
|
|
147
211
|
end
|
|
148
212
|
|
|
213
|
+
# Sets the bits for a hash: the upper 32 bits pick the block, the lower 32 bits (multiplied
|
|
214
|
+
# by each salt) one bit in each of its eight words
|
|
215
|
+
#
|
|
216
|
+
# @param h [Integer] unsigned 64-bit XXH64 hash
|
|
217
|
+
# @return [BloomFilter] self
|
|
149
218
|
def insert_hash(h)
|
|
150
219
|
i = (((h >> 32) * @num_blocks) >> 32) << 3
|
|
151
220
|
x0 = h & 0xFFFF
|
|
@@ -162,10 +231,14 @@ module Herringbone
|
|
|
162
231
|
self
|
|
163
232
|
end
|
|
164
233
|
|
|
234
|
+
# Single-bit masks, +BITS[i] == 1 << i+
|
|
165
235
|
BITS = Array.new(32) { |i| 1 << i }.freeze
|
|
166
236
|
|
|
167
237
|
# Inserts many hashes at once: faster than insert_hash, as the hashes (mostly Bignums) are
|
|
168
238
|
# split into 32-bit halves in bulk and the rest is Fixnum arithmetic
|
|
239
|
+
#
|
|
240
|
+
# @param hashes [Array<Integer>] unsigned 64-bit XXH64 hashes
|
|
241
|
+
# @return [BloomFilter] self
|
|
169
242
|
def insert_hashes(hashes)
|
|
170
243
|
halves = hashes.pack("Q<*").unpack("V*")
|
|
171
244
|
w = @words
|
|
@@ -191,6 +264,10 @@ module Herringbone
|
|
|
191
264
|
self
|
|
192
265
|
end
|
|
193
266
|
|
|
267
|
+
# Whether all the bits for a hash are set (see #insert_hash)
|
|
268
|
+
#
|
|
269
|
+
# @param h [Integer] unsigned 64-bit XXH64 hash
|
|
270
|
+
# @return [Boolean] false when the hashed value is definitely absent
|
|
194
271
|
def might_contain_hash?(h)
|
|
195
272
|
i = (((h >> 32) * @num_blocks) >> 32) << 3
|
|
196
273
|
x0 = h & 0xFFFF
|
|
@@ -207,19 +284,27 @@ module Herringbone
|
|
|
207
284
|
end
|
|
208
285
|
|
|
209
286
|
# The raw bitset (little-endian 32-bit words)
|
|
287
|
+
#
|
|
288
|
+
# @return [String] binary String of #num_bytes bytes
|
|
210
289
|
def bitset = @words.pack("V*")
|
|
211
290
|
|
|
212
291
|
# Header and bitset, as stored in a Parquet file
|
|
292
|
+
#
|
|
293
|
+
# @return [String] Thrift-encoded BloomFilterHeader followed by the bitset (binary)
|
|
213
294
|
def encode
|
|
214
295
|
header.encode << bitset
|
|
215
296
|
end
|
|
216
297
|
|
|
298
|
+
# @return [String] size and, when set, the column's dotted path
|
|
217
299
|
def inspect
|
|
218
300
|
"#<#{self.class.name} #{num_bytes} bytes#{" for #{@column.dotted_path}" if @column}>"
|
|
219
301
|
end
|
|
220
302
|
|
|
221
303
|
private
|
|
222
304
|
|
|
305
|
+
# Header for this filter: split block algorithm, XXH64 hash, uncompressed
|
|
306
|
+
#
|
|
307
|
+
# @return [Format::BloomFilterHeader] the header
|
|
223
308
|
def header
|
|
224
309
|
Format::BloomFilterHeader.new(
|
|
225
310
|
num_bytes: num_bytes,
|
|
@@ -234,6 +319,13 @@ module Herringbone
|
|
|
234
319
|
# Internal (used by reads with where:): the bloom filter of a column chunk. +column+ is a
|
|
235
320
|
# dotted path ("a.b"), an Array path or a Schema::Column. Returns a BloomFilter, or nil when
|
|
236
321
|
# the chunk has none (or one of an unknown kind).
|
|
322
|
+
#
|
|
323
|
+
# @param row_group_index [Integer] index of the row group
|
|
324
|
+
# @param column [String, Array<String, Symbol>, Symbol, Schema::Column] the leaf column
|
|
325
|
+
# @return [BloomFilter, nil] the filter, or nil when there is none or it is unsupported
|
|
326
|
+
# @raise [IndexError] when there is no such row group
|
|
327
|
+
# @raise [ArgumentError] when there is no such column
|
|
328
|
+
# @raise [FormatError] when the filter is truncated or its header cannot be decoded
|
|
237
329
|
def bloom_filter(row_group_index, column)
|
|
238
330
|
col = bloom_filter_column(column)
|
|
239
331
|
rg = row_groups.fetch(row_group_index) { raise IndexError, "No row group #{row_group_index}" }
|
|
@@ -256,10 +348,18 @@ module Herringbone
|
|
|
256
348
|
bitset = @io.read(header.num_bytes)
|
|
257
349
|
raise FormatError, "Truncated bloom filter" if bitset.nil? || bitset.bytesize != header.num_bytes
|
|
258
350
|
BloomFilter.new(bitset: bitset, column: col)
|
|
351
|
+
rescue Thrift::Error => e
|
|
352
|
+
raise FormatError, "Corrupt bloom filter header for #{col.dotted_path}: #{e.message}"
|
|
259
353
|
end
|
|
260
354
|
|
|
261
355
|
private
|
|
262
356
|
|
|
357
|
+
# Resolves the +column+ argument of #bloom_filter to a leaf column
|
|
358
|
+
#
|
|
359
|
+
# @param column [String, Array<String, Symbol>, Symbol, Schema::Column] dotted path, path
|
|
360
|
+
# Array or column
|
|
361
|
+
# @return [Schema::Column] the column
|
|
362
|
+
# @raise [ArgumentError] when the schema has no such column
|
|
263
363
|
def bloom_filter_column(column)
|
|
264
364
|
return column if column.is_a?(Schema::Column)
|
|
265
365
|
col = schema.column(column.is_a?(Array) ? column.map(&:to_s) : column.to_s)
|
|
@@ -17,8 +17,11 @@ module Herringbone
|
|
|
17
17
|
# After this many values, give up on the dictionary if more than half of them are distinct
|
|
18
18
|
CARDINALITY_CHECK_AT = 4096
|
|
19
19
|
|
|
20
|
+
# Whether String#append_as_bytes (Ruby 3.4+) is available for appending without re-encoding
|
|
20
21
|
APPEND_AS_BYTES = "".respond_to?(:append_as_bytes)
|
|
21
22
|
|
|
23
|
+
# @param width [Integer, nil] byte length of each value for FIXED_LEN_BYTE_ARRAY, nil for BYTE_ARRAY
|
|
24
|
+
# @param dictionary [Boolean] start in dictionary mode; false appends raw bytes from the start
|
|
22
25
|
def initialize(width: nil, dictionary: true)
|
|
23
26
|
@width = width
|
|
24
27
|
@bytes = String.new(encoding: Encoding::BINARY)
|
|
@@ -31,14 +34,21 @@ module Herringbone
|
|
|
31
34
|
end
|
|
32
35
|
end
|
|
33
36
|
|
|
37
|
+
# @return [Integer] number of values held
|
|
34
38
|
def size
|
|
35
39
|
@indices ? @indices.size : @count
|
|
36
40
|
end
|
|
37
41
|
|
|
42
|
+
# @return [Boolean] true when no values are held
|
|
38
43
|
def empty? = size.zero?
|
|
39
44
|
|
|
45
|
+
# @return [Boolean] true while still dictionary-encoding (not yet switched to raw bytes)
|
|
40
46
|
def dictionary? = !@indices.nil?
|
|
41
47
|
|
|
48
|
+
# Appends a value, switching to raw bytes when the dictionary limits are exceeded.
|
|
49
|
+
# @param value [String] bytes of the value; for FIXED_LEN_BYTE_ARRAY it must be +width+ bytes
|
|
50
|
+
# long (not checked here)
|
|
51
|
+
# @return [ByteValues] self
|
|
42
52
|
def <<(value)
|
|
43
53
|
if @indices
|
|
44
54
|
index = @dictionary[value]
|
|
@@ -61,16 +71,20 @@ module Herringbone
|
|
|
61
71
|
end
|
|
62
72
|
|
|
63
73
|
# Removes the last value
|
|
74
|
+
# @return [void]
|
|
64
75
|
def pop
|
|
65
76
|
truncate(size - 1) unless size.zero?
|
|
66
77
|
end
|
|
67
78
|
|
|
68
79
|
# Supports the `slice!(n..)` form used to roll back a failed row
|
|
80
|
+
# @param range [Range] endless range; values from +range.begin+ on are dropped
|
|
81
|
+
# @return [void]
|
|
69
82
|
def slice!(range)
|
|
70
83
|
truncate(range.begin)
|
|
71
84
|
end
|
|
72
85
|
|
|
73
86
|
# Approximate memory held, used to size row groups
|
|
87
|
+
# @return [Integer] estimated bytes
|
|
74
88
|
def memory_bytes
|
|
75
89
|
if @indices
|
|
76
90
|
# Each distinct value is a String object plus a Hash entry
|
|
@@ -82,6 +96,11 @@ module Herringbone
|
|
|
82
96
|
|
|
83
97
|
# Returns [:dictionary, values, indices] when the column is worth dictionary-encoding,
|
|
84
98
|
# otherwise [:plain, values]
|
|
99
|
+
#
|
|
100
|
+
# Dictionary encoding is kept unless there are more than 16 values and more than about half
|
|
101
|
+
# of them are distinct.
|
|
102
|
+
# @return [Array(Symbol, Array<String>, Array<Integer>), Array(Symbol, Array<String>)]
|
|
103
|
+
# distinct values and an index per value, or all values in order
|
|
85
104
|
def materialize
|
|
86
105
|
if @indices
|
|
87
106
|
keys = @dictionary.keys
|
|
@@ -93,6 +112,9 @@ module Herringbone
|
|
|
93
112
|
|
|
94
113
|
private
|
|
95
114
|
|
|
115
|
+
# Appends +value+ in raw-bytes mode, recording its length for BYTE_ARRAY.
|
|
116
|
+
# @param value [String] bytes of the value, in any encoding
|
|
117
|
+
# @return [ByteValues] self
|
|
96
118
|
def append_bytes(value)
|
|
97
119
|
if value.encoding == Encoding::BINARY || value.ascii_only?
|
|
98
120
|
@bytes << value
|
|
@@ -106,6 +128,8 @@ module Herringbone
|
|
|
106
128
|
self
|
|
107
129
|
end
|
|
108
130
|
|
|
131
|
+
# Leaves dictionary mode, re-appending every value held so far as raw bytes.
|
|
132
|
+
# @return [void]
|
|
109
133
|
def switch_to_bytes
|
|
110
134
|
keys = @dictionary.keys
|
|
111
135
|
indices = @indices
|
|
@@ -113,6 +137,9 @@ module Herringbone
|
|
|
113
137
|
indices.each { |i| append_bytes(keys[i]) }
|
|
114
138
|
end
|
|
115
139
|
|
|
140
|
+
# Drops all values after the first +n+.
|
|
141
|
+
# @param n [Integer] number of values to keep
|
|
142
|
+
# @return [void]
|
|
116
143
|
def truncate(n)
|
|
117
144
|
if @indices
|
|
118
145
|
@indices.slice!(n..)
|
|
@@ -128,6 +155,8 @@ module Herringbone
|
|
|
128
155
|
end
|
|
129
156
|
end
|
|
130
157
|
|
|
158
|
+
# Splits the raw bytes back into one binary String per value.
|
|
159
|
+
# @return [Array<String>]
|
|
131
160
|
def strings
|
|
132
161
|
if @lengths
|
|
133
162
|
pos = 0
|