herringbone 0.2.0 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: b6e9b4e12f2b83688e6ea7be2752dd389cd1a338489834160efb0df5f796916e
4
- data.tar.gz: c69303ac028259092e2f22c0f5d635f0e428ae1950c1b4b42d4a6379fd6a5a87
3
+ metadata.gz: '08770d40f03a2320d9aa60e6fe63fbfe40ea5ddeddc627bff3be2bf46f431795'
4
+ data.tar.gz: b18fef0b86fd290576d254f338da456ad80dc2600c68eb0c903e6085a024ebe2
5
5
  SHA512:
6
- metadata.gz: 00a7d9246cb73211dcd5c46a5b0dab15b20b0fe1dc0026a2f02dcf3a9e13a11c91b8632bf1891fdd2f5a8deccf42f1c46918ee26ed9bf881208338f731cd12b6
7
- data.tar.gz: 91edc03478f1d9aa24f4d40527b8d2ac8cbed53047e3870e2b17d52ba252bb29a354035f5f072105adf0122bcf93e24c424747552423f9f7ba345fedb98b4219
6
+ metadata.gz: 5c941e4b5f31f278ff819d25d464dab518c46e50d10435c360faf2da5994a9cff891754d00a73d8834aa08f2cd00be0387d9f6c8536459389797ba7daa1a82d0
7
+ data.tar.gz: 4532d47095f0832bf4cebb86ef370882dd7287625d80b9b94f6bc49fffb1ad3317075ecf8c45a1fef3fdb2fa19c3a641c0a6c2c4664bf24f5aa664e5b686e03e
data/README.md CHANGED
@@ -12,12 +12,16 @@ A pure-Ruby reader and writer for [Apache Parquet](https://parquet.apache.org/)
12
12
 
13
13
  ```ruby
14
14
  gem "herringbone"
15
+ gem "snappy" # optional: native Snappy, 2-3x faster reads and writes of typical files
15
16
  gem "zstd-ruby" # optional: ZSTD (faster writes and smaller files than the default Snappy)
16
17
  gem "brotli" # optional: Brotli
17
18
  gem "xxhash" # optional: faster bloom filters
19
+ gem "numo-narray-alt" # optional: read(as: :numo)
18
20
  ```
19
21
 
20
- The only dependency is `bigdecimal`. Snappy, LZ4 and GZIP always work. `Herringbone.codecs` lists
22
+ The only dependency is `bigdecimal`. Snappy, LZ4 and GZIP always work. Snappy, Parquet's most
23
+ common codec, is pure Ruby unless the `snappy` gem is installed (it needs libsnappy or cmake to
24
+ build); Herringbone then uses it automatically. `Herringbone.codecs` lists
21
25
  the codecs this process can use, e.g. `[:none, :snappy, :gzip, :lz4, :lz4_hadoop, :zstd]`. Using
22
26
  a missing one raises `Herringbone::MissingCodecError` naming the gem to add: a writer raises it
23
27
  before writing anything, a reader when it reaches the first such page (the schema and metadata
@@ -76,6 +80,34 @@ are exact. On a 1M-row file with 20k-row pages, looking up one `id` takes 0.1 s
76
80
  Filters help most on columns the data is sorted or clustered by. `reader.scan_plan(where: ...)`
77
81
  shows which row groups and row ranges a read would touch, without reading them.
78
82
 
83
+ ### Numo arrays
84
+
85
+ `as: :numo` returns (or yields, with `each_batch`) a Hash of column name => Numo array, and takes
86
+ the same `columns:`, `where:`, `from:` and `limit:`. Add `gem "numo-narray-alt"` (or
87
+ `numo-narray`) to your Gemfile: Herringbone requires it on first use and raises
88
+ `Herringbone::UnsupportedError` naming the gem when it is missing.
89
+
90
+ ```ruby
91
+ cols = reader.read(as: :numo, columns: %w[id amount]) # { "id" => Numo::Int64, "amount" => Numo::DFloat }
92
+ df = Rover::DataFrame.new(reader.read(as: :numo)) # no conversion needed
93
+ ```
94
+
95
+ Flat numeric and boolean columns are decoded from the page bytes without a Ruby object per value:
96
+ two numeric columns of 1M uncompressed rows read in about 25 ms, against 50 ms with `as: :columns`.
97
+
98
+ | Parquet | Numo |
99
+ |---|---|
100
+ | INT32, INT64 (also TIME) | `Int32`, `Int64` |
101
+ | INT(8/16) signed; INT(8/16/32/64) unsigned | `Int8`, `Int16`; `UInt8` … `UInt64` |
102
+ | FLOAT, FLOAT16, DOUBLE | `SFloat`, `SFloat`, `DFloat` (nulls are NaN) |
103
+ | integers with nulls | `DFloat` with NaN (exact up to 2**53) |
104
+ | BOOLEAN | `Bit`; with nulls `RObject` of true/false/nil |
105
+ | list of numbers, every row the same length and no nulls | 2-D `[rows, length]` (e.g. embeddings) |
106
+ | strings, binary, decimals, dates, timestamps, UUIDs, structs, maps, other lists | `RObject` of the values `read` returns |
107
+
108
+ Whether a column has nulls (or a list column is rectangular) is decided from the rows read, so
109
+ with `each_batch` it can differ between batches.
110
+
79
111
  ## Writing
80
112
 
81
113
  ```ruby
@@ -111,9 +143,10 @@ first 1000 rows unless `schema:` is given; fields declared in a block replace in
111
143
  `Herringbone.write(io, rows, schema: Herringbone::Schema.infer(rows) { json :payload })`.
112
144
 
113
145
  The writer writes to any IO that responds to `#write` (a `File`, `StringIO`, `Tempfile`, socket or
114
- pipe), sequentially, and never seeks, rewinds or closes it. If the `Writer.open` block raises (or
115
- `#abort` is called), no footer is written and what was written so far is left for you to discard.
116
- To replace a file only once it is complete, write to a temporary file and rename it.
146
+ pipe), sequentially, and never seeks, rewinds or closes it (it does switch it to binary mode). If
147
+ the `Writer.open` block raises (or `#abort` is called), no footer is written and what was written
148
+ so far is left for you to discard. To replace a file only once it is complete, write to a temporary
149
+ file and rename it.
117
150
 
118
151
  Column types: `boolean int8 int16 int32 int64 uint8 uint16 uint32 uint64 float double float16
119
152
  string binary json bson enum uuid date int96 time timestamp decimal fixed`, plus `struct`, `list`
@@ -175,6 +208,25 @@ dotted path (`"tags.list.element"`). Hashing is pure Ruby unless the `xxhash` ge
175
208
  a million rows take about 1.2 s longer to write with a filter on an INT64 column and 2.8 s longer
176
209
  with one on a ~22-byte string column, and about 0.5 s longer with `xxhash`.
177
210
 
211
+ ### Writing to S3
212
+
213
+ Parquet keeps its metadata in a footer, so a file can be streamed into an S3 multipart upload
214
+ without a local copy, using `upload_stream` from `aws-sdk-s3`:
215
+
216
+ ```ruby
217
+ s3 = Aws::S3::TransferManager.new # aws-sdk-s3 1.197+; before that, Aws::S3::Object#upload_stream
218
+ s3.upload_stream(bucket: "exports", key: "events.parquet", part_size: 16 * 1024 * 1024) do |io|
219
+ Herringbone::Writer.open(io, schema) do |w|
220
+ events.each { |event| w << event }
221
+ end
222
+ end
223
+ s3.upload_stream(bucket: "exports", key: "orders.parquet") { |io| Herringbone.write(io, Order.all) }
224
+ ```
225
+
226
+ If the block raises, the SDK aborts the multipart upload and raises `Aws::S3::MultipartUploadError`,
227
+ so no partial object is left. S3 allows at most 10,000 parts, which with the default 5MB parts caps
228
+ the file at about 48GB; raise `part_size:` for bigger files.
229
+
178
230
  ## ActiveRecord
179
231
 
180
232
  `Herringbone.write` also takes a model or relation, which it reads with `find_each`, using a schema
@@ -11,13 +11,18 @@ module Herringbone
11
11
  # ActiveRecord is not required: this only uses what a model class exposes
12
12
  # (+columns+, +primary_key+ and, when present, +defined_enums+).
13
13
  #
14
- # only: attribute names to include (Strings or Symbols)
15
- # except: attribute names to leave out
16
- # parquet_enum: Rails enum attributes are string columns holding the labels (the writer
17
- # rejects values outside the enum; stored values such as 0/1 are written as
18
- # their labels). true adds the Parquet ENUM annotation, see Builder#enum.
14
+ # Primary key columns are never nullable. Rails enum attributes are string columns holding the
15
+ # labels (the writer rejects values outside the enum; stored values such as 0/1 are written as
16
+ # their labels).
17
+ #
18
+ # @param model [Class] ActiveRecord model class, or anything exposing +columns+ the same way
19
+ # @param only [Array<String, Symbol>, String, Symbol, nil] attribute names to include
20
+ # @param except [Array<String, Symbol>, String, Symbol, nil] attribute names to leave out
21
+ # @param parquet_enum [Boolean] true adds the Parquet ENUM annotation to enum columns, see Builder#enum
22
+ # @return [Schema] schema with one top-level field per selected column, in model column order
23
+ # @raise [ArgumentError] when no columns are left after +only+ / +except+
19
24
  def self.from_active_record(model, only: nil, except: nil, parquet_enum: false)
20
- only = only && Array(only).map(&:to_s)
25
+ only &&= Array(only).map(&:to_s)
21
26
  except = Array(except).map(&:to_s)
22
27
  defined_enums = model.respond_to?(:defined_enums) ? model.defined_enums.to_h { |k, v| [k.to_s, v] } : {}
23
28
  primary_keys = Array(model.respond_to?(:primary_key) ? model.primary_key : nil).map(&:to_s)
@@ -54,9 +59,14 @@ module Herringbone
54
59
  module ActiveRecordMapping
55
60
  module_function
56
61
 
62
+ # Precision for decimal columns that declare none (Postgres +numeric+ without arguments);
63
+ # 38 digits is the most a 16-byte FIXED_LEN_BYTE_ARRAY decimal holds
57
64
  DEFAULT_DECIMAL_PRECISION = 38
65
+ # Scale used together with DEFAULT_DECIMAL_PRECISION when the column declares neither
58
66
  DEFAULT_DECIMAL_SCALE = 9
59
67
 
68
+ # Bit width of named SQL integer types (MySQL, Postgres and SQLite spellings), keyed by the
69
+ # leading word of the lowercased +sql_type+
60
70
  INTEGER_SQL_TYPES = {
61
71
  "tinyint" => 8, "int1" => 8,
62
72
  "smallint" => 16, "int2" => 16, "smallserial" => 16, "serial2" => 16,
@@ -64,6 +74,16 @@ module Herringbone
64
74
  "bigint" => 64, "int8" => 64, "bigserial" => 64, "serial8" => 64
65
75
  }.freeze
66
76
 
77
+ # Declares +column+ on +builder+: hstore as map<string, string>, array columns as a list of
78
+ # the element type, everything else as the scalar type picked by #scalar_type.
79
+ #
80
+ # @param builder [Builder] builder to add the field to
81
+ # @param column [ActiveRecord::ConnectionAdapters::Column] column metadata (+name+, +type+, and
82
+ # when available +sql_type+, +array+, +limit+, +precision+, +scale+)
83
+ # @param nullable [Boolean] whether the field may be null
84
+ # @param primary [Boolean] whether the column is (part of) the primary key
85
+ # @return [Node] the added schema node
86
+ # @raise [ArgumentError] when the builder rejects the field (e.g. a duplicate name)
67
87
  def add_column(builder, column, nullable, primary)
68
88
  sql_type = column.respond_to?(:sql_type) ? column.sql_type.to_s.downcase : ""
69
89
  array = (column.respond_to?(:array) && column.array) || sql_type.end_with?("[]")
@@ -83,10 +103,18 @@ module Herringbone
83
103
  end
84
104
  end
85
105
 
106
+ # Picks the DSL type for a non-array, non-hstore column, following the table above.
107
+ #
108
+ # @param column [ActiveRecord::ConnectionAdapters::Column] column metadata, read for
109
+ # +limit+, +precision+ and +scale+
110
+ # @param type [Symbol, nil] ActiveRecord's abstract type (+column.type+)
111
+ # @param sql_type [String] lowercased database type with any array suffix removed
112
+ # @param primary [Boolean] whether the column is (part of) the primary key
113
+ # @return [Array(Symbol, Hash{Symbol => Object})] DSL type and its options for Builder#column
86
114
  def scalar_type(column, type, sql_type, primary)
87
115
  case type
88
116
  when :integer, :bigint then [integer_type(column, sql_type, primary), {}]
89
- when :float then [sql_type == "float4" ? :float : :double, {}]
117
+ when :float then [(sql_type == "float4") ? :float : :double, {}]
90
118
  when :decimal, :money
91
119
  precision = column.respond_to?(:precision) ? column.precision : nil
92
120
  scale = column.respond_to?(:scale) ? column.scale : nil
@@ -94,12 +122,12 @@ module Herringbone
94
122
  precision = DEFAULT_DECIMAL_PRECISION
95
123
  scale ||= DEFAULT_DECIMAL_SCALE
96
124
  end
97
- [:decimal, { precision: precision, scale: scale || 0 }]
125
+ [:decimal, {precision: precision, scale: scale || 0}]
98
126
  when :boolean then [:boolean, {}]
99
127
  when :binary then [:binary, {}]
100
128
  when :date then [:date, {}]
101
- when :datetime, :timestamp, :timestamptz then [:timestamp, { unit: :micros, utc: true }]
102
- when :time then [:time, { unit: :micros }]
129
+ when :datetime, :timestamp, :timestamptz then [:timestamp, {unit: :micros, utc: true}]
130
+ when :time then [:time, {unit: :micros}]
103
131
  when :json, :jsonb then [:json, {}]
104
132
  when :uuid then [:uuid, {}]
105
133
  else [:string, {}]
@@ -108,6 +136,11 @@ module Herringbone
108
136
 
109
137
  # Named SQL types (smallint, bigint, ...) decide the width; a generic "integer"/"int"
110
138
  # uses the column limit in bytes, since e.g. SQLite reports `t.integer limit: 2` as "integer(2)".
139
+ #
140
+ # @param column [ActiveRecord::ConnectionAdapters::Column] column metadata, read for +limit+
141
+ # @param sql_type [String] lowercased database type with any array suffix removed
142
+ # @param primary [Boolean] true forces int64, since row ids can outgrow the declared width
143
+ # @return [Symbol] one of +:int8+ .. +:int64+ or +:uint8+ .. +:uint64+
111
144
  def integer_type(column, sql_type, primary)
112
145
  bits = INTEGER_SQL_TYPES[sql_type[/\A[a-z0-9]+/]]
113
146
  bits ||= case (column.respond_to?(:limit) ? column.limit : nil)
@@ -120,9 +153,8 @@ module Herringbone
120
153
  bits = 64 if primary
121
154
  unsigned = sql_type.include?("unsigned")
122
155
  return :"uint#{bits}" if unsigned
123
- { 8 => :int8, 16 => :int16, 32 => :int32, 64 => :int64 }.fetch(bits)
156
+ {8 => :int8, 16 => :int16, 32 => :int32, 64 => :int64}.fetch(bits)
124
157
  end
125
158
  end
126
159
  end
127
160
  end
128
-
@@ -13,19 +13,27 @@ module Herringbone
13
13
  # String hashes the same bytes as the stored value. Nulls are never in a bloom filter.
14
14
  # The writer builds them (bloom_filters: option) and reads with where: consult them.
15
15
  class BloomFilter
16
+ # The spec's eight salt constants, one per word of a block
16
17
  SALT = [0x47b6137b, 0x44974d91, 0x8824ad5b, 0xa2b7289d, 0x705495c7, 0x2df1424b, 0x9efc4947, 0x5c6bfb31].freeze
17
18
  S0, S1, S2, S3, S4, S5, S6, S7 = SALT
18
19
  # Low 16 bits of the salts: (lo * salt) mod 2**32 is computed as
19
20
  # (lo & 0xFFFF) * salt + ((lo >> 16) * (salt & 0xFFFF) << 16), which stays a Fixnum
20
21
  L0, L1, L2, L3, L4, L5, L6, L7 = SALT.map { |s| s & 0xFFFF }
22
+ # Size of a block: eight 32-bit words
21
23
  BLOCK_BYTES = 32
24
+ # Smallest bitset: a single block
22
25
  MIN_BYTES = 32
26
+ # Largest bitset accepted when reading or building (128 MiB, the cap parquet-mr uses)
23
27
  MAX_BYTES = 128 * 1024 * 1024
24
28
  # Default cap for filters sized by the writer (as in parquet-mr)
25
29
  DEFAULT_MAX_BYTES = 1024 * 1024
30
+ # Default false positive probability for filters sized by the writer
26
31
  DEFAULT_FPP = 0.01
32
+ # 32-bit mask
27
33
  M32 = 0xFFFF_FFFF
34
+ # 64-bit mask
28
35
  M64 = 0xFFFF_FFFF_FFFF_FFFF
36
+ # Shorthand for the physical type constants
29
37
  T = Format::Type
30
38
  # Physical types a bloom filter can be built for (the spec does not define BOOLEAN hashing)
31
39
  TYPES = [T::INT32, T::INT64, T::INT96, T::FLOAT, T::DOUBLE, T::BYTE_ARRAY, T::FIXED_LEN_BYTE_ARRAY].freeze
@@ -33,10 +41,17 @@ module Herringbone
33
41
  # Bitset size in bytes for +ndv+ distinct values at false positive probability +fpp+, per the
34
42
  # spec's formula (m = -8 * ndv / ln(1 - fpp ** (1/8)) bits), rounded up to a power of two
35
43
  # and clamped to MIN_BYTES..max_bytes
44
+ #
45
+ # @param ndv [Integer] expected number of distinct values (values below 1 count as 1)
46
+ # @param fpp [Float] false positive probability, strictly between 0 and 1
47
+ # @param max_bytes [Integer] upper bound, clamped to MIN_BYTES..MAX_BYTES and rounded down to a
48
+ # power of two
49
+ # @return [Integer] bitset size in bytes, a power of two
50
+ # @raise [ArgumentError] when +fpp+ is not between 0 and 1
36
51
  def self.optimal_num_bytes(ndv, fpp = DEFAULT_FPP, max_bytes: DEFAULT_MAX_BYTES)
37
52
  fpp = Float(fpp)
38
53
  raise ArgumentError, "fpp must be between 0 and 1, got #{fpp}" unless fpp > 0 && fpp < 1
39
- max_bytes = [[Integer(max_bytes), MAX_BYTES].min, MIN_BYTES].max
54
+ max_bytes = Integer(max_bytes).clamp(MIN_BYTES, MAX_BYTES)
40
55
  max_bytes = 1 << (max_bytes.bit_length - 1) # a power of two
41
56
  ndv = [Integer(ndv), 1].max
42
57
  bits = -8.0 * ndv / Math.log(1 - fpp**(1.0 / 8))
@@ -47,6 +62,12 @@ module Herringbone
47
62
  end
48
63
 
49
64
  # XXH64 of the PLAIN encoding of a physical value of +type+ (as the column encoders produce it)
65
+ #
66
+ # @param value [Integer, Float, String, Array<Integer>] the physical value; for INT96 the
67
+ # +[nanos_of_day, julian_day]+ pair the encoder produces
68
+ # @param type [Integer] Format::Type physical type
69
+ # @return [Integer] the unsigned 64-bit hash
70
+ # @raise [UnsupportedError] for BOOLEAN (or any type not in TYPES)
50
71
  def self.hash_physical(value, type)
51
72
  case type
52
73
  when T::INT32 then XXHash.xxh64_u32(value)
@@ -62,6 +83,12 @@ module Herringbone
62
83
  # Hashes of many physical values of +type+, converting them in bulk. With +distinct+, each
63
84
  # distinct physical value is hashed once (floats are compared by their bytes, so -0.0 and 0.0,
64
85
  # or NaNs with different payloads, stay apart as they hash differently).
86
+ #
87
+ # @param values [Array] physical values, as for .hash_physical
88
+ # @param type [Integer] Format::Type physical type
89
+ # @param distinct [Boolean] hash each distinct value once (the result is then shorter)
90
+ # @return [Array<Integer>] the unsigned 64-bit hashes
91
+ # @raise [UnsupportedError] for BOOLEAN (or any type not in TYPES)
65
92
  def self.hash_physical_all(values, type, distinct: false)
66
93
  case type
67
94
  when T::INT32 then XXHash.xxh64_u32_all(distinct ? values.uniq : values)
@@ -80,6 +107,13 @@ module Herringbone
80
107
 
81
108
  # Reads a filter (header and bitset) from +buf+ at +pos+. Returns nil for algorithms, hashes
82
109
  # or compressions this implementation does not know.
110
+ #
111
+ # @param buf [String] bytes holding the BloomFilterHeader followed by the bitset
112
+ # @param pos [Integer] byte offset of the header in +buf+
113
+ # @param column [Schema::Column, nil] column the filter belongs to, for converting values
114
+ # @return [BloomFilter, nil] the filter, or nil when it is of an unsupported kind
115
+ # @raise [FormatError] when the bitset is shorter than the header announces
116
+ # @raise [Thrift::Error] when the header cannot be decoded
83
117
  def self.decode(buf, pos = 0, column: nil)
84
118
  reader = Thrift::Reader.new(buf, pos)
85
119
  header = reader.read_struct(Format::BloomFilterHeader)
@@ -89,48 +123,78 @@ module Herringbone
89
123
  new(bitset: bitset, column: column)
90
124
  end
91
125
 
126
+ # Whether a header describes a filter this implementation can read: split block algorithm,
127
+ # XXH64, uncompressed, and a size that is a multiple of 32 bytes up to MAX_BYTES
128
+ #
129
+ # @param header [Format::BloomFilterHeader] decoded header
130
+ # @return [Boolean] true when supported
92
131
  def self.supported_header?(header)
93
132
  n = header.num_bytes
94
133
  n.is_a?(Integer) && n >= BLOCK_BYTES && (n % BLOCK_BYTES).zero? && n <= MAX_BYTES &&
95
- header.algorithm&.block && header.hash_function&.xxhash &&
96
- (header.compression.nil? || header.compression.uncompressed)
134
+ !(header.algorithm && header.algorithm.block).nil? &&
135
+ !(header.hash_function && header.hash_function.xxhash).nil? &&
136
+ (header.compression.nil? || !header.compression.uncompressed.nil?)
97
137
  end
98
138
 
139
+ # @return [Schema::Column, nil] column whose encoder converts values, nil for raw Strings
99
140
  attr_reader :column
100
141
 
101
142
  # A filter of +num_bytes+ (a multiple of 32, normally a power of two), or one using an existing
102
143
  # +bitset+ String. With a +column+ (a Schema::Column), values are Ruby values converted with
103
144
  # the column's encoder; without one, values must be Strings and their bytes are hashed.
145
+ #
146
+ # @param num_bytes [Integer, nil] bitset size for an empty filter (MIN_BYTES when nil); ignored
147
+ # with +bitset+
148
+ # @param bitset [String, nil] existing bitset (little-endian 32-bit words)
149
+ # @param column [Schema::Column, nil] column the filter is for
150
+ # @raise [ArgumentError] when the size (or +bitset+'s size) is not a multiple of 32 bytes or
151
+ # exceeds MAX_BYTES
152
+ # @raise [UnsupportedError] when the column's physical type cannot have a bloom filter
104
153
  def initialize(num_bytes = nil, bitset: nil, column: nil)
105
- if bitset
106
- @words = bitset.unpack("V*")
107
- else
108
- num_bytes = Integer(num_bytes || MIN_BYTES)
109
- unless num_bytes >= BLOCK_BYTES && (num_bytes % BLOCK_BYTES).zero? && num_bytes <= MAX_BYTES
110
- raise ArgumentError, "Bloom filter size must be a multiple of 32 bytes up to 128MB, got #{num_bytes}"
111
- end
112
- @words = Array.new(num_bytes / 4, 0)
154
+ num_bytes = bitset ? bitset.bytesize : Integer(num_bytes || MIN_BYTES)
155
+ unless num_bytes >= BLOCK_BYTES && (num_bytes % BLOCK_BYTES).zero? && num_bytes <= MAX_BYTES
156
+ raise ArgumentError, "Bloom filter size must be a multiple of 32 bytes up to 128MB, got #{num_bytes}"
113
157
  end
158
+ @words = bitset ? bitset.unpack("V*") : Array.new(num_bytes / 4, 0)
114
159
  @num_blocks = @words.size / 8
115
- raise ArgumentError, "Bloom filter bitset must be a multiple of 32 bytes" if @num_blocks.zero? || @words.size % 8 != 0
116
160
  @column = column
117
161
  if column && !TYPES.include?(column.type)
118
162
  raise UnsupportedError, "Bloom filters are not supported for #{T::NAMES[column.type]} column #{column.dotted_path}"
119
163
  end
120
164
  end
121
165
 
166
+ # @return [Integer] bitset size in bytes
122
167
  def num_bytes = @words.size * 4
123
168
 
169
+ # Adds a value to the filter
170
+ #
171
+ # @param value [Object] Ruby value, converted with the column's encoder (a String without a
172
+ # column)
173
+ # @return [BloomFilter] self
174
+ # @raise [ArgumentError] for nil or a value the column cannot encode
124
175
  def insert(value)
125
176
  insert_hash(hash_of(value))
126
177
  self
127
178
  end
128
179
 
180
+ # Whether the value may be in the filter. Never false for an inserted value; true for a
181
+ # value that was not inserted with about the false positive probability the filter was sized for.
182
+ #
183
+ # @param value [Object] Ruby value, converted with the column's encoder (a String without a
184
+ # column)
185
+ # @return [Boolean] false when the value is definitely absent
186
+ # @raise [ArgumentError] for nil or a value the column cannot encode
129
187
  def might_contain?(value)
130
188
  might_contain_hash?(hash_of(value))
131
189
  end
132
190
 
133
191
  # XXH64 hash the filter uses for a Ruby +value+
192
+ #
193
+ # @param value [Object] Ruby value, converted with the column's encoder (a String without a
194
+ # column)
195
+ # @return [Integer] the unsigned 64-bit hash
196
+ # @raise [ArgumentError] for nil, a value the column cannot encode, or (without a column) a
197
+ # non-String
134
198
  def hash_of(value)
135
199
  raise ArgumentError, "Nulls are not recorded in bloom filters" if value.nil?
136
200
  if @column
@@ -146,6 +210,11 @@ module Herringbone
146
210
  end
147
211
  end
148
212
 
213
+ # Sets the bits for a hash: the upper 32 bits pick the block, the lower 32 bits (multiplied
214
+ # by each salt) one bit in each of its eight words
215
+ #
216
+ # @param h [Integer] unsigned 64-bit XXH64 hash
217
+ # @return [BloomFilter] self
149
218
  def insert_hash(h)
150
219
  i = (((h >> 32) * @num_blocks) >> 32) << 3
151
220
  x0 = h & 0xFFFF
@@ -162,10 +231,14 @@ module Herringbone
162
231
  self
163
232
  end
164
233
 
234
+ # Single-bit masks, +BITS[i] == 1 << i+
165
235
  BITS = Array.new(32) { |i| 1 << i }.freeze
166
236
 
167
237
  # Inserts many hashes at once: faster than insert_hash, as the hashes (mostly Bignums) are
168
238
  # split into 32-bit halves in bulk and the rest is Fixnum arithmetic
239
+ #
240
+ # @param hashes [Array<Integer>] unsigned 64-bit XXH64 hashes
241
+ # @return [BloomFilter] self
169
242
  def insert_hashes(hashes)
170
243
  halves = hashes.pack("Q<*").unpack("V*")
171
244
  w = @words
@@ -191,6 +264,10 @@ module Herringbone
191
264
  self
192
265
  end
193
266
 
267
+ # Whether all the bits for a hash are set (see #insert_hash)
268
+ #
269
+ # @param h [Integer] unsigned 64-bit XXH64 hash
270
+ # @return [Boolean] false when the hashed value is definitely absent
194
271
  def might_contain_hash?(h)
195
272
  i = (((h >> 32) * @num_blocks) >> 32) << 3
196
273
  x0 = h & 0xFFFF
@@ -207,19 +284,27 @@ module Herringbone
207
284
  end
208
285
 
209
286
  # The raw bitset (little-endian 32-bit words)
287
+ #
288
+ # @return [String] binary String of #num_bytes bytes
210
289
  def bitset = @words.pack("V*")
211
290
 
212
291
  # Header and bitset, as stored in a Parquet file
292
+ #
293
+ # @return [String] Thrift-encoded BloomFilterHeader followed by the bitset (binary)
213
294
  def encode
214
295
  header.encode << bitset
215
296
  end
216
297
 
298
+ # @return [String] size and, when set, the column's dotted path
217
299
  def inspect
218
300
  "#<#{self.class.name} #{num_bytes} bytes#{" for #{@column.dotted_path}" if @column}>"
219
301
  end
220
302
 
221
303
  private
222
304
 
305
+ # Header for this filter: split block algorithm, XXH64 hash, uncompressed
306
+ #
307
+ # @return [Format::BloomFilterHeader] the header
223
308
  def header
224
309
  Format::BloomFilterHeader.new(
225
310
  num_bytes: num_bytes,
@@ -234,6 +319,13 @@ module Herringbone
234
319
  # Internal (used by reads with where:): the bloom filter of a column chunk. +column+ is a
235
320
  # dotted path ("a.b"), an Array path or a Schema::Column. Returns a BloomFilter, or nil when
236
321
  # the chunk has none (or one of an unknown kind).
322
+ #
323
+ # @param row_group_index [Integer] index of the row group
324
+ # @param column [String, Array<String, Symbol>, Symbol, Schema::Column] the leaf column
325
+ # @return [BloomFilter, nil] the filter, or nil when there is none or it is unsupported
326
+ # @raise [IndexError] when there is no such row group
327
+ # @raise [ArgumentError] when there is no such column
328
+ # @raise [FormatError] when the filter is truncated or its header cannot be decoded
237
329
  def bloom_filter(row_group_index, column)
238
330
  col = bloom_filter_column(column)
239
331
  rg = row_groups.fetch(row_group_index) { raise IndexError, "No row group #{row_group_index}" }
@@ -256,10 +348,18 @@ module Herringbone
256
348
  bitset = @io.read(header.num_bytes)
257
349
  raise FormatError, "Truncated bloom filter" if bitset.nil? || bitset.bytesize != header.num_bytes
258
350
  BloomFilter.new(bitset: bitset, column: col)
351
+ rescue Thrift::Error => e
352
+ raise FormatError, "Corrupt bloom filter header for #{col.dotted_path}: #{e.message}"
259
353
  end
260
354
 
261
355
  private
262
356
 
357
+ # Resolves the +column+ argument of #bloom_filter to a leaf column
358
+ #
359
+ # @param column [String, Array<String, Symbol>, Symbol, Schema::Column] dotted path, path
360
+ # Array or column
361
+ # @return [Schema::Column] the column
362
+ # @raise [ArgumentError] when the schema has no such column
263
363
  def bloom_filter_column(column)
264
364
  return column if column.is_a?(Schema::Column)
265
365
  col = schema.column(column.is_a?(Array) ? column.map(&:to_s) : column.to_s)
@@ -17,8 +17,11 @@ module Herringbone
17
17
  # After this many values, give up on the dictionary if more than half of them are distinct
18
18
  CARDINALITY_CHECK_AT = 4096
19
19
 
20
+ # Whether String#append_as_bytes (Ruby 3.4+) is available for appending without re-encoding
20
21
  APPEND_AS_BYTES = "".respond_to?(:append_as_bytes)
21
22
 
23
+ # @param width [Integer, nil] byte length of each value for FIXED_LEN_BYTE_ARRAY, nil for BYTE_ARRAY
24
+ # @param dictionary [Boolean] start in dictionary mode; false appends raw bytes from the start
22
25
  def initialize(width: nil, dictionary: true)
23
26
  @width = width
24
27
  @bytes = String.new(encoding: Encoding::BINARY)
@@ -31,14 +34,21 @@ module Herringbone
31
34
  end
32
35
  end
33
36
 
37
+ # @return [Integer] number of values held
34
38
  def size
35
39
  @indices ? @indices.size : @count
36
40
  end
37
41
 
42
+ # @return [Boolean] true when no values are held
38
43
  def empty? = size.zero?
39
44
 
45
+ # @return [Boolean] true while still dictionary-encoding (not yet switched to raw bytes)
40
46
  def dictionary? = !@indices.nil?
41
47
 
48
+ # Appends a value, switching to raw bytes when the dictionary limits are exceeded.
49
+ # @param value [String] bytes of the value; for FIXED_LEN_BYTE_ARRAY it must be +width+ bytes
50
+ # long (not checked here)
51
+ # @return [ByteValues] self
42
52
  def <<(value)
43
53
  if @indices
44
54
  index = @dictionary[value]
@@ -61,16 +71,20 @@ module Herringbone
61
71
  end
62
72
 
63
73
  # Removes the last value
74
+ # @return [void]
64
75
  def pop
65
76
  truncate(size - 1) unless size.zero?
66
77
  end
67
78
 
68
79
  # Supports the `slice!(n..)` form used to roll back a failed row
80
+ # @param range [Range] endless range; values from +range.begin+ on are dropped
81
+ # @return [void]
69
82
  def slice!(range)
70
83
  truncate(range.begin)
71
84
  end
72
85
 
73
86
  # Approximate memory held, used to size row groups
87
+ # @return [Integer] estimated bytes
74
88
  def memory_bytes
75
89
  if @indices
76
90
  # Each distinct value is a String object plus a Hash entry
@@ -82,6 +96,11 @@ module Herringbone
82
96
 
83
97
  # Returns [:dictionary, values, indices] when the column is worth dictionary-encoding,
84
98
  # otherwise [:plain, values]
99
+ #
100
+ # Dictionary encoding is kept unless there are more than 16 values and more than about half
101
+ # of them are distinct.
102
+ # @return [Array(Symbol, Array<String>, Array<Integer>), Array(Symbol, Array<String>)]
103
+ # distinct values and an index per value, or all values in order
85
104
  def materialize
86
105
  if @indices
87
106
  keys = @dictionary.keys
@@ -93,6 +112,9 @@ module Herringbone
93
112
 
94
113
  private
95
114
 
115
+ # Appends +value+ in raw-bytes mode, recording its length for BYTE_ARRAY.
116
+ # @param value [String] bytes of the value, in any encoding
117
+ # @return [ByteValues] self
96
118
  def append_bytes(value)
97
119
  if value.encoding == Encoding::BINARY || value.ascii_only?
98
120
  @bytes << value
@@ -106,6 +128,8 @@ module Herringbone
106
128
  self
107
129
  end
108
130
 
131
+ # Leaves dictionary mode, re-appending every value held so far as raw bytes.
132
+ # @return [void]
109
133
  def switch_to_bytes
110
134
  keys = @dictionary.keys
111
135
  indices = @indices
@@ -113,6 +137,9 @@ module Herringbone
113
137
  indices.each { |i| append_bytes(keys[i]) }
114
138
  end
115
139
 
140
+ # Drops all values after the first +n+.
141
+ # @param n [Integer] number of values to keep
142
+ # @return [void]
116
143
  def truncate(n)
117
144
  if @indices
118
145
  @indices.slice!(n..)
@@ -128,6 +155,8 @@ module Herringbone
128
155
  end
129
156
  end
130
157
 
158
+ # Splits the raw bytes back into one binary String per value.
159
+ # @return [Array<String>]
131
160
  def strings
132
161
  if @lengths
133
162
  pos = 0