rails-paradedb 0.10.0 → 0.12.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 296e696bfa9711537e984df359c43be2481b53a91f75c880dc33ac36aacf4b1c
4
- data.tar.gz: 78781d5fc912c5b57053818d7f3e76e2d01251fb7d31a586edd5d28b88426e69
3
+ metadata.gz: 6cec3ad6c144476e82c57a3f9cba2b9ff497a838c34c032919b1eb785d4da435
4
+ data.tar.gz: 2af9e08e2bf7b9a12f9c005e99bc73f4a804f8a0c23f50c1fbc57637d6b87109
5
5
  SHA512:
6
- metadata.gz: 5b4fecd2d8620a91e782ea62df980ab6638fec2906ad5a380664e97b77a8d0e1b4b656aa330cc4b9e80a99e18f2223cd2d50da6cea98245c63c1c4fe88ca7450
7
- data.tar.gz: a6c42abf6b75aad3ee2d3b58651cfde545ba86deef7e33df0593cafc94da529818dbe41f290286e0acc3eb9bc62071d4d2b06d7976cb8c44d5a94be7c9a37722
6
+ metadata.gz: 1af9bfba690e6ba816fc5972f576885b72408594e68504e6d3bbf063944f6392debcf3cf7a028b5d08ba626d6acfaaeb7dbcb526db4adafa96cec987902fb6fa
7
+ data.tar.gz: 384b58ba9857bc2edb9a984ab71bd60d455daf5add4b3fc83b0c53458892496d21959b84e71680daddf7acc48f7c2832dbf930185661fb3ff587a22399a4bd84
data/CHANGELOG.md CHANGED
@@ -4,7 +4,20 @@ All notable changes to this project will be documented in this file. The format
4
4
 
5
5
  ## [Unreleased]
6
6
 
7
- [0.10.0] - 2026-08-04
7
+ ## [0.12.0] - 2026-08-13
8
+
9
+ ### Changed
10
+
11
+ - **BREAKING**: Search modifiers are now composable functions: `ParadeDB.boost`, `ParadeDB.constant`, `ParadeDB.fuzzy`, `ParadeDB.slop`, and `ParadeDB.tokenize`. Pass the wrapped value to search methods instead of using modifier keyword arguments.
12
+ - **BREAKING**: Removed the public `ParadeDB::Arel` builder, Arel attribute predications, and `Model.paradedb_arel`. ParadeDB queries now use a private builder composed from standard Active Record nodes.
13
+
14
+ ## [0.11.0] - 2026-08-04
15
+
16
+ ### Added
17
+
18
+ - Vector index build options `centroid_ratio`, `training_samples_per_centroid`, and `cluster_replication` in `ParadeDB::Index` `index_options` and `add_paradedb_index` (pg_search 0.25.0+). They are emitted in the `WITH (...)` clause and round-tripped through the schema dumper.
19
+
20
+ ## [0.10.0] - 2026-08-04
8
21
 
9
22
  ### Added
10
23
 
@@ -20,7 +33,7 @@ All notable changes to this project will be documented in this file. The format
20
33
  - **BREAKING**: Index creation always emits `USING paradedb`, which requires pg_search 0.25.0+. There is no option to select the legacy `bm25` access method.
21
34
  - **BREAKING**: The default index name is now `<table>_search_idx` (previously `<table>_bm25_idx`), and the index generator emits `Create<Model>SearchIndex` migrations named `create_<table>_search_index.rb`.
22
35
 
23
- [0.9.0] - 2026-07-14
36
+ ## [0.9.0] - 2026-07-14
24
37
 
25
38
  ### Added
26
39
 
@@ -170,7 +183,8 @@ All notable changes to this project will be documented in this file. The format
170
183
  - Schema dump/load round-trip for tokenizer configuration and index options
171
184
  (including `target_segment_count`)
172
185
 
173
- [Unreleased]: https://github.com/paradedb/rails-paradedb/compare/v0.10.0...HEAD
186
+ [0.12.0]: https://github.com/paradedb/rails-paradedb/releases/tag/v0.12.0
187
+ [0.11.0]: https://github.com/paradedb/rails-paradedb/releases/tag/v0.11.0
174
188
  [0.10.0]: https://github.com/paradedb/rails-paradedb/releases/tag/v0.10.0
175
189
  [0.9.0]: https://github.com/paradedb/rails-paradedb/releases/tag/v0.9.0
176
190
  [0.8.0]: https://github.com/paradedb/rails-paradedb/releases/tag/v0.8.0
data/README.md CHANGED
@@ -36,47 +36,45 @@
36
36
 
37
37
  ## ParadeDB for Rails
38
38
 
39
- The official ActiveRecord integration for [ParadeDB](https://paradedb.com) (powered by the [`pg_search`](https://github.com/paradedb/paradedb) Postgres extension), including first-class support for managing ParadeDB indexes and running queries using the full ParadeDB API. Follow the [getting started guide](https://docs.paradedb.com/documentation/getting-started/environment#rails) to begin.
39
+ The official [ActiveRecord](https://guides.rubyonrails.org/active_record_basics.html) integration for [ParadeDB](https://paradedb.com) (powered by the [`pg_search`](https://github.com/paradedb/paradedb) Postgres extension), including first-class support for managing ParadeDB indexes and running queries using the full ParadeDB API. The integration covers both [full-text search](https://docs.paradedb.com/documentation/full-text/overview) and [vector search](https://docs.paradedb.com/documentation/vector/overview) over pgvector `vector` types. Follow the [getting started guide](https://docs.paradedb.com/documentation/getting-started/environment#rails) to begin.
40
40
 
41
41
  ## Requirements & Compatibility
42
42
 
43
- | Component | Supported |
44
- | ---------- | ----------------------------------------------------------------- |
45
- | Ruby | 3.2+ |
46
- | Rails | 7.2+ |
47
- | ParadeDB | 0.25.0+ |
48
- | PostgreSQL | 15+ (PostgreSQL adapter with ParadeDB extension) |
49
- | pgvector | Required for vector search; included in the ParadeDB Docker image |
50
-
51
- ## Vector Search
52
-
53
- rails-paradedb supports full-text search and vector search over pgvector `vector(n)` columns. See the [vector search documentation](https://docs.paradedb.com/documentation/vector/overview) for details.
43
+ | Component | Supported |
44
+ | ---------- | ------------------------------------------------------------------ |
45
+ | Ruby | 3.2+ |
46
+ | Rails | 7.2+ |
47
+ | ParadeDB | 0.25.0+ |
48
+ | PostgreSQL | 15+ (with the ParadeDB pg_search extension) |
49
+ | pgvector | Required for vector search (included in the ParadeDB Docker image) |
54
50
 
55
51
  ## Examples
56
52
 
57
53
  - [Quickstart](examples/quickstart/quickstart.rb)
58
54
  - [Vector Search](examples/vector_search/vector_search.rb)
59
55
  - [Faceted Search](examples/faceted_search/faceted_search.rb)
56
+ - [Hybrid Search (RRF)](examples/hybrid_rrf/hybrid_rrf.rb)
57
+ - [Retrieval-Augmented Generation (RAG)](examples/rag/rag.rb)
60
58
  - [Autocomplete](examples/autocomplete/autocomplete.rb)
61
59
  - [More Like This](examples/more_like_this/more_like_this.rb)
62
- - [Hybrid Search (RRF)](examples/hybrid_rrf/hybrid_rrf.rb)
63
- - [RAG](examples/rag/rag.rb)
60
+
61
+ See [examples/README.md](examples/README.md) for setup instructions and a description of each example.
64
62
 
65
63
  ## Contributing
66
64
 
67
- See [CONTRIBUTING.md](CONTRIBUTING.md) for development setup, test commands, linting, and PR workflow.
65
+ See [CONTRIBUTING.md](CONTRIBUTING.md) for development setup, running tests, linting, and the PR workflow.
68
66
 
69
67
  ## Support
70
68
 
71
- If you're missing a feature or have found a bug, open a
69
+ If you're missing a feature or have found a bug, please open a
72
70
  [GitHub Issue](https://github.com/paradedb/rails-paradedb/issues/new/choose).
73
71
 
74
- For community support:
72
+ To get community support, you can:
75
73
 
76
- - Join the [ParadeDB Slack Community](https://paradedb.com/slack)
77
- - Ask in [ParadeDB Discussions](https://github.com/paradedb/paradedb/discussions)
74
+ - Post a question in the [ParadeDB Slack Community](https://paradedb.com/slack)
75
+ - Ask for help on our [GitHub Discussions](https://github.com/paradedb/paradedb/discussions)
78
76
 
79
- For commercial support, contact [sales@paradedb.com](mailto:sales@paradedb.com).
77
+ If you need commercial support, please [contact the ParadeDB team](mailto:sales@paradedb.com).
80
78
 
81
79
  ## Acknowledgments
82
80
 
@@ -7,14 +7,7 @@ module ParadeDB
7
7
  # Backward-compatible reader for code that accessed `filtered_spec.filter`.
8
8
  alias filter agg_filter
9
9
  end
10
- FieldTermFilter = Struct.new(
11
- :field,
12
- :term,
13
- :distance,
14
- :prefix,
15
- :transposition_cost_one,
16
- keyword_init: true
17
- )
10
+ FieldTermFilter = Struct.new(:field, :term, keyword_init: true)
18
11
 
19
12
  TERMS_ORDER = {
20
13
  count_desc: { "_count" => "desc" },
@@ -154,16 +147,9 @@ module ParadeDB
154
147
  { "top_hits" => payload }
155
148
  end
156
149
 
157
- def filtered(spec, filter: nil, field: nil, term: nil, distance: nil, prefix: nil, transposition_cost_one: nil)
150
+ def filtered(spec, filter: nil, field: nil, term: nil)
158
151
  normalized_spec = normalize_spec(spec)
159
- normalized_filter = normalize_filter(
160
- filter: filter,
161
- field: field,
162
- term: term,
163
- distance: distance,
164
- prefix: prefix,
165
- transposition_cost_one: transposition_cost_one
166
- )
152
+ normalized_filter = normalize_filter(filter: filter, field: field, term: term)
167
153
  FilteredSpec.new(spec: normalized_spec, agg_filter: normalized_filter)
168
154
  end
169
155
 
@@ -211,7 +197,7 @@ module ParadeDB
211
197
  end
212
198
  private_class_method :normalize_spec
213
199
 
214
- def normalize_filter(filter:, field:, term:, distance:, prefix:, transposition_cost_one:)
200
+ def normalize_filter(filter:, field:, term:)
215
201
  if filter
216
202
  if !field.nil? || !term.nil?
217
203
  raise ArgumentError, "filtered aggregation accepts either filter: or field/term arguments, not both"
@@ -223,17 +209,7 @@ module ParadeDB
223
209
  raise ArgumentError, "filtered aggregation requires filter: or both field: and term:"
224
210
  end
225
211
 
226
- normalized_distance = distance.nil? ? nil : normalize_non_negative_integer(distance, "distance")
227
- normalized_prefix = normalize_boolean_option(prefix, "prefix")
228
- normalized_transposition = normalize_boolean_option(transposition_cost_one, "transposition_cost_one")
229
-
230
- FieldTermFilter.new(
231
- field: normalize_field(field),
232
- term: term,
233
- distance: normalized_distance,
234
- prefix: normalized_prefix,
235
- transposition_cost_one: normalized_transposition
236
- )
212
+ FieldTermFilter.new(field: normalize_field(field), term: term)
237
213
  end
238
214
  private_class_method :normalize_filter
239
215
 
@@ -318,14 +294,6 @@ module ParadeDB
318
294
  end
319
295
  private_class_method :normalize_sort_direction
320
296
 
321
- def normalize_boolean_option(value, name)
322
- return nil if value.nil?
323
- return value if value == true || value == false
324
-
325
- raise ArgumentError, "#{name} must be true, false, or nil"
326
- end
327
- private_class_method :normalize_boolean_option
328
-
329
297
  def deep_stringify(value)
330
298
  case value
331
299
  when Hash
@@ -83,6 +83,8 @@ module ParadeDB
83
83
  # Consumed by migration helpers; validates and normalizes the DSL class
84
84
  class DefinitionCompiler
85
85
  FIELD_OPTION_KEYS = %i[fast record normalizer expand_dots].freeze
86
+ INDEX_OPTION_KEYS = %i[target_segment_count centroid_ratio training_samples_per_centroid cluster_replication].freeze
87
+ POSITIVE_INTEGER_INDEX_OPTION_KEYS = %i[target_segment_count training_samples_per_centroid cluster_replication].freeze
86
88
 
87
89
  class Compiled
88
90
  attr_reader :table_name, :key_field, :index_name, :entries, :index_options, :field_options, :where
@@ -238,16 +240,26 @@ module ParadeDB
238
240
  memo[key.to_sym] = value
239
241
  end
240
242
 
241
- unknown = normalized.keys - [:target_segment_count]
243
+ unknown = normalized.keys - INDEX_OPTION_KEYS
242
244
  unless unknown.empty?
243
245
  raise InvalidIndexDefinition,
244
246
  "unknown index_options keys: #{unknown.map(&:inspect).join(', ')}"
245
247
  end
246
248
 
247
- if normalized.key?(:target_segment_count)
248
- target = normalized[:target_segment_count]
249
- unless target.is_a?(Integer) && target.positive?
250
- raise InvalidIndexDefinition, "index_options[:target_segment_count] must be an Integer > 0"
249
+ POSITIVE_INTEGER_INDEX_OPTION_KEYS.each do |key|
250
+ next unless normalized.key?(key)
251
+
252
+ value = normalized[key]
253
+ unless value.is_a?(Integer) && value.positive?
254
+ raise InvalidIndexDefinition, "index_options[#{key.inspect}] must be an Integer > 0"
255
+ end
256
+ end
257
+
258
+ if normalized.key?(:centroid_ratio)
259
+ ratio = normalized[:centroid_ratio]
260
+ unless ratio.is_a?(Numeric) && ratio >= 0.000001 && ratio <= 1.0
261
+ raise InvalidIndexDefinition,
262
+ "index_options[:centroid_ratio] must be a Numeric between 0.000001 and 1.0"
251
263
  end
252
264
  end
253
265
 
@@ -100,9 +100,12 @@ module ParadeDB
100
100
  options << "key_field=#{quote(compiled.key_field.to_s)}"
101
101
 
102
102
  compiled.index_options.each do |key, value|
103
- case key.to_sym
104
- when :target_segment_count
105
- options << "target_segment_count=#{Integer(value)}"
103
+ name = key.to_sym
104
+ case name
105
+ when :target_segment_count, :training_samples_per_centroid, :cluster_replication
106
+ options << "#{name}=#{Integer(value)}"
107
+ when :centroid_ratio
108
+ options << "centroid_ratio=#{Float(value)}"
106
109
  else
107
110
  raise ParadeDB::InvalidIndexDefinition, "unsupported index option #{key.inspect}"
108
111
  end
@@ -297,6 +300,7 @@ module ParadeDB
297
300
  c.relname AS index_name,
298
301
  t.relname AS table_name,
299
302
  pg_get_indexdef(c.oid) AS indexdef,
303
+ array_to_json(c.reloptions)::text AS reloptions,
300
304
  pg_get_expr(i.indpred, i.indrelid) AS where_clause
301
305
  FROM pg_class c
302
306
  JOIN pg_namespace n ON n.oid = c.relnamespace
@@ -319,7 +323,7 @@ module ParadeDB
319
323
  name = row["index_name"]
320
324
 
321
325
  key_field = extract_paradedb_key_field(indexdef)
322
- index_options = extract_paradedb_index_options(indexdef)
326
+ index_options = extract_paradedb_index_options(row["reloptions"])
323
327
  fields_sql = extract_paradedb_fields_sql(indexdef)
324
328
  where = normalize_paradedb_where_clause(row["where_clause"])
325
329
 
@@ -369,27 +373,30 @@ module ParadeDB
369
373
  nil
370
374
  end
371
375
 
372
- def extract_paradedb_index_options(indexdef)
373
- with_sql, = extract_paradedb_with_components(indexdef)
374
- options = {}
375
- split_sql_arguments(with_sql).each do |argument|
376
- key, value_sql = split_assignment(argument)
377
- next if key.nil?
378
- next if key == "key_field"
376
+ def extract_paradedb_index_options(reloptions)
377
+ paradedb_reloption_entries(reloptions).each_with_object({}) do |entry, options|
378
+ key, separator, value = entry.to_s.partition("=")
379
+ next if separator.empty?
379
380
 
380
381
  case key
381
- when "target_segment_count"
382
- parsed = parse_sql_literal(value_sql)
383
- if parsed.is_a?(Integer)
384
- options[:target_segment_count] = parsed
385
- elsif parsed.is_a?(String) && parsed.match?(/\A\d+\z/)
386
- options[:target_segment_count] = parsed.to_i
387
- end
382
+ when "target_segment_count", "training_samples_per_centroid", "cluster_replication"
383
+ parsed = Integer(value, 10, exception: false)
384
+ options[key.to_sym] = parsed if parsed
385
+ when "centroid_ratio"
386
+ parsed = Float(value, exception: false)
387
+ options[:centroid_ratio] = parsed if parsed
388
388
  end
389
389
  end
390
- options
391
- rescue
392
- {}
390
+ end
391
+
392
+ def paradedb_reloption_entries(reloptions)
393
+ case reloptions
394
+ when Array then reloptions
395
+ when String then Array(JSON.parse(reloptions))
396
+ else []
397
+ end
398
+ rescue JSON::ParserError
399
+ []
393
400
  end
394
401
 
395
402
  def extract_paradedb_fields_sql(indexdef)
@@ -410,27 +417,6 @@ module ParadeDB
410
417
  indexdef[start..pos - 2]
411
418
  end
412
419
 
413
- def extract_paradedb_with_components(indexdef)
414
- match = indexdef.match(/WITH\s*\(/im)
415
- start = match.end(0)
416
- depth = 1
417
- pos = start
418
- while pos < indexdef.length && depth > 0
419
- case indexdef[pos]
420
- when "(" then depth += 1
421
- when ")" then depth -= 1
422
- end
423
- pos += 1
424
- end
425
- raise "Found invalid index definition `#{indexdef}`" if depth != 0
426
-
427
- with_sql = indexdef[start..pos - 2]
428
- trailing_sql = indexdef[pos..]&.strip
429
- trailing_sql = nil if trailing_sql == ""
430
-
431
- [with_sql, trailing_sql]
432
- end
433
-
434
420
  def normalize_paradedb_where_clause(where)
435
421
  return nil if where.nil?
436
422
 
@@ -11,12 +11,14 @@ module ParadeDB
11
11
  :paradedb_search,
12
12
  :more_like_this,
13
13
  :nearest,
14
+ :l2_distance,
15
+ :cosine_distance,
16
+ :inner_product,
14
17
  :with_facets,
15
18
  :facets,
16
19
  :with_agg,
17
20
  :facets_agg,
18
21
  :aggregate_by,
19
- :paradedb_arel,
20
22
  :paradedb_index,
21
23
  :paradedb_index_class,
22
24
  :paradedb_index_classes,
@@ -74,6 +76,21 @@ module ParadeDB
74
76
  all.extending(SearchMethods).nearest(column, vector, metric: metric)
75
77
  end
76
78
 
79
+ def l2_distance(column, vector)
80
+ ensure_postgres!
81
+ QueryBuilder.new(table_name).vector_distance(column, vector, metric: :l2)
82
+ end
83
+
84
+ def cosine_distance(column, vector)
85
+ ensure_postgres!
86
+ QueryBuilder.new(table_name).vector_distance(column, vector, metric: :cosine)
87
+ end
88
+
89
+ def inner_product(column, vector)
90
+ ensure_postgres!
91
+ QueryBuilder.new(table_name).vector_distance(column, vector, metric: :ip)
92
+ end
93
+
77
94
  def with_facets(*fields, **opts)
78
95
  ensure_postgres!
79
96
  paradedb_validate_index!
@@ -108,11 +125,6 @@ module ParadeDB
108
125
  )
109
126
  end
110
127
 
111
- def paradedb_arel
112
- ensure_postgres!
113
- @paradedb_arel ||= ParadeDB::Arel::Builder.new(table_name)
114
- end
115
-
116
128
  def paradedb_index(index_class)
117
129
  @paradedb_explicit_index_classes ||= []
118
130
  @paradedb_explicit_index_classes << index_class unless @paradedb_explicit_index_classes.include?(index_class)