sequel-duckdb 0.2.0 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (60) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +31 -20
  3. data/lib/sequel/duckdb/version.rb +1 -4
  4. metadata +3 -59
  5. data/.beads/.beads-credential-key +0 -1
  6. data/.beads/.gitignore +0 -66
  7. data/.beads/README.md +0 -85
  8. data/.beads/config.yaml +0 -56
  9. data/.beads/hooks/post-checkout +0 -24
  10. data/.beads/hooks/post-merge +0 -24
  11. data/.beads/hooks/pre-commit +0 -24
  12. data/.beads/hooks/pre-push +0 -24
  13. data/.beads/hooks/prepare-commit-msg +0 -24
  14. data/.beads/metadata.json +0 -7
  15. data/.kiro/specs/advanced-sql-features-implementation/design.md +0 -26
  16. data/.kiro/specs/advanced-sql-features-implementation/requirements.md +0 -43
  17. data/.kiro/specs/advanced-sql-features-implementation/tasks.md +0 -28
  18. data/.kiro/specs/duckdb-sql-syntax-compatibility/design.md +0 -272
  19. data/.kiro/specs/duckdb-sql-syntax-compatibility/requirements.md +0 -84
  20. data/.kiro/specs/duckdb-sql-syntax-compatibility/tasks.md +0 -107
  21. data/.kiro/specs/edge-cases-and-validation-fixes/requirements.md +0 -32
  22. data/.kiro/specs/integration-test-database-setup/design.md +0 -0
  23. data/.kiro/specs/integration-test-database-setup/requirements.md +0 -117
  24. data/.kiro/specs/sequel-duckdb-adapter/design.md +0 -549
  25. data/.kiro/specs/sequel-duckdb-adapter/requirements.md +0 -202
  26. data/.kiro/specs/sequel-duckdb-adapter/tasks.md +0 -292
  27. data/.kiro/specs/sql-expression-handling-fix/design.md +0 -331
  28. data/.kiro/specs/sql-expression-handling-fix/requirements.md +0 -86
  29. data/.kiro/specs/sql-expression-handling-fix/tasks.md +0 -25
  30. data/.kiro/specs/test-infrastructure-improvements/requirements.md +0 -106
  31. data/.kiro/steering/product.md +0 -26
  32. data/.kiro/steering/structure.md +0 -88
  33. data/.kiro/steering/tech.md +0 -137
  34. data/.kiro/steering/testing.md +0 -213
  35. data/.mdformat.toml +0 -2
  36. data/.release-please-manifest.json +0 -3
  37. data/.rubocop.yml +0 -161
  38. data/.rubocop_todo.yml +0 -323
  39. data/.yardopts +0 -8
  40. data/AGENTS.md +0 -154
  41. data/API_DOCUMENTATION.md +0 -943
  42. data/FINAL_STATUS.md +0 -99
  43. data/MIGRATION_EXAMPLES.md +0 -740
  44. data/PERFORMANCE_OPTIMIZATIONS.md +0 -726
  45. data/REFACTORING_SUMMARY.md +0 -264
  46. data/Rakefile +0 -43
  47. data/TASK_10.2_IMPLEMENTATION_SUMMARY.md +0 -182
  48. data/docs/DUCKDB_SQL_PATTERNS.md +0 -448
  49. data/docs/TASK_12_VERIFICATION_SUMMARY.md +0 -135
  50. data/justfile +0 -52
  51. data/plans/date_arithmetic.md +0 -420
  52. data/plans/engineering/Sequel.md +0 -471
  53. data/plans/engineering/duckdb.md +0 -712
  54. data/plans/engineering/sqlite.md +0 -453
  55. data/plans/mock_connection_bug.md +0 -333
  56. data/plans/mock_without_driver_gem.md +0 -371
  57. data/plans/over_engineering_analysis.md +0 -122
  58. data/plans/schema_management.md +0 -383
  59. data/release-please-config.json +0 -14
  60. data/sig/sequel/duckdb.rbs +0 -6
@@ -1,726 +0,0 @@
1
- # Performance Optimization Guide for Sequel-DuckDB
2
-
3
- This guide provides comprehensive strategies for optimizing performance when using Sequel with DuckDB, leveraging DuckDB's unique strengths as an analytical database engine.
4
-
5
- ## Table of Contents
6
-
7
- 1. [Understanding DuckDB's Architecture](#understanding-duckdbs-architecture)
8
- 2. [Query Optimization](#query-optimization)
9
- 3. [Schema Design](#schema-design)
10
- 4. [Bulk Operations](#bulk-operations)
11
- 5. [Memory Management](#memory-management)
12
- 6. [Connection Optimization](#connection-optimization)
13
- 7. [Monitoring and Profiling](#monitoring-and-profiling)
14
- 8. [Best Practices](#best-practices)
15
-
16
- ## Understanding DuckDB's Architecture
17
-
18
- DuckDB is designed as an analytical database with several key characteristics that affect performance optimization:
19
-
20
- ### Columnar Storage
21
-
22
- - Data is stored column-wise, making analytical queries very efficient
23
- - SELECT queries that access few columns are much faster
24
- - Aggregations and analytical functions are highly optimized
25
-
26
- ### Vectorized Execution
27
-
28
- - Operations are performed on batches of data (vectors) rather than row-by-row
29
- - This reduces function call overhead and improves CPU cache utilization
30
- - Particularly beneficial for analytical workloads
31
-
32
- ### In-Memory Processing
33
-
34
- - DuckDB can efficiently process data that fits in memory
35
- - Automatic memory management with spill-to-disk for larger datasets
36
- - Memory-mapped files for efficient file-based database access
37
-
38
- ## Query Optimization
39
-
40
- ### 1. Column Selection Optimization
41
-
42
- **Always select only the columns you need:**
43
-
44
- ```ruby
45
- # ❌ Inefficient - selects all columns
46
- users = db[:users].where(active: true).all
47
-
48
- # ✅ Efficient - selects only needed columns
49
- users = db[:users].select(:id, :name, :email).where(active: true).all
50
-
51
- # ✅ Even better for large result sets
52
- db[:users].select(:id, :name, :email).where(active: true).each do |user|
53
- # Process user
54
- end
55
- ```
56
-
57
- ### 2. Predicate Pushdown
58
-
59
- **Apply filters as early as possible:**
60
-
61
- ```ruby
62
- # ❌ Less efficient - filtering after join
63
- result = db[:users]
64
- .join(:orders, user_id: :id)
65
- .where(users__active: true, orders__status: 'completed')
66
-
67
- # ✅ More efficient - filter before join when possible
68
- active_users = db[:users].where(active: true)
69
- completed_orders = db[:orders].where(status: 'completed')
70
- result = active_users.join(completed_orders, user_id: :id)
71
- ```
72
-
73
- ### 3. Index Utilization
74
-
75
- **Create indexes for frequently queried columns:**
76
-
77
- ```ruby
78
- # Create indexes for common query patterns
79
- db.add_index :users, :email
80
- db.add_index :orders, [:user_id, :status]
81
- db.add_index :products, [:category_id, :active]
82
-
83
- # Composite indexes for multi-column queries
84
- db.add_index :order_items, [:order_id, :product_id]
85
-
86
- # Partial indexes for filtered queries
87
- db.add_index :products, :price, where: { active: true }
88
- ```
89
-
90
- ### 4. Query Plan Analysis
91
-
92
- **Use EXPLAIN to understand query execution:**
93
-
94
- ```ruby
95
- # Analyze query performance
96
- query = db[:users].join(:orders, user_id: :id).where(status: 'completed')
97
- puts query.explain
98
-
99
- # Look for:
100
- # - Index usage
101
- # - Join algorithms
102
- # - Filter pushdown
103
- # - Estimated row counts
104
- ```
105
-
106
- ### 5. Analytical Query Optimization
107
-
108
- **Leverage DuckDB's analytical capabilities:**
109
-
110
- ```ruby
111
- # ✅ Efficient analytical queries
112
- sales_summary = db[:sales]
113
- .select(
114
- :product_category,
115
- Sequel.function(:sum, :amount).as(:total_sales),
116
- Sequel.function(:avg, :amount).as(:avg_sale),
117
- Sequel.function(:count, :id).as(:transaction_count),
118
- Sequel.function(:percentile_cont, 0.5).within_group(:amount).as(:median_sale)
119
- )
120
- .group(:product_category)
121
- .order(Sequel.desc(:total_sales))
122
-
123
- # ✅ Window functions for advanced analytics
124
- monthly_trends = db[:sales]
125
- .select(
126
- :month,
127
- :amount,
128
- Sequel.function(:lag, :amount, 1).over(order: :month).as(:prev_month),
129
- Sequel.function(:sum, :amount).over(order: :month).as(:running_total),
130
- Sequel.function(:rank).over(partition: :category, order: Sequel.desc(:amount)).as(:category_rank)
131
- )
132
- ```
133
-
134
- ## Schema Design
135
-
136
- ### 1. Optimal Data Types
137
-
138
- **Choose appropriate data types for performance:**
139
-
140
- ```ruby
141
- # ✅ Efficient data types
142
- db.create_table :products do
143
- primary_key :id # INTEGER is efficient
144
- String :name, size: 255 # Fixed-size strings when possible
145
- Decimal :price, size: [10, 2] # Precise for monetary values
146
- Integer :stock_quantity # INTEGER for counts
147
- Boolean :active # BOOLEAN is very efficient
148
- Date :created_date # DATE for date-only values
149
- DateTime :created_at # TIMESTAMP for full datetime
150
-
151
- # DuckDB-specific optimized types
152
- column :tags, 'VARCHAR[]' # Arrays for multi-value attributes
153
- column :metadata, 'JSON' # JSON for flexible data
154
- end
155
-
156
- # ❌ Avoid oversized types
157
- # String :description, size: 10000 # Use TEXT instead
158
- # Float :price # Use DECIMAL for money
159
- ```
160
-
161
- ### 2. Partitioning Strategy
162
-
163
- **Design tables for analytical workloads:**
164
-
165
- ```ruby
166
- # ✅ Time-based partitioning pattern
167
- db.create_table :sales_2024_q1 do
168
- primary_key :id
169
- foreign_key :product_id, :products
170
- Decimal :amount, size: [10, 2]
171
- Date :sale_date
172
- DateTime :created_at
173
-
174
- # Constraint to enforce partition bounds
175
- constraint(:date_range) { (sale_date >= '2024-01-01') & (sale_date < '2024-04-01') }
176
- end
177
-
178
- # Create view for unified access
179
- db.run <<~SQL
180
- CREATE VIEW sales AS
181
- SELECT * FROM sales_2024_q1
182
- UNION ALL
183
- SELECT * FROM sales_2024_q2
184
- -- Add more partitions as needed
185
- SQL
186
- ```
187
-
188
- ### 3. Denormalization for Analytics
189
-
190
- **Consider denormalization for read-heavy analytical workloads:**
191
-
192
- ```ruby
193
- # ✅ Denormalized table for analytics
194
- db.create_table :order_analytics do
195
- primary_key :id
196
- Integer :order_id
197
- Integer :user_id
198
- String :user_name # Denormalized from users table
199
- String :user_email # Denormalized from users table
200
- Integer :product_id
201
- String :product_name # Denormalized from products table
202
- String :category_name # Denormalized from categories table
203
- Decimal :unit_price, size: [10, 2]
204
- Integer :quantity
205
- Decimal :total_amount, size: [10, 2]
206
- Date :order_date
207
- DateTime :created_at
208
- end
209
-
210
- # Populate with materialized view pattern
211
- db.run <<~SQL
212
- INSERT INTO order_analytics
213
- SELECT
214
- oi.id,
215
- o.id as order_id,
216
- u.id as user_id,
217
- u.name as user_name,
218
- u.email as user_email,
219
- p.id as product_id,
220
- p.name as product_name,
221
- c.name as category_name,
222
- oi.unit_price,
223
- oi.quantity,
224
- oi.unit_price * oi.quantity as total_amount,
225
- o.created_at::DATE as order_date,
226
- o.created_at
227
- FROM order_items oi
228
- JOIN orders o ON oi.order_id = o.id
229
- JOIN users u ON o.user_id = u.id
230
- JOIN products p ON oi.product_id = p.id
231
- JOIN categories c ON p.category_id = c.id
232
- SQL
233
- ```
234
-
235
- ## Bulk Operations
236
-
237
- ### 1. Efficient Bulk Inserts
238
-
239
- **Use multi_insert for large data loads:**
240
-
241
- ```ruby
242
- # ✅ Efficient bulk insert
243
- data = []
244
- 1000.times do |i|
245
- data << {
246
- name: "User #{i}",
247
- email: "user#{i}@example.com",
248
- created_at: Time.now
249
- }
250
- end
251
-
252
- # Single transaction for all inserts
253
- db.transaction do
254
- db[:users].multi_insert(data)
255
- end
256
-
257
- # ✅ For very large datasets, use batching
258
- def bulk_insert_batched(db, table, data, batch_size = 1000)
259
- data.each_slice(batch_size) do |batch|
260
- db.transaction do
261
- db[table].multi_insert(batch)
262
- end
263
- end
264
- end
265
-
266
- bulk_insert_batched(db, :users, large_dataset, 5000)
267
- ```
268
-
269
- ### 2. Bulk Updates
270
-
271
- **Efficient bulk update patterns:**
272
-
273
- ```ruby
274
- # ✅ Single UPDATE statement for bulk changes
275
- db[:products].where(category_id: 1).update(
276
- active: false,
277
- updated_at: Time.now
278
- )
279
-
280
- # ✅ Conditional bulk updates
281
- db.run <<~SQL
282
- UPDATE products
283
- SET
284
- status = CASE
285
- WHEN stock_quantity = 0 THEN 'out_of_stock'
286
- WHEN stock_quantity < 10 THEN 'low_stock'
287
- ELSE 'in_stock'
288
- END,
289
- updated_at = NOW()
290
- WHERE status != CASE
291
- WHEN stock_quantity = 0 THEN 'out_of_stock'
292
- WHEN stock_quantity < 10 THEN 'low_stock'
293
- ELSE 'in_stock'
294
- END
295
- SQL
296
- ```
297
-
298
- ### 3. Data Loading from Files
299
-
300
- **Leverage DuckDB's file reading capabilities:**
301
-
302
- ```ruby
303
- # ✅ Direct CSV import (very fast)
304
- db.run <<~SQL
305
- CREATE TABLE temp_sales AS
306
- SELECT * FROM read_csv_auto('sales_data.csv')
307
- SQL
308
-
309
- # ✅ Parquet files (excellent for analytics)
310
- db.run <<~SQL
311
- CREATE TABLE sales_archive AS
312
- SELECT * FROM read_parquet('sales_archive.parquet')
313
- SQL
314
-
315
- # ✅ JSON files
316
- db.run <<~SQL
317
- CREATE TABLE user_events AS
318
- SELECT * FROM read_json_auto('user_events.json')
319
- SQL
320
- ```
321
-
322
- ## Memory Management
323
-
324
- ### 1. Connection Configuration
325
-
326
- **Optimize DuckDB memory settings:**
327
-
328
- ```ruby
329
- # ✅ Configure memory limits
330
- db = Sequel.connect(
331
- adapter: 'duckdb',
332
- database: '/path/to/database.duckdb',
333
- config: {
334
- memory_limit: '4GB', # Set appropriate memory limit
335
- threads: 8, # Use multiple threads
336
- max_memory: '8GB', # Maximum memory before spilling
337
- temp_directory: '/tmp/duckdb' # Temporary file location
338
- }
339
- )
340
-
341
- # ✅ Runtime memory configuration
342
- db.run "SET memory_limit='2GB'"
343
- db.run "SET threads=4"
344
- ```
345
-
346
- ### 2. Result Set Management
347
-
348
- **Handle large result sets efficiently:**
349
-
350
- ```ruby
351
- # ❌ Loads entire result set into memory
352
- all_orders = db[:orders].all
353
-
354
- # ✅ Process results in batches
355
- db[:orders].paged_each(rows_per_fetch: 1000) do |order|
356
- # Process each order
357
- process_order(order)
358
- end
359
-
360
- # ✅ Use streaming for very large datasets
361
- db[:large_table].use_cursor.each do |row|
362
- # Process row by row without loading all into memory
363
- process_row(row)
364
- end
365
-
366
- # ✅ Limit result sets when possible
367
- recent_orders = db[:orders]
368
- .where { created_at > Date.today - 30 }
369
- .order(Sequel.desc(:created_at))
370
- .limit(1000)
371
- ```
372
-
373
- ### 3. Connection Pooling
374
-
375
- **Optimize connection management:**
376
-
377
- ```ruby
378
- # ✅ Connection pool configuration
379
- db = Sequel.connect(
380
- adapter: 'duckdb',
381
- database: '/path/to/database.duckdb',
382
- max_connections: 10, # Pool size
383
- pool_timeout: 5, # Connection timeout
384
- pool_sleep_time: 0.001, # Sleep between retries
385
- pool_connection_validation: true
386
- )
387
-
388
- # ✅ Proper connection cleanup
389
- begin
390
- db.transaction do
391
- # Database operations
392
- end
393
- ensure
394
- db.disconnect if db
395
- end
396
- ```
397
-
398
- ## Connection Optimization
399
-
400
- ### 1. Connection Reuse
401
-
402
- **Minimize connection overhead:**
403
-
404
- ```ruby
405
- # ✅ Reuse connections
406
- class DatabaseManager
407
- def self.connection
408
- @connection ||= Sequel.connect('duckdb:///app.duckdb')
409
- end
410
-
411
- def self.disconnect
412
- @connection&.disconnect
413
- @connection = nil
414
- end
415
- end
416
-
417
- # Use throughout application
418
- db = DatabaseManager.connection
419
- ```
420
-
421
- ### 2. Transaction Management
422
-
423
- **Optimize transaction usage:**
424
-
425
- ```ruby
426
- # ✅ Group related operations in transactions
427
- db.transaction do
428
- user_id = db[:users].insert(name: 'John', email: 'john@example.com')
429
- profile_id = db[:profiles].insert(user_id: user_id, bio: 'Developer')
430
- db[:preferences].insert(user_id: user_id, theme: 'dark')
431
- end
432
-
433
- # ✅ Use savepoints for nested operations
434
- db.transaction do
435
- user_id = db[:users].insert(name: 'Jane', email: 'jane@example.com')
436
-
437
- begin
438
- db.transaction(savepoint: true) do
439
- # Risky operation that might fail
440
- db[:audit_log].insert(user_id: user_id, action: 'created')
441
- end
442
- rescue Sequel::DatabaseError
443
- # Continue even if audit logging fails
444
- end
445
- end
446
- ```
447
-
448
- ## Monitoring and Profiling
449
-
450
- ### 1. Query Logging
451
-
452
- **Enable comprehensive logging:**
453
-
454
- ```ruby
455
- # ✅ Enable SQL logging
456
- require 'logger'
457
- db.loggers << Logger.new($stdout)
458
-
459
- # ✅ Custom logger with timing
460
- class PerformanceLogger < Logger
461
- def initialize(*args)
462
- super
463
- @start_times = {}
464
- end
465
-
466
- def info(message)
467
- if message.include?('SELECT') || message.include?('INSERT') || message.include?('UPDATE')
468
- @start_time = Time.now
469
- super("SQL: #{message}")
470
- end
471
- end
472
-
473
- def debug(message)
474
- if @start_time && message.include?('rows')
475
- duration = Time.now - @start_time
476
- super("Duration: #{duration.round(3)}s - #{message}")
477
- @start_time = nil
478
- else
479
- super
480
- end
481
- end
482
- end
483
-
484
- db.loggers << PerformanceLogger.new($stdout)
485
- ```
486
-
487
- ### 2. Performance Monitoring
488
-
489
- **Monitor key performance metrics:**
490
-
491
- ```ruby
492
- # ✅ Query performance monitoring
493
- class QueryMonitor
494
- def self.monitor_query(description, &block)
495
- start_time = Time.now
496
- result = yield
497
- duration = Time.now - start_time
498
-
499
- if duration > 1.0 # Log slow queries
500
- puts "SLOW QUERY (#{duration.round(3)}s): #{description}"
501
- end
502
-
503
- result
504
- end
505
- end
506
-
507
- # Usage
508
- users = QueryMonitor.monitor_query("Fetch active users") do
509
- db[:users].where(active: true).all
510
- end
511
- ```
512
-
513
- ### 3. Database Statistics
514
-
515
- **Monitor database performance:**
516
-
517
- ```ruby
518
- # ✅ Check database statistics
519
- def print_db_stats(db)
520
- # Table sizes
521
- puts "Table Statistics:"
522
- db.tables.each do |table|
523
- count = db[table].count
524
- puts " #{table}: #{count} rows"
525
- end
526
-
527
- # Index usage (if available)
528
- puts "\nIndex Information:"
529
- db.tables.each do |table|
530
- indexes = db.indexes(table)
531
- puts " #{table}: #{indexes.keys.join(', ')}" if indexes.any?
532
- end
533
- end
534
-
535
- print_db_stats(db)
536
- ```
537
-
538
- ## Best Practices
539
-
540
- ### 1. Query Design Patterns
541
-
542
- ```ruby
543
- # ✅ Efficient analytical query pattern
544
- def sales_report(db, start_date, end_date)
545
- db[:order_analytics]
546
- .where(order_date: start_date..end_date)
547
- .select(
548
- :category_name,
549
- Sequel.function(:sum, :total_amount).as(:revenue),
550
- Sequel.function(:count, :order_id).as(:order_count),
551
- Sequel.function(:avg, :total_amount).as(:avg_order_value)
552
- )
553
- .group(:category_name)
554
- .order(Sequel.desc(:revenue))
555
- end
556
-
557
- # ✅ Efficient pagination pattern
558
- def paginated_orders(db, page = 1, per_page = 50)
559
- offset = (page - 1) * per_page
560
-
561
- db[:orders]
562
- .select(:id, :user_id, :total, :status, :created_at)
563
- .order(Sequel.desc(:created_at), :id) # Stable sort
564
- .limit(per_page)
565
- .offset(offset)
566
- end
567
- ```
568
-
569
- ### 2. Caching Strategies
570
-
571
- ```ruby
572
- # ✅ Application-level caching
573
- class CachedQueries
574
- def self.user_stats(db, user_id)
575
- @cache ||= {}
576
- cache_key = "user_stats_#{user_id}"
577
-
578
- @cache[cache_key] ||= db[:orders]
579
- .where(user_id: user_id)
580
- .select(
581
- Sequel.function(:count, :id).as(:order_count),
582
- Sequel.function(:sum, :total).as(:total_spent),
583
- Sequel.function(:avg, :total).as(:avg_order)
584
- )
585
- .first
586
- end
587
-
588
- def self.clear_cache
589
- @cache = {}
590
- end
591
- end
592
- ```
593
-
594
- ### 3. Error Handling and Retry Logic
595
-
596
- ```ruby
597
- # ✅ Robust error handling
598
- def execute_with_retry(db, max_retries = 3)
599
- retries = 0
600
-
601
- begin
602
- yield
603
- rescue Sequel::DatabaseConnectionError => e
604
- retries += 1
605
- if retries <= max_retries
606
- sleep(0.1 * retries) # Exponential backoff
607
- retry
608
- else
609
- raise e
610
- end
611
- end
612
- end
613
-
614
- # Usage
615
- result = execute_with_retry(db) do
616
- db[:users].where(active: true).count
617
- end
618
- ```
619
-
620
- ### 4. Development vs Production Optimization
621
-
622
- ```ruby
623
- # ✅ Environment-specific configuration
624
- class DatabaseConfig
625
- def self.connection_options
626
- if ENV['RAILS_ENV'] == 'production'
627
- {
628
- adapter: 'duckdb',
629
- database: '/var/lib/app/production.duckdb',
630
- config: {
631
- memory_limit: '8GB',
632
- threads: 16,
633
- max_memory: '16GB'
634
- },
635
- max_connections: 20,
636
- pool_timeout: 10
637
- }
638
- else
639
- {
640
- adapter: 'duckdb',
641
- database: ':memory:',
642
- config: {
643
- memory_limit: '1GB',
644
- threads: 4
645
- },
646
- max_connections: 5
647
- }
648
- end
649
- end
650
- end
651
-
652
- db = Sequel.connect(DatabaseConfig.connection_options)
653
- ```
654
-
655
- ## Performance Testing
656
-
657
- ### 1. Benchmarking Queries
658
-
659
- ```ruby
660
- require 'benchmark'
661
-
662
- # ✅ Query performance testing
663
- def benchmark_query(description, iterations = 100)
664
- puts "Benchmarking: #{description}"
665
-
666
- time = Benchmark.measure do
667
- iterations.times { yield }
668
- end
669
-
670
- puts " Total time: #{time.real.round(3)}s"
671
- puts " Average: #{(time.real / iterations * 1000).round(3)}ms per query"
672
- puts " Queries/sec: #{(iterations / time.real).round(1)}"
673
- end
674
-
675
- # Usage
676
- benchmark_query("User lookup by email") do
677
- db[:users].where(email: 'test@example.com').first
678
- end
679
-
680
- benchmark_query("Sales aggregation") do
681
- db[:orders].where(status: 'completed').sum(:total)
682
- end
683
- ```
684
-
685
- ### 2. Load Testing
686
-
687
- ```ruby
688
- # ✅ Concurrent load testing
689
- require 'thread'
690
-
691
- def load_test(db, concurrent_users = 10, queries_per_user = 100)
692
- threads = []
693
- results = Queue.new
694
-
695
- concurrent_users.times do |user_id|
696
- threads << Thread.new do
697
- start_time = Time.now
698
-
699
- queries_per_user.times do
700
- # Simulate user queries
701
- db[:users].where(id: rand(1000)).first
702
- db[:orders].where(user_id: rand(1000)).count
703
- end
704
-
705
- duration = Time.now - start_time
706
- results << { user_id: user_id, duration: duration }
707
- end
708
- end
709
-
710
- threads.each(&:join)
711
-
712
- # Collect results
713
- total_queries = concurrent_users * queries_per_user
714
- total_time = results.size.times.map { results.pop[:duration] }.max
715
-
716
- puts "Load test results:"
717
- puts " #{concurrent_users} concurrent users"
718
- puts " #{total_queries} total queries"
719
- puts " #{total_time.round(3)}s total time"
720
- puts " #{(total_queries / total_time).round(1)} queries/sec"
721
- end
722
-
723
- load_test(db)
724
- ```
725
-
726
- This comprehensive performance optimization guide should help you get the most out of DuckDB's analytical capabilities while using Sequel. Remember that DuckDB excels at analytical workloads, so design your queries and schema to take advantage of its columnar storage and vectorized execution engine.