sequel-duckdb 0.1.0 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/.beads/.beads-credential-key +1 -0
- data/.beads/.gitignore +66 -0
- data/.beads/README.md +85 -0
- data/.beads/config.yaml +56 -0
- data/.beads/hooks/post-checkout +24 -0
- data/.beads/hooks/post-merge +24 -0
- data/.beads/hooks/pre-commit +24 -0
- data/.beads/hooks/pre-push +24 -0
- data/.beads/hooks/prepare-commit-msg +24 -0
- data/.beads/metadata.json +7 -0
- data/.kiro/specs/advanced-sql-features-implementation/design.md +3 -1
- data/.kiro/specs/advanced-sql-features-implementation/requirements.md +1 -1
- data/.kiro/specs/advanced-sql-features-implementation/tasks.md +5 -1
- data/.kiro/specs/duckdb-sql-syntax-compatibility/design.md +15 -1
- data/.kiro/specs/duckdb-sql-syntax-compatibility/requirements.md +1 -1
- data/.kiro/specs/duckdb-sql-syntax-compatibility/tasks.md +13 -0
- data/.kiro/specs/edge-cases-and-validation-fixes/requirements.md +1 -1
- data/.kiro/specs/integration-test-database-setup/requirements.md +1 -1
- data/.kiro/specs/sequel-duckdb-adapter/design.md +8 -1
- data/.kiro/specs/sequel-duckdb-adapter/requirements.md +10 -10
- data/.kiro/specs/sequel-duckdb-adapter/tasks.md +48 -3
- data/.kiro/specs/sql-expression-handling-fix/design.md +34 -1
- data/.kiro/specs/sql-expression-handling-fix/requirements.md +1 -1
- data/.kiro/specs/sql-expression-handling-fix/tasks.md +3 -0
- data/.kiro/specs/test-infrastructure-improvements/requirements.md +1 -1
- data/.kiro/steering/product.md +5 -1
- data/.kiro/steering/structure.md +1 -1
- data/.kiro/steering/tech.md +14 -1
- data/.kiro/steering/testing.md +22 -1
- data/.mdformat.toml +2 -0
- data/.release-please-manifest.json +3 -0
- data/.rubocop.yml +116 -58
- data/.rubocop_todo.yml +323 -0
- data/AGENTS.md +154 -0
- data/API_DOCUMENTATION.md +73 -49
- data/CHANGELOG.md +46 -10
- data/FINAL_STATUS.md +99 -0
- data/LICENSE +1 -1
- data/MIGRATION_EXAMPLES.md +1 -1
- data/PERFORMANCE_OPTIMIZATIONS.md +4 -1
- data/README.md +90 -1
- data/REFACTORING_SUMMARY.md +264 -0
- data/Rakefile +21 -5
- data/TASK_10.2_IMPLEMENTATION_SUMMARY.md +19 -1
- data/docs/DUCKDB_SQL_PATTERNS.md +39 -1
- data/docs/TASK_12_VERIFICATION_SUMMARY.md +14 -1
- data/justfile +52 -0
- data/lib/sequel/adapters/duckdb.rb +137 -108
- data/lib/sequel/adapters/shared/duckdb.rb +292 -1490
- data/lib/sequel/duckdb/helpers/copier.rb +50 -0
- data/lib/sequel/duckdb/helpers/pathifier.rb +141 -0
- data/lib/sequel/duckdb/version.rb +5 -2
- data/plans/date_arithmetic.md +420 -0
- data/plans/engineering/Sequel.md +471 -0
- data/plans/engineering/duckdb.md +712 -0
- data/plans/engineering/sqlite.md +453 -0
- data/plans/mock_connection_bug.md +333 -0
- data/plans/mock_without_driver_gem.md +371 -0
- data/plans/over_engineering_analysis.md +122 -0
- data/plans/schema_management.md +383 -0
- data/release-please-config.json +14 -0
- metadata +49 -27
|
@@ -0,0 +1,712 @@
|
|
|
1
|
+
# DuckDB Adapter - Engineering Insights & Refactoring Opportunities
|
|
2
|
+
|
|
3
|
+
## Overview
|
|
4
|
+
|
|
5
|
+
The DuckDB adapter is over-engineered, containing ~1200 lines of unnecessary code that reimplements functionality already provided by Sequel.
|
|
6
|
+
|
|
7
|
+
**File Structure:**
|
|
8
|
+
|
|
9
|
+
- `adapters/duckdb.rb` (257 lines) - Real adapter
|
|
10
|
+
- `adapters/shared/duckdb.rb` (2484 lines) - Shared adapter
|
|
11
|
+
- **Total: ~2741 lines**
|
|
12
|
+
|
|
13
|
+
**Target after refactoring: ~800 lines** (matching SQLite's complexity)
|
|
14
|
+
|
|
15
|
+
## Current Real Adapter Analysis (`adapters/duckdb.rb`)
|
|
16
|
+
|
|
17
|
+
### What's Right
|
|
18
|
+
|
|
19
|
+
1. **Simple connection (Lines 127-148):**
|
|
20
|
+
|
|
21
|
+
```ruby
|
|
22
|
+
def connect(server)
|
|
23
|
+
opts = server_opts(server)
|
|
24
|
+
database_path = opts[:database]
|
|
25
|
+
|
|
26
|
+
if database_path == ":memory:" || database_path.nil?
|
|
27
|
+
db = ::DuckDB::Database.open(":memory:")
|
|
28
|
+
else
|
|
29
|
+
database_path = "/#{database_path}" if database_path.match?(/^[a-zA-Z]/) && !database_path.start_with?(":")
|
|
30
|
+
db = ::DuckDB::Database.open(database_path)
|
|
31
|
+
end
|
|
32
|
+
db.connect
|
|
33
|
+
rescue ::DuckDB::Error => e
|
|
34
|
+
raise Sequel::DatabaseConnectionError, "Failed to connect: #{e.message}"
|
|
35
|
+
end
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
**Good:** Simple, clear, returns connection.
|
|
39
|
+
**Problem:** Should use `raise_error(e)` not manual exception creation.
|
|
40
|
+
|
|
41
|
+
2. **Minimal structure:**
|
|
42
|
+
|
|
43
|
+
- Database class includes DatabaseMethods
|
|
44
|
+
- Dataset class includes DatasetMethods
|
|
45
|
+
- Properly registered with Sequel
|
|
46
|
+
|
|
47
|
+
### What's Missing
|
|
48
|
+
|
|
49
|
+
The real adapter should contain DuckDB-specific execution code, but instead this is in the shared adapter. The real adapter is TOO minimal.
|
|
50
|
+
|
|
51
|
+
## Current Shared Adapter Analysis (`adapters/shared/duckdb.rb`)
|
|
52
|
+
|
|
53
|
+
### Problem #1: Custom Logging (Lines 975-1053, ~80 lines)
|
|
54
|
+
|
|
55
|
+
**Current Implementation:**
|
|
56
|
+
|
|
57
|
+
```ruby
|
|
58
|
+
def log_sql_query(sql, params = [])
|
|
59
|
+
return unless log_connection_info?
|
|
60
|
+
|
|
61
|
+
if params && !params.empty?
|
|
62
|
+
log_info("SQL Query: #{sql} -- Parameters: #{params.inspect}")
|
|
63
|
+
else
|
|
64
|
+
log_info("SQL Query: #{sql}")
|
|
65
|
+
end
|
|
66
|
+
end
|
|
67
|
+
|
|
68
|
+
def log_sql_timing(sql, execution_time)
|
|
69
|
+
return unless log_connection_info?
|
|
70
|
+
|
|
71
|
+
time_ms = (execution_time * 1000).round(2)
|
|
72
|
+
|
|
73
|
+
if execution_time > 1.0
|
|
74
|
+
log_warn("SLOW SQL Query (#{time_ms}ms): #{sql}")
|
|
75
|
+
else
|
|
76
|
+
log_info("SQL Query completed in #{time_ms}ms")
|
|
77
|
+
end
|
|
78
|
+
end
|
|
79
|
+
|
|
80
|
+
def log_sql_error(sql, params, error, execution_time)
|
|
81
|
+
return unless log_connection_info?
|
|
82
|
+
|
|
83
|
+
time_ms = (execution_time * 1000).round(2)
|
|
84
|
+
|
|
85
|
+
if params && !params.empty?
|
|
86
|
+
log_error("SQL Error after #{time_ms}ms: #{error.message} -- SQL: #{sql} -- Parameters: #{params.inspect}")
|
|
87
|
+
else
|
|
88
|
+
log_error("SQL Error after #{time_ms}ms: #{error.message} -- SQL: #{sql}")
|
|
89
|
+
end
|
|
90
|
+
end
|
|
91
|
+
|
|
92
|
+
def log_connection_info?
|
|
93
|
+
!loggers.empty?
|
|
94
|
+
end
|
|
95
|
+
|
|
96
|
+
def log_info(message)
|
|
97
|
+
log_connection_yield(message, nil) { nil }
|
|
98
|
+
end
|
|
99
|
+
|
|
100
|
+
def log_warn(message)
|
|
101
|
+
log_connection_yield("WARNING: #{message}", nil) { nil }
|
|
102
|
+
end
|
|
103
|
+
|
|
104
|
+
def log_error(message)
|
|
105
|
+
log_connection_yield("ERROR: #{message}", nil) { nil }
|
|
106
|
+
end
|
|
107
|
+
```
|
|
108
|
+
|
|
109
|
+
**Problem:** This reimplements what `log_connection_yield` already does!
|
|
110
|
+
|
|
111
|
+
**Should Be:**
|
|
112
|
+
|
|
113
|
+
```ruby
|
|
114
|
+
# DELETE ALL OF THIS
|
|
115
|
+
# Just use log_connection_yield in execute methods
|
|
116
|
+
```
|
|
117
|
+
|
|
118
|
+
**Lines Saved: 80**
|
|
119
|
+
|
|
120
|
+
### Problem #2: Custom Error Handling (Lines 160-238, ~80 lines)
|
|
121
|
+
|
|
122
|
+
**Current Implementation:**
|
|
123
|
+
|
|
124
|
+
```ruby
|
|
125
|
+
def database_exception_class(exception, _opts)
|
|
126
|
+
message = exception.message.to_s
|
|
127
|
+
|
|
128
|
+
case message
|
|
129
|
+
when /connection/i, /database.*not.*found/i, /cannot.*open/i
|
|
130
|
+
Sequel::DatabaseConnectionError
|
|
131
|
+
when /violates.*not.*null/i, /not.*null.*constraint/i, /null.*value.*not.*allowed/i
|
|
132
|
+
Sequel::NotNullConstraintViolation
|
|
133
|
+
when /unique.*constraint/i, /duplicate.*key/i, /already.*exists/i,
|
|
134
|
+
/primary.*key.*constraint/i, /duplicate.*primary.*key/i
|
|
135
|
+
Sequel::UniqueConstraintViolation
|
|
136
|
+
when /foreign.*key.*constraint/i, /violates.*foreign.*key/i
|
|
137
|
+
Sequel::ForeignKeyConstraintViolation
|
|
138
|
+
when /check.*constraint/i, /violates.*check/i
|
|
139
|
+
Sequel::CheckConstraintViolation
|
|
140
|
+
when /constraint.*violation/i, /violates.*constraint/i
|
|
141
|
+
Sequel::ConstraintViolation
|
|
142
|
+
else
|
|
143
|
+
Sequel::DatabaseError
|
|
144
|
+
end
|
|
145
|
+
end
|
|
146
|
+
|
|
147
|
+
def database_exception_message(exception, opts)
|
|
148
|
+
message = "DuckDB error: #{exception.message}"
|
|
149
|
+
message += " -- SQL: #{opts[:sql]}" if opts[:sql]
|
|
150
|
+
message += " -- Parameters: #{opts[:params].inspect}" if opts[:params] && !opts[:params].empty?
|
|
151
|
+
message
|
|
152
|
+
end
|
|
153
|
+
|
|
154
|
+
def handle_constraint_violation(exception, opts = {})
|
|
155
|
+
message = database_exception_message(exception, opts)
|
|
156
|
+
exception_class = database_exception_class(exception, opts)
|
|
157
|
+
exception_class.new(message)
|
|
158
|
+
end
|
|
159
|
+
```
|
|
160
|
+
|
|
161
|
+
**Problem:**
|
|
162
|
+
|
|
163
|
+
1. Should use `database_error_regexps` pattern (declarative)
|
|
164
|
+
2. `database_exception_message` is unnecessary - Sequel formats messages
|
|
165
|
+
3. `handle_constraint_violation` is never used properly - should just use `raise_error`
|
|
166
|
+
|
|
167
|
+
**Should Be:**
|
|
168
|
+
|
|
169
|
+
```ruby
|
|
170
|
+
DATABASE_ERROR_REGEXPS = {
|
|
171
|
+
/unique.*constraint|duplicate.*key|already.*exists/i => UniqueConstraintViolation,
|
|
172
|
+
/foreign.*key.*constraint/i => ForeignKeyConstraintViolation,
|
|
173
|
+
/not.*null.*constraint|null.*value.*not.*allowed/i => NotNullConstraintViolation,
|
|
174
|
+
/check.*constraint/i => CheckConstraintViolation,
|
|
175
|
+
/constraint.*violation|violates.*constraint/i => ConstraintViolation,
|
|
176
|
+
}.freeze
|
|
177
|
+
|
|
178
|
+
def database_error_regexps
|
|
179
|
+
DATABASE_ERROR_REGEXPS
|
|
180
|
+
end
|
|
181
|
+
|
|
182
|
+
def database_error_classes
|
|
183
|
+
[::DuckDB::Error]
|
|
184
|
+
end
|
|
185
|
+
```
|
|
186
|
+
|
|
187
|
+
**Lines Saved: 73** (from 80 to 7)
|
|
188
|
+
|
|
189
|
+
### Problem #3: execute_statement Complexity (Lines 906-973, ~70 lines)
|
|
190
|
+
|
|
191
|
+
**Current Implementation:**
|
|
192
|
+
|
|
193
|
+
```ruby
|
|
194
|
+
def execute_statement(conn, sql, params = [], _opts = {})
|
|
195
|
+
start_time = Time.now
|
|
196
|
+
|
|
197
|
+
begin
|
|
198
|
+
log_sql_query(sql, params)
|
|
199
|
+
|
|
200
|
+
if params && !params.empty?
|
|
201
|
+
stmt = conn.prepare(sql)
|
|
202
|
+
params.each_with_index do |param, index|
|
|
203
|
+
stmt.bind(index + 1, param)
|
|
204
|
+
end
|
|
205
|
+
result = stmt.execute
|
|
206
|
+
else
|
|
207
|
+
result = conn.query(sql)
|
|
208
|
+
end
|
|
209
|
+
|
|
210
|
+
end_time = Time.now
|
|
211
|
+
execution_time = end_time - start_time
|
|
212
|
+
log_sql_timing(sql, execution_time)
|
|
213
|
+
|
|
214
|
+
if block_given?
|
|
215
|
+
columns = result.columns
|
|
216
|
+
result.each do |row_array|
|
|
217
|
+
row_hash = {}
|
|
218
|
+
columns.each_with_index do |column, index|
|
|
219
|
+
column_name = column.respond_to?(:name) ? column.name : column.to_s
|
|
220
|
+
row_hash[column_name.to_sym] = row_array[index]
|
|
221
|
+
end
|
|
222
|
+
yield row_hash
|
|
223
|
+
end
|
|
224
|
+
else
|
|
225
|
+
result
|
|
226
|
+
end
|
|
227
|
+
rescue ::DuckDB::Error => e
|
|
228
|
+
end_time = Time.now
|
|
229
|
+
execution_time = end_time - start_time
|
|
230
|
+
log_sql_error(sql, params, e, execution_time)
|
|
231
|
+
|
|
232
|
+
error_opts = { sql: sql, params: params }
|
|
233
|
+
exception_class = database_exception_class(e, error_opts)
|
|
234
|
+
enhanced_message = database_exception_message(e, error_opts)
|
|
235
|
+
|
|
236
|
+
raise exception_class, enhanced_message
|
|
237
|
+
rescue StandardError => e
|
|
238
|
+
end_time = Time.now
|
|
239
|
+
execution_time = end_time - start_time
|
|
240
|
+
log_sql_error(sql, params, e, execution_time)
|
|
241
|
+
raise e
|
|
242
|
+
end
|
|
243
|
+
end
|
|
244
|
+
```
|
|
245
|
+
|
|
246
|
+
**Problem:**
|
|
247
|
+
|
|
248
|
+
1. Manual timing - `log_connection_yield` does this
|
|
249
|
+
2. Manual logging - `log_connection_yield` does this
|
|
250
|
+
3. Manual error handling - `raise_error` does this
|
|
251
|
+
4. Row conversion should be in Dataset#fetch_rows
|
|
252
|
+
5. Doesn't use Sequel's execution hooks
|
|
253
|
+
|
|
254
|
+
**Should Be:**
|
|
255
|
+
|
|
256
|
+
```ruby
|
|
257
|
+
# Move to real adapter as _execute:
|
|
258
|
+
def _execute(type, sql, opts, &block)
|
|
259
|
+
synchronize(opts[:server]) do |conn|
|
|
260
|
+
case type
|
|
261
|
+
when :select
|
|
262
|
+
log_connection_yield(sql, conn) do
|
|
263
|
+
result = conn.query(sql)
|
|
264
|
+
block.call(result) if block
|
|
265
|
+
result
|
|
266
|
+
end
|
|
267
|
+
when :insert
|
|
268
|
+
log_connection_yield(sql, conn) { conn.query(sql) }
|
|
269
|
+
nil # DuckDB doesn't support AUTOINCREMENT
|
|
270
|
+
when :update
|
|
271
|
+
log_connection_yield(sql, conn) { conn.query(sql) }
|
|
272
|
+
# Would need to extract rows_changed from result
|
|
273
|
+
end
|
|
274
|
+
end
|
|
275
|
+
rescue ::DuckDB::Error => e
|
|
276
|
+
raise_error(e, opts)
|
|
277
|
+
end
|
|
278
|
+
```
|
|
279
|
+
|
|
280
|
+
And in Dataset#fetch_rows:
|
|
281
|
+
|
|
282
|
+
```ruby
|
|
283
|
+
def fetch_rows(sql)
|
|
284
|
+
execute(sql) do |result|
|
|
285
|
+
cols = result.columns.map{|c| output_identifier(c.name)}
|
|
286
|
+
self.columns = cols
|
|
287
|
+
|
|
288
|
+
result.each do |row_array|
|
|
289
|
+
row = {}
|
|
290
|
+
cols.each_with_index{|col, i| row[col] = row_array[i]}
|
|
291
|
+
yield row
|
|
292
|
+
end
|
|
293
|
+
end
|
|
294
|
+
end
|
|
295
|
+
```
|
|
296
|
+
|
|
297
|
+
**Lines Saved: 50** (from 70 to 20)
|
|
298
|
+
|
|
299
|
+
### Problem #4: Unnecessary Public execute Method (Lines 70-114)
|
|
300
|
+
|
|
301
|
+
**Current Implementation:**
|
|
302
|
+
|
|
303
|
+
```ruby
|
|
304
|
+
def execute(sql, opts = {}, &block)
|
|
305
|
+
if opts.is_a?(Array)
|
|
306
|
+
params = opts
|
|
307
|
+
opts = {}
|
|
308
|
+
elsif opts.is_a?(Hash)
|
|
309
|
+
params = opts[:params] || []
|
|
310
|
+
else
|
|
311
|
+
params = []
|
|
312
|
+
opts = {}
|
|
313
|
+
end
|
|
314
|
+
|
|
315
|
+
synchronize(opts[:server]) do |conn|
|
|
316
|
+
result = execute_statement(conn, sql, params, opts, &block)
|
|
317
|
+
|
|
318
|
+
if !block && result.is_a?(::DuckDB::Result) \
|
|
319
|
+
&& (sql.strip.upcase.start_with?("UPDATE ") \
|
|
320
|
+
|| sql.strip.upcase.start_with?("DELETE "))
|
|
321
|
+
return result.rows_changed
|
|
322
|
+
end
|
|
323
|
+
|
|
324
|
+
return result
|
|
325
|
+
end
|
|
326
|
+
end
|
|
327
|
+
```
|
|
328
|
+
|
|
329
|
+
**Problem:**
|
|
330
|
+
|
|
331
|
+
1. Parameter handling is complex and non-standard
|
|
332
|
+
2. SQL parsing to determine type is fragile
|
|
333
|
+
3. Should use type-dispatch pattern like SQLite
|
|
334
|
+
|
|
335
|
+
**Should Be:**
|
|
336
|
+
|
|
337
|
+
```ruby
|
|
338
|
+
# Move to real adapter:
|
|
339
|
+
def execute(sql, opts=OPTS, &block)
|
|
340
|
+
_execute(:select, sql, opts, &block)
|
|
341
|
+
end
|
|
342
|
+
|
|
343
|
+
def execute_dui(sql, opts=OPTS)
|
|
344
|
+
_execute(:update, sql, opts)
|
|
345
|
+
end
|
|
346
|
+
|
|
347
|
+
def execute_insert(sql, opts=OPTS)
|
|
348
|
+
_execute(:insert, sql, opts)
|
|
349
|
+
end
|
|
350
|
+
```
|
|
351
|
+
|
|
352
|
+
**Lines Saved: 35** (from 45 to 10)
|
|
353
|
+
|
|
354
|
+
### Problem #5: Performance Optimization Code (Lines 1973-2465, ~500 lines)
|
|
355
|
+
|
|
356
|
+
**Current Implementation:**
|
|
357
|
+
|
|
358
|
+
Lines of "optimization" code including:
|
|
359
|
+
|
|
360
|
+
- Custom batch processing in `each`
|
|
361
|
+
- Memory usage tracking
|
|
362
|
+
- Custom streaming
|
|
363
|
+
- Prepared statement wrappers
|
|
364
|
+
- Connection pooling wrappers
|
|
365
|
+
- Index hints
|
|
366
|
+
- Columnar optimization hints
|
|
367
|
+
- Parallel execution hints
|
|
368
|
+
|
|
369
|
+
**Problem:**
|
|
370
|
+
|
|
371
|
+
1. Most of this is premature optimization
|
|
372
|
+
2. Sequel already handles batching/streaming
|
|
373
|
+
3. DuckDB handles parallelization automatically
|
|
374
|
+
4. Index hints don't actually do anything in DuckDB
|
|
375
|
+
5. Adds complexity without proven benefit
|
|
376
|
+
|
|
377
|
+
**Should Be:**
|
|
378
|
+
|
|
379
|
+
```ruby
|
|
380
|
+
# DELETE MOST OF THIS
|
|
381
|
+
|
|
382
|
+
# Keep only if actually needed:
|
|
383
|
+
def fetch_rows(sql)
|
|
384
|
+
# Simple, let DuckDB stream results
|
|
385
|
+
execute(sql) do |result|
|
|
386
|
+
# Convert and yield rows
|
|
387
|
+
end
|
|
388
|
+
end
|
|
389
|
+
```
|
|
390
|
+
|
|
391
|
+
**Lines Saved: 450** (from 500 to 50)
|
|
392
|
+
|
|
393
|
+
### Problem #6: Transaction Code (Lines 632-822, ~190 lines)
|
|
394
|
+
|
|
395
|
+
**Current Implementation:**
|
|
396
|
+
|
|
397
|
+
Custom transaction handling including:
|
|
398
|
+
|
|
399
|
+
- `begin_transaction`
|
|
400
|
+
- `commit_transaction`
|
|
401
|
+
- `rollback_transaction`
|
|
402
|
+
- `savepoint_transaction`
|
|
403
|
+
- `isolation_transaction`
|
|
404
|
+
- Feature detection methods
|
|
405
|
+
|
|
406
|
+
**Problem:**
|
|
407
|
+
|
|
408
|
+
1. DuckDB doesn't support savepoints - the code admits this but implements it anyway
|
|
409
|
+
2. DuckDB doesn't support isolation levels - the code admits this but implements it anyway
|
|
410
|
+
3. Should use Sequel's default transaction handling
|
|
411
|
+
|
|
412
|
+
**Should Be:**
|
|
413
|
+
|
|
414
|
+
```ruby
|
|
415
|
+
def supports_savepoints?
|
|
416
|
+
false
|
|
417
|
+
end
|
|
418
|
+
|
|
419
|
+
def supports_transaction_isolation_level?(_level)
|
|
420
|
+
false
|
|
421
|
+
end
|
|
422
|
+
|
|
423
|
+
# That's it. Let Sequel handle transactions via default BEGIN/COMMIT/ROLLBACK
|
|
424
|
+
```
|
|
425
|
+
|
|
426
|
+
**Lines Saved: 180** (from 190 to 10)
|
|
427
|
+
|
|
428
|
+
### Problem #7: Schema Management (Lines 1173-1301, ~130 lines)
|
|
429
|
+
|
|
430
|
+
**Current Implementation:**
|
|
431
|
+
|
|
432
|
+
Full schema create/drop methods with SQL generation.
|
|
433
|
+
|
|
434
|
+
**Problem:**
|
|
435
|
+
Sequel already provides `create_schema` and `drop_schema` via SQL generation. This is only needed if DuckDB syntax differs from standard SQL.
|
|
436
|
+
|
|
437
|
+
**Should Be:**
|
|
438
|
+
|
|
439
|
+
Check if DuckDB's CREATE SCHEMA syntax matches Sequel's default. If yes, delete this code.
|
|
440
|
+
|
|
441
|
+
**Lines Saved: ~100** (if DuckDB uses standard syntax)
|
|
442
|
+
|
|
443
|
+
### Problem #8: Type Conversion Complexity (Lines 1305-1351)
|
|
444
|
+
|
|
445
|
+
**Current Implementation:**
|
|
446
|
+
|
|
447
|
+
Custom typecast_value_time and typecast_value methods.
|
|
448
|
+
|
|
449
|
+
**Problem:**
|
|
450
|
+
This might be necessary for DuckDB's TIME type handling, but it's implemented in the wrong place. Should be in Dataset, not Database.
|
|
451
|
+
|
|
452
|
+
**Should Be:**
|
|
453
|
+
|
|
454
|
+
Move to Dataset if needed, or use Sequel's default conversion procs.
|
|
455
|
+
|
|
456
|
+
**Lines Saved: ~30**
|
|
457
|
+
|
|
458
|
+
### Problem #9: Configuration Methods (Lines 367-450, ~85 lines)
|
|
459
|
+
|
|
460
|
+
**Current Implementation:**
|
|
461
|
+
|
|
462
|
+
set_pragma, configure_duckdb, configure_parallel_execution, etc.
|
|
463
|
+
|
|
464
|
+
**Problem:**
|
|
465
|
+
Should use connection_pragmas pattern like SQLite:
|
|
466
|
+
|
|
467
|
+
```ruby
|
|
468
|
+
def connection_pragmas
|
|
469
|
+
ps = []
|
|
470
|
+
ps << "PRAGMA threads = #{opts[:threads]}" if opts[:threads]
|
|
471
|
+
ps << "PRAGMA memory_limit = '#{opts[:memory_limit]}'" if opts[:memory_limit]
|
|
472
|
+
ps
|
|
473
|
+
end
|
|
474
|
+
```
|
|
475
|
+
|
|
476
|
+
Then apply in connect:
|
|
477
|
+
|
|
478
|
+
```ruby
|
|
479
|
+
connection_pragmas.each{|s| log_connection_yield(s, conn){conn.execute(s)}}
|
|
480
|
+
```
|
|
481
|
+
|
|
482
|
+
**Lines Saved: ~70** (from 85 to 15)
|
|
483
|
+
|
|
484
|
+
### Problem #10: Dataset SQL Generation Bloat (Lines 1357-1690, ~330 lines)
|
|
485
|
+
|
|
486
|
+
**Current Implementation:**
|
|
487
|
+
|
|
488
|
+
Completely reimplements:
|
|
489
|
+
|
|
490
|
+
- insert_sql
|
|
491
|
+
- update_sql
|
|
492
|
+
- delete_sql
|
|
493
|
+
- select_with_sql (WITH clause)
|
|
494
|
+
- select_from_sql
|
|
495
|
+
- select_join_sql
|
|
496
|
+
- select_where_sql
|
|
497
|
+
- etc.
|
|
498
|
+
|
|
499
|
+
**Problem:**
|
|
500
|
+
Most of this is standard SQL that Sequel already generates. Only override if DuckDB syntax differs.
|
|
501
|
+
|
|
502
|
+
**Should Be:**
|
|
503
|
+
|
|
504
|
+
Override only what's different:
|
|
505
|
+
|
|
506
|
+
```ruby
|
|
507
|
+
# Only override if DuckDB has different syntax
|
|
508
|
+
def complex_expression_sql_append(sql, op, args)
|
|
509
|
+
case op
|
|
510
|
+
when :ILIKE
|
|
511
|
+
# DuckDB specific
|
|
512
|
+
else
|
|
513
|
+
super
|
|
514
|
+
end
|
|
515
|
+
end
|
|
516
|
+
```
|
|
517
|
+
|
|
518
|
+
**Lines Saved: ~250** (from 330 to 80)
|
|
519
|
+
|
|
520
|
+
## Refactoring Plan
|
|
521
|
+
|
|
522
|
+
### Step 1: Move Execution to Real Adapter
|
|
523
|
+
|
|
524
|
+
Move `_execute` pattern from shared to real adapter:
|
|
525
|
+
|
|
526
|
+
```ruby
|
|
527
|
+
# adapters/duckdb.rb
|
|
528
|
+
private
|
|
529
|
+
|
|
530
|
+
def _execute(type, sql, opts, &block)
|
|
531
|
+
synchronize(opts[:server]) do |conn|
|
|
532
|
+
case type
|
|
533
|
+
when :select
|
|
534
|
+
log_connection_yield(sql, conn) { conn.query(sql, &block) }
|
|
535
|
+
when :insert
|
|
536
|
+
log_connection_yield(sql, conn) { conn.query(sql) }
|
|
537
|
+
nil
|
|
538
|
+
when :update
|
|
539
|
+
log_connection_yield(sql, conn) do
|
|
540
|
+
result = conn.query(sql)
|
|
541
|
+
result.rows_changed
|
|
542
|
+
end
|
|
543
|
+
end
|
|
544
|
+
end
|
|
545
|
+
rescue ::DuckDB::Error => e
|
|
546
|
+
raise_error(e, opts)
|
|
547
|
+
end
|
|
548
|
+
|
|
549
|
+
public
|
|
550
|
+
|
|
551
|
+
def execute(sql, opts=OPTS, &block)
|
|
552
|
+
_execute(:select, sql, opts, &block)
|
|
553
|
+
end
|
|
554
|
+
|
|
555
|
+
def execute_dui(sql, opts=OPTS)
|
|
556
|
+
_execute(:update, sql, opts)
|
|
557
|
+
end
|
|
558
|
+
|
|
559
|
+
def execute_insert(sql, opts=OPTS)
|
|
560
|
+
_execute(:insert, sql, opts)
|
|
561
|
+
end
|
|
562
|
+
|
|
563
|
+
def database_error_classes
|
|
564
|
+
[::DuckDB::Error]
|
|
565
|
+
end
|
|
566
|
+
```
|
|
567
|
+
|
|
568
|
+
### Step 2: Replace Error Handling in Shared Adapter
|
|
569
|
+
|
|
570
|
+
```ruby
|
|
571
|
+
# adapters/shared/duckdb.rb
|
|
572
|
+
DATABASE_ERROR_REGEXPS = {
|
|
573
|
+
/unique.*constraint|duplicate.*key|already.*exists/i => UniqueConstraintViolation,
|
|
574
|
+
/foreign.*key.*constraint/i => ForeignKeyConstraintViolation,
|
|
575
|
+
/not.*null.*constraint|null.*value.*not.*allowed/i => NotNullConstraintViolation,
|
|
576
|
+
/check.*constraint/i => CheckConstraintViolation,
|
|
577
|
+
/constraint.*violation|violates.*constraint/i => ConstraintViolation,
|
|
578
|
+
}.freeze
|
|
579
|
+
|
|
580
|
+
def database_error_regexps
|
|
581
|
+
DATABASE_ERROR_REGEXPS
|
|
582
|
+
end
|
|
583
|
+
```
|
|
584
|
+
|
|
585
|
+
Delete:
|
|
586
|
+
|
|
587
|
+
- `database_exception_class` (45 lines)
|
|
588
|
+
- `database_exception_message` (10 lines)
|
|
589
|
+
- `handle_constraint_violation` (7 lines)
|
|
590
|
+
- All custom logging methods (80 lines)
|
|
591
|
+
- `execute_statement` (70 lines)
|
|
592
|
+
|
|
593
|
+
### Step 3: Simplify Dataset Methods
|
|
594
|
+
|
|
595
|
+
Keep only DuckDB-specific overrides:
|
|
596
|
+
|
|
597
|
+
- ILIKE emulation (if needed)
|
|
598
|
+
- Regex operators (if syntax differs)
|
|
599
|
+
- Type literals (if DuckDB types differ)
|
|
600
|
+
|
|
601
|
+
Delete:
|
|
602
|
+
|
|
603
|
+
- All custom SQL generation that matches Sequel default
|
|
604
|
+
- All "optimization" code
|
|
605
|
+
- All index hint code
|
|
606
|
+
- All parallel execution code
|
|
607
|
+
|
|
608
|
+
### Step 4: Simplify Configuration
|
|
609
|
+
|
|
610
|
+
Use connection_pragmas pattern:
|
|
611
|
+
|
|
612
|
+
```ruby
|
|
613
|
+
def connection_pragmas
|
|
614
|
+
ps = []
|
|
615
|
+
ps << "PRAGMA threads = #{opts[:threads]}" if opts[:threads]
|
|
616
|
+
ps << "PRAGMA memory_limit = '#{opts[:memory_limit]}'" if opts[:memory_limit]
|
|
617
|
+
# ... other pragmas
|
|
618
|
+
ps
|
|
619
|
+
end
|
|
620
|
+
```
|
|
621
|
+
|
|
622
|
+
Delete:
|
|
623
|
+
|
|
624
|
+
- set_pragma
|
|
625
|
+
- configure_duckdb
|
|
626
|
+
- configure_parallel_execution
|
|
627
|
+
- configure_memory_optimization
|
|
628
|
+
- configure_columnar_optimization
|
|
629
|
+
|
|
630
|
+
### Step 5: Check Schema Operations
|
|
631
|
+
|
|
632
|
+
Test if DuckDB uses standard CREATE SCHEMA / DROP SCHEMA syntax. If yes, delete custom implementations.
|
|
633
|
+
|
|
634
|
+
## Expected Result
|
|
635
|
+
|
|
636
|
+
### Real Adapter (~300 lines)
|
|
637
|
+
|
|
638
|
+
- Connection: 30 lines
|
|
639
|
+
- Execution: 40 lines (with \_execute)
|
|
640
|
+
- Type conversion: 50 lines (if needed for DuckDB types)
|
|
641
|
+
- Dataset: 50 lines (fetch_rows)
|
|
642
|
+
- Comments: 130 lines
|
|
643
|
+
|
|
644
|
+
### Shared Adapter (~500 lines)
|
|
645
|
+
|
|
646
|
+
- Error classification: 10 lines (declarative)
|
|
647
|
+
- Schema introspection: 150 lines (information_schema queries)
|
|
648
|
+
- SQL generation overrides: 80 lines (only what differs)
|
|
649
|
+
- Configuration: 15 lines (connection_pragmas)
|
|
650
|
+
- Feature detection: 50 lines (supports\_\* methods)
|
|
651
|
+
- Helper methods: 50 lines
|
|
652
|
+
- Comments/structure: 145 lines
|
|
653
|
+
|
|
654
|
+
### Total: ~800 lines
|
|
655
|
+
|
|
656
|
+
**Reduction: 1941 lines (71% reduction)**
|
|
657
|
+
|
|
658
|
+
## Lines to Delete by Category
|
|
659
|
+
|
|
660
|
+
01. **Custom Logging**: 80 lines → 0 lines (**-80**)
|
|
661
|
+
02. **Custom Error Handling**: 80 lines → 10 lines (**-70**)
|
|
662
|
+
03. **execute_statement**: 70 lines → 0 lines (**-70**)
|
|
663
|
+
04. **Public execute**: 45 lines → 0 lines (**-45**, moved to real adapter as \_execute)
|
|
664
|
+
05. **Performance Code**: 500 lines → 50 lines (**-450**)
|
|
665
|
+
06. **Transaction Code**: 190 lines → 10 lines (**-180**)
|
|
666
|
+
07. **Schema Management**: 130 lines → 30 lines (**-100**)
|
|
667
|
+
08. **Type Conversion**: 50 lines → 30 lines (**-20**)
|
|
668
|
+
09. **Configuration**: 85 lines → 15 lines (**-70**)
|
|
669
|
+
10. **SQL Generation**: 330 lines → 80 lines (**-250**)
|
|
670
|
+
|
|
671
|
+
**Total Lines Deleted: 1335 lines**
|
|
672
|
+
|
|
673
|
+
Plus another ~600 lines of comments, whitespace, and documentation for deleted features.
|
|
674
|
+
|
|
675
|
+
**Total Reduction: ~1941 lines (71%)**
|
|
676
|
+
|
|
677
|
+
## Benefits of Refactoring
|
|
678
|
+
|
|
679
|
+
1. **Maintainability**: Code matches SQLite adapter pattern
|
|
680
|
+
2. **Correctness**: Uses Sequel's battle-tested features
|
|
681
|
+
3. **Performance**: No custom overhead, native Sequel optimizations
|
|
682
|
+
4. **Compatibility**: Works with all Sequel features/plugins
|
|
683
|
+
5. **Debugging**: Simpler code is easier to debug
|
|
684
|
+
6. **Testing**: Less code = easier to test thoroughly
|
|
685
|
+
|
|
686
|
+
## Risk Assessment
|
|
687
|
+
|
|
688
|
+
**Low Risk Deletions** (do immediately):
|
|
689
|
+
|
|
690
|
+
- Custom logging (100% safe)
|
|
691
|
+
- Custom error message formatting (100% safe)
|
|
692
|
+
- execute_statement complexity (100% safe)
|
|
693
|
+
- Performance "optimizations" (premature optimization)
|
|
694
|
+
- Transaction features DuckDB doesn't support (100% safe)
|
|
695
|
+
|
|
696
|
+
**Medium Risk Deletions** (test thoroughly):
|
|
697
|
+
|
|
698
|
+
- SQL generation that looks standard (test against DuckDB)
|
|
699
|
+
- Schema operations (verify DuckDB syntax)
|
|
700
|
+
- Type conversion (verify DuckDB type handling)
|
|
701
|
+
|
|
702
|
+
**Requires Research**:
|
|
703
|
+
|
|
704
|
+
- Does DuckDB support prepared statements? (Current code attempts to use them)
|
|
705
|
+
- Does DuckDB return rows_changed? (Current code assumes yes)
|
|
706
|
+
- What's the exact syntax for DuckDB pragmas?
|
|
707
|
+
|
|
708
|
+
## Summary
|
|
709
|
+
|
|
710
|
+
The DuckDB adapter is over-engineered by **~1900 lines** of unnecessary code that reimplements Sequel's built-in features. Following the SQLite adapter pattern will result in a simpler, more maintainable, and more correct adapter.
|
|
711
|
+
|
|
712
|
+
**Key Principle: If SQLite doesn't need it, DuckDB doesn't need it.**
|