dwh 0.5.0 → 0.6.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +16 -0
- data/README.md +19 -3
- data/docs/guides/adapters.md +16 -0
- data/lib/dwh/adapters/big_query.rb +173 -0
- data/lib/dwh/adapters/duck_db.rb +6 -5
- data/lib/dwh/column.rb +2 -2
- data/lib/dwh/factory.rb +53 -30
- data/lib/dwh/settings/bigquery.yml +92 -0
- data/lib/dwh/version.rb +1 -1
- data/lib/dwh.rb +2 -0
- metadata +3 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: ec16e1365da69bdc1778ff349775274d8b3a6d85978a9a69be0d6aa0da44d866
|
|
4
|
+
data.tar.gz: 4fc4ed0d62f7e657f80779265fec1945317cbdbd127e43c654916f2f0c157f88
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: '008e23f7a38eae74baffc20f04b14daf6344dc5e3b9ae8616926fa74fe4bab393b4a21134755dcfc289e999ec967c2c5a2baa98044834ce1f415d6d179046ab4'
|
|
7
|
+
data.tar.gz: 2febcb2fb010b298d70b2e010dc1d2be942f63bacc4ef7e57eaf63e5fb616c27472037687ad144c706478ce80be66a7445f9479d9e774f78683ca9d6a9dc3732
|
data/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,21 @@
|
|
|
1
1
|
## [Unreleased]
|
|
2
2
|
|
|
3
|
+
## [0.6.0] - 2026-09-28
|
|
4
|
+
|
|
5
|
+
### Added
|
|
6
|
+
|
|
7
|
+
- Google BigQuery adapter (`:bigquery`) with dedicated settings and unit/system test coverage. Authenticates with a service-account keyfile or Application Default Credentials; requires the `google-cloud-bigquery` gem.
|
|
8
|
+
|
|
9
|
+
## [0.5.1] - 2026-08-03
|
|
10
|
+
|
|
11
|
+
### Fixed
|
|
12
|
+
|
|
13
|
+
- **Factory#shutdown**: `case pool.class` never matched (every call fell into `else`); undefined `c` NameError when closing live connections; symbol branch called `delete` on the pool instead of the map. Shutdown now matches on the argument, removes the map entry before closing, and tolerates unknown names.
|
|
14
|
+
- **Factory#pool**: check-then-set race could orphan a pool under concurrent first-use; creation is now mutex-guarded.
|
|
15
|
+
- **Factory#start_reaper**: no longer checks out a connection just to log stats (which created connections and reset idle clocks); Idle/Available labels use `pool.idle` / `pool.available`; iterates a snapshot; rescues `PoolShuttingDownError` per pool so a retired pool cannot kill the reaper thread.
|
|
16
|
+
- **DuckDb#close**: use `@connection&.disconnect` so closing an already-closed adapter does not open a new connection.
|
|
17
|
+
- **DuckDb.close_all**: close then clear instead of deleting while iterating (which could skip entries).
|
|
18
|
+
|
|
3
19
|
## [0.5.0] - 2026-06-19
|
|
4
20
|
|
|
5
21
|
### Added
|
data/README.md
CHANGED
|
@@ -33,11 +33,12 @@ The adapter only has 5 core methods (6 including the connection method). A YAML
|
|
|
33
33
|
- **PostgreSQL** - Full-featured RDBMS with advanced SQL support
|
|
34
34
|
- **MySQL** - Popular open-source database
|
|
35
35
|
- **SQL Server** - Microsoft's enterprise database
|
|
36
|
+
- **ClickHouse** - High performance analytical database
|
|
37
|
+
- **Databricks** - Lakehouse SQL warehouse
|
|
38
|
+
- **Google BigQuery** - Google Cloud serverless warehouse
|
|
36
39
|
|
|
37
40
|
## Integrations Coming Soon
|
|
38
41
|
|
|
39
|
-
- **ClickHouse** - High performance analytical db
|
|
40
|
-
- **Databricks** - Big data compute engine
|
|
41
42
|
- **MotherDuck** - Hosted DuckDB service
|
|
42
43
|
|
|
43
44
|
## Quick Start
|
|
@@ -121,6 +122,21 @@ Run tests on druid:
|
|
|
121
122
|
bundle exec rake test:system:druid
|
|
122
123
|
```
|
|
123
124
|
|
|
125
|
+
Cloud warehouse tests (`test/system/cloud_*_test.rb`) run against real accounts and skip
|
|
126
|
+
unless configured through environment variables, so no credentials live in the repo:
|
|
127
|
+
|
|
128
|
+
| Adapter | Variables |
|
|
129
|
+
|---|---|
|
|
130
|
+
| Athena | `ATHENA_S3_OUTPUT`, optional `ATHENA_REGION`, `ATHENA_DATABASE`; AWS credentials from the environment |
|
|
131
|
+
| BigQuery | `BIGQUERY_PROJECT`, optional `BIGQUERY_DATASET`; `GOOGLE_APPLICATION_CREDENTIALS` or gcloud ADC. Creates its own fixture tables |
|
|
132
|
+
| Databricks | `DATABRICKS_HOST`, `DATABRICKS_WAREHOUSE`, `DATABRICKS_CLIENT_ID`, `DATABRICKS_CLIENT_SECRET` |
|
|
133
|
+
| Redshift | `REDSHIFT_HOST`, `REDSHIFT_PASSWORD`, optional `REDSHIFT_USER`, `REDSHIFT_DATABASE`, `REDSHIFT_PORT` |
|
|
134
|
+
| Snowflake | `SNOWFLAKE_ACCOUNT`, `SNOWFLAKE_PAT`, optional `SNOWFLAKE_DATABASE`; key-pair test needs `SNOWFLAKE_USER`, `SNOWFLAKE_PRIVATE_KEY` |
|
|
135
|
+
|
|
136
|
+
```bash
|
|
137
|
+
BIGQUERY_PROJECT=my-project BUNDLE_WITH=development bundle exec ruby -Itest test/system/cloud_bigquery_test.rb
|
|
138
|
+
```
|
|
139
|
+
|
|
124
140
|
## Development
|
|
125
141
|
|
|
126
142
|
After checking out the repo, run `bin/setup` to install dependencies. Then, run `rake test` to run the tests. You can also run `bin/console` for an interactive prompt.
|
|
@@ -137,4 +153,4 @@ This project is available as open source under the terms of the MIT License.
|
|
|
137
153
|
|
|
138
154
|
## Version
|
|
139
155
|
|
|
140
|
-
Current version: 0.
|
|
156
|
+
Current version: 0.6.0
|
data/docs/guides/adapters.md
CHANGED
|
@@ -697,6 +697,22 @@ athena = DWH.create(:athena, {
|
|
|
697
697
|
|
|
698
698
|
See full list of config options here: [athena-api](https://docs.aws.amazon.com/sdk-for-ruby/v2/api/Aws/Athena/Client.html#initialize-instance_method)
|
|
699
699
|
|
|
700
|
+
## Google BigQuery Adapter
|
|
701
|
+
|
|
702
|
+
Requires the `google-cloud-bigquery` gem (`gem install google-cloud-bigquery`, pure Ruby).
|
|
703
|
+
|
|
704
|
+
```ruby
|
|
705
|
+
bq = DWH.create(:bigquery, {
|
|
706
|
+
project_id: 'my-gcp-project',
|
|
707
|
+
dataset: 'analytics', # default dataset for unqualified table names
|
|
708
|
+
keyfile: '/path/to/service-account.json', # optional; omit to use Application Default Credentials
|
|
709
|
+
query_timeout: 300 # optional, seconds
|
|
710
|
+
})
|
|
711
|
+
```
|
|
712
|
+
|
|
713
|
+
Without `keyfile`, credentials resolve via `GOOGLE_APPLICATION_CREDENTIALS` or `gcloud auth application-default login`.
|
|
714
|
+
Extra client options can be passed with `extra_connection_params` (see `Google::Cloud::Bigquery.new`).
|
|
715
|
+
|
|
700
716
|
## Configuration Validation
|
|
701
717
|
|
|
702
718
|
DWH validates configuration parameters at creation time:
|
|
@@ -0,0 +1,173 @@
|
|
|
1
|
+
require 'csv'
|
|
2
|
+
require 'timeout'
|
|
3
|
+
|
|
4
|
+
module DWH
|
|
5
|
+
module Adapters
|
|
6
|
+
# Google BigQuery adapter. Requires the google-cloud-bigquery gem, which is
|
|
7
|
+
# loaded lazily on first connection so the adapter can be created without it.
|
|
8
|
+
#
|
|
9
|
+
# @example Service account keyfile
|
|
10
|
+
# DWH.create(:bigquery, {
|
|
11
|
+
# project_id: 'my-gcp-project',
|
|
12
|
+
# dataset: 'analytics',
|
|
13
|
+
# keyfile: '/path/to/service-account.json'
|
|
14
|
+
# })
|
|
15
|
+
#
|
|
16
|
+
# @example Application Default Credentials (gcloud auth application-default login)
|
|
17
|
+
# DWH.create(:bigquery, { project_id: 'my-gcp-project', dataset: 'analytics' })
|
|
18
|
+
class BigQuery < Adapter
|
|
19
|
+
config :project_id, String, required: true, message: 'GCP project id'
|
|
20
|
+
config :dataset, String, required: true, message: 'default dataset for unqualified table names'
|
|
21
|
+
config :keyfile, String, required: false, default: nil,
|
|
22
|
+
message: 'path to service-account JSON; omit to use Application Default Credentials'
|
|
23
|
+
config :query_timeout, Integer, required: false, default: 300, message: 'query timeout in seconds'
|
|
24
|
+
|
|
25
|
+
# The schema API reports legacy type names; map the ones whose
|
|
26
|
+
# normalized type would otherwise be wrong (INTEGER is 64-bit in BigQuery).
|
|
27
|
+
TYPE_ALIASES = { 'INTEGER' => 'INT64', 'FLOAT' => 'FLOAT64', 'BOOLEAN' => 'BOOL', 'RECORD' => 'STRUCT' }.freeze
|
|
28
|
+
|
|
29
|
+
# (see Adapter#connection)
|
|
30
|
+
def connection
|
|
31
|
+
return @connection if @connection
|
|
32
|
+
|
|
33
|
+
require 'google/cloud/bigquery'
|
|
34
|
+
opts = { project_id: config[:project_id] }
|
|
35
|
+
opts[:credentials] = config[:keyfile] unless config[:keyfile].to_s.empty?
|
|
36
|
+
@connection = Google::Cloud::Bigquery.new(**opts, **extra_connection_params)
|
|
37
|
+
rescue LoadError
|
|
38
|
+
raise ConfigError, <<~MSG
|
|
39
|
+
BigQuery adapter requires the 'google-cloud-bigquery' gem.
|
|
40
|
+
|
|
41
|
+
Install with: gem install google-cloud-bigquery
|
|
42
|
+
|
|
43
|
+
No system libraries required (pure Ruby).
|
|
44
|
+
MSG
|
|
45
|
+
rescue StandardError => e
|
|
46
|
+
raise ConfigError, "Failed to connect to BigQuery: #{e.message}"
|
|
47
|
+
end
|
|
48
|
+
|
|
49
|
+
# (see Adapter#test_connection)
|
|
50
|
+
def test_connection(raise_exception: false)
|
|
51
|
+
raise ConnectionError, "Dataset '#{config[:dataset]}' not found" unless connection.dataset(config[:dataset])
|
|
52
|
+
|
|
53
|
+
true
|
|
54
|
+
rescue StandardError => e
|
|
55
|
+
raise ConnectionError, "BigQuery connection test failed: #{e.message}" if raise_exception
|
|
56
|
+
|
|
57
|
+
false
|
|
58
|
+
end
|
|
59
|
+
|
|
60
|
+
# (see Adapter#tables)
|
|
61
|
+
def tables(**qualifiers)
|
|
62
|
+
dataset_for(qualifiers).tables.all.map(&:table_id)
|
|
63
|
+
end
|
|
64
|
+
|
|
65
|
+
# (see Adapter#stats)
|
|
66
|
+
def stats(table, date_column: nil, **qualifiers)
|
|
67
|
+
sql = 'SELECT COUNT(*) AS row_count'
|
|
68
|
+
sql += ", MIN(#{date_column}) AS date_start, MAX(#{date_column}) AS date_end" if date_column
|
|
69
|
+
row = execute("#{sql} FROM `#{dataset_name(qualifiers)}.#{table}`", format: :object).first || {}
|
|
70
|
+
|
|
71
|
+
TableStats.new(row_count: row['row_count'], date_start: row['date_start'], date_end: row['date_end'])
|
|
72
|
+
end
|
|
73
|
+
|
|
74
|
+
# (see Adapter#metadata)
|
|
75
|
+
# Uses the table schema API instead of INFORMATION_SCHEMA: one call, and
|
|
76
|
+
# precision/scale/length come back as attributes rather than parsed from strings.
|
|
77
|
+
def metadata(table, **qualifiers)
|
|
78
|
+
db_table = Table.new table, schema: dataset_name(qualifiers)
|
|
79
|
+
bq_table = dataset_for(qualifiers).table(db_table.physical_name)
|
|
80
|
+
raise ExecutionError, "Table '#{db_table.schema}.#{db_table.physical_name}' not found" unless bq_table
|
|
81
|
+
|
|
82
|
+
bq_table.schema.fields.each do |field|
|
|
83
|
+
db_table << Column.new(
|
|
84
|
+
name: field.name,
|
|
85
|
+
data_type: TYPE_ALIASES.fetch(field.type, field.type),
|
|
86
|
+
precision: field.precision || 0,
|
|
87
|
+
scale: field.scale || 0,
|
|
88
|
+
max_char_length: field.max_length
|
|
89
|
+
)
|
|
90
|
+
end
|
|
91
|
+
|
|
92
|
+
db_table
|
|
93
|
+
end
|
|
94
|
+
|
|
95
|
+
# (see Adapter#execute)
|
|
96
|
+
def execute(sql, format: :array, retries: 0)
|
|
97
|
+
result = with_debug(sql) { with_retry(retries) { run_query(sql) } }
|
|
98
|
+
|
|
99
|
+
format = format.downcase if format.is_a?(String)
|
|
100
|
+
case format.to_sym
|
|
101
|
+
when :array then result[:rows]
|
|
102
|
+
when :object then result[:rows].map { |row| Hash[result[:headers].zip(row)] }
|
|
103
|
+
when :csv then rows_to_csv(result[:headers], result[:rows])
|
|
104
|
+
when :native then result
|
|
105
|
+
else raise UnsupportedCapability, "Unsupported format: #{format} for BigQuery adapter"
|
|
106
|
+
end
|
|
107
|
+
end
|
|
108
|
+
|
|
109
|
+
# (see Adapter#execute_stream)
|
|
110
|
+
def execute_stream(sql, io, stats: nil, retries: 0)
|
|
111
|
+
with_debug(sql) do
|
|
112
|
+
with_retry(retries) do
|
|
113
|
+
data = query_data(sql)
|
|
114
|
+
io.write(CSV.generate_line(headers_of(data)))
|
|
115
|
+
data.all.each do |row|
|
|
116
|
+
values = row.values
|
|
117
|
+
stats << values unless stats.nil?
|
|
118
|
+
io.write(CSV.generate_line(values))
|
|
119
|
+
end
|
|
120
|
+
end
|
|
121
|
+
end
|
|
122
|
+
|
|
123
|
+
io.rewind
|
|
124
|
+
io
|
|
125
|
+
end
|
|
126
|
+
|
|
127
|
+
# (see Adapter#stream)
|
|
128
|
+
def stream(sql, &block)
|
|
129
|
+
with_debug(sql) { query_data(sql).all.each { |row| block.call(row.values) } }
|
|
130
|
+
end
|
|
131
|
+
|
|
132
|
+
private
|
|
133
|
+
|
|
134
|
+
def dataset_name(qualifiers)
|
|
135
|
+
qualifiers[:dataset] || qualifiers[:schema] || config[:dataset]
|
|
136
|
+
end
|
|
137
|
+
|
|
138
|
+
def dataset_for(qualifiers)
|
|
139
|
+
name = dataset_name(qualifiers)
|
|
140
|
+
connection.dataset(name) || raise(ExecutionError, "Dataset '#{name}' not found")
|
|
141
|
+
end
|
|
142
|
+
|
|
143
|
+
# Runs the query and returns the gem's paged Data object. Rows are hashes
|
|
144
|
+
# keyed by column name in column order, so row.values matches data.fields.
|
|
145
|
+
def query_data(sql)
|
|
146
|
+
Timeout.timeout(config[:query_timeout]) do
|
|
147
|
+
connection.query(sql, dataset: config[:dataset], project: config[:project_id])
|
|
148
|
+
end
|
|
149
|
+
rescue DWHError
|
|
150
|
+
raise
|
|
151
|
+
rescue StandardError => e
|
|
152
|
+
raise ExecutionError, "BigQuery query failed: #{e.message}"
|
|
153
|
+
end
|
|
154
|
+
|
|
155
|
+
def run_query(sql)
|
|
156
|
+
data = query_data(sql)
|
|
157
|
+
{ headers: headers_of(data), rows: data.all.map(&:values) }
|
|
158
|
+
end
|
|
159
|
+
|
|
160
|
+
# DDL/DML statements return no schema.
|
|
161
|
+
def headers_of(data)
|
|
162
|
+
data.schema&.fields&.map(&:name) || []
|
|
163
|
+
end
|
|
164
|
+
|
|
165
|
+
def rows_to_csv(headers, rows)
|
|
166
|
+
CSV.generate do |csv|
|
|
167
|
+
csv << headers
|
|
168
|
+
rows.each { |row| csv << row }
|
|
169
|
+
end
|
|
170
|
+
end
|
|
171
|
+
end
|
|
172
|
+
end
|
|
173
|
+
end
|
data/lib/dwh/adapters/duck_db.rb
CHANGED
|
@@ -56,10 +56,9 @@ module DWH
|
|
|
56
56
|
# we open one instance but many connections. Use this
|
|
57
57
|
# method to close them all.
|
|
58
58
|
def self.close_all
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
end
|
|
59
|
+
# Iterate values then clear — mutating the hash while each-ing skips entries.
|
|
60
|
+
databases.each_value(&:close)
|
|
61
|
+
databases.clear
|
|
63
62
|
end
|
|
64
63
|
|
|
65
64
|
# This disconnects the current connection but
|
|
@@ -68,7 +67,9 @@ module DWH
|
|
|
68
67
|
#
|
|
69
68
|
# (see Adapter#close)
|
|
70
69
|
def close
|
|
71
|
-
connection
|
|
70
|
+
# Use @connection directly: #connection opens a new handle when nil,
|
|
71
|
+
# which would re-open just to close.
|
|
72
|
+
@connection&.disconnect
|
|
72
73
|
@connection = nil
|
|
73
74
|
end
|
|
74
75
|
|
data/lib/dwh/column.rb
CHANGED
|
@@ -36,7 +36,7 @@ module DWH
|
|
|
36
36
|
inner = unwrap_type(data_type)
|
|
37
37
|
|
|
38
38
|
case inner
|
|
39
|
-
when /binary/, 'image'
|
|
39
|
+
when /binary/, 'image', 'bytes'
|
|
40
40
|
'binary'
|
|
41
41
|
when /varchar/, 'string', /text/, /char/, /fixedstring/
|
|
42
42
|
'string'
|
|
@@ -50,7 +50,7 @@ module DWH
|
|
|
50
50
|
when 'bigint', 'bit_int', 'big_integer', /^int64$/, /^int128$/, /^int256$/,
|
|
51
51
|
/^uint64$/, /^uint128$/, /^uint256$/
|
|
52
52
|
'bigint'
|
|
53
|
-
when 'decimal', 'double', 'float', 'real', 'dec', 'numeric', 'money',
|
|
53
|
+
when 'decimal', 'double', 'float', 'real', 'dec', 'numeric', 'bignumeric', 'money',
|
|
54
54
|
/^float32$/, /^float64$/, /^decimal/
|
|
55
55
|
'decimal'
|
|
56
56
|
when 'boolean', 'bit', 'bool'
|
data/lib/dwh/factory.rb
CHANGED
|
@@ -46,6 +46,11 @@ module DWH
|
|
|
46
46
|
@pools ||= {}
|
|
47
47
|
end
|
|
48
48
|
|
|
49
|
+
# Mutex guarding pool creation / map mutation.
|
|
50
|
+
def pool_mutex
|
|
51
|
+
@pool_mutex ||= Mutex.new
|
|
52
|
+
end
|
|
53
|
+
|
|
49
54
|
# The canonical way of creating an adapter instance
|
|
50
55
|
# in DWH.
|
|
51
56
|
# @param adapter_name [String, Symbol]
|
|
@@ -64,62 +69,80 @@ module DWH
|
|
|
64
69
|
# Create a pool of connections for a given name and adapter.
|
|
65
70
|
# Returns existing pool if it was already created.
|
|
66
71
|
#
|
|
67
|
-
# @param name [String] custom name for your pool
|
|
72
|
+
# @param name [String, Symbol] custom name for your pool (stored as String)
|
|
68
73
|
# @param adapter_name [String, Symbol]
|
|
69
74
|
# @param config [Hash] connection options
|
|
70
75
|
# @param timeout [Integer] pool checkout time out
|
|
71
76
|
# @param size [Integer] size of the pool
|
|
72
77
|
def pool(name, adapter_name, config, timeout: 5, size: 10)
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
78
|
+
name = name.to_s
|
|
79
|
+
pool_mutex.synchronize do
|
|
80
|
+
if pools.key?(name)
|
|
81
|
+
pools[name]
|
|
82
|
+
else
|
|
83
|
+
pools[name] = ConnectionPool.new(size: size, timeout: timeout) do
|
|
84
|
+
create(adapter_name, config)
|
|
85
|
+
end
|
|
78
86
|
end
|
|
79
87
|
end
|
|
80
88
|
end
|
|
81
89
|
|
|
82
90
|
# Shutdown a specific pool or all pools
|
|
83
|
-
# @param pool [String, ConnectionPool, nil] pool or name of pool
|
|
91
|
+
# @param pool [String, Symbol, ConnectionPool, nil] pool or name of pool
|
|
84
92
|
# or nil to shut everything down
|
|
85
93
|
def shutdown(pool = nil)
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
94
|
+
# Mutate the map under the same mutex as pool creation so a concurrent
|
|
95
|
+
# create cannot be orphaned by @pools = {} / delete racing it.
|
|
96
|
+
# Close outside the lock — ConnectionPool#shutdown can wait on check-in.
|
|
97
|
+
to_close = pool_mutex.synchronize do
|
|
98
|
+
case pool
|
|
99
|
+
when String, Symbol
|
|
100
|
+
# Delete first so a raising close cannot leave a dead pool in the map.
|
|
101
|
+
removed = pools.delete(pool.to_s)
|
|
102
|
+
removed ? [removed] : []
|
|
103
|
+
when ConnectionPool
|
|
104
|
+
key = pools.key(pool)
|
|
105
|
+
pools.delete(key) if key
|
|
106
|
+
[pool]
|
|
107
|
+
else
|
|
108
|
+
closing = pools.values
|
|
109
|
+
@pools = {}
|
|
110
|
+
closing
|
|
99
111
|
end
|
|
100
|
-
@pools = {}
|
|
101
112
|
end
|
|
113
|
+
to_close.each { |p| p.shutdown { it.close } }
|
|
102
114
|
end
|
|
103
115
|
|
|
104
116
|
# Start reaper that will periodically clean up
|
|
105
117
|
# unused or idle connections.
|
|
106
118
|
# @param frequency [Integer] defaults to 300 seconds
|
|
119
|
+
# @return [Thread] the reaper thread
|
|
107
120
|
def start_reaper(frequency = 300)
|
|
108
121
|
logger.info 'Starting DB Adapter reaper process'
|
|
109
122
|
Thread.new do
|
|
110
123
|
loop do
|
|
111
|
-
|
|
112
|
-
logger.info "DB POOL FOR #{name} STATS:"
|
|
113
|
-
pool.with do
|
|
114
|
-
logger.info "\tSize: #{pool.size}"
|
|
115
|
-
logger.info "\tIdle: #{pool.available}"
|
|
116
|
-
logger.info "\tAvailable: #{pool.available}"
|
|
117
|
-
end
|
|
118
|
-
pool.reap(frequency) { it.close }
|
|
119
|
-
end
|
|
124
|
+
reaper_tick(frequency)
|
|
120
125
|
sleep frequency
|
|
121
126
|
end
|
|
122
127
|
end
|
|
123
128
|
end
|
|
129
|
+
|
|
130
|
+
# One reaper cycle: log pool stats (without checking out) and reap idle connections.
|
|
131
|
+
# Safe to call from tests; rescues per-pool so a shut-down pool cannot kill the loop.
|
|
132
|
+
# @param frequency [Integer] idle threshold passed to ConnectionPool#reap
|
|
133
|
+
def reaper_tick(frequency = 300)
|
|
134
|
+
# Snapshot so concurrent shutdown deletions do not mutate while we iterate.
|
|
135
|
+
pools.to_a.each do |name, pool|
|
|
136
|
+
logger.info "DB POOL FOR #{name} STATS:"
|
|
137
|
+
logger.info "\tSize: #{pool.size}"
|
|
138
|
+
logger.info "\tIdle: #{pool.idle}"
|
|
139
|
+
logger.info "\tAvailable: #{pool.available}"
|
|
140
|
+
pool.reap(frequency) { it.close }
|
|
141
|
+
rescue ConnectionPool::PoolShuttingDownError => e
|
|
142
|
+
logger.info "Skipping reaper for pool #{name}: #{e.class}"
|
|
143
|
+
rescue StandardError => e
|
|
144
|
+
logger.error "Reaper error for pool #{name}: #{e.class}: #{e.message}"
|
|
145
|
+
end
|
|
146
|
+
end
|
|
124
147
|
end
|
|
125
148
|
end
|
|
@@ -0,0 +1,92 @@
|
|
|
1
|
+
# Google BigQuery adapter settings
|
|
2
|
+
# Only overrides that differ from base.yml are listed here.
|
|
3
|
+
|
|
4
|
+
# BigQuery uses backticks for identifier quoting.
|
|
5
|
+
quote: "`@exp`"
|
|
6
|
+
|
|
7
|
+
# FORMAT_DATE uses strftime elements. base.yml's %W/%M mean
|
|
8
|
+
# week-number/minute in BigQuery, so use the standard name codes.
|
|
9
|
+
day_name_format: "%A"
|
|
10
|
+
abbreviated_day_name_format: "%a"
|
|
11
|
+
month_name_format: "%B"
|
|
12
|
+
abbreviated_month_name_format: "%b"
|
|
13
|
+
|
|
14
|
+
# Current time functions require parentheses.
|
|
15
|
+
current_date: "CURRENT_DATE()"
|
|
16
|
+
current_time: "CURRENT_TIME()"
|
|
17
|
+
current_timestamp: "CURRENT_TIMESTAMP()"
|
|
18
|
+
|
|
19
|
+
# Date literals
|
|
20
|
+
date_literal: "DATE '@val'"
|
|
21
|
+
date_time_literal: "TIMESTAMP '@val'"
|
|
22
|
+
|
|
23
|
+
# BigQuery puts the expression first and takes the unit as a bare keyword.
|
|
24
|
+
truncate_date: "DATE_TRUNC(@exp, @unit)"
|
|
25
|
+
date_add: "DATE_ADD(@exp, INTERVAL @val @unit)"
|
|
26
|
+
date_diff: "DATE_DIFF(@end_exp, @start_exp, @unit)"
|
|
27
|
+
date_format_sql: "FORMAT_DATE('@format', @exp)"
|
|
28
|
+
|
|
29
|
+
extract_year: "EXTRACT(YEAR FROM @exp)"
|
|
30
|
+
extract_month: "EXTRACT(MONTH FROM @exp)"
|
|
31
|
+
extract_quarter: "EXTRACT(QUARTER FROM @exp)"
|
|
32
|
+
extract_day_of_year: "EXTRACT(DAYOFYEAR FROM @exp)"
|
|
33
|
+
extract_day_of_month: "EXTRACT(DAY FROM @exp)"
|
|
34
|
+
extract_day_of_week: "EXTRACT(DAYOFWEEK FROM @exp)"
|
|
35
|
+
extract_week_of_year: "EXTRACT(WEEK FROM @exp)"
|
|
36
|
+
extract_hour: "EXTRACT(HOUR FROM @exp)"
|
|
37
|
+
extract_minute: "EXTRACT(MINUTE FROM @exp)"
|
|
38
|
+
extract_year_month: "CAST(FORMAT_DATE('%Y%m', @exp) AS INT64)"
|
|
39
|
+
|
|
40
|
+
# DATE_TRUNC(x, WEEK) is Sunday-based in BigQuery; the explicit forms
|
|
41
|
+
# below are used whenever the requested week_start_day differs.
|
|
42
|
+
default_week_start_day: "sunday"
|
|
43
|
+
sunday_week_start_day: "DATE_TRUNC(@exp, WEEK(SUNDAY))"
|
|
44
|
+
monday_week_start_day: "DATE_TRUNC(@exp, WEEK(MONDAY))"
|
|
45
|
+
|
|
46
|
+
# Null handling
|
|
47
|
+
if_null: "IFNULL(@exp, @when_null)"
|
|
48
|
+
|
|
49
|
+
# Array operations via UNNEST
|
|
50
|
+
array_in_list: "EXISTS(SELECT 1 FROM UNNEST(@exp) AS x WHERE x IN (@list))"
|
|
51
|
+
array_exclude_list: "NOT EXISTS(SELECT 1 FROM UNNEST(@exp) AS x WHERE x IN (@list))"
|
|
52
|
+
array_unnest_join: "CROSS JOIN UNNEST(@exp) AS @alias"
|
|
53
|
+
|
|
54
|
+
# Capabilities
|
|
55
|
+
# Temp tables only live inside a script/session; use CTEs instead.
|
|
56
|
+
supports_temp_tables: false
|
|
57
|
+
temp_table_type: "cte"
|
|
58
|
+
|
|
59
|
+
# BigQuery-specific reserved keywords not in the standard baseline.
|
|
60
|
+
extra_reserved_keywords:
|
|
61
|
+
- struct
|
|
62
|
+
- unnest
|
|
63
|
+
- qualify
|
|
64
|
+
- tablesample
|
|
65
|
+
- window
|
|
66
|
+
- ignore
|
|
67
|
+
- respect
|
|
68
|
+
- int64
|
|
69
|
+
- float64
|
|
70
|
+
- bytes
|
|
71
|
+
- bignumeric
|
|
72
|
+
- bool
|
|
73
|
+
- datetime
|
|
74
|
+
- geography
|
|
75
|
+
- dayofweek
|
|
76
|
+
- dayofyear
|
|
77
|
+
- isoweek
|
|
78
|
+
- isoyear
|
|
79
|
+
|
|
80
|
+
# BigQuery-specific aggregate functions not in the standard baseline.
|
|
81
|
+
extra_aggregate_functions:
|
|
82
|
+
- any_value
|
|
83
|
+
- countif
|
|
84
|
+
- logical_and
|
|
85
|
+
- logical_or
|
|
86
|
+
- array_concat_agg
|
|
87
|
+
- approx_quantiles
|
|
88
|
+
- approx_top_count
|
|
89
|
+
- approx_top_sum
|
|
90
|
+
- bit_xor
|
|
91
|
+
- max_by
|
|
92
|
+
- min_by
|
data/lib/dwh/version.rb
CHANGED
data/lib/dwh.rb
CHANGED
|
@@ -21,6 +21,7 @@ require_relative 'dwh/adapters/athena'
|
|
|
21
21
|
require_relative 'dwh/adapters/redshift'
|
|
22
22
|
require_relative 'dwh/adapters/databricks'
|
|
23
23
|
require_relative 'dwh/adapters/click_house'
|
|
24
|
+
require_relative 'dwh/adapters/big_query'
|
|
24
25
|
|
|
25
26
|
# DWH encapsulates the full functionality of this gem.
|
|
26
27
|
#
|
|
@@ -54,6 +55,7 @@ module DWH
|
|
|
54
55
|
register(:redshift, Adapters::Redshift)
|
|
55
56
|
register(:databricks, Adapters::Databricks)
|
|
56
57
|
register(:clickhouse, Adapters::ClickHouse)
|
|
58
|
+
register(:bigquery, Adapters::BigQuery)
|
|
57
59
|
|
|
58
60
|
# The raw base.yml settings, loaded once. This is the single source of
|
|
59
61
|
# truth for the standard, warehouse-agnostic dialect baseline.
|
metadata
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: dwh
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 0.
|
|
4
|
+
version: 0.6.0
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Ajo Abraham
|
|
@@ -154,6 +154,7 @@ files:
|
|
|
154
154
|
- lib/dwh.rb
|
|
155
155
|
- lib/dwh/adapters.rb
|
|
156
156
|
- lib/dwh/adapters/athena.rb
|
|
157
|
+
- lib/dwh/adapters/big_query.rb
|
|
157
158
|
- lib/dwh/adapters/click_house.rb
|
|
158
159
|
- lib/dwh/adapters/databricks.rb
|
|
159
160
|
- lib/dwh/adapters/druid.rb
|
|
@@ -181,6 +182,7 @@ files:
|
|
|
181
182
|
- lib/dwh/settings.rb
|
|
182
183
|
- lib/dwh/settings/athena.yml
|
|
183
184
|
- lib/dwh/settings/base.yml
|
|
185
|
+
- lib/dwh/settings/bigquery.yml
|
|
184
186
|
- lib/dwh/settings/clickhouse.yml
|
|
185
187
|
- lib/dwh/settings/databricks.yml
|
|
186
188
|
- lib/dwh/settings/druid.yml
|