dwh 0.5.1 → 0.6.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: fd8f7098719fc4a93610c3295a6bcbbf14e8487f07cb8dec86c9b95df7805b4e
4
- data.tar.gz: a18967ab069c582d12bd3eaaa7d97f5242e3380c9f348e1fd2809da4c82ace88
3
+ metadata.gz: 1bb280be5e0be9e979828b1e35d8271dbfd35f6926cedcac3c9c134140ca8b30
4
+ data.tar.gz: 407efc9f63ce81f8bb166d970d5ee2d43e8cd31e1fe41f62802362ddc02653e7
5
5
  SHA512:
6
- metadata.gz: 27599ae5ac8ddd1db8fddf7d9714318eb36f98cb66ff73ed71026dddb0ab4f97e8f2f060dde216e6217485b1627cdc5ce7031254aca92da20395334a61d9bb57
7
- data.tar.gz: a560342b601b038bd8e1d6d30708822433b09d83bef6094a3a1839e5cd0b09dae4bdbffce45e945f43ec1b6fc6e894e1def9d1c3d2a4c76d38a6a2c7cac3f3b6
6
+ metadata.gz: 8773bba3799deb57bd3e3e4738287ded1f0bd8f32fc1225490f346076676f3c962dc58ad43a697e87dc86b41ce10f563af6ac0a1eb3f1ca84836ea9d468cf29d
7
+ data.tar.gz: af66e29a860e89af3c5f8794df099a622d3baa4f7ea5576076c1856399cc91838e3ecdf85ed70c038a8f8061b8da2e65bffc2f49ceca0596ac0457f5b35dfb46
data/CHANGELOG.md CHANGED
@@ -1,5 +1,17 @@
1
1
  ## [Unreleased]
2
2
 
3
+ ## [0.6.1] - 2026-09-30
4
+
5
+ ### Fixed
6
+
7
+ - **BigQuery**: `quote` replaces characters BigQuery rejects in column names (parentheses and similar) so planner-generated aliases such as `Month(Post Date)` execute.
8
+
9
+ ## [0.6.0] - 2026-09-28
10
+
11
+ ### Added
12
+
13
+ - Google BigQuery adapter (`:bigquery`) with dedicated settings and unit/system test coverage. Authenticates with a service-account keyfile or Application Default Credentials; requires the `google-cloud-bigquery` gem.
14
+
3
15
  ## [0.5.1] - 2026-08-03
4
16
 
5
17
  ### Fixed
data/README.md CHANGED
@@ -35,6 +35,7 @@ The adapter only has 5 core methods (6 including the connection method). A YAML
35
35
  - **SQL Server** - Microsoft's enterprise database
36
36
  - **ClickHouse** - High performance analytical database
37
37
  - **Databricks** - Lakehouse SQL warehouse
38
+ - **Google BigQuery** - Google Cloud serverless warehouse
38
39
 
39
40
  ## Integrations Coming Soon
40
41
 
@@ -121,6 +122,21 @@ Run tests on druid:
121
122
  bundle exec rake test:system:druid
122
123
  ```
123
124
 
125
+ Cloud warehouse tests (`test/system/cloud_*_test.rb`) run against real accounts and skip
126
+ unless configured through environment variables, so no credentials live in the repo:
127
+
128
+ | Adapter | Variables |
129
+ |---|---|
130
+ | Athena | `ATHENA_S3_OUTPUT`, optional `ATHENA_REGION`, `ATHENA_DATABASE`; AWS credentials from the environment |
131
+ | BigQuery | `BIGQUERY_PROJECT`, optional `BIGQUERY_DATASET`; `GOOGLE_APPLICATION_CREDENTIALS` or gcloud ADC. Creates its own fixture tables |
132
+ | Databricks | `DATABRICKS_HOST`, `DATABRICKS_WAREHOUSE`, `DATABRICKS_CLIENT_ID`, `DATABRICKS_CLIENT_SECRET` |
133
+ | Redshift | `REDSHIFT_HOST`, `REDSHIFT_PASSWORD`, optional `REDSHIFT_USER`, `REDSHIFT_DATABASE`, `REDSHIFT_PORT` |
134
+ | Snowflake | `SNOWFLAKE_ACCOUNT`, `SNOWFLAKE_PAT`, optional `SNOWFLAKE_DATABASE`; key-pair test needs `SNOWFLAKE_USER`, `SNOWFLAKE_PRIVATE_KEY` |
135
+
136
+ ```bash
137
+ BIGQUERY_PROJECT=my-project BUNDLE_WITH=development bundle exec ruby -Itest test/system/cloud_bigquery_test.rb
138
+ ```
139
+
124
140
  ## Development
125
141
 
126
142
  After checking out the repo, run `bin/setup` to install dependencies. Then, run `rake test` to run the tests. You can also run `bin/console` for an interactive prompt.
@@ -137,4 +153,4 @@ This project is available as open source under the terms of the MIT License.
137
153
 
138
154
  ## Version
139
155
 
140
- Current version: 0.5.0
156
+ Current version: 0.6.1
@@ -697,6 +697,22 @@ athena = DWH.create(:athena, {
697
697
 
698
698
  See full list of config options here: [athena-api](https://docs.aws.amazon.com/sdk-for-ruby/v2/api/Aws/Athena/Client.html#initialize-instance_method)
699
699
 
700
+ ## Google BigQuery Adapter
701
+
702
+ Requires the `google-cloud-bigquery` gem (`gem install google-cloud-bigquery`, pure Ruby).
703
+
704
+ ```ruby
705
+ bq = DWH.create(:bigquery, {
706
+ project_id: 'my-gcp-project',
707
+ dataset: 'analytics', # default dataset for unqualified table names
708
+ keyfile: '/path/to/service-account.json', # optional; omit to use Application Default Credentials
709
+ query_timeout: 300 # optional, seconds
710
+ })
711
+ ```
712
+
713
+ Without `keyfile`, credentials resolve via `GOOGLE_APPLICATION_CREDENTIALS` or `gcloud auth application-default login`.
714
+ Extra client options can be passed with `extra_connection_params` (see `Google::Cloud::Bigquery.new`).
715
+
700
716
  ## Configuration Validation
701
717
 
702
718
  DWH validates configuration parameters at creation time:
@@ -0,0 +1,184 @@
1
+ require 'csv'
2
+ require 'timeout'
3
+
4
+ module DWH
5
+ module Adapters
6
+ # Google BigQuery adapter. Requires the google-cloud-bigquery gem, which is
7
+ # loaded lazily on first connection so the adapter can be created without it.
8
+ #
9
+ # @example Service account keyfile
10
+ # DWH.create(:bigquery, {
11
+ # project_id: 'my-gcp-project',
12
+ # dataset: 'analytics',
13
+ # keyfile: '/path/to/service-account.json'
14
+ # })
15
+ #
16
+ # @example Application Default Credentials (gcloud auth application-default login)
17
+ # DWH.create(:bigquery, { project_id: 'my-gcp-project', dataset: 'analytics' })
18
+ class BigQuery < Adapter
19
+ config :project_id, String, required: true, message: 'GCP project id'
20
+ config :dataset, String, required: true, message: 'default dataset for unqualified table names'
21
+ config :keyfile, String, required: false, default: nil,
22
+ message: 'path to service-account JSON; omit to use Application Default Credentials'
23
+ config :query_timeout, Integer, required: false, default: 300, message: 'query timeout in seconds'
24
+
25
+ # The schema API reports legacy type names; map the ones whose
26
+ # normalized type would otherwise be wrong (INTEGER is 64-bit in BigQuery).
27
+ TYPE_ALIASES = { 'INTEGER' => 'INT64', 'FLOAT' => 'FLOAT64', 'BOOLEAN' => 'BOOL', 'RECORD' => 'STRUCT' }.freeze
28
+
29
+ # BigQuery column names, and therefore SELECT aliases, may not contain these
30
+ # characters (https://cloud.google.com/bigquery/docs/schemas#column_names).
31
+ # Dots are left alone so backtick-quoted `dataset.table` paths still work.
32
+ INVALID_NAME_CHARS = %r{[!"$()*,/;?@\[\\\]^{}~]}
33
+
34
+ # (see Functions#quote) Replaces characters BigQuery rejects in a column
35
+ # name so planner-generated aliases like "Month(Post Date)" execute.
36
+ def quote(exp)
37
+ super(exp.to_s.gsub(INVALID_NAME_CHARS, '_'))
38
+ end
39
+
40
+ # (see Adapter#connection)
41
+ def connection
42
+ return @connection if @connection
43
+
44
+ require 'google/cloud/bigquery'
45
+ opts = { project_id: config[:project_id] }
46
+ opts[:credentials] = config[:keyfile] unless config[:keyfile].to_s.empty?
47
+ @connection = Google::Cloud::Bigquery.new(**opts, **extra_connection_params)
48
+ rescue LoadError
49
+ raise ConfigError, <<~MSG
50
+ BigQuery adapter requires the 'google-cloud-bigquery' gem.
51
+
52
+ Install with: gem install google-cloud-bigquery
53
+
54
+ No system libraries required (pure Ruby).
55
+ MSG
56
+ rescue StandardError => e
57
+ raise ConfigError, "Failed to connect to BigQuery: #{e.message}"
58
+ end
59
+
60
+ # (see Adapter#test_connection)
61
+ def test_connection(raise_exception: false)
62
+ raise ConnectionError, "Dataset '#{config[:dataset]}' not found" unless connection.dataset(config[:dataset])
63
+
64
+ true
65
+ rescue StandardError => e
66
+ raise ConnectionError, "BigQuery connection test failed: #{e.message}" if raise_exception
67
+
68
+ false
69
+ end
70
+
71
+ # (see Adapter#tables)
72
+ def tables(**qualifiers)
73
+ dataset_for(qualifiers).tables.all.map(&:table_id)
74
+ end
75
+
76
+ # (see Adapter#stats)
77
+ def stats(table, date_column: nil, **qualifiers)
78
+ sql = 'SELECT COUNT(*) AS row_count'
79
+ sql += ", MIN(#{date_column}) AS date_start, MAX(#{date_column}) AS date_end" if date_column
80
+ row = execute("#{sql} FROM `#{dataset_name(qualifiers)}.#{table}`", format: :object).first || {}
81
+
82
+ TableStats.new(row_count: row['row_count'], date_start: row['date_start'], date_end: row['date_end'])
83
+ end
84
+
85
+ # (see Adapter#metadata)
86
+ # Uses the table schema API instead of INFORMATION_SCHEMA: one call, and
87
+ # precision/scale/length come back as attributes rather than parsed from strings.
88
+ def metadata(table, **qualifiers)
89
+ db_table = Table.new table, schema: dataset_name(qualifiers)
90
+ bq_table = dataset_for(qualifiers).table(db_table.physical_name)
91
+ raise ExecutionError, "Table '#{db_table.schema}.#{db_table.physical_name}' not found" unless bq_table
92
+
93
+ bq_table.schema.fields.each do |field|
94
+ db_table << Column.new(
95
+ name: field.name,
96
+ data_type: TYPE_ALIASES.fetch(field.type, field.type),
97
+ precision: field.precision || 0,
98
+ scale: field.scale || 0,
99
+ max_char_length: field.max_length
100
+ )
101
+ end
102
+
103
+ db_table
104
+ end
105
+
106
+ # (see Adapter#execute)
107
+ def execute(sql, format: :array, retries: 0)
108
+ result = with_debug(sql) { with_retry(retries) { run_query(sql) } }
109
+
110
+ format = format.downcase if format.is_a?(String)
111
+ case format.to_sym
112
+ when :array then result[:rows]
113
+ when :object then result[:rows].map { |row| Hash[result[:headers].zip(row)] }
114
+ when :csv then rows_to_csv(result[:headers], result[:rows])
115
+ when :native then result
116
+ else raise UnsupportedCapability, "Unsupported format: #{format} for BigQuery adapter"
117
+ end
118
+ end
119
+
120
+ # (see Adapter#execute_stream)
121
+ def execute_stream(sql, io, stats: nil, retries: 0)
122
+ with_debug(sql) do
123
+ with_retry(retries) do
124
+ data = query_data(sql)
125
+ io.write(CSV.generate_line(headers_of(data)))
126
+ data.all.each do |row|
127
+ values = row.values
128
+ stats << values unless stats.nil?
129
+ io.write(CSV.generate_line(values))
130
+ end
131
+ end
132
+ end
133
+
134
+ io.rewind
135
+ io
136
+ end
137
+
138
+ # (see Adapter#stream)
139
+ def stream(sql, &block)
140
+ with_debug(sql) { query_data(sql).all.each { |row| block.call(row.values) } }
141
+ end
142
+
143
+ private
144
+
145
+ def dataset_name(qualifiers)
146
+ qualifiers[:dataset] || qualifiers[:schema] || config[:dataset]
147
+ end
148
+
149
+ def dataset_for(qualifiers)
150
+ name = dataset_name(qualifiers)
151
+ connection.dataset(name) || raise(ExecutionError, "Dataset '#{name}' not found")
152
+ end
153
+
154
+ # Runs the query and returns the gem's paged Data object. Rows are hashes
155
+ # keyed by column name in column order, so row.values matches data.fields.
156
+ def query_data(sql)
157
+ Timeout.timeout(config[:query_timeout]) do
158
+ connection.query(sql, dataset: config[:dataset], project: config[:project_id])
159
+ end
160
+ rescue DWHError
161
+ raise
162
+ rescue StandardError => e
163
+ raise ExecutionError, "BigQuery query failed: #{e.message}"
164
+ end
165
+
166
+ def run_query(sql)
167
+ data = query_data(sql)
168
+ { headers: headers_of(data), rows: data.all.map(&:values) }
169
+ end
170
+
171
+ # DDL/DML statements return no schema.
172
+ def headers_of(data)
173
+ data.schema&.fields&.map(&:name) || []
174
+ end
175
+
176
+ def rows_to_csv(headers, rows)
177
+ CSV.generate do |csv|
178
+ csv << headers
179
+ rows.each { |row| csv << row }
180
+ end
181
+ end
182
+ end
183
+ end
184
+ end
data/lib/dwh/column.rb CHANGED
@@ -36,7 +36,7 @@ module DWH
36
36
  inner = unwrap_type(data_type)
37
37
 
38
38
  case inner
39
- when /binary/, 'image'
39
+ when /binary/, 'image', 'bytes'
40
40
  'binary'
41
41
  when /varchar/, 'string', /text/, /char/, /fixedstring/
42
42
  'string'
@@ -50,7 +50,7 @@ module DWH
50
50
  when 'bigint', 'bit_int', 'big_integer', /^int64$/, /^int128$/, /^int256$/,
51
51
  /^uint64$/, /^uint128$/, /^uint256$/
52
52
  'bigint'
53
- when 'decimal', 'double', 'float', 'real', 'dec', 'numeric', 'money',
53
+ when 'decimal', 'double', 'float', 'real', 'dec', 'numeric', 'bignumeric', 'money',
54
54
  /^float32$/, /^float64$/, /^decimal/
55
55
  'decimal'
56
56
  when 'boolean', 'bit', 'bool'
@@ -0,0 +1,92 @@
1
+ # Google BigQuery adapter settings
2
+ # Only overrides that differ from base.yml are listed here.
3
+
4
+ # BigQuery uses backticks for identifier quoting.
5
+ quote: "`@exp`"
6
+
7
+ # FORMAT_DATE uses strftime elements. base.yml's %W/%M mean
8
+ # week-number/minute in BigQuery, so use the standard name codes.
9
+ day_name_format: "%A"
10
+ abbreviated_day_name_format: "%a"
11
+ month_name_format: "%B"
12
+ abbreviated_month_name_format: "%b"
13
+
14
+ # Current time functions require parentheses.
15
+ current_date: "CURRENT_DATE()"
16
+ current_time: "CURRENT_TIME()"
17
+ current_timestamp: "CURRENT_TIMESTAMP()"
18
+
19
+ # Date literals
20
+ date_literal: "DATE '@val'"
21
+ date_time_literal: "TIMESTAMP '@val'"
22
+
23
+ # BigQuery puts the expression first and takes the unit as a bare keyword.
24
+ truncate_date: "DATE_TRUNC(@exp, @unit)"
25
+ date_add: "DATE_ADD(@exp, INTERVAL @val @unit)"
26
+ date_diff: "DATE_DIFF(@end_exp, @start_exp, @unit)"
27
+ date_format_sql: "FORMAT_DATE('@format', @exp)"
28
+
29
+ extract_year: "EXTRACT(YEAR FROM @exp)"
30
+ extract_month: "EXTRACT(MONTH FROM @exp)"
31
+ extract_quarter: "EXTRACT(QUARTER FROM @exp)"
32
+ extract_day_of_year: "EXTRACT(DAYOFYEAR FROM @exp)"
33
+ extract_day_of_month: "EXTRACT(DAY FROM @exp)"
34
+ extract_day_of_week: "EXTRACT(DAYOFWEEK FROM @exp)"
35
+ extract_week_of_year: "EXTRACT(WEEK FROM @exp)"
36
+ extract_hour: "EXTRACT(HOUR FROM @exp)"
37
+ extract_minute: "EXTRACT(MINUTE FROM @exp)"
38
+ extract_year_month: "CAST(FORMAT_DATE('%Y%m', @exp) AS INT64)"
39
+
40
+ # DATE_TRUNC(x, WEEK) is Sunday-based in BigQuery; the explicit forms
41
+ # below are used whenever the requested week_start_day differs.
42
+ default_week_start_day: "sunday"
43
+ sunday_week_start_day: "DATE_TRUNC(@exp, WEEK(SUNDAY))"
44
+ monday_week_start_day: "DATE_TRUNC(@exp, WEEK(MONDAY))"
45
+
46
+ # Null handling
47
+ if_null: "IFNULL(@exp, @when_null)"
48
+
49
+ # Array operations via UNNEST
50
+ array_in_list: "EXISTS(SELECT 1 FROM UNNEST(@exp) AS x WHERE x IN (@list))"
51
+ array_exclude_list: "NOT EXISTS(SELECT 1 FROM UNNEST(@exp) AS x WHERE x IN (@list))"
52
+ array_unnest_join: "CROSS JOIN UNNEST(@exp) AS @alias"
53
+
54
+ # Capabilities
55
+ # Temp tables only live inside a script/session; use CTEs instead.
56
+ supports_temp_tables: false
57
+ temp_table_type: "cte"
58
+
59
+ # BigQuery-specific reserved keywords not in the standard baseline.
60
+ extra_reserved_keywords:
61
+ - struct
62
+ - unnest
63
+ - qualify
64
+ - tablesample
65
+ - window
66
+ - ignore
67
+ - respect
68
+ - int64
69
+ - float64
70
+ - bytes
71
+ - bignumeric
72
+ - bool
73
+ - datetime
74
+ - geography
75
+ - dayofweek
76
+ - dayofyear
77
+ - isoweek
78
+ - isoyear
79
+
80
+ # BigQuery-specific aggregate functions not in the standard baseline.
81
+ extra_aggregate_functions:
82
+ - any_value
83
+ - countif
84
+ - logical_and
85
+ - logical_or
86
+ - array_concat_agg
87
+ - approx_quantiles
88
+ - approx_top_count
89
+ - approx_top_sum
90
+ - bit_xor
91
+ - max_by
92
+ - min_by
data/lib/dwh/version.rb CHANGED
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module DWH
4
- VERSION = '0.5.1'
4
+ VERSION = '0.6.1'
5
5
  end
data/lib/dwh.rb CHANGED
@@ -21,6 +21,7 @@ require_relative 'dwh/adapters/athena'
21
21
  require_relative 'dwh/adapters/redshift'
22
22
  require_relative 'dwh/adapters/databricks'
23
23
  require_relative 'dwh/adapters/click_house'
24
+ require_relative 'dwh/adapters/big_query'
24
25
 
25
26
  # DWH encapsulates the full functionality of this gem.
26
27
  #
@@ -54,6 +55,7 @@ module DWH
54
55
  register(:redshift, Adapters::Redshift)
55
56
  register(:databricks, Adapters::Databricks)
56
57
  register(:clickhouse, Adapters::ClickHouse)
58
+ register(:bigquery, Adapters::BigQuery)
57
59
 
58
60
  # The raw base.yml settings, loaded once. This is the single source of
59
61
  # truth for the standard, warehouse-agnostic dialect baseline.
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: dwh
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.5.1
4
+ version: 0.6.1
5
5
  platform: ruby
6
6
  authors:
7
7
  - Ajo Abraham
@@ -154,6 +154,7 @@ files:
154
154
  - lib/dwh.rb
155
155
  - lib/dwh/adapters.rb
156
156
  - lib/dwh/adapters/athena.rb
157
+ - lib/dwh/adapters/big_query.rb
157
158
  - lib/dwh/adapters/click_house.rb
158
159
  - lib/dwh/adapters/databricks.rb
159
160
  - lib/dwh/adapters/druid.rb
@@ -181,6 +182,7 @@ files:
181
182
  - lib/dwh/settings.rb
182
183
  - lib/dwh/settings/athena.yml
183
184
  - lib/dwh/settings/base.yml
185
+ - lib/dwh/settings/bigquery.yml
184
186
  - lib/dwh/settings/clickhouse.yml
185
187
  - lib/dwh/settings/databricks.yml
186
188
  - lib/dwh/settings/druid.yml