google-cloud-db-context-engineering 0.7.2__tar.gz → 0.7.3__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {google_cloud_db_context_engineering-0.7.2/src/google_cloud_db_context_engineering.egg-info → google_cloud_db_context_engineering-0.7.3}/PKG-INFO +8 -11
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/README.md +7 -10
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/pyproject.toml +2 -2
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google/cloud/db_context_enrichment/common/context_mutator.py +2 -2
- google_cloud_db_context_engineering-0.7.3/src/google/cloud/db_context_enrichment/common/context_validator.py +188 -0
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google/cloud/db_context_enrichment/evaluate/db_generators/spanner.py +40 -2
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google/cloud/db_context_enrichment/evaluate/evaluate_generator.py +88 -7
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google/cloud/db_context_enrichment/main.py +25 -0
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3/src/google_cloud_db_context_engineering.egg-info}/PKG-INFO +8 -11
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google_cloud_db_context_engineering.egg-info/SOURCES.txt +1 -0
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/LICENSE +0 -0
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/setup.cfg +0 -0
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google/cloud/db_context_enrichment/__init__.py +0 -0
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google/cloud/db_context_enrichment/common/__init__.py +0 -0
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google/cloud/db_context_enrichment/common/config.py +0 -0
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google/cloud/db_context_enrichment/common/context_store_client.py +0 -0
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google/cloud/db_context_enrichment/dataset/__init__.py +0 -0
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google/cloud/db_context_enrichment/dataset/dataset_generator.py +0 -0
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google/cloud/db_context_enrichment/evaluate/__init__.py +0 -0
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google/cloud/db_context_enrichment/evaluate/db_generators/__init__.py +0 -0
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google/cloud/db_context_enrichment/evaluate/db_generators/alloydb.py +0 -0
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google/cloud/db_context_enrichment/evaluate/db_generators/base.py +0 -0
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google/cloud/db_context_enrichment/evaluate/db_generators/mysql.py +0 -0
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google/cloud/db_context_enrichment/evaluate/db_generators/postgres.py +0 -0
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google/cloud/db_context_enrichment/evaluate/result_reader.py +0 -0
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google/cloud/db_context_enrichment/model/__init__.py +0 -0
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google/cloud/db_context_enrichment/model/context.py +0 -0
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google_cloud_db_context_engineering.egg-info/dependency_links.txt +0 -0
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google_cloud_db_context_engineering.egg-info/entry_points.txt +0 -0
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google_cloud_db_context_engineering.egg-info/requires.txt +0 -0
- {google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/src/google_cloud_db_context_engineering.egg-info/top_level.txt +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: google-cloud-db-context-engineering
|
|
3
|
-
Version: 0.7.
|
|
3
|
+
Version: 0.7.3
|
|
4
4
|
Summary: A FastMCP server for generating natural language to SQL templates from database schemas.
|
|
5
5
|
Requires-Python: >=3.12
|
|
6
6
|
Description-Content-Type: text/markdown
|
|
@@ -16,19 +16,17 @@ Requires-Dist: pytest; extra == "test"
|
|
|
16
16
|
Requires-Dist: pytest-asyncio; extra == "test"
|
|
17
17
|
Dynamic: license-file
|
|
18
18
|
|
|
19
|
-
This is not an officially supported Google product. This project is not eligible for the [Google Open Source Software Vulnerability Rewards Program](https://bughunters.google.com/open-source-security), [Google Cloud Platform/SecOps Terms of Service](https://cloud.google.com/terms), [How Gemini for Google Cloud uses your data](https://cloud.google.com/gemini/docs/discover/data-governance). This tool is provided "as is" without warranty of any kind. Users are solely responsible for understanding and managing the tool's interaction with their databases. Use of this tool constitutes acceptance of all risks associated with database access, reading, usage, and modifications.
|
|
20
|
-
|
|
21
19
|
# Context Engineering Agent
|
|
22
20
|
|
|
23
|
-
The **Context Engineering Agent** is an AI coding agent plugin designed to run in developer agent harnesses (such as Claude Code, Antigravity, or Gemini CLI). It generates, evaluates, and iteratively tunes tailored context artifacts (`ContextSets` comprising `Templates`, `Facets`, and `Value Searches`) to enrich database schemas for **Gemini Data Analytics's data agent developer platform tools**, supporting both **relational SQL** and **Graph Query Language (GQL)** across [AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/data-agent-overview), Cloud SQL ([PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/data-agent-overview) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/data-agent-overview)), and [Cloud Spanner (GoogleSQL
|
|
21
|
+
The **Context Engineering Agent** is an AI coding agent plugin designed to run in developer agent harnesses (such as Claude Code, Antigravity, or Gemini CLI). It generates, evaluates, and iteratively tunes tailored context artifacts (`ContextSets` comprising `Templates`, `Facets`, and `Value Searches`) to enrich database schemas for **Gemini Data Analytics's data agent developer platform tools**, supporting both **relational SQL** and **Graph Query Language (GQL)** across [AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/data-agent-overview), Cloud SQL ([PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/data-agent-overview) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/data-agent-overview)), and [Cloud Spanner (GoogleSQL, Spanner Graph, and PostgreSQL)](https://docs.cloud.google.com/gemini/data-agents/querydata/spanner/data-agent-overview).
|
|
24
22
|
|
|
25
23
|
---
|
|
26
24
|
|
|
27
25
|
## Why Context Engineering?
|
|
28
26
|
|
|
29
|
-
When building data agents and natural language analytics interfaces, accurately translating user intent into database queries—whether relational SQL, pure GQL, or hybrid graph queries—is critical.
|
|
27
|
+
When building data agents and natural language analytics interfaces, accurately translating user intent into database queries—whether relational SQL (PostgreSQL, GoogleSQL, MySQL), pure GQL, or hybrid graph queries—is critical.
|
|
30
28
|
|
|
31
|
-
As outlined in **Build Context with Context Engineering Agent** ([AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/build-context-gemini-cli) | Cloud SQL: [PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/build-context-gemini-cli) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/build-context-gemini-cli) | [Spanner
|
|
29
|
+
As outlined in **Build Context with Context Engineering Agent** ([AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/build-context-gemini-cli) | Cloud SQL: [PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/build-context-gemini-cli) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/build-context-gemini-cli) | [Spanner](https://docs.cloud.google.com/gemini/data-agents/querydata/spanner/build-context-gemini-cli)), by optimizing a `ContextSet` to match your application's expected query stream, the **QueryData API** acts as a data agent tool capable of achieving **~100% NL-to-SQL/GQL translation accuracy with low latency**.
|
|
32
30
|
|
|
33
31
|
---
|
|
34
32
|
|
|
@@ -40,7 +38,7 @@ A `ContextSet` is the central artifact generated and managed by the agent, conta
|
|
|
40
38
|
* **Facets**: Reusable, modular query fragments (e.g., parameterized `WHERE` clauses, specialized join filters, or graph `MATCH` traversal patterns) linked to domain vocabulary.
|
|
41
39
|
* **Value Searches**: Specialized mapping queries that dynamically resolve user-supplied values (e.g., *"Lndn"*) to database records (*"London"*) via the capabilities of the underlying database, such as embedding search, AI operators, or trigram search on relational and graph property tables.
|
|
42
40
|
|
|
43
|
-
For full schema details, structure specifications, and dialect-specific JSON representations of `ContextSets`, see the official **Context Sets Overview** ([AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/context-sets-overview) | Cloud SQL: [PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/context-sets-overview) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/context-sets-overview) | [Spanner
|
|
41
|
+
For full schema details, structure specifications, and dialect-specific JSON representations of `ContextSets`, see the official **Context Sets Overview** ([AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/context-sets-overview) | Cloud SQL: [PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/context-sets-overview) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/context-sets-overview) | [Spanner](https://docs.cloud.google.com/gemini/data-agents/querydata/spanner/context-sets-overview)).
|
|
44
42
|
|
|
45
43
|
---
|
|
46
44
|
|
|
@@ -49,13 +47,13 @@ For full schema details, structure specifications, and dialect-specific JSON rep
|
|
|
49
47
|
Before getting started, prepare your GCP environment, required APIs (Data Analytics API, Gemini for Google Cloud API, Dataplex Universal Catalog API), IAM permissions, and database Data API settings.
|
|
50
48
|
|
|
51
49
|
Follow the step-by-step setup guide in the official documentation:
|
|
52
|
-
👉 **Prepare Your Environment**: ([AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/build-context-gemini-cli#prepare-your-environment) | Cloud SQL: [PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/build-context-gemini-cli#prepare-your-environment) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/build-context-gemini-cli#prepare-your-environment) | [Spanner
|
|
50
|
+
👉 **Prepare Your Environment**: ([AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/build-context-gemini-cli#prepare-your-environment) | Cloud SQL: [PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/build-context-gemini-cli#prepare-your-environment) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/build-context-gemini-cli#prepare-your-environment) | [Spanner](https://docs.cloud.google.com/gemini/data-agents/querydata/spanner/build-context-gemini-cli#prepare-your-environment))
|
|
53
51
|
|
|
54
52
|
---
|
|
55
53
|
|
|
56
54
|
## Primary Workflow Phases
|
|
57
55
|
|
|
58
|
-
The agent enables you to craft an optimized context for QueryData API through three primary phases
|
|
56
|
+
The agent enables you to craft an optimized context for QueryData API through three primary phases. Depending on your needs, you may also toggle the agent to skip phases. For example, if you have a dataset already, you can skip directly to context optimization "optimize context using dataset in \<file\>."
|
|
59
57
|
|
|
60
58
|
### Phase 1: Artifact Ingestion
|
|
61
59
|
*Why it matters: Without broader context on the application's goals and scope, AI models generate sterile queries based solely on database column names, missing how your users actually ask for information.*
|
|
@@ -80,8 +78,7 @@ The optimization loop creates an initial `ContextSet` and then iteratively refin
|
|
|
80
78
|
1. **Bootstrap**: Generate an initial baseline context.
|
|
81
79
|
2. **Evaluate**: Measure context effectiveness against a golden dataset.
|
|
82
80
|
3. **Hill-Climbing**: Perform gap analysis on failures and generate automated fixes.
|
|
83
|
-
4. **Iterate**: Apply the improved context and re-run evaluation to continuously improve quality.
|
|
84
|
-
5. **Final Validation** (Optional): Verify mutations against a separated test set to ensure generalization and prevent overfitting.
|
|
81
|
+
4. **Iterate**: Apply the improved context and re-run evaluation to continuously improve quality until we reach an optimal point.
|
|
85
82
|
|
|
86
83
|
*Note: While there is a typical ordering for these CUJs, the agent is flexible in how you want to execute. You can run the full pipeline end-to-end, trigger any individual phase, or ask for targeted changes to the `ContextSet`.*
|
|
87
84
|
|
{google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/README.md
RENAMED
|
@@ -1,16 +1,14 @@
|
|
|
1
|
-
This is not an officially supported Google product. This project is not eligible for the [Google Open Source Software Vulnerability Rewards Program](https://bughunters.google.com/open-source-security), [Google Cloud Platform/SecOps Terms of Service](https://cloud.google.com/terms), [How Gemini for Google Cloud uses your data](https://cloud.google.com/gemini/docs/discover/data-governance). This tool is provided "as is" without warranty of any kind. Users are solely responsible for understanding and managing the tool's interaction with their databases. Use of this tool constitutes acceptance of all risks associated with database access, reading, usage, and modifications.
|
|
2
|
-
|
|
3
1
|
# Context Engineering Agent
|
|
4
2
|
|
|
5
|
-
The **Context Engineering Agent** is an AI coding agent plugin designed to run in developer agent harnesses (such as Claude Code, Antigravity, or Gemini CLI). It generates, evaluates, and iteratively tunes tailored context artifacts (`ContextSets` comprising `Templates`, `Facets`, and `Value Searches`) to enrich database schemas for **Gemini Data Analytics's data agent developer platform tools**, supporting both **relational SQL** and **Graph Query Language (GQL)** across [AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/data-agent-overview), Cloud SQL ([PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/data-agent-overview) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/data-agent-overview)), and [Cloud Spanner (GoogleSQL
|
|
3
|
+
The **Context Engineering Agent** is an AI coding agent plugin designed to run in developer agent harnesses (such as Claude Code, Antigravity, or Gemini CLI). It generates, evaluates, and iteratively tunes tailored context artifacts (`ContextSets` comprising `Templates`, `Facets`, and `Value Searches`) to enrich database schemas for **Gemini Data Analytics's data agent developer platform tools**, supporting both **relational SQL** and **Graph Query Language (GQL)** across [AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/data-agent-overview), Cloud SQL ([PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/data-agent-overview) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/data-agent-overview)), and [Cloud Spanner (GoogleSQL, Spanner Graph, and PostgreSQL)](https://docs.cloud.google.com/gemini/data-agents/querydata/spanner/data-agent-overview).
|
|
6
4
|
|
|
7
5
|
---
|
|
8
6
|
|
|
9
7
|
## Why Context Engineering?
|
|
10
8
|
|
|
11
|
-
When building data agents and natural language analytics interfaces, accurately translating user intent into database queries—whether relational SQL, pure GQL, or hybrid graph queries—is critical.
|
|
9
|
+
When building data agents and natural language analytics interfaces, accurately translating user intent into database queries—whether relational SQL (PostgreSQL, GoogleSQL, MySQL), pure GQL, or hybrid graph queries—is critical.
|
|
12
10
|
|
|
13
|
-
As outlined in **Build Context with Context Engineering Agent** ([AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/build-context-gemini-cli) | Cloud SQL: [PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/build-context-gemini-cli) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/build-context-gemini-cli) | [Spanner
|
|
11
|
+
As outlined in **Build Context with Context Engineering Agent** ([AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/build-context-gemini-cli) | Cloud SQL: [PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/build-context-gemini-cli) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/build-context-gemini-cli) | [Spanner](https://docs.cloud.google.com/gemini/data-agents/querydata/spanner/build-context-gemini-cli)), by optimizing a `ContextSet` to match your application's expected query stream, the **QueryData API** acts as a data agent tool capable of achieving **~100% NL-to-SQL/GQL translation accuracy with low latency**.
|
|
14
12
|
|
|
15
13
|
---
|
|
16
14
|
|
|
@@ -22,7 +20,7 @@ A `ContextSet` is the central artifact generated and managed by the agent, conta
|
|
|
22
20
|
* **Facets**: Reusable, modular query fragments (e.g., parameterized `WHERE` clauses, specialized join filters, or graph `MATCH` traversal patterns) linked to domain vocabulary.
|
|
23
21
|
* **Value Searches**: Specialized mapping queries that dynamically resolve user-supplied values (e.g., *"Lndn"*) to database records (*"London"*) via the capabilities of the underlying database, such as embedding search, AI operators, or trigram search on relational and graph property tables.
|
|
24
22
|
|
|
25
|
-
For full schema details, structure specifications, and dialect-specific JSON representations of `ContextSets`, see the official **Context Sets Overview** ([AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/context-sets-overview) | Cloud SQL: [PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/context-sets-overview) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/context-sets-overview) | [Spanner
|
|
23
|
+
For full schema details, structure specifications, and dialect-specific JSON representations of `ContextSets`, see the official **Context Sets Overview** ([AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/context-sets-overview) | Cloud SQL: [PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/context-sets-overview) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/context-sets-overview) | [Spanner](https://docs.cloud.google.com/gemini/data-agents/querydata/spanner/context-sets-overview)).
|
|
26
24
|
|
|
27
25
|
---
|
|
28
26
|
|
|
@@ -31,13 +29,13 @@ For full schema details, structure specifications, and dialect-specific JSON rep
|
|
|
31
29
|
Before getting started, prepare your GCP environment, required APIs (Data Analytics API, Gemini for Google Cloud API, Dataplex Universal Catalog API), IAM permissions, and database Data API settings.
|
|
32
30
|
|
|
33
31
|
Follow the step-by-step setup guide in the official documentation:
|
|
34
|
-
👉 **Prepare Your Environment**: ([AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/build-context-gemini-cli#prepare-your-environment) | Cloud SQL: [PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/build-context-gemini-cli#prepare-your-environment) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/build-context-gemini-cli#prepare-your-environment) | [Spanner
|
|
32
|
+
👉 **Prepare Your Environment**: ([AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/build-context-gemini-cli#prepare-your-environment) | Cloud SQL: [PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/build-context-gemini-cli#prepare-your-environment) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/build-context-gemini-cli#prepare-your-environment) | [Spanner](https://docs.cloud.google.com/gemini/data-agents/querydata/spanner/build-context-gemini-cli#prepare-your-environment))
|
|
35
33
|
|
|
36
34
|
---
|
|
37
35
|
|
|
38
36
|
## Primary Workflow Phases
|
|
39
37
|
|
|
40
|
-
The agent enables you to craft an optimized context for QueryData API through three primary phases
|
|
38
|
+
The agent enables you to craft an optimized context for QueryData API through three primary phases. Depending on your needs, you may also toggle the agent to skip phases. For example, if you have a dataset already, you can skip directly to context optimization "optimize context using dataset in \<file\>."
|
|
41
39
|
|
|
42
40
|
### Phase 1: Artifact Ingestion
|
|
43
41
|
*Why it matters: Without broader context on the application's goals and scope, AI models generate sterile queries based solely on database column names, missing how your users actually ask for information.*
|
|
@@ -62,8 +60,7 @@ The optimization loop creates an initial `ContextSet` and then iteratively refin
|
|
|
62
60
|
1. **Bootstrap**: Generate an initial baseline context.
|
|
63
61
|
2. **Evaluate**: Measure context effectiveness against a golden dataset.
|
|
64
62
|
3. **Hill-Climbing**: Perform gap analysis on failures and generate automated fixes.
|
|
65
|
-
4. **Iterate**: Apply the improved context and re-run evaluation to continuously improve quality.
|
|
66
|
-
5. **Final Validation** (Optional): Verify mutations against a separated test set to ensure generalization and prevent overfitting.
|
|
63
|
+
4. **Iterate**: Apply the improved context and re-run evaluation to continuously improve quality until we reach an optimal point.
|
|
67
64
|
|
|
68
65
|
*Note: While there is a typical ordering for these CUJs, the agent is flexible in how you want to execute. You can run the full pipeline end-to-end, trigger any individual phase, or ask for targeted changes to the `ContextSet`.*
|
|
69
66
|
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
[project]
|
|
2
2
|
name = "google-cloud-db-context-engineering"
|
|
3
|
-
version = "0.7.
|
|
3
|
+
version = "0.7.3"
|
|
4
4
|
description = "A FastMCP server for generating natural language to SQL templates from database schemas."
|
|
5
5
|
readme = "README.md"
|
|
6
6
|
requires-python = ">=3.12"
|
|
@@ -46,7 +46,7 @@ dev = [
|
|
|
46
46
|
|
|
47
47
|
[tool.db-context-engineering]
|
|
48
48
|
toolbox_version = "1.4.0"
|
|
49
|
-
evalbench_version = "1.
|
|
49
|
+
evalbench_version = "1.17.0"
|
|
50
50
|
|
|
51
51
|
[tool.ruff]
|
|
52
52
|
line-length = 88
|
|
@@ -48,7 +48,7 @@ def mutate_context_set(file_path: str, mutations: list[Mutation]) -> None:
|
|
|
48
48
|
context_set = context.ContextSet()
|
|
49
49
|
else:
|
|
50
50
|
try:
|
|
51
|
-
with open(file_path) as f:
|
|
51
|
+
with open(file_path, encoding="utf-8") as f:
|
|
52
52
|
raw_data = json.load(f)
|
|
53
53
|
context_set = context.ContextSet.model_validate(raw_data)
|
|
54
54
|
except ValidationError as e:
|
|
@@ -126,7 +126,7 @@ def mutate_context_set(file_path: str, mutations: list[Mutation]) -> None:
|
|
|
126
126
|
# 4. Save validated ContextSet
|
|
127
127
|
try:
|
|
128
128
|
os.makedirs(os.path.dirname(os.path.abspath(file_path)), exist_ok=True)
|
|
129
|
-
with open(file_path, "w") as f:
|
|
129
|
+
with open(file_path, "w", encoding="utf-8") as f:
|
|
130
130
|
f.write(context_set.model_dump_json(indent=2, exclude_none=True))
|
|
131
131
|
except OSError as e:
|
|
132
132
|
raise RuntimeError(f"Error saving ContextSet to {file_path}: {e}") from e
|
|
@@ -0,0 +1,188 @@
|
|
|
1
|
+
import json
|
|
2
|
+
from typing import Any
|
|
3
|
+
|
|
4
|
+
from pydantic import ValidationError
|
|
5
|
+
|
|
6
|
+
from google.cloud.db_context_enrichment.model import context
|
|
7
|
+
|
|
8
|
+
_ATTR_TO_TYPE = {
|
|
9
|
+
"templates": "template",
|
|
10
|
+
"facets": "facet",
|
|
11
|
+
"value_searches": "value_search",
|
|
12
|
+
}
|
|
13
|
+
|
|
14
|
+
# ContextSet accepts camelCase aliases (Context Store returns them on download)
|
|
15
|
+
# and the deprecated "fragments" name. Map them back to the canonical attribute
|
|
16
|
+
# so every check below only has to handle one spelling.
|
|
17
|
+
_ALIAS_TO_ATTR = {
|
|
18
|
+
"valueSearches": "value_searches",
|
|
19
|
+
"fragments": "facets",
|
|
20
|
+
}
|
|
21
|
+
|
|
22
|
+
|
|
23
|
+
def validate_context_set(file_path: str) -> dict[str, Any]:
|
|
24
|
+
"""Validate a ContextSet file and return a structured report of issues.
|
|
25
|
+
|
|
26
|
+
Always returns a dict of the shape:
|
|
27
|
+
{
|
|
28
|
+
"valid": bool,
|
|
29
|
+
"issues": [
|
|
30
|
+
{
|
|
31
|
+
"location": {"type": str, "index": int} | None,
|
|
32
|
+
"message": str,
|
|
33
|
+
},
|
|
34
|
+
...
|
|
35
|
+
],
|
|
36
|
+
}
|
|
37
|
+
|
|
38
|
+
File access failures (missing path, permission denied, path is a directory,
|
|
39
|
+
etc.) are surfaced as a single issue rather than raised.
|
|
40
|
+
"""
|
|
41
|
+
try:
|
|
42
|
+
with open(file_path, encoding="utf-8") as f:
|
|
43
|
+
text = f.read()
|
|
44
|
+
except OSError as e:
|
|
45
|
+
return {
|
|
46
|
+
"valid": False,
|
|
47
|
+
"issues": [
|
|
48
|
+
_make_issue(f"Could not read file {file_path}: {type(e).__name__}: {e}")
|
|
49
|
+
],
|
|
50
|
+
}
|
|
51
|
+
|
|
52
|
+
if text.strip() == "":
|
|
53
|
+
return {"valid": True, "issues": []}
|
|
54
|
+
|
|
55
|
+
try:
|
|
56
|
+
raw = json.loads(text)
|
|
57
|
+
except json.JSONDecodeError as e:
|
|
58
|
+
snippet = e.doc[max(0, e.pos - 30) : e.pos + 30]
|
|
59
|
+
return {
|
|
60
|
+
"valid": False,
|
|
61
|
+
"issues": [_make_issue(f"File is not valid JSON: {e}. Near: {snippet!r}")],
|
|
62
|
+
}
|
|
63
|
+
|
|
64
|
+
if not isinstance(raw, dict):
|
|
65
|
+
return {
|
|
66
|
+
"valid": False,
|
|
67
|
+
"issues": [_make_issue("Top-level value must be a JSON object")],
|
|
68
|
+
}
|
|
69
|
+
|
|
70
|
+
raw = _normalize_aliases(raw)
|
|
71
|
+
|
|
72
|
+
issues: list[dict[str, Any]] = []
|
|
73
|
+
issues.extend(_check_pydantic(raw))
|
|
74
|
+
issues.extend(_check_duplicates(raw))
|
|
75
|
+
issues.extend(_check_value_search_value_param(raw))
|
|
76
|
+
|
|
77
|
+
return {"valid": len(issues) == 0, "issues": issues}
|
|
78
|
+
|
|
79
|
+
|
|
80
|
+
def _make_issue(message: str, location: dict[str, Any] | None = None) -> dict[str, Any]:
|
|
81
|
+
return {"location": location, "message": message}
|
|
82
|
+
|
|
83
|
+
|
|
84
|
+
def _normalize_aliases(raw: dict[str, Any]) -> dict[str, Any]:
|
|
85
|
+
"""Rewrite alias keys to their canonical ContextSet attribute names."""
|
|
86
|
+
return {_ALIAS_TO_ATTR.get(key, key): value for key, value in raw.items()}
|
|
87
|
+
|
|
88
|
+
|
|
89
|
+
def _check_pydantic(raw: dict[str, Any]) -> list[dict[str, Any]]:
|
|
90
|
+
issues: list[dict[str, Any]] = []
|
|
91
|
+
try:
|
|
92
|
+
context.ContextSet.model_validate(raw)
|
|
93
|
+
return issues
|
|
94
|
+
except ValidationError as e:
|
|
95
|
+
for err in e.errors():
|
|
96
|
+
loc = err.get("loc", ())
|
|
97
|
+
msg = err.get("msg", "validation error")
|
|
98
|
+
location: dict[str, Any] | None = None
|
|
99
|
+
descriptor = ""
|
|
100
|
+
|
|
101
|
+
if len(loc) >= 1 and loc[0] in _ATTR_TO_TYPE:
|
|
102
|
+
item_type = _ATTR_TO_TYPE[loc[0]]
|
|
103
|
+
if len(loc) >= 2 and isinstance(loc[1], int):
|
|
104
|
+
item_index = loc[1]
|
|
105
|
+
location = {"type": item_type, "index": item_index}
|
|
106
|
+
items = raw.get(loc[0])
|
|
107
|
+
if isinstance(items, list) and 0 <= item_index < len(items):
|
|
108
|
+
descriptor = _describe(item_type, items[item_index])
|
|
109
|
+
|
|
110
|
+
field_path = ".".join(str(p) for p in loc) if loc else ""
|
|
111
|
+
parts = [msg]
|
|
112
|
+
if field_path:
|
|
113
|
+
parts.append(f"at {field_path}")
|
|
114
|
+
if descriptor:
|
|
115
|
+
parts.append(descriptor)
|
|
116
|
+
full_msg = parts[0] + (
|
|
117
|
+
" (" + "; ".join(parts[1:]) + ")" if len(parts) > 1 else ""
|
|
118
|
+
)
|
|
119
|
+
|
|
120
|
+
issues.append(_make_issue(full_msg, location=location))
|
|
121
|
+
return issues
|
|
122
|
+
|
|
123
|
+
|
|
124
|
+
def _check_duplicates(raw: dict[str, Any]) -> list[dict[str, Any]]:
|
|
125
|
+
issues: list[dict[str, Any]] = []
|
|
126
|
+
for attr, item_type in _ATTR_TO_TYPE.items():
|
|
127
|
+
items = raw.get(attr)
|
|
128
|
+
if not isinstance(items, list):
|
|
129
|
+
continue
|
|
130
|
+
seen: dict[str, int] = {}
|
|
131
|
+
for idx, item in enumerate(items):
|
|
132
|
+
try:
|
|
133
|
+
key = _canonical(item)
|
|
134
|
+
except (TypeError, ValueError):
|
|
135
|
+
continue
|
|
136
|
+
if key in seen:
|
|
137
|
+
descriptor = _describe(item_type, item)
|
|
138
|
+
suffix = f" ({descriptor})" if descriptor else ""
|
|
139
|
+
issues.append(
|
|
140
|
+
_make_issue(
|
|
141
|
+
f"Exact duplicate of {item_type} at index {seen[key]}{suffix}",
|
|
142
|
+
location={"type": item_type, "index": idx},
|
|
143
|
+
)
|
|
144
|
+
)
|
|
145
|
+
else:
|
|
146
|
+
seen[key] = idx
|
|
147
|
+
return issues
|
|
148
|
+
|
|
149
|
+
|
|
150
|
+
def _check_value_search_value_param(raw: dict[str, Any]) -> list[dict[str, Any]]:
|
|
151
|
+
issues: list[dict[str, Any]] = []
|
|
152
|
+
items = raw.get("value_searches")
|
|
153
|
+
if not isinstance(items, list):
|
|
154
|
+
return issues
|
|
155
|
+
for idx, item in enumerate(items):
|
|
156
|
+
if not isinstance(item, dict):
|
|
157
|
+
continue
|
|
158
|
+
query = item.get("query")
|
|
159
|
+
if isinstance(query, str) and "$value" not in query:
|
|
160
|
+
descriptor = _describe("value_search", item)
|
|
161
|
+
suffix = f" ({descriptor})" if descriptor else ""
|
|
162
|
+
issues.append(
|
|
163
|
+
_make_issue(
|
|
164
|
+
f"value_search query must reference the $value parameter{suffix}",
|
|
165
|
+
location={"type": "value_search", "index": idx},
|
|
166
|
+
)
|
|
167
|
+
)
|
|
168
|
+
return issues
|
|
169
|
+
|
|
170
|
+
|
|
171
|
+
def _describe(item_type: str, item: Any) -> str:
|
|
172
|
+
"""Short human-readable identifier for an item, e.g. "intent: 'active users'"."""
|
|
173
|
+
if not isinstance(item, dict):
|
|
174
|
+
return ""
|
|
175
|
+
if item_type == "template":
|
|
176
|
+
nl = item.get("nl_query")
|
|
177
|
+
return f"nl_query: {nl!r}" if nl else ""
|
|
178
|
+
if item_type == "facet":
|
|
179
|
+
intent = item.get("intent")
|
|
180
|
+
return f"intent: {intent!r}" if intent else ""
|
|
181
|
+
if item_type == "value_search":
|
|
182
|
+
ct = item.get("concept_type")
|
|
183
|
+
return f"concept_type: {ct!r}" if ct else ""
|
|
184
|
+
return ""
|
|
185
|
+
|
|
186
|
+
|
|
187
|
+
def _canonical(obj: Any) -> str:
|
|
188
|
+
return json.dumps(obj, sort_keys=True, ensure_ascii=False)
|
|
@@ -9,10 +9,10 @@ class SpannerConfigGenerator(BaseDBConfigGenerator):
|
|
|
9
9
|
"""
|
|
10
10
|
Dedicated generator mapping properties to explicit Spanner configuration
|
|
11
11
|
topologies utilized by both EvalBench binaries and GDA Context objects.
|
|
12
|
+
Supports both GoogleSQL and PostgreSQL dialects.
|
|
12
13
|
"""
|
|
13
14
|
|
|
14
15
|
SOURCE_TYPE = "spanner"
|
|
15
|
-
DIALECT = "spanner_gsql"
|
|
16
16
|
REQUIRED_FIELDS = BaseDBConfigGenerator.REQUIRED_FIELDS | {
|
|
17
17
|
"project",
|
|
18
18
|
"instance",
|
|
@@ -25,6 +25,40 @@ class SpannerConfigGenerator(BaseDBConfigGenerator):
|
|
|
25
25
|
self.instance = params.get("instance")
|
|
26
26
|
self.database = params.get("database")
|
|
27
27
|
|
|
28
|
+
raw_dialect = (
|
|
29
|
+
params.get("dialect")
|
|
30
|
+
or params.get("engine")
|
|
31
|
+
or params.get("database_dialect")
|
|
32
|
+
)
|
|
33
|
+
if not raw_dialect and params.get("type") in ("spanner-postgres", "spanner-pg"):
|
|
34
|
+
raw_dialect = "POSTGRESQL"
|
|
35
|
+
|
|
36
|
+
if raw_dialect:
|
|
37
|
+
normalized = str(raw_dialect).strip().lower().replace("-", "_")
|
|
38
|
+
if normalized in (
|
|
39
|
+
"postgresql",
|
|
40
|
+
"postgres",
|
|
41
|
+
"spanner_pg",
|
|
42
|
+
"pg",
|
|
43
|
+
"spanner_postgres",
|
|
44
|
+
):
|
|
45
|
+
self.engine = "POSTGRESQL"
|
|
46
|
+
self._dialect = "spanner_pg"
|
|
47
|
+
elif normalized in ("google_sql", "googlesql", "spanner_gsql", "gsql"):
|
|
48
|
+
self.engine = "GOOGLE_SQL"
|
|
49
|
+
self._dialect = "spanner_gsql"
|
|
50
|
+
else:
|
|
51
|
+
raise ValueError(
|
|
52
|
+
f"Unsupported Spanner dialect/engine: '{raw_dialect}'. Must be 'GOOGLE_SQL' or 'POSTGRESQL'."
|
|
53
|
+
)
|
|
54
|
+
else:
|
|
55
|
+
self.engine = "GOOGLE_SQL"
|
|
56
|
+
self._dialect = "spanner_gsql"
|
|
57
|
+
|
|
58
|
+
@property
|
|
59
|
+
def DIALECT(self) -> str:
|
|
60
|
+
return self._dialect
|
|
61
|
+
|
|
28
62
|
def generate_db_config(self) -> str:
|
|
29
63
|
db_type = "spanner"
|
|
30
64
|
db_path = f"projects/{self.project}/instances/{self.instance}/databases/{self.database}"
|
|
@@ -44,12 +78,16 @@ class SpannerConfigGenerator(BaseDBConfigGenerator):
|
|
|
44
78
|
|
|
45
79
|
def build_datasource_reference(self, context_set_id: str) -> dict[str, Any]:
|
|
46
80
|
database_ref: dict[str, Any] = {
|
|
47
|
-
"engine":
|
|
81
|
+
"engine": self.engine,
|
|
48
82
|
"project_id": self.project,
|
|
49
83
|
"instance_id": self.instance,
|
|
50
84
|
"database_id": self.database,
|
|
51
85
|
}
|
|
52
86
|
if graph_ids := self.params.get("graph_ids"):
|
|
87
|
+
if self.engine == "POSTGRESQL":
|
|
88
|
+
raise ValueError(
|
|
89
|
+
"graph_ids is not supported for Spanner PostgreSQL dialect"
|
|
90
|
+
)
|
|
53
91
|
if not isinstance(graph_ids, list) or not all(
|
|
54
92
|
isinstance(g, str) for g in graph_ids
|
|
55
93
|
):
|
|
@@ -76,10 +76,10 @@ def _extract_toolbox_params(
|
|
|
76
76
|
with open(toolbox_config_path) as f:
|
|
77
77
|
content = f.read()
|
|
78
78
|
interpolated = _interpolate_env_vars(content)
|
|
79
|
-
docs = yaml.safe_load_all(interpolated)
|
|
79
|
+
docs = [doc for doc in yaml.safe_load_all(interpolated) if doc]
|
|
80
|
+
|
|
81
|
+
source_doc = None
|
|
80
82
|
for doc in docs:
|
|
81
|
-
if not doc:
|
|
82
|
-
continue
|
|
83
83
|
if (
|
|
84
84
|
doc.get("kind") == "source"
|
|
85
85
|
and doc.get("name") == toolbox_source_name
|
|
@@ -88,11 +88,24 @@ def _extract_toolbox_params(
|
|
|
88
88
|
raise ValueError(
|
|
89
89
|
f"Selected source '{toolbox_source_name}' is missing the 'type' field."
|
|
90
90
|
)
|
|
91
|
-
|
|
91
|
+
source_doc = doc
|
|
92
|
+
break
|
|
93
|
+
|
|
94
|
+
if not source_doc:
|
|
95
|
+
raise ValueError(
|
|
96
|
+
f"Could not find a 'kind: source' named '{toolbox_source_name}' in {toolbox_config_path}"
|
|
97
|
+
)
|
|
98
|
+
|
|
99
|
+
# For Spanner sources, state.md is the authoritative single source of truth for graph_ids
|
|
100
|
+
# (QueryData API requires explicit graph_ids in model_config.yaml, whereas tools.yaml
|
|
101
|
+
# only configures MCP Toolbox runtime tools and parameters).
|
|
102
|
+
if source_doc.get("type") == "spanner":
|
|
103
|
+
state_md_dir = os.path.dirname(toolbox_config_path)
|
|
104
|
+
state_md_path = os.path.join(state_md_dir, "state.md")
|
|
105
|
+
if graph_ids := _parse_graph_ids_from_state_md(state_md_path):
|
|
106
|
+
source_doc["graph_ids"] = graph_ids
|
|
92
107
|
|
|
93
|
-
|
|
94
|
-
f"Could not find a 'kind: source' named '{toolbox_source_name}' in {toolbox_config_path}"
|
|
95
|
-
)
|
|
108
|
+
return source_doc
|
|
96
109
|
|
|
97
110
|
except FileNotFoundError:
|
|
98
111
|
raise ValueError(f"Config file not found: {toolbox_config_path}")
|
|
@@ -104,6 +117,72 @@ def _extract_toolbox_params(
|
|
|
104
117
|
raise ValueError(f"Failed to parse {toolbox_config_path} as YAML: {e}")
|
|
105
118
|
|
|
106
119
|
|
|
120
|
+
def _parse_graph_ids_from_state_md(state_md_path: str) -> list[str] | None:
|
|
121
|
+
"""Parses graph_ids from state.md.
|
|
122
|
+
|
|
123
|
+
state.md is the authoritative source of truth for the database and graph scope
|
|
124
|
+
because QueryData API requires explicit graph_ids in model_config.yaml to evaluate
|
|
125
|
+
property graphs, whereas tools.yaml only configures MCP Toolbox runtime tools.
|
|
126
|
+
"""
|
|
127
|
+
if not os.path.exists(state_md_path):
|
|
128
|
+
return None
|
|
129
|
+
with open(state_md_path, encoding="utf-8") as f:
|
|
130
|
+
content = f.read()
|
|
131
|
+
match = re.search(
|
|
132
|
+
r"(?:^[ \t]*[-*][ \t]*)?\*\*Graph\s+Ids?:?\*\*:?[ \t]*([^\n]*)",
|
|
133
|
+
content,
|
|
134
|
+
re.MULTILINE | re.IGNORECASE,
|
|
135
|
+
)
|
|
136
|
+
if not match:
|
|
137
|
+
return None
|
|
138
|
+
|
|
139
|
+
val_str = match.group(1).split("#")[0].strip()
|
|
140
|
+
|
|
141
|
+
# If empty on the same line, check for multiline sub-bullets
|
|
142
|
+
if not val_str:
|
|
143
|
+
after_match = content[match.end() :]
|
|
144
|
+
bullet_items = []
|
|
145
|
+
for line in after_match.splitlines():
|
|
146
|
+
line_stripped = line.strip()
|
|
147
|
+
if not line_stripped:
|
|
148
|
+
continue
|
|
149
|
+
if line_stripped.startswith(("-", "*")) and not re.match(
|
|
150
|
+
r"^[-*]\s*\*\*", line_stripped
|
|
151
|
+
):
|
|
152
|
+
item = line_stripped.lstrip("-* ").split("#")[0].strip().strip("'\"`")
|
|
153
|
+
if item and item.lower() not in (
|
|
154
|
+
"none",
|
|
155
|
+
"n/a",
|
|
156
|
+
"null",
|
|
157
|
+
"nil",
|
|
158
|
+
"-",
|
|
159
|
+
):
|
|
160
|
+
bullet_items.append(item)
|
|
161
|
+
elif (
|
|
162
|
+
line_stripped.startswith("#")
|
|
163
|
+
or line_stripped.startswith("- **")
|
|
164
|
+
or line_stripped.startswith("* **")
|
|
165
|
+
):
|
|
166
|
+
break
|
|
167
|
+
else:
|
|
168
|
+
break
|
|
169
|
+
return bullet_items if bullet_items else None
|
|
170
|
+
|
|
171
|
+
# Handle explicit empty / none indicators
|
|
172
|
+
if val_str.lower() in ("none", "n/a", "null", "nil", "-", "[]", ""):
|
|
173
|
+
return None
|
|
174
|
+
|
|
175
|
+
if val_str.startswith("[") and val_str.endswith("]"):
|
|
176
|
+
val_str = val_str[1:-1]
|
|
177
|
+
graphs = [
|
|
178
|
+
g.strip().strip("'\"`")
|
|
179
|
+
for g in val_str.split(",")
|
|
180
|
+
if g.strip().strip("'\"`")
|
|
181
|
+
and g.strip().strip("'\"`").lower() not in ("none", "n/a", "null", "nil", "-")
|
|
182
|
+
]
|
|
183
|
+
return graphs if graphs else None
|
|
184
|
+
|
|
185
|
+
|
|
107
186
|
def _interpolate_env_vars(raw_yaml: str) -> str:
|
|
108
187
|
"""Replaces ${ENV_NAME} or ${ENV_NAME:default_value} with environment variables."""
|
|
109
188
|
# Matches ${VAR_NAME} or ${VAR_NAME:fallback}
|
|
@@ -133,6 +212,8 @@ def _get_db_generator(params: dict[str, Any]) -> BaseDBConfigGenerator:
|
|
|
133
212
|
PostgresConfigGenerator.SOURCE_TYPE: PostgresConfigGenerator,
|
|
134
213
|
MySQLConfigGenerator.SOURCE_TYPE: MySQLConfigGenerator,
|
|
135
214
|
SpannerConfigGenerator.SOURCE_TYPE: SpannerConfigGenerator,
|
|
215
|
+
"spanner-postgres": SpannerConfigGenerator,
|
|
216
|
+
"spanner-pg": SpannerConfigGenerator,
|
|
136
217
|
}
|
|
137
218
|
|
|
138
219
|
if source_type not in generators:
|
|
@@ -6,6 +6,7 @@ from fastmcp import FastMCP
|
|
|
6
6
|
from google.cloud.db_context_enrichment.common import (
|
|
7
7
|
context_mutator,
|
|
8
8
|
context_store_client,
|
|
9
|
+
context_validator,
|
|
9
10
|
)
|
|
10
11
|
from google.cloud.db_context_enrichment.dataset import dataset_generator
|
|
11
12
|
from google.cloud.db_context_enrichment.evaluate import (
|
|
@@ -258,6 +259,30 @@ def mutate_context_set(
|
|
|
258
259
|
return f"Error applying mutations: {str(e)}"
|
|
259
260
|
|
|
260
261
|
|
|
262
|
+
@mcp.tool
|
|
263
|
+
def validate_context_set(file_path: str) -> str:
|
|
264
|
+
"""
|
|
265
|
+
Validate a ContextSet JSON file for structural and convention issues. Reports issues only — does not fix them. The caller (agent) is expected to apply fixes via `mutate_context_set`, then re-run validation until `valid` is true.
|
|
266
|
+
|
|
267
|
+
Args:
|
|
268
|
+
file_path: Absolute path to the ContextSet file.
|
|
269
|
+
|
|
270
|
+
Returns:
|
|
271
|
+
A JSON string of the shape:
|
|
272
|
+
{
|
|
273
|
+
"valid": bool,
|
|
274
|
+
"issues": [
|
|
275
|
+
{
|
|
276
|
+
"location": {"type": "template" | "facet" | "value_search", "index": int} | null,
|
|
277
|
+
"message": str
|
|
278
|
+
},
|
|
279
|
+
...
|
|
280
|
+
]
|
|
281
|
+
}
|
|
282
|
+
"""
|
|
283
|
+
return json.dumps(context_validator.validate_context_set(file_path), indent=2)
|
|
284
|
+
|
|
285
|
+
|
|
261
286
|
@mcp.tool
|
|
262
287
|
async def read_evaluation_result(
|
|
263
288
|
run_folder_path: str, offset: int = 0, batch_size: int = 10
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
Metadata-Version: 2.4
|
|
2
2
|
Name: google-cloud-db-context-engineering
|
|
3
|
-
Version: 0.7.
|
|
3
|
+
Version: 0.7.3
|
|
4
4
|
Summary: A FastMCP server for generating natural language to SQL templates from database schemas.
|
|
5
5
|
Requires-Python: >=3.12
|
|
6
6
|
Description-Content-Type: text/markdown
|
|
@@ -16,19 +16,17 @@ Requires-Dist: pytest; extra == "test"
|
|
|
16
16
|
Requires-Dist: pytest-asyncio; extra == "test"
|
|
17
17
|
Dynamic: license-file
|
|
18
18
|
|
|
19
|
-
This is not an officially supported Google product. This project is not eligible for the [Google Open Source Software Vulnerability Rewards Program](https://bughunters.google.com/open-source-security), [Google Cloud Platform/SecOps Terms of Service](https://cloud.google.com/terms), [How Gemini for Google Cloud uses your data](https://cloud.google.com/gemini/docs/discover/data-governance). This tool is provided "as is" without warranty of any kind. Users are solely responsible for understanding and managing the tool's interaction with their databases. Use of this tool constitutes acceptance of all risks associated with database access, reading, usage, and modifications.
|
|
20
|
-
|
|
21
19
|
# Context Engineering Agent
|
|
22
20
|
|
|
23
|
-
The **Context Engineering Agent** is an AI coding agent plugin designed to run in developer agent harnesses (such as Claude Code, Antigravity, or Gemini CLI). It generates, evaluates, and iteratively tunes tailored context artifacts (`ContextSets` comprising `Templates`, `Facets`, and `Value Searches`) to enrich database schemas for **Gemini Data Analytics's data agent developer platform tools**, supporting both **relational SQL** and **Graph Query Language (GQL)** across [AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/data-agent-overview), Cloud SQL ([PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/data-agent-overview) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/data-agent-overview)), and [Cloud Spanner (GoogleSQL
|
|
21
|
+
The **Context Engineering Agent** is an AI coding agent plugin designed to run in developer agent harnesses (such as Claude Code, Antigravity, or Gemini CLI). It generates, evaluates, and iteratively tunes tailored context artifacts (`ContextSets` comprising `Templates`, `Facets`, and `Value Searches`) to enrich database schemas for **Gemini Data Analytics's data agent developer platform tools**, supporting both **relational SQL** and **Graph Query Language (GQL)** across [AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/data-agent-overview), Cloud SQL ([PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/data-agent-overview) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/data-agent-overview)), and [Cloud Spanner (GoogleSQL, Spanner Graph, and PostgreSQL)](https://docs.cloud.google.com/gemini/data-agents/querydata/spanner/data-agent-overview).
|
|
24
22
|
|
|
25
23
|
---
|
|
26
24
|
|
|
27
25
|
## Why Context Engineering?
|
|
28
26
|
|
|
29
|
-
When building data agents and natural language analytics interfaces, accurately translating user intent into database queries—whether relational SQL, pure GQL, or hybrid graph queries—is critical.
|
|
27
|
+
When building data agents and natural language analytics interfaces, accurately translating user intent into database queries—whether relational SQL (PostgreSQL, GoogleSQL, MySQL), pure GQL, or hybrid graph queries—is critical.
|
|
30
28
|
|
|
31
|
-
As outlined in **Build Context with Context Engineering Agent** ([AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/build-context-gemini-cli) | Cloud SQL: [PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/build-context-gemini-cli) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/build-context-gemini-cli) | [Spanner
|
|
29
|
+
As outlined in **Build Context with Context Engineering Agent** ([AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/build-context-gemini-cli) | Cloud SQL: [PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/build-context-gemini-cli) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/build-context-gemini-cli) | [Spanner](https://docs.cloud.google.com/gemini/data-agents/querydata/spanner/build-context-gemini-cli)), by optimizing a `ContextSet` to match your application's expected query stream, the **QueryData API** acts as a data agent tool capable of achieving **~100% NL-to-SQL/GQL translation accuracy with low latency**.
|
|
32
30
|
|
|
33
31
|
---
|
|
34
32
|
|
|
@@ -40,7 +38,7 @@ A `ContextSet` is the central artifact generated and managed by the agent, conta
|
|
|
40
38
|
* **Facets**: Reusable, modular query fragments (e.g., parameterized `WHERE` clauses, specialized join filters, or graph `MATCH` traversal patterns) linked to domain vocabulary.
|
|
41
39
|
* **Value Searches**: Specialized mapping queries that dynamically resolve user-supplied values (e.g., *"Lndn"*) to database records (*"London"*) via the capabilities of the underlying database, such as embedding search, AI operators, or trigram search on relational and graph property tables.
|
|
42
40
|
|
|
43
|
-
For full schema details, structure specifications, and dialect-specific JSON representations of `ContextSets`, see the official **Context Sets Overview** ([AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/context-sets-overview) | Cloud SQL: [PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/context-sets-overview) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/context-sets-overview) | [Spanner
|
|
41
|
+
For full schema details, structure specifications, and dialect-specific JSON representations of `ContextSets`, see the official **Context Sets Overview** ([AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/context-sets-overview) | Cloud SQL: [PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/context-sets-overview) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/context-sets-overview) | [Spanner](https://docs.cloud.google.com/gemini/data-agents/querydata/spanner/context-sets-overview)).
|
|
44
42
|
|
|
45
43
|
---
|
|
46
44
|
|
|
@@ -49,13 +47,13 @@ For full schema details, structure specifications, and dialect-specific JSON rep
|
|
|
49
47
|
Before getting started, prepare your GCP environment, required APIs (Data Analytics API, Gemini for Google Cloud API, Dataplex Universal Catalog API), IAM permissions, and database Data API settings.
|
|
50
48
|
|
|
51
49
|
Follow the step-by-step setup guide in the official documentation:
|
|
52
|
-
👉 **Prepare Your Environment**: ([AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/build-context-gemini-cli#prepare-your-environment) | Cloud SQL: [PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/build-context-gemini-cli#prepare-your-environment) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/build-context-gemini-cli#prepare-your-environment) | [Spanner
|
|
50
|
+
👉 **Prepare Your Environment**: ([AlloyDB](https://docs.cloud.google.com/gemini/data-agents/querydata/alloydb/build-context-gemini-cli#prepare-your-environment) | Cloud SQL: [PostgreSQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-postgres/build-context-gemini-cli#prepare-your-environment) / [MySQL](https://docs.cloud.google.com/gemini/data-agents/querydata/sql-mysql/build-context-gemini-cli#prepare-your-environment) | [Spanner](https://docs.cloud.google.com/gemini/data-agents/querydata/spanner/build-context-gemini-cli#prepare-your-environment))
|
|
53
51
|
|
|
54
52
|
---
|
|
55
53
|
|
|
56
54
|
## Primary Workflow Phases
|
|
57
55
|
|
|
58
|
-
The agent enables you to craft an optimized context for QueryData API through three primary phases
|
|
56
|
+
The agent enables you to craft an optimized context for QueryData API through three primary phases. Depending on your needs, you may also toggle the agent to skip phases. For example, if you have a dataset already, you can skip directly to context optimization "optimize context using dataset in \<file\>."
|
|
59
57
|
|
|
60
58
|
### Phase 1: Artifact Ingestion
|
|
61
59
|
*Why it matters: Without broader context on the application's goals and scope, AI models generate sterile queries based solely on database column names, missing how your users actually ask for information.*
|
|
@@ -80,8 +78,7 @@ The optimization loop creates an initial `ContextSet` and then iteratively refin
|
|
|
80
78
|
1. **Bootstrap**: Generate an initial baseline context.
|
|
81
79
|
2. **Evaluate**: Measure context effectiveness against a golden dataset.
|
|
82
80
|
3. **Hill-Climbing**: Perform gap analysis on failures and generate automated fixes.
|
|
83
|
-
4. **Iterate**: Apply the improved context and re-run evaluation to continuously improve quality.
|
|
84
|
-
5. **Final Validation** (Optional): Verify mutations against a separated test set to ensure generalization and prevent overfitting.
|
|
81
|
+
4. **Iterate**: Apply the improved context and re-run evaluation to continuously improve quality until we reach an optimal point.
|
|
85
82
|
|
|
86
83
|
*Note: While there is a typical ordering for these CUJs, the agent is flexible in how you want to execute. You can run the full pipeline end-to-end, trigger any individual phase, or ask for targeted changes to the `ContextSet`.*
|
|
87
84
|
|
|
@@ -7,6 +7,7 @@ src/google/cloud/db_context_enrichment/common/__init__.py
|
|
|
7
7
|
src/google/cloud/db_context_enrichment/common/config.py
|
|
8
8
|
src/google/cloud/db_context_enrichment/common/context_mutator.py
|
|
9
9
|
src/google/cloud/db_context_enrichment/common/context_store_client.py
|
|
10
|
+
src/google/cloud/db_context_enrichment/common/context_validator.py
|
|
10
11
|
src/google/cloud/db_context_enrichment/dataset/__init__.py
|
|
11
12
|
src/google/cloud/db_context_enrichment/dataset/dataset_generator.py
|
|
12
13
|
src/google/cloud/db_context_enrichment/evaluate/__init__.py
|
{google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/LICENSE
RENAMED
|
File without changes
|
{google_cloud_db_context_engineering-0.7.2 → google_cloud_db_context_engineering-0.7.3}/setup.cfg
RENAMED
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|
|
File without changes
|