pgvector-template 0.2.2__tar.gz → 0.3.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (29) hide show
  1. pgvector_template-0.3.0/PKG-INFO +152 -0
  2. pgvector_template-0.3.0/README.md +127 -0
  3. {pgvector_template-0.2.2 → pgvector_template-0.3.0}/pgvector_template/core/document.py +3 -1
  4. {pgvector_template-0.2.2 → pgvector_template-0.3.0}/pgvector_template/core/search.py +81 -52
  5. {pgvector_template-0.2.2 → pgvector_template-0.3.0}/pgvector_template/db/document_db.py +2 -1
  6. pgvector_template-0.3.0/pgvector_template/models/search.py +67 -0
  7. {pgvector_template-0.2.2 → pgvector_template-0.3.0}/pgvector_template/service/document_service.py +2 -0
  8. pgvector_template-0.3.0/pgvector_template/utils/__init__.py +0 -0
  9. pgvector_template-0.3.0/pgvector_template/utils/metadata_filter.py +81 -0
  10. pgvector_template-0.3.0/pgvector_template.egg-info/PKG-INFO +152 -0
  11. {pgvector_template-0.2.2 → pgvector_template-0.3.0}/pgvector_template.egg-info/SOURCES.txt +4 -1
  12. {pgvector_template-0.2.2 → pgvector_template-0.3.0}/pyproject.toml +1 -1
  13. pgvector_template-0.2.2/PKG-INFO +0 -46
  14. pgvector_template-0.2.2/README.md +0 -21
  15. pgvector_template-0.2.2/pgvector_template.egg-info/PKG-INFO +0 -46
  16. {pgvector_template-0.2.2 → pgvector_template-0.3.0}/LICENSE +0 -0
  17. {pgvector_template-0.2.2 → pgvector_template-0.3.0}/pgvector_template/__init__.py +0 -0
  18. {pgvector_template-0.2.2 → pgvector_template-0.3.0}/pgvector_template/core/__init__.py +0 -0
  19. {pgvector_template-0.2.2 → pgvector_template-0.3.0}/pgvector_template/core/embedder.py +0 -0
  20. {pgvector_template-0.2.2 → pgvector_template-0.3.0}/pgvector_template/core/manager.py +0 -0
  21. {pgvector_template-0.2.2 → pgvector_template-0.3.0}/pgvector_template/db/__init__.py +0 -0
  22. {pgvector_template-0.2.2 → pgvector_template-0.3.0}/pgvector_template/db/connection.py +0 -0
  23. {pgvector_template-0.2.2/pgvector_template/utils → pgvector_template-0.3.0/pgvector_template/models}/__init__.py +0 -0
  24. {pgvector_template-0.2.2 → pgvector_template-0.3.0}/pgvector_template/service/__init__.py +0 -0
  25. {pgvector_template-0.2.2 → pgvector_template-0.3.0}/pgvector_template/types.py +0 -0
  26. {pgvector_template-0.2.2 → pgvector_template-0.3.0}/pgvector_template.egg-info/dependency_links.txt +0 -0
  27. {pgvector_template-0.2.2 → pgvector_template-0.3.0}/pgvector_template.egg-info/requires.txt +0 -0
  28. {pgvector_template-0.2.2 → pgvector_template-0.3.0}/pgvector_template.egg-info/top_level.txt +0 -0
  29. {pgvector_template-0.2.2 → pgvector_template-0.3.0}/setup.cfg +0 -0
@@ -0,0 +1,152 @@
1
+ Metadata-Version: 2.1
2
+ Name: pgvector-template
3
+ Version: 0.3.0
4
+ Summary: Template library for flexible PGVector RAG implementations
5
+ Author-email: DL <v49t9zpqd@mozmail.com>
6
+ License: MIT
7
+ Project-URL: Homepage, https://github.com/DavidLiuGit/PGVector-Template
8
+ Classifier: Programming Language :: Python :: 3
9
+ Classifier: License :: OSI Approved :: MIT License
10
+ Classifier: Operating System :: OS Independent
11
+ Requires-Python: >=3.11
12
+ Description-Content-Type: text/markdown
13
+ License-File: LICENSE
14
+ Requires-Dist: pgvector>=0.2.0
15
+ Requires-Dist: pydantic<3.0,>=2.11
16
+ Requires-Dist: sqlalchemy>=2.0.0
17
+ Requires-Dist: typing-extensions>=4.0.0
18
+ Provides-Extra: test
19
+ Requires-Dist: psycopg[binary]>=3.1.0; extra == "test"
20
+ Requires-Dist: pytest>=7.0.0; extra == "test"
21
+ Requires-Dist: pytest-cov>=4.0.0; extra == "test"
22
+ Requires-Dist: python-dotenv>=1.0.0; extra == "test"
23
+ Provides-Extra: dev
24
+ Requires-Dist: black>=23.0.0; extra == "dev"
25
+
26
+ # PGVector-Template
27
+
28
+ A flexible, production-ready template library for building Retrieval-Augmented Generation (RAG) applications using PostgreSQL with PGVector extensions.
29
+
30
+ ## Overview
31
+
32
+ PGVector-Template provides a robust foundation for implementing vector-based document storage and retrieval systems. It offers a clean abstraction layer over PostgreSQL's PGVector extension, making it easy to build scalable RAG applications with proper document management, metadata handling, and efficient vector search capabilities.
33
+
34
+ ## Key Features
35
+
36
+ - **Flexible Document Model**: Abstract base classes for customizable document schemas
37
+ - **Vector Search**: Optimized HNSW indexing for fast similarity search
38
+ - **Metadata Management**: JSON-based flexible metadata with GIN indexing
39
+ - **Collection Support**: Organize documents into logical collections
40
+ - **Chunk Management**: Handle long content (refer to as **corpus**) by chunking it into smaller **documents**. Handle recovering the original corpus given its id
41
+ - **Database Abstraction**: Clean SQLAlchemy-based database layer, with an API to create schemas
42
+ - **Type Safety**: Full Pydantic validation and type hints
43
+ - **Production Ready**: Comprehensive testing and error handling
44
+
45
+ ## Architecture
46
+
47
+ The library is organized into several key components:
48
+
49
+ - **Core**: Document models, embedders, search functionality
50
+ - **Database**: Connection management and document database operations
51
+ - **Service**: High-level document service layer
52
+ - **Types**: Shared type definitions and schemas
53
+
54
+ ## Installation
55
+
56
+ ```bash
57
+ pip install pgvector-template
58
+ ```
59
+
60
+ Or add `pgvector-template` to your dependencies
61
+
62
+ ### Prerequisites
63
+
64
+ - Python 3.11+
65
+ - To execute tests: PostgreSQL with PGVector extension
66
+
67
+ ## Configuration
68
+
69
+ ### Database Setup
70
+
71
+ 1. Install PostgreSQL with PGVector extension
72
+ 2. Create your database and enable the vector extension:
73
+
74
+ ```sql
75
+ CREATE EXTENSION IF NOT EXISTS vector;
76
+ ```
77
+
78
+ 3. Set up your connection string in environment variables or pass directly to `DatabaseManager`
79
+
80
+ ### Environment Variables
81
+
82
+ For integration tests, create a `.env` file
83
+
84
+ ```bash
85
+ cp integ-tests/.env.example integ-tests/.env
86
+ ```
87
+
88
+ Specify envvars directly in the .env file. It is loaded automatically for integ tests.
89
+
90
+ ```bash
91
+ DATABASE_URL=postgresql://user:password@localhost:5432/test_db
92
+ ```
93
+
94
+ ## API Reference
95
+
96
+ ### Core Classes
97
+
98
+ - `BaseDocument`: Abstract document model with vector embedding support
99
+ - refer to table schema for explanation of the fields
100
+ - `BaseDocumentOptionalProps`: Optional properties for document creation
101
+ - `DatabaseManager`: Database connection and session management
102
+ - `DocumentDatabaseManager`: High-level document operations
103
+
104
+ ### Key Methods
105
+
106
+ - `BaseDocument.from_props()`: Create document instances from properties
107
+ - `DocumentDatabaseManager.insert_document()`: Store documents
108
+ - `DocumentDatabaseManager.search_similar()`: Vector similarity search
109
+
110
+ ## Testing
111
+
112
+ Install dependencies (preferably in a virtualenv) before running tests:
113
+ ```bash
114
+ pip install -e .[test]
115
+ ```
116
+
117
+ ### Unit Tests
118
+ ```bash
119
+ python -m unittest
120
+ ```
121
+
122
+ ### Integration Tests
123
+
124
+ Integration tests require a PostgreSQL database with PGVector extension. Set up your test database and configure the connection in `integ-tests/.env`:
125
+
126
+ ```bash
127
+ python -m unittest discover -s integ-tests
128
+ ```
129
+
130
+ ## Contributing
131
+
132
+ 1. Fork the repository
133
+ 2. Create a feature branch
134
+ 3. Make your changes with tests
135
+ 4. Run the test suite
136
+ 5. Submit a pull request
137
+
138
+ ### Development Setup
139
+
140
+ ```bash
141
+ pip install -e .[dev,test]
142
+ black . # Format code
143
+ ```
144
+
145
+ ## License
146
+
147
+ MIT License - see [LICENSE](LICENSE) file for details.
148
+
149
+ ## Links
150
+
151
+ - [GitHub Repository](https://github.com/DavidLiuGit/PGVector-Template)
152
+ - [PGVector Documentation](https://github.com/pgvector/pgvector)
@@ -0,0 +1,127 @@
1
+ # PGVector-Template
2
+
3
+ A flexible, production-ready template library for building Retrieval-Augmented Generation (RAG) applications using PostgreSQL with PGVector extensions.
4
+
5
+ ## Overview
6
+
7
+ PGVector-Template provides a robust foundation for implementing vector-based document storage and retrieval systems. It offers a clean abstraction layer over PostgreSQL's PGVector extension, making it easy to build scalable RAG applications with proper document management, metadata handling, and efficient vector search capabilities.
8
+
9
+ ## Key Features
10
+
11
+ - **Flexible Document Model**: Abstract base classes for customizable document schemas
12
+ - **Vector Search**: Optimized HNSW indexing for fast similarity search
13
+ - **Metadata Management**: JSON-based flexible metadata with GIN indexing
14
+ - **Collection Support**: Organize documents into logical collections
15
+ - **Chunk Management**: Handle long content (refer to as **corpus**) by chunking it into smaller **documents**. Handle recovering the original corpus given its id
16
+ - **Database Abstraction**: Clean SQLAlchemy-based database layer, with an API to create schemas
17
+ - **Type Safety**: Full Pydantic validation and type hints
18
+ - **Production Ready**: Comprehensive testing and error handling
19
+
20
+ ## Architecture
21
+
22
+ The library is organized into several key components:
23
+
24
+ - **Core**: Document models, embedders, search functionality
25
+ - **Database**: Connection management and document database operations
26
+ - **Service**: High-level document service layer
27
+ - **Types**: Shared type definitions and schemas
28
+
29
+ ## Installation
30
+
31
+ ```bash
32
+ pip install pgvector-template
33
+ ```
34
+
35
+ Or add `pgvector-template` to your dependencies
36
+
37
+ ### Prerequisites
38
+
39
+ - Python 3.11+
40
+ - To execute tests: PostgreSQL with PGVector extension
41
+
42
+ ## Configuration
43
+
44
+ ### Database Setup
45
+
46
+ 1. Install PostgreSQL with PGVector extension
47
+ 2. Create your database and enable the vector extension:
48
+
49
+ ```sql
50
+ CREATE EXTENSION IF NOT EXISTS vector;
51
+ ```
52
+
53
+ 3. Set up your connection string in environment variables or pass directly to `DatabaseManager`
54
+
55
+ ### Environment Variables
56
+
57
+ For integration tests, create a `.env` file
58
+
59
+ ```bash
60
+ cp integ-tests/.env.example integ-tests/.env
61
+ ```
62
+
63
+ Specify envvars directly in the .env file. It is loaded automatically for integ tests.
64
+
65
+ ```bash
66
+ DATABASE_URL=postgresql://user:password@localhost:5432/test_db
67
+ ```
68
+
69
+ ## API Reference
70
+
71
+ ### Core Classes
72
+
73
+ - `BaseDocument`: Abstract document model with vector embedding support
74
+ - refer to table schema for explanation of the fields
75
+ - `BaseDocumentOptionalProps`: Optional properties for document creation
76
+ - `DatabaseManager`: Database connection and session management
77
+ - `DocumentDatabaseManager`: High-level document operations
78
+
79
+ ### Key Methods
80
+
81
+ - `BaseDocument.from_props()`: Create document instances from properties
82
+ - `DocumentDatabaseManager.insert_document()`: Store documents
83
+ - `DocumentDatabaseManager.search_similar()`: Vector similarity search
84
+
85
+ ## Testing
86
+
87
+ Install dependencies (preferably in a virtualenv) before running tests:
88
+ ```bash
89
+ pip install -e .[test]
90
+ ```
91
+
92
+ ### Unit Tests
93
+ ```bash
94
+ python -m unittest
95
+ ```
96
+
97
+ ### Integration Tests
98
+
99
+ Integration tests require a PostgreSQL database with PGVector extension. Set up your test database and configure the connection in `integ-tests/.env`:
100
+
101
+ ```bash
102
+ python -m unittest discover -s integ-tests
103
+ ```
104
+
105
+ ## Contributing
106
+
107
+ 1. Fork the repository
108
+ 2. Create a feature branch
109
+ 3. Make your changes with tests
110
+ 4. Run the test suite
111
+ 5. Submit a pull request
112
+
113
+ ### Development Setup
114
+
115
+ ```bash
116
+ pip install -e .[dev,test]
117
+ black . # Format code
118
+ ```
119
+
120
+ ## License
121
+
122
+ MIT License - see [LICENSE](LICENSE) file for details.
123
+
124
+ ## Links
125
+
126
+ - [GitHub Repository](https://github.com/DavidLiuGit/PGVector-Template)
127
+ - [PGVector Documentation](https://github.com/pgvector/pgvector)
@@ -86,7 +86,7 @@ class BaseDocument(Base):
86
86
  """Flexible metadata as JSON"""
87
87
  origin_url = Column(String(2048), nullable=True)
88
88
  """Optional source URL"""
89
- language = Column(String(10), default="en")
89
+ language = Column(String(2), default="en")
90
90
  """Language of the content (ISO 639-1 code), e.g., 'en', 'es', 'zh'."""
91
91
  score = Column(Float, nullable=True)
92
92
  """Optional score assigned during ingestion (e.g., relevance, confidence)."""
@@ -190,6 +190,8 @@ class BaseDocumentMetadata(BaseModel):
190
190
  Base metadata structure.
191
191
  It is generally expected that every `BaseDocument`'s metadata follows this exact schema,
192
192
  without any extraneous properties, or any missing properties, to avoid ambiguity.
193
+ It is mandatory to include a description with every field/property.
194
+ May contain nested Pydantic models, but they may not be optional. Instead, set defaults.
193
195
  """
194
196
 
195
197
  document_type: str = Field(
@@ -1,59 +1,23 @@
1
- from dataclasses import dataclass, asdict
2
- from datetime import datetime
3
1
  from logging import getLogger
4
2
  from typing import Any, Type, Sequence
5
3
 
6
- from pydantic import BaseModel, Field, model_validator
7
- from sqlalchemy import text, select, or_, and_
8
- from sqlalchemy.sql import Select
9
- from pgvector.sqlalchemy import Vector
4
+ from pydantic import BaseModel, Field
5
+ from sqlalchemy import select, or_, Integer, Float
6
+ from sqlalchemy.orm import Session
7
+ from sqlalchemy.sql import Select, ColumnElement
10
8
 
11
9
  from pgvector_template.core import (
12
10
  BaseEmbeddingProvider,
13
11
  BaseDocument,
14
12
  BaseDocumentMetadata,
15
13
  )
16
- from sqlalchemy.orm import Session
14
+ from pgvector_template.models.search import SearchQuery, MetadataFilter, RetrievalResult
15
+ from pgvector_template.utils.metadata_filter import validate_metadata_filters
17
16
 
18
17
 
19
18
  logger = getLogger(__name__)
20
19
 
21
20
 
22
- class SearchQuery(BaseModel):
23
- """Standardized search query structure. At least 1 search criterion is required."""
24
-
25
- text: str | None = None
26
- """String to approximate-search (using vector distance) in a semantic search."""
27
- keywords: list[str] | None = None
28
- """List of keywords to exact-match in a keyword search."""
29
- metadata_filters: dict[str, Any] | None = None
30
- """Strict metadata filters that must be matched."""
31
- date_range: tuple[datetime, datetime] | None = None
32
- """Retrieve/limit results based on created_at & updated_at timestamps"""
33
- limit: int = Field(
34
- ...,
35
- ge=1,
36
- )
37
- """Maximum number of results to return."""
38
-
39
- @model_validator(mode="after")
40
- def ensure_criterion(self):
41
- if not any([self.text, self.keywords, self.metadata_filters, self.date_range]):
42
- raise ValueError("At least one search criterion is required")
43
- return self
44
-
45
-
46
- @dataclass
47
- class RetrievalResult:
48
- """Standardized result structure for all retrieval operations"""
49
-
50
- document: BaseDocument
51
- score: float
52
-
53
- def to_dict(self) -> dict[str, Any]:
54
- return asdict(self)
55
-
56
-
57
21
  class BaseSearchClientConfig(BaseModel):
58
22
  """Config obj for `BaseSearchClient`."""
59
23
 
@@ -112,15 +76,15 @@ class BaseSearchClient:
112
76
  if query.text:
113
77
  db_query = self._apply_semantic_search(db_query, query)
114
78
  db_query = self._apply_keyword_search(db_query, query)
115
- # if query.metadata_filters:
116
- # db_query = self._apply_metadata_filters(db_query, query)
79
+ if query.metadata_filters:
80
+ db_query = self._apply_metadata_filters(db_query, query)
117
81
  db_query = db_query.limit(query.limit)
118
82
 
119
83
  # execute query and return results
120
84
  results = self.session.scalars(db_query).all()
121
85
  return self._convert_to_retrieval_results(results)
122
86
 
123
- def _apply_semantic_search(self, query: Select, search_query: SearchQuery) -> Select:
87
+ def _apply_semantic_search(self, db_query: Select, search_query: SearchQuery) -> Select:
124
88
  """Apply semantic (vector) search criteria to the query.
125
89
  `embedding_provider` must be provided at instantiation, or an `ValueError` will be raised.
126
90
  In PGVector, `<=>` operator is used to compare cosine distance. Lower = more similar.
@@ -133,9 +97,11 @@ class BaseSearchClient:
133
97
  Updated SQLAlchemy query with semantic search applied.
134
98
  """
135
99
  if not search_query.text:
136
- return query
100
+ return db_query
137
101
  query_embedding = self.embedding_provider.embed_text(search_query.text)
138
- return query.order_by(self.config.document_cls.embedding.cosine_distance(query_embedding))
102
+ return db_query.order_by(
103
+ self.config.document_cls.embedding.cosine_distance(query_embedding)
104
+ )
139
105
 
140
106
  def _apply_keyword_search(self, db_query: Select, search_query: SearchQuery) -> Select:
141
107
  """Apply keyword (full-text) search criteria to the query.
@@ -154,17 +120,81 @@ class BaseSearchClient:
154
120
  conditions.append(self.config.document_cls.content.ilike(f"%{keyword}%"))
155
121
  return db_query.where(or_(*conditions))
156
122
 
157
- def _apply_metadata_filters(self, query: Select, search_query: SearchQuery) -> Select:
158
- """Apply metadata filters to the query.
123
+ def _apply_metadata_filters(self, db_query: Select, search_query: SearchQuery) -> Select:
124
+ """
125
+ Apply metadata filters to the query. All condtions in `search_query.metadata_filters`,
126
+ if not None, are ANDed together. Metadata filters are applied against
127
+ `BaseDocument.document_metadata` JSONB field.
159
128
 
160
129
  Args:
161
130
  query: The base SQLAlchemy query.
162
- search_query: The search query containing metadata filters.
131
+ search_query: The `SearchQuery` containing metadata filters.
163
132
 
164
133
  Returns:
165
134
  Updated SQLAlchemy query with metadata filters applied.
166
135
  """
167
- raise NotImplementedError
136
+ if not search_query.metadata_filters:
137
+ return db_query
138
+
139
+ # strictly speaking, this validation step is optional. You may override this method to disable
140
+ try:
141
+ validate_metadata_filters(search_query.metadata_filters, self.config.document_metadata_cls)
142
+ except ValueError as e:
143
+ logger.warning(f"Metadata filter validation failed: {e}. Query success not guaranteed!")
144
+
145
+ conditions = [
146
+ self._build_metadata_filter_where_condition(filter_obj)
147
+ for filter_obj in search_query.metadata_filters
148
+ ]
149
+ return db_query.where(*conditions) if conditions else db_query
150
+
151
+ def _build_metadata_filter_where_condition(
152
+ self, filter_obj: MetadataFilter
153
+ ) -> ColumnElement[bool]:
154
+ """Build SQLAlchemy WHERE condition for a metadata filter."""
155
+ field_path = filter_obj.field_name.split(".")
156
+ metadata_col = self.config.document_cls.document_metadata
157
+
158
+ # Navigate to field
159
+ field_ref = metadata_col
160
+ for part in field_path:
161
+ field_ref = field_ref[part]
162
+
163
+ if filter_obj.condition == "eq":
164
+ return (
165
+ field_ref.astext == filter_obj.value
166
+ if isinstance(filter_obj.value, str)
167
+ else field_ref == filter_obj.value
168
+ )
169
+ elif filter_obj.condition in {"gt", "gte", "lt", "lte"}:
170
+ if isinstance(filter_obj.value, str):
171
+ field_text = field_ref.astext
172
+ else:
173
+ cast_type = Integer if isinstance(filter_obj.value, int) else Float
174
+ field_text = field_ref.astext.cast(cast_type)
175
+
176
+ if filter_obj.condition == "gt":
177
+ return field_text > filter_obj.value
178
+ elif filter_obj.condition == "gte":
179
+ return field_text >= filter_obj.value
180
+ elif filter_obj.condition == "lt":
181
+ return field_text < filter_obj.value
182
+ else: # lte
183
+ return field_text <= filter_obj.value
184
+ elif filter_obj.condition == "contains":
185
+ return field_ref.contains([filter_obj.value])
186
+ elif filter_obj.condition == "in":
187
+ return field_ref.astext.in_([str(v) for v in filter_obj.value])
188
+ elif filter_obj.condition == "exists":
189
+ if len(field_path) == 1:
190
+ return metadata_col.has_key(field_path[0])
191
+ else:
192
+ parent_ref = metadata_col
193
+ for part in field_path[:-1]:
194
+ parent_ref = parent_ref[part]
195
+ return parent_ref.has_key(field_path[-1])
196
+ else:
197
+ raise ValueError(f"Unsupported condition: {filter_obj.condition}")
168
198
 
169
199
  def _convert_to_retrieval_results(self, results: Sequence[Any]) -> list[RetrievalResult]:
170
200
  """Convert database results to RetrievalResult objects.
@@ -178,6 +208,5 @@ class BaseSearchClient:
178
208
  """
179
209
  retrieval_results = []
180
210
  for result in results:
181
- doc = result[0] if isinstance(result, tuple) else result
182
211
  retrieval_results.append(RetrievalResult(document=result, score=1.0))
183
212
  return retrieval_results
@@ -37,7 +37,7 @@ class DocumentDatabaseManager(DatabaseManager):
37
37
  self.schema_name = f"{self.SCHEMA_PREFIX}{schema_suffix}"
38
38
  self.document_classes = document_classes
39
39
 
40
- def setup(self) -> None:
40
+ def setup(self) -> str:
41
41
  """One-step setup: initialize connection, create schema and tables for all document classes"""
42
42
  self.initialize()
43
43
  self.create_schema(self.schema_name)
@@ -50,6 +50,7 @@ class DocumentDatabaseManager(DatabaseManager):
50
50
  self.logger.info(
51
51
  f"Document database setup complete for {self.schema_name} with {len(self.document_classes)} tables"
52
52
  )
53
+ return self.schema_name
53
54
 
54
55
 
55
56
  class TempDocumentDatabaseManager(DocumentDatabaseManager):
@@ -0,0 +1,67 @@
1
+ from dataclasses import dataclass, asdict
2
+ from datetime import datetime
3
+ from typing import Any, Literal
4
+
5
+ from pydantic import BaseModel, ConfigDict, Field, model_validator
6
+
7
+ from pgvector_template.core import BaseDocument
8
+
9
+
10
+ class MetadataFilter(BaseModel):
11
+ """
12
+ An object acting as a filter for an arbitrary `Metadata` dictionary/map/object.
13
+ """
14
+
15
+ field_name: str
16
+ """Field path in metadata. Use dot notation for nested fields (e.g., 'publication_info.journal')"""
17
+ condition: Literal["eq", "gt", "gte", "lt", "lte", "contains", "in", "exists"]
18
+ """
19
+ Comparison operator:
20
+ - eq=equal
21
+ - gt/gte=greater than/equal
22
+ - lt/lte=less than/equal
23
+ - contains=array contains values (accepts array)
24
+ - in=value in array
25
+ - exists=field exists
26
+ """
27
+ value: Any
28
+ """Value to compare against. Type should match field type (str, int, float, bool, list)"""
29
+
30
+ model_config = ConfigDict(use_attribute_docstrings=True)
31
+
32
+
33
+ class SearchQuery(BaseModel):
34
+ """Standardized search query structure. At least 1 search criterion is required."""
35
+
36
+ text: str | None = None
37
+ """String to match against using in a semantic search, i.e. using vector distance."""
38
+ keywords: list[str] = []
39
+ """List of keywords to **exact-match** in a keyword search."""
40
+ metadata_filters: list[MetadataFilter] = []
41
+ """List of metadata filter conditions that must be matched."""
42
+ date_range: tuple[datetime, datetime] | None = None
43
+ """Retrieve/limit results based on created_at & updated_at timestamps (i.e. database operations)"""
44
+ limit: int = Field(
45
+ ...,
46
+ ge=1,
47
+ )
48
+ """Maximum number of results to return."""
49
+
50
+ model_config = ConfigDict(use_attribute_docstrings=True)
51
+
52
+ @model_validator(mode="after")
53
+ def ensure_criterion(self):
54
+ if not any([self.text, self.keywords, self.metadata_filters, self.date_range]):
55
+ raise ValueError("At least one search criterion is required")
56
+ return self
57
+
58
+
59
+ @dataclass
60
+ class RetrievalResult:
61
+ """Standardized result structure for all retrieval operations"""
62
+
63
+ document: BaseDocument
64
+ score: float
65
+
66
+ def to_dict(self) -> dict[str, Any]:
67
+ return asdict(self)
@@ -50,8 +50,10 @@ class DocumentServiceConfig(BaseModel):
50
50
  # iff either config is an instance of their respective base config classes (and not a subclass)
51
51
  if type(self.corpus_manager_cfg) is BaseCorpusManagerConfig:
52
52
  self.corpus_manager_cfg.document_cls = self.document_cls
53
+ self.corpus_manager_cfg.document_metadata_cls = self.document_metadata_cls
53
54
  if type(self.search_client_cfg) is BaseSearchClientConfig:
54
55
  self.search_client_cfg.document_cls = self.document_cls
56
+ self.search_client_cfg.document_metadata_cls = self.document_metadata_cls
55
57
 
56
58
  # assign embedding_provider to CorpusManager & SearchClient configs
57
59
  if not self.corpus_manager_cfg.embedding_provider:
@@ -0,0 +1,81 @@
1
+ from typing import Type
2
+
3
+ from pgvector_template.core.document import BaseDocumentMetadata
4
+ from pgvector_template.models.search import MetadataFilter
5
+
6
+
7
+ def validate_metadata_filters(
8
+ filter_obj_list: list[MetadataFilter], metadata_cls: Type[BaseDocumentMetadata]
9
+ ) -> None:
10
+ """Validate a `list[MetadataFilter]` against schema and condition compatibility.
11
+
12
+ Args:
13
+ filter_obj_list (list[MetadataFilter]): _description_
14
+ metadata_cls (Type[BaseDocumentMetadata]): _description_
15
+ """
16
+ for metadata_filter_obj in filter_obj_list:
17
+ validate_metadata_filter(metadata_filter_obj, metadata_cls)
18
+
19
+
20
+ def validate_metadata_filter(
21
+ filter_obj: MetadataFilter, metadata_cls: Type[BaseDocumentMetadata]
22
+ ) -> None:
23
+ """Validate metadata filter against schema and condition compatibility.
24
+
25
+ Note: This validates the schema structure but cannot guarantee runtime data conformity.
26
+ JSONB field access like field_ref[part] will succeed even if the actual data doesn't match the schema.
27
+
28
+ Args:
29
+ filter_obj: The metadata filter to validate
30
+ metadata_cls: The metadata class to validate against
31
+
32
+ Raises:
33
+ ValueError: If field doesn't exist in schema or condition is incompatible with field type
34
+ """
35
+ field_path = filter_obj.field_name.split(".")
36
+ current_field_info = metadata_cls.model_fields
37
+
38
+ # Navigate nested structure
39
+ for i, part in enumerate(field_path):
40
+ if part not in current_field_info:
41
+ raise ValueError(f"Field '{filter_obj.field_name}' not found in metadata schema")
42
+
43
+ field_info = current_field_info[part]
44
+ field_type = field_info.annotation
45
+
46
+ # Handle nested models
47
+ if not field_type:
48
+ raise ValueError(f"Field '{filter_obj.field_name}' not found in metadata schema")
49
+ elif hasattr(field_type, "model_fields"):
50
+ current_field_info = field_type.model_fields
51
+ elif i < len(field_path) - 1:
52
+ raise ValueError(
53
+ f"Cannot navigate into non-model field '{part}' in path '{filter_obj.field_name}'"
54
+ )
55
+
56
+ # Validate condition compatibility with final field type
57
+ validate_condition_compatibility(field_type, filter_obj.condition)
58
+
59
+
60
+ def validate_condition_compatibility(field_type: Type, condition: str) -> None:
61
+ """Validate that condition is compatible with field type."""
62
+ # Extract base type from Optional/Union types
63
+ origin = getattr(field_type, "__origin__", None)
64
+ if origin is not None:
65
+ args = getattr(field_type, "__args__", ())
66
+ if origin is list:
67
+ field_type = list
68
+ elif len(args) > 0:
69
+ field_type = args[0] # First non-None type
70
+
71
+ valid_conditions = {
72
+ str: {"eq", "gt", "gte", "lt", "lte", "in", "exists"},
73
+ int: {"eq", "gt", "gte", "lt", "lte", "exists"},
74
+ float: {"eq", "gt", "gte", "lt", "lte", "exists"},
75
+ bool: {"eq", "exists"},
76
+ list: {"contains", "in", "exists"},
77
+ }
78
+
79
+ allowed = valid_conditions.get(field_type, {"eq", "exists"})
80
+ if condition not in allowed:
81
+ raise ValueError(f"Condition '{condition}' not valid for field type {field_type.__name__}")
@@ -0,0 +1,152 @@
1
+ Metadata-Version: 2.1
2
+ Name: pgvector-template
3
+ Version: 0.3.0
4
+ Summary: Template library for flexible PGVector RAG implementations
5
+ Author-email: DL <v49t9zpqd@mozmail.com>
6
+ License: MIT
7
+ Project-URL: Homepage, https://github.com/DavidLiuGit/PGVector-Template
8
+ Classifier: Programming Language :: Python :: 3
9
+ Classifier: License :: OSI Approved :: MIT License
10
+ Classifier: Operating System :: OS Independent
11
+ Requires-Python: >=3.11
12
+ Description-Content-Type: text/markdown
13
+ License-File: LICENSE
14
+ Requires-Dist: pgvector>=0.2.0
15
+ Requires-Dist: pydantic<3.0,>=2.11
16
+ Requires-Dist: sqlalchemy>=2.0.0
17
+ Requires-Dist: typing-extensions>=4.0.0
18
+ Provides-Extra: test
19
+ Requires-Dist: psycopg[binary]>=3.1.0; extra == "test"
20
+ Requires-Dist: pytest>=7.0.0; extra == "test"
21
+ Requires-Dist: pytest-cov>=4.0.0; extra == "test"
22
+ Requires-Dist: python-dotenv>=1.0.0; extra == "test"
23
+ Provides-Extra: dev
24
+ Requires-Dist: black>=23.0.0; extra == "dev"
25
+
26
+ # PGVector-Template
27
+
28
+ A flexible, production-ready template library for building Retrieval-Augmented Generation (RAG) applications using PostgreSQL with PGVector extensions.
29
+
30
+ ## Overview
31
+
32
+ PGVector-Template provides a robust foundation for implementing vector-based document storage and retrieval systems. It offers a clean abstraction layer over PostgreSQL's PGVector extension, making it easy to build scalable RAG applications with proper document management, metadata handling, and efficient vector search capabilities.
33
+
34
+ ## Key Features
35
+
36
+ - **Flexible Document Model**: Abstract base classes for customizable document schemas
37
+ - **Vector Search**: Optimized HNSW indexing for fast similarity search
38
+ - **Metadata Management**: JSON-based flexible metadata with GIN indexing
39
+ - **Collection Support**: Organize documents into logical collections
40
+ - **Chunk Management**: Handle long content (refer to as **corpus**) by chunking it into smaller **documents**. Handle recovering the original corpus given its id
41
+ - **Database Abstraction**: Clean SQLAlchemy-based database layer, with an API to create schemas
42
+ - **Type Safety**: Full Pydantic validation and type hints
43
+ - **Production Ready**: Comprehensive testing and error handling
44
+
45
+ ## Architecture
46
+
47
+ The library is organized into several key components:
48
+
49
+ - **Core**: Document models, embedders, search functionality
50
+ - **Database**: Connection management and document database operations
51
+ - **Service**: High-level document service layer
52
+ - **Types**: Shared type definitions and schemas
53
+
54
+ ## Installation
55
+
56
+ ```bash
57
+ pip install pgvector-template
58
+ ```
59
+
60
+ Or add `pgvector-template` to your dependencies
61
+
62
+ ### Prerequisites
63
+
64
+ - Python 3.11+
65
+ - To execute tests: PostgreSQL with PGVector extension
66
+
67
+ ## Configuration
68
+
69
+ ### Database Setup
70
+
71
+ 1. Install PostgreSQL with PGVector extension
72
+ 2. Create your database and enable the vector extension:
73
+
74
+ ```sql
75
+ CREATE EXTENSION IF NOT EXISTS vector;
76
+ ```
77
+
78
+ 3. Set up your connection string in environment variables or pass directly to `DatabaseManager`
79
+
80
+ ### Environment Variables
81
+
82
+ For integration tests, create a `.env` file
83
+
84
+ ```bash
85
+ cp integ-tests/.env.example integ-tests/.env
86
+ ```
87
+
88
+ Specify envvars directly in the .env file. It is loaded automatically for integ tests.
89
+
90
+ ```bash
91
+ DATABASE_URL=postgresql://user:password@localhost:5432/test_db
92
+ ```
93
+
94
+ ## API Reference
95
+
96
+ ### Core Classes
97
+
98
+ - `BaseDocument`: Abstract document model with vector embedding support
99
+ - refer to table schema for explanation of the fields
100
+ - `BaseDocumentOptionalProps`: Optional properties for document creation
101
+ - `DatabaseManager`: Database connection and session management
102
+ - `DocumentDatabaseManager`: High-level document operations
103
+
104
+ ### Key Methods
105
+
106
+ - `BaseDocument.from_props()`: Create document instances from properties
107
+ - `DocumentDatabaseManager.insert_document()`: Store documents
108
+ - `DocumentDatabaseManager.search_similar()`: Vector similarity search
109
+
110
+ ## Testing
111
+
112
+ Install dependencies (preferably in a virtualenv) before running tests:
113
+ ```bash
114
+ pip install -e .[test]
115
+ ```
116
+
117
+ ### Unit Tests
118
+ ```bash
119
+ python -m unittest
120
+ ```
121
+
122
+ ### Integration Tests
123
+
124
+ Integration tests require a PostgreSQL database with PGVector extension. Set up your test database and configure the connection in `integ-tests/.env`:
125
+
126
+ ```bash
127
+ python -m unittest discover -s integ-tests
128
+ ```
129
+
130
+ ## Contributing
131
+
132
+ 1. Fork the repository
133
+ 2. Create a feature branch
134
+ 3. Make your changes with tests
135
+ 4. Run the test suite
136
+ 5. Submit a pull request
137
+
138
+ ### Development Setup
139
+
140
+ ```bash
141
+ pip install -e .[dev,test]
142
+ black . # Format code
143
+ ```
144
+
145
+ ## License
146
+
147
+ MIT License - see [LICENSE](LICENSE) file for details.
148
+
149
+ ## Links
150
+
151
+ - [GitHub Repository](https://github.com/DavidLiuGit/PGVector-Template)
152
+ - [PGVector Documentation](https://github.com/pgvector/pgvector)
@@ -16,6 +16,9 @@ pgvector_template/core/search.py
16
16
  pgvector_template/db/__init__.py
17
17
  pgvector_template/db/connection.py
18
18
  pgvector_template/db/document_db.py
19
+ pgvector_template/models/__init__.py
20
+ pgvector_template/models/search.py
19
21
  pgvector_template/service/__init__.py
20
22
  pgvector_template/service/document_service.py
21
- pgvector_template/utils/__init__.py
23
+ pgvector_template/utils/__init__.py
24
+ pgvector_template/utils/metadata_filter.py
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
4
4
 
5
5
  [project]
6
6
  name = "pgvector-template"
7
- version = "0.2.2"
7
+ version = "0.3.0"
8
8
  description = "Template library for flexible PGVector RAG implementations"
9
9
  authors = [{ name="DL", email="v49t9zpqd@mozmail.com" }]
10
10
  license = { text = "MIT" }
@@ -1,46 +0,0 @@
1
- Metadata-Version: 2.1
2
- Name: pgvector-template
3
- Version: 0.2.2
4
- Summary: Template library for flexible PGVector RAG implementations
5
- Author-email: DL <v49t9zpqd@mozmail.com>
6
- License: MIT
7
- Project-URL: Homepage, https://github.com/DavidLiuGit/PGVector-Template
8
- Classifier: Programming Language :: Python :: 3
9
- Classifier: License :: OSI Approved :: MIT License
10
- Classifier: Operating System :: OS Independent
11
- Requires-Python: >=3.11
12
- Description-Content-Type: text/markdown
13
- License-File: LICENSE
14
- Requires-Dist: pgvector>=0.2.0
15
- Requires-Dist: pydantic<3.0,>=2.11
16
- Requires-Dist: sqlalchemy>=2.0.0
17
- Requires-Dist: typing-extensions>=4.0.0
18
- Provides-Extra: test
19
- Requires-Dist: psycopg[binary]>=3.1.0; extra == "test"
20
- Requires-Dist: pytest>=7.0.0; extra == "test"
21
- Requires-Dist: pytest-cov>=4.0.0; extra == "test"
22
- Requires-Dist: python-dotenv>=1.0.0; extra == "test"
23
- Provides-Extra: dev
24
- Requires-Dist: black>=23.0.0; extra == "dev"
25
-
26
- # PGVector-Template
27
-
28
- Template library for flexible PGVector RAG implementations
29
-
30
-
31
- ## Testing
32
-
33
- Install dependencies (preferably in a virtualenv) before running tests:
34
- ```bash
35
- pip install -e .[test]
36
- ```
37
-
38
- ### Unit tests
39
- ```bash
40
- python -m unittest
41
- ```
42
-
43
- ### Integration tests
44
- ```bash
45
- python -m unittest discover -s integ-tests
46
- ```
@@ -1,21 +0,0 @@
1
- # PGVector-Template
2
-
3
- Template library for flexible PGVector RAG implementations
4
-
5
-
6
- ## Testing
7
-
8
- Install dependencies (preferably in a virtualenv) before running tests:
9
- ```bash
10
- pip install -e .[test]
11
- ```
12
-
13
- ### Unit tests
14
- ```bash
15
- python -m unittest
16
- ```
17
-
18
- ### Integration tests
19
- ```bash
20
- python -m unittest discover -s integ-tests
21
- ```
@@ -1,46 +0,0 @@
1
- Metadata-Version: 2.1
2
- Name: pgvector-template
3
- Version: 0.2.2
4
- Summary: Template library for flexible PGVector RAG implementations
5
- Author-email: DL <v49t9zpqd@mozmail.com>
6
- License: MIT
7
- Project-URL: Homepage, https://github.com/DavidLiuGit/PGVector-Template
8
- Classifier: Programming Language :: Python :: 3
9
- Classifier: License :: OSI Approved :: MIT License
10
- Classifier: Operating System :: OS Independent
11
- Requires-Python: >=3.11
12
- Description-Content-Type: text/markdown
13
- License-File: LICENSE
14
- Requires-Dist: pgvector>=0.2.0
15
- Requires-Dist: pydantic<3.0,>=2.11
16
- Requires-Dist: sqlalchemy>=2.0.0
17
- Requires-Dist: typing-extensions>=4.0.0
18
- Provides-Extra: test
19
- Requires-Dist: psycopg[binary]>=3.1.0; extra == "test"
20
- Requires-Dist: pytest>=7.0.0; extra == "test"
21
- Requires-Dist: pytest-cov>=4.0.0; extra == "test"
22
- Requires-Dist: python-dotenv>=1.0.0; extra == "test"
23
- Provides-Extra: dev
24
- Requires-Dist: black>=23.0.0; extra == "dev"
25
-
26
- # PGVector-Template
27
-
28
- Template library for flexible PGVector RAG implementations
29
-
30
-
31
- ## Testing
32
-
33
- Install dependencies (preferably in a virtualenv) before running tests:
34
- ```bash
35
- pip install -e .[test]
36
- ```
37
-
38
- ### Unit tests
39
- ```bash
40
- python -m unittest
41
- ```
42
-
43
- ### Integration tests
44
- ```bash
45
- python -m unittest discover -s integ-tests
46
- ```