pgvector-template 0.2.3__tar.gz → 0.3.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (29) hide show
  1. pgvector_template-0.3.1/PKG-INFO +152 -0
  2. pgvector_template-0.3.1/README.md +127 -0
  3. {pgvector_template-0.2.3 → pgvector_template-0.3.1}/pgvector_template/core/document.py +2 -0
  4. {pgvector_template-0.2.3 → pgvector_template-0.3.1}/pgvector_template/core/search.py +83 -53
  5. pgvector_template-0.3.1/pgvector_template/models/search.py +76 -0
  6. pgvector_template-0.3.1/pgvector_template/utils/__init__.py +0 -0
  7. pgvector_template-0.3.1/pgvector_template/utils/metadata_filter.py +81 -0
  8. pgvector_template-0.3.1/pgvector_template.egg-info/PKG-INFO +152 -0
  9. {pgvector_template-0.2.3 → pgvector_template-0.3.1}/pgvector_template.egg-info/SOURCES.txt +4 -1
  10. {pgvector_template-0.2.3 → pgvector_template-0.3.1}/pyproject.toml +1 -1
  11. pgvector_template-0.2.3/PKG-INFO +0 -46
  12. pgvector_template-0.2.3/README.md +0 -21
  13. pgvector_template-0.2.3/pgvector_template.egg-info/PKG-INFO +0 -46
  14. {pgvector_template-0.2.3 → pgvector_template-0.3.1}/LICENSE +0 -0
  15. {pgvector_template-0.2.3 → pgvector_template-0.3.1}/pgvector_template/__init__.py +0 -0
  16. {pgvector_template-0.2.3 → pgvector_template-0.3.1}/pgvector_template/core/__init__.py +0 -0
  17. {pgvector_template-0.2.3 → pgvector_template-0.3.1}/pgvector_template/core/embedder.py +0 -0
  18. {pgvector_template-0.2.3 → pgvector_template-0.3.1}/pgvector_template/core/manager.py +0 -0
  19. {pgvector_template-0.2.3 → pgvector_template-0.3.1}/pgvector_template/db/__init__.py +0 -0
  20. {pgvector_template-0.2.3 → pgvector_template-0.3.1}/pgvector_template/db/connection.py +0 -0
  21. {pgvector_template-0.2.3 → pgvector_template-0.3.1}/pgvector_template/db/document_db.py +0 -0
  22. {pgvector_template-0.2.3/pgvector_template/utils → pgvector_template-0.3.1/pgvector_template/models}/__init__.py +0 -0
  23. {pgvector_template-0.2.3 → pgvector_template-0.3.1}/pgvector_template/service/__init__.py +0 -0
  24. {pgvector_template-0.2.3 → pgvector_template-0.3.1}/pgvector_template/service/document_service.py +0 -0
  25. {pgvector_template-0.2.3 → pgvector_template-0.3.1}/pgvector_template/types.py +0 -0
  26. {pgvector_template-0.2.3 → pgvector_template-0.3.1}/pgvector_template.egg-info/dependency_links.txt +0 -0
  27. {pgvector_template-0.2.3 → pgvector_template-0.3.1}/pgvector_template.egg-info/requires.txt +0 -0
  28. {pgvector_template-0.2.3 → pgvector_template-0.3.1}/pgvector_template.egg-info/top_level.txt +0 -0
  29. {pgvector_template-0.2.3 → pgvector_template-0.3.1}/setup.cfg +0 -0
@@ -0,0 +1,152 @@
1
+ Metadata-Version: 2.1
2
+ Name: pgvector-template
3
+ Version: 0.3.1
4
+ Summary: Template library for flexible PGVector RAG implementations
5
+ Author-email: DL <v49t9zpqd@mozmail.com>
6
+ License: MIT
7
+ Project-URL: Homepage, https://github.com/DavidLiuGit/PGVector-Template
8
+ Classifier: Programming Language :: Python :: 3
9
+ Classifier: License :: OSI Approved :: MIT License
10
+ Classifier: Operating System :: OS Independent
11
+ Requires-Python: >=3.11
12
+ Description-Content-Type: text/markdown
13
+ License-File: LICENSE
14
+ Requires-Dist: pgvector>=0.2.0
15
+ Requires-Dist: pydantic<3.0,>=2.11
16
+ Requires-Dist: sqlalchemy>=2.0.0
17
+ Requires-Dist: typing-extensions>=4.0.0
18
+ Provides-Extra: test
19
+ Requires-Dist: psycopg[binary]>=3.1.0; extra == "test"
20
+ Requires-Dist: pytest>=7.0.0; extra == "test"
21
+ Requires-Dist: pytest-cov>=4.0.0; extra == "test"
22
+ Requires-Dist: python-dotenv>=1.0.0; extra == "test"
23
+ Provides-Extra: dev
24
+ Requires-Dist: black>=23.0.0; extra == "dev"
25
+
26
+ # PGVector-Template
27
+
28
+ A flexible, production-ready template library for building Retrieval-Augmented Generation (RAG) applications using PostgreSQL with PGVector extensions.
29
+
30
+ ## Overview
31
+
32
+ PGVector-Template provides a robust foundation for implementing vector-based document storage and retrieval systems. It offers a clean abstraction layer over PostgreSQL's PGVector extension, making it easy to build scalable RAG applications with proper document management, metadata handling, and efficient vector search capabilities.
33
+
34
+ ## Key Features
35
+
36
+ - **Flexible Document Model**: Abstract base classes for customizable document schemas
37
+ - **Vector Search**: Optimized HNSW indexing for fast similarity search
38
+ - **Metadata Management**: JSON-based flexible metadata with GIN indexing
39
+ - **Collection Support**: Organize documents into logical collections
40
+ - **Chunk Management**: Handle long content (refer to as **corpus**) by chunking it into smaller **documents**. Handle recovering the original corpus given its id
41
+ - **Database Abstraction**: Clean SQLAlchemy-based database layer, with an API to create schemas
42
+ - **Type Safety**: Full Pydantic validation and type hints
43
+ - **Production Ready**: Comprehensive testing and error handling
44
+
45
+ ## Architecture
46
+
47
+ The library is organized into several key components:
48
+
49
+ - **Core**: Document models, embedders, search functionality
50
+ - **Database**: Connection management and document database operations
51
+ - **Service**: High-level document service layer
52
+ - **Types**: Shared type definitions and schemas
53
+
54
+ ## Installation
55
+
56
+ ```bash
57
+ pip install pgvector-template
58
+ ```
59
+
60
+ Or add `pgvector-template` to your dependencies
61
+
62
+ ### Prerequisites
63
+
64
+ - Python 3.11+
65
+ - To execute tests: PostgreSQL with PGVector extension
66
+
67
+ ## Configuration
68
+
69
+ ### Database Setup
70
+
71
+ 1. Install PostgreSQL with PGVector extension
72
+ 2. Create your database and enable the vector extension:
73
+
74
+ ```sql
75
+ CREATE EXTENSION IF NOT EXISTS vector;
76
+ ```
77
+
78
+ 3. Set up your connection string in environment variables or pass directly to `DatabaseManager`
79
+
80
+ ### Environment Variables
81
+
82
+ For integration tests, create a `.env` file
83
+
84
+ ```bash
85
+ cp integ-tests/.env.example integ-tests/.env
86
+ ```
87
+
88
+ Specify envvars directly in the .env file. It is loaded automatically for integ tests.
89
+
90
+ ```bash
91
+ DATABASE_URL=postgresql://user:password@localhost:5432/test_db
92
+ ```
93
+
94
+ ## API Reference
95
+
96
+ ### Core Classes
97
+
98
+ - `BaseDocument`: Abstract document model with vector embedding support
99
+ - refer to table schema for explanation of the fields
100
+ - `BaseDocumentOptionalProps`: Optional properties for document creation
101
+ - `DatabaseManager`: Database connection and session management
102
+ - `DocumentDatabaseManager`: High-level document operations
103
+
104
+ ### Key Methods
105
+
106
+ - `BaseDocument.from_props()`: Create document instances from properties
107
+ - `DocumentDatabaseManager.insert_document()`: Store documents
108
+ - `DocumentDatabaseManager.search_similar()`: Vector similarity search
109
+
110
+ ## Testing
111
+
112
+ Install dependencies (preferably in a virtualenv) before running tests:
113
+ ```bash
114
+ pip install -e .[test]
115
+ ```
116
+
117
+ ### Unit Tests
118
+ ```bash
119
+ python -m unittest
120
+ ```
121
+
122
+ ### Integration Tests
123
+
124
+ Integration tests require a PostgreSQL database with PGVector extension. Set up your test database and configure the connection in `integ-tests/.env`:
125
+
126
+ ```bash
127
+ python -m unittest discover -s integ-tests
128
+ ```
129
+
130
+ ## Contributing
131
+
132
+ 1. Fork the repository
133
+ 2. Create a feature branch
134
+ 3. Make your changes with tests
135
+ 4. Run the test suite
136
+ 5. Submit a pull request
137
+
138
+ ### Development Setup
139
+
140
+ ```bash
141
+ pip install -e .[dev,test]
142
+ black . # Format code
143
+ ```
144
+
145
+ ## License
146
+
147
+ MIT License - see [LICENSE](LICENSE) file for details.
148
+
149
+ ## Links
150
+
151
+ - [GitHub Repository](https://github.com/DavidLiuGit/PGVector-Template)
152
+ - [PGVector Documentation](https://github.com/pgvector/pgvector)
@@ -0,0 +1,127 @@
1
+ # PGVector-Template
2
+
3
+ A flexible, production-ready template library for building Retrieval-Augmented Generation (RAG) applications using PostgreSQL with PGVector extensions.
4
+
5
+ ## Overview
6
+
7
+ PGVector-Template provides a robust foundation for implementing vector-based document storage and retrieval systems. It offers a clean abstraction layer over PostgreSQL's PGVector extension, making it easy to build scalable RAG applications with proper document management, metadata handling, and efficient vector search capabilities.
8
+
9
+ ## Key Features
10
+
11
+ - **Flexible Document Model**: Abstract base classes for customizable document schemas
12
+ - **Vector Search**: Optimized HNSW indexing for fast similarity search
13
+ - **Metadata Management**: JSON-based flexible metadata with GIN indexing
14
+ - **Collection Support**: Organize documents into logical collections
15
+ - **Chunk Management**: Handle long content (refer to as **corpus**) by chunking it into smaller **documents**. Handle recovering the original corpus given its id
16
+ - **Database Abstraction**: Clean SQLAlchemy-based database layer, with an API to create schemas
17
+ - **Type Safety**: Full Pydantic validation and type hints
18
+ - **Production Ready**: Comprehensive testing and error handling
19
+
20
+ ## Architecture
21
+
22
+ The library is organized into several key components:
23
+
24
+ - **Core**: Document models, embedders, search functionality
25
+ - **Database**: Connection management and document database operations
26
+ - **Service**: High-level document service layer
27
+ - **Types**: Shared type definitions and schemas
28
+
29
+ ## Installation
30
+
31
+ ```bash
32
+ pip install pgvector-template
33
+ ```
34
+
35
+ Or add `pgvector-template` to your dependencies
36
+
37
+ ### Prerequisites
38
+
39
+ - Python 3.11+
40
+ - To execute tests: PostgreSQL with PGVector extension
41
+
42
+ ## Configuration
43
+
44
+ ### Database Setup
45
+
46
+ 1. Install PostgreSQL with PGVector extension
47
+ 2. Create your database and enable the vector extension:
48
+
49
+ ```sql
50
+ CREATE EXTENSION IF NOT EXISTS vector;
51
+ ```
52
+
53
+ 3. Set up your connection string in environment variables or pass directly to `DatabaseManager`
54
+
55
+ ### Environment Variables
56
+
57
+ For integration tests, create a `.env` file
58
+
59
+ ```bash
60
+ cp integ-tests/.env.example integ-tests/.env
61
+ ```
62
+
63
+ Specify envvars directly in the .env file. It is loaded automatically for integ tests.
64
+
65
+ ```bash
66
+ DATABASE_URL=postgresql://user:password@localhost:5432/test_db
67
+ ```
68
+
69
+ ## API Reference
70
+
71
+ ### Core Classes
72
+
73
+ - `BaseDocument`: Abstract document model with vector embedding support
74
+ - refer to table schema for explanation of the fields
75
+ - `BaseDocumentOptionalProps`: Optional properties for document creation
76
+ - `DatabaseManager`: Database connection and session management
77
+ - `DocumentDatabaseManager`: High-level document operations
78
+
79
+ ### Key Methods
80
+
81
+ - `BaseDocument.from_props()`: Create document instances from properties
82
+ - `DocumentDatabaseManager.insert_document()`: Store documents
83
+ - `DocumentDatabaseManager.search_similar()`: Vector similarity search
84
+
85
+ ## Testing
86
+
87
+ Install dependencies (preferably in a virtualenv) before running tests:
88
+ ```bash
89
+ pip install -e .[test]
90
+ ```
91
+
92
+ ### Unit Tests
93
+ ```bash
94
+ python -m unittest
95
+ ```
96
+
97
+ ### Integration Tests
98
+
99
+ Integration tests require a PostgreSQL database with PGVector extension. Set up your test database and configure the connection in `integ-tests/.env`:
100
+
101
+ ```bash
102
+ python -m unittest discover -s integ-tests
103
+ ```
104
+
105
+ ## Contributing
106
+
107
+ 1. Fork the repository
108
+ 2. Create a feature branch
109
+ 3. Make your changes with tests
110
+ 4. Run the test suite
111
+ 5. Submit a pull request
112
+
113
+ ### Development Setup
114
+
115
+ ```bash
116
+ pip install -e .[dev,test]
117
+ black . # Format code
118
+ ```
119
+
120
+ ## License
121
+
122
+ MIT License - see [LICENSE](LICENSE) file for details.
123
+
124
+ ## Links
125
+
126
+ - [GitHub Repository](https://github.com/DavidLiuGit/PGVector-Template)
127
+ - [PGVector Documentation](https://github.com/pgvector/pgvector)
@@ -190,6 +190,8 @@ class BaseDocumentMetadata(BaseModel):
190
190
  Base metadata structure.
191
191
  It is generally expected that every `BaseDocument`'s metadata follows this exact schema,
192
192
  without any extraneous properties, or any missing properties, to avoid ambiguity.
193
+ It is mandatory to include a description with every field/property.
194
+ May contain nested Pydantic models, but they may not be optional. Instead, set defaults.
193
195
  """
194
196
 
195
197
  document_type: str = Field(
@@ -1,60 +1,23 @@
1
- from dataclasses import dataclass, asdict
2
- from datetime import datetime
3
1
  from logging import getLogger
4
2
  from typing import Any, Type, Sequence
5
3
 
6
- from pydantic import BaseModel, ConfigDict, Field, model_validator
7
- from sqlalchemy import select, or_
8
- from sqlalchemy.sql import Select
4
+ from pydantic import BaseModel, Field
5
+ from sqlalchemy import select, or_, Integer, Float
6
+ from sqlalchemy.orm import Session
7
+ from sqlalchemy.sql import Select, ColumnElement
9
8
 
10
9
  from pgvector_template.core import (
11
10
  BaseEmbeddingProvider,
12
11
  BaseDocument,
13
12
  BaseDocumentMetadata,
14
13
  )
15
- from sqlalchemy.orm import Session
14
+ from pgvector_template.models.search import SearchQuery, MetadataFilter, RetrievalResult
15
+ from pgvector_template.utils.metadata_filter import validate_metadata_filters
16
16
 
17
17
 
18
18
  logger = getLogger(__name__)
19
19
 
20
20
 
21
- class SearchQuery(BaseModel):
22
- """Standardized search query structure. At least 1 search criterion is required."""
23
-
24
- text: str | None = None
25
- """String to match against using in a semantic search, i.e. using vector distance."""
26
- keywords: list[str] | None = None
27
- """List of keywords to exact-match in a keyword search."""
28
- metadata_filters: dict[str, Any] | None = None
29
- """Strict metadata filters that must be matched."""
30
- date_range: tuple[datetime, datetime] | None = None
31
- """Retrieve/limit results based on created_at & updated_at timestamps"""
32
- limit: int = Field(
33
- ...,
34
- ge=1,
35
- )
36
- """Maximum number of results to return."""
37
-
38
- model_config = ConfigDict(use_attribute_docstrings=True)
39
-
40
- @model_validator(mode="after")
41
- def ensure_criterion(self):
42
- if not any([self.text, self.keywords, self.metadata_filters, self.date_range]):
43
- raise ValueError("At least one search criterion is required")
44
- return self
45
-
46
-
47
- @dataclass
48
- class RetrievalResult:
49
- """Standardized result structure for all retrieval operations"""
50
-
51
- document: BaseDocument
52
- score: float
53
-
54
- def to_dict(self) -> dict[str, Any]:
55
- return asdict(self)
56
-
57
-
58
21
  class BaseSearchClientConfig(BaseModel):
59
22
  """Config obj for `BaseSearchClient`."""
60
23
 
@@ -113,15 +76,15 @@ class BaseSearchClient:
113
76
  if query.text:
114
77
  db_query = self._apply_semantic_search(db_query, query)
115
78
  db_query = self._apply_keyword_search(db_query, query)
116
- # if query.metadata_filters:
117
- # db_query = self._apply_metadata_filters(db_query, query)
79
+ if query.metadata_filters:
80
+ db_query = self._apply_metadata_filters(db_query, query)
118
81
  db_query = db_query.limit(query.limit)
119
82
 
120
83
  # execute query and return results
121
84
  results = self.session.scalars(db_query).all()
122
85
  return self._convert_to_retrieval_results(results)
123
86
 
124
- def _apply_semantic_search(self, query: Select, search_query: SearchQuery) -> Select:
87
+ def _apply_semantic_search(self, db_query: Select, search_query: SearchQuery) -> Select:
125
88
  """Apply semantic (vector) search criteria to the query.
126
89
  `embedding_provider` must be provided at instantiation, or an `ValueError` will be raised.
127
90
  In PGVector, `<=>` operator is used to compare cosine distance. Lower = more similar.
@@ -134,9 +97,11 @@ class BaseSearchClient:
134
97
  Updated SQLAlchemy query with semantic search applied.
135
98
  """
136
99
  if not search_query.text:
137
- return query
100
+ return db_query
138
101
  query_embedding = self.embedding_provider.embed_text(search_query.text)
139
- return query.order_by(self.config.document_cls.embedding.cosine_distance(query_embedding))
102
+ return db_query.order_by(
103
+ self.config.document_cls.embedding.cosine_distance(query_embedding)
104
+ )
140
105
 
141
106
  def _apply_keyword_search(self, db_query: Select, search_query: SearchQuery) -> Select:
142
107
  """Apply keyword (full-text) search criteria to the query.
@@ -155,17 +120,83 @@ class BaseSearchClient:
155
120
  conditions.append(self.config.document_cls.content.ilike(f"%{keyword}%"))
156
121
  return db_query.where(or_(*conditions))
157
122
 
158
- def _apply_metadata_filters(self, query: Select, search_query: SearchQuery) -> Select:
159
- """Apply metadata filters to the query.
123
+ def _apply_metadata_filters(self, db_query: Select, search_query: SearchQuery) -> Select:
124
+ """
125
+ Apply metadata filters to the query. All condtions in `search_query.metadata_filters`,
126
+ if not None, are ANDed together. Metadata filters are applied against
127
+ `BaseDocument.document_metadata` JSONB field.
160
128
 
161
129
  Args:
162
130
  query: The base SQLAlchemy query.
163
- search_query: The search query containing metadata filters.
131
+ search_query: The `SearchQuery` containing metadata filters.
164
132
 
165
133
  Returns:
166
134
  Updated SQLAlchemy query with metadata filters applied.
167
135
  """
168
- raise NotImplementedError
136
+ if not search_query.metadata_filters:
137
+ return db_query
138
+
139
+ # strictly speaking, this validation step is optional. You may override this method to disable
140
+ try:
141
+ validate_metadata_filters(
142
+ search_query.metadata_filters, self.config.document_metadata_cls
143
+ )
144
+ except ValueError as e:
145
+ logger.warning(f"Metadata filter validation failed: {e}. Query success not guaranteed!")
146
+
147
+ conditions = [
148
+ self._build_metadata_filter_where_condition(filter_obj)
149
+ for filter_obj in search_query.metadata_filters
150
+ ]
151
+ return db_query.where(*conditions) if conditions else db_query
152
+
153
+ def _build_metadata_filter_where_condition(
154
+ self, filter_obj: MetadataFilter
155
+ ) -> ColumnElement[bool]:
156
+ """Build SQLAlchemy WHERE condition for a metadata filter."""
157
+ field_path = filter_obj.field_name.split(".")
158
+ metadata_col = self.config.document_cls.document_metadata
159
+
160
+ # Navigate to field
161
+ field_ref = metadata_col
162
+ for part in field_path:
163
+ field_ref = field_ref[part]
164
+
165
+ if filter_obj.condition == "eq":
166
+ return (
167
+ field_ref.astext == filter_obj.value
168
+ if isinstance(filter_obj.value, str)
169
+ else field_ref == filter_obj.value
170
+ )
171
+ elif filter_obj.condition in {"gt", "gte", "lt", "lte"}:
172
+ if isinstance(filter_obj.value, str):
173
+ field_text = field_ref.astext
174
+ else:
175
+ cast_type = Integer if isinstance(filter_obj.value, int) else Float
176
+ field_text = field_ref.astext.cast(cast_type)
177
+
178
+ if filter_obj.condition == "gt":
179
+ return field_text > filter_obj.value
180
+ elif filter_obj.condition == "gte":
181
+ return field_text >= filter_obj.value
182
+ elif filter_obj.condition == "lt":
183
+ return field_text < filter_obj.value
184
+ else: # lte
185
+ return field_text <= filter_obj.value
186
+ elif filter_obj.condition == "contains":
187
+ return field_ref.contains([filter_obj.value])
188
+ elif filter_obj.condition == "in":
189
+ return field_ref.astext.in_([str(v) for v in filter_obj.value])
190
+ elif filter_obj.condition == "exists":
191
+ if len(field_path) == 1:
192
+ return metadata_col.has_key(field_path[0])
193
+ else:
194
+ parent_ref = metadata_col
195
+ for part in field_path[:-1]:
196
+ parent_ref = parent_ref[part]
197
+ return parent_ref.has_key(field_path[-1])
198
+ else:
199
+ raise ValueError(f"Unsupported condition: {filter_obj.condition}")
169
200
 
170
201
  def _convert_to_retrieval_results(self, results: Sequence[Any]) -> list[RetrievalResult]:
171
202
  """Convert database results to RetrievalResult objects.
@@ -179,6 +210,5 @@ class BaseSearchClient:
179
210
  """
180
211
  retrieval_results = []
181
212
  for result in results:
182
- doc = result[0] if isinstance(result, tuple) else result
183
213
  retrieval_results.append(RetrievalResult(document=result, score=1.0))
184
214
  return retrieval_results
@@ -0,0 +1,76 @@
1
+ from dataclasses import dataclass, asdict
2
+ from datetime import datetime
3
+ from typing import Any, Literal, Type
4
+
5
+ from pydantic import BaseModel, ConfigDict, Field, model_validator
6
+
7
+ from pgvector_template.core import BaseDocument, BaseDocumentMetadata
8
+
9
+
10
+ class MetadataFilter(BaseModel):
11
+ """
12
+ An object acting as a filter for an arbitrary `Metadata` dictionary/map/object.
13
+ """
14
+
15
+ field_name: str
16
+ """Field path in metadata. Use dot notation for nested fields (e.g., 'publication_info.journal')"""
17
+ condition: Literal["eq", "gt", "gte", "lt", "lte", "contains", "in", "exists"]
18
+ """
19
+ Comparison operator:
20
+ - eq=equal
21
+ - gt/gte=greater than/equal
22
+ - lt/lte=less than/equal
23
+ - contains=array contains values (accepts array)
24
+ - in=value in array
25
+ - exists=field exists
26
+ """
27
+ value: Any
28
+ """Value to compare against. Type should match field type (str, int, float, bool, list)"""
29
+
30
+ model_config = ConfigDict(use_attribute_docstrings=True)
31
+
32
+
33
+ class SearchQuery(BaseModel):
34
+ """Standardized search query structure. At least 1 search criterion is required."""
35
+
36
+ text: str | None = None
37
+ """String to match against using in a semantic search, i.e. using vector distance."""
38
+ keywords: list[str] = []
39
+ """List of keywords to **exact-match** in a keyword search."""
40
+ metadata_filters: list[MetadataFilter] = Field(
41
+ default=[],
42
+ json_schema_extra={"metadata_schema": BaseDocumentMetadata.model_json_schema()},
43
+ )
44
+ """
45
+ List of metadata conditions that must be matched.
46
+ Refer to `metadata_schema` for the expected schema, as it exists in the database.
47
+ """
48
+ date_range: tuple[datetime, datetime] | None = None
49
+ """Retrieve/limit results based on created_at & updated_at timestamps (i.e. database operations)"""
50
+ limit: int = Field(
51
+ ...,
52
+ ge=1,
53
+ )
54
+ """Maximum number of results to return."""
55
+
56
+ model_config = ConfigDict(
57
+ use_attribute_docstrings=True,
58
+ arbitrary_types_allowed=True,
59
+ )
60
+
61
+ @model_validator(mode="after")
62
+ def ensure_criterion(self):
63
+ if not any([self.text, self.keywords, self.metadata_filters, self.date_range]):
64
+ raise ValueError("At least one search criterion is required")
65
+ return self
66
+
67
+
68
+ @dataclass
69
+ class RetrievalResult:
70
+ """Standardized result structure for all retrieval operations"""
71
+
72
+ document: BaseDocument
73
+ score: float
74
+
75
+ def to_dict(self) -> dict[str, Any]:
76
+ return asdict(self)
@@ -0,0 +1,81 @@
1
+ from typing import Type
2
+
3
+ from pgvector_template.core.document import BaseDocumentMetadata
4
+ from pgvector_template.models.search import MetadataFilter
5
+
6
+
7
+ def validate_metadata_filters(
8
+ filter_obj_list: list[MetadataFilter], metadata_cls: Type[BaseDocumentMetadata]
9
+ ) -> None:
10
+ """Validate a `list[MetadataFilter]` against schema and condition compatibility.
11
+
12
+ Args:
13
+ filter_obj_list (list[MetadataFilter]): _description_
14
+ metadata_cls (Type[BaseDocumentMetadata]): _description_
15
+ """
16
+ for metadata_filter_obj in filter_obj_list:
17
+ validate_metadata_filter(metadata_filter_obj, metadata_cls)
18
+
19
+
20
+ def validate_metadata_filter(
21
+ filter_obj: MetadataFilter, metadata_cls: Type[BaseDocumentMetadata]
22
+ ) -> None:
23
+ """Validate metadata filter against schema and condition compatibility.
24
+
25
+ Note: This validates the schema structure but cannot guarantee runtime data conformity.
26
+ JSONB field access like field_ref[part] will succeed even if the actual data doesn't match the schema.
27
+
28
+ Args:
29
+ filter_obj: The metadata filter to validate
30
+ metadata_cls: The metadata class to validate against
31
+
32
+ Raises:
33
+ ValueError: If field doesn't exist in schema or condition is incompatible with field type
34
+ """
35
+ field_path = filter_obj.field_name.split(".")
36
+ current_field_info = metadata_cls.model_fields
37
+
38
+ # Navigate nested structure
39
+ for i, part in enumerate(field_path):
40
+ if part not in current_field_info:
41
+ raise ValueError(f"Field '{filter_obj.field_name}' not found in metadata schema")
42
+
43
+ field_info = current_field_info[part]
44
+ field_type = field_info.annotation
45
+
46
+ # Handle nested models
47
+ if not field_type:
48
+ raise ValueError(f"Field '{filter_obj.field_name}' not found in metadata schema")
49
+ elif hasattr(field_type, "model_fields"):
50
+ current_field_info = field_type.model_fields
51
+ elif i < len(field_path) - 1:
52
+ raise ValueError(
53
+ f"Cannot navigate into non-model field '{part}' in path '{filter_obj.field_name}'"
54
+ )
55
+
56
+ # Validate condition compatibility with final field type
57
+ validate_condition_compatibility(field_type, filter_obj.condition)
58
+
59
+
60
+ def validate_condition_compatibility(field_type: Type, condition: str) -> None:
61
+ """Validate that condition is compatible with field type."""
62
+ # Extract base type from Optional/Union types
63
+ origin = getattr(field_type, "__origin__", None)
64
+ if origin is not None:
65
+ args = getattr(field_type, "__args__", ())
66
+ if origin is list:
67
+ field_type = list
68
+ elif len(args) > 0:
69
+ field_type = args[0] # First non-None type
70
+
71
+ valid_conditions = {
72
+ str: {"eq", "gt", "gte", "lt", "lte", "in", "exists"},
73
+ int: {"eq", "gt", "gte", "lt", "lte", "exists"},
74
+ float: {"eq", "gt", "gte", "lt", "lte", "exists"},
75
+ bool: {"eq", "exists"},
76
+ list: {"contains", "in", "exists"},
77
+ }
78
+
79
+ allowed = valid_conditions.get(field_type, {"eq", "exists"})
80
+ if condition not in allowed:
81
+ raise ValueError(f"Condition '{condition}' not valid for field type {field_type.__name__}")
@@ -0,0 +1,152 @@
1
+ Metadata-Version: 2.1
2
+ Name: pgvector-template
3
+ Version: 0.3.1
4
+ Summary: Template library for flexible PGVector RAG implementations
5
+ Author-email: DL <v49t9zpqd@mozmail.com>
6
+ License: MIT
7
+ Project-URL: Homepage, https://github.com/DavidLiuGit/PGVector-Template
8
+ Classifier: Programming Language :: Python :: 3
9
+ Classifier: License :: OSI Approved :: MIT License
10
+ Classifier: Operating System :: OS Independent
11
+ Requires-Python: >=3.11
12
+ Description-Content-Type: text/markdown
13
+ License-File: LICENSE
14
+ Requires-Dist: pgvector>=0.2.0
15
+ Requires-Dist: pydantic<3.0,>=2.11
16
+ Requires-Dist: sqlalchemy>=2.0.0
17
+ Requires-Dist: typing-extensions>=4.0.0
18
+ Provides-Extra: test
19
+ Requires-Dist: psycopg[binary]>=3.1.0; extra == "test"
20
+ Requires-Dist: pytest>=7.0.0; extra == "test"
21
+ Requires-Dist: pytest-cov>=4.0.0; extra == "test"
22
+ Requires-Dist: python-dotenv>=1.0.0; extra == "test"
23
+ Provides-Extra: dev
24
+ Requires-Dist: black>=23.0.0; extra == "dev"
25
+
26
+ # PGVector-Template
27
+
28
+ A flexible, production-ready template library for building Retrieval-Augmented Generation (RAG) applications using PostgreSQL with PGVector extensions.
29
+
30
+ ## Overview
31
+
32
+ PGVector-Template provides a robust foundation for implementing vector-based document storage and retrieval systems. It offers a clean abstraction layer over PostgreSQL's PGVector extension, making it easy to build scalable RAG applications with proper document management, metadata handling, and efficient vector search capabilities.
33
+
34
+ ## Key Features
35
+
36
+ - **Flexible Document Model**: Abstract base classes for customizable document schemas
37
+ - **Vector Search**: Optimized HNSW indexing for fast similarity search
38
+ - **Metadata Management**: JSON-based flexible metadata with GIN indexing
39
+ - **Collection Support**: Organize documents into logical collections
40
+ - **Chunk Management**: Handle long content (refer to as **corpus**) by chunking it into smaller **documents**. Handle recovering the original corpus given its id
41
+ - **Database Abstraction**: Clean SQLAlchemy-based database layer, with an API to create schemas
42
+ - **Type Safety**: Full Pydantic validation and type hints
43
+ - **Production Ready**: Comprehensive testing and error handling
44
+
45
+ ## Architecture
46
+
47
+ The library is organized into several key components:
48
+
49
+ - **Core**: Document models, embedders, search functionality
50
+ - **Database**: Connection management and document database operations
51
+ - **Service**: High-level document service layer
52
+ - **Types**: Shared type definitions and schemas
53
+
54
+ ## Installation
55
+
56
+ ```bash
57
+ pip install pgvector-template
58
+ ```
59
+
60
+ Or add `pgvector-template` to your dependencies
61
+
62
+ ### Prerequisites
63
+
64
+ - Python 3.11+
65
+ - To execute tests: PostgreSQL with PGVector extension
66
+
67
+ ## Configuration
68
+
69
+ ### Database Setup
70
+
71
+ 1. Install PostgreSQL with PGVector extension
72
+ 2. Create your database and enable the vector extension:
73
+
74
+ ```sql
75
+ CREATE EXTENSION IF NOT EXISTS vector;
76
+ ```
77
+
78
+ 3. Set up your connection string in environment variables or pass directly to `DatabaseManager`
79
+
80
+ ### Environment Variables
81
+
82
+ For integration tests, create a `.env` file
83
+
84
+ ```bash
85
+ cp integ-tests/.env.example integ-tests/.env
86
+ ```
87
+
88
+ Specify envvars directly in the .env file. It is loaded automatically for integ tests.
89
+
90
+ ```bash
91
+ DATABASE_URL=postgresql://user:password@localhost:5432/test_db
92
+ ```
93
+
94
+ ## API Reference
95
+
96
+ ### Core Classes
97
+
98
+ - `BaseDocument`: Abstract document model with vector embedding support
99
+ - refer to table schema for explanation of the fields
100
+ - `BaseDocumentOptionalProps`: Optional properties for document creation
101
+ - `DatabaseManager`: Database connection and session management
102
+ - `DocumentDatabaseManager`: High-level document operations
103
+
104
+ ### Key Methods
105
+
106
+ - `BaseDocument.from_props()`: Create document instances from properties
107
+ - `DocumentDatabaseManager.insert_document()`: Store documents
108
+ - `DocumentDatabaseManager.search_similar()`: Vector similarity search
109
+
110
+ ## Testing
111
+
112
+ Install dependencies (preferably in a virtualenv) before running tests:
113
+ ```bash
114
+ pip install -e .[test]
115
+ ```
116
+
117
+ ### Unit Tests
118
+ ```bash
119
+ python -m unittest
120
+ ```
121
+
122
+ ### Integration Tests
123
+
124
+ Integration tests require a PostgreSQL database with PGVector extension. Set up your test database and configure the connection in `integ-tests/.env`:
125
+
126
+ ```bash
127
+ python -m unittest discover -s integ-tests
128
+ ```
129
+
130
+ ## Contributing
131
+
132
+ 1. Fork the repository
133
+ 2. Create a feature branch
134
+ 3. Make your changes with tests
135
+ 4. Run the test suite
136
+ 5. Submit a pull request
137
+
138
+ ### Development Setup
139
+
140
+ ```bash
141
+ pip install -e .[dev,test]
142
+ black . # Format code
143
+ ```
144
+
145
+ ## License
146
+
147
+ MIT License - see [LICENSE](LICENSE) file for details.
148
+
149
+ ## Links
150
+
151
+ - [GitHub Repository](https://github.com/DavidLiuGit/PGVector-Template)
152
+ - [PGVector Documentation](https://github.com/pgvector/pgvector)
@@ -16,6 +16,9 @@ pgvector_template/core/search.py
16
16
  pgvector_template/db/__init__.py
17
17
  pgvector_template/db/connection.py
18
18
  pgvector_template/db/document_db.py
19
+ pgvector_template/models/__init__.py
20
+ pgvector_template/models/search.py
19
21
  pgvector_template/service/__init__.py
20
22
  pgvector_template/service/document_service.py
21
- pgvector_template/utils/__init__.py
23
+ pgvector_template/utils/__init__.py
24
+ pgvector_template/utils/metadata_filter.py
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
4
4
 
5
5
  [project]
6
6
  name = "pgvector-template"
7
- version = "0.2.3"
7
+ version = "0.3.1"
8
8
  description = "Template library for flexible PGVector RAG implementations"
9
9
  authors = [{ name="DL", email="v49t9zpqd@mozmail.com" }]
10
10
  license = { text = "MIT" }
@@ -1,46 +0,0 @@
1
- Metadata-Version: 2.1
2
- Name: pgvector-template
3
- Version: 0.2.3
4
- Summary: Template library for flexible PGVector RAG implementations
5
- Author-email: DL <v49t9zpqd@mozmail.com>
6
- License: MIT
7
- Project-URL: Homepage, https://github.com/DavidLiuGit/PGVector-Template
8
- Classifier: Programming Language :: Python :: 3
9
- Classifier: License :: OSI Approved :: MIT License
10
- Classifier: Operating System :: OS Independent
11
- Requires-Python: >=3.11
12
- Description-Content-Type: text/markdown
13
- License-File: LICENSE
14
- Requires-Dist: pgvector>=0.2.0
15
- Requires-Dist: pydantic<3.0,>=2.11
16
- Requires-Dist: sqlalchemy>=2.0.0
17
- Requires-Dist: typing-extensions>=4.0.0
18
- Provides-Extra: test
19
- Requires-Dist: psycopg[binary]>=3.1.0; extra == "test"
20
- Requires-Dist: pytest>=7.0.0; extra == "test"
21
- Requires-Dist: pytest-cov>=4.0.0; extra == "test"
22
- Requires-Dist: python-dotenv>=1.0.0; extra == "test"
23
- Provides-Extra: dev
24
- Requires-Dist: black>=23.0.0; extra == "dev"
25
-
26
- # PGVector-Template
27
-
28
- Template library for flexible PGVector RAG implementations
29
-
30
-
31
- ## Testing
32
-
33
- Install dependencies (preferably in a virtualenv) before running tests:
34
- ```bash
35
- pip install -e .[test]
36
- ```
37
-
38
- ### Unit tests
39
- ```bash
40
- python -m unittest
41
- ```
42
-
43
- ### Integration tests
44
- ```bash
45
- python -m unittest discover -s integ-tests
46
- ```
@@ -1,21 +0,0 @@
1
- # PGVector-Template
2
-
3
- Template library for flexible PGVector RAG implementations
4
-
5
-
6
- ## Testing
7
-
8
- Install dependencies (preferably in a virtualenv) before running tests:
9
- ```bash
10
- pip install -e .[test]
11
- ```
12
-
13
- ### Unit tests
14
- ```bash
15
- python -m unittest
16
- ```
17
-
18
- ### Integration tests
19
- ```bash
20
- python -m unittest discover -s integ-tests
21
- ```
@@ -1,46 +0,0 @@
1
- Metadata-Version: 2.1
2
- Name: pgvector-template
3
- Version: 0.2.3
4
- Summary: Template library for flexible PGVector RAG implementations
5
- Author-email: DL <v49t9zpqd@mozmail.com>
6
- License: MIT
7
- Project-URL: Homepage, https://github.com/DavidLiuGit/PGVector-Template
8
- Classifier: Programming Language :: Python :: 3
9
- Classifier: License :: OSI Approved :: MIT License
10
- Classifier: Operating System :: OS Independent
11
- Requires-Python: >=3.11
12
- Description-Content-Type: text/markdown
13
- License-File: LICENSE
14
- Requires-Dist: pgvector>=0.2.0
15
- Requires-Dist: pydantic<3.0,>=2.11
16
- Requires-Dist: sqlalchemy>=2.0.0
17
- Requires-Dist: typing-extensions>=4.0.0
18
- Provides-Extra: test
19
- Requires-Dist: psycopg[binary]>=3.1.0; extra == "test"
20
- Requires-Dist: pytest>=7.0.0; extra == "test"
21
- Requires-Dist: pytest-cov>=4.0.0; extra == "test"
22
- Requires-Dist: python-dotenv>=1.0.0; extra == "test"
23
- Provides-Extra: dev
24
- Requires-Dist: black>=23.0.0; extra == "dev"
25
-
26
- # PGVector-Template
27
-
28
- Template library for flexible PGVector RAG implementations
29
-
30
-
31
- ## Testing
32
-
33
- Install dependencies (preferably in a virtualenv) before running tests:
34
- ```bash
35
- pip install -e .[test]
36
- ```
37
-
38
- ### Unit tests
39
- ```bash
40
- python -m unittest
41
- ```
42
-
43
- ### Integration tests
44
- ```bash
45
- python -m unittest discover -s integ-tests
46
- ```