@memberjunction/ai-vector-sync 2.43.0 → 2.45.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +241 -47
- package/package.json +12 -12
package/README.md
CHANGED
|
@@ -1,68 +1,262 @@
|
|
|
1
|
-
#
|
|
1
|
+
# @memberjunction/ai-vector-sync
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
A robust MemberJunction package for synchronizing entities with vector databases by transforming entity records into vector representations using embedding models.
|
|
4
4
|
|
|
5
|
-
|
|
5
|
+
## Overview
|
|
6
6
|
|
|
7
|
-
|
|
7
|
+
The `@memberjunction/ai-vector-sync` package provides a comprehensive solution for:
|
|
8
|
+
- Converting MemberJunction entities into vector embeddings
|
|
9
|
+
- Storing embeddings in vector databases (currently supports Pinecone)
|
|
10
|
+
- Managing the synchronization lifecycle between entities and their vector representations
|
|
11
|
+
- Supporting batch processing for large datasets
|
|
12
|
+
- Providing template-based document generation for vectorization
|
|
8
13
|
|
|
9
|
-
|
|
10
|
-
A SQL database must be configured with the Memberjunction framework installed. This serves as the foundation for managing entity records and templates needed for the vectorization process.
|
|
14
|
+
## Installation
|
|
11
15
|
|
|
12
|
-
|
|
13
|
-
|
|
16
|
+
```bash
|
|
17
|
+
npm install @memberjunction/ai-vector-sync
|
|
18
|
+
```
|
|
14
19
|
|
|
15
|
-
|
|
16
|
-
Acquire an API key for the vector database you plan to use. Currently, Memberjunction supports Pinecone as the vector storage solution, which integrates seamlessly with this package.
|
|
20
|
+
## Prerequisites
|
|
17
21
|
|
|
18
|
-
|
|
19
|
-
Ensure you have an Entity document record defined, along with an associated template. The template specifies the properties of the entity that will be included in the vectorization process, guiding the transformation of data into its vector representation.
|
|
22
|
+
Before using this package, ensure you have:
|
|
20
23
|
|
|
21
|
-
|
|
24
|
+
1. **SQL Database with MemberJunction Framework**
|
|
25
|
+
A properly configured SQL database with the MemberJunction framework installed.
|
|
22
26
|
|
|
23
|
-
|
|
27
|
+
2. **API Keys**
|
|
28
|
+
- Embedding model API key (supports OpenAI, Mistral, etc.)
|
|
29
|
+
- Vector database API key (currently supports Pinecone)
|
|
24
30
|
|
|
25
|
-
|
|
26
|
-
-
|
|
27
|
-
-
|
|
31
|
+
3. **Entity Configuration**
|
|
32
|
+
- Entity Document record defined in MemberJunction
|
|
33
|
+
- Associated template for specifying which entity properties to vectorize
|
|
28
34
|
|
|
29
|
-
|
|
30
|
-
- Once the Entity Document record is obtained, the system identifies the appropriate **embedding model**. This model is critical as it determines how the entity data will be transformed into vector representations.
|
|
31
|
-
- Additionally, the system fetches details regarding the **vector database** where the resulting vectors will be stored. This includes information about the specific database or configuration settings necessary for data insertion.
|
|
35
|
+
## Core Features
|
|
32
36
|
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
- The retrieved records are then vectorized using the selected embedding model. This step involves transforming each record into a high-dimensional vector, which captures the nuances and characteristics of the data.
|
|
37
|
+
### Entity Vectorization
|
|
38
|
+
Transform entity records into high-dimensional vectors that capture the semantic meaning of the data.
|
|
36
39
|
|
|
37
|
-
|
|
38
|
-
|
|
40
|
+
### Batch Processing
|
|
41
|
+
Efficiently handle large datasets with configurable batch sizes for:
|
|
42
|
+
- Record fetching
|
|
43
|
+
- Vectorization
|
|
44
|
+
- Database upsertion
|
|
39
45
|
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
- The EntityRecordDocument ensures traceability and provides a means to reference back to the original data, enabling users to maintain a clear mapping between vectors and their source records.
|
|
43
|
-
-
|
|
44
|
-
# How to Vectorize Records
|
|
46
|
+
### Template-Based Processing
|
|
47
|
+
Use MemberJunction templates to define which entity fields and relationships to include in vectorization.
|
|
45
48
|
|
|
46
|
-
|
|
49
|
+
### Vector Database Integration
|
|
50
|
+
Seamlessly integrate with vector databases through the MemberJunction AI infrastructure.
|
|
47
51
|
|
|
48
|
-
|
|
49
|
-
Add the `Entity Vectorizer` package to your existing project. Alternatively, you can use the `test-vectorization` package available in the Memberjunction repository:
|
|
50
|
-
[Memberjunction Test Vectorization](https://github.com/MemberJunction/MJ/tree/next/test-vectorization)
|
|
52
|
+
## Usage
|
|
51
53
|
|
|
52
|
-
|
|
53
|
-
Ensure that the packages for the embedding model and vector database you plan to use are added to your existing project. These packages must not be tree-shaken out during the build process, as they are required for vectorization to work properly.
|
|
54
|
+
### Basic Entity Vectorization
|
|
54
55
|
|
|
55
|
-
|
|
56
|
-
|
|
56
|
+
```typescript
|
|
57
|
+
import { EntityVectorSyncer } from '@memberjunction/ai-vector-sync';
|
|
58
|
+
import { UserInfo } from '@memberjunction/core';
|
|
57
59
|
|
|
58
|
-
|
|
59
|
-
|
|
60
|
+
// Initialize the syncer
|
|
61
|
+
const syncer = new EntityVectorSyncer();
|
|
60
62
|
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
- **`listBatchCount`** *(optional but recommended)*: Determines the size of the batch of records to fetch from the database at a time.
|
|
64
|
-
- **`listID`** *(optional but recommended)*: If specified, only the records within this list will be vectorized. If omitted, all records within the specified entity will be vectorized.
|
|
65
|
-
-
|
|
66
|
-
5. **Note on Long Running Processes**
|
|
67
|
-
The process of vectorizing records may take a significant amount of time (1+ hour(s)) depending on the size of the dataset and the performance of the configured systems. Although the `VectorizeEntity` function returns a promise, it is often best not to `await` it. Instead, let it run in the background to avoid blocking your application’s main workflow.
|
|
63
|
+
// Configure the syncer (required before first use)
|
|
64
|
+
await syncer.Config(false, contextUser);
|
|
68
65
|
|
|
66
|
+
// Vectorize an entity
|
|
67
|
+
const params = {
|
|
68
|
+
entityID: 'your-entity-id',
|
|
69
|
+
entityDocumentID: 'your-entity-document-id',
|
|
70
|
+
listBatchCount: 50, // Optional: records per batch (default: 50)
|
|
71
|
+
VectorizeBatchCount: 50, // Optional: vectorization batch size (default: 50)
|
|
72
|
+
UpsertBatchCount: 50, // Optional: upsert batch size (default: 50)
|
|
73
|
+
StartingOffset: 0 // Optional: skip records for resuming
|
|
74
|
+
};
|
|
75
|
+
|
|
76
|
+
// Start vectorization (runs asynchronously)
|
|
77
|
+
syncer.VectorizeEntity(params, contextUser);
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
### Vectorizing a Specific List
|
|
81
|
+
|
|
82
|
+
```typescript
|
|
83
|
+
// Vectorize only records within a specific list
|
|
84
|
+
const params = {
|
|
85
|
+
entityID: 'your-entity-id',
|
|
86
|
+
entityDocumentID: 'your-entity-document-id',
|
|
87
|
+
listID: 'your-list-id', // Only vectorize records in this list
|
|
88
|
+
listBatchCount: 100
|
|
89
|
+
};
|
|
90
|
+
|
|
91
|
+
await syncer.VectorizeEntity(params, contextUser);
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
### Working with Entity Documents
|
|
95
|
+
|
|
96
|
+
```typescript
|
|
97
|
+
// Get entity document by ID
|
|
98
|
+
const entityDoc = await syncer.GetEntityDocument('document-id');
|
|
99
|
+
|
|
100
|
+
// Get entity document by name
|
|
101
|
+
const entityDoc = await syncer.GetEntityDocumentByName('Document Name', contextUser);
|
|
102
|
+
|
|
103
|
+
// Get all active entity documents
|
|
104
|
+
const activeDocs = await syncer.GetActiveEntityDocuments();
|
|
105
|
+
|
|
106
|
+
// Get active documents for specific entities
|
|
107
|
+
const specificDocs = await syncer.GetActiveEntityDocuments(['Entity1', 'Entity2']);
|
|
108
|
+
```
|
|
109
|
+
|
|
110
|
+
### Creating Default Entity Documents
|
|
111
|
+
|
|
112
|
+
```typescript
|
|
113
|
+
import { VectorDatabaseEntity, AIModelEntity } from '@memberjunction/core-entities';
|
|
114
|
+
|
|
115
|
+
// Create a default entity document when one doesn't exist
|
|
116
|
+
const entityDoc = await syncer.CreateDefaultEntityDocument(
|
|
117
|
+
entityID,
|
|
118
|
+
vectorDatabase, // VectorDatabaseEntity instance
|
|
119
|
+
aiModel // AIModelEntity instance
|
|
120
|
+
);
|
|
121
|
+
```
|
|
122
|
+
|
|
123
|
+
## API Reference
|
|
124
|
+
|
|
125
|
+
### EntityVectorSyncer
|
|
126
|
+
|
|
127
|
+
The main class for entity vectorization operations.
|
|
128
|
+
|
|
129
|
+
#### Methods
|
|
130
|
+
|
|
131
|
+
##### `Config(forceRefresh: boolean, contextUser?: UserInfo): Promise<void>`
|
|
132
|
+
Configures the syncer and initializes required engines.
|
|
133
|
+
- `forceRefresh`: Force refresh of caches and engines
|
|
134
|
+
- `contextUser`: User context for operations
|
|
135
|
+
|
|
136
|
+
##### `VectorizeEntity(params: VectorizeEntityParams, contextUser?: UserInfo): Promise<VectorizeEntityResponse>`
|
|
137
|
+
Vectorizes entities based on provided parameters.
|
|
138
|
+
- `params`: Configuration for vectorization
|
|
139
|
+
- `contextUser`: Required user context
|
|
140
|
+
|
|
141
|
+
##### `GetEntityDocument(entityDocumentID: string): Promise<EntityDocumentEntity | null>`
|
|
142
|
+
Retrieves an entity document by ID.
|
|
143
|
+
|
|
144
|
+
##### `GetEntityDocumentByName(entityDocumentName: string, contextUser?: UserInfo): Promise<EntityDocumentEntity | null>`
|
|
145
|
+
Retrieves an entity document by name.
|
|
146
|
+
|
|
147
|
+
##### `GetActiveEntityDocuments(entityNames?: string[]): Promise<EntityDocumentEntity[]>`
|
|
148
|
+
Gets all active entity documents, optionally filtered by entity names.
|
|
149
|
+
|
|
150
|
+
##### `CreateDefaultEntityDocument(entityID: string, vectorDatabase: VectorDatabaseEntity, aiModel: AIModelEntity): Promise<EntityDocumentEntity>`
|
|
151
|
+
Creates a default entity document for the specified entity.
|
|
152
|
+
|
|
153
|
+
### Types
|
|
154
|
+
|
|
155
|
+
#### VectorizeEntityParams
|
|
156
|
+
```typescript
|
|
157
|
+
type VectorizeEntityParams = {
|
|
158
|
+
entityID: string; // Required: Entity to vectorize
|
|
159
|
+
entityDocumentID?: string; // Entity document configuration
|
|
160
|
+
listID?: string; // Optional: Specific list to vectorize
|
|
161
|
+
listBatchCount?: number; // Records per fetch batch (default: 50)
|
|
162
|
+
VectorizeBatchCount?: number; // Vectorization batch size (default: 50)
|
|
163
|
+
UpsertBatchCount?: number; // Database upsert batch size (default: 50)
|
|
164
|
+
StartingOffset?: number; // Skip records for resuming
|
|
165
|
+
CurrentUser?: UserInfo; // User context
|
|
166
|
+
options?: any; // Additional options
|
|
167
|
+
}
|
|
168
|
+
```
|
|
169
|
+
|
|
170
|
+
#### EntitySyncConfig
|
|
171
|
+
```typescript
|
|
172
|
+
type EntitySyncConfig = {
|
|
173
|
+
EntityDocumentID: string; // Entity document to use
|
|
174
|
+
Interval: number; // Sync interval in seconds
|
|
175
|
+
RunViewParams: RunViewParams; // View parameters for fetching records
|
|
176
|
+
IncludeInSync: boolean; // Include in sync process
|
|
177
|
+
LastRunDate: string; // Last sync timestamp
|
|
178
|
+
VectorIndexID: number; // Vector index ID
|
|
179
|
+
VectorID: number; // Vector database ID
|
|
180
|
+
}
|
|
181
|
+
```
|
|
182
|
+
|
|
183
|
+
## Architecture
|
|
184
|
+
|
|
185
|
+
### Process Flow
|
|
186
|
+
|
|
187
|
+
1. **Entity Document Retrieval**: Fetches configuration from Entity Document record
|
|
188
|
+
2. **Model and Database Configuration**: Sets up embedding model and vector database
|
|
189
|
+
3. **Data Fetching**: Retrieves entity records in batches
|
|
190
|
+
4. **Vectorization**: Transforms records using embedding model
|
|
191
|
+
5. **Vector Upsertion**: Stores vectors in database
|
|
192
|
+
6. **EntityRecordDocument Creation**: Creates tracking records
|
|
193
|
+
|
|
194
|
+
### Worker Architecture
|
|
195
|
+
|
|
196
|
+
The package uses a multi-worker architecture for efficient processing:
|
|
197
|
+
- **VectorizeTemplates Worker**: Handles template-based text generation and embedding
|
|
198
|
+
- **UpsertVectors Worker**: Manages vector database operations
|
|
199
|
+
- **EntityRecordDocument Worker**: Tracks vector-entity relationships
|
|
200
|
+
|
|
201
|
+
## Configuration
|
|
202
|
+
|
|
203
|
+
### Environment Variables
|
|
204
|
+
|
|
205
|
+
Create a `.env` file with:
|
|
206
|
+
|
|
207
|
+
```env
|
|
208
|
+
# Database Configuration
|
|
209
|
+
DB_HOST=your-database-host
|
|
210
|
+
DB_PORT=1433
|
|
211
|
+
DB_USERNAME=your-username
|
|
212
|
+
DB_PASSWORD=your-password
|
|
213
|
+
DB_DATABASE=your-database
|
|
214
|
+
|
|
215
|
+
# API Keys
|
|
216
|
+
OPENAI_API_KEY=your-openai-key
|
|
217
|
+
MISTRAL_API_KEY=your-mistral-key
|
|
218
|
+
PINECONE_API_KEY=your-pinecone-key
|
|
219
|
+
PINECONE_HOST=your-pinecone-host
|
|
220
|
+
PINECONE_DEFAULT_INDEX=your-default-index
|
|
221
|
+
|
|
222
|
+
# User Configuration
|
|
223
|
+
CURRENT_USER_EMAIL=user@example.com
|
|
224
|
+
```
|
|
225
|
+
|
|
226
|
+
## Performance Considerations
|
|
227
|
+
|
|
228
|
+
- **Long-Running Processes**: Vectorization can take hours for large datasets
|
|
229
|
+
- **Batch Sizes**: Adjust batch sizes based on your system resources
|
|
230
|
+
- **Asynchronous Processing**: Consider running vectorization in background processes
|
|
231
|
+
- **Memory Usage**: Monitor memory usage for large batch sizes
|
|
232
|
+
|
|
233
|
+
## Integration with MemberJunction
|
|
234
|
+
|
|
235
|
+
This package integrates seamlessly with:
|
|
236
|
+
- `@memberjunction/core`: Core entity and metadata functionality
|
|
237
|
+
- `@memberjunction/ai`: AI model abstractions
|
|
238
|
+
- `@memberjunction/ai-vectordb`: Vector database abstractions
|
|
239
|
+
- `@memberjunction/templates`: Template processing engine
|
|
240
|
+
|
|
241
|
+
## Error Handling
|
|
242
|
+
|
|
243
|
+
The package includes comprehensive error handling:
|
|
244
|
+
- Validation of entity documents and templates
|
|
245
|
+
- Graceful handling of API failures
|
|
246
|
+
- Detailed logging through MemberJunction's logging system
|
|
247
|
+
|
|
248
|
+
## Best Practices
|
|
249
|
+
|
|
250
|
+
1. **Start with Small Batches**: Test with small batch sizes before processing large datasets
|
|
251
|
+
2. **Monitor Progress**: Use MemberJunction's logging to track vectorization progress
|
|
252
|
+
3. **Handle Interruptions**: Use `StartingOffset` to resume interrupted processes
|
|
253
|
+
4. **Template Design**: Design templates to include relevant fields for semantic search
|
|
254
|
+
5. **Resource Management**: Consider database and API rate limits when setting batch sizes
|
|
255
|
+
|
|
256
|
+
## License
|
|
257
|
+
|
|
258
|
+
ISC - See LICENSE file for details
|
|
259
|
+
|
|
260
|
+
## Author
|
|
261
|
+
|
|
262
|
+
MemberJunction.com
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@memberjunction/ai-vector-sync",
|
|
3
|
-
"version": "2.
|
|
3
|
+
"version": "2.45.0",
|
|
4
4
|
"description": "MemberJunction: AI Vector/Entity Sync Package - handles synchronization between MemberJunction entities and vector databases",
|
|
5
5
|
"main": "./dist/index.js",
|
|
6
6
|
"types": "./dist/index.d.ts",
|
|
@@ -15,17 +15,17 @@
|
|
|
15
15
|
"author": "MemberJunction.com",
|
|
16
16
|
"license": "ISC",
|
|
17
17
|
"dependencies": {
|
|
18
|
-
"@memberjunction/ai": "2.
|
|
19
|
-
"@memberjunction/ai-vectordb": "2.
|
|
20
|
-
"@memberjunction/ai-vectors": "2.
|
|
21
|
-
"@memberjunction/ai-vectors-pinecone": "2.
|
|
22
|
-
"@memberjunction/ai-mistral": "2.
|
|
23
|
-
"@memberjunction/aiengine": "2.
|
|
24
|
-
"@memberjunction/core": "2.
|
|
25
|
-
"@memberjunction/global": "2.
|
|
26
|
-
"@memberjunction/templates": "2.
|
|
27
|
-
"@memberjunction/templates-base-types": "2.
|
|
28
|
-
"@memberjunction/ai-openai": "2.
|
|
18
|
+
"@memberjunction/ai": "2.45.0",
|
|
19
|
+
"@memberjunction/ai-vectordb": "2.45.0",
|
|
20
|
+
"@memberjunction/ai-vectors": "2.45.0",
|
|
21
|
+
"@memberjunction/ai-vectors-pinecone": "2.45.0",
|
|
22
|
+
"@memberjunction/ai-mistral": "2.45.0",
|
|
23
|
+
"@memberjunction/aiengine": "2.45.0",
|
|
24
|
+
"@memberjunction/core": "2.45.0",
|
|
25
|
+
"@memberjunction/global": "2.45.0",
|
|
26
|
+
"@memberjunction/templates": "2.45.0",
|
|
27
|
+
"@memberjunction/templates-base-types": "2.45.0",
|
|
28
|
+
"@memberjunction/ai-openai": "2.45.0",
|
|
29
29
|
"dotenv": "^16.4.1",
|
|
30
30
|
"typeorm": "^0.3.20"
|
|
31
31
|
},
|