@memberjunction/ai-vector-sync 2.43.0 → 2.45.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (2) hide show
  1. package/README.md +241 -47
  2. package/package.json +12 -12
package/README.md CHANGED
@@ -1,68 +1,262 @@
1
- # Entity Vectorizer Overview
1
+ # @memberjunction/ai-vector-sync
2
2
 
3
- The Entity Vectorizer is a robust software tool designed for transforming entities into vector representations and storing them in a vector database. Built upon the powerful [MemberJunction framework](https://docs.memberjunction.org/), the Entity Vectorizer ensures efficient handling and processing of large batches of data. This overview provides insights into its core functionalities, architecture, and usage.
3
+ A robust MemberJunction package for synchronizing entities with vector databases by transforming entity records into vector representations using embedding models.
4
4
 
5
- # Prerequisites
5
+ ## Overview
6
6
 
7
- Before you can vectorize an entity, ensure the following setup is complete:
7
+ The `@memberjunction/ai-vector-sync` package provides a comprehensive solution for:
8
+ - Converting MemberJunction entities into vector embeddings
9
+ - Storing embeddings in vector databases (currently supports Pinecone)
10
+ - Managing the synchronization lifecycle between entities and their vector representations
11
+ - Supporting batch processing for large datasets
12
+ - Providing template-based document generation for vectorization
8
13
 
9
- 1. **SQL Database with Memberjunction Framework**
10
- A SQL database must be configured with the Memberjunction framework installed. This serves as the foundation for managing entity records and templates needed for the vectorization process.
14
+ ## Installation
11
15
 
12
- 2. **API Key for the Embedding Model**
13
- Obtain an API key for the embedding model of your choice. Memberjunction supports a variety of popular embedding models, such as OpenAI and Mistral. This key is necessary for accessing the model during the vectorization process.
16
+ ```bash
17
+ npm install @memberjunction/ai-vector-sync
18
+ ```
14
19
 
15
- 3. **API Key for the Vector Database**
16
- Acquire an API key for the vector database you plan to use. Currently, Memberjunction supports Pinecone as the vector storage solution, which integrates seamlessly with this package.
20
+ ## Prerequisites
17
21
 
18
- 4. **Entity Document Record and Template**
19
- Ensure you have an Entity document record defined, along with an associated template. The template specifies the properties of the entity that will be included in the vectorization process, guiding the transformation of data into its vector representation.
22
+ Before using this package, ensure you have:
20
23
 
21
- ## Process Overview
24
+ 1. **SQL Database with MemberJunction Framework**
25
+ A properly configured SQL database with the MemberJunction framework installed.
22
26
 
23
- The Entity Vectorizer operates via a systematic process to ensure efficient vectorization and storage of entity data. Below is a detailed breakdown of each step involved:
27
+ 2. **API Keys**
28
+ - Embedding model API key (supports OpenAI, Mistral, etc.)
29
+ - Vector database API key (currently supports Pinecone)
24
30
 
25
- 1. **Entity Document Retrieval**:
26
- - The process begins by taking an **Entity Document ID** as input.
27
- - Using this ID, the system retrieves the corresponding **Entity Document record**. This record contains essential information and metadata about the entity, which is crucial for the subsequent steps.
31
+ 3. **Entity Configuration**
32
+ - Entity Document record defined in MemberJunction
33
+ - Associated template for specifying which entity properties to vectorize
28
34
 
29
- 2. **Model and Database Configuration**:
30
- - Once the Entity Document record is obtained, the system identifies the appropriate **embedding model**. This model is critical as it determines how the entity data will be transformed into vector representations.
31
- - Additionally, the system fetches details regarding the **vector database** where the resulting vectors will be stored. This includes information about the specific database or configuration settings necessary for data insertion.
35
+ ## Core Features
32
36
 
33
- 3. **Data Fetching and Vectorization**:
34
- - With the embedding model and database information in place, the system fetches all records under the specified **Entity**. If a **list ID** is provided, it fetches all records associated with that list.
35
- - The retrieved records are then vectorized using the selected embedding model. This step involves transforming each record into a high-dimensional vector, which captures the nuances and characteristics of the data.
37
+ ### Entity Vectorization
38
+ Transform entity records into high-dimensional vectors that capture the semantic meaning of the data.
36
39
 
37
- 4. **Vector Upsertion**:
38
- - The vectorized data is then **upserted** into the vector database.
40
+ ### Batch Processing
41
+ Efficiently handle large datasets with configurable batch sizes for:
42
+ - Record fetching
43
+ - Vectorization
44
+ - Database upsertion
39
45
 
40
- 5. **EntityRecordDocument Creation**:
41
- - For each vector upserted, a corresponding **EntityRecordDocument record** is generated. This record serves as a link between the vector stored in the database and the original record in MemberJunction.
42
- - The EntityRecordDocument ensures traceability and provides a means to reference back to the original data, enabling users to maintain a clear mapping between vectors and their source records.
43
- -
44
- # How to Vectorize Records
46
+ ### Template-Based Processing
47
+ Use MemberJunction templates to define which entity fields and relationships to include in vectorization.
45
48
 
46
- Follow these steps to use the package:
49
+ ### Vector Database Integration
50
+ Seamlessly integrate with vector databases through the MemberJunction AI infrastructure.
47
51
 
48
- 1. **Add the Package**
49
- Add the `Entity Vectorizer` package to your existing project. Alternatively, you can use the `test-vectorization` package available in the Memberjunction repository:
50
- [Memberjunction Test Vectorization](https://github.com/MemberJunction/MJ/tree/next/test-vectorization)
52
+ ## Usage
51
53
 
52
- 2. **Include Embedding Model and Vector Database Packages**
53
- Ensure that the packages for the embedding model and vector database you plan to use are added to your existing project. These packages must not be tree-shaken out during the build process, as they are required for vectorization to work properly.
54
+ ### Basic Entity Vectorization
54
55
 
55
- 3. **Instantiate the `EntityVectorSyncer` Class**
56
- Create an instance of the `EntityVectorSyncer` class and call its `config` method. This ensures all necessary engines (e.g., embedding and vector databases) are set up properly.
56
+ ```typescript
57
+ import { EntityVectorSyncer } from '@memberjunction/ai-vector-sync';
58
+ import { UserInfo } from '@memberjunction/core';
57
59
 
58
- 4. **Call the `VectorizeEntity` Function**
59
- Use the `VectorizeEntity` function, passing in the required parameters. At a minimum, you must provide:
60
+ // Initialize the syncer
61
+ const syncer = new EntityVectorSyncer();
60
62
 
61
- - **`EntityID`**: The name of the entity you want to vectorize.
62
- - **`EntityDocumentID`**: The ID of the Entity Document to use.
63
- - **`listBatchCount`** *(optional but recommended)*: Determines the size of the batch of records to fetch from the database at a time.
64
- - **`listID`** *(optional but recommended)*: If specified, only the records within this list will be vectorized. If omitted, all records within the specified entity will be vectorized.
65
- -
66
- 5. **Note on Long Running Processes**
67
- The process of vectorizing records may take a significant amount of time (1+ hour(s)) depending on the size of the dataset and the performance of the configured systems. Although the `VectorizeEntity` function returns a promise, it is often best not to `await` it. Instead, let it run in the background to avoid blocking your application’s main workflow.
63
+ // Configure the syncer (required before first use)
64
+ await syncer.Config(false, contextUser);
68
65
 
66
+ // Vectorize an entity
67
+ const params = {
68
+ entityID: 'your-entity-id',
69
+ entityDocumentID: 'your-entity-document-id',
70
+ listBatchCount: 50, // Optional: records per batch (default: 50)
71
+ VectorizeBatchCount: 50, // Optional: vectorization batch size (default: 50)
72
+ UpsertBatchCount: 50, // Optional: upsert batch size (default: 50)
73
+ StartingOffset: 0 // Optional: skip records for resuming
74
+ };
75
+
76
+ // Start vectorization (runs asynchronously)
77
+ syncer.VectorizeEntity(params, contextUser);
78
+ ```
79
+
80
+ ### Vectorizing a Specific List
81
+
82
+ ```typescript
83
+ // Vectorize only records within a specific list
84
+ const params = {
85
+ entityID: 'your-entity-id',
86
+ entityDocumentID: 'your-entity-document-id',
87
+ listID: 'your-list-id', // Only vectorize records in this list
88
+ listBatchCount: 100
89
+ };
90
+
91
+ await syncer.VectorizeEntity(params, contextUser);
92
+ ```
93
+
94
+ ### Working with Entity Documents
95
+
96
+ ```typescript
97
+ // Get entity document by ID
98
+ const entityDoc = await syncer.GetEntityDocument('document-id');
99
+
100
+ // Get entity document by name
101
+ const entityDoc = await syncer.GetEntityDocumentByName('Document Name', contextUser);
102
+
103
+ // Get all active entity documents
104
+ const activeDocs = await syncer.GetActiveEntityDocuments();
105
+
106
+ // Get active documents for specific entities
107
+ const specificDocs = await syncer.GetActiveEntityDocuments(['Entity1', 'Entity2']);
108
+ ```
109
+
110
+ ### Creating Default Entity Documents
111
+
112
+ ```typescript
113
+ import { VectorDatabaseEntity, AIModelEntity } from '@memberjunction/core-entities';
114
+
115
+ // Create a default entity document when one doesn't exist
116
+ const entityDoc = await syncer.CreateDefaultEntityDocument(
117
+ entityID,
118
+ vectorDatabase, // VectorDatabaseEntity instance
119
+ aiModel // AIModelEntity instance
120
+ );
121
+ ```
122
+
123
+ ## API Reference
124
+
125
+ ### EntityVectorSyncer
126
+
127
+ The main class for entity vectorization operations.
128
+
129
+ #### Methods
130
+
131
+ ##### `Config(forceRefresh: boolean, contextUser?: UserInfo): Promise<void>`
132
+ Configures the syncer and initializes required engines.
133
+ - `forceRefresh`: Force refresh of caches and engines
134
+ - `contextUser`: User context for operations
135
+
136
+ ##### `VectorizeEntity(params: VectorizeEntityParams, contextUser?: UserInfo): Promise<VectorizeEntityResponse>`
137
+ Vectorizes entities based on provided parameters.
138
+ - `params`: Configuration for vectorization
139
+ - `contextUser`: Required user context
140
+
141
+ ##### `GetEntityDocument(entityDocumentID: string): Promise<EntityDocumentEntity | null>`
142
+ Retrieves an entity document by ID.
143
+
144
+ ##### `GetEntityDocumentByName(entityDocumentName: string, contextUser?: UserInfo): Promise<EntityDocumentEntity | null>`
145
+ Retrieves an entity document by name.
146
+
147
+ ##### `GetActiveEntityDocuments(entityNames?: string[]): Promise<EntityDocumentEntity[]>`
148
+ Gets all active entity documents, optionally filtered by entity names.
149
+
150
+ ##### `CreateDefaultEntityDocument(entityID: string, vectorDatabase: VectorDatabaseEntity, aiModel: AIModelEntity): Promise<EntityDocumentEntity>`
151
+ Creates a default entity document for the specified entity.
152
+
153
+ ### Types
154
+
155
+ #### VectorizeEntityParams
156
+ ```typescript
157
+ type VectorizeEntityParams = {
158
+ entityID: string; // Required: Entity to vectorize
159
+ entityDocumentID?: string; // Entity document configuration
160
+ listID?: string; // Optional: Specific list to vectorize
161
+ listBatchCount?: number; // Records per fetch batch (default: 50)
162
+ VectorizeBatchCount?: number; // Vectorization batch size (default: 50)
163
+ UpsertBatchCount?: number; // Database upsert batch size (default: 50)
164
+ StartingOffset?: number; // Skip records for resuming
165
+ CurrentUser?: UserInfo; // User context
166
+ options?: any; // Additional options
167
+ }
168
+ ```
169
+
170
+ #### EntitySyncConfig
171
+ ```typescript
172
+ type EntitySyncConfig = {
173
+ EntityDocumentID: string; // Entity document to use
174
+ Interval: number; // Sync interval in seconds
175
+ RunViewParams: RunViewParams; // View parameters for fetching records
176
+ IncludeInSync: boolean; // Include in sync process
177
+ LastRunDate: string; // Last sync timestamp
178
+ VectorIndexID: number; // Vector index ID
179
+ VectorID: number; // Vector database ID
180
+ }
181
+ ```
182
+
183
+ ## Architecture
184
+
185
+ ### Process Flow
186
+
187
+ 1. **Entity Document Retrieval**: Fetches configuration from Entity Document record
188
+ 2. **Model and Database Configuration**: Sets up embedding model and vector database
189
+ 3. **Data Fetching**: Retrieves entity records in batches
190
+ 4. **Vectorization**: Transforms records using embedding model
191
+ 5. **Vector Upsertion**: Stores vectors in database
192
+ 6. **EntityRecordDocument Creation**: Creates tracking records
193
+
194
+ ### Worker Architecture
195
+
196
+ The package uses a multi-worker architecture for efficient processing:
197
+ - **VectorizeTemplates Worker**: Handles template-based text generation and embedding
198
+ - **UpsertVectors Worker**: Manages vector database operations
199
+ - **EntityRecordDocument Worker**: Tracks vector-entity relationships
200
+
201
+ ## Configuration
202
+
203
+ ### Environment Variables
204
+
205
+ Create a `.env` file with:
206
+
207
+ ```env
208
+ # Database Configuration
209
+ DB_HOST=your-database-host
210
+ DB_PORT=1433
211
+ DB_USERNAME=your-username
212
+ DB_PASSWORD=your-password
213
+ DB_DATABASE=your-database
214
+
215
+ # API Keys
216
+ OPENAI_API_KEY=your-openai-key
217
+ MISTRAL_API_KEY=your-mistral-key
218
+ PINECONE_API_KEY=your-pinecone-key
219
+ PINECONE_HOST=your-pinecone-host
220
+ PINECONE_DEFAULT_INDEX=your-default-index
221
+
222
+ # User Configuration
223
+ CURRENT_USER_EMAIL=user@example.com
224
+ ```
225
+
226
+ ## Performance Considerations
227
+
228
+ - **Long-Running Processes**: Vectorization can take hours for large datasets
229
+ - **Batch Sizes**: Adjust batch sizes based on your system resources
230
+ - **Asynchronous Processing**: Consider running vectorization in background processes
231
+ - **Memory Usage**: Monitor memory usage for large batch sizes
232
+
233
+ ## Integration with MemberJunction
234
+
235
+ This package integrates seamlessly with:
236
+ - `@memberjunction/core`: Core entity and metadata functionality
237
+ - `@memberjunction/ai`: AI model abstractions
238
+ - `@memberjunction/ai-vectordb`: Vector database abstractions
239
+ - `@memberjunction/templates`: Template processing engine
240
+
241
+ ## Error Handling
242
+
243
+ The package includes comprehensive error handling:
244
+ - Validation of entity documents and templates
245
+ - Graceful handling of API failures
246
+ - Detailed logging through MemberJunction's logging system
247
+
248
+ ## Best Practices
249
+
250
+ 1. **Start with Small Batches**: Test with small batch sizes before processing large datasets
251
+ 2. **Monitor Progress**: Use MemberJunction's logging to track vectorization progress
252
+ 3. **Handle Interruptions**: Use `StartingOffset` to resume interrupted processes
253
+ 4. **Template Design**: Design templates to include relevant fields for semantic search
254
+ 5. **Resource Management**: Consider database and API rate limits when setting batch sizes
255
+
256
+ ## License
257
+
258
+ ISC - See LICENSE file for details
259
+
260
+ ## Author
261
+
262
+ MemberJunction.com
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@memberjunction/ai-vector-sync",
3
- "version": "2.43.0",
3
+ "version": "2.45.0",
4
4
  "description": "MemberJunction: AI Vector/Entity Sync Package - handles synchronization between MemberJunction entities and vector databases",
5
5
  "main": "./dist/index.js",
6
6
  "types": "./dist/index.d.ts",
@@ -15,17 +15,17 @@
15
15
  "author": "MemberJunction.com",
16
16
  "license": "ISC",
17
17
  "dependencies": {
18
- "@memberjunction/ai": "2.43.0",
19
- "@memberjunction/ai-vectordb": "2.43.0",
20
- "@memberjunction/ai-vectors": "2.43.0",
21
- "@memberjunction/ai-vectors-pinecone": "2.43.0",
22
- "@memberjunction/ai-mistral": "2.43.0",
23
- "@memberjunction/aiengine": "2.43.0",
24
- "@memberjunction/core": "2.43.0",
25
- "@memberjunction/global": "2.43.0",
26
- "@memberjunction/templates": "2.43.0",
27
- "@memberjunction/templates-base-types": "2.43.0",
28
- "@memberjunction/ai-openai": "2.43.0",
18
+ "@memberjunction/ai": "2.45.0",
19
+ "@memberjunction/ai-vectordb": "2.45.0",
20
+ "@memberjunction/ai-vectors": "2.45.0",
21
+ "@memberjunction/ai-vectors-pinecone": "2.45.0",
22
+ "@memberjunction/ai-mistral": "2.45.0",
23
+ "@memberjunction/aiengine": "2.45.0",
24
+ "@memberjunction/core": "2.45.0",
25
+ "@memberjunction/global": "2.45.0",
26
+ "@memberjunction/templates": "2.45.0",
27
+ "@memberjunction/templates-base-types": "2.45.0",
28
+ "@memberjunction/ai-openai": "2.45.0",
29
29
  "dotenv": "^16.4.1",
30
30
  "typeorm": "^0.3.20"
31
31
  },