@memberjunction/ai-prompts 4.0.0 → 4.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (2) hide show
  1. package/README.md +155 -3104
  2. package/package.json +11 -11
package/README.md CHANGED
@@ -1,3195 +1,246 @@
1
1
  # @memberjunction/ai-prompts
2
2
 
3
- Advanced AI prompt execution engine with hierarchical template composition, intelligent model selection, parallel execution, output validation, and comprehensive execution tracking.
3
+ Advanced AI prompt execution engine for MemberJunction. Provides hierarchical template composition, intelligent model selection with failover, parallel execution with judge-based result selection, structured output validation with retry, comprehensive execution tracking, and streaming support. This is the primary interface for executing AI prompts in the MemberJunction framework.
4
4
 
5
- > **Note on Parameters**: This package uses the parameter types defined in `@memberjunction/ai`. For a complete reference of available LLM parameters (temperature, topP, topK, etc.), see the [Parameter Reference](../Core/README.md#parameter-reference) in the AI Core documentation.
5
+ ## Architecture
6
6
 
7
- > **Authentication**: For details on how AI provider credentials are resolved, including the hierarchical credential system, integration with encrypted credentials, and legacy environment variable support, see the [AI Authentication Guide](./AI-AUTHENTICATION.md).
8
-
9
- [![npm version](https://badge.fury.io/js/%40memberjunction%2Fai-prompts.svg)](https://www.npmjs.com/package/@memberjunction/ai-prompts)
10
- [![License: ISC](https://img.shields.io/badge/License-ISC-blue.svg)](https://opensource.org/licenses/ISC)
11
-
12
- ## Key Features
13
-
14
- ### 🎯 Effort Level Control
15
- Granular control over AI model reasoning effort through a 1-100 integer scale. Higher values request more thorough reasoning and analysis from AI models that support effort levels.
7
+ ```mermaid
8
+ graph TD
9
+ subgraph "@memberjunction/ai-prompts"
10
+ PR["AIPromptRunner"]
11
+ style PR fill:#2d8659,stroke:#1a5c3a,color:#fff
16
12
 
17
- #### Effort Level Hierarchy
18
- The effort level is resolved using the following precedence (highest to lowest priority):
13
+ EP["ExecutionPlanner"]
14
+ style EP fill:#7c5295,stroke:#563a6b,color:#fff
19
15
 
20
- 1. **`AIPromptParams.effortLevel`** - Runtime override (highest priority)
21
- 2. **`AIPrompt.EffortLevel`** - Individual prompt setting (lower priority)
22
- 3. **Provider default** - Model's natural behavior (lowest priority)
16
+ PEC["ParallelExecutionCoordinator"]
17
+ style PEC fill:#7c5295,stroke:#563a6b,color:#fff
23
18
 
24
- #### Provider Support
25
- Different AI providers map the 1-100 scale to their specific parameters:
26
- - **OpenAI**: Maps to `reasoning_effort` (1-33=low, 34-66=medium, 67-100=high)
27
- - **Anthropic**: Maps to thinking mode with token budgets (1-100 → 25K-2M tokens)
28
- - **Groq**: Maps to experimental `reasoning_effort` parameter
29
- - **Gemini**: Controls reasoning mode intensity
19
+ PE["ParallelExecution"]
20
+ style PE fill:#7c5295,stroke:#563a6b,color:#fff
21
+ end
30
22
 
31
- ```typescript
32
- const params = new AIPromptParams();
33
- params.prompt = myPrompt;
34
- params.effortLevel = 85; // High effort for thorough analysis
23
+ subgraph "Execution Pipeline"
24
+ T["1. Template Rendering<br/>Handlebars + System Placeholders"]
25
+ style T fill:#b8762f,stroke:#8a5722,color:#fff
35
26
 
36
- const result = await AIPromptRunner.RunPrompt(params);
37
- ```
27
+ MS["2. Model Selection<br/>Default / Specific / ByPower"]
28
+ style MS fill:#b8762f,stroke:#8a5722,color:#fff
38
29
 
39
- ### 🛡️ Model Selection & Intelligent Failover
30
+ EX["3. LLM Execution<br/>With Streaming & Caching"]
31
+ style EX fill:#b8762f,stroke:#8a5722,color:#fff
40
32
 
41
- The AI Prompts system provides sophisticated model selection with instant failover across models and vendors. Configure explicit model/vendor priorities using `AIPromptModel` records, and the system automatically tries all candidates in order when errors occur.
33
+ VAL["4. Output Validation<br/>JSON Schema + Retry"]
34
+ style VAL fill:#b8762f,stroke:#8a5722,color:#fff
42
35
 
43
- #### Selection Strategies
36
+ TRK["5. Execution Tracking<br/>AIPromptRun Records"]
37
+ style TRK fill:#b8762f,stroke:#8a5722,color:#fff
38
+ end
44
39
 
45
- **`SelectionStrategy='Specific'`** (Recommended for production)
46
- - Use explicit `AIPromptModel` configuration for complete control
47
- - Configuration-specific models tried before universal fallbacks
48
- - Priority determines order (higher number = tried first)
49
- - Instant failover to next candidate on any error
40
+ PR --> EP
41
+ PR --> PEC
42
+ PEC --> PE
43
+ PR --> T
44
+ T --> MS
45
+ MS --> EX
46
+ EX --> VAL
47
+ VAL --> TRK
50
48
 
51
- **`SelectionStrategy='ByPower'`**
52
- - Automatically selects models based on `PowerRank`
53
- - Use `PowerPreference`: `Highest`, `Lowest`, or `Balanced`
49
+ subgraph Dependencies
50
+ AI["@memberjunction/ai<br/>BaseLLM"]
51
+ style AI fill:#2d6a9f,stroke:#1a4971,color:#fff
54
52
 
55
- **`SelectionStrategy='Default'`**
56
- - Uses model type filtering and power ranking
53
+ ACP["@memberjunction/ai-core-plus<br/>AIPromptParams"]
54
+ style ACP fill:#2d6a9f,stroke:#1a4971,color:#fff
57
55
 
58
- #### Model Ranking Algorithm (Specific Strategy)
56
+ AIE["@memberjunction/aiengine<br/>AIEngine"]
57
+ style AIE fill:#2d6a9f,stroke:#1a4971,color:#fff
59
58
 
60
- When using `SelectionStrategy='Specific'`, candidates are prioritized using clear, predictable rules:
59
+ TMPL["@memberjunction/templates<br/>TemplateEngine"]
60
+ style TMPL fill:#2d6a9f,stroke:#1a4971,color:#fff
61
61
 
62
- ```mermaid
63
- sequenceDiagram
64
- participant User as User Request
65
- participant Engine as AIPromptRunner
66
- participant DB as AIPromptModel Table
67
- participant Exec as Execution
68
-
69
- User->>Engine: Execute Prompt with ConfigurationID
70
- Engine->>DB: Get AIPromptModel records
71
- DB-->>Engine: Return all records for prompt
72
-
73
- Note over Engine: Filter Phase
74
- Engine->>Engine: Keep: ConfigurationID match OR NULL
75
- Engine->>Engine: Exclude: Different ConfigurationID
76
-
77
- Note over Engine: Sort Phase (2-level)
78
- Engine->>Engine: 1. Config-match before Universal
79
- Engine->>Engine: 2. Priority DESC within group
80
-
81
- Note over Engine: Expand Phase
82
- loop For each AIPromptModel
83
- alt VendorID specified
84
- Engine->>Engine: Create 1 candidate (Model+Vendor)
85
- else VendorID is NULL
86
- Engine->>Engine: Create N candidates (all vendors)
87
- Engine->>Engine: Sort by AIModelVendor.Priority DESC
88
- end
62
+ CRED["@memberjunction/credentials<br/>CredentialEngine"]
63
+ style CRED fill:#2d6a9f,stroke:#1a4971,color:#fff
89
64
  end
90
65
 
91
- Engine->>Exec: Try candidates in order
92
-
93
- loop Instant Failover
94
- Exec->>Exec: Try Candidate N
95
- alt Success
96
- Exec-->>User: Return result
97
- else Recoverable Error
98
- Exec->>Exec: Try Candidate N+1 (instant)
99
- else Fatal Error
100
- Exec-->>User: Fail immediately
101
- end
102
- end
66
+ AI --> PR
67
+ ACP --> PR
68
+ AIE --> PR
69
+ TMPL --> PR
70
+ CRED --> PR
103
71
  ```
104
72
 
105
- #### Configuration Rules
106
-
107
- **Priority Precedence:**
108
- 1. **Configuration-specific** models (matching `ConfigurationID`) - Always tried first
109
- 2. **Universal** models (`ConfigurationID = NULL`) - Fallback options
110
- 3. Within each group: **Higher Priority number** tried first
111
-
112
- **Configuration Filtering:**
113
- - If `ConfigurationID` provided: Use matching config + universal (NULL) models
114
- - If NO `ConfigurationID`: Use ONLY universal (NULL) models
115
- - Models with DIFFERENT `ConfigurationID` are EXCLUDED
116
-
117
- **Vendor Expansion:**
118
- - `AIPromptModel.VendorID` specified → Single candidate (exact model+vendor)
119
- - `AIPromptModel.VendorID = NULL` → Multiple candidates (all vendors for that model, sorted by `AIModelVendor.Priority DESC`)
120
-
121
- #### Example Configuration
122
-
123
- ```sql
124
- -- Example: Production prompt with config-specific and universal fallbacks
125
- INSERT INTO AIPromptModel (PromptID, ModelID, VendorID, ConfigurationID, Priority, Status) VALUES
126
- -- Config-specific models (tried first, regardless of priority number)
127
- (@promptId, @gpt4Id, @openaiId, @prodConfigId, 5, 'Active'),
128
- (@promptId, @gpt4Id, @azureId, @prodConfigId, 3, 'Active'),
129
-
130
- -- Universal fallbacks (tried after config-specific, despite higher priority numbers)
131
- (@promptId, @claudeId, @anthropicId, NULL, 10, 'Active'),
132
- (@promptId, @geminiId, @googleId, NULL, 8, 'Active');
73
+ ## Installation
133
74
 
134
- -- Example: Multi-vendor support for same model
135
- INSERT INTO AIPromptModel (PromptID, ModelID, VendorID, ConfigurationID, Priority, Status) VALUES
136
- (@promptId, @gpt4Id, NULL, @prodConfigId, 10, 'Active');
137
- -- VendorID=NULL expands to all vendors (OpenAI, Azure, Groq)
138
- -- Vendors sorted by AIModelVendor.Priority
75
+ ```bash
76
+ npm install @memberjunction/ai-prompts
139
77
  ```
140
78
 
141
- **Execution order for above config with ConfigurationID=@prodConfigId:**
142
- 1. GPT-4/OpenAI (Config match, Priority 5)
143
- 2. GPT-4/Azure (Config match, Priority 3)
144
- 3. GPT-4/OpenAI (From VendorID=NULL expansion, highest AIModelVendor.Priority)
145
- 4. GPT-4/Azure (From VendorID=NULL expansion)
146
- 5. GPT-4/Groq (From VendorID=NULL expansion, lowest AIModelVendor.Priority)
147
- 6. Claude/Anthropic (Universal fallback, Priority 10)
148
- 7. Gemini/Google (Universal fallback, Priority 8)
149
-
150
- #### Failover Behavior
79
+ ## Key Features
151
80
 
152
- **Instant Failover** - No delays between candidates
153
- - Authentication errors → Filters out all candidates from failed vendor
154
- - Fatal errors → Stops immediately
155
- - Recoverable errors → Tries next candidate instantly
81
+ ### Hierarchical Template Composition
156
82
 
157
- **Validation Retry** - After all candidates exhausted
158
- - If all candidates fail, retries entire list with delays
159
- - Uses `AIPrompt.MaxRetries` and `RetryDelayMode` (Fixed/Linear/Exponential)
83
+ Build complex prompts from reusable sub-templates with unlimited nesting depth:
160
84
 
161
- **Error Handling:**
162
85
  ```typescript
163
- // SelectionStrategy='Specific' with no candidates throws error
164
- if (strategy === 'Specific' && candidates.length === 0) {
165
- throw new Error('Please configure AIPromptModel records for this prompt');
166
- }
167
- ```
168
-
169
- ### 🎯 Dynamic Hierarchical Template Composition
170
-
171
- #### Why Dynamic Template Composition?
172
-
173
- While MemberJunction's template system already supports static template composition (where Template A always includes Templates B and C), the AI Prompts system adds **dynamic template composition** - the ability to inject ANY prompt template into ANY other prompt template at runtime.
86
+ import { AIPromptRunner } from '@memberjunction/ai-prompts';
87
+ import { AIPromptParams, ChildPromptParam } from '@memberjunction/ai-core-plus';
174
88
 
175
- **Static Composition (MJ Templates):** Perfect for fixed relationships like email headers/footers
176
- ```liquid
177
- <!-- Email template always includes same header -->
178
- {% include 'email-header' %}
179
- {{ content }}
180
- {% include 'email-footer' %}
181
- ```
89
+ const runner = new AIPromptRunner();
182
90
 
183
- **Dynamic Composition (AI Prompts):** Essential for flexible runtime relationships
184
- ```typescript
185
- // Inject ANY child prompt into ANY parent prompt at runtime
91
+ // Parent template uses {{ analysis }} and {{ summary }} placeholders
186
92
  const params = new AIPromptParams();
187
- params.prompt = systemPrompt; // e.g., Agent Type's control flow prompt
188
- params.childPrompts = [
189
- new ChildPromptParam(agentPrompt, 'agentInstructions') // Specific agent's prompt
190
- ];
191
- // System prompt can use {{ agentInstructions }} to embed the agent's specific logic
192
- ```
193
-
194
- #### The Agent System Use Case
195
-
196
- This dynamic composition is crucial for AI Agents:
197
- - **Agent Types** have **System Prompts** that control execution flow and response format
198
- - **Individual Agents** have their own **specific prompts** with domain logic
199
- - At runtime, any agent's prompt is dynamically injected into its type's system prompt
200
- - This creates a complete prompt combining the control wrapper with agent-specific instructions
201
-
202
- ```typescript
203
- // Agent Type System Prompt (controls flow)
204
- const systemPrompt = {
205
- templateText: `You are an AI agent. Follow these instructions:
206
-
207
- {{ agentInstructions }} <!-- Dynamically injected at runtime -->
208
-
209
- Respond in JSON format with: { decision: ..., reasoning: ... }`
210
- };
211
-
212
- // Individual Agent Prompt (domain logic)
213
- const dataGatherAgent = {
214
- templateText: `Your role is to gather data from: {{ dataSources }}`
215
- };
216
-
217
- // At runtime, compose them dynamically
93
+ params.prompt = parentPrompt;
218
94
  params.childPrompts = [
219
- new ChildPromptParam(dataGatherAgent, 'agentInstructions')
95
+ new ChildPromptParam(analysisParams, 'analysis'),
96
+ new ChildPromptParam(summaryParams, 'summary')
220
97
  ];
221
- ```
222
-
223
- ### 🔄 System Placeholders
224
- Automatically inject common values into all templates without manual data passing. Includes date/time, user context, prompt metadata, and more.
225
-
226
- ```liquid
227
- Current user: {{ _USER_NAME }}
228
- Date: {{ _CURRENT_DATE }}
229
- Expected output: {{ _OUTPUT_EXAMPLE }}
230
- ```
231
-
232
- ## System Placeholders Reference
233
-
234
- System placeholders are automatically available in all AI prompt templates, providing dynamic values like current date/time, prompt metadata, and user context without requiring manual data passing.
235
-
236
- ### Available System Placeholders
237
-
238
- #### Date/Time Placeholders
239
- - `{{ _CURRENT_DATE }}` - Current date in YYYY-MM-DD format
240
- - `{{ _CURRENT_TIME }}` - Current time in HH:MM AM/PM format with timezone
241
- - `{{ _CURRENT_DATE_AND_TIME }}` - Full timestamp with date and time
242
- - `{{ _CURRENT_DAY_OF_WEEK }}` - Current day name (e.g., Monday, Tuesday)
243
- - `{{ _CURRENT_TIMEZONE }}` - Current timezone identifier
244
- - `{{ _CURRENT_TIMESTAMP_UTC }}` - Current UTC timestamp in ISO format
245
-
246
- #### Prompt Metadata Placeholders
247
- - `{{ _OUTPUT_EXAMPLE }}` - The expected output example from the prompt configuration
248
- - `{{ _PROMPT_NAME }}` - The name of the current prompt
249
- - `{{ _PROMPT_DESCRIPTION }}` - The description of the current prompt
250
- - `{{ _EXPECTED_OUTPUT_TYPE }}` - The expected output type (string, object, number, etc.)
251
- - `{{ _RESPONSE_FORMAT }}` - The expected response format from the prompt
252
-
253
- #### User Context Placeholders
254
- - `{{ _USER_NAME }}` - Current user's full name
255
- - `{{ _USER_EMAIL }}` - Current user's email address
256
- - `{{ _USER_ID }}` - Current user's unique identifier
257
-
258
- #### Environment Placeholders
259
- - `{{ _ENVIRONMENT }}` - Current environment (development, staging, production)
260
- - `{{ _API_VERSION }}` - Current API version
261
-
262
- ### System Placeholder Usage Examples
263
-
264
- #### Example 1: Time-Aware Agent Prompt
265
- ```liquid
266
- You are an AI assistant helping {{ _USER_NAME }} on {{ _CURRENT_DAY_OF_WEEK }}, {{ _CURRENT_DATE }} at {{ _CURRENT_TIME }}.
267
-
268
- User's request: {{ userRequest }}
269
-
270
- Please provide a helpful response considering the current time and day.
271
- ```
272
-
273
- #### Example 2: Agent Type System Prompt with Metadata
274
- ```liquid
275
- # Agent Type: Loop Decision Maker
276
-
277
- Current execution context:
278
- - Date/Time: {{ _CURRENT_DATE_AND_TIME }}
279
- - User: {{ _USER_NAME }} ({{ _USER_EMAIL }})
280
- - Environment: {{ _ENVIRONMENT }}
281
-
282
- ## Expected Output Format
283
- {{ _OUTPUT_EXAMPLE }}
284
-
285
- ## Agent Specific Instructions
286
- {{ agentResponse }}
98
+ params.data = { userInput: 'complex data to process' };
287
99
 
288
- Based on the above agent response and the expected output format ({{ _EXPECTED_OUTPUT_TYPE }}), determine the next step.
289
- ```
290
-
291
- #### Example 3: Debug-Friendly Prompt
292
- ```liquid
293
- [Debug Info]
294
- - Prompt: {{ _PROMPT_NAME }}
295
- - Description: {{ _PROMPT_DESCRIPTION }}
296
- - Expected Output: {{ _EXPECTED_OUTPUT_TYPE }}
297
- - User ID: {{ _USER_ID }}
298
- - Timestamp: {{ _CURRENT_TIMESTAMP_UTC }}
299
-
300
- [Task]
301
- {{ taskDescription }}
302
- ```
303
-
304
- ### Adding Custom System Placeholders
305
-
306
- You can add custom system placeholders programmatically:
307
-
308
- ```typescript
309
- import { SystemPlaceholderManager } from '@memberjunction/ai-prompts';
310
-
311
- // Add a custom placeholder
312
- SystemPlaceholderManager.addPlaceholder({
313
- name: '_ORGANIZATION_NAME',
314
- description: 'Current organization name',
315
- getValue: async (params) => {
316
- // Custom logic to get organization name
317
- return params.contextUser?.OrganizationName || 'Default Organization';
318
- }
319
- });
320
-
321
- // Or add directly to the array
322
- const placeholders = SystemPlaceholderManager.getPlaceholders();
323
- placeholders.push({
324
- name: '_CUSTOM_VALUE',
325
- description: 'My custom value',
326
- getValue: async (params) => 'custom result'
327
- });
328
- ```
329
-
330
- ### Data Merge Priority Order
331
-
332
- When rendering templates, data is merged in this priority order (highest to lowest):
333
- 1. Template-specific data (`templateData` parameter)
334
- 2. Child template renders (for hierarchical template composition)
335
- 3. User-provided data (`data` parameter)
336
- 4. System placeholders (lowest priority)
337
-
338
- This means users can override system placeholders by providing their own values with the same names.
339
-
340
- ### ⚡ Parallel Processing
341
- Multi-model execution with intelligent result selection strategies and AI judge ranking for optimal results.
342
-
343
- ### ✅ Output Validation
344
- JSON schema validation against OutputExample with intelligent retry logic and configurable validation behaviors.
345
-
346
- ### 🚫 Cancellation Support
347
- AbortSignal integration for graceful execution cancellation with proper cleanup and partial result preservation.
348
-
349
- ### 📈 Progress & Streaming
350
- Real-time progress callbacks and streaming response support for responsive user interfaces.
351
-
352
- ### 📊 Comprehensive Tracking
353
- Hierarchical execution logging with the AIPromptRun entity, including token usage, timing, and validation attempts.
354
-
355
- ### 🤖 Agent Integration
356
- Seamless integration with AI Agents through hierarchical prompts and execution tracking.
357
-
358
- ### 💾 Intelligent Caching
359
- Vector similarity matching and TTL-based result caching for performance optimization.
360
-
361
- ### 🔧 Template Integration
362
- Dynamic prompt generation with MemberJunction template system supporting conditionals, loops, and data injection.
363
-
364
- ## Installation
365
-
366
- ```bash
367
- npm install @memberjunction/ai-prompts
100
+ const result = await runner.ExecutePrompt(params);
368
101
  ```
369
102
 
370
- > **Note**: This package uses MemberJunction's class registration system. The package automatically registers its classes on import to ensure proper functionality within the MJ ecosystem.
103
+ Execution order:
104
+ 1. Child prompts render depth-first (children before parents)
105
+ 2. Sibling prompts at each level execute in parallel
106
+ 3. Child results replace placeholders in parent template
107
+ 4. Final composed prompt executes as a single LLM call
371
108
 
372
- ### Type Organization Update (2025)
109
+ ### Model Selection Strategies
373
110
 
374
- As part of improving code organization:
375
- - **This package** now imports base AI types from `@memberjunction/ai` (Core)
376
- - **Prompt-specific types** remain in this package:
377
- - `AIPromptParams`, `AIPromptRunResult`
378
- - `ChildPromptParam`, `SystemPlaceholder`
379
- - Execution callbacks and progress types
380
- - **Agent integration types** are imported from `@memberjunction/ai-agents` when needed
111
+ Three strategies for selecting which AI model executes a prompt:
381
112
 
382
- ## Requirements
113
+ | Strategy | Description |
114
+ |---|---|
115
+ | `Default` | Uses the AI configuration to determine the model based on priority and availability |
116
+ | `Specific` | Uses explicitly associated models from the AIPromptModels table |
117
+ | `ByPower` | Selects the highest PowerRank model matching the prompt's model type |
383
118
 
384
- - Node.js 16+
385
- - MemberJunction Core libraries
386
- - [@memberjunction/ai](../Core/README.md) for base AI types and result structures
387
- - [@memberjunction/aiengine](../Engine/README.md) for model management and basic AI operations
388
- - [@memberjunction/templates](../../Templates/README.md) for template rendering
119
+ Model selection precedence (highest to lowest):
120
+ 1. `AIPromptParams.override` -- Runtime model/vendor override
121
+ 2. `AIPromptParams.modelSelectionPrompt` -- Alternate prompt for model config
122
+ 3. Prompt's own model configuration (strategy + associations)
389
123
 
390
- ## Core Architecture
124
+ ### Parallel Execution with Judging
391
125
 
392
- ### Dynamic vs Static Template Composition
126
+ Execute prompts across multiple models simultaneously and select the best result:
393
127
 
394
- The AI Prompts system introduces **dynamic template composition** that extends beyond MemberJunction's built-in static template features:
128
+ - Configurable execution groups with different models
129
+ - AI judge prompt evaluates and ranks results
130
+ - Automatic selection of best result based on judge scoring
131
+ - Full tracking of all parallel results
395
132
 
396
- #### Static Template Composition (MJ Templates)
397
- MemberJunction's template system supports embedding templates within templates through `{% include %}` directives. This is perfect for fixed relationships:
398
- - Email templates with standard headers/footers
399
- - Report templates with consistent formatting sections
400
- - Any scenario where Template A always includes Templates B and C
133
+ ### Output Validation and Retry
401
134
 
402
- #### Dynamic Template Composition (AI Prompts)
403
- The AI Prompts system adds runtime template composition where relationships are determined dynamically:
404
- - **Runtime Flexibility**: Inject ANY prompt template into ANY other prompt template
405
- - **Context-Aware**: Choose which child templates to inject based on runtime conditions
406
- - **Agent Architecture**: Combine system prompts (control flow) with agent prompts (domain logic)
407
- - **Modular Design**: Build complex prompts from reusable components selected at runtime
135
+ Automatic validation of AI outputs with configurable retry:
408
136
 
409
- **Key Difference**: While MJ Templates handle "Template A always includes B", AI Prompts handle "Template A includes X, where X is determined at runtime"
137
+ - JSON schema validation against `OutputExample` definitions
138
+ - Automatic JSON repair via JSON5 parsing and LLM-based repair
139
+ - Configurable retry count with the original or repaired prompts
140
+ - Validation syntax cleaning (removes `?`, `*`, `:type` markers from JSON keys)
141
+ - Detailed validation attempt tracking
410
142
 
411
- ### AIPromptRunner Class
143
+ ### Streaming Support
412
144
 
413
- The `AIPromptRunner` class is the central component for executing prompts with advanced features:
145
+ Real-time streaming of LLM responses:
414
146
 
415
147
  ```typescript
416
- import { AIPromptRunner, AIPromptParams } from '@memberjunction/ai-prompts';
417
-
418
- // Get a prompt from the system
419
- const prompts = AIEngine.Instance.Prompts;
420
- const summaryPrompt = prompts.find(p => p.Name === 'Document Summarization');
421
-
422
- // Execute the prompt
423
- const params: AIPromptParams = {
424
- prompt: summaryPrompt,
425
- data: {
426
- documentText: "Long document content here...",
427
- targetLength: "2 paragraphs"
428
- },
429
- contextUser: currentUser
148
+ const params = new AIPromptParams();
149
+ params.prompt = myPrompt;
150
+ params.onStreaming = (chunk) => {
151
+ process.stdout.write(chunk.content);
430
152
  };
431
153
 
432
- const runner = new AIPromptRunner();
433
154
  const result = await runner.ExecutePrompt(params);
434
-
435
- if (result.success) {
436
- console.log("Summary:", result.result);
437
- console.log(`Execution time: ${result.executionTimeMS}ms`);
438
- console.log(`Prompt tokens: ${result.promptTokens}`);
439
- console.log(`Completion tokens: ${result.completionTokens}`);
440
- console.log(`Total tokens: ${result.tokensUsed}`);
441
- if (result.cost) {
442
- console.log(`Cost: ${result.cost} ${result.costCurrency || 'USD'}`);
443
- }
444
- } else {
445
- console.error("Error:", result.errorMessage);
446
- }
447
155
  ```
448
156
 
449
- ## Quick Start
157
+ ### Execution Tracking
450
158
 
451
- ### 1. Basic Prompt Execution
159
+ Every prompt execution creates an `AIPromptRun` record with:
160
+ - Model and vendor used
161
+ - Template rendering results
162
+ - Token usage (prompt + completion)
163
+ - Cost tracking
164
+ - Execution time
165
+ - Parent/child relationships for hierarchical prompts
166
+ - Agent run linkage via `agentRunId`
452
167
 
453
- ```typescript
454
- import { AIPromptRunner } from '@memberjunction/ai-prompts';
455
- import { AIEngine } from '@memberjunction/aiengine';
456
-
457
- // Initialize the AI Engine
458
- await AIEngine.Instance.Config(false, currentUser);
459
-
460
- // Find a prompt
461
- const prompt = AIEngine.Instance.Prompts.find(p => p.Name === 'Text Analysis');
462
-
463
- // Execute with data
464
- const runner = new AIPromptRunner();
465
- const result = await runner.ExecutePrompt({
466
- prompt: prompt,
467
- data: {
468
- text: "Analyze this sample text for sentiment and key themes.",
469
- format: "bullet points"
470
- },
471
- contextUser: currentUser
472
- });
473
-
474
- console.log("Analysis:", result.result);
475
- ```
476
-
477
- ### 2. Template-Driven Prompts
478
-
479
- ```typescript
480
- // Prompt templates support dynamic data substitution
481
- const templatePrompt = {
482
- UserMessage: `Analyze the {{entity.EntityType}} record for {{entity.Name}}.
483
- Focus on {{analysisType}} and provide insights about {{entity.Description}}.`
484
- };
485
-
486
- // Data context provides template variables
487
- const result = await runner.ExecutePrompt({
488
- prompt: templatePrompt,
489
- data: {
490
- entity: {
491
- EntityType: "Customer",
492
- Name: "Acme Corp",
493
- Description: "Enterprise software company"
494
- },
495
- analysisType: "growth opportunities"
496
- },
497
- contextUser: currentUser
498
- });
499
- ```
500
-
501
- ### 3. Parallel Execution with Multiple Models
502
-
503
- ```typescript
504
- // Execute the same prompt across multiple models in parallel
505
- const multiModelPrompt = prompts.find(p => p.ParallelizationMode === 'ModelSpecific');
506
-
507
- const result = await runner.ExecutePrompt({
508
- prompt: multiModelPrompt,
509
- data: { query: "Analyze this data pattern" },
510
- contextUser: currentUser
511
- });
512
-
513
- // When using parallel execution, the system automatically selects the best result
514
- console.log(`Final result: ${result.result}`);
515
- console.log(`Execution time: ${result.executionTimeMS}ms`);
516
- console.log(`Total tokens used: ${result.tokensUsed}`);
517
-
518
- // The promptRun entity contains metadata about parallel execution in its Messages field
519
- if (result.promptRun?.Messages) {
520
- const metadata = JSON.parse(result.promptRun.Messages);
521
- if (metadata.parallelExecution) {
522
- console.log(`Parallelization mode: ${metadata.parallelExecution.parallelizationMode}`);
523
- console.log(`Total tasks: ${metadata.parallelExecution.totalTasks}`);
524
- console.log(`Successful tasks: ${metadata.parallelExecution.successfulTasks}`);
525
- }
526
- }
527
- ```
528
-
529
- ### 4. Dynamic Template Composition for AI Agents
168
+ ### Credential Resolution
530
169
 
531
- This example demonstrates the primary use case for dynamic template composition - the AI Agent system:
170
+ Hierarchical credential resolution for API keys:
532
171
 
533
- ```typescript
534
- import { AIPromptRunner, ChildPromptParam } from '@memberjunction/ai-prompts';
535
-
536
- // Agent Type System Prompt - Controls execution flow and response format
537
- const agentTypeSystemPrompt = {
538
- Name: "Data Analysis Agent Type System Prompt",
539
- TemplateID: "system-prompt-template-id",
540
- // Template contains: "You are an AI agent. {{ agentInstructions }} Respond with JSON..."
541
- };
172
+ 1. `AIPromptParams.credentialId` (per-request override)
173
+ 2. `AIPromptModel.CredentialID` (prompt-model specific)
174
+ 3. `AIModelVendor.CredentialID` (model-vendor specific)
175
+ 4. `AIVendor.CredentialID` (vendor default)
176
+ 5. `AIPromptParams.apiKeys[]` (legacy runtime keys)
177
+ 6. `AI_VENDOR_API_KEY__<DRIVER>` environment variables (legacy)
542
178
 
543
- // Individual Agent Prompt - Contains domain-specific logic
544
- const specificAgentPrompt = {
545
- Name: "Customer Churn Analysis Agent",
546
- TemplateID: "churn-agent-template-id",
547
- // Template contains: "Analyze customer data for churn risk factors..."
548
- };
179
+ ### Failover
549
180
 
550
- // At runtime, dynamically compose the prompts
551
- const runner = new AIPromptRunner();
552
- const result = await runner.ExecutePrompt({
553
- prompt: agentTypeSystemPrompt, // Parent template
554
- childPrompts: [
555
- // Dynamically inject the specific agent's instructions
556
- new ChildPromptParam(specificAgentPrompt, 'agentInstructions')
557
- ],
558
- data: {
559
- customerData: analysisData,
560
- thresholds: { churnRisk: 0.7 }
561
- },
562
- contextUser: currentUser
563
- });
564
-
565
- // The system executed ONE prompt that combined:
566
- // 1. System prompt wrapper (control flow)
567
- // 2. Specific agent instructions (domain logic)
568
- // 3. Runtime data
569
- console.log("Agent decision:", result.result);
570
- ```
181
+ When a model fails due to rate limiting, authentication errors, or other transient issues, the runner can automatically retry with alternate models from the selection candidates.
571
182
 
572
- **Why This Matters:**
573
- - Different agents can use the SAME system prompt template
574
- - System prompt enforces consistent response format across all agents
575
- - Agent-specific logic is cleanly separated and reusable
576
- - Runtime composition allows flexible agent architectures
183
+ ## Usage
577
184
 
578
- ### 5. Complete Example with All New Features
185
+ ### Basic Prompt Execution
579
186
 
580
187
  ```typescript
581
188
  import { AIPromptRunner } from '@memberjunction/ai-prompts';
189
+ import { AIPromptParams } from '@memberjunction/ai-core-plus';
582
190
  import { AIEngine } from '@memberjunction/aiengine';
583
191
 
584
- // Complete example showcasing all Phase 6 enhancements
585
- async function comprehensivePromptExecution() {
586
- // Initialize
587
- await AIEngine.Instance.Config(false, currentUser);
588
- const runner = new AIPromptRunner();
589
-
590
- // Set up cancellation (e.g., from user clicking cancel button)
591
- const controller = new AbortController();
592
- const timeoutId = setTimeout(() => {
593
- controller.abort();
594
- console.log('Operation timed out after 2 minutes');
595
- }, 120000);
596
-
597
- try {
598
- const result = await runner.ExecutePrompt({
599
- prompt: complexAnalysisPrompt, // ParallelizationMode: 'ModelSpecific'
600
- data: {
601
- document: largeDocument,
602
- analysisType: 'comprehensive',
603
- outputFormat: 'structured'
604
- },
605
- contextUser: currentUser,
606
-
607
- // Enable cancellation
608
- cancellationToken: controller.signal,
609
-
610
- // Track progress throughout execution
611
- onProgress: (progress) => {
612
- console.log(`[${progress.step}] ${progress.percentage}% - ${progress.message}`);
613
-
614
- // Handle parallel execution progress
615
- if (progress.metadata?.parallelExecution) {
616
- const parallel = progress.metadata.parallelExecution;
617
- console.log(` → Group ${parallel.currentGroup + 1}/${parallel.totalGroups}, Tasks: ${parallel.completedTasks}/${parallel.totalTasks}`);
618
- }
619
-
620
- // Update UI
621
- updateProgressBar(progress.percentage);
622
- updateStatusText(progress.message);
623
- },
624
-
625
- // Receive streaming content updates
626
- onStreaming: (chunk) => {
627
- if (chunk.isComplete) {
628
- console.log(`Streaming complete for ${chunk.modelName}`);
629
- finalizeOutput();
630
- } else {
631
- // Show real-time content generation
632
- console.log(`[${chunk.modelName}]: ${chunk.content.substring(0, 50)}...`);
633
- appendToDisplay(chunk.content, chunk.taskId);
634
- }
635
- }
636
- });
637
-
638
- // Clear timeout since we completed successfully
639
- clearTimeout(timeoutId);
640
-
641
- // Handle different result scenarios
642
- if (result.cancelled) {
643
- console.log(`Execution cancelled: ${result.cancellationReason}`);
644
- // May still have partial results available
645
- if (result.additionalResults && result.additionalResults.length > 0) {
646
- console.log(`${result.additionalResults.length} partial results available`);
647
- }
648
- } else if (result.success) {
649
- console.log('Execution completed successfully!');
650
- console.log(`Primary result from ${result.modelInfo?.modelName}: ${result.result}`);
651
-
652
- // Analyze judge selection if multiple results
653
- if (result.ranking && result.judgeRationale) {
654
- console.log(`Selected as #${result.ranking} by AI judge: ${result.judgeRationale}`);
655
- }
656
-
657
- // Review alternative results from parallel execution
658
- if (result.additionalResults) {
659
- console.log(`${result.additionalResults.length} alternative results ranked by judge:`);
660
- result.additionalResults.forEach((altResult, index) => {
661
- console.log(` ${altResult.ranking}. ${altResult.modelInfo?.modelName}: ${altResult.judgeRationale}`);
662
- });
663
- }
664
-
665
- // Analyze execution performance using hierarchical logging
666
- if (result.promptRun?.RunType === 'ParallelParent') {
667
- await analyzeParallelExecutionPerformance(result.promptRun.ID);
668
- }
669
-
670
- // Check streaming and caching
671
- if (result.wasStreamed) {
672
- console.log('Response was streamed in real-time');
673
- }
674
- if (result.cacheInfo?.cacheHit) {
675
- console.log(`Result served from cache: ${result.cacheInfo.cacheSource}`);
676
- }
677
- } else {
678
- console.error(`Execution failed: ${result.errorMessage}`);
679
- }
680
-
681
- } catch (error) {
682
- clearTimeout(timeoutId);
683
- console.error('Execution error:', error.message);
684
- }
685
- }
686
-
687
- // Helper function to analyze parallel execution performance
688
- async function analyzeParallelExecutionPerformance(parentPromptRunId: string) {
689
- // Query hierarchical logs to understand execution breakdown
690
- console.log('Analyzing parallel execution performance...');
691
-
692
- // This would typically be a database query or API call
693
- // For demonstration, showing the concept:
694
- const analysisQuery = `
695
- SELECT
696
- pr.RunType,
697
- pr.ExecutionOrder,
698
- pr.Success,
699
- pr.ExecutionTimeMS,
700
- pr.TokensUsed,
701
- m.Name as ModelName
702
- FROM AIPromptRun pr
703
- JOIN AIModel m ON pr.ModelID = m.ID
704
- WHERE pr.ParentID = '${parentPromptRunId}' OR pr.ID = '${parentPromptRunId}'
705
- ORDER BY pr.RunType, pr.ExecutionOrder
706
- `;
707
-
708
- console.log('Performance analysis query:', analysisQuery);
709
- // Execute query and analyze results...
710
- }
711
-
712
- // Execute the comprehensive example
713
- comprehensivePromptExecution().catch(console.error);
714
- ```
715
-
716
- ## Advanced Features
192
+ // Get prompt from metadata
193
+ await AIEngine.Instance.Config(false, contextUser);
194
+ const prompt = AIEngine.Instance.Prompts.find(p => p.Name === 'Summarize Content');
717
195
 
718
- ### Intelligent Failover System
719
-
720
- The AI Prompt Runner includes a sophisticated failover system that automatically handles provider outages, rate limits, and service degradation. This ensures your AI-powered applications remain resilient and responsive even when individual providers experience issues.
721
-
722
- #### How Failover Works
723
-
724
- When a prompt execution fails, the system:
725
- 1. **Analyzes the error** using the ErrorAnalyzer to determine if failover is appropriate
726
- 2. **Selects alternative models/vendors** based on the configured strategy
727
- 3. **Applies intelligent delays** with exponential backoff to prevent overwhelming providers
728
- 4. **Tracks all attempts** for debugging and analysis
729
- 5. **Updates the execution** to use the successful model/vendor combination
730
-
731
- #### Failover Configuration
732
-
733
- Configure failover behavior at the prompt level:
734
-
735
- ```typescript
736
- // Database columns added to AIPrompt entity:
737
- FailoverStrategy: 'SameModelDifferentVendor' | 'NextBestModel' | 'PowerRank' | 'None'
738
- FailoverMaxAttempts: number // Maximum failover attempts (default: 3)
739
- FailoverDelaySeconds: number // Initial delay between attempts (default: 1)
740
- FailoverModelStrategy: 'PreferSameModel' | 'PreferDifferentModel' | 'RequireSameModel'
741
- FailoverErrorScope: 'All' | 'NetworkOnly' | 'RateLimitOnly' | 'ServiceErrorOnly'
742
- ```
743
-
744
- #### Failover Strategies Explained
745
-
746
- **SameModelDifferentVendor**: Ideal for multi-cloud deployments
747
- ```typescript
748
- // Example: Claude from different providers
749
- // Primary: Anthropic API
750
- // Failover 1: AWS Bedrock
751
- // Failover 2: Google Vertex AI
752
- ```
753
-
754
- **NextBestModel**: Balances capability and availability
755
- ```typescript
756
- // Example: Gradual capability reduction
757
- // Primary: GPT-4-turbo
758
- // Failover 1: Claude-3-opus
759
- // Failover 2: GPT-3.5-turbo
760
- ```
761
-
762
- **PowerRank**: Uses MemberJunction's model power rankings
763
- ```typescript
764
- // Automatically selects models based on their PowerRank scores
765
- // Ensures you always get the best available model
766
- ```
767
-
768
- #### Error Scope Configuration
769
-
770
- Control which types of errors trigger failover:
771
-
772
- - **All**: Any error triggers failover (most resilient)
773
- - **NetworkOnly**: Only network/connection errors
774
- - **RateLimitOnly**: Only rate limit errors (429 status)
775
- - **ServiceErrorOnly**: Only service errors (500, 503 status)
776
-
777
- #### Failover Tracking
778
-
779
- The system comprehensively tracks failover attempts in the database:
780
-
781
- ```typescript
782
- // AIPromptRun entity tracking fields:
783
- OriginalModelID: string // The initially selected model
784
- OriginalRequestStartTime: Date // When the request started
785
- FailoverAttempts: number // Number of failover attempts made
786
- FailoverErrors: string (JSON) // Detailed error information for each attempt
787
- FailoverDurations: string (JSON) // Duration of each attempt in milliseconds
788
- TotalFailoverDuration: number // Total time spent in failover
789
- ```
790
-
791
- #### Advanced Failover Customization
792
-
793
- The AIPromptRunner exposes protected methods for advanced customization:
794
-
795
- ```typescript
796
- class CustomPromptRunner extends AIPromptRunner {
797
- // Override to implement custom failover configuration
798
- protected getFailoverConfiguration(prompt: AIPromptEntity): FailoverConfiguration {
799
- // Add environment-specific logic
800
- if (process.env.NODE_ENV === 'production') {
801
- return {
802
- strategy: 'SameModelDifferentVendor',
803
- maxAttempts: 5,
804
- delaySeconds: 2,
805
- modelStrategy: 'PreferSameModel',
806
- errorScope: 'NetworkOnly'
807
- };
808
- }
809
- return super.getFailoverConfiguration(prompt);
810
- }
811
-
812
- // Override to implement custom failover decision logic
813
- protected shouldAttemptFailover(
814
- error: Error,
815
- config: FailoverConfiguration,
816
- attemptNumber: number
817
- ): boolean {
818
- // Add custom error analysis
819
- if (error.message.includes('quota_exceeded')) {
820
- return false; // Don't retry quota errors
821
- }
822
- return super.shouldAttemptFailover(error, config, attemptNumber);
823
- }
824
-
825
- // Override to implement custom delay calculation
826
- protected calculateFailoverDelay(
827
- attemptNumber: number,
828
- baseDelaySeconds: number,
829
- previousError?: Error
830
- ): number {
831
- // Custom backoff strategy
832
- if (previousError?.message.includes('rate_limit')) {
833
- return 60000; // 1 minute for rate limits
834
- }
835
- return super.calculateFailoverDelay(attemptNumber, baseDelaySeconds, previousError);
836
- }
837
- }
838
- ```
839
-
840
- #### Failover Best Practices
841
-
842
- 1. **Configure Appropriately**: Use `NetworkOnly` or `RateLimitOnly` for production to avoid retrying invalid requests
843
- 2. **Set Reasonable Attempts**: 3-5 attempts typically sufficient
844
- 3. **Monitor Failover Patterns**: Query the tracking data to identify problematic providers
845
- 4. **Test Failover Scenarios**: Simulate provider outages in development
846
- 5. **Consider Costs**: Failover may route to more expensive providers
847
-
848
- #### Example: Production-Ready Configuration
849
-
850
- ```typescript
851
- const productionPrompt = {
852
- Name: "Customer Service Assistant",
853
- FailoverStrategy: "SameModelDifferentVendor",
854
- FailoverMaxAttempts: 4,
855
- FailoverDelaySeconds: 2,
856
- FailoverModelStrategy: "PreferSameModel",
857
- FailoverErrorScope: "NetworkOnly",
858
- // Ensure failover stays within approved models
859
- MinPowerRank: 85
860
- };
861
-
862
- // Query failover performance
863
- const failoverStats = await runView.RunView({
864
- EntityName: 'MJ: AI Prompt Runs',
865
- ExtraFilter: `FailoverAttempts > 0 AND RunAt >= '2024-01-01'`,
866
- OrderBy: 'RunAt DESC'
867
- });
868
-
869
- // Analyze which vendors are most reliable
870
- SELECT
871
- OriginalModelID,
872
- ModelID as FinalModelID,
873
- COUNT(*) as FailoverCount,
874
- AVG(TotalFailoverDuration) as AvgFailoverTime
875
- FROM AIPromptRun
876
- WHERE FailoverAttempts > 0
877
- GROUP BY OriginalModelID, ModelID
878
- ORDER BY FailoverCount DESC;
879
- ```
880
-
881
- #### Configuration-Aware Failover
882
-
883
- The failover system respects `AIConfiguration` boundaries to ensure environment-specific models stay isolated:
884
-
885
- **How It Works:**
886
- - When you specify a `configurationId`, the system builds a candidate list with two priority tiers:
887
- 1. **Configuration-specific models** (priority 5000+): Models assigned to your configuration
888
- 2. **NULL configuration models** (priority 2000+): Universal fallback models available to all configurations
889
-
890
- **Example Setup:**
891
- ```sql
892
- -- Production Configuration: Only approved production models
893
- INSERT INTO AIPromptModel (PromptID, ModelID, ConfigurationID, Priority)
894
- VALUES
895
- (@PromptID, @Claude35SonnetID, @ProductionConfigID, 100),
896
- (@PromptID, @GPT4ID, @ProductionConfigID, 90);
897
-
898
- -- Development Configuration: Include experimental models
899
- INSERT INTO AIPromptModel (PromptID, ModelID, ConfigurationID, Priority)
900
- VALUES
901
- (@PromptID, @LlamaExperimentalID, @DevelopmentConfigID, 100);
902
-
903
- -- NULL Configuration: Universal fallbacks for all environments
904
- INSERT INTO AIPromptModel (PromptID, ModelID, ConfigurationID, Priority)
905
- VALUES
906
- (@PromptID, @Claude3HaikuID, NULL, 100),
907
- (@PromptID, @GPT35TurboID, NULL, 90);
908
- ```
909
-
910
- **Failover Behavior:**
911
- ```typescript
912
- // Execute with Production configuration
913
- const result = await runner.ExecutePrompt({
914
- prompt: myPrompt,
915
- configurationId: productionConfigID,
916
- data: { query: 'Analyze this' }
917
- });
918
-
919
- // Failover order:
920
- // 1. Try Claude 3.5 Sonnet (Production config, priority 5100)
921
- // 2. Try GPT-4 (Production config, priority 5090)
922
- // 3. Try Claude 3 Haiku (NULL config fallback, priority 2100)
923
- // 4. Try GPT-3.5 Turbo (NULL config fallback, priority 2090)
924
- // ✅ Never crosses to Development config models
925
- ```
926
-
927
- **Key Benefits:**
928
- - **Environment Isolation**: Production models never failover to development/experimental models
929
- - **Controlled Fallback**: Explicit hierarchy from config-specific to universal fallbacks
930
- - **Performance**: Candidate list built once and cached, no rebuilding during failover
931
- - **Consistency**: Same candidate list used for initial selection and all failover attempts
932
-
933
- ### Intelligent Caching
934
-
935
- The prompt system provides sophisticated caching with vector similarity matching:
936
-
937
- ```typescript
938
- // Caching is automatically handled based on prompt configuration:
939
- // - EnableCaching: Whether to use caching for this prompt
940
- // - CacheMatchType: 'Exact' or 'Vector' similarity matching
941
- // - CacheTTLSeconds: Time-to-live for cached results
942
- // - CacheMustMatchModel/Vendor/Agent: Cache constraint options
943
-
944
- // Vector similarity allows reusing results for semantically similar prompts
945
- // even if the exact text differs
946
-
947
- const cachedPrompt = {
948
- Name: "Smart Summary",
949
- EnableCaching: true,
950
- CacheMatchType: "Vector",
951
- CacheTTLSeconds: 3600,
952
- CacheSimilarityThreshold: 0.85,
953
- CacheMustMatchModel: true,
954
- CacheMustMatchVendor: false
955
- };
956
- ```
957
-
958
- ### Parallel Execution Strategies
959
-
960
- The system supports multiple parallelization modes:
961
-
962
- ```typescript
963
- // Prompts can be configured for parallel execution:
964
- // - ParallelizationMode: 'None', 'StaticCount', 'ConfigParam', 'ModelSpecific'
965
- // - ParallelCount: Number of parallel executions
966
- // - ExecutionGroups: Sequential group execution with parallel tasks within groups
967
-
968
- // Example configurations:
969
-
970
- // Static parallel count
971
- const staticParallelPrompt = {
972
- ParallelizationMode: "StaticCount",
973
- ParallelCount: 3
974
- };
975
-
976
- // Configuration-driven count
977
- const configParallelPrompt = {
978
- ParallelizationMode: "ConfigParam",
979
- ParallelConfigParam: "analysis_parallel_count"
980
- };
981
-
982
- // Model-specific configuration
983
- const modelSpecificPrompt = {
984
- ParallelizationMode: "ModelSpecific",
985
- // Uses settings from AIPromptModel entries
986
- };
987
- ```
988
-
989
- ### Result Selection Strategies
990
-
991
- ```typescript
992
- // The engine supports multiple result selection methods:
993
- // - 'First': Use the first successful result
994
- // - 'Random': Randomly select from successful results
995
- // - 'PromptSelector': Use AI to select the best result
996
- // - 'Consensus': Select result with highest agreement
997
-
998
- // Result selector prompts can be configured to intelligently choose
999
- // the best result from parallel executions
1000
-
1001
- const selectorPrompt = {
1002
- Name: "Best Result Selector",
1003
- PromptText: `
1004
- You are evaluating multiple AI responses to select the best one.
1005
- Original query: {{originalQuery}}
1006
-
1007
- Responses:
1008
- {{#each responses}}
1009
- Response {{@index}}: {{this}}
1010
- {{/each}}
1011
-
1012
- Select the response number (0-based) that is most accurate, helpful, and well-written.
1013
- Return only the number.
1014
- `,
1015
- OutputType: "number"
1016
- };
1017
-
1018
- const mainPrompt = {
1019
- ParallelizationMode: "StaticCount",
1020
- ParallelCount: 3,
1021
- ResultSelectorPromptID: selectorPrompt.ID
1022
- };
1023
- ```
1024
-
1025
- ### Output Validation
1026
-
1027
- ```typescript
1028
- // Configure structured output validation
1029
- const validatedPrompt = {
1030
- Name: "Structured Analysis",
1031
- OutputType: "object",
1032
- OutputExample: {
1033
- sentiment: "positive|negative|neutral",
1034
- confidence: 0.95,
1035
- keyThemes: ["theme1", "theme2"],
1036
- summary: "Brief summary text"
1037
- },
1038
- ValidationBehavior: "Strict",
1039
- MaxRetries: 3,
1040
- RetryDelayMS: 1000,
1041
- RetryStrategy: "exponential"
1042
- };
1043
-
1044
- // Validation is automatically applied
1045
- const result = await runner.ExecutePrompt({
1046
- prompt: validatedPrompt,
1047
- data: { text: "Content to analyze" },
1048
- contextUser: currentUser,
1049
- skipValidation: false // Validation enabled
1050
- });
1051
-
1052
- // Result.result will be validated against the expected structure
1053
- ```
1054
-
1055
- ### Validation Syntax Cleaning
1056
-
1057
- When using output validation with JSON responses, the AI Prompt Runner automatically handles validation syntax that AI models might inadvertently include in their JSON keys:
196
+ const runner = new AIPromptRunner();
197
+ const params = new AIPromptParams();
198
+ params.prompt = prompt;
199
+ params.data = { content: documentText, maxLength: 500 };
200
+ params.contextUser = contextUser;
1058
201
 
1059
- ```typescript
1060
- // Validation syntax in prompts:
1061
- // - name?: optional field
1062
- // - items:[2+]: array with minimum 2 items
1063
- // - status:!empty: non-empty required field
1064
- // - count:number: field with type hint
1065
-
1066
- // If the AI returns JSON with validation syntax in keys:
1067
- {
1068
- "name?": "John Doe",
1069
- "items:[2+]": ["apple", "banana", "orange"],
1070
- "status:!empty": "active",
1071
- "count:number": 42
1072
- }
202
+ const result = await runner.ExecutePrompt(params);
1073
203
 
1074
- // The system automatically cleans it to:
1075
- {
1076
- "name": "John Doe",
1077
- "items": ["apple", "banana", "orange"],
1078
- "status": "active",
1079
- "count": 42
204
+ if (result.success) {
205
+ console.log(result.result); // Parsed/validated result
206
+ console.log(result.promptTokens); // Input tokens used
207
+ console.log(result.completionTokens); // Output tokens generated
208
+ console.log(result.executionTimeMS); // Execution duration
1080
209
  }
1081
210
  ```
1082
211
 
1083
- **Automatic Cleaning Behavior:**
1084
- - **Always enabled** when prompt has `ValidationBehavior` set to `"Strict"` or `"Warn"`
1085
- - **Always enabled** when prompt has an `OutputExample` defined
1086
- - **Optional** for prompts with `ValidationBehavior` set to `"None"` (via `cleanValidationSyntax` parameter)
1087
-
1088
- This ensures that validation patterns used in prompt templates don't interfere with the actual JSON structure returned by the AI model.
1089
-
1090
- ```
1091
-
1092
- ### Template Integration
1093
-
1094
- Advanced template features with the MemberJunction template system:
1095
-
1096
- ```typescript
1097
- // Complex template with conditionals and loops
1098
- const advancedTemplate = {
1099
- PromptText: `
1100
- Analyze the following {{entityType}} records:
1101
-
1102
- {{#each records}}
1103
- {{@index + 1}}. {{this.Name}}
1104
- Status: {{this.Status}}
1105
- {{#if this.Priority}}Priority: {{this.Priority}}{{/if}}
1106
- {{#each this.Tags}}
1107
- - Tag: {{this}}
1108
- {{/each}}
1109
- {{/each}}
1110
-
1111
- {{#if includeRecommendations}}
1112
- Please provide recommendations for improvement.
1113
- {{/if}}
1114
-
1115
- Focus on: {{analysisAreas.join(", ")}}
1116
- `
1117
- };
1118
-
1119
- const result = await runner.ExecutePrompt({
1120
- prompt: advancedTemplate,
1121
- data: {
1122
- entityType: "Customer",
1123
- records: customerData,
1124
- includeRecommendations: true,
1125
- analysisAreas: ["revenue potential", "risk factors", "engagement"]
1126
- },
1127
- contextUser: currentUser
1128
- });
1129
- ```
1130
-
1131
- ## Parallel Execution System
1132
-
1133
- The package includes sophisticated parallel execution capabilities through specialized classes that work together to manage complex multi-model executions.
1134
-
1135
- > **Note**: The ExecutionPlanner and ParallelExecutionCoordinator are internal components used by AIPromptRunner. They are not directly exposed in the public API but understanding their operation helps in configuring prompts effectively.
1136
-
1137
- ### ExecutionPlanner (Internal)
1138
-
1139
- The `ExecutionPlanner` class analyzes prompt configuration and creates optimal execution strategies:
1140
-
1141
- **Key Responsibilities:**
1142
- - Analyzes parallelization modes (None, StaticCount, ConfigParam, ModelSpecific)
1143
- - Creates execution groups for coordinated processing
1144
- - Determines optimal task distribution based on model availability
1145
- - Assigns priorities and manages execution order
1146
- - Handles model selection based on power rankings and configuration
1147
-
1148
- **Execution Plan Creation:**
1149
- - For `StaticCount`: Creates N parallel tasks using available models
1150
- - For `ConfigParam`: Uses configuration parameters to determine parallel count
1151
- - For `ModelSpecific`: Uses AIPromptModel entries to define exact model usage
1152
- - Supports execution groups for sequential/parallel hybrid execution
1153
-
1154
- ### ParallelExecutionCoordinator (Internal)
1155
-
1156
- The `ParallelExecutionCoordinator` orchestrates the actual execution of tasks created by the ExecutionPlanner:
1157
-
1158
- **Core Features:**
1159
- - Manages concurrency limits (default: 5 concurrent executions)
1160
- - Implements retry logic with exponential backoff
1161
- - Handles partial result collection when some tasks fail
1162
- - Provides comprehensive execution metrics and timing
1163
- - Supports fail-fast mode for critical operations
1164
-
1165
- **Execution Flow:**
1166
- 1. Groups tasks by execution group number
1167
- 2. Executes groups sequentially (group 0, then 1, then 2, etc.)
1168
- 3. Within each group, executes tasks in parallel up to concurrency limit
1169
- 4. Collects and aggregates results from all executions
1170
- 5. Applies result selection strategy if multiple results available
1171
-
1172
- ### Supported Parallelization Modes
1173
-
1174
- - **None**: Traditional single execution
1175
- - **StaticCount**: Fixed number of parallel executions
1176
- - **ConfigParam**: Dynamic parallel count from configuration
1177
- - **ModelSpecific**: Individual model configurations with execution groups
212
+ ### With Progress Tracking
1178
213
 
1179
214
  ```typescript
1180
- // Example of model-specific parallel configuration
1181
- const modelSpecificExecution = {
1182
- prompt: complexPrompt,
1183
- data: analysisData,
1184
- contextUser: currentUser
215
+ params.onProgress = (progress) => {
216
+ console.log(`[${progress.step}] ${progress.percentage}% - ${progress.message}`);
1185
217
  };
1186
-
1187
- // The system will:
1188
- // 1. Query AIPromptModel entries for this prompt
1189
- // 2. Group executions by ExecutionGroup
1190
- // 3. Execute groups sequentially, models within groups in parallel
1191
- // 4. Apply result selection strategy
1192
- const result = await runner.ExecutePrompt(modelSpecificExecution);
1193
218
  ```
1194
219
 
1195
- ## Performance Monitoring & Analytics
1196
-
1197
- Comprehensive tracking and analytics for prompt executions:
220
+ ### With Effort Level
1198
221
 
1199
222
  ```typescript
1200
- // Execution results include detailed metrics
1201
- const result = await runner.ExecutePrompt(params);
1202
-
1203
- console.log(`Execution time: ${result.executionTimeMS}ms`);
1204
- console.log(`Tokens used: ${result.tokensUsed}`);
1205
-
1206
- // The AIPromptRunResult includes execution tracking
1207
- if (result.promptRun) {
1208
- console.log(`Prompt Run ID: ${result.promptRun.ID}`);
1209
- console.log(`Model used: ${result.promptRun.ModelID}`);
1210
- console.log(`Configuration: ${result.promptRun.ConfigurationID}`);
1211
- }
223
+ params.effortLevel = 85; // High effort for thorough analysis (1-100 scale)
1212
224
  ```
1213
225
 
1214
- ### Early Run ID Callback
1215
-
1216
- Get the PromptRun ID immediately after creation for real-time monitoring:
226
+ ### With Runtime Model Override
1217
227
 
1218
228
  ```typescript
1219
- const params = new AIPromptParams();
1220
- params.prompt = myPrompt;
1221
- params.data = { query: 'Analyze this data' };
1222
-
1223
- // Callback fired immediately after PromptRun record is saved
1224
- params.onPromptRunCreated = async (promptRunId) => {
1225
- console.log(`Prompt run started: ${promptRunId}`);
1226
-
1227
- // Use cases:
1228
- // - Link to parent records (e.g., AIAgentRunStep.TargetLogID)
1229
- // - Send to monitoring systems
1230
- // - Update UI with tracking info
1231
- // - Start real-time log streaming
229
+ params.override = {
230
+ modelId: 'specific-model-id',
231
+ vendorId: 'specific-vendor-id'
1232
232
  };
1233
-
1234
- const result = await runner.ExecutePrompt(params);
1235
233
  ```
1236
234
 
1237
- The callback is invoked:
1238
- - **When**: Right after the AIPromptRun record is created and saved
1239
- - **Before**: The actual AI model execution begins
1240
- - **Error Handling**: Callback errors are logged but don't fail the execution
1241
- - **Async Support**: Can be synchronous or asynchronous
1242
-
1243
- ## AI Prompt Run Logging
1244
-
1245
- The AI Prompt Runner implements a sophisticated hierarchical logging system that tracks all execution activities in the database through the `AIPromptRun` entity. This system provides complete traceability and analytics for both simple and complex parallel executions.
1246
-
1247
- ### Hierarchical Logging Structure
1248
-
1249
- The logging system uses a parent-child relationship model with different `RunType` values to represent the execution hierarchy:
1250
-
1251
- - **`Single`**: Standard single-model execution
1252
- - **`ParallelParent`**: Parent record for parallel execution coordinating multiple models
1253
- - **`ParallelChild`**: Individual model execution within a parallel run
1254
- - **`ResultSelector`**: AI judge execution that selects the best result from parallel executions
1255
-
1256
- ### RunType Values and Relationships
1257
-
1258
- ```typescript
1259
- // Single execution - no parent relationship
1260
- {
1261
- RunType: 'Single',
1262
- ParentID: null,
1263
- ExecutionOrder: null
1264
- }
1265
-
1266
- // Parallel execution creates a hierarchical structure:
1267
- // 1. Parent record coordinates the overall execution
1268
- {
1269
- RunType: 'ParallelParent',
1270
- ParentID: null,
1271
- ExecutionOrder: null
1272
- }
1273
-
1274
- // 2. Child records for each model execution
1275
- {
1276
- RunType: 'ParallelChild',
1277
- ParentID: '12345-parent-id',
1278
- ExecutionOrder: 0 // Order within execution group
1279
- }
1280
-
1281
- // 3. Result selector judges the best result
1282
- {
1283
- RunType: 'ResultSelector',
1284
- ParentID: '12345-parent-id',
1285
- ExecutionOrder: 5 // After all parallel children
1286
- }
1287
- ```
1288
-
1289
- ### Database Schema Fields
1290
-
1291
- Key fields in the `AIPromptRun` entity for hierarchical logging:
1292
-
1293
- ```sql
1294
- -- Core execution tracking
1295
- PromptID uniqueidentifier -- Prompt being executed
1296
- ModelID uniqueidentifier -- AI model used
1297
- VendorID uniqueidentifier -- Vendor providing the model
1298
- RunAt datetime2 -- Execution start time
1299
- CompletedAt datetime2 -- Execution completion time
1300
-
1301
- -- Hierarchical logging fields
1302
- RunType nvarchar(50) -- 'Single', 'ParallelParent', 'ParallelChild', 'ResultSelector'
1303
- ParentID uniqueidentifier -- Parent prompt run ID (NULL for top-level)
1304
- ExecutionOrder int -- Order within parallel execution group
1305
-
1306
- -- Results and metrics
1307
- Success bit -- Whether execution succeeded
1308
- Result nvarchar(max) -- Raw result from AI model
1309
- ErrorMessage nvarchar(500) -- Error message if failed
1310
- ExecutionTimeMS int -- Total execution time
1311
- TokensUsed int -- Total tokens consumed
1312
- TokensPrompt int -- Prompt tokens used
1313
- TokensCompletion int -- Completion tokens generated
1314
-
1315
- -- Cost tracking
1316
- Cost decimal(19,8) -- Cost of this specific execution
1317
- CostCurrency nvarchar(10) -- ISO 4217 currency code (USD, EUR, etc.)
1318
-
1319
- -- Hierarchical rollup fields (NEW)
1320
- TokensUsedRollup int -- Total tokens including all children
1321
- TokensPromptRollup int -- Total prompt tokens including all children
1322
- TokensCompletionRollup int -- Total completion tokens including all children
1323
- -- Note: TotalCost (existing field) serves as the cost rollup
1324
-
1325
- -- Context and configuration
1326
- Messages nvarchar(max) -- JSON with input data and metadata
1327
- ConfigurationID uniqueidentifier -- Environment configuration used
1328
- AgentRunID uniqueidentifier -- Links to parent AIAgentRun if applicable
1329
- ```
1330
-
1331
- ### Hierarchical Token and Cost Tracking
1332
-
1333
- The AI Prompts system implements a sophisticated rollup pattern for tracking token usage and costs across hierarchical prompt executions:
1334
-
1335
- #### Prompt Execution Rollup Pattern
1336
-
1337
- For hierarchical prompt executions (parent prompts with child prompts), each node in the tree contains:
1338
- - **Direct fields** (`TokensPrompt`, `TokensCompletion`, `Cost`): Usage for just that execution
1339
- - **Rollup fields** (`TokensPromptRollup`, `TokensCompletionRollup`, `TotalCost`): Total including all descendants
1340
-
1341
- **Example:**
1342
- ```
1343
- Parent Prompt (100 prompt, 200 completion tokens, $0.05)
1344
- ├── Child A (50 prompt, 100 completion, $0.02)
1345
- └── Child B (75 prompt, 150 completion, $0.03)
1346
-
1347
- Database records:
1348
- - Parent: TokensPrompt=100, TokensPromptRollup=225 (100+50+75)
1349
- TokensCompletion=200, TokensCompletionRollup=450 (200+100+150)
1350
- Cost=0.05, TotalCost=0.10 (0.05+0.02+0.03)
1351
- - Child A: TokensPrompt=50, TokensPromptRollup=50 (leaf node)
1352
- Cost=0.02, TotalCost=0.02 (leaf node)
1353
- - Child B: TokensPrompt=75, TokensPromptRollup=75 (leaf node)
1354
- Cost=0.03, TotalCost=0.03 (leaf node)
1355
- ```
1356
-
1357
- This enables efficient queries like:
1358
- - "What was the total cost of this hierarchical prompt?" → Check root's `TotalCost`
1359
- - "How many tokens did this sub-prompt and its children use?" → Check that node's rollup fields
1360
- - No complex SQL joins or recursive CTEs needed!
1361
-
1362
- #### Agent Run Token Tracking
1363
-
1364
- The `AIAgentRun` entity tracks aggregate token usage across all prompt executions during an agent's lifecycle:
1365
-
1366
- ```sql
1367
- -- New fields in AIAgentRun
1368
- TotalTokensUsed int -- Total tokens (existing)
1369
- TotalPromptTokensUsed int -- Breakdown: prompt tokens (NEW)
1370
- TotalCompletionTokensUsed int -- Breakdown: completion tokens (NEW)
1371
- TotalCost decimal -- Total cost (existing)
1372
-
1373
- -- Hierarchical agent rollup fields (NEW)
1374
- TotalTokensUsedRollup int -- Including sub-agent runs
1375
- TotalPromptTokensUsedRollup int -- Including sub-agent runs
1376
- TotalCompletionTokensUsedRollup int -- Including sub-agent runs
1377
- TotalCostRollup decimal -- Including sub-agent runs
1378
- ```
1379
-
1380
- **Agent Hierarchy Example:**
1381
- ```
1382
- Parent Agent (A)
1383
- ├── Own prompts: 200 prompt, 400 completion tokens
1384
- ├── Sub-Agent (B)
1385
- │ └── Own prompts: 100 prompt, 200 completion tokens
1386
- └── Sub-Agent (C)
1387
- └── Own prompts: 150 prompt, 300 completion tokens
1388
-
1389
- Rollup values:
1390
- - Agent A: TotalPromptTokensUsedRollup = 450 (200+100+150)
1391
- TotalCompletionTokensUsedRollup = 900 (400+200+300)
1392
- - Agent B: TotalPromptTokensUsedRollup = 100 (leaf agent)
1393
- - Agent C: TotalPromptTokensUsedRollup = 150 (leaf agent)
1394
- ```
1395
-
1396
- ### Querying Hierarchical Log Data
1397
-
1398
- The hierarchical structure enables powerful analytics queries:
1399
-
1400
- ```sql
1401
- -- Get all executions for a parallel run
1402
- SELECT
1403
- pr.ID,
1404
- pr.RunType,
1405
- pr.ExecutionOrder,
1406
- pr.Success,
1407
- pr.ExecutionTimeMS,
1408
- pr.TokensUsed,
1409
- m.Name as ModelName,
1410
- p.Name as PromptName
1411
- FROM AIPromptRun pr
1412
- JOIN AIModel m ON pr.ModelID = m.ID
1413
- JOIN AIPrompt p ON pr.PromptID = p.ID
1414
- WHERE pr.ParentID = '12345-parent-id'
1415
- OR pr.ID = '12345-parent-id'
1416
- ORDER BY pr.RunType, pr.ExecutionOrder;
1417
-
1418
- -- Analyze parallel execution performance
1419
- WITH ParallelStats AS (
1420
- SELECT
1421
- ParentID,
1422
- COUNT(*) as TotalChildren,
1423
- SUM(CASE WHEN Success = 1 THEN 1 ELSE 0 END) as SuccessfulChildren,
1424
- AVG(ExecutionTimeMS) as AvgExecutionTime,
1425
- SUM(TokensUsed) as TotalTokens
1426
- FROM AIPromptRun
1427
- WHERE RunType = 'ParallelChild'
1428
- AND ParentID IS NOT NULL
1429
- GROUP BY ParentID
1430
- )
1431
- SELECT
1432
- parent.ID as ParentRunID,
1433
- parent.RunAt,
1434
- parent.ExecutionTimeMS as ParentExecutionTime,
1435
- stats.TotalChildren,
1436
- stats.SuccessfulChildren,
1437
- stats.AvgExecutionTime,
1438
- stats.TotalTokens,
1439
- prompt.Name as PromptName
1440
- FROM AIPromptRun parent
1441
- JOIN ParallelStats stats ON parent.ID = stats.ParentID
1442
- JOIN AIPrompt prompt ON parent.PromptID = prompt.ID
1443
- WHERE parent.RunType = 'ParallelParent'
1444
- ORDER BY parent.RunAt DESC;
1445
-
1446
- -- Find failed executions with context
1447
- SELECT
1448
- pr.ID,
1449
- pr.RunType,
1450
- pr.ParentID,
1451
- pr.ErrorMessage,
1452
- pr.ExecutionTimeMS,
1453
- m.Name as ModelName,
1454
- v.Name as VendorName,
1455
- p.Name as PromptName
1456
- FROM AIPromptRun pr
1457
- JOIN AIModel m ON pr.ModelID = m.ID
1458
- LEFT JOIN AIVendor v ON pr.VendorID = v.ID
1459
- JOIN AIPrompt p ON pr.PromptID = p.ID
1460
- WHERE pr.Success = 0
1461
- ORDER BY pr.RunAt DESC;
1462
- ```
1463
-
1464
- ## Cancellation Support
1465
-
1466
- The AI Prompt Runner provides comprehensive cancellation support through the standard JavaScript `AbortSignal` and `AbortController` pattern, enabling graceful termination of long-running operations.
1467
-
1468
- ### Understanding AbortSignal in Prompt Execution
1469
-
1470
- The `AbortSignal` pattern separates **cancellation control** from **cancellation handling**:
1471
-
1472
- - **Your Code (Controller)**: Creates the `AbortController` and decides **when** to cancel
1473
- - **Prompt Runner (Worker)**: Receives the `AbortSignal` token and handles **how** to cancel gracefully
1474
-
1475
- This separation allows for flexible cancellation from multiple sources (user actions, timeouts, resource limits) while the Prompt Runner handles the complex cleanup across parallel executions, model calls, and result selection.
1476
-
1477
- **The Pattern Flow:**
1478
- ```
1479
- Controller (Your Code) → AbortController.signal → AIPromptRunner
1480
- ↓ ↓ ↓
1481
- Decides WHEN The "Red Phone" Handles HOW
1482
- to cancel Token to stop
1483
- ```
1484
-
1485
- ### Basic Cancellation Usage
1486
-
1487
- ```typescript
1488
- import { AIPromptRunner } from '@memberjunction/ai-prompts';
1489
-
1490
- // Create cancellation controller
1491
- const controller = new AbortController();
1492
- const cancellationToken = controller.signal;
1493
-
1494
- // Set up cancellation after 30 seconds
1495
- setTimeout(() => {
1496
- controller.abort();
1497
- console.log('Prompt execution cancelled due to timeout');
1498
- }, 30000);
1499
-
1500
- // Execute prompt with cancellation support
1501
- const runner = new AIPromptRunner();
1502
- const result = await runner.ExecutePrompt({
1503
- prompt: myPrompt,
1504
- data: { query: 'Long running analysis...' },
1505
- contextUser: currentUser,
1506
- cancellationToken: cancellationToken
1507
- });
1508
-
1509
- // Check if execution was cancelled
1510
- if (result.cancelled) {
1511
- console.log(`Execution cancelled: ${result.cancellationReason}`);
1512
- console.log('Partial results may be available');
1513
- } else if (result.success) {
1514
- console.log('Execution completed successfully');
1515
- }
1516
- ```
1517
-
1518
- ### Cancellation in Parallel Execution
1519
-
1520
- Cancellation works seamlessly with parallel execution, allowing you to stop all running tasks:
1521
-
1522
- ```typescript
1523
- const controller = new AbortController();
1524
-
1525
- // User clicks cancel button
1526
- document.getElementById('cancelButton').onclick = () => {
1527
- controller.abort();
1528
- };
1529
-
1530
- // Execute parallel prompt with multiple models
1531
- const result = await runner.ExecutePrompt({
1532
- prompt: parallelPrompt, // ParallelizationMode: 'ModelSpecific'
1533
- data: analysisData,
1534
- contextUser: currentUser,
1535
- cancellationToken: controller.signal
1536
- });
1537
-
1538
- // Parallel cancellation behavior:
1539
- // - Tasks not yet started will be marked as cancelled
1540
- // - Currently executing tasks will be terminated
1541
- // - Completed tasks remain in the results
1542
- // - Partial results may still be available for analysis
1543
- ```
1544
-
1545
- ### Multiple Cancellation Sources
1546
-
1547
- One of the powerful aspects of the AbortSignal pattern is that multiple sources can cancel the same operation:
1548
-
1549
- ```typescript
1550
- async function intelligentPromptExecution() {
1551
- const controller = new AbortController();
1552
- const signal = controller.signal;
1553
-
1554
- // 1. User cancel button
1555
- document.getElementById('cancelBtn')?.addEventListener('click', () => {
1556
- controller.abort(); // User-initiated cancellation
1557
- console.log('User cancelled the operation');
1558
- });
1559
-
1560
- // 2. Timeout cancellation (prevent runaway prompts)
1561
- const timeout = setTimeout(() => {
1562
- controller.abort(); // Timeout cancellation
1563
- console.log('Operation timed out after 2 minutes');
1564
- }, 120000);
1565
-
1566
- // 3. Resource limit cancellation
1567
- const memoryCheck = setInterval(async () => {
1568
- if (await getMemoryUsage() > MAX_MEMORY_THRESHOLD) {
1569
- controller.abort(); // Resource limit cancellation
1570
- console.log('Cancelled due to memory limits');
1571
- }
1572
- }, 5000);
1573
-
1574
- // 4. Window unload cancellation (cleanup on page close)
1575
- window.addEventListener('beforeunload', () => {
1576
- controller.abort(); // Page closing cancellation
1577
- });
1578
-
1579
- try {
1580
- const result = await runner.ExecutePrompt({
1581
- prompt: complexAnalysisPrompt,
1582
- data: largeDataset,
1583
- cancellationToken: signal // One token, many cancel sources!
1584
- });
1585
-
1586
- // Clean up timers if successful
1587
- clearTimeout(timeout);
1588
- clearInterval(memoryCheck);
1589
-
1590
- return result;
1591
- } catch (error) {
1592
- // The Prompt Runner doesn't know WHY it was cancelled
1593
- // It just knows it should stop gracefully
1594
- console.log('Prompt execution was cancelled:', error.message);
1595
- } finally {
1596
- clearTimeout(timeout);
1597
- clearInterval(memoryCheck);
1598
- }
1599
- }
1600
- ```
1601
-
1602
- ### Cancellation in Component-Based UIs
1603
-
1604
- Perfect for React, Angular, or Vue components:
1605
-
1606
- ```typescript
1607
- class PromptExecutionComponent {
1608
- private currentController: AbortController | null = null;
1609
- private isExecuting: boolean = false;
1610
-
1611
- async executePrompt(prompt: AIPromptEntity, data: any) {
1612
- // Cancel any existing execution
1613
- this.cancelCurrentExecution();
1614
-
1615
- // Create new controller for this execution
1616
- this.currentController = new AbortController();
1617
- this.isExecuting = true;
1618
-
1619
- try {
1620
- const result = await this.runner.ExecutePrompt({
1621
- prompt,
1622
- data,
1623
- cancellationToken: this.currentController.signal,
1624
- onProgress: (progress) => {
1625
- this.updateUI(`${progress.step}: ${progress.percentage}%`);
1626
- },
1627
- onStreaming: (chunk) => {
1628
- this.appendStreamingContent(chunk.content);
1629
- }
1630
- });
1631
-
1632
- this.handleSuccess(result);
1633
- } catch (error) {
1634
- if (error.message.includes('cancelled')) {
1635
- this.handleCancellation();
1636
- } else {
1637
- this.handleError(error);
1638
- }
1639
- } finally {
1640
- this.isExecuting = false;
1641
- this.currentController = null;
1642
- }
1643
- }
1644
-
1645
- // Called when user clicks "Cancel" or navigates away
1646
- cancelCurrentExecution() {
1647
- if (this.currentController && this.isExecuting) {
1648
- this.currentController.abort();
1649
- console.log('Cancelled current prompt execution');
1650
- }
1651
- }
1652
-
1653
- // Component cleanup
1654
- ngOnDestroy() { // Angular example
1655
- this.cancelCurrentExecution();
1656
- }
1657
- }
1658
- ```
1659
-
1660
- ### Integration with BaseLLM Cancellation
1661
-
1662
- The cancellation token is automatically propagated through the entire execution chain:
1663
-
1664
- ```typescript
1665
- // Cancellation Flow in MemberJunction AI Architecture:
1666
- //
1667
- // 1. User Code (AbortController.signal)
1668
- // ↓
1669
- // 2. AIPromptRunner.ExecutePrompt(cancellationToken)
1670
- // ↓
1671
- // 3. ParallelExecutionCoordinator.executeTasksInParallel(cancellationToken)
1672
- // ↓
1673
- // 4. Individual Task Execution with cancellation
1674
- // ↓
1675
- // 5. BaseLLM.ChatCompletion({ cancellationToken })
1676
- // ↓
1677
- // 6. Provider-specific cancellation (fetch signal, Promise.race)
1678
- // ↓
1679
- // 7. AI Model API cancellation (if supported)
1680
-
1681
- // At each level, cancellation is handled appropriately:
1682
- const internalFlow = {
1683
- // Level 1: Prompt Runner checks before major operations
1684
- promptRunner: () => {
1685
- if (cancellationToken?.aborted) {
1686
- return { success: false, cancelled: true };
1687
- }
1688
- },
1689
-
1690
- // Level 2: Parallel coordinator cancels remaining tasks
1691
- parallelCoordinator: () => {
1692
- tasks.forEach(task => {
1693
- if (cancellationToken?.aborted) {
1694
- task.cancelled = true;
1695
- }
1696
- });
1697
- },
1698
-
1699
- // Level 3: BaseLLM uses Promise.race for instant cancellation
1700
- baseLLM: () => {
1701
- return Promise.race([
1702
- actualModelCall(params),
1703
- cancellationPromise(cancellationToken)
1704
- ]);
1705
- },
1706
-
1707
- // Level 4: Native provider cancellation (where supported)
1708
- provider: () => {
1709
- fetch(apiUrl, {
1710
- signal: cancellationToken // Native browser/Node.js cancellation
1711
- });
1712
- }
1713
- };
1714
- ```
1715
-
1716
- ### Cancellation Guarantees
1717
-
1718
- The AI Prompt Runner provides these cancellation guarantees:
1719
-
1720
- 1. **🚫 Instant Recognition**: Cancellation requests are checked at multiple points throughout execution
1721
- 2. **🧹 Graceful Cleanup**: Partial results are preserved and returned when possible
1722
- 3. **📊 Proper Logging**: Cancelled operations are logged with appropriate status and metadata
1723
- 4. **💾 Resource Release**: Network connections and memory are cleaned up promptly
1724
- 5. **🔄 State Consistency**: The system remains in a consistent state after cancellation
1725
-
1726
- **Key Benefits:**
1727
- - **Responsive UI**: Users get immediate feedback when cancelling operations
1728
- - **Resource Efficiency**: Prevents wasted compute and API costs
1729
- - **System Stability**: Avoids memory leaks and hanging operations
1730
- - **Standard Pattern**: Uses native JavaScript APIs - no custom cancellation logic needed
1731
-
1732
- ### Cancellation Result Properties
1733
-
1734
- When execution is cancelled, the result includes detailed cancellation information:
1735
-
1736
- ```typescript
1737
- interface AIPromptRunResult {
1738
- success: boolean;
1739
- cancelled?: boolean; // True if execution was cancelled
1740
- cancellationReason?: CancellationReason; // Why it was cancelled
1741
- status?: ExecutionStatus; // Current execution status
1742
- // ... other properties
1743
- }
1744
-
1745
- type CancellationReason = 'user_requested' | 'timeout' | 'error' | 'resource_limit';
1746
- type ExecutionStatus = 'pending' | 'running' | 'completed' | 'failed' | 'cancelled';
1747
- ```
1748
-
1749
- ## Progress Updates & Streaming
1750
-
1751
- The AI Prompt Runner provides real-time progress updates and streaming support for long-running executions, enabling responsive user interfaces and monitoring dashboards.
1752
-
1753
- ### Progress Callbacks
1754
-
1755
- Track execution progress through different phases:
1756
-
1757
- ```typescript
1758
- const runner = new AIPromptRunner();
1759
-
1760
- const result = await runner.ExecutePrompt({
1761
- prompt: complexPrompt,
1762
- data: { document: longDocument },
1763
- contextUser: currentUser,
1764
-
1765
- // Progress callback receives updates throughout execution
1766
- onProgress: (progress) => {
1767
- console.log(`${progress.step}: ${progress.percentage}% - ${progress.message}`);
1768
-
1769
- // Update UI progress bar
1770
- updateProgressBar(progress.percentage);
1771
- updateStatusMessage(progress.message);
1772
-
1773
- // Access additional metadata
1774
- if (progress.metadata) {
1775
- console.log('Execution metadata:', progress.metadata);
1776
- }
1777
- }
1778
- });
1779
- ```
1780
-
1781
- ### Execution Progress Phases
1782
-
1783
- The progress callback receives updates for these execution phases:
1784
-
1785
- ```typescript
1786
- type ProgressPhase =
1787
- | 'template_rendering' // Rendering prompt template with data
1788
- | 'model_selection' // Selecting appropriate AI model
1789
- | 'execution' // Executing AI model
1790
- | 'validation' // Validating and parsing results
1791
- | 'parallel_coordination' // Coordinating parallel executions
1792
- | 'result_selection'; // AI judge selecting best result
1793
-
1794
- // Example progress updates:
1795
- // template_rendering: 20% - "Rendering prompt template with provided data"
1796
- // model_selection: 40% - "Selected GPT-4 model based on prompt configuration"
1797
- // execution: 60% - "Executing AI model..."
1798
- // validation: 80% - "Validating output against expected format"
1799
- // result_selection: 90% - "AI judge selecting best result from 3 candidates"
1800
- ```
1801
-
1802
- ### Streaming Response Support
1803
-
1804
- Receive real-time content updates as AI models generate responses:
1805
-
1806
- ```typescript
1807
- const result = await runner.ExecutePrompt({
1808
- prompt: streamingPrompt,
1809
- data: { query: 'Generate a detailed report...' },
1810
- contextUser: currentUser,
1811
-
1812
- // Streaming callback receives content chunks as they arrive
1813
- onStreaming: (chunk) => {
1814
- if (chunk.isComplete) {
1815
- console.log('Streaming complete');
1816
- finalizeDocument();
1817
- } else {
1818
- // Append content chunk to UI
1819
- appendToDocument(chunk.content);
1820
-
1821
- // Show which model is generating content (for parallel execution)
1822
- if (chunk.modelName) {
1823
- showActiveModel(chunk.modelName);
1824
- }
1825
- }
1826
- }
1827
- });
1828
- ```
1829
-
1830
- ### Progress Updates in Parallel Execution
1831
-
1832
- Progress tracking works seamlessly with parallel execution:
1833
-
1834
- ```typescript
1835
- const result = await runner.ExecutePrompt({
1836
- prompt: parallelPrompt, // Uses multiple models
1837
- data: analysisData,
1838
- contextUser: currentUser,
1839
-
1840
- onProgress: (progress) => {
1841
- // Parallel execution provides additional metadata
1842
- if (progress.metadata?.parallelExecution) {
1843
- const parallel = progress.metadata.parallelExecution;
1844
- console.log(`Group ${parallel.currentGroup}/${parallel.totalGroups}`);
1845
- console.log(`Tasks: ${parallel.completedTasks}/${parallel.totalTasks}`);
1846
- console.log(`Successful: ${parallel.successfulTasks}`);
1847
- }
1848
- },
1849
-
1850
- onStreaming: (chunk) => {
1851
- // Multiple models may stream simultaneously
1852
- console.log(`${chunk.modelName}: ${chunk.content}`);
1853
-
1854
- // Update model-specific UI sections
1855
- updateModelSection(chunk.taskId, chunk.content);
1856
- }
1857
- });
1858
- ```
1859
-
1860
- ### Advanced Streaming Configuration
1861
-
1862
- Fine-tune streaming behavior for optimal performance:
1863
-
1864
- ```typescript
1865
- // Streaming configuration can be applied globally or per-prompt
1866
- const streamingConfig = {
1867
- enabled: true,
1868
- aggregateParallelUpdates: false, // Separate updates per parallel task
1869
- progressUpdateIntervalMS: 250 // Limit update frequency
1870
- };
1871
-
1872
- // Progress updates are automatically throttled to prevent UI flooding
1873
- // Minimum interval between updates prevents performance issues
1874
- ```
1875
-
1876
- ### Integration with BaseLLM Streaming
1877
-
1878
- The streaming system integrates seamlessly with BaseLLM capabilities:
1879
-
1880
- ```typescript
1881
- // The AI Prompt Runner automatically detects streaming support:
1882
- // 1. Checks if the selected model supports streaming
1883
- // 2. Configures BaseLLM streaming callbacks
1884
- // 3. Aggregates streaming updates from multiple models in parallel execution
1885
- // 4. Provides unified streaming interface regardless of underlying model
1886
-
1887
- // Models that support streaming will automatically use it when callbacks are provided
1888
- // Models without streaming support will provide content in the final result
1889
- ```
1890
-
1891
- ## API Reference
1892
-
1893
- ### Exported Classes and Types
1894
-
1895
- The package exports the following public API:
1896
-
1897
- ```typescript
1898
- // Main classes
1899
- export { AIPromptCategoryEntityExtended } from './AIPromptCategoryExtended';
1900
- export { AIPromptRunner, AIPromptParams, AIPromptRunResult } from './AIPromptRunner';
1901
-
1902
- // Helper types
1903
- export { ChildPromptParam } from './AIPromptRunner';
1904
- export { SystemPlaceholder, SystemPlaceholderManager } from './SystemPlaceholders';
1905
-
1906
- // Callback types
1907
- export type { ExecutionProgressCallback, ExecutionStreamingCallback } from './AIPromptRunner';
1908
- ```
1909
-
1910
- ### Import Examples
1911
-
1912
- ```typescript
1913
- // Import from this package
1914
- import { AIPromptRunner, AIPromptParams, AIPromptRunResult } from '@memberjunction/ai-prompts';
1915
- import { ChildPromptParam, SystemPlaceholderManager } from '@memberjunction/ai-prompts';
1916
-
1917
- // Import base types from Core
1918
- import { ChatResult, ModelUsage, ChatMessage } from '@memberjunction/ai';
1919
-
1920
- // Import entities and engine types
1921
- import { AIPromptEntity } from '@memberjunction/core-entities';
1922
- import { AIEngine } from '@memberjunction/aiengine';
1923
- ```
1924
-
1925
- ### AIPromptRunner Class
1926
-
1927
- Handles execution of AI prompts with advanced parallel processing, template rendering, and result validation.
1928
-
1929
- #### Methods
1930
-
1931
- - `ExecutePrompt(params: AIPromptParams): Promise<AIPromptRunResult>`: Execute a prompt with full feature support including template rendering, model selection, parallel execution, and output validation
1932
-
1933
- #### AIPromptParams Interface
1934
-
1935
- ```typescript
1936
- interface AIPromptParams {
1937
- prompt: AIPromptEntity; // The prompt to execute
1938
- data?: any; // Template and context data
1939
- modelId?: string; // Override model selection
1940
- vendorId?: string; // Override vendor selection
1941
- configurationId?: string; // AI Configuration ID for environment-specific model selection (e.g., Prod vs Dev)
1942
- contextUser?: UserInfo; // User context
1943
- skipValidation?: boolean; // Skip output validation
1944
- templateData?: any; // Additional template data that augments the main data context
1945
- conversationMessages?: ChatMessage[]; // Multi-turn conversation messages
1946
- templateMessageRole?: TemplateMessageRole; // How to use rendered template ('system'|'user'|'none')
1947
- cancellationToken?: AbortSignal; // Cancellation token for aborting execution
1948
- onProgress?: ExecutionProgressCallback; // Progress update callback
1949
- onStreaming?: ExecutionStreamingCallback; // Streaming content callback
1950
- agentRunId?: string; // Optional agent run ID to link prompt executions to parent agent run
1951
- cleanValidationSyntax?: boolean; // Clean validation syntax from JSON responses (auto-enabled for validated prompts)
1952
- }
1953
-
1954
- /**
1955
- * Progress callback function type
1956
- */
1957
- type ExecutionProgressCallback = (progress: {
1958
- step: 'template_rendering' | 'model_selection' | 'execution' | 'validation' | 'parallel_coordination' | 'result_selection';
1959
- percentage: number; // Progress percentage (0-100)
1960
- message: string; // Human-readable status message
1961
- metadata?: Record<string, any>; // Additional metadata about the current step
1962
- }) => void;
1963
-
1964
- /**
1965
- * Streaming callback function type
1966
- */
1967
- type ExecutionStreamingCallback = (chunk: {
1968
- content: string; // The content chunk received
1969
- isComplete: boolean; // Whether this is the final chunk
1970
- taskId?: string; // Which task/model is producing this content (for parallel execution)
1971
- modelName?: string; // Model name producing this content
1972
- }) => void;
1973
-
1974
- /**
1975
- * Template message role type
1976
- */
1977
- type TemplateMessageRole = 'system' | 'user' | 'none';
1978
- ```
1979
-
1980
- ### Extended Entity Classes
1981
-
1982
- #### AIPromptCategoryEntityExtended
1983
-
1984
- Extended prompt category with prompt collection:
1985
-
1986
- ```typescript
1987
- class AIPromptCategoryEntityExtended extends AIPromptCategoryEntity {
1988
- get Prompts(): AIPromptEntity[]; // Prompts in this category
1989
- }
1990
- ```
1991
-
1992
- ### Key Interfaces and Types
1993
-
1994
- ```typescript
1995
- interface AIPromptRunResult<T = unknown> {
1996
- success: boolean; // Whether the execution was successful
1997
- status?: ExecutionStatus; // Current execution status
1998
- cancelled?: boolean; // Whether the execution was cancelled
1999
- cancellationReason?: CancellationReason; // Reason for cancellation if applicable
2000
- rawResult?: string; // The raw result from the AI model
2001
- result?: T; // The parsed/validated result based on OutputType
2002
- errorMessage?: string; // Error message if execution failed
2003
- promptRun?: AIPromptRunEntity; // The AIPromptRun entity that was created for tracking
2004
- executionTimeMS?: number; // Total execution time in milliseconds
2005
-
2006
- // Token tracking (follows ModelUsage convention)
2007
- promptTokens?: number; // Prompt/input tokens for this execution
2008
- completionTokens?: number; // Completion/output tokens for this execution
2009
- tokensUsed?: number; // Total tokens (calculated getter)
2010
-
2011
- // Hierarchical token tracking
2012
- combinedPromptTokens?: number; // Total prompt tokens including all children
2013
- combinedCompletionTokens?: number; // Total completion tokens including all children
2014
- combinedTokensUsed?: number; // Total tokens including all children (calculated)
2015
-
2016
- // Cost tracking
2017
- cost?: number; // Cost of this execution
2018
- costCurrency?: string; // ISO 4217 currency code (USD, EUR, etc.)
2019
- combinedCost?: number; // Total cost including all children
2020
-
2021
- validationResult?: ValidationResult; // Validation result if output validation was performed
2022
- validationAttempts?: ValidationAttempt[]; // Detailed validation attempts
2023
- additionalResults?: AIPromptRunResult<T>[]; // Additional results from parallel execution, ranked by judge
2024
- ranking?: number; // Ranking assigned by judge (1 = best, 2 = second best, etc.)
2025
- judgeRationale?: string; // Judge's rationale for this ranking
2026
- modelInfo?: ModelInfo; // Model information for this result
2027
- judgeMetadata?: JudgeMetadata; // Metadata about the judging process (only present on the main result)
2028
- wasStreamed?: boolean; // Whether streaming was used for this execution
2029
- cacheInfo?: { // Cache information if caching was involved
2030
- cacheHit: boolean;
2031
- cacheKey?: string;
2032
- cacheSource?: string;
2033
- };
2034
- }
2035
-
2036
- // Execution status enumeration
2037
- type ExecutionStatus = 'pending' | 'running' | 'completed' | 'failed' | 'cancelled';
2038
-
2039
- // Cancellation reason enumeration
2040
- type CancellationReason = 'user_requested' | 'timeout' | 'error' | 'resource_limit';
2041
-
2042
- // Model information interface
2043
- interface ModelInfo {
2044
- modelId: string;
2045
- modelName: string;
2046
- vendorId?: string;
2047
- vendorName?: string;
2048
- powerRank?: number;
2049
- modelType?: string;
2050
- }
2051
-
2052
- // Judge metadata interface
2053
- interface JudgeMetadata {
2054
- judgePromptId: string;
2055
- judgeExecutionTimeMS: number;
2056
- judgeTokensUsed?: number;
2057
- judgeCancelled?: boolean;
2058
- judgeErrorMessage?: string;
2059
- }
2060
-
2061
- // Parallelization strategies supported by the system
2062
- type ParallelizationStrategy = 'None' | 'StaticCount' | 'ConfigParam' | 'ModelSpecific';
2063
-
2064
- // Result selection methods for choosing best result from parallel executions
2065
- type ResultSelectionMethod = 'First' | 'Random' | 'PromptSelector' | 'Consensus';
2066
- ```
2067
-
2068
- ## Integration with Other Packages
2069
-
2070
- ### With AI Engine
2071
-
2072
- The Prompts package builds on the AI Engine for basic functionality:
2073
-
2074
- ```typescript
2075
- // AI Engine provides model management and basic operations
2076
- import { AIEngine } from '@memberjunction/aiengine';
2077
- import { AIPromptRunner } from '@memberjunction/ai-prompts';
2078
-
2079
- // Initialize AI Engine first
2080
- await AIEngine.Instance.Config(false, currentUser);
2081
-
2082
- // Access prompts managed by AI Engine
2083
- const prompts = AIEngine.Instance.Prompts;
2084
- const prompt = prompts.find(p => p.Name === 'Your Prompt');
2085
-
2086
- // Use Prompts package for advanced execution
2087
- const runner = new AIPromptRunner();
2088
- const result = await runner.ExecutePrompt({ prompt, data, contextUser });
2089
- ```
2090
-
2091
- ### With AI Agents
2092
-
2093
- AI Agents can leverage the prompt system for sophisticated operations with comprehensive execution tracking:
2094
-
2095
- ```typescript
2096
- // Agents use prompts for their intelligence with hierarchical logging
2097
- import { AgentRunner } from '@memberjunction/ai-agents';
2098
- import { AIPromptRunner } from '@memberjunction/ai-prompts';
2099
-
2100
- class IntelligentAgent extends AgentRunner {
2101
- private promptRunner = new AIPromptRunner();
2102
-
2103
- async execute(context: AgentExecutionContext): Promise<AgentExecutionResult> {
2104
- const prompt = this.getPromptForContext(context);
2105
-
2106
- // Link prompt execution to agent run for comprehensive tracking
2107
- const result = await this.promptRunner.ExecutePrompt({
2108
- prompt: prompt,
2109
- data: context.data,
2110
- contextUser: context.user,
2111
- agentRunId: context.agentRun?.ID // Links prompt to parent agent run
2112
- });
2113
-
2114
- return this.formatAgentResult(result);
2115
- }
2116
- }
2117
- ```
2118
-
2119
- #### Agent-Prompt Integration Features
2120
-
2121
- The AI Prompts system provides seamless integration with AI Agents through the `agentRunId` parameter:
2122
-
2123
- **Hierarchical Execution Tracking:**
2124
- - Prompt executions are linked to their parent agent runs via `AgentRunID` foreign key
2125
- - Provides complete audit trail from agent decision to prompt execution
2126
- - Enables comprehensive resource usage tracking across agent workflows
2127
-
2128
- **Usage Patterns:**
2129
- ```typescript
2130
- // 1. Direct agent-prompt linking
2131
- const result = await promptRunner.ExecutePrompt({
2132
- prompt: myPrompt,
2133
- data: promptData,
2134
- agentRunId: agentRun.ID, // Links to parent agent execution
2135
- contextUser: user
2136
- });
2137
-
2138
- // 2. Parallel execution with agent tracking
2139
- const parallelResult = await promptRunner.ExecutePrompt({
2140
- prompt: parallelPrompt, // ParallelizationMode: 'ModelSpecific'
2141
- data: analysisData,
2142
- agentRunId: agentRun.ID, // All parallel child prompts link to agent
2143
- contextUser: user
2144
- });
2145
-
2146
- // 3. Context compression with agent linking (automatic in AgentRunner)
2147
- // When agents use context compression, compression prompts are automatically
2148
- // linked to the parent agent run for complete execution visibility
2149
- ```
2150
-
2151
- **Database Schema Integration:**
2152
- ```sql
2153
- -- Query agent execution with all related prompts
2154
- SELECT
2155
- ar.ID as AgentRunID,
2156
- ar.Status as AgentStatus,
2157
- ar.StartedAt,
2158
- ar.CompletedAt,
2159
- pr.ID as PromptRunID,
2160
- pr.RunType,
2161
- pr.Success as PromptSuccess,
2162
- pr.ExecutionTimeMS,
2163
- pr.TokensUsed
2164
- FROM AIAgentRun ar
2165
- LEFT JOIN AIPromptRun pr ON ar.ID = pr.AgentRunID
2166
- WHERE ar.ID = 'your-agent-run-id'
2167
- ORDER BY pr.RunAt;
2168
- ```
2169
-
2170
- **Benefits:**
2171
- - **Complete Traceability**: Track all AI model usage from agent decisions to prompt executions
2172
- - **Resource Attribution**: Understand token usage and costs at the agent level
2173
- - **Performance Analysis**: Analyze execution patterns across the agent-prompt hierarchy
2174
- - **Debugging Support**: Full execution history for troubleshooting agent workflows
2175
-
2176
235
  ## Dependencies
2177
236
 
2178
- - `@memberjunction/core` (v2.43.0): MemberJunction core library
2179
- - `@memberjunction/global` (v2.43.0): MemberJunction global utilities
2180
- - `@memberjunction/core-entities` (v2.43.0): MemberJunction entity definitions
2181
- - `@memberjunction/ai` (v2.43.0): Base AI types and result structures (imported for core types)
2182
- - `@memberjunction/aiengine` (v2.43.0): AI model management and basic operations
2183
- - `@memberjunction/templates` (v2.43.0): Template rendering system
2184
- - `dotenv` (^16.4.1): Environment variable management
2185
- - `rxjs` (^7.8.1): Reactive programming support
2186
-
2187
- ## Related Packages
2188
-
2189
- - `@memberjunction/aiengine`: Core AI engine and model management
2190
- - `@memberjunction/ai-agents`: Advanced agent framework built on prompts
2191
- - `@memberjunction/templates`: Template rendering for dynamic content
2192
-
2193
- ## Migration Guide
2194
-
2195
- ### From AI Engine Simple Completions
2196
-
2197
- For cases requiring more sophisticated prompt management:
2198
-
2199
- ```typescript
2200
- // Old: Simple LLM completion (still valid for basic cases)
2201
- const response = await AIEngine.Instance.SimpleLLMCompletion(
2202
- "Analyze this data",
2203
- currentUser,
2204
- "You are a data analyst"
2205
- );
2206
-
2207
- // New: Advanced prompt with caching, validation, and parallel execution
2208
- const prompt = {
2209
- Name: "Data Analysis",
2210
- PromptText: "Analyze this data: {{data}}",
2211
- EnableCaching: true,
2212
- ParallelizationMode: "StaticCount",
2213
- ParallelCount: 2,
2214
- OutputType: "object",
2215
- ValidationBehavior: "Strict"
2216
- };
2217
-
2218
- const runner = new AIPromptRunner();
2219
- const result = await runner.ExecutePrompt({
2220
- prompt: prompt,
2221
- data: { data: "your data here" },
2222
- contextUser: currentUser
2223
- });
2224
- ```
2225
-
2226
- ## Best Practices
2227
-
2228
- 1. **Enable Caching**: Use intelligent caching for expensive operations
2229
- 2. **Validate Outputs**: Always specify expected output types for critical operations
2230
- 3. **Use Templates**: Leverage template system for dynamic prompts
2231
- 4. **Monitor Performance**: Track token usage and execution times
2232
- 5. **Parallel Wisely**: Use parallel execution for independent tasks, not dependent ones
2233
- 6. **Handle Errors**: Implement proper retry logic and error handling
2234
- 7. **Implement Cancellation**: Always provide cancellation tokens for user-facing operations
2235
- 8. **Use Progress Callbacks**: Provide progress feedback for long-running operations
2236
- 9. **Leverage Hierarchical Logging**: Use the logging hierarchy for debugging and analytics
2237
- 10. **Configure Streaming Appropriately**: Enable streaming for responsive user experiences
2238
- 11. **Optimize Judge Selection**: Use efficient judge prompts for parallel result selection
2239
- 12. **Monitor Resource Usage**: Track token consumption and execution times across hierarchical runs
2240
-
2241
- ### Implementation Guidelines
2242
-
2243
- ```typescript
2244
- // Comprehensive prompt execution with all new features
2245
- const controller = new AbortController();
2246
-
2247
- const result = await runner.ExecutePrompt({
2248
- prompt: myPrompt,
2249
- data: executionData,
2250
- contextUser: currentUser,
2251
-
2252
- // Cancellation support
2253
- cancellationToken: controller.signal,
2254
-
2255
- // Progress tracking
2256
- onProgress: (progress) => {
2257
- updateProgressIndicator(progress.percentage, progress.message);
2258
- if (progress.metadata?.parallelExecution) {
2259
- updateParallelStatus(progress.metadata.parallelExecution);
2260
- }
2261
- },
2262
-
2263
- // Streaming for real-time updates
2264
- onStreaming: (chunk) => {
2265
- if (chunk.isComplete) {
2266
- finalizePage();
2267
- } else {
2268
- appendContent(chunk.content);
2269
- }
2270
- }
2271
- });
2272
-
2273
- // Always check for cancellation in results
2274
- if (result.cancelled) {
2275
- handleCancellation(result.cancellationReason);
2276
- } else if (result.success) {
2277
- processResults(result);
2278
-
2279
- // Analyze additional results from parallel execution
2280
- if (result.additionalResults) {
2281
- analyzeAlternativeResults(result.additionalResults);
2282
- }
2283
- }
2284
-
2285
- // Use hierarchical logging data for analytics
2286
- if (result.promptRun) {
2287
- trackExecutionMetrics(result.promptRun);
2288
- if (result.promptRun.RunType === 'ParallelParent') {
2289
- analyzeParallelPerformance(result.promptRun.ID);
2290
- }
2291
- }
2292
- ```
2293
-
2294
- ## Troubleshooting
2295
-
2296
- ### Common Issues
2297
-
2298
- 1. **"No suitable model found" Error**
2299
- - Ensure AIEngine.Instance.Config() is called before using prompts
2300
- - Verify prompt has active AIPromptModel associations or proper model selection configuration
2301
- - Check that models meet MinPowerRank requirements
2302
-
2303
- 2. **Template Rendering Failures**
2304
- - Verify template exists and is associated with the prompt
2305
- - Ensure template data contains all required variables
2306
- - Check template syntax for Handlebars errors
2307
-
2308
- 3. **Parallel Execution Not Working**
2309
- - Confirm ParallelizationMode is set to a value other than 'None'
2310
- - For ModelSpecific mode, ensure AIPromptModel entries exist
2311
- - Check that multiple suitable models are available
2312
-
2313
- 4. **Output Validation Errors**
2314
- - Ensure OutputType matches the expected result format
2315
- - Provide a valid OutputExample for structured data
2316
- - Consider increasing MaxRetries for complex outputs
2317
-
2318
- 5. **Cancellation Not Working**
2319
- - Verify the AbortController is properly created and signal is passed
2320
- - Check that the cancellation token is not already aborted before execution
2321
- - Ensure model implementations support cancellation (older models may not)
2322
- - Review cancellation timing - very fast executions may complete before cancellation
2323
-
2324
- 6. **Progress Updates Not Received**
2325
- - Confirm onProgress callback is properly defined and passed to ExecutePrompt
2326
- - Check that the callback function doesn't throw errors (which can stop updates)
2327
- - Progress updates are throttled - very fast operations may have fewer updates
2328
- - Parallel execution provides more detailed progress metadata
2329
-
2330
- 7. **Streaming Not Working**
2331
- - Verify the selected AI model supports streaming (not all models do)
2332
- - Ensure onStreaming callback is provided in AIPromptParams
2333
- - Check BaseLLM implementation supports streaming for the specific model
2334
- - Review model configuration - some vendors require specific settings for streaming
2335
-
2336
- 8. **Hierarchical Logging Missing**
2337
- - Ensure database schema includes RunType, ParentID, and ExecutionOrder fields
2338
- - Check that user has permissions to create AIPromptRun records
2339
- - Verify prompt run creation isn't being skipped due to errors
2340
- - Review logs for save failures on prompt run entities
2341
-
2342
- 9. **Judge Selection Failing**
2343
- - Confirm ResultSelectorPromptID is set and points to a valid, active prompt
2344
- - Verify the judge prompt returns valid JSON with rankings array
2345
- - Check that judge prompt has proper model associations
2346
- - Review judge prompt timeout settings for complex evaluations
2347
-
2348
- ### Performance Optimization
2349
-
2350
- For optimal performance with the new features:
2351
-
2352
- ```typescript
2353
- // Minimize progress update frequency for high-performance scenarios
2354
- const result = await runner.ExecutePrompt({
2355
- prompt: myPrompt,
2356
- data: myData,
2357
- onProgress: (progress) => {
2358
- // Throttle UI updates
2359
- if (progress.percentage % 10 === 0) {
2360
- updateUI(progress);
2361
- }
2362
- }
2363
- });
2364
-
2365
- // Use cancellation for long-running operations
2366
- const controller = new AbortController();
2367
- setTimeout(() => controller.abort(), 60000); // 1 minute timeout
2368
-
2369
- // Configure parallel execution for optimal throughput
2370
- const parallelPrompt = {
2371
- ParallelizationMode: "ModelSpecific",
2372
- // Configure specific models with different execution groups for coordination
2373
- };
2374
- ```
2375
-
2376
- ### Debugging Hierarchical Logs
2377
-
2378
- Use these queries to troubleshoot execution issues:
2379
-
2380
- ```sql
2381
- -- Find incomplete executions
2382
- SELECT * FROM AIPromptRun
2383
- WHERE CompletedAt IS NULL
2384
- AND RunAt < DATEADD(minute, -5, GETDATE());
2385
-
2386
- -- Check parallel execution hierarchy
2387
- SELECT
2388
- ID, RunType, ParentID, ExecutionOrder, Success, ErrorMessage
2389
- FROM AIPromptRun
2390
- WHERE ParentID = 'your-parent-id' OR ID = 'your-parent-id'
2391
- ORDER BY RunType, ExecutionOrder;
2392
-
2393
- -- Find resource usage patterns
2394
- SELECT
2395
- RunType,
2396
- AVG(ExecutionTimeMS) as AvgTimeMS,
2397
- AVG(TokensUsed) as AvgTokens,
2398
- COUNT(*) as ExecutionCount
2399
- FROM AIPromptRun
2400
- WHERE RunAt > DATEADD(day, -7, GETDATE())
2401
- GROUP BY RunType;
2402
- ```
2403
-
2404
- ## License
2405
-
2406
- ISC
2407
-
2408
- ---
2409
-
2410
- ## Advanced Configuration
2411
-
2412
- ### Cache Configuration
2413
-
2414
- ```typescript
2415
- const cacheOptimizedPrompt = {
2416
- EnableCaching: true,
2417
- CacheMatchType: "Vector", // Vector similarity matching
2418
- CacheTTLSeconds: 3600, // 1 hour cache
2419
- CacheSimilarityThreshold: 0.9, // High similarity required
2420
- CacheMustMatchModel: true, // Model must match
2421
- CacheMustMatchVendor: false, // Vendor can differ
2422
- CacheMustMatchAgent: false, // Agent can differ
2423
- CacheMustMatchConfig: true // Configuration must match
2424
- };
2425
- ```
2426
-
2427
- ### Model Selection Strategies
2428
-
2429
- ```typescript
2430
- // By power ranking
2431
- const powerBasedPrompt = {
2432
- SelectionStrategy: "ByPower",
2433
- PowerPreference: "Highest", // or "Lowest"
2434
- MinPowerRank: 80 // Minimum capability required
2435
- };
2436
-
2437
- // Specific models
2438
- const specificModelsPrompt = {
2439
- SelectionStrategy: "Specific",
2440
- // Models defined in AIPromptModel entries
2441
- };
2442
-
2443
- // Default system selection
2444
- const defaultPrompt = {
2445
- SelectionStrategy: "Default"
2446
- };
2447
- ```
2448
-
2449
- ### Model Priority Behavior
2450
-
2451
- When multiple models are configured for a prompt (using `AIPromptModel` entries), the execution engine uses the `Priority` field to determine the order in which models are tried:
2452
-
2453
- - **Higher priority numbers are executed first** (e.g., Priority 10 before Priority 5)
2454
- - The engine sorts models by `Priority DESC` to establish execution order
2455
- - For models with the same priority, creation date is used as a tiebreaker
2456
- - The first successful model execution is used (fail-fast approach)
2457
-
2458
- ```typescript
2459
- // Example: Models are tried in this order based on Priority
2460
- const promptModels = [
2461
- { ModelID: 'gpt-4-id', Priority: 100 }, // Tried first
2462
- { ModelID: 'claude-3-id', Priority: 50 }, // Tried second
2463
- { ModelID: 'gpt-3.5-id', Priority: 10 } // Tried third
2464
- ];
2465
-
2466
- // The PromptRunner sorts internally using:
2467
- // models.sort((a, b) => b.Priority - a.Priority)
2468
- ```
2469
-
2470
- This priority system allows you to:
2471
- - Set preferred models with higher priorities
2472
- - Configure fallback models with lower priorities
2473
- - Ensure expensive/powerful models are only used when necessary
2474
- - Control the exact execution order for cost optimization
2475
-
2476
- ## Multi-Vendor Model Support
2477
-
2478
- MemberJunction supports multiple inference providers (vendors) for the same AI model, enabling flexible deployment scenarios and vendor failover. This is crucial for:
2479
- - Using the same model from different providers (e.g., Claude from Anthropic API vs AWS Bedrock)
2480
- - Vendor-specific configurations (different token limits, API endpoints, pricing)
2481
- - Failover strategies when primary vendors are unavailable
2482
- - Cost optimization by routing to cheaper providers
2483
-
2484
- ### Vendor Selection Precedence
2485
-
2486
- The AI Prompt Runner uses a sophisticated vendor selection system with clear precedence rules:
2487
-
2488
- 1. **Explicit Runtime Override** (`params.override.vendorId`)
2489
- - Highest precedence - always used when specified
2490
- - If the vendor doesn't provide the selected model, a warning is issued and fallback occurs
2491
-
2492
- 2. **Prompt-Model Association Vendor** (`AIPromptModel.VendorID`)
2493
- - When a prompt has specific model associations, the vendor from the highest priority association is used
2494
- - Configured in the database for prompt-specific vendor preferences
2495
-
2496
- 3. **Highest Priority Vendor** (`AIModelVendor.Priority`)
2497
- - When no explicit vendor is specified, the system uses the vendor with the highest priority value
2498
- - Priority is configured per model-vendor combination in the database
2499
- - **Higher Priority numbers = Higher preference** (e.g., Priority 100 is preferred over Priority 50)
2500
-
2501
- ### Vendor-Specific Configuration
2502
-
2503
- Each vendor can have different configurations for the same model:
2504
-
2505
- ```typescript
2506
- // AIModelVendor entity fields used during execution:
2507
- {
2508
- ModelID: 'claude-3-opus-id',
2509
- VendorID: 'anthropic-id',
2510
- Priority: 100, // Higher = preferred
2511
- DriverClass: 'AnthropicLLM', // Vendor-specific implementation
2512
- APIName: 'claude-3-opus-20240229', // Vendor's API model name
2513
- MaxInputTokens: 200000, // Vendor-specific limits
2514
- MaxOutputTokens: 4096,
2515
- SupportsStreaming: true,
2516
- SupportsEffortLevel: false
2517
- }
2518
- ```
2519
-
2520
- ### Vendor Override Example
2521
-
2522
- ```typescript
2523
- // Execute with specific vendor
2524
- const result = await runner.ExecutePrompt({
2525
- prompt: myPrompt,
2526
- data: { query: 'Analyze this data' },
2527
- contextUser: currentUser,
2528
- override: {
2529
- modelId: 'claude-3-opus-id',
2530
- vendorId: 'aws-bedrock-id' // Use AWS Bedrock instead of default
2531
- }
2532
- });
2533
-
2534
- // The system will:
2535
- // 1. Validate that AWS Bedrock provides Claude 3 Opus
2536
- // 2. Use AWS Bedrock's driver class and configuration
2537
- // 3. Fall back to highest priority vendor if mismatch occurs
2538
- ```
2539
-
2540
- ### Model-Vendor Mismatch Handling
2541
-
2542
- When a specified vendor doesn't provide the requested model:
2543
-
2544
- ```typescript
2545
- // Scenario: User requests GPT-4 from Anthropic (invalid combination)
2546
- const result = await runner.ExecutePrompt({
2547
- prompt: myPrompt,
2548
- override: {
2549
- modelId: 'gpt-4-id',
2550
- vendorId: 'anthropic-id' // Anthropic doesn't provide GPT-4
2551
- }
2552
- });
2553
-
2554
- // System behavior:
2555
- // 1. Logs warning: "⚠️ Warning: Vendor anthropic does not provide model GPT-4. Falling back to highest priority vendor."
2556
- // 2. Finds highest priority vendor for GPT-4 (e.g., OpenAI)
2557
- // 3. Uses the fallback vendor's configuration
2558
- // 4. Execution continues with valid vendor
2559
- ```
2560
-
2561
- ### Database Configuration
2562
-
2563
- Configure multi-vendor support through the MemberJunction metadata:
2564
-
2565
- ```sql
2566
- -- Example: Claude available from multiple vendors
2567
- INSERT INTO [MJ: AI Model Vendors] (ModelID, VendorID, Priority, DriverClass, APIName, MaxInputTokens)
2568
- VALUES
2569
- ('claude-3-id', 'anthropic-id', 100, 'AnthropicLLM', 'claude-3-opus-20240229', 200000),
2570
- ('claude-3-id', 'aws-bedrock-id', 90, 'BedrockLLM', 'anthropic.claude-3-opus', 180000),
2571
- ('claude-3-id', 'vertex-ai-id', 80, 'VertexAILLM', 'claude-3-opus@001', 150000);
2572
-
2573
- -- Query vendor options for a model
2574
- SELECT
2575
- mv.Priority,
2576
- v.Name as VendorName,
2577
- mv.DriverClass,
2578
- mv.APIName,
2579
- mv.MaxInputTokens,
2580
- mv.MaxOutputTokens
2581
- FROM [MJ: AI Model Vendors] mv
2582
- JOIN [MJ: AI Vendors] v ON mv.VendorID = v.ID
2583
- WHERE mv.ModelID = 'claude-3-id'
2584
- AND mv.Status = 'Active'
2585
- ORDER BY mv.Priority DESC;
2586
- ```
2587
-
2588
- ### Benefits of Multi-Vendor Support
2589
-
2590
- 1. **Resilience**: Automatic failover when primary vendor is unavailable
2591
- 2. **Cost Optimization**: Route to cheaper vendors for non-critical tasks
2592
- 3. **Regional Compliance**: Use region-specific vendors for data residency
2593
- 4. **Performance**: Choose vendors with better latency for your location
2594
- 5. **Feature Access**: Some vendors may offer unique features (streaming, tools, etc.)
2595
-
2596
- ### Implementation Details
2597
-
2598
- The multi-vendor support is implemented in the `executeModel` method:
2599
-
2600
- 1. **Vendor Resolution**: Determines which vendor to use based on precedence rules
2601
- 2. **Configuration Loading**: Loads vendor-specific settings from AIModelVendor
2602
- 3. **Driver Selection**: Uses vendor-specific driver class for API communication
2603
- 4. **Fallback Logic**: Handles mismatches gracefully with warnings
2604
- 5. **Execution Tracking**: Records the actual vendor used in AIPromptRun
2605
-
2606
- For additional configuration options and advanced use cases, refer to the source code and entity definitions in the MemberJunction core system.
2607
-
2608
- ## System Prompt Embedding
2609
-
2610
- The AI Prompt Runner provides sophisticated system prompt embedding capabilities for agent architectures through the template engine integration.
2611
-
2612
- ### Architecture Overview
2613
-
2614
- When `systemPromptId` is provided in AIPromptParams, the runner:
2615
- 1. Loads the system prompt template from the database
2616
- 2. Embeds the agent-specific AI prompt using `{% PromptEmbed %}` syntax
2617
- 3. Renders the complete system prompt with agent context
2618
- 4. Uses the rendered system prompt instead of the regular AI prompt template
2619
-
2620
- This enables sophisticated agent architectures where:
2621
- - **System prompts** provide execution control and enforce deterministic JSON response format
2622
- - **Agent prompts** contain domain-specific logic (e.g., DATA_GATHER instructions)
2623
- - **Available actions and sub-agents** are injected for agent decision-making
2624
-
2625
- ### Template Syntax
2626
-
2627
- System prompt templates use the `{% PromptEmbed %}` syntax to embed AI prompts:
2628
-
2629
- ```nunjucks
2630
- # System Prompt Template Example
2631
-
2632
- You are an AI agent with the following specialized instructions:
2633
-
2634
- {% PromptEmbed %}
2635
-
2636
- ## Available Actions
2637
- {{#each availableActions}}
2638
- - **{{this.name}}**: {{this.description}}
2639
- {{/each}}
2640
-
2641
- ## Available Sub-Agents
2642
- {{#each availableSubAgents}}
2643
- - **{{this.name}}**: {{this.description}}
2644
- {{/each}}
2645
-
2646
- ## Response Format
2647
- You must respond with valid JSON following this structure:
2648
- {
2649
- "decision": "execute_action|execute_subagent|complete_task|request_clarification",
2650
- "reasoning": "Explanation of your decision",
2651
- "executionPlan": [
2652
- {
2653
- "type": "action|subagent",
2654
- "targetId": "action-or-agent-id",
2655
- "parameters": {},
2656
- "executionOrder": 1,
2657
- "allowParallel": true
2658
- }
2659
- ],
2660
- "isTaskComplete": false,
2661
- "confidence": 0.95
2662
- }
2663
- ```
2664
-
2665
- ### Validation and Security
2666
-
2667
- The system includes comprehensive validation to ensure proper prompt embedding:
2668
-
2669
- ```typescript
2670
- // Validation process:
2671
- // 1. Verify system prompt exists and has template
2672
- // 2. Check agent-prompt relationships via AIAgentPrompt table
2673
- // 3. Ensure agents using system prompt are linked to current prompt
2674
- // 4. Validate template contains {% PromptEmbed %} syntax
2675
-
2676
- const params = new AIPromptParams();
2677
- params.prompt = agentSpecificPrompt;
2678
- params.systemPromptId = 'system-prompt-id'; // Triggers validation
2679
- params.data = { agentName: 'DataGather', availableActions: [...] };
2680
- ```
2681
-
2682
- ### Integration with AI Agents
2683
-
2684
- The AgentRunner seamlessly uses system prompt embedding:
2685
-
2686
- ```typescript
2687
- // AgentRunner delegates to AIPromptRunner with system prompt embedding
2688
- const promptParams = new AIPromptParams();
2689
- promptParams.prompt = primaryAgentPrompt.prompt;
2690
- promptParams.systemPromptId = this.agentType.SystemPromptID;
2691
- promptParams.data = promptData;
2692
- promptParams.agentRunId = context.agentRun.ID;
2693
-
2694
- const promptResult = await this._promptRunner.ExecutePrompt(promptParams);
2695
- ```
2696
-
2697
- ### Database Schema Integration
2698
-
2699
- The system prompt embedding feature integrates with several database entities:
2700
-
2701
- #### Entity Relationships
2702
-
2703
- ```sql
2704
- -- System prompts are stored as AIPrompt entities with templates
2705
- AIPrompt (SystemPromptID) -> Template -> TemplateContent (contains {% PromptEmbed %})
2706
-
2707
- -- Agent types reference system prompts
2708
- AIAgentType.SystemPromptID -> AIPrompt (system prompt)
2709
-
2710
- -- Agents belong to agent types
2711
- AIAgent.TypeID -> AIAgentType
2712
-
2713
- -- Agent prompts link agents to their specific prompts
2714
- AIAgentPrompt: AgentID + PromptID
2715
-
2716
- -- Validation ensures proper linkage:
2717
- -- Agent -> AgentType -> SystemPrompt
2718
- -- Agent -> AIAgentPrompt -> AIPrompt (to be embedded)
2719
- ```
2720
-
2721
- #### Storage Structure
2722
-
2723
- ```sql
2724
- -- Example system prompt template storage
2725
- INSERT INTO Template (Name, Description)
2726
- VALUES ('AI Agent System Prompt', 'Control wrapper for agent decision-making');
2727
-
2728
- INSERT INTO TemplateContent (TemplateID, TemplateText, Priority)
2729
- VALUES (@TemplateID, 'You are {{agentName}}... {% PromptEmbed %}... Respond with JSON...', 100);
2730
-
2731
- INSERT INTO AIPrompt (Name, Description, TemplateID, Category)
2732
- VALUES ('System Prompt', 'Agent execution control wrapper', @TemplateID, 'System');
2733
-
2734
- UPDATE AIAgentType SET SystemPromptID = @SystemPromptID WHERE Name = 'DataGatherAgent';
2735
- ```
2736
-
2737
- ## API Keys
2738
-
2739
- The AI Prompts system provides flexible API key management for runtime configuration without modifying environment variables or global settings.
2740
-
2741
- ### Environment-Based API Configuration
2742
-
2743
- MemberJunction now supports sophisticated environment-based API key resolution through AIConfigSet entities, allowing different configurations for different environments:
2744
-
2745
- ```typescript
2746
- // Configurations are loaded based on the environment name
2747
- // Default fallback order: process.env.NODE_ENV -> 'production'
2748
-
2749
- // Example: Different API keys per environment
2750
- // Development -> AIConfigSet(Name='development') -> Config Key OPENAI_LLM_APIKEY
2751
- // Production -> AIConfigSet(Name='production') -> Config Key OPENAI_LLM_APIKEY
2752
- ```
2753
-
2754
- #### Configuration Set Priority
2755
-
2756
- When multiple configuration sets exist, they are evaluated in this order:
2757
- 1. **Exact environment match** (e.g., 'development' matches 'development')
2758
- 2. **Priority field** (higher priority values are preferred)
2759
- 3. **Custom resolver logic** (if implemented via subclassing)
2760
-
2761
- ### Using Runtime API Keys
2762
-
2763
- You can provide API keys at prompt execution time, which is useful for:
2764
- - Multi-tenant applications where different users have different API keys
2765
- - Testing with different API providers or accounts
2766
- - Isolating API usage by application or department
2767
- - Temporary API key usage for specific operations
2768
-
2769
- ```typescript
2770
- import { AIPromptRunner, AIPromptParams } from '@memberjunction/ai-prompts';
2771
- import { AIAPIKey } from '@memberjunction/ai';
2772
-
2773
- const runner = new AIPromptRunner();
2774
-
2775
- // Execute with specific API keys
2776
- const result = await runner.ExecutePrompt({
2777
- prompt: myPrompt,
2778
- data: { query: 'Analyze this data' },
2779
- contextUser: currentUser,
2780
- apiKeys: [
2781
- { driverClass: 'OpenAILLM', apiKey: 'sk-user-specific-key' },
2782
- { driverClass: 'AnthropicLLM', apiKey: 'sk-ant-department-key' }
2783
- ]
2784
- });
2785
- ```
2786
-
2787
- ### API Key Precedence
2788
-
2789
- When executing prompts, API keys are resolved in this order:
2790
- 1. **Local API keys** provided in `AIPromptParams.apiKeys` (highest priority)
2791
- 2. **Configuration sets** from database based on environment
2792
- 3. **Environment variables** (traditional dotenv approach)
2793
- 4. **Custom implementations** via AIAPIKeys subclassing
2794
-
2795
- ### Configuration Set Structure
2796
-
2797
- Configuration sets are managed through these entities:
2798
-
2799
- #### AIConfigSet
2800
- - **Name**: Environment name (e.g., 'development', 'production', 'staging')
2801
- - **Description**: Human-readable description
2802
- - **Priority**: Higher values take precedence when multiple sets match
2803
- - **Status**: 'Active' or 'Inactive'
2804
-
2805
- #### AIConfiguration
2806
- - **ConfigSetID**: Links to parent configuration set
2807
- - **ConfigKey**: The configuration key (e.g., 'OPENAI_LLM_APIKEY')
2808
- - **ConfigValue**: The actual value (encrypted for sensitive data)
2809
- - **EncryptedValue**: Whether the value is encrypted
2810
- - **Description**: Documentation for the configuration
2811
-
2812
- ### Environment Variables vs Configuration Sets
2813
-
2814
- ```typescript
2815
- // Traditional environment variable approach (still supported)
2816
- process.env.OPENAI_LLM_APIKEY = 'sk-...';
2817
-
2818
- // New configuration set approach (recommended)
2819
- // Stored in database:
2820
- // ConfigSet: { Name: 'production', Priority: 100 }
2821
- // Config: { ConfigKey: 'OPENAI_LLM_APIKEY', ConfigValue: 'sk-...', Encrypted: true }
2822
-
2823
- // The system automatically uses the configuration set if available
2824
- // Falls back to environment variables if not found in database
2825
- ```
2826
-
2827
- ### Hierarchical API Key Propagation
2828
-
2829
- For hierarchical prompt execution (prompts with child prompts), API keys are automatically propagated:
2830
-
2831
- ```typescript
2832
- const parentParams = new AIPromptParams();
2833
- parentParams.prompt = parentPrompt;
2834
- parentParams.childPrompts = [
2835
- new ChildPromptParam(childPrompt1, 'analysis'),
2836
- new ChildPromptParam(childPrompt2, 'summary')
2837
- ];
2838
- parentParams.apiKeys = [
2839
- { driverClass: 'OpenAILLM', apiKey: 'sk-parent-key' }
2840
- ];
2841
-
2842
- // Child prompts automatically inherit the parent's API keys
2843
- // This ensures consistent API key usage throughout the execution tree
2844
- const result = await runner.ExecutePrompt(parentParams);
2845
- ```
2846
-
2847
- ### Custom Global API Key Management
2848
-
2849
- For advanced scenarios, you can subclass the global `AIAPIKeys` object:
2850
-
2851
- ```typescript
2852
- import { AIAPIKeys, RegisterClass } from '@memberjunction/ai';
2853
-
2854
- @RegisterClass(AIAPIKeys, 'CustomAPIKeys', 2) // Priority 2 overrides default
2855
- export class CustomAPIKeys extends AIAPIKeys {
2856
- public GetAPIKey(AIDriverName: string): string {
2857
- // First check configuration sets
2858
- const configValue = this.getFromConfigSet(AIDriverName);
2859
- if (configValue) return configValue;
2860
-
2861
- // Then check custom logic: database lookup, vault access, etc.
2862
- if (AIDriverName === 'OpenAILLM') {
2863
- return this.getFromVault('openai-key');
2864
- }
2865
-
2866
- // Finally fall back to environment variables
2867
- return super.GetAPIKey(AIDriverName);
2868
- }
2869
-
2870
- private getFromConfigSet(driverName: string): string | null {
2871
- // Implementation would query AIConfiguration entities
2872
- // based on current environment
2873
- const envName = process.env.NODE_ENV || 'production';
2874
- // ... query logic here
2875
- return null;
2876
- }
2877
- }
2878
- ```
2879
-
2880
- ### Multi-Environment Setup Example
2881
-
2882
- ```typescript
2883
- // Development environment setup
2884
- const devConfig = {
2885
- ConfigSet: { Name: 'development', Priority: 100 },
2886
- Configurations: [
2887
- { ConfigKey: 'OPENAI_LLM_APIKEY', ConfigValue: 'sk-dev-...', Encrypted: true },
2888
- { ConfigKey: 'ANTHROPIC_LLM_APIKEY', ConfigValue: 'sk-ant-dev-...', Encrypted: true },
2889
- { ConfigKey: 'LOG_LEVEL', ConfigValue: 'debug', Encrypted: false }
2890
- ]
2891
- };
2892
-
2893
- // Production environment setup
2894
- const prodConfig = {
2895
- ConfigSet: { Name: 'production', Priority: 100 },
2896
- Configurations: [
2897
- { ConfigKey: 'OPENAI_LLM_APIKEY', ConfigValue: 'sk-prod-...', Encrypted: true },
2898
- { ConfigKey: 'ANTHROPIC_LLM_APIKEY', ConfigValue: 'sk-ant-prod-...', Encrypted: true },
2899
- { ConfigKey: 'LOG_LEVEL', ConfigValue: 'error', Encrypted: false }
2900
- ]
2901
- };
2902
-
2903
- // The system automatically loads the correct configuration based on NODE_ENV
2904
- ```
2905
-
2906
- ### Security Best Practices
2907
-
2908
- 1. **Never hardcode API keys** in your source code
2909
- 2. **Use encrypted storage** for sensitive configuration values
2910
- 3. **Separate configurations** by environment (dev, staging, prod)
2911
- 4. **Rotate keys regularly** and update configuration sets
2912
- 5. **Monitor API key usage** to detect unauthorized access
2913
- 6. **Use the principle of least privilege** - give each user/app only the keys they need
2914
- 7. **Audit configuration changes** through MemberJunction's change tracking
2915
-
2916
- ## AI Configuration System
2917
-
2918
- The AI Prompts system supports sophisticated environment-specific model selection through AI Configurations. This allows you to use different sets of models for the same prompts based on the active configuration (e.g., Production vs Development).
2919
-
2920
- ### Configuration Concepts
2921
-
2922
- #### AI Configurations
2923
- - Named configuration sets (e.g., "Production", "Development", "Europe-Region")
2924
- - Filter which models are available for prompt execution
2925
- - Support configuration parameters for dynamic behavior
2926
- - One configuration can be marked as default (`IsDefault = true`)
2927
-
2928
- #### AI Configuration Parameters
2929
- - Name-value pairs stored per configuration
2930
- - Support different data types (string, number, boolean, date, object)
2931
- - Can control dynamic parallelization and other runtime behavior
2932
- - Example: `ParallelExecutions = 5` for development testing
2933
-
2934
- ### Model Selection with Configurations
2935
-
2936
- When executing a prompt with a specific `configurationId`, the system follows a two-phase model selection process that ensures configuration-specific models ALWAYS take precedence:
2937
-
2938
- #### Phase 1: Configuration-Specific Models (Highest Priority)
2939
- ```typescript
2940
- // First, try to find models with matching configuration
2941
- const configModels = promptModels.filter(pm =>
2942
- pm.ConfigurationID === configurationId &&
2943
- (pm.Status === 'Active' || pm.Status === 'Preview')
2944
- );
2945
- ```
2946
-
2947
- If configuration-specific models are found, they are used exclusively. The system will NOT consider models with NULL or mismatched configurations.
2948
-
2949
- #### Phase 2: Default Models (Fallback)
2950
-
2951
- Only if NO models match the specific configuration, the system falls back to models with NULL configuration:
2952
-
2953
- ```typescript
2954
- // Fall back to NULL configuration models only if no matches found
2955
- if (configModels.length === 0) {
2956
- const defaultModels = promptModels.filter(pm =>
2957
- pm.ConfigurationID === null &&
2958
- (pm.Status === 'Active' || pm.Status === 'Preview')
2959
- );
2960
- }
2961
- ```
2962
-
2963
- This ensures:
2964
- - **Configuration-specific models always win**: When you specify a configuration, only models explicitly assigned to that configuration are considered first
2965
- - **Clear environment separation**: Production models never mix with Development models when configurations are used
2966
- - **Explicit fallback behavior**: NULL configuration models serve as defaults only when no configuration-specific models exist
2967
-
2968
- ### Configuration Setup Example
2969
-
2970
- ```sql
2971
- -- Create configurations
2972
- INSERT INTO AIConfiguration (Name, Description, IsDefault, Status)
2973
- VALUES
2974
- ('Production', 'Stable models for production use', 1, 'Active'),
2975
- ('Development', 'Experimental models for testing', 0, 'Active');
2976
-
2977
- -- Assign models to configurations
2978
- -- Production uses GPT-4
2979
- UPDATE AIPromptModel
2980
- SET ConfigurationID = (SELECT ID FROM AIConfiguration WHERE Name = 'Production')
2981
- WHERE PromptID = @PromptID AND ModelID = @GPT4ModelID;
2982
-
2983
- -- Development uses GPT-4-Turbo and Claude-3-Opus
2984
- UPDATE AIPromptModel
2985
- SET ConfigurationID = (SELECT ID FROM AIConfiguration WHERE Name = 'Development')
2986
- WHERE PromptID = @PromptID AND ModelID IN (@GPT4TurboID, @Claude3OpusID);
2987
-
2988
- -- Models with NULL ConfigurationID serve as defaults when no configuration matches
2989
- UPDATE AIPromptModel
2990
- SET ConfigurationID = NULL
2991
- WHERE PromptID = @PromptID AND ModelID = @FallbackModelID;
2992
- ```
2993
-
2994
- ### Using Configurations
2995
-
2996
- #### In Code
2997
- ```typescript
2998
- const result = await promptRunner.ExecutePrompt({
2999
- prompt: myPrompt,
3000
- configurationId: 'dev-config-id', // Optional
3001
- data: { query: 'Analyze this data' },
3002
- contextUser: currentUser
3003
- });
3004
- ```
3005
-
3006
- #### Configuration Precedence
3007
- 1. **Configuration-specific models** (highest priority) - When a configurationId is provided, ONLY models with matching ConfigurationID are considered initially
3008
- 2. **NULL configuration models** (fallback only) - Used only when NO models match the specified configuration
3009
- 3. **Priority within each phase** - Models are ranked by their Priority field (higher number = higher priority)
3010
-
3011
- ### Dynamic Parallelization with Configurations
3012
-
3013
- Configurations can control parallel execution through parameters:
3014
-
3015
- ```typescript
3016
- // Set up configuration parameter
3017
- INSERT INTO AIConfigurationParam (ConfigurationID, Name, Type, Value)
3018
- VALUES (@DevConfigID, 'ParallelExecutions', 'number', '5');
3019
-
3020
- // Use in prompt setup
3021
- UPDATE AIPrompt
3022
- SET ParallelizationMode = 'ConfigParam',
3023
- ParallelConfigParam = 'ParallelExecutions'
3024
- WHERE ID = @PromptID;
3025
- ```
3026
-
3027
- ### Best Practices
3028
-
3029
- 1. **Use NULL ConfigurationID for fallback models** that provide a safety net when no configuration-specific models exist
3030
- 2. **Create environment-specific configurations** for different deployment scenarios (Production, Development, Testing)
3031
- 3. **Document configuration purposes** in the Description field to clarify their intended use
3032
- 4. **Test configuration precedence** to ensure configuration-specific models always take priority over defaults
3033
- 5. **Use configuration parameters** for environment-specific settings beyond just model selection
3034
- 6. **Assign models explicitly to configurations** to ensure clear separation between environments
3035
-
3036
- For more details on API key management, see the [AI Core API Keys documentation](../Core/README.md#api-key-management).
3037
-
3038
- ### Configuration Inheritance (v3.1+)
3039
-
3040
- AI Configurations now support parent-child inheritance relationships. This enables you to create child configurations that inherit prompt-model mappings from parent configurations while overriding specific settings.
3041
-
3042
- #### Use Cases
3043
-
3044
- - **Experimentation**: Create a child config that inherits from "Production" but overrides just 2-3 prompts with experimental models
3045
- - **Regional Variations**: Create "Production-EU" inheriting from "Production" with region-specific model overrides
3046
- - **A/B Testing**: Create test configurations that inherit baseline settings while varying specific prompts
3047
-
3048
- #### How Inheritance Works
3049
-
3050
- When you specify a `configurationId`, the system builds an inheritance chain from child to root:
3051
-
3052
- ```
3053
- Child Config → Parent Config → Grandparent Config → ... → Root Config
3054
- ```
3055
-
3056
- For each prompt, the system walks this chain looking for AIPromptModel matches:
3057
- 1. First checks for models assigned to the child config
3058
- 2. If none found, checks the parent config
3059
- 3. Continues up the chain until a match is found
3060
- 4. Falls back to NULL configuration models as final fallback
3061
-
3062
- #### Setting Up Inheritance
3063
-
3064
- ```sql
3065
- -- Create parent configuration
3066
- INSERT INTO AIConfiguration (ID, Name, Description, ParentID, IsDefault, Status)
3067
- VALUES (NEWID(), 'Production', 'Standard production configuration', NULL, 1, 'Active');
3068
-
3069
- -- Create child configuration that inherits from Production
3070
- INSERT INTO AIConfiguration (ID, Name, Description, ParentID, IsDefault, Status)
3071
- VALUES (NEWID(), 'Production-Experimental', 'Production with experimental models for select prompts',
3072
- (SELECT ID FROM AIConfiguration WHERE Name = 'Production'), 0, 'Active');
3073
- ```
3074
-
3075
- #### Model Override Example
3076
-
3077
- ```sql
3078
- -- Parent (Production) uses GPT-4 for the summarization prompt
3079
- INSERT INTO AIPromptModel (PromptID, ModelID, ConfigurationID, Priority)
3080
- VALUES (@SummarizePromptID, @GPT4ModelID, @ProductionConfigID, 100);
3081
-
3082
- -- Child (Production-Experimental) overrides with Claude for summarization
3083
- INSERT INTO AIPromptModel (PromptID, ModelID, ConfigurationID, Priority)
3084
- VALUES (@SummarizePromptID, @ClaudeModelID, @ExperimentalConfigID, 100);
3085
-
3086
- -- Other prompts in Production-Experimental inherit GPT-4 from parent
3087
- ```
3088
-
3089
- #### Using Inherited Configurations
3090
-
3091
- ```typescript
3092
- // Execute with child configuration - inherits from parent where no override exists
3093
- const result = await promptRunner.ExecutePrompt({
3094
- prompt: summarizePrompt,
3095
- configurationId: experimentalConfigId, // Child config
3096
- data: { text: 'Content to summarize' },
3097
- contextUser: currentUser
3098
- });
3099
-
3100
- // For summarization: Uses Claude (child override)
3101
- // For other prompts: Uses parent's models (inherited)
3102
- ```
3103
-
3104
- #### Parameter Inheritance
3105
-
3106
- Configuration parameters also inherit from parent configurations. Child parameters override parent parameters with the same name:
3107
-
3108
- ```typescript
3109
- // Get all parameters including inherited ones
3110
- const params = AIEngine.Instance.GetConfigurationParamsWithInheritance(childConfigId);
3111
-
3112
- // Parent has: temperature=0.7, maxTokens=4000
3113
- // Child has: temperature=0.9
3114
- // Result: temperature=0.9 (child), maxTokens=4000 (inherited)
3115
- ```
3116
-
3117
- #### Cycle Detection
3118
-
3119
- The system automatically detects circular references in the configuration hierarchy. If a cycle is detected, an error is thrown with a descriptive message showing the problematic chain.
3120
-
3121
- #### Best Practices
3122
-
3123
- 1. **Keep inheritance chains shallow** - 2-3 levels is usually sufficient
3124
- 2. **Document override intentions** - Use Description field to explain why child configs exist
3125
- 3. **Use meaningful names** - Name child configs to indicate their parent (e.g., "Production-EU", "Standard-Experimental")
3126
- 4. **Test inheritance** - Verify that child configs properly inherit unoverridden models
3127
-
3128
- ## Model Selection Tracking (v2.78+)
3129
-
3130
- The AIPromptRunner now provides comprehensive tracking of model selection decisions through the `modelSelectionInfo` property in `AIPromptRunResult`. This feature helps developers understand:
3131
-
3132
- - Which models were considered during selection
3133
- - Why specific models were or weren't available
3134
- - Which configuration influenced the selection
3135
- - What selection strategy was used
3136
-
3137
- ### Enhanced Model Selection Information
3138
-
3139
- ```typescript
3140
- const result = await promptRunner.ExecutePrompt(params);
3141
-
3142
- if (result.modelSelectionInfo) {
3143
- // Access the configuration that was used
3144
- const config = result.modelSelectionInfo.aiConfiguration;
3145
- console.log(`Configuration: ${config?.Name || 'Default'}`);
3146
-
3147
- // See all models that were considered
3148
- for (const candidate of result.modelSelectionInfo.modelsConsidered) {
3149
- console.log(`Model: ${candidate.model.Name}`);
3150
- console.log(` Vendor: ${candidate.vendor?.Name || 'default'}`);
3151
- console.log(` Priority: ${candidate.priority}`);
3152
- console.log(` Available: ${candidate.available}`);
3153
- if (!candidate.available) {
3154
- console.log(` Reason: ${candidate.unavailableReason}`);
3155
- }
3156
- }
3157
-
3158
- // Understand the final selection
3159
- console.log(`Selected: ${result.modelSelectionInfo.modelSelected.Name}`);
3160
- console.log(`Vendor: ${result.modelSelectionInfo.vendorSelected?.Name}`);
3161
- console.log(`Reason: ${result.modelSelectionInfo.selectionReason}`);
3162
- console.log(`Strategy: ${result.modelSelectionInfo.selectionStrategy}`);
3163
- }
3164
- ```
3165
-
3166
- ### Database Storage
3167
-
3168
- Model selection information is stored in the `AIPromptRun.ModelSelection` field as JSON, containing:
3169
- - Model and vendor IDs (not full entities)
3170
- - Configuration ID and name
3171
- - Array of considered models with their availability status
3172
- - Selection reason and strategy
3173
-
3174
- Additional fields track:
3175
- - `SelectionStrategy`: The strategy used ('Default', 'Specific', 'ByPower')
3176
- - `ModelPowerRank`: Power rank of the selected model
3177
- - `Status`: Execution status ('Pending', 'Running', 'Completed', 'Failed', 'Cancelled')
3178
- - `Cancelled`: Boolean flag for cancellation
3179
- - `CancellationReason`: Why execution was cancelled
3180
- - `ErrorDetails`: Detailed error information for failures
3181
-
3182
- ### Benefits
3183
-
3184
- - **Debugging**: Understand why a specific model was selected
3185
- - **Monitoring**: Track which models are being used across prompts
3186
- - **Optimization**: Identify models that are frequently unavailable
3187
- - **Compliance**: Audit model selection for regulatory requirements
3188
-
3189
- ## Version History
3190
-
3191
- - **2.78.0** - Added model selection tracking with full entity objects in results
3192
- - **2.77.0** - Enhanced status tracking and cancellation support
3193
- - **2.76.0** - Added intelligent failover system
3194
- - **2.75.0** - Introduced dynamic template composition
3195
- - **2.50.0** - Initial release with core prompt execution
237
+ - `@memberjunction/ai` -- Core AI abstractions (BaseLLM, ChatParams)
238
+ - `@memberjunction/ai-core-plus` -- AIPromptParams, AIPromptRunResult, extended entities
239
+ - `@memberjunction/ai-engine-base` -- AIEngineBase metadata cache
240
+ - `@memberjunction/aiengine` -- AIEngine server-side operations
241
+ - `@memberjunction/core` -- MJ framework core
242
+ - `@memberjunction/core-entities` -- Generated entity classes
243
+ - `@memberjunction/credentials` -- Credential resolution
244
+ - `@memberjunction/templates` -- Template rendering engine
245
+ - `@memberjunction/templates-base-types` -- Template base types
246
+ - `json5` -- Lenient JSON parsing for repair