@memberjunction/ai-prompts 4.0.0 → 4.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +155 -3104
- package/package.json +11 -11
package/README.md
CHANGED
|
@@ -1,3195 +1,246 @@
|
|
|
1
1
|
# @memberjunction/ai-prompts
|
|
2
2
|
|
|
3
|
-
Advanced AI prompt execution engine
|
|
3
|
+
Advanced AI prompt execution engine for MemberJunction. Provides hierarchical template composition, intelligent model selection with failover, parallel execution with judge-based result selection, structured output validation with retry, comprehensive execution tracking, and streaming support. This is the primary interface for executing AI prompts in the MemberJunction framework.
|
|
4
4
|
|
|
5
|
-
|
|
5
|
+
## Architecture
|
|
6
6
|
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
[
|
|
11
|
-
|
|
12
|
-
## Key Features
|
|
13
|
-
|
|
14
|
-
### 🎯 Effort Level Control
|
|
15
|
-
Granular control over AI model reasoning effort through a 1-100 integer scale. Higher values request more thorough reasoning and analysis from AI models that support effort levels.
|
|
7
|
+
```mermaid
|
|
8
|
+
graph TD
|
|
9
|
+
subgraph "@memberjunction/ai-prompts"
|
|
10
|
+
PR["AIPromptRunner"]
|
|
11
|
+
style PR fill:#2d8659,stroke:#1a5c3a,color:#fff
|
|
16
12
|
|
|
17
|
-
|
|
18
|
-
|
|
13
|
+
EP["ExecutionPlanner"]
|
|
14
|
+
style EP fill:#7c5295,stroke:#563a6b,color:#fff
|
|
19
15
|
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
3. **Provider default** - Model's natural behavior (lowest priority)
|
|
16
|
+
PEC["ParallelExecutionCoordinator"]
|
|
17
|
+
style PEC fill:#7c5295,stroke:#563a6b,color:#fff
|
|
23
18
|
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
27
|
-
- **Anthropic**: Maps to thinking mode with token budgets (1-100 → 25K-2M tokens)
|
|
28
|
-
- **Groq**: Maps to experimental `reasoning_effort` parameter
|
|
29
|
-
- **Gemini**: Controls reasoning mode intensity
|
|
19
|
+
PE["ParallelExecution"]
|
|
20
|
+
style PE fill:#7c5295,stroke:#563a6b,color:#fff
|
|
21
|
+
end
|
|
30
22
|
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
params.effortLevel = 85; // High effort for thorough analysis
|
|
23
|
+
subgraph "Execution Pipeline"
|
|
24
|
+
T["1. Template Rendering<br/>Handlebars + System Placeholders"]
|
|
25
|
+
style T fill:#b8762f,stroke:#8a5722,color:#fff
|
|
35
26
|
|
|
36
|
-
|
|
37
|
-
|
|
27
|
+
MS["2. Model Selection<br/>Default / Specific / ByPower"]
|
|
28
|
+
style MS fill:#b8762f,stroke:#8a5722,color:#fff
|
|
38
29
|
|
|
39
|
-
|
|
30
|
+
EX["3. LLM Execution<br/>With Streaming & Caching"]
|
|
31
|
+
style EX fill:#b8762f,stroke:#8a5722,color:#fff
|
|
40
32
|
|
|
41
|
-
|
|
33
|
+
VAL["4. Output Validation<br/>JSON Schema + Retry"]
|
|
34
|
+
style VAL fill:#b8762f,stroke:#8a5722,color:#fff
|
|
42
35
|
|
|
43
|
-
|
|
36
|
+
TRK["5. Execution Tracking<br/>AIPromptRun Records"]
|
|
37
|
+
style TRK fill:#b8762f,stroke:#8a5722,color:#fff
|
|
38
|
+
end
|
|
44
39
|
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
40
|
+
PR --> EP
|
|
41
|
+
PR --> PEC
|
|
42
|
+
PEC --> PE
|
|
43
|
+
PR --> T
|
|
44
|
+
T --> MS
|
|
45
|
+
MS --> EX
|
|
46
|
+
EX --> VAL
|
|
47
|
+
VAL --> TRK
|
|
50
48
|
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
49
|
+
subgraph Dependencies
|
|
50
|
+
AI["@memberjunction/ai<br/>BaseLLM"]
|
|
51
|
+
style AI fill:#2d6a9f,stroke:#1a4971,color:#fff
|
|
54
52
|
|
|
55
|
-
|
|
56
|
-
|
|
53
|
+
ACP["@memberjunction/ai-core-plus<br/>AIPromptParams"]
|
|
54
|
+
style ACP fill:#2d6a9f,stroke:#1a4971,color:#fff
|
|
57
55
|
|
|
58
|
-
|
|
56
|
+
AIE["@memberjunction/aiengine<br/>AIEngine"]
|
|
57
|
+
style AIE fill:#2d6a9f,stroke:#1a4971,color:#fff
|
|
59
58
|
|
|
60
|
-
|
|
59
|
+
TMPL["@memberjunction/templates<br/>TemplateEngine"]
|
|
60
|
+
style TMPL fill:#2d6a9f,stroke:#1a4971,color:#fff
|
|
61
61
|
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
participant User as User Request
|
|
65
|
-
participant Engine as AIPromptRunner
|
|
66
|
-
participant DB as AIPromptModel Table
|
|
67
|
-
participant Exec as Execution
|
|
68
|
-
|
|
69
|
-
User->>Engine: Execute Prompt with ConfigurationID
|
|
70
|
-
Engine->>DB: Get AIPromptModel records
|
|
71
|
-
DB-->>Engine: Return all records for prompt
|
|
72
|
-
|
|
73
|
-
Note over Engine: Filter Phase
|
|
74
|
-
Engine->>Engine: Keep: ConfigurationID match OR NULL
|
|
75
|
-
Engine->>Engine: Exclude: Different ConfigurationID
|
|
76
|
-
|
|
77
|
-
Note over Engine: Sort Phase (2-level)
|
|
78
|
-
Engine->>Engine: 1. Config-match before Universal
|
|
79
|
-
Engine->>Engine: 2. Priority DESC within group
|
|
80
|
-
|
|
81
|
-
Note over Engine: Expand Phase
|
|
82
|
-
loop For each AIPromptModel
|
|
83
|
-
alt VendorID specified
|
|
84
|
-
Engine->>Engine: Create 1 candidate (Model+Vendor)
|
|
85
|
-
else VendorID is NULL
|
|
86
|
-
Engine->>Engine: Create N candidates (all vendors)
|
|
87
|
-
Engine->>Engine: Sort by AIModelVendor.Priority DESC
|
|
88
|
-
end
|
|
62
|
+
CRED["@memberjunction/credentials<br/>CredentialEngine"]
|
|
63
|
+
style CRED fill:#2d6a9f,stroke:#1a4971,color:#fff
|
|
89
64
|
end
|
|
90
65
|
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
Exec-->>User: Return result
|
|
97
|
-
else Recoverable Error
|
|
98
|
-
Exec->>Exec: Try Candidate N+1 (instant)
|
|
99
|
-
else Fatal Error
|
|
100
|
-
Exec-->>User: Fail immediately
|
|
101
|
-
end
|
|
102
|
-
end
|
|
66
|
+
AI --> PR
|
|
67
|
+
ACP --> PR
|
|
68
|
+
AIE --> PR
|
|
69
|
+
TMPL --> PR
|
|
70
|
+
CRED --> PR
|
|
103
71
|
```
|
|
104
72
|
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
**Priority Precedence:**
|
|
108
|
-
1. **Configuration-specific** models (matching `ConfigurationID`) - Always tried first
|
|
109
|
-
2. **Universal** models (`ConfigurationID = NULL`) - Fallback options
|
|
110
|
-
3. Within each group: **Higher Priority number** tried first
|
|
111
|
-
|
|
112
|
-
**Configuration Filtering:**
|
|
113
|
-
- If `ConfigurationID` provided: Use matching config + universal (NULL) models
|
|
114
|
-
- If NO `ConfigurationID`: Use ONLY universal (NULL) models
|
|
115
|
-
- Models with DIFFERENT `ConfigurationID` are EXCLUDED
|
|
116
|
-
|
|
117
|
-
**Vendor Expansion:**
|
|
118
|
-
- `AIPromptModel.VendorID` specified → Single candidate (exact model+vendor)
|
|
119
|
-
- `AIPromptModel.VendorID = NULL` → Multiple candidates (all vendors for that model, sorted by `AIModelVendor.Priority DESC`)
|
|
120
|
-
|
|
121
|
-
#### Example Configuration
|
|
122
|
-
|
|
123
|
-
```sql
|
|
124
|
-
-- Example: Production prompt with config-specific and universal fallbacks
|
|
125
|
-
INSERT INTO AIPromptModel (PromptID, ModelID, VendorID, ConfigurationID, Priority, Status) VALUES
|
|
126
|
-
-- Config-specific models (tried first, regardless of priority number)
|
|
127
|
-
(@promptId, @gpt4Id, @openaiId, @prodConfigId, 5, 'Active'),
|
|
128
|
-
(@promptId, @gpt4Id, @azureId, @prodConfigId, 3, 'Active'),
|
|
129
|
-
|
|
130
|
-
-- Universal fallbacks (tried after config-specific, despite higher priority numbers)
|
|
131
|
-
(@promptId, @claudeId, @anthropicId, NULL, 10, 'Active'),
|
|
132
|
-
(@promptId, @geminiId, @googleId, NULL, 8, 'Active');
|
|
73
|
+
## Installation
|
|
133
74
|
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
(@promptId, @gpt4Id, NULL, @prodConfigId, 10, 'Active');
|
|
137
|
-
-- VendorID=NULL expands to all vendors (OpenAI, Azure, Groq)
|
|
138
|
-
-- Vendors sorted by AIModelVendor.Priority
|
|
75
|
+
```bash
|
|
76
|
+
npm install @memberjunction/ai-prompts
|
|
139
77
|
```
|
|
140
78
|
|
|
141
|
-
|
|
142
|
-
1. GPT-4/OpenAI (Config match, Priority 5)
|
|
143
|
-
2. GPT-4/Azure (Config match, Priority 3)
|
|
144
|
-
3. GPT-4/OpenAI (From VendorID=NULL expansion, highest AIModelVendor.Priority)
|
|
145
|
-
4. GPT-4/Azure (From VendorID=NULL expansion)
|
|
146
|
-
5. GPT-4/Groq (From VendorID=NULL expansion, lowest AIModelVendor.Priority)
|
|
147
|
-
6. Claude/Anthropic (Universal fallback, Priority 10)
|
|
148
|
-
7. Gemini/Google (Universal fallback, Priority 8)
|
|
149
|
-
|
|
150
|
-
#### Failover Behavior
|
|
79
|
+
## Key Features
|
|
151
80
|
|
|
152
|
-
|
|
153
|
-
- Authentication errors → Filters out all candidates from failed vendor
|
|
154
|
-
- Fatal errors → Stops immediately
|
|
155
|
-
- Recoverable errors → Tries next candidate instantly
|
|
81
|
+
### Hierarchical Template Composition
|
|
156
82
|
|
|
157
|
-
|
|
158
|
-
- If all candidates fail, retries entire list with delays
|
|
159
|
-
- Uses `AIPrompt.MaxRetries` and `RetryDelayMode` (Fixed/Linear/Exponential)
|
|
83
|
+
Build complex prompts from reusable sub-templates with unlimited nesting depth:
|
|
160
84
|
|
|
161
|
-
**Error Handling:**
|
|
162
85
|
```typescript
|
|
163
|
-
|
|
164
|
-
|
|
165
|
-
throw new Error('Please configure AIPromptModel records for this prompt');
|
|
166
|
-
}
|
|
167
|
-
```
|
|
168
|
-
|
|
169
|
-
### 🎯 Dynamic Hierarchical Template Composition
|
|
170
|
-
|
|
171
|
-
#### Why Dynamic Template Composition?
|
|
172
|
-
|
|
173
|
-
While MemberJunction's template system already supports static template composition (where Template A always includes Templates B and C), the AI Prompts system adds **dynamic template composition** - the ability to inject ANY prompt template into ANY other prompt template at runtime.
|
|
86
|
+
import { AIPromptRunner } from '@memberjunction/ai-prompts';
|
|
87
|
+
import { AIPromptParams, ChildPromptParam } from '@memberjunction/ai-core-plus';
|
|
174
88
|
|
|
175
|
-
|
|
176
|
-
```liquid
|
|
177
|
-
<!-- Email template always includes same header -->
|
|
178
|
-
{% include 'email-header' %}
|
|
179
|
-
{{ content }}
|
|
180
|
-
{% include 'email-footer' %}
|
|
181
|
-
```
|
|
89
|
+
const runner = new AIPromptRunner();
|
|
182
90
|
|
|
183
|
-
|
|
184
|
-
```typescript
|
|
185
|
-
// Inject ANY child prompt into ANY parent prompt at runtime
|
|
91
|
+
// Parent template uses {{ analysis }} and {{ summary }} placeholders
|
|
186
92
|
const params = new AIPromptParams();
|
|
187
|
-
params.prompt =
|
|
188
|
-
params.childPrompts = [
|
|
189
|
-
new ChildPromptParam(agentPrompt, 'agentInstructions') // Specific agent's prompt
|
|
190
|
-
];
|
|
191
|
-
// System prompt can use {{ agentInstructions }} to embed the agent's specific logic
|
|
192
|
-
```
|
|
193
|
-
|
|
194
|
-
#### The Agent System Use Case
|
|
195
|
-
|
|
196
|
-
This dynamic composition is crucial for AI Agents:
|
|
197
|
-
- **Agent Types** have **System Prompts** that control execution flow and response format
|
|
198
|
-
- **Individual Agents** have their own **specific prompts** with domain logic
|
|
199
|
-
- At runtime, any agent's prompt is dynamically injected into its type's system prompt
|
|
200
|
-
- This creates a complete prompt combining the control wrapper with agent-specific instructions
|
|
201
|
-
|
|
202
|
-
```typescript
|
|
203
|
-
// Agent Type System Prompt (controls flow)
|
|
204
|
-
const systemPrompt = {
|
|
205
|
-
templateText: `You are an AI agent. Follow these instructions:
|
|
206
|
-
|
|
207
|
-
{{ agentInstructions }} <!-- Dynamically injected at runtime -->
|
|
208
|
-
|
|
209
|
-
Respond in JSON format with: { decision: ..., reasoning: ... }`
|
|
210
|
-
};
|
|
211
|
-
|
|
212
|
-
// Individual Agent Prompt (domain logic)
|
|
213
|
-
const dataGatherAgent = {
|
|
214
|
-
templateText: `Your role is to gather data from: {{ dataSources }}`
|
|
215
|
-
};
|
|
216
|
-
|
|
217
|
-
// At runtime, compose them dynamically
|
|
93
|
+
params.prompt = parentPrompt;
|
|
218
94
|
params.childPrompts = [
|
|
219
|
-
|
|
95
|
+
new ChildPromptParam(analysisParams, 'analysis'),
|
|
96
|
+
new ChildPromptParam(summaryParams, 'summary')
|
|
220
97
|
];
|
|
221
|
-
|
|
222
|
-
|
|
223
|
-
### 🔄 System Placeholders
|
|
224
|
-
Automatically inject common values into all templates without manual data passing. Includes date/time, user context, prompt metadata, and more.
|
|
225
|
-
|
|
226
|
-
```liquid
|
|
227
|
-
Current user: {{ _USER_NAME }}
|
|
228
|
-
Date: {{ _CURRENT_DATE }}
|
|
229
|
-
Expected output: {{ _OUTPUT_EXAMPLE }}
|
|
230
|
-
```
|
|
231
|
-
|
|
232
|
-
## System Placeholders Reference
|
|
233
|
-
|
|
234
|
-
System placeholders are automatically available in all AI prompt templates, providing dynamic values like current date/time, prompt metadata, and user context without requiring manual data passing.
|
|
235
|
-
|
|
236
|
-
### Available System Placeholders
|
|
237
|
-
|
|
238
|
-
#### Date/Time Placeholders
|
|
239
|
-
- `{{ _CURRENT_DATE }}` - Current date in YYYY-MM-DD format
|
|
240
|
-
- `{{ _CURRENT_TIME }}` - Current time in HH:MM AM/PM format with timezone
|
|
241
|
-
- `{{ _CURRENT_DATE_AND_TIME }}` - Full timestamp with date and time
|
|
242
|
-
- `{{ _CURRENT_DAY_OF_WEEK }}` - Current day name (e.g., Monday, Tuesday)
|
|
243
|
-
- `{{ _CURRENT_TIMEZONE }}` - Current timezone identifier
|
|
244
|
-
- `{{ _CURRENT_TIMESTAMP_UTC }}` - Current UTC timestamp in ISO format
|
|
245
|
-
|
|
246
|
-
#### Prompt Metadata Placeholders
|
|
247
|
-
- `{{ _OUTPUT_EXAMPLE }}` - The expected output example from the prompt configuration
|
|
248
|
-
- `{{ _PROMPT_NAME }}` - The name of the current prompt
|
|
249
|
-
- `{{ _PROMPT_DESCRIPTION }}` - The description of the current prompt
|
|
250
|
-
- `{{ _EXPECTED_OUTPUT_TYPE }}` - The expected output type (string, object, number, etc.)
|
|
251
|
-
- `{{ _RESPONSE_FORMAT }}` - The expected response format from the prompt
|
|
252
|
-
|
|
253
|
-
#### User Context Placeholders
|
|
254
|
-
- `{{ _USER_NAME }}` - Current user's full name
|
|
255
|
-
- `{{ _USER_EMAIL }}` - Current user's email address
|
|
256
|
-
- `{{ _USER_ID }}` - Current user's unique identifier
|
|
257
|
-
|
|
258
|
-
#### Environment Placeholders
|
|
259
|
-
- `{{ _ENVIRONMENT }}` - Current environment (development, staging, production)
|
|
260
|
-
- `{{ _API_VERSION }}` - Current API version
|
|
261
|
-
|
|
262
|
-
### System Placeholder Usage Examples
|
|
263
|
-
|
|
264
|
-
#### Example 1: Time-Aware Agent Prompt
|
|
265
|
-
```liquid
|
|
266
|
-
You are an AI assistant helping {{ _USER_NAME }} on {{ _CURRENT_DAY_OF_WEEK }}, {{ _CURRENT_DATE }} at {{ _CURRENT_TIME }}.
|
|
267
|
-
|
|
268
|
-
User's request: {{ userRequest }}
|
|
269
|
-
|
|
270
|
-
Please provide a helpful response considering the current time and day.
|
|
271
|
-
```
|
|
272
|
-
|
|
273
|
-
#### Example 2: Agent Type System Prompt with Metadata
|
|
274
|
-
```liquid
|
|
275
|
-
# Agent Type: Loop Decision Maker
|
|
276
|
-
|
|
277
|
-
Current execution context:
|
|
278
|
-
- Date/Time: {{ _CURRENT_DATE_AND_TIME }}
|
|
279
|
-
- User: {{ _USER_NAME }} ({{ _USER_EMAIL }})
|
|
280
|
-
- Environment: {{ _ENVIRONMENT }}
|
|
281
|
-
|
|
282
|
-
## Expected Output Format
|
|
283
|
-
{{ _OUTPUT_EXAMPLE }}
|
|
284
|
-
|
|
285
|
-
## Agent Specific Instructions
|
|
286
|
-
{{ agentResponse }}
|
|
98
|
+
params.data = { userInput: 'complex data to process' };
|
|
287
99
|
|
|
288
|
-
|
|
289
|
-
```
|
|
290
|
-
|
|
291
|
-
#### Example 3: Debug-Friendly Prompt
|
|
292
|
-
```liquid
|
|
293
|
-
[Debug Info]
|
|
294
|
-
- Prompt: {{ _PROMPT_NAME }}
|
|
295
|
-
- Description: {{ _PROMPT_DESCRIPTION }}
|
|
296
|
-
- Expected Output: {{ _EXPECTED_OUTPUT_TYPE }}
|
|
297
|
-
- User ID: {{ _USER_ID }}
|
|
298
|
-
- Timestamp: {{ _CURRENT_TIMESTAMP_UTC }}
|
|
299
|
-
|
|
300
|
-
[Task]
|
|
301
|
-
{{ taskDescription }}
|
|
302
|
-
```
|
|
303
|
-
|
|
304
|
-
### Adding Custom System Placeholders
|
|
305
|
-
|
|
306
|
-
You can add custom system placeholders programmatically:
|
|
307
|
-
|
|
308
|
-
```typescript
|
|
309
|
-
import { SystemPlaceholderManager } from '@memberjunction/ai-prompts';
|
|
310
|
-
|
|
311
|
-
// Add a custom placeholder
|
|
312
|
-
SystemPlaceholderManager.addPlaceholder({
|
|
313
|
-
name: '_ORGANIZATION_NAME',
|
|
314
|
-
description: 'Current organization name',
|
|
315
|
-
getValue: async (params) => {
|
|
316
|
-
// Custom logic to get organization name
|
|
317
|
-
return params.contextUser?.OrganizationName || 'Default Organization';
|
|
318
|
-
}
|
|
319
|
-
});
|
|
320
|
-
|
|
321
|
-
// Or add directly to the array
|
|
322
|
-
const placeholders = SystemPlaceholderManager.getPlaceholders();
|
|
323
|
-
placeholders.push({
|
|
324
|
-
name: '_CUSTOM_VALUE',
|
|
325
|
-
description: 'My custom value',
|
|
326
|
-
getValue: async (params) => 'custom result'
|
|
327
|
-
});
|
|
328
|
-
```
|
|
329
|
-
|
|
330
|
-
### Data Merge Priority Order
|
|
331
|
-
|
|
332
|
-
When rendering templates, data is merged in this priority order (highest to lowest):
|
|
333
|
-
1. Template-specific data (`templateData` parameter)
|
|
334
|
-
2. Child template renders (for hierarchical template composition)
|
|
335
|
-
3. User-provided data (`data` parameter)
|
|
336
|
-
4. System placeholders (lowest priority)
|
|
337
|
-
|
|
338
|
-
This means users can override system placeholders by providing their own values with the same names.
|
|
339
|
-
|
|
340
|
-
### ⚡ Parallel Processing
|
|
341
|
-
Multi-model execution with intelligent result selection strategies and AI judge ranking for optimal results.
|
|
342
|
-
|
|
343
|
-
### ✅ Output Validation
|
|
344
|
-
JSON schema validation against OutputExample with intelligent retry logic and configurable validation behaviors.
|
|
345
|
-
|
|
346
|
-
### 🚫 Cancellation Support
|
|
347
|
-
AbortSignal integration for graceful execution cancellation with proper cleanup and partial result preservation.
|
|
348
|
-
|
|
349
|
-
### 📈 Progress & Streaming
|
|
350
|
-
Real-time progress callbacks and streaming response support for responsive user interfaces.
|
|
351
|
-
|
|
352
|
-
### 📊 Comprehensive Tracking
|
|
353
|
-
Hierarchical execution logging with the AIPromptRun entity, including token usage, timing, and validation attempts.
|
|
354
|
-
|
|
355
|
-
### 🤖 Agent Integration
|
|
356
|
-
Seamless integration with AI Agents through hierarchical prompts and execution tracking.
|
|
357
|
-
|
|
358
|
-
### 💾 Intelligent Caching
|
|
359
|
-
Vector similarity matching and TTL-based result caching for performance optimization.
|
|
360
|
-
|
|
361
|
-
### 🔧 Template Integration
|
|
362
|
-
Dynamic prompt generation with MemberJunction template system supporting conditionals, loops, and data injection.
|
|
363
|
-
|
|
364
|
-
## Installation
|
|
365
|
-
|
|
366
|
-
```bash
|
|
367
|
-
npm install @memberjunction/ai-prompts
|
|
100
|
+
const result = await runner.ExecutePrompt(params);
|
|
368
101
|
```
|
|
369
102
|
|
|
370
|
-
|
|
103
|
+
Execution order:
|
|
104
|
+
1. Child prompts render depth-first (children before parents)
|
|
105
|
+
2. Sibling prompts at each level execute in parallel
|
|
106
|
+
3. Child results replace placeholders in parent template
|
|
107
|
+
4. Final composed prompt executes as a single LLM call
|
|
371
108
|
|
|
372
|
-
###
|
|
109
|
+
### Model Selection Strategies
|
|
373
110
|
|
|
374
|
-
|
|
375
|
-
- **This package** now imports base AI types from `@memberjunction/ai` (Core)
|
|
376
|
-
- **Prompt-specific types** remain in this package:
|
|
377
|
-
- `AIPromptParams`, `AIPromptRunResult`
|
|
378
|
-
- `ChildPromptParam`, `SystemPlaceholder`
|
|
379
|
-
- Execution callbacks and progress types
|
|
380
|
-
- **Agent integration types** are imported from `@memberjunction/ai-agents` when needed
|
|
111
|
+
Three strategies for selecting which AI model executes a prompt:
|
|
381
112
|
|
|
382
|
-
|
|
113
|
+
| Strategy | Description |
|
|
114
|
+
|---|---|
|
|
115
|
+
| `Default` | Uses the AI configuration to determine the model based on priority and availability |
|
|
116
|
+
| `Specific` | Uses explicitly associated models from the AIPromptModels table |
|
|
117
|
+
| `ByPower` | Selects the highest PowerRank model matching the prompt's model type |
|
|
383
118
|
|
|
384
|
-
|
|
385
|
-
|
|
386
|
-
|
|
387
|
-
|
|
388
|
-
- [@memberjunction/templates](../../Templates/README.md) for template rendering
|
|
119
|
+
Model selection precedence (highest to lowest):
|
|
120
|
+
1. `AIPromptParams.override` -- Runtime model/vendor override
|
|
121
|
+
2. `AIPromptParams.modelSelectionPrompt` -- Alternate prompt for model config
|
|
122
|
+
3. Prompt's own model configuration (strategy + associations)
|
|
389
123
|
|
|
390
|
-
|
|
124
|
+
### Parallel Execution with Judging
|
|
391
125
|
|
|
392
|
-
|
|
126
|
+
Execute prompts across multiple models simultaneously and select the best result:
|
|
393
127
|
|
|
394
|
-
|
|
128
|
+
- Configurable execution groups with different models
|
|
129
|
+
- AI judge prompt evaluates and ranks results
|
|
130
|
+
- Automatic selection of best result based on judge scoring
|
|
131
|
+
- Full tracking of all parallel results
|
|
395
132
|
|
|
396
|
-
|
|
397
|
-
MemberJunction's template system supports embedding templates within templates through `{% include %}` directives. This is perfect for fixed relationships:
|
|
398
|
-
- Email templates with standard headers/footers
|
|
399
|
-
- Report templates with consistent formatting sections
|
|
400
|
-
- Any scenario where Template A always includes Templates B and C
|
|
133
|
+
### Output Validation and Retry
|
|
401
134
|
|
|
402
|
-
|
|
403
|
-
The AI Prompts system adds runtime template composition where relationships are determined dynamically:
|
|
404
|
-
- **Runtime Flexibility**: Inject ANY prompt template into ANY other prompt template
|
|
405
|
-
- **Context-Aware**: Choose which child templates to inject based on runtime conditions
|
|
406
|
-
- **Agent Architecture**: Combine system prompts (control flow) with agent prompts (domain logic)
|
|
407
|
-
- **Modular Design**: Build complex prompts from reusable components selected at runtime
|
|
135
|
+
Automatic validation of AI outputs with configurable retry:
|
|
408
136
|
|
|
409
|
-
|
|
137
|
+
- JSON schema validation against `OutputExample` definitions
|
|
138
|
+
- Automatic JSON repair via JSON5 parsing and LLM-based repair
|
|
139
|
+
- Configurable retry count with the original or repaired prompts
|
|
140
|
+
- Validation syntax cleaning (removes `?`, `*`, `:type` markers from JSON keys)
|
|
141
|
+
- Detailed validation attempt tracking
|
|
410
142
|
|
|
411
|
-
###
|
|
143
|
+
### Streaming Support
|
|
412
144
|
|
|
413
|
-
|
|
145
|
+
Real-time streaming of LLM responses:
|
|
414
146
|
|
|
415
147
|
```typescript
|
|
416
|
-
|
|
417
|
-
|
|
418
|
-
|
|
419
|
-
|
|
420
|
-
const summaryPrompt = prompts.find(p => p.Name === 'Document Summarization');
|
|
421
|
-
|
|
422
|
-
// Execute the prompt
|
|
423
|
-
const params: AIPromptParams = {
|
|
424
|
-
prompt: summaryPrompt,
|
|
425
|
-
data: {
|
|
426
|
-
documentText: "Long document content here...",
|
|
427
|
-
targetLength: "2 paragraphs"
|
|
428
|
-
},
|
|
429
|
-
contextUser: currentUser
|
|
148
|
+
const params = new AIPromptParams();
|
|
149
|
+
params.prompt = myPrompt;
|
|
150
|
+
params.onStreaming = (chunk) => {
|
|
151
|
+
process.stdout.write(chunk.content);
|
|
430
152
|
};
|
|
431
153
|
|
|
432
|
-
const runner = new AIPromptRunner();
|
|
433
154
|
const result = await runner.ExecutePrompt(params);
|
|
434
|
-
|
|
435
|
-
if (result.success) {
|
|
436
|
-
console.log("Summary:", result.result);
|
|
437
|
-
console.log(`Execution time: ${result.executionTimeMS}ms`);
|
|
438
|
-
console.log(`Prompt tokens: ${result.promptTokens}`);
|
|
439
|
-
console.log(`Completion tokens: ${result.completionTokens}`);
|
|
440
|
-
console.log(`Total tokens: ${result.tokensUsed}`);
|
|
441
|
-
if (result.cost) {
|
|
442
|
-
console.log(`Cost: ${result.cost} ${result.costCurrency || 'USD'}`);
|
|
443
|
-
}
|
|
444
|
-
} else {
|
|
445
|
-
console.error("Error:", result.errorMessage);
|
|
446
|
-
}
|
|
447
155
|
```
|
|
448
156
|
|
|
449
|
-
|
|
157
|
+
### Execution Tracking
|
|
450
158
|
|
|
451
|
-
|
|
159
|
+
Every prompt execution creates an `AIPromptRun` record with:
|
|
160
|
+
- Model and vendor used
|
|
161
|
+
- Template rendering results
|
|
162
|
+
- Token usage (prompt + completion)
|
|
163
|
+
- Cost tracking
|
|
164
|
+
- Execution time
|
|
165
|
+
- Parent/child relationships for hierarchical prompts
|
|
166
|
+
- Agent run linkage via `agentRunId`
|
|
452
167
|
|
|
453
|
-
|
|
454
|
-
import { AIPromptRunner } from '@memberjunction/ai-prompts';
|
|
455
|
-
import { AIEngine } from '@memberjunction/aiengine';
|
|
456
|
-
|
|
457
|
-
// Initialize the AI Engine
|
|
458
|
-
await AIEngine.Instance.Config(false, currentUser);
|
|
459
|
-
|
|
460
|
-
// Find a prompt
|
|
461
|
-
const prompt = AIEngine.Instance.Prompts.find(p => p.Name === 'Text Analysis');
|
|
462
|
-
|
|
463
|
-
// Execute with data
|
|
464
|
-
const runner = new AIPromptRunner();
|
|
465
|
-
const result = await runner.ExecutePrompt({
|
|
466
|
-
prompt: prompt,
|
|
467
|
-
data: {
|
|
468
|
-
text: "Analyze this sample text for sentiment and key themes.",
|
|
469
|
-
format: "bullet points"
|
|
470
|
-
},
|
|
471
|
-
contextUser: currentUser
|
|
472
|
-
});
|
|
473
|
-
|
|
474
|
-
console.log("Analysis:", result.result);
|
|
475
|
-
```
|
|
476
|
-
|
|
477
|
-
### 2. Template-Driven Prompts
|
|
478
|
-
|
|
479
|
-
```typescript
|
|
480
|
-
// Prompt templates support dynamic data substitution
|
|
481
|
-
const templatePrompt = {
|
|
482
|
-
UserMessage: `Analyze the {{entity.EntityType}} record for {{entity.Name}}.
|
|
483
|
-
Focus on {{analysisType}} and provide insights about {{entity.Description}}.`
|
|
484
|
-
};
|
|
485
|
-
|
|
486
|
-
// Data context provides template variables
|
|
487
|
-
const result = await runner.ExecutePrompt({
|
|
488
|
-
prompt: templatePrompt,
|
|
489
|
-
data: {
|
|
490
|
-
entity: {
|
|
491
|
-
EntityType: "Customer",
|
|
492
|
-
Name: "Acme Corp",
|
|
493
|
-
Description: "Enterprise software company"
|
|
494
|
-
},
|
|
495
|
-
analysisType: "growth opportunities"
|
|
496
|
-
},
|
|
497
|
-
contextUser: currentUser
|
|
498
|
-
});
|
|
499
|
-
```
|
|
500
|
-
|
|
501
|
-
### 3. Parallel Execution with Multiple Models
|
|
502
|
-
|
|
503
|
-
```typescript
|
|
504
|
-
// Execute the same prompt across multiple models in parallel
|
|
505
|
-
const multiModelPrompt = prompts.find(p => p.ParallelizationMode === 'ModelSpecific');
|
|
506
|
-
|
|
507
|
-
const result = await runner.ExecutePrompt({
|
|
508
|
-
prompt: multiModelPrompt,
|
|
509
|
-
data: { query: "Analyze this data pattern" },
|
|
510
|
-
contextUser: currentUser
|
|
511
|
-
});
|
|
512
|
-
|
|
513
|
-
// When using parallel execution, the system automatically selects the best result
|
|
514
|
-
console.log(`Final result: ${result.result}`);
|
|
515
|
-
console.log(`Execution time: ${result.executionTimeMS}ms`);
|
|
516
|
-
console.log(`Total tokens used: ${result.tokensUsed}`);
|
|
517
|
-
|
|
518
|
-
// The promptRun entity contains metadata about parallel execution in its Messages field
|
|
519
|
-
if (result.promptRun?.Messages) {
|
|
520
|
-
const metadata = JSON.parse(result.promptRun.Messages);
|
|
521
|
-
if (metadata.parallelExecution) {
|
|
522
|
-
console.log(`Parallelization mode: ${metadata.parallelExecution.parallelizationMode}`);
|
|
523
|
-
console.log(`Total tasks: ${metadata.parallelExecution.totalTasks}`);
|
|
524
|
-
console.log(`Successful tasks: ${metadata.parallelExecution.successfulTasks}`);
|
|
525
|
-
}
|
|
526
|
-
}
|
|
527
|
-
```
|
|
528
|
-
|
|
529
|
-
### 4. Dynamic Template Composition for AI Agents
|
|
168
|
+
### Credential Resolution
|
|
530
169
|
|
|
531
|
-
|
|
170
|
+
Hierarchical credential resolution for API keys:
|
|
532
171
|
|
|
533
|
-
|
|
534
|
-
|
|
535
|
-
|
|
536
|
-
|
|
537
|
-
|
|
538
|
-
|
|
539
|
-
TemplateID: "system-prompt-template-id",
|
|
540
|
-
// Template contains: "You are an AI agent. {{ agentInstructions }} Respond with JSON..."
|
|
541
|
-
};
|
|
172
|
+
1. `AIPromptParams.credentialId` (per-request override)
|
|
173
|
+
2. `AIPromptModel.CredentialID` (prompt-model specific)
|
|
174
|
+
3. `AIModelVendor.CredentialID` (model-vendor specific)
|
|
175
|
+
4. `AIVendor.CredentialID` (vendor default)
|
|
176
|
+
5. `AIPromptParams.apiKeys[]` (legacy runtime keys)
|
|
177
|
+
6. `AI_VENDOR_API_KEY__<DRIVER>` environment variables (legacy)
|
|
542
178
|
|
|
543
|
-
|
|
544
|
-
const specificAgentPrompt = {
|
|
545
|
-
Name: "Customer Churn Analysis Agent",
|
|
546
|
-
TemplateID: "churn-agent-template-id",
|
|
547
|
-
// Template contains: "Analyze customer data for churn risk factors..."
|
|
548
|
-
};
|
|
179
|
+
### Failover
|
|
549
180
|
|
|
550
|
-
|
|
551
|
-
const runner = new AIPromptRunner();
|
|
552
|
-
const result = await runner.ExecutePrompt({
|
|
553
|
-
prompt: agentTypeSystemPrompt, // Parent template
|
|
554
|
-
childPrompts: [
|
|
555
|
-
// Dynamically inject the specific agent's instructions
|
|
556
|
-
new ChildPromptParam(specificAgentPrompt, 'agentInstructions')
|
|
557
|
-
],
|
|
558
|
-
data: {
|
|
559
|
-
customerData: analysisData,
|
|
560
|
-
thresholds: { churnRisk: 0.7 }
|
|
561
|
-
},
|
|
562
|
-
contextUser: currentUser
|
|
563
|
-
});
|
|
564
|
-
|
|
565
|
-
// The system executed ONE prompt that combined:
|
|
566
|
-
// 1. System prompt wrapper (control flow)
|
|
567
|
-
// 2. Specific agent instructions (domain logic)
|
|
568
|
-
// 3. Runtime data
|
|
569
|
-
console.log("Agent decision:", result.result);
|
|
570
|
-
```
|
|
181
|
+
When a model fails due to rate limiting, authentication errors, or other transient issues, the runner can automatically retry with alternate models from the selection candidates.
|
|
571
182
|
|
|
572
|
-
|
|
573
|
-
- Different agents can use the SAME system prompt template
|
|
574
|
-
- System prompt enforces consistent response format across all agents
|
|
575
|
-
- Agent-specific logic is cleanly separated and reusable
|
|
576
|
-
- Runtime composition allows flexible agent architectures
|
|
183
|
+
## Usage
|
|
577
184
|
|
|
578
|
-
###
|
|
185
|
+
### Basic Prompt Execution
|
|
579
186
|
|
|
580
187
|
```typescript
|
|
581
188
|
import { AIPromptRunner } from '@memberjunction/ai-prompts';
|
|
189
|
+
import { AIPromptParams } from '@memberjunction/ai-core-plus';
|
|
582
190
|
import { AIEngine } from '@memberjunction/aiengine';
|
|
583
191
|
|
|
584
|
-
//
|
|
585
|
-
|
|
586
|
-
|
|
587
|
-
await AIEngine.Instance.Config(false, currentUser);
|
|
588
|
-
const runner = new AIPromptRunner();
|
|
589
|
-
|
|
590
|
-
// Set up cancellation (e.g., from user clicking cancel button)
|
|
591
|
-
const controller = new AbortController();
|
|
592
|
-
const timeoutId = setTimeout(() => {
|
|
593
|
-
controller.abort();
|
|
594
|
-
console.log('Operation timed out after 2 minutes');
|
|
595
|
-
}, 120000);
|
|
596
|
-
|
|
597
|
-
try {
|
|
598
|
-
const result = await runner.ExecutePrompt({
|
|
599
|
-
prompt: complexAnalysisPrompt, // ParallelizationMode: 'ModelSpecific'
|
|
600
|
-
data: {
|
|
601
|
-
document: largeDocument,
|
|
602
|
-
analysisType: 'comprehensive',
|
|
603
|
-
outputFormat: 'structured'
|
|
604
|
-
},
|
|
605
|
-
contextUser: currentUser,
|
|
606
|
-
|
|
607
|
-
// Enable cancellation
|
|
608
|
-
cancellationToken: controller.signal,
|
|
609
|
-
|
|
610
|
-
// Track progress throughout execution
|
|
611
|
-
onProgress: (progress) => {
|
|
612
|
-
console.log(`[${progress.step}] ${progress.percentage}% - ${progress.message}`);
|
|
613
|
-
|
|
614
|
-
// Handle parallel execution progress
|
|
615
|
-
if (progress.metadata?.parallelExecution) {
|
|
616
|
-
const parallel = progress.metadata.parallelExecution;
|
|
617
|
-
console.log(` → Group ${parallel.currentGroup + 1}/${parallel.totalGroups}, Tasks: ${parallel.completedTasks}/${parallel.totalTasks}`);
|
|
618
|
-
}
|
|
619
|
-
|
|
620
|
-
// Update UI
|
|
621
|
-
updateProgressBar(progress.percentage);
|
|
622
|
-
updateStatusText(progress.message);
|
|
623
|
-
},
|
|
624
|
-
|
|
625
|
-
// Receive streaming content updates
|
|
626
|
-
onStreaming: (chunk) => {
|
|
627
|
-
if (chunk.isComplete) {
|
|
628
|
-
console.log(`Streaming complete for ${chunk.modelName}`);
|
|
629
|
-
finalizeOutput();
|
|
630
|
-
} else {
|
|
631
|
-
// Show real-time content generation
|
|
632
|
-
console.log(`[${chunk.modelName}]: ${chunk.content.substring(0, 50)}...`);
|
|
633
|
-
appendToDisplay(chunk.content, chunk.taskId);
|
|
634
|
-
}
|
|
635
|
-
}
|
|
636
|
-
});
|
|
637
|
-
|
|
638
|
-
// Clear timeout since we completed successfully
|
|
639
|
-
clearTimeout(timeoutId);
|
|
640
|
-
|
|
641
|
-
// Handle different result scenarios
|
|
642
|
-
if (result.cancelled) {
|
|
643
|
-
console.log(`Execution cancelled: ${result.cancellationReason}`);
|
|
644
|
-
// May still have partial results available
|
|
645
|
-
if (result.additionalResults && result.additionalResults.length > 0) {
|
|
646
|
-
console.log(`${result.additionalResults.length} partial results available`);
|
|
647
|
-
}
|
|
648
|
-
} else if (result.success) {
|
|
649
|
-
console.log('Execution completed successfully!');
|
|
650
|
-
console.log(`Primary result from ${result.modelInfo?.modelName}: ${result.result}`);
|
|
651
|
-
|
|
652
|
-
// Analyze judge selection if multiple results
|
|
653
|
-
if (result.ranking && result.judgeRationale) {
|
|
654
|
-
console.log(`Selected as #${result.ranking} by AI judge: ${result.judgeRationale}`);
|
|
655
|
-
}
|
|
656
|
-
|
|
657
|
-
// Review alternative results from parallel execution
|
|
658
|
-
if (result.additionalResults) {
|
|
659
|
-
console.log(`${result.additionalResults.length} alternative results ranked by judge:`);
|
|
660
|
-
result.additionalResults.forEach((altResult, index) => {
|
|
661
|
-
console.log(` ${altResult.ranking}. ${altResult.modelInfo?.modelName}: ${altResult.judgeRationale}`);
|
|
662
|
-
});
|
|
663
|
-
}
|
|
664
|
-
|
|
665
|
-
// Analyze execution performance using hierarchical logging
|
|
666
|
-
if (result.promptRun?.RunType === 'ParallelParent') {
|
|
667
|
-
await analyzeParallelExecutionPerformance(result.promptRun.ID);
|
|
668
|
-
}
|
|
669
|
-
|
|
670
|
-
// Check streaming and caching
|
|
671
|
-
if (result.wasStreamed) {
|
|
672
|
-
console.log('Response was streamed in real-time');
|
|
673
|
-
}
|
|
674
|
-
if (result.cacheInfo?.cacheHit) {
|
|
675
|
-
console.log(`Result served from cache: ${result.cacheInfo.cacheSource}`);
|
|
676
|
-
}
|
|
677
|
-
} else {
|
|
678
|
-
console.error(`Execution failed: ${result.errorMessage}`);
|
|
679
|
-
}
|
|
680
|
-
|
|
681
|
-
} catch (error) {
|
|
682
|
-
clearTimeout(timeoutId);
|
|
683
|
-
console.error('Execution error:', error.message);
|
|
684
|
-
}
|
|
685
|
-
}
|
|
686
|
-
|
|
687
|
-
// Helper function to analyze parallel execution performance
|
|
688
|
-
async function analyzeParallelExecutionPerformance(parentPromptRunId: string) {
|
|
689
|
-
// Query hierarchical logs to understand execution breakdown
|
|
690
|
-
console.log('Analyzing parallel execution performance...');
|
|
691
|
-
|
|
692
|
-
// This would typically be a database query or API call
|
|
693
|
-
// For demonstration, showing the concept:
|
|
694
|
-
const analysisQuery = `
|
|
695
|
-
SELECT
|
|
696
|
-
pr.RunType,
|
|
697
|
-
pr.ExecutionOrder,
|
|
698
|
-
pr.Success,
|
|
699
|
-
pr.ExecutionTimeMS,
|
|
700
|
-
pr.TokensUsed,
|
|
701
|
-
m.Name as ModelName
|
|
702
|
-
FROM AIPromptRun pr
|
|
703
|
-
JOIN AIModel m ON pr.ModelID = m.ID
|
|
704
|
-
WHERE pr.ParentID = '${parentPromptRunId}' OR pr.ID = '${parentPromptRunId}'
|
|
705
|
-
ORDER BY pr.RunType, pr.ExecutionOrder
|
|
706
|
-
`;
|
|
707
|
-
|
|
708
|
-
console.log('Performance analysis query:', analysisQuery);
|
|
709
|
-
// Execute query and analyze results...
|
|
710
|
-
}
|
|
711
|
-
|
|
712
|
-
// Execute the comprehensive example
|
|
713
|
-
comprehensivePromptExecution().catch(console.error);
|
|
714
|
-
```
|
|
715
|
-
|
|
716
|
-
## Advanced Features
|
|
192
|
+
// Get prompt from metadata
|
|
193
|
+
await AIEngine.Instance.Config(false, contextUser);
|
|
194
|
+
const prompt = AIEngine.Instance.Prompts.find(p => p.Name === 'Summarize Content');
|
|
717
195
|
|
|
718
|
-
|
|
719
|
-
|
|
720
|
-
|
|
721
|
-
|
|
722
|
-
|
|
723
|
-
|
|
724
|
-
When a prompt execution fails, the system:
|
|
725
|
-
1. **Analyzes the error** using the ErrorAnalyzer to determine if failover is appropriate
|
|
726
|
-
2. **Selects alternative models/vendors** based on the configured strategy
|
|
727
|
-
3. **Applies intelligent delays** with exponential backoff to prevent overwhelming providers
|
|
728
|
-
4. **Tracks all attempts** for debugging and analysis
|
|
729
|
-
5. **Updates the execution** to use the successful model/vendor combination
|
|
730
|
-
|
|
731
|
-
#### Failover Configuration
|
|
732
|
-
|
|
733
|
-
Configure failover behavior at the prompt level:
|
|
734
|
-
|
|
735
|
-
```typescript
|
|
736
|
-
// Database columns added to AIPrompt entity:
|
|
737
|
-
FailoverStrategy: 'SameModelDifferentVendor' | 'NextBestModel' | 'PowerRank' | 'None'
|
|
738
|
-
FailoverMaxAttempts: number // Maximum failover attempts (default: 3)
|
|
739
|
-
FailoverDelaySeconds: number // Initial delay between attempts (default: 1)
|
|
740
|
-
FailoverModelStrategy: 'PreferSameModel' | 'PreferDifferentModel' | 'RequireSameModel'
|
|
741
|
-
FailoverErrorScope: 'All' | 'NetworkOnly' | 'RateLimitOnly' | 'ServiceErrorOnly'
|
|
742
|
-
```
|
|
743
|
-
|
|
744
|
-
#### Failover Strategies Explained
|
|
745
|
-
|
|
746
|
-
**SameModelDifferentVendor**: Ideal for multi-cloud deployments
|
|
747
|
-
```typescript
|
|
748
|
-
// Example: Claude from different providers
|
|
749
|
-
// Primary: Anthropic API
|
|
750
|
-
// Failover 1: AWS Bedrock
|
|
751
|
-
// Failover 2: Google Vertex AI
|
|
752
|
-
```
|
|
753
|
-
|
|
754
|
-
**NextBestModel**: Balances capability and availability
|
|
755
|
-
```typescript
|
|
756
|
-
// Example: Gradual capability reduction
|
|
757
|
-
// Primary: GPT-4-turbo
|
|
758
|
-
// Failover 1: Claude-3-opus
|
|
759
|
-
// Failover 2: GPT-3.5-turbo
|
|
760
|
-
```
|
|
761
|
-
|
|
762
|
-
**PowerRank**: Uses MemberJunction's model power rankings
|
|
763
|
-
```typescript
|
|
764
|
-
// Automatically selects models based on their PowerRank scores
|
|
765
|
-
// Ensures you always get the best available model
|
|
766
|
-
```
|
|
767
|
-
|
|
768
|
-
#### Error Scope Configuration
|
|
769
|
-
|
|
770
|
-
Control which types of errors trigger failover:
|
|
771
|
-
|
|
772
|
-
- **All**: Any error triggers failover (most resilient)
|
|
773
|
-
- **NetworkOnly**: Only network/connection errors
|
|
774
|
-
- **RateLimitOnly**: Only rate limit errors (429 status)
|
|
775
|
-
- **ServiceErrorOnly**: Only service errors (500, 503 status)
|
|
776
|
-
|
|
777
|
-
#### Failover Tracking
|
|
778
|
-
|
|
779
|
-
The system comprehensively tracks failover attempts in the database:
|
|
780
|
-
|
|
781
|
-
```typescript
|
|
782
|
-
// AIPromptRun entity tracking fields:
|
|
783
|
-
OriginalModelID: string // The initially selected model
|
|
784
|
-
OriginalRequestStartTime: Date // When the request started
|
|
785
|
-
FailoverAttempts: number // Number of failover attempts made
|
|
786
|
-
FailoverErrors: string (JSON) // Detailed error information for each attempt
|
|
787
|
-
FailoverDurations: string (JSON) // Duration of each attempt in milliseconds
|
|
788
|
-
TotalFailoverDuration: number // Total time spent in failover
|
|
789
|
-
```
|
|
790
|
-
|
|
791
|
-
#### Advanced Failover Customization
|
|
792
|
-
|
|
793
|
-
The AIPromptRunner exposes protected methods for advanced customization:
|
|
794
|
-
|
|
795
|
-
```typescript
|
|
796
|
-
class CustomPromptRunner extends AIPromptRunner {
|
|
797
|
-
// Override to implement custom failover configuration
|
|
798
|
-
protected getFailoverConfiguration(prompt: AIPromptEntity): FailoverConfiguration {
|
|
799
|
-
// Add environment-specific logic
|
|
800
|
-
if (process.env.NODE_ENV === 'production') {
|
|
801
|
-
return {
|
|
802
|
-
strategy: 'SameModelDifferentVendor',
|
|
803
|
-
maxAttempts: 5,
|
|
804
|
-
delaySeconds: 2,
|
|
805
|
-
modelStrategy: 'PreferSameModel',
|
|
806
|
-
errorScope: 'NetworkOnly'
|
|
807
|
-
};
|
|
808
|
-
}
|
|
809
|
-
return super.getFailoverConfiguration(prompt);
|
|
810
|
-
}
|
|
811
|
-
|
|
812
|
-
// Override to implement custom failover decision logic
|
|
813
|
-
protected shouldAttemptFailover(
|
|
814
|
-
error: Error,
|
|
815
|
-
config: FailoverConfiguration,
|
|
816
|
-
attemptNumber: number
|
|
817
|
-
): boolean {
|
|
818
|
-
// Add custom error analysis
|
|
819
|
-
if (error.message.includes('quota_exceeded')) {
|
|
820
|
-
return false; // Don't retry quota errors
|
|
821
|
-
}
|
|
822
|
-
return super.shouldAttemptFailover(error, config, attemptNumber);
|
|
823
|
-
}
|
|
824
|
-
|
|
825
|
-
// Override to implement custom delay calculation
|
|
826
|
-
protected calculateFailoverDelay(
|
|
827
|
-
attemptNumber: number,
|
|
828
|
-
baseDelaySeconds: number,
|
|
829
|
-
previousError?: Error
|
|
830
|
-
): number {
|
|
831
|
-
// Custom backoff strategy
|
|
832
|
-
if (previousError?.message.includes('rate_limit')) {
|
|
833
|
-
return 60000; // 1 minute for rate limits
|
|
834
|
-
}
|
|
835
|
-
return super.calculateFailoverDelay(attemptNumber, baseDelaySeconds, previousError);
|
|
836
|
-
}
|
|
837
|
-
}
|
|
838
|
-
```
|
|
839
|
-
|
|
840
|
-
#### Failover Best Practices
|
|
841
|
-
|
|
842
|
-
1. **Configure Appropriately**: Use `NetworkOnly` or `RateLimitOnly` for production to avoid retrying invalid requests
|
|
843
|
-
2. **Set Reasonable Attempts**: 3-5 attempts typically sufficient
|
|
844
|
-
3. **Monitor Failover Patterns**: Query the tracking data to identify problematic providers
|
|
845
|
-
4. **Test Failover Scenarios**: Simulate provider outages in development
|
|
846
|
-
5. **Consider Costs**: Failover may route to more expensive providers
|
|
847
|
-
|
|
848
|
-
#### Example: Production-Ready Configuration
|
|
849
|
-
|
|
850
|
-
```typescript
|
|
851
|
-
const productionPrompt = {
|
|
852
|
-
Name: "Customer Service Assistant",
|
|
853
|
-
FailoverStrategy: "SameModelDifferentVendor",
|
|
854
|
-
FailoverMaxAttempts: 4,
|
|
855
|
-
FailoverDelaySeconds: 2,
|
|
856
|
-
FailoverModelStrategy: "PreferSameModel",
|
|
857
|
-
FailoverErrorScope: "NetworkOnly",
|
|
858
|
-
// Ensure failover stays within approved models
|
|
859
|
-
MinPowerRank: 85
|
|
860
|
-
};
|
|
861
|
-
|
|
862
|
-
// Query failover performance
|
|
863
|
-
const failoverStats = await runView.RunView({
|
|
864
|
-
EntityName: 'MJ: AI Prompt Runs',
|
|
865
|
-
ExtraFilter: `FailoverAttempts > 0 AND RunAt >= '2024-01-01'`,
|
|
866
|
-
OrderBy: 'RunAt DESC'
|
|
867
|
-
});
|
|
868
|
-
|
|
869
|
-
// Analyze which vendors are most reliable
|
|
870
|
-
SELECT
|
|
871
|
-
OriginalModelID,
|
|
872
|
-
ModelID as FinalModelID,
|
|
873
|
-
COUNT(*) as FailoverCount,
|
|
874
|
-
AVG(TotalFailoverDuration) as AvgFailoverTime
|
|
875
|
-
FROM AIPromptRun
|
|
876
|
-
WHERE FailoverAttempts > 0
|
|
877
|
-
GROUP BY OriginalModelID, ModelID
|
|
878
|
-
ORDER BY FailoverCount DESC;
|
|
879
|
-
```
|
|
880
|
-
|
|
881
|
-
#### Configuration-Aware Failover
|
|
882
|
-
|
|
883
|
-
The failover system respects `AIConfiguration` boundaries to ensure environment-specific models stay isolated:
|
|
884
|
-
|
|
885
|
-
**How It Works:**
|
|
886
|
-
- When you specify a `configurationId`, the system builds a candidate list with two priority tiers:
|
|
887
|
-
1. **Configuration-specific models** (priority 5000+): Models assigned to your configuration
|
|
888
|
-
2. **NULL configuration models** (priority 2000+): Universal fallback models available to all configurations
|
|
889
|
-
|
|
890
|
-
**Example Setup:**
|
|
891
|
-
```sql
|
|
892
|
-
-- Production Configuration: Only approved production models
|
|
893
|
-
INSERT INTO AIPromptModel (PromptID, ModelID, ConfigurationID, Priority)
|
|
894
|
-
VALUES
|
|
895
|
-
(@PromptID, @Claude35SonnetID, @ProductionConfigID, 100),
|
|
896
|
-
(@PromptID, @GPT4ID, @ProductionConfigID, 90);
|
|
897
|
-
|
|
898
|
-
-- Development Configuration: Include experimental models
|
|
899
|
-
INSERT INTO AIPromptModel (PromptID, ModelID, ConfigurationID, Priority)
|
|
900
|
-
VALUES
|
|
901
|
-
(@PromptID, @LlamaExperimentalID, @DevelopmentConfigID, 100);
|
|
902
|
-
|
|
903
|
-
-- NULL Configuration: Universal fallbacks for all environments
|
|
904
|
-
INSERT INTO AIPromptModel (PromptID, ModelID, ConfigurationID, Priority)
|
|
905
|
-
VALUES
|
|
906
|
-
(@PromptID, @Claude3HaikuID, NULL, 100),
|
|
907
|
-
(@PromptID, @GPT35TurboID, NULL, 90);
|
|
908
|
-
```
|
|
909
|
-
|
|
910
|
-
**Failover Behavior:**
|
|
911
|
-
```typescript
|
|
912
|
-
// Execute with Production configuration
|
|
913
|
-
const result = await runner.ExecutePrompt({
|
|
914
|
-
prompt: myPrompt,
|
|
915
|
-
configurationId: productionConfigID,
|
|
916
|
-
data: { query: 'Analyze this' }
|
|
917
|
-
});
|
|
918
|
-
|
|
919
|
-
// Failover order:
|
|
920
|
-
// 1. Try Claude 3.5 Sonnet (Production config, priority 5100)
|
|
921
|
-
// 2. Try GPT-4 (Production config, priority 5090)
|
|
922
|
-
// 3. Try Claude 3 Haiku (NULL config fallback, priority 2100)
|
|
923
|
-
// 4. Try GPT-3.5 Turbo (NULL config fallback, priority 2090)
|
|
924
|
-
// ✅ Never crosses to Development config models
|
|
925
|
-
```
|
|
926
|
-
|
|
927
|
-
**Key Benefits:**
|
|
928
|
-
- **Environment Isolation**: Production models never failover to development/experimental models
|
|
929
|
-
- **Controlled Fallback**: Explicit hierarchy from config-specific to universal fallbacks
|
|
930
|
-
- **Performance**: Candidate list built once and cached, no rebuilding during failover
|
|
931
|
-
- **Consistency**: Same candidate list used for initial selection and all failover attempts
|
|
932
|
-
|
|
933
|
-
### Intelligent Caching
|
|
934
|
-
|
|
935
|
-
The prompt system provides sophisticated caching with vector similarity matching:
|
|
936
|
-
|
|
937
|
-
```typescript
|
|
938
|
-
// Caching is automatically handled based on prompt configuration:
|
|
939
|
-
// - EnableCaching: Whether to use caching for this prompt
|
|
940
|
-
// - CacheMatchType: 'Exact' or 'Vector' similarity matching
|
|
941
|
-
// - CacheTTLSeconds: Time-to-live for cached results
|
|
942
|
-
// - CacheMustMatchModel/Vendor/Agent: Cache constraint options
|
|
943
|
-
|
|
944
|
-
// Vector similarity allows reusing results for semantically similar prompts
|
|
945
|
-
// even if the exact text differs
|
|
946
|
-
|
|
947
|
-
const cachedPrompt = {
|
|
948
|
-
Name: "Smart Summary",
|
|
949
|
-
EnableCaching: true,
|
|
950
|
-
CacheMatchType: "Vector",
|
|
951
|
-
CacheTTLSeconds: 3600,
|
|
952
|
-
CacheSimilarityThreshold: 0.85,
|
|
953
|
-
CacheMustMatchModel: true,
|
|
954
|
-
CacheMustMatchVendor: false
|
|
955
|
-
};
|
|
956
|
-
```
|
|
957
|
-
|
|
958
|
-
### Parallel Execution Strategies
|
|
959
|
-
|
|
960
|
-
The system supports multiple parallelization modes:
|
|
961
|
-
|
|
962
|
-
```typescript
|
|
963
|
-
// Prompts can be configured for parallel execution:
|
|
964
|
-
// - ParallelizationMode: 'None', 'StaticCount', 'ConfigParam', 'ModelSpecific'
|
|
965
|
-
// - ParallelCount: Number of parallel executions
|
|
966
|
-
// - ExecutionGroups: Sequential group execution with parallel tasks within groups
|
|
967
|
-
|
|
968
|
-
// Example configurations:
|
|
969
|
-
|
|
970
|
-
// Static parallel count
|
|
971
|
-
const staticParallelPrompt = {
|
|
972
|
-
ParallelizationMode: "StaticCount",
|
|
973
|
-
ParallelCount: 3
|
|
974
|
-
};
|
|
975
|
-
|
|
976
|
-
// Configuration-driven count
|
|
977
|
-
const configParallelPrompt = {
|
|
978
|
-
ParallelizationMode: "ConfigParam",
|
|
979
|
-
ParallelConfigParam: "analysis_parallel_count"
|
|
980
|
-
};
|
|
981
|
-
|
|
982
|
-
// Model-specific configuration
|
|
983
|
-
const modelSpecificPrompt = {
|
|
984
|
-
ParallelizationMode: "ModelSpecific",
|
|
985
|
-
// Uses settings from AIPromptModel entries
|
|
986
|
-
};
|
|
987
|
-
```
|
|
988
|
-
|
|
989
|
-
### Result Selection Strategies
|
|
990
|
-
|
|
991
|
-
```typescript
|
|
992
|
-
// The engine supports multiple result selection methods:
|
|
993
|
-
// - 'First': Use the first successful result
|
|
994
|
-
// - 'Random': Randomly select from successful results
|
|
995
|
-
// - 'PromptSelector': Use AI to select the best result
|
|
996
|
-
// - 'Consensus': Select result with highest agreement
|
|
997
|
-
|
|
998
|
-
// Result selector prompts can be configured to intelligently choose
|
|
999
|
-
// the best result from parallel executions
|
|
1000
|
-
|
|
1001
|
-
const selectorPrompt = {
|
|
1002
|
-
Name: "Best Result Selector",
|
|
1003
|
-
PromptText: `
|
|
1004
|
-
You are evaluating multiple AI responses to select the best one.
|
|
1005
|
-
Original query: {{originalQuery}}
|
|
1006
|
-
|
|
1007
|
-
Responses:
|
|
1008
|
-
{{#each responses}}
|
|
1009
|
-
Response {{@index}}: {{this}}
|
|
1010
|
-
{{/each}}
|
|
1011
|
-
|
|
1012
|
-
Select the response number (0-based) that is most accurate, helpful, and well-written.
|
|
1013
|
-
Return only the number.
|
|
1014
|
-
`,
|
|
1015
|
-
OutputType: "number"
|
|
1016
|
-
};
|
|
1017
|
-
|
|
1018
|
-
const mainPrompt = {
|
|
1019
|
-
ParallelizationMode: "StaticCount",
|
|
1020
|
-
ParallelCount: 3,
|
|
1021
|
-
ResultSelectorPromptID: selectorPrompt.ID
|
|
1022
|
-
};
|
|
1023
|
-
```
|
|
1024
|
-
|
|
1025
|
-
### Output Validation
|
|
1026
|
-
|
|
1027
|
-
```typescript
|
|
1028
|
-
// Configure structured output validation
|
|
1029
|
-
const validatedPrompt = {
|
|
1030
|
-
Name: "Structured Analysis",
|
|
1031
|
-
OutputType: "object",
|
|
1032
|
-
OutputExample: {
|
|
1033
|
-
sentiment: "positive|negative|neutral",
|
|
1034
|
-
confidence: 0.95,
|
|
1035
|
-
keyThemes: ["theme1", "theme2"],
|
|
1036
|
-
summary: "Brief summary text"
|
|
1037
|
-
},
|
|
1038
|
-
ValidationBehavior: "Strict",
|
|
1039
|
-
MaxRetries: 3,
|
|
1040
|
-
RetryDelayMS: 1000,
|
|
1041
|
-
RetryStrategy: "exponential"
|
|
1042
|
-
};
|
|
1043
|
-
|
|
1044
|
-
// Validation is automatically applied
|
|
1045
|
-
const result = await runner.ExecutePrompt({
|
|
1046
|
-
prompt: validatedPrompt,
|
|
1047
|
-
data: { text: "Content to analyze" },
|
|
1048
|
-
contextUser: currentUser,
|
|
1049
|
-
skipValidation: false // Validation enabled
|
|
1050
|
-
});
|
|
1051
|
-
|
|
1052
|
-
// Result.result will be validated against the expected structure
|
|
1053
|
-
```
|
|
1054
|
-
|
|
1055
|
-
### Validation Syntax Cleaning
|
|
1056
|
-
|
|
1057
|
-
When using output validation with JSON responses, the AI Prompt Runner automatically handles validation syntax that AI models might inadvertently include in their JSON keys:
|
|
196
|
+
const runner = new AIPromptRunner();
|
|
197
|
+
const params = new AIPromptParams();
|
|
198
|
+
params.prompt = prompt;
|
|
199
|
+
params.data = { content: documentText, maxLength: 500 };
|
|
200
|
+
params.contextUser = contextUser;
|
|
1058
201
|
|
|
1059
|
-
|
|
1060
|
-
// Validation syntax in prompts:
|
|
1061
|
-
// - name?: optional field
|
|
1062
|
-
// - items:[2+]: array with minimum 2 items
|
|
1063
|
-
// - status:!empty: non-empty required field
|
|
1064
|
-
// - count:number: field with type hint
|
|
1065
|
-
|
|
1066
|
-
// If the AI returns JSON with validation syntax in keys:
|
|
1067
|
-
{
|
|
1068
|
-
"name?": "John Doe",
|
|
1069
|
-
"items:[2+]": ["apple", "banana", "orange"],
|
|
1070
|
-
"status:!empty": "active",
|
|
1071
|
-
"count:number": 42
|
|
1072
|
-
}
|
|
202
|
+
const result = await runner.ExecutePrompt(params);
|
|
1073
203
|
|
|
1074
|
-
|
|
1075
|
-
|
|
1076
|
-
|
|
1077
|
-
|
|
1078
|
-
|
|
1079
|
-
"count": 42
|
|
204
|
+
if (result.success) {
|
|
205
|
+
console.log(result.result); // Parsed/validated result
|
|
206
|
+
console.log(result.promptTokens); // Input tokens used
|
|
207
|
+
console.log(result.completionTokens); // Output tokens generated
|
|
208
|
+
console.log(result.executionTimeMS); // Execution duration
|
|
1080
209
|
}
|
|
1081
210
|
```
|
|
1082
211
|
|
|
1083
|
-
|
|
1084
|
-
- **Always enabled** when prompt has `ValidationBehavior` set to `"Strict"` or `"Warn"`
|
|
1085
|
-
- **Always enabled** when prompt has an `OutputExample` defined
|
|
1086
|
-
- **Optional** for prompts with `ValidationBehavior` set to `"None"` (via `cleanValidationSyntax` parameter)
|
|
1087
|
-
|
|
1088
|
-
This ensures that validation patterns used in prompt templates don't interfere with the actual JSON structure returned by the AI model.
|
|
1089
|
-
|
|
1090
|
-
```
|
|
1091
|
-
|
|
1092
|
-
### Template Integration
|
|
1093
|
-
|
|
1094
|
-
Advanced template features with the MemberJunction template system:
|
|
1095
|
-
|
|
1096
|
-
```typescript
|
|
1097
|
-
// Complex template with conditionals and loops
|
|
1098
|
-
const advancedTemplate = {
|
|
1099
|
-
PromptText: `
|
|
1100
|
-
Analyze the following {{entityType}} records:
|
|
1101
|
-
|
|
1102
|
-
{{#each records}}
|
|
1103
|
-
{{@index + 1}}. {{this.Name}}
|
|
1104
|
-
Status: {{this.Status}}
|
|
1105
|
-
{{#if this.Priority}}Priority: {{this.Priority}}{{/if}}
|
|
1106
|
-
{{#each this.Tags}}
|
|
1107
|
-
- Tag: {{this}}
|
|
1108
|
-
{{/each}}
|
|
1109
|
-
{{/each}}
|
|
1110
|
-
|
|
1111
|
-
{{#if includeRecommendations}}
|
|
1112
|
-
Please provide recommendations for improvement.
|
|
1113
|
-
{{/if}}
|
|
1114
|
-
|
|
1115
|
-
Focus on: {{analysisAreas.join(", ")}}
|
|
1116
|
-
`
|
|
1117
|
-
};
|
|
1118
|
-
|
|
1119
|
-
const result = await runner.ExecutePrompt({
|
|
1120
|
-
prompt: advancedTemplate,
|
|
1121
|
-
data: {
|
|
1122
|
-
entityType: "Customer",
|
|
1123
|
-
records: customerData,
|
|
1124
|
-
includeRecommendations: true,
|
|
1125
|
-
analysisAreas: ["revenue potential", "risk factors", "engagement"]
|
|
1126
|
-
},
|
|
1127
|
-
contextUser: currentUser
|
|
1128
|
-
});
|
|
1129
|
-
```
|
|
1130
|
-
|
|
1131
|
-
## Parallel Execution System
|
|
1132
|
-
|
|
1133
|
-
The package includes sophisticated parallel execution capabilities through specialized classes that work together to manage complex multi-model executions.
|
|
1134
|
-
|
|
1135
|
-
> **Note**: The ExecutionPlanner and ParallelExecutionCoordinator are internal components used by AIPromptRunner. They are not directly exposed in the public API but understanding their operation helps in configuring prompts effectively.
|
|
1136
|
-
|
|
1137
|
-
### ExecutionPlanner (Internal)
|
|
1138
|
-
|
|
1139
|
-
The `ExecutionPlanner` class analyzes prompt configuration and creates optimal execution strategies:
|
|
1140
|
-
|
|
1141
|
-
**Key Responsibilities:**
|
|
1142
|
-
- Analyzes parallelization modes (None, StaticCount, ConfigParam, ModelSpecific)
|
|
1143
|
-
- Creates execution groups for coordinated processing
|
|
1144
|
-
- Determines optimal task distribution based on model availability
|
|
1145
|
-
- Assigns priorities and manages execution order
|
|
1146
|
-
- Handles model selection based on power rankings and configuration
|
|
1147
|
-
|
|
1148
|
-
**Execution Plan Creation:**
|
|
1149
|
-
- For `StaticCount`: Creates N parallel tasks using available models
|
|
1150
|
-
- For `ConfigParam`: Uses configuration parameters to determine parallel count
|
|
1151
|
-
- For `ModelSpecific`: Uses AIPromptModel entries to define exact model usage
|
|
1152
|
-
- Supports execution groups for sequential/parallel hybrid execution
|
|
1153
|
-
|
|
1154
|
-
### ParallelExecutionCoordinator (Internal)
|
|
1155
|
-
|
|
1156
|
-
The `ParallelExecutionCoordinator` orchestrates the actual execution of tasks created by the ExecutionPlanner:
|
|
1157
|
-
|
|
1158
|
-
**Core Features:**
|
|
1159
|
-
- Manages concurrency limits (default: 5 concurrent executions)
|
|
1160
|
-
- Implements retry logic with exponential backoff
|
|
1161
|
-
- Handles partial result collection when some tasks fail
|
|
1162
|
-
- Provides comprehensive execution metrics and timing
|
|
1163
|
-
- Supports fail-fast mode for critical operations
|
|
1164
|
-
|
|
1165
|
-
**Execution Flow:**
|
|
1166
|
-
1. Groups tasks by execution group number
|
|
1167
|
-
2. Executes groups sequentially (group 0, then 1, then 2, etc.)
|
|
1168
|
-
3. Within each group, executes tasks in parallel up to concurrency limit
|
|
1169
|
-
4. Collects and aggregates results from all executions
|
|
1170
|
-
5. Applies result selection strategy if multiple results available
|
|
1171
|
-
|
|
1172
|
-
### Supported Parallelization Modes
|
|
1173
|
-
|
|
1174
|
-
- **None**: Traditional single execution
|
|
1175
|
-
- **StaticCount**: Fixed number of parallel executions
|
|
1176
|
-
- **ConfigParam**: Dynamic parallel count from configuration
|
|
1177
|
-
- **ModelSpecific**: Individual model configurations with execution groups
|
|
212
|
+
### With Progress Tracking
|
|
1178
213
|
|
|
1179
214
|
```typescript
|
|
1180
|
-
|
|
1181
|
-
|
|
1182
|
-
prompt: complexPrompt,
|
|
1183
|
-
data: analysisData,
|
|
1184
|
-
contextUser: currentUser
|
|
215
|
+
params.onProgress = (progress) => {
|
|
216
|
+
console.log(`[${progress.step}] ${progress.percentage}% - ${progress.message}`);
|
|
1185
217
|
};
|
|
1186
|
-
|
|
1187
|
-
// The system will:
|
|
1188
|
-
// 1. Query AIPromptModel entries for this prompt
|
|
1189
|
-
// 2. Group executions by ExecutionGroup
|
|
1190
|
-
// 3. Execute groups sequentially, models within groups in parallel
|
|
1191
|
-
// 4. Apply result selection strategy
|
|
1192
|
-
const result = await runner.ExecutePrompt(modelSpecificExecution);
|
|
1193
218
|
```
|
|
1194
219
|
|
|
1195
|
-
|
|
1196
|
-
|
|
1197
|
-
Comprehensive tracking and analytics for prompt executions:
|
|
220
|
+
### With Effort Level
|
|
1198
221
|
|
|
1199
222
|
```typescript
|
|
1200
|
-
//
|
|
1201
|
-
const result = await runner.ExecutePrompt(params);
|
|
1202
|
-
|
|
1203
|
-
console.log(`Execution time: ${result.executionTimeMS}ms`);
|
|
1204
|
-
console.log(`Tokens used: ${result.tokensUsed}`);
|
|
1205
|
-
|
|
1206
|
-
// The AIPromptRunResult includes execution tracking
|
|
1207
|
-
if (result.promptRun) {
|
|
1208
|
-
console.log(`Prompt Run ID: ${result.promptRun.ID}`);
|
|
1209
|
-
console.log(`Model used: ${result.promptRun.ModelID}`);
|
|
1210
|
-
console.log(`Configuration: ${result.promptRun.ConfigurationID}`);
|
|
1211
|
-
}
|
|
223
|
+
params.effortLevel = 85; // High effort for thorough analysis (1-100 scale)
|
|
1212
224
|
```
|
|
1213
225
|
|
|
1214
|
-
###
|
|
1215
|
-
|
|
1216
|
-
Get the PromptRun ID immediately after creation for real-time monitoring:
|
|
226
|
+
### With Runtime Model Override
|
|
1217
227
|
|
|
1218
228
|
```typescript
|
|
1219
|
-
|
|
1220
|
-
|
|
1221
|
-
|
|
1222
|
-
|
|
1223
|
-
// Callback fired immediately after PromptRun record is saved
|
|
1224
|
-
params.onPromptRunCreated = async (promptRunId) => {
|
|
1225
|
-
console.log(`Prompt run started: ${promptRunId}`);
|
|
1226
|
-
|
|
1227
|
-
// Use cases:
|
|
1228
|
-
// - Link to parent records (e.g., AIAgentRunStep.TargetLogID)
|
|
1229
|
-
// - Send to monitoring systems
|
|
1230
|
-
// - Update UI with tracking info
|
|
1231
|
-
// - Start real-time log streaming
|
|
229
|
+
params.override = {
|
|
230
|
+
modelId: 'specific-model-id',
|
|
231
|
+
vendorId: 'specific-vendor-id'
|
|
1232
232
|
};
|
|
1233
|
-
|
|
1234
|
-
const result = await runner.ExecutePrompt(params);
|
|
1235
233
|
```
|
|
1236
234
|
|
|
1237
|
-
The callback is invoked:
|
|
1238
|
-
- **When**: Right after the AIPromptRun record is created and saved
|
|
1239
|
-
- **Before**: The actual AI model execution begins
|
|
1240
|
-
- **Error Handling**: Callback errors are logged but don't fail the execution
|
|
1241
|
-
- **Async Support**: Can be synchronous or asynchronous
|
|
1242
|
-
|
|
1243
|
-
## AI Prompt Run Logging
|
|
1244
|
-
|
|
1245
|
-
The AI Prompt Runner implements a sophisticated hierarchical logging system that tracks all execution activities in the database through the `AIPromptRun` entity. This system provides complete traceability and analytics for both simple and complex parallel executions.
|
|
1246
|
-
|
|
1247
|
-
### Hierarchical Logging Structure
|
|
1248
|
-
|
|
1249
|
-
The logging system uses a parent-child relationship model with different `RunType` values to represent the execution hierarchy:
|
|
1250
|
-
|
|
1251
|
-
- **`Single`**: Standard single-model execution
|
|
1252
|
-
- **`ParallelParent`**: Parent record for parallel execution coordinating multiple models
|
|
1253
|
-
- **`ParallelChild`**: Individual model execution within a parallel run
|
|
1254
|
-
- **`ResultSelector`**: AI judge execution that selects the best result from parallel executions
|
|
1255
|
-
|
|
1256
|
-
### RunType Values and Relationships
|
|
1257
|
-
|
|
1258
|
-
```typescript
|
|
1259
|
-
// Single execution - no parent relationship
|
|
1260
|
-
{
|
|
1261
|
-
RunType: 'Single',
|
|
1262
|
-
ParentID: null,
|
|
1263
|
-
ExecutionOrder: null
|
|
1264
|
-
}
|
|
1265
|
-
|
|
1266
|
-
// Parallel execution creates a hierarchical structure:
|
|
1267
|
-
// 1. Parent record coordinates the overall execution
|
|
1268
|
-
{
|
|
1269
|
-
RunType: 'ParallelParent',
|
|
1270
|
-
ParentID: null,
|
|
1271
|
-
ExecutionOrder: null
|
|
1272
|
-
}
|
|
1273
|
-
|
|
1274
|
-
// 2. Child records for each model execution
|
|
1275
|
-
{
|
|
1276
|
-
RunType: 'ParallelChild',
|
|
1277
|
-
ParentID: '12345-parent-id',
|
|
1278
|
-
ExecutionOrder: 0 // Order within execution group
|
|
1279
|
-
}
|
|
1280
|
-
|
|
1281
|
-
// 3. Result selector judges the best result
|
|
1282
|
-
{
|
|
1283
|
-
RunType: 'ResultSelector',
|
|
1284
|
-
ParentID: '12345-parent-id',
|
|
1285
|
-
ExecutionOrder: 5 // After all parallel children
|
|
1286
|
-
}
|
|
1287
|
-
```
|
|
1288
|
-
|
|
1289
|
-
### Database Schema Fields
|
|
1290
|
-
|
|
1291
|
-
Key fields in the `AIPromptRun` entity for hierarchical logging:
|
|
1292
|
-
|
|
1293
|
-
```sql
|
|
1294
|
-
-- Core execution tracking
|
|
1295
|
-
PromptID uniqueidentifier -- Prompt being executed
|
|
1296
|
-
ModelID uniqueidentifier -- AI model used
|
|
1297
|
-
VendorID uniqueidentifier -- Vendor providing the model
|
|
1298
|
-
RunAt datetime2 -- Execution start time
|
|
1299
|
-
CompletedAt datetime2 -- Execution completion time
|
|
1300
|
-
|
|
1301
|
-
-- Hierarchical logging fields
|
|
1302
|
-
RunType nvarchar(50) -- 'Single', 'ParallelParent', 'ParallelChild', 'ResultSelector'
|
|
1303
|
-
ParentID uniqueidentifier -- Parent prompt run ID (NULL for top-level)
|
|
1304
|
-
ExecutionOrder int -- Order within parallel execution group
|
|
1305
|
-
|
|
1306
|
-
-- Results and metrics
|
|
1307
|
-
Success bit -- Whether execution succeeded
|
|
1308
|
-
Result nvarchar(max) -- Raw result from AI model
|
|
1309
|
-
ErrorMessage nvarchar(500) -- Error message if failed
|
|
1310
|
-
ExecutionTimeMS int -- Total execution time
|
|
1311
|
-
TokensUsed int -- Total tokens consumed
|
|
1312
|
-
TokensPrompt int -- Prompt tokens used
|
|
1313
|
-
TokensCompletion int -- Completion tokens generated
|
|
1314
|
-
|
|
1315
|
-
-- Cost tracking
|
|
1316
|
-
Cost decimal(19,8) -- Cost of this specific execution
|
|
1317
|
-
CostCurrency nvarchar(10) -- ISO 4217 currency code (USD, EUR, etc.)
|
|
1318
|
-
|
|
1319
|
-
-- Hierarchical rollup fields (NEW)
|
|
1320
|
-
TokensUsedRollup int -- Total tokens including all children
|
|
1321
|
-
TokensPromptRollup int -- Total prompt tokens including all children
|
|
1322
|
-
TokensCompletionRollup int -- Total completion tokens including all children
|
|
1323
|
-
-- Note: TotalCost (existing field) serves as the cost rollup
|
|
1324
|
-
|
|
1325
|
-
-- Context and configuration
|
|
1326
|
-
Messages nvarchar(max) -- JSON with input data and metadata
|
|
1327
|
-
ConfigurationID uniqueidentifier -- Environment configuration used
|
|
1328
|
-
AgentRunID uniqueidentifier -- Links to parent AIAgentRun if applicable
|
|
1329
|
-
```
|
|
1330
|
-
|
|
1331
|
-
### Hierarchical Token and Cost Tracking
|
|
1332
|
-
|
|
1333
|
-
The AI Prompts system implements a sophisticated rollup pattern for tracking token usage and costs across hierarchical prompt executions:
|
|
1334
|
-
|
|
1335
|
-
#### Prompt Execution Rollup Pattern
|
|
1336
|
-
|
|
1337
|
-
For hierarchical prompt executions (parent prompts with child prompts), each node in the tree contains:
|
|
1338
|
-
- **Direct fields** (`TokensPrompt`, `TokensCompletion`, `Cost`): Usage for just that execution
|
|
1339
|
-
- **Rollup fields** (`TokensPromptRollup`, `TokensCompletionRollup`, `TotalCost`): Total including all descendants
|
|
1340
|
-
|
|
1341
|
-
**Example:**
|
|
1342
|
-
```
|
|
1343
|
-
Parent Prompt (100 prompt, 200 completion tokens, $0.05)
|
|
1344
|
-
├── Child A (50 prompt, 100 completion, $0.02)
|
|
1345
|
-
└── Child B (75 prompt, 150 completion, $0.03)
|
|
1346
|
-
|
|
1347
|
-
Database records:
|
|
1348
|
-
- Parent: TokensPrompt=100, TokensPromptRollup=225 (100+50+75)
|
|
1349
|
-
TokensCompletion=200, TokensCompletionRollup=450 (200+100+150)
|
|
1350
|
-
Cost=0.05, TotalCost=0.10 (0.05+0.02+0.03)
|
|
1351
|
-
- Child A: TokensPrompt=50, TokensPromptRollup=50 (leaf node)
|
|
1352
|
-
Cost=0.02, TotalCost=0.02 (leaf node)
|
|
1353
|
-
- Child B: TokensPrompt=75, TokensPromptRollup=75 (leaf node)
|
|
1354
|
-
Cost=0.03, TotalCost=0.03 (leaf node)
|
|
1355
|
-
```
|
|
1356
|
-
|
|
1357
|
-
This enables efficient queries like:
|
|
1358
|
-
- "What was the total cost of this hierarchical prompt?" → Check root's `TotalCost`
|
|
1359
|
-
- "How many tokens did this sub-prompt and its children use?" → Check that node's rollup fields
|
|
1360
|
-
- No complex SQL joins or recursive CTEs needed!
|
|
1361
|
-
|
|
1362
|
-
#### Agent Run Token Tracking
|
|
1363
|
-
|
|
1364
|
-
The `AIAgentRun` entity tracks aggregate token usage across all prompt executions during an agent's lifecycle:
|
|
1365
|
-
|
|
1366
|
-
```sql
|
|
1367
|
-
-- New fields in AIAgentRun
|
|
1368
|
-
TotalTokensUsed int -- Total tokens (existing)
|
|
1369
|
-
TotalPromptTokensUsed int -- Breakdown: prompt tokens (NEW)
|
|
1370
|
-
TotalCompletionTokensUsed int -- Breakdown: completion tokens (NEW)
|
|
1371
|
-
TotalCost decimal -- Total cost (existing)
|
|
1372
|
-
|
|
1373
|
-
-- Hierarchical agent rollup fields (NEW)
|
|
1374
|
-
TotalTokensUsedRollup int -- Including sub-agent runs
|
|
1375
|
-
TotalPromptTokensUsedRollup int -- Including sub-agent runs
|
|
1376
|
-
TotalCompletionTokensUsedRollup int -- Including sub-agent runs
|
|
1377
|
-
TotalCostRollup decimal -- Including sub-agent runs
|
|
1378
|
-
```
|
|
1379
|
-
|
|
1380
|
-
**Agent Hierarchy Example:**
|
|
1381
|
-
```
|
|
1382
|
-
Parent Agent (A)
|
|
1383
|
-
├── Own prompts: 200 prompt, 400 completion tokens
|
|
1384
|
-
├── Sub-Agent (B)
|
|
1385
|
-
│ └── Own prompts: 100 prompt, 200 completion tokens
|
|
1386
|
-
└── Sub-Agent (C)
|
|
1387
|
-
└── Own prompts: 150 prompt, 300 completion tokens
|
|
1388
|
-
|
|
1389
|
-
Rollup values:
|
|
1390
|
-
- Agent A: TotalPromptTokensUsedRollup = 450 (200+100+150)
|
|
1391
|
-
TotalCompletionTokensUsedRollup = 900 (400+200+300)
|
|
1392
|
-
- Agent B: TotalPromptTokensUsedRollup = 100 (leaf agent)
|
|
1393
|
-
- Agent C: TotalPromptTokensUsedRollup = 150 (leaf agent)
|
|
1394
|
-
```
|
|
1395
|
-
|
|
1396
|
-
### Querying Hierarchical Log Data
|
|
1397
|
-
|
|
1398
|
-
The hierarchical structure enables powerful analytics queries:
|
|
1399
|
-
|
|
1400
|
-
```sql
|
|
1401
|
-
-- Get all executions for a parallel run
|
|
1402
|
-
SELECT
|
|
1403
|
-
pr.ID,
|
|
1404
|
-
pr.RunType,
|
|
1405
|
-
pr.ExecutionOrder,
|
|
1406
|
-
pr.Success,
|
|
1407
|
-
pr.ExecutionTimeMS,
|
|
1408
|
-
pr.TokensUsed,
|
|
1409
|
-
m.Name as ModelName,
|
|
1410
|
-
p.Name as PromptName
|
|
1411
|
-
FROM AIPromptRun pr
|
|
1412
|
-
JOIN AIModel m ON pr.ModelID = m.ID
|
|
1413
|
-
JOIN AIPrompt p ON pr.PromptID = p.ID
|
|
1414
|
-
WHERE pr.ParentID = '12345-parent-id'
|
|
1415
|
-
OR pr.ID = '12345-parent-id'
|
|
1416
|
-
ORDER BY pr.RunType, pr.ExecutionOrder;
|
|
1417
|
-
|
|
1418
|
-
-- Analyze parallel execution performance
|
|
1419
|
-
WITH ParallelStats AS (
|
|
1420
|
-
SELECT
|
|
1421
|
-
ParentID,
|
|
1422
|
-
COUNT(*) as TotalChildren,
|
|
1423
|
-
SUM(CASE WHEN Success = 1 THEN 1 ELSE 0 END) as SuccessfulChildren,
|
|
1424
|
-
AVG(ExecutionTimeMS) as AvgExecutionTime,
|
|
1425
|
-
SUM(TokensUsed) as TotalTokens
|
|
1426
|
-
FROM AIPromptRun
|
|
1427
|
-
WHERE RunType = 'ParallelChild'
|
|
1428
|
-
AND ParentID IS NOT NULL
|
|
1429
|
-
GROUP BY ParentID
|
|
1430
|
-
)
|
|
1431
|
-
SELECT
|
|
1432
|
-
parent.ID as ParentRunID,
|
|
1433
|
-
parent.RunAt,
|
|
1434
|
-
parent.ExecutionTimeMS as ParentExecutionTime,
|
|
1435
|
-
stats.TotalChildren,
|
|
1436
|
-
stats.SuccessfulChildren,
|
|
1437
|
-
stats.AvgExecutionTime,
|
|
1438
|
-
stats.TotalTokens,
|
|
1439
|
-
prompt.Name as PromptName
|
|
1440
|
-
FROM AIPromptRun parent
|
|
1441
|
-
JOIN ParallelStats stats ON parent.ID = stats.ParentID
|
|
1442
|
-
JOIN AIPrompt prompt ON parent.PromptID = prompt.ID
|
|
1443
|
-
WHERE parent.RunType = 'ParallelParent'
|
|
1444
|
-
ORDER BY parent.RunAt DESC;
|
|
1445
|
-
|
|
1446
|
-
-- Find failed executions with context
|
|
1447
|
-
SELECT
|
|
1448
|
-
pr.ID,
|
|
1449
|
-
pr.RunType,
|
|
1450
|
-
pr.ParentID,
|
|
1451
|
-
pr.ErrorMessage,
|
|
1452
|
-
pr.ExecutionTimeMS,
|
|
1453
|
-
m.Name as ModelName,
|
|
1454
|
-
v.Name as VendorName,
|
|
1455
|
-
p.Name as PromptName
|
|
1456
|
-
FROM AIPromptRun pr
|
|
1457
|
-
JOIN AIModel m ON pr.ModelID = m.ID
|
|
1458
|
-
LEFT JOIN AIVendor v ON pr.VendorID = v.ID
|
|
1459
|
-
JOIN AIPrompt p ON pr.PromptID = p.ID
|
|
1460
|
-
WHERE pr.Success = 0
|
|
1461
|
-
ORDER BY pr.RunAt DESC;
|
|
1462
|
-
```
|
|
1463
|
-
|
|
1464
|
-
## Cancellation Support
|
|
1465
|
-
|
|
1466
|
-
The AI Prompt Runner provides comprehensive cancellation support through the standard JavaScript `AbortSignal` and `AbortController` pattern, enabling graceful termination of long-running operations.
|
|
1467
|
-
|
|
1468
|
-
### Understanding AbortSignal in Prompt Execution
|
|
1469
|
-
|
|
1470
|
-
The `AbortSignal` pattern separates **cancellation control** from **cancellation handling**:
|
|
1471
|
-
|
|
1472
|
-
- **Your Code (Controller)**: Creates the `AbortController` and decides **when** to cancel
|
|
1473
|
-
- **Prompt Runner (Worker)**: Receives the `AbortSignal` token and handles **how** to cancel gracefully
|
|
1474
|
-
|
|
1475
|
-
This separation allows for flexible cancellation from multiple sources (user actions, timeouts, resource limits) while the Prompt Runner handles the complex cleanup across parallel executions, model calls, and result selection.
|
|
1476
|
-
|
|
1477
|
-
**The Pattern Flow:**
|
|
1478
|
-
```
|
|
1479
|
-
Controller (Your Code) → AbortController.signal → AIPromptRunner
|
|
1480
|
-
↓ ↓ ↓
|
|
1481
|
-
Decides WHEN The "Red Phone" Handles HOW
|
|
1482
|
-
to cancel Token to stop
|
|
1483
|
-
```
|
|
1484
|
-
|
|
1485
|
-
### Basic Cancellation Usage
|
|
1486
|
-
|
|
1487
|
-
```typescript
|
|
1488
|
-
import { AIPromptRunner } from '@memberjunction/ai-prompts';
|
|
1489
|
-
|
|
1490
|
-
// Create cancellation controller
|
|
1491
|
-
const controller = new AbortController();
|
|
1492
|
-
const cancellationToken = controller.signal;
|
|
1493
|
-
|
|
1494
|
-
// Set up cancellation after 30 seconds
|
|
1495
|
-
setTimeout(() => {
|
|
1496
|
-
controller.abort();
|
|
1497
|
-
console.log('Prompt execution cancelled due to timeout');
|
|
1498
|
-
}, 30000);
|
|
1499
|
-
|
|
1500
|
-
// Execute prompt with cancellation support
|
|
1501
|
-
const runner = new AIPromptRunner();
|
|
1502
|
-
const result = await runner.ExecutePrompt({
|
|
1503
|
-
prompt: myPrompt,
|
|
1504
|
-
data: { query: 'Long running analysis...' },
|
|
1505
|
-
contextUser: currentUser,
|
|
1506
|
-
cancellationToken: cancellationToken
|
|
1507
|
-
});
|
|
1508
|
-
|
|
1509
|
-
// Check if execution was cancelled
|
|
1510
|
-
if (result.cancelled) {
|
|
1511
|
-
console.log(`Execution cancelled: ${result.cancellationReason}`);
|
|
1512
|
-
console.log('Partial results may be available');
|
|
1513
|
-
} else if (result.success) {
|
|
1514
|
-
console.log('Execution completed successfully');
|
|
1515
|
-
}
|
|
1516
|
-
```
|
|
1517
|
-
|
|
1518
|
-
### Cancellation in Parallel Execution
|
|
1519
|
-
|
|
1520
|
-
Cancellation works seamlessly with parallel execution, allowing you to stop all running tasks:
|
|
1521
|
-
|
|
1522
|
-
```typescript
|
|
1523
|
-
const controller = new AbortController();
|
|
1524
|
-
|
|
1525
|
-
// User clicks cancel button
|
|
1526
|
-
document.getElementById('cancelButton').onclick = () => {
|
|
1527
|
-
controller.abort();
|
|
1528
|
-
};
|
|
1529
|
-
|
|
1530
|
-
// Execute parallel prompt with multiple models
|
|
1531
|
-
const result = await runner.ExecutePrompt({
|
|
1532
|
-
prompt: parallelPrompt, // ParallelizationMode: 'ModelSpecific'
|
|
1533
|
-
data: analysisData,
|
|
1534
|
-
contextUser: currentUser,
|
|
1535
|
-
cancellationToken: controller.signal
|
|
1536
|
-
});
|
|
1537
|
-
|
|
1538
|
-
// Parallel cancellation behavior:
|
|
1539
|
-
// - Tasks not yet started will be marked as cancelled
|
|
1540
|
-
// - Currently executing tasks will be terminated
|
|
1541
|
-
// - Completed tasks remain in the results
|
|
1542
|
-
// - Partial results may still be available for analysis
|
|
1543
|
-
```
|
|
1544
|
-
|
|
1545
|
-
### Multiple Cancellation Sources
|
|
1546
|
-
|
|
1547
|
-
One of the powerful aspects of the AbortSignal pattern is that multiple sources can cancel the same operation:
|
|
1548
|
-
|
|
1549
|
-
```typescript
|
|
1550
|
-
async function intelligentPromptExecution() {
|
|
1551
|
-
const controller = new AbortController();
|
|
1552
|
-
const signal = controller.signal;
|
|
1553
|
-
|
|
1554
|
-
// 1. User cancel button
|
|
1555
|
-
document.getElementById('cancelBtn')?.addEventListener('click', () => {
|
|
1556
|
-
controller.abort(); // User-initiated cancellation
|
|
1557
|
-
console.log('User cancelled the operation');
|
|
1558
|
-
});
|
|
1559
|
-
|
|
1560
|
-
// 2. Timeout cancellation (prevent runaway prompts)
|
|
1561
|
-
const timeout = setTimeout(() => {
|
|
1562
|
-
controller.abort(); // Timeout cancellation
|
|
1563
|
-
console.log('Operation timed out after 2 minutes');
|
|
1564
|
-
}, 120000);
|
|
1565
|
-
|
|
1566
|
-
// 3. Resource limit cancellation
|
|
1567
|
-
const memoryCheck = setInterval(async () => {
|
|
1568
|
-
if (await getMemoryUsage() > MAX_MEMORY_THRESHOLD) {
|
|
1569
|
-
controller.abort(); // Resource limit cancellation
|
|
1570
|
-
console.log('Cancelled due to memory limits');
|
|
1571
|
-
}
|
|
1572
|
-
}, 5000);
|
|
1573
|
-
|
|
1574
|
-
// 4. Window unload cancellation (cleanup on page close)
|
|
1575
|
-
window.addEventListener('beforeunload', () => {
|
|
1576
|
-
controller.abort(); // Page closing cancellation
|
|
1577
|
-
});
|
|
1578
|
-
|
|
1579
|
-
try {
|
|
1580
|
-
const result = await runner.ExecutePrompt({
|
|
1581
|
-
prompt: complexAnalysisPrompt,
|
|
1582
|
-
data: largeDataset,
|
|
1583
|
-
cancellationToken: signal // One token, many cancel sources!
|
|
1584
|
-
});
|
|
1585
|
-
|
|
1586
|
-
// Clean up timers if successful
|
|
1587
|
-
clearTimeout(timeout);
|
|
1588
|
-
clearInterval(memoryCheck);
|
|
1589
|
-
|
|
1590
|
-
return result;
|
|
1591
|
-
} catch (error) {
|
|
1592
|
-
// The Prompt Runner doesn't know WHY it was cancelled
|
|
1593
|
-
// It just knows it should stop gracefully
|
|
1594
|
-
console.log('Prompt execution was cancelled:', error.message);
|
|
1595
|
-
} finally {
|
|
1596
|
-
clearTimeout(timeout);
|
|
1597
|
-
clearInterval(memoryCheck);
|
|
1598
|
-
}
|
|
1599
|
-
}
|
|
1600
|
-
```
|
|
1601
|
-
|
|
1602
|
-
### Cancellation in Component-Based UIs
|
|
1603
|
-
|
|
1604
|
-
Perfect for React, Angular, or Vue components:
|
|
1605
|
-
|
|
1606
|
-
```typescript
|
|
1607
|
-
class PromptExecutionComponent {
|
|
1608
|
-
private currentController: AbortController | null = null;
|
|
1609
|
-
private isExecuting: boolean = false;
|
|
1610
|
-
|
|
1611
|
-
async executePrompt(prompt: AIPromptEntity, data: any) {
|
|
1612
|
-
// Cancel any existing execution
|
|
1613
|
-
this.cancelCurrentExecution();
|
|
1614
|
-
|
|
1615
|
-
// Create new controller for this execution
|
|
1616
|
-
this.currentController = new AbortController();
|
|
1617
|
-
this.isExecuting = true;
|
|
1618
|
-
|
|
1619
|
-
try {
|
|
1620
|
-
const result = await this.runner.ExecutePrompt({
|
|
1621
|
-
prompt,
|
|
1622
|
-
data,
|
|
1623
|
-
cancellationToken: this.currentController.signal,
|
|
1624
|
-
onProgress: (progress) => {
|
|
1625
|
-
this.updateUI(`${progress.step}: ${progress.percentage}%`);
|
|
1626
|
-
},
|
|
1627
|
-
onStreaming: (chunk) => {
|
|
1628
|
-
this.appendStreamingContent(chunk.content);
|
|
1629
|
-
}
|
|
1630
|
-
});
|
|
1631
|
-
|
|
1632
|
-
this.handleSuccess(result);
|
|
1633
|
-
} catch (error) {
|
|
1634
|
-
if (error.message.includes('cancelled')) {
|
|
1635
|
-
this.handleCancellation();
|
|
1636
|
-
} else {
|
|
1637
|
-
this.handleError(error);
|
|
1638
|
-
}
|
|
1639
|
-
} finally {
|
|
1640
|
-
this.isExecuting = false;
|
|
1641
|
-
this.currentController = null;
|
|
1642
|
-
}
|
|
1643
|
-
}
|
|
1644
|
-
|
|
1645
|
-
// Called when user clicks "Cancel" or navigates away
|
|
1646
|
-
cancelCurrentExecution() {
|
|
1647
|
-
if (this.currentController && this.isExecuting) {
|
|
1648
|
-
this.currentController.abort();
|
|
1649
|
-
console.log('Cancelled current prompt execution');
|
|
1650
|
-
}
|
|
1651
|
-
}
|
|
1652
|
-
|
|
1653
|
-
// Component cleanup
|
|
1654
|
-
ngOnDestroy() { // Angular example
|
|
1655
|
-
this.cancelCurrentExecution();
|
|
1656
|
-
}
|
|
1657
|
-
}
|
|
1658
|
-
```
|
|
1659
|
-
|
|
1660
|
-
### Integration with BaseLLM Cancellation
|
|
1661
|
-
|
|
1662
|
-
The cancellation token is automatically propagated through the entire execution chain:
|
|
1663
|
-
|
|
1664
|
-
```typescript
|
|
1665
|
-
// Cancellation Flow in MemberJunction AI Architecture:
|
|
1666
|
-
//
|
|
1667
|
-
// 1. User Code (AbortController.signal)
|
|
1668
|
-
// ↓
|
|
1669
|
-
// 2. AIPromptRunner.ExecutePrompt(cancellationToken)
|
|
1670
|
-
// ↓
|
|
1671
|
-
// 3. ParallelExecutionCoordinator.executeTasksInParallel(cancellationToken)
|
|
1672
|
-
// ↓
|
|
1673
|
-
// 4. Individual Task Execution with cancellation
|
|
1674
|
-
// ↓
|
|
1675
|
-
// 5. BaseLLM.ChatCompletion({ cancellationToken })
|
|
1676
|
-
// ↓
|
|
1677
|
-
// 6. Provider-specific cancellation (fetch signal, Promise.race)
|
|
1678
|
-
// ↓
|
|
1679
|
-
// 7. AI Model API cancellation (if supported)
|
|
1680
|
-
|
|
1681
|
-
// At each level, cancellation is handled appropriately:
|
|
1682
|
-
const internalFlow = {
|
|
1683
|
-
// Level 1: Prompt Runner checks before major operations
|
|
1684
|
-
promptRunner: () => {
|
|
1685
|
-
if (cancellationToken?.aborted) {
|
|
1686
|
-
return { success: false, cancelled: true };
|
|
1687
|
-
}
|
|
1688
|
-
},
|
|
1689
|
-
|
|
1690
|
-
// Level 2: Parallel coordinator cancels remaining tasks
|
|
1691
|
-
parallelCoordinator: () => {
|
|
1692
|
-
tasks.forEach(task => {
|
|
1693
|
-
if (cancellationToken?.aborted) {
|
|
1694
|
-
task.cancelled = true;
|
|
1695
|
-
}
|
|
1696
|
-
});
|
|
1697
|
-
},
|
|
1698
|
-
|
|
1699
|
-
// Level 3: BaseLLM uses Promise.race for instant cancellation
|
|
1700
|
-
baseLLM: () => {
|
|
1701
|
-
return Promise.race([
|
|
1702
|
-
actualModelCall(params),
|
|
1703
|
-
cancellationPromise(cancellationToken)
|
|
1704
|
-
]);
|
|
1705
|
-
},
|
|
1706
|
-
|
|
1707
|
-
// Level 4: Native provider cancellation (where supported)
|
|
1708
|
-
provider: () => {
|
|
1709
|
-
fetch(apiUrl, {
|
|
1710
|
-
signal: cancellationToken // Native browser/Node.js cancellation
|
|
1711
|
-
});
|
|
1712
|
-
}
|
|
1713
|
-
};
|
|
1714
|
-
```
|
|
1715
|
-
|
|
1716
|
-
### Cancellation Guarantees
|
|
1717
|
-
|
|
1718
|
-
The AI Prompt Runner provides these cancellation guarantees:
|
|
1719
|
-
|
|
1720
|
-
1. **🚫 Instant Recognition**: Cancellation requests are checked at multiple points throughout execution
|
|
1721
|
-
2. **🧹 Graceful Cleanup**: Partial results are preserved and returned when possible
|
|
1722
|
-
3. **📊 Proper Logging**: Cancelled operations are logged with appropriate status and metadata
|
|
1723
|
-
4. **💾 Resource Release**: Network connections and memory are cleaned up promptly
|
|
1724
|
-
5. **🔄 State Consistency**: The system remains in a consistent state after cancellation
|
|
1725
|
-
|
|
1726
|
-
**Key Benefits:**
|
|
1727
|
-
- **Responsive UI**: Users get immediate feedback when cancelling operations
|
|
1728
|
-
- **Resource Efficiency**: Prevents wasted compute and API costs
|
|
1729
|
-
- **System Stability**: Avoids memory leaks and hanging operations
|
|
1730
|
-
- **Standard Pattern**: Uses native JavaScript APIs - no custom cancellation logic needed
|
|
1731
|
-
|
|
1732
|
-
### Cancellation Result Properties
|
|
1733
|
-
|
|
1734
|
-
When execution is cancelled, the result includes detailed cancellation information:
|
|
1735
|
-
|
|
1736
|
-
```typescript
|
|
1737
|
-
interface AIPromptRunResult {
|
|
1738
|
-
success: boolean;
|
|
1739
|
-
cancelled?: boolean; // True if execution was cancelled
|
|
1740
|
-
cancellationReason?: CancellationReason; // Why it was cancelled
|
|
1741
|
-
status?: ExecutionStatus; // Current execution status
|
|
1742
|
-
// ... other properties
|
|
1743
|
-
}
|
|
1744
|
-
|
|
1745
|
-
type CancellationReason = 'user_requested' | 'timeout' | 'error' | 'resource_limit';
|
|
1746
|
-
type ExecutionStatus = 'pending' | 'running' | 'completed' | 'failed' | 'cancelled';
|
|
1747
|
-
```
|
|
1748
|
-
|
|
1749
|
-
## Progress Updates & Streaming
|
|
1750
|
-
|
|
1751
|
-
The AI Prompt Runner provides real-time progress updates and streaming support for long-running executions, enabling responsive user interfaces and monitoring dashboards.
|
|
1752
|
-
|
|
1753
|
-
### Progress Callbacks
|
|
1754
|
-
|
|
1755
|
-
Track execution progress through different phases:
|
|
1756
|
-
|
|
1757
|
-
```typescript
|
|
1758
|
-
const runner = new AIPromptRunner();
|
|
1759
|
-
|
|
1760
|
-
const result = await runner.ExecutePrompt({
|
|
1761
|
-
prompt: complexPrompt,
|
|
1762
|
-
data: { document: longDocument },
|
|
1763
|
-
contextUser: currentUser,
|
|
1764
|
-
|
|
1765
|
-
// Progress callback receives updates throughout execution
|
|
1766
|
-
onProgress: (progress) => {
|
|
1767
|
-
console.log(`${progress.step}: ${progress.percentage}% - ${progress.message}`);
|
|
1768
|
-
|
|
1769
|
-
// Update UI progress bar
|
|
1770
|
-
updateProgressBar(progress.percentage);
|
|
1771
|
-
updateStatusMessage(progress.message);
|
|
1772
|
-
|
|
1773
|
-
// Access additional metadata
|
|
1774
|
-
if (progress.metadata) {
|
|
1775
|
-
console.log('Execution metadata:', progress.metadata);
|
|
1776
|
-
}
|
|
1777
|
-
}
|
|
1778
|
-
});
|
|
1779
|
-
```
|
|
1780
|
-
|
|
1781
|
-
### Execution Progress Phases
|
|
1782
|
-
|
|
1783
|
-
The progress callback receives updates for these execution phases:
|
|
1784
|
-
|
|
1785
|
-
```typescript
|
|
1786
|
-
type ProgressPhase =
|
|
1787
|
-
| 'template_rendering' // Rendering prompt template with data
|
|
1788
|
-
| 'model_selection' // Selecting appropriate AI model
|
|
1789
|
-
| 'execution' // Executing AI model
|
|
1790
|
-
| 'validation' // Validating and parsing results
|
|
1791
|
-
| 'parallel_coordination' // Coordinating parallel executions
|
|
1792
|
-
| 'result_selection'; // AI judge selecting best result
|
|
1793
|
-
|
|
1794
|
-
// Example progress updates:
|
|
1795
|
-
// template_rendering: 20% - "Rendering prompt template with provided data"
|
|
1796
|
-
// model_selection: 40% - "Selected GPT-4 model based on prompt configuration"
|
|
1797
|
-
// execution: 60% - "Executing AI model..."
|
|
1798
|
-
// validation: 80% - "Validating output against expected format"
|
|
1799
|
-
// result_selection: 90% - "AI judge selecting best result from 3 candidates"
|
|
1800
|
-
```
|
|
1801
|
-
|
|
1802
|
-
### Streaming Response Support
|
|
1803
|
-
|
|
1804
|
-
Receive real-time content updates as AI models generate responses:
|
|
1805
|
-
|
|
1806
|
-
```typescript
|
|
1807
|
-
const result = await runner.ExecutePrompt({
|
|
1808
|
-
prompt: streamingPrompt,
|
|
1809
|
-
data: { query: 'Generate a detailed report...' },
|
|
1810
|
-
contextUser: currentUser,
|
|
1811
|
-
|
|
1812
|
-
// Streaming callback receives content chunks as they arrive
|
|
1813
|
-
onStreaming: (chunk) => {
|
|
1814
|
-
if (chunk.isComplete) {
|
|
1815
|
-
console.log('Streaming complete');
|
|
1816
|
-
finalizeDocument();
|
|
1817
|
-
} else {
|
|
1818
|
-
// Append content chunk to UI
|
|
1819
|
-
appendToDocument(chunk.content);
|
|
1820
|
-
|
|
1821
|
-
// Show which model is generating content (for parallel execution)
|
|
1822
|
-
if (chunk.modelName) {
|
|
1823
|
-
showActiveModel(chunk.modelName);
|
|
1824
|
-
}
|
|
1825
|
-
}
|
|
1826
|
-
}
|
|
1827
|
-
});
|
|
1828
|
-
```
|
|
1829
|
-
|
|
1830
|
-
### Progress Updates in Parallel Execution
|
|
1831
|
-
|
|
1832
|
-
Progress tracking works seamlessly with parallel execution:
|
|
1833
|
-
|
|
1834
|
-
```typescript
|
|
1835
|
-
const result = await runner.ExecutePrompt({
|
|
1836
|
-
prompt: parallelPrompt, // Uses multiple models
|
|
1837
|
-
data: analysisData,
|
|
1838
|
-
contextUser: currentUser,
|
|
1839
|
-
|
|
1840
|
-
onProgress: (progress) => {
|
|
1841
|
-
// Parallel execution provides additional metadata
|
|
1842
|
-
if (progress.metadata?.parallelExecution) {
|
|
1843
|
-
const parallel = progress.metadata.parallelExecution;
|
|
1844
|
-
console.log(`Group ${parallel.currentGroup}/${parallel.totalGroups}`);
|
|
1845
|
-
console.log(`Tasks: ${parallel.completedTasks}/${parallel.totalTasks}`);
|
|
1846
|
-
console.log(`Successful: ${parallel.successfulTasks}`);
|
|
1847
|
-
}
|
|
1848
|
-
},
|
|
1849
|
-
|
|
1850
|
-
onStreaming: (chunk) => {
|
|
1851
|
-
// Multiple models may stream simultaneously
|
|
1852
|
-
console.log(`${chunk.modelName}: ${chunk.content}`);
|
|
1853
|
-
|
|
1854
|
-
// Update model-specific UI sections
|
|
1855
|
-
updateModelSection(chunk.taskId, chunk.content);
|
|
1856
|
-
}
|
|
1857
|
-
});
|
|
1858
|
-
```
|
|
1859
|
-
|
|
1860
|
-
### Advanced Streaming Configuration
|
|
1861
|
-
|
|
1862
|
-
Fine-tune streaming behavior for optimal performance:
|
|
1863
|
-
|
|
1864
|
-
```typescript
|
|
1865
|
-
// Streaming configuration can be applied globally or per-prompt
|
|
1866
|
-
const streamingConfig = {
|
|
1867
|
-
enabled: true,
|
|
1868
|
-
aggregateParallelUpdates: false, // Separate updates per parallel task
|
|
1869
|
-
progressUpdateIntervalMS: 250 // Limit update frequency
|
|
1870
|
-
};
|
|
1871
|
-
|
|
1872
|
-
// Progress updates are automatically throttled to prevent UI flooding
|
|
1873
|
-
// Minimum interval between updates prevents performance issues
|
|
1874
|
-
```
|
|
1875
|
-
|
|
1876
|
-
### Integration with BaseLLM Streaming
|
|
1877
|
-
|
|
1878
|
-
The streaming system integrates seamlessly with BaseLLM capabilities:
|
|
1879
|
-
|
|
1880
|
-
```typescript
|
|
1881
|
-
// The AI Prompt Runner automatically detects streaming support:
|
|
1882
|
-
// 1. Checks if the selected model supports streaming
|
|
1883
|
-
// 2. Configures BaseLLM streaming callbacks
|
|
1884
|
-
// 3. Aggregates streaming updates from multiple models in parallel execution
|
|
1885
|
-
// 4. Provides unified streaming interface regardless of underlying model
|
|
1886
|
-
|
|
1887
|
-
// Models that support streaming will automatically use it when callbacks are provided
|
|
1888
|
-
// Models without streaming support will provide content in the final result
|
|
1889
|
-
```
|
|
1890
|
-
|
|
1891
|
-
## API Reference
|
|
1892
|
-
|
|
1893
|
-
### Exported Classes and Types
|
|
1894
|
-
|
|
1895
|
-
The package exports the following public API:
|
|
1896
|
-
|
|
1897
|
-
```typescript
|
|
1898
|
-
// Main classes
|
|
1899
|
-
export { AIPromptCategoryEntityExtended } from './AIPromptCategoryExtended';
|
|
1900
|
-
export { AIPromptRunner, AIPromptParams, AIPromptRunResult } from './AIPromptRunner';
|
|
1901
|
-
|
|
1902
|
-
// Helper types
|
|
1903
|
-
export { ChildPromptParam } from './AIPromptRunner';
|
|
1904
|
-
export { SystemPlaceholder, SystemPlaceholderManager } from './SystemPlaceholders';
|
|
1905
|
-
|
|
1906
|
-
// Callback types
|
|
1907
|
-
export type { ExecutionProgressCallback, ExecutionStreamingCallback } from './AIPromptRunner';
|
|
1908
|
-
```
|
|
1909
|
-
|
|
1910
|
-
### Import Examples
|
|
1911
|
-
|
|
1912
|
-
```typescript
|
|
1913
|
-
// Import from this package
|
|
1914
|
-
import { AIPromptRunner, AIPromptParams, AIPromptRunResult } from '@memberjunction/ai-prompts';
|
|
1915
|
-
import { ChildPromptParam, SystemPlaceholderManager } from '@memberjunction/ai-prompts';
|
|
1916
|
-
|
|
1917
|
-
// Import base types from Core
|
|
1918
|
-
import { ChatResult, ModelUsage, ChatMessage } from '@memberjunction/ai';
|
|
1919
|
-
|
|
1920
|
-
// Import entities and engine types
|
|
1921
|
-
import { AIPromptEntity } from '@memberjunction/core-entities';
|
|
1922
|
-
import { AIEngine } from '@memberjunction/aiengine';
|
|
1923
|
-
```
|
|
1924
|
-
|
|
1925
|
-
### AIPromptRunner Class
|
|
1926
|
-
|
|
1927
|
-
Handles execution of AI prompts with advanced parallel processing, template rendering, and result validation.
|
|
1928
|
-
|
|
1929
|
-
#### Methods
|
|
1930
|
-
|
|
1931
|
-
- `ExecutePrompt(params: AIPromptParams): Promise<AIPromptRunResult>`: Execute a prompt with full feature support including template rendering, model selection, parallel execution, and output validation
|
|
1932
|
-
|
|
1933
|
-
#### AIPromptParams Interface
|
|
1934
|
-
|
|
1935
|
-
```typescript
|
|
1936
|
-
interface AIPromptParams {
|
|
1937
|
-
prompt: AIPromptEntity; // The prompt to execute
|
|
1938
|
-
data?: any; // Template and context data
|
|
1939
|
-
modelId?: string; // Override model selection
|
|
1940
|
-
vendorId?: string; // Override vendor selection
|
|
1941
|
-
configurationId?: string; // AI Configuration ID for environment-specific model selection (e.g., Prod vs Dev)
|
|
1942
|
-
contextUser?: UserInfo; // User context
|
|
1943
|
-
skipValidation?: boolean; // Skip output validation
|
|
1944
|
-
templateData?: any; // Additional template data that augments the main data context
|
|
1945
|
-
conversationMessages?: ChatMessage[]; // Multi-turn conversation messages
|
|
1946
|
-
templateMessageRole?: TemplateMessageRole; // How to use rendered template ('system'|'user'|'none')
|
|
1947
|
-
cancellationToken?: AbortSignal; // Cancellation token for aborting execution
|
|
1948
|
-
onProgress?: ExecutionProgressCallback; // Progress update callback
|
|
1949
|
-
onStreaming?: ExecutionStreamingCallback; // Streaming content callback
|
|
1950
|
-
agentRunId?: string; // Optional agent run ID to link prompt executions to parent agent run
|
|
1951
|
-
cleanValidationSyntax?: boolean; // Clean validation syntax from JSON responses (auto-enabled for validated prompts)
|
|
1952
|
-
}
|
|
1953
|
-
|
|
1954
|
-
/**
|
|
1955
|
-
* Progress callback function type
|
|
1956
|
-
*/
|
|
1957
|
-
type ExecutionProgressCallback = (progress: {
|
|
1958
|
-
step: 'template_rendering' | 'model_selection' | 'execution' | 'validation' | 'parallel_coordination' | 'result_selection';
|
|
1959
|
-
percentage: number; // Progress percentage (0-100)
|
|
1960
|
-
message: string; // Human-readable status message
|
|
1961
|
-
metadata?: Record<string, any>; // Additional metadata about the current step
|
|
1962
|
-
}) => void;
|
|
1963
|
-
|
|
1964
|
-
/**
|
|
1965
|
-
* Streaming callback function type
|
|
1966
|
-
*/
|
|
1967
|
-
type ExecutionStreamingCallback = (chunk: {
|
|
1968
|
-
content: string; // The content chunk received
|
|
1969
|
-
isComplete: boolean; // Whether this is the final chunk
|
|
1970
|
-
taskId?: string; // Which task/model is producing this content (for parallel execution)
|
|
1971
|
-
modelName?: string; // Model name producing this content
|
|
1972
|
-
}) => void;
|
|
1973
|
-
|
|
1974
|
-
/**
|
|
1975
|
-
* Template message role type
|
|
1976
|
-
*/
|
|
1977
|
-
type TemplateMessageRole = 'system' | 'user' | 'none';
|
|
1978
|
-
```
|
|
1979
|
-
|
|
1980
|
-
### Extended Entity Classes
|
|
1981
|
-
|
|
1982
|
-
#### AIPromptCategoryEntityExtended
|
|
1983
|
-
|
|
1984
|
-
Extended prompt category with prompt collection:
|
|
1985
|
-
|
|
1986
|
-
```typescript
|
|
1987
|
-
class AIPromptCategoryEntityExtended extends AIPromptCategoryEntity {
|
|
1988
|
-
get Prompts(): AIPromptEntity[]; // Prompts in this category
|
|
1989
|
-
}
|
|
1990
|
-
```
|
|
1991
|
-
|
|
1992
|
-
### Key Interfaces and Types
|
|
1993
|
-
|
|
1994
|
-
```typescript
|
|
1995
|
-
interface AIPromptRunResult<T = unknown> {
|
|
1996
|
-
success: boolean; // Whether the execution was successful
|
|
1997
|
-
status?: ExecutionStatus; // Current execution status
|
|
1998
|
-
cancelled?: boolean; // Whether the execution was cancelled
|
|
1999
|
-
cancellationReason?: CancellationReason; // Reason for cancellation if applicable
|
|
2000
|
-
rawResult?: string; // The raw result from the AI model
|
|
2001
|
-
result?: T; // The parsed/validated result based on OutputType
|
|
2002
|
-
errorMessage?: string; // Error message if execution failed
|
|
2003
|
-
promptRun?: AIPromptRunEntity; // The AIPromptRun entity that was created for tracking
|
|
2004
|
-
executionTimeMS?: number; // Total execution time in milliseconds
|
|
2005
|
-
|
|
2006
|
-
// Token tracking (follows ModelUsage convention)
|
|
2007
|
-
promptTokens?: number; // Prompt/input tokens for this execution
|
|
2008
|
-
completionTokens?: number; // Completion/output tokens for this execution
|
|
2009
|
-
tokensUsed?: number; // Total tokens (calculated getter)
|
|
2010
|
-
|
|
2011
|
-
// Hierarchical token tracking
|
|
2012
|
-
combinedPromptTokens?: number; // Total prompt tokens including all children
|
|
2013
|
-
combinedCompletionTokens?: number; // Total completion tokens including all children
|
|
2014
|
-
combinedTokensUsed?: number; // Total tokens including all children (calculated)
|
|
2015
|
-
|
|
2016
|
-
// Cost tracking
|
|
2017
|
-
cost?: number; // Cost of this execution
|
|
2018
|
-
costCurrency?: string; // ISO 4217 currency code (USD, EUR, etc.)
|
|
2019
|
-
combinedCost?: number; // Total cost including all children
|
|
2020
|
-
|
|
2021
|
-
validationResult?: ValidationResult; // Validation result if output validation was performed
|
|
2022
|
-
validationAttempts?: ValidationAttempt[]; // Detailed validation attempts
|
|
2023
|
-
additionalResults?: AIPromptRunResult<T>[]; // Additional results from parallel execution, ranked by judge
|
|
2024
|
-
ranking?: number; // Ranking assigned by judge (1 = best, 2 = second best, etc.)
|
|
2025
|
-
judgeRationale?: string; // Judge's rationale for this ranking
|
|
2026
|
-
modelInfo?: ModelInfo; // Model information for this result
|
|
2027
|
-
judgeMetadata?: JudgeMetadata; // Metadata about the judging process (only present on the main result)
|
|
2028
|
-
wasStreamed?: boolean; // Whether streaming was used for this execution
|
|
2029
|
-
cacheInfo?: { // Cache information if caching was involved
|
|
2030
|
-
cacheHit: boolean;
|
|
2031
|
-
cacheKey?: string;
|
|
2032
|
-
cacheSource?: string;
|
|
2033
|
-
};
|
|
2034
|
-
}
|
|
2035
|
-
|
|
2036
|
-
// Execution status enumeration
|
|
2037
|
-
type ExecutionStatus = 'pending' | 'running' | 'completed' | 'failed' | 'cancelled';
|
|
2038
|
-
|
|
2039
|
-
// Cancellation reason enumeration
|
|
2040
|
-
type CancellationReason = 'user_requested' | 'timeout' | 'error' | 'resource_limit';
|
|
2041
|
-
|
|
2042
|
-
// Model information interface
|
|
2043
|
-
interface ModelInfo {
|
|
2044
|
-
modelId: string;
|
|
2045
|
-
modelName: string;
|
|
2046
|
-
vendorId?: string;
|
|
2047
|
-
vendorName?: string;
|
|
2048
|
-
powerRank?: number;
|
|
2049
|
-
modelType?: string;
|
|
2050
|
-
}
|
|
2051
|
-
|
|
2052
|
-
// Judge metadata interface
|
|
2053
|
-
interface JudgeMetadata {
|
|
2054
|
-
judgePromptId: string;
|
|
2055
|
-
judgeExecutionTimeMS: number;
|
|
2056
|
-
judgeTokensUsed?: number;
|
|
2057
|
-
judgeCancelled?: boolean;
|
|
2058
|
-
judgeErrorMessage?: string;
|
|
2059
|
-
}
|
|
2060
|
-
|
|
2061
|
-
// Parallelization strategies supported by the system
|
|
2062
|
-
type ParallelizationStrategy = 'None' | 'StaticCount' | 'ConfigParam' | 'ModelSpecific';
|
|
2063
|
-
|
|
2064
|
-
// Result selection methods for choosing best result from parallel executions
|
|
2065
|
-
type ResultSelectionMethod = 'First' | 'Random' | 'PromptSelector' | 'Consensus';
|
|
2066
|
-
```
|
|
2067
|
-
|
|
2068
|
-
## Integration with Other Packages
|
|
2069
|
-
|
|
2070
|
-
### With AI Engine
|
|
2071
|
-
|
|
2072
|
-
The Prompts package builds on the AI Engine for basic functionality:
|
|
2073
|
-
|
|
2074
|
-
```typescript
|
|
2075
|
-
// AI Engine provides model management and basic operations
|
|
2076
|
-
import { AIEngine } from '@memberjunction/aiengine';
|
|
2077
|
-
import { AIPromptRunner } from '@memberjunction/ai-prompts';
|
|
2078
|
-
|
|
2079
|
-
// Initialize AI Engine first
|
|
2080
|
-
await AIEngine.Instance.Config(false, currentUser);
|
|
2081
|
-
|
|
2082
|
-
// Access prompts managed by AI Engine
|
|
2083
|
-
const prompts = AIEngine.Instance.Prompts;
|
|
2084
|
-
const prompt = prompts.find(p => p.Name === 'Your Prompt');
|
|
2085
|
-
|
|
2086
|
-
// Use Prompts package for advanced execution
|
|
2087
|
-
const runner = new AIPromptRunner();
|
|
2088
|
-
const result = await runner.ExecutePrompt({ prompt, data, contextUser });
|
|
2089
|
-
```
|
|
2090
|
-
|
|
2091
|
-
### With AI Agents
|
|
2092
|
-
|
|
2093
|
-
AI Agents can leverage the prompt system for sophisticated operations with comprehensive execution tracking:
|
|
2094
|
-
|
|
2095
|
-
```typescript
|
|
2096
|
-
// Agents use prompts for their intelligence with hierarchical logging
|
|
2097
|
-
import { AgentRunner } from '@memberjunction/ai-agents';
|
|
2098
|
-
import { AIPromptRunner } from '@memberjunction/ai-prompts';
|
|
2099
|
-
|
|
2100
|
-
class IntelligentAgent extends AgentRunner {
|
|
2101
|
-
private promptRunner = new AIPromptRunner();
|
|
2102
|
-
|
|
2103
|
-
async execute(context: AgentExecutionContext): Promise<AgentExecutionResult> {
|
|
2104
|
-
const prompt = this.getPromptForContext(context);
|
|
2105
|
-
|
|
2106
|
-
// Link prompt execution to agent run for comprehensive tracking
|
|
2107
|
-
const result = await this.promptRunner.ExecutePrompt({
|
|
2108
|
-
prompt: prompt,
|
|
2109
|
-
data: context.data,
|
|
2110
|
-
contextUser: context.user,
|
|
2111
|
-
agentRunId: context.agentRun?.ID // Links prompt to parent agent run
|
|
2112
|
-
});
|
|
2113
|
-
|
|
2114
|
-
return this.formatAgentResult(result);
|
|
2115
|
-
}
|
|
2116
|
-
}
|
|
2117
|
-
```
|
|
2118
|
-
|
|
2119
|
-
#### Agent-Prompt Integration Features
|
|
2120
|
-
|
|
2121
|
-
The AI Prompts system provides seamless integration with AI Agents through the `agentRunId` parameter:
|
|
2122
|
-
|
|
2123
|
-
**Hierarchical Execution Tracking:**
|
|
2124
|
-
- Prompt executions are linked to their parent agent runs via `AgentRunID` foreign key
|
|
2125
|
-
- Provides complete audit trail from agent decision to prompt execution
|
|
2126
|
-
- Enables comprehensive resource usage tracking across agent workflows
|
|
2127
|
-
|
|
2128
|
-
**Usage Patterns:**
|
|
2129
|
-
```typescript
|
|
2130
|
-
// 1. Direct agent-prompt linking
|
|
2131
|
-
const result = await promptRunner.ExecutePrompt({
|
|
2132
|
-
prompt: myPrompt,
|
|
2133
|
-
data: promptData,
|
|
2134
|
-
agentRunId: agentRun.ID, // Links to parent agent execution
|
|
2135
|
-
contextUser: user
|
|
2136
|
-
});
|
|
2137
|
-
|
|
2138
|
-
// 2. Parallel execution with agent tracking
|
|
2139
|
-
const parallelResult = await promptRunner.ExecutePrompt({
|
|
2140
|
-
prompt: parallelPrompt, // ParallelizationMode: 'ModelSpecific'
|
|
2141
|
-
data: analysisData,
|
|
2142
|
-
agentRunId: agentRun.ID, // All parallel child prompts link to agent
|
|
2143
|
-
contextUser: user
|
|
2144
|
-
});
|
|
2145
|
-
|
|
2146
|
-
// 3. Context compression with agent linking (automatic in AgentRunner)
|
|
2147
|
-
// When agents use context compression, compression prompts are automatically
|
|
2148
|
-
// linked to the parent agent run for complete execution visibility
|
|
2149
|
-
```
|
|
2150
|
-
|
|
2151
|
-
**Database Schema Integration:**
|
|
2152
|
-
```sql
|
|
2153
|
-
-- Query agent execution with all related prompts
|
|
2154
|
-
SELECT
|
|
2155
|
-
ar.ID as AgentRunID,
|
|
2156
|
-
ar.Status as AgentStatus,
|
|
2157
|
-
ar.StartedAt,
|
|
2158
|
-
ar.CompletedAt,
|
|
2159
|
-
pr.ID as PromptRunID,
|
|
2160
|
-
pr.RunType,
|
|
2161
|
-
pr.Success as PromptSuccess,
|
|
2162
|
-
pr.ExecutionTimeMS,
|
|
2163
|
-
pr.TokensUsed
|
|
2164
|
-
FROM AIAgentRun ar
|
|
2165
|
-
LEFT JOIN AIPromptRun pr ON ar.ID = pr.AgentRunID
|
|
2166
|
-
WHERE ar.ID = 'your-agent-run-id'
|
|
2167
|
-
ORDER BY pr.RunAt;
|
|
2168
|
-
```
|
|
2169
|
-
|
|
2170
|
-
**Benefits:**
|
|
2171
|
-
- **Complete Traceability**: Track all AI model usage from agent decisions to prompt executions
|
|
2172
|
-
- **Resource Attribution**: Understand token usage and costs at the agent level
|
|
2173
|
-
- **Performance Analysis**: Analyze execution patterns across the agent-prompt hierarchy
|
|
2174
|
-
- **Debugging Support**: Full execution history for troubleshooting agent workflows
|
|
2175
|
-
|
|
2176
235
|
## Dependencies
|
|
2177
236
|
|
|
2178
|
-
- `@memberjunction/
|
|
2179
|
-
- `@memberjunction/
|
|
2180
|
-
- `@memberjunction/
|
|
2181
|
-
- `@memberjunction/
|
|
2182
|
-
- `@memberjunction/
|
|
2183
|
-
- `@memberjunction/
|
|
2184
|
-
- `
|
|
2185
|
-
- `
|
|
2186
|
-
|
|
2187
|
-
|
|
2188
|
-
|
|
2189
|
-
- `@memberjunction/aiengine`: Core AI engine and model management
|
|
2190
|
-
- `@memberjunction/ai-agents`: Advanced agent framework built on prompts
|
|
2191
|
-
- `@memberjunction/templates`: Template rendering for dynamic content
|
|
2192
|
-
|
|
2193
|
-
## Migration Guide
|
|
2194
|
-
|
|
2195
|
-
### From AI Engine Simple Completions
|
|
2196
|
-
|
|
2197
|
-
For cases requiring more sophisticated prompt management:
|
|
2198
|
-
|
|
2199
|
-
```typescript
|
|
2200
|
-
// Old: Simple LLM completion (still valid for basic cases)
|
|
2201
|
-
const response = await AIEngine.Instance.SimpleLLMCompletion(
|
|
2202
|
-
"Analyze this data",
|
|
2203
|
-
currentUser,
|
|
2204
|
-
"You are a data analyst"
|
|
2205
|
-
);
|
|
2206
|
-
|
|
2207
|
-
// New: Advanced prompt with caching, validation, and parallel execution
|
|
2208
|
-
const prompt = {
|
|
2209
|
-
Name: "Data Analysis",
|
|
2210
|
-
PromptText: "Analyze this data: {{data}}",
|
|
2211
|
-
EnableCaching: true,
|
|
2212
|
-
ParallelizationMode: "StaticCount",
|
|
2213
|
-
ParallelCount: 2,
|
|
2214
|
-
OutputType: "object",
|
|
2215
|
-
ValidationBehavior: "Strict"
|
|
2216
|
-
};
|
|
2217
|
-
|
|
2218
|
-
const runner = new AIPromptRunner();
|
|
2219
|
-
const result = await runner.ExecutePrompt({
|
|
2220
|
-
prompt: prompt,
|
|
2221
|
-
data: { data: "your data here" },
|
|
2222
|
-
contextUser: currentUser
|
|
2223
|
-
});
|
|
2224
|
-
```
|
|
2225
|
-
|
|
2226
|
-
## Best Practices
|
|
2227
|
-
|
|
2228
|
-
1. **Enable Caching**: Use intelligent caching for expensive operations
|
|
2229
|
-
2. **Validate Outputs**: Always specify expected output types for critical operations
|
|
2230
|
-
3. **Use Templates**: Leverage template system for dynamic prompts
|
|
2231
|
-
4. **Monitor Performance**: Track token usage and execution times
|
|
2232
|
-
5. **Parallel Wisely**: Use parallel execution for independent tasks, not dependent ones
|
|
2233
|
-
6. **Handle Errors**: Implement proper retry logic and error handling
|
|
2234
|
-
7. **Implement Cancellation**: Always provide cancellation tokens for user-facing operations
|
|
2235
|
-
8. **Use Progress Callbacks**: Provide progress feedback for long-running operations
|
|
2236
|
-
9. **Leverage Hierarchical Logging**: Use the logging hierarchy for debugging and analytics
|
|
2237
|
-
10. **Configure Streaming Appropriately**: Enable streaming for responsive user experiences
|
|
2238
|
-
11. **Optimize Judge Selection**: Use efficient judge prompts for parallel result selection
|
|
2239
|
-
12. **Monitor Resource Usage**: Track token consumption and execution times across hierarchical runs
|
|
2240
|
-
|
|
2241
|
-
### Implementation Guidelines
|
|
2242
|
-
|
|
2243
|
-
```typescript
|
|
2244
|
-
// Comprehensive prompt execution with all new features
|
|
2245
|
-
const controller = new AbortController();
|
|
2246
|
-
|
|
2247
|
-
const result = await runner.ExecutePrompt({
|
|
2248
|
-
prompt: myPrompt,
|
|
2249
|
-
data: executionData,
|
|
2250
|
-
contextUser: currentUser,
|
|
2251
|
-
|
|
2252
|
-
// Cancellation support
|
|
2253
|
-
cancellationToken: controller.signal,
|
|
2254
|
-
|
|
2255
|
-
// Progress tracking
|
|
2256
|
-
onProgress: (progress) => {
|
|
2257
|
-
updateProgressIndicator(progress.percentage, progress.message);
|
|
2258
|
-
if (progress.metadata?.parallelExecution) {
|
|
2259
|
-
updateParallelStatus(progress.metadata.parallelExecution);
|
|
2260
|
-
}
|
|
2261
|
-
},
|
|
2262
|
-
|
|
2263
|
-
// Streaming for real-time updates
|
|
2264
|
-
onStreaming: (chunk) => {
|
|
2265
|
-
if (chunk.isComplete) {
|
|
2266
|
-
finalizePage();
|
|
2267
|
-
} else {
|
|
2268
|
-
appendContent(chunk.content);
|
|
2269
|
-
}
|
|
2270
|
-
}
|
|
2271
|
-
});
|
|
2272
|
-
|
|
2273
|
-
// Always check for cancellation in results
|
|
2274
|
-
if (result.cancelled) {
|
|
2275
|
-
handleCancellation(result.cancellationReason);
|
|
2276
|
-
} else if (result.success) {
|
|
2277
|
-
processResults(result);
|
|
2278
|
-
|
|
2279
|
-
// Analyze additional results from parallel execution
|
|
2280
|
-
if (result.additionalResults) {
|
|
2281
|
-
analyzeAlternativeResults(result.additionalResults);
|
|
2282
|
-
}
|
|
2283
|
-
}
|
|
2284
|
-
|
|
2285
|
-
// Use hierarchical logging data for analytics
|
|
2286
|
-
if (result.promptRun) {
|
|
2287
|
-
trackExecutionMetrics(result.promptRun);
|
|
2288
|
-
if (result.promptRun.RunType === 'ParallelParent') {
|
|
2289
|
-
analyzeParallelPerformance(result.promptRun.ID);
|
|
2290
|
-
}
|
|
2291
|
-
}
|
|
2292
|
-
```
|
|
2293
|
-
|
|
2294
|
-
## Troubleshooting
|
|
2295
|
-
|
|
2296
|
-
### Common Issues
|
|
2297
|
-
|
|
2298
|
-
1. **"No suitable model found" Error**
|
|
2299
|
-
- Ensure AIEngine.Instance.Config() is called before using prompts
|
|
2300
|
-
- Verify prompt has active AIPromptModel associations or proper model selection configuration
|
|
2301
|
-
- Check that models meet MinPowerRank requirements
|
|
2302
|
-
|
|
2303
|
-
2. **Template Rendering Failures**
|
|
2304
|
-
- Verify template exists and is associated with the prompt
|
|
2305
|
-
- Ensure template data contains all required variables
|
|
2306
|
-
- Check template syntax for Handlebars errors
|
|
2307
|
-
|
|
2308
|
-
3. **Parallel Execution Not Working**
|
|
2309
|
-
- Confirm ParallelizationMode is set to a value other than 'None'
|
|
2310
|
-
- For ModelSpecific mode, ensure AIPromptModel entries exist
|
|
2311
|
-
- Check that multiple suitable models are available
|
|
2312
|
-
|
|
2313
|
-
4. **Output Validation Errors**
|
|
2314
|
-
- Ensure OutputType matches the expected result format
|
|
2315
|
-
- Provide a valid OutputExample for structured data
|
|
2316
|
-
- Consider increasing MaxRetries for complex outputs
|
|
2317
|
-
|
|
2318
|
-
5. **Cancellation Not Working**
|
|
2319
|
-
- Verify the AbortController is properly created and signal is passed
|
|
2320
|
-
- Check that the cancellation token is not already aborted before execution
|
|
2321
|
-
- Ensure model implementations support cancellation (older models may not)
|
|
2322
|
-
- Review cancellation timing - very fast executions may complete before cancellation
|
|
2323
|
-
|
|
2324
|
-
6. **Progress Updates Not Received**
|
|
2325
|
-
- Confirm onProgress callback is properly defined and passed to ExecutePrompt
|
|
2326
|
-
- Check that the callback function doesn't throw errors (which can stop updates)
|
|
2327
|
-
- Progress updates are throttled - very fast operations may have fewer updates
|
|
2328
|
-
- Parallel execution provides more detailed progress metadata
|
|
2329
|
-
|
|
2330
|
-
7. **Streaming Not Working**
|
|
2331
|
-
- Verify the selected AI model supports streaming (not all models do)
|
|
2332
|
-
- Ensure onStreaming callback is provided in AIPromptParams
|
|
2333
|
-
- Check BaseLLM implementation supports streaming for the specific model
|
|
2334
|
-
- Review model configuration - some vendors require specific settings for streaming
|
|
2335
|
-
|
|
2336
|
-
8. **Hierarchical Logging Missing**
|
|
2337
|
-
- Ensure database schema includes RunType, ParentID, and ExecutionOrder fields
|
|
2338
|
-
- Check that user has permissions to create AIPromptRun records
|
|
2339
|
-
- Verify prompt run creation isn't being skipped due to errors
|
|
2340
|
-
- Review logs for save failures on prompt run entities
|
|
2341
|
-
|
|
2342
|
-
9. **Judge Selection Failing**
|
|
2343
|
-
- Confirm ResultSelectorPromptID is set and points to a valid, active prompt
|
|
2344
|
-
- Verify the judge prompt returns valid JSON with rankings array
|
|
2345
|
-
- Check that judge prompt has proper model associations
|
|
2346
|
-
- Review judge prompt timeout settings for complex evaluations
|
|
2347
|
-
|
|
2348
|
-
### Performance Optimization
|
|
2349
|
-
|
|
2350
|
-
For optimal performance with the new features:
|
|
2351
|
-
|
|
2352
|
-
```typescript
|
|
2353
|
-
// Minimize progress update frequency for high-performance scenarios
|
|
2354
|
-
const result = await runner.ExecutePrompt({
|
|
2355
|
-
prompt: myPrompt,
|
|
2356
|
-
data: myData,
|
|
2357
|
-
onProgress: (progress) => {
|
|
2358
|
-
// Throttle UI updates
|
|
2359
|
-
if (progress.percentage % 10 === 0) {
|
|
2360
|
-
updateUI(progress);
|
|
2361
|
-
}
|
|
2362
|
-
}
|
|
2363
|
-
});
|
|
2364
|
-
|
|
2365
|
-
// Use cancellation for long-running operations
|
|
2366
|
-
const controller = new AbortController();
|
|
2367
|
-
setTimeout(() => controller.abort(), 60000); // 1 minute timeout
|
|
2368
|
-
|
|
2369
|
-
// Configure parallel execution for optimal throughput
|
|
2370
|
-
const parallelPrompt = {
|
|
2371
|
-
ParallelizationMode: "ModelSpecific",
|
|
2372
|
-
// Configure specific models with different execution groups for coordination
|
|
2373
|
-
};
|
|
2374
|
-
```
|
|
2375
|
-
|
|
2376
|
-
### Debugging Hierarchical Logs
|
|
2377
|
-
|
|
2378
|
-
Use these queries to troubleshoot execution issues:
|
|
2379
|
-
|
|
2380
|
-
```sql
|
|
2381
|
-
-- Find incomplete executions
|
|
2382
|
-
SELECT * FROM AIPromptRun
|
|
2383
|
-
WHERE CompletedAt IS NULL
|
|
2384
|
-
AND RunAt < DATEADD(minute, -5, GETDATE());
|
|
2385
|
-
|
|
2386
|
-
-- Check parallel execution hierarchy
|
|
2387
|
-
SELECT
|
|
2388
|
-
ID, RunType, ParentID, ExecutionOrder, Success, ErrorMessage
|
|
2389
|
-
FROM AIPromptRun
|
|
2390
|
-
WHERE ParentID = 'your-parent-id' OR ID = 'your-parent-id'
|
|
2391
|
-
ORDER BY RunType, ExecutionOrder;
|
|
2392
|
-
|
|
2393
|
-
-- Find resource usage patterns
|
|
2394
|
-
SELECT
|
|
2395
|
-
RunType,
|
|
2396
|
-
AVG(ExecutionTimeMS) as AvgTimeMS,
|
|
2397
|
-
AVG(TokensUsed) as AvgTokens,
|
|
2398
|
-
COUNT(*) as ExecutionCount
|
|
2399
|
-
FROM AIPromptRun
|
|
2400
|
-
WHERE RunAt > DATEADD(day, -7, GETDATE())
|
|
2401
|
-
GROUP BY RunType;
|
|
2402
|
-
```
|
|
2403
|
-
|
|
2404
|
-
## License
|
|
2405
|
-
|
|
2406
|
-
ISC
|
|
2407
|
-
|
|
2408
|
-
---
|
|
2409
|
-
|
|
2410
|
-
## Advanced Configuration
|
|
2411
|
-
|
|
2412
|
-
### Cache Configuration
|
|
2413
|
-
|
|
2414
|
-
```typescript
|
|
2415
|
-
const cacheOptimizedPrompt = {
|
|
2416
|
-
EnableCaching: true,
|
|
2417
|
-
CacheMatchType: "Vector", // Vector similarity matching
|
|
2418
|
-
CacheTTLSeconds: 3600, // 1 hour cache
|
|
2419
|
-
CacheSimilarityThreshold: 0.9, // High similarity required
|
|
2420
|
-
CacheMustMatchModel: true, // Model must match
|
|
2421
|
-
CacheMustMatchVendor: false, // Vendor can differ
|
|
2422
|
-
CacheMustMatchAgent: false, // Agent can differ
|
|
2423
|
-
CacheMustMatchConfig: true // Configuration must match
|
|
2424
|
-
};
|
|
2425
|
-
```
|
|
2426
|
-
|
|
2427
|
-
### Model Selection Strategies
|
|
2428
|
-
|
|
2429
|
-
```typescript
|
|
2430
|
-
// By power ranking
|
|
2431
|
-
const powerBasedPrompt = {
|
|
2432
|
-
SelectionStrategy: "ByPower",
|
|
2433
|
-
PowerPreference: "Highest", // or "Lowest"
|
|
2434
|
-
MinPowerRank: 80 // Minimum capability required
|
|
2435
|
-
};
|
|
2436
|
-
|
|
2437
|
-
// Specific models
|
|
2438
|
-
const specificModelsPrompt = {
|
|
2439
|
-
SelectionStrategy: "Specific",
|
|
2440
|
-
// Models defined in AIPromptModel entries
|
|
2441
|
-
};
|
|
2442
|
-
|
|
2443
|
-
// Default system selection
|
|
2444
|
-
const defaultPrompt = {
|
|
2445
|
-
SelectionStrategy: "Default"
|
|
2446
|
-
};
|
|
2447
|
-
```
|
|
2448
|
-
|
|
2449
|
-
### Model Priority Behavior
|
|
2450
|
-
|
|
2451
|
-
When multiple models are configured for a prompt (using `AIPromptModel` entries), the execution engine uses the `Priority` field to determine the order in which models are tried:
|
|
2452
|
-
|
|
2453
|
-
- **Higher priority numbers are executed first** (e.g., Priority 10 before Priority 5)
|
|
2454
|
-
- The engine sorts models by `Priority DESC` to establish execution order
|
|
2455
|
-
- For models with the same priority, creation date is used as a tiebreaker
|
|
2456
|
-
- The first successful model execution is used (fail-fast approach)
|
|
2457
|
-
|
|
2458
|
-
```typescript
|
|
2459
|
-
// Example: Models are tried in this order based on Priority
|
|
2460
|
-
const promptModels = [
|
|
2461
|
-
{ ModelID: 'gpt-4-id', Priority: 100 }, // Tried first
|
|
2462
|
-
{ ModelID: 'claude-3-id', Priority: 50 }, // Tried second
|
|
2463
|
-
{ ModelID: 'gpt-3.5-id', Priority: 10 } // Tried third
|
|
2464
|
-
];
|
|
2465
|
-
|
|
2466
|
-
// The PromptRunner sorts internally using:
|
|
2467
|
-
// models.sort((a, b) => b.Priority - a.Priority)
|
|
2468
|
-
```
|
|
2469
|
-
|
|
2470
|
-
This priority system allows you to:
|
|
2471
|
-
- Set preferred models with higher priorities
|
|
2472
|
-
- Configure fallback models with lower priorities
|
|
2473
|
-
- Ensure expensive/powerful models are only used when necessary
|
|
2474
|
-
- Control the exact execution order for cost optimization
|
|
2475
|
-
|
|
2476
|
-
## Multi-Vendor Model Support
|
|
2477
|
-
|
|
2478
|
-
MemberJunction supports multiple inference providers (vendors) for the same AI model, enabling flexible deployment scenarios and vendor failover. This is crucial for:
|
|
2479
|
-
- Using the same model from different providers (e.g., Claude from Anthropic API vs AWS Bedrock)
|
|
2480
|
-
- Vendor-specific configurations (different token limits, API endpoints, pricing)
|
|
2481
|
-
- Failover strategies when primary vendors are unavailable
|
|
2482
|
-
- Cost optimization by routing to cheaper providers
|
|
2483
|
-
|
|
2484
|
-
### Vendor Selection Precedence
|
|
2485
|
-
|
|
2486
|
-
The AI Prompt Runner uses a sophisticated vendor selection system with clear precedence rules:
|
|
2487
|
-
|
|
2488
|
-
1. **Explicit Runtime Override** (`params.override.vendorId`)
|
|
2489
|
-
- Highest precedence - always used when specified
|
|
2490
|
-
- If the vendor doesn't provide the selected model, a warning is issued and fallback occurs
|
|
2491
|
-
|
|
2492
|
-
2. **Prompt-Model Association Vendor** (`AIPromptModel.VendorID`)
|
|
2493
|
-
- When a prompt has specific model associations, the vendor from the highest priority association is used
|
|
2494
|
-
- Configured in the database for prompt-specific vendor preferences
|
|
2495
|
-
|
|
2496
|
-
3. **Highest Priority Vendor** (`AIModelVendor.Priority`)
|
|
2497
|
-
- When no explicit vendor is specified, the system uses the vendor with the highest priority value
|
|
2498
|
-
- Priority is configured per model-vendor combination in the database
|
|
2499
|
-
- **Higher Priority numbers = Higher preference** (e.g., Priority 100 is preferred over Priority 50)
|
|
2500
|
-
|
|
2501
|
-
### Vendor-Specific Configuration
|
|
2502
|
-
|
|
2503
|
-
Each vendor can have different configurations for the same model:
|
|
2504
|
-
|
|
2505
|
-
```typescript
|
|
2506
|
-
// AIModelVendor entity fields used during execution:
|
|
2507
|
-
{
|
|
2508
|
-
ModelID: 'claude-3-opus-id',
|
|
2509
|
-
VendorID: 'anthropic-id',
|
|
2510
|
-
Priority: 100, // Higher = preferred
|
|
2511
|
-
DriverClass: 'AnthropicLLM', // Vendor-specific implementation
|
|
2512
|
-
APIName: 'claude-3-opus-20240229', // Vendor's API model name
|
|
2513
|
-
MaxInputTokens: 200000, // Vendor-specific limits
|
|
2514
|
-
MaxOutputTokens: 4096,
|
|
2515
|
-
SupportsStreaming: true,
|
|
2516
|
-
SupportsEffortLevel: false
|
|
2517
|
-
}
|
|
2518
|
-
```
|
|
2519
|
-
|
|
2520
|
-
### Vendor Override Example
|
|
2521
|
-
|
|
2522
|
-
```typescript
|
|
2523
|
-
// Execute with specific vendor
|
|
2524
|
-
const result = await runner.ExecutePrompt({
|
|
2525
|
-
prompt: myPrompt,
|
|
2526
|
-
data: { query: 'Analyze this data' },
|
|
2527
|
-
contextUser: currentUser,
|
|
2528
|
-
override: {
|
|
2529
|
-
modelId: 'claude-3-opus-id',
|
|
2530
|
-
vendorId: 'aws-bedrock-id' // Use AWS Bedrock instead of default
|
|
2531
|
-
}
|
|
2532
|
-
});
|
|
2533
|
-
|
|
2534
|
-
// The system will:
|
|
2535
|
-
// 1. Validate that AWS Bedrock provides Claude 3 Opus
|
|
2536
|
-
// 2. Use AWS Bedrock's driver class and configuration
|
|
2537
|
-
// 3. Fall back to highest priority vendor if mismatch occurs
|
|
2538
|
-
```
|
|
2539
|
-
|
|
2540
|
-
### Model-Vendor Mismatch Handling
|
|
2541
|
-
|
|
2542
|
-
When a specified vendor doesn't provide the requested model:
|
|
2543
|
-
|
|
2544
|
-
```typescript
|
|
2545
|
-
// Scenario: User requests GPT-4 from Anthropic (invalid combination)
|
|
2546
|
-
const result = await runner.ExecutePrompt({
|
|
2547
|
-
prompt: myPrompt,
|
|
2548
|
-
override: {
|
|
2549
|
-
modelId: 'gpt-4-id',
|
|
2550
|
-
vendorId: 'anthropic-id' // Anthropic doesn't provide GPT-4
|
|
2551
|
-
}
|
|
2552
|
-
});
|
|
2553
|
-
|
|
2554
|
-
// System behavior:
|
|
2555
|
-
// 1. Logs warning: "⚠️ Warning: Vendor anthropic does not provide model GPT-4. Falling back to highest priority vendor."
|
|
2556
|
-
// 2. Finds highest priority vendor for GPT-4 (e.g., OpenAI)
|
|
2557
|
-
// 3. Uses the fallback vendor's configuration
|
|
2558
|
-
// 4. Execution continues with valid vendor
|
|
2559
|
-
```
|
|
2560
|
-
|
|
2561
|
-
### Database Configuration
|
|
2562
|
-
|
|
2563
|
-
Configure multi-vendor support through the MemberJunction metadata:
|
|
2564
|
-
|
|
2565
|
-
```sql
|
|
2566
|
-
-- Example: Claude available from multiple vendors
|
|
2567
|
-
INSERT INTO [MJ: AI Model Vendors] (ModelID, VendorID, Priority, DriverClass, APIName, MaxInputTokens)
|
|
2568
|
-
VALUES
|
|
2569
|
-
('claude-3-id', 'anthropic-id', 100, 'AnthropicLLM', 'claude-3-opus-20240229', 200000),
|
|
2570
|
-
('claude-3-id', 'aws-bedrock-id', 90, 'BedrockLLM', 'anthropic.claude-3-opus', 180000),
|
|
2571
|
-
('claude-3-id', 'vertex-ai-id', 80, 'VertexAILLM', 'claude-3-opus@001', 150000);
|
|
2572
|
-
|
|
2573
|
-
-- Query vendor options for a model
|
|
2574
|
-
SELECT
|
|
2575
|
-
mv.Priority,
|
|
2576
|
-
v.Name as VendorName,
|
|
2577
|
-
mv.DriverClass,
|
|
2578
|
-
mv.APIName,
|
|
2579
|
-
mv.MaxInputTokens,
|
|
2580
|
-
mv.MaxOutputTokens
|
|
2581
|
-
FROM [MJ: AI Model Vendors] mv
|
|
2582
|
-
JOIN [MJ: AI Vendors] v ON mv.VendorID = v.ID
|
|
2583
|
-
WHERE mv.ModelID = 'claude-3-id'
|
|
2584
|
-
AND mv.Status = 'Active'
|
|
2585
|
-
ORDER BY mv.Priority DESC;
|
|
2586
|
-
```
|
|
2587
|
-
|
|
2588
|
-
### Benefits of Multi-Vendor Support
|
|
2589
|
-
|
|
2590
|
-
1. **Resilience**: Automatic failover when primary vendor is unavailable
|
|
2591
|
-
2. **Cost Optimization**: Route to cheaper vendors for non-critical tasks
|
|
2592
|
-
3. **Regional Compliance**: Use region-specific vendors for data residency
|
|
2593
|
-
4. **Performance**: Choose vendors with better latency for your location
|
|
2594
|
-
5. **Feature Access**: Some vendors may offer unique features (streaming, tools, etc.)
|
|
2595
|
-
|
|
2596
|
-
### Implementation Details
|
|
2597
|
-
|
|
2598
|
-
The multi-vendor support is implemented in the `executeModel` method:
|
|
2599
|
-
|
|
2600
|
-
1. **Vendor Resolution**: Determines which vendor to use based on precedence rules
|
|
2601
|
-
2. **Configuration Loading**: Loads vendor-specific settings from AIModelVendor
|
|
2602
|
-
3. **Driver Selection**: Uses vendor-specific driver class for API communication
|
|
2603
|
-
4. **Fallback Logic**: Handles mismatches gracefully with warnings
|
|
2604
|
-
5. **Execution Tracking**: Records the actual vendor used in AIPromptRun
|
|
2605
|
-
|
|
2606
|
-
For additional configuration options and advanced use cases, refer to the source code and entity definitions in the MemberJunction core system.
|
|
2607
|
-
|
|
2608
|
-
## System Prompt Embedding
|
|
2609
|
-
|
|
2610
|
-
The AI Prompt Runner provides sophisticated system prompt embedding capabilities for agent architectures through the template engine integration.
|
|
2611
|
-
|
|
2612
|
-
### Architecture Overview
|
|
2613
|
-
|
|
2614
|
-
When `systemPromptId` is provided in AIPromptParams, the runner:
|
|
2615
|
-
1. Loads the system prompt template from the database
|
|
2616
|
-
2. Embeds the agent-specific AI prompt using `{% PromptEmbed %}` syntax
|
|
2617
|
-
3. Renders the complete system prompt with agent context
|
|
2618
|
-
4. Uses the rendered system prompt instead of the regular AI prompt template
|
|
2619
|
-
|
|
2620
|
-
This enables sophisticated agent architectures where:
|
|
2621
|
-
- **System prompts** provide execution control and enforce deterministic JSON response format
|
|
2622
|
-
- **Agent prompts** contain domain-specific logic (e.g., DATA_GATHER instructions)
|
|
2623
|
-
- **Available actions and sub-agents** are injected for agent decision-making
|
|
2624
|
-
|
|
2625
|
-
### Template Syntax
|
|
2626
|
-
|
|
2627
|
-
System prompt templates use the `{% PromptEmbed %}` syntax to embed AI prompts:
|
|
2628
|
-
|
|
2629
|
-
```nunjucks
|
|
2630
|
-
# System Prompt Template Example
|
|
2631
|
-
|
|
2632
|
-
You are an AI agent with the following specialized instructions:
|
|
2633
|
-
|
|
2634
|
-
{% PromptEmbed %}
|
|
2635
|
-
|
|
2636
|
-
## Available Actions
|
|
2637
|
-
{{#each availableActions}}
|
|
2638
|
-
- **{{this.name}}**: {{this.description}}
|
|
2639
|
-
{{/each}}
|
|
2640
|
-
|
|
2641
|
-
## Available Sub-Agents
|
|
2642
|
-
{{#each availableSubAgents}}
|
|
2643
|
-
- **{{this.name}}**: {{this.description}}
|
|
2644
|
-
{{/each}}
|
|
2645
|
-
|
|
2646
|
-
## Response Format
|
|
2647
|
-
You must respond with valid JSON following this structure:
|
|
2648
|
-
{
|
|
2649
|
-
"decision": "execute_action|execute_subagent|complete_task|request_clarification",
|
|
2650
|
-
"reasoning": "Explanation of your decision",
|
|
2651
|
-
"executionPlan": [
|
|
2652
|
-
{
|
|
2653
|
-
"type": "action|subagent",
|
|
2654
|
-
"targetId": "action-or-agent-id",
|
|
2655
|
-
"parameters": {},
|
|
2656
|
-
"executionOrder": 1,
|
|
2657
|
-
"allowParallel": true
|
|
2658
|
-
}
|
|
2659
|
-
],
|
|
2660
|
-
"isTaskComplete": false,
|
|
2661
|
-
"confidence": 0.95
|
|
2662
|
-
}
|
|
2663
|
-
```
|
|
2664
|
-
|
|
2665
|
-
### Validation and Security
|
|
2666
|
-
|
|
2667
|
-
The system includes comprehensive validation to ensure proper prompt embedding:
|
|
2668
|
-
|
|
2669
|
-
```typescript
|
|
2670
|
-
// Validation process:
|
|
2671
|
-
// 1. Verify system prompt exists and has template
|
|
2672
|
-
// 2. Check agent-prompt relationships via AIAgentPrompt table
|
|
2673
|
-
// 3. Ensure agents using system prompt are linked to current prompt
|
|
2674
|
-
// 4. Validate template contains {% PromptEmbed %} syntax
|
|
2675
|
-
|
|
2676
|
-
const params = new AIPromptParams();
|
|
2677
|
-
params.prompt = agentSpecificPrompt;
|
|
2678
|
-
params.systemPromptId = 'system-prompt-id'; // Triggers validation
|
|
2679
|
-
params.data = { agentName: 'DataGather', availableActions: [...] };
|
|
2680
|
-
```
|
|
2681
|
-
|
|
2682
|
-
### Integration with AI Agents
|
|
2683
|
-
|
|
2684
|
-
The AgentRunner seamlessly uses system prompt embedding:
|
|
2685
|
-
|
|
2686
|
-
```typescript
|
|
2687
|
-
// AgentRunner delegates to AIPromptRunner with system prompt embedding
|
|
2688
|
-
const promptParams = new AIPromptParams();
|
|
2689
|
-
promptParams.prompt = primaryAgentPrompt.prompt;
|
|
2690
|
-
promptParams.systemPromptId = this.agentType.SystemPromptID;
|
|
2691
|
-
promptParams.data = promptData;
|
|
2692
|
-
promptParams.agentRunId = context.agentRun.ID;
|
|
2693
|
-
|
|
2694
|
-
const promptResult = await this._promptRunner.ExecutePrompt(promptParams);
|
|
2695
|
-
```
|
|
2696
|
-
|
|
2697
|
-
### Database Schema Integration
|
|
2698
|
-
|
|
2699
|
-
The system prompt embedding feature integrates with several database entities:
|
|
2700
|
-
|
|
2701
|
-
#### Entity Relationships
|
|
2702
|
-
|
|
2703
|
-
```sql
|
|
2704
|
-
-- System prompts are stored as AIPrompt entities with templates
|
|
2705
|
-
AIPrompt (SystemPromptID) -> Template -> TemplateContent (contains {% PromptEmbed %})
|
|
2706
|
-
|
|
2707
|
-
-- Agent types reference system prompts
|
|
2708
|
-
AIAgentType.SystemPromptID -> AIPrompt (system prompt)
|
|
2709
|
-
|
|
2710
|
-
-- Agents belong to agent types
|
|
2711
|
-
AIAgent.TypeID -> AIAgentType
|
|
2712
|
-
|
|
2713
|
-
-- Agent prompts link agents to their specific prompts
|
|
2714
|
-
AIAgentPrompt: AgentID + PromptID
|
|
2715
|
-
|
|
2716
|
-
-- Validation ensures proper linkage:
|
|
2717
|
-
-- Agent -> AgentType -> SystemPrompt
|
|
2718
|
-
-- Agent -> AIAgentPrompt -> AIPrompt (to be embedded)
|
|
2719
|
-
```
|
|
2720
|
-
|
|
2721
|
-
#### Storage Structure
|
|
2722
|
-
|
|
2723
|
-
```sql
|
|
2724
|
-
-- Example system prompt template storage
|
|
2725
|
-
INSERT INTO Template (Name, Description)
|
|
2726
|
-
VALUES ('AI Agent System Prompt', 'Control wrapper for agent decision-making');
|
|
2727
|
-
|
|
2728
|
-
INSERT INTO TemplateContent (TemplateID, TemplateText, Priority)
|
|
2729
|
-
VALUES (@TemplateID, 'You are {{agentName}}... {% PromptEmbed %}... Respond with JSON...', 100);
|
|
2730
|
-
|
|
2731
|
-
INSERT INTO AIPrompt (Name, Description, TemplateID, Category)
|
|
2732
|
-
VALUES ('System Prompt', 'Agent execution control wrapper', @TemplateID, 'System');
|
|
2733
|
-
|
|
2734
|
-
UPDATE AIAgentType SET SystemPromptID = @SystemPromptID WHERE Name = 'DataGatherAgent';
|
|
2735
|
-
```
|
|
2736
|
-
|
|
2737
|
-
## API Keys
|
|
2738
|
-
|
|
2739
|
-
The AI Prompts system provides flexible API key management for runtime configuration without modifying environment variables or global settings.
|
|
2740
|
-
|
|
2741
|
-
### Environment-Based API Configuration
|
|
2742
|
-
|
|
2743
|
-
MemberJunction now supports sophisticated environment-based API key resolution through AIConfigSet entities, allowing different configurations for different environments:
|
|
2744
|
-
|
|
2745
|
-
```typescript
|
|
2746
|
-
// Configurations are loaded based on the environment name
|
|
2747
|
-
// Default fallback order: process.env.NODE_ENV -> 'production'
|
|
2748
|
-
|
|
2749
|
-
// Example: Different API keys per environment
|
|
2750
|
-
// Development -> AIConfigSet(Name='development') -> Config Key OPENAI_LLM_APIKEY
|
|
2751
|
-
// Production -> AIConfigSet(Name='production') -> Config Key OPENAI_LLM_APIKEY
|
|
2752
|
-
```
|
|
2753
|
-
|
|
2754
|
-
#### Configuration Set Priority
|
|
2755
|
-
|
|
2756
|
-
When multiple configuration sets exist, they are evaluated in this order:
|
|
2757
|
-
1. **Exact environment match** (e.g., 'development' matches 'development')
|
|
2758
|
-
2. **Priority field** (higher priority values are preferred)
|
|
2759
|
-
3. **Custom resolver logic** (if implemented via subclassing)
|
|
2760
|
-
|
|
2761
|
-
### Using Runtime API Keys
|
|
2762
|
-
|
|
2763
|
-
You can provide API keys at prompt execution time, which is useful for:
|
|
2764
|
-
- Multi-tenant applications where different users have different API keys
|
|
2765
|
-
- Testing with different API providers or accounts
|
|
2766
|
-
- Isolating API usage by application or department
|
|
2767
|
-
- Temporary API key usage for specific operations
|
|
2768
|
-
|
|
2769
|
-
```typescript
|
|
2770
|
-
import { AIPromptRunner, AIPromptParams } from '@memberjunction/ai-prompts';
|
|
2771
|
-
import { AIAPIKey } from '@memberjunction/ai';
|
|
2772
|
-
|
|
2773
|
-
const runner = new AIPromptRunner();
|
|
2774
|
-
|
|
2775
|
-
// Execute with specific API keys
|
|
2776
|
-
const result = await runner.ExecutePrompt({
|
|
2777
|
-
prompt: myPrompt,
|
|
2778
|
-
data: { query: 'Analyze this data' },
|
|
2779
|
-
contextUser: currentUser,
|
|
2780
|
-
apiKeys: [
|
|
2781
|
-
{ driverClass: 'OpenAILLM', apiKey: 'sk-user-specific-key' },
|
|
2782
|
-
{ driverClass: 'AnthropicLLM', apiKey: 'sk-ant-department-key' }
|
|
2783
|
-
]
|
|
2784
|
-
});
|
|
2785
|
-
```
|
|
2786
|
-
|
|
2787
|
-
### API Key Precedence
|
|
2788
|
-
|
|
2789
|
-
When executing prompts, API keys are resolved in this order:
|
|
2790
|
-
1. **Local API keys** provided in `AIPromptParams.apiKeys` (highest priority)
|
|
2791
|
-
2. **Configuration sets** from database based on environment
|
|
2792
|
-
3. **Environment variables** (traditional dotenv approach)
|
|
2793
|
-
4. **Custom implementations** via AIAPIKeys subclassing
|
|
2794
|
-
|
|
2795
|
-
### Configuration Set Structure
|
|
2796
|
-
|
|
2797
|
-
Configuration sets are managed through these entities:
|
|
2798
|
-
|
|
2799
|
-
#### AIConfigSet
|
|
2800
|
-
- **Name**: Environment name (e.g., 'development', 'production', 'staging')
|
|
2801
|
-
- **Description**: Human-readable description
|
|
2802
|
-
- **Priority**: Higher values take precedence when multiple sets match
|
|
2803
|
-
- **Status**: 'Active' or 'Inactive'
|
|
2804
|
-
|
|
2805
|
-
#### AIConfiguration
|
|
2806
|
-
- **ConfigSetID**: Links to parent configuration set
|
|
2807
|
-
- **ConfigKey**: The configuration key (e.g., 'OPENAI_LLM_APIKEY')
|
|
2808
|
-
- **ConfigValue**: The actual value (encrypted for sensitive data)
|
|
2809
|
-
- **EncryptedValue**: Whether the value is encrypted
|
|
2810
|
-
- **Description**: Documentation for the configuration
|
|
2811
|
-
|
|
2812
|
-
### Environment Variables vs Configuration Sets
|
|
2813
|
-
|
|
2814
|
-
```typescript
|
|
2815
|
-
// Traditional environment variable approach (still supported)
|
|
2816
|
-
process.env.OPENAI_LLM_APIKEY = 'sk-...';
|
|
2817
|
-
|
|
2818
|
-
// New configuration set approach (recommended)
|
|
2819
|
-
// Stored in database:
|
|
2820
|
-
// ConfigSet: { Name: 'production', Priority: 100 }
|
|
2821
|
-
// Config: { ConfigKey: 'OPENAI_LLM_APIKEY', ConfigValue: 'sk-...', Encrypted: true }
|
|
2822
|
-
|
|
2823
|
-
// The system automatically uses the configuration set if available
|
|
2824
|
-
// Falls back to environment variables if not found in database
|
|
2825
|
-
```
|
|
2826
|
-
|
|
2827
|
-
### Hierarchical API Key Propagation
|
|
2828
|
-
|
|
2829
|
-
For hierarchical prompt execution (prompts with child prompts), API keys are automatically propagated:
|
|
2830
|
-
|
|
2831
|
-
```typescript
|
|
2832
|
-
const parentParams = new AIPromptParams();
|
|
2833
|
-
parentParams.prompt = parentPrompt;
|
|
2834
|
-
parentParams.childPrompts = [
|
|
2835
|
-
new ChildPromptParam(childPrompt1, 'analysis'),
|
|
2836
|
-
new ChildPromptParam(childPrompt2, 'summary')
|
|
2837
|
-
];
|
|
2838
|
-
parentParams.apiKeys = [
|
|
2839
|
-
{ driverClass: 'OpenAILLM', apiKey: 'sk-parent-key' }
|
|
2840
|
-
];
|
|
2841
|
-
|
|
2842
|
-
// Child prompts automatically inherit the parent's API keys
|
|
2843
|
-
// This ensures consistent API key usage throughout the execution tree
|
|
2844
|
-
const result = await runner.ExecutePrompt(parentParams);
|
|
2845
|
-
```
|
|
2846
|
-
|
|
2847
|
-
### Custom Global API Key Management
|
|
2848
|
-
|
|
2849
|
-
For advanced scenarios, you can subclass the global `AIAPIKeys` object:
|
|
2850
|
-
|
|
2851
|
-
```typescript
|
|
2852
|
-
import { AIAPIKeys, RegisterClass } from '@memberjunction/ai';
|
|
2853
|
-
|
|
2854
|
-
@RegisterClass(AIAPIKeys, 'CustomAPIKeys', 2) // Priority 2 overrides default
|
|
2855
|
-
export class CustomAPIKeys extends AIAPIKeys {
|
|
2856
|
-
public GetAPIKey(AIDriverName: string): string {
|
|
2857
|
-
// First check configuration sets
|
|
2858
|
-
const configValue = this.getFromConfigSet(AIDriverName);
|
|
2859
|
-
if (configValue) return configValue;
|
|
2860
|
-
|
|
2861
|
-
// Then check custom logic: database lookup, vault access, etc.
|
|
2862
|
-
if (AIDriverName === 'OpenAILLM') {
|
|
2863
|
-
return this.getFromVault('openai-key');
|
|
2864
|
-
}
|
|
2865
|
-
|
|
2866
|
-
// Finally fall back to environment variables
|
|
2867
|
-
return super.GetAPIKey(AIDriverName);
|
|
2868
|
-
}
|
|
2869
|
-
|
|
2870
|
-
private getFromConfigSet(driverName: string): string | null {
|
|
2871
|
-
// Implementation would query AIConfiguration entities
|
|
2872
|
-
// based on current environment
|
|
2873
|
-
const envName = process.env.NODE_ENV || 'production';
|
|
2874
|
-
// ... query logic here
|
|
2875
|
-
return null;
|
|
2876
|
-
}
|
|
2877
|
-
}
|
|
2878
|
-
```
|
|
2879
|
-
|
|
2880
|
-
### Multi-Environment Setup Example
|
|
2881
|
-
|
|
2882
|
-
```typescript
|
|
2883
|
-
// Development environment setup
|
|
2884
|
-
const devConfig = {
|
|
2885
|
-
ConfigSet: { Name: 'development', Priority: 100 },
|
|
2886
|
-
Configurations: [
|
|
2887
|
-
{ ConfigKey: 'OPENAI_LLM_APIKEY', ConfigValue: 'sk-dev-...', Encrypted: true },
|
|
2888
|
-
{ ConfigKey: 'ANTHROPIC_LLM_APIKEY', ConfigValue: 'sk-ant-dev-...', Encrypted: true },
|
|
2889
|
-
{ ConfigKey: 'LOG_LEVEL', ConfigValue: 'debug', Encrypted: false }
|
|
2890
|
-
]
|
|
2891
|
-
};
|
|
2892
|
-
|
|
2893
|
-
// Production environment setup
|
|
2894
|
-
const prodConfig = {
|
|
2895
|
-
ConfigSet: { Name: 'production', Priority: 100 },
|
|
2896
|
-
Configurations: [
|
|
2897
|
-
{ ConfigKey: 'OPENAI_LLM_APIKEY', ConfigValue: 'sk-prod-...', Encrypted: true },
|
|
2898
|
-
{ ConfigKey: 'ANTHROPIC_LLM_APIKEY', ConfigValue: 'sk-ant-prod-...', Encrypted: true },
|
|
2899
|
-
{ ConfigKey: 'LOG_LEVEL', ConfigValue: 'error', Encrypted: false }
|
|
2900
|
-
]
|
|
2901
|
-
};
|
|
2902
|
-
|
|
2903
|
-
// The system automatically loads the correct configuration based on NODE_ENV
|
|
2904
|
-
```
|
|
2905
|
-
|
|
2906
|
-
### Security Best Practices
|
|
2907
|
-
|
|
2908
|
-
1. **Never hardcode API keys** in your source code
|
|
2909
|
-
2. **Use encrypted storage** for sensitive configuration values
|
|
2910
|
-
3. **Separate configurations** by environment (dev, staging, prod)
|
|
2911
|
-
4. **Rotate keys regularly** and update configuration sets
|
|
2912
|
-
5. **Monitor API key usage** to detect unauthorized access
|
|
2913
|
-
6. **Use the principle of least privilege** - give each user/app only the keys they need
|
|
2914
|
-
7. **Audit configuration changes** through MemberJunction's change tracking
|
|
2915
|
-
|
|
2916
|
-
## AI Configuration System
|
|
2917
|
-
|
|
2918
|
-
The AI Prompts system supports sophisticated environment-specific model selection through AI Configurations. This allows you to use different sets of models for the same prompts based on the active configuration (e.g., Production vs Development).
|
|
2919
|
-
|
|
2920
|
-
### Configuration Concepts
|
|
2921
|
-
|
|
2922
|
-
#### AI Configurations
|
|
2923
|
-
- Named configuration sets (e.g., "Production", "Development", "Europe-Region")
|
|
2924
|
-
- Filter which models are available for prompt execution
|
|
2925
|
-
- Support configuration parameters for dynamic behavior
|
|
2926
|
-
- One configuration can be marked as default (`IsDefault = true`)
|
|
2927
|
-
|
|
2928
|
-
#### AI Configuration Parameters
|
|
2929
|
-
- Name-value pairs stored per configuration
|
|
2930
|
-
- Support different data types (string, number, boolean, date, object)
|
|
2931
|
-
- Can control dynamic parallelization and other runtime behavior
|
|
2932
|
-
- Example: `ParallelExecutions = 5` for development testing
|
|
2933
|
-
|
|
2934
|
-
### Model Selection with Configurations
|
|
2935
|
-
|
|
2936
|
-
When executing a prompt with a specific `configurationId`, the system follows a two-phase model selection process that ensures configuration-specific models ALWAYS take precedence:
|
|
2937
|
-
|
|
2938
|
-
#### Phase 1: Configuration-Specific Models (Highest Priority)
|
|
2939
|
-
```typescript
|
|
2940
|
-
// First, try to find models with matching configuration
|
|
2941
|
-
const configModels = promptModels.filter(pm =>
|
|
2942
|
-
pm.ConfigurationID === configurationId &&
|
|
2943
|
-
(pm.Status === 'Active' || pm.Status === 'Preview')
|
|
2944
|
-
);
|
|
2945
|
-
```
|
|
2946
|
-
|
|
2947
|
-
If configuration-specific models are found, they are used exclusively. The system will NOT consider models with NULL or mismatched configurations.
|
|
2948
|
-
|
|
2949
|
-
#### Phase 2: Default Models (Fallback)
|
|
2950
|
-
|
|
2951
|
-
Only if NO models match the specific configuration, the system falls back to models with NULL configuration:
|
|
2952
|
-
|
|
2953
|
-
```typescript
|
|
2954
|
-
// Fall back to NULL configuration models only if no matches found
|
|
2955
|
-
if (configModels.length === 0) {
|
|
2956
|
-
const defaultModels = promptModels.filter(pm =>
|
|
2957
|
-
pm.ConfigurationID === null &&
|
|
2958
|
-
(pm.Status === 'Active' || pm.Status === 'Preview')
|
|
2959
|
-
);
|
|
2960
|
-
}
|
|
2961
|
-
```
|
|
2962
|
-
|
|
2963
|
-
This ensures:
|
|
2964
|
-
- **Configuration-specific models always win**: When you specify a configuration, only models explicitly assigned to that configuration are considered first
|
|
2965
|
-
- **Clear environment separation**: Production models never mix with Development models when configurations are used
|
|
2966
|
-
- **Explicit fallback behavior**: NULL configuration models serve as defaults only when no configuration-specific models exist
|
|
2967
|
-
|
|
2968
|
-
### Configuration Setup Example
|
|
2969
|
-
|
|
2970
|
-
```sql
|
|
2971
|
-
-- Create configurations
|
|
2972
|
-
INSERT INTO AIConfiguration (Name, Description, IsDefault, Status)
|
|
2973
|
-
VALUES
|
|
2974
|
-
('Production', 'Stable models for production use', 1, 'Active'),
|
|
2975
|
-
('Development', 'Experimental models for testing', 0, 'Active');
|
|
2976
|
-
|
|
2977
|
-
-- Assign models to configurations
|
|
2978
|
-
-- Production uses GPT-4
|
|
2979
|
-
UPDATE AIPromptModel
|
|
2980
|
-
SET ConfigurationID = (SELECT ID FROM AIConfiguration WHERE Name = 'Production')
|
|
2981
|
-
WHERE PromptID = @PromptID AND ModelID = @GPT4ModelID;
|
|
2982
|
-
|
|
2983
|
-
-- Development uses GPT-4-Turbo and Claude-3-Opus
|
|
2984
|
-
UPDATE AIPromptModel
|
|
2985
|
-
SET ConfigurationID = (SELECT ID FROM AIConfiguration WHERE Name = 'Development')
|
|
2986
|
-
WHERE PromptID = @PromptID AND ModelID IN (@GPT4TurboID, @Claude3OpusID);
|
|
2987
|
-
|
|
2988
|
-
-- Models with NULL ConfigurationID serve as defaults when no configuration matches
|
|
2989
|
-
UPDATE AIPromptModel
|
|
2990
|
-
SET ConfigurationID = NULL
|
|
2991
|
-
WHERE PromptID = @PromptID AND ModelID = @FallbackModelID;
|
|
2992
|
-
```
|
|
2993
|
-
|
|
2994
|
-
### Using Configurations
|
|
2995
|
-
|
|
2996
|
-
#### In Code
|
|
2997
|
-
```typescript
|
|
2998
|
-
const result = await promptRunner.ExecutePrompt({
|
|
2999
|
-
prompt: myPrompt,
|
|
3000
|
-
configurationId: 'dev-config-id', // Optional
|
|
3001
|
-
data: { query: 'Analyze this data' },
|
|
3002
|
-
contextUser: currentUser
|
|
3003
|
-
});
|
|
3004
|
-
```
|
|
3005
|
-
|
|
3006
|
-
#### Configuration Precedence
|
|
3007
|
-
1. **Configuration-specific models** (highest priority) - When a configurationId is provided, ONLY models with matching ConfigurationID are considered initially
|
|
3008
|
-
2. **NULL configuration models** (fallback only) - Used only when NO models match the specified configuration
|
|
3009
|
-
3. **Priority within each phase** - Models are ranked by their Priority field (higher number = higher priority)
|
|
3010
|
-
|
|
3011
|
-
### Dynamic Parallelization with Configurations
|
|
3012
|
-
|
|
3013
|
-
Configurations can control parallel execution through parameters:
|
|
3014
|
-
|
|
3015
|
-
```typescript
|
|
3016
|
-
// Set up configuration parameter
|
|
3017
|
-
INSERT INTO AIConfigurationParam (ConfigurationID, Name, Type, Value)
|
|
3018
|
-
VALUES (@DevConfigID, 'ParallelExecutions', 'number', '5');
|
|
3019
|
-
|
|
3020
|
-
// Use in prompt setup
|
|
3021
|
-
UPDATE AIPrompt
|
|
3022
|
-
SET ParallelizationMode = 'ConfigParam',
|
|
3023
|
-
ParallelConfigParam = 'ParallelExecutions'
|
|
3024
|
-
WHERE ID = @PromptID;
|
|
3025
|
-
```
|
|
3026
|
-
|
|
3027
|
-
### Best Practices
|
|
3028
|
-
|
|
3029
|
-
1. **Use NULL ConfigurationID for fallback models** that provide a safety net when no configuration-specific models exist
|
|
3030
|
-
2. **Create environment-specific configurations** for different deployment scenarios (Production, Development, Testing)
|
|
3031
|
-
3. **Document configuration purposes** in the Description field to clarify their intended use
|
|
3032
|
-
4. **Test configuration precedence** to ensure configuration-specific models always take priority over defaults
|
|
3033
|
-
5. **Use configuration parameters** for environment-specific settings beyond just model selection
|
|
3034
|
-
6. **Assign models explicitly to configurations** to ensure clear separation between environments
|
|
3035
|
-
|
|
3036
|
-
For more details on API key management, see the [AI Core API Keys documentation](../Core/README.md#api-key-management).
|
|
3037
|
-
|
|
3038
|
-
### Configuration Inheritance (v3.1+)
|
|
3039
|
-
|
|
3040
|
-
AI Configurations now support parent-child inheritance relationships. This enables you to create child configurations that inherit prompt-model mappings from parent configurations while overriding specific settings.
|
|
3041
|
-
|
|
3042
|
-
#### Use Cases
|
|
3043
|
-
|
|
3044
|
-
- **Experimentation**: Create a child config that inherits from "Production" but overrides just 2-3 prompts with experimental models
|
|
3045
|
-
- **Regional Variations**: Create "Production-EU" inheriting from "Production" with region-specific model overrides
|
|
3046
|
-
- **A/B Testing**: Create test configurations that inherit baseline settings while varying specific prompts
|
|
3047
|
-
|
|
3048
|
-
#### How Inheritance Works
|
|
3049
|
-
|
|
3050
|
-
When you specify a `configurationId`, the system builds an inheritance chain from child to root:
|
|
3051
|
-
|
|
3052
|
-
```
|
|
3053
|
-
Child Config → Parent Config → Grandparent Config → ... → Root Config
|
|
3054
|
-
```
|
|
3055
|
-
|
|
3056
|
-
For each prompt, the system walks this chain looking for AIPromptModel matches:
|
|
3057
|
-
1. First checks for models assigned to the child config
|
|
3058
|
-
2. If none found, checks the parent config
|
|
3059
|
-
3. Continues up the chain until a match is found
|
|
3060
|
-
4. Falls back to NULL configuration models as final fallback
|
|
3061
|
-
|
|
3062
|
-
#### Setting Up Inheritance
|
|
3063
|
-
|
|
3064
|
-
```sql
|
|
3065
|
-
-- Create parent configuration
|
|
3066
|
-
INSERT INTO AIConfiguration (ID, Name, Description, ParentID, IsDefault, Status)
|
|
3067
|
-
VALUES (NEWID(), 'Production', 'Standard production configuration', NULL, 1, 'Active');
|
|
3068
|
-
|
|
3069
|
-
-- Create child configuration that inherits from Production
|
|
3070
|
-
INSERT INTO AIConfiguration (ID, Name, Description, ParentID, IsDefault, Status)
|
|
3071
|
-
VALUES (NEWID(), 'Production-Experimental', 'Production with experimental models for select prompts',
|
|
3072
|
-
(SELECT ID FROM AIConfiguration WHERE Name = 'Production'), 0, 'Active');
|
|
3073
|
-
```
|
|
3074
|
-
|
|
3075
|
-
#### Model Override Example
|
|
3076
|
-
|
|
3077
|
-
```sql
|
|
3078
|
-
-- Parent (Production) uses GPT-4 for the summarization prompt
|
|
3079
|
-
INSERT INTO AIPromptModel (PromptID, ModelID, ConfigurationID, Priority)
|
|
3080
|
-
VALUES (@SummarizePromptID, @GPT4ModelID, @ProductionConfigID, 100);
|
|
3081
|
-
|
|
3082
|
-
-- Child (Production-Experimental) overrides with Claude for summarization
|
|
3083
|
-
INSERT INTO AIPromptModel (PromptID, ModelID, ConfigurationID, Priority)
|
|
3084
|
-
VALUES (@SummarizePromptID, @ClaudeModelID, @ExperimentalConfigID, 100);
|
|
3085
|
-
|
|
3086
|
-
-- Other prompts in Production-Experimental inherit GPT-4 from parent
|
|
3087
|
-
```
|
|
3088
|
-
|
|
3089
|
-
#### Using Inherited Configurations
|
|
3090
|
-
|
|
3091
|
-
```typescript
|
|
3092
|
-
// Execute with child configuration - inherits from parent where no override exists
|
|
3093
|
-
const result = await promptRunner.ExecutePrompt({
|
|
3094
|
-
prompt: summarizePrompt,
|
|
3095
|
-
configurationId: experimentalConfigId, // Child config
|
|
3096
|
-
data: { text: 'Content to summarize' },
|
|
3097
|
-
contextUser: currentUser
|
|
3098
|
-
});
|
|
3099
|
-
|
|
3100
|
-
// For summarization: Uses Claude (child override)
|
|
3101
|
-
// For other prompts: Uses parent's models (inherited)
|
|
3102
|
-
```
|
|
3103
|
-
|
|
3104
|
-
#### Parameter Inheritance
|
|
3105
|
-
|
|
3106
|
-
Configuration parameters also inherit from parent configurations. Child parameters override parent parameters with the same name:
|
|
3107
|
-
|
|
3108
|
-
```typescript
|
|
3109
|
-
// Get all parameters including inherited ones
|
|
3110
|
-
const params = AIEngine.Instance.GetConfigurationParamsWithInheritance(childConfigId);
|
|
3111
|
-
|
|
3112
|
-
// Parent has: temperature=0.7, maxTokens=4000
|
|
3113
|
-
// Child has: temperature=0.9
|
|
3114
|
-
// Result: temperature=0.9 (child), maxTokens=4000 (inherited)
|
|
3115
|
-
```
|
|
3116
|
-
|
|
3117
|
-
#### Cycle Detection
|
|
3118
|
-
|
|
3119
|
-
The system automatically detects circular references in the configuration hierarchy. If a cycle is detected, an error is thrown with a descriptive message showing the problematic chain.
|
|
3120
|
-
|
|
3121
|
-
#### Best Practices
|
|
3122
|
-
|
|
3123
|
-
1. **Keep inheritance chains shallow** - 2-3 levels is usually sufficient
|
|
3124
|
-
2. **Document override intentions** - Use Description field to explain why child configs exist
|
|
3125
|
-
3. **Use meaningful names** - Name child configs to indicate their parent (e.g., "Production-EU", "Standard-Experimental")
|
|
3126
|
-
4. **Test inheritance** - Verify that child configs properly inherit unoverridden models
|
|
3127
|
-
|
|
3128
|
-
## Model Selection Tracking (v2.78+)
|
|
3129
|
-
|
|
3130
|
-
The AIPromptRunner now provides comprehensive tracking of model selection decisions through the `modelSelectionInfo` property in `AIPromptRunResult`. This feature helps developers understand:
|
|
3131
|
-
|
|
3132
|
-
- Which models were considered during selection
|
|
3133
|
-
- Why specific models were or weren't available
|
|
3134
|
-
- Which configuration influenced the selection
|
|
3135
|
-
- What selection strategy was used
|
|
3136
|
-
|
|
3137
|
-
### Enhanced Model Selection Information
|
|
3138
|
-
|
|
3139
|
-
```typescript
|
|
3140
|
-
const result = await promptRunner.ExecutePrompt(params);
|
|
3141
|
-
|
|
3142
|
-
if (result.modelSelectionInfo) {
|
|
3143
|
-
// Access the configuration that was used
|
|
3144
|
-
const config = result.modelSelectionInfo.aiConfiguration;
|
|
3145
|
-
console.log(`Configuration: ${config?.Name || 'Default'}`);
|
|
3146
|
-
|
|
3147
|
-
// See all models that were considered
|
|
3148
|
-
for (const candidate of result.modelSelectionInfo.modelsConsidered) {
|
|
3149
|
-
console.log(`Model: ${candidate.model.Name}`);
|
|
3150
|
-
console.log(` Vendor: ${candidate.vendor?.Name || 'default'}`);
|
|
3151
|
-
console.log(` Priority: ${candidate.priority}`);
|
|
3152
|
-
console.log(` Available: ${candidate.available}`);
|
|
3153
|
-
if (!candidate.available) {
|
|
3154
|
-
console.log(` Reason: ${candidate.unavailableReason}`);
|
|
3155
|
-
}
|
|
3156
|
-
}
|
|
3157
|
-
|
|
3158
|
-
// Understand the final selection
|
|
3159
|
-
console.log(`Selected: ${result.modelSelectionInfo.modelSelected.Name}`);
|
|
3160
|
-
console.log(`Vendor: ${result.modelSelectionInfo.vendorSelected?.Name}`);
|
|
3161
|
-
console.log(`Reason: ${result.modelSelectionInfo.selectionReason}`);
|
|
3162
|
-
console.log(`Strategy: ${result.modelSelectionInfo.selectionStrategy}`);
|
|
3163
|
-
}
|
|
3164
|
-
```
|
|
3165
|
-
|
|
3166
|
-
### Database Storage
|
|
3167
|
-
|
|
3168
|
-
Model selection information is stored in the `AIPromptRun.ModelSelection` field as JSON, containing:
|
|
3169
|
-
- Model and vendor IDs (not full entities)
|
|
3170
|
-
- Configuration ID and name
|
|
3171
|
-
- Array of considered models with their availability status
|
|
3172
|
-
- Selection reason and strategy
|
|
3173
|
-
|
|
3174
|
-
Additional fields track:
|
|
3175
|
-
- `SelectionStrategy`: The strategy used ('Default', 'Specific', 'ByPower')
|
|
3176
|
-
- `ModelPowerRank`: Power rank of the selected model
|
|
3177
|
-
- `Status`: Execution status ('Pending', 'Running', 'Completed', 'Failed', 'Cancelled')
|
|
3178
|
-
- `Cancelled`: Boolean flag for cancellation
|
|
3179
|
-
- `CancellationReason`: Why execution was cancelled
|
|
3180
|
-
- `ErrorDetails`: Detailed error information for failures
|
|
3181
|
-
|
|
3182
|
-
### Benefits
|
|
3183
|
-
|
|
3184
|
-
- **Debugging**: Understand why a specific model was selected
|
|
3185
|
-
- **Monitoring**: Track which models are being used across prompts
|
|
3186
|
-
- **Optimization**: Identify models that are frequently unavailable
|
|
3187
|
-
- **Compliance**: Audit model selection for regulatory requirements
|
|
3188
|
-
|
|
3189
|
-
## Version History
|
|
3190
|
-
|
|
3191
|
-
- **2.78.0** - Added model selection tracking with full entity objects in results
|
|
3192
|
-
- **2.77.0** - Enhanced status tracking and cancellation support
|
|
3193
|
-
- **2.76.0** - Added intelligent failover system
|
|
3194
|
-
- **2.75.0** - Introduced dynamic template composition
|
|
3195
|
-
- **2.50.0** - Initial release with core prompt execution
|
|
237
|
+
- `@memberjunction/ai` -- Core AI abstractions (BaseLLM, ChatParams)
|
|
238
|
+
- `@memberjunction/ai-core-plus` -- AIPromptParams, AIPromptRunResult, extended entities
|
|
239
|
+
- `@memberjunction/ai-engine-base` -- AIEngineBase metadata cache
|
|
240
|
+
- `@memberjunction/aiengine` -- AIEngine server-side operations
|
|
241
|
+
- `@memberjunction/core` -- MJ framework core
|
|
242
|
+
- `@memberjunction/core-entities` -- Generated entity classes
|
|
243
|
+
- `@memberjunction/credentials` -- Credential resolution
|
|
244
|
+
- `@memberjunction/templates` -- Template rendering engine
|
|
245
|
+
- `@memberjunction/templates-base-types` -- Template base types
|
|
246
|
+
- `json5` -- Lenient JSON parsing for repair
|