@memstack/core 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +774 -0
- package/dist/index.cjs +1991 -0
- package/dist/index.cjs.map +1 -0
- package/dist/index.d.cts +536 -0
- package/dist/index.d.ts +536 -0
- package/dist/index.js +1974 -0
- package/dist/index.js.map +1 -0
- package/package.json +63 -0
- package/src/adapters/embedding/cohere.ts +77 -0
- package/src/adapters/embedding/openai.ts +65 -0
- package/src/adapters/llm/anthropic.ts +69 -0
- package/src/adapters/llm/groq.ts +185 -0
- package/src/adapters/llm/ollama.ts +85 -0
- package/src/adapters/llm/openai.ts +178 -0
- package/src/adapters/storage/disk.ts +334 -0
- package/src/adapters/storage/memory.ts +147 -0
- package/src/adapters/storage/postgres.ts +310 -0
- package/src/adapters/storage/redis.ts +274 -0
- package/src/client.ts +236 -0
- package/src/errors.ts +47 -0
- package/src/index.ts +52 -0
- package/src/interfaces.ts +154 -0
- package/src/memory/ContextCompiler.ts +158 -0
- package/src/memory/MemoryStore.ts +266 -0
- package/src/memory/Pruner.ts +66 -0
- package/src/memory/Summarizer.ts +46 -0
- package/src/types.ts +50 -0
package/README.md
ADDED
|
@@ -0,0 +1,774 @@
|
|
|
1
|
+
# MemStack
|
|
2
|
+
|
|
3
|
+
> The open-source memory layer for AI agents — store, retrieve, summarize, and prune.
|
|
4
|
+
|
|
5
|
+
[](https://www.npmjs.com/package/@memstack/core)
|
|
6
|
+
[](https://opensource.org/licenses/MIT)
|
|
7
|
+
|
|
8
|
+
```bash
|
|
9
|
+
npm install @memstack/core
|
|
10
|
+
```
|
|
11
|
+
|
|
12
|
+
**The problem:** AI agents forget. Every interaction starts from zero. You either stuff everything into the context window (expensive, slow, degrades output quality) or the agent has no memory of past conversations.
|
|
13
|
+
|
|
14
|
+
**What MemStack does:** A persistent memory pipeline that lives between your agent and the LLM. It stores every interaction, retrieves only what's relevant, summarizes old memories to save tokens, and prunes stale ones automatically. One method call, no infrastructure required.
|
|
15
|
+
|
|
16
|
+
Think of it as the open-source alternative to [Mem0](https://mem0.ai/) — pluggable storage, bring your own LLM, zero vendor lock-in.
|
|
17
|
+
|
|
18
|
+
---
|
|
19
|
+
|
|
20
|
+
## Table of Contents
|
|
21
|
+
|
|
22
|
+
- [Why MemStack](#why-memstack)
|
|
23
|
+
- [Quick Start](#quick-start)
|
|
24
|
+
- [The Memory Pipeline](#the-memory-pipeline)
|
|
25
|
+
- [Store](#1-store)
|
|
26
|
+
- [Retrieve](#2-retrieve)
|
|
27
|
+
- [Compile Context](#3-compile-context)
|
|
28
|
+
- [Summarize](#4-summarize)
|
|
29
|
+
- [Prune](#5-prune)
|
|
30
|
+
- [Real-World Use Cases](#real-world-use-cases)
|
|
31
|
+
- [Support Agent](#support-agent)
|
|
32
|
+
- [RAG Pipeline](#rag-pipeline)
|
|
33
|
+
- [Multi-User Chatbot](#multi-user-chatbot)
|
|
34
|
+
- [Memory Type Reference](#memory-type-reference)
|
|
35
|
+
- [Retrieval Strategies](#retrieval-strategies)
|
|
36
|
+
- [Embeddings](#embeddings)
|
|
37
|
+
- [Adapters](#adapters)
|
|
38
|
+
- [LLM Adapters](#llm-adapters)
|
|
39
|
+
- [Embedding Adapters](#embedding-adapters)
|
|
40
|
+
- [Storage Adapters](#storage-adapters)
|
|
41
|
+
- [Full API Reference](#full-api-reference)
|
|
42
|
+
- [MemStack Client](#memstack-client)
|
|
43
|
+
- [Memory Subsystem](#memory-subsystem)
|
|
44
|
+
- [Export / Import](#export-import)
|
|
45
|
+
- [Health & Close](#health-close)
|
|
46
|
+
- [Configuration](#configuration)
|
|
47
|
+
- [Advanced Usage](#advanced-usage)
|
|
48
|
+
- [Custom Storage](#custom-storage)
|
|
49
|
+
- [Custom LLM / Embedding](#custom-llm-embedding)
|
|
50
|
+
- [Event Hooks](#event-hooks)
|
|
51
|
+
- [Development](#development)
|
|
52
|
+
- [Setup & Tests](#setup-tests)
|
|
53
|
+
- [Debugging](#debugging)
|
|
54
|
+
- [Publishing to npm](#publishing-to-npm)
|
|
55
|
+
- [Contributing](#contributing)
|
|
56
|
+
- [License](#license)
|
|
57
|
+
|
|
58
|
+
---
|
|
59
|
+
|
|
60
|
+
## Why MemStack
|
|
61
|
+
|
|
62
|
+
**LLMs have context windows, not memory.** The difference matters.
|
|
63
|
+
|
|
64
|
+
| Approach | Problem |
|
|
65
|
+
|----------|---------|
|
|
66
|
+
| **Stuff everything in context** | Cost is O(n²). 100 conversations = thousands of tokens = dollars per call. Quality degrades from "lost in the middle" effect. |
|
|
67
|
+
| **Use a vector DB directly** | You get similarity search. You don't get summarization, pruning, recency weighting, deduplication, or token budget management. You're building the pipeline yourself. |
|
|
68
|
+
| **Use Mem0** | Proprietary, cloud-only with their hosted API. You don't control where your data lives. |
|
|
69
|
+
| **Use MemStack** | Full pipeline. Pluggable everything. Your data, your infrastructure. Open source. |
|
|
70
|
+
|
|
71
|
+
**What MemStack handles that raw vector DBs don't:**
|
|
72
|
+
|
|
73
|
+
- **Summarization** — compress 100 old interactions into one paragraph, keep meaning, save tokens
|
|
74
|
+
- **Recency weighting** — recent memories matter more; MemStack sorts them higher
|
|
75
|
+
- **Importance scoring** — not all memories are equal; high-importance ones survive pruning
|
|
76
|
+
- **Deduplication** — identical or near-identical memories are collapsed in context assembly
|
|
77
|
+
- **Token budget** — `compileContext()` tells you how many tokens you're spending before the LLM call
|
|
78
|
+
- **Memory-type routing** — interactions, summaries, observations treated differently at retrieval time
|
|
79
|
+
- **Auto-pruning** — old, low-importance memories clean themselves up
|
|
80
|
+
|
|
81
|
+
---
|
|
82
|
+
|
|
83
|
+
## Quick Start
|
|
84
|
+
|
|
85
|
+
```typescript
|
|
86
|
+
import { MemStack, OpenAILLMAdapter, OpenAIEmbeddingAdapter } from "@memstack/core";
|
|
87
|
+
|
|
88
|
+
const memstack = new MemStack({
|
|
89
|
+
llm: new OpenAILLMAdapter({ apiKey: process.env.OPENAI_API_KEY! }),
|
|
90
|
+
embedding: new OpenAIEmbeddingAdapter({ apiKey: process.env.OPENAI_API_KEY! }),
|
|
91
|
+
});
|
|
92
|
+
|
|
93
|
+
// 1. Store what happened
|
|
94
|
+
await memstack.memory.store({
|
|
95
|
+
actorId: "support-bot-42",
|
|
96
|
+
content: "User reports login failing with error 503 on Chrome 125.",
|
|
97
|
+
tags: ["login", "bug", "chrome"],
|
|
98
|
+
importance: 0.8,
|
|
99
|
+
});
|
|
100
|
+
|
|
101
|
+
// 2. Later, retrieve relevant context
|
|
102
|
+
const memories = await memstack.memory.retrieve({
|
|
103
|
+
actorId: "support-bot-42",
|
|
104
|
+
query: "login error",
|
|
105
|
+
strategy: "hybrid",
|
|
106
|
+
});
|
|
107
|
+
|
|
108
|
+
// 3. Inject into your LLM call
|
|
109
|
+
const ctx = await memstack.memory.compileContext({
|
|
110
|
+
actorId: "support-bot-42",
|
|
111
|
+
maxTokens: 2000,
|
|
112
|
+
});
|
|
113
|
+
|
|
114
|
+
const llmResponse = await llm.complete({
|
|
115
|
+
system: `You are a support bot. Here is what you remember:\n${ctx.systemPrompt}`,
|
|
116
|
+
user: "The user is back and still can't log in. What do you do?",
|
|
117
|
+
});
|
|
118
|
+
|
|
119
|
+
// 4. Every 100 interactions, summarization kicks in automatically.
|
|
120
|
+
// Old interactions are compressed into a paragraph. Token costs stay flat.
|
|
121
|
+
```
|
|
122
|
+
|
|
123
|
+
---
|
|
124
|
+
|
|
125
|
+
## The Memory Pipeline
|
|
126
|
+
|
|
127
|
+
MemStack's core is a five-stage pipeline. Each stage can be used independently.
|
|
128
|
+
|
|
129
|
+
### 1. Store
|
|
130
|
+
|
|
131
|
+
Every agent interaction becomes a `Memory` with metadata that controls how it's retrieved, summarized, and pruned later.
|
|
132
|
+
|
|
133
|
+
```typescript
|
|
134
|
+
interface Memory {
|
|
135
|
+
id: string;
|
|
136
|
+
actorId: string; // Who this memory belongs to (user ID, agent ID, session ID)
|
|
137
|
+
memoryType: MemoryType; // "interaction" | "summary" | "observation"
|
|
138
|
+
content: string; // The actual text
|
|
139
|
+
importance: number; // 0-1 — higher = survives pruning, ranks higher in retrieval
|
|
140
|
+
emotionalValence: number; // -1 to 1 — for tone-aware retrieval
|
|
141
|
+
tags: string[]; // Filter by tag: "bug", "billing", "urgent", etc.
|
|
142
|
+
embedding?: number[]; // Computed automatically if embedding adapter is configured
|
|
143
|
+
metadata?: Record<string, unknown>; // Your custom fields
|
|
144
|
+
expiresAt?: Date; // Auto-pruned after this date
|
|
145
|
+
sourceId?: string; // Link back to the originating event
|
|
146
|
+
createdAt: Date;
|
|
147
|
+
}
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
```typescript
|
|
151
|
+
// Simple store
|
|
152
|
+
await ms.memory.store({
|
|
153
|
+
actorId: "agent-7",
|
|
154
|
+
content: "Customer asked about refund policy for Q2 purchases.",
|
|
155
|
+
tags: ["billing", "refund"],
|
|
156
|
+
});
|
|
157
|
+
|
|
158
|
+
// Batch store — embeddings are batched into one API call for efficiency
|
|
159
|
+
await ms.memory.storeBatch([
|
|
160
|
+
{ actorId: "agent-7", content: "First interaction" },
|
|
161
|
+
{ actorId: "agent-7", content: "Second interaction" },
|
|
162
|
+
{ actorId: "agent-7", content: "Third interaction" },
|
|
163
|
+
]);
|
|
164
|
+
```
|
|
165
|
+
|
|
166
|
+
### 2. Retrieve
|
|
167
|
+
|
|
168
|
+
Pull back what's relevant — by keyword, by meaning (semantic), by recency, or by importance.
|
|
169
|
+
|
|
170
|
+
```typescript
|
|
171
|
+
const memories = await ms.memory.retrieve({
|
|
172
|
+
actorId: "agent-7", // Scope to one actor
|
|
173
|
+
query: "refund policy", // What to search for
|
|
174
|
+
strategy: "hybrid", // How to rank: "recent" | "important" | "semantic" | "hybrid"
|
|
175
|
+
limit: 10, // Max results
|
|
176
|
+
memoryTypes: ["interaction"], // Only certain types
|
|
177
|
+
tags: ["billing"], // Only certain tags
|
|
178
|
+
});
|
|
179
|
+
```
|
|
180
|
+
|
|
181
|
+
**Strategy behavior:**
|
|
182
|
+
|
|
183
|
+
| Strategy | Sorts by | Requires embeddings | Best for |
|
|
184
|
+
|----------|----------|--------------------|----------|
|
|
185
|
+
| `recent` | Newest first | No | Knowing what just happened |
|
|
186
|
+
| `important` | Highest importance first | No | Filtering noise, keeping signal |
|
|
187
|
+
| `semantic` | Cosine similarity to query | Yes | "Find memories about X" |
|
|
188
|
+
| `hybrid` | Semantic + importance blend | Yes | Best of both worlds |
|
|
189
|
+
|
|
190
|
+
No embedding adapter? `semantic` and `hybrid` fall back to keyword matching + importance sort. No API costs, just less precise.
|
|
191
|
+
|
|
192
|
+
### 3. Compile Context
|
|
193
|
+
|
|
194
|
+
The killer feature. `compileContext()` takes the retrieval results and assembles an LLM-ready system prompt — deduplicated, sorted by recency and importance, with a token estimate so you know the cost before calling the LLM.
|
|
195
|
+
|
|
196
|
+
```typescript
|
|
197
|
+
const ctx = await ms.memory.compileContext({
|
|
198
|
+
actorId: "agent-7",
|
|
199
|
+
maxTokens: 2000, // Budget — assembler stops when it hits this
|
|
200
|
+
memoryTypes: ["interaction", "summary"],
|
|
201
|
+
});
|
|
202
|
+
|
|
203
|
+
// ctx.systemPrompt:
|
|
204
|
+
// ## Important Memories
|
|
205
|
+
// - The customer has been attempting login for 3 days. (importance: 0.85)
|
|
206
|
+
// - Refund was processed for order #4521 on Jan 12. (importance: 0.72)
|
|
207
|
+
//
|
|
208
|
+
// ## Recent Interactions
|
|
209
|
+
// - Customer asked about refund policy for Q2 purchases.
|
|
210
|
+
// - Customer reported login error 503 on Chrome 125.
|
|
211
|
+
|
|
212
|
+
console.log(ctx.tokenEstimate); // ~280
|
|
213
|
+
|
|
214
|
+
// Inject into your LLM call
|
|
215
|
+
const response = await llm.complete({
|
|
216
|
+
system: ctx.systemPrompt,
|
|
217
|
+
user: userMessage,
|
|
218
|
+
});
|
|
219
|
+
```
|
|
220
|
+
|
|
221
|
+
`compileContext()` is the difference between "we have a vector DB" and "we have agent memory." It handles deduplication, token budgeting, and the recent-vs-important split that makes context useful.
|
|
222
|
+
|
|
223
|
+
### 4. Summarize
|
|
224
|
+
|
|
225
|
+
When an actor has hundreds of interactions, retrieval gets expensive and context gets bloated. Summarization compresses old interactions into a single paragraph using the configured LLM.
|
|
226
|
+
|
|
227
|
+
```typescript
|
|
228
|
+
const { summary, deletedCount } = await ms.memory.summarize({
|
|
229
|
+
actorId: "agent-7",
|
|
230
|
+
olderThan: new Date(Date.now() - 7 * 86400000), // Older than 7 days
|
|
231
|
+
skipMostRecent: 10, // Never touch the 10 most recent
|
|
232
|
+
targetCount: 50, // Summarize at most 50 memories
|
|
233
|
+
memoryTypes: ["interaction"],
|
|
234
|
+
keepOriginals: false, // Delete originals after summary
|
|
235
|
+
});
|
|
236
|
+
|
|
237
|
+
// summary.content:
|
|
238
|
+
// "Over the past week, the customer reported recurring login failures (error 503)
|
|
239
|
+
// on Chrome 125. Multiple troubleshooting attempts including cache clearing and
|
|
240
|
+
// password reset were unsuccessful. A refund was processed for order #4521."
|
|
241
|
+
|
|
242
|
+
console.log(deletedCount); // 47 — 47 interactions compressed into 1 summary memory
|
|
243
|
+
```
|
|
244
|
+
|
|
245
|
+
**Auto-summarization:** Set `summarizationThreshold` in config (default: 100). Every 100th interaction for an actor triggers summarization automatically.
|
|
246
|
+
|
|
247
|
+
**Warning:** `keepOriginals: false` deletes the summarized memories. Set `keepOriginals: true` to preserve them alongside the summary.
|
|
248
|
+
|
|
249
|
+
**Custom summarization prompt:**
|
|
250
|
+
|
|
251
|
+
```typescript
|
|
252
|
+
const ms = new MemStack({
|
|
253
|
+
llm,
|
|
254
|
+
defaults: {
|
|
255
|
+
summarizationPrompt:
|
|
256
|
+
"You are an enterprise support memory compressor. Highlight: customer name,
|
|
257
|
+
product, severity, resolution status, and any open issues.",
|
|
258
|
+
},
|
|
259
|
+
});
|
|
260
|
+
```
|
|
261
|
+
|
|
262
|
+
### 5. Prune
|
|
263
|
+
|
|
264
|
+
Not all memories deserve to live forever. Pruning removes low-value memories to keep storage and retrieval fast.
|
|
265
|
+
|
|
266
|
+
```typescript
|
|
267
|
+
// Remove memories older than 30 days
|
|
268
|
+
await ms.memory.prune({ type: "byAge", maxAge: 30 * 86400000 });
|
|
269
|
+
|
|
270
|
+
// Keep only memories above importance 0.3
|
|
271
|
+
await ms.memory.prune({ type: "byImportance", minImportance: 0.3 });
|
|
272
|
+
|
|
273
|
+
// Keep at most 500 memories per actor
|
|
274
|
+
await ms.memory.prune({ type: "byCount", maxPerActor: 500 });
|
|
275
|
+
|
|
276
|
+
// Remove specific types
|
|
277
|
+
await ms.memory.prune({ type: "byType", memoryTypes: ["observation"] });
|
|
278
|
+
|
|
279
|
+
// Custom logic
|
|
280
|
+
await ms.memory.prune({
|
|
281
|
+
type: "custom",
|
|
282
|
+
shouldRemove: (memory) => memory.content.includes("[RESOLVED]"),
|
|
283
|
+
});
|
|
284
|
+
|
|
285
|
+
// Dry run first — see what would be removed
|
|
286
|
+
const { wouldPrune, count } = await ms.memory.dryRunPrune({
|
|
287
|
+
type: "byAge",
|
|
288
|
+
maxAge: 86400000,
|
|
289
|
+
});
|
|
290
|
+
console.log(`Would remove ${count} memories:`, wouldPrune);
|
|
291
|
+
```
|
|
292
|
+
|
|
293
|
+
Auto-prune on every `process()` call by setting `pruneStrategy` in config:
|
|
294
|
+
|
|
295
|
+
```typescript
|
|
296
|
+
const ms = new MemStack({
|
|
297
|
+
llm,
|
|
298
|
+
defaults: {
|
|
299
|
+
pruneStrategy: { type: "byImportance", minImportance: 0.05 },
|
|
300
|
+
},
|
|
301
|
+
});
|
|
302
|
+
```
|
|
303
|
+
|
|
304
|
+
---
|
|
305
|
+
|
|
306
|
+
## Real-World Use Cases
|
|
307
|
+
|
|
308
|
+
### Support Agent
|
|
309
|
+
|
|
310
|
+
```typescript
|
|
311
|
+
// Every customer message becomes a memory
|
|
312
|
+
async function handleMessage(customerId: string, message: string) {
|
|
313
|
+
await ms.memory.store({
|
|
314
|
+
actorId: `customer:${customerId}`,
|
|
315
|
+
content: message,
|
|
316
|
+
importance: detectUrgency(message), // NLP heuristic or LLM call
|
|
317
|
+
tags: classifyIntent(message), // "billing", "bug", "account", etc.
|
|
318
|
+
});
|
|
319
|
+
|
|
320
|
+
// Retrieve everything relevant to this customer's history
|
|
321
|
+
const ctx = await ms.memory.compileContext({
|
|
322
|
+
actorId: `customer:${customerId}`,
|
|
323
|
+
maxTokens: 1500,
|
|
324
|
+
});
|
|
325
|
+
|
|
326
|
+
const response = await llm.complete({
|
|
327
|
+
system: `You are a support agent. Customer history:\n${ctx.systemPrompt}`,
|
|
328
|
+
user: message,
|
|
329
|
+
});
|
|
330
|
+
|
|
331
|
+
return response.text;
|
|
332
|
+
}
|
|
333
|
+
|
|
334
|
+
// Every 100th interaction, old history auto-compresses.
|
|
335
|
+
// A customer with 10,000 messages still fits in a $0.02 LLM call.
|
|
336
|
+
```
|
|
337
|
+
|
|
338
|
+
### RAG Pipeline
|
|
339
|
+
|
|
340
|
+
```typescript
|
|
341
|
+
// Index documents as observation memories
|
|
342
|
+
for (const doc of documents) {
|
|
343
|
+
await ms.memory.store({
|
|
344
|
+
actorId: "knowledge-base",
|
|
345
|
+
content: doc.text,
|
|
346
|
+
memoryType: "observation",
|
|
347
|
+
metadata: { source: doc.url, section: doc.section },
|
|
348
|
+
});
|
|
349
|
+
}
|
|
350
|
+
|
|
351
|
+
// Query with semantic search
|
|
352
|
+
const relevantDocs = await ms.memory.retrieve({
|
|
353
|
+
actorId: "knowledge-base",
|
|
354
|
+
query: "How does authentication work?",
|
|
355
|
+
strategy: "semantic",
|
|
356
|
+
limit: 5,
|
|
357
|
+
});
|
|
358
|
+
|
|
359
|
+
const ctx = await ms.memory.compileContext({
|
|
360
|
+
actorId: "knowledge-base",
|
|
361
|
+
memoryTypes: ["observation"],
|
|
362
|
+
});
|
|
363
|
+
|
|
364
|
+
// Prompt the LLM with retrieved context
|
|
365
|
+
const answer = await llm.complete({
|
|
366
|
+
system: `Answer using only these documents:\n${ctx.systemPrompt}`,
|
|
367
|
+
user: "How does authentication work?",
|
|
368
|
+
});
|
|
369
|
+
```
|
|
370
|
+
|
|
371
|
+
### Multi-User Chatbot
|
|
372
|
+
|
|
373
|
+
```typescript
|
|
374
|
+
// Each user gets their own memory space
|
|
375
|
+
async function chat(userId: string, message: string) {
|
|
376
|
+
await ms.memory.store({
|
|
377
|
+
actorId: userId,
|
|
378
|
+
content: message,
|
|
379
|
+
});
|
|
380
|
+
|
|
381
|
+
const ctx = await ms.memory.compileContext({
|
|
382
|
+
actorId: userId,
|
|
383
|
+
maxTokens: 1000,
|
|
384
|
+
});
|
|
385
|
+
|
|
386
|
+
return llm.complete({
|
|
387
|
+
system: `You are a friendly assistant. Conversation history with this user:\n${ctx.systemPrompt}`,
|
|
388
|
+
user: message,
|
|
389
|
+
});
|
|
390
|
+
}
|
|
391
|
+
|
|
392
|
+
// Get stats
|
|
393
|
+
const total = await ms.memory.count();
|
|
394
|
+
const userCount = await ms.memory.count({ actorId: "user-42" });
|
|
395
|
+
```
|
|
396
|
+
|
|
397
|
+
---
|
|
398
|
+
|
|
399
|
+
## Memory Type Reference
|
|
400
|
+
|
|
401
|
+
| Type | Purpose | Example |
|
|
402
|
+
|------|---------|---------|
|
|
403
|
+
| `interaction` | Default. Direct exchanges between agent and user/other agent. | "User asked about billing." |
|
|
404
|
+
| `summary` | Compressed collection of old interactions. Created by `summarize()`. | "Over 3 weeks, user reported 5 login failures..." |
|
|
405
|
+
| `observation` | Passive knowledge — facts, documents, things the agent knows but didn't interact with. | "Company refund policy is 30 days from purchase." |
|
|
406
|
+
| `gossip` | Information about third parties. For multi-agent systems. | "Agent-B told me the user is a power user." |
|
|
407
|
+
|
|
408
|
+
Types control retrieval behavior — `compileContext()` treats `interaction` and `summary` differently from `observation`. Use types to separate "what happened" from "what I know."
|
|
409
|
+
|
|
410
|
+
---
|
|
411
|
+
|
|
412
|
+
## Retrieval Strategies
|
|
413
|
+
|
|
414
|
+
Four strategies, each with a purpose:
|
|
415
|
+
|
|
416
|
+
```typescript
|
|
417
|
+
// "What just happened?" — most recent first
|
|
418
|
+
await ms.memory.retrieve({ actorId: "x", strategy: "recent", limit: 3 });
|
|
419
|
+
|
|
420
|
+
// "What matters most?" — highest importance, ignoring age
|
|
421
|
+
await ms.memory.retrieve({ actorId: "x", strategy: "important" });
|
|
422
|
+
|
|
423
|
+
// "What relates to this query?" — cosine similarity search (needs embeddings)
|
|
424
|
+
await ms.memory.retrieve({ actorId: "x", query: "login bug", strategy: "semantic" });
|
|
425
|
+
|
|
426
|
+
// "Balance relevance and importance" — semantic + importance blend
|
|
427
|
+
await ms.memory.retrieve({ actorId: "x", query: "login bug", strategy: "hybrid" });
|
|
428
|
+
```
|
|
429
|
+
|
|
430
|
+
**Choosing a strategy:**
|
|
431
|
+
- Use `recent` for chatbots, ongoing conversations, anything time-sensitive
|
|
432
|
+
- Use `important` for long-running agents where signal-to-noise matters
|
|
433
|
+
- Use `semantic` for RAG, document search, knowledge base queries
|
|
434
|
+
- Use `hybrid` for most agent memory — it balances meaning with significance
|
|
435
|
+
|
|
436
|
+
---
|
|
437
|
+
|
|
438
|
+
## Embeddings
|
|
439
|
+
|
|
440
|
+
Embeddings power semantic search. They're optional — without them, retrieval uses keyword matching.
|
|
441
|
+
|
|
442
|
+
**With embeddings** (configure an `EmbeddingProvider`): each `store()` computes a vector. `retrieve()` with `"semantic"` or `"hybrid"` uses cosine similarity ranking.
|
|
443
|
+
|
|
444
|
+
**Without embeddings**: everything still works — retrieval falls back to importance + recency + keyword filters. No API costs, no setup.
|
|
445
|
+
|
|
446
|
+
**Batch embedding:** `storeBatch()` sends all texts in one embedding API call, reducing cost and latency.
|
|
447
|
+
|
|
448
|
+
```typescript
|
|
449
|
+
// Disable auto-embedding if you only need keyword search
|
|
450
|
+
const ms = new MemStack({
|
|
451
|
+
llm,
|
|
452
|
+
embedding: new OpenAIEmbeddingAdapter({ apiKey }),
|
|
453
|
+
defaults: { embedOnStore: false },
|
|
454
|
+
});
|
|
455
|
+
```
|
|
456
|
+
|
|
457
|
+
---
|
|
458
|
+
|
|
459
|
+
## Adapters
|
|
460
|
+
|
|
461
|
+
MemStack is provider-agnostic. Every boundary is an interface — bring your own LLM, embedding model, and storage backend.
|
|
462
|
+
|
|
463
|
+
### LLM Adapters
|
|
464
|
+
|
|
465
|
+
Used by `summarize()` and `compileContext()`. Ships with OpenAI and Anthropic built-in.
|
|
466
|
+
|
|
467
|
+
```typescript
|
|
468
|
+
// OpenAI
|
|
469
|
+
import { OpenAILLMAdapter } from "@memstack/core";
|
|
470
|
+
const llm = new OpenAILLMAdapter({
|
|
471
|
+
apiKey: process.env.OPENAI_API_KEY!,
|
|
472
|
+
defaultModel: "gpt-4o-mini", // default
|
|
473
|
+
baseURL: "https://api.openai.com/v1", // for proxies like LiteLLM
|
|
474
|
+
});
|
|
475
|
+
|
|
476
|
+
// Anthropic
|
|
477
|
+
import { AnthropicLLMAdapter } from "@memstack/core";
|
|
478
|
+
const llm = new AnthropicLLMAdapter({
|
|
479
|
+
apiKey: process.env.ANTHROPIC_API_KEY!,
|
|
480
|
+
defaultModel: "claude-sonnet-4-5-20250929",
|
|
481
|
+
});
|
|
482
|
+
|
|
483
|
+
// Ollama (custom — implement LLMProvider)
|
|
484
|
+
import type { LLMProvider } from "@memstack/core";
|
|
485
|
+
class OllamaAdapter implements LLMProvider {
|
|
486
|
+
constructor(private baseURL = "http://localhost:11434") {}
|
|
487
|
+
async complete(req: { system: string; user: string; model?: string }) {
|
|
488
|
+
const res = await fetch(`${this.baseURL}/api/generate`, {
|
|
489
|
+
method: "POST",
|
|
490
|
+
body: JSON.stringify({ model: req.model ?? "llama3.2", prompt: `${req.system}\n\n${req.user}`, stream: false }),
|
|
491
|
+
});
|
|
492
|
+
const data = await res.json() as { response: string };
|
|
493
|
+
return { text: data.response, tokens: { prompt: 0, completion: 0, total: 0 } };
|
|
494
|
+
}
|
|
495
|
+
}
|
|
496
|
+
```
|
|
497
|
+
|
|
498
|
+
### Embedding Adapters
|
|
499
|
+
|
|
500
|
+
Used by semantic retrieval. Ships with OpenAI built-in.
|
|
501
|
+
|
|
502
|
+
```typescript
|
|
503
|
+
import { OpenAIEmbeddingAdapter } from "@memstack/core";
|
|
504
|
+
const embedding = new OpenAIEmbeddingAdapter({
|
|
505
|
+
apiKey: process.env.OPENAI_API_KEY!,
|
|
506
|
+
model: "text-embedding-3-small", // 1536 dimensions (default)
|
|
507
|
+
// model: "text-embedding-3-large", // 3072 dimensions
|
|
508
|
+
});
|
|
509
|
+
```
|
|
510
|
+
|
|
511
|
+
### Storage Adapters
|
|
512
|
+
|
|
513
|
+
Ships with `InMemoryStorage` (zero setup, data lost on restart). For production, implement `StorageProvider` for your database.
|
|
514
|
+
|
|
515
|
+
```typescript
|
|
516
|
+
import { InMemoryStorage } from "@memstack/core";
|
|
517
|
+
const storage = new InMemoryStorage();
|
|
518
|
+
```
|
|
519
|
+
|
|
520
|
+
**Custom storage** — implement `StorageProvider`:
|
|
521
|
+
|
|
522
|
+
```typescript
|
|
523
|
+
import type { StorageProvider, MemoryStoreInput } from "@memstack/core";
|
|
524
|
+
|
|
525
|
+
class PostgresStorage implements StorageProvider {
|
|
526
|
+
async store(input: MemoryStoreInput): Promise<Memory> { /* INSERT */ }
|
|
527
|
+
async get(id: string): Promise<Memory | null> { /* SELECT */ }
|
|
528
|
+
async retrieve(query: MemoryRetrieveQuery, embedding?: number[]): Promise<Memory[]> { /* SELECT + filters */ }
|
|
529
|
+
async count(filter?: MemoryCountFilter): Promise<number> { /* SELECT COUNT */ }
|
|
530
|
+
async delete(id: string): Promise<void> { /* DELETE */ }
|
|
531
|
+
async deleteMany(ids: string[]): Promise<number> { /* DELETE batch */ }
|
|
532
|
+
async storeBatch(inputs: MemoryStoreInput[]): Promise<Memory[]> { /* INSERT batch */ }
|
|
533
|
+
async initialize(): Promise<void> { /* CREATE TABLE */ }
|
|
534
|
+
async close(): Promise<void> { /* close pool */ }
|
|
535
|
+
}
|
|
536
|
+
```
|
|
537
|
+
|
|
538
|
+
See `src/adapters/storage/memory.ts` for a complete reference implementation.
|
|
539
|
+
|
|
540
|
+
---
|
|
541
|
+
|
|
542
|
+
## Full API Reference
|
|
543
|
+
|
|
544
|
+
### MemStack Client
|
|
545
|
+
|
|
546
|
+
```typescript
|
|
547
|
+
import { MemStack } from "@memstack/core";
|
|
548
|
+
|
|
549
|
+
const ms = new MemStack({
|
|
550
|
+
llm: LLMProvider, // Required — for summarization
|
|
551
|
+
embedding?: EmbeddingProvider, // Optional — for semantic search
|
|
552
|
+
storage?: StorageProvider, // Optional — defaults to InMemoryStorage
|
|
553
|
+
defaults?: {
|
|
554
|
+
summarizationThreshold?: number, // Auto-summarize every N interactions. Default: 100
|
|
555
|
+
embedOnStore?: boolean, // Auto-embed on store(). Default: true
|
|
556
|
+
pruneStrategy?: PruneStrategy, // Auto-prune on every store(). Default: disabled
|
|
557
|
+
},
|
|
558
|
+
hooks?: {
|
|
559
|
+
onMemoryStored?: (memory: Memory) => void;
|
|
560
|
+
onMemoryPruned?: (ids: string[]) => void;
|
|
561
|
+
onSummaryCreated?: (summary: Memory, deletedCount: number) => void;
|
|
562
|
+
},
|
|
563
|
+
});
|
|
564
|
+
```
|
|
565
|
+
|
|
566
|
+
### Memory Subsystem
|
|
567
|
+
|
|
568
|
+
All methods accessible via `ms.memory.*`:
|
|
569
|
+
|
|
570
|
+
```typescript
|
|
571
|
+
// Store
|
|
572
|
+
ms.memory.store(input: MemoryStoreInput): Promise<Memory>
|
|
573
|
+
ms.memory.storeBatch(inputs: MemoryStoreInput[]): Promise<Memory[]>
|
|
574
|
+
|
|
575
|
+
// Retrieve
|
|
576
|
+
ms.memory.retrieve(query: MemoryRetrieveQuery): Promise<Memory[]>
|
|
577
|
+
ms.memory.get(id: string): Promise<Memory | null>
|
|
578
|
+
|
|
579
|
+
// Context assembly
|
|
580
|
+
ms.memory.compileContext(options: ContextOptions): Promise<CompiledContext>
|
|
581
|
+
|
|
582
|
+
// Lifecycle
|
|
583
|
+
ms.memory.summarize(options: SummarizeOptions): Promise<{ summary: Memory; deletedCount: number }>
|
|
584
|
+
ms.memory.prune(strategy: PruneStrategy): Promise<{ pruned: string[]; count: number }>
|
|
585
|
+
ms.memory.dryRunPrune(strategy: PruneStrategy): Promise<{ wouldPrune: string[]; count: number }>
|
|
586
|
+
|
|
587
|
+
// Management
|
|
588
|
+
ms.memory.count(filter?: MemoryCountFilter): Promise<number>
|
|
589
|
+
ms.memory.delete(id: string): Promise<void>
|
|
590
|
+
ms.memory.deleteMany(ids: string[]): Promise<number>
|
|
591
|
+
ms.memory.touch(id: string): Promise<void> // bump recency without changing content
|
|
592
|
+
```
|
|
593
|
+
|
|
594
|
+
### Export / Import
|
|
595
|
+
|
|
596
|
+
Snapshot and restore full state for persistence, backups, or migration:
|
|
597
|
+
|
|
598
|
+
```typescript
|
|
599
|
+
// Save
|
|
600
|
+
const snapshot = await ms.export();
|
|
601
|
+
fs.writeFileSync("state.json", JSON.stringify(snapshot, null, 2));
|
|
602
|
+
|
|
603
|
+
// Restore
|
|
604
|
+
const data = JSON.parse(fs.readFileSync("state.json", "utf-8"));
|
|
605
|
+
await ms2.import(data);
|
|
606
|
+
```
|
|
607
|
+
|
|
608
|
+
### Health & Close
|
|
609
|
+
|
|
610
|
+
```typescript
|
|
611
|
+
const status = await ms.health();
|
|
612
|
+
// { storage: true, llm: true, embedding: true }
|
|
613
|
+
|
|
614
|
+
await ms.close(); // graceful shutdown
|
|
615
|
+
```
|
|
616
|
+
|
|
617
|
+
---
|
|
618
|
+
|
|
619
|
+
## Configuration
|
|
620
|
+
|
|
621
|
+
```typescript
|
|
622
|
+
const ms = new MemStack({
|
|
623
|
+
llm: new OpenAILLMAdapter({ apiKey: "..." }),
|
|
624
|
+
|
|
625
|
+
// Defaults control auto-behavior
|
|
626
|
+
defaults: {
|
|
627
|
+
summarizationThreshold: 50, // Summarize every 50 interactions (default: 100)
|
|
628
|
+
embedOnStore: false, // Don't auto-embed — saves API costs
|
|
629
|
+
pruneStrategy: { // Auto-clean on every store()
|
|
630
|
+
type: "byAge",
|
|
631
|
+
maxAge: 90 * 86400000, // 90 days
|
|
632
|
+
},
|
|
633
|
+
},
|
|
634
|
+
|
|
635
|
+
// Hooks for observability
|
|
636
|
+
hooks: {
|
|
637
|
+
onMemoryStored: (m) => logger.debug("memory:stored", { id: m.id, actor: m.actorId }),
|
|
638
|
+
onMemoryPruned: (ids) => logger.info("memory:pruned", { count: ids.length }),
|
|
639
|
+
onSummaryCreated: (summary, n) => logger.info("memory:summarized", { count: n }),
|
|
640
|
+
},
|
|
641
|
+
});
|
|
642
|
+
```
|
|
643
|
+
|
|
644
|
+
---
|
|
645
|
+
|
|
646
|
+
## Advanced Usage
|
|
647
|
+
|
|
648
|
+
### Custom Storage
|
|
649
|
+
|
|
650
|
+
Implement `StorageProvider` for any database. The interface is 9 methods. See the reference section above for the full contract.
|
|
651
|
+
|
|
652
|
+
### Custom LLM / Embedding
|
|
653
|
+
|
|
654
|
+
Implement `LLMProvider` or `EmbeddingProvider` for any service:
|
|
655
|
+
|
|
656
|
+
```typescript
|
|
657
|
+
import type { LLMProvider } from "@memstack/core";
|
|
658
|
+
|
|
659
|
+
class TogetherAIAdapter implements LLMProvider {
|
|
660
|
+
async complete(req: { system: string; user: string; model?: string }) {
|
|
661
|
+
const res = await fetch("https://api.together.xyz/v1/chat/completions", {
|
|
662
|
+
headers: { Authorization: `Bearer ${this.apiKey}`, "Content-Type": "application/json" },
|
|
663
|
+
body: JSON.stringify({ model: req.model, messages: [{ role: "system", content: req.system }, { role: "user", content: req.user }] }),
|
|
664
|
+
});
|
|
665
|
+
const data = await res.json() as any;
|
|
666
|
+
return { text: data.choices[0].message.content, tokens: { prompt: data.usage.prompt_tokens, completion: data.usage.completion_tokens, total: data.usage.total_tokens } };
|
|
667
|
+
}
|
|
668
|
+
}
|
|
669
|
+
```
|
|
670
|
+
|
|
671
|
+
### Event Hooks
|
|
672
|
+
|
|
673
|
+
Monitor memory operations without modifying code:
|
|
674
|
+
|
|
675
|
+
```typescript
|
|
676
|
+
const ms = new MemStack({
|
|
677
|
+
llm,
|
|
678
|
+
hooks: {
|
|
679
|
+
onMemoryStored: (m) => metrics.increment("memory.stored"),
|
|
680
|
+
onSummaryCreated: (_, n) => metrics.gauge("memory.summarized_count", n),
|
|
681
|
+
onMemoryPruned: (ids) => metrics.increment("memory.pruned", ids.length),
|
|
682
|
+
},
|
|
683
|
+
});
|
|
684
|
+
```
|
|
685
|
+
|
|
686
|
+
---
|
|
687
|
+
|
|
688
|
+
---
|
|
689
|
+
|
|
690
|
+
## Development
|
|
691
|
+
|
|
692
|
+
### Setup & Tests
|
|
693
|
+
|
|
694
|
+
```bash
|
|
695
|
+
git clone https://github.com/isiomaC/memstack.git
|
|
696
|
+
cd memstack
|
|
697
|
+
pnpm install
|
|
698
|
+
|
|
699
|
+
pnpm test # 56 tests, no external services needed
|
|
700
|
+
pnpm test:watch # Watch mode
|
|
701
|
+
pnpm build # CJS + ESM + type declarations
|
|
702
|
+
pnpm check # TypeScript type-check only
|
|
703
|
+
```
|
|
704
|
+
|
|
705
|
+
### Debugging
|
|
706
|
+
|
|
707
|
+
Use hooks for observability — MemStack has no built-in logging:
|
|
708
|
+
|
|
709
|
+
```typescript
|
|
710
|
+
const ms = new MemStack({
|
|
711
|
+
llm,
|
|
712
|
+
hooks: {
|
|
713
|
+
onMemoryStored: (m) => console.debug("[memstack] stored:", m.id, m.content.slice(0, 80)),
|
|
714
|
+
onMemoryPruned: (ids) => console.debug("[memstack] pruned:", ids.length),
|
|
715
|
+
},
|
|
716
|
+
});
|
|
717
|
+
```
|
|
718
|
+
|
|
719
|
+
**Common issues:**
|
|
720
|
+
|
|
721
|
+
| Symptom | Cause | Fix |
|
|
722
|
+
|---------|-------|-----|
|
|
723
|
+
| `CONFIG_ERROR: LLM provider is required` | No LLM adapter | Pass any `LLMProvider` to config |
|
|
724
|
+
| Empty retrieval results | Wrong `actorId` or no memories stored | Check `await ms.memory.count({ actorId })` |
|
|
725
|
+
| Semantic search not working | No embedding adapter or `embedOnStore: false` | Add embedding adapter or use `strategy: "recent"` |
|
|
726
|
+
| High memory usage in production | Using InMemoryStorage | Implement `StorageProvider` for Postgres/Redis/etc |
|
|
727
|
+
| Poor summarization quality | Default prompt doesn't match your domain | Use `summarizationPrompt` in `defaults` config |
|
|
728
|
+
|
|
729
|
+
**Inspecting state at runtime:**
|
|
730
|
+
|
|
731
|
+
```typescript
|
|
732
|
+
// How much data do we have?
|
|
733
|
+
const total = await ms.memory.count();
|
|
734
|
+
const perActor = await ms.memory.count({ actorId: "user-42" });
|
|
735
|
+
|
|
736
|
+
// What does one actor's memory look like?
|
|
737
|
+
const snapshot = await ms.export();
|
|
738
|
+
const actorMemories = snapshot.memories.filter(m => m.actorId === "user-42");
|
|
739
|
+
console.log(`User-42: ${actorMemories.length} memories`);
|
|
740
|
+
actorMemories.forEach(m => console.log(` [${m.memoryType}] ${m.content.slice(0, 60)} (imp: ${m.importance})`));
|
|
741
|
+
```
|
|
742
|
+
|
|
743
|
+
---
|
|
744
|
+
|
|
745
|
+
## Publishing to npm
|
|
746
|
+
|
|
747
|
+
```bash
|
|
748
|
+
# Bump version, then:
|
|
749
|
+
pnpm build && pnpm check && pnpm test
|
|
750
|
+
npm login
|
|
751
|
+
npm publish --access public
|
|
752
|
+
```
|
|
753
|
+
|
|
754
|
+
The `@memstack` scope requires `--access public` on first publish.
|
|
755
|
+
|
|
756
|
+
---
|
|
757
|
+
|
|
758
|
+
## Contributing
|
|
759
|
+
|
|
760
|
+
Most needed contributions:
|
|
761
|
+
|
|
762
|
+
- **Storage adapters**: Postgres, Redis, SQLite, filesystem
|
|
763
|
+
- **LLM adapters**: Ollama, Groq, Together AI, Gemini
|
|
764
|
+
- **Embedding adapters**: Cohere, Voyage AI, local transformers.js
|
|
765
|
+
- **Tests**: Edge cases, concurrent access, large-scale benchmarks
|
|
766
|
+
- **Docs**: Architecture diagrams, tutorials
|
|
767
|
+
|
|
768
|
+
Open an issue or PR at [github.com/isiomaC/memstack](https://github.com/isiomaC/memstack).
|
|
769
|
+
|
|
770
|
+
---
|
|
771
|
+
|
|
772
|
+
## License
|
|
773
|
+
|
|
774
|
+
MIT © [MemStack](https://github.com/isiomaC/memstack)
|