@memstack/core 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md ADDED
@@ -0,0 +1,774 @@
1
+ # MemStack
2
+
3
+ > The open-source memory layer for AI agents — store, retrieve, summarize, and prune.
4
+
5
+ [![npm version](https://img.shields.io/npm/v/@memstack/core)](https://www.npmjs.com/package/@memstack/core)
6
+ [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
7
+
8
+ ```bash
9
+ npm install @memstack/core
10
+ ```
11
+
12
+ **The problem:** AI agents forget. Every interaction starts from zero. You either stuff everything into the context window (expensive, slow, degrades output quality) or the agent has no memory of past conversations.
13
+
14
+ **What MemStack does:** A persistent memory pipeline that lives between your agent and the LLM. It stores every interaction, retrieves only what's relevant, summarizes old memories to save tokens, and prunes stale ones automatically. One method call, no infrastructure required.
15
+
16
+ Think of it as the open-source alternative to [Mem0](https://mem0.ai/) — pluggable storage, bring your own LLM, zero vendor lock-in.
17
+
18
+ ---
19
+
20
+ ## Table of Contents
21
+
22
+ - [Why MemStack](#why-memstack)
23
+ - [Quick Start](#quick-start)
24
+ - [The Memory Pipeline](#the-memory-pipeline)
25
+ - [Store](#1-store)
26
+ - [Retrieve](#2-retrieve)
27
+ - [Compile Context](#3-compile-context)
28
+ - [Summarize](#4-summarize)
29
+ - [Prune](#5-prune)
30
+ - [Real-World Use Cases](#real-world-use-cases)
31
+ - [Support Agent](#support-agent)
32
+ - [RAG Pipeline](#rag-pipeline)
33
+ - [Multi-User Chatbot](#multi-user-chatbot)
34
+ - [Memory Type Reference](#memory-type-reference)
35
+ - [Retrieval Strategies](#retrieval-strategies)
36
+ - [Embeddings](#embeddings)
37
+ - [Adapters](#adapters)
38
+ - [LLM Adapters](#llm-adapters)
39
+ - [Embedding Adapters](#embedding-adapters)
40
+ - [Storage Adapters](#storage-adapters)
41
+ - [Full API Reference](#full-api-reference)
42
+ - [MemStack Client](#memstack-client)
43
+ - [Memory Subsystem](#memory-subsystem)
44
+ - [Export / Import](#export-import)
45
+ - [Health & Close](#health-close)
46
+ - [Configuration](#configuration)
47
+ - [Advanced Usage](#advanced-usage)
48
+ - [Custom Storage](#custom-storage)
49
+ - [Custom LLM / Embedding](#custom-llm-embedding)
50
+ - [Event Hooks](#event-hooks)
51
+ - [Development](#development)
52
+ - [Setup & Tests](#setup-tests)
53
+ - [Debugging](#debugging)
54
+ - [Publishing to npm](#publishing-to-npm)
55
+ - [Contributing](#contributing)
56
+ - [License](#license)
57
+
58
+ ---
59
+
60
+ ## Why MemStack
61
+
62
+ **LLMs have context windows, not memory.** The difference matters.
63
+
64
+ | Approach | Problem |
65
+ |----------|---------|
66
+ | **Stuff everything in context** | Cost is O(n²). 100 conversations = thousands of tokens = dollars per call. Quality degrades from "lost in the middle" effect. |
67
+ | **Use a vector DB directly** | You get similarity search. You don't get summarization, pruning, recency weighting, deduplication, or token budget management. You're building the pipeline yourself. |
68
+ | **Use Mem0** | Proprietary, cloud-only with their hosted API. You don't control where your data lives. |
69
+ | **Use MemStack** | Full pipeline. Pluggable everything. Your data, your infrastructure. Open source. |
70
+
71
+ **What MemStack handles that raw vector DBs don't:**
72
+
73
+ - **Summarization** — compress 100 old interactions into one paragraph, keep meaning, save tokens
74
+ - **Recency weighting** — recent memories matter more; MemStack sorts them higher
75
+ - **Importance scoring** — not all memories are equal; high-importance ones survive pruning
76
+ - **Deduplication** — identical or near-identical memories are collapsed in context assembly
77
+ - **Token budget** — `compileContext()` tells you how many tokens you're spending before the LLM call
78
+ - **Memory-type routing** — interactions, summaries, observations treated differently at retrieval time
79
+ - **Auto-pruning** — old, low-importance memories clean themselves up
80
+
81
+ ---
82
+
83
+ ## Quick Start
84
+
85
+ ```typescript
86
+ import { MemStack, OpenAILLMAdapter, OpenAIEmbeddingAdapter } from "@memstack/core";
87
+
88
+ const memstack = new MemStack({
89
+ llm: new OpenAILLMAdapter({ apiKey: process.env.OPENAI_API_KEY! }),
90
+ embedding: new OpenAIEmbeddingAdapter({ apiKey: process.env.OPENAI_API_KEY! }),
91
+ });
92
+
93
+ // 1. Store what happened
94
+ await memstack.memory.store({
95
+ actorId: "support-bot-42",
96
+ content: "User reports login failing with error 503 on Chrome 125.",
97
+ tags: ["login", "bug", "chrome"],
98
+ importance: 0.8,
99
+ });
100
+
101
+ // 2. Later, retrieve relevant context
102
+ const memories = await memstack.memory.retrieve({
103
+ actorId: "support-bot-42",
104
+ query: "login error",
105
+ strategy: "hybrid",
106
+ });
107
+
108
+ // 3. Inject into your LLM call
109
+ const ctx = await memstack.memory.compileContext({
110
+ actorId: "support-bot-42",
111
+ maxTokens: 2000,
112
+ });
113
+
114
+ const llmResponse = await llm.complete({
115
+ system: `You are a support bot. Here is what you remember:\n${ctx.systemPrompt}`,
116
+ user: "The user is back and still can't log in. What do you do?",
117
+ });
118
+
119
+ // 4. Every 100 interactions, summarization kicks in automatically.
120
+ // Old interactions are compressed into a paragraph. Token costs stay flat.
121
+ ```
122
+
123
+ ---
124
+
125
+ ## The Memory Pipeline
126
+
127
+ MemStack's core is a five-stage pipeline. Each stage can be used independently.
128
+
129
+ ### 1. Store
130
+
131
+ Every agent interaction becomes a `Memory` with metadata that controls how it's retrieved, summarized, and pruned later.
132
+
133
+ ```typescript
134
+ interface Memory {
135
+ id: string;
136
+ actorId: string; // Who this memory belongs to (user ID, agent ID, session ID)
137
+ memoryType: MemoryType; // "interaction" | "summary" | "observation"
138
+ content: string; // The actual text
139
+ importance: number; // 0-1 — higher = survives pruning, ranks higher in retrieval
140
+ emotionalValence: number; // -1 to 1 — for tone-aware retrieval
141
+ tags: string[]; // Filter by tag: "bug", "billing", "urgent", etc.
142
+ embedding?: number[]; // Computed automatically if embedding adapter is configured
143
+ metadata?: Record<string, unknown>; // Your custom fields
144
+ expiresAt?: Date; // Auto-pruned after this date
145
+ sourceId?: string; // Link back to the originating event
146
+ createdAt: Date;
147
+ }
148
+ ```
149
+
150
+ ```typescript
151
+ // Simple store
152
+ await ms.memory.store({
153
+ actorId: "agent-7",
154
+ content: "Customer asked about refund policy for Q2 purchases.",
155
+ tags: ["billing", "refund"],
156
+ });
157
+
158
+ // Batch store — embeddings are batched into one API call for efficiency
159
+ await ms.memory.storeBatch([
160
+ { actorId: "agent-7", content: "First interaction" },
161
+ { actorId: "agent-7", content: "Second interaction" },
162
+ { actorId: "agent-7", content: "Third interaction" },
163
+ ]);
164
+ ```
165
+
166
+ ### 2. Retrieve
167
+
168
+ Pull back what's relevant — by keyword, by meaning (semantic), by recency, or by importance.
169
+
170
+ ```typescript
171
+ const memories = await ms.memory.retrieve({
172
+ actorId: "agent-7", // Scope to one actor
173
+ query: "refund policy", // What to search for
174
+ strategy: "hybrid", // How to rank: "recent" | "important" | "semantic" | "hybrid"
175
+ limit: 10, // Max results
176
+ memoryTypes: ["interaction"], // Only certain types
177
+ tags: ["billing"], // Only certain tags
178
+ });
179
+ ```
180
+
181
+ **Strategy behavior:**
182
+
183
+ | Strategy | Sorts by | Requires embeddings | Best for |
184
+ |----------|----------|--------------------|----------|
185
+ | `recent` | Newest first | No | Knowing what just happened |
186
+ | `important` | Highest importance first | No | Filtering noise, keeping signal |
187
+ | `semantic` | Cosine similarity to query | Yes | "Find memories about X" |
188
+ | `hybrid` | Semantic + importance blend | Yes | Best of both worlds |
189
+
190
+ No embedding adapter? `semantic` and `hybrid` fall back to keyword matching + importance sort. No API costs, just less precise.
191
+
192
+ ### 3. Compile Context
193
+
194
+ The killer feature. `compileContext()` takes the retrieval results and assembles an LLM-ready system prompt — deduplicated, sorted by recency and importance, with a token estimate so you know the cost before calling the LLM.
195
+
196
+ ```typescript
197
+ const ctx = await ms.memory.compileContext({
198
+ actorId: "agent-7",
199
+ maxTokens: 2000, // Budget — assembler stops when it hits this
200
+ memoryTypes: ["interaction", "summary"],
201
+ });
202
+
203
+ // ctx.systemPrompt:
204
+ // ## Important Memories
205
+ // - The customer has been attempting login for 3 days. (importance: 0.85)
206
+ // - Refund was processed for order #4521 on Jan 12. (importance: 0.72)
207
+ //
208
+ // ## Recent Interactions
209
+ // - Customer asked about refund policy for Q2 purchases.
210
+ // - Customer reported login error 503 on Chrome 125.
211
+
212
+ console.log(ctx.tokenEstimate); // ~280
213
+
214
+ // Inject into your LLM call
215
+ const response = await llm.complete({
216
+ system: ctx.systemPrompt,
217
+ user: userMessage,
218
+ });
219
+ ```
220
+
221
+ `compileContext()` is the difference between "we have a vector DB" and "we have agent memory." It handles deduplication, token budgeting, and the recent-vs-important split that makes context useful.
222
+
223
+ ### 4. Summarize
224
+
225
+ When an actor has hundreds of interactions, retrieval gets expensive and context gets bloated. Summarization compresses old interactions into a single paragraph using the configured LLM.
226
+
227
+ ```typescript
228
+ const { summary, deletedCount } = await ms.memory.summarize({
229
+ actorId: "agent-7",
230
+ olderThan: new Date(Date.now() - 7 * 86400000), // Older than 7 days
231
+ skipMostRecent: 10, // Never touch the 10 most recent
232
+ targetCount: 50, // Summarize at most 50 memories
233
+ memoryTypes: ["interaction"],
234
+ keepOriginals: false, // Delete originals after summary
235
+ });
236
+
237
+ // summary.content:
238
+ // "Over the past week, the customer reported recurring login failures (error 503)
239
+ // on Chrome 125. Multiple troubleshooting attempts including cache clearing and
240
+ // password reset were unsuccessful. A refund was processed for order #4521."
241
+
242
+ console.log(deletedCount); // 47 — 47 interactions compressed into 1 summary memory
243
+ ```
244
+
245
+ **Auto-summarization:** Set `summarizationThreshold` in config (default: 100). Every 100th interaction for an actor triggers summarization automatically.
246
+
247
+ **Warning:** `keepOriginals: false` deletes the summarized memories. Set `keepOriginals: true` to preserve them alongside the summary.
248
+
249
+ **Custom summarization prompt:**
250
+
251
+ ```typescript
252
+ const ms = new MemStack({
253
+ llm,
254
+ defaults: {
255
+ summarizationPrompt:
256
+ "You are an enterprise support memory compressor. Highlight: customer name,
257
+ product, severity, resolution status, and any open issues.",
258
+ },
259
+ });
260
+ ```
261
+
262
+ ### 5. Prune
263
+
264
+ Not all memories deserve to live forever. Pruning removes low-value memories to keep storage and retrieval fast.
265
+
266
+ ```typescript
267
+ // Remove memories older than 30 days
268
+ await ms.memory.prune({ type: "byAge", maxAge: 30 * 86400000 });
269
+
270
+ // Keep only memories above importance 0.3
271
+ await ms.memory.prune({ type: "byImportance", minImportance: 0.3 });
272
+
273
+ // Keep at most 500 memories per actor
274
+ await ms.memory.prune({ type: "byCount", maxPerActor: 500 });
275
+
276
+ // Remove specific types
277
+ await ms.memory.prune({ type: "byType", memoryTypes: ["observation"] });
278
+
279
+ // Custom logic
280
+ await ms.memory.prune({
281
+ type: "custom",
282
+ shouldRemove: (memory) => memory.content.includes("[RESOLVED]"),
283
+ });
284
+
285
+ // Dry run first — see what would be removed
286
+ const { wouldPrune, count } = await ms.memory.dryRunPrune({
287
+ type: "byAge",
288
+ maxAge: 86400000,
289
+ });
290
+ console.log(`Would remove ${count} memories:`, wouldPrune);
291
+ ```
292
+
293
+ Auto-prune on every `process()` call by setting `pruneStrategy` in config:
294
+
295
+ ```typescript
296
+ const ms = new MemStack({
297
+ llm,
298
+ defaults: {
299
+ pruneStrategy: { type: "byImportance", minImportance: 0.05 },
300
+ },
301
+ });
302
+ ```
303
+
304
+ ---
305
+
306
+ ## Real-World Use Cases
307
+
308
+ ### Support Agent
309
+
310
+ ```typescript
311
+ // Every customer message becomes a memory
312
+ async function handleMessage(customerId: string, message: string) {
313
+ await ms.memory.store({
314
+ actorId: `customer:${customerId}`,
315
+ content: message,
316
+ importance: detectUrgency(message), // NLP heuristic or LLM call
317
+ tags: classifyIntent(message), // "billing", "bug", "account", etc.
318
+ });
319
+
320
+ // Retrieve everything relevant to this customer's history
321
+ const ctx = await ms.memory.compileContext({
322
+ actorId: `customer:${customerId}`,
323
+ maxTokens: 1500,
324
+ });
325
+
326
+ const response = await llm.complete({
327
+ system: `You are a support agent. Customer history:\n${ctx.systemPrompt}`,
328
+ user: message,
329
+ });
330
+
331
+ return response.text;
332
+ }
333
+
334
+ // Every 100th interaction, old history auto-compresses.
335
+ // A customer with 10,000 messages still fits in a $0.02 LLM call.
336
+ ```
337
+
338
+ ### RAG Pipeline
339
+
340
+ ```typescript
341
+ // Index documents as observation memories
342
+ for (const doc of documents) {
343
+ await ms.memory.store({
344
+ actorId: "knowledge-base",
345
+ content: doc.text,
346
+ memoryType: "observation",
347
+ metadata: { source: doc.url, section: doc.section },
348
+ });
349
+ }
350
+
351
+ // Query with semantic search
352
+ const relevantDocs = await ms.memory.retrieve({
353
+ actorId: "knowledge-base",
354
+ query: "How does authentication work?",
355
+ strategy: "semantic",
356
+ limit: 5,
357
+ });
358
+
359
+ const ctx = await ms.memory.compileContext({
360
+ actorId: "knowledge-base",
361
+ memoryTypes: ["observation"],
362
+ });
363
+
364
+ // Prompt the LLM with retrieved context
365
+ const answer = await llm.complete({
366
+ system: `Answer using only these documents:\n${ctx.systemPrompt}`,
367
+ user: "How does authentication work?",
368
+ });
369
+ ```
370
+
371
+ ### Multi-User Chatbot
372
+
373
+ ```typescript
374
+ // Each user gets their own memory space
375
+ async function chat(userId: string, message: string) {
376
+ await ms.memory.store({
377
+ actorId: userId,
378
+ content: message,
379
+ });
380
+
381
+ const ctx = await ms.memory.compileContext({
382
+ actorId: userId,
383
+ maxTokens: 1000,
384
+ });
385
+
386
+ return llm.complete({
387
+ system: `You are a friendly assistant. Conversation history with this user:\n${ctx.systemPrompt}`,
388
+ user: message,
389
+ });
390
+ }
391
+
392
+ // Get stats
393
+ const total = await ms.memory.count();
394
+ const userCount = await ms.memory.count({ actorId: "user-42" });
395
+ ```
396
+
397
+ ---
398
+
399
+ ## Memory Type Reference
400
+
401
+ | Type | Purpose | Example |
402
+ |------|---------|---------|
403
+ | `interaction` | Default. Direct exchanges between agent and user/other agent. | "User asked about billing." |
404
+ | `summary` | Compressed collection of old interactions. Created by `summarize()`. | "Over 3 weeks, user reported 5 login failures..." |
405
+ | `observation` | Passive knowledge — facts, documents, things the agent knows but didn't interact with. | "Company refund policy is 30 days from purchase." |
406
+ | `gossip` | Information about third parties. For multi-agent systems. | "Agent-B told me the user is a power user." |
407
+
408
+ Types control retrieval behavior — `compileContext()` treats `interaction` and `summary` differently from `observation`. Use types to separate "what happened" from "what I know."
409
+
410
+ ---
411
+
412
+ ## Retrieval Strategies
413
+
414
+ Four strategies, each with a purpose:
415
+
416
+ ```typescript
417
+ // "What just happened?" — most recent first
418
+ await ms.memory.retrieve({ actorId: "x", strategy: "recent", limit: 3 });
419
+
420
+ // "What matters most?" — highest importance, ignoring age
421
+ await ms.memory.retrieve({ actorId: "x", strategy: "important" });
422
+
423
+ // "What relates to this query?" — cosine similarity search (needs embeddings)
424
+ await ms.memory.retrieve({ actorId: "x", query: "login bug", strategy: "semantic" });
425
+
426
+ // "Balance relevance and importance" — semantic + importance blend
427
+ await ms.memory.retrieve({ actorId: "x", query: "login bug", strategy: "hybrid" });
428
+ ```
429
+
430
+ **Choosing a strategy:**
431
+ - Use `recent` for chatbots, ongoing conversations, anything time-sensitive
432
+ - Use `important` for long-running agents where signal-to-noise matters
433
+ - Use `semantic` for RAG, document search, knowledge base queries
434
+ - Use `hybrid` for most agent memory — it balances meaning with significance
435
+
436
+ ---
437
+
438
+ ## Embeddings
439
+
440
+ Embeddings power semantic search. They're optional — without them, retrieval uses keyword matching.
441
+
442
+ **With embeddings** (configure an `EmbeddingProvider`): each `store()` computes a vector. `retrieve()` with `"semantic"` or `"hybrid"` uses cosine similarity ranking.
443
+
444
+ **Without embeddings**: everything still works — retrieval falls back to importance + recency + keyword filters. No API costs, no setup.
445
+
446
+ **Batch embedding:** `storeBatch()` sends all texts in one embedding API call, reducing cost and latency.
447
+
448
+ ```typescript
449
+ // Disable auto-embedding if you only need keyword search
450
+ const ms = new MemStack({
451
+ llm,
452
+ embedding: new OpenAIEmbeddingAdapter({ apiKey }),
453
+ defaults: { embedOnStore: false },
454
+ });
455
+ ```
456
+
457
+ ---
458
+
459
+ ## Adapters
460
+
461
+ MemStack is provider-agnostic. Every boundary is an interface — bring your own LLM, embedding model, and storage backend.
462
+
463
+ ### LLM Adapters
464
+
465
+ Used by `summarize()` and `compileContext()`. Ships with OpenAI and Anthropic built-in.
466
+
467
+ ```typescript
468
+ // OpenAI
469
+ import { OpenAILLMAdapter } from "@memstack/core";
470
+ const llm = new OpenAILLMAdapter({
471
+ apiKey: process.env.OPENAI_API_KEY!,
472
+ defaultModel: "gpt-4o-mini", // default
473
+ baseURL: "https://api.openai.com/v1", // for proxies like LiteLLM
474
+ });
475
+
476
+ // Anthropic
477
+ import { AnthropicLLMAdapter } from "@memstack/core";
478
+ const llm = new AnthropicLLMAdapter({
479
+ apiKey: process.env.ANTHROPIC_API_KEY!,
480
+ defaultModel: "claude-sonnet-4-5-20250929",
481
+ });
482
+
483
+ // Ollama (custom — implement LLMProvider)
484
+ import type { LLMProvider } from "@memstack/core";
485
+ class OllamaAdapter implements LLMProvider {
486
+ constructor(private baseURL = "http://localhost:11434") {}
487
+ async complete(req: { system: string; user: string; model?: string }) {
488
+ const res = await fetch(`${this.baseURL}/api/generate`, {
489
+ method: "POST",
490
+ body: JSON.stringify({ model: req.model ?? "llama3.2", prompt: `${req.system}\n\n${req.user}`, stream: false }),
491
+ });
492
+ const data = await res.json() as { response: string };
493
+ return { text: data.response, tokens: { prompt: 0, completion: 0, total: 0 } };
494
+ }
495
+ }
496
+ ```
497
+
498
+ ### Embedding Adapters
499
+
500
+ Used by semantic retrieval. Ships with OpenAI built-in.
501
+
502
+ ```typescript
503
+ import { OpenAIEmbeddingAdapter } from "@memstack/core";
504
+ const embedding = new OpenAIEmbeddingAdapter({
505
+ apiKey: process.env.OPENAI_API_KEY!,
506
+ model: "text-embedding-3-small", // 1536 dimensions (default)
507
+ // model: "text-embedding-3-large", // 3072 dimensions
508
+ });
509
+ ```
510
+
511
+ ### Storage Adapters
512
+
513
+ Ships with `InMemoryStorage` (zero setup, data lost on restart). For production, implement `StorageProvider` for your database.
514
+
515
+ ```typescript
516
+ import { InMemoryStorage } from "@memstack/core";
517
+ const storage = new InMemoryStorage();
518
+ ```
519
+
520
+ **Custom storage** — implement `StorageProvider`:
521
+
522
+ ```typescript
523
+ import type { StorageProvider, MemoryStoreInput } from "@memstack/core";
524
+
525
+ class PostgresStorage implements StorageProvider {
526
+ async store(input: MemoryStoreInput): Promise<Memory> { /* INSERT */ }
527
+ async get(id: string): Promise<Memory | null> { /* SELECT */ }
528
+ async retrieve(query: MemoryRetrieveQuery, embedding?: number[]): Promise<Memory[]> { /* SELECT + filters */ }
529
+ async count(filter?: MemoryCountFilter): Promise<number> { /* SELECT COUNT */ }
530
+ async delete(id: string): Promise<void> { /* DELETE */ }
531
+ async deleteMany(ids: string[]): Promise<number> { /* DELETE batch */ }
532
+ async storeBatch(inputs: MemoryStoreInput[]): Promise<Memory[]> { /* INSERT batch */ }
533
+ async initialize(): Promise<void> { /* CREATE TABLE */ }
534
+ async close(): Promise<void> { /* close pool */ }
535
+ }
536
+ ```
537
+
538
+ See `src/adapters/storage/memory.ts` for a complete reference implementation.
539
+
540
+ ---
541
+
542
+ ## Full API Reference
543
+
544
+ ### MemStack Client
545
+
546
+ ```typescript
547
+ import { MemStack } from "@memstack/core";
548
+
549
+ const ms = new MemStack({
550
+ llm: LLMProvider, // Required — for summarization
551
+ embedding?: EmbeddingProvider, // Optional — for semantic search
552
+ storage?: StorageProvider, // Optional — defaults to InMemoryStorage
553
+ defaults?: {
554
+ summarizationThreshold?: number, // Auto-summarize every N interactions. Default: 100
555
+ embedOnStore?: boolean, // Auto-embed on store(). Default: true
556
+ pruneStrategy?: PruneStrategy, // Auto-prune on every store(). Default: disabled
557
+ },
558
+ hooks?: {
559
+ onMemoryStored?: (memory: Memory) => void;
560
+ onMemoryPruned?: (ids: string[]) => void;
561
+ onSummaryCreated?: (summary: Memory, deletedCount: number) => void;
562
+ },
563
+ });
564
+ ```
565
+
566
+ ### Memory Subsystem
567
+
568
+ All methods accessible via `ms.memory.*`:
569
+
570
+ ```typescript
571
+ // Store
572
+ ms.memory.store(input: MemoryStoreInput): Promise<Memory>
573
+ ms.memory.storeBatch(inputs: MemoryStoreInput[]): Promise<Memory[]>
574
+
575
+ // Retrieve
576
+ ms.memory.retrieve(query: MemoryRetrieveQuery): Promise<Memory[]>
577
+ ms.memory.get(id: string): Promise<Memory | null>
578
+
579
+ // Context assembly
580
+ ms.memory.compileContext(options: ContextOptions): Promise<CompiledContext>
581
+
582
+ // Lifecycle
583
+ ms.memory.summarize(options: SummarizeOptions): Promise<{ summary: Memory; deletedCount: number }>
584
+ ms.memory.prune(strategy: PruneStrategy): Promise<{ pruned: string[]; count: number }>
585
+ ms.memory.dryRunPrune(strategy: PruneStrategy): Promise<{ wouldPrune: string[]; count: number }>
586
+
587
+ // Management
588
+ ms.memory.count(filter?: MemoryCountFilter): Promise<number>
589
+ ms.memory.delete(id: string): Promise<void>
590
+ ms.memory.deleteMany(ids: string[]): Promise<number>
591
+ ms.memory.touch(id: string): Promise<void> // bump recency without changing content
592
+ ```
593
+
594
+ ### Export / Import
595
+
596
+ Snapshot and restore full state for persistence, backups, or migration:
597
+
598
+ ```typescript
599
+ // Save
600
+ const snapshot = await ms.export();
601
+ fs.writeFileSync("state.json", JSON.stringify(snapshot, null, 2));
602
+
603
+ // Restore
604
+ const data = JSON.parse(fs.readFileSync("state.json", "utf-8"));
605
+ await ms2.import(data);
606
+ ```
607
+
608
+ ### Health & Close
609
+
610
+ ```typescript
611
+ const status = await ms.health();
612
+ // { storage: true, llm: true, embedding: true }
613
+
614
+ await ms.close(); // graceful shutdown
615
+ ```
616
+
617
+ ---
618
+
619
+ ## Configuration
620
+
621
+ ```typescript
622
+ const ms = new MemStack({
623
+ llm: new OpenAILLMAdapter({ apiKey: "..." }),
624
+
625
+ // Defaults control auto-behavior
626
+ defaults: {
627
+ summarizationThreshold: 50, // Summarize every 50 interactions (default: 100)
628
+ embedOnStore: false, // Don't auto-embed — saves API costs
629
+ pruneStrategy: { // Auto-clean on every store()
630
+ type: "byAge",
631
+ maxAge: 90 * 86400000, // 90 days
632
+ },
633
+ },
634
+
635
+ // Hooks for observability
636
+ hooks: {
637
+ onMemoryStored: (m) => logger.debug("memory:stored", { id: m.id, actor: m.actorId }),
638
+ onMemoryPruned: (ids) => logger.info("memory:pruned", { count: ids.length }),
639
+ onSummaryCreated: (summary, n) => logger.info("memory:summarized", { count: n }),
640
+ },
641
+ });
642
+ ```
643
+
644
+ ---
645
+
646
+ ## Advanced Usage
647
+
648
+ ### Custom Storage
649
+
650
+ Implement `StorageProvider` for any database. The interface is 9 methods. See the reference section above for the full contract.
651
+
652
+ ### Custom LLM / Embedding
653
+
654
+ Implement `LLMProvider` or `EmbeddingProvider` for any service:
655
+
656
+ ```typescript
657
+ import type { LLMProvider } from "@memstack/core";
658
+
659
+ class TogetherAIAdapter implements LLMProvider {
660
+ async complete(req: { system: string; user: string; model?: string }) {
661
+ const res = await fetch("https://api.together.xyz/v1/chat/completions", {
662
+ headers: { Authorization: `Bearer ${this.apiKey}`, "Content-Type": "application/json" },
663
+ body: JSON.stringify({ model: req.model, messages: [{ role: "system", content: req.system }, { role: "user", content: req.user }] }),
664
+ });
665
+ const data = await res.json() as any;
666
+ return { text: data.choices[0].message.content, tokens: { prompt: data.usage.prompt_tokens, completion: data.usage.completion_tokens, total: data.usage.total_tokens } };
667
+ }
668
+ }
669
+ ```
670
+
671
+ ### Event Hooks
672
+
673
+ Monitor memory operations without modifying code:
674
+
675
+ ```typescript
676
+ const ms = new MemStack({
677
+ llm,
678
+ hooks: {
679
+ onMemoryStored: (m) => metrics.increment("memory.stored"),
680
+ onSummaryCreated: (_, n) => metrics.gauge("memory.summarized_count", n),
681
+ onMemoryPruned: (ids) => metrics.increment("memory.pruned", ids.length),
682
+ },
683
+ });
684
+ ```
685
+
686
+ ---
687
+
688
+ ---
689
+
690
+ ## Development
691
+
692
+ ### Setup & Tests
693
+
694
+ ```bash
695
+ git clone https://github.com/isiomaC/memstack.git
696
+ cd memstack
697
+ pnpm install
698
+
699
+ pnpm test # 56 tests, no external services needed
700
+ pnpm test:watch # Watch mode
701
+ pnpm build # CJS + ESM + type declarations
702
+ pnpm check # TypeScript type-check only
703
+ ```
704
+
705
+ ### Debugging
706
+
707
+ Use hooks for observability — MemStack has no built-in logging:
708
+
709
+ ```typescript
710
+ const ms = new MemStack({
711
+ llm,
712
+ hooks: {
713
+ onMemoryStored: (m) => console.debug("[memstack] stored:", m.id, m.content.slice(0, 80)),
714
+ onMemoryPruned: (ids) => console.debug("[memstack] pruned:", ids.length),
715
+ },
716
+ });
717
+ ```
718
+
719
+ **Common issues:**
720
+
721
+ | Symptom | Cause | Fix |
722
+ |---------|-------|-----|
723
+ | `CONFIG_ERROR: LLM provider is required` | No LLM adapter | Pass any `LLMProvider` to config |
724
+ | Empty retrieval results | Wrong `actorId` or no memories stored | Check `await ms.memory.count({ actorId })` |
725
+ | Semantic search not working | No embedding adapter or `embedOnStore: false` | Add embedding adapter or use `strategy: "recent"` |
726
+ | High memory usage in production | Using InMemoryStorage | Implement `StorageProvider` for Postgres/Redis/etc |
727
+ | Poor summarization quality | Default prompt doesn't match your domain | Use `summarizationPrompt` in `defaults` config |
728
+
729
+ **Inspecting state at runtime:**
730
+
731
+ ```typescript
732
+ // How much data do we have?
733
+ const total = await ms.memory.count();
734
+ const perActor = await ms.memory.count({ actorId: "user-42" });
735
+
736
+ // What does one actor's memory look like?
737
+ const snapshot = await ms.export();
738
+ const actorMemories = snapshot.memories.filter(m => m.actorId === "user-42");
739
+ console.log(`User-42: ${actorMemories.length} memories`);
740
+ actorMemories.forEach(m => console.log(` [${m.memoryType}] ${m.content.slice(0, 60)} (imp: ${m.importance})`));
741
+ ```
742
+
743
+ ---
744
+
745
+ ## Publishing to npm
746
+
747
+ ```bash
748
+ # Bump version, then:
749
+ pnpm build && pnpm check && pnpm test
750
+ npm login
751
+ npm publish --access public
752
+ ```
753
+
754
+ The `@memstack` scope requires `--access public` on first publish.
755
+
756
+ ---
757
+
758
+ ## Contributing
759
+
760
+ Most needed contributions:
761
+
762
+ - **Storage adapters**: Postgres, Redis, SQLite, filesystem
763
+ - **LLM adapters**: Ollama, Groq, Together AI, Gemini
764
+ - **Embedding adapters**: Cohere, Voyage AI, local transformers.js
765
+ - **Tests**: Edge cases, concurrent access, large-scale benchmarks
766
+ - **Docs**: Architecture diagrams, tutorials
767
+
768
+ Open an issue or PR at [github.com/isiomaC/memstack](https://github.com/isiomaC/memstack).
769
+
770
+ ---
771
+
772
+ ## License
773
+
774
+ MIT © [MemStack](https://github.com/isiomaC/memstack)