stratagate-dsh 0.2.71 → 0.2.73
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +11 -0
- package/README.md +324 -213
- package/README.zh-CN.md +326 -213
- package/dist/client.js +3 -3
- package/dist/plugin.js +267 -31
- package/dist/plugin.js.map +1 -1
- package/docs/DSH.md +4 -2
- package/docs/DSH.zh-CN.md +4 -2
- package/docs/assets/41fc676096d0a13337a1c03aaf8f499b.png +0 -0
- package/docs/assets/aaed14b0b43a76334008117f6ca104af.png +0 -0
- package/docs/assets/cecc9d191a4b9bf22a479623e2ebdc1d.png +0 -0
- package/docs/assets/fc07e5b6e1cc07c115faa773a2718aa9.png +0 -0
- package/package.json +25 -24
package/README.md
CHANGED
|
@@ -4,9 +4,9 @@
|
|
|
4
4
|
|
|
5
5
|
# StrataGate
|
|
6
6
|
|
|
7
|
-
###
|
|
7
|
+
### Recent conversations stay detailed. Older memories grow more concise.
|
|
8
8
|
|
|
9
|
-
StrataGate
|
|
9
|
+
StrataGate gradually condenses an AI agent's short-term memory as the conversation progresses, with original details available to expand when needed. Important information becomes Events and a knowledge graph, preserving historical changes and current state for future sessions.
|
|
10
10
|
|
|
11
11
|
[](https://github.com/diqierjia/StrataGate-AgentMemory/actions/workflows/ci.yml)
|
|
12
12
|
[](https://www.npmjs.com/package/stratagate-dsh)
|
|
@@ -22,17 +22,35 @@ StrataGate helps long-running AI agents remember across sessions without turning
|
|
|
22
22
|
|
|
23
23
|
</div>
|
|
24
24
|
|
|
25
|
-
>
|
|
25
|
+
> **In plain words:** StrataGate keeps recent conversations detailed and gradually condenses older ones as the conversation progresses. Important information becomes Events and a knowledge graph for future sessions. Original records remain available to expand and verify when needed.
|
|
26
26
|
|
|
27
27
|
## Why StrataGate?
|
|
28
28
|
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
29
|
+
1. **Short-term memory: details fade as the conversation progresses and expand when needed.**
|
|
30
|
+
|
|
31
|
+
(1) **Recent history stays detailed; older history becomes concise.** Each conversation block has six views, L0–L5, with different levels of detail. As more conversation accumulates, older memories gradually shift from full dialogue to key facts, short summaries, and title indexes, reducing the context occupied by history. → [Layered memory](#layered-memory)
|
|
32
|
+
|
|
33
|
+
(2) **Views shrink while original records remain.** Complete L5 source messages and tool records are preserved. When details need checking, the agent can expand a memory to recover the original wording and context. → [Layered memory](#layered-memory)
|
|
34
|
+
|
|
35
|
+
2. **Long-term memory: an event timeline preserves history, while a knowledge graph organizes current state.**
|
|
36
|
+
|
|
37
|
+
(1) **Events record what happened.** Important decisions, preferences, plans, and changes are extracted from conversations as Events. Each retains its source and distinguishes when something was mentioned from when it happened, so future sessions can retrieve and trace it. → [Event cards](#event-cards)
|
|
38
|
+
|
|
39
|
+
(2) **The knowledge graph represents current state.** Historical Events provide the basis for current information and relationships about people, projects, organizations, tools, and places. New Events can supplement or supersede an earlier state while historical Events and their sources remain preserved. → [Current-state graph](#current-state-graph)
|
|
40
|
+
|
|
41
|
+
(3) **Long-term weights decay too.** As conversations progress, memories that have not been adopted gradually lose weight, affecting their priority during retrieval and automatic recall. Their sources remain available for verification even after their weights decay. → [Weights and adoption-based reinforcement](#use-only-reinforcement)
|
|
42
|
+
|
|
43
|
+
(4) **Bring memories from other AIs.** Imported content can become traceable Events and update the knowledge graph while the original imported text remains preserved. → [External memory import](#external-memory-import)
|
|
44
|
+
|
|
45
|
+
3. **Evidence gate: check whether retrieved evidence is sufficient before answering.**
|
|
46
|
+
|
|
47
|
+
A relevant search result may still be insufficient to answer the question. The agent assesses the evidence and, when needed, searches again, expands Events, or checks the original messages. If it still cannot confirm the answer, it states the uncertainty. → [Evidence gate](#evidence-gate)
|
|
48
|
+
|
|
49
|
+
4. **Reinforce only memories actually used.**
|
|
50
|
+
|
|
51
|
+
Search hits and automatic context injection do not trigger reinforcement. Only evidence recorded as actually used in the final answer increases the adoption count and resets the decay anchor. More adoptions mean slower future decay, preventing a memory from reinforcing itself merely because it is frequently retrieved. → [Use-only reinforcement](#use-only-reinforcement)
|
|
52
|
+
|
|
53
|
+
Get started: → [Quick start](#quick-start-deepseek-harness)
|
|
36
54
|
|
|
37
55
|
## Choose your path
|
|
38
56
|
|
|
@@ -63,209 +81,284 @@ Removing the plugin does not delete the database. For screenshots, configuration
|
|
|
63
81
|
|
|
64
82
|
## The problem behind the design
|
|
65
83
|
|
|
66
|
-
|
|
84
|
+
As conversations accumulate, an AI agent has to fit the current task, earlier discussions, and lasting information into a limited context window.
|
|
67
85
|
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
StrataGate designs long-term memory around four core problems:
|
|
86
|
+
Recent discussions often need their full detail. Older conversations can remain as concise summaries and expand when needed. Meanwhile, user preferences, project decisions, and task plans change, making it necessary to distinguish historical records from current state. StrataGate addresses these needs by managing short-term views, long-term updates, retrieval assessment, and usage feedback separately.
|
|
71
87
|
|
|
72
88
|
| Common problem | How StrataGate handles it |
|
|
73
89
|
| --- | --- |
|
|
74
|
-
|
|
|
75
|
-
|
|
|
76
|
-
| Search finds related
|
|
77
|
-
|
|
|
90
|
+
| Conversations keep growing, and historical details continue to occupy context | Store each conversation block as L0–L5 views; as more conversation accumulates, older content defaults to a more concise view and expands when needed |
|
|
91
|
+
| Important information is scattered across conversations, and old decisions can be confused with current state | Extract Events with sources and timestamps, preserve history in an event timeline, and organize current state in a knowledge graph |
|
|
92
|
+
| Search finds related content but misses details needed to answer | Assess the evidence through the evidence gate; search again, expand memories, or inspect the source when necessary |
|
|
93
|
+
| A memory keeps gaining weight merely because it is frequently retrieved | Separate retrieval from adoption: long-term weights decay as conversations progress, and only memories recorded as actually used receive reinforcement |
|
|
78
94
|
|
|
79
|
-
|
|
95
|
+
Short-term decay gradually reduces the detail that older conversations contribute to current context. Long-term memory and its weighting mechanism help the agent recall information that remains useful in future sessions. Both retain sources so condensed information can still be traced and checked.
|
|
80
96
|
|
|
81
|
-
|
|
97
|
+
<a id="how-stratagate-works"></a>
|
|
98
|
+
|
|
99
|
+
## How it works
|
|
82
100
|
|
|
83
|
-
|
|
101
|
+

|
|
84
102
|
|
|
85
|
-
-
|
|
86
|
-
- 35 sessions;
|
|
87
|
-
- 152 category 1–4 questions;
|
|
88
|
-
- 10 independent Judge evaluations per question.
|
|
103
|
+
*The four-round retrieval budget in Figure 1 is an example evaluation setting. Integrations control their own retrieval loops and budgets. The diagrams use Chinese labels; the accompanying text explains each mechanism in English.*
|
|
89
104
|
|
|
90
|
-
|
|
91
|
-
| --- | ---: | ---: | ---: |
|
|
92
|
-
| Mean accuracy across 10 Judge runs | **80.46%** | 63.22% | **+17.24 percentage points** |
|
|
93
|
-
| Majority-correct | **121 / 152 (79.61%)** | 96 / 152 (63.16%) | **+25 questions** |
|
|
94
|
-
| Temporal | **74.86%** | 34.59% | **+40.27 percentage points** |
|
|
95
|
-
| Single-hop | **89.29%** | 75.14% | **+14.14 percentage points** |
|
|
96
|
-
| Multi-hop | **66.56%** | 61.56% | +5.00 percentage points |
|
|
97
|
-
| Open-domain | 83.08% | **84.62%** | -1.54 percentage points |
|
|
105
|
+
StrataGate's workflow covers memory formation, recall when answering, and feedback after adoption.
|
|
98
106
|
|
|
99
|
-
|
|
107
|
+
### 1. Accumulate conversations into layered short-term memory
|
|
100
108
|
|
|
101
|
-
|
|
109
|
+
Each completed turn and its tool records are saved first. The DeepSeek Harness plugin seals a Block every **6 turns** by default. Users can configure this size; turns below the boundary remain in the current conversation.
|
|
102
110
|
|
|
103
|
-
|
|
111
|
+
Each Block stores six levels of detail, L0–L5, from title indexes and short summaries to complete source records. Decay begins only after background processing finishes: as more Blocks become ready in the same conversation, older Blocks default to shallower views.
|
|
104
112
|
|
|
105
|
-
|
|
106
|
-
- [`benchmarks/locomo-conv26-r8-final.json`](benchmarks/locomo-conv26-r8-final.json)
|
|
113
|
+
This reduces the historical detail presented in context while preserving the complete L5 source records. → [Layered memory](#layered-memory)
|
|
107
114
|
|
|
108
|
-
|
|
115
|
+
### 2. Turn important information into Events and update the knowledge graph
|
|
109
116
|
|
|
110
|
-
|
|
117
|
+
Decisions, preferences, plans, and changes worth keeping become Events. Each Event records content, time, and provenance for future conversations to retrieve.
|
|
118
|
+
|
|
119
|
+
The knowledge graph then organizes current information and relationships about entities such as people, projects, and tools. New Events can supplement or supersede an earlier state while historical Events and their sources remain available. → [Event cards](#event-cards) · [Current-state graph](#current-state-graph)
|
|
120
|
+
|
|
121
|
+
### 3. Activate relevant memories before answering and retrieve more when needed
|
|
122
|
+
|
|
123
|
+
Before each main-model call, the plugin uses the current question and recent conversation to bring in a small set of relevant long-term Events and graph information as background.
|
|
124
|
+
|
|
125
|
+
The current DSH integration includes at most **4 Events and 4 graph nodes**, subject to a total budget of approximately **900 tokens**. The active conversation's history is supplied through layered Blocks and unsealed turns.
|
|
126
|
+
|
|
127
|
+
If the available context is sufficient, the agent can answer directly. Otherwise, it uses memory tools to search Events, the graph, or source messages, expanding details as needed.
|
|
111
128
|
|
|
112
|
-
|
|
129
|
+
### 4. Assess whether actively retrieved evidence answers the question
|
|
113
130
|
|
|
114
|
-
|
|
131
|
+
After finding relevant memories, the agent determines whether they answer the actual question: which evidence supports the answer, what is missing, and whether further retrieval is needed.
|
|
115
132
|
|
|
116
|
-
|
|
117
|
-
2. **Build smaller views.** StrataGate creates layered summaries, Events that describe what happened, and graph facts that describe the current state.
|
|
118
|
-
3. **Search small records first.** The agent starts with compact results and expands an Event, graph node, or source Block only when it needs more detail.
|
|
119
|
-
4. **Check before answering.** The evidence gate decides whether the result is sufficient. If not, the agent searches again or returns to the original messages.
|
|
120
|
-
5. **Reinforce only what helped.** A memory gains long-term weight only after the final answer actually uses it.
|
|
133
|
+
For example, “the project uses pnpm” answers which tool is currently used, but does not explain why pnpm was originally chosen. The agent should expand the relevant Event or inspect the original discussion.
|
|
121
134
|
|
|
122
|
-
|
|
135
|
+
The evidence gate requires an explicit assessment, while code validates evidence references and protocol constraints. If evidence remains insufficient, the agent continues searching or states that it cannot confirm the answer. → [Evidence gate](#evidence-gate)
|
|
136
|
+
|
|
137
|
+
### 5. Record adopted evidence and update long-term weights
|
|
138
|
+
|
|
139
|
+
Search hits and automatic context injection do not reinforce memories.
|
|
140
|
+
|
|
141
|
+
When the agent determines which evidence supports its final answer, it submits a usage receipt. Events recorded as actually used update their adoption counts and decay anchors. More adoptions lead to slower future decay. Unused search results receive no reinforcement from that selection. → [Use-only reinforcement](#use-only-reinforcement)
|
|
142
|
+
|
|
143
|
+
For example, a user first decides to use npm, then switches the project to pnpm. StrataGate preserves both decisions and updates the graph's current state. A later question about what the project uses can be answered from that state; questions about when or why the change happened can follow the Event back to the original conversation.
|
|
144
|
+
|
|
145
|
+
[See a retrieval example that follows an event card back to source messages](#a-real-retrieval-path).
|
|
123
146
|
|
|
124
147
|
## Core design
|
|
125
148
|
|
|
149
|
+
StrataGate separately manages conversation detail, long-term information updates, retrieval assessment, and feedback from adopted memories.
|
|
150
|
+
|
|
151
|
+
**Short-term decay controls how much detail older conversations display. Long-term weights influence information's priority in later recall.** Neither form of decay deletes the original sources.
|
|
152
|
+
|
|
126
153
|
<a id="layered-memory"></a>
|
|
127
154
|
|
|
128
|
-
### 1.
|
|
155
|
+
### 1. Short-term memory: gradually condense older conversations and expand them when needed
|
|
156
|
+
|
|
157
|
+
Recent discussions usually need their full detail. Older conversations can remain in context as concise views. StrataGate stores several levels of detail for the same conversation block and gradually reduces what older Blocks display by default as more conversation accumulates.
|
|
129
158
|
|
|
130
|
-
|
|
159
|
+

|
|
131
160
|
|
|
132
|
-
|
|
161
|
+
**One conversation block, six levels of detail.**
|
|
133
162
|
|
|
134
|
-
|
|
163
|
+
The DeepSeek Harness plugin seals a Block every **6 complete turns** by default. A turn consists of a user question and the assistant's complete response. The Block size is configurable; the core-library default is 12 turns. Content below the boundary remains in the current conversation.
|
|
164
|
+
|
|
165
|
+
Each fully processed Block contains these views:
|
|
135
166
|
|
|
136
167
|
| Level | Contents | Primary use |
|
|
137
168
|
| --- | --- | --- |
|
|
138
|
-
| L0 | Title and tags |
|
|
139
|
-
| L1 | Short summary |
|
|
140
|
-
| L2 | Key facts |
|
|
141
|
-
| L3 |
|
|
142
|
-
| L4 | Readable near-verbatim
|
|
143
|
-
| L5 |
|
|
169
|
+
| L0 | Title and tags | Identify a piece of history with minimal context |
|
|
170
|
+
| L1 | Short summary | Understand the discussion's topic |
|
|
171
|
+
| L2 | Key facts | Review decisions, constraints, plans, and outcomes |
|
|
172
|
+
| L3 | Rule-condensed dialogue | Preserve the discussion while removing bounded redundancy |
|
|
173
|
+
| L4 | Readable, near-verbatim dialogue | Check fuller language context |
|
|
174
|
+
| L5 | Source messages and tool records | Verify provenance and specific details |
|
|
175
|
+
|
|
176
|
+
The model produces L0–L2 summaries. Code generates L3 and L4 deterministically. L3 may condense standalone greetings, pure confirmations, repeated long pasted content, and tool arguments; it does not perform free-form semantic paraphrasing.
|
|
177
|
+
|
|
178
|
+
Source records are saved first, followed by summarization and Event processing. Only a fully processed, ready Block can replace its corresponding native history and participate in decay.
|
|
179
|
+
|
|
180
|
+
**As more conversation accumulates, older Blocks default to less detail.**
|
|
181
|
+
|
|
182
|
+
Display changes follow exponential decay:
|
|
183
|
+
|
|
184
|
+
$$
|
|
185
|
+
w_{\text{block}}(age)=e^{-\lambda_{\text{block}}\,age}
|
|
186
|
+
$$
|
|
187
|
+
|
|
188
|
+
The default $\lambda_{\text{block}}$ is **0.30**. Code maps weight ranges to display levels. Smaller coefficients preserve detail for longer and therefore consume more context.
|
|
189
|
+
|
|
190
|
+
In this formula, `age` is the distance between the current display anchor and the latest ready Block in the same conversation. It measures conversation progress, **not elapsed calendar days**. Unsealed turns and Blocks still awaiting model processing do not advance this decay.
|
|
144
191
|
|
|
145
|
-
|
|
192
|
+
For a Block that starts at L5 and is never expanded again, the default schedule is:
|
|
146
193
|
|
|
147
|
-
|
|
194
|
+
| Additional ready Blocks | Default display level |
|
|
195
|
+
| ---: | --- |
|
|
196
|
+
| 0–1 | L5 |
|
|
197
|
+
| 2 | L4 |
|
|
198
|
+
| 3–4 | L3 |
|
|
199
|
+
| 5–6 | L2 |
|
|
200
|
+
| 7–8 | L1 |
|
|
201
|
+
| 9 or more | L0 |
|
|
148
202
|
|
|
149
|
-
|
|
203
|
+
The figure illustrates the trend. Actual changes depend on both the decay coefficient and level thresholds; a new Block does not necessarily cause a one-level drop.
|
|
150
204
|
|
|
151
|
-
|
|
152
|
-
|
|
205
|
+
**Expand again when details are needed.**
|
|
206
|
+
|
|
207
|
+
Suppose an older conversation currently shows only:
|
|
208
|
+
|
|
209
|
+
> Discussed the project's technical approach and near-term plans.
|
|
210
|
+
|
|
211
|
+
If the user asks why pnpm was chosen, the agent can expand key facts, condensed dialogue, or the complete source records to find the original reason.
|
|
212
|
+
|
|
213
|
+
Expansion can proceed one level at a time or jump directly to a requested level. The selected level and current Block position become the new decay anchor. As subsequent conversation accumulates, the view gradually becomes concise again.
|
|
214
|
+
|
|
215
|
+
Older conversations can therefore remain lightweight during ordinary use while retaining recoverable detail. **L0–L4 are derived views of the source and never overwrite L5.**
|
|
153
216
|
|
|
154
217
|
<a id="event-cards"></a>
|
|
155
218
|
|
|
156
|
-
### 2.
|
|
219
|
+
### 2. Long-term memory: preserve history in Events and organize current state in a knowledge graph
|
|
157
220
|
|
|
158
|
-
|
|
221
|
+
Short-term memory retains a discussion's context. Long-term memory extracts information worth using in future sessions. StrataGate records decisions, preferences, plans, and changes as Events, then uses those Events to organize a knowledge graph.
|
|
159
222
|
|
|
160
|
-
|
|
223
|
+

|
|
161
224
|
|
|
162
|
-
|
|
163
|
-
{
|
|
164
|
-
sourceBlockId,
|
|
165
|
-
sourceMessageIds,
|
|
225
|
+
**Event cards record what happened and retain their sources.**
|
|
166
226
|
|
|
167
|
-
|
|
168
|
-
happenedStart,
|
|
169
|
-
happenedEnd,
|
|
227
|
+
A Block can produce multiple Events or contain no information that needs long-term extraction. In addition to content, each Event retains its source Block, source messages, and whatever temporal information can be established.
|
|
170
228
|
|
|
171
|
-
|
|
172
|
-
participants,
|
|
173
|
-
eventType,
|
|
229
|
+
Two different time axes must be distinguished:
|
|
174
230
|
|
|
175
|
-
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
|
|
231
|
+
| Time | Meaning |
|
|
232
|
+
| --- | --- |
|
|
233
|
+
| Mention time | When the conversation referred to the event |
|
|
234
|
+
| Occurrence time | When the event happened or is planned to happen |
|
|
179
235
|
|
|
180
|
-
|
|
236
|
+
If a user says “Finish the prototype next week” on May 6, May 6 is the mention date, while “next week” describes the planned completion time. The record should retain its planned status and cannot serve as evidence that the prototype is already complete.
|
|
181
237
|
|
|
182
|
-
|
|
183
|
-
- `happenedStart` / `happenedEnd` describe when it actually happened or is expected to happen;
|
|
184
|
-
- `status` distinguishes completed, planned, cancelled, and ongoing events;
|
|
185
|
-
- `supersedesEventIds` and `conflictsWithEventIds` preserve corrections and conflicts.
|
|
238
|
+
When a date cannot be established, the original wording and uncertainty remain available for later source verification.
|
|
186
239
|
|
|
187
|
-
|
|
240
|
+
**New Events update memory through additions, supersession, or conflicts.**
|
|
188
241
|
|
|
189
|
-
|
|
242
|
+
A project discussion might contain these statements over time:
|
|
243
|
+
|
|
244
|
+
> “Use npm for this project.”
|
|
245
|
+
>
|
|
246
|
+
> “Let's switch to pnpm.”
|
|
247
|
+
>
|
|
248
|
+
> “Finish the prototype next week.”
|
|
249
|
+
|
|
250
|
+
They play different roles:
|
|
251
|
+
|
|
252
|
+
- Switching to pnpm updates the package-manager choice; the earlier npm Event remains in history.
|
|
253
|
+
- Finishing the prototype next week adds a plan without changing the package-manager decision.
|
|
254
|
+
- If statements cannot both be true and their validity cannot yet be resolved, a conflict relationship is retained for later verification.
|
|
255
|
+
|
|
256
|
+
New Events do not overwrite the content or provenance of earlier Events. Their validity status and relationships can change as new evidence arrives. This lets the system retrieve current information while also explaining what used to be true and what changed.
|
|
190
257
|
|
|
191
258
|
<a id="current-state-graph"></a>
|
|
192
259
|
|
|
193
|
-
|
|
260
|
+
**The knowledge graph organizes current information and relationships from Events.**
|
|
194
261
|
|
|
195
|
-
|
|
262
|
+
People, projects, organizations, tools, and places become nodes. Connections such as “uses,” “participates in,” and “depends on” become directed edges. Both node attributes and relationships retain their source Events.
|
|
196
263
|
|
|
197
|
-
|
|
264
|
+
In the example, the graph can update the project's current package manager to pnpm while preserving npm as a historical state. Later:
|
|
198
265
|
|
|
199
|
-
|
|
266
|
+
- a question about what the project uses can start with the current graph;
|
|
267
|
+
- a question about when it changed can inspect the change Event;
|
|
268
|
+
- a question about why it changed can expand the Event and return to the original discussion.
|
|
269
|
+
|
|
270
|
+
Graph state must be supported by its sources. “Finish the prototype next week” directly supports a plan. The figure's “prototype in development” state and owner information require additional supporting Events.
|
|
271
|
+
|
|
272
|
+
Graph updates run as independent jobs with persisted progress. A failed update can be retried separately while the committed Events and source records remain available.
|
|
200
273
|
|
|
201
274
|
<a id="evidence-gate"></a>
|
|
202
275
|
|
|
203
|
-
###
|
|
276
|
+
### 3. Evidence gate: check whether retrieved results can answer the actual question
|
|
204
277
|
|
|
205
|
-
|
|
278
|
+
After finding relevant memories, the agent still needs to assess whether they support the answer. StrataGate uses a fixed, compact assessment structure that requires an explicit judgment of sufficiency and the next action.
|
|
206
279
|
|
|
207
|
-
|
|
208
|
-
|
|
209
|
-
|
|
280
|
+
| Field | What it explains |
|
|
281
|
+
| --- | --- |
|
|
282
|
+
| `verdict` | Whether evidence is sufficient, partial, or mismatched |
|
|
283
|
+
| `evidence_refs` | Which retrieved items support the assessment |
|
|
284
|
+
| `fit` | How the evidence matches the question |
|
|
285
|
+
| `missing` | What information is still missing |
|
|
286
|
+
| `next_strategy` | Whether to answer, search again, or expand a memory |
|
|
210
287
|
|
|
211
|
-
|
|
288
|
+
For example, the user asks:
|
|
212
289
|
|
|
213
|
-
|
|
214
|
-
- which results actually support that judgment;
|
|
215
|
-
- how the evidence matches the question;
|
|
216
|
-
- what is still missing;
|
|
217
|
-
- should the next step answer, continue searching, expand an event, or inspect the original messages.
|
|
290
|
+
> Why did we originally switch to pnpm?
|
|
218
291
|
|
|
219
|
-
The
|
|
292
|
+
The only retrieved result says:
|
|
220
293
|
|
|
221
|
-
|
|
222
|
-
2. `next_strategy` is explicitly `answer`;
|
|
223
|
-
3. the judgment uses a fixed, bounded structure instead of an ever-growing private retrieval scratchpad.
|
|
294
|
+
> The project switched from npm to pnpm.
|
|
224
295
|
|
|
225
|
-
|
|
296
|
+
That confirms a change but does not explain the reason. The agent should judge the evidence as partial and expand the Event or inspect the original conversation, rather than infer the historical reason from the tool choice alone.
|
|
226
297
|
|
|
227
|
-
|
|
228
|
-
search_events
|
|
229
|
-
expand_event
|
|
230
|
-
search_graph
|
|
231
|
-
expand_graph_node
|
|
232
|
-
search_raw_memory
|
|
233
|
-
expand_block
|
|
234
|
-
```
|
|
298
|
+
The model assesses semantic sufficiency. Code validates references and protocol constraints: cited evidence must come from the selected retrieval batch, and accepting `sufficient` requires valid evidence references and an explicit choice to answer.
|
|
235
299
|
|
|
236
|
-
|
|
300
|
+
These checks make retrieval traceable and auditable, but the model can still misjudge evidence. When sufficient information is unavailable, the agent should continue searching or state that it cannot confirm the answer.
|
|
301
|
+
|
|
302
|
+
Integrations control active-retrieval loops and budgets. The four-round limit in Figure 1 is an example evaluation setting, not a fixed limit for every integration.
|
|
237
303
|
|
|
238
304
|
<a id="use-only-reinforcement"></a>
|
|
239
305
|
|
|
240
|
-
###
|
|
306
|
+
### 4. Reinforce only memories actually used: more adoptions mean slower future decay
|
|
241
307
|
|
|
242
|
-
|
|
308
|
+
Long-term memories also decay as conversations progress. Here, the changing quantity is an Event's weight, which participates in later recall and ranking. Unlike a short-term Block, an Event does not move through L0–L5 display levels as its weight decays.
|
|
243
309
|
|
|
244
|
-
|
|
310
|
+

|
|
245
311
|
|
|
246
|
-
|
|
247
|
-
await memory.recordMemoryUse({ eventIds, elementIds });
|
|
248
|
-
```
|
|
312
|
+
**New Events have an initial weight that decays when they are not adopted.**
|
|
249
313
|
|
|
250
|
-
|
|
314
|
+
The base Event-weight function is:
|
|
251
315
|
|
|
252
|
-
|
|
316
|
+
$$
|
|
317
|
+
w(t,n)=\max\left(floor,e^{-\lambda(n)t}\right)
|
|
318
|
+
$$
|
|
253
319
|
|
|
254
|
-
|
|
255
|
-
|
|
256
|
-
|
|
257
|
-
|
|
258
|
-
|
|
259
|
-
|
|
260
|
-
|
|
261
|
-
|
|
262
|
-
|
|
320
|
+
$$
|
|
321
|
+
\lambda(n)=\frac{0.15}{1+1.5\ln(n)}
|
|
322
|
+
$$
|
|
323
|
+
|
|
324
|
+
Here:
|
|
325
|
+
|
|
326
|
+
- $t$ is the difference between the current turn and the last adoption turn; new Events start counting from creation;
|
|
327
|
+
- $n$ is the internal adoption count, initialized to 1 and incremented for each recorded adoption;
|
|
328
|
+
- $floor$ is the minimum weight assigned according to the memory's criticality.
|
|
263
329
|
|
|
264
|
-
|
|
330
|
+
Long-term decay also uses conversation turns rather than elapsed wall-clock time. Lower weight may reduce a memory's priority in later recall, but decay does not delete its historical record.
|
|
331
|
+
|
|
332
|
+
**Retrieval does not trigger reinforcement.**
|
|
333
|
+
|
|
334
|
+
A search hit only establishes possible relevance. The system may record when a memory was retrieved, but retrieval does not increase its adoption count or reset its decay anchor.
|
|
335
|
+
|
|
336
|
+
Memories automatically included in context receive no reinforcement merely for being displayed. This prevents a memory from continually gaining weight just because it happened to rank highly and then appeared repeatedly.
|
|
337
|
+
|
|
338
|
+
**Weights update only after recorded adoption.**
|
|
339
|
+
|
|
340
|
+
After selecting evidence for the final answer, the agent submits a usage receipt. Validated Event selections increase their adoption counts and move their decay anchors to the current turn. An ordinary active Event without an additional weight cap returns to weight 1.
|
|
341
|
+
|
|
342
|
+
As the adoption count increases, the decay coefficient decreases. The Event retains more weight over the same number of subsequent turns, so memories that repeatedly help answers decay more slowly.
|
|
343
|
+
|
|
344
|
+
Adoption is based on the agent's submitted evidence selection. Code checks that the evidence belongs to the corresponding batch and has passed a sufficient assessment. Receipts prevent the same operation from being applied twice. Unused results receive no reinforcement from that selection.
|
|
345
|
+
|
|
346
|
+
**Different criticality levels have different minimum weights.**
|
|
347
|
+
|
|
348
|
+
| Memory category | Default minimum weight |
|
|
349
|
+
| --- | ---: |
|
|
350
|
+
| Routine information | 0 |
|
|
351
|
+
| User preference | 0.3 |
|
|
352
|
+
| Identity information | 0.9 |
|
|
353
|
+
| Safety information | 1.0 |
|
|
354
|
+
|
|
355
|
+
Pinned memories have an effective weight of 1. Superseded Events normally receive a low weight cap so older states do not retain excessive priority.
|
|
356
|
+
|
|
357
|
+
These weights express a memory-management policy, not factual accuracy. Even high-weight information must be assessed against the current question, current state, and original source.
|
|
265
358
|
|
|
266
359
|
<a id="external-memory-import"></a>
|
|
267
360
|
|
|
268
|
-
|
|
361
|
+
## Import memory from another AI
|
|
269
362
|
|
|
270
363
|
`importExternalMemory()` can migrate a structured memory summary produced by another AI. The core API extracts candidate Events, compares each candidate with a bounded set of existing Events, and lets a model choose one of five actions: add, merge, supersede, mark a conflict, or ignore. Imported text is also retained as a permanent source Block, so every accepted Event remains traceable to the exact import.
|
|
271
364
|
|
|
@@ -303,78 +396,98 @@ In this path:
|
|
|
303
396
|
|
|
304
397
|
- the event card provides fast location;
|
|
305
398
|
- the source timestamp and original message provide final verification;
|
|
306
|
-
- the evidence gate
|
|
399
|
+
- the evidence gate requires the agent to identify missing information and keep checking when evidence is insufficient.
|
|
307
400
|
|
|
308
|
-
##
|
|
309
|
-
|
|
310
|
-
The current design was not decided in one pass. The most useful result of multiple experiments was not the round number, but the failure mode each round exposed.
|
|
401
|
+
## Experimental results
|
|
311
402
|
|
|
312
|
-
|
|
313
|
-
| --- | --- | --- |
|
|
314
|
-
| Temporal information was compressed into summaries and hard to recover accurately | In the early matched-protocol experiments, adding multiple events per block and explicit occurrence times raised Temporal from 18.92% to 45.95% | Separate mention time from occurrence time, and preserve the original temporal expression and source message |
|
|
315
|
-
| The agent's retrieval scratchpad kept growing | The bounded five-field evidence gate scored 77.63%; expanding it into a larger structured scratchpad reduced the score to 63.82% | Keep the judgment small and bounded, and let code validate its critical constraints |
|
|
316
|
-
| When evidence was insufficient, the agent repeatedly searched the same event cards | An early end-to-end version had 19 questions with at least three event searches and answered only 2 correctly; the current strategy answered 15 of the same questions, including 12 that inspected the source | Change information channels when search adds no new evidence instead of repeating the same search |
|
|
403
|
+
The repository's published R8 comparison uses the LoCoMo conversation sample `conv-26`, containing **419 messages, 35 sessions, and 152 questions** across categories 1–4.
|
|
317
404
|
|
|
318
|
-
|
|
405
|
+
Each system generated answers, and each answer received **10 independent Judge evaluations**. These are repeated evaluations, not ten complete system runs.
|
|
319
406
|
|
|
320
|
-
| Metric |
|
|
407
|
+
| Metric | StrataGate | Mem0 base | Difference |
|
|
321
408
|
| --- | ---: | ---: | ---: |
|
|
322
|
-
| Mean accuracy across 10 Judge runs |
|
|
323
|
-
| Majority-correct |
|
|
324
|
-
|
|
|
325
|
-
|
|
|
326
|
-
|
|
|
409
|
+
| Mean accuracy across 10 Judge runs | **80.46%** | 63.22% | **+17.24 percentage points** |
|
|
410
|
+
| Majority-correct | **121 / 152 (79.61%)** | 96 / 152 (63.16%) | **+25 questions** |
|
|
411
|
+
| Temporal | **74.86%** | 34.59% | **+40.27 percentage points** |
|
|
412
|
+
| Single-hop | **89.29%** | 75.14% | **+14.14 percentage points** |
|
|
413
|
+
| Multi-hop | **66.56%** | 61.56% | +5.00 percentage points |
|
|
414
|
+
| Open-domain | 83.08% | **84.62%** | -1.54 percentage points |
|
|
415
|
+
|
|
416
|
+
Both systems used the same questions, order, answer model, Judge model, evaluation prompt, parser, and evaluation count, and each rebuilt its memory. Memory extraction, retrieval implementation, embedding use, and answer context differed, so this compares two complete system configurations.
|
|
327
417
|
|
|
328
|
-
These results
|
|
418
|
+
These results cover only `conv-26`, not the full LoCoMo dataset. They do not isolate the benefits of short-term decay, the knowledge graph, or the evidence gate. Individual contributions still require ablation experiments.
|
|
329
419
|
|
|
330
|
-
|
|
420
|
+
See the [evaluation document](docs/EVALUATION.md) for the full protocol, per-question results, and Judge variation, and the [machine-readable results](benchmarks/locomo-conv26-r8-final.json) for summary data.
|
|
331
421
|
|
|
332
|
-
|
|
422
|
+
## How these designs emerged
|
|
423
|
+
|
|
424
|
+
Failures observed across several experimental rounds helped StrataGate refine its memory processing and retrieval strategies:
|
|
425
|
+
|
|
426
|
+
- **Store temporal information explicitly.** Summaries alone make event dates hard to recover, so event cards distinguish mention time from occurrence time and retain the original temporal wording and source.
|
|
427
|
+
- **Keep evidence assessment short and explicit.** Five fields require the agent to state its evidence, gaps, and next action while allowing code to validate references.
|
|
428
|
+
- **Change retrieval strategy when repeated searches add no information.** When an event card lacks key details, expanding its source or searching original messages is more useful than repeatedly searching the same cards.
|
|
429
|
+
|
|
430
|
+
The [evaluation document](docs/EVALUATION.md) retains the full R1–R8 history, version differences, and per-question analysis. Multiple changes occurred between rounds; these observations explain design choices but do not independently establish the effect of any single component.
|
|
333
431
|
|
|
334
432
|
## Current limitations and next steps
|
|
335
433
|
|
|
336
|
-
|
|
434
|
+
In the R8 evaluation above, **31 questions were judged incorrect by a majority of evaluators**. Grouped by the observable failure stage:
|
|
337
435
|
|
|
338
|
-
| Failure stage | Questions |
|
|
436
|
+
| Failure stage | Questions | What it indicates |
|
|
339
437
|
| --- | ---: | --- |
|
|
340
|
-
|
|
|
341
|
-
| Evidence
|
|
342
|
-
| Evidence remained
|
|
438
|
+
| Incorrect direct answer without retrieval | 15 | The agent sometimes failed to recognize that historical evidence was needed |
|
|
439
|
+
| Evidence judged sufficient, but the final answer was incorrect | 14 | Evidence could concern a neighboring event or fail to support the complete answer |
|
|
440
|
+
| Evidence remained insufficient at the retrieval budget | 2 | Enough information was not found within the allotted budget |
|
|
343
441
|
|
|
344
|
-
|
|
442
|
+
Within this evaluation, the results point to a need to improve when retrieval starts and whether retrieved evidence actually answers the question. The evidence gate constrains references and assessment procedures, but cannot guarantee the model's semantic judgment or final answer.
|
|
345
443
|
|
|
346
|
-
|
|
444
|
+
Further validation will focus on:
|
|
347
445
|
|
|
348
|
-
1.
|
|
349
|
-
2.
|
|
350
|
-
3.
|
|
351
|
-
4. finally expand to the complete LoCoMo dataset.
|
|
446
|
+
1. Holding the model and memory state fixed while ablating short-term decay and key retrieval mechanisms, comparing accuracy, context use, and call costs.
|
|
447
|
+
2. Supplying known-correct source evidence directly to distinguish retrieval failures from reasoning errors that persist despite correct evidence.
|
|
448
|
+
3. Repeating the same protocol on more conversation samples before extending it to the full LoCoMo dataset.
|
|
352
449
|
|
|
353
450
|
## Current status
|
|
354
451
|
|
|
355
|
-
StrataGate
|
|
452
|
+
StrataGate provides an installable DeepSeek Harness plugin, with memory management implemented through a shared TypeScript core engine.
|
|
356
453
|
|
|
357
|
-
|
|
454
|
+
Current functionality includes:
|
|
358
455
|
|
|
359
|
-
-
|
|
360
|
-
-
|
|
361
|
-
-
|
|
362
|
-
-
|
|
363
|
-
-
|
|
364
|
-
-
|
|
365
|
-
|
|
366
|
-
|
|
367
|
-
|
|
456
|
+
- automatic capture of completed conversations and tool records;
|
|
457
|
+
- L0–L5 short-term memory views, display decay, and on-demand expansion;
|
|
458
|
+
- Events with provenance and time, plus a knowledge graph updated from those Events;
|
|
459
|
+
- automatic activation of relevant long-term memories, active search, and source lookup;
|
|
460
|
+
- retrieval-batch management, evidence assessment, and usage receipts;
|
|
461
|
+
- long-term weight decay, adoption-based reinforcement, and memory import from other AIs.
|
|
462
|
+
|
|
463
|
+
The repository includes automated tests, experiment records, and traceable evaluation results. Public APIs, model integration, and evaluation coverage are still evolving. Custom integrations should pin a version and validate behavior for their own use cases.
|
|
464
|
+
|
|
465
|
+
The standard `StrataGate.open()` entry point uses SQLite persistence; `StrataGate.inMemory()` explicitly selects temporary operation or testing. The storage adapter supports interruption recovery and consistency checks. See the [architecture document](docs/ARCHITECTURE.md) for its constraints.
|
|
466
|
+
|
|
467
|
+
## When StrataGate is a good fit
|
|
368
468
|
|
|
369
|
-
|
|
469
|
+
StrataGate suits agent workflows that need continuity across extended conversations while controlling the context occupied by history. For example:
|
|
370
470
|
|
|
371
|
-
|
|
471
|
+
- **Sustained project work.** Recent discussions stay detailed; older discussions become concise and can expand when the reasoning behind a decision matters.
|
|
472
|
+
- **Continuing work across sessions.** Retrieve project decisions, user preferences, plans, and unfinished work in a new conversation.
|
|
473
|
+
- **Tracking changes.** Understand both current project state and historical changes without mistaking an earlier decision for the current one.
|
|
474
|
+
- **Verifying memory sources.** Trace Events and graph information back to original records when an answer depends on dates, exact wording, or tool results.
|
|
475
|
+
- **Migrating existing memories.** Turn information exported by another AI into Events while retaining the imported source text.
|
|
476
|
+
|
|
477
|
+
Memory can be organized by project, session, or global scope. The knowledge-graph UI primarily supports viewing current information, relationships, and sources; collaborative editing and cloud synchronization across products are not its primary functions.
|
|
478
|
+
|
|
479
|
+
DeepSeek Harness users can install the plugin through the [quick start](#quick-start-deepseek-harness). See the [plugin guide](docs/DSH.md) for configuration, memory tools, and recovery behavior.
|
|
372
480
|
|
|
373
481
|
## Code entry points
|
|
374
482
|
|
|
375
|
-
|
|
483
|
+
This section is for developers. DeepSeek Harness users can follow the [quick start](#quick-start-deepseek-harness) without building the repository themselves.
|
|
484
|
+
|
|
485
|
+
The development environment requires:
|
|
486
|
+
|
|
487
|
+
- Node.js **22.19.0 or later within the 22.x series**, or **24.0.0 or later**;
|
|
488
|
+
- the declared version range is `^22.19.0 || >=24.0.0`.
|
|
376
489
|
|
|
377
|
-
After checking out the repository
|
|
490
|
+
After checking out the repository, run these commands from its root:
|
|
378
491
|
|
|
379
492
|
```bash
|
|
380
493
|
npm install
|
|
@@ -383,31 +496,44 @@ npm test
|
|
|
383
496
|
npm run build
|
|
384
497
|
```
|
|
385
498
|
|
|
386
|
-
The main code
|
|
499
|
+
The main code entry points are:
|
|
387
500
|
|
|
388
|
-
|
|
389
|
-
|
|
390
|
-
-
|
|
391
|
-
|
|
392
|
-
|
|
393
|
-
|
|
394
|
-
|
|
395
|
-
|
|
396
|
-
|
|
397
|
-
|
|
398
|
-
-
|
|
501
|
+
| Entry point | Contents |
|
|
502
|
+
| --- | --- |
|
|
503
|
+
| [Core-engine example](packages/core/examples/basic.ts) | Minimal integration example |
|
|
504
|
+
| [Core state and lifecycle](packages/core/src/store.ts) | Blocks, Events, graph, import, retrieval, and adoption records |
|
|
505
|
+
| [Layering and decay](packages/core/src/blocks.ts) | L0–L5 levels, display decay, and deterministic pruning |
|
|
506
|
+
| [Long-term weights](packages/core/src/weights.ts) | Event decay and minimum-weight calculation |
|
|
507
|
+
| [Event types](packages/core/src/events.ts) | Event-type normalization |
|
|
508
|
+
| [Knowledge graph](packages/core/src/graph.ts) | Provenance checks, graph updates, and state maintenance |
|
|
509
|
+
| [Retrieval ranking](packages/core/src/search.ts) | BM25 ranking and RRF fusion |
|
|
510
|
+
| [Evidence assessment](packages/core/src/retrieval.ts) | Assessment structure and reference constraints |
|
|
511
|
+
| [External-memory import](packages/core/src/external-memory.ts) | Import formats, prompts, and parsing |
|
|
512
|
+
| [Legacy Element-card support](packages/core/src/elements.ts) | Element projection and historical time views |
|
|
513
|
+
|
|
514
|
+
In persistent mode, `recordMemoryUse()` requires a nonempty, stable `receiptId`. Reuse that ID when retrying the same adoption operation to avoid duplicate reinforcement. For example:
|
|
515
|
+
|
|
516
|
+
```ts
|
|
517
|
+
await memory.recordMemoryUse(
|
|
518
|
+
{ eventIds: usedEventIds },
|
|
519
|
+
{ receiptId: usageReceiptId },
|
|
520
|
+
);
|
|
521
|
+
```
|
|
399
522
|
|
|
400
|
-
`
|
|
523
|
+
Here, `usedEventIds` contains the Events the application selected for the answer, and `usageReceiptId` identifies that adoption operation. The DSH plugin handles this through memory tools and its batch protocol; plugin users do not call this API manually.
|
|
524
|
+
|
|
525
|
+
The core example demonstrates API integration. See the [evaluation document](docs/EVALUATION.md) for the model calls, tool loop, and Judge protocol used in the full evaluation.
|
|
401
526
|
|
|
402
527
|
## Documentation and reproduction
|
|
403
528
|
|
|
404
529
|
| Resource | Contents |
|
|
405
530
|
| --- | --- |
|
|
406
|
-
| [
|
|
407
|
-
| [
|
|
408
|
-
| [
|
|
409
|
-
| [
|
|
410
|
-
| [
|
|
531
|
+
| [DeepSeek Harness guide](docs/DSH.md) | Installation, configuration, UI, memory tools, and recovery |
|
|
532
|
+
| [Architecture](docs/ARCHITECTURE.md) | Layering, Events and graph, retrieval, evidence gate, weights, and storage constraints |
|
|
533
|
+
| [External-memory import](docs/EXTERNAL_MEMORY_IMPORT.zh-CN.md) | Export format, import flow, and integration example |
|
|
534
|
+
| [Full evaluation](docs/EVALUATION.md) | Protocol, version history, failure analysis, and result scope |
|
|
535
|
+
| [Evaluation summary data](benchmarks/locomo-conv26-r8-final.json) | Published results, statistics, and artifact information |
|
|
536
|
+
| [Core-engine example](packages/core/examples/basic.ts) | Minimal API integration example |
|
|
411
537
|
|
|
412
538
|
## Repository layout
|
|
413
539
|
|
|
@@ -421,21 +547,6 @@ docs/ DSH usage, architecture, and evaluation
|
|
|
421
547
|
benchmarks/ Machine-readable experiment results
|
|
422
548
|
```
|
|
423
549
|
|
|
424
|
-
## When StrataGate is a good fit
|
|
425
|
-
|
|
426
|
-
Choose StrataGate when you want several of these properties together:
|
|
427
|
-
|
|
428
|
-
- **automatic cross-session memory** for completed conversations and tool results;
|
|
429
|
-
- **local-first storage** in SQLite, without deploying a separate memory service;
|
|
430
|
-
- **project, session, or global isolation** instead of one undifferentiated memory pool;
|
|
431
|
-
- **layered Events and a knowledge graph** that preserve both what happened and what is currently true;
|
|
432
|
-
- **traceable recall** that can expand a memory back to its original turns and tool output;
|
|
433
|
-
- an **evidence-sufficiency gate** before retrieved memory is treated as enough to answer.
|
|
434
|
-
|
|
435
|
-
Consider a different plugin first when the user's main requirement is free-form visual editing of memory records, hosted multi-user synchronization across products, or a minimal manually maintained notes file. StrataGate includes a read-oriented knowledge-graph view, but it is optimized for automatic, local, evidence-traceable memory rather than collaborative knowledge-base editing.
|
|
436
|
-
|
|
437
|
-
For DeepSeek Harness, follow the [quick start](#quick-start-deepseek-harness). The DSH-specific behavior, tools, configuration, and failure semantics are documented in [`docs/DSH.md`](docs/DSH.md).
|
|
438
|
-
|
|
439
550
|
## Contributing
|
|
440
551
|
|
|
441
552
|
Contributions are welcome—whether you are fixing a bug, improving documentation, adding an integration, or exploring a better memory and retrieval strategy.
|