hippo-memory 1.52.8 → 1.52.9
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +158 -98
- package/dist/api.d.ts +51 -18
- package/dist/api.js +121 -76
- package/dist/audit.d.ts +2 -1
- package/dist/audit.js +63 -0
- package/dist/capture.d.ts +37 -0
- package/dist/capture.js +111 -81
- package/dist/cli.js +627 -654
- package/dist/codex-patch.d.ts +12 -0
- package/dist/codex-patch.js +71 -0
- package/dist/config.d.ts +0 -1
- package/dist/config.js +0 -4
- package/dist/connectors/slack/types.d.ts +0 -1
- package/dist/consolidate.js +85 -32
- package/dist/context-render.d.ts +36 -0
- package/dist/context-render.js +154 -0
- package/dist/dag.js +3 -2
- package/dist/db.js +6 -6
- package/dist/dedupe.d.ts +6 -6
- package/dist/dedupe.js +10 -9
- package/dist/doctor.d.ts +1 -1
- package/dist/doctor.js +35 -2
- package/dist/dormant.d.ts +4 -0
- package/dist/dormant.js +17 -2
- package/dist/embedding-provider.d.ts +2 -1
- package/dist/embedding-provider.js +2 -1
- package/dist/embeddings.js +23 -3
- package/dist/extract.js +5 -1
- package/dist/forward-claim-detector.d.ts +1 -1
- package/dist/forward-claim-detector.js +1 -1
- package/dist/graph-recall.d.ts +3 -1
- package/dist/graph-recall.js +5 -3
- package/dist/hooks.d.ts +15 -1
- package/dist/hooks.js +122 -27
- package/dist/importers.js +5 -12
- package/dist/judgment.d.ts +30 -0
- package/dist/judgment.js +122 -0
- package/dist/mcp/server.js +171 -210
- package/dist/merged-row.d.ts +6 -0
- package/dist/merged-row.js +35 -0
- package/dist/multihop.d.ts +2 -1
- package/dist/multihop.js +7 -4
- package/dist/physics-state.d.ts +0 -4
- package/dist/physics-state.js +0 -6
- package/dist/predictions.d.ts +2 -17
- package/dist/predictions.js +2 -15
- package/dist/reject-flow.d.ts +7 -5
- package/dist/reject-flow.js +41 -12
- package/dist/salience.js +12 -5
- package/dist/same-text.d.ts +17 -0
- package/dist/same-text.js +38 -0
- package/dist/scheduler.d.ts +4 -0
- package/dist/scheduler.js +8 -0
- package/dist/search.d.ts +7 -0
- package/dist/search.js +16 -32
- package/dist/secret-detect.d.ts +2 -0
- package/dist/secret-detect.js +6 -0
- package/dist/server-detect.js +9 -33
- package/dist/server.js +6 -62
- package/dist/session-digest.d.ts +79 -0
- package/dist/session-digest.js +528 -0
- package/dist/shared.d.ts +10 -2
- package/dist/shared.js +35 -30
- package/dist/store.d.ts +1 -0
- package/dist/store.js +4 -0
- package/dist/token-ledger.d.ts +46 -8
- package/dist/token-ledger.js +140 -21
- package/dist/version.d.ts +1 -1
- package/dist/version.js +1 -1
- package/extensions/openclaw-plugin/README.md +4 -4
- package/extensions/openclaw-plugin/openclaw.plugin.json +2 -2
- package/extensions/openclaw-plugin/package.json +1 -1
- package/openclaw.plugin.json +2 -2
- package/package.json +2 -2
package/README.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# 🦛 Hippo: memory for AI agents that learns what is wrong
|
|
2
2
|
|
|
3
|
-
**Hippo learns what is wrong and
|
|
3
|
+
**Hippo learns what is wrong and ranks it down.** Good memory is knowing what to forget: what turned out wrong, what got replaced, what nobody used.
|
|
4
4
|
|
|
5
5
|
[](https://npmjs.com/package/hippo-memory)
|
|
6
6
|
[](https://npmjs.com/package/hippo-memory)
|
|
@@ -9,16 +9,18 @@
|
|
|
9
9
|
[](https://hippo-memory.com)
|
|
10
10
|
|
|
11
11
|
<p align="center">
|
|
12
|
-
<img src="https://raw.githubusercontent.com/kitfunso/hippo-memory/master/assets/hippo-init.svg" alt="hippo init
|
|
12
|
+
<img src="https://raw.githubusercontent.com/kitfunso/hippo-memory/master/assets/hippo-init.svg" alt="hippo init adding memory to one project" width="720">
|
|
13
13
|
</p>
|
|
14
14
|
|
|
15
|
-
|
|
15
|
+
Hippo keeps your coding agents' memories in a SQLite store on your machine, with markdown mirrors you can read and commit. Search is BM25 out of the box, with no model and no network call; embeddings are an optional install. `hippo init` installs hooks for Claude Code and OpenCode, adds 2 hooks to Codex's `hooks.json` when Codex is installed (Codex runs them once you trust them in `/hooks`), and adds instructions to an existing `AGENTS.md` for Codex, Cursor, OpenClaw and Pi. Any MCP client can connect. Mark a memory wrong and it ranks lower; run `hippo supersede` and the old fact leaves recall. Zero runtime deps.
|
|
16
|
+
|
|
17
|
+
Install it, then run `hippo init` inside one project. Init creates the project's `.hippo/` store and adds a block to the `CLAUDE.md` or `AGENTS.md` already there. On your machine it adds hooks for the agents it finds, such as Claude Code's in `~/.claude/settings.json`, and a daily 6:15am run. [What hippo init changes](#what-hippo-init-changes) lists all of it and the flags that skip each part.
|
|
16
18
|
|
|
17
19
|
```bash
|
|
18
|
-
npm install -g hippo-memory && hippo init
|
|
20
|
+
npm install -g hippo-memory && hippo init
|
|
19
21
|
```
|
|
20
22
|
|
|
21
|
-
|
|
23
|
+
Setting up every git repo under a folder in one go is a second step. The [Quick start](#quick-start) says what it changes, then gives the command.
|
|
22
24
|
|
|
23
25
|
Having an AI agent install it? Point it at [llms-install.md](llms-install.md): it installs, wires hippo into the agents it finds, and verifies with `hippo doctor`.
|
|
24
26
|
|
|
@@ -35,11 +37,11 @@ Dependencies: Zero runtime deps. Node.js 22.16+. Optional embeddings: bring-you
|
|
|
35
37
|
|
|
36
38
|
## Why this exists
|
|
37
39
|
|
|
38
|
-
Most "AI memory" systems save everything and search later. That's storage with
|
|
40
|
+
Most "AI memory" systems save everything and search later. That's storage with search on top. A note that turned out wrong ranks the same as one that held up, and an old fact sits beside the one that replaced it.
|
|
39
41
|
|
|
40
|
-
Hippo learns from outcomes. When a recalled memory turns out wrong, mark it bad and it drops out of the top results. Memories you keep using get stronger. Those two are the parts we measured helping ([mechanism audit, round 2](https://github.com/kitfunso/hippo-memory/pull/232)). When a fact changes,
|
|
42
|
+
Hippo learns from outcomes. When a recalled memory turns out wrong, mark it bad and it drops out of the top results. Memories you keep using get stronger. Those two are the parts we measured helping retrieval, on a synthetic test ([mechanism audit, round 2](https://github.com/kitfunso/hippo-memory/pull/232)). When a fact changes, run `hippo supersede` and the old version leaves recall; we have not measured whether that helps. The design borrows from the hippocampus (decay, three layers, sleep consolidation), but that is inspiration. In our tests, decay tied with decay switched off and sleep lowered recall ([Benchmarks](#benchmarks)).
|
|
41
43
|
|
|
42
|
-
It also fixes the portability problem. Your ChatGPT memories don't travel to Claude. Your `.cursorrules` don't travel to Codex. Hippo is one
|
|
44
|
+
It also fixes the portability problem. Your ChatGPT memories don't travel to Claude. Your `.cursorrules` don't travel to Codex. Hippo is one store behind every agent. CLAUDE.md, Cursor rules, ChatGPT exports, Slack history, all in one SQLite store, all queryable from any tool that speaks MCP or HTTP.
|
|
43
45
|
|
|
44
46
|
---
|
|
45
47
|
|
|
@@ -51,19 +53,19 @@ pre-registrations kept next to their results, including the runs that failed and
|
|
|
51
53
|
claim we retracted.
|
|
52
54
|
|
|
53
55
|
- **Sequential Learning Benchmark.** [benchmarks/sequential-learning/](benchmarks/sequential-learning/). 50 tasks, 10 buried traps. Measures whether agents learn from past mistakes, not just retrieve text. v0.11.0 informal magnitude RETRACTED v1.7.9; mechanism remains shipped. See [CHANGELOG.md](./CHANGELOG.md) v1.7.9 entry.
|
|
54
|
-
- **R@5 = 74.0%**
|
|
55
|
-
- **R@1 0.41 to 0.62 with `hippo recall "<query>" --reranker jev
|
|
56
|
-
- **
|
|
57
|
-
- **
|
|
56
|
+
- **LongMemEval oracle split, all 500 questions in one pooled store, BM25 only, no embeddings, v0.11: R@5 = 74.0%** ([benchmarks/README.md](benchmarks/README.md), scripts in [benchmarks/longmemeval/](benchmarks/longmemeval/)). A different setup from the per-haystack results under [Benchmarks](#benchmarks), so the two are not a before and after.
|
|
57
|
+
- **On a private 300-query developer store, R@1 0.41 to 0.62 with `hippo recall "<query>" --reranker jev`**, the free local cross-encoder against Jev ([full eval](docs/evals/2026-09-19-jev-reranker.md)). The opt-in [TypeSafe Jev](https://typesafe.ai) reranker, off by default, about 0.0004 USD a recall. 2000-draw paired bootstrap; the margin held in 20 of 20 seeds and a permutation null reached it in 0 of 200 runs. Ranking only: three graded tests on one 150-question LongMemEval set did **not** show a better answer rate than the free local cross-encoder, and that negative result is in the same doc. What it buys today is a shorter context: on that set, 2 memories ranked by Jev answered as well as 5 ranked by the cross-encoder.
|
|
58
|
+
- **Staged Slack corpus, 10 incident scenarios: recall beat transcript replay in 10 of 10** ([benchmarks/e1.3/](benchmarks/e1.3/)). The answers sit mid-channel by design. In every scenario hippo's top 10 results held all the answer messages, and the channel's last 10 messages held none.
|
|
59
|
+
- **Slack connector, 1000-event ingestion smoke: 0 outbound HTTP** ([benchmarks/e1.3/](benchmarks/e1.3/)). Proven by a `globalThis.fetch` spy that throws on call, not a hardcoded zero. Recall makes no network call by default. One default does: `hippo sleep` sends memory text to Anthropic for fact extraction when `ANTHROPIC_API_KEY` is set. Sleep runs at the end of every Claude Code and OpenCode session and in the daily job, so with the key set the call happens without you asking. `{"extraction":{"enabled":false}}` in `.hippo/config.json` turns that off. Opt-in features such as the Jev reranker above, the LLM reranker and the API embedders also call out.
|
|
58
60
|
- **3,500+ tests on a real database.** No module mocks and no mocked store; only paid network calls are stubbed. Project rule. The one mocks-vs-prod divergence that bit us early is now the constraint that kept the next ten releases honest.
|
|
59
|
-
- **dlPFC goal-conditioned cluster discrimination
|
|
61
|
+
- **3-cluster fixture where BM25 alone cannot discriminate: dlPFC goal-conditioned cluster discrimination passes 3 of 3 queries.** Full goal stack with policy weighting and lifespan-windowed outcome propagation, one query per goal; deterministic test in [`benchmarks/micro/results/b3-depth.json`](benchmarks/micro/results/b3-depth.json).
|
|
60
62
|
|
|
61
63
|
---
|
|
62
64
|
|
|
63
65
|
## What it does for your agent
|
|
64
66
|
|
|
65
|
-
- **
|
|
66
|
-
- **Survives tool switches.** Use Claude Code on Monday, Cursor on Tuesday, Codex on Wednesday.
|
|
67
|
+
- **Keeps errors longer.** Tag a failure with `--tag error` and it gets twice the half-life of an ordinary memory, so the lesson is still in the store the next time a recall matches it. In Claude Code, a hook stores failed tool calls as error memories for you.
|
|
68
|
+
- **Survives tool switches.** Use Claude Code on Monday, Cursor on Tuesday, Codex on Wednesday. They all read the same `.hippo/` store, so the memories come with you.
|
|
67
69
|
- **Ingests systems of record.** Slack and GitHub today (`POST /v1/connectors/slack/events`, `POST /v1/connectors/github/events`). Jira and Notion next. Webhooks land as `kind='raw'` memories with full provenance and GDPR-correct deletion.
|
|
68
70
|
- **Knows where every memory came from.** Every row carries `kind`, `scope`, `owner`, and `artifact_ref`. Right-to-be-forgotten is a single API call, not an audit nightmare.
|
|
69
71
|
- **Plays nice with multi-tenant.** API keys, scrypt-hashed. Audit log on every mutation. Tenant A literally cannot see tenant B's memories. Proven by negative test.
|
|
@@ -72,24 +74,27 @@ claim we retracted.
|
|
|
72
74
|
|
|
73
75
|
## Quick start
|
|
74
76
|
|
|
77
|
+
Start in one project. [What hippo init changes](#what-hippo-init-changes) lists everything the second command writes.
|
|
78
|
+
|
|
75
79
|
```bash
|
|
76
80
|
npm install -g hippo-memory
|
|
77
81
|
|
|
78
|
-
#
|
|
82
|
+
# In a project: create its store and wire in the agents it uses
|
|
79
83
|
hippo init
|
|
84
|
+
```
|
|
80
85
|
|
|
81
|
-
|
|
86
|
+
**Optional: many repos at once.** Read what it changes first. `hippo init --scan <folder>` looks for git repos in the folder and up to three levels below it, skipping dot-folders and `node_modules`. Each repo gets a `.hippo/` store, seeded with lessons from the last 365 days of its commits, and is added to the daily run's list. For the agents it finds, it installs the same user-level hooks as `hippo init`: 7 Claude Code hook entries in `~/.claude/settings.json` and the OpenCode plugin. It also sets up the daily 6:15am run, a crontab line on Linux and macOS or a scheduled task on Windows. It adds no block to any repo's `CLAUDE.md` or `AGENTS.md`. `--no-hooks`, `--no-schedule` and `--no-learn` leave out the hooks, the daily run and the history import.
|
|
87
|
+
|
|
88
|
+
```bash
|
|
82
89
|
hippo init --scan ~
|
|
83
90
|
```
|
|
84
91
|
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
After setup, `hippo sleep` runs at session end (via auto-installed agent hooks) and does five things:
|
|
92
|
+
After setup, `hippo sleep` runs when a Claude Code or OpenCode session ends, and in the daily 6:15am job for every project. Codex runs it at session end only if you installed its wrapper. It does five things:
|
|
88
93
|
|
|
89
94
|
1. **Learns** from today's git commits
|
|
90
95
|
2. **Imports** new entries from the project's Claude Code auto memory
|
|
91
96
|
3. **Consolidates** memories (decay, merge, prune)
|
|
92
|
-
4. **Deduplicates**
|
|
97
|
+
4. **Deduplicates** identical memories, keeping the stronger copy
|
|
93
98
|
5. **Shares** high-value lessons to a global store so they surface in every project
|
|
94
99
|
|
|
95
100
|
```bash
|
|
@@ -103,24 +108,32 @@ hippo recall "data pipeline issues" --budget 2000
|
|
|
103
108
|
Full release history: **[CHANGELOG.md](./CHANGELOG.md)** · [GitHub Releases](https://github.com/kitfunso/hippo-memory/releases)
|
|
104
109
|
|
|
105
110
|
|
|
106
|
-
###
|
|
111
|
+
### What hippo init changes
|
|
112
|
+
|
|
113
|
+
Run `hippo init` inside one project. This is everything it writes, in the project and on your machine:
|
|
107
114
|
|
|
108
|
-
|
|
115
|
+
- **The project's store.** A `.hippo/` folder: SQLite plus markdown mirrors. On the first run in a git repo it learns lessons from the last 30 days of commits.
|
|
116
|
+
- **Instruction files.** A block between `<!-- hippo:start -->` and `<!-- hippo:end -->` in the project's `CLAUDE.md` or `AGENTS.md`, only if that file already exists. Codex, Cursor, OpenClaw, OpenCode and Pi read `AGENTS.md`.
|
|
117
|
+
- **Claude Code,** when the project has `CLAUDE.md` or `.claude/settings.json`: 7 hook entries in `~/.claude/settings.json`, one each on SessionEnd, UserPromptSubmit, PreCompact, PostCompact and PostToolUseFailure and two on SessionStart. [Framework Integrations](#framework-integrations) says what each one runs.
|
|
118
|
+
- **OpenCode,** when the project has `.opencode/` or `opencode.json`: a plugin at `~/.config/opencode/plugins/hippo.ts`.
|
|
119
|
+
- **Codex,** when the project has `AGENTS.md` or `.codex` and Codex is installed (`$CODEX_HOME`, else `~/.codex`, exists): 2 hook entries in Codex's `hooks.json`, one on UserPromptSubmit that sends your pinned memories plus the five most recent ones with every prompt and one on SessionStart after a compaction. **Codex runs them only after you trust them once in `/hooks`.** Init also prints `hippo hook install codex`, the opt-in that wraps the Codex launcher to capture sessions; `hippo hook uninstall codex` removes hippo's hooks and the wrapper.
|
|
120
|
+
- **A daily run at 6:15am,** one per machine: a crontab line on Linux and macOS, a scheduled task named `hippo-daily-runner` on Windows. It runs `hippo learn --git --days 1` and then `hippo sleep` in every project listed in `~/.hippo/workspaces.json`, and init adds this project to that list.
|
|
121
|
+
- **Claude Code auto memory.** On the first run, the notes with YAML front matter in this project's own folder under `~/.claude/projects/` are imported into its store. A file that looks like it holds a secret is skipped.
|
|
109
122
|
|
|
110
123
|
```bash
|
|
111
124
|
cd my-project
|
|
112
125
|
hippo init
|
|
113
126
|
|
|
114
|
-
# Initialized Hippo at /my-project
|
|
127
|
+
# Initialized Hippo at /my-project/.hippo
|
|
115
128
|
# Directories: buffer/ episodic/ semantic/ conflicts/
|
|
129
|
+
# Files: hippo.db stats.json
|
|
116
130
|
# Auto-installed claude-code hook in CLAUDE.md
|
|
131
|
+
# Auto-installed hippo session-end SessionEnd hook in claude-code settings
|
|
132
|
+
# (one line per hook entry)
|
|
133
|
+
# Scheduled machine-level daily runner (6:15am) via crontab
|
|
117
134
|
```
|
|
118
135
|
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
It also registers the current project in Hippo's workspace registry and installs one machine-level daily runner (6:15am). That runner sweeps every registered workspace, runs `hippo learn --git --days 1`, then `hippo sleep`. You get strict daily consolidation without creating one OS task per project.
|
|
122
|
-
|
|
123
|
-
To skip: `hippo init --no-hooks --no-schedule`
|
|
136
|
+
To leave parts out: `--no-hooks` skips the instruction files and hooks, `--no-schedule` the daily run, and `--no-learn` the git history and auto memory import. `HIPPO_SKIP_AUTO_INTEGRATIONS=1` skips the same files and hooks that `--no-hooks` does.
|
|
124
137
|
|
|
125
138
|
---
|
|
126
139
|
|
|
@@ -228,7 +241,7 @@ hippo session resume # re-inject latest handoff as context
|
|
|
228
241
|
|
|
229
242
|
### Working memory
|
|
230
243
|
|
|
231
|
-
Working memory is a bounded scratchpad for current-state notes. It's separate from long-term memory
|
|
244
|
+
Working memory is a bounded scratchpad for current-state notes. It's separate from long-term memory. Entries stay until you run `hippo wm flush`.
|
|
232
245
|
|
|
233
246
|
```bash
|
|
234
247
|
hippo wm push --scope repo \
|
|
@@ -258,7 +271,7 @@ hippo recall "data pipeline" --why --limit 5
|
|
|
258
271
|
|
|
259
272
|
## How It Works
|
|
260
273
|
|
|
261
|
-
Input enters the buffer. Important things get encoded into episodic memory. During "sleep,"
|
|
274
|
+
Input enters the buffer. Important things get encoded into episodic memory. During "sleep," related episodes are merged into one semantic memory, by word overlap. Weak memories decay and disappear.
|
|
262
275
|
|
|
263
276
|
The store is SQLite (`.hippo/hippo.db`). The markdown files are mirrors written after each change. `index.json` is no longer refreshed by writes, deletes or recalls: it is written only when you call `rebuildIndex()` from the package, so a copy an older version left on disk goes stale. Read the store through the CLI, the MCP server or the HTTP API.
|
|
264
277
|
|
|
@@ -266,7 +279,7 @@ The store is SQLite (`.hippo/hippo.db`). The markdown files are mirrors written
|
|
|
266
279
|
flowchart TD
|
|
267
280
|
I[New information] --> B[Buffer<br/>session-only, no decay]
|
|
268
281
|
B -->|encode: tags, strength, half-life| E[Episodic Store<br/>timestamped, decay by default<br/>retrieval strengthens, errors stick]
|
|
269
|
-
E -->|hippo sleep<br/>replay + merge| S[Semantic Store<br/>
|
|
282
|
+
E -->|hippo sleep<br/>replay + merge| S[Semantic Store<br/>merged memories, stable<br/>schema-aware]
|
|
270
283
|
E -.->|decay| X[forgotten]
|
|
271
284
|
S -.->|recall| E
|
|
272
285
|
classDef bio fill:#fff4dc,stroke:#a8742d,color:#2b1b00
|
|
@@ -360,14 +373,16 @@ One-off decisions don't repeat, so they can't earn their keep through retrieval
|
|
|
360
373
|
|
|
361
374
|
```bash
|
|
362
375
|
hippo decide "Use PostgreSQL for all new services" --context "JSONB support"
|
|
363
|
-
# Decision recorded:
|
|
376
|
+
# Decision recorded: #1
|
|
377
|
+
# memory: mem_a1b2c3
|
|
364
378
|
|
|
365
379
|
# Later, when the decision changes:
|
|
366
380
|
hippo decide "Use CockroachDB for global services" \
|
|
367
381
|
--context "Need multi-region" \
|
|
368
382
|
--supersedes mem_a1b2c3
|
|
369
|
-
#
|
|
370
|
-
#
|
|
383
|
+
# Decision recorded: #2
|
|
384
|
+
# memory: mem_d4e5f6
|
|
385
|
+
# supersedes memory: mem_a1b2c3 (decision #1 superseded)
|
|
371
386
|
```
|
|
372
387
|
|
|
373
388
|
---
|
|
@@ -407,7 +422,7 @@ When context is generated, confidence is shown inline:
|
|
|
407
422
|
|
|
408
423
|
Agents can see at a glance what's established fact vs. a pattern worth questioning.
|
|
409
424
|
|
|
410
|
-
|
|
425
|
+
A memory not recalled for 30 days is shown as aged when it is read, and recalling it clears that. Pinned and `verified` memories are exempt. Nothing in the store changes.
|
|
411
426
|
|
|
412
427
|
### Conflict tracking
|
|
413
428
|
|
|
@@ -443,7 +458,7 @@ Three modes: `observe` (default), `suggest`, `assert`. Choose based on how direc
|
|
|
443
458
|
|
|
444
459
|
### Sleep consolidation
|
|
445
460
|
|
|
446
|
-
Run `hippo sleep` and episodes
|
|
461
|
+
Run `hippo sleep` and related episodes merge into one memory.
|
|
447
462
|
|
|
448
463
|
```bash
|
|
449
464
|
hippo sleep
|
|
@@ -457,7 +472,7 @@ hippo sleep
|
|
|
457
472
|
# New semantic: 2
|
|
458
473
|
```
|
|
459
474
|
|
|
460
|
-
|
|
475
|
+
Two or more related episodes get merged into a single semantic memory. The originals decay. The pattern survives.
|
|
461
476
|
|
|
462
477
|
Sleep keeps the store tidy. It has not been shown to improve recall. In round 2 of the mechanism audit, a slept LongMemEval store scored 3.6 points lower at hit@5 than the same store never slept, and no scorer showed sleep helping ([PR #232](https://github.com/kitfunso/hippo-memory/pull/232)).
|
|
463
478
|
|
|
@@ -465,12 +480,12 @@ Sleep keeps the store tidy. It has not been shown to improve recall. In round 2
|
|
|
465
480
|
`{"memoryValue":{"enabled":true}}` in `.hippo/config.json`, sleep consults a learned
|
|
466
481
|
linear memory-value scorer before deleting a decayed memory: a memory that scores in the
|
|
467
482
|
top 30% of its tenant by learned value is kept ("rescued") even though its strength fell
|
|
468
|
-
below the decay threshold. The scorer can only rescue, never delete
|
|
483
|
+
below the decay threshold. The scorer can only rescue, never delete: with the flag on,
|
|
469
484
|
sleep deletes a strict subset of what it would delete with the flag off. Every rescue is
|
|
470
485
|
recorded in the audit log (`hippo audit list --op mv_rescue`). The weights were learned
|
|
471
486
|
on the LongMemEval retention benchmark (held-out retention 0.4897 vs 0.4203 for the best
|
|
472
487
|
hand-set baseline); caveat: their usage-feature signs reflect that benchmark's simulated
|
|
473
|
-
usage, NOT real usage value
|
|
488
|
+
usage, NOT real usage value, so treat the flag as an experiment, not a recommendation.
|
|
474
489
|
Tenants with fewer than 10 non-pinned memories never rescue (rank statistics are noise at
|
|
475
490
|
tiny scale).
|
|
476
491
|
|
|
@@ -497,10 +512,17 @@ removes pinned memories or raw receipts (Slack, GitHub, vault imports) either wa
|
|
|
497
512
|
duplicate removal and junk cleanup still delete.
|
|
498
513
|
|
|
499
514
|
**See what memory costs in tokens.** Every block of memory text hippo hands an agent (the
|
|
500
|
-
per-prompt hook,
|
|
501
|
-
|
|
502
|
-
|
|
503
|
-
|
|
515
|
+
per-prompt hook, the block `hippo compact-resume` restores after compaction, `hippo context`,
|
|
516
|
+
`hippo recall`, the MCP tools, the HTTP API) is recorded in a token ledger: counts, surface
|
|
517
|
+
and session, never the text. A block stays in the conversation, so each later model call
|
|
518
|
+
reads it again until the host compacts. When a Claude Code session ends, hippo counts those
|
|
519
|
+
calls in the session's transcript and records the re-read tokens for the per-prompt hook's
|
|
520
|
+
blocks and the compact-resume block, dated by the day of the calls. The other surfaces show sent tokens only: their rows
|
|
521
|
+
cannot tell a sub-agent's call from its parent's. `hippo tokens` shows sent and re-read
|
|
522
|
+
totals for the last 30 days (`--days`, `--json`). A session that is still open, or that
|
|
523
|
+
crashed, shows what was sent only. Re-reads usually bill at the provider's cached-input rate,
|
|
524
|
+
a fraction of the full input price. Counts are estimates (characters / 4), the same estimate
|
|
525
|
+
every budget uses. Rows older than 90 days are pruned.
|
|
504
526
|
|
|
505
527
|
---
|
|
506
528
|
|
|
@@ -521,7 +543,7 @@ hippo outcome --bad
|
|
|
521
543
|
# reward factor decreases, decay accelerates
|
|
522
544
|
```
|
|
523
545
|
|
|
524
|
-
Outcomes are cumulative. A memory with 5 positive outcomes and 0 negative has a reward factor of ~1.42, making its effective half-life 42% longer. A memory with 0 positive and 3 negative has a factor of
|
|
546
|
+
Outcomes are cumulative. A memory with 5 positive outcomes and 0 negative has a reward factor of ~1.42, making its effective half-life 42% longer. A memory with 0 positive and 3 negative has a factor of 0.625, so it decays 1.6 times as fast; each bad mark past the good ones also halves its strength, up to three times, and recall stops strengthening it. Mixed outcomes converge toward neutral (1.0).
|
|
525
547
|
|
|
526
548
|
This is the mechanism with the clearest measured win. On the synthetic E1 test, plain BM25 plus the outcome nudge cut how often a marked-bad memory stayed in the top five from 71.9% to 0.0% ([mechanism audit, round 2](https://github.com/kitfunso/hippo-memory/pull/232)). Every mark in E1 is correct; real marks are noisier, since `--bad` marks the whole recall batch.
|
|
527
549
|
|
|
@@ -544,6 +566,12 @@ hippo recall "api errors" --budget 1000 --json
|
|
|
544
566
|
|
|
545
567
|
Results are ranked by `relevance * strength * recency`. The highest-signal memories fill the budget first.
|
|
546
568
|
|
|
569
|
+
The budget counts the whole block as printed: the heading, each memory's label, date and tags,
|
|
570
|
+
and any snapshot or hint lines, so the token figure in the heading is the size of what the
|
|
571
|
+
model reads. Recall always keeps its first `--min-results` memories (default 1), even one
|
|
572
|
+
larger than the budget; `hippo context` skips a memory that does not fit and keeps filling.
|
|
573
|
+
`--json` returns the memories the text form would print.
|
|
574
|
+
|
|
547
575
|
---
|
|
548
576
|
|
|
549
577
|
### Auto-learn from git
|
|
@@ -588,7 +616,7 @@ hippo watch "npm run build"
|
|
|
588
616
|
|
|
589
617
|
| Command | What it does |
|
|
590
618
|
|---------|-------------|
|
|
591
|
-
| `hippo init` | Create `.hippo
|
|
619
|
+
| `hippo init` | Create `.hippo/`, install agent hooks and the daily run ([what it changes](#what-hippo-init-changes)) |
|
|
592
620
|
| `hippo init --global` | Create global store at `~/.hippo/` |
|
|
593
621
|
| `hippo init --no-hooks` | Create `.hippo/` without auto-installing hooks |
|
|
594
622
|
| `hippo remember "<text>"` | Store a memory |
|
|
@@ -621,7 +649,7 @@ hippo watch "npm run build"
|
|
|
621
649
|
| `hippo capture --stdin` | Extract memories from piped conversation text |
|
|
622
650
|
| `hippo capture --file <path>` | Extract memories from a file |
|
|
623
651
|
| `hippo capture --dry-run` | Preview extraction without writing |
|
|
624
|
-
| `hippo sleep` | Run consolidation (decay + merge
|
|
652
|
+
| `hippo sleep` | Run consolidation (decay + merge) |
|
|
625
653
|
| `hippo sleep --dry-run` | Preview consolidation without writing |
|
|
626
654
|
| `hippo status` | Memory health: counts, strengths, last sleep |
|
|
627
655
|
| `hippo outcome --good` | Strengthen last recalled memories |
|
|
@@ -634,7 +662,7 @@ hippo watch "npm run build"
|
|
|
634
662
|
| `hippo dormant forget <id>` | Delete a dormant memory permanently |
|
|
635
663
|
| `hippo doctor [--json]` | Check the install: Node, store, schema, sleep, agent hooks; each problem names its fix. It never changes `hippo.db`, though SQLite may leave empty `hippo.db-wal` and `hippo.db-shm` files beside it |
|
|
636
664
|
| `hippo support-bundle [--out <file>] [--include-logs]` | Write a redacted JSON file for a support ticket: versions, doctor checks, config, store counts and log names, never memory text; `--include-logs` adds each log's last 200 lines, which can quote it |
|
|
637
|
-
| `hippo tokens [--days n]` | Estimated tokens of memory text handed to agents, per surface, and what skipping unchanged hook blocks saved |
|
|
665
|
+
| `hippo tokens [--days n]` | Estimated tokens of memory text handed to agents, per surface, what later model calls re-read of the hook and compact-resume blocks, and what skipping unchanged hook blocks saved |
|
|
638
666
|
| `hippo failures [--days n]` | Failed tool calls the capture-error hook saw, by outcome, and how many errors first happened in another session |
|
|
639
667
|
| `hippo embed` | Embed all memories for semantic search |
|
|
640
668
|
| `hippo embed --status` | Show embedding coverage |
|
|
@@ -661,7 +689,7 @@ hippo watch "npm run build"
|
|
|
661
689
|
| `hippo decide "<decision>" --context "<why>"` | Include reasoning |
|
|
662
690
|
| `hippo decide "<decision>" --supersedes <id>` | Supersede a previous decision |
|
|
663
691
|
| `hippo hook list` | Show available framework hooks |
|
|
664
|
-
| `hippo hook install <target>` | Install hook (claude-code also adds
|
|
692
|
+
| `hippo hook install <target>` | Install hook (claude-code also adds its 7 settings.json hook entries: session start and end, each prompt, compaction, failed tool calls) |
|
|
665
693
|
| `hippo hook uninstall <target>` | Remove hook |
|
|
666
694
|
| `hippo handoff create --summary "..."` | Create a session handoff |
|
|
667
695
|
| `hippo handoff latest` | Show the most recent handoff |
|
|
@@ -699,22 +727,22 @@ On `heartbeat`, `block`, `review` and `complete`, a given `--run` is checked aga
|
|
|
699
727
|
|
|
700
728
|
| Framework | Detected by | Patches |
|
|
701
729
|
|-----------|------------|---------|
|
|
702
|
-
| Claude Code | `CLAUDE.md` or `.claude/settings.json` | `CLAUDE.md` +
|
|
703
|
-
| Codex | `AGENTS.md` or `.codex` | `AGENTS.md
|
|
730
|
+
| Claude Code | `CLAUDE.md` or `.claude/settings.json` | `CLAUDE.md` + 7 hook entries in `~/.claude/settings.json` (listed below) |
|
|
731
|
+
| Codex | `AGENTS.md` or `.codex` | `AGENTS.md` + `UserPromptSubmit`/`SessionStart(compact)` hooks in Codex's `hooks.json` when Codex is installed (trust them once in `/hooks`); session capture is opt-in with `hippo hook install codex`, which wraps the Codex launcher |
|
|
704
732
|
| Cursor | `AGENTS.md` | `AGENTS.md`, which Cursor reads from the project root |
|
|
705
|
-
| OpenClaw | `.openclaw` or `AGENTS.md` | native
|
|
733
|
+
| OpenClaw | `.openclaw` or `AGENTS.md` | `AGENTS.md`; the native plugin is a separate install: `openclaw plugins install hippo-memory` |
|
|
706
734
|
| OpenCode | `.opencode/` or `opencode.json` | `AGENTS.md` + TS plugin at `~/.config/opencode/plugins/hippo.ts` (subscribes to `session.idle` + `session.created`) |
|
|
707
735
|
| Pi | `.pi` or `.pi/agent` | `AGENTS.md`; copy the [Pi extension](https://github.com/kitfunso/hippo-memory/tree/master/extensions/pi-extension) for session hooks |
|
|
708
736
|
|
|
709
|
-
|
|
737
|
+
Init patches an instruction file only if it already exists. It also sets up a daily run and imports the project's Claude Code auto memory; [What hippo init changes](#what-hippo-init-changes) lists everything.
|
|
710
738
|
|
|
711
739
|
### Manual install
|
|
712
740
|
|
|
713
741
|
If you prefer explicit control:
|
|
714
742
|
|
|
715
743
|
```bash
|
|
716
|
-
hippo hook install claude-code # patches CLAUDE.md + adds
|
|
717
|
-
hippo hook install codex #
|
|
744
|
+
hippo hook install claude-code # patches CLAUDE.md + adds the 7 settings.json hook entries listed below
|
|
745
|
+
hippo hook install codex # patches AGENTS.md + adds hooks to Codex's hooks.json + wraps the detected Codex launcher
|
|
718
746
|
hippo hook install cursor # patches AGENTS.md
|
|
719
747
|
hippo hook install openclaw # patches AGENTS.md
|
|
720
748
|
hippo hook install opencode # patches AGENTS.md + installs the opencode TS plugin
|
|
@@ -723,21 +751,28 @@ hippo hook install opencode # patches AGENTS.md + installs the opencode TS
|
|
|
723
751
|
This adds a `<!-- hippo:start -->` ... `<!-- hippo:end -->` block that tells the agent to:
|
|
724
752
|
1. Run `hippo context --auto --budget 1500` at session start
|
|
725
753
|
2. Run `hippo remember "<what went wrong and why>" --error` the moment it finds out why something failed, never as a closing step
|
|
726
|
-
3.
|
|
754
|
+
3. Everywhere but Claude Code, whose own auto memory does this job: run a plain `hippo remember` the moment it learns something that should outlive the session, leaving out secrets and personal details
|
|
755
|
+
4. Capture a short summary with `hippo capture --stdin` when the session ends, but only where no hook captures the session: Cursor, OpenClaw, OpenCode, Pi, and Codex without its wrapper
|
|
727
756
|
|
|
728
757
|
The block asks for nothing a hook already does, because each extra tool call re-reads the whole context. Re-running `hippo init` swaps a block an older hippo wrote for the current one, as long as nobody edited it. It leaves an edited block alone and says so, and never touches text outside the markers.
|
|
729
758
|
|
|
730
|
-
For Claude Code, it also adds
|
|
731
|
-
- a `SessionEnd` hook that runs `hippo sleep` and then `hippo capture`
|
|
759
|
+
For Claude Code, it also adds 7 hook entries to `~/.claude/settings.json`:
|
|
760
|
+
- a `SessionEnd` hook that runs `hippo sleep` and then `hippo capture` when the session exits. Capture matches the last 20 user and 10 assistant messages of the transcript against word patterns for decisions, rules, errors and preferences. It uses no model and does not read earlier turns, so record lessons with `hippo remember` as you go.
|
|
732
761
|
- a `SessionStart` hook that prints the previous session's consolidation output
|
|
733
|
-
- a `UserPromptSubmit` hook that runs `hippo context --pinned-only --include-recent 5 --format additional-context` every turn. It re-injects pinned memories (`hippo remember <text> --pin`) plus the
|
|
762
|
+
- a `UserPromptSubmit` hook that runs `hippo context --pinned-only --include-recent 5 --format additional-context` every turn. It re-injects pinned memories (`hippo remember <text> --pin`) plus the 5 newest memories in the store, so fresh same-session lessons appear on the next prompt before you pin them. It does not read your prompt unless you set `{"pinnedInject":{"promptRecall":true}}`, which swaps the 5 newest for memories that match the prompt. The block is rendered without live strength percentages, so it stays byte-identical while its memories do not change, and it is sent only when it changed since the session's last prompt: an unchanged block is skipped, resent every 10 skips (`pinnedInject.refreshTurns`, `0` never resends) and resent after compaction. `{"pinnedInject":{"skipUnchanged":false}}` sends it every turn as before. Opt out entirely with `{"pinnedInject":{"enabled":false}}` in `.hippo/config.json`.
|
|
734
763
|
- a `PreCompact` hook that runs `hippo pre-compact` before the transcript gets summarized. It saves a working-state snapshot (task/summary/next step) so mid-session compaction can't drop it; the `SessionEnd` hook still owns extracting durable memories.
|
|
735
|
-
- a second `SessionStart` hook (matcher `compact`) that runs `hippo compact-resume`, printing that snapshot
|
|
764
|
+
- a second `SessionStart` hook (matcher `compact`) that runs `hippo compact-resume`, printing that snapshot back into context right after compaction, if it is under 15 minutes old.
|
|
736
765
|
- a `PostCompact` hook that runs `hippo post-compact`, which tells you what was saved ("Hippo saved your task snapshot before compacting."). It prints nothing when nothing was saved.
|
|
737
766
|
- a `PostToolUseFailure` hook that runs `hippo capture-error`, which stores a failed tool call as an error memory. It skips interrupts, declined permissions and searches that found nothing, and stores a repeated failure once. It also logs every failure, stored or not, for `hippo failures`: the session, the tool and hashes of the error, never its text. A hash is not anonymous, since anyone who guesses an error's text can check it against the hash. The log keeps 90 days.
|
|
738
767
|
|
|
739
768
|
To remove: `hippo hook uninstall claude-code`
|
|
740
769
|
|
|
770
|
+
For Codex, it adds two hooks to `$CODEX_HOME/hooks.json` (else `~/.codex/hooks.json`) and keeps every hook already there:
|
|
771
|
+
- a `UserPromptSubmit` hook that runs the same `hippo context --pinned-only` command as Claude Code's, so your pinned memories plus the five most recent ones reach every prompt as developer context
|
|
772
|
+
- a `SessionStart` hook (matcher `compact`) that runs `hippo compact-resume` after a compaction, so the next prompt sends that block again. Codex gets no `PreCompact` hook from hippo, so it restores a task snapshot only if one was saved with `hippo snapshot save` in the last 15 minutes
|
|
773
|
+
|
|
774
|
+
**Codex runs a new or changed hook only after you trust it, so open `/hooks` in Codex once and trust both;** `hippo doctor` reminds you. The per-prompt hook was checked against a real Codex request; the compaction hook follows Codex's documented `compact` start source and has not been watched end to end in Codex. Each hook also carries a `commandWindows` form (`hippo.cmd ...`), because Codex runs hooks through PowerShell on Windows, where the execution policy can block npm's `hippo.ps1`. hippo only ever appends these two entries and never rewrites one, since Codex treats a changed command as a new hook to trust. To remove: `hippo hook uninstall codex`, which takes out only hippo's exact commands and leaves every other hook, including one of yours that runs hippo.
|
|
775
|
+
|
|
741
776
|
### What the hook adds (Claude Code example)
|
|
742
777
|
|
|
743
778
|
````markdown
|
|
@@ -819,58 +854,75 @@ Hippo's design borrows seven properties of the human hippocampus. This section i
|
|
|
819
854
|
|
|
820
855
|
**Why two stores?** The brain uses a fast hippocampal buffer + a slow neocortical store (Complementary Learning Systems theory, McClelland et al. 1995). If the neocortex learned fast, new information would overwrite old knowledge. The buffer absorbs new episodes; the neocortex extracts patterns over time.
|
|
821
856
|
|
|
822
|
-
**Why decay at all?** In the brain, new neurons born in the dentate gyrus disrupt old memory traces (Frankland et al. 2013), which may reduce interference from outdated information. That is why hippo has decay. In hippo's own tests, age-based decay made no measurable difference to recall: at the 365-day default it tied with decay switched off.
|
|
857
|
+
**Why decay at all?** In the brain, new neurons born in the dentate gyrus disrupt old memory traces (Frankland et al. 2013), which may reduce interference from outdated information. That is why hippo has decay. In hippo's own tests, age-based decay made no measurable difference to recall: at the 365-day default it tied with decay switched off. A bad outcome mark is the forgetting that measured helpful. Supersession, a newer fact replacing an old one, has not been measured.
|
|
823
858
|
|
|
824
859
|
**Why do errors stick?** The amygdala modulates hippocampal consolidation based on emotional significance. Fear and error signals boost encoding. Your first production incident is burned into memory. Your 200th uneventful deploy isn't.
|
|
825
860
|
|
|
826
|
-
**Why does retrieval strengthen?** Recalled memories undergo "reconsolidation" (Nader et al. 2000). The act of retrieval destabilizes the trace, then re-encodes it stronger. This is the testing effect. Hippo
|
|
861
|
+
**Why does retrieval strengthen?** Recalled memories undergo "reconsolidation" (Nader et al. 2000). The act of retrieval destabilizes the trace, then re-encodes it stronger. This is the testing effect. Hippo borrows the idea: each recall adds 2 days to a memory's half-life.
|
|
827
862
|
|
|
828
|
-
**Why does sleep consolidate?** During sleep, the hippocampus replays compressed versions of recent episodes and "teaches" the neocortex by repeatedly activating the same patterns. Hippo's `sleep` command
|
|
863
|
+
**Why does sleep consolidate?** During sleep, the hippocampus replays compressed versions of recent episodes and "teaches" the neocortex by repeatedly activating the same patterns. Hippo's `sleep` command borrows the idea for a consolidation pass. In hippo's own audit it lowered recall ([Sleep consolidation](#sleep-consolidation) has the numbers).
|
|
829
864
|
|
|
830
865
|
The 7 mechanisms in full: [PLAN.md#core-principles](PLAN.md#core-principles)
|
|
831
866
|
|
|
832
867
|
For how these mechanisms connect to LLM training, continual learning, and open research problems: **[RESEARCH.md](RESEARCH.md)**
|
|
833
868
|
|
|
834
|
-
**Why does reward modulate decay?** In spiking neural networks, reward-modulated STDP strengthens synapses that contribute to positive outcomes and weakens those that don't. Hippo's reward-proportional decay (v0.11.0)
|
|
869
|
+
**Why does reward modulate decay?** In spiking neural networks, reward-modulated STDP strengthens synapses that contribute to positive outcomes and weakens those that don't. Hippo's reward-proportional decay (v0.11.0) borrows this idea: memories with consistent positive outcomes decay slower, negatives decay faster, with no fixed deltas. Inspired by [MH-FLOCKE](https://github.com/MarcHesse/mhflocke)'s R-STDP architecture for quadruped locomotion, where the same mechanism produces stable learning with 11.6x lower variance than PPO.
|
|
835
870
|
|
|
836
|
-
**Prior art in agent memory simulation.** The idea that human-like memory produces human-like behavior as an emergent property was explored in IEEE research from 2010-2011 ([5952114](https://ieeexplore.ieee.org/document/5952114), [5548405](https://ieeexplore.ieee.org/document/5548405), [5953964](https://ieeexplore.ieee.org/document/5953964)). Walking between rooms and forgetting why you went there doesn't need direct simulation; it emerges naturally from a memory system with capacity limits and decay. Hippo
|
|
871
|
+
**Prior art in agent memory simulation.** The idea that human-like memory produces human-like behavior as an emergent property was explored in IEEE research from 2010-2011 ([5952114](https://ieeexplore.ieee.org/document/5952114), [5548405](https://ieeexplore.ieee.org/document/5548405), [5953964](https://ieeexplore.ieee.org/document/5953964)). Walking between rooms and forgetting why you went there doesn't need direct simulation; it emerges naturally from a memory system with capacity limits and decay. Hippo takes the mechanisms as design ideas. Whether better agent behavior follows from them has not been shown.
|
|
837
872
|
|
|
838
|
-
**Related work:** [HippoRAG](https://arxiv.org/abs/2405.14831) (Gutierrez et al., 2024) applies hippocampal indexing to RAG via knowledge graphs. [MemPalace](https://github.com/milla-jovovich/mempalace) (Sigman & Jovovich, 2026) organizes memory spatially (wings/halls/rooms) with AAAK compression, achieving 100% on [LongMemEval](https://arxiv.org/abs/2410.10813). [MH-FLOCKE](https://github.com/MarcHesse/mhflocke) (Hesse, 2026) uses spiking neurons with R-STDP for embodied cognition. Each system tackles a different facet: HippoRAG optimizes retrieval quality, MemPalace optimizes retrieval organization, MH-FLOCKE optimizes embodied learning, and Hippo
|
|
873
|
+
**Related work:** [HippoRAG](https://arxiv.org/abs/2405.14831) (Gutierrez et al., 2024) applies hippocampal indexing to RAG via knowledge graphs. [MemPalace](https://github.com/milla-jovovich/mempalace) (Sigman & Jovovich, 2026) organizes memory spatially (wings/halls/rooms) with AAAK compression, achieving 100% on [LongMemEval](https://arxiv.org/abs/2410.10813). [MH-FLOCKE](https://github.com/MarcHesse/mhflocke) (Hesse, 2026) uses spiking neurons with R-STDP for embodied cognition. Each system tackles a different facet: HippoRAG optimizes retrieval quality, MemPalace optimizes retrieval organization, MH-FLOCKE optimizes embodied learning, and Hippo works on the memory lifecycle.
|
|
839
874
|
|
|
840
875
|
---
|
|
841
876
|
|
|
842
877
|
## Comparison
|
|
843
878
|
|
|
844
|
-
The
|
|
879
|
+
Two tables. The first is plain facts: where the data lives, what each tool needs and runs on, and what it has published. Several rows favour the other tools: a managed multi-user service, a native Python library, a full knowledge graph. The second lists design bets, choices a tool made rather than results it measured; what hippo has measured is under [Benchmarks](#benchmarks). Graph-first systems ([gbrain](https://hermesatlas.com/projects/garrytan/gbrain), [Zep](https://www.getzep.com/), [Cognee](https://www.cognee.ai/)), agent-managed systems ([Letta](https://github.com/letta-ai/letta-code)), and version-control or skill-distillation takes ([Memoria](https://github.com/matrixorigin/Memoria), [EverMind](https://evermind.ai/)) solve adjacent problems with different mechanics.
|
|
880
|
+
|
|
881
|
+
### What each tool is
|
|
882
|
+
|
|
883
|
+
The rows from where the data lives to the graph were checked against each tool's own pages on 2026-09-28.
|
|
845
884
|
|
|
846
885
|
| Feature | Hippo | [MemPalace](https://github.com/milla-jovovich/mempalace) | [Mem0](https://github.com/mem0ai/mem0) | [Basic Memory](https://github.com/basicmachines-co/basic-memory) | [gbrain](https://hermesatlas.com/projects/garrytan/gbrain) | [Zep](https://www.getzep.com/) | [Letta](https://github.com/letta-ai/letta-code) | [Cognee](https://www.cognee.ai/) | [Memoria](https://github.com/matrixorigin/Memoria) | [EverMind](https://evermind.ai/) |
|
|
847
886
|
|---------|-------|-----------|------|-------------|--------|-----|-------|--------|---------|----------|
|
|
887
|
+
| Where your data lives | Your machine or your server (SQLite) | Your machine (ChromaDB by default) | Where you run it, or Mem0's cloud | Your machine (Markdown files); cloud optional | Your machine (PGLite) or your Postgres | Zep's cloud (your own cloud on Enterprise) | Your machine; cloud backup with /login | Your machine by default; Cognee Cloud optional | Memoria Cloud, or self-hosted (Docker or embedded) | Your machine by default; EverOS Cloud optional |
|
|
888
|
+
| Managed multi-user service | No (self-hosted, with tenants and API keys; the commercial edition adds hosted SaaS) | ? | Yes (hosted platform) | Yes (Teams) | ? (self-hosted server with OAuth) | Yes (Zep Cloud) | ? (cloud backup with /login) | Yes (Cognee Cloud) | ? (Memoria Cloud) | Yes (EverOS Cloud) |
|
|
889
|
+
| Needs an account or model key | No | No (core path) | Yes (model key; account for the platform) | No (account for the cloud) | No (keyless mode) | Yes (Zep account; Graphiti needs a model key) | Yes (your own model keys) | No (local) | ? (account for Memoria Cloud) | Yes (an LLM key, OpenRouter) |
|
|
890
|
+
| License | MIT | MIT | Apache-2.0 | AGPL-3.0 | MIT | Proprietary cloud (Graphiti: Apache-2.0) | Apache-2.0 | Apache-2.0 | Apache-2.0 | Apache-2.0 (EverOS) + cloud |
|
|
891
|
+
| Runtime | Node.js 22.16+, no runtime deps | Python 3.9+ (ChromaDB by default) | Python or Node.js (server: Postgres + pgvector) | Python (SQLite by default) | Bun (PGLite or Postgres + pgvector) | Managed service (Graphiti: Python + a graph database) | Node.js (npm) | Python (graph and vector stores, or Postgres) | A CLI binary + a MatrixOne database | Python (SQLite + LanceDB) |
|
|
892
|
+
| SDKs and APIs | CLI, HTTP API, Python SDK over HTTP | Python library, CLI | Python and Node.js libraries, a self-hosted server | CLI, cloud app | CLI, HTTP API | Python, TypeScript and Go SDKs | TypeScript SDK, CLI | Python and TypeScript SDKs, REST API, CLI | Python client, REST API, CLI | Python library, HTTP API, CLI |
|
|
893
|
+
| Native Python library | No (the Python SDK calls hippo serve over HTTP) | Yes (mempalace) | Yes (mem0ai) | Yes (basic-memory) | No (installs with Bun) | Yes (Graphiti: graphiti-core) | ? (Letta Code ships on npm) | Yes (cognee) | ? (memoria-client) | Yes (everos) |
|
|
894
|
+
| Graph or entity relations | Partial (entities from decisions and policies; recall --hops, off by default) | Yes (temporal entity graph) | Partial (entity linking; graph memory on Pro) | Yes (wikilinks and observations) | Yes (typed knowledge graph) | Yes (temporal knowledge graph) | No | Yes (knowledge graph) | No (typed claims) | ? (graph view in progress) |
|
|
895
|
+
| Hybrid search (BM25 + embeddings) | Yes (BM25 by default; embeddings are an optional install) | Embeddings + spatial | Yes (semantic + BM25 + entity) | No | Yes (vec + rerank + graph) | Yes (graph + vec) | ? | Yes (GraphRAG) | Yes (vector + full-text) | Yes (mRAG, multi-modal) |
|
|
896
|
+
| MCP server | Yes | Yes | Yes (hosted, needs an account) | Yes | Yes (stdio + HTTP/OAuth) | Yes (hosted, needs an account) | Yes (hosted, needs an API key) | Yes (first-party Claude/LangGraph) | Yes | ? |
|
|
897
|
+
| Multi-agent shared memory | Yes | No | No | No | Yes (brain repo, team mounts) | Yes | Yes (shared memory blocks) | Yes | Yes (branch/merge across sessions) | Yes (multi-agent coordination) |
|
|
898
|
+
| Auto-hook install | Yes (Claude Code hooks, OpenCode plugin) | No | No | No | No | No | No | No | No | No |
|
|
899
|
+
| Cross-tool import (ChatGPT/Claude/Cursor) | Yes | No | No | No | Partial (data sources) | ? | No | Partial (28 data sources) | No (Git ops) | Partial (mRAG: PDFs/images/URLs) |
|
|
900
|
+
| Git-friendly | Yes | No | No | Yes | Yes | No | Yes (memory tracked in git) | No | Yes (Git is the model) | ? |
|
|
901
|
+
| Framework agnostic | Yes | Yes | Partial | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
|
|
902
|
+
| LongMemEval (published) | 85.6% R@5 from hippo recall, 96.8% with the budget lifted (4,000-token default budget; the benchmark scripts' best of five settings: 98.0% local / 99.8% voyage any-evidence, 88.5% local all-evidence; s_cleaned, per-haystack)\* | 96.6% raw / 100% reranked R@5 | 94.4 (hosted platform)\*\* | N/A | 95.53% all-evidence R@5 reranked, 93.19% without (s_cleaned\*) | 90.2% accuracy\*\* (LoCoMo 94.7%) | N/A | N/A | 88.78% overall accuracy w/ reader\*\* | 83.00% overall\*\* (LoCoMo 93.05%, HaluMem 93.04%) |
|
|
903
|
+
|
|
904
|
+
\* Hippo's figures are on `longmemeval_s_cleaned` with a per-question haystack. On a default install, `hippo recall` puts an answer session in its top 5 for 85.6% of the 500 questions inside its default 4,000-token budget. With the budget lifted it scores 96.8%, but then it returns every candidate, so that figure measures the ranking ([result](docs/evals/2026-09-28-recall-cli-longmemeval-result.md)). The 98.0% and 99.8% are the benchmark scripts' retrieval, each the best of five settings, not `hippo recall`. Any-evidence R@5 counts a hit when any answer session is in the top 5, over all 500 questions: 98.0% with the free local MiniLM embedder (an optional install) and 99.8% with voyage-3-large (measured 2026-06-09, not re-run). All-evidence R@5 counts a hit only when every answer session is in the top 5, over the 470 questions that have an answer: 86.8 to 88.5% with MiniLM. gbrain first published 97.6%, an any-evidence score over all 500; its [report](https://github.com/garrytan/gbrain-evals/blob/main/docs/benchmarks/2026-05-07-longmemeval-s.md) now leads with all-evidence, 95.53% (449 of 470) with the paid Voyage rerank-2.5 reranker and 93.19% without it. On all-evidence recall gbrain is ahead. The June 2026 build scored 98.6 any-evidence; [`docs/evals/2026-09-23-longmemeval-reproduction.md`](docs/evals/2026-09-23-longmemeval-reproduction.md) has both runs. An older hippo number, 86.8% R@5 on `longmemeval_oracle` under pooled (non-per-haystack) retrieval, is not comparable to per-haystack figures.
|
|
905
|
+
|
|
906
|
+
\*\* Different metric: these are end-to-end answer scores, not retrieval R@5. Mem0's 94.4 comes from its hosted platform, which its README says includes optimizations the open-source SDK lacks. Zep's 90.2% and 94.7% are accuracy figures from its homepage. Memoria's 88.78% and EverMind's 83% are overall accuracy with a reader LLM. Higher denominator + LLM helps. Not directly comparable to retrieval-only R@5 numbers above. The Mem0, Zep and Letta columns were last checked against each vendor's own pages on 2026-09-28.
|
|
907
|
+
|
|
908
|
+
### Design bets
|
|
909
|
+
|
|
910
|
+
A Yes means the tool made that choice, not that the choice was shown to help. What hippo has measured about its own bets is under [Benchmarks](#benchmarks); decay, for one, tied with decay switched off. Spatial organization and lossless compression are MemPalace's bets.
|
|
911
|
+
|
|
912
|
+
| Design bet | Hippo | [MemPalace](https://github.com/milla-jovovich/mempalace) | [Mem0](https://github.com/mem0ai/mem0) | [Basic Memory](https://github.com/basicmachines-co/basic-memory) | [gbrain](https://hermesatlas.com/projects/garrytan/gbrain) | [Zep](https://www.getzep.com/) | [Letta](https://github.com/letta-ai/letta-code) | [Cognee](https://www.cognee.ai/) | [Memoria](https://github.com/matrixorigin/Memoria) | [EverMind](https://evermind.ai/) |
|
|
913
|
+
|---------|-------|-----------|------|-------------|--------|-----|-------|--------|---------|----------|
|
|
848
914
|
| Decay by default | Yes | No | No | No | No | No | No | No | No | No |
|
|
849
915
|
| Retrieval strengthening | Yes | No | No | No | No | No | No | Partial (recall tuning) | No | Partial (Skill Memory distills patterns) |
|
|
850
916
|
| Reward-proportional decay | Yes | No | No | No | No | No | No | No | No | No |
|
|
851
|
-
|
|
|
852
|
-
| Schema acceleration / knowledge graph | Yes (schema) | No | Partial (entity linking; graph memory on Pro) | No | Yes (typed KG, self-wiring) | Yes (temporal KG) | No | Yes (auto-ontologies) | No (typed claims) | Yes (hierarchical: user/group/agent) |
|
|
917
|
+
| Outcome tracking | Yes | No | No | No | No | No | No | No | No | Partial (Cases: agent trajectories) |
|
|
853
918
|
| Conflict detection + resolution | Yes | No | Partial (hosted platform marks superseded facts) | No | Yes (eval-surfaced) | Yes (auto-invalidate stale facts) | No | No | Yes (auto-detect + quarantine) | Partial (temporal tracking) |
|
|
854
|
-
| Multi-agent shared memory | Yes | No | No | No | Yes (brain repo, team mounts) | Yes | Yes (shared memory blocks) | Yes | Yes (branch/merge across sessions) | Yes (multi-agent coordination) |
|
|
855
919
|
| Transfer scoring | Yes | No | No | No | No | No | No | No | No | No |
|
|
856
|
-
| Outcome tracking | Yes | No | No | No | No | No | No | No | No | Partial (Cases: agent trajectories) |
|
|
857
920
|
| Confidence tiers | Yes | No | No | No | No (typed facts) | No | No | No | No | No |
|
|
921
|
+
| Schema acceleration | Yes | No | No | No | No | No | No | No | No | No |
|
|
858
922
|
| Spatial organization | No | Yes (wings/halls/rooms) | No | No | No | No | No | No | No | No |
|
|
859
923
|
| Lossless compression | No | Yes (AAAK, 30x) | No | No | No | No | No | No | No | No |
|
|
860
|
-
| Cross-tool import (ChatGPT/Claude/Cursor) | Yes | No | No | No | Partial (data sources) | ? | No | Partial (28 data sources) | No (Git ops) | Partial (mRAG: PDFs/images/URLs) |
|
|
861
|
-
| Auto-hook install | Yes | No | No | No | No | No | No | No | No | No |
|
|
862
|
-
| MCP server | Yes | Yes | Yes (hosted, needs an account) | Yes | Yes (stdio + HTTP/OAuth) | Yes (hosted, needs an account) | Yes (hosted, needs an API key) | Yes (first-party Claude/LangGraph) | Yes | ? |
|
|
863
|
-
| Zero runtime deps | Yes | No (ChromaDB) | No | No | No (PGLite or PG+pgvector) | No (managed service) | No (npm deps) | No (Python deps) | Yes (single Rust binary) | No (managed + OSS) |
|
|
864
|
-
| LongMemEval (best published) | 98.0% local / 99.8% voyage any-evidence R@5; 88.5% local all-evidence R@5 (s_cleaned, per-haystack)\* | 96.6% raw / 100% reranked R@5 | 94.4 (hosted platform)\*\* | N/A | 95.53% all-evidence R@5 reranked, 93.19% without (s_cleaned\*) | 90.2% accuracy\*\* (LoCoMo 94.7%) | N/A | N/A | 88.78% overall accuracy w/ reader\*\* | 83.00% overall\*\* (LoCoMo 93.05%, HaluMem 93.04%) |
|
|
865
|
-
| Git-friendly | Yes | No | No | Yes | Yes | No | Yes (memory tracked in git) | No | Yes (Git is the model) | ? |
|
|
866
|
-
| Framework agnostic | Yes | Yes | Partial | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
|
|
867
|
-
| License | MIT | (open) | Apache-2.0 | (open) | MIT | Proprietary cloud (Graphiti: Apache-2.0) | Apache-2.0 | MIT (core) | Apache-2.0 | Apache-2.0 (OSS) + cloud |
|
|
868
924
|
|
|
869
|
-
|
|
870
|
-
|
|
871
|
-
\*\* Different metric: these are end-to-end answer scores, not retrieval R@5. Mem0's 94.4 comes from its hosted platform, which its README says includes optimizations the open-source SDK lacks. Zep's 90.2% and 94.7% are accuracy figures from its homepage. Memoria's 88.78% and EverMind's 83% are overall accuracy with a reader LLM. Higher denominator + LLM helps. Not directly comparable to retrieval-only R@5 numbers above. The Mem0, Zep and Letta columns were last checked against each vendor's own pages on 2026-09-28.
|
|
872
|
-
|
|
873
|
-
Different tools answer different questions. Mem0 and Basic Memory implement "save everything, search later." MemPalace implements "store everything, organize spatially for retrieval." gbrain, Zep, and Cognee implement "extract typed entities and relationships into a knowledge graph." Letta implements "the agent edits its own memory blocks." Memoria implements "Git-style version control over the memory state itself." EverMind implements "self-evolving Skill Memory + multi-modal retrieval over hierarchical scopes." Hippo implements "learn what is wrong and stop repeating it." These are complementary takes, not a single-axis ranking: bio-lifecycle (Hippo) + GraphRAG (gbrain/Cognee/Zep) + agent-self-edit (Letta) + memory-VCS (Memoria) + skill-distillation (EverMind) cover different parts of the same problem.
|
|
925
|
+
Different tools answer different questions. Mem0 and Basic Memory implement "save everything, search later." MemPalace implements "store everything, organize spatially for retrieval." gbrain, Zep, and Cognee implement "extract typed entities and relationships into a knowledge graph." Letta implements "the agent edits its own memory blocks." Memoria implements "Git-style version control over the memory state itself." EverMind implements "self-evolving Skill Memory + multi-modal retrieval over hierarchical scopes." Hippo implements "learn what is wrong and rank it down." These are complementary takes, not a single-axis ranking: memory lifecycle (Hippo) + GraphRAG (gbrain/Cognee/Zep) + agent-self-edit (Letta) + memory-VCS (Memoria) + skill-distillation (EverMind) cover different parts of the same problem.
|
|
874
926
|
|
|
875
927
|
---
|
|
876
928
|
|
|
@@ -882,16 +934,24 @@ Three benchmarks testing three different things. Full details in [`benchmarks/`]
|
|
|
882
934
|
|
|
883
935
|
[LongMemEval](https://arxiv.org/abs/2410.10813) (ICLR 2025) is the industry-standard benchmark: 500 questions across 5 memory abilities, embedded in 115k+ token chat histories.
|
|
884
936
|
|
|
885
|
-
|
|
937
|
+
**`hippo recall` (measured 2026-09-28 on hippo-memory 1.52.5).** On a default install, `hippo recall` puts an answer session in its top 5 for 85.6% of the 500 `_s` questions (95% CI 82.4 to 88.6), and 87.6% with the optional MiniLM embedder (84.6 to 90.4). Most of the gap to the scripts below is the default 4,000-token budget. Each memory in this test is a whole session, about 2,600 tokens at the median, so recall returns a median of 2 sessions. With the budget lifted, the same rankings score 96.8% (95.2 to 98.2) and 97.4% (96.0 to 98.6). A lifted call returns every candidate, a median of 47 sessions and 123,491 tokens per question, so those two figures measure the ranking, not an amount of text an agent could take in. Result: [`docs/evals/2026-09-28-recall-cli-longmemeval-result.md`](docs/evals/2026-09-28-recall-cli-longmemeval-result.md).
|
|
938
|
+
|
|
939
|
+
**The benchmark scripts (`_s` split; MiniLM re-measured 2026-09-23, voyage measured 2026-06-09).** Each question is scored against its own ~48-session haystack, the same way gbrain and other published systems report. The v1.23.0 pluggable embedding provider lets you choose the embedder:
|
|
886
940
|
|
|
887
|
-
| Embedder | Dense-only R@5 | Best
|
|
941
|
+
| Embedder | Dense-only R@5 | Best of five settings, R@5 | R@1 |
|
|
888
942
|
|----------|----------------|-----------------|-----|
|
|
889
943
|
| MiniLM-L6 (local, optional install) | 96.8 | 98.0 | 88.4 |
|
|
890
944
|
| voyage-3-large (opt-in, paid) | 99.8 | 99.8 | 94.6 |
|
|
891
945
|
|
|
892
946
|
These are any-evidence scores: a hit when any answer session is in the top 5, over all 500 questions. Counting a hit only when every answer session is in the top 5, over the 470 questions that have an answer, the MiniLM runs score 86.8 to 88.5. gbrain first published 97.6 any-evidence and now reports 95.53 all-evidence with a paid reranker, 93.19 without, so on all-evidence recall gbrain is ahead. These numbers come from the scripts in `benchmarks/longmemeval/`, which index every turn and fuse BM25 with dense ranks; they are not `hippo recall`, and a default install has no embedder. Re-measure: [`docs/evals/2026-09-23-longmemeval-reproduction.md`](docs/evals/2026-09-23-longmemeval-reproduction.md). Any-evidence recall is near its ceiling on this task; all-evidence recall is not. Method and the global-pool comparison: [`docs/evals/2026-06-09-longmemeval-per-haystack-dual.md`](docs/evals/2026-06-09-longmemeval-per-haystack-dual.md).
|
|
893
947
|
|
|
894
|
-
The differentiator is what happens as one store grows. Point retrieval at a single unified memory of tens of thousands of sessions, with no pre-scoped haystack, and recall stops being free (MiniLM 47, voyage 56 on the 19,195-session `_s` store, June 2026). That is where we expect the memory lifecycle to matter, and it is what hippo measures next (see ROADMAP Part III). It is not shown yet
|
|
948
|
+
The differentiator is what happens as one store grows. Point retrieval at a single unified memory of tens of thousands of sessions, with no pre-scoped haystack, and recall stops being free (MiniLM 47, voyage 56 on the 19,195-session `_s` store, June 2026). That is where we expect the memory lifecycle to matter, and it is what hippo measures next (see ROADMAP Part III). It is not shown yet. Decay and sleep are design choices, and the tests so far do not favour them:
|
|
949
|
+
|
|
950
|
+
- **Decay tied with decay switched off.** On hippo's synthetic lifecycle test, full@365 minus decay-off is -0.7 points [-1.4, 0.1] on currentR5, no measurable effect. That test runs 20 sessions, so a 365-day half-life barely decays inside it ([mechanism audit, round 2](docs/evals/2026-09-23-mechanism-audit-round2-result.md); [decay default](docs/evals/2026-09-24-decay-default-result.md)).
|
|
951
|
+
- **Sleep lowered LongMemEval recall.** The slept store loses 3.6 points of hit@5 [-5.8, -1.4] to the never-slept one under the audit's declared scorer, and no scorer there shows sleep helping recall ([mechanism audit, round 2](docs/evals/2026-09-23-mechanism-audit-round2-result.md)).
|
|
952
|
+
- **Outcome marks helped on the synthetic test.** Plain BM25 plus the fast outcome nudge drops trap persistence from 71.9% to 0.0% (round 2). Round 1 found outcome feedback and retrieval strengthening each help there, in the test's best case: every outcome mark is right, and every scheduled recall repeats the probe's query ([round 1](docs/evals/2026-09-23-mechanism-audit-result.md)). Supersession is not measured yet.
|
|
953
|
+
|
|
954
|
+
None of this shows hippo making agents better at their work; nothing published has shown that.
|
|
895
955
|
|
|
896
956
|
**Hippo v0.28.0 oracle-split results (hybrid BM25 + cosine, full 500 questions, pooled retrieval):**
|
|
897
957
|
|
|
@@ -916,7 +976,7 @@ For context: MemPalace scores 96.6% (raw) using ChromaDB embeddings + spatial in
|
|
|
916
976
|
|
|
917
977
|
Hippo's strongest categories (single-session-assistant 100% R@5, knowledge-update 89.7%) are where keyword overlap between question and stored content is highest. The weakest (preference 20%) involves indirect references that need deeper semantic understanding.
|
|
918
978
|
|
|
919
|
-
> Note: v0.28 R@10 is 1.6pp below v0.11's BM25-only result. The earlier v0.27 benchmark showed an apparent 35pp regression
|
|
979
|
+
> Note: v0.28 R@10 is 1.6pp below v0.11's BM25-only result. The earlier v0.27 benchmark showed an apparent 35pp regression; that was a methodology bug (budget-limited retrieval vs unlimited), fixed in v0.28 with the `minResults` option. See [`evals/README.md`](evals/README.md) for the full investigation and per-type breakdown.
|
|
920
980
|
|
|
921
981
|
```bash
|
|
922
982
|
cd benchmarks/longmemeval
|
|
@@ -948,10 +1008,10 @@ No other public benchmark tests whether memory systems produce learning curves.
|
|
|
948
1008
|
|
|
949
1009
|
50 tasks, 10 trap categories, each appearing 2-3 times across the sequence.
|
|
950
1010
|
|
|
951
|
-
> **v0.11.0 informal results
|
|
1011
|
+
> **v0.11.0 informal results: RETRACTED v1.7.9.** The 78% → 14% magnitude does NOT reproduce on the formal sequential-learning benchmark. Three pre-registered workload variants (v1.7.5 full-late, v1.7.6 budget sweep, v1.7.7 `--restrict-late-to 4`) all returned C2 hippo-base late mean = 0.0% across every seed (the workload's late phase saturates structurally). The mechanism (dlPFC goal-stack: `pushGoal`/`completeGoal` hooks, `--use-goal-stack`) is shipped and exercisable. **The magnitude is RETRACTED. The mechanism is shipped; no magnitude is currently claimed.** v1.8.0 (queued) explores adversarial trap categories as mechanism characterisation under the magnitude-smuggling guard in `docs/RETRACTION.md`. Pre-registration trail: `docs/evals/2026-05-07-v1.7.5-goal-stack-eval-prereg.md`, `docs/evals/2026-05-09-v1.7.6-calibration-result.md`, `docs/evals/2026-05-09-v1.7.7-goal-stack-eval-result.md`. CHANGELOG: see v1.7.9 entry.
|
|
952
1012
|
|
|
953
1013
|
<details>
|
|
954
|
-
<summary>Original v0.11.0 informal numbers (RETRACTED
|
|
1014
|
+
<summary>Original v0.11.0 informal numbers (RETRACTED, preserved as audit trail in git, not reproduced here)</summary>
|
|
955
1015
|
|
|
956
1016
|
v0.11.0 reported a single-run informal headline citing late-phase trap-rate decline on the sequential-learning benchmark. The specific numbers are archived at git tag `v0.11.0` and the corresponding `CHANGELOG.md` historical entry. Retained in version control, not reproduced here, since reproduction risks accidental re-citation. See `git show v0.11.0 -- README.md` for the original wording.
|
|
957
1017
|
|
|
@@ -970,7 +1030,7 @@ node run.mjs --adapter all
|
|
|
970
1030
|
|
|
971
1031
|
### How do I give Claude Code memory between sessions?
|
|
972
1032
|
|
|
973
|
-
Run `npm install -g hippo-memory`, then `hippo init` in the project. If the project has a `CLAUDE.md`, init adds a short block telling Claude to run `hippo context --auto` when a session starts. It also adds
|
|
1033
|
+
Run `npm install -g hippo-memory`, then `hippo init` in the project. If the project has a `CLAUDE.md`, init adds a short block telling Claude to run `hippo context --auto` when a session starts. It also adds 7 hook entries to Claude Code's settings that keep your pinned memories plus the five most recent ones in context, save a task snapshot before compaction, store failed tool calls as lessons, and run `hippo sleep` when the session ends, and it sets up a daily 6:15am run. [What hippo init changes](#what-hippo-init-changes) lists everything. The [Claude Code plugin](https://github.com/kitfunso/hippo-memory/tree/master/extensions/claude-code-plugin) is the alternative to these hooks; use one, not both. To set up every git repo up to three folders below your home directory at once, know what the scan changes first: each repo gets its own store, seeded from a year of its git history, the same hooks go in when one of those repos uses Claude Code, and the daily run is set up, but no block goes into any repo's `CLAUDE.md`. The command is `hippo init --scan ~`; run `hippo init` in the projects where you want the block.
|
|
974
1034
|
|
|
975
1035
|
### How do I give Cursor memory between sessions?
|
|
976
1036
|
|
|
@@ -978,11 +1038,11 @@ Run `npm install -g hippo-memory`, then `hippo init` in the project. If the proj
|
|
|
978
1038
|
|
|
979
1039
|
### How do I give Codex memory across sessions?
|
|
980
1040
|
|
|
981
|
-
`hippo init` adds its instructions to your `AGENTS.md`, which Codex reads before it starts work. Capturing Codex sessions is opt-in: `hippo hook install codex` wraps the Codex launcher, and `hippo hook uninstall codex` removes the wrapper.
|
|
1041
|
+
`hippo init` adds its instructions to your `AGENTS.md`, which Codex reads before it starts work. When Codex is installed, init also adds two hooks to Codex's `hooks.json`: one puts your pinned memories plus the five most recent ones into every prompt, the other makes the next prompt send them again after a compaction. Codex asks you to trust each new hook once in `/hooks`, and skips it until you do. Capturing Codex sessions is opt-in: `hippo hook install codex` wraps the Codex launcher, and `hippo hook uninstall codex` removes the wrapper and the hooks.
|
|
982
1042
|
|
|
983
1043
|
### Which agents does hippo work with?
|
|
984
1044
|
|
|
985
|
-
`hippo init` detects Claude Code, Codex, Cursor, OpenClaw, OpenCode and Pi
|
|
1045
|
+
`hippo init` detects Claude Code, Codex, Cursor, OpenClaw, OpenCode and Pi. It installs hooks for Claude Code and OpenCode, adds 2 hooks to Codex's `hooks.json` when Codex is installed (Codex runs them once you trust them in `/hooks`), and adds instructions to an existing `AGENTS.md` for Codex, Cursor, OpenClaw and Pi. It only patches instruction files that already exist. Any MCP client can use the [MCP server](#mcp-server), and other tools can call the CLI or the HTTP API that `hippo serve` starts.
|
|
986
1046
|
|
|
987
1047
|
### Can I use hippo as an MCP memory server?
|
|
988
1048
|
|
|
@@ -990,7 +1050,7 @@ Yes. `hippo mcp` runs the server over stdio, and `npx -y hippo-memory mcp` runs
|
|
|
990
1050
|
|
|
991
1051
|
### How is hippo different from mem0?
|
|
992
1052
|
|
|
993
|
-
mem0 uses a language model to extract memories, OpenAI by default in its open-source library, and memories stored through its hosted MCP server live in your Mem0 account ([mem0 docs](https://docs.mem0.ai/platform/mem0-mcp), checked 2026-09-28). Hippo stores memories in SQLite on your machine, needs no account and no model, and `hippo init` wires it into the coding agents it finds. mem0's platform
|
|
1053
|
+
mem0 uses a language model to extract memories, OpenAI by default in its open-source library, and memories stored through its hosted MCP server live in your Mem0 account ([mem0 docs](https://docs.mem0.ai/platform/mem0-mcp), checked 2026-09-28). Hippo stores memories in SQLite on your machine, needs no account and no model, and `hippo init` wires it into the coding agents it finds. mem0's platform marks an older fact superseded when a newer one replaces it; in hippo you run `hippo supersede` yourself. Hippo also lets you mark a recalled memory wrong with `hippo outcome --bad`, and it drops out of the top results.
|
|
994
1054
|
|
|
995
1055
|
### Is this just RAG?
|
|
996
1056
|
|
|
@@ -998,7 +1058,7 @@ No. RAG searches a fixed corpus. Hippo's store changes as your agent works: a me
|
|
|
998
1058
|
|
|
999
1059
|
### Does it need embeddings?
|
|
1000
1060
|
|
|
1001
|
-
No. Recall runs on BM25 out of the box, with no model and no network call, and a default install has no embedder. Embeddings are an optional install for hybrid search. On LongMemEval-S, where each question gets its own haystack, the benchmark scripts
|
|
1061
|
+
No. Recall runs on BM25 out of the box, with no model and no network call, and a default install has no embedder. Embeddings are an optional install for hybrid search. On LongMemEval-S, where each question gets its own haystack, `hippo recall` on a default install puts an answer session in its top five for 85.6% of questions inside its 4,000-token budget, and 87.6% with the free local MiniLM embedder; the budget, not the embedder, is most of the gap to the benchmark scripts. Those scripts, which are not `hippo recall`, fuse BM25 with MiniLM and reach 98.0% recall@5 at their best of five settings, counting a hit when any answer session is in the top five. On LongMemEval's oracle split with one pooled store, BM25 alone scored 74.0% recall@5 in v0.11. These runs use different setups, so they are not a before and after.
|
|
1002
1062
|
|
|
1003
1063
|
### Do I still need CLAUDE.md?
|
|
1004
1064
|
|
|
@@ -1014,7 +1074,7 @@ On your machine, in SQLite: `.hippo/hippo.db` in each project, plus a global sto
|
|
|
1014
1074
|
|
|
1015
1075
|
### What does hippo cost?
|
|
1016
1076
|
|
|
1017
|
-
Nothing. Hippo is MIT-licensed and needs no account or API key. Optional features that call an outside provider bill through it: the Jev reranker costs about 0.0004 USD a recall, and API embedders and sleep's fact extraction bill your own keys. Memory text handed to your agent uses context tokens, and `hippo tokens` shows
|
|
1077
|
+
Nothing. Hippo is MIT-licensed and needs no account or API key. Optional features that call an outside provider bill through it: the Jev reranker costs about 0.0004 USD a recall, and API embedders and sleep's fact extraction bill your own keys. Memory text handed to your agent uses context tokens when it is sent, and again, usually at the cheaper cached-input rate, on each later model call until the host compacts. `hippo tokens` shows both for the hook and compact-resume blocks, and the sent tokens for the rest.
|
|
1018
1078
|
|
|
1019
1079
|
### Is it production-ready?
|
|
1020
1080
|
|
|
@@ -1022,7 +1082,7 @@ Judge it by what is tested. 3,500+ tests run against a real database, with no mo
|
|
|
1022
1082
|
|
|
1023
1083
|
### Has hippo been shown to make agents better at their work?
|
|
1024
1084
|
|
|
1025
|
-
|
|
1085
|
+
No. The published numbers measure retrieval: whether the right memory comes back, and whether a memory marked wrong stays out of the results. Decay and sleep are design choices, not measured wins. In the [mechanism audit](https://github.com/kitfunso/hippo-memory/blob/master/docs/evals/2026-09-23-mechanism-audit-round2-result.md), decay had no measurable effect against decay switched off (-0.7 points [-1.4, 0.1] on a 20-session synthetic test, too short for a 365-day half-life to act), and sleep lowered LongMemEval hit@5 by 3.6 points [-5.8, -1.4] under the audit's declared scorer; no scorer there showed sleep helping recall. Every measurement, including failed runs and one retracted claim, is indexed in [docs/evals](https://github.com/kitfunso/hippo-memory/blob/master/docs/evals/README.md).
|
|
1026
1086
|
|
|
1027
1087
|
---
|
|
1028
1088
|
|
|
@@ -1031,7 +1091,7 @@ Not yet. The published numbers measure retrieval: whether the right memory comes
|
|
|
1031
1091
|
Issues and PRs welcome. Before contributing, run `hippo status` in the repo root to see the project's own memory.
|
|
1032
1092
|
|
|
1033
1093
|
The interesting problems:
|
|
1034
|
-
- **LongMemEval retrieval
|
|
1094
|
+
- **LongMemEval retrieval.** `hippo recall` on a default install scores 85.6% R@5 inside its 4,000-token budget and 96.8% with the budget lifted ([result](docs/evals/2026-09-28-recall-cli-longmemeval-result.md)); a budget that fits five median sessions is the next run. The benchmark scripts' best of five settings reach 98.0% with the free local embedder (re-measured 2026-09-23; the June build gave 98.6) and 99.8% with voyage-3-large (measured 2026-06-09), counting a hit when any answer session is in the top 5. That measure is near its ceiling. The all-evidence one, every answer session in the top 5 over the 470 questions with an answer, is not; there the MiniLM runs score 86.8 to 88.5% against gbrain's 95.53%. The lifecycle stress eval (ROADMAP Part III) is the next measurement.
|
|
1035
1095
|
- Better consolidation heuristics (LLM-powered merge vs current text overlap)
|
|
1036
1096
|
- Web UI / dashboard for visualizing decay curves and memory health
|
|
1037
1097
|
- Optimal decay parameter tuning from real usage data
|