hippo-memory 1.52.5 → 1.52.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +56 -33
- package/dist/api.js +2 -1
- package/dist/autolearn.d.ts +1 -4
- package/dist/autolearn.js +3 -5
- package/dist/capture.js +3 -0
- package/dist/cli.js +167 -97
- package/dist/config.js +8 -1
- package/dist/consolidate.d.ts +2 -0
- package/dist/consolidate.js +13 -6
- package/dist/customer-notes.js +3 -2
- package/dist/dag.js +5 -0
- package/dist/decisions.d.ts +1 -1
- package/dist/decisions.js +4 -3
- package/dist/extract.js +3 -0
- package/dist/half-life-migration.d.ts +14 -5
- package/dist/half-life-migration.js +89 -21
- package/dist/hooks.js +1 -1
- package/dist/importers.d.ts +2 -1
- package/dist/importers.js +19 -7
- package/dist/incidents.d.ts +1 -1
- package/dist/incidents.js +4 -3
- package/dist/index.d.ts +4 -1
- package/dist/index.js +6 -1
- package/dist/memory.d.ts +6 -13
- package/dist/memory.js +0 -10
- package/dist/policies.js +3 -2
- package/dist/predictions.js +2 -0
- package/dist/processes.js +3 -2
- package/dist/project-briefs.js +3 -2
- package/dist/skills.js +3 -2
- package/dist/store.d.ts +2 -0
- package/dist/store.js +3 -0
- package/dist/support-bundle.js +35 -14
- package/dist/version.d.ts +1 -1
- package/dist/version.js +1 -1
- package/extensions/openclaw-plugin/openclaw.plugin.json +1 -1
- package/extensions/openclaw-plugin/package.json +1 -1
- package/openclaw.plugin.json +1 -1
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -37,7 +37,7 @@ Dependencies: Zero runtime deps. Node.js 22.16+. Optional embeddings: bring-you
|
|
|
37
37
|
|
|
38
38
|
Most "AI memory" systems save everything and search later. That's storage with semantic search bolted on. It's why your agent kept hitting the same deploy bug last week. And the week before. The system saw the failure four times. It had no way to know it should remember.
|
|
39
39
|
|
|
40
|
-
Hippo learns from outcomes. When a recalled memory turns out wrong, mark it bad and it drops out of the top results.
|
|
40
|
+
Hippo learns from outcomes. When a recalled memory turns out wrong, mark it bad and it drops out of the top results. Memories you keep using get stronger. Those two are the parts we measured helping ([mechanism audit, round 2](https://github.com/kitfunso/hippo-memory/pull/232)). When a fact changes, the new version supersedes the old one; we have not measured whether that helps. The design borrows from the hippocampus (decay, three layers, sleep consolidation), but that is inspiration. We have not measured decay or sleep making recall better.
|
|
41
41
|
|
|
42
42
|
It also fixes the portability problem. Your ChatGPT memories don't travel to Claude. Your `.cursorrules` don't travel to Codex. Hippo is one process behind every agent. CLAUDE.md, Cursor rules, ChatGPT exports, Slack history, all in one SQLite store, all queryable from any tool that speaks MCP or HTTP.
|
|
43
43
|
|
|
@@ -64,7 +64,7 @@ claim we retracted.
|
|
|
64
64
|
|
|
65
65
|
- **Stops repeating mistakes.** Tag a failure with `--tag error` once, the lesson surfaces every time the agent walks back into that part of the code. Errors decay slower than ordinary observations.
|
|
66
66
|
- **Survives tool switches.** Use Claude Code on Monday, Cursor on Tuesday, Codex on Wednesday. Same `.hippo/` store. Same memories. Pick up exactly where you left off.
|
|
67
|
-
- **Ingests systems of record.** Slack today (`POST /v1/connectors/slack/events`).
|
|
67
|
+
- **Ingests systems of record.** Slack and GitHub today (`POST /v1/connectors/slack/events`, `POST /v1/connectors/github/events`). Jira and Notion next. Webhooks land as `kind='raw'` memories with full provenance and GDPR-correct deletion.
|
|
68
68
|
- **Knows where every memory came from.** Every row carries `kind`, `scope`, `owner`, and `artifact_ref`. Right-to-be-forgotten is a single API call, not an audit nightmare.
|
|
69
69
|
- **Plays nice with multi-tenant.** API keys, scrypt-hashed. Audit log on every mutation. Tenant A literally cannot see tenant B's memories. Proven by negative test.
|
|
70
70
|
|
|
@@ -82,7 +82,7 @@ hippo init
|
|
|
82
82
|
hippo init --scan ~
|
|
83
83
|
```
|
|
84
84
|
|
|
85
|
-
`--scan` finds every git repo under your home directory, creates a `.hippo/` store in each one, and seeds it with lessons from the last
|
|
85
|
+
`--scan` finds every git repo under your home directory, creates a `.hippo/` store in each one, and seeds it with lessons from the last 365 days of commit history. One command, instant memory across all your projects. It installs the Claude Code hooks and the OpenCode plugin when it finds those agents, but patches no instruction file; run `hippo init` inside a repo to add the block to its `CLAUDE.md` or `AGENTS.md`.
|
|
86
86
|
|
|
87
87
|
After setup, `hippo sleep` runs at session end (via auto-installed agent hooks) and does five things:
|
|
88
88
|
|
|
@@ -116,7 +116,7 @@ hippo init
|
|
|
116
116
|
# Auto-installed claude-code hook in CLAUDE.md
|
|
117
117
|
```
|
|
118
118
|
|
|
119
|
-
If you have a `CLAUDE.md`, it patches it. `AGENTS.md` for Codex/OpenClaw/OpenCode/Pi.
|
|
119
|
+
If you have a `CLAUDE.md`, it patches it. `AGENTS.md` for Codex/Cursor/OpenClaw/OpenCode/Pi. Your agent starts using Hippo on its next session. For Codex session capture, Hippo wraps the codex launcher only when you explicitly opt in with `hippo hook install codex` (init prints the command when it detects Codex; undo anytime with `hippo hook uninstall codex`).
|
|
120
120
|
|
|
121
121
|
It also registers the current project in Hippo's workspace registry and installs one machine-level daily runner (6:15am). That runner sweeps every registered workspace, runs `hippo learn --git --days 1`, then `hippo sleep`. You get strict daily consolidation without creating one OS task per project.
|
|
122
122
|
|
|
@@ -289,13 +289,13 @@ sequenceDiagram
|
|
|
289
289
|
participant E as Episodic
|
|
290
290
|
participant S as Semantic
|
|
291
291
|
Agent->>B: hippo remember "cache dropped tips_10y" --error
|
|
292
|
-
B->>E: encode (half_life=
|
|
292
|
+
B->>E: encode (half_life=730d, valence=neg)
|
|
293
293
|
Note over E: strength=1.0
|
|
294
294
|
Agent->>E: hippo recall "data pipeline"
|
|
295
295
|
E-->>Agent: returns memory (rank 1)
|
|
296
|
-
Note over E: half_life
|
|
296
|
+
Note over E: half_life 730d → 732d, retrieval_count++
|
|
297
297
|
Agent->>E: hippo outcome --good
|
|
298
|
-
Note over E: reward_factor 1.0 → 1.
|
|
298
|
+
Note over E: reward_factor 1.0 → 1.25
|
|
299
299
|
Agent->>S: hippo sleep
|
|
300
300
|
S->>E: merge 3 related episodic → 1 semantic
|
|
301
301
|
Note over E,S: original episodic decays, pattern survives
|
|
@@ -356,7 +356,7 @@ hippo invalidate "REST API" --reason "migrated to GraphQL"
|
|
|
356
356
|
|
|
357
357
|
### Architectural decisions
|
|
358
358
|
|
|
359
|
-
One-off decisions don't repeat, so they can't earn their keep through retrieval alone. `hippo decide` stores them with
|
|
359
|
+
One-off decisions don't repeat, so they can't earn their keep through retrieval alone. `hippo decide` stores them with verified confidence and the store's default half-life, the same as any other memory, and sleep never retires the memory behind a decision. On a store made before 1.52.7, decisions get 90 days until the store's first `hippo sleep` on 1.52.7 or later moves them to the default.
|
|
360
360
|
|
|
361
361
|
```bash
|
|
362
362
|
hippo decide "Use PostgreSQL for all new services" --context "JSONB support"
|
|
@@ -378,7 +378,7 @@ Tag a memory as an error and it gets 2x the half-life automatically.
|
|
|
378
378
|
|
|
379
379
|
```bash
|
|
380
380
|
hippo remember "deployment failed: forgot to run migrations" --error
|
|
381
|
-
# half_life:
|
|
381
|
+
# half_life: 730d instead of 365d
|
|
382
382
|
# emotional_valence: negative
|
|
383
383
|
# strength formula applies 2.0x multiplier (HIPPO_LOSS_AVERSION_RATIO=0.75 to keep v1.13.4 1.5x)
|
|
384
384
|
|
|
@@ -657,11 +657,11 @@ hippo watch "npm run build"
|
|
|
657
657
|
| `hippo sync` | Pull global memories into local project |
|
|
658
658
|
| `hippo invalidate "<pattern>"` | Actively weaken memories matching an old pattern |
|
|
659
659
|
| `hippo invalidate "<pattern>" --reason "<why>"` | Include what replaced it |
|
|
660
|
-
| `hippo decide "<decision>"` | Record architectural decision
|
|
660
|
+
| `hippo decide "<decision>"` | Record architectural decision |
|
|
661
661
|
| `hippo decide "<decision>" --context "<why>"` | Include reasoning |
|
|
662
662
|
| `hippo decide "<decision>" --supersedes <id>` | Supersede a previous decision |
|
|
663
663
|
| `hippo hook list` | Show available framework hooks |
|
|
664
|
-
| `hippo hook install <target>` | Install hook (claude-code also adds
|
|
664
|
+
| `hippo hook install <target>` | Install hook (claude-code also adds session hooks: sleep at session end, pinned memories, a snapshot before compaction) |
|
|
665
665
|
| `hippo hook uninstall <target>` | Remove hook |
|
|
666
666
|
| `hippo handoff create --summary "..."` | Create a session handoff |
|
|
667
667
|
| `hippo handoff latest` | Show the most recent handoff |
|
|
@@ -701,7 +701,7 @@ On `heartbeat`, `block`, `review` and `complete`, a given `--run` is checked aga
|
|
|
701
701
|
|-----------|------------|---------|
|
|
702
702
|
| Claude Code | `CLAUDE.md` or `.claude/settings.json` | `CLAUDE.md` + `SessionStart`/`SessionEnd` hooks in `settings.json` |
|
|
703
703
|
| Codex | `AGENTS.md` or `.codex` | `AGENTS.md`; session capture is opt-in with `hippo hook install codex`, which wraps the Codex launcher |
|
|
704
|
-
| Cursor |
|
|
704
|
+
| Cursor | `AGENTS.md` | `AGENTS.md`, which Cursor reads from the project root |
|
|
705
705
|
| OpenClaw | `.openclaw` or `AGENTS.md` | native OpenClaw plugin or `AGENTS.md` |
|
|
706
706
|
| OpenCode | `.opencode/` or `opencode.json` | `AGENTS.md` + TS plugin at `~/.config/opencode/plugins/hippo.ts` (subscribes to `session.idle` + `session.created`) |
|
|
707
707
|
| Pi | `.pi` or `.pi/agent` | `AGENTS.md`; copy the [Pi extension](https://github.com/kitfunso/hippo-memory/tree/master/extensions/pi-extension) for session hooks |
|
|
@@ -715,18 +715,20 @@ If you prefer explicit control:
|
|
|
715
715
|
```bash
|
|
716
716
|
hippo hook install claude-code # patches CLAUDE.md + adds SessionStart/SessionEnd + UserPromptSubmit hooks
|
|
717
717
|
hippo hook install codex # optional repair/manual run: patches AGENTS.md + wraps the detected Codex launcher
|
|
718
|
-
hippo hook install cursor # patches .
|
|
718
|
+
hippo hook install cursor # patches AGENTS.md
|
|
719
719
|
hippo hook install openclaw # patches AGENTS.md
|
|
720
720
|
hippo hook install opencode # patches AGENTS.md + installs the opencode TS plugin
|
|
721
721
|
```
|
|
722
722
|
|
|
723
723
|
This adds a `<!-- hippo:start -->` ... `<!-- hippo:end -->` block that tells the agent to:
|
|
724
724
|
1. Run `hippo context --auto --budget 1500` at session start
|
|
725
|
-
2. Run `hippo remember "<
|
|
726
|
-
3.
|
|
725
|
+
2. Run `hippo remember "<what went wrong and why>" --error` the moment it finds out why something failed, never as a closing step
|
|
726
|
+
3. Capture a short summary with `hippo capture --stdin` when the session ends, but only where no hook captures the session: Cursor, OpenClaw, OpenCode, Pi, and Codex without its wrapper
|
|
727
|
+
|
|
728
|
+
The block asks for nothing a hook already does, because each extra tool call re-reads the whole context. Re-running `hippo init` swaps a block an older hippo wrote for the current one, as long as nobody edited it. It leaves an edited block alone and says so, and never touches text outside the markers.
|
|
727
729
|
|
|
728
730
|
For Claude Code, it also adds:
|
|
729
|
-
- a `SessionEnd` hook
|
|
731
|
+
- a `SessionEnd` hook that runs `hippo sleep` and then `hippo capture` on the session's transcript when the session exits
|
|
730
732
|
- a `SessionStart` hook that prints the previous session's consolidation output
|
|
731
733
|
- a `UserPromptSubmit` hook that runs `hippo context --pinned-only --include-recent 5 --format additional-context` every turn. It re-injects pinned memories (`hippo remember <text> --pin`) plus the last 5 writes, so fresh same-session lessons appear on the next prompt before you pin them. The block is rendered without live strength percentages, so it stays byte-identical while its memories do not change, and it is sent only when it changed since the session's last prompt: an unchanged block is skipped, resent every 10 skips (`pinnedInject.refreshTurns`, `0` never resends) and resent after compaction. `{"pinnedInject":{"skipUnchanged":false}}` sends it every turn as before. Opt out entirely with `{"pinnedInject":{"enabled":false}}` in `.hippo/config.json`.
|
|
732
734
|
- a `PreCompact` hook that runs `hippo pre-compact` before the transcript gets summarized. It saves a working-state snapshot (task/summary/next step) so mid-session compaction can't drop it; the `SessionEnd` hook still owns extracting durable memories.
|
|
@@ -738,22 +740,31 @@ To remove: `hippo hook uninstall claude-code`
|
|
|
738
740
|
|
|
739
741
|
### What the hook adds (Claude Code example)
|
|
740
742
|
|
|
741
|
-
|
|
743
|
+
````markdown
|
|
742
744
|
## Project Memory (Hippo)
|
|
743
745
|
|
|
744
|
-
|
|
746
|
+
Pinned rules and recent writes auto-inject at every prompt via the installed
|
|
747
|
+
UserPromptSubmit hook; never re-run that part manually. At the START of a
|
|
748
|
+
task (not per prompt), additionally load task-specific context: git-aware
|
|
749
|
+
recall over the full store that per-prompt injection does not cover. Also
|
|
750
|
+
run it if the hook is not installed:
|
|
751
|
+
```bash
|
|
745
752
|
hippo context --auto --budget 1500
|
|
753
|
+
```
|
|
746
754
|
|
|
747
|
-
When you
|
|
755
|
+
When you find out why something failed, record it right then, while you
|
|
756
|
+
work, never as a closing step:
|
|
757
|
+
```bash
|
|
748
758
|
hippo remember "<what went wrong and why>" --error
|
|
749
|
-
|
|
750
|
-
After completing work successfully:
|
|
751
|
-
hippo outcome --good
|
|
752
759
|
```
|
|
753
760
|
|
|
761
|
+
The installed hooks store failed tool calls and capture the session when it
|
|
762
|
+
ends, so there is nothing to run before you finish.
|
|
763
|
+
````
|
|
764
|
+
|
|
754
765
|
### MCP Server
|
|
755
766
|
|
|
756
|
-
For any MCP-compatible client (Cursor, Windsurf, Cline, Claude Desktop):
|
|
767
|
+
For any MCP-compatible client (Cursor, Windsurf (now Devin Desktop), Cline, Claude Desktop):
|
|
757
768
|
|
|
758
769
|
```bash
|
|
759
770
|
hippo mcp # starts MCP server over stdio
|
|
@@ -850,12 +861,12 @@ The AI-memory category matured fast in 2026. Hippo's specific take — bio-decay
|
|
|
850
861
|
| Auto-hook install | Yes | No | No | No | No | No | No | No | No | No |
|
|
851
862
|
| MCP server | Yes | Yes | Yes (hosted, needs an account) | Yes | Yes (stdio + HTTP/OAuth) | Yes (hosted, needs an account) | Yes (hosted, needs an API key) | Yes (first-party Claude/LangGraph) | Yes | ? |
|
|
852
863
|
| Zero runtime deps | Yes | No (ChromaDB) | No | No | No (PGLite or PG+pgvector) | No (managed service) | No (npm deps) | No (Python deps) | Yes (single Rust binary) | No (managed + OSS) |
|
|
853
|
-
| LongMemEval (best published) | 98.0% local / 99.8% voyage R@5 (s_cleaned, per-haystack)\* | 96.6% raw / 100% reranked R@5 | 94.4 (hosted platform)\*\* | N/A |
|
|
864
|
+
| LongMemEval (best published) | 98.0% local / 99.8% voyage any-evidence R@5; 88.5% local all-evidence R@5 (s_cleaned, per-haystack)\* | 96.6% raw / 100% reranked R@5 | 94.4 (hosted platform)\*\* | N/A | 95.53% all-evidence R@5 reranked, 93.19% without (s_cleaned\*) | 90.2% accuracy\*\* (LoCoMo 94.7%) | N/A | N/A | 88.78% overall accuracy w/ reader\*\* | 83.00% overall\*\* (LoCoMo 93.05%, HaluMem 93.04%) |
|
|
854
865
|
| Git-friendly | Yes | No | No | Yes | Yes | No | Yes (memory tracked in git) | No | Yes (Git is the model) | ? |
|
|
855
866
|
| Framework agnostic | Yes | Yes | Partial | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
|
|
856
867
|
| License | MIT | (open) | Apache-2.0 | (open) | MIT | Proprietary cloud (Graphiti: Apache-2.0) | Apache-2.0 | MIT (core) | Apache-2.0 | Apache-2.0 (OSS) + cloud |
|
|
857
868
|
|
|
858
|
-
\* Hippo's 98.0%
|
|
869
|
+
\* Hippo's figures are on `longmemeval_s_cleaned` with a per-question haystack, each the best of five retrieval settings in the benchmark scripts, not `hippo recall`. Any-evidence R@5 counts a hit when any answer session is in the top 5, over all 500 questions: 98.0% with the free local MiniLM embedder (an optional install) and 99.8% with voyage-3-large (measured 2026-06-09, not re-run). All-evidence R@5 counts a hit only when every answer session is in the top 5, over the 470 questions that have an answer: 86.8 to 88.5% with MiniLM. gbrain first published 97.6%, an any-evidence score over all 500; its [report](https://github.com/garrytan/gbrain-evals/blob/main/docs/benchmarks/2026-05-07-longmemeval-s.md) now leads with all-evidence, 95.53% (449 of 470) with the paid Voyage rerank-2.5 reranker and 93.19% without it. On all-evidence recall gbrain is ahead. The June 2026 build scored 98.6 any-evidence; [`docs/evals/2026-09-23-longmemeval-reproduction.md`](docs/evals/2026-09-23-longmemeval-reproduction.md) has both runs. An older hippo number, 86.8% R@5 on `longmemeval_oracle` under pooled (non-per-haystack) retrieval, is not comparable to per-haystack figures.
|
|
859
870
|
|
|
860
871
|
\*\* Different metric: these are end-to-end answer scores, not retrieval R@5. Mem0's 94.4 comes from its hosted platform, which its README says includes optimizations the open-source SDK lacks. Zep's 90.2% and 94.7% are accuracy figures from its homepage. Memoria's 88.78% and EverMind's 83% are overall accuracy with a reader LLM. Higher denominator + LLM helps. Not directly comparable to retrieval-only R@5 numbers above. The Mem0, Zep and Letta columns were last checked against each vendor's own pages on 2026-09-28.
|
|
861
872
|
|
|
@@ -878,9 +889,9 @@ Three benchmarks testing three different things. Full details in [`benchmarks/`]
|
|
|
878
889
|
| MiniLM-L6 (local, optional install) | 96.8 | 98.0 | 88.4 |
|
|
879
890
|
| voyage-3-large (opt-in, paid) | 99.8 | 99.8 | 94.6 |
|
|
880
891
|
|
|
881
|
-
|
|
892
|
+
These are any-evidence scores: a hit when any answer session is in the top 5, over all 500 questions. Counting a hit only when every answer session is in the top 5, over the 470 questions that have an answer, the MiniLM runs score 86.8 to 88.5. gbrain first published 97.6 any-evidence and now reports 95.53 all-evidence with a paid reranker, 93.19 without, so on all-evidence recall gbrain is ahead. These numbers come from the scripts in `benchmarks/longmemeval/`, which index every turn and fuse BM25 with dense ranks; they are not `hippo recall`, and a default install has no embedder. Re-measure: [`docs/evals/2026-09-23-longmemeval-reproduction.md`](docs/evals/2026-09-23-longmemeval-reproduction.md). Any-evidence recall is near its ceiling on this task; all-evidence recall is not. Method and the global-pool comparison: [`docs/evals/2026-06-09-longmemeval-per-haystack-dual.md`](docs/evals/2026-06-09-longmemeval-per-haystack-dual.md).
|
|
882
893
|
|
|
883
|
-
The differentiator is what happens as one store grows. Point retrieval at a single unified memory of tens of thousands of sessions, with no pre-scoped haystack, and recall stops being free (MiniLM 47, voyage 56 on the 19,195-session `_s` store, June 2026). That is where we expect the memory lifecycle to matter, and it is what hippo measures next (see ROADMAP Part III). It is not shown yet: in our tests so far, decay made no measurable difference and sleep cost recall. Outcome marks and
|
|
894
|
+
The differentiator is what happens as one store grows. Point retrieval at a single unified memory of tens of thousands of sessions, with no pre-scoped haystack, and recall stops being free (MiniLM 47, voyage 56 on the 19,195-session `_s` store, June 2026). That is where we expect the memory lifecycle to matter, and it is what hippo measures next (see ROADMAP Part III). It is not shown yet: in our tests so far, decay made no measurable difference and sleep cost recall. Outcome marks and retrieval strengthening are what measured helpful; supersession is not measured yet.
|
|
884
895
|
|
|
885
896
|
**Hippo v0.28.0 oracle-split results (hybrid BM25 + cosine, full 500 questions, pooled retrieval):**
|
|
886
897
|
|
|
@@ -959,11 +970,11 @@ node run.mjs --adapter all
|
|
|
959
970
|
|
|
960
971
|
### How do I give Claude Code memory between sessions?
|
|
961
972
|
|
|
962
|
-
Run `npm install -g hippo-memory`, then `hippo init` in the project
|
|
973
|
+
Run `npm install -g hippo-memory`, then `hippo init` in the project. If the project has a `CLAUDE.md`, init adds a short block telling Claude to run `hippo context --auto` when a session starts. It also adds hooks to Claude Code's settings that keep your pinned memories in context, save a task snapshot before compaction, and run `hippo sleep` when the session ends. `hippo init --scan ~` gives every git repo under your home folder a store and installs the same hooks, but adds no block to any `CLAUDE.md`. The [Claude Code plugin](https://github.com/kitfunso/hippo-memory/tree/master/extensions/claude-code-plugin) is the alternative to these hooks; use one, not both.
|
|
963
974
|
|
|
964
975
|
### How do I give Cursor memory between sessions?
|
|
965
976
|
|
|
966
|
-
`hippo init` adds its instructions to
|
|
977
|
+
`hippo init` adds its instructions to `AGENTS.md` if the project has one, and Cursor reads that file from the project root. The [MCP server](#mcp-server) gives Cursor's agent tools to recall and store memories once you add `hippo mcp` to `.cursor/mcp.json`. Older hippo versions wrote the block to `.cursorrules`; `hippo hook uninstall cursor` removes it from there, and from `AGENTS.md` only when the block there is Cursor's own and unedited (`hippo hook install cursor` puts it back). A block written for Codex or another agent stays, since Cursor reads it too, and so does an edited block, since hippo cannot tell whose it is. `hippo import --cursor .cursor/rules` turns your existing rules into memories; it reads an older single `.cursorrules` file too.
|
|
967
978
|
|
|
968
979
|
### How do I give Codex memory across sessions?
|
|
969
980
|
|
|
@@ -975,19 +986,27 @@ Run `npm install -g hippo-memory`, then `hippo init` in the project, or `hippo i
|
|
|
975
986
|
|
|
976
987
|
### Can I use hippo as an MCP memory server?
|
|
977
988
|
|
|
978
|
-
Yes. `hippo mcp` runs the server over stdio, and `npx -y hippo-memory mcp` runs it without a global install. Add it to the MCP config of Claude Desktop, Cursor, Windsurf, Cline or any other client (example [above](#mcp-server)); in Claude Code, run `claude mcp add hippo-memory -- hippo mcp`. The agent gets tools such as `hippo_recall`, `hippo_remember` and `hippo_outcome`.
|
|
989
|
+
Yes. `hippo mcp` runs the server over stdio, and `npx -y hippo-memory mcp` runs it without a global install. Add it to the MCP config of Claude Desktop, Cursor, Windsurf (now Devin Desktop), Cline or any other client (example [above](#mcp-server)); in Claude Code, run `claude mcp add hippo-memory -- hippo mcp`. The agent gets tools such as `hippo_recall`, `hippo_remember` and `hippo_outcome`.
|
|
979
990
|
|
|
980
991
|
### How is hippo different from mem0?
|
|
981
992
|
|
|
982
993
|
mem0 uses a language model to extract memories, OpenAI by default in its open-source library, and memories stored through its hosted MCP server live in your Mem0 account ([mem0 docs](https://docs.mem0.ai/platform/mem0-mcp), checked 2026-09-28). Hippo stores memories in SQLite on your machine, needs no account and no model, and `hippo init` wires it into the coding agents it finds. mem0's platform and hippo both mark an older fact superseded when a newer one replaces it. Hippo also lets you mark a recalled memory wrong with `hippo outcome --bad`, and it drops out of the top results.
|
|
983
994
|
|
|
995
|
+
### Is this just RAG?
|
|
996
|
+
|
|
997
|
+
No. RAG searches a fixed corpus. Hippo's store changes as your agent works: a memory marked wrong drops out of the top results, a newer fact supersedes the old one, and memories that keep getting recalled last longer while unused ones fade on a half-life. Recall itself is search: BM25, plus embeddings if you install them.
|
|
998
|
+
|
|
999
|
+
### Does it need embeddings?
|
|
1000
|
+
|
|
1001
|
+
No. Recall runs on BM25 out of the box, with no model and no network call, and a default install has no embedder. Embeddings are an optional install for hybrid search. On LongMemEval-S, where each question gets its own haystack, the benchmark scripts (not `hippo recall`) fuse BM25 with the free local MiniLM embedder and reach 98.0% recall@5, counting a hit when any answer session is in the top five. On LongMemEval's oracle split with one pooled store, BM25 alone scored 74.0% recall@5 in v0.11. The two runs use different setups, so they are not a before and after.
|
|
1002
|
+
|
|
984
1003
|
### Do I still need CLAUDE.md?
|
|
985
1004
|
|
|
986
1005
|
Yes, for short standing rules such as build commands, code style and things never to do. Claude Code loads `CLAUDE.md` and its auto memory into every session, and its [memory docs](https://code.claude.com/docs/en/memory) say that when two rules contradict each other, Claude may pick one arbitrarily. Hippo holds the lessons that pile up, recalls the ones that match the task, and retires the ones marked wrong or replaced. `hippo init` adds its block to `CLAUDE.md`, and `hippo import --claude CLAUDE.md` turns existing notes into memories.
|
|
987
1006
|
|
|
988
1007
|
### What happens when a memory turns out to be wrong?
|
|
989
1008
|
|
|
990
|
-
Mark it, and it
|
|
1009
|
+
Mark it, and it drops out of the top results. `hippo outcome --bad` weakens the memories from the last recall, `hippo supersede <id> "<new fact>"` replaces one with a newer version, and `hippo reject <id> --reason "<why>"` stops that value from returning at all. On the synthetic E1 test, where every mark is correct, plain BM25 plus the outcome mark cut how often a marked-bad memory stayed in the top five from 71.9% to 0.0%. Real marks are noisier, because `--bad` marks the whole recall batch.
|
|
991
1010
|
|
|
992
1011
|
### Where does hippo keep my data?
|
|
993
1012
|
|
|
@@ -997,9 +1016,13 @@ On your machine, in SQLite: `.hippo/hippo.db` in each project, plus a global sto
|
|
|
997
1016
|
|
|
998
1017
|
Nothing. Hippo is MIT-licensed and needs no account or API key. Optional features that call an outside provider bill through it: the Jev reranker costs about 0.0004 USD a recall, and API embedders and sleep's fact extraction bill your own keys. Memory text handed to your agent uses context tokens, and `hippo tokens` shows how many.
|
|
999
1018
|
|
|
1019
|
+
### Is it production-ready?
|
|
1020
|
+
|
|
1021
|
+
Judge it by what is tested. 3,500+ tests run against a real database, with no module mocks and no mocked store, and a negative test checks that one tenant cannot read another's memories. It is MIT-licensed and has zero runtime dependencies. What has not been shown yet is whether agents do better work with it: the published numbers measure retrieval.
|
|
1022
|
+
|
|
1000
1023
|
### Has hippo been shown to make agents better at their work?
|
|
1001
1024
|
|
|
1002
|
-
Not yet. The published numbers measure retrieval: whether the right memory comes back, and whether a memory marked wrong stays out of the results.
|
|
1025
|
+
Not yet. The published numbers measure retrieval: whether the right memory comes back, and whether a memory marked wrong stays out of the results. A paired test that runs real agent sessions with and without hippo is under way. Every measurement, including failed runs and one retracted claim, is indexed in [docs/evals](https://github.com/kitfunso/hippo-memory/blob/master/docs/evals/README.md).
|
|
1003
1026
|
|
|
1004
1027
|
---
|
|
1005
1028
|
|
|
@@ -1008,7 +1031,7 @@ Not yet. The published numbers measure retrieval: whether the right memory comes
|
|
|
1008
1031
|
Issues and PRs welcome. Before contributing, run `hippo status` in the repo root to see the project's own memory.
|
|
1009
1032
|
|
|
1010
1033
|
The interesting problems:
|
|
1011
|
-
- **LongMemEval retrieval (standard task: done).** Per-question-haystack R@5 is 98.0% with the free local embedder (re-measured 2026-09-23; the June build gave 98.6) and 99.8% with voyage-3-large (measured 2026-06-09),
|
|
1034
|
+
- **LongMemEval retrieval (standard task: done).** Per-question-haystack R@5 is 98.0% with the free local embedder (re-measured 2026-09-23; the June build gave 98.6) and 99.8% with voyage-3-large (measured 2026-06-09), counting a hit when any answer session is in the top 5. That measure is near its ceiling. The all-evidence one, every answer session in the top 5 over the 470 questions with an answer, is not; there the MiniLM runs score 86.8 to 88.5% against gbrain's 95.53%. The lifecycle stress eval (ROADMAP Part III) is the next measurement.
|
|
1012
1035
|
- Better consolidation heuristics (LLM-powered merge vs current text overlap)
|
|
1013
1036
|
- Web UI / dashboard for visualizing decay curves and memory health
|
|
1014
1037
|
- Optimal decay parameter tuning from real usage data
|
package/dist/api.js
CHANGED
|
@@ -1180,6 +1180,7 @@ export function supersede(ctx, oldId, newContent) {
|
|
|
1180
1180
|
confidence: 'verified',
|
|
1181
1181
|
tenantId: ctx.tenantId,
|
|
1182
1182
|
scope: old.scope,
|
|
1183
|
+
baseHalfLifeDays: loadConfig(ctx.hippoRoot).defaultHalfLifeDays,
|
|
1183
1184
|
});
|
|
1184
1185
|
// Race-safe transition: open a fresh db handle, BEGIN IMMEDIATE, run all
|
|
1185
1186
|
// three steps (CAS on old + writeEntryDbOnly(new) + supersede audit row)
|
|
@@ -2093,7 +2094,7 @@ export function restoreDormant(ctx, id) {
|
|
|
2093
2094
|
// (The placeholder only satisfies createMemory's minimum length, so a
|
|
2094
2095
|
// legacy row shorter than 3 chars can still be restored.)
|
|
2095
2096
|
const revived = {
|
|
2096
|
-
...createMemory('dormant snapshot defaults'),
|
|
2097
|
+
...createMemory('dormant snapshot defaults', { baseHalfLifeDays: loadConfig(ctx.hippoRoot).defaultHalfLifeDays }),
|
|
2097
2098
|
...dormant.entry,
|
|
2098
2099
|
last_retrieved: now.toISOString(),
|
|
2099
2100
|
};
|
package/dist/autolearn.d.ts
CHANGED
|
@@ -3,10 +3,7 @@
|
|
|
3
3
|
* Agents learn from failures without explicit hippo remember calls.
|
|
4
4
|
*/
|
|
5
5
|
import { MemoryEntry } from './memory.js';
|
|
6
|
-
/**
|
|
7
|
-
* Create a MemoryEntry capturing a command failure.
|
|
8
|
-
* Content format: "Command '<cmd>' failed: <truncated stderr>"
|
|
9
|
-
*/
|
|
6
|
+
/** A memory of a failed command, "Command '<cmd>' failed: <truncated stderr>"; no store is in reach, so `hippo watch` re-derives its half-life from the store's config. */
|
|
10
7
|
export declare function captureError(exitCode: number, stderr: string, command: string, tenantId?: string): MemoryEntry;
|
|
11
8
|
/**
|
|
12
9
|
* Parse git log output for actionable lessons.
|
package/dist/autolearn.js
CHANGED
|
@@ -3,14 +3,11 @@
|
|
|
3
3
|
* Agents learn from failures without explicit hippo remember calls.
|
|
4
4
|
*/
|
|
5
5
|
import { execSync, execFileSync, spawn } from 'child_process';
|
|
6
|
-
import { createMemory, Layer } from './memory.js';
|
|
6
|
+
import { createMemory, Layer, DEFAULT_HALF_LIFE_DAYS } from './memory.js';
|
|
7
7
|
import { loadAllEntries } from './store.js';
|
|
8
8
|
import { textOverlap } from './search.js';
|
|
9
9
|
import { isContentWorthStoring } from './audit.js';
|
|
10
|
-
/**
|
|
11
|
-
* Create a MemoryEntry capturing a command failure.
|
|
12
|
-
* Content format: "Command '<cmd>' failed: <truncated stderr>"
|
|
13
|
-
*/
|
|
10
|
+
/** A memory of a failed command, "Command '<cmd>' failed: <truncated stderr>"; no store is in reach, so `hippo watch` re-derives its half-life from the store's config. */
|
|
14
11
|
export function captureError(exitCode, stderr, command, tenantId) {
|
|
15
12
|
// Truncate to first 500 chars to avoid storing megabytes of build logs
|
|
16
13
|
const wasTruncated = stderr.length > 500;
|
|
@@ -30,6 +27,7 @@ export function captureError(exitCode, stderr, command, tenantId) {
|
|
|
30
27
|
source: 'autolearn',
|
|
31
28
|
confidence: 'observed',
|
|
32
29
|
tenantId,
|
|
30
|
+
baseHalfLifeDays: DEFAULT_HALF_LIFE_DAYS,
|
|
33
31
|
});
|
|
34
32
|
}
|
|
35
33
|
/**
|
package/dist/capture.js
CHANGED
|
@@ -21,6 +21,7 @@ import { defaultPreCompactLogPath } from './hooks.js';
|
|
|
21
21
|
import { redactSecrets } from './secret-detect.js';
|
|
22
22
|
import { RejectedValueError, checkRejectionGuard } from './rejection.js';
|
|
23
23
|
import { openHippoDb, closeHippoDb } from './db.js';
|
|
24
|
+
import { loadConfig } from './config.js';
|
|
24
25
|
// Sentence-level patterns
|
|
25
26
|
//
|
|
26
27
|
// T1 (DF2): each pattern now carries TWO capture groups — group 1 is the
|
|
@@ -790,6 +791,7 @@ function cmdCaptureCore(hippoRoot, options) {
|
|
|
790
791
|
let captured = 0;
|
|
791
792
|
let skipped = 0;
|
|
792
793
|
let rejected = 0;
|
|
794
|
+
const baseHalfLifeDays = loadConfig(targetRoot).defaultHalfLifeDays;
|
|
793
795
|
// AT1 P2 fix (dry-run parity, docs/plans/2026-08-15-at1-rejected-value-tombstone.md):
|
|
794
796
|
// dry-run used to skip the guarded write branch ENTIRELY, so a tombstoned
|
|
795
797
|
// extraction printed as `[capture]` and counted toward `captured` — the
|
|
@@ -823,6 +825,7 @@ function cmdCaptureCore(hippoRoot, options) {
|
|
|
823
825
|
source: 'capture',
|
|
824
826
|
confidence: 'observed',
|
|
825
827
|
tenantId: useGlobal ? undefined : options.tenantId,
|
|
828
|
+
baseHalfLifeDays,
|
|
826
829
|
});
|
|
827
830
|
if (options.dryRun) {
|
|
828
831
|
if (dryRunDb) {
|