hippo-memory 1.52.5 → 1.52.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -37,7 +37,7 @@ Dependencies: Zero runtime deps. Node.js 22.16+. Optional embeddings: bring-you
37
37
 
38
38
  Most "AI memory" systems save everything and search later. That's storage with semantic search bolted on. It's why your agent kept hitting the same deploy bug last week. And the week before. The system saw the failure four times. It had no way to know it should remember.
39
39
 
40
- Hippo learns from outcomes. When a recalled memory turns out wrong, mark it bad and it drops out of the top results. When a fact changes, the new version supersedes the old one. Memories you keep using get stronger. Those are the parts we measured helping ([mechanism audit, round 2](https://github.com/kitfunso/hippo-memory/pull/232)). The design borrows from the hippocampus (decay, three layers, sleep consolidation), but that is inspiration. We have not measured decay or sleep making recall better.
40
+ Hippo learns from outcomes. When a recalled memory turns out wrong, mark it bad and it drops out of the top results. Memories you keep using get stronger. Those two are the parts we measured helping ([mechanism audit, round 2](https://github.com/kitfunso/hippo-memory/pull/232)). When a fact changes, the new version supersedes the old one; we have not measured whether that helps. The design borrows from the hippocampus (decay, three layers, sleep consolidation), but that is inspiration. We have not measured decay or sleep making recall better.
41
41
 
42
42
  It also fixes the portability problem. Your ChatGPT memories don't travel to Claude. Your `.cursorrules` don't travel to Codex. Hippo is one process behind every agent. CLAUDE.md, Cursor rules, ChatGPT exports, Slack history, all in one SQLite store, all queryable from any tool that speaks MCP or HTTP.
43
43
 
@@ -64,7 +64,7 @@ claim we retracted.
64
64
 
65
65
  - **Stops repeating mistakes.** Tag a failure with `--tag error` once, the lesson surfaces every time the agent walks back into that part of the code. Errors decay slower than ordinary observations.
66
66
  - **Survives tool switches.** Use Claude Code on Monday, Cursor on Tuesday, Codex on Wednesday. Same `.hippo/` store. Same memories. Pick up exactly where you left off.
67
- - **Ingests systems of record.** Slack today (`POST /v1/connectors/slack/events`). GitHub, Jira, Notion next. Webhooks land as `kind='raw'` memories with full provenance and GDPR-correct deletion.
67
+ - **Ingests systems of record.** Slack and GitHub today (`POST /v1/connectors/slack/events`, `POST /v1/connectors/github/events`). Jira and Notion next. Webhooks land as `kind='raw'` memories with full provenance and GDPR-correct deletion.
68
68
  - **Knows where every memory came from.** Every row carries `kind`, `scope`, `owner`, and `artifact_ref`. Right-to-be-forgotten is a single API call, not an audit nightmare.
69
69
  - **Plays nice with multi-tenant.** API keys, scrypt-hashed. Audit log on every mutation. Tenant A literally cannot see tenant B's memories. Proven by negative test.
70
70
 
@@ -82,7 +82,7 @@ hippo init
82
82
  hippo init --scan ~
83
83
  ```
84
84
 
85
- `--scan` finds every git repo under your home directory, creates a `.hippo/` store in each one, and seeds it with lessons from the last 30 days of commit history. One command, instant memory across all your projects.
85
+ `--scan` finds every git repo under your home directory, creates a `.hippo/` store in each one, and seeds it with lessons from the last 365 days of commit history. One command, instant memory across all your projects. It installs the Claude Code hooks and the OpenCode plugin when it finds those agents, but patches no instruction file; run `hippo init` inside a repo to add the block to its `CLAUDE.md` or `AGENTS.md`.
86
86
 
87
87
  After setup, `hippo sleep` runs at session end (via auto-installed agent hooks) and does five things:
88
88
 
@@ -116,7 +116,7 @@ hippo init
116
116
  # Auto-installed claude-code hook in CLAUDE.md
117
117
  ```
118
118
 
119
- If you have a `CLAUDE.md`, it patches it. `AGENTS.md` for Codex/OpenClaw/OpenCode/Pi. `.cursorrules` for Cursor. Your agent starts using Hippo on its next session. For Codex session capture, Hippo wraps the codex launcher only when you explicitly opt in with `hippo hook install codex` (init prints the command when it detects Codex; undo anytime with `hippo hook uninstall codex`).
119
+ If you have a `CLAUDE.md`, it patches it. `AGENTS.md` for Codex/Cursor/OpenClaw/OpenCode/Pi. Your agent starts using Hippo on its next session. For Codex session capture, Hippo wraps the codex launcher only when you explicitly opt in with `hippo hook install codex` (init prints the command when it detects Codex; undo anytime with `hippo hook uninstall codex`).
120
120
 
121
121
  It also registers the current project in Hippo's workspace registry and installs one machine-level daily runner (6:15am). That runner sweeps every registered workspace, runs `hippo learn --git --days 1`, then `hippo sleep`. You get strict daily consolidation without creating one OS task per project.
122
122
 
@@ -289,13 +289,13 @@ sequenceDiagram
289
289
  participant E as Episodic
290
290
  participant S as Semantic
291
291
  Agent->>B: hippo remember "cache dropped tips_10y" --error
292
- B->>E: encode (half_life=14d, valence=neg)
292
+ B->>E: encode (half_life=730d, valence=neg)
293
293
  Note over E: strength=1.0
294
294
  Agent->>E: hippo recall "data pipeline"
295
295
  E-->>Agent: returns memory (rank 1)
296
- Note over E: half_life 14d → 16d, retrieval_count++
296
+ Note over E: half_life 730d → 732d, retrieval_count++
297
297
  Agent->>E: hippo outcome --good
298
- Note over E: reward_factor 1.0 → 1.15
298
+ Note over E: reward_factor 1.0 → 1.25
299
299
  Agent->>S: hippo sleep
300
300
  S->>E: merge 3 related episodic → 1 semantic
301
301
  Note over E,S: original episodic decays, pattern survives
@@ -356,7 +356,7 @@ hippo invalidate "REST API" --reason "migrated to GraphQL"
356
356
 
357
357
  ### Architectural decisions
358
358
 
359
- One-off decisions don't repeat, so they can't earn their keep through retrieval alone. `hippo decide` stores them with a 90-day half-life and verified confidence so they survive long enough to matter.
359
+ One-off decisions don't repeat, so they can't earn their keep through retrieval alone. `hippo decide` stores them with verified confidence and the store's default half-life, the same as any other memory, and sleep never retires the memory behind a decision. On a store made before 1.52.7, decisions get 90 days until the store's first `hippo sleep` on 1.52.7 or later moves them to the default.
360
360
 
361
361
  ```bash
362
362
  hippo decide "Use PostgreSQL for all new services" --context "JSONB support"
@@ -378,7 +378,7 @@ Tag a memory as an error and it gets 2x the half-life automatically.
378
378
 
379
379
  ```bash
380
380
  hippo remember "deployment failed: forgot to run migrations" --error
381
- # half_life: 14d instead of 7d
381
+ # half_life: 730d instead of 365d
382
382
  # emotional_valence: negative
383
383
  # strength formula applies 2.0x multiplier (HIPPO_LOSS_AVERSION_RATIO=0.75 to keep v1.13.4 1.5x)
384
384
 
@@ -657,11 +657,11 @@ hippo watch "npm run build"
657
657
  | `hippo sync` | Pull global memories into local project |
658
658
  | `hippo invalidate "<pattern>"` | Actively weaken memories matching an old pattern |
659
659
  | `hippo invalidate "<pattern>" --reason "<why>"` | Include what replaced it |
660
- | `hippo decide "<decision>"` | Record architectural decision (90-day half-life) |
660
+ | `hippo decide "<decision>"` | Record architectural decision |
661
661
  | `hippo decide "<decision>" --context "<why>"` | Include reasoning |
662
662
  | `hippo decide "<decision>" --supersedes <id>` | Supersede a previous decision |
663
663
  | `hippo hook list` | Show available framework hooks |
664
- | `hippo hook install <target>` | Install hook (claude-code also adds Stop hook for auto-sleep) |
664
+ | `hippo hook install <target>` | Install hook (claude-code also adds session hooks: sleep at session end, pinned memories, a snapshot before compaction) |
665
665
  | `hippo hook uninstall <target>` | Remove hook |
666
666
  | `hippo handoff create --summary "..."` | Create a session handoff |
667
667
  | `hippo handoff latest` | Show the most recent handoff |
@@ -701,7 +701,7 @@ On `heartbeat`, `block`, `review` and `complete`, a given `--run` is checked aga
701
701
  |-----------|------------|---------|
702
702
  | Claude Code | `CLAUDE.md` or `.claude/settings.json` | `CLAUDE.md` + `SessionStart`/`SessionEnd` hooks in `settings.json` |
703
703
  | Codex | `AGENTS.md` or `.codex` | `AGENTS.md`; session capture is opt-in with `hippo hook install codex`, which wraps the Codex launcher |
704
- | Cursor | `.cursorrules` or `.cursor/rules` | `.cursorrules` |
704
+ | Cursor | `AGENTS.md` | `AGENTS.md`, which Cursor reads from the project root |
705
705
  | OpenClaw | `.openclaw` or `AGENTS.md` | native OpenClaw plugin or `AGENTS.md` |
706
706
  | OpenCode | `.opencode/` or `opencode.json` | `AGENTS.md` + TS plugin at `~/.config/opencode/plugins/hippo.ts` (subscribes to `session.idle` + `session.created`) |
707
707
  | Pi | `.pi` or `.pi/agent` | `AGENTS.md`; copy the [Pi extension](https://github.com/kitfunso/hippo-memory/tree/master/extensions/pi-extension) for session hooks |
@@ -715,18 +715,20 @@ If you prefer explicit control:
715
715
  ```bash
716
716
  hippo hook install claude-code # patches CLAUDE.md + adds SessionStart/SessionEnd + UserPromptSubmit hooks
717
717
  hippo hook install codex # optional repair/manual run: patches AGENTS.md + wraps the detected Codex launcher
718
- hippo hook install cursor # patches .cursorrules
718
+ hippo hook install cursor # patches AGENTS.md
719
719
  hippo hook install openclaw # patches AGENTS.md
720
720
  hippo hook install opencode # patches AGENTS.md + installs the opencode TS plugin
721
721
  ```
722
722
 
723
723
  This adds a `<!-- hippo:start -->` ... `<!-- hippo:end -->` block that tells the agent to:
724
724
  1. Run `hippo context --auto --budget 1500` at session start
725
- 2. Run `hippo remember "<lesson>" --error` on errors
726
- 3. Run `hippo outcome --good` on completion
725
+ 2. Run `hippo remember "<what went wrong and why>" --error` the moment it finds out why something failed, never as a closing step
726
+ 3. Capture a short summary with `hippo capture --stdin` when the session ends, but only where no hook captures the session: Cursor, OpenClaw, OpenCode, Pi, and Codex without its wrapper
727
+
728
+ The block asks for nothing a hook already does, because each extra tool call re-reads the whole context. Re-running `hippo init` swaps a block an older hippo wrote for the current one, as long as nobody edited it. It leaves an edited block alone and says so, and never touches text outside the markers.
727
729
 
728
730
  For Claude Code, it also adds:
729
- - a `SessionEnd` hook so `hippo sleep` runs automatically when the session exits
731
+ - a `SessionEnd` hook that runs `hippo sleep` and then `hippo capture` on the session's transcript when the session exits
730
732
  - a `SessionStart` hook that prints the previous session's consolidation output
731
733
  - a `UserPromptSubmit` hook that runs `hippo context --pinned-only --include-recent 5 --format additional-context` every turn. It re-injects pinned memories (`hippo remember <text> --pin`) plus the last 5 writes, so fresh same-session lessons appear on the next prompt before you pin them. The block is rendered without live strength percentages, so it stays byte-identical while its memories do not change, and it is sent only when it changed since the session's last prompt: an unchanged block is skipped, resent every 10 skips (`pinnedInject.refreshTurns`, `0` never resends) and resent after compaction. `{"pinnedInject":{"skipUnchanged":false}}` sends it every turn as before. Opt out entirely with `{"pinnedInject":{"enabled":false}}` in `.hippo/config.json`.
732
734
  - a `PreCompact` hook that runs `hippo pre-compact` before the transcript gets summarized. It saves a working-state snapshot (task/summary/next step) so mid-session compaction can't drop it; the `SessionEnd` hook still owns extracting durable memories.
@@ -738,22 +740,31 @@ To remove: `hippo hook uninstall claude-code`
738
740
 
739
741
  ### What the hook adds (Claude Code example)
740
742
 
741
- ```markdown
743
+ ````markdown
742
744
  ## Project Memory (Hippo)
743
745
 
744
- Before starting work, load relevant context:
746
+ Pinned rules and recent writes auto-inject at every prompt via the installed
747
+ UserPromptSubmit hook; never re-run that part manually. At the START of a
748
+ task (not per prompt), additionally load task-specific context: git-aware
749
+ recall over the full store that per-prompt injection does not cover. Also
750
+ run it if the hook is not installed:
751
+ ```bash
745
752
  hippo context --auto --budget 1500
753
+ ```
746
754
 
747
- When you hit an error or discover a gotcha:
755
+ When you find out why something failed, record it right then, while you
756
+ work, never as a closing step:
757
+ ```bash
748
758
  hippo remember "<what went wrong and why>" --error
749
-
750
- After completing work successfully:
751
- hippo outcome --good
752
759
  ```
753
760
 
761
+ The installed hooks store failed tool calls and capture the session when it
762
+ ends, so there is nothing to run before you finish.
763
+ ````
764
+
754
765
  ### MCP Server
755
766
 
756
- For any MCP-compatible client (Cursor, Windsurf, Cline, Claude Desktop):
767
+ For any MCP-compatible client (Cursor, Windsurf (now Devin Desktop), Cline, Claude Desktop):
757
768
 
758
769
  ```bash
759
770
  hippo mcp # starts MCP server over stdio
@@ -850,12 +861,12 @@ The AI-memory category matured fast in 2026. Hippo's specific take — bio-decay
850
861
  | Auto-hook install | Yes | No | No | No | No | No | No | No | No | No |
851
862
  | MCP server | Yes | Yes | Yes (hosted, needs an account) | Yes | Yes (stdio + HTTP/OAuth) | Yes (hosted, needs an account) | Yes (hosted, needs an API key) | Yes (first-party Claude/LangGraph) | Yes | ? |
852
863
  | Zero runtime deps | Yes | No (ChromaDB) | No | No | No (PGLite or PG+pgvector) | No (managed service) | No (npm deps) | No (Python deps) | Yes (single Rust binary) | No (managed + OSS) |
853
- | LongMemEval (best published) | 98.0% local / 99.8% voyage R@5 (s_cleaned, per-haystack)\* | 96.6% raw / 100% reranked R@5 | 94.4 (hosted platform)\*\* | N/A | 97.6-97.9% R@5 (s_cleaned\*) | 90.2% accuracy\*\* (LoCoMo 94.7%) | N/A | N/A | 88.78% overall accuracy w/ reader\*\* | 83.00% overall\*\* (LoCoMo 93.05%, HaluMem 93.04%) |
864
+ | LongMemEval (best published) | 98.0% local / 99.8% voyage any-evidence R@5; 88.5% local all-evidence R@5 (s_cleaned, per-haystack)\* | 96.6% raw / 100% reranked R@5 | 94.4 (hosted platform)\*\* | N/A | 95.53% all-evidence R@5 reranked, 93.19% without (s_cleaned\*) | 90.2% accuracy\*\* (LoCoMo 94.7%) | N/A | N/A | 88.78% overall accuracy w/ reader\*\* | 83.00% overall\*\* (LoCoMo 93.05%, HaluMem 93.04%) |
854
865
  | Git-friendly | Yes | No | No | Yes | Yes | No | Yes (memory tracked in git) | No | Yes (Git is the model) | ? |
855
866
  | Framework agnostic | Yes | Yes | Partial | Yes | Yes | Yes | Yes | Yes | Yes | Yes |
856
867
  | License | MIT | (open) | Apache-2.0 | (open) | MIT | Proprietary cloud (Graphiti: Apache-2.0) | Apache-2.0 | MIT (core) | Apache-2.0 | Apache-2.0 (OSS) + cloud |
857
868
 
858
- \* Hippo's 98.0% uses the free local MiniLM embedder (an optional install) and 99.8% uses voyage-3-large (measured 2026-06-09, not re-run). Both are on `longmemeval_s_cleaned` with a per-question haystack, the split and metric of gbrain's published 97.6%. Each is the best of five retrieval settings in the benchmark scripts, not `hippo recall`. At 500 questions the 95% interval is about ±1.2 points, so 98.0 and 97.6 are a tie. The June 2026 build scored 98.6; [`docs/evals/2026-09-23-longmemeval-reproduction.md`](docs/evals/2026-09-23-longmemeval-reproduction.md) has both runs. An older hippo number, 86.8% R@5 on `longmemeval_oracle` under pooled (non-per-haystack) retrieval, is not comparable to per-haystack figures.
869
+ \* Hippo's figures are on `longmemeval_s_cleaned` with a per-question haystack, each the best of five retrieval settings in the benchmark scripts, not `hippo recall`. Any-evidence R@5 counts a hit when any answer session is in the top 5, over all 500 questions: 98.0% with the free local MiniLM embedder (an optional install) and 99.8% with voyage-3-large (measured 2026-06-09, not re-run). All-evidence R@5 counts a hit only when every answer session is in the top 5, over the 470 questions that have an answer: 86.8 to 88.5% with MiniLM. gbrain first published 97.6%, an any-evidence score over all 500; its [report](https://github.com/garrytan/gbrain-evals/blob/main/docs/benchmarks/2026-05-07-longmemeval-s.md) now leads with all-evidence, 95.53% (449 of 470) with the paid Voyage rerank-2.5 reranker and 93.19% without it. On all-evidence recall gbrain is ahead. The June 2026 build scored 98.6 any-evidence; [`docs/evals/2026-09-23-longmemeval-reproduction.md`](docs/evals/2026-09-23-longmemeval-reproduction.md) has both runs. An older hippo number, 86.8% R@5 on `longmemeval_oracle` under pooled (non-per-haystack) retrieval, is not comparable to per-haystack figures.
859
870
 
860
871
  \*\* Different metric: these are end-to-end answer scores, not retrieval R@5. Mem0's 94.4 comes from its hosted platform, which its README says includes optimizations the open-source SDK lacks. Zep's 90.2% and 94.7% are accuracy figures from its homepage. Memoria's 88.78% and EverMind's 83% are overall accuracy with a reader LLM. Higher denominator + LLM helps. Not directly comparable to retrieval-only R@5 numbers above. The Mem0, Zep and Letta columns were last checked against each vendor's own pages on 2026-09-28.
861
872
 
@@ -878,9 +889,9 @@ Three benchmarks testing three different things. Full details in [`benchmarks/`]
878
889
  | MiniLM-L6 (local, optional install) | 96.8 | 98.0 | 88.4 |
879
890
  | voyage-3-large (opt-in, paid) | 99.8 | 99.8 | 94.6 |
880
891
 
881
- gbrain reports 97.6 R@5 on this split with a paid frontier embedder. Hippo reaches 98.0 with a free local embedder, a tie at 500 questions. These numbers come from the scripts in `benchmarks/longmemeval/`, which index every turn and fuse BM25 with dense ranks; they are not `hippo recall`, and a default install has no embedder. Re-measure: [`docs/evals/2026-09-23-longmemeval-reproduction.md`](docs/evals/2026-09-23-longmemeval-reproduction.md). Retrieval recall on the standard task is effectively saturated, so the embedder is a swappable commodity, not the differentiator. Method and the global-pool comparison: [`docs/evals/2026-06-09-longmemeval-per-haystack-dual.md`](docs/evals/2026-06-09-longmemeval-per-haystack-dual.md).
892
+ These are any-evidence scores: a hit when any answer session is in the top 5, over all 500 questions. Counting a hit only when every answer session is in the top 5, over the 470 questions that have an answer, the MiniLM runs score 86.8 to 88.5. gbrain first published 97.6 any-evidence and now reports 95.53 all-evidence with a paid reranker, 93.19 without, so on all-evidence recall gbrain is ahead. These numbers come from the scripts in `benchmarks/longmemeval/`, which index every turn and fuse BM25 with dense ranks; they are not `hippo recall`, and a default install has no embedder. Re-measure: [`docs/evals/2026-09-23-longmemeval-reproduction.md`](docs/evals/2026-09-23-longmemeval-reproduction.md). Any-evidence recall is near its ceiling on this task; all-evidence recall is not. Method and the global-pool comparison: [`docs/evals/2026-06-09-longmemeval-per-haystack-dual.md`](docs/evals/2026-06-09-longmemeval-per-haystack-dual.md).
882
893
 
883
- The differentiator is what happens as one store grows. Point retrieval at a single unified memory of tens of thousands of sessions, with no pre-scoped haystack, and recall stops being free (MiniLM 47, voyage 56 on the 19,195-session `_s` store, June 2026). That is where we expect the memory lifecycle to matter, and it is what hippo measures next (see ROADMAP Part III). It is not shown yet: in our tests so far, decay made no measurable difference and sleep cost recall. Outcome marks and supersession are what measured helpful.
894
+ The differentiator is what happens as one store grows. Point retrieval at a single unified memory of tens of thousands of sessions, with no pre-scoped haystack, and recall stops being free (MiniLM 47, voyage 56 on the 19,195-session `_s` store, June 2026). That is where we expect the memory lifecycle to matter, and it is what hippo measures next (see ROADMAP Part III). It is not shown yet: in our tests so far, decay made no measurable difference and sleep cost recall. Outcome marks and retrieval strengthening are what measured helpful; supersession is not measured yet.
884
895
 
885
896
  **Hippo v0.28.0 oracle-split results (hybrid BM25 + cosine, full 500 questions, pooled retrieval):**
886
897
 
@@ -959,11 +970,11 @@ node run.mjs --adapter all
959
970
 
960
971
  ### How do I give Claude Code memory between sessions?
961
972
 
962
- Run `npm install -g hippo-memory`, then `hippo init` in the project, or `hippo init --scan ~` for every repo on the machine. If the project has a `CLAUDE.md`, init adds a short block telling Claude to run `hippo context --auto` when a session starts. It also adds hooks to Claude Code's settings that keep your pinned memories in context, save a task snapshot before compaction, and run `hippo sleep` when the session ends. The [Claude Code plugin](https://github.com/kitfunso/hippo-memory/tree/master/extensions/claude-code-plugin) is the alternative to these hooks; use one, not both.
973
+ Run `npm install -g hippo-memory`, then `hippo init` in the project. If the project has a `CLAUDE.md`, init adds a short block telling Claude to run `hippo context --auto` when a session starts. It also adds hooks to Claude Code's settings that keep your pinned memories in context, save a task snapshot before compaction, and run `hippo sleep` when the session ends. `hippo init --scan ~` gives every git repo under your home folder a store and installs the same hooks, but adds no block to any `CLAUDE.md`. The [Claude Code plugin](https://github.com/kitfunso/hippo-memory/tree/master/extensions/claude-code-plugin) is the alternative to these hooks; use one, not both.
963
974
 
964
975
  ### How do I give Cursor memory between sessions?
965
976
 
966
- `hippo init` adds its instructions to `.cursorrules` if the project has one, and the [MCP server](#mcp-server) gives Cursor's agent tools to recall and store memories once you add `hippo mcp` to `.cursor/mcp.json`. `hippo import --cursor .cursorrules` turns your existing rules into memories.
977
+ `hippo init` adds its instructions to `AGENTS.md` if the project has one, and Cursor reads that file from the project root. The [MCP server](#mcp-server) gives Cursor's agent tools to recall and store memories once you add `hippo mcp` to `.cursor/mcp.json`. Older hippo versions wrote the block to `.cursorrules`; `hippo hook uninstall cursor` removes it from there, and from `AGENTS.md` only when the block there is Cursor's own and unedited (`hippo hook install cursor` puts it back). A block written for Codex or another agent stays, since Cursor reads it too, and so does an edited block, since hippo cannot tell whose it is. `hippo import --cursor .cursor/rules` turns your existing rules into memories; it reads an older single `.cursorrules` file too.
967
978
 
968
979
  ### How do I give Codex memory across sessions?
969
980
 
@@ -975,19 +986,27 @@ Run `npm install -g hippo-memory`, then `hippo init` in the project, or `hippo i
975
986
 
976
987
  ### Can I use hippo as an MCP memory server?
977
988
 
978
- Yes. `hippo mcp` runs the server over stdio, and `npx -y hippo-memory mcp` runs it without a global install. Add it to the MCP config of Claude Desktop, Cursor, Windsurf, Cline or any other client (example [above](#mcp-server)); in Claude Code, run `claude mcp add hippo-memory -- hippo mcp`. The agent gets tools such as `hippo_recall`, `hippo_remember` and `hippo_outcome`.
989
+ Yes. `hippo mcp` runs the server over stdio, and `npx -y hippo-memory mcp` runs it without a global install. Add it to the MCP config of Claude Desktop, Cursor, Windsurf (now Devin Desktop), Cline or any other client (example [above](#mcp-server)); in Claude Code, run `claude mcp add hippo-memory -- hippo mcp`. The agent gets tools such as `hippo_recall`, `hippo_remember` and `hippo_outcome`.
979
990
 
980
991
  ### How is hippo different from mem0?
981
992
 
982
993
  mem0 uses a language model to extract memories, OpenAI by default in its open-source library, and memories stored through its hosted MCP server live in your Mem0 account ([mem0 docs](https://docs.mem0.ai/platform/mem0-mcp), checked 2026-09-28). Hippo stores memories in SQLite on your machine, needs no account and no model, and `hippo init` wires it into the coding agents it finds. mem0's platform and hippo both mark an older fact superseded when a newer one replaces it. Hippo also lets you mark a recalled memory wrong with `hippo outcome --bad`, and it drops out of the top results.
983
994
 
995
+ ### Is this just RAG?
996
+
997
+ No. RAG searches a fixed corpus. Hippo's store changes as your agent works: a memory marked wrong drops out of the top results, a newer fact supersedes the old one, and memories that keep getting recalled last longer while unused ones fade on a half-life. Recall itself is search: BM25, plus embeddings if you install them.
998
+
999
+ ### Does it need embeddings?
1000
+
1001
+ No. Recall runs on BM25 out of the box, with no model and no network call, and a default install has no embedder. Embeddings are an optional install for hybrid search. On LongMemEval-S, where each question gets its own haystack, the benchmark scripts (not `hippo recall`) fuse BM25 with the free local MiniLM embedder and reach 98.0% recall@5, counting a hit when any answer session is in the top five. On LongMemEval's oracle split with one pooled store, BM25 alone scored 74.0% recall@5 in v0.11. The two runs use different setups, so they are not a before and after.
1002
+
984
1003
  ### Do I still need CLAUDE.md?
985
1004
 
986
1005
  Yes, for short standing rules such as build commands, code style and things never to do. Claude Code loads `CLAUDE.md` and its auto memory into every session, and its [memory docs](https://code.claude.com/docs/en/memory) say that when two rules contradict each other, Claude may pick one arbitrarily. Hippo holds the lessons that pile up, recalls the ones that match the task, and retires the ones marked wrong or replaced. `hippo init` adds its block to `CLAUDE.md`, and `hippo import --claude CLAUDE.md` turns existing notes into memories.
987
1006
 
988
1007
  ### What happens when a memory turns out to be wrong?
989
1008
 
990
- Mark it, and it stops coming back. `hippo outcome --bad` weakens the memories from the last recall, `hippo supersede <id> "<new fact>"` replaces one with a newer version, and `hippo reject <id> --reason "<why>"` stops that value from returning at all. On the synthetic E1 test, where every mark is correct, plain BM25 plus the outcome mark cut how often a marked-bad memory stayed in the top five from 71.9% to 0.0%. Real marks are noisier, because `--bad` marks the whole recall batch.
1009
+ Mark it, and it drops out of the top results. `hippo outcome --bad` weakens the memories from the last recall, `hippo supersede <id> "<new fact>"` replaces one with a newer version, and `hippo reject <id> --reason "<why>"` stops that value from returning at all. On the synthetic E1 test, where every mark is correct, plain BM25 plus the outcome mark cut how often a marked-bad memory stayed in the top five from 71.9% to 0.0%. Real marks are noisier, because `--bad` marks the whole recall batch.
991
1010
 
992
1011
  ### Where does hippo keep my data?
993
1012
 
@@ -997,9 +1016,13 @@ On your machine, in SQLite: `.hippo/hippo.db` in each project, plus a global sto
997
1016
 
998
1017
  Nothing. Hippo is MIT-licensed and needs no account or API key. Optional features that call an outside provider bill through it: the Jev reranker costs about 0.0004 USD a recall, and API embedders and sleep's fact extraction bill your own keys. Memory text handed to your agent uses context tokens, and `hippo tokens` shows how many.
999
1018
 
1019
+ ### Is it production-ready?
1020
+
1021
+ Judge it by what is tested. 3,500+ tests run against a real database, with no module mocks and no mocked store, and a negative test checks that one tenant cannot read another's memories. It is MIT-licensed and has zero runtime dependencies. What has not been shown yet is whether agents do better work with it: the published numbers measure retrieval.
1022
+
1000
1023
  ### Has hippo been shown to make agents better at their work?
1001
1024
 
1002
- Not yet. The published numbers measure retrieval: whether the right memory comes back, and whether a memory marked wrong stays out of the results. The paired test that runs real agent sessions with and without hippo has not had a scored run yet. Every measurement, including failed runs and one retracted claim, is indexed in [docs/evals](https://github.com/kitfunso/hippo-memory/blob/master/docs/evals/README.md).
1025
+ Not yet. The published numbers measure retrieval: whether the right memory comes back, and whether a memory marked wrong stays out of the results. A paired test that runs real agent sessions with and without hippo is under way. Every measurement, including failed runs and one retracted claim, is indexed in [docs/evals](https://github.com/kitfunso/hippo-memory/blob/master/docs/evals/README.md).
1003
1026
 
1004
1027
  ---
1005
1028
 
@@ -1008,7 +1031,7 @@ Not yet. The published numbers measure retrieval: whether the right memory comes
1008
1031
  Issues and PRs welcome. Before contributing, run `hippo status` in the repo root to see the project's own memory.
1009
1032
 
1010
1033
  The interesting problems:
1011
- - **LongMemEval retrieval (standard task: done).** Per-question-haystack R@5 is 98.0% with the free local embedder (re-measured 2026-09-23; the June build gave 98.6) and 99.8% with voyage-3-large (measured 2026-06-09), level with the published frontier. Retrieval is no longer the gap; the lifecycle stress eval (ROADMAP Part III) is the next measurement.
1034
+ - **LongMemEval retrieval (standard task: done).** Per-question-haystack R@5 is 98.0% with the free local embedder (re-measured 2026-09-23; the June build gave 98.6) and 99.8% with voyage-3-large (measured 2026-06-09), counting a hit when any answer session is in the top 5. That measure is near its ceiling. The all-evidence one, every answer session in the top 5 over the 470 questions with an answer, is not; there the MiniLM runs score 86.8 to 88.5% against gbrain's 95.53%. The lifecycle stress eval (ROADMAP Part III) is the next measurement.
1012
1035
  - Better consolidation heuristics (LLM-powered merge vs current text overlap)
1013
1036
  - Web UI / dashboard for visualizing decay curves and memory health
1014
1037
  - Optimal decay parameter tuning from real usage data
package/dist/api.js CHANGED
@@ -1180,6 +1180,7 @@ export function supersede(ctx, oldId, newContent) {
1180
1180
  confidence: 'verified',
1181
1181
  tenantId: ctx.tenantId,
1182
1182
  scope: old.scope,
1183
+ baseHalfLifeDays: loadConfig(ctx.hippoRoot).defaultHalfLifeDays,
1183
1184
  });
1184
1185
  // Race-safe transition: open a fresh db handle, BEGIN IMMEDIATE, run all
1185
1186
  // three steps (CAS on old + writeEntryDbOnly(new) + supersede audit row)
@@ -2093,7 +2094,7 @@ export function restoreDormant(ctx, id) {
2093
2094
  // (The placeholder only satisfies createMemory's minimum length, so a
2094
2095
  // legacy row shorter than 3 chars can still be restored.)
2095
2096
  const revived = {
2096
- ...createMemory('dormant snapshot defaults'),
2097
+ ...createMemory('dormant snapshot defaults', { baseHalfLifeDays: loadConfig(ctx.hippoRoot).defaultHalfLifeDays }),
2097
2098
  ...dormant.entry,
2098
2099
  last_retrieved: now.toISOString(),
2099
2100
  };
@@ -3,10 +3,7 @@
3
3
  * Agents learn from failures without explicit hippo remember calls.
4
4
  */
5
5
  import { MemoryEntry } from './memory.js';
6
- /**
7
- * Create a MemoryEntry capturing a command failure.
8
- * Content format: "Command '<cmd>' failed: <truncated stderr>"
9
- */
6
+ /** A memory of a failed command, "Command '<cmd>' failed: <truncated stderr>"; no store is in reach, so `hippo watch` re-derives its half-life from the store's config. */
10
7
  export declare function captureError(exitCode: number, stderr: string, command: string, tenantId?: string): MemoryEntry;
11
8
  /**
12
9
  * Parse git log output for actionable lessons.
package/dist/autolearn.js CHANGED
@@ -3,14 +3,11 @@
3
3
  * Agents learn from failures without explicit hippo remember calls.
4
4
  */
5
5
  import { execSync, execFileSync, spawn } from 'child_process';
6
- import { createMemory, Layer } from './memory.js';
6
+ import { createMemory, Layer, DEFAULT_HALF_LIFE_DAYS } from './memory.js';
7
7
  import { loadAllEntries } from './store.js';
8
8
  import { textOverlap } from './search.js';
9
9
  import { isContentWorthStoring } from './audit.js';
10
- /**
11
- * Create a MemoryEntry capturing a command failure.
12
- * Content format: "Command '<cmd>' failed: <truncated stderr>"
13
- */
10
+ /** A memory of a failed command, "Command '<cmd>' failed: <truncated stderr>"; no store is in reach, so `hippo watch` re-derives its half-life from the store's config. */
14
11
  export function captureError(exitCode, stderr, command, tenantId) {
15
12
  // Truncate to first 500 chars to avoid storing megabytes of build logs
16
13
  const wasTruncated = stderr.length > 500;
@@ -30,6 +27,7 @@ export function captureError(exitCode, stderr, command, tenantId) {
30
27
  source: 'autolearn',
31
28
  confidence: 'observed',
32
29
  tenantId,
30
+ baseHalfLifeDays: DEFAULT_HALF_LIFE_DAYS,
33
31
  });
34
32
  }
35
33
  /**
package/dist/capture.js CHANGED
@@ -21,6 +21,7 @@ import { defaultPreCompactLogPath } from './hooks.js';
21
21
  import { redactSecrets } from './secret-detect.js';
22
22
  import { RejectedValueError, checkRejectionGuard } from './rejection.js';
23
23
  import { openHippoDb, closeHippoDb } from './db.js';
24
+ import { loadConfig } from './config.js';
24
25
  // Sentence-level patterns
25
26
  //
26
27
  // T1 (DF2): each pattern now carries TWO capture groups — group 1 is the
@@ -790,6 +791,7 @@ function cmdCaptureCore(hippoRoot, options) {
790
791
  let captured = 0;
791
792
  let skipped = 0;
792
793
  let rejected = 0;
794
+ const baseHalfLifeDays = loadConfig(targetRoot).defaultHalfLifeDays;
793
795
  // AT1 P2 fix (dry-run parity, docs/plans/2026-08-15-at1-rejected-value-tombstone.md):
794
796
  // dry-run used to skip the guarded write branch ENTIRELY, so a tombstoned
795
797
  // extraction printed as `[capture]` and counted toward `captured` — the
@@ -823,6 +825,7 @@ function cmdCaptureCore(hippoRoot, options) {
823
825
  source: 'capture',
824
826
  confidence: 'observed',
825
827
  tenantId: useGlobal ? undefined : options.tenantId,
828
+ baseHalfLifeDays,
826
829
  });
827
830
  if (options.dryRun) {
828
831
  if (dryRunDb) {