hermes-memory-pgvector 0.5.2__tar.gz → 0.5.3__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (41) hide show
  1. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/PKG-INFO +499 -432
  2. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/README.md +466 -399
  3. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/hermes_memory_pgvector.egg-info/PKG-INFO +499 -432
  4. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/hermes_memory_pgvector.egg-info/SOURCES.txt +2 -0
  5. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/hermes_pgvector/__init__.py +1612 -1541
  6. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/hermes_pgvector/__main__.py +471 -450
  7. hermes_memory_pgvector-0.5.3/hermes_pgvector/embed.py +311 -0
  8. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/hermes_pgvector/identity.py +192 -192
  9. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/hermes_pgvector/migrations/001_schema.sql +95 -95
  10. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/hermes_pgvector/migrations/002_agent_attribution.sql +130 -130
  11. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/hermes_pgvector/migrations/003_hybrid_search_fts.sql +51 -51
  12. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/hermes_pgvector/migrations/004_runtime_grants.sql +36 -36
  13. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/hermes_pgvector/plugin.yaml +16 -16
  14. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/hermes_pgvector/store.py +1259 -1255
  15. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/hermes_pgvector/writer.py +192 -192
  16. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/pyproject.toml +83 -83
  17. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/setup.cfg +4 -4
  18. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/tests/test_async_writer.py +110 -110
  19. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/tests/test_config_coercion.py +112 -112
  20. hermes_memory_pgvector-0.5.3/tests/test_embed_config.py +610 -0
  21. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/tests/test_embed_timeouts.py +249 -249
  22. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/tests/test_empty_content.py +332 -332
  23. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/tests/test_exclude_identities_live.py +197 -197
  24. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/tests/test_hybrid_search.py +123 -123
  25. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/tests/test_identity.py +298 -298
  26. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/tests/test_install_shim.py +104 -104
  27. hermes_memory_pgvector-0.5.3/tests/test_loader_embed_clobber.py +188 -0
  28. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/tests/test_read_side_gate.py +187 -187
  29. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/tests/test_save_config_merge.py +68 -68
  30. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/tests/test_session_switch_contract.py +140 -140
  31. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/tests/test_smoke.py +316 -316
  32. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/tests/test_store_v04.py +246 -246
  33. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/tests/test_system_prompt_block.py +51 -51
  34. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/tests/test_tool_args_hardening.py +128 -128
  35. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/tests/test_turn_dedup.py +280 -280
  36. hermes_memory_pgvector-0.5.2/hermes_pgvector/embed.py +0 -156
  37. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/LICENSE +0 -0
  38. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/hermes_memory_pgvector.egg-info/dependency_links.txt +0 -0
  39. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/hermes_memory_pgvector.egg-info/entry_points.txt +0 -0
  40. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/hermes_memory_pgvector.egg-info/requires.txt +0 -0
  41. {hermes_memory_pgvector-0.5.2 → hermes_memory_pgvector-0.5.3}/hermes_memory_pgvector.egg-info/top_level.txt +0 -0
@@ -1,432 +1,499 @@
1
- Metadata-Version: 2.4
2
- Name: hermes-memory-pgvector
3
- Version: 0.5.2
4
- Summary: Postgres + pgvector memory provider plugin for hermes-agent. Multi-agent storage layer with per-minion themes, identity governance, agent attribution + delegation provenance, async writer, no LLM in the memory hot path.
5
- Author: Andrea Borghi
6
- License: BSD-3-Clause
7
- Project-URL: Homepage, https://github.com/andreab67/hermes-memory-pgvector
8
- Project-URL: Source, https://github.com/andreab67/hermes-memory-pgvector
9
- Project-URL: Issues, https://github.com/andreab67/hermes-memory-pgvector/issues
10
- Project-URL: Roadmap, https://github.com/andreab67/hermes-memory-pgvector/blob/main/ROADMAP.md
11
- Keywords: hermes-agent,memory-provider,pgvector,postgres,multi-agent,llm-memory,semantic-search
12
- Classifier: Development Status :: 4 - Beta
13
- Classifier: Intended Audience :: Developers
14
- Classifier: Intended Audience :: System Administrators
15
- Classifier: License :: OSI Approved :: BSD License
16
- Classifier: Operating System :: POSIX :: Linux
17
- Classifier: Programming Language :: Python :: 3
18
- Classifier: Programming Language :: Python :: 3.11
19
- Classifier: Programming Language :: Python :: 3.12
20
- Classifier: Programming Language :: Python :: 3.13
21
- Classifier: Topic :: Database
22
- Classifier: Topic :: Software Development :: Libraries :: Python Modules
23
- Classifier: Topic :: System :: Distributed Computing
24
- Requires-Python: >=3.11
25
- Description-Content-Type: text/markdown
26
- License-File: LICENSE
27
- Requires-Dist: psycopg[binary]<4,>=3.3.5
28
- Requires-Dist: psycopg-pool<4,>=3.3.1
29
- Requires-Dist: PyYAML<7,>=6.0
30
- Provides-Extra: test
31
- Requires-Dist: pytest<9,>=7.4; extra == "test"
32
- Dynamic: license-file
33
-
34
- # hermes-memory-pgvector
35
-
36
- **Postgres + pgvector memory provider for [hermes-agent](https://github.com/NousResearch/hermes-agent).** A shared memory substrate for a fleet of cooperating hermes-agent minions — built on Postgres and a single embedding endpoint you probably already run, with no LLM in the memory hot path.
37
-
38
- ```text
39
- each minion → X-Hermes-Session-Key: <theme>
40
- → hermes-agent gateway
41
- → pgvector plugin
42
- ├── memory_entries (mirrors built-in MEMORY.md / USER.md per theme)
43
- └── conversations (every substantive turn, semantically searchable)
44
- ```
45
-
46
- ## Why it exists
47
-
48
- Existing memory providers each solve a piece of the problem; the gap for **fleet deployments** is wide:
49
-
50
- - **Built-in `memory` tool** persists to per-host `MEMORY.md` / `USER.md`. Two minions on the same host stomp on each other; minions on different hosts have no shared substrate.
51
- - **Honcho** offers cross-session user modelling but requires a full external service, an LLM in the memory hot path for its deriver + dialectic loops, and its own ontology layered on top of the built-in tool. In high-concurrency fleet use it produces retry storms, embedding-endpoint queue backups, and gateway↔Honcho circular dependencies.
52
- - **Holographic** is a fine in-process fact store but uses SQLite — a poor fit for many minions writing concurrently from many hosts.
53
- - **Other providers** (Mem0, Hindsight, OpenViking, ByteRover, RetainDB, Supermemory) all either require a paid cloud, require LLM mediation for memory ops, or both.
54
-
55
- What was missing: a **storage layer** that gives the built-in `memory` model durable, multi-tenant, semantically-searchable backing, with no LLM in the hot path, scoped cleanly per-minion so a marketing agent's notes don't pollute a trading agent's recall. That's what this plugin provides.
56
-
57
- ## Design philosophy
58
-
59
- 1. **Storage layer, not a memory model.** The agent keeps using `memory(action='add', target='memory'|'user', …)`. We mirror those writes via `on_memory_write`. No new ontology for the agent to learn.
60
- 2. **No LLM in the memory hot path.** Embeddings are vector math, not LLM calls. There is no deriver, no dialectic, no dream cycle — the failure modes that hurt Honcho cannot occur here by construction.
61
- 3. **Per-agent themes by default, cross-theme recall on explicit demand.** Every row carries `agent_identity` (resolved from `X-Hermes-Session-Key` header, profile name, workspace, or `'default'`). Recall is scoped to the current theme unless the agent asks for `scope='all'`.
62
- 4. **Fail-soft everywhere.** Embed endpoint down → degrade to text-only writes. Async writer queue full → drop with a one-time warning. DB down → log + skip. No exception escapes into the agent loop.
63
- 5. **Admin/runtime separation.** DDL (`CREATE EXTENSION vector`, `CREATE TABLE`, `CREATE INDEX`) runs once with superuser. The runtime user has DML only on the migrated schema. `ensure_schema()` at runtime is verify-only with a clear `SchemaNotApplied` error if the operator forgot the migration.
64
-
65
- ## Features (v0.3.0)
66
-
67
- | Hook / surface | Behavior |
68
- |---|---|
69
- | `initialize()` | Verifies schema, opens `psycopg_pool.ConnectionPool`, bulk-imports existing `MEMORY.md` + `USER.md` content. |
70
- | `on_memory_write(action, target, content, meta)` | Mirrors built-in `memory` writes into `memory_entries` (add / replace / remove). |
71
- | `sync_turn(user, assistant, session_id)` | Captures every substantive (`>= 40` chars + not boilerplate) chat turn into `conversations`. |
72
- | `prefetch(query)` | Top-K semantically similar `memory_entries` in current theme, injected ambient. |
73
- | `recall_memory(query, scope, target, limit)` tool | Explicit cross-theme search of durable memory entries. |
74
- | `recall_conversation(query, scope, limit)` tool | Explicit search over past chat turns. `scope ∈ {current, session, all, <theme>}`. |
75
-
76
- Internals:
77
-
78
- - **`psycopg_pool.ConnectionPool`** (min=0, max=4, lazy + thread-safe, `max_idle=30s` / `max_lifetime=300s`) shared across the agent thread and the async-writer drain thread. `min_size=0` keeps an idle — or abandoned — pool at **zero** open connections, so a session the gateway never explicitly shuts down cannot strand a Postgres backend (see *Fixed in v0.3.1* below).
79
- - **`AsyncWriter`** — bounded queue + daemon drain thread. Memory write hooks return in microseconds. Worker embeds + writes in the background. Crash-resilient (auto-restart on next enqueue).
80
- - **Single migration** (`hermes_pgvector/migrations/001_schema.sql`) — `memory_entries` + `conversations` + HNSW indexes. Same tuning operators typically use elsewhere.
81
- - **Boilerplate filter** for turn capture — length floor + acknowledgement regex (`"ok"`, `"thanks"`, `"continue"`, …) so the recall table stays high-signal.
82
-
83
- ### Fixed in v0.3.1 — connection-leak hotfix
84
-
85
- A single registered provider has `initialize()` called again for each new session. It previously
86
- reassigned `self._store` / `self._writer` without closing the prior ones, **abandoning a
87
- `ConnectionPool`** whose warm (`min_size=1`) connection lingered in Postgres — committed-but-idle —
88
- until the server's `idle_session_timeout`. Under a burst of concurrent sessions (e.g. a swarm of
89
- systemd-run minions firing on the same minute) these orphaned backends saturated the database's
90
- connection slots. Fixed by:
91
-
92
- 1. **`initialize()` teardown** — drain the prior `AsyncWriter` + close the prior pool before
93
- re-initializing (the call is idempotent and skipped on first init).
94
- 2. **Self-draining pool** — `min_size=0` (an idle or abandoned pool holds *zero* connections) plus
95
- `max_idle=30s` / `max_lifetime=300s`, so connections are short-lived when idle and pooled only
96
- under active load.
97
-
98
- ## New in v0.4.0
99
-
100
- Four capabilities, all storage-layer (still no LLM in the hot path):
101
-
102
- - **Identity governance.** The resolved `agent_identity` is normalized once at init: direct-message session keys like `agent:main:whatsapp:dm:<phone>` collapse to a single `whatsapp-dm` bucket (no PII, no per-contact theme explosion), benchmark traffic (`skill-bench*`) is isolated to `_bench`, and an optional `allowed_themes` allow-list routes typo'd/unknown themes to `default`. The M3 resolution priority is preserved — normalization runs *after* the chain, never at read time. See [`hermes_pgvector/identity.py`](hermes_pgvector/identity.py).
103
- - **Agent attribution + delegation (M4).** Migration `002` adds `memory_agents` (registry) and `memory_agent_edges` (parent→child delegation provenance) plus `conversations.parent_session_id`. The `on_delegation` / `on_session_end` hooks capture which agent delegated what to whom — strictly enqueue-only and fail-soft. **Provenance only** (who/when), never a fact-store ontology. Query it via the `v_agent_memory` view.
104
- - **Embedding backfill + writer resilience.** Rows written text-only during an embed-endpoint outage are no longer permanently unsearchable: `hermes-pgvector backfill` re-embeds `NULL`-embedding rows (idempotent, 768-dim-guarded). The background writer gains a small bounded retry; the hot path stays single-attempt.
105
- - **Conversation TTL + embed policy.** `hermes-pgvector prune --days N` trims old turns (operator-triggered only; `memory_entries` are never pruned). `conversation_embed_policy` (`all` default / `substantive_only` / `none`) tunes embedding cost.
106
-
107
- Maintenance CLI (`hermes-pgvector`, or `python -m hermes_pgvector`): `migrate · stats · backfill · prune · cleanup · remap` — destructive commands default to dry-run. v0.4.0 is a clean upgrade from v0.3.x: apply migration `002` to light up attribution/delegation; without it the new hooks no-op and everything else runs unchanged.
108
-
109
- ## New in v0.4.1 — hybrid recall (vector + full-text)
110
-
111
- `recall_memory` and `recall_conversation` now fuse the HNSW **vector** ranking with a Postgres **full-text** ranking using **Reciprocal Rank Fusion** (RRF, `k=60`). A row surfaces if *either* ranker likes it, which fixes the two blind spots of pure cosine similarity:
112
-
113
- - **Exact-lexical hits** the embedding smooths away — a specific error code, hostname, flag name, or rare identifier the agent quotes verbatim.
114
- - **Text-only rows with a `NULL` embedding** (written while the embed endpoint was down) — invisible to the vector index, but the full-text leg finds them. So hybrid recall doubles as best-effort recovery until the next `backfill`.
115
-
116
- Still a storage-layer feature: **no LLM, no entity graph, no new tables or columns** — just a GIN index over the existing `content` column (migration `003`) and a fused query. It stays inside invariant #1 (a second index over the same text is not a parallel ontology). Fail-soft as ever: a hybrid hiccup degrades to the proven pure-vector path, and a query that *itself* fails to embed degrades to full-text-only instead of erroring. Toggle with `plugins.pgvector.hybrid_search` (default `true`); the ambient `prefetch()` path stays pure-vector. Works without migration `003` — the GIN index only makes the full-text leg faster.
117
-
118
- ## New in v0.4.2 — pip-native install + hardening
119
-
120
- - **`hermes-pgvector install`** — makes a plain `pip install hermes-memory-pgvector` deployable on ANY hermes-agent install: generates the `$HERMES_HOME/plugins/pgvector/` discovery shim (see *Install · Option 1*). No more vendored copies or editable checkouts.
121
- - **Migration `004`** — `hermes-pgvector migrate` now grants the runtime role DML on `memory_entries`/`conversations` itself; the manual OWNER-transfer step is gone (fresh installs previously hit `permission denied` if it was skipped).
122
- - **Correctness fixes** from a full-codebase review: `replace`/`remove` now match `old_text` as a *literal* substring (LIKE `%`/`_`/`\` metacharacters no longer over- or under-match — parity with the built-in tool's `in` semantics); the async writer drains its queue on shutdown instead of silently abandoning up to 255 accepted writes when full; a wrong-dimension embed model now surfaces as `expected 768 dims, got N` instead of a masking 404; DM-key bucketing no longer sweeps ordinary `:signal:`-containing theme names into `whatsapp-dm`; bulk MEMORY.md import circuit-breaks after 3 consecutive embed failures (a hanging endpoint can no longer block session start for minutes); `remap` re-checks its duplicate-drop guard under the advisory lock; tool errors redact credential-looking fragments and preserve `score: null` for full-text-only hybrid hits (with `rrf_score` now included); `recall_memory(scope='session')` returns a helpful error instead of silently matching nothing.
123
-
124
- ## New in v0.4.3 — psycopg 3.3.5 floor
125
-
126
- - **Dependency floor raised**: `psycopg[binary]>=3.3.5` (upstream bugfix release, 2026-08-31: prepared-statement invalidation on `ALTER`/`DISCARD`, DataError fixes for malformed COPY/jsonb data, client-encoding aliases). No code changes.
127
-
128
- ## New in v0.5.0 — import rename (BREAKING), read-side identity gate, config-contract fixes
129
-
130
- > ### Breaking upgrade — read before installing
131
- >
132
- > **1. The import package is renamed** `pgvector` -> `hermes_pgvector`. Unchanged:
133
- > the distribution (`hermes-memory-pgvector`), the CLI (`hermes-pgvector`), and
134
- > the hermes provider name (`pgvector`, i.e. `memory.provider: pgvector`).
135
- >
136
- > Why: the old top-level name is owned by
137
- > [pgvector-python](https://pypi.org/project/pgvector/). Both in one venv meant
138
- > whichever installed last won, and this plugin's shim could import the wrong
139
- > module — taking the fleet's shared memory offline on a single log line.
140
- >
141
- > **2. Any `python -m pgvector ...` command breaks.** It is now
142
- > `hermes-pgvector ...` (or `python -m hermes_pgvector ...`). This matters most
143
- > for scheduled jobs, which fail *silently* — the nightly backfill simply stops,
144
- > and rows written during an embed outage stay permanently unsearchable. Check
145
- > your units before upgrading:
146
- >
147
- > ```bash
148
- > sudo grep -rl 'python -m pgvector' /etc/systemd/system/ /etc/cron.d/ 2>/dev/null
149
- > # e.g. hermes-pgvector-backfill.service:
150
- > # ExecStart=.../python -m pgvector backfill -> -m hermes_pgvector backfill
151
- > sudo systemctl daemon-reload
152
- > ```
153
- >
154
- > **3. Upgrade steps.** An existing shim still reads `from pgvector import ...`;
155
- > after upgrading it fails, and the loader treats that as "plugin absent" and
156
- > falls back to built-in memory.
157
- >
158
- > ```bash
159
- > pip install -U hermes-memory-pgvector==0.5.0
160
- >
161
- > # Preferred, if your hermes-agent reads the `hermes_agent.memory_providers`
162
- > # entry-point group (this package now declares it): drop the shim entirely and
163
- > # let pip discovery take over -- nothing left to go stale on future upgrades.
164
- > hermes-pgvector install --remove
165
- >
166
- > # Otherwise (older host that only scans plugin directories): regenerate it.
167
- > hermes-pgvector install --force
168
- >
169
- > # restart hermes, then verify -- do not skip this:
170
- > hermes memory status # expect: Provider: pgvector; Status: available
171
- > ```
172
- >
173
- > `hermes-pgvector install` now verifies in a clean subprocess that the shim it
174
- > just wrote can actually be imported, and warns if it cannot or if another
175
- > package shadows this one.
176
-
177
- - **Read-side identity gate.** The `whatsapp-dm` and `_bench` sinks were write-side only: `identity.py` stripped PII from the *identity*, but message bodies still live in `content`, and nothing filtered them on read. Any theme could pull DM content into its context via `scope='all'` or by naming the bucket directly — and with turn capture on, the reply quoting it was written back under the *reading* theme, permanently re-attributing DM data. `scope='all'` now excludes those sinks, and naming one explicitly is rejected. An agent that *is* the bucket keeps full access to its own rows, and ordinary cross-theme recall is unaffected. **Scope of the gate:** it excludes by bucket *name*, and bucketing happens at write time — rows are never retroactively rewritten (that is a deliberate invariant: historical rows keep their identity or they become unrecallable). So any row written *before* its key was bucketed still carries the raw identity and is still reachable. Verified on the reference deployment: **zero** such rows exist there (`whatsapp-dm` is already bucketed and no group traffic predates this release). If yours has them, remap them — and note the destructive commands default to **dry-run**, so the first form only *reports*: `hermes-pgvector remap --old <raw> --new whatsapp-dm` to preview, then re-run with `--execute` to actually move the rows (add `--force` if more than 10 duplicates would be dropped). Without `--execute` nothing moves, and it is easy to believe the rows were bucketed when they were not.
178
-
179
-
180
- Correctness release from a full-codebase review. **No schema changes and no new migrations** — but this release is not drop-in: the import package is renamed (see the upgrade box above), and `MemoryStore.search` / `hybrid_search` / `search_turns` / `hybrid_search_turns` gain an `exclude_identities` parameter. The hermes provider name, the CLI, and the on-disk schema are unchanged.
181
-
182
- - **Embed timeouts are configurable, and split by call path.** `timeout` was never plumbed from config at all — every caller silently took a hardcoded 10s. On an endpoint that answers in 6–17s that means a large share of background writes time out, fail soft, and land as rows with a NULL embedding, invisible to recall until `hermes-pgvector backfill` repairs them. Retries could not help: every attempt was capped *below* the latency the endpoint needs. There are now two keys, deliberately asymmetric — `embed_timeout` (default `10.0`) for the agent thread, where a timeout degrades recall to full-text-only and waiting longer would be worse; and `embed_write_timeout` (default `30.0`) for the background writer, where nothing is waiting and giving up costs a permanently unsearchable row. Measured against the reference endpoint: 1/3 writes succeeded at 10s, 3/3 at 30s.
183
- - **Group / channel / thread keys are bucketed.** Multi-party session keys like `agent:main:whatsapp:group:<chat>:<participant>` previously passed through untouched, so with the host's default `group_sessions_per_user` the trailing participant id — a phone number on WhatsApp/SMS/Signal — was stored verbatim as an `agent_identity`. That is the same PII failure the DM bucket exists to prevent, reached through a different `chat_type`. They now collapse to a single `external-group` bucket, which (like `whatsapp-dm` and `_bench`) is excluded from other themes' recall.
184
- - **`allowed_themes` accepts a string again.** The config schema declared it a scalar string while `normalize_identity()` consumed it as a list of names — so a string allow-list was iterated *character by character*, every theme failed the membership test, and the whole fleet was silently routed to `default`. Governance looked configured while doing the opposite. Comma-separated strings and YAML lists both work now.
185
- - **Boolean toggles honour `false` again.** `embed_on_write`, `sync_turns`, `hybrid_search` and `bulk_sync_on_init` are declared by the config schema as the *strings* `"true"`/`"false"`, but were read with plain truthiness — and `bool("false")` is `True`, so turning any of them off via that path did nothing.
186
- - **Conversation turns are no longer written twice.** `sync_turn()` (per exchange) and `on_session_end()` (whole transcript, again on session rotation) both captured the same turns, and `conversations` has no unique constraint. `on_session_end` is now a true backstop: it skips turns already accepted by the writer, and still re-captures ones a full queue dropped.
187
- - **`memory` replace mirrors correctly.** `replace()` issued one bulk `UPDATE` across every substring match, colliding with `UNIQUE(agent_identity, target, content)` — so a `replace` matching two or more entries raised `UniqueViolation` and updated **zero** rows. It now updates the first match, matching the built-in tool.
188
- - **A dead database is no longer silent.** Worker write failures logged at `debug` only, so a Postgres restart after a healthy init discarded every durable write for the rest of the session with no operator signal. The first failure now warns. Relatedly, `system_prompt_block` no longer tells the model "Empty store" when the count query merely *failed*.
189
- - **`save_config` stops deleting your settings.** It replaced the whole `plugins.pgvector` block with schema-declared keys, silently dropping hand-edited ones that are read at runtime (`identity_aliases`, `embed_write_backoff`). It merges now.
190
- - **Fail-soft hardening** (invariant #4): `sync_turn` is wrapped, config casts are guarded, and the recall tools coerce non-string `query`/`scope`/`target` instead of raising `AttributeError` out of the hook. `hermes-pgvector install --remove` now fails closed and requires `--force` on a directory that isn't a generated shim, instead of deleting it outright.
191
-
192
- ## New in v0.5.1 — `memory remove` no longer wipes a whole theme
193
-
194
- Patch release, but **upgrade promptly**: it fixes a data-loss bug. No schema changes, no migrations, no API changes.
195
-
196
- - **`remove` deleted every mirrored entry for a theme, not one.** `_worker` passed `old_text=item.content` — but the built-in tool's remove op carries its target in `old_text` and leaves `content` empty, and the host forwards `old_text` via *metadata*. So `item.content` was always `""`, `store.remove` built `content LIKE '%%'`, and that matches every row: a single `memory remove` deleted the entire mirror for that `(agent_identity, target)`. Verified against Postgres — `DELETE … WHERE c LIKE '%%'` removes all rows. `_worker` now reads `extra["old_text"]`, and `store.remove()` **refuses an empty pattern outright**, so no caller can reach that delete by omission. `remove()` also now deletes **at most one row** (lowest id), matching both the built-in tool — which requires a unique match and errors on ambiguity — and this class's own `replace()`. (The built-in store was never affected; only the pgvector mirror. No loss occurred on the reference deployment: every theme's history is continuous.)
197
-
198
- Found in production: one `memory_entries` row sat with a NULL embedding and **zero-length content**, arrived through the built-in tool's `replace` path. Two defects met there.
199
-
200
- - **Nothing rejected empty content on write.** `on_memory_write` filtered on `target` and `action` but never on content, so an `add`/`replace` carrying nothing created a row that can never be embedded — `embed()` raises `EmbeddingError("empty input")` unconditionally for empty or whitespace text. Such writes are now ignored (`remove` is exempt: it legitimately arrives with empty content and targets the row via `old_text`).
201
- - **`backfill_null_embeddings` retried it forever.** The sweep selected `WHERE embedding IS NULL` with no content filter, so every nightly run re-fetched the row, called `embed()`, failed, and moved on — permanently pinning `failed` above zero and making `remaining == 0` unreachable. That is the damaging half: it destroys the one signal an operator watches, because you can no longer distinguish a permanently-stuck row from a new genuine failure. Un-embeddable rows are now **skipped and reported separately** as `unembeddable` (skipping them silently would be equally misleading), so `remaining` can actually reach zero again.
202
-
203
- ## New in v0.5.2 — documentation only
204
-
205
- **No code changes.** `git diff v0.5.1..v0.5.2` touches only `README.md` and one test file; nothing under `hermes_pgvector/` differs, so the installed behaviour is byte-for-byte identical to v0.5.1. There is no reason to redeploy for this release.
206
-
207
- It exists because PyPI renders a project's README **frozen at upload time**: two fixes that landed after v0.5.1 shipped were visible on GitHub but not on the package page.
208
-
209
- - **The v0.5.1 release notes were out of order.** The section sat before v0.5.0 instead of after it, so the newest release was buried mid-list. These sections run oldest-to-newest.
210
- - A test asserted a `failed` count against a dry-run baseline that is hardcoded to `0`, making the comparison a no-op. Asserted directly now, with the reasoning recorded rather than the misleading framing.
211
-
212
- If you are on v0.5.1 you already have every fix in this release. If you are on **v0.5.0 or earlier, upgrade** — v0.5.1 fixed a data-loss bug where a single `memory remove` deleted a whole theme's mirrored memory.
213
-
214
- ## Multi-agent / per-minion themes
215
-
216
- Each systemd-run minion sets one header on its OpenAI client; everything else flows automatically:
217
-
218
- ```python
219
- client = AsyncOpenAI(
220
- base_url="http://127.0.0.1:8642/v1",
221
- api_key=API_KEY,
222
- default_headers={"X-Hermes-Session-Key": "marketing"}, # ← theme
223
- )
224
- ```
225
-
226
- The gateway plumbs `X-Hermes-Session-Key` through as `gateway_session_key=…` in `MemoryProvider.initialize` kwargs. The plugin reads it with **priority over the profile default**, so `agent_identity='default'` from unprofiled API traffic does not collapse every minion into one shared scope.
227
-
228
- Convention: lowercase, dash-separated, stable. Active themes in this deployment:
229
-
230
- - product/report themes: `marketing`, `sales`, `morning-report`, `morning-report-sr`, `sr-marketing`, `sr-cloud`
231
- - per-worker minions: `agent-trading`, `agent-sre`, `agent-marketing`, `agent-gitlab`, `agent-cloud`, `agent-hermes`
232
- - governed sinks (v0.4): `whatsapp-dm` (collapsed DM/session keys), `_bench` (benchmark traffic), `default` (last resort)
233
-
234
- Set `plugins.pgvector.allowed_themes` to that product/worker list to enforce it — an unknown or typo'd header then falls back to `default` (with a one-time warning) instead of silently minting a new theme.
235
-
236
- ## Install
237
-
238
- ### Option 1: pip + discovery shim (recommended, v0.4.2+)
239
-
240
- ```bash
241
- # 1. Install the package into the SAME environment hermes-agent runs in
242
- pip install hermes-memory-pgvector
243
-
244
- # 2. Create the discovery shim
245
- hermes-pgvector install # writes $HERMES_HOME/plugins/pgvector/
246
-
247
- # 3. Apply ALL migrations (schema + attribution + FTS + runtime grants)
248
- hermes-pgvector migrate --admin-dsn \
249
- "dbname=<your-memory-db> user=postgres host=/var/run/postgresql"
250
-
251
- # 4. Activate + verify
252
- hermes config set memory.provider pgvector
253
- sudo systemctl restart hermes.service
254
- hermes memory status # expect: Provider: pgvector; Status: available
255
- ```
256
-
257
- **Why the shim?** hermes-agent resolves a provider from bundled dirs, then `$HERMES_HOME/plugins/<name>/`, then the `hermes_agent.memory_providers` pip entry-point group. Since v0.5.0 this package **declares that entry point**, so on a host new enough to support it `pip install hermes-memory-pgvector` is sufficient on its own and no shim is needed. The shim remains for older hosts that only scan directories. Note a shim *directory* takes precedence over the entry point when one is present, so a stale shim still wins — which is why the v0.5.0 upgrade tells you to remove it. `hermes-pgvector install` writes a two-line shim whose absolute import resolves to the pip-installed package, so upgrades are just `pip install -U hermes-memory-pgvector` + restart, and rollback is `pip install hermes-memory-pgvector==<prev>` + restart — the shim never changes. `--remove` deletes it; if the package is uninstalled the shim import fails cleanly and hermes falls back to built-in memory.
258
-
259
- ### Option 2: clone + run the installer script (from source)
260
-
261
- ```bash
262
- git clone https://github.com/andreab67/hermes-memory-pgvector.git
263
- cd hermes-memory-pgvector
264
- ./scripts/install.sh
265
- ```
266
-
267
- That:
268
-
269
- 1. `pip install`s `psycopg[binary]`, `psycopg-pool`, `PyYAML` (with the upper-bound pins).
270
- 2. Copies `hermes_pgvector/` into `$HERMES_HOME/plugins/pgvector/` (defaults to `~/.hermes/plugins/pgvector/`).
271
- 3. Prints the admin migration + activation commands you run next.
272
-
273
- ### Option 3: manual
274
-
275
- ```bash
276
- # Python deps
277
- pip install 'psycopg[binary]>=3.3.5,<4' 'psycopg-pool>=3.3.1,<4' 'PyYAML>=6.0,<7'
278
-
279
- # Plugin module
280
- mkdir -p ~/.hermes/plugins
281
- cp -r hermes_pgvector ~/.hermes/plugins/pgvector
282
- ```
283
-
284
- ### Then (admin once)
285
-
286
- ```bash
287
- # Apply the schema migration (CREATE EXTENSION needs superuser)
288
- sudo -u postgres psql -d <your-memory-db> \
289
- -f ~/.hermes/plugins/pgvector/migrations/001_schema.sql
290
-
291
- # v0.4.0: apply the agent-attribution migration too (adds memory_agents /
292
- # memory_agent_edges / conversations.parent_session_id and the GRANTs the
293
- # runtime role needs — it self-grants, so no extra OWNER step for these).
294
- sudo -u postgres psql -d <your-memory-db> \
295
- -f ~/.hermes/plugins/pgvector/migrations/002_agent_attribution.sql
296
-
297
- # v0.4.1: apply the hybrid-search full-text indexes (GIN over content on both
298
- # tables). Optional — hybrid recall works without it, just seq-scans the FTS
299
- # leg. No new tables/columns/GRANTs; needs no OWNER step.
300
- sudo -u postgres psql -d <your-memory-db> \
301
- -f ~/.hermes/plugins/pgvector/migrations/003_hybrid_search_fts.sql
302
-
303
- # v0.4.2: grant the runtime role DML on the core tables (replaces the old
304
- # manual "ALTER TABLE ... OWNER TO hermes" step; skips with a NOTICE if your
305
- # runtime role isn't named 'hermes' — grant manually in that case).
306
- sudo -u postgres psql -d <your-memory-db> \
307
- -f ~/.hermes/plugins/pgvector/migrations/004_runtime_grants.sql
308
- # (or apply every migration in order: hermes-pgvector migrate --admin-dsn "user=postgres host=/var/run/postgresql dbname=<your-memory-db>")
309
-
310
- # Activate
311
- hermes config set memory.provider pgvector
312
- sudo systemctl restart hermes.service # or however you run hermes
313
- hermes memory status # expect: Provider: pgvector; Status: available
314
- ```
315
-
316
- ## Configuration
317
-
318
- Lives in `$HERMES_HOME/config.yaml` under `plugins.pgvector` — every value optional, sensible defaults shown:
319
-
320
- ```yaml
321
- plugins:
322
- pgvector:
323
- dsn: "dbname=hermes_memory user=hermes host=/var/run/postgresql"
324
- embed_url: "http://your-embed-endpoint:11434"
325
- embed_model: "nomic-embed-text"
326
- prefetch_limit: 5
327
- min_similarity: 0.30
328
- embed_on_write: true
329
- scope_default: "current"
330
- write_queue_maxsize: 256
331
- bulk_sync_on_init: true
332
- sync_turns: true
333
- turn_min_chars: 40
334
- # --- v0.4 identity governance + maintenance ---
335
- allowed_themes: [] # empty = governance off; a list enforces an allow-list
336
- bench_mode: "bucket" # bucket -> _bench | reject -> default
337
- conversation_embed_policy: "all" # all | substantive_only | none
338
- ttl_days: 0 # 0 = off; only `pgvector prune` ever deletes (never automatic)
339
- embed_write_retries: 2 # writer-path only; hot path stays single-attempt
340
- ```
341
-
342
- The embed endpoint can be any OpenAI-compatible `/v1/embeddings` or Ollama-native `/api/embed` URL that returns **768-dim vectors** (the schema is hard-coded to `vector(768)` to match `nomic-embed-text`). Use a different model only if it produces 768-dim output, or edit the migration before applying it.
343
-
344
- ## Schema
345
-
346
- ```sql
347
- CREATE TABLE memory_entries (
348
- id BIGSERIAL PRIMARY KEY,
349
- agent_identity TEXT NOT NULL DEFAULT 'default',
350
- target TEXT NOT NULL CHECK (target IN ('memory', 'user')),
351
- content TEXT NOT NULL,
352
- embedding vector(768),
353
- created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
354
- updated_at TIMESTAMPTZ NOT NULL DEFAULT now(),
355
- metadata JSONB NOT NULL DEFAULT '{}'::jsonb,
356
- UNIQUE (agent_identity, target, content)
357
- );
358
-
359
- CREATE TABLE conversations (
360
- id BIGSERIAL PRIMARY KEY,
361
- session_id TEXT NOT NULL,
362
- agent_identity TEXT NOT NULL DEFAULT 'default',
363
- role TEXT NOT NULL CHECK (role IN ('user','assistant','system','tool')),
364
- content TEXT NOT NULL,
365
- ts TIMESTAMPTZ NOT NULL DEFAULT now(),
366
- embedding vector(768),
367
- metadata JSONB NOT NULL DEFAULT '{}'::jsonb
368
- );
369
- ```
370
-
371
- Indexes: HNSW on each `embedding` column (m=16, ef_construction=64) plus per-agent + per-session btree timelines. Full DDL in [`hermes_pgvector/migrations/001_schema.sql`](hermes_pgvector/migrations/001_schema.sql).
372
-
373
- ## Tests
374
-
375
- ```bash
376
- pip install -e ".[test]"
377
-
378
- # Skip mode (no DB, no embed endpoint): everything skips gracefully
379
- pytest tests/
380
-
381
- # Live mode (against a throwaway Postgres + your embed endpoint)
382
- export PG_TEST_DSN='dbname=hermes_test user=postgres host=/var/run/postgresql'
383
- export PG_TEST_EMBED_URL='http://your-embed-endpoint:11434'
384
- pytest tests/
385
- ```
386
-
387
- DB tests skip when `PG_TEST_DSN` is unset; live embed tests skip when `PG_TEST_EMBED_URL` is unset.
388
-
389
- ## Roadmap
390
-
391
- See [`ROADMAP.md`](ROADMAP.md) for the full milestone table. Highlights:
392
-
393
- - **M1 (v0.1, v0.1.1)** ✅ Shared storage with per-agent themes, async writer, connection pool, bulk import from `MEMORY.md`/`USER.md`
394
- - **M2 (v0.2)** ✅ Conversation transcript table with `sync_turn` capture + `recall_conversation` tool
395
- - **M3 (v0.3)** ✅ Identity propagation for stateless API minions via `X-Hermes-Session-Key`
396
- - **M4 (v0.4)** ✅ Identity governance + `on_delegation()`/`on_session_end()` capture + agent attribution (`memory_agents`/`memory_agent_edges`), embedding backfill, conversation TTL, maintenance CLI
397
- - **M5 (v0.5–v0.6)** ⏳ Decay scoring, partial HNSW indexes per-theme, Prometheus metrics, cross-provider bulk-import
398
- - **M6 (v1.0)** ⏳ Stable config schema, full docs, CI coverage
399
-
400
- The roadmap exists so the multi-agent positioning isn't a one-off claim — each milestone has to pass the test *"does this make N cooperating agents more capable?"* before it lands. The `What's not on the roadmap` section in `ROADMAP.md` lists what was deliberately rejected (LLM-mediated dialectic, fact-store ontologies, background derivers, in-plugin RBAC) so the boundaries are explicit.
401
-
402
- ## Rollback
403
-
404
- ```bash
405
- hermes config set memory.provider none
406
- sudo systemctl restart hermes.service
407
-
408
- # Optional — drop the tables (data loss, irreversible)
409
- sudo -u postgres psql -d <your-memory-db> -c "
410
- DROP TABLE IF EXISTS conversations;
411
- DROP TABLE IF EXISTS memory_entries;
412
- "
413
-
414
- # Optional — remove the plugin files
415
- rm -rf ~/.hermes/plugins/pgvector
416
- ```
417
-
418
- ## Why a standalone plugin (not an upstream PR)?
419
-
420
- Per the hermes-agent [`CONTRIBUTING.md`](https://github.com/NousResearch/hermes-agent/blob/main/CONTRIBUTING.md):
421
-
422
- > We are no longer accepting new memory providers into this repo. The set of built-in providers under `plugins/memory/` is closed. If you want to add a new memory backend, publish it as a standalone plugin repo that users install into `~/.hermes/plugins/` (or via a pip entry point).
423
-
424
- The discovery system (`plugins/memory/__init__.py` in hermes-agent) scans `$HERMES_HOME/plugins/<name>/` for any directory whose `__init__.py` calls `register_memory_provider`. This plugin's `hermes_pgvector/__init__.py` does exactly that — no upstream change required.
425
-
426
- ## Contributing
427
-
428
- Bug reports + PRs welcome. Open an issue describing the failure mode + your environment (hermes-agent version, Postgres version, embed endpoint), or a PR with a focused change + test.
429
-
430
- ## License
431
-
432
- [BSD 3-Clause](LICENSE) © 2026 Green Yoga Inc
1
+ Metadata-Version: 2.4
2
+ Name: hermes-memory-pgvector
3
+ Version: 0.5.3
4
+ Summary: Postgres + pgvector memory provider plugin for hermes-agent. Multi-agent storage layer with per-minion themes, identity governance, agent attribution + delegation provenance, async writer, no LLM in the memory hot path.
5
+ Author: Andrea Borghi
6
+ License: BSD-3-Clause
7
+ Project-URL: Homepage, https://github.com/andreab67/hermes-memory-pgvector
8
+ Project-URL: Source, https://github.com/andreab67/hermes-memory-pgvector
9
+ Project-URL: Issues, https://github.com/andreab67/hermes-memory-pgvector/issues
10
+ Project-URL: Roadmap, https://github.com/andreab67/hermes-memory-pgvector/blob/main/ROADMAP.md
11
+ Keywords: hermes-agent,memory-provider,pgvector,postgres,multi-agent,llm-memory,semantic-search
12
+ Classifier: Development Status :: 4 - Beta
13
+ Classifier: Intended Audience :: Developers
14
+ Classifier: Intended Audience :: System Administrators
15
+ Classifier: License :: OSI Approved :: BSD License
16
+ Classifier: Operating System :: POSIX :: Linux
17
+ Classifier: Programming Language :: Python :: 3
18
+ Classifier: Programming Language :: Python :: 3.11
19
+ Classifier: Programming Language :: Python :: 3.12
20
+ Classifier: Programming Language :: Python :: 3.13
21
+ Classifier: Topic :: Database
22
+ Classifier: Topic :: Software Development :: Libraries :: Python Modules
23
+ Classifier: Topic :: System :: Distributed Computing
24
+ Requires-Python: >=3.11
25
+ Description-Content-Type: text/markdown
26
+ License-File: LICENSE
27
+ Requires-Dist: psycopg[binary]<4,>=3.3.5
28
+ Requires-Dist: psycopg-pool<4,>=3.3.1
29
+ Requires-Dist: PyYAML<7,>=6.0
30
+ Provides-Extra: test
31
+ Requires-Dist: pytest<9,>=7.4; extra == "test"
32
+ Dynamic: license-file
33
+
34
+ # hermes-memory-pgvector
35
+
36
+ **Postgres + pgvector memory provider for [hermes-agent](https://github.com/NousResearch/hermes-agent).** A shared memory substrate for a fleet of cooperating hermes-agent minions — built on Postgres and a single embedding endpoint you probably already run, with no LLM in the memory hot path.
37
+
38
+ ```text
39
+ each minion → X-Hermes-Session-Key: <theme>
40
+ → hermes-agent gateway
41
+ → pgvector plugin
42
+ ├── memory_entries (mirrors built-in MEMORY.md / USER.md per theme)
43
+ └── conversations (every substantive turn, semantically searchable)
44
+ ```
45
+
46
+ ## Why it exists
47
+
48
+ Existing memory providers each solve a piece of the problem; the gap for **fleet deployments** is wide:
49
+
50
+ - **Built-in `memory` tool** persists to per-host `MEMORY.md` / `USER.md`. Two minions on the same host stomp on each other; minions on different hosts have no shared substrate.
51
+ - **Honcho** offers cross-session user modelling but requires a full external service, an LLM in the memory hot path for its deriver + dialectic loops, and its own ontology layered on top of the built-in tool. In high-concurrency fleet use it produces retry storms, embedding-endpoint queue backups, and gateway↔Honcho circular dependencies.
52
+ - **Holographic** is a fine in-process fact store but uses SQLite — a poor fit for many minions writing concurrently from many hosts.
53
+ - **Other providers** (Mem0, Hindsight, OpenViking, ByteRover, RetainDB, Supermemory) all either require a paid cloud, require LLM mediation for memory ops, or both.
54
+
55
+ What was missing: a **storage layer** that gives the built-in `memory` model durable, multi-tenant, semantically-searchable backing, with no LLM in the hot path, scoped cleanly per-minion so a marketing agent's notes don't pollute a trading agent's recall. That's what this plugin provides.
56
+
57
+ ## Design philosophy
58
+
59
+ 1. **Storage layer, not a memory model.** The agent keeps using `memory(action='add', target='memory'|'user', …)`. We mirror those writes via `on_memory_write`. No new ontology for the agent to learn.
60
+ 2. **No LLM in the memory hot path.** Embeddings are vector math, not LLM calls. There is no deriver, no dialectic, no dream cycle — the failure modes that hurt Honcho cannot occur here by construction.
61
+ 3. **Per-agent themes by default, cross-theme recall on explicit demand.** Every row carries `agent_identity` (resolved from `X-Hermes-Session-Key` header, profile name, workspace, or `'default'`). Recall is scoped to the current theme unless the agent asks for `scope='all'`.
62
+ 4. **Fail-soft everywhere.** Embed endpoint down → degrade to text-only writes. Async writer queue full → drop with a one-time warning. DB down → log + skip. No exception escapes into the agent loop.
63
+ 5. **Admin/runtime separation.** DDL (`CREATE EXTENSION vector`, `CREATE TABLE`, `CREATE INDEX`) runs once with superuser. The runtime user has DML only on the migrated schema. `ensure_schema()` at runtime is verify-only with a clear `SchemaNotApplied` error if the operator forgot the migration.
64
+
65
+ ## Features (v0.3.0)
66
+
67
+ | Hook / surface | Behavior |
68
+ |---|---|
69
+ | `initialize()` | Verifies schema, opens `psycopg_pool.ConnectionPool`, bulk-imports existing `MEMORY.md` + `USER.md` content. |
70
+ | `on_memory_write(action, target, content, meta)` | Mirrors built-in `memory` writes into `memory_entries` (add / replace / remove). |
71
+ | `sync_turn(user, assistant, session_id)` | Captures every substantive (`>= 40` chars + not boilerplate) chat turn into `conversations`. |
72
+ | `prefetch(query)` | Top-K semantically similar `memory_entries` in current theme, injected ambient. |
73
+ | `recall_memory(query, scope, target, limit)` tool | Explicit cross-theme search of durable memory entries. |
74
+ | `recall_conversation(query, scope, limit)` tool | Explicit search over past chat turns. `scope ∈ {current, session, all, <theme>}`. |
75
+
76
+ Internals:
77
+
78
+ - **`psycopg_pool.ConnectionPool`** (min=0, max=4, lazy + thread-safe, `max_idle=30s` / `max_lifetime=300s`) shared across the agent thread and the async-writer drain thread. `min_size=0` keeps an idle — or abandoned — pool at **zero** open connections, so a session the gateway never explicitly shuts down cannot strand a Postgres backend (see *Fixed in v0.3.1* below).
79
+ - **`AsyncWriter`** — bounded queue + daemon drain thread. Memory write hooks return in microseconds. Worker embeds + writes in the background. Crash-resilient (auto-restart on next enqueue).
80
+ - **Single migration** (`hermes_pgvector/migrations/001_schema.sql`) — `memory_entries` + `conversations` + HNSW indexes. Same tuning operators typically use elsewhere.
81
+ - **Boilerplate filter** for turn capture — length floor + acknowledgement regex (`"ok"`, `"thanks"`, `"continue"`, …) so the recall table stays high-signal.
82
+
83
+ ### Fixed in v0.3.1 — connection-leak hotfix
84
+
85
+ A single registered provider has `initialize()` called again for each new session. It previously
86
+ reassigned `self._store` / `self._writer` without closing the prior ones, **abandoning a
87
+ `ConnectionPool`** whose warm (`min_size=1`) connection lingered in Postgres — committed-but-idle —
88
+ until the server's `idle_session_timeout`. Under a burst of concurrent sessions (e.g. a swarm of
89
+ systemd-run minions firing on the same minute) these orphaned backends saturated the database's
90
+ connection slots. Fixed by:
91
+
92
+ 1. **`initialize()` teardown** — drain the prior `AsyncWriter` + close the prior pool before
93
+ re-initializing (the call is idempotent and skipped on first init).
94
+ 2. **Self-draining pool** — `min_size=0` (an idle or abandoned pool holds *zero* connections) plus
95
+ `max_idle=30s` / `max_lifetime=300s`, so connections are short-lived when idle and pooled only
96
+ under active load.
97
+
98
+ ## New in v0.4.0
99
+
100
+ Four capabilities, all storage-layer (still no LLM in the hot path):
101
+
102
+ - **Identity governance.** The resolved `agent_identity` is normalized once at init: direct-message session keys like `agent:main:whatsapp:dm:<phone>` collapse to a single `whatsapp-dm` bucket (no PII, no per-contact theme explosion), benchmark traffic (`skill-bench*`) is isolated to `_bench`, and an optional `allowed_themes` allow-list routes typo'd/unknown themes to `default`. The M3 resolution priority is preserved — normalization runs *after* the chain, never at read time. See [`hermes_pgvector/identity.py`](hermes_pgvector/identity.py).
103
+ - **Agent attribution + delegation (M4).** Migration `002` adds `memory_agents` (registry) and `memory_agent_edges` (parent→child delegation provenance) plus `conversations.parent_session_id`. The `on_delegation` / `on_session_end` hooks capture which agent delegated what to whom — strictly enqueue-only and fail-soft. **Provenance only** (who/when), never a fact-store ontology. Query it via the `v_agent_memory` view.
104
+ - **Embedding backfill + writer resilience.** Rows written text-only during an embed-endpoint outage are no longer permanently unsearchable: `hermes-pgvector backfill` re-embeds `NULL`-embedding rows (idempotent, 768-dim-guarded). The background writer gains a small bounded retry; the hot path stays single-attempt.
105
+ - **Conversation TTL + embed policy.** `hermes-pgvector prune --days N` trims old turns (operator-triggered only; `memory_entries` are never pruned). `conversation_embed_policy` (`all` default / `substantive_only` / `none`) tunes embedding cost.
106
+
107
+ Maintenance CLI (`hermes-pgvector`, or `python -m hermes_pgvector`): `migrate · stats · backfill · prune · cleanup · remap` — destructive commands default to dry-run. v0.4.0 is a clean upgrade from v0.3.x: apply migration `002` to light up attribution/delegation; without it the new hooks no-op and everything else runs unchanged.
108
+
109
+ ## New in v0.4.1 — hybrid recall (vector + full-text)
110
+
111
+ `recall_memory` and `recall_conversation` now fuse the HNSW **vector** ranking with a Postgres **full-text** ranking using **Reciprocal Rank Fusion** (RRF, `k=60`). A row surfaces if *either* ranker likes it, which fixes the two blind spots of pure cosine similarity:
112
+
113
+ - **Exact-lexical hits** the embedding smooths away — a specific error code, hostname, flag name, or rare identifier the agent quotes verbatim.
114
+ - **Text-only rows with a `NULL` embedding** (written while the embed endpoint was down) — invisible to the vector index, but the full-text leg finds them. So hybrid recall doubles as best-effort recovery until the next `backfill`.
115
+
116
+ Still a storage-layer feature: **no LLM, no entity graph, no new tables or columns** — just a GIN index over the existing `content` column (migration `003`) and a fused query. It stays inside invariant #1 (a second index over the same text is not a parallel ontology). Fail-soft as ever: a hybrid hiccup degrades to the proven pure-vector path, and a query that *itself* fails to embed degrades to full-text-only instead of erroring. Toggle with `plugins.pgvector.hybrid_search` (default `true`); the ambient `prefetch()` path stays pure-vector. Works without migration `003` — the GIN index only makes the full-text leg faster.
117
+
118
+ ## New in v0.4.2 — pip-native install + hardening
119
+
120
+ - **`hermes-pgvector install`** — makes a plain `pip install hermes-memory-pgvector` deployable on ANY hermes-agent install: generates the `$HERMES_HOME/plugins/pgvector/` discovery shim (see *Install · Option 1*). No more vendored copies or editable checkouts.
121
+ - **Migration `004`** — `hermes-pgvector migrate` now grants the runtime role DML on `memory_entries`/`conversations` itself; the manual OWNER-transfer step is gone (fresh installs previously hit `permission denied` if it was skipped).
122
+ - **Correctness fixes** from a full-codebase review: `replace`/`remove` now match `old_text` as a *literal* substring (LIKE `%`/`_`/`\` metacharacters no longer over- or under-match — parity with the built-in tool's `in` semantics); the async writer drains its queue on shutdown instead of silently abandoning up to 255 accepted writes when full; a wrong-dimension embed model now surfaces as `expected 768 dims, got N` instead of a masking 404; DM-key bucketing no longer sweeps ordinary `:signal:`-containing theme names into `whatsapp-dm`; bulk MEMORY.md import circuit-breaks after 3 consecutive embed failures (a hanging endpoint can no longer block session start for minutes); `remap` re-checks its duplicate-drop guard under the advisory lock; tool errors redact credential-looking fragments and preserve `score: null` for full-text-only hybrid hits (with `rrf_score` now included); `recall_memory(scope='session')` returns a helpful error instead of silently matching nothing.
123
+
124
+ ## New in v0.4.3 — psycopg 3.3.5 floor
125
+
126
+ - **Dependency floor raised**: `psycopg[binary]>=3.3.5` (upstream bugfix release, 2026-08-31: prepared-statement invalidation on `ALTER`/`DISCARD`, DataError fixes for malformed COPY/jsonb data, client-encoding aliases). No code changes.
127
+
128
+ ## New in v0.5.0 — import rename (BREAKING), read-side identity gate, config-contract fixes
129
+
130
+ > ### Breaking upgrade — read before installing
131
+ >
132
+ > **1. The import package is renamed** `pgvector` -> `hermes_pgvector`. Unchanged:
133
+ > the distribution (`hermes-memory-pgvector`), the CLI (`hermes-pgvector`), and
134
+ > the hermes provider name (`pgvector`, i.e. `memory.provider: pgvector`).
135
+ >
136
+ > Why: the old top-level name is owned by
137
+ > [pgvector-python](https://pypi.org/project/pgvector/). Both in one venv meant
138
+ > whichever installed last won, and this plugin's shim could import the wrong
139
+ > module — taking the fleet's shared memory offline on a single log line.
140
+ >
141
+ > **2. Any `python -m pgvector ...` command breaks.** It is now
142
+ > `hermes-pgvector ...` (or `python -m hermes_pgvector ...`). This matters most
143
+ > for scheduled jobs, which fail *silently* — the nightly backfill simply stops,
144
+ > and rows written during an embed outage stay permanently unsearchable. Check
145
+ > your units before upgrading:
146
+ >
147
+ > ```bash
148
+ > sudo grep -rl 'python -m pgvector' /etc/systemd/system/ /etc/cron.d/ 2>/dev/null
149
+ > # e.g. hermes-pgvector-backfill.service:
150
+ > # ExecStart=.../python -m pgvector backfill -> -m hermes_pgvector backfill
151
+ > sudo systemctl daemon-reload
152
+ > ```
153
+ >
154
+ > **3. Upgrade steps.** An existing shim still reads `from pgvector import ...`;
155
+ > after upgrading it fails, and the loader treats that as "plugin absent" and
156
+ > falls back to built-in memory.
157
+ >
158
+ > ```bash
159
+ > pip install -U hermes-memory-pgvector==0.5.0
160
+ >
161
+ > # Preferred, if your hermes-agent reads the `hermes_agent.memory_providers`
162
+ > # entry-point group (this package now declares it): drop the shim entirely and
163
+ > # let pip discovery take over -- nothing left to go stale on future upgrades.
164
+ > hermes-pgvector install --remove
165
+ >
166
+ > # Otherwise (older host that only scans plugin directories): regenerate it.
167
+ > hermes-pgvector install --force
168
+ >
169
+ > # restart hermes, then verify -- do not skip this:
170
+ > hermes memory status # expect: Provider: pgvector; Status: available
171
+ > ```
172
+ >
173
+ > `hermes-pgvector install` now verifies in a clean subprocess that the shim it
174
+ > just wrote can actually be imported, and warns if it cannot or if another
175
+ > package shadows this one.
176
+
177
+ - **Read-side identity gate.** The `whatsapp-dm` and `_bench` sinks were write-side only: `identity.py` stripped PII from the *identity*, but message bodies still live in `content`, and nothing filtered them on read. Any theme could pull DM content into its context via `scope='all'` or by naming the bucket directly — and with turn capture on, the reply quoting it was written back under the *reading* theme, permanently re-attributing DM data. `scope='all'` now excludes those sinks, and naming one explicitly is rejected. An agent that *is* the bucket keeps full access to its own rows, and ordinary cross-theme recall is unaffected. **Scope of the gate:** it excludes by bucket *name*, and bucketing happens at write time — rows are never retroactively rewritten (that is a deliberate invariant: historical rows keep their identity or they become unrecallable). So any row written *before* its key was bucketed still carries the raw identity and is still reachable. Verified on the reference deployment: **zero** such rows exist there (`whatsapp-dm` is already bucketed and no group traffic predates this release). If yours has them, remap them — and note the destructive commands default to **dry-run**, so the first form only *reports*: `hermes-pgvector remap --old <raw> --new whatsapp-dm` to preview, then re-run with `--execute` to actually move the rows (add `--force` if more than 10 duplicates would be dropped). Without `--execute` nothing moves, and it is easy to believe the rows were bucketed when they were not.
178
+
179
+
180
+ Correctness release from a full-codebase review. **No schema changes and no new migrations** — but this release is not drop-in: the import package is renamed (see the upgrade box above), and `MemoryStore.search` / `hybrid_search` / `search_turns` / `hybrid_search_turns` gain an `exclude_identities` parameter. The hermes provider name, the CLI, and the on-disk schema are unchanged.
181
+
182
+ - **Embed timeouts are configurable, and split by call path.** `timeout` was never plumbed from config at all — every caller silently took a hardcoded 10s. On an endpoint that answers in 6–17s that means a large share of background writes time out, fail soft, and land as rows with a NULL embedding, invisible to recall until `hermes-pgvector backfill` repairs them. Retries could not help: every attempt was capped *below* the latency the endpoint needs. There are now two keys, deliberately asymmetric — `embed_timeout` (default `10.0`) for the agent thread, where a timeout degrades recall to full-text-only and waiting longer would be worse; and `embed_write_timeout` (default `30.0`) for the background writer, where nothing is waiting and giving up costs a permanently unsearchable row. Measured against the reference endpoint: 1/3 writes succeeded at 10s, 3/3 at 30s.
183
+ - **Group / channel / thread keys are bucketed.** Multi-party session keys like `agent:main:whatsapp:group:<chat>:<participant>` previously passed through untouched, so with the host's default `group_sessions_per_user` the trailing participant id — a phone number on WhatsApp/SMS/Signal — was stored verbatim as an `agent_identity`. That is the same PII failure the DM bucket exists to prevent, reached through a different `chat_type`. They now collapse to a single `external-group` bucket, which (like `whatsapp-dm` and `_bench`) is excluded from other themes' recall.
184
+ - **`allowed_themes` accepts a string again.** The config schema declared it a scalar string while `normalize_identity()` consumed it as a list of names — so a string allow-list was iterated *character by character*, every theme failed the membership test, and the whole fleet was silently routed to `default`. Governance looked configured while doing the opposite. Comma-separated strings and YAML lists both work now.
185
+ - **Boolean toggles honour `false` again.** `embed_on_write`, `sync_turns`, `hybrid_search` and `bulk_sync_on_init` are declared by the config schema as the *strings* `"true"`/`"false"`, but were read with plain truthiness — and `bool("false")` is `True`, so turning any of them off via that path did nothing.
186
+ - **Conversation turns are no longer written twice.** `sync_turn()` (per exchange) and `on_session_end()` (whole transcript, again on session rotation) both captured the same turns, and `conversations` has no unique constraint. `on_session_end` is now a true backstop: it skips turns already accepted by the writer, and still re-captures ones a full queue dropped.
187
+ - **`memory` replace mirrors correctly.** `replace()` issued one bulk `UPDATE` across every substring match, colliding with `UNIQUE(agent_identity, target, content)` — so a `replace` matching two or more entries raised `UniqueViolation` and updated **zero** rows. It now updates the first match, matching the built-in tool.
188
+ - **A dead database is no longer silent.** Worker write failures logged at `debug` only, so a Postgres restart after a healthy init discarded every durable write for the rest of the session with no operator signal. The first failure now warns. Relatedly, `system_prompt_block` no longer tells the model "Empty store" when the count query merely *failed*.
189
+ - **`save_config` stops deleting your settings.** It replaced the whole `plugins.pgvector` block with schema-declared keys, silently dropping hand-edited ones that are read at runtime (`identity_aliases`, `embed_write_backoff`). It merges now.
190
+ - **Fail-soft hardening** (invariant #4): `sync_turn` is wrapped, config casts are guarded, and the recall tools coerce non-string `query`/`scope`/`target` instead of raising `AttributeError` out of the hook. `hermes-pgvector install --remove` now fails closed and requires `--force` on a directory that isn't a generated shim, instead of deleting it outright.
191
+
192
+ ## New in v0.5.1 — `memory remove` no longer wipes a whole theme
193
+
194
+ Patch release, but **upgrade promptly**: it fixes a data-loss bug. No schema changes, no migrations, no API changes.
195
+
196
+ - **`remove` deleted every mirrored entry for a theme, not one.** `_worker` passed `old_text=item.content` — but the built-in tool's remove op carries its target in `old_text` and leaves `content` empty, and the host forwards `old_text` via *metadata*. So `item.content` was always `""`, `store.remove` built `content LIKE '%%'`, and that matches every row: a single `memory remove` deleted the entire mirror for that `(agent_identity, target)`. Verified against Postgres — `DELETE … WHERE c LIKE '%%'` removes all rows. `_worker` now reads `extra["old_text"]`, and `store.remove()` **refuses an empty pattern outright**, so no caller can reach that delete by omission. `remove()` also now deletes **at most one row** (lowest id), matching both the built-in tool — which requires a unique match and errors on ambiguity — and this class's own `replace()`. (The built-in store was never affected; only the pgvector mirror. No loss occurred on the reference deployment: every theme's history is continuous.)
197
+
198
+ Found in production: one `memory_entries` row sat with a NULL embedding and **zero-length content**, arrived through the built-in tool's `replace` path. Two defects met there.
199
+
200
+ - **Nothing rejected empty content on write.** `on_memory_write` filtered on `target` and `action` but never on content, so an `add`/`replace` carrying nothing created a row that can never be embedded — `embed()` raises `EmbeddingError("empty input")` unconditionally for empty or whitespace text. Such writes are now ignored (`remove` is exempt: it legitimately arrives with empty content and targets the row via `old_text`).
201
+ - **`backfill_null_embeddings` retried it forever.** The sweep selected `WHERE embedding IS NULL` with no content filter, so every nightly run re-fetched the row, called `embed()`, failed, and moved on — permanently pinning `failed` above zero and making `remaining == 0` unreachable. That is the damaging half: it destroys the one signal an operator watches, because you can no longer distinguish a permanently-stuck row from a new genuine failure. Un-embeddable rows are now **skipped and reported separately** as `unembeddable` (skipping them silently would be equally misleading), so `remaining` can actually reach zero again.
202
+
203
+ ## New in v0.5.2 — documentation only
204
+
205
+ **No behaviour changes.** `git diff v0.5.1..v0.5.2` shows four files: `README.md`, one test file, and the two version strings (`pyproject.toml` and `hermes_pgvector/plugin.yaml`). The only packaged file that differs is `plugin.yaml`, and only its `version:` line — no logic changed anywhere, so an installed 0.5.2 behaves identically to 0.5.1. There is no reason to redeploy for this release.
206
+
207
+ It exists because PyPI renders a project's README **frozen at upload time**: two fixes that landed after v0.5.1 shipped were visible on GitHub but not on the package page.
208
+
209
+ - **The v0.5.1 release notes were out of order.** The section sat before v0.5.0 instead of after it, so the newest release was buried mid-list. These sections run oldest-to-newest.
210
+ - A test asserted a `failed` count against a dry-run baseline that is hardcoded to `0`, making the comparison a no-op. Asserted directly now, with the reasoning recorded rather than the misleading framing.
211
+
212
+ If you are on v0.5.1 you already have every fix in this release. If you are on **v0.5.0 or earlier, upgrade** — v0.5.1 fixed a data-loss bug where a single `memory remove` deleted a whole theme's mirrored memory.
213
+
214
+ ## New in v0.5.3 - configurable embedding model (dimension, auth, protocol)
215
+
216
+ **Defaults are unchanged:** 768 dimensions, no `Authorization` header, `auto` protocol. No schema changes and no new migrations; a deployment that sets none of the new keys behaves as before, apart from the two fixes at the end of this list.
217
+
218
+ - **The embedding dimension is configuration, not code.** `embed_dim` (default `768`) replaces the literal 768 in the response check, in the backfill dimension guard (`backfill_null_embeddings(expected_dim=...)`), and in the `stats` dry-run. The check was moved, not relaxed: a mismatch still fails fast with `expected N dims (embed_dim), got M`. Changing the value on a database that already holds vectors needs a column migration; see [Changing the embedding dimension](#changing-the-embedding-dimension).
219
+ - **Bearer auth for hosted endpoints.** `embed_api_key_env` holds the *name* of an environment variable (for example `OPENROUTER_API_KEY`). When that variable is set and non-empty, the plugin sends `Authorization: Bearer <value>`. The value is read at call time and is never logged, stored in config, or included in exception messages. Unset or empty means no header, as before.
220
+ - **Explicit protocol selection.** `embed_protocol: openai` uses only `/v1/embeddings`, so a 401 or an unknown-model error from a hosted endpoint is reported as-is instead of being replaced by a 404 from the Ollama-native fallback. `ollama` uses only `/api/embed`. `auto` keeps the old try-OpenAI-then-Ollama behaviour. Unknown values fall back to `auto` with a warning.
221
+ - **One embed path.** Prefetch, both recall tools, the init-time bulk import, the writer drain and `hermes-pgvector backfill` all resolve the endpoint through one helper, so the settings apply the same way everywhere (`stats` reads the same `embed_dim`). Timeouts and retries are unchanged: one attempt on the agent thread, bounded retries on the writer. `backfill` gains `--embed-dim`, `--embed-api-key-env` and `--embed-protocol` (CLI flag > `--config` file > default).
222
+ - **Fixed: embeds broke under hermes-agent's plugin loader.** After it runs the package, `plugins/plugin_loader.py:load_plugin_module` binds every sibling module back onto it, including `setattr(pkg, "embed", <the embed submodule>)`. That replaced the `embed` function the provider called, so every embed raised `TypeError: 'module' object is not callable`. That is not an `EmbeddingError`, so nothing degraded gracefully: prefetch and the recall tools raised out of the hook, and the writer dropped each mirrored write and captured turn outright instead of storing it text-only. Call sites now use a private alias the loader never touches. An external patch that re-binds `embed` inside `register()` is no longer needed, and does no harm if it is still present. `from hermes_pgvector import embed` still works.
223
+ - **Fixed: a read timeout escaped as a bare `TimeoutError`.** urllib wraps errors raised while *sending* a request, but a server that accepts the connection and answers slower than the timeout raises `TimeoutError` from the response read. That slipped past every `except EmbeddingError`: on the agent thread it raised out of prefetch and the recall tools, `auto` never tried its fallback, and on the writer the retries never ran and the write was dropped instead of being stored text-only. It is now an `EmbeddingError`, like every other endpoint failure.
224
+
225
+ ## Multi-agent / per-minion themes
226
+
227
+ Each systemd-run minion sets one header on its OpenAI client; everything else flows automatically:
228
+
229
+ ```python
230
+ client = AsyncOpenAI(
231
+ base_url="http://127.0.0.1:8642/v1",
232
+ api_key=API_KEY,
233
+ default_headers={"X-Hermes-Session-Key": "marketing"}, # ← theme
234
+ )
235
+ ```
236
+
237
+ The gateway plumbs `X-Hermes-Session-Key` through as `gateway_session_key=…` in `MemoryProvider.initialize` kwargs. The plugin reads it with **priority over the profile default**, so `agent_identity='default'` from unprofiled API traffic does not collapse every minion into one shared scope.
238
+
239
+ Convention: lowercase, dash-separated, stable. Active themes in this deployment:
240
+
241
+ - product/report themes: `marketing`, `sales`, `morning-report`, `morning-report-sr`, `sr-marketing`, `sr-cloud`
242
+ - per-worker minions: `agent-trading`, `agent-sre`, `agent-marketing`, `agent-gitlab`, `agent-cloud`, `agent-hermes`
243
+ - governed sinks (v0.4): `whatsapp-dm` (collapsed DM/session keys), `_bench` (benchmark traffic), `default` (last resort)
244
+
245
+ Set `plugins.pgvector.allowed_themes` to that product/worker list to enforce it — an unknown or typo'd header then falls back to `default` (with a one-time warning) instead of silently minting a new theme.
246
+
247
+ ## Install
248
+
249
+ ### Option 1: pip + discovery shim (recommended, v0.4.2+)
250
+
251
+ ```bash
252
+ # 1. Install the package into the SAME environment hermes-agent runs in
253
+ pip install hermes-memory-pgvector
254
+
255
+ # 2. Create the discovery shim
256
+ hermes-pgvector install # writes $HERMES_HOME/plugins/pgvector/
257
+
258
+ # 3. Apply ALL migrations (schema + attribution + FTS + runtime grants)
259
+ hermes-pgvector migrate --admin-dsn \
260
+ "dbname=<your-memory-db> user=postgres host=/var/run/postgresql"
261
+
262
+ # 4. Activate + verify
263
+ hermes config set memory.provider pgvector
264
+ sudo systemctl restart hermes.service
265
+ hermes memory status # expect: Provider: pgvector; Status: available
266
+ ```
267
+
268
+ **Why the shim?** hermes-agent resolves a provider from bundled dirs, then `$HERMES_HOME/plugins/<name>/`, then the `hermes_agent.memory_providers` pip entry-point group. Since v0.5.0 this package **declares that entry point**, so on a host new enough to support it `pip install hermes-memory-pgvector` is sufficient on its own and no shim is needed. The shim remains for older hosts that only scan directories. Note a shim *directory* takes precedence over the entry point when one is present, so a stale shim still wins — which is why the v0.5.0 upgrade tells you to remove it. `hermes-pgvector install` writes a two-line shim whose absolute import resolves to the pip-installed package, so upgrades are just `pip install -U hermes-memory-pgvector` + restart, and rollback is `pip install hermes-memory-pgvector==<prev>` + restart — the shim never changes. `--remove` deletes it; if the package is uninstalled the shim import fails cleanly and hermes falls back to built-in memory.
269
+
270
+ ### Option 2: clone + run the installer script (from source)
271
+
272
+ ```bash
273
+ git clone https://github.com/andreab67/hermes-memory-pgvector.git
274
+ cd hermes-memory-pgvector
275
+ ./scripts/install.sh
276
+ ```
277
+
278
+ That:
279
+
280
+ 1. `pip install`s `psycopg[binary]`, `psycopg-pool`, `PyYAML` (with the upper-bound pins).
281
+ 2. Copies `hermes_pgvector/` into `$HERMES_HOME/plugins/pgvector/` (defaults to `~/.hermes/plugins/pgvector/`).
282
+ 3. Prints the admin migration + activation commands you run next.
283
+
284
+ ### Option 3: manual
285
+
286
+ ```bash
287
+ # Python deps
288
+ pip install 'psycopg[binary]>=3.3.5,<4' 'psycopg-pool>=3.3.1,<4' 'PyYAML>=6.0,<7'
289
+
290
+ # Plugin module
291
+ mkdir -p ~/.hermes/plugins
292
+ cp -r hermes_pgvector ~/.hermes/plugins/pgvector
293
+ ```
294
+
295
+ ### Then (admin once)
296
+
297
+ ```bash
298
+ # Apply the schema migration (CREATE EXTENSION needs superuser)
299
+ sudo -u postgres psql -d <your-memory-db> \
300
+ -f ~/.hermes/plugins/pgvector/migrations/001_schema.sql
301
+
302
+ # v0.4.0: apply the agent-attribution migration too (adds memory_agents /
303
+ # memory_agent_edges / conversations.parent_session_id and the GRANTs the
304
+ # runtime role needs — it self-grants, so no extra OWNER step for these).
305
+ sudo -u postgres psql -d <your-memory-db> \
306
+ -f ~/.hermes/plugins/pgvector/migrations/002_agent_attribution.sql
307
+
308
+ # v0.4.1: apply the hybrid-search full-text indexes (GIN over content on both
309
+ # tables). Optional — hybrid recall works without it, just seq-scans the FTS
310
+ # leg. No new tables/columns/GRANTs; needs no OWNER step.
311
+ sudo -u postgres psql -d <your-memory-db> \
312
+ -f ~/.hermes/plugins/pgvector/migrations/003_hybrid_search_fts.sql
313
+
314
+ # v0.4.2: grant the runtime role DML on the core tables (replaces the old
315
+ # manual "ALTER TABLE ... OWNER TO hermes" step; skips with a NOTICE if your
316
+ # runtime role isn't named 'hermes' — grant manually in that case).
317
+ sudo -u postgres psql -d <your-memory-db> \
318
+ -f ~/.hermes/plugins/pgvector/migrations/004_runtime_grants.sql
319
+ # (or apply every migration in order: hermes-pgvector migrate --admin-dsn "user=postgres host=/var/run/postgresql dbname=<your-memory-db>")
320
+
321
+ # Activate
322
+ hermes config set memory.provider pgvector
323
+ sudo systemctl restart hermes.service # or however you run hermes
324
+ hermes memory status # expect: Provider: pgvector; Status: available
325
+ ```
326
+
327
+ ## Configuration
328
+
329
+ Lives in `$HERMES_HOME/config.yaml` under `plugins.pgvector` — every value optional, sensible defaults shown:
330
+
331
+ ```yaml
332
+ plugins:
333
+ pgvector:
334
+ dsn: "dbname=hermes_memory user=hermes host=/var/run/postgresql"
335
+ embed_url: "http://your-embed-endpoint:11434"
336
+ embed_model: "nomic-embed-text"
337
+ embed_dim: 768 # v0.5.3: must match the model AND the vector(N) columns
338
+ embed_api_key_env: "" # v0.5.3: NAME of an env var holding a bearer token
339
+ embed_protocol: "auto" # v0.5.3: auto | openai | ollama
340
+ prefetch_limit: 5
341
+ min_similarity: 0.30
342
+ embed_on_write: true
343
+ scope_default: "current"
344
+ write_queue_maxsize: 256
345
+ bulk_sync_on_init: true
346
+ sync_turns: true
347
+ turn_min_chars: 40
348
+ # --- v0.4 identity governance + maintenance ---
349
+ allowed_themes: [] # empty = governance off; a list enforces an allow-list
350
+ bench_mode: "bucket" # bucket -> _bench | reject -> default
351
+ conversation_embed_policy: "all" # all | substantive_only | none
352
+ ttl_days: 0 # 0 = off; only `pgvector prune` ever deletes (never automatic)
353
+ embed_write_retries: 2 # writer-path only; hot path stays single-attempt
354
+ ```
355
+
356
+ The embed endpoint can be any OpenAI-compatible `/v1/embeddings` or Ollama-native `/api/embed` URL. Its vectors must be exactly `embed_dim` long, and `embed_dim` must match the database's `vector(N)` columns. Migration 001 creates `vector(768)` to match `nomic-embed-text`, which is why 768 is the default.
357
+
358
+ ### Embedding endpoint keys (v0.5.3)
359
+
360
+ | Key | Default | Meaning |
361
+ |---|---|---|
362
+ | `embed_dim` | `768` | Vector length the model returns. Every embedding is checked against it, and so is the `hermes-pgvector backfill` probe. It must equal the `vector(N)` column size: changing it on an existing database is a migration, see [below](#changing-the-embedding-dimension). |
363
+ | `embed_api_key_env` | unset | **Name** of an environment variable holding a bearer token, e.g. `OPENROUTER_API_KEY`. Never put the token itself in config. The variable is read on every request; when it is set and non-empty the plugin sends `Authorization: Bearer <value>`, otherwise no header. |
364
+ | `embed_protocol` | `auto` | `auto`: try `/v1/embeddings`, then fall back to `/api/embed`. `openai`: `/v1/embeddings` only, so auth and model errors surface as-is (use this for hosted OpenAI-compatible APIs). `ollama`: `/api/embed` only. Unknown values fall back to `auto` with a warning. |
365
+
366
+ Example: OpenAI `text-embedding-3-small` (1536 dimensions) through OpenRouter:
367
+
368
+ ```yaml
369
+ plugins:
370
+ pgvector:
371
+ embed_url: "https://openrouter.ai/api" # the plugin appends /v1/embeddings
372
+ embed_model: "openai/text-embedding-3-small"
373
+ embed_dim: 1536
374
+ embed_api_key_env: "OPENROUTER_API_KEY" # the variable's NAME, not the key
375
+ embed_protocol: "openai"
376
+ ```
377
+
378
+ The variable has to be in the environment of every process that loads the provider (each hermes service) and of any `hermes-pgvector backfill` job. To call OpenAI directly instead, use `embed_url: "https://api.openai.com"`, `embed_model: "text-embedding-3-small"` and a variable holding an OpenAI key.
379
+
380
+ ### Changing the embedding dimension
381
+
382
+ `embed_dim` has to agree with the columns, so switching to a model with a different output size is a migration, not a config edit. Vectors from two different models are not comparable anyway, so every row must be re-embedded. Until config and columns agree, Postgres rejects each write whose vector has the wrong length (`expected 1536 dimensions, not 768`), and the whole row is lost, not stored text-only. Stop the services first.
383
+
384
+ The shipped migration files are not meant to be edited for this. As the table owner:
385
+
386
+ ```sql
387
+ -- 1. With every hermes service that loads the provider stopped:
388
+ BEGIN;
389
+ DROP INDEX IF EXISTS ix_memory_entries_embedding_hnsw;
390
+ DROP INDEX IF EXISTS ix_conversations_embedding_hnsw;
391
+ ALTER TABLE memory_entries ALTER COLUMN embedding TYPE vector(1536) USING NULL::vector(1536);
392
+ ALTER TABLE conversations ALTER COLUMN embedding TYPE vector(1536) USING NULL::vector(1536);
393
+ COMMIT;
394
+ ```
395
+
396
+ 2. Set `embed_model` and `embed_dim` (plus `embed_url`, `embed_api_key_env` and `embed_protocol` as needed), then start the services. New writes are embedded with the new model.
397
+ 3. Re-embed the existing rows, which are all NULL now: `hermes-pgvector backfill --config $HERMES_HOME/config.yaml`. Repeat until every table reports `remaining: 0`. If the endpoint does not return `embed_dim`-length vectors, the run aborts on its first probe, before touching any row, and the logged warning names both sizes.
398
+ 4. Rebuild the HNSW indexes with the shipped tuning. Building them after the backfill is faster than maintaining them during it:
399
+
400
+ ```sql
401
+ CREATE INDEX CONCURRENTLY IF NOT EXISTS ix_memory_entries_embedding_hnsw
402
+ ON memory_entries USING hnsw (embedding vector_cosine_ops) WITH (m = 16, ef_construction = 64);
403
+ CREATE INDEX CONCURRENTLY IF NOT EXISTS ix_conversations_embedding_hnsw
404
+ ON conversations USING hnsw (embedding vector_cosine_ops) WITH (m = 16, ef_construction = 64);
405
+ ```
406
+
407
+ Until step 3 completes, rows without a vector are reachable only through full-text recall (`hybrid_search: true`). pgvector's HNSW index supports `vector` columns of up to 2,000 dimensions. The plugin maintains only `memory_entries` and `conversations`; any other embedding columns in the same database need the same change from whatever writes them.
408
+
409
+ ## Schema
410
+
411
+ ```sql
412
+ CREATE TABLE memory_entries (
413
+ id BIGSERIAL PRIMARY KEY,
414
+ agent_identity TEXT NOT NULL DEFAULT 'default',
415
+ target TEXT NOT NULL CHECK (target IN ('memory', 'user')),
416
+ content TEXT NOT NULL,
417
+ embedding vector(768),
418
+ created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
419
+ updated_at TIMESTAMPTZ NOT NULL DEFAULT now(),
420
+ metadata JSONB NOT NULL DEFAULT '{}'::jsonb,
421
+ UNIQUE (agent_identity, target, content)
422
+ );
423
+
424
+ CREATE TABLE conversations (
425
+ id BIGSERIAL PRIMARY KEY,
426
+ session_id TEXT NOT NULL,
427
+ agent_identity TEXT NOT NULL DEFAULT 'default',
428
+ role TEXT NOT NULL CHECK (role IN ('user','assistant','system','tool')),
429
+ content TEXT NOT NULL,
430
+ ts TIMESTAMPTZ NOT NULL DEFAULT now(),
431
+ embedding vector(768),
432
+ metadata JSONB NOT NULL DEFAULT '{}'::jsonb
433
+ );
434
+ ```
435
+
436
+ Indexes: HNSW on each `embedding` column (m=16, ef_construction=64) plus per-agent + per-session btree timelines. Full DDL in [`hermes_pgvector/migrations/001_schema.sql`](hermes_pgvector/migrations/001_schema.sql).
437
+
438
+ `vector(768)` is the size migration 001 creates. A deployment on a model with a different output size changes both columns and sets `embed_dim` to match; see [Changing the embedding dimension](#changing-the-embedding-dimension).
439
+
440
+ ## Tests
441
+
442
+ ```bash
443
+ pip install -e ".[test]"
444
+
445
+ # Skip mode (no DB, no embed endpoint): everything skips gracefully
446
+ pytest tests/
447
+
448
+ # Live mode (against a throwaway Postgres + your embed endpoint)
449
+ export PG_TEST_DSN='dbname=hermes_test user=postgres host=/var/run/postgresql'
450
+ export PG_TEST_EMBED_URL='http://your-embed-endpoint:11434'
451
+ pytest tests/
452
+ ```
453
+
454
+ DB tests skip when `PG_TEST_DSN` is unset; live embed tests skip when `PG_TEST_EMBED_URL` is unset.
455
+
456
+ ## Roadmap
457
+
458
+ See [`ROADMAP.md`](ROADMAP.md) for the full milestone table. Highlights:
459
+
460
+ - **M1 (v0.1, v0.1.1)** ✅ Shared storage with per-agent themes, async writer, connection pool, bulk import from `MEMORY.md`/`USER.md`
461
+ - **M2 (v0.2)** ✅ Conversation transcript table with `sync_turn` capture + `recall_conversation` tool
462
+ - **M3 (v0.3)** ✅ Identity propagation for stateless API minions via `X-Hermes-Session-Key`
463
+ - **M4 (v0.4)** ✅ Identity governance + `on_delegation()`/`on_session_end()` capture + agent attribution (`memory_agents`/`memory_agent_edges`), embedding backfill, conversation TTL, maintenance CLI
464
+ - **M5 (v0.5–v0.6)** ⏳ Decay scoring, partial HNSW indexes per-theme, Prometheus metrics, cross-provider bulk-import
465
+ - **M6 (v1.0)** ⏳ Stable config schema, full docs, CI coverage
466
+
467
+ The roadmap exists so the multi-agent positioning isn't a one-off claim — each milestone has to pass the test *"does this make N cooperating agents more capable?"* before it lands. The `What's not on the roadmap` section in `ROADMAP.md` lists what was deliberately rejected (LLM-mediated dialectic, fact-store ontologies, background derivers, in-plugin RBAC) so the boundaries are explicit.
468
+
469
+ ## Rollback
470
+
471
+ ```bash
472
+ hermes config set memory.provider none
473
+ sudo systemctl restart hermes.service
474
+
475
+ # Optional — drop the tables (data loss, irreversible)
476
+ sudo -u postgres psql -d <your-memory-db> -c "
477
+ DROP TABLE IF EXISTS conversations;
478
+ DROP TABLE IF EXISTS memory_entries;
479
+ "
480
+
481
+ # Optional — remove the plugin files
482
+ rm -rf ~/.hermes/plugins/pgvector
483
+ ```
484
+
485
+ ## Why a standalone plugin (not an upstream PR)?
486
+
487
+ Per the hermes-agent [`CONTRIBUTING.md`](https://github.com/NousResearch/hermes-agent/blob/main/CONTRIBUTING.md):
488
+
489
+ > We are no longer accepting new memory providers into this repo. The set of built-in providers under `plugins/memory/` is closed. If you want to add a new memory backend, publish it as a standalone plugin repo that users install into `~/.hermes/plugins/` (or via a pip entry point).
490
+
491
+ The discovery system (`plugins/memory/__init__.py` in hermes-agent) scans `$HERMES_HOME/plugins/<name>/` for any directory whose `__init__.py` calls `register_memory_provider`. This plugin's `hermes_pgvector/__init__.py` does exactly that — no upstream change required.
492
+
493
+ ## Contributing
494
+
495
+ Bug reports + PRs welcome. Open an issue describing the failure mode + your environment (hermes-agent version, Postgres version, embed endpoint), or a PR with a focused change + test.
496
+
497
+ ## License
498
+
499
+ [BSD 3-Clause](LICENSE) © 2026 Green Yoga Inc