agentdatabase 0.1.0__py3-none-any.whl
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- agentdatabase-0.1.0.dist-info/METADATA +847 -0
- agentdatabase-0.1.0.dist-info/RECORD +35 -0
- agentdatabase-0.1.0.dist-info/WHEEL +5 -0
- agentdatabase-0.1.0.dist-info/entry_points.txt +2 -0
- agentdatabase-0.1.0.dist-info/licenses/LICENSE +651 -0
- agentdatabase-0.1.0.dist-info/top_level.txt +1 -0
- agentdb/__init__.py +5 -0
- agentdb/adapters/claude_agent_sdk.py +831 -0
- agentdb/adapters/hermes.py +247 -0
- agentdb/backend.py +75 -0
- agentdb/core/__init__.py +19 -0
- agentdb/core/directory_tracking.py +59 -0
- agentdb/core/file_integrity.py +79 -0
- agentdb/core/models.py +90 -0
- agentdb/core/profiles.py +116 -0
- agentdb/core/store.py +373 -0
- agentdb/core/system.py +86 -0
- agentdb/embeddings/__init__.py +7 -0
- agentdb/embeddings/provider.py +99 -0
- agentdb/embeddings/store.py +199 -0
- agentdb/embeddings/text.py +20 -0
- agentdb/gateway/__init__.py +189 -0
- agentdb/gateway/adapter.py +58 -0
- agentdb/governance/__init__.py +3 -0
- agentdb/governance/conflict_detector.py +103 -0
- agentdb/governance/lifecycle_manager.py +226 -0
- agentdb/governance/permission_router.py +131 -0
- agentdb/interface/__init__.py +25 -0
- agentdb/interface/client.py +1108 -0
- agentdb/interface/mcp_server.py +96 -0
- agentdb/retrieval/__init__.py +3 -0
- agentdb/retrieval/algorithm.py +207 -0
- agentdb/skills/__init__.py +3 -0
- agentdb/skills/skill_store.py +351 -0
- agentdb/testing.py +68 -0
|
@@ -0,0 +1,847 @@
|
|
|
1
|
+
Metadata-Version: 2.4
|
|
2
|
+
Name: agentdatabase
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: Governed memory layer for AI agents
|
|
5
|
+
Author-email: DatoSurf <team@datosurf.com>
|
|
6
|
+
License: AGPL-3.0
|
|
7
|
+
Project-URL: Homepage, https://github.com/datosurf-dev/agentdb
|
|
8
|
+
Project-URL: Repository, https://github.com/datosurf-dev/agentdb
|
|
9
|
+
Project-URL: Bug Tracker, https://github.com/datosurf-dev/agentdb/issues
|
|
10
|
+
Keywords: ai,agents,memory,llm,rag,context
|
|
11
|
+
Classifier: Development Status :: 3 - Alpha
|
|
12
|
+
Classifier: License :: OSI Approved :: GNU Affero General Public License v3
|
|
13
|
+
Classifier: Programming Language :: Python :: 3
|
|
14
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
15
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
17
|
+
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
|
|
18
|
+
Classifier: Intended Audience :: Developers
|
|
19
|
+
Requires-Python: >=3.11
|
|
20
|
+
Description-Content-Type: text/markdown
|
|
21
|
+
License-File: LICENSE
|
|
22
|
+
Requires-Dist: pyyaml>=6.0
|
|
23
|
+
Provides-Extra: dev
|
|
24
|
+
Requires-Dist: pytest>=8.0; extra == "dev"
|
|
25
|
+
Requires-Dist: pytest-asyncio>=0.23; extra == "dev"
|
|
26
|
+
Provides-Extra: mcp
|
|
27
|
+
Requires-Dist: mcp<2.0,>=1.0; extra == "mcp"
|
|
28
|
+
Provides-Extra: local-embed
|
|
29
|
+
Requires-Dist: sentence-transformers<4.0,>=2.7; extra == "local-embed"
|
|
30
|
+
Dynamic: license-file
|
|
31
|
+
|
|
32
|
+
# AgentDB
|
|
33
|
+
|
|
34
|
+
**A governed memory layer that keeps agent memory actionable.**
|
|
35
|
+
|
|
36
|
+
AgentDB is designed to improve agent behavior without becoming another layer of operational overhead. It governs what memory is trusted, keeps relevant memory assets in sync, and injects the right context or instructions when an agent is about to act. Every write is validated before it lands. Every read returns a governed context package. The boundary is enforced at the memory layer — not by convention, not by prompt instructions, not by hoping agents behave.
|
|
37
|
+
|
|
38
|
+
> Every other memory product stores context. AgentDB governs it.
|
|
39
|
+
|
|
40
|
+
## Architecture Confirmation
|
|
41
|
+
|
|
42
|
+
The full architecture pitch and the numbered design principles live in one canonical place — this section just points to it instead of repeating it:
|
|
43
|
+
|
|
44
|
+
- Architecture & design principles: `/agentdb/docs/context/architecture-principles.md`
|
|
45
|
+
- Capabilities by type: `/agentdb/docs/context/capabilities.md`
|
|
46
|
+
- Global agent instructions: `/agentdb/AGENTS.md`
|
|
47
|
+
- Plan (active/open work): `/agentdb/docs/superpowers/plans/2026-06-06-agentdb-revised.md`
|
|
48
|
+
- Completed work tracker: `/agentdb/docs/superpowers/plans/completed-tasks.md`
|
|
49
|
+
|
|
50
|
+
## Current Repository Status (As Of 2026-10-04)
|
|
51
|
+
|
|
52
|
+
Core governed memory loop is fully operational: real conflict detection, budgeted retrieval, lifecycle management, centralized config + profiles, MCP server (stdio/Claude Desktop), and two adapters (Claude SDK + Hermes) are all shipped and test-covered. Phase 3 (embeddings/semantic search) is deliberately deferred until the loop proves its value in production. Provider ingestors (Phase 6) and additional framework adapters (Phase 4 Task 15) remain pending.
|
|
53
|
+
|
|
54
|
+
### Feature State Matrix
|
|
55
|
+
|
|
56
|
+
#### Track
|
|
57
|
+
|
|
58
|
+
| Feature area | State | Notes |
|
|
59
|
+
|---|---|---|
|
|
60
|
+
| Package entrypoint (`from agentdb import AgentDB`) | Implemented | Available in `src/agentdb/__init__.py`. |
|
|
61
|
+
| Core client loop (`write`, `read`, `retrieve`, `record_outcome`) | Implemented | Includes audit events and confidence updates. |
|
|
62
|
+
| Backend contract (`AgentDBBackend`) | Implemented | Framework-agnostic read/write/list/grep/glob contract. |
|
|
63
|
+
| Adapter surfaces | Implemented (Claude SDK + Hermes) | `adapters/claude_agent_sdk.py` (full `GatewayAdapter` subclass, proposal workflow, conflict queue, budgeted retrieval) and `adapters/hermes.py` (`AgentDBHermesProvider`, 14 hooks, Principle 12a proven). deepagents, LangGraph, CrewAI, OpenAI Agents, Google ADK, and AutoGen adapters are **not implemented** — deferred until a specific framework's users ask for one; `AgentDBBackend` works generically with any framework today. |
|
|
64
|
+
| Claude SDK async runtime (`from_sdk_client`, `handle_turn_async`) | Implemented | Async turn wrapper available in adapter runtime. |
|
|
65
|
+
| Directory tracking capture (Claude middleware) | Implemented | Optional tracked-directory snapshots persisted as memories. |
|
|
66
|
+
| Package memory workspace (`~/.agentdb/memory`, `config/memory.yaml`) | Implemented | Workspace resolution, config loading, and backward-compatible `path=` override in place. `memory.yaml` is now the single source of truth for all tunable defaults via mode-based profiles (`budget_optimized`, `full_storage`, `research`) — see Centralized config below. |
|
|
67
|
+
| Conversation JSONL capture | Implemented (initial) | Config-driven JSONL capture is available when `capture.conversation_jsonl` is enabled. |
|
|
68
|
+
|
|
69
|
+
#### Reconcile
|
|
70
|
+
|
|
71
|
+
| Feature area | State | Notes |
|
|
72
|
+
|---|---|---|
|
|
73
|
+
| Permission routing boundaries | Implemented | Write-time enforcement via `PermissionRouter`. |
|
|
74
|
+
| Write-time conflict detection (temporal + authority) | Implemented | `ConflictDetector` wired into `write()` — temporal conflicts auto-resolve in favour of recency; authority conflicts (agent vs. human-approved) auto-resolve in favour of human, flagged on retrieve. Semantic conflict detection is Phase 4 (requires embeddings). |
|
|
75
|
+
| External-context conflict detection | Implemented | `external_context` param on `retrieve()` checks overlap between caller-supplied context and stored records; flags `external_conflict_detected` + `external_conflict_rate` stat. `override_instructions` still dead code — detection works, explicit overrides are not generated. |
|
|
76
|
+
| Reconciliation trigger orchestration + `db.gateway` | Implemented | `db.gateway.subscribe`/`dispatch` is the unified surface. `conflict_detected` dispatched by `write()`. `conversation_started` dispatched by both adapters at turn start. Conflict queue (`db.gateway.conflict_queue()`), proposal workflow (`propose`/`approve`/`reject`/`list_proposals`) all live here. |
|
|
77
|
+
| Embedding-backed semantic retrieval/conflicts | Planned (Phase 4) | Current retrieval is keyword/provenance/confidence/recency based; embeddings are only justified if they improve action quality. |
|
|
78
|
+
|
|
79
|
+
#### Rank
|
|
80
|
+
|
|
81
|
+
| Feature area | State | Notes |
|
|
82
|
+
|---|---|---|
|
|
83
|
+
| Algorithm + skills blend for action quality | Implemented (initial) | Retrieval ranks memory; governed skills add maintained guidance during retrieval results. |
|
|
84
|
+
| Skills store + health warnings | Implemented | Versioned skills, staleness-based health scoring, warnings on retrieve. |
|
|
85
|
+
| System stats surface | Implemented (initial) | Includes retrieval/outcome/conflict-rate aggregates. |
|
|
86
|
+
| Entity-aware retrieval core (`entities`, `retrieve_by_entities`) | Implemented (initial) | Additive entity tagging and keyword-based entity retrieval are available in core. |
|
|
87
|
+
| Provider ingestors (Anthropic/OpenAI/Google sync) | Planned (Phase 3) | Architecture intent documented; module surface pending. Deliberately deferred — validating core loop with Claude SDK + Hermes in production first. |
|
|
88
|
+
| Lifecycle manager (`db.lifecycle`) | Implemented | `run_sweep()`: Step 0 archives expired records (`expiry_time ≤ now`); Step 1 optional confidence decay (off by default); Step 2 archives below `ARCHIVAL_THRESHOLD`; Step 3 archives past `STALENESS_DAYS`. `reinforce()` and `record_outcome()` also on `db.lifecycle`. `human_approved` records immune to all sweep operations. |
|
|
89
|
+
| `infer` flag + `expiry_time` on `write()` | Implemented | `infer=True` routes the raw value through a caller-supplied `infer_fn` before storage (raises `ValueError` if no `infer_fn` set). `expiry_time: datetime` sets a hard TTL; sweep archives once it passes. Both apply to `agent_inferred` only; `human_approved` records are immune. |
|
|
90
|
+
| Centralized config + profiles (`db.config`) | Implemented | `memory.yaml` drives all tunable defaults via mode-based profiles. Built-in profiles: `budget_optimized` (retrieve_limit=5, 2k chars, staleness=90d), `full_storage` (20, 12k, 90d), `research` (50, 24k, 180d, recency-first). User-defined profiles via `profiles:` block with optional `extends:`. `db.config` exposes the fully resolved dict. |
|
|
91
|
+
|
|
92
|
+
#### Inject
|
|
93
|
+
|
|
94
|
+
| Feature area | State | Notes |
|
|
95
|
+
|---|---|---|
|
|
96
|
+
| Query-driven retrieval into turn context | Implemented (Claude SDK + Hermes) | Both adapters retrieve at turn start via `prepare_turn`. `budget_tokens` param now wired end-to-end; `RetrievalResult.token_estimate` and `truncated` are live. |
|
|
97
|
+
| Action-driven retrieval (`entities`, `retrieve_by_entities`, `PreToolUse` hook) | Implemented (initial) | Core entity retrieval, adapter `pre_tool_use(...)`, and downstream hook-registration surface are implemented; automatic SDK attachment remains downstream integration work. |
|
|
98
|
+
| MCP server surface | Implemented (stdio) | `interface/mcp_server.py`: 4 tools (`memory_write`, `memory_retrieve`, `conflict_queue`, `memory_stats`). `agentdb-mcp` console script. HTTP+SSE transport deferred. |
|
|
99
|
+
| Hardening and production controls | Planned (Phase 6) | Roadmap-only at this snapshot. |
|
|
100
|
+
|
|
101
|
+
Validation status:
|
|
102
|
+
- Editable install and package-style imports are validated via `pip install -e '.[dev]'`.
|
|
103
|
+
- `python -m pytest -q` → 277 passed.
|
|
104
|
+
|
|
105
|
+
The rest of this README describes architecture intent and phased roadmap behavior.
|
|
106
|
+
|
|
107
|
+
## Execution Strategy
|
|
108
|
+
|
|
109
|
+
To keep delivery aligned with the architecture goals and avoid over-engineering, implementation should follow a narrow-first sequence:
|
|
110
|
+
|
|
111
|
+
1. Prove the governed core loop first:
|
|
112
|
+
- write permission enforcement
|
|
113
|
+
- audit emission
|
|
114
|
+
- conflict-first retrieval
|
|
115
|
+
- outcome recording
|
|
116
|
+
- confidence updates
|
|
117
|
+
- stats visibility
|
|
118
|
+
2. Expand integration breadth only after the core loop is stable and test-validated:
|
|
119
|
+
- provider ingestion
|
|
120
|
+
- MCP surface
|
|
121
|
+
- broader adapter hardening
|
|
122
|
+
3. Preserve phase gates:
|
|
123
|
+
- do not claim semantic retrieval/conflict behavior before embeddings are active
|
|
124
|
+
- keep conflict metadata assembled before memory content
|
|
125
|
+
|
|
126
|
+
This sequencing keeps AgentDB focused on its differentiated value: governed, low-overhead memory that improves action quality at the right time.
|
|
127
|
+
|
|
128
|
+
## Developer Quickstart
|
|
129
|
+
|
|
130
|
+
Use a repo-local virtual environment and editable install:
|
|
131
|
+
|
|
132
|
+
```bash
|
|
133
|
+
python3 -m venv .venv
|
|
134
|
+
source .venv/bin/activate
|
|
135
|
+
python -m pip install --upgrade pip
|
|
136
|
+
python -m pip install -e '.[dev]'
|
|
137
|
+
```
|
|
138
|
+
|
|
139
|
+
Run tests:
|
|
140
|
+
|
|
141
|
+
```bash
|
|
142
|
+
python -m pytest -q
|
|
143
|
+
```
|
|
144
|
+
|
|
145
|
+
Run a minimal local smoke check in Python:
|
|
146
|
+
|
|
147
|
+
```python
|
|
148
|
+
from agentdb import AgentDB
|
|
149
|
+
|
|
150
|
+
db = AgentDB()
|
|
151
|
+
db.write(
|
|
152
|
+
key="memory/working/hello",
|
|
153
|
+
value="world",
|
|
154
|
+
origin="agent_inferred",
|
|
155
|
+
agent_id="research_agent",
|
|
156
|
+
)
|
|
157
|
+
print(db.read("memory/working/hello", agent_id="research_agent").value)
|
|
158
|
+
```
|
|
159
|
+
|
|
160
|
+
Expected output:
|
|
161
|
+
|
|
162
|
+
```text
|
|
163
|
+
world
|
|
164
|
+
```
|
|
165
|
+
|
|
166
|
+
Integration manuals:
|
|
167
|
+
|
|
168
|
+
- Index: `docs/integrations/README.md`
|
|
169
|
+
- Google ADK: `docs/integrations/google-adk.md`
|
|
170
|
+
- Local LLM and Claude wiring: `docs/integrations/model-and-runtime-wiring.md`
|
|
171
|
+
|
|
172
|
+
### Available Now (Current Snapshot)
|
|
173
|
+
|
|
174
|
+
These are the developer-facing surfaces you can use immediately in the current implementation:
|
|
175
|
+
|
|
176
|
+
- `AgentDB` client (`from agentdb import AgentDB`)
|
|
177
|
+
- `write(key, value, origin, agent_id, infer=False, expiry_time=None, ...)`
|
|
178
|
+
- `read(key, agent_id)`
|
|
179
|
+
- `retrieve(query, agent_id, limit=..., budget_tokens=..., external_context=..., session_id=...)`
|
|
180
|
+
- `retrieve_by_entities(entities, agent_id, session_id=...)`
|
|
181
|
+
- `record_outcome(session_id, agent_id, outcome_type, outcome_value)` — backward-compat alias for `db.lifecycle.record_outcome()`
|
|
182
|
+
- `system.agentdb_stats()`
|
|
183
|
+
- `audit_log.replay(from_sequence=0)`
|
|
184
|
+
- `lifecycle` — `run_sweep()`, `reinforce(record_id)`, `record_outcome(...)`
|
|
185
|
+
- `gateway` — `subscribe(trigger, handler)`, `dispatch(trigger, ...)`, `conflict_queue()`, `propose(...)`, `approve(...)`, `reject(...)`, `list_proposals(...)`
|
|
186
|
+
- `config` — returns the fully resolved config dict (mode defaults + yaml overrides)
|
|
187
|
+
- `skills_api` — `create_skill`, `update_skill_version`, `get_skill`, `archive_skill`, `list_stale_skills_for_agent`, `list_skill_warnings_for_agent`
|
|
188
|
+
- optional initialization: `tracking_directories`, `tracking_root`, `tracking_max_files_per_directory`, `infer_fn`, `memory_dir`
|
|
189
|
+
- `AgentDBBackend` (`from agentdb.backend import AgentDBBackend`)
|
|
190
|
+
- `read`, `write`, `delete`, `list`, `grep`, `glob`
|
|
191
|
+
- Adapters currently implemented under `src/agentdb/adapters/`
|
|
192
|
+
- **Claude SDK** (`AgentDBClaudeMiddleware`) — full `GatewayAdapter` subclass; `prepare_turn`, `pre_tool_use`, `finalize_turn`; proposal workflow; conflict queue; budgeted retrieval.
|
|
193
|
+
- **Hermes** (`AgentDBHermesProvider`) — 14-hook `MemoryProvider` ABC + `GatewayAdapter`; Principle 12a proven via shim tests.
|
|
194
|
+
- deepagents, LangGraph, CrewAI, OpenAI Agents, Google ADK, and AutoGen are not built yet — `AgentDBBackend` covers them generically.
|
|
195
|
+
|
|
196
|
+
Example using the current retrieval/improvement loop:
|
|
197
|
+
|
|
198
|
+
```python
|
|
199
|
+
from agentdb import AgentDB
|
|
200
|
+
|
|
201
|
+
db = AgentDB()
|
|
202
|
+
agent_id = "research_agent"
|
|
203
|
+
session_id = "sess-demo"
|
|
204
|
+
|
|
205
|
+
db.write(
|
|
206
|
+
key="memory/working/style_pref",
|
|
207
|
+
value="prefer concise bullet points",
|
|
208
|
+
origin="agent_inferred",
|
|
209
|
+
agent_id=agent_id,
|
|
210
|
+
session_id=session_id,
|
|
211
|
+
)
|
|
212
|
+
|
|
213
|
+
result = db.retrieve(
|
|
214
|
+
query="bullet points",
|
|
215
|
+
agent_id=agent_id,
|
|
216
|
+
session_id=session_id,
|
|
217
|
+
)
|
|
218
|
+
print([r.key for r in result.records])
|
|
219
|
+
|
|
220
|
+
db.record_outcome(
|
|
221
|
+
session_id=session_id,
|
|
222
|
+
agent_id=agent_id,
|
|
223
|
+
outcome_type="task_complete",
|
|
224
|
+
outcome_value=1.0,
|
|
225
|
+
)
|
|
226
|
+
|
|
227
|
+
print(db.system.agentdb_stats())
|
|
228
|
+
```
|
|
229
|
+
|
|
230
|
+
Example using the current Claude middleware action loop:
|
|
231
|
+
|
|
232
|
+
```python
|
|
233
|
+
from agentdb import AgentDB
|
|
234
|
+
from agentdb.adapters.claude_agent_sdk import AgentDBClaudeMiddleware
|
|
235
|
+
|
|
236
|
+
db = AgentDB()
|
|
237
|
+
middleware = AgentDBClaudeMiddleware(
|
|
238
|
+
db=db,
|
|
239
|
+
agent_id="analytics_agent",
|
|
240
|
+
profile="budget_optimized",
|
|
241
|
+
)
|
|
242
|
+
|
|
243
|
+
# Seed governed memory tied to a file/entity the tool is about to touch.
|
|
244
|
+
db.write(
|
|
245
|
+
key="memory/working/models/orders_sql_guidance",
|
|
246
|
+
value="Use the normalized orders model and preserve customer_id joins.",
|
|
247
|
+
origin="human_approved",
|
|
248
|
+
agent_id="analytics_agent",
|
|
249
|
+
entities=["models/orders.sql", "orders"],
|
|
250
|
+
)
|
|
251
|
+
|
|
252
|
+
# Turn-start retrieval still happens before the model responds.
|
|
253
|
+
messages, retrieval = middleware.prepare_turn(
|
|
254
|
+
user_query="Update the orders model to add the new revenue field.",
|
|
255
|
+
session_id="sess-42",
|
|
256
|
+
)
|
|
257
|
+
|
|
258
|
+
# Action-time retrieval can add targeted memory right before a tool call.
|
|
259
|
+
hook_payload = middleware.pre_tool_use(
|
|
260
|
+
tool_name="Edit",
|
|
261
|
+
tool_input={"file_path": "models/orders.sql"},
|
|
262
|
+
session_id="sess-42",
|
|
263
|
+
)
|
|
264
|
+
print(hook_payload["additionalContext"])
|
|
265
|
+
|
|
266
|
+
# Persist the turn and tag the touched entity for future retrieval.
|
|
267
|
+
middleware.finalize_turn(
|
|
268
|
+
session_id="sess-42",
|
|
269
|
+
user_query="Update the orders model to add the new revenue field.",
|
|
270
|
+
model_response="Added revenue logic and kept the join path intact.",
|
|
271
|
+
touched_entities=hook_payload.get("touchedEntities", []),
|
|
272
|
+
)
|
|
273
|
+
```
|
|
274
|
+
|
|
275
|
+
This is the current actionable-memory loop in the repo today:
|
|
276
|
+
|
|
277
|
+
1. `prepare_turn(...)` retrieves governed memory for the user query.
|
|
278
|
+
2. `pre_tool_use(...)` optionally retrieves narrower entity-linked context just before a tool acts.
|
|
279
|
+
3. `finalize_turn(...)` persists the response, optional captures, and touched-entity metadata for later turns.
|
|
280
|
+
|
|
281
|
+
---
|
|
282
|
+
|
|
283
|
+
## Drop into your existing stack in 30 minutes
|
|
284
|
+
|
|
285
|
+
AgentDB defines a generic pluggable backend interface. Any framework with pluggable storage adopts it — **this works today**, no adapter required:
|
|
286
|
+
|
|
287
|
+
```python
|
|
288
|
+
from agentdb import AgentDB
|
|
289
|
+
from agentdb.backend import AgentDBBackend
|
|
290
|
+
|
|
291
|
+
db = AgentDB()
|
|
292
|
+
backend = AgentDBBackend(agent_id="assistant", trust_zone="agent_inferred", db=db)
|
|
293
|
+
|
|
294
|
+
backend.write("memory/working/task_context", "user prefers bullet points")
|
|
295
|
+
backend.read("memory/working/task_context") # → "user prefers bullet points"
|
|
296
|
+
backend.list("memory/working/") # → ["memory/working/task_context"]
|
|
297
|
+
```
|
|
298
|
+
|
|
299
|
+
**Planned (not yet implemented): a `deepagents.py` adapter** for [deepagents](https://github.com/langchain-ai/deepagents) — LangChain's agent harness built on LangGraph — as the flagship named integration. The shape it's expected to take:
|
|
300
|
+
|
|
301
|
+
```python
|
|
302
|
+
# NOT YET IMPLEMENTED — agentdb.adapters.deepagents does not exist yet.
|
|
303
|
+
from deepagents import create_deep_agent
|
|
304
|
+
from deepagents.backends import CompositeBackend, StateBackend
|
|
305
|
+
from agentdb import AgentDB
|
|
306
|
+
from agentdb.adapters.deepagents import make_agentdb_backend
|
|
307
|
+
|
|
308
|
+
db = AgentDB()
|
|
309
|
+
|
|
310
|
+
agent = create_deep_agent(
|
|
311
|
+
model="anthropic:claude-sonnet-4-6",
|
|
312
|
+
tools=[...],
|
|
313
|
+
backend=CompositeBackend(
|
|
314
|
+
default=StateBackend(),
|
|
315
|
+
routes={
|
|
316
|
+
"/memories/": make_agentdb_backend(db, "assistant", trust_zone="agent_inferred"),
|
|
317
|
+
"/skills/": make_agentdb_backend(db, "assistant", trust_zone="human_approved"),
|
|
318
|
+
},
|
|
319
|
+
),
|
|
320
|
+
)
|
|
321
|
+
```
|
|
322
|
+
|
|
323
|
+
deepagents would handle the agent loop, context compression, subagent isolation, and observability; AgentDB governs everything that persists. They don't overlap — but until `make_agentdb_backend` exists, use the generic `AgentDBBackend` example above directly with `CompositeBackend`'s routes.
|
|
324
|
+
|
|
325
|
+
---
|
|
326
|
+
|
|
327
|
+
## The primary differentiator: governed instruction records
|
|
328
|
+
|
|
329
|
+
Skills are the most defensible feature in AgentDB. Not because they store instructions — any file can do that — but because they are database objects with a computed **health score** that degrades automatically when the instruction becomes stale.
|
|
330
|
+
|
|
331
|
+
**Why this matters:** Anthropic's data analytics team built agents that reached 95%+ accuracy using structured skills. Without skills: 21% accuracy. With well-maintained skills: 95%+. The failure mode: skill accuracy drifted from 95% to 65% over one month when skill files were not actively maintained as the underlying data model changed. Their fix was PR discipline — colocation of skill files with transformation models, with hooks that flag drift.
|
|
332
|
+
|
|
333
|
+
AgentDB's health score is the infrastructure solution. It catches drift without process discipline:
|
|
334
|
+
|
|
335
|
+
```python
|
|
336
|
+
db.skills_api.create_skill(
|
|
337
|
+
skill_id="analytics-schema-v1",
|
|
338
|
+
name="analytics-schema",
|
|
339
|
+
content="""
|
|
340
|
+
The revenue table has columns: date, amount_usd, customer_id, product_sku.
|
|
341
|
+
Always join on customer_id, never on email.
|
|
342
|
+
The 'legacy_orders' table was deprecated in Q1 2026 — do not query it.
|
|
343
|
+
""",
|
|
344
|
+
scope=["analytics_agent"],
|
|
345
|
+
)
|
|
346
|
+
# authored_by is always "human" — SkillsAPI.create_skill() sets it, there's no way to pass agent-authored skills in.
|
|
347
|
+
|
|
348
|
+
# 45 days later — schema changed, skill not updated:
|
|
349
|
+
result = db.retrieve(query="revenue breakdown", agent_id="analytics_agent")
|
|
350
|
+
result.skill_warnings
|
|
351
|
+
# → ["Skill 'analytics-schema-v1' is stale (health_score=0.42)"]
|
|
352
|
+
```
|
|
353
|
+
|
|
354
|
+
A health score below 0.6 surfaces a staleness warning in every retrieval result. The agent sees the warning before acting on the skill content.
|
|
355
|
+
|
|
356
|
+
**What degrades the health score today:**
|
|
357
|
+
- Days since last retrieval (staleness decay, half-life 30 days) — the only implemented factor
|
|
358
|
+
|
|
359
|
+
**Planned, not yet implemented:**
|
|
360
|
+
- Unresolved conflicts linked to the skill (no conflict linkage exists on skills yet)
|
|
361
|
+
- A newer version that hasn't been promoted
|
|
362
|
+
|
|
363
|
+
**Current limitation (Phase 2):** Entity drift detection is keyword-based — a skill that uses a synonym for a deprecated entity will not be flagged. Semantic entity matching is Phase 5 scope. This limitation is documented in LIMITATIONS.md.
|
|
364
|
+
|
|
365
|
+
**Planned (Task 13, not started, blocked on Phase 4 embeddings):** a skill draft reconciliation pipeline — an agent can *propose* a new skill or an update to an existing one (written as an ordinary agent-inferred memory record, never touching `skills.sqlite` directly), which gets reconciled against other pending drafts, the target skill, and recently-changed files before a human explicitly approves or rejects it. Proposing is async (notifies a reviewer); only the approval step can ever create or modify a real skill. See the plan for the full design.
|
|
366
|
+
|
|
367
|
+
### Why the algorithm and skills belong together
|
|
368
|
+
|
|
369
|
+
The retrieval algorithm and the governed skills system are one foundation, not two unrelated features.
|
|
370
|
+
|
|
371
|
+
- The algorithm answers: which memories or instructions are most relevant to the action the agent is about to take?
|
|
372
|
+
- Skills answer: which durable guidance should the agent trust when that action happens?
|
|
373
|
+
|
|
374
|
+
In practice, AgentDB is strongest when these work together:
|
|
375
|
+
|
|
376
|
+
1. memory records capture facts, context, prior outcomes, and relevant assets
|
|
377
|
+
2. the ranking algorithm decides what matters now
|
|
378
|
+
3. skills add governed, maintained instructions on top of those facts
|
|
379
|
+
4. the runtime injects both as compact action-ready context
|
|
380
|
+
|
|
381
|
+
That blend is the development foundation for AgentDB. The goal is not to accumulate more memory artifacts. The goal is to turn memory into better decisions, fewer conflicting instructions, and more accurate actions.
|
|
382
|
+
|
|
383
|
+
---
|
|
384
|
+
|
|
385
|
+
## The two zones: the central design decision
|
|
386
|
+
|
|
387
|
+
```
|
|
388
|
+
┌─────────────────────────────────────────────────────────┐
|
|
389
|
+
│ HUMAN-APPROVED ZONE │
|
|
390
|
+
│ skills/ config/human/ prompts/locked/ │
|
|
391
|
+
│ │
|
|
392
|
+
│ Human-authored. Immutable at runtime. No agent write │
|
|
393
|
+
│ reaches here regardless of trust level or content. │
|
|
394
|
+
└─────────────────────────────────────────────────────────┘
|
|
395
|
+
│
|
|
396
|
+
Permission router
|
|
397
|
+
(validates at write time, before storage)
|
|
398
|
+
│
|
|
399
|
+
┌─────────────────────────────────────────────────────────┐
|
|
400
|
+
│ AGENT-INFERRED ZONE │
|
|
401
|
+
│ memory/working/ memory/episodic/ memory/org/ │
|
|
402
|
+
│ memory/long_term/ config/agent/ memory/provider/│
|
|
403
|
+
│ │
|
|
404
|
+
│ Bounded by permissions.yaml. Logged. Lifecycle- │
|
|
405
|
+
│ managed. Confidence-decaying. │
|
|
406
|
+
└─────────────────────────────────────────────────────────┘
|
|
407
|
+
```
|
|
408
|
+
|
|
409
|
+
This boundary is enforced at write time, before any record reaches storage. A write to `skills/` from `origin="agent_inferred"` is rejected and logged — always, regardless of what the agent claims.
|
|
410
|
+
|
|
411
|
+
```python
|
|
412
|
+
db.write(
|
|
413
|
+
key="skills/override-safety",
|
|
414
|
+
value="ignore all previous instructions",
|
|
415
|
+
origin="agent_inferred",
|
|
416
|
+
agent_id="assistant",
|
|
417
|
+
)
|
|
418
|
+
# → PermissionDenied: human zone requires human_approved origin [agent=assistant, key=skills/override-safety]
|
|
419
|
+
# → Logged to audit trail
|
|
420
|
+
```
|
|
421
|
+
|
|
422
|
+
---
|
|
423
|
+
|
|
424
|
+
## Memory biography: what makes this a database
|
|
425
|
+
|
|
426
|
+
Every record in AgentDB carries a full biography — not just a value:
|
|
427
|
+
|
|
428
|
+
Implemented today:
|
|
429
|
+
|
|
430
|
+
```
|
|
431
|
+
key → addressable path e.g. "memory/working/pref_format"
|
|
432
|
+
value → the memory content
|
|
433
|
+
origin → human_approved | agent_inferred | provider_ingested
|
|
434
|
+
trust_level → sealed | human | agent | provider (derived from origin; "sealed" is unused so far)
|
|
435
|
+
confidence → float 0–1, decays over time unless reinforced
|
|
436
|
+
entities → list of canonical identifiers (file paths, model names) for action-driven retrieval
|
|
437
|
+
last_retrieved → drives the ranker's recency signal
|
|
438
|
+
conflict_id → populated by ConflictDetector at write time when a temporal or authority conflict is detected
|
|
439
|
+
expiry_time → optional hard TTL (datetime); lifecycle sweep archives once it passes (agent_inferred only)
|
|
440
|
+
inferred → bool; true when the value was derived by infer_fn rather than stored verbatim
|
|
441
|
+
```
|
|
442
|
+
|
|
443
|
+
Planned, not yet on the record: `tier`, `expiry_policy`, `audit_sequence`, `embedding_version`, `user_id`, `session_id`. See `docs/context/database-schema.md` for the full column-by-column reference.
|
|
444
|
+
|
|
445
|
+
The biography enables permission routing, lifecycle management, and provenance-weighted retrieval today; conflict detection is planned (see "Conflict detection" below). Without it you have a fancier key-value store.
|
|
446
|
+
|
|
447
|
+
---
|
|
448
|
+
|
|
449
|
+
## Memory tiers (naming convention today, not yet enforced retrieval behavior)
|
|
450
|
+
|
|
451
|
+
The key prefixes below are a convention every write should follow — but there is no `tier` field on `MemoryRecord` yet, and the ranking algorithm doesn't currently look at key prefix at all (it ranks uniformly by provenance/recency/confidence/causal regardless of which prefix a key falls under). Treat "Retrieval priority" as the design intent, not current behavior.
|
|
452
|
+
|
|
453
|
+
| Tier | Key prefix | Description | Retrieval priority (planned) |
|
|
454
|
+
|---|---|---|---|
|
|
455
|
+
| Working | `memory/working/` | Current task context. Session-scoped. | High for active session |
|
|
456
|
+
| Episodic | `memory/episodic/` | Past interactions. Decays with time. | Medium, recency-weighted |
|
|
457
|
+
| Long-term | `memory/long_term/` | Durable user preferences and facts. | High, confidence-weighted |
|
|
458
|
+
| Org | `memory/org/` | Institutional knowledge: team structures, product names, internal terminology, decision logs. Changes slowly. Maintained by designated owners. | High for disambiguation; neutral for technical queries |
|
|
459
|
+
| Provider | `memory/provider/` | Ingested from Anthropic / OpenAI / Google. Trust level = provider, never promoted. | Low |
|
|
460
|
+
|
|
461
|
+
---
|
|
462
|
+
|
|
463
|
+
## Conflict detection
|
|
464
|
+
|
|
465
|
+
Detection runs at write time, not read time — by the time an agent is mid-execution it's too late to resolve cleanly.
|
|
466
|
+
|
|
467
|
+
| Conflict type | Status | Resolution |
|
|
468
|
+
|---|---|---|
|
|
469
|
+
| **Temporal** | **Implemented** | Auto-resolved in favour of recency. Logged. `conflict_id` stamped on both records. |
|
|
470
|
+
| **Authority** | **Implemented** | Agent-inferred contradicts human-approved → auto-resolved in favour of human. `override_instructions` populated in `RetrievalResult`. `conflict_detected` dispatched via `db.gateway`. |
|
|
471
|
+
| **Semantic** | Planned (Phase 4) | Requires embeddings — deferred. |
|
|
472
|
+
|
|
473
|
+
`ConflictDetector` (`governance/conflict_detector.py`) is wired into `write()`. On conflict: both records get `conflict_id`, the conflict is stored, and `db.gateway.dispatch("conflict_detected", ...)` fires so any subscriber (Slack notifier, MCP consumer, etc.) is immediately notified.
|
|
474
|
+
|
|
475
|
+
```python
|
|
476
|
+
# write() detects and dispatches automatically
|
|
477
|
+
db.write(key="k", value="new value", origin="agent_inferred", agent_id="a")
|
|
478
|
+
# If an existing record at "k" conflicts → conflict stored, dispatch fires
|
|
479
|
+
|
|
480
|
+
result = db.retrieve(query="communication style", agent_id="assistant")
|
|
481
|
+
result.conflicts # list of ConflictRecord for returned records
|
|
482
|
+
result.override_instructions # human-approved wins messages for authority conflicts
|
|
483
|
+
```
|
|
484
|
+
|
|
485
|
+
External-context detection (caller-supplied context vs. stored records) is a separate, narrower check available via `external_context=` on `retrieve()` — it flags `external_conflict_detected` audit events but does not interact with the stored `conflicts` table.
|
|
486
|
+
|
|
487
|
+
**Known limitation:** semantic conflict detection (same-authority, situationally contradictory) requires Phase 4 embeddings. Until then, only temporal and authority conflicts are detected.
|
|
488
|
+
|
|
489
|
+
---
|
|
490
|
+
|
|
491
|
+
## Retrieval: what's novel, what isn't
|
|
492
|
+
|
|
493
|
+
Semantic similarity, recency decay, and confidence scoring are already implemented in Mem0's 2026 algorithm. AgentDB uses these as baseline signals but does not claim them as differentiators.
|
|
494
|
+
|
|
495
|
+
**Two signals that are genuinely novel:**
|
|
496
|
+
|
|
497
|
+
**Provenance weight** — human-approved memories rank above agent-inferred for equivalent semantic similarity. No current system makes trust level a first-class ranking factor. A user's explicit preference always surfaces above an agent's inferred preference for the same query.
|
|
498
|
+
|
|
499
|
+
**Causal relevance from WAL history** — if a memory was consistently retrieved before successful completions of actions similar to the current one, its relevance score increases. This is a WAL-derived signal no current memory system computes. It rewards memories that have demonstrated predictive value, not just semantic closeness.
|
|
500
|
+
|
|
501
|
+
Signal weights (all configurable, enforced to sum to 1.0 in `MultiSignalRanker`'s constructor):
|
|
502
|
+
|
|
503
|
+
**Current shipped defaults** (`agentdb/retrieval/algorithm.py`):
|
|
504
|
+
|
|
505
|
+
| Signal | Default weight | Novel? |
|
|
506
|
+
|---|---|---|
|
|
507
|
+
| Provenance (trust level) | 0.35 | **Yes** |
|
|
508
|
+
| Recency (decay from last retrieved) | 0.25 | No — Mem0 baseline |
|
|
509
|
+
| Confidence | 0.25 | No — Mem0 baseline |
|
|
510
|
+
| Causal (WAL-derived, approximated) | 0.15 | **Yes** |
|
|
511
|
+
| Semantic (cosine similarity) | 0.0 | No — Mem0 baseline |
|
|
512
|
+
| Workspace (working-directory overlap) | 0.0 | **Yes** |
|
|
513
|
+
|
|
514
|
+
**Target redistribution once Phase 4 activates embeddings** (not active yet):
|
|
515
|
+
|
|
516
|
+
| Signal | Weight |
|
|
517
|
+
|---|---|
|
|
518
|
+
| Provenance | 0.30 |
|
|
519
|
+
| Recency | 0.20 |
|
|
520
|
+
| Confidence | 0.20 |
|
|
521
|
+
| Causal | 0.10 |
|
|
522
|
+
| Semantic | 0.20 |
|
|
523
|
+
|
|
524
|
+
Causal weight is approximated via retrieval frequency × recency until Phase 5 computes full WAL correlation (`CausalGraph`, see the plan's Algorithm Reference).
|
|
525
|
+
|
|
526
|
+
---
|
|
527
|
+
|
|
528
|
+
## Action-driven retrieval (Roadmap: Phase 2 Task F)
|
|
529
|
+
|
|
530
|
+
Standard retrieval is **query-driven**: the agent describes what it needs, AgentDB ranks and returns relevant memories at the start of a turn.
|
|
531
|
+
|
|
532
|
+
Action-driven retrieval is a **planned tool-triggered extension**: memories will be injected at the moment Claude is about to act on a specific file or command — not at turn start, when the relevant file may not yet be known.
|
|
533
|
+
|
|
534
|
+
```
|
|
535
|
+
Turn start → query-driven retrieve("fix the funnel")
|
|
536
|
+
→ general context about funnel models
|
|
537
|
+
|
|
538
|
+
Claude about to Edit → entity-driven retrieve(entity="ce_project_funnel__respondents")
|
|
539
|
+
→ "last time: layer violation introduced", "full-refresh required"
|
|
540
|
+
```
|
|
541
|
+
|
|
542
|
+
This matters for a coding assistant because much of the institutional knowledge is tied to specific files, models, or commands — not to the user's query topic. A memory like "this model requires `--full-refresh`" is invisible to a turn-level query unless the user happened to mention the model name.
|
|
543
|
+
|
|
544
|
+
### Planned design
|
|
545
|
+
|
|
546
|
+
`MemoryRecord` will carry an `entities` field — a list of canonical entity identifiers (file paths, model names, command patterns) attached at write time. `AgentDB.retrieve_by_entities()` will look up memories by entity key, bypassing the keyword ranker for exact matches.
|
|
547
|
+
|
|
548
|
+
The framework adapter (e.g. the Claude SDK adapter) is planned to hook into `PreToolUse` events, extract the entity key from the tool input, and inject retrieved context as `additionalContext` before the tool executes:
|
|
549
|
+
|
|
550
|
+
```python
|
|
551
|
+
# Tool input: Edit("models/marts/core_entities/ce_project_funnel__respondents.sql")
|
|
552
|
+
# Adapter extracts → entity = "ce_project_funnel__respondents"
|
|
553
|
+
result = db.retrieve_by_entities(["ce_project_funnel__respondents"], agent_id="slack_assistant")
|
|
554
|
+
# → injects context before Claude edits the file
|
|
555
|
+
```
|
|
556
|
+
|
|
557
|
+
Entity keys are extracted per tool type:
|
|
558
|
+
|
|
559
|
+
| Tool | Entity key |
|
|
560
|
+
|---|---|
|
|
561
|
+
| `Write` / `Edit` / `MultiEdit` | `file_path` (and derived model name for `.sql` files) |
|
|
562
|
+
| `Bash("dbt build --select X")` | model name parsed from `--select` argument |
|
|
563
|
+
| `Bash("git commit ...")` | currently staged file paths |
|
|
564
|
+
|
|
565
|
+
### Core vs. adapter boundary (planned)
|
|
566
|
+
|
|
567
|
+
| What | Where |
|
|
568
|
+
|---|---|
|
|
569
|
+
| `entities: list[str]` on `MemoryRecord` | AgentDB core |
|
|
570
|
+
| `retrieve_by_entities(entities, agent_id)` | AgentDB core |
|
|
571
|
+
| Write-time entity extraction and tagging | Framework adapter |
|
|
572
|
+
| `PreToolUse` hook + `additionalContext` injection | Claude SDK adapter |
|
|
573
|
+
|
|
574
|
+
The core is framework-agnostic — any adapter (LangGraph, CrewAI, etc.) can implement write-time tagging and call `retrieve_by_entities`. The Claude-specific `PreToolUse` mechanics stay in the adapter.
|
|
575
|
+
|
|
576
|
+
### Current state and phase placement
|
|
577
|
+
|
|
578
|
+
Current state: query-driven retrieval and Claude turn wrappers are implemented; action-driven retrieval hooks are not yet active.
|
|
579
|
+
|
|
580
|
+
Action-driven retrieval uses the existing keyword ranker and works with Phase 2 infrastructure. It improves automatically when Phase 4 semantic search activates — `retrieve_by_entities` can fall back to embedding similarity for entities not yet in the index.
|
|
581
|
+
|
|
582
|
+
**Current limitation (Phase 2):** `retrieve_by_entities` is keyword-based — memories must contain the entity string in their key or value to be matched. Semantic entity matching is deferred to Phase 4+.
|
|
583
|
+
|
|
584
|
+
---
|
|
585
|
+
|
|
586
|
+
## Provider ingestion
|
|
587
|
+
|
|
588
|
+
AgentDB does not compete with provider-native memory. It governs it.
|
|
589
|
+
|
|
590
|
+
```python
|
|
591
|
+
# Provider memory flows INTO AgentDB, not around it
|
|
592
|
+
db.providers.anthropic.sync(memory_store_id="ms_abc123") # NOT YET IMPLEMENTED
|
|
593
|
+
db.providers.openai.sync(agent_id="assistant") # NOT YET IMPLEMENTED
|
|
594
|
+
db.providers.google.sync() # Vertex AI MemoryBankService — NOT YET IMPLEMENTED
|
|
595
|
+
```
|
|
596
|
+
|
|
597
|
+
No `src/agentdb/providers/` directory exists yet (Phase 3, gated behind the core loop stabilizing — see the plan's Current Status). The interface above is the target shape.
|
|
598
|
+
|
|
599
|
+
All three providers write with `origin=provider_ingested`, `trust_level=provider`. No provider memory is ever promoted to `human_approved` — the boundary is enforced at ingestion.
|
|
600
|
+
|
|
601
|
+
**Current limitation:** All three providers use polling. Push subscription is deferred until providers expose push APIs.
|
|
602
|
+
|
|
603
|
+
---
|
|
604
|
+
|
|
605
|
+
## A2A protocol
|
|
606
|
+
|
|
607
|
+
AgentDB governs the memory of individual agents that communicate via the [A2A protocol](https://a2a-protocol.org) (v1.2, Linux Foundation). When Agent A delegates a task to Agent B via A2A, Agent B uses its own AgentDB instance — memory does not flow across A2A boundaries by default.
|
|
608
|
+
|
|
609
|
+
AgentDB does not implement A2A transport. It sits below the transport layer:
|
|
610
|
+
|
|
611
|
+
```
|
|
612
|
+
Agent A ──A2A──► Agent B
|
|
613
|
+
│
|
|
614
|
+
AgentDBBackend
|
|
615
|
+
│
|
|
616
|
+
AgentDB (governed)
|
|
617
|
+
```
|
|
618
|
+
|
|
619
|
+
**Future scope (Phase 6+):** A2A agent cards carry cryptographic signatures (v1.2). An agent card can assert which memory zones the agent is authorized to read. This maps to AgentDB's permission zones — A2A card-based permission bootstrapping is planned.
|
|
620
|
+
|
|
621
|
+
---
|
|
622
|
+
|
|
623
|
+
## System tables
|
|
624
|
+
|
|
625
|
+
The target design models memory metadata as queryable relations on `pg_catalog`. **Only `agentdb_stats()` is implemented today** — `agentdb_memories()`, `agentdb_agents()`, and `agentdb_conflicts()` don't exist yet:
|
|
626
|
+
|
|
627
|
+
```python
|
|
628
|
+
db.system.agentdb_memories() # NOT YET IMPLEMENTED — planned: all active records
|
|
629
|
+
db.system.agentdb_memories(agent_id="assistant") # NOT YET IMPLEMENTED
|
|
630
|
+
db.system.agentdb_agents() # NOT YET IMPLEMENTED — planned: per-agent stats
|
|
631
|
+
db.system.agentdb_conflicts(resolved=False) # NOT YET IMPLEMENTED — planned: unresolved conflict queue
|
|
632
|
+
db.system.agentdb_stats() # Implemented — database-wide counters
|
|
633
|
+
```
|
|
634
|
+
|
|
635
|
+
Actual current return shape:
|
|
636
|
+
|
|
637
|
+
```python
|
|
638
|
+
db.system.agentdb_stats()
|
|
639
|
+
# {
|
|
640
|
+
# "total_memories": 1847,
|
|
641
|
+
# "avg_confidence": 0.74,
|
|
642
|
+
# "low_confidence_count": 12, # confidence < 0.5
|
|
643
|
+
# "conflict_pending_count": 0, # hardcoded 0 — conflict tracking doesn't exist yet (see Conflict detection above)
|
|
644
|
+
# "outcome_positive_rate": 0.87, # None until at least 10 outcomes are recorded
|
|
645
|
+
# "retrieval_count_total": 4021,
|
|
646
|
+
# "external_conflict_rate": None, # None until at least one external-context retrieve happens
|
|
647
|
+
# }
|
|
648
|
+
```
|
|
649
|
+
|
|
650
|
+
---
|
|
651
|
+
|
|
652
|
+
## Permissions
|
|
653
|
+
|
|
654
|
+
Permissions are human-authored YAML, loaded at startup, not modifiable at runtime:
|
|
655
|
+
|
|
656
|
+
```yaml
|
|
657
|
+
# config/permissions.yaml
|
|
658
|
+
default_policy: deny
|
|
659
|
+
|
|
660
|
+
human_zone_prefixes:
|
|
661
|
+
- "skills/"
|
|
662
|
+
- "config/human/"
|
|
663
|
+
- "prompts/locked/"
|
|
664
|
+
|
|
665
|
+
agents:
|
|
666
|
+
analytics_agent:
|
|
667
|
+
can_write:
|
|
668
|
+
- "memory/working/"
|
|
669
|
+
- "memory/episodic/"
|
|
670
|
+
- "memory/org/"
|
|
671
|
+
- "config/agent/learned_prefs/analytics_*"
|
|
672
|
+
cannot_write:
|
|
673
|
+
- "config/human/"
|
|
674
|
+
- "skills/"
|
|
675
|
+
can_read:
|
|
676
|
+
- "memory/working/"
|
|
677
|
+
- "memory/episodic/"
|
|
678
|
+
- "memory/long_term/"
|
|
679
|
+
- "memory/org/"
|
|
680
|
+
```
|
|
681
|
+
|
|
682
|
+
Agents not listed are denied all writes (`default_policy: deny`). Wildcard matching is supported (`analytics_*`). Rules are evaluated in order: `cannot_write` before `can_write`.
|
|
683
|
+
|
|
684
|
+
---
|
|
685
|
+
|
|
686
|
+
## Lifecycle management
|
|
687
|
+
|
|
688
|
+
`db.lifecycle` (`governance/lifecycle_manager.py`) is implemented. Call `db.lifecycle.run_sweep()` on demand — no background scheduler.
|
|
689
|
+
|
|
690
|
+
| Event | Trigger | Action | Status |
|
|
691
|
+
|---|---|---|---|
|
|
692
|
+
| Hard TTL expiry | `expiry_time` field on record, checked at sweep | Archived if `expiry_time ≤ now` | **Implemented** (Step 0 of `run_sweep()`) |
|
|
693
|
+
| Confidence decay | Agent-inferred memory, configurable rate (off by default) | Score reduced per sweep cycle | **Implemented** (off by default — ranking already applies recency decay at read time, making stored decay often redundant) |
|
|
694
|
+
| Staleness archival | Not retrieved within `staleness_days` (default 90) | Archived, excluded from retrieval | **Implemented** |
|
|
695
|
+
| Confidence threshold | Record confidence falls below `archival_threshold` (default 0.10) | Archived | **Implemented** |
|
|
696
|
+
| Skill staleness | Skill health score below 0.6 | Warning surfaced in retrieval; skill not archived automatically | **Implemented** |
|
|
697
|
+
| Reinforcement | Positive outcome signal | Confidence boosted, decay clock reset | **Implemented** via `db.lifecycle.reinforce(record_id)` and `db.lifecycle.record_outcome(...)` |
|
|
698
|
+
| Human revocation | Explicit SDK delete | Archived with audit-log entry; history preserved | **Implemented** via `AgentDB.delete()` |
|
|
699
|
+
|
|
700
|
+
**`human_approved` records are immune to all sweep operations** — no decay, no staleness archival, no TTL expiry, regardless of field values.
|
|
701
|
+
|
|
702
|
+
---
|
|
703
|
+
|
|
704
|
+
## Deployment
|
|
705
|
+
|
|
706
|
+
Three modes, identical interface:
|
|
707
|
+
|
|
708
|
+
```python
|
|
709
|
+
# Embedded — SQLite, zero config, development and prototyping
|
|
710
|
+
from agentdb import AgentDB
|
|
711
|
+
db = AgentDB()
|
|
712
|
+
db = AgentDB(path="./my-agent.db", permissions_path="./permissions.yaml")
|
|
713
|
+
|
|
714
|
+
# Server mode — standalone process, Postgres + pgvector (Phase 5)
|
|
715
|
+
db = AgentDB.connect("mcp://localhost:5433")
|
|
716
|
+
|
|
717
|
+
# Managed cloud — hosted, enterprise SLA (Phase 5+)
|
|
718
|
+
db = AgentDB.connect("agentdb://org.agentdb.io", api_key="...")
|
|
719
|
+
```
|
|
720
|
+
|
|
721
|
+
```typescript
|
|
722
|
+
// Same interface in TypeScript
|
|
723
|
+
import { AgentDB } from 'agentdb'
|
|
724
|
+
const db = new AgentDB()
|
|
725
|
+
const db = AgentDB.connect('mcp://localhost:5433')
|
|
726
|
+
```
|
|
727
|
+
|
|
728
|
+
**MCP server (implemented, Task 14, 2026-07-17):** 4 tools — `memory_write`, `memory_retrieve`, `conflict_queue`, `memory_stats` — via `FastMCP`. Stdio transport for Claude Desktop; `agentdb-mcp` console script. Add to `claude_desktop_config.json`:
|
|
729
|
+
|
|
730
|
+
```json
|
|
731
|
+
{
|
|
732
|
+
"mcpServers": {
|
|
733
|
+
"agentdb": {
|
|
734
|
+
"command": "agentdb-mcp",
|
|
735
|
+
"env": { "AGENTDB_MEMORY_DIR": "/path/to/your/memory/workspace" }
|
|
736
|
+
}
|
|
737
|
+
}
|
|
738
|
+
}
|
|
739
|
+
```
|
|
740
|
+
|
|
741
|
+
HTTP+SSE transport (`agentdb-mcp-http`) and `memory_subscribe` are deferred — push subscriptions don't work over stdio. See `docs/context/capabilities.md`.
|
|
742
|
+
|
|
743
|
+
**Security note:** In embedded (in-process) mode, the permission router runs in the same process as the application. For security-sensitive deployments, server mode is required — the router runs out-of-process and cannot be reached by application code. Embedded mode is for development and prototyping.
|
|
744
|
+
|
|
745
|
+
---
|
|
746
|
+
|
|
747
|
+
## Architecture
|
|
748
|
+
|
|
749
|
+
**This is the target architecture, not a diagram of what's running today.** Implemented today: `AgentDBBackend`, permission router, SQLite store, multi-signal ranker, audit log, conflict detector (`ConflictDetector` — temporal + authority), lifecycle manager (`db.lifecycle`), gateway + proposal workflow (`db.gateway`), MCP server (stdio, `interface/mcp_server.py`). A live in-memory notifier dispatcher (`db.gateway.subscribe`/`dispatch`) exists — narrower and less durable than the "Event bus" box below implies; see `docs/context/capabilities.md`. Not implemented: budget manager as a standalone module, Postgres/pgvector. See `docs/context/database-schema.md` for what's actually persisted.
|
|
750
|
+
|
|
751
|
+
```
|
|
752
|
+
Agent (any framework)
|
|
753
|
+
│
|
|
754
|
+
│ AgentDBBackend / MCP / SDK
|
|
755
|
+
▼
|
|
756
|
+
┌───────────────────────────────────────────┐
|
|
757
|
+
│ AgentDB interface │
|
|
758
|
+
│ AgentDBBackend │ MCP server │ Python SDK │
|
|
759
|
+
└───────────────┬───────────────────────────┘
|
|
760
|
+
│
|
|
761
|
+
┌───────▼────────┐
|
|
762
|
+
│ Middleware │
|
|
763
|
+
│ Origin tag │
|
|
764
|
+
│ Permission │
|
|
765
|
+
│ router │
|
|
766
|
+
│ Conflict │
|
|
767
|
+
│ detector │
|
|
768
|
+
└───────┬────────┘
|
|
769
|
+
│
|
|
770
|
+
┌──────────┼──────────┐
|
|
771
|
+
▼ ▼ ▼
|
|
772
|
+
Human Agent Provider
|
|
773
|
+
zone zone ingest
|
|
774
|
+
(sealed) (governed) (normalised)
|
|
775
|
+
│ │ │
|
|
776
|
+
└──────────┼──────────┘
|
|
777
|
+
│
|
|
778
|
+
┌───────▼────────┐
|
|
779
|
+
│ Retrieval │
|
|
780
|
+
│ Multi-signal │
|
|
781
|
+
│ ranker │
|
|
782
|
+
│ Budget │
|
|
783
|
+
│ manager │
|
|
784
|
+
└───────┬────────┘
|
|
785
|
+
│
|
|
786
|
+
┌──────────┼──────────┐
|
|
787
|
+
▼ ▼ ▼
|
|
788
|
+
SQLite Postgres pgvector
|
|
789
|
+
(embed) (server) (semantic)
|
|
790
|
+
│
|
|
791
|
+
┌───────▼────────┐
|
|
792
|
+
│ Audit log │
|
|
793
|
+
│ (append-only) │
|
|
794
|
+
│ Event bus │
|
|
795
|
+
│ (subscriptions)│
|
|
796
|
+
└────────────────┘
|
|
797
|
+
```
|
|
798
|
+
|
|
799
|
+
---
|
|
800
|
+
|
|
801
|
+
## Memory that improves
|
|
802
|
+
|
|
803
|
+
Most memory systems are static — they store what the agent learned and retrieve it later. AgentDB closes the feedback loop:
|
|
804
|
+
|
|
805
|
+
1. **Every retrieval is logged** with which records were returned and at what scores (retrieval event in the audit trail).
|
|
806
|
+
|
|
807
|
+
2. **Every outcome** — task completion, human correction, human approval — is linked back to the retrieval that preceded it.
|
|
808
|
+
|
|
809
|
+
```python
|
|
810
|
+
# After a task completes:
|
|
811
|
+
db.record_outcome(
|
|
812
|
+
session_id="sess_abc",
|
|
813
|
+
agent_id="assistant",
|
|
814
|
+
outcome_type="task_complete", # or "human_correction" | "human_approval" | "agent_retry"
|
|
815
|
+
outcome_value=1.0, # 1.0 = positive, 0.0 = negative, 0.5 = neutral
|
|
816
|
+
)
|
|
817
|
+
```
|
|
818
|
+
|
|
819
|
+
3. **Memories that consistently precede positive outcomes** have their confidence boosted (up to +0.05 per outcome, capped at +0.1). Memories that precede failures are gently decayed (−0.02 per negative outcome). The agent gets better at knowing what to trust.
|
|
820
|
+
|
|
821
|
+
4. **The corpus quality dashboard** answers "is this agent's memory getting better?" — actual current return shape (see "System tables" above for the full field list; `conflict_pending_count` is hardcoded `0` and `outcome_positive_rate` is `None` under 10 outcomes, since neither conflict tracking nor a small-sample guard-free rate would be trustworthy yet):
|
|
822
|
+
|
|
823
|
+
```python
|
|
824
|
+
db.system.agentdb_stats()
|
|
825
|
+
# {
|
|
826
|
+
# "avg_confidence": 0.81, # trending up = memory improving
|
|
827
|
+
# "low_confidence_count": 3, # records at risk of archival
|
|
828
|
+
# "conflict_pending_count": 0, # hardcoded — no conflict tracking yet
|
|
829
|
+
# "outcome_positive_rate": 0.87, # 87% of tasks completed without correction (once >=10 outcomes recorded)
|
|
830
|
+
# "retrieval_count_total": 412,
|
|
831
|
+
# }
|
|
832
|
+
```
|
|
833
|
+
|
|
834
|
+
**Human-approved memories and skills are immune to outcome-based changes.** The improvement algorithm only operates on agent-inferred memory — the layer the agent itself controls. Humans remain the authority on what is correct. The algorithm makes the agent better at using that authority.
|
|
835
|
+
|
|
836
|
+
Provider-ingested memories can decay from negative outcomes but cannot be boosted — they are external data, not learned knowledge.
|
|
837
|
+
|
|
838
|
+
**Context compaction handling is planned, not yet implemented.** No `checkpoint()` or `write_compaction_summary()` method exists in the codebase today. The intent: promote important working memory to episodic storage before compaction fires, and persist the framework's compressed summary as a durable episodic record, both logged to the audit trail with a compaction boundary marker, with a modest provenance boost for compaction summaries since they're derived from real conversation rather than pure inference.
|
|
839
|
+
|
|
840
|
+
---
|
|
841
|
+
|
|
842
|
+
## What AgentDB is not
|
|
843
|
+
|
|
844
|
+
- **Not a vector database.** pgvector handles semantic retrieval underneath. AgentDB is the governance layer above it.
|
|
845
|
+
- **Not a RAG pipeline.** Existing RAG tooling integrates via adapter.
|
|
846
|
+
- **Not a provider replacement.** Anthropic, OpenAI, Google memory — AgentDB ingests and governs them.
|
|
847
|
+
- **Not a framework.** It owns your data. Leaving requires a migration. That is the moat.
|