mcp-agents-memory 0.9.10 → 0.9.11

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (3) hide show
  1. package/README.md +42 -5
  2. package/build/index.js +245 -44
  3. package/package.json +1 -1
package/README.md CHANGED
@@ -89,6 +89,32 @@
89
89
 
90
90
  ---
91
91
 
92
+ ## 멀티머신 — 서버 / 클라이언트 (콜드패스 처리)
93
+
94
+ 여러 기기가 **하나의 공유 DB**를 쓸 때, Cold Path(태깅·프로필·클러스터링·alias 판정)는 **한 머신에서만** 돌아야 한다 — 안 그러면 같은 row를 여러 기기가 중복 처리하고 클라우드 비용이 배가된다. 같은 패키지를 **config로 역할만** 가른다:
95
+
96
+ | | 클라이언트 | 서버 (처리) |
97
+ |---|---|---|
98
+ | **DB** | 원격 DB 접속 (SSH 터널 등) | DB 호스트 / 직접 접속 |
99
+ | **Cold Path** | `COLD_PATH_ENABLED=false` | 전용 데몬으로 상시 가동 |
100
+ | **하는 일** | `search` / `manage_knowledge`만 | 태깅 · 프로필 · 클러스터링 · alias 판정 |
101
+ | **설정 난이도** | `.env` 몇 줄 (순수 config) | config + 로컬 LLM 인프라 |
102
+
103
+ - **클라이언트**: editor가 띄우는 MCP 서버가 그대로 단말. `.env`에 `COLD_PATH_ENABLED=false`만 추가하면 끝.
104
+ - **서버**: Cold Path를 MCP(=editor) 수명과 분리해 **독립 데몬**으로 상시 가동 (editor를 안 켜도 처리됨):
105
+ ```bash
106
+ mcp-agents-memory coldpath # MCP 서버 없이 Cold Path 워커만 도는 데몬 (systemd 권장)
107
+ ```
108
+ 데몬은 PostgreSQL advisory lock으로 **싱글톤** 보장 — 인스턴스가 몇 개든 락을 잡은 1개만 처리한다(중복 방지·자동 failover).
109
+
110
+ ### Cold Path LLM 백엔드 (config로 교체)
111
+
112
+ `LOCAL_LLM_BASE_URL`로 OpenAI-호환 엔드포인트를 가리키면 로컬/셀프호스트 추론을 쓴다 (llama.cpp, ollama 등). 미설정 시 클라우드(`grok-4-1`) 기본. `LOCAL_GROK_FALLBACK=true`면 로컬 실패 시 grok으로 폴백.
113
+
114
+ > 예) AMD/NVIDIA GPU에 llama.cpp `llama-server`로 Qwen3-14B를 올리고 `LOCAL_LLM_BASE_URL=http://localhost:8080/v1` → 콜드패스 클라우드 비용 ≈ $0. (json_schema 문법 + thinking off로 valid JSON 보장)
115
+
116
+ ---
117
+
92
118
  ## 메모리 로드 룰
93
119
 
94
120
  - **단기 메모리**: 최근 2-3일 raw 그대로, 또는 8000 토큰(약 12000-16000자) 중 먼저 도달하는 것
@@ -140,10 +166,10 @@ manage_knowledge({
140
166
  | 역할 | 사용 기술 |
141
167
  |---|---|
142
168
  | **Embedding** | OpenAI `text-embedding-3-large` (3072 dim) |
143
- | **Tagger (Cold Path)** | Google `gemini-2.5-flash` (predefined + dynamic) |
169
+ | **Cold Path LLM** (tagger / librarian / clusterer / project-alias judge) | 로컬 `Qwen3-14B` (llama.cpp, json_schema 문법 + thinking off → valid JSON 보장) **또는** 클라우드 `grok-4-1-fast-non-reasoning` — `LOCAL_LLM_BASE_URL`로 선택 |
144
170
  | **검색 fallback** | PostgreSQL `ILIKE` (cosine 임계값 미만 시) |
145
171
  | **DB** | PostgreSQL + pgvector |
146
- | **Librarian (memory → user)** | local `qwen3.6:35b-a3b` (Q4_K_M, ollama) — 30 메시지 + 24h 게이트 |
172
+ | **Librarian (memory → user)** | 위 Cold Path 백엔드 공유 — recency-bias 저항 큐레이션(core 정체성 ↔ sub 작업 분리 + null-preserve), 게이트 env tunable |
147
173
  | **Skill 시스템** | TBD (다음 라운드) |
148
174
 
149
175
  ---
@@ -164,11 +190,22 @@ SSH_HOST=...
164
190
 
165
191
  # 모델
166
192
  EMBEDDING_MODEL=text-embedding-3-large
167
- TAGGER_MODEL=gemini-2.5-flash
168
- OPENAI_API_KEY=...
169
- GEMINI_API_KEY=...
193
+ OPENAI_API_KEY=... # embedding (필수)
194
+ XAI_API_KEY=... # grok-4-1 (Cold Path 기본 + 로컬 폴백)
195
+
196
+ # Cold Path LLM 백엔드 — 로컬 추론 쓰려면 OpenAI-호환 엔드포인트 지정 (없으면 클라우드)
197
+ LOCAL_LLM_BASE_URL=http://localhost:8080/v1 # llama.cpp / ollama 등
198
+ LOCAL_GROK_FALLBACK=true # 로컬 실패 시 grok 폴백
199
+ TAGGER_PROVIDER=local # local / xai
200
+ TAGGER_MODEL=qwen3-14b
201
+ LIBRARIAN_PROVIDER=local
202
+ LIBRARIAN_MODEL=qwen3-14b
203
+ LIBRARIAN_ENABLED=true
204
+ LIBRARIAN_MSG_THRESHOLD=30 # 라이브러리언 게이트 (기본 보수적)
205
+ LIBRARIAN_COOLDOWN_HOURS=24
170
206
 
171
207
  # Hot/Cold path 제어
208
+ COLD_PATH_ENABLED=true # false = 단말(Cold Path 안 돎). 멀티머신에선 처리 서버만 true
172
209
  COLD_PATH_INTERVAL_SEC=60 # 1분 단위 스케줄
173
210
  COLD_PATH_BATCH_SIZE=5 # 또는 5메시지 단위
174
211
 
package/build/index.js CHANGED
@@ -55458,16 +55458,38 @@ async function callSpec(spec, opts) {
55458
55458
  }
55459
55459
  case "local": {
55460
55460
  const client2 = getLocalClient();
55461
- const res = await client2.chat.completions.create({
55461
+ let effectiveThinking = opts.enableThinking;
55462
+ if (opts.enableThinking === true && opts.jsonSchema) {
55463
+ console.warn(
55464
+ "[ModelRegistry] enableThinking + jsonSchema both set \u2014 forcing thinking OFF (llama.cpp #20345 silently bypasses grammar when thinking is on)."
55465
+ );
55466
+ effectiveThinking = false;
55467
+ }
55468
+ const responseFormat = opts.jsonSchema ? {
55469
+ type: "json_schema",
55470
+ json_schema: {
55471
+ name: "librarian_profile",
55472
+ strict: true,
55473
+ schema: opts.jsonSchema
55474
+ }
55475
+ } : useJson ? { type: "json_object" } : void 0;
55476
+ const params = {
55462
55477
  model: spec.model_name,
55463
55478
  messages: [
55464
55479
  { role: "system", content: opts.system },
55465
55480
  { role: "user", content: opts.user }
55466
55481
  ],
55467
- ...useJson ? { response_format: { type: "json_object" } } : {},
55482
+ ...responseFormat ? { response_format: responseFormat } : {},
55468
55483
  temperature: 0.1,
55469
55484
  max_tokens: maxTokens
55470
- });
55485
+ };
55486
+ const res = await client2.chat.completions.create(
55487
+ // narrow cast: only this call site is widened to accommodate chat_template_kwargs
55488
+ {
55489
+ ...params,
55490
+ ...effectiveThinking !== void 0 ? { chat_template_kwargs: { enable_thinking: effectiveThinking } } : {}
55491
+ }
55492
+ );
55471
55493
  const raw = res.choices[0]?.message?.content || "";
55472
55494
  return raw.replace(/<think>[\s\S]*?<\/think>/g, "").replace(/```json|```/g, "").trim();
55473
55495
  }
@@ -55539,6 +55561,15 @@ async function listProjectTagCandidates(limit2 = 20) {
55539
55561
  function invalidateCandidateCache() {
55540
55562
  _candidateCache = null;
55541
55563
  }
55564
+ var TAGGER_SCHEMA = {
55565
+ type: "object",
55566
+ properties: {
55567
+ p_tag: { type: ["string", "null"] },
55568
+ d_tag: { type: "array", items: { type: "string" }, maxItems: 5 }
55569
+ },
55570
+ required: ["p_tag", "d_tag"],
55571
+ additionalProperties: false
55572
+ };
55542
55573
  var SYSTEM_PROMPT = `Tagger for one user's personal long-term memory across AI agents.
55543
55574
 
55544
55575
  OUTPUT (strict JSON):
@@ -55598,9 +55629,11 @@ async function tagMessage(input) {
55598
55629
  raw = await callRole("tagger", {
55599
55630
  system: SYSTEM_PROMPT,
55600
55631
  user: userPrompt,
55601
- responseFormat: "json"
55602
- // thinking: false — 태거는 단순 매핑 작업. thinking 켜면 reasoning이 모든 토큰 소비해 content 비어버림.
55603
- // Librarian 등 복잡한 분석 역할에서만 thinking: true 사용.
55632
+ responseFormat: "json",
55633
+ jsonSchema: TAGGER_SCHEMA,
55634
+ enableThinking: false
55635
+ // thinking off — 태거는 단순 매핑 작업. thinking 켜면 reasoning이 모든 토큰 소비해 content 비어버림.
55636
+ // jsonSchema → llama.cpp grammar로 <think> bleed 차단 (Qwen3 bug #20345).
55604
55637
  });
55605
55638
  } catch (err2) {
55606
55639
  if (!localFallbackEnabled) throw err2;
@@ -56218,6 +56251,20 @@ var SOURCE_WEIGHT = {
56218
56251
  name_similarity: 16,
56219
56252
  low_frequency_tag: 8
56220
56253
  };
56254
+ var JUDGE_SCHEMA = {
56255
+ type: "object",
56256
+ properties: {
56257
+ relation: { type: "string", enum: ["rename", "alias", "same_project", "different", "misfile_suspected", "insufficient"] },
56258
+ same_project: { type: "boolean" },
56259
+ source_should_alias_target: { type: "boolean" },
56260
+ confidence: { type: "number" },
56261
+ evidence_memory_ids: { type: "array", items: { type: "integer" } },
56262
+ conflict_memory_ids: { type: "array", items: { type: "integer" } },
56263
+ rationale: { type: "string" }
56264
+ },
56265
+ required: ["relation", "same_project", "source_should_alias_target", "confidence", "evidence_memory_ids", "conflict_memory_ids", "rationale"],
56266
+ additionalProperties: false
56267
+ };
56221
56268
  var JUDGE_SYSTEM_PROMPT = `You are a project-tag alias judge for one user's personal memory system.
56222
56269
 
56223
56270
  Decide whether two CANONICAL project tags refer to the same project identity.
@@ -56771,7 +56818,7 @@ async function judgePair(pair, tagsById, memoriesByTag) {
56771
56818
  const raw = await callSpec(spec, {
56772
56819
  system: JUDGE_SYSTEM_PROMPT,
56773
56820
  user: userPrompt,
56774
- ...spec.provider === "local" ? {} : { responseFormat: "json" },
56821
+ ...spec.provider === "local" ? { jsonSchema: JUDGE_SCHEMA, enableThinking: false } : { responseFormat: "json" },
56775
56822
  maxTokens: envInt("PROJECT_ALIAS_JUDGE_MAX_TOKENS", 8192)
56776
56823
  });
56777
56824
  if (!raw) throw new Error("project_alias_judge returned empty content");
@@ -58900,6 +58947,7 @@ ${tagList}`;
58900
58947
  system: CLUSTER_SYSTEM,
58901
58948
  user: userPrompt,
58902
58949
  responseFormat: "json",
58950
+ enableThinking: false,
58903
58951
  maxTokens: 512
58904
58952
  });
58905
58953
  } catch (err2) {
@@ -58993,35 +59041,63 @@ async function runDtagPromotion() {
58993
59041
 
58994
59042
  // src/librarian.ts
58995
59043
  init_db();
58996
- var LIBRARIAN_MSG_THRESHOLD = 30;
58997
- var LIBRARIAN_COOLDOWN_MS = 24 * 60 * 60 * 1e3;
58998
- var LIBRARIAN_RECENT_LIMIT = 50;
58999
- var LIBRARIAN_MAX_TOKENS = 32768;
59044
+ var LIBRARIAN_MSG_THRESHOLD = Number(process.env.LIBRARIAN_MSG_THRESHOLD ?? 30);
59045
+ var LIBRARIAN_COOLDOWN_MS = Number(process.env.LIBRARIAN_COOLDOWN_HOURS ?? 24) * 60 * 60 * 1e3;
59046
+ var LIBRARIAN_RECENT_SLICE = Number(process.env.LIBRARIAN_RECENT_SLICE ?? 25);
59047
+ var LIBRARIAN_HISTORY_SLICE = Number(process.env.LIBRARIAN_HISTORY_SLICE ?? 25);
59048
+ var LIBRARIAN_MAX_TOKENS = Number(process.env.LIBRARIAN_MAX_TOKENS ?? 2048);
59049
+ var LIBRARIAN_PROFILE_SCHEMA = {
59050
+ type: "object",
59051
+ properties: {
59052
+ core_profile: { type: ["string", "null"] },
59053
+ sub_profile: { type: ["string", "null"] }
59054
+ },
59055
+ required: ["core_profile", "sub_profile"],
59056
+ additionalProperties: false
59057
+ };
59000
59058
  var SYSTEM_PROMPT2 = `You are the Librarian for one user's personal memory system.
59001
59059
 
59002
59060
  YOUR JOB
59003
- Look at the user's recent first-person messages (role='user' only \u2014 ignore
59004
- assistant replies). Identify any STABLE, MEMORABLE facts about WHO THE USER IS
59005
- or HOW THEY WORK that should be promoted to their long-term profile.
59006
-
59007
- OUTPUT TWO SECTIONS:
59008
- 1. core_profile: critically important user identity (name, role, expertise,
59009
- strong preferences). Should be SHORT (5-10 lines max) and high-signal.
59010
- 2. sub_profile: secondary memorable info (tools, environment quirks,
59011
- ongoing project focus, working style preferences). Can be longer but
59012
- still curated.
59013
-
59014
- RULES
59015
- - The user is one person. Output is the WHOLE profile (replacing existing,
59016
- not appending) \u2014 so include relevant existing facts that are still true.
59061
+ Look at the user's first-person messages (role='user' only \u2014 ignore assistant
59062
+ replies). Identify any STABLE, DURABLE facts about WHO THE USER IS that should
59063
+ be promoted to their long-term profile.
59064
+
59065
+ OUTPUT TWO SECTIONS \u2014 these are STRICTLY SEPARATE categories:
59066
+
59067
+ 1. core_profile \u2014 DURABLE IDENTITY ONLY.
59068
+ Who the person IS: name, role, profession, expertise, stable long-term
59069
+ preferences. Should be SHORT (5-10 lines max) and high-signal.
59070
+ Example of the SHAPE only (fictional \u2014 never copy this content): "Backend
59071
+ engineer at a logistics startup; 10+ yrs Python; prefers terse,
59072
+ example-driven answers." Derive the actual content ONLY from the messages.
59073
+
59074
+ 2. sub_profile \u2014 CURRENT WORK AND ACTIVITY.
59075
+ What they are actively doing, building, or focused on: tools, environment
59076
+ details, ongoing projects, recent working style observations. Expected to
59077
+ change often. Can be longer but still curated.
59078
+
59079
+ CRITICAL IDENTITY vs. WORK DISTINCTION
59080
+ A user discussing, building, evaluating, debugging, or working on a topic \u2014
59081
+ including AI models, agent frameworks, or this memory system itself \u2014 is
59082
+ describing their WORK or CURRENT ACTIVITY, NOT their identity. A burst of
59083
+ messages about one subject means they are WORKING on it, not that it defines
59084
+ them. NEVER promote a work or project topic into core_profile. It belongs in
59085
+ sub_profile at most.
59086
+
59087
+ CONSERVATISM / NULL-PRESERVE RULE (most important rule)
59088
+ If the recent window contains NO new durable identity fact \u2014 only project work,
59089
+ meta-tooling, topic evaluation, or session-specific activity \u2014 return
59090
+ core_profile: null to PRESERVE the existing identity unchanged. DO NOT restate,
59091
+ rephrase, or "refresh" an existing core_profile just because you saw it. Null
59092
+ means "keep it as-is." Only set a non-null core_profile when there is an
59093
+ explicit, durable, first-person identity statement that is genuinely new.
59094
+ sub_profile may freely capture current projects and activity.
59095
+
59096
+ OTHER RULES
59017
59097
  - DO NOT invent facts not supported by the messages.
59018
- - DO NOT promote temporary state ("debugging X", "frustrated with Y").
59019
- Only stable identity / preferences.
59020
- - DO NOT promote third-party advice or system hints that show up in
59021
- messages \u2014 only what the user is saying ABOUT THEMSELVES.
59098
+ - DO NOT promote third-party advice or system hints \u2014 only what the user is
59099
+ saying ABOUT THEMSELVES.
59022
59100
  - Korean is fine. Match the language of the user's writing.
59023
- - If recent messages don't add anything new and existing profile is fine,
59024
- output the existing profile unchanged.
59025
59101
 
59026
59102
  FORMAT RULES FOR THE VALUES:
59027
59103
  - Both fields must be PLAIN PROSE TEXT \u2014 no nested JSON, no {}, [], key-value blobs.
@@ -59059,6 +59135,32 @@ async function checkGate() {
59059
59135
  const shouldRun = enoughNewMessages && (neverRan || cooldownPassed);
59060
59136
  return { userId, shouldRun, currentMsgCount, lastRunAt };
59061
59137
  }
59138
+ async function runLibrarianDeepPass(userPrompt, maxTokens) {
59139
+ const analysisSystemPrompt = `${SYSTEM_PROMPT2}
59140
+
59141
+ In this ANALYSIS PASS, think deeply and reason freely about what should go in each profile field.
59142
+ Do NOT output JSON yet \u2014 write your analysis as plain prose reasoning.`;
59143
+ const analysis = await callRole("librarian", {
59144
+ system: analysisSystemPrompt,
59145
+ user: userPrompt,
59146
+ maxTokens,
59147
+ enableThinking: true
59148
+ // jsonSchema 未指定 → thinking ON が有効になる
59149
+ });
59150
+ const pass2User = `${userPrompt}
59151
+
59152
+ ANALYSIS FROM PREVIOUS REASONING PASS (use as additional context):
59153
+ ${analysis}
59154
+
59155
+ Now emit the final JSON profile only.`;
59156
+ return callRole("librarian", {
59157
+ system: SYSTEM_PROMPT2,
59158
+ user: pass2User,
59159
+ maxTokens,
59160
+ jsonSchema: LIBRARIAN_PROFILE_SCHEMA,
59161
+ enableThinking: false
59162
+ });
59163
+ }
59062
59164
  async function runLibrarian() {
59063
59165
  if (librarianRunning) return;
59064
59166
  const gate = await checkGate();
@@ -59082,27 +59184,55 @@ async function runLibrarian() {
59082
59184
  AND is_active = TRUE
59083
59185
  ORDER BY created_at DESC
59084
59186
  LIMIT $2`,
59085
- [gate.userId, LIBRARIAN_RECENT_LIMIT]
59187
+ [gate.userId, LIBRARIAN_RECENT_SLICE]
59086
59188
  );
59087
59189
  if (recentR.rows.length === 0) return;
59088
- const messagesText = recentR.rows.reverse().map(
59089
- (r, i) => `[#${i + 1} @ ${r.created_at?.toISOString().slice(0, 19) ?? ""}] ${r.message}`
59190
+ const historyR = await db.query(
59191
+ `SELECT message, created_at
59192
+ FROM memory
59193
+ WHERE user_id = $1
59194
+ AND role = 'user'
59195
+ AND is_active = TRUE
59196
+ ORDER BY created_at ASC
59197
+ LIMIT $2`,
59198
+ [gate.userId, LIBRARIAN_HISTORY_SLICE]
59199
+ );
59200
+ const recentKeys = new Set(
59201
+ recentR.rows.map(
59202
+ (r) => `${r.created_at?.toISOString() ?? ""}|${r.message}`
59203
+ )
59204
+ );
59205
+ const historicalRows = historyR.rows.filter(
59206
+ (r) => !recentKeys.has(`${r.created_at?.toISOString() ?? ""}|${r.message}`)
59207
+ );
59208
+ const recentSorted = [...recentR.rows].reverse();
59209
+ const historicalSorted = [...historicalRows];
59210
+ const formatRows = (rows, startIdx) => rows.map(
59211
+ (r, i) => `[#${startIdx + i + 1} @ ${r.created_at?.toISOString().slice(0, 19) ?? ""}] ${r.message}`
59090
59212
  ).join("\n\n");
59091
- const userPrompt = `EXISTING PROFILE (subject to update):
59213
+ const historicalText = historicalSorted.length > 0 ? `
59214
+ EARLIER MESSAGES \u2014 historical context (identity anchors, NOT current activity):
59215
+ ${formatRows(historicalSorted, 0)}
59216
+ ` : "";
59217
+ const recentText = `
59218
+ RECENT MESSAGES \u2014 most recent ${recentSorted.length} messages:
59219
+ ${formatRows(recentSorted, historicalSorted.length)}
59220
+ `;
59221
+ const userPrompt = `EXISTING PROFILE (preserve unless a new durable identity fact appears):
59092
59222
  core_profile:
59093
59223
  ${before.core_profile ?? "(empty)"}
59094
59224
 
59095
59225
  sub_profile:
59096
59226
  ${before.sub_profile ?? "(empty)"}
59097
-
59098
- RECENT USER MESSAGES (most recent ${recentR.rows.length}, role='user'):
59099
- ${messagesText}
59100
-
59227
+ ${historicalText}${recentText}
59101
59228
  Task: produce updated core_profile and sub_profile JSON per the system prompt.`;
59102
- const raw = await callRole("librarian", {
59229
+ const deepMode = process.env.LIBRARIAN_DEEP_THINKING === "true";
59230
+ const raw = deepMode ? await runLibrarianDeepPass(userPrompt, LIBRARIAN_MAX_TOKENS) : await callRole("librarian", {
59103
59231
  system: SYSTEM_PROMPT2,
59104
59232
  user: userPrompt,
59105
- maxTokens: LIBRARIAN_MAX_TOKENS
59233
+ maxTokens: LIBRARIAN_MAX_TOKENS,
59234
+ jsonSchema: LIBRARIAN_PROFILE_SCHEMA,
59235
+ enableThinking: false
59106
59236
  });
59107
59237
  if (!raw) {
59108
59238
  throw new Error("Librarian returned empty content (reasoning token budget exceeded?)");
@@ -59138,8 +59268,9 @@ Task: produce updated core_profile and sub_profile JSON per the system prompt.`;
59138
59268
  `UPDATE users SET librarian_msg_count_at_run = $1 WHERE user_id = $2`,
59139
59269
  [gate.currentMsgCount, gate.userId]
59140
59270
  );
59271
+ const totalMsgs = recentSorted.length + historicalSorted.length;
59141
59272
  console.error(
59142
- `\u{1F4DA} [Librarian] done \u2014 ${recentR.rows.length} msgs, profile ${changed ? "updated" : "unchanged"}`
59273
+ `\u{1F4DA} [Librarian] done \u2014 ${totalMsgs} msgs (${recentSorted.length} recent + ${historicalSorted.length} historical), profile ${changed ? "updated" : "unchanged"}`
59143
59274
  );
59144
59275
  } catch (err2) {
59145
59276
  console.error("\u26A0\uFE0F [Librarian] run failed (retries in 24h):", err2);
@@ -59330,11 +59461,56 @@ async function maybeRunLibrarian() {
59330
59461
  console.error("\u26A0\uFE0F [Librarian] unhandled error (non-blocking):", err2);
59331
59462
  }
59332
59463
  }
59333
- function startColdPathWorker() {
59464
+ var COLDPATH_ADVISORY_LOCK_KEY = 4242000017;
59465
+ var lockClient = null;
59466
+ async function acquireColdPathLock() {
59467
+ if (lockClient) return true;
59468
+ const client2 = await db.getClient();
59469
+ try {
59470
+ const res = await client2.query("SELECT pg_try_advisory_lock($1) AS got", [COLDPATH_ADVISORY_LOCK_KEY]);
59471
+ if (res.rows[0]?.got === true) {
59472
+ lockClient = client2;
59473
+ return true;
59474
+ }
59475
+ client2.release();
59476
+ return false;
59477
+ } catch (err2) {
59478
+ try {
59479
+ client2.release();
59480
+ } catch {
59481
+ }
59482
+ throw err2;
59483
+ }
59484
+ }
59485
+ async function releaseColdPathLock() {
59486
+ const c = lockClient;
59487
+ if (!c) return;
59488
+ lockClient = null;
59489
+ try {
59490
+ await c.query("SELECT pg_advisory_unlock($1)", [COLDPATH_ADVISORY_LOCK_KEY]);
59491
+ } catch {
59492
+ }
59493
+ try {
59494
+ c.release();
59495
+ } catch {
59496
+ }
59497
+ }
59498
+ async function startColdPathWorker() {
59334
59499
  if (process.env.COLD_PATH_ENABLED === "false") {
59335
59500
  console.error("\u{1F535} [ColdPath] disabled (COLD_PATH_ENABLED=false)");
59336
59501
  return;
59337
59502
  }
59503
+ let got = false;
59504
+ try {
59505
+ got = await acquireColdPathLock();
59506
+ } catch (err2) {
59507
+ console.error("\u274C [ColdPath] advisory lock acquire failed \u2014 not starting:", err2);
59508
+ return;
59509
+ }
59510
+ if (!got) {
59511
+ console.error("\u{1F512} [ColdPath] another instance holds the lock \u2014 this process will not process");
59512
+ return;
59513
+ }
59338
59514
  const intervalSec = envInt3("COLD_PATH_INTERVAL_SEC", DEFAULT_INTERVAL_SEC);
59339
59515
  const warmupSec = envInt3("COLD_PATH_WARMUP_SEC", DEFAULT_WARMUP_SEC);
59340
59516
  console.error(`\u{1F535} [ColdPath] starting \u2014 warmup ${warmupSec}s, interval ${intervalSec}s`);
@@ -59356,6 +59532,7 @@ function stopColdPathWorker() {
59356
59532
  if (intervalTimer) clearInterval(intervalTimer);
59357
59533
  warmupTimer = null;
59358
59534
  intervalTimer = null;
59535
+ void releaseColdPathLock();
59359
59536
  console.error("\u{1F535} [ColdPath] stopped");
59360
59537
  }
59361
59538
  async function drainColdPath(maxTicks = 12) {
@@ -60187,6 +60364,7 @@ function printHelp() {
60187
60364
 
60188
60365
  Usage:
60189
60366
  mcp-agents-memory Run the MCP server (stdio).
60367
+ mcp-agents-memory coldpath Run ONLY the cold-path worker as a standalone always-on daemon (no MCP server). For the processing/GPU machine via systemd.
60190
60368
  mcp-agents-memory setup Interactive setup \u2014 write config to ~/.config/mcp-agents-memory/.env and run migrations.
60191
60369
  mcp-agents-memory migrate Apply any pending DB migrations against the configured database.
60192
60370
  mcp-agents-memory help Show this message.
@@ -60236,13 +60414,36 @@ async function runMcpServer() {
60236
60414
  console.error("\u274C Fatal error during startup:", err2);
60237
60415
  process.exit(1);
60238
60416
  }
60239
- startColdPathWorker();
60417
+ void startColdPathWorker().catch((err2) => console.error("\u274C [ColdPath] start failed:", err2));
60418
+ }
60419
+ async function runColdPathDaemon() {
60420
+ process.env.COLD_PATH_ENABLED = "true";
60421
+ process.on("SIGTERM", () => {
60422
+ void shutdown("SIGTERM");
60423
+ });
60424
+ process.on("SIGINT", () => {
60425
+ void shutdown("SIGINT");
60426
+ });
60427
+ process.on("SIGHUP", () => {
60428
+ void shutdown("SIGHUP");
60429
+ });
60430
+ try {
60431
+ await db.connect();
60432
+ } catch (err2) {
60433
+ console.error("\u274C [ColdPathDaemon] DB connect failed \u2014 exiting (systemd will retry):", err2);
60434
+ process.exit(1);
60435
+ }
60436
+ console.error(`\u{1F9CA} Cold-path daemon (v${PACKAGE_VERSION2}) started \u2014 no MCP server`);
60437
+ await startColdPathWorker();
60240
60438
  }
60241
60439
  async function cli() {
60242
60440
  const cmd = process.argv[2];
60243
60441
  if (!cmd || cmd === "serve") {
60244
60442
  return runMcpServer();
60245
60443
  }
60444
+ if (cmd === "coldpath") {
60445
+ return runColdPathDaemon();
60446
+ }
60246
60447
  if (cmd === "help" || cmd === "--help" || cmd === "-h") {
60247
60448
  printHelp();
60248
60449
  return;
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "mcp-agents-memory",
3
- "version": "0.9.10",
3
+ "version": "0.9.11",
4
4
  "description": "Long-term memory for AI agents with provenance tracking",
5
5
  "main": "build/index.js",
6
6
  "type": "module",