mcp-agents-memory 0.9.10 → 0.9.12

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (3) hide show
  1. package/README.md +42 -5
  2. package/build/index.js +274 -53
  3. package/package.json +1 -1
package/README.md CHANGED
@@ -89,6 +89,32 @@
89
89
 
90
90
  ---
91
91
 
92
+ ## 멀티머신 — 서버 / 클라이언트 (콜드패스 처리)
93
+
94
+ 여러 기기가 **하나의 공유 DB**를 쓸 때, Cold Path(태깅·프로필·클러스터링·alias 판정)는 **한 머신에서만** 돌아야 한다 — 안 그러면 같은 row를 여러 기기가 중복 처리하고 클라우드 비용이 배가된다. 같은 패키지를 **config로 역할만** 가른다:
95
+
96
+ | | 클라이언트 | 서버 (처리) |
97
+ |---|---|---|
98
+ | **DB** | 원격 DB 접속 (SSH 터널 등) | DB 호스트 / 직접 접속 |
99
+ | **Cold Path** | `COLD_PATH_ENABLED=false` | 전용 데몬으로 상시 가동 |
100
+ | **하는 일** | `search` / `manage_knowledge`만 | 태깅 · 프로필 · 클러스터링 · alias 판정 |
101
+ | **설정 난이도** | `.env` 몇 줄 (순수 config) | config + 로컬 LLM 인프라 |
102
+
103
+ - **클라이언트**: editor가 띄우는 MCP 서버가 그대로 단말. `.env`에 `COLD_PATH_ENABLED=false`만 추가하면 끝.
104
+ - **서버**: Cold Path를 MCP(=editor) 수명과 분리해 **독립 데몬**으로 상시 가동 (editor를 안 켜도 처리됨):
105
+ ```bash
106
+ mcp-agents-memory coldpath # MCP 서버 없이 Cold Path 워커만 도는 데몬 (systemd 권장)
107
+ ```
108
+ 데몬은 PostgreSQL advisory lock으로 **싱글톤** 보장 — 인스턴스가 몇 개든 락을 잡은 1개만 처리한다(중복 방지·자동 failover).
109
+
110
+ ### Cold Path LLM 백엔드 (config로 교체)
111
+
112
+ `LOCAL_LLM_BASE_URL`로 OpenAI-호환 엔드포인트를 가리키면 로컬/셀프호스트 추론을 쓴다 (llama.cpp, ollama 등). 미설정 시 클라우드(`grok-4-1`) 기본. `LOCAL_GROK_FALLBACK=true`면 로컬 실패 시 grok으로 폴백.
113
+
114
+ > 예) AMD/NVIDIA GPU에 llama.cpp `llama-server`로 Qwen3-14B를 올리고 `LOCAL_LLM_BASE_URL=http://localhost:8080/v1` → 콜드패스 클라우드 비용 ≈ $0. (json_schema 문법 + thinking off로 valid JSON 보장)
115
+
116
+ ---
117
+
92
118
  ## 메모리 로드 룰
93
119
 
94
120
  - **단기 메모리**: 최근 2-3일 raw 그대로, 또는 8000 토큰(약 12000-16000자) 중 먼저 도달하는 것
@@ -140,10 +166,10 @@ manage_knowledge({
140
166
  | 역할 | 사용 기술 |
141
167
  |---|---|
142
168
  | **Embedding** | OpenAI `text-embedding-3-large` (3072 dim) |
143
- | **Tagger (Cold Path)** | Google `gemini-2.5-flash` (predefined + dynamic) |
169
+ | **Cold Path LLM** (tagger / librarian / clusterer / project-alias judge) | 로컬 `Qwen3-14B` (llama.cpp, json_schema 문법 + thinking off → valid JSON 보장) **또는** 클라우드 `grok-4-1-fast-non-reasoning` — `LOCAL_LLM_BASE_URL`로 선택 |
144
170
  | **검색 fallback** | PostgreSQL `ILIKE` (cosine 임계값 미만 시) |
145
171
  | **DB** | PostgreSQL + pgvector |
146
- | **Librarian (memory → user)** | local `qwen3.6:35b-a3b` (Q4_K_M, ollama) — 30 메시지 + 24h 게이트 |
172
+ | **Librarian (memory → user)** | 위 Cold Path 백엔드 공유 — recency-bias 저항 큐레이션(core 정체성 ↔ sub 작업 분리 + null-preserve), 게이트 env tunable |
147
173
  | **Skill 시스템** | TBD (다음 라운드) |
148
174
 
149
175
  ---
@@ -164,11 +190,22 @@ SSH_HOST=...
164
190
 
165
191
  # 모델
166
192
  EMBEDDING_MODEL=text-embedding-3-large
167
- TAGGER_MODEL=gemini-2.5-flash
168
- OPENAI_API_KEY=...
169
- GEMINI_API_KEY=...
193
+ OPENAI_API_KEY=... # embedding (필수)
194
+ XAI_API_KEY=... # grok-4-1 (Cold Path 기본 + 로컬 폴백)
195
+
196
+ # Cold Path LLM 백엔드 — 로컬 추론 쓰려면 OpenAI-호환 엔드포인트 지정 (없으면 클라우드)
197
+ LOCAL_LLM_BASE_URL=http://localhost:8080/v1 # llama.cpp / ollama 등
198
+ LOCAL_GROK_FALLBACK=true # 로컬 실패 시 grok 폴백
199
+ TAGGER_PROVIDER=local # local / xai
200
+ TAGGER_MODEL=qwen3-14b
201
+ LIBRARIAN_PROVIDER=local
202
+ LIBRARIAN_MODEL=qwen3-14b
203
+ LIBRARIAN_ENABLED=true
204
+ LIBRARIAN_MSG_THRESHOLD=30 # 라이브러리언 게이트 (기본 보수적)
205
+ LIBRARIAN_COOLDOWN_HOURS=24
170
206
 
171
207
  # Hot/Cold path 제어
208
+ COLD_PATH_ENABLED=true # false = 단말(Cold Path 안 돎). 멀티머신에선 처리 서버만 true
172
209
  COLD_PATH_INTERVAL_SEC=60 # 1분 단위 스케줄
173
210
  COLD_PATH_BATCH_SIZE=5 # 또는 5메시지 단위
174
211
 
package/build/index.js CHANGED
@@ -55458,16 +55458,38 @@ async function callSpec(spec, opts) {
55458
55458
  }
55459
55459
  case "local": {
55460
55460
  const client2 = getLocalClient();
55461
- const res = await client2.chat.completions.create({
55461
+ let effectiveThinking = opts.enableThinking;
55462
+ if (opts.enableThinking === true && opts.jsonSchema) {
55463
+ console.warn(
55464
+ "[ModelRegistry] enableThinking + jsonSchema both set \u2014 forcing thinking OFF (llama.cpp #20345 silently bypasses grammar when thinking is on)."
55465
+ );
55466
+ effectiveThinking = false;
55467
+ }
55468
+ const responseFormat = opts.jsonSchema ? {
55469
+ type: "json_schema",
55470
+ json_schema: {
55471
+ name: "librarian_profile",
55472
+ strict: true,
55473
+ schema: opts.jsonSchema
55474
+ }
55475
+ } : useJson ? { type: "json_object" } : void 0;
55476
+ const params = {
55462
55477
  model: spec.model_name,
55463
55478
  messages: [
55464
55479
  { role: "system", content: opts.system },
55465
55480
  { role: "user", content: opts.user }
55466
55481
  ],
55467
- ...useJson ? { response_format: { type: "json_object" } } : {},
55482
+ ...responseFormat ? { response_format: responseFormat } : {},
55468
55483
  temperature: 0.1,
55469
55484
  max_tokens: maxTokens
55470
- });
55485
+ };
55486
+ const res = await client2.chat.completions.create(
55487
+ // narrow cast: only this call site is widened to accommodate chat_template_kwargs
55488
+ {
55489
+ ...params,
55490
+ ...effectiveThinking !== void 0 ? { chat_template_kwargs: { enable_thinking: effectiveThinking } } : {}
55491
+ }
55492
+ );
55471
55493
  const raw = res.choices[0]?.message?.content || "";
55472
55494
  return raw.replace(/<think>[\s\S]*?<\/think>/g, "").replace(/```json|```/g, "").trim();
55473
55495
  }
@@ -55539,6 +55561,15 @@ async function listProjectTagCandidates(limit2 = 20) {
55539
55561
  function invalidateCandidateCache() {
55540
55562
  _candidateCache = null;
55541
55563
  }
55564
+ var TAGGER_SCHEMA = {
55565
+ type: "object",
55566
+ properties: {
55567
+ p_tag: { type: ["string", "null"] },
55568
+ d_tag: { type: "array", items: { type: "string" }, maxItems: 5 }
55569
+ },
55570
+ required: ["p_tag", "d_tag"],
55571
+ additionalProperties: false
55572
+ };
55542
55573
  var SYSTEM_PROMPT = `Tagger for one user's personal long-term memory across AI agents.
55543
55574
 
55544
55575
  OUTPUT (strict JSON):
@@ -55598,9 +55629,11 @@ async function tagMessage(input) {
55598
55629
  raw = await callRole("tagger", {
55599
55630
  system: SYSTEM_PROMPT,
55600
55631
  user: userPrompt,
55601
- responseFormat: "json"
55602
- // thinking: false — 태거는 단순 매핑 작업. thinking 켜면 reasoning이 모든 토큰 소비해 content 비어버림.
55603
- // Librarian 등 복잡한 분석 역할에서만 thinking: true 사용.
55632
+ responseFormat: "json",
55633
+ jsonSchema: TAGGER_SCHEMA,
55634
+ enableThinking: false
55635
+ // thinking off — 태거는 단순 매핑 작업. thinking 켜면 reasoning이 모든 토큰 소비해 content 비어버림.
55636
+ // jsonSchema → llama.cpp grammar로 <think> bleed 차단 (Qwen3 bug #20345).
55604
55637
  });
55605
55638
  } catch (err2) {
55606
55639
  if (!localFallbackEnabled) throw err2;
@@ -56218,6 +56251,20 @@ var SOURCE_WEIGHT = {
56218
56251
  name_similarity: 16,
56219
56252
  low_frequency_tag: 8
56220
56253
  };
56254
+ var JUDGE_SCHEMA = {
56255
+ type: "object",
56256
+ properties: {
56257
+ relation: { type: "string", enum: ["rename", "alias", "same_project", "different", "misfile_suspected", "insufficient"] },
56258
+ same_project: { type: "boolean" },
56259
+ source_should_alias_target: { type: "boolean" },
56260
+ confidence: { type: "number" },
56261
+ evidence_memory_ids: { type: "array", items: { type: "integer" } },
56262
+ conflict_memory_ids: { type: "array", items: { type: "integer" } },
56263
+ rationale: { type: "string" }
56264
+ },
56265
+ required: ["relation", "same_project", "source_should_alias_target", "confidence", "evidence_memory_ids", "conflict_memory_ids", "rationale"],
56266
+ additionalProperties: false
56267
+ };
56221
56268
  var JUDGE_SYSTEM_PROMPT = `You are a project-tag alias judge for one user's personal memory system.
56222
56269
 
56223
56270
  Decide whether two CANONICAL project tags refer to the same project identity.
@@ -56771,7 +56818,7 @@ async function judgePair(pair, tagsById, memoriesByTag) {
56771
56818
  const raw = await callSpec(spec, {
56772
56819
  system: JUDGE_SYSTEM_PROMPT,
56773
56820
  user: userPrompt,
56774
- ...spec.provider === "local" ? {} : { responseFormat: "json" },
56821
+ ...spec.provider === "local" ? { jsonSchema: JUDGE_SCHEMA, enableThinking: false } : { responseFormat: "json" },
56775
56822
  maxTokens: envInt("PROJECT_ALIAS_JUDGE_MAX_TOKENS", 8192)
56776
56823
  });
56777
56824
  if (!raw) throw new Error("project_alias_judge returned empty content");
@@ -58877,11 +58924,10 @@ var CLUSTER_SYSTEM = `You are a keyword clustering assistant for a personal memo
58877
58924
  Given a list of d_tags (short hyphenated keywords) with their occurrence counts,
58878
58925
  group semantically similar tags that refer to the same project or topic.
58879
58926
 
58880
- OUTPUT strict JSON array:
58881
- [
58882
- { "canonical": "<best-slug>", "members": ["<tag1>", "<tag2>", ...] },
58883
- ...
58884
- ]
58927
+ OUTPUT strict JSON object (NOT a bare array):
58928
+ { "clusters": [
58929
+ { "canonical": "<best-slug>", "members": ["<tag1>", "<tag2>", ...] }
58930
+ ] }
58885
58931
 
58886
58932
  Rules:
58887
58933
  - canonical must be one of the input tags (pick the most descriptive one) or a clean slug if none fit
@@ -58889,6 +58935,25 @@ Rules:
58889
58935
  - Only group tags that clearly refer to the same project/topic
58890
58936
  - Tags with no similar counterparts become their own single-member cluster
58891
58937
  - Do NOT merge unrelated topics just because they share one word`;
58938
+ var CLUSTER_SCHEMA = {
58939
+ type: "object",
58940
+ properties: {
58941
+ clusters: {
58942
+ type: "array",
58943
+ items: {
58944
+ type: "object",
58945
+ properties: {
58946
+ canonical: { type: "string" },
58947
+ members: { type: "array", items: { type: "string" } }
58948
+ },
58949
+ required: ["canonical", "members"],
58950
+ additionalProperties: false
58951
+ }
58952
+ }
58953
+ },
58954
+ required: ["clusters"],
58955
+ additionalProperties: false
58956
+ };
58892
58957
  async function clusterDTags(tags) {
58893
58958
  if (tags.length === 0) return [];
58894
58959
  const tagList = tags.map((t) => `${t.tag} (${t.cnt}x)`).join(", ");
@@ -58899,8 +58964,10 @@ ${tagList}`;
58899
58964
  raw = await callRole("clusterer", {
58900
58965
  system: CLUSTER_SYSTEM,
58901
58966
  user: userPrompt,
58902
- responseFormat: "json",
58903
- maxTokens: 512
58967
+ jsonSchema: CLUSTER_SCHEMA,
58968
+ enableThinking: false,
58969
+ maxTokens: 4096
58970
+ // was 512 — 50개 태그 클러스터링 출력이 잘려 invalid JSON 폴백됨 (ctx 8192 내 여유)
58904
58971
  });
58905
58972
  } catch (err2) {
58906
58973
  console.error("\u26A0\uFE0F [DTagPromoter] clusterer call failed, falling back to no clustering:", err2);
@@ -58908,8 +58975,9 @@ ${tagList}`;
58908
58975
  }
58909
58976
  let parsed;
58910
58977
  try {
58911
- parsed = JSON.parse(raw);
58912
- if (!Array.isArray(parsed)) throw new Error("not an array");
58978
+ const obj = JSON.parse(raw);
58979
+ parsed = obj?.clusters;
58980
+ if (!Array.isArray(parsed)) throw new Error("no clusters array");
58913
58981
  } catch {
58914
58982
  console.error("\u26A0\uFE0F [DTagPromoter] clusterer returned invalid JSON, falling back:", raw.slice(0, 200));
58915
58983
  return tags.map((t) => ({ canonical: t.tag, members: [t.tag], total: t.cnt }));
@@ -58993,35 +59061,63 @@ async function runDtagPromotion() {
58993
59061
 
58994
59062
  // src/librarian.ts
58995
59063
  init_db();
58996
- var LIBRARIAN_MSG_THRESHOLD = 30;
58997
- var LIBRARIAN_COOLDOWN_MS = 24 * 60 * 60 * 1e3;
58998
- var LIBRARIAN_RECENT_LIMIT = 50;
58999
- var LIBRARIAN_MAX_TOKENS = 32768;
59064
+ var LIBRARIAN_MSG_THRESHOLD = Number(process.env.LIBRARIAN_MSG_THRESHOLD ?? 30);
59065
+ var LIBRARIAN_COOLDOWN_MS = Number(process.env.LIBRARIAN_COOLDOWN_HOURS ?? 24) * 60 * 60 * 1e3;
59066
+ var LIBRARIAN_RECENT_SLICE = Number(process.env.LIBRARIAN_RECENT_SLICE ?? 25);
59067
+ var LIBRARIAN_HISTORY_SLICE = Number(process.env.LIBRARIAN_HISTORY_SLICE ?? 25);
59068
+ var LIBRARIAN_MAX_TOKENS = Number(process.env.LIBRARIAN_MAX_TOKENS ?? 2048);
59069
+ var LIBRARIAN_PROFILE_SCHEMA = {
59070
+ type: "object",
59071
+ properties: {
59072
+ core_profile: { type: ["string", "null"] },
59073
+ sub_profile: { type: ["string", "null"] }
59074
+ },
59075
+ required: ["core_profile", "sub_profile"],
59076
+ additionalProperties: false
59077
+ };
59000
59078
  var SYSTEM_PROMPT2 = `You are the Librarian for one user's personal memory system.
59001
59079
 
59002
59080
  YOUR JOB
59003
- Look at the user's recent first-person messages (role='user' only \u2014 ignore
59004
- assistant replies). Identify any STABLE, MEMORABLE facts about WHO THE USER IS
59005
- or HOW THEY WORK that should be promoted to their long-term profile.
59006
-
59007
- OUTPUT TWO SECTIONS:
59008
- 1. core_profile: critically important user identity (name, role, expertise,
59009
- strong preferences). Should be SHORT (5-10 lines max) and high-signal.
59010
- 2. sub_profile: secondary memorable info (tools, environment quirks,
59011
- ongoing project focus, working style preferences). Can be longer but
59012
- still curated.
59013
-
59014
- RULES
59015
- - The user is one person. Output is the WHOLE profile (replacing existing,
59016
- not appending) \u2014 so include relevant existing facts that are still true.
59081
+ Look at the user's first-person messages (role='user' only \u2014 ignore assistant
59082
+ replies). Identify any STABLE, DURABLE facts about WHO THE USER IS that should
59083
+ be promoted to their long-term profile.
59084
+
59085
+ OUTPUT TWO SECTIONS \u2014 these are STRICTLY SEPARATE categories:
59086
+
59087
+ 1. core_profile \u2014 DURABLE IDENTITY ONLY.
59088
+ Who the person IS: name, role, profession, expertise, stable long-term
59089
+ preferences. Should be SHORT (5-10 lines max) and high-signal.
59090
+ Example of the SHAPE only (fictional \u2014 never copy this content): "Backend
59091
+ engineer at a logistics startup; 10+ yrs Python; prefers terse,
59092
+ example-driven answers." Derive the actual content ONLY from the messages.
59093
+
59094
+ 2. sub_profile \u2014 CURRENT WORK AND ACTIVITY.
59095
+ What they are actively doing, building, or focused on: tools, environment
59096
+ details, ongoing projects, recent working style observations. Expected to
59097
+ change often. Can be longer but still curated.
59098
+
59099
+ CRITICAL IDENTITY vs. WORK DISTINCTION
59100
+ A user discussing, building, evaluating, debugging, or working on a topic \u2014
59101
+ including AI models, agent frameworks, or this memory system itself \u2014 is
59102
+ describing their WORK or CURRENT ACTIVITY, NOT their identity. A burst of
59103
+ messages about one subject means they are WORKING on it, not that it defines
59104
+ them. NEVER promote a work or project topic into core_profile. It belongs in
59105
+ sub_profile at most.
59106
+
59107
+ CONSERVATISM / NULL-PRESERVE RULE (most important rule)
59108
+ If the recent window contains NO new durable identity fact \u2014 only project work,
59109
+ meta-tooling, topic evaluation, or session-specific activity \u2014 return
59110
+ core_profile: null to PRESERVE the existing identity unchanged. DO NOT restate,
59111
+ rephrase, or "refresh" an existing core_profile just because you saw it. Null
59112
+ means "keep it as-is." Only set a non-null core_profile when there is an
59113
+ explicit, durable, first-person identity statement that is genuinely new.
59114
+ sub_profile may freely capture current projects and activity.
59115
+
59116
+ OTHER RULES
59017
59117
  - DO NOT invent facts not supported by the messages.
59018
- - DO NOT promote temporary state ("debugging X", "frustrated with Y").
59019
- Only stable identity / preferences.
59020
- - DO NOT promote third-party advice or system hints that show up in
59021
- messages \u2014 only what the user is saying ABOUT THEMSELVES.
59118
+ - DO NOT promote third-party advice or system hints \u2014 only what the user is
59119
+ saying ABOUT THEMSELVES.
59022
59120
  - Korean is fine. Match the language of the user's writing.
59023
- - If recent messages don't add anything new and existing profile is fine,
59024
- output the existing profile unchanged.
59025
59121
 
59026
59122
  FORMAT RULES FOR THE VALUES:
59027
59123
  - Both fields must be PLAIN PROSE TEXT \u2014 no nested JSON, no {}, [], key-value blobs.
@@ -59059,6 +59155,32 @@ async function checkGate() {
59059
59155
  const shouldRun = enoughNewMessages && (neverRan || cooldownPassed);
59060
59156
  return { userId, shouldRun, currentMsgCount, lastRunAt };
59061
59157
  }
59158
+ async function runLibrarianDeepPass(userPrompt, maxTokens) {
59159
+ const analysisSystemPrompt = `${SYSTEM_PROMPT2}
59160
+
59161
+ In this ANALYSIS PASS, think deeply and reason freely about what should go in each profile field.
59162
+ Do NOT output JSON yet \u2014 write your analysis as plain prose reasoning.`;
59163
+ const analysis = await callRole("librarian", {
59164
+ system: analysisSystemPrompt,
59165
+ user: userPrompt,
59166
+ maxTokens,
59167
+ enableThinking: true
59168
+ // jsonSchema 未指定 → thinking ON が有効になる
59169
+ });
59170
+ const pass2User = `${userPrompt}
59171
+
59172
+ ANALYSIS FROM PREVIOUS REASONING PASS (use as additional context):
59173
+ ${analysis}
59174
+
59175
+ Now emit the final JSON profile only.`;
59176
+ return callRole("librarian", {
59177
+ system: SYSTEM_PROMPT2,
59178
+ user: pass2User,
59179
+ maxTokens,
59180
+ jsonSchema: LIBRARIAN_PROFILE_SCHEMA,
59181
+ enableThinking: false
59182
+ });
59183
+ }
59062
59184
  async function runLibrarian() {
59063
59185
  if (librarianRunning) return;
59064
59186
  const gate = await checkGate();
@@ -59082,27 +59204,55 @@ async function runLibrarian() {
59082
59204
  AND is_active = TRUE
59083
59205
  ORDER BY created_at DESC
59084
59206
  LIMIT $2`,
59085
- [gate.userId, LIBRARIAN_RECENT_LIMIT]
59207
+ [gate.userId, LIBRARIAN_RECENT_SLICE]
59086
59208
  );
59087
59209
  if (recentR.rows.length === 0) return;
59088
- const messagesText = recentR.rows.reverse().map(
59089
- (r, i) => `[#${i + 1} @ ${r.created_at?.toISOString().slice(0, 19) ?? ""}] ${r.message}`
59210
+ const historyR = await db.query(
59211
+ `SELECT message, created_at
59212
+ FROM memory
59213
+ WHERE user_id = $1
59214
+ AND role = 'user'
59215
+ AND is_active = TRUE
59216
+ ORDER BY created_at ASC
59217
+ LIMIT $2`,
59218
+ [gate.userId, LIBRARIAN_HISTORY_SLICE]
59219
+ );
59220
+ const recentKeys = new Set(
59221
+ recentR.rows.map(
59222
+ (r) => `${r.created_at?.toISOString() ?? ""}|${r.message}`
59223
+ )
59224
+ );
59225
+ const historicalRows = historyR.rows.filter(
59226
+ (r) => !recentKeys.has(`${r.created_at?.toISOString() ?? ""}|${r.message}`)
59227
+ );
59228
+ const recentSorted = [...recentR.rows].reverse();
59229
+ const historicalSorted = [...historicalRows];
59230
+ const formatRows = (rows, startIdx) => rows.map(
59231
+ (r, i) => `[#${startIdx + i + 1} @ ${r.created_at?.toISOString().slice(0, 19) ?? ""}] ${r.message}`
59090
59232
  ).join("\n\n");
59091
- const userPrompt = `EXISTING PROFILE (subject to update):
59233
+ const historicalText = historicalSorted.length > 0 ? `
59234
+ EARLIER MESSAGES \u2014 historical context (identity anchors, NOT current activity):
59235
+ ${formatRows(historicalSorted, 0)}
59236
+ ` : "";
59237
+ const recentText = `
59238
+ RECENT MESSAGES \u2014 most recent ${recentSorted.length} messages:
59239
+ ${formatRows(recentSorted, historicalSorted.length)}
59240
+ `;
59241
+ const userPrompt = `EXISTING PROFILE (preserve unless a new durable identity fact appears):
59092
59242
  core_profile:
59093
59243
  ${before.core_profile ?? "(empty)"}
59094
59244
 
59095
59245
  sub_profile:
59096
59246
  ${before.sub_profile ?? "(empty)"}
59097
-
59098
- RECENT USER MESSAGES (most recent ${recentR.rows.length}, role='user'):
59099
- ${messagesText}
59100
-
59247
+ ${historicalText}${recentText}
59101
59248
  Task: produce updated core_profile and sub_profile JSON per the system prompt.`;
59102
- const raw = await callRole("librarian", {
59249
+ const deepMode = process.env.LIBRARIAN_DEEP_THINKING === "true";
59250
+ const raw = deepMode ? await runLibrarianDeepPass(userPrompt, LIBRARIAN_MAX_TOKENS) : await callRole("librarian", {
59103
59251
  system: SYSTEM_PROMPT2,
59104
59252
  user: userPrompt,
59105
- maxTokens: LIBRARIAN_MAX_TOKENS
59253
+ maxTokens: LIBRARIAN_MAX_TOKENS,
59254
+ jsonSchema: LIBRARIAN_PROFILE_SCHEMA,
59255
+ enableThinking: false
59106
59256
  });
59107
59257
  if (!raw) {
59108
59258
  throw new Error("Librarian returned empty content (reasoning token budget exceeded?)");
@@ -59138,8 +59288,9 @@ Task: produce updated core_profile and sub_profile JSON per the system prompt.`;
59138
59288
  `UPDATE users SET librarian_msg_count_at_run = $1 WHERE user_id = $2`,
59139
59289
  [gate.currentMsgCount, gate.userId]
59140
59290
  );
59291
+ const totalMsgs = recentSorted.length + historicalSorted.length;
59141
59292
  console.error(
59142
- `\u{1F4DA} [Librarian] done \u2014 ${recentR.rows.length} msgs, profile ${changed ? "updated" : "unchanged"}`
59293
+ `\u{1F4DA} [Librarian] done \u2014 ${totalMsgs} msgs (${recentSorted.length} recent + ${historicalSorted.length} historical), profile ${changed ? "updated" : "unchanged"}`
59143
59294
  );
59144
59295
  } catch (err2) {
59145
59296
  console.error("\u26A0\uFE0F [Librarian] run failed (retries in 24h):", err2);
@@ -59330,11 +59481,56 @@ async function maybeRunLibrarian() {
59330
59481
  console.error("\u26A0\uFE0F [Librarian] unhandled error (non-blocking):", err2);
59331
59482
  }
59332
59483
  }
59333
- function startColdPathWorker() {
59484
+ var COLDPATH_ADVISORY_LOCK_KEY = 4242000017;
59485
+ var lockClient = null;
59486
+ async function acquireColdPathLock() {
59487
+ if (lockClient) return true;
59488
+ const client2 = await db.getClient();
59489
+ try {
59490
+ const res = await client2.query("SELECT pg_try_advisory_lock($1) AS got", [COLDPATH_ADVISORY_LOCK_KEY]);
59491
+ if (res.rows[0]?.got === true) {
59492
+ lockClient = client2;
59493
+ return true;
59494
+ }
59495
+ client2.release();
59496
+ return false;
59497
+ } catch (err2) {
59498
+ try {
59499
+ client2.release();
59500
+ } catch {
59501
+ }
59502
+ throw err2;
59503
+ }
59504
+ }
59505
+ async function releaseColdPathLock() {
59506
+ const c = lockClient;
59507
+ if (!c) return;
59508
+ lockClient = null;
59509
+ try {
59510
+ await c.query("SELECT pg_advisory_unlock($1)", [COLDPATH_ADVISORY_LOCK_KEY]);
59511
+ } catch {
59512
+ }
59513
+ try {
59514
+ c.release();
59515
+ } catch {
59516
+ }
59517
+ }
59518
+ async function startColdPathWorker() {
59334
59519
  if (process.env.COLD_PATH_ENABLED === "false") {
59335
59520
  console.error("\u{1F535} [ColdPath] disabled (COLD_PATH_ENABLED=false)");
59336
59521
  return;
59337
59522
  }
59523
+ let got = false;
59524
+ try {
59525
+ got = await acquireColdPathLock();
59526
+ } catch (err2) {
59527
+ console.error("\u274C [ColdPath] advisory lock acquire failed \u2014 not starting:", err2);
59528
+ return;
59529
+ }
59530
+ if (!got) {
59531
+ console.error("\u{1F512} [ColdPath] another instance holds the lock \u2014 this process will not process");
59532
+ return;
59533
+ }
59338
59534
  const intervalSec = envInt3("COLD_PATH_INTERVAL_SEC", DEFAULT_INTERVAL_SEC);
59339
59535
  const warmupSec = envInt3("COLD_PATH_WARMUP_SEC", DEFAULT_WARMUP_SEC);
59340
59536
  console.error(`\u{1F535} [ColdPath] starting \u2014 warmup ${warmupSec}s, interval ${intervalSec}s`);
@@ -59356,6 +59552,7 @@ function stopColdPathWorker() {
59356
59552
  if (intervalTimer) clearInterval(intervalTimer);
59357
59553
  warmupTimer = null;
59358
59554
  intervalTimer = null;
59555
+ void releaseColdPathLock();
59359
59556
  console.error("\u{1F535} [ColdPath] stopped");
59360
59557
  }
59361
59558
  async function drainColdPath(maxTicks = 12) {
@@ -60187,6 +60384,7 @@ function printHelp() {
60187
60384
 
60188
60385
  Usage:
60189
60386
  mcp-agents-memory Run the MCP server (stdio).
60387
+ mcp-agents-memory coldpath Run ONLY the cold-path worker as a standalone always-on daemon (no MCP server). For the processing/GPU machine via systemd.
60190
60388
  mcp-agents-memory setup Interactive setup \u2014 write config to ~/.config/mcp-agents-memory/.env and run migrations.
60191
60389
  mcp-agents-memory migrate Apply any pending DB migrations against the configured database.
60192
60390
  mcp-agents-memory help Show this message.
@@ -60236,13 +60434,36 @@ async function runMcpServer() {
60236
60434
  console.error("\u274C Fatal error during startup:", err2);
60237
60435
  process.exit(1);
60238
60436
  }
60239
- startColdPathWorker();
60437
+ void startColdPathWorker().catch((err2) => console.error("\u274C [ColdPath] start failed:", err2));
60438
+ }
60439
+ async function runColdPathDaemon() {
60440
+ process.env.COLD_PATH_ENABLED = "true";
60441
+ process.on("SIGTERM", () => {
60442
+ void shutdown("SIGTERM");
60443
+ });
60444
+ process.on("SIGINT", () => {
60445
+ void shutdown("SIGINT");
60446
+ });
60447
+ process.on("SIGHUP", () => {
60448
+ void shutdown("SIGHUP");
60449
+ });
60450
+ try {
60451
+ await db.connect();
60452
+ } catch (err2) {
60453
+ console.error("\u274C [ColdPathDaemon] DB connect failed \u2014 exiting (systemd will retry):", err2);
60454
+ process.exit(1);
60455
+ }
60456
+ console.error(`\u{1F9CA} Cold-path daemon (v${PACKAGE_VERSION2}) started \u2014 no MCP server`);
60457
+ await startColdPathWorker();
60240
60458
  }
60241
60459
  async function cli() {
60242
60460
  const cmd = process.argv[2];
60243
60461
  if (!cmd || cmd === "serve") {
60244
60462
  return runMcpServer();
60245
60463
  }
60464
+ if (cmd === "coldpath") {
60465
+ return runColdPathDaemon();
60466
+ }
60246
60467
  if (cmd === "help" || cmd === "--help" || cmd === "-h") {
60247
60468
  printHelp();
60248
60469
  return;
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "mcp-agents-memory",
3
- "version": "0.9.10",
3
+ "version": "0.9.12",
4
4
  "description": "Long-term memory for AI agents with provenance tracking",
5
5
  "main": "build/index.js",
6
6
  "type": "module",