pipecat-memorysync 1.1.0__tar.gz → 1.1.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -2,6 +2,19 @@
2
2
 
3
3
  All notable changes to `pipecat-memorysync` are documented here.
4
4
 
5
+ ## 1.1.1 — 2026-08-29
6
+
7
+ - **Immediate, loss-proof capture.** 1.1.0's turn-complete merging
8
+ deferred the current utterance until the turn completed — on a browser
9
+ disconnect that deferral raced the brief cancel salvage window and
10
+ could lose the newest (usually most important) turn over slow
11
+ networks. Capture is immediate again: every new user/assistant message
12
+ is in flight the moment its frame passes, BEFORE the LLM replies.
13
+ Junk filtering and fact extraction now happen server-side (the
14
+ platform's low-value gate + conversational ingestion), so client-side
15
+ merging is unnecessary. Suite: 17 checks on Pipecat v1.8.1, including
16
+ the exact disconnect sequence that previously lost data.
17
+
5
18
  ## 1.1.0 — 2026-08-29
6
19
 
7
20
  - **Turn-complete capture.** Voice aggregators split one utterance across
@@ -1,6 +1,6 @@
1
1
  Metadata-Version: 2.5
2
2
  Name: pipecat-memorysync
3
- Version: 1.1.0
3
+ Version: 1.1.1
4
4
  Summary: MemorySync for Pipecat: budgeted memory recall that never stalls a voice reply, delta-only conversation persistence with idempotency seeds, and a pipeline that cannot be broken by a memory outage.
5
5
  Project-URL: Homepage, https://docs.memorysync.io/guides/pipecat
6
6
  Project-URL: Documentation, https://docs.memorysync.io/guides/pipecat
@@ -81,10 +81,12 @@ enriched or not, on time.
81
81
  - **Budgeted recall.** Enrichment runs under a hard timeout (default
82
82
  **1.2 s**). A slow or dead memory backend means an unenriched frame, never a
83
83
  stalled voice reply.
84
- - **Turn-complete capture.** Voice aggregators split one utterance across
85
- several context messages at speech pauses; consecutive fragments merge
86
- into ONE stored turn when the assistant's reply completes it (the
87
- in-progress tail flushes at end of call) — no per-fragment junk rows.
84
+ - **Immediate, loss-proof capture.** Every new message is sent the moment
85
+ its frame passes — the current utterance is in flight BEFORE the LLM
86
+ replies, so a disconnect can never lose it. Junk filtering and fact
87
+ extraction happen server-side (the platform's low-value gate +
88
+ conversational ingestion): the memory store receives curated facts,
89
+ not transcript clutter.
88
90
  - **Delta-only capture.** Only messages *not seen before* are stored, tracked
89
91
  by deterministic idempotency seeds. Growing a 50-message context does not
90
92
  re-store 50 messages per turn.
@@ -57,10 +57,12 @@ enriched or not, on time.
57
57
  - **Budgeted recall.** Enrichment runs under a hard timeout (default
58
58
  **1.2 s**). A slow or dead memory backend means an unenriched frame, never a
59
59
  stalled voice reply.
60
- - **Turn-complete capture.** Voice aggregators split one utterance across
61
- several context messages at speech pauses; consecutive fragments merge
62
- into ONE stored turn when the assistant's reply completes it (the
63
- in-progress tail flushes at end of call) — no per-fragment junk rows.
60
+ - **Immediate, loss-proof capture.** Every new message is sent the moment
61
+ its frame passes — the current utterance is in flight BEFORE the LLM
62
+ replies, so a disconnect can never lose it. Junk filtering and fact
63
+ extraction happen server-side (the platform's low-value gate +
64
+ conversational ingestion): the memory store receives curated facts,
65
+ not transcript clutter.
64
66
  - **Delta-only capture.** Only messages *not seen before* are stored, tracked
65
67
  by deterministic idempotency seeds. Growing a 50-message context does not
66
68
  re-store 50 messages per turn.
@@ -1,3 +1,3 @@
1
1
  """Version for pipecat-memorysync. Single source of truth."""
2
2
 
3
- __version__ = "1.1.0"
3
+ __version__ = "1.1.1"
@@ -7,13 +7,13 @@ predecessors don't:
7
7
  1. **Recall is budgeted.** Enrichment waits at most ``recall_timeout``
8
8
  seconds (default 1.2). On timeout or failure the context frame passes
9
9
  through unenriched — a voice reply is never stalled by a slow network.
10
- 2. **Capture is turn-complete and delta-only.** Voice aggregators split
11
- one utterance across several context messages at speech pauses; this
12
- service merges consecutive new user fragments and stores them as ONE
13
- verbatim turn once the assistant's reply marks the turn complete (the
14
- in-progress tail is flushed at end of call). Only NEW messages are
15
- considered each frame — never the whole conversation re-sent — and
16
- every stored turn carries a cross-adapter fnv1a64 idempotency seed.
10
+ 2. **Capture is immediate and delta-only.** Every NEW user/assistant
11
+ message is sent to the platform the moment it appears in a context
12
+ frame — the current utterance is already in flight BEFORE the LLM
13
+ replies, so a disconnect can never lose it. Junk filtering and fact
14
+ extraction are the platform's job (the low-value gate + conversational
15
+ ingestion), not the client's. Only new messages are sent each frame,
16
+ with cross-adapter fnv1a64 idempotency seeds.
17
17
  3. **Nothing raises, nothing is dropped.** Every failure path logs and
18
18
  pushes the ORIGINAL frame through. The pipeline cannot stall and the
19
19
  LLM always gets its context.
@@ -141,14 +141,9 @@ class MemorySyncMemoryService(FrameProcessor):
141
141
  self.params = params
142
142
 
143
143
  self._stored_seeds: Set[str] = set()
144
- # Message-level consumption ledger: which context messages have
145
- # been merged into a stored turn. Distinct from _stored_seeds
146
- # (which dedups at merged-turn granularity) so a failed store can
147
- # release its messages for re-capture on the next frame.
144
+ # Message-level ledger: which context messages have been sent.
145
+ # A failed store releases its message so the next frame retries.
148
146
  self._seen_messages: Set[str] = set()
149
- # Snapshot of the in-progress user utterance (fragments after the
150
- # last assistant reply) — flushed as one turn at end of call.
151
- self._tail_pending: List[tuple[str, str]] = []
152
147
  self._store_tasks: Set[asyncio.Task] = set()
153
148
  self._last_query: Optional[str] = None
154
149
 
@@ -249,25 +244,21 @@ class MemorySyncMemoryService(FrameProcessor):
249
244
  body = "\n".join(lines)
250
245
  return f"{self.params.system_prompt}\n{body}\n\n{CONTEXT_GUARD}"
251
246
 
252
- # ── capture: turn-complete, delta-only, background ─────────────────
247
+ # ── capture: immediate, delta-only, background ─────────────────────
253
248
 
254
249
  def _capture_delta(self, context: Any) -> None:
255
- """Queue storage for completed turns built from NEW messages only.
256
-
257
- Voice aggregators append each speech fragment as its own user
258
- message ("Do you know that my favorite", "dinner", "and", …).
259
- Storing them individually floods memory with junk rows, so:
260
-
261
- - consecutive NEW user messages merge into ONE turn, stored the
262
- moment an assistant message follows them (turn completed);
263
- - user messages after the last assistant reply are an in-progress
264
- utterance — they stay pending (snapshotted for the end-of-call
265
- flush) instead of being stored piecemeal;
266
- - NEW assistant messages store immediately, flushing any pending
267
- user group first so ordering survives.
250
+ """Queue storage for NEW messages the moment they appear.
251
+
252
+ The platform owns junk filtering (the low-value gate) and fact
253
+ extraction (conversational ingestion), so the adapter's single
254
+ duty is delivering every turn reliably and EARLY. The current
255
+ user utterance is in the frame BEFORE the LLM replies — sending
256
+ it immediately means a mid-call disconnect can never lose it.
257
+ (1.1.0 deferred the current utterance until the turn completed;
258
+ that deferral raced the disconnect salvage window and could drop
259
+ the newest — usually most important — turn. Never again.)
268
260
  """
269
261
  header = self.params.system_prompt
270
- entries: List[tuple[str, str]] = [] # (speaker_role, trimmed_text)
271
262
  for message in context.get_messages():
272
263
  role = message.get("role")
273
264
  if role not in ("user", "assistant"):
@@ -277,57 +268,12 @@ class MemorySyncMemoryService(FrameProcessor):
277
268
  continue # our own injection (user-role mode) never re-enters
278
269
  speaker_role = "human" if role == "user" else "ai"
279
270
  trimmed = text if len(text) <= MAX_TURN_CHARS else text[:MAX_TURN_CHARS] + "…"
280
- entries.append((speaker_role, trimmed))
281
-
282
- last_ai = -1
283
- for index, (speaker_role, _text) in enumerate(entries):
284
- if speaker_role == "ai":
285
- last_ai = index
286
-
287
- group_texts: List[str] = []
288
- group_keys: List[str] = []
289
- for index, (speaker_role, text) in enumerate(entries):
290
- key = f"{speaker_role}:{text}"
291
- if speaker_role == "human":
292
- if index > last_ai:
293
- continue # in-progress utterance — tail snapshot below
294
- if key in self._seen_messages:
295
- continue
296
- group_texts.append(text)
297
- group_keys.append(key)
298
- else:
299
- if group_texts:
300
- self._store_merged(group_texts, group_keys)
301
- group_texts, group_keys = [], []
302
- if key not in self._seen_messages:
303
- self._seen_messages.add(key)
304
- self._bound_seen()
305
- self._queue_store("ai", text, [key])
306
-
307
- self._tail_pending = [
308
- (text, f"human:{text}")
309
- for speaker_role, text in entries[last_ai + 1:]
310
- if speaker_role == "human" and f"human:{text}" not in self._seen_messages
311
- ]
312
-
313
- def _store_merged(self, texts: List[str], keys: List[str]) -> None:
314
- """One utterance from its fragments: mark consumed, merge, store."""
315
- for key in keys:
271
+ key = f"{speaker_role}:{trimmed}"
272
+ if key in self._seen_messages:
273
+ continue
316
274
  self._seen_messages.add(key)
317
- self._bound_seen()
318
- merged = " ".join(texts)
319
- if len(merged) > MAX_TURN_CHARS:
320
- merged = merged[:MAX_TURN_CHARS] + "…"
321
- self._queue_store("human", merged, keys)
322
-
323
- def _flush_tail(self) -> None:
324
- """Store the in-progress utterance (end-of-call, no reply coming)."""
325
- if not self._tail_pending:
326
- return
327
- texts = [text for text, _key in self._tail_pending]
328
- keys = [key for _text, key in self._tail_pending]
329
- self._tail_pending = []
330
- self._store_merged(texts, keys)
275
+ self._bound_seen()
276
+ self._queue_store(speaker_role, trimmed, [key])
331
277
 
332
278
  def _bound_seen(self) -> None:
333
279
  if len(self._seen_messages) > 4096:
@@ -360,7 +306,7 @@ class MemorySyncMemoryService(FrameProcessor):
360
306
  )
361
307
  except Exception as exc: # noqa: BLE001
362
308
  # The write never landed: release both ledgers so the next
363
- # frame re-captures these messages and retries the store.
309
+ # frame re-captures this message and retries the store.
364
310
  self._stored_seeds.discard(seed)
365
311
  for key in msg_keys:
366
312
  self._seen_messages.discard(key)
@@ -392,12 +338,7 @@ class MemorySyncMemoryService(FrameProcessor):
392
338
  # ── lifecycle ──────────────────────────────────────────────────────
393
339
 
394
340
  async def _flush(self, *, timeout: float) -> None:
395
- """Store the in-progress utterance, then let queued stores land,
396
- bounded. Never raises."""
397
- try:
398
- self._flush_tail()
399
- except Exception: # noqa: BLE001
400
- pass
341
+ """Let queued stores land, bounded. Never raises."""
401
342
  pending = [t for t in self._store_tasks if not t.done()]
402
343
  if not pending:
403
344
  return
@@ -252,61 +252,47 @@ async def test_missing_user_id_is_loud_at_construction(mock):
252
252
  MemorySyncMemoryService(api_key="ms_x", user_id="", transport=mock.transport())
253
253
 
254
254
 
255
- # ── capture: turn-complete merging (voice fragmenting) ────────────────
255
+ # ── capture: immediate and loss-proof ─────────────────────────────────
256
256
 
257
257
 
258
- async def test_speech_fragments_merge_into_one_turn(mock, make_service):
259
- """VAD pauses split one utterance across context messages — they must
260
- store as ONE merged row when the assistant reply completes the turn,
261
- never as per-fragment junk rows ("and", "dinner", …)."""
258
+ async def test_current_utterance_stores_before_any_reply(mock, make_service):
259
+ """The user's words must be in flight the moment the frame passes —
260
+ BEFORE the LLM replies — so a disconnect can never lose them."""
262
261
  service = make_service()
263
- fragments = [
264
- "Do you know that my favorite",
265
- "dinner",
266
- "and",
267
- "dinner is my favorite dish is biryani.",
268
- ]
269
- # The aggregator's view grows frame by frame while the user speaks…
270
- growing = [
271
- context_of([{"role": "user", "content": f} for f in fragments[: i + 1]])
272
- for i in range(len(fragments))
273
- ]
274
- # …and the turn completes when the assistant replies.
275
- completed = context_of(
276
- [{"role": "user", "content": f} for f in fragments]
277
- + [{"role": "assistant", "content": "Got it, biryani it is!"}]
278
- )
279
- await drive(service, *growing, completed)
280
- await drain(mock, 2)
281
-
282
- texts = sorted(r["text"] for r in mock.rows)
283
- assert texts == [
284
- "ai: Got it, biryani it is!",
285
- "human: Do you know that my favorite dinner and dinner is my favorite dish is biryani.",
262
+ ctx = context_of([
263
+ {"role": "user", "content": "My age is twenty two and I completed my bachelor's."},
264
+ ])
265
+ await drive(service, ctx)
266
+ await drain(mock, 1)
267
+ assert [r["text"] for r in mock.rows] == [
268
+ "human: My age is twenty two and I completed my bachelor's."
286
269
  ]
287
- assert mock.add_turn_calls() == 2, "four fragments + one reply = exactly two rows"
288
270
 
289
271
 
290
- async def test_in_progress_utterance_flushes_once_at_end_of_call(mock, make_service):
291
- """A call that ends mid-utterance (no reply coming) must not lose the
292
- tail — and must store it merged, not per fragment."""
272
+ async def test_demo_disconnect_sequence_loses_nothing(mock, make_service):
273
+ """The exact sequence that lost data in 1.1.0: greeting frame, then a
274
+ frame carrying the reply + the important utterance, then immediate
275
+ CancelFrame (browser disconnect). Every message must already be
276
+ stored — nothing may depend on a post-cancel flush window."""
293
277
  service = make_service()
294
- ctx1 = context_of([{"role": "user", "content": "my name is Abdullah"}])
295
- ctx2 = context_of([
296
- {"role": "user", "content": "my name is Abdullah"},
297
- {"role": "user", "content": "and my favorite color is black"},
278
+ f1 = context_of([{"role": "user", "content": "Hello there, anyone home?"}])
279
+ f2 = context_of([
280
+ {"role": "user", "content": "Hello there, anyone home?"},
281
+ {"role": "assistant", "content": "Hi! How can I help you today?"},
282
+ {"role": "user", "content": "My age is twenty two and I completed my bachelor's."},
298
283
  ])
299
- await drive(service, ctx1, ctx2) # run_test's EndFrame triggers the flush
300
- await drain(mock, 1)
301
-
302
- assert [r["text"] for r in mock.rows] == [
303
- "human: my name is Abdullah and my favorite color is black"
284
+ await drive(service, f1, f2) # run_test tears the pipeline down right after
285
+ await drain(mock, 3)
286
+ texts = sorted(r["text"] for r in mock.rows)
287
+ assert texts == [
288
+ "ai: Hi! How can I help you today?",
289
+ "human: Hello there, anyone home?",
290
+ "human: My age is twenty two and I completed my bachelor's.",
304
291
  ]
305
- assert mock.add_turn_calls() == 1
306
292
 
307
293
 
308
- async def test_failed_store_releases_fragments_for_retry(mock, make_service):
309
- """A store that never landed must not consume its messages — the next
294
+ async def test_failed_store_releases_message_for_retry(mock, make_service):
295
+ """A store that never landed must not consume its message — the next
310
296
  frame re-captures and retries, converging on one row."""
311
297
  service = make_service()
312
298
  mock.fail_next = 503