@johpaz/hive-sdk 0.4.8 → 0.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,6 +1,57 @@
1
1
  # Changelog
2
2
 
3
- ## Sin publicar
3
+ ## 0.5.0
4
+
5
+ ### Jev — plano de decisión (OpenRouter Decisions)
6
+
7
+ Portado de hive 1.1.0. Jev decide por turno qué historial, herramientas,
8
+ skills, notas y reglas del playbook entran al contexto; entre iteraciones poda
9
+ resultados viejos de herramientas y sugiere la siguiente acción; decide si un
10
+ lote de herramientas corre en paralelo, y conoce el mapa del enjambre
11
+ (especialistas y estado de cada MCP) para recomendar a quién delegar.
12
+ **Sin clave de OpenRouter, Jev no existe y todo corre igual que antes.**
13
+
14
+ - **Clave inyectable por llamada**: `jev?: { apiKey, mcpSettingsPath? } | false`
15
+ en `AgentLoopOptions`, `compileContext`, `IsolatedAgentOptions`,
16
+ `runRoleSwarm` y `runSwarm`, igual que `credentials`. `false` lo apaga;
17
+ sin la opción decide la fila `openrouter` del inquilino actual. Con un
18
+ inquilino activo **nunca** se usa `OPENROUTER_API_KEY` ni la caché de
19
+ secretos del proceso: la clave de la plataforma no se usa en nombre de un
20
+ cliente.
21
+ - **Estado por inquilino**: fallos, cooldown y totales se llevan por
22
+ `currentTenant()`; una clave inválida de un cliente no pone en fallback a
23
+ los demás.
24
+ - **Evento para el host**: cada decisión llega por `onStep` como
25
+ `StepEvent` `jev_decision` (`jev`: agente, tipo, resumen, tokens ahorrados,
26
+ latencia, costo, especialista recomendado, MCP apagados). También se emite
27
+ `canvas:jev_decision` / `canvas:jev_status` para hosts tipo hive.
28
+ - **Uso y costo**: `recordJevDecision` y los campos `jev*` de
29
+ `UsageRollupDoc`; `getUsageStats()` devuelve `jev` con el total y el
30
+ desglose por agente. El ahorro se estima (caracteres/4) y se cotiza con el
31
+ modelo del agente asesorado.
32
+ - **Catálogo**: modelo `openrouter/typesafe/jev-1.13` con `modelType:
33
+ "decision"`, excluido de `get_available_models` (y de `getDefaultLLM`, que
34
+ sólo toma modelos `llm`).
35
+ - Herramienta nueva `conversation_read`: recupera mensajes o notas que Jev
36
+ dejó fuera del contexto, siempre dentro del hilo actual.
37
+ - `executeToolBatch` acepta `parallelToolCalls` por lote.
38
+ - `NarrationEventDoc.kind` suma `"decision"`: un host puede anotar las
39
+ decisiones de Jev en `narrationEvents` para sus vistas de actividad.
40
+ `shouldDeliverToChannel` nunca lo entrega a un canal.
41
+ - `loadDurableProviderApiKey(id)`: lee la clave sólo de la colección
42
+ `secrets` (particionada por inquilino), sin pasar por la caché de proceso ni
43
+ el llavero del SO.
44
+
45
+ **Privacidad.** Con Jev activo se envían a la API de decisiones de OpenRouter
46
+ (`https://openrouter.ai/api/alpha/decisions`), con la clave del workspace:
47
+ el objetivo del turno (hasta 3 500 caracteres), extractos de hasta 450
48
+ caracteres de mensajes previos del hilo, nombre y descripción de herramientas
49
+ y skills candidatas, notas del scratchpad y reglas del playbook (hasta 350
50
+ caracteres cada una), extractos de resultados de herramientas (hasta 650
51
+ caracteres), argumentos de llamadas en lote (hasta 700) y el mapa del enjambre
52
+ (ids, nombres y descripciones de especialistas, nombres y estado de los MCP).
53
+ No se envían ids de MCP con prefijo de inquilino, credenciales ni adjuntos
54
+ binarios. Un host que no quiera enviar nada pasa `jev: false`.
4
55
 
5
56
  ### Plataforma y seguridad
6
57
 
@@ -579,6 +630,29 @@
579
630
  job que genera un scaffold con `create-app` y lo typechequea contra el SDK de
580
631
  ese commit.
581
632
 
633
+ ## 0.4.9
634
+
635
+ ### Corregido
636
+
637
+ - **La compactación de historial dejaba de ser excepcional y corría casi en cada
638
+ turno.** `agent.context.compactionThreshold` es una proporción de la ventana
639
+ del modelo —su valor por defecto es `0.8`, o sea el 80 %— pero se leía como un
640
+ número de tokens: el umbral efectivo quedaba en 0.8 tokens. Cualquier hilo con
641
+ más de cinco mensajes se resumía en cada turno, lo que cuesta una llamada extra
642
+ al modelo por mensaje y reemplaza el historial por un resumen desde el primer
643
+ intercambio. Ahora un valor menor o igual a 1 se aplica sobre la ventana del
644
+ modelo y uno mayor se sigue leyendo como tokens, para quien fijó un número
645
+ absoluto. El umbral se calcula además con la ventana del modelo que corre el
646
+ turno, no con la del coordinador.
647
+ - **El resumen se pedía con una credencial global.** `compactThread` resolvía el
648
+ modelo con `getDefaultLLM()` y llamaba a `resolveProviderConfig` sin
649
+ credenciales, así que la llave salía del secret store, del llavero del sistema
650
+ o del entorno del proceso. En una instalación de un solo usuario da igual; en
651
+ una multi-inquilino significa resumir la conversación de un cliente con la
652
+ llave de la plataforma o de otro cliente. `maybeCompact` y `compactThread`
653
+ aceptan ahora el modelo y las credenciales del turno (`CompactionLLM`), y el
654
+ agent loop les pasa los suyos. Sin ese dato se comportan como antes.
655
+
582
656
  ## 0.1.5
583
657
 
584
658
  Sincronización del SDK con el runtime de agentes de `hive`. **Trae rupturas de
package/README.md CHANGED
@@ -271,4 +271,4 @@ npm view @johpaz/hive-sdk dist-tags # verificar después del release
271
271
 
272
272
  ---
273
273
 
274
- *Hive SDK v0.4.8 — MIT*
274
+ *Hive SDK v0.5.0 — MIT*
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@johpaz/hive-sdk",
3
- "version": "0.4.8",
3
+ "version": "0.5.0",
4
4
  "private": false,
5
5
  "description": "Hive SDK — The Agent Harness SDK. Build, deploy, and scale AI agent applications with multi-channel support, context engineering, and swarm orchestration.",
6
6
  "license": "MIT",
@@ -23,9 +23,12 @@ import { callLLM, resolveProviderConfig, getDefaultLLM, type LLMMessage, type Pr
23
23
  import { addMessage } from "./conversation-store.ts"
24
24
  import { saveTrace, recordLLMUsage } from "./tracer.ts"
25
25
  import { maybeCompact, clearOldToolResults } from "./compaction.ts"
26
- import { emitCanvas } from "../canvas/emitter.ts"
26
+ import { emitCanvas, type CanvasJevDecision } from "../canvas/emitter.ts"
27
27
  import type { MCPClientManager } from "../mcp/index.ts"
28
28
  import { compileContext } from "./context-compiler.ts"
29
+ import { MINIMAL_TOOLS } from "./minimal-loadout.ts"
30
+ import { jevWantsParallel, planJevIteration } from "./jev-planner.ts"
31
+ import { emitJevDecision, type JevOption } from "./jev-decisions.ts"
29
32
  import { formatToolResult } from "../utils/toon.ts"
30
33
  import { redactBinaryStrings } from "../utils/redact-binary.ts"
31
34
  import { resolveUserId, resolveAgentId } from "../storage/onboarding.ts"
@@ -52,6 +55,10 @@ import { getNarration } from "../events/tool-narration.ts"
52
55
 
53
56
  const log = logger.child("agent-loop")
54
57
 
58
+ const JEV_ACTION_LABELS: Record<string, string> = {
59
+ continue: "Continuar", delegate: "Delegar", discover: "Descubrir", finish: "Cerrar",
60
+ }
61
+
55
62
  // Per-operation budget for a single LLM call — NOT an aggregate deadline for the
56
63
  // whole turn. Each call gets its own fresh window; a slow-but-healthy multi-step
57
64
  // turn (many quick operations) is never killed just for taking a while overall.
@@ -226,6 +233,12 @@ export interface AgentLoopOptions {
226
233
  * dos inquilinos concurrentes en el mismo proceso compartían credencial.
227
234
  */
228
235
  credentials?: ProviderCredentials
236
+ /**
237
+ * Jev (OpenRouter Decisions) for this run. `{ apiKey }` uses that key,
238
+ * `false` turns it off, undefined reads the current tenant's `openrouter`
239
+ * provider row. Travels with the run like `credentials`.
240
+ */
241
+ jev?: JevOption
229
242
  /** Whether to resume from a previously saved checkpoint */
230
243
  resume?: boolean
231
244
  /** Run budget — overrides agent.max_iterations when set */
@@ -258,12 +271,17 @@ export interface AgentLoopOptions {
258
271
  export type { StepEvent as AgentStepEvent }
259
272
 
260
273
  export interface StepEvent {
261
- type: "text" | "tool_call" | "tool_result"
274
+ type: "text" | "tool_call" | "tool_result" | "jev_decision"
262
275
  message: string
263
276
  toolName?: string
264
277
  isError?: boolean
278
+ /** Present on `jev_decision`: what Jev decided for this run, its cost and the estimated savings. */
279
+ jev?: JevStepDecision
265
280
  }
266
281
 
282
+ /** One Jev decision as the host receives it through `onStep`. */
283
+ export type JevStepDecision = CanvasJevDecision
284
+
267
285
  // ─── Stream chunk types (compatible with providers/index.ts) ─────────────────
268
286
 
269
287
  export interface StreamChunk {
@@ -357,12 +375,20 @@ export async function* runAgent(
357
375
  channel: opts.channel,
358
376
  source: opts.historySource ?? "message",
359
377
  })
360
- // Run compaction if conversation history is getting large
378
+ // Run compaction if conversation history is getting large.
379
+ // El modelo del turno viaja con sus credenciales: el resumen es una llamada
380
+ // al modelo como cualquier otra y tiene que cobrarse a la misma cuenta.
361
381
  await maybeCompact(
362
382
  opts.threadId,
363
383
  opts.channel && opts.userId
364
384
  ? { channel: opts.channel, userId: opts.userId }
365
- : undefined
385
+ : undefined,
386
+ {
387
+ provider: providerCfg.provider,
388
+ model: providerCfg.model,
389
+ credentials: opts.credentials,
390
+ contextWindow: providerCfg.contextWindow,
391
+ }
366
392
  )
367
393
  }
368
394
 
@@ -377,8 +403,25 @@ export async function* runAgent(
377
403
  taskContext: opts.taskContext,
378
404
  userId: opts.userId,
379
405
  causalStreamId,
406
+ skipJev: !!opts.resume,
407
+ jev: opts.jev,
380
408
  })
381
409
 
410
+ // Every decision goes to the canvas (hosts like hive) and to onStep (hosts
411
+ // that drive runAgent themselves, like hive-cloud).
412
+ const publishJev = async (decision: Parameters<typeof emitJevDecision>[0]): Promise<void> => {
413
+ const event = emitJevDecision(decision)
414
+ if (!opts.onStep) return
415
+ try {
416
+ await opts.onStep({ type: "jev_decision", message: event.summary, jev: event })
417
+ } catch (err) {
418
+ log.warn(`[agent-loop] onStep(jev_decision) failed: ${(err as Error).message}`)
419
+ }
420
+ }
421
+ if (ctx.jevDecision) {
422
+ await publishJev({ ...ctx.jevDecision, agentId: opts.agentId, kind: "context", provider: providerCfg.provider, model: providerCfg.model })
423
+ }
424
+
382
425
  // Force extra tools into the loadout (tests/evals)
383
426
  if (opts.extraTools?.length) {
384
427
  const existingNames = new Set(ctx.tools.map((t: any) => t.function?.name))
@@ -409,9 +452,12 @@ export async function* runAgent(
409
452
  if (opts.isolated) {
410
453
  messages.push({ role: "user", content: opts.userMessage })
411
454
  }
455
+ const jevObjective = typeof opts.userMessage === "string" ? opts.userMessage :
456
+ opts.userMessage.filter((part) => part.type === "text").map((part) => (part as { text: string }).text).join("\n")
412
457
 
413
458
  // ── Resume from checkpoint ─────────────────────────────────────────────────
414
- let injectedToolNames: string[] = []
459
+ // Seeded with the compiled loadout so a checkpoint records the tools Jev chose.
460
+ let injectedToolNames: string[] = ctx.tools.map(t => t.function.name).filter(name => !MINIMAL_TOOLS.has(name))
415
461
  let systemPromptSkillSections: string[] = []
416
462
  let resumedFromPending = false
417
463
  let iterations = 0
@@ -432,6 +478,15 @@ export async function* runAgent(
432
478
  if (restored) {
433
479
  messages = restored.messages
434
480
  injectedToolNames = restored.injectedToolNames ?? []
481
+ // A resume skips Jev: restore the loadout the checkpoint recorded.
482
+ const currentTools = new Set(ctx.tools.map(t => t.function.name))
483
+ for (const name of injectedToolNames) {
484
+ const tool = ctx.allTools.find(t => t.name === name)
485
+ if (tool && !currentTools.has(name)) {
486
+ ctx.tools.push({ type: "function", function: { name: tool.name, description: tool.description, parameters: tool.parameters } })
487
+ currentTools.add(name)
488
+ }
489
+ }
435
490
  systemPromptSkillSections = restored.systemPromptSkillSections ?? []
436
491
  iterations = restored.iterations ?? 0
437
492
  totalInputTokens = restored.totalInputTokens ?? 0
@@ -517,11 +572,27 @@ export async function* runAgent(
517
572
  : null
518
573
  let streamedThisCall = false
519
574
  let response: Awaited<ReturnType<typeof callLLM>>
575
+ const jevIteration = await planJevIteration({ objective: jevObjective, messages, tools: ctx.tools, jev: opts.jev })
576
+ .catch((err) => { log.warn(`[agent-loop] Jev iteration fallback: ${(err as Error).message}`); return null })
577
+ const callMessages = jevIteration?.messages ?? messages
578
+ const callTools = jevIteration?.tools ?? ctx.tools
579
+ if (jevIteration) {
580
+ log.info(`[agent-loop] Jev action=${jevIteration.action} omitted_results=${jevIteration.omittedResults} tools=${callTools.map(t => t.function.name).join(",")}`)
581
+ // Measured on what the provider actually receives, after the usual truncation.
582
+ const payloadChars = (msgs: LLMMessage[], tools: typeof ctx.tools) =>
583
+ JSON.stringify(clearOldToolResults(msgs)).length + JSON.stringify(tools).length
584
+ await publishJev({
585
+ agentId: opts.agentId, kind: "iteration", provider: providerCfg.provider, model: providerCfg.model,
586
+ summary: `${JEV_ACTION_LABELS[jevIteration.action] ?? jevIteration.action} · ${jevIteration.omittedResults} resultado(s) omitido(s) · ${callTools.length}/${ctx.tools.length} herramientas`,
587
+ savedTokens: Math.round((payloadChars(messages, ctx.tools) - payloadChars(callMessages, callTools)) / 4),
588
+ latencyMs: jevIteration.decision.latencyMs, costUsd: jevIteration.decision.costUsd,
589
+ })
590
+ }
520
591
  try {
521
592
  response = await withTimeout(() => callLLM({
522
593
  ...providerCfg,
523
- messages: clearOldToolResults(messages) as LLMMessage[],
524
- tools: ctx.tools.length > 0 ? ctx.tools : undefined,
594
+ messages: clearOldToolResults(callMessages) as LLMMessage[],
595
+ tools: callTools.length > 0 ? callTools : undefined,
525
596
  signal: opts.signal,
526
597
  sessionId: opts.threadId,
527
598
  onToken: opts.onToken && !delegationGroupAtCall
@@ -681,6 +752,16 @@ export async function* runAgent(
681
752
  }
682
753
  }
683
754
 
755
+ const jevParallel = await jevWantsParallel(response.tool_calls, opts.jev)
756
+ .catch((err) => { log.warn(`[agent-loop] Jev parallel fallback: ${(err as Error).message}`); return null })
757
+ if (jevParallel?.decision) {
758
+ log.info(`[agent-loop] Jev parallel=${jevParallel.parallel} calls=${response.tool_calls.length}`)
759
+ await publishJev({
760
+ agentId: opts.agentId, kind: "parallel", provider: providerCfg.provider, model: providerCfg.model,
761
+ summary: `${response.tool_calls.length} herramientas ${jevParallel.parallel ? "en paralelo" : "en secuencia"}`,
762
+ savedTokens: 0, latencyMs: jevParallel.decision.latencyMs, costUsd: jevParallel.decision.costUsd,
763
+ })
764
+ }
684
765
  const toolResults = await executeToolBatch({
685
766
  toolCalls: response.tool_calls,
686
767
  allTools: ctx.allTools,
@@ -701,6 +782,7 @@ export async function* runAgent(
701
782
  },
702
783
  hiveConfig,
703
784
  workerPool: hiveConfig.tools?.workerPool,
785
+ parallelToolCalls: jevParallel?.parallel,
704
786
  signal: opts.signal,
705
787
  })
706
788
 
@@ -1297,6 +1379,8 @@ export interface IsolatedAgentOptions {
1297
1379
  * reabría justo en el camino de delegación.
1298
1380
  */
1299
1381
  credentials?: ProviderCredentials
1382
+ /** Jev for the worker, inherited from the delegating turn like `credentials`. */
1383
+ jev?: JevOption
1300
1384
  }
1301
1385
 
1302
1386
  export async function runAgentIsolatedDetailed(
@@ -1321,6 +1405,7 @@ export async function runAgentIsolatedDetailed(
1321
1405
  channel: opts.channel,
1322
1406
  sessionId: opts.sessionId,
1323
1407
  credentials: opts.credentials,
1408
+ jev: opts.jev,
1324
1409
  })) {
1325
1410
  if (chunk.agent?.messages?.[0]?.content) {
1326
1411
  lastContent = chunk.agent.messages[0].content
@@ -26,7 +26,10 @@ import {
26
26
  type StoredMessage,
27
27
  } from "./conversation-store.ts"
28
28
  import { estimateTokens } from "../utils/toon.ts"
29
- import { callLLM, resolveProviderConfig, getDefaultLLM, type ContentPart } from "./llm-client.ts"
29
+ import {
30
+ callLLM, resolveProviderConfig, getDefaultLLM,
31
+ type ContentPart, type ProviderCredentials,
32
+ } from "./llm-client.ts"
30
33
  import { col, fromIndexable } from "../storage/hive.ts"
31
34
  import type { AgentDoc, ModelDoc } from "../storage/collections.ts"
32
35
  import { loadConfig } from "../config/loader.ts"
@@ -41,6 +44,61 @@ const KEEP_LAST_N_MESSAGES = 5 // always keep most recent N messages
41
44
  const TOOL_RESULT_MAX_CHARS = 200 // max chars for old tool results after clearing
42
45
  const MAX_TRANSCRIPT_MSGS = 30 // cap messages sent to summarizer (avoids OOM on small models)
43
46
  const MAX_MSG_CHARS = 300 // chars per message in transcript
47
+ /** Ventana asumida cuando no se conoce la del modelo: `COMPACT_TOKEN_THRESHOLD` es su 25 %. */
48
+ const ASSUMED_CONTEXT_WINDOW = 128_000
49
+ const DEFAULT_CONTEXT_RATIO = 0.25
50
+
51
+ /**
52
+ * El modelo con el que se pide el resumen: el del turno que disparó la
53
+ * compactación, con SUS credenciales.
54
+ */
55
+ export interface CompactionLLM {
56
+ provider?: string
57
+ model?: string
58
+ /** En multi-inquilino, la llave del cliente. Sin esto se usaría una global. */
59
+ credentials?: ProviderCredentials
60
+ /** La ventana del modelo, si quien llama ya la resolvió. */
61
+ contextWindow?: number
62
+ }
63
+
64
+ /**
65
+ * A partir de cuántos tokens de historial se compacta.
66
+ *
67
+ * `agent.context.compactionThreshold` es una PROPORCIÓN de la ventana del
68
+ * modelo —su valor por defecto es 0.8, o sea el 80 %—, pero se leía como si
69
+ * fueran tokens. Con la configuración por defecto el umbral quedaba en 0.8
70
+ * tokens: cualquier hilo con más de cinco mensajes se resumía en cada turno,
71
+ * pagando una llamada extra al modelo y reemplazando el historial por un
72
+ * resumen desde el primer intercambio. Un valor mayor que 1 se sigue leyendo
73
+ * como tokens, que es lo que espera quien fijó un número absoluto.
74
+ */
75
+ export function resolveCompactionThreshold(configured: number | undefined, contextWindow?: number): number {
76
+ const known = contextWindow && contextWindow > 0 ? contextWindow : undefined
77
+ if (typeof configured === "number" && Number.isFinite(configured) && configured > 0) {
78
+ return configured <= 1
79
+ ? Math.floor((known ?? ASSUMED_CONTEXT_WINDOW) * configured)
80
+ : Math.floor(configured)
81
+ }
82
+ return known ? Math.floor(known * DEFAULT_CONTEXT_RATIO) : COMPACT_TOKEN_THRESHOLD
83
+ }
84
+
85
+ /** La ventana del modelo del turno; si no se sabe cuál es, la del coordinador. */
86
+ async function modelContextWindow(modelId?: string): Promise<number | undefined> {
87
+ try {
88
+ const modelsCol = await col<ModelDoc>("models")
89
+ if (modelId) return (await modelsCol.get(modelId))?.doc.context_window || undefined
90
+ const agentsCol = await col<AgentDoc>("agents")
91
+ const coordinators = await agentsCol.findBy("role", "coordinator", { limit: 1 })
92
+ // El id se busca completo: recortar el primer segmento rompía cualquier
93
+ // modelo cuyo nombre lleve barra (meta/llama-3.3-70b-instruct buscaba
94
+ // "llama-3.3-70b-instruct", no encontraba nada y caía al default).
95
+ const id = fromIndexable(coordinators[0]?.doc.model_id ?? null)
96
+ if (!id) return undefined
97
+ return (await modelsCol.get(id))?.doc.context_window || undefined
98
+ } catch {
99
+ return undefined
100
+ }
101
+ }
44
102
 
45
103
  /**
46
104
  * Check if compaction is needed and run it if so.
@@ -48,37 +106,16 @@ const MAX_MSG_CHARS = 300 // chars per message in transcript
48
106
  */
49
107
  export async function maybeCompact(
50
108
  threadId: string,
51
- notify?: { channel: string; userId: string }
109
+ notify?: { channel: string; userId: string },
110
+ llm?: CompactionLLM
52
111
  ): Promise<void> {
53
112
  try {
54
113
  const totalTokens = await getTotalTokens(threadId)
55
-
56
- // Orden de precedencia: lo que el usuario configuró gana sobre lo que se
57
- // deduce del modelo, y eso gana sobre la constante.
58
- //
59
- // `agent.context.compactionThreshold` estaba en el esquema de configuración
60
- // y **no lo leía nadie**: alguien podía ajustarlo y no pasaba nada. Una
61
- // opción que no hace nada es peor que no tenerla, porque el usuario cree
62
- // que cambió algo.
63
- let effectiveThreshold = COMPACT_TOKEN_THRESHOLD
64
- const configurado = loadConfig().agent?.context?.compactionThreshold
65
- try {
66
- const agentsCol = await col<AgentDoc>("agents")
67
- const coordinators = await agentsCol.findBy("role", "coordinator", { limit: 1 })
68
- const modelId = fromIndexable(coordinators[0]?.doc.model_id ?? null)
69
- if (modelId) {
70
- const modelsCol = await col<ModelDoc>("models")
71
- // El id se busca completo: recortar el primer segmento rompía cualquier
72
- // modelo cuyo nombre lleve barra (meta/llama-3.3-70b-instruct buscaba
73
- // "llama-3.3-70b-instruct", no encontraba nada y caía al default).
74
- const modelEntry = await modelsCol.get(modelId)
75
- if (modelEntry?.doc.context_window) {
76
- effectiveThreshold = Math.floor(modelEntry.doc.context_window * 0.25)
77
- }
78
- }
79
- } catch { /* use default threshold */ }
80
-
81
- if (configurado && configurado > 0) effectiveThreshold = configurado
114
+ const contextWindow = llm?.contextWindow ?? (await modelContextWindow(llm?.model))
115
+ const effectiveThreshold = resolveCompactionThreshold(
116
+ loadConfig().agent?.context?.compactionThreshold,
117
+ contextWindow,
118
+ )
82
119
 
83
120
  if (totalTokens < effectiveThreshold) return
84
121
 
@@ -96,8 +133,8 @@ export async function maybeCompact(
96
133
  // Already summarized up to near the current state
97
134
  if (summary && summary.last_message_id > totalMessages - KEEP_LAST_N_MESSAGES) return
98
135
 
99
- log.info(`[compaction] Compacting thread=${threadId} tokens=${totalTokens}`)
100
- await compactThread(threadId, notify)
136
+ log.info(`[compaction] Compacting thread=${threadId} tokens=${totalTokens} threshold=${effectiveThreshold}`)
137
+ await compactThread(threadId, notify, llm)
101
138
  } catch (err) {
102
139
  log.warn("[compaction] Error during compaction check:", err)
103
140
  }
@@ -139,7 +176,8 @@ export function renderTranscript(rows: StoredMessage[], maxMsgChars = MAX_MSG_CH
139
176
  */
140
177
  export async function compactThread(
141
178
  threadId: string,
142
- notify?: { channel: string; userId: string }
179
+ notify?: { channel: string; userId: string },
180
+ llm?: CompactionLLM
143
181
  ): Promise<void> {
144
182
  const allMessages = await getHistory(threadId)
145
183
  if (allMessages.length <= KEEP_LAST_N_MESSAGES) return
@@ -162,10 +200,18 @@ export async function compactThread(
162
200
  const capped = toSummarize.slice(-MAX_TRANSCRIPT_MSGS)
163
201
  const transcript = renderTranscript(capped)
164
202
 
165
- const defaultLLM = await getDefaultLLM()
166
- if (!defaultLLM) throw new Error("No active LLM providers/models configured in the database")
203
+ // El modelo del turno y SUS credenciales. Antes el resumen se pedía siempre
204
+ // con `getDefaultLLM()` y sin credenciales, así que `resolveProviderConfig`
205
+ // caía al secret store, al llavero del sistema o al entorno: en una
206
+ // instalación multi-inquilino eso resume la conversación de un cliente con
207
+ // la llave de la plataforma (o de otro cliente). Sin `llm` se comporta como
208
+ // antes, que es lo que necesita una instalación de un solo usuario.
209
+ const target = llm?.provider && llm.model
210
+ ? { provider: llm.provider, model: llm.model }
211
+ : await getDefaultLLM()
212
+ if (!target) throw new Error("No active LLM providers/models configured in the database")
167
213
 
168
- const providerCfg = await resolveProviderConfig(defaultLLM.provider, defaultLLM.model)
214
+ const providerCfg = await resolveProviderConfig(target.provider, target.model, llm?.credentials)
169
215
 
170
216
  const summaryResponse = await callLLM({
171
217
  ...providerCfg,