@johpaz/hive-sdk 0.4.8 → 0.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +75 -1
- package/README.md +1 -1
- package/package.json +1 -1
- package/packages/core/src/agent/agent-loop.ts +92 -7
- package/packages/core/src/agent/compaction.ts +81 -35
- package/packages/core/src/agent/context-compiler.ts +122 -23
- package/packages/core/src/agent/index.ts +2 -0
- package/packages/core/src/agent/jev-decisions.ts +199 -0
- package/packages/core/src/agent/jev-planner.ts +298 -0
- package/packages/core/src/agent/tool-selector.ts +1 -0
- package/packages/core/src/canvas/emitter.ts +19 -0
- package/packages/core/src/events/channel-narration.ts +3 -0
- package/packages/core/src/services/swarms.ts +4 -0
- package/packages/core/src/storage/collections.ts +9 -2
- package/packages/core/src/storage/crypto.ts +15 -7
- package/packages/core/src/storage/seed.ts +4 -0
- package/packages/core/src/storage/usage.ts +55 -2
- package/packages/core/src/swarm/RoleSwarm.ts +6 -0
- package/packages/core/src/tool-runtime/index.ts +20 -1
- package/packages/core/src/tools/agents/get-available-models.ts +3 -1
- package/packages/core/src/tools/core/index.ts +33 -2
package/CHANGELOG.md
CHANGED
|
@@ -1,6 +1,57 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
-
##
|
|
3
|
+
## 0.5.0
|
|
4
|
+
|
|
5
|
+
### Jev — plano de decisión (OpenRouter Decisions)
|
|
6
|
+
|
|
7
|
+
Portado de hive 1.1.0. Jev decide por turno qué historial, herramientas,
|
|
8
|
+
skills, notas y reglas del playbook entran al contexto; entre iteraciones poda
|
|
9
|
+
resultados viejos de herramientas y sugiere la siguiente acción; decide si un
|
|
10
|
+
lote de herramientas corre en paralelo, y conoce el mapa del enjambre
|
|
11
|
+
(especialistas y estado de cada MCP) para recomendar a quién delegar.
|
|
12
|
+
**Sin clave de OpenRouter, Jev no existe y todo corre igual que antes.**
|
|
13
|
+
|
|
14
|
+
- **Clave inyectable por llamada**: `jev?: { apiKey, mcpSettingsPath? } | false`
|
|
15
|
+
en `AgentLoopOptions`, `compileContext`, `IsolatedAgentOptions`,
|
|
16
|
+
`runRoleSwarm` y `runSwarm`, igual que `credentials`. `false` lo apaga;
|
|
17
|
+
sin la opción decide la fila `openrouter` del inquilino actual. Con un
|
|
18
|
+
inquilino activo **nunca** se usa `OPENROUTER_API_KEY` ni la caché de
|
|
19
|
+
secretos del proceso: la clave de la plataforma no se usa en nombre de un
|
|
20
|
+
cliente.
|
|
21
|
+
- **Estado por inquilino**: fallos, cooldown y totales se llevan por
|
|
22
|
+
`currentTenant()`; una clave inválida de un cliente no pone en fallback a
|
|
23
|
+
los demás.
|
|
24
|
+
- **Evento para el host**: cada decisión llega por `onStep` como
|
|
25
|
+
`StepEvent` `jev_decision` (`jev`: agente, tipo, resumen, tokens ahorrados,
|
|
26
|
+
latencia, costo, especialista recomendado, MCP apagados). También se emite
|
|
27
|
+
`canvas:jev_decision` / `canvas:jev_status` para hosts tipo hive.
|
|
28
|
+
- **Uso y costo**: `recordJevDecision` y los campos `jev*` de
|
|
29
|
+
`UsageRollupDoc`; `getUsageStats()` devuelve `jev` con el total y el
|
|
30
|
+
desglose por agente. El ahorro se estima (caracteres/4) y se cotiza con el
|
|
31
|
+
modelo del agente asesorado.
|
|
32
|
+
- **Catálogo**: modelo `openrouter/typesafe/jev-1.13` con `modelType:
|
|
33
|
+
"decision"`, excluido de `get_available_models` (y de `getDefaultLLM`, que
|
|
34
|
+
sólo toma modelos `llm`).
|
|
35
|
+
- Herramienta nueva `conversation_read`: recupera mensajes o notas que Jev
|
|
36
|
+
dejó fuera del contexto, siempre dentro del hilo actual.
|
|
37
|
+
- `executeToolBatch` acepta `parallelToolCalls` por lote.
|
|
38
|
+
- `NarrationEventDoc.kind` suma `"decision"`: un host puede anotar las
|
|
39
|
+
decisiones de Jev en `narrationEvents` para sus vistas de actividad.
|
|
40
|
+
`shouldDeliverToChannel` nunca lo entrega a un canal.
|
|
41
|
+
- `loadDurableProviderApiKey(id)`: lee la clave sólo de la colección
|
|
42
|
+
`secrets` (particionada por inquilino), sin pasar por la caché de proceso ni
|
|
43
|
+
el llavero del SO.
|
|
44
|
+
|
|
45
|
+
**Privacidad.** Con Jev activo se envían a la API de decisiones de OpenRouter
|
|
46
|
+
(`https://openrouter.ai/api/alpha/decisions`), con la clave del workspace:
|
|
47
|
+
el objetivo del turno (hasta 3 500 caracteres), extractos de hasta 450
|
|
48
|
+
caracteres de mensajes previos del hilo, nombre y descripción de herramientas
|
|
49
|
+
y skills candidatas, notas del scratchpad y reglas del playbook (hasta 350
|
|
50
|
+
caracteres cada una), extractos de resultados de herramientas (hasta 650
|
|
51
|
+
caracteres), argumentos de llamadas en lote (hasta 700) y el mapa del enjambre
|
|
52
|
+
(ids, nombres y descripciones de especialistas, nombres y estado de los MCP).
|
|
53
|
+
No se envían ids de MCP con prefijo de inquilino, credenciales ni adjuntos
|
|
54
|
+
binarios. Un host que no quiera enviar nada pasa `jev: false`.
|
|
4
55
|
|
|
5
56
|
### Plataforma y seguridad
|
|
6
57
|
|
|
@@ -579,6 +630,29 @@
|
|
|
579
630
|
job que genera un scaffold con `create-app` y lo typechequea contra el SDK de
|
|
580
631
|
ese commit.
|
|
581
632
|
|
|
633
|
+
## 0.4.9
|
|
634
|
+
|
|
635
|
+
### Corregido
|
|
636
|
+
|
|
637
|
+
- **La compactación de historial dejaba de ser excepcional y corría casi en cada
|
|
638
|
+
turno.** `agent.context.compactionThreshold` es una proporción de la ventana
|
|
639
|
+
del modelo —su valor por defecto es `0.8`, o sea el 80 %— pero se leía como un
|
|
640
|
+
número de tokens: el umbral efectivo quedaba en 0.8 tokens. Cualquier hilo con
|
|
641
|
+
más de cinco mensajes se resumía en cada turno, lo que cuesta una llamada extra
|
|
642
|
+
al modelo por mensaje y reemplaza el historial por un resumen desde el primer
|
|
643
|
+
intercambio. Ahora un valor menor o igual a 1 se aplica sobre la ventana del
|
|
644
|
+
modelo y uno mayor se sigue leyendo como tokens, para quien fijó un número
|
|
645
|
+
absoluto. El umbral se calcula además con la ventana del modelo que corre el
|
|
646
|
+
turno, no con la del coordinador.
|
|
647
|
+
- **El resumen se pedía con una credencial global.** `compactThread` resolvía el
|
|
648
|
+
modelo con `getDefaultLLM()` y llamaba a `resolveProviderConfig` sin
|
|
649
|
+
credenciales, así que la llave salía del secret store, del llavero del sistema
|
|
650
|
+
o del entorno del proceso. En una instalación de un solo usuario da igual; en
|
|
651
|
+
una multi-inquilino significa resumir la conversación de un cliente con la
|
|
652
|
+
llave de la plataforma o de otro cliente. `maybeCompact` y `compactThread`
|
|
653
|
+
aceptan ahora el modelo y las credenciales del turno (`CompactionLLM`), y el
|
|
654
|
+
agent loop les pasa los suyos. Sin ese dato se comportan como antes.
|
|
655
|
+
|
|
582
656
|
## 0.1.5
|
|
583
657
|
|
|
584
658
|
Sincronización del SDK con el runtime de agentes de `hive`. **Trae rupturas de
|
package/README.md
CHANGED
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@johpaz/hive-sdk",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.5.0",
|
|
4
4
|
"private": false,
|
|
5
5
|
"description": "Hive SDK — The Agent Harness SDK. Build, deploy, and scale AI agent applications with multi-channel support, context engineering, and swarm orchestration.",
|
|
6
6
|
"license": "MIT",
|
|
@@ -23,9 +23,12 @@ import { callLLM, resolveProviderConfig, getDefaultLLM, type LLMMessage, type Pr
|
|
|
23
23
|
import { addMessage } from "./conversation-store.ts"
|
|
24
24
|
import { saveTrace, recordLLMUsage } from "./tracer.ts"
|
|
25
25
|
import { maybeCompact, clearOldToolResults } from "./compaction.ts"
|
|
26
|
-
import { emitCanvas } from "../canvas/emitter.ts"
|
|
26
|
+
import { emitCanvas, type CanvasJevDecision } from "../canvas/emitter.ts"
|
|
27
27
|
import type { MCPClientManager } from "../mcp/index.ts"
|
|
28
28
|
import { compileContext } from "./context-compiler.ts"
|
|
29
|
+
import { MINIMAL_TOOLS } from "./minimal-loadout.ts"
|
|
30
|
+
import { jevWantsParallel, planJevIteration } from "./jev-planner.ts"
|
|
31
|
+
import { emitJevDecision, type JevOption } from "./jev-decisions.ts"
|
|
29
32
|
import { formatToolResult } from "../utils/toon.ts"
|
|
30
33
|
import { redactBinaryStrings } from "../utils/redact-binary.ts"
|
|
31
34
|
import { resolveUserId, resolveAgentId } from "../storage/onboarding.ts"
|
|
@@ -52,6 +55,10 @@ import { getNarration } from "../events/tool-narration.ts"
|
|
|
52
55
|
|
|
53
56
|
const log = logger.child("agent-loop")
|
|
54
57
|
|
|
58
|
+
const JEV_ACTION_LABELS: Record<string, string> = {
|
|
59
|
+
continue: "Continuar", delegate: "Delegar", discover: "Descubrir", finish: "Cerrar",
|
|
60
|
+
}
|
|
61
|
+
|
|
55
62
|
// Per-operation budget for a single LLM call — NOT an aggregate deadline for the
|
|
56
63
|
// whole turn. Each call gets its own fresh window; a slow-but-healthy multi-step
|
|
57
64
|
// turn (many quick operations) is never killed just for taking a while overall.
|
|
@@ -226,6 +233,12 @@ export interface AgentLoopOptions {
|
|
|
226
233
|
* dos inquilinos concurrentes en el mismo proceso compartían credencial.
|
|
227
234
|
*/
|
|
228
235
|
credentials?: ProviderCredentials
|
|
236
|
+
/**
|
|
237
|
+
* Jev (OpenRouter Decisions) for this run. `{ apiKey }` uses that key,
|
|
238
|
+
* `false` turns it off, undefined reads the current tenant's `openrouter`
|
|
239
|
+
* provider row. Travels with the run like `credentials`.
|
|
240
|
+
*/
|
|
241
|
+
jev?: JevOption
|
|
229
242
|
/** Whether to resume from a previously saved checkpoint */
|
|
230
243
|
resume?: boolean
|
|
231
244
|
/** Run budget — overrides agent.max_iterations when set */
|
|
@@ -258,12 +271,17 @@ export interface AgentLoopOptions {
|
|
|
258
271
|
export type { StepEvent as AgentStepEvent }
|
|
259
272
|
|
|
260
273
|
export interface StepEvent {
|
|
261
|
-
type: "text" | "tool_call" | "tool_result"
|
|
274
|
+
type: "text" | "tool_call" | "tool_result" | "jev_decision"
|
|
262
275
|
message: string
|
|
263
276
|
toolName?: string
|
|
264
277
|
isError?: boolean
|
|
278
|
+
/** Present on `jev_decision`: what Jev decided for this run, its cost and the estimated savings. */
|
|
279
|
+
jev?: JevStepDecision
|
|
265
280
|
}
|
|
266
281
|
|
|
282
|
+
/** One Jev decision as the host receives it through `onStep`. */
|
|
283
|
+
export type JevStepDecision = CanvasJevDecision
|
|
284
|
+
|
|
267
285
|
// ─── Stream chunk types (compatible with providers/index.ts) ─────────────────
|
|
268
286
|
|
|
269
287
|
export interface StreamChunk {
|
|
@@ -357,12 +375,20 @@ export async function* runAgent(
|
|
|
357
375
|
channel: opts.channel,
|
|
358
376
|
source: opts.historySource ?? "message",
|
|
359
377
|
})
|
|
360
|
-
// Run compaction if conversation history is getting large
|
|
378
|
+
// Run compaction if conversation history is getting large.
|
|
379
|
+
// El modelo del turno viaja con sus credenciales: el resumen es una llamada
|
|
380
|
+
// al modelo como cualquier otra y tiene que cobrarse a la misma cuenta.
|
|
361
381
|
await maybeCompact(
|
|
362
382
|
opts.threadId,
|
|
363
383
|
opts.channel && opts.userId
|
|
364
384
|
? { channel: opts.channel, userId: opts.userId }
|
|
365
|
-
: undefined
|
|
385
|
+
: undefined,
|
|
386
|
+
{
|
|
387
|
+
provider: providerCfg.provider,
|
|
388
|
+
model: providerCfg.model,
|
|
389
|
+
credentials: opts.credentials,
|
|
390
|
+
contextWindow: providerCfg.contextWindow,
|
|
391
|
+
}
|
|
366
392
|
)
|
|
367
393
|
}
|
|
368
394
|
|
|
@@ -377,8 +403,25 @@ export async function* runAgent(
|
|
|
377
403
|
taskContext: opts.taskContext,
|
|
378
404
|
userId: opts.userId,
|
|
379
405
|
causalStreamId,
|
|
406
|
+
skipJev: !!opts.resume,
|
|
407
|
+
jev: opts.jev,
|
|
380
408
|
})
|
|
381
409
|
|
|
410
|
+
// Every decision goes to the canvas (hosts like hive) and to onStep (hosts
|
|
411
|
+
// that drive runAgent themselves, like hive-cloud).
|
|
412
|
+
const publishJev = async (decision: Parameters<typeof emitJevDecision>[0]): Promise<void> => {
|
|
413
|
+
const event = emitJevDecision(decision)
|
|
414
|
+
if (!opts.onStep) return
|
|
415
|
+
try {
|
|
416
|
+
await opts.onStep({ type: "jev_decision", message: event.summary, jev: event })
|
|
417
|
+
} catch (err) {
|
|
418
|
+
log.warn(`[agent-loop] onStep(jev_decision) failed: ${(err as Error).message}`)
|
|
419
|
+
}
|
|
420
|
+
}
|
|
421
|
+
if (ctx.jevDecision) {
|
|
422
|
+
await publishJev({ ...ctx.jevDecision, agentId: opts.agentId, kind: "context", provider: providerCfg.provider, model: providerCfg.model })
|
|
423
|
+
}
|
|
424
|
+
|
|
382
425
|
// Force extra tools into the loadout (tests/evals)
|
|
383
426
|
if (opts.extraTools?.length) {
|
|
384
427
|
const existingNames = new Set(ctx.tools.map((t: any) => t.function?.name))
|
|
@@ -409,9 +452,12 @@ export async function* runAgent(
|
|
|
409
452
|
if (opts.isolated) {
|
|
410
453
|
messages.push({ role: "user", content: opts.userMessage })
|
|
411
454
|
}
|
|
455
|
+
const jevObjective = typeof opts.userMessage === "string" ? opts.userMessage :
|
|
456
|
+
opts.userMessage.filter((part) => part.type === "text").map((part) => (part as { text: string }).text).join("\n")
|
|
412
457
|
|
|
413
458
|
// ── Resume from checkpoint ─────────────────────────────────────────────────
|
|
414
|
-
|
|
459
|
+
// Seeded with the compiled loadout so a checkpoint records the tools Jev chose.
|
|
460
|
+
let injectedToolNames: string[] = ctx.tools.map(t => t.function.name).filter(name => !MINIMAL_TOOLS.has(name))
|
|
415
461
|
let systemPromptSkillSections: string[] = []
|
|
416
462
|
let resumedFromPending = false
|
|
417
463
|
let iterations = 0
|
|
@@ -432,6 +478,15 @@ export async function* runAgent(
|
|
|
432
478
|
if (restored) {
|
|
433
479
|
messages = restored.messages
|
|
434
480
|
injectedToolNames = restored.injectedToolNames ?? []
|
|
481
|
+
// A resume skips Jev: restore the loadout the checkpoint recorded.
|
|
482
|
+
const currentTools = new Set(ctx.tools.map(t => t.function.name))
|
|
483
|
+
for (const name of injectedToolNames) {
|
|
484
|
+
const tool = ctx.allTools.find(t => t.name === name)
|
|
485
|
+
if (tool && !currentTools.has(name)) {
|
|
486
|
+
ctx.tools.push({ type: "function", function: { name: tool.name, description: tool.description, parameters: tool.parameters } })
|
|
487
|
+
currentTools.add(name)
|
|
488
|
+
}
|
|
489
|
+
}
|
|
435
490
|
systemPromptSkillSections = restored.systemPromptSkillSections ?? []
|
|
436
491
|
iterations = restored.iterations ?? 0
|
|
437
492
|
totalInputTokens = restored.totalInputTokens ?? 0
|
|
@@ -517,11 +572,27 @@ export async function* runAgent(
|
|
|
517
572
|
: null
|
|
518
573
|
let streamedThisCall = false
|
|
519
574
|
let response: Awaited<ReturnType<typeof callLLM>>
|
|
575
|
+
const jevIteration = await planJevIteration({ objective: jevObjective, messages, tools: ctx.tools, jev: opts.jev })
|
|
576
|
+
.catch((err) => { log.warn(`[agent-loop] Jev iteration fallback: ${(err as Error).message}`); return null })
|
|
577
|
+
const callMessages = jevIteration?.messages ?? messages
|
|
578
|
+
const callTools = jevIteration?.tools ?? ctx.tools
|
|
579
|
+
if (jevIteration) {
|
|
580
|
+
log.info(`[agent-loop] Jev action=${jevIteration.action} omitted_results=${jevIteration.omittedResults} tools=${callTools.map(t => t.function.name).join(",")}`)
|
|
581
|
+
// Measured on what the provider actually receives, after the usual truncation.
|
|
582
|
+
const payloadChars = (msgs: LLMMessage[], tools: typeof ctx.tools) =>
|
|
583
|
+
JSON.stringify(clearOldToolResults(msgs)).length + JSON.stringify(tools).length
|
|
584
|
+
await publishJev({
|
|
585
|
+
agentId: opts.agentId, kind: "iteration", provider: providerCfg.provider, model: providerCfg.model,
|
|
586
|
+
summary: `${JEV_ACTION_LABELS[jevIteration.action] ?? jevIteration.action} · ${jevIteration.omittedResults} resultado(s) omitido(s) · ${callTools.length}/${ctx.tools.length} herramientas`,
|
|
587
|
+
savedTokens: Math.round((payloadChars(messages, ctx.tools) - payloadChars(callMessages, callTools)) / 4),
|
|
588
|
+
latencyMs: jevIteration.decision.latencyMs, costUsd: jevIteration.decision.costUsd,
|
|
589
|
+
})
|
|
590
|
+
}
|
|
520
591
|
try {
|
|
521
592
|
response = await withTimeout(() => callLLM({
|
|
522
593
|
...providerCfg,
|
|
523
|
-
messages: clearOldToolResults(
|
|
524
|
-
tools:
|
|
594
|
+
messages: clearOldToolResults(callMessages) as LLMMessage[],
|
|
595
|
+
tools: callTools.length > 0 ? callTools : undefined,
|
|
525
596
|
signal: opts.signal,
|
|
526
597
|
sessionId: opts.threadId,
|
|
527
598
|
onToken: opts.onToken && !delegationGroupAtCall
|
|
@@ -681,6 +752,16 @@ export async function* runAgent(
|
|
|
681
752
|
}
|
|
682
753
|
}
|
|
683
754
|
|
|
755
|
+
const jevParallel = await jevWantsParallel(response.tool_calls, opts.jev)
|
|
756
|
+
.catch((err) => { log.warn(`[agent-loop] Jev parallel fallback: ${(err as Error).message}`); return null })
|
|
757
|
+
if (jevParallel?.decision) {
|
|
758
|
+
log.info(`[agent-loop] Jev parallel=${jevParallel.parallel} calls=${response.tool_calls.length}`)
|
|
759
|
+
await publishJev({
|
|
760
|
+
agentId: opts.agentId, kind: "parallel", provider: providerCfg.provider, model: providerCfg.model,
|
|
761
|
+
summary: `${response.tool_calls.length} herramientas ${jevParallel.parallel ? "en paralelo" : "en secuencia"}`,
|
|
762
|
+
savedTokens: 0, latencyMs: jevParallel.decision.latencyMs, costUsd: jevParallel.decision.costUsd,
|
|
763
|
+
})
|
|
764
|
+
}
|
|
684
765
|
const toolResults = await executeToolBatch({
|
|
685
766
|
toolCalls: response.tool_calls,
|
|
686
767
|
allTools: ctx.allTools,
|
|
@@ -701,6 +782,7 @@ export async function* runAgent(
|
|
|
701
782
|
},
|
|
702
783
|
hiveConfig,
|
|
703
784
|
workerPool: hiveConfig.tools?.workerPool,
|
|
785
|
+
parallelToolCalls: jevParallel?.parallel,
|
|
704
786
|
signal: opts.signal,
|
|
705
787
|
})
|
|
706
788
|
|
|
@@ -1297,6 +1379,8 @@ export interface IsolatedAgentOptions {
|
|
|
1297
1379
|
* reabría justo en el camino de delegación.
|
|
1298
1380
|
*/
|
|
1299
1381
|
credentials?: ProviderCredentials
|
|
1382
|
+
/** Jev for the worker, inherited from the delegating turn like `credentials`. */
|
|
1383
|
+
jev?: JevOption
|
|
1300
1384
|
}
|
|
1301
1385
|
|
|
1302
1386
|
export async function runAgentIsolatedDetailed(
|
|
@@ -1321,6 +1405,7 @@ export async function runAgentIsolatedDetailed(
|
|
|
1321
1405
|
channel: opts.channel,
|
|
1322
1406
|
sessionId: opts.sessionId,
|
|
1323
1407
|
credentials: opts.credentials,
|
|
1408
|
+
jev: opts.jev,
|
|
1324
1409
|
})) {
|
|
1325
1410
|
if (chunk.agent?.messages?.[0]?.content) {
|
|
1326
1411
|
lastContent = chunk.agent.messages[0].content
|
|
@@ -26,7 +26,10 @@ import {
|
|
|
26
26
|
type StoredMessage,
|
|
27
27
|
} from "./conversation-store.ts"
|
|
28
28
|
import { estimateTokens } from "../utils/toon.ts"
|
|
29
|
-
import {
|
|
29
|
+
import {
|
|
30
|
+
callLLM, resolveProviderConfig, getDefaultLLM,
|
|
31
|
+
type ContentPart, type ProviderCredentials,
|
|
32
|
+
} from "./llm-client.ts"
|
|
30
33
|
import { col, fromIndexable } from "../storage/hive.ts"
|
|
31
34
|
import type { AgentDoc, ModelDoc } from "../storage/collections.ts"
|
|
32
35
|
import { loadConfig } from "../config/loader.ts"
|
|
@@ -41,6 +44,61 @@ const KEEP_LAST_N_MESSAGES = 5 // always keep most recent N messages
|
|
|
41
44
|
const TOOL_RESULT_MAX_CHARS = 200 // max chars for old tool results after clearing
|
|
42
45
|
const MAX_TRANSCRIPT_MSGS = 30 // cap messages sent to summarizer (avoids OOM on small models)
|
|
43
46
|
const MAX_MSG_CHARS = 300 // chars per message in transcript
|
|
47
|
+
/** Ventana asumida cuando no se conoce la del modelo: `COMPACT_TOKEN_THRESHOLD` es su 25 %. */
|
|
48
|
+
const ASSUMED_CONTEXT_WINDOW = 128_000
|
|
49
|
+
const DEFAULT_CONTEXT_RATIO = 0.25
|
|
50
|
+
|
|
51
|
+
/**
|
|
52
|
+
* El modelo con el que se pide el resumen: el del turno que disparó la
|
|
53
|
+
* compactación, con SUS credenciales.
|
|
54
|
+
*/
|
|
55
|
+
export interface CompactionLLM {
|
|
56
|
+
provider?: string
|
|
57
|
+
model?: string
|
|
58
|
+
/** En multi-inquilino, la llave del cliente. Sin esto se usaría una global. */
|
|
59
|
+
credentials?: ProviderCredentials
|
|
60
|
+
/** La ventana del modelo, si quien llama ya la resolvió. */
|
|
61
|
+
contextWindow?: number
|
|
62
|
+
}
|
|
63
|
+
|
|
64
|
+
/**
|
|
65
|
+
* A partir de cuántos tokens de historial se compacta.
|
|
66
|
+
*
|
|
67
|
+
* `agent.context.compactionThreshold` es una PROPORCIÓN de la ventana del
|
|
68
|
+
* modelo —su valor por defecto es 0.8, o sea el 80 %—, pero se leía como si
|
|
69
|
+
* fueran tokens. Con la configuración por defecto el umbral quedaba en 0.8
|
|
70
|
+
* tokens: cualquier hilo con más de cinco mensajes se resumía en cada turno,
|
|
71
|
+
* pagando una llamada extra al modelo y reemplazando el historial por un
|
|
72
|
+
* resumen desde el primer intercambio. Un valor mayor que 1 se sigue leyendo
|
|
73
|
+
* como tokens, que es lo que espera quien fijó un número absoluto.
|
|
74
|
+
*/
|
|
75
|
+
export function resolveCompactionThreshold(configured: number | undefined, contextWindow?: number): number {
|
|
76
|
+
const known = contextWindow && contextWindow > 0 ? contextWindow : undefined
|
|
77
|
+
if (typeof configured === "number" && Number.isFinite(configured) && configured > 0) {
|
|
78
|
+
return configured <= 1
|
|
79
|
+
? Math.floor((known ?? ASSUMED_CONTEXT_WINDOW) * configured)
|
|
80
|
+
: Math.floor(configured)
|
|
81
|
+
}
|
|
82
|
+
return known ? Math.floor(known * DEFAULT_CONTEXT_RATIO) : COMPACT_TOKEN_THRESHOLD
|
|
83
|
+
}
|
|
84
|
+
|
|
85
|
+
/** La ventana del modelo del turno; si no se sabe cuál es, la del coordinador. */
|
|
86
|
+
async function modelContextWindow(modelId?: string): Promise<number | undefined> {
|
|
87
|
+
try {
|
|
88
|
+
const modelsCol = await col<ModelDoc>("models")
|
|
89
|
+
if (modelId) return (await modelsCol.get(modelId))?.doc.context_window || undefined
|
|
90
|
+
const agentsCol = await col<AgentDoc>("agents")
|
|
91
|
+
const coordinators = await agentsCol.findBy("role", "coordinator", { limit: 1 })
|
|
92
|
+
// El id se busca completo: recortar el primer segmento rompía cualquier
|
|
93
|
+
// modelo cuyo nombre lleve barra (meta/llama-3.3-70b-instruct buscaba
|
|
94
|
+
// "llama-3.3-70b-instruct", no encontraba nada y caía al default).
|
|
95
|
+
const id = fromIndexable(coordinators[0]?.doc.model_id ?? null)
|
|
96
|
+
if (!id) return undefined
|
|
97
|
+
return (await modelsCol.get(id))?.doc.context_window || undefined
|
|
98
|
+
} catch {
|
|
99
|
+
return undefined
|
|
100
|
+
}
|
|
101
|
+
}
|
|
44
102
|
|
|
45
103
|
/**
|
|
46
104
|
* Check if compaction is needed and run it if so.
|
|
@@ -48,37 +106,16 @@ const MAX_MSG_CHARS = 300 // chars per message in transcript
|
|
|
48
106
|
*/
|
|
49
107
|
export async function maybeCompact(
|
|
50
108
|
threadId: string,
|
|
51
|
-
notify?: { channel: string; userId: string }
|
|
109
|
+
notify?: { channel: string; userId: string },
|
|
110
|
+
llm?: CompactionLLM
|
|
52
111
|
): Promise<void> {
|
|
53
112
|
try {
|
|
54
113
|
const totalTokens = await getTotalTokens(threadId)
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
// y **no lo leía nadie**: alguien podía ajustarlo y no pasaba nada. Una
|
|
61
|
-
// opción que no hace nada es peor que no tenerla, porque el usuario cree
|
|
62
|
-
// que cambió algo.
|
|
63
|
-
let effectiveThreshold = COMPACT_TOKEN_THRESHOLD
|
|
64
|
-
const configurado = loadConfig().agent?.context?.compactionThreshold
|
|
65
|
-
try {
|
|
66
|
-
const agentsCol = await col<AgentDoc>("agents")
|
|
67
|
-
const coordinators = await agentsCol.findBy("role", "coordinator", { limit: 1 })
|
|
68
|
-
const modelId = fromIndexable(coordinators[0]?.doc.model_id ?? null)
|
|
69
|
-
if (modelId) {
|
|
70
|
-
const modelsCol = await col<ModelDoc>("models")
|
|
71
|
-
// El id se busca completo: recortar el primer segmento rompía cualquier
|
|
72
|
-
// modelo cuyo nombre lleve barra (meta/llama-3.3-70b-instruct buscaba
|
|
73
|
-
// "llama-3.3-70b-instruct", no encontraba nada y caía al default).
|
|
74
|
-
const modelEntry = await modelsCol.get(modelId)
|
|
75
|
-
if (modelEntry?.doc.context_window) {
|
|
76
|
-
effectiveThreshold = Math.floor(modelEntry.doc.context_window * 0.25)
|
|
77
|
-
}
|
|
78
|
-
}
|
|
79
|
-
} catch { /* use default threshold */ }
|
|
80
|
-
|
|
81
|
-
if (configurado && configurado > 0) effectiveThreshold = configurado
|
|
114
|
+
const contextWindow = llm?.contextWindow ?? (await modelContextWindow(llm?.model))
|
|
115
|
+
const effectiveThreshold = resolveCompactionThreshold(
|
|
116
|
+
loadConfig().agent?.context?.compactionThreshold,
|
|
117
|
+
contextWindow,
|
|
118
|
+
)
|
|
82
119
|
|
|
83
120
|
if (totalTokens < effectiveThreshold) return
|
|
84
121
|
|
|
@@ -96,8 +133,8 @@ export async function maybeCompact(
|
|
|
96
133
|
// Already summarized up to near the current state
|
|
97
134
|
if (summary && summary.last_message_id > totalMessages - KEEP_LAST_N_MESSAGES) return
|
|
98
135
|
|
|
99
|
-
log.info(`[compaction] Compacting thread=${threadId} tokens=${totalTokens}`)
|
|
100
|
-
await compactThread(threadId, notify)
|
|
136
|
+
log.info(`[compaction] Compacting thread=${threadId} tokens=${totalTokens} threshold=${effectiveThreshold}`)
|
|
137
|
+
await compactThread(threadId, notify, llm)
|
|
101
138
|
} catch (err) {
|
|
102
139
|
log.warn("[compaction] Error during compaction check:", err)
|
|
103
140
|
}
|
|
@@ -139,7 +176,8 @@ export function renderTranscript(rows: StoredMessage[], maxMsgChars = MAX_MSG_CH
|
|
|
139
176
|
*/
|
|
140
177
|
export async function compactThread(
|
|
141
178
|
threadId: string,
|
|
142
|
-
notify?: { channel: string; userId: string }
|
|
179
|
+
notify?: { channel: string; userId: string },
|
|
180
|
+
llm?: CompactionLLM
|
|
143
181
|
): Promise<void> {
|
|
144
182
|
const allMessages = await getHistory(threadId)
|
|
145
183
|
if (allMessages.length <= KEEP_LAST_N_MESSAGES) return
|
|
@@ -162,10 +200,18 @@ export async function compactThread(
|
|
|
162
200
|
const capped = toSummarize.slice(-MAX_TRANSCRIPT_MSGS)
|
|
163
201
|
const transcript = renderTranscript(capped)
|
|
164
202
|
|
|
165
|
-
|
|
166
|
-
|
|
203
|
+
// El modelo del turno y SUS credenciales. Antes el resumen se pedía siempre
|
|
204
|
+
// con `getDefaultLLM()` y sin credenciales, así que `resolveProviderConfig`
|
|
205
|
+
// caía al secret store, al llavero del sistema o al entorno: en una
|
|
206
|
+
// instalación multi-inquilino eso resume la conversación de un cliente con
|
|
207
|
+
// la llave de la plataforma (o de otro cliente). Sin `llm` se comporta como
|
|
208
|
+
// antes, que es lo que necesita una instalación de un solo usuario.
|
|
209
|
+
const target = llm?.provider && llm.model
|
|
210
|
+
? { provider: llm.provider, model: llm.model }
|
|
211
|
+
: await getDefaultLLM()
|
|
212
|
+
if (!target) throw new Error("No active LLM providers/models configured in the database")
|
|
167
213
|
|
|
168
|
-
const providerCfg = await resolveProviderConfig(
|
|
214
|
+
const providerCfg = await resolveProviderConfig(target.provider, target.model, llm?.credentials)
|
|
169
215
|
|
|
170
216
|
const summaryResponse = await callLLM({
|
|
171
217
|
...providerCfg,
|