open-agents-ai 0.39.1 → 0.40.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (3) hide show
  1. package/README.md +30 -7
  2. package/dist/index.js +60 -11
  3. package/package.json +1 -1
package/README.md CHANGED
@@ -109,7 +109,7 @@ Long conversations consume context window tokens. Open Agents uses progressive c
109
109
 
110
110
  ### How It Works
111
111
 
112
- Compaction triggers automatically when estimated token usage reaches 75% of the model's context window. The system:
112
+ Compaction triggers automatically when estimated token usage reaches a tier-proportional threshold of the model's context window. The system:
113
113
 
114
114
  1. **Preserves** the system prompt and initial user task (head messages)
115
115
  2. **Summarizes** middle messages (tool calls, results, exploration) into a structured digest
@@ -131,13 +131,36 @@ Six strategies are available via `/compact <strategy>`:
131
131
 
132
132
  ### Automatic Compaction
133
133
 
134
- Compaction thresholds scale dynamically with model size:
134
+ Compaction thresholds scale **proportionally** with the model's actual context window size:
135
135
 
136
- | Model Tier | Threshold | Recent Messages Kept |
137
- |------------|-----------|---------------------|
138
- | Large (30B+) | 40,000 tokens (or 75% of context) | 12 messages |
139
- | Medium (8-29B) | 24,000 tokens (or 75% of context) | 8 messages |
140
- | Small (≤7B) | 12,000 tokens (or 75% of context) | 4-6 messages |
136
+ | Model Tier | Normal Mode | Deep Context Mode | Recent Messages Kept |
137
+ |------------|-------------|-------------------|---------------------|
138
+ | Large (30B+) | 75% of context window | 85% of context window | 4-12 (normal) / 4-24 (deep) |
139
+ | Medium (8-29B) | 70% of context window | 85% of context window | 4-12 (normal) / 4-24 (deep) |
140
+ | Small (≤7B) | 65% of context window | 85% of context window | 4-12 (normal) / 4-24 (deep) |
141
+
142
+ For example, a 128K-context large model compacts at ~96K tokens in normal mode (75%) or ~109K tokens in deep mode (85%) — instead of the previous fixed 40K threshold that wasted 69% of available context.
143
+
144
+ ### Deep Context Mode (`/deep`)
145
+
146
+ Toggle with `/deep` — relaxes compaction so large models leverage more of their context window for complex multi-step reasoning.
147
+
148
+ When deep context is active:
149
+ - **Compaction fires at 85%** of context instead of 65-75% — the model retains much more working memory
150
+ - **Double the recent messages** (up to 24 instead of 12) preserved after compaction
151
+ - **Richer summaries** — compression budget increased from 20% to 30% of context
152
+ - **Larger tool outputs** — cap raised from 8K to 16K chars per tool result
153
+ - **Relaxed output folding** — more head/tail lines preserved (50/25 instead of 20/10 for large models)
154
+
155
+ This mirrors how human cognition works during deep problem-solving: situationally-relevant memories are transiently activated to occupy a larger portion of working memory, with the most relevant details in high-attention positions while supporting context backs them up. LLM attention mechanisms work similarly — earlier relevant context still influences generation even at lower positional weight.
156
+
157
+ Use deep context for:
158
+ - Complex multi-file refactoring or debugging
159
+ - Architecture analysis across many files
160
+ - Long debugging sessions where error context from earlier is critical
161
+ - Tasks where the agent needs to reason about patterns across many files
162
+
163
+ The setting persists to `.oa/settings.json`. Deep context is particularly valuable for models with 64K+ context windows (Qwen3.5-122B, Llama 3.1 70B, etc.) where the default thresholds were leaving significant capacity unused.
141
164
 
142
165
  ### Status Bar Context Tracking (`Ctx:`)
143
166
 
package/dist/index.js CHANGED
@@ -13180,6 +13180,7 @@ Rules:
13180
13180
  taskTimeoutMs: options?.taskTimeoutMs ?? 36e5,
13181
13181
  // 60 min
13182
13182
  compactionThreshold: options?.compactionThreshold ?? 4e4,
13183
+ deepContext: options?.deepContext ?? false,
13183
13184
  dynamicContext: options?.dynamicContext ?? "",
13184
13185
  streamEnabled: options?.streamEnabled ?? false,
13185
13186
  bruteForce: options?.bruteForce ?? true,
@@ -13225,19 +13226,47 @@ Rules:
13225
13226
  /**
13226
13227
  * Compute all context-dependent limits from the current contextWindowSize
13227
13228
  * and modelTier. Returns sensible defaults when contextWindowSize is 0.
13229
+ *
13230
+ * Key insight: LLM attention is most effective when the most relevant content
13231
+ * sits in high-attention positions (start and end of context). Rather than
13232
+ * compacting aggressively and losing detail, we scale thresholds proportionally
13233
+ * to the model's actual context window so larger models can leverage more of
13234
+ * their capacity. The deepContext flag further relaxes limits for complex
13235
+ * multi-step reasoning where earlier context carries critical signal.
13236
+ *
13237
+ * Context utilization targets:
13238
+ * - Small models: compact at ~65% (limited attention span, aggressive helps)
13239
+ * - Medium models: compact at ~70% (balanced)
13240
+ * - Large models: compact at ~75% (plenty of room, let attention work)
13241
+ * - Deep context: compact at ~85% (trust the model's full attention window)
13228
13242
  */
13229
13243
  contextLimits() {
13230
13244
  const ctx = this.options.contextWindowSize;
13231
13245
  const tier = this.options.modelTier ?? "large";
13232
- const compactionThreshold = ctx > 0 ? Math.min(this.options.compactionThreshold, Math.floor(ctx * 0.75)) : this.options.compactionThreshold;
13233
- const keepRecentDivisor = tier === "small" ? 2e3 : tier === "medium" ? 3e3 : 4e3;
13234
- const keepRecent = ctx > 0 ? Math.max(4, Math.min(12, Math.floor(ctx / keepRecentDivisor))) : 12;
13246
+ const deep = this.options.deepContext ?? false;
13247
+ let compactionThreshold;
13248
+ if (ctx > 0) {
13249
+ const factor = deep ? 0.85 : tier === "small" ? 0.65 : tier === "medium" ? 0.7 : 0.75;
13250
+ compactionThreshold = Math.floor(ctx * factor);
13251
+ } else {
13252
+ compactionThreshold = deep ? 8e4 : this.options.compactionThreshold;
13253
+ }
13254
+ const keepRecentMax = deep ? 24 : 12;
13255
+ let keepRecent;
13256
+ if (ctx > 0) {
13257
+ const keepRecentDivisor = tier === "small" ? 2e3 : tier === "medium" ? 3e3 : 4e3;
13258
+ keepRecent = Math.max(4, Math.min(keepRecentMax, Math.floor(ctx / keepRecentDivisor)));
13259
+ } else {
13260
+ keepRecent = deep ? 20 : 12;
13261
+ }
13235
13262
  const maxOutputTokens = ctx > 0 ? Math.min(this.options.maxTokens, Math.max(2048, Math.floor(ctx * 0.25))) : this.options.maxTokens;
13236
- const toolOutputMaxChars = ctx > 0 ? Math.max(2e3, Math.min(8e3, Math.floor(ctx * 0.5))) : 8e3;
13237
- const foldLineThreshold = tier === "small" ? 30 : tier === "medium" ? 35 : 40;
13238
- const foldHeadLines = tier === "small" ? 12 : tier === "medium" ? 16 : 20;
13239
- const foldTailLines = tier === "small" ? 5 : tier === "medium" ? 8 : 10;
13240
- const maxSummaryChars = ctx > 0 ? Math.max(2e3, Math.min(8e3, Math.floor(ctx * 0.2))) : 4e3;
13263
+ const toolOutputCeiling = deep ? 16e3 : 8e3;
13264
+ const toolOutputMaxChars = ctx > 0 ? Math.max(2e3, Math.min(toolOutputCeiling, Math.floor(ctx * (deep ? 0.8 : 0.5)))) : toolOutputCeiling;
13265
+ const foldLineThreshold = deep ? tier === "small" ? 50 : tier === "medium" ? 80 : 120 : tier === "small" ? 30 : tier === "medium" ? 35 : 40;
13266
+ const foldHeadLines = deep ? tier === "small" ? 20 : tier === "medium" ? 30 : 50 : tier === "small" ? 12 : tier === "medium" ? 16 : 20;
13267
+ const foldTailLines = deep ? tier === "small" ? 10 : tier === "medium" ? 15 : 25 : tier === "small" ? 5 : tier === "medium" ? 8 : 10;
13268
+ const maxSummaryCeiling = deep ? 16e3 : 8e3;
13269
+ const maxSummaryChars = ctx > 0 ? Math.max(2e3, Math.min(maxSummaryCeiling, Math.floor(ctx * (deep ? 0.3 : 0.2)))) : deep ? 8e3 : 4e3;
13241
13270
  const repetitionWindow = tier === "small" ? 6 : tier === "medium" ? 8 : 10;
13242
13271
  return {
13243
13272
  compactionThreshold,
@@ -16914,6 +16943,7 @@ function renderSlashHelp() {
16914
16943
  ["/compact", "Force context compaction now (default strategy)"],
16915
16944
  ["/compact <strategy>", "Compact with strategy: aggressive, decisions, errors, summary, structured"],
16916
16945
  ["/bruteforce", "Toggle brute-force mode (auto re-engage on turn limit)"],
16946
+ ["/deep", "Toggle deep context \u2014 relaxes compaction so large models use 85% of context"],
16917
16947
  ["/tools", "List agent-created custom tools"],
16918
16948
  ["/skills", "List available AIWG skills"],
16919
16949
  ["/skills <keyword>", "Filter skills by name or trigger"],
@@ -19451,6 +19481,18 @@ async function handleSlashCommand(input, ctx) {
19451
19481
  renderInfo(`Brute-force mode: ${isOn ? "on" : "off"}${hasLocal ? " (project-local)" : ""}` + (isOn ? " \u2014 agent will auto re-engage when turn limit is hit, reassess and try creative strategies" : ""));
19452
19482
  return "handled";
19453
19483
  }
19484
+ case "deep":
19485
+ case "deep-context": {
19486
+ if (!ctx.deepContextToggle) {
19487
+ renderWarning("Deep context not available");
19488
+ return "handled";
19489
+ }
19490
+ const isOn = ctx.deepContextToggle();
19491
+ const save = hasLocal ? ctx.saveLocalSettings.bind(ctx) : ctx.saveSettings.bind(ctx);
19492
+ save({ deepContext: isOn });
19493
+ renderInfo(`Deep context mode: ${isOn ? "ON" : "OFF"}${hasLocal ? " (project-local)" : ""}` + (isOn ? " \u2014 compaction relaxed to 85% of context window. Large models will retain more working memory for complex reasoning." : " \u2014 compaction restored to default thresholds."));
19494
+ return "handled";
19495
+ }
19454
19496
  case "emojis":
19455
19497
  case "emoji": {
19456
19498
  const current = ctx.getEmojis?.() ?? true;
@@ -25552,7 +25594,7 @@ Use task_status("${taskId}") or task_output("${taskId}") to check progress.`
25552
25594
  }
25553
25595
  };
25554
25596
  }
25555
- function startTask(task, config, repoRoot, voice, stream, taskStores, bruteForce, statusBar, sudoCallback, costTracker, onComplete, taskType, contextWindowSize, modelCaps, personality) {
25597
+ function startTask(task, config, repoRoot, voice, stream, taskStores, bruteForce, statusBar, sudoCallback, costTracker, onComplete, taskType, contextWindowSize, modelCaps, personality, deepContext) {
25556
25598
  const voiceStyleMap = {
25557
25599
  concise: 1,
25558
25600
  balanced: 3,
@@ -25576,6 +25618,7 @@ function startTask(task, config, repoRoot, voice, stream, taskStores, bruteForce
25576
25618
  taskTimeoutMs: 36e5,
25577
25619
  // 60 minutes — never give up prematurely
25578
25620
  compactionThreshold,
25621
+ deepContext: deepContext ?? false,
25579
25622
  dynamicContext,
25580
25623
  modelTier,
25581
25624
  streamEnabled: stream?.enabled ?? false,
@@ -25924,6 +25967,7 @@ async function startInteractive(config, repoPath) {
25924
25967
  let streamEnabled = savedSettings.stream ?? false;
25925
25968
  let bruteForceEnabled = savedSettings.bruteforce ?? true;
25926
25969
  let currentStyle = PRESET_NAMES.includes(savedSettings.style) ? savedSettings.style : "balanced";
25970
+ let deepContextEnabled = savedSettings.deepContext ?? false;
25927
25971
  if (savedSettings.emojis !== void 0)
25928
25972
  setEmojisEnabled(savedSettings.emojis);
25929
25973
  if (savedSettings.colors !== void 0)
@@ -26175,6 +26219,7 @@ Rationale: ${proposal.rationale}${provenanceNote}`;
26175
26219
  "/telegram",
26176
26220
  "/compact",
26177
26221
  "/gc",
26222
+ "/deep",
26178
26223
  "/style",
26179
26224
  "/personality"
26180
26225
  ];
@@ -26336,6 +26381,10 @@ Rationale: ${proposal.rationale}${provenanceNote}`;
26336
26381
  bruteForceEnabled = !bruteForceEnabled;
26337
26382
  return bruteForceEnabled;
26338
26383
  },
26384
+ deepContextToggle() {
26385
+ deepContextEnabled = !deepContextEnabled;
26386
+ return deepContextEnabled;
26387
+ },
26339
26388
  setStyle(preset) {
26340
26389
  currentStyle = preset;
26341
26390
  },
@@ -26875,7 +26924,7 @@ Execute this skill now. Follow the behavioral guidance above.`;
26875
26924
  toolPatternStore: toolPatternStore ?? void 0
26876
26925
  }, bruteForceEnabled, statusBar, handleSudoRequest, costTracker, (summary) => {
26877
26926
  lastCompletedSummary = summary;
26878
- }, currentTaskType, resolvedContextWindowSize, resolvedCaps, currentStyle);
26927
+ }, currentTaskType, resolvedContextWindowSize, resolvedCaps, currentStyle, deepContextEnabled);
26879
26928
  activeTask = task;
26880
26929
  showPrompt();
26881
26930
  await task.promise;
@@ -26987,7 +27036,7 @@ NEW TASK: ${fullInput}`;
26987
27036
  toolPatternStore: toolPatternStore ?? void 0
26988
27037
  }, bruteForceEnabled, statusBar, handleSudoRequest, costTracker, (summary) => {
26989
27038
  lastCompletedSummary = summary;
26990
- }, currentTaskType, resolvedContextWindowSize, resolvedCaps, currentStyle);
27039
+ }, currentTaskType, resolvedContextWindowSize, resolvedCaps, currentStyle, deepContextEnabled);
26991
27040
  activeTask = task;
26992
27041
  showPrompt();
26993
27042
  await task.promise;
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "open-agents-ai",
3
- "version": "0.39.1",
3
+ "version": "0.40.0",
4
4
  "description": "AI coding agent powered by open-source models (Ollama/vLLM) — interactive TUI with agentic tool-calling loop",
5
5
  "type": "module",
6
6
  "main": "./dist/index.js",