mixdog 1.0.5 → 1.0.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +206 -475
- package/package.json +11 -4
- package/src/defaults/skills/setup/references/actions.md +26 -6
- package/src/defaults/skills/setup/references/surfaces.md +8 -3
- package/src/runtime/agent/orchestrator/runtime-core/builtin-features.mjs +16 -7
- package/src/runtime/agent/orchestrator/runtime-core/config-helpers.mjs +6 -2
- package/src/runtime/agent/orchestrator/session/manager/session-record.mjs +0 -3
- package/src/runtime/agent/orchestrator/session/store-summary-index.mjs +0 -2
- package/src/runtime/agent/orchestrator/session/store-summary-reader.mjs +24 -2
- package/src/runtime/agent/orchestrator/session/store-transcript-cache.mjs +25 -0
- package/src/runtime/agent/orchestrator/session/store-transcript-worker.mjs +14 -1
- package/src/runtime/agent/orchestrator/tools/builtin/runtime-capabilities.mjs +3 -45
- package/src/runtime/channels/lib/config.mjs +0 -1
- package/src/runtime/channels/lib/owned-runtime/config-reload.mjs +1 -1
- package/src/runtime/channels/lib/scheduler.mjs +15 -56
- package/src/runtime/channels/lib/worker-main.mjs +2 -9
- package/src/runtime/shared/config.mjs +66 -1
- package/src/runtime/shared/llm/usage-ledger-quota.mjs +13 -2
- package/src/runtime/shared/path-executable.mjs +58 -0
- package/src/session-runtime/boot/apis.mjs +6 -2
- package/src/session-runtime/boot/begin.mjs +0 -2
- package/src/session-runtime/boot/providers.mjs +3 -0
- package/src/session-runtime/boot/tools.mjs +0 -1
- package/src/session-runtime/config-lifecycle.mjs +11 -0
- package/src/session-runtime/internal-tool-executor.mjs +0 -2
- package/src/session-runtime/runtime-feature-gates.mjs +2 -3
- package/src/session-runtime/schedule-session-run.mjs +0 -1
- package/src/session-runtime/services/channel-admin.mjs +8 -26
- package/src/session-runtime/settings-compaction-api.mjs +5 -1
- package/src/session-runtime/settings-system-api.mjs +2 -20
- package/src/session-runtime/setup-tool/executor.mjs +61 -3
- package/src/session-runtime/setup-tool/extended-actions.mjs +31 -6
- package/src/session-runtime/setup-tool/settings-contract.mjs +27 -3
- package/src/session-runtime/setup-tool/tool-defs.mjs +35 -5
- package/src/session-runtime/webhook-session-run.mjs +0 -2
- package/src/session-runtime/workflow-agents-api/agent-editor.mjs +13 -3
- package/src/standalone/daemon-stored-session-views.mjs +4 -0
- package/src/standalone/session-protocol.mjs +0 -1
- package/src/standalone/session-service/session-calls/session-views.mjs +10 -4
- package/src/standalone/session-service/stored-reader.mjs +20 -2
- package/src/standalone/session-service/viewers.mjs +3 -3
- package/src/standalone/session-service.mjs +3 -1
- package/src/tui/App.jsx +6 -1
- package/src/tui/app/create-app-pickers.mjs +1 -0
- package/src/tui/app/doctor.mjs +6 -14
- package/src/tui/app/maintenance-pickers/auto-clear-picker.mjs +8 -3
- package/src/tui/app/settings-picker/settings-rows.mjs +5 -1
- package/src/tui/app/use-ui-open-request.mjs +13 -3
- package/src/tui/dist/index.mjs +41 -9
- package/src/tui/session/session-api/settings.mjs +0 -6
package/README.md
CHANGED
|
@@ -1,521 +1,252 @@
|
|
|
1
|
-
|
|
2
|
-
|
|
3
|
-
|
|
4
|
-

|
|
5
|
-

|
|
6
|
-
|
|
7
|
-
## More work. Less cost. Less complexity.
|
|
8
|
-
|
|
9
|
-
Get more from your models and budget with an efficient AI coding
|
|
10
|
-
harness—and intuitive controls for managing sessions, agents, and your
|
|
11
|
-
entire workflow.
|
|
12
|
-
|
|
13
|
-
- **More work for your budget.** Cache-aware context, focused
|
|
14
|
-
tools, and compaction reduce overhead so more of your budget goes toward
|
|
15
|
-
the task. The published same-model Terminal-Bench comparisons below show comparable
|
|
16
|
-
or better results with smaller contexts and lower costs at the same API rates.
|
|
17
|
-
- **Easy to start. Simple to manage.** Guided setup and visual controls
|
|
18
|
-
help you choose models, assign agent roles, and configure workflows without
|
|
19
|
-
becoming an expert in agent infrastructure or building your own stack.
|
|
20
|
-
- **One workspace, your way.** In Desktop, organize parallel sessions with
|
|
21
|
-
tabs and split panes, customize agents and workflows, and keep token
|
|
22
|
-
statistics and supported provider limits in view.
|
|
23
|
-
|
|
24
|
-
Use supported subscription accounts, API keys, or Mixdog's built-in Local Provider.
|
|
25
|
-
Take the same agent beyond code into browsers, Windows apps, documents,
|
|
26
|
-
images, and video—and continue live sessions across terminal, Desktop,
|
|
27
|
-
and a paired browser on your computer or phone.
|
|
1
|
+
<p align="center">
|
|
2
|
+
<img src="https://raw.githubusercontent.com/tribgames/mixdog/main/apps/desktop/build/mixdog.png" alt="" width="96">
|
|
3
|
+
</p>
|
|
28
4
|
|
|
29
|
-
|
|
5
|
+
<h1 align="center">Mixdog</h1>
|
|
30
6
|
|
|
31
|
-
|
|
7
|
+
<p align="center">
|
|
8
|
+
<b>Same model. Same performance. 63% fewer tokens.</b><br>
|
|
9
|
+
Free, open-source desktop coding agent for Windows.
|
|
10
|
+
</p>
|
|
32
11
|
|
|
33
12
|
<p align="center">
|
|
34
13
|
<a href="https://github.com/tribgames/mixdog/releases/latest/download/mixdog-desktop-win-x64.exe">
|
|
35
|
-
<img src="https://
|
|
14
|
+
<img src="https://raw.githubusercontent.com/tribgames/mixdog/main/docs/assets/download-windows.svg" alt="Download Mixdog for Windows x64" width="320">
|
|
36
15
|
</a>
|
|
37
16
|
</p>
|
|
38
17
|
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
Requires Node.js 22.19+ (22.x) or 24+.
|
|
45
|
-
|
|
46
|
-
```bash
|
|
47
|
-
npm install -g mixdog
|
|
48
|
-
mixdog
|
|
49
|
-
```
|
|
50
|
-
|
|
51
|
-
First run guides you through provider authentication, model selection, and
|
|
52
|
-
workflow setup.
|
|
53
|
-
|
|
54
|
-
## One workspace for your sessions and agents
|
|
55
|
-
|
|
56
|
-
Run multiple AI sessions side by side and manage your agents in Desktop.
|
|
57
|
-
Combine tabs and split panes, customize how you work, and keep token usage
|
|
58
|
-
and supported provider limits in view.
|
|
59
|
-
|
|
60
|
-
- **Multiple sessions, manageable agents.** Keep separate tasks in separate
|
|
61
|
-
sessions and work on them in parallel. Manage agent definitions, assign
|
|
62
|
-
models by role, and configure workflows in one app instead of assembling
|
|
63
|
-
your own agent stack.
|
|
64
|
-
- **Tabs and split panes—together.** Use tabs to organize sessions and split
|
|
65
|
-
panes to follow several side by side. Each pane can hold its own tabs, so
|
|
66
|
-
you do not have to choose between quick switching and a simultaneous view.
|
|
67
|
-
- **Make the workspace your own.** Visual controls put layout, providers,
|
|
68
|
-
models, agent rules, workflows, and extensions within easy reach. The model
|
|
69
|
-
picker shows pricing, context limits, and capability metadata to help you
|
|
70
|
-
choose—not just a list of model names.
|
|
71
|
-
- **See where your tokens go.** Usage statistics break down token totals by
|
|
72
|
-
provider and model, including input, output, cache hits, and cache hit rate,
|
|
73
|
-
with trends and cost figures. Subscription values use list prices; API
|
|
74
|
-
costs may be estimates. Neither is an invoice.
|
|
75
|
-
- **Keep remaining usage in sight.** The usage panel brings supported
|
|
76
|
-
providers' quota windows and reset times together, reducing trips to
|
|
77
|
-
separate account dashboards. Available figures depend on the provider.
|
|
78
|
-
- **Less window switching.** Chat, a Monaco code editor, Git, terminals, and
|
|
79
|
-
a file explorer share one workspace, keeping the conversation close to the
|
|
80
|
-
files and changes you are working on.
|
|
81
|
-
- **Pick up on another screen.** Continue the same live session from Desktop,
|
|
82
|
-
TUI, or a paired browser on your computer or phone, without starting a
|
|
83
|
-
separate conversation.
|
|
84
|
-
|
|
85
|
-
## Benchmarks
|
|
86
|
-
|
|
87
|
-
Terminal-Bench 2.1 — same model, same 89 tasks, same official verifier, with
|
|
88
|
-
only the harness changed. Against the native CLI of each model family, Mixdog
|
|
89
|
-
delivers the same results at the same speed — on a fraction of the context,
|
|
90
|
-
for far less cost.
|
|
91
|
-
|
|
92
|
-
### GPT-5.6 Sol xhigh — Mixdog vs Codex CLI
|
|
93
|
-
|
|
94
|
-

|
|
95
|
-
|
|
96
|
-
- **39%** lower priced cost — $0.476 vs $0.782 per trial
|
|
97
|
-
- **46%** smaller median final context — 18.5k vs 34.3k tokens
|
|
98
|
-
- **86.5%** (385/445) vs Codex CLI's **86.1%** (383/445) — full `k=5` on both
|
|
99
|
-
sides, pass@5 **96.6%** vs 95.5%
|
|
100
|
-
- Matched speed — 415s vs 437s wall time per trial
|
|
101
|
-
|
|
102
|
-
### Claude Opus 5 — Mixdog vs Claude Code
|
|
18
|
+
<p align="center">
|
|
19
|
+
<a href="https://github.com/tribgames/mixdog/releases/latest"><img src="https://img.shields.io/github/v/release/tribgames/mixdog?label=release" alt="Latest release"></a>
|
|
20
|
+
<a href="LICENSE"><img src="https://img.shields.io/badge/license-Apache--2.0-blue" alt="Apache-2.0 license"></a>
|
|
21
|
+
<img src="https://img.shields.io/badge/platform-Windows%20x64-0078D4" alt="Windows x64">
|
|
22
|
+
</p>
|
|
103
23
|
|
|
104
|
-
|
|
24
|
+
<p align="center">
|
|
25
|
+
<sub>The installer is not code-signed yet. If Windows SmartScreen appears,
|
|
26
|
+
choose <b>More info → Run anyway</b>.</sub>
|
|
27
|
+
</p>
|
|
105
28
|
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
- **1.16×** faster — 610s vs 708s wall time per trial
|
|
110
|
-
|
|
111
|
-
These published runs use the official Harbor verifier with fast mode off;
|
|
112
|
-
task failures and agent timeouts are never retried. The Sol comparison
|
|
113
|
-
follows the protocol the official Terminal-Bench leaderboard requires on both
|
|
114
|
-
sides — all 89 tasks repeated five times (`k=5`, 445 trials each); the
|
|
115
|
-
Opus-side runs are single passes (`k=1`, 89 trials each). Speed is the full
|
|
116
|
-
trial wall clock, and cost values both sides at the same current API list
|
|
117
|
-
rates, not actual subscription charges or invoices. These are measurements of
|
|
118
|
-
the pinned source revision, not a new benchmark of every subsequent release.
|
|
119
|
-
|
|
120
|
-
The leaderboard is not accepting community submissions, so every run here ships
|
|
121
|
-
its raw artifacts instead — Harbor verdicts, official verifier output, pinned
|
|
122
|
-
task checksums, and the usage snapshots behind every cost figure — alongside
|
|
123
|
-
the harness, presets, and metric scripts that recompute each number above:
|
|
124
|
-
[`benchmarks/terminal-bench-2.1/`](benchmarks/terminal-bench-2.1/).
|
|
125
|
-
|
|
126
|
-
## Less overhead. More budget for the work.
|
|
127
|
-
|
|
128
|
-
Mixdog reduces the overhead of repeatedly sending context, re-explaining
|
|
129
|
-
requirements, and rediscovering prior work. Focused tools keep unnecessary
|
|
130
|
-
text out of the prompt, while provider-aware caching reuses stable input.
|
|
131
|
-
|
|
132
|
-
Compaction keeps long conversations manageable with a handoff for continuing
|
|
133
|
-
the task. Optional idle-time compaction reduces the history resent after
|
|
134
|
-
long breaks, when provider caches may have expired. Approved memory and
|
|
135
|
-
past-work retrieval help carry earlier decisions and requirements forward
|
|
136
|
-
without loading the entire conversation archive into every prompt.
|
|
137
|
-
|
|
138
|
-
You do not have to use the same high-cost model for every role. Choose models
|
|
139
|
-
by role and workflow to focus your budget on the work that needs them.
|
|
140
|
-
|
|
141
|
-
The benchmarks above measure single-model, single-session runs without
|
|
142
|
-
personal memory, sub-agent delegation, or helper-model lookups. Their cost
|
|
143
|
-
figures already account for cache usage; savings in ongoing work depend on
|
|
144
|
-
the provider, workload, and configuration.
|
|
145
|
-
|
|
146
|
-
## How Mixdog keeps context lean
|
|
147
|
-
|
|
148
|
-
Efficiency comes from several layers working together, not just a shorter
|
|
149
|
-
prompt or a larger context window:
|
|
150
|
-
|
|
151
|
-
1. **Lightweight system instructions** — continuously refined rules keep
|
|
152
|
-
operating guidance focused without repeating the same policy.
|
|
153
|
-
2. **Purpose-built tools** — scoped queries, batched calls, and bounded
|
|
154
|
-
results retrieve the evidence a task needs instead of dumping whole files.
|
|
155
|
-
3. **Built-in ast-grep and code graphs** — parsed symbols, signatures, calls,
|
|
156
|
-
and imports answer structural questions without repeated text searches.
|
|
157
|
-
4. **Provider-aware caching** — stable prompt layers and provider-specific
|
|
158
|
-
cache controls help reuse context that has already been processed.
|
|
159
|
-
5. **Structured compaction** — a task handoff, the latest request, and a
|
|
160
|
-
bounded execution history keep long sessions moving.
|
|
161
|
-
6. **Idle-time context reduction** — configurable automatic compaction
|
|
162
|
-
reduces the context sent after long idle gaps, when caches may be cold.
|
|
163
|
-
7. **On-demand prompt loading** — skill bodies and deferred tool schemas
|
|
164
|
-
load when needed; stable instructions stay separate from changing state.
|
|
165
|
-
8. **Database-backed long-term memory** — retrieve relevant history instead
|
|
166
|
-
of injecting the whole archive into every session.
|
|
167
|
-
9. **Tool-result reduction** — repeated results become short references,
|
|
168
|
-
while large outputs can be saved separately and returned as previews.
|
|
169
|
-
|
|
170
|
-
Caching can reduce repeated processing and input cost, but cached tokens
|
|
171
|
-
still count toward the model's context limit. Compaction, selective retrieval,
|
|
172
|
-
and output reduction reduce the amount of context the model needs. See
|
|
173
|
-
[Context efficiency](docs/context-efficiency.md) for the mechanisms,
|
|
174
|
-
implementation references, and limits.
|
|
175
|
-
|
|
176
|
-
## What you can do
|
|
177
|
-
|
|
178
|
-
### Build, test, and review
|
|
179
|
-
|
|
180
|
-
Search repositories with text and AST-based tools, edit files, run tests
|
|
181
|
-
and background commands, and review changes. Desktop brings the agent together
|
|
182
|
-
with a Monaco editor, Git, terminals, and a file explorer. Use workflows and
|
|
183
|
-
role-specific models to organize work, and extend the toolset with MCP
|
|
184
|
-
servers, skills, hooks, and plugins.
|
|
185
|
-
|
|
186
|
-
**Code graph.** Inspect exports, signatures, and nested members; locate
|
|
187
|
-
declarations and references; trace calls and imports; and assess which files
|
|
188
|
-
a change may affect. Call relationships come from parsed call sites rather than
|
|
189
|
-
text matches, and identifier references exclude comment-only mentions.
|
|
190
|
-
The native engine embeds tree-sitter and ast-grep, parses 31 languages, and
|
|
191
|
-
extracts symbols and imports for 24. Capabilities vary by language; this is
|
|
192
|
-
structural navigation, not a replacement for a compiler's type analysis.
|
|
193
|
-
|
|
194
|
-
**Code Tidy.** Install the built-in capability to format, lint, and check
|
|
195
|
-
structural rules through the agent. It respects project configuration and
|
|
196
|
-
uses project-local, system-installed, or supported managed engines.
|
|
197
|
-
Fixes are previewed without changing files unless explicitly applied. See
|
|
198
|
-
[Code Tidy](docs/code-tidy.md) for engine setup and rule coverage.
|
|
199
|
-
|
|
200
|
-
The GitHub integration manages repositories, issues, pull requests and reviews,
|
|
201
|
-
Actions, releases, and notifications. Source Control commits use a manually
|
|
202
|
-
entered summary and optional description; there is no built-in AI commit-message
|
|
203
|
-
generator. See [Git & GitHub](docs/git-github-integration.md) for supported
|
|
204
|
-
operations and permission requirements.
|
|
205
|
-
|
|
206
|
-
### Keep longer work moving
|
|
207
|
-
|
|
208
|
-
Resume saved chats and recall prior work through local semantic and lexical
|
|
209
|
-
search. Long-term memory separates searchable conversation history (`recall`)
|
|
210
|
-
from approved shared or project-scoped preferences (`memory`); generated
|
|
211
|
-
conversation summaries do not become standing instructions.
|
|
212
|
-
|
|
213
|
-
Compaction keeps long conversations manageable while retaining the latest
|
|
214
|
-
request and the context needed to continue. Configurable idle-time compaction
|
|
215
|
-
can also reduce input cost when resuming after a long idle period, when the
|
|
216
|
-
provider's cache may have expired.
|
|
217
|
-
For an explicitly requested longer-running objective, **Goals** track
|
|
218
|
-
completion conditions and tasks, support time limits and automatic
|
|
219
|
-
continuation, and let you pause or resume the work.
|
|
220
|
-
|
|
221
|
-
### Work beyond the repository
|
|
222
|
-
|
|
223
|
-
- **Browser Use** — operate signed-in Chromium pages, forms, tabs, and
|
|
224
|
-
downloads. On Windows, import a Chrome profile, including cookies and
|
|
225
|
-
passwords; cookie and password import require administrator approval and
|
|
226
|
-
a build with the native importer. Session cookies are encrypted with the
|
|
227
|
-
OS keychain and restored on launch when encryption is available.
|
|
228
|
-
Developer controls share the same pages and sign-in through
|
|
229
|
-
`browser_devtools`.
|
|
230
|
-
- **Computer Use on Windows** — operate native apps through accessibility,
|
|
231
|
-
screenshots, OCR, keyboard, and pointer input with guarded execution.
|
|
232
|
-
An overlay provides Stop and Resume controls.
|
|
233
|
-
- **Documents** — create and edit Word, Excel, and PowerPoint files, work
|
|
234
|
-
with PDFs, and review rendered previews alongside automated checks.
|
|
235
|
-
Use portable OOXML editing without Microsoft Office, or Microsoft Office
|
|
236
|
-
automation on Windows. Rendering and spreadsheet recalculation depend on
|
|
237
|
-
the available engines. See [Office runtime](src/runtime/office/README.md).
|
|
238
|
-
- **Image and video Studio** — generate and edit images, generate short video
|
|
239
|
-
clips, and keep the results in a persistent local gallery. Continue a clip
|
|
240
|
-
by using its last frame as the reference for a new generation; this carries
|
|
241
|
-
over the pose, not the original motion or camera trajectory. Available
|
|
242
|
-
models and controls depend on your signed-in provider routes.
|
|
243
|
-
|
|
244
|
-
Browser Use and Computer Use are opt-in capabilities. In interactive sessions,
|
|
245
|
-
each asks for approval before its first live call by default; approval covers
|
|
246
|
-
the rest of that session, and a restart asks again. Headless and agent-owned
|
|
247
|
-
sessions without an approval UI are not gated by this first-use prompt.
|
|
248
|
-
|
|
249
|
-
### Continue from another screen
|
|
250
|
-
|
|
251
|
-
Desktop, TUI, and paired browsers share live sessions rather than starting
|
|
252
|
-
independent copies. The installable remote web app connects to Desktop over
|
|
253
|
-
authenticated end-to-end encryption, so you can follow and continue work from
|
|
254
|
-
a computer or phone.
|
|
255
|
-
|
|
256
|
-
## Providers
|
|
257
|
-
|
|
258
|
-
Mixdog supports subscription OAuth and API-key routes, including:
|
|
29
|
+
<p align="center">
|
|
30
|
+
<img src="https://raw.githubusercontent.com/tribgames/mixdog/main/docs/assets/desktop.png" alt="Mixdog Desktop" width="860">
|
|
31
|
+
</p>
|
|
259
32
|
|
|
260
|
-
|
|
261
|
-
|
|
262
|
-
- Google Gemini API keys
|
|
263
|
-
- xAI API keys and Grok account OAuth
|
|
264
|
-
- OpenRouter API keys and its unified model catalog
|
|
265
|
-
- DeepSeek and OpenCode Go
|
|
266
|
-
- Mixdog's built-in Local Provider
|
|
33
|
+
<a name="benchmarks"></a>
|
|
34
|
+
## Same performance, 63% fewer tokens
|
|
267
35
|
|
|
268
|
-
|
|
269
|
-
|
|
270
|
-
OAuth risks account restrictions.
|
|
36
|
+
Mixdog and Codex CLI ran Terminal-Bench 2.1 on the same model, GPT-5.6 Sol
|
|
37
|
+
xhigh. Both passed the same share of tasks. Mixdog used 63% fewer tokens.
|
|
271
38
|
|
|
272
|
-
|
|
273
|
-
|
|
39
|
+
| | Mixdog | Codex CLI | |
|
|
40
|
+
| --- | --- | --- | --- |
|
|
41
|
+
| **Total tokens** (incl. cached input) | **156.5M** | 421.4M | **63% fewer** |
|
|
42
|
+
| Success rate | **86.5%** (385/445) | 86.1% (383/445) | matched |
|
|
43
|
+
| Priced cost per trial | **$0.476** | $0.782 | 39% lower |
|
|
44
|
+
| Median first request | **4.7k** | 14.4k | 67% smaller |
|
|
45
|
+
| Median final context | **18.5k** | 34.3k | 46% smaller |
|
|
46
|
+
| Wall time per trial | **415s** | 437s | matched |
|
|
274
47
|
|
|
275
|
-
|
|
276
|
-
endpoint registry. The former Ollama and LM Studio routes have been retired.
|
|
48
|
+

|
|
277
49
|
|
|
278
|
-
|
|
50
|
+
- **The official leaderboard protocol.** All 89 tasks, five trials each — 445
|
|
51
|
+
trials per side — with unmodified task timeouts and resources, scored by the
|
|
52
|
+
official Harbor verifier.
|
|
53
|
+
- **Same conditions on both sides.** Same model and reasoning level, the same
|
|
54
|
+
kind of subscription sign-in, fast mode off, no retries of task failures or
|
|
55
|
+
agent timeouts.
|
|
56
|
+
- **Everything is published.** Verdicts, verifier output, and usage snapshots
|
|
57
|
+
for both sides, plus the scripts that recompute every number, are in
|
|
58
|
+
[`benchmarks/terminal-bench-2.1/`](benchmarks/terminal-bench-2.1/).
|
|
279
59
|
|
|
280
|
-
|
|
281
|
-
server. The managed runtime currently requires **Windows x64 and an NVIDIA
|
|
282
|
-
GPU**, with enough VRAM for the selected model.
|
|
60
|
+
### Across the 89 tasks
|
|
283
61
|
|
|
284
|
-
|
|
285
|
-
installation. **Extensions → Plugin → Local Provider** manages installed
|
|
286
|
-
models, download progress and resumption, and automatic unloading when idle.
|
|
287
|
-
Installed models are available through the `mixdog-local` provider.
|
|
62
|
+
Mixdog used fewer tokens on 87 of the 89 tasks. The median task used 68% fewer.
|
|
288
63
|
|
|
289
|
-
|
|
64
|
+
| Token change | Tasks |
|
|
65
|
+
| --- | --- |
|
|
66
|
+
| 75% fewer or better | 24 |
|
|
67
|
+
| 50–75% fewer | 42 |
|
|
68
|
+
| 25–50% fewer | 19 |
|
|
69
|
+
| 0–25% fewer | 2 |
|
|
70
|
+
| More tokens | 2 |
|
|
71
|
+
|
|
72
|
+
Pass counts were equal on 60 tasks; Mixdog passed more trials on 14 and Codex
|
|
73
|
+
CLI on 15. Mixdog used more tokens on `mteb-leaderboard` (6.8×) and
|
|
74
|
+
`crack-7z-hash` (1.4×). Every task is listed in
|
|
75
|
+
[`results.md`](benchmarks/terminal-bench-2.1/results.md).
|
|
76
|
+
|
|
77
|
+
<details>
|
|
78
|
+
<summary><b>Claude Opus 5 — Mixdog vs Claude Code</b> (single pass, 89 trials each)</summary>
|
|
79
|
+
|
|
80
|
+
| | Mixdog | Claude Code | |
|
|
81
|
+
| --- | --- | --- | --- |
|
|
82
|
+
| Solved | **79/89** | 77/89 | +2 tasks |
|
|
83
|
+
| Priced cost per run | **$104.29** | $129.21 | 19% lower |
|
|
84
|
+
| Median final context | **27.6k** | 38.2k | 28% smaller |
|
|
85
|
+
| Wall time per trial | **610s** | 708s | 1.16× faster |
|
|
290
86
|
|
|
291
|
-
|
|
292
|
-
# Start in the current project
|
|
293
|
-
mixdog
|
|
87
|
+

|
|
294
88
|
|
|
295
|
-
|
|
296
|
-
|
|
89
|
+
One trial per task, so per-task differences sit within run-to-run variance.
|
|
90
|
+
|
|
91
|
+
</details>
|
|
92
|
+
|
|
93
|
+
<sub>Codex CLI 0.151.0. Mixdog ran single-model, single-session — no
|
|
94
|
+
sub-agents or helper models. Cost values both sides at the same API list
|
|
95
|
+
rates, not subscription charges. Results measure the pinned source revision.
|
|
96
|
+
The leaderboard is not accepting community submissions, so the runs are
|
|
97
|
+
published here instead.</sub>
|
|
98
|
+
|
|
99
|
+
## How it uses fewer tokens
|
|
100
|
+
|
|
101
|
+
- **A light start.** System instructions and tool descriptions are kept to
|
|
102
|
+
what the model needs.
|
|
103
|
+
*67% smaller first request — 4.7k tokens against 14.4k.*
|
|
104
|
+
- **Fewer round trips.** The agent picks the right tool first and runs
|
|
105
|
+
independent actions in one batch.
|
|
106
|
+
*33% fewer model requests — a median of 10 per trial against 16.*
|
|
107
|
+
- **Smaller requests.** Tools return only the part that was asked for, and
|
|
108
|
+
95% of terminal output is filtered before it reaches the model.
|
|
109
|
+
*46% less input per request — 24.3k tokens against 44.9k.*
|
|
110
|
+
- **Shorter output.** Less filler and repetition in replies.
|
|
111
|
+
*8% fewer output tokens.*
|
|
112
|
+
|
|
113
|
+
The benchmark ran one model in one session. In everyday use, more applies on
|
|
114
|
+
top of that:
|
|
115
|
+
|
|
116
|
+
- **Auto-clear.** After a long break, the conversation is compacted before
|
|
117
|
+
work resumes.
|
|
118
|
+
- **Light compaction.** Long sessions continue from a structured handoff
|
|
119
|
+
instead of the full history.
|
|
120
|
+
- **Database memory.** Past work is stored and searched, so the prompt does
|
|
121
|
+
not grow with the archive.
|
|
122
|
+
- **Code graph.** Query a project's structure instead of reading whole files.
|
|
123
|
+
- **Orchestration.** Hand routine work to cheaper models and keep the
|
|
124
|
+
expensive one for the steps that need it.
|
|
125
|
+
|
|
126
|
+
See [Context efficiency](docs/context-efficiency.md) for how each layer works.
|
|
127
|
+
|
|
128
|
+
## A model for each job
|
|
129
|
+
|
|
130
|
+
Not every step needs your most expensive model.
|
|
131
|
+
|
|
132
|
+
- **Mix models by role.** Let a strong model lead and plan while cheaper ones
|
|
133
|
+
search, edit, and review.
|
|
134
|
+
- **Dial delegation from Solo to Swarm.** The lead can do the work itself or
|
|
135
|
+
coordinate a team of agents working in parallel.
|
|
136
|
+
- **Workflows you can read.** Workflows and agents are plain Markdown
|
|
137
|
+
(`WORKFLOW.md`, `AGENT.md`) with visual editors in the app.
|
|
138
|
+
- **Every account in one place.** Claude, ChatGPT, and Grok subscriptions, API
|
|
139
|
+
keys, and local models side by side.
|
|
140
|
+
- **Run models on your own GPU.** The built-in Local Provider downloads and
|
|
141
|
+
runs models inside the app.
|
|
142
|
+
- **Know what it costs.** Usage by provider and model — input, output, cache
|
|
143
|
+
hits, and cost — plus quota windows and resets.
|
|
144
|
+
|
|
145
|
+
## Everything in one window
|
|
146
|
+
|
|
147
|
+
- **A full workspace.** Tabs and split panes, Monaco editor, terminals, file
|
|
148
|
+
explorer, and language servers that start when you open a file.
|
|
149
|
+
- **Git and GitHub built in.** Review changes, commit, and handle pull
|
|
150
|
+
requests, issues, Actions, and releases without leaving the app.
|
|
151
|
+
- **Code that stays clean.** The code graph maps a project's structure, and
|
|
152
|
+
Code Tidy formats, lints, and applies structural fixes.
|
|
153
|
+
- **Memory that lasts.** Mixdog remembers what matters across sessions and
|
|
154
|
+
searches past work on demand.
|
|
155
|
+
- **Extend it.** Add skills, MCP servers, and plugins from one Extensions view.
|
|
156
|
+
|
|
157
|
+
## It keeps working when you step away
|
|
158
|
+
|
|
159
|
+
- **Goals.** Set a Goal and the session keeps working toward it — pause and
|
|
160
|
+
resume whenever you like.
|
|
161
|
+
- **Schedules and webhooks.** Run a task on a timer or whenever a URL is
|
|
162
|
+
called.
|
|
163
|
+
- **Your phone is the remote.** Scan a QR code to continue the same live
|
|
164
|
+
session from your phone, end-to-end encrypted, with a notification when the
|
|
165
|
+
task finishes.
|
|
166
|
+
|
|
167
|
+
## Beyond code
|
|
168
|
+
|
|
169
|
+
- **Browser Use.** Operate real, signed-in web pages, with Chrome profile
|
|
170
|
+
import.
|
|
171
|
+
- **Computer Use.** Operate native Windows apps with guarded input and an
|
|
172
|
+
on-screen Stop control.
|
|
173
|
+
- **Documents.** Create and edit Word, Excel, PowerPoint, and PDF with
|
|
174
|
+
rendered previews.
|
|
175
|
+
- **Image and video Studio.** Generate and edit with a local gallery.
|
|
176
|
+
|
|
177
|
+
Browser Use and Computer Use are opt-in and ask for approval before their
|
|
178
|
+
first live call in each session.
|
|
297
179
|
|
|
298
|
-
|
|
299
|
-
mixdog --workflow default
|
|
180
|
+
## Get started
|
|
300
181
|
|
|
301
|
-
|
|
302
|
-
|
|
182
|
+
[Download the installer](https://github.com/tribgames/mixdog/releases/latest/download/mixdog-desktop-win-x64.exe)
|
|
183
|
+
and sign in. No config files, no YAML, no terminal.
|
|
303
184
|
|
|
304
|
-
|
|
305
|
-
|
|
185
|
+
- **One sign-in and you are working.** Use the ChatGPT, Claude, or Grok
|
|
186
|
+
subscription you already pay for, or paste an API key. Mixdog picks the
|
|
187
|
+
model for you.
|
|
188
|
+
- **Advanced setups, one switch each.** A team of agents, a model per role,
|
|
189
|
+
long-term memory, phone access — each is a toggle, not a config file.
|
|
190
|
+
- **You won't get lost.** A short tutorial gets you set up in five quick
|
|
191
|
+
steps.
|
|
192
|
+
- **Batteries included.** Git, Memory, Browser, and Office tools ship with the
|
|
193
|
+
app. Turn on what you need.
|
|
306
194
|
|
|
307
|
-
|
|
308
|
-
|
|
309
|
-
```
|
|
195
|
+
<details>
|
|
196
|
+
<summary><b>Supported providers</b></summary>
|
|
310
197
|
|
|
311
|
-
|
|
198
|
+
- Anthropic API keys and Claude account OAuth
|
|
199
|
+
- OpenAI API keys and ChatGPT/Codex account OAuth
|
|
200
|
+
- Google Gemini API keys
|
|
201
|
+
- xAI API keys and Grok account OAuth
|
|
202
|
+
- OpenRouter, DeepSeek, and OpenCode Go
|
|
203
|
+
- Built-in **Local Provider** — download and run models in the app
|
|
204
|
+
(Windows x64 with an NVIDIA GPU)
|
|
312
205
|
|
|
313
|
-
|
|
206
|
+
Cursor and Antigravity (Gemini) OAuth are off by default under
|
|
207
|
+
**Settings → Developer**; using them through OAuth risks account restrictions.
|
|
314
208
|
|
|
315
|
-
|
|
316
|
-
configuration and no agent delegation. It requires an explicit provider and
|
|
317
|
-
model. It does not load the host's behavioral configuration, personal memory,
|
|
318
|
-
prior sessions, user profile, skills, MCP servers, or plugins:
|
|
209
|
+
</details>
|
|
319
210
|
|
|
320
|
-
|
|
321
|
-
|
|
322
|
-
mixdog exec --provider openai-oauth --model gpt-5.6-sol --effort xhigh --fast "review the current diff"
|
|
323
|
-
mixdog exec --provider openai-oauth --model gpt-5.6-sol --json "fix the failing test"
|
|
324
|
-
```
|
|
211
|
+
<details>
|
|
212
|
+
<summary><b>Command line</b></summary>
|
|
325
213
|
|
|
326
|
-
|
|
327
|
-
when needed:
|
|
214
|
+
Mixdog also runs in the terminal. Requires Node.js 22.19+ (22.x) or 24+.
|
|
328
215
|
|
|
329
216
|
```bash
|
|
330
|
-
|
|
331
|
-
|
|
332
|
-
|
|
333
|
-
Disabling web search does **not** block ordinary shell networking: package
|
|
334
|
-
managers, Git clients, and other commands can still access the network.
|
|
335
|
-
Headless exec is not an offline sandbox.
|
|
336
|
-
|
|
337
|
-
`--memory`, `--workflow`, `--readonly`, `--remote`, and `--onboarding` are not
|
|
338
|
-
supported by `mixdog exec`. Use an interactive session for personal memory and
|
|
339
|
-
saved-work continuation. `--json` emits timestamped JSONL events to stdout;
|
|
340
|
-
diagnostics remain on stderr.
|
|
341
|
-
|
|
342
|
-
## TUI commands
|
|
343
|
-
|
|
344
|
-
```text
|
|
345
|
-
/clear start a fresh chat
|
|
346
|
-
/project switch the current project
|
|
347
|
-
/resume resume a saved chat
|
|
348
|
-
/inherit carry this conversation into a new session on the current model
|
|
349
|
-
/compact compact older conversation context
|
|
350
|
-
/goal start, inspect, pause, or resume a durable session Goal
|
|
351
|
-
/autoclear manage idle-time context compaction
|
|
352
|
-
/context inspect the current context surface
|
|
353
|
-
/usage show provider quota and balance
|
|
354
|
-
/providers configure provider authentication
|
|
355
|
-
/model choose the main provider and model
|
|
356
|
-
/websearch choose the web search route
|
|
357
|
-
/workflow choose the active workflow
|
|
358
|
-
/agents inspect agents and model overrides
|
|
359
|
-
/effort set reasoning effort
|
|
360
|
-
/fast toggle supported model fast mode
|
|
361
|
-
/OutputStyle choose the Lead response style
|
|
362
|
-
/theme change the TUI color theme
|
|
363
|
-
/memory inspect and edit core memory
|
|
364
|
-
/mcp manage MCP servers and tools
|
|
365
|
-
/skills choose a skill for the next request
|
|
366
|
-
/plugins manage local plugin integrations
|
|
367
|
-
/setting open runtime settings
|
|
368
|
-
/profile set your title, development experience, and response language
|
|
369
|
-
/update check for updates
|
|
370
|
-
/doctor diagnose installation health
|
|
371
|
-
/quit quit the TUI
|
|
372
|
-
```
|
|
373
|
-
|
|
374
|
-
Workflows and agents are Markdown definition packs (`WORKFLOW.md`, `AGENT.md`).
|
|
375
|
-
Built-in packs ship with Mixdog; custom packs live under the Mixdog data
|
|
376
|
-
directory.
|
|
377
|
-
|
|
378
|
-
By default the Lead does the work without delegating (**Solo**). Each
|
|
379
|
-
session's orchestration mode sets how much it delegates to agents, up to
|
|
380
|
-
**Swarm** for maximum delegation. Running multiple independent Desktop
|
|
381
|
-
sessions is separate from delegating work to agents within one session.
|
|
382
|
-
|
|
383
|
-
To start a time-bounded Goal, for example:
|
|
384
|
-
|
|
385
|
-
```text
|
|
386
|
-
/goal Fix the failing tests --time 1h
|
|
387
|
-
/goal status
|
|
388
|
-
/goal pause
|
|
389
|
-
/goal resume
|
|
217
|
+
npm install -g mixdog
|
|
218
|
+
mixdog
|
|
390
219
|
```
|
|
391
220
|
|
|
392
|
-
|
|
393
|
-
|
|
394
|
-
Mixdog Desktop runs the same agent runtime as the CLI. In the **Sessions**
|
|
395
|
-
panel, choose **New task** for agent work or **New Studio** for image and video
|
|
396
|
-
work. The **Projects** panel has **Project** and **Workflow** tabs for managing
|
|
397
|
-
repositories, workflow packs, and agent definitions.
|
|
398
|
-
|
|
399
|
-
The workspace includes:
|
|
400
|
-
|
|
401
|
-
- Split panes for parallel, independently routed agent sessions
|
|
402
|
-
- Live session handoff between the TUI, desktop windows, and paired browsers
|
|
403
|
-
- Monaco editor, LSP integration, diffs, and turn-by-turn edit review
|
|
404
|
-
- Built-in code graph navigation and installable Code Tidy checks and fixes
|
|
405
|
-
- Git staging, commits with manually entered messages, and branches
|
|
406
|
-
- GitHub repositories, issues, pull requests, reviews, Actions, releases,
|
|
407
|
-
and notifications
|
|
408
|
-
- File explorer with previews, thumbnails, search, and drag-and-drop
|
|
409
|
-
- Integrated terminal tabs using the local system shell
|
|
410
|
-
- Browser Use pane with agent control and Chrome profile import on Windows
|
|
411
|
-
- Computer Use on Windows with guarded native input
|
|
412
|
-
- Word, Excel, PowerPoint, and PDF tools with rendered previews
|
|
413
|
-
- Image and video generation Studio with a persistent local gallery
|
|
414
|
-
- Visual workflow, agent, schedule, and webhook editors, plus session Goal
|
|
415
|
-
progress and controls
|
|
416
|
-
- Voice dictation with an optional local transcription runtime
|
|
417
|
-
- Extensions hub with guided setup for Git & GitHub, Memory, Browser Use,
|
|
418
|
-
Computer Use, Office, Code Tidy, Local Provider, and voice
|
|
419
|
-
- Provider setup, usage, git identity, and remote pairing settings
|
|
420
|
-
|
|
421
|
-
In **Extensions**, the **Plugin** tab manages integrations and built-in
|
|
422
|
-
capabilities; the **Skill** tab lets you add and manage skills and MCP servers.
|
|
423
|
-
Language servers start on demand when installed locally or on your system;
|
|
424
|
-
Mixdog does not download them automatically. See
|
|
425
|
-
[language server setup](docs/language-servers.md).
|
|
426
|
-
|
|
427
|
-
The paired remote web app is installable on desktop and mobile browsers. It
|
|
428
|
-
uses an authenticated end-to-end encrypted connection before session state,
|
|
429
|
-
terminal data, files, or operation requests cross the relay, and adds mobile
|
|
430
|
-
share-target intake, push notifications, and remote Browser Use.
|
|
431
|
-
|
|
432
|
-
## Data and configuration
|
|
433
|
-
|
|
434
|
-
Mixdog uses `~/.mixdog` as its home root and `~/.mixdog/data` for runtime data
|
|
435
|
-
by default.
|
|
221
|
+
`mixdog exec` runs one non-interactive, single-model session without personal
|
|
222
|
+
memory, prior sessions, skills, MCP servers, or plugins:
|
|
436
223
|
|
|
437
224
|
```bash
|
|
438
|
-
|
|
439
|
-
MIXDOG_DATA_DIR=/path/to/data mixdog
|
|
225
|
+
mixdog exec --provider openai-oauth --model gpt-5.6-sol --effort xhigh "fix the failing test"
|
|
440
226
|
```
|
|
441
227
|
|
|
442
|
-
|
|
443
|
-
|
|
444
|
-
- `MIXDOG_TUI_MOUSE=0` — use terminal-native mouse behavior.
|
|
445
|
-
- `MIXDOG_DISABLE_MODEL_PREFETCH=1` — disable provider model prefetch.
|
|
446
|
-
- `MIXDOG_MODE=ship|dev` — select shipping or development diagnostics.
|
|
447
|
-
- `MIXDOG_DIAGNOSTICS=1` — force diagnostic trace and log output.
|
|
448
|
-
|
|
449
|
-
## Core technology
|
|
228
|
+
Web search is off by default (`--web-search` enables it). This does not block
|
|
229
|
+
shell networking — headless exec is not an offline sandbox.
|
|
450
230
|
|
|
451
|
-
|
|
452
|
-
| --- | --- |
|
|
453
|
-
| Shared agent runtime | Node.js and ECMAScript modules, shared by CLI and Desktop |
|
|
454
|
-
| Terminal UI | React and Ink |
|
|
455
|
-
| Desktop workspace | Electron, React, TypeScript, Monaco, and xterm.js |
|
|
456
|
-
| Native code tools | Rust, tree-sitter, and embedded ast-grep for parsing, structural queries, and rule-based checks |
|
|
457
|
-
| Browser automation | Chromium and the Chrome DevTools Protocol (CDP) |
|
|
458
|
-
| Long-term memory | Managed local PostgreSQL with pgvector and full-text search |
|
|
459
|
-
| Documents | Portable OOXML and PDF tooling, plus Microsoft Office automation on Windows |
|
|
231
|
+
Run `mixdog --help` for the full reference.
|
|
460
232
|
|
|
461
|
-
|
|
462
|
-
retrieval keeps historical context selective, and provider-specific caching
|
|
463
|
-
reduces repeated model processing. See [Context efficiency](docs/context-efficiency.md)
|
|
464
|
-
for how they work with prompt management and compaction.
|
|
233
|
+
</details>
|
|
465
234
|
|
|
466
|
-
##
|
|
235
|
+
## Docs
|
|
467
236
|
|
|
468
|
-
|
|
469
|
-
|
|
470
|
-
|
|
471
|
-
|
|
472
|
-
|
|
473
|
-
|
|
474
|
-
|
|
475
|
-
npm test -- src/runtime/memory # narrow to one path
|
|
476
|
-
npm run test:slow # *.slow.test.mjs
|
|
477
|
-
npm run test:live # built-artifact or live-system checks
|
|
478
|
-
npm run build:tui
|
|
479
|
-
npm run audit:models
|
|
480
|
-
```
|
|
237
|
+
- [Context efficiency](docs/context-efficiency.md)
|
|
238
|
+
- [Benchmarks](benchmarks/terminal-bench-2.1/)
|
|
239
|
+
- [Code Tidy](docs/code-tidy.md)
|
|
240
|
+
- [Git & GitHub](docs/git-github-integration.md)
|
|
241
|
+
- [Language servers](docs/language-servers.md)
|
|
242
|
+
- [Office runtime](src/runtime/office/README.md)
|
|
243
|
+
- [Development, configuration, and testing](docs/development.md)
|
|
481
244
|
|
|
482
|
-
|
|
245
|
+
## Feedback
|
|
483
246
|
|
|
484
|
-
|
|
485
|
-
|
|
486
|
-
|
|
487
|
-
npm run dev
|
|
488
|
-
```
|
|
489
|
-
|
|
490
|
-
Desktop development runs in an isolated dev profile, including its own data,
|
|
491
|
-
daemon and tool connections. It never reads or changes the installed app's
|
|
492
|
-
settings, sign-ins or sessions. `npm run dev` keeps one profile
|
|
493
|
-
(`%LOCALAPPDATA%\mixdog-dev\default` on Windows, `~/.mixdog-dev/default`
|
|
494
|
-
elsewhere), so settings, sign-ins and sessions made in the dev app survive
|
|
495
|
-
restarts; select another kept profile with `npm run dev -- --profile <name>`.
|
|
496
|
-
`npm run dev:fresh` starts from an empty throwaway profile, retained in the
|
|
497
|
-
system temporary directory for diagnosis. CDP uses port `9342`; if it is
|
|
498
|
-
occupied, reuse the running dev app or choose another port with
|
|
499
|
-
`npm run dev -- --port 9343`.
|
|
500
|
-
On Windows, `npm run e2e:direct` and `npm run e2e:direct:source` also use isolated
|
|
501
|
-
profiles and port `9342` (override with `-- -Port 9343`).
|
|
502
|
-
|
|
503
|
-
Both packages discover `*.test.mjs` and `*-test.mjs` under their `src/` and
|
|
504
|
-
`scripts/` directories. Fast, slow, and live tests run in separate lanes;
|
|
505
|
-
live checks need their corresponding built artifacts or services. See
|
|
506
|
-
[testing practices](docs/testing.md) for details.
|
|
507
|
-
|
|
508
|
-
Main directories:
|
|
509
|
-
|
|
510
|
-
```text
|
|
511
|
-
src/ CLI, TUI, runtime, workflows, agents, and rules
|
|
512
|
-
apps/desktop/ desktop app
|
|
513
|
-
apps/relay/ remote web app and relay
|
|
514
|
-
native/ native process, search, graph, patch, and support binaries
|
|
515
|
-
scripts/ tests, diagnostics, benchmarks, and build scripts
|
|
516
|
-
benchmarks/ reproducible benchmark harnesses, results, and raw artifacts
|
|
517
|
-
src/vendor/ vendored runtime components
|
|
518
|
-
```
|
|
247
|
+
Found a bug or missing a feature?
|
|
248
|
+
[Open an issue](https://github.com/tribgames/mixdog/issues). If Mixdog saves
|
|
249
|
+
you tokens, a star helps others find it.
|
|
519
250
|
|
|
520
251
|
## License
|
|
521
252
|
|