omnilane 0.10.3 → 0.10.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +38 -1
- package/README.ja.md +18 -1
- package/README.ko.md +17 -1
- package/README.md +20 -1
- package/README.zh-CN.md +15 -1
- package/README.zh-TW.md +15 -1
- package/VERSION +1 -1
- package/package.json +1 -1
- package/routing.yaml +1 -1
package/CHANGELOG.md
CHANGED
|
@@ -6,6 +6,42 @@ semantic version tags.
|
|
|
6
6
|
|
|
7
7
|
## [Unreleased]
|
|
8
8
|
|
|
9
|
+
## [0.10.4] - 2026-07-26
|
|
10
|
+
|
|
11
|
+
No lane ordering changes. This release corrects documentation that could send
|
|
12
|
+
work to the wrong model, and records the evidence behind the shipped defaults.
|
|
13
|
+
|
|
14
|
+
### Changed
|
|
15
|
+
|
|
16
|
+
- Narrowed the `long-context` lane description in `routing.yaml` and in all five
|
|
17
|
+
localized README lane tables. It previously called the lane long-document
|
|
18
|
+
*synthesis* while shipping Gemini first; published multi-needle scores at 1M
|
|
19
|
+
favour Claude by roughly threefold, while Gemini leads single-needle
|
|
20
|
+
retrieval. These are different capabilities with different leaders, and the
|
|
21
|
+
old wording pointed multi-hop work at the wrong candidate. The lane now
|
|
22
|
+
describes retrieval and volume sweeps, and names the Claude candidate for
|
|
23
|
+
integration across scattered sources. Ordering is deliberately unchanged: the
|
|
24
|
+
supporting evidence is secondary and covers prior model generations, which is
|
|
25
|
+
not a sufficient basis for moving a shipped default.
|
|
26
|
+
|
|
27
|
+
### Fixed
|
|
28
|
+
|
|
29
|
+
- `docs/model-capabilities-2026-07.md` quoted the Artificial Analysis Coding
|
|
30
|
+
Agent Index at v1.1 while the index had re-based twice more. The same model
|
|
31
|
+
reads 80, 78 or 67 depending on the version and harness a source used, so the
|
|
32
|
+
figure is not portable across versions. The section now records every observed
|
|
33
|
+
value with its provenance, documents the v1.3 composition, and states that the
|
|
34
|
+
index may be cited for ordering but never for a number.
|
|
35
|
+
- Documented the writing evidence behind `taste-final`, which previously rested
|
|
36
|
+
entirely on general and agentic indexes that do not measure prose. Added
|
|
37
|
+
EQ-Bench Creative Writing v3, EQ-Bench Longform, and the Lech Mazur
|
|
38
|
+
story-writing benchmark, each read from the publisher. Recorded that Mazur has
|
|
39
|
+
not yet evaluated Claude Opus 5 and that Claude Fable 5 leads that board, so
|
|
40
|
+
the question stays open rather than being presented as settled.
|
|
41
|
+
- Added per-effort cost and throughput to the Intelligence Index table, which is
|
|
42
|
+
what justifies the shipped effort levels: `xhigh` reaches the same index score
|
|
43
|
+
as `max` for 30% (Opus 5) and 53% (GPT-5.6 Sol) less per task.
|
|
44
|
+
|
|
9
45
|
## [0.10.3] - 2026-07-26
|
|
10
46
|
|
|
11
47
|
### Changed
|
|
@@ -411,7 +447,8 @@ semantic version tags.
|
|
|
411
447
|
- Initial shared routing table, cross-vendor dispatcher, runners, installer,
|
|
412
448
|
and baseline lint fixes.
|
|
413
449
|
|
|
414
|
-
[Unreleased]: https://github.com/Seraphim0916/omnilane/compare/v0.10.
|
|
450
|
+
[Unreleased]: https://github.com/Seraphim0916/omnilane/compare/v0.10.4...HEAD
|
|
451
|
+
[0.10.4]: https://github.com/Seraphim0916/omnilane/compare/v0.10.3...v0.10.4
|
|
415
452
|
[0.10.3]: https://github.com/Seraphim0916/omnilane/compare/v0.10.2...v0.10.3
|
|
416
453
|
[0.10.2]: https://github.com/Seraphim0916/omnilane/compare/v0.10.1...v0.10.2
|
|
417
454
|
[0.10.1]: https://github.com/Seraphim0916/omnilane/compare/v0.9.1...v0.10.1
|
package/README.ja.md
CHANGED
|
@@ -119,7 +119,7 @@ flowchart LR
|
|
|
119
119
|
| ✒️ taste-final | Claude Opus 5 (high) | GPT-5.6 Sol (max) | 対外文章、prompt/ドキュメント推敲、スタイル最終審 |
|
|
120
120
|
| 💬 consult | 明示指定したベンダー/モデル | —(フォールバックなし) | 自然言語で直接相談。`--vendor` を必ず維持 |
|
|
121
121
|
| 🎨 ui-draft | GPT-5.6 Sol (xhigh) | Claude Opus 5 (high) | デザインシステム/参考画像がある場合の UI ドラフト |
|
|
122
|
-
| 📚 long-context | Gemini 3.1 Pro (High) | Claude Opus 5 (high) | 100
|
|
122
|
+
| 📚 long-context | Gemini 3.1 Pro (High) | Claude Opus 5 (high) | 100 万トークン級の走査と検索。複数箇所をまたぐ統合には Claude 候補を、高速反復ループは Flash を優先 |
|
|
123
123
|
| ⚡ fast-agentic | Gemini 3.6 Flash (High) | GPT-5.6 Luna (high) | 高速なマルチステップ agentic ループ、マルチモーダル確認 |
|
|
124
124
|
| 📡 live-search | Grok 4.5 | —(off) | リアルタイム X/ウェブ検索とソーシャル文脈 |
|
|
125
125
|
| 🚰 coding-overflow | Grok 4.5 | Kimi K3 → Qwen3 Coder Plus → OpenCode | Codex クォータ逼迫時の中級コーディング逃し弁 |
|
|
@@ -503,6 +503,23 @@ scripts/dispatch.sh --dry-run hardest-coding "…" # 解決済みプラン、
|
|
|
503
503
|
|
|
504
504
|
## 📜 リリース履歴
|
|
505
505
|
|
|
506
|
+
## v0.10.4 の新機能
|
|
507
|
+
|
|
508
|
+
- **`long-context` がマルチホップ作業を誤ったモデルに向けなくなりました** — この
|
|
509
|
+
レーンは長文*統合*を名乗りながら Gemini を第一候補にしていましたが、公開されて
|
|
510
|
+
いる 1M トークンのマルチニードル評価では Claude が約 3 倍のスコアを示し、Gemini
|
|
511
|
+
が強いのはシングルニードル検索です。レーンの説明を走査と検索に改め、複数箇所を
|
|
512
|
+
またぐ統合には Claude 候補を案内します。順序は意図的に据え置き — 根拠が二次情報
|
|
513
|
+
であり、前世代モデルの測定に基づくためです。
|
|
514
|
+
- **Coding Agent Index を数値として引用しなくなりました** — 同一モデルがバージョン
|
|
515
|
+
とハーネス次第で 80、78、67 と読み取れます。今後は順序の参照のみに用い、観測値
|
|
516
|
+
ごとの出所を記録しています。
|
|
517
|
+
- **`taste-final` に文章特化の根拠を追加** — 従来は散文を測らない汎用・エージェント
|
|
518
|
+
指標のみで順序を決めていました。EQ-Bench Creative Writing v3、EQ-Bench Longform、
|
|
519
|
+
Lech Mazur の 3 ボードを、いずれも公開元から直接取得して追加しました。
|
|
520
|
+
- **努力度ごとのコストとスループットを追加**。既定が `xhigh` である理由を示します:
|
|
521
|
+
`max` と同じ指数スコアを、1 タスクあたり 30-53% 安く得られます。
|
|
522
|
+
|
|
506
523
|
## v0.10.3 の新機能
|
|
507
524
|
|
|
508
525
|
- **5 言語すべての README を再構成** — 冒頭で「これは何か、なぜ欲しくなるのか」を
|
package/README.ko.md
CHANGED
|
@@ -117,7 +117,7 @@ flowchart LR
|
|
|
117
117
|
| ✒️ taste-final | Claude Opus 5 (high) | GPT-5.6 Sol (max) | 대외 문장, prompt/문서 다듬기, 스타일 최종심 |
|
|
118
118
|
| 💬 consult | 명시적으로 지정한 벤더/모델 | —(폴백 없음) | 자연어 직접 상담. `--vendor` 를 반드시 유지 |
|
|
119
119
|
| 🎨 ui-draft | GPT-5.6 Sol (xhigh) | Claude Opus 5 (high) | 디자인 시스템/참고 이미지가 있을 때의 UI 초안 |
|
|
120
|
-
| 📚 long-context | Gemini 3.1 Pro (High) | Claude Opus 5 (high) | 100만 토큰급
|
|
120
|
+
| 📚 long-context | Gemini 3.1 Pro (High) | Claude Opus 5 (high) | 100만 토큰급 훑기와 검색. 여러 곳을 잇는 다중 홉 통합은 Claude 후보를, 빠른 반복 루프는 Flash 우선 |
|
|
121
121
|
| ⚡ fast-agentic | Gemini 3.6 Flash (High) | GPT-5.6 Luna (high) | 빠른 멀티스텝 agentic 루프, 멀티모달 확인 |
|
|
122
122
|
| 📡 live-search | Grok 4.5 | —(off) | 실시간 X/웹 검색과 소셜 맥락 |
|
|
123
123
|
| 🚰 coding-overflow | Grok 4.5 | Kimi K3 → Qwen3 Coder Plus → OpenCode | Codex 쿼터 소진 시 중급 코딩 안전 밸브 |
|
|
@@ -487,6 +487,22 @@ scripts/dispatch.sh --dry-run hardest-coding "…" # 완전히 해석된 계
|
|
|
487
487
|
|
|
488
488
|
## 📜 릴리스 기록
|
|
489
489
|
|
|
490
|
+
## v0.10.4 새 기능
|
|
491
|
+
|
|
492
|
+
- **`long-context` 가 더 이상 다중 홉 작업을 잘못된 모델로 보내지 않습니다** — 이
|
|
493
|
+
레인은 장문 *통합*을 표방하면서 Gemini 를 1순위로 두었지만, 공개된 1M 토큰 멀티
|
|
494
|
+
니들 점수는 Claude 가 약 3배 앞서고 Gemini 가 강한 쪽은 싱글 니들 검색입니다.
|
|
495
|
+
레인 설명을 훑기와 검색으로 좁히고, 여러 곳을 잇는 통합에는 Claude 후보를
|
|
496
|
+
안내합니다. 순서는 의도적으로 유지 — 근거가 2차 자료이고 이전 세대 모델을 측정한
|
|
497
|
+
것이기 때문입니다.
|
|
498
|
+
- **Coding Agent Index 를 수치로 인용하지 않습니다** — 같은 모델이 버전과 하네스에
|
|
499
|
+
따라 80, 78, 67 로 읽힙니다. 이제 순서 참조로만 쓰고, 관측값마다 출처를 기록합니다.
|
|
500
|
+
- **`taste-final` 에 글쓰기 전용 근거 추가** — 이전에는 산문을 측정하지 않는 범용·
|
|
501
|
+
에이전트 지표만으로 순서를 정했습니다. EQ-Bench Creative Writing v3, EQ-Bench
|
|
502
|
+
Longform, Lech Mazur 세 보드를 모두 발행처에서 직접 가져와 추가했습니다.
|
|
503
|
+
- **노력 수준별 비용과 처리량 추가**. 기본값이 `xhigh` 인 이유를 보여줍니다:
|
|
504
|
+
`max` 와 같은 지수 점수를 작업당 30-53% 저렴하게 얻습니다.
|
|
505
|
+
|
|
490
506
|
## v0.10.3 새 기능
|
|
491
507
|
|
|
492
508
|
- **5개 언어 README 전면 재구성** — 문서 첫머리에서 "이게 무엇이고 왜 필요한가"를
|
package/README.md
CHANGED
|
@@ -122,7 +122,7 @@ actually resolves.
|
|
|
122
122
|
| ✒️ taste-final | Claude Opus 5 (high) | GPT-5.6 Sol (max) | User-facing prose, prompt/doc polish, style arbitration |
|
|
123
123
|
| 💬 consult | Explicit named vendor/model | — (no fallback) | Direct natural-language consultation; always keep `--vendor` |
|
|
124
124
|
| 🎨 ui-draft | GPT-5.6 Sol (xhigh) | Claude Opus 5 (high) | UI drafts only WITH a design system / reference images |
|
|
125
|
-
| 📚 long-context | Gemini 3.1 Pro (High) | Claude Opus 5 (high) | 1M-token
|
|
125
|
+
| 📚 long-context | Gemini 3.1 Pro (High) | Claude Opus 5 (high) | 1M-token sweeps and retrieval; for multi-hop synthesis prefer the Claude candidate, and Flash for fast repeated loops |
|
|
126
126
|
| ⚡ fast-agentic | Gemini 3.6 Flash (High) | GPT-5.6 Luna (high) | Fast multi-step agentic loops, multimodal checks |
|
|
127
127
|
| 📡 live-search | Grok 4.5 | — (off) | Realtime X/web search and social context |
|
|
128
128
|
| 🚰 coding-overflow | Grok 4.5 | Kimi K3 → Qwen3 Coder Plus → OpenCode | Codex-quota relief valve for mid-tier coding |
|
|
@@ -533,6 +533,25 @@ working notes, including per-benchmark caveats, live in
|
|
|
533
533
|
|
|
534
534
|
## 📜 Release history
|
|
535
535
|
|
|
536
|
+
## What's new in v0.10.4
|
|
537
|
+
|
|
538
|
+
- **`long-context` no longer points multi-hop work at the wrong model** — the
|
|
539
|
+
lane called itself long-document *synthesis* while shipping Gemini first, but
|
|
540
|
+
published multi-needle scores at 1M favour Claude by roughly threefold while
|
|
541
|
+
Gemini leads single-needle retrieval. The lane now describes retrieval and
|
|
542
|
+
volume sweeps and names the Claude candidate for integration work. Ordering is
|
|
543
|
+
unchanged; the evidence is secondary and covers prior model generations.
|
|
544
|
+
- **The Coding Agent Index is no longer quoted as a number** — the same model
|
|
545
|
+
reads 80, 78 or 67 depending on index version and harness. It is now cited for
|
|
546
|
+
ordering only, with every observed value and its provenance recorded.
|
|
547
|
+
- **`taste-final` has writing evidence behind it** — previously ordered from
|
|
548
|
+
general and agentic indexes that do not measure prose. Added EQ-Bench Creative
|
|
549
|
+
Writing v3, EQ-Bench Longform, and the Lech Mazur benchmark, read from the
|
|
550
|
+
publishers.
|
|
551
|
+
- **Per-effort cost and throughput** added to the model notes, showing why the
|
|
552
|
+
defaults use `xhigh`: it reaches the same index score as `max` for 30-53% less
|
|
553
|
+
per task.
|
|
554
|
+
|
|
536
555
|
## What's new in v0.10.3
|
|
537
556
|
|
|
538
557
|
- **Restructured READMEs in all five languages** — the reader now meets a plain
|
package/README.zh-CN.md
CHANGED
|
@@ -110,7 +110,7 @@ flowchart LR
|
|
|
110
110
|
| ✒️ taste-final | Claude Opus 5 (high) | GPT-5.6 Sol (max) | 对外文字、prompt 与文档打磨、风格终审 |
|
|
111
111
|
| 💬 consult | 明确指定的厂商/模型 | —(不降级) | 自然语言直接咨询;必须保留 `--vendor` |
|
|
112
112
|
| 🎨 ui-draft | GPT-5.6 Sol (xhigh) | Claude Opus 5 (high) | 有设计规范/参考图时的 UI 出稿;开放式视觉品味交给 taste-final |
|
|
113
|
-
| 📚 long-context | Gemini 3.1 Pro (High) | Claude Opus 5 (high) | 百万 token
|
|
113
|
+
| 📚 long-context | Gemini 3.1 Pro (High) | Claude Opus 5 (high) | 百万 token 扫读与检索;要跨段落多跳整合请改用 Claude 候选,高速重复循环仍优先 Flash |
|
|
114
114
|
| ⚡ fast-agentic | Gemini 3.6 Flash (High) | GPT-5.6 Luna (high) | 快速多步骤 agentic 循环、多模态检查 |
|
|
115
115
|
| 📡 live-search | Grok 4.5 | —(off) | 实时 X/网络搜索与社群脉络 |
|
|
116
116
|
| 🚰 coding-overflow | Grok 4.5 | Kimi K3 → Qwen3 Coder Plus → OpenCode | Codex 额度吃紧时的中量级编码溢流道;事实性声明须另行查证 |
|
|
@@ -462,6 +462,20 @@ plan 模式,或只给只读工具集)。要改文件必须同时给 `--mode work
|
|
|
462
462
|
|
|
463
463
|
## 📜 版本历程
|
|
464
464
|
|
|
465
|
+
## v0.10.4 新功能
|
|
466
|
+
|
|
467
|
+
- **`long-context` 不再把多跳任务指向错的模型**——这条通道原本自称长文*整合*,
|
|
468
|
+
首选却是 Gemini;但公开的百万 token 多针分数显示 Claude 领先约三倍,Gemini
|
|
469
|
+
强的是单针检索。通道说明已改为扫读与检索,跨来源整合请改用 Claude 候选。
|
|
470
|
+
排序刻意不动:那批证据是二手且测的是上一代模型,不足以移动已发布的默认值。
|
|
471
|
+
- **Coding Agent Index 不再被当数字引用**——同一个模型在不同版本与 harness 下
|
|
472
|
+
读出 80、78、67 三种值。现在只用来看排序,并记录每个观测值的出处。
|
|
473
|
+
- **`taste-final` 补上写作专项证据**——此前完全靠不测文笔的通用与 agentic 指标
|
|
474
|
+
排序。已加入 EQ-Bench Creative Writing v3、EQ-Bench Longform 与 Lech Mazur
|
|
475
|
+
三个榜,全部取自发布者第一手。
|
|
476
|
+
- **补上逐档位成本与吞吐**,说明默认为何用 `xhigh`:拿到与 `max` 相同的指数
|
|
477
|
+
分数,每任务却便宜 30-53%。
|
|
478
|
+
|
|
465
479
|
## v0.10.3 新功能
|
|
466
480
|
|
|
467
481
|
- **五种语言的 README 全面重整**——文档开头改成先讲清楚「这是什么、我为什么会
|
package/README.zh-TW.md
CHANGED
|
@@ -110,7 +110,7 @@ flowchart LR
|
|
|
110
110
|
| ✒️ taste-final | Claude Opus 5 (high) | GPT-5.6 Sol (max) | 對外文字、prompt 與文件打磨、風格終審 |
|
|
111
111
|
| 💬 consult | 明確點名的廠商/模型 | —(不降級) | 自然語言直接諮詢;必須保留 `--vendor` |
|
|
112
112
|
| 🎨 ui-draft | GPT-5.6 Sol (xhigh) | Claude Opus 5 (high) | 有設計規範/參考圖時的 UI 出稿;開放式視覺品味交給 taste-final |
|
|
113
|
-
| 📚 long-context | Gemini 3.1 Pro (High) | Claude Opus 5 (high) | 百萬 token
|
|
113
|
+
| 📚 long-context | Gemini 3.1 Pro (High) | Claude Opus 5 (high) | 百萬 token 掃讀與檢索;要跨段落多跳整合請改用 Claude 候選,高速重複迴圈仍優先 Flash |
|
|
114
114
|
| ⚡ fast-agentic | Gemini 3.6 Flash (High) | GPT-5.6 Luna (high) | 快速多步驟 agentic 迴圈、多模態檢查 |
|
|
115
115
|
| 📡 live-search | Grok 4.5 | —(off) | 即時 X/網路搜尋與社群脈絡 |
|
|
116
116
|
| 🚰 coding-overflow | Grok 4.5 | Kimi K3 → Qwen3 Coder Plus → OpenCode | Codex 額度吃緊時的中量級編碼溢流道;事實性宣稱須另行查證 |
|
|
@@ -472,6 +472,20 @@ plan 模式,或只給唯讀工具集)。要改檔必須同時給 `--mode work`
|
|
|
472
472
|
|
|
473
473
|
## 📜 版本歷程
|
|
474
474
|
|
|
475
|
+
## v0.10.4 新功能
|
|
476
|
+
|
|
477
|
+
- **`long-context` 不再把多跳任務指向錯的模型**——這條通道原本自稱長文*整合*,
|
|
478
|
+
首選卻是 Gemini;但公開的百萬 token 多針分數顯示 Claude 領先約三倍,Gemini
|
|
479
|
+
強的是單針檢索。通道說明已改為掃讀與檢索,跨來源整合請改用 Claude 候選。
|
|
480
|
+
排序刻意不動:那批證據是二手且測的是上一代模型,不足以移動已出貨的預設。
|
|
481
|
+
- **Coding Agent Index 不再被當數字引用**——同一個模型在不同版本與 harness 下
|
|
482
|
+
讀出 80、78、67 三種值。現在只用來看排序,並記錄每個觀測值的出處。
|
|
483
|
+
- **`taste-final` 補上寫作專項證據**——先前完全靠不測文筆的通用與 agentic 指標
|
|
484
|
+
排序。已加入 EQ-Bench Creative Writing v3、EQ-Bench Longform 與 Lech Mazur
|
|
485
|
+
三個榜,全部取自發布者第一手。
|
|
486
|
+
- **補上逐檔位成本與吞吐**,說明預設為何用 `xhigh`:拿到與 `max` 相同的指數
|
|
487
|
+
分數,每任務卻便宜 30-53%。
|
|
488
|
+
|
|
475
489
|
## v0.10.3 新功能
|
|
476
490
|
|
|
477
491
|
- **五種語言的 README 全面重整**——文件開頭改成先講清楚「這是什麼、我為什麼會
|
package/VERSION
CHANGED
|
@@ -1 +1 @@
|
|
|
1
|
-
0.10.
|
|
1
|
+
0.10.4
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "omnilane",
|
|
3
|
-
"version": "0.10.
|
|
3
|
+
"version": "0.10.4",
|
|
4
4
|
"description": "One routing table, every harness — classify subtasks into lanes and dispatch each lane to the best vendor's agentic CLI (Codex, Claude, Gemini, Grok) using your existing subscription logins.",
|
|
5
5
|
"bin": {
|
|
6
6
|
"omnilane": "bin/omnilane"
|
package/routing.yaml
CHANGED
|
@@ -24,7 +24,7 @@ hard-judgment: claude claude-opus-5 xhigh | codex gpt-5.6-sol max # AA Intel
|
|
|
24
24
|
taste-final: claude claude-opus-5 high | codex gpt-5.6-sol max # user-facing prose, prompt/doc polish, Chinese phrasing, style arbitration
|
|
25
25
|
consult: codex gpt-5.6-sol max | claude claude-opus-5 high | grok grok-4.5 - | gemini "Gemini 3.1 Pro (High)" - # direct named-model consultation; use --vendor to prevent fallback
|
|
26
26
|
ui-draft: codex gpt-5.6-sol xhigh | claude claude-opus-5 high # only with a design system / reference images; open-ended visual taste -> taste-final
|
|
27
|
-
long-context: gemini "Gemini 3.1 Pro (High)" - | claude claude-opus-5 high | codex gpt-5.6-sol high # all have 1M context;
|
|
27
|
+
long-context: gemini "Gemini 3.1 Pro (High)" - | claude claude-opus-5 high | codex gpt-5.6-sol high # all have 1M context; Gemini leads single-needle retrieval at full length and is the cheapest way to sweep volume, Flash for fast loops. For multi-hop synthesis across a large corpus prefer the claude candidate: published multi-needle scores at 1M favour Claude by a wide margin (see docs/model-capabilities-2026-07.md)
|
|
28
28
|
fast-agentic: gemini "Gemini 3.6 Flash (High)" - | codex gpt-5.6-luna high # speed + agentic tool loops; 3.6 Flash (released 2026-07-21): AA Intelligence 50, #1 output speed 303.6 tok/s, -17% output tokens vs 3.5 Flash per Google
|
|
29
29
|
live-search: grok grok-4.5 - | off # native X/web search lane; no real substitute
|
|
30
30
|
coding-overflow: grok grok-4.5 - | kimi kimi-k3 - | qwen qwen3-coder-plus - | opencode - - | off # codex-quota relief valve: mid-tier coding; Grok 4.5 (GA 2026-07-16) Terminal-Bench 83.3 / SWE-Bench Pro 64.7, but AA hallucination 54% — verify factual claims. qwen3-coder-plus = 2025-09-23 snapshot alias (Qwen 3.6 Plus exists; re-evaluate before swapping). kimi/qwen model fields are CLI aliases — adjust to your login. opencode "-" model = its own configured default.
|