omnilane 0.10.2 → 0.10.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -6,6 +6,81 @@ semantic version tags.
6
6
 
7
7
  ## [Unreleased]
8
8
 
9
+ ## [0.10.4] - 2026-07-26
10
+
11
+ No lane ordering changes. This release corrects documentation that could send
12
+ work to the wrong model, and records the evidence behind the shipped defaults.
13
+
14
+ ### Changed
15
+
16
+ - Narrowed the `long-context` lane description in `routing.yaml` and in all five
17
+ localized README lane tables. It previously called the lane long-document
18
+ *synthesis* while shipping Gemini first; published multi-needle scores at 1M
19
+ favour Claude by roughly threefold, while Gemini leads single-needle
20
+ retrieval. These are different capabilities with different leaders, and the
21
+ old wording pointed multi-hop work at the wrong candidate. The lane now
22
+ describes retrieval and volume sweeps, and names the Claude candidate for
23
+ integration across scattered sources. Ordering is deliberately unchanged: the
24
+ supporting evidence is secondary and covers prior model generations, which is
25
+ not a sufficient basis for moving a shipped default.
26
+
27
+ ### Fixed
28
+
29
+ - `docs/model-capabilities-2026-07.md` quoted the Artificial Analysis Coding
30
+ Agent Index at v1.1 while the index had re-based twice more. The same model
31
+ reads 80, 78 or 67 depending on the version and harness a source used, so the
32
+ figure is not portable across versions. The section now records every observed
33
+ value with its provenance, documents the v1.3 composition, and states that the
34
+ index may be cited for ordering but never for a number.
35
+ - Documented the writing evidence behind `taste-final`, which previously rested
36
+ entirely on general and agentic indexes that do not measure prose. Added
37
+ EQ-Bench Creative Writing v3, EQ-Bench Longform, and the Lech Mazur
38
+ story-writing benchmark, each read from the publisher. Recorded that Mazur has
39
+ not yet evaluated Claude Opus 5 and that Claude Fable 5 leads that board, so
40
+ the question stays open rather than being presented as settled.
41
+ - Added per-effort cost and throughput to the Intelligence Index table, which is
42
+ what justifies the shipped effort levels: `xhigh` reaches the same index score
43
+ as `max` for 30% (Opus 5) and 53% (GPT-5.6 Sol) less per task.
44
+
45
+ ## [0.10.3] - 2026-07-26
46
+
47
+ ### Changed
48
+
49
+ - Restructured all five localized READMEs for first-time readers: a plain
50
+ "what is this and why would I want it" section now opens the document,
51
+ version history is consolidated into a single section at the bottom instead
52
+ of interrupting the introduction, and a new FAQ answers the recurring
53
+ questions — whether every subscription is required, where code is sent, why
54
+ Claude Fable 5 is absent from the defaults, why the Claude lanes use `xhigh`
55
+ rather than `max`, what happens when a lane's first-choice CLI is missing,
56
+ and whether a dispatched worker can edit files.
57
+ - Documented the Claude Fable 5 routing decision with measurements rather than
58
+ cost alone: Artificial Analysis calls Opus 5 (61) and Fable 5 (60)
59
+ "effectively tied" on the Intelligence Index and Epoch AI ranks them the
60
+ other way, but Opus 5 leads AA-Briefcase by 146 Elo and GDPval-AA v2 by 114
61
+ Elo at 20% lower cost per task, so the top Claude tier stays a main-loop
62
+ choice rather than a dispatch target.
63
+
64
+ ### Fixed
65
+
66
+ - `plugin.json` and `.claude-plugin/plugin.json` still advertised `0.10.0`
67
+ after the 0.10.1 and 0.10.2 releases, so plugin installations reported a
68
+ stale version. Both now track `VERSION`, and the release version test that
69
+ covers them passes again.
70
+ - `routing.local.yaml.example` shipped starter profiles pinned to retired
71
+ models: seven `claude-opus-4-8` references now use `claude-opus-5` with
72
+ lane-appropriate effort, and the Gemini 3.5 Flash candidates move to 3.6
73
+ Flash, matching the defaults since 0.10.0.
74
+ - Corrected the Artificial Analysis Intelligence Index figures in
75
+ `docs/model-capabilities-2026-07.md` against the source article (index
76
+ points, not percentages; Fable 5 60 and GPT-5.6 Sol 59), added the
77
+ AA-Briefcase / GDPval-AA v2 comparison, and recorded the two results that
78
+ cut against the defaults: Fable 5 leads on factual knowledge, and GPT-5.6 Sol
79
+ leads on presentation quality.
80
+ - Replaced the stale "Anthropic positions Fable 5 above Opus 4.8" note in
81
+ `routing.yaml` and `skills/omnilane/SKILL.md`, which had not been updated
82
+ when Opus 5 shipped.
83
+
9
84
  ## [0.10.2] - 2026-07-26
10
85
 
11
86
  ### Changed
@@ -372,7 +447,9 @@ semantic version tags.
372
447
  - Initial shared routing table, cross-vendor dispatcher, runners, installer,
373
448
  and baseline lint fixes.
374
449
 
375
- [Unreleased]: https://github.com/Seraphim0916/omnilane/compare/v0.10.2...HEAD
450
+ [Unreleased]: https://github.com/Seraphim0916/omnilane/compare/v0.10.4...HEAD
451
+ [0.10.4]: https://github.com/Seraphim0916/omnilane/compare/v0.10.3...v0.10.4
452
+ [0.10.3]: https://github.com/Seraphim0916/omnilane/compare/v0.10.2...v0.10.3
376
453
  [0.10.2]: https://github.com/Seraphim0916/omnilane/compare/v0.10.1...v0.10.2
377
454
  [0.10.1]: https://github.com/Seraphim0916/omnilane/compare/v0.9.1...v0.10.1
378
455
  [0.10.0]: https://github.com/Seraphim0916/omnilane/compare/v0.9.1...1ea55e5
package/README.ja.md CHANGED
@@ -21,48 +21,30 @@
21
21
 
22
22
  ---
23
23
 
24
- ## 2026-07-25 更新
24
+ ## 🤔 omnilane とは
25
25
 
26
- - デフォルトルーティングに `claude-opus-5` を追加しました。`hard-judgment` と `taste-final` の第一候補となり、最難関のコーディング作業でもフォールバックとして利用できます。
27
- - `omnilane configure` を全 13 プロバイダーへ拡張しました。選択可能なモデルは 106 件で、CodexClaude Code、Grok Build、Antigravity の最新カタログに加え、検証済みの OpenRouter/OpenCode ショートカットを収録しています。`c` によるカスタムモデル ID の入力も引き続き利用できます。
26
+ **何が問題か。** すでに AI コーディングアシスタント——**Claude Code、Codex、
27
+ CursorGemini CLI** など——を使っていますよね。どれも一つのモデルファミリー
28
+ としかやり取りしません。つまり、頼んだ作業はすべて同じモデルで走ります。適任か
29
+ どうかに関係なく——使い捨てのファイル名変更が最も高価なモデルを消費し、本当に
30
+ 難しい設計上の問いは、たまたま開いていたモデルに当たります。
28
31
 
29
- ---
30
- ## 👋 はじめての方へ
31
-
32
- すでに AI コーディングアシスタント——**Claude Code、Codex、Cursor、Gemini
33
- CLI** など——を使っていますよね。どれも一度に一つの AI モデルとやり取りし、
34
- 「どのタスクにどのモデルが最適か」の判断はあなた任せです。
35
-
36
- **omnilane がその判断を代行します。** 一つひとつの作業を、それに最も強い(そして最も安い)
37
- モデルへ自動で振り分けます——難しいコーディングはトップコーダーへ、ちょっとした確認は
38
- 速くて安いモデルへ、長い文書は大コンテキストモデルへ——すべてあなたが既に契約している
39
- サブスクと API キーで。組み込みのデフォルトのまま使うも、小さな設定ファイルを一つ調整するも
40
- 自由。新しく面倒を見る対象は増えず(既存ツールの裏で動きます)、`./install.sh --uninstall`
41
- できれいに削除できます。
42
-
43
- **[⬇ 60 秒クイックスタートへ](#-60-秒クイックスタート)**
44
-
45
- ## v0.10.0 の新機能
32
+ **omnilane がすること。** アシスタントにルーティング表を渡します。作業は
33
+ **レーン**——最難関のコーディング、機械的な物量、トリアージ、難しい判断、
34
+ 最終的な仕上げ——に振り分けられ、各レーンにはそれに最も強く(そして最も安い)
35
+ モデルが指定されています。アシスタントは自分の得意なレーンを自分で処理し、
36
+ 残りは既存のログインを使って別ベンダーの CLI にバックグラウンドで渡します。
46
37
 
47
- - **Gemini 3.6 Flash を既定に** — `fast-agentic`・`triage`・`bulk-mechanical`
48
- の gemini 候補(および `Gemini Flash` エイリアス)を 2026-07-21 リリースの
49
- Gemini 3.6 Flash に更新。出力トークンが減り、出力単価も下がり、Artificial
50
- Analysis 計測の出力速度は首位です。
51
- - **エビデンス再監査** — ルーティングのコメント、モデル能力ノート、Gemini
52
- 価格表を公式ソースに合わせて更新(2026-07-21/22)。
53
-
54
- ## v0.9.1 の新機能
55
-
56
- - **修正**: `configure set` が `routing.local.yaml` の手書きコメントを削除しなく
57
- なりました。書き換えるのは自身のスタンプ行と置き換え対象のレーンだけです。
38
+ **omnilane ではないもの。** プロキシでも、新しいサブスクでも、常時面倒を見る
39
+ サービスでもありません。表一枚とディスパッチスクリプト一本が、既存ツールの裏で
40
+ 動くだけです。`./install.sh --uninstall` で痕跡なく削除できます。
58
41
 
59
- ## v0.9.0 の新機能
42
+ **すべてのサブスクは不要です。** 各レーンはフォールバックチェーンであり、実際に
43
+ インストール済みの最初の候補が選ばれます。CLI が一つでも七つでも動作し、どれも
44
+ 無いレーンは失敗せず単にオフになります。サブスク一つでも、デフォルト表はその
45
+ ベンダーに収束します。
60
46
 
61
- - **OpenAI 互換の direct-API ベンダーを 5 つ追加** — `deepseek`、`zai`(GLM)
62
- `mistral`、`groq`、`cerebras` が `openrouter` と同じく CLI 不要のレーンに
63
- (curl と `<VENDOR>_API_KEY` だけ)。`lib/common.sh` のレジストリに 1 行で
64
- 追加でき、モデル能力の比較は [`docs/model-capabilities-2026-07.md`](docs/model-capabilities-2026-07.md) を参照。
65
- - **Fish シェル補完** — `omnilane completion fish | source`。
47
+ **[⬇ 60 秒クイックスタートへ](#-60-秒クイックスタート)** · **[❓ FAQ を読む](#-faq)**
66
48
 
67
49
  ## ⚡ 60 秒クイックスタート
68
50
 
@@ -137,22 +119,18 @@ flowchart LR
137
119
  | ✒️ taste-final | Claude Opus 5 (high) | GPT-5.6 Sol (max) | 対外文章、prompt/ドキュメント推敲、スタイル最終審 |
138
120
  | 💬 consult | 明示指定したベンダー/モデル | —(フォールバックなし) | 自然言語で直接相談。`--vendor` を必ず維持 |
139
121
  | 🎨 ui-draft | GPT-5.6 Sol (xhigh) | Claude Opus 5 (high) | デザインシステム/参考画像がある場合の UI ドラフト |
140
- | 📚 long-context | Gemini 3.1 Pro (High) | Claude Opus 5 (high) | 100 万トークン級の長文統合。Pro agentic 対応、高速反復ループは Flash を優先 |
122
+ | 📚 long-context | Gemini 3.1 Pro (High) | Claude Opus 5 (high) | 100 万トークン級の走査と検索。複数箇所をまたぐ統合には Claude 候補を、高速反復ループは Flash を優先 |
141
123
  | ⚡ fast-agentic | Gemini 3.6 Flash (High) | GPT-5.6 Luna (high) | 高速なマルチステップ agentic ループ、マルチモーダル確認 |
142
124
  | 📡 live-search | Grok 4.5 | —(off) | リアルタイム X/ウェブ検索とソーシャル文脈 |
143
125
  | 🚰 coding-overflow | Grok 4.5 | Kimi K3 → Qwen3 Coder Plus → OpenCode | Codex クォータ逼迫時の中級コーディング逃し弁 |
144
126
  | 🗳️ arbitrate | off(オプトイン) | — | 内蔵オピニオンパネル(重大な判断用)——デフォルト無効。`routing.local.yaml` で有効化;投票者×ラウンドごとに 1 コール消費 |
145
127
 
146
128
  **バックアップ**はチェーンの次の候補——第一候補のベンダー CLI が未インストールの
147
- ときにディスパッチが降格する先です。
129
+ ときにディスパッチが降格する先です。どのレーンもこうしたチェーンで、チェーン内に
130
+ 何もインストールされていなければレーンは `off` に降格します。
148
131
 
149
- > **Claude Fable 5 はどこ?** 意図的にデフォルト表に入れていません:Claude の
150
- > 最上位ティアは通常*メインループ自身*であり、ディスパッチされるワーカーでは
151
- > ないため(価格も Opus より上)。これはコスト/ガードレール/メインループ方針で
152
- > あり、能力評価ではありません。Anthropic は Fable 5 を Opus 5 より上位に
153
- > 位置付けています。設定メニューのモデル一覧には載っているので、
154
- > 使いたければ自分でルーティングできます(例:`routing.local.yaml` に
155
- > `taste-final: claude claude-fable-5 high`)。
132
+ > **Claude Fable 5 はどこ?** 意図的にデフォルト表に入れていません——理由と
133
+ > 実測データは [FAQ](#-faq) にまとめてあります。
156
134
 
157
135
  ### 自然言語コンサルテーション
158
136
 
@@ -377,12 +355,139 @@ configure.sh set|get|unset|list|diff LANE [SPEC] # routing.local.yaml を非
377
355
  記録し、`jobs.sh status` が `dead` を報告。
378
356
  - **ペイロード上限** — 巨大なタスクテキストは自動で頭尾トランケート。
379
357
 
358
+ ## ❓ FAQ
359
+
360
+ <details>
361
+ <summary><b>これらのサブスクは全部必要ですか?</b></summary>
362
+
363
+ <br/>
364
+
365
+ いいえ。各レーンはフォールバックチェーンで、実際にインストール済みの最初の候補が
366
+ 使われます。サブスクが一つなら表全体がそのベンダーに収束し、チェーン内に何も無い
367
+ レーンはエラーではなく単にオフになります。`omnilane doctor` で今このマシンが実際に
368
+ 到達できる先が分かり、`routing.local.yaml.example` にはよくある状況向けの
369
+ スタータープロファイル(Claude のみ、Codex 中心、Codex 無し)が入っています。
370
+
371
+ </details>
372
+
373
+ <details>
374
+ <summary><b>omnilane は私のコードを新しい送信先へ送りますか?</b></summary>
375
+
376
+ <br/>
377
+
378
+ 新しい送信先は増えません。ディスパッチはすでにインストールしログイン済みの
379
+ ベンダー CLI を呼ぶだけなので、コードが届くのは元から使っているベンダーだけです。
380
+ ランナーはサブスク制 CLI を呼ぶ前に API キーの環境変数を取り除くため、残っていた
381
+ キーによって従量課金へ黙って切り替わることもありません。唯一の例外は direct-API
382
+ ベンダー群(`openrouter`、`deepseek`、`zai`、`mistral`、`groq`、`cerebras`)で、
383
+ これらは定義上あなたが設定したキーでそのプロバイダーの API を呼びます——いずれも
384
+ advise 専用で、ファイルを編集しません。
385
+
386
+ </details>
387
+
388
+ <details>
389
+ <summary><b>Claude Fable 5 はどこ?なぜデフォルト表に無いのですか?</b></summary>
390
+
391
+ <br/>
392
+
393
+ **Claude の最上位ティアは通常メインループ自身であり、ディスパッチされるワーカー
394
+ ではないからです。** レーンは「今あなたが動かしているモデル以外」に作業を送るために
395
+ あります。Fable 5 がメインループなら、判断や文章を Fable 5 に戻すのは呼び出しが
396
+ 一回増えるだけで得るものがありません——だからこそ上の「メインモデルを選ぶ」一覧では
397
+ Fable 5 に**ドライバー**として独立した行があり、hard-judgment、taste-final、
398
+ 正確性が要の最難関修正を自分で処理します。
399
+
400
+ **計測データもワーカーとしての採用を支持しません。** Artificial Analysis の
401
+ Intelligence Index(2026-07-24)では Opus 5(max)が 61、Fable 5(max)が 60 —— AA 自身が
402
+ 「実質的に同点」と表現し、Epoch AI の Capability Index は順位が逆です
403
+ (Fable 5 161、Opus 5 159)。総合的な知能は引き分けと見てよいでしょう。差が付くのは
404
+ エージェント的な専門アウトプットで、その差は小さくありません:
405
+
406
+ | ベンチマーク | Claude Opus 5 (max) | Claude Fable 5 | |
407
+ |---|---:|---:|---|
408
+ | AA-Briefcase(エージェント的知識労働、Elo) | 1720 | 1574 | **+146** |
409
+ | GDPval-AA v2(Elo) | 1861 | 1747 | **+114** |
410
+ | AA-Briefcase の 1 タスク単価 | $17.79 | $22.30 | **-20%** |
411
+ | API 価格、入力/出力 1M あたり | $5 / $25 | $10 / $50 | **半額** |
412
+
413
+ Opus 5 の max・xhigh・high の三ティアが AA-Briefcase の上位三席を占め、`high`
414
+ ティアですら 1 タスク単価が半分以下で Fable 5 を上回ります。つまり Fable 5 は
415
+ 価格が二倍でありながら、レーンが重視するどの軸でも優位を買えていません。
416
+
417
+ **Fable 5 が実際に優れている点**:知識の広さです。AA-Omniscience では今も Opus 5 を
418
+ 上回っており(モデルのサイズクラスから見て妥当)、一方の Opus 5 は不確かなときでも
419
+ 答えがちで、ハルシネーション率は 50%(Opus 4.8 比 +14 ポイント)です。実行より
420
+ 想起が中心のタスクなら、明示的に指名してください:
421
+
422
+ ```bash
423
+ dispatch.sh --vendor claude --model claude-fable-5 --effort high consult "…"
424
+ ```
425
+
426
+ **これはコストとメインループの方針上の選択であり、能力評価ではありません。**
427
+ Fable 5 は設定メニューのモデル一覧に載っており、`routing.local.yaml` の 1 行で
428
+ デフォルトを上書きできます:
429
+
430
+ ```yaml
431
+ taste-final: claude claude-fable-5 high
432
+ ```
433
+
434
+ </details>
435
+
436
+ <details>
437
+ <summary><b>Claude のレーンはなぜ <code>max</code> ではなく <code>xhigh</code> なのですか?</b></summary>
438
+
439
+ <br/>
440
+
441
+ 努力度は高ければ高いほど良い、というものではないからです。Anthropic は `xhigh` を
442
+ コーディングとエージェント作業の出発点、`high` をそれ以外の知能を要する作業の下限、
443
+ `max` を正確性がコストに優先する場合の設定として文書化しています。第三者の計測も
444
+ 一致しており、Vals.ai の Vibe Code Bench では Opus 5 は `high` で 89.8%、`xhigh` で
445
+ 88.3%、`max` で 88.4% —— 上位ティアはより手の込んだ解を出し、その分だけ失敗も増えます。
446
+ ワークロードが違うならレーン単位で上げてください:
447
+
448
+ ```bash
449
+ omnilane configure set hard-judgment "claude claude-opus-5 max"
450
+ ```
451
+
452
+ </details>
453
+
454
+ <details>
455
+ <summary><b>レーンの第一候補 CLI が無いときはどうなりますか?</b></summary>
456
+
457
+ <br/>
458
+
459
+ ディスパッチはチェーンを辿り、手元にある最初のベンダーを使います。コールを消費
460
+ せずに判断を確認できます:
461
+
462
+ ```bash
463
+ scripts/dispatch.sh --explain hardest-coding # 候補ごとのトレース
464
+ scripts/dispatch.sh --list # 実効表の全体
465
+ scripts/dispatch.sh --dry-run hardest-coding "…" # 解決済みプラン、プロバイダー呼び出し無し
466
+ ```
467
+
468
+ </details>
469
+
470
+ <details>
471
+ <summary><b>ディスパッチされたワーカーはファイルを編集できますか?</b></summary>
472
+
473
+ <br/>
474
+
475
+ 依頼した場合のみです。ディスパッチの既定は読み取り専用の `advise` で、ベンダーごとに
476
+ 実装されています(読み取り専用サンドボックス、plan モード、あるいは読み取り専用の
477
+ ツールセット)。編集には `--mode work` と明示的な `--workdir` の両方が必要です。
478
+ ワーカー自身は再ディスパッチできません——深度ガードが終了コード 86 で入れ子の
479
+ ファンアウトを拒否するため、一つのコマンドがエージェントの連鎖に膨らんでクォータを
480
+ 食い潰すことはありません。
481
+
482
+ </details>
483
+
380
484
  ## 📊 デフォルト値と出典
381
485
 
382
486
  デフォルトのレーン割当は Artificial Analysis の 2026-07 スナップショット
383
487
  (AA サイトの生レコードと各社公式価格ページで照合済み)と公開の比較レビューに
384
488
  基づきます。これは意見であって法則ではありません——設定メニューと
385
- `routing.local.yaml` はそのためにあります。
489
+ `routing.local.yaml` はそのためにあります。ベンチマークごとの但し書きを含む
490
+ 作業ノートは [`docs/model-capabilities-2026-07.md`](docs/model-capabilities-2026-07.md) にあります。
386
491
 
387
492
  ## ⚠️ 既知の制限
388
493
 
@@ -398,8 +503,84 @@ configure.sh set|get|unset|list|diff LANE [SPEC] # routing.local.yaml を非
398
503
 
399
504
  ## 📜 リリース履歴
400
505
 
506
+ ## v0.10.4 の新機能
507
+
508
+ - **`long-context` がマルチホップ作業を誤ったモデルに向けなくなりました** — この
509
+ レーンは長文*統合*を名乗りながら Gemini を第一候補にしていましたが、公開されて
510
+ いる 1M トークンのマルチニードル評価では Claude が約 3 倍のスコアを示し、Gemini
511
+ が強いのはシングルニードル検索です。レーンの説明を走査と検索に改め、複数箇所を
512
+ またぐ統合には Claude 候補を案内します。順序は意図的に据え置き — 根拠が二次情報
513
+ であり、前世代モデルの測定に基づくためです。
514
+ - **Coding Agent Index を数値として引用しなくなりました** — 同一モデルがバージョン
515
+ とハーネス次第で 80、78、67 と読み取れます。今後は順序の参照のみに用い、観測値
516
+ ごとの出所を記録しています。
517
+ - **`taste-final` に文章特化の根拠を追加** — 従来は散文を測らない汎用・エージェント
518
+ 指標のみで順序を決めていました。EQ-Bench Creative Writing v3、EQ-Bench Longform、
519
+ Lech Mazur の 3 ボードを、いずれも公開元から直接取得して追加しました。
520
+ - **努力度ごとのコストとスループットを追加**。既定が `xhigh` である理由を示します:
521
+ `max` と同じ指数スコアを、1 タスクあたり 30-53% 安く得られます。
522
+
523
+ ## v0.10.3 の新機能
524
+
525
+ - **5 言語すべての README を再構成** — 冒頭で「これは何か、なぜ欲しくなるのか」を
526
+ まず説明し、バージョン履歴は導入部を分断せず最下部にまとめました。さらに FAQ を
527
+ 新設し、繰り返し寄せられた疑問に答えています:サブスクは全部必要か、コードは
528
+ どこへ送られるか、なぜ Fable 5 がデフォルト表に無いのか、なぜ `max` ではなく
529
+ `xhigh` なのか、CLI が無いときはどうなるか、ワーカーはファイルを編集できるのか。
530
+ - **修正: プラグインマニフェストのバージョンが古いままだった** — `plugin.json` と
531
+ `.claude-plugin/plugin.json` が 0.10.1・0.10.2 リリース後も `0.10.0` を表示して
532
+ いたため、プラグインのインストールで誤ったバージョンが出ていました。
533
+ - **修正: `routing.local.yaml.example` が退役モデルを指していた** — スタータープロ
534
+ ファイル内の `claude-opus-4-8` をすべて `claude-opus-5`(レーンに応じた努力度付き)
535
+ に、Gemini 3.5 Flash 候補を 3.6 Flash に更新し、0.10.0 以降のデフォルトと揃えました。
536
+ - **Intelligence Index の数値を出典に照合して訂正**
537
+ (`docs/model-capabilities-2026-07.md`):パーセントではなく指数ポイントです。
538
+ AA-Briefcase / GDPval-AA v2 の比較を追加し、デフォルトと逆向きの二つの結果も
539
+ 記録しました:事実知識では Fable 5 が、プレゼン品質では GPT-5.6 Sol が上回ります。
540
+
541
+ ## v0.10.2 の新機能
542
+
543
+ - **`hardest-coding` と `hard-judgment` の Claude 努力度を `max` から `xhigh` へ**。
544
+ Claude Opus 5 に関する Anthropic の文書化された指針に合わせました:コーディングと
545
+ エージェント作業は `xhigh` から始め、それ以外の知能を要する作業の下限は `high`、
546
+ `max` は正確性がコストに優先する場合に限る、というものです。戻す場合は
547
+ `omnilane configure set <lane> "<spec>"` を使ってください。
548
+ - **CHANGELOG の壊れた比較リンクを 2 件修正** — 公開されなかった `v0.10.0` タグを
549
+ 指していました。
550
+
551
+ ## v0.10.1 の新機能
552
+
553
+ - **デフォルトルーティングに `claude-opus-5` を追加**。`hard-judgment` と
554
+ `taste-final` の第一候補となり、最難関のコーディング作業でもフォールバックとして
555
+ 利用できます。
556
+ - **`omnilane configure` を全 13 プロバイダーへ拡張**。選択可能なモデルは 106 件で、
557
+ Codex、Claude Code、Grok Build、Antigravity の最新カタログに加え、検証済みの
558
+ OpenRouter/OpenCode ショートカットを収録しています。`c` によるカスタムモデル ID の
559
+ 入力も引き続き利用できます。
560
+
401
561
  <details>
402
- <summary>過去のリリース(v0.8.3 以前)</summary>
562
+ <summary>過去のリリース(v0.10.0 以前)</summary>
563
+
564
+ ## v0.10.0 の新機能
565
+
566
+ - **Gemini 3.6 Flash を既定に** — `fast-agentic`・`triage`・`bulk-mechanical`
567
+ の gemini 候補(および `Gemini Flash` エイリアス)を Gemini 3.6 Flash に更新。
568
+ 出力トークンが減り、出力単価も下がり、Artificial Analysis 計測の出力速度は首位です。
569
+ - **エビデンス再監査** — ルーティングのコメント、モデル能力ノート、Gemini
570
+ 価格表を公式ソースに合わせて更新。
571
+
572
+ ## v0.9.1 の新機能
573
+
574
+ - **修正**: `configure set` が `routing.local.yaml` の手書きコメントを削除しなく
575
+ なりました。書き換えるのは自身のスタンプ行と置き換え対象のレーンだけです。
576
+
577
+ ## v0.9.0 の新機能
578
+
579
+ - **OpenAI 互換の direct-API ベンダーを 5 つ追加** — `deepseek`、`zai`(GLM)、
580
+ `mistral`、`groq`、`cerebras` が `openrouter` と同じく CLI 不要のレーンに
581
+ (curl と `<VENDOR>_API_KEY` だけ)。`lib/common.sh` のレジストリに 1 行で
582
+ 追加でき、モデル能力の比較は [`docs/model-capabilities-2026-07.md`](docs/model-capabilities-2026-07.md) を参照。
583
+ - **Fish シェル補完** — `omnilane completion fish | source`。
403
584
 
404
585
  ## v0.8.3 の新機能
405
586