@modular-prompt/experiment 0.4.32 → 0.4.34

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (2) hide show
  1. package/README.md +70 -1
  2. package/package.json +5 -5
package/README.md CHANGED
@@ -69,7 +69,76 @@ npx modular-experiment experiment.yaml --evaluate # 評価付き
69
69
  npx modular-experiment experiment.yaml --repeat 10 # 複数回実行
70
70
  ```
71
71
 
72
- 設定ファイルの詳細、評価器の書き方、プログラマティックAPIについては `skills/experiment/SKILL.md` を参照。
72
+ #### CLIオプション
73
+
74
+ | オプション | 説明 | デフォルト |
75
+ |-----------|------|-----------|
76
+ | `--test-case <name>` | テストケース名フィルター(指定した名前のみ実行) | all |
77
+ | `--model <provider>` | モデルプロバイダーフィルター(例: `mlx`, `vertexai`, `googlegenai`) | すべての有効なモデル |
78
+ | `--modules <names>` | カンマ区切りのモジュール名(指定したモジュールのみテスト) | all |
79
+ | `--repeat <count>` | 実行回数(統計的な評価に有用) | 1 |
80
+ | `--evaluate` | AI評価器を有効化(評価フェーズを実行) | false |
81
+ | `--evaluators <names>` | カンマ区切りの評価器名(指定した評価器のみ使用) | all |
82
+ | `--dry-run` | 実行計画の表示のみ(実験を実行しない) | false |
83
+ | `--verbose` | 詳細なログ出力(内部処理の表示) | false |
84
+ | `--log-file <path>` | 詳細ログのJSONL出力先ファイルパス | なし |
85
+ | `--trace-dir <dir>` | 構造化された実行ログの出力ディレクトリ(prefix/contextごとにファイル分割、summary.json付き) | なし |
86
+ | `--output <path>` | 実験結果のJSON出力先ファイルパス(メタデータと結果を含む) | なし |
87
+
88
+ **ログ・トレースオプション詳細:**
89
+ - `--log-file`: 全実行ログをJSONL形式で1ファイルに出力。各行が1ログエントリ。
90
+ - `--trace-dir`: ログをprefix/context別にファイル分割して出力。summary.jsonに全体統計を含む。読みやすい形式(タイムスタンプ + レベル + メッセージ)。
91
+ - `--output`: 実験結果を構造化JSON形式で保存。タイムスタンプ、使用モデル、繰り返し回数などのメタデータを含む。
92
+
93
+ ## 設定ファイルの詳細
94
+
95
+ ### モデル指定: DriverSet記法
96
+
97
+ テストケースの `models` フィールドでは、単一のモデル名(文字列)または役割別モデルのマッピング(オブジェクト)を指定できます。
98
+
99
+ #### 単一モデル(文字列)
100
+
101
+ すべての役割で同じモデルを使用します。
102
+
103
+ ```yaml
104
+ testCases:
105
+ - name: 基本テスト
106
+ input:
107
+ query: "質問内容"
108
+ models:
109
+ - gpt4o # すべての役割でgpt4oを使用
110
+ ```
111
+
112
+ #### DriverSet(役割別モデル)
113
+
114
+ 役割ごとに異なるモデルを指定できます。`default` は必須で、他の役割(`thinking`、`instruct`、`chat`、`plan`)はオプションです。未指定の役割は自動的に `default` にフォールバックします。
115
+
116
+ ```yaml
117
+ models:
118
+ gemma4:
119
+ provider: mlx
120
+ model: gemma4-26b-a4b
121
+ qwen:
122
+ provider: mlx
123
+ model: qwen3.5-9b
124
+
125
+ testCases:
126
+ - name: 役割別モデルテスト
127
+ input:
128
+ query: "質問内容"
129
+ models:
130
+ - default: gemma4 # 通常のタスクはgemma4
131
+ thinking: qwen # thinking役割はqwen
132
+ ```
133
+
134
+ **利用可能なModelRole:**
135
+ - `default`: 必須。メインのモデル
136
+ - `thinking`: 推論タスク用(オプション)
137
+ - `instruct`: 指示実行用(オプション)
138
+ - `chat`: 対話用(オプション)
139
+ - `plan`: 計画立案用(オプション)
140
+
141
+ **注記:** 値には設定ファイルのトップレベル `models:` セクションで定義されたモデル名を指定します。
73
142
 
74
143
  ## Skills (for Claude Code)
75
144
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@modular-prompt/experiment",
3
- "version": "0.4.32",
3
+ "version": "0.4.34",
4
4
  "description": "Experiment framework for comparing and evaluating prompt modules",
5
5
  "type": "module",
6
6
  "main": "./dist/index.js",
@@ -25,10 +25,10 @@
25
25
  "jiti": "2.6.1",
26
26
  "yaml": "2.8.2",
27
27
  "zod": "3.25.76",
28
- "@modular-prompt/core": "0.2.2",
29
- "@modular-prompt/process": "0.5.0",
30
- "@modular-prompt/driver": "0.12.0",
31
- "@modular-prompt/utils": "0.3.4"
28
+ "@modular-prompt/driver": "0.13.1",
29
+ "@modular-prompt/process": "0.5.2",
30
+ "@modular-prompt/core": "0.3.0",
31
+ "@modular-prompt/utils": "0.3.5"
32
32
  },
33
33
  "devDependencies": {
34
34
  "@eslint/js": "9.39.2",