@molecule/api-resource-ai-models 1.4.0 → 1.5.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -3,7 +3,7 @@ AUTO-GENERATED — DO NOT EDIT THIS FILE.
3
3
  Generated by `mlcl sync-docs` from the package's src/index.ts JSDoc + mlcl/registry.json.
4
4
  Edits here are overwritten on the next commit (molecule's pre-commit hook regenerates).
5
5
  To change this document, edit the module-level JSDoc in src/index.ts.
6
- Generated: 2026-09-01T19:40:01.931Z
6
+ Generated: 2026-09-06T11:38:29.688Z
7
7
  -->
8
8
 
9
9
  # @molecule/api-resource-ai-models
@@ -846,13 +846,22 @@ Sources (verified 2026-07-28; OpenAI re-verified 2026-07-31 after the
846
846
  at the post-promo list rate ($1.50/$7.50, cache read $0.15) under the same
847
847
  never-under-charge policy, each with a KNOWN_DIVERGENCES entry expiring
848
848
  2026-12-31. gemini-3.5-flash carries no promo: still $1.50/$9.00, $0.15)
849
+ (re-verified 2026-09-02: gemini-3.8-flash "New Stable" — supersedes
850
+ 3.7-flash, which the models page now calls "previous-generation". THIRD id
851
+ on the same flash-tier promo row, verbatim "$0.75 through December 31,
852
+ 2026. $1.50 starting January 1, 2027." / "$3.75 … $7.50" / "$0.075 …
853
+ $0.15", so it too is cataloged at list with a KNOWN_DIVERGENCES entry.
854
+ Specs from /docs/models/gemini-3.8-flash: 1M ctx / 65,536 out, thinking
855
+ low|medium|high, vision, tools, caching, search grounding, code execution,
856
+ url context — identical surface to 3.7-flash. Knowledge cutoff is published
857
+ by neither Google nor models.dev for this id)
849
858
  - xAI: https://docs.x.ai/developers/models + /developers/grok-4-5
850
859
  (grok-4.5 flagship 2026-07-08: $2/$6, 500K ctx, ≥200K prompts bill 2× —
851
860
  not modeled; reasoning_effort low|medium|high default high, image input;
852
861
  grok-4.3 still served at $1.25/$2.50 with the bigger 1M window;
853
862
  grok-code-fast-1 no longer listed — retires 2026-08-15)
854
863
  - DeepSeek: https://api-docs.deepseek.com/quick_start/pricing (verified
855
- 2026-08-31; legacy deepseek-chat/-reasoner ids fully retired 2026-07-24 —
864
+ 2026-09-06; legacy deepseek-chat/-reasoner ids fully retired 2026-07-24 —
856
865
  never in this catalog. The V4-Pro-GA price RISE effective 2026-08-16T16:00Z
857
866
  has LANDED and is folded into the base fields, along with the peak-hour 2×
858
867
  the same card introduced; every rate re-read on the card 2026-08-31 and
@@ -861,7 +870,10 @@ Sources (verified 2026-07-28; OpenAI re-verified 2026-07-31 after the
861
870
  where on 2026-08-18 it carried no day qualifier at all, so the windows are
862
871
  `daysOfWeekUtc`-restricted rather than daily. `deepseek-v4-flash-vision-exp`
863
872
  also appears on the card at flash's rates: EXPERIMENTAL and vision-only-new,
864
- deliberately not catalogued.)
873
+ deliberately not catalogued. The 2026-09-06 re-read found every native rate
874
+ unchanged; what moved was the US re-host — DeepInfra cut
875
+ `DeepSeek-V4-Flash-0731` from $0.08 to $0.06 input, cache read $0.016 to
876
+ $0.015, output held at $0.18. See that entry's `regionPricing`.)
865
877
  - Moonshot: https://platform.kimi.ai/docs/models + DeepInfra's model API for
866
878
  the US re-host (kimi-k3 flagship 2026-07-16
867
879
  — 2.8T MoE, 1M ctx, $3/$15 — NOT added: thinking is forced-on with
package/dist/models.d.ts CHANGED
@@ -114,13 +114,22 @@ import type { ModelDefinition } from './types.js';
114
114
  * at the post-promo list rate ($1.50/$7.50, cache read $0.15) under the same
115
115
  * never-under-charge policy, each with a KNOWN_DIVERGENCES entry expiring
116
116
  * 2026-12-31. gemini-3.5-flash carries no promo: still $1.50/$9.00, $0.15)
117
+ * (re-verified 2026-09-02: gemini-3.8-flash "New Stable" — supersedes
118
+ * 3.7-flash, which the models page now calls "previous-generation". THIRD id
119
+ * on the same flash-tier promo row, verbatim "$0.75 through December 31,
120
+ * 2026. $1.50 starting January 1, 2027." / "$3.75 … $7.50" / "$0.075 …
121
+ * $0.15", so it too is cataloged at list with a KNOWN_DIVERGENCES entry.
122
+ * Specs from /docs/models/gemini-3.8-flash: 1M ctx / 65,536 out, thinking
123
+ * low|medium|high, vision, tools, caching, search grounding, code execution,
124
+ * url context — identical surface to 3.7-flash. Knowledge cutoff is published
125
+ * by neither Google nor models.dev for this id)
117
126
  * - xAI: https://docs.x.ai/developers/models + /developers/grok-4-5
118
127
  * (grok-4.5 flagship 2026-07-08: $2/$6, 500K ctx, ≥200K prompts bill 2× —
119
128
  * not modeled; reasoning_effort low|medium|high default high, image input;
120
129
  * grok-4.3 still served at $1.25/$2.50 with the bigger 1M window;
121
130
  * grok-code-fast-1 no longer listed — retires 2026-08-15)
122
131
  * - DeepSeek: https://api-docs.deepseek.com/quick_start/pricing (verified
123
- * 2026-08-31; legacy deepseek-chat/-reasoner ids fully retired 2026-07-24 —
132
+ * 2026-09-06; legacy deepseek-chat/-reasoner ids fully retired 2026-07-24 —
124
133
  * never in this catalog. The V4-Pro-GA price RISE effective 2026-08-16T16:00Z
125
134
  * has LANDED and is folded into the base fields, along with the peak-hour 2×
126
135
  * the same card introduced; every rate re-read on the card 2026-08-31 and
@@ -129,7 +138,10 @@ import type { ModelDefinition } from './types.js';
129
138
  * where on 2026-08-18 it carried no day qualifier at all, so the windows are
130
139
  * `daysOfWeekUtc`-restricted rather than daily. `deepseek-v4-flash-vision-exp`
131
140
  * also appears on the card at flash's rates: EXPERIMENTAL and vision-only-new,
132
- * deliberately not catalogued.)
141
+ * deliberately not catalogued. The 2026-09-06 re-read found every native rate
142
+ * unchanged; what moved was the US re-host — DeepInfra cut
143
+ * `DeepSeek-V4-Flash-0731` from $0.08 to $0.06 input, cache read $0.016 to
144
+ * $0.015, output held at $0.18. See that entry's `regionPricing`.)
133
145
  * - Moonshot: https://platform.kimi.ai/docs/models + DeepInfra's model API for
134
146
  * the US re-host (kimi-k3 flagship 2026-07-16
135
147
  * — 2.8T MoE, 1M ctx, $3/$15 — NOT added: thinking is forced-on with
@@ -1 +1 @@
1
- {"version":3,"file":"models.d.ts","sourceRoot":"","sources":["../src/models.ts"],"names":[],"mappings":"AAAA;;;;;;;;GAQG;AAEH,OAAO,KAAK,EAAE,eAAe,EAAE,MAAM,YAAY,CAAA;AAQjD;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;GAwLG;AACH,eAAO,MAAM,MAAM,EAAE,SAAS,eAAe,EAupDnC,CAAA"}
1
+ {"version":3,"file":"models.d.ts","sourceRoot":"","sources":["../src/models.ts"],"names":[],"mappings":"AAAA;;;;;;;;GAQG;AAEH,OAAO,KAAK,EAAE,eAAe,EAAE,MAAM,YAAY,CAAA;AAQjD;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;GAoMG;AACH,eAAO,MAAM,MAAM,EAAE,SAAS,eAAe,EA6sDnC,CAAA"}
package/dist/models.js CHANGED
@@ -118,13 +118,22 @@ const WEEKDAYS_UTC = [1, 2, 3, 4, 5];
118
118
  * at the post-promo list rate ($1.50/$7.50, cache read $0.15) under the same
119
119
  * never-under-charge policy, each with a KNOWN_DIVERGENCES entry expiring
120
120
  * 2026-12-31. gemini-3.5-flash carries no promo: still $1.50/$9.00, $0.15)
121
+ * (re-verified 2026-09-02: gemini-3.8-flash "New Stable" — supersedes
122
+ * 3.7-flash, which the models page now calls "previous-generation". THIRD id
123
+ * on the same flash-tier promo row, verbatim "$0.75 through December 31,
124
+ * 2026. $1.50 starting January 1, 2027." / "$3.75 … $7.50" / "$0.075 …
125
+ * $0.15", so it too is cataloged at list with a KNOWN_DIVERGENCES entry.
126
+ * Specs from /docs/models/gemini-3.8-flash: 1M ctx / 65,536 out, thinking
127
+ * low|medium|high, vision, tools, caching, search grounding, code execution,
128
+ * url context — identical surface to 3.7-flash. Knowledge cutoff is published
129
+ * by neither Google nor models.dev for this id)
121
130
  * - xAI: https://docs.x.ai/developers/models + /developers/grok-4-5
122
131
  * (grok-4.5 flagship 2026-07-08: $2/$6, 500K ctx, ≥200K prompts bill 2× —
123
132
  * not modeled; reasoning_effort low|medium|high default high, image input;
124
133
  * grok-4.3 still served at $1.25/$2.50 with the bigger 1M window;
125
134
  * grok-code-fast-1 no longer listed — retires 2026-08-15)
126
135
  * - DeepSeek: https://api-docs.deepseek.com/quick_start/pricing (verified
127
- * 2026-08-31; legacy deepseek-chat/-reasoner ids fully retired 2026-07-24 —
136
+ * 2026-09-06; legacy deepseek-chat/-reasoner ids fully retired 2026-07-24 —
128
137
  * never in this catalog. The V4-Pro-GA price RISE effective 2026-08-16T16:00Z
129
138
  * has LANDED and is folded into the base fields, along with the peak-hour 2×
130
139
  * the same card introduced; every rate re-read on the card 2026-08-31 and
@@ -133,7 +142,10 @@ const WEEKDAYS_UTC = [1, 2, 3, 4, 5];
133
142
  * where on 2026-08-18 it carried no day qualifier at all, so the windows are
134
143
  * `daysOfWeekUtc`-restricted rather than daily. `deepseek-v4-flash-vision-exp`
135
144
  * also appears on the card at flash's rates: EXPERIMENTAL and vision-only-new,
136
- * deliberately not catalogued.)
145
+ * deliberately not catalogued. The 2026-09-06 re-read found every native rate
146
+ * unchanged; what moved was the US re-host — DeepInfra cut
147
+ * `DeepSeek-V4-Flash-0731` from $0.08 to $0.06 input, cache read $0.016 to
148
+ * $0.015, output held at $0.18. See that entry's `regionPricing`.)
137
149
  * - Moonshot: https://platform.kimi.ai/docs/models + DeepInfra's model API for
138
150
  * the US re-host (kimi-k3 flagship 2026-07-16
139
151
  * — 2.8T MoE, 1M ctx, $3/$15 — NOT added: thinking is forced-on with
@@ -776,11 +788,48 @@ export const MODELS = [
776
788
  // replace outright: the google bond has never been implemented/wired, so no
777
789
  // historical usage can reference the old ids.
778
790
  // ---------------------------------------------------------------------------
791
+ {
792
+ id: 'gemini-3.8-flash',
793
+ provider: 'google',
794
+ label: 'Gemini 3.8 Flash',
795
+ description: 'Google agentic flagship — long-horizon engineering & autonomous agents',
796
+ // Verified against /docs/models/gemini-3.8-flash (2026-09-02) — same limits
797
+ // and capability surface as 3.7-flash.
798
+ contextWindow: 1_048_576,
799
+ maxOutputTokens: 65_536,
800
+ supportsThinking: true,
801
+ thinkingBudgetTokens: 10_000,
802
+ thinkingConfigurable: true,
803
+ // thinking_level low|medium|high — minimal NOT supported on this model.
804
+ supportedEffortLevels: ['low', 'medium', 'high'],
805
+ defaultEffortLevel: 'medium',
806
+ supportsVision: true,
807
+ supportsPromptCaching: true,
808
+ supportsTools: true,
809
+ webSearchToolType: 'google_search',
810
+ codeExecutionToolType: 'code_execution',
811
+ webFetchToolType: 'url_context',
812
+ // "New Stable" 2026-09-02. LIST price $1.50/$7.50 — same card as 3.7- and
813
+ // 3.6-flash, and on the same flash-tier launch promo ($0.75/$3.75, cache
814
+ // read $0.075) through 2026-12-31; billed here at standard list so metering
815
+ // never under-charges (identical policy and expiry to the two entries
816
+ // below — see the matching KNOWN_DIVERGENCES entry in
817
+ // scripts/check-model-freshness.mjs).
818
+ inputPricePerMTok: 1.5,
819
+ outputPricePerMTok: 7.5,
820
+ // Gemini context cache: read $0.15/M (0.1× input), no write premium
821
+ // (storage billed separately per hour — not modeled).
822
+ cacheReadPricePerMTok: 0.15,
823
+ cacheWritePricePerMTok: 1.5,
824
+ // Published by neither Google's model page nor models.dev for this id —
825
+ // carried from 3.7-flash (same generation) as a best-effort estimate.
826
+ knowledgeCutoff: '2026-03-01',
827
+ },
779
828
  {
780
829
  id: 'gemini-3.7-flash',
781
830
  provider: 'google',
782
831
  label: 'Gemini 3.7 Flash',
783
- description: 'Google agentic flagship — complex coding & multi-step execution',
832
+ description: 'Previous Google agentic flagship — complex coding & multi-step execution',
784
833
  // Verified against /docs/models/gemini-3.7-flash (2026-08-13).
785
834
  contextWindow: 1_048_576,
786
835
  maxOutputTokens: 65_536,
@@ -810,6 +859,11 @@ export const MODELS = [
810
859
  cacheWritePricePerMTok: 1.5,
811
860
  // Not on Google's docs — models.dev reports 2026-03 (lead, not authority).
812
861
  knowledgeCutoff: '2026-03-01',
862
+ // Superseded by gemini-3.8-flash (2026-09-02) — same flash tier, same list
863
+ // price; Google's own models page now calls 3.7 "previous-generation".
864
+ // Still served upstream, so it stays priceable.
865
+ deprecatedAt: '2026-09-02',
866
+ supersededBy: 'gemini-3.8-flash',
813
867
  },
814
868
  {
815
869
  id: 'gemini-3.6-flash',
@@ -850,9 +904,11 @@ export const MODELS = [
850
904
  knowledgeCutoff: '2026-01-01',
851
905
  // Superseded by gemini-3.7-flash (2026-08-13) — same flash tier, same list
852
906
  // price; Google's own models page now calls 3.6 "previous-generation".
853
- // Still served upstream, so it stays priceable.
907
+ // Still served upstream, so it stays priceable. Re-pointed to
908
+ // gemini-3.8-flash on 2026-09-02 when 3.7 was itself deprecated —
909
+ // supersededBy must target a SELECTABLE model, never a chain.
854
910
  deprecatedAt: '2026-08-13',
855
- supersededBy: 'gemini-3.7-flash',
911
+ supersededBy: 'gemini-3.8-flash',
856
912
  },
857
913
  {
858
914
  id: 'gemini-3.5-flash',
@@ -883,11 +939,11 @@ export const MODELS = [
883
939
  cacheWritePricePerMTok: 1.5,
884
940
  knowledgeCutoff: '2025-01-01',
885
941
  // Superseded within the flash tier (first by 3.6-flash on 2026-07-21, now
886
- // pointed one hop to gemini-3.7-flash — supersededBy must target a
942
+ // pointed one hop to gemini-3.8-flash — supersededBy must target a
887
943
  // SELECTABLE model, never a chain). Still served upstream, so it stays
888
944
  // priceable.
889
945
  deprecatedAt: '2026-07-21',
890
- supersededBy: 'gemini-3.7-flash',
946
+ supersededBy: 'gemini-3.8-flash',
891
947
  },
892
948
  {
893
949
  id: 'gemini-3.1-pro-preview',
@@ -1221,23 +1277,29 @@ export const MODELS = [
1221
1277
  // US (DeepInfra) DEFAULT as of 2026-08-16 — flipped from CN when DeepSeek's
1222
1278
  // rise landed (owner decision 2026-08-14). CN was cheaper on real traffic
1223
1279
  // only because of its cache reads; the rise takes those from $0.0028 to
1224
- // $0.007 (peak $0.014) against DeepInfra's flat $0.016, which is no longer
1225
- // enough to carry the 1.6-3.1x it now loses on fresh input and output. On
1280
+ // $0.007 (peak $0.014) against DeepInfra's flat $0.015, which is no longer
1281
+ // enough to carry the 3.7x CN now loses on BOTH fresh input (0.22 vs 0.06)
1282
+ // and output (0.66 vs 0.18). On
1226
1283
  // the agentic mix this model actually serves (~94% cache hits) US is
1227
- // cheaper at EVERY hour: 0.114c/turn flat vs 0.152c off-peak and 0.303c at
1284
+ // cheaper at EVERY hour: 0.103c/turn flat vs 0.152c off-peak and 0.303c at
1228
1285
  // peak. It is also flat-rate, so free-tier cost stops varying by Beijing
1229
1286
  // business hours. Re-derive if the cache-hit ratio drops much below ~90%,
1230
1287
  // where CN's cheaper reads start winning again. This deliberately splits
1231
1288
  // the plan/execute pair across regions — Pro stays CN because its US
1232
1289
  // re-host is ~2.3x its own native rate even after the rise.
1233
1290
  regions: ['us', 'cn'],
1234
- // US = DeepInfra, verified 2026-08-13 against the id the bond actually
1291
+ // US = DeepInfra, verified 2026-09-06 against the id the bond actually
1235
1292
  // sends: `deepseek-ai/DeepSeek-V4-Flash-0731`, the official release that
1236
1293
  // supersedes the preview weights still served under the un-dated id
1237
- // (cents_per_input_token 0.000008, cents_per_output_token 0.000018,
1238
- // rate_per_input_token_cached 0.2 → cache read = 0.2 × input).
1294
+ // (cents_per_input_token 0.000006, cents_per_output_token 0.000018,
1295
+ // rate_per_input_token_cached 0.25 → cache read = 0.25 × input).
1296
+ // DeepInfra CUT the 0731 input rate 0.08 → 0.06 (cache read 0.016 →
1297
+ // 0.015); output held at 0.18. The un-dated `DeepSeek-V4-Flash` id moved
1298
+ // the other way (0.09 input, cache 0.2× = 0.018), so the two ids no longer
1299
+ // price alike — this entry tracks the dated one the modelMap sends, and the
1300
+ // gap widens the case for the US default this model already carries.
1239
1301
  regionPricing: {
1240
- us: { inputPricePerMTok: 0.08, outputPricePerMTok: 0.18, cacheReadPricePerMTok: 0.016 },
1302
+ us: { inputPricePerMTok: 0.06, outputPricePerMTok: 0.18, cacheReadPricePerMTok: 0.015 },
1241
1303
  },
1242
1304
  // Peak-hour surcharge live since 2026-08-16T16:00Z, Mon-Fri (see
1243
1305
  // deepseek-v4-pro). It applies to the NATIVE CN card only — this model
@@ -1743,9 +1805,13 @@ export const MODELS = [
1743
1805
  // GLM context cache: read ≈0.19× input, no write premium.
1744
1806
  cacheReadPricePerMTok: 0.26,
1745
1807
  cacheWritePricePerMTok: 1.4,
1746
- // No US re-host exists — DeepInfra serves GLM-5.2 and GLM-5.3-Flash but
1747
- // returns "model not found" for zai-org/GLM-5.3 (checked 2026-08-26), so
1748
- // this is pinned to the native host and bills the list card above.
1808
+ // Pinned to the native host, billing the list card above. DeepInfra has
1809
+ // since STARTED serving zai-org/GLM-5.3 (it 404'd when this was cataloged
1810
+ // on 2026-08-26; live 2026-09-02 at $1.20/$4.00, cache read 0.1× = $0.12,
1811
+ // i.e. below native). Adding 'us' here is not a catalog-only edit — it
1812
+ // needs a matching entry in molecule-dev's region-model-maps.ts and a
1813
+ // regionPricing block, moved together, or dispatch sends the canonical id
1814
+ // and 404s. Left as a human decision.
1749
1815
  regions: ['cn'],
1750
1816
  // Same base weights as glm-5.2, so the same best-effort estimate — Z.ai
1751
1817
  // publishes no cutoff.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@molecule/api-resource-ai-models",
3
- "version": "1.4.0",
3
+ "version": "1.5.1",
4
4
  "description": "AI model catalog — server-side source of truth plus an authentication-gated discovery endpoint",
5
5
  "type": "module",
6
6
  "main": "dist/index.js",