llm.rb 15.4.0 → 15.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: de1f4eeee4762548e75eac2b5d4112de49cf446944b30c545b76e62750aad6c3
4
- data.tar.gz: 4abc1d2d707717fde266b48bfa626f06d553514761653db8c3cd190cfcffb864
3
+ metadata.gz: df1a9b7fbf194f568b95f3f1b2d0314bb703027bc89e4432109d118f238ea2ba
4
+ data.tar.gz: 679938a2f6538f9c1999d31870462498393e1b3a7c7203630c5e8706c893e9ce
5
5
  SHA512:
6
- metadata.gz: c28cadd2a412e2623b472d3be7814aa51a462076782fc847bc73cdc57a601a630d909bb9b40c29d31dd7107b487bf169bdeb245e9803581df0d90d642694bc16
7
- data.tar.gz: 2fd3c1905d2db4212958e98e537db76d37f0f30f5cee4dda3cbf351c7f0fffe1ce6eaa5b38553e80b9441b029e2edf78c4e9ea75c200b731638dbbe49bd947df
6
+ metadata.gz: 49151bad60a22a06b119e08062e58c6d35727d59d10e07d9921166a23b5817083ea11fe942a090bc0e64b68a06340b9cf175a87ecfa81775dffcc7d88a8abb0c
7
+ data.tar.gz: 92ddc3cd73f7af70f25e735fd8ac63dcbca6ed48fd78ba3510be0171ea180d006f399ff60e215bdf1b9159e52b721916c59fcb7278c322fb0ffe789db5900602
data/CHANGELOG.md CHANGED
@@ -17,6 +17,86 @@
17
17
 
18
18
  *No unreleased changes yet. Check back after the next release.*
19
19
 
20
+ ## v15.5.0
21
+
22
+ Changes since `v15.4.1`.
23
+
24
+ This release makes the exec-backed tools report how long a command ran, calls a
25
+ tracer's `on_exit` hook once when the last scope ends, and resolves a `Proc`
26
+ parameter type when a schema is serialized. It also fixes `LLM::Schema.to_s`
27
+ for a schema with a deferred type.
28
+
29
+ ### Tools
30
+
31
+ * **tools: report how long a command ran** <br>
32
+ [`LLM::Tool::Exec#call`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool/Exec.html#call-instance_method)
33
+ now includes a `duration` field in its result, a string such as `"0.4 seconds"`
34
+ that reports how long the command took. The tools that route through `exec`
35
+ (`git`, `rg`, `mkdir`, `ruby`, and `bundle`) return it as well. A command that
36
+ was not found still returns the error hash, which has no `duration`.
37
+
38
+ ### Tracers
39
+
40
+ * **tracer: call `on_exit` once, when the last scope ends** <br>
41
+ [`LLM::Tracer#on_exit`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer.html#on_exit-instance_method)
42
+ is a new hook, called once when the last scope that is open for a tracer
43
+ ends. A tracer that holds a connection, a file, or a span releases it here.
44
+ The count of open scopes belongs to the tracer rather than to the thread that
45
+ opened one, because a tool runs on a thread of its own and scopes the turn's
46
+ tracer while it does. Before this, a tool's scope looked like the outermost
47
+ one: a tracer that released its resource in `on_exit` lost it in the middle
48
+ of the turn, and the rest of the trace was written without it.
49
+
50
+ ### Schema
51
+
52
+ * **schema: resolve a parameter type when the schema is serialized** <br>
53
+ [`LLM::Schema`](https://r.uby.dev/api-docs/llm.rb/LLM/Schema.html) and
54
+ [`LLM::Tool`](https://r.uby.dev/api-docs/llm.rb/LLM/Tool.html) accept a `Proc`
55
+ as the type of a parameter or a property, and the proc is called when the
56
+ schema is serialized rather than when the class that declares it is defined.
57
+ A type that is only known at runtime no longer has to be known at load time:
58
+ `parameter :kind, proc { Enum[*Article.kinds] }, "The article kind"` asks the
59
+ model for the values a store holds right now, and the proc runs again on the
60
+ next request, so it sees the values that are current then. Whatever the proc
61
+ returns is resolved in turn, so it may be a leaf, a class, or an array of
62
+ either, and the description, `required`, `default`, and `enum` given with the
63
+ parameter are applied to it.
64
+
65
+ * **schema: keep `to_s` working for a deferred type** <br>
66
+ Fix a bug where a schema with a proc-typed property raised from
67
+ [`LLM::Schema.to_s`](https://r.uby.dev/api-docs/llm.rb/LLM/Schema.html#to_s-class_method)
68
+ (and so from `inspect` and `p`) because the placeholder resolved its leaf
69
+ through a private method that shadowed the one `LLM::Schema::Leaf#default`
70
+ uses. Serialization, and therefore the request sent to a provider, was
71
+ unaffected.
72
+
73
+ ## v15.4.1
74
+
75
+ Changes since `v15.4.0`.
76
+
77
+ This release removes the gemspec's post-install message, so installing the gem
78
+ no longer prints the r.uby.dev notice. It also refreshes the model registry
79
+ with current model listings, limits, and pricing.
80
+
81
+ ### Core
82
+
83
+ * **remove the gemspec post install message** <br>
84
+ The gemspec no longer sets `post_install_message`, so installing the
85
+ gem no longer prints the r.uby.dev website notice.
86
+
87
+ ### Registry
88
+
89
+ * **refresh model metadata** <br>
90
+ Update `data/` with current model listings, limits, and pricing for the
91
+ Alibaba, Bedrock, DeepSeek, Google, Moonshot, and OpenRouter registries.
92
+ Bedrock adds the Kimi K3 and GPT-6 Sol and GPT-6 Luna families in both
93
+ the global and US regions and raises the context limit to 1M tokens for
94
+ two models, while Moonshot raises the Kimi K3 output limit to 1M tokens.
95
+ OpenRouter adds `qwen/qwen3.8-max-prime`, `z-ai/glm-5.3-prime`,
96
+ `upstage/solar-mini4`, the Aion 3.5 models, and `stealth/space-bunny-alpha`,
97
+ drops `mistralai/devstral-2512` and a free Ling 3.0 Flash VL entry, and
98
+ reprices several DeepSeek and Mistral models.
99
+
20
100
  ## v15.4.0
21
101
 
22
102
  Changes since `v15.3.0`.
data/README.md CHANGED
@@ -19,10 +19,13 @@ on CRuby. It has zero runtime dependencies by default, supports
19
19
  concurrent and parallel tool execution and has a single coherent API
20
20
  that spans 14+ providers.
21
21
 
22
- The most effective way to learn about llm.rb is to ask [the r.uby.dev chatbot](https://r.uby.dev)
23
- a question. It is connected to the llm.rb GitHub repository, backed by
24
- ActiveRecord and uses the builtin MCP feature to connect to GitHub. The chatbot
25
- is an llm.rb agent that is deployed with [roda-llm](https://github.com/r-uby-dev/roda-llm#readme).
22
+ It is possible to see llm.rb in action on the
23
+ [the r.uby.dev website](https://r.uby.dev) where
24
+ I am working on building an agentic platform that
25
+ users can use to manage multiple agents that are
26
+ specialized in different areas, and have access to
27
+ different services (eg GitHub, etc). Check it out if
28
+ curious. Still in early development.
26
29
 
27
30
  ## Install
28
31
 
@@ -362,9 +365,6 @@ for both Rack-based / Rails-based applications. On databases
362
365
  where it is supported, such as PostgreSQL, the column can be optimized by using
363
366
  the `jsonb` type.
364
367
 
365
- The following example is based on the agent used to power the
366
- [r.uby.dev chatbot](https://r.uby.dev).
367
-
368
368
  ```ruby
369
369
  require "active_record"
370
370
  require "llm"
@@ -1030,10 +1030,8 @@ IO.copy_stream res.images[0], "rocket-with-dog.svg"
1030
1030
  <summary>Where can I see llm.rb in action?</summary>
1031
1031
  <br>
1032
1032
 
1033
- The [r.uby.dev](https://r.uby.dev) website deploys
1034
- multiple llm.rb agents that guests can interact with
1035
- and there is even an agent that is connected to this
1036
- GitHub repository.
1033
+ The [r.uby.dev](https://r.uby.dev) website.
1034
+
1037
1035
  </details>
1038
1036
  <details>
1039
1037
  <summary>What about local LLM support?</summary>
@@ -1094,9 +1092,9 @@ web</a>.
1094
1092
 
1095
1093
  The llm.rb project was started more than three
1096
1094
  years ago by
1097
- [@0x1eef](https://github.com/0x1eef) and
1098
- [@antaz](https://github.com/0x1eef). The primary
1099
- maintainer is [@0x1eef](https://github.com/0x1eef).
1095
+ [@altruby](https://github.com/altruby) and
1096
+ [@antaz](https://github.com/altruby). The primary
1097
+ maintainer is [@altruby](https://github.com/altruby).
1100
1098
  Over those three years multiple other contributors have
1101
1099
  contributed to llm.rb as well, and new contributors are
1102
1100
  always welcome.
data/data/alibaba.json CHANGED
@@ -38,7 +38,7 @@
38
38
  "open_weights": false,
39
39
  "limit": {
40
40
  "context": 1000000,
41
- "output": 65536
41
+ "output": 131072
42
42
  },
43
43
  "cost": {
44
44
  "input": 2.5,
@@ -1848,7 +1848,7 @@
1848
1848
  "open_weights": false,
1849
1849
  "limit": {
1850
1850
  "context": 1000000,
1851
- "output": 65536
1851
+ "output": 131072
1852
1852
  },
1853
1853
  "cost": {
1854
1854
  "input": 0.5,
data/data/bedrock.json CHANGED
@@ -566,7 +566,7 @@
566
566
  },
567
567
  "open_weights": false,
568
568
  "limit": {
569
- "context": 272000,
569
+ "context": 1000000,
570
570
  "output": 128000
571
571
  },
572
572
  "provider": {
@@ -886,7 +886,7 @@
886
886
  },
887
887
  "open_weights": false,
888
888
  "limit": {
889
- "context": 272000,
889
+ "context": 1000000,
890
890
  "output": 128000
891
891
  },
892
892
  "provider": {
@@ -1587,6 +1587,41 @@
1587
1587
  "cache_write": 6.875
1588
1588
  }
1589
1589
  },
1590
+ "global.moonshotai.kimi-k3": {
1591
+ "id": "global.moonshotai.kimi-k3",
1592
+ "name": "Kimi K3 (Global)",
1593
+ "description": "Multimodal Kimi model with 1M context and toggleable max-effort thinking for long-horizon agent work",
1594
+ "family": "kimi-k3",
1595
+ "attachment": true,
1596
+ "reasoning": true,
1597
+ "reasoning_options": [],
1598
+ "tool_call": true,
1599
+ "interleaved": true,
1600
+ "structured_output": true,
1601
+ "temperature": false,
1602
+ "release_date": "2026-07-16",
1603
+ "last_updated": "2026-07-16",
1604
+ "modalities": {
1605
+ "input": [
1606
+ "text",
1607
+ "image"
1608
+ ],
1609
+ "output": [
1610
+ "text"
1611
+ ]
1612
+ },
1613
+ "open_weights": true,
1614
+ "limit": {
1615
+ "context": 1048576,
1616
+ "output": 131072
1617
+ },
1618
+ "cost": {
1619
+ "input": 3,
1620
+ "output": 15,
1621
+ "cache_read": 0.3,
1622
+ "cache_write": 3.75
1623
+ }
1624
+ },
1590
1625
  "apac.amazon.nova-pro-v1:0": {
1591
1626
  "id": "apac.amazon.nova-pro-v1:0",
1592
1627
  "name": "Nova Pro (APAC)",
@@ -1675,6 +1710,41 @@
1675
1710
  "cache_write": 4.125
1676
1711
  }
1677
1712
  },
1713
+ "us.moonshotai.kimi-k3": {
1714
+ "id": "us.moonshotai.kimi-k3",
1715
+ "name": "Kimi K3 (US)",
1716
+ "description": "Multimodal Kimi model with 1M context and toggleable max-effort thinking for long-horizon agent work",
1717
+ "family": "kimi-k3",
1718
+ "attachment": true,
1719
+ "reasoning": true,
1720
+ "reasoning_options": [],
1721
+ "tool_call": true,
1722
+ "interleaved": true,
1723
+ "structured_output": true,
1724
+ "temperature": false,
1725
+ "release_date": "2026-07-16",
1726
+ "last_updated": "2026-07-16",
1727
+ "modalities": {
1728
+ "input": [
1729
+ "text",
1730
+ "image"
1731
+ ],
1732
+ "output": [
1733
+ "text"
1734
+ ]
1735
+ },
1736
+ "open_weights": true,
1737
+ "limit": {
1738
+ "context": 1048576,
1739
+ "output": 131072
1740
+ },
1741
+ "cost": {
1742
+ "input": 3.3,
1743
+ "output": 16.5,
1744
+ "cache_read": 0.33,
1745
+ "cache_write": 4.125
1746
+ }
1747
+ },
1678
1748
  "apac.amazon.nova-lite-v1:0": {
1679
1749
  "id": "apac.amazon.nova-lite-v1:0",
1680
1750
  "name": "Nova Lite (APAC)",
@@ -1933,6 +2003,72 @@
1933
2003
  "output": 1.2
1934
2004
  }
1935
2005
  },
2006
+ "us.openai.gpt-6-luna": {
2007
+ "id": "us.openai.gpt-6-luna",
2008
+ "name": "GPT-6 Luna (US)",
2009
+ "description": "OpenAI's most efficient model for focused, high-volume tasks",
2010
+ "family": "gpt-luna",
2011
+ "attachment": true,
2012
+ "reasoning": true,
2013
+ "reasoning_options": [
2014
+ {
2015
+ "type": "effort",
2016
+ "values": [
2017
+ "none",
2018
+ "low",
2019
+ "medium",
2020
+ "high",
2021
+ "xhigh",
2022
+ "max"
2023
+ ]
2024
+ }
2025
+ ],
2026
+ "tool_call": true,
2027
+ "structured_output": true,
2028
+ "temperature": false,
2029
+ "knowledge": "2026-05-18",
2030
+ "release_date": "2026-09-22",
2031
+ "last_updated": "2026-09-22",
2032
+ "modalities": {
2033
+ "input": [
2034
+ "text",
2035
+ "image"
2036
+ ],
2037
+ "output": [
2038
+ "text"
2039
+ ]
2040
+ },
2041
+ "open_weights": false,
2042
+ "limit": {
2043
+ "context": 1050000,
2044
+ "input": 922000,
2045
+ "output": 128000
2046
+ },
2047
+ "cost": {
2048
+ "input": 0.11,
2049
+ "output": 0.55,
2050
+ "cache_read": 0.011,
2051
+ "cache_write": 0.1375,
2052
+ "tiers": [
2053
+ {
2054
+ "input": 0.22,
2055
+ "output": 0.825,
2056
+ "cache_read": 0.022,
2057
+ "cache_write": 0.275,
2058
+ "tier": {
2059
+ "type": "context",
2060
+ "size": 272000
2061
+ }
2062
+ }
2063
+ ],
2064
+ "context_over_200k": {
2065
+ "input": 0.22,
2066
+ "output": 0.825,
2067
+ "cache_read": 0.022,
2068
+ "cache_write": 0.275
2069
+ }
2070
+ }
2071
+ },
1936
2072
  "meta.llama3-3-70b-instruct-v1:0": {
1937
2073
  "id": "meta.llama3-3-70b-instruct-v1:0",
1938
2074
  "name": "Llama 3.3 70B Instruct",
@@ -3258,6 +3394,72 @@
3258
3394
  "cache_write": 0.33
3259
3395
  }
3260
3396
  },
3397
+ "global.openai.gpt-6-sol": {
3398
+ "id": "global.openai.gpt-6-sol",
3399
+ "name": "GPT-6 Sol (Global)",
3400
+ "description": "OpenAI model for complex coding and agentic workflows",
3401
+ "family": "gpt-sol",
3402
+ "attachment": true,
3403
+ "reasoning": true,
3404
+ "reasoning_options": [
3405
+ {
3406
+ "type": "effort",
3407
+ "values": [
3408
+ "none",
3409
+ "low",
3410
+ "medium",
3411
+ "high",
3412
+ "xhigh",
3413
+ "max"
3414
+ ]
3415
+ }
3416
+ ],
3417
+ "tool_call": true,
3418
+ "structured_output": true,
3419
+ "temperature": false,
3420
+ "knowledge": "2026-04-20",
3421
+ "release_date": "2026-09-22",
3422
+ "last_updated": "2026-09-22",
3423
+ "modalities": {
3424
+ "input": [
3425
+ "text",
3426
+ "image"
3427
+ ],
3428
+ "output": [
3429
+ "text"
3430
+ ]
3431
+ },
3432
+ "open_weights": false,
3433
+ "limit": {
3434
+ "context": 1050000,
3435
+ "input": 922000,
3436
+ "output": 128000
3437
+ },
3438
+ "cost": {
3439
+ "input": 2,
3440
+ "output": 10,
3441
+ "cache_read": 0.2,
3442
+ "cache_write": 2.5,
3443
+ "tiers": [
3444
+ {
3445
+ "input": 4,
3446
+ "output": 15,
3447
+ "cache_read": 0.4,
3448
+ "cache_write": 5,
3449
+ "tier": {
3450
+ "type": "context",
3451
+ "size": 272000
3452
+ }
3453
+ }
3454
+ ],
3455
+ "context_over_200k": {
3456
+ "input": 4,
3457
+ "output": 15,
3458
+ "cache_read": 0.4,
3459
+ "cache_write": 5
3460
+ }
3461
+ }
3462
+ },
3261
3463
  "global.xai.grok-4.6": {
3262
3464
  "id": "global.xai.grok-4.6",
3263
3465
  "name": "Grok 4.6 (Global)",
@@ -4572,6 +4774,72 @@
4572
4774
  }
4573
4775
  }
4574
4776
  },
4777
+ "global.openai.gpt-6-luna": {
4778
+ "id": "global.openai.gpt-6-luna",
4779
+ "name": "GPT-6 Luna (Global)",
4780
+ "description": "OpenAI's most efficient model for focused, high-volume tasks",
4781
+ "family": "gpt-luna",
4782
+ "attachment": true,
4783
+ "reasoning": true,
4784
+ "reasoning_options": [
4785
+ {
4786
+ "type": "effort",
4787
+ "values": [
4788
+ "none",
4789
+ "low",
4790
+ "medium",
4791
+ "high",
4792
+ "xhigh",
4793
+ "max"
4794
+ ]
4795
+ }
4796
+ ],
4797
+ "tool_call": true,
4798
+ "structured_output": true,
4799
+ "temperature": false,
4800
+ "knowledge": "2026-05-18",
4801
+ "release_date": "2026-09-22",
4802
+ "last_updated": "2026-09-22",
4803
+ "modalities": {
4804
+ "input": [
4805
+ "text",
4806
+ "image"
4807
+ ],
4808
+ "output": [
4809
+ "text"
4810
+ ]
4811
+ },
4812
+ "open_weights": false,
4813
+ "limit": {
4814
+ "context": 1050000,
4815
+ "input": 922000,
4816
+ "output": 128000
4817
+ },
4818
+ "cost": {
4819
+ "input": 0.1,
4820
+ "output": 0.5,
4821
+ "cache_read": 0.01,
4822
+ "cache_write": 0.125,
4823
+ "tiers": [
4824
+ {
4825
+ "input": 0.2,
4826
+ "output": 0.75,
4827
+ "cache_read": 0.02,
4828
+ "cache_write": 0.25,
4829
+ "tier": {
4830
+ "type": "context",
4831
+ "size": 272000
4832
+ }
4833
+ }
4834
+ ],
4835
+ "context_over_200k": {
4836
+ "input": 0.2,
4837
+ "output": 0.75,
4838
+ "cache_read": 0.02,
4839
+ "cache_write": 0.25
4840
+ }
4841
+ }
4842
+ },
4575
4843
  "us.anthropic.claude-fable-5-1": {
4576
4844
  "id": "us.anthropic.claude-fable-5-1",
4577
4845
  "name": "Claude Fable 5.1 (US)",
@@ -5349,6 +5617,72 @@
5349
5617
  "cache_write": 0.06
5350
5618
  }
5351
5619
  },
5620
+ "us.openai.gpt-6-sol": {
5621
+ "id": "us.openai.gpt-6-sol",
5622
+ "name": "GPT-6 Sol (US)",
5623
+ "description": "OpenAI model for complex coding and agentic workflows",
5624
+ "family": "gpt-sol",
5625
+ "attachment": true,
5626
+ "reasoning": true,
5627
+ "reasoning_options": [
5628
+ {
5629
+ "type": "effort",
5630
+ "values": [
5631
+ "none",
5632
+ "low",
5633
+ "medium",
5634
+ "high",
5635
+ "xhigh",
5636
+ "max"
5637
+ ]
5638
+ }
5639
+ ],
5640
+ "tool_call": true,
5641
+ "structured_output": true,
5642
+ "temperature": false,
5643
+ "knowledge": "2026-04-20",
5644
+ "release_date": "2026-09-22",
5645
+ "last_updated": "2026-09-22",
5646
+ "modalities": {
5647
+ "input": [
5648
+ "text",
5649
+ "image"
5650
+ ],
5651
+ "output": [
5652
+ "text"
5653
+ ]
5654
+ },
5655
+ "open_weights": false,
5656
+ "limit": {
5657
+ "context": 1050000,
5658
+ "input": 922000,
5659
+ "output": 128000
5660
+ },
5661
+ "cost": {
5662
+ "input": 2.2,
5663
+ "output": 11,
5664
+ "cache_read": 0.22,
5665
+ "cache_write": 2.75,
5666
+ "tiers": [
5667
+ {
5668
+ "input": 4.4,
5669
+ "output": 16.5,
5670
+ "cache_read": 0.44,
5671
+ "cache_write": 5.5,
5672
+ "tier": {
5673
+ "type": "context",
5674
+ "size": 272000
5675
+ }
5676
+ }
5677
+ ],
5678
+ "context_over_200k": {
5679
+ "input": 4.4,
5680
+ "output": 16.5,
5681
+ "cache_read": 0.44,
5682
+ "cache_write": 5.5
5683
+ }
5684
+ }
5685
+ },
5352
5686
  "global.anthropic.claude-opus-4-7": {
5353
5687
  "id": "global.anthropic.claude-opus-4-7",
5354
5688
  "name": "Claude Opus 4.7 (Global)",
data/data/deepseek.json CHANGED
@@ -49,7 +49,7 @@
49
49
  "open_weights": true,
50
50
  "limit": {
51
51
  "context": 1000000,
52
- "output": 384000
52
+ "output": 393216
53
53
  },
54
54
  "status": "deprecated",
55
55
  "cost": {
@@ -100,7 +100,7 @@
100
100
  "open_weights": true,
101
101
  "limit": {
102
102
  "context": 1000000,
103
- "output": 384000
103
+ "output": 393216
104
104
  },
105
105
  "status": "deprecated",
106
106
  "cost": {
@@ -149,7 +149,7 @@
149
149
  "open_weights": true,
150
150
  "limit": {
151
151
  "context": 1000000,
152
- "output": 384000
152
+ "output": 393216
153
153
  },
154
154
  "cost": {
155
155
  "input": 0.435,
@@ -199,7 +199,7 @@
199
199
  "open_weights": true,
200
200
  "limit": {
201
201
  "context": 1000000,
202
- "output": 384000
202
+ "output": 393216
203
203
  },
204
204
  "cost": {
205
205
  "input": 0.15,
data/data/google.json CHANGED
@@ -138,7 +138,7 @@
138
138
  "open_weights": false,
139
139
  "limit": {
140
140
  "context": 65536,
141
- "output": 65536
141
+ "output": 4096
142
142
  },
143
143
  "cost": {
144
144
  "input": 0.25,
@@ -933,8 +933,8 @@
933
933
  },
934
934
  "open_weights": false,
935
935
  "limit": {
936
- "context": 65536,
937
- "output": 65536
936
+ "context": 131072,
937
+ "output": 32768
938
938
  },
939
939
  "cost": {
940
940
  "input": 0.5,
data/data/moonshot.json CHANGED
@@ -167,7 +167,7 @@
167
167
  "open_weights": true,
168
168
  "limit": {
169
169
  "context": 1048576,
170
- "output": 131072
170
+ "output": 1048576
171
171
  },
172
172
  "cost": {
173
173
  "input": 3,