llm.rb 13.0.0 → 14.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +505 -14
- data/README.md +484 -50
- data/bin/llm.rb +148 -0
- data/data/anthropic.json +206 -263
- data/data/bedrock.json +2138 -1860
- data/data/deepinfra.json +1003 -624
- data/data/deepseek.json +38 -34
- data/data/google.json +1079 -371
- data/data/mistral.json +448 -368
- data/data/moonshot.json +384 -0
- data/data/openai.json +974 -1343
- data/data/xai.json +154 -126
- data/data/zai.json +191 -191
- data/lib/llm/agent.rb +123 -20
- data/lib/llm/context.rb +71 -88
- data/lib/llm/cost.rb +23 -17
- data/lib/llm/error.rb +0 -8
- data/lib/llm/function/array.rb +3 -3
- data/lib/llm/function/async/task.rb +2 -0
- data/lib/llm/function/fiber/task.rb +2 -0
- data/lib/llm/function/fork/task.rb +2 -0
- data/lib/llm/function/ractor/task.rb +2 -0
- data/lib/llm/function/sequential/group.rb +4 -1
- data/lib/llm/function/sequential/task.rb +1 -1
- data/lib/llm/function/task.rb +4 -0
- data/lib/llm/function/thread/task.rb +2 -0
- data/lib/llm/function.rb +33 -6
- data/lib/llm/guard/loop.rb +89 -0
- data/lib/llm/guard/null.rb +19 -0
- data/lib/llm/guard.rb +61 -0
- data/lib/llm/provider.rb +36 -0
- data/lib/llm/providers/anthropic/stream_parser.rb +1 -0
- data/lib/llm/providers/anthropic.rb +2 -9
- data/lib/llm/providers/bedrock/request_adapter.rb +1 -1
- data/lib/llm/providers/bedrock/stream_parser.rb +1 -0
- data/lib/llm/providers/bedrock.rb +1 -8
- data/lib/llm/providers/google/stream_parser.rb +1 -0
- data/lib/llm/providers/google.rb +1 -8
- data/lib/llm/providers/mistral.rb +1 -1
- data/lib/llm/providers/moonshot.rb +76 -0
- data/lib/llm/providers/ollama.rb +2 -9
- data/lib/llm/providers/openai/responses/stream_parser.rb +1 -0
- data/lib/llm/providers/openai/responses.rb +7 -9
- data/lib/llm/providers/openai/stream_parser.rb +1 -0
- data/lib/llm/providers/openai.rb +4 -11
- data/lib/llm/repl/bar.rb +4 -3
- data/lib/llm/repl/{transcript.rb → buffer.rb} +69 -29
- data/lib/llm/repl/color.rb +78 -0
- data/lib/llm/repl/command.rb +12 -5
- data/lib/llm/repl/commands/compact.rb +2 -2
- data/lib/llm/repl/commands/help.rb +3 -5
- data/lib/llm/repl/input/char.rb +46 -0
- data/lib/llm/repl/input/row.rb +39 -0
- data/lib/llm/repl/input.rb +251 -66
- data/lib/llm/repl/markdown/table.rb +11 -3
- data/lib/llm/repl/markdown.rb +34 -8
- data/lib/llm/repl/node.rb +37 -0
- data/lib/llm/repl/status.rb +42 -7
- data/lib/llm/repl/stream.rb +18 -6
- data/lib/llm/repl/walker.rb +3 -2
- data/lib/llm/repl/window.rb +54 -35
- data/lib/llm/repl.rb +74 -32
- data/lib/llm/skill.rb +20 -4
- data/lib/llm/stream.rb +8 -7
- data/lib/llm/tool.rb +29 -0
- data/lib/llm/tools/{swap_text.rb → edit-file.rb} +3 -3
- data/lib/llm/tools/git.rb +3 -0
- data/lib/llm/tools/mkdir.rb +3 -0
- data/lib/llm/tools/rg.rb +3 -0
- data/lib/llm/tools/ruby.rb +46 -0
- data/lib/llm/tools/shell.rb +3 -0
- data/lib/llm/tracer/pretty_logger.rb +127 -0
- data/lib/llm/tracer.rb +1 -0
- data/lib/llm/transformer/null.rb +21 -0
- data/lib/llm/transformer.rb +55 -0
- data/lib/llm/version.rb +1 -1
- data/lib/llm.rb +12 -2
- data/llm.gemspec +9 -2
- data/resources/deepdive/advanced/cancellation.md +74 -0
- data/resources/deepdive/advanced/compaction.md +83 -0
- data/resources/deepdive/advanced/context.md +267 -0
- data/resources/deepdive/advanced/guard.md +371 -0
- data/resources/deepdive/advanced/tracer.md +180 -0
- data/resources/deepdive/advanced/transformer.md +67 -0
- data/resources/deepdive/advanced/transports.md +45 -0
- data/resources/deepdive/everything_else/audio.md +122 -0
- data/resources/deepdive/everything_else/cost.md +99 -0
- data/resources/deepdive/everything_else/images.md +89 -0
- data/resources/deepdive/everything_else/object.md +108 -0
- data/resources/deepdive/everything_else/ocr.md +48 -0
- data/resources/deepdive/fundamentals/agents.md +202 -0
- data/resources/deepdive/fundamentals/builtin_tools.md +191 -0
- data/resources/deepdive/fundamentals/concurrency.md +104 -0
- data/resources/deepdive/fundamentals/database.md +449 -0
- data/resources/deepdive/fundamentals/embeddings.md +157 -0
- data/resources/deepdive/fundamentals/repl.md +87 -0
- data/resources/deepdive/fundamentals/schema.md +61 -0
- data/resources/deepdive/fundamentals/skills.md +106 -0
- data/resources/deepdive/fundamentals/stream.md +110 -0
- data/resources/deepdive/fundamentals/tools.md +265 -0
- data/resources/deepdive/protocols/a2a.md +106 -0
- data/resources/deepdive/protocols/mcp.md +111 -0
- data/resources/deepdive.md +58 -1792
- metadata +51 -7
- data/lib/llm/loop_guard.rb +0 -107
|
@@ -0,0 +1,180 @@
|
|
|
1
|
+
|
|
2
|
+
## Tracer
|
|
3
|
+
|
|
4
|
+
### Introduction
|
|
5
|
+
|
|
6
|
+
#### Overview
|
|
7
|
+
|
|
8
|
+
Tracers let you observe what the runtime is doing. They hook into
|
|
9
|
+
requests, tool calls, compactions, and other events. Debug a
|
|
10
|
+
misbehaving agent, monitor request latency, or export spans to
|
|
11
|
+
an observability backend. A provider-wide tracer intercepts every
|
|
12
|
+
request through that provider. An agent-local tracer only covers
|
|
13
|
+
requests made by that agent.
|
|
14
|
+
|
|
15
|
+
#### How it works
|
|
16
|
+
|
|
17
|
+
When you want to observe runtime events, subclass
|
|
18
|
+
[`LLM::Tracer`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer.html)
|
|
19
|
+
and implement the hooks you need, or use one of the built-in
|
|
20
|
+
tracers. The built-in tracers include
|
|
21
|
+
[`LLM::Tracer::Logger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html)
|
|
22
|
+
(writes structured JSON to stdout or a file),
|
|
23
|
+
[`LLM::Tracer::PrettyLogger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/PrettyLogger.html)
|
|
24
|
+
(writes human-readable single-line logs to stderr), and
|
|
25
|
+
[`LLM::Tracer::Telemetry`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Telemetry.html)
|
|
26
|
+
(exports spans via OTLP for OpenTelemetry). Attach a tracer to a
|
|
27
|
+
provider or an agent:
|
|
28
|
+
|
|
29
|
+
```ruby
|
|
30
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
31
|
+
llm.tracer = LLM::Tracer::PrettyLogger.new(llm)
|
|
32
|
+
agent = LLM::Agent.new(llm)
|
|
33
|
+
agent.talk "Hello"
|
|
34
|
+
```
|
|
35
|
+
|
|
36
|
+
#### Why would I use it?
|
|
37
|
+
|
|
38
|
+
Tracers give you visibility into what the runtime is doing. Debug
|
|
39
|
+
a misbehaving agent by tracing every request it makes. Monitor
|
|
40
|
+
request latency and token usage across providers. Export spans to
|
|
41
|
+
OpenTelemetry for integration with existing observability pipelines.
|
|
42
|
+
[`LLM::Tracer::PrettyLogger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/PrettyLogger.html)
|
|
43
|
+
is the best choice during development for compact, human-readable
|
|
44
|
+
output.
|
|
45
|
+
[`LLM::Tracer::Logger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html)
|
|
46
|
+
provides structured JSON, and
|
|
47
|
+
[`LLM::Tracer::Telemetry`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Telemetry.html)
|
|
48
|
+
exports spans to OpenTelemetry for production observability.
|
|
49
|
+
|
|
50
|
+
#### Notes
|
|
51
|
+
|
|
52
|
+
The tracer is extensible. You can implement custom hooks for any
|
|
53
|
+
runtime event. The scope can be an individual agent or every
|
|
54
|
+
request a provider makes. Three built-in tracers are available:
|
|
55
|
+
[`LLM::Tracer::PrettyLogger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/PrettyLogger.html)
|
|
56
|
+
(human-readable),
|
|
57
|
+
[`LLM::Tracer::Logger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html)
|
|
58
|
+
(structured JSON), and
|
|
59
|
+
[`LLM::Tracer::Telemetry`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Telemetry.html)
|
|
60
|
+
(OpenTelemetry).
|
|
61
|
+
|
|
62
|
+
### Provider
|
|
63
|
+
|
|
64
|
+
#### Overview
|
|
65
|
+
|
|
66
|
+
A provider-wide tracer intercepts every request made through that
|
|
67
|
+
provider. All agents sharing the same provider share the same
|
|
68
|
+
tracer. Use this to trace at the infrastructure level without
|
|
69
|
+
configuring each agent individually.
|
|
70
|
+
|
|
71
|
+
#### How it works
|
|
72
|
+
|
|
73
|
+
When you want every request through a provider to be traced, set
|
|
74
|
+
the tracer on the provider directly. Every request made through
|
|
75
|
+
that provider, regardless of which agent initiates it, flows
|
|
76
|
+
through the same tracer hooks. The provider holds a reference to the
|
|
77
|
+
tracer and passes it to every new context it creates. This ensures
|
|
78
|
+
consistent observability without configuring each agent individually.
|
|
79
|
+
|
|
80
|
+
```ruby
|
|
81
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
82
|
+
llm.tracer = LLM::Tracer::Logger.new(llm, io: $stdout)
|
|
83
|
+
```
|
|
84
|
+
|
|
85
|
+
#### Why would I use it?
|
|
86
|
+
|
|
87
|
+
A provider-wide tracer captures every request at the infrastructure level.
|
|
88
|
+
All agents sharing the same provider share the same tracer.
|
|
89
|
+
|
|
90
|
+
#### Notes
|
|
91
|
+
|
|
92
|
+
The tracer can also write to a file with the `path:` option to
|
|
93
|
+
[`LLM::Tracer::Logger.new`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html#initialize-instance_method).
|
|
94
|
+
|
|
95
|
+
### Agent
|
|
96
|
+
|
|
97
|
+
#### Overview
|
|
98
|
+
|
|
99
|
+
An agent-local tracer only covers requests made by that agent.
|
|
100
|
+
Attach it via the `tracer:` keyword argument to
|
|
101
|
+
[`LLM::Agent.new`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#initialize-instance_method)
|
|
102
|
+
and it follows that agent wherever it goes. Different agents can
|
|
103
|
+
have different tracers.
|
|
104
|
+
|
|
105
|
+
#### How it works
|
|
106
|
+
|
|
107
|
+
When you want a tracer for a specific agent, pass it to the agent
|
|
108
|
+
on creation. Only requests made by
|
|
109
|
+
that agent flow through the tracer, leaving other agents on the
|
|
110
|
+
same provider unaffected.
|
|
111
|
+
|
|
112
|
+
```ruby
|
|
113
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
114
|
+
agent = LLM::Agent.new(llm, tracer: LLM::Tracer::Logger.new(llm, io: $stdout))
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
#### Why would I use it?
|
|
118
|
+
|
|
119
|
+
Agent-local tracers let each agent log differently.
|
|
120
|
+
One agent might log to stdout, another to a file, a third to
|
|
121
|
+
OpenTelemetry.
|
|
122
|
+
|
|
123
|
+
#### Notes
|
|
124
|
+
|
|
125
|
+
The tracer can also write to a file with the `path:` option to
|
|
126
|
+
[`LLM::Tracer::Logger.new`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html#initialize-instance_method).
|
|
127
|
+
|
|
128
|
+
### PrettyLogger
|
|
129
|
+
|
|
130
|
+
#### Overview
|
|
131
|
+
|
|
132
|
+
[`LLM::Tracer::PrettyLogger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/PrettyLogger.html)
|
|
133
|
+
writes human-readable single-line logs to stderr. Unlike
|
|
134
|
+
[`LLM::Tracer::Logger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html)
|
|
135
|
+
(which emits structured JSON), the pretty logger is designed for
|
|
136
|
+
interactive development sessions where you want to see request and
|
|
137
|
+
tool-call activity at a glance.
|
|
138
|
+
|
|
139
|
+
#### How it works
|
|
140
|
+
|
|
141
|
+
Each request and tool call produces a single line on stderr with
|
|
142
|
+
the model, duration, and a summary of the activity. The logger
|
|
143
|
+
accepts an `io:` option to redirect output.
|
|
144
|
+
|
|
145
|
+
##### Provider-wide
|
|
146
|
+
|
|
147
|
+
```ruby
|
|
148
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
149
|
+
llm.tracer = LLM::Tracer::PrettyLogger.new(llm)
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
##### Agent-local
|
|
153
|
+
|
|
154
|
+
```ruby
|
|
155
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
156
|
+
agent = LLM::Agent.new(llm, tracer: LLM::Tracer::PrettyLogger.new(llm))
|
|
157
|
+
```
|
|
158
|
+
|
|
159
|
+
##### Custom output
|
|
160
|
+
|
|
161
|
+
```ruby
|
|
162
|
+
tracer = LLM::Tracer::PrettyLogger.new(llm, io: $stdout)
|
|
163
|
+
tracer = LLM::Tracer::PrettyLogger.new(llm, io: File.open("trace.log", "a"))
|
|
164
|
+
```
|
|
165
|
+
|
|
166
|
+
#### Why would I use it?
|
|
167
|
+
|
|
168
|
+
The pretty logger is the best choice for development. The output is
|
|
169
|
+
compact enough to follow in real time while still showing the model
|
|
170
|
+
name, duration, and tool calls. Switch to
|
|
171
|
+
[`LLM::Tracer::Logger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html)
|
|
172
|
+
when you need structured JSON for programmatic analysis, or to
|
|
173
|
+
[`LLM::Tracer::Telemetry`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Telemetry.html)
|
|
174
|
+
when you need OpenTelemetry exports.
|
|
175
|
+
|
|
176
|
+
#### Notes
|
|
177
|
+
|
|
178
|
+
The pretty logger writes to `$stderr` by default. Set `io:` to
|
|
179
|
+
redirect output. All three built-in tracers share the same interface,
|
|
180
|
+
so switching between them requires changing only the class name.
|
|
@@ -0,0 +1,67 @@
|
|
|
1
|
+
|
|
2
|
+
## Transformer
|
|
3
|
+
|
|
4
|
+
### Introduction
|
|
5
|
+
|
|
6
|
+
#### Overview
|
|
7
|
+
|
|
8
|
+
[`LLM::Transformer`](https://r.uby.dev/api-docs/llm.rb/LLM/Transformer.html)
|
|
9
|
+
is the superclass for message transformers. A transformer is bound
|
|
10
|
+
to a context and rewrites a single message before it is sent to the
|
|
11
|
+
provider. This lets you scrub sensitive data, inject context, or
|
|
12
|
+
otherwise modify outgoing messages without changing your prompt
|
|
13
|
+
code.
|
|
14
|
+
|
|
15
|
+
#### How it works
|
|
16
|
+
|
|
17
|
+
A transformer is a subclass of
|
|
18
|
+
[`LLM::Transformer`](https://r.uby.dev/api-docs/llm.rb/LLM/Transformer.html)
|
|
19
|
+
that implements
|
|
20
|
+
[`LLM::Transformer#call`](https://r.uby.dev/api-docs/llm.rb/LLM/Transformer.html#call-instance_method).
|
|
21
|
+
The method receives the message to transform and returns a
|
|
22
|
+
message.
|
|
23
|
+
You can mutate the message in place or return a new one; either way,
|
|
24
|
+
the returned message is what gets sent.
|
|
25
|
+
|
|
26
|
+
Configure the transformer on a context with the `transformer:` option,
|
|
27
|
+
passing a class rather than an instance. The runtime instantiates it
|
|
28
|
+
once per turn. Options passed through `transformer_options:` are
|
|
29
|
+
forwarded to `call` as keyword arguments. The transformer runs on
|
|
30
|
+
the most recent message in both chat completions and Responses API
|
|
31
|
+
turns, before the request reaches the provider:
|
|
32
|
+
|
|
33
|
+
```ruby
|
|
34
|
+
class RedactEmails < LLM::Transformer
|
|
35
|
+
def call(message:)
|
|
36
|
+
content = message.content.to_s.gsub(/[\w.+-]+@[\w-]+\.[\w.]+/, "[EMAIL]")
|
|
37
|
+
LLM::Message.new(message.role, content, message.extra)
|
|
38
|
+
end
|
|
39
|
+
end
|
|
40
|
+
|
|
41
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
42
|
+
ctx = LLM::Context.new(
|
|
43
|
+
llm,
|
|
44
|
+
transformer: RedactEmails
|
|
45
|
+
)
|
|
46
|
+
ctx.talk "Contact support@example.com for help"
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
#### Why would I use it?
|
|
50
|
+
|
|
51
|
+
Transformers give you a single hook point for all outgoing messages.
|
|
52
|
+
Common uses include redacting PII before it leaves your process,
|
|
53
|
+
injecting a timestamp or request ID, or normalizing content for a
|
|
54
|
+
particular provider. Because the transformer runs automatically on
|
|
55
|
+
every turn, you never need to remember to apply the transform in
|
|
56
|
+
your prompt code.
|
|
57
|
+
|
|
58
|
+
#### Notes
|
|
59
|
+
|
|
60
|
+
[`LLM::Transformer::Null`](https://r.uby.dev/api-docs/llm.rb/LLM/Transformer/Null.html)
|
|
61
|
+
is the default transformer; it returns the message unchanged. The
|
|
62
|
+
`transformer_options:` hash is passed to `call` on every turn.
|
|
63
|
+
Streams can observe transformation through the
|
|
64
|
+
[`LLM::Stream#on_transform`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_transform-instance_method)
|
|
65
|
+
and
|
|
66
|
+
[`LLM::Stream#on_transform_finish`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_transform_finish-instance_method)
|
|
67
|
+
callbacks, which receive the transformer instance.
|
|
@@ -0,0 +1,45 @@
|
|
|
1
|
+
|
|
2
|
+
## Transports
|
|
3
|
+
|
|
4
|
+
### Introduction
|
|
5
|
+
|
|
6
|
+
#### Overview
|
|
7
|
+
|
|
8
|
+
The transport option selects which HTTP library is used for network
|
|
9
|
+
communication.
|
|
10
|
+
[`LLM::Provider`](https://r.uby.dev/api-docs/llm.rb/LLM/Provider.html),
|
|
11
|
+
[`LLM::MCP`](https://r.uby.dev/api-docs/llm.rb/LLM/MCP.html), and
|
|
12
|
+
[`LLM::A2A`](https://r.uby.dev/api-docs/llm.rb/LLM/A2A.html)
|
|
13
|
+
all accept this option. Three backends are available out of the box:
|
|
14
|
+
`net/http` is always there, `net/http/persistent` pools connections,
|
|
15
|
+
and `curb` wraps libcurl. Each implements the same internal interface
|
|
16
|
+
so switching between them is a one-word change.
|
|
17
|
+
|
|
18
|
+
#### How it works
|
|
19
|
+
|
|
20
|
+
**Net/HTTP** (`:net_http`) is the
|
|
21
|
+
default and is always available. **Net/HTTP/Persistent**
|
|
22
|
+
(`:net_http_persistent`) maintains a connection pool so the cost
|
|
23
|
+
of tearing down and setting up connections is kept low. **Curb**
|
|
24
|
+
(`:curb`) provides bindings for libcurl.
|
|
25
|
+
|
|
26
|
+
```ruby
|
|
27
|
+
llm = LLM.deepseek(key: "...", transport: :net_http)
|
|
28
|
+
mcp = LLM::MCP.http(url: "...", transport: :net_http_persistent)
|
|
29
|
+
a2a = LLM::A2A.rest(url: "...", transport: :curb)
|
|
30
|
+
```
|
|
31
|
+
|
|
32
|
+
#### Why would I use it?
|
|
33
|
+
|
|
34
|
+
Different environments have different HTTP needs.
|
|
35
|
+
|
|
36
|
+
The default `net/http` transport is always available and works
|
|
37
|
+
everywhere. The persistent transport reduces connection overhead
|
|
38
|
+
when your agent makes many requests to the same provider in quick
|
|
39
|
+
succession. Curb gives you libcurl bindings for environments where
|
|
40
|
+
that is already configured or preferred.
|
|
41
|
+
|
|
42
|
+
#### Notes
|
|
43
|
+
|
|
44
|
+
The persistent transport is built on top of `net/http`. Curb
|
|
45
|
+
requires the `curb` gem.
|
|
@@ -0,0 +1,122 @@
|
|
|
1
|
+
|
|
2
|
+
## Audio
|
|
3
|
+
|
|
4
|
+
### Introduction
|
|
5
|
+
|
|
6
|
+
#### Overview
|
|
7
|
+
|
|
8
|
+
The audio interface covers three things: turning text into speech,
|
|
9
|
+
transcribing audio into text, and translating spoken language.
|
|
10
|
+
OpenAI supports all three. Google and DeepInfra support subsets.
|
|
11
|
+
Each method follows the same pattern: pass input, get output,
|
|
12
|
+
copy the result somewhere useful.
|
|
13
|
+
|
|
14
|
+
#### How it works
|
|
15
|
+
|
|
16
|
+
When you want to convert text to speech, call
|
|
17
|
+
[`LLM::Audio#create_speech`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_speech).
|
|
18
|
+
The provider returns an audio clip as a
|
|
19
|
+
[`LLM::URIData`](https://r.uby.dev/api-docs/llm.rb/LLM/URIData.html)
|
|
20
|
+
object. The generated audio can be copied
|
|
21
|
+
to a file or streamed directly. Each provider supports different
|
|
22
|
+
output formats and voice options:
|
|
23
|
+
|
|
24
|
+
```ruby
|
|
25
|
+
llm = LLM.openai(key: ENV["KEY"])
|
|
26
|
+
res = llm.audio.create_speech(input: "Hello world")
|
|
27
|
+
IO.copy_stream res.audio.decoded, "helloworld.mp3"
|
|
28
|
+
```
|
|
29
|
+
|
|
30
|
+
#### Why would I use it?
|
|
31
|
+
|
|
32
|
+
Audio support lets you build voice interfaces, add accessibility
|
|
33
|
+
features, and work across languages without wiring up a separate
|
|
34
|
+
speech service. Generate audio for notifications, narrate written
|
|
35
|
+
content, or add speech output to an existing application with a
|
|
36
|
+
single method call.
|
|
37
|
+
|
|
38
|
+
#### Notes
|
|
39
|
+
|
|
40
|
+
OpenAI has full audio support. Google and DeepInfra have partial
|
|
41
|
+
support. The
|
|
42
|
+
[`LLM::Audio#create_speech`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_speech)
|
|
43
|
+
method returns a
|
|
44
|
+
[`LLM::URIData`](https://r.uby.dev/api-docs/llm.rb/LLM/URIData.html)
|
|
45
|
+
object.
|
|
46
|
+
|
|
47
|
+
### Transcription
|
|
48
|
+
|
|
49
|
+
#### Overview
|
|
50
|
+
|
|
51
|
+
Transcription turns an audio file into text. Pass a file path
|
|
52
|
+
to
|
|
53
|
+
[`LLM::Audio#create_transcription`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_transcription)
|
|
54
|
+
and get back the spoken content as
|
|
55
|
+
a string. OpenAI, Google, and DeepInfra support it. Transcribe
|
|
56
|
+
meeting notes, voice memos, or podcast episodes for search and
|
|
57
|
+
processing. The response is plain text you can feed into any
|
|
58
|
+
downstream pipeline.
|
|
59
|
+
|
|
60
|
+
#### How it works
|
|
61
|
+
|
|
62
|
+
When you want to transcribe audio into text, call
|
|
63
|
+
[`LLM::Audio#create_transcription`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_transcription).
|
|
64
|
+
The provider processes the audio and returns the transcribed text.
|
|
65
|
+
The file can be a local path or a URL depending on provider support:
|
|
66
|
+
|
|
67
|
+
```ruby
|
|
68
|
+
llm = LLM.google(key: ENV["KEY"])
|
|
69
|
+
res = llm.audio.create_transcription(file: "helloworld.mp3")
|
|
70
|
+
res.text # => "Hello world"
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
#### Why would I use it?
|
|
74
|
+
|
|
75
|
+
Transcribe recorded meetings, voice memos, or podcast episodes for
|
|
76
|
+
search and processing. The returned text feeds directly into search
|
|
77
|
+
indexes, summarization pipelines, or downstream extraction. Spoken
|
|
78
|
+
content becomes as queryable as written text.
|
|
79
|
+
|
|
80
|
+
#### Notes
|
|
81
|
+
|
|
82
|
+
OpenAI has full audio support. Google and DeepInfra have partial
|
|
83
|
+
support.
|
|
84
|
+
|
|
85
|
+
### Translation
|
|
86
|
+
|
|
87
|
+
#### Overview
|
|
88
|
+
|
|
89
|
+
Translation transcribes audio and translates it into English in
|
|
90
|
+
one step. Pass a file to
|
|
91
|
+
[`LLM::Audio#create_translation`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_translation)
|
|
92
|
+
and get back the
|
|
93
|
+
translated text. OpenAI and Google support it. Translate
|
|
94
|
+
multilingual podcasts, interviews, or any audio where you need
|
|
95
|
+
the content in English without running a separate translation
|
|
96
|
+
pipeline.
|
|
97
|
+
|
|
98
|
+
#### How it works
|
|
99
|
+
|
|
100
|
+
When you want to translate spoken audio into English, call
|
|
101
|
+
[`LLM::Audio#create_translation`](https://r.uby.dev/api-docs/llm.rb/LLM/Audio.html#create_translation).
|
|
102
|
+
The provider transcribes the spoken language and translates the
|
|
103
|
+
result into English in a single operation. The returned text is the
|
|
104
|
+
English translation:
|
|
105
|
+
|
|
106
|
+
```ruby
|
|
107
|
+
llm = LLM.google(key: ENV["KEY"])
|
|
108
|
+
res = llm.audio.create_translation(file: "bomdia.mp3")
|
|
109
|
+
res.text # => "Good day"
|
|
110
|
+
```
|
|
111
|
+
|
|
112
|
+
#### Why would I use it?
|
|
113
|
+
|
|
114
|
+
Translate a podcast or interview into another language without a
|
|
115
|
+
separate speech service. The provider handles both transcription and
|
|
116
|
+
translation in a single call, so you get English text from any
|
|
117
|
+
supported source language without chaining two operations together.
|
|
118
|
+
|
|
119
|
+
#### Notes
|
|
120
|
+
|
|
121
|
+
Each method works independently and each provider supports a
|
|
122
|
+
different subset.
|
|
@@ -0,0 +1,99 @@
|
|
|
1
|
+
|
|
2
|
+
## LLM::Cost
|
|
3
|
+
|
|
4
|
+
### Introduction
|
|
5
|
+
|
|
6
|
+
#### Overview
|
|
7
|
+
|
|
8
|
+
[`LLM::Cost`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html)
|
|
9
|
+
represents the approximate cost of a conversation. It breaks the
|
|
10
|
+
total down by token type, so you can see how much was spent on input,
|
|
11
|
+
output, cached tokens, reasoning, audio, and images. Cost is computed
|
|
12
|
+
from token usage and the pricing data shipped in the model registry.
|
|
13
|
+
|
|
14
|
+
#### How it works
|
|
15
|
+
|
|
16
|
+
When you want to know what a conversation cost so far, call
|
|
17
|
+
[`LLM::Context#cost`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#cost-instance_method)
|
|
18
|
+
(or
|
|
19
|
+
[`LLM::Agent#cost`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#cost-instance_method))
|
|
20
|
+
and read the breakdown. The REPL shows this live in its status bar
|
|
21
|
+
after every turn:
|
|
22
|
+
|
|
23
|
+
```ruby
|
|
24
|
+
llm = LLM.deepseek(key: ENV["KEY"])
|
|
25
|
+
ctx = LLM::Context.new(llm)
|
|
26
|
+
ctx.talk "Hello"
|
|
27
|
+
|
|
28
|
+
cost = ctx.cost
|
|
29
|
+
cost.input # => 0.0000042
|
|
30
|
+
cost.output # => 0.0000084
|
|
31
|
+
cost.total # => 0.0000126
|
|
32
|
+
cost.to_s # => "0.0000126"
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
#### Why would I use it?
|
|
36
|
+
|
|
37
|
+
Cost tracking matters in production. Monitoring spend per
|
|
38
|
+
conversation, per agent, or per provider tells you which workflows
|
|
39
|
+
are expensive and when to switch models. Log a structured breakdown
|
|
40
|
+
with
|
|
41
|
+
[`LLM::Cost#to_h`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#to_h-instance_method),
|
|
42
|
+
or read
|
|
43
|
+
[`LLM::Cost#total`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#total-instance_method)
|
|
44
|
+
for a single number.
|
|
45
|
+
|
|
46
|
+
#### Notes
|
|
47
|
+
|
|
48
|
+
Cost is an approximation based on the pricing in the model registry.
|
|
49
|
+
[`LLM::Cost.from`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#from-class_method)
|
|
50
|
+
returns an empty cost when the model or registry cannot be found, so
|
|
51
|
+
a missing model never crashes your code.
|
|
52
|
+
|
|
53
|
+
### Reading the breakdown
|
|
54
|
+
|
|
55
|
+
#### Overview
|
|
56
|
+
|
|
57
|
+
[`LLM::Cost`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html)
|
|
58
|
+
exposes each cost component as a reader, plus
|
|
59
|
+
[`LLM::Cost#total`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#total-instance_method),
|
|
60
|
+
[`LLM::Cost#to_h`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#to_h-instance_method),
|
|
61
|
+
and
|
|
62
|
+
[`LLM::Cost#to_s`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#to_s-instance_method).
|
|
63
|
+
|
|
64
|
+
#### How it works
|
|
65
|
+
|
|
66
|
+
Each component is a Float, or `nil` when no tokens of that type were
|
|
67
|
+
used. The
|
|
68
|
+
[`LLM::Cost#to_h`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#to_h-instance_method)
|
|
69
|
+
method returns a Hash with only the non-nil components and the total:
|
|
70
|
+
|
|
71
|
+
```ruby
|
|
72
|
+
cost = ctx.cost
|
|
73
|
+
|
|
74
|
+
cost.input
|
|
75
|
+
cost.output
|
|
76
|
+
cost.cache_read
|
|
77
|
+
cost.cache_write
|
|
78
|
+
cost.reasoning
|
|
79
|
+
cost.input_audio
|
|
80
|
+
cost.output_audio
|
|
81
|
+
cost.input_image
|
|
82
|
+
|
|
83
|
+
cost.to_h # => {input: 4.2e-06, output: 8.4e-06, total: 1.26e-05}
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
#### Why would I use it?
|
|
87
|
+
|
|
88
|
+
The per-component breakdown shows where the money goes. High cache
|
|
89
|
+
read costs suggest a conversation benefits from prompt caching.
|
|
90
|
+
High reasoning costs point at a model that thinks a lot. Log
|
|
91
|
+
[`LLM::Cost#to_h`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#to_h-instance_method)
|
|
92
|
+
at the end of a session to keep a spend trail.
|
|
93
|
+
|
|
94
|
+
#### Notes
|
|
95
|
+
|
|
96
|
+
[`LLM::Cost#to_s`](https://r.uby.dev/api-docs/llm.rb/LLM/Cost.html#to_s-instance_method)
|
|
97
|
+
returns the total in a compact, human-friendly format
|
|
98
|
+
(`"0.0000126"`). Components that were not used are `nil`, so sum
|
|
99
|
+
them with `compact` if you aggregate across conversations.
|
|
@@ -0,0 +1,89 @@
|
|
|
1
|
+
|
|
2
|
+
## Images
|
|
3
|
+
|
|
4
|
+
### Introduction
|
|
5
|
+
|
|
6
|
+
#### Overview
|
|
7
|
+
|
|
8
|
+
A handful of providers can generate images from a text prompt.
|
|
9
|
+
OpenAI, Google, xAI, DeepInfra, and DeepSeek all support it.
|
|
10
|
+
OpenAI, xAI, and DeepInfra also let you edit existing images.
|
|
11
|
+
The API is the same across providers, so switching between them
|
|
12
|
+
requires no code changes.
|
|
13
|
+
|
|
14
|
+
#### How it works
|
|
15
|
+
|
|
16
|
+
The
|
|
17
|
+
[`LLM::Images#create`](https://r.uby.dev/api-docs/llm.rb/LLM/Images.html#create)
|
|
18
|
+
method sends a prompt to the provider and
|
|
19
|
+
returns the result as a
|
|
20
|
+
[`LLM::URIData`](https://r.uby.dev/api-docs/llm.rb/LLM/URIData.html)
|
|
21
|
+
object. The same API works
|
|
22
|
+
across providers: swap
|
|
23
|
+
[`LLM.openai`](https://r.uby.dev/api-docs/llm.rb/LLM.html#openai-class_method) for
|
|
24
|
+
[`LLM.xai`](https://r.uby.dev/api-docs/llm.rb/LLM.html#xai-class_method) and the rest
|
|
25
|
+
of the code is identical.
|
|
26
|
+
|
|
27
|
+
```ruby
|
|
28
|
+
require "llm"
|
|
29
|
+
|
|
30
|
+
llm = LLM.openai(key: ENV["KEY"])
|
|
31
|
+
res = llm.images.create(prompt: "a dog on a rocket to the moon")
|
|
32
|
+
IO.copy_stream res.images[0], "dogrocket.png"
|
|
33
|
+
```
|
|
34
|
+
|
|
35
|
+
#### Why would I use it?
|
|
36
|
+
|
|
37
|
+
Image generation and editing let the model produce visual output
|
|
38
|
+
directly from your prompts. The API is the same across providers,
|
|
39
|
+
so switching between OpenAI and xAI requires changing one line.
|
|
40
|
+
Prototype visual concepts, generate assets, or augment datasets
|
|
41
|
+
without wiring up a separate image API.
|
|
42
|
+
|
|
43
|
+
#### Notes
|
|
44
|
+
|
|
45
|
+
Google only supports image generation, not edits. DeepSeek
|
|
46
|
+
generates SVGs rather than raster images. DeepSeek can also
|
|
47
|
+
maintain a session across multiple generations through the
|
|
48
|
+
`agent` parameter on the response object.
|
|
49
|
+
|
|
50
|
+
### Editing
|
|
51
|
+
|
|
52
|
+
#### Overview
|
|
53
|
+
|
|
54
|
+
Image editing takes an existing image and a text prompt, then
|
|
55
|
+
produces a modified version. You can add objects, change colours,
|
|
56
|
+
or alter the scene while keeping the original composition.
|
|
57
|
+
OpenAI, xAI, and DeepInfra support raster image edits. DeepSeek
|
|
58
|
+
generates SVG documents which can be refined with follow-up
|
|
59
|
+
text-to-image prompts, giving you iterative vector editing.
|
|
60
|
+
|
|
61
|
+
#### How it works
|
|
62
|
+
|
|
63
|
+
The
|
|
64
|
+
[`LLM::Images#edit`](https://r.uby.dev/api-docs/llm.rb/LLM/Images.html#edit)
|
|
65
|
+
method takes a prompt and an image path. It
|
|
66
|
+
returns a modified image as a
|
|
67
|
+
[`LLM::URIData`](https://r.uby.dev/api-docs/llm.rb/LLM/URIData.html)
|
|
68
|
+
object. Copy the result
|
|
69
|
+
to a file the same way you would with generated images.
|
|
70
|
+
|
|
71
|
+
```ruby
|
|
72
|
+
llm = LLM.openai(key: ENV["KEY"])
|
|
73
|
+
res = llm.images.edit(prompt: "add a mustache", image: "self.jpg")
|
|
74
|
+
IO.copy_stream res.images[0], "mustache.png"
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
#### Why would I use it?
|
|
78
|
+
|
|
79
|
+
Editing lets the model modify existing images rather than starting
|
|
80
|
+
from scratch. DeepSeek's SVG output is particularly useful here
|
|
81
|
+
because vector graphics can be refined iteratively. Make targeted
|
|
82
|
+
adjustments: add objects, change colours, or alter the scene,
|
|
83
|
+
all while keeping the original composition intact.
|
|
84
|
+
|
|
85
|
+
#### Notes
|
|
86
|
+
|
|
87
|
+
Google does not support image edits. DeepSeek can maintain a
|
|
88
|
+
session across multiple generations through the `agent` parameter
|
|
89
|
+
on the response object.
|