llm.rb 13.1.0 → 14.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (90) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +320 -0
  3. data/README.md +340 -31
  4. data/bin/llm.rb +36 -12
  5. data/data/anthropic.json +206 -263
  6. data/data/bedrock.json +2138 -1860
  7. data/data/deepinfra.json +1003 -624
  8. data/data/deepseek.json +38 -34
  9. data/data/google.json +1079 -371
  10. data/data/mistral.json +448 -368
  11. data/data/moonshot.json +384 -0
  12. data/data/openai.json +974 -1343
  13. data/data/xai.json +154 -126
  14. data/data/zai.json +191 -191
  15. data/lib/llm/agent.rb +47 -14
  16. data/lib/llm/context.rb +71 -88
  17. data/lib/llm/cost.rb +23 -17
  18. data/lib/llm/error.rb +0 -8
  19. data/lib/llm/function/async/task.rb +2 -0
  20. data/lib/llm/function/fiber/task.rb +2 -0
  21. data/lib/llm/function/fork/task.rb +2 -0
  22. data/lib/llm/function/ractor/task.rb +2 -0
  23. data/lib/llm/function/sequential/group.rb +4 -1
  24. data/lib/llm/function/sequential/task.rb +1 -1
  25. data/lib/llm/function/task.rb +4 -0
  26. data/lib/llm/function/thread/task.rb +2 -0
  27. data/lib/llm/function.rb +32 -4
  28. data/lib/llm/guard/loop.rb +89 -0
  29. data/lib/llm/guard/null.rb +19 -0
  30. data/lib/llm/guard.rb +61 -0
  31. data/lib/llm/provider.rb +36 -0
  32. data/lib/llm/providers/anthropic/stream_parser.rb +1 -0
  33. data/lib/llm/providers/anthropic.rb +1 -8
  34. data/lib/llm/providers/bedrock/stream_parser.rb +1 -0
  35. data/lib/llm/providers/bedrock.rb +1 -8
  36. data/lib/llm/providers/google/stream_parser.rb +1 -0
  37. data/lib/llm/providers/google.rb +1 -8
  38. data/lib/llm/providers/moonshot.rb +76 -0
  39. data/lib/llm/providers/ollama.rb +1 -8
  40. data/lib/llm/providers/openai/responses/stream_parser.rb +1 -0
  41. data/lib/llm/providers/openai/responses.rb +6 -8
  42. data/lib/llm/providers/openai/stream_parser.rb +1 -0
  43. data/lib/llm/providers/openai.rb +3 -10
  44. data/lib/llm/repl/bar.rb +4 -3
  45. data/lib/llm/repl/buffer.rb +42 -15
  46. data/lib/llm/repl/color.rb +78 -0
  47. data/lib/llm/repl/input/char.rb +46 -0
  48. data/lib/llm/repl/input/row.rb +39 -0
  49. data/lib/llm/repl/input.rb +251 -66
  50. data/lib/llm/repl/markdown/table.rb +6 -2
  51. data/lib/llm/repl/markdown.rb +31 -5
  52. data/lib/llm/repl/status.rb +38 -3
  53. data/lib/llm/repl/stream.rb +16 -4
  54. data/lib/llm/repl/walker.rb +3 -2
  55. data/lib/llm/repl/window.rb +25 -5
  56. data/lib/llm/repl.rb +29 -13
  57. data/lib/llm/stream.rb +8 -7
  58. data/lib/llm/tool.rb +29 -0
  59. data/lib/llm/transformer/null.rb +21 -0
  60. data/lib/llm/transformer.rb +55 -0
  61. data/lib/llm/version.rb +1 -1
  62. data/lib/llm.rb +12 -2
  63. data/llm.gemspec +1 -0
  64. data/resources/deepdive/advanced/cancellation.md +74 -0
  65. data/resources/deepdive/advanced/compaction.md +83 -0
  66. data/resources/deepdive/advanced/context.md +267 -0
  67. data/resources/deepdive/advanced/guard.md +371 -0
  68. data/resources/deepdive/advanced/tracer.md +180 -0
  69. data/resources/deepdive/advanced/transformer.md +67 -0
  70. data/resources/deepdive/advanced/transports.md +45 -0
  71. data/resources/deepdive/everything_else/audio.md +122 -0
  72. data/resources/deepdive/everything_else/cost.md +99 -0
  73. data/resources/deepdive/everything_else/images.md +89 -0
  74. data/resources/deepdive/everything_else/object.md +108 -0
  75. data/resources/deepdive/everything_else/ocr.md +48 -0
  76. data/resources/deepdive/fundamentals/agents.md +202 -0
  77. data/resources/deepdive/fundamentals/builtin_tools.md +191 -0
  78. data/resources/deepdive/fundamentals/concurrency.md +104 -0
  79. data/resources/deepdive/fundamentals/database.md +449 -0
  80. data/resources/deepdive/fundamentals/embeddings.md +157 -0
  81. data/resources/deepdive/fundamentals/repl.md +87 -0
  82. data/resources/deepdive/fundamentals/schema.md +61 -0
  83. data/resources/deepdive/fundamentals/skills.md +106 -0
  84. data/resources/deepdive/fundamentals/stream.md +110 -0
  85. data/resources/deepdive/fundamentals/tools.md +265 -0
  86. data/resources/deepdive/protocols/a2a.md +106 -0
  87. data/resources/deepdive/protocols/mcp.md +111 -0
  88. data/resources/deepdive.md +7 -1
  89. metadata +36 -3
  90. data/lib/llm/loop_guard.rb +0 -107
@@ -0,0 +1,371 @@
1
+
2
+ ## Guard
3
+
4
+ ### Introduction
5
+
6
+ #### Overview
7
+
8
+ [`LLM::Guard`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard.html)
9
+ is the superclass for context-level supervisors. A guard is bound
10
+ to a context and inspects each pending tool call before it runs.
11
+ It can let the call through, cancel it, block it with an error, or
12
+ answer for it with a synthesized result. Beyond loop detection,
13
+ guards handle policy, validation, quotas, cost control, caching,
14
+ and approval workflows.
15
+
16
+ #### How it works
17
+
18
+ A guard is a subclass of
19
+ [`LLM::Guard`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard.html)
20
+ that implements
21
+ [`LLM::Guard#call`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard.html#call-instance_method).
22
+ The guard is stamped onto the functions the context binds, and when a
23
+ function's task runs the guard is checked on the calling thread before
24
+ the tool is handed to the strategy. `call` receives the pending
25
+ `function:` and returns an
26
+ [`LLM::Function::Return`](https://r.uby.dev/api-docs/llm.rb/LLM/Function/Return.html)
27
+ to close that single call, or `nil` to let it run. The guard inspects
28
+ the current conversation through the `messages` helper and reads the
29
+ pending call through the function's own accessors.
30
+
31
+ Configure a guard by passing a class through the `guard:` option;
32
+ options given in `guard_options:` are forwarded to `call` as keyword
33
+ arguments:
34
+
35
+ ```ruby
36
+ class RateLimitGuard < LLM::Guard
37
+ def call(function:, limit: 10)
38
+ if messages.count(&:tool_return?) >= limit
39
+ function.return(error: true, type: "guard_error",
40
+ message: "too many tool calls")
41
+ end
42
+ end
43
+ end
44
+
45
+ agent = LLM::Agent.new(
46
+ llm,
47
+ guard: RateLimitGuard,
48
+ guard_options: {limit: 3}
49
+ )
50
+ agent.talk "Research the market", tools: [FetchNews, FetchStocks]
51
+ ```
52
+
53
+ #### Why would I use it?
54
+
55
+ A guard is the one hook that sees every pending tool call and the
56
+ full conversation at once. It can observe what the model is about
57
+ to do, cancel a call that needs approval, block one that violates
58
+ policy, answer a cheap question without running a tool, or stop
59
+ work when a budget is spent. Because the guard runs before the
60
+ tool, anything it intercepts never executes.
61
+
62
+ #### Notes
63
+
64
+ [`LLM::Guard::Null`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard/Null.html)
65
+ is the default guard for
66
+ [`LLM::Context`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html);
67
+ it never blocks tool work.
68
+ [`LLM::Context#guard`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#guard)
69
+ returns the configured guard class. Because `call` returns an
70
+ [`LLM::Function::Return`](https://r.uby.dev/api-docs/llm.rb/LLM/Function/Return.html)
71
+ rather than a warning string, a guard can block an individual tool
72
+ call while the rest of the batch still executes. The guard is stamped
73
+ onto every function the context binds, so it runs whenever a task is
74
+ spawned — including tool calls a stream queues itself during a
75
+ streaming turn. A blocked call yields its return without executing.
76
+
77
+ The runtime binds a guard instance to the context and stamps it onto
78
+ the functions it resolves, so a guard cannot carry state in instance
79
+ variables between calls. Anything a guard needs to remember, like
80
+ how many calls already ran, must come from the conversation
81
+ (`messages`) or from class-level state.
82
+
83
+ ### Inspect
84
+
85
+ #### Overview
86
+
87
+ The guard receives the pending
88
+ [`LLM::Function`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html)
89
+ and can read what the model is about to do before anything runs.
90
+ The function exposes its `name`, `arguments`, and parameter schema,
91
+ so a guard can base its decision on the actual call rather than on
92
+ the whole conversation. The guard also holds the context, so the
93
+ message history, token usage, and cost are visible too.
94
+
95
+ #### How it works
96
+
97
+ When you want to inspect a pending call before it runs, read the
98
+ function's accessors inside `call`.
99
+ [`LLM::Function#name`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html#name)
100
+ is the tool name, and
101
+ [`LLM::Function#arguments`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html#arguments)
102
+ is an
103
+ [`LLM::Object`](https://r.uby.dev/api-docs/llm.rb/LLM/Object.html)
104
+ with the parsed arguments. The full parameter schema is available
105
+ through
106
+ [`LLM::Function#params`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html#params).
107
+ Returning `nil` lets the call run, so an inspection-only guard is a
108
+ pure observer:
109
+
110
+ ```ruby
111
+ class AuditGuard < LLM::Guard
112
+ def call(function:)
113
+ warn "pending: #{function.name}(#{function.arguments.inspect})"
114
+ nil # let it run
115
+ end
116
+ end
117
+ ```
118
+
119
+ #### Why would I use it?
120
+
121
+ Inspecting the pending call is the basis for policy, validation,
122
+ and audit. A guard can log every call into a telemetry stream, or
123
+ reject a single call because its arguments violate a rule while
124
+ letting every other call through. The same read underpins the
125
+ richer decisions in the rest of this document.
126
+
127
+ #### Notes
128
+
129
+ The guard sees the parsed arguments exactly as the model requested
130
+ them. Reading them costs nothing and never executes the tool.
131
+ Returning `nil` means the call proceeds normally.
132
+
133
+ ### Cancel
134
+
135
+ #### Overview
136
+
137
+ Cancelling is distinct from blocking with an error.
138
+ [`LLM::Function#cancel`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html#cancel)
139
+ produces a return the model reads as a declined call, not a failed
140
+ one. The conversation stays honest: the model asked for something,
141
+ and the runtime declined it with a reason.
142
+
143
+ #### How it works
144
+
145
+ When you want to cancel a single pending call, call the
146
+ [`LLM::Function#cancel`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html#cancel)
147
+ method with a `reason:` and return the result from `call`. The
148
+ pending call never executes, and the rest of the batch still runs.
149
+ A guard can cancel one call in a two-call batch and the other call
150
+ still executes:
151
+
152
+ ```ruby
153
+ class ApprovalGuard < LLM::Guard
154
+ def call(function:)
155
+ if function.name == "delete-file"
156
+ function.cancel(reason: "delete requires human approval")
157
+ end
158
+ end
159
+ end
160
+ ```
161
+
162
+ #### Why would I use it?
163
+
164
+ Cancellation suits approval workflows and environment rules. A call
165
+ that needs a human in the loop, a tool that is disabled for a
166
+ particular user, or an operation that must not run in production
167
+ can all be declined without pretending the tool failed. Because
168
+ cancellation is per call, it declines only the offending tool and
169
+ leaves the rest of the batch intact.
170
+
171
+ #### Notes
172
+
173
+ The model receives
174
+ `{cancelled: true, reason: "delete requires human approval"}` as
175
+ that call's result and can react, for example by asking for
176
+ permission or skipping the operation. A cancelled call still counts
177
+ as resolved, like any tool return, so the model keeps moving. The
178
+ reason string is what the model sees, so write it as guidance
179
+ ("ask the user first") rather than a raw error dump.
180
+
181
+ ### Block
182
+
183
+ #### Overview
184
+
185
+ Blocking is how a guard enforces policy. A blocked call never runs,
186
+ and the model receives an in-band error explaining why. Unlike a
187
+ cancellation, an error tells the model the tool could not produce
188
+ a result at all.
189
+
190
+ #### How it works
191
+
192
+ When you want to block a call, return an
193
+ [`LLM::Function::Return`](https://r.uby.dev/api-docs/llm.rb/LLM/Function/Return.html)
194
+ with `error: true`. The `type` and `message` are free-form, and the
195
+ model sees them inside the return value. Returning `nil` lets the
196
+ call through:
197
+
198
+ ```ruby
199
+ class PolicyGuard < LLM::Guard
200
+ def call(function:)
201
+ if function.name == "shell"
202
+ function.return(error: true, type: "policy_error",
203
+ message: "shell is disabled")
204
+ end
205
+ end
206
+ end
207
+ ```
208
+
209
+ #### Why would I use it?
210
+
211
+ Blocking denies a dangerous tool, rejects an out-of-policy
212
+ argument, or stops a quota violation before it happens. The model
213
+ receives the return and can adapt, so the conversation stays valid.
214
+
215
+ #### Notes
216
+
217
+ A blocked call never executes. The return it produces is sent back
218
+ through the model like any tool result, so the model sees why the
219
+ call was blocked and can change course.
220
+
221
+ ### Answer
222
+
223
+ #### Overview
224
+
225
+ A guard's return is injected into the conversation as if the tool
226
+ had executed, so a guard can answer for a tool that never runs.
227
+ From the model's side, a synthesized result is indistinguishable
228
+ from a real one.
229
+
230
+ #### How it works
231
+
232
+ When you want to answer a pending call yourself, return a value
233
+ through the
234
+ [`LLM::Function#return`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html)
235
+ helper. The value you pass becomes the tool's result as-is, so it
236
+ must look like a plausible answer:
237
+
238
+ ```ruby
239
+ class CacheGuard < LLM::Guard
240
+ CACHE = {"get-weather:tokyo" => {forecast: "sunny"}}
241
+
242
+ def call(function:)
243
+ key = "#{function.name}:#{function.arguments[:city]}"
244
+ function.return(CACHE[key]) if CACHE.key?(key)
245
+ end
246
+ end
247
+ ```
248
+
249
+ #### Why would I use it?
250
+
251
+ Standing in for a tool is how you cache expensive calls, mock tools
252
+ in tests, or degrade gracefully when a service is down. The same
253
+ hook can return a fixed answer for tools that should not run in a
254
+ given environment.
255
+
256
+ #### Notes
257
+
258
+ [`LLM::Function#unavailable`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html)
259
+ marks a tool as not found, and
260
+ [`LLM::Function#budget_spent`](https://r.uby.dev/api-docs/llm.rb/LLM/Function.html)
261
+ reports that the tool budget is exhausted. Both are returns too,
262
+ so they flow back to the model like any tool result. A guard that
263
+ synthesizes a result runs instead of the tool, so side effects the
264
+ tool would have performed, like a database write, are skipped.
265
+
266
+ ### Budget
267
+
268
+ #### Overview
269
+
270
+ A guard can read the accumulated usage and cost of the conversation
271
+ through the context. That makes it the natural place to enforce a
272
+ hard ceiling: stop calling tools once a cost or token budget is
273
+ spent, or once the context window is nearly full.
274
+
275
+ #### How it works
276
+
277
+ When you want to enforce a budget, compare
278
+ [`LLM::Context#usage`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#usage)
279
+ or
280
+ [`LLM::Context#cost`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#cost-instance_method)
281
+ against a limit inside `call`. Pass the limit through
282
+ `guard_options:` so it can be tuned per context. Returning a
283
+ cancellation or an error closes the call before it runs:
284
+
285
+ ```ruby
286
+ class BudgetGuard < LLM::Guard
287
+ def call(function:, limit: 0.05)
288
+ if ctx.cost.total >= limit
289
+ function.cancel(reason: "cost ceiling reached, ask before continuing")
290
+ end
291
+ end
292
+ end
293
+
294
+ ctx = LLM::Context.new(
295
+ llm,
296
+ guard: BudgetGuard,
297
+ guard_options: {limit: 0.01}
298
+ )
299
+ ```
300
+
301
+ #### Why would I use it?
302
+
303
+ Budgets matter for long autonomous runs where the model decides how
304
+ many tools to call. A cost ceiling keeps a runaway agent from
305
+ spending money, and a quota derived from `messages` (as in the
306
+ rate-limit example above) keeps a single user from exhausting
307
+ shared resources. The same comparison works with
308
+ [`LLM::Context#usage`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#usage)
309
+ against
310
+ [`LLM::Context#context_window`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#context_window)
311
+ to keep the conversation inside the model's window.
312
+
313
+ #### Notes
314
+
315
+ The cost reported by
316
+ [`LLM::Context#cost`](https://r.uby.dev/api-docs/llm.rb/LLM/Context.html#cost-instance_method)
317
+ reflects the conversation so far, so a limit is enforced per call:
318
+ once the ceiling is crossed, the next pending call is declined. The
319
+ guard is not a replacement for
320
+ [`LLM::Agent.tool_budget`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#tool_budget-class_method),
321
+ which caps the number of tool calls in a single turn. The two
322
+ compose: the budget caps call count, and a guard enforces cost.
323
+
324
+ ### Loop
325
+
326
+ #### Overview
327
+
328
+ [`LLM::Guard::Loop`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard/Loop.html)
329
+ is the built-in loop-detection guard. It reduces each assistant
330
+ tool call to a `[tool name, arguments]` signature and checks whether
331
+ the tail of the sequence is repeating.
332
+
333
+ #### How it works
334
+
335
+ When you want to detect repeated tool-call patterns, enable
336
+ [`LLM::Guard::Loop`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard/Loop.html)
337
+ and tune the `threshold:` option, which is the number of repeated
338
+ patterns required before the guard intervenes (default `3`). When
339
+ the guard detects a repeat, it returns an in-band
340
+ [`LLM::Function::Return`](https://r.uby.dev/api-docs/llm.rb/LLM/Function/Return.html)
341
+ with type `"guard_error"` and a message that tells the model it is
342
+ stuck and should change approach:
343
+
344
+ ```ruby
345
+ ctx = LLM::Context.new(
346
+ llm,
347
+ guard: LLM::Guard::Loop,
348
+ guard_options: {threshold: 2}
349
+ )
350
+ ctx.talk "Research the market", tools: [FetchNews, FetchStocks]
351
+ ```
352
+
353
+ #### Why would I use it?
354
+
355
+ Loop detection matters for long, autonomous agent runs. Without it,
356
+ a model that repeats a tool call with the same arguments can
357
+ bounce between calls forever. The guard turns that into a bounded
358
+ conversation: after the threshold, the model receives a message
359
+ telling it to stop and try a different strategy.
360
+
361
+ #### Notes
362
+
363
+ [`LLM::Agent`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html)
364
+ enables
365
+ [`LLM::Guard::Loop`](https://r.uby.dev/api-docs/llm.rb/LLM/Guard/Loop.html)
366
+ by default, so agents get loop protection without configuration.
367
+ A custom guard can be passed through the `guard:` option to replace
368
+ the loop guard entirely. Guards and the agent's tool budget
369
+ complement each other: a guard blocks work that looks stuck, while
370
+ [`LLM::Agent.tool_budget`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#tool_budget-class_method)
371
+ caps the total number of tool calls in a single turn.
@@ -0,0 +1,180 @@
1
+
2
+ ## Tracer
3
+
4
+ ### Introduction
5
+
6
+ #### Overview
7
+
8
+ Tracers let you observe what the runtime is doing. They hook into
9
+ requests, tool calls, compactions, and other events. Debug a
10
+ misbehaving agent, monitor request latency, or export spans to
11
+ an observability backend. A provider-wide tracer intercepts every
12
+ request through that provider. An agent-local tracer only covers
13
+ requests made by that agent.
14
+
15
+ #### How it works
16
+
17
+ When you want to observe runtime events, subclass
18
+ [`LLM::Tracer`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer.html)
19
+ and implement the hooks you need, or use one of the built-in
20
+ tracers. The built-in tracers include
21
+ [`LLM::Tracer::Logger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html)
22
+ (writes structured JSON to stdout or a file),
23
+ [`LLM::Tracer::PrettyLogger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/PrettyLogger.html)
24
+ (writes human-readable single-line logs to stderr), and
25
+ [`LLM::Tracer::Telemetry`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Telemetry.html)
26
+ (exports spans via OTLP for OpenTelemetry). Attach a tracer to a
27
+ provider or an agent:
28
+
29
+ ```ruby
30
+ llm = LLM.deepseek(key: ENV["KEY"])
31
+ llm.tracer = LLM::Tracer::PrettyLogger.new(llm)
32
+ agent = LLM::Agent.new(llm)
33
+ agent.talk "Hello"
34
+ ```
35
+
36
+ #### Why would I use it?
37
+
38
+ Tracers give you visibility into what the runtime is doing. Debug
39
+ a misbehaving agent by tracing every request it makes. Monitor
40
+ request latency and token usage across providers. Export spans to
41
+ OpenTelemetry for integration with existing observability pipelines.
42
+ [`LLM::Tracer::PrettyLogger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/PrettyLogger.html)
43
+ is the best choice during development for compact, human-readable
44
+ output.
45
+ [`LLM::Tracer::Logger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html)
46
+ provides structured JSON, and
47
+ [`LLM::Tracer::Telemetry`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Telemetry.html)
48
+ exports spans to OpenTelemetry for production observability.
49
+
50
+ #### Notes
51
+
52
+ The tracer is extensible. You can implement custom hooks for any
53
+ runtime event. The scope can be an individual agent or every
54
+ request a provider makes. Three built-in tracers are available:
55
+ [`LLM::Tracer::PrettyLogger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/PrettyLogger.html)
56
+ (human-readable),
57
+ [`LLM::Tracer::Logger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html)
58
+ (structured JSON), and
59
+ [`LLM::Tracer::Telemetry`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Telemetry.html)
60
+ (OpenTelemetry).
61
+
62
+ ### Provider
63
+
64
+ #### Overview
65
+
66
+ A provider-wide tracer intercepts every request made through that
67
+ provider. All agents sharing the same provider share the same
68
+ tracer. Use this to trace at the infrastructure level without
69
+ configuring each agent individually.
70
+
71
+ #### How it works
72
+
73
+ When you want every request through a provider to be traced, set
74
+ the tracer on the provider directly. Every request made through
75
+ that provider, regardless of which agent initiates it, flows
76
+ through the same tracer hooks. The provider holds a reference to the
77
+ tracer and passes it to every new context it creates. This ensures
78
+ consistent observability without configuring each agent individually.
79
+
80
+ ```ruby
81
+ llm = LLM.deepseek(key: ENV["KEY"])
82
+ llm.tracer = LLM::Tracer::Logger.new(llm, io: $stdout)
83
+ ```
84
+
85
+ #### Why would I use it?
86
+
87
+ A provider-wide tracer captures every request at the infrastructure level.
88
+ All agents sharing the same provider share the same tracer.
89
+
90
+ #### Notes
91
+
92
+ The tracer can also write to a file with the `path:` option to
93
+ [`LLM::Tracer::Logger.new`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html#initialize-instance_method).
94
+
95
+ ### Agent
96
+
97
+ #### Overview
98
+
99
+ An agent-local tracer only covers requests made by that agent.
100
+ Attach it via the `tracer:` keyword argument to
101
+ [`LLM::Agent.new`](https://r.uby.dev/api-docs/llm.rb/LLM/Agent.html#initialize-instance_method)
102
+ and it follows that agent wherever it goes. Different agents can
103
+ have different tracers.
104
+
105
+ #### How it works
106
+
107
+ When you want a tracer for a specific agent, pass it to the agent
108
+ on creation. Only requests made by
109
+ that agent flow through the tracer, leaving other agents on the
110
+ same provider unaffected.
111
+
112
+ ```ruby
113
+ llm = LLM.deepseek(key: ENV["KEY"])
114
+ agent = LLM::Agent.new(llm, tracer: LLM::Tracer::Logger.new(llm, io: $stdout))
115
+ ```
116
+
117
+ #### Why would I use it?
118
+
119
+ Agent-local tracers let each agent log differently.
120
+ One agent might log to stdout, another to a file, a third to
121
+ OpenTelemetry.
122
+
123
+ #### Notes
124
+
125
+ The tracer can also write to a file with the `path:` option to
126
+ [`LLM::Tracer::Logger.new`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html#initialize-instance_method).
127
+
128
+ ### PrettyLogger
129
+
130
+ #### Overview
131
+
132
+ [`LLM::Tracer::PrettyLogger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/PrettyLogger.html)
133
+ writes human-readable single-line logs to stderr. Unlike
134
+ [`LLM::Tracer::Logger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html)
135
+ (which emits structured JSON), the pretty logger is designed for
136
+ interactive development sessions where you want to see request and
137
+ tool-call activity at a glance.
138
+
139
+ #### How it works
140
+
141
+ Each request and tool call produces a single line on stderr with
142
+ the model, duration, and a summary of the activity. The logger
143
+ accepts an `io:` option to redirect output.
144
+
145
+ ##### Provider-wide
146
+
147
+ ```ruby
148
+ llm = LLM.deepseek(key: ENV["KEY"])
149
+ llm.tracer = LLM::Tracer::PrettyLogger.new(llm)
150
+ ```
151
+
152
+ ##### Agent-local
153
+
154
+ ```ruby
155
+ llm = LLM.deepseek(key: ENV["KEY"])
156
+ agent = LLM::Agent.new(llm, tracer: LLM::Tracer::PrettyLogger.new(llm))
157
+ ```
158
+
159
+ ##### Custom output
160
+
161
+ ```ruby
162
+ tracer = LLM::Tracer::PrettyLogger.new(llm, io: $stdout)
163
+ tracer = LLM::Tracer::PrettyLogger.new(llm, io: File.open("trace.log", "a"))
164
+ ```
165
+
166
+ #### Why would I use it?
167
+
168
+ The pretty logger is the best choice for development. The output is
169
+ compact enough to follow in real time while still showing the model
170
+ name, duration, and tool calls. Switch to
171
+ [`LLM::Tracer::Logger`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Logger.html)
172
+ when you need structured JSON for programmatic analysis, or to
173
+ [`LLM::Tracer::Telemetry`](https://r.uby.dev/api-docs/llm.rb/LLM/Tracer/Telemetry.html)
174
+ when you need OpenTelemetry exports.
175
+
176
+ #### Notes
177
+
178
+ The pretty logger writes to `$stderr` by default. Set `io:` to
179
+ redirect output. All three built-in tracers share the same interface,
180
+ so switching between them requires changing only the class name.
@@ -0,0 +1,67 @@
1
+
2
+ ## Transformer
3
+
4
+ ### Introduction
5
+
6
+ #### Overview
7
+
8
+ [`LLM::Transformer`](https://r.uby.dev/api-docs/llm.rb/LLM/Transformer.html)
9
+ is the superclass for message transformers. A transformer is bound
10
+ to a context and rewrites a single message before it is sent to the
11
+ provider. This lets you scrub sensitive data, inject context, or
12
+ otherwise modify outgoing messages without changing your prompt
13
+ code.
14
+
15
+ #### How it works
16
+
17
+ A transformer is a subclass of
18
+ [`LLM::Transformer`](https://r.uby.dev/api-docs/llm.rb/LLM/Transformer.html)
19
+ that implements
20
+ [`LLM::Transformer#call`](https://r.uby.dev/api-docs/llm.rb/LLM/Transformer.html#call-instance_method).
21
+ The method receives the message to transform and returns a
22
+ message.
23
+ You can mutate the message in place or return a new one; either way,
24
+ the returned message is what gets sent.
25
+
26
+ Configure the transformer on a context with the `transformer:` option,
27
+ passing a class rather than an instance. The runtime instantiates it
28
+ once per turn. Options passed through `transformer_options:` are
29
+ forwarded to `call` as keyword arguments. The transformer runs on
30
+ the most recent message in both chat completions and Responses API
31
+ turns, before the request reaches the provider:
32
+
33
+ ```ruby
34
+ class RedactEmails < LLM::Transformer
35
+ def call(message:)
36
+ content = message.content.to_s.gsub(/[\w.+-]+@[\w-]+\.[\w.]+/, "[EMAIL]")
37
+ LLM::Message.new(message.role, content, message.extra)
38
+ end
39
+ end
40
+
41
+ llm = LLM.deepseek(key: ENV["KEY"])
42
+ ctx = LLM::Context.new(
43
+ llm,
44
+ transformer: RedactEmails
45
+ )
46
+ ctx.talk "Contact support@example.com for help"
47
+ ```
48
+
49
+ #### Why would I use it?
50
+
51
+ Transformers give you a single hook point for all outgoing messages.
52
+ Common uses include redacting PII before it leaves your process,
53
+ injecting a timestamp or request ID, or normalizing content for a
54
+ particular provider. Because the transformer runs automatically on
55
+ every turn, you never need to remember to apply the transform in
56
+ your prompt code.
57
+
58
+ #### Notes
59
+
60
+ [`LLM::Transformer::Null`](https://r.uby.dev/api-docs/llm.rb/LLM/Transformer/Null.html)
61
+ is the default transformer; it returns the message unchanged. The
62
+ `transformer_options:` hash is passed to `call` on every turn.
63
+ Streams can observe transformation through the
64
+ [`LLM::Stream#on_transform`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_transform-instance_method)
65
+ and
66
+ [`LLM::Stream#on_transform_finish`](https://r.uby.dev/api-docs/llm.rb/LLM/Stream.html#on_transform_finish-instance_method)
67
+ callbacks, which receive the transformer instance.
@@ -0,0 +1,45 @@
1
+
2
+ ## Transports
3
+
4
+ ### Introduction
5
+
6
+ #### Overview
7
+
8
+ The transport option selects which HTTP library is used for network
9
+ communication.
10
+ [`LLM::Provider`](https://r.uby.dev/api-docs/llm.rb/LLM/Provider.html),
11
+ [`LLM::MCP`](https://r.uby.dev/api-docs/llm.rb/LLM/MCP.html), and
12
+ [`LLM::A2A`](https://r.uby.dev/api-docs/llm.rb/LLM/A2A.html)
13
+ all accept this option. Three backends are available out of the box:
14
+ `net/http` is always there, `net/http/persistent` pools connections,
15
+ and `curb` wraps libcurl. Each implements the same internal interface
16
+ so switching between them is a one-word change.
17
+
18
+ #### How it works
19
+
20
+ **Net/HTTP** (`:net_http`) is the
21
+ default and is always available. **Net/HTTP/Persistent**
22
+ (`:net_http_persistent`) maintains a connection pool so the cost
23
+ of tearing down and setting up connections is kept low. **Curb**
24
+ (`:curb`) provides bindings for libcurl.
25
+
26
+ ```ruby
27
+ llm = LLM.deepseek(key: "...", transport: :net_http)
28
+ mcp = LLM::MCP.http(url: "...", transport: :net_http_persistent)
29
+ a2a = LLM::A2A.rest(url: "...", transport: :curb)
30
+ ```
31
+
32
+ #### Why would I use it?
33
+
34
+ Different environments have different HTTP needs.
35
+
36
+ The default `net/http` transport is always available and works
37
+ everywhere. The persistent transport reduces connection overhead
38
+ when your agent makes many requests to the same provider in quick
39
+ succession. Curb gives you libcurl bindings for environments where
40
+ that is already configured or preferred.
41
+
42
+ #### Notes
43
+
44
+ The persistent transport is built on top of `net/http`. Curb
45
+ requires the `curb` gem.