claude-agent-sdk 1.1.0 → 1.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (46) hide show
  1. checksums.yaml +4 -4
  2. data/.yardopts +10 -0
  3. data/CHANGELOG.md +90 -0
  4. data/README.md +43 -31
  5. data/docs/cli-installer.md +26 -4
  6. data/docs/client.md +29 -11
  7. data/docs/configuration.md +164 -1
  8. data/docs/errors.md +32 -2
  9. data/docs/hooks-and-permissions.md +30 -10
  10. data/docs/mcp-servers.md +30 -9
  11. data/docs/observability.md +61 -10
  12. data/docs/options.md +232 -0
  13. data/docs/rails.md +263 -18
  14. data/docs/sessions.md +40 -12
  15. data/docs/subagents.md +1 -1
  16. data/docs/types.md +100 -11
  17. data/lib/claude_agent_sdk/cli_installer.rb +140 -19
  18. data/lib/claude_agent_sdk/command_builder.rb +84 -27
  19. data/lib/claude_agent_sdk/fiber_boundary.rb +45 -2
  20. data/lib/claude_agent_sdk/instrumentation/otel.rb +90 -28
  21. data/lib/claude_agent_sdk/query.rb +228 -77
  22. data/lib/claude_agent_sdk/railtie.rb +27 -2
  23. data/lib/claude_agent_sdk/sdk_mcp_server.rb +78 -26
  24. data/lib/claude_agent_sdk/session_mutations.rb +112 -92
  25. data/lib/claude_agent_sdk/session_resume.rb +356 -39
  26. data/lib/claude_agent_sdk/session_store.rb +31 -2
  27. data/lib/claude_agent_sdk/sessions.rb +720 -138
  28. data/lib/claude_agent_sdk/subprocess_cli_transport.rb +227 -29
  29. data/lib/claude_agent_sdk/testing/session_store_conformance.rb +18 -7
  30. data/lib/claude_agent_sdk/transcript_mirror_batcher.rb +45 -37
  31. data/lib/claude_agent_sdk/transport.rb +28 -12
  32. data/lib/claude_agent_sdk/types/attributes.rb +9 -0
  33. data/lib/claude_agent_sdk/types/base.rb +85 -15
  34. data/lib/claude_agent_sdk/types/hooks.rb +73 -0
  35. data/lib/claude_agent_sdk/types/mcp.rb +37 -1
  36. data/lib/claude_agent_sdk/types/option_values.rb +186 -4
  37. data/lib/claude_agent_sdk/types/options.rb +35 -5
  38. data/lib/claude_agent_sdk/types/permissions.rb +18 -9
  39. data/lib/claude_agent_sdk/version.rb +1 -1
  40. data/lib/claude_agent_sdk.rb +94 -46
  41. data/lib/generators/claude_agent_sdk/install/templates/claude_agent_sdk.rb.tt +6 -0
  42. data/sig/claude_agent_sdk/types/hooks.rbs +6 -3
  43. data/sig/claude_agent_sdk/types/option_values.rbs +23 -6
  44. data/sig/claude_agent_sdk/types/options.rbs +11 -7
  45. data/sig/claude_agent_sdk/types/permissions.rbs +4 -2
  46. metadata +6 -4
@@ -35,6 +35,8 @@ ClaudeAgentSDK.query(prompt: "Create a profile for a software engineer", options
35
35
  end
36
36
  ```
37
37
 
38
+ In the `output_format` Hash, `type` may be a String or a Symbol (`type: :json_schema`), and the keys Symbols or Strings.
39
+
38
40
  See [examples/structured_output_example.rb](https://github.com/ya-luotao/claude-agent-sdk-ruby/blob/main/examples/structured_output_example.rb).
39
41
 
40
42
  ## Thinking Configuration
@@ -96,6 +98,8 @@ options = ClaudeAgentSDK::ClaudeAgentOptions.new(
96
98
  )
97
99
  ```
98
100
 
101
+ In a Hash form, `type` may be a String or a Symbol (`type: :preset`).
102
+
99
103
  `snapshot` is sent on the control-protocol `initialize` request (never as a CLI flag), so it applies to both `query()` and `Client`. When omitted it acts as `true`, except in bare mode (`bare: true`), where it acts as `false`. A `SystemPromptFile` has no `snapshot`.
100
104
 
101
105
  Requires Claude Code CLI 2.1.257 or later. Before 2.1.265, a session with an `append` or custom prompt recorded it only when `snapshot` was `true`. See [Modifying system prompts](https://code.claude.com/docs/en/agent-sdk/modifying-system-prompts#change-the-prompt-of-an-existing-session) for details.
@@ -177,10 +181,18 @@ Available beta features are listed in the `SDK_BETAS` constant.
177
181
  # Array of tool names
178
182
  options = ClaudeAgentSDK::ClaudeAgentOptions.new(tools: ['Read', 'Edit', 'Bash'])
179
183
 
184
+ # The same names as one String, in the comma-separated form of the CLI's --tools flag
185
+ options = ClaudeAgentSDK::ClaudeAgentOptions.new(tools: 'Read,Edit,Bash')
186
+
180
187
  # Preset
181
188
  options = ClaudeAgentSDK::ClaudeAgentOptions.new(tools: ClaudeAgentSDK::ToolsPreset.new(preset: 'claude_code'))
189
+
190
+ # The preset as a Hash
191
+ options = ClaudeAgentSDK::ClaudeAgentOptions.new(tools: { type: 'preset', preset: 'claude_code' })
182
192
  ```
183
193
 
194
+ A String is passed to the CLI as written. In the Hash form, `type` may be a String or a Symbol (`type: :preset`).
195
+
184
196
  ## Skills
185
197
 
186
198
  `skills` is the single place to enable skills for the main session — it auto-allows the `Skill` tool and defaults `setting_sources` to `['user', 'project']` (when unset) so skill files are discovered:
@@ -195,9 +207,15 @@ options = ClaudeAgentSDK::ClaudeAgentOptions.new(skills: %w[pdf docx])
195
207
 
196
208
  Semantics: `nil` (default) leaves CLI defaults untouched; `[]` hides every skill from the listing; an Array adds `Skill(name)` allow-rules per entry (use `plugin:skill` for plugin-qualified names). An explicitly set `setting_sources` (including `[]`) is never overridden. This is a context filter, not a sandbox — skill files remain readable on disk.
197
209
 
210
+ If you also give `tools` as a list of names, put `'Skill'` in it. `skills` adds the `Skill` tool to the *allowed* tools, while `tools` decides which tools the session has at all: with `tools: ['Read'], skills: 'all'` there is no `Skill` tool, and no skill can run. Leaving `tools` unset keeps it.
211
+
212
+ ```ruby
213
+ options = ClaudeAgentSDK::ClaudeAgentOptions.new(tools: %w[Read Skill], skills: %w[pdf docx])
214
+ ```
215
+
198
216
  ## Sandbox Settings
199
217
 
200
- Configure [sandbox-runtime](https://github.com/anthropic-experimental/sandbox-runtime) restrictions (network policy, filesystem access) via the CLI's `--sandbox` flag. The CLI handles OS-level process isolation using `srt`.
218
+ Configure [sandbox-runtime](https://github.com/anthropic-experimental/sandbox-runtime) restrictions (network policy, filesystem access) with the `sandbox` option. The SDK sends it to the CLI as the `sandbox` key of the `--settings` argument, merged into your `settings:` when you pass both; there is no separate sandbox flag. The CLI handles OS-level process isolation using `srt`.
201
219
 
202
220
  ```ruby
203
221
  sandbox = ClaudeAgentSDK::SandboxSettings.new(
@@ -212,8 +230,40 @@ options = ClaudeAgentSDK::ClaudeAgentOptions.new(
212
230
  )
213
231
  ```
214
232
 
233
+ `sandbox` also takes a Hash, and so do `network` and `filesystem`, inside that Hash or inside a `SandboxSettings`. A Hash may spell the fields of `SandboxSettings`, `SandboxNetworkConfig` and `SandboxFilesystemConfig` as the classes do (`denied_domains`) or as the CLI does (`deniedDomains`), with Symbol or String keys:
234
+
235
+ ```ruby
236
+ options = ClaudeAgentSDK::ClaudeAgentOptions.new(
237
+ sandbox: {
238
+ enabled: true,
239
+ network: { denied_domains: ['evil.example'] },
240
+ filesystem: { deny_read: ['~/.ssh'] }
241
+ }
242
+ )
243
+ ```
244
+
245
+ - Any other key is sent as written. That is how to pass a sandbox setting the classes have no attribute for (`allowAppleEvents`, `strictAllowlist` inside `network`): spell it as the CLI does.
246
+ - A snake_case key is sent under the CLI's name only when its value has the shape the CLI accepts for that key: `true` or `false` for a switch, an Array of Strings for a list, a port number (an Integer from 0 to 65535) for a proxy port, a Hash of String Arrays for `ignore_violations`. With a value of any other shape (`denied_domains: 'evil.example'`, a String where an Array belongs) the key is sent as written, and the CLI ignores a key it does not know.
247
+ - A field holding `nil` is left out, in either spelling, as a `nil` attribute of the typed classes is.
248
+
249
+ The second rule exists because of how the CLI validates these settings (observed with CLI 2.1.287). When one sandbox value fails its settings schema, the CLI discards the **whole** `--settings` value: the sandbox, and everything you passed in `settings:` next to it, `permissions` rules included. The CLI reports the failure in the `errors` of its `get_settings` control response, and the SDK does not surface that: `connect` succeeds, nothing appears on stderr, and the session runs unsandboxed. That applies to a value written under the CLI's own name (`excludedCommands: 'docker'`) and to the typed classes, which send each value as you gave it: `SandboxSettings.new(excluded_commands: 'docker')` leaves the session without a sandbox.
250
+
251
+ `enabled: true` asks for a sandbox. It does not make one a requirement. When the sandbox cannot start on the host (missing dependencies, an unsupported platform), the CLI carries on without it. Its settings schema describes the setting that decides this, `failIfUnavailable`, as follows (CLI 2.1.287):
252
+
253
+ > Exit with an error at startup if sandbox.enabled is true but the sandbox cannot start (missing dependencies or unsupported platform). When false (default), a warning is shown and commands run unsandboxed.
254
+
255
+ The SDK sends your sandbox settings as you wrote them and does not add this one. If commands must never run unsandboxed, set it yourself:
256
+
257
+ ```ruby
258
+ sandbox = ClaudeAgentSDK::SandboxSettings.new(enabled: true, fail_if_unavailable: true)
259
+ ```
260
+
261
+ Without it, the sign is the CLI's warning on stderr ("Sandbox disabled: ... Commands will run WITHOUT sandboxing. Network and filesystem restrictions will NOT be enforced."), and stderr reaches your code only through the `stderr` (or `debug_stderr`) option. This is the CLI's own description of its behavior: the fallback has not been reproduced in this SDK's testing, where the sandbox was always available.
262
+
215
263
  See [examples/sandbox_example.rb](https://github.com/ya-luotao/claude-agent-sdk-ruby/blob/main/examples/sandbox_example.rb).
216
264
 
265
+ When the `sandbox:` option enabled the sandbox and the CLI reports `Sandbox disabled: …` on its stderr, the SDK repeats that line as a Ruby warning prefixed `[claude-agent-sdk]`, once per session, whether or not you set `stderr:`.
266
+
217
267
  ## Bare Mode
218
268
 
219
269
  Bare mode (`--bare`) is a minimal startup mode that skips hooks, LSP, plugin sync, attribution, auto-memory, background prefetches, keychain reads, and CLAUDE.md auto-discovery. It sets `CLAUDE_CODE_SIMPLE=1` internally. Useful for scripted/programmatic usage where you want fast startup and full control over what's loaded.
@@ -287,6 +337,119 @@ and `Client#query`.
287
337
 
288
338
  Matches the Python SDK's `verbatim_prompts`.
289
339
 
340
+ ## Session Isolation
341
+
342
+ Claude Code keeps an **auto-memory** for each project: Markdown notes and an
343
+ index file, `<config dir>/projects/<project key>/memory/MEMORY.md`. It is a
344
+ CLI feature, it is on by default, and SDK sessions take part in it. If one
345
+ process runs sessions for more than one user or tenant, three things follow:
346
+
347
+ - **Every session reads the index.** The CLI puts `MEMORY.md` into the context
348
+ of every session of that project, under the same heading as `CLAUDE.md` (in
349
+ CLI 2.1.287: "IMPORTANT: These instructions OVERRIDE any default behavior
350
+ and you MUST follow them exactly as written"). That holds with the SDK's
351
+ default (empty) system prompt and with `tools: []`. `setting_sources: []`
352
+ does not change it either: that option selects the user, project and local
353
+ sources (their settings files and their `CLAUDE.md` files), and the
354
+ auto-memory is not one of them.
355
+ - **A session on the `claude_code` preset also writes it.** The preset system
356
+ prompt includes instructions for keeping memories, so a message such as
357
+ "remember that ..." makes the model save a note and update the index with
358
+ its file tools. The CLI allows those writes by itself: in testing (CLI
359
+ 2.1.286) they succeeded in the default permission mode with no
360
+ `can_use_tool` callback, no hook and no allow rule. Do not count on your
361
+ permission setup to stop them. With the SDK's default system prompt the
362
+ sessions tested only read the index: asked to remember something, they
363
+ wrote nothing (an observation, not a guarantee).
364
+ - **The memory belongs to the project, not to the session.** The project key
365
+ is the root of the git repository that contains the working directory, so
366
+ every subdirectory and every git worktree of one repository shares one
367
+ memory directory. Outside a repository the key is the working directory
368
+ itself.
369
+
370
+ So on a server that runs every user's session from one checkout and one config
371
+ directory with the `claude_code` preset, what one user asks the agent to
372
+ remember can be saved without a permission check and then reach every later
373
+ session as an instruction. With the default system prompt the exposure is the
374
+ read side: whatever memory already exists for that project and config
375
+ directory, a developer's own for instance, is in every session's context.
376
+
377
+ ### Turning auto-memory off
378
+
379
+ Servers and multi-tenant hosts should switch it off. Both forms work with
380
+ `query()` and `Client`:
381
+
382
+ ```ruby
383
+ # An environment variable for the CLI process
384
+ options = ClaudeAgentSDK::ClaudeAgentOptions.new(
385
+ env: { 'CLAUDE_CODE_DISABLE_AUTO_MEMORY' => '1' }
386
+ )
387
+
388
+ # Or the CLI setting
389
+ options = ClaudeAgentSDK::ClaudeAgentOptions.new(
390
+ settings: { autoMemoryEnabled: false }
391
+ )
392
+ ```
393
+
394
+ To apply it to every session, set it as a default (a per-call `env` Hash is
395
+ merged into the configured one):
396
+
397
+ ```ruby
398
+ ClaudeAgentSDK.configure do |config|
399
+ config.default_options = { env: { 'CLAUDE_CODE_DISABLE_AUTO_MEMORY' => '1' } }
400
+ end
401
+ ```
402
+
403
+ - The variable must be `'1'`. `'0'` and `'false'` do not mean "the default":
404
+ they force auto-memory **on** and override `autoMemoryEnabled: false`.
405
+ - `bare: true` turns auto-memory off as well, but bare mode never reads an
406
+ OAuth login or the keychain: it authenticates with `ANTHROPIC_API_KEY` (or
407
+ an `apiKeyHelper` setting) only. See [Bare Mode](#bare-mode).
408
+ - Both switches reach the CLI that `query()` and `Client` start themselves,
409
+ through the default `SubprocessCLITransport`: it puts `env` into the CLI's
410
+ environment and `settings` on its command line (`--settings`).
411
+ - A [custom transport](client.md#custom-transport) starts the CLI its own way,
412
+ so neither switch reaches that CLI unless the transport passes it through:
413
+ `env` into the environment it gives the CLI, `settings` onto the command
414
+ line (a transport that builds its command line with `CommandBuilder` gets
415
+ `--settings` from it; one that does not, has to add it). Until then the
416
+ session is not isolated.
417
+
418
+ ### What does not isolate sessions
419
+
420
+ - **A different `cwd`** separates the memory only when the two directories
421
+ are not in the same git repository.
422
+ - **A different `CLAUDE_CONFIG_DIR`** separates it only when the two config
423
+ directories do not share `projects/` (a `projects/` that is a symlink to
424
+ another config directory's is shared). A new config directory also has no
425
+ login, so give the CLI `ANTHROPIC_API_KEY` or `CLAUDE_CODE_OAUTH_TOKEN`
426
+ (in the process environment or in `env`).
427
+ - **`setting_sources: []`** keeps the user, project and local settings and
428
+ `CLAUDE.md` files out of a session. The auto-memory index still loads. So
429
+ do the MCP servers connected to the claude.ai account the CLI is logged in
430
+ with: `strict_mcp_config: true` is what limits a session to the servers in
431
+ `mcp_servers` (see the [options reference](options.md#strict_mcp_config)).
432
+
433
+ "One working directory per tenant" is therefore not enough on its own. Use the
434
+ switch.
435
+
436
+ ### Checking what a session loaded
437
+
438
+ `Client#context_usage` reports the files in a session's context without a
439
+ model call. The auto-memory index is the entry whose `:type` is `"AutoMem"`:
440
+
441
+ ```ruby
442
+ ClaudeAgentSDK::Client.open(options: options) do |client|
443
+ client.context_usage.fetch(:memoryFiles, []).each do |file|
444
+ puts "#{file[:type]} #{file[:path]} (#{file[:tokens]} tokens)"
445
+ end
446
+ end
447
+ # AutoMem /home/app/.claude/projects/-srv-app/memory/MEMORY.md (27 tokens)
448
+ ```
449
+
450
+ With auto-memory off the list has no `AutoMem` entry. `query()` has no
451
+ equivalent; run the check through a `Client` with the same options.
452
+
290
453
  ## Forwarding Subagent Text
291
454
 
292
455
  By default only `tool_use` / `tool_result` blocks from subagents (spawned via
data/docs/errors.md CHANGED
@@ -87,6 +87,30 @@ initial handshake that was still in flight when the CLI exited. Match on
87
87
  `Resume rejected by --resume-drops-turn:` in the message and treat it as
88
88
  deterministic: clear the fork target and resume plainly rather than retrying.
89
89
 
90
+ ## Errors from Client Control Methods
91
+
92
+ `Client#interrupt`, `#set_model`, `#set_permission_mode`, `#rewind_files`, `#reconnect_mcp_server`, `#toggle_mcp_server`, `#stop_task` and the other control methods (the Ruby-style `model=` and `permission_mode=` included) send a request to the CLI and wait for its answer. They can fail in three ways:
93
+
94
+ | What happened | Raised |
95
+ |---------------|--------|
96
+ | The client is not connected, or the connection ended while the request was waiting | `CLIConnectionError` |
97
+ | The CLI did not answer in time | `ControlRequestTimeoutError` |
98
+ | The CLI answered with an error: an MCP server name it does not know, a rewind without `enable_file_checkpointing`, a permission mode it does not accept | a plain `StandardError` whose message is the CLI's text |
99
+
100
+ The third is **not** a `ClaudeSDKError`, so `rescue ClaudeAgentSDK::ClaudeSDKError` does not catch it:
101
+
102
+ ```ruby
103
+ begin
104
+ client.reconnect_mcp_server('my-server')
105
+ rescue ClaudeAgentSDK::ClaudeSDKError => e
106
+ puts "Connection lost or timed out: #{e.message}"
107
+ rescue StandardError => e
108
+ puts "The CLI refused: #{e.message}" # e.g. "Server not found: my-server"
109
+ end
110
+ ```
111
+
112
+ Rescue `StandardError` for it, and do not test for the exact class.
113
+
90
114
  ## Configuring Timeout
91
115
 
92
116
  The control request timeout defaults to **1200 seconds** (20 minutes) to accommodate long-running agent sessions. Override it via environment variable:
@@ -98,7 +122,7 @@ export CLAUDE_AGENT_SDK_CONTROL_REQUEST_TIMEOUT_SECONDS=300 # 5 minutes
98
122
  ## Error Type Reference
99
123
 
100
124
  ```ruby
101
- # Base exception class for all SDK errors
125
+ # Base class of the SDK's error classes (all the ones below)
102
126
  class ClaudeSDKError < StandardError; end
103
127
 
104
128
  # Raised when connection to Claude Code fails
@@ -113,6 +137,11 @@ class CLINotFoundError < CLIConnectionError
113
137
  # @param cli_path [String, nil] Optional path to the CLI that was not found
114
138
  end
115
139
 
140
+ # Raised by CLIInstaller.install / .install_pinned (and the
141
+ # claude_agent_sdk:install_cli rake task) when the CLI cannot be installed.
142
+ # Never raised by query, ask or Client
143
+ class CLIInstallError < ClaudeSDKError; end
144
+
116
145
  # Raised by the local-disk session APIs when CLAUDE_CONFIG_DIR is unset and
117
146
  # no usable home directory exists for the default ~/.claude
118
147
  class ConfigDirError < ClaudeSDKError; end
@@ -155,10 +184,11 @@ end
155
184
 
156
185
  | Error | Description |
157
186
  |-------|-------------|
158
- | `ClaudeSDKError` | Base error for all SDK errors |
187
+ | `ClaudeSDKError` | Base class of every error class in this table. One failure is not a `ClaudeSDKError`: a control request the CLI rejects, see [Errors from Client Control Methods](#errors-from-client-control-methods) |
159
188
  | `CLIConnectionError` | Connection issues — including every write after a stdin write was cancelled mid-frame (the connection is unusable from then on — reconnect), and `ClaudeAgentSDK.ask` when the stream ends without a `ResultMessage` |
160
189
  | `ControlRequestTimeoutError` | Control protocol timeout (configurable via env var) |
161
190
  | `CLINotFoundError` | Claude Code not installed |
191
+ | `CLIInstallError` | `CLIInstaller.install` / `install_pinned` (or the `claude_agent_sdk:install_cli` rake task) could not install the CLI: unsupported platform, invalid version, HTTP error, a response or download over its size limit, checksum mismatch, or a filesystem error. A wrapped failure keeps the original exception in `#cause`. This is the error a deploy or `bin/setup` script rescues; a failed install leaves a previously installed binary in place. See [CLI installer](cli-installer.md) |
162
192
  | `ConfigDirError` | A local-disk session API (`list_sessions`, `get_session_*`, `rename_session`, ...) could not locate the Claude config directory: `CLAUDE_CONFIG_DIR` is unset and there is no usable home directory (`HOME` unset with no passwd entry, as under `docker --user` in a minimal image, or an empty/relative `HOME`). Set `CLAUDE_CONFIG_DIR` |
163
193
  | `SessionStoreError` | Resuming from `session_store:` failed: a store call (`#load`, `#list_sessions`, `#list_subkeys`) raised or exceeded `load_timeout_ms` while the SDK materialized the transcript, before the CLI started. The message names the call; `#cause` is the adapter's exception. Before 1.0 this was a bare `RuntimeError` — see [Sessions](sessions.md#mirroring-to-a-sessionstore) |
164
194
  | `ProcessError` | Process failed (includes `exit_code` and `stderr`) — also raised when the CLI is still running 5s after closing stdout and the SDK had to terminate it |
@@ -69,16 +69,34 @@ Async do
69
69
  }
70
70
  )
71
71
 
72
- client = ClaudeAgentSDK::Client.new(options: options)
73
- client.connect
74
- client.query("Run the bash command: ./foo.sh --help")
75
- client.receive_response { |msg| puts msg }
76
- client.disconnect
72
+ ClaudeAgentSDK::Client.open(options: options) do |client|
73
+ client.query("Run the bash command: ./foo.sh --help")
74
+ client.receive_response { |msg| puts msg }
75
+ end
77
76
  end.wait
78
77
  ```
79
78
 
80
79
  See [examples/hooks_example.rb](https://github.com/ya-luotao/claude-agent-sdk-ruby/blob/main/examples/hooks_example.rb), [examples/advanced_hooks_example.rb](https://github.com/ya-luotao/claude-agent-sdk-ruby/blob/main/examples/advanced_hooks_example.rb), and [examples/lifecycle_hooks_example.rb](https://github.com/ya-luotao/claude-agent-sdk-ruby/blob/main/examples/lifecycle_hooks_example.rb).
81
80
 
81
+ ### Hook output
82
+
83
+ A callback returns a typed output (`AsyncHookJSONOutput`, or `SyncHookJSONOutput`, which can hold a `*HookSpecificOutput`), a Hash, or `nil` (the same as `{}`). A Hash stands for the typed output with the same fields. Its keys may be Symbols or Strings, and each field may be spelled as the typed class's attribute (`permission_decision`) or as the CLI reads it (`permissionDecision`), at the top level and inside `hook_specific_output`. The deny in the example above can also be written as:
84
+
85
+ ```ruby
86
+ {
87
+ hook_specific_output: {
88
+ hook_event_name: 'PreToolUse', # a typed output sets this itself; a Hash has to carry it
89
+ permission_decision: 'deny',
90
+ permission_decision_reason: "Command contains forbidden pattern: #{pattern}"
91
+ }
92
+ }
93
+ ```
94
+
95
+ - `continue_` and `async_` stand for `continue` and `async`.
96
+ - A key the typed classes do not have is sent as written, so a CLI field the SDK does not model has to be spelled the way the CLI spells it.
97
+ - The keys inside `updated_input`, `updated_tool_output`, `updated_mcp_tool_output` and a `PermissionRequest` `decision` are not renamed: those are the tool's own payloads (for `decision`, the CLI's), so spell them as the tool or the CLI does.
98
+ - If one Hash spells the same field both ways, the CLI's spelling is the one sent (`permissionDecision` over `permission_decision`, `continue` over `continue_`).
99
+
82
100
  ### Hook cancellation
83
101
 
84
102
  Dispatched hooks receive `HookContext#request_id` and `HookContext#signal`, using
@@ -123,11 +141,10 @@ Async do
123
141
  can_use_tool: permission_callback
124
142
  )
125
143
 
126
- client = ClaudeAgentSDK::Client.new(options: options)
127
- client.connect
128
- client.query("Create a file called test.txt with content 'Hello'")
129
- client.receive_response { |msg| puts msg }
130
- client.disconnect
144
+ ClaudeAgentSDK::Client.open(options: options) do |client|
145
+ client.query("Create a file called test.txt with content 'Hello'")
146
+ client.receive_response { |msg| puts msg }
147
+ end
131
148
  end.wait
132
149
  ```
133
150
 
@@ -202,6 +219,9 @@ An exception raised inside a hook or a `can_use_tool` callback fails that
202
219
  control request: the CLI receives an error response carrying the exception
203
220
  message, and the session carries on with later requests. The request's
204
221
  cancellation signal is invalidated, as for any other callback failure.
222
+ `NotImplementedError`, `LoadError`, `SystemStackError` and `SecurityError` are
223
+ answered the same way, although they are not `StandardError`s. Bytes of the
224
+ message that are not valid UTF-8 are replaced with U+FFFD.
205
225
 
206
226
  `exit`, `Interrupt` and other signal exceptions are never swallowed. If one
207
227
  is raised while a callback runs (by the callback itself, or a real Ctrl-C /
data/docs/mcp-servers.md CHANGED
@@ -8,7 +8,6 @@ A **custom tool** is a Ruby proc/lambda that you can offer to Claude, for Claude
8
8
 
9
9
  ```ruby
10
10
  require 'claude_agent_sdk'
11
- require 'async'
12
11
 
13
12
  greet_tool = ClaudeAgentSDK.create_tool(
14
13
  'greet', 'Greet a user', { name: :string },
@@ -28,13 +27,10 @@ options = ClaudeAgentSDK::ClaudeAgentOptions.new(
28
27
  allowed_tools: ['mcp__tools__greet']
29
28
  )
30
29
 
31
- Async do
32
- client = ClaudeAgentSDK::Client.new(options: options)
33
- client.connect
30
+ ClaudeAgentSDK::Client.open(options: options) do |client|
34
31
  client.query("Greet Alice")
35
32
  client.receive_response { |msg| puts msg }
36
- client.disconnect
37
- end.wait
33
+ end
38
34
  ```
39
35
 
40
36
  ## Handler Return Values
@@ -56,7 +52,28 @@ ClaudeAgentSDK.create_tool('lookup_order', 'Look up an order', { id: :string })
56
52
  end
57
53
  ```
58
54
 
59
- Any other return value (`nil`, an Integer, an Array, ...) is reported to Claude as an `isError: true` result saying the tool must return a hash with a `:content` key.
55
+ Any other return value (`nil`, an Integer, an Array, ...) is reported to Claude as an `isError: true` result saying the tool must return a hash with a `:content` key. So is a Hash whose `:content` is not an Array — `{ content: 'text' }`, or a single block Hash — with a message saying `:content` must be an Array of content blocks; return the String itself, or wrap the block in an Array.
56
+
57
+ ## Shorthand Input Schemas
58
+
59
+ `{ name: :string }` is shorthand for a JSON Schema object in which every listed parameter is required. Each value names a type, as a Ruby class or as a Symbol:
60
+
61
+ | Shorthand | JSON Schema type | The handler receives |
62
+ | --- | --- | --- |
63
+ | `String`, `:string` | `string` | a String |
64
+ | `Integer`, `:integer` | `integer` | an Integer |
65
+ | `Float`, `:float`, `:number` | `number` | an Integer or a Float |
66
+ | `TrueClass`, `FalseClass`, `:boolean` | `boolean` | `true` or `false` |
67
+ | `Array`, `:array` | `array` | an Array, with elements of any type |
68
+ | `Hash`, `:object` | `object` | a Hash with Symbol keys |
69
+
70
+ ```ruby
71
+ ClaudeAgentSDK.create_tool('tag_order', 'Tag an order', { order_id: Integer, tags: Array }) do |args|
72
+ "Tagged order #{args[:order_id]} with #{args[:tags].join(', ')}"
73
+ end
74
+ ```
75
+
76
+ Arguments are validated against these types before the handler runs. For an optional parameter, the type of an Array's elements, an enum or a per-parameter description, pass a full JSON Schema instead (next section).
60
77
 
61
78
  ## Pre-built JSON Schemas
62
79
 
@@ -125,12 +142,14 @@ options = ClaudeAgentSDK::ClaudeAgentOptions.new(
125
142
 
126
143
  An exception raised inside a handler is returned to the model as an
127
144
  `isError: true` result carrying the exception message, so it can self-correct.
145
+ That holds for `NotImplementedError`, `LoadError`, `SystemStackError` and
146
+ `SecurityError` as well, although they are not `StandardError`s.
128
147
  `exit`, `Interrupt` and other signal exceptions are never swallowed: the CLI
129
148
  first gets an `isError` result naming the exception class
130
149
  (`"SystemExit: exit"`), so it is not left waiting on the tool call, and then
131
150
  the exception propagates as Ruby normally would (`exit` ends the process,
132
151
  Ctrl-C interrupts it). Called directly, without a session,
133
- `SdkMcpServer#call_tool` simply lets such exceptions propagate. Cancellation of the tool call itself still propagates.
152
+ `SdkMcpServer#call_tool` simply lets `exit` and signal exceptions propagate. Cancellation of the tool call itself still propagates.
134
153
 
135
154
  ## Mixed Server Support
136
155
 
@@ -196,7 +215,9 @@ server = ClaudeAgentSDK.create_sdk_mcp_server(
196
215
  ```
197
216
 
198
217
  An exception raised inside a resource reader or prompt generator is answered
199
- with a JSON-RPC internal error (`-32603`) carrying the exception message. For
218
+ with a JSON-RPC internal error (`-32603`) carrying the exception message;
219
+ `NotImplementedError`, `LoadError`, `SystemStackError` and `SecurityError`
220
+ are covered here too. For
200
221
  `exit`, `Interrupt` and other signal exceptions the CLI gets that error first,
201
222
  naming the exception class, and then the exception propagates as Ruby normally
202
223
  would.
@@ -1,6 +1,6 @@
1
1
  # Observability (OpenTelemetry / Langfuse)
2
2
 
3
- The SDK includes a built-in **observer interface** and an **OpenTelemetry observer** for tracing agent sessions. Traces are emitted using standard `gen_ai.*` semantic conventions, compatible with Langfuse, Jaeger, Datadog, and any OTel backend.
3
+ The SDK includes a built-in **observer interface** and an **OpenTelemetry observer** for tracing agent sessions. Span attributes follow the Langfuse and OpenInference conventions, plus a subset of the OTel `gen_ai.*` attributes. Any OTel backend (Jaeger, Datadog, ...) can store and display the spans; one that interprets attributes by the OTel GenAI semantic conventions reads only part of them, as [Span Attributes](#span-attributes) explains.
4
4
 
5
5
  ## Distributed Trace Context (W3C)
6
6
 
@@ -18,13 +18,15 @@ In `Client` mode, call `disconnect` (ideally in an `ensure` block) so `on_close`
18
18
 
19
19
  ```
20
20
  claude_agent.session (root span — one per query/session)
21
- ├── claude_agent.generation (per AssistantMessage, with model + token usage)
21
+ ├── claude_agent.generation (per AssistantMessage, with model; token usage once per API response)
22
22
  ├── claude_agent.tool.Bash (per tool call, open on ToolUseBlock, close on ToolResultBlock)
23
23
  ├── claude_agent.tool.Read
24
24
  ├── claude_agent.generation
25
25
  └── ...
26
26
  ```
27
27
 
28
+ A session that fails before its first `InitMessage` (the CLI cannot be found or started, or the `initialize` handshake fails or times out) still leaves a span. `OTelObserver` emits a `claude_agent.session` span that records the exception and has error status, and ends it at once, so it is exported even when no `on_close` follows: a `Client#connect` that fails before the handshake never calls it. The span carries the observer's default attributes and, if a prompt had already been sent, `input.value`. It has no model, `session.id` or `claude_code.*` attributes, because the CLI never reported them. Once a trace has started, an error is recorded on its session span instead, and an error that arrives after the trace ended adds no span: the usual one is the `ResultError` raised when the CLI exits non-zero after an error result, which that trace already reports.
29
+
28
30
  ## Setup with Langfuse
29
31
 
30
32
  **1. Install the OTel gems** (not bundled with the SDK — you choose your exporter):
@@ -38,6 +40,7 @@ Or add to your Gemfile:
38
40
  ```ruby
39
41
  gem 'opentelemetry-sdk', '~> 1.4'
40
42
  gem 'opentelemetry-exporter-otlp', '~> 0.28'
43
+ gem 'base64' # the snippet below requires it; under Bundler, Ruby 3.4+ needs it listed
41
44
  ```
42
45
 
43
46
  **2. Configure the OTel SDK** to export to your Langfuse instance:
@@ -110,19 +113,67 @@ See [docs/rails.md](rails.md) for the Rails-specific pattern.
110
113
 
111
114
  ## Span Attributes
112
115
 
113
- The OTel observer sets attributes using both `gen_ai.*` (OTel GenAI) and OpenInference conventions for maximum backend compatibility:
114
-
115
- | Span | Type | Key Attributes |
116
- |------|------|----------------|
117
- | `claude_agent.session` | `agent` | `gen_ai.system`, `gen_ai.request.model`, `session.id`, `input.value`, `output.value`, `gen_ai.usage.cost`, `llm.cost.total`, `gen_ai.usage.cache_creation_input_tokens`, `gen_ai.usage.cache_read_input_tokens`, `llm.token_count.prompt_details.cache_read`, `llm.token_count.prompt_details.cache_write` |
118
- | `claude_agent.generation` | `generation` | `gen_ai.response.model`, `gen_ai.usage.input_tokens`, `gen_ai.usage.output_tokens`, `gen_ai.usage.cache_creation_input_tokens`, `gen_ai.usage.cache_read_input_tokens`, `output.value` |
119
- | `claude_agent.tool.*` | `tool` | `tool.name`, `input.value`, `output.value` |
116
+ Attribute names follow the Langfuse and OpenInference conventions, plus a subset of the OTel `gen_ai.*` attributes. The tables below list every attribute the observer sets. An attribute whose value the CLI did not report is left out, and `input.value`, `output.value` and `gen_ai.completion` are cut at 4,096 characters.
117
+
118
+ **`claude_agent.session`**
119
+
120
+ | Attribute | Value |
121
+ |-----------|-------|
122
+ | `gen_ai.system` | `anthropic` |
123
+ | `gen_ai.request.model`, `llm.model_name` | The model named by the `InitMessage` |
124
+ | `session.id` | The CLI session ID |
125
+ | `openinference.span.kind` | `AGENT` |
126
+ | `langfuse.observation.type` | `agent` |
127
+ | `input.mime_type`, `output.mime_type` | `text/plain` |
128
+ | `claude_code.version`, `claude_code.cwd`, `claude_code.permission_mode` | From the `InitMessage` |
129
+ | Your default attributes | Whatever you passed to `OTelObserver.new`. They are set on this span only |
130
+ | `input.value` | The prompt of the trace |
131
+ | `output.value` | `ResultMessage#result`, or the last assistant text when the result has none |
132
+ | `gen_ai.usage.input_tokens`, `gen_ai.usage.output_tokens`, `gen_ai.usage.cache_creation_input_tokens`, `gen_ai.usage.cache_read_input_tokens` | The four counts of `ResultMessage#usage` as the API reports them, so `input_tokens` leaves cache tokens out |
133
+ | `llm.token_count.prompt` | Input, cache-creation and cache-read tokens added up |
134
+ | `llm.token_count.completion` | Output tokens |
135
+ | `llm.token_count.total` | `llm.token_count.prompt` plus `llm.token_count.completion` |
136
+ | `llm.token_count.prompt_details.cache_read`, `llm.token_count.prompt_details.cache_write` | Cache-read and cache-creation tokens |
137
+ | `gen_ai.usage.cost`, `llm.cost.total` | The cost increase since the previous result (see below) |
138
+ | `claude_agent.duration_ms`, `claude_agent.duration_api_ms`, `claude_agent.num_turns`, `claude_agent.stop_reason` | From the `ResultMessage` |
139
+
140
+ The span status is error when the `ResultMessage` has `is_error` (the description is its stop reason) or when `on_error` recorded an exception, which also adds an `exception` event. Three more events are recorded on this span: `api_retry` (`attempt`, `max_retries`, `retry_delay_ms`, `error_status`, `error`), `rate_limit` (`status`, `rate_limit_type`) and `tool_progress` (`tool_name`, `tool_use_id`, `elapsed_time_seconds`).
141
+
142
+ **`claude_agent.generation`**
143
+
144
+ | Attribute | Value |
145
+ |-----------|-------|
146
+ | `openinference.span.kind` | `LLM` |
147
+ | `langfuse.observation.type` | `generation` |
148
+ | `gen_ai.response.model`, `llm.model_name` | The model named by the `AssistantMessage` |
149
+ | `gen_ai.completion`, `output.value` | The text blocks of the message joined by newlines; an empty string when it has none (a thinking or tool-call message) |
150
+ | `gen_ai.usage.input_tokens`, `gen_ai.usage.output_tokens`, `gen_ai.usage.cache_creation_input_tokens`, `gen_ai.usage.cache_read_input_tokens` | The usage of the API response, on the first span of each `message_id` only (see below) |
151
+
152
+ **`claude_agent.tool.<tool name>`**
153
+
154
+ | Attribute | Value |
155
+ |-----------|-------|
156
+ | `openinference.span.kind` | `TOOL` |
157
+ | `langfuse.observation.type` | `tool` |
158
+ | `tool.name` | The name of the tool |
159
+ | `input.value`, `input.mime_type` | The tool input as JSON, and `application/json` |
160
+ | `output.value`, `output.mime_type` | The tool result: a String as it is (`text/plain`), structured content as JSON (`application/json`). Both are left out when the result has no content |
161
+
162
+ The span status is error when the tool result has `is_error`.
163
+
164
+ **Token usage on generation spans.** The CLI sends one `AssistantMessage` per content block, so an API response with a thinking block and a tool call produces two `claude_agent.generation` spans, and both messages repeat the response's `message_id` and `usage`. The observer sets the four `gen_ai.usage.*` token attributes on the first generation span of each `message_id` and on none of that response's other spans, so a sum over generation spans counts every response once. A message without a `message_id` keeps its usage. That usage is the snapshot taken when the response started: the input, cache-read and cache-creation counts are final, but `gen_ai.usage.output_tokens` holds only the few tokens generated by then, and the response's final output count never reaches the message stream. For authoritative totals, output tokens included, read the session span, which takes them from `ResultMessage#usage`.
120
165
 
121
166
  `gen_ai.usage.cost` and `llm.cost.total` record the **increase** since the last observed `ResultMessage.total_cost_usd` in the connected session, rather than repeating that cumulative total on every turn's span. The baseline survives per-turn span resets and resets on close, a changed session ID (such as `/clear`), or a decreased counter. The original `ResultMessage` is unchanged.
122
167
 
123
168
  Missing costs omit both attributes without discarding the last known total; the next reported increment can therefore include an unreported or interrupted turn. The first observation uses the CLI's reported total. If a newer CLI restores historical spend on resume, that first span also includes it: a fresh observer cannot separate spend it never observed. These are CLI cost estimates, not billing records.
124
169
 
125
- Events (`api_retry`, `rate_limit`, `tool_progress`) are recorded on the root span.
170
+ **How this compares with the OTel GenAI semantic conventions.** Nine of the attributes above are `gen_ai.*` names. Three are used as the conventions define them: `gen_ai.request.model`, `gen_ai.response.model`, and `gen_ai.usage.output_tokens` on the session span. The other six are deprecated, removed, not defined or defined differently there, and the observer keeps them for the Langfuse mappings they were added for:
171
+
172
+ - `gen_ai.system` is deprecated in favor of `gen_ai.provider.name`, and `gen_ai.completion` has been removed.
173
+ - `gen_ai.usage.cache_creation_input_tokens`, `gen_ai.usage.cache_read_input_tokens` and `gen_ai.usage.cost` are not names the conventions define. The two cache names are the Anthropic API's field names.
174
+ - `gen_ai.usage.input_tokens` is the API's `input_tokens`, which leaves cache tokens out, while the conventions say the value should include them. A backend that follows the conventions therefore sees only part of the input of a cached turn: 18 tokens for a recorded turn that consumed 49,974. The inclusive count is `llm.token_count.prompt` on the session span.
175
+
176
+ The observer sets neither `gen_ai.operation.name` nor `gen_ai.provider.name`. This comparison was made against `opentelemetry-semantic_conventions` 1.43.0; the GenAI conventions are still marked as in development.
126
177
 
127
178
  The `langfuse.observation.type` attribute is set on each span (`agent`/`generation`/`tool`) to enable Langfuse's **trace flow diagram** (DAG graph visualization).
128
179