turnkit 0.7.2 → 0.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: fb1f58cfea02169c68e433d8a8b84dc5eb2a8a54f18ea9e89853e4ad85445ccd
4
- data.tar.gz: '085c56655a5e11670528b157926a6003a04a074944adaf7cf61730cc18c5d6a3'
3
+ metadata.gz: 73e85ee51e846b0d62828d964ad53b00fd0f6ce1e11d56611b47b45462d49d9f
4
+ data.tar.gz: f4c4591902e4a9a9fc4c2a50c6e768ad6a4191ac242418434ff5d8b22e99143b
5
5
  SHA512:
6
- metadata.gz: 38ef2c605e8ee0d5a15417db75a6efce91b045103f68f86a241f7db788806b63ab6dbdd16cc40b1bee24b54ed93f6b83fb9331e13dbf6d06900bc79f4e7305d8
7
- data.tar.gz: a9a752449f11422b599b9eaba4598441a295366fc4ab1dddccb136689d9279c2e8b426de9d1f2f8b45fe74011209bd22942e388033892547b5f8355a65d8e56c
6
+ metadata.gz: cd39ba14aa1e48ce133d8b4cffc76d63939494a90466cfd62a5b70cb612f76f2caf8a5ffafc371272f00644ff1842a9b96768fea0724788721c45c3d36163655
7
+ data.tar.gz: a77ae5606a43be2790b54cf2a6c8eb78225b5f8b41d0f7b166b8c67b6568a357e65d1ae8f690b446c55afde370aab33d1d3f21dad30a7aec493eb2eb6007224c
data/CHANGELOG.md CHANGED
@@ -1,5 +1,31 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.9.0 - 2026-09-19
4
+
5
+ - Add `Turn#internal_evaluation` and a dependency-free Cloudflare Jev adapter for
6
+ Noul/Choice/Score judgments, separate from chat. Persist request/attempt receipts
7
+ with fenced atomic usage/cost accounting and bounded opt-in retries.
8
+ - Add optional runtime-owned `output_metadata` to child results so output-audit
9
+ callables can report assessments without changing strict generated schemas.
10
+ - Preserve unknown evaluation costs in cost summaries and retain total-only
11
+ provider charges when aggregating them with itemized costs.
12
+
13
+ ## 0.8.0 - 2026-09-18
14
+
15
+ - Add typed delegation to `SubAgentTool`: subclasses declare their own
16
+ parameters, an `agent` class macro (or instance `agent`), and `task_for` to
17
+ build the child task inside the runtime so bulk data never enters the parent
18
+ context. Emit `sub_agent.delegated` with `task_chars` per delegation, and
19
+ auto-register agents owned by sub-agent tools alongside `sub_agents`.
20
+ - Add `Agent#tool_policy`, a per-agent routing/cost gate that runs after
21
+ authorization and returns `:allow` or `[:block, reason]`; blocked calls return
22
+ the reason to the model with `details["tool_policy_blocked"]`.
23
+ - Add `examples/shunt`, a Spotify Portal-style context router with
24
+ `bulk_read`/`code_write` worker agents, a large-file `read_file` gate, and
25
+ measured savings.
26
+ - Breaking: `SubAgentTool#build_child` is an instance method, and the `task`
27
+ parameter is declared only on `SubAgentTool.for` classes.
28
+
3
29
  ## 0.7.2 - 2026-09-10
4
30
 
5
31
  - Add explicit `Tool.budget_completion!` for replay-safe local terminal saves
data/README.md CHANGED
@@ -18,6 +18,10 @@ For GPT-6 Astra tools, opt into the pinned RubyLLM 2 release candidate and
18
18
  The [live validation app](examples/interactive_validation/README.md) exercises
19
19
  these controls with Rails 8.1, PostgreSQL, Sidekiq, and actual Astra/xhigh requests.
20
20
 
21
+ For non-generative judgments, use [structured evaluations](docs/structured-evaluations.md).
22
+ Cloudflare Jev supports Noul, Choice, and Score through a separate typed API with
23
+ durable receipts, usage/cost accounting, and output-audit integration—not chat messages.
24
+
21
25
  ## Installation
22
26
 
23
27
  Add this line to your application's **Gemfile**:
@@ -590,6 +594,28 @@ puts turn.output_text
590
594
 
591
595
  Rely on TurnKit to validate tools and model-provided arguments.
592
596
 
597
+ #### Tool policies
598
+
599
+ Gate tool calls per agent for routing or cost, separately from identity
600
+ authorization:
601
+
602
+ ```ruby
603
+ agent = TurnKit::Agent.new(
604
+ name: "reporter",
605
+ tools: [ReadFile, BulkRead],
606
+ tool_policy: lambda do |tool:, arguments:, context:|
607
+ next :allow unless tool.is_a?(ReadFile) && File.foreach(arguments["path"]).count > 350
608
+ [:block, "File is large. Use `bulk_read` with a question, or re-read with `offset`/`limit`."]
609
+ end
610
+ )
611
+ ```
612
+
613
+ The policy runs after authorization and before the tool executes. Return
614
+ `:allow` (or `nil`) to proceed, or `[:block, reason]` to return the reason to
615
+ the model as a tool error with `details["tool_policy_blocked"] = true`. Keep
616
+ `authorization_policy` for who may call what; use `tool_policy` for how much and
617
+ which way.
618
+
593
619
  ### Images
594
620
 
595
621
  Generate images inside a durable turn with `turn.paint`. The image call uses the
@@ -839,6 +865,32 @@ puts turn.output_text
839
865
 
840
866
  Use sub-agents for isolated child conversations.
841
867
 
868
+ #### Typed delegation
869
+
870
+ Subclass `SubAgentTool` to build the child task from typed arguments so bulk
871
+ data never enters the parent's context:
872
+
873
+ ```ruby
874
+ class BulkRead < TurnKit::SubAgentTool
875
+ agent reader
876
+ description "Read files and answer a question about them."
877
+
878
+ parameter :question, :string, required: true
879
+ parameter :paths, :array, required: true, items: :string
880
+
881
+ def task_for(question:, paths:)
882
+ files = paths.map { |path| "<file path=\"#{path}\">\n#{File.read(path)}\n</file>" }
883
+ "#{question}\n\n#{files.join("\n")}"
884
+ end
885
+ end
886
+ ```
887
+
888
+ The parent model supplies `question` and `paths`; `task_for` runs inside the
889
+ runtime, and only the child's answer returns to the parent. Override `agent` on
890
+ an instance when the child agent is configured at runtime. Each delegation emits
891
+ `sub_agent.delegated` with `task_chars`, so avoided parent context is
892
+ measurable. See [`examples/shunt`](examples/shunt) for a complete routing setup.
893
+
842
894
  #### Oracle- and Librarian-style specialists
843
895
 
844
896
  Use ordinary agents, not a second specialist runtime. Give each specialist its
@@ -0,0 +1,76 @@
1
+ # frozen_string_literal: true
2
+
3
+ require "net/http"
4
+ require "timeout"
5
+
6
+ module TurnKit
7
+ module Adapters
8
+ class CloudflareJev
9
+ MODEL = "typesafe/jev"
10
+
11
+ def initialize(account_id:, api_token:)
12
+ raise ConfigError, "Cloudflare account ID is required" unless account_id.to_s.match?(/\A[a-zA-Z0-9_-]+\z/)
13
+ raise ConfigError, "Cloudflare API token is required" if api_token.to_s.empty?
14
+ @account_id, @api_token = account_id, api_token
15
+ end
16
+
17
+ def inspect = "#<#{self.class.name}>"
18
+
19
+ # Include route/account, but never the bearer credential, in receipt identity.
20
+ def identity = { "provider" => "cloudflare", "account" => @account_id, "schema" => 1 }
21
+ def cost_model(model:) = "cloudflare/#{model}"
22
+
23
+ def validate!(model:, state:, questions:)
24
+ raise InputError, "unsupported Cloudflare evaluation model" unless model == MODEL
25
+ Evaluation.input!(state: state, questions: questions)
26
+ end
27
+
28
+ def evaluate(model:, state:, questions:, timeout:)
29
+ raise ArgumentError, "timeout must be positive and finite" unless timeout.is_a?(Numeric) && timeout.finite? && timeout.positive?
30
+ input = validate!(model: model, state: state, questions: questions)
31
+ uri = URI("https://api.cloudflare.com/client/v4/accounts/#{@account_id}/ai/run")
32
+ request = Net::HTTP::Post.new(uri)
33
+ request["Authorization"] = "Bearer #{@api_token}"
34
+ request["Content-Type"] = "application/json"
35
+ request.body = JSON.generate(model: model, input: input)
36
+ response = Timeout.timeout(timeout) do
37
+ Net::HTTP.start(uri.host, uri.port, use_ssl: true, open_timeout: timeout, read_timeout: timeout, write_timeout: timeout) do |http|
38
+ http.max_retries = 0
39
+ http.request(request)
40
+ end
41
+ end
42
+ code = response.code.to_i
43
+ unless code.between?(200, 299)
44
+ transient = [408, 429, 500, 502, 503, 504].include?(code)
45
+ # Quota exhaustion is not transient capacity pressure.
46
+ if code == 429
47
+ errors = JSON.parse(response.body).fetch("errors", []) rescue []
48
+ transient = false if Array(errors).any? { |error| error.is_a?(Hash) && error["code"] == 3036 }
49
+ end
50
+ delay = retry_delay(response["Retry-After"])
51
+ raise EvaluationError.new(code >= 500 || code == 408 ? :uncertain : :unavailable,
52
+ retryable: transient, retry_after: delay, http_status: code)
53
+ end
54
+ value = JSON.parse(response.body)
55
+ if value.is_a?(Hash) && value.key?("result")
56
+ raise EvaluationError.new(:unavailable) unless value["success"] == true
57
+ value = value["result"]
58
+ end
59
+ Evaluation.result!(value, questions: input.fetch("questions"))
60
+ rescue JSON::ParserError
61
+ raise EvaluationError.new(:malformed), cause: nil
62
+ rescue Timeout::Error, IOError, SystemCallError, OpenSSL::SSL::SSLError, SocketError, Net::HTTPBadResponse, Net::ProtocolError
63
+ raise EvaluationError.new(:uncertain, retryable: true), cause: nil
64
+ end
65
+
66
+ private
67
+ def retry_delay(value)
68
+ return unless value
69
+ return value.to_f if value.match?(/\A\d+(\.\d+)?\z/)
70
+ [Time.httpdate(value) - Clock.now, 0].max
71
+ rescue ArgumentError
72
+ nil
73
+ end
74
+ end
75
+ end
76
+ end
data/lib/turnkit/agent.rb CHANGED
@@ -24,10 +24,10 @@ module TurnKit
24
24
  attr_reader :name, :description, :model, :instructions, :tools, :skills, :available_skills, :sub_agents
25
25
  attr_reader :client, :store, :max_iterations, :timeout, :max_spend, :max_depth, :max_tool_executions, :max_tool_executions_by_name
26
26
  attr_reader :prompt_sections, :system_prompt, :prompt_mode, :thinking, :compaction, :output_schema, :input_schema, :on_event
27
- attr_reader :output_policy, :output_policy_mode, :output_policy_model, :output_retries, :context_contributors
27
+ attr_reader :output_policy, :output_policy_mode, :output_policy_model, :output_retries, :context_contributors, :tool_policy
28
28
 
29
29
  def initialize(name:, description: "", model: nil, instructions: "", orchestrator: false, tools: [], skills: [], available_skills: [], sub_agents: [],
30
- system_prompt: nil, prompt_sections: nil, prompt_mode: nil, client: nil, store: nil,
30
+ tool_policy: nil, system_prompt: nil, prompt_sections: nil, prompt_mode: nil, client: nil, store: nil,
31
31
  max_iterations: nil, timeout: nil, max_spend: nil, max_depth: nil, max_tool_executions: nil, max_tool_executions_by_name: nil, thinking: nil, compaction: nil,
32
32
  output_schema: nil, input_schema: nil, output_policy: nil, output_policy_mode: nil, output_policy_model: nil, output_policy_thinking: nil, output_retries: 0, on_event: nil, context_contributors: [], inherit_globals: true)
33
33
  @name = name.to_s
@@ -39,6 +39,7 @@ module TurnKit
39
39
  @skills = Array(skills).dup.freeze
40
40
  @available_skills = ((inherit_globals ? Array(TurnKit.available_skills) : []) + Array(available_skills)).uniq { |skill| skill.key }.freeze
41
41
  @sub_agents = Array(sub_agents).dup.freeze
42
+ @tool_policy = tool_policy
42
43
  @system_prompt = system_prompt
43
44
  @prompt_sections = prompt_sections
44
45
  @prompt_mode = prompt_mode&.to_sym || (:task if @orchestrator)
@@ -264,7 +264,7 @@ module TurnKit
264
264
  loaded = Background.load_turn(record.fetch("id"), store: store)
265
265
  tool = loaded.agent.effective_tools(turn: loaded).find { |candidate| candidate.tool_name == execution["tool_name"] }
266
266
  recovery = tool.is_a?(Class) ? tool.recovery : tool&.class&.recovery
267
- ordinary = tool && !(tool.is_a?(Class) && tool < SubAgentTool) &&
267
+ ordinary = tool && !SubAgentTool.delegates?(tool) &&
268
268
  ![WaitTool, LaunchAgentTool, SendMessageTool].include?(tool)
269
269
  if ordinary && recovery == :replay_safe
270
270
  store.claim_tool_execution(execution.fetch("id"), to: "pending", started_at: nil)
@@ -39,7 +39,7 @@ module TurnKit
39
39
  if existing
40
40
  child = existing
41
41
  else
42
- built = SubAgentTool.for(agent).build_child(task: task, context: context)
42
+ built = SubAgentTool.for(agent).new.build_child(task: task, context: context)
43
43
  options = parent.store.load_turn(built.id).fetch("options")
44
44
  options = options.merge("callback_conversation_id" => parent.conversation.id) if callback
45
45
  child = parent.store.update_turn(built.id, submitted_at: Clock.now, options: options)
data/lib/turnkit/cost.rb CHANGED
@@ -10,6 +10,13 @@ module TurnKit
10
10
  def self.aggregate(costs)
11
11
  costs = costs.compact
12
12
  return new unless costs.any?
13
+ return new(unknown: true) if costs.any?(&:unknown?)
14
+ # A provider-supplied total cannot be assigned to a token component.
15
+ # Preserve it when combining chat totals with itemized evaluation prices.
16
+ if costs.any? { |cost| cost.total && COMPONENTS.all? { |component| cost.public_send(component).nil? } }
17
+ totals = costs.map(&:total)
18
+ return totals.any?(&:nil?) ? new : new(total: totals.sum)
19
+ end
13
20
 
14
21
  if costs.any? { |cost| COMPONENTS.any? { |component| !cost.public_send(component).nil? } }
15
22
  values = COMPONENTS.to_h do |component|
@@ -41,6 +48,10 @@ module TurnKit
41
48
 
42
49
  def self.from_record(record)
43
50
  attrs = record.transform_keys(&:to_s)
51
+ evaluations = attrs.dig("options", "state", "evaluations") || {}
52
+ if evaluations.values.any? { |receipt| receipt.fetch("attempts", []).any? { |attempt| attempt["cost"].nil? } }
53
+ return new(unknown: true)
54
+ end
44
55
  usage = attrs["usage"] || {}
45
56
  return from_hash(usage["cost_details"] || usage[:cost_details]) if usage["cost_details"] || usage[:cost_details]
46
57
  return new(total: attrs["cost"]) if attrs["cost"]
@@ -93,7 +104,8 @@ module TurnKit
93
104
  cache_read: hash[:cache_read],
94
105
  cache_write: hash[:cache_write],
95
106
  thinking: hash[:thinking],
96
- total: hash[:total]
107
+ total: hash[:total],
108
+ unknown: hash[:unknown] || false
97
109
  )
98
110
  end
99
111
 
@@ -120,7 +132,7 @@ module TurnKit
120
132
  tokens.to_i * price.to_f / PER_MILLION
121
133
  end
122
134
 
123
- def initialize(input: nil, output: nil, cache_read: nil, cache_write: nil, thinking: nil, total: nil, strict: false)
135
+ def initialize(input: nil, output: nil, cache_read: nil, cache_write: nil, thinking: nil, total: nil, strict: false, unknown: false)
124
136
  @input = number(input)
125
137
  @output = number(output)
126
138
  @cache_read = number(cache_read)
@@ -128,9 +140,13 @@ module TurnKit
128
140
  @thinking = number(thinking)
129
141
  @total = number(total)
130
142
  @strict = strict
143
+ @unknown = unknown
131
144
  end
132
145
 
146
+ def unknown? = @unknown
147
+
133
148
  def total
149
+ return nil if unknown?
134
150
  return @total if @total
135
151
  return nil if @strict && COMPONENTS.any? { |component| public_send(component).nil? }
136
152
 
@@ -145,7 +161,8 @@ module TurnKit
145
161
  "cache_read" => cache_read,
146
162
  "cache_write" => cache_write,
147
163
  "thinking" => thinking,
148
- "total" => total
164
+ "total" => total,
165
+ "unknown" => (true if unknown?)
149
166
  }.compact
150
167
  end
151
168
 
@@ -0,0 +1,119 @@
1
+ # frozen_string_literal: true
2
+
3
+ module TurnKit
4
+ # Deliberately separate from Result: evaluations have no message parts or text.
5
+ class EvaluationResult
6
+ attr_reader :model, :answers, :usage, :receipt_id
7
+
8
+ def initialize(model:, answers:, usage:, receipt_id: nil)
9
+ @model, @answers, @usage, @receipt_id = model, answers, usage, receipt_id
10
+ end
11
+
12
+ def to_h
13
+ { "model" => model, "answers" => answers, "usage" => usage.to_h }
14
+ end
15
+
16
+ def self.from_h(value, receipt_id: nil)
17
+ new(model: value.fetch("model"), answers: value.fetch("answers"),
18
+ usage: Usage.from_h(value.fetch("usage")), receipt_id: receipt_id)
19
+ end
20
+ end
21
+
22
+ class EvaluationError < Error
23
+ attr_reader :status, :retry_after, :usage, :model, :http_status
24
+
25
+ def initialize(status, retryable: false, retry_after: nil, usage: nil, model: nil, http_status: nil)
26
+ @status, @retryable, @retry_after, @usage, @model = status.to_s, retryable, retry_after, usage, model
27
+ @http_status = http_status
28
+ # Never include provider bodies, URLs, state, question names, or credentials.
29
+ super("evaluation #{@status}")
30
+ end
31
+
32
+ def retryable? = @retryable
33
+ end
34
+
35
+ module Evaluation
36
+ module_function
37
+
38
+ # Canonical JSON also rejects Ruby objects which JSON.generate would stringify.
39
+ def json(value)
40
+ case value
41
+ when Hash
42
+ raise InputError, "evaluation keys must be strings or symbols" unless value.keys.all? { |k| k.is_a?(String) || k.is_a?(Symbol) }
43
+ raise InputError, "duplicate evaluation keys" unless value.keys.map(&:to_s).uniq.size == value.size
44
+ value.transform_keys(&:to_s).sort.to_h.transform_values { |v| json(v) }
45
+ when Array then value.map { |v| json(v) }
46
+ when String, Integer, TrueClass, FalseClass, NilClass then value
47
+ when Float
48
+ raise InputError, "evaluation numbers must be finite" unless value.finite?
49
+ value
50
+ else raise InputError, "evaluation values must be JSON"
51
+ end
52
+ end
53
+
54
+ def text?(value) = value.nil? || value.is_a?(String) || value.is_a?(Hash) || value.is_a?(Array)
55
+
56
+ def input!(state:, questions:)
57
+ input = json({ state: state, questions: questions })
58
+ valid = text?(input["state"]) && input["questions"].is_a?(Hash)
59
+ raise InputError, "invalid evaluation input" unless valid
60
+ input["questions"].each do |id, question|
61
+ valid = !id.empty? && question.is_a?(Hash) && (question.keys - %w[type instructions criteria]).empty? &&
62
+ question.key?("instructions") && text?(question["instructions"])
63
+ raise InputError, "invalid evaluation question" unless valid
64
+ criteria = question["criteria"]
65
+ valid = case question["type"]
66
+ when "noul"
67
+ criteria.nil? || (criteria.is_a?(Hash) && (criteria.keys - %w[true false]).empty? && criteria.values.all? { |v| text?(v) })
68
+ when "choice"
69
+ criteria.is_a?(Hash) && criteria.values.all? { |v| text?(v) }
70
+ when "score"
71
+ criteria.is_a?(Array) && criteria.size >= 2 && criteria.all? { |v| text?(v) }
72
+ else false
73
+ end
74
+ raise InputError, "invalid evaluation criteria" unless valid
75
+ end
76
+ input
77
+ end
78
+
79
+ def probability?(value) = value.is_a?(Numeric) && value.finite? && value.between?(0, 1)
80
+
81
+ def usage(value)
82
+ return unless value.is_a?(Hash) && value.keys.sort == %w[input_tokens output_tokens]
83
+ return unless value.values.all? { |v| v.is_a?(Integer) && v.between?(0, 9_007_199_254_740_991) }
84
+ Usage.new(input_tokens: value["input_tokens"], output_tokens: value["output_tokens"])
85
+ end
86
+
87
+ def result!(value, questions:)
88
+ observed = usage(value["usage"]) if value.is_a?(Hash)
89
+ model = value["model"] if value.is_a?(Hash) && value["model"].is_a?(String) && !value["model"].empty?
90
+ invalid = -> { raise EvaluationError.new(:malformed, usage: observed, model: model) }
91
+ invalid.call unless value.is_a?(Hash) && value.keys.sort == %w[answers model usage] && model && observed
92
+ answers = value["answers"]
93
+ invalid.call unless answers.is_a?(Hash) && answers.keys.sort == questions.keys.sort
94
+ answers.each do |id, answer|
95
+ question = questions.fetch(id)
96
+ type = question.fetch("type")
97
+ invalid.call unless answer.is_a?(Hash) && answer["type"] == type
98
+ if type == "noul"
99
+ invalid.call unless answer.keys.sort == %w[noul type] && probability?(answer["noul"])
100
+ next
101
+ end
102
+ keys = type == "choice" ? %w[choice confidence probabilities type] : %w[confidence legend probabilities score type]
103
+ invalid.call unless answer.keys.sort == keys && probability?(answer["confidence"])
104
+ expected = type == "choice" ? question.fetch("criteria").keys : question.fetch("criteria").each_index.map(&:to_s)
105
+ probabilities = answer["probabilities"]
106
+ invalid.call unless probabilities.is_a?(Hash) && probabilities.keys.sort == expected.sort &&
107
+ probabilities.values.all? { |p| probability?(p) } && (probabilities.values.sum - 1).abs <= 0.01
108
+ if type == "choice"
109
+ invalid.call unless expected.include?(answer["choice"])
110
+ else
111
+ score, legend = answer.values_at("score", "legend")
112
+ invalid.call unless score.is_a?(Numeric) && score.finite? && score.between?(0, expected.size - 1) &&
113
+ legend.is_a?(Hash) && legend.keys.sort == expected.sort && legend.values.all? { |v| v.is_a?(String) }
114
+ end
115
+ end
116
+ EvaluationResult.new(model: model, answers: answers, usage: observed)
117
+ end
118
+ end
119
+ end
@@ -0,0 +1,142 @@
1
+ # frozen_string_literal: true
2
+
3
+ module TurnKit
4
+ module InternalEvaluation
5
+ # Call from a tool or output-audit callable, never from a persistence callback.
6
+ # before_dispatch is application transport/spend policy, not model input.
7
+ def internal_evaluation(evaluator:, model:, purpose:, state:, questions:, policy_version:,
8
+ candidate: nil, max_attempts: 1, timeout: 30, before_dispatch: nil)
9
+ raise LostClaim, "evaluation requires an owned execution" unless store.is_a?(ExecutionStore)
10
+ raise ArgumentError, "max_attempts must be 1..3" unless max_attempts.is_a?(Integer) && (1..3).include?(max_attempts)
11
+ raise ArgumentError, "timeout must be positive and finite" unless timeout.is_a?(Numeric) && timeout.finite? && timeout.positive?
12
+ input = evaluator.validate!(model: model, state: state, questions: questions)
13
+ reload
14
+ identity = Evaluation.json(provider: evaluator.identity, model: model, purpose: purpose.to_s,
15
+ policy_version: policy_version.to_s, candidate: candidate, input: input, schema: 1,
16
+ output_candidate: @record.dig("options", "state", "candidate"),
17
+ output_data: @record.dig("options", "state", "output_data"),
18
+ timeout: timeout, max_attempts: max_attempts)
19
+ key = Digest::SHA256.hexdigest(JSON.generate(identity))
20
+ cost_model = evaluator.cost_model(model: model)
21
+ deadline = Clock.now + timeout
22
+ loop do
23
+ receipt = store.atomic do
24
+ reload
25
+ raise LostClaim, "evaluation requires an owned execution" unless store.is_a?(ExecutionStore) && running?
26
+ (evaluation_receipts[key] || { "id" => key, "model" => model,
27
+ "provider" => identity["provider"], "purpose" => purpose.to_s,
28
+ "policy_version" => policy_version.to_s, "cost_model" => cost_model, "attempts" => [] })
29
+ end
30
+ return EvaluationResult.from_h(receipt.fetch("result"), receipt_id: key) if receipt["status"] == "completed"
31
+ last = receipt["attempts"].last
32
+ if last && (!last["retryable"] || receipt["attempts"].size >= max_attempts)
33
+ raise EvaluationError.new(last.fetch("status"), http_status: last["http_status"])
34
+ end
35
+ if last && last["owner"] == @record["claim_token"] && !last["finished_at"]
36
+ raise EvaluationError.new(:in_progress)
37
+ end
38
+ delay = last ? [last.fetch("retry_at", 0) - Clock.now.to_f, 0].max : 0
39
+ remaining = check_evaluation_dispatch!(deadline)
40
+ raise BudgetError, "evaluation deadline exceeded" if delay >= remaining
41
+ sleep(delay) if delay.positive?
42
+ Authorization.authorize!(:evaluate, principal: @record.dig("options", "principal"),
43
+ turn: self, evaluator: evaluator, model: model, purpose: purpose)
44
+ before_dispatch.call(turn: self, model: model, purpose: purpose) if before_dispatch
45
+ attempt = nil
46
+ remaining = nil
47
+ store.atomic do
48
+ reload
49
+ remaining = check_evaluation_dispatch!(deadline)
50
+ current = evaluation_receipts[key]
51
+ # A second caller may have committed or dispatched while policy ran.
52
+ raise EvaluationError.new(:in_progress) if current && current != receipt
53
+ attempt = { "id" => SecureRandom.uuid, "owner" => @record["claim_token"],
54
+ "status" => "uncertain", "retryable" => true, "started_at" => Clock.now.iso8601(6),
55
+ "cost" => nil }
56
+ receipt["attempts"] << attempt
57
+ receipt["status"] = "uncertain"
58
+ save_evaluation_receipt!(key, receipt)
59
+ end
60
+ event = { receipt_id: key, attempt_id: attempt["id"], model: model, purpose: purpose.to_s,
61
+ question_count: input.fetch("questions").size }
62
+ emit("evaluation.requested", event)
63
+ result = nil
64
+ failure = nil
65
+ begin
66
+ result = evaluator.evaluate(model: model, state: input.fetch("state"),
67
+ questions: input.fetch("questions"), timeout: remaining)
68
+ rescue EvaluationError => error
69
+ failure = error
70
+ end
71
+ observed_usage = result ? result.usage : failure.usage
72
+ observed_cost = Cost.from_usage(observed_usage, model: cost_model) if observed_usage
73
+ store.atomic do
74
+ reload
75
+ attempt.merge!("status" => result ? "completed" : failure.status,
76
+ "finished_at" => Clock.now.iso8601(6), "retryable" => failure&.retryable? || false,
77
+ "http_status" => failure&.http_status,
78
+ "usage" => observed_usage&.to_h, "cost" => observed_cost&.total,
79
+ "returned_model" => result ? result.model : failure.model)
80
+ if failure&.retryable?
81
+ attempt["retry_at"] = Clock.now.to_f + (failure.retry_after || rand * (2 ** (receipt["attempts"].size - 1)))
82
+ end
83
+ receipt["status"] = attempt["status"]
84
+ receipt["result"] = result.to_h if result
85
+ add_usage!(observed_usage, cost: observed_cost) if observed_usage
86
+ save_evaluation_receipt!(key, receipt)
87
+ end
88
+ emit(result ? "evaluation.completed" : "evaluation.failed", event.merge(
89
+ status: attempt["status"], returned_model: attempt["returned_model"], http_status: attempt["http_status"],
90
+ usage: observed_usage&.to_h, cost: observed_cost&.to_h,
91
+ duration: Clock.now - Time.iso8601(attempt["started_at"])))
92
+ # Keep the shared inline budget in sync without turning a paid successful
93
+ # assessment into a retry just because its own cost exhausted the budget.
94
+ begin
95
+ budget.add_cost!(observed_cost&.total)
96
+ rescue BudgetError
97
+ # The next dispatch checks the persisted root ledger.
98
+ end
99
+ root_budget = execution_budget
100
+ root_budget.check!(depth: depth, allow_exhausted_spend: true)
101
+ raise BudgetError, "evaluation deadline exceeded" if Clock.now >= deadline
102
+ return EvaluationResult.from_h(receipt.fetch("result"), receipt_id: key) if result
103
+ raise failure unless failure.retryable? && receipt["attempts"].size < max_attempts
104
+ end
105
+ end
106
+
107
+ # Copies, not references to live state. Attempts without usage have unknown
108
+ # cost, even if the turn's known-cost subtotal is zero.
109
+ def evaluation_receipts
110
+ JSON.parse(JSON.generate(@record.dig("options", "state", "evaluations") || {}))
111
+ end
112
+
113
+ def output_metadata = JSON.parse(JSON.generate(@record.dig("options", "state", "output_metadata") || {}))
114
+
115
+ def output_metadata=(value)
116
+ value = Evaluation.json(value)
117
+ raise InputError, "output metadata must be an object" unless value.is_a?(Hash)
118
+ raise LostClaim, "output metadata requires an owned execution" unless store.is_a?(ExecutionStore)
119
+ store.atomic do
120
+ reload
121
+ raise LostClaim, "output metadata requires an owned execution" unless store.is_a?(ExecutionStore) && running?
122
+ update_state!("output_metadata" => value)
123
+ end
124
+ end
125
+
126
+ private
127
+ def save_evaluation_receipt!(key, receipt)
128
+ update_state!("evaluations" => evaluation_receipts.merge(key => receipt))
129
+ end
130
+
131
+ def check_evaluation_dispatch!(deadline)
132
+ current_budget = execution_budget
133
+ current_budget.check!(depth: depth)
134
+ raise LostClaim, "evaluation execution is no longer running" unless running?
135
+ raise EvaluationError.new(:interrupted) if @record.dig("options", "controls", "pause_requested")
136
+ root_deadline = current_budget.root_started_at + current_budget.timeout if current_budget.timeout
137
+ remaining = [(deadline - Clock.now), (root_deadline - Clock.now if root_deadline)].compact.min
138
+ raise BudgetError, "evaluation deadline exceeded" unless remaining.positive?
139
+ remaining
140
+ end
141
+ end
142
+ end
@@ -1,24 +1,49 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module TurnKit
4
+ # Runs a child agent in a fresh conversation and returns only its final result.
5
+ # `SubAgentTool.for(agent)` exposes an agent as a tool taking `task`. Subclasses
6
+ # declare their own parameters and override `task_for` to assemble the task in
7
+ # Ruby, so bulk data (file contents, records) reaches the child without ever
8
+ # entering the parent model's context. The child comes from the `agent` class
9
+ # macro or an instance-level `agent` override.
4
10
  class SubAgentTool < Tool
5
- parameter :task, :string, required: true, description: "The complete task for the sub-agent, including all relevant context."
6
-
7
- def self.for(agent)
8
- Class.new(self) do
9
- @agent = agent
10
- tool_name agent.name
11
- description agent.description.empty? ? "Delegate work to #{agent.name}." : agent.description
12
- usage_hint "Use when work can be delegated independently to #{agent.name}. Pass a complete task and only relevant context."
11
+ class << self
12
+ def agent(value = nil)
13
+ @agent = value if value
14
+ @agent || (superclass < SubAgentTool ? superclass.agent : nil)
15
+ end
13
16
 
14
- class << self
15
- attr_reader :agent
17
+ def for(agent)
18
+ sub_agent = agent
19
+ Class.new(self) do
20
+ agent sub_agent
21
+ tool_name sub_agent.name
22
+ description sub_agent.description.empty? ? "Delegate work to #{sub_agent.name}." : sub_agent.description
23
+ usage_hint "Use when work can be delegated independently to #{sub_agent.name}. Pass a complete task and only relevant context."
24
+ parameter :task, :string, required: true, description: "The complete task for the sub-agent, including all relevant context."
16
25
  end
17
26
  end
27
+
28
+ def delegates?(tool)
29
+ tool.is_a?(self) || (tool.is_a?(Class) && tool <= self)
30
+ end
31
+
32
+ def result(record)
33
+ { "conversation_id" => record.fetch("conversation_id"), "turn_id" => record.fetch("id"),
34
+ "status" => record.fetch("status"), "result" => record["output_text"].to_s,
35
+ "output_metadata" => record.dig("options", "state", "output_metadata"),
36
+ "output_data" => record["output_data"], "error" => record["error"] }.compact
37
+ end
18
38
  end
19
39
 
20
- def self.build_child(task:, context:)
21
- sub_agent = agent
40
+ def agent = self.class.agent
41
+
42
+ def task_for(**arguments)
43
+ arguments.fetch(:task)
44
+ end
45
+
46
+ def build_child(task:, context:)
22
47
  parent_turn = context.turn
23
48
  lineage = {
24
49
  "parent_conversation_id" => parent_turn.conversation.id,
@@ -27,32 +52,28 @@ module TurnKit
27
52
  "principal" => context.principal
28
53
  }
29
54
  store = parent_turn.store
30
- record = store.create_conversation("agent_name" => sub_agent.name, "model" => sub_agent.effective_model, "metadata" => lineage)
31
- conversation = Conversation.new(agent: sub_agent, record: record, store: store, model: sub_agent.effective_model, metadata: lineage)
55
+ record = store.create_conversation("agent_name" => agent.name, "model" => agent.effective_model, "metadata" => lineage)
56
+ conversation = Conversation.new(agent: agent, record: record, store: store, model: agent.effective_model, metadata: lineage)
32
57
  trigger = conversation.say(task, metadata: lineage)
58
+ parent_turn.emit("sub_agent.delegated", id: context.execution.tool_call_id, name: agent.name,
59
+ conversation_id: record.fetch("id"), task_chars: task.length)
33
60
  conversation.build_turn(
34
61
  trigger_message_id: trigger.id,
35
62
  budget: parent_turn.budget,
36
63
  parent_turn: parent_turn,
37
64
  parent_tool_execution: context.execution,
38
65
  depth: parent_turn.depth + 1,
39
- model: sub_agent.effective_model,
40
- agent: sub_agent,
66
+ model: agent.effective_model,
67
+ agent: agent,
41
68
  principal: context.principal,
42
69
  on_event: parent_turn.agent.effective_on_event
43
70
  )
44
71
  end
45
72
 
46
- def self.result(record)
47
- { "conversation_id" => record.fetch("conversation_id"), "turn_id" => record.fetch("id"),
48
- "status" => record.fetch("status"), "result" => record["output_text"].to_s,
49
- "output_data" => record["output_data"], "error" => record["error"] }.compact
50
- end
51
-
52
- def call(task:, context:)
73
+ def call(context:, **arguments)
53
74
  Authorization.authorize!(:launch_agent, principal: context.principal, turn: context.turn,
54
- agent: self.class.agent, arguments: { "task" => task })
55
- child = self.class.build_child(task: task, context: context)
75
+ agent: agent, arguments: arguments.transform_keys(&:to_s))
76
+ child = build_child(task: task_for(**arguments), context: context)
56
77
  child.run!
57
78
  SubAgentTool.result(child.store.load_turn(child.id))
58
79
  end
@@ -93,6 +93,8 @@ module TurnKit
93
93
  context = ToolContext.new(turn: turn, execution: execution)
94
94
  payload = begin
95
95
  Authorization.authorize!(:tool, principal: context.principal, turn: turn, tool: tool, arguments: tool_call.arguments)
96
+ blocked = tool_policy_block(tool, tool_call.arguments, context)
97
+ return finish_error(execution, tool_call, blocked, details: { "tool_policy_blocked" => true }) if blocked
96
98
  # Observe cancellation/reconciliation immediately before crossing the
97
99
  # external-effect boundary. Calls already sent cannot be recalled.
98
100
  control = turn.control_boundary!
@@ -191,20 +193,30 @@ module TurnKit
191
193
  end
192
194
 
193
195
  def subagent?(tool)
194
- tool.is_a?(Class) && tool < SubAgentTool
196
+ SubAgentTool.delegates?(tool)
197
+ end
198
+
199
+ # Agent-owned routing/cost policy, distinct from identity authorization.
200
+ # Returns the block reason the model sees, or nil to proceed.
201
+ def tool_policy_block(tool, arguments, context)
202
+ decision, reason = turn.agent.tool_policy&.call(tool: tool, arguments: arguments, context: context)
203
+ return reason.to_s if decision == :block
204
+ raise ArgumentError, "tool_policy must return :allow or [:block, reason]" unless [ nil, :allow ].include?(decision)
195
205
  end
196
206
 
197
207
  def delegate(tool, call, context)
198
- arguments = tool.validate_arguments(call.arguments)
208
+ tool = tool.new if tool.is_a?(Class)
209
+ arguments = tool.class.validate_arguments(call.arguments)
199
210
  Authorization.authorize!(:launch_agent, principal: context.principal, turn: turn, agent: tool.agent, arguments: arguments)
200
211
  TurnKit.resolve_agent(tool.agent.name)
212
+ task = tool.task_for(**arguments.transform_keys(&:to_sym))
201
213
  child = turn.store.atomic_graph do
202
214
  turn.store.atomic(Background.root_conversation(turn.store, turn.store.load_turn(turn.id))) do
203
215
  control = turn.control_boundary!
204
216
  next control if control
205
217
  row = turn.store.list_turns(root_turn_id: turn.root_turn_id).find { |candidate| candidate["parent_tool_execution_id"] == context.execution.id }
206
218
  unless row
207
- built = tool.build_child(task: arguments.fetch("task"), context: context)
219
+ built = tool.build_child(task: task, context: context)
208
220
  row = turn.store.update_turn(built.id, submitted_at: Clock.now)
209
221
  end
210
222
  Background.wait(turn, [row.fetch("id")])
data/lib/turnkit/turn.rb CHANGED
@@ -3,6 +3,7 @@
3
3
  module TurnKit
4
4
  class Turn
5
5
  include TurnControls
6
+ include InternalEvaluation
6
7
  STATUSES = Record::TURN_STATUSES
7
8
 
8
9
  attr_reader :agent, :conversation, :store, :budget, :depth
@@ -325,6 +326,7 @@ module TurnKit
325
326
  add_usage!(result.usage, cost: cost)
326
327
  persist_assistant_message(result)
327
328
  update_state!("phase" => result.tool_calls? ? "tools" : "output", "parts" => result.parts,
329
+ "output_metadata" => nil,
328
330
  "candidate" => result.text, "output_data" => result.output_data, "terminal_tool_name" => nil,
329
331
  "budget_completion_call_id" => select_budget_completion(result))
330
332
  end
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module TurnKit
4
- VERSION = "0.7.2"
4
+ VERSION = "0.9.0"
5
5
  end
data/lib/turnkit.rb CHANGED
@@ -46,9 +46,12 @@ require_relative "turnkit/load_skill_tool"
46
46
  require_relative "turnkit/message_projection"
47
47
  require_relative "turnkit/tool_runner"
48
48
  require_relative "turnkit/turn_controls"
49
+ require_relative "turnkit/evaluation"
50
+ require_relative "turnkit/internal_evaluation"
49
51
  require_relative "turnkit/turn"
50
52
  require_relative "turnkit/usage"
51
53
  require_relative "turnkit/run"
54
+ require_relative "turnkit/adapters/cloudflare_jev"
52
55
  require_relative "turnkit/adapters/codex"
53
56
  require_relative "turnkit/adapters/ruby_llm"
54
57
  require_relative "turnkit/active_record_store"
@@ -78,7 +81,8 @@ module TurnKit
78
81
 
79
82
  def self.register(agent)
80
83
  @agents[agent.name] = agent
81
- agent.sub_agents.each { |child| register(child) }
84
+ tools = agent.effective_tools + agent.available_skills.flat_map(&:tools)
85
+ tools.each { |tool| register(tool.agent) if SubAgentTool.delegates?(tool) }
82
86
  agent
83
87
  end
84
88
 
metadata CHANGED
@@ -1,14 +1,14 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: turnkit
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.7.2
4
+ version: 0.9.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Sam Couch
8
8
  autorequire:
9
9
  bindir: bin
10
10
  cert_chain: []
11
- date: 2026-09-10 00:00:00.000000000 Z
11
+ date: 2026-09-19 00:00:00.000000000 Z
12
12
  dependencies: []
13
13
  description: TurnKit is a Ruby/Rails agent runtime for durable AI conversations, application
14
14
  runs, orchestrator agents, tool calling, skills, sub-agents, context compaction,
@@ -37,6 +37,7 @@ files:
37
37
  - lib/generators/turnkit/upgrade_generator.rb
38
38
  - lib/turnkit.rb
39
39
  - lib/turnkit/active_record_store.rb
40
+ - lib/turnkit/adapters/cloudflare_jev.rb
40
41
  - lib/turnkit/adapters/codex.rb
41
42
  - lib/turnkit/adapters/ruby_llm.rb
42
43
  - lib/turnkit/agent.rb
@@ -50,11 +51,13 @@ files:
50
51
  - lib/turnkit/coordination_tools.rb
51
52
  - lib/turnkit/cost.rb
52
53
  - lib/turnkit/error.rb
54
+ - lib/turnkit/evaluation.rb
53
55
  - lib/turnkit/event.rb
54
56
  - lib/turnkit/execution_store.rb
55
57
  - lib/turnkit/id.rb
56
58
  - lib/turnkit/image_result.rb
57
59
  - lib/turnkit/image_tool.rb
60
+ - lib/turnkit/internal_evaluation.rb
58
61
  - lib/turnkit/job.rb
59
62
  - lib/turnkit/load_skill_tool.rb
60
63
  - lib/turnkit/media_analysis_result.rb