turnkit 0.7.2 → 0.9.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +26 -0
- data/README.md +52 -0
- data/lib/turnkit/adapters/cloudflare_jev.rb +76 -0
- data/lib/turnkit/agent.rb +3 -2
- data/lib/turnkit/background.rb +1 -1
- data/lib/turnkit/coordination_tools.rb +1 -1
- data/lib/turnkit/cost.rb +20 -3
- data/lib/turnkit/evaluation.rb +119 -0
- data/lib/turnkit/internal_evaluation.rb +142 -0
- data/lib/turnkit/sub_agent_tool.rb +46 -25
- data/lib/turnkit/tool_runner.rb +15 -3
- data/lib/turnkit/turn.rb +2 -0
- data/lib/turnkit/version.rb +1 -1
- data/lib/turnkit.rb +5 -1
- metadata +5 -2
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 73e85ee51e846b0d62828d964ad53b00fd0f6ce1e11d56611b47b45462d49d9f
|
|
4
|
+
data.tar.gz: f4c4591902e4a9a9fc4c2a50c6e768ad6a4191ac242418434ff5d8b22e99143b
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: cd39ba14aa1e48ce133d8b4cffc76d63939494a90466cfd62a5b70cb612f76f2caf8a5ffafc371272f00644ff1842a9b96768fea0724788721c45c3d36163655
|
|
7
|
+
data.tar.gz: a77ae5606a43be2790b54cf2a6c8eb78225b5f8b41d0f7b166b8c67b6568a357e65d1ae8f690b446c55afde370aab33d1d3f21dad30a7aec493eb2eb6007224c
|
data/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,31 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.9.0 - 2026-09-19
|
|
4
|
+
|
|
5
|
+
- Add `Turn#internal_evaluation` and a dependency-free Cloudflare Jev adapter for
|
|
6
|
+
Noul/Choice/Score judgments, separate from chat. Persist request/attempt receipts
|
|
7
|
+
with fenced atomic usage/cost accounting and bounded opt-in retries.
|
|
8
|
+
- Add optional runtime-owned `output_metadata` to child results so output-audit
|
|
9
|
+
callables can report assessments without changing strict generated schemas.
|
|
10
|
+
- Preserve unknown evaluation costs in cost summaries and retain total-only
|
|
11
|
+
provider charges when aggregating them with itemized costs.
|
|
12
|
+
|
|
13
|
+
## 0.8.0 - 2026-09-18
|
|
14
|
+
|
|
15
|
+
- Add typed delegation to `SubAgentTool`: subclasses declare their own
|
|
16
|
+
parameters, an `agent` class macro (or instance `agent`), and `task_for` to
|
|
17
|
+
build the child task inside the runtime so bulk data never enters the parent
|
|
18
|
+
context. Emit `sub_agent.delegated` with `task_chars` per delegation, and
|
|
19
|
+
auto-register agents owned by sub-agent tools alongside `sub_agents`.
|
|
20
|
+
- Add `Agent#tool_policy`, a per-agent routing/cost gate that runs after
|
|
21
|
+
authorization and returns `:allow` or `[:block, reason]`; blocked calls return
|
|
22
|
+
the reason to the model with `details["tool_policy_blocked"]`.
|
|
23
|
+
- Add `examples/shunt`, a Spotify Portal-style context router with
|
|
24
|
+
`bulk_read`/`code_write` worker agents, a large-file `read_file` gate, and
|
|
25
|
+
measured savings.
|
|
26
|
+
- Breaking: `SubAgentTool#build_child` is an instance method, and the `task`
|
|
27
|
+
parameter is declared only on `SubAgentTool.for` classes.
|
|
28
|
+
|
|
3
29
|
## 0.7.2 - 2026-09-10
|
|
4
30
|
|
|
5
31
|
- Add explicit `Tool.budget_completion!` for replay-safe local terminal saves
|
data/README.md
CHANGED
|
@@ -18,6 +18,10 @@ For GPT-6 Astra tools, opt into the pinned RubyLLM 2 release candidate and
|
|
|
18
18
|
The [live validation app](examples/interactive_validation/README.md) exercises
|
|
19
19
|
these controls with Rails 8.1, PostgreSQL, Sidekiq, and actual Astra/xhigh requests.
|
|
20
20
|
|
|
21
|
+
For non-generative judgments, use [structured evaluations](docs/structured-evaluations.md).
|
|
22
|
+
Cloudflare Jev supports Noul, Choice, and Score through a separate typed API with
|
|
23
|
+
durable receipts, usage/cost accounting, and output-audit integration—not chat messages.
|
|
24
|
+
|
|
21
25
|
## Installation
|
|
22
26
|
|
|
23
27
|
Add this line to your application's **Gemfile**:
|
|
@@ -590,6 +594,28 @@ puts turn.output_text
|
|
|
590
594
|
|
|
591
595
|
Rely on TurnKit to validate tools and model-provided arguments.
|
|
592
596
|
|
|
597
|
+
#### Tool policies
|
|
598
|
+
|
|
599
|
+
Gate tool calls per agent for routing or cost, separately from identity
|
|
600
|
+
authorization:
|
|
601
|
+
|
|
602
|
+
```ruby
|
|
603
|
+
agent = TurnKit::Agent.new(
|
|
604
|
+
name: "reporter",
|
|
605
|
+
tools: [ReadFile, BulkRead],
|
|
606
|
+
tool_policy: lambda do |tool:, arguments:, context:|
|
|
607
|
+
next :allow unless tool.is_a?(ReadFile) && File.foreach(arguments["path"]).count > 350
|
|
608
|
+
[:block, "File is large. Use `bulk_read` with a question, or re-read with `offset`/`limit`."]
|
|
609
|
+
end
|
|
610
|
+
)
|
|
611
|
+
```
|
|
612
|
+
|
|
613
|
+
The policy runs after authorization and before the tool executes. Return
|
|
614
|
+
`:allow` (or `nil`) to proceed, or `[:block, reason]` to return the reason to
|
|
615
|
+
the model as a tool error with `details["tool_policy_blocked"] = true`. Keep
|
|
616
|
+
`authorization_policy` for who may call what; use `tool_policy` for how much and
|
|
617
|
+
which way.
|
|
618
|
+
|
|
593
619
|
### Images
|
|
594
620
|
|
|
595
621
|
Generate images inside a durable turn with `turn.paint`. The image call uses the
|
|
@@ -839,6 +865,32 @@ puts turn.output_text
|
|
|
839
865
|
|
|
840
866
|
Use sub-agents for isolated child conversations.
|
|
841
867
|
|
|
868
|
+
#### Typed delegation
|
|
869
|
+
|
|
870
|
+
Subclass `SubAgentTool` to build the child task from typed arguments so bulk
|
|
871
|
+
data never enters the parent's context:
|
|
872
|
+
|
|
873
|
+
```ruby
|
|
874
|
+
class BulkRead < TurnKit::SubAgentTool
|
|
875
|
+
agent reader
|
|
876
|
+
description "Read files and answer a question about them."
|
|
877
|
+
|
|
878
|
+
parameter :question, :string, required: true
|
|
879
|
+
parameter :paths, :array, required: true, items: :string
|
|
880
|
+
|
|
881
|
+
def task_for(question:, paths:)
|
|
882
|
+
files = paths.map { |path| "<file path=\"#{path}\">\n#{File.read(path)}\n</file>" }
|
|
883
|
+
"#{question}\n\n#{files.join("\n")}"
|
|
884
|
+
end
|
|
885
|
+
end
|
|
886
|
+
```
|
|
887
|
+
|
|
888
|
+
The parent model supplies `question` and `paths`; `task_for` runs inside the
|
|
889
|
+
runtime, and only the child's answer returns to the parent. Override `agent` on
|
|
890
|
+
an instance when the child agent is configured at runtime. Each delegation emits
|
|
891
|
+
`sub_agent.delegated` with `task_chars`, so avoided parent context is
|
|
892
|
+
measurable. See [`examples/shunt`](examples/shunt) for a complete routing setup.
|
|
893
|
+
|
|
842
894
|
#### Oracle- and Librarian-style specialists
|
|
843
895
|
|
|
844
896
|
Use ordinary agents, not a second specialist runtime. Give each specialist its
|
|
@@ -0,0 +1,76 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
require "net/http"
|
|
4
|
+
require "timeout"
|
|
5
|
+
|
|
6
|
+
module TurnKit
|
|
7
|
+
module Adapters
|
|
8
|
+
class CloudflareJev
|
|
9
|
+
MODEL = "typesafe/jev"
|
|
10
|
+
|
|
11
|
+
def initialize(account_id:, api_token:)
|
|
12
|
+
raise ConfigError, "Cloudflare account ID is required" unless account_id.to_s.match?(/\A[a-zA-Z0-9_-]+\z/)
|
|
13
|
+
raise ConfigError, "Cloudflare API token is required" if api_token.to_s.empty?
|
|
14
|
+
@account_id, @api_token = account_id, api_token
|
|
15
|
+
end
|
|
16
|
+
|
|
17
|
+
def inspect = "#<#{self.class.name}>"
|
|
18
|
+
|
|
19
|
+
# Include route/account, but never the bearer credential, in receipt identity.
|
|
20
|
+
def identity = { "provider" => "cloudflare", "account" => @account_id, "schema" => 1 }
|
|
21
|
+
def cost_model(model:) = "cloudflare/#{model}"
|
|
22
|
+
|
|
23
|
+
def validate!(model:, state:, questions:)
|
|
24
|
+
raise InputError, "unsupported Cloudflare evaluation model" unless model == MODEL
|
|
25
|
+
Evaluation.input!(state: state, questions: questions)
|
|
26
|
+
end
|
|
27
|
+
|
|
28
|
+
def evaluate(model:, state:, questions:, timeout:)
|
|
29
|
+
raise ArgumentError, "timeout must be positive and finite" unless timeout.is_a?(Numeric) && timeout.finite? && timeout.positive?
|
|
30
|
+
input = validate!(model: model, state: state, questions: questions)
|
|
31
|
+
uri = URI("https://api.cloudflare.com/client/v4/accounts/#{@account_id}/ai/run")
|
|
32
|
+
request = Net::HTTP::Post.new(uri)
|
|
33
|
+
request["Authorization"] = "Bearer #{@api_token}"
|
|
34
|
+
request["Content-Type"] = "application/json"
|
|
35
|
+
request.body = JSON.generate(model: model, input: input)
|
|
36
|
+
response = Timeout.timeout(timeout) do
|
|
37
|
+
Net::HTTP.start(uri.host, uri.port, use_ssl: true, open_timeout: timeout, read_timeout: timeout, write_timeout: timeout) do |http|
|
|
38
|
+
http.max_retries = 0
|
|
39
|
+
http.request(request)
|
|
40
|
+
end
|
|
41
|
+
end
|
|
42
|
+
code = response.code.to_i
|
|
43
|
+
unless code.between?(200, 299)
|
|
44
|
+
transient = [408, 429, 500, 502, 503, 504].include?(code)
|
|
45
|
+
# Quota exhaustion is not transient capacity pressure.
|
|
46
|
+
if code == 429
|
|
47
|
+
errors = JSON.parse(response.body).fetch("errors", []) rescue []
|
|
48
|
+
transient = false if Array(errors).any? { |error| error.is_a?(Hash) && error["code"] == 3036 }
|
|
49
|
+
end
|
|
50
|
+
delay = retry_delay(response["Retry-After"])
|
|
51
|
+
raise EvaluationError.new(code >= 500 || code == 408 ? :uncertain : :unavailable,
|
|
52
|
+
retryable: transient, retry_after: delay, http_status: code)
|
|
53
|
+
end
|
|
54
|
+
value = JSON.parse(response.body)
|
|
55
|
+
if value.is_a?(Hash) && value.key?("result")
|
|
56
|
+
raise EvaluationError.new(:unavailable) unless value["success"] == true
|
|
57
|
+
value = value["result"]
|
|
58
|
+
end
|
|
59
|
+
Evaluation.result!(value, questions: input.fetch("questions"))
|
|
60
|
+
rescue JSON::ParserError
|
|
61
|
+
raise EvaluationError.new(:malformed), cause: nil
|
|
62
|
+
rescue Timeout::Error, IOError, SystemCallError, OpenSSL::SSL::SSLError, SocketError, Net::HTTPBadResponse, Net::ProtocolError
|
|
63
|
+
raise EvaluationError.new(:uncertain, retryable: true), cause: nil
|
|
64
|
+
end
|
|
65
|
+
|
|
66
|
+
private
|
|
67
|
+
def retry_delay(value)
|
|
68
|
+
return unless value
|
|
69
|
+
return value.to_f if value.match?(/\A\d+(\.\d+)?\z/)
|
|
70
|
+
[Time.httpdate(value) - Clock.now, 0].max
|
|
71
|
+
rescue ArgumentError
|
|
72
|
+
nil
|
|
73
|
+
end
|
|
74
|
+
end
|
|
75
|
+
end
|
|
76
|
+
end
|
data/lib/turnkit/agent.rb
CHANGED
|
@@ -24,10 +24,10 @@ module TurnKit
|
|
|
24
24
|
attr_reader :name, :description, :model, :instructions, :tools, :skills, :available_skills, :sub_agents
|
|
25
25
|
attr_reader :client, :store, :max_iterations, :timeout, :max_spend, :max_depth, :max_tool_executions, :max_tool_executions_by_name
|
|
26
26
|
attr_reader :prompt_sections, :system_prompt, :prompt_mode, :thinking, :compaction, :output_schema, :input_schema, :on_event
|
|
27
|
-
attr_reader :output_policy, :output_policy_mode, :output_policy_model, :output_retries, :context_contributors
|
|
27
|
+
attr_reader :output_policy, :output_policy_mode, :output_policy_model, :output_retries, :context_contributors, :tool_policy
|
|
28
28
|
|
|
29
29
|
def initialize(name:, description: "", model: nil, instructions: "", orchestrator: false, tools: [], skills: [], available_skills: [], sub_agents: [],
|
|
30
|
-
system_prompt: nil, prompt_sections: nil, prompt_mode: nil, client: nil, store: nil,
|
|
30
|
+
tool_policy: nil, system_prompt: nil, prompt_sections: nil, prompt_mode: nil, client: nil, store: nil,
|
|
31
31
|
max_iterations: nil, timeout: nil, max_spend: nil, max_depth: nil, max_tool_executions: nil, max_tool_executions_by_name: nil, thinking: nil, compaction: nil,
|
|
32
32
|
output_schema: nil, input_schema: nil, output_policy: nil, output_policy_mode: nil, output_policy_model: nil, output_policy_thinking: nil, output_retries: 0, on_event: nil, context_contributors: [], inherit_globals: true)
|
|
33
33
|
@name = name.to_s
|
|
@@ -39,6 +39,7 @@ module TurnKit
|
|
|
39
39
|
@skills = Array(skills).dup.freeze
|
|
40
40
|
@available_skills = ((inherit_globals ? Array(TurnKit.available_skills) : []) + Array(available_skills)).uniq { |skill| skill.key }.freeze
|
|
41
41
|
@sub_agents = Array(sub_agents).dup.freeze
|
|
42
|
+
@tool_policy = tool_policy
|
|
42
43
|
@system_prompt = system_prompt
|
|
43
44
|
@prompt_sections = prompt_sections
|
|
44
45
|
@prompt_mode = prompt_mode&.to_sym || (:task if @orchestrator)
|
data/lib/turnkit/background.rb
CHANGED
|
@@ -264,7 +264,7 @@ module TurnKit
|
|
|
264
264
|
loaded = Background.load_turn(record.fetch("id"), store: store)
|
|
265
265
|
tool = loaded.agent.effective_tools(turn: loaded).find { |candidate| candidate.tool_name == execution["tool_name"] }
|
|
266
266
|
recovery = tool.is_a?(Class) ? tool.recovery : tool&.class&.recovery
|
|
267
|
-
ordinary = tool && !
|
|
267
|
+
ordinary = tool && !SubAgentTool.delegates?(tool) &&
|
|
268
268
|
![WaitTool, LaunchAgentTool, SendMessageTool].include?(tool)
|
|
269
269
|
if ordinary && recovery == :replay_safe
|
|
270
270
|
store.claim_tool_execution(execution.fetch("id"), to: "pending", started_at: nil)
|
|
@@ -39,7 +39,7 @@ module TurnKit
|
|
|
39
39
|
if existing
|
|
40
40
|
child = existing
|
|
41
41
|
else
|
|
42
|
-
built = SubAgentTool.for(agent).build_child(task: task, context: context)
|
|
42
|
+
built = SubAgentTool.for(agent).new.build_child(task: task, context: context)
|
|
43
43
|
options = parent.store.load_turn(built.id).fetch("options")
|
|
44
44
|
options = options.merge("callback_conversation_id" => parent.conversation.id) if callback
|
|
45
45
|
child = parent.store.update_turn(built.id, submitted_at: Clock.now, options: options)
|
data/lib/turnkit/cost.rb
CHANGED
|
@@ -10,6 +10,13 @@ module TurnKit
|
|
|
10
10
|
def self.aggregate(costs)
|
|
11
11
|
costs = costs.compact
|
|
12
12
|
return new unless costs.any?
|
|
13
|
+
return new(unknown: true) if costs.any?(&:unknown?)
|
|
14
|
+
# A provider-supplied total cannot be assigned to a token component.
|
|
15
|
+
# Preserve it when combining chat totals with itemized evaluation prices.
|
|
16
|
+
if costs.any? { |cost| cost.total && COMPONENTS.all? { |component| cost.public_send(component).nil? } }
|
|
17
|
+
totals = costs.map(&:total)
|
|
18
|
+
return totals.any?(&:nil?) ? new : new(total: totals.sum)
|
|
19
|
+
end
|
|
13
20
|
|
|
14
21
|
if costs.any? { |cost| COMPONENTS.any? { |component| !cost.public_send(component).nil? } }
|
|
15
22
|
values = COMPONENTS.to_h do |component|
|
|
@@ -41,6 +48,10 @@ module TurnKit
|
|
|
41
48
|
|
|
42
49
|
def self.from_record(record)
|
|
43
50
|
attrs = record.transform_keys(&:to_s)
|
|
51
|
+
evaluations = attrs.dig("options", "state", "evaluations") || {}
|
|
52
|
+
if evaluations.values.any? { |receipt| receipt.fetch("attempts", []).any? { |attempt| attempt["cost"].nil? } }
|
|
53
|
+
return new(unknown: true)
|
|
54
|
+
end
|
|
44
55
|
usage = attrs["usage"] || {}
|
|
45
56
|
return from_hash(usage["cost_details"] || usage[:cost_details]) if usage["cost_details"] || usage[:cost_details]
|
|
46
57
|
return new(total: attrs["cost"]) if attrs["cost"]
|
|
@@ -93,7 +104,8 @@ module TurnKit
|
|
|
93
104
|
cache_read: hash[:cache_read],
|
|
94
105
|
cache_write: hash[:cache_write],
|
|
95
106
|
thinking: hash[:thinking],
|
|
96
|
-
total: hash[:total]
|
|
107
|
+
total: hash[:total],
|
|
108
|
+
unknown: hash[:unknown] || false
|
|
97
109
|
)
|
|
98
110
|
end
|
|
99
111
|
|
|
@@ -120,7 +132,7 @@ module TurnKit
|
|
|
120
132
|
tokens.to_i * price.to_f / PER_MILLION
|
|
121
133
|
end
|
|
122
134
|
|
|
123
|
-
def initialize(input: nil, output: nil, cache_read: nil, cache_write: nil, thinking: nil, total: nil, strict: false)
|
|
135
|
+
def initialize(input: nil, output: nil, cache_read: nil, cache_write: nil, thinking: nil, total: nil, strict: false, unknown: false)
|
|
124
136
|
@input = number(input)
|
|
125
137
|
@output = number(output)
|
|
126
138
|
@cache_read = number(cache_read)
|
|
@@ -128,9 +140,13 @@ module TurnKit
|
|
|
128
140
|
@thinking = number(thinking)
|
|
129
141
|
@total = number(total)
|
|
130
142
|
@strict = strict
|
|
143
|
+
@unknown = unknown
|
|
131
144
|
end
|
|
132
145
|
|
|
146
|
+
def unknown? = @unknown
|
|
147
|
+
|
|
133
148
|
def total
|
|
149
|
+
return nil if unknown?
|
|
134
150
|
return @total if @total
|
|
135
151
|
return nil if @strict && COMPONENTS.any? { |component| public_send(component).nil? }
|
|
136
152
|
|
|
@@ -145,7 +161,8 @@ module TurnKit
|
|
|
145
161
|
"cache_read" => cache_read,
|
|
146
162
|
"cache_write" => cache_write,
|
|
147
163
|
"thinking" => thinking,
|
|
148
|
-
"total" => total
|
|
164
|
+
"total" => total,
|
|
165
|
+
"unknown" => (true if unknown?)
|
|
149
166
|
}.compact
|
|
150
167
|
end
|
|
151
168
|
|
|
@@ -0,0 +1,119 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module TurnKit
|
|
4
|
+
# Deliberately separate from Result: evaluations have no message parts or text.
|
|
5
|
+
class EvaluationResult
|
|
6
|
+
attr_reader :model, :answers, :usage, :receipt_id
|
|
7
|
+
|
|
8
|
+
def initialize(model:, answers:, usage:, receipt_id: nil)
|
|
9
|
+
@model, @answers, @usage, @receipt_id = model, answers, usage, receipt_id
|
|
10
|
+
end
|
|
11
|
+
|
|
12
|
+
def to_h
|
|
13
|
+
{ "model" => model, "answers" => answers, "usage" => usage.to_h }
|
|
14
|
+
end
|
|
15
|
+
|
|
16
|
+
def self.from_h(value, receipt_id: nil)
|
|
17
|
+
new(model: value.fetch("model"), answers: value.fetch("answers"),
|
|
18
|
+
usage: Usage.from_h(value.fetch("usage")), receipt_id: receipt_id)
|
|
19
|
+
end
|
|
20
|
+
end
|
|
21
|
+
|
|
22
|
+
class EvaluationError < Error
|
|
23
|
+
attr_reader :status, :retry_after, :usage, :model, :http_status
|
|
24
|
+
|
|
25
|
+
def initialize(status, retryable: false, retry_after: nil, usage: nil, model: nil, http_status: nil)
|
|
26
|
+
@status, @retryable, @retry_after, @usage, @model = status.to_s, retryable, retry_after, usage, model
|
|
27
|
+
@http_status = http_status
|
|
28
|
+
# Never include provider bodies, URLs, state, question names, or credentials.
|
|
29
|
+
super("evaluation #{@status}")
|
|
30
|
+
end
|
|
31
|
+
|
|
32
|
+
def retryable? = @retryable
|
|
33
|
+
end
|
|
34
|
+
|
|
35
|
+
module Evaluation
|
|
36
|
+
module_function
|
|
37
|
+
|
|
38
|
+
# Canonical JSON also rejects Ruby objects which JSON.generate would stringify.
|
|
39
|
+
def json(value)
|
|
40
|
+
case value
|
|
41
|
+
when Hash
|
|
42
|
+
raise InputError, "evaluation keys must be strings or symbols" unless value.keys.all? { |k| k.is_a?(String) || k.is_a?(Symbol) }
|
|
43
|
+
raise InputError, "duplicate evaluation keys" unless value.keys.map(&:to_s).uniq.size == value.size
|
|
44
|
+
value.transform_keys(&:to_s).sort.to_h.transform_values { |v| json(v) }
|
|
45
|
+
when Array then value.map { |v| json(v) }
|
|
46
|
+
when String, Integer, TrueClass, FalseClass, NilClass then value
|
|
47
|
+
when Float
|
|
48
|
+
raise InputError, "evaluation numbers must be finite" unless value.finite?
|
|
49
|
+
value
|
|
50
|
+
else raise InputError, "evaluation values must be JSON"
|
|
51
|
+
end
|
|
52
|
+
end
|
|
53
|
+
|
|
54
|
+
def text?(value) = value.nil? || value.is_a?(String) || value.is_a?(Hash) || value.is_a?(Array)
|
|
55
|
+
|
|
56
|
+
def input!(state:, questions:)
|
|
57
|
+
input = json({ state: state, questions: questions })
|
|
58
|
+
valid = text?(input["state"]) && input["questions"].is_a?(Hash)
|
|
59
|
+
raise InputError, "invalid evaluation input" unless valid
|
|
60
|
+
input["questions"].each do |id, question|
|
|
61
|
+
valid = !id.empty? && question.is_a?(Hash) && (question.keys - %w[type instructions criteria]).empty? &&
|
|
62
|
+
question.key?("instructions") && text?(question["instructions"])
|
|
63
|
+
raise InputError, "invalid evaluation question" unless valid
|
|
64
|
+
criteria = question["criteria"]
|
|
65
|
+
valid = case question["type"]
|
|
66
|
+
when "noul"
|
|
67
|
+
criteria.nil? || (criteria.is_a?(Hash) && (criteria.keys - %w[true false]).empty? && criteria.values.all? { |v| text?(v) })
|
|
68
|
+
when "choice"
|
|
69
|
+
criteria.is_a?(Hash) && criteria.values.all? { |v| text?(v) }
|
|
70
|
+
when "score"
|
|
71
|
+
criteria.is_a?(Array) && criteria.size >= 2 && criteria.all? { |v| text?(v) }
|
|
72
|
+
else false
|
|
73
|
+
end
|
|
74
|
+
raise InputError, "invalid evaluation criteria" unless valid
|
|
75
|
+
end
|
|
76
|
+
input
|
|
77
|
+
end
|
|
78
|
+
|
|
79
|
+
def probability?(value) = value.is_a?(Numeric) && value.finite? && value.between?(0, 1)
|
|
80
|
+
|
|
81
|
+
def usage(value)
|
|
82
|
+
return unless value.is_a?(Hash) && value.keys.sort == %w[input_tokens output_tokens]
|
|
83
|
+
return unless value.values.all? { |v| v.is_a?(Integer) && v.between?(0, 9_007_199_254_740_991) }
|
|
84
|
+
Usage.new(input_tokens: value["input_tokens"], output_tokens: value["output_tokens"])
|
|
85
|
+
end
|
|
86
|
+
|
|
87
|
+
def result!(value, questions:)
|
|
88
|
+
observed = usage(value["usage"]) if value.is_a?(Hash)
|
|
89
|
+
model = value["model"] if value.is_a?(Hash) && value["model"].is_a?(String) && !value["model"].empty?
|
|
90
|
+
invalid = -> { raise EvaluationError.new(:malformed, usage: observed, model: model) }
|
|
91
|
+
invalid.call unless value.is_a?(Hash) && value.keys.sort == %w[answers model usage] && model && observed
|
|
92
|
+
answers = value["answers"]
|
|
93
|
+
invalid.call unless answers.is_a?(Hash) && answers.keys.sort == questions.keys.sort
|
|
94
|
+
answers.each do |id, answer|
|
|
95
|
+
question = questions.fetch(id)
|
|
96
|
+
type = question.fetch("type")
|
|
97
|
+
invalid.call unless answer.is_a?(Hash) && answer["type"] == type
|
|
98
|
+
if type == "noul"
|
|
99
|
+
invalid.call unless answer.keys.sort == %w[noul type] && probability?(answer["noul"])
|
|
100
|
+
next
|
|
101
|
+
end
|
|
102
|
+
keys = type == "choice" ? %w[choice confidence probabilities type] : %w[confidence legend probabilities score type]
|
|
103
|
+
invalid.call unless answer.keys.sort == keys && probability?(answer["confidence"])
|
|
104
|
+
expected = type == "choice" ? question.fetch("criteria").keys : question.fetch("criteria").each_index.map(&:to_s)
|
|
105
|
+
probabilities = answer["probabilities"]
|
|
106
|
+
invalid.call unless probabilities.is_a?(Hash) && probabilities.keys.sort == expected.sort &&
|
|
107
|
+
probabilities.values.all? { |p| probability?(p) } && (probabilities.values.sum - 1).abs <= 0.01
|
|
108
|
+
if type == "choice"
|
|
109
|
+
invalid.call unless expected.include?(answer["choice"])
|
|
110
|
+
else
|
|
111
|
+
score, legend = answer.values_at("score", "legend")
|
|
112
|
+
invalid.call unless score.is_a?(Numeric) && score.finite? && score.between?(0, expected.size - 1) &&
|
|
113
|
+
legend.is_a?(Hash) && legend.keys.sort == expected.sort && legend.values.all? { |v| v.is_a?(String) }
|
|
114
|
+
end
|
|
115
|
+
end
|
|
116
|
+
EvaluationResult.new(model: model, answers: answers, usage: observed)
|
|
117
|
+
end
|
|
118
|
+
end
|
|
119
|
+
end
|
|
@@ -0,0 +1,142 @@
|
|
|
1
|
+
# frozen_string_literal: true
|
|
2
|
+
|
|
3
|
+
module TurnKit
|
|
4
|
+
module InternalEvaluation
|
|
5
|
+
# Call from a tool or output-audit callable, never from a persistence callback.
|
|
6
|
+
# before_dispatch is application transport/spend policy, not model input.
|
|
7
|
+
def internal_evaluation(evaluator:, model:, purpose:, state:, questions:, policy_version:,
|
|
8
|
+
candidate: nil, max_attempts: 1, timeout: 30, before_dispatch: nil)
|
|
9
|
+
raise LostClaim, "evaluation requires an owned execution" unless store.is_a?(ExecutionStore)
|
|
10
|
+
raise ArgumentError, "max_attempts must be 1..3" unless max_attempts.is_a?(Integer) && (1..3).include?(max_attempts)
|
|
11
|
+
raise ArgumentError, "timeout must be positive and finite" unless timeout.is_a?(Numeric) && timeout.finite? && timeout.positive?
|
|
12
|
+
input = evaluator.validate!(model: model, state: state, questions: questions)
|
|
13
|
+
reload
|
|
14
|
+
identity = Evaluation.json(provider: evaluator.identity, model: model, purpose: purpose.to_s,
|
|
15
|
+
policy_version: policy_version.to_s, candidate: candidate, input: input, schema: 1,
|
|
16
|
+
output_candidate: @record.dig("options", "state", "candidate"),
|
|
17
|
+
output_data: @record.dig("options", "state", "output_data"),
|
|
18
|
+
timeout: timeout, max_attempts: max_attempts)
|
|
19
|
+
key = Digest::SHA256.hexdigest(JSON.generate(identity))
|
|
20
|
+
cost_model = evaluator.cost_model(model: model)
|
|
21
|
+
deadline = Clock.now + timeout
|
|
22
|
+
loop do
|
|
23
|
+
receipt = store.atomic do
|
|
24
|
+
reload
|
|
25
|
+
raise LostClaim, "evaluation requires an owned execution" unless store.is_a?(ExecutionStore) && running?
|
|
26
|
+
(evaluation_receipts[key] || { "id" => key, "model" => model,
|
|
27
|
+
"provider" => identity["provider"], "purpose" => purpose.to_s,
|
|
28
|
+
"policy_version" => policy_version.to_s, "cost_model" => cost_model, "attempts" => [] })
|
|
29
|
+
end
|
|
30
|
+
return EvaluationResult.from_h(receipt.fetch("result"), receipt_id: key) if receipt["status"] == "completed"
|
|
31
|
+
last = receipt["attempts"].last
|
|
32
|
+
if last && (!last["retryable"] || receipt["attempts"].size >= max_attempts)
|
|
33
|
+
raise EvaluationError.new(last.fetch("status"), http_status: last["http_status"])
|
|
34
|
+
end
|
|
35
|
+
if last && last["owner"] == @record["claim_token"] && !last["finished_at"]
|
|
36
|
+
raise EvaluationError.new(:in_progress)
|
|
37
|
+
end
|
|
38
|
+
delay = last ? [last.fetch("retry_at", 0) - Clock.now.to_f, 0].max : 0
|
|
39
|
+
remaining = check_evaluation_dispatch!(deadline)
|
|
40
|
+
raise BudgetError, "evaluation deadline exceeded" if delay >= remaining
|
|
41
|
+
sleep(delay) if delay.positive?
|
|
42
|
+
Authorization.authorize!(:evaluate, principal: @record.dig("options", "principal"),
|
|
43
|
+
turn: self, evaluator: evaluator, model: model, purpose: purpose)
|
|
44
|
+
before_dispatch.call(turn: self, model: model, purpose: purpose) if before_dispatch
|
|
45
|
+
attempt = nil
|
|
46
|
+
remaining = nil
|
|
47
|
+
store.atomic do
|
|
48
|
+
reload
|
|
49
|
+
remaining = check_evaluation_dispatch!(deadline)
|
|
50
|
+
current = evaluation_receipts[key]
|
|
51
|
+
# A second caller may have committed or dispatched while policy ran.
|
|
52
|
+
raise EvaluationError.new(:in_progress) if current && current != receipt
|
|
53
|
+
attempt = { "id" => SecureRandom.uuid, "owner" => @record["claim_token"],
|
|
54
|
+
"status" => "uncertain", "retryable" => true, "started_at" => Clock.now.iso8601(6),
|
|
55
|
+
"cost" => nil }
|
|
56
|
+
receipt["attempts"] << attempt
|
|
57
|
+
receipt["status"] = "uncertain"
|
|
58
|
+
save_evaluation_receipt!(key, receipt)
|
|
59
|
+
end
|
|
60
|
+
event = { receipt_id: key, attempt_id: attempt["id"], model: model, purpose: purpose.to_s,
|
|
61
|
+
question_count: input.fetch("questions").size }
|
|
62
|
+
emit("evaluation.requested", event)
|
|
63
|
+
result = nil
|
|
64
|
+
failure = nil
|
|
65
|
+
begin
|
|
66
|
+
result = evaluator.evaluate(model: model, state: input.fetch("state"),
|
|
67
|
+
questions: input.fetch("questions"), timeout: remaining)
|
|
68
|
+
rescue EvaluationError => error
|
|
69
|
+
failure = error
|
|
70
|
+
end
|
|
71
|
+
observed_usage = result ? result.usage : failure.usage
|
|
72
|
+
observed_cost = Cost.from_usage(observed_usage, model: cost_model) if observed_usage
|
|
73
|
+
store.atomic do
|
|
74
|
+
reload
|
|
75
|
+
attempt.merge!("status" => result ? "completed" : failure.status,
|
|
76
|
+
"finished_at" => Clock.now.iso8601(6), "retryable" => failure&.retryable? || false,
|
|
77
|
+
"http_status" => failure&.http_status,
|
|
78
|
+
"usage" => observed_usage&.to_h, "cost" => observed_cost&.total,
|
|
79
|
+
"returned_model" => result ? result.model : failure.model)
|
|
80
|
+
if failure&.retryable?
|
|
81
|
+
attempt["retry_at"] = Clock.now.to_f + (failure.retry_after || rand * (2 ** (receipt["attempts"].size - 1)))
|
|
82
|
+
end
|
|
83
|
+
receipt["status"] = attempt["status"]
|
|
84
|
+
receipt["result"] = result.to_h if result
|
|
85
|
+
add_usage!(observed_usage, cost: observed_cost) if observed_usage
|
|
86
|
+
save_evaluation_receipt!(key, receipt)
|
|
87
|
+
end
|
|
88
|
+
emit(result ? "evaluation.completed" : "evaluation.failed", event.merge(
|
|
89
|
+
status: attempt["status"], returned_model: attempt["returned_model"], http_status: attempt["http_status"],
|
|
90
|
+
usage: observed_usage&.to_h, cost: observed_cost&.to_h,
|
|
91
|
+
duration: Clock.now - Time.iso8601(attempt["started_at"])))
|
|
92
|
+
# Keep the shared inline budget in sync without turning a paid successful
|
|
93
|
+
# assessment into a retry just because its own cost exhausted the budget.
|
|
94
|
+
begin
|
|
95
|
+
budget.add_cost!(observed_cost&.total)
|
|
96
|
+
rescue BudgetError
|
|
97
|
+
# The next dispatch checks the persisted root ledger.
|
|
98
|
+
end
|
|
99
|
+
root_budget = execution_budget
|
|
100
|
+
root_budget.check!(depth: depth, allow_exhausted_spend: true)
|
|
101
|
+
raise BudgetError, "evaluation deadline exceeded" if Clock.now >= deadline
|
|
102
|
+
return EvaluationResult.from_h(receipt.fetch("result"), receipt_id: key) if result
|
|
103
|
+
raise failure unless failure.retryable? && receipt["attempts"].size < max_attempts
|
|
104
|
+
end
|
|
105
|
+
end
|
|
106
|
+
|
|
107
|
+
# Copies, not references to live state. Attempts without usage have unknown
|
|
108
|
+
# cost, even if the turn's known-cost subtotal is zero.
|
|
109
|
+
def evaluation_receipts
|
|
110
|
+
JSON.parse(JSON.generate(@record.dig("options", "state", "evaluations") || {}))
|
|
111
|
+
end
|
|
112
|
+
|
|
113
|
+
def output_metadata = JSON.parse(JSON.generate(@record.dig("options", "state", "output_metadata") || {}))
|
|
114
|
+
|
|
115
|
+
def output_metadata=(value)
|
|
116
|
+
value = Evaluation.json(value)
|
|
117
|
+
raise InputError, "output metadata must be an object" unless value.is_a?(Hash)
|
|
118
|
+
raise LostClaim, "output metadata requires an owned execution" unless store.is_a?(ExecutionStore)
|
|
119
|
+
store.atomic do
|
|
120
|
+
reload
|
|
121
|
+
raise LostClaim, "output metadata requires an owned execution" unless store.is_a?(ExecutionStore) && running?
|
|
122
|
+
update_state!("output_metadata" => value)
|
|
123
|
+
end
|
|
124
|
+
end
|
|
125
|
+
|
|
126
|
+
private
|
|
127
|
+
def save_evaluation_receipt!(key, receipt)
|
|
128
|
+
update_state!("evaluations" => evaluation_receipts.merge(key => receipt))
|
|
129
|
+
end
|
|
130
|
+
|
|
131
|
+
def check_evaluation_dispatch!(deadline)
|
|
132
|
+
current_budget = execution_budget
|
|
133
|
+
current_budget.check!(depth: depth)
|
|
134
|
+
raise LostClaim, "evaluation execution is no longer running" unless running?
|
|
135
|
+
raise EvaluationError.new(:interrupted) if @record.dig("options", "controls", "pause_requested")
|
|
136
|
+
root_deadline = current_budget.root_started_at + current_budget.timeout if current_budget.timeout
|
|
137
|
+
remaining = [(deadline - Clock.now), (root_deadline - Clock.now if root_deadline)].compact.min
|
|
138
|
+
raise BudgetError, "evaluation deadline exceeded" unless remaining.positive?
|
|
139
|
+
remaining
|
|
140
|
+
end
|
|
141
|
+
end
|
|
142
|
+
end
|
|
@@ -1,24 +1,49 @@
|
|
|
1
1
|
# frozen_string_literal: true
|
|
2
2
|
|
|
3
3
|
module TurnKit
|
|
4
|
+
# Runs a child agent in a fresh conversation and returns only its final result.
|
|
5
|
+
# `SubAgentTool.for(agent)` exposes an agent as a tool taking `task`. Subclasses
|
|
6
|
+
# declare their own parameters and override `task_for` to assemble the task in
|
|
7
|
+
# Ruby, so bulk data (file contents, records) reaches the child without ever
|
|
8
|
+
# entering the parent model's context. The child comes from the `agent` class
|
|
9
|
+
# macro or an instance-level `agent` override.
|
|
4
10
|
class SubAgentTool < Tool
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
tool_name agent.name
|
|
11
|
-
description agent.description.empty? ? "Delegate work to #{agent.name}." : agent.description
|
|
12
|
-
usage_hint "Use when work can be delegated independently to #{agent.name}. Pass a complete task and only relevant context."
|
|
11
|
+
class << self
|
|
12
|
+
def agent(value = nil)
|
|
13
|
+
@agent = value if value
|
|
14
|
+
@agent || (superclass < SubAgentTool ? superclass.agent : nil)
|
|
15
|
+
end
|
|
13
16
|
|
|
14
|
-
|
|
15
|
-
|
|
17
|
+
def for(agent)
|
|
18
|
+
sub_agent = agent
|
|
19
|
+
Class.new(self) do
|
|
20
|
+
agent sub_agent
|
|
21
|
+
tool_name sub_agent.name
|
|
22
|
+
description sub_agent.description.empty? ? "Delegate work to #{sub_agent.name}." : sub_agent.description
|
|
23
|
+
usage_hint "Use when work can be delegated independently to #{sub_agent.name}. Pass a complete task and only relevant context."
|
|
24
|
+
parameter :task, :string, required: true, description: "The complete task for the sub-agent, including all relevant context."
|
|
16
25
|
end
|
|
17
26
|
end
|
|
27
|
+
|
|
28
|
+
def delegates?(tool)
|
|
29
|
+
tool.is_a?(self) || (tool.is_a?(Class) && tool <= self)
|
|
30
|
+
end
|
|
31
|
+
|
|
32
|
+
def result(record)
|
|
33
|
+
{ "conversation_id" => record.fetch("conversation_id"), "turn_id" => record.fetch("id"),
|
|
34
|
+
"status" => record.fetch("status"), "result" => record["output_text"].to_s,
|
|
35
|
+
"output_metadata" => record.dig("options", "state", "output_metadata"),
|
|
36
|
+
"output_data" => record["output_data"], "error" => record["error"] }.compact
|
|
37
|
+
end
|
|
18
38
|
end
|
|
19
39
|
|
|
20
|
-
def self.
|
|
21
|
-
|
|
40
|
+
def agent = self.class.agent
|
|
41
|
+
|
|
42
|
+
def task_for(**arguments)
|
|
43
|
+
arguments.fetch(:task)
|
|
44
|
+
end
|
|
45
|
+
|
|
46
|
+
def build_child(task:, context:)
|
|
22
47
|
parent_turn = context.turn
|
|
23
48
|
lineage = {
|
|
24
49
|
"parent_conversation_id" => parent_turn.conversation.id,
|
|
@@ -27,32 +52,28 @@ module TurnKit
|
|
|
27
52
|
"principal" => context.principal
|
|
28
53
|
}
|
|
29
54
|
store = parent_turn.store
|
|
30
|
-
record = store.create_conversation("agent_name" =>
|
|
31
|
-
conversation = Conversation.new(agent:
|
|
55
|
+
record = store.create_conversation("agent_name" => agent.name, "model" => agent.effective_model, "metadata" => lineage)
|
|
56
|
+
conversation = Conversation.new(agent: agent, record: record, store: store, model: agent.effective_model, metadata: lineage)
|
|
32
57
|
trigger = conversation.say(task, metadata: lineage)
|
|
58
|
+
parent_turn.emit("sub_agent.delegated", id: context.execution.tool_call_id, name: agent.name,
|
|
59
|
+
conversation_id: record.fetch("id"), task_chars: task.length)
|
|
33
60
|
conversation.build_turn(
|
|
34
61
|
trigger_message_id: trigger.id,
|
|
35
62
|
budget: parent_turn.budget,
|
|
36
63
|
parent_turn: parent_turn,
|
|
37
64
|
parent_tool_execution: context.execution,
|
|
38
65
|
depth: parent_turn.depth + 1,
|
|
39
|
-
model:
|
|
40
|
-
agent:
|
|
66
|
+
model: agent.effective_model,
|
|
67
|
+
agent: agent,
|
|
41
68
|
principal: context.principal,
|
|
42
69
|
on_event: parent_turn.agent.effective_on_event
|
|
43
70
|
)
|
|
44
71
|
end
|
|
45
72
|
|
|
46
|
-
def
|
|
47
|
-
{ "conversation_id" => record.fetch("conversation_id"), "turn_id" => record.fetch("id"),
|
|
48
|
-
"status" => record.fetch("status"), "result" => record["output_text"].to_s,
|
|
49
|
-
"output_data" => record["output_data"], "error" => record["error"] }.compact
|
|
50
|
-
end
|
|
51
|
-
|
|
52
|
-
def call(task:, context:)
|
|
73
|
+
def call(context:, **arguments)
|
|
53
74
|
Authorization.authorize!(:launch_agent, principal: context.principal, turn: context.turn,
|
|
54
|
-
agent:
|
|
55
|
-
child =
|
|
75
|
+
agent: agent, arguments: arguments.transform_keys(&:to_s))
|
|
76
|
+
child = build_child(task: task_for(**arguments), context: context)
|
|
56
77
|
child.run!
|
|
57
78
|
SubAgentTool.result(child.store.load_turn(child.id))
|
|
58
79
|
end
|
data/lib/turnkit/tool_runner.rb
CHANGED
|
@@ -93,6 +93,8 @@ module TurnKit
|
|
|
93
93
|
context = ToolContext.new(turn: turn, execution: execution)
|
|
94
94
|
payload = begin
|
|
95
95
|
Authorization.authorize!(:tool, principal: context.principal, turn: turn, tool: tool, arguments: tool_call.arguments)
|
|
96
|
+
blocked = tool_policy_block(tool, tool_call.arguments, context)
|
|
97
|
+
return finish_error(execution, tool_call, blocked, details: { "tool_policy_blocked" => true }) if blocked
|
|
96
98
|
# Observe cancellation/reconciliation immediately before crossing the
|
|
97
99
|
# external-effect boundary. Calls already sent cannot be recalled.
|
|
98
100
|
control = turn.control_boundary!
|
|
@@ -191,20 +193,30 @@ module TurnKit
|
|
|
191
193
|
end
|
|
192
194
|
|
|
193
195
|
def subagent?(tool)
|
|
194
|
-
|
|
196
|
+
SubAgentTool.delegates?(tool)
|
|
197
|
+
end
|
|
198
|
+
|
|
199
|
+
# Agent-owned routing/cost policy, distinct from identity authorization.
|
|
200
|
+
# Returns the block reason the model sees, or nil to proceed.
|
|
201
|
+
def tool_policy_block(tool, arguments, context)
|
|
202
|
+
decision, reason = turn.agent.tool_policy&.call(tool: tool, arguments: arguments, context: context)
|
|
203
|
+
return reason.to_s if decision == :block
|
|
204
|
+
raise ArgumentError, "tool_policy must return :allow or [:block, reason]" unless [ nil, :allow ].include?(decision)
|
|
195
205
|
end
|
|
196
206
|
|
|
197
207
|
def delegate(tool, call, context)
|
|
198
|
-
|
|
208
|
+
tool = tool.new if tool.is_a?(Class)
|
|
209
|
+
arguments = tool.class.validate_arguments(call.arguments)
|
|
199
210
|
Authorization.authorize!(:launch_agent, principal: context.principal, turn: turn, agent: tool.agent, arguments: arguments)
|
|
200
211
|
TurnKit.resolve_agent(tool.agent.name)
|
|
212
|
+
task = tool.task_for(**arguments.transform_keys(&:to_sym))
|
|
201
213
|
child = turn.store.atomic_graph do
|
|
202
214
|
turn.store.atomic(Background.root_conversation(turn.store, turn.store.load_turn(turn.id))) do
|
|
203
215
|
control = turn.control_boundary!
|
|
204
216
|
next control if control
|
|
205
217
|
row = turn.store.list_turns(root_turn_id: turn.root_turn_id).find { |candidate| candidate["parent_tool_execution_id"] == context.execution.id }
|
|
206
218
|
unless row
|
|
207
|
-
built = tool.build_child(task:
|
|
219
|
+
built = tool.build_child(task: task, context: context)
|
|
208
220
|
row = turn.store.update_turn(built.id, submitted_at: Clock.now)
|
|
209
221
|
end
|
|
210
222
|
Background.wait(turn, [row.fetch("id")])
|
data/lib/turnkit/turn.rb
CHANGED
|
@@ -3,6 +3,7 @@
|
|
|
3
3
|
module TurnKit
|
|
4
4
|
class Turn
|
|
5
5
|
include TurnControls
|
|
6
|
+
include InternalEvaluation
|
|
6
7
|
STATUSES = Record::TURN_STATUSES
|
|
7
8
|
|
|
8
9
|
attr_reader :agent, :conversation, :store, :budget, :depth
|
|
@@ -325,6 +326,7 @@ module TurnKit
|
|
|
325
326
|
add_usage!(result.usage, cost: cost)
|
|
326
327
|
persist_assistant_message(result)
|
|
327
328
|
update_state!("phase" => result.tool_calls? ? "tools" : "output", "parts" => result.parts,
|
|
329
|
+
"output_metadata" => nil,
|
|
328
330
|
"candidate" => result.text, "output_data" => result.output_data, "terminal_tool_name" => nil,
|
|
329
331
|
"budget_completion_call_id" => select_budget_completion(result))
|
|
330
332
|
end
|
data/lib/turnkit/version.rb
CHANGED
data/lib/turnkit.rb
CHANGED
|
@@ -46,9 +46,12 @@ require_relative "turnkit/load_skill_tool"
|
|
|
46
46
|
require_relative "turnkit/message_projection"
|
|
47
47
|
require_relative "turnkit/tool_runner"
|
|
48
48
|
require_relative "turnkit/turn_controls"
|
|
49
|
+
require_relative "turnkit/evaluation"
|
|
50
|
+
require_relative "turnkit/internal_evaluation"
|
|
49
51
|
require_relative "turnkit/turn"
|
|
50
52
|
require_relative "turnkit/usage"
|
|
51
53
|
require_relative "turnkit/run"
|
|
54
|
+
require_relative "turnkit/adapters/cloudflare_jev"
|
|
52
55
|
require_relative "turnkit/adapters/codex"
|
|
53
56
|
require_relative "turnkit/adapters/ruby_llm"
|
|
54
57
|
require_relative "turnkit/active_record_store"
|
|
@@ -78,7 +81,8 @@ module TurnKit
|
|
|
78
81
|
|
|
79
82
|
def self.register(agent)
|
|
80
83
|
@agents[agent.name] = agent
|
|
81
|
-
agent.
|
|
84
|
+
tools = agent.effective_tools + agent.available_skills.flat_map(&:tools)
|
|
85
|
+
tools.each { |tool| register(tool.agent) if SubAgentTool.delegates?(tool) }
|
|
82
86
|
agent
|
|
83
87
|
end
|
|
84
88
|
|
metadata
CHANGED
|
@@ -1,14 +1,14 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: turnkit
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 0.
|
|
4
|
+
version: 0.9.0
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Sam Couch
|
|
8
8
|
autorequire:
|
|
9
9
|
bindir: bin
|
|
10
10
|
cert_chain: []
|
|
11
|
-
date: 2026-09-
|
|
11
|
+
date: 2026-09-19 00:00:00.000000000 Z
|
|
12
12
|
dependencies: []
|
|
13
13
|
description: TurnKit is a Ruby/Rails agent runtime for durable AI conversations, application
|
|
14
14
|
runs, orchestrator agents, tool calling, skills, sub-agents, context compaction,
|
|
@@ -37,6 +37,7 @@ files:
|
|
|
37
37
|
- lib/generators/turnkit/upgrade_generator.rb
|
|
38
38
|
- lib/turnkit.rb
|
|
39
39
|
- lib/turnkit/active_record_store.rb
|
|
40
|
+
- lib/turnkit/adapters/cloudflare_jev.rb
|
|
40
41
|
- lib/turnkit/adapters/codex.rb
|
|
41
42
|
- lib/turnkit/adapters/ruby_llm.rb
|
|
42
43
|
- lib/turnkit/agent.rb
|
|
@@ -50,11 +51,13 @@ files:
|
|
|
50
51
|
- lib/turnkit/coordination_tools.rb
|
|
51
52
|
- lib/turnkit/cost.rb
|
|
52
53
|
- lib/turnkit/error.rb
|
|
54
|
+
- lib/turnkit/evaluation.rb
|
|
53
55
|
- lib/turnkit/event.rb
|
|
54
56
|
- lib/turnkit/execution_store.rb
|
|
55
57
|
- lib/turnkit/id.rb
|
|
56
58
|
- lib/turnkit/image_result.rb
|
|
57
59
|
- lib/turnkit/image_tool.rb
|
|
60
|
+
- lib/turnkit/internal_evaluation.rb
|
|
58
61
|
- lib/turnkit/job.rb
|
|
59
62
|
- lib/turnkit/load_skill_tool.rb
|
|
60
63
|
- lib/turnkit/media_analysis_result.rb
|