agent-harness 0.43.0 → 0.44.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 40d84099230c75237b23dca08330a46b739029768ce345b57cb232ba0e66a935
4
- data.tar.gz: 47d36a8e160bcd29f9d50f6eabc1d0ded6795d6e05f2e45e93038d1987cd352c
3
+ metadata.gz: 4692f9efdcfe0158cba3afcff4b077fdcc67dc3c9f28198569be1e338add3b74
4
+ data.tar.gz: 3911441fb063e7b52ed74f3fad737d2ef9189bb122897a44e65c1ebf449fcaa8
5
5
  SHA512:
6
- metadata.gz: f059d918fed14691fe444dba2e5e93aaeb2cdb185785a580b8a7341db8198616a0a830fd43461b4dfceffc9258f71ce25a8cbb69351747af46980ba0bd783757
7
- data.tar.gz: 2f09db993da48a0ca0ed852eee9653898f474e7c61a6e0603dde3588457d226c8f51de71a8b322e7eba46b7e64203cc61a7607aa31764ed63188ef8b38f7c1e5
6
+ metadata.gz: b933a56fb9905d8e007903ff05245d29885b574bd78db06621d135da1e605f8141ad13e32c616614f0183b4757d88f314aa4750737755d73f587af89955ecf20
7
+ data.tar.gz: 43e7cc1a88f2984ad163b788dd56a798429a4cecd102e02c1c7b11dc5247bcbd56811cdf6aa2a2eaca87da59d0f679ecb27b64c99e3b7fa972d021962467501a
@@ -1,3 +1,3 @@
1
1
  {
2
- ".": "0.43.0"
2
+ ".": "0.44.1"
3
3
  }
data/CHANGELOG.md CHANGED
@@ -2,10 +2,25 @@
2
2
 
3
3
  ### Features
4
4
 
5
+ * verify RDR-072's published capability artifacts and migration limits, retain the application chat loop over normalized transport, and remove the private RubyLLM streaming-refusal patch in favor of an explicit non-streaming schema boundary ([#437](https://github.com/viamin/agent-harness/issues/437)).
5
6
  * expose normalized API chat transport and schema-constrained responses with preserved JSON text, locally validated parsed values, and explicit refusal, truncation, invalid JSON, schema mismatch, and unsupported-mode outcomes; this combined delivery supersedes the separate chat-transport work item ([#433](https://github.com/viamin/agent-harness/issues/433), [#434](https://github.com/viamin/agent-harness/issues/434)).
6
7
  * add runner model compatibility contract (`AgentHarness.model_compatibility`) with structured `ModelCompatibility::Result` outcomes. Codex exposes static facts for CLI-gated models (e.g. `gpt-5.5` requires Codex CLI `>= 0.116.0`), a baseline supported-model list, supported auth modes, and a `DEFAULT_COMPATIBLE_MODEL_ID` fallback so downstream orchestrators can validate tier/model assignments before scheduling agent runs ([#259](https://github.com/viamin/agent-harness/issues/259)).
7
8
  * **auth:** add provider-owned PKCE code-exchange API for Claude OAuth (`AgentHarness::Authentication.exchange_code`). Takes an authorization code plus PKCE verifier (and `redirect_uri`/`client_id`), posts an `authorization_code` grant to the Claude token endpoint, and persists the resulting access/refresh tokens in the native `claudeAiOauth` shape. Adds `exchange_code_supported?` and a `code_exchange` key to `auth_capabilities` ([#266](https://github.com/viamin/agent-harness/issues/266)).
8
9
 
10
+ ## [0.44.1](https://github.com/viamin/agent-harness/compare/agent-harness/v0.44.0...agent-harness/v0.44.1) (2026-09-27)
11
+
12
+
13
+ ### Bug Fixes
14
+
15
+ * **deps-dev:** bump @anthropic-ai/claude-code from 2.1.275 to 2.1.282 in /vendor/pins/claude ([#453](https://github.com/viamin/agent-harness/issues/453)) ([1848bbb](https://github.com/viamin/agent-harness/commit/1848bbb03653e0b64524fbffbcc939eb8542ad2f))
16
+
17
+ ## [0.44.0](https://github.com/viamin/agent-harness/compare/agent-harness/v0.43.0...agent-harness/v0.44.0) (2026-09-25)
18
+
19
+
20
+ ### Features
21
+
22
+ * Evaluate resumable loop delegation and implement only if simpler (RDR-072) ([#448](https://github.com/viamin/agent-harness/issues/448)) ([02dedfb](https://github.com/viamin/agent-harness/commit/02dedfb08096631e59f326333df61feb57eba395))
23
+
9
24
  ## [0.43.0](https://github.com/viamin/agent-harness/compare/agent-harness/v0.42.0...agent-harness/v0.43.0) (2026-09-25)
10
25
 
11
26
 
data/README.md CHANGED
@@ -806,8 +806,9 @@ Health checks run five steps per provider: registration, CLI availability, authe
806
806
  The proposed provider-neutral API execution boundary for embeddings, chat,
807
807
  structured output, usage, and optional persistence is documented in the
808
808
  [Provider-Neutral API Execution Contract](docs/provider-neutral-api-execution-contract.md).
809
- It is a docs-only design boundary; capability support requires the release
810
- evidence described there.
809
+ That document records the shipped capability versions, migration examples,
810
+ verified limits, and the evidence-backed decision to retain the application
811
+ conversation loop over the normalized transport.
811
812
 
812
813
  ```bash
813
814
  # Install dependencies
@@ -9,9 +9,12 @@ usage, and optional conversation-persistence work.
9
9
 
10
10
  RDR-072's rollout guard was **docs-only** for the design phase. The normalized
11
11
  chat, attempt-accounting, schema, and embedding capabilities described below
12
- are now implemented. Other capabilities remain design contracts and each still needs
13
- its own failing-first contract tests, implementation, release evidence, and
14
- downstream adoption evidence before a caller enables it.
12
+ are now implemented. The resumable-loop delegation evaluation is also
13
+ complete: it closed with a **retained-loop outcome** (see "Resumable-loop
14
+ delegation evaluation"), so no loop runtime was added. Other capabilities
15
+ remain design contracts and each still needs its own failing-first contract
16
+ tests, implementation, release evidence, and downstream adoption evidence
17
+ before a caller enables it.
15
18
 
16
19
  Existing CLI and subscription behavior remains the default. Existing
17
20
  `TextTransport`, `OpenAICompatibleTransport`, `Conversation`, and `Response`
@@ -553,6 +556,7 @@ records.
553
556
  | Attempt usage | `usage.ruby_llm` per physical attempt | Useful facts, but the public payload has no stable attempt ID; harness must add one |
554
557
  | Plain Ruby resume | transcript can be reconstructed manually | No documented state export/import API; implement normalized export/import outside RubyLLM |
555
558
  | Rails resume | `acts_as_chat` transcript plus supporting records | Technically restart-safe at checkpoints; at-least-once side effects remain |
559
+ | Loop controls | `Chat#step`, `#complete`, `#run_tools`, `#approve`, `#deny`, `Tool.requires_approval`, `#cancel` | Evaluated for delegation; retained-loop outcome (see below) |
556
560
 
557
561
  ### Optional Rails supporting tables
558
562
 
@@ -591,14 +595,96 @@ historical/pending-conversation tests. Reverting the gem is not a data rollback.
591
595
  2. **Use RubyLLM tool-call and usage tables selectively.** Potentially removes
592
596
  bookkeeping, but only after stable-attempt mapping, tenant-scoped access,
593
597
  audit, and migration tests are complete.
594
- 3. **Delegate the full loop and all supporting tables.** Not recommended now.
595
- It does not yet demonstrate reduced maintenance across both repositories and
596
- increases migration and recovery coupling.
598
+ 3. **Delegate the full loop and all supporting tables.** Evaluated and
599
+ rejected for now: the delegation review below found maintenance increases
600
+ across both repositories. Retaining Paid's loop over the normalized
601
+ transport is the completed outcome and creates no future delegation
602
+ obligation.
597
603
 
598
604
  The state investigation is therefore positive for checkpoint-based Rails
599
605
  restoration and normalized plain Ruby reconstruction, and negative for
600
606
  exactly-once recovery or an off-the-shelf plain Ruby export/import mechanism.
601
607
 
608
+ ## Resumable-loop delegation evaluation
609
+
610
+ RDR-072 permits delegating the chat loop only when behavior is preserved and
611
+ maintenance decreases across both repositories, counting adapters,
612
+ persistence, and recovery code. Retaining Paid's loop over the normalized
613
+ transport is an acceptable completed outcome. This section records the
614
+ evaluation against RubyLLM 2.0.0's public loop controls
615
+ (`Chat#step`, `#complete`, `#run_tools`, `#complete?`, `#awaiting_approval?`,
616
+ `#pending_approvals`, `#approve`, `#deny`, `#cancel`, and
617
+ `Tool.requires_approval`) and its conclusion.
618
+
619
+ ### Verified behaviors
620
+
621
+ `spec/ruby_llm_loop_delegation_evaluation_spec.rb` drives a real
622
+ `RubyLLM::Chat` through the public API with stubbed Anthropic Messages
623
+ responses and pins each fact with a contract test:
624
+
625
+ | Required behavior | Result | Evidence |
626
+ | --- | --- | --- |
627
+ | Single-step execution | Verified | `#step` advances one move; `#complete?` flips only on a final answer |
628
+ | Mixed read/write batches | Verified | Reads without approval execute; `requires_approval` writes stay pending (`#awaiting_approval?`) |
629
+ | Multiple pending decisions | Verified | `#pending_approvals` lists every undecided write; `#approve`/`#deny` resolve them independently |
630
+ | Denial | Verified | A denied call receives a structured denial result and the model continues |
631
+ | Completion | Verified | `#complete?` reports the terminal state |
632
+ | Cancellation | Verified | `#cancel` raises `CancelledError` at the next loop checkpoint, then clears |
633
+ | Completed tool results preserved | Verified | `#run_tools` skips calls that already carry results, so a resumed round executes only the remainder |
634
+
635
+ ### Gaps that fail the delegation criteria
636
+
637
+ | Criterion | Finding | Evidence |
638
+ | --- | --- | --- |
639
+ | Stable tool IDs | `ToolCall#id` is the provider wire id verbatim; no caller-stable identity is generated, so the contract's "provider_id is not a durable application identifier" rule needs a new mapping adapter | provider-id example |
640
+ | Iteration limits | The loop has no bound: `#complete` steps until `#complete?` or `#awaiting_approval?`, and `tool_options` exposes only choice/calls/concurrency; bounding stays caller-owned | unbounded-iteration example |
641
+ | Restart restoration without Rails | Decisions recorded with `#approve` live in per-chat in-memory state; a plain Ruby chat rebuilt from the same messages is awaiting approval again. Durable decisions require `acts_as_chat` (Active Record), which plain Ruby consumers must not require | reconstruction example |
642
+ | Crash recovery | Same as above: checkpoint restoration is Rails-only, so a plain Ruby consumer re-implements decision and transcript persistence | reconstruction example |
643
+ | Runner-isolated tool execution | `Tool#execute` runs in the calling process; the loop has no dispatch boundary. Paid executes tools in runner processes under caller-owned authorization and atomic claims, so delegation needs new dispatch adapters rather than removing code | in-process execution throughout |
644
+ | Bounded, non-nested retries | Each loop `#generate` retries beneath the caller through Faraday middleware by default (`max_retries` defaults to three), violating the single-retry-owner and shared `max_attempts` budget rules; even with retries disabled, per-step budget/fallback sequencing is absent | hidden-retry example |
645
+ | Attempt accounting | `usage.ruby_llm` payloads carry only operation/provider/model/status/tokens/cost — no attempt or request identity — so the harness `AttemptReport` ledger with stable `attempt_id` would be bypassed | usage-payload example |
646
+
647
+ ### Maintenance comparison
648
+
649
+ Delegating the loop would remove Paid's turn-taking mechanics (step dispatch
650
+ and approve/deny plumbing) while keeping Paid's authorization, atomic claims,
651
+ runner dispatch, iteration bounds, workflow recovery, durable accounting, and
652
+ cross-process cancellation. It would add, across the two repositories:
653
+
654
+ - a provider-id to stable-id mapping adapter;
655
+ - in-process `Tool#execute` to runner-process dispatch adapters;
656
+ - a per-step retry, shared-budget, and fallback sequencing wrapper;
657
+ - a caller-side iteration-limit driver replacing `#complete`;
658
+ - plain Ruby decision/transcript persistence and export/import (Rails
659
+ `acts_as_chat` cannot be a dependency); and
660
+ - attempt-accounting extraction with harness-side attempt identity.
661
+
662
+ The additions exceed the removals, so maintenance increases across both
663
+ repositories and the criteria fail.
664
+
665
+ ### Outcome
666
+
667
+ **Retained loop.** Paid keeps its conversation loop over
668
+ `Api::ChatTransport`; the harness keeps one normalized response per call
669
+ with bounded retries, fallback, cancellation, and the attempt ledger. Tool
670
+ execution, approval decisions, authorization, atomic claims, iteration
671
+ limits, and crash recovery remain caller responsibilities under the existing
672
+ ownership boundary. No loop runtime, adapter, persistence, or release was
673
+ added, so there is no migration cost and no consumer activation step. This
674
+ closes the evaluation without creating a future delegation obligation.
675
+
676
+ Re-evaluate only if RubyLLM later documents a public, plain-Ruby-resumable
677
+ loop with caller-stable tool identity, an external tool-dispatch boundary, a
678
+ bounded `#complete`, and per-attempt accounting identity; the evaluation
679
+ spec's examples are the tripwire that detects such changes on upgrade.
680
+
681
+ ### Communication
682
+
683
+ The retained-loop outcome, the behavior-test evidence above, and the zero
684
+ migration cost are the message for the Paid adoption and closeout tracking
685
+ (viamin/paid#4014 and the parent viamin/agent-harness#430). No scope beyond
686
+ the accepted RDR-072 alternatives is requested or implied.
687
+
602
688
  ## Current harness gaps and incremental delivery
603
689
 
604
690
  The current transports already normalize basic text, tool calls, token totals,
@@ -615,7 +701,8 @@ Ship capabilities independently in this order:
615
701
  3. normalized non-streaming and streaming chat transport;
616
702
  4. structured output for verified model/protocol combinations;
617
703
  5. plain Ruby state round trips; and
618
- 6. optional Rails persistence evaluation, then loop evaluation.
704
+ 6. optional Rails persistence evaluation, then loop evaluation (closed with
705
+ the retained-loop outcome recorded above; no runtime change).
619
706
 
620
707
  Each implementation issue starts with failing contract tests for request-local
621
708
  credential isolation, custom endpoints/headers, unsupported outcomes,
@@ -625,40 +712,117 @@ types stay behind the harness boundary.
625
712
 
626
713
  ## Compatibility and release evidence
627
714
 
628
- ### Attempt-accounting capability evidence
629
-
630
- - Release: pending the first published version containing issue #435; a Git
631
- branch or tag alone is not downstream adoption evidence.
632
- - Scope: normalized chat through `Api::ChatTransport` for Anthropic Messages
633
- and OpenAI Responses/Chat Completions (including compatible endpoints), with
634
- request-local API-key authentication. It is stacked on the normalized chat
635
- capability from #433.
636
- - Verification: the API contract specs cover request-local credentials,
637
- endpoint/header isolation, error classification, bounded non-nested retry,
638
- fallback, cancellation, partial usage, cache usage, repeated identity,
639
- observer redaction, and JSON reload with preserved pricing. The full upstream
640
- suite and lint run on the repository's supported Ruby environment.
641
- - Migration: no Rails tables or migrations are loaded or required. Paid and
642
- agent-image consumer versions remain unverified and MUST NOT adopt this
643
- capability until their integration suites record the exact released gem and
644
- image versions.
645
- - Retained paths: CLI/subscription providers, legacy HTTP transports,
646
- embeddings, and token trackers remain outside this ledger and require
647
- separate migration issues.
648
-
649
- ### Schema capability release evidence
650
-
651
- - Publication: unreleased; record the first installable version before
652
- downstream adoption.
653
- - Verified scopes: `:schema` with Anthropic Messages, OpenAI Responses, and
654
- OpenAI Chat Completions using API-key authentication, including compatible
655
- endpoints that explicitly select Chat Completions.
656
- - Contract coverage: valid and required-field schemas, classified failures,
657
- bounded retries, cancellation, refusal, truncation, malformed JSON, and
658
- schema mismatch.
659
- - Retained paths: CLI and subscription execution remain on existing provider
660
- interfaces. JSON-only mode and model-specific capability discovery are not
661
- migrated.
715
+ The following RubyGems releases are the first installable artifacts for each
716
+ capability. Pin the capability's minimum version during migration; do not infer
717
+ support from an issue, branch, or Git tag.
718
+
719
+ | Capability | First installable version | Contract / delivery | Install check |
720
+ | --- | ---: | ---: | --- |
721
+ | Native embeddings | `0.40.0` | [#432](https://github.com/viamin/agent-harness/issues/432) / [#439](https://github.com/viamin/agent-harness/issues/439) | `gem install agent-harness -v 0.40.0` |
722
+ | Normalized chat, tools, and streaming | `0.41.0` | [#434](https://github.com/viamin/agent-harness/issues/434) / [#441](https://github.com/viamin/agent-harness/issues/441) | `gem install agent-harness -v 0.41.0` |
723
+ | Attempt usage and cost | `0.42.0` | [#435](https://github.com/viamin/agent-harness/issues/435) / [#443](https://github.com/viamin/agent-harness/issues/443) | `gem install agent-harness -v 0.42.0` |
724
+ | Schema-constrained responses | `0.43.0` | [#436](https://github.com/viamin/agent-harness/issues/436) / [#442](https://github.com/viamin/agent-harness/issues/442) | `gem install agent-harness -v 0.43.0` |
725
+
726
+ All four are plain-Ruby capabilities. They load without Rails or Active Record,
727
+ create no tables, and require no migration. The runtime floor is Ruby 3.2 and
728
+ RubyLLM 2.x. The four released versions above still pin `ruby_llm = 2.0.0` and
729
+ prepend the private `AgentHarness::Api::RubyLlmResponsesStreamingRefusal`
730
+ module into `RubyLLM::Protocols::Responses` process-wide. The prepend-free,
731
+ public-API-only boundary (RubyLLM's public context, chat, message, tool, token
732
+ and cost APIs only; no private method calls, no downstream monkey patch, and a
733
+ loosened `ruby_llm ~> 2.0` dependency) lands in the first release published
734
+ after [#437](https://github.com/viamin/agent-harness/issues/437); record that
735
+ version here once it exists, because support must not be inferred from an
736
+ issue, branch, or Git tag.
737
+
738
+ ### Migration examples and limits
739
+
740
+ - **Embeddings (`>= 0.40.0`):** replace a downstream direct/proxy HTTP patch
741
+ with `AgentHarness.embed`, passing credentials, endpoint and headers on each
742
+ request. Remove the old retry layer after direct and proxy contract tests
743
+ pass. The operation is OpenAI-compatible, API-key only, has batch-total usage
744
+ only, and does not estimate missing tokens.
745
+ - **Chat (`>= 0.41.0`):** construct `AgentHarness::Api::ChatTransport` and send
746
+ one normalized request for each model step. Keep the application loop and
747
+ append completed tool results to the next request. Verified protocols are
748
+ Anthropic Messages and OpenAI Responses/Chat Completions; custom compatible
749
+ endpoints must explicitly choose Chat Completions. Media, custom auth modes,
750
+ protocol probing, and request-local connect timeouts are unsupported.
751
+ - **Attempt accounting (`>= 0.42.0`):** persist each entry in `result[:attempts]`
752
+ once by `attempt_id`, including failed and cancelled attempts. Do not add
753
+ aggregate `result[:usage]` again, convert unknown usage to zero, or wrap the
754
+ call in another provider-request retry loop. Paid still owns durable budget
755
+ accounting, workflow recovery and runner changes.
756
+ - **Schema (`>= 0.43.0`):** set `operation: :schema`, `schema_mode:
757
+ :json_schema`, and provide a named JSON Schema. The harness preserves the
758
+ provider text and validates it locally; use a non-streaming schema request.
759
+ JSON-only mode, schema repair, and model-specific capability discovery are
760
+ also unsupported. Rejection of streamed schema requests
761
+ (`unsupported/structured_output_not_supported`) and removal of the private
762
+ RubyLLM streaming-refusal shim land in the first release published after
763
+ this change; the released `0.43.0` still contains that shim and executes
764
+ streamed schema requests.
765
+
766
+ For every migration, run the consumer contract suite with tenant-specific
767
+ credentials and headers, custom endpoints, auth and transient failures,
768
+ cancellation, and the configured attempt bound. Inspect captured logs for
769
+ credentials and request bodies. Record the exact Paid host and agent-image
770
+ artifact versions before enabling that scope.
771
+
772
+ ### Verification and retained paths
773
+
774
+ The [#431 contract audit](https://github.com/viamin/agent-harness/issues/431)
775
+ is implemented by the upstream specs under `spec/agent_harness/embeddings_spec.rb`
776
+ and `spec/agent_harness/api/`. Those specs cover request-local credential, endpoint and header
777
+ isolation; reserved authentication headers; classified errors; bounded retries
778
+ with RubyLLM retries disabled; cancellation; partial streams; stable attempt
779
+ identity; usage/cost serialization; secret-free observer data; schema refusal,
780
+ truncation and validation; and the public-only RubyLLM integration. The normal
781
+ provider, command-executor, model-discovery and authentication suites remain the
782
+ CLI/subscription regression coverage.
783
+
784
+ CLI and subscription execution, legacy `TextTransport` and
785
+ `OpenAICompatibleTransport`, provider selection, workflow recovery, and Paid's
786
+ durable accounting are intentionally unchanged. No credential or authentication
787
+ mode is selected implicitly. Rails persistence and conversation state round
788
+ trips were not adopted.
789
+
790
+ ### Retained-loop closeout
791
+
792
+ The completed boundary is normalized transport with Paid's loop retained.
793
+ `Api::ChatTransport` performs one assistant response and never executes tools;
794
+ Paid continues to own authorization, confirmation policy, tool side effects,
795
+ transcript persistence, approval resumption and crash recovery. Delegating the
796
+ loop would require new persistence, tenant/audit adapters and recovery glue in
797
+ both repositories without removing those application responsibilities. There
798
+ is therefore no demonstrated net maintenance reduction or safe migration of
799
+ pending conversations. Under RDR-072 this evidence-backed retained-loop result
800
+ is complete, not deferred delegation.
801
+
802
+ Downstream adoption is tracked in
803
+ [viamin/paid#4014](https://github.com/viamin/paid/issues/4014); preservation of
804
+ Codex subscription discovery/recovery during dependency adoption is tracked in
805
+ [viamin/paid#3995](https://github.com/viamin/paid/issues/3995). Close parent
806
+ [agent-harness#430](https://github.com/viamin/agent-harness/issues/430) only
807
+ after those consumers record their exact host/image artifacts and passing
808
+ integration evidence for every scope they enable.
809
+
810
+ ### Loop-delegation evaluation evidence (retained loop)
811
+
812
+ - Publication: none required. The evaluation changed documentation and tests
813
+ only; no runtime path, adapter, table, or public API was added or altered,
814
+ so every released agent-harness version already carries the outcome.
815
+ - Verification: `spec/ruby_llm_loop_delegation_evaluation_spec.rb` exercises
816
+ only the public RubyLLM 2.0 API (pinned `ruby_llm = 2.0.0`) against stubbed
817
+ Anthropic Messages responses and runs with the upstream suite and lint on
818
+ the repository's supported Ruby environment. It records the verified loop
819
+ behaviors and the gaps that failed the delegation criteria.
820
+ - Migration: none. Paid and other consumers keep their existing loops over
821
+ `Api::ChatTransport`; no consumer activation, feature flag, or rollout step
822
+ applies.
823
+ - Retained responsibilities: documented in "Resumable-loop delegation
824
+ evaluation" above; loop ownership stays with Paid under RDR-072's accepted
825
+ alternatives.
662
826
 
663
827
  AgentHarness currently supports Ruby 3.2 and later and must remain usable as a
664
828
  plain Ruby gem. RubyLLM 2.0.0 itself supports Ruby 3.1.3 and later and adds
@@ -683,6 +847,6 @@ For every capability, release evidence must name:
683
847
  - retained paths and follow-up issues for combinations not migrated.
684
848
 
685
849
  Paid issue `viamin/paid#4014` should receive this compatibility result, the
686
- state-restoration conclusion, the stable-ID gap, and the per-capability release
687
- evidence. Downstream adoption cannot proceed from this design issue closing or
688
- from a Git tag alone.
850
+ state-restoration conclusion, the stable-ID gap, the per-capability release
851
+ evidence, and the retained-loop outcome of the delegation evaluation.
852
+ Downstream adoption cannot proceed from this closeout or from a Git tag alone.
@@ -373,7 +373,9 @@ module AgentHarness
373
373
  end
374
374
 
375
375
  def schema_mode_supported?
376
- !schema_operation? || !request[:schema_mode] || request[:schema_mode].to_sym == :json_schema
376
+ return true unless schema_operation?
377
+
378
+ request[:stream] != true && (!request[:schema_mode] || request[:schema_mode].to_sym == :json_schema)
377
379
  end
378
380
 
379
381
  def schema_payload
@@ -5,22 +5,6 @@ require "ruby_llm"
5
5
 
6
6
  module AgentHarness
7
7
  module Api
8
- # RubyLLM 2.0.0 flattens Responses API refusal deltas into ordinary text
9
- # chunks and offers no public access to their event type. The gemspec pins
10
- # that exact release so this compatibility shim cannot silently outlive the
11
- # private parser shape it targets. Remove it when RubyLLM exposes refusals.
12
- module RubyLlmResponsesStreamingRefusal
13
- module RefusalChunk; end
14
-
15
- def build_chunk(data)
16
- super.tap do |chunk|
17
- chunk.extend(RefusalChunk) if data["type"] == "response.refusal.delta"
18
- end
19
- end
20
- end
21
-
22
- RubyLLM::Protocols::Responses.prepend(RubyLlmResponsesStreamingRefusal)
23
-
24
8
  # Translates the normalized public chat values to RubyLLM public objects.
25
9
  class RubyLlmChatAdapter
26
10
  class UnsupportedOptionError < StandardError; end
@@ -38,12 +22,12 @@ module AgentHarness
38
22
  provider_usage_reported = false
39
23
  chat = prepared_chat || prepare(candidate:, messages:, tools:, max_output_tokens:, temperature:, timeout:,
40
24
  schema:, on_accounting:, provider_usage: -> { provider_usage_reported })
41
- response, streamed_refusal = generate(chat, stream, cancellation) do |event|
25
+ response = generate(chat, stream, cancellation) do |event|
42
26
  provider_usage_reported = true if event[:type] == :usage_updated
43
27
  on_event&.call(event)
44
28
  end
45
29
  emit_completed_tool_calls(response, &on_event) if stream
46
- normalize_response(response, streamed_refusal: streamed_refusal)
30
+ normalize_response(response)
47
31
  end
48
32
 
49
33
  def prepare(candidate:, messages:, tools:, max_output_tokens:, temperature:, timeout:, on_accounting: nil,
@@ -61,7 +45,6 @@ module AgentHarness
61
45
  # duplicate cumulative reports are not re-emitted.
62
46
  class StreamState
63
47
  attr_accessor :input_tokens, :output_tokens
64
- attr_reader :refusal
65
48
 
66
49
  def initialize
67
50
  @provider_id_by_key = {}
@@ -69,7 +52,6 @@ module AgentHarness
69
52
  @latest_provider_id = nil
70
53
  @input_tokens = nil
71
54
  @output_tokens = nil
72
- @refusal = false
73
55
  end
74
56
 
75
57
  # Links a stream chunk key to its provider call id, returning true
@@ -90,10 +72,6 @@ module AgentHarness
90
72
  total = (input_tokens + output_tokens) if input_tokens && output_tokens
91
73
  {input_tokens: input_tokens, output_tokens: output_tokens, total_tokens: total}
92
74
  end
93
-
94
- def observe(chunk)
95
- @refusal ||= chunk.is_a?(RubyLlmResponsesStreamingRefusal::RefusalChunk)
96
- end
97
75
  end
98
76
 
99
77
  private
@@ -182,10 +160,10 @@ module AgentHarness
182
160
  end
183
161
 
184
162
  def generate(chat, stream, cancellation)
185
- return [generate_without_events(chat, cancellation), false] unless stream
163
+ return generate_without_events(chat, cancellation) unless stream
186
164
 
187
165
  state = StreamState.new
188
- response = chat.generate do |chunk|
166
+ chat.generate do |chunk|
189
167
  if cancelled?(cancellation)
190
168
  chat.cancel
191
169
  raise RubyLLM::CancelledError
@@ -193,7 +171,6 @@ module AgentHarness
193
171
 
194
172
  stream_events(chunk, state).each { |event| yield event }
195
173
  end
196
- [response, state.refusal]
197
174
  end
198
175
 
199
176
  def generate_without_events(chat, cancellation)
@@ -259,7 +236,6 @@ module AgentHarness
259
236
  end
260
237
 
261
238
  def stream_events(chunk, state)
262
- state.observe(chunk)
263
239
  [text_event(chunk), *tool_call_events(chunk, state), usage_event(chunk, state)].compact
264
240
  end
265
241
 
@@ -325,14 +301,14 @@ module AgentHarness
325
301
  end
326
302
  end
327
303
 
328
- def normalize_response(response, streamed_refusal: false)
304
+ def normalize_response(response)
329
305
  {
330
306
  content: response.content || "",
331
307
  model: response.model,
332
308
  finish_reason: response.finish_reason,
333
309
  usage: normalize_usage(response.tokens),
334
310
  tool_calls: Array(response.tool_calls&.values).map { |call| normalize_tool_call(call) },
335
- refusal: streamed_refusal || refusal?(response)
311
+ refusal: refusal?(response)
336
312
  }
337
313
  end
338
314
 
@@ -27,7 +27,7 @@ module AgentHarness
27
27
 
28
28
  # Model name pattern for Anthropic Claude models
29
29
  MODEL_PATTERN = /^claude-[\d.-]+-(?:opus|sonnet|haiku)(?:-\d{8})?$/i
30
- SUPPORTED_CLI_VERSION = "2.1.275"
30
+ SUPPORTED_CLI_VERSION = "2.1.282"
31
31
  SUPPORTED_CLI_REQUIREMENT = Gem::Requirement.new(">= #{SUPPORTED_CLI_VERSION}", "< 2.2.0").freeze
32
32
 
33
33
  # Matches semver (e.g. "2.1.92"), optional pre-release (e.g. "2.1.92-beta.1"),
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module AgentHarness
4
- VERSION = "0.43.0"
4
+ VERSION = "0.44.1"
5
5
  end
metadata CHANGED
@@ -1,7 +1,7 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: agent-harness
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.43.0
4
+ version: 0.44.1
5
5
  platform: ruby
6
6
  authors:
7
7
  - Bart Agapinan
@@ -47,16 +47,16 @@ dependencies:
47
47
  name: ruby_llm
48
48
  requirement: !ruby/object:Gem::Requirement
49
49
  requirements:
50
- - - '='
50
+ - - "~>"
51
51
  - !ruby/object:Gem::Version
52
- version: 2.0.0
52
+ version: '2.0'
53
53
  type: :runtime
54
54
  prerelease: false
55
55
  version_requirements: !ruby/object:Gem::Requirement
56
56
  requirements:
57
- - - '='
57
+ - - "~>"
58
58
  - !ruby/object:Gem::Version
59
- version: 2.0.0
59
+ version: '2.0'
60
60
  - !ruby/object:Gem::Dependency
61
61
  name: rake
62
62
  requirement: !ruby/object:Gem::Requirement