agent-harness 0.43.0 → 0.44.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/.release-please-manifest.json +1 -1
- data/CHANGELOG.md +15 -0
- data/README.md +3 -2
- data/docs/provider-neutral-api-execution-contract.md +208 -44
- data/lib/agent_harness/api/chat_transport.rb +3 -1
- data/lib/agent_harness/api/ruby_llm_chat_adapter.rb +6 -30
- data/lib/agent_harness/providers/anthropic.rb +1 -1
- data/lib/agent_harness/version.rb +1 -1
- metadata +5 -5
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 4692f9efdcfe0158cba3afcff4b077fdcc67dc3c9f28198569be1e338add3b74
|
|
4
|
+
data.tar.gz: 3911441fb063e7b52ed74f3fad737d2ef9189bb122897a44e65c1ebf449fcaa8
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: b933a56fb9905d8e007903ff05245d29885b574bd78db06621d135da1e605f8141ad13e32c616614f0183b4757d88f314aa4750737755d73f587af89955ecf20
|
|
7
|
+
data.tar.gz: 43e7cc1a88f2984ad163b788dd56a798429a4cecd102e02c1c7b11dc5247bcbd56811cdf6aa2a2eaca87da59d0f679ecb27b64c99e3b7fa972d021962467501a
|
data/CHANGELOG.md
CHANGED
|
@@ -2,10 +2,25 @@
|
|
|
2
2
|
|
|
3
3
|
### Features
|
|
4
4
|
|
|
5
|
+
* verify RDR-072's published capability artifacts and migration limits, retain the application chat loop over normalized transport, and remove the private RubyLLM streaming-refusal patch in favor of an explicit non-streaming schema boundary ([#437](https://github.com/viamin/agent-harness/issues/437)).
|
|
5
6
|
* expose normalized API chat transport and schema-constrained responses with preserved JSON text, locally validated parsed values, and explicit refusal, truncation, invalid JSON, schema mismatch, and unsupported-mode outcomes; this combined delivery supersedes the separate chat-transport work item ([#433](https://github.com/viamin/agent-harness/issues/433), [#434](https://github.com/viamin/agent-harness/issues/434)).
|
|
6
7
|
* add runner model compatibility contract (`AgentHarness.model_compatibility`) with structured `ModelCompatibility::Result` outcomes. Codex exposes static facts for CLI-gated models (e.g. `gpt-5.5` requires Codex CLI `>= 0.116.0`), a baseline supported-model list, supported auth modes, and a `DEFAULT_COMPATIBLE_MODEL_ID` fallback so downstream orchestrators can validate tier/model assignments before scheduling agent runs ([#259](https://github.com/viamin/agent-harness/issues/259)).
|
|
7
8
|
* **auth:** add provider-owned PKCE code-exchange API for Claude OAuth (`AgentHarness::Authentication.exchange_code`). Takes an authorization code plus PKCE verifier (and `redirect_uri`/`client_id`), posts an `authorization_code` grant to the Claude token endpoint, and persists the resulting access/refresh tokens in the native `claudeAiOauth` shape. Adds `exchange_code_supported?` and a `code_exchange` key to `auth_capabilities` ([#266](https://github.com/viamin/agent-harness/issues/266)).
|
|
8
9
|
|
|
10
|
+
## [0.44.1](https://github.com/viamin/agent-harness/compare/agent-harness/v0.44.0...agent-harness/v0.44.1) (2026-09-27)
|
|
11
|
+
|
|
12
|
+
|
|
13
|
+
### Bug Fixes
|
|
14
|
+
|
|
15
|
+
* **deps-dev:** bump @anthropic-ai/claude-code from 2.1.275 to 2.1.282 in /vendor/pins/claude ([#453](https://github.com/viamin/agent-harness/issues/453)) ([1848bbb](https://github.com/viamin/agent-harness/commit/1848bbb03653e0b64524fbffbcc939eb8542ad2f))
|
|
16
|
+
|
|
17
|
+
## [0.44.0](https://github.com/viamin/agent-harness/compare/agent-harness/v0.43.0...agent-harness/v0.44.0) (2026-09-25)
|
|
18
|
+
|
|
19
|
+
|
|
20
|
+
### Features
|
|
21
|
+
|
|
22
|
+
* Evaluate resumable loop delegation and implement only if simpler (RDR-072) ([#448](https://github.com/viamin/agent-harness/issues/448)) ([02dedfb](https://github.com/viamin/agent-harness/commit/02dedfb08096631e59f326333df61feb57eba395))
|
|
23
|
+
|
|
9
24
|
## [0.43.0](https://github.com/viamin/agent-harness/compare/agent-harness/v0.42.0...agent-harness/v0.43.0) (2026-09-25)
|
|
10
25
|
|
|
11
26
|
|
data/README.md
CHANGED
|
@@ -806,8 +806,9 @@ Health checks run five steps per provider: registration, CLI availability, authe
|
|
|
806
806
|
The proposed provider-neutral API execution boundary for embeddings, chat,
|
|
807
807
|
structured output, usage, and optional persistence is documented in the
|
|
808
808
|
[Provider-Neutral API Execution Contract](docs/provider-neutral-api-execution-contract.md).
|
|
809
|
-
|
|
810
|
-
evidence
|
|
809
|
+
That document records the shipped capability versions, migration examples,
|
|
810
|
+
verified limits, and the evidence-backed decision to retain the application
|
|
811
|
+
conversation loop over the normalized transport.
|
|
811
812
|
|
|
812
813
|
```bash
|
|
813
814
|
# Install dependencies
|
|
@@ -9,9 +9,12 @@ usage, and optional conversation-persistence work.
|
|
|
9
9
|
|
|
10
10
|
RDR-072's rollout guard was **docs-only** for the design phase. The normalized
|
|
11
11
|
chat, attempt-accounting, schema, and embedding capabilities described below
|
|
12
|
-
are now implemented.
|
|
13
|
-
|
|
14
|
-
|
|
12
|
+
are now implemented. The resumable-loop delegation evaluation is also
|
|
13
|
+
complete: it closed with a **retained-loop outcome** (see "Resumable-loop
|
|
14
|
+
delegation evaluation"), so no loop runtime was added. Other capabilities
|
|
15
|
+
remain design contracts and each still needs its own failing-first contract
|
|
16
|
+
tests, implementation, release evidence, and downstream adoption evidence
|
|
17
|
+
before a caller enables it.
|
|
15
18
|
|
|
16
19
|
Existing CLI and subscription behavior remains the default. Existing
|
|
17
20
|
`TextTransport`, `OpenAICompatibleTransport`, `Conversation`, and `Response`
|
|
@@ -553,6 +556,7 @@ records.
|
|
|
553
556
|
| Attempt usage | `usage.ruby_llm` per physical attempt | Useful facts, but the public payload has no stable attempt ID; harness must add one |
|
|
554
557
|
| Plain Ruby resume | transcript can be reconstructed manually | No documented state export/import API; implement normalized export/import outside RubyLLM |
|
|
555
558
|
| Rails resume | `acts_as_chat` transcript plus supporting records | Technically restart-safe at checkpoints; at-least-once side effects remain |
|
|
559
|
+
| Loop controls | `Chat#step`, `#complete`, `#run_tools`, `#approve`, `#deny`, `Tool.requires_approval`, `#cancel` | Evaluated for delegation; retained-loop outcome (see below) |
|
|
556
560
|
|
|
557
561
|
### Optional Rails supporting tables
|
|
558
562
|
|
|
@@ -591,14 +595,96 @@ historical/pending-conversation tests. Reverting the gem is not a data rollback.
|
|
|
591
595
|
2. **Use RubyLLM tool-call and usage tables selectively.** Potentially removes
|
|
592
596
|
bookkeeping, but only after stable-attempt mapping, tenant-scoped access,
|
|
593
597
|
audit, and migration tests are complete.
|
|
594
|
-
3. **Delegate the full loop and all supporting tables.**
|
|
595
|
-
|
|
596
|
-
|
|
598
|
+
3. **Delegate the full loop and all supporting tables.** Evaluated and
|
|
599
|
+
rejected for now: the delegation review below found maintenance increases
|
|
600
|
+
across both repositories. Retaining Paid's loop over the normalized
|
|
601
|
+
transport is the completed outcome and creates no future delegation
|
|
602
|
+
obligation.
|
|
597
603
|
|
|
598
604
|
The state investigation is therefore positive for checkpoint-based Rails
|
|
599
605
|
restoration and normalized plain Ruby reconstruction, and negative for
|
|
600
606
|
exactly-once recovery or an off-the-shelf plain Ruby export/import mechanism.
|
|
601
607
|
|
|
608
|
+
## Resumable-loop delegation evaluation
|
|
609
|
+
|
|
610
|
+
RDR-072 permits delegating the chat loop only when behavior is preserved and
|
|
611
|
+
maintenance decreases across both repositories, counting adapters,
|
|
612
|
+
persistence, and recovery code. Retaining Paid's loop over the normalized
|
|
613
|
+
transport is an acceptable completed outcome. This section records the
|
|
614
|
+
evaluation against RubyLLM 2.0.0's public loop controls
|
|
615
|
+
(`Chat#step`, `#complete`, `#run_tools`, `#complete?`, `#awaiting_approval?`,
|
|
616
|
+
`#pending_approvals`, `#approve`, `#deny`, `#cancel`, and
|
|
617
|
+
`Tool.requires_approval`) and its conclusion.
|
|
618
|
+
|
|
619
|
+
### Verified behaviors
|
|
620
|
+
|
|
621
|
+
`spec/ruby_llm_loop_delegation_evaluation_spec.rb` drives a real
|
|
622
|
+
`RubyLLM::Chat` through the public API with stubbed Anthropic Messages
|
|
623
|
+
responses and pins each fact with a contract test:
|
|
624
|
+
|
|
625
|
+
| Required behavior | Result | Evidence |
|
|
626
|
+
| --- | --- | --- |
|
|
627
|
+
| Single-step execution | Verified | `#step` advances one move; `#complete?` flips only on a final answer |
|
|
628
|
+
| Mixed read/write batches | Verified | Reads without approval execute; `requires_approval` writes stay pending (`#awaiting_approval?`) |
|
|
629
|
+
| Multiple pending decisions | Verified | `#pending_approvals` lists every undecided write; `#approve`/`#deny` resolve them independently |
|
|
630
|
+
| Denial | Verified | A denied call receives a structured denial result and the model continues |
|
|
631
|
+
| Completion | Verified | `#complete?` reports the terminal state |
|
|
632
|
+
| Cancellation | Verified | `#cancel` raises `CancelledError` at the next loop checkpoint, then clears |
|
|
633
|
+
| Completed tool results preserved | Verified | `#run_tools` skips calls that already carry results, so a resumed round executes only the remainder |
|
|
634
|
+
|
|
635
|
+
### Gaps that fail the delegation criteria
|
|
636
|
+
|
|
637
|
+
| Criterion | Finding | Evidence |
|
|
638
|
+
| --- | --- | --- |
|
|
639
|
+
| Stable tool IDs | `ToolCall#id` is the provider wire id verbatim; no caller-stable identity is generated, so the contract's "provider_id is not a durable application identifier" rule needs a new mapping adapter | provider-id example |
|
|
640
|
+
| Iteration limits | The loop has no bound: `#complete` steps until `#complete?` or `#awaiting_approval?`, and `tool_options` exposes only choice/calls/concurrency; bounding stays caller-owned | unbounded-iteration example |
|
|
641
|
+
| Restart restoration without Rails | Decisions recorded with `#approve` live in per-chat in-memory state; a plain Ruby chat rebuilt from the same messages is awaiting approval again. Durable decisions require `acts_as_chat` (Active Record), which plain Ruby consumers must not require | reconstruction example |
|
|
642
|
+
| Crash recovery | Same as above: checkpoint restoration is Rails-only, so a plain Ruby consumer re-implements decision and transcript persistence | reconstruction example |
|
|
643
|
+
| Runner-isolated tool execution | `Tool#execute` runs in the calling process; the loop has no dispatch boundary. Paid executes tools in runner processes under caller-owned authorization and atomic claims, so delegation needs new dispatch adapters rather than removing code | in-process execution throughout |
|
|
644
|
+
| Bounded, non-nested retries | Each loop `#generate` retries beneath the caller through Faraday middleware by default (`max_retries` defaults to three), violating the single-retry-owner and shared `max_attempts` budget rules; even with retries disabled, per-step budget/fallback sequencing is absent | hidden-retry example |
|
|
645
|
+
| Attempt accounting | `usage.ruby_llm` payloads carry only operation/provider/model/status/tokens/cost — no attempt or request identity — so the harness `AttemptReport` ledger with stable `attempt_id` would be bypassed | usage-payload example |
|
|
646
|
+
|
|
647
|
+
### Maintenance comparison
|
|
648
|
+
|
|
649
|
+
Delegating the loop would remove Paid's turn-taking mechanics (step dispatch
|
|
650
|
+
and approve/deny plumbing) while keeping Paid's authorization, atomic claims,
|
|
651
|
+
runner dispatch, iteration bounds, workflow recovery, durable accounting, and
|
|
652
|
+
cross-process cancellation. It would add, across the two repositories:
|
|
653
|
+
|
|
654
|
+
- a provider-id to stable-id mapping adapter;
|
|
655
|
+
- in-process `Tool#execute` to runner-process dispatch adapters;
|
|
656
|
+
- a per-step retry, shared-budget, and fallback sequencing wrapper;
|
|
657
|
+
- a caller-side iteration-limit driver replacing `#complete`;
|
|
658
|
+
- plain Ruby decision/transcript persistence and export/import (Rails
|
|
659
|
+
`acts_as_chat` cannot be a dependency); and
|
|
660
|
+
- attempt-accounting extraction with harness-side attempt identity.
|
|
661
|
+
|
|
662
|
+
The additions exceed the removals, so maintenance increases across both
|
|
663
|
+
repositories and the criteria fail.
|
|
664
|
+
|
|
665
|
+
### Outcome
|
|
666
|
+
|
|
667
|
+
**Retained loop.** Paid keeps its conversation loop over
|
|
668
|
+
`Api::ChatTransport`; the harness keeps one normalized response per call
|
|
669
|
+
with bounded retries, fallback, cancellation, and the attempt ledger. Tool
|
|
670
|
+
execution, approval decisions, authorization, atomic claims, iteration
|
|
671
|
+
limits, and crash recovery remain caller responsibilities under the existing
|
|
672
|
+
ownership boundary. No loop runtime, adapter, persistence, or release was
|
|
673
|
+
added, so there is no migration cost and no consumer activation step. This
|
|
674
|
+
closes the evaluation without creating a future delegation obligation.
|
|
675
|
+
|
|
676
|
+
Re-evaluate only if RubyLLM later documents a public, plain-Ruby-resumable
|
|
677
|
+
loop with caller-stable tool identity, an external tool-dispatch boundary, a
|
|
678
|
+
bounded `#complete`, and per-attempt accounting identity; the evaluation
|
|
679
|
+
spec's examples are the tripwire that detects such changes on upgrade.
|
|
680
|
+
|
|
681
|
+
### Communication
|
|
682
|
+
|
|
683
|
+
The retained-loop outcome, the behavior-test evidence above, and the zero
|
|
684
|
+
migration cost are the message for the Paid adoption and closeout tracking
|
|
685
|
+
(viamin/paid#4014 and the parent viamin/agent-harness#430). No scope beyond
|
|
686
|
+
the accepted RDR-072 alternatives is requested or implied.
|
|
687
|
+
|
|
602
688
|
## Current harness gaps and incremental delivery
|
|
603
689
|
|
|
604
690
|
The current transports already normalize basic text, tool calls, token totals,
|
|
@@ -615,7 +701,8 @@ Ship capabilities independently in this order:
|
|
|
615
701
|
3. normalized non-streaming and streaming chat transport;
|
|
616
702
|
4. structured output for verified model/protocol combinations;
|
|
617
703
|
5. plain Ruby state round trips; and
|
|
618
|
-
6. optional Rails persistence evaluation, then loop evaluation
|
|
704
|
+
6. optional Rails persistence evaluation, then loop evaluation (closed with
|
|
705
|
+
the retained-loop outcome recorded above; no runtime change).
|
|
619
706
|
|
|
620
707
|
Each implementation issue starts with failing contract tests for request-local
|
|
621
708
|
credential isolation, custom endpoints/headers, unsupported outcomes,
|
|
@@ -625,40 +712,117 @@ types stay behind the harness boundary.
|
|
|
625
712
|
|
|
626
713
|
## Compatibility and release evidence
|
|
627
714
|
|
|
628
|
-
|
|
629
|
-
|
|
630
|
-
|
|
631
|
-
|
|
632
|
-
|
|
633
|
-
|
|
634
|
-
|
|
635
|
-
|
|
636
|
-
-
|
|
637
|
-
|
|
638
|
-
|
|
639
|
-
|
|
640
|
-
|
|
641
|
-
|
|
642
|
-
|
|
643
|
-
|
|
644
|
-
|
|
645
|
-
|
|
646
|
-
|
|
647
|
-
|
|
648
|
-
|
|
649
|
-
|
|
650
|
-
|
|
651
|
-
|
|
652
|
-
|
|
653
|
-
-
|
|
654
|
-
|
|
655
|
-
|
|
656
|
-
|
|
657
|
-
|
|
658
|
-
|
|
659
|
-
|
|
660
|
-
|
|
661
|
-
|
|
715
|
+
The following RubyGems releases are the first installable artifacts for each
|
|
716
|
+
capability. Pin the capability's minimum version during migration; do not infer
|
|
717
|
+
support from an issue, branch, or Git tag.
|
|
718
|
+
|
|
719
|
+
| Capability | First installable version | Contract / delivery | Install check |
|
|
720
|
+
| --- | ---: | ---: | --- |
|
|
721
|
+
| Native embeddings | `0.40.0` | [#432](https://github.com/viamin/agent-harness/issues/432) / [#439](https://github.com/viamin/agent-harness/issues/439) | `gem install agent-harness -v 0.40.0` |
|
|
722
|
+
| Normalized chat, tools, and streaming | `0.41.0` | [#434](https://github.com/viamin/agent-harness/issues/434) / [#441](https://github.com/viamin/agent-harness/issues/441) | `gem install agent-harness -v 0.41.0` |
|
|
723
|
+
| Attempt usage and cost | `0.42.0` | [#435](https://github.com/viamin/agent-harness/issues/435) / [#443](https://github.com/viamin/agent-harness/issues/443) | `gem install agent-harness -v 0.42.0` |
|
|
724
|
+
| Schema-constrained responses | `0.43.0` | [#436](https://github.com/viamin/agent-harness/issues/436) / [#442](https://github.com/viamin/agent-harness/issues/442) | `gem install agent-harness -v 0.43.0` |
|
|
725
|
+
|
|
726
|
+
All four are plain-Ruby capabilities. They load without Rails or Active Record,
|
|
727
|
+
create no tables, and require no migration. The runtime floor is Ruby 3.2 and
|
|
728
|
+
RubyLLM 2.x. The four released versions above still pin `ruby_llm = 2.0.0` and
|
|
729
|
+
prepend the private `AgentHarness::Api::RubyLlmResponsesStreamingRefusal`
|
|
730
|
+
module into `RubyLLM::Protocols::Responses` process-wide. The prepend-free,
|
|
731
|
+
public-API-only boundary (RubyLLM's public context, chat, message, tool, token
|
|
732
|
+
and cost APIs only; no private method calls, no downstream monkey patch, and a
|
|
733
|
+
loosened `ruby_llm ~> 2.0` dependency) lands in the first release published
|
|
734
|
+
after [#437](https://github.com/viamin/agent-harness/issues/437); record that
|
|
735
|
+
version here once it exists, because support must not be inferred from an
|
|
736
|
+
issue, branch, or Git tag.
|
|
737
|
+
|
|
738
|
+
### Migration examples and limits
|
|
739
|
+
|
|
740
|
+
- **Embeddings (`>= 0.40.0`):** replace a downstream direct/proxy HTTP patch
|
|
741
|
+
with `AgentHarness.embed`, passing credentials, endpoint and headers on each
|
|
742
|
+
request. Remove the old retry layer after direct and proxy contract tests
|
|
743
|
+
pass. The operation is OpenAI-compatible, API-key only, has batch-total usage
|
|
744
|
+
only, and does not estimate missing tokens.
|
|
745
|
+
- **Chat (`>= 0.41.0`):** construct `AgentHarness::Api::ChatTransport` and send
|
|
746
|
+
one normalized request for each model step. Keep the application loop and
|
|
747
|
+
append completed tool results to the next request. Verified protocols are
|
|
748
|
+
Anthropic Messages and OpenAI Responses/Chat Completions; custom compatible
|
|
749
|
+
endpoints must explicitly choose Chat Completions. Media, custom auth modes,
|
|
750
|
+
protocol probing, and request-local connect timeouts are unsupported.
|
|
751
|
+
- **Attempt accounting (`>= 0.42.0`):** persist each entry in `result[:attempts]`
|
|
752
|
+
once by `attempt_id`, including failed and cancelled attempts. Do not add
|
|
753
|
+
aggregate `result[:usage]` again, convert unknown usage to zero, or wrap the
|
|
754
|
+
call in another provider-request retry loop. Paid still owns durable budget
|
|
755
|
+
accounting, workflow recovery and runner changes.
|
|
756
|
+
- **Schema (`>= 0.43.0`):** set `operation: :schema`, `schema_mode:
|
|
757
|
+
:json_schema`, and provide a named JSON Schema. The harness preserves the
|
|
758
|
+
provider text and validates it locally; use a non-streaming schema request.
|
|
759
|
+
JSON-only mode, schema repair, and model-specific capability discovery are
|
|
760
|
+
also unsupported. Rejection of streamed schema requests
|
|
761
|
+
(`unsupported/structured_output_not_supported`) and removal of the private
|
|
762
|
+
RubyLLM streaming-refusal shim land in the first release published after
|
|
763
|
+
this change; the released `0.43.0` still contains that shim and executes
|
|
764
|
+
streamed schema requests.
|
|
765
|
+
|
|
766
|
+
For every migration, run the consumer contract suite with tenant-specific
|
|
767
|
+
credentials and headers, custom endpoints, auth and transient failures,
|
|
768
|
+
cancellation, and the configured attempt bound. Inspect captured logs for
|
|
769
|
+
credentials and request bodies. Record the exact Paid host and agent-image
|
|
770
|
+
artifact versions before enabling that scope.
|
|
771
|
+
|
|
772
|
+
### Verification and retained paths
|
|
773
|
+
|
|
774
|
+
The [#431 contract audit](https://github.com/viamin/agent-harness/issues/431)
|
|
775
|
+
is implemented by the upstream specs under `spec/agent_harness/embeddings_spec.rb`
|
|
776
|
+
and `spec/agent_harness/api/`. Those specs cover request-local credential, endpoint and header
|
|
777
|
+
isolation; reserved authentication headers; classified errors; bounded retries
|
|
778
|
+
with RubyLLM retries disabled; cancellation; partial streams; stable attempt
|
|
779
|
+
identity; usage/cost serialization; secret-free observer data; schema refusal,
|
|
780
|
+
truncation and validation; and the public-only RubyLLM integration. The normal
|
|
781
|
+
provider, command-executor, model-discovery and authentication suites remain the
|
|
782
|
+
CLI/subscription regression coverage.
|
|
783
|
+
|
|
784
|
+
CLI and subscription execution, legacy `TextTransport` and
|
|
785
|
+
`OpenAICompatibleTransport`, provider selection, workflow recovery, and Paid's
|
|
786
|
+
durable accounting are intentionally unchanged. No credential or authentication
|
|
787
|
+
mode is selected implicitly. Rails persistence and conversation state round
|
|
788
|
+
trips were not adopted.
|
|
789
|
+
|
|
790
|
+
### Retained-loop closeout
|
|
791
|
+
|
|
792
|
+
The completed boundary is normalized transport with Paid's loop retained.
|
|
793
|
+
`Api::ChatTransport` performs one assistant response and never executes tools;
|
|
794
|
+
Paid continues to own authorization, confirmation policy, tool side effects,
|
|
795
|
+
transcript persistence, approval resumption and crash recovery. Delegating the
|
|
796
|
+
loop would require new persistence, tenant/audit adapters and recovery glue in
|
|
797
|
+
both repositories without removing those application responsibilities. There
|
|
798
|
+
is therefore no demonstrated net maintenance reduction or safe migration of
|
|
799
|
+
pending conversations. Under RDR-072 this evidence-backed retained-loop result
|
|
800
|
+
is complete, not deferred delegation.
|
|
801
|
+
|
|
802
|
+
Downstream adoption is tracked in
|
|
803
|
+
[viamin/paid#4014](https://github.com/viamin/paid/issues/4014); preservation of
|
|
804
|
+
Codex subscription discovery/recovery during dependency adoption is tracked in
|
|
805
|
+
[viamin/paid#3995](https://github.com/viamin/paid/issues/3995). Close parent
|
|
806
|
+
[agent-harness#430](https://github.com/viamin/agent-harness/issues/430) only
|
|
807
|
+
after those consumers record their exact host/image artifacts and passing
|
|
808
|
+
integration evidence for every scope they enable.
|
|
809
|
+
|
|
810
|
+
### Loop-delegation evaluation evidence (retained loop)
|
|
811
|
+
|
|
812
|
+
- Publication: none required. The evaluation changed documentation and tests
|
|
813
|
+
only; no runtime path, adapter, table, or public API was added or altered,
|
|
814
|
+
so every released agent-harness version already carries the outcome.
|
|
815
|
+
- Verification: `spec/ruby_llm_loop_delegation_evaluation_spec.rb` exercises
|
|
816
|
+
only the public RubyLLM 2.0 API (pinned `ruby_llm = 2.0.0`) against stubbed
|
|
817
|
+
Anthropic Messages responses and runs with the upstream suite and lint on
|
|
818
|
+
the repository's supported Ruby environment. It records the verified loop
|
|
819
|
+
behaviors and the gaps that failed the delegation criteria.
|
|
820
|
+
- Migration: none. Paid and other consumers keep their existing loops over
|
|
821
|
+
`Api::ChatTransport`; no consumer activation, feature flag, or rollout step
|
|
822
|
+
applies.
|
|
823
|
+
- Retained responsibilities: documented in "Resumable-loop delegation
|
|
824
|
+
evaluation" above; loop ownership stays with Paid under RDR-072's accepted
|
|
825
|
+
alternatives.
|
|
662
826
|
|
|
663
827
|
AgentHarness currently supports Ruby 3.2 and later and must remain usable as a
|
|
664
828
|
plain Ruby gem. RubyLLM 2.0.0 itself supports Ruby 3.1.3 and later and adds
|
|
@@ -683,6 +847,6 @@ For every capability, release evidence must name:
|
|
|
683
847
|
- retained paths and follow-up issues for combinations not migrated.
|
|
684
848
|
|
|
685
849
|
Paid issue `viamin/paid#4014` should receive this compatibility result, the
|
|
686
|
-
state-restoration conclusion, the stable-ID gap,
|
|
687
|
-
evidence
|
|
688
|
-
from a Git tag alone.
|
|
850
|
+
state-restoration conclusion, the stable-ID gap, the per-capability release
|
|
851
|
+
evidence, and the retained-loop outcome of the delegation evaluation.
|
|
852
|
+
Downstream adoption cannot proceed from this closeout or from a Git tag alone.
|
|
@@ -373,7 +373,9 @@ module AgentHarness
|
|
|
373
373
|
end
|
|
374
374
|
|
|
375
375
|
def schema_mode_supported?
|
|
376
|
-
|
|
376
|
+
return true unless schema_operation?
|
|
377
|
+
|
|
378
|
+
request[:stream] != true && (!request[:schema_mode] || request[:schema_mode].to_sym == :json_schema)
|
|
377
379
|
end
|
|
378
380
|
|
|
379
381
|
def schema_payload
|
|
@@ -5,22 +5,6 @@ require "ruby_llm"
|
|
|
5
5
|
|
|
6
6
|
module AgentHarness
|
|
7
7
|
module Api
|
|
8
|
-
# RubyLLM 2.0.0 flattens Responses API refusal deltas into ordinary text
|
|
9
|
-
# chunks and offers no public access to their event type. The gemspec pins
|
|
10
|
-
# that exact release so this compatibility shim cannot silently outlive the
|
|
11
|
-
# private parser shape it targets. Remove it when RubyLLM exposes refusals.
|
|
12
|
-
module RubyLlmResponsesStreamingRefusal
|
|
13
|
-
module RefusalChunk; end
|
|
14
|
-
|
|
15
|
-
def build_chunk(data)
|
|
16
|
-
super.tap do |chunk|
|
|
17
|
-
chunk.extend(RefusalChunk) if data["type"] == "response.refusal.delta"
|
|
18
|
-
end
|
|
19
|
-
end
|
|
20
|
-
end
|
|
21
|
-
|
|
22
|
-
RubyLLM::Protocols::Responses.prepend(RubyLlmResponsesStreamingRefusal)
|
|
23
|
-
|
|
24
8
|
# Translates the normalized public chat values to RubyLLM public objects.
|
|
25
9
|
class RubyLlmChatAdapter
|
|
26
10
|
class UnsupportedOptionError < StandardError; end
|
|
@@ -38,12 +22,12 @@ module AgentHarness
|
|
|
38
22
|
provider_usage_reported = false
|
|
39
23
|
chat = prepared_chat || prepare(candidate:, messages:, tools:, max_output_tokens:, temperature:, timeout:,
|
|
40
24
|
schema:, on_accounting:, provider_usage: -> { provider_usage_reported })
|
|
41
|
-
response
|
|
25
|
+
response = generate(chat, stream, cancellation) do |event|
|
|
42
26
|
provider_usage_reported = true if event[:type] == :usage_updated
|
|
43
27
|
on_event&.call(event)
|
|
44
28
|
end
|
|
45
29
|
emit_completed_tool_calls(response, &on_event) if stream
|
|
46
|
-
normalize_response(response
|
|
30
|
+
normalize_response(response)
|
|
47
31
|
end
|
|
48
32
|
|
|
49
33
|
def prepare(candidate:, messages:, tools:, max_output_tokens:, temperature:, timeout:, on_accounting: nil,
|
|
@@ -61,7 +45,6 @@ module AgentHarness
|
|
|
61
45
|
# duplicate cumulative reports are not re-emitted.
|
|
62
46
|
class StreamState
|
|
63
47
|
attr_accessor :input_tokens, :output_tokens
|
|
64
|
-
attr_reader :refusal
|
|
65
48
|
|
|
66
49
|
def initialize
|
|
67
50
|
@provider_id_by_key = {}
|
|
@@ -69,7 +52,6 @@ module AgentHarness
|
|
|
69
52
|
@latest_provider_id = nil
|
|
70
53
|
@input_tokens = nil
|
|
71
54
|
@output_tokens = nil
|
|
72
|
-
@refusal = false
|
|
73
55
|
end
|
|
74
56
|
|
|
75
57
|
# Links a stream chunk key to its provider call id, returning true
|
|
@@ -90,10 +72,6 @@ module AgentHarness
|
|
|
90
72
|
total = (input_tokens + output_tokens) if input_tokens && output_tokens
|
|
91
73
|
{input_tokens: input_tokens, output_tokens: output_tokens, total_tokens: total}
|
|
92
74
|
end
|
|
93
|
-
|
|
94
|
-
def observe(chunk)
|
|
95
|
-
@refusal ||= chunk.is_a?(RubyLlmResponsesStreamingRefusal::RefusalChunk)
|
|
96
|
-
end
|
|
97
75
|
end
|
|
98
76
|
|
|
99
77
|
private
|
|
@@ -182,10 +160,10 @@ module AgentHarness
|
|
|
182
160
|
end
|
|
183
161
|
|
|
184
162
|
def generate(chat, stream, cancellation)
|
|
185
|
-
return
|
|
163
|
+
return generate_without_events(chat, cancellation) unless stream
|
|
186
164
|
|
|
187
165
|
state = StreamState.new
|
|
188
|
-
|
|
166
|
+
chat.generate do |chunk|
|
|
189
167
|
if cancelled?(cancellation)
|
|
190
168
|
chat.cancel
|
|
191
169
|
raise RubyLLM::CancelledError
|
|
@@ -193,7 +171,6 @@ module AgentHarness
|
|
|
193
171
|
|
|
194
172
|
stream_events(chunk, state).each { |event| yield event }
|
|
195
173
|
end
|
|
196
|
-
[response, state.refusal]
|
|
197
174
|
end
|
|
198
175
|
|
|
199
176
|
def generate_without_events(chat, cancellation)
|
|
@@ -259,7 +236,6 @@ module AgentHarness
|
|
|
259
236
|
end
|
|
260
237
|
|
|
261
238
|
def stream_events(chunk, state)
|
|
262
|
-
state.observe(chunk)
|
|
263
239
|
[text_event(chunk), *tool_call_events(chunk, state), usage_event(chunk, state)].compact
|
|
264
240
|
end
|
|
265
241
|
|
|
@@ -325,14 +301,14 @@ module AgentHarness
|
|
|
325
301
|
end
|
|
326
302
|
end
|
|
327
303
|
|
|
328
|
-
def normalize_response(response
|
|
304
|
+
def normalize_response(response)
|
|
329
305
|
{
|
|
330
306
|
content: response.content || "",
|
|
331
307
|
model: response.model,
|
|
332
308
|
finish_reason: response.finish_reason,
|
|
333
309
|
usage: normalize_usage(response.tokens),
|
|
334
310
|
tool_calls: Array(response.tool_calls&.values).map { |call| normalize_tool_call(call) },
|
|
335
|
-
refusal:
|
|
311
|
+
refusal: refusal?(response)
|
|
336
312
|
}
|
|
337
313
|
end
|
|
338
314
|
|
|
@@ -27,7 +27,7 @@ module AgentHarness
|
|
|
27
27
|
|
|
28
28
|
# Model name pattern for Anthropic Claude models
|
|
29
29
|
MODEL_PATTERN = /^claude-[\d.-]+-(?:opus|sonnet|haiku)(?:-\d{8})?$/i
|
|
30
|
-
SUPPORTED_CLI_VERSION = "2.1.
|
|
30
|
+
SUPPORTED_CLI_VERSION = "2.1.282"
|
|
31
31
|
SUPPORTED_CLI_REQUIREMENT = Gem::Requirement.new(">= #{SUPPORTED_CLI_VERSION}", "< 2.2.0").freeze
|
|
32
32
|
|
|
33
33
|
# Matches semver (e.g. "2.1.92"), optional pre-release (e.g. "2.1.92-beta.1"),
|
metadata
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: agent-harness
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 0.
|
|
4
|
+
version: 0.44.1
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Bart Agapinan
|
|
@@ -47,16 +47,16 @@ dependencies:
|
|
|
47
47
|
name: ruby_llm
|
|
48
48
|
requirement: !ruby/object:Gem::Requirement
|
|
49
49
|
requirements:
|
|
50
|
-
- -
|
|
50
|
+
- - "~>"
|
|
51
51
|
- !ruby/object:Gem::Version
|
|
52
|
-
version: 2.0
|
|
52
|
+
version: '2.0'
|
|
53
53
|
type: :runtime
|
|
54
54
|
prerelease: false
|
|
55
55
|
version_requirements: !ruby/object:Gem::Requirement
|
|
56
56
|
requirements:
|
|
57
|
-
- -
|
|
57
|
+
- - "~>"
|
|
58
58
|
- !ruby/object:Gem::Version
|
|
59
|
-
version: 2.0
|
|
59
|
+
version: '2.0'
|
|
60
60
|
- !ruby/object:Gem::Dependency
|
|
61
61
|
name: rake
|
|
62
62
|
requirement: !ruby/object:Gem::Requirement
|