ollama-client 1.3.0 → 1.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/.env.example +11 -0
- data/API_CONTRACT.md +79 -8
- data/CHANGELOG.md +45 -0
- data/CONTRIBUTING.md +8 -0
- data/README.md +205 -273
- data/ROADMAP.md +23 -0
- data/docs/API_GAPS.md +18 -141
- data/docs/ARCHITECTURE.md +141 -0
- data/docs/AREAS_FOR_CONSIDERATION.md +13 -1
- data/docs/CLOUD.md +19 -0
- data/docs/CONSOLE_IMPROVEMENTS.md +20 -1
- data/docs/ECOSYSTEM_GEMS.md +41 -0
- data/docs/ECOSYSTEM_STRATEGY.md +151 -0
- data/docs/GETTING_STARTED.md +16 -8
- data/docs/INTEGRATION_TESTING.md +23 -5
- data/docs/PRODUCTION_FIXES.md +15 -3
- data/docs/QUICK_START.md +1 -1
- data/docs/README.md +1 -0
- data/docs/RUBYLLM_ADOPTION_MATRIX.md +273 -0
- data/docs/adr/001-openai-boundary.md +12 -0
- data/docs/adr/002-transport-abstraction.md +12 -0
- data/docs/adr/003-response-normalization.md +12 -0
- data/docs/adr/004-mock-transport.md +16 -0
- data/docs/adr/005-error-taxonomy.md +16 -0
- data/docs/adr/006-stream-runtime.md +16 -0
- data/docs/ecosystem/BOUNDARIES.md +19 -0
- data/docs/ecosystem/DEPENDENCY_GRAPH.md +14 -0
- data/docs/ecosystem/DESIGN_PRINCIPLES.md +9 -0
- data/docs/ecosystem/EXISTING_REPOS.md +31 -0
- data/docs/ecosystem/EXPERIMENTAL_LABS.md +24 -0
- data/docs/ecosystem/OVERVIEW.md +28 -0
- data/docs/ecosystem/RELEASE_ORDER.md +17 -0
- data/docs/observability/README.md +8 -0
- data/docs/rails/README.md +7 -0
- data/docs/rfcs/0001-stream-runtime.md +6 -0
- data/docs/rfcs/0002-schema-system.md +6 -0
- data/docs/rfcs/0003-observability-hooks.md +6 -0
- data/docs/rfcs/0004-async-runtime.md +6 -0
- data/docs/rfcs/README.md +16 -0
- data/docs/runtime/ERROR_CONTRACT.md +14 -0
- data/docs/runtime/SCHEMA_CONTRACT.md +28 -0
- data/docs/runtime/STREAM_CONTRACT.md +17 -0
- data/docs/runtime/STREAM_RUNTIME.md +20 -0
- data/docs/runtime/TRANSPORT_CONTRACT.md +26 -0
- data/docs/schema/README.md +8 -0
- data/docs/schema/STRUCTURED_OUTPUTS.md +22 -0
- data/docs/streaming/README.md +8 -0
- data/docs/testing/README.md +7 -0
- data/docs/testing/REPLAY_SYSTEM.md +17 -0
- data/docs/transport/README.md +7 -0
- data/examples/README.md +44 -0
- data/examples/basic_chat.rb +9 -0
- data/examples/cloud_models.rb +162 -0
- data/examples/embeddings.rb +10 -0
- data/examples/free_catalog.json +290 -0
- data/examples/generate.rb +9 -0
- data/examples/streaming.rb +12 -0
- data/examples/structured_tools.rb +90 -0
- data/examples/tool_calling_direct.rb +101 -0
- data/examples/tool_dto_example.rb +94 -0
- data/exe/ollama-client +5 -1
- data/lib/ollama/agent/executor.rb +243 -0
- data/lib/ollama/agent/messages.rb +31 -0
- data/lib/ollama/agent/planner.rb +45 -0
- data/lib/ollama/api_key_pool.rb +61 -0
- data/lib/ollama/attachment.rb +74 -0
- data/lib/ollama/capabilities.rb +1 -1
- data/lib/ollama/chat_response.rb +33 -0
- data/lib/ollama/client/chat/request_preparer.rb +82 -0
- data/lib/ollama/client/chat.rb +58 -68
- data/lib/ollama/client/chat_stream_processor.rb +85 -25
- data/lib/ollama/client/generate/request_preparer.rb +118 -0
- data/lib/ollama/client/generate/response_formatter.rb +95 -0
- data/lib/ollama/client/generate.rb +83 -165
- data/lib/ollama/client/model_management.rb +182 -84
- data/lib/ollama/client/openai_compat.rb +185 -0
- data/lib/ollama/client/raw.rb +66 -0
- data/lib/ollama/client/tool_intent.rb +38 -0
- data/lib/ollama/client/web_search.rb +39 -0
- data/lib/ollama/client.rb +83 -35
- data/lib/ollama/config.rb +122 -17
- data/lib/ollama/embeddings.rb +67 -29
- data/lib/ollama/errors.rb +55 -1
- data/lib/ollama/events.rb +55 -0
- data/lib/ollama/generate_stream_handler.rb +23 -4
- data/lib/ollama/http_error_handler.rb +42 -0
- data/lib/ollama/messages.rb +109 -0
- data/lib/ollama/middleware/cache.rb +73 -0
- data/lib/ollama/middleware/logger.rb +74 -0
- data/lib/ollama/middleware/metrics.rb +84 -0
- data/lib/ollama/middleware/tracing.rb +101 -0
- data/lib/ollama/middleware.rb +39 -0
- data/lib/ollama/model_profile.rb +1 -1
- data/lib/ollama/openai.rb +16 -0
- data/lib/ollama/options.rb +63 -1
- data/lib/ollama/params.rb +139 -0
- data/lib/ollama/parsers/base.rb +24 -0
- data/lib/ollama/parsers/chat.rb +22 -0
- data/lib/ollama/parsers/embeddings.rb +23 -0
- data/lib/ollama/parsers/generate.rb +38 -0
- data/lib/ollama/parsers/list_running.rb +15 -0
- data/lib/ollama/parsers/show_model.rb +14 -0
- data/lib/ollama/parsers/version.rb +15 -0
- data/lib/ollama/pipeline.rb +169 -0
- data/lib/ollama/plugins.rb +83 -0
- data/lib/ollama/policies/auto_pull.rb +69 -0
- data/lib/ollama/policies/base.rb +70 -0
- data/lib/ollama/policies/capability_validation.rb +88 -0
- data/lib/ollama/policies/fallback.rb +57 -0
- data/lib/ollama/policies/rate_limit.rb +115 -0
- data/lib/ollama/policies/repair_json.rb +137 -0
- data/lib/ollama/policies/retry/strategies/exponential.rb +27 -0
- data/lib/ollama/policies/retry/strategies/fixed.rb +27 -0
- data/lib/ollama/policies/retry/strategies/jitter.rb +30 -0
- data/lib/ollama/policies/retry/strategies/linear.rb +27 -0
- data/lib/ollama/policies/retry.rb +152 -0
- data/lib/ollama/policies/schema_repair.rb +180 -0
- data/lib/ollama/policies/timeout.rb +43 -0
- data/lib/ollama/policies.rb +24 -0
- data/lib/ollama/prompt.rb +86 -0
- data/lib/ollama/prompt_adapters/base.rb +3 -2
- data/lib/ollama/prompt_adapters/gemma4.rb +22 -14
- data/lib/ollama/prompts/tool_planner.rb +35 -0
- data/lib/ollama/providers/base.rb +70 -0
- data/lib/ollama/providers/llama_cpp.rb +131 -0
- data/lib/ollama/providers/ollama.rb +54 -0
- data/lib/ollama/providers/openai.rb +134 -0
- data/lib/ollama/providers.rb +28 -0
- data/lib/ollama/rate_limit_handler.rb +48 -0
- data/lib/ollama/request.rb +169 -0
- data/lib/ollama/response.rb +3 -2
- data/lib/ollama/responses/base.rb +11 -0
- data/lib/ollama/responses/chat.rb +11 -0
- data/lib/ollama/responses/embeddings.rb +14 -0
- data/lib/ollama/responses/generate.rb +22 -0
- data/lib/ollama/schema_dsl.rb +97 -0
- data/lib/ollama/schema_validator.rb +90 -53
- data/lib/ollama/schemas/tool_intent.json +15 -0
- data/lib/ollama/schemas/tool_intent.rb +16 -0
- data/lib/ollama/serializers/base.rb +44 -0
- data/lib/ollama/serializers/chat.rb +42 -0
- data/lib/ollama/serializers/embeddings.rb +36 -0
- data/lib/ollama/serializers/generate.rb +47 -0
- data/lib/ollama/streaming_observer.rb +22 -0
- data/lib/ollama/testing.rb +103 -0
- data/lib/ollama/tool/function/parameters/property.rb +72 -0
- data/lib/ollama/tool/function/parameters.rb +101 -0
- data/lib/ollama/tool/function.rb +78 -0
- data/lib/ollama/tool.rb +60 -0
- data/lib/ollama/tool_dsl.rb +93 -0
- data/lib/ollama/tool_intent.rb +19 -0
- data/lib/ollama/transport/base.rb +42 -0
- data/lib/ollama/transport/mock.rb +48 -0
- data/lib/ollama/transport/net_http.rb +76 -0
- data/lib/ollama/transport/request.rb +20 -0
- data/lib/ollama/transport/response.rb +41 -0
- data/lib/ollama/transport.rb +26 -0
- data/lib/ollama/version.rb +1 -1
- data/lib/ollama_client.rb +35 -0
- data/script/live_branch_smoke/chat_exercises.rb +232 -0
- data/script/live_branch_smoke/generation_exercises.rb +128 -0
- data/script/live_branch_smoke/model_exercises.rb +105 -0
- data/script/live_branch_smoke/utility_exercises.rb +375 -0
- data/script/live_branch_smoke.rb +159 -174
- data/test_all_features.rb +315 -0
- metadata +158 -24
- data/.cursor/.gitignore +0 -1
- data/RELEASE_NOTES_v0.2.6.md +0 -41
- data/devagent_proper.rb +0 -430
- data/docs/TESTING.md +0 -508
- data/examples/agent_loop.rb +0 -120
- data/examples/failure_modes/invalid_json_repair.rb +0 -42
- data/examples/production/rails_agent.rb +0 -62
- data/market.jpg +0 -0
- data/print_capabilities.rb +0 -20
- data/schema.json +0 -1
- data/test_tool.rb +0 -26
|
@@ -0,0 +1,151 @@
|
|
|
1
|
+
# Ollama Ruby Ecosystem Strategy & Architectural Blueprint
|
|
2
|
+
|
|
3
|
+
This document defines the overarching architectural strategy, dependency contracts, and governance model for the Ollama Ruby ecosystem. It establishes the foundational guardrails ensuring that as the ecosystem expands across multiple specialized gems, it maintains strict determinism, modularity, and architectural integrity.
|
|
4
|
+
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
## 1. Executive Summary & Core Philosophy
|
|
8
|
+
|
|
9
|
+
The Ollama Ruby ecosystem is designed around a **deterministic kernel** (`ollama-client`) surrounded by **modular companion gems**.
|
|
10
|
+
|
|
11
|
+
### The Core Philosophy
|
|
12
|
+
1. **Deterministic Kernel**: `ollama-client` is the canonical source of truth for all low-level transport, retry semantics, connection pooling, raw endpoint definitions, and error taxonomy.
|
|
13
|
+
2. **Composition over Inheritance**: Companion gems (`ollama-openai`, `ollama-observability`, `ollama-stream`, `ollama-agent`, `ollama-rails`) build upon the core client through well-defined public interfaces and callback hooks rather than monkey-patching or redefining internal transport logic.
|
|
14
|
+
3. **Domain Isolation**: Each companion gem encapsulates a single, cohesive domain (e.g., OpenAI compatibility, OpenTelemetry instrumentation, advanced streaming, agent orchestration, Rails integration).
|
|
15
|
+
|
|
16
|
+
---
|
|
17
|
+
|
|
18
|
+
## 2. The Dependency Hierarchy & Directionality
|
|
19
|
+
|
|
20
|
+
To prevent architectural inversion and circular dependencies, all gems in the ecosystem must adhere strictly to the following unidirectional dependency graph:
|
|
21
|
+
|
|
22
|
+
```
|
|
23
|
+
┌────────────────────────────────────────────────────────┐
|
|
24
|
+
│ ollama-client (Core Kernel) │
|
|
25
|
+
│ (Transport, Retries, Error Taxonomy, Raw Schemas) │
|
|
26
|
+
└───────────────────────────▲────────────────────────────┘
|
|
27
|
+
│
|
|
28
|
+
┌──────────────────┼──────────────────┐
|
|
29
|
+
│ │ │
|
|
30
|
+
┌────────┴────────┐┌────────┴────────┐┌────────┴────────┐
|
|
31
|
+
│ ollama-openai ││ollama-observab. ││ ollama-stream │
|
|
32
|
+
│(OpenAI Facade) ││ (OTel Telemetry)││ (SSE & WebSockets│
|
|
33
|
+
└────────▲────────┘└─────────────────┘└────────▲────────┘
|
|
34
|
+
│ │
|
|
35
|
+
└──────────────────┬──────────────────┘
|
|
36
|
+
│
|
|
37
|
+
┌──────────┴──────────┐
|
|
38
|
+
│ ollama-agent │
|
|
39
|
+
│(Tools, Memory, Plan)│
|
|
40
|
+
└──────────▲──────────┘
|
|
41
|
+
│
|
|
42
|
+
┌──────────┴──────────┐
|
|
43
|
+
│ ollama-rails │
|
|
44
|
+
│(ActiveJob, Turbo, │
|
|
45
|
+
│ ActionCable) │
|
|
46
|
+
└─────────────────────┘
|
|
47
|
+
```
|
|
48
|
+
|
|
49
|
+
### Critical Dependency Rules
|
|
50
|
+
- **Rule 1**: `ollama-client` must NEVER depend on any companion gem or higher-level concept (e.g., agents, Rails, OpenTelemetry).
|
|
51
|
+
- **Rule 2**: Companion gems must declare `ollama-client` as their primary dependency and utilize its public API or hook system.
|
|
52
|
+
- **Rule 3**: Higher-level orchestration gems (`ollama-agent`, `ollama-rails`) may compose multiple lower-level companion gems (e.g., `ollama-agent` utilizing `ollama-stream` for tool streaming or `ollama-openai` for LLM interop).
|
|
53
|
+
|
|
54
|
+
---
|
|
55
|
+
|
|
56
|
+
## 3. Companion Gem Specifications & Responsibilities
|
|
57
|
+
|
|
58
|
+
### 1. `ollama-client` (The Kernel)
|
|
59
|
+
- **Role**: Canonical infrastructure layer.
|
|
60
|
+
- **Responsibilities**: Faraday HTTP/Faraday WebSocket transport, connection pooling, exponential backoff retries, raw API endpoint mapping (`/api/chat`, `/api/generate`, `/api/embeddings`, `/api/pull`), configuration management (`Ollama::Config`), and canonical error taxonomy (`Ollama::Error`, `Ollama::ConnectionError`, `Ollama::TimeoutError`, `Ollama::RateLimitError`).
|
|
61
|
+
- **Prohibited**: High-level workflow orchestration, third-party framework coupling.
|
|
62
|
+
|
|
63
|
+
### 2. `ollama-openai` (Protocol Interop)
|
|
64
|
+
- **Role**: Drop-in OpenAI compatibility layer.
|
|
65
|
+
- **Responsibilities**: Translating OpenAI-style requests (`client.chat(parameters: {})`) into native Ollama payloads, mapping OpenAI function definitions to Ollama tool schemas, normalizing Ollama responses into OpenAI JSON structures (`chatcmpl-*`), and companion support for frameworks like LangChain, Vercel AI SDK, and ruby-openai consumers.
|
|
66
|
+
- **Prohibited**: Custom HTTP transport implementations.
|
|
67
|
+
|
|
68
|
+
### 3. `ollama-observability` (Telemetry & Logging)
|
|
69
|
+
- **Role**: Enterprise-grade observability layer.
|
|
70
|
+
- **Responsibilities**: Subscribing to `ollama-client` hooks (`on_response`, `on_token`, `on_error`) to generate OpenTelemetry tracing spans, tracking metrics (Time-to-First-Token, latency histograms, prompt/completion token counters), and emitting structured JSON logs. Supports payload redaction for PII/PHI compliance.
|
|
71
|
+
- **Prohibited**: Modifying raw API responses or interfering with inference execution.
|
|
72
|
+
|
|
73
|
+
### 4. `ollama-stream` (Advanced Streaming & Transport)
|
|
74
|
+
- **Role**: High-performance streaming runtime.
|
|
75
|
+
- **Responsibilities**: Encapsulating SSE streams into formal `Ollama::Stream::StreamObject` instances supporting pause, resume, and cancellation; managing persistent bidirectional WebSocket sessions; providing backpressure and queue bounding (`FlowController`); implementing incremental JSON fragment recovery (`IncrementalParser`); and offering a Rack-compatible SSE proxy adapter.
|
|
76
|
+
- **Prohibited**: Re-implementing base Faraday connection logic.
|
|
77
|
+
|
|
78
|
+
### 5. `ollama-agent` (Orchestration & Tooling)
|
|
79
|
+
- **Role**: Autonomous agent execution framework.
|
|
80
|
+
- **Responsibilities**: Defining standardized tool contracts (`Ollama::Agent::Tool`), managing tool registries, maintaining conversation memory (`WindowMemory`, `SummaryMemory`), executing autonomous ReAct/Plan-and-Solve agent loops (`Executor`), and providing structured JSON output parsing.
|
|
81
|
+
- **Prohibited**: Direct HTTP transport management.
|
|
82
|
+
|
|
83
|
+
### 6. `ollama-rails` (Rails Integration)
|
|
84
|
+
- **Role**: Idiomatic Ruby on Rails integration.
|
|
85
|
+
- **Responsibilities**: Providing Rails Railtie for zero-config initialization, wrapping async inference in ActiveJob (`Ollama::Rails::GenerateJob`, `ChatJob`), broadcasting live stream tokens over ActionCable/Turbo Streams (`Ollama::Rails::BroadcastHelpers`), offering ActiveRecord mixins (`Ollama::Rails::Embeddable` for pgvector/neighbor integration), and providing Rake tasks for model management (`ollama:pull`, `ollama:list`).
|
|
86
|
+
- **Prohibited**: Modifying core Ruby runtime behavior outside the Rails application context.
|
|
87
|
+
|
|
88
|
+
---
|
|
89
|
+
|
|
90
|
+
## 4. Architectural Guardrails & Extension Mechanisms
|
|
91
|
+
|
|
92
|
+
### The Hook System
|
|
93
|
+
To enable companion gems to extend functionality without monkey-patching, `ollama-client` exposes a robust hook and middleware architecture:
|
|
94
|
+
|
|
95
|
+
```ruby
|
|
96
|
+
# Example of Hook Subscription in Companion Gems
|
|
97
|
+
config.on_response = ->(raw_response, metadata) {
|
|
98
|
+
# Used by ollama-observability for metrics and logging
|
|
99
|
+
Telemetry.record_latency(metadata[:duration])
|
|
100
|
+
}
|
|
101
|
+
|
|
102
|
+
client.chat(
|
|
103
|
+
model: "llama3",
|
|
104
|
+
messages: history,
|
|
105
|
+
hooks: {
|
|
106
|
+
on_token: ->(chunk, logprobs) { # Used by ollama-stream },
|
|
107
|
+
on_tool_call: ->(tool_call) { # Used by ollama-agent },
|
|
108
|
+
on_error: ->(error) { # Used by ollama-observability }
|
|
109
|
+
}
|
|
110
|
+
)
|
|
111
|
+
```
|
|
112
|
+
|
|
113
|
+
### Error Handling & Taxonomy
|
|
114
|
+
All companion gems must catch base `Ollama::Error` exceptions from `ollama-client` and either allow them to bubble up or wrap them in domain-specific exceptions that inherit from `Ollama::Error` (e.g., `Ollama::Openai::Error < Ollama::Error`).
|
|
115
|
+
|
|
116
|
+
---
|
|
117
|
+
|
|
118
|
+
## 5. Development & Workspace Governance
|
|
119
|
+
|
|
120
|
+
### Local Dependency Mapping
|
|
121
|
+
During development, all companion gems must reference the local `ollama-client` kernel in their `Gemfile` to ensure continuous integration and contract verification across the workspace:
|
|
122
|
+
|
|
123
|
+
```ruby
|
|
124
|
+
# Gemfile for Companion Gems
|
|
125
|
+
source "https://rubygems.org"
|
|
126
|
+
gemspec
|
|
127
|
+
|
|
128
|
+
# Enforce local workspace dependency
|
|
129
|
+
gem "ollama-client", path: "../ollama-client"
|
|
130
|
+
```
|
|
131
|
+
|
|
132
|
+
### Verification & Testing
|
|
133
|
+
Every companion gem must maintain an independent RSpec test suite verifying its specific domain logic against mock `Ollama::Client` instances. Contract tests must ensure that changes in `ollama-client` do not break companion gem facades.
|
|
134
|
+
|
|
135
|
+
---
|
|
136
|
+
|
|
137
|
+
## 6. Roadmap & Future Evolution
|
|
138
|
+
|
|
139
|
+
1. **Phase 1: Foundation (Completed)**
|
|
140
|
+
- Core transport, retry semantics, error taxonomy, and configuration in `ollama-client`.
|
|
141
|
+
- OpenAI interop facade in `ollama-openai`.
|
|
142
|
+
- Enterprise telemetry in `ollama-observability`.
|
|
143
|
+
- Advanced streaming primitives in `ollama-stream`.
|
|
144
|
+
|
|
145
|
+
2. **Phase 2: Orchestration & Frameworks (Current)**
|
|
146
|
+
- Autonomous agent loops and tool calling in `ollama-agent`.
|
|
147
|
+
- Idiomatic Rails integrations in `ollama-rails`.
|
|
148
|
+
|
|
149
|
+
3. **Phase 3: Advanced Optimization & Ecosystem Scaling (Future)**
|
|
150
|
+
- Native C/Rust extensions for high-throughput incremental JSON parsing.
|
|
151
|
+
- Distributed multi-node Ollama cluster balancing and routing gems.
|
data/docs/GETTING_STARTED.md
CHANGED
|
@@ -7,11 +7,13 @@ This guide shows you step-by-step how to create a client object to use all featu
|
|
|
7
7
|
### Option A: Using Bundler (Recommended)
|
|
8
8
|
|
|
9
9
|
Add to your `Gemfile`:
|
|
10
|
+
|
|
10
11
|
```ruby
|
|
11
12
|
gem "ollama-client"
|
|
12
13
|
```
|
|
13
14
|
|
|
14
15
|
Then run:
|
|
16
|
+
|
|
15
17
|
```bash
|
|
16
18
|
bundle install
|
|
17
19
|
```
|
|
@@ -46,7 +48,7 @@ client = Ollama::Client.new
|
|
|
46
48
|
|
|
47
49
|
# Defaults:
|
|
48
50
|
# - base_url: "http://localhost:11434"
|
|
49
|
-
# - model: "
|
|
51
|
+
# - model: "qwen3.5:4b"
|
|
50
52
|
# - timeout: 20 seconds
|
|
51
53
|
# - retries: 2
|
|
52
54
|
# - temperature: 0.2
|
|
@@ -82,14 +84,16 @@ To use [Ollama Cloud](https://docs.ollama.com/cloud) models (hosted at ollama.co
|
|
|
82
84
|
```ruby
|
|
83
85
|
config = Ollama::Config.new
|
|
84
86
|
config.base_url = "https://ollama.com"
|
|
85
|
-
config.
|
|
87
|
+
config.api_keys = ENV["OLLAMA_API_KEYS"] # optional comma-separated key pool
|
|
88
|
+
config.api_key = ENV["OLLAMA_API_KEY"] if config.api_keys.empty? # single-key fallback
|
|
89
|
+
config.enable_multi_key_concurrency = Ollama::Config.truthy_env?(ENV["ENABLE_MULTI_KEY_CONCURRENCY"])
|
|
86
90
|
client = Ollama::Client.new(config: config)
|
|
87
91
|
|
|
88
92
|
# Use a cloud model (e.g. gpt-oss:120b-cloud)
|
|
89
93
|
client.chat(messages: [{ role: "user", content: "Why is the sky blue?" }], model: "gpt-oss:120b-cloud")
|
|
90
94
|
```
|
|
91
95
|
|
|
92
|
-
All requests will send `Authorization: Bearer <api_key>` and use HTTPS. The same client works for chat, generate, embeddings, and model listing.
|
|
96
|
+
All requests will send `Authorization: Bearer <api_key>` and use HTTPS. When multiple keys are configured, HTTP 429 responses rotate to the next key; if the full pool remains rate-limited through all retry cycles, the client raises `Ollama::RateLimitExhaustedError`. The same client works for chat, generate, embeddings, and model listing.
|
|
93
97
|
|
|
94
98
|
### Option D: Client from Environment Variables
|
|
95
99
|
|
|
@@ -98,7 +102,9 @@ The gem automatically loads `.env` file. You can set these environment variables
|
|
|
98
102
|
```bash
|
|
99
103
|
# In your .env file or shell environment
|
|
100
104
|
OLLAMA_BASE_URL=http://localhost:11434
|
|
101
|
-
OLLAMA_API_KEY=your_key_for_ollama_cloud # optional, for https://ollama.com
|
|
105
|
+
OLLAMA_API_KEY=your_key_for_ollama_cloud # optional single-key fallback, for https://ollama.com
|
|
106
|
+
OLLAMA_API_KEYS=key_abc123,key_xyz789 # optional multi-key pool, takes precedence
|
|
107
|
+
ENABLE_MULTI_KEY_CONCURRENCY=false # optional; true/1 round-robins initial keys across threads
|
|
102
108
|
OLLAMA_MODEL=qwen2.5:14b
|
|
103
109
|
OLLAMA_TEMPERATURE=0.1
|
|
104
110
|
```
|
|
@@ -111,7 +117,9 @@ require "ollama_client"
|
|
|
111
117
|
# Create config and read from environment
|
|
112
118
|
config = Ollama::Config.new
|
|
113
119
|
config.base_url = ENV["OLLAMA_BASE_URL"] if ENV["OLLAMA_BASE_URL"]
|
|
114
|
-
config.
|
|
120
|
+
config.api_keys = ENV["OLLAMA_API_KEYS"] if ENV["OLLAMA_API_KEYS"]
|
|
121
|
+
config.api_key = ENV["OLLAMA_API_KEY"] if config.api_keys.empty? && ENV["OLLAMA_API_KEY"]
|
|
122
|
+
config.enable_multi_key_concurrency = Ollama::Config.truthy_env?(ENV["ENABLE_MULTI_KEY_CONCURRENCY"])
|
|
115
123
|
config.model = ENV["OLLAMA_MODEL"] if ENV["OLLAMA_MODEL"]
|
|
116
124
|
config.temperature = ENV["OLLAMA_TEMPERATURE"].to_f if ENV["OLLAMA_TEMPERATURE"]
|
|
117
125
|
|
|
@@ -125,7 +133,7 @@ Create a `config.json` file:
|
|
|
125
133
|
```json
|
|
126
134
|
{
|
|
127
135
|
"base_url": "http://localhost:11434",
|
|
128
|
-
"model": "
|
|
136
|
+
"model": "qwen3.5:4b",
|
|
129
137
|
"timeout": 30,
|
|
130
138
|
"retries": 3,
|
|
131
139
|
"temperature": 0.2,
|
|
@@ -329,7 +337,7 @@ require "ollama_client"
|
|
|
329
337
|
# Step 2: Create client (using environment variables from .env)
|
|
330
338
|
config = Ollama::Config.new
|
|
331
339
|
config.base_url = ENV["OLLAMA_BASE_URL"] || "http://localhost:11434"
|
|
332
|
-
config.model = ENV["OLLAMA_MODEL"] || "
|
|
340
|
+
config.model = ENV["OLLAMA_MODEL"] || "qwen3.5:4b"
|
|
333
341
|
config.temperature = ENV["OLLAMA_TEMPERATURE"].to_f if ENV["OLLAMA_TEMPERATURE"]
|
|
334
342
|
|
|
335
343
|
client = Ollama::Client.new(config: config)
|
|
@@ -361,7 +369,7 @@ end
|
|
|
361
369
|
| Option | Default | Description |
|
|
362
370
|
|--------|---------|-------------|
|
|
363
371
|
| `base_url` | `"http://localhost:11434"` | Ollama server URL |
|
|
364
|
-
| `model` | `"
|
|
372
|
+
| `model` | `"qwen3.5:4b"` | Default model to use |
|
|
365
373
|
| `timeout` | `20` | Request timeout in seconds |
|
|
366
374
|
| `retries` | `2` | Number of retry attempts on failure |
|
|
367
375
|
| `temperature` | `0.2` | Model temperature (0.0-2.0) |
|
data/docs/INTEGRATION_TESTING.md
CHANGED
|
@@ -10,7 +10,7 @@ bundle exec rspec --tag integration
|
|
|
10
10
|
|
|
11
11
|
# Run with custom configuration
|
|
12
12
|
OLLAMA_URL=http://localhost:11434 \
|
|
13
|
-
OLLAMA_MODEL=
|
|
13
|
+
OLLAMA_MODEL=qwen3.5:4b \
|
|
14
14
|
bundle exec rspec --tag integration
|
|
15
15
|
```
|
|
16
16
|
|
|
@@ -27,18 +27,21 @@ Environment variables are documented in the header of `script/live_branch_smoke.
|
|
|
27
27
|
## Prerequisites
|
|
28
28
|
|
|
29
29
|
1. **Ollama server running** (default: `http://localhost:11434`)
|
|
30
|
+
|
|
30
31
|
```bash
|
|
31
32
|
# Start Ollama server
|
|
32
33
|
ollama serve
|
|
33
34
|
```
|
|
34
35
|
|
|
35
36
|
2. **At least one model installed**
|
|
37
|
+
|
|
36
38
|
```bash
|
|
37
39
|
# Install a model
|
|
38
|
-
ollama pull
|
|
40
|
+
ollama pull qwen3.5:4b
|
|
39
41
|
```
|
|
40
42
|
|
|
41
43
|
3. **Optional: Embedding model** (for embedding tests)
|
|
44
|
+
|
|
42
45
|
```bash
|
|
43
46
|
ollama pull nomic-embed-text:latest
|
|
44
47
|
```
|
|
@@ -48,6 +51,7 @@ Environment variables are documented in the header of `script/live_branch_smoke.
|
|
|
48
51
|
Integration tests verify:
|
|
49
52
|
|
|
50
53
|
### ✅ Core Client Methods
|
|
54
|
+
|
|
51
55
|
- `#list_models` - Lists available models
|
|
52
56
|
- `#generate` - Structured JSON output with schema
|
|
53
57
|
- `#generate` - Plain text output without schema
|
|
@@ -57,54 +61,62 @@ Integration tests verify:
|
|
|
57
61
|
- `#chat_raw` - Tool calling (if model supports it)
|
|
58
62
|
|
|
59
63
|
### ✅ Embeddings
|
|
64
|
+
|
|
60
65
|
- Single text embeddings
|
|
61
66
|
- Multiple text embeddings
|
|
62
67
|
- Error handling for missing/unsupported models
|
|
63
68
|
|
|
64
69
|
### ✅ Agent Components
|
|
70
|
+
|
|
65
71
|
- `Ollama::Agent::Planner` - Planning decisions
|
|
66
72
|
- `Ollama::Agent::Executor` - Tool execution loops (if model supports tools)
|
|
67
73
|
|
|
68
74
|
### ✅ Chat Session
|
|
75
|
+
|
|
69
76
|
- Session management
|
|
70
77
|
- Conversation state
|
|
71
78
|
- Message history
|
|
72
79
|
|
|
73
80
|
### ✅ Error Handling
|
|
81
|
+
|
|
74
82
|
- `NotFoundError` for non-existent models
|
|
75
83
|
- Proper error propagation
|
|
76
84
|
|
|
77
85
|
## Running Tests
|
|
78
86
|
|
|
79
87
|
### Run All Integration Tests
|
|
88
|
+
|
|
80
89
|
```bash
|
|
81
90
|
bundle exec rspec --tag integration
|
|
82
91
|
```
|
|
83
92
|
|
|
84
93
|
### Run Specific Test File
|
|
94
|
+
|
|
85
95
|
```bash
|
|
86
96
|
bundle exec rspec spec/integration/ollama_client_integration_spec.rb --tag integration
|
|
87
97
|
```
|
|
88
98
|
|
|
89
99
|
### Run Specific Test
|
|
100
|
+
|
|
90
101
|
```bash
|
|
91
102
|
bundle exec rspec spec/integration/ollama_client_integration_spec.rb:32 --tag integration
|
|
92
103
|
```
|
|
93
104
|
|
|
94
105
|
### With Environment Variables
|
|
106
|
+
|
|
95
107
|
```bash
|
|
96
108
|
# Custom Ollama URL
|
|
97
109
|
OLLAMA_URL=http://remote-server:11434 bundle exec rspec --tag integration
|
|
98
110
|
|
|
99
111
|
# Custom model
|
|
100
|
-
OLLAMA_MODEL=
|
|
112
|
+
OLLAMA_MODEL=qwen3.5:4b bundle exec rspec --tag integration
|
|
101
113
|
|
|
102
114
|
# Custom embedding model
|
|
103
115
|
OLLAMA_EMBEDDING_MODEL=nomic-embed-text:latest bundle exec rspec --tag integration
|
|
104
116
|
|
|
105
117
|
# All together
|
|
106
118
|
OLLAMA_URL=http://localhost:11434 \
|
|
107
|
-
OLLAMA_MODEL=
|
|
119
|
+
OLLAMA_MODEL=qwen3.5:4b \
|
|
108
120
|
OLLAMA_EMBEDDING_MODEL=nomic-embed-text:latest \
|
|
109
121
|
bundle exec rspec --tag integration
|
|
110
122
|
```
|
|
@@ -112,11 +124,13 @@ bundle exec rspec --tag integration
|
|
|
112
124
|
## Test Behavior
|
|
113
125
|
|
|
114
126
|
### Automatic Skipping
|
|
127
|
+
|
|
115
128
|
- Tests automatically skip if Ollama server is not available
|
|
116
129
|
- Tests skip if required models are not installed
|
|
117
130
|
- Tests skip if models don't support certain features (e.g., tool calling)
|
|
118
131
|
|
|
119
132
|
### Expected Results
|
|
133
|
+
|
|
120
134
|
- **Passing**: Client correctly communicates with Ollama
|
|
121
135
|
- **Pending/Skipped**: Expected when models/features unavailable
|
|
122
136
|
- **Failing**: Indicates actual client issues (rare)
|
|
@@ -143,20 +157,24 @@ bundle exec rspec --tag integration
|
|
|
143
157
|
## Troubleshooting
|
|
144
158
|
|
|
145
159
|
### "Ollama server not available"
|
|
160
|
+
|
|
146
161
|
- Start Ollama: `ollama serve`
|
|
147
162
|
- Check URL: `OLLAMA_URL=http://localhost:11434`
|
|
148
163
|
- Verify connection: `curl http://localhost:11434/api/tags`
|
|
149
164
|
|
|
150
165
|
### "Model not found"
|
|
151
|
-
|
|
166
|
+
|
|
167
|
+
- Install model: `ollama pull qwen3.5:4b`
|
|
152
168
|
- Set model: `OLLAMA_MODEL=your-model`
|
|
153
169
|
|
|
154
170
|
### "Empty embedding returned"
|
|
171
|
+
|
|
155
172
|
- Install embedding model: `ollama pull nomic-embed-text:latest`
|
|
156
173
|
- Verify model supports embeddings
|
|
157
174
|
- Set model: `OLLAMA_EMBEDDING_MODEL=nomic-embed-text:latest`
|
|
158
175
|
|
|
159
176
|
### "HTTP 400: Bad Request" (tool calling)
|
|
177
|
+
|
|
160
178
|
- Some models don't support tool calling
|
|
161
179
|
- Test will skip automatically
|
|
162
180
|
- Try a different model that supports tools
|
data/docs/PRODUCTION_FIXES.md
CHANGED
|
@@ -9,8 +9,9 @@ This document summarizes all critical fixes applied to make `ollama-client` prod
|
|
|
9
9
|
**Issue:** LLMs can return JSON wrapped in markdown, prefixed with text, or with unicode garbage.
|
|
10
10
|
|
|
11
11
|
**Fix:** Enhanced `parse_json_response()` to:
|
|
12
|
+
|
|
12
13
|
- Handle JSON arrays (not just objects)
|
|
13
|
-
- Strip markdown code fences (```json
|
|
14
|
+
- Strip markdown code fences (```json ...```)
|
|
14
15
|
- Normalize unicode and whitespace
|
|
15
16
|
- Extract nested JSON if first attempt fails
|
|
16
17
|
- Better error messages with context
|
|
@@ -22,6 +23,7 @@ This document summarizes all critical fixes applied to make `ollama-client` prod
|
|
|
22
23
|
**Issue:** Need explicit retry rules for different HTTP status codes.
|
|
23
24
|
|
|
24
25
|
**Fix:** Made retry policy explicit and documented:
|
|
26
|
+
|
|
25
27
|
- **Retry:** 408 (Request Timeout), 429 (Too Many Requests), 500 (Internal Server Error), 503 (Service Unavailable)
|
|
26
28
|
- **Never retry:** All other 4xx and 5xx errors
|
|
27
29
|
|
|
@@ -32,11 +34,13 @@ This document summarizes all critical fixes applied to make `ollama-client` prod
|
|
|
32
34
|
**Issue:** Global configuration is not thread-safe, but this wasn't clearly communicated.
|
|
33
35
|
|
|
34
36
|
**Fix:**
|
|
37
|
+
|
|
35
38
|
- Added warnings in `OllamaClient.configure` when used from multiple threads
|
|
36
39
|
- Added documentation comments in `Config` class
|
|
37
40
|
- Warns users to use per-client configuration for concurrent agents
|
|
38
41
|
|
|
39
42
|
**Location:**
|
|
43
|
+
|
|
40
44
|
- `lib/ollama_client.rb:14-16`
|
|
41
45
|
- `lib/ollama/config.rb:7-15`
|
|
42
46
|
|
|
@@ -45,6 +49,7 @@ This document summarizes all critical fixes applied to make `ollama-client` prod
|
|
|
45
49
|
**Issue:** Schemas allow extra properties by default, letting LLMs add unexpected fields.
|
|
46
50
|
|
|
47
51
|
**Fix:** Enforce `additionalProperties: false` by default:
|
|
52
|
+
|
|
48
53
|
- Automatically adds `additionalProperties: false` to object schemas
|
|
49
54
|
- Recursively applies to nested objects and array items
|
|
50
55
|
- Only if not explicitly set (allows opt-out if needed)
|
|
@@ -56,6 +61,7 @@ This document summarizes all critical fixes applied to make `ollama-client` prod
|
|
|
56
61
|
**Issue:** `chat()` should require explicit opt-in for agent usage.
|
|
57
62
|
|
|
58
63
|
**Fix:**
|
|
64
|
+
|
|
59
65
|
- Added `strict:` parameter to `chat()`
|
|
60
66
|
- Warns when `strict: false` (default)
|
|
61
67
|
- In strict mode, doesn't retry on schema violations
|
|
@@ -68,6 +74,7 @@ This document summarizes all critical fixes applied to make `ollama-client` prod
|
|
|
68
74
|
**Issue:** Need a variant that fails fast on schema violations without retries.
|
|
69
75
|
|
|
70
76
|
**Fix:** Added `generate_strict!()` method:
|
|
77
|
+
|
|
71
78
|
- No retries on schema violations
|
|
72
79
|
- Immediate failure for guaranteed contract enforcement
|
|
73
80
|
- Useful for strict agent contracts
|
|
@@ -79,10 +86,12 @@ This document summarizes all critical fixes applied to make `ollama-client` prod
|
|
|
79
86
|
**Issue:** No way to track latency, attempts, or model used per request.
|
|
80
87
|
|
|
81
88
|
**Fix:** Added `include_meta:` parameter to `generate()` and `chat()`:
|
|
89
|
+
|
|
82
90
|
- Returns `{ "data": ..., "meta": { "latency_ms": ..., "model": ..., "attempts": ... } }`
|
|
83
91
|
- Enables logging, metrics, and debugging
|
|
84
92
|
|
|
85
93
|
**Location:**
|
|
94
|
+
|
|
86
95
|
- `lib/ollama/client.rb:58-86` (generate)
|
|
87
96
|
- `lib/ollama/client.rb:27-70` (chat)
|
|
88
97
|
|
|
@@ -91,6 +100,7 @@ This document summarizes all critical fixes applied to make `ollama-client` prod
|
|
|
91
100
|
**Issue:** No way to verify Ollama server is reachable before making requests.
|
|
92
101
|
|
|
93
102
|
**Fix:** Added `health()` method:
|
|
103
|
+
|
|
94
104
|
- Returns `{ status: "healthy|unhealthy", latency_ms: ..., error: ... }`
|
|
95
105
|
- Short timeout (5s) for quick health checks
|
|
96
106
|
- Useful for auto-restart systems
|
|
@@ -100,6 +110,7 @@ This document summarizes all critical fixes applied to make `ollama-client` prod
|
|
|
100
110
|
## 📊 Impact
|
|
101
111
|
|
|
102
112
|
These fixes address **~70% of real-world Ollama failures** by:
|
|
113
|
+
|
|
103
114
|
- Robust JSON extraction (handles model quirks)
|
|
104
115
|
- Explicit retry policies (prevents wasted retries)
|
|
105
116
|
- Strict schemas (catches unexpected fields)
|
|
@@ -133,7 +144,7 @@ result = client.generate(
|
|
|
133
144
|
|
|
134
145
|
puts result["meta"]["latency_ms"] # 245.32
|
|
135
146
|
puts result["meta"]["attempts"] # 1
|
|
136
|
-
puts result["meta"]["model"] # "
|
|
147
|
+
puts result["meta"]["model"] # "qwen3.5:4b"
|
|
137
148
|
```
|
|
138
149
|
|
|
139
150
|
### Health Check
|
|
@@ -162,11 +173,12 @@ result = client.generate_strict!(
|
|
|
162
173
|
**Breaking Changes:** None - all changes are backward compatible.
|
|
163
174
|
|
|
164
175
|
**New Defaults:**
|
|
176
|
+
|
|
165
177
|
- Schemas now reject extra properties by default (can opt-out by setting `additionalProperties: true`)
|
|
166
178
|
- `chat()` now warns unless `strict: true` is passed
|
|
167
179
|
|
|
168
180
|
**Recommended Updates:**
|
|
181
|
+
|
|
169
182
|
- Use `include_meta: true` for production logging
|
|
170
183
|
- Use per-client config for concurrent agents
|
|
171
184
|
- Use `generate_strict!` when you need guaranteed contracts
|
|
172
|
-
|
data/docs/QUICK_START.md
CHANGED
|
@@ -12,7 +12,7 @@ client = Ollama::Client.new
|
|
|
12
12
|
|
|
13
13
|
# Or with custom config
|
|
14
14
|
config = Ollama::Config.new
|
|
15
|
-
config.model = ENV["OLLAMA_MODEL"] || "
|
|
15
|
+
config.model = ENV["OLLAMA_MODEL"] || "qwen3.5:4b"
|
|
16
16
|
config.base_url = ENV["OLLAMA_BASE_URL"] || "http://localhost:11434"
|
|
17
17
|
client = Ollama::Client.new(config: config)
|
|
18
18
|
```
|
data/docs/README.md
CHANGED
|
@@ -9,6 +9,7 @@ This directory contains internal development documentation for the ollama-client
|
|
|
9
9
|
## Contents
|
|
10
10
|
|
|
11
11
|
### Design Documentation
|
|
12
|
+
- **[RUBYLLM_ADOPTION_MATRIX.md](RUBYLLM_ADOPTION_MATRIX.md)** - Feature-by-feature comparison against RubyLLM's public API, with copy/improve/reject/ecosystem verdicts and a prioritized build order
|
|
12
13
|
- **[HANDLERS_ANALYSIS.md](HANDLERS_ANALYSIS.md)** - Analysis of handler architecture decisions (why we didn't adopt ollama-ruby's handler pattern)
|
|
13
14
|
- **[FEATURES_ADDED.md](FEATURES_ADDED.md)** - Features integrated from ollama-ruby that align with our agent-first philosophy
|
|
14
15
|
- **[PRODUCTION_FIXES.md](PRODUCTION_FIXES.md)** - Production-ready fixes for hybrid agents (JSON parsing, retry policy, etc.)
|