ollama-client 1.3.0 → 1.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (178) hide show
  1. checksums.yaml +4 -4
  2. data/.env.example +11 -0
  3. data/API_CONTRACT.md +79 -8
  4. data/CHANGELOG.md +45 -0
  5. data/CONTRIBUTING.md +8 -0
  6. data/README.md +205 -273
  7. data/ROADMAP.md +23 -0
  8. data/docs/API_GAPS.md +18 -141
  9. data/docs/ARCHITECTURE.md +141 -0
  10. data/docs/AREAS_FOR_CONSIDERATION.md +13 -1
  11. data/docs/CLOUD.md +19 -0
  12. data/docs/CONSOLE_IMPROVEMENTS.md +20 -1
  13. data/docs/ECOSYSTEM_GEMS.md +41 -0
  14. data/docs/ECOSYSTEM_STRATEGY.md +151 -0
  15. data/docs/GETTING_STARTED.md +16 -8
  16. data/docs/INTEGRATION_TESTING.md +23 -5
  17. data/docs/PRODUCTION_FIXES.md +15 -3
  18. data/docs/QUICK_START.md +1 -1
  19. data/docs/README.md +1 -0
  20. data/docs/RUBYLLM_ADOPTION_MATRIX.md +273 -0
  21. data/docs/adr/001-openai-boundary.md +12 -0
  22. data/docs/adr/002-transport-abstraction.md +12 -0
  23. data/docs/adr/003-response-normalization.md +12 -0
  24. data/docs/adr/004-mock-transport.md +16 -0
  25. data/docs/adr/005-error-taxonomy.md +16 -0
  26. data/docs/adr/006-stream-runtime.md +16 -0
  27. data/docs/ecosystem/BOUNDARIES.md +19 -0
  28. data/docs/ecosystem/DEPENDENCY_GRAPH.md +14 -0
  29. data/docs/ecosystem/DESIGN_PRINCIPLES.md +9 -0
  30. data/docs/ecosystem/EXISTING_REPOS.md +31 -0
  31. data/docs/ecosystem/EXPERIMENTAL_LABS.md +24 -0
  32. data/docs/ecosystem/OVERVIEW.md +28 -0
  33. data/docs/ecosystem/RELEASE_ORDER.md +17 -0
  34. data/docs/observability/README.md +8 -0
  35. data/docs/rails/README.md +7 -0
  36. data/docs/rfcs/0001-stream-runtime.md +6 -0
  37. data/docs/rfcs/0002-schema-system.md +6 -0
  38. data/docs/rfcs/0003-observability-hooks.md +6 -0
  39. data/docs/rfcs/0004-async-runtime.md +6 -0
  40. data/docs/rfcs/README.md +16 -0
  41. data/docs/runtime/ERROR_CONTRACT.md +14 -0
  42. data/docs/runtime/SCHEMA_CONTRACT.md +28 -0
  43. data/docs/runtime/STREAM_CONTRACT.md +17 -0
  44. data/docs/runtime/STREAM_RUNTIME.md +20 -0
  45. data/docs/runtime/TRANSPORT_CONTRACT.md +26 -0
  46. data/docs/schema/README.md +8 -0
  47. data/docs/schema/STRUCTURED_OUTPUTS.md +22 -0
  48. data/docs/streaming/README.md +8 -0
  49. data/docs/testing/README.md +7 -0
  50. data/docs/testing/REPLAY_SYSTEM.md +17 -0
  51. data/docs/transport/README.md +7 -0
  52. data/examples/README.md +44 -0
  53. data/examples/basic_chat.rb +9 -0
  54. data/examples/cloud_models.rb +162 -0
  55. data/examples/embeddings.rb +10 -0
  56. data/examples/free_catalog.json +290 -0
  57. data/examples/generate.rb +9 -0
  58. data/examples/streaming.rb +12 -0
  59. data/examples/structured_tools.rb +90 -0
  60. data/examples/tool_calling_direct.rb +101 -0
  61. data/examples/tool_dto_example.rb +94 -0
  62. data/exe/ollama-client +5 -1
  63. data/lib/ollama/agent/executor.rb +243 -0
  64. data/lib/ollama/agent/messages.rb +31 -0
  65. data/lib/ollama/agent/planner.rb +45 -0
  66. data/lib/ollama/api_key_pool.rb +61 -0
  67. data/lib/ollama/attachment.rb +74 -0
  68. data/lib/ollama/capabilities.rb +1 -1
  69. data/lib/ollama/chat_response.rb +33 -0
  70. data/lib/ollama/client/chat/request_preparer.rb +82 -0
  71. data/lib/ollama/client/chat.rb +58 -68
  72. data/lib/ollama/client/chat_stream_processor.rb +85 -25
  73. data/lib/ollama/client/generate/request_preparer.rb +118 -0
  74. data/lib/ollama/client/generate/response_formatter.rb +95 -0
  75. data/lib/ollama/client/generate.rb +83 -165
  76. data/lib/ollama/client/model_management.rb +182 -84
  77. data/lib/ollama/client/openai_compat.rb +185 -0
  78. data/lib/ollama/client/raw.rb +66 -0
  79. data/lib/ollama/client/tool_intent.rb +38 -0
  80. data/lib/ollama/client/web_search.rb +39 -0
  81. data/lib/ollama/client.rb +83 -35
  82. data/lib/ollama/config.rb +122 -17
  83. data/lib/ollama/embeddings.rb +67 -29
  84. data/lib/ollama/errors.rb +55 -1
  85. data/lib/ollama/events.rb +55 -0
  86. data/lib/ollama/generate_stream_handler.rb +23 -4
  87. data/lib/ollama/http_error_handler.rb +42 -0
  88. data/lib/ollama/messages.rb +109 -0
  89. data/lib/ollama/middleware/cache.rb +73 -0
  90. data/lib/ollama/middleware/logger.rb +74 -0
  91. data/lib/ollama/middleware/metrics.rb +84 -0
  92. data/lib/ollama/middleware/tracing.rb +101 -0
  93. data/lib/ollama/middleware.rb +39 -0
  94. data/lib/ollama/model_profile.rb +1 -1
  95. data/lib/ollama/openai.rb +16 -0
  96. data/lib/ollama/options.rb +63 -1
  97. data/lib/ollama/params.rb +139 -0
  98. data/lib/ollama/parsers/base.rb +24 -0
  99. data/lib/ollama/parsers/chat.rb +22 -0
  100. data/lib/ollama/parsers/embeddings.rb +23 -0
  101. data/lib/ollama/parsers/generate.rb +38 -0
  102. data/lib/ollama/parsers/list_running.rb +15 -0
  103. data/lib/ollama/parsers/show_model.rb +14 -0
  104. data/lib/ollama/parsers/version.rb +15 -0
  105. data/lib/ollama/pipeline.rb +169 -0
  106. data/lib/ollama/plugins.rb +83 -0
  107. data/lib/ollama/policies/auto_pull.rb +69 -0
  108. data/lib/ollama/policies/base.rb +70 -0
  109. data/lib/ollama/policies/capability_validation.rb +88 -0
  110. data/lib/ollama/policies/fallback.rb +57 -0
  111. data/lib/ollama/policies/rate_limit.rb +115 -0
  112. data/lib/ollama/policies/repair_json.rb +137 -0
  113. data/lib/ollama/policies/retry/strategies/exponential.rb +27 -0
  114. data/lib/ollama/policies/retry/strategies/fixed.rb +27 -0
  115. data/lib/ollama/policies/retry/strategies/jitter.rb +30 -0
  116. data/lib/ollama/policies/retry/strategies/linear.rb +27 -0
  117. data/lib/ollama/policies/retry.rb +152 -0
  118. data/lib/ollama/policies/schema_repair.rb +180 -0
  119. data/lib/ollama/policies/timeout.rb +43 -0
  120. data/lib/ollama/policies.rb +24 -0
  121. data/lib/ollama/prompt.rb +86 -0
  122. data/lib/ollama/prompt_adapters/base.rb +3 -2
  123. data/lib/ollama/prompt_adapters/gemma4.rb +22 -14
  124. data/lib/ollama/prompts/tool_planner.rb +35 -0
  125. data/lib/ollama/providers/base.rb +70 -0
  126. data/lib/ollama/providers/llama_cpp.rb +131 -0
  127. data/lib/ollama/providers/ollama.rb +54 -0
  128. data/lib/ollama/providers/openai.rb +134 -0
  129. data/lib/ollama/providers.rb +28 -0
  130. data/lib/ollama/rate_limit_handler.rb +48 -0
  131. data/lib/ollama/request.rb +169 -0
  132. data/lib/ollama/response.rb +3 -2
  133. data/lib/ollama/responses/base.rb +11 -0
  134. data/lib/ollama/responses/chat.rb +11 -0
  135. data/lib/ollama/responses/embeddings.rb +14 -0
  136. data/lib/ollama/responses/generate.rb +22 -0
  137. data/lib/ollama/schema_dsl.rb +97 -0
  138. data/lib/ollama/schema_validator.rb +90 -53
  139. data/lib/ollama/schemas/tool_intent.json +15 -0
  140. data/lib/ollama/schemas/tool_intent.rb +16 -0
  141. data/lib/ollama/serializers/base.rb +44 -0
  142. data/lib/ollama/serializers/chat.rb +42 -0
  143. data/lib/ollama/serializers/embeddings.rb +36 -0
  144. data/lib/ollama/serializers/generate.rb +47 -0
  145. data/lib/ollama/streaming_observer.rb +22 -0
  146. data/lib/ollama/testing.rb +103 -0
  147. data/lib/ollama/tool/function/parameters/property.rb +72 -0
  148. data/lib/ollama/tool/function/parameters.rb +101 -0
  149. data/lib/ollama/tool/function.rb +78 -0
  150. data/lib/ollama/tool.rb +60 -0
  151. data/lib/ollama/tool_dsl.rb +93 -0
  152. data/lib/ollama/tool_intent.rb +19 -0
  153. data/lib/ollama/transport/base.rb +42 -0
  154. data/lib/ollama/transport/mock.rb +48 -0
  155. data/lib/ollama/transport/net_http.rb +76 -0
  156. data/lib/ollama/transport/request.rb +20 -0
  157. data/lib/ollama/transport/response.rb +41 -0
  158. data/lib/ollama/transport.rb +26 -0
  159. data/lib/ollama/version.rb +1 -1
  160. data/lib/ollama_client.rb +35 -0
  161. data/script/live_branch_smoke/chat_exercises.rb +232 -0
  162. data/script/live_branch_smoke/generation_exercises.rb +128 -0
  163. data/script/live_branch_smoke/model_exercises.rb +105 -0
  164. data/script/live_branch_smoke/utility_exercises.rb +375 -0
  165. data/script/live_branch_smoke.rb +159 -174
  166. data/test_all_features.rb +315 -0
  167. metadata +158 -24
  168. data/.cursor/.gitignore +0 -1
  169. data/RELEASE_NOTES_v0.2.6.md +0 -41
  170. data/devagent_proper.rb +0 -430
  171. data/docs/TESTING.md +0 -508
  172. data/examples/agent_loop.rb +0 -120
  173. data/examples/failure_modes/invalid_json_repair.rb +0 -42
  174. data/examples/production/rails_agent.rb +0 -62
  175. data/market.jpg +0 -0
  176. data/print_capabilities.rb +0 -20
  177. data/schema.json +0 -1
  178. data/test_tool.rb +0 -26
@@ -0,0 +1,151 @@
1
+ # Ollama Ruby Ecosystem Strategy & Architectural Blueprint
2
+
3
+ This document defines the overarching architectural strategy, dependency contracts, and governance model for the Ollama Ruby ecosystem. It establishes the foundational guardrails ensuring that as the ecosystem expands across multiple specialized gems, it maintains strict determinism, modularity, and architectural integrity.
4
+
5
+ ---
6
+
7
+ ## 1. Executive Summary & Core Philosophy
8
+
9
+ The Ollama Ruby ecosystem is designed around a **deterministic kernel** (`ollama-client`) surrounded by **modular companion gems**.
10
+
11
+ ### The Core Philosophy
12
+ 1. **Deterministic Kernel**: `ollama-client` is the canonical source of truth for all low-level transport, retry semantics, connection pooling, raw endpoint definitions, and error taxonomy.
13
+ 2. **Composition over Inheritance**: Companion gems (`ollama-openai`, `ollama-observability`, `ollama-stream`, `ollama-agent`, `ollama-rails`) build upon the core client through well-defined public interfaces and callback hooks rather than monkey-patching or redefining internal transport logic.
14
+ 3. **Domain Isolation**: Each companion gem encapsulates a single, cohesive domain (e.g., OpenAI compatibility, OpenTelemetry instrumentation, advanced streaming, agent orchestration, Rails integration).
15
+
16
+ ---
17
+
18
+ ## 2. The Dependency Hierarchy & Directionality
19
+
20
+ To prevent architectural inversion and circular dependencies, all gems in the ecosystem must adhere strictly to the following unidirectional dependency graph:
21
+
22
+ ```
23
+ ┌────────────────────────────────────────────────────────┐
24
+ │ ollama-client (Core Kernel) │
25
+ │ (Transport, Retries, Error Taxonomy, Raw Schemas) │
26
+ └───────────────────────────▲────────────────────────────┘
27
+ │
28
+ ┌──────────────────┼──────────────────┐
29
+ │ │ │
30
+ ┌────────┴────────┐┌────────┴────────┐┌────────┴────────┐
31
+ │ ollama-openai ││ollama-observab. ││ ollama-stream │
32
+ │(OpenAI Facade) ││ (OTel Telemetry)││ (SSE & WebSockets│
33
+ └────────▲────────┘└─────────────────┘└────────▲────────┘
34
+ │ │
35
+ └──────────────────┬──────────────────┘
36
+ │
37
+ ┌──────────┴──────────┐
38
+ │ ollama-agent │
39
+ │(Tools, Memory, Plan)│
40
+ └──────────▲──────────┘
41
+ │
42
+ ┌──────────┴──────────┐
43
+ │ ollama-rails │
44
+ │(ActiveJob, Turbo, │
45
+ │ ActionCable) │
46
+ └─────────────────────┘
47
+ ```
48
+
49
+ ### Critical Dependency Rules
50
+ - **Rule 1**: `ollama-client` must NEVER depend on any companion gem or higher-level concept (e.g., agents, Rails, OpenTelemetry).
51
+ - **Rule 2**: Companion gems must declare `ollama-client` as their primary dependency and utilize its public API or hook system.
52
+ - **Rule 3**: Higher-level orchestration gems (`ollama-agent`, `ollama-rails`) may compose multiple lower-level companion gems (e.g., `ollama-agent` utilizing `ollama-stream` for tool streaming or `ollama-openai` for LLM interop).
53
+
54
+ ---
55
+
56
+ ## 3. Companion Gem Specifications & Responsibilities
57
+
58
+ ### 1. `ollama-client` (The Kernel)
59
+ - **Role**: Canonical infrastructure layer.
60
+ - **Responsibilities**: Faraday HTTP/Faraday WebSocket transport, connection pooling, exponential backoff retries, raw API endpoint mapping (`/api/chat`, `/api/generate`, `/api/embeddings`, `/api/pull`), configuration management (`Ollama::Config`), and canonical error taxonomy (`Ollama::Error`, `Ollama::ConnectionError`, `Ollama::TimeoutError`, `Ollama::RateLimitError`).
61
+ - **Prohibited**: High-level workflow orchestration, third-party framework coupling.
62
+
63
+ ### 2. `ollama-openai` (Protocol Interop)
64
+ - **Role**: Drop-in OpenAI compatibility layer.
65
+ - **Responsibilities**: Translating OpenAI-style requests (`client.chat(parameters: {})`) into native Ollama payloads, mapping OpenAI function definitions to Ollama tool schemas, normalizing Ollama responses into OpenAI JSON structures (`chatcmpl-*`), and companion support for frameworks like LangChain, Vercel AI SDK, and ruby-openai consumers.
66
+ - **Prohibited**: Custom HTTP transport implementations.
67
+
68
+ ### 3. `ollama-observability` (Telemetry & Logging)
69
+ - **Role**: Enterprise-grade observability layer.
70
+ - **Responsibilities**: Subscribing to `ollama-client` hooks (`on_response`, `on_token`, `on_error`) to generate OpenTelemetry tracing spans, tracking metrics (Time-to-First-Token, latency histograms, prompt/completion token counters), and emitting structured JSON logs. Supports payload redaction for PII/PHI compliance.
71
+ - **Prohibited**: Modifying raw API responses or interfering with inference execution.
72
+
73
+ ### 4. `ollama-stream` (Advanced Streaming & Transport)
74
+ - **Role**: High-performance streaming runtime.
75
+ - **Responsibilities**: Encapsulating SSE streams into formal `Ollama::Stream::StreamObject` instances supporting pause, resume, and cancellation; managing persistent bidirectional WebSocket sessions; providing backpressure and queue bounding (`FlowController`); implementing incremental JSON fragment recovery (`IncrementalParser`); and offering a Rack-compatible SSE proxy adapter.
76
+ - **Prohibited**: Re-implementing base Faraday connection logic.
77
+
78
+ ### 5. `ollama-agent` (Orchestration & Tooling)
79
+ - **Role**: Autonomous agent execution framework.
80
+ - **Responsibilities**: Defining standardized tool contracts (`Ollama::Agent::Tool`), managing tool registries, maintaining conversation memory (`WindowMemory`, `SummaryMemory`), executing autonomous ReAct/Plan-and-Solve agent loops (`Executor`), and providing structured JSON output parsing.
81
+ - **Prohibited**: Direct HTTP transport management.
82
+
83
+ ### 6. `ollama-rails` (Rails Integration)
84
+ - **Role**: Idiomatic Ruby on Rails integration.
85
+ - **Responsibilities**: Providing Rails Railtie for zero-config initialization, wrapping async inference in ActiveJob (`Ollama::Rails::GenerateJob`, `ChatJob`), broadcasting live stream tokens over ActionCable/Turbo Streams (`Ollama::Rails::BroadcastHelpers`), offering ActiveRecord mixins (`Ollama::Rails::Embeddable` for pgvector/neighbor integration), and providing Rake tasks for model management (`ollama:pull`, `ollama:list`).
86
+ - **Prohibited**: Modifying core Ruby runtime behavior outside the Rails application context.
87
+
88
+ ---
89
+
90
+ ## 4. Architectural Guardrails & Extension Mechanisms
91
+
92
+ ### The Hook System
93
+ To enable companion gems to extend functionality without monkey-patching, `ollama-client` exposes a robust hook and middleware architecture:
94
+
95
+ ```ruby
96
+ # Example of Hook Subscription in Companion Gems
97
+ config.on_response = ->(raw_response, metadata) {
98
+ # Used by ollama-observability for metrics and logging
99
+ Telemetry.record_latency(metadata[:duration])
100
+ }
101
+
102
+ client.chat(
103
+ model: "llama3",
104
+ messages: history,
105
+ hooks: {
106
+ on_token: ->(chunk, logprobs) { # Used by ollama-stream },
107
+ on_tool_call: ->(tool_call) { # Used by ollama-agent },
108
+ on_error: ->(error) { # Used by ollama-observability }
109
+ }
110
+ )
111
+ ```
112
+
113
+ ### Error Handling & Taxonomy
114
+ All companion gems must catch base `Ollama::Error` exceptions from `ollama-client` and either allow them to bubble up or wrap them in domain-specific exceptions that inherit from `Ollama::Error` (e.g., `Ollama::Openai::Error < Ollama::Error`).
115
+
116
+ ---
117
+
118
+ ## 5. Development & Workspace Governance
119
+
120
+ ### Local Dependency Mapping
121
+ During development, all companion gems must reference the local `ollama-client` kernel in their `Gemfile` to ensure continuous integration and contract verification across the workspace:
122
+
123
+ ```ruby
124
+ # Gemfile for Companion Gems
125
+ source "https://rubygems.org"
126
+ gemspec
127
+
128
+ # Enforce local workspace dependency
129
+ gem "ollama-client", path: "../ollama-client"
130
+ ```
131
+
132
+ ### Verification & Testing
133
+ Every companion gem must maintain an independent RSpec test suite verifying its specific domain logic against mock `Ollama::Client` instances. Contract tests must ensure that changes in `ollama-client` do not break companion gem facades.
134
+
135
+ ---
136
+
137
+ ## 6. Roadmap & Future Evolution
138
+
139
+ 1. **Phase 1: Foundation (Completed)**
140
+ - Core transport, retry semantics, error taxonomy, and configuration in `ollama-client`.
141
+ - OpenAI interop facade in `ollama-openai`.
142
+ - Enterprise telemetry in `ollama-observability`.
143
+ - Advanced streaming primitives in `ollama-stream`.
144
+
145
+ 2. **Phase 2: Orchestration & Frameworks (Current)**
146
+ - Autonomous agent loops and tool calling in `ollama-agent`.
147
+ - Idiomatic Rails integrations in `ollama-rails`.
148
+
149
+ 3. **Phase 3: Advanced Optimization & Ecosystem Scaling (Future)**
150
+ - Native C/Rust extensions for high-throughput incremental JSON parsing.
151
+ - Distributed multi-node Ollama cluster balancing and routing gems.
@@ -7,11 +7,13 @@ This guide shows you step-by-step how to create a client object to use all featu
7
7
  ### Option A: Using Bundler (Recommended)
8
8
 
9
9
  Add to your `Gemfile`:
10
+
10
11
  ```ruby
11
12
  gem "ollama-client"
12
13
  ```
13
14
 
14
15
  Then run:
16
+
15
17
  ```bash
16
18
  bundle install
17
19
  ```
@@ -46,7 +48,7 @@ client = Ollama::Client.new
46
48
 
47
49
  # Defaults:
48
50
  # - base_url: "http://localhost:11434"
49
- # - model: "llama3.2:3b"
51
+ # - model: "qwen3.5:4b"
50
52
  # - timeout: 20 seconds
51
53
  # - retries: 2
52
54
  # - temperature: 0.2
@@ -82,14 +84,16 @@ To use [Ollama Cloud](https://docs.ollama.com/cloud) models (hosted at ollama.co
82
84
  ```ruby
83
85
  config = Ollama::Config.new
84
86
  config.base_url = "https://ollama.com"
85
- config.api_key = ENV["OLLAMA_API_KEY"] # or your API key
87
+ config.api_keys = ENV["OLLAMA_API_KEYS"] # optional comma-separated key pool
88
+ config.api_key = ENV["OLLAMA_API_KEY"] if config.api_keys.empty? # single-key fallback
89
+ config.enable_multi_key_concurrency = Ollama::Config.truthy_env?(ENV["ENABLE_MULTI_KEY_CONCURRENCY"])
86
90
  client = Ollama::Client.new(config: config)
87
91
 
88
92
  # Use a cloud model (e.g. gpt-oss:120b-cloud)
89
93
  client.chat(messages: [{ role: "user", content: "Why is the sky blue?" }], model: "gpt-oss:120b-cloud")
90
94
  ```
91
95
 
92
- All requests will send `Authorization: Bearer <api_key>` and use HTTPS. The same client works for chat, generate, embeddings, and model listing.
96
+ All requests will send `Authorization: Bearer <api_key>` and use HTTPS. When multiple keys are configured, HTTP 429 responses rotate to the next key; if the full pool remains rate-limited through all retry cycles, the client raises `Ollama::RateLimitExhaustedError`. The same client works for chat, generate, embeddings, and model listing.
93
97
 
94
98
  ### Option D: Client from Environment Variables
95
99
 
@@ -98,7 +102,9 @@ The gem automatically loads `.env` file. You can set these environment variables
98
102
  ```bash
99
103
  # In your .env file or shell environment
100
104
  OLLAMA_BASE_URL=http://localhost:11434
101
- OLLAMA_API_KEY=your_key_for_ollama_cloud # optional, for https://ollama.com
105
+ OLLAMA_API_KEY=your_key_for_ollama_cloud # optional single-key fallback, for https://ollama.com
106
+ OLLAMA_API_KEYS=key_abc123,key_xyz789 # optional multi-key pool, takes precedence
107
+ ENABLE_MULTI_KEY_CONCURRENCY=false # optional; true/1 round-robins initial keys across threads
102
108
  OLLAMA_MODEL=qwen2.5:14b
103
109
  OLLAMA_TEMPERATURE=0.1
104
110
  ```
@@ -111,7 +117,9 @@ require "ollama_client"
111
117
  # Create config and read from environment
112
118
  config = Ollama::Config.new
113
119
  config.base_url = ENV["OLLAMA_BASE_URL"] if ENV["OLLAMA_BASE_URL"]
114
- config.api_key = ENV["OLLAMA_API_KEY"] if ENV["OLLAMA_API_KEY"]
120
+ config.api_keys = ENV["OLLAMA_API_KEYS"] if ENV["OLLAMA_API_KEYS"]
121
+ config.api_key = ENV["OLLAMA_API_KEY"] if config.api_keys.empty? && ENV["OLLAMA_API_KEY"]
122
+ config.enable_multi_key_concurrency = Ollama::Config.truthy_env?(ENV["ENABLE_MULTI_KEY_CONCURRENCY"])
115
123
  config.model = ENV["OLLAMA_MODEL"] if ENV["OLLAMA_MODEL"]
116
124
  config.temperature = ENV["OLLAMA_TEMPERATURE"].to_f if ENV["OLLAMA_TEMPERATURE"]
117
125
 
@@ -125,7 +133,7 @@ Create a `config.json` file:
125
133
  ```json
126
134
  {
127
135
  "base_url": "http://localhost:11434",
128
- "model": "llama3.2:3b",
136
+ "model": "qwen3.5:4b",
129
137
  "timeout": 30,
130
138
  "retries": 3,
131
139
  "temperature": 0.2,
@@ -329,7 +337,7 @@ require "ollama_client"
329
337
  # Step 2: Create client (using environment variables from .env)
330
338
  config = Ollama::Config.new
331
339
  config.base_url = ENV["OLLAMA_BASE_URL"] || "http://localhost:11434"
332
- config.model = ENV["OLLAMA_MODEL"] || "llama3.2:3b"
340
+ config.model = ENV["OLLAMA_MODEL"] || "qwen3.5:4b"
333
341
  config.temperature = ENV["OLLAMA_TEMPERATURE"].to_f if ENV["OLLAMA_TEMPERATURE"]
334
342
 
335
343
  client = Ollama::Client.new(config: config)
@@ -361,7 +369,7 @@ end
361
369
  | Option | Default | Description |
362
370
  |--------|---------|-------------|
363
371
  | `base_url` | `"http://localhost:11434"` | Ollama server URL |
364
- | `model` | `"llama3.2:3b"` | Default model to use |
372
+ | `model` | `"qwen3.5:4b"` | Default model to use |
365
373
  | `timeout` | `20` | Request timeout in seconds |
366
374
  | `retries` | `2` | Number of retry attempts on failure |
367
375
  | `temperature` | `0.2` | Model temperature (0.0-2.0) |
@@ -10,7 +10,7 @@ bundle exec rspec --tag integration
10
10
 
11
11
  # Run with custom configuration
12
12
  OLLAMA_URL=http://localhost:11434 \
13
- OLLAMA_MODEL=llama3.2:3b \
13
+ OLLAMA_MODEL=qwen3.5:4b \
14
14
  bundle exec rspec --tag integration
15
15
  ```
16
16
 
@@ -27,18 +27,21 @@ Environment variables are documented in the header of `script/live_branch_smoke.
27
27
  ## Prerequisites
28
28
 
29
29
  1. **Ollama server running** (default: `http://localhost:11434`)
30
+
30
31
  ```bash
31
32
  # Start Ollama server
32
33
  ollama serve
33
34
  ```
34
35
 
35
36
  2. **At least one model installed**
37
+
36
38
  ```bash
37
39
  # Install a model
38
- ollama pull llama3.2:3b
40
+ ollama pull qwen3.5:4b
39
41
  ```
40
42
 
41
43
  3. **Optional: Embedding model** (for embedding tests)
44
+
42
45
  ```bash
43
46
  ollama pull nomic-embed-text:latest
44
47
  ```
@@ -48,6 +51,7 @@ Environment variables are documented in the header of `script/live_branch_smoke.
48
51
  Integration tests verify:
49
52
 
50
53
  ### ✅ Core Client Methods
54
+
51
55
  - `#list_models` - Lists available models
52
56
  - `#generate` - Structured JSON output with schema
53
57
  - `#generate` - Plain text output without schema
@@ -57,54 +61,62 @@ Integration tests verify:
57
61
  - `#chat_raw` - Tool calling (if model supports it)
58
62
 
59
63
  ### ✅ Embeddings
64
+
60
65
  - Single text embeddings
61
66
  - Multiple text embeddings
62
67
  - Error handling for missing/unsupported models
63
68
 
64
69
  ### ✅ Agent Components
70
+
65
71
  - `Ollama::Agent::Planner` - Planning decisions
66
72
  - `Ollama::Agent::Executor` - Tool execution loops (if model supports tools)
67
73
 
68
74
  ### ✅ Chat Session
75
+
69
76
  - Session management
70
77
  - Conversation state
71
78
  - Message history
72
79
 
73
80
  ### ✅ Error Handling
81
+
74
82
  - `NotFoundError` for non-existent models
75
83
  - Proper error propagation
76
84
 
77
85
  ## Running Tests
78
86
 
79
87
  ### Run All Integration Tests
88
+
80
89
  ```bash
81
90
  bundle exec rspec --tag integration
82
91
  ```
83
92
 
84
93
  ### Run Specific Test File
94
+
85
95
  ```bash
86
96
  bundle exec rspec spec/integration/ollama_client_integration_spec.rb --tag integration
87
97
  ```
88
98
 
89
99
  ### Run Specific Test
100
+
90
101
  ```bash
91
102
  bundle exec rspec spec/integration/ollama_client_integration_spec.rb:32 --tag integration
92
103
  ```
93
104
 
94
105
  ### With Environment Variables
106
+
95
107
  ```bash
96
108
  # Custom Ollama URL
97
109
  OLLAMA_URL=http://remote-server:11434 bundle exec rspec --tag integration
98
110
 
99
111
  # Custom model
100
- OLLAMA_MODEL=llama3.2:3b bundle exec rspec --tag integration
112
+ OLLAMA_MODEL=qwen3.5:4b bundle exec rspec --tag integration
101
113
 
102
114
  # Custom embedding model
103
115
  OLLAMA_EMBEDDING_MODEL=nomic-embed-text:latest bundle exec rspec --tag integration
104
116
 
105
117
  # All together
106
118
  OLLAMA_URL=http://localhost:11434 \
107
- OLLAMA_MODEL=llama3.2:3b \
119
+ OLLAMA_MODEL=qwen3.5:4b \
108
120
  OLLAMA_EMBEDDING_MODEL=nomic-embed-text:latest \
109
121
  bundle exec rspec --tag integration
110
122
  ```
@@ -112,11 +124,13 @@ bundle exec rspec --tag integration
112
124
  ## Test Behavior
113
125
 
114
126
  ### Automatic Skipping
127
+
115
128
  - Tests automatically skip if Ollama server is not available
116
129
  - Tests skip if required models are not installed
117
130
  - Tests skip if models don't support certain features (e.g., tool calling)
118
131
 
119
132
  ### Expected Results
133
+
120
134
  - **Passing**: Client correctly communicates with Ollama
121
135
  - **Pending/Skipped**: Expected when models/features unavailable
122
136
  - **Failing**: Indicates actual client issues (rare)
@@ -143,20 +157,24 @@ bundle exec rspec --tag integration
143
157
  ## Troubleshooting
144
158
 
145
159
  ### "Ollama server not available"
160
+
146
161
  - Start Ollama: `ollama serve`
147
162
  - Check URL: `OLLAMA_URL=http://localhost:11434`
148
163
  - Verify connection: `curl http://localhost:11434/api/tags`
149
164
 
150
165
  ### "Model not found"
151
- - Install model: `ollama pull llama3.2:3b`
166
+
167
+ - Install model: `ollama pull qwen3.5:4b`
152
168
  - Set model: `OLLAMA_MODEL=your-model`
153
169
 
154
170
  ### "Empty embedding returned"
171
+
155
172
  - Install embedding model: `ollama pull nomic-embed-text:latest`
156
173
  - Verify model supports embeddings
157
174
  - Set model: `OLLAMA_EMBEDDING_MODEL=nomic-embed-text:latest`
158
175
 
159
176
  ### "HTTP 400: Bad Request" (tool calling)
177
+
160
178
  - Some models don't support tool calling
161
179
  - Test will skip automatically
162
180
  - Try a different model that supports tools
@@ -9,8 +9,9 @@ This document summarizes all critical fixes applied to make `ollama-client` prod
9
9
  **Issue:** LLMs can return JSON wrapped in markdown, prefixed with text, or with unicode garbage.
10
10
 
11
11
  **Fix:** Enhanced `parse_json_response()` to:
12
+
12
13
  - Handle JSON arrays (not just objects)
13
- - Strip markdown code fences (```json ... ```)
14
+ - Strip markdown code fences (```json ...```)
14
15
  - Normalize unicode and whitespace
15
16
  - Extract nested JSON if first attempt fails
16
17
  - Better error messages with context
@@ -22,6 +23,7 @@ This document summarizes all critical fixes applied to make `ollama-client` prod
22
23
  **Issue:** Need explicit retry rules for different HTTP status codes.
23
24
 
24
25
  **Fix:** Made retry policy explicit and documented:
26
+
25
27
  - **Retry:** 408 (Request Timeout), 429 (Too Many Requests), 500 (Internal Server Error), 503 (Service Unavailable)
26
28
  - **Never retry:** All other 4xx and 5xx errors
27
29
 
@@ -32,11 +34,13 @@ This document summarizes all critical fixes applied to make `ollama-client` prod
32
34
  **Issue:** Global configuration is not thread-safe, but this wasn't clearly communicated.
33
35
 
34
36
  **Fix:**
37
+
35
38
  - Added warnings in `OllamaClient.configure` when used from multiple threads
36
39
  - Added documentation comments in `Config` class
37
40
  - Warns users to use per-client configuration for concurrent agents
38
41
 
39
42
  **Location:**
43
+
40
44
  - `lib/ollama_client.rb:14-16`
41
45
  - `lib/ollama/config.rb:7-15`
42
46
 
@@ -45,6 +49,7 @@ This document summarizes all critical fixes applied to make `ollama-client` prod
45
49
  **Issue:** Schemas allow extra properties by default, letting LLMs add unexpected fields.
46
50
 
47
51
  **Fix:** Enforce `additionalProperties: false` by default:
52
+
48
53
  - Automatically adds `additionalProperties: false` to object schemas
49
54
  - Recursively applies to nested objects and array items
50
55
  - Only if not explicitly set (allows opt-out if needed)
@@ -56,6 +61,7 @@ This document summarizes all critical fixes applied to make `ollama-client` prod
56
61
  **Issue:** `chat()` should require explicit opt-in for agent usage.
57
62
 
58
63
  **Fix:**
64
+
59
65
  - Added `strict:` parameter to `chat()`
60
66
  - Warns when `strict: false` (default)
61
67
  - In strict mode, doesn't retry on schema violations
@@ -68,6 +74,7 @@ This document summarizes all critical fixes applied to make `ollama-client` prod
68
74
  **Issue:** Need a variant that fails fast on schema violations without retries.
69
75
 
70
76
  **Fix:** Added `generate_strict!()` method:
77
+
71
78
  - No retries on schema violations
72
79
  - Immediate failure for guaranteed contract enforcement
73
80
  - Useful for strict agent contracts
@@ -79,10 +86,12 @@ This document summarizes all critical fixes applied to make `ollama-client` prod
79
86
  **Issue:** No way to track latency, attempts, or model used per request.
80
87
 
81
88
  **Fix:** Added `include_meta:` parameter to `generate()` and `chat()`:
89
+
82
90
  - Returns `{ "data": ..., "meta": { "latency_ms": ..., "model": ..., "attempts": ... } }`
83
91
  - Enables logging, metrics, and debugging
84
92
 
85
93
  **Location:**
94
+
86
95
  - `lib/ollama/client.rb:58-86` (generate)
87
96
  - `lib/ollama/client.rb:27-70` (chat)
88
97
 
@@ -91,6 +100,7 @@ This document summarizes all critical fixes applied to make `ollama-client` prod
91
100
  **Issue:** No way to verify Ollama server is reachable before making requests.
92
101
 
93
102
  **Fix:** Added `health()` method:
103
+
94
104
  - Returns `{ status: "healthy|unhealthy", latency_ms: ..., error: ... }`
95
105
  - Short timeout (5s) for quick health checks
96
106
  - Useful for auto-restart systems
@@ -100,6 +110,7 @@ This document summarizes all critical fixes applied to make `ollama-client` prod
100
110
  ## 📊 Impact
101
111
 
102
112
  These fixes address **~70% of real-world Ollama failures** by:
113
+
103
114
  - Robust JSON extraction (handles model quirks)
104
115
  - Explicit retry policies (prevents wasted retries)
105
116
  - Strict schemas (catches unexpected fields)
@@ -133,7 +144,7 @@ result = client.generate(
133
144
 
134
145
  puts result["meta"]["latency_ms"] # 245.32
135
146
  puts result["meta"]["attempts"] # 1
136
- puts result["meta"]["model"] # "llama3.2:3b"
147
+ puts result["meta"]["model"] # "qwen3.5:4b"
137
148
  ```
138
149
 
139
150
  ### Health Check
@@ -162,11 +173,12 @@ result = client.generate_strict!(
162
173
  **Breaking Changes:** None - all changes are backward compatible.
163
174
 
164
175
  **New Defaults:**
176
+
165
177
  - Schemas now reject extra properties by default (can opt-out by setting `additionalProperties: true`)
166
178
  - `chat()` now warns unless `strict: true` is passed
167
179
 
168
180
  **Recommended Updates:**
181
+
169
182
  - Use `include_meta: true` for production logging
170
183
  - Use per-client config for concurrent agents
171
184
  - Use `generate_strict!` when you need guaranteed contracts
172
-
data/docs/QUICK_START.md CHANGED
@@ -12,7 +12,7 @@ client = Ollama::Client.new
12
12
 
13
13
  # Or with custom config
14
14
  config = Ollama::Config.new
15
- config.model = ENV["OLLAMA_MODEL"] || "llama3.2:3b"
15
+ config.model = ENV["OLLAMA_MODEL"] || "qwen3.5:4b"
16
16
  config.base_url = ENV["OLLAMA_BASE_URL"] || "http://localhost:11434"
17
17
  client = Ollama::Client.new(config: config)
18
18
  ```
data/docs/README.md CHANGED
@@ -9,6 +9,7 @@ This directory contains internal development documentation for the ollama-client
9
9
  ## Contents
10
10
 
11
11
  ### Design Documentation
12
+ - **[RUBYLLM_ADOPTION_MATRIX.md](RUBYLLM_ADOPTION_MATRIX.md)** - Feature-by-feature comparison against RubyLLM's public API, with copy/improve/reject/ecosystem verdicts and a prioritized build order
12
13
  - **[HANDLERS_ANALYSIS.md](HANDLERS_ANALYSIS.md)** - Analysis of handler architecture decisions (why we didn't adopt ollama-ruby's handler pattern)
13
14
  - **[FEATURES_ADDED.md](FEATURES_ADDED.md)** - Features integrated from ollama-ruby that align with our agent-first philosophy
14
15
  - **[PRODUCTION_FIXES.md](PRODUCTION_FIXES.md)** - Production-ready fixes for hybrid agents (JSON parsing, retry policy, etc.)