llm_classifier 0.2.0 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 2740f212b3f80944530c9b0ca84d18828499d8cc6d66de231bac734d2f83fc43
4
- data.tar.gz: 1a9c8211890f2a74c16d6883a28a58c8a6be4d7340e62d8c1b8aafc89746fe7e
3
+ metadata.gz: ae693cd210a9623a0a4eb201458683ab6d1a407d80a5e8b8f75ad7c08ad21f0c
4
+ data.tar.gz: ef0f86cfce90ba06a6acd264ef9219c0ed5fa61ae739fd4f992d79365b3ee03a
5
5
  SHA512:
6
- metadata.gz: 8332595d0ecb1390cda51139c745be5bc2f3c407e545594f2b9c57e22cd52f7ec0c42d2d08cf252c4e2d08a9def41b34ae31aed8062e45f4eee484776be8b4f2
7
- data.tar.gz: 3bd39aaf2842079e629046a8bf3afda9ec01d41e850eaf9be39a5d958513f63c0e51603dada50aeac1f560f3d469740d2fddeea59c9e1acc559c1dc34e4cc5e4
6
+ metadata.gz: db55a9e72578fdb7ff9f1ccf0321ddfb08f63cf60c7774e0f53968edf4e9c75f625ecab82a2a269f771e00744f9545fca1c28d6405ec564ecbfd883dc2c84eef
7
+ data.tar.gz: e69e4f57501a6d8371af21ec29d2edf76147da230fab8066121fb5471bb7217473e1c26193b0f5daadbafe0d4473d3f1c091bd3fe5aecec49b7bb2fa6bd5299a
data/CHANGELOG.md CHANGED
@@ -7,6 +7,46 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
7
7
 
8
8
  ## [Unreleased]
9
9
 
10
+ ## [0.3.0] - 2026-09-29
11
+
12
+ ### Added
13
+ - Structured output: every request sends a JSON Schema generated from the classifier's
14
+ categories (`Classifier.output_schema`), enforced by the provider via ruby_llm's `with_schema`.
15
+ Single-label classifiers get a `category` enum; multi-label get a `categories` enum array.
16
+ - `output_field` DSL to declare extra response fields (e.g. `output_field :evidence, description: "..."`).
17
+ Their values are returned in `Result#metadata`.
18
+ - `Result#model` reports the model ruby_llm actually used when the classifier doesn't set one.
19
+ - Support for ruby_llm 2.x (1.14+ remains supported).
20
+
21
+ ### Changed
22
+ - **Breaking:** `ruby_llm` (>= 1.14, < 3) is now a runtime dependency and the only built-in adapter.
23
+ Configure provider API keys with `RubyLLM.configure`.
24
+ - **Breaking:** the model must support structured outputs (for Anthropic: Claude Haiku 4.5, Sonnet 4.5,
25
+ Opus 4.5 and newer). Older models such as `claude-sonnet-4-20250514` reject the request, and the
26
+ classifier returns a failed `Result`.
27
+ - **Breaking:** `Adapters::Base#chat` now takes a `schema:` keyword. Custom adapters must accept it.
28
+ - **Breaking:** an adapter's Hash return value is treated as the `{ content:, input_tokens:, output_tokens: }`
29
+ wrapper only when it has a `:content` key; any other Hash is read as the parsed response itself.
30
+ (Previously every Hash was treated as the wrapper.)
31
+ - **Breaking:** single-label classifiers can no longer abstain. `category` is a required enum, so the
32
+ model always picks one; 0.2.0 returned a failed `Result` when the model returned no category. Use
33
+ `multi_label true` with `require_categories true` if "none of these" must be possible.
34
+ - `config.default_model` defaults to `nil`, deferring to `RubyLLM.config.default_model`
35
+ (was `"gpt-4o-mini"`).
36
+ - **Breaking:** the schema forbids undeclared fields, so extra fields a prompt asks for (which used
37
+ to land in `Result#metadata`) must now be declared with `output_field`.
38
+ - **Breaking:** single-label responses use a `category` key rather than a one-element `categories`
39
+ array. Prompts that describe the old JSON format can drop it; the schema takes precedence.
40
+ - Categories are matched case-insensitively and returned as defined.
41
+ - The default system prompt no longer includes JSON format instructions, and tells multi-label
42
+ classifiers that no category is a valid answer. The schema's `categories` description says the same,
43
+ so custom prompts get the hint too.
44
+
45
+ ### Removed
46
+ - **Breaking:** the direct `:openai` and `:anthropic` adapters, `config.openai_api_key` /
47
+ `config.anthropic_api_key`, and `Configuration#adapter_class`. Selecting a removed adapter returns a failed `Result` explaining the change.
48
+ - Markdown code-fence stripping of responses (unnecessary with schema-constrained output).
49
+
10
50
  ## [0.1.0] - 2024-12-02
11
51
 
12
52
  ### Added
data/CLAUDE.md ADDED
@@ -0,0 +1,81 @@
1
+ # CLAUDE.md
2
+
3
+ LlmClassifier - Ruby gem for building LLM-powered classifiers with a clean DSL. Talks to LLMs through ruby_llm (>= 1.14, < 3) with schema-constrained structured output, plus optional Rails integration.
4
+
5
+ - Ruby >= 3.2, RSpec, RuboCop, Zeitwerk autoloading
6
+ - No Rails dependency in core; Rails integration is opt-in via `lib/llm_classifier/rails/`
7
+ - CI tests against Ruby 3.4 and 4.0, each with ruby_llm 1.14.0 (the floor), `~> 1.16`, and `~> 2.0` (`RUBY_LLM_VERSION` env var pins the version in the Gemfile)
8
+
9
+ ## Development with Docker
10
+
11
+ Ruby is not installed on the host. Use Docker to run tests and linting:
12
+
13
+ ```bash
14
+ # Run tests and rubocop (Ruby 3.4)
15
+ docker.exe run --rm -v "$(wslpath -w "$(pwd)"):/app" -w /app ruby:3.4-slim \
16
+ bash -c "apt-get update -qq && apt-get install -y -qq build-essential git 2>/dev/null && \
17
+ gem install bundler --no-document && bundle install --quiet && \
18
+ bundle exec rspec && bundle exec rubocop"
19
+
20
+ # Rubocop only
21
+ docker.exe run --rm -v "$(wslpath -w "$(pwd)"):/app" -w /app ruby:3.4-slim \
22
+ bash -c "apt-get update -qq && apt-get install -y -qq build-essential git 2>/dev/null && \
23
+ gem install bundler --no-document && bundle install --quiet && \
24
+ bundle exec rubocop"
25
+
26
+ # Single spec file
27
+ docker.exe run --rm -v "$(wslpath -w "$(pwd)"):/app" -w /app ruby:3.4-slim \
28
+ bash -c "apt-get update -qq && apt-get install -y -qq build-essential git 2>/dev/null && \
29
+ gem install bundler --no-document && bundle install --quiet && \
30
+ bundle exec rspec spec/llm_classifier/classifier_spec.rb"
31
+ ```
32
+
33
+ Docker Desktop must be running on Windows. The `docker.exe` command is used because Docker runs via WSL2 integration. The `-v` flag bind-mounts the project so edits on host are immediately visible.
34
+
35
+ A `.devcontainer/` setup also exists for VS Code Dev Containers.
36
+
37
+ ## Quick Commands (inside Docker)
38
+
39
+ ```bash
40
+ bundle exec rspec # all tests
41
+ bundle exec rspec spec/llm_classifier/classifier_spec.rb # single file
42
+ bundle exec rubocop # all files
43
+ bundle exec rubocop -a # auto-correct
44
+ gem build llm_classifier.gemspec # build gem
45
+ ```
46
+
47
+ ## Project Structure
48
+
49
+ All sibling projects are located in `/home/axium/projects/`. The `prospector` gem depends on `llm_classifier`.
50
+
51
+ ## Code Standards
52
+
53
+ - Double-quoted strings (enforced by RuboCop)
54
+ - Max line length: 120 characters
55
+ - Max method length: 20 lines
56
+ - RSpec example max: 15 lines, max 6 expectations per example
57
+ - `Style/HashExcept` disabled (requires ActiveSupport)
58
+ - `Metrics/ClassLength` exempted for `classifier.rb` and `content_fetchers/web.rb`
59
+
60
+ ## Git Workflow
61
+
62
+ - Never push directly to main. Always create a feature branch and PR.
63
+ - Run the full test suite and rubocop before creating a PR.
64
+ - Version bumps in `lib/llm_classifier/version.rb` go in the feature PR, not separately.
65
+
66
+ ## Key Classes
67
+
68
+ - `LlmClassifier::Classifier` - Core DSL and classification pipeline; `.output_schema` builds the JSON Schema sent with every request
69
+ - `LlmClassifier::Result` - Value object returned from every classification
70
+ - `LlmClassifier::Knowledge` - Domain knowledge DSL container (`method_missing`-based)
71
+ - `LlmClassifier::Configuration` - Global config (adapter, default model, web fetch, queue). API keys live in `RubyLLM.configure`
72
+ - `LlmClassifier::Adapters::Base` - Abstract adapter interface
73
+ - `LlmClassifier::ContentFetchers::Web` - HTTP fetcher with SSRF protection
74
+ - `LlmClassifier::Rails::Concerns::Classifiable` - ActiveRecord integration
75
+
76
+ ## Component Documentation
77
+
78
+ - [lib/llm_classifier/adapters/CLAUDE.md](lib/llm_classifier/adapters/CLAUDE.md) - LLM adapter contract and implementations
79
+ - [lib/llm_classifier/content_fetchers/CLAUDE.md](lib/llm_classifier/content_fetchers/CLAUDE.md) - Content fetchers and SSRF protection
80
+ - [lib/llm_classifier/rails/CLAUDE.md](lib/llm_classifier/rails/CLAUDE.md) - Rails integration (Zeitwerk-excluded)
81
+ - [spec/CLAUDE.md](spec/CLAUDE.md) - Testing conventions
data/README.md CHANGED
@@ -1,6 +1,6 @@
1
1
  # LlmClassifier
2
2
 
3
- A flexible Ruby gem for building LLM-powered classifiers. Define categories, system prompts, and domain knowledge using a clean DSL. Supports multiple LLM backends and integrates seamlessly with Rails.
3
+ A flexible Ruby gem for building LLM-powered classifiers. Define categories, system prompts, and domain knowledge using a clean DSL. Responses are constrained to a JSON Schema generated from your categories, so the model can only answer with a category you defined. Works with any provider [ruby_llm](https://rubyllm.com) supports and integrates with Rails.
4
4
 
5
5
  ## Installation
6
6
 
@@ -8,10 +8,16 @@ Add this line to your application's Gemfile:
8
8
 
9
9
  ```ruby
10
10
  gem 'llm_classifier'
11
+ ```
12
+
13
+ `llm_classifier` talks to LLMs through [ruby_llm](https://rubyllm.com) (1.14+ or 2.x), which is installed as a dependency. Configure your provider credentials there:
11
14
 
12
- # Add your preferred LLM adapter
13
- gem 'ruby_llm' # recommended
14
- # or use direct API adapters (no additional gem needed)
15
+ ```ruby
16
+ # config/initializers/ruby_llm.rb
17
+ RubyLLM.configure do |config|
18
+ config.anthropic_api_key = ENV["ANTHROPIC_API_KEY"]
19
+ config.default_model = "claude-opus-5-5"
20
+ end
15
21
  ```
16
22
 
17
23
  And then execute:
@@ -41,17 +47,12 @@ class SentimentClassifier < LlmClassifier::Classifier
41
47
  - positive: Expresses satisfaction, happiness, or approval
42
48
  - negative: Expresses dissatisfaction, unhappiness, or criticism
43
49
  - neutral: Neither positive nor negative, factual or balanced
44
-
45
- Respond with ONLY a JSON object:
46
- {
47
- "categories": ["category"],
48
- "confidence": 0.0-1.0,
49
- "reasoning": "Brief explanation"
50
- }
51
50
  PROMPT
52
51
  end
53
52
  ```
54
53
 
54
+ You don't need to describe the response format in the prompt. The gem sends a JSON Schema with every request (see [Structured Output](#structured-output)).
55
+
55
56
  ### 2. Use It
56
57
 
57
58
  ```ruby
@@ -68,15 +69,8 @@ result.reasoning # => "Strong positive language with 'love' and 'absolutely'"
68
69
  ```ruby
69
70
  # config/initializers/llm_classifier.rb
70
71
  LlmClassifier.configure do |config|
71
- # LLM adapter: :ruby_llm (default), :openai, :anthropic
72
- config.adapter = :ruby_llm
73
-
74
- # Default model for classification
75
- config.default_model = "gpt-4o-mini"
76
-
77
- # API keys (reads from ENV by default)
78
- config.openai_api_key = ENV["OPENAI_API_KEY"]
79
- config.anthropic_api_key = ENV["ANTHROPIC_API_KEY"]
72
+ # Default model for classification. nil (the default) uses RubyLLM.config.default_model.
73
+ config.default_model = "claude-opus-5-5"
80
74
 
81
75
  # Content fetching settings
82
76
  config.web_fetch_timeout = 10
@@ -86,6 +80,44 @@ end
86
80
 
87
81
  ## Features
88
82
 
83
+ ### Structured Output
84
+
85
+ Every request carries a JSON Schema built from the classifier's categories. ruby_llm passes it to the provider's native structured-output feature (for example, Anthropic's `output_config.format`), and the provider enforces it during generation. The model can't return malformed JSON or a category you didn't define. A refused or truncated response still comes back as a failed `Result`.
86
+
87
+ ```ruby
88
+ SentimentClassifier.output_schema
89
+ # => {
90
+ # type: "object",
91
+ # properties: {
92
+ # reasoning: { type: "string", ... },
93
+ # category: { type: "string", enum: ["positive", "negative", "neutral"] },
94
+ # confidence: { type: "number", ... }
95
+ # },
96
+ # required: ["reasoning", "category", "confidence"],
97
+ # additionalProperties: false
98
+ # }
99
+ ```
100
+
101
+ ### Extra Output Fields
102
+
103
+ The schema doesn't allow fields you haven't declared. To have the model return more than the category, declare each extra field with `output_field`. Options are JSON Schema keywords, and `type` defaults to `"string"`. The values come back in `result.metadata`:
104
+
105
+ ```ruby
106
+ class BusinessClassifier < LlmClassifier::Classifier
107
+ categories :dealership, :mechanic, :parts
108
+ multi_label true
109
+ output_field :evidence, description: "Words or brands that confirmed the category"
110
+ output_field :brands, type: "array", items: { type: "string" }
111
+ end
112
+
113
+ result = BusinessClassifier.classify("Joe's Harley-Davidson Service")
114
+ result.metadata # => { "evidence" => "Harley-Davidson, Service", "brands" => ["Harley-Davidson"] }
115
+ ```
116
+
117
+ Single-label classifiers get a `category` string, so the model must pick exactly one. Multi-label classifiers get a `categories` array, which may be empty. Categories are matched case-insensitively, because providers guarantee enum membership but not capitalization.
118
+
119
+ The model needs to support structured outputs. Current Claude and OpenAI models do.
120
+
89
121
  ### Multi-label Classification
90
122
 
91
123
  ```ruby
@@ -159,16 +191,20 @@ class AuditedClassifier < LlmClassifier::Classifier
159
191
  end
160
192
  ```
161
193
 
162
- ### Override Adapter Per-Classifier
194
+ ### Override Model Per-Classifier
163
195
 
164
196
  ```ruby
165
197
  class CriticalClassifier < LlmClassifier::Classifier
166
198
  categories :high, :medium, :low
167
- adapter :anthropic # Use Anthropic for this classifier
168
- model "claude-sonnet-4-20250514" # Specific model
199
+ model "claude-opus-5-5"
169
200
  end
201
+
202
+ # Or per call
203
+ CriticalClassifier.classify(text, model: "claude-haiku-4-5")
170
204
  ```
171
205
 
206
+ If ruby_llm raises `ModelNotFoundError` for a newly released model, refresh its registry with `RubyLLM.models.refresh!`.
207
+
172
208
  ## Rails Integration
173
209
 
174
210
  ### ActiveRecord Concern
@@ -247,22 +283,20 @@ Features:
247
283
 
248
284
  ## Adapters
249
285
 
250
- ### Built-in Adapters
251
-
252
- - **`:ruby_llm`** - Uses the [ruby_llm](https://github.com/crmne/ruby_llm) gem (recommended)
253
- - **`:openai`** - Direct OpenAI API integration
254
- - **`:anthropic`** - Direct Anthropic API integration
286
+ The built-in `:ruby_llm` adapter (the default) routes requests through [ruby_llm](https://rubyllm.com), so any provider it supports works.
255
287
 
256
288
  ### Custom Adapter
257
289
 
258
290
  ```ruby
259
291
  class MyCustomAdapter < LlmClassifier::Adapters::Base
260
- def chat(model:, system_prompt:, user_prompt:)
261
- # Make API call and return response text
292
+ def chat(model:, system_prompt:, user_prompt:, schema:)
293
+ # `schema` is the classifier's JSON Schema. Constrain the response to it and
294
+ # return the parsed Hash or the JSON string.
262
295
  MyLlmClient.complete(
263
296
  model: model,
264
297
  system: system_prompt,
265
- prompt: user_prompt
298
+ prompt: user_prompt,
299
+ json_schema: schema
266
300
  )
267
301
  end
268
302
  end
@@ -285,8 +319,9 @@ result.category # => "primary_category" (first)
285
319
  result.categories # => ["cat1", "cat2"] (all)
286
320
  result.confidence # => 0.95
287
321
  result.reasoning # => "Explanation from LLM"
288
- result.raw_response # => Original JSON string
289
- result.metadata # => Additional data from response
322
+ result.raw_response # => Response JSON string
323
+ result.model # => Model used (the classifier's, or ruby_llm's default)
324
+ result.metadata # => Values of declared output_field entries
290
325
  result.error # => Error message if failed
291
326
  result.to_h # => Hash representation
292
327
  ```
@@ -0,0 +1,24 @@
1
+ # Adapters
2
+
3
+ LLM provider adapters. All inherit from `Adapters::Base` and implement `#chat(model:, system_prompt:, user_prompt:, schema:)`.
4
+
5
+ ## Inventory
6
+
7
+ - `Base` - Abstract interface. Provides `#config` helper for accessing `LlmClassifier.configuration`
8
+ - `RubyLlm` - The only built-in adapter. Calls `RubyLLM.chat(...).with_instructions(...).with_schema(...).ask(...)` and returns a Hash with `:content`, `:input_tokens`, `:output_tokens`, `:model` (`chat.model.id`, the model actually used)
9
+
10
+ ## Conventions
11
+
12
+ - `schema:` is the classifier's `output_schema` (JSON Schema, symbol keys), including any `output_field` declarations. Adapters must constrain the response to it
13
+ - `#chat` returns the content (a parsed Hash or a JSON String) or a wrapper Hash `{ content:, input_tokens:, output_tokens:, model: }`. `Classifier#extract_response_data` treats a Hash as the wrapper only when it has a `:content` key
14
+ - ruby_llm 1.x returns structured content as a Hash; 2.x returns a JSON String (`response.parsed` holds the Hash). `Classifier#parse_response` accepts both
15
+ - Token counts come from `response.tokens.input` / `.output`, which exist in ruby_llm 1.13+ and 2.x (`input_tokens` readers were removed in 2.0)
16
+ - `model: nil` lets ruby_llm fall back to `RubyLLM.config.default_model`
17
+ - Provider credentials are configured in `RubyLLM.configure`, not in this gem
18
+ - Custom adapters are passed as a Class to `config.adapter` or the `adapter` DSL
19
+ - Anthropic schema limits: no `minimum`/`maximum`, no `maxItems`, `minItems` only 0 or 1, `additionalProperties: false` required. Don't add `minItems: 1` for `require_categories`: an empty array is how the model signals "no match"
20
+
21
+ ## Related
22
+
23
+ - [../content_fetchers/CLAUDE.md](../content_fetchers/CLAUDE.md) - Content fetchers
24
+ - [../../spec/CLAUDE.md](../../spec/CLAUDE.md) - Testing conventions
@@ -4,7 +4,10 @@ module LlmClassifier
4
4
  module Adapters
5
5
  # Base adapter class for LLM providers
6
6
  class Base
7
- def chat(model:, system_prompt:, user_prompt:)
7
+ # schema is a JSON Schema Hash the response must conform to. Return the response
8
+ # content (a parsed Hash or a JSON String), or a wrapper Hash with a :content key:
9
+ # { content:, input_tokens:, output_tokens:, model: }.
10
+ def chat(model:, system_prompt:, user_prompt:, schema:)
8
11
  raise NotImplementedError, "Subclasses must implement #chat"
9
12
  end
10
13
 
@@ -2,33 +2,25 @@
2
2
 
3
3
  module LlmClassifier
4
4
  module Adapters
5
- # Adapter for the ruby_llm gem
5
+ # Adapter for the ruby_llm gem. Provider credentials and the fallback model come from
6
+ # RubyLLM's own configuration.
6
7
  class RubyLlm < Base
7
- def chat(model:, system_prompt:, user_prompt:)
8
- ensure_ruby_llm_loaded!
8
+ def chat(model:, system_prompt:, user_prompt:, schema:)
9
+ require "ruby_llm" unless defined?(::RubyLLM)
9
10
 
10
- chat_instance = ::RubyLLM.chat(model: model)
11
- chat_instance.with_instructions(system_prompt)
12
- response = chat_instance.ask(user_prompt)
11
+ chat = ::RubyLLM.chat(model: model)
12
+ response = chat.with_instructions(system_prompt)
13
+ .with_schema(name: "classification", schema: schema, strict: true)
14
+ .ask(user_prompt)
13
15
 
16
+ # ruby_llm 1.x returns structured content as a Hash, 2.x as a JSON String.
14
17
  {
15
18
  content: response.content,
16
- input_tokens: response.input_tokens,
17
- output_tokens: response.output_tokens
19
+ input_tokens: response.tokens&.input,
20
+ output_tokens: response.tokens&.output,
21
+ model: chat.model&.id
18
22
  }
19
23
  end
20
-
21
- private
22
-
23
- def ensure_ruby_llm_loaded!
24
- return if defined?(::RubyLLM)
25
-
26
- begin
27
- require "ruby_llm"
28
- rescue LoadError
29
- raise AdapterError, "ruby_llm gem is not installed. Add it to your Gemfile: gem 'ruby_llm'"
30
- end
31
- end
32
24
  end
33
25
  end
34
26
  end
@@ -5,6 +5,12 @@ require "json"
5
5
  module LlmClassifier
6
6
  # Base classifier class that provides a DSL for defining LLM-powered classifiers
7
7
  class Classifier
8
+ # Response fields the classifier itself defines.
9
+ BUILT_IN_FIELDS = %w[reasoning category categories confidence].freeze
10
+ # Names output_field can't take. "content" would make an adapter's bare parsed Hash
11
+ # look like the { content: } response wrapper.
12
+ RESERVED_FIELDS = (BUILT_IN_FIELDS + %w[content]).freeze
13
+
8
14
  class << self
9
15
  attr_reader :defined_categories, :defined_system_prompt, :defined_model,
10
16
  :defined_adapter, :defined_multi_label, :defined_require_categories,
@@ -67,6 +73,19 @@ module LlmClassifier
67
73
  @defined_knowledge
68
74
  end
69
75
 
76
+ # Declares an extra response field, returned in Result#metadata. Options are JSON Schema
77
+ # keywords for the field (type defaults to "string").
78
+ def output_field(name, **schema)
79
+ name = name.to_s
80
+ raise ArgumentError, "#{name} is a reserved output field name" if RESERVED_FIELDS.include?(name)
81
+
82
+ output_fields[name] = { type: "string" }.merge(schema)
83
+ end
84
+
85
+ def output_fields
86
+ @output_fields ||= {}
87
+ end
88
+
70
89
  def before_classify(&block)
71
90
  @before_classify_callbacks ||= []
72
91
  @before_classify_callbacks << block
@@ -80,6 +99,29 @@ module LlmClassifier
80
99
  def classify(input, **)
81
100
  new(input, **).classify
82
101
  end
102
+
103
+ # JSON Schema the LLM response is constrained to. Reasoning comes first so the
104
+ # model explains itself before committing to a label.
105
+ def output_schema
106
+ label_key = multi_label ? :categories : :category
107
+ properties = {
108
+ reasoning: { type: "string", description: "Brief explanation of the classification" },
109
+ label_key => label_schema,
110
+ confidence: { type: "number", description: "Confidence from 0.0 to 1.0" }
111
+ }.merge(output_fields.transform_keys(&:to_sym))
112
+
113
+ { type: "object", properties: properties, required: properties.keys.map(&:to_s), additionalProperties: false }
114
+ end
115
+
116
+ private
117
+
118
+ def label_schema
119
+ category = { type: "string" }
120
+ category[:enum] = categories if categories.any?
121
+ return category unless multi_label
122
+
123
+ { type: "array", items: category, description: "Every category that applies. Empty if none apply." }
124
+ end
83
125
  end
84
126
 
85
127
  attr_reader :input, :options
@@ -116,32 +158,32 @@ module LlmClassifier
116
158
  response = adapter_instance.chat(
117
159
  model: resolved_model,
118
160
  system_prompt: build_system_prompt,
119
- user_prompt: build_user_prompt(processed_input)
161
+ user_prompt: build_user_prompt(processed_input),
162
+ schema: self.class.output_schema
120
163
  )
121
164
 
122
- content, token_data = extract_response_data(response)
123
- parse_response(content, resolved_model, token_data)
165
+ content, response_meta = extract_response_data(response)
166
+ parse_response(content, resolved_model || response_meta[:model], response_meta)
124
167
  end
125
168
 
169
+ # Adapters return the content itself, or wrap it as { content:, input_tokens:, output_tokens:, model: }.
126
170
  def extract_response_data(response)
127
- if response.is_a?(Hash)
128
- [response[:content], { input_tokens: response[:input_tokens], output_tokens: response[:output_tokens] }]
129
- else
130
- [response, {}]
131
- end
171
+ return [response, {}] unless response.is_a?(Hash) && response.key?(:content)
172
+
173
+ [response[:content], response.slice(:input_tokens, :output_tokens, :model)]
132
174
  end
133
175
 
134
176
  def build_adapter
135
177
  adapter_name = self.class.adapter
136
- adapter_class = case adapter_name
137
- when :ruby_llm then Adapters::RubyLlm
138
- when :openai then Adapters::OpenAI
139
- when :anthropic then Adapters::Anthropic
140
- when Class then adapter_name
141
- else
142
- raise AdapterError, "Unknown adapter: #{adapter_name}"
143
- end
144
- adapter_class.new
178
+ case adapter_name
179
+ when :ruby_llm then Adapters::RubyLlm.new
180
+ when Class then adapter_name.new
181
+ when :openai, :anthropic
182
+ raise AdapterError, "The :#{adapter_name} adapter was removed in 0.3.0. " \
183
+ "Configure the provider in RubyLLM and use the :ruby_llm adapter."
184
+ else
185
+ raise AdapterError, "Unknown adapter: #{adapter_name}"
186
+ end
145
187
  end
146
188
 
147
189
  def build_system_prompt
@@ -157,16 +199,8 @@ module LlmClassifier
157
199
  categories = self.class.categories.join(", ")
158
200
  multi = self.class.multi_label
159
201
 
160
- <<~PROMPT
161
- You are a classifier. Classify the given input into #{multi ? "one or more of" : "exactly one of"} these categories: #{categories}.
162
-
163
- Respond with ONLY a JSON object in this format:
164
- {
165
- "categories": [#{multi ? '"category1", "category2"' : '"category"'}],
166
- "confidence": 0.0-1.0,
167
- "reasoning": "Brief explanation"
168
- }
169
- PROMPT
202
+ scope = multi ? "every category that applies (none if none apply)" : "exactly one of these categories"
203
+ "You are a classifier. Classify the given input into #{scope}: #{categories}."
170
204
  end
171
205
 
172
206
  def build_user_prompt(processed_input)
@@ -180,24 +214,24 @@ module LlmClassifier
180
214
  end
181
215
  end
182
216
 
183
- def parse_response(response, resolved_model = nil, token_data = {})
184
- json = JSON.parse(strip_code_fences(response))
217
+ # Adapters return structured output either already parsed (a Hash) or as JSON text.
218
+ def parse_response(content, resolved_model = nil, token_data = {})
219
+ json = content.is_a?(Hash) ? content.transform_keys(&:to_s) : JSON.parse(content.to_s)
220
+ raw_response = content.is_a?(String) ? content : JSON.generate(json)
185
221
  valid_categories = extract_valid_categories(json)
186
222
 
187
- return build_failure_result(response, json) if should_fail?(valid_categories)
223
+ return build_failure_result(raw_response, json) if should_fail?(valid_categories)
188
224
 
189
- build_success_result(json, valid_categories, response, resolved_model, token_data)
225
+ build_success_result(json, valid_categories, raw_response, resolved_model, token_data)
190
226
  rescue JSON::ParserError => e
191
- Result.failure(error: "Failed to parse response: #{e.message}", raw_response: response)
192
- end
193
-
194
- def strip_code_fences(text)
195
- text.sub(/\A\s*```\w*\R?/, "").sub(/\R?```\s*\z/, "")
227
+ Result.failure(error: "Failed to parse response: #{e.message}", raw_response: content)
196
228
  end
197
229
 
230
+ # Structured outputs guarantee enum membership but not capitalization, so match
231
+ # case-insensitively and return the category as defined.
198
232
  def extract_valid_categories(json)
199
- raw_categories = Array(json["categories"] || json["category"])
200
- raw_categories.select { |c| self.class.categories.include?(c.to_s) }
233
+ defined = self.class.categories.to_h { |c| [c.downcase, c] }
234
+ Array(json["categories"] || json["category"]).filter_map { |c| defined[c.to_s.downcase] }.uniq
201
235
  end
202
236
 
203
237
  def should_fail?(valid_categories)
@@ -217,8 +251,7 @@ module LlmClassifier
217
251
 
218
252
  def build_success_result(json, valid_categories, response, resolved_model = nil, token_data = {})
219
253
  categories = self.class.multi_label ? valid_categories : [valid_categories.first].compact
220
- excluded_keys = %w[categories category confidence reasoning]
221
- metadata = json.reject { |k, _| excluded_keys.include?(k) }
254
+ metadata = json.reject { |k, _| BUILT_IN_FIELDS.include?(k) }
222
255
 
223
256
  Result.success(
224
257
  categories: categories,
@@ -5,34 +5,16 @@ require "logger"
5
5
  module LlmClassifier
6
6
  # Configuration object for LlmClassifier settings
7
7
  class Configuration
8
- attr_accessor :adapter, :default_model, :openai_api_key, :anthropic_api_key,
9
- :web_fetch_timeout, :web_fetch_user_agent, :default_queue,
10
- :logger
8
+ attr_accessor :adapter, :default_model, :web_fetch_timeout, :web_fetch_user_agent,
9
+ :default_queue, :logger
11
10
 
12
11
  def initialize
13
12
  @adapter = :ruby_llm
14
- @default_model = "gpt-4o-mini"
15
- @openai_api_key = ENV.fetch("OPENAI_API_KEY", nil)
16
- @anthropic_api_key = ENV.fetch("ANTHROPIC_API_KEY", nil)
13
+ @default_model = nil # nil defers to RubyLLM.config.default_model
17
14
  @web_fetch_timeout = 10
18
15
  @web_fetch_user_agent = "LlmClassifier/#{VERSION}"
19
16
  @default_queue = :classification
20
17
  @logger = defined?(::Rails) ? ::Rails.logger : Logger.new($stdout)
21
18
  end
22
-
23
- def adapter_class
24
- case adapter
25
- when :ruby_llm
26
- Adapters::RubyLlm
27
- when :openai
28
- Adapters::OpenAI
29
- when :anthropic
30
- Adapters::Anthropic
31
- when Class
32
- adapter
33
- else
34
- raise ConfigurationError, "Unknown adapter: #{adapter}"
35
- end
36
- end
37
19
  end
38
20
  end
@@ -0,0 +1,27 @@
1
+ # Content Fetchers
2
+
3
+ Utilities for fetching external content to use as classification input. Not wired into `Classifier` automatically -- callers fetch content and pass it in.
4
+
5
+ ## Inventory
6
+
7
+ - `Base` - Abstract interface. Subclasses implement `#fetch(source)`
8
+ - `Web` - HTTP fetcher with SSRF protection, redirect following, and HTML text extraction
9
+ - `Null` - No-op fetcher, always returns `nil`
10
+
11
+ ## SSRF Protection (`Web`)
12
+
13
+ - Validates resolved IPs against private/loopback CIDR ranges before connecting
14
+ - Follows up to 3 redirects, re-validating each redirect target
15
+ - `normalize_redirect_url` handles relative and absolute redirect URLs
16
+ - Uses `nil? || empty?` guards (not ActiveSupport `.blank?`) to avoid the dependency
17
+
18
+ ## HTML Processing (`Web`)
19
+
20
+ - Nokogiri is lazily loaded (`require "nokogiri"` inside the method) since it's an optional dependency
21
+ - Strips `<script>`, `<style>`, `<nav>`, `<footer>`, `<header>` elements
22
+ - Truncates extracted text to 2000 characters
23
+
24
+ ## Related
25
+
26
+ - [../adapters/CLAUDE.md](../adapters/CLAUDE.md) - LLM adapters
27
+ - [../../spec/CLAUDE.md](../../spec/CLAUDE.md) - Testing conventions
@@ -0,0 +1,25 @@
1
+ # Rails Integration
2
+
3
+ This entire subtree is excluded from Zeitwerk autoloading (`loader.ignore` in `lib/llm_classifier.rb`) and loaded manually only when `Rails::Railtie` is defined. This keeps Rails as an optional dependency.
4
+
5
+ ## Inventory
6
+
7
+ - `Railtie` - Sets `Rails.logger` as the default LlmClassifier logger
8
+ - `Concerns::Classifiable` - ActiveRecord concern adding a `classifies` macro
9
+ - `Generators::InstallGenerator` - `rails g llm_classifier:install` scaffolds an initializer
10
+ - `Generators::ClassifierGenerator` - `rails g llm_classifier:classifier Name cat1 cat2` scaffolds a classifier and spec
11
+
12
+ ## Classifiable Concern
13
+
14
+ The `classifies` macro defines three instance methods per classification:
15
+
16
+ - `classify_<attr>!` - Runs classification and stores the result
17
+ - `<attr>_category` / `<attr>_categories` - Reads stored category data
18
+ - `<attr>_classification` - Returns the full stored classification hash
19
+
20
+ Results are written into a JSONB column (via `store_in:`) or a transient instance variable if no column is specified.
21
+
22
+ ## Related
23
+
24
+ - [../adapters/CLAUDE.md](../adapters/CLAUDE.md) - LLM adapters
25
+ - [../content_fetchers/CLAUDE.md](../content_fetchers/CLAUDE.md) - Content fetchers
@@ -14,16 +14,10 @@ module LlmClassifier
14
14
  create_file "config/initializers/llm_classifier.rb", <<~RUBY
15
15
  # frozen_string_literal: true
16
16
 
17
+ # Provider API keys are configured in RubyLLM (config/initializers/ruby_llm.rb).
17
18
  LlmClassifier.configure do |config|
18
- # LLM adapter to use. Options: :ruby_llm, :openai, :anthropic
19
- config.adapter = :ruby_llm
20
-
21
- # Default model for classification
22
- config.default_model = "gpt-4o-mini"
23
-
24
- # API keys (reads from ENV by default)
25
- # config.openai_api_key = ENV["OPENAI_API_KEY"]
26
- # config.anthropic_api_key = ENV["ANTHROPIC_API_KEY"]
19
+ # Default model for classification. nil uses RubyLLM.config.default_model.
20
+ # config.default_model = "claude-opus-5-5"
27
21
 
28
22
  # Content fetching settings
29
23
  config.web_fetch_timeout = 10
@@ -45,7 +39,7 @@ module LlmClassifier
45
39
  say "LlmClassifier installed successfully!", :green
46
40
  say "\n"
47
41
  say "Next steps:"
48
- say " 1. Configure your API keys in config/initializers/llm_classifier.rb"
42
+ say " 1. Configure your provider API keys with RubyLLM.configure"
49
43
  say " 2. Generate a classifier: rails g llm_classifier:classifier SentimentClassifier"
50
44
  say "\n"
51
45
  end
@@ -7,10 +7,7 @@ class <%= class_name %> < LlmClassifier::Classifier
7
7
  # multi_label true
8
8
 
9
9
  # Uncomment to override the default model
10
- # model "gpt-4o-mini"
11
-
12
- # Uncomment to override the default adapter
13
- # adapter :openai
10
+ # model "claude-opus-5-5"
14
11
 
15
12
  system_prompt <<~PROMPT
16
13
  You are a classifier. Analyze the given input and classify it into the appropriate category.
@@ -19,13 +16,6 @@ class <%= class_name %> < LlmClassifier::Classifier
19
16
  <% categories_array.each do |cat| -%>
20
17
  - <%= cat %>: [describe what this category means]
21
18
  <% end -%>
22
-
23
- Respond with ONLY a JSON object:
24
- {
25
- "categories": ["category"],
26
- "confidence": 0.0-1.0,
27
- "reasoning": "Brief explanation"
28
- }
29
19
  PROMPT
30
20
 
31
21
  # Uncomment to add domain knowledge
@@ -1,5 +1,5 @@
1
1
  # frozen_string_literal: true
2
2
 
3
3
  module LlmClassifier
4
- VERSION = "0.2.0"
4
+ VERSION = "0.3.0"
5
5
  end
@@ -3,10 +3,7 @@
3
3
  require "zeitwerk"
4
4
 
5
5
  loader = Zeitwerk::Loader.for_gem
6
- loader.inflector.inflect(
7
- "openai" => "OpenAI",
8
- "ruby_llm" => "RubyLlm"
9
- )
6
+ loader.inflector.inflect("ruby_llm" => "RubyLlm")
10
7
  loader.ignore("#{__dir__}/llm_classifier/rails")
11
8
  loader.setup
12
9
 
metadata CHANGED
@@ -1,15 +1,34 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: llm_classifier
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.2.0
4
+ version: 0.3.0
5
5
  platform: ruby
6
6
  authors:
7
7
  - Dmitry Sychev
8
- autorequire:
9
8
  bindir: exe
10
9
  cert_chain: []
11
- date: 2026-04-05 00:00:00.000000000 Z
10
+ date: 1980-01-02 00:00:00.000000000 Z
12
11
  dependencies:
12
+ - !ruby/object:Gem::Dependency
13
+ name: ruby_llm
14
+ requirement: !ruby/object:Gem::Requirement
15
+ requirements:
16
+ - - ">="
17
+ - !ruby/object:Gem::Version
18
+ version: '1.14'
19
+ - - "<"
20
+ - !ruby/object:Gem::Version
21
+ version: '3'
22
+ type: :runtime
23
+ prerelease: false
24
+ version_requirements: !ruby/object:Gem::Requirement
25
+ requirements:
26
+ - - ">="
27
+ - !ruby/object:Gem::Version
28
+ version: '1.14'
29
+ - - "<"
30
+ - !ruby/object:Gem::Version
31
+ version: '3'
13
32
  - !ruby/object:Gem::Dependency
14
33
  name: zeitwerk
15
34
  requirement: !ruby/object:Gem::Requirement
@@ -25,8 +44,8 @@ dependencies:
25
44
  - !ruby/object:Gem::Version
26
45
  version: '2.6'
27
46
  description: A flexible Ruby gem for building LLM-based classifiers. Define categories,
28
- system prompts, and domain knowledge using a clean DSL. Supports multiple LLM backends
29
- (ruby_llm, OpenAI, Anthropic) and integrates seamlessly with Rails.
47
+ system prompts, and domain knowledge using a clean DSL. Responses are schema-constrained
48
+ via ruby_llm structured outputs, with optional Rails integration.
30
49
  email:
31
50
  - dmitry.sychev@axiumfoundry.com
32
51
  executables: []
@@ -39,20 +58,22 @@ files:
39
58
  - ".rspec"
40
59
  - ".rubocop.yml"
41
60
  - CHANGELOG.md
61
+ - CLAUDE.md
42
62
  - LICENSE.txt
43
63
  - README.md
44
64
  - Rakefile
45
65
  - lib/llm_classifier.rb
46
- - lib/llm_classifier/adapters/anthropic.rb
66
+ - lib/llm_classifier/adapters/CLAUDE.md
47
67
  - lib/llm_classifier/adapters/base.rb
48
- - lib/llm_classifier/adapters/openai.rb
49
68
  - lib/llm_classifier/adapters/ruby_llm.rb
50
69
  - lib/llm_classifier/classifier.rb
51
70
  - lib/llm_classifier/configuration.rb
71
+ - lib/llm_classifier/content_fetchers/CLAUDE.md
52
72
  - lib/llm_classifier/content_fetchers/base.rb
53
73
  - lib/llm_classifier/content_fetchers/null.rb
54
74
  - lib/llm_classifier/content_fetchers/web.rb
55
75
  - lib/llm_classifier/knowledge.rb
76
+ - lib/llm_classifier/rails/CLAUDE.md
56
77
  - lib/llm_classifier/rails/concerns/classifiable.rb
57
78
  - lib/llm_classifier/rails/generators/classifier_generator.rb
58
79
  - lib/llm_classifier/rails/generators/install_generator.rb
@@ -69,7 +90,6 @@ metadata:
69
90
  source_code_uri: https://github.com/AxiumFoundry/llm_classifier
70
91
  changelog_uri: https://github.com/AxiumFoundry/llm_classifier/blob/main/CHANGELOG.md
71
92
  rubygems_mfa_required: 'true'
72
- post_install_message:
73
93
  rdoc_options: []
74
94
  require_paths:
75
95
  - lib
@@ -84,8 +104,7 @@ required_rubygems_version: !ruby/object:Gem::Requirement
84
104
  - !ruby/object:Gem::Version
85
105
  version: '0'
86
106
  requirements: []
87
- rubygems_version: 3.4.20
88
- signing_key:
107
+ rubygems_version: 3.6.9
89
108
  specification_version: 4
90
109
  summary: LLM-powered classification for Ruby with pluggable adapters and Rails integration
91
110
  test_files: []
@@ -1,72 +0,0 @@
1
- # frozen_string_literal: true
2
-
3
- require "net/http"
4
- require "json"
5
- require "uri"
6
-
7
- module LlmClassifier
8
- module Adapters
9
- # Adapter for Anthropic API
10
- class Anthropic < Base
11
- API_URL = "https://api.anthropic.com/v1/messages"
12
- API_VERSION = "2023-06-01"
13
-
14
- def chat(model:, system_prompt:, user_prompt:)
15
- api_key = validate_api_key
16
- response = send_request(model, system_prompt, user_prompt, api_key)
17
- parse_response(response)
18
- end
19
-
20
- private
21
-
22
- def validate_api_key
23
- api_key = config.anthropic_api_key
24
- raise ConfigurationError, "Anthropic API key not configured" unless api_key
25
-
26
- api_key
27
- end
28
-
29
- def send_request(model, system_prompt, user_prompt, api_key)
30
- uri = URI(API_URL)
31
- http = build_http_client(uri)
32
- request = build_request(uri, api_key, model, system_prompt, user_prompt)
33
- http.request(request)
34
- end
35
-
36
- def build_http_client(uri)
37
- http = Net::HTTP.new(uri.host, uri.port)
38
- http.use_ssl = true
39
- http
40
- end
41
-
42
- def build_request(uri, api_key, model, system_prompt, user_prompt)
43
- request = Net::HTTP::Post.new(uri)
44
- request["Content-Type"] = "application/json"
45
- request["x-api-key"] = api_key
46
- request["anthropic-version"] = API_VERSION
47
- request.body = build_request_body(model, system_prompt, user_prompt)
48
- request
49
- end
50
-
51
- def build_request_body(model, system_prompt, user_prompt)
52
- {
53
- model: model,
54
- max_tokens: 1024,
55
- system: system_prompt,
56
- messages: [
57
- { role: "user", content: user_prompt }
58
- ]
59
- }.to_json
60
- end
61
-
62
- def parse_response(response)
63
- unless response.is_a?(Net::HTTPSuccess)
64
- raise AdapterError, "Anthropic API error: #{response.code} - #{response.body}"
65
- end
66
-
67
- parsed = JSON.parse(response.body)
68
- parsed.dig("content", 0, "text")
69
- end
70
- end
71
- end
72
- end
@@ -1,70 +0,0 @@
1
- # frozen_string_literal: true
2
-
3
- require "net/http"
4
- require "json"
5
- require "uri"
6
-
7
- module LlmClassifier
8
- module Adapters
9
- # Adapter for OpenAI API
10
- class OpenAI < Base
11
- API_URL = "https://api.openai.com/v1/chat/completions"
12
-
13
- def chat(model:, system_prompt:, user_prompt:)
14
- api_key = validate_api_key
15
- response = send_request(model, system_prompt, user_prompt, api_key)
16
- parse_response(response)
17
- end
18
-
19
- private
20
-
21
- def validate_api_key
22
- api_key = config.openai_api_key
23
- raise ConfigurationError, "OpenAI API key not configured" unless api_key
24
-
25
- api_key
26
- end
27
-
28
- def send_request(model, system_prompt, user_prompt, api_key)
29
- uri = URI(API_URL)
30
- http = build_http_client(uri)
31
- request = build_request(uri, api_key, model, system_prompt, user_prompt)
32
- http.request(request)
33
- end
34
-
35
- def build_http_client(uri)
36
- http = Net::HTTP.new(uri.host, uri.port)
37
- http.use_ssl = true
38
- http
39
- end
40
-
41
- def build_request(uri, api_key, model, system_prompt, user_prompt)
42
- request = Net::HTTP::Post.new(uri)
43
- request["Content-Type"] = "application/json"
44
- request["Authorization"] = "Bearer #{api_key}"
45
- request.body = build_request_body(model, system_prompt, user_prompt)
46
- request
47
- end
48
-
49
- def build_request_body(model, system_prompt, user_prompt)
50
- {
51
- model: model,
52
- messages: [
53
- { role: "system", content: system_prompt },
54
- { role: "user", content: user_prompt }
55
- ],
56
- temperature: 0.3
57
- }.to_json
58
- end
59
-
60
- def parse_response(response)
61
- unless response.is_a?(Net::HTTPSuccess)
62
- raise AdapterError, "OpenAI API error: #{response.code} - #{response.body}"
63
- end
64
-
65
- parsed = JSON.parse(response.body)
66
- parsed.dig("choices", 0, "message", "content")
67
- end
68
- end
69
- end
70
- end