llm_classifier 0.1.0 → 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/.rubocop.yml +3 -3
- data/CHANGELOG.md +40 -0
- data/CLAUDE.md +81 -0
- data/README.md +90 -34
- data/lib/llm_classifier/adapters/CLAUDE.md +24 -0
- data/lib/llm_classifier/adapters/base.rb +4 -1
- data/lib/llm_classifier/adapters/ruby_llm.rb +15 -19
- data/lib/llm_classifier/classifier.rb +99 -37
- data/lib/llm_classifier/configuration.rb +3 -21
- data/lib/llm_classifier/content_fetchers/CLAUDE.md +27 -0
- data/lib/llm_classifier/content_fetchers/web.rb +1 -1
- data/lib/llm_classifier/rails/CLAUDE.md +25 -0
- data/lib/llm_classifier/rails/generators/install_generator.rb +4 -10
- data/lib/llm_classifier/rails/generators/templates/classifier.rb.erb +1 -11
- data/lib/llm_classifier/result.rb +19 -5
- data/lib/llm_classifier/version.rb +1 -1
- data/lib/llm_classifier.rb +1 -4
- metadata +28 -6
- data/lib/llm_classifier/adapters/anthropic.rb +0 -72
- data/lib/llm_classifier/adapters/openai.rb +0 -70
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: ae693cd210a9623a0a4eb201458683ab6d1a407d80a5e8b8f75ad7c08ad21f0c
|
|
4
|
+
data.tar.gz: ef0f86cfce90ba06a6acd264ef9219c0ed5fa61ae739fd4f992d79365b3ee03a
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: db55a9e72578fdb7ff9f1ccf0321ddfb08f63cf60c7774e0f53968edf4e9c75f625ecab82a2a269f771e00744f9545fca1c28d6405ec564ecbfd883dc2c84eef
|
|
7
|
+
data.tar.gz: e69e4f57501a6d8371af21ec29d2edf76147da230fab8066121fb5471bb7217473e1c26193b0f5daadbafe0d4473d3f1c091bd3fe5aecec49b7bb2fa6bd5299a
|
data/.rubocop.yml
CHANGED
|
@@ -1,8 +1,8 @@
|
|
|
1
|
-
|
|
1
|
+
plugins:
|
|
2
2
|
- rubocop-rspec
|
|
3
3
|
|
|
4
4
|
AllCops:
|
|
5
|
-
TargetRubyVersion: 3.
|
|
5
|
+
TargetRubyVersion: 3.2
|
|
6
6
|
NewCops: enable
|
|
7
7
|
SuggestExtensions: false
|
|
8
8
|
Exclude:
|
|
@@ -42,4 +42,4 @@ RSpec/ExampleLength:
|
|
|
42
42
|
Max: 15
|
|
43
43
|
|
|
44
44
|
RSpec/MultipleExpectations:
|
|
45
|
-
Max:
|
|
45
|
+
Max: 6
|
data/CHANGELOG.md
CHANGED
|
@@ -7,6 +7,46 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|
|
7
7
|
|
|
8
8
|
## [Unreleased]
|
|
9
9
|
|
|
10
|
+
## [0.3.0] - 2026-09-29
|
|
11
|
+
|
|
12
|
+
### Added
|
|
13
|
+
- Structured output: every request sends a JSON Schema generated from the classifier's
|
|
14
|
+
categories (`Classifier.output_schema`), enforced by the provider via ruby_llm's `with_schema`.
|
|
15
|
+
Single-label classifiers get a `category` enum; multi-label get a `categories` enum array.
|
|
16
|
+
- `output_field` DSL to declare extra response fields (e.g. `output_field :evidence, description: "..."`).
|
|
17
|
+
Their values are returned in `Result#metadata`.
|
|
18
|
+
- `Result#model` reports the model ruby_llm actually used when the classifier doesn't set one.
|
|
19
|
+
- Support for ruby_llm 2.x (1.14+ remains supported).
|
|
20
|
+
|
|
21
|
+
### Changed
|
|
22
|
+
- **Breaking:** `ruby_llm` (>= 1.14, < 3) is now a runtime dependency and the only built-in adapter.
|
|
23
|
+
Configure provider API keys with `RubyLLM.configure`.
|
|
24
|
+
- **Breaking:** the model must support structured outputs (for Anthropic: Claude Haiku 4.5, Sonnet 4.5,
|
|
25
|
+
Opus 4.5 and newer). Older models such as `claude-sonnet-4-20250514` reject the request, and the
|
|
26
|
+
classifier returns a failed `Result`.
|
|
27
|
+
- **Breaking:** `Adapters::Base#chat` now takes a `schema:` keyword. Custom adapters must accept it.
|
|
28
|
+
- **Breaking:** an adapter's Hash return value is treated as the `{ content:, input_tokens:, output_tokens: }`
|
|
29
|
+
wrapper only when it has a `:content` key; any other Hash is read as the parsed response itself.
|
|
30
|
+
(Previously every Hash was treated as the wrapper.)
|
|
31
|
+
- **Breaking:** single-label classifiers can no longer abstain. `category` is a required enum, so the
|
|
32
|
+
model always picks one; 0.2.0 returned a failed `Result` when the model returned no category. Use
|
|
33
|
+
`multi_label true` with `require_categories true` if "none of these" must be possible.
|
|
34
|
+
- `config.default_model` defaults to `nil`, deferring to `RubyLLM.config.default_model`
|
|
35
|
+
(was `"gpt-4o-mini"`).
|
|
36
|
+
- **Breaking:** the schema forbids undeclared fields, so extra fields a prompt asks for (which used
|
|
37
|
+
to land in `Result#metadata`) must now be declared with `output_field`.
|
|
38
|
+
- **Breaking:** single-label responses use a `category` key rather than a one-element `categories`
|
|
39
|
+
array. Prompts that describe the old JSON format can drop it; the schema takes precedence.
|
|
40
|
+
- Categories are matched case-insensitively and returned as defined.
|
|
41
|
+
- The default system prompt no longer includes JSON format instructions, and tells multi-label
|
|
42
|
+
classifiers that no category is a valid answer. The schema's `categories` description says the same,
|
|
43
|
+
so custom prompts get the hint too.
|
|
44
|
+
|
|
45
|
+
### Removed
|
|
46
|
+
- **Breaking:** the direct `:openai` and `:anthropic` adapters, `config.openai_api_key` /
|
|
47
|
+
`config.anthropic_api_key`, and `Configuration#adapter_class`. Selecting a removed adapter returns a failed `Result` explaining the change.
|
|
48
|
+
- Markdown code-fence stripping of responses (unnecessary with schema-constrained output).
|
|
49
|
+
|
|
10
50
|
## [0.1.0] - 2024-12-02
|
|
11
51
|
|
|
12
52
|
### Added
|
data/CLAUDE.md
ADDED
|
@@ -0,0 +1,81 @@
|
|
|
1
|
+
# CLAUDE.md
|
|
2
|
+
|
|
3
|
+
LlmClassifier - Ruby gem for building LLM-powered classifiers with a clean DSL. Talks to LLMs through ruby_llm (>= 1.14, < 3) with schema-constrained structured output, plus optional Rails integration.
|
|
4
|
+
|
|
5
|
+
- Ruby >= 3.2, RSpec, RuboCop, Zeitwerk autoloading
|
|
6
|
+
- No Rails dependency in core; Rails integration is opt-in via `lib/llm_classifier/rails/`
|
|
7
|
+
- CI tests against Ruby 3.4 and 4.0, each with ruby_llm 1.14.0 (the floor), `~> 1.16`, and `~> 2.0` (`RUBY_LLM_VERSION` env var pins the version in the Gemfile)
|
|
8
|
+
|
|
9
|
+
## Development with Docker
|
|
10
|
+
|
|
11
|
+
Ruby is not installed on the host. Use Docker to run tests and linting:
|
|
12
|
+
|
|
13
|
+
```bash
|
|
14
|
+
# Run tests and rubocop (Ruby 3.4)
|
|
15
|
+
docker.exe run --rm -v "$(wslpath -w "$(pwd)"):/app" -w /app ruby:3.4-slim \
|
|
16
|
+
bash -c "apt-get update -qq && apt-get install -y -qq build-essential git 2>/dev/null && \
|
|
17
|
+
gem install bundler --no-document && bundle install --quiet && \
|
|
18
|
+
bundle exec rspec && bundle exec rubocop"
|
|
19
|
+
|
|
20
|
+
# Rubocop only
|
|
21
|
+
docker.exe run --rm -v "$(wslpath -w "$(pwd)"):/app" -w /app ruby:3.4-slim \
|
|
22
|
+
bash -c "apt-get update -qq && apt-get install -y -qq build-essential git 2>/dev/null && \
|
|
23
|
+
gem install bundler --no-document && bundle install --quiet && \
|
|
24
|
+
bundle exec rubocop"
|
|
25
|
+
|
|
26
|
+
# Single spec file
|
|
27
|
+
docker.exe run --rm -v "$(wslpath -w "$(pwd)"):/app" -w /app ruby:3.4-slim \
|
|
28
|
+
bash -c "apt-get update -qq && apt-get install -y -qq build-essential git 2>/dev/null && \
|
|
29
|
+
gem install bundler --no-document && bundle install --quiet && \
|
|
30
|
+
bundle exec rspec spec/llm_classifier/classifier_spec.rb"
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
Docker Desktop must be running on Windows. The `docker.exe` command is used because Docker runs via WSL2 integration. The `-v` flag bind-mounts the project so edits on host are immediately visible.
|
|
34
|
+
|
|
35
|
+
A `.devcontainer/` setup also exists for VS Code Dev Containers.
|
|
36
|
+
|
|
37
|
+
## Quick Commands (inside Docker)
|
|
38
|
+
|
|
39
|
+
```bash
|
|
40
|
+
bundle exec rspec # all tests
|
|
41
|
+
bundle exec rspec spec/llm_classifier/classifier_spec.rb # single file
|
|
42
|
+
bundle exec rubocop # all files
|
|
43
|
+
bundle exec rubocop -a # auto-correct
|
|
44
|
+
gem build llm_classifier.gemspec # build gem
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
## Project Structure
|
|
48
|
+
|
|
49
|
+
All sibling projects are located in `/home/axium/projects/`. The `prospector` gem depends on `llm_classifier`.
|
|
50
|
+
|
|
51
|
+
## Code Standards
|
|
52
|
+
|
|
53
|
+
- Double-quoted strings (enforced by RuboCop)
|
|
54
|
+
- Max line length: 120 characters
|
|
55
|
+
- Max method length: 20 lines
|
|
56
|
+
- RSpec example max: 15 lines, max 6 expectations per example
|
|
57
|
+
- `Style/HashExcept` disabled (requires ActiveSupport)
|
|
58
|
+
- `Metrics/ClassLength` exempted for `classifier.rb` and `content_fetchers/web.rb`
|
|
59
|
+
|
|
60
|
+
## Git Workflow
|
|
61
|
+
|
|
62
|
+
- Never push directly to main. Always create a feature branch and PR.
|
|
63
|
+
- Run the full test suite and rubocop before creating a PR.
|
|
64
|
+
- Version bumps in `lib/llm_classifier/version.rb` go in the feature PR, not separately.
|
|
65
|
+
|
|
66
|
+
## Key Classes
|
|
67
|
+
|
|
68
|
+
- `LlmClassifier::Classifier` - Core DSL and classification pipeline; `.output_schema` builds the JSON Schema sent with every request
|
|
69
|
+
- `LlmClassifier::Result` - Value object returned from every classification
|
|
70
|
+
- `LlmClassifier::Knowledge` - Domain knowledge DSL container (`method_missing`-based)
|
|
71
|
+
- `LlmClassifier::Configuration` - Global config (adapter, default model, web fetch, queue). API keys live in `RubyLLM.configure`
|
|
72
|
+
- `LlmClassifier::Adapters::Base` - Abstract adapter interface
|
|
73
|
+
- `LlmClassifier::ContentFetchers::Web` - HTTP fetcher with SSRF protection
|
|
74
|
+
- `LlmClassifier::Rails::Concerns::Classifiable` - ActiveRecord integration
|
|
75
|
+
|
|
76
|
+
## Component Documentation
|
|
77
|
+
|
|
78
|
+
- [lib/llm_classifier/adapters/CLAUDE.md](lib/llm_classifier/adapters/CLAUDE.md) - LLM adapter contract and implementations
|
|
79
|
+
- [lib/llm_classifier/content_fetchers/CLAUDE.md](lib/llm_classifier/content_fetchers/CLAUDE.md) - Content fetchers and SSRF protection
|
|
80
|
+
- [lib/llm_classifier/rails/CLAUDE.md](lib/llm_classifier/rails/CLAUDE.md) - Rails integration (Zeitwerk-excluded)
|
|
81
|
+
- [spec/CLAUDE.md](spec/CLAUDE.md) - Testing conventions
|
data/README.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# LlmClassifier
|
|
2
2
|
|
|
3
|
-
A flexible Ruby gem for building LLM-powered classifiers. Define categories, system prompts, and domain knowledge using a clean DSL.
|
|
3
|
+
A flexible Ruby gem for building LLM-powered classifiers. Define categories, system prompts, and domain knowledge using a clean DSL. Responses are constrained to a JSON Schema generated from your categories, so the model can only answer with a category you defined. Works with any provider [ruby_llm](https://rubyllm.com) supports and integrates with Rails.
|
|
4
4
|
|
|
5
5
|
## Installation
|
|
6
6
|
|
|
@@ -8,10 +8,16 @@ Add this line to your application's Gemfile:
|
|
|
8
8
|
|
|
9
9
|
```ruby
|
|
10
10
|
gem 'llm_classifier'
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
`llm_classifier` talks to LLMs through [ruby_llm](https://rubyllm.com) (1.14+ or 2.x), which is installed as a dependency. Configure your provider credentials there:
|
|
11
14
|
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
+
```ruby
|
|
16
|
+
# config/initializers/ruby_llm.rb
|
|
17
|
+
RubyLLM.configure do |config|
|
|
18
|
+
config.anthropic_api_key = ENV["ANTHROPIC_API_KEY"]
|
|
19
|
+
config.default_model = "claude-opus-5-5"
|
|
20
|
+
end
|
|
15
21
|
```
|
|
16
22
|
|
|
17
23
|
And then execute:
|
|
@@ -41,17 +47,12 @@ class SentimentClassifier < LlmClassifier::Classifier
|
|
|
41
47
|
- positive: Expresses satisfaction, happiness, or approval
|
|
42
48
|
- negative: Expresses dissatisfaction, unhappiness, or criticism
|
|
43
49
|
- neutral: Neither positive nor negative, factual or balanced
|
|
44
|
-
|
|
45
|
-
Respond with ONLY a JSON object:
|
|
46
|
-
{
|
|
47
|
-
"categories": ["category"],
|
|
48
|
-
"confidence": 0.0-1.0,
|
|
49
|
-
"reasoning": "Brief explanation"
|
|
50
|
-
}
|
|
51
50
|
PROMPT
|
|
52
51
|
end
|
|
53
52
|
```
|
|
54
53
|
|
|
54
|
+
You don't need to describe the response format in the prompt. The gem sends a JSON Schema with every request (see [Structured Output](#structured-output)).
|
|
55
|
+
|
|
55
56
|
### 2. Use It
|
|
56
57
|
|
|
57
58
|
```ruby
|
|
@@ -68,15 +69,8 @@ result.reasoning # => "Strong positive language with 'love' and 'absolutely'"
|
|
|
68
69
|
```ruby
|
|
69
70
|
# config/initializers/llm_classifier.rb
|
|
70
71
|
LlmClassifier.configure do |config|
|
|
71
|
-
#
|
|
72
|
-
config.
|
|
73
|
-
|
|
74
|
-
# Default model for classification
|
|
75
|
-
config.default_model = "gpt-4o-mini"
|
|
76
|
-
|
|
77
|
-
# API keys (reads from ENV by default)
|
|
78
|
-
config.openai_api_key = ENV["OPENAI_API_KEY"]
|
|
79
|
-
config.anthropic_api_key = ENV["ANTHROPIC_API_KEY"]
|
|
72
|
+
# Default model for classification. nil (the default) uses RubyLLM.config.default_model.
|
|
73
|
+
config.default_model = "claude-opus-5-5"
|
|
80
74
|
|
|
81
75
|
# Content fetching settings
|
|
82
76
|
config.web_fetch_timeout = 10
|
|
@@ -86,6 +80,44 @@ end
|
|
|
86
80
|
|
|
87
81
|
## Features
|
|
88
82
|
|
|
83
|
+
### Structured Output
|
|
84
|
+
|
|
85
|
+
Every request carries a JSON Schema built from the classifier's categories. ruby_llm passes it to the provider's native structured-output feature (for example, Anthropic's `output_config.format`), and the provider enforces it during generation. The model can't return malformed JSON or a category you didn't define. A refused or truncated response still comes back as a failed `Result`.
|
|
86
|
+
|
|
87
|
+
```ruby
|
|
88
|
+
SentimentClassifier.output_schema
|
|
89
|
+
# => {
|
|
90
|
+
# type: "object",
|
|
91
|
+
# properties: {
|
|
92
|
+
# reasoning: { type: "string", ... },
|
|
93
|
+
# category: { type: "string", enum: ["positive", "negative", "neutral"] },
|
|
94
|
+
# confidence: { type: "number", ... }
|
|
95
|
+
# },
|
|
96
|
+
# required: ["reasoning", "category", "confidence"],
|
|
97
|
+
# additionalProperties: false
|
|
98
|
+
# }
|
|
99
|
+
```
|
|
100
|
+
|
|
101
|
+
### Extra Output Fields
|
|
102
|
+
|
|
103
|
+
The schema doesn't allow fields you haven't declared. To have the model return more than the category, declare each extra field with `output_field`. Options are JSON Schema keywords, and `type` defaults to `"string"`. The values come back in `result.metadata`:
|
|
104
|
+
|
|
105
|
+
```ruby
|
|
106
|
+
class BusinessClassifier < LlmClassifier::Classifier
|
|
107
|
+
categories :dealership, :mechanic, :parts
|
|
108
|
+
multi_label true
|
|
109
|
+
output_field :evidence, description: "Words or brands that confirmed the category"
|
|
110
|
+
output_field :brands, type: "array", items: { type: "string" }
|
|
111
|
+
end
|
|
112
|
+
|
|
113
|
+
result = BusinessClassifier.classify("Joe's Harley-Davidson Service")
|
|
114
|
+
result.metadata # => { "evidence" => "Harley-Davidson, Service", "brands" => ["Harley-Davidson"] }
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
Single-label classifiers get a `category` string, so the model must pick exactly one. Multi-label classifiers get a `categories` array, which may be empty. Categories are matched case-insensitively, because providers guarantee enum membership but not capitalization.
|
|
118
|
+
|
|
119
|
+
The model needs to support structured outputs. Current Claude and OpenAI models do.
|
|
120
|
+
|
|
89
121
|
### Multi-label Classification
|
|
90
122
|
|
|
91
123
|
```ruby
|
|
@@ -100,6 +132,27 @@ result = TopicClassifier.classify("Building a Rails API with React frontend")
|
|
|
100
132
|
result.categories # => ["rails", "javascript"]
|
|
101
133
|
```
|
|
102
134
|
|
|
135
|
+
### Requiring Categories
|
|
136
|
+
|
|
137
|
+
By default, multi-label classifiers return `Result.success` even when no categories match (empty array). Use `require_categories` to treat empty results as failures:
|
|
138
|
+
|
|
139
|
+
```ruby
|
|
140
|
+
class StrictClassifier < LlmClassifier::Classifier
|
|
141
|
+
categories :mechanic, :instructor, :gear
|
|
142
|
+
multi_label true
|
|
143
|
+
require_categories true # Result.failure when no categories match
|
|
144
|
+
|
|
145
|
+
system_prompt "Classify this business..."
|
|
146
|
+
end
|
|
147
|
+
|
|
148
|
+
result = StrictClassifier.classify("Joe's Pizza Shop")
|
|
149
|
+
result.success? # => false (no motorcycle categories matched)
|
|
150
|
+
result.failure? # => true
|
|
151
|
+
result.error # => "No valid categories returned"
|
|
152
|
+
```
|
|
153
|
+
|
|
154
|
+
This is useful when classification is a filtering step and you need to distinguish "no match" from "classification succeeded."
|
|
155
|
+
|
|
103
156
|
### Domain Knowledge
|
|
104
157
|
|
|
105
158
|
Inject domain-specific knowledge into your prompts:
|
|
@@ -138,16 +191,20 @@ class AuditedClassifier < LlmClassifier::Classifier
|
|
|
138
191
|
end
|
|
139
192
|
```
|
|
140
193
|
|
|
141
|
-
### Override
|
|
194
|
+
### Override Model Per-Classifier
|
|
142
195
|
|
|
143
196
|
```ruby
|
|
144
197
|
class CriticalClassifier < LlmClassifier::Classifier
|
|
145
198
|
categories :high, :medium, :low
|
|
146
|
-
|
|
147
|
-
model "claude-sonnet-4-20250514" # Specific model
|
|
199
|
+
model "claude-opus-5-5"
|
|
148
200
|
end
|
|
201
|
+
|
|
202
|
+
# Or per call
|
|
203
|
+
CriticalClassifier.classify(text, model: "claude-haiku-4-5")
|
|
149
204
|
```
|
|
150
205
|
|
|
206
|
+
If ruby_llm raises `ModelNotFoundError` for a newly released model, refresh its registry with `RubyLLM.models.refresh!`.
|
|
207
|
+
|
|
151
208
|
## Rails Integration
|
|
152
209
|
|
|
153
210
|
### ActiveRecord Concern
|
|
@@ -226,22 +283,20 @@ Features:
|
|
|
226
283
|
|
|
227
284
|
## Adapters
|
|
228
285
|
|
|
229
|
-
|
|
230
|
-
|
|
231
|
-
- **`:ruby_llm`** - Uses the [ruby_llm](https://github.com/crmne/ruby_llm) gem (recommended)
|
|
232
|
-
- **`:openai`** - Direct OpenAI API integration
|
|
233
|
-
- **`:anthropic`** - Direct Anthropic API integration
|
|
286
|
+
The built-in `:ruby_llm` adapter (the default) routes requests through [ruby_llm](https://rubyllm.com), so any provider it supports works.
|
|
234
287
|
|
|
235
288
|
### Custom Adapter
|
|
236
289
|
|
|
237
290
|
```ruby
|
|
238
291
|
class MyCustomAdapter < LlmClassifier::Adapters::Base
|
|
239
|
-
def chat(model:, system_prompt:, user_prompt:)
|
|
240
|
-
#
|
|
292
|
+
def chat(model:, system_prompt:, user_prompt:, schema:)
|
|
293
|
+
# `schema` is the classifier's JSON Schema. Constrain the response to it and
|
|
294
|
+
# return the parsed Hash or the JSON string.
|
|
241
295
|
MyLlmClient.complete(
|
|
242
296
|
model: model,
|
|
243
297
|
system: system_prompt,
|
|
244
|
-
prompt: user_prompt
|
|
298
|
+
prompt: user_prompt,
|
|
299
|
+
json_schema: schema
|
|
245
300
|
)
|
|
246
301
|
end
|
|
247
302
|
end
|
|
@@ -264,8 +319,9 @@ result.category # => "primary_category" (first)
|
|
|
264
319
|
result.categories # => ["cat1", "cat2"] (all)
|
|
265
320
|
result.confidence # => 0.95
|
|
266
321
|
result.reasoning # => "Explanation from LLM"
|
|
267
|
-
result.raw_response # =>
|
|
268
|
-
result.
|
|
322
|
+
result.raw_response # => Response JSON string
|
|
323
|
+
result.model # => Model used (the classifier's, or ruby_llm's default)
|
|
324
|
+
result.metadata # => Values of declared output_field entries
|
|
269
325
|
result.error # => Error message if failed
|
|
270
326
|
result.to_h # => Hash representation
|
|
271
327
|
```
|
|
@@ -281,7 +337,7 @@ This project includes a [Dev Container](https://containers.dev/) configuration f
|
|
|
281
337
|
3. Press `Cmd+Shift+P` and select "Dev Containers: Reopen in Container"
|
|
282
338
|
4. Wait for the container to build and start
|
|
283
339
|
|
|
284
|
-
The container includes Ruby
|
|
340
|
+
The container includes Ruby, GitHub CLI, and useful VS Code extensions.
|
|
285
341
|
|
|
286
342
|
### Local Setup
|
|
287
343
|
|
|
@@ -0,0 +1,24 @@
|
|
|
1
|
+
# Adapters
|
|
2
|
+
|
|
3
|
+
LLM provider adapters. All inherit from `Adapters::Base` and implement `#chat(model:, system_prompt:, user_prompt:, schema:)`.
|
|
4
|
+
|
|
5
|
+
## Inventory
|
|
6
|
+
|
|
7
|
+
- `Base` - Abstract interface. Provides `#config` helper for accessing `LlmClassifier.configuration`
|
|
8
|
+
- `RubyLlm` - The only built-in adapter. Calls `RubyLLM.chat(...).with_instructions(...).with_schema(...).ask(...)` and returns a Hash with `:content`, `:input_tokens`, `:output_tokens`, `:model` (`chat.model.id`, the model actually used)
|
|
9
|
+
|
|
10
|
+
## Conventions
|
|
11
|
+
|
|
12
|
+
- `schema:` is the classifier's `output_schema` (JSON Schema, symbol keys), including any `output_field` declarations. Adapters must constrain the response to it
|
|
13
|
+
- `#chat` returns the content (a parsed Hash or a JSON String) or a wrapper Hash `{ content:, input_tokens:, output_tokens:, model: }`. `Classifier#extract_response_data` treats a Hash as the wrapper only when it has a `:content` key
|
|
14
|
+
- ruby_llm 1.x returns structured content as a Hash; 2.x returns a JSON String (`response.parsed` holds the Hash). `Classifier#parse_response` accepts both
|
|
15
|
+
- Token counts come from `response.tokens.input` / `.output`, which exist in ruby_llm 1.13+ and 2.x (`input_tokens` readers were removed in 2.0)
|
|
16
|
+
- `model: nil` lets ruby_llm fall back to `RubyLLM.config.default_model`
|
|
17
|
+
- Provider credentials are configured in `RubyLLM.configure`, not in this gem
|
|
18
|
+
- Custom adapters are passed as a Class to `config.adapter` or the `adapter` DSL
|
|
19
|
+
- Anthropic schema limits: no `minimum`/`maximum`, no `maxItems`, `minItems` only 0 or 1, `additionalProperties: false` required. Don't add `minItems: 1` for `require_categories`: an empty array is how the model signals "no match"
|
|
20
|
+
|
|
21
|
+
## Related
|
|
22
|
+
|
|
23
|
+
- [../content_fetchers/CLAUDE.md](../content_fetchers/CLAUDE.md) - Content fetchers
|
|
24
|
+
- [../../spec/CLAUDE.md](../../spec/CLAUDE.md) - Testing conventions
|
|
@@ -4,7 +4,10 @@ module LlmClassifier
|
|
|
4
4
|
module Adapters
|
|
5
5
|
# Base adapter class for LLM providers
|
|
6
6
|
class Base
|
|
7
|
-
|
|
7
|
+
# schema is a JSON Schema Hash the response must conform to. Return the response
|
|
8
|
+
# content (a parsed Hash or a JSON String), or a wrapper Hash with a :content key:
|
|
9
|
+
# { content:, input_tokens:, output_tokens:, model: }.
|
|
10
|
+
def chat(model:, system_prompt:, user_prompt:, schema:)
|
|
8
11
|
raise NotImplementedError, "Subclasses must implement #chat"
|
|
9
12
|
end
|
|
10
13
|
|
|
@@ -2,28 +2,24 @@
|
|
|
2
2
|
|
|
3
3
|
module LlmClassifier
|
|
4
4
|
module Adapters
|
|
5
|
-
# Adapter for the ruby_llm gem
|
|
5
|
+
# Adapter for the ruby_llm gem. Provider credentials and the fallback model come from
|
|
6
|
+
# RubyLLM's own configuration.
|
|
6
7
|
class RubyLlm < Base
|
|
7
|
-
def chat(model:, system_prompt:, user_prompt:)
|
|
8
|
-
|
|
8
|
+
def chat(model:, system_prompt:, user_prompt:, schema:)
|
|
9
|
+
require "ruby_llm" unless defined?(::RubyLLM)
|
|
9
10
|
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
11
|
+
chat = ::RubyLLM.chat(model: model)
|
|
12
|
+
response = chat.with_instructions(system_prompt)
|
|
13
|
+
.with_schema(name: "classification", schema: schema, strict: true)
|
|
14
|
+
.ask(user_prompt)
|
|
13
15
|
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
22
|
-
begin
|
|
23
|
-
require "ruby_llm"
|
|
24
|
-
rescue LoadError
|
|
25
|
-
raise AdapterError, "ruby_llm gem is not installed. Add it to your Gemfile: gem 'ruby_llm'"
|
|
26
|
-
end
|
|
16
|
+
# ruby_llm 1.x returns structured content as a Hash, 2.x as a JSON String.
|
|
17
|
+
{
|
|
18
|
+
content: response.content,
|
|
19
|
+
input_tokens: response.tokens&.input,
|
|
20
|
+
output_tokens: response.tokens&.output,
|
|
21
|
+
model: chat.model&.id
|
|
22
|
+
}
|
|
27
23
|
end
|
|
28
24
|
end
|
|
29
25
|
end
|
|
@@ -5,9 +5,16 @@ require "json"
|
|
|
5
5
|
module LlmClassifier
|
|
6
6
|
# Base classifier class that provides a DSL for defining LLM-powered classifiers
|
|
7
7
|
class Classifier
|
|
8
|
+
# Response fields the classifier itself defines.
|
|
9
|
+
BUILT_IN_FIELDS = %w[reasoning category categories confidence].freeze
|
|
10
|
+
# Names output_field can't take. "content" would make an adapter's bare parsed Hash
|
|
11
|
+
# look like the { content: } response wrapper.
|
|
12
|
+
RESERVED_FIELDS = (BUILT_IN_FIELDS + %w[content]).freeze
|
|
13
|
+
|
|
8
14
|
class << self
|
|
9
15
|
attr_reader :defined_categories, :defined_system_prompt, :defined_model,
|
|
10
|
-
:defined_adapter, :defined_multi_label, :
|
|
16
|
+
:defined_adapter, :defined_multi_label, :defined_require_categories,
|
|
17
|
+
:defined_knowledge,
|
|
11
18
|
:before_classify_callbacks, :after_classify_callbacks
|
|
12
19
|
|
|
13
20
|
def categories(*cats)
|
|
@@ -50,6 +57,14 @@ module LlmClassifier
|
|
|
50
57
|
end
|
|
51
58
|
end
|
|
52
59
|
|
|
60
|
+
def require_categories(value = nil)
|
|
61
|
+
if value.nil?
|
|
62
|
+
@defined_require_categories || false
|
|
63
|
+
else
|
|
64
|
+
@defined_require_categories = value
|
|
65
|
+
end
|
|
66
|
+
end
|
|
67
|
+
|
|
53
68
|
def knowledge(&)
|
|
54
69
|
if block_given?
|
|
55
70
|
@defined_knowledge = Knowledge.new
|
|
@@ -58,6 +73,19 @@ module LlmClassifier
|
|
|
58
73
|
@defined_knowledge
|
|
59
74
|
end
|
|
60
75
|
|
|
76
|
+
# Declares an extra response field, returned in Result#metadata. Options are JSON Schema
|
|
77
|
+
# keywords for the field (type defaults to "string").
|
|
78
|
+
def output_field(name, **schema)
|
|
79
|
+
name = name.to_s
|
|
80
|
+
raise ArgumentError, "#{name} is a reserved output field name" if RESERVED_FIELDS.include?(name)
|
|
81
|
+
|
|
82
|
+
output_fields[name] = { type: "string" }.merge(schema)
|
|
83
|
+
end
|
|
84
|
+
|
|
85
|
+
def output_fields
|
|
86
|
+
@output_fields ||= {}
|
|
87
|
+
end
|
|
88
|
+
|
|
61
89
|
def before_classify(&block)
|
|
62
90
|
@before_classify_callbacks ||= []
|
|
63
91
|
@before_classify_callbacks << block
|
|
@@ -68,8 +96,31 @@ module LlmClassifier
|
|
|
68
96
|
@after_classify_callbacks << block
|
|
69
97
|
end
|
|
70
98
|
|
|
71
|
-
def classify(input, **
|
|
72
|
-
new(input, **
|
|
99
|
+
def classify(input, **)
|
|
100
|
+
new(input, **).classify
|
|
101
|
+
end
|
|
102
|
+
|
|
103
|
+
# JSON Schema the LLM response is constrained to. Reasoning comes first so the
|
|
104
|
+
# model explains itself before committing to a label.
|
|
105
|
+
def output_schema
|
|
106
|
+
label_key = multi_label ? :categories : :category
|
|
107
|
+
properties = {
|
|
108
|
+
reasoning: { type: "string", description: "Brief explanation of the classification" },
|
|
109
|
+
label_key => label_schema,
|
|
110
|
+
confidence: { type: "number", description: "Confidence from 0.0 to 1.0" }
|
|
111
|
+
}.merge(output_fields.transform_keys(&:to_sym))
|
|
112
|
+
|
|
113
|
+
{ type: "object", properties: properties, required: properties.keys.map(&:to_s), additionalProperties: false }
|
|
114
|
+
end
|
|
115
|
+
|
|
116
|
+
private
|
|
117
|
+
|
|
118
|
+
def label_schema
|
|
119
|
+
category = { type: "string" }
|
|
120
|
+
category[:enum] = categories if categories.any?
|
|
121
|
+
return category unless multi_label
|
|
122
|
+
|
|
123
|
+
{ type: "array", items: category, description: "Every category that applies. Empty if none apply." }
|
|
73
124
|
end
|
|
74
125
|
end
|
|
75
126
|
|
|
@@ -103,26 +154,36 @@ module LlmClassifier
|
|
|
103
154
|
|
|
104
155
|
def perform_classification(processed_input)
|
|
105
156
|
adapter_instance = build_adapter
|
|
157
|
+
resolved_model = options[:model] || self.class.model
|
|
106
158
|
response = adapter_instance.chat(
|
|
107
|
-
model:
|
|
159
|
+
model: resolved_model,
|
|
108
160
|
system_prompt: build_system_prompt,
|
|
109
|
-
user_prompt: build_user_prompt(processed_input)
|
|
161
|
+
user_prompt: build_user_prompt(processed_input),
|
|
162
|
+
schema: self.class.output_schema
|
|
110
163
|
)
|
|
111
164
|
|
|
112
|
-
|
|
165
|
+
content, response_meta = extract_response_data(response)
|
|
166
|
+
parse_response(content, resolved_model || response_meta[:model], response_meta)
|
|
167
|
+
end
|
|
168
|
+
|
|
169
|
+
# Adapters return the content itself, or wrap it as { content:, input_tokens:, output_tokens:, model: }.
|
|
170
|
+
def extract_response_data(response)
|
|
171
|
+
return [response, {}] unless response.is_a?(Hash) && response.key?(:content)
|
|
172
|
+
|
|
173
|
+
[response[:content], response.slice(:input_tokens, :output_tokens, :model)]
|
|
113
174
|
end
|
|
114
175
|
|
|
115
176
|
def build_adapter
|
|
116
177
|
adapter_name = self.class.adapter
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
|
|
122
|
-
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
|
|
178
|
+
case adapter_name
|
|
179
|
+
when :ruby_llm then Adapters::RubyLlm.new
|
|
180
|
+
when Class then adapter_name.new
|
|
181
|
+
when :openai, :anthropic
|
|
182
|
+
raise AdapterError, "The :#{adapter_name} adapter was removed in 0.3.0. " \
|
|
183
|
+
"Configure the provider in RubyLLM and use the :ruby_llm adapter."
|
|
184
|
+
else
|
|
185
|
+
raise AdapterError, "Unknown adapter: #{adapter_name}"
|
|
186
|
+
end
|
|
126
187
|
end
|
|
127
188
|
|
|
128
189
|
def build_system_prompt
|
|
@@ -138,16 +199,8 @@ module LlmClassifier
|
|
|
138
199
|
categories = self.class.categories.join(", ")
|
|
139
200
|
multi = self.class.multi_label
|
|
140
201
|
|
|
141
|
-
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
Respond with ONLY a JSON object in this format:
|
|
145
|
-
{
|
|
146
|
-
"categories": [#{multi ? '"category1", "category2"' : '"category"'}],
|
|
147
|
-
"confidence": 0.0-1.0,
|
|
148
|
-
"reasoning": "Brief explanation"
|
|
149
|
-
}
|
|
150
|
-
PROMPT
|
|
202
|
+
scope = multi ? "every category that applies (none if none apply)" : "exactly one of these categories"
|
|
203
|
+
"You are a classifier. Classify the given input into #{scope}: #{categories}."
|
|
151
204
|
end
|
|
152
205
|
|
|
153
206
|
def build_user_prompt(processed_input)
|
|
@@ -161,24 +214,31 @@ module LlmClassifier
|
|
|
161
214
|
end
|
|
162
215
|
end
|
|
163
216
|
|
|
164
|
-
|
|
165
|
-
|
|
217
|
+
# Adapters return structured output either already parsed (a Hash) or as JSON text.
|
|
218
|
+
def parse_response(content, resolved_model = nil, token_data = {})
|
|
219
|
+
json = content.is_a?(Hash) ? content.transform_keys(&:to_s) : JSON.parse(content.to_s)
|
|
220
|
+
raw_response = content.is_a?(String) ? content : JSON.generate(json)
|
|
166
221
|
valid_categories = extract_valid_categories(json)
|
|
167
222
|
|
|
168
|
-
return build_failure_result(
|
|
223
|
+
return build_failure_result(raw_response, json) if should_fail?(valid_categories)
|
|
169
224
|
|
|
170
|
-
build_success_result(json, valid_categories,
|
|
225
|
+
build_success_result(json, valid_categories, raw_response, resolved_model, token_data)
|
|
171
226
|
rescue JSON::ParserError => e
|
|
172
|
-
Result.failure(error: "Failed to parse response: #{e.message}", raw_response:
|
|
227
|
+
Result.failure(error: "Failed to parse response: #{e.message}", raw_response: content)
|
|
173
228
|
end
|
|
174
229
|
|
|
230
|
+
# Structured outputs guarantee enum membership but not capitalization, so match
|
|
231
|
+
# case-insensitively and return the category as defined.
|
|
175
232
|
def extract_valid_categories(json)
|
|
176
|
-
|
|
177
|
-
|
|
233
|
+
defined = self.class.categories.to_h { |c| [c.downcase, c] }
|
|
234
|
+
Array(json["categories"] || json["category"]).filter_map { |c| defined[c.to_s.downcase] }.uniq
|
|
178
235
|
end
|
|
179
236
|
|
|
180
237
|
def should_fail?(valid_categories)
|
|
181
|
-
|
|
238
|
+
return false if valid_categories.any?
|
|
239
|
+
return false if self.class.categories.empty?
|
|
240
|
+
|
|
241
|
+
!self.class.multi_label || self.class.require_categories
|
|
182
242
|
end
|
|
183
243
|
|
|
184
244
|
def build_failure_result(response, json)
|
|
@@ -189,17 +249,19 @@ module LlmClassifier
|
|
|
189
249
|
)
|
|
190
250
|
end
|
|
191
251
|
|
|
192
|
-
def build_success_result(json, valid_categories, response)
|
|
252
|
+
def build_success_result(json, valid_categories, response, resolved_model = nil, token_data = {})
|
|
193
253
|
categories = self.class.multi_label ? valid_categories : [valid_categories.first].compact
|
|
194
|
-
|
|
195
|
-
metadata = json.reject { |k, _| excluded_keys.include?(k) }
|
|
254
|
+
metadata = json.reject { |k, _| BUILT_IN_FIELDS.include?(k) }
|
|
196
255
|
|
|
197
256
|
Result.success(
|
|
198
257
|
categories: categories,
|
|
199
258
|
confidence: json["confidence"]&.to_f,
|
|
200
259
|
reasoning: json["reasoning"],
|
|
201
260
|
raw_response: response,
|
|
202
|
-
metadata: metadata
|
|
261
|
+
metadata: metadata,
|
|
262
|
+
model: resolved_model,
|
|
263
|
+
input_tokens: token_data[:input_tokens],
|
|
264
|
+
output_tokens: token_data[:output_tokens]
|
|
203
265
|
)
|
|
204
266
|
end
|
|
205
267
|
end
|
|
@@ -5,34 +5,16 @@ require "logger"
|
|
|
5
5
|
module LlmClassifier
|
|
6
6
|
# Configuration object for LlmClassifier settings
|
|
7
7
|
class Configuration
|
|
8
|
-
attr_accessor :adapter, :default_model, :
|
|
9
|
-
:
|
|
10
|
-
:logger
|
|
8
|
+
attr_accessor :adapter, :default_model, :web_fetch_timeout, :web_fetch_user_agent,
|
|
9
|
+
:default_queue, :logger
|
|
11
10
|
|
|
12
11
|
def initialize
|
|
13
12
|
@adapter = :ruby_llm
|
|
14
|
-
@default_model =
|
|
15
|
-
@openai_api_key = ENV.fetch("OPENAI_API_KEY", nil)
|
|
16
|
-
@anthropic_api_key = ENV.fetch("ANTHROPIC_API_KEY", nil)
|
|
13
|
+
@default_model = nil # nil defers to RubyLLM.config.default_model
|
|
17
14
|
@web_fetch_timeout = 10
|
|
18
15
|
@web_fetch_user_agent = "LlmClassifier/#{VERSION}"
|
|
19
16
|
@default_queue = :classification
|
|
20
17
|
@logger = defined?(::Rails) ? ::Rails.logger : Logger.new($stdout)
|
|
21
18
|
end
|
|
22
|
-
|
|
23
|
-
def adapter_class
|
|
24
|
-
case adapter
|
|
25
|
-
when :ruby_llm
|
|
26
|
-
Adapters::RubyLlm
|
|
27
|
-
when :openai
|
|
28
|
-
Adapters::OpenAI
|
|
29
|
-
when :anthropic
|
|
30
|
-
Adapters::Anthropic
|
|
31
|
-
when Class
|
|
32
|
-
adapter
|
|
33
|
-
else
|
|
34
|
-
raise ConfigurationError, "Unknown adapter: #{adapter}"
|
|
35
|
-
end
|
|
36
|
-
end
|
|
37
19
|
end
|
|
38
20
|
end
|
|
@@ -0,0 +1,27 @@
|
|
|
1
|
+
# Content Fetchers
|
|
2
|
+
|
|
3
|
+
Utilities for fetching external content to use as classification input. Not wired into `Classifier` automatically -- callers fetch content and pass it in.
|
|
4
|
+
|
|
5
|
+
## Inventory
|
|
6
|
+
|
|
7
|
+
- `Base` - Abstract interface. Subclasses implement `#fetch(source)`
|
|
8
|
+
- `Web` - HTTP fetcher with SSRF protection, redirect following, and HTML text extraction
|
|
9
|
+
- `Null` - No-op fetcher, always returns `nil`
|
|
10
|
+
|
|
11
|
+
## SSRF Protection (`Web`)
|
|
12
|
+
|
|
13
|
+
- Validates resolved IPs against private/loopback CIDR ranges before connecting
|
|
14
|
+
- Follows up to 3 redirects, re-validating each redirect target
|
|
15
|
+
- `normalize_redirect_url` handles relative and absolute redirect URLs
|
|
16
|
+
- Uses `nil? || empty?` guards (not ActiveSupport `.blank?`) to avoid the dependency
|
|
17
|
+
|
|
18
|
+
## HTML Processing (`Web`)
|
|
19
|
+
|
|
20
|
+
- Nokogiri is lazily loaded (`require "nokogiri"` inside the method) since it's an optional dependency
|
|
21
|
+
- Strips `<script>`, `<style>`, `<nav>`, `<footer>`, `<header>` elements
|
|
22
|
+
- Truncates extracted text to 2000 characters
|
|
23
|
+
|
|
24
|
+
## Related
|
|
25
|
+
|
|
26
|
+
- [../adapters/CLAUDE.md](../adapters/CLAUDE.md) - LLM adapters
|
|
27
|
+
- [../../spec/CLAUDE.md](../../spec/CLAUDE.md) - Testing conventions
|
|
@@ -0,0 +1,25 @@
|
|
|
1
|
+
# Rails Integration
|
|
2
|
+
|
|
3
|
+
This entire subtree is excluded from Zeitwerk autoloading (`loader.ignore` in `lib/llm_classifier.rb`) and loaded manually only when `Rails::Railtie` is defined. This keeps Rails as an optional dependency.
|
|
4
|
+
|
|
5
|
+
## Inventory
|
|
6
|
+
|
|
7
|
+
- `Railtie` - Sets `Rails.logger` as the default LlmClassifier logger
|
|
8
|
+
- `Concerns::Classifiable` - ActiveRecord concern adding a `classifies` macro
|
|
9
|
+
- `Generators::InstallGenerator` - `rails g llm_classifier:install` scaffolds an initializer
|
|
10
|
+
- `Generators::ClassifierGenerator` - `rails g llm_classifier:classifier Name cat1 cat2` scaffolds a classifier and spec
|
|
11
|
+
|
|
12
|
+
## Classifiable Concern
|
|
13
|
+
|
|
14
|
+
The `classifies` macro defines three instance methods per classification:
|
|
15
|
+
|
|
16
|
+
- `classify_<attr>!` - Runs classification and stores the result
|
|
17
|
+
- `<attr>_category` / `<attr>_categories` - Reads stored category data
|
|
18
|
+
- `<attr>_classification` - Returns the full stored classification hash
|
|
19
|
+
|
|
20
|
+
Results are written into a JSONB column (via `store_in:`) or a transient instance variable if no column is specified.
|
|
21
|
+
|
|
22
|
+
## Related
|
|
23
|
+
|
|
24
|
+
- [../adapters/CLAUDE.md](../adapters/CLAUDE.md) - LLM adapters
|
|
25
|
+
- [../content_fetchers/CLAUDE.md](../content_fetchers/CLAUDE.md) - Content fetchers
|
|
@@ -14,16 +14,10 @@ module LlmClassifier
|
|
|
14
14
|
create_file "config/initializers/llm_classifier.rb", <<~RUBY
|
|
15
15
|
# frozen_string_literal: true
|
|
16
16
|
|
|
17
|
+
# Provider API keys are configured in RubyLLM (config/initializers/ruby_llm.rb).
|
|
17
18
|
LlmClassifier.configure do |config|
|
|
18
|
-
#
|
|
19
|
-
config.
|
|
20
|
-
|
|
21
|
-
# Default model for classification
|
|
22
|
-
config.default_model = "gpt-4o-mini"
|
|
23
|
-
|
|
24
|
-
# API keys (reads from ENV by default)
|
|
25
|
-
# config.openai_api_key = ENV["OPENAI_API_KEY"]
|
|
26
|
-
# config.anthropic_api_key = ENV["ANTHROPIC_API_KEY"]
|
|
19
|
+
# Default model for classification. nil uses RubyLLM.config.default_model.
|
|
20
|
+
# config.default_model = "claude-opus-5-5"
|
|
27
21
|
|
|
28
22
|
# Content fetching settings
|
|
29
23
|
config.web_fetch_timeout = 10
|
|
@@ -45,7 +39,7 @@ module LlmClassifier
|
|
|
45
39
|
say "LlmClassifier installed successfully!", :green
|
|
46
40
|
say "\n"
|
|
47
41
|
say "Next steps:"
|
|
48
|
-
say " 1. Configure your API keys
|
|
42
|
+
say " 1. Configure your provider API keys with RubyLLM.configure"
|
|
49
43
|
say " 2. Generate a classifier: rails g llm_classifier:classifier SentimentClassifier"
|
|
50
44
|
say "\n"
|
|
51
45
|
end
|
|
@@ -7,10 +7,7 @@ class <%= class_name %> < LlmClassifier::Classifier
|
|
|
7
7
|
# multi_label true
|
|
8
8
|
|
|
9
9
|
# Uncomment to override the default model
|
|
10
|
-
# model "
|
|
11
|
-
|
|
12
|
-
# Uncomment to override the default adapter
|
|
13
|
-
# adapter :openai
|
|
10
|
+
# model "claude-opus-5-5"
|
|
14
11
|
|
|
15
12
|
system_prompt <<~PROMPT
|
|
16
13
|
You are a classifier. Analyze the given input and classify it into the appropriate category.
|
|
@@ -19,13 +16,6 @@ class <%= class_name %> < LlmClassifier::Classifier
|
|
|
19
16
|
<% categories_array.each do |cat| -%>
|
|
20
17
|
- <%= cat %>: [describe what this category means]
|
|
21
18
|
<% end -%>
|
|
22
|
-
|
|
23
|
-
Respond with ONLY a JSON object:
|
|
24
|
-
{
|
|
25
|
-
"categories": ["category"],
|
|
26
|
-
"confidence": 0.0-1.0,
|
|
27
|
-
"reasoning": "Brief explanation"
|
|
28
|
-
}
|
|
29
19
|
PROMPT
|
|
30
20
|
|
|
31
21
|
# Uncomment to add domain knowledge
|
|
@@ -3,15 +3,21 @@
|
|
|
3
3
|
module LlmClassifier
|
|
4
4
|
# Result object returned from classification operations
|
|
5
5
|
class Result
|
|
6
|
-
attr_reader :categories, :confidence, :reasoning, :raw_response, :metadata, :error
|
|
6
|
+
attr_reader :categories, :confidence, :reasoning, :raw_response, :metadata, :error, :model,
|
|
7
|
+
:input_tokens, :output_tokens
|
|
7
8
|
|
|
8
|
-
def initialize(categories: [], confidence: nil, reasoning: nil,
|
|
9
|
+
def initialize(categories: [], confidence: nil, reasoning: nil,
|
|
10
|
+
raw_response: nil, error: nil, metadata: {},
|
|
11
|
+
model: nil, input_tokens: nil, output_tokens: nil)
|
|
9
12
|
@categories = Array(categories)
|
|
10
13
|
@confidence = confidence
|
|
11
14
|
@reasoning = reasoning
|
|
12
15
|
@raw_response = raw_response
|
|
13
16
|
@metadata = metadata
|
|
14
17
|
@error = error
|
|
18
|
+
@model = model
|
|
19
|
+
@input_tokens = input_tokens
|
|
20
|
+
@output_tokens = output_tokens
|
|
15
21
|
end
|
|
16
22
|
|
|
17
23
|
def success?
|
|
@@ -38,18 +44,26 @@ module LlmClassifier
|
|
|
38
44
|
confidence: @confidence,
|
|
39
45
|
reasoning: @reasoning,
|
|
40
46
|
metadata: @metadata,
|
|
41
|
-
error: @error
|
|
47
|
+
error: @error,
|
|
48
|
+
model: @model,
|
|
49
|
+
input_tokens: @input_tokens,
|
|
50
|
+
output_tokens: @output_tokens
|
|
42
51
|
}
|
|
43
52
|
end
|
|
44
53
|
|
|
45
54
|
class << self
|
|
46
|
-
def success(categories:, confidence: nil, reasoning: nil,
|
|
55
|
+
def success(categories:, confidence: nil, reasoning: nil,
|
|
56
|
+
raw_response: nil, metadata: {},
|
|
57
|
+
model: nil, input_tokens: nil, output_tokens: nil)
|
|
47
58
|
new(
|
|
48
59
|
categories: categories,
|
|
49
60
|
confidence: confidence,
|
|
50
61
|
reasoning: reasoning,
|
|
51
62
|
raw_response: raw_response,
|
|
52
|
-
metadata: metadata
|
|
63
|
+
metadata: metadata,
|
|
64
|
+
model: model,
|
|
65
|
+
input_tokens: input_tokens,
|
|
66
|
+
output_tokens: output_tokens
|
|
53
67
|
)
|
|
54
68
|
end
|
|
55
69
|
|
data/lib/llm_classifier.rb
CHANGED
|
@@ -3,10 +3,7 @@
|
|
|
3
3
|
require "zeitwerk"
|
|
4
4
|
|
|
5
5
|
loader = Zeitwerk::Loader.for_gem
|
|
6
|
-
loader.inflector.inflect(
|
|
7
|
-
"openai" => "OpenAI",
|
|
8
|
-
"ruby_llm" => "RubyLlm"
|
|
9
|
-
)
|
|
6
|
+
loader.inflector.inflect("ruby_llm" => "RubyLlm")
|
|
10
7
|
loader.ignore("#{__dir__}/llm_classifier/rails")
|
|
11
8
|
loader.setup
|
|
12
9
|
|
metadata
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: llm_classifier
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 0.
|
|
4
|
+
version: 0.3.0
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Dmitry Sychev
|
|
@@ -9,6 +9,26 @@ bindir: exe
|
|
|
9
9
|
cert_chain: []
|
|
10
10
|
date: 1980-01-02 00:00:00.000000000 Z
|
|
11
11
|
dependencies:
|
|
12
|
+
- !ruby/object:Gem::Dependency
|
|
13
|
+
name: ruby_llm
|
|
14
|
+
requirement: !ruby/object:Gem::Requirement
|
|
15
|
+
requirements:
|
|
16
|
+
- - ">="
|
|
17
|
+
- !ruby/object:Gem::Version
|
|
18
|
+
version: '1.14'
|
|
19
|
+
- - "<"
|
|
20
|
+
- !ruby/object:Gem::Version
|
|
21
|
+
version: '3'
|
|
22
|
+
type: :runtime
|
|
23
|
+
prerelease: false
|
|
24
|
+
version_requirements: !ruby/object:Gem::Requirement
|
|
25
|
+
requirements:
|
|
26
|
+
- - ">="
|
|
27
|
+
- !ruby/object:Gem::Version
|
|
28
|
+
version: '1.14'
|
|
29
|
+
- - "<"
|
|
30
|
+
- !ruby/object:Gem::Version
|
|
31
|
+
version: '3'
|
|
12
32
|
- !ruby/object:Gem::Dependency
|
|
13
33
|
name: zeitwerk
|
|
14
34
|
requirement: !ruby/object:Gem::Requirement
|
|
@@ -24,8 +44,8 @@ dependencies:
|
|
|
24
44
|
- !ruby/object:Gem::Version
|
|
25
45
|
version: '2.6'
|
|
26
46
|
description: A flexible Ruby gem for building LLM-based classifiers. Define categories,
|
|
27
|
-
system prompts, and domain knowledge using a clean DSL.
|
|
28
|
-
|
|
47
|
+
system prompts, and domain knowledge using a clean DSL. Responses are schema-constrained
|
|
48
|
+
via ruby_llm structured outputs, with optional Rails integration.
|
|
29
49
|
email:
|
|
30
50
|
- dmitry.sychev@axiumfoundry.com
|
|
31
51
|
executables: []
|
|
@@ -38,20 +58,22 @@ files:
|
|
|
38
58
|
- ".rspec"
|
|
39
59
|
- ".rubocop.yml"
|
|
40
60
|
- CHANGELOG.md
|
|
61
|
+
- CLAUDE.md
|
|
41
62
|
- LICENSE.txt
|
|
42
63
|
- README.md
|
|
43
64
|
- Rakefile
|
|
44
65
|
- lib/llm_classifier.rb
|
|
45
|
-
- lib/llm_classifier/adapters/
|
|
66
|
+
- lib/llm_classifier/adapters/CLAUDE.md
|
|
46
67
|
- lib/llm_classifier/adapters/base.rb
|
|
47
|
-
- lib/llm_classifier/adapters/openai.rb
|
|
48
68
|
- lib/llm_classifier/adapters/ruby_llm.rb
|
|
49
69
|
- lib/llm_classifier/classifier.rb
|
|
50
70
|
- lib/llm_classifier/configuration.rb
|
|
71
|
+
- lib/llm_classifier/content_fetchers/CLAUDE.md
|
|
51
72
|
- lib/llm_classifier/content_fetchers/base.rb
|
|
52
73
|
- lib/llm_classifier/content_fetchers/null.rb
|
|
53
74
|
- lib/llm_classifier/content_fetchers/web.rb
|
|
54
75
|
- lib/llm_classifier/knowledge.rb
|
|
76
|
+
- lib/llm_classifier/rails/CLAUDE.md
|
|
55
77
|
- lib/llm_classifier/rails/concerns/classifiable.rb
|
|
56
78
|
- lib/llm_classifier/rails/generators/classifier_generator.rb
|
|
57
79
|
- lib/llm_classifier/rails/generators/install_generator.rb
|
|
@@ -75,7 +97,7 @@ required_ruby_version: !ruby/object:Gem::Requirement
|
|
|
75
97
|
requirements:
|
|
76
98
|
- - ">="
|
|
77
99
|
- !ruby/object:Gem::Version
|
|
78
|
-
version: 3.
|
|
100
|
+
version: 3.2.0
|
|
79
101
|
required_rubygems_version: !ruby/object:Gem::Requirement
|
|
80
102
|
requirements:
|
|
81
103
|
- - ">="
|
|
@@ -1,72 +0,0 @@
|
|
|
1
|
-
# frozen_string_literal: true
|
|
2
|
-
|
|
3
|
-
require "net/http"
|
|
4
|
-
require "json"
|
|
5
|
-
require "uri"
|
|
6
|
-
|
|
7
|
-
module LlmClassifier
|
|
8
|
-
module Adapters
|
|
9
|
-
# Adapter for Anthropic API
|
|
10
|
-
class Anthropic < Base
|
|
11
|
-
API_URL = "https://api.anthropic.com/v1/messages"
|
|
12
|
-
API_VERSION = "2023-06-01"
|
|
13
|
-
|
|
14
|
-
def chat(model:, system_prompt:, user_prompt:)
|
|
15
|
-
api_key = validate_api_key
|
|
16
|
-
response = send_request(model, system_prompt, user_prompt, api_key)
|
|
17
|
-
parse_response(response)
|
|
18
|
-
end
|
|
19
|
-
|
|
20
|
-
private
|
|
21
|
-
|
|
22
|
-
def validate_api_key
|
|
23
|
-
api_key = config.anthropic_api_key
|
|
24
|
-
raise ConfigurationError, "Anthropic API key not configured" unless api_key
|
|
25
|
-
|
|
26
|
-
api_key
|
|
27
|
-
end
|
|
28
|
-
|
|
29
|
-
def send_request(model, system_prompt, user_prompt, api_key)
|
|
30
|
-
uri = URI(API_URL)
|
|
31
|
-
http = build_http_client(uri)
|
|
32
|
-
request = build_request(uri, api_key, model, system_prompt, user_prompt)
|
|
33
|
-
http.request(request)
|
|
34
|
-
end
|
|
35
|
-
|
|
36
|
-
def build_http_client(uri)
|
|
37
|
-
http = Net::HTTP.new(uri.host, uri.port)
|
|
38
|
-
http.use_ssl = true
|
|
39
|
-
http
|
|
40
|
-
end
|
|
41
|
-
|
|
42
|
-
def build_request(uri, api_key, model, system_prompt, user_prompt)
|
|
43
|
-
request = Net::HTTP::Post.new(uri)
|
|
44
|
-
request["Content-Type"] = "application/json"
|
|
45
|
-
request["x-api-key"] = api_key
|
|
46
|
-
request["anthropic-version"] = API_VERSION
|
|
47
|
-
request.body = build_request_body(model, system_prompt, user_prompt)
|
|
48
|
-
request
|
|
49
|
-
end
|
|
50
|
-
|
|
51
|
-
def build_request_body(model, system_prompt, user_prompt)
|
|
52
|
-
{
|
|
53
|
-
model: model,
|
|
54
|
-
max_tokens: 1024,
|
|
55
|
-
system: system_prompt,
|
|
56
|
-
messages: [
|
|
57
|
-
{ role: "user", content: user_prompt }
|
|
58
|
-
]
|
|
59
|
-
}.to_json
|
|
60
|
-
end
|
|
61
|
-
|
|
62
|
-
def parse_response(response)
|
|
63
|
-
unless response.is_a?(Net::HTTPSuccess)
|
|
64
|
-
raise AdapterError, "Anthropic API error: #{response.code} - #{response.body}"
|
|
65
|
-
end
|
|
66
|
-
|
|
67
|
-
parsed = JSON.parse(response.body)
|
|
68
|
-
parsed.dig("content", 0, "text")
|
|
69
|
-
end
|
|
70
|
-
end
|
|
71
|
-
end
|
|
72
|
-
end
|
|
@@ -1,70 +0,0 @@
|
|
|
1
|
-
# frozen_string_literal: true
|
|
2
|
-
|
|
3
|
-
require "net/http"
|
|
4
|
-
require "json"
|
|
5
|
-
require "uri"
|
|
6
|
-
|
|
7
|
-
module LlmClassifier
|
|
8
|
-
module Adapters
|
|
9
|
-
# Adapter for OpenAI API
|
|
10
|
-
class OpenAI < Base
|
|
11
|
-
API_URL = "https://api.openai.com/v1/chat/completions"
|
|
12
|
-
|
|
13
|
-
def chat(model:, system_prompt:, user_prompt:)
|
|
14
|
-
api_key = validate_api_key
|
|
15
|
-
response = send_request(model, system_prompt, user_prompt, api_key)
|
|
16
|
-
parse_response(response)
|
|
17
|
-
end
|
|
18
|
-
|
|
19
|
-
private
|
|
20
|
-
|
|
21
|
-
def validate_api_key
|
|
22
|
-
api_key = config.openai_api_key
|
|
23
|
-
raise ConfigurationError, "OpenAI API key not configured" unless api_key
|
|
24
|
-
|
|
25
|
-
api_key
|
|
26
|
-
end
|
|
27
|
-
|
|
28
|
-
def send_request(model, system_prompt, user_prompt, api_key)
|
|
29
|
-
uri = URI(API_URL)
|
|
30
|
-
http = build_http_client(uri)
|
|
31
|
-
request = build_request(uri, api_key, model, system_prompt, user_prompt)
|
|
32
|
-
http.request(request)
|
|
33
|
-
end
|
|
34
|
-
|
|
35
|
-
def build_http_client(uri)
|
|
36
|
-
http = Net::HTTP.new(uri.host, uri.port)
|
|
37
|
-
http.use_ssl = true
|
|
38
|
-
http
|
|
39
|
-
end
|
|
40
|
-
|
|
41
|
-
def build_request(uri, api_key, model, system_prompt, user_prompt)
|
|
42
|
-
request = Net::HTTP::Post.new(uri)
|
|
43
|
-
request["Content-Type"] = "application/json"
|
|
44
|
-
request["Authorization"] = "Bearer #{api_key}"
|
|
45
|
-
request.body = build_request_body(model, system_prompt, user_prompt)
|
|
46
|
-
request
|
|
47
|
-
end
|
|
48
|
-
|
|
49
|
-
def build_request_body(model, system_prompt, user_prompt)
|
|
50
|
-
{
|
|
51
|
-
model: model,
|
|
52
|
-
messages: [
|
|
53
|
-
{ role: "system", content: system_prompt },
|
|
54
|
-
{ role: "user", content: user_prompt }
|
|
55
|
-
],
|
|
56
|
-
temperature: 0.3
|
|
57
|
-
}.to_json
|
|
58
|
-
end
|
|
59
|
-
|
|
60
|
-
def parse_response(response)
|
|
61
|
-
unless response.is_a?(Net::HTTPSuccess)
|
|
62
|
-
raise AdapterError, "OpenAI API error: #{response.code} - #{response.body}"
|
|
63
|
-
end
|
|
64
|
-
|
|
65
|
-
parsed = JSON.parse(response.body)
|
|
66
|
-
parsed.dig("choices", 0, "message", "content")
|
|
67
|
-
end
|
|
68
|
-
end
|
|
69
|
-
end
|
|
70
|
-
end
|