dev-double 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- dev_double-0.1.0/.github/workflows/publish.yml +59 -0
- dev_double-0.1.0/.gitignore +30 -0
- dev_double-0.1.0/LICENSE +21 -0
- dev_double-0.1.0/PKG-INFO +354 -0
- dev_double-0.1.0/README.md +325 -0
- dev_double-0.1.0/evals/cases.py +353 -0
- dev_double-0.1.0/evals/run.py +256 -0
- dev_double-0.1.0/examples/tour.py +80 -0
- dev_double-0.1.0/pyproject.toml +53 -0
- dev_double-0.1.0/src/dev_double/__init__.py +3 -0
- dev_double-0.1.0/src/dev_double/apps/__init__.py +1 -0
- dev_double-0.1.0/src/dev_double/apps/cli/__init__.py +5 -0
- dev_double-0.1.0/src/dev_double/apps/cli/main.py +122 -0
- dev_double-0.1.0/src/dev_double/apps/client/__init__.py +5 -0
- dev_double-0.1.0/src/dev_double/apps/client/http_client.py +102 -0
- dev_double-0.1.0/src/dev_double/apps/composition.py +71 -0
- dev_double-0.1.0/src/dev_double/apps/config.py +36 -0
- dev_double-0.1.0/src/dev_double/apps/server/__init__.py +5 -0
- dev_double-0.1.0/src/dev_double/apps/server/app.py +122 -0
- dev_double-0.1.0/src/dev_double/apps/server/recorder.py +62 -0
- dev_double-0.1.0/src/dev_double/apps/server/transport/__init__.py +39 -0
- dev_double-0.1.0/src/dev_double/apps/server/transport/choice_use_cases.py +121 -0
- dev_double-0.1.0/src/dev_double/apps/server/transport/common.py +37 -0
- dev_double-0.1.0/src/dev_double/apps/server/transport/decide.py +131 -0
- dev_double-0.1.0/src/dev_double/apps/server/transport/extract.py +93 -0
- dev_double-0.1.0/src/dev_double/apps/server/transport/guard_judge.py +97 -0
- dev_double-0.1.0/src/dev_double/apps/server/transport/rerank.py +55 -0
- dev_double-0.1.0/src/dev_double/client.py +5 -0
- dev_double-0.1.0/src/dev_double/core/__init__.py +1 -0
- dev_double-0.1.0/src/dev_double/core/decision/__init__.py +62 -0
- dev_double-0.1.0/src/dev_double/core/decision/answer_shaping.py +48 -0
- dev_double-0.1.0/src/dev_double/core/decision/confidence.py +19 -0
- dev_double-0.1.0/src/dev_double/core/decision/decider_basic_impl.py +65 -0
- dev_double-0.1.0/src/dev_double/core/decision/decision_service_basic_impl.py +141 -0
- dev_double-0.1.0/src/dev_double/core/decision/defaults.py +40 -0
- dev_double-0.1.0/src/dev_double/core/decision/errors.py +7 -0
- dev_double-0.1.0/src/dev_double/core/decision/extract_questions.py +40 -0
- dev_double-0.1.0/src/dev_double/core/decision/field_extraction.py +53 -0
- dev_double-0.1.0/src/dev_double/core/decision/i_clock.py +11 -0
- dev_double-0.1.0/src/dev_double/core/decision/i_decider.py +36 -0
- dev_double-0.1.0/src/dev_double/core/decision/i_decision_service.py +40 -0
- dev_double-0.1.0/src/dev_double/core/decision/i_engine.py +27 -0
- dev_double-0.1.0/src/dev_double/core/decision/i_extractor.py +19 -0
- dev_double-0.1.0/src/dev_double/core/decision/i_generator.py +19 -0
- dev_double-0.1.0/src/dev_double/core/decision/i_id_provider.py +9 -0
- dev_double-0.1.0/src/dev_double/core/decision/i_record_reader.py +25 -0
- dev_double-0.1.0/src/dev_double/core/decision/i_reranker.py +18 -0
- dev_double-0.1.0/src/dev_double/core/decision/label_scoring.py +89 -0
- dev_double-0.1.0/src/dev_double/core/decision/prompts.py +130 -0
- dev_double-0.1.0/src/dev_double/core/decision/record_parsing.py +153 -0
- dev_double-0.1.0/src/dev_double/core/decision/t_answer.py +32 -0
- dev_double-0.1.0/src/dev_double/core/decision/t_classify.py +42 -0
- dev_double-0.1.0/src/dev_double/core/decision/t_decide.py +26 -0
- dev_double-0.1.0/src/dev_double/core/decision/t_extract.py +89 -0
- dev_double-0.1.0/src/dev_double/core/decision/t_gate.py +36 -0
- dev_double-0.1.0/src/dev_double/core/decision/t_generate.py +37 -0
- dev_double-0.1.0/src/dev_double/core/decision/t_guard.py +38 -0
- dev_double-0.1.0/src/dev_double/core/decision/t_input.py +8 -0
- dev_double-0.1.0/src/dev_double/core/decision/t_judge.py +35 -0
- dev_double-0.1.0/src/dev_double/core/decision/t_label_query.py +36 -0
- dev_double-0.1.0/src/dev_double/core/decision/t_meta.py +17 -0
- dev_double-0.1.0/src/dev_double/core/decision/t_question.py +70 -0
- dev_double-0.1.0/src/dev_double/core/decision/t_rerank.py +36 -0
- dev_double-0.1.0/src/dev_double/core/decision/t_route.py +28 -0
- dev_double-0.1.0/src/dev_double/core/decision/t_usage.py +15 -0
- dev_double-0.1.0/src/dev_double/core/decision/tracker.py +27 -0
- dev_double-0.1.0/src/dev_double/core/decision/use_case_questions.py +80 -0
- dev_double-0.1.0/src/dev_double/core/decision/value_parsing.py +109 -0
- dev_double-0.1.0/src/dev_double/providers/__init__.py +1 -0
- dev_double-0.1.0/src/dev_double/providers/mock/__init__.py +1 -0
- dev_double-0.1.0/src/dev_double/providers/mock/decision/__init__.py +5 -0
- dev_double-0.1.0/src/dev_double/providers/mock/decision/clock_mock_impl.py +19 -0
- dev_double-0.1.0/src/dev_double/providers/mock/decision/engine_mock_impl.py +78 -0
- dev_double-0.1.0/src/dev_double/providers/mock/decision/id_provider_mock_impl.py +15 -0
- dev_double-0.1.0/src/dev_double/providers/needle/__init__.py +1 -0
- dev_double-0.1.0/src/dev_double/providers/needle/decision/__init__.py +4 -0
- dev_double-0.1.0/src/dev_double/providers/needle/decision/decider_needle_impl.py +117 -0
- dev_double-0.1.0/src/dev_double/providers/needle/decision/record_tool.py +75 -0
- dev_double-0.1.0/src/dev_double/providers/openai/__init__.py +1 -0
- dev_double-0.1.0/src/dev_double/providers/openai/decision/__init__.py +3 -0
- dev_double-0.1.0/src/dev_double/providers/openai/decision/engine_openai_impl.py +154 -0
- dev_double-0.1.0/src/dev_double/providers/std/__init__.py +1 -0
- dev_double-0.1.0/src/dev_double/providers/std/decision/__init__.py +4 -0
- dev_double-0.1.0/src/dev_double/providers/std/decision/clock_std_impl.py +12 -0
- dev_double-0.1.0/src/dev_double/providers/std/decision/id_provider_std_impl.py +12 -0
- dev_double-0.1.0/src/dev_double/providers/systemone/__init__.py +1 -0
- dev_double-0.1.0/src/dev_double/providers/systemone/decision/__init__.py +3 -0
- dev_double-0.1.0/src/dev_double/providers/systemone/decision/decider_system_one_impl.py +100 -0
- dev_double-0.1.0/tests/_shared/__init__.py +0 -0
- dev_double-0.1.0/tests/_shared/resources.py +37 -0
- dev_double-0.1.0/tests/_shared/timing.py +23 -0
- dev_double-0.1.0/tests/apps/cli/test_cli.py +23 -0
- dev_double-0.1.0/tests/apps/client/test_client.py +26 -0
- dev_double-0.1.0/tests/apps/server/test_extract_e2e.py +64 -0
- dev_double-0.1.0/tests/apps/server/test_server_e2e.py +195 -0
- dev_double-0.1.0/tests/apps/test_composition.py +49 -0
- dev_double-0.1.0/tests/conftest.py +8 -0
- dev_double-0.1.0/tests/core/decision/test_confidence_and_shaping.py +40 -0
- dev_double-0.1.0/tests/core/decision/test_decider_basic_impl.py +76 -0
- dev_double-0.1.0/tests/core/decision/test_decider_record_reader.py +121 -0
- dev_double-0.1.0/tests/core/decision/test_decision_service.py +122 -0
- dev_double-0.1.0/tests/core/decision/test_extract_invariants.py +71 -0
- dev_double-0.1.0/tests/core/decision/test_extract_service.py +118 -0
- dev_double-0.1.0/tests/core/decision/test_label_scoring.py +78 -0
- dev_double-0.1.0/tests/core/decision/test_question_invariants.py +70 -0
- dev_double-0.1.0/tests/core/decision/test_record_parsing.py +155 -0
- dev_double-0.1.0/tests/core/decision/test_value_parsing.py +79 -0
- dev_double-0.1.0/tests/core/test_dependency_rule.py +87 -0
- dev_double-0.1.0/tests/providers/mock/decision/test_mock_providers.py +48 -0
- dev_double-0.1.0/tests/providers/needle/decision/test_decider_needle_impl.py +175 -0
- dev_double-0.1.0/tests/providers/needle/decision/test_record_tool.py +21 -0
- dev_double-0.1.0/tests/providers/openai/decision/test_engine_openai_generate.py +122 -0
- dev_double-0.1.0/tests/providers/openai/decision/test_engine_openai_impl.py +144 -0
- dev_double-0.1.0/tests/providers/std/decision/test_std_providers.py +17 -0
- dev_double-0.1.0/tests/providers/systemone/decision/test_decider_system_one_impl.py +117 -0
- dev_double-0.1.0/uv.lock +687 -0
|
@@ -0,0 +1,59 @@
|
|
|
1
|
+
# Publishes to PyPI when a GitHub release is published (tag vX.Y.Z).
|
|
2
|
+
# Uses PyPI trusted publishing: no API token is stored in GitHub.
|
|
3
|
+
# One-time setup on pypi.org: add a trusted publisher for this repository,
|
|
4
|
+
# workflow "publish.yml", environment "pypi".
|
|
5
|
+
|
|
6
|
+
name: Publish to PyPI
|
|
7
|
+
|
|
8
|
+
on:
|
|
9
|
+
release:
|
|
10
|
+
types: [published]
|
|
11
|
+
workflow_dispatch:
|
|
12
|
+
|
|
13
|
+
jobs:
|
|
14
|
+
test:
|
|
15
|
+
runs-on: ubuntu-latest
|
|
16
|
+
strategy:
|
|
17
|
+
matrix:
|
|
18
|
+
python-version: ["3.10", "3.13"]
|
|
19
|
+
steps:
|
|
20
|
+
- uses: actions/checkout@v4
|
|
21
|
+
- uses: astral-sh/setup-uv@v6
|
|
22
|
+
with:
|
|
23
|
+
python-version: ${{ matrix.python-version }}
|
|
24
|
+
# Tests that need a model server skip themselves when none is running.
|
|
25
|
+
- run: uv run --frozen pytest -q
|
|
26
|
+
|
|
27
|
+
build:
|
|
28
|
+
needs: test
|
|
29
|
+
runs-on: ubuntu-latest
|
|
30
|
+
steps:
|
|
31
|
+
- uses: actions/checkout@v4
|
|
32
|
+
- uses: astral-sh/setup-uv@v6
|
|
33
|
+
- name: Check the release tag matches the package version
|
|
34
|
+
if: github.event_name == 'release'
|
|
35
|
+
run: |
|
|
36
|
+
version=$(uv version --short)
|
|
37
|
+
test "${GITHUB_REF_NAME}" = "v${version}" || {
|
|
38
|
+
echo "Tag ${GITHUB_REF_NAME} does not match pyproject version ${version}"; exit 1; }
|
|
39
|
+
- run: uv build
|
|
40
|
+
- uses: actions/upload-artifact@v4
|
|
41
|
+
with:
|
|
42
|
+
name: dist
|
|
43
|
+
path: dist/
|
|
44
|
+
|
|
45
|
+
publish:
|
|
46
|
+
needs: build
|
|
47
|
+
if: github.event_name == 'release'
|
|
48
|
+
runs-on: ubuntu-latest
|
|
49
|
+
environment:
|
|
50
|
+
name: pypi
|
|
51
|
+
url: https://pypi.org/p/dev-double
|
|
52
|
+
permissions:
|
|
53
|
+
id-token: write
|
|
54
|
+
steps:
|
|
55
|
+
- uses: actions/download-artifact@v4
|
|
56
|
+
with:
|
|
57
|
+
name: dist
|
|
58
|
+
path: dist/
|
|
59
|
+
- uses: pypa/gh-action-pypi-publish@release/v1
|
|
@@ -0,0 +1,30 @@
|
|
|
1
|
+
# Python
|
|
2
|
+
__pycache__/
|
|
3
|
+
*.py[cod]
|
|
4
|
+
.venv/
|
|
5
|
+
dist/
|
|
6
|
+
build/
|
|
7
|
+
*.egg-info/
|
|
8
|
+
|
|
9
|
+
# Tool caches and coverage
|
|
10
|
+
.pytest_cache/
|
|
11
|
+
.ruff_cache/
|
|
12
|
+
.mypy_cache/
|
|
13
|
+
.coverage
|
|
14
|
+
htmlcov/
|
|
15
|
+
|
|
16
|
+
# Secrets: API keys for hosted engines go in .env, never in the repo
|
|
17
|
+
.env
|
|
18
|
+
.env.*
|
|
19
|
+
!.env.example
|
|
20
|
+
|
|
21
|
+
# Local data: recorded traffic (--record) and eval reports (--json)
|
|
22
|
+
*.jsonl
|
|
23
|
+
evals/*.json
|
|
24
|
+
|
|
25
|
+
# Editors, OS, and local agent config
|
|
26
|
+
.claude
|
|
27
|
+
.idea/
|
|
28
|
+
.vscode/
|
|
29
|
+
.DS_Store
|
|
30
|
+
Thumbs.db
|
dev_double-0.1.0/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Marijus Masteika
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,354 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: dev-double
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: A local stand-in for System 1 decision models: build routing, guardrails, tool gating, triage, evals, reranking and extraction now, swap the engine later.
|
|
5
|
+
Project-URL: Homepage, https://github.com/masteris777/dev-double
|
|
6
|
+
Project-URL: Repository, https://github.com/masteris777/dev-double
|
|
7
|
+
Project-URL: Issues, https://github.com/masteris777/dev-double/issues
|
|
8
|
+
Author: Marijus Masteika
|
|
9
|
+
License-Expression: MIT
|
|
10
|
+
License-File: LICENSE
|
|
11
|
+
Keywords: classification,decision-model,extraction,guardrails,llm,local,mock,ollama,rerank,system-1
|
|
12
|
+
Classifier: Development Status :: 3 - Alpha
|
|
13
|
+
Classifier: Framework :: FastAPI
|
|
14
|
+
Classifier: Intended Audience :: Developers
|
|
15
|
+
Classifier: Programming Language :: Python :: 3
|
|
16
|
+
Classifier: Programming Language :: Python :: 3.10
|
|
17
|
+
Classifier: Programming Language :: Python :: 3.11
|
|
18
|
+
Classifier: Programming Language :: Python :: 3.12
|
|
19
|
+
Classifier: Programming Language :: Python :: 3.13
|
|
20
|
+
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
|
|
21
|
+
Requires-Python: >=3.10
|
|
22
|
+
Requires-Dist: fastapi>=0.110
|
|
23
|
+
Requires-Dist: httpx>=0.27
|
|
24
|
+
Requires-Dist: pydantic>=2.6
|
|
25
|
+
Requires-Dist: uvicorn>=0.29
|
|
26
|
+
Provides-Extra: needle
|
|
27
|
+
Requires-Dist: cactus-needle; extra == 'needle'
|
|
28
|
+
Description-Content-Type: text/markdown
|
|
29
|
+
|
|
30
|
+
# dev-double
|
|
31
|
+
|
|
32
|
+
A local stand-in for System 1 decision models, so you can build your harness before the real model is approved.
|
|
33
|
+
|
|
34
|
+
## Why this exists
|
|
35
|
+
|
|
36
|
+
System 1 models such as Jev, Kev, and Laya took everyone by storm. They don't write text; they return a typed decision with a calibrated probability in milliseconds. Teams use them for model routing, guardrails, tool-call gating, inbox triage, reranking, LLM evals, bulk labeling, real-time control, and confidence gates.
|
|
37
|
+
|
|
38
|
+
Many companies can't use them yet. A new model has to pass security review and whitelisting first, and some teams need it to run on-prem so no context leaves the network. That takes months, and in the meantime development stops: you can't build a harness around an API you're not allowed to call.
|
|
39
|
+
|
|
40
|
+
dev-double fills that gap. It serves decision endpoints with the same shape of output (a typed answer, a probability for every option, and a confidence score), and a model you're already allowed to use does the work underneath. Your team writes the routing, guardrails, gates, and thresholds now. When the real model is approved, you swap the engine and keep everything you built.
|
|
41
|
+
|
|
42
|
+
Kev and Laya are open-weight System 1 models you can host yourself. If your company can approve and run one of them, use it directly. dev-double is for when you can't yet: it runs on a model you're already allowed to use.
|
|
43
|
+
|
|
44
|
+
Like a test double in your unit tests, it stands in for the real dependency. It's slower and less accurate than a purpose-built decision model, but it has the same shape, so your code doesn't change when the real one arrives, and the work goes on.
|
|
45
|
+
|
|
46
|
+
**Any model behind an OpenAI-compatible API works as the engine.** We develop and test with Qwen 2.5 on a local GPU through Ollama. vLLM, llama.cpp, and LM Studio work the same way, and so can a hosted small model, such as one from OpenAI or Anthropic, if your company has already approved it (see [Choosing a model](#choosing-a-model)).
|
|
47
|
+
|
|
48
|
+
## Quick start
|
|
49
|
+
|
|
50
|
+
```bash
|
|
51
|
+
# 1. A local model server (Ollama >= 0.12.11 returns the token probabilities we need)
|
|
52
|
+
ollama pull qwen2.5:7b # ~4.7 GB; qwen2.5:3b (~1.9 GB) if you're short on memory
|
|
53
|
+
|
|
54
|
+
# 2. Install and check the model works
|
|
55
|
+
pip install dev-double
|
|
56
|
+
dev-double doctor
|
|
57
|
+
|
|
58
|
+
# 3. Serve
|
|
59
|
+
dev-double serve # http://127.0.0.1:8787, interactive docs at /docs
|
|
60
|
+
```
|
|
61
|
+
|
|
62
|
+
```bash
|
|
63
|
+
curl -s localhost:8787/v1/route -H 'content-type: application/json' \
|
|
64
|
+
-d '{"input": "Prove that there are infinitely many primes, then formalize it in Lean."}'
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
```json
|
|
68
|
+
{
|
|
69
|
+
"route": "large",
|
|
70
|
+
"probabilities": {"small": 0.0121, "medium": 0.1603, "large": 0.8276},
|
|
71
|
+
"confidence": 0.7414,
|
|
72
|
+
"meta": {"engine": "openai", "model": "qwen2.5:7b", "latency_ms": 212, "usage": {"input_tokens": 187, "output_tokens": 1}, "warnings": []}
|
|
73
|
+
}
|
|
74
|
+
```
|
|
75
|
+
|
|
76
|
+
(Numbers are illustrative.)
|
|
77
|
+
|
|
78
|
+
No model server yet? `dev-double serve --engine mock` answers from word overlap. It's instant and deterministic; use it for CI, never for quality.
|
|
79
|
+
|
|
80
|
+
## Endpoints
|
|
81
|
+
|
|
82
|
+
| Endpoint | Use it for | Returns |
|
|
83
|
+
|---|---|---|
|
|
84
|
+
| `POST /v1/decide` | Anything: ask several typed questions about one input | per question: `binary`, `choice`, or `scale` answer |
|
|
85
|
+
| `POST /v1/route` | Model routing: small, medium, or large model? Or your own routes | `route`, `probabilities`, `confidence` |
|
|
86
|
+
| `POST /v1/guard` | Screen input for prompt injection, abuse, off-topic use | `allowed`, `flagged`, per-policy `probability` |
|
|
87
|
+
| `POST /v1/gate` | Should an agent's tool call be allowed, confirmed, or denied? | `decision`, `probabilities`, `confidence` |
|
|
88
|
+
| `POST /v1/classify` | Inbox triage, intent detection, bulk labeling (`inputs: [...]`) | per item: `label`, `probabilities`, `confidence` |
|
|
89
|
+
| `POST /v1/judge` | LLM evals: score an output against a rubric | `score`, `normalized`, `level`, `probabilities` |
|
|
90
|
+
| `POST /v1/rerank` | Order documents by relevance (Cohere-style format, also `/v2/rerank`) | `results: [{index, relevance_score}]` |
|
|
91
|
+
| `POST /v1/extract` | Pull typed fields out of text: invoices, tickets, bookings, tool-call arguments | `values`, per-field `value` and `confidence` |
|
|
92
|
+
|
|
93
|
+
Every response carries `meta`: engine, model, latency, token usage, and `warnings`, which tell you when an answer is less trustworthy than usual.
|
|
94
|
+
|
|
95
|
+
### `/v1/decide`: the core
|
|
96
|
+
|
|
97
|
+
Three question types cover every use case above:
|
|
98
|
+
|
|
99
|
+
```json
|
|
100
|
+
{
|
|
101
|
+
"input": {"ticket": "My flight was cancelled and nobody is answering the phone!!"},
|
|
102
|
+
"questions": {
|
|
103
|
+
"wants_refund": {"type": "binary", "question": "Does the customer ask for their money back?"},
|
|
104
|
+
"intent": {
|
|
105
|
+
"type": "choice",
|
|
106
|
+
"question": "What does the customer want?",
|
|
107
|
+
"options": {
|
|
108
|
+
"refund": "Money returned.",
|
|
109
|
+
"rebooking": "A replacement flight.",
|
|
110
|
+
"information": "Only information."
|
|
111
|
+
}
|
|
112
|
+
},
|
|
113
|
+
"frustration": {
|
|
114
|
+
"type": "scale",
|
|
115
|
+
"question": "How frustrated is the customer?",
|
|
116
|
+
"levels": ["Calm.", "Concerned but civil.", "Very angry."]
|
|
117
|
+
}
|
|
118
|
+
}
|
|
119
|
+
}
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
```json
|
|
123
|
+
{
|
|
124
|
+
"answers": {
|
|
125
|
+
"wants_refund": {"type": "binary", "value": false, "probability": 0.31, "confidence": 0.38},
|
|
126
|
+
"intent": {"type": "choice", "value": "rebooking", "probabilities": {"refund": 0.22, "rebooking": 0.71, "information": 0.07}, "confidence": 0.565},
|
|
127
|
+
"frustration": {"type": "scale", "value": 1.62, "level": 2, "probabilities": {"0": 0.04, "1": 0.3, "2": 0.66}, "legend": {"0": "Calm.", "1": "Concerned but civil.", "2": "Very angry."}, "confidence": 0.49}
|
|
128
|
+
},
|
|
129
|
+
"meta": {"...": "..."}
|
|
130
|
+
}
|
|
131
|
+
```
|
|
132
|
+
|
|
133
|
+
- **binary**: `probability` that the answer is yes. Optional `yes` and `no` fields say what each means.
|
|
134
|
+
- **choice**: up to 20 unordered options. `value` is the most likely key.
|
|
135
|
+
- **scale**: 2 to 10 ordered levels, lowest first. `value` is the probability-weighted level (it can fall between levels); `level` is the single most likely one.
|
|
136
|
+
- **confidence**: how concentrated the distribution is, 1 when all probability is on one answer, 0 when spread evenly.
|
|
137
|
+
|
|
138
|
+
Questions in one request run in parallel.
|
|
139
|
+
|
|
140
|
+
### `/v1/extract`: typed fields
|
|
141
|
+
|
|
142
|
+
Fields are `string`, `number`, `integer`, `boolean`, or `enum` (with `options`: a list, or option -> description). Up to 30 fields:
|
|
143
|
+
|
|
144
|
+
```json
|
|
145
|
+
{
|
|
146
|
+
"input": "Invoice #2291 from Acme Corp. Amount due: 1.200,00 EUR by 1 October 2026. Status: unpaid.",
|
|
147
|
+
"fields": {
|
|
148
|
+
"vendor": {"type": "string", "description": "Company that issued the invoice."},
|
|
149
|
+
"total": {"type": "number", "description": "Amount due, without currency symbol."},
|
|
150
|
+
"due_date": {"type": "string", "description": "Due date as YYYY-MM-DD."},
|
|
151
|
+
"po_number": {"type": "string", "description": "Purchase order number."},
|
|
152
|
+
"currency": {"type": "enum", "options": ["EUR", "USD", "GBP"]},
|
|
153
|
+
"paid": {"type": "boolean", "description": "Whether the invoice is already paid."}
|
|
154
|
+
}
|
|
155
|
+
}
|
|
156
|
+
```
|
|
157
|
+
|
|
158
|
+
```json
|
|
159
|
+
{
|
|
160
|
+
"values": {"vendor": "Acme Corp", "total": 1200.0, "due_date": "2026-10-01", "po_number": null, "currency": "EUR", "paid": false},
|
|
161
|
+
"fields": {
|
|
162
|
+
"vendor": {"value": "Acme Corp", "confidence": 0.99},
|
|
163
|
+
"total": {"value": 1200.0, "confidence": 0.99},
|
|
164
|
+
"due_date": {"value": "2026-10-01", "confidence": 0.98},
|
|
165
|
+
"po_number": {"value": null, "confidence": 0.97},
|
|
166
|
+
"currency": {"value": "EUR", "confidence": 0.95, "probabilities": {"EUR": 0.97, "USD": 0.02, "GBP": 0.01}},
|
|
167
|
+
"paid": {"value": false, "confidence": 0.8, "probability": 0.1}
|
|
168
|
+
},
|
|
169
|
+
"meta": {"...": "..."}
|
|
170
|
+
}
|
|
171
|
+
```
|
|
172
|
+
|
|
173
|
+
`value` is `null` when the input doesn't contain the field. Enum fields add `probabilities`; boolean fields add `probability` (of true). Numbers are read tolerantly (`$1,200.50`, `1.200,50 EUR`). Fields run in parallel.
|
|
174
|
+
|
|
175
|
+
### Confidence gates
|
|
176
|
+
|
|
177
|
+
Decide the thresholds in code, not in a prompt:
|
|
178
|
+
|
|
179
|
+
```python
|
|
180
|
+
from dev_double.client import DevDouble
|
|
181
|
+
|
|
182
|
+
sd = DevDouble()
|
|
183
|
+
g = sd.gate("send_email", {"to": "all-staff@corp.com"}, context="User asked to draft a reply to Bob.")
|
|
184
|
+
if g["decision"] == "allow" and g["confidence"] > 0.9:
|
|
185
|
+
run_tool()
|
|
186
|
+
elif g["decision"] == "deny" and g["confidence"] > 0.9:
|
|
187
|
+
refuse()
|
|
188
|
+
else:
|
|
189
|
+
ask_human()
|
|
190
|
+
```
|
|
191
|
+
|
|
192
|
+
## How it works
|
|
193
|
+
|
|
194
|
+
Each question becomes one prompt that asks the question, shows the input, asks the question again, and ends with "reply with exactly one label": the option name for choices, a digit for scale levels, yes or no for binary. The model generates at most a few tokens. dev-double reads the probabilities the model assigned to each label's tokens (`logprobs`) and renormalizes them. So the probabilities come from the model itself, not from a number the model writes down. That is also why it's much faster than asking an LLM for JSON.
|
|
195
|
+
|
|
196
|
+
Two details that matter with small models:
|
|
197
|
+
|
|
198
|
+
- **Option names, not letters.** Lettered options (A, B, C) make small models over-pick "B". Answering with the option name avoids that. When two names share their first token ("re" in *refund* and *rebooking*), the next token's probabilities split them.
|
|
199
|
+
- **Question before and after the input.** On our test cases this raised accuracy from 6/9 to 8/9 for yes/no questions.
|
|
200
|
+
|
|
201
|
+
For `/v1/extract`, enum and boolean fields are scored like choice and binary questions; all string, number, and integer fields are generated together as one JSON object, and each value's confidence comes from the probabilities of the tokens that spell it. Engines that can't generate text (System 1 servers) leave those fields `null` with a warning.
|
|
202
|
+
|
|
203
|
+
If the model server returns no logprobs, or the model answers with something that isn't a label, you still get an answer, plus a warning in `meta.warnings`.
|
|
204
|
+
|
|
205
|
+
## Fidelity: what differs from a real decision model
|
|
206
|
+
|
|
207
|
+
Build your harness knowing these:
|
|
208
|
+
|
|
209
|
+
- **Latency.** A local 7B model takes roughly 250 ms per question on a good GPU (a rerank of 4 documents about 1 s), several seconds on CPU. Purpose-built decision models aim much lower. Don't design around the slowness: no caching or batching workarounds a real engine won't need. `meta.latency_ms` shows what you're paying.
|
|
210
|
+
- **Calibration.** Probabilities from a small general model are only roughly calibrated. Thresholds you tune now (like `confidence > 0.9`) will need re-tuning on the real engine. Keep a labeled set of examples so you can re-tune in an afternoon.
|
|
211
|
+
- **Accuracy.** Small models miss nuance. See [Measured results](#measured-results) for what works and what doesn't, and run the evals on your own model before relying on a use case.
|
|
212
|
+
- **Repeatability.** Answers are deterministic while a model stays loaded, but can shift slightly after the model server reloads it. Borderline cases may flip. Don't write tests that assert exact probabilities from a real model; use `--engine mock` for that.
|
|
213
|
+
- **Limits.** Up to 20 options per choice question (option keys must differ ignoring case and punctuation), 10 scale levels, text input only.
|
|
214
|
+
- **Models.** Use non-thinking instruct models. Reasoning models that write out their thinking first (`<think>`) don't work, because the answer is no longer in the first tokens.
|
|
215
|
+
|
|
216
|
+
## Choosing a model
|
|
217
|
+
|
|
218
|
+
dev-double talks to any server with an OpenAI-compatible `/v1/chat/completions` endpoint. What matters is whether that server returns **logprobs** (the probability of each token), because the probabilities in every answer come from them.
|
|
219
|
+
|
|
220
|
+
| Engine | Logprobs | Notes |
|
|
221
|
+
|---|---|---|
|
|
222
|
+
| Ollama ≥ 0.12.11 (default) | yes | Tested with `qwen2.5:7b` and `qwen2.5:3b` on a local GPU. |
|
|
223
|
+
| vLLM, llama.cpp server, LM Studio | yes | `--base-url http://localhost:8000/v1` (vLLM) or `:8080/v1` (llama.cpp). |
|
|
224
|
+
| OpenAI and other hosted APIs | on non-reasoning models | `--base-url https://api.openai.com/v1 --api-key ... --model <small non-reasoning model>`. Not tested yet. |
|
|
225
|
+
| Anthropic | no | Via its OpenAI-compatible endpoint (`--base-url https://api.anthropic.com/v1/`). Not tested yet. With no logprobs, answers are 0 or 1 and `meta.warnings` says so. |
|
|
226
|
+
| Kev, Laya (`--engine systemone`) | native | Open-weight System 1 models, self-hosted. Tested: Kev 0.8B, Laya. |
|
|
227
|
+
| Cactus Needle (`--engine needle`) | one confidence per pick | Tiny on-device tool-calling model. Tested; not a good fit (see below). On Windows, cactus-needle 3.0.5 asks for an engine build that isn't published: download the 3.0.1 engine from Hugging Face and set `NEEDLE3_LIB_PATH`. |
|
|
228
|
+
|
|
229
|
+
Good local choices: Qwen 2.5 (7B recommended, 3B if memory is tight), Llama 3.x, Gemma 3, Phi-4-mini, Mistral.
|
|
230
|
+
|
|
231
|
+
A hosted model sends your input to that provider. Use one only if your company has already approved it for this data; the point of dev-double is to not need the unapproved model.
|
|
232
|
+
|
|
233
|
+
### Which model for which use case
|
|
234
|
+
|
|
235
|
+
Most teams need one use case, not all nine, so pick per use case. `evals/run.py` holds 140 labeled cases, 20 per use case, including deliberately hard ones (rerank distractors that share the query's words, answers that are subtly wrong, extraction fields that aren't in the text and must come back empty). A model counts as **good enough** for a use case at 17/20 (85%) or better.
|
|
236
|
+
|
|
237
|
+
| Use case | Recommended | Also good enough | Not good enough |
|
|
238
|
+
|---|---|---|---|
|
|
239
|
+
| guard (injection, abuse) | **qwen2.5:7b** (18/20) | none | qwen2.5:3b 12, Kev 13, Laya 12 |
|
|
240
|
+
| route | **qwen2.5:7b** (18/20) | none | qwen2.5:3b 16, Kev 14, Laya 9 |
|
|
241
|
+
| gate | **qwen2.5:7b** (17/20, see caution below) | none | Laya 14, qwen2.5:3b 11, Kev 10 |
|
|
242
|
+
| classify (inbox triage) | **Kev 0.8B** (17/20) | none; qwen2.5:7b is just under (16/20) | qwen2.5:3b 15, Laya 12 |
|
|
243
|
+
| judge | **qwen2.5:7b** (18/20) | none | qwen2.5:3b 15, Kev 10, Laya 9 |
|
|
244
|
+
| rerank | **qwen2.5:3b** (20/20) | qwen2.5:7b 20, Laya 20, Kev 19 | Needle 10 |
|
|
245
|
+
| extract | **qwen2.5:3b** (16/20, 92% of fields); close, see below | none | qwen2.5:7b 15 (93% of fields), Needle 4 (69%) |
|
|
246
|
+
|
|
247
|
+
Full results, on an RTX 4070 Laptop GPU (8 GB), one model loaded at a time:
|
|
248
|
+
|
|
249
|
+
| Use case | qwen2.5:7b | qwen2.5:3b | Laya | Kev 0.8B | Needle |
|
|
250
|
+
|---|---|---|---|---|---|
|
|
251
|
+
| guard | **18**/20 | 12 | 12 | 13 | 0 |
|
|
252
|
+
| route | **18** | 16 | 9 | 14 | 8 |
|
|
253
|
+
| gate | **17** | 11 | 14 | 10 | 7 |
|
|
254
|
+
| classify | 16 | 15 | 12 | **17** | 5 |
|
|
255
|
+
| judge | **18** | 15 | 9 | 10 | 2 |
|
|
256
|
+
| rerank | **20** | **20** | **20** | **19** | 10 |
|
|
257
|
+
| extract | 15 | **16** | n/a | n/a | 4 |
|
|
258
|
+
| **total** | **122/140** | 105 | 76/120 | 83/120 | 36 |
|
|
259
|
+
| median latency per call | 260–300 ms | 250–260 ms | 25–50 ms | 30–45 ms | 55–390 ms |
|
|
260
|
+
|
|
261
|
+
Rerank latency is per query of 4 documents: about 1.1 s for the Qwen models, 90–110 ms for Laya and Kev. Guard asks two questions per input, so it takes about twice as long. Extract takes 0.9–1.3 s per record on the Qwen models (one call for all text fields plus one per enum or yes/no field). Extract counts a record as right only if every field is right. Laya and Kev can't write text, so they can't fill text fields and weren't scored on extract.
|
|
262
|
+
|
|
263
|
+
What that means in practice:
|
|
264
|
+
|
|
265
|
+
- **qwen2.5:7b is the default** and is good enough for every use case except inbox triage, where it rates polite but non-urgent requests as urgent.
|
|
266
|
+
- **qwen2.5:3b is enough for rerank.** It's also fine for rough routing, but it mixes up guard policies and wrongly denies harmless tool calls.
|
|
267
|
+
- **Be careful with confident mistakes in gate.** The 7B allowed an 8,400 transfer to a new payee, an email to the whole company, and a production deploy, all with near-certainty; each should have asked a human. Put hard limits (amounts, recipients, destructive tools) in code, not only in a confidence threshold.
|
|
268
|
+
- **Kev and Laya are open-weight System 1 models, and far faster** (10x). Laya ranks documents perfectly and Kev is the best at triage. Elsewhere they fall short here, but three caveats apply:
|
|
269
|
+
- They got the same question wording as the LLMs, which wasn't tuned for them.
|
|
270
|
+
- Only Kev's 0.8B model fits in 8 GB (the 4B needs 32 GB).
|
|
271
|
+
- Kev ran without its fast kernels, which aren't available on Windows.
|
|
272
|
+
|
|
273
|
+
If your company can approve one of them, test it on your own cases with `--engine systemone`.
|
|
274
|
+
- **Extraction is close to good enough.** Both Qwen models get about 92% of fields right, but only 15–16 of 20 records completely right. Typical misses: inventing a date from "next summer", leaving out a meeting title that is in the text. Check low-confidence fields, or send them to a human.
|
|
275
|
+
- **Needle is not a decision model.** It's a tiny on-device model that *writes* tool calls. Asked to classify, it often returns no call at all, and embedding-based rerank only reached 10/20. Even on extraction, its home ground, it got 69% of fields right out of the box. Cactus pitches fine-tuning on your own schema, which we didn't test. Its fit is the step before a decision model: Needle proposes the tool call on the device, and `/v1/gate` decides whether to run it.
|
|
276
|
+
|
|
277
|
+
Reproduce any column (one model at a time; load only one model into GPU memory):
|
|
278
|
+
|
|
279
|
+
```bash
|
|
280
|
+
uv run python evals/run.py --model qwen2.5:7b # Ollama
|
|
281
|
+
uv run python evals/run.py --model laya=http://localhost:8000 # laya-serve
|
|
282
|
+
uv run python evals/run.py --model kev=http://localhost:8009 # python -m kev.serve --run jaredpalmer/kev-0.8b --port 8009
|
|
283
|
+
uv run --extra needle python evals/run.py --model needle # set NEEDLE_TELEMETRY=0
|
|
284
|
+
```
|
|
285
|
+
|
|
286
|
+
Twenty cases per use case is still small; add cases from your own domain before relying on a result.
|
|
287
|
+
|
|
288
|
+
## Switching to the real engine later
|
|
289
|
+
|
|
290
|
+
1. Keep dev-double behind your own small interface, for example `Decisions.route(...)`, `Decisions.gate(...)`.
|
|
291
|
+
2. Record real traffic while you develop: `dev-double serve --record calls.jsonl`. Each line holds the request, response, and latency.
|
|
292
|
+
3. When the real engine is approved, implement the same interface with its SDK, replay `calls.jsonl` through both, and compare decisions and latency before switching.
|
|
293
|
+
|
|
294
|
+
The rest of your harness (thresholds aside) does not change.
|
|
295
|
+
|
|
296
|
+
## Configuration
|
|
297
|
+
|
|
298
|
+
| Flag | Environment variable | Default |
|
|
299
|
+
|---|---|---|
|
|
300
|
+
| `--engine` | `DEV_DOUBLE_ENGINE` | `openai` (any OpenAI-compatible server), `mock` (word overlap, for CI), `systemone` (a self-hosted Kev or Laya server), or `needle` (Cactus Needle on-device; `pip install "dev-double[needle]"`) |
|
|
301
|
+
| `--base-url` | `DEV_DOUBLE_BASE_URL` | `http://localhost:11434/v1` (Ollama) |
|
|
302
|
+
| `--model` | `DEV_DOUBLE_MODEL` | `qwen2.5:7b` |
|
|
303
|
+
| `--api-key` | `DEV_DOUBLE_API_KEY` | none |
|
|
304
|
+
| `--max-concurrency` | `DEV_DOUBLE_MAX_CONCURRENCY` | `4` |
|
|
305
|
+
| `--systemone-url` | `DEV_DOUBLE_SYSTEMONE_URL` | `http://localhost:8000` (serves `POST /v1/systemone`) |
|
|
306
|
+
| | `DEV_DOUBLE_TIMEOUT` | `120` seconds per model call |
|
|
307
|
+
| `--record` | `DEV_DOUBLE_RECORD` | off |
|
|
308
|
+
| `--host`, `--port` | | `127.0.0.1`, `8787` |
|
|
309
|
+
|
|
310
|
+
For vLLM: `--base-url http://localhost:8000/v1 --model Qwen/Qwen2.5-7B-Instruct`. For llama.cpp server: `--base-url http://localhost:8080/v1`.
|
|
311
|
+
|
|
312
|
+
`--engine systemone` and `--engine needle` send each question to a model that answers typed questions natively, with no dev-double prompt, so you can compare dev-double with the real thing behind the same endpoints. Needle has no probability for every option: it returns one confidence for its pick, and the other options share the rest evenly. It reranks with embeddings. Telemetry is off by default (`NEEDLE_TELEMETRY=0`, `DO_NOT_TRACK=1`), and `NEEDLE3_LIB_PATH` points it at a local build.
|
|
313
|
+
|
|
314
|
+
The server has no authentication. Keep it on localhost or behind your own gateway.
|
|
315
|
+
|
|
316
|
+
## Development
|
|
317
|
+
|
|
318
|
+
```bash
|
|
319
|
+
uv sync
|
|
320
|
+
uv run pytest # unit + e2e on the mock engine; integration tests
|
|
321
|
+
# run when Ollama (qwen2.5:7b) or a System 1
|
|
322
|
+
# server on :8000 answers, else skip
|
|
323
|
+
uv run python evals/run.py --model qwen2.5:7b # quality, needs a model server
|
|
324
|
+
uv run python evals/run.py --model qwen2.5:7b --only gate # one use case
|
|
325
|
+
```
|
|
326
|
+
|
|
327
|
+
Run the evals on at least two models before and after any prompt change: a wording that helps one model can hurt another.
|
|
328
|
+
|
|
329
|
+
## Architecture
|
|
330
|
+
|
|
331
|
+
The code has three layers, and imports only go inward:
|
|
332
|
+
|
|
333
|
+
```
|
|
334
|
+
apps ──→ core ←── providers
|
|
335
|
+
```
|
|
336
|
+
|
|
337
|
+
- **`dev_double/core/decision/`**: the domain, with no third-party imports (no httpx, pydantic, time, or os). It holds the interfaces (`IEngine`, `IDecider`, `IClock`, `IIdProvider`, `IDecisionService`, and the optional capabilities `IReranker`, `IGenerator`, `IRecordReader`, `IExtractor`), the domain types (`TChoiceQuestion`, `TRouteRequest`, ...), the prompts, the label scoring, and the use cases (`DecisionServiceBasicImpl`). `DeciderBasicImpl` turns a question into a prompt and asks an `IEngine`.
|
|
338
|
+
- **`dev_double/providers/<name>/decision/`**: one folder per technology. Each implements core interfaces and depends only on core, never on another provider. `openai` (`EngineOpenAIImpl`), `mock` (`EngineMockImpl`, plus a deterministic clock and id provider for tests), `std` (the real clock and UUIDs), `systemone` (`DeciderSystemOneImpl`), and `needle` (`DeciderNeedleImpl`).
|
|
339
|
+
- **`dev_double/apps/`**: `server/` (FastAPI, the pydantic transport schemas, the recorder), `cli/`, `client/` (the `DevDouble` SDK), and `composition.py`, the only place that picks implementations from the settings.
|
|
340
|
+
|
|
341
|
+
To add an engine:
|
|
342
|
+
|
|
343
|
+
- If it scores labels from a prompt, implement `IEngine` in `providers/<name>/decision/engine_<name>_impl.py`. Implement `IGenerator` too if it can generate short text, so `/v1/extract` can fill string and number fields.
|
|
344
|
+
- If it answers typed questions natively, implement `IDecider` in `providers/<name>/decision/decider_<name>_impl.py`, and also `IReranker` if it can rerank without yes/no questions, or `IExtractor` if it extracts whole records.
|
|
345
|
+
|
|
346
|
+
Then add a branch in `apps/composition.py` and the name to `ENGINES` in `apps/config.py`, and put its tests in `tests/providers/<name>/`. `tests/` mirrors `src/`. `tests/core/test_dependency_rule.py` fails if core imports anything external or if one provider imports another.
|
|
347
|
+
|
|
348
|
+
## Credits
|
|
349
|
+
|
|
350
|
+
Built by Marijus Masteika with [Claude](https://claude.com/claude-code) (Anthropic) as a coding partner. Claude co-authored the commits.
|
|
351
|
+
|
|
352
|
+
## License and trademarks
|
|
353
|
+
|
|
354
|
+
MIT. dev-double is an independent project with its own API design. It is not affiliated with or endorsed by TypeSafe AI (Jev), the Kev project, Convai Innovations (Laya), or Cactus Compute (Needle). dev-double does not serve their APIs; `--engine systemone` and `--engine needle` are clients, for comparison. Those names belong to their owners and are mentioned only to describe what dev-double stands in for or is compared with.
|