llmxray 0.4.9 → 0.5.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (60) hide show
  1. package/README.ar.md +302 -301
  2. package/README.fr.md +298 -297
  3. package/README.md +298 -297
  4. package/README.sr.md +298 -297
  5. package/README.zh-CN.md +298 -297
  6. package/dist/assets/{AITrainingPage-E1g8rBQe.js → AITrainingPage-Cb771S8Z.js} +1 -1
  7. package/dist/assets/{AnalyticsPage-B-axxCjX.js → AnalyticsPage-D4bgNoXM.js} +1 -1
  8. package/dist/assets/BenchmarkPage-BTfBJXIj.js +31 -0
  9. package/dist/assets/{ComparisonPage-phs59_G2.js → ComparisonPage-DMcD9JL0.js} +1 -1
  10. package/dist/assets/{CostDashboardPage-DmXvNxoW.js → CostDashboardPage-BjUiVjlg.js} +1 -1
  11. package/dist/assets/DashboardPage-xgeJtl-Q.js +114 -0
  12. package/dist/assets/EmbeddingsPage-DY1JKxCN.js +1 -0
  13. package/dist/assets/{FimPlaygroundPage-DuzUaxaO.js → FimPlaygroundPage-BWSlc3k7.js} +1 -1
  14. package/dist/assets/{GoogleCallbackPage-B_wKPxTY.js → GoogleCallbackPage-C9xQpfMW.js} +1 -1
  15. package/dist/assets/{JsonTreeNode-C3GvpFIv.js → JsonTreeNode-D9YzXU8h.js} +1 -1
  16. package/dist/assets/{ModelCapabilityIcons.vue_vue_type_script_setup_true_lang-tkg1H--1.js → ModelCapabilityIcons.vue_vue_type_script_setup_true_lang-CPvepOP2.js} +1 -1
  17. package/dist/assets/ProtocolObservatoryPage-D7szMZuY.js +2 -0
  18. package/dist/assets/{RAGPage-DGFy-b60.js → RAGPage-C7awGOKp.js} +1 -1
  19. package/dist/assets/SessionPage-8CBF4JLG.js +15 -0
  20. package/dist/assets/{SettingsPage-CFSn9lUW.js → SettingsPage-MEHyheCY.js} +1 -1
  21. package/dist/assets/{StatusBadge.vue_vue_type_script_setup_true_lang-W3vS9xKq.js → StatusBadge.vue_vue_type_script_setup_true_lang-BPTPrdQf.js} +1 -1
  22. package/dist/assets/{StorageGauge.vue_vue_type_style_index_0_lang-NsFsiU2e.js → StorageGauge.vue_vue_type_style_index_0_lang-a4pXCCmZ.js} +1 -1
  23. package/dist/assets/{SystemPage-tBgC5G3-.js → SystemPage-mElnkuG4.js} +2 -2
  24. package/dist/assets/{TabBar.vue_vue_type_script_setup_true_lang-DL7akl1T.js → TabBar.vue_vue_type_script_setup_true_lang-g3X_zm3P.js} +1 -1
  25. package/dist/assets/{TokenStreamDisplay.vue_vue_type_script_setup_true_lang-B2xlTJuS.js → TokenStreamDisplay.vue_vue_type_script_setup_true_lang-DQBSq5z2.js} +1 -1
  26. package/dist/assets/{ToolWorkshopPage-BAQPeIGP.js → ToolWorkshopPage-3Za0iXSV.js} +1 -1
  27. package/dist/assets/{agent-store-CPoOtsp1.js → agent-store-Cboc3nde.js} +1 -1
  28. package/dist/assets/{canvas-ai-db-DkGslnpm.js → canvas-ai-db-Dy65jrST.js} +1 -1
  29. package/dist/assets/{download-ChpH3xF_.js → download-Cipln2ej.js} +1 -1
  30. package/dist/assets/format-CKxSMkvr.js +1 -0
  31. package/dist/assets/{generate-service-CzzxTSzU.js → generate-service-C8to1Kz5.js} +1 -1
  32. package/dist/assets/{google-auth-store-CWJyfJot.js → google-auth-store-8ShUecGK.js} +1 -1
  33. package/dist/assets/index-DAqgYcEq.js +36 -0
  34. package/dist/assets/{index-BkYEc6g3.js → index-DlTtfQCm.js} +1 -1
  35. package/dist/assets/{index-BqaWqS5R.js → index-NGfHHJ7k.js} +1 -1
  36. package/dist/assets/{index-aU4tMz4Z.css → index-UYM0582r.css} +1 -1
  37. package/dist/assets/{info-D5-JTLE0.js → info-BzNBY8Xz.js} +1 -1
  38. package/dist/assets/{metrics-store--HW_Yktg.js → metrics-store-DLxxEHpn.js} +1 -1
  39. package/dist/assets/{papaparse.min-B9Ijd5eE.js → papaparse.min-BMNThItx.js} +1 -1
  40. package/dist/assets/{pencil-CsKrNJEo.js → pencil-c-VIQAV1.js} +1 -1
  41. package/dist/assets/{play-nicSMl5X.js → play-CyV4NPEG.js} +1 -1
  42. package/dist/assets/{rag-store-CEuRiXlx.js → rag-store-7BNJiZff.js} +3 -3
  43. package/dist/assets/{session-store-Dt5jLOgV.js → session-store-DJJV_fRl.js} +1 -1
  44. package/dist/assets/{share-2-BdqfbjHE.js → share-2-DTMUvOzC.js} +1 -1
  45. package/dist/assets/{square-KUkPfYFT.js → square-CVjmHbTf.js} +1 -1
  46. package/dist/assets/{storage-store-ColSqJM3.js → storage-store-CAp_Ass2.js} +1 -1
  47. package/dist/assets/stream-handler-DVtwM8Et.js +4 -0
  48. package/dist/assets/{tool-workshop-store-CpWKGa76.js → tool-workshop-store-C_snrdcI.js} +1 -1
  49. package/dist/assets/{toolcall-store-C3ghhPw3.js → toolcall-store-DVrrMX0Y.js} +1 -1
  50. package/dist/assets/{trash-2-DnAmbFuB.js → trash-2-DfEdg6Up.js} +1 -1
  51. package/dist/index.html +2 -2
  52. package/package.json +1 -1
  53. package/dist/assets/BenchmarkPage-CMa0Q0wR.js +0 -31
  54. package/dist/assets/DashboardPage-CFydBLae.js +0 -114
  55. package/dist/assets/EmbeddingsPage-Dr4zAK3P.js +0 -1
  56. package/dist/assets/ProtocolObservatoryPage-C5NU7kv4.js +0 -2
  57. package/dist/assets/SessionPage-C4oWNW2u.js +0 -15
  58. package/dist/assets/format-BqVZ9rwM.js +0 -1
  59. package/dist/assets/index-CE4LM5vl.js +0 -36
  60. package/dist/assets/stream-handler-DHKXR5XU.js +0 -4
package/README.md CHANGED
@@ -1,297 +1,298 @@
1
- <p align="center">
2
- <img src="https://raw.githubusercontent.com/LogneBudo/llmxray/master/public/favicon.svg" alt="LLMxRay" width="80" />
3
- </p>
4
-
5
- <h1 align="center">LLMxRay</h1>
6
- <p align="center"><strong>See what your AI is actually doing.</strong></p>
7
- <p align="center">
8
- Real-time token streaming, quality analysis, performance profiling, and cost tracking<br/>
9
- for local LLMs. No cloud. No API keys. No cost.
10
- </p>
11
-
12
- <p align="center">
13
- <a href="https://www.npmjs.com/package/llmxray"><img src="https://img.shields.io/npm/v/llmxray?color=cb3837&logo=npm&logoColor=white" alt="npm" /></a>
14
- <a href="https://hub.docker.com/r/djovaneli/llmxray"><img src="https://img.shields.io/docker/pulls/djovaneli/llmxray?color=2496ED&logo=docker&logoColor=white" alt="Docker" /></a>
15
- <img src="https://img.shields.io/badge/license-Apache%202.0-blue" alt="License" />
16
- <img src="https://img.shields.io/badge/ollama-local-000?logo=ollama&logoColor=white" alt="Ollama" />
17
- </p>
18
-
19
- <p align="center">
20
- 🌐 <strong>English</strong> &bull;
21
- <a href="README.fr.md">Français</a> &bull;
22
- <a href="README.zh-CN.md">中文</a> &bull;
23
- <a href="README.ar.md">العربية</a> &bull;
24
- <a href="README.sr.md">Srpski</a>
25
- </p>
26
-
27
- <p align="center">
28
- <a href="#quick-start">Quick Start</a> &bull;
29
- <a href="#features">Features</a> &bull;
30
- <a href="#screenshots">Screenshots</a> &bull;
31
- <a href="#who-is-this-for">Who Is This For</a> &bull;
32
- <a href="CHANGELOG.md">Changelog</a>
33
- </p>
34
-
35
- <p align="center">
36
- <img src="docs/public/screenshots/demo.gif" alt="LLMxRay demo — real-time token streaming with confidence coloring" width="800" />
37
- </p>
38
-
39
- ---
40
-
41
- ## Quick Start
42
-
43
- **One command. 30 seconds.**
44
-
45
- ```bash
46
- npx llmxray
47
- ```
48
-
49
- Or with Docker:
50
-
51
- ```bash
52
- docker run -p 5174:5174 djovaneli/llmxray
53
- ```
54
-
55
- Open **http://localhost:5174** and start chatting. That's it.
56
-
57
- > **Prerequisite:** [Ollama](https://ollama.com/download) running locally with at least one model pulled (`ollama pull llama3.2`).
58
-
59
- ---
60
-
61
- ## Why LLMxRay?
62
-
63
- You run a local LLM. You chat with it. But what actually happened?
64
-
65
- - How fast was each token? Which ones was the model confident about?
66
- - Is the response quality degrading over long conversations?
67
- - What would this have cost if you ran it in the cloud?
68
- - Is the model repeating itself? Refusing? Generating gibberish?
69
- - How does temperature 0.3 compare to 0.9 on the *same* prompt?
70
-
71
- **LLMxRay answers all of these, visually, in real time, for free.**
72
-
73
- ---
74
-
75
- ## Features
76
-
77
- ### Real-Time Chat with Token Intelligence
78
- Chat with any Ollama model and watch tokens arrive with **confidence coloring** — each token is tinted based on generation speed. Supports markdown, multi-turn conversations, file attachments, vision models, and slash commands.
79
-
80
- ### Response Quality Gates
81
- Every response is automatically analyzed. Colored badges appear only when something is wrong:
82
- - **Repetition** — excessive repeated phrases (4-gram analysis)
83
- - **Refusal** — "as an AI language model" and 7 other patterns
84
- - **Gibberish** — high non-ASCII ratio
85
- - **Empty** — fewer than 10 words
86
- - **Truncation** — hit the token limit without finishing
87
-
88
- ### Model Comparison Workbench
89
- Up to **4 slots** with independent model, temperature, and system prompt. Features include side-by-side streaming, word-level diff highlighting, metrics comparison, and one-click presets (Temperature Sweep, Deterministic Pair, Language Compare with Token Tax visualization).
90
-
91
- ### Performance Analytics
92
- - **Latency percentiles** (P50/P95/P99) for duration and TTFT
93
- - **Error intelligence** — 7-category classifier with timeline
94
- - **Usage heatmap** — 7x24 grid of your active hours
95
- - **Settings impact** — temperature vs tokens/sec scatter plots
96
- - **Cold vs warm start** tracking with model load history
97
-
98
- ### Cost Dashboard
99
- Token usage per model/day with estimated cloud-equivalent pricing. See what you're *saving* by running locally.
100
-
101
- ### Surgical Benchmark
102
- Test model knowledge with multi-choice question suites. Uses real logprobs via OpenAI-compatible endpoint for accurate confidence measurement. Build custom suites visually or let AI generate them from a topic.
103
-
104
- ### Embeddings Lab & RAG Pipeline
105
- Embed text, visualize vectors, measure cosine similarity. Build a local knowledge base from PDFs, DOCX, and CSV — chunked, embedded, and searchable. All stored in IndexedDB. Zero cost.
106
-
107
- ### Tool Workshop (Visual Canvas)
108
- Drag-and-drop node canvas for building tool definitions. Bidirectional code sync (edit nodes or TypeScript — both update). Probe APIs, auto-generate schemas, test with live execution.
109
-
110
- ### Fill-in-the-Middle Playground *(new in v0.4.7)*
111
- Code completion for Qwen-Coder, CodeLlama, Codestral, DeepSeek-Coder, and StarCoder. Two textareas (prefix / suffix), the model fills the gap. Uses Ollama's `suffix` field on `/api/generate`. Stitched preview shows the result as it would appear in your editor.
112
-
113
- ### Protocol Observatory *(new in v0.4.7)*
114
- Fire the same prompt through Ollama's three serving protocols — **native** `/api/chat`, **OpenAI-compat** `/v1/chat/completions`, and **Anthropic-compat** `/v1/messages` — in parallel against your local model. Side-by-side streaming, per-protocol metrics, and an envelope-diff tab that shows how each protocol frames finish reasons, token counts, and error envelopes. No cloud, no API keys — all three endpoints are local on `localhost:11434`.
115
-
116
- ### AI Training Pipeline
117
- Curate training data from your conversations. Tag, review, and export as JSONL for fine-tuning.
118
-
119
- ### Local AI History Database
120
- Every experiment (benchmarks, comparisons, chats, training pairs) is automatically archived in a queryable IndexedDB database with filters, trends, exports, and retention policies.
121
-
122
- ### Multilingual
123
- Full translations in English, French, Serbian (Latin + Cyrillic), Chinese, and Arabic. RTL layout support. Community scaffolds for Hebrew and Japanese.
124
-
125
- ---
126
-
127
- ## Ollama Compatibility
128
-
129
- Tested and verified against **Ollama 0.24.0** (the current latest stable as of May 2026). LLMxRay uses these Ollama endpoints:
130
-
131
- | Endpoint | Used for |
132
- |---|---|
133
- | `/api/chat` | Streaming chat (NDJSON, with `tools`, `think`, `format` schema) |
134
- | `/api/generate` | Generation + Fill-in-the-Middle via `suffix` |
135
- | `/api/tags`, `/api/show` | Model list + capability detection (`thinking`, `tools`, `vision`) |
136
- | `/api/embed` | Vector embeddings for RAG |
137
- | `/api/pull`, `/api/delete`, `/api/ps`, `/api/version` | Model management + status |
138
- | `/v1/chat/completions` | OpenAI-compat path used by Surgical Benchmark for real logprobs |
139
- | `/v1/messages` | Anthropic-compat path used by Protocol Observatory |
140
-
141
- **Compatible with:** Ollama 0.20 and newer (older versions work for chat/generate but lack `think` and JSON-schema `format`). **Recommended:** Ollama 0.24+ for full feature parity including the Anthropic-compat endpoint (added in 0.23) and the `think: "max"` mode (added in 0.21.3).
142
-
143
- ---
144
-
145
- ## Screenshots
146
-
147
- <table>
148
- <tr>
149
- <td width="50%">
150
-
151
- **Chat with token streaming and confidence**
152
- ![Chat](docs/public/screenshots/chat-diagnostics.png)
153
-
154
- </td>
155
- <td width="50%">
156
-
157
- **Model comparison — side by side**
158
- ![Compare](docs/public/screenshots/compare-sidebyside.png)
159
-
160
- </td>
161
- </tr>
162
- <tr>
163
- <td width="50%">
164
-
165
- **Session deep dive — metrics and timing**
166
- ![Session](docs/public/screenshots/session-details.png)
167
-
168
- </td>
169
- <td width="50%">
170
-
171
- **Benchmark with confidence radar**
172
- ![Benchmark](docs/public/screenshots/benchmark.png)
173
-
174
- </td>
175
- </tr>
176
- <tr>
177
- <td width="50%">
178
-
179
- **Embeddings — cosine similarity**
180
- ![Embeddings](docs/public/screenshots/embed-similarity.png)
181
-
182
- </td>
183
- <td width="50%">
184
-
185
- **System monitor — hardware and Ollama status**
186
- ![System](docs/public/screenshots/my-system.png)
187
-
188
- </td>
189
- </tr>
190
- </table>
191
-
192
- ---
193
-
194
- ## Who Is This For
195
-
196
- | You are... | LLMxRay helps you... |
197
- |---|---|
198
- | **Developer** | Debug prompts, profile latency, compare models, inspect tool calls, track costs |
199
- | **Researcher** | Run controlled experiments with consistent settings across models and temperatures |
200
- | **Student / Educator** | Explore model behavior visually built-in Educators Kit with 9 interactive modules |
201
- | **AI team lead** | Understand quality trends, error patterns, and resource usage across your local fleet |
202
-
203
- ---
204
-
205
- ## Install Options
206
-
207
- ### npx (recommended)
208
- ```bash
209
- npx llmxray
210
- npx llmxray --port 3000
211
- npx llmxray --ollama-url http://192.168.1.50:11434
212
- ```
213
-
214
- ### Docker
215
- ```bash
216
- docker run -p 5174:5174 djovaneli/llmxray
217
- docker run -p 5174:5174 -e OLLAMA_URL=http://host.docker.internal:11434 djovaneli/llmxray
218
- ```
219
-
220
- ### From source
221
- ```bash
222
- git clone https://github.com/LogneBudo/llmxray.git
223
- cd llmxray
224
- npm install
225
- npm run dev # http://localhost:5173
226
- ```
227
-
228
- ---
229
-
230
- ## Tech Stack
231
-
232
- | Layer | Technology |
233
- |---|---|
234
- | Framework | Vue 3.5 + Composition API |
235
- | Language | TypeScript 5.9 (strict) |
236
- | Build | Vite 7.3 |
237
- | Styling | Tailwind CSS 4.2 |
238
- | State | Pinia 3 (store-per-concern) |
239
- | Charts | Chart.js 4, D3.js 7 |
240
- | Canvas | Vue Flow (visual node editor) |
241
- | Code Editor | CodeMirror 6 |
242
- | Storage | IndexedDB (browser-native) |
243
- | LLM Backend | Ollama (local) |
244
-
245
- ---
246
-
247
- ## Architecture
248
-
249
- **Streaming** — Reads Ollama NDJSON via `fetch()` + `ReadableStream`. Tokens update the UI reactively through Pinia stores.
250
-
251
- **Token confidence** — Approximated from inter-token latency (faster = more confident). Clearly labeled as approximation. Benchmarks use real logprobs via OpenAI-compatible endpoint.
252
-
253
- **Store-per-concern** — Each domain has its own Pinia store: tokens, sessions, metrics, reasoning, comparison, embeddings, quality, cost, and more.
254
-
255
- **Hardware detection** — Custom Vite plugin queries the OS directly (PowerShell/proc/sysctl) for accurate hardware specs.
256
-
257
- ---
258
-
259
- ## Development
260
-
261
- | Command | What it does |
262
- |---|---|
263
- | `npm run dev` | Dev server (port 5173) |
264
- | `npm run build` | Type-check + production build |
265
- | `npm run test` | Unit tests (Vitest) |
266
- | `npm run test:e2e` | End-to-end (Playwright) |
267
-
268
- ---
269
-
270
- ## Contributing
271
-
272
- Contributions welcome! See [CONTRIBUTING.md](CONTRIBUTING.md) for setup and guidelines.
273
-
274
- **Community translations especially welcome** — scaffold files ready for Hebrew and Japanese.
275
-
276
- ---
277
-
278
- ## License
279
-
280
- [Apache License 2.0](LICENSE)
281
-
282
- ## Trademark
283
-
284
- **LLMxRay** is a trademark of Ivan Stankovic ([LogneBudo](https://github.com/LogneBudo)). See [TRADEMARK.md](TRADEMARK.md).
285
-
286
- ---
287
-
288
- <p align="center">
289
- <strong>If LLMxRay helps you understand your AI better, consider giving it a star.</strong><br/>
290
- It helps others discover the project.
291
- </p>
292
-
293
- <p align="center">
294
- <a href="https://github.com/LogneBudo/llmxray">
295
- <img src="https://img.shields.io/github/stars/LogneBudo/llmxray?style=social" alt="GitHub stars" />
296
- </a>
297
- </p>
1
+ <p align="center">
2
+ <img src="https://raw.githubusercontent.com/LogneBudo/llmxray/master/public/favicon.svg" alt="LLMxRay" width="80" />
3
+ </p>
4
+
5
+ <h1 align="center">LLMxRay</h1>
6
+ <p align="center"><strong>See what your AI is actually doing.</strong></p>
7
+ <p align="center">
8
+ Real-time token streaming, quality analysis, performance profiling, and cost tracking<br/>
9
+ for local LLMs. No cloud. No API keys. No cost.
10
+ </p>
11
+
12
+ <p align="center">
13
+ <a href="https://www.npmjs.com/package/llmxray"><img src="https://img.shields.io/npm/v/llmxray?color=cb3837&logo=npm&logoColor=white" alt="npm" /></a>
14
+ <a href="https://hub.docker.com/r/djovaneli/llmxray"><img src="https://img.shields.io/docker/pulls/djovaneli/llmxray?color=2496ED&logo=docker&logoColor=white" alt="Docker" /></a>
15
+ <img src="https://img.shields.io/badge/license-Apache%202.0-blue" alt="License" />
16
+ <img src="https://img.shields.io/badge/ollama-local-000?logo=ollama&logoColor=white" alt="Ollama" />
17
+ </p>
18
+
19
+ <p align="center">
20
+ 🌐 <strong>English</strong> &bull;
21
+ <a href="README.fr.md">Français</a> &bull;
22
+ <a href="README.zh-CN.md">中文</a> &bull;
23
+ <a href="README.ar.md">العربية</a> &bull;
24
+ <a href="README.sr.md">Srpski</a>
25
+ </p>
26
+
27
+ <p align="center">
28
+ <a href="#quick-start">Quick Start</a> &bull;
29
+ <a href="#features">Features</a> &bull;
30
+ <a href="#screenshots">Screenshots</a> &bull;
31
+ <a href="#who-is-this-for">Who Is This For</a> &bull;
32
+ <a href="CHANGELOG.md">Changelog</a>
33
+ </p>
34
+
35
+ <p align="center">
36
+ <img src="docs/public/screenshots/demo.gif" alt="LLMxRay demo — real-time token streaming with confidence coloring" width="800" />
37
+ </p>
38
+
39
+ ---
40
+
41
+ ## Quick Start
42
+
43
+ **One command. 30 seconds.**
44
+
45
+ ```bash
46
+ npx llmxray
47
+ ```
48
+
49
+ Or with Docker:
50
+
51
+ ```bash
52
+ docker run -p 5174:5174 djovaneli/llmxray
53
+ ```
54
+
55
+ Open **http://localhost:5174** and start chatting. That's it.
56
+
57
+ > **Prerequisite:** [Ollama](https://ollama.com/download) running locally with at least one model pulled (`ollama pull llama3.2`).
58
+
59
+ ---
60
+
61
+ ## Why LLMxRay?
62
+
63
+ You run a local LLM. You chat with it. But what actually happened?
64
+
65
+ - How fast was each token? Which ones was the model confident about?
66
+ - Is the response quality degrading over long conversations?
67
+ - What would this have cost if you ran it in the cloud?
68
+ - Is the model repeating itself? Refusing? Generating gibberish?
69
+ - How does temperature 0.3 compare to 0.9 on the *same* prompt?
70
+
71
+ **LLMxRay answers all of these, visually, in real time, for free.**
72
+
73
+ ---
74
+
75
+ ## Features
76
+
77
+ ### Real-Time Chat with Token Intelligence
78
+ Chat with any Ollama model and watch tokens arrive with **confidence coloring** — each token is tinted based on generation speed. Supports markdown, multi-turn conversations, file attachments, vision models, and slash commands. For reasoning models, set the thinking budget per conversation — off, model's choice, or an explicit low / medium / high / max effort.
79
+
80
+ ### Response Quality Gates
81
+ Every response is automatically analyzed. Colored badges appear only when something is wrong:
82
+ - **Repetition** — excessive repeated phrases (4-gram analysis)
83
+ - **Refusal** — "as an AI language model" and 7 other patterns
84
+ - **Gibberish** — high non-ASCII ratio
85
+ - **Empty** — fewer than 10 words
86
+ - **Truncation** — hit the token limit without finishing
87
+
88
+ ### Model Comparison Workbench
89
+ Up to **4 slots** with independent model, temperature, and system prompt. Features include side-by-side streaming, word-level diff highlighting, metrics comparison, and one-click presets (Temperature Sweep, Deterministic Pair, Language Compare with Token Tax visualization).
90
+
91
+ ### Performance Analytics
92
+ - **Latency percentiles** (P50/P95/P99) for duration and TTFT
93
+ - **Error intelligence** — 7-category classifier with timeline
94
+ - **Usage heatmap** — 7x24 grid of your active hours
95
+ - **Settings impact** — temperature vs tokens/sec scatter plots
96
+ - **Cold vs warm start** tracking with model load history
97
+
98
+ ### Cost Dashboard
99
+ Token usage per model/day with estimated cloud-equivalent pricing. See what you're *saving* by running locally.
100
+
101
+ ### Surgical Benchmark
102
+ Test model knowledge with multi-choice question suites. Uses real logprobs via OpenAI-compatible endpoint for accurate confidence measurement. Build custom suites visually or let AI generate them from a topic.
103
+
104
+ ### Embeddings Lab & RAG Pipeline
105
+ Embed text, visualize vectors, measure cosine similarity. Request a narrower output vector to see what Matryoshka truncation costs in similarity. Build a local knowledge base from PDFs, DOCX, and CSV — chunked, embedded, and searchable. All stored in IndexedDB. Zero cost.
106
+
107
+ ### Tool Workshop (Visual Canvas)
108
+ Drag-and-drop node canvas for building tool definitions. Bidirectional code sync (edit nodes or TypeScript — both update). Probe APIs, auto-generate schemas, test with live execution.
109
+
110
+ ### Fill-in-the-Middle Playground *(new in v0.4.7)*
111
+ Code completion for Qwen-Coder, CodeLlama, Codestral, DeepSeek-Coder, and StarCoder. Two textareas (prefix / suffix), the model fills the gap. Uses Ollama's `suffix` field on `/api/generate`. Stitched preview shows the result as it would appear in your editor.
112
+
113
+ ### Protocol Observatory *(new in v0.4.7)*
114
+ Fire the same prompt through Ollama's three serving protocols — **native** `/api/chat`, **OpenAI-compat** `/v1/chat/completions`, and **Anthropic-compat** `/v1/messages` — in parallel against your local model. Side-by-side streaming, per-protocol metrics, and an envelope-diff tab that shows how each protocol frames finish reasons, token counts, and error envelopes. No cloud, no API keys — all three endpoints are local on `localhost:11434`.
115
+
116
+ ### AI Training Pipeline
117
+ Curate training data from your conversations. Tag, review, and export as JSONL for fine-tuning.
118
+
119
+ ### Local AI History Database
120
+ Every experiment (benchmarks, comparisons, chats, training pairs) is automatically archived in a queryable IndexedDB database with filters, trends, exports, and retention policies.
121
+
122
+ ### Multilingual
123
+ Full translations in English, French, Serbian (Latin + Cyrillic), Chinese, and Arabic. RTL layout support. Community scaffolds for Hebrew and Japanese.
124
+
125
+ ---
126
+
127
+ ## Ollama Compatibility
128
+
129
+ Tested and verified against **Ollama 0.33.x** (verified on 0.33.3, September 2026). LLMxRay uses these Ollama endpoints:
130
+
131
+ | Endpoint | Used for |
132
+ |---|---|
133
+ | `/api/chat` | Streaming chat (NDJSON, with `tools`, `think` effort levels, `format` schema) |
134
+ | `/api/generate` | Generation + Fill-in-the-Middle via `suffix` |
135
+ | `/api/tags` | Model list + capabilities, context length, and embedding width |
136
+ | `/api/show` | Parameters, template, license, and architecture metadata |
137
+ | `/api/embed` | Vector embeddings for RAG, with optional `dimensions` truncation |
138
+ | `/api/pull`, `/api/delete`, `/api/ps`, `/api/version` | Model management + status |
139
+ | `/v1/chat/completions` | OpenAI-compat path used by Surgical Benchmark for real logprobs and usage totals |
140
+ | `/v1/messages` | Anthropic-compat path used by Protocol Observatory |
141
+
142
+ **Compatible with:** Ollama 0.20 and newer (older versions work for chat/generate but lack `think` and JSON-schema `format`). **Recommended:** Ollama 0.33.3+ — prompt-cache reuse is reported (`prompt_eval_cached_count`, and `usage.prompt_tokens_details.cached_tokens` on the OpenAI-compatible endpoint), so prefill throughput is measured over the tokens actually evaluated. From 0.32: capabilities and context length arrive with the model listing, `think` accepts graded effort levels, and embeddings accept a `dimensions` width.
143
+
144
+ ---
145
+
146
+ ## Screenshots
147
+
148
+ <table>
149
+ <tr>
150
+ <td width="50%">
151
+
152
+ **Chat with token streaming and confidence**
153
+ ![Chat](docs/public/screenshots/chat-diagnostics.png)
154
+
155
+ </td>
156
+ <td width="50%">
157
+
158
+ **Model comparison — side by side**
159
+ ![Compare](docs/public/screenshots/compare-sidebyside.png)
160
+
161
+ </td>
162
+ </tr>
163
+ <tr>
164
+ <td width="50%">
165
+
166
+ **Session deep dive — metrics and timing**
167
+ ![Session](docs/public/screenshots/session-details.png)
168
+
169
+ </td>
170
+ <td width="50%">
171
+
172
+ **Benchmark with confidence radar**
173
+ ![Benchmark](docs/public/screenshots/benchmark.png)
174
+
175
+ </td>
176
+ </tr>
177
+ <tr>
178
+ <td width="50%">
179
+
180
+ **Embeddings — cosine similarity**
181
+ ![Embeddings](docs/public/screenshots/embed-similarity.png)
182
+
183
+ </td>
184
+ <td width="50%">
185
+
186
+ **System monitor — hardware and Ollama status**
187
+ ![System](docs/public/screenshots/my-system.png)
188
+
189
+ </td>
190
+ </tr>
191
+ </table>
192
+
193
+ ---
194
+
195
+ ## Who Is This For
196
+
197
+ | You are... | LLMxRay helps you... |
198
+ |---|---|
199
+ | **Developer** | Debug prompts, profile latency, compare models, inspect tool calls, track costs |
200
+ | **Researcher** | Run controlled experiments with consistent settings across models and temperatures |
201
+ | **Student / Educator** | Explore model behavior visually built-in Educators Kit with 9 interactive modules |
202
+ | **AI team lead** | Understand quality trends, error patterns, and resource usage across your local fleet |
203
+
204
+ ---
205
+
206
+ ## Install Options
207
+
208
+ ### npx (recommended)
209
+ ```bash
210
+ npx llmxray
211
+ npx llmxray --port 3000
212
+ npx llmxray --ollama-url http://192.168.1.50:11434
213
+ ```
214
+
215
+ ### Docker
216
+ ```bash
217
+ docker run -p 5174:5174 djovaneli/llmxray
218
+ docker run -p 5174:5174 -e OLLAMA_URL=http://host.docker.internal:11434 djovaneli/llmxray
219
+ ```
220
+
221
+ ### From source
222
+ ```bash
223
+ git clone https://github.com/LogneBudo/llmxray.git
224
+ cd llmxray
225
+ npm install
226
+ npm run dev # http://localhost:5173
227
+ ```
228
+
229
+ ---
230
+
231
+ ## Tech Stack
232
+
233
+ | Layer | Technology |
234
+ |---|---|
235
+ | Framework | Vue 3.5 + Composition API |
236
+ | Language | TypeScript 5.9 (strict) |
237
+ | Build | Vite 7.3 |
238
+ | Styling | Tailwind CSS 4.2 |
239
+ | State | Pinia 3 (store-per-concern) |
240
+ | Charts | Chart.js 4, D3.js 7 |
241
+ | Canvas | Vue Flow (visual node editor) |
242
+ | Code Editor | CodeMirror 6 |
243
+ | Storage | IndexedDB (browser-native) |
244
+ | LLM Backend | Ollama (local) |
245
+
246
+ ---
247
+
248
+ ## Architecture
249
+
250
+ **Streaming** — Reads Ollama NDJSON via `fetch()` + `ReadableStream`. Tokens update the UI reactively through Pinia stores.
251
+
252
+ **Token confidence** — Approximated from inter-token latency (faster = more confident). Clearly labeled as approximation. Benchmarks use real logprobs via OpenAI-compatible endpoint.
253
+
254
+ **Store-per-concern** — Each domain has its own Pinia store: tokens, sessions, metrics, reasoning, comparison, embeddings, quality, cost, and more.
255
+
256
+ **Hardware detection** — Custom Vite plugin queries the OS directly (PowerShell/proc/sysctl) for accurate hardware specs.
257
+
258
+ ---
259
+
260
+ ## Development
261
+
262
+ | Command | What it does |
263
+ |---|---|
264
+ | `npm run dev` | Dev server (port 5173) |
265
+ | `npm run build` | Type-check + production build |
266
+ | `npm run test` | Unit tests (Vitest) |
267
+ | `npm run test:e2e` | End-to-end (Playwright) |
268
+
269
+ ---
270
+
271
+ ## Contributing
272
+
273
+ Contributions welcome! See [CONTRIBUTING.md](CONTRIBUTING.md) for setup and guidelines.
274
+
275
+ **Community translations especially welcome** — scaffold files ready for Hebrew and Japanese.
276
+
277
+ ---
278
+
279
+ ## License
280
+
281
+ [Apache License 2.0](LICENSE)
282
+
283
+ ## Trademark
284
+
285
+ **LLMxRay** is a trademark of Ivan Stankovic ([LogneBudo](https://github.com/LogneBudo)). See [TRADEMARK.md](TRADEMARK.md).
286
+
287
+ ---
288
+
289
+ <p align="center">
290
+ <strong>If LLMxRay helps you understand your AI better, consider giving it a star.</strong><br/>
291
+ It helps others discover the project.
292
+ </p>
293
+
294
+ <p align="center">
295
+ <a href="https://github.com/LogneBudo/llmxray">
296
+ <img src="https://img.shields.io/github/stars/LogneBudo/llmxray?style=social" alt="GitHub stars" />
297
+ </a>
298
+ </p>