llmxray 0.4.0 → 0.4.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +265 -244
- package/dist/assets/{AITrainingPage-Ck1bDWWB.js → AITrainingPage-VU-Nu5i0.js} +1 -1
- package/dist/assets/{BenchmarkPage-VzMNGUEy.js → BenchmarkPage-CQUTAU0c.js} +1 -1
- package/dist/assets/{ComparisonPage-BBliPm7K.js → ComparisonPage-DSO_-JZF.js} +1 -1
- package/dist/assets/CostDashboardPage-BLzCE-bT.js +1 -0
- package/dist/assets/DashboardPage-lqAFJvOo.js +116 -0
- package/dist/assets/{EmbeddingsPage-BIHHeJq6.js → EmbeddingsPage-DDBfFe8z.js} +1 -1
- package/dist/assets/{GoogleCallbackPage-2gGq0Swg.js → GoogleCallbackPage-H8X6YDz9.js} +1 -1
- package/dist/assets/{JsonTreeNode-CDR0az7J.js → JsonTreeNode-DQ_R5Euc.js} +1 -1
- package/dist/assets/{ModelCapabilityIcons.vue_vue_type_script_setup_true_lang-DkuvrY_5.js → ModelCapabilityIcons.vue_vue_type_script_setup_true_lang-CV939MUZ.js} +1 -1
- package/dist/assets/{RAGPage-BdXerRDW.js → RAGPage-DbZxWHmi.js} +1 -1
- package/dist/assets/SessionPage-DvtFE4GS.js +15 -0
- package/dist/assets/{SettingsPage-B1wm-L6T.js → SettingsPage-fmgr8zF9.js} +1 -1
- package/dist/assets/{StatusBadge.vue_vue_type_script_setup_true_lang-J8rYaobd.js → StatusBadge.vue_vue_type_script_setup_true_lang-BbHpA7hy.js} +1 -1
- package/dist/assets/{StorageGauge.vue_vue_type_style_index_0_lang-Begwajgo.js → StorageGauge.vue_vue_type_style_index_0_lang-Dd3BsCg6.js} +1 -1
- package/dist/assets/{SystemPage-DDZrsYds.js → SystemPage-Bd3SBsWI.js} +2 -2
- package/dist/assets/{TabBar.vue_vue_type_script_setup_true_lang-DVCeD1qL.js → TabBar.vue_vue_type_script_setup_true_lang-DlSEPHs6.js} +1 -1
- package/dist/assets/{TokenStreamDisplay.vue_vue_type_script_setup_true_lang-rJpx5sU7.js → TokenStreamDisplay.vue_vue_type_script_setup_true_lang-P-I0kJTd.js} +1 -1
- package/dist/assets/{ToolWorkshopPage-DnHp32vX.js → ToolWorkshopPage-D9doB8PF.js} +1 -1
- package/dist/assets/{agent-store-BWO6HamI.js → agent-store-CuCtjbhV.js} +1 -1
- package/dist/assets/{canvas-ai-db-anWOo2F0.js → canvas-ai-db-DBUT-nyD.js} +1 -1
- package/dist/assets/{download-BqPPd3_X.js → download-C3s6K1NW.js} +1 -1
- package/dist/assets/{generate-service-C3Z8NZaT.js → generate-service-DrRYe_-5.js} +1 -1
- package/dist/assets/{google-auth-store-CgUWvtRU.js → google-auth-store-B0WvuFWG.js} +1 -1
- package/dist/assets/index-Bawj8X2v.css +1 -0
- package/dist/assets/{index-KhnkHbiC.js → index-CUNpOQX9.js} +1 -1
- package/dist/assets/index-DYHBy1xj.js +8 -0
- package/dist/assets/{index-B7GdQF1S.js → index-DoMyvAkM.js} +1 -1
- package/dist/assets/info-D-oZIv7P.js +1 -0
- package/dist/assets/{metrics-store-BMhkZ9B4.js → metrics-store-DRErjO4D.js} +1 -1
- package/dist/assets/{papaparse.min-z5sNvmdo.js → papaparse.min-DJzuA_ek.js} +1 -1
- package/dist/assets/{pencil-Csf0-tn3.js → pencil-MO2uWQbW.js} +1 -1
- package/dist/assets/{rag-store-Ch0N5LTe.js → rag-store-CBWvD3Va.js} +3 -3
- package/dist/assets/{share-2-DIi3V2dc.js → share-2-CnM5sBDm.js} +1 -1
- package/dist/assets/{storage-store-uPTARmbg.js → storage-store-cx_2Hb_J.js} +1 -1
- package/dist/assets/{stream-handler-BuwD741F.js → stream-handler-CXdBM2li.js} +1 -1
- package/dist/assets/{tool-workshop-store-koPpdoxb.js → tool-workshop-store-Cb4eb3e4.js} +1 -1
- package/dist/assets/{toolcall-store-JKmpIFj3.js → toolcall-store-xQEQUq5f.js} +1 -1
- package/dist/assets/{trash-2-aeVGvgIz.js → trash-2-Cj-NOsVZ.js} +1 -1
- package/dist/index.html +2 -2
- package/package.json +1 -1
- package/dist/assets/DashboardPage-CH15t_MS.js +0 -116
- package/dist/assets/SessionPage-B39x2A8D.js +0 -15
- package/dist/assets/index-ClE7gpZb.js +0 -8
- package/dist/assets/index-Dr1ztyT3.css +0 -1
package/README.md
CHANGED
|
@@ -1,244 +1,265 @@
|
|
|
1
|
-
<p align="center">
|
|
2
|
-
<img src="https://raw.githubusercontent.com/LogneBudo/llmxray/master/public/favicon.svg" alt="LLMxRay" width="80" />
|
|
3
|
-
</p>
|
|
4
|
-
|
|
5
|
-
<h1 align="center">LLMxRay</h1>
|
|
6
|
-
<p align="center"><strong>Local LLM Observatory</strong></p>
|
|
7
|
-
<p align="center">
|
|
8
|
-
See what your AI is <em>actually</em> doing — token by token, layer by layer.
|
|
9
|
-
</p>
|
|
10
|
-
|
|
11
|
-
<p align="center">
|
|
12
|
-
<img src="https://img.shields.io/badge/vue-3.5-42b883?logo=vuedotjs&logoColor=white" alt="Vue 3.5" />
|
|
13
|
-
<img src="https://img.shields.io/badge/vite-7.3-646cff?logo=vite&logoColor=white" alt="Vite 7.3" />
|
|
14
|
-
<img src="https://img.shields.io/badge/typescript-5.9-3178c6?logo=typescript&logoColor=white" alt="TypeScript 5.9" />
|
|
15
|
-
<img src="https://img.shields.io/badge/tailwind-4.2-06b6d4?logo=tailwindcss&logoColor=white" alt="Tailwind 4.2" />
|
|
16
|
-
<img src="https://img.shields.io/badge/ollama-local-000?logo=ollama&logoColor=white" alt="Ollama" />
|
|
17
|
-
</p>
|
|
18
|
-
|
|
19
|
-
---
|
|
20
|
-
|
|
21
|
-
## What is LLMxRay?
|
|
22
|
-
|
|
23
|
-
LLMxRay is a **free, local-first** dashboard that connects to [Ollama](https://ollama.com) running on your machine. It lets you chat with any model you've downloaded and then **inspect everything that happened behind the scenes**: how fast each token arrived, what the model might have been "thinking", how different settings change the output, and much more.
|
|
24
|
-
|
|
25
|
-
**No cloud. No API keys. No cost.** Everything runs on your hardware.
|
|
26
|
-
|
|
27
|
-
### Who is this for?
|
|
28
|
-
|
|
29
|
-
| You are... | LLMxRay helps you... |
|
|
30
|
-
|---|---|
|
|
31
|
-
| **Curious beginner** | See AI responses form in real-time and learn what "temperature" or "tokens" actually mean |
|
|
32
|
-
| **Student / educator** | Explore model behavior visually — great for AI/ML coursework and demos |
|
|
33
|
-
| **Developer** | Debug prompts, compare models, profile latency, inspect tool calls |
|
|
34
|
-
| **Researcher** | Run controlled experiments: same prompt, different settings, side-by-side results |
|
|
35
|
-
|
|
36
|
-
---
|
|
37
|
-
|
|
38
|
-
## Features at a Glance
|
|
39
|
-
|
|
40
|
-
### Chat with Real-Time Token Streaming
|
|
41
|
-
Start a conversation with any Ollama model. Watch tokens appear one by one with **confidence coloring** — each token is tinted based on how quickly the model produced it (faster = more confident). Supports markdown rendering, multi-turn conversation, file attachments, and slash commands.
|
|
42
|
-
|
|
43
|
-
### Session Deep Dive
|
|
44
|
-
Click any past session to explore six tabs of detail:
|
|
45
|
-
|
|
46
|
-
- **Stream** — Every token with timing data, plus a metrics dashboard (time-to-first-token, tokens/sec, latency chart)
|
|
47
|
-
- **Reasoning** — If you're running a reasoning model like DeepSeek-R1, the `<think>` blocks are parsed and displayed step-by-step
|
|
48
|
-
- **Introspection** — Visualizations of layer activations, attention heatmaps, and model architecture (illustrative)
|
|
49
|
-
- **Tools** — Timeline of any tool calls the model made, with parameters and results
|
|
50
|
-
- **Agent** — State-flow graph showing how an agent-style prompt progressed
|
|
51
|
-
- **Prompt** — Anatomy breakdown of your prompt: sections, token counts, structure
|
|
52
|
-
|
|
53
|
-
### Compare Models (and Settings)
|
|
54
|
-
The comparison workbench goes beyond "Model A vs Model B". Create up to **4 slots**, each with its own model, temperature, system prompt, and sampling parameters. Compare the *same* model at different temperatures to see how creativity changes. Features include:
|
|
55
|
-
|
|
56
|
-
- **Grid view** — Side-by-side streaming results with per-slot settings pills
|
|
57
|
-
- **Diff view** — Word-level highlighting of what changed between two outputs
|
|
58
|
-
- **Metrics bar** — Visual comparison of TTFT, tokens/sec, and total tokens
|
|
59
|
-
- **Quick presets** — "Temperature Sweep" (3 temps) and "Deterministic Pair" (same seed) one-click setups
|
|
60
|
-
- Embedding models are automatically filtered out — only chat-capable models appear
|
|
61
|
-
|
|
62
|
-
### Embeddings Lab
|
|
63
|
-
Embed any text and visualize the resulting vector. Compare two texts with a **cosine similarity meter** to see how semantically close they are. A hands-on way to understand what embeddings actually represent.
|
|
64
|
-
|
|
65
|
-
### RAG Pipeline
|
|
66
|
-
Build a local knowledge base from your documents:
|
|
67
|
-
|
|
68
|
-
1. **Upload** PDFs, Word docs (.docx), or CSVs
|
|
69
|
-
2. **Chunk & embed** automatically using your chosen embedding model
|
|
70
|
-
3. **Search** with natural language — results ranked by semantic similarity
|
|
71
|
-
|
|
72
|
-
Everything is stored in **IndexedDB** (your browser's built-in database). Zero cost, zero setup, zero external services.
|
|
73
|
-
|
|
74
|
-
### Tool Workshop (Visual Canvas)
|
|
75
|
-
Build, edit, and test tool definitions on an interactive **node-based canvas** powered by Vue Flow:
|
|
76
|
-
|
|
77
|
-
- **Drag-and-drop nodes** — Each tool is a visual node showing name, description, parameters, and implementation body
|
|
78
|
-
- **Inline code editing** — Full CodeMirror 6 editors with TypeScript syntax highlighting directly on each node
|
|
79
|
-
- **Bidirectional code sync** — Open the Code Panel to see all tools as combined TypeScript source. Edit code, nodes update. Edit nodes, code updates. Powered by a Recast AST parser
|
|
80
|
-
- **Schema viewer** — Auto-generated OpenAI-compatible JSON schemas with one-click copy
|
|
81
|
-
- **Probe & Pick** — Point at any API URL, inspect the response JSON tree, and auto-generate fetch code + parameter mappings
|
|
82
|
-
- **OpenAPI discovery** — Auto-detect and parse OpenAPI/Swagger specs to pick endpoints visually
|
|
83
|
-
- **Live execution overlays** — During chat, tool nodes pulse when the model calls them and show results inline
|
|
84
|
-
- **Templates** — Start from 15+ built-in templates (web fetch, calculator, Google Calendar/Gmail, regex tester, and more)
|
|
85
|
-
- **Persistent layout** — Node positions, mappings, and probe configs survive across sessions
|
|
86
|
-
|
|
87
|
-
### Tool Call Optimizer
|
|
88
|
-
When the model calls a tool during chat, an **"Optimize this Tool"** button appears on the result. Click it to open the Response Optimizer Drawer:
|
|
89
|
-
|
|
90
|
-
- Visualize the API response as an interactive JSON tree
|
|
91
|
-
- Select only the fields the model actually needs
|
|
92
|
-
- Auto-generate optimized fetch code with field extraction
|
|
93
|
-
- One click to create a new optimized tool in the Workshop
|
|
94
|
-
|
|
95
|
-
###
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
|
|
122
|
-
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
|
|
131
|
-
|
|
132
|
-
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
|
|
140
|
-
|
|
141
|
-
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
|
|
145
|
-
|
|
146
|
-
|
|
147
|
-
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
152
|
-
|
|
153
|
-
|
|
154
|
-
|
|
155
|
-
|
|
156
|
-
|
|
157
|
-
|
|
158
|
-
|
|
159
|
-
|
|
160
|
-
|
|
161
|
-
|
|
162
|
-
|
|
163
|
-
|
|
164
|
-
|
|
165
|
-
|
|
166
|
-
|
|
167
|
-
|
|
168
|
-
|
|
169
|
-
|
|
170
|
-
|
|
171
|
-
|
|
172
|
-
|
|
173
|
-
|
|
174
|
-
|
|
175
|
-
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
|
|
179
|
-
|
|
180
|
-
|
|
181
|
-
|
|
182
|
-
|
|
183
|
-
|
|
184
|
-
|
|
185
|
-
|
|
186
|
-
|
|
187
|
-
|
|
188
|
-
|
|
189
|
-
|
|
190
|
-
|
|
191
|
-
|
|
192
|
-
|
|
193
|
-
|
|
194
|
-
|
|
195
|
-
|
|
196
|
-
|
|
197
|
-
|
|
198
|
-
|
|
199
|
-
|
|
200
|
-
|
|
201
|
-
|
|
202
|
-
|
|
203
|
-
|
|
204
|
-
|
|
205
|
-
|
|
206
|
-
|
|
207
|
-
|
|
208
|
-
|
|
209
|
-
|
|
210
|
-
|
|
211
|
-
|
|
212
|
-
|
|
213
|
-
|
|
214
|
-
|
|
215
|
-
|
|
216
|
-
|
|
217
|
-
|
|
218
|
-
|
|
219
|
-
|
|
220
|
-
|
|
221
|
-
|
|
222
|
-
|
|
223
|
-
|
|
224
|
-
|
|
225
|
-
|
|
226
|
-
|
|
227
|
-
|
|
228
|
-
|
|
229
|
-
|
|
230
|
-
|
|
231
|
-
|
|
232
|
-
|
|
233
|
-
|
|
234
|
-
|
|
235
|
-
|
|
236
|
-
|
|
237
|
-
|
|
238
|
-
|
|
239
|
-
|
|
240
|
-
|
|
241
|
-
|
|
242
|
-
|
|
243
|
-
|
|
244
|
-
|
|
1
|
+
<p align="center">
|
|
2
|
+
<img src="https://raw.githubusercontent.com/LogneBudo/llmxray/master/public/favicon.svg" alt="LLMxRay" width="80" />
|
|
3
|
+
</p>
|
|
4
|
+
|
|
5
|
+
<h1 align="center">LLMxRay</h1>
|
|
6
|
+
<p align="center"><strong>Local LLM Observatory</strong></p>
|
|
7
|
+
<p align="center">
|
|
8
|
+
See what your AI is <em>actually</em> doing — token by token, layer by layer.
|
|
9
|
+
</p>
|
|
10
|
+
|
|
11
|
+
<p align="center">
|
|
12
|
+
<img src="https://img.shields.io/badge/vue-3.5-42b883?logo=vuedotjs&logoColor=white" alt="Vue 3.5" />
|
|
13
|
+
<img src="https://img.shields.io/badge/vite-7.3-646cff?logo=vite&logoColor=white" alt="Vite 7.3" />
|
|
14
|
+
<img src="https://img.shields.io/badge/typescript-5.9-3178c6?logo=typescript&logoColor=white" alt="TypeScript 5.9" />
|
|
15
|
+
<img src="https://img.shields.io/badge/tailwind-4.2-06b6d4?logo=tailwindcss&logoColor=white" alt="Tailwind 4.2" />
|
|
16
|
+
<img src="https://img.shields.io/badge/ollama-local-000?logo=ollama&logoColor=white" alt="Ollama" />
|
|
17
|
+
</p>
|
|
18
|
+
|
|
19
|
+
---
|
|
20
|
+
|
|
21
|
+
## What is LLMxRay?
|
|
22
|
+
|
|
23
|
+
LLMxRay is a **free, local-first** dashboard that connects to [Ollama](https://ollama.com) running on your machine. It lets you chat with any model you've downloaded and then **inspect everything that happened behind the scenes**: how fast each token arrived, what the model might have been "thinking", how different settings change the output, and much more.
|
|
24
|
+
|
|
25
|
+
**No cloud. No API keys. No cost.** Everything runs on your hardware.
|
|
26
|
+
|
|
27
|
+
### Who is this for?
|
|
28
|
+
|
|
29
|
+
| You are... | LLMxRay helps you... |
|
|
30
|
+
|---|---|
|
|
31
|
+
| **Curious beginner** | See AI responses form in real-time and learn what "temperature" or "tokens" actually mean |
|
|
32
|
+
| **Student / educator** | Explore model behavior visually — great for AI/ML coursework and demos |
|
|
33
|
+
| **Developer** | Debug prompts, compare models, profile latency, inspect tool calls |
|
|
34
|
+
| **Researcher** | Run controlled experiments: same prompt, different settings, side-by-side results |
|
|
35
|
+
|
|
36
|
+
---
|
|
37
|
+
|
|
38
|
+
## Features at a Glance
|
|
39
|
+
|
|
40
|
+
### Chat with Real-Time Token Streaming
|
|
41
|
+
Start a conversation with any Ollama model. Watch tokens appear one by one with **confidence coloring** — each token is tinted based on how quickly the model produced it (faster = more confident). Supports markdown rendering, multi-turn conversation, file attachments, and slash commands.
|
|
42
|
+
|
|
43
|
+
### Session Deep Dive
|
|
44
|
+
Click any past session to explore six tabs of detail:
|
|
45
|
+
|
|
46
|
+
- **Stream** — Every token with timing data, plus a metrics dashboard (time-to-first-token, tokens/sec, latency chart)
|
|
47
|
+
- **Reasoning** — If you're running a reasoning model like DeepSeek-R1, the `<think>` blocks are parsed and displayed step-by-step
|
|
48
|
+
- **Introspection** — Visualizations of layer activations, attention heatmaps, and model architecture (illustrative)
|
|
49
|
+
- **Tools** — Timeline of any tool calls the model made, with parameters and results
|
|
50
|
+
- **Agent** — State-flow graph showing how an agent-style prompt progressed
|
|
51
|
+
- **Prompt** — Anatomy breakdown of your prompt: sections, token counts, structure
|
|
52
|
+
|
|
53
|
+
### Compare Models (and Settings)
|
|
54
|
+
The comparison workbench goes beyond "Model A vs Model B". Create up to **4 slots**, each with its own model, temperature, system prompt, and sampling parameters. Compare the *same* model at different temperatures to see how creativity changes. Features include:
|
|
55
|
+
|
|
56
|
+
- **Grid view** — Side-by-side streaming results with per-slot settings pills
|
|
57
|
+
- **Diff view** — Word-level highlighting of what changed between two outputs
|
|
58
|
+
- **Metrics bar** — Visual comparison of TTFT, tokens/sec, and total tokens
|
|
59
|
+
- **Quick presets** — "Temperature Sweep" (3 temps) and "Deterministic Pair" (same seed) one-click setups
|
|
60
|
+
- Embedding models are automatically filtered out — only chat-capable models appear
|
|
61
|
+
|
|
62
|
+
### Embeddings Lab
|
|
63
|
+
Embed any text and visualize the resulting vector. Compare two texts with a **cosine similarity meter** to see how semantically close they are. A hands-on way to understand what embeddings actually represent.
|
|
64
|
+
|
|
65
|
+
### RAG Pipeline
|
|
66
|
+
Build a local knowledge base from your documents:
|
|
67
|
+
|
|
68
|
+
1. **Upload** PDFs, Word docs (.docx), or CSVs
|
|
69
|
+
2. **Chunk & embed** automatically using your chosen embedding model
|
|
70
|
+
3. **Search** with natural language — results ranked by semantic similarity
|
|
71
|
+
|
|
72
|
+
Everything is stored in **IndexedDB** (your browser's built-in database). Zero cost, zero setup, zero external services.
|
|
73
|
+
|
|
74
|
+
### Tool Workshop (Visual Canvas)
|
|
75
|
+
Build, edit, and test tool definitions on an interactive **node-based canvas** powered by Vue Flow:
|
|
76
|
+
|
|
77
|
+
- **Drag-and-drop nodes** — Each tool is a visual node showing name, description, parameters, and implementation body
|
|
78
|
+
- **Inline code editing** — Full CodeMirror 6 editors with TypeScript syntax highlighting directly on each node
|
|
79
|
+
- **Bidirectional code sync** — Open the Code Panel to see all tools as combined TypeScript source. Edit code, nodes update. Edit nodes, code updates. Powered by a Recast AST parser
|
|
80
|
+
- **Schema viewer** — Auto-generated OpenAI-compatible JSON schemas with one-click copy
|
|
81
|
+
- **Probe & Pick** — Point at any API URL, inspect the response JSON tree, and auto-generate fetch code + parameter mappings
|
|
82
|
+
- **OpenAPI discovery** — Auto-detect and parse OpenAPI/Swagger specs to pick endpoints visually
|
|
83
|
+
- **Live execution overlays** — During chat, tool nodes pulse when the model calls them and show results inline
|
|
84
|
+
- **Templates** — Start from 15+ built-in templates (web fetch, calculator, Google Calendar/Gmail, regex tester, and more)
|
|
85
|
+
- **Persistent layout** — Node positions, mappings, and probe configs survive across sessions
|
|
86
|
+
|
|
87
|
+
### Tool Call Optimizer
|
|
88
|
+
When the model calls a tool during chat, an **"Optimize this Tool"** button appears on the result. Click it to open the Response Optimizer Drawer:
|
|
89
|
+
|
|
90
|
+
- Visualize the API response as an interactive JSON tree
|
|
91
|
+
- Select only the fields the model actually needs
|
|
92
|
+
- Auto-generate optimized fetch code with field extraction
|
|
93
|
+
- One click to create a new optimized tool in the Workshop
|
|
94
|
+
|
|
95
|
+
### Response Quality Gates
|
|
96
|
+
Every assistant response is automatically analyzed for common quality issues. Small colored badges appear below the response metrics when problems are detected:
|
|
97
|
+
|
|
98
|
+
- **Repetition** — Flags responses with excessive repeated 4-gram phrases (>50% = fail, >30% = warn)
|
|
99
|
+
- **Refusal** — Detects 8 common refusal patterns ("as an AI language model", "I cannot help", etc.)
|
|
100
|
+
- **Gibberish** — Warns when non-ASCII characters exceed 40% of the response
|
|
101
|
+
- **Empty** — Flags responses with fewer than 10 words
|
|
102
|
+
- **Truncation** — Warns when the response hit the token limit or used >90% of budget without clean ending
|
|
103
|
+
|
|
104
|
+
No news is good news — badges only appear when something is wrong.
|
|
105
|
+
|
|
106
|
+
### Cost Dashboard
|
|
107
|
+
Track token usage across all your sessions with estimated cloud-equivalent costs. Navigate to the **Costs** page in the sidebar to see:
|
|
108
|
+
|
|
109
|
+
- **Summary cards** — Total tokens, sessions, estimated cost, average cost per session
|
|
110
|
+
- **Token usage by model** — Stacked bar chart showing prompt vs completion tokens per model
|
|
111
|
+
- **Daily usage trends** — Line chart with dual axes (tokens + estimated cost over time)
|
|
112
|
+
- **Model breakdown table** — Detailed per-model statistics with pricing source transparency
|
|
113
|
+
|
|
114
|
+
Costs are estimates based on equivalent cloud API pricing (Groq, Together AI, Google, Mistral, etc.). Ollama runs locally at zero cost — the dashboard shows what you're saving.
|
|
115
|
+
|
|
116
|
+
### Model Browser
|
|
117
|
+
See every model installed in Ollama with details like parameter count, quantization level, family, and format. Includes architecture diagrams showing the model's structure.
|
|
118
|
+
|
|
119
|
+
### System Monitor
|
|
120
|
+
Real hardware specs (not browser estimates) — CPU model, total RAM with live usage, GPU with driver version, storage. Plus live Ollama status: running models, memory allocation, inference settings.
|
|
121
|
+
|
|
122
|
+
### Settings
|
|
123
|
+
Configure your Ollama connection URL with a live connection tester. Set default **temperature** and **context length** with visual scales and educational tooltips that explain what each setting does in plain language.
|
|
124
|
+
|
|
125
|
+
---
|
|
126
|
+
|
|
127
|
+
## Quick Start
|
|
128
|
+
|
|
129
|
+
### Prerequisites
|
|
130
|
+
|
|
131
|
+
1. **Node.js 18+** — [Download](https://nodejs.org)
|
|
132
|
+
2. **Ollama** running locally — [Download](https://ollama.com/download)
|
|
133
|
+
3. At least one model pulled:
|
|
134
|
+
```bash
|
|
135
|
+
ollama pull llama3.2
|
|
136
|
+
```
|
|
137
|
+
|
|
138
|
+
### Install and Run
|
|
139
|
+
|
|
140
|
+
```bash
|
|
141
|
+
git clone https://github.com/LogneBudo/llmxray.git
|
|
142
|
+
cd llmxray
|
|
143
|
+
npm install
|
|
144
|
+
npm run dev
|
|
145
|
+
```
|
|
146
|
+
|
|
147
|
+
Open **http://localhost:5173** in your browser. That's it.
|
|
148
|
+
|
|
149
|
+
> LLMxRay's dev server automatically proxies API calls to Ollama at `localhost:11434`. If Ollama is running on a different port or machine, change it in **Settings**.
|
|
150
|
+
|
|
151
|
+
### Build for Production
|
|
152
|
+
|
|
153
|
+
```bash
|
|
154
|
+
npm run build # Output in dist/
|
|
155
|
+
npm run preview # Preview the build locally
|
|
156
|
+
```
|
|
157
|
+
|
|
158
|
+
---
|
|
159
|
+
|
|
160
|
+
## Tech Stack
|
|
161
|
+
|
|
162
|
+
| Layer | Technology |
|
|
163
|
+
|---|---|
|
|
164
|
+
| Framework | Vue 3.5 + Composition API (`<script setup>`) |
|
|
165
|
+
| Language | TypeScript 5.9 (strict) |
|
|
166
|
+
| Build | Vite 7.3 |
|
|
167
|
+
| Styling | Tailwind CSS 4.2 (custom dark theme) |
|
|
168
|
+
| State | Pinia 3 (one store per concern) |
|
|
169
|
+
| Routing | Vue Router 5 |
|
|
170
|
+
| Charts | Chart.js 4 + vue-chartjs, D3.js 7 |
|
|
171
|
+
| Canvas | Vue Flow 1.x (node-based visual editor) |
|
|
172
|
+
| Code Editor | CodeMirror 6 (TypeScript + JSON highlighting) |
|
|
173
|
+
| AST Parser | Recast + @babel/parser (bidirectional code sync) |
|
|
174
|
+
| Markdown | marked |
|
|
175
|
+
| Diffing | diff (word-level) |
|
|
176
|
+
| Documents | pdfjs-dist (lazy), mammoth (DOCX), papaparse (CSV) |
|
|
177
|
+
| Storage | IndexedDB (browser-native, zero-cost) |
|
|
178
|
+
| IDs | nanoid |
|
|
179
|
+
| LLM Backend | Ollama (local, via `/api` proxy) |
|
|
180
|
+
|
|
181
|
+
---
|
|
182
|
+
|
|
183
|
+
## Architecture Highlights
|
|
184
|
+
|
|
185
|
+
**Streaming** — LLMxRay reads Ollama's NDJSON response streams via `fetch()` + `ReadableStream`. Tokens arrive one by one and update the UI reactively through Pinia stores.
|
|
186
|
+
|
|
187
|
+
**Token confidence** — Ollama doesn't expose logprobs, so confidence is approximated from inter-token latency. Faster tokens = higher confidence. This is labeled clearly in the UI as an approximation.
|
|
188
|
+
|
|
189
|
+
**Introspection data** — Layer activations and attention heatmaps are synthetic (illustrative). They demonstrate what these visualizations *would* look like with real data. Clearly labeled as "Illustrative" in the UI.
|
|
190
|
+
|
|
191
|
+
**Store-per-concern** — Each domain has its own Pinia store: tokens, sessions, metrics, reasoning, comparison, embeddings, RAG, models, and more. This keeps state management modular and testable.
|
|
192
|
+
|
|
193
|
+
**Hardware detection** — The System page uses a custom Vite plugin (`vite-plugin-system-info.ts`) that queries the OS directly via PowerShell (Windows), `/proc` + `lspci` (Linux), or `sysctl` (macOS) for accurate hardware specs.
|
|
194
|
+
|
|
195
|
+
---
|
|
196
|
+
|
|
197
|
+
## Development
|
|
198
|
+
|
|
199
|
+
### Scripts
|
|
200
|
+
|
|
201
|
+
| Command | What it does |
|
|
202
|
+
|---|---|
|
|
203
|
+
| `npm run dev` | Start dev server (port 5173) |
|
|
204
|
+
| `npm run build` | Type-check + production build |
|
|
205
|
+
| `npm run preview` | Preview production build |
|
|
206
|
+
| `npm run test` | Run unit tests (Vitest) |
|
|
207
|
+
| `npm run test:watch` | Tests in watch mode |
|
|
208
|
+
| `npm run test:coverage` | Coverage report |
|
|
209
|
+
| `npm run test:e2e` | Playwright end-to-end tests |
|
|
210
|
+
| `npm run test:e2e:headed` | E2E with visible browser |
|
|
211
|
+
| `npm run test:e2e:live` | E2E against live Ollama |
|
|
212
|
+
|
|
213
|
+
### Project Structure
|
|
214
|
+
|
|
215
|
+
```
|
|
216
|
+
src/
|
|
217
|
+
pages/ 8 page components (Dashboard, Compare, RAG, etc.)
|
|
218
|
+
components/ 50+ components organized by feature
|
|
219
|
+
chat/ Chat UI, token stream, attachments
|
|
220
|
+
comparison/ Slot configurator, grid, diff view, metrics bar
|
|
221
|
+
metrics/ Dashboard, charts, session history
|
|
222
|
+
reasoning/ Think-block viewer
|
|
223
|
+
introspection/ Layer activations, attention, architecture
|
|
224
|
+
rag/ Document upload, search, ingest
|
|
225
|
+
embeddings/ Vector viz, similarity meter
|
|
226
|
+
tool-canvas/ Visual canvas, node editor, CodeMirror wrapper
|
|
227
|
+
tool-optimizer/ Response optimizer drawer, JSON tree
|
|
228
|
+
tool-calls/ Tool call timeline, definitions
|
|
229
|
+
agent-graph/ Agent state flow
|
|
230
|
+
common/ Layout, sidebar, shared components
|
|
231
|
+
stores/ Pinia stores (one per concern)
|
|
232
|
+
services/ Ollama client, streaming, generation, RAG, AST parser, probe
|
|
233
|
+
types/ TypeScript interfaces
|
|
234
|
+
utils/ Formatting, color scales, slot labels
|
|
235
|
+
composables/ Vue composables (markdown, etc.)
|
|
236
|
+
router/ Route definitions
|
|
237
|
+
```
|
|
238
|
+
|
|
239
|
+
---
|
|
240
|
+
|
|
241
|
+
## Troubleshooting
|
|
242
|
+
|
|
243
|
+
| Problem | Solution |
|
|
244
|
+
|---|---|
|
|
245
|
+
| "Disconnected" in Settings | Make sure Ollama is running: `ollama serve` |
|
|
246
|
+
| No models in dropdowns | Pull a model first: `ollama pull llama3.2` |
|
|
247
|
+
| System page shows "Restart dev server" | Stop and restart `npm run dev` (the hardware plugin loads at startup) |
|
|
248
|
+
| Slow first response | Normal — Ollama loads the model into memory on first use |
|
|
249
|
+
| High RAM/VRAM usage | Use smaller quantized models (Q4) or reduce context length in Settings |
|
|
250
|
+
|
|
251
|
+
---
|
|
252
|
+
|
|
253
|
+
## License
|
|
254
|
+
|
|
255
|
+
Licensed under the [Apache License 2.0](LICENSE). You are free to use, modify, and distribute this software under the terms of the license.
|
|
256
|
+
|
|
257
|
+
## Trademark
|
|
258
|
+
|
|
259
|
+
**LLMxRay** is a trademark of Ivan Stankovic ([LogneBudo](https://github.com/LogneBudo)). See [TRADEMARK.md](TRADEMARK.md) for usage guidelines.
|
|
260
|
+
|
|
261
|
+
---
|
|
262
|
+
|
|
263
|
+
<p align="center">
|
|
264
|
+
Built with curiosity by <a href="https://github.com/LogneBudo">LogneBudo</a>
|
|
265
|
+
</p>
|