llmxray 0.4.0 → 0.4.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (45) hide show
  1. package/README.md +265 -244
  2. package/dist/assets/{AITrainingPage-Ck1bDWWB.js → AITrainingPage-VU-Nu5i0.js} +1 -1
  3. package/dist/assets/{BenchmarkPage-VzMNGUEy.js → BenchmarkPage-CQUTAU0c.js} +1 -1
  4. package/dist/assets/{ComparisonPage-BBliPm7K.js → ComparisonPage-DSO_-JZF.js} +1 -1
  5. package/dist/assets/CostDashboardPage-BLzCE-bT.js +1 -0
  6. package/dist/assets/DashboardPage-lqAFJvOo.js +116 -0
  7. package/dist/assets/{EmbeddingsPage-BIHHeJq6.js → EmbeddingsPage-DDBfFe8z.js} +1 -1
  8. package/dist/assets/{GoogleCallbackPage-2gGq0Swg.js → GoogleCallbackPage-H8X6YDz9.js} +1 -1
  9. package/dist/assets/{JsonTreeNode-CDR0az7J.js → JsonTreeNode-DQ_R5Euc.js} +1 -1
  10. package/dist/assets/{ModelCapabilityIcons.vue_vue_type_script_setup_true_lang-DkuvrY_5.js → ModelCapabilityIcons.vue_vue_type_script_setup_true_lang-CV939MUZ.js} +1 -1
  11. package/dist/assets/{RAGPage-BdXerRDW.js → RAGPage-DbZxWHmi.js} +1 -1
  12. package/dist/assets/SessionPage-DvtFE4GS.js +15 -0
  13. package/dist/assets/{SettingsPage-B1wm-L6T.js → SettingsPage-fmgr8zF9.js} +1 -1
  14. package/dist/assets/{StatusBadge.vue_vue_type_script_setup_true_lang-J8rYaobd.js → StatusBadge.vue_vue_type_script_setup_true_lang-BbHpA7hy.js} +1 -1
  15. package/dist/assets/{StorageGauge.vue_vue_type_style_index_0_lang-Begwajgo.js → StorageGauge.vue_vue_type_style_index_0_lang-Dd3BsCg6.js} +1 -1
  16. package/dist/assets/{SystemPage-DDZrsYds.js → SystemPage-Bd3SBsWI.js} +2 -2
  17. package/dist/assets/{TabBar.vue_vue_type_script_setup_true_lang-DVCeD1qL.js → TabBar.vue_vue_type_script_setup_true_lang-DlSEPHs6.js} +1 -1
  18. package/dist/assets/{TokenStreamDisplay.vue_vue_type_script_setup_true_lang-rJpx5sU7.js → TokenStreamDisplay.vue_vue_type_script_setup_true_lang-P-I0kJTd.js} +1 -1
  19. package/dist/assets/{ToolWorkshopPage-DnHp32vX.js → ToolWorkshopPage-D9doB8PF.js} +1 -1
  20. package/dist/assets/{agent-store-BWO6HamI.js → agent-store-CuCtjbhV.js} +1 -1
  21. package/dist/assets/{canvas-ai-db-anWOo2F0.js → canvas-ai-db-DBUT-nyD.js} +1 -1
  22. package/dist/assets/{download-BqPPd3_X.js → download-C3s6K1NW.js} +1 -1
  23. package/dist/assets/{generate-service-C3Z8NZaT.js → generate-service-DrRYe_-5.js} +1 -1
  24. package/dist/assets/{google-auth-store-CgUWvtRU.js → google-auth-store-B0WvuFWG.js} +1 -1
  25. package/dist/assets/index-Bawj8X2v.css +1 -0
  26. package/dist/assets/{index-KhnkHbiC.js → index-CUNpOQX9.js} +1 -1
  27. package/dist/assets/index-DYHBy1xj.js +8 -0
  28. package/dist/assets/{index-B7GdQF1S.js → index-DoMyvAkM.js} +1 -1
  29. package/dist/assets/info-D-oZIv7P.js +1 -0
  30. package/dist/assets/{metrics-store-BMhkZ9B4.js → metrics-store-DRErjO4D.js} +1 -1
  31. package/dist/assets/{papaparse.min-z5sNvmdo.js → papaparse.min-DJzuA_ek.js} +1 -1
  32. package/dist/assets/{pencil-Csf0-tn3.js → pencil-MO2uWQbW.js} +1 -1
  33. package/dist/assets/{rag-store-Ch0N5LTe.js → rag-store-CBWvD3Va.js} +3 -3
  34. package/dist/assets/{share-2-DIi3V2dc.js → share-2-CnM5sBDm.js} +1 -1
  35. package/dist/assets/{storage-store-uPTARmbg.js → storage-store-cx_2Hb_J.js} +1 -1
  36. package/dist/assets/{stream-handler-BuwD741F.js → stream-handler-CXdBM2li.js} +1 -1
  37. package/dist/assets/{tool-workshop-store-koPpdoxb.js → tool-workshop-store-Cb4eb3e4.js} +1 -1
  38. package/dist/assets/{toolcall-store-JKmpIFj3.js → toolcall-store-xQEQUq5f.js} +1 -1
  39. package/dist/assets/{trash-2-aeVGvgIz.js → trash-2-Cj-NOsVZ.js} +1 -1
  40. package/dist/index.html +2 -2
  41. package/package.json +1 -1
  42. package/dist/assets/DashboardPage-CH15t_MS.js +0 -116
  43. package/dist/assets/SessionPage-B39x2A8D.js +0 -15
  44. package/dist/assets/index-ClE7gpZb.js +0 -8
  45. package/dist/assets/index-Dr1ztyT3.css +0 -1
package/README.md CHANGED
@@ -1,244 +1,265 @@
1
- <p align="center">
2
- <img src="https://raw.githubusercontent.com/LogneBudo/llmxray/master/public/favicon.svg" alt="LLMxRay" width="80" />
3
- </p>
4
-
5
- <h1 align="center">LLMxRay</h1>
6
- <p align="center"><strong>Local LLM Observatory</strong></p>
7
- <p align="center">
8
- See what your AI is <em>actually</em> doing &mdash; token by token, layer by layer.
9
- </p>
10
-
11
- <p align="center">
12
- <img src="https://img.shields.io/badge/vue-3.5-42b883?logo=vuedotjs&logoColor=white" alt="Vue 3.5" />
13
- <img src="https://img.shields.io/badge/vite-7.3-646cff?logo=vite&logoColor=white" alt="Vite 7.3" />
14
- <img src="https://img.shields.io/badge/typescript-5.9-3178c6?logo=typescript&logoColor=white" alt="TypeScript 5.9" />
15
- <img src="https://img.shields.io/badge/tailwind-4.2-06b6d4?logo=tailwindcss&logoColor=white" alt="Tailwind 4.2" />
16
- <img src="https://img.shields.io/badge/ollama-local-000?logo=ollama&logoColor=white" alt="Ollama" />
17
- </p>
18
-
19
- ---
20
-
21
- ## What is LLMxRay?
22
-
23
- LLMxRay is a **free, local-first** dashboard that connects to [Ollama](https://ollama.com) running on your machine. It lets you chat with any model you've downloaded and then **inspect everything that happened behind the scenes**: how fast each token arrived, what the model might have been "thinking", how different settings change the output, and much more.
24
-
25
- **No cloud. No API keys. No cost.** Everything runs on your hardware.
26
-
27
- ### Who is this for?
28
-
29
- | You are... | LLMxRay helps you... |
30
- |---|---|
31
- | **Curious beginner** | See AI responses form in real-time and learn what "temperature" or "tokens" actually mean |
32
- | **Student / educator** | Explore model behavior visually &mdash; great for AI/ML coursework and demos |
33
- | **Developer** | Debug prompts, compare models, profile latency, inspect tool calls |
34
- | **Researcher** | Run controlled experiments: same prompt, different settings, side-by-side results |
35
-
36
- ---
37
-
38
- ## Features at a Glance
39
-
40
- ### Chat with Real-Time Token Streaming
41
- Start a conversation with any Ollama model. Watch tokens appear one by one with **confidence coloring** &mdash; each token is tinted based on how quickly the model produced it (faster = more confident). Supports markdown rendering, multi-turn conversation, file attachments, and slash commands.
42
-
43
- ### Session Deep Dive
44
- Click any past session to explore six tabs of detail:
45
-
46
- - **Stream** &mdash; Every token with timing data, plus a metrics dashboard (time-to-first-token, tokens/sec, latency chart)
47
- - **Reasoning** &mdash; If you're running a reasoning model like DeepSeek-R1, the `<think>` blocks are parsed and displayed step-by-step
48
- - **Introspection** &mdash; Visualizations of layer activations, attention heatmaps, and model architecture (illustrative)
49
- - **Tools** &mdash; Timeline of any tool calls the model made, with parameters and results
50
- - **Agent** &mdash; State-flow graph showing how an agent-style prompt progressed
51
- - **Prompt** &mdash; Anatomy breakdown of your prompt: sections, token counts, structure
52
-
53
- ### Compare Models (and Settings)
54
- The comparison workbench goes beyond "Model A vs Model B". Create up to **4 slots**, each with its own model, temperature, system prompt, and sampling parameters. Compare the *same* model at different temperatures to see how creativity changes. Features include:
55
-
56
- - **Grid view** &mdash; Side-by-side streaming results with per-slot settings pills
57
- - **Diff view** &mdash; Word-level highlighting of what changed between two outputs
58
- - **Metrics bar** &mdash; Visual comparison of TTFT, tokens/sec, and total tokens
59
- - **Quick presets** &mdash; "Temperature Sweep" (3 temps) and "Deterministic Pair" (same seed) one-click setups
60
- - Embedding models are automatically filtered out &mdash; only chat-capable models appear
61
-
62
- ### Embeddings Lab
63
- Embed any text and visualize the resulting vector. Compare two texts with a **cosine similarity meter** to see how semantically close they are. A hands-on way to understand what embeddings actually represent.
64
-
65
- ### RAG Pipeline
66
- Build a local knowledge base from your documents:
67
-
68
- 1. **Upload** PDFs, Word docs (.docx), or CSVs
69
- 2. **Chunk & embed** automatically using your chosen embedding model
70
- 3. **Search** with natural language &mdash; results ranked by semantic similarity
71
-
72
- Everything is stored in **IndexedDB** (your browser's built-in database). Zero cost, zero setup, zero external services.
73
-
74
- ### Tool Workshop (Visual Canvas)
75
- Build, edit, and test tool definitions on an interactive **node-based canvas** powered by Vue Flow:
76
-
77
- - **Drag-and-drop nodes** &mdash; Each tool is a visual node showing name, description, parameters, and implementation body
78
- - **Inline code editing** &mdash; Full CodeMirror 6 editors with TypeScript syntax highlighting directly on each node
79
- - **Bidirectional code sync** &mdash; Open the Code Panel to see all tools as combined TypeScript source. Edit code, nodes update. Edit nodes, code updates. Powered by a Recast AST parser
80
- - **Schema viewer** &mdash; Auto-generated OpenAI-compatible JSON schemas with one-click copy
81
- - **Probe & Pick** &mdash; Point at any API URL, inspect the response JSON tree, and auto-generate fetch code + parameter mappings
82
- - **OpenAPI discovery** &mdash; Auto-detect and parse OpenAPI/Swagger specs to pick endpoints visually
83
- - **Live execution overlays** &mdash; During chat, tool nodes pulse when the model calls them and show results inline
84
- - **Templates** &mdash; Start from 15+ built-in templates (web fetch, calculator, Google Calendar/Gmail, regex tester, and more)
85
- - **Persistent layout** &mdash; Node positions, mappings, and probe configs survive across sessions
86
-
87
- ### Tool Call Optimizer
88
- When the model calls a tool during chat, an **"Optimize this Tool"** button appears on the result. Click it to open the Response Optimizer Drawer:
89
-
90
- - Visualize the API response as an interactive JSON tree
91
- - Select only the fields the model actually needs
92
- - Auto-generate optimized fetch code with field extraction
93
- - One click to create a new optimized tool in the Workshop
94
-
95
- ### Model Browser
96
- See every model installed in Ollama with details like parameter count, quantization level, family, and format. Includes architecture diagrams showing the model's structure.
97
-
98
- ### System Monitor
99
- Real hardware specs (not browser estimates) &mdash; CPU model, total RAM with live usage, GPU with driver version, storage. Plus live Ollama status: running models, memory allocation, inference settings.
100
-
101
- ### Settings
102
- Configure your Ollama connection URL with a live connection tester. Set default **temperature** and **context length** with visual scales and educational tooltips that explain what each setting does in plain language.
103
-
104
- ---
105
-
106
- ## Quick Start
107
-
108
- ### Prerequisites
109
-
110
- 1. **Node.js 18+** &mdash; [Download](https://nodejs.org)
111
- 2. **Ollama** running locally &mdash; [Download](https://ollama.com/download)
112
- 3. At least one model pulled:
113
- ```bash
114
- ollama pull llama3.2
115
- ```
116
-
117
- ### Install and Run
118
-
119
- ```bash
120
- git clone https://github.com/LogneBudo/llmxray.git
121
- cd llmxray
122
- npm install
123
- npm run dev
124
- ```
125
-
126
- Open **http://localhost:5173** in your browser. That's it.
127
-
128
- > LLMxRay's dev server automatically proxies API calls to Ollama at `localhost:11434`. If Ollama is running on a different port or machine, change it in **Settings**.
129
-
130
- ### Build for Production
131
-
132
- ```bash
133
- npm run build # Output in dist/
134
- npm run preview # Preview the build locally
135
- ```
136
-
137
- ---
138
-
139
- ## Tech Stack
140
-
141
- | Layer | Technology |
142
- |---|---|
143
- | Framework | Vue 3.5 + Composition API (`<script setup>`) |
144
- | Language | TypeScript 5.9 (strict) |
145
- | Build | Vite 7.3 |
146
- | Styling | Tailwind CSS 4.2 (custom dark theme) |
147
- | State | Pinia 3 (one store per concern) |
148
- | Routing | Vue Router 5 |
149
- | Charts | Chart.js 4 + vue-chartjs, D3.js 7 |
150
- | Canvas | Vue Flow 1.x (node-based visual editor) |
151
- | Code Editor | CodeMirror 6 (TypeScript + JSON highlighting) |
152
- | AST Parser | Recast + @babel/parser (bidirectional code sync) |
153
- | Markdown | marked |
154
- | Diffing | diff (word-level) |
155
- | Documents | pdfjs-dist (lazy), mammoth (DOCX), papaparse (CSV) |
156
- | Storage | IndexedDB (browser-native, zero-cost) |
157
- | IDs | nanoid |
158
- | LLM Backend | Ollama (local, via `/api` proxy) |
159
-
160
- ---
161
-
162
- ## Architecture Highlights
163
-
164
- **Streaming** &mdash; LLMxRay reads Ollama's NDJSON response streams via `fetch()` + `ReadableStream`. Tokens arrive one by one and update the UI reactively through Pinia stores.
165
-
166
- **Token confidence** &mdash; Ollama doesn't expose logprobs, so confidence is approximated from inter-token latency. Faster tokens = higher confidence. This is labeled clearly in the UI as an approximation.
167
-
168
- **Introspection data** &mdash; Layer activations and attention heatmaps are synthetic (illustrative). They demonstrate what these visualizations *would* look like with real data. Clearly labeled as "Illustrative" in the UI.
169
-
170
- **Store-per-concern** &mdash; Each domain has its own Pinia store: tokens, sessions, metrics, reasoning, comparison, embeddings, RAG, models, and more. This keeps state management modular and testable.
171
-
172
- **Hardware detection** &mdash; The System page uses a custom Vite plugin (`vite-plugin-system-info.ts`) that queries the OS directly via PowerShell (Windows), `/proc` + `lspci` (Linux), or `sysctl` (macOS) for accurate hardware specs.
173
-
174
- ---
175
-
176
- ## Development
177
-
178
- ### Scripts
179
-
180
- | Command | What it does |
181
- |---|---|
182
- | `npm run dev` | Start dev server (port 5173) |
183
- | `npm run build` | Type-check + production build |
184
- | `npm run preview` | Preview production build |
185
- | `npm run test` | Run unit tests (Vitest) |
186
- | `npm run test:watch` | Tests in watch mode |
187
- | `npm run test:coverage` | Coverage report |
188
- | `npm run test:e2e` | Playwright end-to-end tests |
189
- | `npm run test:e2e:headed` | E2E with visible browser |
190
- | `npm run test:e2e:live` | E2E against live Ollama |
191
-
192
- ### Project Structure
193
-
194
- ```
195
- src/
196
- pages/ 8 page components (Dashboard, Compare, RAG, etc.)
197
- components/ 50+ components organized by feature
198
- chat/ Chat UI, token stream, attachments
199
- comparison/ Slot configurator, grid, diff view, metrics bar
200
- metrics/ Dashboard, charts, session history
201
- reasoning/ Think-block viewer
202
- introspection/ Layer activations, attention, architecture
203
- rag/ Document upload, search, ingest
204
- embeddings/ Vector viz, similarity meter
205
- tool-canvas/ Visual canvas, node editor, CodeMirror wrapper
206
- tool-optimizer/ Response optimizer drawer, JSON tree
207
- tool-calls/ Tool call timeline, definitions
208
- agent-graph/ Agent state flow
209
- common/ Layout, sidebar, shared components
210
- stores/ Pinia stores (one per concern)
211
- services/ Ollama client, streaming, generation, RAG, AST parser, probe
212
- types/ TypeScript interfaces
213
- utils/ Formatting, color scales, slot labels
214
- composables/ Vue composables (markdown, etc.)
215
- router/ Route definitions
216
- ```
217
-
218
- ---
219
-
220
- ## Troubleshooting
221
-
222
- | Problem | Solution |
223
- |---|---|
224
- | "Disconnected" in Settings | Make sure Ollama is running: `ollama serve` |
225
- | No models in dropdowns | Pull a model first: `ollama pull llama3.2` |
226
- | System page shows "Restart dev server" | Stop and restart `npm run dev` (the hardware plugin loads at startup) |
227
- | Slow first response | Normal &mdash; Ollama loads the model into memory on first use |
228
- | High RAM/VRAM usage | Use smaller quantized models (Q4) or reduce context length in Settings |
229
-
230
- ---
231
-
232
- ## License
233
-
234
- Licensed under the [Apache License 2.0](LICENSE). You are free to use, modify, and distribute this software under the terms of the license.
235
-
236
- ## Trademark
237
-
238
- **LLMxRay** is a trademark of Ivan Stankovic ([LogneBudo](https://github.com/LogneBudo)). See [TRADEMARK.md](TRADEMARK.md) for usage guidelines.
239
-
240
- ---
241
-
242
- <p align="center">
243
- Built with curiosity by <a href="https://github.com/LogneBudo">LogneBudo</a>
244
- </p>
1
+ <p align="center">
2
+ <img src="https://raw.githubusercontent.com/LogneBudo/llmxray/master/public/favicon.svg" alt="LLMxRay" width="80" />
3
+ </p>
4
+
5
+ <h1 align="center">LLMxRay</h1>
6
+ <p align="center"><strong>Local LLM Observatory</strong></p>
7
+ <p align="center">
8
+ See what your AI is <em>actually</em> doing &mdash; token by token, layer by layer.
9
+ </p>
10
+
11
+ <p align="center">
12
+ <img src="https://img.shields.io/badge/vue-3.5-42b883?logo=vuedotjs&logoColor=white" alt="Vue 3.5" />
13
+ <img src="https://img.shields.io/badge/vite-7.3-646cff?logo=vite&logoColor=white" alt="Vite 7.3" />
14
+ <img src="https://img.shields.io/badge/typescript-5.9-3178c6?logo=typescript&logoColor=white" alt="TypeScript 5.9" />
15
+ <img src="https://img.shields.io/badge/tailwind-4.2-06b6d4?logo=tailwindcss&logoColor=white" alt="Tailwind 4.2" />
16
+ <img src="https://img.shields.io/badge/ollama-local-000?logo=ollama&logoColor=white" alt="Ollama" />
17
+ </p>
18
+
19
+ ---
20
+
21
+ ## What is LLMxRay?
22
+
23
+ LLMxRay is a **free, local-first** dashboard that connects to [Ollama](https://ollama.com) running on your machine. It lets you chat with any model you've downloaded and then **inspect everything that happened behind the scenes**: how fast each token arrived, what the model might have been "thinking", how different settings change the output, and much more.
24
+
25
+ **No cloud. No API keys. No cost.** Everything runs on your hardware.
26
+
27
+ ### Who is this for?
28
+
29
+ | You are... | LLMxRay helps you... |
30
+ |---|---|
31
+ | **Curious beginner** | See AI responses form in real-time and learn what "temperature" or "tokens" actually mean |
32
+ | **Student / educator** | Explore model behavior visually &mdash; great for AI/ML coursework and demos |
33
+ | **Developer** | Debug prompts, compare models, profile latency, inspect tool calls |
34
+ | **Researcher** | Run controlled experiments: same prompt, different settings, side-by-side results |
35
+
36
+ ---
37
+
38
+ ## Features at a Glance
39
+
40
+ ### Chat with Real-Time Token Streaming
41
+ Start a conversation with any Ollama model. Watch tokens appear one by one with **confidence coloring** &mdash; each token is tinted based on how quickly the model produced it (faster = more confident). Supports markdown rendering, multi-turn conversation, file attachments, and slash commands.
42
+
43
+ ### Session Deep Dive
44
+ Click any past session to explore six tabs of detail:
45
+
46
+ - **Stream** &mdash; Every token with timing data, plus a metrics dashboard (time-to-first-token, tokens/sec, latency chart)
47
+ - **Reasoning** &mdash; If you're running a reasoning model like DeepSeek-R1, the `<think>` blocks are parsed and displayed step-by-step
48
+ - **Introspection** &mdash; Visualizations of layer activations, attention heatmaps, and model architecture (illustrative)
49
+ - **Tools** &mdash; Timeline of any tool calls the model made, with parameters and results
50
+ - **Agent** &mdash; State-flow graph showing how an agent-style prompt progressed
51
+ - **Prompt** &mdash; Anatomy breakdown of your prompt: sections, token counts, structure
52
+
53
+ ### Compare Models (and Settings)
54
+ The comparison workbench goes beyond "Model A vs Model B". Create up to **4 slots**, each with its own model, temperature, system prompt, and sampling parameters. Compare the *same* model at different temperatures to see how creativity changes. Features include:
55
+
56
+ - **Grid view** &mdash; Side-by-side streaming results with per-slot settings pills
57
+ - **Diff view** &mdash; Word-level highlighting of what changed between two outputs
58
+ - **Metrics bar** &mdash; Visual comparison of TTFT, tokens/sec, and total tokens
59
+ - **Quick presets** &mdash; "Temperature Sweep" (3 temps) and "Deterministic Pair" (same seed) one-click setups
60
+ - Embedding models are automatically filtered out &mdash; only chat-capable models appear
61
+
62
+ ### Embeddings Lab
63
+ Embed any text and visualize the resulting vector. Compare two texts with a **cosine similarity meter** to see how semantically close they are. A hands-on way to understand what embeddings actually represent.
64
+
65
+ ### RAG Pipeline
66
+ Build a local knowledge base from your documents:
67
+
68
+ 1. **Upload** PDFs, Word docs (.docx), or CSVs
69
+ 2. **Chunk & embed** automatically using your chosen embedding model
70
+ 3. **Search** with natural language &mdash; results ranked by semantic similarity
71
+
72
+ Everything is stored in **IndexedDB** (your browser's built-in database). Zero cost, zero setup, zero external services.
73
+
74
+ ### Tool Workshop (Visual Canvas)
75
+ Build, edit, and test tool definitions on an interactive **node-based canvas** powered by Vue Flow:
76
+
77
+ - **Drag-and-drop nodes** &mdash; Each tool is a visual node showing name, description, parameters, and implementation body
78
+ - **Inline code editing** &mdash; Full CodeMirror 6 editors with TypeScript syntax highlighting directly on each node
79
+ - **Bidirectional code sync** &mdash; Open the Code Panel to see all tools as combined TypeScript source. Edit code, nodes update. Edit nodes, code updates. Powered by a Recast AST parser
80
+ - **Schema viewer** &mdash; Auto-generated OpenAI-compatible JSON schemas with one-click copy
81
+ - **Probe & Pick** &mdash; Point at any API URL, inspect the response JSON tree, and auto-generate fetch code + parameter mappings
82
+ - **OpenAPI discovery** &mdash; Auto-detect and parse OpenAPI/Swagger specs to pick endpoints visually
83
+ - **Live execution overlays** &mdash; During chat, tool nodes pulse when the model calls them and show results inline
84
+ - **Templates** &mdash; Start from 15+ built-in templates (web fetch, calculator, Google Calendar/Gmail, regex tester, and more)
85
+ - **Persistent layout** &mdash; Node positions, mappings, and probe configs survive across sessions
86
+
87
+ ### Tool Call Optimizer
88
+ When the model calls a tool during chat, an **"Optimize this Tool"** button appears on the result. Click it to open the Response Optimizer Drawer:
89
+
90
+ - Visualize the API response as an interactive JSON tree
91
+ - Select only the fields the model actually needs
92
+ - Auto-generate optimized fetch code with field extraction
93
+ - One click to create a new optimized tool in the Workshop
94
+
95
+ ### Response Quality Gates
96
+ Every assistant response is automatically analyzed for common quality issues. Small colored badges appear below the response metrics when problems are detected:
97
+
98
+ - **Repetition** &mdash; Flags responses with excessive repeated 4-gram phrases (>50% = fail, >30% = warn)
99
+ - **Refusal** &mdash; Detects 8 common refusal patterns ("as an AI language model", "I cannot help", etc.)
100
+ - **Gibberish** &mdash; Warns when non-ASCII characters exceed 40% of the response
101
+ - **Empty** &mdash; Flags responses with fewer than 10 words
102
+ - **Truncation** &mdash; Warns when the response hit the token limit or used >90% of budget without clean ending
103
+
104
+ No news is good news &mdash; badges only appear when something is wrong.
105
+
106
+ ### Cost Dashboard
107
+ Track token usage across all your sessions with estimated cloud-equivalent costs. Navigate to the **Costs** page in the sidebar to see:
108
+
109
+ - **Summary cards** &mdash; Total tokens, sessions, estimated cost, average cost per session
110
+ - **Token usage by model** &mdash; Stacked bar chart showing prompt vs completion tokens per model
111
+ - **Daily usage trends** &mdash; Line chart with dual axes (tokens + estimated cost over time)
112
+ - **Model breakdown table** &mdash; Detailed per-model statistics with pricing source transparency
113
+
114
+ Costs are estimates based on equivalent cloud API pricing (Groq, Together AI, Google, Mistral, etc.). Ollama runs locally at zero cost &mdash; the dashboard shows what you're saving.
115
+
116
+ ### Model Browser
117
+ See every model installed in Ollama with details like parameter count, quantization level, family, and format. Includes architecture diagrams showing the model's structure.
118
+
119
+ ### System Monitor
120
+ Real hardware specs (not browser estimates) &mdash; CPU model, total RAM with live usage, GPU with driver version, storage. Plus live Ollama status: running models, memory allocation, inference settings.
121
+
122
+ ### Settings
123
+ Configure your Ollama connection URL with a live connection tester. Set default **temperature** and **context length** with visual scales and educational tooltips that explain what each setting does in plain language.
124
+
125
+ ---
126
+
127
+ ## Quick Start
128
+
129
+ ### Prerequisites
130
+
131
+ 1. **Node.js 18+** &mdash; [Download](https://nodejs.org)
132
+ 2. **Ollama** running locally &mdash; [Download](https://ollama.com/download)
133
+ 3. At least one model pulled:
134
+ ```bash
135
+ ollama pull llama3.2
136
+ ```
137
+
138
+ ### Install and Run
139
+
140
+ ```bash
141
+ git clone https://github.com/LogneBudo/llmxray.git
142
+ cd llmxray
143
+ npm install
144
+ npm run dev
145
+ ```
146
+
147
+ Open **http://localhost:5173** in your browser. That's it.
148
+
149
+ > LLMxRay's dev server automatically proxies API calls to Ollama at `localhost:11434`. If Ollama is running on a different port or machine, change it in **Settings**.
150
+
151
+ ### Build for Production
152
+
153
+ ```bash
154
+ npm run build # Output in dist/
155
+ npm run preview # Preview the build locally
156
+ ```
157
+
158
+ ---
159
+
160
+ ## Tech Stack
161
+
162
+ | Layer | Technology |
163
+ |---|---|
164
+ | Framework | Vue 3.5 + Composition API (`<script setup>`) |
165
+ | Language | TypeScript 5.9 (strict) |
166
+ | Build | Vite 7.3 |
167
+ | Styling | Tailwind CSS 4.2 (custom dark theme) |
168
+ | State | Pinia 3 (one store per concern) |
169
+ | Routing | Vue Router 5 |
170
+ | Charts | Chart.js 4 + vue-chartjs, D3.js 7 |
171
+ | Canvas | Vue Flow 1.x (node-based visual editor) |
172
+ | Code Editor | CodeMirror 6 (TypeScript + JSON highlighting) |
173
+ | AST Parser | Recast + @babel/parser (bidirectional code sync) |
174
+ | Markdown | marked |
175
+ | Diffing | diff (word-level) |
176
+ | Documents | pdfjs-dist (lazy), mammoth (DOCX), papaparse (CSV) |
177
+ | Storage | IndexedDB (browser-native, zero-cost) |
178
+ | IDs | nanoid |
179
+ | LLM Backend | Ollama (local, via `/api` proxy) |
180
+
181
+ ---
182
+
183
+ ## Architecture Highlights
184
+
185
+ **Streaming** &mdash; LLMxRay reads Ollama's NDJSON response streams via `fetch()` + `ReadableStream`. Tokens arrive one by one and update the UI reactively through Pinia stores.
186
+
187
+ **Token confidence** &mdash; Ollama doesn't expose logprobs, so confidence is approximated from inter-token latency. Faster tokens = higher confidence. This is labeled clearly in the UI as an approximation.
188
+
189
+ **Introspection data** &mdash; Layer activations and attention heatmaps are synthetic (illustrative). They demonstrate what these visualizations *would* look like with real data. Clearly labeled as "Illustrative" in the UI.
190
+
191
+ **Store-per-concern** &mdash; Each domain has its own Pinia store: tokens, sessions, metrics, reasoning, comparison, embeddings, RAG, models, and more. This keeps state management modular and testable.
192
+
193
+ **Hardware detection** &mdash; The System page uses a custom Vite plugin (`vite-plugin-system-info.ts`) that queries the OS directly via PowerShell (Windows), `/proc` + `lspci` (Linux), or `sysctl` (macOS) for accurate hardware specs.
194
+
195
+ ---
196
+
197
+ ## Development
198
+
199
+ ### Scripts
200
+
201
+ | Command | What it does |
202
+ |---|---|
203
+ | `npm run dev` | Start dev server (port 5173) |
204
+ | `npm run build` | Type-check + production build |
205
+ | `npm run preview` | Preview production build |
206
+ | `npm run test` | Run unit tests (Vitest) |
207
+ | `npm run test:watch` | Tests in watch mode |
208
+ | `npm run test:coverage` | Coverage report |
209
+ | `npm run test:e2e` | Playwright end-to-end tests |
210
+ | `npm run test:e2e:headed` | E2E with visible browser |
211
+ | `npm run test:e2e:live` | E2E against live Ollama |
212
+
213
+ ### Project Structure
214
+
215
+ ```
216
+ src/
217
+ pages/ 8 page components (Dashboard, Compare, RAG, etc.)
218
+ components/ 50+ components organized by feature
219
+ chat/ Chat UI, token stream, attachments
220
+ comparison/ Slot configurator, grid, diff view, metrics bar
221
+ metrics/ Dashboard, charts, session history
222
+ reasoning/ Think-block viewer
223
+ introspection/ Layer activations, attention, architecture
224
+ rag/ Document upload, search, ingest
225
+ embeddings/ Vector viz, similarity meter
226
+ tool-canvas/ Visual canvas, node editor, CodeMirror wrapper
227
+ tool-optimizer/ Response optimizer drawer, JSON tree
228
+ tool-calls/ Tool call timeline, definitions
229
+ agent-graph/ Agent state flow
230
+ common/ Layout, sidebar, shared components
231
+ stores/ Pinia stores (one per concern)
232
+ services/ Ollama client, streaming, generation, RAG, AST parser, probe
233
+ types/ TypeScript interfaces
234
+ utils/ Formatting, color scales, slot labels
235
+ composables/ Vue composables (markdown, etc.)
236
+ router/ Route definitions
237
+ ```
238
+
239
+ ---
240
+
241
+ ## Troubleshooting
242
+
243
+ | Problem | Solution |
244
+ |---|---|
245
+ | "Disconnected" in Settings | Make sure Ollama is running: `ollama serve` |
246
+ | No models in dropdowns | Pull a model first: `ollama pull llama3.2` |
247
+ | System page shows "Restart dev server" | Stop and restart `npm run dev` (the hardware plugin loads at startup) |
248
+ | Slow first response | Normal &mdash; Ollama loads the model into memory on first use |
249
+ | High RAM/VRAM usage | Use smaller quantized models (Q4) or reduce context length in Settings |
250
+
251
+ ---
252
+
253
+ ## License
254
+
255
+ Licensed under the [Apache License 2.0](LICENSE). You are free to use, modify, and distribute this software under the terms of the license.
256
+
257
+ ## Trademark
258
+
259
+ **LLMxRay** is a trademark of Ivan Stankovic ([LogneBudo](https://github.com/LogneBudo)). See [TRADEMARK.md](TRADEMARK.md) for usage guidelines.
260
+
261
+ ---
262
+
263
+ <p align="center">
264
+ Built with curiosity by <a href="https://github.com/LogneBudo">LogneBudo</a>
265
+ </p>