llmxray 0.4.2 → 0.4.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (192) hide show
  1. package/README.md +265 -276
  2. package/dist/assets/{AITrainingPage-DEU4UAR7.js → AITrainingPage-GE_hvc8i.js} +1 -1
  3. package/dist/assets/{AnalyticsPage-C7mNUAil.js → AnalyticsPage-CUpkcOcS.js} +1 -1
  4. package/dist/assets/{BenchmarkPage-jiUqdHmg.js → BenchmarkPage-DQuv9EdH.js} +2 -2
  5. package/dist/assets/{ComparisonPage-BoBWSObY.js → ComparisonPage-DXh2a8TJ.js} +3 -3
  6. package/dist/assets/{CostDashboardPage-BKQDsZhx.js → CostDashboardPage-C4IVmHn1.js} +1 -1
  7. package/dist/assets/{DashboardPage-CV8Ejgxv.js → DashboardPage-DhvUf3u6.js} +6 -6
  8. package/dist/assets/{EmbeddingsPage-DXjVT4eU.js → EmbeddingsPage-tMH5Unz8.js} +1 -1
  9. package/dist/assets/GoogleCallbackPage-Bljvoj2X.js +1 -0
  10. package/dist/assets/{JsonTreeNode-C0UNUwZm.js → JsonTreeNode-srBpwnIe.js} +1 -1
  11. package/dist/assets/{ModelCapabilityIcons.vue_vue_type_script_setup_true_lang-vQNvYFzV.js → ModelCapabilityIcons.vue_vue_type_script_setup_true_lang-CBo5u7cq.js} +1 -1
  12. package/dist/assets/{RAGPage-Dv4UDEyz.js → RAGPage-Br6DZPuU.js} +1 -1
  13. package/dist/assets/{SessionPage-B0fMQGSB.js → SessionPage-CVpgkmrd.js} +2 -2
  14. package/dist/assets/SettingsPage-DY1zfkza.js +1 -0
  15. package/dist/assets/{StatusBadge.vue_vue_type_script_setup_true_lang-BiB0M6lv.js → StatusBadge.vue_vue_type_script_setup_true_lang-PDuG6Cj4.js} +1 -1
  16. package/dist/assets/{StorageGauge.vue_vue_type_style_index_0_lang-DToi3R0q.js → StorageGauge.vue_vue_type_style_index_0_lang-OYUc8hU5.js} +1 -1
  17. package/dist/assets/{SystemPage-BPK15qwe.js → SystemPage-D-tjkE5U.js} +3 -3
  18. package/dist/assets/{TabBar.vue_vue_type_script_setup_true_lang-BtGCdMZK.js → TabBar.vue_vue_type_script_setup_true_lang-DvubG_tm.js} +1 -1
  19. package/dist/assets/{TokenStreamDisplay.vue_vue_type_script_setup_true_lang-B6juPZsY.js → TokenStreamDisplay.vue_vue_type_script_setup_true_lang-GmRAK2xe.js} +1 -1
  20. package/dist/assets/{ToolWorkshopPage-DmpoHHaD.js → ToolWorkshopPage-ChG2v784.js} +2 -2
  21. package/dist/assets/{ToolWorkshopPage-7TRbWVVL.css → ToolWorkshopPage-Duf-VZbs.css} +1 -1
  22. package/dist/assets/ad-B18i8NGa.svg +150 -0
  23. package/dist/assets/ad-Blhdm5jl.svg +148 -0
  24. package/dist/assets/af-Bc2fqp73.svg +81 -0
  25. package/dist/assets/af-C77Rf6cE.svg +81 -0
  26. package/dist/assets/{agent-store-C6w1zV86.js → agent-store-qLxji-w9.js} +1 -1
  27. package/dist/assets/arab-C-KgnQEz.svg +109 -0
  28. package/dist/assets/arab-C4CYPgyC.svg +109 -0
  29. package/dist/assets/as-BTEVCXG-.svg +73 -0
  30. package/dist/assets/as-Dekqy8Of.svg +72 -0
  31. package/dist/assets/aw-CLCX8uk5.svg +186 -0
  32. package/dist/assets/aw-W0PWLK5p.svg +186 -0
  33. package/dist/assets/bm-BeYgB2z9.svg +97 -0
  34. package/dist/assets/bm-DvNWWcPM.svg +97 -0
  35. package/dist/assets/bn-B6T3O78g.svg +36 -0
  36. package/dist/assets/bn-CPQcA8Ol.svg +36 -0
  37. package/dist/assets/bo-CcUiMqkJ.svg +673 -0
  38. package/dist/assets/bo-Dry0C6UA.svg +674 -0
  39. package/dist/assets/br-Cu5YU29T.svg +45 -0
  40. package/dist/assets/br-Dr5rMAMb.svg +45 -0
  41. package/dist/assets/bt-BTo4qm10.svg +89 -0
  42. package/dist/assets/bt-SxWnbWW0.svg +89 -0
  43. package/dist/assets/bz-BCKHR4_q.svg +145 -0
  44. package/dist/assets/bz-CoBdB-p8.svg +145 -0
  45. package/dist/assets/{canvas-ai-db-BpqhUp6k.js → canvas-ai-db-DIu3Cyqr.js} +1 -1
  46. package/dist/assets/cy-DJKnEFYW.svg +6 -0
  47. package/dist/assets/cy-bZuP8hmf.svg +6 -0
  48. package/dist/assets/dg-CJPJrjiZ.svg +130 -0
  49. package/dist/assets/dg-DqkWLbnk.svg +130 -0
  50. package/dist/assets/dm-Cbhezfe1.svg +152 -0
  51. package/dist/assets/dm-DPPHwW2M.svg +152 -0
  52. package/dist/assets/do-B86d445t.svg +121 -0
  53. package/dist/assets/do-DeRnbj4d.svg +122 -0
  54. package/dist/assets/{download-pCIjWOn5.js → download-DBk31Sbj.js} +1 -1
  55. package/dist/assets/eac-CwGQsyAM.svg +48 -0
  56. package/dist/assets/eac-h4QKADRE.svg +48 -0
  57. package/dist/assets/ec-CaVOFQ3t.svg +138 -0
  58. package/dist/assets/ec-cwfBJlvF.svg +138 -0
  59. package/dist/assets/eg-DwOkwyQ0.svg +38 -0
  60. package/dist/assets/eg-YC70hswZ.svg +38 -0
  61. package/dist/assets/es-BuSGTZm_.svg +547 -0
  62. package/dist/assets/es-d5m8M5h8.svg +544 -0
  63. package/dist/assets/es-ga-D9xG2hYr.svg +187 -0
  64. package/dist/assets/es-ga-DXhVZ333.svg +187 -0
  65. package/dist/assets/fj-DEAVMg38.svg +120 -0
  66. package/dist/assets/fj-u3dAPoew.svg +123 -0
  67. package/dist/assets/fk-B-RvQ4Hz.svg +89 -0
  68. package/dist/assets/fk-nuUF_Ak3.svg +90 -0
  69. package/dist/assets/gb-nir-D4gikpNq.svg +132 -0
  70. package/dist/assets/gb-nir-vEp1ZXy6.svg +131 -0
  71. package/dist/assets/gb-wls-Bxz9hxvX.svg +9 -0
  72. package/dist/assets/gb-wls-CK0XlKT-.svg +9 -0
  73. package/dist/assets/generate-service-BIV4odla.js +3 -0
  74. package/dist/assets/{google-auth-store-Dgv2FihL.js → google-auth-store-Cxj4bIXD.js} +1 -1
  75. package/dist/assets/gq-CPnMO1hT.svg +23 -0
  76. package/dist/assets/gq-Cag8QTk2.svg +23 -0
  77. package/dist/assets/gs-DOgYbHsY.svg +132 -0
  78. package/dist/assets/gs-DiiNa0F5.svg +133 -0
  79. package/dist/assets/gt-BLpn5qMn.svg +204 -0
  80. package/dist/assets/gt-CJo5DI-7.svg +204 -0
  81. package/dist/assets/gu-Di1JYREk.svg +19 -0
  82. package/dist/assets/gu-SbvrH0uZ.svg +19 -0
  83. package/dist/assets/hr-BpiVVBoV.svg +56 -0
  84. package/dist/assets/hr-fzLfaANM.svg +58 -0
  85. package/dist/assets/ht-DIMg4gti.svg +116 -0
  86. package/dist/assets/ht-pweRl6ZP.svg +116 -0
  87. package/dist/assets/im--VPIqfkF.svg +36 -0
  88. package/dist/assets/im-Dd9p-0-T.svg +36 -0
  89. package/dist/assets/index-B0RKZCO8.css +1 -0
  90. package/dist/assets/index-BI1ok0nK.js +8 -0
  91. package/dist/assets/{index-DHuj6tKB.js → index-CIiugJtV.js} +1 -1
  92. package/dist/assets/index-lybAETDi.js +3 -0
  93. package/dist/assets/{info-Eg1tUw1x.js → info-Cxpjao0U.js} +1 -1
  94. package/dist/assets/io-13HOfeJD.svg +130 -0
  95. package/dist/assets/io-BImhNBcd.svg +130 -0
  96. package/dist/assets/ir-Q03Mij62.svg +219 -0
  97. package/dist/assets/ir-cCIgaNf6.svg +219 -0
  98. package/dist/assets/je-DyWbhIiC.svg +62 -0
  99. package/dist/assets/je-vXe0Dr49.svg +62 -0
  100. package/dist/assets/kg-B0FsxZiL.svg +4 -0
  101. package/dist/assets/kg-CjfitMyT.svg +4 -0
  102. package/dist/assets/kh-BBvObpUS.svg +61 -0
  103. package/dist/assets/kh-BeWfuE30.svg +61 -0
  104. package/dist/assets/ki-fuIMkEYQ.svg +36 -0
  105. package/dist/assets/ki-p_fAQGbS.svg +36 -0
  106. package/dist/assets/ky-BqaZHuhf.svg +103 -0
  107. package/dist/assets/ky-Dpsu1myA.svg +103 -0
  108. package/dist/assets/kz-CwKXYZ8s.svg +36 -0
  109. package/dist/assets/kz-Dkyx6q-p.svg +36 -0
  110. package/dist/assets/li-CHdhvNcr.svg +43 -0
  111. package/dist/assets/li-CMlf0YU8.svg +43 -0
  112. package/dist/assets/lk-DSQoDxn_.svg +22 -0
  113. package/dist/assets/lk-DUkgV9Tq.svg +22 -0
  114. package/dist/assets/md-DRlxvNwm.svg +70 -0
  115. package/dist/assets/md-DTi94M3M.svg +71 -0
  116. package/dist/assets/me-CfGorN3b.svg +118 -0
  117. package/dist/assets/me-Cv4Gwqah.svg +116 -0
  118. package/dist/assets/{metrics-store-7m97Sb7w.js → metrics-store-BZ8E63Rc.js} +1 -1
  119. package/dist/assets/mp-CrOApEqW.svg +86 -0
  120. package/dist/assets/mp-CuaQmCLf.svg +86 -0
  121. package/dist/assets/ms-B-w7hFKu.svg +25 -0
  122. package/dist/assets/ms-DxciGbUu.svg +29 -0
  123. package/dist/assets/mt-YDa8zgzO.svg +56 -0
  124. package/dist/assets/mt-YqzKx9xl.svg +58 -0
  125. package/dist/assets/mx-Cc8Ccfe8.svg +382 -0
  126. package/dist/assets/mx-CvCwYHGF.svg +377 -0
  127. package/dist/assets/nf-DGrQb42O.svg +11 -0
  128. package/dist/assets/nf-Dl00mlk2.svg +9 -0
  129. package/dist/assets/ni-BX2WCaNt.svg +129 -0
  130. package/dist/assets/ni-CcFCSQxm.svg +129 -0
  131. package/dist/assets/om-DcqxRdQL.svg +115 -0
  132. package/dist/assets/om-nN8zP2Bu.svg +115 -0
  133. package/dist/assets/{papaparse.min-3FeAYh4d.js → papaparse.min-kBINl9ZG.js} +1 -1
  134. package/dist/assets/{pencil-CZp68oBU.js → pencil-DuvF5FdW.js} +1 -1
  135. package/dist/assets/pn-BPAlH32D.svg +53 -0
  136. package/dist/assets/pn-DgxdtieE.svg +53 -0
  137. package/dist/assets/pt-BTevY6N2.svg +57 -0
  138. package/dist/assets/pt-DZ2ADgIR.svg +57 -0
  139. package/dist/assets/py-BKi5dxWt.svg +156 -0
  140. package/dist/assets/py-mNzh0mZC.svg +157 -0
  141. package/dist/assets/{rag-store-Bm0nxdec.js → rag-store-fluSCOWr.js} +3 -3
  142. package/dist/assets/rs-BfwKwXtn.svg +292 -0
  143. package/dist/assets/rs-CnTO3ehk.svg +296 -0
  144. package/dist/assets/sa-Dh79zbT9.svg +25 -0
  145. package/dist/assets/sa-DnlyVVKx.svg +25 -0
  146. package/dist/assets/{session-store-BQ0TF24S.js → session-store-DQW7koXz.js} +1 -1
  147. package/dist/assets/sh-ac-D-aE2xRW.svg +690 -0
  148. package/dist/assets/sh-ac-FjwY7RYr.svg +689 -0
  149. package/dist/assets/sh-hl-CgxUDvtv.svg +164 -0
  150. package/dist/assets/sh-hl-CqtQPzWZ.svg +164 -0
  151. package/dist/assets/sh-ta-BFo5zkKU.svg +76 -0
  152. package/dist/assets/sh-ta-CPJublpi.svg +76 -0
  153. package/dist/assets/{share-2-CJJGZsZO.js → share-2-NJrUNieP.js} +1 -1
  154. package/dist/assets/sm-BKrUHzrq.svg +73 -0
  155. package/dist/assets/sm-DGBIRFB_.svg +75 -0
  156. package/dist/assets/{storage-store-6qFwk_C5.js → storage-store-BVsJ2m6T.js} +1 -1
  157. package/dist/assets/{stream-handler-BM5egg14.js → stream-handler-BJnVDKi0.js} +1 -1
  158. package/dist/assets/sv-CJIHhYwF.svg +593 -0
  159. package/dist/assets/sv-RZ39q5hO.svg +593 -0
  160. package/dist/assets/sx-RKKs0ph6.svg +56 -0
  161. package/dist/assets/sx-nDhIaDNb.svg +56 -0
  162. package/dist/assets/sz-D39eIL5d.svg +34 -0
  163. package/dist/assets/sz-qxMwa2gs.svg +34 -0
  164. package/dist/assets/tc-CJHJmJj1.svg +50 -0
  165. package/dist/assets/tc-dtelpZmc.svg +50 -0
  166. package/dist/assets/tm-C_WSgUcv.svg +204 -0
  167. package/dist/assets/tm-DGBJvQay.svg +205 -0
  168. package/dist/assets/{tool-workshop-store-9_FKfdnT.js → tool-workshop-store-C-j6ilp2.js} +1 -1
  169. package/dist/assets/{toolcall-store-C4Trr01T.js → toolcall-store-CMBLYe4d.js} +1 -1
  170. package/dist/assets/{trash-2-DjfJXqVK.js → trash-2-CCIzK_vg.js} +1 -1
  171. package/dist/assets/un-Bqg4Cbbh.svg +16 -0
  172. package/dist/assets/un-DabL4p35.svg +16 -0
  173. package/dist/assets/va-B9-hqIE-.svg +190 -0
  174. package/dist/assets/va-s7kyhqIM.svg +190 -0
  175. package/dist/assets/vg-C7xY6pic.svg +59 -0
  176. package/dist/assets/vg-ClZ-0KpG.svg +59 -0
  177. package/dist/assets/vi-BC_zcciE.svg +28 -0
  178. package/dist/assets/vi-BSdsyIxY.svg +28 -0
  179. package/dist/assets/xk-Bj15g7cp.svg +5 -0
  180. package/dist/assets/xk-Cdz2uTvR.svg +5 -0
  181. package/dist/assets/zm-BmsW91ne.svg +27 -0
  182. package/dist/assets/zm-D8B-0kdx.svg +27 -0
  183. package/dist/assets/zw-CSuuaw9K.svg +21 -0
  184. package/dist/assets/zw-U0m7oJ5e.svg +21 -0
  185. package/dist/index.html +3 -2
  186. package/package.json +89 -88
  187. package/dist/assets/GoogleCallbackPage-BUIkYpZv.js +0 -1
  188. package/dist/assets/SettingsPage-DTJP-4xU.js +0 -1
  189. package/dist/assets/generate-service-DmtXqq_h.js +0 -3
  190. package/dist/assets/index-534TcCbY.js +0 -8
  191. package/dist/assets/index-B72Sw9uL.css +0 -1
  192. package/dist/assets/index-ClITfJ3t.js +0 -3
package/README.md CHANGED
@@ -1,276 +1,265 @@
1
- <p align="center">
2
- <img src="https://raw.githubusercontent.com/LogneBudo/llmxray/master/public/favicon.svg" alt="LLMxRay" width="80" />
3
- </p>
4
-
5
- <h1 align="center">LLMxRay</h1>
6
- <p align="center"><strong>Local LLM Observatory</strong></p>
7
- <p align="center">
8
- See what your AI is <em>actually</em> doing &mdash; token by token, layer by layer.
9
- </p>
10
-
11
- <p align="center">
12
- <img src="https://img.shields.io/badge/vue-3.5-42b883?logo=vuedotjs&logoColor=white" alt="Vue 3.5" />
13
- <img src="https://img.shields.io/badge/vite-7.3-646cff?logo=vite&logoColor=white" alt="Vite 7.3" />
14
- <img src="https://img.shields.io/badge/typescript-5.9-3178c6?logo=typescript&logoColor=white" alt="TypeScript 5.9" />
15
- <img src="https://img.shields.io/badge/tailwind-4.2-06b6d4?logo=tailwindcss&logoColor=white" alt="Tailwind 4.2" />
16
- <img src="https://img.shields.io/badge/ollama-local-000?logo=ollama&logoColor=white" alt="Ollama" />
17
- </p>
18
-
19
- ---
20
-
21
- ## What is LLMxRay?
22
-
23
- LLMxRay is a **free, local-first** dashboard that connects to [Ollama](https://ollama.com) running on your machine. It lets you chat with any model you've downloaded and then **inspect everything that happened behind the scenes**: how fast each token arrived, what the model might have been "thinking", how different settings change the output, and much more.
24
-
25
- **No cloud. No API keys. No cost.** Everything runs on your hardware.
26
-
27
- ### Who is this for?
28
-
29
- | You are... | LLMxRay helps you... |
30
- |---|---|
31
- | **Curious beginner** | See AI responses form in real-time and learn what "temperature" or "tokens" actually mean |
32
- | **Student / educator** | Explore model behavior visually &mdash; great for AI/ML coursework and demos |
33
- | **Developer** | Debug prompts, compare models, profile latency, inspect tool calls |
34
- | **Researcher** | Run controlled experiments: same prompt, different settings, side-by-side results |
35
-
36
- ---
37
-
38
- ## Features at a Glance
39
-
40
- ### Chat with Real-Time Token Streaming
41
- Start a conversation with any Ollama model. Watch tokens appear one by one with **confidence coloring** &mdash; each token is tinted based on how quickly the model produced it (faster = more confident). Supports markdown rendering, multi-turn conversation, file attachments, and slash commands.
42
-
43
- ### Session Deep Dive
44
- Click any past session to explore six tabs of detail:
45
-
46
- - **Stream** &mdash; Every token with timing data, plus a metrics dashboard (time-to-first-token, tokens/sec, latency chart)
47
- - **Reasoning** &mdash; If you're running a reasoning model like DeepSeek-R1, the `<think>` blocks are parsed and displayed step-by-step
48
- - **Introspection** &mdash; Visualizations of layer activations, attention heatmaps, and model architecture (illustrative)
49
- - **Tools** &mdash; Timeline of any tool calls the model made, with parameters and results
50
- - **Agent** &mdash; State-flow graph showing how an agent-style prompt progressed
51
- - **Prompt** &mdash; Anatomy breakdown of your prompt: sections, token counts, structure
52
-
53
- ### Compare Models (and Settings)
54
- The comparison workbench goes beyond "Model A vs Model B". Create up to **4 slots**, each with its own model, temperature, system prompt, and sampling parameters. Compare the *same* model at different temperatures to see how creativity changes. Features include:
55
-
56
- - **Grid view** &mdash; Side-by-side streaming results with per-slot settings pills
57
- - **Diff view** &mdash; Word-level highlighting of what changed between two outputs
58
- - **Metrics bar** &mdash; Visual comparison of TTFT, tokens/sec, and total tokens
59
- - **Quick presets** &mdash; "Temperature Sweep" (3 temps) and "Deterministic Pair" (same seed) one-click setups
60
- - Embedding models are automatically filtered out &mdash; only chat-capable models appear
61
-
62
- ### Embeddings Lab
63
- Embed any text and visualize the resulting vector. Compare two texts with a **cosine similarity meter** to see how semantically close they are. A hands-on way to understand what embeddings actually represent.
64
-
65
- ### RAG Pipeline
66
- Build a local knowledge base from your documents:
67
-
68
- 1. **Upload** PDFs, Word docs (.docx), or CSVs
69
- 2. **Chunk & embed** automatically using your chosen embedding model
70
- 3. **Search** with natural language &mdash; results ranked by semantic similarity
71
-
72
- Everything is stored in **IndexedDB** (your browser's built-in database). Zero cost, zero setup, zero external services.
73
-
74
- ### Tool Workshop (Visual Canvas)
75
- Build, edit, and test tool definitions on an interactive **node-based canvas** powered by Vue Flow:
76
-
77
- - **Drag-and-drop nodes** &mdash; Each tool is a visual node showing name, description, parameters, and implementation body
78
- - **Inline code editing** &mdash; Full CodeMirror 6 editors with TypeScript syntax highlighting directly on each node
79
- - **Bidirectional code sync** &mdash; Open the Code Panel to see all tools as combined TypeScript source. Edit code, nodes update. Edit nodes, code updates. Powered by a Recast AST parser
80
- - **Schema viewer** &mdash; Auto-generated OpenAI-compatible JSON schemas with one-click copy
81
- - **Probe & Pick** &mdash; Point at any API URL, inspect the response JSON tree, and auto-generate fetch code + parameter mappings
82
- - **OpenAPI discovery** &mdash; Auto-detect and parse OpenAPI/Swagger specs to pick endpoints visually
83
- - **Live execution overlays** &mdash; During chat, tool nodes pulse when the model calls them and show results inline
84
- - **Templates** &mdash; Start from 15+ built-in templates (web fetch, calculator, Google Calendar/Gmail, regex tester, and more)
85
- - **Persistent layout** &mdash; Node positions, mappings, and probe configs survive across sessions
86
-
87
- ### Tool Call Optimizer
88
- When the model calls a tool during chat, an **"Optimize this Tool"** button appears on the result. Click it to open the Response Optimizer Drawer:
89
-
90
- - Visualize the API response as an interactive JSON tree
91
- - Select only the fields the model actually needs
92
- - Auto-generate optimized fetch code with field extraction
93
- - One click to create a new optimized tool in the Workshop
94
-
95
- ### Response Quality Gates
96
- Every assistant response is automatically analyzed for common quality issues. Small colored badges appear below the response metrics when problems are detected:
97
-
98
- - **Repetition** &mdash; Flags responses with excessive repeated 4-gram phrases (>50% = fail, >30% = warn)
99
- - **Refusal** &mdash; Detects 8 common refusal patterns ("as an AI language model", "I cannot help", etc.)
100
- - **Gibberish** &mdash; Warns when non-ASCII characters exceed 40% of the response
101
- - **Empty** &mdash; Flags responses with fewer than 10 words
102
- - **Truncation** &mdash; Warns when the response hit the token limit or used >90% of budget without clean ending
103
-
104
- No news is good news &mdash; badges only appear when something is wrong.
105
-
106
- ### Cost Dashboard
107
- Track token usage across all your sessions with estimated cloud-equivalent costs. Navigate to the **Costs** page in the sidebar to see:
108
-
109
- - **Summary cards** &mdash; Total tokens, sessions, estimated cost, average cost per session
110
- - **Token usage by model** &mdash; Stacked bar chart showing prompt vs completion tokens per model
111
- - **Daily usage trends** &mdash; Line chart with dual axes (tokens + estimated cost over time)
112
- - **Model breakdown table** &mdash; Detailed per-model statistics with pricing source transparency
113
-
114
- Costs are estimates based on equivalent cloud API pricing (Groq, Together AI, Google, Mistral, etc.). Ollama runs locally at zero cost &mdash; the dashboard shows what you're saving.
115
-
116
- ### Performance Analytics
117
- A dedicated **Analytics** page (`/analytics`) provides deep insight into your usage patterns:
118
-
119
- - **Latency percentiles** (P50/P95/P99) for request duration and time-to-first-token
120
- - **Quality over turns** &mdash; per-turn quality scoring across multi-turn conversations
121
- - **Error intelligence** &mdash; error classification (7 categories) with timeline and per-model breakdown
122
- - **Usage heatmap** &mdash; 7-day x 24-hour grid showing your active hours
123
- - **Model distribution** &mdash; doughnut chart of request volume per model
124
- - **Settings impact** &mdash; scatter plots correlating temperature with generation speed
125
- - **Model load history** &mdash; timeline of cold/warm starts with load durations
126
-
127
- ### Model Browser
128
- See every model installed in Ollama with details like parameter count, quantization level, family, and format. Includes architecture diagrams showing the model's structure.
129
-
130
- ### System Monitor
131
- Real hardware specs (not browser estimates) &mdash; CPU model, total RAM with live usage, GPU with driver version, storage. Plus live Ollama status: running models, memory allocation, inference settings.
132
-
133
- ### Settings
134
- Configure your Ollama connection URL with a live connection tester. Set default **temperature** and **context length** with visual scales and educational tooltips that explain what each setting does in plain language.
135
-
136
- ---
137
-
138
- ## Quick Start
139
-
140
- ### Prerequisites
141
-
142
- 1. **Node.js 18+** &mdash; [Download](https://nodejs.org)
143
- 2. **Ollama** running locally &mdash; [Download](https://ollama.com/download)
144
- 3. At least one model pulled:
145
- ```bash
146
- ollama pull llama3.2
147
- ```
148
-
149
- ### Install and Run
150
-
151
- ```bash
152
- git clone https://github.com/LogneBudo/llmxray.git
153
- cd llmxray
154
- npm install
155
- npm run dev
156
- ```
157
-
158
- Open **http://localhost:5173** in your browser. That's it.
159
-
160
- > LLMxRay's dev server automatically proxies API calls to Ollama at `localhost:11434`. If Ollama is running on a different port or machine, change it in **Settings**.
161
-
162
- ### Build for Production
163
-
164
- ```bash
165
- npm run build # Output in dist/
166
- npm run preview # Preview the build locally
167
- ```
168
-
169
- ---
170
-
171
- ## Tech Stack
172
-
173
- | Layer | Technology |
174
- |---|---|
175
- | Framework | Vue 3.5 + Composition API (`<script setup>`) |
176
- | Language | TypeScript 5.9 (strict) |
177
- | Build | Vite 7.3 |
178
- | Styling | Tailwind CSS 4.2 (custom dark theme) |
179
- | State | Pinia 3 (one store per concern) |
180
- | Routing | Vue Router 5 |
181
- | Charts | Chart.js 4 + vue-chartjs, D3.js 7 |
182
- | Canvas | Vue Flow 1.x (node-based visual editor) |
183
- | Code Editor | CodeMirror 6 (TypeScript + JSON highlighting) |
184
- | AST Parser | Recast + @babel/parser (bidirectional code sync) |
185
- | Markdown | marked |
186
- | Diffing | diff (word-level) |
187
- | Documents | pdfjs-dist (lazy), mammoth (DOCX), papaparse (CSV) |
188
- | Storage | IndexedDB (browser-native, zero-cost) |
189
- | IDs | nanoid |
190
- | LLM Backend | Ollama (local, via `/api` proxy) |
191
-
192
- ---
193
-
194
- ## Architecture Highlights
195
-
196
- **Streaming** &mdash; LLMxRay reads Ollama's NDJSON response streams via `fetch()` + `ReadableStream`. Tokens arrive one by one and update the UI reactively through Pinia stores.
197
-
198
- **Token confidence** &mdash; Ollama doesn't expose logprobs, so confidence is approximated from inter-token latency. Faster tokens = higher confidence. This is labeled clearly in the UI as an approximation.
199
-
200
- **Introspection data** &mdash; Layer activations and attention heatmaps are synthetic (illustrative). They demonstrate what these visualizations *would* look like with real data. Clearly labeled as "Illustrative" in the UI.
201
-
202
- **Store-per-concern** &mdash; Each domain has its own Pinia store: tokens, sessions, metrics, reasoning, comparison, embeddings, RAG, models, and more. This keeps state management modular and testable.
203
-
204
- **Hardware detection** &mdash; The System page uses a custom Vite plugin (`vite-plugin-system-info.ts`) that queries the OS directly via PowerShell (Windows), `/proc` + `lspci` (Linux), or `sysctl` (macOS) for accurate hardware specs.
205
-
206
- ---
207
-
208
- ## Development
209
-
210
- ### Scripts
211
-
212
- | Command | What it does |
213
- |---|---|
214
- | `npm run dev` | Start dev server (port 5173) |
215
- | `npm run build` | Type-check + production build |
216
- | `npm run preview` | Preview production build |
217
- | `npm run test` | Run unit tests (Vitest) |
218
- | `npm run test:watch` | Tests in watch mode |
219
- | `npm run test:coverage` | Coverage report |
220
- | `npm run test:e2e` | Playwright end-to-end tests |
221
- | `npm run test:e2e:headed` | E2E with visible browser |
222
- | `npm run test:e2e:live` | E2E against live Ollama |
223
-
224
- ### Project Structure
225
-
226
- ```
227
- src/
228
- pages/ 8 page components (Dashboard, Compare, RAG, etc.)
229
- components/ 50+ components organized by feature
230
- chat/ Chat UI, token stream, attachments
231
- comparison/ Slot configurator, grid, diff view, metrics bar
232
- metrics/ Dashboard, charts, session history
233
- reasoning/ Think-block viewer
234
- introspection/ Layer activations, attention, architecture
235
- rag/ Document upload, search, ingest
236
- embeddings/ Vector viz, similarity meter
237
- tool-canvas/ Visual canvas, node editor, CodeMirror wrapper
238
- tool-optimizer/ Response optimizer drawer, JSON tree
239
- tool-calls/ Tool call timeline, definitions
240
- agent-graph/ Agent state flow
241
- common/ Layout, sidebar, shared components
242
- stores/ Pinia stores (one per concern)
243
- services/ Ollama client, streaming, generation, RAG, AST parser, probe
244
- types/ TypeScript interfaces
245
- utils/ Formatting, color scales, slot labels
246
- composables/ Vue composables (markdown, etc.)
247
- router/ Route definitions
248
- ```
249
-
250
- ---
251
-
252
- ## Troubleshooting
253
-
254
- | Problem | Solution |
255
- |---|---|
256
- | "Disconnected" in Settings | Make sure Ollama is running: `ollama serve` |
257
- | No models in dropdowns | Pull a model first: `ollama pull llama3.2` |
258
- | System page shows "Restart dev server" | Stop and restart `npm run dev` (the hardware plugin loads at startup) |
259
- | Slow first response | Normal &mdash; Ollama loads the model into memory on first use |
260
- | High RAM/VRAM usage | Use smaller quantized models (Q4) or reduce context length in Settings |
261
-
262
- ---
263
-
264
- ## License
265
-
266
- Licensed under the [Apache License 2.0](LICENSE). You are free to use, modify, and distribute this software under the terms of the license.
267
-
268
- ## Trademark
269
-
270
- **LLMxRay** is a trademark of Ivan Stankovic ([LogneBudo](https://github.com/LogneBudo)). See [TRADEMARK.md](TRADEMARK.md) for usage guidelines.
271
-
272
- ---
273
-
274
- <p align="center">
275
- Built with curiosity by <a href="https://github.com/LogneBudo">LogneBudo</a>
276
- </p>
1
+ <p align="center">
2
+ <img src="https://raw.githubusercontent.com/LogneBudo/llmxray/master/public/favicon.svg" alt="LLMxRay" width="80" />
3
+ </p>
4
+
5
+ <h1 align="center">LLMxRay</h1>
6
+ <p align="center"><strong>See what your AI is actually doing.</strong></p>
7
+ <p align="center">
8
+ Real-time token streaming, quality analysis, performance profiling, and cost tracking<br/>
9
+ for local LLMs. No cloud. No API keys. No cost.
10
+ </p>
11
+
12
+ <p align="center">
13
+ <a href="https://www.npmjs.com/package/llmxray"><img src="https://img.shields.io/npm/v/llmxray?color=cb3837&logo=npm&logoColor=white" alt="npm" /></a>
14
+ <a href="https://hub.docker.com/r/djovaneli/llmxray"><img src="https://img.shields.io/docker/pulls/djovaneli/llmxray?color=2496ED&logo=docker&logoColor=white" alt="Docker" /></a>
15
+ <img src="https://img.shields.io/badge/license-Apache%202.0-blue" alt="License" />
16
+ <img src="https://img.shields.io/badge/ollama-local-000?logo=ollama&logoColor=white" alt="Ollama" />
17
+ </p>
18
+
19
+ <p align="center">
20
+ <a href="#quick-start">Quick Start</a> &bull;
21
+ <a href="#features">Features</a> &bull;
22
+ <a href="#screenshots">Screenshots</a> &bull;
23
+ <a href="#who-is-this-for">Who Is This For</a> &bull;
24
+ <a href="CHANGELOG.md">Changelog</a>
25
+ </p>
26
+
27
+ <p align="center">
28
+ <img src="docs/public/screenshots/demo.gif" alt="LLMxRay demo — real-time token streaming with confidence coloring" width="800" />
29
+ </p>
30
+
31
+ ---
32
+
33
+ ## Quick Start
34
+
35
+ **One command. 30 seconds.**
36
+
37
+ ```bash
38
+ npx llmxray
39
+ ```
40
+
41
+ Or with Docker:
42
+
43
+ ```bash
44
+ docker run -p 5174:5174 djovaneli/llmxray
45
+ ```
46
+
47
+ Open **http://localhost:5174** and start chatting. That's it.
48
+
49
+ > **Prerequisite:** [Ollama](https://ollama.com/download) running locally with at least one model pulled (`ollama pull llama3.2`).
50
+
51
+ ---
52
+
53
+ ## Why LLMxRay?
54
+
55
+ You run a local LLM. You chat with it. But what actually happened?
56
+
57
+ - How fast was each token? Which ones was the model confident about?
58
+ - Is the response quality degrading over long conversations?
59
+ - What would this have cost if you ran it in the cloud?
60
+ - Is the model repeating itself? Refusing? Generating gibberish?
61
+ - How does temperature 0.3 compare to 0.9 on the *same* prompt?
62
+
63
+ **LLMxRay answers all of these, visually, in real time, for free.**
64
+
65
+ ---
66
+
67
+ ## Features
68
+
69
+ ### Real-Time Chat with Token Intelligence
70
+ Chat with any Ollama model and watch tokens arrive with **confidence coloring** each token is tinted based on generation speed. Supports markdown, multi-turn conversations, file attachments, vision models, and slash commands.
71
+
72
+ ### Response Quality Gates
73
+ Every response is automatically analyzed. Colored badges appear only when something is wrong:
74
+ - **Repetition** excessive repeated phrases (4-gram analysis)
75
+ - **Refusal** "as an AI language model" and 7 other patterns
76
+ - **Gibberish** — high non-ASCII ratio
77
+ - **Empty** fewer than 10 words
78
+ - **Truncation** hit the token limit without finishing
79
+
80
+ ### Model Comparison Workbench
81
+ Up to **4 slots** with independent model, temperature, and system prompt. Features include side-by-side streaming, word-level diff highlighting, metrics comparison, and one-click presets (Temperature Sweep, Deterministic Pair, Language Compare with Token Tax visualization).
82
+
83
+ ### Performance Analytics
84
+ - **Latency percentiles** (P50/P95/P99) for duration and TTFT
85
+ - **Error intelligence** 7-category classifier with timeline
86
+ - **Usage heatmap** — 7x24 grid of your active hours
87
+ - **Settings impact** — temperature vs tokens/sec scatter plots
88
+ - **Cold vs warm start** tracking with model load history
89
+
90
+ ### Cost Dashboard
91
+ Token usage per model/day with estimated cloud-equivalent pricing. See what you're *saving* by running locally.
92
+
93
+ ### Surgical Benchmark
94
+ Test model knowledge with multi-choice question suites. Uses real logprobs via OpenAI-compatible endpoint for accurate confidence measurement. Build custom suites visually or let AI generate them from a topic.
95
+
96
+ ### Embeddings Lab & RAG Pipeline
97
+ Embed text, visualize vectors, measure cosine similarity. Build a local knowledge base from PDFs, DOCX, and CSV — chunked, embedded, and searchable. All stored in IndexedDB. Zero cost.
98
+
99
+ ### Tool Workshop (Visual Canvas)
100
+ Drag-and-drop node canvas for building tool definitions. Bidirectional code sync (edit nodes or TypeScript — both update). Probe APIs, auto-generate schemas, test with live execution.
101
+
102
+ ### AI Training Pipeline
103
+ Curate training data from your conversations. Tag, review, and export as JSONL for fine-tuning.
104
+
105
+ ### Local AI History Database
106
+ Every experiment (benchmarks, comparisons, chats, training pairs) is automatically archived in a queryable IndexedDB database with filters, trends, exports, and retention policies.
107
+
108
+ ### Multilingual
109
+ Full translations in English, French, Chinese, and Arabic. RTL layout support. Community scaffolds for Hebrew and Japanese.
110
+
111
+ ---
112
+
113
+ ## Screenshots
114
+
115
+ <table>
116
+ <tr>
117
+ <td width="50%">
118
+
119
+ **Chat with token streaming and confidence**
120
+ ![Chat](docs/public/screenshots/chat-diagnostics.png)
121
+
122
+ </td>
123
+ <td width="50%">
124
+
125
+ **Model comparison side by side**
126
+ ![Compare](docs/public/screenshots/compare-sidebyside.png)
127
+
128
+ </td>
129
+ </tr>
130
+ <tr>
131
+ <td width="50%">
132
+
133
+ **Session deep dive — metrics and timing**
134
+ ![Session](docs/public/screenshots/session-details.png)
135
+
136
+ </td>
137
+ <td width="50%">
138
+
139
+ **Benchmark with confidence radar**
140
+ ![Benchmark](docs/public/screenshots/benchmark.png)
141
+
142
+ </td>
143
+ </tr>
144
+ <tr>
145
+ <td width="50%">
146
+
147
+ **Embeddings — cosine similarity**
148
+ ![Embeddings](docs/public/screenshots/embed-similarity.png)
149
+
150
+ </td>
151
+ <td width="50%">
152
+
153
+ **System monitor — hardware and Ollama status**
154
+ ![System](docs/public/screenshots/my-system.png)
155
+
156
+ </td>
157
+ </tr>
158
+ </table>
159
+
160
+ ---
161
+
162
+ ## Who Is This For
163
+
164
+ | You are... | LLMxRay helps you... |
165
+ |---|---|
166
+ | **Developer** | Debug prompts, profile latency, compare models, inspect tool calls, track costs |
167
+ | **Researcher** | Run controlled experiments with consistent settings across models and temperatures |
168
+ | **Student / Educator** | Explore model behavior visually — built-in Educators Kit with 9 interactive modules |
169
+ | **AI team lead** | Understand quality trends, error patterns, and resource usage across your local fleet |
170
+
171
+ ---
172
+
173
+ ## Install Options
174
+
175
+ ### npx (recommended)
176
+ ```bash
177
+ npx llmxray
178
+ npx llmxray --port 3000
179
+ npx llmxray --ollama-url http://192.168.1.50:11434
180
+ ```
181
+
182
+ ### Docker
183
+ ```bash
184
+ docker run -p 5174:5174 djovaneli/llmxray
185
+ docker run -p 5174:5174 -e OLLAMA_URL=http://host.docker.internal:11434 djovaneli/llmxray
186
+ ```
187
+
188
+ ### From source
189
+ ```bash
190
+ git clone https://github.com/LogneBudo/llmxray.git
191
+ cd llmxray
192
+ npm install
193
+ npm run dev # http://localhost:5173
194
+ ```
195
+
196
+ ---
197
+
198
+ ## Tech Stack
199
+
200
+ | Layer | Technology |
201
+ |---|---|
202
+ | Framework | Vue 3.5 + Composition API |
203
+ | Language | TypeScript 5.9 (strict) |
204
+ | Build | Vite 7.3 |
205
+ | Styling | Tailwind CSS 4.2 |
206
+ | State | Pinia 3 (store-per-concern) |
207
+ | Charts | Chart.js 4, D3.js 7 |
208
+ | Canvas | Vue Flow (visual node editor) |
209
+ | Code Editor | CodeMirror 6 |
210
+ | Storage | IndexedDB (browser-native) |
211
+ | LLM Backend | Ollama (local) |
212
+
213
+ ---
214
+
215
+ ## Architecture
216
+
217
+ **Streaming** — Reads Ollama NDJSON via `fetch()` + `ReadableStream`. Tokens update the UI reactively through Pinia stores.
218
+
219
+ **Token confidence** Approximated from inter-token latency (faster = more confident). Clearly labeled as approximation. Benchmarks use real logprobs via OpenAI-compatible endpoint.
220
+
221
+ **Store-per-concern** Each domain has its own Pinia store: tokens, sessions, metrics, reasoning, comparison, embeddings, quality, cost, and more.
222
+
223
+ **Hardware detection** — Custom Vite plugin queries the OS directly (PowerShell/proc/sysctl) for accurate hardware specs.
224
+
225
+ ---
226
+
227
+ ## Development
228
+
229
+ | Command | What it does |
230
+ |---|---|
231
+ | `npm run dev` | Dev server (port 5173) |
232
+ | `npm run build` | Type-check + production build |
233
+ | `npm run test` | Unit tests (Vitest) |
234
+ | `npm run test:e2e` | End-to-end (Playwright) |
235
+
236
+ ---
237
+
238
+ ## Contributing
239
+
240
+ Contributions welcome! See [CONTRIBUTING.md](CONTRIBUTING.md) for setup and guidelines.
241
+
242
+ **Community translations especially welcome** — scaffold files ready for Hebrew and Japanese.
243
+
244
+ ---
245
+
246
+ ## License
247
+
248
+ [Apache License 2.0](LICENSE)
249
+
250
+ ## Trademark
251
+
252
+ **LLMxRay** is a trademark of Ivan Stankovic ([LogneBudo](https://github.com/LogneBudo)). See [TRADEMARK.md](TRADEMARK.md).
253
+
254
+ ---
255
+
256
+ <p align="center">
257
+ <strong>If LLMxRay helps you understand your AI better, consider giving it a star.</strong><br/>
258
+ It helps others discover the project.
259
+ </p>
260
+
261
+ <p align="center">
262
+ <a href="https://github.com/LogneBudo/llmxray">
263
+ <img src="https://img.shields.io/github/stars/LogneBudo/llmxray?style=social" alt="GitHub stars" />
264
+ </a>
265
+ </p>