webai-scanner 0.1.0__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- webai_scanner-0.1.0/ARCHITECTURE.md +228 -0
- webai_scanner-0.1.0/PKG-INFO +212 -0
- webai_scanner-0.1.0/README.md +203 -0
- webai_scanner-0.1.0/WebAI_specification.md +86 -0
- webai_scanner-0.1.0/benchmark_results.csv +3 -0
- webai_scanner-0.1.0/benchmark_runner.py +326 -0
- webai_scanner-0.1.0/catalog_scanner.py +204 -0
- webai_scanner-0.1.0/doc_updater.py +112 -0
- webai_scanner-0.1.0/docs/ARCHITECTURE.md +228 -0
- webai_scanner-0.1.0/docs/WebAI_specification.md +122 -0
- webai_scanner-0.1.0/docs/decisions.md +272 -0
- webai_scanner-0.1.0/docs/research.md +111 -0
- webai_scanner-0.1.0/docs/starter.md +100 -0
- webai_scanner-0.1.0/download_models.py +99 -0
- webai_scanner-0.1.0/llms-full.txt +137 -0
- webai_scanner-0.1.0/llms.txt +33 -0
- webai_scanner-0.1.0/pyproject.toml +21 -0
- webai_scanner-0.1.0/run.sh +47 -0
- webai_scanner-0.1.0/src/webai/__init__.py +5 -0
- webai_scanner-0.1.0/src/webai/catalog_scanner.py +204 -0
- webai_scanner-0.1.0/src/webai/cli.py +87 -0
- webai_scanner-0.1.0/src/webai/steam_client.py +434 -0
- webai_scanner-0.1.0/starter.md +100 -0
- webai_scanner-0.1.0/steam_client.py +434 -0
- webai_scanner-0.1.0/storage4gaming_client.py +303 -0
- webai_scanner-0.1.0/telemetry_db.py +219 -0
- webai_scanner-0.1.0/test_telemetry.py +93 -0
- webai_scanner-0.1.0/webai_benchmarks.db +0 -0
|
@@ -0,0 +1,228 @@
|
|
|
1
|
+
# WebAI Architecture & System Design
|
|
2
|
+
|
|
3
|
+
**WebAI** is an open architectural standard and benchmarking suite designed to establish an **AI Agent Readability & Identity Standard** for the modern web.
|
|
4
|
+
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
## 1. The Core Problem: Human-First vs. Agent-First Web
|
|
8
|
+
|
|
9
|
+
### The 57% Reality
|
|
10
|
+
Cloudflare and major internet infrastructure providers have documented that **over 57% of all web traffic** is generated by automated bots and AI agents. Despite this majority, the internet remains engineered exclusively **human-first**:
|
|
11
|
+
- Web pages are dominated by presentation markup: nested DOM `<div>`s, CSS rules, stylesheets, layout boilerplate, navigation bars, cookie banners, tracking scripts, and SVGs.
|
|
12
|
+
- A human views only the rendered graphical output. However, an AI agent interacting with the site must ingest the entire raw DOM payload into its transformer context window.
|
|
13
|
+
- As an example, the Storage4gaming calculator page contains over **60,000 lines of HTML (~3.2 MB)**. Feeding this markup to an LLM wastes thousands of input tokens and imposes significant latency and compute overhead.
|
|
14
|
+
- This creates massive downstream economic and societal waste: compute cycles that could solve complex tasks are instead squandered reading visual layout code.
|
|
15
|
+
|
|
16
|
+
### WebAI's Dual-Web Solution
|
|
17
|
+
WebAI establishes an **agent-ready web operating in parallel with the human web**:
|
|
18
|
+
- Essential factual data is surfaced through compact JSON, semantic JSON-LD, or structured Markdown (e.g. `llms.txt`).
|
|
19
|
+
- A 60k-line DOM is distilled down into a ~100-byte semantic object.
|
|
20
|
+
- **Result:** ~75–95% input token reduction, 1.5x–2.5x speedup in agent inference, and zero loss of extraction accuracy.
|
|
21
|
+
|
|
22
|
+
```
|
|
23
|
+
┌────────────────────────────────────────────────────────┐
|
|
24
|
+
│ Human-First Web: ~1,800 - 3,200,000 bytes │
|
|
25
|
+
│ [DOM wrappers, styling, tracking, UI boilerplate] │
|
|
26
|
+
│ ↳ Ingested by Agent = 600 - 50,000+ tokens │
|
|
27
|
+
└────────────────────────────────────────────────────────┘
|
|
28
|
+
vs
|
|
29
|
+
┌────────────────────────────────────────────────────────┐
|
|
30
|
+
│ Agent-First Web: ~100 bytes │
|
|
31
|
+
│ [Structured JSON / Semantic Key-Values / llms.txt] │
|
|
32
|
+
│ ↳ Ingested by Agent = ~100 - 150 tokens │
|
|
33
|
+
│ ↳ 75% - 95% Token Reduction, 2x+ Latency Improvement │
|
|
34
|
+
└────────────────────────────────────────────────────────┘
|
|
35
|
+
```
|
|
36
|
+
|
|
37
|
+
---
|
|
38
|
+
|
|
39
|
+
## 2. System Architecture & Comparative Data Flow
|
|
40
|
+
|
|
41
|
+
The following diagram illustrates how WebAI evaluates, benchmarks, and logs token telemetry across local open-weight models on Apple Silicon:
|
|
42
|
+
|
|
43
|
+
```mermaid
|
|
44
|
+
flowchart TD
|
|
45
|
+
subgraph Data Sources & Case Studies
|
|
46
|
+
S4G_HTML[Storage4gaming Live DOM<br/>Human-First Payload: 3.2MB]
|
|
47
|
+
S4G_Agent[Storage4gaming llms.txt & JSON<br/>Agent-First: ~100 bytes]
|
|
48
|
+
Steam_HTML[Steam Storefront Page DOM<br/>Human-First Payload: ~250KB]
|
|
49
|
+
Steam_API[Steam Store API appdetails<br/>Agent-First: ~115 bytes]
|
|
50
|
+
Catalog[Cross-Platform Steam Discovery<br/>Windows & macOS Local Filesystem]
|
|
51
|
+
end
|
|
52
|
+
|
|
53
|
+
subgraph Client & Calculator Layer
|
|
54
|
+
S4G_Client[storage4gaming_client.py<br/>S4G Spec Client & llms.txt Generator]
|
|
55
|
+
Steam_Client[steam_client.py<br/>Steam API Client & Hardware Calculator]
|
|
56
|
+
|
|
57
|
+
S4G_HTML --> S4G_Client
|
|
58
|
+
S4G_Agent --> S4G_Client
|
|
59
|
+
Steam_HTML --> Steam_Client
|
|
60
|
+
Steam_API --> Steam_Client
|
|
61
|
+
Catalog --> Steam_Client
|
|
62
|
+
end
|
|
63
|
+
|
|
64
|
+
subgraph Inference & Benchmarking Engine
|
|
65
|
+
Runner[benchmark_runner.py<br/>MLX-LM Sequential Harness]
|
|
66
|
+
S4G_Client --> Runner
|
|
67
|
+
Steam_Client --> Runner
|
|
68
|
+
|
|
69
|
+
Runner --> M1[Bucket 1: Sub-4B Models<br/>Llama 3.2, Qwen 2.5, Phi 4]
|
|
70
|
+
Runner --> M2[Bucket 2: 4B-12B Models<br/>Gemma 4 12B, Gemma 4 E4B, DeepSeek R1]
|
|
71
|
+
end
|
|
72
|
+
|
|
73
|
+
subgraph Telemetry & Persistence Layer
|
|
74
|
+
Runner --> Telemetry[telemetry_db.py<br/>SQLite Telemetry Engine]
|
|
75
|
+
Telemetry --> DB[(webai_benchmarks.db)]
|
|
76
|
+
Telemetry --> CSV[benchmark_results.csv]
|
|
77
|
+
Telemetry --> DocSync[doc_updater.py<br/>Auto-Doc Sync]
|
|
78
|
+
DocSync --> Readme[README.md & ARCHITECTURE.md]
|
|
79
|
+
end
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
---
|
|
83
|
+
|
|
84
|
+
## 3. The AIAID (AI Agent Identity) Architecture
|
|
85
|
+
|
|
86
|
+
To bring transparency and trusted authentication to agent web traversal, WebAI introduces the **AIAID (AI Agent ID)** protocol.
|
|
87
|
+
|
|
88
|
+
### Why AIAID?
|
|
89
|
+
Today, devices have MAC addresses and serial numbers. However, AI agents traverse the internet anonymously, leading websites to deploy aggressive bot blockers (Cloudflare challenge walls, CAPTCHAs).
|
|
90
|
+
|
|
91
|
+
With AIAID:
|
|
92
|
+
1. Legitimate agents identify themselves transparently.
|
|
93
|
+
2. Participating sites grant authenticated agents **instant access to high-density semantic endpoints (`llms.txt`, JSON API)** without CAPTCHAs.
|
|
94
|
+
3. Unverified scrapers continue to receive standard human-facing HTML or bot challenge walls.
|
|
95
|
+
|
|
96
|
+
```mermaid
|
|
97
|
+
sequenceDiagram
|
|
98
|
+
autonumber
|
|
99
|
+
actor User as Human User
|
|
100
|
+
participant Agent as AI Agent (with AIAID)
|
|
101
|
+
participant Site as WebAI-Ready Site (Storage4gaming)
|
|
102
|
+
participant Registry as Public AIAID Depository / Registry
|
|
103
|
+
|
|
104
|
+
User->>Agent: "Calculate required storage for my Steam library"
|
|
105
|
+
Agent->>Site: HTTP GET /calculator (Header: X-AIAID: aiaid_7f9c2b...)
|
|
106
|
+
Note over Site: Site detects Agent seeking Agent-First Endpoint
|
|
107
|
+
Site->>Registry: Verify AIAID(aiaid_7f9c2b...)
|
|
108
|
+
Registry->>Agent: Issue Cryptographic Challenge Nonce
|
|
109
|
+
Agent-->>Registry: Signed Nonce Response (Private Key)
|
|
110
|
+
Registry-->>Site: AIAID Verified: Active, Legitimate, Tier: Standard
|
|
111
|
+
Site-->>Agent: HTTP 200 OK (Clean Semantic JSON / llms.txt)
|
|
112
|
+
Note over Agent: Consumes ~100 tokens in 0.4s!
|
|
113
|
+
Agent-->>User: "You need a 1TB NVMe SSD and 16GB RAM."
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
---
|
|
117
|
+
|
|
118
|
+
## 4. Storage4gaming Rewrite Blueprint (The Reference Site)
|
|
119
|
+
|
|
120
|
+
As the founding real-world testbed of WebAI, **Storage4gaming.com** will undergo a complete code rewrite into an **AI-Agent-Ready Architecture**:
|
|
121
|
+
|
|
122
|
+
```
|
|
123
|
+
storage4gaming.com/
|
|
124
|
+
├── Human Interface (Web Browser)
|
|
125
|
+
│ ├── /calculator/ -> Interactive Visual Hardware Sizing Tool (React / Tailwind)
|
|
126
|
+
│ ├── /games/[slug]/ -> Human-readable hardware review with interactive charts
|
|
127
|
+
│ └── /guides/ -> SSD & RAM buying guides
|
|
128
|
+
│
|
|
129
|
+
├── Agent Interface (AI Ready)
|
|
130
|
+
│ ├── /llms.txt -> Lightweight Markdown index of verified game hardware specs
|
|
131
|
+
│ ├── /llms-full.txt -> Complete agent knowledge base with sizing formulas
|
|
132
|
+
│ ├── /api/games/[slug].json -> Ultra-compact JSON payload (~100 bytes)
|
|
133
|
+
│ └── /api/auth/aiaid -> Verification endpoint for AIAID agent handshake
|
|
134
|
+
│
|
|
135
|
+
└── Smart Content Negotiation Layer
|
|
136
|
+
└── When Accept: application/json or X-AIAID is present -> Auto-routes to Agent Interface
|
|
137
|
+
```
|
|
138
|
+
|
|
139
|
+
---
|
|
140
|
+
|
|
141
|
+
## 5. Component Breakdown
|
|
142
|
+
|
|
143
|
+
### A. Site Integration Client (`storage4gaming_client.py`)
|
|
144
|
+
- Interfaces with Storage4gaming (`https://www.storage4gaming.com`).
|
|
145
|
+
- Generates three parallel representations for PC game titles:
|
|
146
|
+
1. `human_first_html`: Real-world DOM cards with CSS layout classes, badge wrappers, and page context (~1,800 to 3,200,000 bytes).
|
|
147
|
+
2. `agent_first_json`: High-density semantic JSON containing only verified hardware metrics (~100 bytes).
|
|
148
|
+
3. `agent_first_llmstxt`: Pure markdown entry adhering to the `llms.txt` standard (~120 bytes).
|
|
149
|
+
|
|
150
|
+
### B. Steam Store Client & Calculator (`steam_client.py`)
|
|
151
|
+
- Interfaces with the official Steam Store API (`store.steampowered.com/api/appdetails`).
|
|
152
|
+
- Evaluates the Steam Storefront desktop HTML page against the official JSON API.
|
|
153
|
+
- Implements the **Hardware Capacity Calculator**:
|
|
154
|
+
- Parses numeric RAM (GB), Storage (GB), and Drive Type (NVMe SSD, SSD, HDD).
|
|
155
|
+
- Computes raw storage footprint across all discovered games.
|
|
156
|
+
- Adds 20% safety margin for patches and swap files.
|
|
157
|
+
- Recommends commercial SSD drive sizing (500GB, 1TB, 2TB, 4TB).
|
|
158
|
+
- Recommends system RAM tier (16GB vs 32GB).
|
|
159
|
+
|
|
160
|
+
### C. Universal MLX Benchmark Harness (`benchmark_runner.py`)
|
|
161
|
+
- Built on Apple's `mlx-lm` framework for native Apple Silicon unified memory acceleration.
|
|
162
|
+
- Enforces **strict sequential execution** (`batch_size=1`, single model resident in RAM at any time).
|
|
163
|
+
- Between model runs, it calls `del model`, `del tokenizer`, `gc.collect()`, and `mx.clear_cache()` to prevent memory leakage.
|
|
164
|
+
- Captures exact tokenizer token counts, prompt eval speed (tok/sec), generation latency, and peak memory.
|
|
165
|
+
|
|
166
|
+
### D. Telemetry Engine & SQL Database (`telemetry_db.py`)
|
|
167
|
+
- Embedded SQLite engine requiring zero external daemon or background service.
|
|
168
|
+
- Database file: `webai_benchmarks.db`.
|
|
169
|
+
- Automatically computes `token_savings_pct` by pairing the agent-first run against the human-first baseline.
|
|
170
|
+
- Automatically synchronizes all database records to `benchmark_results.csv`.
|
|
171
|
+
|
|
172
|
+
### E. Catalog Scanner (`catalog_scanner.py`)
|
|
173
|
+
- Offline, zero-auth Steam discovery scanner with native support for both **macOS** and **Windows**:
|
|
174
|
+
- **macOS:** Scans `~/Library/Application Support/Steam/steamapps/libraryfolders.vdf`, `userdata/<id>/config/localconfig.vdf`, and binary `appcache/appinfo.vdf`.
|
|
175
|
+
- **Windows:** Auto-detects Steam via Windows Registry (`HKCU\Software\Valve\Steam\SteamPath`) or `C:\Program Files (x86)\Steam`, multi-drive library folders, user account cache, and `appcache\appinfo.vdf`.
|
|
176
|
+
|
|
177
|
+
### F. Documentation Auto-Updater (`doc_updater.py`)
|
|
178
|
+
- Automatically reads live data from `webai_benchmarks.db` and the Hugging Face cache.
|
|
179
|
+
- Directly updates the markdown tables in `README.md` and `ARCHITECTURE.md` to keep documentation synchronised with real experimental results.
|
|
180
|
+
|
|
181
|
+
---
|
|
182
|
+
|
|
183
|
+
## 6. Live Benchmark Performance Summary
|
|
184
|
+
|
|
185
|
+
The table below is dynamically synchronized with the local SQLite database (`webai_benchmarks.db`):
|
|
186
|
+
|
|
187
|
+
<!-- BENCHMARK_TABLE_START -->
|
|
188
|
+
|
|
189
|
+
| Model | Target Game | Human Tokens | Agent Tokens | Token Savings | Human Latency | Agent Latency | Speedup |
|
|
190
|
+
|------------------------------|------------------|----------------|----------------|-----------------|-----------------|-----------------|-------------|
|
|
191
|
+
| `Llama-3.2-1B-Instruct-4bit` | Apex Legends | 565 tok | 132 tok | **76.64%** | 0.56s | 0.42s | 1.3x faster |
|
|
192
|
+
| `Llama-3.2-1B-Instruct-4bit` | Counter-Strike 2 | 567 tok | 136 tok | **76.01%** | 0.99s | 0.45s | 2.2x faster |
|
|
193
|
+
| `Llama-3.2-1B-Instruct-4bit` | Halo Infinite | 565 tok | 132 tok | **76.64%** | 0.50s | 0.42s | 1.2x faster |
|
|
194
|
+
| `gemma-4-e4b-it-OptiQ-4bit` | Apex Legends | 635 tok | 125 tok | **80.31%** | 4.69s | 3.45s | 1.4x faster |
|
|
195
|
+
| `gemma-4-e4b-it-OptiQ-4bit` | Counter-Strike 2 | 637 tok | 132 tok | **79.28%** | 6.96s | 4.77s | 1.5x faster |
|
|
196
|
+
| `gemma-4-e4b-it-OptiQ-4bit` | Halo Infinite | 636 tok | 126 tok | **80.19%** | 3.61s | 3.31s | 1.1x faster |
|
|
197
|
+
|
|
198
|
+
<!-- BENCHMARK_TABLE_END -->
|
|
199
|
+
|
|
200
|
+
---
|
|
201
|
+
|
|
202
|
+
## 7. Model Evaluation Matrix
|
|
203
|
+
|
|
204
|
+
<!-- MODEL_STATUS_START -->
|
|
205
|
+
|
|
206
|
+
| Status | Bucket | Est. RAM (4-bit) | Hugging Face Model ID |
|
|
207
|
+
|--------------|----------|--------------------|--------------------------------------------------|
|
|
208
|
+
| **[CACHED]** | <4B | ~1.0 GB | `mlx-community/Llama-3.2-1B-Instruct-4bit` |
|
|
209
|
+
| [NOT CACHED] | <4B | ~2.2 GB | `mlx-community/Llama-3.2-3B-Instruct-4bit` |
|
|
210
|
+
| [NOT CACHED] | <4B | ~2.1 GB | `mlx-community/Qwen2.5-3B-Instruct-4bit` |
|
|
211
|
+
| [NOT CACHED] | <4B | ~2.5 GB | `mlx-community/Phi-4-mini-instruct-4bit` |
|
|
212
|
+
| **[CACHED]** | 4B-12B | ~7.5 GB | `mlx-community/gemma-4-12B-it-qat-4bit` |
|
|
213
|
+
| **[CACHED]** | 4B-12B | ~3.5 GB | `mlx-community/gemma-4-e4b-it-OptiQ-4bit` |
|
|
214
|
+
| [NOT CACHED] | 4B-12B | ~4.8 GB | `mlx-community/Qwen2.5-7B-Instruct-4bit` |
|
|
215
|
+
| [NOT CACHED] | 4B-12B | ~5.2 GB | `mlx-community/Llama-3.1-8B-Instruct-4bit` |
|
|
216
|
+
| [NOT CACHED] | 4B-12B | ~5.2 GB | `mlx-community/DeepSeek-R1-Distill-Qwen-8B-4bit` |
|
|
217
|
+
|
|
218
|
+
<!-- MODEL_STATUS_END -->
|
|
219
|
+
|
|
220
|
+
---
|
|
221
|
+
|
|
222
|
+
## 8. Hardware & Memory Profile (16 GB Unified RAM)
|
|
223
|
+
|
|
224
|
+
The WebAI suite is strictly engineered for Apple Silicon (macOS) with 16 GB unified memory:
|
|
225
|
+
- **macOS & System Reserve:** ~3.5 - 4.5 GB RAM.
|
|
226
|
+
- **Sub-4B Models:** ~0.9 - 2.5 GB peak active memory.
|
|
227
|
+
- **Medium (4B - 12B) Models:** ~3.5 - 7.5 GB peak active memory.
|
|
228
|
+
- **Headroom:** Always maintains >= 4 GB of free headroom, preventing swap file churn and thermal throttling.
|
|
@@ -0,0 +1,212 @@
|
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
|
+
Name: webai-scanner
|
|
3
|
+
Version: 0.1.0
|
|
4
|
+
Summary: AI Agent Readability auditor & local Steam game discovery engine
|
|
5
|
+
Author: Alexander Markovski
|
|
6
|
+
License: MIT
|
|
7
|
+
Requires-Python: >=3.10
|
|
8
|
+
Description-Content-Type: text/markdown
|
|
9
|
+
|
|
10
|
+
# WebAI: The AI-Ready Internet Standard & Token Benchmarking Suite
|
|
11
|
+
|
|
12
|
+
**WebAI** is an open architectural standard, benchmarking suite, and agent identity framework designed to establish an **AI-Agent-Ready Internet** operating alongside the modern human-first web.
|
|
13
|
+
|
|
14
|
+
---
|
|
15
|
+
|
|
16
|
+
## 🌍 The WebAI Manifesto: Why Build This Now?
|
|
17
|
+
|
|
18
|
+
* **The 57% Reality:** Major internet infrastructure providers like Cloudflare have documented that **over 57% of all web traffic** is now generated by automated systems and AI agents.
|
|
19
|
+
* **The Structural Flaw:** Despite AI being the majority of internet consumers, the web remains engineered exclusively for human visual perception (deeply nested HTML DOM trees, CSS stylesheets, layout wrappers, and tracking scripts).
|
|
20
|
+
* **The Downstream Impact:** AI agents waste up to **95%+ of their I/O token budgets** simply parsing presentation boilerplate. This inflates task completion times, drives up inference costs, and needlessly ties up global GPU clusters that could be solving valuable societal tasks.
|
|
21
|
+
* **The Dual-Web Solution:** Humans keep their rich visual web interfaces, while websites provide a parallel **Agent-Ready Web** with zero presentation bloat.
|
|
22
|
+
|
|
23
|
+
### The Two Pillars of WebAI
|
|
24
|
+
1. **Pillar 1: Content & Representation Standards**
|
|
25
|
+
* Standardized `llms.txt` and semantic JSON contracts.
|
|
26
|
+
* **Storage4gaming.com:** Complete rewrite into the reference AI-agent-ready site.
|
|
27
|
+
* **Empirical Benchmarks:** Quantifying token and latency savings across open-weight models on Storage4gaming and the Steam Storefront.
|
|
28
|
+
2. **Pillar 2: AI Agent Identity & Authentication (AIAID)**
|
|
29
|
+
* Cryptographic machine identities for AI agents (analogous to MAC addresses / UUIDs).
|
|
30
|
+
* A public depository and challenge-response handshake that unlocks high-speed semantic endpoints for verified agents while filtering malicious scrapers.
|
|
31
|
+
|
|
32
|
+
---
|
|
33
|
+
|
|
34
|
+
## 🌟 Key Highlights
|
|
35
|
+
|
|
36
|
+
- **~76% - 95% Token Reduction:** Cuts prompt tokens from thousands of tokens down to ~100 tokens.
|
|
37
|
+
- **1.5x - 2.5x Faster Latency:** Accelerates task completion by stripping away visual rendering layers.
|
|
38
|
+
- **Embedded SQLite Telemetry:** All benchmark runs, tokenizer token counts, throughput (tok/s), latency, and peak RAM are persisted to `webai_benchmarks.db` and synchronized with `benchmark_results.csv`.
|
|
39
|
+
- **Apple Silicon Optimized:** 100% offline inference on Apple Silicon using `mlx-lm` with sequential execution and automatic memory cache clearing to protect 16 GB Unified Memory.
|
|
40
|
+
- **Auto-Updating Documentation:** Live benchmark tables and model cache statuses in this README and `ARCHITECTURE.md` stay automatically synchronized with your local benchmark runs.
|
|
41
|
+
|
|
42
|
+
---
|
|
43
|
+
|
|
44
|
+
## 📊 Live Benchmark Performance Summary
|
|
45
|
+
|
|
46
|
+
> The table below is automatically synchronized with [`webai_benchmarks.db`](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/webai_benchmarks.db) by `doc_updater.py`:
|
|
47
|
+
|
|
48
|
+
<!-- BENCHMARK_TABLE_START -->
|
|
49
|
+
|
|
50
|
+
| Model | Target Game | Human Tokens | Agent Tokens | Token Savings | Human Latency | Agent Latency | Speedup |
|
|
51
|
+
|------------------------------|------------------|----------------|----------------|-----------------|-----------------|-----------------|-------------|
|
|
52
|
+
| `Llama-3.2-1B-Instruct-4bit` | Apex Legends | 565 tok | 132 tok | **76.64%** | 0.56s | 0.42s | 1.3x faster |
|
|
53
|
+
| `Llama-3.2-1B-Instruct-4bit` | Counter-Strike 2 | 567 tok | 136 tok | **76.01%** | 0.99s | 0.45s | 2.2x faster |
|
|
54
|
+
| `Llama-3.2-1B-Instruct-4bit` | Halo Infinite | 565 tok | 132 tok | **76.64%** | 0.50s | 0.42s | 1.2x faster |
|
|
55
|
+
| `gemma-4-e4b-it-OptiQ-4bit` | Apex Legends | 635 tok | 125 tok | **80.31%** | 4.69s | 3.45s | 1.4x faster |
|
|
56
|
+
| `gemma-4-e4b-it-OptiQ-4bit` | Counter-Strike 2 | 637 tok | 132 tok | **79.28%** | 6.96s | 4.77s | 1.5x faster |
|
|
57
|
+
| `gemma-4-e4b-it-OptiQ-4bit` | Halo Infinite | 636 tok | 126 tok | **80.19%** | 3.61s | 3.31s | 1.1x faster |
|
|
58
|
+
|
|
59
|
+
<!-- BENCHMARK_TABLE_END -->
|
|
60
|
+
|
|
61
|
+
---
|
|
62
|
+
|
|
63
|
+
## 🤖 Supported Model Matrix & Local Cache Status
|
|
64
|
+
|
|
65
|
+
The suite evaluates models across two buckets, all optimized in 4-bit quantization for 16 GB Unified Memory:
|
|
66
|
+
- **Bucket 1 (< 4B):** Ultra-lightweight models for fast edge inference.
|
|
67
|
+
- **Bucket 2 (4B – 12B):** Desktop-class models including Google's **Gemma 4 12B** and **Gemma 4 E4B (8B)**.
|
|
68
|
+
|
|
69
|
+
> Check your local cache anytime with `python download_models.py --list`:
|
|
70
|
+
|
|
71
|
+
<!-- MODEL_STATUS_START -->
|
|
72
|
+
|
|
73
|
+
| Status | Bucket | Est. RAM (4-bit) | Hugging Face Model ID |
|
|
74
|
+
|--------------|----------|--------------------|--------------------------------------------------|
|
|
75
|
+
| **[CACHED]** | <4B | ~1.0 GB | `mlx-community/Llama-3.2-1B-Instruct-4bit` |
|
|
76
|
+
| [NOT CACHED] | <4B | ~2.2 GB | `mlx-community/Llama-3.2-3B-Instruct-4bit` |
|
|
77
|
+
| [NOT CACHED] | <4B | ~2.1 GB | `mlx-community/Qwen2.5-3B-Instruct-4bit` |
|
|
78
|
+
| [NOT CACHED] | <4B | ~2.5 GB | `mlx-community/Phi-4-mini-instruct-4bit` |
|
|
79
|
+
| **[CACHED]** | 4B-12B | ~7.5 GB | `mlx-community/gemma-4-12B-it-qat-4bit` |
|
|
80
|
+
| **[CACHED]** | 4B-12B | ~3.5 GB | `mlx-community/gemma-4-e4b-it-OptiQ-4bit` |
|
|
81
|
+
| [NOT CACHED] | 4B-12B | ~4.8 GB | `mlx-community/Qwen2.5-7B-Instruct-4bit` |
|
|
82
|
+
| [NOT CACHED] | 4B-12B | ~5.2 GB | `mlx-community/Llama-3.1-8B-Instruct-4bit` |
|
|
83
|
+
| [NOT CACHED] | 4B-12B | ~5.2 GB | `mlx-community/DeepSeek-R1-Distill-Qwen-8B-4bit` |
|
|
84
|
+
|
|
85
|
+
<!-- MODEL_STATUS_END -->
|
|
86
|
+
|
|
87
|
+
---
|
|
88
|
+
|
|
89
|
+
## 🚀 Quickstart
|
|
90
|
+
|
|
91
|
+
### Prerequisites
|
|
92
|
+
- macOS on Apple Silicon (M-series).
|
|
93
|
+
- Homebrew installed (`/opt/homebrew/bin/brew`).
|
|
94
|
+
- Python 3.11 (`brew install python@3.11`).
|
|
95
|
+
|
|
96
|
+
### One-Command Setup & Benchmark
|
|
97
|
+
Execute the master runner to set up the environment, run unit tests, and launch a benchmark:
|
|
98
|
+
```bash
|
|
99
|
+
chmod +x run.sh
|
|
100
|
+
./run.sh
|
|
101
|
+
```
|
|
102
|
+
|
|
103
|
+
---
|
|
104
|
+
|
|
105
|
+
## 🛠️ CLI Usage Guide
|
|
106
|
+
|
|
107
|
+
### 1. Download / Install Models Without Running Inference
|
|
108
|
+
Download model weights directly to your local Hugging Face cache (`~/.cache/huggingface/hub/`) without loading them into RAM:
|
|
109
|
+
```bash
|
|
110
|
+
# Check cache status of all models:
|
|
111
|
+
.venv/bin/python3 download_models.py --list
|
|
112
|
+
|
|
113
|
+
# Download a specific model (e.g. Gemma 4 12B):
|
|
114
|
+
.venv/bin/python3 download_models.py --models mlx-community/gemma-4-12B-it-qat-4bit
|
|
115
|
+
|
|
116
|
+
# Download all models in a bucket:
|
|
117
|
+
.venv/bin/python3 download_models.py --bucket "<4B"
|
|
118
|
+
```
|
|
119
|
+
|
|
120
|
+
### 2. Run Comparative Token Benchmarks
|
|
121
|
+
Run side-by-side comparative evaluation between human-first HTML and agent-first JSON:
|
|
122
|
+
```bash
|
|
123
|
+
# Benchmark specific models:
|
|
124
|
+
.venv/bin/python3 benchmark_runner.py --models mlx-community/gemma-4-e4b-it-OptiQ-4bit mlx-community/Llama-3.2-1B-Instruct-4bit
|
|
125
|
+
|
|
126
|
+
# Benchmark a specific game title:
|
|
127
|
+
.venv/bin/python3 benchmark_runner.py --game "Cyberpunk 2077"
|
|
128
|
+
|
|
129
|
+
# Enable live web scraping from https://www.storage4gaming.com:
|
|
130
|
+
.venv/bin/python3 benchmark_runner.py --live
|
|
131
|
+
```
|
|
132
|
+
|
|
133
|
+
### 3. Scan Local Steam Game Catalog
|
|
134
|
+
Discover locally installed games and owned account library titles from Steam (Windows & macOS):
|
|
135
|
+
```bash
|
|
136
|
+
# Standard discovery:
|
|
137
|
+
.venv/bin/python3 catalog_scanner.py
|
|
138
|
+
|
|
139
|
+
# Test with synthetic sample catalog:
|
|
140
|
+
.venv/bin/python3 catalog_scanner.py --sample
|
|
141
|
+
```
|
|
142
|
+
|
|
143
|
+
### 4. Steam Storefront Specs & Hardware Capacity Calculator
|
|
144
|
+
Query the Steam Store API vs. storefront HTML, and calculate total library storage and RAM:
|
|
145
|
+
```bash
|
|
146
|
+
# Benchmark single Steam game (e.g. Apex Legends):
|
|
147
|
+
.venv/bin/python3 steam_client.py --appid 1172470
|
|
148
|
+
|
|
149
|
+
# Calculate aggregate storage & RAM for your entire scanned Steam library:
|
|
150
|
+
.venv/bin/python3 steam_client.py --scan
|
|
151
|
+
```
|
|
152
|
+
|
|
153
|
+
### 5. Update Documentation Automatically
|
|
154
|
+
Keep `README.md` and `ARCHITECTURE.md` updated with the latest SQLite benchmark records and model cache status:
|
|
155
|
+
```bash
|
|
156
|
+
.venv/bin/python3 doc_updater.py
|
|
157
|
+
```
|
|
158
|
+
|
|
159
|
+
---
|
|
160
|
+
|
|
161
|
+
## 🗄️ Database & Telemetry Inspection
|
|
162
|
+
|
|
163
|
+
All metrics are stored in SQLite (`webai_benchmarks.db`) and CSV (`benchmark_results.csv`).
|
|
164
|
+
|
|
165
|
+
### Query Recent Runs
|
|
166
|
+
```bash
|
|
167
|
+
sqlite3 webai_benchmarks.db "SELECT model_id, site_format, prompt_tokens, token_savings_pct, total_time_sec, peak_memory_mb FROM benchmarks ORDER BY id DESC LIMIT 6;"
|
|
168
|
+
```
|
|
169
|
+
|
|
170
|
+
### View CSV Log
|
|
171
|
+
```bash
|
|
172
|
+
cat benchmark_results.csv
|
|
173
|
+
```
|
|
174
|
+
|
|
175
|
+
---
|
|
176
|
+
|
|
177
|
+
## 📚 Central Documentation (`docs/`)
|
|
178
|
+
|
|
179
|
+
All architectural designs, research, and design decision logs are organized in the [`docs/`](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs) folder:
|
|
180
|
+
* **[docs/decisions.md](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs/decisions.md)**: Design log covering Skills vs. Scripts, `llms.txt`, local filesystem discovery, PyPI packaging, readability scoring metrics, and the AIAID protocol.
|
|
181
|
+
* **[docs/research.md](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs/research.md)**: Research on publishing platforms (GitHub, Hugging Face Datasets & Spaces, Model Context Protocol / MCP, `llms.txt` directories, and PyPI).
|
|
182
|
+
* **[docs/ARCHITECTURE.md](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs/ARCHITECTURE.md)**: Full system design, Mermaid data flow, AIAID authentication handshake, and Storage4gaming rewrite blueprint.
|
|
183
|
+
* **[docs/WebAI_specification.md](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs/WebAI_specification.md)**: Master project specification and criteria.
|
|
184
|
+
* **[docs/starter.md](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs/starter.md)**: Task requirements and model testing matrix.
|
|
185
|
+
|
|
186
|
+
---
|
|
187
|
+
|
|
188
|
+
## 📂 Repository Layout
|
|
189
|
+
|
|
190
|
+
```
|
|
191
|
+
├── docs/ # Central documentation folder
|
|
192
|
+
│ ├── decisions.md # Pending design decisions and architectural trade-offs
|
|
193
|
+
│ ├── research.md # Research on publishing platforms (GitHub, Hugging Face, MCP)
|
|
194
|
+
│ ├── ARCHITECTURE.md # Technical design, data flow, and database schema
|
|
195
|
+
│ ├── WebAI_specification.md # Core project specification
|
|
196
|
+
│ └── starter.md # Task requirements and model evaluation matrix
|
|
197
|
+
├── llms.txt # Standard machine-readable AI agent spec index
|
|
198
|
+
├── llms-full.txt # Comprehensive agent knowledge base and sizing formulas
|
|
199
|
+
├── README.md # Project overview, manifesto, quickstart, and live scoreboard
|
|
200
|
+
├── ARCHITECTURE.md # Root link to technical design
|
|
201
|
+
├── telemetry_db.py # SQLite telemetry database engine and CSV exporter
|
|
202
|
+
├── storage4gaming_client.py # Storage4gaming comparative payload generator (HTML vs JSON vs llms.txt)
|
|
203
|
+
├── steam_client.py # Steam Storefront API client & aggregate hardware calculator
|
|
204
|
+
├── benchmark_runner.py # MLX sequential inference and benchmarking harness
|
|
205
|
+
├── download_models.py # Standalone model downloader/installer
|
|
206
|
+
├── catalog_scanner.py # Cross-platform Steam game library detector (Windows & macOS)
|
|
207
|
+
├── doc_updater.py # Documentation auto-synchronizer
|
|
208
|
+
├── test_telemetry.py # Automated test suite for database and telemetry
|
|
209
|
+
├── run.sh # Master setup and execution script
|
|
210
|
+
├── webai_benchmarks.db # SQLite database storing benchmark records
|
|
211
|
+
└── benchmark_results.csv # Exported benchmark CSV dataset
|
|
212
|
+
```
|
|
@@ -0,0 +1,203 @@
|
|
|
1
|
+
# WebAI: The AI-Ready Internet Standard & Token Benchmarking Suite
|
|
2
|
+
|
|
3
|
+
**WebAI** is an open architectural standard, benchmarking suite, and agent identity framework designed to establish an **AI-Agent-Ready Internet** operating alongside the modern human-first web.
|
|
4
|
+
|
|
5
|
+
---
|
|
6
|
+
|
|
7
|
+
## 🌍 The WebAI Manifesto: Why Build This Now?
|
|
8
|
+
|
|
9
|
+
* **The 57% Reality:** Major internet infrastructure providers like Cloudflare have documented that **over 57% of all web traffic** is now generated by automated systems and AI agents.
|
|
10
|
+
* **The Structural Flaw:** Despite AI being the majority of internet consumers, the web remains engineered exclusively for human visual perception (deeply nested HTML DOM trees, CSS stylesheets, layout wrappers, and tracking scripts).
|
|
11
|
+
* **The Downstream Impact:** AI agents waste up to **95%+ of their I/O token budgets** simply parsing presentation boilerplate. This inflates task completion times, drives up inference costs, and needlessly ties up global GPU clusters that could be solving valuable societal tasks.
|
|
12
|
+
* **The Dual-Web Solution:** Humans keep their rich visual web interfaces, while websites provide a parallel **Agent-Ready Web** with zero presentation bloat.
|
|
13
|
+
|
|
14
|
+
### The Two Pillars of WebAI
|
|
15
|
+
1. **Pillar 1: Content & Representation Standards**
|
|
16
|
+
* Standardized `llms.txt` and semantic JSON contracts.
|
|
17
|
+
* **Storage4gaming.com:** Complete rewrite into the reference AI-agent-ready site.
|
|
18
|
+
* **Empirical Benchmarks:** Quantifying token and latency savings across open-weight models on Storage4gaming and the Steam Storefront.
|
|
19
|
+
2. **Pillar 2: AI Agent Identity & Authentication (AIAID)**
|
|
20
|
+
* Cryptographic machine identities for AI agents (analogous to MAC addresses / UUIDs).
|
|
21
|
+
* A public depository and challenge-response handshake that unlocks high-speed semantic endpoints for verified agents while filtering malicious scrapers.
|
|
22
|
+
|
|
23
|
+
---
|
|
24
|
+
|
|
25
|
+
## 🌟 Key Highlights
|
|
26
|
+
|
|
27
|
+
- **~76% - 95% Token Reduction:** Cuts prompt tokens from thousands of tokens down to ~100 tokens.
|
|
28
|
+
- **1.5x - 2.5x Faster Latency:** Accelerates task completion by stripping away visual rendering layers.
|
|
29
|
+
- **Embedded SQLite Telemetry:** All benchmark runs, tokenizer token counts, throughput (tok/s), latency, and peak RAM are persisted to `webai_benchmarks.db` and synchronized with `benchmark_results.csv`.
|
|
30
|
+
- **Apple Silicon Optimized:** 100% offline inference on Apple Silicon using `mlx-lm` with sequential execution and automatic memory cache clearing to protect 16 GB Unified Memory.
|
|
31
|
+
- **Auto-Updating Documentation:** Live benchmark tables and model cache statuses in this README and `ARCHITECTURE.md` stay automatically synchronized with your local benchmark runs.
|
|
32
|
+
|
|
33
|
+
---
|
|
34
|
+
|
|
35
|
+
## 📊 Live Benchmark Performance Summary
|
|
36
|
+
|
|
37
|
+
> The table below is automatically synchronized with [`webai_benchmarks.db`](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/webai_benchmarks.db) by `doc_updater.py`:
|
|
38
|
+
|
|
39
|
+
<!-- BENCHMARK_TABLE_START -->
|
|
40
|
+
|
|
41
|
+
| Model | Target Game | Human Tokens | Agent Tokens | Token Savings | Human Latency | Agent Latency | Speedup |
|
|
42
|
+
|------------------------------|------------------|----------------|----------------|-----------------|-----------------|-----------------|-------------|
|
|
43
|
+
| `Llama-3.2-1B-Instruct-4bit` | Apex Legends | 565 tok | 132 tok | **76.64%** | 0.56s | 0.42s | 1.3x faster |
|
|
44
|
+
| `Llama-3.2-1B-Instruct-4bit` | Counter-Strike 2 | 567 tok | 136 tok | **76.01%** | 0.99s | 0.45s | 2.2x faster |
|
|
45
|
+
| `Llama-3.2-1B-Instruct-4bit` | Halo Infinite | 565 tok | 132 tok | **76.64%** | 0.50s | 0.42s | 1.2x faster |
|
|
46
|
+
| `gemma-4-e4b-it-OptiQ-4bit` | Apex Legends | 635 tok | 125 tok | **80.31%** | 4.69s | 3.45s | 1.4x faster |
|
|
47
|
+
| `gemma-4-e4b-it-OptiQ-4bit` | Counter-Strike 2 | 637 tok | 132 tok | **79.28%** | 6.96s | 4.77s | 1.5x faster |
|
|
48
|
+
| `gemma-4-e4b-it-OptiQ-4bit` | Halo Infinite | 636 tok | 126 tok | **80.19%** | 3.61s | 3.31s | 1.1x faster |
|
|
49
|
+
|
|
50
|
+
<!-- BENCHMARK_TABLE_END -->
|
|
51
|
+
|
|
52
|
+
---
|
|
53
|
+
|
|
54
|
+
## 🤖 Supported Model Matrix & Local Cache Status
|
|
55
|
+
|
|
56
|
+
The suite evaluates models across two buckets, all optimized in 4-bit quantization for 16 GB Unified Memory:
|
|
57
|
+
- **Bucket 1 (< 4B):** Ultra-lightweight models for fast edge inference.
|
|
58
|
+
- **Bucket 2 (4B – 12B):** Desktop-class models including Google's **Gemma 4 12B** and **Gemma 4 E4B (8B)**.
|
|
59
|
+
|
|
60
|
+
> Check your local cache anytime with `python download_models.py --list`:
|
|
61
|
+
|
|
62
|
+
<!-- MODEL_STATUS_START -->
|
|
63
|
+
|
|
64
|
+
| Status | Bucket | Est. RAM (4-bit) | Hugging Face Model ID |
|
|
65
|
+
|--------------|----------|--------------------|--------------------------------------------------|
|
|
66
|
+
| **[CACHED]** | <4B | ~1.0 GB | `mlx-community/Llama-3.2-1B-Instruct-4bit` |
|
|
67
|
+
| [NOT CACHED] | <4B | ~2.2 GB | `mlx-community/Llama-3.2-3B-Instruct-4bit` |
|
|
68
|
+
| [NOT CACHED] | <4B | ~2.1 GB | `mlx-community/Qwen2.5-3B-Instruct-4bit` |
|
|
69
|
+
| [NOT CACHED] | <4B | ~2.5 GB | `mlx-community/Phi-4-mini-instruct-4bit` |
|
|
70
|
+
| **[CACHED]** | 4B-12B | ~7.5 GB | `mlx-community/gemma-4-12B-it-qat-4bit` |
|
|
71
|
+
| **[CACHED]** | 4B-12B | ~3.5 GB | `mlx-community/gemma-4-e4b-it-OptiQ-4bit` |
|
|
72
|
+
| [NOT CACHED] | 4B-12B | ~4.8 GB | `mlx-community/Qwen2.5-7B-Instruct-4bit` |
|
|
73
|
+
| [NOT CACHED] | 4B-12B | ~5.2 GB | `mlx-community/Llama-3.1-8B-Instruct-4bit` |
|
|
74
|
+
| [NOT CACHED] | 4B-12B | ~5.2 GB | `mlx-community/DeepSeek-R1-Distill-Qwen-8B-4bit` |
|
|
75
|
+
|
|
76
|
+
<!-- MODEL_STATUS_END -->
|
|
77
|
+
|
|
78
|
+
---
|
|
79
|
+
|
|
80
|
+
## 🚀 Quickstart
|
|
81
|
+
|
|
82
|
+
### Prerequisites
|
|
83
|
+
- macOS on Apple Silicon (M-series).
|
|
84
|
+
- Homebrew installed (`/opt/homebrew/bin/brew`).
|
|
85
|
+
- Python 3.11 (`brew install python@3.11`).
|
|
86
|
+
|
|
87
|
+
### One-Command Setup & Benchmark
|
|
88
|
+
Execute the master runner to set up the environment, run unit tests, and launch a benchmark:
|
|
89
|
+
```bash
|
|
90
|
+
chmod +x run.sh
|
|
91
|
+
./run.sh
|
|
92
|
+
```
|
|
93
|
+
|
|
94
|
+
---
|
|
95
|
+
|
|
96
|
+
## 🛠️ CLI Usage Guide
|
|
97
|
+
|
|
98
|
+
### 1. Download / Install Models Without Running Inference
|
|
99
|
+
Download model weights directly to your local Hugging Face cache (`~/.cache/huggingface/hub/`) without loading them into RAM:
|
|
100
|
+
```bash
|
|
101
|
+
# Check cache status of all models:
|
|
102
|
+
.venv/bin/python3 download_models.py --list
|
|
103
|
+
|
|
104
|
+
# Download a specific model (e.g. Gemma 4 12B):
|
|
105
|
+
.venv/bin/python3 download_models.py --models mlx-community/gemma-4-12B-it-qat-4bit
|
|
106
|
+
|
|
107
|
+
# Download all models in a bucket:
|
|
108
|
+
.venv/bin/python3 download_models.py --bucket "<4B"
|
|
109
|
+
```
|
|
110
|
+
|
|
111
|
+
### 2. Run Comparative Token Benchmarks
|
|
112
|
+
Run side-by-side comparative evaluation between human-first HTML and agent-first JSON:
|
|
113
|
+
```bash
|
|
114
|
+
# Benchmark specific models:
|
|
115
|
+
.venv/bin/python3 benchmark_runner.py --models mlx-community/gemma-4-e4b-it-OptiQ-4bit mlx-community/Llama-3.2-1B-Instruct-4bit
|
|
116
|
+
|
|
117
|
+
# Benchmark a specific game title:
|
|
118
|
+
.venv/bin/python3 benchmark_runner.py --game "Cyberpunk 2077"
|
|
119
|
+
|
|
120
|
+
# Enable live web scraping from https://www.storage4gaming.com:
|
|
121
|
+
.venv/bin/python3 benchmark_runner.py --live
|
|
122
|
+
```
|
|
123
|
+
|
|
124
|
+
### 3. Scan Local Steam Game Catalog
|
|
125
|
+
Discover locally installed games and owned account library titles from Steam (Windows & macOS):
|
|
126
|
+
```bash
|
|
127
|
+
# Standard discovery:
|
|
128
|
+
.venv/bin/python3 catalog_scanner.py
|
|
129
|
+
|
|
130
|
+
# Test with synthetic sample catalog:
|
|
131
|
+
.venv/bin/python3 catalog_scanner.py --sample
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
### 4. Steam Storefront Specs & Hardware Capacity Calculator
|
|
135
|
+
Query the Steam Store API vs. storefront HTML, and calculate total library storage and RAM:
|
|
136
|
+
```bash
|
|
137
|
+
# Benchmark single Steam game (e.g. Apex Legends):
|
|
138
|
+
.venv/bin/python3 steam_client.py --appid 1172470
|
|
139
|
+
|
|
140
|
+
# Calculate aggregate storage & RAM for your entire scanned Steam library:
|
|
141
|
+
.venv/bin/python3 steam_client.py --scan
|
|
142
|
+
```
|
|
143
|
+
|
|
144
|
+
### 5. Update Documentation Automatically
|
|
145
|
+
Keep `README.md` and `ARCHITECTURE.md` updated with the latest SQLite benchmark records and model cache status:
|
|
146
|
+
```bash
|
|
147
|
+
.venv/bin/python3 doc_updater.py
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
---
|
|
151
|
+
|
|
152
|
+
## 🗄️ Database & Telemetry Inspection
|
|
153
|
+
|
|
154
|
+
All metrics are stored in SQLite (`webai_benchmarks.db`) and CSV (`benchmark_results.csv`).
|
|
155
|
+
|
|
156
|
+
### Query Recent Runs
|
|
157
|
+
```bash
|
|
158
|
+
sqlite3 webai_benchmarks.db "SELECT model_id, site_format, prompt_tokens, token_savings_pct, total_time_sec, peak_memory_mb FROM benchmarks ORDER BY id DESC LIMIT 6;"
|
|
159
|
+
```
|
|
160
|
+
|
|
161
|
+
### View CSV Log
|
|
162
|
+
```bash
|
|
163
|
+
cat benchmark_results.csv
|
|
164
|
+
```
|
|
165
|
+
|
|
166
|
+
---
|
|
167
|
+
|
|
168
|
+
## 📚 Central Documentation (`docs/`)
|
|
169
|
+
|
|
170
|
+
All architectural designs, research, and design decision logs are organized in the [`docs/`](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs) folder:
|
|
171
|
+
* **[docs/decisions.md](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs/decisions.md)**: Design log covering Skills vs. Scripts, `llms.txt`, local filesystem discovery, PyPI packaging, readability scoring metrics, and the AIAID protocol.
|
|
172
|
+
* **[docs/research.md](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs/research.md)**: Research on publishing platforms (GitHub, Hugging Face Datasets & Spaces, Model Context Protocol / MCP, `llms.txt` directories, and PyPI).
|
|
173
|
+
* **[docs/ARCHITECTURE.md](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs/ARCHITECTURE.md)**: Full system design, Mermaid data flow, AIAID authentication handshake, and Storage4gaming rewrite blueprint.
|
|
174
|
+
* **[docs/WebAI_specification.md](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs/WebAI_specification.md)**: Master project specification and criteria.
|
|
175
|
+
* **[docs/starter.md](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs/starter.md)**: Task requirements and model testing matrix.
|
|
176
|
+
|
|
177
|
+
---
|
|
178
|
+
|
|
179
|
+
## 📂 Repository Layout
|
|
180
|
+
|
|
181
|
+
```
|
|
182
|
+
├── docs/ # Central documentation folder
|
|
183
|
+
│ ├── decisions.md # Pending design decisions and architectural trade-offs
|
|
184
|
+
│ ├── research.md # Research on publishing platforms (GitHub, Hugging Face, MCP)
|
|
185
|
+
│ ├── ARCHITECTURE.md # Technical design, data flow, and database schema
|
|
186
|
+
│ ├── WebAI_specification.md # Core project specification
|
|
187
|
+
│ └── starter.md # Task requirements and model evaluation matrix
|
|
188
|
+
├── llms.txt # Standard machine-readable AI agent spec index
|
|
189
|
+
├── llms-full.txt # Comprehensive agent knowledge base and sizing formulas
|
|
190
|
+
├── README.md # Project overview, manifesto, quickstart, and live scoreboard
|
|
191
|
+
├── ARCHITECTURE.md # Root link to technical design
|
|
192
|
+
├── telemetry_db.py # SQLite telemetry database engine and CSV exporter
|
|
193
|
+
├── storage4gaming_client.py # Storage4gaming comparative payload generator (HTML vs JSON vs llms.txt)
|
|
194
|
+
├── steam_client.py # Steam Storefront API client & aggregate hardware calculator
|
|
195
|
+
├── benchmark_runner.py # MLX sequential inference and benchmarking harness
|
|
196
|
+
├── download_models.py # Standalone model downloader/installer
|
|
197
|
+
├── catalog_scanner.py # Cross-platform Steam game library detector (Windows & macOS)
|
|
198
|
+
├── doc_updater.py # Documentation auto-synchronizer
|
|
199
|
+
├── test_telemetry.py # Automated test suite for database and telemetry
|
|
200
|
+
├── run.sh # Master setup and execution script
|
|
201
|
+
├── webai_benchmarks.db # SQLite database storing benchmark records
|
|
202
|
+
└── benchmark_results.csv # Exported benchmark CSV dataset
|
|
203
|
+
```
|