webai-scanner 0.1.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,228 @@
1
+ # WebAI Architecture & System Design
2
+
3
+ **WebAI** is an open architectural standard and benchmarking suite designed to establish an **AI Agent Readability & Identity Standard** for the modern web.
4
+
5
+ ---
6
+
7
+ ## 1. The Core Problem: Human-First vs. Agent-First Web
8
+
9
+ ### The 57% Reality
10
+ Cloudflare and major internet infrastructure providers have documented that **over 57% of all web traffic** is generated by automated bots and AI agents. Despite this majority, the internet remains engineered exclusively **human-first**:
11
+ - Web pages are dominated by presentation markup: nested DOM `<div>`s, CSS rules, stylesheets, layout boilerplate, navigation bars, cookie banners, tracking scripts, and SVGs.
12
+ - A human views only the rendered graphical output. However, an AI agent interacting with the site must ingest the entire raw DOM payload into its transformer context window.
13
+ - As an example, the Storage4gaming calculator page contains over **60,000 lines of HTML (~3.2 MB)**. Feeding this markup to an LLM wastes thousands of input tokens and imposes significant latency and compute overhead.
14
+ - This creates massive downstream economic and societal waste: compute cycles that could solve complex tasks are instead squandered reading visual layout code.
15
+
16
+ ### WebAI's Dual-Web Solution
17
+ WebAI establishes an **agent-ready web operating in parallel with the human web**:
18
+ - Essential factual data is surfaced through compact JSON, semantic JSON-LD, or structured Markdown (e.g. `llms.txt`).
19
+ - A 60k-line DOM is distilled down into a ~100-byte semantic object.
20
+ - **Result:** ~75–95% input token reduction, 1.5x–2.5x speedup in agent inference, and zero loss of extraction accuracy.
21
+
22
+ ```
23
+ ┌────────────────────────────────────────────────────────┐
24
+ │ Human-First Web: ~1,800 - 3,200,000 bytes │
25
+ │ [DOM wrappers, styling, tracking, UI boilerplate] │
26
+ │ ↳ Ingested by Agent = 600 - 50,000+ tokens │
27
+ └────────────────────────────────────────────────────────┘
28
+ vs
29
+ ┌────────────────────────────────────────────────────────┐
30
+ │ Agent-First Web: ~100 bytes │
31
+ │ [Structured JSON / Semantic Key-Values / llms.txt] │
32
+ │ ↳ Ingested by Agent = ~100 - 150 tokens │
33
+ │ ↳ 75% - 95% Token Reduction, 2x+ Latency Improvement │
34
+ └────────────────────────────────────────────────────────┘
35
+ ```
36
+
37
+ ---
38
+
39
+ ## 2. System Architecture & Comparative Data Flow
40
+
41
+ The following diagram illustrates how WebAI evaluates, benchmarks, and logs token telemetry across local open-weight models on Apple Silicon:
42
+
43
+ ```mermaid
44
+ flowchart TD
45
+ subgraph Data Sources & Case Studies
46
+ S4G_HTML[Storage4gaming Live DOM<br/>Human-First Payload: 3.2MB]
47
+ S4G_Agent[Storage4gaming llms.txt & JSON<br/>Agent-First: ~100 bytes]
48
+ Steam_HTML[Steam Storefront Page DOM<br/>Human-First Payload: ~250KB]
49
+ Steam_API[Steam Store API appdetails<br/>Agent-First: ~115 bytes]
50
+ Catalog[Cross-Platform Steam Discovery<br/>Windows & macOS Local Filesystem]
51
+ end
52
+
53
+ subgraph Client & Calculator Layer
54
+ S4G_Client[storage4gaming_client.py<br/>S4G Spec Client & llms.txt Generator]
55
+ Steam_Client[steam_client.py<br/>Steam API Client & Hardware Calculator]
56
+
57
+ S4G_HTML --> S4G_Client
58
+ S4G_Agent --> S4G_Client
59
+ Steam_HTML --> Steam_Client
60
+ Steam_API --> Steam_Client
61
+ Catalog --> Steam_Client
62
+ end
63
+
64
+ subgraph Inference & Benchmarking Engine
65
+ Runner[benchmark_runner.py<br/>MLX-LM Sequential Harness]
66
+ S4G_Client --> Runner
67
+ Steam_Client --> Runner
68
+
69
+ Runner --> M1[Bucket 1: Sub-4B Models<br/>Llama 3.2, Qwen 2.5, Phi 4]
70
+ Runner --> M2[Bucket 2: 4B-12B Models<br/>Gemma 4 12B, Gemma 4 E4B, DeepSeek R1]
71
+ end
72
+
73
+ subgraph Telemetry & Persistence Layer
74
+ Runner --> Telemetry[telemetry_db.py<br/>SQLite Telemetry Engine]
75
+ Telemetry --> DB[(webai_benchmarks.db)]
76
+ Telemetry --> CSV[benchmark_results.csv]
77
+ Telemetry --> DocSync[doc_updater.py<br/>Auto-Doc Sync]
78
+ DocSync --> Readme[README.md & ARCHITECTURE.md]
79
+ end
80
+ ```
81
+
82
+ ---
83
+
84
+ ## 3. The AIAID (AI Agent Identity) Architecture
85
+
86
+ To bring transparency and trusted authentication to agent web traversal, WebAI introduces the **AIAID (AI Agent ID)** protocol.
87
+
88
+ ### Why AIAID?
89
+ Today, devices have MAC addresses and serial numbers. However, AI agents traverse the internet anonymously, leading websites to deploy aggressive bot blockers (Cloudflare challenge walls, CAPTCHAs).
90
+
91
+ With AIAID:
92
+ 1. Legitimate agents identify themselves transparently.
93
+ 2. Participating sites grant authenticated agents **instant access to high-density semantic endpoints (`llms.txt`, JSON API)** without CAPTCHAs.
94
+ 3. Unverified scrapers continue to receive standard human-facing HTML or bot challenge walls.
95
+
96
+ ```mermaid
97
+ sequenceDiagram
98
+ autonumber
99
+ actor User as Human User
100
+ participant Agent as AI Agent (with AIAID)
101
+ participant Site as WebAI-Ready Site (Storage4gaming)
102
+ participant Registry as Public AIAID Depository / Registry
103
+
104
+ User->>Agent: "Calculate required storage for my Steam library"
105
+ Agent->>Site: HTTP GET /calculator (Header: X-AIAID: aiaid_7f9c2b...)
106
+ Note over Site: Site detects Agent seeking Agent-First Endpoint
107
+ Site->>Registry: Verify AIAID(aiaid_7f9c2b...)
108
+ Registry->>Agent: Issue Cryptographic Challenge Nonce
109
+ Agent-->>Registry: Signed Nonce Response (Private Key)
110
+ Registry-->>Site: AIAID Verified: Active, Legitimate, Tier: Standard
111
+ Site-->>Agent: HTTP 200 OK (Clean Semantic JSON / llms.txt)
112
+ Note over Agent: Consumes ~100 tokens in 0.4s!
113
+ Agent-->>User: "You need a 1TB NVMe SSD and 16GB RAM."
114
+ ```
115
+
116
+ ---
117
+
118
+ ## 4. Storage4gaming Rewrite Blueprint (The Reference Site)
119
+
120
+ As the founding real-world testbed of WebAI, **Storage4gaming.com** will undergo a complete code rewrite into an **AI-Agent-Ready Architecture**:
121
+
122
+ ```
123
+ storage4gaming.com/
124
+ ├── Human Interface (Web Browser)
125
+ │ ├── /calculator/ -> Interactive Visual Hardware Sizing Tool (React / Tailwind)
126
+ │ ├── /games/[slug]/ -> Human-readable hardware review with interactive charts
127
+ │ └── /guides/ -> SSD & RAM buying guides
128
+ │
129
+ ├── Agent Interface (AI Ready)
130
+ │ ├── /llms.txt -> Lightweight Markdown index of verified game hardware specs
131
+ │ ├── /llms-full.txt -> Complete agent knowledge base with sizing formulas
132
+ │ ├── /api/games/[slug].json -> Ultra-compact JSON payload (~100 bytes)
133
+ │ └── /api/auth/aiaid -> Verification endpoint for AIAID agent handshake
134
+ │
135
+ └── Smart Content Negotiation Layer
136
+ └── When Accept: application/json or X-AIAID is present -> Auto-routes to Agent Interface
137
+ ```
138
+
139
+ ---
140
+
141
+ ## 5. Component Breakdown
142
+
143
+ ### A. Site Integration Client (`storage4gaming_client.py`)
144
+ - Interfaces with Storage4gaming (`https://www.storage4gaming.com`).
145
+ - Generates three parallel representations for PC game titles:
146
+ 1. `human_first_html`: Real-world DOM cards with CSS layout classes, badge wrappers, and page context (~1,800 to 3,200,000 bytes).
147
+ 2. `agent_first_json`: High-density semantic JSON containing only verified hardware metrics (~100 bytes).
148
+ 3. `agent_first_llmstxt`: Pure markdown entry adhering to the `llms.txt` standard (~120 bytes).
149
+
150
+ ### B. Steam Store Client & Calculator (`steam_client.py`)
151
+ - Interfaces with the official Steam Store API (`store.steampowered.com/api/appdetails`).
152
+ - Evaluates the Steam Storefront desktop HTML page against the official JSON API.
153
+ - Implements the **Hardware Capacity Calculator**:
154
+ - Parses numeric RAM (GB), Storage (GB), and Drive Type (NVMe SSD, SSD, HDD).
155
+ - Computes raw storage footprint across all discovered games.
156
+ - Adds 20% safety margin for patches and swap files.
157
+ - Recommends commercial SSD drive sizing (500GB, 1TB, 2TB, 4TB).
158
+ - Recommends system RAM tier (16GB vs 32GB).
159
+
160
+ ### C. Universal MLX Benchmark Harness (`benchmark_runner.py`)
161
+ - Built on Apple's `mlx-lm` framework for native Apple Silicon unified memory acceleration.
162
+ - Enforces **strict sequential execution** (`batch_size=1`, single model resident in RAM at any time).
163
+ - Between model runs, it calls `del model`, `del tokenizer`, `gc.collect()`, and `mx.clear_cache()` to prevent memory leakage.
164
+ - Captures exact tokenizer token counts, prompt eval speed (tok/sec), generation latency, and peak memory.
165
+
166
+ ### D. Telemetry Engine & SQL Database (`telemetry_db.py`)
167
+ - Embedded SQLite engine requiring zero external daemon or background service.
168
+ - Database file: `webai_benchmarks.db`.
169
+ - Automatically computes `token_savings_pct` by pairing the agent-first run against the human-first baseline.
170
+ - Automatically synchronizes all database records to `benchmark_results.csv`.
171
+
172
+ ### E. Catalog Scanner (`catalog_scanner.py`)
173
+ - Offline, zero-auth Steam discovery scanner with native support for both **macOS** and **Windows**:
174
+ - **macOS:** Scans `~/Library/Application Support/Steam/steamapps/libraryfolders.vdf`, `userdata/<id>/config/localconfig.vdf`, and binary `appcache/appinfo.vdf`.
175
+ - **Windows:** Auto-detects Steam via Windows Registry (`HKCU\Software\Valve\Steam\SteamPath`) or `C:\Program Files (x86)\Steam`, multi-drive library folders, user account cache, and `appcache\appinfo.vdf`.
176
+
177
+ ### F. Documentation Auto-Updater (`doc_updater.py`)
178
+ - Automatically reads live data from `webai_benchmarks.db` and the Hugging Face cache.
179
+ - Directly updates the markdown tables in `README.md` and `ARCHITECTURE.md` to keep documentation synchronised with real experimental results.
180
+
181
+ ---
182
+
183
+ ## 6. Live Benchmark Performance Summary
184
+
185
+ The table below is dynamically synchronized with the local SQLite database (`webai_benchmarks.db`):
186
+
187
+ <!-- BENCHMARK_TABLE_START -->
188
+
189
+ | Model | Target Game | Human Tokens | Agent Tokens | Token Savings | Human Latency | Agent Latency | Speedup |
190
+ |------------------------------|------------------|----------------|----------------|-----------------|-----------------|-----------------|-------------|
191
+ | `Llama-3.2-1B-Instruct-4bit` | Apex Legends | 565 tok | 132 tok | **76.64%** | 0.56s | 0.42s | 1.3x faster |
192
+ | `Llama-3.2-1B-Instruct-4bit` | Counter-Strike 2 | 567 tok | 136 tok | **76.01%** | 0.99s | 0.45s | 2.2x faster |
193
+ | `Llama-3.2-1B-Instruct-4bit` | Halo Infinite | 565 tok | 132 tok | **76.64%** | 0.50s | 0.42s | 1.2x faster |
194
+ | `gemma-4-e4b-it-OptiQ-4bit` | Apex Legends | 635 tok | 125 tok | **80.31%** | 4.69s | 3.45s | 1.4x faster |
195
+ | `gemma-4-e4b-it-OptiQ-4bit` | Counter-Strike 2 | 637 tok | 132 tok | **79.28%** | 6.96s | 4.77s | 1.5x faster |
196
+ | `gemma-4-e4b-it-OptiQ-4bit` | Halo Infinite | 636 tok | 126 tok | **80.19%** | 3.61s | 3.31s | 1.1x faster |
197
+
198
+ <!-- BENCHMARK_TABLE_END -->
199
+
200
+ ---
201
+
202
+ ## 7. Model Evaluation Matrix
203
+
204
+ <!-- MODEL_STATUS_START -->
205
+
206
+ | Status | Bucket | Est. RAM (4-bit) | Hugging Face Model ID |
207
+ |--------------|----------|--------------------|--------------------------------------------------|
208
+ | **[CACHED]** | <4B | ~1.0 GB | `mlx-community/Llama-3.2-1B-Instruct-4bit` |
209
+ | [NOT CACHED] | <4B | ~2.2 GB | `mlx-community/Llama-3.2-3B-Instruct-4bit` |
210
+ | [NOT CACHED] | <4B | ~2.1 GB | `mlx-community/Qwen2.5-3B-Instruct-4bit` |
211
+ | [NOT CACHED] | <4B | ~2.5 GB | `mlx-community/Phi-4-mini-instruct-4bit` |
212
+ | **[CACHED]** | 4B-12B | ~7.5 GB | `mlx-community/gemma-4-12B-it-qat-4bit` |
213
+ | **[CACHED]** | 4B-12B | ~3.5 GB | `mlx-community/gemma-4-e4b-it-OptiQ-4bit` |
214
+ | [NOT CACHED] | 4B-12B | ~4.8 GB | `mlx-community/Qwen2.5-7B-Instruct-4bit` |
215
+ | [NOT CACHED] | 4B-12B | ~5.2 GB | `mlx-community/Llama-3.1-8B-Instruct-4bit` |
216
+ | [NOT CACHED] | 4B-12B | ~5.2 GB | `mlx-community/DeepSeek-R1-Distill-Qwen-8B-4bit` |
217
+
218
+ <!-- MODEL_STATUS_END -->
219
+
220
+ ---
221
+
222
+ ## 8. Hardware & Memory Profile (16 GB Unified RAM)
223
+
224
+ The WebAI suite is strictly engineered for Apple Silicon (macOS) with 16 GB unified memory:
225
+ - **macOS & System Reserve:** ~3.5 - 4.5 GB RAM.
226
+ - **Sub-4B Models:** ~0.9 - 2.5 GB peak active memory.
227
+ - **Medium (4B - 12B) Models:** ~3.5 - 7.5 GB peak active memory.
228
+ - **Headroom:** Always maintains >= 4 GB of free headroom, preventing swap file churn and thermal throttling.
@@ -0,0 +1,212 @@
1
+ Metadata-Version: 2.5
2
+ Name: webai-scanner
3
+ Version: 0.1.0
4
+ Summary: AI Agent Readability auditor & local Steam game discovery engine
5
+ Author: Alexander Markovski
6
+ License: MIT
7
+ Requires-Python: >=3.10
8
+ Description-Content-Type: text/markdown
9
+
10
+ # WebAI: The AI-Ready Internet Standard & Token Benchmarking Suite
11
+
12
+ **WebAI** is an open architectural standard, benchmarking suite, and agent identity framework designed to establish an **AI-Agent-Ready Internet** operating alongside the modern human-first web.
13
+
14
+ ---
15
+
16
+ ## 🌍 The WebAI Manifesto: Why Build This Now?
17
+
18
+ * **The 57% Reality:** Major internet infrastructure providers like Cloudflare have documented that **over 57% of all web traffic** is now generated by automated systems and AI agents.
19
+ * **The Structural Flaw:** Despite AI being the majority of internet consumers, the web remains engineered exclusively for human visual perception (deeply nested HTML DOM trees, CSS stylesheets, layout wrappers, and tracking scripts).
20
+ * **The Downstream Impact:** AI agents waste up to **95%+ of their I/O token budgets** simply parsing presentation boilerplate. This inflates task completion times, drives up inference costs, and needlessly ties up global GPU clusters that could be solving valuable societal tasks.
21
+ * **The Dual-Web Solution:** Humans keep their rich visual web interfaces, while websites provide a parallel **Agent-Ready Web** with zero presentation bloat.
22
+
23
+ ### The Two Pillars of WebAI
24
+ 1. **Pillar 1: Content & Representation Standards**
25
+ * Standardized `llms.txt` and semantic JSON contracts.
26
+ * **Storage4gaming.com:** Complete rewrite into the reference AI-agent-ready site.
27
+ * **Empirical Benchmarks:** Quantifying token and latency savings across open-weight models on Storage4gaming and the Steam Storefront.
28
+ 2. **Pillar 2: AI Agent Identity & Authentication (AIAID)**
29
+ * Cryptographic machine identities for AI agents (analogous to MAC addresses / UUIDs).
30
+ * A public depository and challenge-response handshake that unlocks high-speed semantic endpoints for verified agents while filtering malicious scrapers.
31
+
32
+ ---
33
+
34
+ ## 🌟 Key Highlights
35
+
36
+ - **~76% - 95% Token Reduction:** Cuts prompt tokens from thousands of tokens down to ~100 tokens.
37
+ - **1.5x - 2.5x Faster Latency:** Accelerates task completion by stripping away visual rendering layers.
38
+ - **Embedded SQLite Telemetry:** All benchmark runs, tokenizer token counts, throughput (tok/s), latency, and peak RAM are persisted to `webai_benchmarks.db` and synchronized with `benchmark_results.csv`.
39
+ - **Apple Silicon Optimized:** 100% offline inference on Apple Silicon using `mlx-lm` with sequential execution and automatic memory cache clearing to protect 16 GB Unified Memory.
40
+ - **Auto-Updating Documentation:** Live benchmark tables and model cache statuses in this README and `ARCHITECTURE.md` stay automatically synchronized with your local benchmark runs.
41
+
42
+ ---
43
+
44
+ ## 📊 Live Benchmark Performance Summary
45
+
46
+ > The table below is automatically synchronized with [`webai_benchmarks.db`](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/webai_benchmarks.db) by `doc_updater.py`:
47
+
48
+ <!-- BENCHMARK_TABLE_START -->
49
+
50
+ | Model | Target Game | Human Tokens | Agent Tokens | Token Savings | Human Latency | Agent Latency | Speedup |
51
+ |------------------------------|------------------|----------------|----------------|-----------------|-----------------|-----------------|-------------|
52
+ | `Llama-3.2-1B-Instruct-4bit` | Apex Legends | 565 tok | 132 tok | **76.64%** | 0.56s | 0.42s | 1.3x faster |
53
+ | `Llama-3.2-1B-Instruct-4bit` | Counter-Strike 2 | 567 tok | 136 tok | **76.01%** | 0.99s | 0.45s | 2.2x faster |
54
+ | `Llama-3.2-1B-Instruct-4bit` | Halo Infinite | 565 tok | 132 tok | **76.64%** | 0.50s | 0.42s | 1.2x faster |
55
+ | `gemma-4-e4b-it-OptiQ-4bit` | Apex Legends | 635 tok | 125 tok | **80.31%** | 4.69s | 3.45s | 1.4x faster |
56
+ | `gemma-4-e4b-it-OptiQ-4bit` | Counter-Strike 2 | 637 tok | 132 tok | **79.28%** | 6.96s | 4.77s | 1.5x faster |
57
+ | `gemma-4-e4b-it-OptiQ-4bit` | Halo Infinite | 636 tok | 126 tok | **80.19%** | 3.61s | 3.31s | 1.1x faster |
58
+
59
+ <!-- BENCHMARK_TABLE_END -->
60
+
61
+ ---
62
+
63
+ ## 🤖 Supported Model Matrix & Local Cache Status
64
+
65
+ The suite evaluates models across two buckets, all optimized in 4-bit quantization for 16 GB Unified Memory:
66
+ - **Bucket 1 (< 4B):** Ultra-lightweight models for fast edge inference.
67
+ - **Bucket 2 (4B – 12B):** Desktop-class models including Google's **Gemma 4 12B** and **Gemma 4 E4B (8B)**.
68
+
69
+ > Check your local cache anytime with `python download_models.py --list`:
70
+
71
+ <!-- MODEL_STATUS_START -->
72
+
73
+ | Status | Bucket | Est. RAM (4-bit) | Hugging Face Model ID |
74
+ |--------------|----------|--------------------|--------------------------------------------------|
75
+ | **[CACHED]** | <4B | ~1.0 GB | `mlx-community/Llama-3.2-1B-Instruct-4bit` |
76
+ | [NOT CACHED] | <4B | ~2.2 GB | `mlx-community/Llama-3.2-3B-Instruct-4bit` |
77
+ | [NOT CACHED] | <4B | ~2.1 GB | `mlx-community/Qwen2.5-3B-Instruct-4bit` |
78
+ | [NOT CACHED] | <4B | ~2.5 GB | `mlx-community/Phi-4-mini-instruct-4bit` |
79
+ | **[CACHED]** | 4B-12B | ~7.5 GB | `mlx-community/gemma-4-12B-it-qat-4bit` |
80
+ | **[CACHED]** | 4B-12B | ~3.5 GB | `mlx-community/gemma-4-e4b-it-OptiQ-4bit` |
81
+ | [NOT CACHED] | 4B-12B | ~4.8 GB | `mlx-community/Qwen2.5-7B-Instruct-4bit` |
82
+ | [NOT CACHED] | 4B-12B | ~5.2 GB | `mlx-community/Llama-3.1-8B-Instruct-4bit` |
83
+ | [NOT CACHED] | 4B-12B | ~5.2 GB | `mlx-community/DeepSeek-R1-Distill-Qwen-8B-4bit` |
84
+
85
+ <!-- MODEL_STATUS_END -->
86
+
87
+ ---
88
+
89
+ ## 🚀 Quickstart
90
+
91
+ ### Prerequisites
92
+ - macOS on Apple Silicon (M-series).
93
+ - Homebrew installed (`/opt/homebrew/bin/brew`).
94
+ - Python 3.11 (`brew install python@3.11`).
95
+
96
+ ### One-Command Setup & Benchmark
97
+ Execute the master runner to set up the environment, run unit tests, and launch a benchmark:
98
+ ```bash
99
+ chmod +x run.sh
100
+ ./run.sh
101
+ ```
102
+
103
+ ---
104
+
105
+ ## 🛠️ CLI Usage Guide
106
+
107
+ ### 1. Download / Install Models Without Running Inference
108
+ Download model weights directly to your local Hugging Face cache (`~/.cache/huggingface/hub/`) without loading them into RAM:
109
+ ```bash
110
+ # Check cache status of all models:
111
+ .venv/bin/python3 download_models.py --list
112
+
113
+ # Download a specific model (e.g. Gemma 4 12B):
114
+ .venv/bin/python3 download_models.py --models mlx-community/gemma-4-12B-it-qat-4bit
115
+
116
+ # Download all models in a bucket:
117
+ .venv/bin/python3 download_models.py --bucket "<4B"
118
+ ```
119
+
120
+ ### 2. Run Comparative Token Benchmarks
121
+ Run side-by-side comparative evaluation between human-first HTML and agent-first JSON:
122
+ ```bash
123
+ # Benchmark specific models:
124
+ .venv/bin/python3 benchmark_runner.py --models mlx-community/gemma-4-e4b-it-OptiQ-4bit mlx-community/Llama-3.2-1B-Instruct-4bit
125
+
126
+ # Benchmark a specific game title:
127
+ .venv/bin/python3 benchmark_runner.py --game "Cyberpunk 2077"
128
+
129
+ # Enable live web scraping from https://www.storage4gaming.com:
130
+ .venv/bin/python3 benchmark_runner.py --live
131
+ ```
132
+
133
+ ### 3. Scan Local Steam Game Catalog
134
+ Discover locally installed games and owned account library titles from Steam (Windows & macOS):
135
+ ```bash
136
+ # Standard discovery:
137
+ .venv/bin/python3 catalog_scanner.py
138
+
139
+ # Test with synthetic sample catalog:
140
+ .venv/bin/python3 catalog_scanner.py --sample
141
+ ```
142
+
143
+ ### 4. Steam Storefront Specs & Hardware Capacity Calculator
144
+ Query the Steam Store API vs. storefront HTML, and calculate total library storage and RAM:
145
+ ```bash
146
+ # Benchmark single Steam game (e.g. Apex Legends):
147
+ .venv/bin/python3 steam_client.py --appid 1172470
148
+
149
+ # Calculate aggregate storage & RAM for your entire scanned Steam library:
150
+ .venv/bin/python3 steam_client.py --scan
151
+ ```
152
+
153
+ ### 5. Update Documentation Automatically
154
+ Keep `README.md` and `ARCHITECTURE.md` updated with the latest SQLite benchmark records and model cache status:
155
+ ```bash
156
+ .venv/bin/python3 doc_updater.py
157
+ ```
158
+
159
+ ---
160
+
161
+ ## 🗄️ Database & Telemetry Inspection
162
+
163
+ All metrics are stored in SQLite (`webai_benchmarks.db`) and CSV (`benchmark_results.csv`).
164
+
165
+ ### Query Recent Runs
166
+ ```bash
167
+ sqlite3 webai_benchmarks.db "SELECT model_id, site_format, prompt_tokens, token_savings_pct, total_time_sec, peak_memory_mb FROM benchmarks ORDER BY id DESC LIMIT 6;"
168
+ ```
169
+
170
+ ### View CSV Log
171
+ ```bash
172
+ cat benchmark_results.csv
173
+ ```
174
+
175
+ ---
176
+
177
+ ## 📚 Central Documentation (`docs/`)
178
+
179
+ All architectural designs, research, and design decision logs are organized in the [`docs/`](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs) folder:
180
+ * **[docs/decisions.md](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs/decisions.md)**: Design log covering Skills vs. Scripts, `llms.txt`, local filesystem discovery, PyPI packaging, readability scoring metrics, and the AIAID protocol.
181
+ * **[docs/research.md](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs/research.md)**: Research on publishing platforms (GitHub, Hugging Face Datasets & Spaces, Model Context Protocol / MCP, `llms.txt` directories, and PyPI).
182
+ * **[docs/ARCHITECTURE.md](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs/ARCHITECTURE.md)**: Full system design, Mermaid data flow, AIAID authentication handshake, and Storage4gaming rewrite blueprint.
183
+ * **[docs/WebAI_specification.md](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs/WebAI_specification.md)**: Master project specification and criteria.
184
+ * **[docs/starter.md](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs/starter.md)**: Task requirements and model testing matrix.
185
+
186
+ ---
187
+
188
+ ## 📂 Repository Layout
189
+
190
+ ```
191
+ ├── docs/ # Central documentation folder
192
+ │ ├── decisions.md # Pending design decisions and architectural trade-offs
193
+ │ ├── research.md # Research on publishing platforms (GitHub, Hugging Face, MCP)
194
+ │ ├── ARCHITECTURE.md # Technical design, data flow, and database schema
195
+ │ ├── WebAI_specification.md # Core project specification
196
+ │ └── starter.md # Task requirements and model evaluation matrix
197
+ ├── llms.txt # Standard machine-readable AI agent spec index
198
+ ├── llms-full.txt # Comprehensive agent knowledge base and sizing formulas
199
+ ├── README.md # Project overview, manifesto, quickstart, and live scoreboard
200
+ ├── ARCHITECTURE.md # Root link to technical design
201
+ ├── telemetry_db.py # SQLite telemetry database engine and CSV exporter
202
+ ├── storage4gaming_client.py # Storage4gaming comparative payload generator (HTML vs JSON vs llms.txt)
203
+ ├── steam_client.py # Steam Storefront API client & aggregate hardware calculator
204
+ ├── benchmark_runner.py # MLX sequential inference and benchmarking harness
205
+ ├── download_models.py # Standalone model downloader/installer
206
+ ├── catalog_scanner.py # Cross-platform Steam game library detector (Windows & macOS)
207
+ ├── doc_updater.py # Documentation auto-synchronizer
208
+ ├── test_telemetry.py # Automated test suite for database and telemetry
209
+ ├── run.sh # Master setup and execution script
210
+ ├── webai_benchmarks.db # SQLite database storing benchmark records
211
+ └── benchmark_results.csv # Exported benchmark CSV dataset
212
+ ```
@@ -0,0 +1,203 @@
1
+ # WebAI: The AI-Ready Internet Standard & Token Benchmarking Suite
2
+
3
+ **WebAI** is an open architectural standard, benchmarking suite, and agent identity framework designed to establish an **AI-Agent-Ready Internet** operating alongside the modern human-first web.
4
+
5
+ ---
6
+
7
+ ## 🌍 The WebAI Manifesto: Why Build This Now?
8
+
9
+ * **The 57% Reality:** Major internet infrastructure providers like Cloudflare have documented that **over 57% of all web traffic** is now generated by automated systems and AI agents.
10
+ * **The Structural Flaw:** Despite AI being the majority of internet consumers, the web remains engineered exclusively for human visual perception (deeply nested HTML DOM trees, CSS stylesheets, layout wrappers, and tracking scripts).
11
+ * **The Downstream Impact:** AI agents waste up to **95%+ of their I/O token budgets** simply parsing presentation boilerplate. This inflates task completion times, drives up inference costs, and needlessly ties up global GPU clusters that could be solving valuable societal tasks.
12
+ * **The Dual-Web Solution:** Humans keep their rich visual web interfaces, while websites provide a parallel **Agent-Ready Web** with zero presentation bloat.
13
+
14
+ ### The Two Pillars of WebAI
15
+ 1. **Pillar 1: Content & Representation Standards**
16
+ * Standardized `llms.txt` and semantic JSON contracts.
17
+ * **Storage4gaming.com:** Complete rewrite into the reference AI-agent-ready site.
18
+ * **Empirical Benchmarks:** Quantifying token and latency savings across open-weight models on Storage4gaming and the Steam Storefront.
19
+ 2. **Pillar 2: AI Agent Identity & Authentication (AIAID)**
20
+ * Cryptographic machine identities for AI agents (analogous to MAC addresses / UUIDs).
21
+ * A public depository and challenge-response handshake that unlocks high-speed semantic endpoints for verified agents while filtering malicious scrapers.
22
+
23
+ ---
24
+
25
+ ## 🌟 Key Highlights
26
+
27
+ - **~76% - 95% Token Reduction:** Cuts prompt tokens from thousands of tokens down to ~100 tokens.
28
+ - **1.5x - 2.5x Faster Latency:** Accelerates task completion by stripping away visual rendering layers.
29
+ - **Embedded SQLite Telemetry:** All benchmark runs, tokenizer token counts, throughput (tok/s), latency, and peak RAM are persisted to `webai_benchmarks.db` and synchronized with `benchmark_results.csv`.
30
+ - **Apple Silicon Optimized:** 100% offline inference on Apple Silicon using `mlx-lm` with sequential execution and automatic memory cache clearing to protect 16 GB Unified Memory.
31
+ - **Auto-Updating Documentation:** Live benchmark tables and model cache statuses in this README and `ARCHITECTURE.md` stay automatically synchronized with your local benchmark runs.
32
+
33
+ ---
34
+
35
+ ## 📊 Live Benchmark Performance Summary
36
+
37
+ > The table below is automatically synchronized with [`webai_benchmarks.db`](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/webai_benchmarks.db) by `doc_updater.py`:
38
+
39
+ <!-- BENCHMARK_TABLE_START -->
40
+
41
+ | Model | Target Game | Human Tokens | Agent Tokens | Token Savings | Human Latency | Agent Latency | Speedup |
42
+ |------------------------------|------------------|----------------|----------------|-----------------|-----------------|-----------------|-------------|
43
+ | `Llama-3.2-1B-Instruct-4bit` | Apex Legends | 565 tok | 132 tok | **76.64%** | 0.56s | 0.42s | 1.3x faster |
44
+ | `Llama-3.2-1B-Instruct-4bit` | Counter-Strike 2 | 567 tok | 136 tok | **76.01%** | 0.99s | 0.45s | 2.2x faster |
45
+ | `Llama-3.2-1B-Instruct-4bit` | Halo Infinite | 565 tok | 132 tok | **76.64%** | 0.50s | 0.42s | 1.2x faster |
46
+ | `gemma-4-e4b-it-OptiQ-4bit` | Apex Legends | 635 tok | 125 tok | **80.31%** | 4.69s | 3.45s | 1.4x faster |
47
+ | `gemma-4-e4b-it-OptiQ-4bit` | Counter-Strike 2 | 637 tok | 132 tok | **79.28%** | 6.96s | 4.77s | 1.5x faster |
48
+ | `gemma-4-e4b-it-OptiQ-4bit` | Halo Infinite | 636 tok | 126 tok | **80.19%** | 3.61s | 3.31s | 1.1x faster |
49
+
50
+ <!-- BENCHMARK_TABLE_END -->
51
+
52
+ ---
53
+
54
+ ## 🤖 Supported Model Matrix & Local Cache Status
55
+
56
+ The suite evaluates models across two buckets, all optimized in 4-bit quantization for 16 GB Unified Memory:
57
+ - **Bucket 1 (< 4B):** Ultra-lightweight models for fast edge inference.
58
+ - **Bucket 2 (4B – 12B):** Desktop-class models including Google's **Gemma 4 12B** and **Gemma 4 E4B (8B)**.
59
+
60
+ > Check your local cache anytime with `python download_models.py --list`:
61
+
62
+ <!-- MODEL_STATUS_START -->
63
+
64
+ | Status | Bucket | Est. RAM (4-bit) | Hugging Face Model ID |
65
+ |--------------|----------|--------------------|--------------------------------------------------|
66
+ | **[CACHED]** | <4B | ~1.0 GB | `mlx-community/Llama-3.2-1B-Instruct-4bit` |
67
+ | [NOT CACHED] | <4B | ~2.2 GB | `mlx-community/Llama-3.2-3B-Instruct-4bit` |
68
+ | [NOT CACHED] | <4B | ~2.1 GB | `mlx-community/Qwen2.5-3B-Instruct-4bit` |
69
+ | [NOT CACHED] | <4B | ~2.5 GB | `mlx-community/Phi-4-mini-instruct-4bit` |
70
+ | **[CACHED]** | 4B-12B | ~7.5 GB | `mlx-community/gemma-4-12B-it-qat-4bit` |
71
+ | **[CACHED]** | 4B-12B | ~3.5 GB | `mlx-community/gemma-4-e4b-it-OptiQ-4bit` |
72
+ | [NOT CACHED] | 4B-12B | ~4.8 GB | `mlx-community/Qwen2.5-7B-Instruct-4bit` |
73
+ | [NOT CACHED] | 4B-12B | ~5.2 GB | `mlx-community/Llama-3.1-8B-Instruct-4bit` |
74
+ | [NOT CACHED] | 4B-12B | ~5.2 GB | `mlx-community/DeepSeek-R1-Distill-Qwen-8B-4bit` |
75
+
76
+ <!-- MODEL_STATUS_END -->
77
+
78
+ ---
79
+
80
+ ## 🚀 Quickstart
81
+
82
+ ### Prerequisites
83
+ - macOS on Apple Silicon (M-series).
84
+ - Homebrew installed (`/opt/homebrew/bin/brew`).
85
+ - Python 3.11 (`brew install python@3.11`).
86
+
87
+ ### One-Command Setup & Benchmark
88
+ Execute the master runner to set up the environment, run unit tests, and launch a benchmark:
89
+ ```bash
90
+ chmod +x run.sh
91
+ ./run.sh
92
+ ```
93
+
94
+ ---
95
+
96
+ ## 🛠️ CLI Usage Guide
97
+
98
+ ### 1. Download / Install Models Without Running Inference
99
+ Download model weights directly to your local Hugging Face cache (`~/.cache/huggingface/hub/`) without loading them into RAM:
100
+ ```bash
101
+ # Check cache status of all models:
102
+ .venv/bin/python3 download_models.py --list
103
+
104
+ # Download a specific model (e.g. Gemma 4 12B):
105
+ .venv/bin/python3 download_models.py --models mlx-community/gemma-4-12B-it-qat-4bit
106
+
107
+ # Download all models in a bucket:
108
+ .venv/bin/python3 download_models.py --bucket "<4B"
109
+ ```
110
+
111
+ ### 2. Run Comparative Token Benchmarks
112
+ Run side-by-side comparative evaluation between human-first HTML and agent-first JSON:
113
+ ```bash
114
+ # Benchmark specific models:
115
+ .venv/bin/python3 benchmark_runner.py --models mlx-community/gemma-4-e4b-it-OptiQ-4bit mlx-community/Llama-3.2-1B-Instruct-4bit
116
+
117
+ # Benchmark a specific game title:
118
+ .venv/bin/python3 benchmark_runner.py --game "Cyberpunk 2077"
119
+
120
+ # Enable live web scraping from https://www.storage4gaming.com:
121
+ .venv/bin/python3 benchmark_runner.py --live
122
+ ```
123
+
124
+ ### 3. Scan Local Steam Game Catalog
125
+ Discover locally installed games and owned account library titles from Steam (Windows & macOS):
126
+ ```bash
127
+ # Standard discovery:
128
+ .venv/bin/python3 catalog_scanner.py
129
+
130
+ # Test with synthetic sample catalog:
131
+ .venv/bin/python3 catalog_scanner.py --sample
132
+ ```
133
+
134
+ ### 4. Steam Storefront Specs & Hardware Capacity Calculator
135
+ Query the Steam Store API vs. storefront HTML, and calculate total library storage and RAM:
136
+ ```bash
137
+ # Benchmark single Steam game (e.g. Apex Legends):
138
+ .venv/bin/python3 steam_client.py --appid 1172470
139
+
140
+ # Calculate aggregate storage & RAM for your entire scanned Steam library:
141
+ .venv/bin/python3 steam_client.py --scan
142
+ ```
143
+
144
+ ### 5. Update Documentation Automatically
145
+ Keep `README.md` and `ARCHITECTURE.md` updated with the latest SQLite benchmark records and model cache status:
146
+ ```bash
147
+ .venv/bin/python3 doc_updater.py
148
+ ```
149
+
150
+ ---
151
+
152
+ ## 🗄️ Database & Telemetry Inspection
153
+
154
+ All metrics are stored in SQLite (`webai_benchmarks.db`) and CSV (`benchmark_results.csv`).
155
+
156
+ ### Query Recent Runs
157
+ ```bash
158
+ sqlite3 webai_benchmarks.db "SELECT model_id, site_format, prompt_tokens, token_savings_pct, total_time_sec, peak_memory_mb FROM benchmarks ORDER BY id DESC LIMIT 6;"
159
+ ```
160
+
161
+ ### View CSV Log
162
+ ```bash
163
+ cat benchmark_results.csv
164
+ ```
165
+
166
+ ---
167
+
168
+ ## 📚 Central Documentation (`docs/`)
169
+
170
+ All architectural designs, research, and design decision logs are organized in the [`docs/`](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs) folder:
171
+ * **[docs/decisions.md](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs/decisions.md)**: Design log covering Skills vs. Scripts, `llms.txt`, local filesystem discovery, PyPI packaging, readability scoring metrics, and the AIAID protocol.
172
+ * **[docs/research.md](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs/research.md)**: Research on publishing platforms (GitHub, Hugging Face Datasets & Spaces, Model Context Protocol / MCP, `llms.txt` directories, and PyPI).
173
+ * **[docs/ARCHITECTURE.md](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs/ARCHITECTURE.md)**: Full system design, Mermaid data flow, AIAID authentication handshake, and Storage4gaming rewrite blueprint.
174
+ * **[docs/WebAI_specification.md](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs/WebAI_specification.md)**: Master project specification and criteria.
175
+ * **[docs/starter.md](file:///Users/alexandermarkovski/Coding/In_Progress/WebAI/docs/starter.md)**: Task requirements and model testing matrix.
176
+
177
+ ---
178
+
179
+ ## 📂 Repository Layout
180
+
181
+ ```
182
+ ├── docs/ # Central documentation folder
183
+ │ ├── decisions.md # Pending design decisions and architectural trade-offs
184
+ │ ├── research.md # Research on publishing platforms (GitHub, Hugging Face, MCP)
185
+ │ ├── ARCHITECTURE.md # Technical design, data flow, and database schema
186
+ │ ├── WebAI_specification.md # Core project specification
187
+ │ └── starter.md # Task requirements and model evaluation matrix
188
+ ├── llms.txt # Standard machine-readable AI agent spec index
189
+ ├── llms-full.txt # Comprehensive agent knowledge base and sizing formulas
190
+ ├── README.md # Project overview, manifesto, quickstart, and live scoreboard
191
+ ├── ARCHITECTURE.md # Root link to technical design
192
+ ├── telemetry_db.py # SQLite telemetry database engine and CSV exporter
193
+ ├── storage4gaming_client.py # Storage4gaming comparative payload generator (HTML vs JSON vs llms.txt)
194
+ ├── steam_client.py # Steam Storefront API client & aggregate hardware calculator
195
+ ├── benchmark_runner.py # MLX sequential inference and benchmarking harness
196
+ ├── download_models.py # Standalone model downloader/installer
197
+ ├── catalog_scanner.py # Cross-platform Steam game library detector (Windows & macOS)
198
+ ├── doc_updater.py # Documentation auto-synchronizer
199
+ ├── test_telemetry.py # Automated test suite for database and telemetry
200
+ ├── run.sh # Master setup and execution script
201
+ ├── webai_benchmarks.db # SQLite database storing benchmark records
202
+ └── benchmark_results.csv # Exported benchmark CSV dataset
203
+ ```