@thinkingos/vsl-sdk 0.4.1 β†’ 0.4.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (2) hide show
  1. package/README.md +115 -57
  2. package/package.json +1 -1
package/README.md CHANGED
@@ -1,84 +1,142 @@
1
- 🧾 Defensive Publication: Visual Scene Language (VSL)
1
+ # Visual Scene Language (VSL) SDK
2
2
 
3
- Author: Maxim Zhadobin
4
- Publication Date: 28.05.2025
5
- Version: v0.1
3
+ Visual Scene Language (VSL) is a JSON-based format and TypeScript SDK for AI agents that interact with screen content. It provides a structured, semantic representation of visual scenes β€” replacing raw pixel screenshots with compact, machine-readable JSON that language models can interpret directly.
6
4
 
7
- βΈ»
5
+ VSL is not a UI application or a rendering engine. It is an infrastructure layer (middleware) that sits between the screen and the LLM, enabling computer use agents to understand, navigate, and act on visual interfaces through semantic structure rather than pixel analysis.
8
6
 
9
- πŸ“˜ Invention Title
7
+ ---
10
8
 
11
- Method for Representing and Editing Visual Scenes for Language Models via a Structured Format: Visual Scene Language (VSL)
9
+ ## How It Works
12
10
 
13
- βΈ»
11
+ Instead of sending full-page screenshots (1-2 MB) to an LLM on every interaction cycle, VSL extracts the semantic structure of the screen through DOM and accessibility APIs, then produces a compact JSON document (10-100 KB) containing:
14
12
 
15
- πŸ“Œ Application Domains
16
- β€’ Artificial Intelligence and Large Language Models (LLMs)
17
- β€’ Image generation and editing
18
- β€’ UX/UI design
19
- β€’ Architecture, Engineering, and Construction (AEC)
20
- β€’ 2D/3D modeling, AR/VR
13
+ - A canvas descriptor (viewport dimensions, coordinate system)
14
+ - A tree of semantic objects (buttons, text fields, containers, links, etc.) with types, positions, sizes, states, and available actions
15
+ - Visual fragments only for elements that cannot be described through structural data alone (images, canvas elements, custom widgets)
21
16
 
22
- βΈ»
17
+ The LLM receives this JSON, reasons about the scene, and returns an action (click, type, scroll, navigate). The SDK executes the action against the DOM, and the cycle repeats β€” with subsequent snapshots transmitted as diffs rather than full documents.
23
18
 
24
- 🧠 Abstract
19
+ ---
25
20
 
26
- A method is proposed for describing visual scenes in a machine-readable format interpretable by language models (LLMs). The format is a structured JSON schema containing:
27
- β€’ canvas parameters (dimensions, measurement units, background),
28
- β€’ a list of objects (type, size, coordinates, anchor points),
29
- β€’ optional styles and object relationships.
21
+ ## Repository Structure
30
22
 
31
- An LLM can use this structure:
32
- β€’ to understand the image as a meaningful scene,
33
- β€’ to edit the scene based on textual commands,
34
- β€’ to generate a new scene,
35
- β€’ to reconstruct an image from the JSON.
23
+ visual-scene-language/
24
+ src/ Core SDK source code
25
+ extension/ Chrome browser extension (Manifest V3)
26
+ packages/
27
+ mcp-server/ Model Context Protocol server for VSL
28
+ tests/ Integration and extension tests
29
+ scripts/ Demo and utility scripts
30
+ ### Core SDK (`src/`)
36
31
 
37
- βΈ»
32
+ The SDK is organized into six modules, each corresponding to a stage of the capture-to-action pipeline:
38
33
 
39
- πŸ”„ Technological Workflow
40
- 1. An image is converted into a JSON-based scene structure (VSL).
41
- 2. The LLM interprets and modifies the JSON in response to a text prompt.
42
- 3. The modified JSON is rendered into a new image.
43
- 4. The cycle may repeat iteratively.
34
+ | Module | Path | Responsibility |
35
+ |---|---|---|
36
+ | **Capture** | `src/capture/` | DOM tree extraction, element metadata, computed styles, bounding rectangles |
37
+ | **Segmentation** | `src/segmentation/` | Multi-level semantic classification (tag-based, ARIA role, CSS heuristic, structural, vision-assisted) |
38
+ | **Builder** | `src/builder/` | Assembles segmented elements into a VSL JSON document |
39
+ | **Cache & Diff** | `src/cache/`, `src/diff/`, `src/session/` | Content-hash and coordinate-hash caching, mutation observer invalidation, diff engine, snapshot sessions |
40
+ | **Action Executor** | `src/executor/` | Resolves target IDs to DOM elements and executes actions (click, type, scroll, navigate, select, etc.) |
41
+ | **Vision** | `src/vision/` | Visual fragment extraction, vision-based enrichment, LLM vision classifier |
44
42
 
45
- βΈ»
43
+ ### MCP Server (`packages/mcp-server/`)
46
44
 
47
- 🧩 Distinctive Features
48
- β€’ Unlike generative models (DALLΒ·E, Midjourney), this method separates the semantic structure of the scene from its visual rendering.
49
- β€’ Unlike scene graphs in computer graphics, VSL is optimized for language models and semantic processing.
50
- β€’ The method does not require a visual interface β€” interaction occurs through structure and natural language commands.
45
+ A standalone Model Context Protocol server that exposes VSL capabilities as tools for any MCP-compatible AI agent. Available tools:
51
46
 
52
- βΈ»
47
+ - `vsl_get_snapshot` β€” capture the current page as a VSL JSON document
48
+ - `vsl_get_diff` β€” retrieve only the changes since the last snapshot
49
+ - `vsl_execute_action` β€” execute an action (click, type, scroll, navigate) on the page
50
+ - `vsl_navigate` β€” navigate the browser to a specified URL
51
+ - `vsl_read_page` β€” hybrid page reader (HTTP-first for static pages, headless browser for SPAs)
52
+ - `vsl_get_text_block` β€” retrieve a cached long-form text block by reference ID
53
+ - `vsl_get_visual` β€” retrieve a visual fragment (image/canvas) by reference ID
54
+ - `vsl_get_full_json` β€” retrieve the complete VSL document without diff compression
55
+ - `vsl_clear_cache` β€” invalidate the snapshot cache
56
+ - `vsl_download` β€” download a resource from the page
53
57
 
54
- πŸ’‘ Example Applications
55
- β€’ Editing user interfaces, slides, and illustrations without a visual editor.
56
- β€’ Generating AR/VR scenes from textual scenarios.
57
- β€’ Interactive assistants working with visual objects.
58
- β€’ Robotics: LLM perceives the environment through VSL and gives structured commands.
59
- β€’ Architectural and engineering design: users describe a building or structure via text; the LLM generates a corresponding VSL structure. A renderer builds a 2D/3D model from it. Edits are made via natural language and reflected in the scene. The system can also validate design choices against regulatory standards.
58
+ ### Browser Extension (`extension/`)
60
59
 
61
- βΈ»
60
+ A Chrome extension (Manifest V3) that runs VSL capture and action execution directly in the browser. It provides a content script for DOM interaction and a popup interface for controlling the agent session.
62
61
 
63
- πŸ“œ Legal Status
62
+ ---
64
63
 
65
- This document is published with the intent to prevent patent claims by third parties. The author waives exclusive patent rights in favor of public use but formally asserts authorship, concept origin, and publication date.
64
+ ## Installation
66
65
 
67
- βΈ»
66
+ npm install @thinkingos/vsl-sdk
67
+ Requires Node.js 18 or later.
68
68
 
69
- πŸ“Ž Attachments (to be included in GitHub repository)
70
- β€’ vsl_scene_example.json: example scene with a red rectangle
71
- β€’ llm_prompts_examples.md: prompt examples for modifying the scene
72
- β€’ vsl_to_image.py: Python script to visualize the JSON scene
73
- β€’ vsl_editor_mockup.ipynb: (optional) Jupyter Notebook mock editor
74
- β€’ Extended format specs: styles, materials, animation, 3D coordinates, behavioral logic
69
+ ---
75
70
 
76
- βΈ»
71
+ ## Package Exports
77
72
 
78
- πŸ”– License
73
+ The SDK publishes dual CJS/ESM builds with TypeScript declarations:
74
+
75
+ {
76
+ "import": "./dist/index.mjs",
77
+ "require": "./dist/index.js",
78
+ "types": "./dist/index.d.ts"
79
+ }
80
+ ---
81
+
82
+ ## Key Capabilities
83
+
84
+ ### Snapshot Generation
85
+
86
+ The capture pipeline extracts the DOM tree, classifies each element through a multi-level segmentation algorithm (HTML tag mapping, ARIA role resolution, CSS heuristics, structural analysis, and optional vision model enrichment), and builds a VSL JSON document with semantic object types, relative coordinates, states, and available actions.
87
+
88
+ ### Cache and Diff
89
+
90
+ Static elements are cached by content hash and coordinate hash after the first snapshot. Subsequent snapshots transmit only the differences β€” added, modified, and removed objects β€” while unchanged elements are referenced by ID. Mutation observers provide point invalidation when the DOM changes between snapshots. URL navigation triggers a full cache reset; viewport resize resets only coordinate-dependent entries.
91
+
92
+ ### Visual Fragments and Screenshots
93
+
94
+ When structural data alone is not sufficient, the AI agent can request a screenshot of the full screen or of a specific element identified in the VSL document. Each element in the VSL JSON carries a unique ID β€” the agent passes this ID to retrieve a visual fragment (image, canvas, or custom widget) as a base64-encoded WebP image. This allows the agent to fall back to pixel-level inspection only when needed, while keeping the default workflow fully semantic.
95
+ ### Action Execution
96
+
97
+ Actions returned by the LLM are resolved against the live DOM. The executor maps semantic object IDs to actual DOM elements, performs the requested interaction, and returns a structured result indicating success or failure.
98
+
99
+ ### Humanization Layer
100
+
101
+ To operate on sites with anti-bot detection (LinkedIn, banking portals, etc.), VSL includes a Humanization Layer that sits between the Action Executor and the browser. It makes agent actions indistinguishable from human behavior by introducing configurable timing delays, Bezier-curve mouse movements, scroll randomization, human-like typing simulation, session management, and fingerprint rotation. Configuration is per-site β€” each domain can have its own timing profiles, action rate limits, and typing speed ranges. The overhead (+300–3000ms per action) is intentional: it mimics real human speed and rhythm.
102
+
103
+ ### Prompt Injection Filter
104
+
105
+ VSL intercepts all text extracted from the screen before it reaches the LLM. A pattern-based filter runs at the entry point of the Capture Layer β€” immediately after text extraction from any source (DOM, raw HTML) and before segmentation β€” scanning for hidden instructions embedded in `display:none` elements, alt texts, meta tags, HTML comments, and CSS content properties. Detected injections are stripped, logged, or blocked depending on severity. The pattern library supports three tiers: bundled (shipped with the SDK), remote (auto-loaded from CDN, cached locally), and custom (user-defined via configuration). This provides a defense-in-depth security layer at the middleware level.
106
+
107
+ ---
108
+
109
+ ## Development
110
+
111
+ npm install # install dependencies
112
+ npm run build # build SDK (tsup)
113
+ npm run test # run unit tests (jest)
114
+ npm run lint # lint source (eslint)
115
+ npm run typecheck # type-check without emit (tsc --noEmit)
116
+ npm run test:e2e # end-to-end tests (playwright)
117
+ npm run demo:cache-diff # cache-diff demonstration
118
+ ---
119
+
120
+ ## Documentation
121
+
122
+ | Document | Description |
123
+ |---|---|
124
+ | `ARCHITECTURE.md` | System architecture: pipeline stages, data flow, caching strategy, diff format, action model |
125
+ | `PRODUCT_CONCEPT.md` | Product concept: problem statement, solution overview, use cases, competitive positioning |
126
+ | `DESIGN_SYSTEM.md` | JSON schema design principles, naming conventions, data model, API patterns |
127
+ | `DECISIONS.md` | Architectural decisions with context and rationale |
128
+ | `ROADMAP.md` | Implementation roadmap organized by phases and milestones |
129
+ | `TARGET_AUDIENCE.md` | Target audience definitions and usage scenarios |
130
+ | `README_AI.md` | Project bible for AI agents (single-document context ingestion) |
131
+
132
+ ---
133
+
134
+ ## License
79
135
 
80
136
  Creative Commons Attribution 4.0 International (CC BY 4.0)
81
137
 
82
- βΈ»
138
+ ---
139
+
140
+ ## Author
83
141
 
84
- β€œVisual Scene Language is not just a way to describe an image β€” it is a language for spatial thinking by artificial intelligence.”
142
+ Maxim Zhadobin
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@thinkingos/vsl-sdk",
3
- "version": "0.4.1",
3
+ "version": "0.4.3",
4
4
  "description": "Visual Scene Language (VSL) SDK β€” M1.1 Snapshot generation (DOM β†’ VSL JSON)",
5
5
  "license": "SEE LICENSE IN LICENSE.txt",
6
6
  "engines": {