@thinkingos/vsl-sdk 0.4.2 β 0.4.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +115 -57
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -1,84 +1,142 @@
|
|
|
1
|
-
|
|
1
|
+
# Visual Scene Language (VSL) SDK
|
|
2
2
|
|
|
3
|
-
|
|
4
|
-
Publication Date: 28.05.2025
|
|
5
|
-
Version: v0.1
|
|
3
|
+
Visual Scene Language (VSL) is a JSON-based format and TypeScript SDK for AI agents that interact with screen content. It provides a structured, semantic representation of visual scenes β replacing raw pixel screenshots with compact, machine-readable JSON that language models can interpret directly.
|
|
6
4
|
|
|
7
|
-
|
|
5
|
+
VSL is not a UI application or a rendering engine. It is an infrastructure layer (middleware) that sits between the screen and the LLM, enabling computer use agents to understand, navigate, and act on visual interfaces through semantic structure rather than pixel analysis.
|
|
8
6
|
|
|
9
|
-
|
|
7
|
+
---
|
|
10
8
|
|
|
11
|
-
|
|
9
|
+
## How It Works
|
|
12
10
|
|
|
13
|
-
|
|
11
|
+
Instead of sending full-page screenshots (1-2 MB) to an LLM on every interaction cycle, VSL extracts the semantic structure of the screen through DOM and accessibility APIs, then produces a compact JSON document (10-100 KB) containing:
|
|
14
12
|
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
β’ UX/UI design
|
|
19
|
-
β’ Architecture, Engineering, and Construction (AEC)
|
|
20
|
-
β’ 2D/3D modeling, AR/VR
|
|
13
|
+
- A canvas descriptor (viewport dimensions, coordinate system)
|
|
14
|
+
- A tree of semantic objects (buttons, text fields, containers, links, etc.) with types, positions, sizes, states, and available actions
|
|
15
|
+
- Visual fragments only for elements that cannot be described through structural data alone (images, canvas elements, custom widgets)
|
|
21
16
|
|
|
22
|
-
|
|
17
|
+
The LLM receives this JSON, reasons about the scene, and returns an action (click, type, scroll, navigate). The SDK executes the action against the DOM, and the cycle repeats β with subsequent snapshots transmitted as diffs rather than full documents.
|
|
23
18
|
|
|
24
|
-
|
|
19
|
+
---
|
|
25
20
|
|
|
26
|
-
|
|
27
|
-
β’ canvas parameters (dimensions, measurement units, background),
|
|
28
|
-
β’ a list of objects (type, size, coordinates, anchor points),
|
|
29
|
-
β’ optional styles and object relationships.
|
|
21
|
+
## Repository Structure
|
|
30
22
|
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
23
|
+
visual-scene-language/
|
|
24
|
+
src/ Core SDK source code
|
|
25
|
+
extension/ Chrome browser extension (Manifest V3)
|
|
26
|
+
packages/
|
|
27
|
+
mcp-server/ Model Context Protocol server for VSL
|
|
28
|
+
tests/ Integration and extension tests
|
|
29
|
+
scripts/ Demo and utility scripts
|
|
30
|
+
### Core SDK (`src/`)
|
|
36
31
|
|
|
37
|
-
|
|
32
|
+
The SDK is organized into six modules, each corresponding to a stage of the capture-to-action pipeline:
|
|
38
33
|
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
34
|
+
| Module | Path | Responsibility |
|
|
35
|
+
|---|---|---|
|
|
36
|
+
| **Capture** | `src/capture/` | DOM tree extraction, element metadata, computed styles, bounding rectangles |
|
|
37
|
+
| **Segmentation** | `src/segmentation/` | Multi-level semantic classification (tag-based, ARIA role, CSS heuristic, structural, vision-assisted) |
|
|
38
|
+
| **Builder** | `src/builder/` | Assembles segmented elements into a VSL JSON document |
|
|
39
|
+
| **Cache & Diff** | `src/cache/`, `src/diff/`, `src/session/` | Content-hash and coordinate-hash caching, mutation observer invalidation, diff engine, snapshot sessions |
|
|
40
|
+
| **Action Executor** | `src/executor/` | Resolves target IDs to DOM elements and executes actions (click, type, scroll, navigate, select, etc.) |
|
|
41
|
+
| **Vision** | `src/vision/` | Visual fragment extraction, vision-based enrichment, LLM vision classifier |
|
|
44
42
|
|
|
45
|
-
|
|
43
|
+
### MCP Server (`packages/mcp-server/`)
|
|
46
44
|
|
|
47
|
-
|
|
48
|
-
β’ Unlike generative models (DALLΒ·E, Midjourney), this method separates the semantic structure of the scene from its visual rendering.
|
|
49
|
-
β’ Unlike scene graphs in computer graphics, VSL is optimized for language models and semantic processing.
|
|
50
|
-
β’ The method does not require a visual interface β interaction occurs through structure and natural language commands.
|
|
45
|
+
A standalone Model Context Protocol server that exposes VSL capabilities as tools for any MCP-compatible AI agent. Available tools:
|
|
51
46
|
|
|
52
|
-
|
|
47
|
+
- `vsl_get_snapshot` β capture the current page as a VSL JSON document
|
|
48
|
+
- `vsl_get_diff` β retrieve only the changes since the last snapshot
|
|
49
|
+
- `vsl_execute_action` β execute an action (click, type, scroll, navigate) on the page
|
|
50
|
+
- `vsl_navigate` β navigate the browser to a specified URL
|
|
51
|
+
- `vsl_read_page` β hybrid page reader (HTTP-first for static pages, headless browser for SPAs)
|
|
52
|
+
- `vsl_get_text_block` β retrieve a cached long-form text block by reference ID
|
|
53
|
+
- `vsl_get_visual` β retrieve a visual fragment (image/canvas) by reference ID
|
|
54
|
+
- `vsl_get_full_json` β retrieve the complete VSL document without diff compression
|
|
55
|
+
- `vsl_clear_cache` β invalidate the snapshot cache
|
|
56
|
+
- `vsl_download` β download a resource from the page
|
|
53
57
|
|
|
54
|
-
|
|
55
|
-
β’ Editing user interfaces, slides, and illustrations without a visual editor.
|
|
56
|
-
β’ Generating AR/VR scenes from textual scenarios.
|
|
57
|
-
β’ Interactive assistants working with visual objects.
|
|
58
|
-
β’ Robotics: LLM perceives the environment through VSL and gives structured commands.
|
|
59
|
-
β’ Architectural and engineering design: users describe a building or structure via text; the LLM generates a corresponding VSL structure. A renderer builds a 2D/3D model from it. Edits are made via natural language and reflected in the scene. The system can also validate design choices against regulatory standards.
|
|
58
|
+
### Browser Extension (`extension/`)
|
|
60
59
|
|
|
61
|
-
|
|
60
|
+
A Chrome extension (Manifest V3) that runs VSL capture and action execution directly in the browser. It provides a content script for DOM interaction and a popup interface for controlling the agent session.
|
|
62
61
|
|
|
63
|
-
|
|
62
|
+
---
|
|
64
63
|
|
|
65
|
-
|
|
64
|
+
## Installation
|
|
66
65
|
|
|
67
|
-
|
|
66
|
+
npm install @thinkingos/vsl-sdk
|
|
67
|
+
Requires Node.js 18 or later.
|
|
68
68
|
|
|
69
|
-
|
|
70
|
-
β’ vsl_scene_example.json: example scene with a red rectangle
|
|
71
|
-
β’ llm_prompts_examples.md: prompt examples for modifying the scene
|
|
72
|
-
β’ vsl_to_image.py: Python script to visualize the JSON scene
|
|
73
|
-
β’ vsl_editor_mockup.ipynb: (optional) Jupyter Notebook mock editor
|
|
74
|
-
β’ Extended format specs: styles, materials, animation, 3D coordinates, behavioral logic
|
|
69
|
+
---
|
|
75
70
|
|
|
76
|
-
|
|
71
|
+
## Package Exports
|
|
77
72
|
|
|
78
|
-
|
|
73
|
+
The SDK publishes dual CJS/ESM builds with TypeScript declarations:
|
|
74
|
+
|
|
75
|
+
{
|
|
76
|
+
"import": "./dist/index.mjs",
|
|
77
|
+
"require": "./dist/index.js",
|
|
78
|
+
"types": "./dist/index.d.ts"
|
|
79
|
+
}
|
|
80
|
+
---
|
|
81
|
+
|
|
82
|
+
## Key Capabilities
|
|
83
|
+
|
|
84
|
+
### Snapshot Generation
|
|
85
|
+
|
|
86
|
+
The capture pipeline extracts the DOM tree, classifies each element through a multi-level segmentation algorithm (HTML tag mapping, ARIA role resolution, CSS heuristics, structural analysis, and optional vision model enrichment), and builds a VSL JSON document with semantic object types, relative coordinates, states, and available actions.
|
|
87
|
+
|
|
88
|
+
### Cache and Diff
|
|
89
|
+
|
|
90
|
+
Static elements are cached by content hash and coordinate hash after the first snapshot. Subsequent snapshots transmit only the differences β added, modified, and removed objects β while unchanged elements are referenced by ID. Mutation observers provide point invalidation when the DOM changes between snapshots. URL navigation triggers a full cache reset; viewport resize resets only coordinate-dependent entries.
|
|
91
|
+
|
|
92
|
+
### Visual Fragments and Screenshots
|
|
93
|
+
|
|
94
|
+
When structural data alone is not sufficient, the AI agent can request a screenshot of the full screen or of a specific element identified in the VSL document. Each element in the VSL JSON carries a unique ID β the agent passes this ID to retrieve a visual fragment (image, canvas, or custom widget) as a base64-encoded WebP image. This allows the agent to fall back to pixel-level inspection only when needed, while keeping the default workflow fully semantic.
|
|
95
|
+
### Action Execution
|
|
96
|
+
|
|
97
|
+
Actions returned by the LLM are resolved against the live DOM. The executor maps semantic object IDs to actual DOM elements, performs the requested interaction, and returns a structured result indicating success or failure.
|
|
98
|
+
|
|
99
|
+
### Humanization Layer
|
|
100
|
+
|
|
101
|
+
To operate on sites with anti-bot detection (LinkedIn, banking portals, etc.), VSL includes a Humanization Layer that sits between the Action Executor and the browser. It makes agent actions indistinguishable from human behavior by introducing configurable timing delays, Bezier-curve mouse movements, scroll randomization, human-like typing simulation, session management, and fingerprint rotation. Configuration is per-site β each domain can have its own timing profiles, action rate limits, and typing speed ranges. The overhead (+300β3000ms per action) is intentional: it mimics real human speed and rhythm.
|
|
102
|
+
|
|
103
|
+
### Prompt Injection Filter
|
|
104
|
+
|
|
105
|
+
VSL intercepts all text extracted from the screen before it reaches the LLM. A pattern-based filter runs at the entry point of the Capture Layer β immediately after text extraction from any source (DOM, raw HTML) and before segmentation β scanning for hidden instructions embedded in `display:none` elements, alt texts, meta tags, HTML comments, and CSS content properties. Detected injections are stripped, logged, or blocked depending on severity. The pattern library supports three tiers: bundled (shipped with the SDK), remote (auto-loaded from CDN, cached locally), and custom (user-defined via configuration). This provides a defense-in-depth security layer at the middleware level.
|
|
106
|
+
|
|
107
|
+
---
|
|
108
|
+
|
|
109
|
+
## Development
|
|
110
|
+
|
|
111
|
+
npm install # install dependencies
|
|
112
|
+
npm run build # build SDK (tsup)
|
|
113
|
+
npm run test # run unit tests (jest)
|
|
114
|
+
npm run lint # lint source (eslint)
|
|
115
|
+
npm run typecheck # type-check without emit (tsc --noEmit)
|
|
116
|
+
npm run test:e2e # end-to-end tests (playwright)
|
|
117
|
+
npm run demo:cache-diff # cache-diff demonstration
|
|
118
|
+
---
|
|
119
|
+
|
|
120
|
+
## Documentation
|
|
121
|
+
|
|
122
|
+
| Document | Description |
|
|
123
|
+
|---|---|
|
|
124
|
+
| `ARCHITECTURE.md` | System architecture: pipeline stages, data flow, caching strategy, diff format, action model |
|
|
125
|
+
| `PRODUCT_CONCEPT.md` | Product concept: problem statement, solution overview, use cases, competitive positioning |
|
|
126
|
+
| `DESIGN_SYSTEM.md` | JSON schema design principles, naming conventions, data model, API patterns |
|
|
127
|
+
| `DECISIONS.md` | Architectural decisions with context and rationale |
|
|
128
|
+
| `ROADMAP.md` | Implementation roadmap organized by phases and milestones |
|
|
129
|
+
| `TARGET_AUDIENCE.md` | Target audience definitions and usage scenarios |
|
|
130
|
+
| `README_AI.md` | Project bible for AI agents (single-document context ingestion) |
|
|
131
|
+
|
|
132
|
+
---
|
|
133
|
+
|
|
134
|
+
## License
|
|
79
135
|
|
|
80
136
|
Creative Commons Attribution 4.0 International (CC BY 4.0)
|
|
81
137
|
|
|
82
|
-
|
|
138
|
+
---
|
|
139
|
+
|
|
140
|
+
## Author
|
|
83
141
|
|
|
84
|
-
|
|
142
|
+
Maxim Zhadobin
|