vibes-plug 2.14.1 → 3.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (151) hide show
  1. package/.claude/rules/vibes-plug-core.md +5 -0
  2. package/.cursor/rules/vibes-plug-core.mdc +7 -2
  3. package/.cursorrules +8 -2
  4. package/AGENTS.md +23 -2
  5. package/CHANGELOG.md +114 -0
  6. package/CLAUDE.md +10 -3
  7. package/README.md +216 -611
  8. package/bin/vibes.mjs +1104 -0
  9. package/package.json +11 -3
  10. package/plugin.json +4 -3
  11. package/scripts/check-anti-slop.mjs +53 -0
  12. package/scripts/install.js +3 -1
  13. package/scripts/update_skills.js +1 -1
  14. package/scripts/update_skills.mjs +86 -0
  15. package/scripts/validate-skills.mjs +111 -0
  16. package/skills/accessibility-testing-expert/SKILL.md +117 -116
  17. package/skills/affective-computing-emotion-ai/SKILL.md +83 -0
  18. package/skills/agentic-coding-workflow-expert/SKILL.md +297 -0
  19. package/skills/agentic-memory-architect/SKILL.md +52 -0
  20. package/skills/agentic-micro-economy-architect/SKILL.md +92 -0
  21. package/skills/ai-llm-integration-expert/SKILL.md +330 -194
  22. package/skills/ai-media-generation-expert/SKILL.md +173 -172
  23. package/skills/ai-prompt-engineering-expert/SKILL.md +204 -134
  24. package/skills/ai-safety-governance-expert/SKILL.md +223 -0
  25. package/skills/angular-expert/SKILL.md +149 -148
  26. package/skills/anti-slop/SKILL.md +134 -133
  27. package/skills/api-design-expert/SKILL.md +4 -3
  28. package/skills/api-gateway-proxy-expert/SKILL.md +3 -2
  29. package/skills/app-analyzer-optimizer/SKILL.md +4 -3
  30. package/skills/apple-ecosystem-expert/SKILL.md +6 -5
  31. package/skills/astro-framework-expert/SKILL.md +201 -200
  32. package/skills/async-queue-temporal-expert/SKILL.md +218 -217
  33. package/skills/authentication-identity-expert/SKILL.md +174 -173
  34. package/skills/autonomous-red-teamer/SKILL.md +338 -203
  35. package/skills/autonomous-tdd-debugger/SKILL.md +6 -5
  36. package/skills/biome-linter-formatter-expert/SKILL.md +90 -89
  37. package/skills/blockchain-web3-expert/SKILL.md +116 -115
  38. package/skills/brainstorming/SKILL.md +392 -377
  39. package/skills/browser-automation-expert/SKILL.md +260 -222
  40. package/skills/bun-runtime-expert/SKILL.md +5 -4
  41. package/skills/chatbot-messaging-expert/SKILL.md +115 -114
  42. package/skills/ci-cd-devops-architect/SKILL.md +3 -2
  43. package/skills/cloud-hosting-expert/SKILL.md +5 -4
  44. package/skills/coderabbit/SKILL.md +5 -4
  45. package/skills/compliance-gdpr-privacy-expert/SKILL.md +3 -2
  46. package/skills/composable-mach-architect/SKILL.md +338 -0
  47. package/skills/cron-scheduler-expert/SKILL.md +5 -4
  48. package/skills/data-pipeline-etl-expert/SKILL.md +3 -2
  49. package/skills/data-telemetry-expert/SKILL.md +5 -4
  50. package/skills/data-visualization-expert/SKILL.md +155 -154
  51. package/skills/database-orm-expert/SKILL.md +166 -165
  52. package/skills/deep-research-analyst/SKILL.md +182 -136
  53. package/skills/dependency-upgrade-migrator/SKILL.md +11 -10
  54. package/skills/design-system-architect/SKILL.md +4 -3
  55. package/skills/desktop-electron-expert/SKILL.md +129 -128
  56. package/skills/documentation-site-expert/SKILL.md +60 -59
  57. package/skills/doku-mcp-server/SKILL.md +5 -4
  58. package/skills/doku-payment-gateway/SKILL.md +250 -232
  59. package/skills/domain-driven-design-expert/SKILL.md +3 -2
  60. package/skills/e2e-testing-expert/SKILL.md +5 -4
  61. package/skills/ecommerce-expert/SKILL.md +88 -87
  62. package/skills/email-notification-expert/SKILL.md +5 -4
  63. package/skills/ephemeral-generative-ui-architect/SKILL.md +88 -0
  64. package/skills/error-resilience-expert/SKILL.md +14 -13
  65. package/skills/event-driven-architect/SKILL.md +5 -4
  66. package/skills/feature-flag-analytics-expert/SKILL.md +3 -2
  67. package/skills/file-upload-media-expert/SKILL.md +5 -4
  68. package/skills/firebase-security-expert/SKILL.md +5 -4
  69. package/skills/form-validation-expert/SKILL.md +7 -6
  70. package/skills/frontier-ai-models-expert/SKILL.md +116 -0
  71. package/skills/fullstack-expert/SKILL.md +185 -184
  72. package/skills/gemini-agent-booster/SKILL.md +248 -172
  73. package/skills/geospatial-maps-expert/SKILL.md +81 -80
  74. package/skills/global-a11y-i18n-expert/SKILL.md +5 -4
  75. package/skills/glsl-shader-expert/SKILL.md +191 -190
  76. package/skills/go-programming-expert/SKILL.md +5 -4
  77. package/skills/graph-rag-knowledge-expert/SKILL.md +201 -200
  78. package/skills/graphql-apollo-expert/SKILL.md +5 -4
  79. package/skills/headless-cms-expert/SKILL.md +182 -181
  80. package/skills/hig/SKILL.md +5 -4
  81. package/skills/js-backend-expert/SKILL.md +219 -218
  82. package/skills/legacy-code-translator/SKILL.md +6 -5
  83. package/skills/llm-finops-router/SKILL.md +52 -0
  84. package/skills/local-slm-edge-ai-expert/SKILL.md +168 -167
  85. package/skills/logging-error-tracking-expert/SKILL.md +5 -4
  86. package/skills/mcp-server-architect/SKILL.md +315 -307
  87. package/skills/micro-frontend-architect/SKILL.md +5 -4
  88. package/skills/mobile-expo-expert/SKILL.md +5 -4
  89. package/skills/modern-css-native-expert/SKILL.md +190 -189
  90. package/skills/monorepo-architect/SKILL.md +5 -4
  91. package/skills/mpa-orchestrator/SKILL.md +41 -4
  92. package/skills/multi-agent-orchestration/SKILL.md +388 -254
  93. package/skills/mvc-expert/SKILL.md +5 -4
  94. package/skills/n8n-automation-expert/SKILL.md +90 -89
  95. package/skills/nextjs-app-router-expert/SKILL.md +3 -2
  96. package/skills/openapi-swagger-codegen-expert/SKILL.md +4 -3
  97. package/skills/payment-gateway-expert/SKILL.md +131 -128
  98. package/skills/pdf-document-generation-expert/SKILL.md +92 -91
  99. package/skills/performance-web-vitals/SKILL.md +5 -4
  100. package/skills/post-quantum-crypto-migrator/SKILL.md +3 -2
  101. package/skills/prd-architect/SKILL.md +183 -182
  102. package/skills/proactive-background-watcher/SKILL.md +5 -4
  103. package/skills/production-ready-hardener/SKILL.md +10 -9
  104. package/skills/pwa-offline-first-expert/SKILL.md +227 -226
  105. package/skills/pydantic-ai-expert/SKILL.md +162 -161
  106. package/skills/python-programming-expert/SKILL.md +5 -4
  107. package/skills/rate-limit-abuse-prevention/SKILL.md +5 -4
  108. package/skills/realtime-collaboration-expert/SKILL.md +3 -2
  109. package/skills/rich-text-editor-expert/SKILL.md +178 -177
  110. package/skills/rust-programming-expert/SKILL.md +5 -4
  111. package/skills/saas-architect/SKILL.md +155 -154
  112. package/skills/saas-billing/SKILL.md +394 -382
  113. package/skills/saas-multi-tenant/SKILL.md +7 -6
  114. package/skills/scalability-clean-code/SKILL.md +5 -4
  115. package/skills/search-engine-expert/SKILL.md +90 -89
  116. package/skills/self-healing-cloud-orchestrator/SKILL.md +3 -2
  117. package/skills/senior-frontend/SKILL.md +14 -9
  118. package/skills/seo/SKILL.md +4 -4
  119. package/skills/session-memory-manager/SKILL.md +129 -128
  120. package/skills/solidjs-expert/SKILL.md +81 -80
  121. package/skills/spa-orchestrator/SKILL.md +5 -4
  122. package/skills/sse-websocket-streaming-expert/SKILL.md +3 -2
  123. package/skills/state-management-expert/SKILL.md +5 -4
  124. package/skills/supabase-security-expert/SKILL.md +5 -4
  125. package/skills/svelte-sveltekit-expert/SKILL.md +92 -91
  126. package/skills/svg-animation-motion-expert/SKILL.md +3 -2
  127. package/skills/synthetic-data-finetuning-expert/SKILL.md +156 -155
  128. package/skills/tailwind-expert/SKILL.md +62 -5
  129. package/skills/tanstack-query-expert/SKILL.md +5 -4
  130. package/skills/tauri-expert/SKILL.md +5 -4
  131. package/skills/typescript-expert/SKILL.md +5 -4
  132. package/skills/ui-ux-pro-max/SKILL.md +7 -6
  133. package/skills/vector-db-rag-expert/SKILL.md +209 -208
  134. package/skills/vercel-ai-sdk-expert/SKILL.md +226 -181
  135. package/skills/voice-ai-realtime-agent/SKILL.md +243 -242
  136. package/skills/vue-frontend-expert/SKILL.md +5 -4
  137. package/skills/wasm-edge-computing-expert/SKILL.md +3 -2
  138. package/skills/web-3d-graphics-expert/SKILL.md +314 -313
  139. package/skills/web-game-engine-expert/SKILL.md +330 -329
  140. package/skills/web-scraper/SKILL.md +158 -157
  141. package/skills/website-design-cloner/SKILL.md +5 -4
  142. package/skills/webxr-ar-vr-expert/SKILL.md +163 -162
  143. package/skills/wordpress-headless-expert/SKILL.md +145 -144
  144. package/skills/zero-tech-debt-auditor/SKILL.md +115 -0
  145. package/skills/zero-to-prod-orchestrator/SKILL.md +281 -229
  146. package/skills/zero-trust-secret-vault/SKILL.md +3 -2
  147. package/BLUEPRINT.md +0 -319
  148. package/skills/bootstrap-to-modern/SKILL.md +0 -94
  149. package/skills/multiple-entry-points/SKILL.md +0 -91
  150. package/skills/secure-fuzz-testing/SKILL.md +0 -207
  151. package/skills/visual-qa-vision-agent/SKILL.md +0 -71
@@ -1,157 +1,158 @@
1
- ---
2
- name: web-scraper
3
- description: "Smart agentic web data extraction with multi-strategy scraping (Crawl4AI v4, Firecrawl), LLM extraction loops, anti-bot bypass, and structured export / Ekstraksi data web cerdas dan agentic dengan scraping multi-strategi (Crawl4AI v4, Firecrawl), ekstraksi LLM, bypass anti-bot, dan ekspor terstruktur."
4
- author: "Roedy Rustam"
5
- ---
6
-
7
- # Agentic Web Scraper Expert (2026 Edition)
8
-
9
- [English](#english) | [Bahasa Indonesia](#bahasa-indonesia)
10
-
11
- ---
12
-
13
- <a name="english"></a>
14
- ## English
15
-
16
- ### Orchestration & Integration
17
- Connects and orchestrates with relevant domain skills like `browser-automation-expert`, `ai-llm-integration-expert`, `brainstorming`, and `zero-to-prod-orchestrator` to ensure cohesive agentic execution.
18
-
19
- ### Description
20
- Advanced Agentic Web Scraping utilizing modern multi-strategy data extraction. Leverages **Crawl4AI v4** and **Firecrawl** to convert raw DOMs into LLM-friendly Markdown. Implements Agentic Extraction loops where the LLM guides the scraper dynamically based on page state. Incorporates strategies for bypassing anti-bot measures (Cloudflare Turnstile, Datadome) and navigating dynamic Shadow DOMs.
21
-
22
- ### Trigger Conditions
23
- - Extracting structured data from websites for analysis, training data, or content pipelines.
24
- - Scraping dynamic JavaScript-rendered pages and complex SPAs.
25
- - Converting web pages to clean Markdown for LLM context or RAG pipelines.
26
- - Dealing with anti-bot protections or complex Shadow DOM architectures during scraping.
27
- - Implementing an automated agentic data extraction loop.
28
-
29
- ### Extracting DOM into LLM-Friendly Markdown
30
- Use **Crawl4AI v4** for high-performance async extraction and **Firecrawl** for seamless LLM-ready conversion.
31
-
32
- **Crawl4AI v4 (Async Python):**
33
- ```python
34
- import asyncio
35
- from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode
36
-
37
- async def extract_markdown(url: str):
38
- config = BrowserConfig(headless=True, bypass_csp=True)
39
- run_config = CrawlerRunConfig(
40
- cache_mode=CacheMode.ENABLED,
41
- remove_overlay_elements=True,
42
- word_count_threshold=50
43
- )
44
-
45
- async with AsyncWebCrawler(config=config) as crawler:
46
- result = await crawler.arun(url=url, config=run_config)
47
- # Returns clean, AI-optimized markdown ready for LLM consumption
48
- return result.markdown.fit_markdown
49
- ```
50
-
51
- **Firecrawl (Managed API):**
52
- ```python
53
- from firecrawl import FirecrawlApp
54
- from pydantic import BaseModel
55
-
56
- app = FirecrawlApp(api_key="fc-xxxx")
57
-
58
- class ExtractionSchema(BaseModel):
59
- title: str
60
- content: str
61
- key_metrics: list[str]
62
-
63
- # Single API call to extract structured data based on JSON schema
64
- result = app.scrape_url(
65
- "https://example.com/data",
66
- formats=["extract", "markdown"],
67
- extract={"schema": ExtractionSchema.model_json_schema()}
68
- )
69
- print(result.markdown) # Clean markdown
70
- print(result.extract) # Structured JSON
71
- ```
72
-
73
- ### Anti-Bot Bypass & Shadow DOMs
74
- Scraping modern web apps requires bypassing anti-bot measures like Cloudflare Turnstile and Datadome, as well as accessing deeply nested elements.
75
-
76
- 1. **Anti-Bot Bypass (Cloudflare Turnstile, Datadome):**
77
- - **Residential Proxies:** Rotate high-quality residential IPs to avoid datacenter IP bans.
78
- - **Browser Fingerprinting:** Use tools like `playwright-stealth` or specialized stealth browsers (e.g., Undetected ChromeDriver, Curl-Impersonate) to mask automated fingerprints (WebGL, Canvas, User-Agent).
79
- - **Human-like Interaction:** Introduce random delays, simulate realistic mouse movements, and handle CAPTCHAs via third-party solving services only when necessary.
80
- 2. **Dynamic Shadow DOMs:**
81
- - Use CSS piercing selectors or JavaScript execution to penetrate the Shadow Root.
82
- - Example (Playwright): `await page.locator('my-web-component >> css=.internal-element').text_content()`
83
- - Recursively traverse the DOM tree injecting scripts to extract content from encapsulated components.
84
-
85
- ### Agentic Extraction Loops
86
- Implement an autonomous loop where an LLM guides the scraper based on the current page state, rather than relying on brittle CSS selectors.
87
-
88
- 1. **Observe:** The scraper extracts the current DOM into clean Markdown.
89
- 2. **Analyze:** The LLM analyzes the Markdown to identify necessary data or the next interaction step (e.g., "Click the 'Load More' button").
90
- 3. **Act:** The LLM issues a command (extract data, navigate, click, fill form).
91
- 4. **Loop:** Repeat until the extraction goal is met.
92
-
93
- ```python
94
- async def agentic_scrape_loop(url: str, goal: str):
95
- current_url = url
96
- while True:
97
- markdown_content = await extract_markdown(current_url)
98
- # LLM analyzes state and decides next action
99
- action = await llm_decide_action(markdown_content, goal)
100
-
101
- if action.type == "COMPLETE":
102
- return action.extracted_data
103
- elif action.type == "CLICK":
104
- await click_element(action.target_selector)
105
- elif action.type == "NAVIGATE":
106
- current_url = action.new_url
107
- ```
108
-
109
- ### Ethical Scraping Checklist
110
- - [ ] Check `robots.txt` and respect `Disallow` rules.
111
- - [ ] Implement rate limiting.
112
- - [ ] Use descriptive `User-Agent` headers.
113
- - [ ] Do not scrape personal/private data without consent.
114
-
115
- ---
116
-
117
- <a name="bahasa-indonesia"></a>
118
- ## Bahasa Indonesia
119
-
120
- ### Integrasi Orkestrasi
121
- Terhubung dan mengorkestrasi skill domain yang relevan seperti `browser-automation-expert`, `ai-llm-integration-expert`, `brainstorming`, dan `zero-to-prod-orchestrator` untuk memastikan eksekusi agentic yang kohesif.
122
-
123
- ### Deskripsi
124
- Scraping Web Agentic tingkat lanjut menggunakan ekstraksi data multi-strategi modern. Memanfaatkan **Crawl4AI v4** dan **Firecrawl** untuk mengubah DOM mentah menjadi Markdown yang ramah LLM. Mengimplementasikan loop Ekstraksi Agentic di mana LLM memandu scraper secara dinamis berdasarkan status halaman. Menggabungkan strategi untuk melewati tindakan anti-bot (Cloudflare Turnstile, Datadome) dan menavigasi Shadow DOM yang dinamis.
125
-
126
- ### Kondisi Pemicu
127
- - Mengekstrak data terstruktur dari situs web untuk analisis, data pelatihan, atau pipeline konten.
128
- - Scraping halaman yang dirender JavaScript secara dinamis dan SPA kompleks.
129
- - Mengonversi halaman web menjadi Markdown bersih untuk konteks LLM atau pipeline RAG.
130
- - Menghadapi perlindungan anti-bot atau arsitektur Shadow DOM yang kompleks saat scraping.
131
- - Mengimplementasikan loop ekstraksi data agentic otomatis.
132
-
133
- ### Mengekstrak DOM menjadi Markdown Ramah LLM
134
- Gunakan **Crawl4AI v4** untuk ekstraksi async berperforma tinggi dan **Firecrawl** untuk konversi siap LLM yang mulus. (Lihat contoh kode di bagian bahasa Inggris).
135
-
136
- ### Bypass Anti-Bot & Shadow DOM
137
- 1. **Bypass Anti-Bot (Cloudflare Turnstile, Datadome):**
138
- - **Proxy Residensial:** Rotasi IP residensial berkualitas tinggi untuk menghindari pemblokiran IP datacenter.
139
- - **Browser Fingerprinting:** Gunakan alat seperti `playwright-stealth` atau browser stealth khusus untuk menyembunyikan sidik jari otomatis.
140
- - **Interaksi Mirip Manusia:** Tambahkan penundaan acak, simulasikan gerakan mouse yang realistis.
141
- 2. **Shadow DOM Dinamis:**
142
- - Gunakan selektor penembus CSS atau eksekusi JavaScript untuk menembus Shadow Root.
143
- - Telusuri pohon DOM secara rekursif dengan menyuntikkan skrip untuk mengekstrak konten.
144
-
145
- ### Loop Ekstraksi Agentic
146
- Implementasikan loop otonom di mana LLM memandu scraper berdasarkan status halaman saat ini, bukan bergantung pada selektor CSS yang rentan rusak.
147
-
148
- 1. **Observasi:** Scraper mengekstrak DOM saat ini menjadi Markdown yang bersih.
149
- 2. **Analisis:** LLM menganalisis Markdown untuk mengidentifikasi data yang diperlukan atau langkah interaksi selanjutnya (misal: "Klik tombol 'Muat Lebih Banyak'").
150
- 3. **Aksi:** LLM mengeluarkan perintah (ekstrak data, navigasi, klik, isi form).
151
- 4. **Loop:** Ulangi hingga tujuan ekstraksi tercapai.
152
-
153
- ### Checklist Scraping Etis
154
- - [ ] Periksa `robots.txt` dan hormati aturan `Disallow`.
155
- - [ ] Implementasikan rate limiting.
156
- - [ ] Gunakan header `User-Agent` yang deskriptif.
157
- - [ ] Jangan scraping data pribadi/privat tanpa izin.
1
+ ---
2
+ name: web-scraper
3
+ description: "Smart agentic web data extraction with multi-strategy scraping (Crawl4AI v4, Firecrawl), LLM extraction loops, anti-bot bypass, and structured export / Ekstraksi data web cerdas dan agentic dengan scraping multi-strategi (Crawl4AI v4, Firecrawl), ekstraksi LLM, bypass anti-bot, dan ekspor terstruktur."
4
+ author: "Roedy Rustam"
5
+ version: "3.0.0"
6
+ ---
7
+
8
+ # Agentic Web Scraper Expert (2026 Edition)
9
+
10
+ [English](#english) | [Bahasa Indonesia](#bahasa-indonesia)
11
+
12
+ ---
13
+
14
+ <a name="english"></a>
15
+ ## English
16
+
17
+ ### Orchestration & Integration
18
+ Connects and orchestrates with relevant domain skills like `browser-automation-expert`, `ai-llm-integration-expert`, `brainstorming`, and `zero-to-prod-orchestrator` to ensure cohesive agentic execution.
19
+
20
+ ### Description
21
+ Advanced Agentic Web Scraping utilizing modern multi-strategy data extraction. Leverages **Crawl4AI v4** and **Firecrawl** to convert raw DOMs into LLM-friendly Markdown. Implements Agentic Extraction loops where the LLM guides the scraper dynamically based on page state. Incorporates strategies for bypassing anti-bot measures (Cloudflare Turnstile, Datadome) and navigating dynamic Shadow DOMs.
22
+
23
+ ### Trigger Conditions
24
+ - Extracting structured data from websites for analysis, training data, or content pipelines.
25
+ - Scraping dynamic JavaScript-rendered pages and complex SPAs.
26
+ - Converting web pages to clean Markdown for LLM context or RAG pipelines.
27
+ - Dealing with anti-bot protections or complex Shadow DOM architectures during scraping.
28
+ - Implementing an automated agentic data extraction loop.
29
+
30
+ ### Extracting DOM into LLM-Friendly Markdown
31
+ Use **Crawl4AI v4** for high-performance async extraction and **Firecrawl** for seamless LLM-ready conversion.
32
+
33
+ **Crawl4AI v4 (Async Python):**
34
+ ```python
35
+ import asyncio
36
+ from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode
37
+
38
+ async def extract_markdown(url: str):
39
+ config = BrowserConfig(headless=True, bypass_csp=True)
40
+ run_config = CrawlerRunConfig(
41
+ cache_mode=CacheMode.ENABLED,
42
+ remove_overlay_elements=True,
43
+ word_count_threshold=50
44
+ )
45
+
46
+ async with AsyncWebCrawler(config=config) as crawler:
47
+ result = await crawler.arun(url=url, config=run_config)
48
+ # Returns clean, AI-optimized markdown ready for LLM consumption
49
+ return result.markdown.fit_markdown
50
+ ```
51
+
52
+ **Firecrawl (Managed API):**
53
+ ```python
54
+ from firecrawl import FirecrawlApp
55
+ from pydantic import BaseModel
56
+
57
+ app = FirecrawlApp(api_key="fc-xxxx")
58
+
59
+ class ExtractionSchema(BaseModel):
60
+ title: str
61
+ content: str
62
+ key_metrics: list[str]
63
+
64
+ # Single API call to extract structured data based on JSON schema
65
+ result = app.scrape_url(
66
+ "https://example.com/data",
67
+ formats=["extract", "markdown"],
68
+ extract={"schema": ExtractionSchema.model_json_schema()}
69
+ )
70
+ print(result.markdown) # Clean markdown
71
+ print(result.extract) # Structured JSON
72
+ ```
73
+
74
+ ### Anti-Bot Bypass & Shadow DOMs
75
+ Scraping modern web apps requires bypassing anti-bot measures like Cloudflare Turnstile and Datadome, as well as accessing deeply nested elements.
76
+
77
+ 1. **Anti-Bot Bypass (Cloudflare Turnstile, Datadome):**
78
+ - **Residential Proxies:** Rotate high-quality residential IPs to avoid datacenter IP bans.
79
+ - **Browser Fingerprinting:** Use tools like `playwright-stealth` or specialized stealth browsers (e.g., Undetected ChromeDriver, Curl-Impersonate) to mask automated fingerprints (WebGL, Canvas, User-Agent).
80
+ - **Human-like Interaction:** Introduce random delays, simulate realistic mouse movements, and handle CAPTCHAs via third-party solving services only when necessary.
81
+ 2. **Dynamic Shadow DOMs:**
82
+ - Use CSS piercing selectors or JavaScript execution to penetrate the Shadow Root.
83
+ - Example (Playwright): `await page.locator('my-web-component >> css=.internal-element').text_content()`
84
+ - Recursively traverse the DOM tree injecting scripts to extract content from encapsulated components.
85
+
86
+ ### Agentic Extraction Loops
87
+ Implement an autonomous loop where an LLM guides the scraper based on the current page state, rather than relying on brittle CSS selectors.
88
+
89
+ 1. **Observe:** The scraper extracts the current DOM into clean Markdown.
90
+ 2. **Analyze:** The LLM analyzes the Markdown to identify necessary data or the next interaction step (e.g., "Click the 'Load More' button").
91
+ 3. **Act:** The LLM issues a command (extract data, navigate, click, fill form).
92
+ 4. **Loop:** Repeat until the extraction goal is met.
93
+
94
+ ```python
95
+ async def agentic_scrape_loop(url: str, goal: str):
96
+ current_url = url
97
+ while True:
98
+ markdown_content = await extract_markdown(current_url)
99
+ # LLM analyzes state and decides next action
100
+ action = await llm_decide_action(markdown_content, goal)
101
+
102
+ if action.type == "COMPLETE":
103
+ return action.extracted_data
104
+ elif action.type == "CLICK":
105
+ await click_element(action.target_selector)
106
+ elif action.type == "NAVIGATE":
107
+ current_url = action.new_url
108
+ ```
109
+
110
+ ### Ethical Scraping Checklist
111
+ - [ ] Check `robots.txt` and respect `Disallow` rules.
112
+ - [ ] Implement rate limiting.
113
+ - [ ] Use descriptive `User-Agent` headers.
114
+ - [ ] Do not scrape personal/private data without consent.
115
+
116
+ ---
117
+
118
+ <a name="bahasa-indonesia"></a>
119
+ ## Bahasa Indonesia
120
+
121
+ ### Integrasi Orkestrasi
122
+ Terhubung dan mengorkestrasi skill domain yang relevan seperti `browser-automation-expert`, `ai-llm-integration-expert`, `brainstorming`, dan `zero-to-prod-orchestrator` untuk memastikan eksekusi agentic yang kohesif.
123
+
124
+ ### Deskripsi
125
+ Scraping Web Agentic tingkat lanjut menggunakan ekstraksi data multi-strategi modern. Memanfaatkan **Crawl4AI v4** dan **Firecrawl** untuk mengubah DOM mentah menjadi Markdown yang ramah LLM. Mengimplementasikan loop Ekstraksi Agentic di mana LLM memandu scraper secara dinamis berdasarkan status halaman. Menggabungkan strategi untuk melewati tindakan anti-bot (Cloudflare Turnstile, Datadome) dan menavigasi Shadow DOM yang dinamis.
126
+
127
+ ### Kondisi Pemicu
128
+ - Mengekstrak data terstruktur dari situs web untuk analisis, data pelatihan, atau pipeline konten.
129
+ - Scraping halaman yang dirender JavaScript secara dinamis dan SPA kompleks.
130
+ - Mengonversi halaman web menjadi Markdown bersih untuk konteks LLM atau pipeline RAG.
131
+ - Menghadapi perlindungan anti-bot atau arsitektur Shadow DOM yang kompleks saat scraping.
132
+ - Mengimplementasikan loop ekstraksi data agentic otomatis.
133
+
134
+ ### Mengekstrak DOM menjadi Markdown Ramah LLM
135
+ Gunakan **Crawl4AI v4** untuk ekstraksi async berperforma tinggi dan **Firecrawl** untuk konversi siap LLM yang mulus. (Lihat contoh kode di bagian bahasa Inggris).
136
+
137
+ ### Bypass Anti-Bot & Shadow DOM
138
+ 1. **Bypass Anti-Bot (Cloudflare Turnstile, Datadome):**
139
+ - **Proxy Residensial:** Rotasi IP residensial berkualitas tinggi untuk menghindari pemblokiran IP datacenter.
140
+ - **Browser Fingerprinting:** Gunakan alat seperti `playwright-stealth` atau browser stealth khusus untuk menyembunyikan sidik jari otomatis.
141
+ - **Interaksi Mirip Manusia:** Tambahkan penundaan acak, simulasikan gerakan mouse yang realistis.
142
+ 2. **Shadow DOM Dinamis:**
143
+ - Gunakan selektor penembus CSS atau eksekusi JavaScript untuk menembus Shadow Root.
144
+ - Telusuri pohon DOM secara rekursif dengan menyuntikkan skrip untuk mengekstrak konten.
145
+
146
+ ### Loop Ekstraksi Agentic
147
+ Implementasikan loop otonom di mana LLM memandu scraper berdasarkan status halaman saat ini, bukan bergantung pada selektor CSS yang rentan rusak.
148
+
149
+ 1. **Observasi:** Scraper mengekstrak DOM saat ini menjadi Markdown yang bersih.
150
+ 2. **Analisis:** LLM menganalisis Markdown untuk mengidentifikasi data yang diperlukan atau langkah interaksi selanjutnya (misal: "Klik tombol 'Muat Lebih Banyak'").
151
+ 3. **Aksi:** LLM mengeluarkan perintah (ekstrak data, navigasi, klik, isi form).
152
+ 4. **Loop:** Ulangi hingga tujuan ekstraksi tercapai.
153
+
154
+ ### Checklist Scraping Etis
155
+ - [ ] Periksa `robots.txt` dan hormati aturan `Disallow`.
156
+ - [ ] Implementasikan rate limiting.
157
+ - [ ] Gunakan header `User-Agent` yang deskriptif.
158
+ - [ ] Jangan scraping data pribadi/privat tanpa izin.
@@ -1,7 +1,8 @@
1
1
  ---
2
2
  name: website-design-cloner
3
3
  description: "Analyzes and reverse-engineers website designs directly from a target URL, extracting layout structures, design tokens (colors, typography, spacing), component hierarchies, visual assets, and responsive behaviors to enable full 1:1 duplication into modern code (Tailwind CSS v4, React/Next.js, HTML/CSS). / Mempelajari dan merekayasa balik desain situs web langsung dari URL target, mengekstrak struktur layout, design token (warna, tipografi, spacing), hierarki komponen, aset visual, dan perilaku responsif untuk duplikasi 1:1 penuh ke kode modern."
4
- author: "Roedy Rustam"
4
+ author: "Roedy Rustam"
5
+ version: "3.0.0"
5
6
  ---
6
7
 
7
8
  # Website Design Cloner & Reverse Engineering Expert (2026 Edition)
@@ -14,7 +15,7 @@ author: "Roedy Rustam"
14
15
  ## English
15
16
 
16
17
  ### Orchestration & Integration
17
- Connects and orchestrates with relevant domain skills like `brainstorming`, `zero-to-prod-orchestrator`, and `project-context-mapper` to ensure cohesive execution.
18
+ Connects and orchestrates with relevant domain skills like `brainstorming`, `zero-to-prod-orchestrator`, and `session-memory-manager` to ensure cohesive execution.
18
19
 
19
20
  ### Description
20
21
  `website-design-cloner` is an advanced URL-to-Code visual reverse engineering skill. It enables AI agents to inspect any target website URL, analyze its visual aesthetic, layout grid, CSS design tokens (OKLCH/HEX colors, typography system, container bounds, shadow tiers, border-radii), DOM structure, and interactive components, and synthesize production-ready code (Tailwind CSS v4, React 19, Next.js 15, HTML5/Vanilla CSS) to achieve a full 1:1 duplication.
@@ -135,7 +136,7 @@ export function HeroSection() {
135
136
  ## Bahasa Indonesia
136
137
 
137
138
  ### Integrasi Orkestrasi
138
- Terhubung dan mengorkestrasi skill domain yang relevan seperti `brainstorming`, `zero-to-prod-orchestrator`, dan `project-context-mapper` untuk memastikan eksekusi yang kohesif.
139
+ Terhubung dan mengorkestrasi skill domain yang relevan seperti `brainstorming`, `zero-to-prod-orchestrator`, dan `session-memory-manager` untuk memastikan eksekusi yang kohesif.
139
140
 
140
141
  ### Deskripsi
141
142
  `website-design-cloner` adalah skill rekayasa balik (*reverse engineering*) visual dari URL ke kode. Skill ini memungkinkan agen AI mempelajari situs web target dari URL, mengaudit estetika visual, grid layout, design token CSS (warna OKLCH/HEX, sistem tipografi, batas kontainer, bayangan, radius border), struktur DOM, dan komponen interaktif, lalu merekonstruksi kode siap produksi (Tailwind CSS v4, React 19, Next.js 15, HTML5/CSS3) untuk mencapai duplikasi 1:1 penuh.
@@ -177,4 +178,4 @@ Terhubung dan mengorkestrasi skill domain yang relevan seperti `brainstorming`,
177
178
  | `design-system-architect` | Menyusun token visual hasil ekstraksi ke dalam design system enterprise berbasis OKLCH & Radix/Base UI. |
178
179
  | `ui-ux-pro-max` | Memberikan acuan BM25 visual style, pasangan font Google Fonts, dan checklist aksesibilitas WCAG 2.2. |
179
180
  | `senior-frontend` | Mengimplementasikan kode komponen React 19 / Next.js 15 App Router siap produksi. |
180
- | `tailwind-expert` | Mengatur konfigurasi Tailwind CSS v4 `@theme` dan utilitas responsif. |
181
+ | `tailwind-expert` | Mengatur konfigurasi Tailwind CSS v4 `@theme` dan utilitas responsif. |