vibes-plug 2.5.0 → 2.14.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (187) hide show
  1. package/.claude/rules/vibes-plug-core.md +32 -0
  2. package/.cursor/rules/vibes-plug-core.mdc +51 -0
  3. package/.cursorrules +42 -0
  4. package/AGENTS.md +37 -7
  5. package/BLUEPRINT.md +319 -217
  6. package/CHANGELOG.md +170 -1
  7. package/CLAUDE.md +70 -0
  8. package/LICENSE +1 -1
  9. package/README.md +611 -415
  10. package/index.js +19 -0
  11. package/package.json +44 -8
  12. package/plugin.json +24 -7
  13. package/scripts/check-anti-slop.js +53 -0
  14. package/scripts/generate_swarm_gif.py +295 -0
  15. package/scripts/install.js +201 -0
  16. package/skills/accessibility-testing-expert/SKILL.md +116 -0
  17. package/skills/ai-llm-integration-expert/SKILL.md +192 -178
  18. package/skills/ai-media-generation-expert/SKILL.md +172 -0
  19. package/skills/ai-prompt-engineering-expert/SKILL.md +134 -0
  20. package/skills/angular-expert/SKILL.md +148 -0
  21. package/skills/anti-slop/SKILL.md +133 -0
  22. package/skills/api-design-expert/SKILL.md +6 -3
  23. package/skills/api-gateway-proxy-expert/SKILL.md +81 -0
  24. package/skills/app-analyzer-optimizer/SKILL.md +6 -3
  25. package/skills/apple-ecosystem-expert/SKILL.md +144 -141
  26. package/skills/astro-framework-expert/SKILL.md +200 -0
  27. package/skills/async-queue-temporal-expert/SKILL.md +189 -26
  28. package/skills/authentication-identity-expert/SKILL.md +78 -181
  29. package/skills/autonomous-red-teamer/SKILL.md +172 -28
  30. package/skills/autonomous-tdd-debugger/SKILL.md +70 -64
  31. package/skills/biome-linter-formatter-expert/SKILL.md +89 -0
  32. package/skills/blockchain-web3-expert/SKILL.md +115 -0
  33. package/skills/bootstrap-to-modern/SKILL.md +9 -6
  34. package/skills/brainstorming/SKILL.md +60 -52
  35. package/skills/browser-automation-expert/SKILL.md +197 -21
  36. package/skills/bun-runtime-expert/SKILL.md +7 -1
  37. package/skills/chatbot-messaging-expert/SKILL.md +114 -0
  38. package/skills/ci-cd-devops-architect/SKILL.md +45 -36
  39. package/skills/cloud-hosting-expert/SKILL.md +7 -1
  40. package/skills/coderabbit/SKILL.md +7 -1
  41. package/skills/compliance-gdpr-privacy-expert/SKILL.md +85 -0
  42. package/skills/cron-scheduler-expert/SKILL.md +303 -297
  43. package/skills/data-pipeline-etl-expert/SKILL.md +84 -0
  44. package/skills/data-telemetry-expert/SKILL.md +7 -1
  45. package/skills/data-visualization-expert/SKILL.md +154 -0
  46. package/skills/database-orm-expert/SKILL.md +165 -294
  47. package/skills/deep-research-analyst/SKILL.md +136 -0
  48. package/skills/dependency-upgrade-migrator/SKILL.md +300 -294
  49. package/skills/design-system-architect/SKILL.md +309 -260
  50. package/skills/desktop-electron-expert/SKILL.md +128 -0
  51. package/skills/documentation-site-expert/SKILL.md +59 -0
  52. package/skills/doku-mcp-server/SKILL.md +7 -1
  53. package/skills/doku-payment-gateway/SKILL.md +7 -1
  54. package/skills/domain-driven-design-expert/SKILL.md +82 -0
  55. package/skills/e2e-testing-expert/SKILL.md +7 -1
  56. package/skills/ecommerce-expert/SKILL.md +87 -0
  57. package/skills/email-notification-expert/SKILL.md +393 -360
  58. package/skills/error-resilience-expert/SKILL.md +506 -479
  59. package/skills/event-driven-architect/SKILL.md +7 -1
  60. package/skills/feature-flag-analytics-expert/SKILL.md +65 -45
  61. package/skills/file-upload-media-expert/SKILL.md +436 -430
  62. package/skills/firebase-security-expert/SKILL.md +7 -1
  63. package/skills/form-validation-expert/SKILL.md +406 -400
  64. package/skills/fullstack-expert/SKILL.md +184 -202
  65. package/skills/gemini-agent-booster/SKILL.md +173 -135
  66. package/skills/geospatial-maps-expert/SKILL.md +80 -0
  67. package/skills/global-a11y-i18n-expert/SKILL.md +7 -1
  68. package/skills/glsl-shader-expert/SKILL.md +158 -69
  69. package/skills/go-programming-expert/SKILL.md +21 -15
  70. package/skills/graph-rag-knowledge-expert/SKILL.md +200 -0
  71. package/skills/graphql-apollo-expert/SKILL.md +113 -107
  72. package/skills/headless-cms-expert/SKILL.md +181 -0
  73. package/skills/hig/SKILL.md +7 -1
  74. package/skills/js-backend-expert/SKILL.md +218 -216
  75. package/skills/legacy-code-translator/SKILL.md +70 -64
  76. package/skills/local-slm-edge-ai-expert/SKILL.md +167 -0
  77. package/skills/logging-error-tracking-expert/SKILL.md +343 -337
  78. package/skills/mcp-server-architect/SKILL.md +308 -194
  79. package/skills/micro-frontend-architect/SKILL.md +111 -105
  80. package/skills/mobile-expo-expert/SKILL.md +8 -2
  81. package/skills/modern-css-native-expert/SKILL.md +189 -0
  82. package/skills/monorepo-architect/SKILL.md +7 -1
  83. package/skills/mpa-orchestrator/SKILL.md +20 -1
  84. package/skills/multi-agent-orchestration/SKILL.md +254 -234
  85. package/skills/multiple-entry-points/SKILL.md +37 -1
  86. package/skills/mvc-expert/SKILL.md +7 -1
  87. package/skills/n8n-automation-expert/SKILL.md +89 -0
  88. package/skills/nextjs-app-router-expert/SKILL.md +148 -0
  89. package/skills/openapi-swagger-codegen-expert/SKILL.md +67 -0
  90. package/skills/payment-gateway-expert/SKILL.md +85 -1
  91. package/skills/pdf-document-generation-expert/SKILL.md +91 -0
  92. package/skills/performance-web-vitals/SKILL.md +7 -1
  93. package/skills/post-quantum-crypto-migrator/SKILL.md +57 -57
  94. package/skills/prd-architect/SKILL.md +182 -201
  95. package/skills/proactive-background-watcher/SKILL.md +67 -61
  96. package/skills/production-ready-hardener/SKILL.md +459 -456
  97. package/skills/pwa-offline-first-expert/SKILL.md +226 -0
  98. package/skills/pydantic-ai-expert/SKILL.md +161 -0
  99. package/skills/python-programming-expert/SKILL.md +407 -401
  100. package/skills/rate-limit-abuse-prevention/SKILL.md +376 -370
  101. package/skills/realtime-collaboration-expert/SKILL.md +55 -1
  102. package/skills/rich-text-editor-expert/SKILL.md +177 -0
  103. package/skills/rust-programming-expert/SKILL.md +7 -1
  104. package/skills/saas-architect/SKILL.md +154 -0
  105. package/skills/saas-billing/SKILL.md +7 -1
  106. package/skills/saas-multi-tenant/SKILL.md +7 -1
  107. package/skills/scalability-clean-code/SKILL.md +7 -1
  108. package/skills/search-engine-expert/SKILL.md +89 -0
  109. package/skills/secure-fuzz-testing/SKILL.md +7 -1
  110. package/skills/self-healing-cloud-orchestrator/SKILL.md +57 -57
  111. package/skills/senior-frontend/SKILL.md +139 -161
  112. package/skills/senior-frontend/scripts/frontend_scaffolder.py +1 -1
  113. package/skills/seo/SKILL.md +41 -17
  114. package/skills/session-memory-manager/SKILL.md +128 -0
  115. package/skills/solidjs-expert/SKILL.md +80 -0
  116. package/skills/spa-orchestrator/SKILL.md +20 -1
  117. package/skills/sse-websocket-streaming-expert/SKILL.md +93 -0
  118. package/skills/state-management-expert/SKILL.md +7 -1
  119. package/skills/supabase-security-expert/SKILL.md +7 -1
  120. package/skills/svelte-sveltekit-expert/SKILL.md +91 -0
  121. package/skills/svg-animation-motion-expert/SKILL.md +115 -0
  122. package/skills/synthetic-data-finetuning-expert/SKILL.md +155 -0
  123. package/skills/tailwind-expert/SKILL.md +88 -136
  124. package/skills/tanstack-query-expert/SKILL.md +7 -1
  125. package/skills/tauri-expert/SKILL.md +7 -1
  126. package/skills/typescript-expert/SKILL.md +12 -6
  127. package/skills/ui-ux-pro-max/SKILL.md +25 -3
  128. package/skills/vector-db-rag-expert/SKILL.md +175 -19
  129. package/skills/vercel-ai-sdk-expert/SKILL.md +181 -0
  130. package/skills/visual-qa-vision-agent/SKILL.md +70 -64
  131. package/skills/voice-ai-realtime-agent/SKILL.md +242 -0
  132. package/skills/vue-frontend-expert/SKILL.md +131 -125
  133. package/skills/wasm-edge-computing-expert/SKILL.md +97 -0
  134. package/skills/web-3d-graphics-expert/SKILL.md +262 -80
  135. package/skills/web-game-engine-expert/SKILL.md +281 -48
  136. package/skills/web-scraper/SKILL.md +157 -207
  137. package/skills/website-design-cloner/SKILL.md +179 -173
  138. package/skills/webxr-ar-vr-expert/SKILL.md +108 -63
  139. package/skills/wordpress-headless-expert/SKILL.md +144 -0
  140. package/skills/zero-to-prod-orchestrator/SKILL.md +65 -38
  141. package/skills/zero-trust-secret-vault/SKILL.md +87 -39
  142. package/.github/ISSUE_TEMPLATE/feature_request.md +0 -20
  143. package/.github/workflows/publish.yml +0 -20
  144. package/CONTRIBUTING.md +0 -199
  145. package/SECURITY.md +0 -21
  146. package/banner.png +0 -0
  147. package/skills/ai-cost-token-optimizer/SKILL.md +0 -52
  148. package/skills/asisten_ramah/SKILL.md +0 -41
  149. package/skills/auto-doc-updater/SKILL.md +0 -214
  150. package/skills/autonomous-chaos-monkey/SKILL.md +0 -63
  151. package/skills/autonomous-swarm-director/SKILL.md +0 -69
  152. package/skills/edge-serverless-db-expert/SKILL.md +0 -43
  153. package/skills/hyper-context-synthesizer/SKILL.md +0 -55
  154. package/skills/llm-cost-arbitrage-router/SKILL.md +0 -59
  155. package/skills/mcp-client-orchestrator/SKILL.md +0 -70
  156. package/skills/mobile-push-notification-expert/SKILL.md +0 -51
  157. package/skills/monday-design-aesthetic/SKILL.md +0 -67
  158. package/skills/project-context-mapper/SKILL.md +0 -79
  159. package/skills/saas-mvp-launcher/SKILL.md +0 -241
  160. package/skills/saas-transformer/SKILL.md +0 -489
  161. package/skills/saas-transformer/references/billing_integration_guide.md +0 -401
  162. package/skills/self-evolving-memory-graph/SKILL.md +0 -75
  163. package/skills/senior-fullstack/SKILL.md +0 -167
  164. package/skills/senior-fullstack/references/architecture_patterns.md +0 -160
  165. package/skills/senior-fullstack/references/development_workflows.md +0 -222
  166. package/skills/senior-fullstack/references/tech_stack_guide.md +0 -190
  167. package/skills/senior-fullstack/scripts/code_quality_analyzer.py +0 -114
  168. package/skills/senior-fullstack/scripts/fullstack_scaffolder.py +0 -114
  169. package/skills/senior-fullstack/scripts/project_scaffolder.py +0 -114
  170. package/skills/seo-aeo-landing-page-writer/SKILL.md +0 -97
  171. package/skills/seo-geo/SKILL.md +0 -188
  172. package/skills/session-context-loader/SKILL.md +0 -77
  173. package/skills/session-handoff-resume/SKILL.md +0 -158
  174. package/skills/skill_baru/SKILL.md +0 -172
  175. package/skills/supabase-migration/SKILL.md +0 -45
  176. package/skills/token-saver/SKILL.md +0 -119
  177. package/skills/ui-components-expert/SKILL.md +0 -280
  178. package/skills/ui-ux-pro-max/scripts/__pycache__/core.cpython-310.pyc +0 -0
  179. package/skills/ui-ux-pro-max/scripts/__pycache__/core.cpython-312.pyc +0 -0
  180. package/skills/ui-ux-pro-max/scripts/__pycache__/design_system.cpython-310.pyc +0 -0
  181. package/skills/ui-ux-pro-max/scripts/__pycache__/design_system.cpython-312.pyc +0 -0
  182. package/skills/ui_ux_expert/SKILL.md +0 -125
  183. package/skills/vibe-code-gardener/SKILL.md +0 -181
  184. package/vibes-swarm-demo.gif +0 -0
  185. /package/skills/{saas-transformer → saas-architect}/references/feature_gating_patterns.md +0 -0
  186. /package/skills/{saas-transformer → saas-architect}/references/saas_transformation_checklist.md +0 -0
  187. /package/skills/{saas-transformer → saas-architect}/scripts/saas_transformation_scanner.py +0 -0
@@ -1,207 +1,157 @@
1
- ---
2
- name: web-scraper
3
- description: "Smart web data extraction capability with multi-strategy scraping (Crawl4AI, Playwright, BeautifulSoup), LLM extraction, pagination support, and structured export / Kemampuan ekstraksi data web cerdas dengan strategi scraping modern (Crawl4AI, Playwright, BeautifulSoup), ekstraksi LLM, paginasi, dan ekspor terstruktur."
4
- author: "Roedy Rustam"
5
- ---
6
-
7
- # Web Scraper Expert (2026 Edition)
8
-
9
- [English](#english) | [Bahasa Indonesia](#bahasa-indonesia)
10
-
11
- ---
12
-
13
- <a name="english"></a>
14
- ## English
15
-
16
- ### Description
17
- Smart web data extraction using modern scraping strategies — **Firecrawl** (LLM-ready structured extraction), **Jina Reader API** (clean markdown from any URL), **Crawl4AI** (async Python scraping), **Playwright** (browser automation), and **BeautifulSoup** (lightweight HTML parsing). Supports pagination, dynamic content, structured export (JSON, CSV, Markdown), and AI-powered data extraction.
18
-
19
- ### Trigger Conditions
20
- - Extracting structured data from websites for analysis, training data, or content pipelines.
21
- - Scraping dynamic JavaScript-rendered pages (SPAs, React apps).
22
- - Converting web pages to clean Markdown for LLM context or RAG pipelines.
23
- - Analyzing or cloning website UI design, templates, and components directly from a URL (Handoff to `website-design-cloner`).
24
- - Building an automated data pipeline that scrapes, transforms, and stores data.
25
- - Extracting data at scale with rate limiting and proxy rotation.
26
-
27
- ### Strategy Selection Guide
28
-
29
- | Strategy | Tool | Best For | JS Required |
30
- |---|---|---|---|
31
- | **Managed API** | Firecrawl / Jina Reader | LLM-ready content, clean Markdown | Handled |
32
- | **Async Python** | Crawl4AI | High-volume, AI-powered extraction | Optional |
33
- | **Browser Automation** | Playwright | Complex SPAs, auth-required pages | ✅ |
34
- | **Lightweight HTML** | BeautifulSoup + httpx | Static HTML, fast extraction | ❌ |
35
-
36
- ### Firecrawl — LLM-Ready Web Scraping (2026 Standard)
37
- Firecrawl converts any URL to clean, LLM-ready Markdown with a single API call:
38
- ```python
39
- from firecrawl import FirecrawlApp
40
-
41
- app = FirecrawlApp(api_key="fc-xxxx")
42
-
43
- # Single page — clean Markdown
44
- result = app.scrape_url(
45
- "https://example.com/article",
46
- formats=["markdown", "html"],
47
- only_main_content=True, # removes nav, footer, ads
48
- )
49
- print(result.markdown)
50
-
51
- # Full site crawl
52
- crawl_result = app.crawl_url(
53
- "https://docs.example.com",
54
- limit=50,
55
- scrape_options={"formats": ["markdown"]},
56
- )
57
-
58
- # LLM-powered structured extraction with JSON schema
59
- from pydantic import BaseModel
60
-
61
- class ProductInfo(BaseModel):
62
- name: str
63
- price: float
64
- description: str
65
- in_stock: bool
66
-
67
- result = app.scrape_url(
68
- "https://shop.example.com/product/123",
69
- formats=["extract"],
70
- extract={"schema": ProductInfo.model_json_schema()},
71
- )
72
- product = ProductInfo(**result.extract)
73
- ```
74
-
75
- ### Jina Reader API — Clean Markdown from Any URL
76
- ```python
77
- import httpx
78
-
79
- async def url_to_markdown(url: str) -> str:
80
- """Convert any URL to clean Markdown via Jina Reader API."""
81
- async with httpx.AsyncClient() as client:
82
- response = await client.get(
83
- f"https://r.jina.ai/{url}",
84
- headers={
85
- "Accept": "application/json",
86
- "X-Return-Format": "markdown",
87
- "Authorization": f"Bearer {JINA_API_KEY}",
88
- }
89
- )
90
- data = response.json()
91
- return data["data"]["content"]
92
-
93
- # Search and get results as Markdown
94
- async def search_to_markdown(query: str) -> str:
95
- async with httpx.AsyncClient() as client:
96
- response = await client.get(
97
- f"https://s.jina.ai/{query}",
98
- headers={"Accept": "application/json"}
99
- )
100
- return response.json()["data"]
101
- ```
102
-
103
- ### Crawl4AI — Async Python Web Scraper
104
- ```python
105
- import asyncio
106
- from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode
107
-
108
- async def scrape_with_ai_extraction():
109
- browser_config = BrowserConfig(headless=True, verbose=False)
110
-
111
- async with AsyncWebCrawler(config=browser_config) as crawler:
112
- result = await crawler.arun(
113
- url="https://news.example.com",
114
- config=CrawlerRunConfig(
115
- cache_mode=CacheMode.ENABLED, # cache results
116
- word_count_threshold=50, # skip short content
117
- exclude_external_links=True,
118
- remove_overlay_elements=True, # remove popups/modals
119
- ),
120
- )
121
-
122
- print(result.markdown.fit_markdown) # cleaned, AI-optimized Markdown
123
- print(result.links) # extracted links
124
- ```
125
-
126
- ### Playwright — Complex Dynamic Pages
127
- ```python
128
- from playwright.async_api import async_playwright
129
- import asyncio
130
-
131
- async def scrape_spa(url: str) -> dict:
132
- async with async_playwright() as pw:
133
- browser = await pw.chromium.launch(headless=True)
134
- page = await browser.new_page()
135
-
136
- # Block images/CSS to speed up
137
- await page.route("**/*.{png,jpg,jpeg,gif,css,woff2}", lambda r: r.abort())
138
-
139
- await page.goto(url, wait_until="networkidle")
140
-
141
- # Wait for dynamic content
142
- await page.wait_for_selector("[data-testid='product-list']")
143
-
144
- # Extract structured data via JavaScript
145
- products = await page.evaluate("""() => {
146
- return Array.from(document.querySelectorAll('.product-card')).map(card => ({
147
- name: card.querySelector('.product-name')?.textContent?.trim(),
148
- price: card.querySelector('.price')?.textContent?.trim(),
149
- }));
150
- }""")
151
-
152
- await browser.close()
153
- return {"products": products, "url": url}
154
- ```
155
-
156
- ### Ethical Scraping Checklist
157
- - [ ] Check `robots.txt` and respect `Disallow` rules.
158
- - [ ] Implement rate limiting — minimum 1-2 seconds between requests.
159
- - [ ] Use descriptive `User-Agent` header with contact email.
160
- - [ ] Cache results to avoid repeated requests.
161
- - [ ] Prefer public APIs or official data feeds when available.
162
- - [ ] Do not scrape personal/private data without consent.
163
-
164
- ---
165
-
166
- <a name="bahasa-indonesia"></a>
167
- ## Bahasa Indonesia
168
-
169
- ### Deskripsi
170
- Ekstraksi data web yang cerdas menggunakan strategi scraping modern — **Firecrawl** (ekstraksi terstruktur siap-LLM), **Jina Reader API** (Markdown bersih dari URL manapun), **Crawl4AI** (scraping Python async), **Playwright** (otomasi browser), dan **BeautifulSoup** (parsing HTML ringan). Mendukung paginasi, konten dinamis, ekspor terstruktur, dan ekstraksi data bertenaga AI.
171
-
172
- ### Kondisi Pemicu
173
- - Mengekstrak data terstruktur dari website untuk analisis, data pelatihan, atau pipeline konten.
174
- - Scraping halaman yang dirender JavaScript secara dinamis (SPA, aplikasi React).
175
- - Mengonversi halaman web menjadi Markdown bersih untuk konteks LLM atau pipeline RAG.
176
- - Mempelajari atau menduplikasi desain UI, template, dan komponen situs web langsung dari URL (Delegasikan ke `website-design-cloner`).
177
- - Membangun pipeline data otomatis yang melakukan scraping, transformasi, dan penyimpanan.
178
- - Mengekstrak data dalam skala besar dengan rate limiting dan rotasi proxy.
179
-
180
- ### Panduan Pemilihan Strategi
181
-
182
- | Strategi | Tool | Terbaik Untuk | Perlu JS |
183
- |---|---|---|---|
184
- | **API Terkelola** | Firecrawl / Jina Reader | Konten siap LLM, Markdown bersih | Ditangani |
185
- | **Python Async** | Crawl4AI | Volume tinggi, ekstraksi AI | Opsional |
186
- | **Otomasi Browser** | Playwright | SPA kompleks, halaman butuh auth | ✅ |
187
- | **HTML Ringan** | BeautifulSoup + httpx | HTML statis, ekstraksi cepat | ❌ |
188
-
189
- ### Firecrawl — Scraping Web Siap LLM
190
- Firecrawl mengonversi URL apapun menjadi Markdown bersih siap LLM dengan satu panggilan API. Mendukung crawling seluruh situs, ekstraksi terstruktur berbasis skema JSON, dan penghapusan konten yang tidak relevan (navigasi, footer, iklan).
191
-
192
- ### Jina Reader API — Markdown Bersih dari URL Manapun
193
- Jina Reader (`r.jina.ai/{url}`) mengonversi halaman web manapun menjadi Markdown yang dioptimalkan untuk LLM. Jina Search (`s.jina.ai/{query}`) melakukan pencarian web dan mengembalikan hasilnya sebagai Markdown.
194
-
195
- ### Crawl4AI — Web Scraper Python Async
196
- Crawl4AI mendukung caching hasil, ekstraksi Markdown yang dioptimalkan AI, dan konfigurasi browser yang fleksibel untuk scraping konten dinamis.
197
-
198
- ### Playwright — Halaman Dinamis Kompleks
199
- Gunakan Playwright untuk scraping SPA yang memerlukan eksekusi JavaScript, menunggu elemen dinamis, atau interaksi dengan halaman (klik, scroll, isi form).
200
-
201
- ### Checklist Scraping Etis
202
- - [ ] Periksa `robots.txt` dan hormati aturan `Disallow`.
203
- - [ ] Implementasikan rate limiting — minimal 1-2 detik antar permintaan.
204
- - [ ] Gunakan header `User-Agent` deskriptif dengan email kontak.
205
- - [ ] Cache hasil untuk menghindari permintaan berulang.
206
- - [ ] Utamakan API publik atau feed data resmi jika tersedia.
207
- - [ ] Jangan scraping data pribadi/privat tanpa izin.
1
+ ---
2
+ name: web-scraper
3
+ description: "Smart agentic web data extraction with multi-strategy scraping (Crawl4AI v4, Firecrawl), LLM extraction loops, anti-bot bypass, and structured export / Ekstraksi data web cerdas dan agentic dengan scraping multi-strategi (Crawl4AI v4, Firecrawl), ekstraksi LLM, bypass anti-bot, dan ekspor terstruktur."
4
+ author: "Roedy Rustam"
5
+ ---
6
+
7
+ # Agentic Web Scraper Expert (2026 Edition)
8
+
9
+ [English](#english) | [Bahasa Indonesia](#bahasa-indonesia)
10
+
11
+ ---
12
+
13
+ <a name="english"></a>
14
+ ## English
15
+
16
+ ### Orchestration & Integration
17
+ Connects and orchestrates with relevant domain skills like `browser-automation-expert`, `ai-llm-integration-expert`, `brainstorming`, and `zero-to-prod-orchestrator` to ensure cohesive agentic execution.
18
+
19
+ ### Description
20
+ Advanced Agentic Web Scraping utilizing modern multi-strategy data extraction. Leverages **Crawl4AI v4** and **Firecrawl** to convert raw DOMs into LLM-friendly Markdown. Implements Agentic Extraction loops where the LLM guides the scraper dynamically based on page state. Incorporates strategies for bypassing anti-bot measures (Cloudflare Turnstile, Datadome) and navigating dynamic Shadow DOMs.
21
+
22
+ ### Trigger Conditions
23
+ - Extracting structured data from websites for analysis, training data, or content pipelines.
24
+ - Scraping dynamic JavaScript-rendered pages and complex SPAs.
25
+ - Converting web pages to clean Markdown for LLM context or RAG pipelines.
26
+ - Dealing with anti-bot protections or complex Shadow DOM architectures during scraping.
27
+ - Implementing an automated agentic data extraction loop.
28
+
29
+ ### Extracting DOM into LLM-Friendly Markdown
30
+ Use **Crawl4AI v4** for high-performance async extraction and **Firecrawl** for seamless LLM-ready conversion.
31
+
32
+ **Crawl4AI v4 (Async Python):**
33
+ ```python
34
+ import asyncio
35
+ from crawl4ai import AsyncWebCrawler, BrowserConfig, CrawlerRunConfig, CacheMode
36
+
37
+ async def extract_markdown(url: str):
38
+ config = BrowserConfig(headless=True, bypass_csp=True)
39
+ run_config = CrawlerRunConfig(
40
+ cache_mode=CacheMode.ENABLED,
41
+ remove_overlay_elements=True,
42
+ word_count_threshold=50
43
+ )
44
+
45
+ async with AsyncWebCrawler(config=config) as crawler:
46
+ result = await crawler.arun(url=url, config=run_config)
47
+ # Returns clean, AI-optimized markdown ready for LLM consumption
48
+ return result.markdown.fit_markdown
49
+ ```
50
+
51
+ **Firecrawl (Managed API):**
52
+ ```python
53
+ from firecrawl import FirecrawlApp
54
+ from pydantic import BaseModel
55
+
56
+ app = FirecrawlApp(api_key="fc-xxxx")
57
+
58
+ class ExtractionSchema(BaseModel):
59
+ title: str
60
+ content: str
61
+ key_metrics: list[str]
62
+
63
+ # Single API call to extract structured data based on JSON schema
64
+ result = app.scrape_url(
65
+ "https://example.com/data",
66
+ formats=["extract", "markdown"],
67
+ extract={"schema": ExtractionSchema.model_json_schema()}
68
+ )
69
+ print(result.markdown) # Clean markdown
70
+ print(result.extract) # Structured JSON
71
+ ```
72
+
73
+ ### Anti-Bot Bypass & Shadow DOMs
74
+ Scraping modern web apps requires bypassing anti-bot measures like Cloudflare Turnstile and Datadome, as well as accessing deeply nested elements.
75
+
76
+ 1. **Anti-Bot Bypass (Cloudflare Turnstile, Datadome):**
77
+ - **Residential Proxies:** Rotate high-quality residential IPs to avoid datacenter IP bans.
78
+ - **Browser Fingerprinting:** Use tools like `playwright-stealth` or specialized stealth browsers (e.g., Undetected ChromeDriver, Curl-Impersonate) to mask automated fingerprints (WebGL, Canvas, User-Agent).
79
+ - **Human-like Interaction:** Introduce random delays, simulate realistic mouse movements, and handle CAPTCHAs via third-party solving services only when necessary.
80
+ 2. **Dynamic Shadow DOMs:**
81
+ - Use CSS piercing selectors or JavaScript execution to penetrate the Shadow Root.
82
+ - Example (Playwright): `await page.locator('my-web-component >> css=.internal-element').text_content()`
83
+ - Recursively traverse the DOM tree injecting scripts to extract content from encapsulated components.
84
+
85
+ ### Agentic Extraction Loops
86
+ Implement an autonomous loop where an LLM guides the scraper based on the current page state, rather than relying on brittle CSS selectors.
87
+
88
+ 1. **Observe:** The scraper extracts the current DOM into clean Markdown.
89
+ 2. **Analyze:** The LLM analyzes the Markdown to identify necessary data or the next interaction step (e.g., "Click the 'Load More' button").
90
+ 3. **Act:** The LLM issues a command (extract data, navigate, click, fill form).
91
+ 4. **Loop:** Repeat until the extraction goal is met.
92
+
93
+ ```python
94
+ async def agentic_scrape_loop(url: str, goal: str):
95
+ current_url = url
96
+ while True:
97
+ markdown_content = await extract_markdown(current_url)
98
+ # LLM analyzes state and decides next action
99
+ action = await llm_decide_action(markdown_content, goal)
100
+
101
+ if action.type == "COMPLETE":
102
+ return action.extracted_data
103
+ elif action.type == "CLICK":
104
+ await click_element(action.target_selector)
105
+ elif action.type == "NAVIGATE":
106
+ current_url = action.new_url
107
+ ```
108
+
109
+ ### Ethical Scraping Checklist
110
+ - [ ] Check `robots.txt` and respect `Disallow` rules.
111
+ - [ ] Implement rate limiting.
112
+ - [ ] Use descriptive `User-Agent` headers.
113
+ - [ ] Do not scrape personal/private data without consent.
114
+
115
+ ---
116
+
117
+ <a name="bahasa-indonesia"></a>
118
+ ## Bahasa Indonesia
119
+
120
+ ### Integrasi Orkestrasi
121
+ Terhubung dan mengorkestrasi skill domain yang relevan seperti `browser-automation-expert`, `ai-llm-integration-expert`, `brainstorming`, dan `zero-to-prod-orchestrator` untuk memastikan eksekusi agentic yang kohesif.
122
+
123
+ ### Deskripsi
124
+ Scraping Web Agentic tingkat lanjut menggunakan ekstraksi data multi-strategi modern. Memanfaatkan **Crawl4AI v4** dan **Firecrawl** untuk mengubah DOM mentah menjadi Markdown yang ramah LLM. Mengimplementasikan loop Ekstraksi Agentic di mana LLM memandu scraper secara dinamis berdasarkan status halaman. Menggabungkan strategi untuk melewati tindakan anti-bot (Cloudflare Turnstile, Datadome) dan menavigasi Shadow DOM yang dinamis.
125
+
126
+ ### Kondisi Pemicu
127
+ - Mengekstrak data terstruktur dari situs web untuk analisis, data pelatihan, atau pipeline konten.
128
+ - Scraping halaman yang dirender JavaScript secara dinamis dan SPA kompleks.
129
+ - Mengonversi halaman web menjadi Markdown bersih untuk konteks LLM atau pipeline RAG.
130
+ - Menghadapi perlindungan anti-bot atau arsitektur Shadow DOM yang kompleks saat scraping.
131
+ - Mengimplementasikan loop ekstraksi data agentic otomatis.
132
+
133
+ ### Mengekstrak DOM menjadi Markdown Ramah LLM
134
+ Gunakan **Crawl4AI v4** untuk ekstraksi async berperforma tinggi dan **Firecrawl** untuk konversi siap LLM yang mulus. (Lihat contoh kode di bagian bahasa Inggris).
135
+
136
+ ### Bypass Anti-Bot & Shadow DOM
137
+ 1. **Bypass Anti-Bot (Cloudflare Turnstile, Datadome):**
138
+ - **Proxy Residensial:** Rotasi IP residensial berkualitas tinggi untuk menghindari pemblokiran IP datacenter.
139
+ - **Browser Fingerprinting:** Gunakan alat seperti `playwright-stealth` atau browser stealth khusus untuk menyembunyikan sidik jari otomatis.
140
+ - **Interaksi Mirip Manusia:** Tambahkan penundaan acak, simulasikan gerakan mouse yang realistis.
141
+ 2. **Shadow DOM Dinamis:**
142
+ - Gunakan selektor penembus CSS atau eksekusi JavaScript untuk menembus Shadow Root.
143
+ - Telusuri pohon DOM secara rekursif dengan menyuntikkan skrip untuk mengekstrak konten.
144
+
145
+ ### Loop Ekstraksi Agentic
146
+ Implementasikan loop otonom di mana LLM memandu scraper berdasarkan status halaman saat ini, bukan bergantung pada selektor CSS yang rentan rusak.
147
+
148
+ 1. **Observasi:** Scraper mengekstrak DOM saat ini menjadi Markdown yang bersih.
149
+ 2. **Analisis:** LLM menganalisis Markdown untuk mengidentifikasi data yang diperlukan atau langkah interaksi selanjutnya (misal: "Klik tombol 'Muat Lebih Banyak'").
150
+ 3. **Aksi:** LLM mengeluarkan perintah (ekstrak data, navigasi, klik, isi form).
151
+ 4. **Loop:** Ulangi hingga tujuan ekstraksi tercapai.
152
+
153
+ ### Checklist Scraping Etis
154
+ - [ ] Periksa `robots.txt` dan hormati aturan `Disallow`.
155
+ - [ ] Implementasikan rate limiting.
156
+ - [ ] Gunakan header `User-Agent` yang deskriptif.
157
+ - [ ] Jangan scraping data pribadi/privat tanpa izin.