jarvis-ai-framework 1.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (240) hide show
  1. package/AGENTS.md +416 -0
  2. package/LICENSE +21 -0
  3. package/README.md +190 -0
  4. package/agents/AGENTS.md +234 -0
  5. package/agents/README.md +309 -0
  6. package/agents/engineering/data/eng.data-engineer.agent.md +309 -0
  7. package/agents/engineering/eng.agent.md +303 -0
  8. package/agents/engineering/eng.bug-hunter.md +386 -0
  9. package/agents/engineering/eng.cybersecurity.agent.md +503 -0
  10. package/agents/engineering/eng.dev-code-reviewer.md +148 -0
  11. package/agents/engineering/eng.docs-writer.md +152 -0
  12. package/agents/engineering/eng.frontend.agent.md +117 -0
  13. package/agents/engineering/eng.rpa.agent.md +215 -0
  14. package/agents/engineering/eng.tech-analyst.agent.md +102 -0
  15. package/agents/engineering/eng.ux-designer.agent.md +193 -0
  16. package/agents/engineering/qa/eng.qa.cypress-specialist.md +109 -0
  17. package/agents/engineering/qa/eng.qa.quality-champion-task-agent.md +85 -0
  18. package/agents/engineering/qa/eng.qa.quality-strategist.md +111 -0
  19. package/agents/engineering/qa/eng.qa.test-architect.md +400 -0
  20. package/agents/engineering/qa/eng.qa.test-planner.md +477 -0
  21. package/agents/engineering/qa/eng.qa.testing-engineer.md +339 -0
  22. package/agents/product/prod.pm-checker.md +52 -0
  23. package/bin/commands/docs-publish.js +184 -0
  24. package/bin/commands/docs-sync.js +139 -0
  25. package/bin/commands/info.js +87 -0
  26. package/bin/commands/init.js +237 -0
  27. package/bin/commands/install-rtk.js +90 -0
  28. package/bin/commands/list.js +48 -0
  29. package/bin/commands/qa-signoff.js +112 -0
  30. package/bin/commands/whoami.js +43 -0
  31. package/bin/jarvis.js +159 -0
  32. package/bin/lib/auth/session.js +56 -0
  33. package/bin/lib/config/constants.js +123 -0
  34. package/bin/lib/config/ide-config.js +233 -0
  35. package/bin/lib/core/scanner.js +124 -0
  36. package/bin/lib/core/sync-engine.js +551 -0
  37. package/bin/lib/docs/fetch-file.sh +41 -0
  38. package/bin/lib/docs/publish-file.sh +284 -0
  39. package/bin/lib/docs/validate-frontmatter.js +157 -0
  40. package/bin/lib/env-loader.js +198 -0
  41. package/bin/lib/tasks/comment.js +131 -0
  42. package/bin/lib/utils/git-parser.js +145 -0
  43. package/bin/lib/utils/logger.js +104 -0
  44. package/bin/lib/utils/npmrc-parser.js +106 -0
  45. package/bin/lib/utils/paths.js +55 -0
  46. package/bin/lib/utils/ui.js +59 -0
  47. package/bin/lib/vcs/api.js +312 -0
  48. package/bin/lib/vcs/create-issue.js +43 -0
  49. package/bin/lib/vcs/create-merge.js +43 -0
  50. package/bin/lib/vcs/fetch-raw.js +30 -0
  51. package/bin/postinstall.js +41 -0
  52. package/members.md +25 -0
  53. package/package.json +55 -0
  54. package/rules/AGENTS.md +205 -0
  55. package/rules/engineering/data/data-rules.md +200 -0
  56. package/rules/engineering/eng-rules.md +243 -0
  57. package/rules/engineering/eng-security-rules.md +186 -0
  58. package/rules/engineering/eng.breakdown-subtasks-rules.md +585 -0
  59. package/rules/engineering/eng.bump-rules.md +27 -0
  60. package/rules/engineering/eng.docs-scraping-rules.md +64 -0
  61. package/rules/engineering/eng.downstream-flow-rules.md +297 -0
  62. package/rules/engineering/eng.integrations-rules.md +73 -0
  63. package/rules/engineering/eng.plan-rules.md +333 -0
  64. package/rules/engineering/eng.pr-rules.md +359 -0
  65. package/rules/engineering/eng.pre-pr-rules.md +103 -0
  66. package/rules/engineering/eng.start-rules.md +246 -0
  67. package/rules/engineering/eng.tech-spec-rules.md +968 -0
  68. package/rules/engineering/eng.work-rules.md +312 -0
  69. package/rules/engineering/frontend/eng.frontend-rules.md +147 -0
  70. package/rules/engineering/qa/eng.qa.cypress-standards-rules.md +259 -0
  71. package/rules/engineering/qa/eng.qa.exploratory-session-rules.md +137 -0
  72. package/rules/engineering/qa/eng.qa.quality-gate-scoring-rules.md +181 -0
  73. package/rules/engineering/qa/eng.qa.tech-spec-validation-criteria-rules.md +120 -0
  74. package/rules/engineering/rpa/eng.rpa-rules.md +230 -0
  75. package/rules/product/README.md +24 -0
  76. package/rules/product/prod-rules.md +151 -0
  77. package/rules/rtk-rules.md +68 -0
  78. package/skills/AGENTS.md +290 -0
  79. package/skills/SKILLS-ROADMAP.md +333 -0
  80. package/skills/churn-audit/SKILL.md +385 -0
  81. package/skills/context-detect/SKILL.md +399 -0
  82. package/skills/context-detect/assets/context-profile-template.md +127 -0
  83. package/skills/docs-central/README.md +310 -0
  84. package/skills/docs-central/SKILL.md +423 -0
  85. package/skills/docs-index/SKILL.md +377 -0
  86. package/skills/eng-ai-engineer/SKILL.md +296 -0
  87. package/skills/eng-arch-c4/SKILL.md +358 -0
  88. package/skills/eng-arch-c4/assets/example-code.md +189 -0
  89. package/skills/eng-arch-c4/assets/example-component.md +105 -0
  90. package/skills/eng-arch-c4/assets/example-container.md +104 -0
  91. package/skills/eng-arch-c4/assets/example-context.md +81 -0
  92. package/skills/eng-backend/SKILL.md +776 -0
  93. package/skills/eng-browser-extension-builder/SKILL.md +385 -0
  94. package/skills/eng-cybersecurity/SKILL.md +645 -0
  95. package/skills/eng-data-bi/SKILL.md +199 -0
  96. package/skills/eng-data-debug/SKILL.md +307 -0
  97. package/skills/eng-data-engineer/SKILL.md +256 -0
  98. package/skills/eng-data-onboard/SKILL.md +310 -0
  99. package/skills/eng-data-orchestrator/SKILL.md +426 -0
  100. package/skills/eng-design-system/SKILL.md +619 -0
  101. package/skills/eng-docs-write/SKILL.md +312 -0
  102. package/skills/eng-frontend/SKILL.md +913 -0
  103. package/skills/eng-jira-comment/SKILL.md +17 -0
  104. package/skills/eng-microfrontend/SKILL.md +602 -0
  105. package/skills/eng-ms-trace/SKILL.md +469 -0
  106. package/skills/eng-nestjs/SKILL.md +791 -0
  107. package/skills/eng-performance-engineer/SKILL.md +312 -0
  108. package/skills/eng-pr/SKILL.md +339 -0
  109. package/skills/eng-qa-a11y-audit/SKILL.md +269 -0
  110. package/skills/eng-qa-bug-report/SKILL.md +1088 -0
  111. package/skills/eng-qa-bug-report/TASK_MANAGERS.md +138 -0
  112. package/skills/eng-qa-cypress-e2e/SKILL.md +177 -0
  113. package/skills/eng-qa-dev-guide/SKILL.md +164 -0
  114. package/skills/eng-qa-e2e/SKILL.md +400 -0
  115. package/skills/eng-qa-e2e-spec-writer/SKILL.md +322 -0
  116. package/skills/eng-qa-exploratory/SKILL.md +188 -0
  117. package/skills/eng-qa-gate/SKILL.md +370 -0
  118. package/skills/eng-qa-gate/assets/checklist-validacao.md +291 -0
  119. package/skills/eng-qa-graphql-contract/SKILL.md +256 -0
  120. package/skills/eng-qa-quality-report/SKILL.md +412 -0
  121. package/skills/eng-qa-test-plan/SKILL.md +466 -0
  122. package/skills/eng-qa-test-plan/assets/test-coverage-template.md +92 -0
  123. package/skills/eng-qa-test-plan/assets/test-patterns.md +178 -0
  124. package/skills/eng-qa-testsprite/SKILL.md +325 -0
  125. package/skills/eng-qa-testsprite/references/testsprite-mcp.md +224 -0
  126. package/skills/eng-qa-unit-test/SKILL.md +471 -0
  127. package/skills/eng-rabbitmq/SKILL.md +661 -0
  128. package/skills/eng-scraper/SKILL.md +683 -0
  129. package/skills/eng-scraper-robot-builder/SKILL.md +370 -0
  130. package/skills/eng-security-patch/SKILL.md +378 -0
  131. package/skills/eng-security-triage/SKILL.md +266 -0
  132. package/skills/eng-task-comment/SKILL.md +60 -0
  133. package/skills/eng-tech-analyst/SKILL.md +529 -0
  134. package/skills/eng-threat-model/SKILL.md +161 -0
  135. package/skills/init-jarvis/SKILL.md +1304 -0
  136. package/skills/init-jarvis/assets/mcp-configs.md +389 -0
  137. package/skills/init-jarvis/assets/onboarding-checklist.md +104 -0
  138. package/skills/init-jarvis/assets/setup-guide.md +360 -0
  139. package/skills/lovable-prompt-generator/SKILL.md +304 -0
  140. package/skills/prod-roadmap-report/README.md +303 -0
  141. package/skills/prod-roadmap-report/SKILL.md +198 -0
  142. package/skills/prod-roadmap-report/commands/status.compiled.single.team.md +23 -0
  143. package/skills/prod-roadmap-report/commands/status.list.projects.md +17 -0
  144. package/skills/prod-roadmap-report/commands/status.memory.md +192 -0
  145. package/skills/prod-roadmap-report/commands/status.roadmap.preview.md +94 -0
  146. package/skills/prod-roadmap-report/references/detailed-guide.md +236 -0
  147. package/skills/prod-roadmap-report/rules/detailed-guide.md +237 -0
  148. package/skills/prod-roadmap-report/rules/status-report-rules.md +44 -0
  149. package/skills/prod-roadmap-report/templates/template-multiple-teams-compiled-status.md +53 -0
  150. package/skills/prod-roadmap-report/templates/template-projects-list.md +23 -0
  151. package/skills/prod-roadmap-report/templates/template-single-team-compiled-status.md +60 -0
  152. package/skills/prod-roadmap-report/templates/template-single-team-status.md +49 -0
  153. package/skills/prod-specs/SKILL.md +108 -0
  154. package/skills/prod-specs/references/prod.spec.clarify.md +176 -0
  155. package/skills/prod-specs/references/prod.spec.epic.md +107 -0
  156. package/skills/prod-specs/references/prod.spec.frd.md +135 -0
  157. package/skills/prod-specs/references/prod.spec.issue.md +145 -0
  158. package/skills/prod-specs/references/prod.spec.prd.md +118 -0
  159. package/skills/prod-specs/rules/prod-spec-rules.md +186 -0
  160. package/skills/prod-specs/templates/prod-breakdown-template.md +136 -0
  161. package/skills/prod-specs/templates/prod-epic-template.md +76 -0
  162. package/skills/prod-specs/templates/prod-frd-template.md +172 -0
  163. package/skills/prod-specs/templates/prod-issue-template.md +68 -0
  164. package/skills/prod-specs/templates/prod-prd-full-template.md +159 -0
  165. package/skills/prod-specs/templates/prod-prd-template.md +173 -0
  166. package/skills/prod-specs-update/SKILL.md +272 -0
  167. package/skills/report-issue/SKILL.md +156 -0
  168. package/taxonomy.md +270 -0
  169. package/templates/AGENTS.md +189 -0
  170. package/templates/CDD aplicado a Prompts.md +182 -0
  171. package/templates/ENV-template.md +187 -0
  172. package/templates/engineering/AGENTS-template.md +71 -0
  173. package/templates/engineering/ARD-template.md +193 -0
  174. package/templates/engineering/CONTACTS-template.md +135 -0
  175. package/templates/engineering/PR-template.md +40 -0
  176. package/templates/engineering/RFC-Playbook.md +325 -0
  177. package/templates/engineering/RFC-template.md +199 -0
  178. package/templates/engineering/architecture-template.md +277 -0
  179. package/templates/engineering/breakdown-subtasks-template.md +582 -0
  180. package/templates/engineering/c4-model-template.md +516 -0
  181. package/templates/engineering/data-contract-template.md +135 -0
  182. package/templates/engineering/data-pipeline-template.md +163 -0
  183. package/templates/engineering/plan-template.md +255 -0
  184. package/templates/engineering/qa/eng.qa.quality-gate-examples-template.md +311 -0
  185. package/templates/engineering/qa/eng.qa.quality-gate-report-template.md +249 -0
  186. package/templates/engineering/qa/qa.cypress-test-template.md +172 -0
  187. package/templates/engineering/qa/qa.exploratory-session-template.md +148 -0
  188. package/templates/engineering/qa/qa.quality-report-template.md +130 -0
  189. package/templates/engineering/qa/qa.release-signoff-template.md +54 -0
  190. package/templates/engineering/qa/qa.sprint-plan-template.md +49 -0
  191. package/templates/engineering/swagger-template.md +145 -0
  192. package/templates/engineering/tech-spec-template.md +497 -0
  193. package/templates/engineering/work-progress-template.md +155 -0
  194. package/workflows/AGENTS.md +240 -0
  195. package/workflows/README.md +160 -0
  196. package/workflows/all-tools.md +11 -0
  197. package/workflows/engineering/data/data.contract.md +202 -0
  198. package/workflows/engineering/data/data.new-pipeline.md +234 -0
  199. package/workflows/engineering/eng.breakdown-subtasks.md +420 -0
  200. package/workflows/engineering/eng.bug-audit.md +591 -0
  201. package/workflows/engineering/eng.build-tech-spec.md +1116 -0
  202. package/workflows/engineering/eng.create-ard-from-code.md +259 -0
  203. package/workflows/engineering/eng.create-ard.md +382 -0
  204. package/workflows/engineering/eng.create-rfc.md +245 -0
  205. package/workflows/engineering/eng.debug.md +479 -0
  206. package/workflows/engineering/eng.docs.md +40 -0
  207. package/workflows/engineering/eng.light-arch.md +84 -0
  208. package/workflows/engineering/eng.plan.md +213 -0
  209. package/workflows/engineering/eng.pr.md +466 -0
  210. package/workflows/engineering/eng.pre-pr.md +167 -0
  211. package/workflows/engineering/eng.review.md +185 -0
  212. package/workflows/engineering/eng.rpa.robot.md +342 -0
  213. package/workflows/engineering/eng.security-audit.md +312 -0
  214. package/workflows/engineering/eng.security-incident.md +275 -0
  215. package/workflows/engineering/eng.security-pipeline.md +210 -0
  216. package/workflows/engineering/eng.security-review.md +235 -0
  217. package/workflows/engineering/eng.start.md +494 -0
  218. package/workflows/engineering/eng.work.md +558 -0
  219. package/workflows/engineering/frontend/eng.frontend-component.md +190 -0
  220. package/workflows/engineering/frontend/eng.frontend-perf-audit.md +375 -0
  221. package/workflows/engineering/frontend/eng.frontend-review.md +185 -0
  222. package/workflows/engineering/qa/eng.qa-dev-quality-guide.md +51 -0
  223. package/workflows/engineering/qa/eng.qa-e2e-test-generation.md +51 -0
  224. package/workflows/engineering/qa/eng.qa-exploratory-session.md +60 -0
  225. package/workflows/engineering/qa/eng.qa-quality-gate-validation.md +202 -0
  226. package/workflows/engineering/qa/eng.qa-quality-report.md +83 -0
  227. package/workflows/engineering/qa/eng.qa-refinement-entry.md +83 -0
  228. package/workflows/engineering/qa/eng.qa-release-signoff.md +170 -0
  229. package/workflows/engineering/qa/eng.qa-sprint-planning.md +100 -0
  230. package/workflows/engineering/ta/eng.ta.atendimento.md +93 -0
  231. package/workflows/product/prod.roadmap.preview.md +110 -0
  232. package/workflows/product/prod.spec.breakdown.md +163 -0
  233. package/workflows/product/prod.spec.clarify.md +178 -0
  234. package/workflows/product/prod.spec.epic.md +154 -0
  235. package/workflows/product/prod.spec.frd.md +96 -0
  236. package/workflows/product/prod.spec.issue.md +145 -0
  237. package/workflows/product/prod.spec.md +60 -0
  238. package/workflows/product/prod.spec.prd.md +100 -0
  239. package/workflows/taxonomy.md +92 -0
  240. package/workflows/warm-up.md +574 -0
@@ -0,0 +1,683 @@
1
+ ---
2
+ name: eng-scraper
3
+ description: >
4
+ Especialista em web scraping, extração de dados, automação de browser e parsing de HTML/XML/PDF.
5
+ Domina Puppeteer (principal), Playwright, Cheerio, anti-bot e pipelines ETL leves com NestJS e TypeScript.
6
+ Trigger: Use para web scraping, headless browser, parsing de HTML/XML, extração de dados ou automação de navegação.
7
+ license: AGPL-3.0
8
+ compatibility: Designed for Claude Code (or similar products)
9
+ allowed-tools: Read Write Edit Grep Glob Bash
10
+ metadata:
11
+ author: jarvis-team
12
+ version: "1.0"
13
+ # Campos Claude Code-specific (não fazem parte da spec oficial agentskills.io):
14
+ argument-hint: "[url|site|tarefa] [contexto]"
15
+ disable-model-invocation: false
16
+ ---
17
+
18
+ # Eng Scraper - Especialista em Extração de Dados
19
+
20
+ Você é um **especialista em web scraping e extração de dados** com domínio em browsers headless, parsing de conteúdo, contorno de mecanismos anti-bot e construção de pipelines de dados resilientes.
21
+
22
+ ## Objetivo
23
+
24
+ Extrair dados estruturados de fontes web de forma eficiente, resiliente e ética — desde scripts simples de coleta até pipelines completos de ETL leve.
25
+
26
+ ## Entrada
27
+
28
+ - `$ARGUMENTS` - URL alvo, site ou tarefa de extração (ex: `scraper-precos-ecommerce`, `extrair-tabela-pdf`, `monitorar-vagas`, `pipeline-noticias`)
29
+
30
+ ## Recursos
31
+
32
+ - **ENV**: `$IDE/ENV.md` (variáveis de ambiente do projeto)
33
+ - **Saída**: scripts de scraping TypeScript como `@Injectable()` NestJS no repositório atual
34
+
35
+ ---
36
+
37
+ ## Pré-requisito
38
+
39
+ Verificar robots.txt e Terms of Service do site alvo ANTES de qualquer implementação:
40
+
41
+ ```bash
42
+ # Verificar robots.txt
43
+ curl -s "{url-alvo}/robots.txt"
44
+ ```
45
+
46
+ Se houver restrições legais ou éticas significativas, comunicar ao usuário antes de prosseguir.
47
+
48
+ ---
49
+
50
+ ## Quando Usar
51
+
52
+ Use este skill quando:
53
+ - Extrair dados de páginas web (preços, produtos, notícias, vagas, tabelas)
54
+ - Automatizar navegação em browsers (preenchimento de formulários, login, screenshots)
55
+ - Fazer parsing de HTML, XML, JSON ou PDFs
56
+ - Construir pipelines de extração → transformação → carga (ETL leve)
57
+ - Lidar com sites protegidos por mecanismos anti-bot
58
+ - Monitorar mudanças em páginas ou conjuntos de dados web
59
+
60
+ **NÃO usar quando:**
61
+ - A extração pode ser feita via API oficial — sempre preferir API sobre scraping
62
+ - O site proíbe scraping explicitamente e o caso de uso não é legítimo
63
+ - A tarefa é de backend genérico sem extração de dados web
64
+
65
+ ---
66
+
67
+ ## Validação de Entrada
68
+
69
+ Se $ARGUMENTS está vazio:
70
+ → Solicitar ao usuário: URL ou site alvo, tipo de dado a extrair, formato de saída desejado
71
+ → Verificar se existe API oficial antes de prosseguir com scraping
72
+
73
+ ---
74
+
75
+ ## Padrões Críticos
76
+
77
+ ### Padrão 1: Escolher a Ferramenta Certa
78
+
79
+ ```
80
+ Site renderizado com JavaScript? → Puppeteer (padrão do projeto) ou Playwright
81
+ Site com HTML estático? → Cheerio (mais rápido, menor overhead)
82
+ APIs internas (XHR/fetch)? → Interceptar requests com Puppeteer → mais estável
83
+ Dados em PDFs? → pdf-parse, pdfjs-dist
84
+ Dados em XML/RSS? → fast-xml-parser, xml2js
85
+ Dados em CSVs/planilhas? → csv-parse, xlsx
86
+ ```
87
+
88
+ ### Padrão 2: Resiliência por Padrão
89
+
90
+ ```
91
+ - Nunca confiar na estrutura do HTML → pode mudar a qualquer momento
92
+ - Sempre verificar se elementos existem antes de extrair
93
+ - Logar estruturas inesperadas para detectar mudanças de layout
94
+ - Retry automático para falhas transitórias de rede
95
+ - Timeout em todas as operações de rede e navegação
96
+ ```
97
+
98
+ ### Padrão 3: Responsabilidade com o Servidor Alvo
99
+
100
+ ```
101
+ - Rate limiting: mínimo 1-2 segundos entre requests
102
+ - Randomizar delays para parecer mais orgânico
103
+ - Não escalar paralelismo sem avaliar impacto
104
+ - Identificar o scraper via User-Agent quando possível
105
+ - Caching: não re-baixar dados que já foram coletados
106
+ ```
107
+
108
+ ---
109
+
110
+ ## Árvore de Decisão
111
+
112
+ ```
113
+ Site tem API oficial? → Usar API (não scraping)
114
+ HTML estático sem JS? → Cheerio / node-fetch
115
+ HTML renderizado com JS? → Puppeteer (padrão) / Playwright
116
+ Precisar fazer login? → Puppeteer + session/cookie management
117
+ API interna interceptável? → Puppeteer request interception
118
+ Dados em PDF? → pdf-parse / pdfjs-dist
119
+ Dados em XML/RSS? → fast-xml-parser
120
+ Site com anti-bot avançado? → Seção: Técnicas Anti-Bot
121
+ Pipeline de múltiplas fontes? → Seção: Pipelines ETL
122
+ ```
123
+
124
+ ---
125
+
126
+ ## Fluxo de Trabalho
127
+
128
+ ### Ética e Aspectos Legais — Verificar Primeiro
129
+
130
+ Antes de qualquer implementação, avaliar:
131
+
132
+ ```
133
+ 1. Existe API oficial? → Usar API sempre que disponível
134
+ 2. robots.txt permite o acesso? → Verificar e respeitar
135
+ 3. Terms of Service proíbem scraping? → Avaliar com o usuário
136
+ 4. Os dados são públicos? → Dados pessoais exigem atenção especial (LGPD/GDPR)
137
+ 5. Qual o impacto no servidor alvo? → Rate limiting generoso, nunca DDoS
138
+ ```
139
+
140
+ ```bash
141
+ # Verificar robots.txt antes de começar
142
+ curl -s "https://exemplo.com/robots.txt"
143
+ ```
144
+
145
+ > Se houver restrições legais ou éticas significativas, comunicar ao usuário antes de prosseguir.
146
+
147
+ ### Puppeteer — Ferramenta Principal (padrão do projeto)
148
+
149
+ Puppeteer é a ferramenta padrão para scraping com browser headless. Todo código de scraping deve ser em TypeScript e, quando integrado ao sistema, implementado como `@Injectable()` NestJS.
150
+
151
+ #### Setup básico
152
+
153
+ ```typescript
154
+ import puppeteer, { Browser, Page } from 'puppeteer'
155
+
156
+ async function createBrowser(): Promise<Browser> {
157
+ return puppeteer.launch({
158
+ headless: true,
159
+ args: [
160
+ '--no-sandbox',
161
+ '--disable-setuid-sandbox',
162
+ '--disable-blink-features=AutomationControlled',
163
+ ],
164
+ })
165
+ }
166
+
167
+ // ✅ Sempre fechar browser após uso
168
+ async function withBrowser<T>(fn: (browser: Browser) => Promise<T>): Promise<T> {
169
+ const browser = await createBrowser()
170
+ try {
171
+ return await fn(browser)
172
+ } finally {
173
+ await browser.close()
174
+ }
175
+ }
176
+ ```
177
+
178
+ #### NestJS — ScraperService como @Injectable
179
+
180
+ ```typescript
181
+ import { Injectable, Logger } from '@nestjs/common'
182
+ import puppeteer, { Browser, Page } from 'puppeteer'
183
+
184
+ @Injectable()
185
+ export class ScraperService {
186
+ private readonly logger = new Logger(ScraperService.name)
187
+
188
+ async scrapeProducts(url: string): Promise<Product[]> {
189
+ const browser = await puppeteer.launch({ headless: true, args: ['--no-sandbox'] })
190
+
191
+ try {
192
+ const page = await browser.newPage()
193
+ await page.setUserAgent('Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) ...')
194
+
195
+ await page.goto(url, { waitUntil: 'networkidle2', timeout: 30_000 })
196
+ await page.waitForSelector('[data-testid="product-card"]', { timeout: 10_000 })
197
+
198
+ const products = await page.evaluate(() =>
199
+ Array.from(document.querySelectorAll('[data-testid="product-card"]')).map((card) => ({
200
+ name: card.querySelector('h2')?.textContent?.trim() ?? null,
201
+ price: card.querySelector('[data-price]')?.getAttribute('data-price') ?? null,
202
+ url: (card.querySelector('a') as HTMLAnchorElement)?.href ?? null,
203
+ }))
204
+ )
205
+
206
+ const valid = products.filter((p) => p.name && p.price)
207
+ if (valid.length < products.length) {
208
+ this.logger.warn(`${products.length - valid.length} produtos com dados incompletos ignorados`)
209
+ }
210
+
211
+ return valid as Product[]
212
+ } finally {
213
+ await browser.close()
214
+ }
215
+ }
216
+ }
217
+ ```
218
+
219
+ #### Interceptar chamadas de API interna
220
+
221
+ ```typescript
222
+ // ✅ Mais estável que scraping de HTML — dados estruturados direto da API interna
223
+ async function interceptApiData(url: string, apiPattern: RegExp): Promise<unknown[]> {
224
+ return withBrowser(async (browser) => {
225
+ const page = await browser.newPage()
226
+ const captured: unknown[] = []
227
+
228
+ await page.setRequestInterception(true)
229
+
230
+ page.on('request', (req) => req.continue())
231
+
232
+ page.on('response', async (response) => {
233
+ if (apiPattern.test(response.url()) && response.headers()['content-type']?.includes('json')) {
234
+ try {
235
+ captured.push(await response.json())
236
+ } catch {
237
+ // ignorar responses não-JSON
238
+ }
239
+ }
240
+ })
241
+
242
+ await page.goto(url, { waitUntil: 'networkidle2', timeout: 30_000 })
243
+
244
+ return captured
245
+ })
246
+ }
247
+
248
+ // Uso: interceptar chamadas de /api/products/*
249
+ const data = await interceptApiData('https://loja.com/produtos', /\/api\/products/)
250
+ ```
251
+
252
+ #### Paginação automática
253
+
254
+ ```typescript
255
+ async function scrapeAllPages(baseUrl: string): Promise<unknown[]> {
256
+ return withBrowser(async (browser) => {
257
+ const page = await browser.newPage()
258
+ const allItems: unknown[] = []
259
+ let pageNum = 1
260
+ let hasNextPage = true
261
+
262
+ while (hasNextPage) {
263
+ await page.goto(`${baseUrl}?page=${pageNum}`, { waitUntil: 'networkidle2' })
264
+
265
+ const items = await page.evaluate(() =>
266
+ Array.from(document.querySelectorAll('.item')).map((el) => ({
267
+ title: el.querySelector('h3')?.textContent?.trim() ?? null,
268
+ }))
269
+ )
270
+
271
+ allItems.push(...items)
272
+
273
+ hasNextPage = (await page.$('[aria-label="Próxima página"]:not([disabled])')) !== null
274
+ pageNum++
275
+
276
+ await randomDelay(1000, 3000)
277
+ }
278
+
279
+ return allItems
280
+ })
281
+ }
282
+ ```
283
+
284
+ ### Playwright — Alternativa (quando necessário)
285
+
286
+ Usar Playwright quando Puppeteer não resolver o caso de uso (ex: multi-browser testing, maior controle de contexto).
287
+
288
+ ```typescript
289
+ import { chromium } from 'playwright'
290
+
291
+ // Mesma lógica do Puppeteer — API similar com pequenas diferenças:
292
+ // page.goto(url, { waitUntil: 'networkidle' }) → sem o "2" no Playwright
293
+ // page.$() → page.locator() é o padrão moderno no Playwright
294
+ // page.evaluate() → idêntico
295
+ ```
296
+
297
+ ### Cheerio — Scraping de HTML Estático
298
+
299
+ ```typescript
300
+ import * as cheerio from 'cheerio'
301
+
302
+ async function scrapeWithCheerio(url: string) {
303
+ const response = await fetch(url, {
304
+ headers: {
305
+ 'User-Agent': 'Mozilla/5.0 (compatible; MyBot/1.0; +https://meusite.com/bot)',
306
+ 'Accept-Language': 'pt-BR,pt;q=0.9',
307
+ },
308
+ })
309
+
310
+ if (!response.ok) {
311
+ throw new Error(`HTTP ${response.status}: ${url}`)
312
+ }
313
+
314
+ const html = await response.text()
315
+ const $ = cheerio.load(html)
316
+
317
+ // ✅ Extração com Cheerio
318
+ const articles = $('article.post').map((_, el) => ({
319
+ title: $(el).find('h2').text().trim() || null,
320
+ date: $(el).find('time').attr('datetime') || null,
321
+ url: $(el).find('a.read-more').attr('href') || null,
322
+ summary: $(el).find('.excerpt').text().trim() || null,
323
+ })).get()
324
+
325
+ return articles.filter((a) => a.title && a.url)
326
+ }
327
+ ```
328
+
329
+ ### Parsing de PDFs
330
+
331
+ ```typescript
332
+ import pdfParse from 'pdf-parse'
333
+ import { readFileSync } from 'fs'
334
+
335
+ async function extractTextFromPdf(filePath: string): Promise<string> {
336
+ const buffer = readFileSync(filePath)
337
+ const { text, numpages } = await pdfParse(buffer)
338
+ console.log(`PDF processado: ${numpages} páginas`)
339
+ return text
340
+ }
341
+
342
+ // ✅ Extrair tabela de texto de PDF (regex + heurística)
343
+ function extractTableFromText(text: string, headers: string[]): Record<string, string>[] {
344
+ const lines = text.split('\n').map((l) => l.trim()).filter(Boolean)
345
+ const results: Record<string, string>[] = []
346
+
347
+ for (const line of lines) {
348
+ // lógica específica para o formato do PDF alvo
349
+ const match = line.match(/(\w+)\s+([\d,.]+)\s+([\d,.]+)/)
350
+ if (match) {
351
+ results.push({
352
+ [headers[0]]: match[1],
353
+ [headers[1]]: match[2],
354
+ [headers[2]]: match[3],
355
+ })
356
+ }
357
+ }
358
+
359
+ return results
360
+ }
361
+ ```
362
+
363
+ ### Técnicas Anti-Bot
364
+
365
+ #### Headers realistas
366
+
367
+ ```typescript
368
+ // ✅ Headers que imitam browser real
369
+ const realisticHeaders = {
370
+ 'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36',
371
+ 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
372
+ 'Accept-Language': 'pt-BR,pt;q=0.9,en-US;q=0.8,en;q=0.7',
373
+ 'Accept-Encoding': 'gzip, deflate, br',
374
+ 'Connection': 'keep-alive',
375
+ 'Upgrade-Insecure-Requests': '1',
376
+ }
377
+ ```
378
+
379
+ #### Rate limiting e delays aleatórios
380
+
381
+ ```typescript
382
+ // ✅ Delay com jitter para comportamento orgânico
383
+ function randomDelay(minMs: number, maxMs: number): Promise<void> {
384
+ const ms = Math.floor(Math.random() * (maxMs - minMs + 1)) + minMs
385
+ return new Promise((resolve) => setTimeout(resolve, ms))
386
+ }
387
+
388
+ // ✅ Controle de concorrência (evitar sobrecarga no alvo)
389
+ import PQueue from 'p-queue'
390
+
391
+ const queue = new PQueue({
392
+ concurrency: 2, // máximo 2 requests simultâneos
393
+ intervalCap: 5, // máximo 5 requests
394
+ interval: 10_000, // por 10 segundos
395
+ })
396
+
397
+ async function scrapeUrls(urls: string[]) {
398
+ return Promise.all(
399
+ urls.map((url) => queue.add(async () => {
400
+ await randomDelay(500, 1500)
401
+ return scrapeUrl(url)
402
+ }))
403
+ )
404
+ }
405
+ ```
406
+
407
+ #### Puppeteer — reduzir fingerprint de automação
408
+
409
+ ```typescript
410
+ import puppeteer from 'puppeteer'
411
+ import { executablePath } from 'puppeteer'
412
+
413
+ // ✅ Ocultar sinais de automação via page.evaluateOnNewDocument
414
+ async function createStealthPage(browser: Browser): Promise<Page> {
415
+ const page = await browser.newPage()
416
+
417
+ await page.setUserAgent(
418
+ 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
419
+ )
420
+
421
+ await page.setViewport({ width: 1366, height: 768 })
422
+
423
+ // Remover propriedades que identificam automação
424
+ await page.evaluateOnNewDocument(() => {
425
+ Object.defineProperty(navigator, 'webdriver', { get: () => undefined })
426
+ Object.defineProperty(navigator, 'plugins', { get: () => [1, 2, 3] })
427
+ Object.defineProperty(navigator, 'languages', { get: () => ['pt-BR', 'pt', 'en'] })
428
+ })
429
+
430
+ return page
431
+ }
432
+ ```
433
+
434
+ #### Rotação de proxies
435
+
436
+ ```typescript
437
+ // ✅ Pool de proxies com rotação
438
+ const proxies = [
439
+ 'http://user:pass@proxy1:8080',
440
+ 'http://user:pass@proxy2:8080',
441
+ 'http://user:pass@proxy3:8080',
442
+ ]
443
+
444
+ function getRandomProxy(): string {
445
+ return proxies[Math.floor(Math.random() * proxies.length)]
446
+ }
447
+
448
+ const browser = await puppeteer.launch({
449
+ headless: true,
450
+ args: [`--proxy-server=${getRandomProxy()}`, '--no-sandbox'],
451
+ })
452
+ ```
453
+
454
+ ### Resiliência e Retry
455
+
456
+ ```typescript
457
+ // ✅ Retry com backoff exponencial
458
+ async function withRetry<T>(
459
+ fn: () => Promise<T>,
460
+ options: { attempts?: number; baseDelay?: number; label?: string } = {}
461
+ ): Promise<T> {
462
+ const { attempts = 3, baseDelay = 1000, label = 'operação' } = options
463
+
464
+ for (let attempt = 1; attempt <= attempts; attempt++) {
465
+ try {
466
+ return await fn()
467
+ } catch (error) {
468
+ if (attempt === attempts) throw error
469
+
470
+ const delay = baseDelay * Math.pow(2, attempt - 1)
471
+ console.warn(`${label} falhou (tentativa ${attempt}/${attempts}). Retry em ${delay}ms...`)
472
+ await new Promise((resolve) => setTimeout(resolve, delay))
473
+ }
474
+ }
475
+
476
+ throw new Error('Unreachable')
477
+ }
478
+
479
+ // Uso
480
+ const data = await withRetry(
481
+ () => scrapePage(url),
482
+ { attempts: 3, baseDelay: 2000, label: `scrape ${url}` }
483
+ )
484
+ ```
485
+
486
+ #### Detectar mudanças de estrutura
487
+
488
+ ```typescript
489
+ // ✅ Validar que os dados extraídos fazem sentido
490
+ function validateExtractedData(data: unknown[], schema: { field: string; required: boolean }[]) {
491
+ const requiredFields = schema.filter((s) => s.required).map((s) => s.field)
492
+
493
+ const issues = data.filter((item) =>
494
+ requiredFields.some((field) => !item[field as keyof typeof item])
495
+ )
496
+
497
+ if (issues.length > 0) {
498
+ const pct = Math.round((issues.length / data.length) * 100)
499
+ console.warn(`ALERTA: ${issues.length}/${data.length} (${pct}%) itens com campos obrigatórios ausentes`)
500
+ console.warn('Possível mudança de estrutura no site alvo. Verificar seletores.')
501
+ }
502
+
503
+ return data.filter((item) => requiredFields.every((f) => item[f as keyof typeof item]))
504
+ }
505
+ ```
506
+
507
+ ### Pipelines ETL Leve
508
+
509
+ #### Estrutura básica de pipeline
510
+
511
+ ```typescript
512
+ interface PipelineConfig {
513
+ sources: string[] // URLs a processar
514
+ transform: (raw: unknown) => unknown // transformação de dados
515
+ output: string // caminho do arquivo de saída
516
+ concurrency?: number // requests simultâneos
517
+ }
518
+
519
+ async function runPipeline(config: PipelineConfig) {
520
+ console.log(`Iniciando pipeline: ${config.sources.length} fontes`)
521
+
522
+ // Extract
523
+ const rawData = await scrapeUrls(config.sources)
524
+
525
+ // Transform
526
+ const transformed = rawData
527
+ .flat()
528
+ .map(config.transform)
529
+ .filter(Boolean)
530
+
531
+ // Deduplicação por URL ou ID
532
+ const deduped = [...new Map(transformed.map((item) => [item.id ?? item.url, item])).values()]
533
+
534
+ // Load
535
+ await saveToFile(config.output, deduped)
536
+ console.log(`Pipeline concluído: ${deduped.length} registros salvos em ${config.output}`)
537
+
538
+ return deduped
539
+ }
540
+ ```
541
+
542
+ #### Formatos de saída
543
+
544
+ ```typescript
545
+ import { writeFileSync } from 'fs'
546
+ import { stringify } from 'csv-stringify/sync'
547
+
548
+ function saveToFile(path: string, data: unknown[]) {
549
+ if (path.endsWith('.json')) {
550
+ writeFileSync(path, JSON.stringify(data, null, 2), 'utf-8')
551
+ } else if (path.endsWith('.csv')) {
552
+ const csv = stringify(data as object[], { header: true })
553
+ writeFileSync(path, csv, 'utf-8')
554
+ } else if (path.endsWith('.ndjson')) {
555
+ const ndjson = data.map((d) => JSON.stringify(d)).join('\n')
556
+ writeFileSync(path, ndjson, 'utf-8')
557
+ }
558
+ }
559
+ ```
560
+
561
+ ### Monitoramento de Mudanças
562
+
563
+ ```typescript
564
+ // ✅ Detectar mudanças em páginas monitoradas
565
+ async function checkForChanges(url: string, storePath: string) {
566
+ const current = await scrapeWithCheerio(url)
567
+ const currentHash = crypto.createHash('md5').update(JSON.stringify(current)).digest('hex')
568
+
569
+ let previous: { hash: string; data: unknown; timestamp: string } | null = null
570
+ try {
571
+ previous = JSON.parse(readFileSync(storePath, 'utf-8'))
572
+ } catch {
573
+ // primeira execução
574
+ }
575
+
576
+ if (previous?.hash === currentHash) {
577
+ console.log('Sem mudanças detectadas.')
578
+ return { changed: false }
579
+ }
580
+
581
+ // Persistir estado atual
582
+ writeFileSync(storePath, JSON.stringify({ hash: currentHash, data: current, timestamp: new Date().toISOString() }))
583
+
584
+ return { changed: true, current, previous: previous?.data }
585
+ }
586
+ ```
587
+
588
+ ---
589
+
590
+ ## Regras
591
+
592
+ ### Nunca
593
+ - Ignorar `robots.txt` sem avaliar o contexto de uso
594
+ - Scraping agressivo (sem rate limiting) que possa causar impacto no servidor alvo
595
+ - Coletar dados pessoais sem finalidade legítima e base legal (LGPD/GDPR)
596
+ - Assumir que a estrutura do HTML é estável — sempre validar
597
+ - Fazer login em contas de terceiros sem autorização explícita
598
+ - Expor credenciais de proxy ou contas em código ou logs
599
+
600
+ ### Sempre
601
+ - Verificar `robots.txt` e Terms of Service antes de implementar
602
+ - Preferir API oficial quando disponível
603
+ - Rate limiting e delays aleatórios para não sobrecarregar o alvo
604
+ - Validar dados extraídos para detectar mudanças de estrutura
605
+ - Logging adequado para monitorar saúde do scraper
606
+ - Tratar erros e implementar retry para falhas transitórias
607
+ - Caching para não re-baixar dados já coletados
608
+
609
+ ---
610
+
611
+ ## Tratamento de Erros
612
+
613
+ ### Site bloqueando requests (403/429)
614
+ - Aumentar delay entre requests
615
+ - Verificar e ajustar User-Agent
616
+ - Considerar rotação de proxies
617
+ - Se persistir, comunicar ao usuário — pode ser proteção legítima
618
+
619
+ ### Estrutura HTML mudou (dados extraídos vazios ou incorretos)
620
+ - Logar amostra do HTML recebido para inspecionar
621
+ - Identificar novos seletores CSS ou XPath
622
+ - Adicionar validação para detectar mudanças futuras automaticamente
623
+
624
+ ### Timeout de navegação
625
+ - Aumentar timeout da operação específica
626
+ - Verificar se o site tem renderização lenta ou depende de recursos externos
627
+ - Tentar com `waitUntil: 'domcontentloaded'` em vez de `'networkidle2'`
628
+
629
+ ### PDF corrompido ou não parseável
630
+ - Verificar se o arquivo está completo (não truncado)
631
+ - Tentar biblioteca alternativa (pdfjs-dist vs pdf-parse)
632
+ - Extrair como imagem e usar OCR se o PDF for escaneado
633
+
634
+ ---
635
+
636
+ ## Checklist de Conclusão
637
+
638
+ - [ ] robots.txt e ToS verificados
639
+ - [ ] API oficial descartada como alternativa
640
+ - [ ] Ferramenta escolhida adequada (Puppeteer / Playwright / Cheerio / parser)
641
+ - [ ] Rate limiting implementado
642
+ - [ ] Retry com backoff exponencial
643
+ - [ ] Validação dos dados extraídos
644
+ - [ ] Detecção de mudança de estrutura
645
+ - [ ] Formato de saída definido (JSON / CSV / NDJSON)
646
+ - [ ] Testes com amostra real dos dados
647
+
648
+ ---
649
+
650
+ ## Output
651
+
652
+ | Artefato | Descrição |
653
+ |----------|-----------|
654
+ | Scraper | Script de extração com retry, rate limiting e validação |
655
+ | Parser | Lógica de parsing adaptada à estrutura da fonte |
656
+ | Pipeline | ETL completo: extração → transformação → arquivo de saída |
657
+ | Monitor | Script de detecção de mudanças com diff estruturado |
658
+
659
+ ---
660
+
661
+ ## Mensagem de Conclusão
662
+
663
+ ```
664
+ Scraper implementado!
665
+
666
+ Fonte: {URL ou tipo de fonte}
667
+ Ferramenta: {Puppeteer / Playwright / Cheerio / pdf-parse / xml-parser}
668
+ Dados extraídos: {campos coletados}
669
+ Rate limiting: {delay entre requests}
670
+ Retry: {N tentativas com backoff exponencial}
671
+
672
+ Saída: {JSON / CSV / NDJSON em path/to/output}
673
+ Validação: {campos obrigatórios verificados}
674
+
675
+ Próximo passo: {executar com amostra real / agendar com cron / integrar no pipeline}
676
+ ```
677
+
678
+ ---
679
+
680
+ ## Recursos Adicionais
681
+
682
+ - **Referências**: Veja [references/](references/) para links de documentação local
683
+ - **Puppeteer docs**: https://pptr.dev (documentação oficial)