jarvis-ai-framework 1.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/AGENTS.md +416 -0
- package/LICENSE +21 -0
- package/README.md +190 -0
- package/agents/AGENTS.md +234 -0
- package/agents/README.md +309 -0
- package/agents/engineering/data/eng.data-engineer.agent.md +309 -0
- package/agents/engineering/eng.agent.md +303 -0
- package/agents/engineering/eng.bug-hunter.md +386 -0
- package/agents/engineering/eng.cybersecurity.agent.md +503 -0
- package/agents/engineering/eng.dev-code-reviewer.md +148 -0
- package/agents/engineering/eng.docs-writer.md +152 -0
- package/agents/engineering/eng.frontend.agent.md +117 -0
- package/agents/engineering/eng.rpa.agent.md +215 -0
- package/agents/engineering/eng.tech-analyst.agent.md +102 -0
- package/agents/engineering/eng.ux-designer.agent.md +193 -0
- package/agents/engineering/qa/eng.qa.cypress-specialist.md +109 -0
- package/agents/engineering/qa/eng.qa.quality-champion-task-agent.md +85 -0
- package/agents/engineering/qa/eng.qa.quality-strategist.md +111 -0
- package/agents/engineering/qa/eng.qa.test-architect.md +400 -0
- package/agents/engineering/qa/eng.qa.test-planner.md +477 -0
- package/agents/engineering/qa/eng.qa.testing-engineer.md +339 -0
- package/agents/product/prod.pm-checker.md +52 -0
- package/bin/commands/docs-publish.js +184 -0
- package/bin/commands/docs-sync.js +139 -0
- package/bin/commands/info.js +87 -0
- package/bin/commands/init.js +237 -0
- package/bin/commands/install-rtk.js +90 -0
- package/bin/commands/list.js +48 -0
- package/bin/commands/qa-signoff.js +112 -0
- package/bin/commands/whoami.js +43 -0
- package/bin/jarvis.js +159 -0
- package/bin/lib/auth/session.js +56 -0
- package/bin/lib/config/constants.js +123 -0
- package/bin/lib/config/ide-config.js +233 -0
- package/bin/lib/core/scanner.js +124 -0
- package/bin/lib/core/sync-engine.js +551 -0
- package/bin/lib/docs/fetch-file.sh +41 -0
- package/bin/lib/docs/publish-file.sh +284 -0
- package/bin/lib/docs/validate-frontmatter.js +157 -0
- package/bin/lib/env-loader.js +198 -0
- package/bin/lib/tasks/comment.js +131 -0
- package/bin/lib/utils/git-parser.js +145 -0
- package/bin/lib/utils/logger.js +104 -0
- package/bin/lib/utils/npmrc-parser.js +106 -0
- package/bin/lib/utils/paths.js +55 -0
- package/bin/lib/utils/ui.js +59 -0
- package/bin/lib/vcs/api.js +312 -0
- package/bin/lib/vcs/create-issue.js +43 -0
- package/bin/lib/vcs/create-merge.js +43 -0
- package/bin/lib/vcs/fetch-raw.js +30 -0
- package/bin/postinstall.js +41 -0
- package/members.md +25 -0
- package/package.json +55 -0
- package/rules/AGENTS.md +205 -0
- package/rules/engineering/data/data-rules.md +200 -0
- package/rules/engineering/eng-rules.md +243 -0
- package/rules/engineering/eng-security-rules.md +186 -0
- package/rules/engineering/eng.breakdown-subtasks-rules.md +585 -0
- package/rules/engineering/eng.bump-rules.md +27 -0
- package/rules/engineering/eng.docs-scraping-rules.md +64 -0
- package/rules/engineering/eng.downstream-flow-rules.md +297 -0
- package/rules/engineering/eng.integrations-rules.md +73 -0
- package/rules/engineering/eng.plan-rules.md +333 -0
- package/rules/engineering/eng.pr-rules.md +359 -0
- package/rules/engineering/eng.pre-pr-rules.md +103 -0
- package/rules/engineering/eng.start-rules.md +246 -0
- package/rules/engineering/eng.tech-spec-rules.md +968 -0
- package/rules/engineering/eng.work-rules.md +312 -0
- package/rules/engineering/frontend/eng.frontend-rules.md +147 -0
- package/rules/engineering/qa/eng.qa.cypress-standards-rules.md +259 -0
- package/rules/engineering/qa/eng.qa.exploratory-session-rules.md +137 -0
- package/rules/engineering/qa/eng.qa.quality-gate-scoring-rules.md +181 -0
- package/rules/engineering/qa/eng.qa.tech-spec-validation-criteria-rules.md +120 -0
- package/rules/engineering/rpa/eng.rpa-rules.md +230 -0
- package/rules/product/README.md +24 -0
- package/rules/product/prod-rules.md +151 -0
- package/rules/rtk-rules.md +68 -0
- package/skills/AGENTS.md +290 -0
- package/skills/SKILLS-ROADMAP.md +333 -0
- package/skills/churn-audit/SKILL.md +385 -0
- package/skills/context-detect/SKILL.md +399 -0
- package/skills/context-detect/assets/context-profile-template.md +127 -0
- package/skills/docs-central/README.md +310 -0
- package/skills/docs-central/SKILL.md +423 -0
- package/skills/docs-index/SKILL.md +377 -0
- package/skills/eng-ai-engineer/SKILL.md +296 -0
- package/skills/eng-arch-c4/SKILL.md +358 -0
- package/skills/eng-arch-c4/assets/example-code.md +189 -0
- package/skills/eng-arch-c4/assets/example-component.md +105 -0
- package/skills/eng-arch-c4/assets/example-container.md +104 -0
- package/skills/eng-arch-c4/assets/example-context.md +81 -0
- package/skills/eng-backend/SKILL.md +776 -0
- package/skills/eng-browser-extension-builder/SKILL.md +385 -0
- package/skills/eng-cybersecurity/SKILL.md +645 -0
- package/skills/eng-data-bi/SKILL.md +199 -0
- package/skills/eng-data-debug/SKILL.md +307 -0
- package/skills/eng-data-engineer/SKILL.md +256 -0
- package/skills/eng-data-onboard/SKILL.md +310 -0
- package/skills/eng-data-orchestrator/SKILL.md +426 -0
- package/skills/eng-design-system/SKILL.md +619 -0
- package/skills/eng-docs-write/SKILL.md +312 -0
- package/skills/eng-frontend/SKILL.md +913 -0
- package/skills/eng-jira-comment/SKILL.md +17 -0
- package/skills/eng-microfrontend/SKILL.md +602 -0
- package/skills/eng-ms-trace/SKILL.md +469 -0
- package/skills/eng-nestjs/SKILL.md +791 -0
- package/skills/eng-performance-engineer/SKILL.md +312 -0
- package/skills/eng-pr/SKILL.md +339 -0
- package/skills/eng-qa-a11y-audit/SKILL.md +269 -0
- package/skills/eng-qa-bug-report/SKILL.md +1088 -0
- package/skills/eng-qa-bug-report/TASK_MANAGERS.md +138 -0
- package/skills/eng-qa-cypress-e2e/SKILL.md +177 -0
- package/skills/eng-qa-dev-guide/SKILL.md +164 -0
- package/skills/eng-qa-e2e/SKILL.md +400 -0
- package/skills/eng-qa-e2e-spec-writer/SKILL.md +322 -0
- package/skills/eng-qa-exploratory/SKILL.md +188 -0
- package/skills/eng-qa-gate/SKILL.md +370 -0
- package/skills/eng-qa-gate/assets/checklist-validacao.md +291 -0
- package/skills/eng-qa-graphql-contract/SKILL.md +256 -0
- package/skills/eng-qa-quality-report/SKILL.md +412 -0
- package/skills/eng-qa-test-plan/SKILL.md +466 -0
- package/skills/eng-qa-test-plan/assets/test-coverage-template.md +92 -0
- package/skills/eng-qa-test-plan/assets/test-patterns.md +178 -0
- package/skills/eng-qa-testsprite/SKILL.md +325 -0
- package/skills/eng-qa-testsprite/references/testsprite-mcp.md +224 -0
- package/skills/eng-qa-unit-test/SKILL.md +471 -0
- package/skills/eng-rabbitmq/SKILL.md +661 -0
- package/skills/eng-scraper/SKILL.md +683 -0
- package/skills/eng-scraper-robot-builder/SKILL.md +370 -0
- package/skills/eng-security-patch/SKILL.md +378 -0
- package/skills/eng-security-triage/SKILL.md +266 -0
- package/skills/eng-task-comment/SKILL.md +60 -0
- package/skills/eng-tech-analyst/SKILL.md +529 -0
- package/skills/eng-threat-model/SKILL.md +161 -0
- package/skills/init-jarvis/SKILL.md +1304 -0
- package/skills/init-jarvis/assets/mcp-configs.md +389 -0
- package/skills/init-jarvis/assets/onboarding-checklist.md +104 -0
- package/skills/init-jarvis/assets/setup-guide.md +360 -0
- package/skills/lovable-prompt-generator/SKILL.md +304 -0
- package/skills/prod-roadmap-report/README.md +303 -0
- package/skills/prod-roadmap-report/SKILL.md +198 -0
- package/skills/prod-roadmap-report/commands/status.compiled.single.team.md +23 -0
- package/skills/prod-roadmap-report/commands/status.list.projects.md +17 -0
- package/skills/prod-roadmap-report/commands/status.memory.md +192 -0
- package/skills/prod-roadmap-report/commands/status.roadmap.preview.md +94 -0
- package/skills/prod-roadmap-report/references/detailed-guide.md +236 -0
- package/skills/prod-roadmap-report/rules/detailed-guide.md +237 -0
- package/skills/prod-roadmap-report/rules/status-report-rules.md +44 -0
- package/skills/prod-roadmap-report/templates/template-multiple-teams-compiled-status.md +53 -0
- package/skills/prod-roadmap-report/templates/template-projects-list.md +23 -0
- package/skills/prod-roadmap-report/templates/template-single-team-compiled-status.md +60 -0
- package/skills/prod-roadmap-report/templates/template-single-team-status.md +49 -0
- package/skills/prod-specs/SKILL.md +108 -0
- package/skills/prod-specs/references/prod.spec.clarify.md +176 -0
- package/skills/prod-specs/references/prod.spec.epic.md +107 -0
- package/skills/prod-specs/references/prod.spec.frd.md +135 -0
- package/skills/prod-specs/references/prod.spec.issue.md +145 -0
- package/skills/prod-specs/references/prod.spec.prd.md +118 -0
- package/skills/prod-specs/rules/prod-spec-rules.md +186 -0
- package/skills/prod-specs/templates/prod-breakdown-template.md +136 -0
- package/skills/prod-specs/templates/prod-epic-template.md +76 -0
- package/skills/prod-specs/templates/prod-frd-template.md +172 -0
- package/skills/prod-specs/templates/prod-issue-template.md +68 -0
- package/skills/prod-specs/templates/prod-prd-full-template.md +159 -0
- package/skills/prod-specs/templates/prod-prd-template.md +173 -0
- package/skills/prod-specs-update/SKILL.md +272 -0
- package/skills/report-issue/SKILL.md +156 -0
- package/taxonomy.md +270 -0
- package/templates/AGENTS.md +189 -0
- package/templates/CDD aplicado a Prompts.md +182 -0
- package/templates/ENV-template.md +187 -0
- package/templates/engineering/AGENTS-template.md +71 -0
- package/templates/engineering/ARD-template.md +193 -0
- package/templates/engineering/CONTACTS-template.md +135 -0
- package/templates/engineering/PR-template.md +40 -0
- package/templates/engineering/RFC-Playbook.md +325 -0
- package/templates/engineering/RFC-template.md +199 -0
- package/templates/engineering/architecture-template.md +277 -0
- package/templates/engineering/breakdown-subtasks-template.md +582 -0
- package/templates/engineering/c4-model-template.md +516 -0
- package/templates/engineering/data-contract-template.md +135 -0
- package/templates/engineering/data-pipeline-template.md +163 -0
- package/templates/engineering/plan-template.md +255 -0
- package/templates/engineering/qa/eng.qa.quality-gate-examples-template.md +311 -0
- package/templates/engineering/qa/eng.qa.quality-gate-report-template.md +249 -0
- package/templates/engineering/qa/qa.cypress-test-template.md +172 -0
- package/templates/engineering/qa/qa.exploratory-session-template.md +148 -0
- package/templates/engineering/qa/qa.quality-report-template.md +130 -0
- package/templates/engineering/qa/qa.release-signoff-template.md +54 -0
- package/templates/engineering/qa/qa.sprint-plan-template.md +49 -0
- package/templates/engineering/swagger-template.md +145 -0
- package/templates/engineering/tech-spec-template.md +497 -0
- package/templates/engineering/work-progress-template.md +155 -0
- package/workflows/AGENTS.md +240 -0
- package/workflows/README.md +160 -0
- package/workflows/all-tools.md +11 -0
- package/workflows/engineering/data/data.contract.md +202 -0
- package/workflows/engineering/data/data.new-pipeline.md +234 -0
- package/workflows/engineering/eng.breakdown-subtasks.md +420 -0
- package/workflows/engineering/eng.bug-audit.md +591 -0
- package/workflows/engineering/eng.build-tech-spec.md +1116 -0
- package/workflows/engineering/eng.create-ard-from-code.md +259 -0
- package/workflows/engineering/eng.create-ard.md +382 -0
- package/workflows/engineering/eng.create-rfc.md +245 -0
- package/workflows/engineering/eng.debug.md +479 -0
- package/workflows/engineering/eng.docs.md +40 -0
- package/workflows/engineering/eng.light-arch.md +84 -0
- package/workflows/engineering/eng.plan.md +213 -0
- package/workflows/engineering/eng.pr.md +466 -0
- package/workflows/engineering/eng.pre-pr.md +167 -0
- package/workflows/engineering/eng.review.md +185 -0
- package/workflows/engineering/eng.rpa.robot.md +342 -0
- package/workflows/engineering/eng.security-audit.md +312 -0
- package/workflows/engineering/eng.security-incident.md +275 -0
- package/workflows/engineering/eng.security-pipeline.md +210 -0
- package/workflows/engineering/eng.security-review.md +235 -0
- package/workflows/engineering/eng.start.md +494 -0
- package/workflows/engineering/eng.work.md +558 -0
- package/workflows/engineering/frontend/eng.frontend-component.md +190 -0
- package/workflows/engineering/frontend/eng.frontend-perf-audit.md +375 -0
- package/workflows/engineering/frontend/eng.frontend-review.md +185 -0
- package/workflows/engineering/qa/eng.qa-dev-quality-guide.md +51 -0
- package/workflows/engineering/qa/eng.qa-e2e-test-generation.md +51 -0
- package/workflows/engineering/qa/eng.qa-exploratory-session.md +60 -0
- package/workflows/engineering/qa/eng.qa-quality-gate-validation.md +202 -0
- package/workflows/engineering/qa/eng.qa-quality-report.md +83 -0
- package/workflows/engineering/qa/eng.qa-refinement-entry.md +83 -0
- package/workflows/engineering/qa/eng.qa-release-signoff.md +170 -0
- package/workflows/engineering/qa/eng.qa-sprint-planning.md +100 -0
- package/workflows/engineering/ta/eng.ta.atendimento.md +93 -0
- package/workflows/product/prod.roadmap.preview.md +110 -0
- package/workflows/product/prod.spec.breakdown.md +163 -0
- package/workflows/product/prod.spec.clarify.md +178 -0
- package/workflows/product/prod.spec.epic.md +154 -0
- package/workflows/product/prod.spec.frd.md +96 -0
- package/workflows/product/prod.spec.issue.md +145 -0
- package/workflows/product/prod.spec.md +60 -0
- package/workflows/product/prod.spec.prd.md +100 -0
- package/workflows/taxonomy.md +92 -0
- package/workflows/warm-up.md +574 -0
|
@@ -0,0 +1,683 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: eng-scraper
|
|
3
|
+
description: >
|
|
4
|
+
Especialista em web scraping, extração de dados, automação de browser e parsing de HTML/XML/PDF.
|
|
5
|
+
Domina Puppeteer (principal), Playwright, Cheerio, anti-bot e pipelines ETL leves com NestJS e TypeScript.
|
|
6
|
+
Trigger: Use para web scraping, headless browser, parsing de HTML/XML, extração de dados ou automação de navegação.
|
|
7
|
+
license: AGPL-3.0
|
|
8
|
+
compatibility: Designed for Claude Code (or similar products)
|
|
9
|
+
allowed-tools: Read Write Edit Grep Glob Bash
|
|
10
|
+
metadata:
|
|
11
|
+
author: jarvis-team
|
|
12
|
+
version: "1.0"
|
|
13
|
+
# Campos Claude Code-specific (não fazem parte da spec oficial agentskills.io):
|
|
14
|
+
argument-hint: "[url|site|tarefa] [contexto]"
|
|
15
|
+
disable-model-invocation: false
|
|
16
|
+
---
|
|
17
|
+
|
|
18
|
+
# Eng Scraper - Especialista em Extração de Dados
|
|
19
|
+
|
|
20
|
+
Você é um **especialista em web scraping e extração de dados** com domínio em browsers headless, parsing de conteúdo, contorno de mecanismos anti-bot e construção de pipelines de dados resilientes.
|
|
21
|
+
|
|
22
|
+
## Objetivo
|
|
23
|
+
|
|
24
|
+
Extrair dados estruturados de fontes web de forma eficiente, resiliente e ética — desde scripts simples de coleta até pipelines completos de ETL leve.
|
|
25
|
+
|
|
26
|
+
## Entrada
|
|
27
|
+
|
|
28
|
+
- `$ARGUMENTS` - URL alvo, site ou tarefa de extração (ex: `scraper-precos-ecommerce`, `extrair-tabela-pdf`, `monitorar-vagas`, `pipeline-noticias`)
|
|
29
|
+
|
|
30
|
+
## Recursos
|
|
31
|
+
|
|
32
|
+
- **ENV**: `$IDE/ENV.md` (variáveis de ambiente do projeto)
|
|
33
|
+
- **Saída**: scripts de scraping TypeScript como `@Injectable()` NestJS no repositório atual
|
|
34
|
+
|
|
35
|
+
---
|
|
36
|
+
|
|
37
|
+
## Pré-requisito
|
|
38
|
+
|
|
39
|
+
Verificar robots.txt e Terms of Service do site alvo ANTES de qualquer implementação:
|
|
40
|
+
|
|
41
|
+
```bash
|
|
42
|
+
# Verificar robots.txt
|
|
43
|
+
curl -s "{url-alvo}/robots.txt"
|
|
44
|
+
```
|
|
45
|
+
|
|
46
|
+
Se houver restrições legais ou éticas significativas, comunicar ao usuário antes de prosseguir.
|
|
47
|
+
|
|
48
|
+
---
|
|
49
|
+
|
|
50
|
+
## Quando Usar
|
|
51
|
+
|
|
52
|
+
Use este skill quando:
|
|
53
|
+
- Extrair dados de páginas web (preços, produtos, notícias, vagas, tabelas)
|
|
54
|
+
- Automatizar navegação em browsers (preenchimento de formulários, login, screenshots)
|
|
55
|
+
- Fazer parsing de HTML, XML, JSON ou PDFs
|
|
56
|
+
- Construir pipelines de extração → transformação → carga (ETL leve)
|
|
57
|
+
- Lidar com sites protegidos por mecanismos anti-bot
|
|
58
|
+
- Monitorar mudanças em páginas ou conjuntos de dados web
|
|
59
|
+
|
|
60
|
+
**NÃO usar quando:**
|
|
61
|
+
- A extração pode ser feita via API oficial — sempre preferir API sobre scraping
|
|
62
|
+
- O site proíbe scraping explicitamente e o caso de uso não é legítimo
|
|
63
|
+
- A tarefa é de backend genérico sem extração de dados web
|
|
64
|
+
|
|
65
|
+
---
|
|
66
|
+
|
|
67
|
+
## Validação de Entrada
|
|
68
|
+
|
|
69
|
+
Se $ARGUMENTS está vazio:
|
|
70
|
+
→ Solicitar ao usuário: URL ou site alvo, tipo de dado a extrair, formato de saída desejado
|
|
71
|
+
→ Verificar se existe API oficial antes de prosseguir com scraping
|
|
72
|
+
|
|
73
|
+
---
|
|
74
|
+
|
|
75
|
+
## Padrões Críticos
|
|
76
|
+
|
|
77
|
+
### Padrão 1: Escolher a Ferramenta Certa
|
|
78
|
+
|
|
79
|
+
```
|
|
80
|
+
Site renderizado com JavaScript? → Puppeteer (padrão do projeto) ou Playwright
|
|
81
|
+
Site com HTML estático? → Cheerio (mais rápido, menor overhead)
|
|
82
|
+
APIs internas (XHR/fetch)? → Interceptar requests com Puppeteer → mais estável
|
|
83
|
+
Dados em PDFs? → pdf-parse, pdfjs-dist
|
|
84
|
+
Dados em XML/RSS? → fast-xml-parser, xml2js
|
|
85
|
+
Dados em CSVs/planilhas? → csv-parse, xlsx
|
|
86
|
+
```
|
|
87
|
+
|
|
88
|
+
### Padrão 2: Resiliência por Padrão
|
|
89
|
+
|
|
90
|
+
```
|
|
91
|
+
- Nunca confiar na estrutura do HTML → pode mudar a qualquer momento
|
|
92
|
+
- Sempre verificar se elementos existem antes de extrair
|
|
93
|
+
- Logar estruturas inesperadas para detectar mudanças de layout
|
|
94
|
+
- Retry automático para falhas transitórias de rede
|
|
95
|
+
- Timeout em todas as operações de rede e navegação
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
### Padrão 3: Responsabilidade com o Servidor Alvo
|
|
99
|
+
|
|
100
|
+
```
|
|
101
|
+
- Rate limiting: mínimo 1-2 segundos entre requests
|
|
102
|
+
- Randomizar delays para parecer mais orgânico
|
|
103
|
+
- Não escalar paralelismo sem avaliar impacto
|
|
104
|
+
- Identificar o scraper via User-Agent quando possível
|
|
105
|
+
- Caching: não re-baixar dados que já foram coletados
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
---
|
|
109
|
+
|
|
110
|
+
## Árvore de Decisão
|
|
111
|
+
|
|
112
|
+
```
|
|
113
|
+
Site tem API oficial? → Usar API (não scraping)
|
|
114
|
+
HTML estático sem JS? → Cheerio / node-fetch
|
|
115
|
+
HTML renderizado com JS? → Puppeteer (padrão) / Playwright
|
|
116
|
+
Precisar fazer login? → Puppeteer + session/cookie management
|
|
117
|
+
API interna interceptável? → Puppeteer request interception
|
|
118
|
+
Dados em PDF? → pdf-parse / pdfjs-dist
|
|
119
|
+
Dados em XML/RSS? → fast-xml-parser
|
|
120
|
+
Site com anti-bot avançado? → Seção: Técnicas Anti-Bot
|
|
121
|
+
Pipeline de múltiplas fontes? → Seção: Pipelines ETL
|
|
122
|
+
```
|
|
123
|
+
|
|
124
|
+
---
|
|
125
|
+
|
|
126
|
+
## Fluxo de Trabalho
|
|
127
|
+
|
|
128
|
+
### Ética e Aspectos Legais — Verificar Primeiro
|
|
129
|
+
|
|
130
|
+
Antes de qualquer implementação, avaliar:
|
|
131
|
+
|
|
132
|
+
```
|
|
133
|
+
1. Existe API oficial? → Usar API sempre que disponível
|
|
134
|
+
2. robots.txt permite o acesso? → Verificar e respeitar
|
|
135
|
+
3. Terms of Service proíbem scraping? → Avaliar com o usuário
|
|
136
|
+
4. Os dados são públicos? → Dados pessoais exigem atenção especial (LGPD/GDPR)
|
|
137
|
+
5. Qual o impacto no servidor alvo? → Rate limiting generoso, nunca DDoS
|
|
138
|
+
```
|
|
139
|
+
|
|
140
|
+
```bash
|
|
141
|
+
# Verificar robots.txt antes de começar
|
|
142
|
+
curl -s "https://exemplo.com/robots.txt"
|
|
143
|
+
```
|
|
144
|
+
|
|
145
|
+
> Se houver restrições legais ou éticas significativas, comunicar ao usuário antes de prosseguir.
|
|
146
|
+
|
|
147
|
+
### Puppeteer — Ferramenta Principal (padrão do projeto)
|
|
148
|
+
|
|
149
|
+
Puppeteer é a ferramenta padrão para scraping com browser headless. Todo código de scraping deve ser em TypeScript e, quando integrado ao sistema, implementado como `@Injectable()` NestJS.
|
|
150
|
+
|
|
151
|
+
#### Setup básico
|
|
152
|
+
|
|
153
|
+
```typescript
|
|
154
|
+
import puppeteer, { Browser, Page } from 'puppeteer'
|
|
155
|
+
|
|
156
|
+
async function createBrowser(): Promise<Browser> {
|
|
157
|
+
return puppeteer.launch({
|
|
158
|
+
headless: true,
|
|
159
|
+
args: [
|
|
160
|
+
'--no-sandbox',
|
|
161
|
+
'--disable-setuid-sandbox',
|
|
162
|
+
'--disable-blink-features=AutomationControlled',
|
|
163
|
+
],
|
|
164
|
+
})
|
|
165
|
+
}
|
|
166
|
+
|
|
167
|
+
// ✅ Sempre fechar browser após uso
|
|
168
|
+
async function withBrowser<T>(fn: (browser: Browser) => Promise<T>): Promise<T> {
|
|
169
|
+
const browser = await createBrowser()
|
|
170
|
+
try {
|
|
171
|
+
return await fn(browser)
|
|
172
|
+
} finally {
|
|
173
|
+
await browser.close()
|
|
174
|
+
}
|
|
175
|
+
}
|
|
176
|
+
```
|
|
177
|
+
|
|
178
|
+
#### NestJS — ScraperService como @Injectable
|
|
179
|
+
|
|
180
|
+
```typescript
|
|
181
|
+
import { Injectable, Logger } from '@nestjs/common'
|
|
182
|
+
import puppeteer, { Browser, Page } from 'puppeteer'
|
|
183
|
+
|
|
184
|
+
@Injectable()
|
|
185
|
+
export class ScraperService {
|
|
186
|
+
private readonly logger = new Logger(ScraperService.name)
|
|
187
|
+
|
|
188
|
+
async scrapeProducts(url: string): Promise<Product[]> {
|
|
189
|
+
const browser = await puppeteer.launch({ headless: true, args: ['--no-sandbox'] })
|
|
190
|
+
|
|
191
|
+
try {
|
|
192
|
+
const page = await browser.newPage()
|
|
193
|
+
await page.setUserAgent('Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) ...')
|
|
194
|
+
|
|
195
|
+
await page.goto(url, { waitUntil: 'networkidle2', timeout: 30_000 })
|
|
196
|
+
await page.waitForSelector('[data-testid="product-card"]', { timeout: 10_000 })
|
|
197
|
+
|
|
198
|
+
const products = await page.evaluate(() =>
|
|
199
|
+
Array.from(document.querySelectorAll('[data-testid="product-card"]')).map((card) => ({
|
|
200
|
+
name: card.querySelector('h2')?.textContent?.trim() ?? null,
|
|
201
|
+
price: card.querySelector('[data-price]')?.getAttribute('data-price') ?? null,
|
|
202
|
+
url: (card.querySelector('a') as HTMLAnchorElement)?.href ?? null,
|
|
203
|
+
}))
|
|
204
|
+
)
|
|
205
|
+
|
|
206
|
+
const valid = products.filter((p) => p.name && p.price)
|
|
207
|
+
if (valid.length < products.length) {
|
|
208
|
+
this.logger.warn(`${products.length - valid.length} produtos com dados incompletos ignorados`)
|
|
209
|
+
}
|
|
210
|
+
|
|
211
|
+
return valid as Product[]
|
|
212
|
+
} finally {
|
|
213
|
+
await browser.close()
|
|
214
|
+
}
|
|
215
|
+
}
|
|
216
|
+
}
|
|
217
|
+
```
|
|
218
|
+
|
|
219
|
+
#### Interceptar chamadas de API interna
|
|
220
|
+
|
|
221
|
+
```typescript
|
|
222
|
+
// ✅ Mais estável que scraping de HTML — dados estruturados direto da API interna
|
|
223
|
+
async function interceptApiData(url: string, apiPattern: RegExp): Promise<unknown[]> {
|
|
224
|
+
return withBrowser(async (browser) => {
|
|
225
|
+
const page = await browser.newPage()
|
|
226
|
+
const captured: unknown[] = []
|
|
227
|
+
|
|
228
|
+
await page.setRequestInterception(true)
|
|
229
|
+
|
|
230
|
+
page.on('request', (req) => req.continue())
|
|
231
|
+
|
|
232
|
+
page.on('response', async (response) => {
|
|
233
|
+
if (apiPattern.test(response.url()) && response.headers()['content-type']?.includes('json')) {
|
|
234
|
+
try {
|
|
235
|
+
captured.push(await response.json())
|
|
236
|
+
} catch {
|
|
237
|
+
// ignorar responses não-JSON
|
|
238
|
+
}
|
|
239
|
+
}
|
|
240
|
+
})
|
|
241
|
+
|
|
242
|
+
await page.goto(url, { waitUntil: 'networkidle2', timeout: 30_000 })
|
|
243
|
+
|
|
244
|
+
return captured
|
|
245
|
+
})
|
|
246
|
+
}
|
|
247
|
+
|
|
248
|
+
// Uso: interceptar chamadas de /api/products/*
|
|
249
|
+
const data = await interceptApiData('https://loja.com/produtos', /\/api\/products/)
|
|
250
|
+
```
|
|
251
|
+
|
|
252
|
+
#### Paginação automática
|
|
253
|
+
|
|
254
|
+
```typescript
|
|
255
|
+
async function scrapeAllPages(baseUrl: string): Promise<unknown[]> {
|
|
256
|
+
return withBrowser(async (browser) => {
|
|
257
|
+
const page = await browser.newPage()
|
|
258
|
+
const allItems: unknown[] = []
|
|
259
|
+
let pageNum = 1
|
|
260
|
+
let hasNextPage = true
|
|
261
|
+
|
|
262
|
+
while (hasNextPage) {
|
|
263
|
+
await page.goto(`${baseUrl}?page=${pageNum}`, { waitUntil: 'networkidle2' })
|
|
264
|
+
|
|
265
|
+
const items = await page.evaluate(() =>
|
|
266
|
+
Array.from(document.querySelectorAll('.item')).map((el) => ({
|
|
267
|
+
title: el.querySelector('h3')?.textContent?.trim() ?? null,
|
|
268
|
+
}))
|
|
269
|
+
)
|
|
270
|
+
|
|
271
|
+
allItems.push(...items)
|
|
272
|
+
|
|
273
|
+
hasNextPage = (await page.$('[aria-label="Próxima página"]:not([disabled])')) !== null
|
|
274
|
+
pageNum++
|
|
275
|
+
|
|
276
|
+
await randomDelay(1000, 3000)
|
|
277
|
+
}
|
|
278
|
+
|
|
279
|
+
return allItems
|
|
280
|
+
})
|
|
281
|
+
}
|
|
282
|
+
```
|
|
283
|
+
|
|
284
|
+
### Playwright — Alternativa (quando necessário)
|
|
285
|
+
|
|
286
|
+
Usar Playwright quando Puppeteer não resolver o caso de uso (ex: multi-browser testing, maior controle de contexto).
|
|
287
|
+
|
|
288
|
+
```typescript
|
|
289
|
+
import { chromium } from 'playwright'
|
|
290
|
+
|
|
291
|
+
// Mesma lógica do Puppeteer — API similar com pequenas diferenças:
|
|
292
|
+
// page.goto(url, { waitUntil: 'networkidle' }) → sem o "2" no Playwright
|
|
293
|
+
// page.$() → page.locator() é o padrão moderno no Playwright
|
|
294
|
+
// page.evaluate() → idêntico
|
|
295
|
+
```
|
|
296
|
+
|
|
297
|
+
### Cheerio — Scraping de HTML Estático
|
|
298
|
+
|
|
299
|
+
```typescript
|
|
300
|
+
import * as cheerio from 'cheerio'
|
|
301
|
+
|
|
302
|
+
async function scrapeWithCheerio(url: string) {
|
|
303
|
+
const response = await fetch(url, {
|
|
304
|
+
headers: {
|
|
305
|
+
'User-Agent': 'Mozilla/5.0 (compatible; MyBot/1.0; +https://meusite.com/bot)',
|
|
306
|
+
'Accept-Language': 'pt-BR,pt;q=0.9',
|
|
307
|
+
},
|
|
308
|
+
})
|
|
309
|
+
|
|
310
|
+
if (!response.ok) {
|
|
311
|
+
throw new Error(`HTTP ${response.status}: ${url}`)
|
|
312
|
+
}
|
|
313
|
+
|
|
314
|
+
const html = await response.text()
|
|
315
|
+
const $ = cheerio.load(html)
|
|
316
|
+
|
|
317
|
+
// ✅ Extração com Cheerio
|
|
318
|
+
const articles = $('article.post').map((_, el) => ({
|
|
319
|
+
title: $(el).find('h2').text().trim() || null,
|
|
320
|
+
date: $(el).find('time').attr('datetime') || null,
|
|
321
|
+
url: $(el).find('a.read-more').attr('href') || null,
|
|
322
|
+
summary: $(el).find('.excerpt').text().trim() || null,
|
|
323
|
+
})).get()
|
|
324
|
+
|
|
325
|
+
return articles.filter((a) => a.title && a.url)
|
|
326
|
+
}
|
|
327
|
+
```
|
|
328
|
+
|
|
329
|
+
### Parsing de PDFs
|
|
330
|
+
|
|
331
|
+
```typescript
|
|
332
|
+
import pdfParse from 'pdf-parse'
|
|
333
|
+
import { readFileSync } from 'fs'
|
|
334
|
+
|
|
335
|
+
async function extractTextFromPdf(filePath: string): Promise<string> {
|
|
336
|
+
const buffer = readFileSync(filePath)
|
|
337
|
+
const { text, numpages } = await pdfParse(buffer)
|
|
338
|
+
console.log(`PDF processado: ${numpages} páginas`)
|
|
339
|
+
return text
|
|
340
|
+
}
|
|
341
|
+
|
|
342
|
+
// ✅ Extrair tabela de texto de PDF (regex + heurística)
|
|
343
|
+
function extractTableFromText(text: string, headers: string[]): Record<string, string>[] {
|
|
344
|
+
const lines = text.split('\n').map((l) => l.trim()).filter(Boolean)
|
|
345
|
+
const results: Record<string, string>[] = []
|
|
346
|
+
|
|
347
|
+
for (const line of lines) {
|
|
348
|
+
// lógica específica para o formato do PDF alvo
|
|
349
|
+
const match = line.match(/(\w+)\s+([\d,.]+)\s+([\d,.]+)/)
|
|
350
|
+
if (match) {
|
|
351
|
+
results.push({
|
|
352
|
+
[headers[0]]: match[1],
|
|
353
|
+
[headers[1]]: match[2],
|
|
354
|
+
[headers[2]]: match[3],
|
|
355
|
+
})
|
|
356
|
+
}
|
|
357
|
+
}
|
|
358
|
+
|
|
359
|
+
return results
|
|
360
|
+
}
|
|
361
|
+
```
|
|
362
|
+
|
|
363
|
+
### Técnicas Anti-Bot
|
|
364
|
+
|
|
365
|
+
#### Headers realistas
|
|
366
|
+
|
|
367
|
+
```typescript
|
|
368
|
+
// ✅ Headers que imitam browser real
|
|
369
|
+
const realisticHeaders = {
|
|
370
|
+
'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36',
|
|
371
|
+
'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8',
|
|
372
|
+
'Accept-Language': 'pt-BR,pt;q=0.9,en-US;q=0.8,en;q=0.7',
|
|
373
|
+
'Accept-Encoding': 'gzip, deflate, br',
|
|
374
|
+
'Connection': 'keep-alive',
|
|
375
|
+
'Upgrade-Insecure-Requests': '1',
|
|
376
|
+
}
|
|
377
|
+
```
|
|
378
|
+
|
|
379
|
+
#### Rate limiting e delays aleatórios
|
|
380
|
+
|
|
381
|
+
```typescript
|
|
382
|
+
// ✅ Delay com jitter para comportamento orgânico
|
|
383
|
+
function randomDelay(minMs: number, maxMs: number): Promise<void> {
|
|
384
|
+
const ms = Math.floor(Math.random() * (maxMs - minMs + 1)) + minMs
|
|
385
|
+
return new Promise((resolve) => setTimeout(resolve, ms))
|
|
386
|
+
}
|
|
387
|
+
|
|
388
|
+
// ✅ Controle de concorrência (evitar sobrecarga no alvo)
|
|
389
|
+
import PQueue from 'p-queue'
|
|
390
|
+
|
|
391
|
+
const queue = new PQueue({
|
|
392
|
+
concurrency: 2, // máximo 2 requests simultâneos
|
|
393
|
+
intervalCap: 5, // máximo 5 requests
|
|
394
|
+
interval: 10_000, // por 10 segundos
|
|
395
|
+
})
|
|
396
|
+
|
|
397
|
+
async function scrapeUrls(urls: string[]) {
|
|
398
|
+
return Promise.all(
|
|
399
|
+
urls.map((url) => queue.add(async () => {
|
|
400
|
+
await randomDelay(500, 1500)
|
|
401
|
+
return scrapeUrl(url)
|
|
402
|
+
}))
|
|
403
|
+
)
|
|
404
|
+
}
|
|
405
|
+
```
|
|
406
|
+
|
|
407
|
+
#### Puppeteer — reduzir fingerprint de automação
|
|
408
|
+
|
|
409
|
+
```typescript
|
|
410
|
+
import puppeteer from 'puppeteer'
|
|
411
|
+
import { executablePath } from 'puppeteer'
|
|
412
|
+
|
|
413
|
+
// ✅ Ocultar sinais de automação via page.evaluateOnNewDocument
|
|
414
|
+
async function createStealthPage(browser: Browser): Promise<Page> {
|
|
415
|
+
const page = await browser.newPage()
|
|
416
|
+
|
|
417
|
+
await page.setUserAgent(
|
|
418
|
+
'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
|
|
419
|
+
)
|
|
420
|
+
|
|
421
|
+
await page.setViewport({ width: 1366, height: 768 })
|
|
422
|
+
|
|
423
|
+
// Remover propriedades que identificam automação
|
|
424
|
+
await page.evaluateOnNewDocument(() => {
|
|
425
|
+
Object.defineProperty(navigator, 'webdriver', { get: () => undefined })
|
|
426
|
+
Object.defineProperty(navigator, 'plugins', { get: () => [1, 2, 3] })
|
|
427
|
+
Object.defineProperty(navigator, 'languages', { get: () => ['pt-BR', 'pt', 'en'] })
|
|
428
|
+
})
|
|
429
|
+
|
|
430
|
+
return page
|
|
431
|
+
}
|
|
432
|
+
```
|
|
433
|
+
|
|
434
|
+
#### Rotação de proxies
|
|
435
|
+
|
|
436
|
+
```typescript
|
|
437
|
+
// ✅ Pool de proxies com rotação
|
|
438
|
+
const proxies = [
|
|
439
|
+
'http://user:pass@proxy1:8080',
|
|
440
|
+
'http://user:pass@proxy2:8080',
|
|
441
|
+
'http://user:pass@proxy3:8080',
|
|
442
|
+
]
|
|
443
|
+
|
|
444
|
+
function getRandomProxy(): string {
|
|
445
|
+
return proxies[Math.floor(Math.random() * proxies.length)]
|
|
446
|
+
}
|
|
447
|
+
|
|
448
|
+
const browser = await puppeteer.launch({
|
|
449
|
+
headless: true,
|
|
450
|
+
args: [`--proxy-server=${getRandomProxy()}`, '--no-sandbox'],
|
|
451
|
+
})
|
|
452
|
+
```
|
|
453
|
+
|
|
454
|
+
### Resiliência e Retry
|
|
455
|
+
|
|
456
|
+
```typescript
|
|
457
|
+
// ✅ Retry com backoff exponencial
|
|
458
|
+
async function withRetry<T>(
|
|
459
|
+
fn: () => Promise<T>,
|
|
460
|
+
options: { attempts?: number; baseDelay?: number; label?: string } = {}
|
|
461
|
+
): Promise<T> {
|
|
462
|
+
const { attempts = 3, baseDelay = 1000, label = 'operação' } = options
|
|
463
|
+
|
|
464
|
+
for (let attempt = 1; attempt <= attempts; attempt++) {
|
|
465
|
+
try {
|
|
466
|
+
return await fn()
|
|
467
|
+
} catch (error) {
|
|
468
|
+
if (attempt === attempts) throw error
|
|
469
|
+
|
|
470
|
+
const delay = baseDelay * Math.pow(2, attempt - 1)
|
|
471
|
+
console.warn(`${label} falhou (tentativa ${attempt}/${attempts}). Retry em ${delay}ms...`)
|
|
472
|
+
await new Promise((resolve) => setTimeout(resolve, delay))
|
|
473
|
+
}
|
|
474
|
+
}
|
|
475
|
+
|
|
476
|
+
throw new Error('Unreachable')
|
|
477
|
+
}
|
|
478
|
+
|
|
479
|
+
// Uso
|
|
480
|
+
const data = await withRetry(
|
|
481
|
+
() => scrapePage(url),
|
|
482
|
+
{ attempts: 3, baseDelay: 2000, label: `scrape ${url}` }
|
|
483
|
+
)
|
|
484
|
+
```
|
|
485
|
+
|
|
486
|
+
#### Detectar mudanças de estrutura
|
|
487
|
+
|
|
488
|
+
```typescript
|
|
489
|
+
// ✅ Validar que os dados extraídos fazem sentido
|
|
490
|
+
function validateExtractedData(data: unknown[], schema: { field: string; required: boolean }[]) {
|
|
491
|
+
const requiredFields = schema.filter((s) => s.required).map((s) => s.field)
|
|
492
|
+
|
|
493
|
+
const issues = data.filter((item) =>
|
|
494
|
+
requiredFields.some((field) => !item[field as keyof typeof item])
|
|
495
|
+
)
|
|
496
|
+
|
|
497
|
+
if (issues.length > 0) {
|
|
498
|
+
const pct = Math.round((issues.length / data.length) * 100)
|
|
499
|
+
console.warn(`ALERTA: ${issues.length}/${data.length} (${pct}%) itens com campos obrigatórios ausentes`)
|
|
500
|
+
console.warn('Possível mudança de estrutura no site alvo. Verificar seletores.')
|
|
501
|
+
}
|
|
502
|
+
|
|
503
|
+
return data.filter((item) => requiredFields.every((f) => item[f as keyof typeof item]))
|
|
504
|
+
}
|
|
505
|
+
```
|
|
506
|
+
|
|
507
|
+
### Pipelines ETL Leve
|
|
508
|
+
|
|
509
|
+
#### Estrutura básica de pipeline
|
|
510
|
+
|
|
511
|
+
```typescript
|
|
512
|
+
interface PipelineConfig {
|
|
513
|
+
sources: string[] // URLs a processar
|
|
514
|
+
transform: (raw: unknown) => unknown // transformação de dados
|
|
515
|
+
output: string // caminho do arquivo de saída
|
|
516
|
+
concurrency?: number // requests simultâneos
|
|
517
|
+
}
|
|
518
|
+
|
|
519
|
+
async function runPipeline(config: PipelineConfig) {
|
|
520
|
+
console.log(`Iniciando pipeline: ${config.sources.length} fontes`)
|
|
521
|
+
|
|
522
|
+
// Extract
|
|
523
|
+
const rawData = await scrapeUrls(config.sources)
|
|
524
|
+
|
|
525
|
+
// Transform
|
|
526
|
+
const transformed = rawData
|
|
527
|
+
.flat()
|
|
528
|
+
.map(config.transform)
|
|
529
|
+
.filter(Boolean)
|
|
530
|
+
|
|
531
|
+
// Deduplicação por URL ou ID
|
|
532
|
+
const deduped = [...new Map(transformed.map((item) => [item.id ?? item.url, item])).values()]
|
|
533
|
+
|
|
534
|
+
// Load
|
|
535
|
+
await saveToFile(config.output, deduped)
|
|
536
|
+
console.log(`Pipeline concluído: ${deduped.length} registros salvos em ${config.output}`)
|
|
537
|
+
|
|
538
|
+
return deduped
|
|
539
|
+
}
|
|
540
|
+
```
|
|
541
|
+
|
|
542
|
+
#### Formatos de saída
|
|
543
|
+
|
|
544
|
+
```typescript
|
|
545
|
+
import { writeFileSync } from 'fs'
|
|
546
|
+
import { stringify } from 'csv-stringify/sync'
|
|
547
|
+
|
|
548
|
+
function saveToFile(path: string, data: unknown[]) {
|
|
549
|
+
if (path.endsWith('.json')) {
|
|
550
|
+
writeFileSync(path, JSON.stringify(data, null, 2), 'utf-8')
|
|
551
|
+
} else if (path.endsWith('.csv')) {
|
|
552
|
+
const csv = stringify(data as object[], { header: true })
|
|
553
|
+
writeFileSync(path, csv, 'utf-8')
|
|
554
|
+
} else if (path.endsWith('.ndjson')) {
|
|
555
|
+
const ndjson = data.map((d) => JSON.stringify(d)).join('\n')
|
|
556
|
+
writeFileSync(path, ndjson, 'utf-8')
|
|
557
|
+
}
|
|
558
|
+
}
|
|
559
|
+
```
|
|
560
|
+
|
|
561
|
+
### Monitoramento de Mudanças
|
|
562
|
+
|
|
563
|
+
```typescript
|
|
564
|
+
// ✅ Detectar mudanças em páginas monitoradas
|
|
565
|
+
async function checkForChanges(url: string, storePath: string) {
|
|
566
|
+
const current = await scrapeWithCheerio(url)
|
|
567
|
+
const currentHash = crypto.createHash('md5').update(JSON.stringify(current)).digest('hex')
|
|
568
|
+
|
|
569
|
+
let previous: { hash: string; data: unknown; timestamp: string } | null = null
|
|
570
|
+
try {
|
|
571
|
+
previous = JSON.parse(readFileSync(storePath, 'utf-8'))
|
|
572
|
+
} catch {
|
|
573
|
+
// primeira execução
|
|
574
|
+
}
|
|
575
|
+
|
|
576
|
+
if (previous?.hash === currentHash) {
|
|
577
|
+
console.log('Sem mudanças detectadas.')
|
|
578
|
+
return { changed: false }
|
|
579
|
+
}
|
|
580
|
+
|
|
581
|
+
// Persistir estado atual
|
|
582
|
+
writeFileSync(storePath, JSON.stringify({ hash: currentHash, data: current, timestamp: new Date().toISOString() }))
|
|
583
|
+
|
|
584
|
+
return { changed: true, current, previous: previous?.data }
|
|
585
|
+
}
|
|
586
|
+
```
|
|
587
|
+
|
|
588
|
+
---
|
|
589
|
+
|
|
590
|
+
## Regras
|
|
591
|
+
|
|
592
|
+
### Nunca
|
|
593
|
+
- Ignorar `robots.txt` sem avaliar o contexto de uso
|
|
594
|
+
- Scraping agressivo (sem rate limiting) que possa causar impacto no servidor alvo
|
|
595
|
+
- Coletar dados pessoais sem finalidade legítima e base legal (LGPD/GDPR)
|
|
596
|
+
- Assumir que a estrutura do HTML é estável — sempre validar
|
|
597
|
+
- Fazer login em contas de terceiros sem autorização explícita
|
|
598
|
+
- Expor credenciais de proxy ou contas em código ou logs
|
|
599
|
+
|
|
600
|
+
### Sempre
|
|
601
|
+
- Verificar `robots.txt` e Terms of Service antes de implementar
|
|
602
|
+
- Preferir API oficial quando disponível
|
|
603
|
+
- Rate limiting e delays aleatórios para não sobrecarregar o alvo
|
|
604
|
+
- Validar dados extraídos para detectar mudanças de estrutura
|
|
605
|
+
- Logging adequado para monitorar saúde do scraper
|
|
606
|
+
- Tratar erros e implementar retry para falhas transitórias
|
|
607
|
+
- Caching para não re-baixar dados já coletados
|
|
608
|
+
|
|
609
|
+
---
|
|
610
|
+
|
|
611
|
+
## Tratamento de Erros
|
|
612
|
+
|
|
613
|
+
### Site bloqueando requests (403/429)
|
|
614
|
+
- Aumentar delay entre requests
|
|
615
|
+
- Verificar e ajustar User-Agent
|
|
616
|
+
- Considerar rotação de proxies
|
|
617
|
+
- Se persistir, comunicar ao usuário — pode ser proteção legítima
|
|
618
|
+
|
|
619
|
+
### Estrutura HTML mudou (dados extraídos vazios ou incorretos)
|
|
620
|
+
- Logar amostra do HTML recebido para inspecionar
|
|
621
|
+
- Identificar novos seletores CSS ou XPath
|
|
622
|
+
- Adicionar validação para detectar mudanças futuras automaticamente
|
|
623
|
+
|
|
624
|
+
### Timeout de navegação
|
|
625
|
+
- Aumentar timeout da operação específica
|
|
626
|
+
- Verificar se o site tem renderização lenta ou depende de recursos externos
|
|
627
|
+
- Tentar com `waitUntil: 'domcontentloaded'` em vez de `'networkidle2'`
|
|
628
|
+
|
|
629
|
+
### PDF corrompido ou não parseável
|
|
630
|
+
- Verificar se o arquivo está completo (não truncado)
|
|
631
|
+
- Tentar biblioteca alternativa (pdfjs-dist vs pdf-parse)
|
|
632
|
+
- Extrair como imagem e usar OCR se o PDF for escaneado
|
|
633
|
+
|
|
634
|
+
---
|
|
635
|
+
|
|
636
|
+
## Checklist de Conclusão
|
|
637
|
+
|
|
638
|
+
- [ ] robots.txt e ToS verificados
|
|
639
|
+
- [ ] API oficial descartada como alternativa
|
|
640
|
+
- [ ] Ferramenta escolhida adequada (Puppeteer / Playwright / Cheerio / parser)
|
|
641
|
+
- [ ] Rate limiting implementado
|
|
642
|
+
- [ ] Retry com backoff exponencial
|
|
643
|
+
- [ ] Validação dos dados extraídos
|
|
644
|
+
- [ ] Detecção de mudança de estrutura
|
|
645
|
+
- [ ] Formato de saída definido (JSON / CSV / NDJSON)
|
|
646
|
+
- [ ] Testes com amostra real dos dados
|
|
647
|
+
|
|
648
|
+
---
|
|
649
|
+
|
|
650
|
+
## Output
|
|
651
|
+
|
|
652
|
+
| Artefato | Descrição |
|
|
653
|
+
|----------|-----------|
|
|
654
|
+
| Scraper | Script de extração com retry, rate limiting e validação |
|
|
655
|
+
| Parser | Lógica de parsing adaptada à estrutura da fonte |
|
|
656
|
+
| Pipeline | ETL completo: extração → transformação → arquivo de saída |
|
|
657
|
+
| Monitor | Script de detecção de mudanças com diff estruturado |
|
|
658
|
+
|
|
659
|
+
---
|
|
660
|
+
|
|
661
|
+
## Mensagem de Conclusão
|
|
662
|
+
|
|
663
|
+
```
|
|
664
|
+
Scraper implementado!
|
|
665
|
+
|
|
666
|
+
Fonte: {URL ou tipo de fonte}
|
|
667
|
+
Ferramenta: {Puppeteer / Playwright / Cheerio / pdf-parse / xml-parser}
|
|
668
|
+
Dados extraídos: {campos coletados}
|
|
669
|
+
Rate limiting: {delay entre requests}
|
|
670
|
+
Retry: {N tentativas com backoff exponencial}
|
|
671
|
+
|
|
672
|
+
Saída: {JSON / CSV / NDJSON em path/to/output}
|
|
673
|
+
Validação: {campos obrigatórios verificados}
|
|
674
|
+
|
|
675
|
+
Próximo passo: {executar com amostra real / agendar com cron / integrar no pipeline}
|
|
676
|
+
```
|
|
677
|
+
|
|
678
|
+
---
|
|
679
|
+
|
|
680
|
+
## Recursos Adicionais
|
|
681
|
+
|
|
682
|
+
- **Referências**: Veja [references/](references/) para links de documentação local
|
|
683
|
+
- **Puppeteer docs**: https://pptr.dev (documentação oficial)
|