@bacnh85/pi-web 0.5.6 → 0.6.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +13 -0
- package/README.md +14 -8
- package/extensions/index.ts +16 -10
- package/extensions/lib/agy.ts +201 -0
- package/extensions/lib/extract.ts +44 -5
- package/package.json +2 -2
- package/skills/pi-web/SKILL.md +12 -5
package/CHANGELOG.md
CHANGED
|
@@ -1,7 +1,20 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.6.0 (2026-08-07)
|
|
4
|
+
|
|
5
|
+
### Features
|
|
6
|
+
|
|
7
|
+
- **agy extraction backend:** `web_extract` gains a new `agy` mode that uses the Antigravity CLI (Gemini/Claude) native `read_url` web tool to fetch bot-protected and anti-AI-scraping pages that block Firecrawl/Crawl4AI. `auto` mode now falls back static → dynamic → full → agy; explicit `mode: "agy"` forces it. Structured extraction (`prompt`/`schema`) is supported. `web_status` reports `agy.installed`.
|
|
8
|
+
- agy is optional and self-contained: if the CLI is not installed, `auto` mode skips it silently and existing flows are unchanged. Install: `curl -fsSL https://antigravity.google/cli/install.sh | bash`, then authenticate once with `agy`.
|
|
9
|
+
|
|
3
10
|
All notable changes to `pi-web` will be documented in this file.
|
|
4
11
|
|
|
12
|
+
## 0.5.7 (2026-08-05)
|
|
13
|
+
|
|
14
|
+
### Improvements
|
|
15
|
+
|
|
16
|
+
- Patch version bump for release sync and package documentation update.
|
|
17
|
+
|
|
5
18
|
## 0.5.6 (2026-08-01)
|
|
6
19
|
|
|
7
20
|
### Features
|
package/README.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
|
|
3
3
|
Pi extension for **unified web search, content extraction, site crawling, and page capture**.
|
|
4
4
|
|
|
5
|
-
Auto-selects the best backend from SearXNG (self-hosted), Brave Search, Firecrawl, and
|
|
5
|
+
Auto-selects the best backend from SearXNG (self-hosted), Brave Search, Firecrawl, Crawl4AI, and agy (Gemini/Claude, when installed) — so agents don't have to know which backend to use. Search selection is adaptive: broad discovery prefers self-hosted SearXNG, while precision-sensitive searches and inline content prefer Brave.
|
|
6
6
|
|
|
7
7
|
## Install
|
|
8
8
|
|
|
@@ -79,12 +79,13 @@ Use `backend` parameter to force a specific backend when needed.
|
|
|
79
79
|
|
|
80
80
|
### `web_extract` — Unified content extraction
|
|
81
81
|
|
|
82
|
-
Extracts readable content from a URL. Auto-selects backend: static (JSDOM) → dynamic (Firecrawl) → full (Crawl4AI), with extraction diagnostics showing fallback attempts.
|
|
82
|
+
Extracts readable content from a URL. Auto-selects backend: static (JSDOM) → dynamic (Firecrawl) → full (Crawl4AI) → agy (model-backed), with extraction diagnostics showing fallback attempts.
|
|
83
83
|
|
|
84
84
|
```
|
|
85
85
|
web_extract url="https://docs.ansible.com/..."
|
|
86
86
|
web_extract url="https://riven.tv/" mode=static
|
|
87
87
|
web_extract url="https://example.com" mode=dynamic prompt="Extract pricing plans"
|
|
88
|
+
web_extract url="https://blocked.example.com" mode=agy
|
|
88
89
|
```
|
|
89
90
|
|
|
90
91
|
Parameters:
|
|
@@ -92,9 +93,9 @@ Parameters:
|
|
|
92
93
|
| Parameter | Type | Default | Description |
|
|
93
94
|
|---|---|---|---|
|
|
94
95
|
| `url` | string | — | URL to extract |
|
|
95
|
-
| `mode` | string | `auto` | `auto`, `static`, `dynamic`, or `
|
|
96
|
-
| `prompt` | string | — | Prompt for JSON extraction (dynamic
|
|
97
|
-
| `schema` | any | — | JSON schema for structured extraction (dynamic
|
|
96
|
+
| `mode` | string | `auto` | `auto`, `static`, `dynamic`, `full`, or `agy` |
|
|
97
|
+
| `prompt` | string | — | Prompt for JSON extraction (dynamic/agy modes) |
|
|
98
|
+
| `schema` | any | — | JSON schema for structured extraction (dynamic/agy modes) |
|
|
98
99
|
| `content_chars` | number | 20000 | Max content chars |
|
|
99
100
|
| `wait_for` | number | — | Milliseconds to wait for Firecrawl dynamic rendering. Crawl4AI `/md` full mode may ignore this. |
|
|
100
101
|
| `mobile` | boolean | false | Emulate mobile viewport (dynamic mode) |
|
|
@@ -106,11 +107,14 @@ Parameters:
|
|
|
106
107
|
| `static` | JSDOM+Readability | Simple static pages, blog posts, docs | No |
|
|
107
108
|
| `dynamic` | Firecrawl Scrape | JS-rendered pages, dynamic content | Maybe |
|
|
108
109
|
| `full` | Crawl4AI | JS-heavy SPA, complex rendering | Maybe |
|
|
109
|
-
| `
|
|
110
|
+
| `agy` | agy (Gemini/Claude) | Bot-protected / anti-AI-scraping pages | agy CLI installed |
|
|
111
|
+
| `auto` (default) | static → dynamic → full → agy | Unknown page type | Maybe |
|
|
110
112
|
|
|
111
113
|
In `auto` mode, fallbacks are noted in the output (e.g., `[Extraction fell back to Firecrawl Scrape (dynamic mode)]`). If `static` extraction fails, the tool gracefully escalates to heavier backends.
|
|
112
114
|
|
|
113
|
-
> ⚠️ **Note on Firecrawl Scrape**: Fails on bot-protected sites (Ansible docs, many CDN-backed doc sites). Falls back to `full` mode (Crawl4AI) in `auto` mode.
|
|
115
|
+
> ⚠️ **Note on Firecrawl Scrape**: Fails on bot-protected sites (Ansible docs, many CDN-backed doc sites). Falls back to `full` mode (Crawl4AI) in `auto` mode, and to `agy` mode as a last resort.
|
|
116
|
+
|
|
117
|
+
> **`agy` mode (optional)**: Uses the [Antigravity CLI](https://antigravity.google/) with Gemini/Claude — its native `read_url` browser tool can fetch pages that block Firecrawl/Crawl4AI. Install with `curl -fsSL https://antigravity.google/cli/install.sh | bash`, authenticate once with `agy`, then `auto` mode falls back to it automatically. If agy is not installed, `auto` mode skips it silently; `web_status` reports `agy.installed`.
|
|
114
118
|
|
|
115
119
|
### `web_map` — Site URL discovery
|
|
116
120
|
|
|
@@ -172,7 +176,8 @@ Typical output:
|
|
|
172
176
|
"baseUrl": "http://172.30.55.22:11235",
|
|
173
177
|
...
|
|
174
178
|
"health": { "status": "healthy", "version": "0.5.0", ... }
|
|
175
|
-
}
|
|
179
|
+
},
|
|
180
|
+
"agy": { "installed": true }
|
|
176
181
|
}
|
|
177
182
|
```
|
|
178
183
|
|
|
@@ -188,6 +193,7 @@ Typical output:
|
|
|
188
193
|
| `lib/searxng.ts` | SearXNG metasearch fetch client (internal) |
|
|
189
194
|
| `lib/firecrawl.ts` | Firecrawl API fetch client with v2→v1 fallback (internal) |
|
|
190
195
|
| `lib/crawl4ai.ts` | Crawl4AI Docker API fetch client (internal) |
|
|
196
|
+
| `lib/agy.ts` | agy (Antigravity CLI) spawn helper — `read_url` extraction via Gemini/Claude |
|
|
191
197
|
| `lib/search.ts` | Unified search orchestrator — probes backends, fallback chain |
|
|
192
198
|
| `lib/extract.ts` | Unified extraction orchestrator — mode-based backend selection |
|
|
193
199
|
|
package/extensions/index.ts
CHANGED
|
@@ -62,7 +62,8 @@ The pi-web extension provides 7 unified tools that auto-select the best backend:
|
|
|
62
62
|
- **\`web_search\`** — Search the web (auto: SearXNG → Brave → Firecrawl). Use
|
|
63
63
|
\`backend\` for explicit control, \`engines\` for SearXNG tuning.
|
|
64
64
|
- **\`web_extract\`** — Extract readable content from a URL (auto: static JSDOM
|
|
65
|
-
→ dynamic Firecrawl → full Crawl4AI). Use \`mode\` for
|
|
65
|
+
→ dynamic Firecrawl → full Crawl4AI → agy model-backed). Use \`mode\` for
|
|
66
|
+
explicit control.
|
|
66
67
|
- **\`web_map\`** — Discover URLs from a site (Firecrawl Map).
|
|
67
68
|
- **\`web_crawl\`** — Crawl multiple pages. \`mode: "light"\` (Firecrawl) or
|
|
68
69
|
\`mode: "full"\` (Crawl4AI).
|
|
@@ -74,7 +75,8 @@ Backend selection rules:
|
|
|
74
75
|
- Firecrawl Search has poor semantic accuracy on domain-specific queries; prefer
|
|
75
76
|
SearXNG or Brave for precision (force via \`backend\`).
|
|
76
77
|
- Firecrawl Scrape fails on bot-protected sites (e.g. Ansible docs); Crawl4AI
|
|
77
|
-
handles those (force via \`mode: "full"\`)
|
|
78
|
+
handles those (force via \`mode: "full"\`), and agy (Gemini/Claude read_url)
|
|
79
|
+
handles the rest as a last-resort fallback (force via \`mode: "agy"\`).
|
|
78
80
|
- Always cite source URLs when web results materially support an answer.`;
|
|
79
81
|
|
|
80
82
|
|
|
@@ -136,23 +138,24 @@ export default function piWebExtension(pi: ExtensionAPI) {
|
|
|
136
138
|
name: "web_extract",
|
|
137
139
|
label: "Web Content Extraction",
|
|
138
140
|
description:
|
|
139
|
-
"Extract readable content from a URL. Auto mode: static\u2192dynamic\u2192full.",
|
|
141
|
+
"Extract readable content from a URL. Auto mode: static\u2192dynamic\u2192full\u2192agy.",
|
|
140
142
|
promptSnippet: "Extract readable webpage content as markdown",
|
|
141
143
|
promptGuidelines: [
|
|
142
144
|
"Clean markdown from a known URL.",
|
|
143
|
-
"mode: 'static' (no API key, JSDOM), 'dynamic' (Firecrawl JS), 'full' (Crawl4AI).",
|
|
144
|
-
"'auto' tries static\u2192dynamic\u2192full; see diagnostics for fallback chain.",
|
|
145
|
-
"
|
|
145
|
+
"mode: 'static' (no API key, JSDOM), 'dynamic' (Firecrawl JS), 'full' (Crawl4AI), 'agy' (Gemini/Claude via agy).",
|
|
146
|
+
"'auto' tries static\u2192dynamic\u2192full\u2192agy; see diagnostics for fallback chain.",
|
|
147
|
+
"mode: 'agy' uses agy's native read_url for bot-protected/JS-heavy pages \u2014 last-resort fallback in auto.",
|
|
148
|
+
"Use prompt+schema for structured JSON extraction (dynamic/agy modes).",
|
|
146
149
|
"Cite the source URL.",
|
|
147
150
|
],
|
|
148
151
|
parameters: Type.Object({
|
|
149
152
|
url: Type.String(),
|
|
150
153
|
mode: Type.Optional(Type.Union(
|
|
151
|
-
[Type.Literal("auto"), Type.Literal("static"), Type.Literal("dynamic"), Type.Literal("full")],
|
|
152
|
-
{ default: "auto", description: "auto, static, dynamic, full." },
|
|
154
|
+
[Type.Literal("auto"), Type.Literal("static"), Type.Literal("dynamic"), Type.Literal("full"), Type.Literal("agy")],
|
|
155
|
+
{ default: "auto", description: "auto, static, dynamic, full, agy." },
|
|
153
156
|
)),
|
|
154
|
-
prompt: Type.Optional(Type.String({ description: "Prompt for structured JSON extraction (dynamic
|
|
155
|
-
schema: Type.Optional(Type.Any({ description: "JSON schema for structured extraction (dynamic
|
|
157
|
+
prompt: Type.Optional(Type.String({ description: "Prompt for structured JSON extraction (dynamic/agy modes)." })),
|
|
158
|
+
schema: Type.Optional(Type.Any({ description: "JSON schema for structured extraction (dynamic/agy modes)." })),
|
|
156
159
|
content_chars: Type.Optional(Type.Number({ default: 20000 })),
|
|
157
160
|
wait_for: Type.Optional(Type.Number({ description: "Ms to wait for Firecrawl render before extraction." })),
|
|
158
161
|
mobile: Type.Optional(Type.Boolean({ default: false, description: "Mobile viewport (dynamic mode only)." })),
|
|
@@ -407,6 +410,8 @@ export default function piWebExtension(pi: ExtensionAPI) {
|
|
|
407
410
|
const c4aiUrl = findEnvValue("CRAWL4AI_API_URL", cwd, trusted);
|
|
408
411
|
const c4aiToken = findEnvValue("CRAWL4AI_API_TOKEN", cwd, trusted);
|
|
409
412
|
|
|
413
|
+
const { isAgyInstalled } = await import("./lib/agy");
|
|
414
|
+
|
|
410
415
|
const fcBaseUrl = normalizeFirecrawlBaseUrl(fireUrl.value);
|
|
411
416
|
const fcHosted = !fireUrl.value || fcBaseUrl.startsWith(HOSTED_FIRECRAWL_BASE_URL);
|
|
412
417
|
|
|
@@ -427,6 +432,7 @@ export default function piWebExtension(pi: ExtensionAPI) {
|
|
|
427
432
|
apiTokenFound: Boolean(c4aiToken.value),
|
|
428
433
|
apiTokenSource: c4aiToken.value ? c4aiToken.source : "not set",
|
|
429
434
|
},
|
|
435
|
+
agy: { installed: isAgyInstalled() },
|
|
430
436
|
};
|
|
431
437
|
|
|
432
438
|
// Crawl4AI health check
|
|
@@ -0,0 +1,201 @@
|
|
|
1
|
+
// agy-based web extraction — uses agy's native read_url tool via Gemini/Claude.
|
|
2
|
+
// Self-contained: does NOT import from pi-agy (different concern, no coupling).
|
|
3
|
+
//
|
|
4
|
+
// Runs agy in --mode plan (read-only: no file writes). Verified (agy 1.1.11):
|
|
5
|
+
// plan mode auto-approves read_url in headless -p WITHOUT
|
|
6
|
+
// --dangerously-skip-permissions, so no broad permission bypass is needed.
|
|
7
|
+
|
|
8
|
+
import { createRequire } from "node:module";
|
|
9
|
+
|
|
10
|
+
import { sanitizeError } from "./format";
|
|
11
|
+
|
|
12
|
+
// ponytail: require (not ESM import) so tests can patch cp.spawn/spawnSync
|
|
13
|
+
// without mocking the module graph — same pattern as pi-agy.
|
|
14
|
+
const _require = createRequire(import.meta.url);
|
|
15
|
+
const cp = _require("node:child_process") as typeof import("node:child_process");
|
|
16
|
+
|
|
17
|
+
const AGY_FETCH_TIMEOUT_MS = 90_000; // agy needs time for model call + web fetch
|
|
18
|
+
const AGY_PROBE_TIMEOUT_MS = 5_000;
|
|
19
|
+
const AGY_MAX_OUTPUT_BYTES = 200_000; // bound output to protect Pi context
|
|
20
|
+
export const AGY_MODEL = "gemini-3.6-flash-medium"; // ponytail: fixed default; users needing model control use agy_execute
|
|
21
|
+
|
|
22
|
+
// Cache install status with a TTL — spawnSync blocks the event loop up to
|
|
23
|
+
// AGY_PROBE_TIMEOUT_MS, and web_status/extract can call this repeatedly.
|
|
24
|
+
let agyInstalledCache: { ok: boolean; at: number } | null = null;
|
|
25
|
+
const AGY_INSTALL_CACHE_TTL_MS = 60_000;
|
|
26
|
+
|
|
27
|
+
export function isAgyInstalled(): boolean {
|
|
28
|
+
if (agyInstalledCache && Date.now() - agyInstalledCache.at < AGY_INSTALL_CACHE_TTL_MS) {
|
|
29
|
+
return agyInstalledCache.ok;
|
|
30
|
+
}
|
|
31
|
+
// ponytail: spawnSync is the simplest reliable probe; result is cached so the
|
|
32
|
+
// event-loop block happens at most once per 60s.
|
|
33
|
+
try {
|
|
34
|
+
const r = cp.spawnSync("agy", ["--version"], { timeout: AGY_PROBE_TIMEOUT_MS, stdio: "ignore" });
|
|
35
|
+
agyInstalledCache = { ok: r.status === 0, at: Date.now() };
|
|
36
|
+
return agyInstalledCache.ok;
|
|
37
|
+
} catch {
|
|
38
|
+
agyInstalledCache = { ok: false, at: Date.now() };
|
|
39
|
+
return false;
|
|
40
|
+
}
|
|
41
|
+
}
|
|
42
|
+
|
|
43
|
+
// Test-only: clear the cached install status.
|
|
44
|
+
export function resetAgyInstalledCache(): void {
|
|
45
|
+
agyInstalledCache = null;
|
|
46
|
+
}
|
|
47
|
+
|
|
48
|
+
// Validate + sanitize a URL before it is interpolated into the agy model prompt.
|
|
49
|
+
// Prompt-injection guard: reject non-http(s) schemes and strip control chars/newlines
|
|
50
|
+
// so a crafted URL cannot break out of the "fetch this URL" framing.
|
|
51
|
+
export function sanitizeAgyUrl(url: string): string {
|
|
52
|
+
const trimmed = url.trim();
|
|
53
|
+
let parsed: URL;
|
|
54
|
+
try {
|
|
55
|
+
parsed = new URL(trimmed);
|
|
56
|
+
} catch {
|
|
57
|
+
throw new Error(`Invalid URL for agy extraction: ${trimmed.slice(0, 200)}`);
|
|
58
|
+
}
|
|
59
|
+
if (parsed.protocol !== "http:" && parsed.protocol !== "https:") {
|
|
60
|
+
throw new Error(`Unsupported URL scheme for agy extraction: ${parsed.protocol}`);
|
|
61
|
+
}
|
|
62
|
+
// Neutralize any control characters/newlines so the URL stays on one line
|
|
63
|
+
// inside the prompt and cannot smuggle instructions.
|
|
64
|
+
return trimmed.replace(/[\x00-\x1f\x7f]/g, "");
|
|
65
|
+
}
|
|
66
|
+
|
|
67
|
+
// Build the agy CLI args for a web-fetch task. Pure — tested without mocking.
|
|
68
|
+
export function buildAgyFetchArgs(url: string, prompt?: string, schema?: unknown): string[] {
|
|
69
|
+
const safeUrl = sanitizeAgyUrl(url);
|
|
70
|
+
const structured = prompt || schema !== undefined;
|
|
71
|
+
const jsonInstruction = schema !== undefined
|
|
72
|
+
? `\nReturn as JSON matching this schema: ${JSON.stringify(schema)}`
|
|
73
|
+
: "\nReturn the result as JSON.";
|
|
74
|
+
const fetchInstruction = structured
|
|
75
|
+
? (prompt ? `Then extract this information: ${prompt}` : "Then extract the requested fields.") + jsonInstruction
|
|
76
|
+
: "Return the full page content as clean markdown.";
|
|
77
|
+
|
|
78
|
+
const agyPrompt = `Use your read_url web tool to fetch this URL: ${safeUrl}\n\n${fetchInstruction}\n\nReturn ONLY the result, no commentary.`;
|
|
79
|
+
|
|
80
|
+
return [
|
|
81
|
+
"--model",
|
|
82
|
+
AGY_MODEL,
|
|
83
|
+
"--mode",
|
|
84
|
+
"plan", // read-only: no file writes
|
|
85
|
+
"--print-timeout",
|
|
86
|
+
"90s",
|
|
87
|
+
// Verified (agy 1.1.11): plan mode auto-approves read_url in headless -p
|
|
88
|
+
// WITHOUT --dangerously-skip-permissions — pi-agy's flag is only needed for
|
|
89
|
+
// write modes (accept-edits/sandbox). Omitting it keeps the auto-approval
|
|
90
|
+
// surface at read-only web tools only.
|
|
91
|
+
"--output-format",
|
|
92
|
+
"json", // structured response for clean parsing
|
|
93
|
+
"-p",
|
|
94
|
+
agyPrompt,
|
|
95
|
+
];
|
|
96
|
+
}
|
|
97
|
+
|
|
98
|
+
export async function extractViaAgy(params: {
|
|
99
|
+
url: string;
|
|
100
|
+
prompt?: string; // structured extraction prompt
|
|
101
|
+
schema?: unknown; // JSON schema for structured extraction
|
|
102
|
+
contentChars?: number;
|
|
103
|
+
signal?: AbortSignal;
|
|
104
|
+
}): Promise<string> {
|
|
105
|
+
const { url, prompt, schema, contentChars, signal } = params;
|
|
106
|
+
const output = await spawnAgyRaw(buildAgyFetchArgs(url, prompt, schema), signal);
|
|
107
|
+
return parseAgyResponse(output, contentChars ?? 20000);
|
|
108
|
+
}
|
|
109
|
+
|
|
110
|
+
// Parse structured JSON out of agy's model output — used when a prompt/schema was
|
|
111
|
+
// requested so the structured result is first-class, matching dynamic mode's
|
|
112
|
+
// `structured` field. Handles both fenced (```json ... ```) and bare JSON, which
|
|
113
|
+
// the model produces nondeterministically. Pure — tested without mocking.
|
|
114
|
+
export function parseAgyStructured(markdown: string): unknown {
|
|
115
|
+
const trimmed = markdown.trim();
|
|
116
|
+
if (!trimmed) return undefined;
|
|
117
|
+
// Fenced block first, then bare JSON object/array.
|
|
118
|
+
const m = trimmed.match(/^```(?:json)?\s*([\s\S]*?)```/);
|
|
119
|
+
const candidate = m ? m[1] : trimmed.startsWith("{") || trimmed.startsWith("[") ? trimmed : undefined;
|
|
120
|
+
if (candidate === undefined) return undefined;
|
|
121
|
+
try {
|
|
122
|
+
return JSON.parse(candidate);
|
|
123
|
+
} catch {
|
|
124
|
+
return undefined;
|
|
125
|
+
}
|
|
126
|
+
}
|
|
127
|
+
|
|
128
|
+
// Spawn agy with bounded output collection + proper error handling.
|
|
129
|
+
async function spawnAgyRaw(args: string[], signal?: AbortSignal): Promise<string> {
|
|
130
|
+
// ponytail: agy is a Go binary in PATH (official installer); no shell wrapper.
|
|
131
|
+
const child = cp.spawn("agy", args, {
|
|
132
|
+
stdio: ["ignore", "pipe", "pipe"],
|
|
133
|
+
signal,
|
|
134
|
+
timeout: AGY_FETCH_TIMEOUT_MS + 5_000, // let agy report its own print timeout first
|
|
135
|
+
});
|
|
136
|
+
|
|
137
|
+
const stdout: Buffer[] = [];
|
|
138
|
+
const stderr: Buffer[] = [];
|
|
139
|
+
let stdoutBytes = 0;
|
|
140
|
+
let stderrBytes = 0;
|
|
141
|
+
|
|
142
|
+
child.stdout.on("data", (d: Buffer) => {
|
|
143
|
+
// Capture all bytes up to the cap, including the partial final chunk
|
|
144
|
+
// (mirrors pi-agy's appendBounded behavior).
|
|
145
|
+
if (stdoutBytes < AGY_MAX_OUTPUT_BYTES) {
|
|
146
|
+
stdout.push(d.subarray(0, AGY_MAX_OUTPUT_BYTES - stdoutBytes));
|
|
147
|
+
}
|
|
148
|
+
stdoutBytes += d.length;
|
|
149
|
+
});
|
|
150
|
+
child.stderr.on("data", (d: Buffer) => {
|
|
151
|
+
stderrBytes += d.length;
|
|
152
|
+
if (stderrBytes <= 16 * 1024) stderr.push(d);
|
|
153
|
+
});
|
|
154
|
+
|
|
155
|
+
return new Promise<string>((resolve, reject) => {
|
|
156
|
+
let settled = false;
|
|
157
|
+
const done = (fn: () => void) => {
|
|
158
|
+
if (settled) return;
|
|
159
|
+
settled = true;
|
|
160
|
+
fn();
|
|
161
|
+
};
|
|
162
|
+
|
|
163
|
+
child.on("error", (err: Error) => {
|
|
164
|
+
done(() => {
|
|
165
|
+
if (signal?.aborted) reject(new Error("agy was cancelled"));
|
|
166
|
+
else if ((err as NodeJS.ErrnoException).code === "ENOENT") {
|
|
167
|
+
reject(new Error("Antigravity CLI (agy) not found in PATH. Install: curl -fsSL https://antigravity.google/cli/install.sh | bash"));
|
|
168
|
+
} else reject(new Error(`agy spawn failed: ${sanitizeError(err)}`));
|
|
169
|
+
});
|
|
170
|
+
});
|
|
171
|
+
|
|
172
|
+
child.on("close", (code: number | null, sig: string | null) => {
|
|
173
|
+
done(() => {
|
|
174
|
+
const out = Buffer.concat(stdout).toString("utf8");
|
|
175
|
+
const err = Buffer.concat(stderr).toString("utf8");
|
|
176
|
+
|
|
177
|
+
if (sig === "SIGTERM" || sig === "SIGKILL" || code === null) {
|
|
178
|
+
reject(new Error(`agy was cancelled (${sig || "timeout"})`));
|
|
179
|
+
} else if (code !== 0) {
|
|
180
|
+
reject(new Error(`agy exited with code ${code}: ${sanitizeError((err || out).slice(0, 1000).trim() || "(no output)")}`));
|
|
181
|
+
} else {
|
|
182
|
+
// stdout only — stderr may carry warnings that would corrupt JSON envelope parsing
|
|
183
|
+
resolve(out);
|
|
184
|
+
}
|
|
185
|
+
});
|
|
186
|
+
});
|
|
187
|
+
});
|
|
188
|
+
}
|
|
189
|
+
|
|
190
|
+
// Extract .response from agy's JSON envelope, fall back to raw text. Pure — tested without mocking.
|
|
191
|
+
export function parseAgyResponse(raw: string, maxChars: number): string {
|
|
192
|
+
let text = raw;
|
|
193
|
+
try {
|
|
194
|
+
const parsed = JSON.parse(raw);
|
|
195
|
+
if (typeof parsed.response === "string") text = parsed.response;
|
|
196
|
+
else if (parsed.response !== undefined) text = JSON.stringify(parsed.response);
|
|
197
|
+
} catch {
|
|
198
|
+
// not JSON — use raw
|
|
199
|
+
}
|
|
200
|
+
return text.slice(0, maxChars);
|
|
201
|
+
}
|
|
@@ -3,16 +3,18 @@
|
|
|
3
3
|
// "static" → JSDOM+Readability (no external API)
|
|
4
4
|
// "dynamic" → Firecrawl Scrape (JS rendering/structured JSON)
|
|
5
5
|
// "full" → Crawl4AI markdown endpoint
|
|
6
|
-
// "
|
|
6
|
+
// "agy" → agy (Gemini/Claude) native read_url — bot-protected/JS-heavy pages
|
|
7
|
+
// "auto" → static → dynamic → full → agy
|
|
7
8
|
|
|
8
9
|
import { cwdFromContext, includeProjectEnv } from "./config";
|
|
9
10
|
import { loadFirecrawlConfig, loadCrawl4aiConfig, type FirecrawlConfig, type Crawl4aiConfig } from "./config";
|
|
10
11
|
import { fetchReadableContent } from "./content";
|
|
11
12
|
import { firecrawlRequest, type FirecrawlResult } from "./firecrawl";
|
|
12
13
|
import { fetchCrawl4aiMarkdown } from "./crawl4ai";
|
|
14
|
+
import { isAgyInstalled, extractViaAgy, parseAgyStructured } from "./agy";
|
|
13
15
|
import { formatFirecrawlScrape, sanitizeError } from "./format";
|
|
14
16
|
|
|
15
|
-
export type ExtractMode = "auto" | "static" | "dynamic" | "full";
|
|
17
|
+
export type ExtractMode = "auto" | "static" | "dynamic" | "full" | "agy";
|
|
16
18
|
export type ExtractAttemptStatus = "success" | "empty" | "error";
|
|
17
19
|
|
|
18
20
|
export interface ExtractParams {
|
|
@@ -142,13 +144,49 @@ async function extractFull(params: ExtractParams, ctx?: Record<string, unknown>)
|
|
|
142
144
|
function modesFor(params: ExtractParams): Array<Exclude<ExtractMode, "auto">> {
|
|
143
145
|
const mode = params.mode ?? "auto";
|
|
144
146
|
if (mode !== "auto") return [mode];
|
|
145
|
-
return ["static", "dynamic", "full"];
|
|
147
|
+
return ["static", "dynamic", "full", "agy"];
|
|
146
148
|
}
|
|
147
149
|
|
|
148
150
|
async function runExtractor(mode: Exclude<ExtractMode, "auto">, params: ExtractParams, ctx?: Record<string, unknown>): Promise<ExtractResult | null> {
|
|
149
151
|
if (mode === "static") return extractStatic(params);
|
|
150
152
|
if (mode === "dynamic") return extractDynamic(params, ctx);
|
|
151
|
-
return extractFull(params, ctx);
|
|
153
|
+
if (mode === "full") return extractFull(params, ctx);
|
|
154
|
+
return extractAgy(params);
|
|
155
|
+
}
|
|
156
|
+
|
|
157
|
+
async function extractAgy(params: ExtractParams): Promise<ExtractResult | null> {
|
|
158
|
+
if (!isAgyInstalled()) return null; // graceful skip — recorded as skipped attempt
|
|
159
|
+
const run = () => extractViaAgy({
|
|
160
|
+
url: params.url,
|
|
161
|
+
prompt: params.prompt,
|
|
162
|
+
schema: params.schema,
|
|
163
|
+
contentChars: params.content_chars,
|
|
164
|
+
signal: params.signal,
|
|
165
|
+
});
|
|
166
|
+
// ponytail: the model occasionally returns empty output (transient); one retry
|
|
167
|
+
// materially improves reliability at the cost of one extra spawn.
|
|
168
|
+
let markdown = await run();
|
|
169
|
+
if (!markdown.trim()) markdown = await run();
|
|
170
|
+
let structured: unknown;
|
|
171
|
+
// Match dynamic mode: prompt OR schema requests structured output.
|
|
172
|
+
if (params.prompt || params.schema !== undefined) {
|
|
173
|
+
structured = parseAgyStructured(markdown);
|
|
174
|
+
if (structured !== undefined) {
|
|
175
|
+
// Strip the JSON (fenced or bare) so the result is clean markdown; the
|
|
176
|
+
// structured payload is surfaced separately (rendered like dynamic mode).
|
|
177
|
+
const fenced = /```(?:json)?\s*[\s\S]*?```/.test(markdown);
|
|
178
|
+
let clean = fenced ? markdown.replace(/```(?:json)?\s*[\s\S]*?```/, "").trim() : "";
|
|
179
|
+
if (!fenced && markdown.trim().length > 0) {
|
|
180
|
+
// Bare JSON (no fence) — the whole output is JSON, so nothing remains.
|
|
181
|
+
clean = "";
|
|
182
|
+
}
|
|
183
|
+
// When the model returned only JSON, don't duplicate it as plain text body
|
|
184
|
+
// (it is already surfaced via structuredSection below).
|
|
185
|
+
const body = clean || "(Structured extraction only — see JSON below)";
|
|
186
|
+
return { title: "", markdown: body + structuredSection(structured), backend: "agy", structured };
|
|
187
|
+
}
|
|
188
|
+
}
|
|
189
|
+
return { title: "", markdown, backend: "agy" };
|
|
152
190
|
}
|
|
153
191
|
|
|
154
192
|
export async function extractWithDiagnostics(params: ExtractParams): Promise<ExtractDiagnostics> {
|
|
@@ -162,7 +200,8 @@ export async function extractWithDiagnostics(params: ExtractParams): Promise<Ext
|
|
|
162
200
|
const result = await runExtractor(mode, params, ctx);
|
|
163
201
|
if (isUseful(result, mode, explicit)) {
|
|
164
202
|
if (!explicit && attempts.length > 0) {
|
|
165
|
-
|
|
203
|
+
const backendLabel = mode === "dynamic" ? "Firecrawl Scrape (dynamic mode)" : mode === "full" ? "Crawl4AI (full browser mode)" : "agy (model-backed browser)";
|
|
204
|
+
result.markdown = `[Extraction fell back to ${backendLabel}]\n\n${result.markdown}`;
|
|
166
205
|
}
|
|
167
206
|
attempts.push({ mode, backend: result.backend, status: "success", message: `Selected ${mode}`, contentLength: result.markdown.length });
|
|
168
207
|
return { result, attempts, selectedMode: mode, fallbackUsed: attempts.length > 1 };
|
package/package.json
CHANGED
|
@@ -1,13 +1,13 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@bacnh85/pi-web",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.6.0",
|
|
4
4
|
"description": "Pi extension for web search, page extraction, Firecrawl scraping/crawling, and Crawl4AI headless browser crawling.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"license": "MIT",
|
|
7
7
|
"publishConfig": {
|
|
8
8
|
"access": "public"
|
|
9
9
|
},
|
|
10
|
-
"homepage": "https://github.com/bacnh85/pi-extensions
|
|
10
|
+
"homepage": "https://github.com/bacnh85/pi-extensions/tree/main/pi-web",
|
|
11
11
|
"repository": {
|
|
12
12
|
"type": "git",
|
|
13
13
|
"url": "git+https://github.com/bacnh85/pi-extensions.git",
|
package/skills/pi-web/SKILL.md
CHANGED
|
@@ -5,14 +5,14 @@ description: Web search, content extraction, site crawling, and page capture via
|
|
|
5
5
|
|
|
6
6
|
# pi-web — Unified Web Tools
|
|
7
7
|
|
|
8
|
-
Use the **7 unified tools** from the `pi-web` extension for all web-related tasks. These tools automatically select the best backend from SearXNG, Brave Search, Firecrawl, and
|
|
8
|
+
Use the **7 unified tools** from the `pi-web` extension for all web-related tasks. These tools automatically select the best backend from SearXNG, Brave Search, Firecrawl, Crawl4AI, and agy (when installed) — you don't need to know which backend to use. Search selection is adaptive: broad discovery prefers self-hosted SearXNG, while precision-sensitive queries and inline content prefer Brave.
|
|
9
9
|
|
|
10
10
|
## Quick Reference
|
|
11
11
|
|
|
12
12
|
| Tool | Purpose | Auto-selection |
|
|
13
13
|
|---|---|---|
|
|
14
14
|
| `web_search` | Search the web for sources, docs, facts | SearXNG → Brave → Firecrawl |
|
|
15
|
-
| `web_extract` | Extract readable content from a URL | Static (JSDOM) → Dynamic (Firecrawl) → Full (Crawl4AI) |
|
|
15
|
+
| `web_extract` | Extract readable content from a URL | Static (JSDOM) → Dynamic (Firecrawl) → Full (Crawl4AI) → agy (model-backed) |
|
|
16
16
|
| `web_map` | Discover URLs from a site | Firecrawl Map (only option) |
|
|
17
17
|
| `web_crawl` | Crawl multiple pages from a site | Light (Firecrawl) or Full (Crawl4AI) |
|
|
18
18
|
| `web_screenshot` | Capture page screenshot as PNG | Crawl4AI (only option) |
|
|
@@ -35,7 +35,8 @@ What do you need?
|
|
|
35
35
|
│ ├─ static page (blog, docs): mode=static (fastest, no API key)
|
|
36
36
|
│ ├─ dynamic page (JS-rendered): mode=dynamic
|
|
37
37
|
│ ├─ JS-heavy SPA: mode=full
|
|
38
|
-
│
|
|
38
|
+
│ ├─ bot-protected / blocked to scrapers: mode=agy (Gemini/Claude via agy)
|
|
39
|
+
│ └─ auto (default): tries static > dynamic > full > agy
|
|
39
40
|
│
|
|
40
41
|
├── Site URL discovery (find pages on a site)
|
|
41
42
|
│ → web_map
|
|
@@ -78,6 +79,9 @@ What do you need?
|
|
|
78
79
|
- ❌ Fails on bot-protected sites (Ansible docs, many CDN-backed doc sites). Falls through Crawl4AI in `auto` mode.
|
|
79
80
|
- ✅ Supported: prompt-based JSON extraction, schema-based structured extraction.
|
|
80
81
|
3. **full** (Crawl4AI headless browser) — handles all content types. **Resource-intensive** (launches a full headless browser). Use only when static and dynamic modes fail, or when explicitly needed.
|
|
82
|
+
4. **agy** (Gemini/Claude via agy CLI) — last-resort fallback. Uses agy's native `read_url` web tool with a model-driven browser, so it can fetch pages that block Firecrawl/Crawl4AI (bot protection, anti-AI scraping). Requires the `agy` CLI installed and authenticated. Also supports prompt/schema-based structured extraction.
|
|
83
|
+
|
|
84
|
+
> **Requirement for agy mode**: install and authenticate the Antigravity CLI: `curl -fsSL https://antigravity.google/cli/install.sh | bash` then run `agy` once interactively. If agy is not installed, `auto` mode skips it silently; explicit `mode: "agy"` reports the install hint.
|
|
81
85
|
|
|
82
86
|
## Fallback Strategy
|
|
83
87
|
|
|
@@ -85,8 +89,9 @@ If one tool fails, try the next option in the chain:
|
|
|
85
89
|
|
|
86
90
|
- **Search issues**: `web_search` auto-fallbacks and reports backend diagnostics. If all backends fail, configure at least one via env vars (check `web_status`).
|
|
87
91
|
- **Extraction issues**: `web_extract` auto-fallbacks in `auto` mode. If all modes fail:
|
|
88
|
-
1. Try `
|
|
89
|
-
2.
|
|
92
|
+
1. Try `mode: "agy"` explicitly — agy's model-backed browser often gets pages that block Firecrawl/Crawl4AI.
|
|
93
|
+
2. Try `web_screenshot` for a visual snapshot — may work when extraction is blocked.
|
|
94
|
+
3. The page may require interactive login, CAPTCHA, or be a non-HTML resource.
|
|
90
95
|
- **Tool not found**: Ensure `pi-web` extension is installed (`pi install ./extensions/pi-web`).
|
|
91
96
|
|
|
92
97
|
## Cross-tool Decision Guide
|
|
@@ -95,6 +100,7 @@ If one tool fails, try the next option in the chain:
|
|
|
95
100
|
|---|---|---|
|
|
96
101
|
| A few specific pages from a site | `web_map` + `web_extract` on each URL | `web_crawl` (heavier than needed) |
|
|
97
102
|
| Content from a JS-heavy page that fails in `auto` mode | `web_extract` with `mode: "full"` | Retrying `auto` mode repeatedly |
|
|
103
|
+
| Content from a bot-protected page that blocks Firecrawl/Crawl4AI | `web_extract` with `mode: "agy"` | Retrying `web_extract` with all modes |
|
|
98
104
|
| A visual of a bot-protected page | `web_screenshot` | Retrying `web_extract` with all modes |
|
|
99
105
|
| Content alongside search results | `web_search` with `include_content: true` (auto prefers Brave) or `backend: "brave"` | Search snippets alone |
|
|
100
106
|
| Printable/archivable page | `web_pdf` | Taking a screenshot and converting |
|
|
@@ -105,4 +111,5 @@ If one tool fails, try the next option in the chain:
|
|
|
105
111
|
- **Always cite source URLs** when using web content in answers.
|
|
106
112
|
- `web_status` shows which backends are configured without printing secrets. For Firecrawl, `apiKeyFound: false` is normal for self-hosted instances without auth — check the `ready` field to see if Firecrawl is actually usable.
|
|
107
113
|
- The `backend` and `mode` parameters give explicit control when auto-selection is not desired.
|
|
114
|
+
- `mode: "agy"` requires the `agy` CLI (Antigravity) installed and authenticated. `web_status` reports `agy.installed` so you can check availability without guessing.
|
|
108
115
|
- Backend-specific config (API keys, URLs) comes from environment variables, not tool parameters. Use `web_status` to verify configuration.
|