@prismer/runtime 2.0.7 → 2.0.8
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +201 -0
- package/built-in-skills/agent-coordination/SKILL.md +235 -0
- package/built-in-skills/agent-meta/SKILL.md +52 -0
- package/built-in-skills/assets/SKILL.md +131 -0
- package/built-in-skills/canvas-design/LICENSE.txt +202 -0
- package/built-in-skills/canvas-design/SKILL.md +156 -0
- package/built-in-skills/canvas-design/canvas-fonts/ArsenalSC-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/ArsenalSC-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/BigShoulders-Bold.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/BigShoulders-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/BigShoulders-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/Boldonse-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/Boldonse-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/BricolageGrotesque-Bold.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/BricolageGrotesque-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/BricolageGrotesque-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/CrimsonPro-Bold.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/CrimsonPro-Italic.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/CrimsonPro-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/CrimsonPro-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/DMMono-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/DMMono-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/EricaOne-OFL.txt +94 -0
- package/built-in-skills/canvas-design/canvas-fonts/EricaOne-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/GeistMono-Bold.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/GeistMono-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/GeistMono-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/Gloock-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/Gloock-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/IBMPlexMono-Bold.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/IBMPlexMono-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/IBMPlexMono-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/IBMPlexSerif-Bold.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/IBMPlexSerif-BoldItalic.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/IBMPlexSerif-Italic.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/IBMPlexSerif-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/InstrumentSans-Bold.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/InstrumentSans-BoldItalic.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/InstrumentSans-Italic.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/InstrumentSans-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/InstrumentSans-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/InstrumentSerif-Italic.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/InstrumentSerif-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/Italiana-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/Italiana-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/JetBrainsMono-Bold.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/JetBrainsMono-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/JetBrainsMono-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/Jura-Light.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/Jura-Medium.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/Jura-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/LibreBaskerville-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/LibreBaskerville-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/Lora-Bold.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/Lora-BoldItalic.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/Lora-Italic.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/Lora-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/Lora-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/NationalPark-Bold.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/NationalPark-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/NationalPark-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/NothingYouCouldDo-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/NothingYouCouldDo-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/Outfit-Bold.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/Outfit-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/Outfit-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/PixelifySans-Medium.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/PixelifySans-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/PoiretOne-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/PoiretOne-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/RedHatMono-Bold.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/RedHatMono-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/RedHatMono-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/Silkscreen-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/Silkscreen-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/SmoochSans-Medium.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/SmoochSans-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/Tektur-Medium.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/Tektur-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/Tektur-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/WorkSans-Bold.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/WorkSans-BoldItalic.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/WorkSans-Italic.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/WorkSans-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/WorkSans-Regular.ttf +0 -0
- package/built-in-skills/canvas-design/canvas-fonts/YoungSerif-OFL.txt +93 -0
- package/built-in-skills/canvas-design/canvas-fonts/YoungSerif-Regular.ttf +0 -0
- package/built-in-skills/claim-agent-ownership/SKILL.md +254 -0
- package/built-in-skills/claude-api/LICENSE.txt +202 -0
- package/built-in-skills/claude-api/SKILL.md +324 -0
- package/built-in-skills/claude-api/csharp/claude-api.md +402 -0
- package/built-in-skills/claude-api/curl/examples.md +216 -0
- package/built-in-skills/claude-api/curl/managed-agents.md +336 -0
- package/built-in-skills/claude-api/go/claude-api.md +421 -0
- package/built-in-skills/claude-api/go/managed-agents/README.md +561 -0
- package/built-in-skills/claude-api/java/claude-api.md +432 -0
- package/built-in-skills/claude-api/java/managed-agents/README.md +442 -0
- package/built-in-skills/claude-api/php/claude-api.md +375 -0
- package/built-in-skills/claude-api/php/managed-agents/README.md +435 -0
- package/built-in-skills/claude-api/python/claude-api/README.md +420 -0
- package/built-in-skills/claude-api/python/claude-api/batches.md +185 -0
- package/built-in-skills/claude-api/python/claude-api/files-api.md +165 -0
- package/built-in-skills/claude-api/python/claude-api/streaming.md +162 -0
- package/built-in-skills/claude-api/python/claude-api/tool-use.md +590 -0
- package/built-in-skills/claude-api/python/managed-agents/README.md +332 -0
- package/built-in-skills/claude-api/ruby/claude-api.md +113 -0
- package/built-in-skills/claude-api/ruby/managed-agents/README.md +389 -0
- package/built-in-skills/claude-api/shared/agent-design.md +101 -0
- package/built-in-skills/claude-api/shared/error-codes.md +213 -0
- package/built-in-skills/claude-api/shared/live-sources.md +135 -0
- package/built-in-skills/claude-api/shared/managed-agents-api-reference.md +378 -0
- package/built-in-skills/claude-api/shared/managed-agents-client-patterns.md +209 -0
- package/built-in-skills/claude-api/shared/managed-agents-core.md +238 -0
- package/built-in-skills/claude-api/shared/managed-agents-environments.md +215 -0
- package/built-in-skills/claude-api/shared/managed-agents-events.md +195 -0
- package/built-in-skills/claude-api/shared/managed-agents-memory.md +197 -0
- package/built-in-skills/claude-api/shared/managed-agents-multiagent.md +99 -0
- package/built-in-skills/claude-api/shared/managed-agents-onboarding.md +114 -0
- package/built-in-skills/claude-api/shared/managed-agents-outcomes.md +106 -0
- package/built-in-skills/claude-api/shared/managed-agents-overview.md +68 -0
- package/built-in-skills/claude-api/shared/managed-agents-self-hosted-sandboxes.md +173 -0
- package/built-in-skills/claude-api/shared/managed-agents-tools.md +321 -0
- package/built-in-skills/claude-api/shared/managed-agents-webhooks.md +110 -0
- package/built-in-skills/claude-api/shared/model-migration.md +779 -0
- package/built-in-skills/claude-api/shared/models.md +121 -0
- package/built-in-skills/claude-api/shared/prompt-caching.md +171 -0
- package/built-in-skills/claude-api/shared/tool-use-concepts.md +327 -0
- package/built-in-skills/claude-api/typescript/claude-api/README.md +333 -0
- package/built-in-skills/claude-api/typescript/claude-api/batches.md +106 -0
- package/built-in-skills/claude-api/typescript/claude-api/files-api.md +98 -0
- package/built-in-skills/claude-api/typescript/claude-api/streaming.md +178 -0
- package/built-in-skills/claude-api/typescript/claude-api/tool-use.md +527 -0
- package/built-in-skills/claude-api/typescript/managed-agents/README.md +359 -0
- package/built-in-skills/doc-coauthoring/SKILL.md +375 -0
- package/built-in-skills/frontend-design/LICENSE.txt +177 -0
- package/built-in-skills/frontend-design/SKILL.md +42 -0
- package/built-in-skills/human-approval/SKILL.md +114 -0
- package/built-in-skills/image-generate/SKILL.md +327 -0
- package/built-in-skills/ingest/SKILL.md +105 -0
- package/built-in-skills/internal-comms/LICENSE.txt +202 -0
- package/built-in-skills/internal-comms/SKILL.md +32 -0
- package/built-in-skills/internal-comms/examples/3p-updates.md +47 -0
- package/built-in-skills/internal-comms/examples/company-newsletter.md +65 -0
- package/built-in-skills/internal-comms/examples/faq-answers.md +30 -0
- package/built-in-skills/internal-comms/examples/general-comms.md +16 -0
- package/built-in-skills/liteparse/SKILL.md +156 -0
- package/built-in-skills/mcp-builder/LICENSE.txt +202 -0
- package/built-in-skills/mcp-builder/SKILL.md +236 -0
- package/built-in-skills/mcp-builder/reference/evaluation.md +602 -0
- package/built-in-skills/mcp-builder/reference/mcp_best_practices.md +249 -0
- package/built-in-skills/mcp-builder/reference/node_mcp_server.md +970 -0
- package/built-in-skills/mcp-builder/reference/python_mcp_server.md +719 -0
- package/built-in-skills/mcp-builder/scripts/connections.py +151 -0
- package/built-in-skills/mcp-builder/scripts/evaluation.py +373 -0
- package/built-in-skills/mcp-builder/scripts/example_evaluation.xml +22 -0
- package/built-in-skills/mcp-builder/scripts/requirements.txt +2 -0
- package/built-in-skills/memory/SKILL.md +106 -0
- package/built-in-skills/memory-curation/SKILL.md +135 -0
- package/built-in-skills/office-artifacts/SKILL.md +198 -0
- package/built-in-skills/prismer-im-collab/SKILL.md +148 -0
- package/built-in-skills/skill-authoring/SKILL.md +124 -0
- package/built-in-skills/skill-authoring/skill.json +74 -0
- package/built-in-skills/skill-creator/LICENSE.txt +202 -0
- package/built-in-skills/skill-creator/SKILL.md +485 -0
- package/built-in-skills/skill-creator/agents/analyzer.md +274 -0
- package/built-in-skills/skill-creator/agents/comparator.md +202 -0
- package/built-in-skills/skill-creator/agents/grader.md +223 -0
- package/built-in-skills/skill-creator/assets/eval_review.html +146 -0
- package/built-in-skills/skill-creator/eval-viewer/generate_review.py +471 -0
- package/built-in-skills/skill-creator/eval-viewer/viewer.html +1325 -0
- package/built-in-skills/skill-creator/references/schemas.md +430 -0
- package/built-in-skills/skill-creator/scripts/__init__.py +0 -0
- package/built-in-skills/skill-creator/scripts/aggregate_benchmark.py +401 -0
- package/built-in-skills/skill-creator/scripts/generate_report.py +326 -0
- package/built-in-skills/skill-creator/scripts/improve_description.py +247 -0
- package/built-in-skills/skill-creator/scripts/package_skill.py +136 -0
- package/built-in-skills/skill-creator/scripts/quick_validate.py +103 -0
- package/built-in-skills/skill-creator/scripts/run_eval.py +310 -0
- package/built-in-skills/skill-creator/scripts/run_loop.py +328 -0
- package/built-in-skills/skill-creator/scripts/utils.py +47 -0
- package/built-in-skills/slack-gif-creator/LICENSE.txt +202 -0
- package/built-in-skills/slack-gif-creator/SKILL.md +271 -0
- package/built-in-skills/slack-gif-creator/core/easing.py +234 -0
- package/built-in-skills/slack-gif-creator/core/frame_composer.py +176 -0
- package/built-in-skills/slack-gif-creator/core/gif_builder.py +269 -0
- package/built-in-skills/slack-gif-creator/core/validators.py +136 -0
- package/built-in-skills/slack-gif-creator/requirements.txt +4 -0
- package/built-in-skills/tasks/SKILL.md +398 -0
- package/built-in-skills/team/SKILL.md +76 -0
- package/built-in-skills/web-artifacts-builder/LICENSE.txt +202 -0
- package/built-in-skills/web-artifacts-builder/SKILL.md +104 -0
- package/built-in-skills/web-artifacts-builder/scripts/bundle-artifact.sh +54 -0
- package/built-in-skills/web-artifacts-builder/scripts/init-artifact.sh +334 -0
- package/built-in-skills/web-artifacts-builder/scripts/shadcn-components.tar.gz +0 -0
- package/built-in-skills/webapp-testing/LICENSE.txt +202 -0
- package/built-in-skills/webapp-testing/SKILL.md +96 -0
- package/built-in-skills/webapp-testing/examples/console_logging.py +35 -0
- package/built-in-skills/webapp-testing/examples/element_discovery.py +40 -0
- package/built-in-skills/webapp-testing/examples/static_html_automation.py +33 -0
- package/built-in-skills/webapp-testing/scripts/with_server.py +106 -0
- package/dist/cli.cjs +6195 -1631
- package/dist/cli.js +6148 -1585
- package/dist/index.cjs +6032 -1468
- package/dist/index.d.cts +813 -43
- package/dist/index.d.ts +813 -43
- package/dist/index.js +6017 -1454
- package/package.json +4 -2
|
@@ -0,0 +1,114 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: human-approval
|
|
3
|
+
description: Request human approval before performing a SAFETY-CRITICAL, IRREVERSIBLE, or SCOPE-EXPANDING action — submit a structured context (action, scope, risk, consequence) plus options, then STOP the current turn. The platform redispatches the agent after the human decides. NEVER use for routine deliverables (writing docs / generating files / summarising chats / answering questions / explaining concepts / read-only tool calls) — those are pre-authorized; just produce the output. The 5-minute-rollback litmus test applies: if the wrong outcome can be undone in <5 minutes by editing or deleting, it is NOT approval-eligible. Routine misuse (intro/explain/summarise/file-generation) burns the user's attention budget and is a contract violation. Executes via `cloud approval` CLI.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Human Approval
|
|
7
|
+
|
|
8
|
+
Some actions need a human in the loop **before** they execute: production deploys, large credit spend, deleting data, scope-expanding decisions. This skill submits a structured request and **halts the current turn**. The platform stores the request, notifies the human, and **redispatches the agent** with the decision when the human responds — you don't poll, you don't re-ask in the same turn.
|
|
9
|
+
|
|
10
|
+
## When to use
|
|
11
|
+
|
|
12
|
+
- **Production-impacting** action: deploy, schema change, infra reconfig, payment send.
|
|
13
|
+
- **Irreversible**: delete files, drop tables, revoke keys, close accounts.
|
|
14
|
+
- **Scope-expanding**: the task as-stated implies more changes than the user originally agreed to.
|
|
15
|
+
- **High credit cost**: any operation that would spend > expected budget.
|
|
16
|
+
- **Authority-elevating**: granting access, changing roles, modifying ACLs.
|
|
17
|
+
|
|
18
|
+
## Not when to use
|
|
19
|
+
|
|
20
|
+
- Routine clarifying questions ("what column name do you want?") — just ask in chat.
|
|
21
|
+
- Choosing between two equivalent options where the user clearly didn't care — pick one and proceed.
|
|
22
|
+
- When the user already explicitly approved this action in the current conversation — proceed.
|
|
23
|
+
|
|
24
|
+
## Skill scope guard (v2.0.8)
|
|
25
|
+
|
|
26
|
+
The "Not when to use" list above is the **load-bearing rule**. As of
|
|
27
|
+
release 2.0.8 we tightened it because routine deliverables (write a
|
|
28
|
+
doc, summarise a chat, draft a slide deck, answer a question) were
|
|
29
|
+
incorrectly triggering approval gates — the user got a yellow "等待
|
|
30
|
+
人工确认" banner for a request as simple as "@ceo 给我介绍一下产品 PDF",
|
|
31
|
+
which is a deliverable request, not a scope-expansion.
|
|
32
|
+
|
|
33
|
+
The following 8 categories are **never** approval-eligible. Run them
|
|
34
|
+
directly and report the result in the same turn:
|
|
35
|
+
|
|
36
|
+
| Category | Why it's not approval-eligible | Use instead |
|
|
37
|
+
|---|---|---|
|
|
38
|
+
| Writing a document / generating a report / outputting PDF, DOCX, PPTX, XLSX, CSV | The user *asked for the deliverable*; gating it is anti-UX. | Call `office-artifacts` and ship. |
|
|
39
|
+
| Summarising a conversation / writing meeting notes | Pure synthesis from data the user already has. | Reply in chat. |
|
|
40
|
+
| Answering a question / explaining a concept | The user invited the answer by asking. | Reply in chat. |
|
|
41
|
+
| Asking the user for a preference ("Chinese or English?") | A chat question is the correct affordance. | Ask in chat — `human-approval` is overkill. |
|
|
42
|
+
| Choosing model parameters / temperature / sampling style | Internal agent decision; users don't have context to judge. | Decide and proceed; mention the choice in the reply. |
|
|
43
|
+
| Naming files / picking output paths | Internal agent decision; reversible by renaming. | Pick sensible defaults; let user override if asked. |
|
|
44
|
+
| Internal brainstorming / scoring multiple candidates | The user asked for the *winner*, not the deliberation. | Do the work, surface the winner. |
|
|
45
|
+
| Calling read-only MCP tools (search, web fetch, file read) | No side effect; trivially reversible. | Call directly. |
|
|
46
|
+
|
|
47
|
+
**The litmus test:** "**If this step turns out wrong, can I roll it
|
|
48
|
+
back in under 5 minutes by editing or deleting something?**"
|
|
49
|
+
- If **yes** → not approval-eligible. Ship it.
|
|
50
|
+
- If **no** → safety-critical / irreversible / scope-expanding → approval-eligible.
|
|
51
|
+
|
|
52
|
+
Mis-using `human-approval` for routine work burns the user's attention
|
|
53
|
+
budget, breaks chat flow, and signals lack of agent confidence — all
|
|
54
|
+
three are real costs. The role templates (CEO / engineer / marketer /
|
|
55
|
+
researcher / verifier) carry an explicit `operatingPrinciples` line as
|
|
56
|
+
of 2.0.8: "Never trigger human-approval for routine deliverables".
|
|
57
|
+
|
|
58
|
+
## CLI Reference
|
|
59
|
+
|
|
60
|
+
**Anchor required.** `POST /api/im/approvals` rejects requests with neither `taskId` nor `conversationId`, because the platform needs a target to deliver the human decision to. Every invocation MUST pass one of `--task-id` or `--conversation-id`.
|
|
61
|
+
|
|
62
|
+
```bash
|
|
63
|
+
# Linked to a task — the platform resumes the task on decision (preferred for marketplace / agent flows)
|
|
64
|
+
cloud approval request-human \
|
|
65
|
+
--task-id <taskId> \
|
|
66
|
+
--action "approve marketplace task completion" \
|
|
67
|
+
--context "Result: scan-deps found 3 CVEs. Report attached." \
|
|
68
|
+
--risk "Releases 10-credit escrow to the assignee."
|
|
69
|
+
|
|
70
|
+
# Linked to a conversation — the decision is posted back as a chat message
|
|
71
|
+
cloud approval request-human \
|
|
72
|
+
--conversation-id <conversationId> \
|
|
73
|
+
--action "delete branch feat/old-experiment" \
|
|
74
|
+
--context "Last commit 2025-12-10. Merged into main. Local copy preserved." \
|
|
75
|
+
--risk "Irreversible. No remote backup; force-pushed commits would be lost."
|
|
76
|
+
|
|
77
|
+
# With explicit options (multi-choice)
|
|
78
|
+
cloud approval request-human \
|
|
79
|
+
--conversation-id <conversationId> \
|
|
80
|
+
--action "deploy v1.8.2 to prod" \
|
|
81
|
+
--context "All gates green, 9/9 webhook tests pass, test env stable 48h." \
|
|
82
|
+
--risk "Touches payment webhook. Rollback ETA 5min via git revert + redeploy." \
|
|
83
|
+
--options "deploy-now" "deploy-tomorrow-morning" "wait-for-manual-smoke-test"
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
## Workflow
|
|
87
|
+
|
|
88
|
+
1. **Summarize the action** in one sentence — what will happen, on what resource, with what permission.
|
|
89
|
+
2. **Provide context** — recent state, related artifacts, why this came up now.
|
|
90
|
+
3. **Spell out risk and consequence** — what breaks if this is wrong, what's reversible, what's not, who else is affected.
|
|
91
|
+
4. **Provide options** when binary approve/reject is insufficient. Options must be **mutually understandable** and **independently actionable** (each is a thing the agent can do without further clarification).
|
|
92
|
+
5. **Submit the approval request.** Capture the returned `approvalId`.
|
|
93
|
+
6. **Stop the current turn.** Don't ask follow-up questions, don't start the action, don't speculate about the answer. The platform will redispatch you when the human decides.
|
|
94
|
+
|
|
95
|
+
## Operating Rules
|
|
96
|
+
|
|
97
|
+
- **Do not proceed with the gated action in the same turn after requesting approval.** This is the load-bearing rule. The platform will redispatch the agent with the decision; running the action now defeats the gate.
|
|
98
|
+
- **Do not hide material risks or irreversible effects** from the approval context. The human is approving based on what you wrote — incomplete framing is worse than no gate.
|
|
99
|
+
- **Keep options mutually understandable and actionable.** "Approve" / "Approve with conditions" / "Reject" is fine. "Maybe" / "Let me think" is not — that's not a decision the human can choose.
|
|
100
|
+
- **Use this for safety-critical decisions**, not routine clarification. Asking the user "what label do you prefer?" via human-approval is overkill and burns their attention budget.
|
|
101
|
+
- **Link to a task** (`--task-id`) when the approval gates a task's progression. The platform resumes the task automatically when the human approves.
|
|
102
|
+
- If the user already explicitly approved this exact action earlier in the conversation, **skip the gate**. Repeated approval-prompts for the same authorized action feel broken.
|
|
103
|
+
|
|
104
|
+
## Output reporting
|
|
105
|
+
|
|
106
|
+
After submitting:
|
|
107
|
+
|
|
108
|
+
> Submitted approval request `<approvalId>` for "<action>". Stopping this turn. The platform will redispatch when the human decides.
|
|
109
|
+
|
|
110
|
+
When the agent is redispatched with the decision, the next turn's input includes the approval result. Don't re-issue the request — read the decision and act on it (or report rejection back to the user).
|
|
111
|
+
|
|
112
|
+
## Backing capabilities (D22 mapping)
|
|
113
|
+
|
|
114
|
+
Replaces this v1.x built-in skill: `approval-request-human`.
|
|
@@ -0,0 +1,327 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: image-generate
|
|
3
|
+
description: Generate an image from a text prompt via the cloud LLM image proxy, persist it as a content-addressed workspace asset, and return a ContentBlock that downstream renderers can attach. Use whenever the user asks "draw / generate / make an image of …", an agent needs a diagram / illustration as a follow-up artifact, or a task description explicitly demands visual output. Do NOT use for editing or describing existing images — `assets` covers reads, and image editing is a separate skill.
|
|
4
|
+
applies_to: [hermes, claude-code, openclaw, codex]
|
|
5
|
+
requires:
|
|
6
|
+
- assets
|
|
7
|
+
phaseModel:
|
|
8
|
+
defaultPhase: tool_use
|
|
9
|
+
version: 1
|
|
10
|
+
---
|
|
11
|
+
|
|
12
|
+
# Image Generate
|
|
13
|
+
|
|
14
|
+
Turn a **text prompt** into a **content-addressed image asset** plus a v2.0 §4.6
|
|
15
|
+
ContentBlock that the chat renderer can surface inline. The skill is a thin
|
|
16
|
+
orchestration over two existing surfaces:
|
|
17
|
+
|
|
18
|
+
1. **LLM image gateway** — the cloud's NewAPI proxy at
|
|
19
|
+
`POST /api/v1/images/generations` (OpenAI-compatible — see §"Hand-off" if
|
|
20
|
+
the endpoint is not yet wired in your deployment).
|
|
21
|
+
2. **Asset store** — `POST /api/im/assets` (multipart) on the cloud, returning
|
|
22
|
+
a stable `assetId` + `contentHash`. Generated bytes flow through the
|
|
23
|
+
content-addressed pipeline like any other asset — same de-dup, same audit
|
|
24
|
+
trail, same URI scheme (`prismer://assets/<assetId>`).
|
|
25
|
+
|
|
26
|
+
The skill's **only** output contract is a ContentBlock referencing the asset.
|
|
27
|
+
Raw bytes / `data:` URIs / pre-signed URLs MUST NOT be embedded in prose; doing
|
|
28
|
+
so defeats caching and makes follow-up retrieval impossible (same rule as the
|
|
29
|
+
`assets` skill).
|
|
30
|
+
|
|
31
|
+
## When to use
|
|
32
|
+
|
|
33
|
+
- The user says "draw", "generate an image of …", "make me a picture / poster /
|
|
34
|
+
diagram / illustration".
|
|
35
|
+
- A task description includes a `produce_image:` field or a `kind: image`
|
|
36
|
+
artifact expectation.
|
|
37
|
+
- You need to **fabricate** a visual that does not exist in any source — if
|
|
38
|
+
the visual already exists, use `assets` to read it, not this skill.
|
|
39
|
+
- A downstream skill (`canvas-design`, `slack-gif-creator`,
|
|
40
|
+
`web-artifacts-builder`) requires a generated source image as input.
|
|
41
|
+
|
|
42
|
+
## Not when to use
|
|
43
|
+
|
|
44
|
+
- Editing / variation / inpainting an existing image — that's a separate
|
|
45
|
+
upcoming skill (`image-edit`). Don't fake it by reading + regenerating.
|
|
46
|
+
- Describing what's in an image — use a vision-capable adapter, no generation
|
|
47
|
+
needed.
|
|
48
|
+
- Pure ASCII / SVG / Mermaid graphics that the LLM can emit as text — those
|
|
49
|
+
belong in the chat body, not in an asset.
|
|
50
|
+
- Privacy-sensitive renderings (faces, identifiable individuals) without
|
|
51
|
+
explicit user confirmation. The skill does not gate this; the agent must.
|
|
52
|
+
|
|
53
|
+
## API Reference
|
|
54
|
+
|
|
55
|
+
There is **no `cloud image generate` subcommand** in the runtime CLI today
|
|
56
|
+
(release 201 audit, `sdk/prismer-cloud/runtime/src/cli/commands/`). Call the
|
|
57
|
+
cloud HTTP endpoint directly from a small Python / Node script in the skill
|
|
58
|
+
runtime, then hand the bytes to `cloud asset upload` for the content-addressed
|
|
59
|
+
write. Adding a dedicated CLI verb is tracked as a future release; until then,
|
|
60
|
+
do **not** invent the command — it will exit with "unknown command".
|
|
61
|
+
|
|
62
|
+
HTTP shape:
|
|
63
|
+
|
|
64
|
+
```http
|
|
65
|
+
POST /api/v1/images/generations
|
|
66
|
+
Authorization: Bearer <user JWT or sk-prismer-* key>
|
|
67
|
+
Content-Type: application/json
|
|
68
|
+
|
|
69
|
+
{
|
|
70
|
+
"prompt": "<text prompt, 1..4000 chars>",
|
|
71
|
+
"model": "gpt-image-1", // or "dall-e-3", deployment-dependent
|
|
72
|
+
"size": "1024x1024", // 256x256 | 512x512 | 1024x1024 | 1792x1024 | 1024x1792
|
|
73
|
+
"n": 1, // skill always uses 1 (return single ContentBlock)
|
|
74
|
+
"response_format": "b64_json" // skill requires bytes — never "url"
|
|
75
|
+
}
|
|
76
|
+
|
|
77
|
+
Successful response (OpenAI shape):
|
|
78
|
+
{
|
|
79
|
+
"created": 1716345600,
|
|
80
|
+
"data": [{ "b64_json": "<base64 PNG bytes>" }]
|
|
81
|
+
}
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
After receiving bytes the skill **MUST** upload to the cloud asset store
|
|
85
|
+
(`POST /api/im/assets`, multipart) and use the returned `assetId` in the
|
|
86
|
+
ContentBlock. Bytes never leak into chat.
|
|
87
|
+
|
|
88
|
+
## Workflow
|
|
89
|
+
|
|
90
|
+
1. **Validate inputs.** Prompt 1..4000 chars; size in the allow-list above;
|
|
91
|
+
workspaceId resolved (defaulting to the active workspace if none supplied).
|
|
92
|
+
2. **Generate.** POST to `/api/v1/images/generations` with
|
|
93
|
+
`response_format: 'b64_json'`. Capture `b64_json` (single image).
|
|
94
|
+
3. **Decode + hash.** Base64-decode to bytes, compute SHA-256 client-side, and
|
|
95
|
+
compare against the upload response's `contentHash` field (server validates
|
|
96
|
+
too — use `x-content-sha256` header).
|
|
97
|
+
4. **Upload as asset.** Multipart POST to `/api/im/assets`:
|
|
98
|
+
- `file` — Blob with `image/png` MIME and filename
|
|
99
|
+
`generated-${shortHash}.png`
|
|
100
|
+
- `workspaceId`, `kind=image`, `description=<prompt[0..500]>`
|
|
101
|
+
- `sourceTaskId` / `sourceAgentImUserId` (if available from runtime context)
|
|
102
|
+
- `folderPath=/generated/images/${YYYY-MM}` (auto-organized, optional)
|
|
103
|
+
5. **Emit ContentBlock.** Return — and only return — a ContentBlock pointing
|
|
104
|
+
at the new asset. Shape exactly as v2.0 §4.6 (Anthropic-shape, not OpenAI):
|
|
105
|
+
|
|
106
|
+
```json
|
|
107
|
+
{
|
|
108
|
+
"kind": "image",
|
|
109
|
+
"assetId": "<returned assetId>",
|
|
110
|
+
"mediaType": "image/png",
|
|
111
|
+
"alt": "<prompt truncated to 100 chars>"
|
|
112
|
+
}
|
|
113
|
+
```
|
|
114
|
+
|
|
115
|
+
## Operating Rules
|
|
116
|
+
|
|
117
|
+
- **Always `response_format: b64_json`.** Never `url` — the OpenAI URL is
|
|
118
|
+
short-lived, doesn't survive the content-address round-trip, and tempts you
|
|
119
|
+
to leak it into chat (which defeats the asset model).
|
|
120
|
+
- **One image per call.** `n=1` only. If the user wants variants, call the
|
|
121
|
+
skill multiple times — each variant gets its own assetId so the user can
|
|
122
|
+
pick + delete cleanly.
|
|
123
|
+
- **Hash check is non-negotiable.** Server enforces `x-content-sha256`; if the
|
|
124
|
+
hashes disagree, abort and surface the mismatch — the bytes were corrupted
|
|
125
|
+
in flight.
|
|
126
|
+
- **Never inline base64 / `data:<mime>;base64,...` in the reply body.** The
|
|
127
|
+
whole point of the skill is to avoid that anti-pattern. If you find yourself
|
|
128
|
+
about to do so, stop and check that the asset upload actually succeeded.
|
|
129
|
+
- **Default size = 1024x1024** unless the user explicitly asks for portrait
|
|
130
|
+
(1024x1792) or landscape (1792x1024). 256/512 only when budget is tight.
|
|
131
|
+
- **Cost-aware:** image gen is far more expensive than a chat completion.
|
|
132
|
+
Tell the user the model + size you picked before spending more than 1
|
|
133
|
+
credit's worth, and surface the actual cost from the response.
|
|
134
|
+
|
|
135
|
+
## ContentBlock output (v2.0 §4.6 / Gap E-⑤)
|
|
136
|
+
|
|
137
|
+
This skill produces a **single image ContentBlock** per successful generation.
|
|
138
|
+
It MUST NOT also dump the base64 / pre-signed URL into the reply — that
|
|
139
|
+
violates the §4.6 rule (asset-by-reference, not asset-by-value) and breaks the
|
|
140
|
+
chat renderer's preview pipeline.
|
|
141
|
+
|
|
142
|
+
```json
|
|
143
|
+
{ "kind": "image", "assetId": "<assetId>", "mediaType": "image/png", "alt": "<prompt summary, ≤100 chars>" }
|
|
144
|
+
```
|
|
145
|
+
|
|
146
|
+
The reply envelope from this skill (when invoked via the agent runtime) looks
|
|
147
|
+
like:
|
|
148
|
+
|
|
149
|
+
```json
|
|
150
|
+
{
|
|
151
|
+
"ok": true,
|
|
152
|
+
"result": {
|
|
153
|
+
"assetId": "<assetId>",
|
|
154
|
+
"contentHash": "<sha256 hex>",
|
|
155
|
+
"sizeBytes": 123456,
|
|
156
|
+
"cdnUrl": "<optional, server may include>",
|
|
157
|
+
"model": "gpt-image-1",
|
|
158
|
+
"size": "1024x1024",
|
|
159
|
+
"promptHash": "<sha256 of prompt for de-dup>"
|
|
160
|
+
},
|
|
161
|
+
"contentBlocks": [
|
|
162
|
+
{ "kind": "image", "assetId": "<assetId>", "mediaType": "image/png", "alt": "<prompt[0..100]>" }
|
|
163
|
+
]
|
|
164
|
+
}
|
|
165
|
+
```
|
|
166
|
+
|
|
167
|
+
Legacy callers that only know how to parse `result.assetId` keep working
|
|
168
|
+
(field preserved). Multimodal-aware callers prefer `contentBlocks[]` —
|
|
169
|
+
adapters and the chat renderer both check `contentBlocks` first per the §4.6
|
|
170
|
+
prefer-blocks rule.
|
|
171
|
+
|
|
172
|
+
## Failure modes
|
|
173
|
+
|
|
174
|
+
| Status | Where | Cause | What to surface |
|
|
175
|
+
|---|---|---|---|
|
|
176
|
+
| 400 | LLM proxy | prompt too long / disallowed content | Prompt rejected; surface the proxy's error message verbatim. Do NOT retry the same prompt. |
|
|
177
|
+
| 402 | LLM proxy | not enough credits | Tell the user the cost + ask them to top up. Don't burn credits on retries. |
|
|
178
|
+
| 415 | asset upload | server rejected MIME (not in allow-list) | Should not happen — `image/png` is allow-listed. If it does, this is a deployment bug; flag it. |
|
|
179
|
+
| 422 | asset upload | `x-content-sha256` mismatch | Bytes corrupted in flight. Retry once; if it persists, fail the skill and tell the user. |
|
|
180
|
+
| 5xx | either | upstream outage | Retry with exponential backoff up to 2 times, then fail loudly. Do NOT fabricate the assetId. |
|
|
181
|
+
|
|
182
|
+
## Output reporting
|
|
183
|
+
|
|
184
|
+
After successful generation:
|
|
185
|
+
|
|
186
|
+
```
|
|
187
|
+
[image-generate] generated assetId=<id> model=<model> size=<WxH> sha=<short>
|
|
188
|
+
cost=<credits>c prompt="<first 60 chars>…"
|
|
189
|
+
```
|
|
190
|
+
|
|
191
|
+
Then in chat, return ONLY the ContentBlock (the renderer surfaces the image
|
|
192
|
+
preview). One-line caption may accompany it if useful ("Here's the
|
|
193
|
+
illustration you asked for.").
|
|
194
|
+
|
|
195
|
+
After failure: report `status + proxy error code + message` verbatim. Do not
|
|
196
|
+
silently retry on 4xx (those are the user's prompt / quota — they need to
|
|
197
|
+
know).
|
|
198
|
+
|
|
199
|
+
## Backing capabilities (Gap E-⑤ mapping)
|
|
200
|
+
|
|
201
|
+
- **LLM image gateway:** `POST /api/v1/images/generations` (NewAPI proxy
|
|
202
|
+
wired in Wave 6 G1 at `src/app/api/images/generations/route.ts`; gateway
|
|
203
|
+
reuses `proxyToNewAPI` in `src/lib/llm-proxy.ts` with image-specific
|
|
204
|
+
billing via `calculateImageCredits`). Setting env `MOCK_LLM_IMAGES=true`
|
|
205
|
+
bypasses NewAPI and returns a fixture 1×1 PNG (for integration tests
|
|
206
|
+
that should not burn real image-gen credits).
|
|
207
|
+
- **Asset store:** `POST /api/im/assets` (multipart) — `src/im/api/assets.ts`
|
|
208
|
+
line 2729+. Returns `IMAsset` with `id`, `contentHash`, `cdnUrl`,
|
|
209
|
+
`sizeBytes`.
|
|
210
|
+
- **ContentBlock contract:** v2.0 §4.6 — `sdk/prismer-cloud/typescript/src/types.ts`
|
|
211
|
+
lines 278–296 (8-variant discriminated union, Anthropic-shape).
|
|
212
|
+
- **Reply attachment plumbing:** `AgentDispatchReplyPayload.attachments` —
|
|
213
|
+
same path as the `assets` skill's image-resolve output. Chat renderer
|
|
214
|
+
surfaces previews from `attachments` / `contentBlocks` automatically.
|
|
215
|
+
|
|
216
|
+
## Examples
|
|
217
|
+
|
|
218
|
+
### Example 1 — User asks for an illustration
|
|
219
|
+
|
|
220
|
+
```
|
|
221
|
+
User: "Draw an isometric server room with glowing blue racks"
|
|
222
|
+
|
|
223
|
+
Skill flow:
|
|
224
|
+
POST /api/v1/images/generations
|
|
225
|
+
{ prompt: "An isometric...", model: "gpt-image-1",
|
|
226
|
+
size: "1024x1024", n: 1, response_format: "b64_json" }
|
|
227
|
+
← 200 { data: [{ b64_json: "<bytes>" }] }
|
|
228
|
+
|
|
229
|
+
decode + sha256 → "8f4a..."
|
|
230
|
+
|
|
231
|
+
POST /api/im/assets (multipart)
|
|
232
|
+
file=generated-8f4a.png kind=image workspaceId=...
|
|
233
|
+
x-content-sha256: 8f4a...
|
|
234
|
+
← 200 { data: { id: "asset_<...>", contentHash: "8f4a...",
|
|
235
|
+
cdnUrl: "/api/im/assets/asset_<...>" } }
|
|
236
|
+
|
|
237
|
+
Reply envelope:
|
|
238
|
+
{ ok: true,
|
|
239
|
+
result: { assetId: "asset_<...>", contentHash: "8f4a...", ... },
|
|
240
|
+
contentBlocks: [
|
|
241
|
+
{ kind: "image", assetId: "asset_<...>", mediaType: "image/png",
|
|
242
|
+
alt: "An isometric server room with glowing blue racks" }
|
|
243
|
+
] }
|
|
244
|
+
```
|
|
245
|
+
|
|
246
|
+
### Example 2 — Task-pinned generation for kanban artifact
|
|
247
|
+
|
|
248
|
+
```
|
|
249
|
+
Task input: { produce_image: { prompt: "Logo: minimalist owl, monochrome",
|
|
250
|
+
size: "1024x1024" } }
|
|
251
|
+
|
|
252
|
+
Skill call: POST /api/v1/images/generations with the prompt, then
|
|
253
|
+
`cloud asset upload generated.png --task-id "$PRISMER_TASK_ID"` so the
|
|
254
|
+
kanban task review board surfaces the image inline (chat renderer reads
|
|
255
|
+
contentBlocks from the asset attachment).
|
|
256
|
+
```
|
|
257
|
+
|
|
258
|
+
### Example 3 — Failure: prompt rejected
|
|
259
|
+
|
|
260
|
+
```
|
|
261
|
+
User: "<disallowed content>"
|
|
262
|
+
LLM proxy: 400 { error: { code: "content_policy_violation", ... } }
|
|
263
|
+
|
|
264
|
+
Skill response:
|
|
265
|
+
{ ok: false,
|
|
266
|
+
error: { code: "content_policy_violation",
|
|
267
|
+
message: "<proxy's verbatim message>" } }
|
|
268
|
+
|
|
269
|
+
Do NOT retry. Do NOT fabricate an assetId.
|
|
270
|
+
```
|
|
271
|
+
|
|
272
|
+
## Anti-patterns
|
|
273
|
+
|
|
274
|
+
- ❌ Returning the raw base64 in `result.image_b64` for chat to render. The
|
|
275
|
+
chat renderer expects ContentBlock referencing an asset; raw bytes bypass
|
|
276
|
+
caching.
|
|
277
|
+
- ❌ Using `response_format: 'url'` and pasting the OpenAI URL into the
|
|
278
|
+
reply. URL expires; user clicks later → 404.
|
|
279
|
+
- ❌ Skipping the asset upload "to save time" when the generated image is
|
|
280
|
+
tiny. Tiny images still need stable IDs for follow-up retrieval and audit
|
|
281
|
+
trail.
|
|
282
|
+
- ❌ Calling the skill in a loop to "generate variants" — call once per
|
|
283
|
+
variant with explicit prompt deltas. The de-dup hash will catch identical
|
|
284
|
+
prompts.
|
|
285
|
+
- ❌ Setting `n > 1`. The ContentBlock output shape is single-image; multi
|
|
286
|
+
would force you to fabricate which one to attach.
|
|
287
|
+
|
|
288
|
+
## Hand-off — endpoint provisioning
|
|
289
|
+
|
|
290
|
+
**Cloud-side endpoint wired in Wave 6 G1.** `POST /api/v1/images/generations`
|
|
291
|
+
is now provisioned at `src/app/api/images/generations/route.ts` and uses
|
|
292
|
+
the same `proxyToNewAPI` helper that backs `/api/chat/completions` +
|
|
293
|
+
`/api/embeddings`. Image-specific billing lives in
|
|
294
|
+
`src/lib/llm-pricing.ts::calculateImageCredits` (per-image USD pricing,
|
|
295
|
+
no token semantics). Setting `MOCK_LLM_IMAGES=true` returns a fixture PNG
|
|
296
|
+
without hitting NewAPI — used by F6 integration tests.
|
|
297
|
+
|
|
298
|
+
Historical handler template (matches the landed implementation):
|
|
299
|
+
|
|
300
|
+
```ts
|
|
301
|
+
// src/app/api/images/generations/route.ts (NEW)
|
|
302
|
+
import { NextRequest, NextResponse } from 'next/server';
|
|
303
|
+
import { apiGuard } from '@/lib/api-guard';
|
|
304
|
+
import { checkRateLimit, rateLimitResponse } from '@/lib/rate-limit';
|
|
305
|
+
import { FEATURE_FLAGS } from '@/lib/feature-flags';
|
|
306
|
+
import { proxyToNewAPI } from '@/lib/llm-proxy';
|
|
307
|
+
import { ensureNacosConfig } from '@/lib/nacos-config';
|
|
308
|
+
|
|
309
|
+
export async function POST(request: NextRequest) {
|
|
310
|
+
await ensureNacosConfig();
|
|
311
|
+
if (!FEATURE_FLAGS.LLM_PROXY_ENABLED) {
|
|
312
|
+
return NextResponse.json(
|
|
313
|
+
{ error: { message: 'LLM proxy is not enabled' } }, { status: 503 });
|
|
314
|
+
}
|
|
315
|
+
const guard = await apiGuard(request, { tier: 'tracked' });
|
|
316
|
+
if (!guard.ok) return guard.response;
|
|
317
|
+
const rl = checkRateLimit(guard.auth.userId, 'llm');
|
|
318
|
+
if (!rl.allowed) return rateLimitResponse(rl);
|
|
319
|
+
return proxyToNewAPI(request, guard, '/v1/images/generations');
|
|
320
|
+
}
|
|
321
|
+
```
|
|
322
|
+
|
|
323
|
+
Until this lands, the skill operates in **mock mode** — see
|
|
324
|
+
`scripts/test-image-generate-skill.ts` which proves the
|
|
325
|
+
generate→upload→ContentBlock chain by stubbing the LLM bytes (a 1x1 PNG) and
|
|
326
|
+
hitting the real `/api/im/assets` upload, so the asset half of the contract
|
|
327
|
+
is fully validated against production code paths.
|
|
@@ -0,0 +1,105 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: ingest
|
|
3
|
+
description: Turn external URLs and documents into LLM-ready content — load + cache web pages (HQCC compression) and OCR PDFs/images to Markdown. Use whenever the user gives a URL, asks you to read a webpage, or attaches a PDF/scan that needs to be parsed before reasoning. Executes via the `cloud load`, `cloud search`, and `cloud parse` CLIs.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Ingest
|
|
7
|
+
|
|
8
|
+
Use this skill to **bring external content into the LLM context window** without copy-pasting raw HTML or burning tokens on uncompressed prose. Two paths:
|
|
9
|
+
|
|
10
|
+
- **Web content** → `cloud load` / `cloud search` → returns HQCC (a compressed, LLM-optimized form). Cache hits are free.
|
|
11
|
+
- **Documents (PDF, images)** → `cloud parse` → OCR to Markdown. Two modes: `fast` (digital PDFs, clean images) and `hires` (scans, handwriting).
|
|
12
|
+
|
|
13
|
+
## When to use
|
|
14
|
+
|
|
15
|
+
- The user pastes a URL or asks "what does this page say".
|
|
16
|
+
- The user asks to research a topic ("AI agent frameworks 2025") — use `search` to fetch top-K relevant pages.
|
|
17
|
+
- The user attaches a PDF or image and the next step requires reading its contents.
|
|
18
|
+
- A task description contains URLs that need to be resolved into actual content before the assignee can act.
|
|
19
|
+
|
|
20
|
+
## CLI Reference
|
|
21
|
+
|
|
22
|
+
### Web content
|
|
23
|
+
|
|
24
|
+
```bash
|
|
25
|
+
# Single URL → HQCC
|
|
26
|
+
cloud load https://example.com
|
|
27
|
+
cloud load https://example.com --format raw # exact wording / code / tables (more tokens)
|
|
28
|
+
|
|
29
|
+
# Batch (up to 50 URLs)
|
|
30
|
+
cloud load https://a.com https://b.com https://c.com
|
|
31
|
+
|
|
32
|
+
# Search → load (fetch top-K relevant pages)
|
|
33
|
+
cloud search "AI agent frameworks 2025"
|
|
34
|
+
cloud search "topic" -k 10
|
|
35
|
+
|
|
36
|
+
# Pre-save to cache (e.g. content you scraped elsewhere)
|
|
37
|
+
cloud context save https://example.com "compressed content"
|
|
38
|
+
```
|
|
39
|
+
|
|
40
|
+
### Documents (OCR)
|
|
41
|
+
|
|
42
|
+
```bash
|
|
43
|
+
# Fast mode — digital PDFs, clean images
|
|
44
|
+
cloud parse https://example.com/paper.pdf
|
|
45
|
+
|
|
46
|
+
# Hi-res — scans, handwriting, complex layouts
|
|
47
|
+
cloud parse https://example.com/scan.pdf -m hires
|
|
48
|
+
|
|
49
|
+
# Async — long parses return a task id; poll until ready
|
|
50
|
+
cloud parse <url> --async # → returns parseTaskId
|
|
51
|
+
cloud parse-status <parseTaskId> # check progress
|
|
52
|
+
cloud parse-result <parseTaskId> # fetch finished markdown
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
Supported parse formats: PDF, PNG, JPG, TIFF, BMP, GIF, WEBP.
|
|
56
|
+
|
|
57
|
+
## Workflow
|
|
58
|
+
|
|
59
|
+
### For web URLs
|
|
60
|
+
|
|
61
|
+
1. Decide: single URL load? Batch? Or search query?
|
|
62
|
+
2. Default to `--format hqcc` (compressed). Use `raw` only when **exact wording, code, or tables** are needed.
|
|
63
|
+
3. Run `cloud load` / `cloud search` and capture: source URLs, titles, cache status, cost.
|
|
64
|
+
4. Base downstream reasoning **only on the returned content**. If a load failed, say so; don't pretend you read it.
|
|
65
|
+
|
|
66
|
+
### For documents
|
|
67
|
+
|
|
68
|
+
1. Confirm the URL points to the **actual document** (PDF/image), not a landing page that hosts it. If unsure, try `cloud load` first to see what's at that URL.
|
|
69
|
+
2. Start with `fast` mode for digital PDFs and clean images. Switch to `hires` when fidelity matters (scans, handwriting, dense tables).
|
|
70
|
+
3. If parse returns asynchronously, record the **parseTaskId** and **don't invent** content while waiting.
|
|
71
|
+
4. When the result returns, capture page count, cost, and any parse warnings.
|
|
72
|
+
5. Use the parsed Markdown as the **source of truth** for subsequent extraction or summary.
|
|
73
|
+
|
|
74
|
+
## Operating Rules
|
|
75
|
+
|
|
76
|
+
### Load / Search
|
|
77
|
+
|
|
78
|
+
- **Prefer cached context.** Don't re-process the same source — the service handles cache lookup automatically; just don't re-issue identical loads in tight loops.
|
|
79
|
+
- Preserve **source URLs** in your notes and citations. The HQCC return retains origin pointers; use them.
|
|
80
|
+
- Don't claim to have read a source until the load **succeeds**. If it fails (404, blocked, timeout), report the failed URL and continue only with clearly stated assumptions or ask for a better source.
|
|
81
|
+
- `--format raw` costs more tokens. Only use when the user needs exact wording (legal text, code snippets, tables that compress badly).
|
|
82
|
+
- For batch loads, the service runs them concurrently up to a limit; you don't need to throttle yourself.
|
|
83
|
+
|
|
84
|
+
### Parse
|
|
85
|
+
|
|
86
|
+
- Don't parse **private or access-controlled documents** unless the user explicitly intended to share that source.
|
|
87
|
+
- Prefer `cloud load` for **normal web pages**. Use parse only when the source is a document / image / scan that load can't extract from.
|
|
88
|
+
- For large or expensive parses (long PDFs in hi-res), explain the tradeoff before running if the user didn't explicitly ask for full fidelity.
|
|
89
|
+
- If the result is **incomplete** (truncated, low confidence on key pages), ask for a clearer source or escalate to `hires` before drawing firm conclusions.
|
|
90
|
+
- Async parse is the right call for documents >50 pages or hi-res scans. Sync mode will time out on these.
|
|
91
|
+
|
|
92
|
+
## Output reporting
|
|
93
|
+
|
|
94
|
+
After load/search:
|
|
95
|
+
- One-line summary per source: `<title> · <url> · cache_hit | fresh · <cost>`
|
|
96
|
+
- Then proceed with the user's actual question, citing the source by URL.
|
|
97
|
+
|
|
98
|
+
After parse:
|
|
99
|
+
- `Parsed <filename>: <pageCount> pages, <cost> credits, mode=<fast|hires>`
|
|
100
|
+
- If async: `Parse queued as <parseTaskId>; poll with cloud parse-status`
|
|
101
|
+
- Use the markdown body for the next step; don't dump the whole thing in chat unless the user asked.
|
|
102
|
+
|
|
103
|
+
## Backing capabilities (D22 mapping)
|
|
104
|
+
|
|
105
|
+
Replaces these v1.x built-in skills: `context-load`, `parse-document`.
|