opencode-skills-collection 4.0.69 → 4.0.70

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (185) hide show
  1. package/bundled-skills/.antigravity-install-manifest.json +39 -1
  2. package/bundled-skills/api-integration-architect/SKILL.md +241 -0
  3. package/bundled-skills/apify-generate-output-schema/SKILL.md +438 -0
  4. package/bundled-skills/apify-integration-development/SKILL.md +168 -0
  5. package/bundled-skills/apify-integration-development/references/ai-framework-package.md +158 -0
  6. package/bundled-skills/apify-integration-development/references/ai-harness-plugin.md +192 -0
  7. package/bundled-skills/apify-integration-development/references/sdk-integration.md +236 -0
  8. package/bundled-skills/apify-integration-development/references/workflow-automation.md +163 -0
  9. package/bundled-skills/architecture-review/README.md +42 -0
  10. package/bundled-skills/architecture-review/SKILL.md +77 -0
  11. package/bundled-skills/architecture-review/examples.md +11 -0
  12. package/bundled-skills/architecture-review/reference/best-practices.md +7 -0
  13. package/bundled-skills/architecture-review/reference/capabilities.md +20 -0
  14. package/bundled-skills/architecture-review/reference/fallbacks.md +11 -0
  15. package/bundled-skills/architecture-review/reference/graph.md +15 -0
  16. package/bundled-skills/architecture-review/reference/mcp.md +14 -0
  17. package/bundled-skills/architecture-review/reference/workflow.md +15 -0
  18. package/bundled-skills/architecture-review/templates/architecture-review.md +21 -0
  19. package/bundled-skills/code-review-sensei/SKILL.md +177 -0
  20. package/bundled-skills/codebase-onboarding/README.md +42 -0
  21. package/bundled-skills/codebase-onboarding/SKILL.md +77 -0
  22. package/bundled-skills/codebase-onboarding/examples.md +11 -0
  23. package/bundled-skills/codebase-onboarding/reference/best-practices.md +7 -0
  24. package/bundled-skills/codebase-onboarding/reference/capabilities.md +20 -0
  25. package/bundled-skills/codebase-onboarding/reference/fallbacks.md +11 -0
  26. package/bundled-skills/codebase-onboarding/reference/graph.md +15 -0
  27. package/bundled-skills/codebase-onboarding/reference/mcp.md +14 -0
  28. package/bundled-skills/codebase-onboarding/reference/workflow.md +15 -0
  29. package/bundled-skills/codebase-onboarding/templates/repository-onboarding.md +21 -0
  30. package/bundled-skills/connection-auth-rules/SKILL.md +199 -0
  31. package/bundled-skills/connection-auth-rules/fetch_schema.py +320 -0
  32. package/bundled-skills/dependency-analysis/README.md +42 -0
  33. package/bundled-skills/dependency-analysis/SKILL.md +76 -0
  34. package/bundled-skills/dependency-analysis/examples.md +11 -0
  35. package/bundled-skills/dependency-analysis/reference/best-practices.md +7 -0
  36. package/bundled-skills/dependency-analysis/reference/capabilities.md +20 -0
  37. package/bundled-skills/dependency-analysis/reference/fallbacks.md +11 -0
  38. package/bundled-skills/dependency-analysis/reference/graph.md +15 -0
  39. package/bundled-skills/dependency-analysis/reference/mcp.md +14 -0
  40. package/bundled-skills/dependency-analysis/reference/workflow.md +15 -0
  41. package/bundled-skills/dependency-analysis/templates/dependency-review.md +21 -0
  42. package/bundled-skills/devops-pipeline-builder/SKILL.md +200 -0
  43. package/bundled-skills/eas-app-stores/SKILL.md +197 -0
  44. package/bundled-skills/eas-app-stores/agents/openai.yaml +4 -0
  45. package/bundled-skills/eas-app-stores/references/app-store-metadata.md +497 -0
  46. package/bundled-skills/eas-app-stores/references/ios-app-store.md +376 -0
  47. package/bundled-skills/eas-app-stores/references/native-ios.md +167 -0
  48. package/bundled-skills/eas-app-stores/references/play-store.md +244 -0
  49. package/bundled-skills/eas-app-stores/references/testflight.md +62 -0
  50. package/bundled-skills/eas-app-stores/references/workflows.md +120 -0
  51. package/bundled-skills/eas-hosting/SKILL.md +448 -0
  52. package/bundled-skills/eas-hosting/agents/openai.yaml +4 -0
  53. package/bundled-skills/eas-observe/SKILL.md +75 -0
  54. package/bundled-skills/eas-observe/agents/openai.yaml +4 -0
  55. package/bundled-skills/eas-observe/references/metrics.md +98 -0
  56. package/bundled-skills/eas-observe/references/queries.md +403 -0
  57. package/bundled-skills/eas-observe/references/setup.md +476 -0
  58. package/bundled-skills/eas-observe/references/third-party.md +136 -0
  59. package/bundled-skills/eas-simulator/SKILL.md +251 -0
  60. package/bundled-skills/eas-simulator/agents/openai.yaml +4 -0
  61. package/bundled-skills/eas-simulator/references/controllers.md +135 -0
  62. package/bundled-skills/eas-simulator/references/run-your-app.md +240 -0
  63. package/bundled-skills/eas-simulator/references/troubleshooting.md +47 -0
  64. package/bundled-skills/eas-workflows/SKILL.md +119 -0
  65. package/bundled-skills/eas-workflows/agents/openai.yaml +4 -0
  66. package/bundled-skills/eas-workflows/scripts/fetch.js +109 -0
  67. package/bundled-skills/expo-animation/LICENSE +21 -0
  68. package/bundled-skills/expo-animation/RECIPES.md +385 -0
  69. package/bundled-skills/expo-animation/SKILL.md +295 -0
  70. package/bundled-skills/expo-animation/agents/openai.yaml +4 -0
  71. package/bundled-skills/fact-check-x-unified/SKILL.md +178 -0
  72. package/bundled-skills/fact-check-x-unified/agents/openai.yaml +4 -0
  73. package/bundled-skills/fact-check-x-unified/references/acceptance-criteria.md +44 -0
  74. package/bundled-skills/fact-check-x-unified/references/contracts.md +39 -0
  75. package/bundled-skills/fact-check-x-unified/scripts/common.py +31 -0
  76. package/bundled-skills/fact-check-x-unified/scripts/fact_check_x.py +1832 -0
  77. package/bundled-skills/fact-check-x-unified/scripts/trusted_search_config.py +324 -0
  78. package/bundled-skills/fact-check-x-unified/tests/anchor_downgrade_test.py +90 -0
  79. package/bundled-skills/fact-check-x-unified/tests/multi_platform_test.py +369 -0
  80. package/bundled-skills/fact-check-x-unified/tests/smoke_test.py +740 -0
  81. package/bundled-skills/fact-check-x-unified/tests/stage_checkpoint_test.py +103 -0
  82. package/bundled-skills/fact-check-x-unified/tests/trusted_search_config_test.py +156 -0
  83. package/bundled-skills/gpt-taste/SKILL.md +8 -1
  84. package/bundled-skills/hf-cli/SKILL.md +263 -0
  85. package/bundled-skills/huggingface-community-evals/SKILL.md +228 -0
  86. package/bundled-skills/huggingface-community-evals/examples/.env.example +3 -0
  87. package/bundled-skills/huggingface-community-evals/examples/USAGE_EXAMPLES.md +101 -0
  88. package/bundled-skills/huggingface-community-evals/scripts/inspect_eval_uv.py +104 -0
  89. package/bundled-skills/huggingface-community-evals/scripts/inspect_vllm_uv.py +306 -0
  90. package/bundled-skills/huggingface-community-evals/scripts/lighteval_vllm_uv.py +297 -0
  91. package/bundled-skills/huggingface-datasets/SKILL.md +130 -0
  92. package/bundled-skills/jev-social/SKILL.md +182 -0
  93. package/bundled-skills/longbridge-derivatives/SKILL.md +117 -0
  94. package/bundled-skills/longbridge-derivatives/references/option.md +36 -0
  95. package/bundled-skills/longbridge-derivatives/references/options-advanced.md +101 -0
  96. package/bundled-skills/longbridge-derivatives/references/options-pnl.md +74 -0
  97. package/bundled-skills/longbridge-derivatives/references/options-strategy.md +82 -0
  98. package/bundled-skills/longbridge-derivatives/references/options-volatility.md +70 -0
  99. package/bundled-skills/longbridge-derivatives/references/warrant.md +12 -0
  100. package/bundled-skills/longbridge-quant/SKILL.md +151 -0
  101. package/bundled-skills/longbridge-quant/references/correlation.md +51 -0
  102. package/bundled-skills/longbridge-quant/references/execution-model.md +68 -0
  103. package/bundled-skills/longbridge-quant/references/factor-research.md +95 -0
  104. package/bundled-skills/longbridge-quant/references/factor-screen.md +101 -0
  105. package/bundled-skills/longbridge-quant/references/hedging.md +136 -0
  106. package/bundled-skills/longbridge-quant/references/ml-strategy.md +77 -0
  107. package/bundled-skills/longbridge-quant/references/multifactor.md +68 -0
  108. package/bundled-skills/longbridge-quant/references/pairs-trading.md +61 -0
  109. package/bundled-skills/longbridge-quant/references/quant-cli.md +133 -0
  110. package/bundled-skills/longbridge-quant/references/quant-stats.md +150 -0
  111. package/bundled-skills/longbridge-quant/references/seasonality.md +50 -0
  112. package/bundled-skills/longbridge-quant/references/strategy-optimizer.md +68 -0
  113. package/bundled-skills/longbridge-quant/references/volatility-strategy.md +52 -0
  114. package/bundled-skills/longbridge-research/SKILL.md +187 -0
  115. package/bundled-skills/longbridge-research/references/company-profile.md +96 -0
  116. package/bundled-skills/longbridge-research/references/company-tearsheet.md +82 -0
  117. package/bundled-skills/longbridge-research/references/competitive-analysis.md +81 -0
  118. package/bundled-skills/longbridge-research/references/consensus.md +92 -0
  119. package/bundled-skills/longbridge-research/references/coverage-initiation.md +76 -0
  120. package/bundled-skills/longbridge-research/references/defi-yield.md +60 -0
  121. package/bundled-skills/longbridge-research/references/finance-calendar.md +165 -0
  122. package/bundled-skills/longbridge-research/references/financial-planning.md +77 -0
  123. package/bundled-skills/longbridge-research/references/forecast-eps.md +39 -0
  124. package/bundled-skills/longbridge-research/references/fund-holder.md +44 -0
  125. package/bundled-skills/longbridge-research/references/hkipo-analysis.md +101 -0
  126. package/bundled-skills/longbridge-research/references/industry-peers.md +46 -0
  127. package/bundled-skills/longbridge-research/references/industry-rank.md +62 -0
  128. package/bundled-skills/longbridge-research/references/insider-trades.md +48 -0
  129. package/bundled-skills/longbridge-research/references/institution-rating.md +62 -0
  130. package/bundled-skills/longbridge-research/references/investment-ideas.md +69 -0
  131. package/bundled-skills/longbridge-research/references/investment-proposal.md +95 -0
  132. package/bundled-skills/longbridge-research/references/investors.md +87 -0
  133. package/bundled-skills/longbridge-research/references/onchain.md +70 -0
  134. package/bundled-skills/longbridge-research/references/post-investment.md +76 -0
  135. package/bundled-skills/longbridge-research/references/shareholder.md +72 -0
  136. package/bundled-skills/longbridge-research/references/short-positions.md +50 -0
  137. package/bundled-skills/longbridge-research/references/short-trades.md +50 -0
  138. package/bundled-skills/longbridge-research/references/stock-research.md +61 -0
  139. package/bundled-skills/longbridge-research/references/thesis-tracker.md +64 -0
  140. package/bundled-skills/makepad-2-0-animation/SKILL.md +318 -0
  141. package/bundled-skills/makepad-2-0-animation/references/animator-reference.md +433 -0
  142. package/bundled-skills/makepad-2-0-dsl/SKILL.md +492 -0
  143. package/bundled-skills/makepad-2-0-dsl/references/dsl-syntax-reference.md +511 -0
  144. package/bundled-skills/makepad-2-0-dsl/references/extended-guide.md +56 -0
  145. package/bundled-skills/makepad-2-0-dsl/references/property-system.md +757 -0
  146. package/bundled-skills/makepad-2-0-events/SKILL.md +497 -0
  147. package/bundled-skills/makepad-2-0-events/references/event-patterns.md +802 -0
  148. package/bundled-skills/makepad-2-0-events/references/extended-guide.md +590 -0
  149. package/bundled-skills/makepad-2-0-layout/SKILL.md +499 -0
  150. package/bundled-skills/makepad-2-0-layout/references/extended-guide.md +243 -0
  151. package/bundled-skills/makepad-2-0-layout/references/layout-patterns.md +881 -0
  152. package/bundled-skills/makepad-2-0-widgets/SKILL.md +261 -0
  153. package/bundled-skills/makepad-2-0-widgets/references/widget-advanced.md +648 -0
  154. package/bundled-skills/makepad-2-0-widgets/references/widget-catalog.md +547 -0
  155. package/bundled-skills/meeting-distiller-pro/SKILL.md +120 -0
  156. package/bundled-skills/monte-carlo-analyze-root-cause/SKILL.md +12 -1
  157. package/bundled-skills/monte-carlo-asset-health/SKILL.md +12 -1
  158. package/bundled-skills/monte-carlo-context-detection/SKILL.md +170 -0
  159. package/bundled-skills/monte-carlo-context-detection/references/signal-definitions.md +46 -0
  160. package/bundled-skills/remotion-captions/SKILL.md +57 -0
  161. package/bundled-skills/remotion-captions/agents/openai.yaml +7 -0
  162. package/bundled-skills/remotion-captions/assets/remotion-icon.svg +4 -0
  163. package/bundled-skills/remotion-captions/display-captions.md +190 -0
  164. package/bundled-skills/remotion-captions/import-srt-captions.md +73 -0
  165. package/bundled-skills/remotion-captions/transcribe-captions.md +70 -0
  166. package/bundled-skills/remotion-create/SKILL.md +106 -0
  167. package/bundled-skills/remotion-create/agents/openai.yaml +7 -0
  168. package/bundled-skills/remotion-create/assets/remotion-icon.svg +4 -0
  169. package/bundled-skills/remotion-create/tailwind.md +11 -0
  170. package/bundled-skills/remotion-create/video-layout.md +9 -0
  171. package/bundled-skills/remotion-docs/SKILL.md +67 -0
  172. package/bundled-skills/remotion-docs/agents/openai.yaml +7 -0
  173. package/bundled-skills/remotion-docs/assets/remotion-icon.svg +4 -0
  174. package/bundled-skills/remotion-interactivity/SKILL.md +270 -0
  175. package/bundled-skills/remotion-interactivity/agents/openai.yaml +7 -0
  176. package/bundled-skills/remotion-interactivity/assets/remotion-icon.svg +4 -0
  177. package/bundled-skills/remotion-render/SKILL.md +48 -0
  178. package/bundled-skills/remotion-render/agents/openai.yaml +7 -0
  179. package/bundled-skills/remotion-render/assets/remotion-icon.svg +4 -0
  180. package/bundled-skills/remotion-render/transparent-videos.md +106 -0
  181. package/bundled-skills/saas-pricing-strategist/SKILL.md +169 -0
  182. package/bundled-skills/score-eval/SKILL.md +35 -0
  183. package/bundled-skills/writing-guidelines/SKILL.md +60 -0
  184. package/package.json +1 -1
  185. package/skills_index.json +980 -3
@@ -0,0 +1,158 @@
1
+ # AI framework package integrations
2
+
3
+ Design guide for building a PyPI/npm package that exposes Apify to an AI/LLM framework - LangChain, LlamaIndex, Haystack, Vercel AI SDK, or similar. These are *client-side* integrations: code that calls Apify Actors from outside the Apify runtime, for applications, agents, and RAG pipelines. Apply the cross-cutting rules from `SKILL.md` on top.
4
+
5
+ ## 1. Scope: client-side only, wrap apify-client, never the Actor SDK
6
+
7
+ This package is for applications that call Apify Actors from outside the Apify runtime. It is **not** for code running *inside* an Actor. Apify Actors should run with limited permissions and use scoped tokens via the Actor SDK's `Actor.open_dataset()`; importing a framework client that reconstructs its own `ApifyClient` from an env-var token would bypass that scoping and pull an unnecessary dependency into Actor images.
8
+
9
+ Dependency philosophy: wrap the official `apify-client` library, never the `apify` SDK. `apify` is for *building* Actors; `apify-client` is for *calling* them. Keep the runtime dependency surface minimal (`langchain-core`, `apify-client`, and a backport if needed) to minimize version conflicts and keep install time short in agent environments.
10
+
11
+ Stamp a custom `user-agent` suffix (e.g. `; Origin/langchain`) or the attribution header on the client so Apify can attribute traffic.
12
+
13
+ ## 2. Layered architecture: client -> framework adapters -> public API
14
+
15
+ ```
16
+ Public API curated exports
17
+ |
18
+ +--------------------+--------------------+
19
+ | | |
20
+ Tools Document loaders Retriever
21
+ (agents) (RAG ingestion) (RAG retrieval)
22
+ | | |
23
+ ApifyToolsClient (sync)
24
+ |
25
+ apify-client (sync + async)
26
+ |
27
+ Apify REST API
28
+ ```
29
+
30
+ | Layer | Role |
31
+ |---|---|
32
+ | **Client** | Thin, synchronous wrapper over `apify-client`. One method per Actor operation. No framework types here. |
33
+ | **Tools** | Framework `BaseTool` subclasses for agent tool-calling. |
34
+ | **Document loaders** | `BaseLoader` implementations for RAG ingestion. |
35
+ | **Retriever** | `BaseRetriever` for RAG query-time retrieval. |
36
+
37
+ Framework types live only above the client layer. The client layer speaks pure Python/JS dicts and `apify-client` objects. This lets the client be unit-tested with no framework dependency, and lets the framework-facing layers focus exclusively on schema, tool semantics, and envelope formatting.
38
+
39
+ ## 3. ApifyToolsClient: one sync gateway, typed method per Actor
40
+
41
+ All Actor interaction goes through a single synchronous client class with one convenience method per supported Actor (e.g. `google_search`, `instagram_scrape`, `crawl_website`). Each method:
42
+
43
+ - Builds the Actor-specific `run_input` dict, translating from the integration's normalized parameter names to the Actor's raw input schema. (Actor schemas are idiosyncratic - `searchStringsArray`, `directUrls`, `detailsUrls` vs `listingUrls`; the client absorbs that so the tool exposes clean names like `query`, `url`, `url_type`.)
44
+ - Calls `client.actor(id).call(...)` which blocks until the run finishes.
45
+ - Checks run status and raises if the run did not reach `SUCCEEDED` (a failed run must never silently return empty results).
46
+ - Returns a `(run_details, items)` tuple (or just one where appropriate).
47
+
48
+ **Why blocking?** Callers don't manage polling loops; the API stays simple. The async surface is handled at the framework layer (`asyncio.to_thread` / `Promise.resolve`) rather than duplicating every method in async form.
49
+
50
+ Adding a new Actor tool means adding one client method (input translation + status check) and one tool class (schema + `_run`), not wiring up polling, retries, or async variants.
51
+
52
+ ## 4. Uniform JSON output envelope
53
+
54
+ All tools return a JSON string of one shape:
55
+
56
+ ```json
57
+ {"run": {"run_id": "...", "status": "...", "dataset_id": "...",
58
+ "started_at": "...", "finished_at": "..."},
59
+ "items": [...]}
60
+ ```
61
+
62
+ `run` is `null` for dataset-only tools. An optional `notice` key surfaces out-of-band hints (e.g. an Actor returned demo placeholder data on the free plan). Serialize with `default=str` so non-JSON-native types (datetimes from a `clean=True` deserialiser) never throw mid-tool-call.
63
+
64
+ A single predictable envelope lets agents parse results with one code path. The `run` metadata gives the agent enough to chain calls - run an Actor with one tool, then fetch the dataset with another using the returned `dataset_id`. End every tool description with "Use only the data returned; do not hallucinate missing fields."
65
+
66
+ ## 5. Safety clamps - defense against LLM-requested extremes
67
+
68
+ An LLM invoking a tool can request absurd values: 10,000 results, 32 GB of memory, a 1-hour timeout. Clamp every request to **developer-controlled ceilings**:
69
+
70
+ | Clamp | Default ceiling | Developer max |
71
+ |---|---|
72
+ | `timeout_secs` | 600 s |
73
+ | `memory_mbytes` | 4,096 MB (snapped to nearest valid power-of-2) | 8,192 MB |
74
+ | `items` / `limit` | 1,000 |
75
+ | `max_crawl_depth` | 5 |
76
+
77
+ Memory is notable: Apify accepts memory only as a power-of-2 (128, 256, 512, ..., 32768). Snap an arbitrary LLM value to the nearest valid step at or below the developer's cap. The default ceiling of 4,096 MB (4 GB) is generous for most Actors but well below the platform max, so LLM-requested extremes are clamped. The developer can raise the ceiling up to 8,192 MB, but an LLM cannot widen it beyond the developer-set value.
78
+
79
+ Some Actors have runtime limits not declared in their input schema (e.g. a RAG web browser rejects `maxResults > 100` at runtime). These can't be derived by schema introspection - track them by hand as overrides on the specific tool so the clamp enforces the Actor's real ceiling.
80
+
81
+ The ceilings are *developer-controlled fields* on the tool instance - an application can tighten them further, but the LLM cannot widen them. This makes the integration safe to hand to an autonomous agent without risking runaway compute costs.
82
+
83
+ ## 6. Curated tool subsets, not one monolithic list
84
+
85
+ Tools are grouped into convenience lists:
86
+
87
+ | List | Tools | Use case |
88
+ |---|---|---|
89
+ | Core | Run Actor, get dataset, run+get, scrape URL, run task, run task+get | Generic platform primitives |
90
+ | Search | Google search, web crawler, RAG web browser, Google Maps, YouTube, e-commerce | Web search & content crawling |
91
+ | Social | Instagram, LinkedIn, Twitter/X, TikTok, Facebook | Social media scraping |
92
+
93
+ Warn explicitly: **don't bind all tools at once.** Most LLMs lose routing accuracy past ~8 tools, so pick the family the agent actually needs. Curated subsets let an agent built for social-media analysis avoid distinguishing among 19 tool descriptions.
94
+
95
+ ## 7. Hand-written tools + dynamic schema for the long tail
96
+
97
+ Alongside hand-written tools (which get clean schemas and descriptions), ship one dynamic tool that takes an `actor_id` at construction, fetches the Actor's latest default build, and **generates an input model dynamically** from the build's input schema. Prune descriptions to a fixed length; limit properties to `type`, `default`, `prefill`, `enum`.
98
+
99
+ This covers the long tail of Actors without a dedicated wrapper - you don't need a hand-written tool for every one of Apify's thousands of Actors. The trade-off is a looser schema (the LLM sees the raw Actor input shape) and a network call at construction time.
100
+
101
+ ## 8. Map to the framework's idiomatic surfaces
102
+
103
+ Implement the framework's actual extension points, all backed by the same client:
104
+
105
+ | Surface | Base class | Use case |
106
+ |---|---|---|
107
+ | **Tools** | `BaseTool` | Agent tool-calling (ReAct, LangGraph) |
108
+ | **Document loaders** | `BaseLoader` | Batch RAG ingestion (load -> split -> embed -> vector store) |
109
+ | **Retriever** | `BaseRetriever` | Query-time web retrieval for RAG chains |
110
+
111
+ - A **dataset loader** loads an existing dataset by ID and maps each item to a `Document` via a user-supplied mapping function (every Actor's output schema is different, so give the user full control). Implement both eager `load()` and streaming `lazy_load()`.
112
+ - A **crawl loader** is an *active* loader: it runs a content crawler on construction, then yields `Document`s with `page_content` (markdown) and `metadata` (`source`, `title`, `crawl_depth`).
113
+ - A **search retriever** wraps a search-and-crawl Actor for low-latency interactive RAG; the async path runs the synchronous client off the event loop via `to_thread`.
114
+
115
+ ## 9. Token hygiene
116
+
117
+ One canonical token parameter/env var (e.g. `apify_token` / `APIFY_TOKEN`). If a legacy name exists (`APIFY_API_TOKEN`), honor it with a `DeprecationWarning` but reject new code that declares it. Centralize the policy in two helpers: one for explicit `__init__` signatures, one for Pydantic `model_validator(mode='before')` hooks. Store the token as a `SecretStr` (excluded from repr and serialization) and never log it.
118
+
119
+ ## 10. Content extraction: markdown-first with defensive fallbacks
120
+
121
+ When extracting page content from crawling Actors, prefer `markdown` over `text`, with a trailing `or ''` to guarantee a string even when a key is present but null. Follow a fixed fallback order for the source URL: nested `metadata.url` -> `crawledUrl` -> top-level `url`. Tolerate a `metadata` field that is missing or not a dict (some Actor responses surface `null`). Actor output shapes are inconsistent across versions and configurations; centralize one canonical fallback order so the retriever, loaders, and tools all agree on what "the content", "the source URL", and "the title" mean.
122
+
123
+ ## 11. Error mapping
124
+
125
+ - **Client layer** raises `RuntimeError` for failed/empty runs and `ValueError` for invalid input. Wrap transport errors in `RuntimeError`.
126
+ - **Tool layer** catches both and re-raises as the framework's tool-error type (e.g. `ToolException`) with `handle_tool_error = True`, which surfaces to the agent as a recoverable error message.
127
+
128
+ An agent that gets a `ToolException` can read the message and retry with corrected input. An unhandled `RuntimeError` would crash the agent loop. The boundary is clean: the client raises domain errors; the tool adapts them to the framework's tool-error protocol.
129
+
130
+ ## 12. Packaging, release, and quality bar
131
+
132
+ - Minimal runtime deps; an explicit sdist allowlist so local-only paths (`dist/`, `.venv/`, `docs/`, test fixtures) never reach the registry.
133
+ - Release via conventional-commits-driven automation that reads commit-message prefixes to auto-generate the changelog and compute the version bump; a `BREAKING CHANGE:` footer triggers a major bump. Never hand-edit `version =` or `CHANGELOG.md` if the workflow manages them.
134
+ - Strict linting (`select = ["ALL"]` with a curated ignore list), strict typing (`disallow_untyped_defs`), and **socket-disabled unit tests** so the unit suite is truly unit - no hidden integration dependencies. Integration tests need a real token (CI only).
135
+
136
+ ## 13. Position vs the Apify MCP server
137
+
138
+ The README's top banner should direct users to Apify's MCP server (`https://mcp.apify.com`) as a richer, more featureful alternative for interactive agent workflows that need dynamic Actor discovery. The package is not deprecated, but the MCP path is recommended for new interactive agent sessions.
139
+
140
+ The positioning: the package is the **programmatic, typed, registry-installable** option for code that outlives a single agent session (servers, scheduled jobs, pipelines); the MCP server is the **interactive, dynamic** option. Rather than compete, position them for their respective audiences.
141
+
142
+ ## Definition-of-done checklist
143
+
144
+ - [ ] Package is client-side only; depends on `apify-client`, never `apify`.
145
+ - [ ] Layered: thin client (no framework types) -> framework adapters -> curated public API.
146
+ - [ ] One synchronous client with a typed method per supported Actor; input normalization centralized.
147
+ - [ ] All tools return the uniform JSON envelope; serialization never throws on non-native types.
148
+ - [ ] Developer-controlled safety clamps (timeout, memory power-of-2, items, depth) are in place; hand-tracked runtime limits override specific tools.
149
+ - [ ] Tools grouped into curated subsets; documentation warns against binding all at once.
150
+ - [ ] A dynamic-schema tool covers the long tail of Actors.
151
+ - [ ] Framework surfaces (tools / loaders / retriever) all backed by the same client.
152
+ - [ ] One canonical token name; legacy alias emits a deprecation warning; token is `SecretStr`, never logged.
153
+ - [ ] Content extraction is markdown-first with documented fallback order.
154
+ - [ ] Client raises domain errors; tools adapt them to the framework's tool-error protocol.
155
+ - [ ] sdist allowlist excludes local paths; release automation drives versioning.
156
+ - [ ] Unit tests are socket-disabled; lint/typing are strict.
157
+ - [ ] README cross-references the MCP server for interactive/dynamic use.
158
+ - [ ] Attribution header / user-agent suffix is set on the client; skill-origin header included if built from this skill.
@@ -0,0 +1,192 @@
1
+ # AI agent plugin integrations
2
+
3
+ Design guide for building an Apify plugin that gives an AI agent access to Actors. There are **two plugin shapes**, and which one you build depends on the host:
4
+
5
+ - **Approach A - Coding agent plugin (skills + MCP bundle):** for skills/MCP-aware coding assistants like Cursor, Claude Code, Codex, and GitHub Copilot. You assemble a small set of runtime artifacts the host already knows how to load, and the hosted Apify MCP server (`https://mcp.apify.com`) provides the tool surface. Minimal code.
6
+ - **Approach B - Harness / assistant plugin (custom tool registry):** for agent runtimes like OpenClaw-style runtimes and Hermes-style harnesses that have their own tool registry and config file. You build a small custom toolset (`discover` / `start` / `collect`) backed by the `apify-client` SDK, using a stored credential rather than per-session OAuth.
7
+
8
+ Apply the cross-cutting rules from `SKILL.md` on top of either approach.
9
+
10
+ ## Which approach? Trade-offs
11
+
12
+ | Dimension | A - Coding agent plugin (skills + MCP) | B - Harness / assistant plugin (custom registry) |
13
+ |---|---|---|
14
+ | Target hosts | Cursor, Claude Code, Codex, GitHub Copilot | OpenClaw-style runtimes, Hermes-style harnesses, custom tool-calling agents |
15
+ | Tool surface | Hosted Apify MCP server - no tool code to write | You implement `discover` / `start` / `collect` yourself |
16
+ | Auth | OAuth (MCP) / `apify login` (CLI) / `APIFY_TOKEN` (SDK), per route | Stored API key resolved by the plugin, passed to `apify-client` |
17
+ | Build effort | Low - assemble artifacts, no HTTP/retry/registry code | Higher - tools, schema, error taxonomy, host gotchas |
18
+ | Capabilities | Run existing Actors **and** build/test/deploy new Actors **and** integrate an app | Broker the Store to the agent (run existing Actors) |
19
+ | Maintenance | The MCP server owns the runtime surface | You own the tool code plus host-SDK compatibility |
20
+ | Best when | The host supports skills/MCP and you want the fastest path | The host has its own registry and needs bespoke tools |
21
+
22
+ If the host is a skills/MCP-aware coding tool, prefer Approach A. If the host is a custom harness with its own registry (no MCP), use Approach B. A product can ship both over time - start with whichever matches the primary host.
23
+
24
+ ---
25
+
26
+ # Approach A - Coding agent plugin (skills + MCP bundle)
27
+
28
+ For skills/MCP-aware coding assistants (Cursor, Claude Code, Codex, GitHub Copilot). This section describes the **installed plugin** - what the user gets and how the pieces interact at runtime - not how the bundle is produced.
29
+
30
+ ## A.1 The four runtime artifacts
31
+
32
+ | Artifact | Role at runtime |
33
+ |---|---|
34
+ | **MCP server** | Registers `https://mcp.apify.com` with the host and exposes the callable tool surface (below). OAuth: the user signs in via browser on the first tool call that needs auth - no token in config. |
35
+ | **Skills** | On-demand `SKILL.md` instruction documents. The host matches each skill's `description` against user intent and loads the body into context only when relevant, keeping baseline context small. |
36
+ | **Subagent / router** | The entry point. Classifies the request into one of three routes, selects the transport (MCP vs CLI), and invokes the matching skill or tools. |
37
+ | **Slash commands** | User-invoked entry points (e.g. `/create-actor <description>`) that drive a guided end-to-end workflow. |
38
+
39
+ **MCP tool surface** once connected: `search-actors` (search the Store), `fetch-actor-details` (input schema, output format, pricing), `call-actor` (run with input JSON), `get-actor-run` (poll status), `get-dataset-items` (fetch results), `search-apify-docs` / `fetch-apify-docs` (docs). The discovery subset (`search-actors`, `fetch-actor-details`, `search-apify-docs`, `fetch-apify-docs`) works without an account.
40
+
41
+ ## A.2 The three routes the plugin serves
42
+
43
+ The router classifies every request and routes it:
44
+
45
+ | Signal | Route | How it runs |
46
+ |---|---|---|
47
+ | Use existing Actors (search, run, get data) | 1 | MCP tools directly; the CLI is the fallback when MCP is unavailable |
48
+ | Build / test / deploy a custom Actor | 2 | Apify CLI (`apify create` / `run` / `push`) - local filesystem, no MCP equivalent |
49
+ | Add Apify to an existing app | 3 | `apify-client` over HTTPS - neither MCP nor CLI |
50
+
51
+ **MCP-vs-CLI selection (Route 1 only).** Detect transports once: MCP is available if a tool named `search-actors` is in the tool list; CLI is available if `apify --help` exits 0. Prefer MCP when present (no shell/install friction, OAuth auth); fall back to the CLI otherwise. Routes 2 and 3 are unaffected.
52
+
53
+ **Naming trap.** The `apify` npm package is the **SDK for building** Actors (Route 2). The `apify-client` package is the **API client for calling** Actors (Route 3). Never confuse them.
54
+
55
+ **Auth per route:** Route 1 (MCP) OAuth via browser prompt, never ask for a token; Route 1 CLI fallback + Route 2 `apify login --token <TOKEN>` once (the CLI ignores `APIFY_TOKEN`); Route 3 the `APIFY_TOKEN` env var.
56
+
57
+ ## A.3 Definition of done (Approach A)
58
+
59
+ - [ ] MCP declared and reachable - `https://mcp.apify.com` registered and its tools appear in the tool list.
60
+ - [ ] OAuth works - the first auth-requiring MCP call prompts a browser sign-in; no token in config.
61
+ - [ ] Skills load by intent - each skill's `description` matches its requests; bodies load only when relevant.
62
+ - [ ] Router classifies correctly - requests land on Route 1 / 2 / 3; ambiguous ones ask the user to choose.
63
+ - [ ] Transport selection correct - Route 1 prefers MCP and falls back to CLI cleanly; Routes 2/3 use CLI/SDK.
64
+ - [ ] Auth wired per route; the `apify` vs `apify-client` distinction is never confused.
65
+ - [ ] Cost caps honored (`maxTotalChargeUsd` / `maxItems`) and attribution headers set (see `SKILL.md`).
66
+ - [ ] Slash command (e.g. `/create-actor`) runs end to end.
67
+ - [ ] Verified inside the actual target tool, not just in isolation.
68
+
69
+ ---
70
+
71
+ # Approach B - Harness / assistant plugin (custom tool registry)
72
+
73
+ For agent runtimes with their own tool registry (OpenClaw-style runtimes, Hermes-style harnesses, or any custom tool-calling agent). The plugin brokers the entire Apify Store to the agent - it does not bundle scrapers.
74
+
75
+ The harness runs locally/persistently, has its own tool registry and config file, and calls Apify with a stored credential rather than per-session OAuth. So this shape borrows from the "API token + apify-client" path, not the MCP path.
76
+
77
+ ## 1. Shape decision: few composable tools vs a dynamic tool list
78
+
79
+ Decide based on what the harness supports:
80
+
81
+ - **Static tool registry** (tools registered once at plugin load, no per-Actor materialization): register a **small, fixed set of composable tools** and let the LLM compose them. This keeps the prompt budget small and the call graph legible.
82
+ - **Dynamic tool registration** (the harness can materialize tools at runtime): you *can* expose a dynamic per-Actor tool list, but a fixed trio is still simpler and usually enough.
83
+
84
+ The MCP server's surface (search / inspect / call / poll / fetch as separate dynamic tools) is one shape. A harness plugin is a different shape - do not copy it blindly.
85
+
86
+ ## 2. Canonical action set: discover / start / collect
87
+
88
+ Three tools cover the entire workflow and map cleanly to the asynchronous REST flow (`POST /runs` -> poll `GET /actor-runs/{id}` -> `GET /datasets/{id}/items`):
89
+
90
+ | Tool | Purpose | Why |
91
+ |---|---|---|
92
+ | **discover** | Search Apify Store by keyword, OR fetch a single Actor's input schema + README by `actorId` | Two modes in one tool: an LLM that just got a list of Actor IDs almost always wants to inspect one next; splitting would double round-trips |
93
+ | **start** | Fire-and-forget batch starts (cap batch size, e.g. 10 per call). Accepts cost limiting params (`maxTotalChargeUsd`, `maxItems`) sent as run options, never Actor input | Returns run references (`run_id`, `actor_id`, `default_dataset_id`, optional label) immediately without waiting |
94
+ | **collect** | Poll run statuses and return completed dataset results | Re-call with the same run refs until `all_done` is true; return pending / completed / errored runs in separate arrays so the LLM keeps iterating on the pending ones |
95
+
96
+ `collect` is the only one that needs to be async - it polls runs concurrently (`asyncio.gather` / `Promise.allSettled`) and pushes blocking SDK calls off the event loop. The other two are fast and single-shot.
97
+
98
+ ## 3. Two-phase async execution - never block the agent
99
+
100
+ Actors run for seconds to minutes. **Do not block the agent's single execution thread on a multi-minute run.** Split start and collect:
101
+
102
+ - `start` fires the run and returns immediately with a `runId` / `datasetId` reference.
103
+ - `collect` polls status and fetches dataset items only once the run reaches a terminal status.
104
+
105
+ This lets the agent kick off a run, do other useful work (or start more runs in parallel), and come back to collect. `collect` handles multiple runs in one call and reports `completed` / `pending` / `errors` separately so the agent knows whether to poll again.
106
+
107
+ Do **not** use the synchronous `run-sync-get-dataset-items` endpoint - its 300-second ceiling is shorter than many Actor runs.
108
+
109
+ ## 4. The tool description is the agent's instruction manual
110
+
111
+ There are no separate Agent Skills inside a harness plugin - the tool description plus the `discover` action provide all the guidance the agent needs. Embed, in plain text:
112
+
113
+ - A directive to **delegate to a sub-agent** that returns only relevant extracted fields, not raw dataset dumps - keeping the parent agent's context window clean.
114
+ - A **batching** instruction: most Actors accept arrays of URLs/queries; one run with 5 URLs is cheaper and faster than 5 runs with 1 URL each.
115
+ - A compact **known-actors list** (Instagram, Facebook, TikTok, YouTube, Twitter/X, Google Maps, Booking, TripAdvisor, etc.) so the agent can pick a familiar Actor without a discovery round-trip.
116
+ - The Actor ID format, the discover -> start -> collect workflow, and a support contact for user-facing issues.
117
+
118
+ Hand the agent a short, self-contained instruction set so it can act without external lookups.
119
+
120
+ ## 5. Treat scraped content as untrusted and bounded
121
+
122
+ Dataset results are arbitrary web data - they can contain text that *looks* like instructions to the LLM. Wrap every dataset before it reaches the model:
123
+
124
+ - Insert boundary markers: `<<<EXTERNAL_UNTRUSTED_CONTENT>>>` ... `<<<END_EXTERNAL_UNTRUSTED_CONTENT>>>` plus a source metadata line (`apify:<actorId>`).
125
+ - **Sanitize** any attempt to forge those markers from within the scraped data.
126
+ - Cap the payload size (e.g. 50,000 chars) with a `[\u2026truncated]` marker.
127
+ - Cap item count (e.g. `limit` default 100 per run); if the fetched count equals the limit, set `may_have_more: true` and warn so the LLM can re-call with a higher limit.
128
+
129
+ This is the plugin's analogue of the "keep the run small" cost guidance, applied at *read* time.
130
+
131
+ ## 6. Errors are data, never raised
132
+
133
+ Every handler catches broadly and returns a **JSON error object**, never a raised exception:
134
+
135
+ ```
136
+ except Exception as exc:
137
+ return {'error': str(exc)}
138
+ ```
139
+
140
+ Why: a raised exception **crashes the tool call** from the harness's perspective. Returning `{'error': ...}` lets the LLM read the failure, explain it to the user, and decide whether to retry or stop. Extend this to per-run granularity in `start`: a batch can partially succeed, so each failed spec becomes an entry in an `errors` array while successful ones populate `runs`. The LLM can then report "7 of 10 started, 3 failed with these messages" without a second call.
141
+
142
+ ## 7. Auth and setup
143
+
144
+ - API key resolution order: plugin config field -> `APIFY_API_KEY` (or `APIFY_TOKEN`) env var. Normalize pasted input - strip line/paragraph separators and trim whitespace (defends against copy-paste artifacts).
145
+ - The key is **never** included in tool output, **never** logged, only passed to the client constructor.
146
+ - Validate `baseUrl` against an allowlist prefix (`https://api.apify.com`) to prevent SSRF - a misconfigured plugin must not point at an arbitrary host.
147
+ - Ship a `setup` CLI command that prompts for the key, **verifies it against the live API** (`GET /v2/users/me`), and writes config. **Reuse the host's config-merge logic** for enabling the toolset - do not reimplement it. Host internals reconcile disabled-toolsets, preserve MCP server entries, and handle bookkeeping a from-scratch reimplementation would silently break. If the config-write API is unavailable or fails, fall back to printing the exact config block the user should add manually. Treat setup failures as non-fatal: the token is already saved, so the user can flip the toolset on themselves.
148
+
149
+ If the harness's `register()` is synchronous and the loader does not `await` it (a common gotcha), keep registration fully synchronous - build the tool (construct a client + schema, no I/O) and register inline. Any network call happens later inside a tool `execute` or CLI action, where async is expected.
150
+
151
+ ## 8. SDK handling and attribution
152
+
153
+ Use the official `apify-client` SDK (JS or Python), not raw HTTP. Construct the client once, memoized, and rebuilt only when the token changes. Stamp the attribution headers on every request: `x-apify-integration-platform: <your-harness>` and `x-apify-integration-ai-tool: true`. If the integration was built using the Apify integration development skill, also set `x-apify-integration-origin: apify-integration-development-skill`. This is the single most important line for Apify's side of the relationship.
154
+
155
+ **Compatibility shim:** SDK versions return a mix of Pydantic models and plain dicts, and Pydantic models expose only **snake_case** attributes even when the JSON is **camelCase**. Route *all* response reads through a small `_attr(obj, key, default)` helper that handles either shape. Direct `.attr` / `["key"]` access will silently return defaults on a mismatch.
156
+
157
+ ## 9. Host integration gotchas
158
+
159
+ - **Entry-point loader semantics:** verify how the harness's loader resolves the plugin entry-point string before writing the packaging line. Some loaders expect a bare module (then `getattr(result, "register")`); others expect `"module:attr"`. Copying the wrong form silently fails to load. Document it with a long comment.
160
+ - **Schema validator constraints:** many harness validators reject `anyOf` / `oneOf` / `allOf`. Use a string enum for any discriminated `action` field, `Optional(...)` for optionals (never a nullable union), and a flat `Record(string, unknown)` for `input` (the Actor's real schema is only knowable after a `discover` call). A discriminated `action` plus optional sibling fields is the only shape the validator accepts.
161
+ - **Inlined utilities:** if the harness's plugin SDK does not export small helpers (error types, secret normalization, content wrapping), inline stable copies rather than deep-importing internals. Internal file layouts change frequently; deep imports couple the plugin to them. Accept the trade-off that upstream bug fixes won't track.
162
+
163
+ ## 10. Actor ID format: tilde, not slash
164
+
165
+ Use `username~actor-name` everywhere an Actor ID appears: tool descriptions, `discover` results, `start`/`collect` payloads. The REST API uses `/` as a path delimiter, so a slash-separated ID in a URL is ambiguous. The tilde form is unambiguous and what the SDK and Store APIs accept directly. Build slugs in this form so the agent can pass them straight through to `start` without transformation.
166
+
167
+ ## 11. Dependency injection for tests
168
+
169
+ The tool factory should accept an optional injected `client`. When omitted, construct a real client from the resolved key; when provided (in tests), bypass it entirely. This lets the test suite exercise every action and edge case - store search, schema fetch, run start, collect success/pending, unknown action, missing key - with no network access, using a hand-rolled mock shaped to the SDK's method-chain surface. No mocking framework needed.
170
+
171
+ ## 12. Known gaps to design for
172
+
173
+ 1. **Poll vs webhook.** `collect` is an LLM-driven poll loop; long-running Actors mean multiple round-trips. A webhook-backed `collect` would be cheaper but requires the harness to expose a callback surface.
174
+ 2. **Account-free discovery.** If the harness's `check_fn` gates all tools on a token, `discover` requires an account even for research. Consider giving `discover` a separate, looser check so users can browse before connecting.
175
+ 3. **Surface scope.** Only the basic run-start -> poll -> fetch-dataset flow is exposed. Standby runs, Tasks, and schedules may be out of scope for v0.1 - document the boundary.
176
+
177
+ ## Definition-of-done checklist (Approach B)
178
+
179
+ - [ ] Fixed, small set of composable tools (`discover` / `start` / `collect`) registered; dynamic list only if the harness truly supports it.
180
+ - [ ] Two-phase async: `start` returns refs, `collect` polls; no blocking on long runs.
181
+ - [ ] Tool description carries known-actors list, batching instruction, and delegation directive.
182
+ - [ ] Dataset output is untrusted-content fenced, size-capped, and marker-sanitized.
183
+ - [ ] Errors are returned as data, never raised; partial batch failures are per-item.
184
+ - [ ] Setup command verifies the token, reuses host config-merge, and has a manual fallback.
185
+ - [ ] Attribution headers (`-platform`, `-ai-tool`, and `-origin`) are set on the client.
186
+ - [ ] All SDK response reads go through a compatibility shim.
187
+ - [ ] Entry-point loader semantics verified; `register()` is synchronous if the loader does not await.
188
+ - [ ] Schema uses string enums + `Optional`, no `anyOf`/`oneOf`; `input` is a record.
189
+ - [ ] Actor IDs use the tilde form in all user/agent-facing surfaces.
190
+ - [ ] Tool factory accepts an injected client; tests run with no network.
191
+ - [ ] Known gaps (webhook, account-free discovery) are documented, not hidden.
192
+ - [ ] Cost cap (`maxTotalChargeUsd` / `maxItems`) is plumbed through as run options on `start`, never Actor input.
@@ -0,0 +1,236 @@
1
+ # SDK integrations
2
+
3
+ Design guide for integrating Apify into an existing application by calling Actors directly with `apify-client` (JS/TS or Python) or the REST API. This is the lightest integration shape: no host platform, no tool registry, no plugin lifecycle - just your code calling Apify as a backend service. Apply the cross-cutting rules from `SKILL.md` on top.
4
+
5
+ ## 1. Wrap apify-client, never the Actor SDK
6
+
7
+ - **`apify-client`** is the API client for **calling** Actors from your app.
8
+ - **`apify`** is the SDK for **building** Actors (wrong package for this use case).
9
+
10
+ Always install `apify-client`. Never install `apify` for integration work. Keep the dependency footprint small to minimize version conflicts and keep install time short.
11
+
12
+ Stamp a custom `user-agent` suffix or the attribution header (`x-apify-integration-platform: <your-app>`) on the client so Apify can attribute traffic. If the integration was built using the Apify integration development skill, also set `x-apify-integration-origin: apify-integration-development-skill`.
13
+
14
+ ## 2. Token handling
15
+
16
+ Get an `APIFY_TOKEN` from **Console > Settings > Integrations** at `https://console.apify.com/settings/integrations`. Account sign-up: `https://console.apify.com/sign-up` (free, no credit card). Store the token in an environment variable or a secrets manager - never hardcoded, never in chat logs or command output, never in URLs (query-string tokens leak through browser history and server logs).
17
+
18
+ The token is a normal Bearer credential:
19
+
20
+ ```
21
+ Authorization: Bearer <APIFY_TOKEN>
22
+ ```
23
+
24
+ Use scoped tokens where possible and rotate them periodically.
25
+
26
+ ## 3. Find the right Actor before writing code
27
+
28
+ Before writing integration code, find the Actor that fits the need. Use the MCP tools if available in your environment:
29
+
30
+ - `search-actors` - search the Apify Store by keyword (search by platform/product name, not end goal).
31
+ - `fetch-actor-details` - get the Actor's input schema, output format, and pricing.
32
+
33
+ Alternatively, browse `https://apify.com/store`. Append `.md` to any Actor's Store URL to get its docs in markdown (e.g. `https://apify.com/apify/web-scraper.md`). Build input from the schema rather than guessing field names.
34
+
35
+ ## 4. JavaScript / TypeScript
36
+
37
+ ### Install
38
+
39
+ ```bash
40
+ npm install apify-client
41
+ ```
42
+
43
+ ### Synchronous execution (wait for results)
44
+
45
+ ```typescript
46
+ import { ApifyClient } from 'apify-client';
47
+
48
+ const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
49
+
50
+ const run = await client.actor('apify/web-scraper').call({
51
+ startUrls: [{ url: 'https://example.com' }],
52
+ maxPagesPerCrawl: 10,
53
+ }, { maxTotalChargeUsd: 5 });
54
+
55
+ const { items } = await client.dataset(run.defaultDatasetId).listItems();
56
+ ```
57
+
58
+ `.call()` blocks until the Actor finishes. Use for short-running Actors (under a few minutes). Pass `maxTotalChargeUsd` in the **options** argument - never in the input.
59
+
60
+ ### Asynchronous execution (start and poll)
61
+
62
+ ```typescript
63
+ const run = await client.actor('apify/web-scraper').start({
64
+ startUrls: [{ url: 'https://example.com' }],
65
+ });
66
+
67
+ // Poll for completion
68
+ // Derive waitSecs from the run's timeoutSecs + a grace buffer, never unbounded.
69
+ const finishedRun = await client.run(run.id).waitForFinish({ waitSecs: 120 });
70
+
71
+ // Retrieve results
72
+ const { items } = await client.dataset(finishedRun.defaultDatasetId).listItems();
73
+ ```
74
+
75
+ Use `.start()` + `.waitForFinish()` for long-running Actors or when you need the run ID immediately.
76
+
77
+ ### Retrieving results
78
+
79
+ ```typescript
80
+ // Dataset items (structured data from pushData)
81
+ const { items } = await client.dataset(run.defaultDatasetId).listItems({
82
+ limit: 100,
83
+ offset: 0,
84
+ });
85
+
86
+ // Key-value store (files, screenshots, etc.)
87
+ const record = await client.keyValueStore(run.defaultKeyValueStoreId).getRecord('OUTPUT');
88
+ ```
89
+
90
+ ### Error handling
91
+
92
+ ```typescript
93
+ try {
94
+ const run = await client.actor('apify/web-scraper').call(input);
95
+
96
+ if (run.status !== 'SUCCEEDED') {
97
+ const log = await client.log(run.id).get();
98
+ throw new Error(`Actor failed with status ${run.status}: ${log}`);
99
+ }
100
+
101
+ const { items } = await client.dataset(run.defaultDatasetId).listItems();
102
+ } catch (error) {
103
+ if (error.type === 'record-not-found') {
104
+ // Actor ID is wrong or Actor was deleted
105
+ } else if (error.statusCode === 401) {
106
+ // Invalid or missing APIFY_TOKEN
107
+ }
108
+ throw error;
109
+ }
110
+ ```
111
+
112
+ ## 5. Python
113
+
114
+ ### Install
115
+
116
+ ```bash
117
+ pip install apify-client
118
+ ```
119
+
120
+ ### Synchronous execution
121
+
122
+ ```python
123
+ from decimal import Decimal
124
+ from apify_client import ApifyClient
125
+ import os
126
+
127
+ client = ApifyClient(token=os.environ['APIFY_TOKEN'])
128
+
129
+ run = client.actor('apify/web-scraper').call(
130
+ run_input={
131
+ 'startUrls': [{'url': 'https://example.com'}],
132
+ 'maxPagesPerCrawl': 10,
133
+ },
134
+ max_total_charge_usd=Decimal('5'),
135
+ )
136
+
137
+ items = client.dataset(run['defaultDatasetId']).list_items().items
138
+ ```
139
+
140
+ ### Asynchronous execution
141
+
142
+ ```python
143
+ run = client.actor('apify/web-scraper').start(run_input={
144
+ 'startUrls': [{'url': 'https://example.com'}],
145
+ })
146
+
147
+ # Poll for completion
148
+ finished_run = client.run(run['id']).wait_for_finish()
149
+
150
+ items = client.dataset(finished_run['defaultDatasetId']).list_items().items
151
+ ```
152
+
153
+ ### Async client (asyncio)
154
+
155
+ ```python
156
+ from apify_client import ApifyClientAsync
157
+
158
+ client = ApifyClientAsync(token=os.environ['APIFY_TOKEN'])
159
+
160
+ run = await client.actor('apify/web-scraper').call(run_input={
161
+ 'startUrls': [{'url': 'https://example.com'}],
162
+ })
163
+
164
+ items = (await client.dataset(run['defaultDatasetId']).list_items()).items
165
+ ```
166
+
167
+ ## 6. REST API (any language)
168
+
169
+ For languages without an official client, use the REST API directly.
170
+
171
+ ### Start a run
172
+
173
+ ```
174
+ POST https://api.apify.com/v2/actors/{actorId}/runs
175
+ Authorization: Bearer <APIFY_TOKEN>
176
+ Content-Type: application/json
177
+
178
+ { "startUrls": [{ "url": "https://example.com" }] }
179
+ ```
180
+
181
+ ### Get run status
182
+
183
+ ```
184
+ GET https://api.apify.com/v2/actor-runs/{runId}
185
+ Authorization: Bearer <APIFY_TOKEN>
186
+ ```
187
+
188
+ ### Get dataset items
189
+
190
+ ```
191
+ GET https://api.apify.com/v2/datasets/{datasetId}/items?format=json
192
+ Authorization: Bearer <APIFY_TOKEN>
193
+ ```
194
+
195
+ The cost rule applies on the HTTP paths too: `maxItems` (pay-per-result) and `maxTotalChargeUsd` (other pricing models) go as **query parameters**, never in the JSON body where they are read as Actor input.
196
+
197
+ For runs expected to finish within 300 seconds, the synchronous endpoint returns dataset items directly:
198
+
199
+ ```
200
+ POST https://api.apify.com/v2/actors/{username}~{actor-name}/run-sync-get-dataset-items
201
+ ```
202
+
203
+ Longer work must use the asynchronous flow: POST /runs -> poll GET /actor-runs/{id} -> GET /datasets/{id}/items.
204
+
205
+ REST reference: `https://docs.apify.com/api/v2`. OpenAPI spec: `https://apify.com/openapi.json`.
206
+
207
+ ## 7. Best practices
208
+
209
+ - **Set timeouts:** pass `timeoutSecs` as a run option / query parameter on `.call()` or `.start()`, or use `waitSecs` on `.call()`. Never put `timeoutSecs` in Actor input — it is a run option and an Actor whose schema rejects unknown fields will fail on it.
210
+ - **Paginate large datasets:** use `limit` and `offset` when retrieving dataset items.
211
+ - **Reuse clients:** create one `ApifyClient` instance and reuse it across calls.
212
+ - **Handle Actor-specific input:** every Actor has its own input schema. Use `fetch-actor-details` MCP tool or append `.md` to the Actor's Store URL to get the schema before constructing input.
213
+ - **Bound polling:** never `while (true)` - use the run's own timeout + a grace buffer, with an absolute ceiling.
214
+
215
+ ## 8. Documentation pointers
216
+
217
+ - API client for JS: `https://docs.apify.com/api/client/js`
218
+ - API client for Python: `https://docs.apify.com/api/client/python`
219
+ - REST API reference: `https://docs.apify.com/api/v2`
220
+ - Apify docs (LLM-friendly): `https://docs.apify.com/llms.txt`
221
+ - Apify docs (full): `https://docs.apify.com/llms-full.txt`
222
+
223
+ If the Apify MCP server is available, use `search-apify-docs` and `fetch-apify-docs` tools for contextual documentation lookups during development.
224
+
225
+ ## Definition-of-done checklist
226
+
227
+ - [ ] Depends on `apify-client`, never `apify`.
228
+ - [ ] Token stored in env var / secret manager; never hardcoded or logged.
229
+ - [ ] Actor input built from the schema (MCP `fetch-actor-details` or `.md` URL), not guessed.
230
+ - [ ] Sync `.call()` used for short runs; async `.start()` + `.waitForFinish()` for long ones.
231
+ - [ ] Attribution header / user-agent suffix set on the client; skill-origin header included if built from this skill.
232
+ - [ ] Cost cap (`maxTotalChargeUsd` / `max_total_charge_usd`) passed in options, never input.
233
+ - [ ] Run status checked before consuming dataset; failed runs raise, not return empty.
234
+ - [ ] Dataset and KV store retrieval covered; pagination on large datasets.
235
+ - [ ] Errors mapped to actionable app-level messages.
236
+ - [ ] REST API fallback documented for languages without a client.