ottoport 1.5.2 → 1.7.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/README.md +28 -5
- package/cli/ottoport.mjs +54 -4
- package/commands/image.md +4 -1
- package/mcp/server.mjs +49 -7
- package/package.json +1 -1
- package/skills/ottoport/SKILL.md +9 -2
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "ottoport",
|
|
3
3
|
"description": "One API for every model. Call chat, image, video, speech and music models through the OttoPort gateway — as MCP tools, slash commands, or the bundled CLI.",
|
|
4
|
-
"version": "1.
|
|
4
|
+
"version": "1.7.1",
|
|
5
5
|
"author": {
|
|
6
6
|
"name": "LITBOX LLC",
|
|
7
7
|
"email": "support@ottoport.ai"
|
package/README.md
CHANGED
|
@@ -20,12 +20,14 @@ ottoport models [--modality chat|image|video|tts|music] [--json]
|
|
|
20
20
|
ottoport chat "<prompt>" [--model claude-sonnet-5] [--system "..."] [--no-stream]
|
|
21
21
|
[--temperature 0.7] [--max-tokens 512]
|
|
22
22
|
ottoport image "<prompt>" [--model gpt-image-2] [--size 1024x1024 | --resolution 2k]
|
|
23
|
-
[--aspect-ratio 16:9] [--n 1] [--seed 7]
|
|
23
|
+
[--quality high] [--aspect-ratio 16:9] [--n 1] [--seed 7]
|
|
24
24
|
[--image <url>] [--ref <url>[,<url>…]] [--out file.png]
|
|
25
25
|
ottoport image --model qwen-image-layered --image <url> [--layers 4]
|
|
26
26
|
ottoport video "<prompt>" [--model kling-3.0] [--duration 5] [--resolution 1080p]
|
|
27
27
|
[--aspect-ratio 16:9] [--image <url>] [--last-frame <url>]
|
|
28
28
|
[--ref <url>[,<url>…]] [--out clip.mp4]
|
|
29
|
+
[--webhook-url https://…] [--no-wait]
|
|
30
|
+
ottoport webhook status|retry <job-id> # delivery attempts / redeliver now
|
|
29
31
|
ottoport speech "<text>" [--model gpt-4o-mini-tts] [--voice alloy] [--out speech.mp3]
|
|
30
32
|
ottoport music "<prompt>" [--model suno-v5] [--duration 15] [--out song.mp3]
|
|
31
33
|
|
|
@@ -75,7 +77,8 @@ forwarded from the environment rather than written into any config file.
|
|
|
75
77
|
## MCP tools
|
|
76
78
|
|
|
77
79
|
`ottoport_list_models`, `ottoport_chat`, `ottoport_generate_image`,
|
|
78
|
-
`ottoport_generate_video`, `
|
|
80
|
+
`ottoport_generate_video`, `ottoport_video_webhook`, `ottoport_generate_speech`,
|
|
81
|
+
`ottoport_generate_music`.
|
|
79
82
|
The server speaks stdio JSON-RPC and is launched by the host as a subprocess, so
|
|
80
83
|
it needs Node 20+ on the PATH of whatever starts that host.
|
|
81
84
|
|
|
@@ -87,6 +90,11 @@ first-and-last-frame (`image_url` + `last_frame_url`), and reference-to-video
|
|
|
87
90
|
(`reference_image_urls`). Both take a `resolution`: `480p`…`4k` for video,
|
|
88
91
|
`1k`/`2k`/`4k` for images, and a tier above a model's base rate costs more.
|
|
89
92
|
|
|
93
|
+
The GPT Image family reads a `quality` on top of that — `low`, `medium` (the
|
|
94
|
+
default every catalog rate is measured at), `high`, and on `gpt-image-2.5` and
|
|
95
|
+
`gpt-image-2.5-flare` also `xhigh` and `max`. It is how much compute the model
|
|
96
|
+
spends, and the ladder is steep: `max` bills sixteen times `medium`.
|
|
97
|
+
|
|
90
98
|
Two image models take a picture apart instead: `qwen-image-layered` and
|
|
91
99
|
`seedream-5.0-layers` split `image_url` into RGBA layers and return one URL per
|
|
92
100
|
layer, bottom first (the prompt is an optional caption; `num_layers`, 2–10, is
|
|
@@ -99,6 +107,17 @@ reads, pass `model` to that tool or run `ottoport models <model-id>`. Asking
|
|
|
99
107
|
for something a model does not offer is refused with its actual list, before a
|
|
100
108
|
generation is spent.
|
|
101
109
|
|
|
110
|
+
### Webhooks
|
|
111
|
+
|
|
112
|
+
Pass `webhook_url` (MCP) or `--webhook-url` (CLI) with a video to have OttoPort
|
|
113
|
+
POST the finished job to that HTTPS endpoint as `video.completed` or
|
|
114
|
+
`video.failed`. Deliveries are signed:
|
|
115
|
+
`x-ottoport-signature: t=<unix>,v1=<hex HMAC-SHA256(secret, "<t>.<raw body>")>`.
|
|
116
|
+
Failed deliveries are retried after 1, 5, 30 and 120 minutes, 5 attempts in
|
|
117
|
+
all (redirects are not followed); `ottoport webhook retry <job-id>` or `ottoport_video_webhook` sends
|
|
118
|
+
one again at any time. Verification examples are at
|
|
119
|
+
https://ottoport.ai/docs#webhooks.
|
|
120
|
+
|
|
102
121
|
## Configuration
|
|
103
122
|
|
|
104
123
|
- `OTTOPORT_API_KEY` — your `op-…` key. Export it from your shell profile: the
|
|
@@ -108,15 +127,19 @@ generation is spent.
|
|
|
108
127
|
only for local development or a self-hosted gateway.
|
|
109
128
|
- `OTTOPORT_BASE_URL` configures the MCP server only; it does not silently
|
|
110
129
|
redirect CLI traffic.
|
|
130
|
+
- `OTTOPORT_WEBHOOK_SECRET` — the secret video webhooks are signed with, used
|
|
131
|
+
when a webhook URL is given without an explicit secret. Keep it in the
|
|
132
|
+
environment rather than on the command line.
|
|
111
133
|
|
|
112
134
|
## Notes
|
|
113
135
|
|
|
114
136
|
- Image, video, and audio calls return **URLs**, and provider URLs expire.
|
|
115
137
|
Download anything worth keeping.
|
|
116
138
|
- Video generation takes a minute or more and the call blocks until the job is
|
|
117
|
-
terminal
|
|
118
|
-
|
|
119
|
-
|
|
139
|
+
terminal, unless a webhook is given or `--no-wait` is passed. A retry is a
|
|
140
|
+
second billable generation, not a resumption.
|
|
141
|
+
- `resolution` and `quality` are price multipliers, not formatting hints. Leave
|
|
142
|
+
them unset to bill at the model's base tier.
|
|
120
143
|
- Requests are billed against the prepaid balance on your OttoPort account.
|
|
121
144
|
|
|
122
145
|
Docs: <https://ottoport.ai/docs> · Support: support@ottoport.ai
|
package/cli/ottoport.mjs
CHANGED
|
@@ -236,6 +236,7 @@ async function cmdImage(prompt, flags) {
|
|
|
236
236
|
...(last(flags.layers) ? { num_layers: Number(last(flags.layers)) } : {}),
|
|
237
237
|
...(last(flags.size) ? { size: last(flags.size) } : {}),
|
|
238
238
|
...(last(flags.resolution) ? { resolution: last(flags.resolution) } : {}),
|
|
239
|
+
...(last(flags.quality) ? { quality: last(flags.quality) } : {}),
|
|
239
240
|
...(last(flags["aspect-ratio"]) ? { aspect_ratio: flags["aspect-ratio"] } : {}),
|
|
240
241
|
...(last(flags.n) ? { n: Number(last(flags.n)) } : {}),
|
|
241
242
|
...(last(flags.seed) ? { seed: Number(last(flags.seed)) } : {}),
|
|
@@ -257,11 +258,13 @@ async function cmdImage(prompt, flags) {
|
|
|
257
258
|
}
|
|
258
259
|
|
|
259
260
|
async function cmdVideo(prompt, flags) {
|
|
260
|
-
|
|
261
|
+
// A depth model needs only --video; the gateway refuses a prompt there.
|
|
262
|
+
if (!prompt && !last(flags.video)) die("video requires a prompt: ottoport video \"a drone shot\" (or --video <url> on a depth model)");
|
|
261
263
|
const { baseUrl, apiKey } = config(flags);
|
|
262
264
|
const body = {
|
|
263
265
|
model: last(flags.model) || "kling-3.0",
|
|
264
|
-
prompt,
|
|
266
|
+
...(prompt ? { prompt } : {}),
|
|
267
|
+
...(last(flags.video) ? { video_url: last(flags.video) } : {}),
|
|
265
268
|
...(last(flags.duration) ? { duration: Number(last(flags.duration)) } : {}),
|
|
266
269
|
...(last(flags.resolution) ? { resolution: last(flags.resolution) } : {}),
|
|
267
270
|
...(last(flags["aspect-ratio"]) ? { aspect_ratio: flags["aspect-ratio"] } : {}),
|
|
@@ -269,6 +272,7 @@ async function cmdVideo(prompt, flags) {
|
|
|
269
272
|
...(last(flags.image) ? { image_url: flags.image } : {}),
|
|
270
273
|
...(last(flags["last-frame"]) ? { last_frame_url: flags["last-frame"] } : {}),
|
|
271
274
|
...(references(flags) ? { reference_image_urls: references(flags) } : {}),
|
|
275
|
+
...webhook(flags),
|
|
272
276
|
};
|
|
273
277
|
console.error("submitting video job (this can take a minute)…");
|
|
274
278
|
const res = await fetch(`${baseUrl}/api/v1/videos/generations`, {
|
|
@@ -279,7 +283,8 @@ async function cmdVideo(prompt, flags) {
|
|
|
279
283
|
if (!res.ok) die(await readError(res));
|
|
280
284
|
let job = await res.json();
|
|
281
285
|
if (job.status === "failed") die(job.error || "video generation failed", 2);
|
|
282
|
-
if (
|
|
286
|
+
if (body.webhook_url) process.stderr.write(`webhook: ${body.webhook_url} will be notified when the job finishes\n`);
|
|
287
|
+
if (flags.wait !== false) {
|
|
283
288
|
process.stderr.write(`job ${job.id} accepted; waiting for completion…\n`);
|
|
284
289
|
while (job.status === "queued" || job.status === "processing") {
|
|
285
290
|
await new Promise((resolve) => setTimeout(resolve, 3_000));
|
|
@@ -298,6 +303,42 @@ async function cmdVideo(prompt, flags) {
|
|
|
298
303
|
}
|
|
299
304
|
}
|
|
300
305
|
|
|
306
|
+
/**
|
|
307
|
+
* `--webhook-url` registers the job's webhook. The secret comes from
|
|
308
|
+
* `--webhook-secret` or OTTOPORT_WEBHOOK_SECRET; flags are visible in shell
|
|
309
|
+
* history and `ps`, so the environment is the better home for it.
|
|
310
|
+
*/
|
|
311
|
+
function webhook(flags) {
|
|
312
|
+
const url = last(flags["webhook-url"]);
|
|
313
|
+
if (!url) return {};
|
|
314
|
+
const secret = last(flags["webhook-secret"]) || process.env.OTTOPORT_WEBHOOK_SECRET;
|
|
315
|
+
if (!secret) die("--webhook-url needs a signing secret: set OTTOPORT_WEBHOOK_SECRET or pass --webhook-secret");
|
|
316
|
+
return { webhook_url: url, webhook_secret: secret };
|
|
317
|
+
}
|
|
318
|
+
|
|
319
|
+
function describeDelivery(d) {
|
|
320
|
+
const response = d.response_status ? `HTTP ${d.response_status}` : d.attempts ? "no response" : "not sent";
|
|
321
|
+
return `${d.event} ${d.status} ${d.attempts}/${d.max_attempts} attempts ${response}${d.response_body ? ` ${d.response_body.slice(0, 120)}` : ""}`;
|
|
322
|
+
}
|
|
323
|
+
|
|
324
|
+
async function cmdWebhook(action, jobId, flags) {
|
|
325
|
+
if (!["status", "retry"].includes(action) || !jobId) die("usage: ottoport webhook status|retry <job-id> [--json]");
|
|
326
|
+
const { baseUrl, apiKey } = config(flags);
|
|
327
|
+
const path = `${baseUrl}/api/v1/videos/generations/${encodeURIComponent(jobId)}`;
|
|
328
|
+
const res = action === "retry"
|
|
329
|
+
? await fetch(`${path}/webhook`, { method: "POST", headers: headers(apiKey, false) })
|
|
330
|
+
: await fetch(path, { headers: headers(apiKey, false) });
|
|
331
|
+
if (!res.ok) die(await readError(res));
|
|
332
|
+
const body = await res.json();
|
|
333
|
+
if (flags.json) return void console.log(JSON.stringify(body, null, 2));
|
|
334
|
+
const deliveries = action === "retry" ? body.deliveries : body.webhook?.deliveries;
|
|
335
|
+
if (action === "status" && !body.webhook) return void console.log(`job ${jobId} (${body.status}) was submitted without a webhook`);
|
|
336
|
+
if (action === "status") console.log(`job ${jobId} (${body.status}) → ${body.webhook.url}`);
|
|
337
|
+
if (!deliveries?.length) return void console.log("no deliveries yet — one is sent when the job finishes");
|
|
338
|
+
for (const d of deliveries) console.log(describeDelivery(d));
|
|
339
|
+
if (action === "retry" && deliveries[0].status !== "delivered") process.exitCode = 2;
|
|
340
|
+
}
|
|
341
|
+
|
|
301
342
|
async function cmdAudio(kind, prompt, flags) {
|
|
302
343
|
if (!prompt) die(`${kind} requires a prompt: ottoport ${kind} "hello"`);
|
|
303
344
|
const { baseUrl, apiKey } = config(flags);
|
|
@@ -337,7 +378,7 @@ Usage:
|
|
|
337
378
|
ottoport models <model-id> [--json] what that one model accepts
|
|
338
379
|
ottoport chat "<prompt>" [--model claude-sonnet-5] [--system "..."] [--no-stream]
|
|
339
380
|
[--temperature 0.7] [--max-tokens 512]
|
|
340
|
-
ottoport image "<prompt>" [--model gpt-image-2] [--size 1024x1024 | --resolution 2k]
|
|
381
|
+
ottoport image "<prompt>" [--model gpt-image-2] [--size 1024x1024 | --resolution 2k] [--quality high]
|
|
341
382
|
[--aspect-ratio 16:9] [--n 1] [--seed 7]
|
|
342
383
|
[--image <url>] [--ref <url>[,<url>…]] [--out file.png]
|
|
343
384
|
ottoport image --model qwen-image-layered --image <url> [--layers 4]
|
|
@@ -345,7 +386,12 @@ Usage:
|
|
|
345
386
|
ottoport video "<prompt>" [--model kling-3.0] [--duration 5] [--resolution 1080p]
|
|
346
387
|
[--aspect-ratio 16:9] [--seed 7] [--image <url>]
|
|
347
388
|
[--last-frame <url>] [--ref <url>[,<url>…]]
|
|
389
|
+
[--webhook-url https://…] [--webhook-secret <s>]
|
|
348
390
|
[--no-wait] [--out clip.mp4]
|
|
391
|
+
ottoport video --model video-depth-anything --video <url> [--out depth.mp4]
|
|
392
|
+
turn a video (≤60s) into a grayscale depth video
|
|
393
|
+
ottoport webhook status <job-id> [--json] delivery attempts and responses
|
|
394
|
+
ottoport webhook retry <job-id> [--json] redeliver a finished job's webhook now
|
|
349
395
|
ottoport speech "<text>" [--model gpt-4o-mini-tts] [--voice alloy] [--format mp3] [--out speech.mp3]
|
|
350
396
|
ottoport music "<prompt>" [--model suno-v5] [--duration 15] [--format mp3] [--out song.mp3]
|
|
351
397
|
ottoport mcp
|
|
@@ -356,6 +402,7 @@ Usage:
|
|
|
356
402
|
Global flags:
|
|
357
403
|
--url <base> override https://ottoport.ai (local/self-hosted development)
|
|
358
404
|
--key <key> OttoPort API key (env OTTOPORT_API_KEY)
|
|
405
|
+
--webhook-secret <s> signs webhook deliveries (env OTTOPORT_WEBHOOK_SECRET)
|
|
359
406
|
|
|
360
407
|
Examples:
|
|
361
408
|
export OTTOPORT_API_KEY=op-...
|
|
@@ -365,6 +412,8 @@ Examples:
|
|
|
365
412
|
ottoport image "isometric city at dusk" --model gpt-image-2 --out city.png
|
|
366
413
|
ottoport video "cinematic product reveal" --model seedance-2.0-fast --resolution 1080p
|
|
367
414
|
ottoport video "the logo unfolds" --image start.png --last-frame end.png
|
|
415
|
+
ottoport video "drone over fjords" --webhook-url https://example.com/hooks/ottoport --no-wait
|
|
416
|
+
ottoport webhook retry 7c1e9a52-…
|
|
368
417
|
ottoport image "her, on a beach" --ref face.png,style.png
|
|
369
418
|
ottoport image --model seedream-5.0-layers --image https://…/poster.png
|
|
370
419
|
ottoport speech "Welcome to OttoPort" --voice alloy --out welcome.mp3
|
|
@@ -384,6 +433,7 @@ async function main() {
|
|
|
384
433
|
case "chat": return await cmdChat(prompt, flags);
|
|
385
434
|
case "image": return await cmdImage(prompt, flags);
|
|
386
435
|
case "video": return await cmdVideo(prompt, flags);
|
|
436
|
+
case "webhook": return await cmdWebhook(rest[0], rest[1], flags);
|
|
387
437
|
case "speech": return await cmdAudio("speech", prompt, flags);
|
|
388
438
|
case "music": return await cmdAudio("music", prompt, flags);
|
|
389
439
|
case "mcp": return await import(new URL("../mcp/server.mjs", import.meta.url));
|
package/commands/image.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
description: Generate an image through the OttoPort gateway
|
|
3
|
-
argument-hint: "<prompt> [--model <id>] [--resolution 2k] [--ref <url>…]"
|
|
3
|
+
argument-hint: "<prompt> [--model <id>] [--resolution 2k] [--quality high] [--ref <url>…]"
|
|
4
4
|
allowed-tools: mcp__ottoport__ottoport_generate_image, mcp__ottoport__ottoport_list_models, Bash(curl:*)
|
|
5
5
|
---
|
|
6
6
|
|
|
@@ -15,6 +15,9 @@ Request: $ARGUMENTS
|
|
|
15
15
|
- Detail is `resolution` (`1k`/`2k`/`4k`), optionally with `aspect_ratio`; pass
|
|
16
16
|
`size` instead when the user named exact pixels. A higher tier costs more, so
|
|
17
17
|
send one only when the request asks for it.
|
|
18
|
+
- On the GPT Image models, `quality` buys compute rather than pixels — `low`,
|
|
19
|
+
`medium` (the default), `high`, and `xhigh`/`max` on `gpt-image-2.5`. It is
|
|
20
|
+
billed steeply, so send it only when the user asked for a draft or a finish.
|
|
18
21
|
- Models differ in how many references and which resolutions they take. If a
|
|
19
22
|
request needs more than one reference, or 4k, check `ottoport_list_models`
|
|
20
23
|
first — it prints each model's menu — rather than spending a failed call.
|
package/mcp/server.mjs
CHANGED
|
@@ -177,17 +177,25 @@ async function generateImage({ prompt, model, ...rest }) {
|
|
|
177
177
|
return lines.length ? lines.join("\n") : "(no image returned)";
|
|
178
178
|
}
|
|
179
179
|
|
|
180
|
-
async function generateVideo({ prompt, model, ...rest }) {
|
|
181
|
-
|
|
180
|
+
async function generateVideo({ prompt, model, webhook_url, webhook_secret, ...rest }) {
|
|
181
|
+
// A depth model takes only `video_url`; the gateway says which models need a prompt.
|
|
182
|
+
if (!prompt && !rest.video_url) throw new Error("`prompt` is required");
|
|
183
|
+
// The secret is best left in the server's environment, so an agent never has
|
|
184
|
+
// to carry it through a conversation.
|
|
185
|
+
const secret = webhook_secret || process.env.OTTOPORT_WEBHOOK_SECRET;
|
|
186
|
+
if (webhook_url && !secret) throw new Error("`webhook_url` needs a signing secret: set OTTOPORT_WEBHOOK_SECRET for the MCP server or pass `webhook_secret`");
|
|
182
187
|
const res = await fetch(`${BASE}/api/v1/videos/generations`, {
|
|
183
188
|
method: "POST",
|
|
184
189
|
headers: { ...headers(), "idempotency-key": crypto.randomUUID() },
|
|
185
190
|
// Passed through; see `generateImage` for why this layer names nothing.
|
|
186
|
-
body: JSON.stringify({ model: model || "kling-3.0", prompt, ...rest }),
|
|
191
|
+
body: JSON.stringify({ model: model || "kling-3.0", ...(prompt ? { prompt } : {}), ...rest, ...(webhook_url ? { webhook_url, webhook_secret: secret } : {}) }),
|
|
187
192
|
});
|
|
188
193
|
if (!res.ok) throw new Error(await gwError(res));
|
|
189
194
|
let job = await res.json();
|
|
190
195
|
if (job.status === "failed") throw new Error(job.error || "video generation failed");
|
|
196
|
+
// With a webhook the caller has said where the result should go; holding the
|
|
197
|
+
// tool call open for minutes as well would only block the agent.
|
|
198
|
+
if (webhook_url) return `Video job ${job.id} is ${job.status}. ${webhook_url} will receive video.completed or video.failed when it finishes; check deliveries with ottoport_video_webhook.`;
|
|
191
199
|
// The public video endpoint is deliberately asynchronous. MCP tools should
|
|
192
200
|
// still fulfil their promise of returning usable output, so wait for the
|
|
193
201
|
// job rather than returning a queued-job JSON blob to the agent.
|
|
@@ -204,6 +212,18 @@ async function generateVideo({ prompt, model, ...rest }) {
|
|
|
204
212
|
return url || JSON.stringify(job);
|
|
205
213
|
}
|
|
206
214
|
|
|
215
|
+
async function videoWebhook({ job_id, action = "status" }) {
|
|
216
|
+
if (!job_id) throw new Error("`job_id` is required");
|
|
217
|
+
const path = `${BASE}/api/v1/videos/generations/${encodeURIComponent(job_id)}`;
|
|
218
|
+
const res = action === "retry"
|
|
219
|
+
? await fetch(`${path}/webhook`, { method: "POST", headers: headers(false) })
|
|
220
|
+
: await fetch(path, { headers: headers(false) });
|
|
221
|
+
if (!res.ok) throw new Error(await gwError(res));
|
|
222
|
+
const body = await res.json();
|
|
223
|
+
if (action === "retry") return JSON.stringify(body.deliveries ?? [], null, 2);
|
|
224
|
+
return JSON.stringify({ id: body.id, status: body.status, webhook: body.webhook ?? null }, null, 2);
|
|
225
|
+
}
|
|
226
|
+
|
|
207
227
|
async function generateAudio(kind, { prompt, model, voice, duration, format }) {
|
|
208
228
|
if (!prompt) throw new Error("`prompt` is required");
|
|
209
229
|
const res = await fetch(`${BASE}/api/v1/audio/${kind === "tts" ? "speech" : "music"}`, {
|
|
@@ -270,23 +290,27 @@ const TOOLS = [
|
|
|
270
290
|
// at best. The model's own menu is in `ottoport_list_models`.
|
|
271
291
|
resolution: { type: "string", description: 'Output detail — "1k", "2k" or "4k", but only the tiers this model lists in ottoport_list_models. Priced accordingly. Pair with `aspect_ratio` when you know the shape but not the pixel vocabulary; an explicit `size` wins.' },
|
|
272
292
|
aspect_ratio: { type: "string", description: 'e.g. "16:9". Read only when `size` is absent.' },
|
|
293
|
+
quality: { type: "string", description: 'How much compute the model spends, on the GPT Image family only — "low", "medium" (the default), "high", and on gpt-image-2.5 "xhigh" and "max". Priced by tier, and the ladder is steep: `max` costs sixteen times `medium`. Check the model\'s `qualities` in ottoport_list_models.' },
|
|
273
294
|
n: { type: "number" },
|
|
274
295
|
seed: { type: "number" },
|
|
275
296
|
image_url: { type: "string", description: "One reference image, for editing and image-to-image — or, on a layer-decomposition model, the image to split." },
|
|
276
297
|
reference_image_urls: { type: "array", items: { type: "string" }, description: "Several references — a character sheet, a product plus a scene — up to the model's maxInputImages. Call ottoport_list_models to see it." },
|
|
277
298
|
num_layers: { type: "number", description: "Layer models that let you choose (qwen-image-layered, 2–10): how many RGBA layers to split into. Each bills separately. seedream-5.0-layers decides for itself and refuses this." },
|
|
299
|
+
horizontal_angle: { type: "number", description: "Camera models only (qwen-image-angles): orbit in degrees, 0–360. Needs image_url; prompt optional." },
|
|
300
|
+
vertical_angle: { type: "number", description: "Camera models only: elevation in degrees, -30 (below) to 90 (overhead)." },
|
|
301
|
+
zoom: { type: "number", description: "Camera models only: 0 (wide) to 10 (close)." },
|
|
278
302
|
},
|
|
279
303
|
},
|
|
280
304
|
handler: generateImage,
|
|
281
305
|
},
|
|
282
306
|
{
|
|
283
307
|
name: "ottoport_generate_video",
|
|
284
|
-
description: "Generate a video through an OttoPort video model (Veo, Kling, Seedance, …).
|
|
308
|
+
description: "Generate a video through an OttoPort video model (Veo, Kling, Seedance, …). Modes: text-to-video, image-to-video (`image_url`), first-and-last-frame (plus `last_frame_url`), reference-to-video (`reference_image_urls`), and on some models video edit/extend (`video_url`). video-depth-anything instead turns `video_url` into a grayscale depth video and takes no prompt. Which a model accepts is in its capabilities — call ottoport_list_models. Blocks on the queue and returns a video URL, unless `webhook_url` is given.",
|
|
285
309
|
inputSchema: {
|
|
286
310
|
type: "object",
|
|
287
311
|
properties: {
|
|
288
|
-
prompt: { type: "string" },
|
|
289
|
-
model: { type: "string", description: "Model id, e.g. veo-3.1, kling-3.0, seedance-2.0-fast. Default kling-3.0." },
|
|
312
|
+
prompt: { type: "string", description: "Required, except on video-depth-anything, which takes none." },
|
|
313
|
+
model: { type: "string", description: "Model id, e.g. veo-3.1, kling-3.0, seedance-2.0-fast, video-depth-anything. Default kling-3.0." },
|
|
290
314
|
duration: { type: "number", description: "Seconds. The main lever on what this costs." },
|
|
291
315
|
// Per-model, not universal — see the note on the image tool. kling-3.0
|
|
292
316
|
// has no 480p and no 4k; gemini-omni-flash has only 720p.
|
|
@@ -296,11 +320,29 @@ const TOOLS = [
|
|
|
296
320
|
image_url: { type: "string", description: "First frame, for image-to-video." },
|
|
297
321
|
last_frame_url: { type: "string", description: "Final frame. With `image_url`, this is first-and-last-frame generation." },
|
|
298
322
|
reference_image_urls: { type: "array", items: { type: "string" }, description: "Subjects the shot carries through — a character, a product, a style — without being a frame of it." },
|
|
323
|
+
keyframes: { type: "array", items: { type: "object", properties: { image_url: { type: "string" }, time: { type: "number", description: "Seconds from the start; omit to spread evenly." } }, required: ["image_url"] }, description: "Images pinned in time, for keyframes-to-video." },
|
|
324
|
+
video_url: { type: "string", description: "A video to edit or extend, on models listing video-edit / video-extend — or the only input to video-depth-anything. Edits and depth bill per second of this video (at most 60s)." },
|
|
325
|
+
video_task: { type: "string", enum: ["edit", "extend"], description: "edit (default) restyles video_url; extend continues it for `duration` more seconds." },
|
|
326
|
+
webhook_url: { type: "string", description: "Public HTTPS endpoint to POST the finished job to (events video.completed / video.failed, signed with HMAC-SHA256). With it the tool returns the job id at once instead of waiting." },
|
|
327
|
+
webhook_secret: { type: "string", description: "Signing secret for webhook_url. Omit to use the server's OTTOPORT_WEBHOOK_SECRET." },
|
|
299
328
|
},
|
|
300
|
-
required: [
|
|
329
|
+
required: [],
|
|
301
330
|
},
|
|
302
331
|
handler: generateVideo,
|
|
303
332
|
},
|
|
333
|
+
{
|
|
334
|
+
name: "ottoport_video_webhook",
|
|
335
|
+
description: "Inspect or redeliver a video job's webhook. `status` returns the job and every delivery attempt with the endpoint's response; `retry` sends a finished job's webhook again now (resetting its 5 automatic attempts) and returns the result.",
|
|
336
|
+
inputSchema: {
|
|
337
|
+
type: "object",
|
|
338
|
+
properties: {
|
|
339
|
+
job_id: { type: "string", description: "The video job id returned by ottoport_generate_video." },
|
|
340
|
+
action: { type: "string", enum: ["status", "retry"], description: "Default status." },
|
|
341
|
+
},
|
|
342
|
+
required: ["job_id"],
|
|
343
|
+
},
|
|
344
|
+
handler: videoWebhook,
|
|
345
|
+
},
|
|
304
346
|
{
|
|
305
347
|
name: "ottoport_generate_speech",
|
|
306
348
|
description: "Convert text to speech through an OttoPort TTS model. Returns an audio URL or data URL.",
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "ottoport",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.7.1",
|
|
4
4
|
"description": "Claude Code plugin, CLI and MCP server for OttoPort — one OpenAI-compatible API for every LLM, image, video, and speech model.",
|
|
5
5
|
"homepage": "https://ottoport.ai",
|
|
6
6
|
"repository": {
|
package/skills/ottoport/SKILL.md
CHANGED
|
@@ -16,8 +16,9 @@ and do not call the HTTP API, when a tool covers the job:
|
|
|
16
16
|
| --- | --- |
|
|
17
17
|
| `ottoport_list_models` | The catalog: what each model is for, and its rate in credits and USD. Filter with `modality`, or pass `model` for ONE model in full — its modes, its resolution tiers with prices, and every parameter it reads. |
|
|
18
18
|
| `ottoport_chat` | `prompt`, plus optional `model`, `system`, `temperature`, `max_tokens`. |
|
|
19
|
-
| `ottoport_generate_image` | `prompt`, plus optional `model`, `size` or `resolution`, `aspect_ratio`, `n`, `seed`, `image_url` (one reference), `reference_image_urls` (several). Layer decomposition too: `model` `qwen-image-layered` or `seedream-5.0-layers` with `image_url` alone (prompt optional, `num_layers` on Qwen) returns one URL per RGBA layer. |
|
|
20
|
-
| `ottoport_generate_video` | `prompt`, plus optional `model`, `duration`, `resolution`, `aspect_ratio`, `seed`, `image_url` (first frame), `last_frame_url`, `reference_image_urls
|
|
19
|
+
| `ottoport_generate_image` | `prompt`, plus optional `model`, `size` or `resolution`, `quality` (GPT Image only), `aspect_ratio`, `n`, `seed`, `image_url` (one reference), `reference_image_urls` (several). Layer decomposition too: `model` `qwen-image-layered` or `seedream-5.0-layers` with `image_url` alone (prompt optional, `num_layers` on Qwen) returns one URL per RGBA layer. |
|
|
20
|
+
| `ottoport_generate_video` | `prompt`, plus optional `model`, `duration`, `resolution`, `aspect_ratio`, `seed`, `image_url` (first frame), `last_frame_url`, `reference_image_urls`, `webhook_url` (returns the job id at once and POSTs the result there; the secret comes from `webhook_secret` or the server's `OTTOPORT_WEBHOOK_SECRET`). |
|
|
21
|
+
| `ottoport_video_webhook` | `job_id`, plus `action` `status` (every delivery attempt and the endpoint's response) or `retry` (redeliver a finished job's webhook now). |
|
|
21
22
|
| `ottoport_generate_speech` | `prompt` (the text), plus optional `model`, `voice`, `format`. |
|
|
22
23
|
| `ottoport_generate_music` | `prompt`, plus optional `model`, `duration`, `format`. |
|
|
23
24
|
|
|
@@ -54,6 +55,12 @@ video and `"1k"`/`"2k"`/`"4k"` for images, but **no model offers all of it**:
|
|
|
54
55
|
`gpt-image-1.5` only 1k. Same for how many references a model takes —
|
|
55
56
|
`nano-banana-lite` accepts 3, `gpt-image-2` accepts 16.
|
|
56
57
|
|
|
58
|
+
`quality` is a second priced dial, on the GPT Image family alone: `low`,
|
|
59
|
+
`medium` (the default, and the tier every catalog rate was measured at),
|
|
60
|
+
`high`, and `xhigh`/`max` on `gpt-image-2.5` and `gpt-image-2.5-flare`. It buys
|
|
61
|
+
compute, not pixels, and the ladder is steep — `max` bills sixteen times
|
|
62
|
+
`medium`. A model's tiers are its `qualities` in `ottoport_list_models`.
|
|
63
|
+
|
|
57
64
|
So read the model's own menu before naming a mode, a resolution or a second
|
|
58
65
|
reference image: `ottoport_list_models` with `model` set to the one you are
|
|
59
66
|
about to call answers with exactly that — its modes, what each resolution tier
|