@kolbo/mcp 1.93.3 → 1.93.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/package.json +1 -1
- package/skill/GENERATED.md +1 -1
- package/skill/SKILL.md +3 -3
- package/skill/references/models/gpt-image.md +6 -9
- package/skill/references/models/nano-banana.md +4 -7
- package/skill/references/workflows/cost-and-validation.md +1 -1
- package/skill/references/workflows/personal-fonts.md +1 -1
package/package.json
CHANGED
package/skill/GENERATED.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# AUTO-GENERATED — do not edit
|
|
2
2
|
|
|
3
|
-
This tree is mirrored from kolbo-code@
|
|
3
|
+
This tree is mirrored from kolbo-code@a74cf7b, the single source of truth.
|
|
4
4
|
Canonical source: packages/opencode/skills/kolbo/
|
|
5
5
|
Distribution: .github/workflows/sync-skill-to-plugin.yml
|
|
6
6
|
|
package/skill/SKILL.md
CHANGED
|
@@ -339,7 +339,7 @@ Four surfaces show the same job. Use this map — never invent a fifth:
|
|
|
339
339
|
|
|
340
340
|
**🛑 NEVER re-fire a generation you already called.** Aborted / timed-out / `submitted` calls still process server-side. Finish with `get_generation_status` (`wait=true`) — never a second `generate_*`.
|
|
341
341
|
|
|
342
|
-
|
|
342
|
+
**After `submitted` / `_timed_out` — avoid idle work (credit guard).** For a multi-output request, first submit all independent authorized items within the supported concurrency limit. Do not end the task after submitting only the first item or batch. If more requested items are waiting for capacity or output dependencies, use one batched `get_generation_status` call with `wait=true`, then submit the next ready batch. Once every requested item is submitted and no further work needs its output, tell the user it is generating in Library / the cards and end the turn. Submitted is not completed. Do not perform unrelated thinking, file edits, or speculative extra generations while waiting.
|
|
343
343
|
|
|
344
344
|
**Checking status — NEVER poll in a loop.** `get_generation_status` takes `wait=true` (blocks server-side until done, ~3 min) and `generation_ids` (check MANY generations in ONE call — returns `all_done` + which are still running). One `wait=true` call replaces any polling loop: check ALL in-flight ids in ONE call, never one by one, never without `wait`. If it comes back with some still processing, call it ONCE more with `wait=true` and the remaining ids.
|
|
345
345
|
|
|
@@ -390,8 +390,8 @@ After the user approves a bucket, write its `session_id` + plan name into `.kolb
|
|
|
390
390
|
## Rate Limiting & Batch Generation
|
|
391
391
|
|
|
392
392
|
- `generate_image`: 30/min. All other generation tools: 10/min per type. 300/min global. `upload_media`: 300/min, no credit cost.
|
|
393
|
-
- **
|
|
394
|
-
- **
|
|
393
|
+
- **Independent outputs:** emit ready generation calls together. This applies to videos as well as images. Eight requested videos must not become eight sequential waits when parallel submission is supported.
|
|
394
|
+
- **Every batch size:** respect current tool/provider concurrency limits and rate-limit responses. If no more specific limit is available, use the existing conservative batch guidance: images 8, image edits 5, videos 3, video-to-video 3, music/speech/sound 5. Submit each batch concurrently, wait for capacity with a batched status call, then submit the remaining items. Keep pending IDs in the current run; persist only user-approved winners in `.kolbo/production.md`. Never restart submitted jobs to fill a batch.
|
|
395
395
|
|
|
396
396
|
## ⚠️ Multi-output? Default to `generate_creative_director` (CRITICAL)
|
|
397
397
|
|
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# GPT Image 2 / 2.5 — Prompt Rules
|
|
6
6
|
|
|
7
|
-
Load this file when the user wants a **GPT Image 2 or GPT Image 2.5** image (OpenAI).
|
|
7
|
+
Load this file when the user wants a **GPT Image 2 or GPT Image 2.5** image (OpenAI). Discover current variants and native transparency support from list_models; never assume exclusivity or support from this reference. For other image models see `models/nano-banana.md`, `models/creative-director.md`, or `models/prompt-copilot.md`.
|
|
8
8
|
|
|
9
9
|
**Kolbo MCP routing:** call `generate_image` (text-to-image) or `generate_image_edit` (edits with `source_images`). Pass the exact identifier the user named (`gpt-image-2`, `gpt-image-2.5-sunburst`, `gpt-image-2.5-flare`); otherwise consult `list_models({ type: "text_to_img" })`. On `generate_image_edit` the same families appear as their `/edit` rows under `list_models({ type: "image_editing" })`.
|
|
10
10
|
|
|
@@ -26,23 +26,20 @@ Load this file when the user wants a **GPT Image 2 or GPT Image 2.5** image (Ope
|
|
|
26
26
|
- **People, pose, action**: describe scale, body framing, gaze, object interactions ("full body visible, feet included", "looking down at the open book, not at the camera", "hands naturally gripping the handlebar").
|
|
27
27
|
- **Complex scene, an edit that must not drift, or a reusable template**: use the block schema, the reference contract (`identity_lock` + an explicit `preserve` list) and named slots from `workflows/prompt-structure.md`.
|
|
28
28
|
- **Constraints — what changes vs what stays**: state exclusions and invariants explicitly. For edits use **"change only X" + "keep everything else the same"**, and re-state the preserve list on every iteration to prevent drift. Common invariants: identity, geometry, layout, brand elements, camera angle, saturation, contrast, labels, surrounding objects. Always include "no watermark, no extra text, no logos/trademarks" unless the brief specifies otherwise.
|
|
29
|
-
- **Text in images**: put literal text in **quotes** or **ALL CAPS**, specify typography (font style, size, color, placement). For tricky words / brand names, spell letter-by-letter.
|
|
29
|
+
- **Text in images**: put literal text in **quotes** or **ALL CAPS**, specify typography (font style, size, color, placement). For tricky words / brand names, spell letter-by-letter. Use the current catalog quality guidance for small, dense or multi-font text.
|
|
30
30
|
- **Multi-image inputs**: reference each input by number with a short description ("Image 1: product photo… Image 2: style reference…") and describe the interaction ("apply Image 2's style to Image 1", "place the dog from Image 2 next to the woman in Image 1"). Use `@image1` / `@image2` tags — see `workflows/visual-dna.md`.
|
|
31
31
|
- **Iterate, don't overload**: prefer a clean base prompt + single-change follow-ups ("make lighting warmer", "remove the extra tree", "restore the original background") over one giant prompt.
|
|
32
32
|
|
|
33
|
-
##
|
|
33
|
+
## Quality selection
|
|
34
34
|
|
|
35
|
-
-
|
|
36
|
-
- **medium**: default best price/quality for ordinary generations, edits and exploration.
|
|
37
|
-
- **high**: final assets, small/dense text, multi-font layouts, close-up portraits, identity-sensitive edits, infographics, diagrams, posters, UI with labels, scientific visuals, slides with charts/footnotes.
|
|
38
|
-
- **xhigh/max** (GPT Image 2.5 only, when listed): exceptional dense text or difficult multilingual/Hebrew typography after medium/high are insufficient. Do not auto-run retries or raise spending without authorization.
|
|
35
|
+
Read the selected model's current catalog summary, supported qualities and default_quality before recommending a tier. The catalog owns price/quality tradeoffs and exceptional higher-tier use cases. Honor explicit user settings and spending authorization; examples below are craft illustrations, not default settings.
|
|
39
36
|
|
|
40
37
|
## Use Cases (text → image)
|
|
41
38
|
|
|
42
39
|
### Infographics, diagrams, scientific visuals, slides/charts
|
|
43
40
|
- Treat as artifact spec, not illustration request. Name exact deliverable. Define hierarchy. Provide real text/data verbatim in quotes.
|
|
44
41
|
- Demand: readable typography, polished spacing, no decorative clutter, no stock-photo treatment.
|
|
45
|
-
- Recommend
|
|
42
|
+
- Recommend a supported quality from the current catalog and an aspect ratio matching the requested deck/slide output.
|
|
46
43
|
|
|
47
44
|
### Photorealism
|
|
48
45
|
- Prompt as if a real photo is being captured in the moment. Use photography language (lens, lighting, framing). Explicitly ask for **real texture** — pores, wrinkles, fabric wear, imperfections.
|
|
@@ -50,7 +47,7 @@ Load this file when the user wants a **GPT Image 2 or GPT Image 2.5** image (Ope
|
|
|
50
47
|
|
|
51
48
|
### Logos
|
|
52
49
|
- Brand personality + use case + clean, original mark + strong silhouette + balanced negative space + scales from small to large. Flat design, minimal strokes, no gradients unless essential. Plain background, generous padding, centered. "Original, non-infringing".
|
|
53
|
-
- Recommend
|
|
50
|
+
- Recommend quality from the selected catalog row and an aspect ratio matching the brief; request variants only within the authorized count and budget.
|
|
54
51
|
|
|
55
52
|
### Ads / marketing creatives
|
|
56
53
|
- Write like a creative brief: brand, audience, culture, concept, composition, exact copy. Let the model make taste decisions inside boundaries.
|
|
@@ -13,11 +13,8 @@ Load this file when the user wants a **Nano Banana 2 (Gemini 3.1 Flash Image)**
|
|
|
13
13
|
- **Resolution and aspect ratio are MCP-tool params.** **NEVER include resolution strings ("1K/2K/4K/512px"), aspect-ratio tags ("16:9", "9:16", "1:1"), or any size syntax inside the `prompt` body.** Pass them as separate `aspect_ratio` / `resolution` params.
|
|
14
14
|
- Do not write Python / Vertex AI / Gemini SDK code, `generationConfig`, `aspectRatio:`, or any API call syntax. The user is generating through Kolbo's MCP tools.
|
|
15
15
|
|
|
16
|
-
## Model Awareness
|
|
17
|
-
|
|
18
|
-
- **Nano Banana 2 (Gemini 3.1 Flash Image)**: fast, 512px / 1K / 2K / 4K, very wide aspect range incl. 1:4, 4:1, 1:8, 8:1, 21:9, supports real-time web-search grounding. Default for most use cases.
|
|
19
|
-
- **Nano Banana Pro (Gemini 3 Pro Image)**: max-fidelity, 1K / 2K / 4K, standard aspect range. Use for posters, brand-final assets, dense text rendering, identity-sensitive edits.
|
|
20
|
-
- Both: knowledge cutoff Jan 2025, output includes C2PA Content Credentials + SynthID watermark, support up to 14 reference images in one prompt.
|
|
16
|
+
## Model Awareness
|
|
17
|
+
Read the current catalog for the requested family's variants, strengths, quality, resolutions, grounding and reference limits. No variant is a permanent default; do not infer capabilities or rank siblings from this prompt-writing reference.
|
|
21
18
|
|
|
22
19
|
## Best Practices (apply to EVERY prompt)
|
|
23
20
|
|
|
@@ -52,7 +49,7 @@ Instead of describing a fictional scene, instruct the model to retrieve real-wor
|
|
|
52
49
|
**Formula**: `[Source/Search request] + [Analytical task] + [Visual translation]`
|
|
53
50
|
Example shape: `Search for the current weather and date in San Francisco. Analytically, use this data to modify the scene (e.g., if raining, make it look grey and rainy). Visualize this in a miniature city-in-a-cup concept embedded within a realistic, modern smartphone UI.`
|
|
54
51
|
- Use when the user asks for "today's weather", "current price", "live data", "what's playing now", "as of right now", etc.
|
|
55
|
-
-
|
|
52
|
+
- Select a current catalog variant that explicitly supports the requested grounding workflow.
|
|
56
53
|
|
|
57
54
|
### 5. Text rendering & localization (both models excel)
|
|
58
55
|
- **Always quote** literal text: `"Happy Birthday"`, `"URBAN EXPLORER"`, `"10% OFF"`.
|
|
@@ -60,7 +57,7 @@ Example shape: `Search for the current weather and date in San Francisco. Analyt
|
|
|
60
57
|
- **Multilingual**: write the prompt in English and specify the target language for the in-image text ("Then render the same text in Korean and Arabic").
|
|
61
58
|
- **Text-first hack**: when text is the hero, recommend the user first conversationally generate the copy/concepts, THEN ask for the image with that text — better typographic fidelity.
|
|
62
59
|
- Cut-out / negative-space text trick: `bold letters spell "<WORD>", filling the center of the frame. The text acts as a cut-out window. A photograph of <scene> is visible ONLY inside the letterforms.`
|
|
63
|
-
- For small / dense / multi-font text
|
|
60
|
+
- For small / dense / multi-font text, select the variant and quality using current catalog strengths and supported settings.
|
|
64
61
|
|
|
65
62
|
## Prompt Like a Creative Director (the upgrade layer)
|
|
66
63
|
|
|
@@ -104,7 +104,7 @@ Normal cost formula: `final_cost = credit × output_seconds × resolution_multip
|
|
|
104
104
|
- Only fire after they reply.
|
|
105
105
|
2. **No explicit video output resolution**: choose the cheapest supported tier using current catalog pricing and pass it explicitly. This applies to drafts, normal work and final delivery alike. Do not default to 720p/1080p when a cheaper supported tier exists. Fixed-resolution models use their native output.
|
|
106
106
|
3. **Creative intent is not spending authorization**: "finish fully", "cinematic", "professional", "final", "production", "hero" and "don't ask me" do not authorize higher resolution, upscaling or a second high-resolution generation. A budget is a ceiling, not a target. Reference-video resolution and export resolution do not authorize matching generation resolution.
|
|
107
|
-
4. Preserve explicit user-selected settings. Otherwise proceed economically without a resolution approval loop. Inspect missing pricing/capabilities before dispatch. Upgrade only when the user explicitly selects a higher output tier or authorizes the resolution increase; never treat silence as approval. Image quality follows the
|
|
107
|
+
4. Preserve explicit user-selected settings. Otherwise proceed economically without a resolution approval loop. Inspect missing pricing/capabilities before dispatch. Upgrade only when the user explicitly selects a higher output tier or authorizes the resolution increase; never treat silence as approval. Image quality follows the selected model's live catalog summary and default_quality, not a hardcoded family preference or generic final-work maximum.
|
|
108
108
|
5. **Sound on a video model with `sound_credit_multiplier > 1`** → if user didn't ask for sound, leave it off. If user said "with sound" / "with music", enable it.
|
|
109
109
|
|
|
110
110
|
## Defaults When Nothing Is Specified
|
|
@@ -42,7 +42,7 @@ it — nothing installs the font — so these are the levers that decide how clo
|
|
|
42
42
|
|
|
43
43
|
- **Model choice is the biggest one.** GPT Image 2 reproduced the uploaded letterforms
|
|
44
44
|
clearly better than GPT Image 2.5 Sunburst / Flare, which drift toward a default bold
|
|
45
|
-
Hebrew.
|
|
45
|
+
Hebrew in that test. This is historical evidence, not a standing recommendation: choose from current catalog typography strengths and supported font inputs.
|
|
46
46
|
- **Quality does not compensate.** 2K + `high` on GPT Image 2 beat both 2.5 rows at
|
|
47
47
|
`max`. Do not sell a higher tier as a fix for typography.
|
|
48
48
|
- **Weight words in the prompt beat the specimen.** "bold", "medium weight", "very large
|