pi-codex-image-gen 0.1.11 → 0.1.12
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +13 -0
- package/README.md +11 -8
- package/extensions/index.ts +211 -33
- package/package.json +4 -3
- package/skills/imagegen/SKILL.md +18 -15
package/CHANGELOG.md
CHANGED
|
@@ -1,11 +1,24 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.1.12 - 2026-07-17
|
|
4
|
+
|
|
3
5
|
All notable changes to this project will be documented in this file.
|
|
4
6
|
|
|
5
7
|
This project follows the spirit of [Keep a Changelog](https://keepachangelog.com/en/1.1.0/) and uses semantic versioning for releases.
|
|
6
8
|
|
|
7
9
|
## [Unreleased]
|
|
8
10
|
|
|
11
|
+
### Added
|
|
12
|
+
|
|
13
|
+
- Add native edits using up to five local or recent conversation images.
|
|
14
|
+
- Honor bounded `Retry-After` delays and make retry backoff abort-aware.
|
|
15
|
+
|
|
16
|
+
### Fixed
|
|
17
|
+
|
|
18
|
+
- Strictly validate base64 data and output-format signatures before returning or saving images.
|
|
19
|
+
- Preserve valid inline images and report a warning when disk persistence fails.
|
|
20
|
+
- Keep original generated artifacts when copying selected outputs elsewhere by default.
|
|
21
|
+
|
|
9
22
|
## [0.1.11] - 2026-07-01
|
|
10
23
|
|
|
11
24
|
### Changed
|
package/README.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# pi-codex-image-gen
|
|
2
2
|
|
|
3
|
-
Image generation for [Pi](https://pi.dev) using the ChatGPT Images 2.0 model via the OpenAI Codex Responses backend.
|
|
3
|
+
Image generation and editing for [Pi](https://pi.dev) using the ChatGPT Images 2.0 model via the OpenAI Codex Responses backend.
|
|
4
4
|
|
|
5
5
|
## Install
|
|
6
6
|
|
|
@@ -34,7 +34,7 @@ In a Pi session:
|
|
|
34
34
|
> Generate a pixel-art sword icon, 32×32, with a blue blade and gold hilt
|
|
35
35
|
```
|
|
36
36
|
|
|
37
|
-
The agent will invoke `codex_generate_image` with your prompt, stream the response from the Codex backend, and save the resulting image to disk. The `model` parameter controls the Codex routing model; image generation is always performed by **gpt-image-2** on the backend.
|
|
37
|
+
The agent will invoke `codex_generate_image` with your prompt, optionally include up to five local or recent conversation images for editing, stream the response from the Codex backend, and save the resulting image to disk. The `model` parameter controls the Codex routing model; image generation is always performed by **gpt-image-2** on the backend.
|
|
38
38
|
|
|
39
39
|
## Authentication
|
|
40
40
|
|
|
@@ -55,7 +55,7 @@ Create a JSON config file at one (or both) of these locations:
|
|
|
55
55
|
| Global | `~/.pi/agent/extensions/codex-image-gen.json` |
|
|
56
56
|
| Project | `<project-root>/.pi/extensions/codex-image-gen.json` |
|
|
57
57
|
|
|
58
|
-
Project config overrides global config. Example:
|
|
58
|
+
Project config overrides global config only when project trust is active. If project trust is declined or otherwise inactive, project config is ignored and global config still applies. Example:
|
|
59
59
|
|
|
60
60
|
```json
|
|
61
61
|
{
|
|
@@ -89,7 +89,7 @@ Project config overrides global config. Example:
|
|
|
89
89
|
| `none` | Image is returned inline but not written to disk. |
|
|
90
90
|
| `project` | Saves to `<project>/.pi/generated-images/<session-id>/`. |
|
|
91
91
|
| `global` | Saves to `~/.pi/agent/generated-images/<session-id>/`. |
|
|
92
|
-
| `custom` | Saves to a user-specified directory (requires `saveDir` or env). |
|
|
92
|
+
| `custom` | Saves to a user-specified directory (requires `saveDir` or env). `~` and `~/...` expand to the current user's home directory. |
|
|
93
93
|
|
|
94
94
|
## Tool parameters
|
|
95
95
|
|
|
@@ -100,15 +100,18 @@ Project config overrides global config. Example:
|
|
|
100
100
|
| `outputFormat` | string | — | `png` (default), `jpeg`, or `webp`. |
|
|
101
101
|
| `save` | string | — | Override save mode for this call. |
|
|
102
102
|
| `saveDir` | string | — | Directory when `save=custom`. Relative paths resolve under CWD. |
|
|
103
|
+
| `referencedImagePaths` | string[] | — | Up to five local images to edit. Relative paths resolve under CWD. |
|
|
104
|
+
| `numLastImagesToInclude` | integer | — | Include the most recent one to five conversation images for editing. Mutually exclusive with `referencedImagePaths`. |
|
|
103
105
|
|
|
104
106
|
## How it works
|
|
105
107
|
|
|
106
108
|
1. Resolves auth via Pi's `openai-codex` provider (ChatGPT session token).
|
|
107
109
|
2. Sends a Codex Responses API request to the routing model (default `gpt-5.5`) with the `image_generation` tool enabled.
|
|
108
|
-
3.
|
|
109
|
-
4.
|
|
110
|
-
5.
|
|
111
|
-
6.
|
|
110
|
+
3. For edits, attaches the selected local or conversation images to the request.
|
|
111
|
+
4. The backend invokes **gpt-image-2** to generate or edit the image.
|
|
112
|
+
5. Parses the SSE stream and strictly validates the returned base64 and image format.
|
|
113
|
+
6. Saves the image according to the active save mode; persistence failures produce a warning without discarding a valid inline image.
|
|
114
|
+
7. Returns the image data inline plus metadata (model, format, path, revised prompt, usage).
|
|
112
115
|
|
|
113
116
|
## Troubleshooting
|
|
114
117
|
|
package/extensions/index.ts
CHANGED
|
@@ -7,7 +7,8 @@
|
|
|
7
7
|
*/
|
|
8
8
|
|
|
9
9
|
import { readFileSync } from "node:fs";
|
|
10
|
-
import { mkdir, writeFile } from "node:fs/promises";
|
|
10
|
+
import { mkdir, readFile, writeFile } from "node:fs/promises";
|
|
11
|
+
import { homedir } from "node:os";
|
|
11
12
|
import { isAbsolute, join, resolve } from "node:path";
|
|
12
13
|
import { StringEnum } from "@earendil-works/pi-ai";
|
|
13
14
|
import { type ExtensionAPI, getAgentDir, withFileMutationQueue } from "@earendil-works/pi-coding-agent";
|
|
@@ -23,6 +24,8 @@ const DEFAULT_SAVE_MODE = "global";
|
|
|
23
24
|
const OPENAI_BETA_HEADER = "responses=experimental";
|
|
24
25
|
const MAX_RETRIES = 3;
|
|
25
26
|
const BASE_DELAY_MS = 1000;
|
|
27
|
+
const MAX_RETRY_DELAY_MS = 30_000;
|
|
28
|
+
const MAX_EDIT_IMAGES = 5;
|
|
26
29
|
|
|
27
30
|
const SAVE_MODES = ["none", "project", "global", "custom"] as const;
|
|
28
31
|
type SaveMode = (typeof SAVE_MODES)[number];
|
|
@@ -37,9 +40,50 @@ function isRetryableStatus(status: number, errorText: string): boolean {
|
|
|
37
40
|
return /rate.?limit|overloaded|service.?unavailable|upstream.?connect|connection.?refused/i.test(errorText);
|
|
38
41
|
}
|
|
39
42
|
|
|
40
|
-
function
|
|
41
|
-
|
|
42
|
-
|
|
43
|
+
export function parseRetryAfter(value: string | null, nowMs = Date.now()): number | undefined {
|
|
44
|
+
if (!value) return undefined;
|
|
45
|
+
const trimmed = value.trim();
|
|
46
|
+
if (/^\d+(?:\.\d+)?$/.test(trimmed)) {
|
|
47
|
+
const milliseconds = Number(trimmed) * 1000;
|
|
48
|
+
return Number.isFinite(milliseconds) ? Math.min(milliseconds, MAX_RETRY_DELAY_MS) : undefined;
|
|
49
|
+
}
|
|
50
|
+
const dateMs = Date.parse(trimmed);
|
|
51
|
+
if (!Number.isFinite(dateMs) || dateMs <= nowMs) return undefined;
|
|
52
|
+
return Math.min(dateMs - nowMs, MAX_RETRY_DELAY_MS);
|
|
53
|
+
}
|
|
54
|
+
|
|
55
|
+
export function retryDelayMs(
|
|
56
|
+
attempt: number,
|
|
57
|
+
retryAfter: string | null,
|
|
58
|
+
random = Math.random,
|
|
59
|
+
nowMs = Date.now(),
|
|
60
|
+
): number {
|
|
61
|
+
const serverDelay = parseRetryAfter(retryAfter, nowMs);
|
|
62
|
+
if (serverDelay !== undefined) {
|
|
63
|
+
return Math.floor(Math.min(serverDelay * (1 + random() * 0.1), MAX_RETRY_DELAY_MS));
|
|
64
|
+
}
|
|
65
|
+
const exponential = Math.min(BASE_DELAY_MS * 2 ** (attempt - 1), MAX_RETRY_DELAY_MS);
|
|
66
|
+
return Math.floor(exponential * (0.9 + random() * 0.2));
|
|
67
|
+
}
|
|
68
|
+
|
|
69
|
+
export function abortableDelay(milliseconds: number, signal?: AbortSignal): Promise<void> {
|
|
70
|
+
if (signal?.aborted) return Promise.reject(new Error("Image generation was aborted."));
|
|
71
|
+
return new Promise<void>((resolve, reject) => {
|
|
72
|
+
const timer = setTimeout(finish, milliseconds);
|
|
73
|
+
function cleanup() {
|
|
74
|
+
clearTimeout(timer);
|
|
75
|
+
signal?.removeEventListener("abort", abort);
|
|
76
|
+
}
|
|
77
|
+
function finish() {
|
|
78
|
+
cleanup();
|
|
79
|
+
resolve();
|
|
80
|
+
}
|
|
81
|
+
function abort() {
|
|
82
|
+
cleanup();
|
|
83
|
+
reject(new Error("Image generation was aborted."));
|
|
84
|
+
}
|
|
85
|
+
signal?.addEventListener("abort", abort, { once: true });
|
|
86
|
+
});
|
|
43
87
|
}
|
|
44
88
|
|
|
45
89
|
// --- Tool parameter schema ---
|
|
@@ -56,6 +100,19 @@ const TOOL_PARAMS = Type.Object({
|
|
|
56
100
|
description: "Directory to save the image when save=custom. Relative paths resolve under the current workspace.",
|
|
57
101
|
}),
|
|
58
102
|
),
|
|
103
|
+
referencedImagePaths: Type.Optional(
|
|
104
|
+
Type.Array(Type.String(), {
|
|
105
|
+
maxItems: MAX_EDIT_IMAGES,
|
|
106
|
+
description: "Up to five local image paths to edit. Relative paths resolve under the current workspace.",
|
|
107
|
+
}),
|
|
108
|
+
),
|
|
109
|
+
numLastImagesToInclude: Type.Optional(
|
|
110
|
+
Type.Integer({
|
|
111
|
+
minimum: 1,
|
|
112
|
+
maximum: MAX_EDIT_IMAGES,
|
|
113
|
+
description: "Use the most recent one to five images from the current conversation as edit inputs.",
|
|
114
|
+
}),
|
|
115
|
+
),
|
|
59
116
|
});
|
|
60
117
|
|
|
61
118
|
type ToolParams = Static<typeof TOOL_PARAMS>;
|
|
@@ -87,6 +144,11 @@ interface ParsedCodexResponse {
|
|
|
87
144
|
usage?: unknown;
|
|
88
145
|
}
|
|
89
146
|
|
|
147
|
+
interface InputImage {
|
|
148
|
+
data: string;
|
|
149
|
+
mimeType: string;
|
|
150
|
+
}
|
|
151
|
+
|
|
90
152
|
// --- #11: Typed SSE event discriminated union ---
|
|
91
153
|
|
|
92
154
|
type CodexSseEvent =
|
|
@@ -143,15 +205,18 @@ function readConfigFile(path: string): ExtensionConfig {
|
|
|
143
205
|
}
|
|
144
206
|
}
|
|
145
207
|
|
|
146
|
-
function loadConfig(cwd: string): ExtensionConfig {
|
|
147
|
-
const globalConfig = readConfigFile(join(
|
|
208
|
+
export function loadConfig(cwd: string, projectTrusted: boolean, agentDir = getAgentDir()): ExtensionConfig {
|
|
209
|
+
const globalConfig = readConfigFile(join(agentDir, "extensions", "codex-image-gen.json"));
|
|
210
|
+
if (!projectTrusted) return globalConfig;
|
|
148
211
|
const projectConfig = readConfigFile(join(cwd, ".pi", "extensions", "codex-image-gen.json"));
|
|
149
212
|
return { ...globalConfig, ...projectConfig };
|
|
150
213
|
}
|
|
151
214
|
|
|
152
215
|
// --- Path helpers ---
|
|
153
216
|
|
|
154
|
-
function resolveUnderCwd(cwd: string, path: string): string {
|
|
217
|
+
export function resolveUnderCwd(cwd: string, path: string, homeDir = homedir()): string {
|
|
218
|
+
if (path === "~") return homeDir;
|
|
219
|
+
if (path.startsWith("~/")) return resolve(homeDir, path.slice(2));
|
|
155
220
|
return isAbsolute(path) ? path : resolve(cwd, path);
|
|
156
221
|
}
|
|
157
222
|
|
|
@@ -199,22 +264,115 @@ function mimeForFormat(outputFormat: OutputFormat): string {
|
|
|
199
264
|
return outputFormat === "jpeg" ? "image/jpeg" : `image/${outputFormat}`;
|
|
200
265
|
}
|
|
201
266
|
|
|
202
|
-
|
|
267
|
+
function imagePath(outputFormat: OutputFormat, outputDir: string, imageCallId: string): string {
|
|
203
268
|
const filename = `${sanitizePathPart(imageCallId, "image_generation")}.${extensionForFormat(outputFormat)}`;
|
|
204
|
-
|
|
269
|
+
return join(outputDir, filename);
|
|
270
|
+
}
|
|
271
|
+
|
|
272
|
+
export function decodeImageData(base64Data: string, outputFormat: OutputFormat): Buffer {
|
|
273
|
+
const value = base64Data.trim();
|
|
274
|
+
if (!value || value.length % 4 !== 0 || !/^(?:[A-Za-z0-9+/]{4})*(?:[A-Za-z0-9+/]{2}==|[A-Za-z0-9+/]{3}=)?$/.test(value)) {
|
|
275
|
+
throw new Error("Codex returned invalid base64 image data.");
|
|
276
|
+
}
|
|
277
|
+
const bytes = Buffer.from(value, "base64");
|
|
278
|
+
if (bytes.length === 0 || bytes.toString("base64") !== value) {
|
|
279
|
+
throw new Error("Codex returned invalid base64 image data.");
|
|
280
|
+
}
|
|
281
|
+
const validSignature =
|
|
282
|
+
(outputFormat === "png" && bytes.length >= 8 && bytes.subarray(0, 8).equals(Buffer.from([0x89, 0x50, 0x4e, 0x47, 0x0d, 0x0a, 0x1a, 0x0a]))) ||
|
|
283
|
+
(outputFormat === "jpeg" && bytes.length >= 3 && bytes[0] === 0xff && bytes[1] === 0xd8 && bytes[2] === 0xff) ||
|
|
284
|
+
(outputFormat === "webp" && bytes.length >= 12 && bytes.toString("ascii", 0, 4) === "RIFF" && bytes.toString("ascii", 8, 12) === "WEBP");
|
|
285
|
+
if (!validSignature) throw new Error(`Codex returned image data that does not match ${outputFormat}.`);
|
|
286
|
+
return bytes;
|
|
287
|
+
}
|
|
288
|
+
|
|
289
|
+
async function saveImage(
|
|
290
|
+
bytes: Buffer,
|
|
291
|
+
outputFormat: OutputFormat,
|
|
292
|
+
outputDir: string,
|
|
293
|
+
imageCallId: string,
|
|
294
|
+
): Promise<string> {
|
|
295
|
+
const filePath = imagePath(outputFormat, outputDir, imageCallId);
|
|
205
296
|
await withFileMutationQueue(filePath, async () => {
|
|
206
297
|
await mkdir(outputDir, { recursive: true });
|
|
207
|
-
await writeFile(filePath,
|
|
298
|
+
await writeFile(filePath, bytes);
|
|
208
299
|
});
|
|
209
300
|
return filePath;
|
|
210
301
|
}
|
|
211
302
|
|
|
303
|
+
export function selectRecentImages(messages: unknown[], count: number): InputImage[] {
|
|
304
|
+
const images: InputImage[] = [];
|
|
305
|
+
for (let index = messages.length - 1; index >= 0 && images.length < count; index--) {
|
|
306
|
+
const message = messages[index] as { content?: unknown };
|
|
307
|
+
if (!Array.isArray(message?.content)) continue;
|
|
308
|
+
for (let contentIndex = message.content.length - 1; contentIndex >= 0 && images.length < count; contentIndex--) {
|
|
309
|
+
const block = message.content[contentIndex] as { type?: unknown; data?: unknown; mimeType?: unknown };
|
|
310
|
+
if (block?.type === "image" && typeof block.data === "string" && typeof block.mimeType === "string") {
|
|
311
|
+
images.push({ data: block.data, mimeType: block.mimeType });
|
|
312
|
+
}
|
|
313
|
+
}
|
|
314
|
+
}
|
|
315
|
+
return images.reverse();
|
|
316
|
+
}
|
|
317
|
+
|
|
318
|
+
function mimeFromBytes(bytes: Buffer, path: string): string {
|
|
319
|
+
if (bytes.length >= 8 && bytes.subarray(0, 8).equals(Buffer.from([0x89, 0x50, 0x4e, 0x47, 0x0d, 0x0a, 0x1a, 0x0a]))) return "image/png";
|
|
320
|
+
if (bytes.length >= 3 && bytes[0] === 0xff && bytes[1] === 0xd8 && bytes[2] === 0xff) return "image/jpeg";
|
|
321
|
+
if (bytes.length >= 12 && bytes.toString("ascii", 0, 4) === "RIFF" && bytes.toString("ascii", 8, 12) === "WEBP") return "image/webp";
|
|
322
|
+
throw new Error(`Referenced image is unavailable or unsupported: ${path}`);
|
|
323
|
+
}
|
|
324
|
+
|
|
325
|
+
export async function resolveInputImages(
|
|
326
|
+
params: ToolParams,
|
|
327
|
+
cwd: string,
|
|
328
|
+
messages: unknown[],
|
|
329
|
+
): Promise<InputImage[]> {
|
|
330
|
+
const paths = params.referencedImagePaths ?? [];
|
|
331
|
+
if (paths.length > 0 && params.numLastImagesToInclude !== undefined) {
|
|
332
|
+
throw new Error("Provide only one of referencedImagePaths or numLastImagesToInclude.");
|
|
333
|
+
}
|
|
334
|
+
if (paths.length > MAX_EDIT_IMAGES) throw new Error(`referencedImagePaths accepts at most ${MAX_EDIT_IMAGES} paths.`);
|
|
335
|
+
if (paths.length > 0) {
|
|
336
|
+
return Promise.all(
|
|
337
|
+
paths.map(async (path) => {
|
|
338
|
+
const normalized = path.startsWith("@") ? path.slice(1) : path;
|
|
339
|
+
const absolutePath = resolveUnderCwd(cwd, normalized);
|
|
340
|
+
let bytes: Buffer;
|
|
341
|
+
try {
|
|
342
|
+
bytes = await readFile(absolutePath);
|
|
343
|
+
} catch (error) {
|
|
344
|
+
throw new Error(`Unable to read referenced image at ${absolutePath}: ${error instanceof Error ? error.message : String(error)}`);
|
|
345
|
+
}
|
|
346
|
+
return { data: bytes.toString("base64"), mimeType: mimeFromBytes(bytes, absolutePath) };
|
|
347
|
+
}),
|
|
348
|
+
);
|
|
349
|
+
}
|
|
350
|
+
if (params.numLastImagesToInclude !== undefined) {
|
|
351
|
+
const count = params.numLastImagesToInclude;
|
|
352
|
+
if (!Number.isInteger(count) || count < 1 || count > MAX_EDIT_IMAGES) {
|
|
353
|
+
throw new Error(`numLastImagesToInclude must be between 1 and ${MAX_EDIT_IMAGES}.`);
|
|
354
|
+
}
|
|
355
|
+
const images = selectRecentImages(messages, count);
|
|
356
|
+
if (images.length !== count) {
|
|
357
|
+
throw new Error(`Requested the last ${count} conversation images, but only ${images.length} were available.`);
|
|
358
|
+
}
|
|
359
|
+
return images;
|
|
360
|
+
}
|
|
361
|
+
return [];
|
|
362
|
+
}
|
|
363
|
+
|
|
212
364
|
// --- Request building ---
|
|
213
365
|
// #2: prompt_cache_key set to sessionId
|
|
214
366
|
// #7: parallel_tool_calls: false
|
|
215
367
|
// #14: include removed (not needed without reasoning)
|
|
216
368
|
|
|
217
|
-
function buildRequestBody(
|
|
369
|
+
export function buildRequestBody(
|
|
370
|
+
params: ToolParams,
|
|
371
|
+
model: string,
|
|
372
|
+
outputFormat: OutputFormat,
|
|
373
|
+
sessionId: string,
|
|
374
|
+
inputImages: InputImage[] = [],
|
|
375
|
+
) {
|
|
218
376
|
return {
|
|
219
377
|
model,
|
|
220
378
|
store: false,
|
|
@@ -225,7 +383,13 @@ function buildRequestBody(params: ToolParams, model: string, outputFormat: Outpu
|
|
|
225
383
|
input: [
|
|
226
384
|
{
|
|
227
385
|
role: "user",
|
|
228
|
-
content: [
|
|
386
|
+
content: [
|
|
387
|
+
{ type: "input_text", text: params.prompt },
|
|
388
|
+
...inputImages.map((image) => ({
|
|
389
|
+
type: "input_image",
|
|
390
|
+
image_url: `data:${image.mimeType};base64,${image.data}`,
|
|
391
|
+
})),
|
|
392
|
+
],
|
|
229
393
|
},
|
|
230
394
|
],
|
|
231
395
|
tools: [{ type: "image_generation", output_format: outputFormat }],
|
|
@@ -346,9 +510,10 @@ async function requestImage(
|
|
|
346
510
|
model: string,
|
|
347
511
|
outputFormat: OutputFormat,
|
|
348
512
|
sessionId: string,
|
|
513
|
+
inputImages: InputImage[],
|
|
349
514
|
signal?: AbortSignal,
|
|
350
515
|
): Promise<ParsedCodexResponse> {
|
|
351
|
-
const body = JSON.stringify(buildRequestBody(params, model, outputFormat, sessionId));
|
|
516
|
+
const body = JSON.stringify(buildRequestBody(params, model, outputFormat, sessionId, inputImages));
|
|
352
517
|
const headers: Record<string, string> = {
|
|
353
518
|
Authorization: `Bearer ${token}`,
|
|
354
519
|
"chatgpt-account-id": accountId,
|
|
@@ -371,8 +536,8 @@ async function requestImage(
|
|
|
371
536
|
if (!response.ok) {
|
|
372
537
|
const errorText = await response.text();
|
|
373
538
|
if (attempt <= MAX_RETRIES && isRetryableStatus(response.status, errorText)) {
|
|
374
|
-
const delay =
|
|
375
|
-
await
|
|
539
|
+
const delay = retryDelayMs(attempt, response.headers.get("retry-after"));
|
|
540
|
+
await abortableDelay(delay, signal);
|
|
376
541
|
continue;
|
|
377
542
|
}
|
|
378
543
|
throw new Error(`Codex image generation request failed (${response.status}): ${errorText}`);
|
|
@@ -393,17 +558,18 @@ export default function codexImageGen(pi: ExtensionAPI) {
|
|
|
393
558
|
name: "codex_generate_image",
|
|
394
559
|
label: "Codex Image",
|
|
395
560
|
description:
|
|
396
|
-
"Generate an image with the OpenAI Codex ChatGPT backend built-in image_generation tool (gpt-image-2). Uses the existing openai-codex login; does not require OPENAI_API_KEY.",
|
|
397
|
-
promptSnippet: "Generate bitmap images via the OpenAI Codex ChatGPT backend gpt-image-2 image_generation tool.",
|
|
561
|
+
"Generate or edit an image with the OpenAI Codex ChatGPT backend built-in image_generation tool (gpt-image-2). Accepts up to five local or recent conversation images. Uses the existing openai-codex login; does not require OPENAI_API_KEY.",
|
|
562
|
+
promptSnippet: "Generate or edit bitmap images via the OpenAI Codex ChatGPT backend gpt-image-2 image_generation tool.",
|
|
398
563
|
promptGuidelines: [
|
|
399
|
-
"Use codex_generate_image when the user asks to generate a raster image
|
|
564
|
+
"Use codex_generate_image when the user asks to generate or edit a raster image with OpenAI/Codex image generation.",
|
|
400
565
|
"Do not use codex_generate_image without a clear image-generation request, because it consumes the user's Codex image quota.",
|
|
401
566
|
],
|
|
402
567
|
parameters: TOOL_PARAMS,
|
|
403
568
|
executionMode: "parallel", // #4: safe to run concurrently — no shared state, saves serialized per-path
|
|
404
569
|
async execute(toolCallId, params: ToolParams, signal, onUpdate, ctx) {
|
|
405
570
|
const outputFormat = params.outputFormat || "png";
|
|
406
|
-
const
|
|
571
|
+
const projectTrusted = typeof ctx.isProjectTrusted === "function" && ctx.isProjectTrusted();
|
|
572
|
+
const config = loadConfig(ctx.cwd, projectTrusted); // #5: load once, pass to resolveSaveConfig
|
|
407
573
|
const requestedModel = params.model || config.model || DEFAULT_MODEL;
|
|
408
574
|
const model = ctx.modelRegistry.find(PROVIDER, requestedModel)?.id || requestedModel; // #6: removed dead FALLBACK_MODEL
|
|
409
575
|
const token = await ctx.modelRegistry.getApiKeyForProvider(PROVIDER);
|
|
@@ -412,32 +578,40 @@ export default function codexImageGen(pi: ExtensionAPI) {
|
|
|
412
578
|
}
|
|
413
579
|
const accountId = extractChatGptAccountId(token);
|
|
414
580
|
const sessionId = ctx.sessionManager.getSessionId();
|
|
581
|
+
const messages: unknown[] = [];
|
|
582
|
+
for (const entry of ctx.sessionManager.getBranch()) {
|
|
583
|
+
if (entry.type === "message") messages.push(entry.message);
|
|
584
|
+
if (entry.type === "custom_message") messages.push(entry);
|
|
585
|
+
}
|
|
586
|
+
const inputImages = await resolveInputImages(params, ctx.cwd, messages);
|
|
415
587
|
|
|
416
588
|
onUpdate?.({
|
|
417
|
-
content: [{ type: "text", text: `Requesting gpt-image-2 generation through ${PROVIDER}/${model}...` }],
|
|
418
|
-
details: { provider: PROVIDER, model, outputFormat },
|
|
589
|
+
content: [{ type: "text", text: `Requesting gpt-image-2 ${inputImages.length > 0 ? "edit" : "generation"} through ${PROVIDER}/${model}...` }],
|
|
590
|
+
details: { provider: PROVIDER, model, outputFormat, inputImageCount: inputImages.length },
|
|
419
591
|
});
|
|
420
592
|
|
|
421
|
-
const parsed = await requestImage(params, token, accountId, model, outputFormat, sessionId, signal);
|
|
593
|
+
const parsed = await requestImage(params, token, accountId, model, outputFormat, sessionId, inputImages, signal);
|
|
422
594
|
if (!parsed.image) {
|
|
423
595
|
const text = parsed.text.join("").trim();
|
|
424
596
|
throw new Error(text ? `Codex did not return an image. Response text: ${text}` : "Codex did not return an image.");
|
|
425
597
|
}
|
|
426
598
|
|
|
599
|
+
const imageBytes = decodeImageData(parsed.image.result, outputFormat);
|
|
427
600
|
const saveConfig = resolveSaveConfig(params, ctx.cwd, sessionId, config);
|
|
428
601
|
let savedPath: string | undefined;
|
|
602
|
+
let attemptedPath: string | undefined;
|
|
603
|
+
let saveWarning: string | undefined;
|
|
429
604
|
if (saveConfig.mode !== "none" && saveConfig.outputDir) {
|
|
430
|
-
|
|
431
|
-
|
|
432
|
-
|
|
433
|
-
|
|
434
|
-
|
|
435
|
-
provider: PROVIDER,
|
|
436
|
-
|
|
437
|
-
|
|
438
|
-
|
|
439
|
-
|
|
440
|
-
});
|
|
605
|
+
attemptedPath = imagePath(outputFormat, saveConfig.outputDir, parsed.image.id || toolCallId);
|
|
606
|
+
try {
|
|
607
|
+
savedPath = await saveImage(imageBytes, outputFormat, saveConfig.outputDir, parsed.image.id || toolCallId);
|
|
608
|
+
onUpdate?.({
|
|
609
|
+
content: [{ type: "text", text: `Image saved to ${savedPath}.` }],
|
|
610
|
+
details: { provider: PROVIDER, model, savedPath, byteCount: imageBytes.length },
|
|
611
|
+
});
|
|
612
|
+
} catch (error) {
|
|
613
|
+
saveWarning = `Image generation succeeded, but the image could not be saved to disk: ${error instanceof Error ? error.message : String(error)}`;
|
|
614
|
+
}
|
|
441
615
|
}
|
|
442
616
|
|
|
443
617
|
const summary = [
|
|
@@ -445,6 +619,7 @@ export default function codexImageGen(pi: ExtensionAPI) {
|
|
|
445
619
|
`Status: ${parsed.image.status}.`,
|
|
446
620
|
parsed.image.revisedPrompt ? `Revised prompt: ${parsed.image.revisedPrompt}` : undefined,
|
|
447
621
|
savedPath ? `Saved image to: ${savedPath}` : "Image was not saved to disk.",
|
|
622
|
+
saveWarning ? `Warning: ${saveWarning}` : undefined,
|
|
448
623
|
]
|
|
449
624
|
.filter(Boolean)
|
|
450
625
|
.join(" ");
|
|
@@ -461,6 +636,9 @@ export default function codexImageGen(pi: ExtensionAPI) {
|
|
|
461
636
|
outputFormat,
|
|
462
637
|
saveMode: saveConfig.mode,
|
|
463
638
|
savedPath,
|
|
639
|
+
attemptedPath,
|
|
640
|
+
saveWarning,
|
|
641
|
+
inputImageCount: inputImages.length,
|
|
464
642
|
responseId: parsed.responseId,
|
|
465
643
|
imageGenerationId: parsed.image.id,
|
|
466
644
|
revisedPrompt: parsed.image.revisedPrompt,
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "pi-codex-image-gen",
|
|
3
|
-
"version": "0.1.
|
|
3
|
+
"version": "0.1.12",
|
|
4
4
|
"description": "Image generation for Pi using the ChatGPT Images 2.0 model.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"license": "Apache-2.0",
|
|
@@ -41,6 +41,7 @@
|
|
|
41
41
|
"scripts": {
|
|
42
42
|
"check": "tsc --noEmit",
|
|
43
43
|
"typecheck": "tsc --noEmit",
|
|
44
|
+
"test": "rm -rf .test-dist && tsc --noEmit false --outDir .test-dist && node --test tests/*.test.mjs; status=$?; rm -rf .test-dist; exit $status",
|
|
44
45
|
"pack:dry-run": "npm pack --dry-run"
|
|
45
46
|
},
|
|
46
47
|
"pi": {
|
|
@@ -59,8 +60,8 @@
|
|
|
59
60
|
"devDependencies": {
|
|
60
61
|
"@earendil-works/pi-ai": "^0.80.0",
|
|
61
62
|
"@earendil-works/pi-coding-agent": "^0.80.0",
|
|
62
|
-
"typebox": "^1.
|
|
63
|
-
"@types/node": "^
|
|
63
|
+
"typebox": "^1.3.4",
|
|
64
|
+
"@types/node": "^26.1.0",
|
|
64
65
|
"typescript": "^6.0.3"
|
|
65
66
|
},
|
|
66
67
|
"publishConfig": {
|
package/skills/imagegen/SKILL.md
CHANGED
|
@@ -7,7 +7,7 @@ description: "Generate or edit raster images when the task benefits from AI-crea
|
|
|
7
7
|
|
|
8
8
|
> Adapted from OpenAI Codex's `imagegen` skill for Pi.
|
|
9
9
|
> Original source: https://github.com/openai/codex/tree/main/codex-rs/skills/src/assets/samples/imagegen
|
|
10
|
-
> Required Pi modifications: use Pi's `codex_generate_image` tool name, Pi artifact paths, and the bundled helper path
|
|
10
|
+
> Required Pi modifications: use Pi's `codex_generate_image` tool name, Pi artifact paths, and the bundled helper path.
|
|
11
11
|
|
|
12
12
|
Generates or edits images for the current project (for example website assets, game assets, UI mockups, product mockups, wireframes, logo design, photorealistic images, or infographics).
|
|
13
13
|
|
|
@@ -15,7 +15,7 @@ Generates or edits images for the current project (for example website assets, g
|
|
|
15
15
|
|
|
16
16
|
This skill has exactly two top-level modes:
|
|
17
17
|
|
|
18
|
-
- **Default Pi tool mode (preferred):** Pi `codex_generate_image` tool for
|
|
18
|
+
- **Default Pi tool mode (preferred):** Pi `codex_generate_image` tool for new image generation, edits using up to five local or recent conversation images, reference variants, and simple transparent-image requests. Does not require `OPENAI_API_KEY`.
|
|
19
19
|
- **Fallback CLI mode:** `scripts/image_gen.py` CLI. Use when the user explicitly asks for the CLI/API/model path, or after the user explicitly confirms a true model-native transparency fallback with `gpt-image-1.5`. Requires `OPENAI_API_KEY`.
|
|
20
20
|
|
|
21
21
|
Within CLI fallback, the CLI exposes three subcommands:
|
|
@@ -25,8 +25,9 @@ Within CLI fallback, the CLI exposes three subcommands:
|
|
|
25
25
|
- `generate-batch`
|
|
26
26
|
|
|
27
27
|
Rules:
|
|
28
|
-
- Use the Pi `codex_generate_image` tool by default for
|
|
29
|
-
-
|
|
28
|
+
- Use the Pi `codex_generate_image` tool by default for new image generation requests.
|
|
29
|
+
- Use `referencedImagePaths` for edits when every target has a local path. Use `numLastImagesToInclude` only when a target is available solely in recent conversation history. Never provide both selectors. Masks and advanced CLI-only controls still require confirmed CLI fallback.
|
|
30
|
+
- Do not switch to CLI fallback for ordinary generation quality, size, or output file-path control.
|
|
30
31
|
- If the user explicitly asks for a transparent image/background, stay on Pi `codex_generate_image` first: prompt for a flat removable chroma-key background, then remove it locally with the installed helper at `scripts/remove_chroma_key.py`.
|
|
31
32
|
- Never silently switch from Pi `codex_generate_image` or CLI `gpt-image-2` to CLI `gpt-image-1.5`. Treat this as a model/path downgrade and ask the user before doing it, unless the user has already explicitly requested `gpt-image-1.5`, `scripts/image_gen.py`, or CLI fallback.
|
|
32
33
|
- If a transparent request appears too complex for clean chroma-key removal, asks for true/native transparency, or local removal fails validation, explain that true transparency requires CLI `gpt-image-1.5 --background transparent --output-format png` because `gpt-image-2` does not support `background=transparent`, then ask whether to proceed. Run the CLI fallback only after the user confirms.
|
|
@@ -38,12 +39,13 @@ Rules:
|
|
|
38
39
|
Pi tool save-path policy:
|
|
39
40
|
- In Pi tool mode, generated images are saved under Pi's agent directory by default: `<pi-agent-dir>/generated-images/<pi-session-id>/<image-call-id>.*`. The default Pi agent directory is `~/.pi/agent`, but it can be overridden with `PI_CODING_AGENT_DIR`; use Pi's configured agent directory, not a hardcoded home path.
|
|
40
41
|
- Do not describe or rely on OS temp as the default Pi tool destination.
|
|
41
|
-
- Do not describe or rely on a destination-path argument (if any) on the Pi `codex_generate_image` tool. If a specific location is needed, generate first and then
|
|
42
|
+
- Do not describe or rely on a destination-path argument (if any) on the Pi `codex_generate_image` tool. If a specific location is needed, generate first and then copy the selected output from `<pi-agent-dir>/generated-images/<pi-session-id>/<image-call-id>.*`.
|
|
42
43
|
- Save-path precedence in Pi tool mode:
|
|
43
|
-
1. If the user names a destination,
|
|
44
|
-
2. If the image is meant for the current project,
|
|
44
|
+
1. If the user names a destination, copy the selected output there and leave the original in place.
|
|
45
|
+
2. If the image is meant for the current project, copy the final selected image into the workspace before finishing and leave the original in place.
|
|
45
46
|
3. If the image is only for preview or brainstorming, render it inline; the underlying file can remain at the default `<pi-agent-dir>/generated-images/<pi-session-id>/*` path.
|
|
46
47
|
- Never leave a project-referenced asset only at the default `<pi-agent-dir>/generated-images/<pi-session-id>/*` path.
|
|
48
|
+
- Move or delete a Pi-generated original only when the user explicitly requests it.
|
|
47
49
|
- Do not overwrite an existing asset unless the user explicitly asked for replacement; otherwise create a sibling versioned filename such as `hero-v2.png` or `item-icon-edited.png`.
|
|
48
50
|
|
|
49
51
|
Shared prompt guidance for both modes lives in `references/prompting.md` and `references/sample-prompts.md`.
|
|
@@ -82,9 +84,10 @@ Intent:
|
|
|
82
84
|
- If the user provides no images, treat the request as **generate**.
|
|
83
85
|
|
|
84
86
|
Pi edit semantics:
|
|
85
|
-
-
|
|
86
|
-
-
|
|
87
|
-
-
|
|
87
|
+
- Use `referencedImagePaths` when all edit targets have readable local paths, with at most five paths.
|
|
88
|
+
- Use `numLastImagesToInclude` for the smallest recent-conversation window containing all targets, from one to five images.
|
|
89
|
+
- Never provide both image selectors. If neither can include every target, ask the user to attach or provide the missing image.
|
|
90
|
+
- Use confirmed CLI fallback only for masks or other explicit CLI-only parameters.
|
|
88
91
|
- For edits, preserve invariants aggressively and save non-destructively by default.
|
|
89
92
|
|
|
90
93
|
Execution strategy:
|
|
@@ -95,7 +98,7 @@ Execution strategy:
|
|
|
95
98
|
Assume the user wants a new image unless they clearly ask to change an existing one.
|
|
96
99
|
|
|
97
100
|
## Workflow
|
|
98
|
-
1. Decide the top-level mode: Pi tool by default,
|
|
101
|
+
1. Decide the top-level mode: Pi tool by default for generation, supported edits, and simple transparent-output requests; fallback CLI only if explicitly requested or after the user confirms an unsupported edit control or transparent-output fallback.
|
|
99
102
|
2. Decide the intent: `generate` or `edit`.
|
|
100
103
|
3. Decide whether the output is preview-only or meant to be consumed by the current project.
|
|
101
104
|
4. Decide the execution strategy: single asset vs repeated Pi tool calls vs CLI `generate-batch`.
|
|
@@ -104,17 +107,17 @@ Assume the user wants a new image unless they clearly ask to change an existing
|
|
|
104
107
|
- reference image
|
|
105
108
|
- edit target
|
|
106
109
|
- supporting insert/style/compositing input
|
|
107
|
-
7.
|
|
110
|
+
7. For local edit targets, pass up to five paths through `referencedImagePaths`. For pathless conversation images, use the smallest valid `numLastImagesToInclude`.
|
|
108
111
|
8. If the user asked for a photo, illustration, sprite, product image, banner, or other explicitly raster-style asset, use `codex_generate_image` rather than substituting SVG/HTML/CSS placeholders. If the request is for an icon, logo, or UI graphic that should match existing repo-native SVG/vector/code assets, prefer editing those directly instead.
|
|
109
112
|
9. Augment the prompt based on specificity:
|
|
110
113
|
- If the user's prompt is already specific and detailed, normalize it into a clear spec without adding creative requirements.
|
|
111
114
|
- If the user's prompt is generic, add tasteful augmentation only when it materially improves output quality.
|
|
112
|
-
10. Use the Pi `codex_generate_image` tool by default.
|
|
115
|
+
10. Use the Pi `codex_generate_image` tool by default for generation and supported existing-image edits. Ask for CLI fallback confirmation only when the request requires unsupported controls such as masks.
|
|
113
116
|
11. For transparent-output requests, follow the transparent image guidance below: generate with Pi `codex_generate_image` on a flat chroma-key background, copy the selected output into the workspace or `tmp/imagegen/`, run the installed `scripts/remove_chroma_key.py` helper, and validate the alpha result before using it. If this path looks unsuitable or fails, ask before switching to CLI `gpt-image-1.5`.
|
|
114
117
|
12. Inspect outputs and validate: subject, style, composition, text accuracy, and invariants/avoid items.
|
|
115
118
|
13. Iterate with a single targeted change, then re-check.
|
|
116
119
|
14. For preview-only work, render the image inline; the underlying file may remain at the default `<pi-agent-dir>/generated-images/<pi-session-id>/<image-call-id>.*` path.
|
|
117
|
-
15. For project-bound work,
|
|
120
|
+
15. For project-bound work, copy the selected artifact into the workspace, leave the original in place, and update any consuming code or references. Never leave a project-referenced asset only at the default `<pi-agent-dir>/generated-images/<pi-session-id>/<image-call-id>.*` path.
|
|
118
121
|
16. For batches or multi-asset requests, persist every requested deliverable final in the workspace unless the user explicitly asked to keep outputs preview-only. Discarded variants do not need to be kept unless requested.
|
|
119
122
|
17. If the user explicitly chooses or confirms the CLI fallback, then use the fallback-only docs for model, quality, size, `input_fidelity`, masks, output format, output paths, and network setup.
|
|
120
123
|
18. Always report the final saved path(s) for any workspace-bound asset(s), plus the final prompt or prompt set and whether the Pi tool or fallback CLI mode was used.
|
|
@@ -126,7 +129,7 @@ Transparent-image requests still use Pi `codex_generate_image` first. Because th
|
|
|
126
129
|
Default sequence:
|
|
127
130
|
1. Use Pi `codex_generate_image` to generate the requested subject on a perfectly flat solid chroma-key background.
|
|
128
131
|
2. Choose a key color that is unlikely to appear in the subject: default `#00ff00`, use `#ff00ff` for green subjects, and avoid `#0000ff` for blue subjects.
|
|
129
|
-
3. After generation,
|
|
132
|
+
3. After generation, copy the selected source image from `<pi-agent-dir>/generated-images/<pi-session-id>/<image-call-id>.*` into the workspace or `tmp/imagegen/` and leave the original in place.
|
|
130
133
|
4. Run the bundled helper from this skill directory:
|
|
131
134
|
```bash
|
|
132
135
|
python "scripts/remove_chroma_key.py" \
|