pi-codex-image-gen 0.1.11 → 0.1.12

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,11 +1,24 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.1.12 - 2026-07-17
4
+
3
5
  All notable changes to this project will be documented in this file.
4
6
 
5
7
  This project follows the spirit of [Keep a Changelog](https://keepachangelog.com/en/1.1.0/) and uses semantic versioning for releases.
6
8
 
7
9
  ## [Unreleased]
8
10
 
11
+ ### Added
12
+
13
+ - Add native edits using up to five local or recent conversation images.
14
+ - Honor bounded `Retry-After` delays and make retry backoff abort-aware.
15
+
16
+ ### Fixed
17
+
18
+ - Strictly validate base64 data and output-format signatures before returning or saving images.
19
+ - Preserve valid inline images and report a warning when disk persistence fails.
20
+ - Keep original generated artifacts when copying selected outputs elsewhere by default.
21
+
9
22
  ## [0.1.11] - 2026-07-01
10
23
 
11
24
  ### Changed
package/README.md CHANGED
@@ -1,6 +1,6 @@
1
1
  # pi-codex-image-gen
2
2
 
3
- Image generation for [Pi](https://pi.dev) using the ChatGPT Images 2.0 model via the OpenAI Codex Responses backend.
3
+ Image generation and editing for [Pi](https://pi.dev) using the ChatGPT Images 2.0 model via the OpenAI Codex Responses backend.
4
4
 
5
5
  ## Install
6
6
 
@@ -34,7 +34,7 @@ In a Pi session:
34
34
  > Generate a pixel-art sword icon, 32×32, with a blue blade and gold hilt
35
35
  ```
36
36
 
37
- The agent will invoke `codex_generate_image` with your prompt, stream the response from the Codex backend, and save the resulting image to disk. The `model` parameter controls the Codex routing model; image generation is always performed by **gpt-image-2** on the backend.
37
+ The agent will invoke `codex_generate_image` with your prompt, optionally include up to five local or recent conversation images for editing, stream the response from the Codex backend, and save the resulting image to disk. The `model` parameter controls the Codex routing model; image generation is always performed by **gpt-image-2** on the backend.
38
38
 
39
39
  ## Authentication
40
40
 
@@ -55,7 +55,7 @@ Create a JSON config file at one (or both) of these locations:
55
55
  | Global | `~/.pi/agent/extensions/codex-image-gen.json` |
56
56
  | Project | `<project-root>/.pi/extensions/codex-image-gen.json` |
57
57
 
58
- Project config overrides global config. Example:
58
+ Project config overrides global config only when project trust is active. If project trust is declined or otherwise inactive, project config is ignored and global config still applies. Example:
59
59
 
60
60
  ```json
61
61
  {
@@ -89,7 +89,7 @@ Project config overrides global config. Example:
89
89
  | `none` | Image is returned inline but not written to disk. |
90
90
  | `project` | Saves to `<project>/.pi/generated-images/<session-id>/`. |
91
91
  | `global` | Saves to `~/.pi/agent/generated-images/<session-id>/`. |
92
- | `custom` | Saves to a user-specified directory (requires `saveDir` or env). |
92
+ | `custom` | Saves to a user-specified directory (requires `saveDir` or env). `~` and `~/...` expand to the current user's home directory. |
93
93
 
94
94
  ## Tool parameters
95
95
 
@@ -100,15 +100,18 @@ Project config overrides global config. Example:
100
100
  | `outputFormat` | string | — | `png` (default), `jpeg`, or `webp`. |
101
101
  | `save` | string | — | Override save mode for this call. |
102
102
  | `saveDir` | string | — | Directory when `save=custom`. Relative paths resolve under CWD. |
103
+ | `referencedImagePaths` | string[] | — | Up to five local images to edit. Relative paths resolve under CWD. |
104
+ | `numLastImagesToInclude` | integer | — | Include the most recent one to five conversation images for editing. Mutually exclusive with `referencedImagePaths`. |
103
105
 
104
106
  ## How it works
105
107
 
106
108
  1. Resolves auth via Pi's `openai-codex` provider (ChatGPT session token).
107
109
  2. Sends a Codex Responses API request to the routing model (default `gpt-5.5`) with the `image_generation` tool enabled.
108
- 3. The backend invokes **gpt-image-2** to generate the image.
109
- 4. Parses the SSE stream for `response.output_item.done` events containing the base64 image.
110
- 5. Saves the image to disk according to the active save mode.
111
- 6. Returns the image data inline plus metadata (model, format, path, revised prompt, usage).
110
+ 3. For edits, attaches the selected local or conversation images to the request.
111
+ 4. The backend invokes **gpt-image-2** to generate or edit the image.
112
+ 5. Parses the SSE stream and strictly validates the returned base64 and image format.
113
+ 6. Saves the image according to the active save mode; persistence failures produce a warning without discarding a valid inline image.
114
+ 7. Returns the image data inline plus metadata (model, format, path, revised prompt, usage).
112
115
 
113
116
  ## Troubleshooting
114
117
 
@@ -7,7 +7,8 @@
7
7
  */
8
8
 
9
9
  import { readFileSync } from "node:fs";
10
- import { mkdir, writeFile } from "node:fs/promises";
10
+ import { mkdir, readFile, writeFile } from "node:fs/promises";
11
+ import { homedir } from "node:os";
11
12
  import { isAbsolute, join, resolve } from "node:path";
12
13
  import { StringEnum } from "@earendil-works/pi-ai";
13
14
  import { type ExtensionAPI, getAgentDir, withFileMutationQueue } from "@earendil-works/pi-coding-agent";
@@ -23,6 +24,8 @@ const DEFAULT_SAVE_MODE = "global";
23
24
  const OPENAI_BETA_HEADER = "responses=experimental";
24
25
  const MAX_RETRIES = 3;
25
26
  const BASE_DELAY_MS = 1000;
27
+ const MAX_RETRY_DELAY_MS = 30_000;
28
+ const MAX_EDIT_IMAGES = 5;
26
29
 
27
30
  const SAVE_MODES = ["none", "project", "global", "custom"] as const;
28
31
  type SaveMode = (typeof SAVE_MODES)[number];
@@ -37,9 +40,50 @@ function isRetryableStatus(status: number, errorText: string): boolean {
37
40
  return /rate.?limit|overloaded|service.?unavailable|upstream.?connect|connection.?refused/i.test(errorText);
38
41
  }
39
42
 
40
- function backoffMs(attempt: number): number {
41
- const jitter = 0.9 + Math.random() * 0.2; // matches codex-rs jitter range
42
- return BASE_DELAY_MS * 2 ** (attempt - 1) * jitter;
43
+ export function parseRetryAfter(value: string | null, nowMs = Date.now()): number | undefined {
44
+ if (!value) return undefined;
45
+ const trimmed = value.trim();
46
+ if (/^\d+(?:\.\d+)?$/.test(trimmed)) {
47
+ const milliseconds = Number(trimmed) * 1000;
48
+ return Number.isFinite(milliseconds) ? Math.min(milliseconds, MAX_RETRY_DELAY_MS) : undefined;
49
+ }
50
+ const dateMs = Date.parse(trimmed);
51
+ if (!Number.isFinite(dateMs) || dateMs <= nowMs) return undefined;
52
+ return Math.min(dateMs - nowMs, MAX_RETRY_DELAY_MS);
53
+ }
54
+
55
+ export function retryDelayMs(
56
+ attempt: number,
57
+ retryAfter: string | null,
58
+ random = Math.random,
59
+ nowMs = Date.now(),
60
+ ): number {
61
+ const serverDelay = parseRetryAfter(retryAfter, nowMs);
62
+ if (serverDelay !== undefined) {
63
+ return Math.floor(Math.min(serverDelay * (1 + random() * 0.1), MAX_RETRY_DELAY_MS));
64
+ }
65
+ const exponential = Math.min(BASE_DELAY_MS * 2 ** (attempt - 1), MAX_RETRY_DELAY_MS);
66
+ return Math.floor(exponential * (0.9 + random() * 0.2));
67
+ }
68
+
69
+ export function abortableDelay(milliseconds: number, signal?: AbortSignal): Promise<void> {
70
+ if (signal?.aborted) return Promise.reject(new Error("Image generation was aborted."));
71
+ return new Promise<void>((resolve, reject) => {
72
+ const timer = setTimeout(finish, milliseconds);
73
+ function cleanup() {
74
+ clearTimeout(timer);
75
+ signal?.removeEventListener("abort", abort);
76
+ }
77
+ function finish() {
78
+ cleanup();
79
+ resolve();
80
+ }
81
+ function abort() {
82
+ cleanup();
83
+ reject(new Error("Image generation was aborted."));
84
+ }
85
+ signal?.addEventListener("abort", abort, { once: true });
86
+ });
43
87
  }
44
88
 
45
89
  // --- Tool parameter schema ---
@@ -56,6 +100,19 @@ const TOOL_PARAMS = Type.Object({
56
100
  description: "Directory to save the image when save=custom. Relative paths resolve under the current workspace.",
57
101
  }),
58
102
  ),
103
+ referencedImagePaths: Type.Optional(
104
+ Type.Array(Type.String(), {
105
+ maxItems: MAX_EDIT_IMAGES,
106
+ description: "Up to five local image paths to edit. Relative paths resolve under the current workspace.",
107
+ }),
108
+ ),
109
+ numLastImagesToInclude: Type.Optional(
110
+ Type.Integer({
111
+ minimum: 1,
112
+ maximum: MAX_EDIT_IMAGES,
113
+ description: "Use the most recent one to five images from the current conversation as edit inputs.",
114
+ }),
115
+ ),
59
116
  });
60
117
 
61
118
  type ToolParams = Static<typeof TOOL_PARAMS>;
@@ -87,6 +144,11 @@ interface ParsedCodexResponse {
87
144
  usage?: unknown;
88
145
  }
89
146
 
147
+ interface InputImage {
148
+ data: string;
149
+ mimeType: string;
150
+ }
151
+
90
152
  // --- #11: Typed SSE event discriminated union ---
91
153
 
92
154
  type CodexSseEvent =
@@ -143,15 +205,18 @@ function readConfigFile(path: string): ExtensionConfig {
143
205
  }
144
206
  }
145
207
 
146
- function loadConfig(cwd: string): ExtensionConfig {
147
- const globalConfig = readConfigFile(join(getAgentDir(), "extensions", "codex-image-gen.json"));
208
+ export function loadConfig(cwd: string, projectTrusted: boolean, agentDir = getAgentDir()): ExtensionConfig {
209
+ const globalConfig = readConfigFile(join(agentDir, "extensions", "codex-image-gen.json"));
210
+ if (!projectTrusted) return globalConfig;
148
211
  const projectConfig = readConfigFile(join(cwd, ".pi", "extensions", "codex-image-gen.json"));
149
212
  return { ...globalConfig, ...projectConfig };
150
213
  }
151
214
 
152
215
  // --- Path helpers ---
153
216
 
154
- function resolveUnderCwd(cwd: string, path: string): string {
217
+ export function resolveUnderCwd(cwd: string, path: string, homeDir = homedir()): string {
218
+ if (path === "~") return homeDir;
219
+ if (path.startsWith("~/")) return resolve(homeDir, path.slice(2));
155
220
  return isAbsolute(path) ? path : resolve(cwd, path);
156
221
  }
157
222
 
@@ -199,22 +264,115 @@ function mimeForFormat(outputFormat: OutputFormat): string {
199
264
  return outputFormat === "jpeg" ? "image/jpeg" : `image/${outputFormat}`;
200
265
  }
201
266
 
202
- async function saveImage(base64Data: string, outputFormat: OutputFormat, outputDir: string, imageCallId: string): Promise<string> {
267
+ function imagePath(outputFormat: OutputFormat, outputDir: string, imageCallId: string): string {
203
268
  const filename = `${sanitizePathPart(imageCallId, "image_generation")}.${extensionForFormat(outputFormat)}`;
204
- const filePath = join(outputDir, filename);
269
+ return join(outputDir, filename);
270
+ }
271
+
272
+ export function decodeImageData(base64Data: string, outputFormat: OutputFormat): Buffer {
273
+ const value = base64Data.trim();
274
+ if (!value || value.length % 4 !== 0 || !/^(?:[A-Za-z0-9+/]{4})*(?:[A-Za-z0-9+/]{2}==|[A-Za-z0-9+/]{3}=)?$/.test(value)) {
275
+ throw new Error("Codex returned invalid base64 image data.");
276
+ }
277
+ const bytes = Buffer.from(value, "base64");
278
+ if (bytes.length === 0 || bytes.toString("base64") !== value) {
279
+ throw new Error("Codex returned invalid base64 image data.");
280
+ }
281
+ const validSignature =
282
+ (outputFormat === "png" && bytes.length >= 8 && bytes.subarray(0, 8).equals(Buffer.from([0x89, 0x50, 0x4e, 0x47, 0x0d, 0x0a, 0x1a, 0x0a]))) ||
283
+ (outputFormat === "jpeg" && bytes.length >= 3 && bytes[0] === 0xff && bytes[1] === 0xd8 && bytes[2] === 0xff) ||
284
+ (outputFormat === "webp" && bytes.length >= 12 && bytes.toString("ascii", 0, 4) === "RIFF" && bytes.toString("ascii", 8, 12) === "WEBP");
285
+ if (!validSignature) throw new Error(`Codex returned image data that does not match ${outputFormat}.`);
286
+ return bytes;
287
+ }
288
+
289
+ async function saveImage(
290
+ bytes: Buffer,
291
+ outputFormat: OutputFormat,
292
+ outputDir: string,
293
+ imageCallId: string,
294
+ ): Promise<string> {
295
+ const filePath = imagePath(outputFormat, outputDir, imageCallId);
205
296
  await withFileMutationQueue(filePath, async () => {
206
297
  await mkdir(outputDir, { recursive: true });
207
- await writeFile(filePath, Buffer.from(base64Data, "base64"));
298
+ await writeFile(filePath, bytes);
208
299
  });
209
300
  return filePath;
210
301
  }
211
302
 
303
+ export function selectRecentImages(messages: unknown[], count: number): InputImage[] {
304
+ const images: InputImage[] = [];
305
+ for (let index = messages.length - 1; index >= 0 && images.length < count; index--) {
306
+ const message = messages[index] as { content?: unknown };
307
+ if (!Array.isArray(message?.content)) continue;
308
+ for (let contentIndex = message.content.length - 1; contentIndex >= 0 && images.length < count; contentIndex--) {
309
+ const block = message.content[contentIndex] as { type?: unknown; data?: unknown; mimeType?: unknown };
310
+ if (block?.type === "image" && typeof block.data === "string" && typeof block.mimeType === "string") {
311
+ images.push({ data: block.data, mimeType: block.mimeType });
312
+ }
313
+ }
314
+ }
315
+ return images.reverse();
316
+ }
317
+
318
+ function mimeFromBytes(bytes: Buffer, path: string): string {
319
+ if (bytes.length >= 8 && bytes.subarray(0, 8).equals(Buffer.from([0x89, 0x50, 0x4e, 0x47, 0x0d, 0x0a, 0x1a, 0x0a]))) return "image/png";
320
+ if (bytes.length >= 3 && bytes[0] === 0xff && bytes[1] === 0xd8 && bytes[2] === 0xff) return "image/jpeg";
321
+ if (bytes.length >= 12 && bytes.toString("ascii", 0, 4) === "RIFF" && bytes.toString("ascii", 8, 12) === "WEBP") return "image/webp";
322
+ throw new Error(`Referenced image is unavailable or unsupported: ${path}`);
323
+ }
324
+
325
+ export async function resolveInputImages(
326
+ params: ToolParams,
327
+ cwd: string,
328
+ messages: unknown[],
329
+ ): Promise<InputImage[]> {
330
+ const paths = params.referencedImagePaths ?? [];
331
+ if (paths.length > 0 && params.numLastImagesToInclude !== undefined) {
332
+ throw new Error("Provide only one of referencedImagePaths or numLastImagesToInclude.");
333
+ }
334
+ if (paths.length > MAX_EDIT_IMAGES) throw new Error(`referencedImagePaths accepts at most ${MAX_EDIT_IMAGES} paths.`);
335
+ if (paths.length > 0) {
336
+ return Promise.all(
337
+ paths.map(async (path) => {
338
+ const normalized = path.startsWith("@") ? path.slice(1) : path;
339
+ const absolutePath = resolveUnderCwd(cwd, normalized);
340
+ let bytes: Buffer;
341
+ try {
342
+ bytes = await readFile(absolutePath);
343
+ } catch (error) {
344
+ throw new Error(`Unable to read referenced image at ${absolutePath}: ${error instanceof Error ? error.message : String(error)}`);
345
+ }
346
+ return { data: bytes.toString("base64"), mimeType: mimeFromBytes(bytes, absolutePath) };
347
+ }),
348
+ );
349
+ }
350
+ if (params.numLastImagesToInclude !== undefined) {
351
+ const count = params.numLastImagesToInclude;
352
+ if (!Number.isInteger(count) || count < 1 || count > MAX_EDIT_IMAGES) {
353
+ throw new Error(`numLastImagesToInclude must be between 1 and ${MAX_EDIT_IMAGES}.`);
354
+ }
355
+ const images = selectRecentImages(messages, count);
356
+ if (images.length !== count) {
357
+ throw new Error(`Requested the last ${count} conversation images, but only ${images.length} were available.`);
358
+ }
359
+ return images;
360
+ }
361
+ return [];
362
+ }
363
+
212
364
  // --- Request building ---
213
365
  // #2: prompt_cache_key set to sessionId
214
366
  // #7: parallel_tool_calls: false
215
367
  // #14: include removed (not needed without reasoning)
216
368
 
217
- function buildRequestBody(params: ToolParams, model: string, outputFormat: OutputFormat, sessionId: string) {
369
+ export function buildRequestBody(
370
+ params: ToolParams,
371
+ model: string,
372
+ outputFormat: OutputFormat,
373
+ sessionId: string,
374
+ inputImages: InputImage[] = [],
375
+ ) {
218
376
  return {
219
377
  model,
220
378
  store: false,
@@ -225,7 +383,13 @@ function buildRequestBody(params: ToolParams, model: string, outputFormat: Outpu
225
383
  input: [
226
384
  {
227
385
  role: "user",
228
- content: [{ type: "input_text", text: params.prompt }],
386
+ content: [
387
+ { type: "input_text", text: params.prompt },
388
+ ...inputImages.map((image) => ({
389
+ type: "input_image",
390
+ image_url: `data:${image.mimeType};base64,${image.data}`,
391
+ })),
392
+ ],
229
393
  },
230
394
  ],
231
395
  tools: [{ type: "image_generation", output_format: outputFormat }],
@@ -346,9 +510,10 @@ async function requestImage(
346
510
  model: string,
347
511
  outputFormat: OutputFormat,
348
512
  sessionId: string,
513
+ inputImages: InputImage[],
349
514
  signal?: AbortSignal,
350
515
  ): Promise<ParsedCodexResponse> {
351
- const body = JSON.stringify(buildRequestBody(params, model, outputFormat, sessionId));
516
+ const body = JSON.stringify(buildRequestBody(params, model, outputFormat, sessionId, inputImages));
352
517
  const headers: Record<string, string> = {
353
518
  Authorization: `Bearer ${token}`,
354
519
  "chatgpt-account-id": accountId,
@@ -371,8 +536,8 @@ async function requestImage(
371
536
  if (!response.ok) {
372
537
  const errorText = await response.text();
373
538
  if (attempt <= MAX_RETRIES && isRetryableStatus(response.status, errorText)) {
374
- const delay = backoffMs(attempt);
375
- await new Promise<void>((resolve) => setTimeout(resolve, delay));
539
+ const delay = retryDelayMs(attempt, response.headers.get("retry-after"));
540
+ await abortableDelay(delay, signal);
376
541
  continue;
377
542
  }
378
543
  throw new Error(`Codex image generation request failed (${response.status}): ${errorText}`);
@@ -393,17 +558,18 @@ export default function codexImageGen(pi: ExtensionAPI) {
393
558
  name: "codex_generate_image",
394
559
  label: "Codex Image",
395
560
  description:
396
- "Generate an image with the OpenAI Codex ChatGPT backend built-in image_generation tool (gpt-image-2). Uses the existing openai-codex login; does not require OPENAI_API_KEY.",
397
- promptSnippet: "Generate bitmap images via the OpenAI Codex ChatGPT backend gpt-image-2 image_generation tool.",
561
+ "Generate or edit an image with the OpenAI Codex ChatGPT backend built-in image_generation tool (gpt-image-2). Accepts up to five local or recent conversation images. Uses the existing openai-codex login; does not require OPENAI_API_KEY.",
562
+ promptSnippet: "Generate or edit bitmap images via the OpenAI Codex ChatGPT backend gpt-image-2 image_generation tool.",
398
563
  promptGuidelines: [
399
- "Use codex_generate_image when the user asks to generate a raster image, illustration, photo, sprite, icon draft, banner, or other bitmap asset with OpenAI/Codex image generation.",
564
+ "Use codex_generate_image when the user asks to generate or edit a raster image with OpenAI/Codex image generation.",
400
565
  "Do not use codex_generate_image without a clear image-generation request, because it consumes the user's Codex image quota.",
401
566
  ],
402
567
  parameters: TOOL_PARAMS,
403
568
  executionMode: "parallel", // #4: safe to run concurrently — no shared state, saves serialized per-path
404
569
  async execute(toolCallId, params: ToolParams, signal, onUpdate, ctx) {
405
570
  const outputFormat = params.outputFormat || "png";
406
- const config = loadConfig(ctx.cwd); // #5: load once, pass to resolveSaveConfig
571
+ const projectTrusted = typeof ctx.isProjectTrusted === "function" && ctx.isProjectTrusted();
572
+ const config = loadConfig(ctx.cwd, projectTrusted); // #5: load once, pass to resolveSaveConfig
407
573
  const requestedModel = params.model || config.model || DEFAULT_MODEL;
408
574
  const model = ctx.modelRegistry.find(PROVIDER, requestedModel)?.id || requestedModel; // #6: removed dead FALLBACK_MODEL
409
575
  const token = await ctx.modelRegistry.getApiKeyForProvider(PROVIDER);
@@ -412,32 +578,40 @@ export default function codexImageGen(pi: ExtensionAPI) {
412
578
  }
413
579
  const accountId = extractChatGptAccountId(token);
414
580
  const sessionId = ctx.sessionManager.getSessionId();
581
+ const messages: unknown[] = [];
582
+ for (const entry of ctx.sessionManager.getBranch()) {
583
+ if (entry.type === "message") messages.push(entry.message);
584
+ if (entry.type === "custom_message") messages.push(entry);
585
+ }
586
+ const inputImages = await resolveInputImages(params, ctx.cwd, messages);
415
587
 
416
588
  onUpdate?.({
417
- content: [{ type: "text", text: `Requesting gpt-image-2 generation through ${PROVIDER}/${model}...` }],
418
- details: { provider: PROVIDER, model, outputFormat },
589
+ content: [{ type: "text", text: `Requesting gpt-image-2 ${inputImages.length > 0 ? "edit" : "generation"} through ${PROVIDER}/${model}...` }],
590
+ details: { provider: PROVIDER, model, outputFormat, inputImageCount: inputImages.length },
419
591
  });
420
592
 
421
- const parsed = await requestImage(params, token, accountId, model, outputFormat, sessionId, signal);
593
+ const parsed = await requestImage(params, token, accountId, model, outputFormat, sessionId, inputImages, signal);
422
594
  if (!parsed.image) {
423
595
  const text = parsed.text.join("").trim();
424
596
  throw new Error(text ? `Codex did not return an image. Response text: ${text}` : "Codex did not return an image.");
425
597
  }
426
598
 
599
+ const imageBytes = decodeImageData(parsed.image.result, outputFormat);
427
600
  const saveConfig = resolveSaveConfig(params, ctx.cwd, sessionId, config);
428
601
  let savedPath: string | undefined;
602
+ let attemptedPath: string | undefined;
603
+ let saveWarning: string | undefined;
429
604
  if (saveConfig.mode !== "none" && saveConfig.outputDir) {
430
- savedPath = await saveImage(parsed.image.result, outputFormat, saveConfig.outputDir, parsed.image.id || toolCallId);
431
- // #12: second onUpdate after save with path + byte count
432
- onUpdate?.({
433
- content: [{ type: "text", text: `Image saved to ${savedPath}.` }],
434
- details: {
435
- provider: PROVIDER,
436
- model,
437
- savedPath,
438
- byteCount: Buffer.byteLength(parsed.image.result, "base64"),
439
- },
440
- });
605
+ attemptedPath = imagePath(outputFormat, saveConfig.outputDir, parsed.image.id || toolCallId);
606
+ try {
607
+ savedPath = await saveImage(imageBytes, outputFormat, saveConfig.outputDir, parsed.image.id || toolCallId);
608
+ onUpdate?.({
609
+ content: [{ type: "text", text: `Image saved to ${savedPath}.` }],
610
+ details: { provider: PROVIDER, model, savedPath, byteCount: imageBytes.length },
611
+ });
612
+ } catch (error) {
613
+ saveWarning = `Image generation succeeded, but the image could not be saved to disk: ${error instanceof Error ? error.message : String(error)}`;
614
+ }
441
615
  }
442
616
 
443
617
  const summary = [
@@ -445,6 +619,7 @@ export default function codexImageGen(pi: ExtensionAPI) {
445
619
  `Status: ${parsed.image.status}.`,
446
620
  parsed.image.revisedPrompt ? `Revised prompt: ${parsed.image.revisedPrompt}` : undefined,
447
621
  savedPath ? `Saved image to: ${savedPath}` : "Image was not saved to disk.",
622
+ saveWarning ? `Warning: ${saveWarning}` : undefined,
448
623
  ]
449
624
  .filter(Boolean)
450
625
  .join(" ");
@@ -461,6 +636,9 @@ export default function codexImageGen(pi: ExtensionAPI) {
461
636
  outputFormat,
462
637
  saveMode: saveConfig.mode,
463
638
  savedPath,
639
+ attemptedPath,
640
+ saveWarning,
641
+ inputImageCount: inputImages.length,
464
642
  responseId: parsed.responseId,
465
643
  imageGenerationId: parsed.image.id,
466
644
  revisedPrompt: parsed.image.revisedPrompt,
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-codex-image-gen",
3
- "version": "0.1.11",
3
+ "version": "0.1.12",
4
4
  "description": "Image generation for Pi using the ChatGPT Images 2.0 model.",
5
5
  "type": "module",
6
6
  "license": "Apache-2.0",
@@ -41,6 +41,7 @@
41
41
  "scripts": {
42
42
  "check": "tsc --noEmit",
43
43
  "typecheck": "tsc --noEmit",
44
+ "test": "rm -rf .test-dist && tsc --noEmit false --outDir .test-dist && node --test tests/*.test.mjs; status=$?; rm -rf .test-dist; exit $status",
44
45
  "pack:dry-run": "npm pack --dry-run"
45
46
  },
46
47
  "pi": {
@@ -59,8 +60,8 @@
59
60
  "devDependencies": {
60
61
  "@earendil-works/pi-ai": "^0.80.0",
61
62
  "@earendil-works/pi-coding-agent": "^0.80.0",
62
- "typebox": "^1.2.10",
63
- "@types/node": "^25.9.3",
63
+ "typebox": "^1.3.4",
64
+ "@types/node": "^26.1.0",
64
65
  "typescript": "^6.0.3"
65
66
  },
66
67
  "publishConfig": {
@@ -7,7 +7,7 @@ description: "Generate or edit raster images when the task benefits from AI-crea
7
7
 
8
8
  > Adapted from OpenAI Codex's `imagegen` skill for Pi.
9
9
  > Original source: https://github.com/openai/codex/tree/main/codex-rs/skills/src/assets/samples/imagegen
10
- > Required Pi modifications: use Pi's `codex_generate_image` tool name, Pi artifact paths, and the bundled helper path; direct image editing remains CLI fallback unless the Pi tool grows edit support.
10
+ > Required Pi modifications: use Pi's `codex_generate_image` tool name, Pi artifact paths, and the bundled helper path.
11
11
 
12
12
  Generates or edits images for the current project (for example website assets, game assets, UI mockups, product mockups, wireframes, logo design, photorealistic images, or infographics).
13
13
 
@@ -15,7 +15,7 @@ Generates or edits images for the current project (for example website assets, g
15
15
 
16
16
  This skill has exactly two top-level modes:
17
17
 
18
- - **Default Pi tool mode (preferred):** Pi `codex_generate_image` tool for normal image generation, editing, and simple transparent-image requests. Does not require `OPENAI_API_KEY`.
18
+ - **Default Pi tool mode (preferred):** Pi `codex_generate_image` tool for new image generation, edits using up to five local or recent conversation images, reference variants, and simple transparent-image requests. Does not require `OPENAI_API_KEY`.
19
19
  - **Fallback CLI mode:** `scripts/image_gen.py` CLI. Use when the user explicitly asks for the CLI/API/model path, or after the user explicitly confirms a true model-native transparency fallback with `gpt-image-1.5`. Requires `OPENAI_API_KEY`.
20
20
 
21
21
  Within CLI fallback, the CLI exposes three subcommands:
@@ -25,8 +25,9 @@ Within CLI fallback, the CLI exposes three subcommands:
25
25
  - `generate-batch`
26
26
 
27
27
  Rules:
28
- - Use the Pi `codex_generate_image` tool by default for normal image generation and editing requests.
29
- - Do not switch to CLI fallback for ordinary quality, size, or file-path control.
28
+ - Use the Pi `codex_generate_image` tool by default for new image generation requests.
29
+ - Use `referencedImagePaths` for edits when every target has a local path. Use `numLastImagesToInclude` only when a target is available solely in recent conversation history. Never provide both selectors. Masks and advanced CLI-only controls still require confirmed CLI fallback.
30
+ - Do not switch to CLI fallback for ordinary generation quality, size, or output file-path control.
30
31
  - If the user explicitly asks for a transparent image/background, stay on Pi `codex_generate_image` first: prompt for a flat removable chroma-key background, then remove it locally with the installed helper at `scripts/remove_chroma_key.py`.
31
32
  - Never silently switch from Pi `codex_generate_image` or CLI `gpt-image-2` to CLI `gpt-image-1.5`. Treat this as a model/path downgrade and ask the user before doing it, unless the user has already explicitly requested `gpt-image-1.5`, `scripts/image_gen.py`, or CLI fallback.
32
33
  - If a transparent request appears too complex for clean chroma-key removal, asks for true/native transparency, or local removal fails validation, explain that true transparency requires CLI `gpt-image-1.5 --background transparent --output-format png` because `gpt-image-2` does not support `background=transparent`, then ask whether to proceed. Run the CLI fallback only after the user confirms.
@@ -38,12 +39,13 @@ Rules:
38
39
  Pi tool save-path policy:
39
40
  - In Pi tool mode, generated images are saved under Pi's agent directory by default: `<pi-agent-dir>/generated-images/<pi-session-id>/<image-call-id>.*`. The default Pi agent directory is `~/.pi/agent`, but it can be overridden with `PI_CODING_AGENT_DIR`; use Pi's configured agent directory, not a hardcoded home path.
40
41
  - Do not describe or rely on OS temp as the default Pi tool destination.
41
- - Do not describe or rely on a destination-path argument (if any) on the Pi `codex_generate_image` tool. If a specific location is needed, generate first and then move or copy the selected output from `<pi-agent-dir>/generated-images/<pi-session-id>/<image-call-id>.*`.
42
+ - Do not describe or rely on a destination-path argument (if any) on the Pi `codex_generate_image` tool. If a specific location is needed, generate first and then copy the selected output from `<pi-agent-dir>/generated-images/<pi-session-id>/<image-call-id>.*`.
42
43
  - Save-path precedence in Pi tool mode:
43
- 1. If the user names a destination, move or copy the selected output there.
44
- 2. If the image is meant for the current project, move or copy the final selected image into the workspace before finishing.
44
+ 1. If the user names a destination, copy the selected output there and leave the original in place.
45
+ 2. If the image is meant for the current project, copy the final selected image into the workspace before finishing and leave the original in place.
45
46
  3. If the image is only for preview or brainstorming, render it inline; the underlying file can remain at the default `<pi-agent-dir>/generated-images/<pi-session-id>/*` path.
46
47
  - Never leave a project-referenced asset only at the default `<pi-agent-dir>/generated-images/<pi-session-id>/*` path.
48
+ - Move or delete a Pi-generated original only when the user explicitly requests it.
47
49
  - Do not overwrite an existing asset unless the user explicitly asked for replacement; otherwise create a sibling versioned filename such as `hero-v2.png` or `item-icon-edited.png`.
48
50
 
49
51
  Shared prompt guidance for both modes lives in `references/prompting.md` and `references/sample-prompts.md`.
@@ -82,9 +84,10 @@ Intent:
82
84
  - If the user provides no images, treat the request as **generate**.
83
85
 
84
86
  Pi edit semantics:
85
- - The current Pi `codex_generate_image` tool is for new image generation. Do not promise arbitrary filesystem-path editing through the Pi tool.
86
- - If the user wants to edit an existing image, use the explicit CLI fallback only when the user asks for it or confirms it.
87
- - If a local file needs direct file-path control, masks, or other explicit CLI-only parameters, use the explicit CLI fallback only after confirmation.
87
+ - Use `referencedImagePaths` when all edit targets have readable local paths, with at most five paths.
88
+ - Use `numLastImagesToInclude` for the smallest recent-conversation window containing all targets, from one to five images.
89
+ - Never provide both image selectors. If neither can include every target, ask the user to attach or provide the missing image.
90
+ - Use confirmed CLI fallback only for masks or other explicit CLI-only parameters.
88
91
  - For edits, preserve invariants aggressively and save non-destructively by default.
89
92
 
90
93
  Execution strategy:
@@ -95,7 +98,7 @@ Execution strategy:
95
98
  Assume the user wants a new image unless they clearly ask to change an existing one.
96
99
 
97
100
  ## Workflow
98
- 1. Decide the top-level mode: Pi tool by default, including simple transparent-output requests; fallback CLI only if explicitly requested or after the user explicitly confirms a transparent-output fallback.
101
+ 1. Decide the top-level mode: Pi tool by default for generation, supported edits, and simple transparent-output requests; fallback CLI only if explicitly requested or after the user confirms an unsupported edit control or transparent-output fallback.
99
102
  2. Decide the intent: `generate` or `edit`.
100
103
  3. Decide whether the output is preview-only or meant to be consumed by the current project.
101
104
  4. Decide the execution strategy: single asset vs repeated Pi tool calls vs CLI `generate-batch`.
@@ -104,17 +107,17 @@ Assume the user wants a new image unless they clearly ask to change an existing
104
107
  - reference image
105
108
  - edit target
106
109
  - supporting insert/style/compositing input
107
- 7. If the edit target is only on the local filesystem, use CLI fallback for direct edits only after the user asks for or confirms fallback mode.
110
+ 7. For local edit targets, pass up to five paths through `referencedImagePaths`. For pathless conversation images, use the smallest valid `numLastImagesToInclude`.
108
111
  8. If the user asked for a photo, illustration, sprite, product image, banner, or other explicitly raster-style asset, use `codex_generate_image` rather than substituting SVG/HTML/CSS placeholders. If the request is for an icon, logo, or UI graphic that should match existing repo-native SVG/vector/code assets, prefer editing those directly instead.
109
112
  9. Augment the prompt based on specificity:
110
113
  - If the user's prompt is already specific and detailed, normalize it into a clear spec without adding creative requirements.
111
114
  - If the user's prompt is generic, add tasteful augmentation only when it materially improves output quality.
112
- 10. Use the Pi `codex_generate_image` tool by default.
115
+ 10. Use the Pi `codex_generate_image` tool by default for generation and supported existing-image edits. Ask for CLI fallback confirmation only when the request requires unsupported controls such as masks.
113
116
  11. For transparent-output requests, follow the transparent image guidance below: generate with Pi `codex_generate_image` on a flat chroma-key background, copy the selected output into the workspace or `tmp/imagegen/`, run the installed `scripts/remove_chroma_key.py` helper, and validate the alpha result before using it. If this path looks unsuitable or fails, ask before switching to CLI `gpt-image-1.5`.
114
117
  12. Inspect outputs and validate: subject, style, composition, text accuracy, and invariants/avoid items.
115
118
  13. Iterate with a single targeted change, then re-check.
116
119
  14. For preview-only work, render the image inline; the underlying file may remain at the default `<pi-agent-dir>/generated-images/<pi-session-id>/<image-call-id>.*` path.
117
- 15. For project-bound work, move or copy the selected artifact into the workspace and update any consuming code or references. Never leave a project-referenced asset only at the default `<pi-agent-dir>/generated-images/<pi-session-id>/<image-call-id>.*` path.
120
+ 15. For project-bound work, copy the selected artifact into the workspace, leave the original in place, and update any consuming code or references. Never leave a project-referenced asset only at the default `<pi-agent-dir>/generated-images/<pi-session-id>/<image-call-id>.*` path.
118
121
  16. For batches or multi-asset requests, persist every requested deliverable final in the workspace unless the user explicitly asked to keep outputs preview-only. Discarded variants do not need to be kept unless requested.
119
122
  17. If the user explicitly chooses or confirms the CLI fallback, then use the fallback-only docs for model, quality, size, `input_fidelity`, masks, output format, output paths, and network setup.
120
123
  18. Always report the final saved path(s) for any workspace-bound asset(s), plus the final prompt or prompt set and whether the Pi tool or fallback CLI mode was used.
@@ -126,7 +129,7 @@ Transparent-image requests still use Pi `codex_generate_image` first. Because th
126
129
  Default sequence:
127
130
  1. Use Pi `codex_generate_image` to generate the requested subject on a perfectly flat solid chroma-key background.
128
131
  2. Choose a key color that is unlikely to appear in the subject: default `#00ff00`, use `#ff00ff` for green subjects, and avoid `#0000ff` for blue subjects.
129
- 3. After generation, move or copy the selected source image from `<pi-agent-dir>/generated-images/<pi-session-id>/<image-call-id>.*` into the workspace or `tmp/imagegen/`.
132
+ 3. After generation, copy the selected source image from `<pi-agent-dir>/generated-images/<pi-session-id>/<image-call-id>.*` into the workspace or `tmp/imagegen/` and leave the original in place.
130
133
  4. Run the bundled helper from this skill directory:
131
134
  ```bash
132
135
  python "scripts/remove_chroma_key.py" \