Image Playground
A personal image canvas with generation, point edits, visual references, and version history.
What it does
Image Playground opens with a responsive gallery of image stacks, most recently edited first, and a + New image tile.
Each original generation or upload starts a separate stack. Global and point edits stay in that stack,
including edits made from older versions. The gallery shows only the latest version, with a stack symbol
and version count; all retained versions remain accessible in the editor history. This applies to web,
iOS, and macOS. Existing histories without stack IDs group consecutive edits with the preceding original;
old branches cannot be reconstructed precisely because their source was not stored.
Select a thumbnail to reopen the latest version in the editor, or select New image to open an empty canvas.
Use Back to gallery to return to the overview. Starting a new image or editing an older version
preserves the other gallery entries, within the existing limit of 80 retained versions.
Each gallery tile has a hover/focus Delete image control with an inline confirmation, matching the
Video Studio gallery. Select multiple in the header switches the gallery into selection mode: tiles
toggle a checkmark, a selection bar shows the count with Select all, Clear selection, and
Delete selected, and one inline confirmation deletes the whole selection. Deleting a gallery tile removes all versions in that stack; multi-selection counts stacks, not
hidden versions. Both confirmation flows explain this. Deleting removes the
versions from the playground state and, unless another remaining version still reuses the same asset,
deletes the generated image itself through DELETE /api/generated-images/:id (database row, workspace
file, and the shared media-library entry). The state is saved once per batch; cancelling changes nothing,
and a failed delete keeps the versions that were not reached and shows an error. The iOS and macOS gallery
offers the same single and multi-select deletes with native confirmation dialogs.
Enter a prompt, optionally attach reference images, and generate the first image. The global prompt bar stays visible at the bottom for subsequent changes.
Right-click the image to open a point editor at that coordinate, with a coordinate hint, optional
feathered mask, and optional reference image. Each result replaces the canvas image. Version history
supports undo, redo, and restoring an earlier thumbnail.
The module API persists the current user's state through GET/PUT/DELETE /api/modules/image-playground/api/state, using per-user JSON files under the persistent workspace at
<workspace>/.clapilotaicore/module-state/image-playground/ (the workspace root comes from
CLAPILOT_WORKSPACE_DIR/OPENCLAW_WORKSPACE_DIR, default /app/workspace). State is never written
inside the module folder, because the bundled-modules tree is part of the Docker image and is replaced
on every update. Reopening the module displays the gallery from that state; a corrupt state file is treated as empty
and replaced by the next save. If the state request itself fails (for example during a container
restart), the module locks every editing control, shows a reload action, and never saves until a
later load succeeds, so a transient outage cannot overwrite the stored versions. The last debounced
save is flushed with a keepalive request when the module is closed. Image assets remain in generated_images;
GET /api/modules/image-playground/api/health checks module health.
Model selection
Like Whiteboard, the model picker reads /api/generated-images/models with cache: "no-store".
Selecting a model sends its exact model and optional provider_slug; Default sends neither.
Each catalog entry under Settings -> ClapilotAICore -> AI Media -> Usable image models carries a
use attribute: Text to image, Image edit, or Both (imageOperations: generate, edit,
both; absent means both). The web module and the Apple editor only offer models that fit the
current step: a fresh canvas lists text-to-image and both, an existing image lists edit and both.
The server enforces the same rule for the HTTP endpoints and the agent tools and answers with a
localized error instead of switching models. An explicitly selected provider is binding: edits are
never rerouted to another provider, and a provider without an /images/edits endpoint fails with a
clear error. When no model is selected, the starred default is used if it supports the operation,
otherwise the first catalog entry that does. Every generation or edit call is recorded in the
ClapilotAICore request log with kind image, the provider that actually received the request, and
the upstream base URL.
How point edits work
The design uses option a1: the selected coordinates are encoded in the provider prompt. In addition,
the client generates a PNG mask with a transparent, feathered circle and imports it as a generated-image
asset, then sends its ID as mask_image_id. Transparent pixels identify the editable region.
OpenAI API-key routing sends this as a real inpainting mask in the /images/edits mask field;
Codex/Gemini routing receives it as a guidance image, and xAI keeps the mask ahead of extra
references when its three-image input limit is reached. The mask is rendered at the exact pixel size
of the source image, because the OpenAI mask must match the image dimensions. Sources beyond the
browser canvas budget (longest edge over 4096 px or more than 16.7 megapixels) skip the mask with a
notice and rely on the coordinate hint alone. The mask can be switched off for each edit.
The coordinate hint remains part of the prompt.
Mask and reference uploads are imported with import_source set to image-playground-mask,
image-playground-reference, or image-playground-upload in the asset metadata, so throwaway masks
can be told apart from user drops and pruned later. Each attached reference is imported once and its
asset ID is reused across retries; a point mask is likewise reused while the base image and point stay
the same, so a failing provider call does not multiply stored assets. Cleanup itself is not part of this version.
Reference images
Multiple visual references use reference_image_ids on the existing generated-image endpoints.
For edits, the first source image is always the image being edited, followed by the extra references.
Both endpoints echo only the references the provider actually forwarded, so the client can tell when a
provider cap dropped one; the version history records that forwarded count, not the requested one.
Each version also stores the asset MIME type so downloads carry a matching file extension.
Generation with references uses the edit pipeline. Uploads are persisted through
POST /api/generated-images/import. Edits use size: "auto" to preserve the source ratio;
first generation can use an explicitly selected size.
Data visibility
Playground editor state remains personal. Generated module work products are now registered in the workspace-global media library, using authenticated asset references without public tokens. The library indexes existing generated versions from saved Playground histories and automatically registers new generation/edit results. Temporary masks, reference uploads, and personal chat-only generations remain outside the library.
Agent access
The native agent uses the existing images_generate and images_edit tools for image work.
Playground-specific state tools are intentionally not shipped in this version, and native tool arguments
are unchanged. Agent image work does not manage the playground's canvas state or version history.
Touch fallback
Long-press the image for 500 ms to open the point editor. The bottom global prompt bar stays available on touch devices as well.
Limits
- Up to 4 reference images for generation and 3 extra references for edits.
- Up to 20 MB per upload.
- Up to 80 versions retained in playground state.
Testing
Open Modules -> Image Playground. Generate an image with and without references, apply a global change, and try right-click and touch long-press point edits with the mask enabled and disabled. Check undo, redo, thumbnail restore, and persistence after reopening the module. Verify another user's state and image assets remain inaccessible.
Native Apple editor
The iPhone, iPad, and macOS apps include a native SwiftUI Image Playground entry when the module is installed. All Apple clients start with the same thumbnail gallery and New image tile; opening a tile shows the full editor with a Back to gallery action. The editor uses the same personal state endpoint and generated-image assets as the web module: model/provider selection, generation, global edits, reference imports, version selection, undo/redo, and image export work across clients. The native section publishes /modules/image-playground as its page route. While a request runs, the native editor shows the Clapilot loading animation behind the canvas (the current image fades, an empty canvas shows a "Generating image" label) instead of a blocking progress overlay; the composer and canvas stay locked until the request finishes.
Enable Edit region to paint an edit mask. On iPad, PencilKit captures Apple Pencil strokes; Draw with finger optionally enables finger input. On Mac, drag with the pointer. The brush width and clear-selection controls apply to the mask. Marks select the area to change; they are not drawn onto the output image. The native editor converts the marks into transparent regions of a source-sized PNG and submits mask_image_id through the existing edit endpoint. Provider-specific mask behavior and the 4096-pixel/16.7-megapixel limit described above still apply. Mask drawing is temporary and clears when switching images.
Native errors remain visible with a retry action. If saving history fails after generation, retry saves the returned asset without generating again. Do not navigate away before an unsaved-history error has been resolved. State is shared with the web editor using the existing last-write-wins contract; simultaneous edits in multiple clients are not merged.
Native chat/live page context includes view (gallery or editor) and the active image_id and prompt only while the editor is open. Returning to the gallery clears the active image context. Existing images_generate and images_edit tools remain available; agents do not operate the Pencil canvas or modify playground version history. This is the same intentional boundary as the web module.
Apple verification: build macOS and iOS Simulator targets, run ImagePlaygroundTests, then check generation/import, references, region edits, history restoration, and export against an authenticated instance. Physical Pencil pressure and palm rejection must be checked on an iPad with Apple Pencil.
Shared media library
The Media library action opens generated images and videos organized into folders and nested subfolders. Playground output defaults to Images/Image Playground; Social Media output defaults to Images/Social Media; completed Video Studio exports and scene clips appear under Videos/Video Studio. Folder names are library organization, not physical filesystem moves. Original source files stay in place, and moving an item never breaks existing project or history references.
The web editor, iOS, and macOS can select a library image and import it as an editable source. Library images can also be attached as references instead of a local upload: the prompt bar and the point editor each have a Reference image from media library button next to the local attach button on the web (the Apple attach menu offers Reference from media library). Library references behave exactly like local ones, are draft-only until generation, and count toward the same reference limit. Generated results are automatically registered through library_source: "image-playground" on the generate/edit APIs. Imported originals and temporary masks are not registered as generated output. Existing saved history is indexed when the shared library is browsed; historical images no longer referenced by saved history are not automatically made shared.
Every library tile also has a hover/focus Delete permanently control with an inline confirmation, and Select multiple enables checkboxes with a selection bar (Select all, Clear selection, Delete selected) that deletes the whole selection after one confirmation. Library deletion is real and workspace-global: the underlying module output is removed for everyone. Generated images lose their file, record, and the owner's playground gallery version; Video Studio scene clips lose their file and record and the scene falls back to frame-ready so it can be regenerated; rendered Video Studio videos lose the file. iOS and macOS offer the same single delete next to Open and a Select multiple mode with a bulk delete action.
Social Media can attach library images or videos to a draft. Livestream can import Video Studio output into its mounted media folder, then queue it through the existing explicit queue action. Neither selection nor folder organization publishes anything. File Explorer and Video Studio expose the same folder browser. The native Livestream and File Explorer surfaces use the existing module/web routes; there is no separate SwiftUI Livestream or File Explorer implementation.
Agents use media_library_list, media_library_update, and social_media_attach_media with source: "media_library" and asset_id. media_library_update with action: "delete_asset" or "delete_assets" performs the same permanent, workspace-global deletion and should only run after an explicit user request naming the items; the playground's own DELETE /api/generated-images/:id stays UI-only, like Video Studio video deletion. Personal chat generations can be explicitly shared using media_library_update with action: "share_image" and the generated image ID. List results are authenticated URLs, not public distribution links.
Image Playground generation jobs
Image Playground text-to-image requests (including reference images), global edits, and point edits now use short HTTP requests and a durable PostgreSQL job. Web, iOS, and macOS wait for status updates instead of keeping a provider request open through a proxy. Slow providers can finish after the proxy's request timeout. Reopening the Playground recovers recent workspace jobs and the latest older output explicitly shared to the Image Playground media-library folder; private chat images and reference/mask uploads are excluded. Existing gallery entries are deduplicated by asset ID. The personal canvas/history layout remains unchanged.
POST /api/generated-images/generateandPOST /api/generated-images/editaccept their existing fields plusasync: true(or?async=true),library_source: "image-playground", and a stablerequest_id(16–128 letters, digits, underscores or hyphens; UUID recommended). They return HTTP 202 with{ job: { id, status, kind, prompt, result } }before calling the provider;kindisgenerateoredit. Retries with the same creator/request ID return the original job and ignore replacement payloads. Edit jobs keepimage_id,reference_image_idsandmask_image_idin the stored request and run through the same edit executor as the synchronous endpoint.GET /api/generated-images/jobs?id=<uuid>returns{ jobs: [...] }withqueued,running,ready, orfailed. Onlyreadyhas a completedresult.assetandresult.markdown; failures haveresult.error. Responses are authenticated and uncached. Omittingidrecovers up to 80 recent Image Playground jobs, plus the latest legacy shared asset when capacity allows. Jobs/output are workspace-global; creator IDs are used for attribution and idempotency, not read filtering.- The web client polls every three seconds and retains the request ID across ambiguous POST failures. Poll/network/proxy failures keep it waiting instead of re-submitting a generation. Apple clients offer reload/recovery after connection errors and lock new generation until that recovery succeeds. All status/error UI is localized in German, English and Italian.
images_generatesupportsasync: truewithrequest_idfor Image Playground output, andjob_idfor subsequent status checks (the requiredpromptis ignored for status checks). Agent responses must describe queued/running jobs as still generating and only present a result once ready. Synchronous generation for private chat and Notes insertion keeps its existing contract, andimages_editstays synchronous for agents.
Deployment requires migration 323_generated_image_jobs.sql before updated clients/API.
Generation is executed after the HTTP response in the persistent Node server, protected
by a PostgreSQL session advisory lock. Polling also resumes queued jobs. A process restart
reconciles a running job against generated_images.metadata.generation_job_id; a saved
asset is recovered, otherwise the job fails with generation_interrupted instead of
silently repeating a potentially chargeable provider call. A server restart cannot resume
an upstream provider call whose result had not yet reached local storage.
Async Image Playground recovery compatibility
Generation accepts async: true in the body or ?async=true for existing clients.
Clients may send display_prompt separately from the provider prompt; recovery uses
that display text and removes the legacy reference-only instruction when absent.
Job reads remain workspace-global, but only the creator receives private reference
IDs and source paths. Other members receive the output asset without source IDs or
metadata. Web, iOS, and macOS restore completed jobs on opening the playground and
skip active workspace jobs during initial load so another user's provider cannot
lock the gallery. Reopen after completion to recover those jobs; newly started
jobs continue polling and retry transient connection failures with the same ID.
Gallery regression check: create two images, edit the first (including a point edit and an edit from an older version), and verify there are two tiles with the latest edit on top. Reopen the module and verify the stack count persists. Delete one stack and confirm the other remains; test multi-select and cancellation on web, iOS, and macOS.
