Video Studio

Bundled module for HTML product videos and AI storyboard videos — characters, templates, Bausteine, Feinschliff, and agent-driven rendering.

What it does

Video Studio is a bundled module for creating and watching Clapilot-branded demo videos. Videos are authored as HyperFrames HTML compositions (real clapilot-* UI rendered as live DOM — camera zoom, typing, animated logo intro/outro) and rendered server-side to MP4. You describe the video in an embedded chat; the agent composes it from reusable scene blocks (Bausteine), renders it in the background, and the finished MP4 appears in the module gallery.

How to open / enable it

  • Open Video Studio from the module navigation (/modules/video-studio).
  • The open view is URL-addressable: ?view=<gallery|templates|library|styles|characters|chat>, ?project=<slug> for an open Feinschliff editor project, and ?aiView=<create|board> plus ?aiProject=<id> for the AI storyboard views. The module mirrors its state into these query parameters shallowly (announced via the clapilot:shallow-url-change window event), so the Clapilot Tab Layout restores the exact view — e.g. an open storyboard — when switching back to a Video Studio tab, and the URLs work as deep links.
  • Video Studio is generally available to signed-in users when the bundled module is installed. The ADMIN_ONLY_MODULE_SLUGS mechanism remains in place for other modules, but video-studio is not in that set.
  • Module manifest: bundled-modules/video-studio/module.json — slug video-studio, entry index.html, renderer react (rendered by src/components/modules/video-studio-module.tsx; the legacy index.html iframe view is kept as a fallback), icon video (app-tile public/assets/module-app-icons/video-studio.png, nav public/icons/navigation/videostudio.png).
  • On iOS/macOS the native section (clients/apple/ClapilotApple/Sources/Clapilot/Views/VideoStudioView.swift plus VideoStudioComposerToolbar.swift, MainAppSection.videoStudio) appears in the side menu only when a probe of the module API succeeds (AppModel.refreshVideoStudioAvailabilityIfNeeded); an unavailable or uninstalled module stays hidden.

Speech models and scene voice-over

In Settings → ClapilotAICore → Audio, add multiple TTS and STT provider/model entries and mark one default per capability. Credentials stay with the existing provider configurations; add additional accounts in the Providers tab. Each TTS entry can store a voice. ElevenLabs voices load from the selected account, including all pages of its voice list. Compatible providers can expose /audio/voices; other providers keep an editable voice ID. Failed voice discovery shows an error and retry action.

A ready storyboard scene's Voiceover editor selects a configured TTS provider, model, and voice. Generating a preview saves the narration and selection in metadata.voiceoverText / metadata.voiceoverSpeech; synthesis honors that selection. The workspace default is used when the selection is cleared. The project's dubbing voice is separate and does not override narration. The same settings and scene controls are available in the shared iOS/macOS client. Review the audio preview before applying it over the scene's original audio.

Chat and live agents discover configured speech models with video_studio_list_voices(include_models=true) and fetch a selected provider's voices with provider_slug. They set voiceover_speech: {providerSlug, model, voice} with video_studio_update_scene or video_studio_voiceover_scene. Existing media_tts_speak and media_stt_transcribe accept the discovered provider/model overrides directly. HTML scene MP4s remain cached while narration, dubbing, or other scenes change.

Key workflows

Create a video (Neues Video)

Open the single "+ Neues Video" menu and choose one of two creation paths:

  • KI-Video opens the full-screen storyboard creation form over the gallery.
  • HTML-Video opens the chat-driven HTML/HyperFrames flow. Optionally drag a few Bausteine into the composer, toggle Voiceover/Musik, and describe the video (e.g. "Erstelle ein Video über den E-Mail-Flow."). The agent composes, renders (and muxes audio if requested), and the finished MP4 appears in the gallery.

The HTML-Video entry opens the full Clapilot chat in-module — it embeds the real ChatFloatingWidget (the same component as the floating dock: pet animation, tool-call log, response timing, timestamps, markdown), not a separate chat implementation. The widget is embeddable via an additive embedded config ({ sessionId, composerAccessory, onDidSend }): with no config it is the unchanged floating dock, and with the config it drops the floating shell and renders just the thread + composer pinned to a forced session. The composer accessory provides:

  • Dedicated session — the module gets-or-creates a stable, non-main "Video Studio" chat session (id chat-video-studio-<user>) via POST /api/chat/sessions { ensureScope: "video-studio" } (getOrCreateNamedChatSession), so video chat history stays separate from the user's Hauptchat.
  • Drag-and-drop Bausteine — the library blocks render as chips above the composer that can be dragged (or clicked) to pin them. Pinned slugs are published in the chat pageContext as selectedBlocks; on send the module prompt (src/app/api/chat/route.ts) instructs the agent to run compose-video.mjs with exactly those blocks, then the pins clear (onDidSend).
  • Voiceover & Musik controls — app-styled toggles with provider/voice and provider/model selectors, published in pageContext (voiceover = provider:voice, music = provider_slug:model):
    • Voiceover → media_tts_speak (provider openai | gemini | openai_compatible; voices openai alloy/echo/fable/nova/shimmer/onyx, gemini Kore/Puck/Charon/Fenrir/Leda, or the configured Spark/self-hosted TTS runtime).
    • Musik → livestream_generate_music; the web and Apple composers load the enabled provider/model choices from the curated AI Media music catalog when it is non-empty, otherwise they use the enabled provider's configured models and default.
  • AI Video provider controls use the same AI media video defaults as chat-generated videos (videos_generate/videos_status) and can select xai-grok-video (grok-imagine-video-1.5/grok-imagine-video) alongside Gemini/Kie video rows configured under Settings -> ClapilotAICore -> AI Media. Live Stream Studio-specific jobs still use livestream_generate_video.
  • Format controls — aspect ratio (16:9 / 9:16), a portrait fill mode (shown only for 9:16), and output resolution (1080p / 4K), published in pageContext as aspectRatio, portraitFill, resolution:
    • compose-video.mjs takes --aspect 16:9|9:16 and --portrait-fill frame|bleed. frame scales the whole 16:9 stage into the portrait width and centers it with the brand gradient above/below (works with every existing Baustein); bleed emits a bare 1080×1920 composition that the agent then re-lays-out per scene.
    • The render uses HyperFrames' --resolution preset matching the aspect: landscape / landscape-4k for 16:9, portrait / portrait-4k for 9:16 (the composition is unchanged; Chrome renders at a higher device-pixel-ratio, and the preset aspect must match the composition).

KI-Storyboard-Videos

AI studio data is workspace-global: every authenticated user sees and can edit or delete all KI storyboard projects, reusable characters, and character voices. The owner_user_id columns on these rows are kept purely as creator attribution (who created the row) and are never used as a read or write filter — the same model as the Wiki and Kunden modules. Generated scene clips and start frames are likewise readable by any authenticated user so a shared storyboard renders for non-creators.

The Galerie and Charaktere tabs expose the KI storyboard path beside HTML/HyperFrames product videos. The gallery toolbar has a search field on web, iOS, and macOS matching video title, file name, and KI project title/prompt. Unfinished or failed KI project cards show the start frame of the lowest-index scene with a frame as soon as it exists. The active query is published in chat pageContext as gallerySearch only while non-empty, alongside galleryFilter and gallerySort. There is no separate KI-Videos tab: the gallery combines rendered videos and unfinished or failed KI projects, with a compact Alle / HTML / KI filter. Every rendered video card shows the actual first video frame on web, iOS, and macOS, including older files without a saved thumbnail. Frames are extracted locally on demand and cached by file revision; portrait frames remain fully visible. Ready KI results also show a KI badge, and a secondary action that reopens the storyboard. Project cards show status, scene count, last update, and the error hint when generation failed. Unfinished and failed project cards can be deleted with an inline confirmation; active generation warns that pending results will be discarded. The storyboard header exposes the same project deletion for every status. Deleting a ready project removes its storyboard rows but keeps the final MP4 in the gallery.

Reusable characters hold a stable appearance prompt, canonical portrait, an optional uploaded voice sample, and an optional provider preset voiceId. When creating a character From image, drop one PNG, JPG/JPEG, or WebP onto the reference-image picker or select it with the existing picker. The web and native iOS/macOS editors highlight the drop target and keep the current selection if an unsupported, empty, or multiple-file drop is rejected. The selected image appears immediately in a local preview, can be replaced or removed, and is submitted only when saved. Use image keeps the original image bytes as the ready portrait; it does not call an image provider, derive an appearance description, resize, or re-encode the upload. A name is required; the description is optional. The Generate a new portrait with AI toggle is off by default for uploaded images. Enable it only to intentionally create a new portrait based on the reference. Prompt-based creation still generates a portrait. Changes to an original character's text keep its portrait unchanged; explicit portrait regeneration remains available.

The same original-image mode is available to chat/live agents through video_studio_create_character with image_id and generate_portrait: false (the default when an image is supplied). Explicit AI regeneration uses generate_portrait: true on creation or regenerate_portrait: true on an agent update. Upload errors retain their specific localized message. JPEG provider references carry a matching JPEG filename and MIME type. The voice sample can be uploaded or replaced in the character editor, accepts MP3, WAV, M4A, OGG, or FLAC up to 25 MB, and is stored under the character creator's directory in .clapilot/video-studio-voices/ on the shared workspace volume (the directory name is attribution only; any user can manage the sample). During AI scene generation, Clapilot matches the scene's explicit characters (or its script speakers). Provider-native voice conditioning is enabled by default per project in metadata.voiceConditioning: direct Spark/OpenAI-compatible video receives up to three workspace reference_audio_paths plus a speaker-reference prompt preamble; Kie Seedance receives up to three ordered, tokenized public input.reference_audio_urls when a public base URL is available; and xAI receives one top-level voice_id only when exactly one scene character has a preset. Unsupported or ambiguous provider paths omit voice conditioning. Creating a character from a prompt, opting into AI portrait generation, or explicitly regenerating a portrait persists the character immediately and builds its portrait in the background, so the editor resets for the next entry instead of waiting on image generation. Character cards show generating, ready, or failed, refresh every six seconds only while work is active, and offer a retry after failure. A generation left active for more than ten minutes (for example after a server restart) is reconciled to failed. Characters whose portraits are still being created remain selectable: frame generation continues without the optional portrait reference and the finished portrait appears in the catalog when ready. A project then turns a creative brief into a storyboard whose scenes contain a complete video prompt, a dedicated start-frame prompt, explicit dialogue lines with speaker assignments, duration, and character references. After review, Clapilot creates missing start frames, starts deterministic per-scene provider jobs, polls their real status, and normalizes and concatenates the finished clips with ffmpeg. Each completed concat creates a new numbered MP4/thumbnail pair; the latest version lands in the same Video Studio gallery as HTML videos while the storyboard retains playable version history.

Project progress advances server-side. An in-process reconciler started with the web server checks active projects every 60 seconds by default, updates scene state from persisted provider jobs, applies stale/stall failure rules, and claims final concatenation through the same single-winner transition used by the status API. Browser and native polling only observe that durable progress; closing a client or restarting it is not required to finish a project. Set CLAPILOT_VIDEO_STUDIO_RECONCILE_SECONDS=0 to disable the worker. The same reconcile path runs a generation watchdog with a CLAPILOT_VIDEO_STUDIO_GENERATION_STALE_MINUTES window (default 45, minimum 5, 0 disables it) so no project can stay in generating or concatenating indefinitely without a diagnosis (see Troubleshooting stalled generation).

Voice consistency

The storyboard voice menu controls two independent, composable metadata-only project settings. Provider voice references (metadata.voiceConditioning, default true) applies the native provider contracts described above so clips speak with the right voices natively. Replace voices (dubbing) (metadata.dubbing, strictly opt-in and default false; enabling it requires a configured ElevenLabs provider) is post-production voice conversion: after a scene reaches clip_ready, the scene clip's existing voice audio is extracted, converted through the shared voice-conversion engine (currently ElevenLabs speech-to-speech, src/lib/video-studio-voice-conversion.ts — future dubbing providers plug in there), and remuxed as <clip>-dub.mp4 whose audio track fully replaces the original. Dubbing never synthesizes speech from text — text-to-speech narration is the separate, additive per-scene voiceover feature below. The target voice per scene is the project dub-voice override (metadata.dubVoiceId) when set, otherwise the ElevenLabs voice assigned to the scene's speaking character (first speaker with an assigned voice, then any linked scene character with one); a scene with no resolvable target voice fails its dub with api.videoStudio.ai.error.dubVoiceRequired. Because voice conversion applies one voice per clip, the one-speaker-per-scene rule below keeps automatic dubbing unambiguous.

Dubbing does not change the scene's main status. metadata.dubStatus advances pending → dubbing → done|failed, and stores dubPath only after the remux is complete. Final concatenation waits for done on every scene that has dialogue and uses dubPath instead of the provider clip. The server reconciler claims pending work after restarts; a dubbing claim older than 15 minutes becomes failed. Resubmitting or regenerating a scene video clears all prior dub metadata so the new clip must be dubbed again. Native iOS/macOS storyboard menus show both project settings read-only; the web storyboard can change them while the project is editable.

Each clip_ready scene also has a user-driven Voiceover action on the web — an additive narration track, not a dub. The scene card's voiceover popover edits a per-scene Voiceover text field (persisted as metadata.voiceoverText via the scene PATCH/video_studio_update_scene contract); it is intended for non-conversation clips that should just carry some spoken text. Generate preview synthesizes that text as one narration track (using the project dub-voice override when set, otherwise the TTS runtime default voice — an ElevenLabs provider always wins for routing) and writes only a standalone M4A under the project's ai-build/dubs workspace directory. The asynchronous state is stored in scene.metadata.dubPreview as generating → ready|failed with mode=voiceover; a generating preview older than ten minutes fails during normal project polling. A ready preview can be heard alone or started approximately in sync with the muted scene video. Apply to scene mixes that exact preview OVER the clip's existing audio — the original track is kept, ducked to half volume underneath (clips without audio get the narration alone); the voice-over is always additive and never replaces the track — records metadata.dubApplied, and makes the mixed MP4 the scene's active clip and concat source. Replacing spoken voices is exclusively the dubbing (voice conversion) feature: the project-level dubbing pipeline above or the per-scene Stimme ändern action below. Applying to an already-ready project reopens it as storyboard_ready so the next generation creates a new final version. Discard removes only the preview file and dubPreview metadata; regenerating the scene video clears both preview and applied voice-over state. Native iOS/macOS continues to display the resulting active clip read-only.

Stimme ändern (ElevenLabs Voice Changer)

Every clip_ready scene additionally offers a Stimme ändern action (waveform icon) once an ElevenLabs provider is configured (Einstellungen → ClapilotAICore → Provider → ElevenLabs, API key; the legacy global ElevenLabs key keeps working as a fallback). The popover lists the account's own/custom voices (cloned, generated, professional) first and the provider default voices below. Applying a voice extracts the scene's current audio (the applied dub clip when a dub exists, otherwise the generated clip), converts it with the ElevenLabs speech-to-speech voice changer while preserving content and timing, remuxes the converted track over the clip, and stores the replacement under the project's ai-build/voice directory. The asynchronous state lives in scene.metadata.voiceChange (generating → ready|failed, stale after ten minutes); final concatenation prefers a ready voice-change clip over the dub/raw clip, and applying to a ready project reopens it as storyboard_ready. Original wiederherstellen deletes the converted clip and restores the previous audio. Characters can carry an assigned ElevenLabs voice (dropdown in the character dialog, visible once ElevenLabs is configured; stored as metadata.elevenLabsVoiceId/elevenLabsVoiceName): the change-voice popover preselects the scene speaker's assigned voice, video_studio_change_scene_voice resolves it automatically when voice_id is omitted, and the automatic project dubbing uses the same assigned voice as its per-scene fallback target. Stimme ändern and project dubbing are the same voice-conversion mechanism — Stimme ändern is the manual per-scene form, project dubbing the automatic pipeline form. Regenerating the scene video or applying a new dub clears the voice change. The same flow is available to agents via video_studio_list_voices and video_studio_change_scene_voice. ElevenLabs can also be selected as the workspace TTS and STT runtime (Einstellungen → ClapilotAICore → Audio) once the provider is configured.

Start- und Endbilder

Projects can opt into Start- + Endbilder (toggle in the creation form and in the storyboard audio/continuity settings, use_end_frames). When enabled, every AI scene additionally generates an end frame from its start frame — same location, lighting, and wardrobe, but the scene's final instant — stored in end_frame_image_id and shown as a small overlay on the scene card. Supported providers (Kie Kling 3.0 image_urls[1], Kie Seedance last_frame_url, Kie Veo FIRST_AND_LAST_FRAMES_2_VIDEO, and OpenAI-compatible last_frame) then receive first and last frames for each clip; other providers fall back to first-frame-only generation. Scene continuity also improves: a scene marked An vorherige Szene anknüpfen uses the previous scene's end frame (when present) as its visual reference instead of the previous start frame. End-frame generation is best-effort — a failed end frame degrades that scene to first-frame-only generation instead of failing it.

Hintergrundmusik

Projects can enable Hintergrundmusik in the storyboard settings panel (metadata.musicEnabled, PATCH/tool field music, default off). During bundling the server generates one music track for the entire stitched video through the enabled music-capable AI Media provider (Kie or Gemini), keyed by a hash of the effective prompt so an unchanged prompt reuses the already-generated track across new versions. The style prompt (metadata.musicPrompt, field music_prompt) is optional — when empty, an automatic instrumental prompt derived from the project prompt is used. ffmpeg then loops or trims the track to the video length with a two-second fade-out and mixes it under the existing audio, ducked by sidechain compression whenever voices or dubs speak (see Endbearbeitung below for the master loudness). Music generation is best-effort: while the track is still generating, bundling waits and retries on the next poll; a failed or stale (15 min) generation finishes the video without music instead of failing the project. A finished video can get music after the fact: the play menu's Musik generieren + zusammenfügen option (generate flag add_music) enables music, generates a fresh track, and bundles the cached clips with it into a new version without regenerating any clips. The storyboard settings panel shows compact title-only toggle rows; under the music toggle a Musik generieren button generates the track standalone (play + regenerate buttons replace it once the track is ready, streamed from the project music endpoint), and under the dubbing toggle a Dubbing generieren button re-dubs all clips in the background — its voice popover offers the assigned character voices (default) or one ElevenLabs voice for every scene (stored as metadata.dubVoiceId); it requires a configured ElevenLabs provider.

Endbearbeitung (finishing, clip checks, exports)

The storyboard header's wand button opens the Endbearbeitung panel. Its settings live in metadata.finishing (PATCH/tool field finishing, a partial object merged over the current values) and apply to the next final version:

SettingDefaultEffect
modebasicbasic stitches with ffmpeg; polished renders the picture as one HyperFrames composition with the overlays below.
transitionsonUses each scene's metadata.transition (the transition INTO the scene); off = hard cuts.
titleCard, headlines, captions, nameChips, logoOutrooffPolished overlays: project title over the first scene, each scene's metadata.headline, the script lines as captions, a name tag the first time a character appears, and a 3-second logo outro.
outroName, outroUrlemptyOutro text; empty uses the instance branding (NEXT_PUBLIC_BRANDING_*, logo from public/).
beatSync, musicBpmoff, 120Asks the music model for a steady tempo and trims AI shots so every cut lands on a beat.
loudnessLufs-14EBU R128 loudness of the master (-14 web/social, -16 App Store previews).
clipQa, clipReview, clipQaAutoRetryonChecks every generated clip, has a vision model review its contact sheet, and regenerates a failing clip once.
exportsnoneExport profiles rendered automatically after every final version.

Transitions. The storyboard model proposes a transition and an optional short headline per scene (cut for continuing shots; fade/dissolve for calm moments; flash for energy; wipes/slides for topic changes; fade through black for time jumps). The scene card edits both (PATCH /api/video-studio/ai/scenes/:id with transition {type, duration} and headline). Types: cut, fade, fadeblack, flash, wipeleft, wiperight, slideleft, slideright, smoothleft, circleopen, zoomin, dissolve; durations are 0.2–1.5 s and clamped to a third of the shorter neighbouring clip. The first scene always cuts in. ffmpeg chains xfade for the picture and acrossfade for the sound on one shared timeline; muted scenes are silenced over their visible span on that timeline.

Audio master. Every final version goes through one audio pass: the project music (looped, faded out) is ducked under the stitched audio with sidechaincompress, muted ranges are applied, and loudnorm brings the mix to the target loudness (true peak -1.5 dBTP).

Polished mode. The same timeline is written as a HyperFrames composition: each normalized clip is a muted <video> layer, transitions are GSAP tweens, overlays and the outro are HTML layers, and the stitched ffmpeg audio is muxed onto the render (padded under the outro). If the render fails, the version is finished with the plain stitch and the project banner says so.

Beat sync. The music's tempo, beat phase and downbeat are estimated from its onset envelope. Each AI shot keeps at least 60% of its length and is trimmed so the middle of its outgoing transition lands on a beat, preferring downbeats and moments of peak motion in the shot (ffmpeg scene scores). Scenes with dialogue and HTML scenes are never trimmed. An estimate with low confidence leaves the timing unchanged.

Clip checks. When a clip is ready, the scene stays clip_generating while it is checked: duration against the request, hard cuts inside the clip (select='gt(scene,0.3)'), black stretches (blackdetect), frozen picture (freezedetect), and a 4×2 contact sheet that the vision model (the rag_document_vision_model, then video_storyboard_model, then agent:main) reviews for prompt mismatch, extra people, identity changes, logos or text, and deformities. A failing clip is regenerated once; a clip that fails again is kept and flagged. The verdict is stored in scene.metadata.clipQa (checking, passed, retrying, flagged with issues[]), shown on the scene card, and the sheet is served by GET /api/video-studio/ai/scenes/:id/qa-sheet. If the check itself cannot run, the clip counts as passed.

Exports. Each final version can carry deliverables in metadata.versions[].exports, rendered in the background from the master into video-studio/exports/<slug>/v<N>/ (removed when the version is deleted):

ProfileOutput
webH.264 CRF 23, +faststart, project size, -14 LUFS.
social_vertical1080×1920; a landscape master is shown whole over a blurred fill.
square1080×1080 with the same blurred fill.
app_store_preview886×1920 (portrait) or 1920×886 (landscape), center-cropped, 30 fps, at most 30 s with a fade-out, H.264 10 Mbit/s, AAC 256 kbit/s, -16 LUFS.
postersThree JPEG stills from the opening, middle and ending.

Start exports from a version tile's Exportieren… menu, with POST /api/video-studio/ai/projects/:id/versions/:version/exports {profiles}, or with the agent tool video_studio_export_version. Download with GET …/exports?profile=<profile>&file=<index>&download=1.

Blocks designed natively for portrait can declare "aspect": "9:16" in their block.json; when a 9:16 project renders a block scene whose blocks all declare 9:16, the composition bleeds full-frame (portraitFill: "bleed") instead of framing the 16:9 stage with brand bars. The storyboard's "Neue Szene" tile opens a three-option menu: AI scene, building-block scene, or linking an existing gallery video as a scene.

Storyboards additionally follow a one-speaker-per-scene rule: multiple characters may be visible in a scene, but all script lines of one scene come from a single speaker; conversations are split across consecutive scenes. This keeps per-scene dubbing and voice changes unambiguous.

  • Characters first — create recurring people in Charaktere, then attach their ids to the project and scenes. Assigning a character directly to a scene automatically links that character to the project as well, including assignments made through the Video Studio agent tools. Frame generation receives their stored visual identity and portrait references so a series does not rely on names alone.
  • Storyboard review — selecting a KI project in the gallery opens the scene board as a full-screen gallery takeover. Prompts and dialogue stay editable per scene, and a start frame can be generated again or edited while it remains linked to the scene. Editable storyboards can append, delete, and drag-reorder scenes on the web; the native list provides add, swipe-delete, and move actions. Structural changes are blocked during storyboard or clip generation/concatenation. Changing a ready storyboard returns it to storyboard_ready, after which the existing new-version generation flow creates the next output.
  • Project aspect ratio — the storyboard header always shows 16:9 or 9:16. Web users can change it while the project is draft, storyboard_ready, or failed; the native storyboard header displays the stored value read-only. Existing frames remain reviewable and should be regenerated for the best composition after a change.
  • Per-scene clip preview — as soon as an AI scene reaches clip_ready, its card and native detail view can play the authenticated generated clip over the start-frame poster. This makes each moving scene reviewable before the final concatenated gallery video is ready.
  • Scene continuity — each AI scene after scene 1 can be marked as continuing the immediately previous setting. Its start-frame generation then uses the previous AI scene's ready start frame as the first environment reference for location, lighting, props, and wardrobe, followed by portraits of characters in the current scene. Image edits keep the current frame first and add the previous frame when the four-reference limit allows. HTML scenes and missing or failed previous frames break the chain without blocking generation. Reordering does not rewrite this flag: continues_previous stays attached to the moved scene, so “previous” means whichever scene precedes it in the new order.
  • HTML scene mixing — a scene can reference an existing gallery slug or select ordered Bausteine directly from the library. The picker shows looping previews, metadata duration defaults, curated per-block text fields, and an AI content prompt for every selection. The prompt is disabled with an explanation only for an explicitly slotless block. Rendering uses the project aspect through the existing module manifest/HyperFrames handler and saves a versioned per-scene MP4 immediately in the storyboard. Both block scenes and linked gallery videos have playable previews. Saving unchanged blocks reuses the existing MP4 and preserves applied audio; changed selections, duration, content, or aspect render a new clip, and a missing MP4 is repaired. Dubbing, voice-over text edits, and regenerating another scene keep the saved HTML visuals. Final assembly reuses those clips and their normalized intermediates. Re-rendering asks for confirmation. Native clients identify both kinds and play ready clips, while creation/editing is web+agent only.
  • Provider and model — a project stores its exact video provider and model. If neither is selected, creation uses the enabled video provider carrying the configured default model (or the first enabled video provider). Explicit choices are preserved; generation does not silently replace them with another default.
  • Start-frame image model — creation stores the selected image provider/model in metadata.imageModel. The project header can change that selection, and frame regeneration can override it for one call. Generation and edit requests pass the pair through the normal image resolver; a selected model that cannot edit falls back to an edit-capable configured provider without changing the stored project choice. Every generation/edit prompt also states the project orientation explicitly. Clapilot probes the returned image with ffprobe; a frame outside the five-percent canonical-ratio tolerance is scaled to cover and center-cropped with ffmpeg (never padded) to 1536x864 for 16:9 projects or 864x1536 for 9:16 projects (image models are asked for their nearest size, 1536x1024/1024x1536, which video models would otherwise stretch). Reference images go the other way: before an edit request the first reference (portrait or character image) is extended to the project's aspect ratio with a blurred, darkened copy of itself behind it, so the model receives the whole person instead of a crop. Frame prompts forbid text overlays and brand logos or trademarks. The corrected PNG is stored as a new generated_images asset whose metadata records aspectNormalizedFrom with the original provider asset id, and only that corrected asset is linked to the scene.
  • Failed start frames and resuming — a start frame that the image provider rejects marks only that scene failed with a localized "start frame generation failed" message that wraps the provider's concrete error (for example an OpenAI-Codex HTTP 400 body); HTML scenes and already rendered clips are untouched. Project creation, video_studio_create_ai_project, video_studio_start_generation, and video_studio_get_ai_project report such scenes explicitly (failedScenes[] plus a message naming the scenes) instead of claiming success. Starting generation on that project first claims it from storyboard_ready/failed into storyboard_ready (a project that is still generating is rejected before any scene is changed), then regenerates the missing start frames in the background with the stored image selection and chains the clip fan-out behind them, so an existing project is resumed after a provider fix rather than duplicated and the request never blocks on image calls. With retryFailedOnly, only failed scenes get a new frame; a separately added pending scene is left alone. While frames regenerate the project stays storyboard_ready; the fan-out claims generating afterwards and a scene whose frame still fails ends the project as failed once no clip is in flight. OpenAI-Codex image generation picks its Responses host chat model from the provider's configured Codex models and the shipped defaults, skipping models the ChatGPT-account Codex backend rejects as unsupported (detected from the rejection wording about the model, whether or not the body echoes the model name).
  • Final-video versions — every concat writes <output>-v<N>.mp4 and <output>-v<N>.jpg. Ordered entries in metadata.versions record the duration and exact video provider/model, while the legacy final-path columns keep pointing to the latest file. The web storyboard can play and delete non-latest versions; native clients list and play every version but intentionally leave deletion to the web app. An existing unversioned final is recorded as v1 when the next concat writes v2.
  • Storyboard model — app_settings.video_storyboard_model optionally selects the text model that produces the strict storyboard JSON. An empty value falls back to agent:main; the model actually used is stored on the project.

Project state machine:

draft → storyboard_generating → storyboard_ready → generating → concatenating → ready

Adding, deleting, or reordering a scene on a ready project transitions ready → storyboard_ready; the user then reviews the changed structure and uses the normal new-version generation flow.

failed is a terminal/error branch from storyboard generation, scene generation, provider-unreachable stall detection, or concatenation. A reviewed, failed, actively generating, or ready project can be restarted. Restart preserves scene start frames, detaches every prior AI clip, and submits every AI scene as a fresh provider request; for a ready project this is the Generate new version flow and earlier final files remain untouched. Cancelling active clip generation returns the project to storyboard_ready and resets in-flight scenes.

Scene state machine:

pending → frame_ready → clip_generating → clip_ready

failed records a missing frame, provider failure, or block-render failure. Linked HTML scenes become clip_ready after validation; block scenes use clip_generating → clip_ready|failed. A persisted block render older than 15 minutes without a result reconciles to a localized failed state.

Troubleshooting stalled generation

The project status endpoint polls provider jobs and then reconciles the persisted asset state even when one provider status request throws. Network errors are stored per generated-video asset with an error count, first-error timestamp, last check, and bounded retry time. The normal configured poll failure window still moves a persistently unreachable asset to failed. Independently, a clip_generating scene is marked failed with a localized provider-unreachable message once its asset has at least three poll errors spanning five minutes, so the UI never remains silently stuck. Use Neu starten / Restart (or video_studio_start_generation { restart: true }) after the provider recovers; this creates fresh provider jobs and may spend provider budget again.

A provider that answers every poll with "processing" but never finishes, a submitting process that died before a scene received its clip job, or a container restart during ffmpeg concatenation would otherwise leave a project in generating/concatenating forever without any error. The generation watchdog (part of every reconcile tick and status poll; window CLAPILOT_VIDEO_STUDIO_GENERATION_STALE_MINUTES, default 45) closes those gaps with a per-case diagnosis instead of a silent stall:

  • Stale clip job (clipStale): a clip_generating scene whose linked generated-video job is still unfinished and was submitted more than the window ago is marked failed. The message names the provider, model, and last provider job status. Job age is measured from the job's creation time, not from its last poll, so a job that keeps reporting "pending" still trips the watchdog.
  • Missing clip job (clipJobMissing): a clip_generating scene whose generated-video row no longer exists (for example deleted from the gallery) fails immediately.
  • Stalled scene (sceneStalled): an AI scene still in pending/frame_ready (or clip_generating without a job id) while both the project row and the scene row are older than the window is marked failed, because no worker will ever submit it. The project claim bumps updated_at, so a fresh fan-out with old untouched scene rows is never treated as a stall.
  • Interrupted concatenation (concatStale): a concatenating claim older than the window transitions to failed; the finished clips stay clip_ready and can be bundled again.
  • Project diagnosis (generationStale): when the watchdog failed at least one scene and nothing is in flight any more, the project becomes failed with a message naming the window and the number of affected scenes; the concrete cause is on each scene. A generating project without any scene fails the same way once the claim is older than the window.

Each watchdog action logs a structured video_studio_generation_watchdog event with reason (clip_job_stale, clip_job_missing, scene_stalled, concat_stale, no_scenes), the project/scene ids, and the provider job details where available. Afterwards the normal recovery actions apply: Fehlgeschlagene erneut versuchen / Retry failed resubmits only the failed scenes, per-scene regeneration rebuilds one clip, and Zusammenfügen / Bundle re-runs an interrupted concatenation. A retry-failed-only run also submits scenes that already have a start frame but never received a clip (frame_ready), so a partial retry cannot itself leave a scene behind.

The storyboard header model menu can override the stored video provider/model at generation start or restart. The selection is persisted before scene requests are submitted, and scene durations are normalized again for the selected model. A ready or failed scene can also be regenerated individually after budget confirmation; the project returns to generating, then the normal single-winner reconciler appends a new numbered final and updates the gallery pointer.

Install content from the bundle store

The studio ships general: only the generic skill engine (scripts/ + a generic SKILL.md) is seeded into new instances. Branded scenes, templates, base shells and design styleguides live in a bundle catalog that ships with the module but is not auto-seeded:

  • catalog: bundled-modules/video-studio/bundles/<id>/ — each with bundle.json (id, name, description, version, includesBase, blocks[], templates[], styles[], extras[], legacyPaths[]) + blocks/ (incl. _base) + templates/ + the extras (styles/, agents/).
  • shipped bundles:
    • clapilot-video-bausteine — the 6 Clapilot Bausteine (brand-intro, brand-outro, email-workspace, ai-draft, decision, chat-scene), 2 templates, the shared base shell, and the clapilot design style.
    • nordlicht-dark — a dark studio set (aurora-intro, aurora-metrics, aurora-outro) with its own base shell and the nordlicht design style.

A Store button (bottom-right on the Bibliothek/Stile tabs) opens a popup listing bundles with install/uninstall (each row shows the shipped Bausteine/Vorlagen/Stile counts). Install copies the bundle's blocks/_base/templates and each extras item into the seeded skill dir; uninstall removes them plus any legacyPaths left by older bundle layouts. A bundle reads as installed when its blocks are present in the workspace, so existing live instances (already seeded before the store existed) show it installed automatically — nothing is removed from their volumes. New instances start empty and install on demand. Extensible for more bundles and, later, user uploads.

Migration note: the entrypoint seed (seed_workspace_tree_if_missing) syncs files present in workspace-seed but never deletes volume-only files. Moving content out of workspace-seed leaves existing volumes' content intact; the seeded SKILL.md is auto-synced to the generic version, and the styles layout lands on an existing instance when the bundle is (re)installed from the store (uninstall also cleans the legacy top-level STYLEGUIDE.md/references//assets/).

Bibliothek (templates, Bausteine, assets)

The Bibliothek tab surfaces the building-blocks library the agent composes from:

  • Vorlagen — full templates in skills/clapilot-video-styleguide/templates/ (render-as-is or customize).
  • Bausteine — reusable scene blocks in skills/clapilot-video-styleguide/blocks/<slug>/ (block.json + block.html + block.css + block.js, scoped under .b-<slug>). Each card shows an inline looping preview (a small muted preview.mp4 rendered per block). The agent compounds a video from them with scripts/compose-video.mjs --blocks a,b,c --name <slug>, and can create new blocks (which then appear here). Template blocks declare editable copy in block.json.textSlots[] as { key, label, find, defaultValue, scope, required?, multiline? }, where scope is html, js, or both. The block picker shows the label and defaultValue, and persists edits through the per-scene { find, replace, scope? } compose contract. The composer replaces the first literal occurrence in the selected scope; omitted override scopes still target both sources for backward compatibility. Intentionally slotless blocks declare an explicit empty array. Third-party blocks with no textSlots key derive up to twelve deduplicated slots from visible HTML text nodes. See the skill's SKILL.md → Bausteine. Regenerate previews after adding/changing a block with scripts/render-block-previews.mjs (runs in the runtime container — composes each block solo, renders, then downscales to a 480p/15fps muted clip into blocks/<slug>/preview.mp4).
  • Assets — shared images (the Clapilot logo) from blocks/_base/assets/ plus custom uploads: drag images (or use the upload tile) into the assets grid to add your own logos/graphics. Custom assets live in skills/clapilot-video-styleguide/custom-assets/ (outside the bundle dirs, so store install/uninstall never touches them), are badged Eigene with a delete action, and compose-video.mjs copies them into every composed project's assets/ — scenes reference them as assets/<file> (the agent is told about them in the module prompt). The grid additionally lists the workspace-global styleguide brand assets (Settings → Styleguide, stored under <workspace>/.clapilot/styleguide-assets/) badged Brand; they are read-only here (managed via /api/styleguide/assets) and compose-video.mjs copies them into composed projects alongside the custom uploads.

Design styles (Stile tab)

Design styles are first-class entities at skills/clapilot-video-styleguide/styles/<slug>/:

  • style.json — { slug, title, description, source, version } (source = shipping bundle id or custom).
  • STYLEGUIDE.md — the style's design rules; the agent reads and follows it when the style is selected.
  • optional tokens.css (inlined deterministically via compose-video.mjs --style <slug>), references/, assets/.

A style may also ship a human-facing HTML reference page under references/*.html (the Clapilot style ships clapilot-video-styleguide.html); the Stile card then shows an HTML-Referenz button that opens it in a sandboxed iframe popup, streamed via style-reference.

Surfaced in the UI as the Stile tab (cards with source badge and a rendered styleguide popup). The "Neues Video" composer has a Stil selector (defaults to the first installed style) published as pageContext.style; the chat route instructs the agent to read + follow that style's STYLEGUIDE.md. Bundles ship styles under their styles/ dir; new styles are created by the agent in chat (writing styles/<slug>/style.json + STYLEGUIDE.md with source: "custom"), and new Bausteine are tagged with their style via block.json.style (shown on the Bibliothek cards).

Fine-tune a video (Feinschliff editor)

Gallery videos with a matching project show a pencil action that opens the Feinschliff editor — a native, in-module rebuild of the essential HyperFrames-Studio workflow (no preview server, no proxy, docker-friendly):

  • Preview + transport — the project's index.html loads in a same-origin iframe (served via project/…); because HyperFrames compositions expose a paused, seek-safe GSAP timeline at window.__timelines, the editor drives play/pause/seek directly for frame-accurate scrubbing, with a time ruler and clickable scene lanes.
  • Scene edits — per-scene duration inputs and reorder controls, plus per-scene text fields (leaf text nodes).
  • Post-production voiceover — finished block-based projects expose a scene-level voiceover editor in Feinschliff. The initial action synthesizes one TTS file per scene through the configured Clapilot TTS runtime (Gemini/Kore by default). Text and voice remain editable; regenerating one scene only calls TTS for that segment and remuxes audio onto the existing MP4 without rendering the visual composition again. Audio is aligned to the scene start, padded when short, and moderately accelerated (up to 1.35x) then clipped when it exceeds the scene. Voiceover metadata and relative segment paths are stored under voiceover in compose.json. Choosing an uploaded character voice sends that character as voice_character_id, locks the provider to the OpenAI-compatible Spark TTS path, and uses the same sample as the ref_audio voice-clone reference for every synthesized scene segment.
  • Persistence model — compose.json (emitted by compose-video.mjs for every composed project) is the single editable source: { blocks, durations, aspect, portraitFill, style, textOverrides }. Saving writes the manifest and recomposes index.html from it (--manifest); text overrides are first-occurrence string replacements per scene, re-applied on every recompose. Duration semantics: shorter scenes clip at their end, longer ones hold the final state (tween offsets stay local to the scene).
  • Re-render — render { fromProject: true } renders the recomposed project over the existing MP4.
  • Legacy migration — composed projects from before the manifest existed are auto-migrated on first open: project-manifest parses blocks/durations/aspect/style back out of index.html and diffs each scene against its pristine block to preserve agent-made text edits as textOverrides, then persists the compose.json.
  • Clip emulation — raw compositions stack all scenes absolutely (the HyperFrames runner normally toggles them per frame), so the editor applies clip windows itself on load/seek/playback — exactly one scene visible.
  • Hand-authored projects (no compose.json, not composer-shaped) open in preview-only mode (scrub + re-render, no scene edits).

Manage the gallery

Search the gallery by title, file name, or KI prompt; all search terms must match, ignoring case and diacritics, and clearing the search restores the filtered gallery.

The Alle / HTML / KI filter separates non-KI rendered videos from matched KI outputs and KI project cards. Rendered outputs match projects by the gallery slug and the project's outputSlug/finalVideoPath. Each video card exposes a trash action with in-place confirmation. video-delete removes only the rendered MP4 — the composition project stays on disk, so the video can be recomposed and re-rendered later. KI project deletion is separate: DELETE /api/video-studio/ai/projects/:id removes the project (any authenticated user may delete any project) and cascading storyboard rows. It never removes the finished gallery MP4. The native iOS/macOS gallery and storyboard use the same endpoint through destructive confirmation dialogs.

How the agent can drive it

When chatContext.moduleSlug === "video-studio", src/app/api/chat/route.ts injects a module system prompt instructing the agent to either compound a video from Bausteine (scripts/compose-video.mjs) or scaffold a template (scripts/new-clapilot-video.mjs), edit the German scene copy, lint, and render in the background with npx hyperframes render --output /app/workspace/video-studio/videos/<slug>.mp4 via exec_command (detached so it beats the exec time limit). The MP4 surfaces in the gallery on the next poll. UI selections (pinned Bausteine, voiceover/music, style, aspect/fill/resolution) reach the agent through the chat pageContext keys described above.

Vertonung pipeline — when the user requests a voiceover and/or music, the agent generates the audio with the first-party media tools (media_tts_speak → absolutePath; livestream_generate_music → audioPath, polling livestream_get_asset for the async kie-ai-music provider), then renders silently to projects/<slug>/silent.mp4 and muxes the tracks into the final videos/<slug>.mp4 with scripts/add-audio.mjs (voiceover at full volume, music looped + ducked, output trimmed to the video length). See the skill SKILL.md → Audio — Voiceover & Musik.

Neues Video chat blocks + format + audio compose-video.mjs project + compose.json hyperframes render detached, Chromium Gallery videos/*.mp4 TTS + Musik tools optional Vertonung add-audio.mjs mux

Configuration & limits

Runtime data and storage

All compositions and rendered videos live on the shared workspace volume so both the web container (module API) and the agent container (exec_command renders) see them:

  • rendered videos: /app/workspace/video-studio/videos/<slug>.mp4
  • AI project exports: /app/workspace/video-studio/exports/<slug>/v<N>/
  • AI project build files (normalized clips, clip-check contact sheets, polished compositions): /app/workspace/video-studio/projects/<slug>/ai-build/
  • composition projects: /app/workspace/video-studio/projects/<slug>/
  • character voice samples: /app/workspace/.clapilot/video-studio-voices/<owner-user-id>/<character-id>.<ext>
  • skill engine (always seeded): /app/workspace/skills/clapilot-video-styleguide/ — a generic SKILL.md (compose/render/audio mechanics, points to the installed bundle's STYLEGUIDE.md for the house style) + scripts/. No house style of its own. The content AND the design styleguide (base shell + Bausteine + templates + STYLEGUIDE.md + references/ + brand assets/ + agents/) are NOT seeded; they install from a bundle (see the bundle store above).
  • installed-bundle marker: /app/workspace/video-studio/installed-bundles.json.

The gallery can be filtered by video source and sorted by newest or oldest modification date, or alphabetically by title in either direction. The same controls and ordering are available in the web, iOS, and macOS clients.

Implementation: bundled-modules/video-studio/api/handler.mjs.

Runtime requirements

  • HyperFrames CLI is installed globally in the runtime image (Dockerfile global npm install -g).
  • The image already ships Node 22, ffmpeg, and Chromium; HyperFrames renders with system Chromium (--disable-dev-shm-usage, so the default 64 MB /dev/shm is sufficient).

Module API endpoints

Base: /api/modules/video-studio/api

MethodEndpointPurpose
GETlistList rendered MP4s in videos/ with title, size, mtime, url.
GETtemplatesList full video templates from the seeded skill (clapilot-feature-flow, clapilot-chat-interface).
GETblocksList composable scene blocks (Bausteine) from the skill blocks/; each block includes previewUrl when a preview.mp4 exists and normalized textSlots[] (key, label, exact compose find, displayed defaultValue, scope, required, multiline). First-party slots come from the bundled catalog so existing installs receive curation updates; a third-party manifest with no slots key falls back to visible HTML text extraction, while an explicit empty array stays slotless.
GETassetsList video assets: bundle assets (blocks/_base/assets/) plus user-uploaded custom assets, each with a source flag.
GETasset?path=<name>Stream an asset image (custom assets shadow bundle names).
POSTasset-uploadMultipart upload (files) of custom images (PNG/JPG/WEBP/SVG, ≤20 MB each) into custom-assets/ — sanitized + deduped names; the Bibliothek exposes it via drag-and-drop + a file picker.
POSTasset-deleteBody { name } — delete a custom asset (bundle assets are managed by the store).
GETblock-preview?slug=<slug>Stream a block's looping inline preview MP4 (blocks/<slug>/preview.mp4, HTTP Range, short-cached).
GETfile?path=videos/<name>Stream an MP4 (supports HTTP Range) with the session cookie.
GETstylesList installed design styles ({ slug, title, description, source, version, hasGuide }; legacy top-level STYLEGUIDE.md is synthesized as a legacy style).
GETstyle?slug=<slug>Style detail incl. the full STYLEGUIDE.md text (guide).
GETstyle-reference?slug=&file=Stream a style's bundled HTML reference page (e.g. the Clapilot video styleguide HTML) for the UI iframe popup.
GETbundlesList the bundle catalog with install status ({ id, name, description, version, blockCount, templateCount, styleCount, installed }).
POSTbundles/installBody { id } — copy a bundle's blocks/_base/templates into the skill dir.
POSTbundles/uninstallBody { id } — remove a bundle's blocks/_base/templates from the skill dir.
GETproject/<slug>/<...file>Serve a project file (index.html + relative assets) for the Feinschliff editor's same-origin iframe preview.
GETproject-manifest?slug=The project's editable compose.json ({ hasManifest, manifest }; hand-authored projects report hasManifest: false).
POSTproject-editBody { slug, blocks?, durations?, textOverrides?, title? } — merge into compose.json and recompose index.html via compose-video.mjs --manifest.
POSTrenderRender { name, html }, { name, template }, { name, blocks, durations?, textOverrides?, aspect?, style? }, or { name, fromProject: true }. Block input writes the normal manifest and invokes compose-video.mjs --manifest; per-block picker/agent overrides are flattened to the manifest's scene-indexed { find, replace } entries before this call. Every form then runs HyperFrames into videos/.
POSTvideo-deleteBody { name } — delete a rendered MP4 from the gallery (the project stays for re-rendering).

Apple client (iOS/macOS)

  • Galerie / Vorlagen / Bibliothek — native grids/lists over the module API. MP4s and the looping Baustein previews are fetched authenticated through ClapilotAPI.fetchVideoStudioMedia (session cookie), cached as temp files (VideoStudioMediaCache), and played with AVPlayer/AVPlayerLooper.
  • AI project status — the Galerie shows a compact native strip for active projects and recently failed projects. It polls the workspace-global project list only while work is in progress. Native clients create AI projects (VideoStudioAICreationView.swift), manage characters with portraits and voice samples (VideoStudioCharactersView.swift), and edit the storyboard (VideoStudioStoryboardView.swift): add, delete and reorder scenes, edit prompts and scripts, transitions and headlines, regenerate frames and clips, start or cancel generation, set the finishing options, and create and download exports.
  • Mixed-scene parity — native storyboard rows show linked-gallery versus block-rendered kind badges, continue polling in-flight block renders, and play authenticated ready clips. Authoring remains web-and-agent only.
  • Neues Video — a fourth segmented tab that embeds the real chat: ChatView(presentation: .embedded) — a chrome-free presentation of the standard chat (full markdown/canvas rendering, tool-call log, attachments, the real composer) — pinned to the dedicated session (AppModel.activateVideoStudioChatSession, POST /api/chat/sessions { ensureScope: "video-studio" }; the previously active session is restored when leaving the tab). The video controls (tap-to-pin Baustein chips, voiceover/music menus, 16:9/9:16 + Füllung + 1080p/4K) render as a composerAccessory (VideoStudioComposerToolbar.swift) above the composer and publish via the live context. Every Apple text-chat send attaches the client context (AppModel.currentClientContextPayload, snapshotted onto the queued message at submit time to avoid racing UI state) — web parity for module-aware prompts.

Troubleshooting

  • compose-video.mjs --list shows no Bausteine / the Bibliothek is empty — new instances start without content; install a bundle from the Store popup first.
  • Portrait (9:16) output does not match the selected fill — aspect ratio and resolution are deterministic, but the fill style is agent-guided: in practice the agent often prefers to adapt the scene layout to fill the portrait frame (stacking the 16:9 columns) even when frame is selected — this usually looks more native, though quality can vary. The strict branded frame wrapper is always available via compose-video.mjs --portrait-fill frame; nudge the agent with a follow-up if you want it.
  • The video does not appear immediately — renders run detached in the agent container and typically take ~30–60 s; the gallery picks the MP4 up on its next poll.
  • The module is missing on iOS/macOS — the native section is hidden when the module API probe returns 403 (non-admin users).

File context menu

Right-click a rendered video or storyboard for Rename, Copy, Share, Download and Delete. Sharing offers Email and Chat (Channel or Direct message); it opens a draft and does not send. Videos download in their original format, storyboards as JSON. Video rename preserves file paths used by project versions. Copying a storyboard creates a new editable storyboard without starting generation.

The Apple clients use native context menus (right-click on Mac, long-press on iPhone/iPad). Sharing opens the authenticated Clapilot review/composer flow. Channel and direct-message sharing links to the original authenticated module/item and retains its permissions; no export or public link is created. Email continues to use a revocable public document link. No message is sent automatically.

Internal links in chat automatically show authenticated inline media or a clickable reference to the original item. Videos have playback controls, images have previews, and other documents/module items use document-style references. This also applies to existing messages and the Apple clients; no exported copy or public link is generated.

Shared generated media

The Media library action browses workspace-global generated output, with folder selection, nested folder creation, search, previews, and moving assets between folders. Existing files are referenced in place. Image Playground results can be reused in Social Media; Video Studio exports and completed scene clips can be imported into Social Media or Livestream. Livestream imports copy into the existing streamer media mount and do not enqueue or start playback. Social Media imports attach to the current draft and do not publish it.

The same library is available in web, iOS, and macOS. See Image Playground for visibility rules and agent access. Deployment requires 311_media_library.sql. Source removal can make a library reference unavailable; moving a library item never moves the source file.

Explicit scene silence and STS reset

Use video_studio_mute_scene (or POST /api/video-studio/ai/scenes/:id/audio with mode=muted) when a scene must be fully silent. This removes embedded source speech as well as applied TTS/STS overlays, without changing the source file or any previous final version. It clears active narration state and returns the silent preview on the next project read. Merely clearing the script or using change_scene_voice(reset=true) cannot remove embedded audio; the latter only undoes the STS conversion.

The project must be idle. Muting preserves actual source video timing: a 3.008s HTML source requested as 5s remains 3.008s, without a frozen tail or spilled narration. bundle_only=true stitches current clips into a new version, without provider, TTS, STS or music generation, and honors mute even under an existing project music bed. V1/V2 remain available; production output and audio QA must be explicitly performed after deployment, never inferred from tool success. Originals/derived files are retained; this operation does not promise an automatic unmute or speech/music separation.

The operation is available to chat/Live/CLI agents. Web and Apple clients use the same updated scene clip URL for preview; no new client-only editor control is introduced. Bearer access uses the same workspace member/role boundary as browser access, including specialist and automation sessions with a verified actor. No project-owner impersonation or public media exception is needed.

Scene mute lifecycle

Mute applies to the current scene clip. Regenerating/replacing an AI clip or HTML source clears the override; editing script or voiceover text also clears it so narration can be generated again. There is no standalone unmute mode; STS reset only discards voice conversion. Existing final versions remain unchanged. A missing silent derivative returns 404. Bundle-only suppression applies only during an active generation/concatenation run and does not suppress later narration edits. Web, iOS, macOS and agent operations share these server rules; no client controls change.

Automatische Videotitel

Die Galerie zeigt zuerst einen manuell vergebenen Titel. Andernfalls verwendet sie den Titel der Komposition oder des HTML-Dokuments; bei älteren Kompositionen mit technischem Standardtitel wird die sichtbare Überschrift verwendet. KI-Videos übernehmen den inhaltlichen Storyboard-Titel beim Export. Ein bereits vergebener Projekttitel bleibt bei der Storyboard-Erstellung erhalten. Dateinamen und Projektverknüpfungen bleiben unverändert. Diese Auflösung gilt gemeinsam für Web, iOS, macOS und den Modul-API-Zugriff durch Agenten. Ohne verfügbare Inhaltsmetadaten bleibt der bisherige Dateiname der Fallback.