Video Studio
Bundled module for HTML product videos and AI storyboard videos — characters, templates, Bausteine, Feinschliff, and agent-driven rendering.
What it does
Video Studio is a bundled module for creating and watching Clapilot-branded demo videos. Videos are
authored as HyperFrames HTML compositions (real clapilot-*
UI rendered as live DOM — camera zoom, typing, animated logo intro/outro) and rendered server-side to MP4.
You describe the video in an embedded chat; the agent composes it from reusable scene blocks (Bausteine),
renders it in the background, and the finished MP4 appears in the module gallery.
How to open / enable it
- Open Video Studio from the module navigation (
/modules/video-studio). - The open view is URL-addressable:
?view=<gallery|templates|library|styles|characters|chat>,?project=<slug>for an open Feinschliff editor project, and?aiView=<create|board>plus?aiProject=<id>for the AI storyboard views. The module mirrors its state into these query parameters shallowly (announced via theclapilot:shallow-url-changewindow event), so the Clapilot Tab Layout restores the exact view — e.g. an open storyboard — when switching back to a Video Studio tab, and the URLs work as deep links. - Video Studio is generally available to signed-in users when the bundled module is installed. The
ADMIN_ONLY_MODULE_SLUGSmechanism remains in place for other modules, butvideo-studiois not in that set. - Module manifest:
bundled-modules/video-studio/module.json— slugvideo-studio, entryindex.html, rendererreact(rendered bysrc/components/modules/video-studio-module.tsx; the legacyindex.htmliframe view is kept as a fallback), iconvideo(app-tilepublic/assets/module-app-icons/video-studio.png, navpublic/icons/navigation/videostudio.png). - On iOS/macOS the native section (
clients/apple/ClapilotApple/Sources/Clapilot/Views/VideoStudioView.swiftplusVideoStudioComposerToolbar.swift,MainAppSection.videoStudio) appears in the side menu only when a probe of the module API succeeds (AppModel.refreshVideoStudioAvailabilityIfNeeded); an unavailable or uninstalled module stays hidden.
Speech models and scene voice-over
In Settings → ClapilotAICore → Audio, add multiple TTS and STT provider/model entries and mark one default per capability. Credentials stay with the existing provider configurations; add additional accounts in the Providers tab. Each TTS entry can store a voice. ElevenLabs voices load from the selected account, including all pages of its voice list. Compatible providers can expose /audio/voices; other providers keep an editable voice ID. Failed voice discovery shows an error and retry action.
A ready storyboard scene's Voiceover editor selects a configured TTS provider, model, and voice. Generating a preview saves the narration and selection in metadata.voiceoverText / metadata.voiceoverSpeech; synthesis honors that selection. The workspace default is used when the selection is cleared. The project's dubbing voice is separate and does not override narration. The same settings and scene controls are available in the shared iOS/macOS client. Review the audio preview before applying it over the scene's original audio.
Chat and live agents discover configured speech models with video_studio_list_voices(include_models=true) and fetch a selected provider's voices with provider_slug. They set voiceover_speech: {providerSlug, model, voice} with video_studio_update_scene or video_studio_voiceover_scene. Existing media_tts_speak and media_stt_transcribe accept the discovered provider/model overrides directly. HTML scene MP4s remain cached while narration, dubbing, or other scenes change.
Key workflows
Create a video (Neues Video)
Open the single "+ Neues Video" menu and choose one of two creation paths:
- KI-Video opens the full-screen storyboard creation form over the gallery.
- HTML-Video opens the chat-driven HTML/HyperFrames flow. Optionally drag a few Bausteine into the composer, toggle Voiceover/Musik, and describe the video (e.g. "Erstelle ein Video über den E-Mail-Flow."). The agent composes, renders (and muxes audio if requested), and the finished MP4 appears in the gallery.
The HTML-Video entry opens the full Clapilot chat in-module — it embeds the real
ChatFloatingWidget (the same component as the floating dock: pet animation, tool-call log, response timing,
timestamps, markdown), not a separate chat implementation. The widget is embeddable via an additive embedded
config ({ sessionId, composerAccessory, onDidSend }): with no config it is the unchanged floating dock, and
with the config it drops the floating shell and renders just the thread + composer pinned to a forced session.
The composer accessory provides:
- Dedicated session — the module gets-or-creates a stable, non-main "Video Studio" chat session
(id
chat-video-studio-<user>) viaPOST /api/chat/sessions { ensureScope: "video-studio" }(getOrCreateNamedChatSession), so video chat history stays separate from the user's Hauptchat. - Drag-and-drop Bausteine — the library blocks render as chips above the composer that can be dragged (or
clicked) to pin them. Pinned slugs are published in the chat
pageContextasselectedBlocks; on send the module prompt (src/app/api/chat/route.ts) instructs the agent to runcompose-video.mjswith exactly those blocks, then the pins clear (onDidSend). - Voiceover & Musik controls — app-styled toggles with provider/voice and provider/model selectors,
published in
pageContext(voiceover=provider:voice,music=provider_slug:model):- Voiceover →
media_tts_speak(provideropenai|gemini|openai_compatible; voices openaialloy/echo/fable/nova/shimmer/onyx, geminiKore/Puck/Charon/Fenrir/Leda, or the configured Spark/self-hosted TTS runtime). - Musik →
livestream_generate_music; the web and Apple composers load the enabled provider/model choices from the curated AI Media music catalog when it is non-empty, otherwise they use the enabled provider's configured models and default.
- Voiceover →
- AI Video provider controls use the same AI media video defaults as chat-generated videos (
videos_generate/videos_status) and can selectxai-grok-video(grok-imagine-video-1.5/grok-imagine-video) alongside Gemini/Kie video rows configured underSettings -> ClapilotAICore -> AI Media. Live Stream Studio-specific jobs still uselivestream_generate_video. - Format controls — aspect ratio (
16:9/9:16), a portrait fill mode (shown only for 9:16), and output resolution (1080p/4K), published inpageContextasaspectRatio,portraitFill,resolution:compose-video.mjstakes--aspect 16:9|9:16and--portrait-fill frame|bleed. frame scales the whole 16:9 stage into the portrait width and centers it with the brand gradient above/below (works with every existing Baustein); bleed emits a bare 1080×1920 composition that the agent then re-lays-out per scene.- The render uses HyperFrames'
--resolutionpreset matching the aspect:landscape/landscape-4kfor 16:9,portrait/portrait-4kfor 9:16 (the composition is unchanged; Chrome renders at a higher device-pixel-ratio, and the preset aspect must match the composition).
KI-Storyboard-Videos
AI studio data is workspace-global: every authenticated user sees and can edit or delete all KI storyboard
projects, reusable characters, and character voices. The owner_user_id columns on these rows are kept purely as
creator attribution (who created the row) and are never used as a read or write filter — the same model as the Wiki
and Kunden modules. Generated scene clips and start frames are likewise readable by any authenticated user so a
shared storyboard renders for non-creators.
The Galerie and Charaktere tabs expose the KI storyboard path beside HTML/HyperFrames product videos.
The gallery toolbar has a search field on web, iOS, and macOS matching video title, file name, and KI project title/prompt.
Unfinished or failed KI project cards show the start frame of the lowest-index scene with a frame as soon as it exists.
The active query is published in chat pageContext as gallerySearch only while non-empty, alongside galleryFilter and gallerySort.
There is no separate KI-Videos tab: the gallery combines rendered videos and unfinished or failed KI projects,
with a compact Alle / HTML / KI filter. Every rendered video card shows the actual first video frame on web,
iOS, and macOS, including older files without a saved thumbnail. Frames are extracted locally on demand and cached
by file revision; portrait frames remain fully visible. Ready KI results also show a KI badge, and a secondary action that reopens the storyboard. Project cards show status, scene count,
last update, and the error hint when generation failed. Unfinished and failed project cards can be deleted with an
inline confirmation; active generation warns that pending results will be discarded. The storyboard header exposes
the same project deletion for every status. Deleting a ready project removes its storyboard rows but keeps the final
MP4 in the gallery.
Reusable characters hold a stable appearance prompt, canonical portrait, an optional uploaded voice sample, and an
optional provider preset voiceId.
When creating a character From image, drop one PNG, JPG/JPEG, or WebP onto the reference-image
picker or select it with the existing picker. The web and native iOS/macOS editors highlight the drop
target and keep the current selection if an unsupported, empty, or multiple-file drop is rejected.
The selected image appears immediately in a local preview, can be replaced or removed, and is submitted
only when saved. Use image keeps the original image bytes as the ready portrait; it does not call an
image provider, derive an appearance description, resize, or re-encode the upload. A name is required;
the description is optional. The Generate a new portrait with AI toggle is off by default for uploaded
images. Enable it only to intentionally create a new portrait based on the reference. Prompt-based
creation still generates a portrait. Changes to an original character's text keep its portrait unchanged;
explicit portrait regeneration remains available.
The same original-image mode is available to chat/live agents through video_studio_create_character
with image_id and generate_portrait: false (the default when an image is supplied). Explicit AI
regeneration uses generate_portrait: true on creation or regenerate_portrait: true on an agent update.
Upload errors retain their specific localized message. JPEG provider references carry a matching JPEG
filename and MIME type.
The voice sample can be uploaded or replaced in the character editor, accepts MP3, WAV, M4A, OGG, or FLAC up to
25 MB, and is stored under the character creator's directory in .clapilot/video-studio-voices/ on the shared
workspace volume (the directory name is attribution only; any user can manage the sample). During AI
scene generation, Clapilot matches the scene's explicit characters (or its script speakers). Provider-native voice
conditioning is enabled by default per project in metadata.voiceConditioning: direct Spark/OpenAI-compatible video
receives up to three workspace reference_audio_paths plus a speaker-reference prompt preamble; Kie Seedance receives
up to three ordered, tokenized public input.reference_audio_urls when a public base URL is available; and xAI receives
one top-level voice_id only when exactly one scene character has a preset. Unsupported or ambiguous provider paths
omit voice conditioning. Creating a character from a prompt, opting into AI portrait generation, or
explicitly regenerating a portrait persists the character immediately and builds its portrait in the background, so the editor resets for the next
entry instead of waiting on image generation. Character cards show generating, ready, or failed, refresh every
six seconds only while work is active, and offer a retry after failure. A generation left active for more than ten
minutes (for example after a server restart) is reconciled to failed. Characters whose portraits are still being
created remain selectable: frame generation continues without the optional portrait reference and the finished
portrait appears in the catalog when ready. A project then turns a creative brief
into a storyboard whose scenes contain a complete video prompt, a dedicated start-frame prompt, explicit dialogue
lines with speaker assignments, duration, and character references. After review, Clapilot creates missing start
frames, starts deterministic per-scene provider jobs, polls their real status, and normalizes and concatenates the
finished clips with ffmpeg. Each completed concat creates a new numbered MP4/thumbnail pair; the latest version lands
in the same Video Studio gallery as HTML videos while the storyboard retains playable version history.
Project progress advances server-side. An in-process reconciler started with the web server checks active projects
every 60 seconds by default, updates scene state from persisted provider jobs, applies stale/stall failure rules, and
claims final concatenation through the same single-winner transition used by the status API. Browser and native
polling only observe that durable progress; closing a client or restarting it is not required to finish a project.
Set CLAPILOT_VIDEO_STUDIO_RECONCILE_SECONDS=0 to disable the worker. The same reconcile path runs a generation
watchdog with a CLAPILOT_VIDEO_STUDIO_GENERATION_STALE_MINUTES window (default 45, minimum 5, 0 disables it)
so no project can stay in generating or concatenating indefinitely without a diagnosis (see
Troubleshooting stalled generation).
Voice consistency
The storyboard voice menu controls two independent, composable metadata-only project settings. Provider voice
references (metadata.voiceConditioning, default true) applies the native provider contracts described above so
clips speak with the right voices natively. Replace voices (dubbing) (metadata.dubbing, strictly opt-in and
default false; enabling it requires a configured ElevenLabs provider) is post-production voice conversion: after a
scene reaches clip_ready, the scene clip's existing voice audio is extracted, converted through the shared
voice-conversion engine (currently ElevenLabs speech-to-speech, src/lib/video-studio-voice-conversion.ts — future
dubbing providers plug in there), and remuxed as <clip>-dub.mp4 whose audio track fully replaces the original.
Dubbing never synthesizes speech from text — text-to-speech narration is the separate, additive per-scene voiceover
feature below. The target voice per scene is the project dub-voice override (metadata.dubVoiceId) when set,
otherwise the ElevenLabs voice assigned to the scene's speaking character (first speaker with an assigned voice, then
any linked scene character with one); a scene with no resolvable target voice fails its dub with
api.videoStudio.ai.error.dubVoiceRequired. Because voice conversion applies one voice per clip, the
one-speaker-per-scene rule below keeps automatic dubbing unambiguous.
Dubbing does not change the scene's main status. metadata.dubStatus advances
pending → dubbing → done|failed, and stores dubPath only after the remux is complete. Final concatenation waits for
done on every scene that has dialogue and uses dubPath instead of the provider clip. The server reconciler claims
pending work after restarts; a dubbing claim older than 15 minutes becomes failed. Resubmitting or regenerating a
scene video clears all prior dub metadata so the new clip must be dubbed again. Native iOS/macOS storyboard menus show
both project settings read-only; the web storyboard can change them while the project is editable.
Each clip_ready scene also has a user-driven Voiceover action on the web — an additive narration track,
not a dub. The scene card's voiceover popover edits a per-scene Voiceover text field (persisted as
metadata.voiceoverText via the scene PATCH/video_studio_update_scene contract); it is intended for
non-conversation clips that should just carry some spoken text. Generate preview synthesizes that text as one
narration track (using the project dub-voice override when set, otherwise the TTS runtime default voice — an
ElevenLabs provider always wins for routing) and writes only a standalone M4A under the project's ai-build/dubs
workspace directory. The asynchronous state is stored in scene.metadata.dubPreview as generating → ready|failed
with mode=voiceover; a generating preview older than ten minutes fails during normal project polling. A ready
preview can be heard alone or started approximately in sync with the muted scene video. Apply to scene mixes
that exact preview OVER the clip's existing audio — the original track is kept, ducked to half volume underneath
(clips without audio get the narration alone); the voice-over is always additive and never replaces the track —
records metadata.dubApplied, and makes the mixed MP4 the scene's active clip and concat source.
Replacing spoken voices is exclusively the dubbing (voice conversion) feature: the project-level dubbing pipeline
above or the per-scene Stimme ändern action below.
Applying to an already-ready project reopens it as storyboard_ready so the next generation creates a new final
version. Discard removes only the preview file and dubPreview metadata; regenerating the scene video clears both
preview and applied voice-over state. Native iOS/macOS continues to display the resulting active clip read-only.
Stimme ändern (ElevenLabs Voice Changer)
Every clip_ready scene additionally offers a Stimme ändern action (waveform icon) once an ElevenLabs provider is
configured (Einstellungen → ClapilotAICore → Provider → ElevenLabs, API key; the legacy global ElevenLabs key keeps
working as a fallback). The popover lists the account's own/custom voices (cloned, generated, professional) first and
the provider default voices below. Applying a voice extracts the scene's current audio (the applied dub clip when a
dub exists, otherwise the generated clip), converts it with the ElevenLabs speech-to-speech voice changer while
preserving content and timing, remuxes the converted track over the clip, and stores the replacement under the
project's ai-build/voice directory. The asynchronous state lives in scene.metadata.voiceChange
(generating → ready|failed, stale after ten minutes); final concatenation prefers a ready voice-change clip over the
dub/raw clip, and applying to a ready project reopens it as storyboard_ready. Original wiederherstellen deletes
the converted clip and restores the previous audio. Characters can carry an assigned ElevenLabs voice
(dropdown in the character dialog, visible once ElevenLabs is configured; stored as
metadata.elevenLabsVoiceId/elevenLabsVoiceName): the change-voice popover preselects the scene speaker's
assigned voice, video_studio_change_scene_voice resolves it automatically when voice_id is omitted, and the
automatic project dubbing uses the same assigned voice as its per-scene fallback target. Stimme ändern and project
dubbing are the same voice-conversion mechanism — Stimme ändern is the manual per-scene form, project dubbing the
automatic pipeline form. Regenerating the scene video or applying a new dub clears the
voice change. The same flow is available to agents via video_studio_list_voices and
video_studio_change_scene_voice. ElevenLabs can also be selected as the workspace TTS and STT runtime
(Einstellungen → ClapilotAICore → Audio) once the provider is configured.
Start- und Endbilder
Projects can opt into Start- + Endbilder (toggle in the creation form and in the storyboard audio/continuity
settings, use_end_frames). When enabled, every AI scene additionally generates an end frame from its start frame —
same location, lighting, and wardrobe, but the scene's final instant — stored in end_frame_image_id and shown as a
small overlay on the scene card. Supported providers (Kie Kling 3.0 image_urls[1], Kie Seedance last_frame_url,
Kie Veo FIRST_AND_LAST_FRAMES_2_VIDEO, and OpenAI-compatible last_frame) then receive first and last frames for
each clip; other providers fall back to first-frame-only generation. Scene continuity also improves: a scene marked
An vorherige Szene anknüpfen uses the previous scene's end frame (when present) as its visual reference instead of
the previous start frame. End-frame generation is best-effort — a failed end frame degrades that scene to
first-frame-only generation instead of failing it.
Hintergrundmusik
Projects can enable Hintergrundmusik in the storyboard settings panel (metadata.musicEnabled, PATCH/tool field
music, default off). During bundling the server generates one music track for the entire stitched video through the
enabled music-capable AI Media provider (Kie or Gemini), keyed by a hash of the effective prompt so an unchanged prompt
reuses the already-generated track across new versions. The style prompt (metadata.musicPrompt, field music_prompt)
is optional — when empty, an automatic instrumental prompt derived from the project prompt is used. ffmpeg then loops or
trims the track to the video length with a two-second fade-out and mixes it under the existing audio, ducked by
sidechain compression whenever voices or dubs speak (see Endbearbeitung below for the master loudness). Music generation is best-effort: while the track is still
generating, bundling waits and retries on the next poll; a failed or stale (15 min) generation finishes the video
without music instead of failing the project. A finished video can get music after the fact: the play menu's Musik generieren + zusammenfügen option (generate flag add_music) enables music, generates a fresh track, and bundles the cached clips with it into a new version without regenerating any clips. The storyboard settings panel shows compact title-only toggle rows; under the music toggle a Musik generieren button generates the track standalone (play + regenerate buttons replace it once the track is ready, streamed from the project music endpoint), and under the dubbing toggle a Dubbing generieren button re-dubs all clips in the background — its voice popover offers the assigned character voices (default) or one ElevenLabs voice for every scene (stored as metadata.dubVoiceId); it requires a configured ElevenLabs provider.
Endbearbeitung (finishing, clip checks, exports)
The storyboard header's wand button opens the Endbearbeitung panel. Its settings live in
metadata.finishing (PATCH/tool field finishing, a partial object merged over the current values) and apply to the
next final version:
| Setting | Default | Effect |
|---|---|---|
mode | basic | basic stitches with ffmpeg; polished renders the picture as one HyperFrames composition with the overlays below. |
transitions | on | Uses each scene's metadata.transition (the transition INTO the scene); off = hard cuts. |
titleCard, headlines, captions, nameChips, logoOutro | off | Polished overlays: project title over the first scene, each scene's metadata.headline, the script lines as captions, a name tag the first time a character appears, and a 3-second logo outro. |
outroName, outroUrl | empty | Outro text; empty uses the instance branding (NEXT_PUBLIC_BRANDING_*, logo from public/). |
beatSync, musicBpm | off, 120 | Asks the music model for a steady tempo and trims AI shots so every cut lands on a beat. |
loudnessLufs | -14 | EBU R128 loudness of the master (-14 web/social, -16 App Store previews). |
clipQa, clipReview, clipQaAutoRetry | on | Checks every generated clip, has a vision model review its contact sheet, and regenerates a failing clip once. |
exports | none | Export profiles rendered automatically after every final version. |
Transitions. The storyboard model proposes a transition and an optional short headline per scene (cut for
continuing shots; fade/dissolve for calm moments; flash for energy; wipes/slides for topic changes; fade through
black for time jumps). The scene card edits both (PATCH /api/video-studio/ai/scenes/:id with transition
{type, duration} and headline). Types: cut, fade, fadeblack, flash, wipeleft, wiperight, slideleft,
slideright, smoothleft, circleopen, zoomin, dissolve; durations are 0.2–1.5 s and clamped to a third of the
shorter neighbouring clip. The first scene always cuts in. ffmpeg chains xfade for the picture and acrossfade
for the sound on one shared timeline; muted scenes are silenced over their visible span on that timeline.
Audio master. Every final version goes through one audio pass: the project music (looped, faded out) is ducked
under the stitched audio with sidechaincompress, muted ranges are applied, and loudnorm brings the mix to the
target loudness (true peak -1.5 dBTP).
Polished mode. The same timeline is written as a HyperFrames composition: each normalized clip is a muted
<video> layer, transitions are GSAP tweens, overlays and the outro are HTML layers, and the stitched ffmpeg audio
is muxed onto the render (padded under the outro). If the render fails, the version is finished with the plain
stitch and the project banner says so.
Beat sync. The music's tempo, beat phase and downbeat are estimated from its onset envelope. Each AI shot keeps at least 60% of its length and is trimmed so the middle of its outgoing transition lands on a beat, preferring downbeats and moments of peak motion in the shot (ffmpeg scene scores). Scenes with dialogue and HTML scenes are never trimmed. An estimate with low confidence leaves the timing unchanged.
Clip checks. When a clip is ready, the scene stays clip_generating while it is checked: duration against the
request, hard cuts inside the clip (select='gt(scene,0.3)'), black stretches (blackdetect), frozen picture
(freezedetect), and a 4×2 contact sheet that the vision model (the rag_document_vision_model, then
video_storyboard_model, then agent:main) reviews for prompt mismatch, extra people, identity changes, logos or
text, and deformities. A failing clip is regenerated once; a clip that fails again is kept and flagged. The verdict
is stored in scene.metadata.clipQa (checking, passed, retrying, flagged with issues[]), shown on the
scene card, and the sheet is served by GET /api/video-studio/ai/scenes/:id/qa-sheet. If the check itself cannot
run, the clip counts as passed.
Exports. Each final version can carry deliverables in metadata.versions[].exports, rendered in the background
from the master into video-studio/exports/<slug>/v<N>/ (removed when the version is deleted):
| Profile | Output |
|---|---|
web | H.264 CRF 23, +faststart, project size, -14 LUFS. |
social_vertical | 1080×1920; a landscape master is shown whole over a blurred fill. |
square | 1080×1080 with the same blurred fill. |
app_store_preview | 886×1920 (portrait) or 1920×886 (landscape), center-cropped, 30 fps, at most 30 s with a fade-out, H.264 10 Mbit/s, AAC 256 kbit/s, -16 LUFS. |
posters | Three JPEG stills from the opening, middle and ending. |
Start exports from a version tile's Exportieren… menu, with
POST /api/video-studio/ai/projects/:id/versions/:version/exports {profiles}, or with the agent tool
video_studio_export_version. Download with GET …/exports?profile=<profile>&file=<index>&download=1.
Blocks designed natively for portrait can declare "aspect": "9:16" in their block.json; when a 9:16
project renders a block scene whose blocks all declare 9:16, the composition bleeds full-frame
(portraitFill: "bleed") instead of framing the 16:9 stage with brand bars. The storyboard's "Neue Szene"
tile opens a three-option menu: AI scene, building-block scene, or linking an existing gallery video as a
scene.
Storyboards additionally follow a one-speaker-per-scene rule: multiple characters may be visible in a scene, but all script lines of one scene come from a single speaker; conversations are split across consecutive scenes. This keeps per-scene dubbing and voice changes unambiguous.
- Characters first — create recurring people in Charaktere, then attach their ids to the project and scenes. Assigning a character directly to a scene automatically links that character to the project as well, including assignments made through the Video Studio agent tools. Frame generation receives their stored visual identity and portrait references so a series does not rely on names alone.
- Storyboard review — selecting a KI project in the gallery opens the scene board as a full-screen gallery
takeover. Prompts and dialogue stay editable per scene, and a start frame can be generated again or edited while
it remains linked to the scene. Editable storyboards can append, delete, and drag-reorder scenes on the web; the
native list provides add, swipe-delete, and move actions. Structural changes are blocked during storyboard or
clip generation/concatenation. Changing a ready storyboard returns it to
storyboard_ready, after which the existing new-version generation flow creates the next output. - Project aspect ratio — the storyboard header always shows
16:9or9:16. Web users can change it while the project isdraft,storyboard_ready, orfailed; the native storyboard header displays the stored value read-only. Existing frames remain reviewable and should be regenerated for the best composition after a change. - Per-scene clip preview — as soon as an AI scene reaches
clip_ready, its card and native detail view can play the authenticated generated clip over the start-frame poster. This makes each moving scene reviewable before the final concatenated gallery video is ready. - Scene continuity — each AI scene after scene 1 can be marked as continuing the immediately previous setting.
Its start-frame generation then uses the previous AI scene's ready start frame as the first environment reference
for location, lighting, props, and wardrobe, followed by portraits of characters in the current scene. Image edits
keep the current frame first and add the previous frame when the four-reference limit allows. HTML scenes and
missing or failed previous frames break the chain without blocking generation.
Reordering does not rewrite this flag:
continues_previousstays attached to the moved scene, so “previous” means whichever scene precedes it in the new order. - HTML scene mixing — a scene can reference an existing gallery slug or select ordered Bausteine directly from the library. The picker shows looping previews, metadata duration defaults, curated per-block text fields, and an AI content prompt for every selection. The prompt is disabled with an explanation only for an explicitly slotless block. Rendering uses the project aspect through the existing module manifest/HyperFrames handler and saves a versioned per-scene MP4 immediately in the storyboard. Both block scenes and linked gallery videos have playable previews. Saving unchanged blocks reuses the existing MP4 and preserves applied audio; changed selections, duration, content, or aspect render a new clip, and a missing MP4 is repaired. Dubbing, voice-over text edits, and regenerating another scene keep the saved HTML visuals. Final assembly reuses those clips and their normalized intermediates. Re-rendering asks for confirmation. Native clients identify both kinds and play ready clips, while creation/editing is web+agent only.
- Provider and model — a project stores its exact video provider and model. If neither is selected, creation uses the enabled video provider carrying the configured default model (or the first enabled video provider). Explicit choices are preserved; generation does not silently replace them with another default.
- Start-frame image model — creation stores the selected image provider/model in
metadata.imageModel. The project header can change that selection, and frame regeneration can override it for one call. Generation and edit requests pass the pair through the normal image resolver; a selected model that cannot edit falls back to an edit-capable configured provider without changing the stored project choice. Every generation/edit prompt also states the project orientation explicitly. Clapilot probes the returned image with ffprobe; a frame outside the five-percent canonical-ratio tolerance is scaled to cover and center-cropped with ffmpeg (never padded) to1536x864for16:9projects or864x1536for9:16projects (image models are asked for their nearest size,1536x1024/1024x1536, which video models would otherwise stretch). Reference images go the other way: before an edit request the first reference (portrait or character image) is extended to the project's aspect ratio with a blurred, darkened copy of itself behind it, so the model receives the whole person instead of a crop. Frame prompts forbid text overlays and brand logos or trademarks. The corrected PNG is stored as a newgenerated_imagesasset whose metadata recordsaspectNormalizedFromwith the original provider asset id, and only that corrected asset is linked to the scene. - Failed start frames and resuming — a start frame that the image provider rejects marks only that scene
failedwith a localized "start frame generation failed" message that wraps the provider's concrete error (for example an OpenAI-Codex HTTP 400 body); HTML scenes and already rendered clips are untouched. Project creation,video_studio_create_ai_project,video_studio_start_generation, andvideo_studio_get_ai_projectreport such scenes explicitly (failedScenes[]plus a message naming the scenes) instead of claiming success. Starting generation on that project first claims it fromstoryboard_ready/failedintostoryboard_ready(a project that is stillgeneratingis rejected before any scene is changed), then regenerates the missing start frames in the background with the stored image selection and chains the clip fan-out behind them, so an existing project is resumed after a provider fix rather than duplicated and the request never blocks on image calls. WithretryFailedOnly, onlyfailedscenes get a new frame; a separately addedpendingscene is left alone. While frames regenerate the project staysstoryboard_ready; the fan-out claimsgeneratingafterwards and a scene whose frame still fails ends the project asfailedonce no clip is in flight. OpenAI-Codex image generation picks its Responses host chat model from the provider's configured Codex models and the shipped defaults, skipping models the ChatGPT-account Codex backend rejects as unsupported (detected from the rejection wording about the model, whether or not the body echoes the model name). - Final-video versions — every concat writes
<output>-v<N>.mp4and<output>-v<N>.jpg. Ordered entries inmetadata.versionsrecord the duration and exact video provider/model, while the legacy final-path columns keep pointing to the latest file. The web storyboard can play and delete non-latest versions; native clients list and play every version but intentionally leave deletion to the web app. An existing unversioned final is recorded as v1 when the next concat writes v2. - Storyboard model —
app_settings.video_storyboard_modeloptionally selects the text model that produces the strict storyboard JSON. An empty value falls back toagent:main; the model actually used is stored on the project.
Project state machine:
draft → storyboard_generating → storyboard_ready → generating → concatenating → ready
Adding, deleting, or reordering a scene on a ready project transitions ready → storyboard_ready; the user then
reviews the changed structure and uses the normal new-version generation flow.
failed is a terminal/error branch from storyboard generation, scene generation, provider-unreachable stall
detection, or concatenation. A reviewed, failed, actively generating, or ready project can be restarted. Restart
preserves scene start frames, detaches every prior AI clip, and submits every AI scene as a fresh provider request;
for a ready project this is the Generate new version flow and earlier final files remain untouched. Cancelling active clip generation returns the project to storyboard_ready and resets
in-flight scenes.
Scene state machine:
pending → frame_ready → clip_generating → clip_ready
failed records a missing frame, provider failure, or block-render failure. Linked HTML scenes become clip_ready
after validation; block scenes use clip_generating → clip_ready|failed. A persisted block render older than 15
minutes without a result reconciles to a localized failed state.
Troubleshooting stalled generation
The project status endpoint polls provider jobs and then reconciles the persisted asset state even when one provider
status request throws. Network errors are stored per generated-video asset with an error count, first-error timestamp,
last check, and bounded retry time. The normal configured poll failure window still moves a persistently unreachable
asset to failed. Independently, a clip_generating scene is marked failed with a localized provider-unreachable
message once its asset has at least three poll errors spanning five minutes, so the UI never remains silently stuck.
Use Neu starten / Restart (or video_studio_start_generation { restart: true }) after the provider recovers; this
creates fresh provider jobs and may spend provider budget again.
A provider that answers every poll with "processing" but never finishes, a submitting process that died before a
scene received its clip job, or a container restart during ffmpeg concatenation would otherwise leave a project in
generating/concatenating forever without any error. The generation watchdog (part of every reconcile tick and
status poll; window CLAPILOT_VIDEO_STUDIO_GENERATION_STALE_MINUTES, default 45) closes those gaps with a
per-case diagnosis instead of a silent stall:
- Stale clip job (
clipStale): aclip_generatingscene whose linked generated-video job is still unfinished and was submitted more than the window ago is markedfailed. The message names the provider, model, and last provider job status. Job age is measured from the job's creation time, not from its last poll, so a job that keeps reporting "pending" still trips the watchdog. - Missing clip job (
clipJobMissing): aclip_generatingscene whose generated-video row no longer exists (for example deleted from the gallery) fails immediately. - Stalled scene (
sceneStalled): an AI scene still inpending/frame_ready(orclip_generatingwithout a job id) while both the project row and the scene row are older than the window is markedfailed, because no worker will ever submit it. The project claim bumpsupdated_at, so a fresh fan-out with old untouched scene rows is never treated as a stall. - Interrupted concatenation (
concatStale): aconcatenatingclaim older than the window transitions tofailed; the finished clips stayclip_readyand can be bundled again. - Project diagnosis (
generationStale): when the watchdog failed at least one scene and nothing is in flight any more, the project becomesfailedwith a message naming the window and the number of affected scenes; the concrete cause is on each scene. Ageneratingproject without any scene fails the same way once the claim is older than the window.
Each watchdog action logs a structured video_studio_generation_watchdog event with reason
(clip_job_stale, clip_job_missing, scene_stalled, concat_stale, no_scenes), the project/scene ids, and the
provider job details where available. Afterwards the normal recovery actions apply: Fehlgeschlagene erneut
versuchen / Retry failed resubmits only the failed scenes, per-scene regeneration rebuilds one clip, and
Zusammenfügen / Bundle re-runs an interrupted concatenation. A retry-failed-only run also submits scenes that
already have a start frame but never received a clip (frame_ready), so a partial retry cannot itself leave a scene
behind.
The storyboard header model menu can override the stored video provider/model at generation start or restart. The
selection is persisted before scene requests are submitted, and scene durations are normalized again for the selected
model. A ready or failed scene can also be regenerated individually after budget confirmation; the project returns to
generating, then the normal single-winner reconciler appends a new numbered final and updates the gallery pointer.
Install content from the bundle store
The studio ships general: only the generic skill engine (scripts/ + a generic SKILL.md) is seeded into
new instances. Branded scenes, templates, base shells and design styleguides live in a bundle catalog
that ships with the module but is not auto-seeded:
- catalog:
bundled-modules/video-studio/bundles/<id>/— each withbundle.json(id,name,description,version,includesBase,blocks[],templates[],styles[],extras[],legacyPaths[]) +blocks/(incl._base) +templates/+ theextras(styles/,agents/). - shipped bundles:
clapilot-video-bausteine— the 6 Clapilot Bausteine (brand-intro,brand-outro,email-workspace,ai-draft,decision,chat-scene), 2 templates, the shared base shell, and theclapilotdesign style.nordlicht-dark— a dark studio set (aurora-intro,aurora-metrics,aurora-outro) with its own base shell and thenordlichtdesign style.
A Store button (bottom-right on the Bibliothek/Stile tabs) opens a popup listing bundles with
install/uninstall (each row shows the shipped Bausteine/Vorlagen/Stile counts). Install copies the bundle's
blocks/_base/templates and each extras item into the seeded skill dir; uninstall removes them plus
any legacyPaths left by older bundle layouts. A bundle reads as installed when its blocks are present in the
workspace, so existing live instances (already seeded before the store existed) show it installed
automatically — nothing is removed from their volumes. New instances start empty and install on demand.
Extensible for more bundles and, later, user uploads.
Migration note: the entrypoint seed (
seed_workspace_tree_if_missing) syncs files present inworkspace-seedbut never deletes volume-only files. Moving content out ofworkspace-seedleaves existing volumes' content intact; the seededSKILL.mdis auto-synced to the generic version, and the styles layout lands on an existing instance when the bundle is (re)installed from the store (uninstall also cleans the legacy top-levelSTYLEGUIDE.md/references//assets/).
Bibliothek (templates, Bausteine, assets)
The Bibliothek tab surfaces the building-blocks library the agent composes from:
- Vorlagen — full templates in
skills/clapilot-video-styleguide/templates/(render-as-is or customize). - Bausteine — reusable scene blocks in
skills/clapilot-video-styleguide/blocks/<slug>/(block.json+block.html+block.css+block.js, scoped under.b-<slug>). Each card shows an inline looping preview (a small mutedpreview.mp4rendered per block). The agent compounds a video from them withscripts/compose-video.mjs --blocks a,b,c --name <slug>, and can create new blocks (which then appear here). Template blocks declare editable copy inblock.json.textSlots[]as{ key, label, find, defaultValue, scope, required?, multiline? }, wherescopeishtml,js, orboth. The block picker shows the label anddefaultValue, and persists edits through the per-scene{ find, replace, scope? }compose contract. The composer replaces the first literal occurrence in the selected scope; omitted override scopes still target both sources for backward compatibility. Intentionally slotless blocks declare an explicit empty array. Third-party blocks with notextSlotskey derive up to twelve deduplicated slots from visible HTML text nodes. See the skill'sSKILL.md→ Bausteine. Regenerate previews after adding/changing a block withscripts/render-block-previews.mjs(runs in the runtime container — composes each block solo, renders, then downscales to a 480p/15fps muted clip intoblocks/<slug>/preview.mp4). - Assets — shared images (the Clapilot logo) from
blocks/_base/assets/plus custom uploads: drag images (or use the upload tile) into the assets grid to add your own logos/graphics. Custom assets live inskills/clapilot-video-styleguide/custom-assets/(outside the bundle dirs, so store install/uninstall never touches them), are badged Eigene with a delete action, andcompose-video.mjscopies them into every composed project'sassets/— scenes reference them asassets/<file>(the agent is told about them in the module prompt). The grid additionally lists the workspace-global styleguide brand assets (Settings → Styleguide, stored under<workspace>/.clapilot/styleguide-assets/) badged Brand; they are read-only here (managed via/api/styleguide/assets) andcompose-video.mjscopies them into composed projects alongside the custom uploads.
Design styles (Stile tab)
Design styles are first-class entities at skills/clapilot-video-styleguide/styles/<slug>/:
style.json—{ slug, title, description, source, version }(source= shipping bundle id orcustom).STYLEGUIDE.md— the style's design rules; the agent reads and follows it when the style is selected.- optional
tokens.css(inlined deterministically viacompose-video.mjs --style <slug>),references/,assets/.
A style may also ship a human-facing HTML reference page under references/*.html (the Clapilot style ships
clapilot-video-styleguide.html); the Stile card then shows an HTML-Referenz button that opens it in a
sandboxed iframe popup, streamed via style-reference.
Surfaced in the UI as the Stile tab (cards with source badge and a rendered styleguide popup). The
"Neues Video" composer has a Stil selector (defaults to the first installed style) published as
pageContext.style; the chat route instructs the agent to read + follow that style's STYLEGUIDE.md. Bundles
ship styles under their styles/ dir; new styles are created by the agent in chat (writing
styles/<slug>/style.json + STYLEGUIDE.md with source: "custom"), and new Bausteine are tagged with their
style via block.json.style (shown on the Bibliothek cards).
Fine-tune a video (Feinschliff editor)
Gallery videos with a matching project show a pencil action that opens the Feinschliff editor — a native, in-module rebuild of the essential HyperFrames-Studio workflow (no preview server, no proxy, docker-friendly):
- Preview + transport — the project's
index.htmlloads in a same-origin iframe (served viaproject/…); because HyperFrames compositions expose a paused, seek-safe GSAP timeline atwindow.__timelines, the editor drivesplay/pause/seekdirectly for frame-accurate scrubbing, with a time ruler and clickable scene lanes. - Scene edits — per-scene duration inputs and reorder controls, plus per-scene text fields (leaf text nodes).
- Post-production voiceover — finished block-based projects expose a scene-level voiceover editor in
Feinschliff. The initial action synthesizes one TTS file per scene through the configured Clapilot TTS runtime
(Gemini/Kore by default). Text and voice remain editable; regenerating one scene only calls TTS for that segment
and remuxes audio onto the existing MP4 without rendering the visual composition again. Audio is aligned to the
scene start, padded when short, and moderately accelerated (up to 1.35x) then clipped when it exceeds the scene.
Voiceover metadata and relative segment paths are stored under
voiceoverincompose.json. Choosing an uploaded character voice sends that character asvoice_character_id, locks the provider to the OpenAI-compatible Spark TTS path, and uses the same sample as theref_audiovoice-clone reference for every synthesized scene segment. - Persistence model —
compose.json(emitted bycompose-video.mjsfor every composed project) is the single editable source:{ blocks, durations, aspect, portraitFill, style, textOverrides }. Saving writes the manifest and recomposesindex.htmlfrom it (--manifest); text overrides are first-occurrence string replacements per scene, re-applied on every recompose. Duration semantics: shorter scenes clip at their end, longer ones hold the final state (tween offsets stay local to the scene). - Re-render —
render { fromProject: true }renders the recomposed project over the existing MP4. - Legacy migration — composed projects from before the manifest existed are auto-migrated on first open:
project-manifestparses blocks/durations/aspect/style back out ofindex.htmland diffs each scene against its pristine block to preserve agent-made text edits astextOverrides, then persists thecompose.json. - Clip emulation — raw compositions stack all scenes absolutely (the HyperFrames runner normally toggles them per frame), so the editor applies clip windows itself on load/seek/playback — exactly one scene visible.
- Hand-authored projects (no
compose.json, not composer-shaped) open in preview-only mode (scrub + re-render, no scene edits).
Manage the gallery
Search the gallery by title, file name, or KI prompt; all search terms must match, ignoring case and diacritics, and clearing the search restores the filtered gallery.
The Alle / HTML / KI filter separates non-KI rendered videos from matched KI outputs and KI project cards.
Rendered outputs match projects by the gallery slug and the project's outputSlug/finalVideoPath. Each video
card exposes a trash action with in-place confirmation. video-delete removes only the rendered MP4 — the
composition project stays on disk, so the video can be recomposed and re-rendered later.
KI project deletion is separate: DELETE /api/video-studio/ai/projects/:id removes the project (any authenticated
user may delete any project) and cascading storyboard rows. It never removes the finished gallery MP4. The native iOS/macOS gallery and
storyboard use the same endpoint through destructive confirmation dialogs.
How the agent can drive it
When chatContext.moduleSlug === "video-studio", src/app/api/chat/route.ts injects a module system
prompt instructing the agent to either compound a video from Bausteine (scripts/compose-video.mjs) or
scaffold a template (scripts/new-clapilot-video.mjs), edit the German scene copy, lint, and render in the
background with npx hyperframes render --output /app/workspace/video-studio/videos/<slug>.mp4 via
exec_command (detached so it beats the exec time limit). The MP4 surfaces in the gallery on the next poll.
UI selections (pinned Bausteine, voiceover/music, style, aspect/fill/resolution) reach the agent through the
chat pageContext keys described above.
Vertonung pipeline — when the user requests a voiceover and/or music, the agent generates the audio with
the first-party media tools (media_tts_speak → absolutePath; livestream_generate_music → audioPath,
polling livestream_get_asset for the async kie-ai-music provider), then renders silently to
projects/<slug>/silent.mp4 and muxes the tracks into the final videos/<slug>.mp4 with
scripts/add-audio.mjs (voiceover at full volume, music looped + ducked, output trimmed to the video length).
See the skill SKILL.md → Audio — Voiceover & Musik.
Configuration & limits
Runtime data and storage
All compositions and rendered videos live on the shared workspace volume so both the web container
(module API) and the agent container (exec_command renders) see them:
- rendered videos:
/app/workspace/video-studio/videos/<slug>.mp4 - AI project exports:
/app/workspace/video-studio/exports/<slug>/v<N>/ - AI project build files (normalized clips, clip-check contact sheets, polished compositions):
/app/workspace/video-studio/projects/<slug>/ai-build/ - composition projects:
/app/workspace/video-studio/projects/<slug>/ - character voice samples:
/app/workspace/.clapilot/video-studio-voices/<owner-user-id>/<character-id>.<ext> - skill engine (always seeded):
/app/workspace/skills/clapilot-video-styleguide/— a genericSKILL.md(compose/render/audio mechanics, points to the installed bundle'sSTYLEGUIDE.mdfor the house style) +scripts/. No house style of its own. The content AND the design styleguide (base shell + Bausteine + templates +STYLEGUIDE.md+references/+ brandassets/+agents/) are NOT seeded; they install from a bundle (see the bundle store above). - installed-bundle marker:
/app/workspace/video-studio/installed-bundles.json.
The gallery can be filtered by video source and sorted by newest or oldest modification date, or alphabetically by title in either direction. The same controls and ordering are available in the web, iOS, and macOS clients.
Implementation: bundled-modules/video-studio/api/handler.mjs.
Runtime requirements
- HyperFrames CLI is installed globally in the runtime image (
Dockerfileglobalnpm install -g). - The image already ships Node 22, ffmpeg, and Chromium; HyperFrames renders with system Chromium
(
--disable-dev-shm-usage, so the default 64 MB/dev/shmis sufficient).
Module API endpoints
Base: /api/modules/video-studio/api
| Method | Endpoint | Purpose |
|---|---|---|
| GET | list | List rendered MP4s in videos/ with title, size, mtime, url. |
| GET | templates | List full video templates from the seeded skill (clapilot-feature-flow, clapilot-chat-interface). |
| GET | blocks | List composable scene blocks (Bausteine) from the skill blocks/; each block includes previewUrl when a preview.mp4 exists and normalized textSlots[] (key, label, exact compose find, displayed defaultValue, scope, required, multiline). First-party slots come from the bundled catalog so existing installs receive curation updates; a third-party manifest with no slots key falls back to visible HTML text extraction, while an explicit empty array stays slotless. |
| GET | assets | List video assets: bundle assets (blocks/_base/assets/) plus user-uploaded custom assets, each with a source flag. |
| GET | asset?path=<name> | Stream an asset image (custom assets shadow bundle names). |
| POST | asset-upload | Multipart upload (files) of custom images (PNG/JPG/WEBP/SVG, ≤20 MB each) into custom-assets/ — sanitized + deduped names; the Bibliothek exposes it via drag-and-drop + a file picker. |
| POST | asset-delete | Body { name } — delete a custom asset (bundle assets are managed by the store). |
| GET | block-preview?slug=<slug> | Stream a block's looping inline preview MP4 (blocks/<slug>/preview.mp4, HTTP Range, short-cached). |
| GET | file?path=videos/<name> | Stream an MP4 (supports HTTP Range) with the session cookie. |
| GET | styles | List installed design styles ({ slug, title, description, source, version, hasGuide }; legacy top-level STYLEGUIDE.md is synthesized as a legacy style). |
| GET | style?slug=<slug> | Style detail incl. the full STYLEGUIDE.md text (guide). |
| GET | style-reference?slug=&file= | Stream a style's bundled HTML reference page (e.g. the Clapilot video styleguide HTML) for the UI iframe popup. |
| GET | bundles | List the bundle catalog with install status ({ id, name, description, version, blockCount, templateCount, styleCount, installed }). |
| POST | bundles/install | Body { id } — copy a bundle's blocks/_base/templates into the skill dir. |
| POST | bundles/uninstall | Body { id } — remove a bundle's blocks/_base/templates from the skill dir. |
| GET | project/<slug>/<...file> | Serve a project file (index.html + relative assets) for the Feinschliff editor's same-origin iframe preview. |
| GET | project-manifest?slug= | The project's editable compose.json ({ hasManifest, manifest }; hand-authored projects report hasManifest: false). |
| POST | project-edit | Body { slug, blocks?, durations?, textOverrides?, title? } — merge into compose.json and recompose index.html via compose-video.mjs --manifest. |
| POST | render | Render { name, html }, { name, template }, { name, blocks, durations?, textOverrides?, aspect?, style? }, or { name, fromProject: true }. Block input writes the normal manifest and invokes compose-video.mjs --manifest; per-block picker/agent overrides are flattened to the manifest's scene-indexed { find, replace } entries before this call. Every form then runs HyperFrames into videos/. |
| POST | video-delete | Body { name } — delete a rendered MP4 from the gallery (the project stays for re-rendering). |
Apple client (iOS/macOS)
- Galerie / Vorlagen / Bibliothek — native grids/lists over the module API. MP4s and the looping Baustein
previews are fetched authenticated through
ClapilotAPI.fetchVideoStudioMedia(session cookie), cached as temp files (VideoStudioMediaCache), and played withAVPlayer/AVPlayerLooper. - AI project status — the Galerie shows a compact native strip for active projects and recently failed projects.
It polls the workspace-global project list only while work is in progress. Native clients create AI projects
(
VideoStudioAICreationView.swift), manage characters with portraits and voice samples (VideoStudioCharactersView.swift), and edit the storyboard (VideoStudioStoryboardView.swift): add, delete and reorder scenes, edit prompts and scripts, transitions and headlines, regenerate frames and clips, start or cancel generation, set the finishing options, and create and download exports. - Mixed-scene parity — native storyboard rows show linked-gallery versus block-rendered kind badges, continue polling in-flight block renders, and play authenticated ready clips. Authoring remains web-and-agent only.
- Neues Video — a fourth segmented tab that embeds the real chat:
ChatView(presentation: .embedded)— a chrome-free presentation of the standard chat (full markdown/canvas rendering, tool-call log, attachments, the real composer) — pinned to the dedicated session (AppModel.activateVideoStudioChatSession,POST /api/chat/sessions { ensureScope: "video-studio" }; the previously active session is restored when leaving the tab). The video controls (tap-to-pin Baustein chips, voiceover/music menus, 16:9/9:16 + Füllung + 1080p/4K) render as acomposerAccessory(VideoStudioComposerToolbar.swift) above the composer and publish via the live context. Every Apple text-chat send attaches the client context (AppModel.currentClientContextPayload, snapshotted onto the queued message at submit time to avoid racing UI state) — web parity for module-aware prompts.
Troubleshooting
compose-video.mjs --listshows no Bausteine / the Bibliothek is empty — new instances start without content; install a bundle from the Store popup first.- Portrait (9:16) output does not match the selected fill — aspect ratio and resolution are deterministic,
but the fill style is agent-guided: in practice the agent often prefers to adapt the scene layout to fill
the portrait frame (stacking the 16:9 columns) even when
frameis selected — this usually looks more native, though quality can vary. The strict brandedframewrapper is always available viacompose-video.mjs --portrait-fill frame; nudge the agent with a follow-up if you want it. - The video does not appear immediately — renders run detached in the agent container and typically take ~30–60 s; the gallery picks the MP4 up on its next poll.
- The module is missing on iOS/macOS — the native section is hidden when the module API probe returns 403 (non-admin users).
File context menu
Right-click a rendered video or storyboard for Rename, Copy, Share, Download and Delete. Sharing offers Email and Chat (Channel or Direct message); it opens a draft and does not send. Videos download in their original format, storyboards as JSON. Video rename preserves file paths used by project versions. Copying a storyboard creates a new editable storyboard without starting generation.
The Apple clients use native context menus (right-click on Mac, long-press on iPhone/iPad). Sharing opens the authenticated Clapilot review/composer flow. Channel and direct-message sharing links to the original authenticated module/item and retains its permissions; no export or public link is created. Email continues to use a revocable public document link. No message is sent automatically.
Internal links in chat automatically show authenticated inline media or a clickable reference to the original item. Videos have playback controls, images have previews, and other documents/module items use document-style references. This also applies to existing messages and the Apple clients; no exported copy or public link is generated.
Shared generated media
The Media library action browses workspace-global generated output, with folder selection, nested folder creation, search, previews, and moving assets between folders. Existing files are referenced in place. Image Playground results can be reused in Social Media; Video Studio exports and completed scene clips can be imported into Social Media or Livestream. Livestream imports copy into the existing streamer media mount and do not enqueue or start playback. Social Media imports attach to the current draft and do not publish it.
The same library is available in web, iOS, and macOS. See Image Playground for visibility rules and agent access. Deployment requires 311_media_library.sql. Source removal can make a library reference unavailable; moving a library item never moves the source file.
Explicit scene silence and STS reset
Use video_studio_mute_scene (or POST /api/video-studio/ai/scenes/:id/audio
with mode=muted) when a scene must be fully silent. This removes embedded
source speech as well as applied TTS/STS overlays, without changing the
source file or any previous final version. It clears active narration state
and returns the silent preview on the next project read. Merely clearing the
script or using change_scene_voice(reset=true) cannot remove embedded audio;
the latter only undoes the STS conversion.
The project must be idle. Muting preserves actual source video timing: a 3.008s
HTML source requested as 5s remains 3.008s, without a frozen tail or spilled
narration. bundle_only=true stitches current clips into a new version,
without provider, TTS, STS or music generation, and honors mute even under an
existing project music bed. V1/V2 remain available; production output and audio
QA must be explicitly performed after deployment, never inferred from tool
success. Originals/derived files are retained; this operation does not promise
an automatic unmute or speech/music separation.
The operation is available to chat/Live/CLI agents. Web and Apple clients use the same updated scene clip URL for preview; no new client-only editor control is introduced. Bearer access uses the same workspace member/role boundary as browser access, including specialist and automation sessions with a verified actor. No project-owner impersonation or public media exception is needed.
Scene mute lifecycle
Mute applies to the current scene clip. Regenerating/replacing an AI clip or HTML source clears the override; editing script or voiceover text also clears it so narration can be generated again. There is no standalone unmute mode; STS reset only discards voice conversion. Existing final versions remain unchanged. A missing silent derivative returns 404. Bundle-only suppression applies only during an active generation/concatenation run and does not suppress later narration edits. Web, iOS, macOS and agent operations share these server rules; no client controls change.
Automatische Videotitel
Die Galerie zeigt zuerst einen manuell vergebenen Titel. Andernfalls verwendet sie den Titel der Komposition oder des HTML-Dokuments; bei älteren Kompositionen mit technischem Standardtitel wird die sichtbare Überschrift verwendet. KI-Videos übernehmen den inhaltlichen Storyboard-Titel beim Export. Ein bereits vergebener Projekttitel bleibt bei der Storyboard-Erstellung erhalten. Dateinamen und Projektverknüpfungen bleiben unverändert. Diese Auflösung gilt gemeinsam für Web, iOS, macOS und den Modul-API-Zugriff durch Agenten. Ohne verfügbare Inhaltsmetadaten bleibt der bisherige Dateiname der Fallback.
