Image-to-Video with FLUX 3
Animate a still image into a 5-20 second video with synchronized audio using FLUX 3 by Black Forest Labs. The clip starts from your uploaded image, and you can optionally pin a last frame so the video transitions exactly where you want it to end — at 720p or 1080p.
What makes FLUX 3 different
Your image becomes the first frame
Upload a still — a photo, a render, or an image generated earlier on the same board — and FLUX 3 animates it into a clip of 5 to 20 seconds. The prompt describes what happens next: the motion, the camera move, the sound. Because the prompt is interpreted and expanded before generation, plain language is enough to direct the shot.
Pin the last frame for exact transitions
The optional Last Frame input lets you give FLUX 3 both endpoints: the video starts on your first image and lands precisely on your second. That turns image-to-video into a transition engine — morph between two product angles, two scenes, or two poses — with the in-between motion generated for you.
Synchronized audio included in the rate
Like the text-to-video mode, FLUX 3 image-to-video generates matching ambient sound, speech, and effects at no extra charge — $0.17 per second at 720p and $0.29 at 1080p whether audio is on or off. Note one model rule: runs with image inputs are capped at safety tolerance 2.
Chain it after an image model
On a Scenetra board, FLUX 3 image-to-video slots naturally after an image node: generate a still with FLUX.2 Pro or Nano Banana 2, wire it into the video node, and animate the keeper. For cheap motion tests there's FLUX 3 Draft image-to-video at $0.06 per second — same inputs, preview pricing.
Playground
A timelapse of a flower blooming in a sunlit meadow, cinematic quality
Drop images or click to upload
Try FLUX 3 in Scenetra
Open PlaygroundParameters
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt* | text | Describe the scene, action, camera movement, and any audio you want. Plain language works — the prompt is interpreted and expanded before generation. | — |
| Image* | image upload | First-frame image. The video starts from this image. | — |
| Last Frame | image upload | Optional last-frame image. The video transitions from the first frame to this one. | — |
| Duration | select | Duration of the video in seconds (5-20).5 · 6 · 7 · 8 · 9 · 10 · 11 · 12 · 13 · 14 · 15 · 16 · 17 · 18 · 19 · 20 | 5 |
| Resolution | select | Output resolution.720p · 1080p | 720p |
| Aspect Ratio | select | Aspect ratio of the generated video. Auto lets the model choose from your prompt and inputs.Auto · 21:9 · 2:1 · 16:9 · 4:3 · 1:1 · 3:4 · 9:16 | auto |
| Generate Audio | select | Generate synchronized audio (ambient sound, speech, effects). Off for a silent clip.false · true | true |
| Safety Tolerance | select | Moderation strictness, 0 (strictest) to 4 (most permissive). Runs with image inputs are capped at 2 by the model.0 · 1 · 2 · 3 · 4 | 4 |
Pricing
From $0.17 per second
Pay per generation — only pay for what you use.
720p
$0.17/s
per second
1080p
$0.29/s
per second
Audio
included
per second
What a video costs
| Duration | 720p | 1080p |
|---|---|---|
| 5 seconds | $0.85 | $1.45 |
| 10 seconds | $1.70 | $2.90 |
Use cases
Product & e-commerce
- Animate product photography
- Two-angle transition shots
- Listing videos with ambience
- 1080p hero footage from stills
Film & previz
- Animate concept art
- First-to-last-frame shot design
- Establishing shots from stills
- 20-second scene extensions
Social content
- Bring photos to life with sound
- Vertical 9:16 story clips
- Before/after transition posts
- Animated key visuals
Design & branding
- Animate brand imagery
- Logo scene reveals
- Moodboard stills into motion
- Campaign variants per ratio
Related models
FLUX 3 Draft
Animate a still image into a 720p draft video at $0.06 per second with FLUX 3 Draft, the rapid-iteration tier of Black Forest Labs' FLUX 3. Same inputs as the standard model — first frame, optional pinned last frame, 5-20 second durations, synchronized audio — at roughly a third of the price, so you can test how an image moves before paying for the full-quality render.
View model →Flux 2 Pro
Black Forest Labs' professional-grade image model delivering high-quality outputs with precise control. Flux 2 Pro supports text-to-image generation and image editing with flexible sizing, safety controls, and simple flat per-image pricing.
View model →Seedance 2.5 Image-to-Video
Animate any image into a cinematic video up to 30 seconds with Seedance 2.5 by ByteDance. Start from a first-frame image, optionally set a last frame to control the ending, and get synchronized audio by default. Output preserves your image's aspect ratio.
View model →Kling 3.0 Pro
Kuaishou's flagship video generation model delivering stunning visual fidelity. Kling 3.0 Pro supports text-to-video, image-to-video with start/end frames, element references for character consistency, and native audio generation.
View model →Frequently asked questions
How does FLUX 3 image-to-video work?+
You upload an image that becomes the video's first frame, write a prompt describing the motion, camera, and sound you want, and FLUX 3 generates a 5-20 second clip from there at 720p or 1080p, with synchronized audio. On Scenetra it runs as a node — upload directly or wire in an image from another node on the board.
What is the last-frame input for?+
It's an optional second image that pins where the video ends: the clip starts on your first image and transitions to the last one, with FLUX 3 generating the motion in between. It's the tool for controlled transitions — between poses, camera angles, scenes, or product states.
How much does FLUX 3 image-to-video cost?+
The same as text-to-video: $0.17 per second at 720p and $0.29 per second at 1080p, billed on output duration, with audio included. A 5-second 720p animation is $0.85; 10 seconds at 1080p is $2.90. Scenetra bills at provider cost with 0% markup.
Does the animated video have sound?+
Yes, if you leave Generate Audio on (the default) — FLUX 3 produces ambient sound, speech, and effects synchronized to the motion, at no extra cost. Switch it off for a silent clip.
Why is safety tolerance limited with image inputs?+
FLUX 3 exposes a safety tolerance from 0 (strictest) to 4 (most permissive), but the model itself caps runs that include image inputs at level 2. That's a Black Forest Labs rule for image-conditioned generation, and it applies regardless of where you run the model.
Is there a cheaper way to test animations before the final render?+
Yes — FLUX 3 Draft image-to-video takes the same image, last-frame, and prompt inputs but renders 720p previews at $0.06 per second, about a third of the standard rate. Iterate there, then re-run the final on standard FLUX 3. New Scenetra accounts also get free welcome credits and a 7-day trial.
Can I generate the input image and the video in one workflow?+
That's exactly what Scenetra's node-based editor is for: put an image model (FLUX.2 Pro, Nano Banana 2, Seedream) and a FLUX 3 image-to-video node on the same board, connect them, and each new still can be animated with one click — no downloading and re-uploading between tools.
Start creating with FLUX 3
Use FLUX 3 alongside 50+ other AI models in Scenetra's visual workflow editor. No setup required.
Get Started Free