Image to Video: create clips with motion, speech and sound 📹

Image to Video turns a still image into a short clip. Animate an outfit, add motion to a product photo, or create a UGC-style talking video with generated speech and lip-sync. You can also describe the sounds you want in the scene.


Videos are generated at 720p. Choose 5 or 10 seconds. A 5-second video costs 60 credits; check the displayed cost for your selected duration before generating. Creating a starting image in another Claid tool costs separately.


Choose your starting image

Use a clear image that already contains the person, outfit or product you want in the video.

  • Already have a photo of someone wearing the outfit? Use it directly for a walk, small turn, fit check or talking clip.
  • Only have a clothing photo? First use AI Fashion Models to create a photo of a person wearing it. Then use that result in Image to Video. These are two separate generation steps.
  • Making a talking product clip? Start with a photo that already shows the person with the product. Keep their face and mouth visible.
  • Animating a product without a presenter? Use a product photo or a finished scene from AI Photoshoot.

Leave room in the image for the movement you want. For an outfit clip, make sure the parts of the garment you want to show are visible.


Create your video

  1. Open Image to Video in Claid Studio and sign in.
  2. Upload or select the image you want to animate.
  3. Write a prompt describing what should happen. Include the main movement, camera direction and any speech or sound.
  4. If you want help writing the prompt, use Prompt Assistant. Give it a short direction, then review the detailed prompt it produces before generating.
  5. Choose a 5-second or 10-second duration and check the credit cost.
  6. Generate the clip and find it in the Videos tab. Review the movement, product details and audio before using it.



Start with one simple action. A slow turn or small gesture is a better first test than a fast spin, several camera moves and a long spoken line together.


Prompts to try


Fashion movement

Starting image: a person wearing the outfit in front of a mirror.

Model adjusts the fit while looking in the mirror.


Generated video:


You can use this short direction with Prompt Assistant. For more control, specify the pace and camera movement in your own prompt.


A talking product clip

Starting image: a person holding the product, with their mouth clearly visible.

Hold the product toward the camera and say in a warm, casual voice: "This is the detail I wanted to show you."


Generated video:


Spoken line: "This is the detail I wanted to show you."
Keep the line short enough to fit naturally into the selected duration. You don't need to record a voice: Claid generates speech and lip-sync with the video. Check the words, pronunciation and mouth movement in the result.


Product motion with sound

Starting image: a hand holding an unopened drink can.

Open the can slowly, with a clear click and a short hiss. Keep the camera still. No speech.

Describe the visible action and its sound together. Avoid asking for many actions in one short clip.


Speech, language and quiet clips

For talking videos, put the words you want spoken in quotation marks and describe the tone. If you want a particular language, specify it and provide the line in that language. Review the output rather than assuming the speech will match perfectly.
If you don't request sound, Prompt Assistant directs the clip toward near-silence. This isn't a guaranteed mute switch. For a quiet clip, explicitly request no speech, music or sound effects, then check the result.
UGC-style means AI-generated content with a creator-like presentation. It isn't a real customer's testimonial. Write lines about visible details or demonstrations, not invented personal experiences or product results.


Common questions and fixes

Why does a 5-second video now cost 60 credits instead of 35?

The upgraded model requires more computing resources. It brings improved motion and supports generated speech, lip-sync and sound. The current price is 60 credits for a 5-second video; creating the starting image is charged separately.


The outfit, product or movement looks wrong

Try a smaller movement and fewer simultaneous actions. Make sure the starting image shows the important details clearly. AI can still change garment details, hands, text or labels, so inspect the full clip before publishing it. Another generation may produce a different result and uses additional credits.


Speech is rushed, incomplete or inaccurate

Shorten the line and simplify the movement. Give the speaker one brief sentence rather than a script with several points. Review the new take for both speech and lip-sync.


I can't hear anything

First check whether playback is muted and turn up the volume. Then check whether your prompt requested speech or sound. If you used Prompt Assistant without asking for audio, it may have directed the clip toward near-silence.


I can't generate, or a video appears stuck

Check that an image is selected, a prompt is entered and you have enough credits for the chosen duration. Follow any validation message shown in Studio. If a submitted video appears pending, check the Videos tab before submitting it again. If the problem continues, contact support through the help center with the error message and when you submitted the generation.


Can I use Image to Video through the API?

Yes. Image-to-video generation is also available through the Claid API. API requests use the prompt you provide; the Studio Prompt Assistant is a separate app feature.

Updated on: 15/09/2026