IMAGEGEN STUDIO · PRACTICAL GUIDE
Text-to-Video vs. Image-to-Video: Choose Your First Frame
The key distinction is the starting input. Text-to-video begins from a written prompt; image-to-video begins with an opening image and asks for motion. Neither is universally better. Choose based on how important the first composition is, then confirm the selected endpoint’s settings.
Text-to-video: describe a new opening scene
Use a written subject, setting and action when the first frame can be invented. Wan 3.0 Prime and Seedance 2.0 each have text-to-video endpoints. Their duration, resolution, aspect ratio and audio controls are endpoint-specific; read the selected route’s current values instead of assuming they match.
Image-to-video: anchor the first frame
Use a source still when a product, crop, artwork or approved composition should form the opening. Wan 3.0 Prime has a dedicated I2V endpoint requiring an opening image. Vidu Q3 Turbo is also documented as I2V, not T2V, and its notes describe 1–16-second clips with 540p/720p choices.
Write the prompt for the workflow
For T2V, describe the full scene and one main movement. For I2V, describe how the visible elements should move and what needs review. Do not ask an image-to-video model to reveal exact hidden surfaces that are absent from the starting frame.
Compare two routes fairly
If you test both, use a consistent shot concept and score framing, motion and detail separately. An image-to-video test necessarily includes a source still, so record it. One comparison supports a decision for that shot, not a universal model-ranking claim.
Check facts and generated details
Review face, text, logo, geometry, edges and temporal changes throughout the clip. A starting image does not guarantee those details remain pixel-identical. For product or documentary use, verify each factual claim against trusted source material.
Prompt or planning example
If the product’s opening crop and label placement matter, begin with an authorized still and try an I2V route with subtle motion. If the setting and composition are still open, test a T2V route with the same one-action brief.
Adapt this starting point to the source image, destination and controls shown in the selected route. Keep the original where relevant, then compare the result against the specific details named in your brief. If a detail is factual, verify it from a trusted source before using the image.
Further reading
- spicyapi.ai/models/wan-3-0-prime/text-to-video
Wan 3.0 Prime text-to-video starts from a written prompt; current published duration and resolution options include 2–30 seconds and 480p/720p/1080p, with generated-audio controls. Accessed 2026-10-08.
- spicyapi.ai/models/wan-3-0-prime/image-to-video
Image-to-video requires an opening image and uses a different workflow from text-to-video; motion instructions, duration, resolution and optional audio are endpoint-specific. Accessed 2026-10-08.
- spicyapi.ai/models/vidu-q3-turbo
Vidu Q3 Turbo is image-to-video, not text-to-video; it needs an opening image and prompt and currently documents 1–16-second clips with 540p/720p choices. Accessed 2026-10-08.
Related tools and guides
Model availability and exposed controls may change. Check the current form and source documentation before relying on a specific setting. Review generated material for factual accuracy and suitability before use.