IMAGEGEN STUDIO · PRACTICAL GUIDE
Text-to-Video Prompt Structure: Subject, Action, Camera
A text-to-video prompt is a short shot brief. Identify the subject, one observable action, the setting and a camera view that supports the action. Add duration or audio only after checking that the selected endpoint provides those options. Clear structure helps review a clip; it does not guarantee exact motion.
A four-part shot brief
1) Subject and location. 2) One action with a beginning and end. 3) Camera framing or movement. 4) Visual details to retain. For example: “A small paper boat on a rain puddle; it drifts slowly from left to center; locked medium view; keep the curb and reflection visible.”
Avoid conflicting direction
Do not request a locked frame and a fast orbit at once. Avoid multiple primary actions unless the sequence is ordered and the clip has room. Use concrete movement rather than “make it cinematic.” If the scene is complex, split it into separate shots.
Add duration and resolution from the active route
Wan 3.0 Prime text-to-video currently documents 2–30 seconds and 480p, 720p or 1080p choices. Seedance 2.0 has its own endpoint-specific duration, resolution, aspect-ratio and audio settings; consult its current parameter table. Check the ImageGen Studio form because exposed controls may differ.
Decide whether audio belongs in the shot
If the endpoint has audio controls, describe only sounds that support the action. Keep dialogue short and review timing and intelligibility. Wan 3.0 Prime and Seedance 2.0 text-to-video each document audio controls, but the controls are not necessarily identical.
Review the clip as a sequence
Check opening, middle and ending for subject drift, unwanted camera motion, altered text and confusing transitions. If a failure is isolated, change only the relevant instruction and compare. Keep the result labeled as illustrative when it is not a factual depiction.
Prompt or planning example
A red kite rises above a grassy hill in a steady breeze. One person holds the string in the lower foreground. Wide fixed view; the kite climbs slowly and remains visible. No cut, orbit or extra people. Select a duration only from the active endpoint’s options.
Adapt this starting point to the source image, destination and controls shown in the selected route. Keep the original where relevant, then compare the result against the specific details named in your brief. If a detail is factual, verify it from a trusted source before using the image.
Further reading
- spicyapi.ai/models/wan-3-0-prime/text-to-video
Wan 3.0 Prime text-to-video starts from a written prompt; current published duration and resolution options include 2–30 seconds and 480p/720p/1080p, with generated-audio controls. Accessed 2026-10-08.
- spicyapi.ai/models/seedance-2-0/text-to-video
Seedance 2.0 has a text-to-video route and endpoint-specific duration, resolution, aspect-ratio and audio controls. Read its current parameter table before naming values. Accessed 2026-10-08.
- developers.google.com/search/docs/fundamentals/using-gen-ai-content
AI assistance is not automatically disallowed, but publishing many pages without added user value can violate scaled-content-abuse policies; prioritize accuracy, quality and relevance. Accessed 2026-10-08.
Related tools and guides
Model availability and exposed controls may change. Check the current form and source documentation before relying on a specific setting. Review generated material for factual accuracy and suitability before use.