Text to Video
- Best for
- Discovering a new scene directly from a written direction.
- Choose when
- The subject, composition, or opening frame has not been approved yet.
- Key difference
- The prompt must define both the opening visual state and the motion.
Text-to-video starts with language. Image-to-video starts with a frame. The better choice depends on whether the job needs visual discovery or tighter control over an existing subject, composition, or product shot.
Method: Model and workflow specifications are drawn from the current RedVideo configuration catalog. Live availability, supported settings, and generation quotes can change; confirm them in Studio before generating.
Both routes create AI video, but they solve different starting problems. Pick the side that matches what is already decided before comparing individual models.
| Capability | Text to VideoStart with text | Image to VideoStart with an image |
|---|---|---|
| Starting input | Written scene direction | One compatible source image |
| Strongest use | Visual discovery and new scenes | Continuity from an approved frame |
| Prompt focus | Subject, composition, action, camera, and atmosphere | Motion, camera behavior, pace, and intended change |
| Opening composition | Invented by the model | Anchored by the source image |
| Current access | Current models require Pro or Ultra | Includes eligible Core settings plus paid model routes |
Swipe horizontally to compare every configuration.
Text-to-video is useful when the scene exists primarily as an idea. A prompt can establish subject, environment, camera direction, mood, and action without requiring a finished source frame.
Image-to-video is useful when the first frame already matters. A portrait, illustration, product image, or generated keyframe gives the model a stronger visual anchor while the prompt directs motion and camera behavior.
A text-to-video prompt carries more of the visual specification, so clear descriptions of framing, action, lighting, and camera movement can reduce ambiguity. Results can still vary because the model is inventing the opening composition as well as the motion.
With image-to-video, the frame supplies appearance and layout. The prompt can concentrate on what changes over time: the movement, pace, expression, atmosphere, and camera path. This can make controlled variations easier, although source-image quality still shapes the result.
For concept exploration, story beats, and shots without a fixed visual identity, begin with text to video. For product motion, character continuity, artwork animation, and shots that must begin from an approved frame, begin with image to video.
RedVideo exposes both workflows across multiple model configurations. Availability, duration, output settings, and quotes vary by model, so the studio remains the source of truth for a specific generation.
Select the model and workflow that fit the idea, review the live settings and quote, then start creating.