SpaceXAI Adds References to Grok 'Imagine Video 1.5' for Consistent Face, Voice, and Scene Control

By the summer of 2026 most generative video systems had settled into a familiar set of strengths and constraints. 

They could turn a still image into a short clip with synchronized sound effects, ambient noise and dialogue in a single pass. Motion had become more coherent over several seconds, and physics looked less broken than in earlier models. 

Yet three practical problems remained common. 

Starting from text alone often produced weaker results than beginning with a carefully chosen image. Keeping a specific face, product or location consistent across multiple shots required heavy post-processing or repeated trial and error. 

And output resolution rarely rose above 720p without additional upscaling steps that introduced their own artifacts.

xAI’s Imagine Video 1.5, released in general form on 16 June 2026, addressed several of the earlier shortcomings. 

The model introduced improvements in the believability of weight and momentum, reduced warping over the length of a clip, and generated clearer speech that stayed better aligned with lip movement and action. 

Audio, including sound effects and ambient layers, was produced together with the visuals rather than added later. 

A Fast variant cut generation time for a six-second 720p clip to roughly 25 seconds, down from more than 40 seconds in the prior version. 

Clips could run from one to about fifteen seconds at 24 frames per second, with aspect-ratio options that covered the usual landscape, portrait and square formats. At launch the highest native resolution was 720p, and the primary workflow remained image-to-video. 

The model also ranked near the top of independent image-to-video arena evaluations at the time.

On 31 July the same model received a further set of capabilities

Text-to-video became available, allowing a written prompt to produce a clip without any starting image. 

Native 1080p output was added for both the text-to-video and image-to-video paths. Image references, limited to a maximum of seven per generation, let users lock individual elements such as a face, a product or a background location. 

Video file

These references can be mixed: one character can be held constant while the setting changes, a scene can stay fixed while different characters appear, or both character and scene can remain stable while only the action varies. 

Voice references can be supplied alongside a character image so that the same voice continues across successive shots. The combination is intended to reduce the need for manual consistency fixes when building short multi-scene sequences.

Access to the new image and voice reference features began in the United States for SuperGrok Plus and SuperGrok Heavy subscribers on the web interface at grok.com/imagine and on the iOS app, with a broader rollout to remaining tiers planned over the following days. 

Text-to-video and native 1080p became generally available across web, iOS and Android at the same time. 

In the xAI API the model continues under the name grok-imagine-video-1.5. Image references, text-to-video and 1080p are supported directly; voice references remain available on request. Generation parameters include duration (commonly six to fifteen seconds), aspect ratio and resolution selection.

The practical effect is most visible when users want a recurring character or product to appear in different environments or with different actions. 

Video file

Even with these additions, the update shows a limit, as it remains focused on short-form video generation, producing completed clips that last only a few seconds rather than extended sequences.

Beyond these new control features, the core capabilities introduced in the June release remain largely unchanged. Motion quality, physics simulation, and synchronized audio generation continue to perform at a similar level, indicating that the underlying model has not undergone a fundamental overhaul.

Instead, the update reflects a broader trend of iterative refinement rather than a complete redesign. 

Early examples shared publicly demonstrate noticeably stronger consistency, with the same character retaining facial identity and voice across different environments, or a single object remaining stable as camera angles, movement, and lighting change around it. Since each generated clip is still relatively short, creating longer videos continues to rely on stitching together multiple generations, often by using the final frame of one clip as the starting point for the next.

Published