2026-09-04
Why Do My Midjourney and Sora Prompts Look So Inconsistent?
Image and video models generate a new interpretation every time. Without explicit scene, subject, style, and composition details, each generation is a fresh guess - which is why two runs of a product shot of my skincare bottle can look like they came from different photographers.
Text prompts have some tolerance for vagueness because a reader fills gaps with reasonable assumptions. Image and video prompts do not have that cushion. Every unstated detail becomes a random visual choice.
Before
product shot of a skincare bottle, premium
After
Editorial product photography. A matte white 50ml glass dropper bottle with a black cap, no visible label text, centered on a light gray marble surface. Soft natural window light from the left, shallow depth of field, muted tones, no harsh shadows. Bottle occupies the lower two-thirds of the frame, negative space above for text overlay.
The four things that need to be explicit
Subject specifics. Not a skincare bottle - a matte white 50ml glass dropper bottle with a black cap, no visible label text. Vague subject descriptions get a different bottle shape, material, and cap every single run.
Scene and lighting. On a marble surface, soft natural window light from the left, shallow depth of field. Without this, lighting and background shift wildly between generations, which is why a set of matching product shots often does not match at all.
Style reference. Editorial product photography, muted tones, no harsh shadows gives the model a visual category to anchor to. Without a style anchor, tone can swing from glossy commercial to flat stock-photo between runs.
Composition. Centered, bottle occupying the lower two-thirds of frame, negative space above for text overlay. If you need the output to work inside a specific layout (an ad, a landing page hero), composition has to be stated, not assumed.
Why this matters more for video
Sora and other video tools compound the same problem across every frame. A vague prompt does not just risk one inconsistent image. It risks visual drift within a single generation, since there is no single correct interpretation the model is holding onto throughout the clip.
Compile a rough image or video idea in Studio. Vault can save the style anchor so the fifth product shot starts from the same visual foundation as the first.
FAQ
Related problems
Same class of brief failure, different search phrasing.