Generating realistic fire, water and smoke used to mean hours of simulation and a specialist. Now a well-phrased prompt can get you most of the way there. But "most of the way" is the whole story with text-to-VFX in 2026 — it's genuinely useful, and genuinely not a magic button. Here's where it earns its place.
What text-to-VFX does well today
Models trained on physics can now produce convincing elemental effects — flame, splash, smoke, dust, sparks — from a description, and drop them into a shot far faster than a from-scratch simulation. For establishing energy, atmosphere and background elements, it's a real time-saver, and it makes previs of effect-heavy sequences almost instant.
Where it still needs a hand
Precise interaction is the hard part: fire that has to wrap a specific product, water that must respect a real object's geometry, smoke that reacts to a character's exact movement. Generated effects are great at looking right in isolation and less reliable at integrating with a plate at pixel level. That's still where a compositor's craft decides whether the shot sells.
The realistic workflow
We treat text-to-VFX as a fast first layer, not the finish. Generate the element, then art-direct, relight and composite it into the shot by hand so it obeys the scene's light and physics. On effect-driven sequences, that hybrid — AI for speed, humans for integration — is dramatically faster than the old way without the "floating effect" look that gives cheap work away.
The bottom line
Text-to-VFX is real and worth using — as an accelerator, not an autopilot. Lean on it for atmosphere, background elements and previs; keep a human compositor on anything that has to interact with your product or hero, and the seams stay invisible.

