I Made an AI Short Film in 12 Hours. Vegas Let Loose.
- Jesse Hayes
- Mar 27
- 15 min read
Everyone has an opinion about AI-generated media. Most of those opinions are formed from screenshots, hot takes, AI slop exposure and second hand dread. I recently found myself reading a long comment thread on LinkedIn, an entertaining mix of denial, acceptance, end of the world foreshadowing and just general dismissiveness; a colourful mix of human responses to our new AI overlords. Yet, I wanted my opinion to come from somewhere messier and more useful, from actually doing it. The video is posted down the bottom if you want to skip the field report.
So this is my first formal step into the arena, or into the Matrix as you will. No grand plans, no complex intent; just get in there and see what it's all about and come out with an informed opinion

There's a version of this post where I wax poetic about the future of cinema, drop some Godard reference or wrap it all up in a neat manifesto to make myself sound more like an artistic film buff. That's not this, at all. This is a raw field report. Specific, honest, and a little rough around the edges, because the little short film I made was too.
Vegas Let Loose is a three-minute car chase set in Las Vegas. Two government operatives. Four hitmen. The Las Vegas Strip at night. The car of my dreams. A score that owes something to Fast and the Furious: Tokyo Drift and makes absolutely no apology for it. It's cinematic. It moves. It has atmosphere. And action. Yet, It also has characters whose likenesses drift in and out of consistency, a lead actor who occasionally bears absolutely no resemblance to his real world counter part, and a supporting actor that drifts to the likeness of his partner, cuts that don't quite hold, and some physics that operate by their own rules entirely.
"The flaws in your first project are the curriculum for your second one. You need to make mistakes, find the boundaries, the limitations, and the work-arounds."

And now I'm showing it off anyway. Proudly, actually. Because the point here was never to make something perfect, it was to jump into the deep end, late to the AI content creation party, make something real, learn from it, and tell the truth about what that experience is like from the inside. Better late than never!
"Good Luck, Have Fun, Just Try." Just get in there and have a go.

Why 12 Hours? Why Now?
I gave myself a hard ceiling of ten to twelve hours. Not because I was short on time, I wasn't particularly; but because the point was to explore only the surface level, as if I were the every-person. Someone with no real experience creating AI content or media, who woke up one morning with an idea and decided to just make a thing. The time constraint was philosophical as much as practical: force the decisions, get it done, resist the rabbit hole. Get in, have some fun, don't take it too seriously. Good luck, have fun, just try.
AI creative tools are advancing fast enough that the trap right now is the pursuit of perfect before you understand the fundamentals. Anyone who's spent five hours tweaking Kling prompts for shot consistency on their first project (or experiment) knows exactly what I mean. You find yourself optimising for the wrong things when you don't have enough baseline experience to know what actually matters, or even how the processes work. It's natural to try to apply logic to fix some of the craziness that comes out of the AI. Why oh why, is this car driving backwards, and why oh why did the gunmen just warp through space and time into the backseat of the Mercedes to have a fist fight I didn't ask for!?
Without understanding how these exciting new tool work, you'll quickly find yourself doing expert-level problem solving without any of the experience required to identify whether the problem you have is even worth solving, or even the right problem. That's not overly efficient or effective. You wouldn't do it in any other discipline, neither should you here.
So: twelve hours hard limit Jesse, do not take the red pill; yet. Foundational tools only, no advanced workflows, no spiralling into technique rendering credit debt. Build the pipeline, get reps on the tools, identify the actual gaps versus the perceived ones. Create something fun, and write the report.
The Stack — What I Used and What I Actually Thought
Here's an honest breakdown of every tool in my infant pipeline, not a sponsored overview, just what it was like going hands-on with each one of these tools under a time constraint and no prior workflow.
Concept
Written Narrative
Scene beats, shot list, sequence logic. This is where any project lives or dies, no tool compensates for a weak concept brief. In my case, the brief was offensively simple: "Sexy car, Las Vegas car chase, small shootout, no overarching story. Lock, stock and two smoking tires." Crude, but it worked. The point is, you need something anchoring every decision downstream. Even a napkin-level brief keeps you from drifting. Simple word document, this happens then this happens then that happens; end.
Image Generation
Google Banana Pro 4k
Truly remarkable technology. The implications of what's sitting in this tool alone are staggering, and that's a completely different blog post. For my purposes here: environments, vehicles, establishing shots as start frames, excellent! Where it fell apart was character and object consistency across generation passes, and that failure was entirely mine. Start with poor technique, try to cut corners on reference or concept imagery that informs the world building, and you will be fighting it at every single pass with compounding frustration. The tool didn't fail me. Appropriate workflow didn't exist at this time.
Image Refinement
Adobe Photoshop
Still essential. Resizing frames, micro adjustments, removing strange little artefacts that defy any reality I can comprehend, background objects that seem to have developed their own narrative ambitions, fingers that took creative liberties (Rare these days, but it happens). AI generation at this stage can produce things that genuinely need a human eye and a clone stamp. This is the step most AI filmmaking breakdowns gloss over. They shouldn't. A small chunk of my hours went here, and every minute was worth it. Skip it and the problems don't just persist, they multiply downstream once you feed those images into video generation. Quality assurance, ignore it at your own peril!
The Platform
Higgsfield AI
This is where I found the most untapped depth in the entire pipeline, and honestly where I exercised the most restraint, because the rabbit hole here is real. I fought my way through a half dozen other platforms, ended up with some subscriptions I really don't need. But this is where I landed, and it's awesome. Angle variation and shot development is excellent; the breadth of visual options you can generate from a single image is genuinely impressive. But the real power behind Higgsfield, the tools; Cinema Studio, advanced motion tools, voice, performance direction and the rest. I didn't touch any of it, very intentionally, however tempted. That's all Experiment #02 territory. What I will say is: if you're using Higgsfield at a surface level and it feels limited and that results in undesired output, that's almost certainly a you problem, not a tool problem. The depth and the control is in there.
Video Generation Model
Kling 3.0
The anchor of the pipeline. Single-shot prompting is intuitive; multi-shot auto-generation gives you useful rough coverage quickly. Motion quality at 1080p is genuinely impressive, vehicle movement especially in most cases. But prompt craft matters enormously here, and I mean that in the most literal way. A vague prompt produces vague video. A precise one, specific camera angle, character positioning, lighting, approximate focal length, motion intention; produces something cinematic as long as you don't confuse the absolute hell out of the poor little AI agent. The intent behind the prompt, and understanding the language of the reality the AI exists/or doesn't exist in, is everything.
Worth noting: I tested several other video models during the process. But landed on Kling 3.0. Arguably Higgsfield Cinema Studio 3.0 should have been my go to, but i considered that to be an advanced tool that would have broken my exposure to the basics. The pace of development in this space means model superiority is measured in months, not years. Seedream 2.0 will be the new leader, by far; notwithstanding the legal issues restricting its release.
What we're going to see is a decade of relentless industry bunny-hopping as the technology just keeps going, evolving rapidly from iteration to iteration .
Video Edit
Adobe Premiere Pro
Assembly, pacing, basic LUT grading. This is still human craft, full stop. The edit is where the film actually becomes a film, the structural and rhythmic decisions are entirely human, and I suspect they will be for a while longer. Nothing grand in my simple edit here; I had about four hours left on the clock by the time I reached this stage. Consider it a speed run. Some might call it scrappy. I'd call it an edit made by someone who had already spent eight hours learning what they didn't know in a platform they've never used.
Score
Eleven Labs
No surprise here, but it earned its praise. The custom music track landed well for what the film needed: good bass, low-end focus so as not to drown out engines, tires, and the occasional action one-liner ("Cash out!"). The skill here is giving the platform a useful brief on scene timing and tonal intent. The output quality ceiling is higher than I expected. My audio application in this film is about as basic as it gets, essentially a music video approach; so I'm not going to pretend I explored what proper cinematic sound staging with these tools looks like. That's a completely different discipline.
Upscale
Topaz Video AI
Quiet workhorse, genuinely impressive, non-negotiable. AI video is outputting at 1080p right now and will likely stay there for a while on account of not turning the entire Earth into a server farm. The upscale to 4K is the step that gives finished work actual visual weight. Where some outputs have insane detail straight out of the model, others can be soft in ways that only reveal themselves at full scale. The Topaz pass recovers enough to matter. Don't skip it. But it does indicate an underlying skill I need to circle back to, understanding what prompts of lack of workflow did in fact cause any quality drop off.
Export
HandBrake
Web-optimised 4K and 1080p file delivery. Fast, clean, no surprises, no drama. Exactly what you want at the end of two six-hour content creation sessions, to finish.
What the Process Feels Like From the Inside:
Here's the rub, and the part I think most ill-informed commentary on AI filmmaking at this stage in its relative infancy gets wrong: firstly, good to understand that this technology is still in its relative infancy, and second; it's not effortless. It's a different kind of effort. You get out what you put into it. Respect the workflow, and you'll get consistency. Lazy effort, and you'll have a lovely bowl of AI slop, delicious early AI era Will Smith spaghetti. If you think the technology will do the fundamental work of creative governance and direction, then you're going to have a very expensive flow of content that may make sense as a single frame, but mostly disconnected from itself and useless when put together.
Traditional filmmaking front-loads production, much of the grind and the art is in the work before a single camera rolls. Coordination, screen writing, logistics planning etc; and that all comes after the initial planning of the script and the not so glamorous planning work. AI filmmaking at least now, front-loads the discipline required, then back-loads selection, correction, and editorial judgment; automating the in between.
"The output is only as strong as the quality of decisions you make at every node. The tools amplify clarity; they also amplify vagueness."
When I wrote a vague prompt, I got vague video. When I wrote a precise one, with specific camera angle, character positioning, lighting direction, focal length approximation, and motion description, I got something more cinematic, and predictable. The gap between those two outcomes is enormous. The tool didn't change. The intent behind the prompt did. And the level of respect I had for the tool I was using. This isn't going to change any time soon, AI opens a lot of doors, but the imperative need for a human director, who acts as both cast and crew; is the difference between content of value, and content that exists only for the sake of existing. And honestly, my 3 minute video falls into that latter category. As it should. I have to earn it, work for it, invest the time, and get the skills; to be able to create something worthy.

The Honest Breakdown: What Worked, What Didn't
Where It (Or I) Delivered
Cinematic atmosphere delivery after locking in the start frames
Kling motion quality on vehicle shots was great at times
Score tone and energy from Eleven Labs custom track
Higgsfield angle/shots variation breadth and control over image generation
Topaz upscale quality for 4K delivery
The AI having patience for my lack of workflow discipline
Where It (Or I) Broke Down
AI being driven by a driver that doesn't have a licence
Character consistency shot-to-shot absent character designation
Background vehicle/object coherence
The cuts that don't and cannot land together on inconsistent narrative flow
Limited scene variety from image-to-video pipeline
All of the above, easily resolved with tools available
Character consistency is the dominant technical challenge I faced in AI filmmaking in this first experiment. But, this is really not on account of the technology by any means. Higgsfield contains more advanced tools such as character creation which allows you do lock in character appearance, more feature rich video direction tools, and most of all; as a first time entrant I simply didn't have the knowledge or the experience to understand the correct workflow and discipline to get the output I was seeking.
Every tool in the stack I used operates in its own world to a degree. So with no persistent model of my characters, or the world, or the objects in that world; Higgsfield, and Kling creating stable outcomes that can see beyond its direct reference material isn't something that someone should expect of it. It's not magic, It's a process. So both I and the Ai tools, were really just approximating visual coherence through reference images, an undisciplined shotgun approach hoping for the best with some guideline prompts and some boundaries. It works, up to a point. And that point is very noticeable.
The solution isn't a better prompt, neither is it a new magic wand, sudden evolution of the technology or the models. It's a better, or should I say; proper workflow. Back to the basics, without irony, as real world cinema has mastered over the century. Character sheets, dedicated reference and or concept art, custom fine-tuned models for recurring characters, start and end frames locked down to a repeatable process, more detailed instruction of individual shots, and just good old fashion knowledge and experience in understanding that a tool must be respected; it doesn't matter if you're using a paintbrush, a pencil or these new AI tools.
"The solution isn't a better prompt, and it's not a new magic wand, or a sudden evolution of the models. It's a proper workflow — the same discipline real-world cinema has built over a century."

We've all already had the conversation: "Is this going to replace filmmakers?"
It's the wrong question, at the moment. But it's the one people can't help asking, so let's actually engage with it properly because the answer is more interesting and more nuanced than either the panic or the dismissal allows for.
What these tools do is collapse the gap between idea and image. That gap used to require millions of dollars of capital, industry connections, equipment, crew, locations, permits, insurance, time, or a significant amount of industry access. Hollywood, as it has operated for the better part of a century, has one of the highest bars to entry of any industry. There are gatekeepers. They control what gets made and who gets to make it. That's not a cynical take; it's just the structural reality of an industry built around enormously expensive physical production where there's a finite list of actors repeatedly getting the acting chops.
Sure, Netflix and the independent film movement dropped the entry point meaningfully. The relentless societal demand for a constant, glutinous supply of content to consume has forced an unprecedented volume of production. And we've watched what happens when that demand outpaces creative intent, as what Hollywood has largely delivered over the past decade is, with some notable exceptions, quite dismal and at times depressing.
The Marvel era of increasingly pointless blockbusters, in essence each the exact same movie and plot line, with different characters. Need I mentioned the systematic dismantling of IP like Star Wars, Star Trek, Alien, Indiana Jones and the many other poorly treated and commercially abused intellectual properties. Content produced not because someone had a story that had to be told, but because quarterly projection required it.
What most ill-informed commentary on AI filmmaking gets wrong: it's not effortless. It's a different kind of effort. You get out what you put in.
AI doesn't fix that problem. A tool that makes production cheaper and faster in the hands of people with no creative vision just produces cheaper, faster, slop. The gatekeeping problem wasn't really about resources in the first place, it was about the conflation of resources and taste, as though the ability to finance a film was evidence of the ability to make a meaningful one. Just because you can, doesn't mean you should.
What actually changes with AI is access for the people who were never the problem to begin with. The millions of people out there, everyday people of varied and entirely unrecognised creative ability, who have spent their entire lives digesting films and television and may have something genuinely worth saying, but have never had the means to say it in moving images. Those people can now wake up on a Tuesday morning and start making the film that's been living in their head for years. The creative vision was never their problem. The production machine simply presented an near impossible, or intimidating barrier for entry.
These new tools do the heavy physical lifting that used to gate entry. But the creative judgment is still entirely human, and in some ways it's more demanding now, because you're making those decisions faster, with more options, and with fewer of the natural friction points that used to force deliberate thinking. The decisions that a production schedule and budget used to make for you now have to come from somewhere else. They have to come from you. As well as the self imposed discipline to get the job done, and done well.
"These tools don't replace filmmakers. They remove the production ceiling that was keeping most filmmakers from making films."
Will AI also produce an enormous amount of mediocre content? Obviously. It already is, terrifyingly so. But here's the thing, mediocre content isn't new. Social media birthed into the world a broad range of toxic and fundamentally detrimental socioeconomic and psychological detriments upon us, mostly on account of the orgy of absolutely mind numbingly pointless content. So I'm not going to assert that AI slop is any more dangerous than human slop. Arguably, if the world continues becoming more saturated with pointless and soulless content, there's now a very real possibility that with the potential content AI can deliver; truly exceptional work might actually stand out a great deal more against the background noise.
The noise floor rises, but so does the contrast of anything genuinely worth watching. One more thread worth pulling on: Is AI going to replace actors? That's an entirely different and considerably more fraught conversation, with real human stakes and legitimate grievances on multiple sides. I'll leave that for another post, where it deserves the full weight of attention. That may be a more worth point for experiment #02?!
What's Coming Next Then?
I won't overreach here. Experiment #01 aka Vegas Let Loose wasn't designed to be a map to anywhere, it was designed only to give me a starting baseline. A lick, a taste, a whiff of what these tools are actually capable of when a person with a limited brief and a self-imposed deadline sits down and uses them.
In short: "Hey look! I made a car chase with explosions with AI in twelve hours". It's flawed, but it's mine, and it was genuinely constructive. I understand the landscape in a way I didn't before I did it. And I now have an informed albeit basic position on the technical side. I have specific, evidence-based knowledge of what to do differently next time. That's exactly what first experiments are for. Not validation, initial exposure.
You do one thing, then the next thing, and after you've created enough things; you may very well end up with something truly inspired, that you love, that other people might also love too; and in the process you've amassed new skills, new knowledge, capabilities and a maybe a new source of passion.
At the end of the day.... That's the point, the tools change, but the core disciplines haven't. So just have fun out there. If you're aspiring to make something that other people should take seriously, best you take creating it seriously too; that begins with good old fashion discipline, workflow, a tangible investment of time, respecting processes, and making sure you know exactly why you're making whatever it is you're making from the get go.
Some things never change.
If you're sitting on an AI project idea and waiting until the tools are better, that will do the entire thing for you, or waiting until the creative process does indeed become entirely automated, you've probably missed the point. AI is a tool, using it is how you get to better. There is real human skill required. It's not magic. You just have to give it a crack, make the thing, and enjoy it. I did.... I hope you enjoy watching Vegas Let Loose in the optimum environment, in 4k with some full range speakers or decent headphones. I don't know about you, but I do love some bass!
Thank you for taking the time to read this field report. If you enjoyed it and the video, or even if you didn't; please feel free to drop me a comment or a reaction in the LinkedIn thread. Also happy to have a conversation if this has inspired you to have a go or if it just sparked interest. Follow along if you're into it, what comes next? Experiment #02:
Author & AI Filmmaking Test Subject - Jesse Hayes | March 2026





