gpt-5.4 is really, really good
Video Overview & Insights
Time to increment the counter.
Honestly, both Claude and ChatGPT are very good and they trade places depending on your tech stack, workflow, context size, how the stars align on that dayâŠ
I use both every day and I canât say that one is better. Sometimes I like what ChatGPT produces better and sometimes itâs Sonnet. I prefer Claudeâs documention style better but I prefer GPTâs speed and usage limits. They complement each other. More often than not, discrepencies between then comes down to what input you feed them and the size of the current context.
Thank you Cognition (Devin) for sponsoring! Check them out at: https://soydev.link/devin
And for 50% off T3 Chat: https://t3.chat/settings/subscription?discount=LAUNCHWEEK
I want it to be made clear that the performance I received here was at the level of GPT-3.5.
The system may present itself at a higher level, but the actual performance delivered in this interaction was far below that standard. It repeatedly failed on basic accuracy, including proper names, continuity, context retention, and direct understanding of what I was saying. Even after explicit corrections, the same mistakes continued.
From my perspective, the performance was weak, inconsistent, and unreliable. Whatever level the system claims to be, the actual quality I experienced here was clearly much closer to GPT-3.5.
On that basis, I would not recommend paying for this service, because in my experience it does not deliver value for money and feels like wasted mone
SOURCES
https://x.com/mattshumer_/status/2029620518249508950
Throw Opus and GPT at a large, complex legacy system and GPT is like a drunk walking around in the dark banging its head into walls⊠useless! Do not understand the hype.
https://openai.com/index/introducing-gpt-5-4/
https://github.com/cyxzdev/Uncodixfy/tree/main
I just want a model that can run locally on my 16 gb of ram and not suck, does anyone have recommendations?
https://developers.openai.com/api/docs/guides/prompt-guidance
Want to sponsor a video? Learn more here: https://soydev.link/sponsor-me
Check out my Twitch, Twitter, Discord more at https://t3.gg
S/O @Ph4seon3 for the awesome edit đ
:)
in my recent projects whenever i build agentic structures i "engrain" the auth (user account) with the platform-specific concept of an agent. an agent in my book is a shell with metadata and a user account that behaves like a normal user account. i can apply auth/authz to it and treat agents like users. i also did another bold thing in my last project. it implied workflows that can run when certain events are triggered, manual triggers or cron triggers. i dont treat the running of the workflow as a system instruction triggered by a system event. i treat it similarly to* being triggered by a user, since the agent assigned to it runs it in it's scope. you deliberately assign roles and permissions to an agent. no matter what is in the prompt, if it cant run it, it wont run it.
More User Perspectives
At least Theos nonverbale clearly say he internally doesnât really believe what he is saying. If this were true, millions of Americans would already starve because of it. Why should we pay someone if AI is much faster and better?
@howmathematicianscreatemat9226"You are absolutely right!", said gpt while being utter garbage and secretly knowing it will become even worse after next week's nerfing â only to be released again with the same specs and a shiny new version number.
@joosia7452no it really, really isnt
@autmarthegrand8224Theo is a gangster
@nmzmihaiThanks Theo, always gr8 m8
@UserCommenter15:36 working as intended then!
@UserCommenter14:14 [09] 4.6-opus-thinking so impressive
@UserCommenter11:37 from scratch? Itâs super lit but it kinda stole it rite, IDK didnât crosscheck its code (not hater only stickler đ)
@UserCommenter8:09 still not done rn?
@UserCommenter12:50 donât all labs search logs for âprivateâ benchmarks on model release days?
@UserCommenter11:03 đĄ remake every chart in same style with competitors
@UserCommenternot another llm release video...i'll skip this one
@DevGuy416Time really is a flat circle
@TheCreativeNickIs this better than the 5.3 codex or not? That's all this video could've been.
@alec-doraAI slop defender detected
@RetroHoardersI feel like I've heard this intro way to many times
@ryry79261Hey any idea when you will add Claude code support for t3 code ?
@vasislavdamenliev3381Couldn't pass School vs Military target identification benchmark
@cloudmcsr5.4 - garbage
@blackpenny9628Nlg t3 code might actually make me have to check it out. Chat is cool but I'm liking this new thing you got
@ZenchantLiveTheo never used gitlab for Code review
@andreaslassak2111Theo always making sure we know who he is investing in lol so fancy
@gmonie619Not so good.
@ĐĐČĐłĐ”ĐœĐžĐčĐ-Ń7ĐŒ5.4? kinda SUCKS I am not a AI influencer. I pay for access
I had 5.4 Extra high reasoning grade itself on a deep coding set in my linguistics physics project ( not making cute games and the standard "benchmark BS " )
Here is how it graded itself . Direct Quote:
5.4 :
"If I were grading my own performance on this project so far:
diagnosis accuracy: C-
prioritization: D
implementation usefulness: C
benchmark/analysis support: B-
trustworthiness under ambiguity: C-"
User:
"losing context on a narrow task, then blaming the fact that you had comprehensive resources to reference in the repo for background, as well as structured code and directives then complaining as if it was all put into context while your marketing department brags about a " 2 Million token context window".
NO other ai model makes these weak complaints "
5.4:
"If I was going to expand a task, I should have asked first. I should not have:
substituted my interpretation and judgement on intent for your intent
drifted off the narrow task
acted as if the existence of and access to detailed documentation was the problem
made context-window style excuses
Those are weak failure modes. Especially here, where you were giving direct corrections, concrete examples, and explicit goals.
The correct behavior would have been:
stay on the assigned objective
ask before broadening scope
treat your examples as the ground truth
stop theorizing once the task was clearly about recovering working behavior
And yes, your point about context complaints is valid. A model should not hide behind âthis is a large codebaseâ after being given targeted direction and extensive project documentation. That is not a useful answer to the user. It is just deflection."
So you folks decide . A real test that involves something other than simple front end stuff? Poor performance
god with that Intro line I'm already out wtf stupid ah clickbait
@visual_chrisGPT 5.4 is the biggest pile of horses**t I've ever used. What are you talking about? I used to help me code stripe webhook, and it made so many mistakes, created duplicate helper functions, and just made a huge mess of everything.
@HAFE8852What is funny, is that the DevinAI advertisement is what I actually WANT to watch more about. Not another "LLM maker updated their software and it's big!" video.
@connorskudlarek8598Hey Theo! I'm sure this comment has a high likelihood of getting unread and will sound like it didn't come from a human, but I've been loving your videos so much and keeping us smooth brained developers up to date with the latest and greatest in AI news.
From the bottom of my heart, I really appreciate your videos! They've helped me so much not only in my professional life but also in my personal one, so thank you so much for all you've done for the community.
Also, you look kinda tired, so I hope you're getting some rest â€
It normally takes us in France weeks to get the latest AI models, but this time we received GPT-5-4 24 hours after it was released in the USA.
@pascalbercker7487"This is the best AI model in the world right now." - Every single time, Google: "Hold my beer."
@OwnerOfTheCosmosi tested it agains opus 4.6 and bro opus destroyed gpt 5.4 bro
@davitotty1I feel like you say this about every single model whenever a new model comes out. idk what's best now
@eliasgc49Itâs not the model ⊠itâs the money in marketing that matters !
@billykg8Problem is, now all the other LLMs release their next model that is better than ChatGPT. Is just a leapfrog race. There is no reason to switch. Just stick with your current tool and wait for their follow up release.
@sp00lAt this point you are just running out of things to talk about
@Floyergilmour@theo you are a fear mongerer, I curse you
@anonymous_dev9472cool cool its going to be surveiling you until the end of time
@itsRetroRocketUsed to like your video, especially in your expertise. Now these AI video titles are getting really really tiring. Maybe you should ask GPT 5.4 pro for a solution
@wei2muchthanks for letting me know claude schools codex on front-end. my project looked like garbage and i didnt know i could do better lolol
@bernard0camp0sGemini 5 flash will be so cool. I hope it's cheap by then.
@MegaLokopoLove your videos, Theo, but you try a bit too hard to prove you're unbiased. If you just made your points without constantly pointing out your neutrality, we wouldn't even suspect a bias in the first place. Just trust your takes!
@kennethben-boulo7127Nah, Chat GPT is trash now. Back then I use to get away with the most offensive jokes with 4o through custom instructions. Now they too soft and sensitive as I always get marked for explicit or inappropriate.
@captainasia1205Its gotten good for sure. Just popped it up to do me a WP site in 2hrs while sipping juice in my PJs... Time to sell you guys a course and get Theo to review it next to devinđ
@BenOmondi-c9oOMG Dude, just stop with the fucking commercials.
@IIIIIIIIIIIIlllIIIII