free web page hit counter
đŸ›Ąïž
Copyright Notice: This video is officially sourced and embedded from YouTube. For all copyright inquiries, reports, or removals, please contact YouTube's legal team here.
Theo - t3 gg

Theo - t3 gg

553,000 subscribers

⏱ 👁 154,014 views

gpt-5.4 is really, really good

Video Overview & Insights

Time to increment the counter.

Honestly, both Claude and ChatGPT are very good and they trade places depending on your tech stack, workflow, context size, how the stars align on that day


I use both every day and I can’t say that one is better. Sometimes I like what ChatGPT produces better and sometimes it’s Sonnet. I prefer Claude’s documention style better but I prefer GPT’s speed and usage limits. They complement each other. More often than not, discrepencies between then comes down to what input you feed them and the size of the current context.

— @jonathansaindon788

Thank you Cognition (Devin) for sponsoring! Check them out at: https://soydev.link/devin

And for 50% off T3 Chat: https://t3.chat/settings/subscription?discount=LAUNCHWEEK

I want it to be made clear that the performance I received here was at the level of GPT-3.5.
The system may present itself at a higher level, but the actual performance delivered in this interaction was far below that standard. It repeatedly failed on basic accuracy, including proper names, continuity, context retention, and direct understanding of what I was saying. Even after explicit corrections, the same mistakes continued.
From my perspective, the performance was weak, inconsistent, and unreliable. Whatever level the system claims to be, the actual quality I experienced here was clearly much closer to GPT-3.5.
On that basis, I would not recommend paying for this service, because in my experience it does not deliver value for money and feels like wasted mone

— @Infinity-roamer

SOURCES

https://x.com/mattshumer_/status/2029620518249508950

Throw Opus and GPT at a large, complex legacy system and GPT is like a drunk walking around in the dark banging its head into walls
 useless! Do not understand the hype.

— @ps170cyber

https://openai.com/index/introducing-gpt-5-4/

https://github.com/cyxzdev/Uncodixfy/tree/main

I just want a model that can run locally on my 16 gb of ram and not suck, does anyone have recommendations?

— @franciscosilva2135

https://developers.openai.com/api/docs/guides/prompt-guidance

Want to sponsor a video? Learn more here: https://soydev.link/sponsor-me

28:40 "Searching web for "Sea Shanty" puzzle solutions" 💀

— @Happ1ness

Check out my Twitch, Twitter, Discord more at https://t3.gg

S/O @Ph4seon3 for the awesome edit 🙏

:)
in my recent projects whenever i build agentic structures i "engrain" the auth (user account) with the platform-specific concept of an agent. an agent in my book is a shell with metadata and a user account that behaves like a normal user account. i can apply auth/authz to it and treat agents like users. i also did another bold thing in my last project. it implied workflows that can run when certain events are triggered, manual triggers or cron triggers. i dont treat the running of the workflow as a system instruction triggered by a system event. i treat it similarly to* being triggered by a user, since the agent assigned to it runs it in it's scope. you deliberately assign roles and permissions to an agent. no matter what is in the prompt, if it cant run it, it wont run it.

— @silviuiacob4013

More User Perspectives

@

At least Theos nonverbale clearly say he internally doesn’t really believe what he is saying. If this were true, millions of Americans would already starve because of it. Why should we pay someone if AI is much faster and better?

@howmathematicianscreatemat9226
@

"You are absolutely right!", said gpt while being utter garbage and secretly knowing it will become even worse after next week's nerfing — only to be released again with the same specs and a shiny new version number.

@joosia7452
@

no it really, really isnt

@autmarthegrand8224
@

Theo is a gangster

@nmzmihai
@

Thanks Theo, always gr8 m8

@UserCommenter
@

15:36 working as intended then!

@UserCommenter
@

14:14 [09] 4.6-opus-thinking so impressive

@UserCommenter
@

11:37 from scratch? It’s super lit but it kinda stole it rite, IDK didn’t crosscheck its code (not hater only stickler 😇)

@UserCommenter
@

8:09 still not done rn?

@UserCommenter
@

12:50 don’t all labs search logs for ‘private’ benchmarks on model release days?

@UserCommenter
@

11:03 💡 remake every chart in same style with competitors

@UserCommenter
@

not another llm release video...i'll skip this one

@DevGuy416
@

Time really is a flat circle

@TheCreativeNick
@

Is this better than the 5.3 codex or not? That's all this video could've been.

@alec-dora
@

AI slop defender detected

@RetroHoarders
@

I feel like I've heard this intro way to many times

@ryry79261
@

Hey any idea when you will add Claude code support for t3 code ?

@vasislavdamenliev3381
@

Couldn't pass School vs Military target identification benchmark

@cloudmcsr
@

5.4 - garbage

@blackpenny9628
@

Nlg t3 code might actually make me have to check it out. Chat is cool but I'm liking this new thing you got

@ZenchantLive
@

Theo never used gitlab for Code review

@andreaslassak2111
@

Theo always making sure we know who he is investing in lol so fancy

@gmonie619
@

Not so good.

@ЕĐČĐłĐ”ĐœĐžĐčЗ-э7ĐŒ
@

5.4? kinda SUCKS I am not a AI influencer. I pay for access
I had 5.4 Extra high reasoning grade itself on a deep coding set in my linguistics physics project ( not making cute games and the standard "benchmark BS " )
Here is how it graded itself . Direct Quote:
5.4 :
"If I were grading my own performance on this project so far:

diagnosis accuracy: C-
prioritization: D
implementation usefulness: C
benchmark/analysis support: B-
trustworthiness under ambiguity: C-"

User:
"losing context on a narrow task, then blaming the fact that you had comprehensive resources to reference in the repo for background, as well as structured code and directives then complaining as if it was all put into context while your marketing department brags about a " 2 Million token context window".
NO other ai model makes these weak complaints "
5.4:
"If I was going to expand a task, I should have asked first. I should not have:

substituted my interpretation and judgement on intent for your intent
drifted off the narrow task
acted as if the existence of and access to detailed documentation was the problem
made context-window style excuses
Those are weak failure modes. Especially here, where you were giving direct corrections, concrete examples, and explicit goals.

The correct behavior would have been:

stay on the assigned objective
ask before broadening scope
treat your examples as the ground truth
stop theorizing once the task was clearly about recovering working behavior
And yes, your point about context complaints is valid. A model should not hide behind “this is a large codebase” after being given targeted direction and extensive project documentation. That is not a useful answer to the user. It is just deflection."

So you folks decide . A real test that involves something other than simple front end stuff? Poor performance

@hitmusicworldwide
@

god with that Intro line I'm already out wtf stupid ah clickbait

@visual_chris
@

GPT 5.4 is the biggest pile of horses**t I've ever used. What are you talking about? I used to help me code stripe webhook, and it made so many mistakes, created duplicate helper functions, and just made a huge mess of everything.

@HAFE8852
@

What is funny, is that the DevinAI advertisement is what I actually WANT to watch more about. Not another "LLM maker updated their software and it's big!" video.

@connorskudlarek8598
@

Hey Theo! I'm sure this comment has a high likelihood of getting unread and will sound like it didn't come from a human, but I've been loving your videos so much and keeping us smooth brained developers up to date with the latest and greatest in AI news.

From the bottom of my heart, I really appreciate your videos! They've helped me so much not only in my professional life but also in my personal one, so thank you so much for all you've done for the community.

Also, you look kinda tired, so I hope you're getting some rest ❀

@JohnestOfPauls
@

It normally takes us in France weeks to get the latest AI models, but this time we received GPT-5-4 24 hours after it was released in the USA.

@pascalbercker7487
@

"This is the best AI model in the world right now." - Every single time, Google: "Hold my beer."

@OwnerOfTheCosmos
@

i tested it agains opus 4.6 and bro opus destroyed gpt 5.4 bro

@davitotty1
@

I feel like you say this about every single model whenever a new model comes out. idk what's best now

@eliasgc49
@

It’s not the model 
 it’s the money in marketing that matters !

@billykg8
@

Problem is, now all the other LLMs release their next model that is better than ChatGPT. Is just a leapfrog race. There is no reason to switch. Just stick with your current tool and wait for their follow up release.

@sp00l
@

At this point you are just running out of things to talk about

@Floyergilmour
@

@theo you are a fear mongerer, I curse you

@anonymous_dev9472
@

cool cool its going to be surveiling you until the end of time

@itsRetroRocket
@

Used to like your video, especially in your expertise. Now these AI video titles are getting really really tiring. Maybe you should ask GPT 5.4 pro for a solution

@wei2much
@

thanks for letting me know claude schools codex on front-end. my project looked like garbage and i didnt know i could do better lolol

@bernard0camp0s
@

Gemini 5 flash will be so cool. I hope it's cheap by then.

@MegaLokopo
@

Love your videos, Theo, but you try a bit too hard to prove you're unbiased. If you just made your points without constantly pointing out your neutrality, we wouldn't even suspect a bias in the first place. Just trust your takes!

@kennethben-boulo7127
@

Nah, Chat GPT is trash now. Back then I use to get away with the most offensive jokes with 4o through custom instructions. Now they too soft and sensitive as I always get marked for explicit or inappropriate.

@captainasia1205
@

Its gotten good for sure. Just popped it up to do me a WP site in 2hrs while sipping juice in my PJs... Time to sell you guys a course and get Theo to review it next to devin😅

@BenOmondi-c9o
@

OMG Dude, just stop with the fucking commercials.

@IIIIIIIIIIIIlllIIIII