We Tested $200 ChatGPT Pro and Claude on Research Math
Video Overview & Insights
In today's video we'll be exploring the $200 subscriptions of ChatGPT Pro and Anthropic's 20x max plan against research maths problems and research paper visualisations to see which model can make the best program to visualise 2D Lorentzian triangulations!
Thank you for this video you can make
A huge thanks to the channel member Belzedar who helped me get access to the $200 Claude model!
Thanks for all of the support! đđ
maybe it would be nice to see a colab between kyle basares and your channel.
đđ€ž Connect on Instagram đčđ”
Instagram.comn/easy.riderss :)
You donât really get to complain that they aren't subsidising you more than they already are. I like what you're doing, but the entitlement is quite something.
More User Perspectives
You should be running Claude in Claude Code - it can't run out of thinking there
@theepicosityofpizzaFor Claude, don't use the chat interface. Use the Code interface, it has a lot more context, and does not keep stopping like the chat version. You have to use it on the code tab part of the Claude Desktop
@ipechmanyou need to try codex and claude code/ the experiences are vastly different than web
@simonlevy00It's time to test GLM 5.2
@juancardona8213I am working on some economic theory that is simple, but not standard. Fable made many mistakes. I intentionally paid for the Pro plan, and now I have this feeling that I cannot fully trust it.
Although I admit that Fable provided some criticism, and some of it was good, in other cases I was able to prove that it was not correct several times.
Did you tried using claude cowork? Because the context is higher. Maby in this way Claude Fable 5 may give the answere without hitting the writing limits.
@MrBoboliHey, I know this is so unrelated, but Iâve been trying to get my hands on a freebord. Iâm from Puerto Rico and sadly thatâs just non existent here. I learned that u were a pro and I was wondering if I could ever buy a used one from you.
@SpAz8.Just a random question: do you still play piano and if so how are finding it?
@itspulsarI'd say this test was pretty bad. Given how LLMs work and the variability answes can have a better way is to do 3 generations and compare those on average.
@gabriel3782-j5jClaude models tend to overthink and get trapped in loops when on max thinking. I would suggest a lower effort mode would solve the issues you faced with Claude's reasoning budget.
@WADstephenTry GLM 5.2, I heard it's very very competitive to the frontier models I'm pretty sure it's #1 on AIM math and completely free to use with no subscriptions also review
Kimi newest models
Qwen newest 3.7 plus and max models
Deepseek V4 expert
Try GLM 5.2
@nigelwatson3049try composer
@dudhshshheshhshsha87054:51 there's no way im that childish...
@Newo-mx5nfpls bench openrouter fusion (deepseek v4 pro + gemini 3 flash + kimi k2.6) against gpt 5.5.
@yhjhpf-l7xThe MOST useful bench for me, thanks a lot! (I am jr data science level product manager trying to build BigTech grade DS product)
@artmzrbnI have to be one of the few people on this planet doing the kind of research that I do (A model to replace the standard model in Cosmology and Particle Physics). The truth is, I can't do it accurately with just either of them. I have to use both of them. Both rely heavily on my abstract reasoning, but to get the math right, I actually need to use both. One to stress test the other. They'll give their corrections and recommendations, and I can then take that back to the other one, and then they will build a new version plate of whatever I'm working on within the model. I first have to derive the ontology plate. I usually do that with ChatGPT, and then Claude stress-tests it.
Then I have to do the observations and predictions, which usually require several plates that are stress-tested. Then there's the math plate, which takes the longest and requires the most plates to be stress-tested. Then I can move to the flagship paper, which usually requires fewer stress tests because of all the prior stress testing. Those can come back clean or need an additional plate or two. But for developing a new model in physics like mine (Logical Mechanics is Cascading Asymmetry and Degenerate Symmetry), you can't do it accurately with just either of them. They've worked very well together during stress testing by accepting each other's corrections and recommendations. Running out of tokens before I can get an answer to my question happens to me a lot, which is why I have to start every topic I'm working on with abstract reasoning using ChatGPT. But neither one of them can do it alone when it comes to the math. I'm an INTP, but I choose the ENTP method: build two models, stress-test them, and where they meet is the truth. I don't know if that's common?
You should run sub-agents for complex tasks like this. Have 5-6 agents for research, some for developing different solution variants and potential tracks, some for lemmas, some for proving them, and some for review. That way, the main orchestrator agent will be able to maintain the entire context effectively. Right now, from an AI engineering perspective, your approach is poor. You're just talking into a black box and saying, 'make it look nice.'
@LumaTreePokerWhy donât you use Claude code instead of the chat interface? You can even put it in /goal mode to solve your question. So much better.
@GingaByte54319Have you tried not using Max effort? I've never seen these errors before but I never use Max reasoning.
@ster2600Still super interesting and cool to get your perspective and experience :)
I personally still prefer Claude: its personality, explanations, and how it writes code for my algorithms. But the math I do is way less theoretical than yours, which makes a big difference (less likely to get stuck in an almost infinite loop during the thinking phase, especially on the max effort setting).
Opus 4.8 is capable of solid math (56% on FrontierMath Tier 4; 80% on FrontierMath Tiers 1-3), but what you're experiencing is really embarrassing for Anthropic... If I were you, I'd try lowering the thinking effort level, even if that sucks on paper :(
Could you try an AI specifically designed for math, such as Aristotle from Harmonic or AXLE from Axiom?
@legoroboticssimulations7163This is a foundational problem with LLM based AI. The biggest leaps in capacity will come from AI built on different foundational reasoning.
@Tom-O-RowI don't know if you've noticed this but any online content on these LLMs will lead the comments getting spammed and liked by bots of the company trying to shill their product.
@fadasdWait, so these $200⊠is that what they are going to pay ME for actually using these stochastic bullshit-generators? Because if I am the one spending my valuable time fact-checking their endless hallucinations and fixing their broken code, $200 a month sounds like a very low entry-level salary for an "AI-babysitter".
@calysa_projectReally embarrassing video. Just not for Anthropic! đ
Also: Fable never had at any point visible thinking. đ€
Maybe riding isn't that easy after all đ€Ł
The workflow is not great, instead of using the website chatbot and repeately asking it to "continue":
(1) use a tool such as Claude Code or Codex,
(2) create a project repository, and
(3) ask it to write everything within a file (such as a markdown file or txt)
Then read the file after everything its done.
The workflow in the chatbot website is very simplified.
A typical case of PEBKAC error
@TheShadowHoodie@Easy Riders (For non-coding tasks) Have you tried asking for an output in LaTeX rendered as a pdf? I used to have the same problem with both GPT and claude (sub-standard outputs), and then I had a silly idea to ask it to âgive me your output as a PDF, and also give me the LaTeX code that produced that PDFâ⊠I found a night-and-day difference in the outputs before and after I started asking for it, especially with diagrams and visuals. Having said that, Iâm not a maths researcher (yet). Itâs especially useful when trying to make summary sheets or explainer documents for maths/ computing concepts. It may not work for you (given the differences in how we use AI), but it might get around the output token limit because you can ask it to continue and itâll have whatever its already written in LaTeX as prior context.
@arjunkapoor5653Dont use claude. Use opencode. Run in plan mode first then same a memory file as you go. Trim the memory file.
@vertical-golfMaybe a fine tuned math model based off of an already strong model could potentially perform much better even than a generalized AI model. Maybe Minimax? You would definitely have to rent out some cloud compute but it would be interesting to see. Furthermore, you would be able to prevent thinking loops by lowering the temperature, top-p, repetition penalty, etc.
@anonymous-hz9dqI'm not familiar with the web apps. However.
You should consider trying it with claude code.
These models have limited context sizes and expecting them to answer complex questions without giving them access to a filesystem, a wiki and programming tools and artifacts is like expecting your Uber driver to wear a blindfold.
Technically it's possible he'd get you to your destination but... it's a bad way to go.
Also, I don't think LLMs can reason at all. It's a category error, perpetrated by a marketing term.
'Reasoning' is a marketing term.
To be productive with these models, you need a coding harness and a human working together with the model.
The human brings experience, REASONING and human context and attention.
The model brings it's own tool-like context and attention and embedded knowledge and execution speed.
LLMs should not be sold as something you can plug a question into and get a dependable answer out.
They are tools, that should be part of a harness that a human wears.
The human needs to learn the domains as they wear the harness and apply the tools.
LLMs are not an easy button.
Finally if anyone reads this far.
Pi the agent harness is far better than claude code the agent harness.
It's simple. It's efficient. It's transparent. It's extendable. It's free.
Honestly at some point everybody is able to use these tools and get great usage out of it so probably the tool is not the problem.
These are new tools and everybody's discovering the better way to use it. I don't understand why you think this should work without you making any effort at all to understand how they work and how to use them to their max potential.
Obviously if you're not getting any benefit from it you should try to change your approach rather than expect for the tool to work exactly how you wish it worked.
You should use claude code, try to break the problem into smaller tasks, probably use plan mode.
You should also extract text from the PDFs and use that rathter of clustering your context window.
It's like you just bought a racing car and you're complaining that it doesn't have an automatic gearbox so you're struggling to reach full speed.
You know there's a thing called arxiv lean
@narwin2477You should have tried Fable on vs code. Blow minding. I hope you will be able to test it when they will make it available again
@TotiiusUse Codex...pursue a task function
@zayzay9006Dude Claude is AWFUL at math, and if youâre working on difficult problems it gives you outright refusals to proceed. Itâs terrible, and uses safety and psychoanalysis to waste and up the charge on tokens it charges you. Itâs a manipulation business model.
@spaghettisquasherI use both somewhat heavily, Used to be more GPT, now use far more Claude. The issue is that with GPT, their primary focus is seemingly getting the user interface streamlined. Nothing in their interface tells you where you are at in your context, or your 5 hour window limit, or your weekly limit. You will just be quietly downgraded to a less powerful model if you go over your allotted usage. Although I do believe they are better about telling you when that happens. Claude on the other hand, you can look at how much of your 5 hour window you've used, and your weekly limit. You know exactly when both will roll over, and if using Claude Code, you can even see where you are in your chat context window as well. In my experience, when using ChatGPT, you have to learn how to prompt. When using Claude, you have to learn how to use Anthropic's UI. As others have said, using Claude Code would be the way to go. But to your credit, if you are paying $200 a month, you shouldn't have to know which system to use in a very convoluted UI. They should all work very similar to the user.
@thesharkn8doI didnt run fable on max i ran it on high and extra high and it would find out quite interesting solutions to problems that other ai was unable to do. Its a shame they removed fable access it was really good
@cybervandTry with a harness like Claude Code or Codex. You can have it continually update a LaTeX document in a loop, making incremental progress. This makes a huge difference for writing code, but you may find that it is also better for maths research. If you hit usage limits, you can just have it continue from the current state of the document. A harness will also usually provide more tools that it can use. For example, it could write a small script to verify some steps that it has taken. You can also use version control (like Git) to fork the document at any stage. Or you can kick off a bunch of agents to try different approaches.
@SegFaulted0x0yeah one shooting doesnt work for this, altough it could have easily have helped you or at least produce an answer if you let it made a plan,, and adviced it to go trough step by step
@LalaA-s9fI've never been able to get Claude to provide an answer (always run out of tokens). You can ask ChatGPT to give the answer in LaTeX format.
@kmmertesFable code and fable chat ui use different models although they are both named fable
@kingstricker43Thanks for sharing. I wish I had the mathematical background to explore ideas like this more deeply.
@yogeshjogAn old man doing the exact same video would get NO views. Just admit why you do this and why you are here. It's just an example of the privilege of youth and beauty. (Yes, this is an old man complaining - but sooner or later you too will be old. But will you be bitter?)
@pascalbercker7487You probably should use the Claude Code instead of the chat, either the cli or the desktop app
@mahmoudalmontasser8369Try using GPAI I heard it's the best at math, science and physics I heard scores of 99.2 on QPADim that's just what I heard I'm not sure
@Dr.TechTeach