Yes, but most of the paid “LLMs” are not actually just LLMs, they have specific aritimetic modules that end up being activated for most math questions that actual LLMs suck at, however are pretty easy for arithmetic modules, like simple calculations, unit conversions, etc…
So yeah, any edge “LLM” tool currently will always get the original question asked here CORRECT even on the free tiers, you can test it yourself on Google Search itself in “AI mode”, on ChatGPT, or DeepSeek…
So you are technically correctly in saying that LLMs suck and cannot be trusted for math in almost any case. However you are wrong on implying that ChatGPT, or other commercial models are just LLMs, they are not.
P.S.: I’m not condoning the use of AI to resolve math related questiona, nor saying that their replies should be trusted, just stating the fact that actual “AI” products use more than just a large language model to generate their responses, and for math related questions most “AI” today will get arithmetic question answered correctly NOT DEPENDING on having training for the exact questiona being asked since they can identify a math arithmetic question being asked and instead of trying to fully generate the answer with the model itself, which would lack the training data to be able to process such wide range of possible question, they generate placeholders internally and let a arithmetic like module do the calculations, which take negligible processing power to be cacoulated by CPUs.
Can confirm; I teach statistics and allow take home exams. I don’t even really need guardrails on the math-- if you cheat via LLM, your answer is almost always hilariously wrong.
For simple math that could work, and as long as the question is close enough to an exact match with plenty of published examples to copy.
A good rule of thumb is that the script it will come up with is about as likely to be correct as blindly taking the highest voted answer to the most similar question on Stack Overflow.
If the question is simple and common, the odds are quite good. If the question is nuanced or rare, the odds of a correct result drop off aggressively.
cannot confirm. have used chatgpt many times to check my physics homework (after doing it myself first) and it was basically correct about 90% of the time.
I can’t say that particularly shocks me. I imagine that Chat GPT has probably swallowed dozens of Physics textbooks.
The bigger part of the problem is knowing where we are on the human knowledge novelty / training data available curve, or rather, when we need to get off of the ride.
Edit: Any chance your chosen AI has access to Wolfram Alpha as an MCP Tool? Because that would be very different, too.
If it’s successful at math, there’s likely much more than a learning model in play.
Which brings us back to the opacity problem. :(
If the answer machine has been cleverly rigged (and weighted) to pass off math questions to more capable software, things can be great.
But a person who just hears that AI can do math now can have a very bad time if they go blindly use a brute force pleasant answer machine on a math problem. :(
Yesterday I literally had to count 14 thousand pounds of fire extinguishers (as in count each extinguisher in a bunch of crates), one of my coworkers came over and said “Why don’t you just use AI to count them.” They then pulled out their phone, took a picture, asked AI to count them, and was off buy almost 60% on the first crate. Then went “oh, well that’s just one example,”
They then tried 4 more times, being between 75% and 50% off each time. They then went “Oh, well, ummm, nevermind.” and wandered off again.
(I had to count the extinguishers to make a Bill of Lading for them for shipping them out, you need an accurate count otherwise it will be held by the shipping company until you provide a correct BoL with the correct count due to legal reasons. I couldn’t just weigh them and divide the total wieght by the weight of a single extinguisher as they were a bunch of different sizes mixed together. NGL it was the easiest money I’ve ever made at the at job.)
Oh it sucked for sure, I kept getting interrupted which meant whatever crate I was counting at the time had to be started over.
But the biggest perk was not having to deal with the other idiots on our loading dock for the entire time I was counting.
Though I’ve joked about it in the past I really should get a shirt that says “I’m trying to focus on my job, please don’t distract me” or something like that.
I know this is a meme, but for the record math is one of the things LLMs are famously worst at
were before they started just writing a mini-python script for every calculation and executing that.
Yes, but most of the paid “LLMs” are not actually just LLMs, they have specific aritimetic modules that end up being activated for most math questions that actual LLMs suck at, however are pretty easy for arithmetic modules, like simple calculations, unit conversions, etc…
So yeah, any edge “LLM” tool currently will always get the original question asked here CORRECT even on the free tiers, you can test it yourself on Google Search itself in “AI mode”, on ChatGPT, or DeepSeek…
So you are technically correctly in saying that LLMs suck and cannot be trusted for math in almost any case. However you are wrong on implying that ChatGPT, or other commercial models are just LLMs, they are not.
P.S.: I’m not condoning the use of AI to resolve math related questiona, nor saying that their replies should be trusted, just stating the fact that actual “AI” products use more than just a large language model to generate their responses, and for math related questions most “AI” today will get arithmetic question answered correctly NOT DEPENDING on having training for the exact questiona being asked since they can identify a math arithmetic question being asked and instead of trying to fully generate the answer with the model itself, which would lack the training data to be able to process such wide range of possible question, they generate placeholders internally and let a arithmetic like module do the calculations, which take negligible processing power to be cacoulated by CPUs.
Can confirm; I teach statistics and allow take home exams. I don’t even really need guardrails on the math-- if you cheat via LLM, your answer is almost always hilariously wrong.
I’ve not used gpt for quite a while, but can’t you ask to python script the evaluation?
For simple math that could work, and as long as the question is close enough to an exact match with plenty of published examples to copy.
A good rule of thumb is that the script it will come up with is about as likely to be correct as blindly taking the highest voted answer to the most similar question on Stack Overflow.
If the question is simple and common, the odds are quite good. If the question is nuanced or rare, the odds of a correct result drop off aggressively.
cannot confirm. have used chatgpt many times to check my physics homework (after doing it myself first) and it was basically correct about 90% of the time.
these were not standard exercises (i think)
Thanks. I think every data point helps people.
I can’t say that particularly shocks me. I imagine that Chat GPT has probably swallowed dozens of Physics textbooks.
The bigger part of the problem is knowing where we are on the human knowledge novelty / training data available curve, or rather, when we need to get off of the ride.
Edit: Any chance your chosen AI has access to Wolfram Alpha as an MCP Tool? Because that would be very different, too.
If it’s successful at math, there’s likely much more than a learning model in play.
Which brings us back to the opacity problem. :(
If the answer machine has been cleverly rigged (and weighted) to pass off math questions to more capable software, things can be great.
But a person who just hears that AI can do math now can have a very bad time if they go blindly use a brute force pleasant answer machine on a math problem. :(
Yesterday I literally had to count 14 thousand pounds of fire extinguishers (as in count each extinguisher in a bunch of crates), one of my coworkers came over and said “Why don’t you just use AI to count them.” They then pulled out their phone, took a picture, asked AI to count them, and was off buy almost 60% on the first crate. Then went “oh, well that’s just one example,”
They then tried 4 more times, being between 75% and 50% off each time. They then went “Oh, well, ummm, nevermind.” and wandered off again.
(I had to count the extinguishers to make a Bill of Lading for them for shipping them out, you need an accurate count otherwise it will be held by the shipping company until you provide a correct BoL with the correct count due to legal reasons. I couldn’t just weigh them and divide the total wieght by the weight of a single extinguisher as they were a bunch of different sizes mixed together. NGL it was the easiest money I’ve ever made at the at job.)
$10 says they still think AI is great at counting things.
Easy money for you. For me that work would have been torture.
Oh it sucked for sure, I kept getting interrupted which meant whatever crate I was counting at the time had to be started over.
But the biggest perk was not having to deal with the other idiots on our loading dock for the entire time I was counting.
Though I’ve joked about it in the past I really should get a shirt that says “I’m trying to focus on my job, please don’t distract me” or something like that.