What sort of maths are LLMs good at?

(gowers.wordpress.com)

84 points | by ColinWright 1 hour ago

6 comments

  • scronkfinkle 32 minutes ago
    > A good sign that LLMs have reached human level for a much wider class of problems will be if they start proving theorems using methods that, like much of the very best human mathematics, are new and surprising but that with hindsight come to seem beautiful and natural. They should also be methods that are difficult to stumble on by accident. It is hard to say precisely what would count as such a proof, but I think we’ll recognise it when we see it

    Agreed. I find that after seeing these results from OpenAI we undeniably have a machine that has:

    * General knowledge of nearly every subject humanity has ever learned

    * The ability to simulate reasoning (albeit sometimes not very well) with that knowledge

    * The ability to reference across the domains of knowledge

    To me, this is more or less what I would think "Artificial General Intelligence" is. It's the cumulative knowledge of all general human intelligence, baked into an artificial form, which can then use that knowledge to achieve novel goals.

    In many cases of mathematical breakthroughs there is an insight that comes from just happening to know a combination of already existing ideas and then combining them to solve that problem. This is where having that general knowledge seems particularly strong because we can run these machines for weeks on end effectively trying to brute force.

    That being said, I could never imagine an LLM in its current form inventing something as elegant as the Fourier transform.

  • h_mirin 42 minutes ago
    This is really an argument about test-time scaling, even though the post never uses the term.

    These days "test-time scaling" mostly means letting the model talk to itself for longer, but the first genuinely surprising results came from plain sampling. Google's AlphaCode generated millions of candidate programs and filtered them down to a handful of submissions, which beat the average human programmer in 2022, before ChatGPT even showed up.

    Sampling is what AI is good at. Making examples and doing LeetCode are similar in that verification is clear and cheap. Compared to that, "proof" is still a vague concept, except where Lean works. See the fuss over the ABC conjecture. So humans are still needed.

    The interesting question to me is what happens after enough learning from "sampling." Isn't AlphaGo's move 37 an AI's nose? If that happens in mathematics, we may end up with results that are correct, machine checkable, and not explainable in any way we find satisfying.

    • laszlojamf 18 minutes ago
      for somebody who's out of the loop: what's the fuss over the ABC conjecture?
  • n4r9 1 hour ago
    A thoughtful and measured post, as usual from Gowers. The final note is neat and worth pasting out here in full:

    > A good sign that LLMs have reached human level for a much wider class of problems will be if they start proving theorems using methods that, like much of the very best human mathematics, are new and surprising but that with hindsight come to seem beautiful and natural. They should also be methods that are difficult to stumble on by accident. It is hard to say precisely what would count as such a proof, but I think we’ll recognise it when we see it.

    • tcp_handshaker 46 minutes ago
      >>A good sign that LLMs have reached human level for a much wider class >> of problems will be if they start proving theorems using methods that, like much of the very best human mathematics, are new and surprising but that with hindsight come to seem beautiful and natural.

      I must be taking crazy pills and the AGI surely will pass me by... But TODAY, middle August 2026...And in the context of testing and evaluating the capabilities of current SOTA models to implement an Agentic application for job search, here is some simple inhouse built evals I run today, since I don´t trust LLM vendors published benchmarks...

      Models tested: GPT-5.6 Sol in Extra High mode and Opus 4.8 Max.

      TASK REQUEST: Clear, not too long not too short prompt, for LLMs to go out and research freelance consulting gigs for one specific IT domain, and in one specific country in Europe, including maybe opportunities driven from temp agencies based in geographically close countries.

      RESULT: Models go out, fetch the data, and completely misunderstand the task...offering on first results, permanent roles instead of freelance, and based on the country where the agencies are, not in the one it was request for. Think for example IT jobs in Ireland, while freelance agency in London.

      ANALYSIS: No intelligence I can call it shown by models, adding cognitive effort for human in the loop to detect subtle factors, and therefore totally useless for agentic app...Best practices would be I guess to add agents on top of agents but although in the p95 of cases that will reduce the errors...for the remaining 5% that could have hallucinations or logic hallucinations like these ones, compounding on top of other logic hallucinations.

      I dont care about the theorems being proven. At the end we will found out what most mathematicians were doing, was just exploring the same combinatorial and abstraction patterns. And because of that I am sure LLMs will make mince meat of a lot of mathematical domains.

      But right now, what we call intelligence is not existing where it matters, and Ed Zitron is right its a parlour trick.

      • CodeCompost 6 minutes ago
        LLMs have no sense of geography. They measure distances between parts of words, not distances between parts of world.
      • criley2 10 minutes ago
        In my opinion, your usage of AI appears extremely sophomoric, and your results seem to follow your own skill level.

        There is this persistent belief that AI is a great leveler and you just type "pls get me a job kthx" and it should perform literal miracles.

        Try some context engineering. Try customizing your harness. Try having the harness improve itself. These are AI 101 lessons you can find in many tutorials. Anthropic has a great set of tutorials on how to use claude that go into a lot more depth for beginners.

        This reminds me of how when Juniors use AI, they produce offensive slop, but when Principals use AI, they produce some truly beautiful systems.

        AI is not a great leveler, it's a skill based tool. If your results suck, before you blame the tool, consider other possibilities.

        But of course, bias confirmation that AI is just a big scam sounds a lot easier than admitting a skill deficit and spending real time and effort learning.

      • m348e912 39 minutes ago
        >TASK REQUEST: Clear, not too long not too short prompt, for LLMs to go out and research freelance consulting gigs for one specific IT domain, and in one specific country in Europe, including maybe opportunities driven from temp agencies based in geographically close countries.

        As a human, not an LLM, I could interpret "including maybe opportunities driven from temp agencies based in geographically close countries" as meaning "including opportunities in nearby countries outside of Ireland" (that happen to be driven by temp agencies).

        Before writing off LLM as simply a "stochastic parrot" or a "parlour trick" remember it can't read your mind, not yet anyway.

        • tcp_handshaker 38 minutes ago
          I am describing the contents of the prompt that was not the prompt. The prompt was very clear to the model that some freelance opportunities in country A, the only one in consideration could be available via agencies in country B and C. And it was a clear prompt.

          So what happen is a prompt said for example, find freelance opportunities in Ireland but keep in mind some of these might be available via temp agencies in London.

          If you offer me not freelance but permanent roles, and not in Ireland in London...that is a logic failure.

          Its this type of complexity with the normal world, that these SOTA constructions so badly fail at, and so spectacularly fail at the margins... despite maxing all benchmarks...Parlour trick.

          • xnorswap 22 minutes ago
            I found it very difficult to parse your description, "Clear, not too long not too short prompt, for LLMs to go out and research freelance consulting gigs for one specific IT domain, and in one specific country in Europe, including maybe opportunities driven from temp agencies based in geographically close countries."

            ( I was trying to quote a single sentence and then realised it ran on for the whole paragraph. )

            Given how difficult I found that to follow, are you sure your prompt is actually "Clear, not too long not too short"? We now only have your word for it. I too had assumed that was a prompt given to an LLM to further prompt agents.

            • geon 10 minutes ago
              That was not the prompt.
            • tcp_handshaker 9 minutes ago
              You can easily test this yourself with the SOTA models....or read the corroborating literature...

              "General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks" https://arxiv.org/abs/2604.11778

              "...General365, a benchmark specifically designed to assess general reasoning in LLMs. By restricting background knowledge to a K-12 level, General365 explicitly decouples reasoning from specialized expertise. The benchmark comprises 365 seed problems and 1,095 variant problems across eight categories, ensuring both high difficulty and diversity. Evaluations across 26 leading LLMs reveal that even the top-performing model achieves only 62.8% accuracy, in stark contrast to the near-perfect performances of LLMs in math and physics benchmarks..."

          • geon 11 minutes ago
            Would the llm work better if it was given the job ad and asked where the job was located?

            It seems to me that such simplified tasks tend work better. The rest of the loop is just scraping websites, which doesn’t really have a reason to rely on ai agents.

          • someguyiguess 23 minutes ago
            Based on the rest of your writing I’m going to assume that the prompt was the problem.
            • coldtea 20 minutes ago
              He was perfectly clear in both cases.

              If a human misunderstood this, they'd be a dumb human.

            • tcp_handshaker 10 minutes ago
              Keep deluding yourself, unless you work for an LLM provider...

              "Frontier LLMs Still Struggle with Simple Reasoning Tasks" https://arxiv.org/abs/2507.07313

              "General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks" https://arxiv.org/abs/2604.11778

              "...General365, a benchmark specifically designed to assess general reasoning in LLMs. By restricting background knowledge to a K-12 level, General365 explicitly decouples reasoning from specialized expertise. The benchmark comprises 365 seed problems and 1,095 variant problems across eight categories, ensuring both high difficulty and diversity. Evaluations across 26 leading LLMs reveal that even the top-performing model achieves only 62.8% accuracy, in stark contrast to the near-perfect performances of LLMs in math and physics benchmarks..."

  • dataviz1000 31 minutes ago
    If you want to peek inside how a model solves a math problem have a look at some data visualizations I made solving basic multiplication.[0]

    I wanted to demonstrate capacity (how well it does a thing) instead of capability (which things it does, like drawing a pelican on a bicycle with SVG or solving a Rubik's Cube). To understand how LLMs solve math, look at the simplest case of multiplication. I deconstructed and classified the thinking token output. It is very important that model training yields thinking token output that structurally follows an observe, orient, decide, act (do the multiplication), and observe again loop.

    [0] https://adamsohn.com/reasoning-grid/

  • pinkmoonx 57 minutes ago
    How interesting is it that in the same way the human brain unconsciously does calculus and linear algebra, but struggles in the conscious space (we have to go learn it, it’s not easy) the same is true of LLMs.

    They are algebra, and yet kinda suck at it without training

  • krupkinmaxim 10 minutes ago
    [flagged]