The research community is in uproar after OpenAI released a trove of more than 700 mathematical preprints entirely generated by AI on 6 October. The San Francisco, California-based maker of ChatGPT posted the preprints on the software repository Github.

Although some mathematicians celebrated the solution of longstanding problems, others took to social media to complain about being scooped. Some were incensed at what one physicist called a ‘slopocalypse’, even if the mathematical content could end up being formally correct.

  • zener_diode@feddit.org
    link
    fedilink
    English
    arrow-up
    9
    ·
    12 hours ago

    I think there are three main issues (though I might be overlooking something):

    1. AI taking over peoples work is as much an issue for mathematicians here, as it is for programers, artists, writers, etc. elsewhere.
    2. We still need to check if these 700 papers are correct (AI makes mistakes at least as often as humans, usually more often). This involves a huge amount of effort. And if it turns out that they contain a bunch of mistakes, OpenAI is unlikely to care. To them, these papers have already served their purpose (generating headlines).
    3. Solving a problem in mathematics means being able to (formally) convince other people that your opinion is correct (massively simplified, but imo that is what mathematics boils down to), by getting them to understand why your opinion is correct. Having a pile of linear algebra spit out 700 texts about certain math problems removes the whole “human understanding” part of that.

    However, it’s not entirely unprecedented for a computer to provide a proof that we don’t completely understand. I forgot the details (and I’ll edit them in if I can find it), but there was at least one problem that was solved using a computer and brute force trying millions of cases. That produces a proof much longer than any human could read in a lifetime, and iirc there where quite a few mathematicians unhappy about it at the time too.

    • Eq0@literature.cafe
      link
      fedilink
      English
      arrow-up
      9
      ·
      12 hours ago

      I can add about your last paragraph(it’s literally the core of what I do). There is a field of computer assisted proofs, that is mathematical proofs that need a computer to be completed.

      The first and most well known example is the four color theorem. How many colors do you need to color “a map”. Answer: 4. Proof: very very long. By hand it is possible to prove that there are only 1834 options, and a computer was used to color all these explicit options. At the time, it was a scandal. Nowadays, other computer assisted proofs are accepted, such as the ones relying on validated computing: if a computer (with some restrictions and guardrails) can show that a certain value is over/under a given threshold then something else is true.

      Then, there is the validated proofs approach. This is where Lean comes into play, if you have heard about it. You can ask a computer to check your proof. This is helpful for confusing, long proofs (most of math). You input all the logical steps you took and Lean confirms that all is logically sound. Many AI proofs are “Lean verified”, but there is controversy if they are proving what they claim they are proving. It’s also a massive chunk of code that nobody can understand.

      • Grimy@lemmy.world
        link
        fedilink
        English
        arrow-up
        3
        ·
        11 hours ago

        Have you looked at some of the problem solved? I’m way out of my depth here, so I’m having trouble understanding how much 700 is. If it is proven and understandable, how much of an advancement would this represent? Sorry if it’s a random question, but you seem more knowledgeable than most.

        • Eq0@literature.cafe
          link
          fedilink
          English
          arrow-up
          6
          ·
          10 hours ago

          Just to give a comparison, a high-output mathematician publishes between 2 and 5 papers a year (depending on branch of math and dividing by co-authors).

          Most are incremental work, so a little step towards solving a problem or a conjecture. A lot is just a “hey, look at this neat trick”. Solving “big problems” is usually the work of a decade or more, in which the mathematician is working on other stuff as well. So let’s roughly say that solving a big problem usually takes some 10 years of work and some 10-30 papers (assuming working roughly half time on it).

          So on one hand, 1 paper per big problem is too little to actually understand what’s going on, on the other 700 papers are the output of more than a hundred mathematicians over a year. A math department in a university is around 50 mathematicians, I would say? So two medium-sized departments.

          The other part of your comment. “If its proven and understandable”

          At the moment, the preprints are assumed to be a shit sandwich. (Aka “would you eat a sandwich is there could be shit in it?” ). Using the results is just too risky without understanding if they are correct and how they are build.

          The “understandable” bit is also really hard. I picked a random one in my field (nothing I directly worked on) and it was unreadable - mostly because of notation used without defining it and no explanation of what is going on, no overview or intuition. So to me it seems more shit than sandwich. As a reviewer, I would never accept such a paper.

          Then finally, let is assume it’s all perfect and good. The goal of proving something in math is to develop understanding and a method to apply to other cases. So once all these papers are studied and understood, poop discarded, rest of the sandwich saved, ideally we will have new understandings of whole sections of math, new connection between items we’re weren’t aware of. What I think AI did in this context is to chain things that were already known, but there was no one whose knowledge spanned wide enough to know that all the pieces were already laid out. So it could be groundbreaking - once we remove the shit. How much shit there is is anyone’s guess.

          • Grimy@lemmy.world
            link
            fedilink
            English
            arrow-up
            3
            ·
            edit-2
            9 hours ago

            Thanks! That really puts things into perspective. It’s less impressive than I first thought, especially since it seems like spaghetti that needs a lot of untangling. 700 struck me as a huge number at first.