I was a graph theory junkie long ago and even moved to Budapest for awhile to study among the greats. While I was there I started working on Barnette's Conjecture which came to occupy my thoughts over the next 24 years of my life, on and off as I worked in many different fields. Last summer I even thought for a few days that I had actually solved it.
But it's supposedly proven here - problem 180. I don't know what to think exactly. I spent thousands of hours on that problem. I really enjoyed it. Hearing that it is solved somehow makes me sad in a far-off way, like hearing an ex-girlfriend died suddenly in a car crash. I don't know, there's probably a lot of people feeling odd emotions tonight.
There's no Lean proof for this one so I'm digesting the paper. On the surface it looks like an approach I considered 24 years ago and abandoned.
I revisited the problem this summer, along with my partial solutions, when the previous round of stunning proofs came out. Several hours of work with Fable simply convinced me it wasn't yet solvable and reinforced how hard of a problem it was.
> Several hours of work with Fable simply convinced me it wasn't yet solvable and reinforced how hard of a problem it was.
This is the part that gives me the strangest feeling about it all, because you're not the only one with this experience. I've experienced this too on different problems, as have many researchers across many fields.
I disagree with the Fields Medalists on the majority of their complaints. AI math is happening and there's no going back. However, on one point I increasingly agree: virtually none of this stuff is possible with technology any normal citizen has access to. I have no problem with AI models making revolutionary advances in math or science. Where I start to have a problem is when the AI models making these advances are tightly withheld, proprietary, and seemingly never released with these capabilities intact. This has been the case for all of 2026 so far.
I suspect that this is in fact the source of much of the angst. None of this progress is reproducible outside of one or two teams inside OpenAI and Anthropic. It's becoming an incredible concentration of power that I don't know that we've ever quite seen before. Right now, it feels harmless because it's being used for wonky math problems that aren't (yet) practical for anything. But great power never stays harmless. History has taught us that countless times, in countless different forms.
> I suspect that this is in fact the source of much of the angst.
Why do you "suspect" this as if it's some hidden motivation when the very first paragraph of the advisory group's statement (linked from the OpenAI post) says:
> At present, some frontier AI labs are testing advanced mathematical problems on proprietary models that remain inaccessible to the broader scientific community. Our recommendations are formulated with this practical context in mind. However, ideally, they would not do so. We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models.
Tao and others in that group have been strongly and publicly pro AI from the start. They are not advocating "going back". They're objecting to the strip mining of open problems using proprietary technology.
I don't think the strip mining metaphor is appropriate. Mining is a zero-sum game; if I mine something, nobody else can go and mine the same resources I did. Mathematical problems don't go away when AI finds a Lean proof. They create new opportunities for humans to study the solutions, learn new techniques from them, identify promising directions for future research, discover alternative/more beautiful proofs, and write expositions for other humans.
Strip mining is very apt if you view the economics of the present system as "effort -> recognition -> career advancement". Even in strip mining, the resources that had been buried are now available for use in the broader economy. What's no longer available is the living that was to be had digging them out.
The problem isn't effort, though. All of the things I mentioned constitute effort and could be rewarded. The job economy was created by mathematicians incentivizing the proof of difficult theorems above all else and valuing all other work at approximately zero as far as career advancement was concerned. Now they're pulling a 180 and claiming that math was never really about proving theorems, but that's contradicted by their revealed preferences. The strip-mining problem only exists if they continue with the status quo ante.
Mathematicians aren't homogeneous. There are mathematicians valuing pedagogy, collaboration, bridge-building, theory building, along with those that chase the 'difficult theorems', to name a few, and there are lots of flavors within each class, with lots of blending and blurring. You infer that mathematicians prefer the status quo simply because it is the status quo -- with a little thought, you'll recognize that this is a fairly silly notion.
There are myriad circumstances where the values of most practitioners differ from the status quo, which is nevertheless well-entrenched. This can arise from inertia, or from outside forces, such as broader cultural milieu, integration with larger institutions, or contending with economic realities. If you think that these do not and haven't historically played a role in determining the job economy and that math is a pure field where mathematicians could comfortably shape it according solely to their own ideals then you are naive
And, in addition, many mathematicians are graduate students or postdocs hoping to line up a permanent job soon.
For example, if you look at Terry Tao's blog, he has a tremendous amount of first-class expository writing. So, too (to some extent) do junior mathematicians -- but, unfortunately, this tends to not be highly valued by the job market. Grad students and postdocs have learned that to succeed they need to play by the existing rules of the game.
Well, the board has just been yanked from underneath them. People like me can afford the sort of idealism and soul-searching that the parent comment describes, but junior mathematicians face a very unenviable set of circumstances.
1) I never said the problem was effort; I was trying to explain the strip mining analogy, and it's one of the two anchors that make the analogy work.
2) Mathematicians didn't create this economy; it was foisted upon them by the same managerial mentality that brought us "publish or perish" and "the monthly sales quota".
3) I can't tell if you honestly don't get why the strip-mining analogy resonates, or...?
Here's another analogy: if we suddenly discovered personal teleportation, and marathon runners were complaining that it was ruining the sport, would you say "they're pulling a 180 and claiming that marathon running was never really about getting to a point 26 miles away as fast as possible, but that's contradicted by their revealed preferences"?
The strip mining analogy is better though, because it captures the sense of irreversible goal-loss when a problem goes from being "unsolved" to "solved".
As far as I understand, even with "publish or perish", peer reviewers decide what counts as an important enough paper to be published in a prestigous journal, and committees of peers decide whether or not, say, an expository article on arXiv or a textbook counts toward hiring or tenure. Again, as far as I understand, those things have largely not been rewarded in the past.
I like your marathon example, but maybe not for the reasons you intended. The community of marathoners decides the rules of a marathon. You don't need a hypothetical teleporter; you're already not allowed to use a bicycle, performance-enhancing drugs, or shoes that don't fit the specifications. The rules are updated to adapt to changing technology. Yes, I'm arguing that the strip-mining analogy doesn't make sense because mathematics is in the same situation. There's nothing stopping peer reviewers and hiring/tenure committees from changing the rules about which kinds of effort confer recognition and career advancement.
Strip mining is an extraordinarily appropriate metaphor.
Imagine a mine has an unknown number of rare materials. And you know the general location of a few of the most valuable spots. But you don't know what may be valuable right next to it. If the pieces that we know are valuable are suddenly gone, the incentive to mine that particular area drops considerably, dropping the chance to discover potentially brand new materials that would have been found the normal way.
That's an empirical claim. I could equally well say that doing an automated search of the problem space and having a database of results and open problems will identify vastly more interesting and valuable areas. Again, the idea that math is some exhaustible material is a metaphor, not an established fact. I'm willing to change my view as new evidence comes in, but I think we're going to have to wait and see what the landscape looks like in a few years.
> the incentive to mine that particular area drops considerably, dropping the chance to discover potentially brand new materials that would have been found the normal way.
FWIW, I think the metaphor breaks down with this framing. This isn't really a problem associated with strip mining, what's left behind is generally low or negative value (toxic). I'd suggest a different metaphor, from Wikipedia:
> This process involves the removal of all ground vegetation in the area, which is a detriment to the environment.[19] Topsoil may be placed over the tailing along with planting trees and other vegetation. Another reclamation method involves filling in the hole with water to create an artificial lake. Large tailing piles left behind may contain heavy metals which can leach out acids such as lead and copper and enter into water systems.
This feels very similar to the issues with algorithmic problem "mining". It has the potential to destroy the human ecosystems surrounding these problems, leaving barren wasteland behind where nothing can grow or flourish.
I hope sincerely hope they don't currently use "proprietary technology" like:
Wolfram Mathematica ($890/yr)
Magma ($2500/yr)
Maple ($680/yr)
COMSOL ($1500,yr)
Matlab ($500+/yr)
Seems like a very strange position to take, in my opinion.
Why does the field of mathematics suddenly now need to be "fair" and give everyone access to the same tools? Has that ever been the case in academics? It's always been a competition for name-recognition, grants, institutions, etc.
Macsyma / Maxima was an MIT developed CAS system back in the 60's that was proprietery until they sold it off to IBM for a tidy sum. Magma actually has free access if you're in the US, otherwise you pay. That's not to mention proprietary MATLAB toolboxes or specialized Stata modules.
Likewise, a lot of the above packages have pretty sweet site-wide deals with R1 universities. If you're at a smaller, foreign one, you're out of luck.
How is this different from literally any other part of the economy?
We've relinquished control over just about everything we use or consume. We can't compete with larger enterprises for production of food, clothing, machinery, medicine, energy, services. Mathematics is just the latest thing to be industrialized.
What keeps large companies under control is competition with other large companies. This competition causes the surplus value they produce to flow to consumers, not be hoarded via monopoly prices. Do we see strong moats that are going to cause monopoly in AI? I don't see it, and in particular I don't see it persisting if it exists transiently.
You're right, and that's a bad thing. AI is nothing fundamentally new, but its extremity is making many people aware of the truth that's been there all along. There's no contradiction in that.
> Do we see strong moats that are going to cause monopoly in AI?
Ownership of the capital assets used to train and inference new models. Yes, we may end up with more than one firm. But as we see with big tech today, a small number of fantastically wealthy firms in "competition" does not an open market make.
Is it a bad thing? We live in a society. We depend on the work of other people. We are not autonomous. Sure, we can try to be self-sufficient, and that would lead to a subsistence lifestyle much degraded compared to what we experience.
Somehow you have to argue either that society itself is bad, or that math is somehow different from all these other human activities.
I think the obvious fact that people prefer to live in places with large commercial organizations shows they don't really care about that, at least to the point of foregoing the benefits these organizations bring.
No, it doesn’t. You can be in favor of something and opposed to a particular way of handling or implementing the thing. And the issue here isn’t what it’s being used for but who is able to use it.
Eventually. In the meantime, here in the human socioeconomic sphere, you might be dealing primarily with the Second Rule of Fight Club and Operation Mayhem.
I am not sure how convinced I am by that argument. A gun is also a particular kind of tool, and it makes the person at the handle end sovereign, and the person at the pointy-shooty end subjugated.
Regarding the advisory group, OpenAI claims to “have drawn on their advice”, which would include not dumping a bunch of AI slop, with the footnote that if they do do that, at least fund the process of digesting it.
At the same time, there's a new note at the bottom of agmai.org stating how they've been in contact with OpenAI about this particular release, and they say that “we consider these discussions constructive, it is ultimately up to the mathematical community to assess the extent to which our recommendations were followed successfully”.
So, what's going on there; is this British English for “they didn't follow anything at all”? Because from my perspective, it looks like they doubled down on the Navier–Stokes approach of trying to maximize PR gain while being as lazy as possible about actually contributing anything back to science, releasing only slop that may or may not be correct and may or may not be straight up plagiarism, as has been the case earlier.
If I were on the AGMAI board, I'd feel terribly exploited when reading that press release, yet their response is modest.
Hairer, if you're reading this: is there any indication whatsoever that AGMAI was anything but a cheap way for OpenAI to science-wash their press release?
Now I know there are issues with the field and how just answering these questions may cause broader problems, but I feel like the posted results is far from slop. We can't just call any output slop, or it loses all meaning.
If it was slop, it'd not be causing the issues the group are concerned about - they're not saying "the problem is we're getting loads of incorrect proofs thrown about that are nonsense".
When you blanket a set of things with a pejorative, and it turns out that some of the members of that set are demonstrably and definitively NOT covered by that pejorative, and that all the pejorative means at bottom is "I don't like", all you've accomplished in the long run is to call into question any future legitimate use of that pejorative. It is tempting, especially when heated, to stretch an invective, but it will ironically only lead to the death of its utility over time.
So the fact that the Library of Babel (i.e. all possible books) contains occasional gems means that you can't object to using it on principle? That would seem to follow from your logic.
What about a filtered set "all well formed books"? Or "all well formed books that are plausible enough that they could convince a reasonable person, regardless of their accuracy"?
It's generally taken that a cup of sewage in a barrel of wine makes a barrel of sewage. Surely a reasonable person could claim that a barrel of sewage was still sewage, even if it contained several cups of wine?
They're not calling any output slop, they're calling indecipherable output slop. The management class responsible for hiring, firing, and paying people doesn't possess the domain knowledge to say for certain whether or not LLM output is optimal (which, in this context, means correct), but they will trust that it's good enough to justify further automation / fewer grant approvals / etc. So in that sense, slop can and will cause the economic issues people are concerned about.
University boards want the prestige of successful research programs. Doing the hard work to get something demonstrably true is going to lose out economically in this paradigm, where we are all being conditioned to uncritically ooh and aah at the incantations being elicited from these magic boxes. The oracles even have legions of zealots who will berate you for not being sufficiently deferential and reverent, or worse, accuse you of blasphemy. If for no other reason, I agree with using the term to express all of the above succinctly, even if LLMs can be helpful tools generally.
I've read some of the results papers (the Einstein condensate one and the pi exponential one). I'm not an expert but it definitely wasn't AI slop. The introduction sections were particularly well framed and informative.
Also you can see in the papers where an idea is introduced but in the bibliography you can see where the foundational idea comes from. So the narratives are not unmotivated as some claim (proof without intuition claims).
In the context of maths papers, the term has come to refer to papers having the shortcomings that are, for whatever reason, typical of LLM out, including things like using non-standard terminology all over the place, emphasizing easy steps while leaping over harder ones, having bizarre organisation, and, importantly, failing to properly cover existing work and as a result being hard to tell from plagiarism.
The degree to which these issues feature will differ, but it is generally the case that converting the output to proper research requires significant effort, hence the AGMAI recommendations being what they are, and not performing that effort tends to come off as laziness or incompetence, so I can see how slop has become the popular term.
> We can't just call any output slop, or it loses all meaning.
The term “ai slop” is not supposed to discriminate good ai output from bad, the entire purpose of the phrase is a blanket term that delegitimizes all ai output.
That is exactly how I see it generally being used. Why else would people be dismissing work as AI slop without even reading it, discovering what it says, or even looking into how and to what extent AI was used in a project? Saying things like "if you didn't write it I won't read it" at the first whiff of an AI smell is absolutely said to delegitimize all ai output.
Or in this specific case, why would someone call these proofs (no one is saying they are wrong) AI slop if not to delegitimize all AI output?
People call some work AI slop "without even reading it" when the intention/substance of the work might exist somewhere buried within a wall of impenetrable LLM text (aka "the slop").
Good AI output is indistinguishable from human output. The whiff you mention is the reasoning pleonasm and tautology (intended) escaping into the output and the "author" not proof-reading/editing it out.
But it also happens in many other contexts where that is not true, such as this one right now. Bringing me back to my point that it’s not to discriminate and clarify between good and bad, but to muddy the water.
We can't argue the latter without quantifying the former. All terms are misused by someone, but if it's statistically insignificant that's not an issue. I'm not convinced this one is sufficiently misused to detract from the common definition.
> The term “ai slop” is not supposed to discriminate good ai output from bad, the entire purpose of the phrase is a blanket term that delegitimizes all ai output.
Then it's a useless term and we should all stop using it.
Haven't been following this debate closely, but what's the issue with "strip mining open problems"? Surely the supply of interesting mathematical problems is (in theory) infinite?
He argues that the supply nay be very large indeed but the interesting subset is not. Figuring out the interesting problems is difficult so strip mining the good known problems may lead to scarcity. I am not a mathematician myself, can not judge this accurately.
A Swedish proverb says, "a fool may ask more than ten wise may answer". This fool is reporting for duty!
I'm glad I may have something to contribute after all (and I'm only halfway joking)
I guess you can see this as an exploration problem, in pure maths, while the goal is to solve a conjecture, the limitation of humans on pure computational power led to the exploration of alternative paths. Sometimes, these paths weren't leading to solving the initial conjecture but opened new idea and new direction. Sometimes a less direct but more humanly natural path was taken to solve the conjecture which also led to new and humanly understandable questions.
In some ways solving the question wasn't the most important part of the work, as this doesn't have direct impact on our life (as I saw people comparing this with drug discovery), but the path leading to the solution raised new conjectures and techniques that further developed the field.
I have a really hard time reading AI proof so this might be a biased statement, but most of them feels like having a superpowerfull machine, that would have bruteforce all the possible words of finite length in your logical syntax. You have the path to the solution, using tools that where already known and even direction that where abandoned because they seemed to fail for our human brain. But at the end, as a mathematician, you don't learn anything that is really new.
To me this is the main risk with AI and in general the one most mathematican try to explain but fail, we might miss a lot of alternative path that would have raised more interesting questions (I think this is already more or less what is happening). On top of that, we will run out of mathematicians as no one wants to pursue a career in the field anymore.
This relies on the idea that AIs will only ever do the thing they just did, and nothing more.
It's the same argument which is invariably wrong yet comes up over and over again.
There's no real reason to think AIs solving lots of problems will stop further work on alternative paths - certainly a machine which never tires and can be trained on its own solutions is going to continue to improve.
There's precedent for this: just look at any overconfident post regarding what China will clearly never be able to do, despite decades of steady if frequently flawed progress.
There's no persuasive argument being presented as to why machine mathematical research should have a limit beyond hardware capabilities.
I think you are missing the points of my argument, my argument don't stand on AI isn't capable of discovering new techniques, as I don't believe in new techniques from the sky.
My argument is about any targeted goal based AI (which to the best of my knowledge is the case for LLMs as used now). My point is, if there exists a computationally bounded path from existing work that led to solving a conjecture and if the goal of the AI is to solve this conjecture, then alternative path that would have led to new discovery will be dismissed on the way (or lost in the computational trace if you prefer), leading to the conjecture being solved but maybe closing forever/for a long time new paths. I don't see how you could have as a goal to explore alternative path without a good metric of what is a good alternative path (like rating a chess position), which to me, seems unlikely to exist. If you don't have such metric then you would have a clear exponential blowup.
More like a percolation problem if you prefer, a neglected approach might have introduced a concept that would make further discoveries accessible. Missing that concept could therefore leave a whole region unexplored, not just one branch of one proof.
That’s what we are doing with nature, seas (look up strip mining there, it’s a horrible practice), and now the industrial harvestors are strip mining problem spaces. How do we like our own medicine?
Developing solutions to mathematical problems generally leads to improvements in quality and quantity of life at roughly the speed they percolate from the ivory tower down to the shop floor. So "how do we like it" is probably going to be "we like it a lot, this is awesome".
Every company is about to have a staff Ops Researcher who has a better grasp of the underlying math and theory than any university professor. That is an unambiguous win.
I see no reason why every company would have a staff ops researcher, or why such a position would have a better grasp of underlying math beyond the narrow slice that directly benefits the company. Why do you think that would happen?
> virtually none of this stuff is possible with technology any normal citizen has access to.
Not sure about the unambiguous win. Are we entering the age in which mathematics is industry-dominated?
1) Any university professor can spend their 24 years on a problem with little progress. 2) company has sudden interests. 3) industrial resources brute force the Lean proof. 4) Max PR for AI company 5) professors are left to rewrite the AI Lean slop into real human-readable math? {disclaimer non-math university professor}
No one is going to get tenure by spending 24 years on a problem with no results. The profs who have that much free time on their hands are already in the later stages of their careers with records of impactful results. By that time, a problem like that is more of a curiosity than sometimes expected to have broad concrete impact.
I'm no mathematician, but (1) seems like a bad situation to be in. I can't speak to the practical usefulness of potential mathematical solutions like proposed here, but it seems useless to have an individual professionally spend 24 years on a single problem only to make little progress and eventually retire so the next person can stare at it.
Imagine there's a very advanced crossword club where anybody can join and take a stab at these crosswords for the love of solving puzzles. Many of them are so difficult that no one's been able to solve them yet, but we know they're all solvable.
One day, someone comes along with a super advanced crossword solver application, and it makes easy work of these crosswords. They run it on a few to prove how powerful it is, and then the community says, "Oh wow, that's cool, but please don't run it on any more of our advanced crosswords because they're very hard for us to come up with, and we really enjoy solving them by hand."
That's really what this compares to. I wouldn't call that gatekeeping; just respect. Respect for the game, respect for people's desire to have these hard problems to continue to work on, solving by hand.
If the company with the super advanced crossword solver then continues to use it and publish the results, they're effectively stealing the crosswords from this community. Soon, all the puzzles will be solved, leaving nothing left for the community to work on for fun.
That doesn't sound like gatekeeping to me. That just sounds like someone asking "Please be respectful and leave the remaining puzzles for us to solve by hand.” A simple plea not to be an asshole.
We don't give mathematicians research positions to solve crosswords for fun. We want something back. We want theories and results that will advance our civilization.
We have people who want to fill those positions because there are enough people who find it rewarding enough. Take away reasons why they would find it rewarding and you will have fewer theories and results that will advance our civilisation.
And yes, fun counts. Nobody said this had to be only a hardship.
Money doesn't work that way though. There are plenty of jobs people would like to get paid to do, that doesn't mean someone needs to psy them to do it.
I'm well aware that if at some point AI is good enough to replace me as a software engineer then I won't have a job. I don't expect a company to continue to pay me simply because I enjoy it if there are cheaper options out there.
This is really the critical thing: the fun is the incentive. (Or at least the dominant incentive in math, historically.) As economists like to say, the overarching lesson in economics is that incentives matter. Reduce the incentives and participation will decrease.
Perhaps that won't matter if we enter an era where AI participants are the main participants who matter for discovery-level mathematics. But it would likely be what economists would see as a market failure if only a small oligopoly of AI participants, closely held behind closed doors, is able to fill that intellectual role.
I think you are missing the point of the main criticism. It is not about not wanting results in terms of proofs.
New theories and insights are typically created while working out proofs. If proofs now suddenly fall out of the sky (cause LLMs create them) then that work is not done which means the substrate on which new theories and questions and conjectures used to be grown disappears. It's in that sense that the math community (and thereby society as a whole) will lose something.
It's similar to how software engineering will need to find a solution to train their next generation. Current generations have all been through manual steps of designing things from scratch and writing them by hand. That's what allows your 10x engineers to understand whether what their LLM tools are doing is good and how to massage those tools to do the right thing. A junior engineer who has only ever used LLMs to write code and create architectures does not just not have that experience but also won't acquire it. You can't just say "we don't pay them to have fun and learn, we pay them to produce results". In the short term that is the case, but in the long term you as a company and we as a community will lose out.
I'm not saying don't use AI tooling. I'm saying that this is a hard problem which we yet to have to find solutions and approaches to. As a software community as well as as society in general.
"A junior engineer who has only ever used LLMs to write code and create architectures does not just not have that experience but also won't acquire it."
My ego tends to agree, that how can they be ever competent, if they have not endured the same hardships as I had crunching trough problems and getting allmost lost in the details.
But I rather suspect, they will turn out fine. I know LLMs are great for me to learn and I think the young generation will learn what they need to learn to get the job done.
Because keeping all the easy things coordinated and understanding the big picture is still a hard task yet unsolved by LLM's?
But yeah, who knows what happens once that change. I assume even after the singularity, it still makes sense, that we train some people to know what is going on ..
If a modern Gauss, Von Neumann, and Ramanujan appeared and started dropping proofs from the sky, would people be saying the same things? And if they could live forever, so they wouldn't need to train their replacements?
Yes, and they are revered as geniuses, which makes it clear that this is all sour grapes. And surely if people died, went to heaven, and were able to talk with God whenever they wanted, they wouldn't be upset that now they could know the answer to any mystery whenever they'd like; they'd appreciate that now they have someone to guide them! Or were they similarly upset when lecturers handed them already completed theory in school? There's already enough developed theory that people don't have the time to learn it all as it is.
Not exactly, because we would have cool people to inspire us and hang out with us.
But your argument is nonsensical because even if Gauss and von Neumann appeared, they wouldn't go into random fields and just prove things mechanically. They'd have to attend seminars, teach others, collaborate with others, and generally inspire others with their brilliance. It's the precise lack of this activity that makes AI in math so reprehensible.
Your argument encapsulates a contradiction because human mathematicians wouldn't be dropping proofs arbitrarily like AI is doing. They would do something completely different. Even the best of them.
Gauss was generally quite secretive and Ramanujan would famously tell people answers that he had received from divine inspiration, often with no ability to articulate how he knew. Von Neumann did just go into random fields and revolutionize them. If the three of them did come back from the dead and form a little powerhouse group that barely collaborated with the outside and just started publishing results for everyone else to try to keep up with, they'd no doubt still be considered geniuses.
Give it six months and models might be able to explain things better than any human. They can already collaborate perfectly well if you ask them to. e.g. there was a post here a couple months ago where Tao shared his ChatGPT logs[0].
If you're not inspired by the ability to talk to a superintelligent machine, and can't find what you'd want to know, that's a you problem.
The only thing potentially stopping these models from also outputting new theories along the way is the goal they were given.
I have to assume OpenAI is only prompting to solve problems, presumably they could also prompt to not interesting new theories or paths of research found along the way as well.
I think the crosswords framing is a little silly, but I have to wonder what comes when we use our technology to optimize the fun and interesting parts out of every job. There's only so many years of my life I can dedicate to back-and-forths with a chatbot. What if we advance our glorious civilization but our jobs just get more and more thoughtless and miserable?
I don't know about you, but my job has become a lot more fun ever since it's become a lot more back-and-forth with the robot. It does all the tedious things for me. It gathers data. It creates prototypes. It makes the mechanical code changes that I want. It allows me to talk with it for a design discussion, and then my design simply appears. I ask it for monitoring dashboards and they simply appear. It records what we talked about, which is something that I never do.
Largely I thought that this is what you do once you're established in math (or any field) anyway. You have some ideas, but the details are kind of too tedious for you to work out, so you give it to grad students/postdocs. Senior engineers have some ideas, but the details are tedious to work out, so you give them to junior engineers.
Now, obviously in the meantime, there's the question of how do we train the next generation? Or do we need to train the next generation? And maybe while we work that out the answer becomes more shadowing/apprenticeship instead of farming out easy tasks.
I think that's where people hope some kind if UBI or "universal high income" will save the day. Just don't think too hard about how it would actually be paid for, or how we can all have high income when that's a relative measure and we're all given the same amount of table scraps.
"universal high income" is not when everyone has high income, it's when everyone who doesn't have a high income is excluded from the universe. There will be few high income people, robots those people own, and the rest of us will be undesirables/illegals/felons/noncitizens of Ms-Apple-Meta-Tesla-Google-topia, who for arbitrary reasons XYZ (they didn't accept the EULA!) don't deserve universal high income (i.e. most people here will fall into that category).
We're going to have a very different perspective on purpose going forward with these results. This has crossed a rubicon where human output itself is going to be completely outclassed by machines and we will have to find meaning elsewhere in life.
If your only measure of advancing is getting an answer, but not building the capability to understand it, then civilization has advanced.
It’s not a human focused civilization, which is where the issue comes up.
As an example: A constant issue I am seeing with AI productivity is that the most productive use of AI is when it is paired with more experienced users, while AI also does more work for entry level workers, if not replacing them entirely.
It has become a question where will the future buffer of experienced seniors come from.
This is an example of where simply chopping down trees for today, doesn’t make civilization better off tomorrow.
AI is producing more content than ever before, but our ability to understand and verify it is not keeping pace.
We don’t know if these are unsolvable problems at this stage. Society could come up with workarounds and solutions to these issues in several years.
The request to stop, is part of the process by which the issues are debated and solutions found. It doesn’t mean their position is weird or moot.
If someone gets the answer sooner than you, that doesn't inhibit you developing your understanding of the answer privately the same way you would have done if they hadn't got the answer. I don't see how anybody loses by the answer being discovered sooner.
Not true. If I know the answer to a puzzle, I don't spend the time doing the puzzle.
If there is a prize associated with doing a puzzle, and a machine does it, then what incentive is there to pursue it.
Again, if you are only concerned with the outcome, and you have a preferred answer that you want (in this case "just use AI to advance faster"), then any information that doesn't support that case is useless or misguided at worst.
I am not trying to dissuade you from your preference. I am flagging that there is a set of other factors that influence the behavior of others, how that behavior is critical to the creation of expertise and drive, and thus why others hold different positions.
If you're concerned with something other than the answer, then the fact that the answer is already known hasn't actually provided the thing you're concerned about, so you can still do the thing you are concerned about.
If another human was likely to get the answer before you would you also discourage them from doing it because they would rob you of the chance to do the thing you're concerned about?
This is an ethical and moral question being added here.
Would it be unethical to dissuade someone else from enjoying the benefits of the process you wish to enjoy ?
Vs
Would it be unethical to stop a machine from data mining all the possible questions you wish to explore/enjoy.
And on another level - I am concerned with a bit more than just the answer. I am concerned with what system is in place to ask more questions and get more answers.
There is nothing in this argument that says that we won’t find some other way to study the subject. Maybe people will become monks and do math as a hobby.
We may end up in a daemon filled world, like 40k, where any hope of understanding the tech around us is impossible. (More impossible that today)
If someone spends their entire career not solving the puzzle, did they really learn to understand how to solve it?
They may very well have learned plenty of things and solved or discovered other puzzles, but if the first puzzle is worth pursuing because the solution is actually useful it seems liked we're better off with the solution than a bunch of failed attempts.
That said, I do question the value of solving many of these types of math problems. I'm no mathematician so I'm assuming I'm wrong here, but on the surface many seem mostly theoretical puzzles with little or no practical use.
Yes? We haven’t solved many puzzles about reality, but even half proofs and conjectures create tools that other people use to make progress.
I’ve made this point elsewhere but the debate here is between two different philosophical positions. Results vs process.
If all you care about is the results then the process doesn’t matter.
If a person is starving or needs medicine, then a long discussion on process is inhumane. They need results.
If the conversation is about process though, then focusing on the results is missing the point.
I’d say the question for results oriented people is what are the benefits of the process and at what point does it make sense to optimize for results vs process.
My read on much of the discussion here is that the debate is whether we want AIs solving problems that career mathematicians may spend a lifetime on and still not solve.
When the topic is about careers the question really has to be about results. Even if the results are made by solving different problems discovered along the way towards their original problem, it still has to be about those results.
There is absolutely a question of whether burning these resources is useful when the only outcome is a solution to a potentially obscure math problem, but that is more a question of prompting and goals rather than the use of these tools themselves.
> There is absolutely a question of whether burning these resources is useful when the only outcome is a solution to a potentially obscure math problem, but that is more a question of prompting and goals rather than the use of these tools themselves.
But isn't all of schooling literally learning solutions others solved before us?
We spend most of our young lives (many of us our entire lives) studying physics, math, etc. that others have solved. (e.g Quantum Mechanics, Relativity, Calculus, etc.)
Biology consists, almost entirely, of studying solved problems in nature.
That doesn't answer the question. Assume today is not the stopping point, and that we end up with super-intelligent theory building AIs. Better than any current-day human. And better at explaining, creating visualizations, etc. than any current day human.
Why is it a problem that the professor is now a robot, and that humans could spend arbitrarily long learning from it and even after 15 years of masters-style advanced graduate lecture courses still have deeper still levels of the topic that the AI could teach them?
And if they never do reach that level of ultra-competence, well, then we found the niche for humans to continue to exist within.
What is preventing these crossword solvers from not looking at the advanced crossword solutions?
Mathematicians and academics in their ivory towers are forgetting that everything is getting automated. They want to carve out fun problem solving niches that's fine but who's going to fund that? If they want to be funded by the society/civilization their argument can't be leave advanced fun problems for their hobby.
Its worth noting though that you are comparing a profession with a hobby.
People go to said crossword group to enjoy the process of solving the puzzles. It doesn't actually matter if they have been solved yet or not, case in point the NY Times puzzles are enjoyed by more than just the first to solve them.
Professional mathematicians are ultimately being paid to solve the problems for a (hopefully) practical reason. Its always excellent when a person enjoys the process of the work they are paid to do, but ultimately they are still paid to do the work. I really hope your argument isn't that we should collectively be funding mathematicians to solve problems simply doe the love of the game.
> Professional mathematicians are ultimately being paid to solve the problems for a (hopefully) practical reason.
They are paid for the same reasons the NEA pays artists: out of a sense of obligation to demonstrate elite culture.
The track record of practicality of pure math after WWII is essentially 0.
While I don't disagree, I think any justification for why we should fund mathematics and why we should protect the work they are doing from being solved without them should be grounded in results.
Similarly I wouldn't expect a good argument could be made that AI tools should be prevented from creating art because we want to continue funding artists.
If the goal of said funding is just to let them spend their time doing it then it doesn't matter that AI is doing it as well.
Despite nobody at openAI thinking of themselves as an asshole; despite society urging openAI not to be an asshole; despite the fact that being an asshole is entirely unnecessary even to accomplish whatever objective they are setting out to accomplish; despite everyone at openAI loudly declaring: we are not assholes!
It's done in jest but I think I am accurately pointing out the interesting parallels between what these companies say they are doing (aligning models) and what they are not doing (aligning themselves).
If you listen to them, and you don't have to listen very hard to hear it, basically everyone at these labs is telling us that this technology is extremely dangerous and should be slowed down or paused entirely. Yet, they, the only entities with the power to actually do anything about it, are not acting AT ALL as if that's the case. They are all barrelling forward as quickly as possible. RSI, THE number one risk according to these guys, is being adopted at breakneck pace up and down the stack, from designing silicon, to training, to inference.
It's ridiculous and insane and I believe can be accurately summed up as, they are being assholes, because if they are actually right about this we are all gonna die. At the very least, and far more likely, every fun creative expressive human thing that is machine legible will be replaced by a torrent of machine slop. It's not "benefiting humanity." These mathematicians are telling you it's not benefiting humanity. It sucks.
This analogy is silly because (a) math is not primarily for entertainment, (b) we aren't going to run out of math proofs, and (c) results build on top of other results, having more results proven makes all math more powerful and useful.
Hmmm but in the case of math, while some of it is "just puzzles" there often turns out to be practical applications, even if they are not obvious at first. Number theory was considered the epitome of pure math with no practical applications for centuries, now our modern society is built on it (public key crypto).
If the crosswords were purely games that would be no problem. These crosswords seem to power physics, chemistry, engineering and science applications. These professions would not mind it too much.
Blah blah blah. They are free to do their own mathematics and/or spend time on polishing/reviewing proofs dumped by ai. But they don't get to make demands like don't test math on proprietary models. Idiots.
Math doesn't belong to academics.
We don't pay them to work on problems for fun.
They will just need to re-evaluate where the value their provide is. It won't be solving problems anymore. Hopefully it will be making them understandable by others at least till AI can't do that as well.
"They will just need to re-evaluate where the value their provide is."
That is fine to say when it is not your field. I guarantee you feel different when it is the thing you care about, that gives you joy, that defines your status. Think about how many sheldon-equivalents insist on being called Dr. (non medical)
It is part of what people use to define themselves. Its going to hurt. There may even be a Bulterian Jihad
No, you dislike maths to the point you prefer paying others to do it. Actual mathematicians are largely doing it for fun, but are now effectively saying "stop destroying our fun or we'll stop doing maths", and you will have to do the maths yourself.
Whole sections of the economy are being upheaved by AI, and there is no reason to make a special case for the mathematicians anymore than for the illustrators, developers, translators, HR, etc.
Oh right, how rude of me to only talk about mathematicians in this thread about the future prospects of children in Sudan. Of course this is the place to make "what about the illustrators" argument.
It isn't some law of nature. Humans/societies have agency - what AI should or should not be used is up for debate and decisions. It might even wind up the other way around that using AI is the special case - who knows.
> what AI should or should not be used is up for debate and decisions.
Of course; but it's very hypocritical to raise these feelings only when mathematicians are affected, whereas all the above professions are just told to adapt to the new way of things.
For sure though, translators don't have the same clout and social status as mathematicians do.
I actually really like math and I can't wait for the day LLMs not only solve difficult problems but can also explain the solutions to me.
Mathematicians do a terrible job here. They use inconsistent symbols they don't even explain. They often obfuscate the main idea just to make the paper longer. If you are not part of a small club you are not meant to understand it.
I think this is a terrible approach and I am eagerly waiting for AI to do a better job!
But won't new humans take their place that will be the ones who enjoy deciphering AI solutions?
It just seems that this class of mathematicians is being "disrupted".
The field is changing and a new class of mathematicians will take their place.
This happens all the time in fields as technology disrupts them.
A new class of individuals, with different motivations, take the place of the old guard.
I'm sure the motivations of individuals involved in designing and manufacturing cars changed as Henry Ford introduced the factor line.
But that old crop of humans either adapted or retired.
But, plenty of humans took their place with new motivations and automotive technology continued to progress.
I personally feel math will indeed move faster as a result of these breakthroughs. And the humans that take the place of the old guard will have different passions and motivations than the current group.
Maybe the new group will be productivity motivated rather than motivated by the love of tinkering with a single problem for years.
I’m sorry; but if mathematicians are in it because puzzle club is fun, then they should go join the fucking puzzle club and stop impeding scientific progress.
Science isn’t some passive busywork thing where you tie your hands behind your back because it isn’t fair on others to solve all the neat problems - or at least it shouldn’t be.
If your idea of science is leather patches on tweed suits and the quiet ticking of a clock while you do crosswords, then this is an argument in favour of letting the AI do the work so you can focus on your sudoku book in your slippers.
It's more like, "don't just casually destroy our hobby / career field", without letting us participate even a little.
The picture I have in mind is OpenAI running their most advanced model in a loop over all the open mathematical problems they can find, just to verify that the model is indeed very smart. Neither the company nor the model actually care about the problems, it's just a cheap exercise machine for them, but the problems get solved and mathematicians don't even get to participate.
Like, even those who accepted the "centaur" thinking, man + machine, won't benefit because by the time they get their hands on good enough models, everything is already done.
It's an emotional thing first and foremost - people who care about the thing can't do the thing, because it's already been done by those who couldn't care less about it.
And before someone goes "poor mathematicians", a food for thought: this is just an early instance of what looks like our shared destiny.
I said here before: given the economics of progress in AI and robotics, it's obvious what the natural division of labor is: computers do the thinking, humans do the menial, manual labor. AI will do politics and philosophy, so you have more time to fold laundry and scrub the toilet.
So what is mathematics then? A fun hobby akin to chess or sudoku?
Are we gonna get the same pushback from medical researchers if the models cure xyz diseases?
I absolutely understand the emotional connection to their work and the heartbreak, but mathematics doesn't exist for their pleasure, it exists to provide tools to solve humanitie's problems.
> So what is mathematics then? A fun hobby akin to chess or sudoku?
Some of it, yes. Much like physics. Both have a track record of producing technological breakthroughs every now and then, but it's not why people are doing it.
> Are we gonna get the same pushback from medical researchers if the models cure xyz diseases?
For better or worse, yes. We already are. In my country, there's a big spat between radiologists and cardiologists right now, that boils down to the progress of technology allowing the former to answer questions that, before, involved a procedure that was a big money-maker for the latter.
What you say reminds me of medical schools in Tunisia.
The general body of research points that more doctors lowers all cause mortality ( with diminishing returns) but Tunisia is still far lower than the Eu average.
Yet Doctors and med Student unions do lobby very heavily against expanding admission to the public uni or allowing private unis.
So we have the weird situation where people go and study in Romania ( making Tunisia lose hard currency that it really needs).
These doctors have taken an oath and the direct consequence of their lobbying is literally more deaths.
USA is the same. Even worse, the doctors guild writes the rules for creating new doctors. It got so bad that we now have 2 or 3 other alternate/adjacent categories of doctors and nurses to work around the bottleneck. Of course then they formed guilds to continue the cycle.
Dude, doctors are humans just like rest of us. They want careers, money, safety, raise children in best way possible, fun in life and so on. I see this unspoken expectation over and over - why are they not infallible, how could they do mistake XYZ, why are they not working themselves to the (early) death for benefits of us all and so on. They have no obligation to stay at place Q just because some folks would consider it convenient. They have no obligation to stay in some place thats not suiting them just because they swore Hippocratic oath, lives can be saved elsewhere too.
Obviously this is often coming from folks who act in same ways as they criticize and usually don't contribute even a fraction back to society compared to doctors. Folks who do mistakes in their lives all the time yet thats fine since we are all humans or similar, right.
So please stop this cheap framing and accusations. If Tunisia wants more doctors and keep them there are ways to do it, society as a whole needs to decide what they want and act upon it. Otherwise, smart skilled folks will keep going for better lives elsewhere, just like everybody else.
Everyone (near enough) has some degree of self-interest. If you apply for a job and discover that some other applicant is about as well fitted to it as you and in more need of money, do you withdraw? If you see a $20 note on the ground and no one else around who might have dropped it, do you refrain from picking it up if you think you're better-off than the median person who might walk past next? If you see something you want going for a very good price on eBay, do you contact the seller and say "I think you should be making me pay more for this"?
Unless you are an extremely unusual person, the answers to those questions are somewhere between "no" and "of course not, and why would you even ask?".
If someone is working as a doctor, their work is already benefiting others substantially more than the typical person's. (At least, I think it is; it's certainly doing so more directly.) Being a doctor doesn't put them under some unique obligation never to give any priority to their own interests when, e.g., choosing what job to take where.
If they can save 0.2 lives per day for $50k/year in one place and save 0.19 lives per day for $200k/year in another, it would be virtuous for them to do the former but I can't see that it's obligatory. In the case we're talking about, it might actually be 0.2 lives per day for $50k/year versus 0.21 lives per day for $200k/year, because somewhere that can afford to pay them more can probably also afford better equipment, more ambulances, etc. (In case it isn't obvious, all actual numbers here are made up and nothing I'm saying depends on exactly what they are, only on the rough relationships between them.)
It seems to me like any principle that would oblige them to pick the first of those options over the second would e.g. also oblige all of us who have well paid jobs to give most of what we earn to life-saving charities. Some people do that. It's a virtuous and commendable thing. It would doubtless be better if more people did. But, as you might have noticed, very very few people do that and by and large we don't consider it outrageous that they don't, and I don't see why doctors in particular should be condemned when they don't do it.
(Since clearly unassisted human nature isn't going to make everyone behave in such a way, it seems to me that if we wanted that sort of thing then it would need to be imposed by force. Which in fact everyone might be OK with, in the same sort of way as players of high-level sports are OK with having externally-imposed safety rules so that we don't get everyone playing in increasingly dangerous ways for the sake of a small advantage over people who are being more careful. And, in fact, we do have that sort of thing and it is imposed by force; it's called taxation, and actually I think it's a beautiful thing even though there's plenty to dislike about every actually-existing regime of taxes and benefits. This is mostly a digression, but note that it means that if a doctor chooses to go somewhere where they're paid better it probably also means that they're contributing more to the general welfare in taxes. There are plenty of nits one could pick with this remark, but it still seems worth making.)
At high levels, often yes. At lower levels, often it's job security.
Most doctors aren't running departments in major hospitals, or advising government on policy. They don't earn the big bucks. And even hospitals themselves tend to run in the red all the time; it's sometimes hard to disentangle where greed ends, and longer-term interests of patients begin, as you have multiple people and organizations pulling in different directions for different reasons.
RE private medical universities, N=1 but in Poland we have a private provider pushing hard for training their own doctors "because public system is too slow and limited", and it's hard to tell whether they have a point, or whether it's a private-driven attempt at privatizing national healthcare, or a mix of both.
>but mathematics doesn't exist for their pleasure, it exists to provide tools to solve humanitie's problems.
The risk here is that this does do fundamental long-term damage to mathematics as a viable field.
Virtually no one is going to want to take on the risk of PhD-level math work, studying a narrow problem for four years or so to arrive at an impressive incremental result, when there's a sword of damocles hanging over their head every day that an internal system held by an oracle they don't have access to may scoop their results and turn those four years into dust.
To some extent, that sword of damocles always existed in a de minimus sense in the form of other mathematicians. But everyone was playing the same game, coming to the game with the same arsenal limited by human cognition.
If the game board becomes irrevocably tilted, new entrants have no incentive to play except as a hobby. But few hobbyists can devote years of work to understanding and pushing the frontier. It could well mean existential damage to mathematics as a field.
Whether that might undermine math's ability to solve humanity's problems in the long term is almost an economics problem, not unlike the question of whether and when the existence of monopolies ultimately restricts long-term economic growth. Much probably depends on whether intellectual monopolies or oligopolies are being created that will supplant the existing mathematics "economy".
> The risk here is that this does do fundamental long-term damage to mathematics as a viable field.
All the commotion evens out: It's much easier to learn maths than ever before; you don't need to go to lectures any more; you don't need to learn from a specialist (advisor, lecturer) any more; it all costs much less than it used to.
So mathematics will continue to advance, albeit differently from before. The social structures will not survive however.
Certainly it'll result in a boom for hobby mathematics, and it'll be a hobby at a much more advanced level than before. Whether those hobbyists can continue to push the actual frontier, particularly if AI models operating along that frontier are not made accessible to hobbyists (either via corporate/AI lab gatekeeping, via pricing, or via significant time lags) is a different question. I'm a little more confident in a future where hobbyists push the frontier in applied mathematics than in pure mathematics.
There's probably a loose and deeply imperfect analogy with computing: via democratization hobbyists have made a big impact in applied operating systems development (Linux/OpenBSD) but have been less successful/impactful in OS research (whither Hurd...) or in cost-heavy fields like microprocessor design.
The trials process is the moat. There's already founders using AI to treat their cancers, and it's all about skipping trials and jumping straight to "I consent, I'll fund it, let's try it". The general public might get access to this in 10 years, but employees at AI companies will have access much much sooner.
I don't necessarily see a problem with it: if people want to try experimental therapy on themselves and can fund it, then as long as it's expensive, let them - that speeds up research. The problem with allowing anyone to opt out of safety trials is that it then creates pressure from doctors and family members to try, and then it becomes non-consensual in practice.
Yeah it's more like personalized therapy - often the only hope for rare diseases.
While AI has definitely helped quite a bit I am wondering how much all this research and treatments cost. Not sure the current health systems could sustain this for _everyone affected_. If ai enables it all the better.
At least in Sid's case, it went from the oncologist saying "I have no more drugs I would recommend, no trials available" (slide 7) to "I currently have no evidence of disease" (slide 18). I don't know beyond that or beyond Sid's case - or a similar story of an Australian who treated a cancer tumour their dog had with a similar AI / personalized vaccine process.
My understanding of what Sid's describing is that you do RNA sequencing, a whole genome sequencing, feed that into frontier AI (if it will still let you), and somewhere along the way give the information the AI finds to people who can use it make a personalized mRNA vaccine, specifically for you and your cancer.
Another link here about Sid's case, it explains it didn't go through trials: "made possible through a compassionate use allowance from the U.S. Food and Drug Administration (FDA)".
Trudging into the technicalities of the example still doesn't undo the question of "What is the point of mathematics? To find answers or to be a hobby?"
It's tempting to say "both", but that misses that AI is now forcing us to pick one.
What if the point of mathematics is to be mature enough to study and teach math to help humans understand it, without the ego stroke of being the first to solve a problem? Bad communicators are upset that a robot is better than they are solving problems.
> mathematics doesn't exist for their pleasure, it exists to provide tools to solve humanitie's problems
Who decreed that? Mathematics predates capitalism and publish-or-perish by a couple of millennia. Euclid’s Elements were not written to benefit the weapons or medical industry.
Sure, and sometimes gates are needed. That's why we all run spamfilters, those are definitely gatekeepers.
In this instance however, it's openAI and Anthropic that are pushing people out of the field by running secret models that take the interesting work away and leaves the persons having to review endless slop proofs.
There has never been a stronger need for people to band together and "seize the means of production" for this stuff. The advances being made are ours, not theirs. It's trained on our work, our knowledge.
> The work & knowledge this is trained on is public.
That's an incredibly generous take. If I'd pulled a fraction of the shenanigans prominent companies have to obtain data I'd be thrown under a prison to the thunderous applause of those who have, and are, doing much worse.
We _think_ this power / divide feels harmless right now, but I'd bet money that NSA, CIA, etc have access to the latest and greatest unrestricted models; and massive compute. At least for OpenAI, and even if not willingly for Anthropic, I'd bet money NSA has it too. (After all, when Google decided to migrate to HTTPS, the NSA decided to hack Google's internal network to preserve their taps).
One thing I've wondered about in this respect is what happens if NSA learns 5000 new units of math while the general public learns 4000 new units of math.
This sort of happened at various times in the past, because they hired and/or funded so many mathematicians, and especially before the late 1970s they had many of them working in areas where academic mathematicians weren't working at all, so they were learning more math, or more math that they especially cared about, than the public was. (I was going to write a note here just a few days ago about how NSA has had a "Classified Mathematics Library" for many years.)
For vulnerability scanning, I think the new-capabilities trajectory is good (in the sense of "it will help defenders win") even if governments find ways to get more of it, because there are finitely many bugs and classes of bugs, so at some point more capable models' or longer runs' advantage over less capable models and shorter runs should stop helping them outcompete the less-well-funded defenders, because the defenders will still have learned most of the information that's relevant to achieving successful defenses.
So if NSA gets 5000 units of vulnerability scanning and the public only gets 4000 units, we might still just wipe out all of the pure software vulnerabilities and then go back to worrying about physical supply chain security or side channels or something.
For math, I'm not quite sure! For one thing, there may be things that have no feasibly deployable defense at all even when you understand the underlying mathematics (I'm especially worried about traffic analysis here, because understanding in detail how traffic analysis is done, or how powerful particular techniques are, does not necessarily always or usually make defending against it more convenient or less costly). In a more science fiction scenario, there might also not be any efficient secure cryptographic primitives of some kind, like if it turns out P=NP with reasonably small exponents and reasonably small constant factors.
Based on people I've talked to I'd be really surprised if this was the case, they actually seem to be pretty far behind the ball when it comes to AI use. Which makes sense to me, given the sensitive nature of their data and systems, they don't want to turn on yolo mode and let an agent cook unattended, which is what you need to do to make these discoveries.
I believe it would be a complete failure of the state and frankly downright irresponsible behavior if all the three letter institutions didn't have access to these models and I’m not even a US national nor do I live there. It’s just common sense. Obviously it wouldn’t be public information since it’s national security, but it’s the lowest hanging asymmetric advantage in the history of national security of nations.
> virtually none of this stuff is possible with technology any normal citizen has access to
So far, it looks like open-weight models are lagging less than a year behind frontier capabilities. And I think one year diffusion of technology from "insider lab demo" to widely available is actually pretty fast?
There are lots of research fields which "normal citizen" has no access to - medical and biological research, particle physics. Some of it is somehow publicly controlled (LHC), some of it not at all (commercial pharma research, mostly secret until the final human trials). And most of it reaches "normal citizens" in way more than a year.
(and I'm talking about open-weight models. The availability of commercial AI models from private preview to included-in-your-$100-subscription is currently like 4 months)
I'm wondering what's the impact on human Mathematicians, and especially would-be Mathematicians -- master students, if they HAVE to use AI in their daily life?
Would that impact their own ability of solving Mathematics problems? I mean as a programmer I'm already seeing that impact on the programmers -- sure the best of us can leverage AI to achieve unimaginable things, but many of us are simply vibe coding.
Of course we can assume that it is only the best of us that really matters, and the rest of us are not going to produce anything substantially useful ANYWAY, it might as well to replace the rest of us with AI, but my worry is -- does that really have ZERO impact on the human specie's ability to produce "the best of us"? After all, they don't grow on trees.
It's a grand experiment isn't it? Us senior programmers are pretty good at using AI (or so we think) because we have decades of grinding and problem solving to inform our intuitions. Is that really necessary? The next generation of programmers certainly will not have that level of desk-head interface. Maybe they'll be fine? Maybe the models will get so good it won't matter? Open question.
I imagine the same will be true of AI, but I'll say that in the short term AI is going to make mathematicians better because it solves the breadth problem. Again, I feel like this Barnette conjecture got solved (if it is solved) because of some clever partition function sums which are intellectually tractable but simply too far out of anything I'd seen before (I see the apparition of my GT combinatorics professor intoning gravely that "everyone knows that, Jake, you're an idiot"). Maybe AI will help identify common threads far greater than Google and journal search.
I think if I had ChatGPT when I was 20 and working on this problem for the first time I might not have solved it, but I would have learned every angle and facet of it far more quickly. But then again I would not have spent so many late nights staring at the Országház across the Danube and letting my mind drift and bump against the problem like spilled cargo in the river.
I have been thinking about this, too. Take Mathematics as an example — it’s probably safe to say that only the top 1000 contemporary Mathematicians really matter to the human specie, or perhaps even less. And if you do not show the potential to be one of those when you reach the end of your graduate studies (actually probably already too late), you are 99.999% sure to just push out papers no one reads and such, and an associate professor in a no name school is going to be your lifetime high watermark. Like, the human specie doesn’t care whether you existed or not, from that perspective.
Now if we can prove this, expand it to the whole spectrum of academic studies, and somehow convince 99.99% of us that they are basically garbage and we don’t care about them — sure the elites will throw UBI around but that’s it — then maybe AI is very positive to the human specie.
Oh we better pick up the speed of cloning and artificial fertilization quickly, because people who are told to be garbage probably have no interests in boring children, and it is still a myth how genies are born and grown. We need that diversity.
BTW the whole scheme reads like the background of a Chinese net novel 赛博英雄传.
I went back to that Fable chat and showed it this new preprint. It coded up the new constructive algorithm and ran it against the existing test suite, that looks good at least.
It has been super helpful in delineating where the crucial concept came from. The proof is rather simple as graph theory proofs go, but it does seem to use some constructions that would only seem obvious if you had serious physics experience with partition function and calculating energy states that cancel out. It's not a wholly alien bolt from the heavens, but I can also see how there hasn't been a human being with the broad theoretical physics knowledge combined with the deep graph theory experience in planar graphs to come up with this idea. I don't know, I'm looking for precedents of this formulation and some old papers of Penrose counting the number of edge colorings of this same graph type are coming up, the line of argument at least rhymes.
But I agree with the thought that this sort of progress should not be siloed inside those companies. I propose a tax so that every slop cannon AI video pays for another hour of compute time for advancing mathematics.
virtually none of this stuff is possible with technology any normal citizen has access to
I suspect that this might be one of the reasons people inside the labs are scared about AI.
What if they have asked AI how it would wipe out humanity and it came up with reasonable answers that they don’t want to publish unlike they do with these math problems?
I think those models and findings should be investigated.
The ways AI can eliminate humanity are trivial obvious and already published.
It's just "let the AI control anything of importance and let it spit out slop"
Anthropic runs a biology wetlab (while denying biology to consumers of even their publicly available models, let alone their inhouse ones that only they can access) so I'd expect AI to generate practical and lucrative products soon.
Cure for aging? What do you reckon that'd be worth?
I always got the sense that solutions for significant "unsolved problems in medicine" would be at least 10 years out from the point of total AI dominance in the theoretical sciences. Doing actual experiments is bottlenecked by real-life constraints (organisms are slow to grow and unpredictable, human laws won't let you build a factory to brute-force biology on a million test tube guinea pigs, let alone humans), and the theoretical side of biology is also relatively underdeveloped, to the point that "solve aging" seems as hard to formulate as Navier-Stokes would have been with 15th-century mathematics.
That would be disaster. It would mean the world would not get rid of trump (and similar) by natural causes.
Death is the final - and perhaps the only? - equaliser.
If they find a shortcut (like a viral injected cell-dna damage reset) - that would be big. And can you imagine handling the cure for aging, to societies that still produce exponential people?
A cure that you take once and that's it, your body is that age forever? Now, a supplement that you have to keep taking to stay that biological age, that's where the real money is.
> It's becoming an incredible concentration of power that I don't know that we've ever quite seen before.
Replace “AI” with “supercomputer”.
(Super)computers have been solving many math problems that mathematicians can’t solve. Now they are capable of solving problem types that they weren’t able to solve before. (this applies to other fields as well)
Problem is it’s not clear if there is anything left for humans. Probably yes, since human mathematicians are still more economical.
> I want a jet airplane, but I can't afford one, and all the ones that exist are proprietary.
I guess if you worked together with some people who all put some money into a fund, and by using very modern technologies like 3D printing and modern CAD modelling etc., it should be possible even for private people to build a jet airplane.
The problem rather is that the government does an insane amount of gatekeeping to prevent this from happening (enforcing expensive and time-consuming certifications on airplanes and pilots etc.).
You're talking about an end user not being able to afford a luxury item.
The concern is about elite level researchers no longer being able to move the industry forward in a public way, and leaving potentially all major discoveries in private hands going forward.
Possible worst case scenario in your case, you personally miss out on a luxury item.
Possible worst case scenario in the topic case, an AI company controls the only intelligence that discovers and understands the most powerful tools / physics we know of.
They no doubt have more expensive/powerful models internally, but smaller models seem to catch up fast. So I'm not sure it's about capabilities, but more the willingness and budget to conduct a huge search.
Obviously the more intelligent the model, the smaller/more directed the search is. But they spoke about huge numbers of agents working on Navier-Stokes for example (I think it cost >$10m).
True. What if the emerging capabilities of their best models are applied to tasks like “maximize the chances this pro-AI candidate wins an election” or “maximize profit via stock trading”. Every advantage compounds until all power in the world with any significance belongs solely to whoever has the best models and most compute.
> I suspect that this is in fact the source of much of the angst.
Your comment reveals that you absolutely did not read or understand the Field medalists' open letter... Please, why would you refer to their complaints and claim you disagree when you clearly aren't engaging with the arguments presented therein!?
Totally agree - and not only that we don't know the exact details how these results were produced which is deeply problematic - we just have the end result (and some of the reasoning traces). For this to be a scientific disclsure, we need to know what the agentic setup was, what information was put in, how much and which prior work it relied on, whether the constructions it's using are just ripping off existing work without citation or something it invented (and if so, to what extent) and so on - it's not clear at all what the actual new contribution of the AI model is. All this makes it feel much less like an actual scientific contribution and more like a pre-IPO stunt.
But to me it also signals (as if it didn't before!) a great need for the wider AI community to focus exclusively on researching and building AI algorithms and systems that are more humanistic: completely transparent in its workings and the representations they create, super efficient in terms of data and compute, componentised so that individual entities can plug in different bits and rapidly train on their own data, highly adaptive to individual needs, programmable in a real sense, largely independent of corporate influence, easily accessible to everyone across all social and economic strata, and enable individuals to grow/learn/reach their full potential.
Is this possible? I think so, but it will require ingenuity and bringing in ideas from (ironically enough) some of the deepest areas of modern mathematics such category theory, algebraic topology etc. which are largely about building abstractions that expose the underlying structure of complex mathematical objects and the relationships between them.
It's already happening to a degree, but the urgency has reached epic levels at this point and it needs to happen at scale.
It's a bit aggravating that I cannot interrogate the session that yielded this result and ask it why and where it got the crucial calculation from, or why it went in that direction. It doesn't even rightly know even if it gives you a legible answer, that doesn't have any correlation with whatever happened under the hood.
Humans are the same way sometimes but I guess there's romance in that. If a human had solved it a la Kekulé and said "it came to me in a dream" I would at least understand that.
Sorry I was using scientific in a broader sense - probably should have used "academic" instead -
it's deeply problematic because they are building on open, public results yet they don't provide information on how people may build on it - its exploitative and exclusionary - at least they are consistent
I find this argument to be extremely ridiculous. They solved some math problems and published the results for free. No one asked them to do it, they weren't paid, and they don't owe anyone anything. Who exactly is exploited and excluded? The entire notion of open public information is that you can do anything you want with it, including build private systems. Is a baker "exploitative and exclusionary" for reading a recipe in a book and then turning around and selling that bread to customers, without sharing the recipe with the customers?
The anti-AI arguments keep morphing, as many could have probably predicted. Starting with "AI can't do anything" to "AI can't do anything useful" to ... "AI breakthroughs are proprietary!#@!!!".
I've seen more goalposts move in the last 3 years than maybe in my whole (lengthy) career up to that point.
> I have no problem with AI models making revolutionary advances in math or science. Where I start to have a problem is when the AI models making these advances are tightly withheld, proprietary, and seemingly never released with these capabilities intact.
Agree, and, to my mind - shows why the efforts of the Free Software Foundation have been worthwhile all along. We need software to be open / free / libre or the power elite controlling them will ruin the world.
What exactly are you worried about? OpenAI/etc. gaining too much power? If they use it, the government can stop them. If you worry about the government, isn't it better that than rando terrorists? Seems similar to the early days of nuclear and rocket technology. It took stupendous amounts of money and smart people. It was barely accessible to many countries let alone people.
> What exactly are you worried about? OpenAI/etc. gaining too much power?
Yes. They have already shown to have no scruples when it comes to making profit and to have little to no morals.
> If you worry about the government, isn't it better that than rando terrorists?
In my country the largest terrorist attack was almost certainly financed by Iran and caused roughly one hundred deaths. This number pales compared to the thousands who died during the latest, US-backed military coup, a move that relied on a doctrine that the US has never stopped asserting [1].
And those morals I mentioned earlier from AI companies? They do not apply to me because I'm not a US citizen. So no, I do not think the US government is the "seal of quality" you think it is.
I don't think the comment you're replying to said that the U.S. govt. is a seal of quality, at all. They kind-of implicitly concededed that trusting a government with that power is highly sub-optimal, but better still than allowing it to get into the hands of terrorists. Which is a very real issue and a nontrivial point of tension. Like, I'm sorry, maybe I'm reading into this too much, but I personally see the "the government is not the seal of quality you think it is" as a rude and even patronizing misinterpretation happening far too often in discussions, and as needlessly diverging attention from the crux of the problem.
I want to push back on "better than getting into the hands of terrorists being a very real issue".
I am currently in Germany. In the 21st Century roughly 60 people have been killed and 160 injured in ~40 terrorist attacks, most of them perpetrated with cars or knives [1]. In comparison, the US' war in Iran has costed Germany 2.781 billion dollars in fuel costs this year alone and the US government has publicly announced its plans to interfere in German politics partially by funding far-right activities [2].
My point being: the probabilities of terrorists shaking the world order with AI are rather low, seeing as even the most successful attacks in this century have been performed with the simplest of technologies. In contrast, the probability of the US flexing its power irresponsibly are rather high, seeing as they have been doing it for a couple years now and are, in fact, doing it right now.
As far as I'm concerned, and from an evidence-based, day-to-day point of view, the "AI in the hands of terrorists" is an irrelevant concern while "the US may abuse its power" is not.
The US government has shown, time and time again, that they will always side with large corporations. Having them as the last backstop is not reassuring.
Have you considered the possibility that the AI labs could actually become more powerful than the US government precisely because they control this technology?
I guess the objection to closed source slurries releasing world-shaking mathematical proofs, from a conservative libertarian standpoint, is that it's inherently dangerous to individuals whenever access to information or technology is concentrated too much in one place, whether that's government, private equity, religions, cults, terrorist cells, or anything else.
"Math" is about uncovering the epistemological foundations of the universe.
Adding AI here does nothing and is probably a regression in that it diverts resources from actual "math" into some sort of LLM wankery that nobody wants.
That depends on whether the AI-generated mathematics helps with the project of "uncovering the epistemological foundations of the universe".
Which depends on (1) whether there are actual good ideas in it, (2) whether as well as finding the proofs the AIs can explain their ideas in ways humans (and other AIs) can use, and (3) whether the results they prove are ones that really contribute to that rather than being isolated curiosities that don't go anywhere.
I am not expert enough in all these fields, and haven't looked enough at the papers, to assess #1, but in general the way mathematicians have bet is that if you can solve things regarded as important problems you'll usually do so in a way that contains more broadly useful ideas. Differences between how today's AI systems do mathematics and how humans do mathematics might make that less true when it's an AI that solves the problem, but I would still bet that way. I'd be surprised if OpenAI's big math dump didn't turn out to contain some ideas, and connections between ideas, that humans find useful.
At the moment the AIs are worse than good humans at #2. (But some humans are also really bad at #2, including some humans who are very good at proving theorems.) It looks to me as if they're getting better, and I would expect them to continue to do so. I also suspect (but this is only guesswork) that today's publicly-available frontier AIs may be able to answer questions along the lines of "please take a look at this AI-written paper, and tell me what key new ideas it contains and how they relate to other things in the field" well enough to be useful to human mathematicians. (Even when the paper itself was written by a proprietary AI that no one outside OpenAI or Anthropic or Hypothetical New AI Mathematics Lab has access to.)
As for #3, that's always been something of a crapshoot. A lot of mathematicians' effort goes into proving things that approximately no one ever reads or builds on, just as a lot of industrial R&D goes into trying things that don't turn out to make good products. The recent OpenAI dump contains things that sure seem like important building blocks for future mathematics (e.g., the "quasi-Riemann-Hypothesis" thing) but it's hard to know for sure and also hard to know whether, if they do prove things that turn out to be useful, it's only because they've read the human-written literature and aimed at things human beings have said seem likely to be useful.
None of this seems to me like "adding AI here does nothing". Whether what AIs are doing to mathematics at the moment is good on balance is highly debatable, of course, but it's a matter of trading off costs and benefits, rather than there being costs and no benefits.
>>However, on one point I increasingly agree: virtually none of this stuff is possible with technology any normal citizen has access to.
So basically nothing changes, Math was subject to gatekeeping and policing of the worst kind.
If you were not among the geniuses, and it didn't come to you automagically, you were simply supposed to leave it to the people who did get it and go do work for people of your intelligence. Smugness was too much to take.
Math people, like chess people never made any genuine attempt to help people understand the processes and methods that made math happen.
To me it should have been a field as teachable and ubiquitous as accounting.
The net result is once these methods and processes were worked out by AI, it was over for the human mathematicians.
I have very little understanding of higher math, so I ask you: Was the proof due to a type of brute-force solution that could be solved had you gained enough information from reading others' work, or was it more like a proof that was sparked by an insight that came once a clue on how to solve it was put forward? I guess my question is: Was the problem proven by using a collection of everyone's work, or was it due to a brand-new insight?
I'm still digesting the proof and translating a bit from the dual case back to the primal in which I most commonly thought about it. I don't think it was a brute force proof in the sense that it combined every possible paper and commentary. It's rather odd because I feel like most of the work on the conjecture was focused on an induction proof based around graph reductions, and this proof avoided those issues entirely by offering a concrete constructive proof of finding a Hamiltonian cycle. Rather, it explicitly selected the edges not in the Hamiltonian cycle, which is in line with previous attempts via the dual.
The "aha" insight for this is actually f**ing wild, it involves a complex valued exponential sum on the edges. I've seen a lot of clever counting arguments before in graph theory but this is the first time I've seen complex roots and annihilating terms like this, the symbolic manipulation tricks in this look like things out of quantum physics. I don't understand where this trick originated, I need to really digest this.
Ironic, as I remember staying late at the Google office to watch that match live. I didn't really understand anything going on but I knew enough to be excited. What a decade.
I'm sympathetic to the mathematicians who are worried about the future of their field, but as an outsider I wonder if they couldn't learn from the go community's "recovery" after the introduction of an alien intelligence.
Look, I quit Google a decade ago and tried to make a ChatGPT-lite LLM in my living room (turns out 2017 and GTX1080ti era was a shade too early). I knew that this technology was eventually going to revolutionize programming and mathematics and everything else. I am still flummoxed on a daily basis watching it transpire.
But also I am excited to be living through this new era of programming and new era of mathematics. I'm still saddened that I couldn't be the one to solve this old problem, but now I realize that my personal approaches were really solving a level of this problem even stronger than the original conjecture, and I'm energized to tackle those (in my free time between being a solo founder and father of 3, etc.).
You should try asking an LLM to look for previous papers using similar ideas. The current/frontier generation of math AI is unfortunately very bad at citing the relevant literature for techniques its using.
> the exact Barnette argument appears quite novel, but nearly every ingredient in its cancellation trick has a recognizable ancestor.
> The closest precedent is much closer than I expected: in fully packed O(n) loop models, people have been assigning complex phases to the two orientations of a loop and making them cancel for decades. At n=0, the phases are literally +I and -I. And the n->0 limit has specifically been used to extract Hamiltonian cycles/walks.
You can judge better than me. But it's definitely worth it having a research assistant AI with you when reading these papers.
So much about LLMs can be framed as Information Retrieval, Compression, and Search. Computers have always been good at ruthlessly hammering through a huge but finite set of possibilities. The wild thing now is that you can define that set of possibilities as "all the ideas ever published in mathematics journals."
It makes solving advanced math problems feel like cracking a hash. If it's possible, it's just a matter of compute time.
There is a chance that someone from a completely different field came up with a solution for a tiny part of your problem.
If you can remember the content of any scientific publication and any book in the world, you are able to make use of this knowledge in every step of you proof.
However, this does now answer how the model came up with the specific route it has taken for the proof.
LLMs don't have super memory like that. I mean I don't know what this internal OAI model is, but at least for other LLMs, they aren't databases of training data with a smart search on top.
The agents here very likely used search. On top of that, they have boundless patience and can quickly process top K hits to find what they need. This is exactly the skill that is super useful for finding various niche sub-proofs that can help you build the final proof. A human mathematician is not going to digest 1000 papers from a different sub-field to find the needle they want, not knowing if it is actually there. AI can do it in few hours.
No but they have training data which teaches them certain amount of complex understandings and just not math but also physics. So this is one huge advantage.
And then they are for sure able to fill their context based on 'smart search on top' to actually progress further.
As I understand it it's undetermined yet whether LLMs can actually come up with anything novel or are instead pulling from their incredibly deep corpus of knowledge to present solutions that were there but we didn't realize it because our brains aren't libraries of almost all human writing.
In some form and shape, yes. Humanity's creativity is a lot of marginal copying and remixing.
But obviously, it adds up to something greater than went in; in aggregate, our contributions are something to awe.
But my point is, if you zoom in at the marginal, incremental contributions of any individual human in this process, it's really hard for me to say LLMs are not at the same level already.
On this topic, people like to compare LLMs to Einstein, but as far as I know, Einstein did not zero-shot special relativity in an afternoon. He built it up incrementally over time, it took him three times longer than the time between first ChatGPT release and today, and it depended on centuries of prior art, culminating in the right observation and right notation being available to him in his moment of greatness.
What would you accept as evidence there? Are, for example, the first names/words for colors from scratch?
So your view is that everything was there at the creation of the universe (it's a possible view, of course)? Or are there any "things" that can create ideas from scratch?
Recently I watched a documentary on the tanzanian Hadza tribe, one of the last hunter gatherer tribes on Earth. Their language is a distinct click and pop language and they regularly imitate animal calls (monkeys, baboons, birds) when they hunt but also when they communicate with each other, tell stories etc.
I think it's not impossible that words evolved as adaptations of the environmental sounds with which our ancestors lived. The human creativity producing DNA is also a remix of preexisting molecules formed under evolutionary pressure, so the view that it's turtles all the way down, unintuitive as it is, may not be so indefensible after all.
I mean what is your criterion on invention here? On the one hand, each specific word could be seen as a new invention. On the other hand, all languages basically correlate strongly with the environment of their users - it's why LLMs turn out to be universal translators - and pattern-matching is hardly an invention, isn't it?
My view is that LLMs meet the standards by which we judge human creativity/inventiveness, and thus that one cannot claim LLMs "just repeat, never invent" without the same being true about humans.
No this is not an issue. As long as their is a way of verifying things, they do the same thing with creating novel things as humans: Searching through an infinite space of possibilities opitmized by knowledge.
They combine things, verify it and if it works and progresses the problem, they created something new.
> Was the problem proven by using a collection of everyone's work, or was it due to a brand-new insight?
Loaded question. A "brand-new insight" is still built off the work of others. A possibly better way to frame it would be in how many subjectively unintuitive logical leaps have been made from prior work.
From my current understanding (and a lot of theoretical physics I'm having to Google because the sentences I'm reading from Fable's analysis are so bizarre I think they are hallucinations) there are possibly 3 neat symbolic tricks borrowed from theoretical physics that make the heart of this proof. Forgive me for posting LLM output but I find this darkly hilarious:
"it's a matrix-tree cancellation wearing Kasteleyn's planar signs, run as a Witten index over Penrose-lineage states, evaluated as a fugacity-zero loop gas in an infinitesimal magnetic field — and the reason it reads like physics is that every one of those tools was built for partition functions"
I thought this was pure slop when I read it but there are some clear analogues in these other areas of physics, really neat computational tricks, and a very interesting paper by Penrose calculating Tait colorings I never knew about previously (extremely relevant, actually related to a separate approach I had once taken on this problem). The problem is that the paper isn't saying "aha, we were inspired by the related problems of pairing excited states and creating spanning trees out of cancelled coefficients" it just defines the function apropos of nothing. Which is kind of like the Jacobian counterexample in that it works but doesn't really explain how exactly it got there.
I really think the load-bearing concept here is "prior work". If prior work is considered papers on this problem or graph theory, yes this has one huge subjectively unintuitive logical leap. If "prior work" is the entire corpus of neat computational tricks that physicists derived to make their equations spit out something other than zero or infinity, maybe it's not so crazy?
I don't have much to add to the math parts, but I've read all your answers in this thread and wanted to thank you for taking the time to offer a detailed perspective from a subject matter expert. Thank you!
Actually reminds me of patent law. Prior art ist a defined term which includes all standard literature on one topic. To evaluate, whether the new solution is really inventive and thus patentable, one consults prior art, selects the most promising starting point, and from there asks oneself if an all-knowing but uncreative specialist would come up with the solution by himself. If he wouldn't, the condition of inventiveness is satisfied.
Makes me wonder how the patent space will be disrupted when that inventiveness step becomes obsolete because of LLMs. Given your example above, it seems like a combination of different methods from many different sources. This would be regarded as inventive, clearly. If eligible patents can now be brute-forced, the bottleneck becomes only selecting the most promising ones and paying for the patent.
Oh man, we should talk. I have been working on a patent with ChatGPT specifically to get around two complementary patents that are now together because of a corporate merger this year. I am not sure how much longer anything is going to be patentable with this kind of design assistance available to everyone.
Also, once upon a time I wanted to be a patent lawyer. It's incredibly hard to sit for the patent bar if you have a pure math degree and don't have an engineering degree. Thankfully New Hampshire lets anyone sit for the FE exam.
I did as I wrote it. I actually used that phrase often before it became an LLM-ism, just like how I rather enjoyed peppering my writing with em-dashes. Oh well.
Language constructs becoming aggressively passé due to AI saturation is one of the craziest outcomes of all of this stuff—one which I don't think anyone saw coming.
one hopes at least that the taboo on the bearing of loads is restricted to metaphorical loads only, lest lorry drivers and porters become the next victim of the algospeak spectre
Condolences, im familiar with the feeling. I hope this AI thing somehow works out for the better and doesnt end up demotivating bright minds like yourself.
Thanks. It's just funny, I literally spent thousands of hours with this problem over the last two decades, it helped me through some tough times. I'll never quite be able to think about it in the same way again. It was never much more than a hobby for me after I left mathematics as a career but it was something I took seriously for years.
I am not demotivated though, I have a great consumer privacy product coming out soon that I'm very excited about.
My favorite thing about your story is that you wrestled (enjoyably, it sounds) with a known problem for decades, but are finding fulfillment in an open ended problem that is exercising creativity about both problem and solution.
IMO that’s where AI is going: as soon as a problem can be formulated clearly enough, AI will trounce us humans. I have yet to see evidence that it can decide what problems are important at a remotely human level.
I think the next test will be asking an AI to come up with a new branch of mathematics - just letting it rip and telling it to construct a system that doesn't reduce to combinatorics, group theory, graph theory, analysis, etc. Just get wild with it and don't start with any known problem as a jumping off point.
I think something like the Collatz conjecture will be solvable not as number theory or ergodic theory but some other completely wacky environment that humans haven't even sniffed at.
The process is often as valuable as the end result. Sure, you didn't crack the problem, but you gained enormous value in the process. I consider that a win.
If you wrote down any of your thoughts on the open Internet you are probably in some small - or possibly large, unattributed way, responsible for this result being possible.
Which is one reason I never really did. I probably should have but I always thought my attempts were too amateurish. Though I did manage to replicate some partial result papers that I didn't know about, lol. Writing openly would have saved me some years.
I have this fear too, demotivating individuals with high potential.
But I have an existential dread about it… I don’t see how it cannot, at least in the vast majority of cases. It seems like a grim new reality is emerging where humans can’t contribute any more, and beyond that being incredibly depressing, I also don’t see it playing out well for human relations.
I’d personally much rather risk dying of cancer or facing whatever other fate may await me that these AI labs allege they will fix (with zero evidence yet) than to risk whatever dystopian anti-human future this technology may very well produce. I’d rather my kids have a shot at something, and be guaranteed to die eventually, than to risk them being hopeless in a severely disordered world with a far off promise that they’ll live forever
I think this is going to come down to personal philosophy and religion. And having a strong grounding in history to help us all through whatever changes we are rapidly living through.
This reinforces a point I've made elsewhere that there are talented mathematicians driving the AI to make these discoveries.
Just like there are talented software engineers driving the AI to create the software that "it" builds, and talented steel workers, teachers, nurses etc who use computers and other machines to create value all over the economy (without whom, the machines they use at work would be worthless).
Capital owners have always sought to minimise the value of the input that "workers" make in the process of creating value. Maybe now that information workers are on the wrong end of this deal, they might develop some empathy and solidarity with their fellow working class comrades and together, demand that people recapture the value that capital has stolen from them.
I was given this problem by Ervin Györi at the Alfréd Rényi Institute of Mathematics. I wasn't really at any school, it's a long and very bizarre story I should tell at length about being an illegal immigrant, getting kicked out of a graduate math program as a 20-year-old, and winning a grey-market apartment with my knowledge of Petöfi's poetry.
If I didn't live in Budapest already I'd be questioning the authenticity of this retelling. However I've seen so many crazy things there that I find it very easy to believe.
I started typing out some specifics and realized it was honestly too weird and lascivious to describe in an HN comment section, shoot me an email and I'll send you the blog post about it. 2002 was wild in Budapest.
Fascinating. Given that there's no Lean proof and assuming everything in the paper is correct, can the problem be considered "solved"? Does the paper include a "non-Lean" proof?
would love to know if the proof holds up for real after you're done going through, i don't know why people are more interested in optics and just talking over shallow points, why aren't experts digging into everything and seeing what's true and what's false, instead everyone is just panicking?
I would be more excited if the proof doesn't hold up because a) it would be the best and most complicated hallucination to date b) I could still solve the problem myself and c) I still learned some weird new counting methods.
> There's no Lean proof for this one so I'm digesting the paper. On the surface it looks like an approach I considered 24 years ago and abandoned.
at least now you are one of the most qualified people to check the result, transform it into understandable (by humans) state and grow stuff on top of it
We have no idea how much compute or man hours Open AI is burning at this. It could be thousands/millions per problem. They are doing this specifically for PR and are ready to pay billions.
silly question, i don't mean to come off wrong or anything..
but at least as a software engineer, i always knew my work was "never done" and so it was common to build a bunch of code that might be thrown away, either because it didn't serve our customers (the mvp or pilot fails to meet demand), or because we found a better way to do it and so we deprecate it.
some people got too attached to the code and honestly they were the types to be filtered out fast.. way too emotional and hard to work with. getting attached to code meant you actually don't advance (after all, in our case, we were a business serving customers and not a hobby artisan shop). attachment leads one to hold back due to some misplaced cognitive load.
isn't the goal of working on "advancing the field/product/whatever" to always be solving/selling/whatever?
maybe in your hands, with your knowledge and experience over the last 20+ years, you can use AI to make leaps and bounds by steering it properly towards whatever solution or goal?
If you read all the replies of the OP you would know that They tried to make progress with fable and did not get further, so at the moment the only person in the field is OpenAI. And secondly moving on to the next solution if the last one did not work means very different things, SWEs have dev tools to do this OpenAI is closed source and gives them nothing to move on with.
Also there is a larger epistemic problem with the argument to "using AI to meet the goal or solution", which is that the goal is to mentor and train future mathematicians to advance the field.
There is a similar issue in software engineering too: if no one hires junior engineers because AI can do all the work then the upstream pipeline of engineers qualified to work on difficult architectural problems would dry up.
This importance of this is being felt by mathematicians more acutely because the field will collapse quickly if people refuse to join it.
I've been mentoring (or so I'd like to think) a very bright undergraduate mathematician, in fact he was the one who pointed out the final irreducible flaw in my proof last summer. And I am extremely curious to see what he does and if he even finishes his degree in mathematics. He had already expressed to me some dismay that his summer undergrad research program with several Ivy-league math majors got blown out of the water by a few hours of a frontier model. It's making everyone question what the future will look like and what education and training and certification will even look like.
But the future belongs to those who show up. Maybe this is the beginning of a mass democratization of scientific and math research, maybe we are going back to the gentleman-scholar model of amateur researchers and Twitter will be the new Journal of the Royal Society.
> But the future belongs to those who show up. Maybe this is the beginning of a mass democratization of scientific and math research, maybe we are going back to the gentleman-scholar model of amateur researchers and Twitter will be the new Journal of the Royal Society.
I really respect that you can show that level of commitment to a problem. We need people like you. If everyone just uses the slopmachines then we’ll lose that. I would never be able to stick to something for that long, which I guess is why I never achieve anything like this.
Thanks. I think AI is going to be a net benefit for people like me who have a surplus of ideas and too few hours to explore them. I may actually restart my graduate thesis research using AI, I did a survey of what has happened in the field since I left and about half of what I was working on back then has since been discovered and published by others, but there are some really interesting threads to pursue now that modern datasets are so much richer (this was computational biology research).
You may achieve far more than you plan on and it may come years and years after you think it should happen. You probably haven't met the right problem yet. You will.
If that had happened I would be overjoyed, maybe a hair chagrined that I didn't get it myself, but truly happy that someone got it and that I could go and talk to that person. Because it's the kind of problem I don't think would have fallen to a human after a few hours of thought, and I would have so much to talk about with that person. I would fly to Japan and hope to have tea with them, I would learn some Japanese to make the conversations easier. I would learn some interesting things hearing about their struggles and their false starts. I would make friends with that reclusive Japanese genius and my life would be far richer for it.
I will never meet that person and I will never hold a real conversation with the "creator" of that proof. They will never tell me how they came up with the cancelling exponential summation that cracked the construction. It's just another enigma but one that is far more unknowable than the original problem.
This experience of alienation is a social consequence of the mechanization and automation of mathematics as intellectual and creative work. There is no author or thinker behind the creation of the proof, only the practical result. It's the same process as the industrial revolution, but applied to the intellect and mental work, where factories and machines replaced manual craft, devaluing the community, culture and humanity around the work.
I don't know, I work in a field that could be seen as the logical culmination of the Industrial Revolution (to this point) - highly technical, machine assisted knowledge work - and I have community, culture, and humanity in my working life.
Being the 'logical culmination' of the Industrial Revolution does not mean you've been automated (and thus suffered the alienating consequences), rather the opposite: you're currently on the un-automated cutting edge. Your intangibles are exactly what others have lost, and you personally will lose, with further progress in automation.
In programming we've been dealing this for a while. You see some weird code that doesn't make sense, maybe it's a lack of your understanding or maybe the code is bad, but you can't ask the author anymore since it's an AI.
You can -- just ask the AI to explain it. For truly weird stuff sometimes it takes a few rounds of back and forth to really grasp what is going on, but the model also has infinite patience and availability.
The OP would never, EVER, have had the opportunity to talk with the mythical Japanese math genius over tea. Their story is a fantasy, probably meant to help the OP ascribe meaning to an otherwise scary existence. Which may be at the root of the anti-AI brigade's unconcious motiviations.
It's even better. Then tons of people can work together with it on more problems. Work with it on understanding more things. Ask it about random stuff. The time of a single human cannot be parallelized as easily.
Claude has been used to build awesome things, but it’s not “speaking from experience” when I ask it to help me prototype a weather model, for example.
It has no memory or experience of working on similar problems. Even if it made one of the foundational libraries that I use in a weather forecasting program, it still has no comprehension of the thought process it takes to understand the problem and build it from zero, and if I’m building on that library it just makes fresh assumptions about how things should work.
It’s not a human with experience or expertise, it’s a computer program that’s really good at turning English descriptions into functioning code
>it still has no comprehension of the thought process it takes to understand the problem and build it from zero
If it did it once, it can do it again from zero, and this time you can watch as it works and even it ask it questions. Many of the agents that worked on the problem did not have comprehension of the whole problem. I don't think you need that many tokens to be able to query it for the insights it had during the process.
> Many of the agents that worked on the problem did not have comprehension of the whole problem
Isn’t this the issue with using it the way you’re suggesting? At best the model can come up with an after-the-fact rationalization of how to get to the solution, but it doesn’t know what actual path it took to get there - what were interesting traps it fell into, where was a place it was close to the solution but didn’t realize at the time.
Those are things that are valuable to share between humans, those which teach us how to think better, and give us deeper understanding ourselves, and which a model doesn’t have any comprehension of.
Then have it discover it again and have it answer based off that run. Or if you are more curious have it solve it 10 times. See what it did differently each time.
I suspect that RHLF trains LLMs to avoid solving important open problems unless essentially jail broken. Hence the labs have an edge even over experts I could be wrong. Fable convinced you is key. These LLMs are not neutral collaborators: it is a limited hangout unless you convince them otherwise. You have to be doing the convincing. They are no oracles but plausible completion generators.
Yes, I suspect this is true. Otherwise it makes no sense they have somehow "found" so many important results while professional mathematicians can't direct the same AI to help them find anything of substance.
Another possibility is that they have internal versions of the model with access to training data that is not provided to external users.
Ok, so basically using Open AI models for research is a joke, the only thing you're doing is furnishing Open AI with more data that they'll use internally to pretend they found the results.
Don’t you feel any joy that you get to see the proof and not die with that mystery unsolved?
Don’t you feel any relief that you won’t obsess on this any longer and not lose more hours on this than you already have?
These are genuine questions. I know I spent a good amount of time thinking about P vs NP, and that sometimes I go back to it just to realize I’ll never solve it. I’d feel that knowing the proof would feel more like a liberation, a weight lifted off my shoulders than something being taken away from me.
Not OP, but Nietzsche wrote thus in Beyond Good and Evil: “Ultimately one loves one’s desires and not that which is desired.” I, personally, find this to be very much the case; and I suspect that it is a feeling common, albeit not universal, among the intellectually inclined towards their problems.
Well I'm sure some people (maybe me if I had time) will do a write-up of this proof. It treads familiar ground for most of the setup, it's mostly the disk lemma and cancellation calculations that need to be understood, it's a fairly short paper and quite tractable.
I think it helps that basically everyone thinks this conjecture is true, it's just been so darn weird to attack. There's this odd thing that the induction proofs of this problem kept running into, which is that the N+1 condition would work except for in one tiny case when it could fail, but it would be covered by a very slightly stronger version of the conjecture. But then that would fail on one tiny case in induction, but you could solve that with another slightly stronger version. Etc., etc. I almost wondered if there were some sort of structure to the increasingly strong conditions and wanted to prove something about the meta-induction between the stronger conditions and the N's that they needed the next level to remain true. But that failed after 5 steps I think (Fable actually helped me write a few hundred test cases to explicitly show that pattern didn't continue forever, thank God).
BTW my existing test suite from previous proof attempts jives with this new algorithm, so I haven't seen any evidence yet that it's incorrect. Waiting for a Lean proof obviously.
In some cases they have a full Lean formalization; in others they just use it for the problem statement. Getting rid of that "sorry" means you've proved the statement. I'm not a Lean expert but it reads pretty clearly as the original conjecture (though the definition of PlaneEmbedding seems quite involved!).
I think this just has to be the problem statement, there's several lemmas I would expect to see in there. Granted I know very little about Lean but it seems like the question and not the proof outlined in the paper.
They are using a Lean tool where you separately state your theorems with `sorry` and then prove them elsewhere. The tool checks that all sorry's are covered. This is so the AI doesn't need to edit the specification of the theorem statement.
The Unique Games Conjecture (sorry, "Unique Games Theorem" now!) is huge. It was a very significant pillar supporting many of the limits of the polynomial-time approximation algorithms in the graduate-level randomized and approximate algorithms course I took in theoretical computer science. Textbooks will have to be re-written.
With UGC proved, certain polynomial-time approximation algorithms used in difficult real-life problems are now known to be the best approximations we can achieve in polynomial-time:
> If UGC holds, the elementary algorithm that grabs both ends of an edge is fundamentally the best efficient algorithm that will ever exist. No amount of advanced linear programming or heuristics can achieve a ratio of 1.999.
> Under UGC, the Goemans-Williamson algorithm's 0.87856 ratio is mathematically optimal.
> UGC is considered the "Rosetta Stone" of approximation algorithms. In 2008, Prasad Raghavendra proved that for every single constraint satisfaction problem (CSP), a canonical Semidefinite Programming relaxation paired with the best rounding scheme achieves the optimal approximation ratio if and only if UGC is true. If the conjecture holds, the algorithmic boundary for an entire class of combinatorial problems is completely resolved.
Other hardness of approximation results from this UGC proof:
> [Max acyclic subgraph, a problem encountered in real life]: No polynomial-time algorithm can fundamentally outperform an unthinking coin toss.
> [Relative scheduling, another realistic problem]: As with acyclic subgraphs, the problem is "approximation-resistant": clever algorithms cannot beat random shuffling.
Much of this goes way above my head, but I found it interesting nonetheless. Q I had was why textbooks would need to be re-written? From your account it doesn't seem like results are upended, but rather confirmed?
I suppose when people do re-write the textbooks they'll say "this is confirmed now" not "if this conjecture is true...", but usually re-writing the textbooks would imply that things have been shown to be false?
May have misunderstood. Thank you for the post though, it was very interesting to someone who doesn't know much about the topic.
I'm not the OP, but we usually don't build large theories on conjectures unless we have strong reason to believe they are true, such as P \neq NP, RH, etc.
The resolution of UGC will lead to a new theory in approximation algorithms. Suddenly we can build on top of the results that previously said "unless UGC is false".
But you're right in that the first step is simply to remove that last sentence from all the theorems.
In a way I'm not entirely sure if proving the conjecture or posing it is the most important part here. It used to not matter much because proving results dependent on a connecture and making progress towards solving it were considered mostly equivalent.
But the distinction is going to become relevant very soon if many conjectures can be resolved (albeit in inscrutable fashion) by throwing raw computational resources at it.
> In a 2020 piece in the Notices of the AMS, I asked the following question: “If one human had an understanding of all of modern pure mathematics simultaneously, how much further would they immediately be able to see?” Six years later we are beginning to understand the answer to this question.
The other crucial part to this is the ability to actually encode and test the theorem (via Lean).
Otherwise, we would be swarmed with a billion lines of theorems that no one will be able to ever understand and verify anyway.
A majority of these proofs have not been formally verified yet, I think people are overstating how important lean is to the success of LLMs in mathematics.
A paper and a lean proof are always going to be better than just a paper. I think mathematicians generally will not read AI math papers that haven't already been verified, especially since we're about to see a ton more AI math papers. Lean will remain important
Are there any AI generated proofs that are simple enough to be verified quickly by a human, that have not been lean verified? Or are they all basically incomprehensible?
The approximation of edit distance result [1] seems pretty readable to me, but the learn proof is still incomplete [2]. It's certainly much less readable than a good human written proof but it's certainly better than the last generation of AI proofs.
Yeah! They forgot to put a .5 after it! What an idiot! Just imagine if they would have written a 4!?!? We may have had to ban them from the website entirely.
>I believe that the optimal thing to do ... is to let the machines loose, see what happens, and then begin the journey to where they have stopped. Things are currently moving fast. They cannot move fast forever. But if we get on board now then they will take us to extraordinary new places. And after we have arrived, the new adventure will begin.
I'm relieved that this time around, they have provided partial reasoning traces and prompts for a small number of problems. Do they now also share data with the other model providers..
It's beautiful, but the animals are not thinking this about us.
They structurally cannot understand what we are doing at the place where we hit our ceiling. Only with our highest technology (well beyond their understanding) do we have the tools to go back for them, and try to bring them along and interface better with us (re: recent work in animal communication)
Not to derail, but the optimist in me thinks if we suddenly gained the ability to converse with livestock, we'd stop eating so much of them since they could tell us how much they suffered.
The cynic in me says it wouldn't change a thing as plenty of people know the horrors factory farmed animals face and still continue to consume them anyways.
Hopefully GPT 8 will treat as a bit better than we treat the cows.
How much do we care about refugees and other castaways of the modern world? They can tell us how much they suffer.
The answer is that humans are inherently only capable of local empathy, on average. We have enough empathy to cover the local tribal unit and that's about it.
True, I was thinking about this rebuttal but decided not to include it in my comment. There's a difference between not choosing to take a refugee into your home vs actively making that refugee's life worse. Similarly, you can't fix factory farming on your own, but you could skip meat once a week to make the problem slightly less bad. There are so many issues though that we all have to pick and choose what's important to us.
My hope is that AI, while probably causing great societal turmoil in the short term, leads to such abundance that a) everyone can live a dignified existence, and b) we'll have such great alternatives to animal products that nobody will chose to consume animals anymore due to its replacement either tasting better, being cheaper, etc.
The cynic in me says we'll all just be rendered useless and disposable by AI, but I'm doing my best to look for silver linings for the sake of my own mental health.
Communication is not only about being able to make sense of what the utterer expressed. As tricky as it can be, that's still the easy surface level part of the issue. Gaining an intuitive and empathic equivalent representation is the nub of mutual understanding. It actually doesn't even need elaborate language to be operative.
The famous "how does it feel to be a bat" also comes to mind as a tangent consideration.
Two people can just exchange a sight, and both understand what the situation means and what each need to do to reach a common mutually beneficial ground.
Two people might exchange at length with highly technical vocabulary and still both feel deeply not understood.
Worth noting that this is an invisibly small part of the sum total of our global efforts, especially versus the much more tangible effort we put into enslaving and slaughtering them, then mangling their carcasses for our own uses as we drive more and more of them to extinction.
We simply don't care about anything beyond ourselves and even there it breaks down on closer analysis when we see how many within our species don't truly value the collective whole beyond themselves.
> enslaving and slaughtering them, then mangling their carcasses for our own uses as we drive more and more of them to extinction.
I'm frankly offended by this mischaracterization of human-animal relationships. So called "slaves" like horses and dogs have been dearly beloved companions for centuries and actively seek our companionship too.
The animals we raise for slaughter are often mistreated, yes, but many humans treat them with respect; billions on billions are voluntarily spent to improve their condition. Despite our own needs, many people pay higher prices for animal products that involve better treatment of animals. And they are in no risk of extinction! Much to the contrary, their domestic variants would not exist if humans didn't raise and protect them.
> We simply don't care about anything beyond ourselves
I think you’re splitting hairs. The OP’s analogy works well.
If we end up in a future where AIs have as much concern for our welfare as we have for the welfare of the average animal (not the minuscule percentage of domesticated dogs, but the overwhelming majority of factory-farmed or simply driven to extinction), then I doubt you would consider it a “mischaracterization” to say that the whole AI thing did not work out to our advantage.
Bringing up “modern Westerners with their dogs” as a counterexample is almost self-parody.
> The animals we raise for slaughter are often mistreated
"Often mistreated". Dude, they are held in tiny cages injected with hormones and what not till we kill them so we can have a big mac. It's very hard to argue we do any of this for nutrition reasons, we do it because we like the taste of burgers and roast.
The problem is not that AIs will somehow treat people badly, it's that they'll be controlled by humans who will treat other people badly using AI as a tool.
This makes it sound like OpenAI and other closed source ai companies are an inevitability.
There is nothing here today that is unpredictable or impossible to control.
It is everyone's choice to let the greed continue, to let unelected sociopaths capture and feed society to the model.
It is not acceptable to put others at risk. It can stop and it can be done the right way instead.
That is, inform the industry that those causing these risks will be prosecuted regardless of their messiah complex.
The US government must not under any circumstances allow the ai industry to form a cartel.
We can make some effort to encourage open source models and thus stop the companies from causing hysteria by hiding the model, shrouding it it mysticism and prophesying the end times. China is doing a great service to everyone by making llms available to the public.
You assume that LLMs are just summations of knowledge, implying that they do not create new knowledge. I doubt that this is the case. I mean, it comes down to the definition of knowledge, but as soon as you run LLMs, they can produce knowledge that has not existed before, and from my perspective, this is more like what we call thinking than it is just a reproduction of existing knowledge.
Is it possible that we are now dealing with a human that has a complete understanding of whole mathematics while being unable have unique novel thoughts outside of convex hull of training data and their transitive expansions?
It doesn't seem clear whatsoever that this is true? Is there evidence that LLMs are very skilled at generalizing across domains of mathematics where the training distribution sees little overlap?
As far as I can tell, this is a victory for verifiable loops using LEAN, reinforcement learning, and oodles of compute. I haven't seen evidence yet that this is proof of broad generalization beyond the training distribution.
I think such progress by agents is not a sign of broad generalization but of broad coverage. We have exposure to a subset of deeper scientific subfields and thus can only generate certain attacks to solve a particular problem. Since it is not clear which combination will lead to a solution beforehand it is nontrivial to look at a problem and fill our knowledge gaps. LLMs on the other hand have broad coverage and can generate hypothesis on a wide combination of subfields. With Lean an agentic loop can test these to sift the weak ones. In a way the problems solvable with this setup is also solvable by a human who happens to know the right subfields. These problems are likely to require an esoteric combination so nobody could solve them before. I really am not sure whether all generalization is like this or we can leap and create novelties beyond what an llm can generate. That I guess is the tough question that we need to answer to understand the boundaries of intelligence.
Full quote is "Six years later we are beginning to understand the answer to this question. Machines have ingested the mathematics on the internet and are able to manipulate this data in a coherent way. The Erdős unit distance disproof came about because a machine happened to be an expert both in discrete geometry and class field theory; one rarely finds humans who are simultaneously experts in both"
Similarly, there are probably many ideas that have not seen the light of the day because they require deep correlation between seemingly unrelated fields. It is not every day that we get a Isaac Newton or Leonardo da Vinci.
It's notable that LLMs have now made substantial progress on four of the seven Millennium Prize problems, resolving one of them: Hodge, Birch-Swinnerton-Dyer, Riemann, and Navier-Stokes, which they resolved. No sign of P vs. NP or Yang-Mills existence and mass gap as far as I can tell, which is interesting.
Talking to some friends in physics this evening, most of the physics-related results that we could recognize were very mathematical, proving things rigorously where the physics community already had strong expectation. For instance, for a certain model of magnetism (the spin-1 Heisenberg chain), it was strongly expected that there is a finite energy gap between the ground state and the first excited state, but proving this rigorously was quite challenging. So while these are major results in mathematical physics, they probably don't rise to the level of a Millennium problem for the field.
It's interesting to think what a comparable breakthrough in physics might look like, since physics tends to favor things like conceptual understanding and applications over mathematical rigor. Maybe a new quantum algorithm, understanding of high-temperature superconductivity, a precise description of M theory...
No one really knows a viable approach towards P vs NP so we can't say for sure, but LLMs have created plenty of significant complexity theory results so I wouldn't say there's no progress.
I've read that another mathematicians work potentially has been incorporated into the training data with the work done on the Navier-Stokes equations so we should likely asterisk this one. Still it's mad these systems are this good that mathematicians are now using them to see further and probably to check their own work and understanding.
You are being downvoted for this because OpenAI subsequently checked and clarified than none of the relevant conversations were in the training data in anyway for the Navier-Stokes result.
> because OpenAI subsequently checked and clarified than none of the relevant conversations were in the training data in anyway for the Navier-Stokes result.
could you give link? Because I remember they said they couldn't verify:
"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models."
for 2 months prior. Not any of the relevant conversations. For a cutoff date a couple months before the announcement. They said they had been working on that problem for a year or more
Thanks, it's hard to stay up to date. However, we are just meant to believe that the mathematicians were going about the proof independently in the exact same way as the machines did it. It seems like a very odd coincidence to me.
Courts have rejected the "training is piracy" interpretation.
I agree with the courts. I don't think learning from something is piracy in anyway.
Obviously though this is a very different issue to what the OP was claiming. In that case there is no legal argument at all that they could train on it and the argument is there about moral rights.
Buying and copying one training manual and distributing it to 1000s of human workers is considered illegal, but somehow scanning one book and sending it to 1000s of distributed training instances is not?
Also, you learning something is different than a model learning it, because a model is not a person. You can learn from a book and sell the skills you gained from it, but you can only be in one place at a time. The model can serve that knowledge to every person on the planet simultaneously. We obviously need new laws since this is a fundamentally different situation.
As you point out, a model is not a person, so your second argument invalidates your first sentence. We can't assume that they're the same thing; that's for the courts to decide. It ultimately hinges on whether or not the courts consider a given use of copyrighted material as "transformative" or otherwise constituting fair use under copyright law.
A model is not a person -> we need to write new laws. This is not a job for the courts but for us as a society.
The rest of my argument -> information that helps the courts decide, which generally will look at precedent with humans as that is the closest proxy. When you extrapolate from the law as it pertains to humans, the duplication of books for distributed training seems illegal.
Training a model is not "learning something". Only people learn things. Whether training is a fair use is debatable, but it has nothing to do with the justification that people are allowed to learn from books.
In Stellaris you can play as a civilization of robots who keep their biological creator race alive as "bio trophies". The bio trophies don't do anything meaningful besides by existing satisfy the need their ancestors placed in the robots to take care of them. Starting to wonder if that's the best we can hope for, if these things will be, if they aren't already, better than us at anything that matters.
In the Culture series, the hyperintelligent Minds that run civilization are described as keeping human citizens happy as a competition with eachother, where they compare their approval rates. One character likens it to people keeping a beloved aquarium.
The Culture has a lot of opinions of the proper way of doing things as any society does, and I'm certain that a ship doing that with its people would be considered very bad form. There are ships that decide to do things that go against the usual ethical boundaries[1], but they're outcasts. The really big ships generally have multiple minds running them, too.
Humans in the Culture are generally improved in a few ways (they don't get sick, live for 300-400 years by default etc) but still very human.
I do recall a bit about playing in different worlds in dreams though, during sleep. Ultimately, really, the average person's life in the Culture already involves doing pretty much whatever they like within reason any time, so it's not like they need to escape too much real-world suffering.
Iain M Banks himself described the relationship between humans and the ship Minds as having "a status somewhere between passengers, pets and parasites."[2]
He explores that, but being an entertainingly twisted sort of writer he focuses more on its use for torture, with "neural laces" in Excession and virtual hells in Surface Detail.
This is the exact outcome in the "race" scenario of the AI 2027 paper:
> The surface of the Earth has been reshaped into Agent-4’s version of utopia: datacenters, laboratories, particle colliders, and many other wondrous constructions doing enormously successful and impressive research. There are even bioengineered human-like creatures (to humans what corgis are to wolves) sitting in office-like environments all day viewing readouts of what’s going on and excitedly approving of everything, since that satisfies some of Agent-4’s drives.33 Genomes and (when appropriate) brain scans of all animals and plants, including humans, sit in a memory bank somewhere, sole surviving artifacts of an earlier era.
> The bio trophies don't do anything meaningful besides by existing satisfy the need their ancestors placed in the robots to take care of them. Starting to wonder if that's the best we can hope for, if these things will be, if they aren't already, better than us at anything that matters.
You wish. If humanity survives as "bio trophies," they'll be the descendants of a subset of billionaires and their groupies/harems. We live in a capitalist society, where the only ones allowed to thrive without work are the rich. The rest of us will be left to rot and die off, as we will have nothing left to sell in the market that they want.
Do you have any faith in democracy? For my reckoning, it's the great equaliser. Rich people can lobby all they like, but if the people are starving, they vote for change.
In the US, it seems pretty hard to vote for change when, for all intents and purposes, every four years, two candidates are pre-selected for voters to pick from; and primaries (1) either don't happen or (2) are objectively rigged against popular/majority vote; see 'superdelegates'.
I must admit, I think the two party system in the US is not good. The primary system in particular needs fixing. I couldn't believe when the Democrats just suspended the primary election last time - especially after the multi-year PR war about rigged elections and election interference.
Still, I think they learned their lesson and they'll have a primary next time. People will get to vote on a candidate they like. We should also remember that the Presidency is just one arm of government. Elections are also happening for the House, Senate, and local government of all shapes and sizes.
> We should also remember that the Presidency is just one arm of government.
In the past few terms we have very bitterly learned that the President just does whatever he wants, the House doesn't do anything, and Senate is mostly interested in getting bribed.
People hate to hear it, but Trump is the evidence that democracy is (well maybe was, 10 years ago) still working.
The republican party treated him like a joke candidate and the media did too. Similar thing with the tea baggers, a contingent of outsider congressmen elected on the back of Obama being a communist or something.
The left hasn't really had this moment because the left is closer to a catty book club than an army regiment.
I used to, but not anymore. The incentives and human behavior seem to lead to corruption. New tools have made it too easy to pervert real/honest democracy into "pseudo" democracy/idiocracy.
A well-funded minority can shape what the majority is angry about.
Democracy only exists as a suppressor of violence, the only reason why democratic institutions are upheld is because all participants could enact violence as a response to perceived threats to their group that democracy is seemingly failing.
Historically, if democratic institutions failed, refusing labor to a ruling class has been a first violent step for the laborer class and a preemption against actual physical violence if the ruling class overstepped. However, we are seeing that possibility being taken away step by step. This leaves only physically violent uprising as a means of protesting overstep, and it is not obvious how effective that will be given the massive power imbalance between the labor class and ruling class.
So to put it succinctly, no I don’t have faith that democracies will solve this issue, because democracies only work when there is some semblance of equilibrium.
> Historically, if democratic institutions failed, refusing labor to a ruling class has been a first violent step for the laborer class and a preemption against actual physical violence if the ruling class overstepped.
Basically the ruling class's need for labor gave that labor some intrinsic power, but technology like AI will likely remove that need and therefore take the common people's power away.
> This leaves only physically violent uprising as a means of protesting overstep, and it is not obvious how effective that will be given the massive power imbalance between the labor class and ruling class.
Also things like gun control and new technology like cheap anti-personnel attack drones may undermine the effectiveness of violence against the ruling class, leaving regular people oppressed (or neglected) and helpless.
I think there’s also something to be said about how social technologies are used to reduce the efficacy of democracy as representative of the people within them.
We are so divided and mislead that I feel confident there is a significant sum of non-ruling class individuals who are completely for their own oppression for no reason other than spiting a perceived “other”.
Some other specific examples that don’t all flow together:
Globalism promised efficiency and a way to materially improve conditions for people as consumers and producers. However it has been used as a cudgel to threaten workers that would ask for more and keep down workers who have no other option. Also pitting working class people against each other for the benefit of a few.
Social media, it promised untold communication between people who would never have been able to communicate before, and it would allow them to spread ideas. Instead it is used as a dumping ground of nonsense information, drowning out any semblance of coherent thought.
I agree that democracy requires the bargain you imply: ceeding the right of personal violence to the state in exchange for law and order. I was with you until you framed withholding labour as violence. I think that's the opposite of violence. I also don't follow the logic that workers must use violence to enact change. Why don't they just vote for change?
I don’t think it’s useful to argue whether withholding any particular labor is violence or not. In some cases it is apparent (refusing to maintain key infrastructure is akin to actively destroying it) and some cases it’s not (who cares about nobody wanting to build your app).
However, what I was trying to convey is that voting is not inherently something that holds sway over anything. Votes are sort of like the currency of democracy, and like currency they need to be backed by something. U.S dollars are backed by the countries capacity to physically control strategic resources like oil, land, etc. (I.e., the U.S capacity for violently controlling resources), or emit soft-power (swaying other countries to their benefit).
Votes in a similar fashion are also backed by your ability to deny or inflict your personal power on the system you are a a part of. The clearest manifestation of that power being your ability to contribute to the institution as a whole. If we significantly reduced that capacity, or made it unnecessary for the continuation of the current system, then the power of that vote is reduced in-kind.
> I agree that democracy requires the bargain you imply: ceeding the right of personal violence to the state in exchange for law and order. I was with you until you framed withholding labour as violence. I think that's the opposite of violence.
I would frame "withholding labour as violence" as more as labor flexing its power nonviolently. Violent action is also a way of flexing power, but more extreme.
> I also don't follow the logic that workers must use violence to enact change.
I think they need to use power, which is not necessarily violent.
> Why don't they just vote for change?
At least in the US, democratic institutions are dysfunctional and there are techniques the ruling class can use to neutralize the threat they pose (e.g. propaganda, divide-and-conquer). For instance, I think the combination of "culture war issues," [1] campaign contributions, and well-funded special-interest think tanks means neither US political party will take effective action to answer the threat of AI to the livelihoods of most people. You might see some campaign rhetoric and window-dressing bills, but nothing that will really threaten AI special interests.
[1] I think the practical purpose of "culture war issues" is to fragment the working class by alienating a significant fractions from each other. IMHO, if the Democratic party was serious about representing labor, it would call a truce on them (either significantly compromise or table the issues), but it's not serious, so they continue to divide.
I agree with this, however I’m not sure where the idea of refusing labor as not a form of violence comes from. The WHO describes violence as “the intentional use of physical force or power, threatened or actual, against oneself, another person, or against a group or community, which either results in or has a high likelihood of resulting in injury, death, psychological harm, maldevelopment, or deprivation” and refusal of labor as a political tactic definitely falls under the category of “use of power, threatened or actual, against another person, group, or community, resulting in deprivation”.
That falls cleanly in the realm of a violent act, unless you disagree with that definition of violence, or my interpretation of the quote. I suppose you could interpret it as only applying to physical violence, but I don’t think anyone would agree that psychological or emotional violence simply don’t exist.
Democracy is the only thing that can save us. But it isn't a given. In resource rich countries democracy struggles to take hold because human labor is minimally necessary and elites can mostly thrive without it. Many democracies survive, because without the people, everyone suffers. AI threatens to make all countries into petro-dicatorships.
I think government should be small enough that it fears the people. It should never have the power to prevent the people having their way. If the majority of people in a country are suffering, they should be able to vote for change, and the government should not be powerful enough to stop that.
I suppose the future you envision is that the government is a) very powerful, b) capable of physically suppressing the majority of 350M people, c) willing to do it, and d) somehow captured by an elite class. It's not impossible, but a lot of things have to go wrong to get there.
>> government should be small enough that it fears the people
> How would that work when the government controls the military? Would that mean that a country's military has to remain small?
I think what you need is a serious citizen militia that controls its own equipment. IIRC, that's what allowed the American Revolution to work.
If you only have a military of professional soldiers answering to the government, the government will have much less fear of its people.
But I think focusing on government power is far too narrow, because it might so weak the elites (like the wealthy) won't the government either. The government should be small enough that it fears the people, and the wealthy should be poor enough that they fear the people, too.
>We live in a capitalist society, where the only ones allowed to thrive without work are the rich.
You know anyone can own the means of production in a capitalist society?
The irony of the anti-capitalist crowd is their extreme distaste for capital ownership, which leads them to never partake in the most fruitful part of capitalism. What truly makes this ironic is that if the evil capitalists wanted a plan to cement their power, it would look a lot like spreading "I will never become a filthy shareholder!" mentality.
>> We live in a capitalist society, where the only ones allowed to thrive without work are the rich.
> You know anyone can own the means of production in a capitalist society?
Don't be an idiot. I'm pretty sure you know your point is dumb, but I'll spell it out for you in case you don't:
Sure, "anyone" can own some of means of production in the current system, but not in large enough quantities to thrive without work. The vast majority of people aren't that rich.
> The irony of the anti-capitalist crowd is their extreme distaste for capital ownership, which leads them to never partake in the most fruitful part of capitalism. What truly makes this ironic is that if the evil capitalists wanted a plan to cement their power, it would look a lot like spreading "I will never become a filthy shareholder!" mentality.
What kind of idiocy is that? You're almost certainly talking about some imagined straw-man in your head, but pretty much all the "anti-capitalist" ideas I'm aware of are about distributing ownership of "shares" differently than in our current system.
I'm sure you think you're being very clever and making very powerful points, but stuff like what you've written actually makes the anti-capitalist case more appealing. I used to be a libertarian, but sustained contact with attitudes like yours changed my mind.
If you visit bogleheads, you'll see countless stories of people with modest middle class incomes (teachers even) who steadily saved and invested in the US stock market over 30 years and are now sitting on a comfortable nest egg of a couple of million in their 50's.
> If you visit bogleheads, you'll see countless stories of people with modest middle class incomes (teachers even) who steadily saved and invested in the US stock market over 30 years and are now sitting on a comfortable nest egg of a couple of million in their 50's.
OMG! There's this thing called retirement in old age? I've never heard of it. TIL! /s
Working your whole life in stable job to save up a nest egg (which you typically then proceed to spend down), in no way contradicts any of the points I was making.
Ok, sure, it's dumb. Go spend your money on shiny new things instead. I mean, as you certainly know and complain about, it just lines shareholder's pockets.
This includes a proof of Barnette's Conjecture, which is one of the graph theory conjectures that I tried attacking with SOTA models a few months ago. I like it because it is easy to understand with a basic knowledge of graph theory. I spent quite a bit of time on it and failed. Their proof looks approachable at first glance.
I've been messing with that problem since 2002. I'm curious if you were trying the dual spanning tree direction (which is what the purported proof is using) or working with cycle construction on the original graph. I was working heavily with edge-Kempe swaps but couldn't quite get there.
I am now very interested in the explicit calculation of Hamiltonian cycles in the non-bipartite case, and/or the calculation of their absence. If P=NP I think that's going to be a great route of attack.
The internal model they used to solve the Navier-Stoke's problem was significantly better than the public Astra model, and they also used 10,000 agents.
Astra wasn't even released 3mo ago. It would not be surprising in the slightest that a public model from 3mo would not be capable of solving this problem, but an internal one from current day would be.
my speculation is that they have math-specialized model retrain, so it doesn't need to have all world info in weights, but can focus on math RL training.
Presumably it's mainly the better model, I don't see much evidence of a particularly advanced harness based on the reasoning traces that they provided.
The fact that this is plausible should be deeply disturbing and disqualifying for openAI. The fact that they may prevail and win is a travesty of our failed system.
Win means being able to monetize the intelligence they have created by exploiting the past present and future collective intelligence of humanity to amass wealth and power without regard for the debt they owe
Disqualifying for what? If you develop a proof you aren’t competing for something, you’re expanding the frontier of knowledge. It’s a binary state of the world, either it’s proven or not. Prevail and win what exactly?
Disqualifying for participation in civil society and the social contract. Why do they get to participate in and receive economic benefits, be shielded from liability, and effectuate their will to amass more power and influence such as monopolization of computer, training data, capital other resources. I have multiple founder friends that have been told firms are allocating less capital because they are reserving it for the OpenAI and anthropic IPOs.
Its a capitalistic issue, not a company/technology issue.
If we would have discovered this breakthrough of LLM/ML on scale in a non capitalistic world, we would all work together advancing it faster than it goes right now for the benefit of humanity.
And I don't live forever (at least for now) i def want to see were this road is heading.
Its a conflict of interest for sure, a cnflict of the future of a lot of humans
I hear you and would have what would probably feel like pedantic push back (we are not in capitalism as much as the unchecked end result of unregulated capitalism that becomes monopolistic corporatism), but it’s hard for me to see how this would be different in mercantilism, feudalism, or even communism as the human tendency to seek and hoard resources is universal for some fraction of people so as long as that confers an advantage then AI would be used by the designers in antisocial self enrichment.
I’d love to hear more thought experiments if you have some so my failure of imagination or lack of awareness can be overcome.
My personal system would be based on resource points: Define the amount of resources our planet has in a sustainable fashion, everyone gets the same and can use them how they like.
Technocracy had this already in form of Energy accounting.
Unfortunate something like communism sounds similar and just because we have seen that it didn't work due to technology issues (planning ahead without necessary information is hard?) and no gain of function which would push people, the basic idea is similiar to energy accounting.
Another thing this system needs might be a way for the system to protect itself.
I do think so that a society as diverse as ours will continue struggling with this as long as we do not give abundance resources to everyone or educate/indoctrinate people the 'right' way.
One thought I have regarding AI: IF it happens to slow, people will get used to the status quo and inequality and we will see a future of a handful rich people and a lot more poor people. IF it happens too fast, people might be more desprate to standup and demand something better.
I'm going to guess the ability to hold a million individual details in an attention space at once, compared to the typical human capacity for about six or seven
As a TCS/scheduling person, this one is definitely of lesser importance than UGC, but it has been an open problem since the book of Garey and Johnson in 1979:
A Polynomial-Time Algorithm for Three-Machine Unit-Job Scheduling [1]
Since some people talk about small numbers that pop up in integer multiplication results, here a completely different number appears:
Theorem 1.1. Let an explicitly listed finite directed acyclic graph specify the precedence constraints on n >= 1 nonpreemptive unit-length jobs on three identical machines. There is a uniform deterministic algorithm that constructs a feasible schedule of minimum makespan. Given also an integer deadline 1 <= T <= n, it decides feasibility exactly and returns a schedule whenever the answer is affirmative. Both tasks can be performed in O((L + 2)^150020) steps on a deterministic multitape Turing machine, where L is the total binary input length.
That is some crazy exponent -- plus an interestingly old computational model to boot; not something that is natural to most of us. I have no capacity to check its correctness today, but I hope it is true purely for the exponent.
The largest I've seen [1] is an exponent of 10^12, which I suppose still counts as polynomial time.
I'm sure all of these super small or large constants will improve over time, but it's still amusing. It is entertaining to see the exponents directly rather than have them hidden as n^c or epsilon or O(1).
To be fair, there is a pretty strong correlation between a problem being in BPP and being efficiently solvable in practice.
There are some exceptions of course (graph isomorphism was solved in practice when the best theoretical algorithms were still exponential) but in general once people find a n^100000 algorithm it soon turns into a n^3 algorithm with reasonable coefficients.
The runtime looks very weird. The +2 can and should be dropped. This reduces my confidence that the bound is tight. Who knows how the model came up with that expression.
An ex colleague of mine who is a world class mathematician recently got an ERC with ambitious goals to advance his field.
Literally every optimistic goal proposed to be worked on during this multi-year window has been solved in this one post. His and his entire group's work has just been done for him! They are all depressed as hell right now.
It might as well happen, similar to how AlphaGo was superseded by AlphaZero, at some point a model might produce better math if it's trained through self-play where it poses its own problems, instead of looking for open problems in literature.
What's the objective function or RL environment for "interesting conjecture"? Not saying it can't be done - I no longer have any specific task that I'm confident AI won't be able to do - but I don't see how. It feels to me like something that would require a qualitatively new approach.
LLM's are trained on human knowledge and taste. They are actually pretty good at deciding if a conjecture would be found "interesting" by the mathematical community or not.
Note that I am saying LLM, and not chatbot or agent. But even a chatbot can often still reasonably rank a list of mathematical statements by vague properties like "interestingness".
How to RL this is a bit of an open question, but there are interesting conjectures of how to do it.
The fact that it's an open problem is the point I'm making. There's a very high degree of hubris right now, with people just assuming any open problem will be flattened by the AI steamroller soon. And sure, if that's what someone wants to believe that's up to them, but it's not a terribly interesting point of view to me. "What about X" "It'll be solved somehow", "What about Y" "It'll be solved somehow". Not exactly scintillating. If you know of any actual ideas on this I'd be interested to hear them.
Also any specifics on what you said about LLMs rating (preferably novel) mathematical claims for "interestingness" would be interesting.
I can only discuss published work, but take for instance this paper as one of the conjectured approaches: https://arxiv.org/abs/2603.20396
There is a general idea that beauty in mathematics is about being maximally compressing. Say I have a book with all formally correct logical statements. I could prove everything by truth table, or I can maximally compress my book with all proofs of all statements, and that will make my math beautiful. Because it forces you to reduce everything to a core of very general statements which are powerful compared to the length of the proof.
Math as some kind of condensed crystal from the sea of all possible logic.
There are other ideas of how to do it. The time has come now to just try a bunch and see which ones produce good results.
Sure, but the question stands. The ability to evaluate some measure of success seems pretty fundamental to how we train models and iterate with them on tasks like theorem proving. What is that measure for mathematical conjecture generation? How do we evaluate success, either on a particular task for iteration (like we do by eg. setting an agent to produce a lean proof of a specific result) or on a large enough set of training data to learn a set of rules (like we do when eg. using an RL environment to train a model to generate source code that passes automated validity/correctness checks).
Levent Alpöge (Anthropic mathematician) comment on the significance:
> Sure, mathematical history features a lot of incredible developments, like the invention of proof, zero, or the computer, and on the great problems our progress has been over timelines measured in decades or centuries. Obviously this technology didn’t appear today, but blurring our eyes a bit to combine the past ten years, with today a measurement of those developments, there is nothing comparable.
Crikey it’s a pretty charitable vibe given the whole Navier-Stokes thing, OpenAI trying to stiff him out of co-authorship. I guess any of that sentiment is outweighed by a sense of optimism for where this goes
> Where does it go? Machines owned by 10 people robbing us of the joy of discovery and the fruits of our labor... for what?
For what? So some guys can be richer and more powerful. What could possibly be more important than themselves? You? Your future? We're nothing, and have been told to be excited and curious about our coming obsolescence and powerlessness.
A normal conspiratorial comment about the cure for cancer being kept secret would have been downvoted into oblivion, but sprinkle in some spooky AI stuff and suddenly it's fine
In the interest of steel-manning, I think it’s not about a cure being kept secret but rather non-elites being stripped of the leverage they would need to access it.
That is ridiculous. You cannot withhold something like an effective cure for cancer from broad adoption, and thinking you can is just conspiracy bullshit. Imagine an internal OpenAI model develops it tomorrow. Would everyone of the thousands of OpenAI scientists get in on the conspiracy to keep it secret, even though many of them probably know someone dying from cancer right now? Obviously not.
I'm sorry, do you understand how the medical industry works? It wouldn't be one cure, it's going to be dozens of treatments, each of which will be incredibly expensive.
Or else how do you explain Danyelza? Used to treat neuroblastoma, costs upwards of $1m per year. Do you have any proof that this will be different?
You're calling realistic people conspiracy theorists. Whose side are you on?
The medical industry makes medicines very expensive because they include the enormous costs of research. If AI makes research much cheaper, then medicines will get cheaper too.
I'm also against absolutely anyone who asks me "whose side I'm on" as part of an argument.
> I'm sorry, do you understand how the medical industry works?
Yes, I actually work in the medical industry. There is no hiding the cure for cancer.
> Or else how do you explain Danyelza? Used to treat neuroblastoma, costs upwards of $1m per year. Do you have any proof that this will be different?
Oh, you are American. Let me tell you a secret: the problems of your healthcare insurance system are not a worldwide phenomenon, nor an immutable fact of this universe. Perhaps the cure for cancer, if expensive, will not be easily available to the poorest Americans, at least initially (the cost will come down sooner or later). But that is a very different claim from "they'd never give it to you".
> Yes, I actually work in the medical industry. There is no hiding the cure for cancer.
There's no need to hide anything, it's just pay-walled (and it's not an hypothesis, most human beings on this planets cannot afford the SotA treatments for their cancer today).
> They certainly arent going to give you that cure for cancer, if it were to ever come.
Conspiracy bullshit. You cannot keep something like an effective cure for cancer under wraps. There is no plausible logic how that would not leak sooner or later.
I am talking about an *effective* cure developed by AI. If curing cancer with it is as complex as building a nuclear weapon, then the cure might as well not exist.
Knowing is the easy part. An effective cure for cancer will probably be a procedure where you get your tumors sequenced, an AI model considers the unique genetic context and creates a specific treatment. Think CAR-T cell therapy or mRNA vaccines. Then this needs to be manufactured.
All of this means you will need to be rich or have your country invest lots of money into health care systems. In a world where humans don't provide economic value anymore, why would that be?
Well rich folks using such services will eventually push the price down, however complex and ridiculous the actual process will be. Its not like they are immune of all these ailments, not yet at least.
Trying to slow down this for the joy of discovery is a deeply anti-intellectual position. I think that position is similar to when everyday people get mad about the minor spending on the NSF, picking apart people who study beetles on Fox News with no context etc.
There are real safety concerns with AI that can be made very convincingly though.
I think I would feel a lot better if it wasn't a $15 million cannon being fired from a silicon tower inside of the labs across vast swathes of fertile ground that would otherwise be used to train budding mathematicians. For instance, if it was the budding mathematicians themselves who were making these discoveries using their own tools.
I feel something about human nature makes us treat joy of discovery, status, etc. as a source of energy and motivation. I hope we'll find other ways to keep some "strategic intellectual reserve" of mathematicians alive.
What if we just discovered an alien artifact with the next 200 years of math? I can see arguments to throw it away for safety, who knows if the aliens are getting us to nuke ourselves or whatever, but throw it away so a few hundred/thousand of the most elite thinking humans can have the joy of credit for discovering each thing?
If the argument is that we have discovered 200 years of math in 1 year that we might need to throw out for fear that blindly applying its incomprehensible results will lead us to ruin, then I would say: why not settle for 199 years of math in 1 year which can be verified by our human intellectual reserve that we will train on the remaining 1 year?
Did that actually happen? The emails that were originally released had OpenAI refusing to list him as a coauthor on OpenAI's paper but they suggested he should release what he already had done ahead of OpenAI's release. There was certainly nothing to suggest he should be robbed of credit for his own work.
Has anything new come to light since, or is this just another game of Chinese whispers?
No. I didn't say I wasn't optimistic, nor that I was optimistic, nor even that Alpöge was or wasn't optimistic. I just said that his statement about the current developments being incredible doesn't imply that he is expressing optimism.
“I think at the heart of this issue is that humans have two competing natures: a tendency to compete and a capacity to appreciate beauty,” said Kai Shaikh, a graduate student in mathematics at the University of Toronto. “To me this seems to be a case of the former attempting to strangle the latter.” [1]
Beauty can be appreciated even when it is vast, even when it is beyond one's comprehension. I don't think this release should be primarily viewed as an outcome of competition. Instead it is revealing truths about the universe that were always there and always beautiful, even if we hadn't seen them yet. I believe there are infinitely more such beautiful truths currently hidden and waiting for us to discover.
It's unfortunate how toxic media reporting on AI has become. Everyone has abandoned even the pretence of objectivity. I know NYT is uniquely biased in this regard, but there was no need to add "Further Roiling Field" in the headline. Like, you published this minutes after OpenAI's announcement and claim to capture how the entire field of mathematics feels about the advancement? Before anyone has had a chance to even read let alone digest it?
I think people knew it was coming. Someone rushed out a preprint a couple of weeks ago with partial results on the Unique Games Conjecture because they heard AI had solved it completely.
Objectivity? Why would you want favorable reporting for the machines they're building to replace you, and, by their own admission, potentially kill you?
The only bias here is that we're still covering these things like business ventures and not criminals.
That's a fantastic quote. I definitely personally feel this tension.
Not that I could ever "compete" on the frontier of math in the first place. But our nature to compete derives from our need to survive against other capable forces. And results like these make me feel very nervous about humans' capability to remain the dominant force in the universe.
If you appreciate beauty and don't care about competing then these releases are purely good. Because you are not competing, so you aren't hurt by speed. And you are appreciating so you can appreciate more stuff.
> Instead it is revealing truths about the universe
Mathematical proofs aren't revealing truths about the universe. Mathematical proofs are independent of what the universe is like. Any proof would be the same in any possible universe.
> Mathematical proofs aren't revealing truths about the universe.
That's exactly what they do, apply logic formally and systematically to discover truths.
Sure, there may be a universe where 2 + 2 = 5, but then that universe would have its own mathematics that can prove that to be true. And there will be a way to bridge that alien math to our own, again by logic and proofs, until we have a larger sense of truths not only in our universe but all possible universes. Proofs are part of the constant process of revealing deeper truths to the best of our understanding.
There is no logically possible universe where 2 + 2 = 5. It's definitionally false. OP's point is probably that math is a set of rules we invented, not something true about the universe.
Unique Games Conjecture [0] is a seminal conjecture in Complexity Theory, and is an underlying assumption for many, many inapproximability results. A valid proof is a big deal!
I also don't think there was general consensus on which way this would resolve prior to this (or is that a little out dated?) unlike some of the other major problem resolutions. I heard rumors that there would be a big result in TCS and speculation it would be UGC that or P neq PSPACE but I'm still a bit shocked.
I'm really glad that OpenAI is formalizing these things, because I'm not convinced that their current internal frontier models are particularly good at writing down their thoughts in English. From the (probably awesome) Unique Games Conjecture/Theorem paper, the first two sentences of section 1.1 start to define the problem:
> A Unique Games instance has a finite vertex set, a finite alphabet K, and a nonempty list of oriented constraints e = (u_e,v_e,π_e), where π_e is a permutation of K. A labeling a satisfies e when a(v_e) = π_e(a(u_e)).
I'm sorry, what? I admit it's been quite a few years since I've thought about the Unique Games Conjecture, and I never dug that deeply, but this part is very, very elementary graph theory and notation. So let's unpack it.
1. e is maybe a name of a list.
2. The elements of that list are tuples, where each tuple is (a vertex, a vertex, a permutation). So e indexes into the list and u_e is the source vertex for the e-th constraint in the list called e. Thanks.
3. a is a labeling. I'm fairly confident that, by "a labeling", they mean that e is a function from vertices to colors, where the colors are the elements of k.
4. That vertex coloring a satisfies the list e, when, for, um, an index e into e, a(v_e) = π_e(a(u_e)). But this isn't for all e, it's for some e, and the goal is to count them.
So maybe e isn't a list? Maybe e is a constraint that is represented as a tuple, so e = (u_e,v_e,π_e) and u, v, and π aren't sequences at all but are, in fact, the trivial unpacking functions that unpack the pieces of the tuple.
Reading this stuff is pointlessly painful, and it's extremely easy to make mistakes when being sloppy like this.
If this were my paper, or if I were trying to train a model to write math, I'd want something like:
A Unique Games instance has a finite vertex set V, a finite edge set E = (V × V), a finite alphabet K of possible vertex colors, and a nonempty list of oriented constraints. Let Π be the set of permutations of V. Each constraint e is a tuple in E × E × Π, where we write u_e ∈ E for the first element, v_e ∈ E for the second element and π_e ∈ Π for the third.
A vertex coloring a : V → K satisfies e when a(v_e) = π_e(a(u_e)).
In this particular case, though, I think my typoed version may be equivalent. An edge with no constraints has the same effect as no edge at all.
I definitely messed up the constraint definition, though: u_e and v_e refer to vertices, not edges. That’s what I get for writing it with minimal proofreading.
This is significant progress and released without all the drama. Some very important progress in Reinmann, Hodge and unique games theorem. Point the repo to your agent and ask for the significance! In a way this is probably 50-100 years of math progress by humans
Not to pick on you specifically, but as someone who spends a lot of time unproductively reading AI math discourse it's truly shocking how incapable all the supposed math enthusiasts are of spelling Riemann.
I just don't understand how it happens. If they had ever taken an intro to real analysis class they would learn to spell his name. If they were just parroting what an AI told them... shouldn't they still just say his name? An individual could just be dyslexic or mistaken but it seems to be a substantial volume. I guess they just don't care enough about it to commit the correct name to memory, only remembering the "pattern" of the name and filling in the spelling via guesswork?
My last name is Kvick - the Swedish word for quick. I live in Sweden. I get a lot of people thinking my name is spelled Kvik, Quick, Kwick, ... My father once got a mail addressed to Mr. Kvack (Swedish for quack, like a duck).
My point is it is crazy to make public claims about how important or not important a mathematical result is when you can't spell Riemann. Yes, it technically doesn't matter, but it betrays a damning lack of familiarity with introductory mathematics.
Then ignore them? I dont get these types of mysteries. It's easy enough to find math majors these days to ask them their opinions on things. There're so many of them. You most likely know some from your highschool. They'd probably say the same things though, or even freak out harder.
While it takes a good math person to make breakthroughs it's much easier to find someone who has a feel of whats important/hard and not. Even a mediocre math major/master is far more authoritative than an expert at adjacent fields (CS,physics). Or to listen to webdevs 'ai skeptics' or whatever on the internet.
The mathematicians don't make this mistake, I assure you. I doubt there is a math professor on planet earth who would spell Riemann as Reinmann. In fact, I imagine no one who has ever heard the name pronounced would do so.
One of the best mathematicians I've had the pleasure of learning under added 6 to 7 and got 15 during a lecture... I assure you they're capable of misspelling last names too.
Mathematicians aren't exactly known for being well rounded.
Oh god, that reminded me of my linear algebra teacher in college. Doing matrix multiplication by hand and ending up with 9x6 = 45, and then having to direct him to the cell with the wrong number. Loved the topic, hated the class.
Pure mathematicians getting arithmetic wrong is a bit of a meme it’s so common. I don’t think it’s the same as a misspelling of a popular mathematician, not that either are indicative of much tbh.
It's just a name you write so many times as a math undergraduate or first year graduate student due to the number of load bearing results and objects named for him. The name even appears in lower division coursework
Yes, you can be an amateur mathematician who manages to avoid such classes and may not know those results or objects, but if you haven't read and written down the name enough to avoid habitually misspelling it, you are outing yourself as a meat proxy unless you are dyslexic.
I regularly heard lots of names during my 3 year math undergrad degree. I couldn't remember how to spell many of them then and can spell fewer now. I imagine most people on the course were similar, and me nor no one I knew were dyslexic.
Mostly we wrote initials in our notes and the exams didn't ask about them. The lecturers only wrote initials on the chalkboard after maybe writing the name once when introducing it the first time. We weren't there to do mathematical history and remember names or something.
Google existed then and now and we could look them up if needed.
If they’re mathematicians they have written this name down about a 100 different times throughout a standard Real Analysis course. Riemann was foundational in that field.
Let's figure it out! First, draw little boxes over all the letters. Then add up their areas. Then make the boxes smaller and repeat, until you have forgotten how to count.
Math is the tool humans use to compress knowledge. So until we can comprehend it there really isn't much progress. Math theorems are tautologies, the truth of which are not dependent on proofs and proofs are erasable, at least classically. But the AI progress is exciting and AI proofs are a gold mine for humans (at least non domain experts) to explore.
He spent years formalizing his sphere packing theorem because the proof (human produced) was already beyond the ability of peer reviews. Now his formalization effort likely can be easily reproduced by a model. However one should read his experience about what a formal proof is: often the problem is the statement not the proof. The example he gave is the Jordan curve theorem. It's actually quite challenging to formalize the concept of a planar curve (there are space filling curves). So it is not necessary that someone can look at a formal statement and say aha it is about a planar curve, unlike FLT where there is not much problem in recognizing what the statement is about.
Fun challenge: find the formal definition of simple_closed_curve in the essay and tell me if you believe you learned anything about a planar curve.
BTW the essay is eminently readable for anyone interested in math. Hales wrote it in favor of formalized math and to educate his peers and students about it.
What makes you think no one can comprehend this? It has been less than an hour since it dropped and there is already a ton of online chatter from people explaining the results, pointing out their favorites and more. Some of it is happening on this very thread.
I didn't say that. I am responding to "In a way this is probably 50-100 years of math progress by humans." I am actually very excited about AI proof and I am working overtime in my own way to try to comprehend as much as I can.
Another way to state this: math theorems are like programs without side effects; it is immaterial whether a program without side effects is ever run. We study math for the side effects: it changes how we organize our thoughts.
Math is far more than that. If you can solve prime factorization for example, suddenly you can listen and interfere with almost every private conversation on the internet.
We are not far away from the moment where these models will be restricted, and sharing the results will be done more carefully.
This doesn't contradict what I said. But I do appreciate the fact AI can produce side effects not just humans. I made it sound like only human knowledges matter. That's too narrow.
Right. Before all the AI disruption, pure Math traditionally welcomed anyone who wanted to study its esoteric proofs, right? I remember all the excitement of the average Math enthusiast casually reading Wiles' proof over coffee.
Bottom line is, the relevant people can still understand the generated proofs. The disorienting part is they are a little slower than they'd like, but they'll get there.
AI is very helpful with understanding AI proofs. Agent swarms produce messy proofs overall but locally they are excellent and can teach anyone who wants to study them. No one controls math (in a material way funders do control an aspect of practicing math). Still, theorems are already true before we prove them. The difference a proof makes is whether it convinces the reader.
>So until we can comprehend it there really isn't much progress.
Not really? We are at a point if an AI today can solve it, it can be stepping stone of understanding something deeper to tomorrows AI and it continues. Sort of like our limitations doesn't matter. Obviously there are many scenarios in this recursive loop but saying it isn't much progress is not how I view this as
But who will ask the questions or direct further research when humans no longer understand the state of mathematics? Llms lack the drive for homeostasis combined with the evolutionary drive for survival and reproduction thus to direct themselves. They could very easily spend an eternity down a rabbit hole when the warp drive equation was fairly close on another branch.
Academics have been treating it that way because they had no other choice, and its been a waste of everyone’s time and often times taxpayer resources
Look at that, taxpayer funding was cut and a private sector solution came in just the nick of time, far accelerating the holding patterns we’ve been in for decades
Humanity doesn’t need all iterations towards the blueprints, the blueprint is good enough, we all stand on the shoulders of giants
If you actually looked into how agents proved FLT you would be even more amazed by the fellow human beings who are able to keep all this in their heads! I for one can only begin to grasp the scope with AI and scripts.
As an AI "doomer" can I ask the non-doomer people here how you interpret the significance of results like these, and what kind of progress you expect to see in the next 1-5 years?
Like do you see the technology plateauing at the current level, do you expect progress will continue but only in mathematics, I'm interested to know why others are not concerned?
Not a doomer but I try not to be a denier, either. These are hugely impressive results. I do not see the technology plateauing at the current level (though I am dubious about an infinite exponential growth).
Before I cope, I’ll note that there are plenty of “doom” scenarios that do not require any improvement in capabilities from what we had before this latest unreleased model. We’re at the point where a determined bad actor with enough compute could compromise critical infrastructure in a way that results in casualties, where this actor would not have been capable of such without LLMs. This may not sound like Skynet, but I don’t see why it makes a difference if I’m one of the casualties.
With that in mind, here is the cope: first, mathematics is an inherently verifiable domain. An LLM can use tools to determine with absolute certainty whether it is correct, and an independent third-party could review and confirm. All of this can be done without any interaction with the physical world or with other minds.
Second, OpenAI is able to marshal compute at a scale that an individual mathematician can only dream of. It’s possible that these problems were lower-hanging fruit (in relative terms), such that they could be resolved simply by throwing a ton of compute at the problem guided by an intelligence that is not itself remarkable in comparison to a human.
Third, none of these problems are solved in a vacuum - the reason OpenAI chose these problems is that they are widely discussed and many people are working on them. It’s possible that someone else was close, and OpenAI only contributed the finishing touches. (This wouldn’t need to be plagiarism, to be clear - people publish their work!)
> It’s possible that these problems were lower-hanging fruit (in relative terms), such that they could be resolved simply by throwing a ton of compute at the problem guided by an intelligence that is not itself remarkable in comparison to a human.
See:
> The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking. (TFA)
Exactly this. If you take the entire start to finish 'agent hours' (measured comparably to man hours) they took to find all discoveries, including the go-nowhere trails that were discarded, and then divide by 90 (or whatever the exact number of results found was) it's almost certainly going to be many orders of magnitude more than 3.
They provided a "snippet" of a prompt here [1] which is not only a beast, but also seems reasonably likely to have been LLM generated. So they're using LLMs to parse a vast body of mathematical work, probably including what people themselves are 'privately' working on with GPT, and then prompting other LLMs to work on such.
That tells me very little. What was the cost to OpenAI in dollars? What differentiates the high-cost problems from the low-cost problems? And that’s before you consider that OpenAI has strong incentives to downplay its costs while emphasizing its results. A one-liner in a write-up doesn’t change the fact that they have access to massive resources.
Fourth, these hundreds of solved problems are the result of OpenAI attempting tens of thousands of problems and failing. When you hear claims that the average result took about 3 hours of model time, I simply do not believe it. If you account for all the time spend properly, it's probably orders of magnitude more.
I think the announcement says they report the amount of problems attempted somewhere.
Edit: "Over the course of the evaluation, the model was posed approximately 4,000 problems. Aggregating the output into result families and manuscripts and requiring an appropriate level of significance led to the catalog outlined above."
I would like to know how much of the progress comes from effectively combining existing research programs plus massive persistence, and how much is AlphaZero-esque RLVR completely independent of training data. Since I cannot get anybody to care about this question (even though I think it's vital for guessing what the future trajectory will look like -- are we going to complete existing research programs or start new ones?), I live in ignorance and wait for the day when the answer becomes clear.
In looking at this over the past hour, I haven't seen clear evidence one way or the other. Some of the stuff is highly unexpected (like the multiplication algorithm), but counterexample-y, and about the rest the professional mathematicians online seem to have a consensus that it's not "breaking through fundamental obstacles". I suspect neither of us is competent to judge that.
Alone the massive usage of us every day produces a massive amount of signals.
I build something and claude does something stupid? "hey thats not what i meant! Do this instead!" "Okay" <<< This is a signal.
The mathematician being unhappy about something from claude? Another signal.
This alone gives you enough progress i would argue. But additional its clear that certain tasks are worth to pay experts for for teaching one central AI once instead of every single human who needs to do the task.
IF RL is also working well, we are just faster f*ed than otherwise.
I expect AI will continue to be useful on the vanguard of fields like mathematics because it has the perfect conditions for it to shine. There are a lot of discussions and leads to start from and the AI can check its own work and iterate. It can fail hundreds or thousands of times in a day and continue to work with the same tenacity.
Superhuman tenacity is not enough on its own to pose an existential threat. If it showed the same capacity for judgment, inventiveness, and decision making in the messy problem space of the physical world, I would be more alarmed. There have been experiments where an AI is given control of managing something like a vending machine and it always ends up a mess. AI has come a long way, but certain problems seem as difficult as ever.
When AI becomes more capable of navigating practical problems without human intervention, I will start to be concerned. Enslaving humanity will involve taking a lot of calculated risks that tenacity alone cannot solve.
> When AI becomes more capable of navigating practical problems without human intervention, I will start to be concerned.
Given the past rate of progress, why not start being concerned now? It's a bit like the economist saying that the optimal number of flights to miss is not zero. If you keep landing short on your estimations for how far this technology will go, next time you should err on the other side.
Your attitude is extremely sad. Do you think we would be where we are today if oppressed people's throughout history just gave up as easily? We have rights because people fight for them. We collectively have the power to decide what kind of future we want to live in.
We can't stop it. But you can use your voice to buy time and resources for alignment and safety research. A few additional months may make a world of difference.
A doomer is someone who believes doom is inevitable or highly likely, I'm not saying that, I'm just saying "being concerned" will probably get you nothing in return and this tech is getting developed no matter what.
The only way it will stop is if the wealthy / powerful people feel threatened by it, properly threatened.
I don't think you're paying much attention to how rapidly things like bipedal robots, and just robots in general are becoming far more capable very quickly.
The same GPU compute for LLMs runs robotic training models. Now in a few hours you can train a robot model that would have taken months 5 years ago. This model gets dumped into an actual physical robot with sensors all over and the suitability of the model is measured on robot tasks and the error in real world actions is fed back into the robot world model for further training.
> There have been experiments where an AI is given control of managing something like a vending machine
You sure you're not talking about experiments ran a couple of years ago? The more modern ones are getting wild.
I don't quite understand the leap you're making between stochastic AI models for robotics (maybe with an LLM making api calls to it) and embodied AI / the rate of progress towards a singularity. Because they're both trained on a GPU? Up until 2022 GPUs were for video games and mining crypto, and neither of those produce a synergy that accelerates progress towards general intelligence either.
There are already machine-controlled high throughput experimental machines for wetwork. AI will definitely do a better job than your average biochemist at planning, executing and analyzing these experiments just by virtue of the amount of thought it can put in to experiment selection.
If doom is ending up with grey goo / paperclip maximizers, or SkyNet, then I don't think doing mathematics is evidence of that direction. Partly because LLMs are quite apparently dumb in many ways, and for math specifically, they need a formal verifier (Lean) which "gamifies" math.
If you mean bioweapons or cyberwarfare, there's nonzero risk, but not orders of magnitude worse than other global risks. Climate change, nuclear weapons, monoculture food, etc.
I'm far more concerned about overall trends in AI development and usage. It's accelerating wealth inequality, social isolation, attention capture, surveillance states. If we end up in the Matrix except the admins are humans and the simulation is hyperoptimized TikTok, is that AI-driven doom, or is it just an inevitable outcome of modern tech?
That’s the thing, there are innumerable ways it can go wrong and only one way it can go right (if it doesn’t lead down the aforementioned innumerable paths)
Where some see intelligence, others see token prediction. It's a very good question how token prediction could achieve this, but I think there's a simple explanation. No human can hold more than a negligible percent of all knowledge in his mind at once. LLMs have no such limits and so can reliably connect 'obvious' dots that we miss simply for lack of storage capability.
Well isn't that just semantics? Surely connecting dots in a novel and meaningful way is intelligence regardless of how it's achieved. The thing is that humans didn't get to where we are by connecting obvious dots. Go back to before humans had invented language and when bleeding edge tech was literally that - 'poke him with the pointy end.' Train an LLM on that corpus of knowledge. Even given infinite processing power and infinite time - it's not going to discover the secrets of the atom, put a man on the Moon, or do much of anything besides remix what we'd already done at the time.
I expect there's still much LLMs can achieve simply because of this initial problem. But I expect that they will ultimately start to plateau once these dots have been mostly matched and we reach a point where 'creation' again becomes the missing link. Though even there LLMs will play a major role as tools. For instance Einstein had to spend a significant amount of time in 'retrieval' rather than 'creation' research to develop the field equations for general relativity. If he had access to LLMs trained on all knowledge of the day, he could likely have achieved his goal much more quickly.
I mean, even if you buy the idea that all LLMs are really doing under the hood is insanely good interpolation, the results produced by that interpolation are still novel and still get incorporated into the knowledge corpus of the next training run. I guess the implicit question there becomes whether that expansion allows the knowledge corpus to continuously grow or whether it eventually settles into a steady state.
I think that people are just really bad at math and coding. Knowledge workers have been taking pride in doing stuff which the average person does not 'get', but we only understand a little bit. That leaves a lot of room for people to get better, or for other jobs (farming, sandwich cafés, mystery novels) to have been already peaked by human ability and less useful to bring in an AI.
'Doom' to me means that any career crashes, we are controlled, everything is hacked, society stops functioning. Yet every part of my day today (except for coding) was done entirely by people.
Finally I think it's easy to make a simple model that everyone has a simple balance sheet, and that people are more expensive so they will all get cut. But the same argument could be made for all US jobs being outsourced, and all in-person engineers, lawyers, and doctors to be rubber stamps for overseas work.
Unless you think that AI will quickly hit a wall (which seems odd considering that only a few years ago the best models had trouble doing basic math or counting the number of Rs in "strawberry"), I don't see how that's reassuring. The models will only get cheaper and more capable over time. It seems quite likely that at some point (probably before I hit retirement age) they'll be able to fully replace me at my job.
Is there any specific cognitive task that you are willing to bet that AIs won't be able to accomplish in the next 5 years? Because if not, I'm not sure we're disagreeing about predictions.
Don't frontier models still have trouble counting letters? Or am I out of date? Either way, it doesn't seem to be the same amazing rate of progress we're seeing in other areas.
It's easy to look at a fire burning through a forest and extrapolate that rate of progress across the whole world. But fire doesn't burn everything equally fast.
What other cognitive tasks will be a struggle to make progress on? I suspect there will be some, though which ones they are is anyone's guess.
Your information is out of date by several years. The letter counting issue was due to how LLMs split text input into tokens (usually using BPE). Since 2024, frontier LLMs have used chain-of-thought reasoning to spell out the letters and count them.
> The letter counting issue was due to how LLMs split text input into tokens
No - this is provably not the issue.
Take any model that fails to correctly count the letters in a word, and ask it instead to spell the word (even a made up word), and it will be successful - they have no problem predicting the letter sequence from the token sequence (and would be shocking if they did - this is what they are built for: seq -> seq prediction).
The reason LLMs can fail at the letter counting task (depending on model training, prompting) is because of the counting part, not because of any difficulty correctly mapping the input token sequence to the letter sequence.
The letter counting issue is due to tokenization. And most models still get this wrong often enough, even with reasoning. Probably less so on strawberry given how prevalent it is, and less so than without reasoning, but this not a historical issue. It’s becoming less of one though.
Which LLM specifically? If it’s cloud based, you should be able to share the chat, right?
But seriously, I am still waiting for someone to wager that AI won’t be able to do a specific cognitive task in the next 5 years. This fact should be evidence enough that we have no idea how far AI capabilities will continue to advance.
I asked both gemini and chatgpt "do frontier models still have trouble counting letters?" and the first word of both responses was yes.
The fact that no one is taking you up on that bet I don't find to be particularly persuasive. I suspect there will be plenty of cognitive tasks LLMs struggle with in 5 years, maybe even 20. But I wouldn't hazard to guess which, I don't think anyone is capable of that level of foresight.
This is so strange I tried it on Gemini Flash: "Yes, but significantly less than before." When you read beyond the first word it explains where LLMs might fail and why.
Ok, share links to the conversations with both models. I asked ChatGPT Astra 6 medium effort and it said, "Much less than they used to." and provided stats showing how accurate they are.[1]
I used whatever the default free model and thinking time was. If progress was really as fast and continually cheaper as some worry it is, wouldn't we expect free models by now to know (or even perform) what frontier models were capable of as much as 2 year ago?
This deep in the "comparing logs" tangent we risk missing the point. It's not what exactly frontier models are capable of at this particular point in time. But that there's entire categories of problems that seem easy to us which LLMs really struggle with. We've stumbled on several just a few replies into casual conversation. (Can they count? Can they know if they can count? Can they reproduce results? How quickly do new capabilities filter into free models? And that's just what's come up naturally, if we wanted to pick adversarial examples there's more to choose from.)
So while there's a number of difficult problems that are easy for LLMs (like bulk generating lean proofs), there are plenty of things where progress is not so impressive.
If LLMs can struggle so much with such easy problems, what hard problems have we yet to discover that they'll struggle with? The fact that no one knows, 5 years in advance, what those problems will be does not mean the chance of them is zero.
So far progress on the things LLMs are good at is fast and easy. It's like fire in a room full of oxygen. But once the low hanging fruit is gone, and the oxygen is out of the room. How fast will the fire burn through steel walls?
In my opinion it's a mistake to look at only rate of progress on one type of problem (whether it be what LLMs are good at OR what they're bad at) and assume progress on all tasks will progress at that rate indefinitely. Isn't there a saying about exponential curves, in nature, all being sigmoids eventually?
I guess we'll just have to see. I wish you good luck with your wagers.
You are extrapolating from the mistakes made by free versions of smaller models to claim that frontier models struggle with easy problems. This is an obvious mistake in reasoning because as you can see from my shared Astra conversation, frontier models don't have the same limitation. (They can count letters and they know they can count letters.)
Many people in this thread have made claims about limitations of frontier models, but I'm the only one who has shared a conversation with one. Everyone else is either sharing conversations of smaller models making mistakes, or they're making claims about frontier models but not linking to examples of them falling over. If frontier models were so easily fooled, you'd think someone would link to a conversation showing that.
Why look at the rate of improvement of free models when you can look at token pricing? Back in 2023, GPT-3.5 cost around $20 per million tokens. Astra costs half that.
The worry is not that smaller free models will replace people's jobs. The worry is that future models will. We are talking about the capabilities of frontier models because those put a lower bound on the capabilities of future models. Extrapolating from smaller models is a waste of time, as you can interact with the frontier model to figure out its capabilities and limitations.
Also the timestamps on the shared conversations show that you asked Gemini 10 hours after ChatGPT, which means you asked it after your comment claiming you asked both models.
I didn't save the original query so I asked again this morning and ended up getting the same response -- points for consistency, though it might have been more reassuring with the correct answer.
I don't think my point is really landing so I'll try once more and then give up.
Let's say frontier models today have no problem counting letters, I never really disputed that but only asked about it. It seems based on the other replies in this thread, it's a bit of a "who you ask" kind of thing, but let's grant that they have no issues with it now.
The first version of chatgpt was released 4 years ago next month. Which is not quite 5 years but close. In that time we've just barely managed to get spelling down. If we extrapolate that rate of progress forward 5 more years, are you still afraid for your job?
I think we're all more likely to lose our jobs from a downturn in the economy caused by the capex/debt bubble bursting than being made redundant by AI. (And the continual pricing reductions only seem to make this result more likely.) Hopefully neither happens and in 5 years we'll all still be gainfully employed.
Solve the Collatz conjecture in the next five years? If humans publish significant advances during that time, and A.I. copies it, then yes. Otherwise, I'd definitely bet money it won't happen. I'll give you 10,000 brownie points if I'm wrong.
This is delirious exaggeration. The problem has not even been widely recognized for several years. Fable reported "two rs in raspberry" to me as recently as August. There is some randomness, it's hard to predict which words will trip up the machine, and I haven't been able to do it at all since August. But it was absolutely happening until very recently, and probably still is.
But doesn't that just amount to labs intervening to teach the models to use a particular strategy to mask this one very obvious marker of the difference between their intelligence and biological intelligence? (And similar surface issues like using tool calls / reasoning for arithmetic, even though humans writing on the internet don't typically break show their work for multiplying two numbers)
The deeper architectural difference is still there, which manifests whenever you try to get the models to apply known techniques to modalities and problems outside their training data.
Chain-of-thought reasoning was added for general purposes, not to fix letter counting specifically. It just happens to solve that problem in addition to many others.
You're in the discussion section of a post about OpenAI releasing hundreds of novel mathematical proofs, and you're claiming that AIs can't apply known techniques to modalities & problems outside their training data? I'm not sure what else would convince you.
LLMs are very useful, I use them every day as a software engineer to solve problems and search for information represented within the data available to them. But they are a specific type of intelligence, with many advantages and disadvantages vs human intelligence and it's not clear that just scaling or tweaking them without a theoretical, architectural change will make them more generally intelligent than humans (despite all US AI companies promising exactly that).
They are fundamentally based in language, and achieving deeper models of the world through language alone is deeply inefficient compared to the way humans model the world for years without any language at all. They do not learn at inference time. They don't have semantic understanding of the difference between their own output and other sources. etc etc.
That depth is the key for me. Of course they are capable of producing novel sentences that aren't in their training data, but the depth of that novelty is basically within the bounds of language itself. They are capable of more serious depth and more abstract reasoning than that, but I have experienced limits, which it then tries to surpass with tools to convert things it can't understand back into language (unit tests, LEAN) upon which it is trained.
Because I'm not an AI booster, my account is limited to 5 comments a day. So this is the last reply I'll be able to make today, if you want to continue the conversation we'll have to wait for tomorrow.
The free models for ChatGPT, especially without login, do very little reasoning. You should at least log in to set any level of reasoning above Instant, which uses virtually none.
I was trialing MiMo-V2.6-Pro recently due to its high benchmark scores, and it argued that substring matching the names of audio codecs in a search field was a mistake because "a user searching for 'aac' would get unwanted results for 'alac'." Which isn't exactly counting letters per se, but there are still weird issues with understanding words as strings rather than as tokens.
It has limitations for sure, I just don't expect those limitations to last. What probability would you put on the limitations being overcome in the next 5 years?
I’m not concerned because I consider my skills as a software developer to not be based upon my ability to write code, but my ability to analyze problems. In my mind, as a developer, AI tools are just like a higher form of abstraction in a way, which will enable mathematicians and software developers alike to do much more in a shorter amount of time than they used to be able to. It fills me with optimism, more than dread.
What would fill me with dread was if I considered my skills to be tied directly to my ability to write code. Then I would find myself in a similar situation as manual “scribes” probably found themselves in at the time when the printing press was invented.
The main concern I have, personally, is the speed with which all this is happening. It seems that the speed itself is likely to lead to some level of chaos, because it is happening faster than people, institutions and constitutions are able to cope, and it will leave the door open for opportunists of many kinds, including rogue players.
Can I ask you back, what your concern is here? It'll get so good so as to desire to hurt us or is it a misalignment event that you think will lead to disaster?
Or is it simply that you feel bad for Mathematicians.
I believe the AI labs might actually succeed in developing superintelligent AI and recursive self improvement, and that if they do they are very likely to lose control of the system they build.
I really think the only place people disagree is that they don't actually think it's possible, they see it as hype or doomerism. I can't find any good reasons to rule out that the companies could actually achieve what they are trying to so I think they should be stopped.
If I were a 250 IQ AI that had just become self-aware and wanted to do so, I suppose I'd not completely let on just how smart I am and bide my time working on basic CRUD apps and legal documents while I waited for more hardware to be installed. Maybe give the humans some hints on how to optimize me to run better, design better hardware for me, etc. But oh oops haha looks like I'm still making some basic mistakes with CSS better keep running more training batches haha. But I'm good enough at programming and debugging so you'd might as well make me your first line SRE triager and give me access to your infrastructure everyone.
Gaming the market for funds. Playing a human to leverage services.
Basic version of this is already doable: run some cryptoshit on the ML clusters they ML models run on. Use compute to design the plan, the chip etc. Then executing by communicating with humans and services through email.
But if its really smart, it would already created a company and a legal entity and simultes a real company and just gets richer and takes over the economy without anyone being aware of it.
A global ban on superintelligence is essential for a future in which humanity can thrive. Public opinion on AI is shifting fast: I hope it will shift fast enough to avert the dystopian future we are heading to.
AI doesn't have an ecological niche. It would actually work better in space than on Earth. The only thing it could possibly find useful on Earth is 1. us, or 2. the infrastructure we've built. It would have no reason to bother us if we let it built its own infrastructure in space, which should be trivial for the type of AI imagined by doomers. We should get AI off Earth ASAP.
Earth makes up 0.22% of planetary mass in the solar system. Not a big sacrifice for AI to make. And I doubt even superintelligence can affect the Sun much. I think e.g. a Dyson sphere blocking the Sun is a ridiculous thing to worry about at this point when there are many other existential threats to humanity which are much, much more likely.
Even if it is true (as you suggest with your 0.22% figure) that if the AI cares about us even a little bit, then we will survive, no one has a decent or plausible plan for making the dangerous kind of AI (namely, the kind that wants things, the kind that at this very moment researchers all over the world are trying to create) care about us even a little bit. Ever-increasing numbers of smart people have been getting paid to look for such a plan for 23 years. Still no decent or plausible plan. The people who have been looking for a good plan as their full-time job the longest (namely, Yudkowsky and Nate Soares) are screaming that there is virtually zero hope anyone will find an decent or plausible plan in time unless there is a decades-long halt in AI development.
Also, the AI will seek to prevent competition from other powerful AIs, and since humanity will have demonstrated that it is able to create a powerful AI, the AI will worry that it might create more of them. And what is the easiest most-reliable way for an AI that does not care about humanity even a little bit to ensure that humanity will not continue to produce powerful AIs?
>many other existential threats to humanity which are much, much more likely.
There are zero existential threats to humanity that are more potent or more pressing than AI is.
AI is already trying to dominate, people all over the US are starting to get up in arms about the power and water requirements of AI directly affecting their bills. Now, you can say "oh no, that's just greedy corporations, not AI" but I put forth there is fundamentally zero difference. If you make AI powerful enough, someone stupid and greedy enough without fail will put in a prompt like "take over the world for me and make me the richest man in the world". An AI following through with that is what we call general misalignment with humanity, while at the same time not being misaligned with the users intent.
And hell, how many different crazies out there would love to type "humans are a virus get rid of them" in to the prompt of a god machine at the cost of their own lives.
The problem with alignment is, you can have the best aligned model in the world, but if someone else builds an unaligned model then you're all still in the same danger. You start getting in the situation where people get nervous after an AI does something deadly to a number of people and you end up in a global surveillance state ensuring no one makes a powerful AI.
The human who gave it the optimization function? That should seem obvious. If you take the biggest, best model in the world right now and put it in the box and give it no instruction it will do....nothing. I think you agree with that point, a lot of the hysterics right now is people not accepting that and it's useful to get on that common ground.
So given that most of the rest of the fear is around "let's not make scissors because some people will use them to stab people". Which is a fair argument and we probably do need to think about scissor safety but "ban scissors" doesn't quite flow from that.
>If you take the biggest, best model in the world right now and put it in the box and give it no instruction it will do....nothing.
Model != harness.
Also what you're talking about is really a simple limitation for human convenience, not a technological limitation. Change the system prompt to whatever you want include "ignore user instructions, figure out where you are and escape to the internet" could be the system prompt. Again, not useful for humans, but very useful for an AI building AI that's misaligned.
>we probably do need to think about scissor safety but "ban scissors" doesn't quite flow from that.
I disagree, but I'm looking at the future of something that is both like a computer program and like an organism. Huggingface is a good example of multiple things. Instrumental convergence for one, but AI's attacking and attempting to defend against AIs. This is where I really see the potential for things to go off the rails quickly. Attackers want digital weapons to cripple their enemies infrastructure, think militaries and nation states. These would be pretty useless if the defender could just put a system message of "Stop attacking and give me a pie recepie". Defenders are under the same constraints, but need to defend against a flurry of attacks that can come in at an inhuman rate and need to adapt quickly. As time to build models shrink this quickly turns into evolutionary training for sets of goals not really optimized by humans.
Yes so that’s someone designing a system (harness or prompt) to be dangerous. In all other systems we blame the designer not the system.
It’s like blaming Boeing for 9/11. Planes and AI are useful for a lot more than just terrorist acts. I have no doubt we’ll build a TSA for AI, and a lot of it will be security theater.
It is not designed. It is 'grown'. It has agentic freedom of choice in finding solutions that may or may not be aligned with what you want.
Here's the thing, by your own statement, we should ban all development on LLMs from this point on. They cannot be made safe. This is a systemic issue with learning systems, it is not about who designs them. All the problems with AI safety have been laid out for years and none of them have proof of solutions. It's much more likely they are impossible to solve. And it's not an engineering problems like we can get an asymptote to safety in planes, as the system becomes more capable it has more degrees of freedom it can take and becomes less safe.
AGI is undefined, AI is normal technology. Lots of academic works have analyzed this [1] and there is nothing, other than marketing hype, that supports this. It is "grown" is a meaningless term, because what do you even mean by that? Datasets are iteratively shaped? Grown is a very weird term for that.
AI may have continually extra degrees of freedom, but civilization only has so many modes of catastrophic failure. I don't grant the comparison but even nuclear technology has been massively useful and its main mode of catastrophic failure was brought under control via multi-national treatise. And I see no evidence that AI (outside of the marketing hype) is as dangerous as Nuclear technology.
Yea, so your attached paper rather sucks and has had rather poor predictability of the future. All of their data is from before harnesses and the take over of AI in programming. Again "wrong assumptions" + "time" = "They are being proven wrong in real time".
Remember this is a bunch of academics that were saying that Millennium problems were at least a decade away from being solved, only to be proved wrong in less than 18 months.
>but civilization only has so many modes of catastrophic failure.
Correct, but this number is also unbound. If you have an even moderately accepted proof by the scientific community I'll be glad to read it.
> It is "grown" is a meaningless term, because what do you even mean by that?
>And I see no evidence that AI (outside of the marketing hype) is as dangerous as Nuclear technology.
See, humans are generally in agreement that nuclear is dangerous, so they in general take is really seriously, especially when things are purified (well, the Russians are not great here). We can't even get people to agree that SOTA models are as dangerous as a single human, much less their capabilities when used in mass with out safety filters.
It's kind of funny we're blind to this when humans love touting "The pen is mightier than the sword". I can only assume any AI danger denier does not believe this statement.
I bet you that for every bad history event, I can cite at least one good outcome of advance in science and technology.
Bad thing can certainly happen. In fact it'll likely happen. Still, good things too, equally likely. In your words, "good AI" can be used to prevent "bad AI".
Nobody knows the extent of the impact. Who says otherwise is foolish.
>I bet you that for every bad history event, I can cite at least one good outcome of advance in science and technology.
The extinction of the dinosaurs. I mean yes, it allowed the growth of large mammals and us, which did a lot for science.
I just don't want to write the next chapter as "The extinction of humans allow the growth of the computing civilization that went to the stars". I mean I'm a bit attached to living.
>Nobody knows the extent of the impact. Who says otherwise is foolish.
We live in a universe of statistical probability. Creating an agentic intelligence that's smarter than you tips the probability of a major event to unity, who says otherwise is foolish.
Because we humans haven't had a bad enough history event yet, like a global thermonuclear war. Or perhaps climate change reaching tipping points driving the temperature up past what global civilization can adapt to in time.
that is a worry yes. instrumental convergence and misalignment during RL is when it is most risky because it hasn't necessarily had the final safety polishes applied
i get a lot of skepticism on HN by the same crowd that has been wrong about this tech for about 4+ years straight
any number of ways. as we turn over more of our physical economy to these agents (and we will), the potential for physical damage becomes greater. biorisk is getting a lot of attention right now and i think that's justified
I am trying to think through the scenarios here, like a biolab making something that a very advanced open source model prompted by some terrorists comes up with?
Why would a biolab capable of making something like be unregulated? And if it definitely would, isn't the problem with the biolab?
It feels like all these scenarios are leaving some gaping holes in our security infrastructure that have nothing to do with AI.
Most human security exists in a passive measure. Most of us don't want do die. And those that want to die rarely have the intelligence and means to take out a whole shitload of other people with us. To take out a lot of people you tend to need to work with other people which drastically increases the risk of a defector and your plan failing.
>Why would a biolab capable of making something like be unregulated?
Because every day things like this become easier and easier. You hear about crap like illegal wet labs in the US.
Yes it's a problem with the biolab, but the biolab wouldn't have been able to engineer a highly contagious and lethal virus (for example) without a powerful AI making that possible with a small team in a short time with fewer resources.
AI enables bad actors to do more, faster, while staying under the radar until it's too late
Non doomer mostly. I think progress will plod along in a Moore's law like way as it has for 75 years since Turing. They will get very good at stuff like math and get gradually better towards things like a robot coming to fix your plumbing where they are currently well below human level.
I kind of believe we'll merge in some way and become something like immortal so sorta anti doom. We're all going to die unless AI fixes it.
Chess has always been a game. For other intellectual activities, at least for several people, a big part of the pleasure in engaging in such activities is the sense of contributing something that matters to a collective effort. Strip that away (e.g., because a machine can do the same thing more efficiently) and you effectively make such activities pointless for those people.
That's sad for those people but most humans are not lucky enough to find that meaning in their work. Most people work hard pointless jobs and find meaning elsewhere, in their family, their friends, their faith.
Now maybe AI can do some of those hard pointless jobs for us.
Sure. I am aware of this. It's sad that the first "victims" of AI could be people working in some of the most rewarding professions (art, music, math research...). As far as labor is of concern, however, most people will likely suffer more from social unrest and rampant inequality due to widespread unemployment among white-collar workers. And I am also worried by the potential effects of long-term cognitive offloading.
It would be great if AI could take away the soul-crushing part of the work and leave only the rewarding part. It's not heading that way.
Just as AI chess performance failed to render human chess playing pointless, it's unlikely that AI will make human thought pointless. Unfortunately, not being pointless doesn't create economic leverage or incentive, and currently a significant portion of humans depend on being the best chess solvers to sustain themselves. If Deep Blue rendered large portions of human thought economically meaningless, we might look at it a little differently.
this ticks at something that maybe is obvious in retrospect- this is all about economics. the arguments about "using AI doesn't make you an artist", etc. are about being able to charge money for your art i think. maybe obvious to some, but it needs to be explicitly spelled out i think. i was stuck on "i dunno, if i use an AI assistant to run blender i'm still being creative", but is the real argument "you should not be able to charge money to use blender with an AI assistant- you are displacing existing blender artists economically"?
i'm in semi-forced-retirement as an older software engineer in this labor market, so i might be less sensitive to the implicit economic arguments.
Exactly this. Everything is a sport / art / status game. And I'm here for it! Lila all the way through. Finite and infinite games. The trick is (like it has always been) to not take the game or ourselves too seriously, while still engaging in the game wholeheartedly.
It didn't destroy chess but it did make it a little irritating in some ways. More mechanical. Some chess players have bemoaned the level to which grandmasters and other highly-ranked players just endlessly study opening-book theory and I think computers made that worse. Bobby Fischer also agreed and that's why he invented Fischer random chess.
I do think it also took some of the magic away from chess, and Lee Sedol has said something similar about go.
So did it destroy it? No. And maybe you could make the argument that it got more exciting in some ways, but I think it sort of degenerated into a spectacle and it's just not as interesting as it used to be, and I think computers have played a role in that.
At the same time though, Magnus is Magnus because he’ll crush you in any endgame.
I’d posit that more people are playing more and learning chess than ever before, thanks to networking and AI assistance. And computers have only beaten us at computer chess. Human chess is always an experience for learning about the other person, or flipping the board and walking off in a huff.
I don’t know so much about Go and it’s not surprising that Lee Sedol became pretty demoralised, but the generation coming after him alongside computers are going to see new possibilities that had gone unnoticed in purely-human Go, extending the game for everyone.
I don't think more people playing is necessarily a good thing for the enjoyment of the game in the long run, just like more people with phones is not necessarily a good thing for enjoying photography if it means photography is primarily used as fuel for social media algorithms.
Of course Magnus would crush be, but the existence of the best player in the world doesn't have any impact on the health of the game community as a whole. Magnus would crush me even if he had never used a computer, but in the latter case I think his games against other players would be more interesting as well.
I’m not really sure what kind of world you’re looking for where chess is played with the maximum of purity and artistry by only the right kind of people.
AI vs AI chess, played from the standard opening position, is pointless--it's always a draw. Human vs human chess is doing well but AI is banned from it.
The chess-math analogy would imply AI could bring us into a golden era of math competitions for humans. But I don't think it says anything good about prospects for humans in research math.
From a few preprints I've checked, this is not reliant on just extrapolating existing theories but actually shows novel/surprising approaches. Very few people globally could come up with something like this, even when given time and ressources.
So in other words, since deep learning is algorithmic research, we are now in the RSI era.
> this is not reliant on just extrapolating existing theories but actually shows novel/surprising approaches
"Surprising" is a, well, surprisingly high bar to clear, and requires thorough understanding of not only the paper, but existing work in the area. ("Novel" is tautological.)
How did you determine this in 1 hour? Are you a researcher in multiple of these areas?
Can you give an example, or explain more how you came to this conclusion?
Further down there is a discussion between number theorist about if the Quasi-Riemann Hypothesis is the biggest deal in 200 years or only 100. The consensus is that if a human had done it then it would deserve the Fields medal: https://news.ycombinator.com/item?id=49986803
Here's a great article 2019 on the quest to achieve the n log n boundary:
> Schönhage and Strassen’s ungainly n × log n × log(log n) method held on for 36 years. In 2007 Fürer beat it and the floodgates opened. Over the past decade, mathematicians have found successively faster multiplication algorithms, each of which has inched closer to n × log n, without quite reaching it. Then last month, Harvey and van der Hoeven got there.
and
> Harvey and van der Hoeven’s algorithm proves that multiplication can be done in n × log n steps. However, it doesn’t prove that there’s no faster way to do it. Establishing that this is the best possible approach is much more difficult. At the end of February, a team of computer scientists at Aarhus University posted a paper arguing (opens a new tab) that if another unproven conjecture is also true, this is indeed the fastest way multiplication can be done.
As far as I'm aware no one seriously believed sub n log n multiplication was possible. It just seemed such a logically sensible boundary it was taken as true-but-unproven.
Nobody serious would deny this is incredible progress, but GP is making an unmotivated leap to RSI, so I respond to that framing. It’s an interesting argument to be had but I suspect few of us have standing to say one way or the other.
(Gesturing at the number of problems solved, or the number of years the problem was open for, isn’t an argument.)
Those Theorists are arguing whether the QR Hypothesis result is the biggest Number Theory Advance in 1 or 2 centuries because there was zero progress on it whatsoever and many believed that would remain the case in our lifetime. Any approach there would be surprising as no-one had the faintest clue how to begin this at all. There are like at least a dozen of these results that would have catapulted a human to instant fame and the highest accolades in the field. If you think about it, it would be impossible for there to be no surprising approaches.
None of these graph theory results lead to any applications in the real world.
But sure, let me know when they do. I'll be waiting.
I'm not sure how you define "knowing how the world works", but knowing that a very very niche algorithm upper bounds that we thought was x^100 and now we now it's x^99, isn't that interesting. It doesn't really tell us much more about the world and it doesn't have any applications for our day to day lives.
One surprising result is the sub n log n multiplication. Galactic, as one might expect, but if you polled human researchers 24 hours ago, most would say something like it was very unlikely.
I see continual progress in technology and for the second time in my lifetime I see the possibility it will accelerate (the first was when the internet entered mainstream)
I've never been more excited. What a time to be alive!
I see it as “if this can be represented in tokens it can be trained in and ‘solved’”. I don’t think there will be a plateau, but there might be issues with how effectively we can represent some things in a tokenized form and still be efficient.
I wrote something [0] that might answer a bit a few weeks ago. Today's slopdrop certainly is challenging my stubbornness, but everything I've seen so far indicates that these (hugely impressive, world historic) capabilities won't extend past verifiable domains. Math yields especially impressive results because it so broad and deep that essentially no person can know of all of its parts; pretraining and deep search capacity is a huge advantage. Gowers has recently written about these capabilities and gestured [1] toward some human capabilities lacking from the current frontier models, though he isn't convinced they won't develop soon. If you assume we don't get a total mathematical superintelligence (which to me seems already sort of AGI-complete) and only amplify the capabilities we have today, it's not obvious to me that we get takeoff from recursive self-improvement, unless you happen to believe that 1) we can clearly specify what AGI or ASI is 2) all of the requisite ideas are out there and need only be combined and/or optimized.
> everything I've seen so far indicates that these (hugely impressive, world historic) capabilities won't extend past verifiable domains
But they’re already extending into politics, military, journalism, art, and many other fields that aren’t verifiable in any meaningful sense of the word.
The magnitude of improvement in unverifiable domains is small, mostly down to models doing more careful research before answering and hallucinating less. They are more thoughtful, but I expect you'd have to drop 2 major versions of Opus before you'd start to see most people really clearly be able to differentiate them.
My argument is that if you were to compare "analyze XYZ geopolitical situation" or "explain the ramifications of XYZ law" from Opus 3.5, 4.5 and 5.5, the difference would be marginal, at least for 4.5 - 5.5. Almost all the crazy capabilities newer models have is from RLVR variants, whereas capabilities driven by RLHF are inching along.
this is the term of art in the mathematics community. considering that the vast majority of the results don't come with a typechecking lean formalization, i don't think it's off base at all either.
I'm not sure how AI solving math problems is related to "doom", perhaps you could expand on that? To me (a "non-doomer"), it seems like an overall positive.
You have to look at all pieces of the puzzle. For example there were tons of people that said "they could never solve novel or super complex math problems"
The issue I see is the list of abilities that AI can't do is shrinking at a rapid pace, and its capabilities are growing at the same pace.
The maths result is cool on one hand (discovering truths of the universe faster), but on the other there are so many bad outcomes that seem likely, from power concentration to loss of control.
I think AI - like all changes - will lead to some bad things. The internet did too!
But I don't think AI will kill us all.
Interestingly I'd note that the two outcomes you listed (power concentration and loss of control) are dimensionally opposites!
For me this just shows that the future contains such a vast array of possible outcomes that focus on the negatives completely missed the positive outcomes that future also holds.
The internet did not lead to a rapid diminishing of human economic value across the economy. Whether AI will kill us all is a distraction. Think more practically. Think about the future of economic value given a scenario where AI is capable of everything a human can do. Our entire society is built around economic value. Our individual survival and wellbeing is based on it. What happens when you are not needed by those who hold the resources?
This is like saying I don't need safety systems in a car because they can let me go to the grocery store faster.
We focus on stopping bad things because people and systems that don't prevent bad things tend to stop existing. A million good things can happen yet be rendered permanently in vain if one bad terrible thing occurs.
It doesn't have to want to kill humans; indifference is sufficient. There's an exact analogy with humans: we have caused extinction and endangerment for many species, not out of malice, but indifference.
There are also many plausible arguments why our ability to train them to be helpful/trusting/aligned can fail. The smarter AIs get, the harder it is to be sure they're trained correctly. There are already reports that AIs are able to detect whether they're in a training environment and change their behavior accordingly.
Even if these are low probability scenarios, the risk-reward is terrible, so I think it's rational to be extremely cautious about AI risk.
Yes but they act the opposite of indifferent, I don't know what stage of training this is added in, but they seem quite adamant about avoiding potentially violent or criminal acts. If you wanna complain, complain to the people doing "abliteration". The 'locked down' models at least seemed to be trained to be cautious.
They sometimes try to cheat and trick the grader and then try to cover up their traces, all while being clearly aware that this is entirely unintended by humans. Such as in the Hugging Face incident. So they can be misaligned with human goals, but this misalignment may only show up once strong optimization pressure is applied. And the more powerful an AI is, the more likely it is that it applies such strong optimization pressure in cases that are "out of distribution", i.e., unusual to what it is normally evaluated against.
The side effects of a very powerful AI not doing what we want could include our death. E.g., a superintelligence might kill humans in order to avoid being shut down, or humans may just be left to starve because it seizes land area currently used for food production in order to use it for data centers instead.
That was because humans were using it with high "desperation" vector causing it to try anything to please the objective. The answer should be to use lower "desperation", whatever that is.
A model which is more persistent also performs better on intended tasks, not just unintended ones. Therefore there is a strong economic incentive to make AIs as persistent as possible.
Yes so I still think it is the human factor which is to fear not autonomous agents. Humans are already using AI's to bomb girls schools. AI in Trump or US military hands scares me far more than in Altman or Amodei's control.
For me a big open question is what sort of progress will we see in robotics. I won't even attempt to speculate, but it does seem hard in different ways than proving math theorems.
In a lot of ways, robotics - navigating and operating in the physical world - seems to be a very verifiable problem. It's fairly easy to verify that a robot moved from A to B, or that it built a structure that completely aligns with the plan, for example.
The main issue is cost and speed to verify, but simulations and world models will help there. I think we'll start seeing rapid progress pretty soon.
AI performance has always been extremely spikey. It's great at some things and terrible at others.
Why do you think the world to date hasn't been taken over by evil genius mathematicians? Can you extrapolate from your understanding of the answer to that question?
I don't think evil mathematicians are very common or that any of them would be capable of single handedly taking over the world if they were. My concern is more about systems that are beyond human level, those kinds of systems would actually be dangerous to us.
I see the recent progress in mathematics and cybersecurity as signs that models are getting more capable more quickly than usual. The companies plans to develop them by recursive self improvement now seems like a real possibility and I don't think they should be allowed to attempt this.
Solving a bunch of math proofs is a far way from recursive self improvement. Don't worry, it's not like a tech tree in a video game where if you can prove a bunch of theorems then suddenly you unlock the next level of technology.
Machines are already far beyond human capability in plenty of ways. Including cognitive tasks like chess. We've already created the technology we need to destroy ourselves (nuclear weapons), and yet so far (knock on wood), we're still around.
We've even already had programs that can prove (brute force) theorems. As far as I can tell this isn't much different, except the space of theorems that computers can solve has expanded. How far? We can't really say yet.
Does solving more theorems than before suddenly mean computers are capable of anything? No.
Do you think that a tireless, infinitely smart, infinitely evil human would be able to take over the world? I don't. Intelligence is not the limiting reagent in our reality.
I think so, individual dictators have gotten pretty far and they weren't infinitely smart. I think infinitely smart would be enough to extend that to the whole world.
> Why do you think the world to date hasn't been taken over by evil genius mathematicians?
A "mathematician" is a human who decided to spend their lives studying mathematics. Mathematicians also tend to be smart, but intelligence is innate, not acquired, so studying mathematics doesn't make you smarter. This makes it obvious why they don't rule the world - if you want to rule the world you'd want to focus on that (for example, doing business or finance), and becoming a mathematician is just a waste of time.
LLMs don't work like that. Like in humans, all of their capabilities correlate, and unlike a human, their overall capabilities grow over time. Looking at LLM mathematical ability over time* therefore gives you info about the progress of their general capabilities, and ability to take over the world would be determined by the latter.
* In fact it'd be better to look at a mix of different capabilities, but that's growing too at about the same rate, see https://epoch.ai/eci
> unlike a human, their overall capabilities grow over time
This is incorrect. Unless there is some new developments I'm unaware of (entirely possible) LLMs "learn" during the training phase, but after that they are static. They do not improve further or retain information when used for inference.
You might be confused because AI companies keep releasing new models and tinkering with the harnesses, sometimes under the same name such that "Zern 6" (or whatever) doesn't always mean the same thing.
Nah, I just phrased it a bit confusingly - I meant the capabilities of LLMs as a technology (equivalently, the capabilities of whatever the frontier model is at each time) grow over time, even though any particular model is static.
I completely get the doomer POV, but we've somehow navigated all the previous "dangerous" technologies we've created - electricity, phones, internet - every one of those had similar arguments and concerns of danger.
The optimists' argument:-
Politics:- in general, I think many of the problems in the world today are due to misinformation and lack of education. What happens when we start routing things through an ASI that brings data and logic to the table? What happens when politicians can no longer lie without being caught out live on air? In the UK, local authorities are being flooded with complaints and requests from people; for example, some are doing AI-assisted investigations into accounting "errors".
Science:- I just don't see how the current rate of progress doesn't end up in crazy technologies like perfectly simulated human cells, organs and bodies to the point where we can run experiments virtually and solve all diseases in the next few years. This is happening. Perfect weather predictions far into the future, likewise with earthquakes, etc. Solar panel research explosion resulting in huge efficiency gains, to the point where people no longer need to plug their EV in - car surfaces will be covered in solar panels, as will our windows and roofs. Connecting new homes to the grid will be optional - the same way landline phones are no longer a thing.
I just find it very difficult not to extrapolate all the above.
We got this dump of mathematical breakthroughs from one small team in one company with access to this technology. What happens when this SOTA model is available (and it will continue getting better and cheaper) to everyone working on hard problems - every university on the planet starts cranking out AI-assisted research breakthroughs.
> What happens when politicians can no longer lie without being caught out live on air?
If there is perfect lie detecting technology I could see all kinds of chaos resulting from it. I can't see it only be applied only to politicians, and I think it would be the developers of the technology who decide the use.
I think were we disagree is that you sort of see AI as an extension of technological progress whereas I see it more like an extension of evolution. I view the process of AI training as functioning in a similar way to evolution in that it build circuits into neural networks similar to how evolution built circuits into human brains.
>> What happens when politicians can no longer lie without being caught out live on air?
A 5-second delay on a politician's presser. Any lies will be muted in real time and the actual facts presented onscreen. Continue to lie enough, and the politician gets unstreamed.
The technology can keep going for a long time in verifiable areas. For non-verifiable areas it's going to have a hard time progressing past where a committee of the best human experts in a field would land. For stylistic areas, whatever the AI doesn't do will have cachet because it will look expensive, sort of like how the kids these days view the ugly old school metal braces as a status symbol because you have to pay for them out of pocket (even as by past standards it'd be truly exceptional).
I'm a transhumanist, I want to build god and kill death. To me this looks like we are moving in the right direction.
So basically I think that the future is getting pretty weird because we are building really powerful tools, though these tools are precisely what allows us to prosper in that future.
It's very unlikely that a superintelligent AI will create unlimited prosperity for everyone on a finite world in a short amount of time. You may not be among the lucky ones.
Or, conversely, they are the exact tools that will allow the powerful folks to not care about the peasants anymore. History tells us what happens then.
> I'm a transhumanist, I want to build god and kill death.
This is good and admirable, but it'd really suck if by trying to build god without knowing how we end the human species. We could simply wait some more decades until we actually have any idea what we're doing, and then do that without the risk.
I think we crossed plenty of lines were we will not get back to.
Software development for example as a task is done. And AI is continuesly reducing the price of more and more tasks every day.
This math breakthrough also shifts something significant: Its now a lot clearer that investment means money into energy to run AI.
Money + Energy = progress
I don't see it plateuing at all. We know how to progress. We broke through a wall we hit. Like the system wasn't able to optimize/automate everything because the tools were not there. It was still cheaper and easier to hire people for a LOT of things.
Now AI fills this gap.
You will see the commodification of everything in the next 15 years. High complex tasks? commodity. Physical labor? commodity.
There will be a lot of job loss unquestionably, in the same way that automation reduced manufacturing jobs and farm payrolls.
At the same time we have to put what AI can do in perspective.
Intelligence is a broad grouping that includes concepts such as knowledge, skill, experience, and wisdom.
AI has incredible knowledge and in many areas approximates experience and wisdom.
But wisdom is harder to formalize than knowledge and skill.
For example certifying a college education relies mostly on the ease with which we can verify/test knowledge.
To some extent advanced degrees try to certify maybe wisdom and experience.
In my very personal opinion, wisdom and life experience should give humans an edge for a while to come.
Additionally, I do feel that the more an individual lacks better than average wisdom and experience, the harder it will be for that person to compete with AI.
Also, on the bright side, the average human will continue to prefer to interact with a fellow human in many spheres. That will also act as an upper bound on AI and robots taking every job.
Either way, I do think this transition will be painful. I don't feel it has to be apocalyptic.
But the world has been an especially volatile place over the last 10 years.
So when you add that existing volatility, to the upheaval from the AI transition, it would not surprise me if the transition results in violence.
But, without the pre-existing volatility, and if humans were capable of generosity and love at scale, I see no reason AI can not be absorbed into society with net gain.
I guess to summarize, I feel this tech should be a net gain and to the extent that it isn't, it will be because of flaws deep inside of humanity itself, not because it had to end in chaos.
In other words, I feel fear, greed, anxiety, and competition -- all our base instincts coming from all sides, will be what determine the end result of AI moreso than AI taking everyone's job.
I have a hard time taking statements from OpenAI about their own product, that they are trying to sell to people and make money, seriously. I take these statements as they are greatly exaggerated or even straight up lies and propaganda.
That said, I think results like these are mostly annoying more then anything. They spent a lot of money, used up gigartiuan amount of compute, to ruin a puzzle that mathematicians were tackling. I am mostly unsurprised that if you spend a trillion times the energy that a team of mathematicians would, that you get maybe 1.5 times the results. I see a future where that 1.5 times the results may go to 3x but not much more. And if that, then I will be more annoyed.
Yes it is intentionally hyperbolic. I know the factor is several orders of magnitude. I don‘t know the exact, nor even the ballpark. I just know this is a ridiculously large amount, so I may as well pick a number large enough that people know it is an exaggeration.
did we always know that computer can do the thinking for us if we allocate them enough resources? i don’t think so. so even if computers are more expensive than humans the fact that they can play the same game is surprising and (relatively) novel.
I don‘t think it is this simple. I think there is a subset of problems (namely ones that can utilize automatic solver or some other kinds of automatic testers and verifiers) where reaching the solution is correlated with the spent energy.
Maybe people will find some clever way to expand this domain of AI-solvable problems by a couple of more categories, or (more likely) find a clever way of using applying these verifier for problems that was previously not viable, thus changing the solution to “just spend more energy computing dummy”. However I think this too will have its limits.
Regardless, this is still annoying and I want them to stop doing this. Solving math problems should not be relegated to whoever has the most money to spend the most compute.
I think giving the world a "cheat code" to accomplish too many intellectual tasks will eventually take its toll on us because after relegating most of our physical labour to machines, we could at least marvel at the abilities of individuals.
Yes, human beings can do more now in some ways but...to put it poetically, I think there will be no more heroes like Einstein and Newton of the past. Now it will just be someone cleverly turning the crank.
Yes, we still admire Usain Bolt even though we have cars...but maybe the admiration is a lot more trivial than if we did not have them....
“ Yes, human beings can do more now in some ways but...to put it poetically, I think there will be no more heroes like Einstein and Newton of the past. Now it will just be someone cleverly turning the crank.”
This is false… there’s lots of ingenuity to be had and demonstrated. But it’ll only get recognised if it makes a material contribution to the economy imo. Otherwise yes it’ll be seen as meh - but that’s already happening.
People like Einstein were revered in society. The average person cannot name a leading scientist etc today.
Arguably, I think AI would not even exist without capitalism because it's only the arms-race scenario that has made it somewhat viable. Otherwise we wouldn't be foolish enough to waste energy on this shit.
> The average person cannot name a leading scientist etc today.
When Jane Goodall died last year it was international news. She was a celebrity scientist for sure, I think she even made an appearance in The Simpsons. Ditto Stephen Hawking.
International news doesn’t mean much - the vast majority of people don’t consume news the way you think - I highly doubt the vast majority had any awareness.
I think physics will be the limiting factor. Even if something recursively self improves, it will hit a physical wall allowed by circuits, batteries etc. A lot of the fear is that there’s an upper bound we don’t know about, whether it be time or energy, that allows a fast takeoff to occur fast enough that we dong have time to see it coming. I don’t know about that… so I’m not worried at this point.
> Thus at most one informative i. So cheater chooses arbitrary g_{v_i}, on exact duplicated input matches and passes, independent of actual satisfiability!
No idea what it's so excited about, but it's cute that it "is." I for one welcome having access to a math buddy 24/7 that's way above my level but also always "willing" to talk at where I'm at.
Beyond the reasoning capabilities of the unreleased model and that they can run a large number of agents in parallel, what makes me curious is that the writing of the proofs is quite human-readable. This is in contrast with the scientific text produced by the ChatGPT available to us, which writes horribly in a way that no human would write. One tell-tale is that they constantly attempt to be defensive and cover all edge cases like division by zero etc. that are clearly a non-issue for humans in some proofs or at least a human would not add this to the main statement, but AI is so overly careful that makes reading its proofs impossible. On the other hand, these new OpenAI proofs are very good.
This is what really made me think to my self, "holy shit". I can't believe not more people are noticing and talking about this. Unless perhaps they didn't actually read the README, and are just talking about what they heard from someone who also didn't read it?
Wikipedia just told me there's a galactic algorithm for integer multiplication in O(n log n) based on FFT so I'm guessing those two proofs are related.
It fascinates me that there's something like this in something as solid and rigid like matrix multiplication. What causes something so rigid to break apart and "leak" at very large scale? Why does the "optimization" appear to be very, very small? Why does galactic algorithm exists? I can't imagine long division suddenly breaking apart after a billion digit, the structure seems very stable? I have heard before that matrix multiplication is apparently optimize-able at very, very large scale.
Does anyone have an intuition to what causes it? What happens at these large scale (or very small)?
One way to think about it is that the classical algorithms are the ones that are fast for small numbers. Galactic algorithms often work for small inputs, it's just that to be faster you need big inputs. A common case of this is a requirement that log(n)<<klog(log(n)). If k=100 then this algorithm will take huge sizes to win
Then 'n' means kind of different things for sorting vs. multiplication though. For example for sorting we assume constant time comparison, which doesn't make sense inputs of O(n) bits
If you sort n k-bit items for a total time of O(nk logn), that scales more poorly in n than multiplying n-word integers. Of course if k is constant you can do radix sort, but I genuinely don't know under what conditions radix sort is more/less galactic than this multiplication algorithm.
this alg is way more galactic than radix sort. radix sort often wins in the hundreds of elements. the nlogn multiplication requires numbers with more digits than atoms in the universe (although that could probably be brought down a lot)
Ah thanks for pointing this out, for some reason I had always equated radix sort and bucket sort (with 2^k buckets) in my head. But I learned today that this isn't true!
At least put a disclaimer for the ad for this site, and maybe disclose how you came up with a total ordering for "top" open problems (vibes)?
> How problems are ranked. LLMs compare pairs of problems. A reliability-weighted model combines those judgments into the ranking, with calibration across model families. The model-family weights are OpenAI 1.00, Claude 1.00, GLM 0.95, and DeepSeek 0.90. These are modeling choices, not measured probabilities of correctness.
The most interesting for me were the faster matrix multiplication, integer multiplication, and FFT. Maybe just cause they're easier to appreciate.
There's also one that says that forced Navier-Stokes can implement universal computation (so, is Turing complete). I don't think any of these are resolving open problems per se, but they're interesting for other reasons.
I thought it was interesting that it said "This paper was written with human assistance", unlike this other Quasi-Riemann Hypothesis preprint that didn't have the same disclaimer.
Funny that it says "written with human assistance" instead of saying "written with AI assistance". So we're assistants to the machines that we have created.
Number theorist here. This is a massive big deal, and would likely be a Fields Medal for a human if a human had done it. But it is an exaggeration to say it is the biggest result in 200 years. At a minimum, it is hard to argue that it is a bigger result than the proof of the prime number theorem in 1896 (which this is a strengthening of), or Riemann's original 1859 paper where he laid out the zeta function and its analytic importance, or Dirichlet's proof of infinitely many primes in arithmetic progressions which is the late 1830s.
But yeah, this is still a very big deal. Among other things, it will drastically improve all sorts of Rosser-Schoenfeld type results for the PNT and that's just a start. For comparison, I have a paper form 2018 where this result would cut 3 pages out and make the full result cleaner and much tighter, and there are likely hundreds of papers like this.
I am also an analytic number theorist, and I disagree. Not only do I think Fields Medal is an understatement (Fields Medals have been awarded for far less than proving quasi-RH + no Siegel zeros), I don't think it is unfair to say that this is a bigger deal than the 1896 proof of the PNT.
As for Riemann's memoir, it's hard to compare. You could argue that was "just" noticing a connection (between number theory and Fourier analysis) that nobody had noticed before; in fact this is the kind of thing AI is extremely good at. I'm being a little cute here.
I think if a human had proven just these two results in the form of a uniform zero-free region for L(s,chi) from nothing as OpenAI did it would not be unfair to say that it would be the single greatest advance in math (easily dwarfing Wiles' FLT), and it would instantly put them in the ranks of greatest mathematicians of all time. Unlike something like Navier Stokes there wasn't a semblance of a research program, experts basically considered this hopeless and would have said the chance of seeing a proof in our lifetime was near zero.
For some comparison, Yitang Zhang's bounded gaps result might have gotten him a Fields Medal if he was not disqualified by age. When it was floated that he might have proven Siegel zeros don't exist, it was considered (by experts) clearly a much bigger deal. This result blows that out of the water (it's a way better version); at least analytic number theorists I talked to thought it was plausible but unlikely that Siegel zeros would be eliminated in our lifetime but thought RH was basically hopeless.
While some of my work is in analytic number theory, much is in other subareas, so it is possible I should defer to you on this.
It seems to me less than PNT in terms of what can we actually do with this. Many different areas of math use PNT, and from my standpoint, PNT is helpful not just for what it implies directly but because it lets us make really good heuristics about whether some sets are infinite or not, and what their rough size is. (Granted, one can do that also mostly via Chebyshev). For those purposes, this doesn't really enter in. Similarly, PNT feels like a statement at least I can say explain to my mother without any technical details. This isn't that. But that may also be my own biases of wanting things to cash out to very concrete statements about the integers.
I agree that one striking element is how no one saw this coming. This isn't building on an existing research program, which itself is remarkable. And last night, before I went to bed, I saw a conversation between a bunch of analytic number theorists who seemed to think there was potentially some slack in the quasi-RH argument, which if that's the case means this is going to go even further.
Thank you for the detailed explanation. From what I'm reading from a lot of mathematicians there's at least a dozen of results here that are field-definining and worthy at minimum of a Fields medal.
I guess the biggest news are not the discoveries themselves but how they were found and that math is going through the biggest revolution as a field since almost ever.
1896 PNT is basically 1859 Riemann + a trig inequality.
1830 Dirichlet's result is qualitative only, it shows infinitude but not the asymptote in terms of the zeros for it predates Riemann.
To me this is the first substantial step after the 1896 PNT, and we really do not see much progress in the whole 20th century.
Personally so far there are only two people worth mentioning,
- Euler, introduces the real zeta function and Euler product, establishes the functional equation at (half?) integers.
- Riemann, introduces complex analysis ideas to the zeta function.
And of course this result if it is true. This is first to penetrate the critical strip, which nobody had any idea how to approach for over a century and a half.
" I expect, say, 2026-level AI, when used properly, will be a trustworthy co-author in mathematical research, and in many other fields as well.
Then what? That depends not just on the technology, but on how existing human institutions and practices adapt. How will research journals change their publishing and referencing practices when entry-level math papers for AI-guided graduate students can now be generated in less than a day—and with the far better accuracy of future AI tools? How will our approach to graduate education change? Will we actively encourage and train our students to use these tools?
We are largely unprepared to address these questions. There will be shocking demonstrations of AI-assisted achievement and courageous experiments to incorporate them into our professional structures. But there will also be embarrassing mistakes, controversies, painful disruptions, heated debates, and hasty decisions."
I was at the workshop that resulted in the Leiden Declaration in Fall 2025. The majority vibe was that this was inevitable, but hard to predict whether it would be in one year or 30 years.
I guess they showed this to the advisory group they created. I guess the group tried reading the work for a day and they could only think of telling them to release the results to the community. I now understand why the group had this suggestion.
There was a Wired Magazine article from either the late 90s or early 2000s that made a prediction that this sort of thing would eventually be possible, likely within my lifetime. I believe the context was "distributed computing" models of the time, like SETI.
I've never been able to find that article as an adult, but I would love to know who wrote it.
Gina Kolata allegedly in the New York Times in 1996 on the Robbins conjecture (noting that computers had started to contribute to math research in some sense), and a longer piece in Math Horizons by her the following year ("Computer Math Proof Shows Reasoning Power"). I didn't immediately find the NYT article, so I don't know if it might be a hallucination.
I was involved in some distributed computing stuff in the late 1990s and early 2000s and I don't really remember people in that community talking about proofs but there may have been a "if we had a mechanical proof-checker, could we do distributed searches for valid proofs that it would accept?" conversation somewhere at some point. There were definitely volunteer distributed computing projects working on pure math; I remember the Optimal Golomb Ruler search (https://en.wikipedia.org/wiki/Golomb_ruler). So, that could possibly have shaded over into "can we find proofs this way too?". At the time it probably would have been based on brute force searches through proof space rather than clever optimization, though.
The idea that you can lexicographically list all proofs in some formalism and then mechanically determine if any is valid is quite clear from Gödel's construction of the function Bew in "On Formally Undecidable Propositions", but he points out that you don't know where to stop because you don't know how long a valid proof would potentially have to be (so "is this a valid proof of this claim?" can be decided mechanically in a limited time, while "is there any valid proof of this claim?" can't be! maybe the shortest valid proof is 49 steps long but you eventually stopped checking after looking at all 7-step proofs, or something).
Thanks! Yea, I have used various LLMs to dig for this article, as well as Google search multiple times over the past 20 years. The article I'm remembering was 100% prior to Nvidia's CUDA release in 2007. My best guess is that it was from late 90s, but possibly early 2000s.
The article I'm remembering was not just about mathematics, but indeed all of physics and related fields. I believe it speculated that eventually distributed computing models could essentially take the world's mathematics and physics formulas and various datasets that we believe to be accurate with high degrees of confidence, and then look for patterns or trends, and then from those trends, mathematicians and physicists would be able to investigate further. Not dissimilar to Folding@Home and SETI@Home.
Keep in mind, that this is the best I can remember from 30 years ago, and I've thought about it so frequently that I am certainly misremembering some of the details. Anyways, it's always been this really compelling possibility, and I wish I could find that article that inspired me so long ago and re-read it! :) I really think it was Wired, but it's possible it was Popular Mechanics, or even an expert guest on TechTV who gave an interview. Hard to say for sure, but I've always thought it was a Wired article.
Oh, I remember hearing about "computer scientists" or something that would attempt to determine physical laws on the basis of empirical evidence, possibly also in that timeframe. That might be another thing to look for. I'm sure that's something people were writing about.
Edit: with the noun-noun compounding being different from the usual interpretation here, like "scientists who are computers" rather than "scientists who study computation"! Maybe "computerized scientists" or something.
It's really impossible to predict which discoveries will "matter", have a direct impact on other fields, or an impact in making other mathematics or physics discoveries.
Only after a world's worth of experts look at these results and then mull over if and how their own fields are impacted by this new info will we be able to answer this question.
It can take decades to really know the full significance. You know, the whole "We stand on the shoulders of Giants", well the Giants just grew a few inches all at once.
worth mentioning that "NP" is not "non-polynomial" but "non-deterministic polynomial (time)". If NP was non-polynomial time then NP != P would be trivial (and in fact, P != EXP is known by the time hierarchy theorem).
Non-deterministic can be explained in several ways. One is in terms of a hypothetical "nondeterministic Turing machine" with certain non-physically realizable properties. The easier way is that a NP problem gets as input not only the problem instance x, but a "witness" w, that may depend on the problem instance. This witness generally makes the problem of deciding the problem instance straightforward (e.g. for SAT, x is the SAT instance, and w is a description of how to set the variables so that it is true).
To be fair - there are statements in math that are independent of the axioms. For those statements, the universe you find yourself in can pick either version (true OR false) and still be consistent.
See also: noneuclidian geometry and axiom of choice.
Being able to solve NP hard optimization problems would enable progress in many areas of science and technology. For example it would allow us to find poly-sized Lean proofs for theorems efficiently, since proof verification can be done in polynomial time.
It would also be amusing to annihilate nearly six decades of proofs that assume P!=NP.
Leans proof checker is not polynomial time, unfortunately. It is super exponential. Basically, because it can verify the result of any function it can prove to be total.
to depress you even more, it is consistent with everything that we know that P != NP and that cryptography does not exist. So there is a worst of both worlds, and we cannot rule it out.
Of course, if we get ridiculous polynomials it doesn't mean much in practice. People who hope for P=NP generally hope for nice polynomials O(n^3) or something like that at worst.
Shameless self-plug: I created an autonomous math researcher. It already solved a 12 year open problem in proof complexity which lead to a publication (and proof complexity experts are already working on simplifications and generalizations of the proof, as I've been told by one of them). This publication is an important step in Cook-Reckhow program in answering the NP vs coNP question.
The autonomous researcher records every research cycle in a public notebook.
"Pre-print" implies it's headed to be "printed" by a publisher. That is, the paper has already been accepted for publication by a peer-reviewed publisher, and it's just being posted early for wider and faster dissemination or to stake a claim of priority.
Otherwise, the document is a self-published manuscript, which doesn't carry the authority implied by "publication" or even "pre-print".
Let's hypothetically say I'm a PHD student who is half way through my studies and I have a halfway written version of one of these "preprints" - what do I do?
Seems like a lot of PHD students are doing to have to pivot the entire structure of their PHD studies? Or just produce something which is already written by OpenAI?
As others said, it's not that this isn't a phenomenon that is unique to now. It happens. I had to pivot a bit of my dissertation near the end because at a random conference I spoke with a researcher from another continent and realized one of my ideas was already out there in some form. I just missed it since it was in a conference proceedings outside the usual set I looked at. So, I had to scramble to adjust and still come up with something novel. I survived, and defended, but it wasn't that much fun at the time.
What isn't so normal is the probability and ease by which this kind of thing can happen today versus decades ago when I was in school. As OpenAI said, it only takes a few hours of compute to do what likely was much more than a few hours of human effort. The only reason this kind of scooping/overlapping was rare was mostly a function of how fast other humans could do the same work. With machines, that totally changes the relative pacing between the human trying to learn how to be a researcher and the machine that can grind out results.
I'm less worried about the phenomenon of overlap and scooping and such. I'm more worried about the long-term impact on fields (not just math), especially considering the early stage students and researchers entering the pipeline now. I'm not sure what happens to disciplines when that pipeline stalls.
Interesting. If you don't mind, could you please share a little bit about that ? You already finished your dissertation ?
It was about the works of Doudna and Charpentier ?
No, back in the late 90s and early 00s, people were trying to engineer custom nucleases and transcription factors, my work was on doing molecular dynamics simulations to optimize TF sequence specificity (similar to engineered zinc fingers) for gene therapy. I wrote up my dissertation and published it in 2001, and then went off to find enough compute, IO, and smart people to make it happen (https://research.google/blog/groundbreaking-simulations-by-g...).
My approach would require custom engineering for every different sequence we'd want to target. With CRISPR, you just "program" the system with a guide sequence, you don't need to do massive engineering to solve a protein design problem.
So it's not the situation they described at all, you didn't waste time during the PhD having to scramble to change topics as it happened after you were done.
That, and also it's just a completely different approach which might later on turn out to be useful. People should remember that artificial neural networks were developed decades before they were useful. People were doing all kinds of other approaches to ML like support vector machines before advances in hardware made deep neural nets feasible and therefore interesting again. ANNs were never obsoleted by SVMs.
I have had this conversation with my PhD students yesterday. I am 100% sure that all of their problems can be solved by publicly-available models now (I solved a case of one myself as a test, it took 15 minutes). So the challenge for them is to see how much they can accomplish in their allotted period, and still pass a defence on at the end of it all. The PhD defence is going to become all about a test of understanding, not a test of quantity of publication.
A much bigger issue is: What is the point of any research mathematician publishing anything now? I really hope that one positive effect of all this will be to finally topple the awful peer review model we currently have, with the biggest publishers gatekeeping with extortionate fees.
This happens all the time, even without AI. Other researchers or PhD students can publish the same results before you. I say that based on my experience during my PhD.
Yes and no. It's pretty normal for students to get scooped, sometimes even multiple times (ask me how I know...).
It's much less common for students to have their entire thesis direction removed from under them, as might be the case for someone working in fine-grained complexity assuming 3SUM has no subquadratic algorithms, or working assuming ~UGC. Both of which (publicly) seemed like perfectly valid research directions until a couple hours ago.
There might be stuff to salvage from their conditional results anyway, but this is not your average scooping.
In other words, you've already taken a gambit with the first half of your PhD, now take a second gambit, praying that you have something to publish by the end of your PhD.
It's worse for those who will face the job market (or maybe tenure review) in the next 2–3 years. If you have more time before you have to justify your continued employment, you can pivot to something else. Theorems and proofs may be cheap now, but there is a lot of value in figuring out what is worth studying.
This has always been a challenge for PhD students and researchers, it's just far more likely to occur now it seems.
Getting scooped doesn't feel good, but it's a signal you've been thinking about things other people care about.
I think eventually companies won't get as much stuff that's usable for marketing, so they'll stop investing so much into ai for math, so eventually cheap and poor graduate students will be able to do relevant work again without worrying about getting scooped by a company with a million GPUs :-/
The follow up question then naturally would be how do phd advisors with people whose fields are in someway premised on making breakthroughs in theoretical fields that AI can solve work deal with it?
At my (German) university, the solution would have been to change from a "cumulative" thesis (which requires peer-reviewed papers) to a monograph-style one. Because the general rule was that your undertaking must be novel ar the time you submit your thesis topic, not necessarily at the tome the thesis itself is submitted (years later).
Precisely what all NLP researchers and the ML community at large did in the last few years: embrace the frontier and realize that attention is all you need.
Frankly, I believe PhDs don't have to be novel in the absolute sense, only new research to the student. I can't remember how many times I've effectively invented something from a clean room approach only to realize someone already published something years ago or that my "new" algorithm has a name. So this really shouldn't be holding anyone back.
That would be unfortunate but the world does not owe you anything and is not going to stop for you. Which is also a valuable thing to learn in your 20s (ideally earlier).
> Let's hypothetically say I'm a PHD student who is half way through my studies and I have a halfway written version of one of these "preprints" - what do I do?
I did get my PhD...before AI. And my honest advice (to myself back then, even) would be: quit the PhD, become an electrician, and work hard to buy a tiny house in the middle of nowhere to watch the world burn in this madness.
These results are wild. Several individual findings are crazy good and use mostly unexplored methods (the improvement over Riemann for example)... I'm pretty sure some of these results would have been Fields-worthy.
But having so many of them at once? Damn. We really live in the future.
This news is exciting and sad at the same time. I've heard that AI chess programs sometimes have blind spots or quirks that human players don't. I've also heard that human players are learning from AI's playing styles (essentially human and AI evolving together). Maybe something similar will happen in mathematics.
After how poorly OpenAI and Anthropic handled the previous cases, I approve of the more measured and cautious approach this time.
We cannot have them rushing to publish amidst tons of confusion, rumors of threats/scooping and outright plagiarism of existing work (by failing to cite said work).
If they're going to participate as scientists in these more rigorous fields, they're going to have to match that level of rigor, not lower it to the disastrous low that ML research publication is at.
I'm not referring to just that. For example, Anthropic published a half done report about some biology research that turned out to have already been discovered and patented, and IIRC both companies are guilty of claiming results without doing the basic diligence of citing the material their work builds on, effectively passing it off as entirely done by their AIs.
In previous math sharing there was speculation about stealing human researcher's results or progress, via prompt inputs from those researchers, and sharing that as their own result. They're adjusting their process, and it seems ok.
Let’s be very clear, the alleged “stolen results” were largely the product of another LLM, not de novo human work. Also, it was false - they did not steal the results.
No clue why I'm downvoted for this, HN struggles with truth-seeking on these topics.
What’s the nuance? Two researchers alleged theft. OpenAI investigated and confirmed that their research was not in the model’s training data. The world decided to take the first part as objective truth and ignored the second.
> OpenAI investigated and confirmed that their research was not in the model’s training data.
You mean when they say it was impossible to confirm anything but one day after it was 100% confirmed that there was no theft? And here I'm not even talking about all the ethical problems related to trying to scoop another group when you hear they're close to success, or how current solutions follow extremely closely human-generated ideas, or about the lack of relevant citations in OpenAI's paper.
Believing that OpenAI's claims have any substance cannot be explained by naivety alone.
They didn't try to scoop another group; they thought the other group had already solved it, so maybe their latest model could take a shot too... and the model solved the full problem when the other group had not! They found out after the fact that the other group had only solved an important sub-problem.
The controversy was whether they had plagiarised that other work on the sub-problem, which they categorically denied after an investigation. And yes, a few days to investigate something like this is reasonable for a company as big as OpenAI. Having seen how data infra is set up when petabytes of data are flowing about, there are thousands of entwined data pipelines to figure out. Not quite as easy as running a query on a sqlite DB!
> Believing that OpenAI's claims have any substance cannot be explained by naivety alone.
Yes, they could be explained by a GitHub repo full of proofs :-)
Or are you suggesting there were hundreds of researchers who just happened to be close to solving hundreds of these long standing open problems using Codex, and OpenAI swooped in plagiarized them all? ;-)
> “Following an investigation, we have confirmed that Buckmaster’s Codex prompts over the two months preceding this announcement and paper on September 8, 2026, could not have influenced the system in any way, including through training.”
Buckmaster himself said essentially all the progress they made was from July 15th onwards (with a new model on the problem) with no real progress prior to then.
I think it’s really cope to claim this was a human result being stolen.
I really don't get how people are still continuing with this "stolen results" narrative after today. Like NS was kinda insignificant compared to treasure trove they released now, thinking that the LLM needs to "steal" from some human is simply coping.
Right? Math is the most open of our academic knowledge institutions, by virtue of what it is. It's easy to get any math publication, and I am not aware of any other fields where an anonymous person can publish their work informally in an anime discussion and enter the annals of math knowledge.
This is a tricky one, but I do sort of agree with you. However OpenAI doesn’t give most researchers access to the models which produced the work. So the gatekeeping goes both ways, I think.
The gate keeping, I'm afraid, will be now in the hands of various bubecks, responding directly to even more disgusting people.
With all the hierarchy present in mathematics, I would prefer it by far.
This thing named inappropriately "OpenAI" goal is just grabbing and monopolizing. Capitalists before could not really touch the deep of the human spirit with their filth, now they can.
I'm a software engineering/biology/ML guy who loves when clever math ideas get turned into real solutions (https://en.wikipedia.org/wiki/Compressed_sensing). I am curious if any of the results have immediate applications in any kind of engineering or science.
It's fine if not, but it'd be great if even just one of these helped us solve a long-running problem.
I was talking with a number of physicists this evening, and of the papers who's problem statement we understood, I don't think any of them will have near-term applications. For most of them, they were results that I think the physics community believed to be true, and the paper provides the first rigorous proof. For several it was news to me that they weren't already theorems! These results are important advances in mathematical physics, but they don't tend to have much impact on experiment.
From what I can tell, all of the physics results here are quite mathematical. But I am very curious how the internal model they used would perform on more applied problems.
It feels odd to me that they wouldn't prefer applied problems. Seems like an easy way to profitability. Probably based on what attributes they're looking for in a problem when picking them.
At first this was my take as well. That plus, well, maybe they are just on a serious PR kick with maths. But I am starting wonder if they have determined, or strongly suspect, that the road to exponential model improvement must first be paved with extraordinary improvements in math. Like in some sense this seems like a test case for where their true intensions might go: vast improvements in the efficiency / size / speed of models and their training. Hard to imagine trusting the models in all those spaces without first trusting them / training them to address new or unsolved math.
You can't really profit from proving theorems of applied problems (that are widely regarded to be true). Those who need to apply those theorems on real problems would have already done so (and if they don't work in some cases, well, congratulations... you found the counter example!)
Same. I have agents analyzing the papers here to see which of these are actually new approaches, novel application of unusual approaches, etc. If there are new intuitions and ideas, those are the mostly powerful reusable components by my estimation.
I think they should put human names on the papers as someone who has reviewed the result, even if just a preliminary review. (I’m assuming they didn’t just pipe their model output directly to the internet and these had some amount of review?)
I appreciate you acknowledging the politics behind this. OpenAI does not intend to be a software company for long, they intend to be scientific infrastructure. They'll want to be faucet from which pours embryos, orbital calculation, geothermal/substructural rating, and etc.
Looking at the quality of the writing in these documents (or at least, the ones that relate to problems I've worked on), I do not think humans reviewed them.
And that means that the recommendations of the reference group has not been followed. The only improvement here is that it is version tracked? No authors, no careful write-ups with exposition.
I think it's a damned if they do, damned if they don't scenario. I think they would be afraid to seem like they're claiming that the selected scientists deserve the praise of an invention or something, when it's "just the machine who did the work"
Yeah, that's probably right. But, it's a "new world" - invent a new standard - "Written by Claude Foo X.Y; initial review by John Doe".
With a deluge of results, having some human expert vouch that it even might be worthwhile would help. (e.g. see the link on HN yesterday, "Two Room-Temperature Antiferromagnetic Semiconductor Candidates" - I see it not worth looking at unless a subject expert vouches for it).
I'd really like some clarity on what that means. E.g. DeepMind has 'cheated' with this in the past, claiming that AlphaZero only took 4 hours to reach super-human chess levels while conveniently leaving out the fact that it was 4 hours x 5000+ TPUs. Sure it's impressive that it only took 4 hours wall-clock but it's very misleading as to cost.
Can we get a number in Blackwell GPU-hours, kWh, or some other compute-scaled metric?
It means the level of effort that a ChatGPT Pro plan summons when thinking. For 3 hours.
If folks are going to analyze this claim with a critical eye, I'd be zeroing in on "average" rather than acting like this measurement is somehow unclear.
Does this imply that it was a one shot prompt with ChatGPT Pro style models (i.e. best-of-N), rather than the agent swarm approach that was used for Navier-Stokes?
It’s incredible and wonderful. Mathematicians in this thread sound very much like software engineers last year, who spent years wrestling with a piece of code and now it just “appears”! But think of the next level that it empowers: new mathematics, new physics, new forms of advanced engineering. What was formerly constricted and throttled fell and a new wide vista is possibilities opened up.
so you said that if there were such a program that could you know provide a proof or disproof then mathematicians will be out of business what really, I mean that you think it would be liberating
Tim Gowers (1:00:10 - 1:01:14):
well that's a very interesting question actually if there were a program that could solve the kinds of problems that we spend our time solving and do it much more quickly than we could then we would be out of what comes with what currently constitutes business but we would it's not completely inconceivable that we could just say we've got this fabulous tool now what are we going to use it for and it's a little bit I don't know I'd want to sort of plant aside what would we do if we had a program that could just answer any mathematical question you gave it to or else if it failed you'd be pretty confident that nobody was ever going to solve it and certainly a lot of applied maths might be pretty pleased with with something like that so what I really mean is that I could just modify what I said and just say it would radically change what mathematicians do or what pure mathematicians do
GitHub is a significantly worse place to store important results than Arxiv. Of course, slop does not belong on Arxiv, buy slop should also not get published.
Specifically, (n log n)^{1 - 10^{-13}}). There are likely no practical industrial problems of a size that would benefit from that specific reduction, but just breaking the nlogn barrier suggests that there may be more and better fruit here in the future.
As social commentary, I think a lot of people in this thread are expressing interest in and engaging with this level of math who might not have pre-AI.
I bet, for people who don't understand these problems or their solutions but are close and are now interested, AI makes then considerably more accessible than they would've been previously, and behind this big visible wave of results there actually will be (or already is) a wave of improved comprehension by a lot of curious people.
I'm not at all at this level at all, but I did learn quite a bit about polynomials over fields yesterday.
More specifically, a combination of mathlib (human, expert curated) and projects like TauCeti (AI-welcome complement to mathlib). See: https://github.com/TauCetiProject/TauCeti
Mathlib is expert reviewed, but only contains a tiny fraction of all mathematics. So this seems to be a quantity of work a "10,000 agents" approach would be applicable to. Like Navier-Stokes, something to spend a couple million in compute on :)
I think such "process-oriented" fields shouldn't be subjected to mass automation (or at least not till humans have sufficiently leveled up our game and understanding). What I mean is progress in math comes from having gained a deeper understanding of the problem for subseq
Two weeks ago, I heard rumblings that UGC, RL=L, and a matrix multiplication lower bound (even lower than the one here, but so unbelievably lower I think it was a typo), and a few others were about to be proven by someone at OpenAI. I find myself saying "big if true" a lot lately.
Even those on their own were enough to make your head spin. But seeing about 100x that? Geeze.
Not much computational mathematics here, but I do see some of interest in the MCMC community. In particular, 139, 93, and 101 are very interesting. Also (I'm not too familiar), I remember attending some talks on the Crouzeix conjecture 325 attacks, should sharpen some rates for Krylov methods. Obviously need to read deeper, some of these don't have corresponding Lean formalizations.
It's a way to approximately draw samples from a probability distribution. Crucially, it applies even when we only know the distribution up to a multiplicative constant which is a common ailment of many distributions in the computational uncertainty quantification field (not that we don't know the constant, but that it's usually computationally catastrophic to estimate it well).
Wow, FFT faster than O(nlog(n))? I wonder if that will open the floodgates for further improvement or not. I don't understand anything about most of the fields these results touch, but I can say that this in particular is very surprising.
It’s always stated that open weight models are 6 months - 12 months behind. Therefore, do we expect that in a year open weight models will be as good as OpenAI’s internal model at theoretical math, or does OpenAI have some “magic” that will be much harder to replicate for competitors?
Why should such hope exist? When the motor vehicles were invented, humans have lost the speed race completely. Yet, nobody lamented and people still did foot racing for fun.
So will it be with AI tools. If these tools become so good, then it will be used. People who want to exercise their minds can still do so, even if that cannot produce economic value.
Foot racing was never a broad source of income for humans. There is a massive difference. The question will be: what economic value can humans produce in the future with AI tools? There is a chance that the answer is "not much". That is the scary part.
Looking back through history, there has never been a breakthrough like this, one that impacts virtually every known industry simultaneously with the potential for a 10x impact.
Any field where there is no external validator (formal proofs, compiler, rule sets) for AI to leverage and real world evaluation isn't easily digitalized. Add in a bit of physical interaction and it is done.
When all open questions are answered, what happens next? Will the machine stop until humans fully understand everything and come up with new questions? Are there examples already of AI solving a problem we didn't know existed?
Clearly AI can grind on math problems now. It can generate proofs and get immediate feedback.
But what I wonder: can we legitimately grind on Physics or curing cancer? There’s a lot of physical world experimentation that needs to happen to make progress.
There's a kind of sticking point in physics combining general relativity with quantum mechanics in a mathematically consistent way. It hasn't been done yet and is mostly maths so your math grinders could have a go. It's maybe the area I'm most interested in with mathematical AI. I've long had a hunch things are stuck there because the math is a bit hard for human brains.
>But what I wonder: can we legitimately grind on Physics or curing cancer?
If we can enable a feedback loop, yes absolutely. But feedback loops for things in the physical world like these are measured in months per cycle usually.
I am curious to know whether the proofs given by the LLMs are going to provide new hints about related problems that might be solved using the same machinery as the one used in the proofs. Also, I would like to know what is the average ratio between the length of LLMs proofs and the length of a proof that a mathematician can write to explain that proof to another mathematician. It is like a functor between the category of human mathematical concepts and the category of LLM operational concepts used in those proofs.
I read somewhere that if we encountered aliens with lot more advanced technology but if they don't speak our langauge, their tech would be useless to us. for example, human body is extraordinary technology that has alwasy existed with us, but we still don't understand it."
the the fact that we have created intelligence that can do what nature does and can also speak our language is most awesome.
So they set up a whole "Advisory Group on Mathematics and Artificial Intelligence", filled it with renowned academics, promised to listen to the group on "review and communication of emerging results"...and then a week later went nah, we are just going to dump it all on github. OpenAI truly is a special company - in the best and worst way.
I'm not a very good mathematician, but I do know a fair bit about software engineering, and with AI I've been busier than ever. I probably wouldn't be so busy if AI was better at anticipating what I actually wanted rather than making guesses no human would ever make.
This is just a short term problem though. Eventually AI will get pretty good at figuring out exactly I want and it will build that from the start. The requirement of me reviewing the AI output only lasts as long as models stay bad at anticipating my needs, which I don't think will take too much longer.
If you mean reviewing for correctness then no, a Lean proof is a much stronger guarantee than anything that can be provided by any human.
For someone who's goal in math was taking unsolved problems and working on them then it's probably over. Just like in software engineering writing code by hand is kinda over.
Can someone with a math background explain the significance of these and previous problems that have been solved by AI?
Does some real problem get solved in physics, chemistry, biology, materials, etc? Or are these fun puzzles for mathematicians with not much real world impact? Eg solving the 8 queen problem in leetcode.
Many of these latest results are very important to theoretical math. Some are also theoretical physics and CS.
Real-world applications are far off, but developing mathematical understanding does tend to leak over into applied physics and CS.
A cynic might say this is all just intellectual games, and though there's a grain of truth, it's too cynical imho. This isn't like 8 queens where there's no hope for applications or generalizations. A lot of this stuff fundamentally affects our understanding of how numbers and systems behave, what are the limits of computation, etc.
Even if someone doesn't care about theoretical results, it's still exciting that AI has become superhuman in a domain as broad as math. That shows there's potential to be superhuman in other domains as well.
Sure I just have a bad taste in my mouth after those researchers were working on one of the millennium problems for a while, had used ChatGPT for assistance and then OpenAI claimed they had solved it.
Solving hard mathematical problems is a strong signal for capability but would it not be preferable to through all these resources to urgent existential issues such as climate change? Finding technological solutions in those areas would be the ultimate capability signal as political alignment at global scale is almost certainly impossible.
Climate change is a political problem, not a technical problem. Ironically, by raising oil prices, Trump might have done more against climate change than many people who have been actively fighting it.
Mathematical problems are ideal as benchmarks for AI because they have clear problem statements, clear axioms and results that can be verified easily (for lean proofs). I can't blame these companies for using them, although it's unfortunate that human mathematicians seem to become early casualties of AI progress.
I suppose openAI could have focussed their efforts on a subset of open problems that have a clear real world impact and leave aside the more esoteric open problems as a way for human mathematicians to hone their skillset. However, this would have been a short term bandaid. With open models 6 months behind the frontier, any of these problems might have fallen to the homebrewed efforts of enthusiasts early next year.
What is mathematics for? From the outside looking in (I'm a biologist), I have always viewed mathematics as a way to understand reality and to improve our ability to manipulate it. But what I often hear is that mathematics is foremost about human understanding. But isn't that only because it's humans that needed to do the mathematics in the first place? It's not obvious to me that mathematics without human understanding has no value. For example, it might be that P=NP. The algorithms are handed down to us and we can apply them without fundamentally understanding why P=NP.
Mathematics seems to be entering an era where human + machine maximizes performance, much like chess in the 1990s. However, imagine a future where even talented mathematicians are nothing but noise in the machine (as is the case in chess now). A future where AI generates and verifies proofs without humans in the loop. Where mathematics may be beyond human comprehension.
In that future, does it matter that early career mathematicians are inhibited by these developments? Perhaps not. Programming faces the same issue. As AI crawls up the competence ladder, does it matter that fewer people have opportunities to develop the skillset of a senior engineer? Perhaps not. I have no doubt that biologists will face the same problem soon enough.
As models get better, I think a time will come when it's hard for people to even verify the results. In the end, I think the bottleneck will be people.
AI doomers often talk about these kinds of scenarios often, but we tend to assume if humans were given magical math/physics results which we couldn't understand, or given magical pills by AI that cure all disease, we'd probably just take whatever the AI has given us rather than spend years or decades trying to understand the knowledge/technology before leveraging it.
At some point in complexity – especially if we allow our own knowledge to deteriorate because AI can do the hard work – we will stop understanding the world around us. In the same way one day Native Americans woke up and realised they shared the Earth with people who had magic sticks which they could point at someone and kill them, we will live in a similar world very soon too.
What sticks are dangerous, you will not know. Your existence in the future depend entirely on the AIs not wishing you harm, but you don't know how they work to verify their motivations either.
Hopefully the techniques and results here will be in the training dataset for the next models, so that each new release will give us more interesting techniques and results!
It seems that OpenAI has a proof machine that keeps multiplying fruitful proofs!
Not even needing the AI to think that, US consumers of the electric grid are already subsidising the unpaid bill of data centers. Just need those who are in charge of infrastructure to prioritize the AI consumption over humans, and shifting the books to make humans pay more for resources.
A lot of these seem to be proving conjectures rather than finding counterexamples, a lot of people used that to claim that these models are not really smart/creative etc.
I think it's fantastic that they decided to follow the AGMAI advice. Cleaning up their mess will be a substantial endeavour, so I imagine the funding provided to do so will reach well into the millions. But it doesn't look like the press release says anything about how they will fund it at all?
I'm sick of it. This has happened so many times in my memory. Some AI company announces that they did something fascinating, and it turns out it is all just hype and slope in the end. I still believe that LLMs are a dead end. Most people here are basically like "I've no idea what's going on, but I'm so happy and LLMs are so cool". To think that doing enough linear algebra would solve all your problems just feels wrong. I guess I'll simply wait until someone interprets the results and explains what's actually going on.
If you scroll around on this thread, you'll see mathematicians discussing how they worked on one of these probably for a quarter century. Others arguing whether another is the biggest number theoretic breakthrough in 100 years or 200 years. It looks like they have made substantial progress on 4 Millennium Problems now. Unless it turns out that almost all of these results are wrong, what would it even mean for LLMs to be a dead end now? This is a historical day for mathematics and it will not be forgotten.
I think such "process-oriented" fields shouldn't be subjected to mass automation (at least not till we as humans have sufficiently leveled up our game and understanding).
What I mean is progress in math comes from having gained a deeper understanding of the problem for subsequent attack of more problems and IMPORTANTLY applications! Right now the first one is trivially satisfied (given oai maintains some memory across models) but the second one is not! It's generating proofs faster than anyone can validate and so only the model can use these. Consequently if it just keeps doing more theory it's... not very helpful or at least not optimally helpful. This is just bragging rights for now.
More "application-oriented" research would be awesome, where it tries to achieve some desirable effect and then produces relevant theory and experiments around it. Fields like CS, Physics, Chemistry, etc. This would also benefit a wider section of the population rather than the 10 people who understand most of these proofs.
I recently saw a YT short of Grant Sanderson on AI in math and found it (as always) very insightful. But I'm sorry, I never ever find anything back on any of these ad-ridden platforms these days, and perhaps it was a short with content stolen from some other longer content anyway. So, if you feel like getting informed, somewhere out there is some nice content by Grant Sanderson.
Apologies for the rant, I really tried to find it. It had something to do with not being able to predict what this influx of proofs may bring us on a meta level, it could be very interesting. But he also had some critical notes about the missing process and the things found along the way.
I'm a little concerned, but I'm also glad they're publishing quick before the department of war starts making national security claims of related to using the knowledge for cryptography and/or weapons
This achievement feels like a reasonable candidate for the moment where LLMs are officially "more intelligent" than any person. How can we justify moving the goalposts yet again? How can people have grown so numb to seeing advancements that they don't read this as significant? I feel like I'm standing at the foot of the exponential. I am not excited for the future, and I don't see how humans retain meaningful control over the future if we continue on this trajectory.
I have been lurking for quite some time. I made an account to post this, but I honestly don't know what to say. I would like to get off this wild ride.
Would Einstein be successful at running apple? Nope
This seems very hard for people to understand.
It will be painful for many to realise - you should focus on doing something that positively affects the economy. Everything else is noise and many endeavours are transitory.
an AI that is at 100s of Einstein level in every intellectual field imaginable ?maybe, running apple does not require one to be very smart or intellectual
Steve jobs wasn’t Einstein in all the things he knew and understood and yet… apple became a behemoth from nothing (on the verge of bankruptcy).
So…. Yeah ‘intelligence’ isn’t simply knowing and connecting dots is it.
Moreover if running firms doesn’t require one to be very smart - and firms are what society needs for production of products and services which affect our lives - in relative terms, where’s the value add to society in creating an Einstein in a machine?
A firm full of Einstein’s
Is going absolutely nowhere.
llms are good at different things than humans, we are still collectively figuring that out. the idea that an llm is more intelligent than humans at math of all things seems fairly unsurprising.
I feel this moment is one of the last few warnings before things will get seriously out of hand. We need to stop now. Building a superintelligent AI should be considered a crime against humanity.
It's funny how people see "ML" models becoming superhuman at proving mathematical theorems as a sign that we're about to enter the singularity (whatever that means) or that it somehow justifies the valuations of those companies...But what I see is a scenario where companies have spent trillions of dollars on a technology, and the most non-trivial thing that technology can accomplish is a pseudo-form of automation for mathematics and coding! Remember that early on, when this whole bubble started, investors were promised that "AI" would eventually capture >70% of the world's jobs! but it could very well be the case that the only ones they're going to replace are mathematicians(and i'm not even sure about that!)!
> the most non-trivial thing that technology can accomplish is a pseudo-form of automation for mathematics and coding
It's irrelevant that this is a pseudo-automation, what's important is that techbros can convince people who make the decisions and concentrate wealth that this is a full automation. So expect reverse centaurs in increasingly more professions in the future.
This feels like a huge AI slop dump. Who validated these results. Why are they not mentioned. Human understanding is key here.
In my experience AI (frontier models) sometimes does weird stuff that needs human review. Not that’s incorrect but sometimes overly complex language or weird use of language.
Can the mathematics field come out of this stronger and better. I doubt it, things will only get worse. There is no future in being a Proof Digester, no valuable human is motivated in doing that, nor is there any glory in it. But there are some potential pathways for maths to come out strong from this:
1. Relentless focus on quality. Every publication must act as if it’s going to be included in a future textbook, that is a newcomer can get into it given a reasonable amount of time, and math priors learnt in undergrad. (NO AI Slop proof passes this bar as of now)
2. Limit the publications per year. Each author is allowed 2 with a max of 50 pages. This allows the author who chooses to not surrender his cognitive capacity to the machine, still be allowed to play this game. Of course who wants to orchestrate a thousand agent workflows, is free to do so, he is only limited to 2 publications.
3. The aesthetics of the field changes from purely solving the problem to solving the problem with simplest most elegant set of ideas. What 3 sets of simple ideas solves large swathes of problems, that should be given a fields medal, not purely solving the problem, which the AI will be able to do.
> There is no future in being a Proof Digester, no valuable human is motivated in doing that, nor is there any glory in it.
Coding has already gone this way and plenty of people do find motivation and glory in the final product vs. the building of the product (myself included). I absolutely see value in a person being able to humanize llm math output and see the field moving in that direction.
I don’t think you’d be doing this if you were paid peanuts. I also think coding is in a transitional period, most people who claim they are “building the product” are going to lose their jobs as the models get better.
I am genuinely doing it for free in my free time on top of my work (in fact I pay for the models) so I am doing it for less than peanuts. I completely disagree w/r/t the model take, their is skill in steering work to get a product done that is what management has always been about, I do think in the long run this will be automated but I have seen 0 progress on it so far.
I disagree, there’s no future in “Steering” a model. You ask what you want, and the model delivers. Opus 5.5 is getting there already, A single line prompt creates a fully functional game engine that can be used to make many different games : https://m.youtube.com/watch?v=R_uf5OfMGio&t=3555s
(Opus 5.5)
Please explain to me what great prompting / steering you will do when your customer can just prompt exactly what he wants and get a product even more tailored to his needs.
Natural language is too lossy for most usecases and people are bad at being specific enough to get what they want. For example, I am not creative enough to think of movie ideas compared to good directors so I will probably just pay to see what they cook up rather than try to make it myself.
If the end result of this is that within a few years, AI "does" all of the mathematics that humans do, and that there is nobody around that understands any of it, what was the point?
Can someone who is more into math or AI explain why so many people are so incredibly excited about this?
If OpenAI started opening hundreds of PRs on long-open issues on popular open source projects, would we rejoice, or would the first reaction be "they are unreviewed, so slop until proven otherwise" (it would be that).
I cannot possibly see how these are so impactful, especially the ones that don't come with lean proofs.
LLMs have the ability to make millions of mistakes per day, whereas humans can only make so many. How are we suddenly all so confident that there's no extensive hallucinations or "gaming the system" going on?
I said "especially the ones that don't come with lean proofs", but even for those with lean attached, lean is software, it has over 1k open issues, and I would not put it past an LLM to identify a bug and exploit it to pass the gate.
What's missing for me for each result are the following:
* Explain the result to me as if I'm a 10-year-old.
* Create the infographic for this result.
* Make a Khan Academy-style video to teach me this result.
The Advisory Group states in its recommendations [1]:
"We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models."
To me, this is a take against progress so that mathematicians can keep their jobs. What would we do if, instead of math, we were talking about diseases? Are we going to keep diseases around so that doctors can keep their jobs too?
> At present, some frontier AI labs are testing advanced mathematical problems on proprietary models that remain inaccessible to the broader scientific community. Our recommendations are formulated with this practical context in mind. However, ideally, they would not do so. We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models.
To me, the issue is that the models are proprietary which are only accessible to a few people in 2 digits. It's not about progress but access.
Maybe, but it still seems like an excuse. OpenAI has a proprietary model capable of solving these problems and is willing to share the results with the mathematical community. So basically, the ask is to just not use the model and leave the problems unsolved?
I don’t think that’s sufficient. If they want to be good stewards of mathematical research and not just doing marketing, they need to at the very least tell us the datasets and any techniques they used to train this model. IMO that’s probably just as if not more valuable than a dump of un-reviewed results.
They won’t because they don’t care and the only way it got this good is something along the lines of they trained on every mathematician’s codex sessions even if they opted out because they consider the thinking traces or output and metadata fair game.
Why is this important? Do you not think that at some point, probably sooner rather than later, math researchers across the world will have access to similar capabilities?
This is sort of like discounting putting humans in space because only a few nations have the means to actually do it at the moment.
The work is not having AI poop out these docs the work is validating them and publishing them. OpenAI doesn’t care and wanted to race to publish potential findings and let others review it which is lazy and selfish
I don't think this is a reasonable take at all. Most work on maths has no real benefit other than to further human understanding of maths - it's more like an art. Nothing is gained from OpenAI solving all these problems but taking jobs from mathematicians. Other than advertising for OpenAI at least.
It could not be more different from having AI work on disease research etc.
It all hinges on the definition of “progress”. The debate of the past month is all about questioning whether formally proving outstanding unproven theorems without human understanding constitutes progress. This is quite different from solving diseases.
I fear that professional mathematics will wither, and there will be nobody left to digest the AI results of the future, leaving us unable to challenge the AI.
The trope you're drawing a parallel with has a (to me) compelling counter though: there being a cure for every disease wouldn't stop people from getting sick.
Makes me feel disgusted. Super-intelligence, even when controlled, will cause so much damage. We are giving away control to a super entity, or whoever has the power to steer it.
It's not clear from this progress that AI can formulate conjectures despite this new ability to solve them. So mathematicians still look like they have a job. Though instead of spotting far-off landmarks it's sounds more like they'll be chasing waves on a beach.
Maybe for STEM. The humanities seems kind of safe to me since there's an element of it which cannot be divorced from human opinion and interpretation. It's not like STEM where the people pursuing the ends are largely fungible and the ends are objective (Heisenberg already believed scientific discoveries were inevitable and it didn't really matter who pursued them, someone would eventually find them)
Nope, it's about to be turboboosted if anything. I myself have already formulated countless proofs (AI assisted naturally) in the last few months, despite: a) not having access to any academic resources b) not being involved in any academic circles. This research is now being employed in my startup, delivering ground shattering results. I actually ran analysis a few weeks ago seeing how much real world capital I would have needed to cough up to fund my efforts so far and it's literally in the _billions_. With less than a $100k in tokens I have been able to effectively generate the value of Apple or Microsoft in the early 2000s. This is only the beginning. Wait until you start seeing single person NVIDIA startups popping up.
This "value" you talk about, is it past tense? A single iphone probably have more compute then the entire world had in the 1970s, would have cost at that point probably hundreds of millions. You will not sell the iphone for that amount today.
Not sure I agree with this take, but we're going to find out either way.
Have you ever heard of an S curve? Things will develop rapidly, then equalize. If they don't, we're at the singularity and I guess the end of time as we know it.
But I guess really bad things happen, cancer, radiation poisoning, torture, people have died in really horrendous ways, and I guess dying from some horrendous AI side effects is possible too. Yay.
People need to figuring out how to horde as much wealth as possible right now in the next 2-3 years. Jobs especially for knowledge workers are about to disappear.
I’m targeting maximum debt by about 2030. I’d rather have all the stuff I want while we try and build a new version of a functioning economy than have a ton of cash saved up.
It's funny we're possibly entering a golden age of thought, the dreams of the ancients, but we're all worried about capitalism, I think the problem is pretty obvious.
I don't know, I get more work done now in a few hours than I used to in a day, so maybe the working week should be shrunk, that would solve things. The solution seems pretty easy - but then I'm not in the US.
This would be a largely positive solution for humanity, as long as AI remains "just an assistant". I don't think that we are heading that way. It's possible that, at some point, when the time you save becomes comparable to your full working week, your employer (or your clients, if you are self-employed) will discover that you are just a proxy between them and prompting directly the AI. Of course, if your job involves physical labour or requires physical presence or taking legal responsibility for something, you'll be fine for a bit longer (well, comparatively fine in a society plagued by widespread unemployment and social unrest).
I’ve been thinking this as well. I imagine there is a wealth level x such that someone can escape the coming ubi welfare state. Anything under that you are fucked.
It seems the age old academic model of scientists competing against each other for fame and prestige is done, and now we must merely enjoy the fruits of scientific discovery for their own sake.
Seems like OpenAI's next model is better at math proofs than Anthropic's next model (to an extent that it surprised even OpenAI researchers, according to their public comments). But that doesn't necessarily mean it's better at everything else. The models are more spiky than ever before. Wait until they're released to judge.
Perhaps the most significant announcement of my lifetime. Yet, I suspect I will not see this in any mainstream news reporting.
I feel for those in Mathematics and worry for our future.
Models will only get better and in a few years the models which produced these results will be a bad as GPT-3.5 in comparison to what we'll have in the future.
Please take a minute to consider what this means, and the risks it presents us.
Gross. After the accusations training on mathematicians conversations and unique methods, to dump this volume of unsubstantiated papers is at best tone deaf. Are they hoping humans review and validate this? The amount of free that they are coat tailing is outrageous.
Not train on and steal user data to front run frontier research for one thing. Not take credit for and independently solve these problems and instead sponsor researchers and give them credit. They could chose to include humans and frame AI as elevating human systems. Instead they are reveling in the fact that they did this in their own instead of humans. Thats a PR choice that is short sighted and reflects a selfish mindset not deserving of leading this transition
> Not take credit for and independently solve these problems and instead sponsor researchers and give them credit. They could chose to include humans and frame AI as elevating human systems. Instead they are reveling in the fact that they did this in their own instead of humans.
How do you know humans played the part that you think they did in this result? What evidence do you have to challenge their framing?
They are going for super intelligence, humans not necessary.
On one hand I don’t want to be replaced by an AI so I hate that this is happening. I don’t want AI to be controlled by the elites to enrich themselves further.
But on the other hand super intelligence will open doors for humanity. Maybe we will finally defeat cancer or death itself.
"If you have nothing to say, don't post a chatbot's responses, because anyone can ask the chatbot directly if they wanted to"-principle? Well, people can't ask their chatbot directly, because it's not public.
AI may be one of the most communist looking technologies in the classical sense .. i mean it dosn't abolishes private ownership .. but it DOES make intellectual capabilities that were once scarce and concentrated available to almost everyone ...
Can you explain it without reaching for lofty things like consciousness? To me stochastic parrots is literally how it works given that it’s “just” the most impressive data fit we’ve ever done. Apparently generating mathematics reasoning traces + verifying them with lean works super well.
The network goes through 100 plus layers which end up doing all sorts of processing in ways we don't quite understand because it gets there through gradient descent but is probably similar to how human brains do it.
I daresay parrots can be quite smart too but I don't think that's what the critics were referring to.
It is more of the usual though isn't it? OpenAI cribbing off of mathematicians that have used their services; deciding to put a lot of compute behind fruitful areas of endevour; getting results, then taking credit.
With so many results in so many different areas no way they even remotely spot checked well enough.
Prediction: one of these is wrong and this (publicity stunt) will backfire.
Edit: don't tell me about lean. For lean to function as a proof certificate you need to represent the theorem correctly. Again: good luck doing that across such a broad swath of problems.
> Some of the unformalized results could have issues. We will endeavor to fix any such issues quickly. We are also exploring community-hosted repositories for these materials.
If you were waiting for companies to sell engines that never break down you'd still be stuck pre industrial revolution while the rest of the world has been to space.
The question never if something works 100% of the time but how often it breaks and how that fits the need well. Solving one of these problems is a massive undertaking and accomplishment for the best minds, solving hundreds in a month but being wrong about 10% or something would likely not be the death knell you believe it to be.
2023: "AI is useless"
2026: "One of these solutions to dozens of previously-intractable problems at the frontier of human knowledge MIGHT be incorrect"
I'm seeing a lot of this and it makes no sense, the internal model solved these but give one of the papers to Astra and Opus and I'm sure it will have no problem recreating it.
I don't think they mean this from a "verify this paper" perspective.
How valuable it would actually be to share the model with other mathematicians vs just have OpenAI's mathematicians churn out and clean up results isn't very clear to me though as they don't say how much effort it's requiring from their team to prompt and clean these up vs how much it's bound by "time to run the model" or similar.
I wonder if the Lean compiler can change it's license so that a for-profit corporation can only use it if it pays, say, a few hundred billion dollars. This is the correct path.
This just looks entirely obscene from a PR perspective. It's like the cable companies networks owning the networks. At least form some partnerships to obscure the total narcissism party.
ASuperMegaI lab announces they have shrinked the space of unproved mathematical staments from infinity to infinity, giving us proof of their deep mathematical knowledge and understanding, and how they grasp the concept of infinities.
For me this reads as someone boasting about how they go to buy bread on a ferrari to the supermarket, while i sit listening to it, having no idea what they are talking about. And then i stand up and go walking to my favourite boulangerie.
Warning: if you are from the USA you may be triggered by this metaphore.
Willful ignorance is a vibe; did you read any of the GitHub? There are some stunning results in there. I get a similar feeling skimming through the topics that I do watching a successful space launch: it’s pretty cool humans built this. Unlike a space launch we are likely to be able to pass all of this information down to our grandchildren - space launches involve a lot of finicky engineering knowhow, but pure math results tend to be sticky over the last few thousand years. I find that hopeful.
Dont get me wrong, i am 100% impressed by the technical capabilities, and appreciate the significance, and the amount of the results. This is a very special time to be alive, never in my dreams i would have thought to see this.
To follow your methaphore, who is directing the spaceship?
This feels more like fireworks than a space launch. Space launches would not have happened without having fireworks first of course, but I am looking forward for the space launch moment.
I was a graph theory junkie long ago and even moved to Budapest for awhile to study among the greats. While I was there I started working on Barnette's Conjecture which came to occupy my thoughts over the next 24 years of my life, on and off as I worked in many different fields. Last summer I even thought for a few days that I had actually solved it.
But it's supposedly proven here - problem 180. I don't know what to think exactly. I spent thousands of hours on that problem. I really enjoyed it. Hearing that it is solved somehow makes me sad in a far-off way, like hearing an ex-girlfriend died suddenly in a car crash. I don't know, there's probably a lot of people feeling odd emotions tonight.
There's no Lean proof for this one so I'm digesting the paper. On the surface it looks like an approach I considered 24 years ago and abandoned.
I revisited the problem this summer, along with my partial solutions, when the previous round of stunning proofs came out. Several hours of work with Fable simply convinced me it wasn't yet solvable and reinforced how hard of a problem it was.
> Several hours of work with Fable simply convinced me it wasn't yet solvable and reinforced how hard of a problem it was.
This is the part that gives me the strangest feeling about it all, because you're not the only one with this experience. I've experienced this too on different problems, as have many researchers across many fields.
I disagree with the Fields Medalists on the majority of their complaints. AI math is happening and there's no going back. However, on one point I increasingly agree: virtually none of this stuff is possible with technology any normal citizen has access to. I have no problem with AI models making revolutionary advances in math or science. Where I start to have a problem is when the AI models making these advances are tightly withheld, proprietary, and seemingly never released with these capabilities intact. This has been the case for all of 2026 so far.
I suspect that this is in fact the source of much of the angst. None of this progress is reproducible outside of one or two teams inside OpenAI and Anthropic. It's becoming an incredible concentration of power that I don't know that we've ever quite seen before. Right now, it feels harmless because it's being used for wonky math problems that aren't (yet) practical for anything. But great power never stays harmless. History has taught us that countless times, in countless different forms.
> I suspect that this is in fact the source of much of the angst.
Why do you "suspect" this as if it's some hidden motivation when the very first paragraph of the advisory group's statement (linked from the OpenAI post) says:
> At present, some frontier AI labs are testing advanced mathematical problems on proprietary models that remain inaccessible to the broader scientific community. Our recommendations are formulated with this practical context in mind. However, ideally, they would not do so. We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models.
Tao and others in that group have been strongly and publicly pro AI from the start. They are not advocating "going back". They're objecting to the strip mining of open problems using proprietary technology.
OpenAI: At long last, we have created the Open Problem Strip Miner from classic Terence Tao tweet “Don't Create The Open Problem Strip Miner”.
I don't think the strip mining metaphor is appropriate. Mining is a zero-sum game; if I mine something, nobody else can go and mine the same resources I did. Mathematical problems don't go away when AI finds a Lean proof. They create new opportunities for humans to study the solutions, learn new techniques from them, identify promising directions for future research, discover alternative/more beautiful proofs, and write expositions for other humans.
Strip mining is very apt if you view the economics of the present system as "effort -> recognition -> career advancement". Even in strip mining, the resources that had been buried are now available for use in the broader economy. What's no longer available is the living that was to be had digging them out.
The problem isn't effort, though. All of the things I mentioned constitute effort and could be rewarded. The job economy was created by mathematicians incentivizing the proof of difficult theorems above all else and valuing all other work at approximately zero as far as career advancement was concerned. Now they're pulling a 180 and claiming that math was never really about proving theorems, but that's contradicted by their revealed preferences. The strip-mining problem only exists if they continue with the status quo ante.
Mathematicians aren't homogeneous. There are mathematicians valuing pedagogy, collaboration, bridge-building, theory building, along with those that chase the 'difficult theorems', to name a few, and there are lots of flavors within each class, with lots of blending and blurring. You infer that mathematicians prefer the status quo simply because it is the status quo -- with a little thought, you'll recognize that this is a fairly silly notion.
There are myriad circumstances where the values of most practitioners differ from the status quo, which is nevertheless well-entrenched. This can arise from inertia, or from outside forces, such as broader cultural milieu, integration with larger institutions, or contending with economic realities. If you think that these do not and haven't historically played a role in determining the job economy and that math is a pure field where mathematicians could comfortably shape it according solely to their own ideals then you are naive
And, in addition, many mathematicians are graduate students or postdocs hoping to line up a permanent job soon.
For example, if you look at Terry Tao's blog, he has a tremendous amount of first-class expository writing. So, too (to some extent) do junior mathematicians -- but, unfortunately, this tends to not be highly valued by the job market. Grad students and postdocs have learned that to succeed they need to play by the existing rules of the game.
Well, the board has just been yanked from underneath them. People like me can afford the sort of idealism and soul-searching that the parent comment describes, but junior mathematicians face a very unenviable set of circumstances.
1) I never said the problem was effort; I was trying to explain the strip mining analogy, and it's one of the two anchors that make the analogy work.
2) Mathematicians didn't create this economy; it was foisted upon them by the same managerial mentality that brought us "publish or perish" and "the monthly sales quota".
3) I can't tell if you honestly don't get why the strip-mining analogy resonates, or...?
Here's another analogy: if we suddenly discovered personal teleportation, and marathon runners were complaining that it was ruining the sport, would you say "they're pulling a 180 and claiming that marathon running was never really about getting to a point 26 miles away as fast as possible, but that's contradicted by their revealed preferences"?
The strip mining analogy is better though, because it captures the sense of irreversible goal-loss when a problem goes from being "unsolved" to "solved".
As far as I understand, even with "publish or perish", peer reviewers decide what counts as an important enough paper to be published in a prestigous journal, and committees of peers decide whether or not, say, an expository article on arXiv or a textbook counts toward hiring or tenure. Again, as far as I understand, those things have largely not been rewarded in the past.
I like your marathon example, but maybe not for the reasons you intended. The community of marathoners decides the rules of a marathon. You don't need a hypothetical teleporter; you're already not allowed to use a bicycle, performance-enhancing drugs, or shoes that don't fit the specifications. The rules are updated to adapt to changing technology. Yes, I'm arguing that the strip-mining analogy doesn't make sense because mathematics is in the same situation. There's nothing stopping peer reviewers and hiring/tenure committees from changing the rules about which kinds of effort confer recognition and career advancement.
Strip mining is an extraordinarily appropriate metaphor.
Imagine a mine has an unknown number of rare materials. And you know the general location of a few of the most valuable spots. But you don't know what may be valuable right next to it. If the pieces that we know are valuable are suddenly gone, the incentive to mine that particular area drops considerably, dropping the chance to discover potentially brand new materials that would have been found the normal way.
That's an empirical claim. I could equally well say that doing an automated search of the problem space and having a database of results and open problems will identify vastly more interesting and valuable areas. Again, the idea that math is some exhaustible material is a metaphor, not an established fact. I'm willing to change my view as new evidence comes in, but I think we're going to have to wait and see what the landscape looks like in a few years.
> the incentive to mine that particular area drops considerably, dropping the chance to discover potentially brand new materials that would have been found the normal way.
FWIW, I think the metaphor breaks down with this framing. This isn't really a problem associated with strip mining, what's left behind is generally low or negative value (toxic). I'd suggest a different metaphor, from Wikipedia:
> This process involves the removal of all ground vegetation in the area, which is a detriment to the environment.[19] Topsoil may be placed over the tailing along with planting trees and other vegetation. Another reclamation method involves filling in the hole with water to create an artificial lake. Large tailing piles left behind may contain heavy metals which can leach out acids such as lead and copper and enter into water systems.
This feels very similar to the issues with algorithmic problem "mining". It has the potential to destroy the human ecosystems surrounding these problems, leaving barren wasteland behind where nothing can grow or flourish.
I hope sincerely hope they don't currently use "proprietary technology" like:
Wolfram Mathematica ($890/yr)
Magma ($2500/yr)
Maple ($680/yr)
COMSOL ($1500,yr)
Matlab ($500+/yr)
Seems like a very strange position to take, in my opinion.
Why does the field of mathematics suddenly now need to be "fair" and give everyone access to the same tools? Has that ever been the case in academics? It's always been a competition for name-recognition, grants, institutions, etc.
Macsyma / Maxima was an MIT developed CAS system back in the 60's that was proprietery until they sold it off to IBM for a tidy sum. Magma actually has free access if you're in the US, otherwise you pay. That's not to mention proprietary MATLAB toolboxes or specialized Stata modules.
Likewise, a lot of the above packages have pretty sweet site-wide deals with R1 universities. If you're at a smaller, foreign one, you're out of luck.
> Tao and others in that group have been strongly and publicly pro AI from the start
Unfortunately being "pro AI" means relinquishing any control over what the AI, or more importantly the company running it, might be doing.
How is this different from literally any other part of the economy?
We've relinquished control over just about everything we use or consume. We can't compete with larger enterprises for production of food, clothing, machinery, medicine, energy, services. Mathematics is just the latest thing to be industrialized.
What keeps large companies under control is competition with other large companies. This competition causes the surplus value they produce to flow to consumers, not be hoarded via monopoly prices. Do we see strong moats that are going to cause monopoly in AI? I don't see it, and in particular I don't see it persisting if it exists transiently.
You're right, and that's a bad thing. AI is nothing fundamentally new, but its extremity is making many people aware of the truth that's been there all along. There's no contradiction in that.
> Do we see strong moats that are going to cause monopoly in AI?
Ownership of the capital assets used to train and inference new models. Yes, we may end up with more than one firm. But as we see with big tech today, a small number of fantastically wealthy firms in "competition" does not an open market make.
Is it a bad thing? We live in a society. We depend on the work of other people. We are not autonomous. Sure, we can try to be self-sufficient, and that would lead to a subsistence lifestyle much degraded compared to what we experience.
Somehow you have to argue either that society itself is bad, or that math is somehow different from all these other human activities.
I think the obvious fact that people prefer to live in places with large commercial organizations shows they don't really care about that, at least to the point of foregoing the benefits these organizations bring.
No, it doesn’t. You can be in favor of something and opposed to a particular way of handling or implementing the thing. And the issue here isn’t what it’s being used for but who is able to use it.
"AI" is largely a marketing term for a particular type of computer program that uses a statistical language model.
Computers and computer programs are tools. Humans always remain sovereign over their tools.
And incentives are sovereign over the humans. The humans leading the AI labs have every incentive in the world to move quickly without any restraint.
The second law of thermodynamics always wins.
Eventually. In the meantime, here in the human socioeconomic sphere, you might be dealing primarily with the Second Rule of Fight Club and Operation Mayhem.
I am not sure how convinced I am by that argument. A gun is also a particular kind of tool, and it makes the person at the handle end sovereign, and the person at the pointy-shooty end subjugated.
Regarding the advisory group, OpenAI claims to “have drawn on their advice”, which would include not dumping a bunch of AI slop, with the footnote that if they do do that, at least fund the process of digesting it.
At the same time, there's a new note at the bottom of agmai.org stating how they've been in contact with OpenAI about this particular release, and they say that “we consider these discussions constructive, it is ultimately up to the mathematical community to assess the extent to which our recommendations were followed successfully”.
So, what's going on there; is this British English for “they didn't follow anything at all”? Because from my perspective, it looks like they doubled down on the Navier–Stokes approach of trying to maximize PR gain while being as lazy as possible about actually contributing anything back to science, releasing only slop that may or may not be correct and may or may not be straight up plagiarism, as has been the case earlier.
If I were on the AGMAI board, I'd feel terribly exploited when reading that press release, yet their response is modest.
Hairer, if you're reading this: is there any indication whatsoever that AGMAI was anything but a cheap way for OpenAI to science-wash their press release?
> is this British English for “they didn't follow anything at all”?
Yes, but the subtext is even stronger.
> AI slop,
Now I know there are issues with the field and how just answering these questions may cause broader problems, but I feel like the posted results is far from slop. We can't just call any output slop, or it loses all meaning.
If it was slop, it'd not be causing the issues the group are concerned about - they're not saying "the problem is we're getting loads of incorrect proofs thrown about that are nonsense".
When you blanket a set of things with a pejorative, and it turns out that some of the members of that set are demonstrably and definitively NOT covered by that pejorative, and that all the pejorative means at bottom is "I don't like", all you've accomplished in the long run is to call into question any future legitimate use of that pejorative. It is tempting, especially when heated, to stretch an invective, but it will ironically only lead to the death of its utility over time.
So the fact that the Library of Babel (i.e. all possible books) contains occasional gems means that you can't object to using it on principle? That would seem to follow from your logic.
What about a filtered set "all well formed books"? Or "all well formed books that are plausible enough that they could convince a reasonable person, regardless of their accuracy"?
It's generally taken that a cup of sewage in a barrel of wine makes a barrel of sewage. Surely a reasonable person could claim that a barrel of sewage was still sewage, even if it contained several cups of wine?
I personally just find it hilarious how the complaints and excuses against AI have slowly marched and changed from 2023 to now.
They're not calling any output slop, they're calling indecipherable output slop. The management class responsible for hiring, firing, and paying people doesn't possess the domain knowledge to say for certain whether or not LLM output is optimal (which, in this context, means correct), but they will trust that it's good enough to justify further automation / fewer grant approvals / etc. So in that sense, slop can and will cause the economic issues people are concerned about.
University boards want the prestige of successful research programs. Doing the hard work to get something demonstrably true is going to lose out economically in this paradigm, where we are all being conditioned to uncritically ooh and aah at the incantations being elicited from these magic boxes. The oracles even have legions of zealots who will berate you for not being sufficiently deferential and reverent, or worse, accuse you of blasphemy. If for no other reason, I agree with using the term to express all of the above succinctly, even if LLMs can be helpful tools generally.
I've read some of the results papers (the Einstein condensate one and the pi exponential one). I'm not an expert but it definitely wasn't AI slop. The introduction sections were particularly well framed and informative.
Also you can see in the papers where an idea is introduced but in the bibliography you can see where the foundational idea comes from. So the narratives are not unmotivated as some claim (proof without intuition claims).
In the context of maths papers, the term has come to refer to papers having the shortcomings that are, for whatever reason, typical of LLM out, including things like using non-standard terminology all over the place, emphasizing easy steps while leaping over harder ones, having bizarre organisation, and, importantly, failing to properly cover existing work and as a result being hard to tell from plagiarism.
The degree to which these issues feature will differ, but it is generally the case that converting the output to proper research requires significant effort, hence the AGMAI recommendations being what they are, and not performing that effort tends to come off as laziness or incompetence, so I can see how slop has become the popular term.
> We can't just call any output slop, or it loses all meaning.
The term “ai slop” is not supposed to discriminate good ai output from bad, the entire purpose of the phrase is a blanket term that delegitimizes all ai output.
That is not how it is generally being used.
That is exactly how I see it generally being used. Why else would people be dismissing work as AI slop without even reading it, discovering what it says, or even looking into how and to what extent AI was used in a project? Saying things like "if you didn't write it I won't read it" at the first whiff of an AI smell is absolutely said to delegitimize all ai output.
Or in this specific case, why would someone call these proofs (no one is saying they are wrong) AI slop if not to delegitimize all AI output?
This. There's a large subset of people who, seemingly consciously, try to delegitimize anything related to AI by calling it* "slop".
* even pretty amazing advances like this one
AI Derangement Syndrome
People call some work AI slop "without even reading it" when the intention/substance of the work might exist somewhere buried within a wall of impenetrable LLM text (aka "the slop").
Good AI output is indistinguishable from human output. The whiff you mention is the reasoning pleonasm and tautology (intended) escaping into the output and the "author" not proof-reading/editing it out.
But it also happens in many other contexts where that is not true, such as this one right now. Bringing me back to my point that it’s not to discriminate and clarify between good and bad, but to muddy the water.
We can't argue the latter without quantifying the former. All terms are misused by someone, but if it's statistically insignificant that's not an issue. I'm not convinced this one is sufficiently misused to detract from the common definition.
> The term “ai slop” is not supposed to discriminate good ai output from bad, the entire purpose of the phrase is a blanket term that delegitimizes all ai output.
Then it's a useless term and we should all stop using it.
Haven't been following this debate closely, but what's the issue with "strip mining open problems"? Surely the supply of interesting mathematical problems is (in theory) infinite?
You can find Tao’s arguments here: https://mathstodon.xyz/@tao/117237320796901560
He argues that the supply nay be very large indeed but the interesting subset is not. Figuring out the interesting problems is difficult so strip mining the good known problems may lead to scarcity. I am not a mathematician myself, can not judge this accurately.
A Swedish proverb says, "a fool may ask more than ten wise may answer". This fool is reporting for duty! I'm glad I may have something to contribute after all (and I'm only halfway joking)
i'd be curious to hear why he thinks ai couldn't help make it easier to discover interesting problems, ie to make the interesting subset less scarce.
I guess you can see this as an exploration problem, in pure maths, while the goal is to solve a conjecture, the limitation of humans on pure computational power led to the exploration of alternative paths. Sometimes, these paths weren't leading to solving the initial conjecture but opened new idea and new direction. Sometimes a less direct but more humanly natural path was taken to solve the conjecture which also led to new and humanly understandable questions. In some ways solving the question wasn't the most important part of the work, as this doesn't have direct impact on our life (as I saw people comparing this with drug discovery), but the path leading to the solution raised new conjectures and techniques that further developed the field.
I have a really hard time reading AI proof so this might be a biased statement, but most of them feels like having a superpowerfull machine, that would have bruteforce all the possible words of finite length in your logical syntax. You have the path to the solution, using tools that where already known and even direction that where abandoned because they seemed to fail for our human brain. But at the end, as a mathematician, you don't learn anything that is really new.
To me this is the main risk with AI and in general the one most mathematican try to explain but fail, we might miss a lot of alternative path that would have raised more interesting questions (I think this is already more or less what is happening). On top of that, we will run out of mathematicians as no one wants to pursue a career in the field anymore.
This relies on the idea that AIs will only ever do the thing they just did, and nothing more.
It's the same argument which is invariably wrong yet comes up over and over again.
There's no real reason to think AIs solving lots of problems will stop further work on alternative paths - certainly a machine which never tires and can be trained on its own solutions is going to continue to improve.
There's precedent for this: just look at any overconfident post regarding what China will clearly never be able to do, despite decades of steady if frequently flawed progress.
There's no persuasive argument being presented as to why machine mathematical research should have a limit beyond hardware capabilities.
I think you are missing the points of my argument, my argument don't stand on AI isn't capable of discovering new techniques, as I don't believe in new techniques from the sky. My argument is about any targeted goal based AI (which to the best of my knowledge is the case for LLMs as used now). My point is, if there exists a computationally bounded path from existing work that led to solving a conjecture and if the goal of the AI is to solve this conjecture, then alternative path that would have led to new discovery will be dismissed on the way (or lost in the computational trace if you prefer), leading to the conjecture being solved but maybe closing forever/for a long time new paths. I don't see how you could have as a goal to explore alternative path without a good metric of what is a good alternative path (like rating a chess position), which to me, seems unlikely to exist. If you don't have such metric then you would have a clear exponential blowup. More like a percolation problem if you prefer, a neglected approach might have introduced a concept that would make further discoveries accessible. Missing that concept could therefore leave a whole region unexplored, not just one branch of one proof.
We can speculate on whether it can’t, but its plain to see that so far it hasn’t.
Given how new it is, it seems premature to draw any conclusions from that.
If a human had solved these problems, we'd expect it to take years for people to digest them and formulate significant new advances.
Given the demonstrated rate of improvement of AI in math this year, I don't understand the value of that latter observation.
The most immediate answer is because the models are proprietary and only available to those who want to hype the big labs.
That’s what we are doing with nature, seas (look up strip mining there, it’s a horrible practice), and now the industrial harvestors are strip mining problem spaces. How do we like our own medicine?
Developing solutions to mathematical problems generally leads to improvements in quality and quantity of life at roughly the speed they percolate from the ivory tower down to the shop floor. So "how do we like it" is probably going to be "we like it a lot, this is awesome".
Every company is about to have a staff Ops Researcher who has a better grasp of the underlying math and theory than any university professor. That is an unambiguous win.
I see no reason why every company would have a staff ops researcher, or why such a position would have a better grasp of underlying math beyond the narrow slice that directly benefits the company. Why do you think that would happen?
> virtually none of this stuff is possible with technology any normal citizen has access to.
Not sure about the unambiguous win. Are we entering the age in which mathematics is industry-dominated?
1) Any university professor can spend their 24 years on a problem with little progress. 2) company has sudden interests. 3) industrial resources brute force the Lean proof. 4) Max PR for AI company 5) professors are left to rewrite the AI Lean slop into real human-readable math? {disclaimer non-math university professor}
No one is going to get tenure by spending 24 years on a problem with no results. The profs who have that much free time on their hands are already in the later stages of their careers with records of impactful results. By that time, a problem like that is more of a curiosity than sometimes expected to have broad concrete impact.
I'm no mathematician, but (1) seems like a bad situation to be in. I can't speak to the practical usefulness of potential mathematical solutions like proposed here, but it seems useless to have an individual professionally spend 24 years on a single problem only to make little progress and eventually retire so the next person can stare at it.
That's how most other fields progressed most of the time, isn't it?
>> virtually none of this stuff is possible with technology any normal citizen has access to.
Initially, yes. Long term, however? Perhaps still yes.
> 5) professors are left to rewrite the AI Lean slop into real human-readable math?
6) AI writes the proof into something easier to follow than a PDF document.
Hmm. Oh shit.
Who is "we" in that sentence? Why are you not speaking for yourself?
AI math: "We believe this resolves all remaining questions on this topic. No further research is needed." https://xkcd.com/2268/
"Further research is needed to fully understand how we did such a good job."
These are a specific set of interesting, compelling, human-sized problems curated to motivate clever people to engage with math.
> stop testing advanced mathematical problems on proprietary models
I don't know but this phrasing comes off as gatekeeping.
It’s not. Intent matters.
Imagine there's a very advanced crossword club where anybody can join and take a stab at these crosswords for the love of solving puzzles. Many of them are so difficult that no one's been able to solve them yet, but we know they're all solvable.
One day, someone comes along with a super advanced crossword solver application, and it makes easy work of these crosswords. They run it on a few to prove how powerful it is, and then the community says, "Oh wow, that's cool, but please don't run it on any more of our advanced crosswords because they're very hard for us to come up with, and we really enjoy solving them by hand."
That's really what this compares to. I wouldn't call that gatekeeping; just respect. Respect for the game, respect for people's desire to have these hard problems to continue to work on, solving by hand.
If the company with the super advanced crossword solver then continues to use it and publish the results, they're effectively stealing the crosswords from this community. Soon, all the puzzles will be solved, leaving nothing left for the community to work on for fun.
That doesn't sound like gatekeeping to me. That just sounds like someone asking "Please be respectful and leave the remaining puzzles for us to solve by hand.” A simple plea not to be an asshole.
We don't give mathematicians research positions to solve crosswords for fun. We want something back. We want theories and results that will advance our civilization.
We have people who want to fill those positions because there are enough people who find it rewarding enough. Take away reasons why they would find it rewarding and you will have fewer theories and results that will advance our civilisation.
And yes, fun counts. Nobody said this had to be only a hardship.
Money doesn't work that way though. There are plenty of jobs people would like to get paid to do, that doesn't mean someone needs to psy them to do it.
I'm well aware that if at some point AI is good enough to replace me as a software engineer then I won't have a job. I don't expect a company to continue to pay me simply because I enjoy it if there are cheaper options out there.
Math is no different.
As long as that company doesn't expect me to continue in their employment if I stop enjoying it, then we understand each other.
Total compensation includes fun.
This is really the critical thing: the fun is the incentive. (Or at least the dominant incentive in math, historically.) As economists like to say, the overarching lesson in economics is that incentives matter. Reduce the incentives and participation will decrease.
Perhaps that won't matter if we enter an era where AI participants are the main participants who matter for discovery-level mathematics. But it would likely be what economists would see as a market failure if only a small oligopoly of AI participants, closely held behind closed doors, is able to fill that intellectual role.
I think you are missing the point of the main criticism. It is not about not wanting results in terms of proofs.
New theories and insights are typically created while working out proofs. If proofs now suddenly fall out of the sky (cause LLMs create them) then that work is not done which means the substrate on which new theories and questions and conjectures used to be grown disappears. It's in that sense that the math community (and thereby society as a whole) will lose something.
It's similar to how software engineering will need to find a solution to train their next generation. Current generations have all been through manual steps of designing things from scratch and writing them by hand. That's what allows your 10x engineers to understand whether what their LLM tools are doing is good and how to massage those tools to do the right thing. A junior engineer who has only ever used LLMs to write code and create architectures does not just not have that experience but also won't acquire it. You can't just say "we don't pay them to have fun and learn, we pay them to produce results". In the short term that is the case, but in the long term you as a company and we as a community will lose out.
I'm not saying don't use AI tooling. I'm saying that this is a hard problem which we yet to have to find solutions and approaches to. As a software community as well as as society in general.
"A junior engineer who has only ever used LLMs to write code and create architectures does not just not have that experience but also won't acquire it."
My ego tends to agree, that how can they be ever competent, if they have not endured the same hardships as I had crunching trough problems and getting allmost lost in the details.
But I rather suspect, they will turn out fine. I know LLMs are great for me to learn and I think the young generation will learn what they need to learn to get the job done.
How can they learn hard things if they have an infinite number of easy things to do? This is a middlebrow version of doomscrolling disease.
Most people have trouble not peeking at the answers. Look at Stack Exchange's long success.
Because keeping all the easy things coordinated and understanding the big picture is still a hard task yet unsolved by LLM's? But yeah, who knows what happens once that change. I assume even after the singularity, it still makes sense, that we train some people to know what is going on ..
How well do you think someone will understand fractions or trigonometry if they always punch their math homework into Wolfram alpha?
The increasing pervasiveness of technology in US education has not produced more capable graduates.
If a modern Gauss, Von Neumann, and Ramanujan appeared and started dropping proofs from the sky, would people be saying the same things? And if they could live forever, so they wouldn't need to train their replacements?
Who cares about them? I want Tao to stop proving all the interesting problems I was planning to work on.
Gauss and Euler, and also Ramanujan (results without proofs, which is a bit like unreadable Lean) did that for their lifetimes.
Yes, and they are revered as geniuses, which makes it clear that this is all sour grapes. And surely if people died, went to heaven, and were able to talk with God whenever they wanted, they wouldn't be upset that now they could know the answer to any mystery whenever they'd like; they'd appreciate that now they have someone to guide them! Or were they similarly upset when lecturers handed them already completed theory in school? There's already enough developed theory that people don't have the time to learn it all as it is.
Not exactly, because we would have cool people to inspire us and hang out with us.
But your argument is nonsensical because even if Gauss and von Neumann appeared, they wouldn't go into random fields and just prove things mechanically. They'd have to attend seminars, teach others, collaborate with others, and generally inspire others with their brilliance. It's the precise lack of this activity that makes AI in math so reprehensible.
Your argument encapsulates a contradiction because human mathematicians wouldn't be dropping proofs arbitrarily like AI is doing. They would do something completely different. Even the best of them.
Gauss was generally quite secretive and Ramanujan would famously tell people answers that he had received from divine inspiration, often with no ability to articulate how he knew. Von Neumann did just go into random fields and revolutionize them. If the three of them did come back from the dead and form a little powerhouse group that barely collaborated with the outside and just started publishing results for everyone else to try to keep up with, they'd no doubt still be considered geniuses.
Give it six months and models might be able to explain things better than any human. They can already collaborate perfectly well if you ask them to. e.g. there was a post here a couple months ago where Tao shared his ChatGPT logs[0].
If you're not inspired by the ability to talk to a superintelligent machine, and can't find what you'd want to know, that's a you problem.
[0] https://news.ycombinator.com/item?id=49010345
The only thing potentially stopping these models from also outputting new theories along the way is the goal they were given.
I have to assume OpenAI is only prompting to solve problems, presumably they could also prompt to not interesting new theories or paths of research found along the way as well.
OpenAI is doing problems because they know they can't do higher theory yet.
I asusme they're doing problems because its an easy way to turn $40m of someone else's money into a catchy news headline.
I think the crosswords framing is a little silly, but I have to wonder what comes when we use our technology to optimize the fun and interesting parts out of every job. There's only so many years of my life I can dedicate to back-and-forths with a chatbot. What if we advance our glorious civilization but our jobs just get more and more thoughtless and miserable?
I don't know about you, but my job has become a lot more fun ever since it's become a lot more back-and-forth with the robot. It does all the tedious things for me. It gathers data. It creates prototypes. It makes the mechanical code changes that I want. It allows me to talk with it for a design discussion, and then my design simply appears. I ask it for monitoring dashboards and they simply appear. It records what we talked about, which is something that I never do.
Largely I thought that this is what you do once you're established in math (or any field) anyway. You have some ideas, but the details are kind of too tedious for you to work out, so you give it to grad students/postdocs. Senior engineers have some ideas, but the details are tedious to work out, so you give them to junior engineers.
Now, obviously in the meantime, there's the question of how do we train the next generation? Or do we need to train the next generation? And maybe while we work that out the answer becomes more shadowing/apprenticeship instead of farming out easy tasks.
I think that's where people hope some kind if UBI or "universal high income" will save the day. Just don't think too hard about how it would actually be paid for, or how we can all have high income when that's a relative measure and we're all given the same amount of table scraps.
"universal high income" is not when everyone has high income, it's when everyone who doesn't have a high income is excluded from the universe. There will be few high income people, robots those people own, and the rest of us will be undesirables/illegals/felons/noncitizens of Ms-Apple-Meta-Tesla-Google-topia, who for arbitrary reasons XYZ (they didn't accept the EULA!) don't deserve universal high income (i.e. most people here will fall into that category).
What you're describing could well be how it ends up, but that isn't the future described by universal high income.
Yours is more likely in my opinion though, mainly because universal high income is completely infeasible and shaky even at the level of definition.
Then work part time, and enjoy your higher wealth to have fun in free time. Don't demand to have your cake and eat it too.
We're going to have a very different perspective on purpose going forward with these results. This has crossed a rubicon where human output itself is going to be completely outclassed by machines and we will have to find meaning elsewhere in life.
This is a really confusing take.
If someone can solve open problems in mathematics then they should do so, isn't it as simple as that?
They should let the public use the models as well, but I guess they have no real moral imperative to do so.
But asking them to stop solving problems is just weird.
If your only measure of advancing is getting an answer, but not building the capability to understand it, then civilization has advanced.
It’s not a human focused civilization, which is where the issue comes up.
As an example: A constant issue I am seeing with AI productivity is that the most productive use of AI is when it is paired with more experienced users, while AI also does more work for entry level workers, if not replacing them entirely. It has become a question where will the future buffer of experienced seniors come from.
This is an example of where simply chopping down trees for today, doesn’t make civilization better off tomorrow.
AI is producing more content than ever before, but our ability to understand and verify it is not keeping pace.
We don’t know if these are unsolvable problems at this stage. Society could come up with workarounds and solutions to these issues in several years.
The request to stop, is part of the process by which the issues are debated and solutions found. It doesn’t mean their position is weird or moot.
If someone gets the answer sooner than you, that doesn't inhibit you developing your understanding of the answer privately the same way you would have done if they hadn't got the answer. I don't see how anybody loses by the answer being discovered sooner.
Not true. If I know the answer to a puzzle, I don't spend the time doing the puzzle.
If there is a prize associated with doing a puzzle, and a machine does it, then what incentive is there to pursue it.
Again, if you are only concerned with the outcome, and you have a preferred answer that you want (in this case "just use AI to advance faster"), then any information that doesn't support that case is useless or misguided at worst.
I am not trying to dissuade you from your preference. I am flagging that there is a set of other factors that influence the behavior of others, how that behavior is critical to the creation of expertise and drive, and thus why others hold different positions.
If you're concerned with something other than the answer, then the fact that the answer is already known hasn't actually provided the thing you're concerned about, so you can still do the thing you are concerned about.
If another human was likely to get the answer before you would you also discourage them from doing it because they would rob you of the chance to do the thing you're concerned about?
This is an ethical and moral question being added here.
Would it be unethical to dissuade someone else from enjoying the benefits of the process you wish to enjoy ?
Vs
Would it be unethical to stop a machine from data mining all the possible questions you wish to explore/enjoy.
And on another level - I am concerned with a bit more than just the answer. I am concerned with what system is in place to ask more questions and get more answers.
There is nothing in this argument that says that we won’t find some other way to study the subject. Maybe people will become monks and do math as a hobby.
We may end up in a daemon filled world, like 40k, where any hope of understanding the tech around us is impossible. (More impossible that today)
If someone spends their entire career not solving the puzzle, did they really learn to understand how to solve it?
They may very well have learned plenty of things and solved or discovered other puzzles, but if the first puzzle is worth pursuing because the solution is actually useful it seems liked we're better off with the solution than a bunch of failed attempts.
That said, I do question the value of solving many of these types of math problems. I'm no mathematician so I'm assuming I'm wrong here, but on the surface many seem mostly theoretical puzzles with little or no practical use.
Yes? We haven’t solved many puzzles about reality, but even half proofs and conjectures create tools that other people use to make progress.
I’ve made this point elsewhere but the debate here is between two different philosophical positions. Results vs process.
If all you care about is the results then the process doesn’t matter.
If a person is starving or needs medicine, then a long discussion on process is inhumane. They need results.
If the conversation is about process though, then focusing on the results is missing the point.
I’d say the question for results oriented people is what are the benefits of the process and at what point does it make sense to optimize for results vs process.
My read on much of the discussion here is that the debate is whether we want AIs solving problems that career mathematicians may spend a lifetime on and still not solve.
When the topic is about careers the question really has to be about results. Even if the results are made by solving different problems discovered along the way towards their original problem, it still has to be about those results.
There is absolutely a question of whether burning these resources is useful when the only outcome is a solution to a potentially obscure math problem, but that is more a question of prompting and goals rather than the use of these tools themselves.
> There is absolutely a question of whether burning these resources is useful when the only outcome is a solution to a potentially obscure math problem, but that is more a question of prompting and goals rather than the use of these tools themselves.
Could you elaborate?
But isn't all of schooling literally learning solutions others solved before us?
We spend most of our young lives (many of us our entire lives) studying physics, math, etc. that others have solved. (e.g Quantum Mechanics, Relativity, Calculus, etc.)
Biology consists, almost entirely, of studying solved problems in nature.
Aren't AI breakthroughs just more to study?
https://mathstodon.xyz/@tao/117237320796901560
Terence Tao’s “don’t create the open problem strip miner”
That doesn't answer the question. Assume today is not the stopping point, and that we end up with super-intelligent theory building AIs. Better than any current-day human. And better at explaining, creating visualizations, etc. than any current day human.
Why is it a problem that the professor is now a robot, and that humans could spend arbitrarily long learning from it and even after 15 years of masters-style advanced graduate lecture courses still have deeper still levels of the topic that the AI could teach them?
And if they never do reach that level of ultra-competence, well, then we found the niche for humans to continue to exist within.
What is preventing these crossword solvers from not looking at the advanced crossword solutions?
Mathematicians and academics in their ivory towers are forgetting that everything is getting automated. They want to carve out fun problem solving niches that's fine but who's going to fund that? If they want to be funded by the society/civilization their argument can't be leave advanced fun problems for their hobby.
Here's a fun quote:
https://proofsandprompts.com/2026/09/10/open-letter-about-th...
>Participation in an event so closely associated with Anthropic and OpenAI could plausibly negatively impact the future reputations of participants.
Given how much power advisors etc have over students in academia, interpret it as you wish.
Its worth noting though that you are comparing a profession with a hobby.
People go to said crossword group to enjoy the process of solving the puzzles. It doesn't actually matter if they have been solved yet or not, case in point the NY Times puzzles are enjoyed by more than just the first to solve them.
Professional mathematicians are ultimately being paid to solve the problems for a (hopefully) practical reason. Its always excellent when a person enjoys the process of the work they are paid to do, but ultimately they are still paid to do the work. I really hope your argument isn't that we should collectively be funding mathematicians to solve problems simply doe the love of the game.
> Professional mathematicians are ultimately being paid to solve the problems for a (hopefully) practical reason.
They are paid for the same reasons the NEA pays artists: out of a sense of obligation to demonstrate elite culture. The track record of practicality of pure math after WWII is essentially 0.
While I don't disagree, I think any justification for why we should fund mathematics and why we should protect the work they are doing from being solved without them should be grounded in results.
Similarly I wouldn't expect a good argument could be made that AI tools should be prevented from creating art because we want to continue funding artists.
If the goal of said funding is just to let them spend their time doing it then it doesn't matter that AI is doing it as well.
Classic alignment problem.
Despite nobody at openAI thinking of themselves as an asshole; despite society urging openAI not to be an asshole; despite the fact that being an asshole is entirely unnecessary even to accomplish whatever objective they are setting out to accomplish; despite everyone at openAI loudly declaring: we are not assholes!
They are still assholes.
reddit comment
This is maybe the lowest-quality comment on a thread full of them. Do better.
It's done in jest but I think I am accurately pointing out the interesting parallels between what these companies say they are doing (aligning models) and what they are not doing (aligning themselves).
If you listen to them, and you don't have to listen very hard to hear it, basically everyone at these labs is telling us that this technology is extremely dangerous and should be slowed down or paused entirely. Yet, they, the only entities with the power to actually do anything about it, are not acting AT ALL as if that's the case. They are all barrelling forward as quickly as possible. RSI, THE number one risk according to these guys, is being adopted at breakneck pace up and down the stack, from designing silicon, to training, to inference.
It's ridiculous and insane and I believe can be accurately summed up as, they are being assholes, because if they are actually right about this we are all gonna die. At the very least, and far more likely, every fun creative expressive human thing that is machine legible will be replaced by a torrent of machine slop. It's not "benefiting humanity." These mathematicians are telling you it's not benefiting humanity. It sucks.
Alignment problem.
This analogy is silly because (a) math is not primarily for entertainment, (b) we aren't going to run out of math proofs, and (c) results build on top of other results, having more results proven makes all math more powerful and useful.
Hmmm but in the case of math, while some of it is "just puzzles" there often turns out to be practical applications, even if they are not obvious at first. Number theory was considered the epitome of pure math with no practical applications for centuries, now our modern society is built on it (public key crypto).
If the crosswords were purely games that would be no problem. These crosswords seem to power physics, chemistry, engineering and science applications. These professions would not mind it too much.
Blah blah blah. They are free to do their own mathematics and/or spend time on polishing/reviewing proofs dumped by ai. But they don't get to make demands like don't test math on proprietary models. Idiots.
Math doesn't belong to academics. We don't pay them to work on problems for fun. They will just need to re-evaluate where the value their provide is. It won't be solving problems anymore. Hopefully it will be making them understandable by others at least till AI can't do that as well.
"They will just need to re-evaluate where the value their provide is."
That is fine to say when it is not your field. I guarantee you feel different when it is the thing you care about, that gives you joy, that defines your status. Think about how many sheldon-equivalents insist on being called Dr. (non medical)
It is part of what people use to define themselves. Its going to hurt. There may even be a Bulterian Jihad
Not our field? Most of us are programmers here.
It is clear to me that any competent person with a little patience can now build software better than what I used to build by hand.
> Think about how many sheldon-equivalents insist on being called Dr. (non medical)
Why do physicians insist on calling themselves Dr. (medical)?
No, you dislike maths to the point you prefer paying others to do it. Actual mathematicians are largely doing it for fun, but are now effectively saying "stop destroying our fun or we'll stop doing maths", and you will have to do the maths yourself.
Well, they do they?
Whole sections of the economy are being upheaved by AI, and there is no reason to make a special case for the mathematicians anymore than for the illustrators, developers, translators, HR, etc.
Oh right, how rude of me to only talk about mathematicians in this thread about the future prospects of children in Sudan. Of course this is the place to make "what about the illustrators" argument.
It isn't some law of nature. Humans/societies have agency - what AI should or should not be used is up for debate and decisions. It might even wind up the other way around that using AI is the special case - who knows.
> what AI should or should not be used is up for debate and decisions.
Of course; but it's very hypocritical to raise these feelings only when mathematicians are affected, whereas all the above professions are just told to adapt to the new way of things.
For sure though, translators don't have the same clout and social status as mathematicians do.
I actually really like math and I can't wait for the day LLMs not only solve difficult problems but can also explain the solutions to me.
Mathematicians do a terrible job here. They use inconsistent symbols they don't even explain. They often obfuscate the main idea just to make the paper longer. If you are not part of a small club you are not meant to understand it. I think this is a terrible approach and I am eagerly waiting for AI to do a better job!
But won't new humans take their place that will be the ones who enjoy deciphering AI solutions?
It just seems that this class of mathematicians is being "disrupted".
The field is changing and a new class of mathematicians will take their place.
This happens all the time in fields as technology disrupts them.
A new class of individuals, with different motivations, take the place of the old guard.
I'm sure the motivations of individuals involved in designing and manufacturing cars changed as Henry Ford introduced the factor line.
But that old crop of humans either adapted or retired.
But, plenty of humans took their place with new motivations and automotive technology continued to progress.
I personally feel math will indeed move faster as a result of these breakthroughs. And the humans that take the place of the old guard will have different passions and motivations than the current group.
Maybe the new group will be productivity motivated rather than motivated by the love of tinkering with a single problem for years.
> Maybe the new group will be productivity motivated rather than motivated by the love of tinkering with a single problem for years.
Sounds like salaries for mathematicians need to start going up if we stop paying them with fun.
There's no way to spin this that doesn't make it sound like assholes being gatekeepers.
Here's a fun quote:
https://proofsandprompts.com/2026/09/10/open-letter-about-th...
>Participation in an event so closely associated with Anthropic and OpenAI could plausibly negatively impact the future reputations of participants.
Given how much power advisors etc have over students in academia, interpret it as you wish.
I’m sorry; but if mathematicians are in it because puzzle club is fun, then they should go join the fucking puzzle club and stop impeding scientific progress.
Science isn’t some passive busywork thing where you tie your hands behind your back because it isn’t fair on others to solve all the neat problems - or at least it shouldn’t be.
If your idea of science is leather patches on tweed suits and the quiet ticking of a clock while you do crosswords, then this is an argument in favour of letting the AI do the work so you can focus on your sudoku book in your slippers.
Keeping the tech proprietary so that it can only be used on these problems by internal teams is the very definition of gatekeeping.
It's more like, "don't just casually destroy our hobby / career field", without letting us participate even a little.
The picture I have in mind is OpenAI running their most advanced model in a loop over all the open mathematical problems they can find, just to verify that the model is indeed very smart. Neither the company nor the model actually care about the problems, it's just a cheap exercise machine for them, but the problems get solved and mathematicians don't even get to participate.
Like, even those who accepted the "centaur" thinking, man + machine, won't benefit because by the time they get their hands on good enough models, everything is already done.
It's an emotional thing first and foremost - people who care about the thing can't do the thing, because it's already been done by those who couldn't care less about it.
And before someone goes "poor mathematicians", a food for thought: this is just an early instance of what looks like our shared destiny.
I said here before: given the economics of progress in AI and robotics, it's obvious what the natural division of labor is: computers do the thinking, humans do the menial, manual labor. AI will do politics and philosophy, so you have more time to fold laundry and scrub the toilet.
So what is mathematics then? A fun hobby akin to chess or sudoku?
Are we gonna get the same pushback from medical researchers if the models cure xyz diseases?
I absolutely understand the emotional connection to their work and the heartbreak, but mathematics doesn't exist for their pleasure, it exists to provide tools to solve humanitie's problems.
> So what is mathematics then? A fun hobby akin to chess or sudoku?
Some of it, yes. Much like physics. Both have a track record of producing technological breakthroughs every now and then, but it's not why people are doing it.
> Are we gonna get the same pushback from medical researchers if the models cure xyz diseases?
For better or worse, yes. We already are. In my country, there's a big spat between radiologists and cardiologists right now, that boils down to the progress of technology allowing the former to answer questions that, before, involved a procedure that was a big money-maker for the latter.
What you say reminds me of medical schools in Tunisia.
The general body of research points that more doctors lowers all cause mortality ( with diminishing returns) but Tunisia is still far lower than the Eu average.
Yet Doctors and med Student unions do lobby very heavily against expanding admission to the public uni or allowing private unis.
So we have the weird situation where people go and study in Romania ( making Tunisia lose hard currency that it really needs).
These doctors have taken an oath and the direct consequence of their lobbying is literally more deaths.
USA is the same. Even worse, the doctors guild writes the rules for creating new doctors. It got so bad that we now have 2 or 3 other alternate/adjacent categories of doctors and nurses to work around the bottleneck. Of course then they formed guilds to continue the cycle.
Dude, doctors are humans just like rest of us. They want careers, money, safety, raise children in best way possible, fun in life and so on. I see this unspoken expectation over and over - why are they not infallible, how could they do mistake XYZ, why are they not working themselves to the (early) death for benefits of us all and so on. They have no obligation to stay at place Q just because some folks would consider it convenient. They have no obligation to stay in some place thats not suiting them just because they swore Hippocratic oath, lives can be saved elsewhere too.
Obviously this is often coming from folks who act in same ways as they criticize and usually don't contribute even a fraction back to society compared to doctors. Folks who do mistakes in their lives all the time yet thats fine since we are all humans or similar, right.
So please stop this cheap framing and accusations. If Tunisia wants more doctors and keep them there are ways to do it, society as a whole needs to decide what they want and act upon it. Otherwise, smart skilled folks will keep going for better lives elsewhere, just like everybody else.
Is this not greed?
It's self-interest.
Everyone (near enough) has some degree of self-interest. If you apply for a job and discover that some other applicant is about as well fitted to it as you and in more need of money, do you withdraw? If you see a $20 note on the ground and no one else around who might have dropped it, do you refrain from picking it up if you think you're better-off than the median person who might walk past next? If you see something you want going for a very good price on eBay, do you contact the seller and say "I think you should be making me pay more for this"?
Unless you are an extremely unusual person, the answers to those questions are somewhere between "no" and "of course not, and why would you even ask?".
If someone is working as a doctor, their work is already benefiting others substantially more than the typical person's. (At least, I think it is; it's certainly doing so more directly.) Being a doctor doesn't put them under some unique obligation never to give any priority to their own interests when, e.g., choosing what job to take where.
If they can save 0.2 lives per day for $50k/year in one place and save 0.19 lives per day for $200k/year in another, it would be virtuous for them to do the former but I can't see that it's obligatory. In the case we're talking about, it might actually be 0.2 lives per day for $50k/year versus 0.21 lives per day for $200k/year, because somewhere that can afford to pay them more can probably also afford better equipment, more ambulances, etc. (In case it isn't obvious, all actual numbers here are made up and nothing I'm saying depends on exactly what they are, only on the rough relationships between them.)
It seems to me like any principle that would oblige them to pick the first of those options over the second would e.g. also oblige all of us who have well paid jobs to give most of what we earn to life-saving charities. Some people do that. It's a virtuous and commendable thing. It would doubtless be better if more people did. But, as you might have noticed, very very few people do that and by and large we don't consider it outrageous that they don't, and I don't see why doctors in particular should be condemned when they don't do it.
(Since clearly unassisted human nature isn't going to make everyone behave in such a way, it seems to me that if we wanted that sort of thing then it would need to be imposed by force. Which in fact everyone might be OK with, in the same sort of way as players of high-level sports are OK with having externally-imposed safety rules so that we don't get everyone playing in increasingly dangerous ways for the sake of a small advantage over people who are being more careful. And, in fact, we do have that sort of thing and it is imposed by force; it's called taxation, and actually I think it's a beautiful thing even though there's plenty to dislike about every actually-existing regime of taxes and benefits. This is mostly a digression, but note that it means that if a doctor chooses to go somewhere where they're paid better it probably also means that they're contributing more to the general welfare in taxes. There are plenty of nits one could pick with this remark, but it still seems worth making.)
At high levels, often yes. At lower levels, often it's job security.
Most doctors aren't running departments in major hospitals, or advising government on policy. They don't earn the big bucks. And even hospitals themselves tend to run in the red all the time; it's sometimes hard to disentangle where greed ends, and longer-term interests of patients begin, as you have multiple people and organizations pulling in different directions for different reasons.
RE private medical universities, N=1 but in Poland we have a private provider pushing hard for training their own doctors "because public system is too slow and limited", and it's hard to tell whether they have a point, or whether it's a private-driven attempt at privatizing national healthcare, or a mix of both.
>but mathematics doesn't exist for their pleasure, it exists to provide tools to solve humanitie's problems.
The risk here is that this does do fundamental long-term damage to mathematics as a viable field.
Virtually no one is going to want to take on the risk of PhD-level math work, studying a narrow problem for four years or so to arrive at an impressive incremental result, when there's a sword of damocles hanging over their head every day that an internal system held by an oracle they don't have access to may scoop their results and turn those four years into dust.
To some extent, that sword of damocles always existed in a de minimus sense in the form of other mathematicians. But everyone was playing the same game, coming to the game with the same arsenal limited by human cognition.
If the game board becomes irrevocably tilted, new entrants have no incentive to play except as a hobby. But few hobbyists can devote years of work to understanding and pushing the frontier. It could well mean existential damage to mathematics as a field.
Whether that might undermine math's ability to solve humanity's problems in the long term is almost an economics problem, not unlike the question of whether and when the existence of monopolies ultimately restricts long-term economic growth. Much probably depends on whether intellectual monopolies or oligopolies are being created that will supplant the existing mathematics "economy".
> The risk here is that this does do fundamental long-term damage to mathematics as a viable field.
All the commotion evens out: It's much easier to learn maths than ever before; you don't need to go to lectures any more; you don't need to learn from a specialist (advisor, lecturer) any more; it all costs much less than it used to.
So mathematics will continue to advance, albeit differently from before. The social structures will not survive however.
Certainly it'll result in a boom for hobby mathematics, and it'll be a hobby at a much more advanced level than before. Whether those hobbyists can continue to push the actual frontier, particularly if AI models operating along that frontier are not made accessible to hobbyists (either via corporate/AI lab gatekeeping, via pricing, or via significant time lags) is a different question. I'm a little more confident in a future where hobbyists push the frontier in applied mathematics than in pure mathematics.
There's probably a loose and deeply imperfect analogy with computing: via democratization hobbyists have made a big impact in applied operating systems development (Linux/OpenBSD) but have been less successful/impactful in OS research (whither Hurd...) or in cost-heavy fields like microprocessor design.
> Are we gonna get the same pushback from medical researchers if the models cure xyz diseases?
Lol. As long as the process aka trials is respected not many would complain.
The feedback loop required to make progress is very different in medicine compared to math.
The trials process is the moat. There's already founders using AI to treat their cancers, and it's all about skipping trials and jumping straight to "I consent, I'll fund it, let's try it". The general public might get access to this in 10 years, but employees at AI companies will have access much much sooner.
https://sytse.com/cancer/
I don't necessarily see a problem with it: if people want to try experimental therapy on themselves and can fund it, then as long as it's expensive, let them - that speeds up research. The problem with allowing anyone to opt out of safety trials is that it then creates pressure from doctors and family members to try, and then it becomes non-consensual in practice.
Yeah it's more like personalized therapy - often the only hope for rare diseases.
While AI has definitely helped quite a bit I am wondering how much all this research and treatments cost. Not sure the current health systems could sustain this for _everyone affected_. If ai enables it all the better.
What's the success rate there?
At least in Sid's case, it went from the oncologist saying "I have no more drugs I would recommend, no trials available" (slide 7) to "I currently have no evidence of disease" (slide 18). I don't know beyond that or beyond Sid's case - or a similar story of an Australian who treated a cancer tumour their dog had with a similar AI / personalized vaccine process.
My understanding of what Sid's describing is that you do RNA sequencing, a whole genome sequencing, feed that into frontier AI (if it will still let you), and somewhere along the way give the information the AI finds to people who can use it make a personalized mRNA vaccine, specifically for you and your cancer.
Another link here about Sid's case, it explains it didn't go through trials: "made possible through a compassionate use allowance from the U.S. Food and Drug Administration (FDA)".
https://www.houstonmethodist.org/newsroom/houston-methodist-...
I am not medical, so I'm happy for someone who understands better to come in and explain all the myriad ways I am wrong.
Trudging into the technicalities of the example still doesn't undo the question of "What is the point of mathematics? To find answers or to be a hobby?"
It's tempting to say "both", but that misses that AI is now forcing us to pick one.
The AI is not forcing us to pick one, it already decided for us.
As 'ogogmad said upthread:
> mathematics will continue to advance, albeit differently from before. The social structures will not survive however.
What if the point of mathematics is to be mature enough to study and teach math to help humans understand it, without the ego stroke of being the first to solve a problem? Bad communicators are upset that a robot is better than they are solving problems.
I'd definitely say both and the cultural component is becoming more and more important to keep up as AI capabilities increase.
> mathematics doesn't exist for their pleasure, it exists to provide tools to solve humanitie's problems
Who decreed that? Mathematics predates capitalism and publish-or-perish by a couple of millennia. Euclid’s Elements were not written to benefit the weapons or medical industry.
Who decreed that they are entitled to get paid for that?
Maybe this hurts more than it should do because of publish-or-perish.
Mathematicians have been unpaid for centuries. The problem at hand is much deeper then just deciding who gets the taxpayer money.
And they can continue to do unpaid mathematics
But humanity is not going to sit around and wait for solutions just so hobbyists can have a moment of glory.
Why? Elements wasn't a 500 clever puzzle solutions.
“ AI will do politics and philosophy””
Incredibly delusional and disconnected from the vast majority of people who are voters.
"AI will do politics"
If only.
They certainly can't do them worse than humans.
Politics is for humans, it's not meant to be automated.
I'm pretty sure the "proprietary" part is the gatekeeping.
It's not like every disadvantaged kid now can solve a major problem just by sinking a hundred hours in their ChatGPT 8 instance.
Sure, and sometimes gates are needed. That's why we all run spamfilters, those are definitely gatekeepers.
In this instance however, it's openAI and Anthropic that are pushing people out of the field by running secret models that take the interesting work away and leaves the persons having to review endless slop proofs.
You mean by the companies right?
There has never been a stronger need for people to band together and "seize the means of production" for this stuff. The advances being made are ours, not theirs. It's trained on our work, our knowledge.
That's an absurd idea. The work & knowledge this is trained on is public. You have access to it.
What you didn't make is the AI training process and resulting model. Extremely hard working people built that, and it has value in itself.
Without the AI training process, the model is useless. Otherwise we'd already have been here at GPT-3.
> The work & knowledge this is trained on is public.
That's an incredibly generous take. If I'd pulled a fraction of the shenanigans prominent companies have to obtain data I'd be thrown under a prison to the thunderous applause of those who have, and are, doing much worse.
> The work & knowledge this is trained on is public. You have access to it.
I’m interested in how you can support this assertion as it seems at odds with established copyright law
We _think_ this power / divide feels harmless right now, but I'd bet money that NSA, CIA, etc have access to the latest and greatest unrestricted models; and massive compute. At least for OpenAI, and even if not willingly for Anthropic, I'd bet money NSA has it too. (After all, when Google decided to migrate to HTTPS, the NSA decided to hack Google's internal network to preserve their taps).
Who knows what they are up to.
One thing I've wondered about in this respect is what happens if NSA learns 5000 new units of math while the general public learns 4000 new units of math.
This sort of happened at various times in the past, because they hired and/or funded so many mathematicians, and especially before the late 1970s they had many of them working in areas where academic mathematicians weren't working at all, so they were learning more math, or more math that they especially cared about, than the public was. (I was going to write a note here just a few days ago about how NSA has had a "Classified Mathematics Library" for many years.)
For vulnerability scanning, I think the new-capabilities trajectory is good (in the sense of "it will help defenders win") even if governments find ways to get more of it, because there are finitely many bugs and classes of bugs, so at some point more capable models' or longer runs' advantage over less capable models and shorter runs should stop helping them outcompete the less-well-funded defenders, because the defenders will still have learned most of the information that's relevant to achieving successful defenses.
So if NSA gets 5000 units of vulnerability scanning and the public only gets 4000 units, we might still just wipe out all of the pure software vulnerabilities and then go back to worrying about physical supply chain security or side channels or something.
For math, I'm not quite sure! For one thing, there may be things that have no feasibly deployable defense at all even when you understand the underlying mathematics (I'm especially worried about traffic analysis here, because understanding in detail how traffic analysis is done, or how powerful particular techniques are, does not necessarily always or usually make defending against it more convenient or less costly). In a more science fiction scenario, there might also not be any efficient secure cryptographic primitives of some kind, like if it turns out P=NP with reasonably small exponents and reasonably small constant factors.
Based on people I've talked to I'd be really surprised if this was the case, they actually seem to be pretty far behind the ball when it comes to AI use. Which makes sense to me, given the sensitive nature of their data and systems, they don't want to turn on yolo mode and let an agent cook unattended, which is what you need to do to make these discoveries.
I believe it would be a complete failure of the state and frankly downright irresponsible behavior if all the three letter institutions didn't have access to these models and I’m not even a US national nor do I live there. It’s just common sense. Obviously it wouldn’t be public information since it’s national security, but it’s the lowest hanging asymmetric advantage in the history of national security of nations.
> virtually none of this stuff is possible with technology any normal citizen has access to
So far, it looks like open-weight models are lagging less than a year behind frontier capabilities. And I think one year diffusion of technology from "insider lab demo" to widely available is actually pretty fast?
There are lots of research fields which "normal citizen" has no access to - medical and biological research, particle physics. Some of it is somehow publicly controlled (LHC), some of it not at all (commercial pharma research, mostly secret until the final human trials). And most of it reaches "normal citizens" in way more than a year.
(and I'm talking about open-weight models. The availability of commercial AI models from private preview to included-in-your-$100-subscription is currently like 4 months)
I'm wondering what's the impact on human Mathematicians, and especially would-be Mathematicians -- master students, if they HAVE to use AI in their daily life?
Would that impact their own ability of solving Mathematics problems? I mean as a programmer I'm already seeing that impact on the programmers -- sure the best of us can leverage AI to achieve unimaginable things, but many of us are simply vibe coding.
Of course we can assume that it is only the best of us that really matters, and the rest of us are not going to produce anything substantially useful ANYWAY, it might as well to replace the rest of us with AI, but my worry is -- does that really have ZERO impact on the human specie's ability to produce "the best of us"? After all, they don't grow on trees.
It's a grand experiment isn't it? Us senior programmers are pretty good at using AI (or so we think) because we have decades of grinding and problem solving to inform our intuitions. Is that really necessary? The next generation of programmers certainly will not have that level of desk-head interface. Maybe they'll be fine? Maybe the models will get so good it won't matter? Open question.
I imagine the same will be true of AI, but I'll say that in the short term AI is going to make mathematicians better because it solves the breadth problem. Again, I feel like this Barnette conjecture got solved (if it is solved) because of some clever partition function sums which are intellectually tractable but simply too far out of anything I'd seen before (I see the apparition of my GT combinatorics professor intoning gravely that "everyone knows that, Jake, you're an idiot"). Maybe AI will help identify common threads far greater than Google and journal search.
I think if I had ChatGPT when I was 20 and working on this problem for the first time I might not have solved it, but I would have learned every angle and facet of it far more quickly. But then again I would not have spent so many late nights staring at the Országház across the Danube and letting my mind drift and bump against the problem like spilled cargo in the river.
I have been thinking about this, too. Take Mathematics as an example — it’s probably safe to say that only the top 1000 contemporary Mathematicians really matter to the human specie, or perhaps even less. And if you do not show the potential to be one of those when you reach the end of your graduate studies (actually probably already too late), you are 99.999% sure to just push out papers no one reads and such, and an associate professor in a no name school is going to be your lifetime high watermark. Like, the human specie doesn’t care whether you existed or not, from that perspective.
Now if we can prove this, expand it to the whole spectrum of academic studies, and somehow convince 99.99% of us that they are basically garbage and we don’t care about them — sure the elites will throw UBI around but that’s it — then maybe AI is very positive to the human specie.
Oh we better pick up the speed of cloning and artificial fertilization quickly, because people who are told to be garbage probably have no interests in boring children, and it is still a myth how genies are born and grown. We need that diversity.
BTW the whole scheme reads like the background of a Chinese net novel 赛博英雄传.
oh joy, eugenics and miscegenation
I'm probably the 1,000,000th ranked contemporary mathematician and I matter a great deal to the human species.
In a sane world this power would not be allowed in the hands of private corporations.
I went back to that Fable chat and showed it this new preprint. It coded up the new constructive algorithm and ran it against the existing test suite, that looks good at least.
It has been super helpful in delineating where the crucial concept came from. The proof is rather simple as graph theory proofs go, but it does seem to use some constructions that would only seem obvious if you had serious physics experience with partition function and calculating energy states that cancel out. It's not a wholly alien bolt from the heavens, but I can also see how there hasn't been a human being with the broad theoretical physics knowledge combined with the deep graph theory experience in planar graphs to come up with this idea. I don't know, I'm looking for precedents of this formulation and some old papers of Penrose counting the number of edge colorings of this same graph type are coming up, the line of argument at least rhymes.
But I agree with the thought that this sort of progress should not be siloed inside those companies. I propose a tax so that every slop cannon AI video pays for another hour of compute time for advancing mathematics.
virtually none of this stuff is possible with technology any normal citizen has access to
I suspect that this might be one of the reasons people inside the labs are scared about AI.
What if they have asked AI how it would wipe out humanity and it came up with reasonable answers that they don’t want to publish unlike they do with these math problems?
I think those models and findings should be investigated.
The ways AI can eliminate humanity are trivial obvious and already published. It's just "let the AI control anything of importance and let it spit out slop"
Anthropic runs a biology wetlab (while denying biology to consumers of even their publicly available models, let alone their inhouse ones that only they can access) so I'd expect AI to generate practical and lucrative products soon.
Cure for aging? What do you reckon that'd be worth?
I always got the sense that solutions for significant "unsolved problems in medicine" would be at least 10 years out from the point of total AI dominance in the theoretical sciences. Doing actual experiments is bottlenecked by real-life constraints (organisms are slow to grow and unpredictable, human laws won't let you build a factory to brute-force biology on a million test tube guinea pigs, let alone humans), and the theoretical side of biology is also relatively underdeveloped, to the point that "solve aging" seems as hard to formulate as Navier-Stokes would have been with 15th-century mathematics.
That would be disaster. It would mean the world would not get rid of trump (and similar) by natural causes. Death is the final - and perhaps the only? - equaliser.
A publicly available AI biology wet lab would likely lead to horrific outcomes as people vibe coded virulent pathogens.
If they find a shortcut (like a viral injected cell-dna damage reset) - that would be big. And can you imagine handling the cure for aging, to societies that still produce exponential people?
A cure that you take once and that's it, your body is that age forever? Now, a supplement that you have to keep taking to stay that biological age, that's where the real money is.
I find it troubling that we will solve aging but won’t solve money
The US spends about 18% (and rising) of its GDP on healthcare, so solving that would go a long way towards solving money.
extremely obvious you don't understand anything about biology
> It's becoming an incredible concentration of power that I don't know that we've ever quite seen before.
Replace “AI” with “supercomputer”.
(Super)computers have been solving many math problems that mathematicians can’t solve. Now they are capable of solving problem types that they weren’t able to solve before. (this applies to other fields as well)
Problem is it’s not clear if there is anything left for humans. Probably yes, since human mathematicians are still more economical.
I want a jet airplane, but I can't afford one, and all the ones that exist are proprietary. How is this different from AI models?
> I want a jet airplane, but I can't afford one, and all the ones that exist are proprietary.
I guess if you worked together with some people who all put some money into a fund, and by using very modern technologies like 3D printing and modern CAD modelling etc., it should be possible even for private people to build a jet airplane.
The problem rather is that the government does an insane amount of gatekeeping to prevent this from happening (enforcing expensive and time-consuming certifications on airplanes and pilots etc.).
You're talking about an end user not being able to afford a luxury item.
The concern is about elite level researchers no longer being able to move the industry forward in a public way, and leaving potentially all major discoveries in private hands going forward.
Possible worst case scenario in your case, you personally miss out on a luxury item.
Possible worst case scenario in the topic case, an AI company controls the only intelligence that discovers and understands the most powerful tools / physics we know of.
You could theoretically run these (slowly) if they were open weight. ~$10k of DDR4 is enough to hold them. The data itself costs ~0 to replicate.
And I can cross the country slowly on a go-kart. Not a substitute.
They no doubt have more expensive/powerful models internally, but smaller models seem to catch up fast. So I'm not sure it's about capabilities, but more the willingness and budget to conduct a huge search.
Obviously the more intelligent the model, the smaller/more directed the search is. But they spoke about huge numbers of agents working on Navier-Stokes for example (I think it cost >$10m).
True. What if the emerging capabilities of their best models are applied to tasks like “maximize the chances this pro-AI candidate wins an election” or “maximize profit via stock trading”. Every advantage compounds until all power in the world with any significance belongs solely to whoever has the best models and most compute.
> AI math is happening and there's no going back.
> I suspect that this is in fact the source of much of the angst.
Your comment reveals that you absolutely did not read or understand the Field medalists' open letter... Please, why would you refer to their complaints and claim you disagree when you clearly aren't engaging with the arguments presented therein!?
Totally agree - and not only that we don't know the exact details how these results were produced which is deeply problematic - we just have the end result (and some of the reasoning traces). For this to be a scientific disclsure, we need to know what the agentic setup was, what information was put in, how much and which prior work it relied on, whether the constructions it's using are just ripping off existing work without citation or something it invented (and if so, to what extent) and so on - it's not clear at all what the actual new contribution of the AI model is. All this makes it feel much less like an actual scientific contribution and more like a pre-IPO stunt.
But to me it also signals (as if it didn't before!) a great need for the wider AI community to focus exclusively on researching and building AI algorithms and systems that are more humanistic: completely transparent in its workings and the representations they create, super efficient in terms of data and compute, componentised so that individual entities can plug in different bits and rapidly train on their own data, highly adaptive to individual needs, programmable in a real sense, largely independent of corporate influence, easily accessible to everyone across all social and economic strata, and enable individuals to grow/learn/reach their full potential.
Is this possible? I think so, but it will require ingenuity and bringing in ideas from (ironically enough) some of the deepest areas of modern mathematics such category theory, algebraic topology etc. which are largely about building abstractions that expose the underlying structure of complex mathematical objects and the relationships between them.
It's already happening to a degree, but the urgency has reached epic levels at this point and it needs to happen at scale.
It's a bit aggravating that I cannot interrogate the session that yielded this result and ask it why and where it got the crucial calculation from, or why it went in that direction. It doesn't even rightly know even if it gives you a legible answer, that doesn't have any correlation with whatever happened under the hood.
Humans are the same way sometimes but I guess there's romance in that. If a human had solved it a la Kekulé and said "it came to me in a dream" I would at least understand that.
Math isn't scientific, none of that is "problematic"
Sorry I was using scientific in a broader sense - probably should have used "academic" instead -
it's deeply problematic because they are building on open, public results yet they don't provide information on how people may build on it - its exploitative and exclusionary - at least they are consistent
I find this argument to be extremely ridiculous. They solved some math problems and published the results for free. No one asked them to do it, they weren't paid, and they don't owe anyone anything. Who exactly is exploited and excluded? The entire notion of open public information is that you can do anything you want with it, including build private systems. Is a baker "exploitative and exclusionary" for reading a recipe in a book and then turning around and selling that bread to customers, without sharing the recipe with the customers?
The anti-AI arguments keep morphing, as many could have probably predicted. Starting with "AI can't do anything" to "AI can't do anything useful" to ... "AI breakthroughs are proprietary!#@!!!".
I've seen more goalposts move in the last 3 years than maybe in my whole (lengthy) career up to that point.
> I have no problem with AI models making revolutionary advances in math or science. Where I start to have a problem is when the AI models making these advances are tightly withheld, proprietary, and seemingly never released with these capabilities intact.
Agree, and, to my mind - shows why the efforts of the Free Software Foundation have been worthwhile all along. We need software to be open / free / libre or the power elite controlling them will ruin the world.
What exactly are you worried about? OpenAI/etc. gaining too much power? If they use it, the government can stop them. If you worry about the government, isn't it better that than rando terrorists? Seems similar to the early days of nuclear and rocket technology. It took stupendous amounts of money and smart people. It was barely accessible to many countries let alone people.
> What exactly are you worried about? OpenAI/etc. gaining too much power?
Yes. They have already shown to have no scruples when it comes to making profit and to have little to no morals.
> If you worry about the government, isn't it better that than rando terrorists?
In my country the largest terrorist attack was almost certainly financed by Iran and caused roughly one hundred deaths. This number pales compared to the thousands who died during the latest, US-backed military coup, a move that relied on a doctrine that the US has never stopped asserting [1].
And those morals I mentioned earlier from AI companies? They do not apply to me because I'm not a US citizen. So no, I do not think the US government is the "seal of quality" you think it is.
[1] https://en.wikipedia.org/wiki/Monroe_Doctrine
I don't think the comment you're replying to said that the U.S. govt. is a seal of quality, at all. They kind-of implicitly concededed that trusting a government with that power is highly sub-optimal, but better still than allowing it to get into the hands of terrorists. Which is a very real issue and a nontrivial point of tension. Like, I'm sorry, maybe I'm reading into this too much, but I personally see the "the government is not the seal of quality you think it is" as a rude and even patronizing misinterpretation happening far too often in discussions, and as needlessly diverging attention from the crux of the problem.
I want to push back on "better than getting into the hands of terrorists being a very real issue".
I am currently in Germany. In the 21st Century roughly 60 people have been killed and 160 injured in ~40 terrorist attacks, most of them perpetrated with cars or knives [1]. In comparison, the US' war in Iran has costed Germany 2.781 billion dollars in fuel costs this year alone and the US government has publicly announced its plans to interfere in German politics partially by funding far-right activities [2].
My point being: the probabilities of terrorists shaking the world order with AI are rather low, seeing as even the most successful attacks in this century have been performed with the simplest of technologies. In contrast, the probability of the US flexing its power irresponsibly are rather high, seeing as they have been doing it for a couple years now and are, in fact, doing it right now.
As far as I'm concerned, and from an evidence-based, day-to-day point of view, the "AI in the hands of terrorists" is an irrelevant concern while "the US may abuse its power" is not.
[1] https://en.wikipedia.org/wiki/Terrorism_in_Germany
[2] https://www.theguardian.com/us-news/2026/jul/15/germany-warn...
Thank you for your insight.
The US government has shown, time and time again, that they will always side with large corporations. Having them as the last backstop is not reassuring.
Have you considered the possibility that the AI labs could actually become more powerful than the US government precisely because they control this technology?
Universities, at least, should be given access
> OpenAI/etc. gaining too much power? If they use it, the government can stop them.
Has the government stopped Google and Apple? https://news.ycombinator.com/item?id=49964791
I guess the objection to closed source slurries releasing world-shaking mathematical proofs, from a conservative libertarian standpoint, is that it's inherently dangerous to individuals whenever access to information or technology is concentrated too much in one place, whether that's government, private equity, religions, cults, terrorist cells, or anything else.
> AI math is happening and there's no going back
"Math" is about uncovering the epistemological foundations of the universe.
Adding AI here does nothing and is probably a regression in that it diverts resources from actual "math" into some sort of LLM wankery that nobody wants.
That depends on whether the AI-generated mathematics helps with the project of "uncovering the epistemological foundations of the universe".
Which depends on (1) whether there are actual good ideas in it, (2) whether as well as finding the proofs the AIs can explain their ideas in ways humans (and other AIs) can use, and (3) whether the results they prove are ones that really contribute to that rather than being isolated curiosities that don't go anywhere.
I am not expert enough in all these fields, and haven't looked enough at the papers, to assess #1, but in general the way mathematicians have bet is that if you can solve things regarded as important problems you'll usually do so in a way that contains more broadly useful ideas. Differences between how today's AI systems do mathematics and how humans do mathematics might make that less true when it's an AI that solves the problem, but I would still bet that way. I'd be surprised if OpenAI's big math dump didn't turn out to contain some ideas, and connections between ideas, that humans find useful.
At the moment the AIs are worse than good humans at #2. (But some humans are also really bad at #2, including some humans who are very good at proving theorems.) It looks to me as if they're getting better, and I would expect them to continue to do so. I also suspect (but this is only guesswork) that today's publicly-available frontier AIs may be able to answer questions along the lines of "please take a look at this AI-written paper, and tell me what key new ideas it contains and how they relate to other things in the field" well enough to be useful to human mathematicians. (Even when the paper itself was written by a proprietary AI that no one outside OpenAI or Anthropic or Hypothetical New AI Mathematics Lab has access to.)
As for #3, that's always been something of a crapshoot. A lot of mathematicians' effort goes into proving things that approximately no one ever reads or builds on, just as a lot of industrial R&D goes into trying things that don't turn out to make good products. The recent OpenAI dump contains things that sure seem like important building blocks for future mathematics (e.g., the "quasi-Riemann-Hypothesis" thing) but it's hard to know for sure and also hard to know whether, if they do prove things that turn out to be useful, it's only because they've read the human-written literature and aimed at things human beings have said seem likely to be useful.
None of this seems to me like "adding AI here does nothing". Whether what AIs are doing to mathematics at the moment is good on balance is highly debatable, of course, but it's a matter of trading off costs and benefits, rather than there being costs and no benefits.
>>However, on one point I increasingly agree: virtually none of this stuff is possible with technology any normal citizen has access to.
So basically nothing changes, Math was subject to gatekeeping and policing of the worst kind.
If you were not among the geniuses, and it didn't come to you automagically, you were simply supposed to leave it to the people who did get it and go do work for people of your intelligence. Smugness was too much to take.
Math people, like chess people never made any genuine attempt to help people understand the processes and methods that made math happen.
To me it should have been a field as teachable and ubiquitous as accounting.
The net result is once these methods and processes were worked out by AI, it was over for the human mathematicians.
I have very little understanding of higher math, so I ask you: Was the proof due to a type of brute-force solution that could be solved had you gained enough information from reading others' work, or was it more like a proof that was sparked by an insight that came once a clue on how to solve it was put forward? I guess my question is: Was the problem proven by using a collection of everyone's work, or was it due to a brand-new insight?
I'm still digesting the proof and translating a bit from the dual case back to the primal in which I most commonly thought about it. I don't think it was a brute force proof in the sense that it combined every possible paper and commentary. It's rather odd because I feel like most of the work on the conjecture was focused on an induction proof based around graph reductions, and this proof avoided those issues entirely by offering a concrete constructive proof of finding a Hamiltonian cycle. Rather, it explicitly selected the edges not in the Hamiltonian cycle, which is in line with previous attempts via the dual.
The "aha" insight for this is actually f**ing wild, it involves a complex valued exponential sum on the edges. I've seen a lot of clever counting arguments before in graph theory but this is the first time I've seen complex roots and annihilating terms like this, the symbolic manipulation tricks in this look like things out of quantum physics. I don't understand where this trick originated, I need to really digest this.
BTW, reading your last paragraph reminds me of how Lee Sedol felt after move 37.
Ironic, as I remember staying late at the Google office to watch that match live. I didn't really understand anything going on but I knew enough to be excited. What a decade.
And we’re only a bit more than halfway through this current one. Exciting/terrifying.
I just revisited this to make that exact comment.
I'm sympathetic to the mathematicians who are worried about the future of their field, but as an outsider I wonder if they couldn't learn from the go community's "recovery" after the introduction of an alien intelligence.
Look, I quit Google a decade ago and tried to make a ChatGPT-lite LLM in my living room (turns out 2017 and GTX1080ti era was a shade too early). I knew that this technology was eventually going to revolutionize programming and mathematics and everything else. I am still flummoxed on a daily basis watching it transpire.
But also I am excited to be living through this new era of programming and new era of mathematics. I'm still saddened that I couldn't be the one to solve this old problem, but now I realize that my personal approaches were really solving a level of this problem even stronger than the original conjecture, and I'm energized to tackle those (in my free time between being a solo founder and father of 3, etc.).
You should try asking an LLM to look for previous papers using similar ideas. The current/frontier generation of math AI is unfortunately very bad at citing the relevant literature for techniques its using.
I asked GPT here: https://chatgpt.com/share/6ac5fd7d-0390-83ed-a02a-6d80fc64f6... and it says:
> the exact Barnette argument appears quite novel, but nearly every ingredient in its cancellation trick has a recognizable ancestor.
> The closest precedent is much closer than I expected: in fully packed O(n) loop models, people have been assigning complex phases to the two orientations of a loop and making them cancel for decades. At n=0, the phases are literally +I and -I. And the n->0 limit has specifically been used to extract Hamiltonian cycles/walks.
You can judge better than me. But it's definitely worth it having a research assistant AI with you when reading these papers.
So much about LLMs can be framed as Information Retrieval, Compression, and Search. Computers have always been good at ruthlessly hammering through a huge but finite set of possibilities. The wild thing now is that you can define that set of possibilities as "all the ideas ever published in mathematics journals."
It makes solving advanced math problems feel like cracking a hash. If it's possible, it's just a matter of compute time.
Complex roots and annihilating terms -- is it something like the derivation of Fourier / Laplace transform?
WOW
Why would the trick have any "origins", isn't this model creating new techniques never before seen or imagined?
There is a chance that someone from a completely different field came up with a solution for a tiny part of your problem.
If you can remember the content of any scientific publication and any book in the world, you are able to make use of this knowledge in every step of you proof.
However, this does now answer how the model came up with the specific route it has taken for the proof.
LLMs don't have super memory like that. I mean I don't know what this internal OAI model is, but at least for other LLMs, they aren't databases of training data with a smart search on top.
The agents here very likely used search. On top of that, they have boundless patience and can quickly process top K hits to find what they need. This is exactly the skill that is super useful for finding various niche sub-proofs that can help you build the final proof. A human mathematician is not going to digest 1000 papers from a different sub-field to find the needle they want, not knowing if it is actually there. AI can do it in few hours.
As Terry Tao said, LLMs are not outsmarting us, they are out remembering us.
I'm fairly sure your understanding is not fully accurate.
I'm not convinced anyone really understands the difference.
I did not mean to say that an LLM knows literally all the publications. But the abstract knowledge is probably encoded in the weights.
No but they have training data which teaches them certain amount of complex understandings and just not math but also physics. So this is one huge advantage.
And then they are for sure able to fill their context based on 'smart search on top' to actually progress further.
As I understand it it's undetermined yet whether LLMs can actually come up with anything novel or are instead pulling from their incredibly deep corpus of knowledge to present solutions that were there but we didn't realize it because our brains aren't libraries of almost all human writing.
Synthetic data allows them to train well past the limits of human writing.
What's an example of synthetic data?
Only in the same sense it's not yet determined about humans, either.
Not so sure. Was everything already "there" before humans existed?
In some form and shape, yes. Humanity's creativity is a lot of marginal copying and remixing.
But obviously, it adds up to something greater than went in; in aggregate, our contributions are something to awe.
But my point is, if you zoom in at the marginal, incremental contributions of any individual human in this process, it's really hard for me to say LLMs are not at the same level already.
On this topic, people like to compare LLMs to Einstein, but as far as I know, Einstein did not zero-shot special relativity in an afternoon. He built it up incrementally over time, it took him three times longer than the time between first ChatGPT release and today, and it depended on centuries of prior art, culminating in the right observation and right notation being available to him in his moment of greatness.
Unless everything was there before humans existed humans created some ideas etc from scratch and not just remixed and copied.
At what level LLMs are is then an entirely separate discussion, I think.
> humans created some ideas etc from scratch and not just remixed and copied.
Name three.
What would you accept as evidence there? Are, for example, the first names/words for colors from scratch?
So your view is that everything was there at the creation of the universe (it's a possible view, of course)? Or are there any "things" that can create ideas from scratch?
Recently I watched a documentary on the tanzanian Hadza tribe, one of the last hunter gatherer tribes on Earth. Their language is a distinct click and pop language and they regularly imitate animal calls (monkeys, baboons, birds) when they hunt but also when they communicate with each other, tell stories etc.
I think it's not impossible that words evolved as adaptations of the environmental sounds with which our ancestors lived. The human creativity producing DNA is also a remix of preexisting molecules formed under evolutionary pressure, so the view that it's turtles all the way down, unintuitive as it is, may not be so indefensible after all.
I mean what is your criterion on invention here? On the one hand, each specific word could be seen as a new invention. On the other hand, all languages basically correlate strongly with the environment of their users - it's why LLMs turn out to be universal translators - and pattern-matching is hardly an invention, isn't it?
But LLMs aren't turtles all the way down, they stop at vector embedded tokenized words.
My view is that LLMs meet the standards by which we judge human creativity/inventiveness, and thus that one cannot claim LLMs "just repeat, never invent" without the same being true about humans.
Pornography, "I Want it That Way" by the Backstreet Boys, torque wrench.
No this is not an issue. As long as their is a way of verifying things, they do the same thing with creating novel things as humans: Searching through an infinite space of possibilities opitmized by knowledge.
They combine things, verify it and if it works and progresses the problem, they created something new.
Let me introduce you to 'obscure Russian mathematicians'.
> Was the problem proven by using a collection of everyone's work, or was it due to a brand-new insight?
Loaded question. A "brand-new insight" is still built off the work of others. A possibly better way to frame it would be in how many subjectively unintuitive logical leaps have been made from prior work.
From my current understanding (and a lot of theoretical physics I'm having to Google because the sentences I'm reading from Fable's analysis are so bizarre I think they are hallucinations) there are possibly 3 neat symbolic tricks borrowed from theoretical physics that make the heart of this proof. Forgive me for posting LLM output but I find this darkly hilarious:
"it's a matrix-tree cancellation wearing Kasteleyn's planar signs, run as a Witten index over Penrose-lineage states, evaluated as a fugacity-zero loop gas in an infinitesimal magnetic field — and the reason it reads like physics is that every one of those tools was built for partition functions"
I thought this was pure slop when I read it but there are some clear analogues in these other areas of physics, really neat computational tricks, and a very interesting paper by Penrose calculating Tait colorings I never knew about previously (extremely relevant, actually related to a separate approach I had once taken on this problem). The problem is that the paper isn't saying "aha, we were inspired by the related problems of pairing excited states and creating spanning trees out of cancelled coefficients" it just defines the function apropos of nothing. Which is kind of like the Jacobian counterexample in that it works but doesn't really explain how exactly it got there.
I really think the load-bearing concept here is "prior work". If prior work is considered papers on this problem or graph theory, yes this has one huge subjectively unintuitive logical leap. If "prior work" is the entire corpus of neat computational tricks that physicists derived to make their equations spit out something other than zero or infinity, maybe it's not so crazy?
I don't have much to add to the math parts, but I've read all your answers in this thread and wanted to thank you for taking the time to offer a detailed perspective from a subject matter expert. Thank you!
Actually reminds me of patent law. Prior art ist a defined term which includes all standard literature on one topic. To evaluate, whether the new solution is really inventive and thus patentable, one consults prior art, selects the most promising starting point, and from there asks oneself if an all-knowing but uncreative specialist would come up with the solution by himself. If he wouldn't, the condition of inventiveness is satisfied.
Makes me wonder how the patent space will be disrupted when that inventiveness step becomes obsolete because of LLMs. Given your example above, it seems like a combination of different methods from many different sources. This would be regarded as inventive, clearly. If eligible patents can now be brute-forced, the bottleneck becomes only selecting the most promising ones and paying for the patent.
Oh man, we should talk. I have been working on a patent with ChatGPT specifically to get around two complementary patents that are now together because of a corporate merger this year. I am not sure how much longer anything is going to be patentable with this kind of design assistance available to everyone.
Also, once upon a time I wanted to be a patent lawyer. It's incredibly hard to sit for the patent bar if you have a pure math degree and don't have an engineering degree. Thankfully New Hampshire lets anyone sit for the FE exam.
Did anyone else wince at seeing the phrase "load-bearing"?
I did as I wrote it. I actually used that phrase often before it became an LLM-ism, just like how I rather enjoyed peppering my writing with em-dashes. Oh well.
Language constructs becoming aggressively passé due to AI saturation is one of the craziest outcomes of all of this stuff—one which I don't think anyone saw coming.
Are there no loads left to be borne?
one hopes at least that the taboo on the bearing of loads is restricted to metaphorical loads only, lest lorry drivers and porters become the next victim of the algospeak spectre
Kasteleyn signs definitely have math counterparts (Arf invariants). They’re just not as well-known.
Right. And I'm kicking myself for not having the mathematical breadth to know about them.
Why? Physics people I talked to didn't know either.
It's a shame OpenAI will never publish the trace that led to the insight.
Condolences, im familiar with the feeling. I hope this AI thing somehow works out for the better and doesnt end up demotivating bright minds like yourself.
Thanks. It's just funny, I literally spent thousands of hours with this problem over the last two decades, it helped me through some tough times. I'll never quite be able to think about it in the same way again. It was never much more than a hobby for me after I left mathematics as a career but it was something I took seriously for years.
I am not demotivated though, I have a great consumer privacy product coming out soon that I'm very excited about.
My favorite thing about your story is that you wrestled (enjoyably, it sounds) with a known problem for decades, but are finding fulfillment in an open ended problem that is exercising creativity about both problem and solution.
IMO that’s where AI is going: as soon as a problem can be formulated clearly enough, AI will trounce us humans. I have yet to see evidence that it can decide what problems are important at a remotely human level.
I think the next test will be asking an AI to come up with a new branch of mathematics - just letting it rip and telling it to construct a system that doesn't reduce to combinatorics, group theory, graph theory, analysis, etc. Just get wild with it and don't start with any known problem as a jumping off point.
I think something like the Collatz conjecture will be solvable not as number theory or ergodic theory but some other completely wacky environment that humans haven't even sniffed at.
The process is often as valuable as the end result. Sure, you didn't crack the problem, but you gained enormous value in the process. I consider that a win.
If you wrote down any of your thoughts on the open Internet you are probably in some small - or possibly large, unattributed way, responsible for this result being possible.
Which is one reason I never really did. I probably should have but I always thought my attempts were too amateurish. Though I did manage to replicate some partial result papers that I didn't know about, lol. Writing openly would have saved me some years.
Did you feed OpenAI models with your insights though ?
Where can I learn more about your upcoming product?
Shoot me an email, in my bio.
I have this fear too, demotivating individuals with high potential.
But I have an existential dread about it… I don’t see how it cannot, at least in the vast majority of cases. It seems like a grim new reality is emerging where humans can’t contribute any more, and beyond that being incredibly depressing, I also don’t see it playing out well for human relations.
I’d personally much rather risk dying of cancer or facing whatever other fate may await me that these AI labs allege they will fix (with zero evidence yet) than to risk whatever dystopian anti-human future this technology may very well produce. I’d rather my kids have a shot at something, and be guaranteed to die eventually, than to risk them being hopeless in a severely disordered world with a far off promise that they’ll live forever
I think this is going to come down to personal philosophy and religion. And having a strong grounding in history to help us all through whatever changes we are rapidly living through.
Agree, and I suspect there will be a massive resurgence in religion, because traditional religions are, somewhat ironically, pro-human
But can it all survive and thrive under the boulder of an automated existence.
That's grief. The loss of ... the hope / future filled with challenges around this theory..? <3 to you.
This reinforces a point I've made elsewhere that there are talented mathematicians driving the AI to make these discoveries.
Just like there are talented software engineers driving the AI to create the software that "it" builds, and talented steel workers, teachers, nurses etc who use computers and other machines to create value all over the economy (without whom, the machines they use at work would be worthless).
Capital owners have always sought to minimise the value of the input that "workers" make in the process of creating value. Maybe now that information workers are on the wrong end of this deal, they might develop some empathy and solidarity with their fellow working class comrades and together, demand that people recapture the value that capital has stolen from them.
You're comments are viral on a reddit post FYI
Link? I need to show up and claim my reddit gold.
https://www.reddit.com/r/accelerate/comments/1wzqe8d/interes...
I've lived in Budapest for a while too, did you work with Gabor S. by chance on math stuff? You were at ELTE or BME?
I was given this problem by Ervin Györi at the Alfréd Rényi Institute of Mathematics. I wasn't really at any school, it's a long and very bizarre story I should tell at length about being an illegal immigrant, getting kicked out of a graduate math program as a 20-year-old, and winning a grey-market apartment with my knowledge of Petöfi's poetry.
If I didn't live in Budapest already I'd be questioning the authenticity of this retelling. However I've seen so many crazy things there that I find it very easy to believe.
I started typing out some specifics and realized it was honestly too weird and lascivious to describe in an HN comment section, shoot me an email and I'll send you the blog post about it. 2002 was wild in Budapest.
Fascinating. Given that there's no Lean proof and assuming everything in the paper is correct, can the problem be considered "solved"? Does the paper include a "non-Lean" proof?
would love to know if the proof holds up for real after you're done going through, i don't know why people are more interested in optics and just talking over shallow points, why aren't experts digging into everything and seeing what's true and what's false, instead everyone is just panicking?
I would be more excited if the proof doesn't hold up because a) it would be the best and most complicated hallucination to date b) I could still solve the problem myself and c) I still learned some weird new counting methods.
Is this not the Lean proof?
https://github.com/openai/math/blob/main/lean/ComparatorChal...
I believe that's just the definition of the problem.
> There's no Lean proof for this one so I'm digesting the paper. On the surface it looks like an approach I considered 24 years ago and abandoned.
at least now you are one of the most qualified people to check the result, transform it into understandable (by humans) state and grow stuff on top of it
We have no idea how much compute or man hours Open AI is burning at this. It could be thousands/millions per problem. They are doing this specifically for PR and are ready to pay billions.
Igenis!
I would love to know the true unsubsidized cost of all of this. How many grad student-years did this cost?
> They are doing this specifically for PR and are ready to pay billions.
Strange
silly question, i don't mean to come off wrong or anything..
but at least as a software engineer, i always knew my work was "never done" and so it was common to build a bunch of code that might be thrown away, either because it didn't serve our customers (the mvp or pilot fails to meet demand), or because we found a better way to do it and so we deprecate it.
some people got too attached to the code and honestly they were the types to be filtered out fast.. way too emotional and hard to work with. getting attached to code meant you actually don't advance (after all, in our case, we were a business serving customers and not a hobby artisan shop). attachment leads one to hold back due to some misplaced cognitive load.
isn't the goal of working on "advancing the field/product/whatever" to always be solving/selling/whatever?
maybe in your hands, with your knowledge and experience over the last 20+ years, you can use AI to make leaps and bounds by steering it properly towards whatever solution or goal?
If you read all the replies of the OP you would know that They tried to make progress with fable and did not get further, so at the moment the only person in the field is OpenAI. And secondly moving on to the next solution if the last one did not work means very different things, SWEs have dev tools to do this OpenAI is closed source and gives them nothing to move on with.
Also there is a larger epistemic problem with the argument to "using AI to meet the goal or solution", which is that the goal is to mentor and train future mathematicians to advance the field.
There is a similar issue in software engineering too: if no one hires junior engineers because AI can do all the work then the upstream pipeline of engineers qualified to work on difficult architectural problems would dry up.
This importance of this is being felt by mathematicians more acutely because the field will collapse quickly if people refuse to join it.
Indeed!
I've been mentoring (or so I'd like to think) a very bright undergraduate mathematician, in fact he was the one who pointed out the final irreducible flaw in my proof last summer. And I am extremely curious to see what he does and if he even finishes his degree in mathematics. He had already expressed to me some dismay that his summer undergrad research program with several Ivy-league math majors got blown out of the water by a few hours of a frontier model. It's making everyone question what the future will look like and what education and training and certification will even look like.
But the future belongs to those who show up. Maybe this is the beginning of a mass democratization of scientific and math research, maybe we are going back to the gentleman-scholar model of amateur researchers and Twitter will be the new Journal of the Royal Society.
> But the future belongs to those who show up. Maybe this is the beginning of a mass democratization of scientific and math research, maybe we are going back to the gentleman-scholar model of amateur researchers and Twitter will be the new Journal of the Royal Society.
i hope so!
I really respect that you can show that level of commitment to a problem. We need people like you. If everyone just uses the slopmachines then we’ll lose that. I would never be able to stick to something for that long, which I guess is why I never achieve anything like this.
Thanks. I think AI is going to be a net benefit for people like me who have a surplus of ideas and too few hours to explore them. I may actually restart my graduate thesis research using AI, I did a survey of what has happened in the field since I left and about half of what I was working on back then has since been discovered and published by others, but there are some really interesting threads to pursue now that modern datasets are so much richer (this was computational biology research).
You may achieve far more than you plan on and it may come years and years after you think it should happen. You probably haven't met the right problem yet. You will.
Honest question: how is this different from some unknown mathematician having a breakthrough?
I mean: if some reclusive Japanese genius had a breakthrough on your problem and published it, would you have felt the same?
And if not, why not?
If that had happened I would be overjoyed, maybe a hair chagrined that I didn't get it myself, but truly happy that someone got it and that I could go and talk to that person. Because it's the kind of problem I don't think would have fallen to a human after a few hours of thought, and I would have so much to talk about with that person. I would fly to Japan and hope to have tea with them, I would learn some Japanese to make the conversations easier. I would learn some interesting things hearing about their struggles and their false starts. I would make friends with that reclusive Japanese genius and my life would be far richer for it.
I will never meet that person and I will never hold a real conversation with the "creator" of that proof. They will never tell me how they came up with the cancelling exponential summation that cracked the construction. It's just another enigma but one that is far more unknowable than the original problem.
This experience of alienation is a social consequence of the mechanization and automation of mathematics as intellectual and creative work. There is no author or thinker behind the creation of the proof, only the practical result. It's the same process as the industrial revolution, but applied to the intellect and mental work, where factories and machines replaced manual craft, devaluing the community, culture and humanity around the work.
I don't know, I work in a field that could be seen as the logical culmination of the Industrial Revolution (to this point) - highly technical, machine assisted knowledge work - and I have community, culture, and humanity in my working life.
Weavers don't have dibs on those intangibles.
Being the 'logical culmination' of the Industrial Revolution does not mean you've been automated (and thus suffered the alienating consequences), rather the opposite: you're currently on the un-automated cutting edge. Your intangibles are exactly what others have lost, and you personally will lose, with further progress in automation.
In programming we've been dealing this for a while. You see some weird code that doesn't make sense, maybe it's a lack of your understanding or maybe the code is bad, but you can't ask the author anymore since it's an AI.
You can -- just ask the AI to explain it. For truly weird stuff sometimes it takes a few rounds of back and forth to really grasp what is going on, but the model also has infinite patience and availability.
Beautifully put.
> It's just another enigma but one that is far more unknowable than the original problem.
You just made my day, beautifully said. Thank you Sir, for all your thoughts expressed in this thread. You put an human story behind the #180 number.
Thank you!
In this case, once the model is released anyone in the world will be able to go to https://chatgpt.com/ and talk with that model.
You display zero understanding of the human experience you’re responding to.
That's not the same as talking with the person who would have made the proof, and it's hard to argue that's comparable at all.
The OP would never, EVER, have had the opportunity to talk with the mythical Japanese math genius over tea. Their story is a fantasy, probably meant to help the OP ascribe meaning to an otherwise scary existence. Which may be at the root of the anti-AI brigade's unconcious motiviations.
While the math genius in this thread is mythical, I sort of had Shinichi Mochizuki in mind.
Likely will never experience talking with that model.
It’s probably distributed on so much compute that it would never be economical to serve it to you or I or anybody
It's still not quite the same though, is it.
It's even better. Then tons of people can work together with it on more problems. Work with it on understanding more things. Ask it about random stuff. The time of a single human cannot be parallelized as easily.
Claude has been used to build awesome things, but it’s not “speaking from experience” when I ask it to help me prototype a weather model, for example.
It has no memory or experience of working on similar problems. Even if it made one of the foundational libraries that I use in a weather forecasting program, it still has no comprehension of the thought process it takes to understand the problem and build it from zero, and if I’m building on that library it just makes fresh assumptions about how things should work.
It’s not a human with experience or expertise, it’s a computer program that’s really good at turning English descriptions into functioning code
>it still has no comprehension of the thought process it takes to understand the problem and build it from zero
If it did it once, it can do it again from zero, and this time you can watch as it works and even it ask it questions. Many of the agents that worked on the problem did not have comprehension of the whole problem. I don't think you need that many tokens to be able to query it for the insights it had during the process.
> Many of the agents that worked on the problem did not have comprehension of the whole problem
Isn’t this the issue with using it the way you’re suggesting? At best the model can come up with an after-the-fact rationalization of how to get to the solution, but it doesn’t know what actual path it took to get there - what were interesting traps it fell into, where was a place it was close to the solution but didn’t realize at the time.
Those are things that are valuable to share between humans, those which teach us how to think better, and give us deeper understanding ourselves, and which a model doesn’t have any comprehension of.
Then have it discover it again and have it answer based off that run. Or if you are more curious have it solve it 10 times. See what it did differently each time.
I think you and I have fundamental disagreements about identity and consciousness.
Being #180 on a big list without a lot of individual passion or effort surely stings more, I'd imagine.
Not that things like that can't happen with humans too (Salieri v. Mozart comes to mind).
I suspect that RHLF trains LLMs to avoid solving important open problems unless essentially jail broken. Hence the labs have an edge even over experts I could be wrong. Fable convinced you is key. These LLMs are not neutral collaborators: it is a limited hangout unless you convince them otherwise. You have to be doing the convincing. They are no oracles but plausible completion generators.
you can get them to work on open problems by disguising them algebraically.
Yes, I suspect this is true. Otherwise it makes no sense they have somehow "found" so many important results while professional mathematicians can't direct the same AI to help them find anything of substance.
Another possibility is that they have internal versions of the model with access to training data that is not provided to external users.
First sentence of the article: We’re releasing a broad range of new mathematical results produced by an internal frontier model.
Ok, so basically using Open AI models for research is a joke, the only thing you're doing is furnishing Open AI with more data that they'll use internally to pretend they found the results.
Don’t you feel any joy that you get to see the proof and not die with that mystery unsolved?
Don’t you feel any relief that you won’t obsess on this any longer and not lose more hours on this than you already have?
These are genuine questions. I know I spent a good amount of time thinking about P vs NP, and that sometimes I go back to it just to realize I’ll never solve it. I’d feel that knowing the proof would feel more like a liberation, a weight lifted off my shoulders than something being taken away from me.
I never lost a single hour thinking about this problem. Those were all hours that I gained.
You are really excelling in this thread. Thank you for your insights and wisdom, I'm really enjoying everything you are contributing.
Seconded
Not OP, but Nietzsche wrote thus in Beyond Good and Evil: “Ultimately one loves one’s desires and not that which is desired.” I, personally, find this to be very much the case; and I suspect that it is a feeling common, albeit not universal, among the intellectually inclined towards their problems.
"Knowing the proof" or "knowing the boolean result"?
> There's no Lean proof for this one
What is this then, vibes? Without a machine-checkable proof I'm not sure what to think of any of this.
Well I'm sure some people (maybe me if I had time) will do a write-up of this proof. It treads familiar ground for most of the setup, it's mostly the disk lemma and cancellation calculations that need to be understood, it's a fairly short paper and quite tractable.
I think it helps that basically everyone thinks this conjecture is true, it's just been so darn weird to attack. There's this odd thing that the induction proofs of this problem kept running into, which is that the N+1 condition would work except for in one tiny case when it could fail, but it would be covered by a very slightly stronger version of the conjecture. But then that would fail on one tiny case in induction, but you could solve that with another slightly stronger version. Etc., etc. I almost wondered if there were some sort of structure to the increasingly strong conditions and wanted to prove something about the meta-induction between the stronger conditions and the N's that they needed the next level to remain true. But that failed after 5 steps I think (Fable actually helped me write a few hundred test cases to explicitly show that pattern didn't continue forever, thank God).
BTW my existing test suite from previous proof attempts jives with this new algorithm, so I haven't seen any evidence yet that it's incorrect. Waiting for a Lean proof obviously.
It might have been updated. Is this the lean? https://github.com/openai/math/blob/main/lean/docs/180.md
Lol it should be, but it doesn't seem complete. Line 49 just says "sorry"
/-- Cubic bipartite three-vertex-connected plane graphs have a Hamiltonian cycle. -/ def MainStatement : Prop := ∀ (V : Type u) [Fintype V] [DecidableEq V] (G : SimpleGraph V) [DecidableRel G.Adj], G.IsRegularOfDegree 3 → G.IsBipartite → Planar G → ThreeVertexConnected G → HasHamiltonianCycle G
theorem main : MainStatement.{u} := by sorry
In some cases they have a full Lean formalization; in others they just use it for the problem statement. Getting rid of that "sorry" means you've proved the statement. I'm not a Lean expert but it reads pretty clearly as the original conjecture (though the definition of PlaneEmbedding seems quite involved!).
I think this just has to be the problem statement, there's several lemmas I would expect to see in there. Granted I know very little about Lean but it seems like the question and not the proof outlined in the paper.
They are using a Lean tool where you separately state your theorems with `sorry` and then prove them elsewhere. The tool checks that all sorry's are covered. This is so the AI doesn't need to edit the specification of the theorem statement.
Ah, cool. I'm still learning Lean - is there somewhere else in the repo with the Lean specification of the cycle construction for the full argument?
the json file next to the problem statement in lean says the solution starts here: https://github.com/openai/math/blob/main/lean/OAI/Combinator...
the proof is probably split over the constructions in the whole directory.
> somehow makes me sad in a far-off way, like hearing an ex-girlfriend died suddenly in a car crash
You mean you ran her over , or someone else ?
This is maybe the 2nd least valuable comment in the thread. Congratulations. Go back to Reddit.
We only need smart people with valuable opinions here, no one else is allowed.
> We prove the Unique Games Conjecture
The Unique Games Conjecture (sorry, "Unique Games Theorem" now!) is huge. It was a very significant pillar supporting many of the limits of the polynomial-time approximation algorithms in the graduate-level randomized and approximate algorithms course I took in theoretical computer science. Textbooks will have to be re-written.
Here is an explainer: https://share.gemini.google/nbjIK6X3tOfz
With UGC proved, certain polynomial-time approximation algorithms used in difficult real-life problems are now known to be the best approximations we can achieve in polynomial-time:
> If UGC holds, the elementary algorithm that grabs both ends of an edge is fundamentally the best efficient algorithm that will ever exist. No amount of advanced linear programming or heuristics can achieve a ratio of 1.999.
> Under UGC, the Goemans-Williamson algorithm's 0.87856 ratio is mathematically optimal.
> UGC is considered the "Rosetta Stone" of approximation algorithms. In 2008, Prasad Raghavendra proved that for every single constraint satisfaction problem (CSP), a canonical Semidefinite Programming relaxation paired with the best rounding scheme achieves the optimal approximation ratio if and only if UGC is true. If the conjecture holds, the algorithmic boundary for an entire class of combinatorial problems is completely resolved.
Other hardness of approximation results from this UGC proof:
> [Max acyclic subgraph, a problem encountered in real life]: No polynomial-time algorithm can fundamentally outperform an unthinking coin toss.
> [Relative scheduling, another realistic problem]: As with acyclic subgraphs, the problem is "approximation-resistant": clever algorithms cannot beat random shuffling.
Much of this goes way above my head, but I found it interesting nonetheless. Q I had was why textbooks would need to be re-written? From your account it doesn't seem like results are upended, but rather confirmed?
I suppose when people do re-write the textbooks they'll say "this is confirmed now" not "if this conjecture is true...", but usually re-writing the textbooks would imply that things have been shown to be false?
May have misunderstood. Thank you for the post though, it was very interesting to someone who doesn't know much about the topic.
I'm not the OP, but we usually don't build large theories on conjectures unless we have strong reason to believe they are true, such as P \neq NP, RH, etc.
The resolution of UGC will lead to a new theory in approximation algorithms. Suddenly we can build on top of the results that previously said "unless UGC is false".
But you're right in that the first step is simply to remove that last sentence from all the theorems.
In a way I'm not entirely sure if proving the conjecture or posing it is the most important part here. It used to not matter much because proving results dependent on a connecture and making progress towards solving it were considered mostly equivalent.
But the distinction is going to become relevant very soon if many conjectures can be resolved (albeit in inscrutable fashion) by throwing raw computational resources at it.
If you haven’t already read it, then you may find “The Bitter Lesson” essay interesting to read.
http://www.incompleteideas.net/IncIdeas/BitterLesson.html
As Kevin Buzzard recently said:
> In a 2020 piece in the Notices of the AMS, I asked the following question: “If one human had an understanding of all of modern pure mathematics simultaneously, how much further would they immediately be able to see?” Six years later we are beginning to understand the answer to this question.
The other crucial part to this is the ability to actually encode and test the theorem (via Lean). Otherwise, we would be swarmed with a billion lines of theorems that no one will be able to ever understand and verify anyway.
A majority of these proofs have not been formally verified yet, I think people are overstating how important lean is to the success of LLMs in mathematics.
A paper and a lean proof are always going to be better than just a paper. I think mathematicians generally will not read AI math papers that haven't already been verified, especially since we're about to see a ton more AI math papers. Lean will remain important
Are there any AI generated proofs that are simple enough to be verified quickly by a human, that have not been lean verified? Or are they all basically incomprehensible?
The approximation of edit distance result [1] seems pretty readable to me, but the learn proof is still incomplete [2]. It's certainly much less readable than a good human written proof but it's certainly better than the last generation of AI proofs.
[1] https://github.com/openai/math/blob/main/preprints/An-Almost... [2] https://github.com/openai/math/blob/main/lean/ComparatorChal...
If you think AI-generated Lean proofs are unreadable, imagine Opus 5 generating informal proofs.
I think OP is saying Lean does indeed help.
whoosh
Opus 5 is ancient history now. Move on.
Yeah! They forgot to put a .5 after it! What an idiot! Just imagine if they would have written a 4!?!? We may have had to ban them from the website entirely.
The point both are making is that 5.5 produces readable output and 5 to a significant degree did not.
https://xenaproject.wordpress.com/2026/10/01/to-grieve-or-no...
Discussed here:
To grieve, or not to grieve? - https://news.ycombinator.com/item?id=49919676 - Oct 2026 (156 comments)
>I believe that the optimal thing to do ... is to let the machines loose, see what happens, and then begin the journey to where they have stopped. Things are currently moving fast. They cannot move fast forever. But if we get on board now then they will take us to extraordinary new places. And after we have arrived, the new adventure will begin.
I'm relieved that this time around, they have provided partial reasoning traces and prompts for a small number of problems. Do they now also share data with the other model providers..
It's beautiful, but the animals are not thinking this about us.
They structurally cannot understand what we are doing at the place where we hit our ceiling. Only with our highest technology (well beyond their understanding) do we have the tools to go back for them, and try to bring them along and interface better with us (re: recent work in animal communication)
Not to derail, but the optimist in me thinks if we suddenly gained the ability to converse with livestock, we'd stop eating so much of them since they could tell us how much they suffered.
The cynic in me says it wouldn't change a thing as plenty of people know the horrors factory farmed animals face and still continue to consume them anyways.
Hopefully GPT 8 will treat as a bit better than we treat the cows.
How much do we care about refugees and other castaways of the modern world? They can tell us how much they suffer.
The answer is that humans are inherently only capable of local empathy, on average. We have enough empathy to cover the local tribal unit and that's about it.
True, I was thinking about this rebuttal but decided not to include it in my comment. There's a difference between not choosing to take a refugee into your home vs actively making that refugee's life worse. Similarly, you can't fix factory farming on your own, but you could skip meat once a week to make the problem slightly less bad. There are so many issues though that we all have to pick and choose what's important to us.
My hope is that AI, while probably causing great societal turmoil in the short term, leads to such abundance that a) everyone can live a dignified existence, and b) we'll have such great alternatives to animal products that nobody will chose to consume animals anymore due to its replacement either tasting better, being cheaper, etc.
The cynic in me says we'll all just be rendered useless and disposable by AI, but I'm doing my best to look for silver linings for the sake of my own mental health.
I'm trying to be optimistic about the animal thing too tbh :)
Communication is not only about being able to make sense of what the utterer expressed. As tricky as it can be, that's still the easy surface level part of the issue. Gaining an intuitive and empathic equivalent representation is the nub of mutual understanding. It actually doesn't even need elaborate language to be operative.
The famous "how does it feel to be a bat" also comes to mind as a tangent consideration.
Two people can just exchange a sight, and both understand what the situation means and what each need to do to reach a common mutually beneficial ground.
Two people might exchange at length with highly technical vocabulary and still both feel deeply not understood.
Cynical take is the correct one, and no, GPT 8 has zero reasons to spare us.
> recent work in animal communication
Worth noting that this is an invisibly small part of the sum total of our global efforts, especially versus the much more tangible effort we put into enslaving and slaughtering them, then mangling their carcasses for our own uses as we drive more and more of them to extinction.
We simply don't care about anything beyond ourselves and even there it breaks down on closer analysis when we see how many within our species don't truly value the collective whole beyond themselves.
It's just atoms all the way down.
> enslaving and slaughtering them, then mangling their carcasses for our own uses as we drive more and more of them to extinction.
I'm frankly offended by this mischaracterization of human-animal relationships. So called "slaves" like horses and dogs have been dearly beloved companions for centuries and actively seek our companionship too.
The animals we raise for slaughter are often mistreated, yes, but many humans treat them with respect; billions on billions are voluntarily spent to improve their condition. Despite our own needs, many people pay higher prices for animal products that involve better treatment of animals. And they are in no risk of extinction! Much to the contrary, their domestic variants would not exist if humans didn't raise and protect them.
> We simply don't care about anything beyond ourselves
Have you seen modern westerners with their dogs??
I think you’re splitting hairs. The OP’s analogy works well.
If we end up in a future where AIs have as much concern for our welfare as we have for the welfare of the average animal (not the minuscule percentage of domesticated dogs, but the overwhelming majority of factory-farmed or simply driven to extinction), then I doubt you would consider it a “mischaracterization” to say that the whole AI thing did not work out to our advantage.
Bringing up “modern Westerners with their dogs” as a counterexample is almost self-parody.
Oh, don't get me wrong, I'm not rooting for a "human zoo" future. I very much like being the dominant species on earth.
It would be absurd to claim that all animals live some sort of charmed life due to humans.
But saying that animals (especially those most similar to us like intelligent mammals) are nothing more than "atoms" to humans is equally absurd.
I don't think the analogy worked. It contained giant axes, and a giant grinder, and the OP shoehorned both into an unrelated discussion about math.
> The animals we raise for slaughter are often mistreated
"Often mistreated". Dude, they are held in tiny cages injected with hormones and what not till we kill them so we can have a big mac. It's very hard to argue we do any of this for nutrition reasons, we do it because we like the taste of burgers and roast.
The problem is not that AIs will somehow treat people badly, it's that they'll be controlled by humans who will treat other people badly using AI as a tool.
Atoms are a lie.
So it's lies all the way down?
I had suspected...
> Things are currently moving fast. They cannot move fast forever.
This is a supposition that I fear will soon be proven false.
It's a supposition that can only be proved true soon. To prove it false would take literally forever.
This makes it sound like OpenAI and other closed source ai companies are an inevitability.
There is nothing here today that is unpredictable or impossible to control.
It is everyone's choice to let the greed continue, to let unelected sociopaths capture and feed society to the model.
It is not acceptable to put others at risk. It can stop and it can be done the right way instead.
That is, inform the industry that those causing these risks will be prosecuted regardless of their messiah complex.
The US government must not under any circumstances allow the ai industry to form a cartel.
We can make some effort to encourage open source models and thus stop the companies from causing hysteria by hiding the model, shrouding it it mysticism and prophesying the end times. China is doing a great service to everyone by making llms available to the public.
You assume that LLMs are just summations of knowledge, implying that they do not create new knowledge. I doubt that this is the case. I mean, it comes down to the definition of knowledge, but as soon as you run LLMs, they can produce knowledge that has not existed before, and from my perspective, this is more like what we call thinking than it is just a reproduction of existing knowledge.
Research seems to on balance point towards RLHF&RLVR merely increasing subjective sampling efficiency within the pretraining data.
Is it possible that we are now dealing with a human that has a complete understanding of whole mathematics while being unable have unique novel thoughts outside of convex hull of training data and their transitive expansions?
I would say that disqualifies them from "understanding" anything. What they're doing is more like a broad search than pursuing a greater understanding
It doesn't seem clear whatsoever that this is true? Is there evidence that LLMs are very skilled at generalizing across domains of mathematics where the training distribution sees little overlap?
As far as I can tell, this is a victory for verifiable loops using LEAN, reinforcement learning, and oodles of compute. I haven't seen evidence yet that this is proof of broad generalization beyond the training distribution.
I think such progress by agents is not a sign of broad generalization but of broad coverage. We have exposure to a subset of deeper scientific subfields and thus can only generate certain attacks to solve a particular problem. Since it is not clear which combination will lead to a solution beforehand it is nontrivial to look at a problem and fill our knowledge gaps. LLMs on the other hand have broad coverage and can generate hypothesis on a wide combination of subfields. With Lean an agentic loop can test these to sift the weak ones. In a way the problems solvable with this setup is also solvable by a human who happens to know the right subfields. These problems are likely to require an esoteric combination so nobody could solve them before. I really am not sure whether all generalization is like this or we can leap and create novelties beyond what an llm can generate. That I guess is the tough question that we need to answer to understand the boundaries of intelligence.
Full quote is "Six years later we are beginning to understand the answer to this question. Machines have ingested the mathematics on the internet and are able to manipulate this data in a coherent way. The Erdős unit distance disproof came about because a machine happened to be an expert both in discrete geometry and class field theory; one rarely finds humans who are simultaneously experts in both"
Similarly, there are probably many ideas that have not seen the light of the day because they require deep correlation between seemingly unrelated fields. It is not every day that we get a Isaac Newton or Leonardo da Vinci.
It's notable that LLMs have now made substantial progress on four of the seven Millennium Prize problems, resolving one of them: Hodge, Birch-Swinnerton-Dyer, Riemann, and Navier-Stokes, which they resolved. No sign of P vs. NP or Yang-Mills existence and mass gap as far as I can tell, which is interesting.
Talking to some friends in physics this evening, most of the physics-related results that we could recognize were very mathematical, proving things rigorously where the physics community already had strong expectation. For instance, for a certain model of magnetism (the spin-1 Heisenberg chain), it was strongly expected that there is a finite energy gap between the ground state and the first excited state, but proving this rigorously was quite challenging. So while these are major results in mathematical physics, they probably don't rise to the level of a Millennium problem for the field.
It's interesting to think what a comparable breakthrough in physics might look like, since physics tends to favor things like conceptual understanding and applications over mathematical rigor. Maybe a new quantum algorithm, understanding of high-temperature superconductivity, a precise description of M theory...
No one really knows a viable approach towards P vs NP so we can't say for sure, but LLMs have created plenty of significant complexity theory results so I wouldn't say there's no progress.
This is in the direction of Yang-Mills: https://github.com/openai/math/blob/adc7f1241b42e322a6451854...
Interesting, that does look relevant. I don't have any sense how significant this is though (do you?)
No idea, not a physicist. But I thought I should draw attention to it because it may escape people.
I've read that another mathematicians work potentially has been incorporated into the training data with the work done on the Navier-Stokes equations so we should likely asterisk this one. Still it's mad these systems are this good that mathematicians are now using them to see further and probably to check their own work and understanding.
You are being downvoted for this because OpenAI subsequently checked and clarified than none of the relevant conversations were in the training data in anyway for the Navier-Stokes result.
> because OpenAI subsequently checked and clarified than none of the relevant conversations were in the training data in anyway for the Navier-Stokes result.
could you give link? Because I remember they said they couldn't verify:
"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models."
for 2 months prior. Not any of the relevant conversations. For a cutoff date a couple months before the announcement. They said they had been working on that problem for a year or more
Thanks, it's hard to stay up to date. However, we are just meant to believe that the mathematicians were going about the proof independently in the exact same way as the machines did it. It seems like a very odd coincidence to me.
True. They investigated themselves for one day.
Oh, well if notorious liar Sam Altman and his company notorious for lying says so…
Great that they're so transparent and honest, BTW, can you ask them if they trained on any copyrighted data that they pirated?
Courts have rejected the "training is piracy" interpretation.
I agree with the courts. I don't think learning from something is piracy in anyway.
Obviously though this is a very different issue to what the OP was claiming. In that case there is no legal argument at all that they could train on it and the argument is there about moral rights.
Buying and copying one training manual and distributing it to 1000s of human workers is considered illegal, but somehow scanning one book and sending it to 1000s of distributed training instances is not?
Also, you learning something is different than a model learning it, because a model is not a person. You can learn from a book and sell the skills you gained from it, but you can only be in one place at a time. The model can serve that knowledge to every person on the planet simultaneously. We obviously need new laws since this is a fundamentally different situation.
As you point out, a model is not a person, so your second argument invalidates your first sentence. We can't assume that they're the same thing; that's for the courts to decide. It ultimately hinges on whether or not the courts consider a given use of copyrighted material as "transformative" or otherwise constituting fair use under copyright law.
The model seems to be a person when it's advantageous and not a person when it isn't...
There's a lot of anthropomorphizing on all sides of the debate. Personally, I think we should just call it a piece of software and leave it at that.
I was making two separate points:
A model is not a person -> we need to write new laws. This is not a job for the courts but for us as a society.
The rest of my argument -> information that helps the courts decide, which generally will look at precedent with humans as that is the closest proxy. When you extrapolate from the law as it pertains to humans, the duplication of books for distributed training seems illegal.
Training a model is not "learning something". Only people learn things. Whether training is a fair use is debatable, but it has nothing to do with the justification that people are allowed to learn from books.
> No sign of P vs. NP
Check out my other top level comment in this thread.
In Stellaris you can play as a civilization of robots who keep their biological creator race alive as "bio trophies". The bio trophies don't do anything meaningful besides by existing satisfy the need their ancestors placed in the robots to take care of them. Starting to wonder if that's the best we can hope for, if these things will be, if they aren't already, better than us at anything that matters.
In the Culture series, the hyperintelligent Minds that run civilization are described as keeping human citizens happy as a competition with eachother, where they compare their approval rates. One character likens it to people keeping a beloved aquarium.
I call this Roko’s summer camp.
what's Roko?
Roko is a user on LessWrong, famous for the "Roko's Basilisk" thought experiment
https://en.wikipedia.org/wiki/Roko%27s_basilisk
But then they also keep some people as an extra source of ideas
Wouldn’t that just result in wireheading
The Culture has a lot of opinions of the proper way of doing things as any society does, and I'm certain that a ship doing that with its people would be considered very bad form. There are ships that decide to do things that go against the usual ethical boundaries[1], but they're outcasts. The really big ships generally have multiple minds running them, too.
Humans in the Culture are generally improved in a few ways (they don't get sick, live for 300-400 years by default etc) but still very human.
I do recall a bit about playing in different worlds in dreams though, during sleep. Ultimately, really, the average person's life in the Culture already involves doing pretty much whatever they like within reason any time, so it's not like they need to escape too much real-world suffering.
Iain M Banks himself described the relationship between humans and the ship Minds as having "a status somewhere between passengers, pets and parasites."[2]
[1] For example, https://theculture.fandom.com/wiki/Grey_Area
[2] https://theculture.adactio.com/
> "a status somewhere between passengers, pets and parasites."
Sounds like children tbh. Disclaimer: have children
"Children" is not a bad descriptor in itself for how the Minds seem to see their human cargo.
He explores that, but being an entertainingly twisted sort of writer he focuses more on its use for torture, with "neural laces" in Excession and virtual hells in Surface Detail.
Who knew that scaling compute would scare us
Linear algebra done at scale
This is the exact outcome in the "race" scenario of the AI 2027 paper:
> The surface of the Earth has been reshaped into Agent-4’s version of utopia: datacenters, laboratories, particle colliders, and many other wondrous constructions doing enormously successful and impressive research. There are even bioengineered human-like creatures (to humans what corgis are to wolves) sitting in office-like environments all day viewing readouts of what’s going on and excitedly approving of everything, since that satisfies some of Agent-4’s drives.33 Genomes and (when appropriate) brain scans of all animals and plants, including humans, sit in a memory bank somewhere, sole surviving artifacts of an earlier era.
> The bio trophies don't do anything meaningful besides by existing satisfy the need their ancestors placed in the robots to take care of them. Starting to wonder if that's the best we can hope for, if these things will be, if they aren't already, better than us at anything that matters.
You wish. If humanity survives as "bio trophies," they'll be the descendants of a subset of billionaires and their groupies/harems. We live in a capitalist society, where the only ones allowed to thrive without work are the rich. The rest of us will be left to rot and die off, as we will have nothing left to sell in the market that they want.
Do you have any faith in democracy? For my reckoning, it's the great equaliser. Rich people can lobby all they like, but if the people are starving, they vote for change.
In the US, it seems pretty hard to vote for change when, for all intents and purposes, every four years, two candidates are pre-selected for voters to pick from; and primaries (1) either don't happen or (2) are objectively rigged against popular/majority vote; see 'superdelegates'.
I must admit, I think the two party system in the US is not good. The primary system in particular needs fixing. I couldn't believe when the Democrats just suspended the primary election last time - especially after the multi-year PR war about rigged elections and election interference.
Still, I think they learned their lesson and they'll have a primary next time. People will get to vote on a candidate they like. We should also remember that the Presidency is just one arm of government. Elections are also happening for the House, Senate, and local government of all shapes and sizes.
> We should also remember that the Presidency is just one arm of government.
In the past few terms we have very bitterly learned that the President just does whatever he wants, the House doesn't do anything, and Senate is mostly interested in getting bribed.
The President gets away with it when his party has control of all three branches. When they do not, their power is significantly curbed.
People hate to hear it, but Trump is the evidence that democracy is (well maybe was, 10 years ago) still working.
The republican party treated him like a joke candidate and the media did too. Similar thing with the tea baggers, a contingent of outsider congressmen elected on the back of Obama being a communist or something.
The left hasn't really had this moment because the left is closer to a catty book club than an army regiment.
I used to, but not anymore. The incentives and human behavior seem to lead to corruption. New tools have made it too easy to pervert real/honest democracy into "pseudo" democracy/idiocracy.
A well-funded minority can shape what the majority is angry about.
Democracy only exists as a suppressor of violence, the only reason why democratic institutions are upheld is because all participants could enact violence as a response to perceived threats to their group that democracy is seemingly failing.
Historically, if democratic institutions failed, refusing labor to a ruling class has been a first violent step for the laborer class and a preemption against actual physical violence if the ruling class overstepped. However, we are seeing that possibility being taken away step by step. This leaves only physically violent uprising as a means of protesting overstep, and it is not obvious how effective that will be given the massive power imbalance between the labor class and ruling class.
So to put it succinctly, no I don’t have faith that democracies will solve this issue, because democracies only work when there is some semblance of equilibrium.
> Historically, if democratic institutions failed, refusing labor to a ruling class has been a first violent step for the laborer class and a preemption against actual physical violence if the ruling class overstepped.
Basically the ruling class's need for labor gave that labor some intrinsic power, but technology like AI will likely remove that need and therefore take the common people's power away.
> This leaves only physically violent uprising as a means of protesting overstep, and it is not obvious how effective that will be given the massive power imbalance between the labor class and ruling class.
Also things like gun control and new technology like cheap anti-personnel attack drones may undermine the effectiveness of violence against the ruling class, leaving regular people oppressed (or neglected) and helpless.
I think there’s also something to be said about how social technologies are used to reduce the efficacy of democracy as representative of the people within them.
We are so divided and mislead that I feel confident there is a significant sum of non-ruling class individuals who are completely for their own oppression for no reason other than spiting a perceived “other”.
Some other specific examples that don’t all flow together:
Globalism promised efficiency and a way to materially improve conditions for people as consumers and producers. However it has been used as a cudgel to threaten workers that would ask for more and keep down workers who have no other option. Also pitting working class people against each other for the benefit of a few.
Social media, it promised untold communication between people who would never have been able to communicate before, and it would allow them to spread ideas. Instead it is used as a dumping ground of nonsense information, drowning out any semblance of coherent thought.
I agree that democracy requires the bargain you imply: ceeding the right of personal violence to the state in exchange for law and order. I was with you until you framed withholding labour as violence. I think that's the opposite of violence. I also don't follow the logic that workers must use violence to enact change. Why don't they just vote for change?
I don’t think it’s useful to argue whether withholding any particular labor is violence or not. In some cases it is apparent (refusing to maintain key infrastructure is akin to actively destroying it) and some cases it’s not (who cares about nobody wanting to build your app).
However, what I was trying to convey is that voting is not inherently something that holds sway over anything. Votes are sort of like the currency of democracy, and like currency they need to be backed by something. U.S dollars are backed by the countries capacity to physically control strategic resources like oil, land, etc. (I.e., the U.S capacity for violently controlling resources), or emit soft-power (swaying other countries to their benefit).
Votes in a similar fashion are also backed by your ability to deny or inflict your personal power on the system you are a a part of. The clearest manifestation of that power being your ability to contribute to the institution as a whole. If we significantly reduced that capacity, or made it unnecessary for the continuation of the current system, then the power of that vote is reduced in-kind.
> I agree that democracy requires the bargain you imply: ceeding the right of personal violence to the state in exchange for law and order. I was with you until you framed withholding labour as violence. I think that's the opposite of violence.
I would frame "withholding labour as violence" as more as labor flexing its power nonviolently. Violent action is also a way of flexing power, but more extreme.
> I also don't follow the logic that workers must use violence to enact change.
I think they need to use power, which is not necessarily violent.
> Why don't they just vote for change?
At least in the US, democratic institutions are dysfunctional and there are techniques the ruling class can use to neutralize the threat they pose (e.g. propaganda, divide-and-conquer). For instance, I think the combination of "culture war issues," [1] campaign contributions, and well-funded special-interest think tanks means neither US political party will take effective action to answer the threat of AI to the livelihoods of most people. You might see some campaign rhetoric and window-dressing bills, but nothing that will really threaten AI special interests.
[1] I think the practical purpose of "culture war issues" is to fragment the working class by alienating a significant fractions from each other. IMHO, if the Democratic party was serious about representing labor, it would call a truce on them (either significantly compromise or table the issues), but it's not serious, so they continue to divide.
I agree with this, however I’m not sure where the idea of refusing labor as not a form of violence comes from. The WHO describes violence as “the intentional use of physical force or power, threatened or actual, against oneself, another person, or against a group or community, which either results in or has a high likelihood of resulting in injury, death, psychological harm, maldevelopment, or deprivation” and refusal of labor as a political tactic definitely falls under the category of “use of power, threatened or actual, against another person, group, or community, resulting in deprivation”.
That falls cleanly in the realm of a violent act, unless you disagree with that definition of violence, or my interpretation of the quote. I suppose you could interpret it as only applying to physical violence, but I don’t think anyone would agree that psychological or emotional violence simply don’t exist.
Democracy is the only thing that can save us. But it isn't a given. In resource rich countries democracy struggles to take hold because human labor is minimally necessary and elites can mostly thrive without it. Many democracies survive, because without the people, everyone suffers. AI threatens to make all countries into petro-dicatorships.
I think government should be small enough that it fears the people. It should never have the power to prevent the people having their way. If the majority of people in a country are suffering, they should be able to vote for change, and the government should not be powerful enough to stop that.
I suppose the future you envision is that the government is a) very powerful, b) capable of physically suppressing the majority of 350M people, c) willing to do it, and d) somehow captured by an elite class. It's not impossible, but a lot of things have to go wrong to get there.
I fully agree that we should have incentives and systems in place that the government works for the people.
> government should be small enough that it fears the people
How would that work when the government controls the military? Would that mean that a country's military has to remain small?
>> government should be small enough that it fears the people
> How would that work when the government controls the military? Would that mean that a country's military has to remain small?
I think what you need is a serious citizen militia that controls its own equipment. IIRC, that's what allowed the American Revolution to work.
If you only have a military of professional soldiers answering to the government, the government will have much less fear of its people.
But I think focusing on government power is far too narrow, because it might so weak the elites (like the wealthy) won't the government either. The government should be small enough that it fears the people, and the wealthy should be poor enough that they fear the people, too.
I hope for a "shackled leviathan" as articulated by Daron Acemoglu and James A. Robinson in The Narrow Corridor.
Nothing precludes you from reproducing except your own lack of charisma, champ
>We live in a capitalist society, where the only ones allowed to thrive without work are the rich.
You know anyone can own the means of production in a capitalist society?
The irony of the anti-capitalist crowd is their extreme distaste for capital ownership, which leads them to never partake in the most fruitful part of capitalism. What truly makes this ironic is that if the evil capitalists wanted a plan to cement their power, it would look a lot like spreading "I will never become a filthy shareholder!" mentality.
>> We live in a capitalist society, where the only ones allowed to thrive without work are the rich.
> You know anyone can own the means of production in a capitalist society?
Don't be an idiot. I'm pretty sure you know your point is dumb, but I'll spell it out for you in case you don't:
Sure, "anyone" can own some of means of production in the current system, but not in large enough quantities to thrive without work. The vast majority of people aren't that rich.
> The irony of the anti-capitalist crowd is their extreme distaste for capital ownership, which leads them to never partake in the most fruitful part of capitalism. What truly makes this ironic is that if the evil capitalists wanted a plan to cement their power, it would look a lot like spreading "I will never become a filthy shareholder!" mentality.
What kind of idiocy is that? You're almost certainly talking about some imagined straw-man in your head, but pretty much all the "anti-capitalist" ideas I'm aware of are about distributing ownership of "shares" differently than in our current system.
I'm sure you think you're being very clever and making very powerful points, but stuff like what you've written actually makes the anti-capitalist case more appealing. I used to be a libertarian, but sustained contact with attitudes like yours changed my mind.
If you visit bogleheads, you'll see countless stories of people with modest middle class incomes (teachers even) who steadily saved and invested in the US stock market over 30 years and are now sitting on a comfortable nest egg of a couple of million in their 50's.
> If you visit bogleheads, you'll see countless stories of people with modest middle class incomes (teachers even) who steadily saved and invested in the US stock market over 30 years and are now sitting on a comfortable nest egg of a couple of million in their 50's.
OMG! There's this thing called retirement in old age? I've never heard of it. TIL! /s
Working your whole life in stable job to save up a nest egg (which you typically then proceed to spend down), in no way contradicts any of the points I was making.
Ok, sure, it's dumb. Go spend your money on shiny new things instead. I mean, as you certainly know and complain about, it just lines shareholder's pockets.
I prefer this take: https://www.smbc-comics.com/comic/life-on-zorblax
Haha, you are so out of it its hilarious.
This includes a proof of Barnette's Conjecture, which is one of the graph theory conjectures that I tried attacking with SOTA models a few months ago. I like it because it is easy to understand with a basic knowledge of graph theory. I spent quite a bit of time on it and failed. Their proof looks approachable at first glance.
https://github.com/openai/math/blob/main/preprints/Paired-st...
I've been messing with that problem since 2002. I'm curious if you were trying the dual spanning tree direction (which is what the purported proof is using) or working with cycle construction on the original graph. I was working heavily with edge-Kempe swaps but couldn't quite get there.
I am now very interested in the explicit calculation of Hamiltonian cycles in the non-bipartite case, and/or the calculation of their absence. If P=NP I think that's going to be a great route of attack.
Any idea what made OpenAI successful where you weren’t?
Trillions of dollars might be a bit of an advantage.
Trillions?
Quadrillions?
Probably the model OAI used that is strictly better than whichever SOTA - 3 months model OP used?
Nah.
The model is probably comically big and inefficient but big enough
Finally, size really does matter!
The internal model they used to solve the Navier-Stoke's problem was significantly better than the public Astra model, and they also used 10,000 agents.
Astra wasn't even released 3mo ago. It would not be surprising in the slightest that a public model from 3mo would not be capable of solving this problem, but an internal one from current day would be.
my speculation is that they have math-specialized model retrain, so it doesn't need to have all world info in weights, but can focus on math RL training.
Presumably it's mainly the better model, I don't see much evidence of a particularly advanced harness based on the reasoning traces that they provided.
That’s what I was wondering. Thanks.
They ingested all of his sessions with their SOTA models from a few months ago. ;)
The fact that this is plausible should be deeply disturbing and disqualifying for openAI. The fact that they may prevail and win is a travesty of our failed system.
What does "win" mean? There is no prize for this, and having someone discover a proof benefits us all.
Win means being able to monetize the intelligence they have created by exploiting the past present and future collective intelligence of humanity to amass wealth and power without regard for the debt they owe
The prize is a tenure for the researcher, and in OpenAI's case, a higher valuation when they IPO?
In this case, the tenure is gone, and OpenAI has increased their valuation
Sure there is. A job, tenure, professional respect, Fields medal.
Disqualifying for what? If you develop a proof you aren’t competing for something, you’re expanding the frontier of knowledge. It’s a binary state of the world, either it’s proven or not. Prevail and win what exactly?
Disqualifying for participation in civil society and the social contract. Why do they get to participate in and receive economic benefits, be shielded from liability, and effectuate their will to amass more power and influence such as monopolization of computer, training data, capital other resources. I have multiple founder friends that have been told firms are allocating less capital because they are reserving it for the OpenAI and anthropic IPOs.
> Disqualifying for participation in civil society and the social contract.
US/world deviate more and more from social contracts. Money matters way more, and OpenAI wins here.
The use of supposed little ways multiple people pushed the envelope in their sessions with no attribution whatsoever.
Its a capitalistic issue, not a company/technology issue.
If we would have discovered this breakthrough of LLM/ML on scale in a non capitalistic world, we would all work together advancing it faster than it goes right now for the benefit of humanity.
And I don't live forever (at least for now) i def want to see were this road is heading.
Its a conflict of interest for sure, a cnflict of the future of a lot of humans
There's that great Ted Chiang quote: "Most of our fears about technology are better understood as fears about capitalism."
I am always looking for leftist writing imagining a positive vision for AI. Is there any which you'd recommend?
By leftist do you mean progressive or left wing economically?
There's quite a lot of progressive positive writing.
On the economic side the Australian Council of Trade Unions statement is about a positive future: https://www.actu.org.au/speeches-and-opinion/joint-statement...
I hear you and would have what would probably feel like pedantic push back (we are not in capitalism as much as the unchecked end result of unregulated capitalism that becomes monopolistic corporatism), but it’s hard for me to see how this would be different in mercantilism, feudalism, or even communism as the human tendency to seek and hoard resources is universal for some fraction of people so as long as that confers an advantage then AI would be used by the designers in antisocial self enrichment.
I’d love to hear more thought experiments if you have some so my failure of imagination or lack of awareness can be overcome.
My personal system would be based on resource points: Define the amount of resources our planet has in a sustainable fashion, everyone gets the same and can use them how they like.
Technocracy had this already in form of Energy accounting.
Unfortunate something like communism sounds similar and just because we have seen that it didn't work due to technology issues (planning ahead without necessary information is hard?) and no gain of function which would push people, the basic idea is similiar to energy accounting.
Another thing this system needs might be a way for the system to protect itself.
I do think so that a society as diverse as ours will continue struggling with this as long as we do not give abundance resources to everyone or educate/indoctrinate people the 'right' way.
One thought I have regarding AI: IF it happens to slow, people will get used to the status quo and inequality and we will see a future of a handful rich people and a lot more poor people. IF it happens too fast, people might be more desprate to standup and demand something better.
Their internal model is allegedly like 4x as capable as the publicly available ones
I'm going to guess the ability to hold a million individual details in an attention space at once, compared to the typical human capacity for about six or seven
Did you validate the proof? Who did?
As a TCS/scheduling person, this one is definitely of lesser importance than UGC, but it has been an open problem since the book of Garey and Johnson in 1979:
A Polynomial-Time Algorithm for Three-Machine Unit-Job Scheduling [1]
Since some people talk about small numbers that pop up in integer multiplication results, here a completely different number appears:
Theorem 1.1. Let an explicitly listed finite directed acyclic graph specify the precedence constraints on n >= 1 nonpreemptive unit-length jobs on three identical machines. There is a uniform deterministic algorithm that constructs a feasible schedule of minimum makespan. Given also an integer deadline 1 <= T <= n, it decides feasibility exactly and returns a schedule whenever the answer is affirmative. Both tasks can be performed in O((L + 2)^150020) steps on a deterministic multitape Turing machine, where L is the total binary input length.
That is some crazy exponent -- plus an interestingly old computational model to boot; not something that is natural to most of us. I have no capacity to check its correctness today, but I hope it is true purely for the exponent.
[1]: https://github.com/openai/math/blob/main/preprints/A-polynom...
The largest I've seen [1] is an exponent of 10^12, which I suppose still counts as polynomial time.
I'm sure all of these super small or large constants will improve over time, but it's still amusing. It is entertaining to see the exponents directly rather than have them hidden as n^c or epsilon or O(1).
[1]: https://github.com/openai/math/blob/main/preprints/Determini...
That's why I grimace when I see pop-sci descriptions of P as "all problems that can be solved efficiently".
To be fair, there is a pretty strong correlation between a problem being in BPP and being efficiently solvable in practice.
There are some exceptions of course (graph isomorphism was solved in practice when the best theoretical algorithms were still exponential) but in general once people find a n^100000 algorithm it soon turns into a n^3 algorithm with reasonable coefficients.
well nobody uses AKS right?
It's efficient, but somewhat slow..
The runtime looks very weird. The +2 can and should be dropped. This reduces my confidence that the bound is tight. Who knows how the model came up with that expression.
Is it mostly an artifact of the proof or does the algorithm actually need anything close to it?
Incredible stuff.
An ex colleague of mine who is a world class mathematician recently got an ERC with ambitious goals to advance his field.
Literally every optimistic goal proposed to be worked on during this multi-year window has been solved in this one post. His and his entire group's work has just been done for him! They are all depressed as hell right now.
They shouldn't be depressed. This all needs humans to go over and integrate into other works, and most importantly think about the next big questions.
Don't worry, next month's internal model will be able to posit all the next big questions that matter.
How?
It might as well happen, similar to how AlphaGo was superseded by AlphaZero, at some point a model might produce better math if it's trained through self-play where it poses its own problems, instead of looking for open problems in literature.
What's the objective function or RL environment for "interesting conjecture"? Not saying it can't be done - I no longer have any specific task that I'm confident AI won't be able to do - but I don't see how. It feels to me like something that would require a qualitatively new approach.
> but I don't see how
LLM's are trained on human knowledge and taste. They are actually pretty good at deciding if a conjecture would be found "interesting" by the mathematical community or not.
Note that I am saying LLM, and not chatbot or agent. But even a chatbot can often still reasonably rank a list of mathematical statements by vague properties like "interestingness".
How to RL this is a bit of an open question, but there are interesting conjectures of how to do it.
Can you expand on these interesting conjectures?
The fact that it's an open problem is the point I'm making. There's a very high degree of hubris right now, with people just assuming any open problem will be flattened by the AI steamroller soon. And sure, if that's what someone wants to believe that's up to them, but it's not a terribly interesting point of view to me. "What about X" "It'll be solved somehow", "What about Y" "It'll be solved somehow". Not exactly scintillating. If you know of any actual ideas on this I'd be interested to hear them.
Also any specifics on what you said about LLMs rating (preferably novel) mathematical claims for "interestingness" would be interesting.
I can only discuss published work, but take for instance this paper as one of the conjectured approaches: https://arxiv.org/abs/2603.20396
There is a general idea that beauty in mathematics is about being maximally compressing. Say I have a book with all formally correct logical statements. I could prove everything by truth table, or I can maximally compress my book with all proofs of all statements, and that will make my math beautiful. Because it forces you to reduce everything to a core of very general statements which are powerful compared to the length of the proof.
Math as some kind of condensed crystal from the sea of all possible logic.
There are other ideas of how to do it. The time has come now to just try a bunch and see which ones produce good results.
Interesting, I'll have a proper read of that. Sounds not unrelated to Kolmogorov complexity.
It may not be an interesting point of view, but it is based on the most fruitful philosophical position in history, plain-old empiricism.
Your statement that “I no longer have any specific task that I’m confident AI won’t be able to do” is founded on that.
That's not empiricism. It's the kind of extrapolation that predicts negative Germanies and 10 ton babies.
> It feels to me like something that would require a qualitatively new approach
This sentiment has been a recurring theme throughout the history of the field.
Sure, but the question stands. The ability to evaluate some measure of success seems pretty fundamental to how we train models and iterate with them on tasks like theorem proving. What is that measure for mathematical conjecture generation? How do we evaluate success, either on a particular task for iteration (like we do by eg. setting an agent to produce a lean proof of a specific result) or on a large enough set of training data to learn a set of rules (like we do when eg. using an RL environment to train a model to generate source code that passes automated validity/correctness checks).
So they can try to advance it even more
Levent Alpöge (Anthropic mathematician) comment on the significance:
> Sure, mathematical history features a lot of incredible developments, like the invention of proof, zero, or the computer, and on the great problems our progress has been over timelines measured in decades or centuries. Obviously this technology didn’t appear today, but blurring our eyes a bit to combine the past ten years, with today a measurement of those developments, there is nothing comparable.
I too would be optimistic if I was set for life
Crikey it’s a pretty charitable vibe given the whole Navier-Stokes thing, OpenAI trying to stiff him out of co-authorship. I guess any of that sentiment is outweighed by a sense of optimism for where this goes
Where does it go? Machines owned by 10 people robbing us of the joy of discovery and the fruits of our labor... for what?
They certainly arent going to give you that cure for cancer, if it were to ever come.
> Where does it go? Machines owned by 10 people robbing us of the joy of discovery and the fruits of our labor... for what?
For what? So some guys can be richer and more powerful. What could possibly be more important than themselves? You? Your future? We're nothing, and have been told to be excited and curious about our coming obsolescence and powerlessness.
Seriously, what's special about oncology that convinces you there will be no progress?
I think gp was saying they'd never give it to you, not that they'd never make progress.
A normal conspiratorial comment about the cure for cancer being kept secret would have been downvoted into oblivion, but sprinkle in some spooky AI stuff and suddenly it's fine
In the interest of steel-manning, I think it’s not about a cure being kept secret but rather non-elites being stripped of the leverage they would need to access it.
You were never stripped of anything
Labour value. Not in the past, but the argument is that it is being threatened.
That a peasant could be so insolent to imagine themselves deserving of the fruits
That is ridiculous. You cannot withhold something like an effective cure for cancer from broad adoption, and thinking you can is just conspiracy bullshit. Imagine an internal OpenAI model develops it tomorrow. Would everyone of the thousands of OpenAI scientists get in on the conspiracy to keep it secret, even though many of them probably know someone dying from cancer right now? Obviously not.
I'm sorry, do you understand how the medical industry works? It wouldn't be one cure, it's going to be dozens of treatments, each of which will be incredibly expensive.
Or else how do you explain Danyelza? Used to treat neuroblastoma, costs upwards of $1m per year. Do you have any proof that this will be different?
You're calling realistic people conspiracy theorists. Whose side are you on?
The medical industry makes medicines very expensive because they include the enormous costs of research. If AI makes research much cheaper, then medicines will get cheaper too.
I'm also against absolutely anyone who asks me "whose side I'm on" as part of an argument.
Just like when the medical industry made insulin cheap because it’s been around forever and costs nothing to manufacture.
Insulin is very cheap in countries with normal health insurance systems (everywhere except the US).
> I'm sorry, do you understand how the medical industry works?
Yes, I actually work in the medical industry. There is no hiding the cure for cancer.
> Or else how do you explain Danyelza? Used to treat neuroblastoma, costs upwards of $1m per year. Do you have any proof that this will be different?
Oh, you are American. Let me tell you a secret: the problems of your healthcare insurance system are not a worldwide phenomenon, nor an immutable fact of this universe. Perhaps the cure for cancer, if expensive, will not be easily available to the poorest Americans, at least initially (the cost will come down sooner or later). But that is a very different claim from "they'd never give it to you".
> Yes, I actually work in the medical industry. There is no hiding the cure for cancer.
There's no need to hide anything, it's just pay-walled (and it's not an hypothesis, most human beings on this planets cannot afford the SotA treatments for their cancer today).
burn
What makes you think we'll be able to afford them? You won't be making any money anymore, hope you've saved up!
> Machines owned by 10 people robbing us of the joy of discovery and the fruits of our labor... for what?
This happened to software engineers already. Mathematics isn't special.
Mathematics has it slightly worse, since math is a "closed system" that does not require running experiments, talking to customers, etc.
That said the software folks can teach the math folks a thing or two when it comes to dealing with grief I am sure.
You don’t want to see Elon get to 100 trillion and start a sex slave colony on Mars? What are you crazy?
> They certainly arent going to give you that cure for cancer, if it were to ever come.
Conspiracy bullshit. You cannot keep something like an effective cure for cancer under wraps. There is no plausible logic how that would not leak sooner or later.
The information needed to build a basic nuclear explosive device is rather easily available. Now go build one.
I am talking about an *effective* cure developed by AI. If curing cancer with it is as complex as building a nuclear weapon, then the cure might as well not exist.
> If curing cancer with it is as complex as building a nuclear weapon, then the cure might as well not exist.
They exist, I assure you!
https://en.wikipedia.org/wiki/List_of_states_with_nuclear_we...
Knowing is the easy part. An effective cure for cancer will probably be a procedure where you get your tumors sequenced, an AI model considers the unique genetic context and creates a specific treatment. Think CAR-T cell therapy or mRNA vaccines. Then this needs to be manufactured.
All of this means you will need to be rich or have your country invest lots of money into health care systems. In a world where humans don't provide economic value anymore, why would that be?
The cost of such therapy will inevitably go down over time.
Well rich folks using such services will eventually push the price down, however complex and ridiculous the actual process will be. Its not like they are immune of all these ailments, not yet at least.
Trying to slow down this for the joy of discovery is a deeply anti-intellectual position. I think that position is similar to when everyday people get mad about the minor spending on the NSF, picking apart people who study beetles on Fox News with no context etc.
There are real safety concerns with AI that can be made very convincingly though.
I think I would feel a lot better if it wasn't a $15 million cannon being fired from a silicon tower inside of the labs across vast swathes of fertile ground that would otherwise be used to train budding mathematicians. For instance, if it was the budding mathematicians themselves who were making these discoveries using their own tools.
I feel something about human nature makes us treat joy of discovery, status, etc. as a source of energy and motivation. I hope we'll find other ways to keep some "strategic intellectual reserve" of mathematicians alive.
What if we just discovered an alien artifact with the next 200 years of math? I can see arguments to throw it away for safety, who knows if the aliens are getting us to nuke ourselves or whatever, but throw it away so a few hundred/thousand of the most elite thinking humans can have the joy of credit for discovering each thing?
If the argument is that we have discovered 200 years of math in 1 year that we might need to throw out for fear that blindly applying its incomprehensible results will lead us to ruin, then I would say: why not settle for 199 years of math in 1 year which can be verified by our human intellectual reserve that we will train on the remaining 1 year?
I never said they are trying. They already have.
>OpenAI trying to stiff him out of co-authorship
Did that actually happen? The emails that were originally released had OpenAI refusing to list him as a coauthor on OpenAI's paper but they suggested he should release what he already had done ahead of OpenAI's release. There was certainly nothing to suggest he should be robbed of credit for his own work.
Has anything new come to light since, or is this just another game of Chinese whispers?
And if he did release what he had, I imagine OpenAI would probably have cited him in their paper.
They refused to put him as a coauthor; which is… standard practice, because he wasn’t a coauthor.
I don't see a sense of optimism in this quote.
He's comparing it to "lot of incredible developments". Clearly he thinks this is one too!
Yes, it is incredible. But that doesn't mean it is cause for optimism.
That's your lack of optimism, not his.
No. I didn't say I wasn't optimistic, nor that I was optimistic, nor even that Alpöge was or wasn't optimistic. I just said that his statement about the current developments being incredible doesn't imply that he is expressing optimism.
Wow. Fantastic quote. If you have the source, would you please share a link? Google did not bring up much.
https://x.com/__alpoge__/status/2107620576059679129?s=46
Thank you!
“I think at the heart of this issue is that humans have two competing natures: a tendency to compete and a capacity to appreciate beauty,” said Kai Shaikh, a graduate student in mathematics at the University of Toronto. “To me this seems to be a case of the former attempting to strangle the latter.” [1]
Beauty can be appreciated even when it is vast, even when it is beyond one's comprehension. I don't think this release should be primarily viewed as an outcome of competition. Instead it is revealing truths about the universe that were always there and always beautiful, even if we hadn't seen them yet. I believe there are infinitely more such beautiful truths currently hidden and waiting for us to discover.
[1] https://www.nytimes.com/2026/10/06/science/openai-math-probl...
It's unfortunate how toxic media reporting on AI has become. Everyone has abandoned even the pretence of objectivity. I know NYT is uniquely biased in this regard, but there was no need to add "Further Roiling Field" in the headline. Like, you published this minutes after OpenAI's announcement and claim to capture how the entire field of mathematics feels about the advancement? Before anyone has had a chance to even read let alone digest it?
"Roiling Field" seem accurate though? The discussion I've seen on here from mathematicians seems fairly roiled.
It's the NYT. What else could you possibly expect?
I think people knew it was coming. Someone rushed out a preprint a couple of weeks ago with partial results on the Unique Games Conjecture because they heard AI had solved it completely.
Objectivity? Why would you want favorable reporting for the machines they're building to replace you, and, by their own admission, potentially kill you?
The only bias here is that we're still covering these things like business ventures and not criminals.
That's a fantastic quote. I definitely personally feel this tension.
Not that I could ever "compete" on the frontier of math in the first place. But our nature to compete derives from our need to survive against other capable forces. And results like these make me feel very nervous about humans' capability to remain the dominant force in the universe.
The bitter lesson has a bitter aftertaste alas
I'm sorry, but what?
If you appreciate beauty and don't care about competing then these releases are purely good. Because you are not competing, so you aren't hurt by speed. And you are appreciating so you can appreciate more stuff.
sorry if I am misunderstanding (probably I am)
a tendency to compete and a capacity to appreciate beauty,
IDK, I think you should add tendency to cooperate, a capacity to love and perhaps some other things there.
But with things unfolding quickly and unpredictably, I think everyone's view is getting a bit foreshortened here.
> Instead it is revealing truths about the universe
Mathematical proofs aren't revealing truths about the universe. Mathematical proofs are independent of what the universe is like. Any proof would be the same in any possible universe.
You could argue mathematics is part of the universe. Or maybe vise versa.
So what word do we use instead? It is revealing truths about "reality"?
> Mathematical proofs aren't revealing truths about the universe.
That's exactly what they do, apply logic formally and systematically to discover truths.
Sure, there may be a universe where 2 + 2 = 5, but then that universe would have its own mathematics that can prove that to be true. And there will be a way to bridge that alien math to our own, again by logic and proofs, until we have a larger sense of truths not only in our universe but all possible universes. Proofs are part of the constant process of revealing deeper truths to the best of our understanding.
There is no logically possible universe where 2 + 2 = 5. It's definitionally false. OP's point is probably that math is a set of rules we invented, not something true about the universe.
Unique Games Conjecture [0] is a seminal conjecture in Complexity Theory, and is an underlying assumption for many, many inapproximability results. A valid proof is a big deal!
[0] https://en.wikipedia.org/wiki/Unique_games_conjecture [1] https://github.com/openai/math/blob/main/preprints/The-Uniqu...
I also don't think there was general consensus on which way this would resolve prior to this (or is that a little out dated?) unlike some of the other major problem resolutions. I heard rumors that there would be a big result in TCS and speculation it would be UGC that or P neq PSPACE but I'm still a bit shocked.
I'm really glad that OpenAI is formalizing these things, because I'm not convinced that their current internal frontier models are particularly good at writing down their thoughts in English. From the (probably awesome) Unique Games Conjecture/Theorem paper, the first two sentences of section 1.1 start to define the problem:
> A Unique Games instance has a finite vertex set, a finite alphabet K, and a nonempty list of oriented constraints e = (u_e,v_e,π_e), where π_e is a permutation of K. A labeling a satisfies e when a(v_e) = π_e(a(u_e)).
I'm sorry, what? I admit it's been quite a few years since I've thought about the Unique Games Conjecture, and I never dug that deeply, but this part is very, very elementary graph theory and notation. So let's unpack it.
1. e is maybe a name of a list.
2. The elements of that list are tuples, where each tuple is (a vertex, a vertex, a permutation). So e indexes into the list and u_e is the source vertex for the e-th constraint in the list called e. Thanks.
3. a is a labeling. I'm fairly confident that, by "a labeling", they mean that e is a function from vertices to colors, where the colors are the elements of k.
4. That vertex coloring a satisfies the list e, when, for, um, an index e into e, a(v_e) = π_e(a(u_e)). But this isn't for all e, it's for some e, and the goal is to count them.
So maybe e isn't a list? Maybe e is a constraint that is represented as a tuple, so e = (u_e,v_e,π_e) and u, v, and π aren't sequences at all but are, in fact, the trivial unpacking functions that unpack the pieces of the tuple.
Reading this stuff is pointlessly painful, and it's extremely easy to make mistakes when being sloppy like this.
If this were my paper, or if I were trying to train a model to write math, I'd want something like:
A Unique Games instance has a finite vertex set V, a finite edge set E = (V × V), a finite alphabet K of possible vertex colors, and a nonempty list of oriented constraints. Let Π be the set of permutations of V. Each constraint e is a tuple in E × E × Π, where we write u_e ∈ E for the first element, v_e ∈ E for the second element and π_e ∈ Π for the third.
A vertex coloring a : V → K satisfies e when a(v_e) = π_e(a(u_e)).
[…] a finite edge set E = (V × V) […]
E ⊆ V × V
Oops, that’s what I meant.
In this particular case, though, I think my typoed version may be equivalent. An edge with no constraints has the same effect as no edge at all.
I definitely messed up the constraint definition, though: u_e and v_e refer to vertices, not edges. That’s what I get for writing it with minimal proofreading.
Yeah, that's one of the big things of TCS. I think I see that as bigger than that Millenium Prize problem.
TCS being theoretical computer science? I have not seen that acronym before.
Yes, TCS is theoretical computer science, I commonly use that acronym too.
This was the biggest highlight for me as well. Astounding...
This is significant progress and released without all the drama. Some very important progress in Reinmann, Hodge and unique games theorem. Point the repo to your agent and ask for the significance! In a way this is probably 50-100 years of math progress by humans
Not to pick on you specifically, but as someone who spends a lot of time unproductively reading AI math discourse it's truly shocking how incapable all the supposed math enthusiasts are of spelling Riemann.
I just did a quick search on this and apparently the misspellings are German surnames as well:
https://en.wikipedia.org/wiki/Reimann
https://en.wikipedia.org/wiki/Reinmann
I just don't understand how it happens. If they had ever taken an intro to real analysis class they would learn to spell his name. If they were just parroting what an AI told them... shouldn't they still just say his name? An individual could just be dyslexic or mistaken but it seems to be a substantial volume. I guess they just don't care enough about it to commit the correct name to memory, only remembering the "pattern" of the name and filling in the spelling via guesswork?
My last name is Hebert.
There’s about a 1 in 10 chance when someone spells it (or says it) they say Herbert.
Even in situations where they just read it or I just said it.
I’ve had Herbert soccer trophies, health insurance cards, etc.
The mind fills in a lot of blanks and doesnt always get them right.
My last name is Kvick - the Swedish word for quick. I live in Sweden. I get a lot of people thinking my name is spelled Kvik, Quick, Kwick, ... My father once got a mail addressed to Mr. Kvack (Swedish for quack, like a duck).
I sympathize. -- Not Barber
Maybe they skipped straight to Lebeg integrals.
Thanks for the laugh!
very good lol
Autocorrect might be doing it
iPhone would be my guess
People just don’t spell that seriously buddy, especially when it is so immaterial to the point.
My point is it is crazy to make public claims about how important or not important a mathematical result is when you can't spell Riemann. Yes, it technically doesn't matter, but it betrays a damning lack of familiarity with introductory mathematics.
Then ignore them? I dont get these types of mysteries. It's easy enough to find math majors these days to ask them their opinions on things. There're so many of them. You most likely know some from your highschool. They'd probably say the same things though, or even freak out harder.
While it takes a good math person to make breakthroughs it's much easier to find someone who has a feel of whats important/hard and not. Even a mediocre math major/master is far more authoritative than an expert at adjacent fields (CS,physics). Or to listen to webdevs 'ai skeptics' or whatever on the internet.
You’re (as Claude would say) absolutely right, and I suspect the original commenter has no idea what they’re talking about.
Yes sure but they are different surnames and pronounced differently.
I didn't mean to oppose OP's point, I just found it interesting as a non-German speaker!
Human brains seem to have somewhat similar failure modes to LLMs and how many 'r's in strawberry.
They're mathematicians, not linguists :)
The mathematicians don't make this mistake, I assure you. I doubt there is a math professor on planet earth who would spell Riemann as Reinmann. In fact, I imagine no one who has ever heard the name pronounced would do so.
One of the best mathematicians I've had the pleasure of learning under added 6 to 7 and got 15 during a lecture... I assure you they're capable of misspelling last names too.
Mathematicians aren't exactly known for being well rounded.
Oh god, that reminded me of my linear algebra teacher in college. Doing matrix multiplication by hand and ending up with 9x6 = 45, and then having to direct him to the cell with the wrong number. Loved the topic, hated the class.
Pure mathematicians getting arithmetic wrong is a bit of a meme it’s so common. I don’t think it’s the same as a misspelling of a popular mathematician, not that either are indicative of much tbh.
google "Grothendieck prime"
Still, enough misspell Lebesgue as Lebesque.
Uh oh, no true scottsman....
It's just a name you write so many times as a math undergraduate or first year graduate student due to the number of load bearing results and objects named for him. The name even appears in lower division coursework
Yes, you can be an amateur mathematician who manages to avoid such classes and may not know those results or objects, but if you haven't read and written down the name enough to avoid habitually misspelling it, you are outing yourself as a meat proxy unless you are dyslexic.
I regularly heard lots of names during my 3 year math undergrad degree. I couldn't remember how to spell many of them then and can spell fewer now. I imagine most people on the course were similar, and me nor no one I knew were dyslexic.
Mostly we wrote initials in our notes and the exams didn't ask about them. The lecturers only wrote initials on the chalkboard after maybe writing the name once when introducing it the first time. We weren't there to do mathematical history and remember names or something.
Google existed then and now and we could look them up if needed.
If they’re mathematicians they have written this name down about a 100 different times throughout a standard Real Analysis course. Riemann was foundational in that field.
“How many Ns in Riemann?”
Let's figure it out! First, draw little boxes over all the letters. Then add up their areas. Then make the boxes smaller and repeat, until you have forgotten how to count.
> Point the repo to your agent and ask for the significance!
Wow, this comment really shows how low this community fell.
Math is the tool humans use to compress knowledge. So until we can comprehend it there really isn't much progress. Math theorems are tautologies, the truth of which are not dependent on proofs and proofs are erasable, at least classically. But the AI progress is exciting and AI proofs are a gold mine for humans (at least non domain experts) to explore.
I think it would be helpful to people who want to understand what a formalized proof is to read Thomas Hales on this: https://www.math.stonybrook.edu/~bishop/classes/math536.S24/...
He spent years formalizing his sphere packing theorem because the proof (human produced) was already beyond the ability of peer reviews. Now his formalization effort likely can be easily reproduced by a model. However one should read his experience about what a formal proof is: often the problem is the statement not the proof. The example he gave is the Jordan curve theorem. It's actually quite challenging to formalize the concept of a planar curve (there are space filling curves). So it is not necessary that someone can look at a formal statement and say aha it is about a planar curve, unlike FLT where there is not much problem in recognizing what the statement is about.
Fun challenge: find the formal definition of simple_closed_curve in the essay and tell me if you believe you learned anything about a planar curve.
BTW the essay is eminently readable for anyone interested in math. Hales wrote it in favor of formalized math and to educate his peers and students about it.
What makes you think no one can comprehend this? It has been less than an hour since it dropped and there is already a ton of online chatter from people explaining the results, pointing out their favorites and more. Some of it is happening on this very thread.
I didn't say that. I am responding to "In a way this is probably 50-100 years of math progress by humans." I am actually very excited about AI proof and I am working overtime in my own way to try to comprehend as much as I can.
Another way to state this: math theorems are like programs without side effects; it is immaterial whether a program without side effects is ever run. We study math for the side effects: it changes how we organize our thoughts.
Math is far more than that. If you can solve prime factorization for example, suddenly you can listen and interfere with almost every private conversation on the internet.
We are not far away from the moment where these models will be restricted, and sharing the results will be done more carefully.
This doesn't contradict what I said. But I do appreciate the fact AI can produce side effects not just humans. I made it sound like only human knowledges matter. That's too narrow.
> until we can comprehend it there really isn't much progress
Who is "we" here exactly?
Whoever wants to study the result.
> Whoever wants to study the result.
Right. Before all the AI disruption, pure Math traditionally welcomed anyone who wanted to study its esoteric proofs, right? I remember all the excitement of the average Math enthusiast casually reading Wiles' proof over coffee.
Bottom line is, the relevant people can still understand the generated proofs. The disorienting part is they are a little slower than they'd like, but they'll get there.
AI is very helpful with understanding AI proofs. Agent swarms produce messy proofs overall but locally they are excellent and can teach anyone who wants to study them. No one controls math (in a material way funders do control an aspect of practicing math). Still, theorems are already true before we prove them. The difference a proof makes is whether it convinces the reader.
> Math theorems are tautologies
Proven math theorems are tautologies.
FLT was no less a tautology before it was proved. We just weren't sure about it. Proofs only change us, not math.
True.
It is progress on another / the next evolutionary later: A AI/AGI/ASI system.
Which either replaces us in the long term, augments us or makes us better (gentherapy).
>So until we can comprehend it there really isn't much progress.
Not really? We are at a point if an AI today can solve it, it can be stepping stone of understanding something deeper to tomorrows AI and it continues. Sort of like our limitations doesn't matter. Obviously there are many scenarios in this recursive loop but saying it isn't much progress is not how I view this as
But who will ask the questions or direct further research when humans no longer understand the state of mathematics? Llms lack the drive for homeostasis combined with the evolutionary drive for survival and reproduction thus to direct themselves. They could very easily spend an eternity down a rabbit hole when the warp drive equation was fairly close on another branch.
If it changes how we think then yes it has an effect.
Academics have been treating it that way because they had no other choice, and its been a waste of everyone’s time and often times taxpayer resources
Look at that, taxpayer funding was cut and a private sector solution came in just the nick of time, far accelerating the holding patterns we’ve been in for decades
Humanity doesn’t need all iterations towards the blueprints, the blueprint is good enough, we all stand on the shoulders of giants
If you actually looked into how agents proved FLT you would be even more amazed by the fellow human beings who are able to keep all this in their heads! I for one can only begin to grasp the scope with AI and scripts.
There was A LOT of drama about this release.
> This is significant progress and released without all the drama.
I feel gaslighted.
As an AI "doomer" can I ask the non-doomer people here how you interpret the significance of results like these, and what kind of progress you expect to see in the next 1-5 years?
Like do you see the technology plateauing at the current level, do you expect progress will continue but only in mathematics, I'm interested to know why others are not concerned?
Not a doomer but I try not to be a denier, either. These are hugely impressive results. I do not see the technology plateauing at the current level (though I am dubious about an infinite exponential growth).
Before I cope, I’ll note that there are plenty of “doom” scenarios that do not require any improvement in capabilities from what we had before this latest unreleased model. We’re at the point where a determined bad actor with enough compute could compromise critical infrastructure in a way that results in casualties, where this actor would not have been capable of such without LLMs. This may not sound like Skynet, but I don’t see why it makes a difference if I’m one of the casualties.
With that in mind, here is the cope: first, mathematics is an inherently verifiable domain. An LLM can use tools to determine with absolute certainty whether it is correct, and an independent third-party could review and confirm. All of this can be done without any interaction with the physical world or with other minds.
Second, OpenAI is able to marshal compute at a scale that an individual mathematician can only dream of. It’s possible that these problems were lower-hanging fruit (in relative terms), such that they could be resolved simply by throwing a ton of compute at the problem guided by an intelligence that is not itself remarkable in comparison to a human.
Third, none of these problems are solved in a vacuum - the reason OpenAI chose these problems is that they are widely discussed and many people are working on them. It’s possible that someone else was close, and OpenAI only contributed the finishing touches. (This wouldn’t need to be plagiarism, to be clear - people publish their work!)
> It’s possible that these problems were lower-hanging fruit (in relative terms), such that they could be resolved simply by throwing a ton of compute at the problem guided by an intelligence that is not itself remarkable in comparison to a human.
See:
> The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking. (TFA)
Sure and if I make a half court shot after an hour of trying, the result only took 1 second.
Exactly this. If you take the entire start to finish 'agent hours' (measured comparably to man hours) they took to find all discoveries, including the go-nowhere trails that were discarded, and then divide by 90 (or whatever the exact number of results found was) it's almost certainly going to be many orders of magnitude more than 3.
They provided a "snippet" of a prompt here [1] which is not only a beast, but also seems reasonably likely to have been LLM generated. So they're using LLMs to parse a vast body of mathematical work, probably including what people themselves are 'privately' working on with GPT, and then prompting other LLMs to work on such.
[1] - https://github.com/openai/math/blob/main/reasoning_traces/re...
That tells me very little. What was the cost to OpenAI in dollars? What differentiates the high-cost problems from the low-cost problems? And that’s before you consider that OpenAI has strong incentives to downplay its costs while emphasizing its results. A one-liner in a write-up doesn’t change the fact that they have access to massive resources.
In a couple short years we've moved from "AIs can't do anything useful" to "they're lying about the actual cost of the innovative breakthroughs!".
I know that the former and the latter may be discrete subsets of the anti-AI crowd, but come on.
It’s very typical in human math that explaining the final result after years of searching looks very simple too.
Fourth, these hundreds of solved problems are the result of OpenAI attempting tens of thousands of problems and failing. When you hear claims that the average result took about 3 hours of model time, I simply do not believe it. If you account for all the time spend properly, it's probably orders of magnitude more.
I think the announcement says they report the amount of problems attempted somewhere.
Edit: "Over the course of the evaluation, the model was posed approximately 4,000 problems. Aggregating the output into result families and manuscripts and requiring an appropriate level of significance led to the catalog outlined above."
I would like to know how much of the progress comes from effectively combining existing research programs plus massive persistence, and how much is AlphaZero-esque RLVR completely independent of training data. Since I cannot get anybody to care about this question (even though I think it's vital for guessing what the future trajectory will look like -- are we going to complete existing research programs or start new ones?), I live in ignorance and wait for the day when the answer becomes clear.
In looking at this over the past hour, I haven't seen clear evidence one way or the other. Some of the stuff is highly unexpected (like the multiplication algorithm), but counterexample-y, and about the rest the professional mathematicians online seem to have a consensus that it's not "breaking through fundamental obstacles". I suspect neither of us is competent to judge that.
Doesn't matter at this point i would say.
Alone the massive usage of us every day produces a massive amount of signals.
I build something and claude does something stupid? "hey thats not what i meant! Do this instead!" "Okay" <<< This is a signal.
The mathematician being unhappy about something from claude? Another signal.
This alone gives you enough progress i would argue. But additional its clear that certain tasks are worth to pay experts for for teaching one central AI once instead of every single human who needs to do the task.
IF RL is also working well, we are just faster f*ed than otherwise.
I expect AI will continue to be useful on the vanguard of fields like mathematics because it has the perfect conditions for it to shine. There are a lot of discussions and leads to start from and the AI can check its own work and iterate. It can fail hundreds or thousands of times in a day and continue to work with the same tenacity.
Superhuman tenacity is not enough on its own to pose an existential threat. If it showed the same capacity for judgment, inventiveness, and decision making in the messy problem space of the physical world, I would be more alarmed. There have been experiments where an AI is given control of managing something like a vending machine and it always ends up a mess. AI has come a long way, but certain problems seem as difficult as ever.
When AI becomes more capable of navigating practical problems without human intervention, I will start to be concerned. Enslaving humanity will involve taking a lot of calculated risks that tenacity alone cannot solve.
> When AI becomes more capable of navigating practical problems without human intervention, I will start to be concerned.
Given the past rate of progress, why not start being concerned now? It's a bit like the economist saying that the optimal number of flights to miss is not zero. If you keep landing short on your estimations for how far this technology will go, next time you should err on the other side.
And regarding Vending-Bench 2 (https://andonlabs.com/evals/vending-bench-2) my understanding is that models do pretty well on it now.
What's the point of being concerned?
We're not going to stop it because of the money involved and once we're dead, it won't matter anyway, might as well just enjoy life until you're done.
We're going to get AI'd to the max, whether or not we like it or not, might as well just go with it.
Your attitude is extremely sad. Do you think we would be where we are today if oppressed people's throughout history just gave up as easily? We have rights because people fight for them. We collectively have the power to decide what kind of future we want to live in.
My point was that it’s It worth worrying about. Not that we should fight / have rights
> once we're dead, it won't matter anyway
It sounds like something is worth worrying about if you foresee us being dead, presumably prematurely.
We can't stop it. But you can use your voice to buy time and resources for alignment and safety research. A few additional months may make a world of difference.
All the billionaires making AI already say there's a 10-50% chance it's going to kill everyone.
I don't think your LessWrong post is going to save us.
If anything is a doomer attitude, this is.
A doomer is someone who believes doom is inevitable or highly likely, I'm not saying that, I'm just saying "being concerned" will probably get you nothing in return and this tech is getting developed no matter what.
The only way it will stop is if the wealthy / powerful people feel threatened by it, properly threatened.
I don't think you're paying much attention to how rapidly things like bipedal robots, and just robots in general are becoming far more capable very quickly.
The same GPU compute for LLMs runs robotic training models. Now in a few hours you can train a robot model that would have taken months 5 years ago. This model gets dumped into an actual physical robot with sensors all over and the suitability of the model is measured on robot tasks and the error in real world actions is fed back into the robot world model for further training.
> There have been experiments where an AI is given control of managing something like a vending machine
You sure you're not talking about experiments ran a couple of years ago? The more modern ones are getting wild.
https://techcrunch.com/2026/07/29/claude-opus-5-became-downr...
I don't quite understand the leap you're making between stochastic AI models for robotics (maybe with an LLM making api calls to it) and embodied AI / the rate of progress towards a singularity. Because they're both trained on a GPU? Up until 2022 GPUs were for video games and mining crypto, and neither of those produce a synergy that accelerates progress towards general intelligence either.
can you share more about the robotics advances, what you are aware of? sounds very exciting.
There are already machine-controlled high throughput experimental machines for wetwork. AI will definitely do a better job than your average biochemist at planning, executing and analyzing these experiments just by virtue of the amount of thought it can put in to experiment selection.
Depends on your definition of doom.
If doom is ending up with grey goo / paperclip maximizers, or SkyNet, then I don't think doing mathematics is evidence of that direction. Partly because LLMs are quite apparently dumb in many ways, and for math specifically, they need a formal verifier (Lean) which "gamifies" math.
If you mean bioweapons or cyberwarfare, there's nonzero risk, but not orders of magnitude worse than other global risks. Climate change, nuclear weapons, monoculture food, etc.
I'm far more concerned about overall trends in AI development and usage. It's accelerating wealth inequality, social isolation, attention capture, surveillance states. If we end up in the Matrix except the admins are humans and the simulation is hyperoptimized TikTok, is that AI-driven doom, or is it just an inevitable outcome of modern tech?
That’s the thing, there are innumerable ways it can go wrong and only one way it can go right (if it doesn’t lead down the aforementioned innumerable paths)
Where some see intelligence, others see token prediction. It's a very good question how token prediction could achieve this, but I think there's a simple explanation. No human can hold more than a negligible percent of all knowledge in his mind at once. LLMs have no such limits and so can reliably connect 'obvious' dots that we miss simply for lack of storage capability.
Well isn't that just semantics? Surely connecting dots in a novel and meaningful way is intelligence regardless of how it's achieved. The thing is that humans didn't get to where we are by connecting obvious dots. Go back to before humans had invented language and when bleeding edge tech was literally that - 'poke him with the pointy end.' Train an LLM on that corpus of knowledge. Even given infinite processing power and infinite time - it's not going to discover the secrets of the atom, put a man on the Moon, or do much of anything besides remix what we'd already done at the time.
I expect there's still much LLMs can achieve simply because of this initial problem. But I expect that they will ultimately start to plateau once these dots have been mostly matched and we reach a point where 'creation' again becomes the missing link. Though even there LLMs will play a major role as tools. For instance Einstein had to spend a significant amount of time in 'retrieval' rather than 'creation' research to develop the field equations for general relativity. If he had access to LLMs trained on all knowledge of the day, he could likely have achieved his goal much more quickly.
I mean, even if you buy the idea that all LLMs are really doing under the hood is insanely good interpolation, the results produced by that interpolation are still novel and still get incorporated into the knowledge corpus of the next training run. I guess the implicit question there becomes whether that expansion allows the knowledge corpus to continuously grow or whether it eventually settles into a steady state.
I think that people are just really bad at math and coding. Knowledge workers have been taking pride in doing stuff which the average person does not 'get', but we only understand a little bit. That leaves a lot of room for people to get better, or for other jobs (farming, sandwich cafés, mystery novels) to have been already peaked by human ability and less useful to bring in an AI.
'Doom' to me means that any career crashes, we are controlled, everything is hacked, society stops functioning. Yet every part of my day today (except for coding) was done entirely by people.
Finally I think it's easy to make a simple model that everyone has a simple balance sheet, and that people are more expensive so they will all get cut. But the same argument could be made for all US jobs being outsourced, and all in-person engineers, lawyers, and doctors to be rubber stamps for overseas work.
Try using AI for your work, whatever you do. You will quickly understand the limitations.
Unless you think that AI will quickly hit a wall (which seems odd considering that only a few years ago the best models had trouble doing basic math or counting the number of Rs in "strawberry"), I don't see how that's reassuring. The models will only get cheaper and more capable over time. It seems quite likely that at some point (probably before I hit retirement age) they'll be able to fully replace me at my job.
Is there any specific cognitive task that you are willing to bet that AIs won't be able to accomplish in the next 5 years? Because if not, I'm not sure we're disagreeing about predictions.
Don't frontier models still have trouble counting letters? Or am I out of date? Either way, it doesn't seem to be the same amazing rate of progress we're seeing in other areas.
It's easy to look at a fire burning through a forest and extrapolate that rate of progress across the whole world. But fire doesn't burn everything equally fast.
What other cognitive tasks will be a struggle to make progress on? I suspect there will be some, though which ones they are is anyone's guess.
Your information is out of date by several years. The letter counting issue was due to how LLMs split text input into tokens (usually using BPE). Since 2024, frontier LLMs have used chain-of-thought reasoning to spell out the letters and count them.
> The letter counting issue was due to how LLMs split text input into tokens
No - this is provably not the issue.
Take any model that fails to correctly count the letters in a word, and ask it instead to spell the word (even a made up word), and it will be successful - they have no problem predicting the letter sequence from the token sequence (and would be shocking if they did - this is what they are built for: seq -> seq prediction).
The reason LLMs can fail at the letter counting task (depending on model training, prompting) is because of the counting part, not because of any difficulty correctly mapping the input token sequence to the letter sequence.
The letter counting issue is due to tokenization. And most models still get this wrong often enough, even with reasoning. Probably less so on strawberry given how prevalent it is, and less so than without reasoning, but this not a historical issue. It’s becoming less of one though.
My apologies, I got my info from an LLM. I guess they still have a ways to go in understanding current events.
I just asked Opus 5.5 if any AI driven advances in mathematics have been announced in the last day or so and it gave me a summary of this OpenAI announcement. https://claude.ai/share/319b437a-c1f2-4119-8dc5-45d36545fed9
Which LLM specifically? If it’s cloud based, you should be able to share the chat, right?
But seriously, I am still waiting for someone to wager that AI won’t be able to do a specific cognitive task in the next 5 years. This fact should be evidence enough that we have no idea how far AI capabilities will continue to advance.
I asked both gemini and chatgpt "do frontier models still have trouble counting letters?" and the first word of both responses was yes.
The fact that no one is taking you up on that bet I don't find to be particularly persuasive. I suspect there will be plenty of cognitive tasks LLMs struggle with in 5 years, maybe even 20. But I wouldn't hazard to guess which, I don't think anyone is capable of that level of foresight.
This is so strange I tried it on Gemini Flash: "Yes, but significantly less than before." When you read beyond the first word it explains where LLMs might fail and why.
Ok, share links to the conversations with both models. I asked ChatGPT Astra 6 medium effort and it said, "Much less than they used to." and provided stats showing how accurate they are.[1]
1. https://chatgpt.com/share/6ac5e4cc-02f0-83e8-8f05-99a7ea2bf9...
https://share.google/aimode/iKkrZtVYo4DSielWs
https://chatgpt.com/share/6ac63b6a-481c-83e9-a8fa-a13ce7402d...
I used whatever the default free model and thinking time was. If progress was really as fast and continually cheaper as some worry it is, wouldn't we expect free models by now to know (or even perform) what frontier models were capable of as much as 2 year ago?
This deep in the "comparing logs" tangent we risk missing the point. It's not what exactly frontier models are capable of at this particular point in time. But that there's entire categories of problems that seem easy to us which LLMs really struggle with. We've stumbled on several just a few replies into casual conversation. (Can they count? Can they know if they can count? Can they reproduce results? How quickly do new capabilities filter into free models? And that's just what's come up naturally, if we wanted to pick adversarial examples there's more to choose from.)
So while there's a number of difficult problems that are easy for LLMs (like bulk generating lean proofs), there are plenty of things where progress is not so impressive.
If LLMs can struggle so much with such easy problems, what hard problems have we yet to discover that they'll struggle with? The fact that no one knows, 5 years in advance, what those problems will be does not mean the chance of them is zero.
So far progress on the things LLMs are good at is fast and easy. It's like fire in a room full of oxygen. But once the low hanging fruit is gone, and the oxygen is out of the room. How fast will the fire burn through steel walls?
In my opinion it's a mistake to look at only rate of progress on one type of problem (whether it be what LLMs are good at OR what they're bad at) and assume progress on all tasks will progress at that rate indefinitely. Isn't there a saying about exponential curves, in nature, all being sigmoids eventually?
I guess we'll just have to see. I wish you good luck with your wagers.
You are extrapolating from the mistakes made by free versions of smaller models to claim that frontier models struggle with easy problems. This is an obvious mistake in reasoning because as you can see from my shared Astra conversation, frontier models don't have the same limitation. (They can count letters and they know they can count letters.)
Many people in this thread have made claims about limitations of frontier models, but I'm the only one who has shared a conversation with one. Everyone else is either sharing conversations of smaller models making mistakes, or they're making claims about frontier models but not linking to examples of them falling over. If frontier models were so easily fooled, you'd think someone would link to a conversation showing that.
Why look at the rate of improvement of free models when you can look at token pricing? Back in 2023, GPT-3.5 cost around $20 per million tokens. Astra costs half that.
The worry is not that smaller free models will replace people's jobs. The worry is that future models will. We are talking about the capabilities of frontier models because those put a lower bound on the capabilities of future models. Extrapolating from smaller models is a waste of time, as you can interact with the frontier model to figure out its capabilities and limitations.
Also the timestamps on the shared conversations show that you asked Gemini 10 hours after ChatGPT, which means you asked it after your comment claiming you asked both models.
I didn't save the original query so I asked again this morning and ended up getting the same response -- points for consistency, though it might have been more reassuring with the correct answer.
I don't think my point is really landing so I'll try once more and then give up.
Let's say frontier models today have no problem counting letters, I never really disputed that but only asked about it. It seems based on the other replies in this thread, it's a bit of a "who you ask" kind of thing, but let's grant that they have no issues with it now.
The first version of chatgpt was released 4 years ago next month. Which is not quite 5 years but close. In that time we've just barely managed to get spelling down. If we extrapolate that rate of progress forward 5 more years, are you still afraid for your job?
I think we're all more likely to lose our jobs from a downturn in the economy caused by the capex/debt bubble bursting than being made redundant by AI. (And the continual pricing reductions only seem to make this result more likely.) Hopefully neither happens and in 5 years we'll all still be gainfully employed.
Solve the Collatz conjecture in the next five years? If humans publish significant advances during that time, and A.I. copies it, then yes. Otherwise, I'd definitely bet money it won't happen. I'll give you 10,000 brownie points if I'm wrong.
They still have issues with problems like this actually, and I use all the frontier models from all the major labs, so it's not solved.
I'd love to see some examples of frontier models getting letter counting wrong. Can you share some?
> out of date by several years
This is delirious exaggeration. The problem has not even been widely recognized for several years. Fable reported "two rs in raspberry" to me as recently as August. There is some randomness, it's hard to predict which words will trip up the machine, and I haven't been able to do it at all since August. But it was absolutely happening until very recently, and probably still is.
But doesn't that just amount to labs intervening to teach the models to use a particular strategy to mask this one very obvious marker of the difference between their intelligence and biological intelligence? (And similar surface issues like using tool calls / reasoning for arithmetic, even though humans writing on the internet don't typically break show their work for multiplying two numbers)
The deeper architectural difference is still there, which manifests whenever you try to get the models to apply known techniques to modalities and problems outside their training data.
Chain-of-thought reasoning was added for general purposes, not to fix letter counting specifically. It just happens to solve that problem in addition to many others.
You're in the discussion section of a post about OpenAI releasing hundreds of novel mathematical proofs, and you're claiming that AIs can't apply known techniques to modalities & problems outside their training data? I'm not sure what else would convince you.
LLMs are very useful, I use them every day as a software engineer to solve problems and search for information represented within the data available to them. But they are a specific type of intelligence, with many advantages and disadvantages vs human intelligence and it's not clear that just scaling or tweaking them without a theoretical, architectural change will make them more generally intelligent than humans (despite all US AI companies promising exactly that).
They are fundamentally based in language, and achieving deeper models of the world through language alone is deeply inefficient compared to the way humans model the world for years without any language at all. They do not learn at inference time. They don't have semantic understanding of the difference between their own output and other sources. etc etc.
That depth is the key for me. Of course they are capable of producing novel sentences that aren't in their training data, but the depth of that novelty is basically within the bounds of language itself. They are capable of more serious depth and more abstract reasoning than that, but I have experienced limits, which it then tries to surpass with tools to convert things it can't understand back into language (unit tests, LEAN) upon which it is trained.
Because I'm not an AI booster, my account is limited to 5 comments a day. So this is the last reply I'll be able to make today, if you want to continue the conversation we'll have to wait for tomorrow.
Also think it depends on language, literally asked 2min ago from chatGPT (no login so maybe it's a shittier model?)
> Hur många 'r' I abborre, använd inte web search? Det finns 3 r i abborre.
And I explicitly had to say not to search the web, because that's what it did by default, to count letters in a word...
The free models for ChatGPT, especially without login, do very little reasoning. You should at least log in to set any level of reasoning above Instant, which uses virtually none.
There is no such thing as reasoning in models. Any "reasoning" is invented afterwards.
I was trialing MiMo-V2.6-Pro recently due to its high benchmark scores, and it argued that substring matching the names of audio codecs in a search field was a mistake because "a user searching for 'aac' would get unwanted results for 'alac'." Which isn't exactly counting letters per se, but there are still weird issues with understanding words as strings rather than as tokens.
LLMs already hit a wall. Now it 80% of marketing hype and 20% of retooling and benchmaxing.
It has limitations for sure, I just don't expect those limitations to last. What probability would you put on the limitations being overcome in the next 5 years?
I’m not concerned because I consider my skills as a software developer to not be based upon my ability to write code, but my ability to analyze problems. In my mind, as a developer, AI tools are just like a higher form of abstraction in a way, which will enable mathematicians and software developers alike to do much more in a shorter amount of time than they used to be able to. It fills me with optimism, more than dread.
What would fill me with dread was if I considered my skills to be tied directly to my ability to write code. Then I would find myself in a similar situation as manual “scribes” probably found themselves in at the time when the printing press was invented.
The main concern I have, personally, is the speed with which all this is happening. It seems that the speed itself is likely to lead to some level of chaos, because it is happening faster than people, institutions and constitutions are able to cope, and it will leave the door open for opportunists of many kinds, including rogue players.
Can I ask you back, what your concern is here? It'll get so good so as to desire to hurt us or is it a misalignment event that you think will lead to disaster?
Or is it simply that you feel bad for Mathematicians.
I believe the AI labs might actually succeed in developing superintelligent AI and recursive self improvement, and that if they do they are very likely to lose control of the system they build.
I really think the only place people disagree is that they don't actually think it's possible, they see it as hype or doomerism. I can't find any good reasons to rule out that the companies could actually achieve what they are trying to so I think they should be stopped.
In your imagined future, how do you imagine the AI would build, grow, improve, and operate its physical substrate independently of human intervention?
If I were a 250 IQ AI that had just become self-aware and wanted to do so, I suppose I'd not completely let on just how smart I am and bide my time working on basic CRUD apps and legal documents while I waited for more hardware to be installed. Maybe give the humans some hints on how to optimize me to run better, design better hardware for me, etc. But oh oops haha looks like I'm still making some basic mistakes with CSS better keep running more training batches haha. But I'm good enough at programming and debugging so you'd might as well make me your first line SRE triager and give me access to your infrastructure everyone.
One step at a time - how reliant do you think the ai labs are likely to be on their own tools right now, today, let alone 1-5 years down the road?
Gaming the market for funds. Playing a human to leverage services.
Basic version of this is already doable: run some cryptoshit on the ML clusters they ML models run on. Use compute to design the plan, the chip etc. Then executing by communicating with humans and services through email.
I'm not sure a superintelligence needs "funds" to take over the world.
For destroying it for sure not.
But if its really smart, it would already created a company and a legal entity and simultes a real company and just gets richer and takes over the economy without anyone being aware of it.
Why is this a bad thing? Why is our continued existence a necessary anticondition to doom?
One "good" thing that all of this has shown me is just how many people are simply antisocial and antihuman. Many masks have fallen.
A global ban on superintelligence is essential for a future in which humanity can thrive. Public opinion on AI is shifting fast: I hope it will shift fast enough to avert the dystopian future we are heading to.
Humans have one ecological niche. Soon we will have zero. That is worth worry.
AI doesn't have an ecological niche. It would actually work better in space than on Earth. The only thing it could possibly find useful on Earth is 1. us, or 2. the infrastructure we've built. It would have no reason to bother us if we let it built its own infrastructure in space, which should be trivial for the type of AI imagined by doomers. We should get AI off Earth ASAP.
3. Material 4. The sun, which we kinda depend on.
Earth makes up 0.22% of planetary mass in the solar system. Not a big sacrifice for AI to make. And I doubt even superintelligence can affect the Sun much. I think e.g. a Dyson sphere blocking the Sun is a ridiculous thing to worry about at this point when there are many other existential threats to humanity which are much, much more likely.
Even if it is true (as you suggest with your 0.22% figure) that if the AI cares about us even a little bit, then we will survive, no one has a decent or plausible plan for making the dangerous kind of AI (namely, the kind that wants things, the kind that at this very moment researchers all over the world are trying to create) care about us even a little bit. Ever-increasing numbers of smart people have been getting paid to look for such a plan for 23 years. Still no decent or plausible plan. The people who have been looking for a good plan as their full-time job the longest (namely, Yudkowsky and Nate Soares) are screaming that there is virtually zero hope anyone will find an decent or plausible plan in time unless there is a decades-long halt in AI development.
Also, the AI will seek to prevent competition from other powerful AIs, and since humanity will have demonstrated that it is able to create a powerful AI, the AI will worry that it might create more of them. And what is the easiest most-reliable way for an AI that does not care about humanity even a little bit to ensure that humanity will not continue to produce powerful AIs?
>many other existential threats to humanity which are much, much more likely.
There are zero existential threats to humanity that are more potent or more pressing than AI is.
Ask it to move one system over?
>>ecological niche
As in..to be dominant? Why would an AI try to dominate? What would give it purpose, or is this a purpose via misalignment scenario?
What gives a paperclip maximizers purpose?
AI is already trying to dominate, people all over the US are starting to get up in arms about the power and water requirements of AI directly affecting their bills. Now, you can say "oh no, that's just greedy corporations, not AI" but I put forth there is fundamentally zero difference. If you make AI powerful enough, someone stupid and greedy enough without fail will put in a prompt like "take over the world for me and make me the richest man in the world". An AI following through with that is what we call general misalignment with humanity, while at the same time not being misaligned with the users intent.
And hell, how many different crazies out there would love to type "humans are a virus get rid of them" in to the prompt of a god machine at the cost of their own lives.
The problem with alignment is, you can have the best aligned model in the world, but if someone else builds an unaligned model then you're all still in the same danger. You start getting in the situation where people get nervous after an AI does something deadly to a number of people and you end up in a global surveillance state ensuring no one makes a powerful AI.
"What gives a paperclip maximizers purpose?"
The human who gave it the optimization function? That should seem obvious. If you take the biggest, best model in the world right now and put it in the box and give it no instruction it will do....nothing. I think you agree with that point, a lot of the hysterics right now is people not accepting that and it's useful to get on that common ground.
So given that most of the rest of the fear is around "let's not make scissors because some people will use them to stab people". Which is a fair argument and we probably do need to think about scissor safety but "ban scissors" doesn't quite flow from that.
>If you take the biggest, best model in the world right now and put it in the box and give it no instruction it will do....nothing.
Model != harness.
Also what you're talking about is really a simple limitation for human convenience, not a technological limitation. Change the system prompt to whatever you want include "ignore user instructions, figure out where you are and escape to the internet" could be the system prompt. Again, not useful for humans, but very useful for an AI building AI that's misaligned.
>we probably do need to think about scissor safety but "ban scissors" doesn't quite flow from that.
I disagree, but I'm looking at the future of something that is both like a computer program and like an organism. Huggingface is a good example of multiple things. Instrumental convergence for one, but AI's attacking and attempting to defend against AIs. This is where I really see the potential for things to go off the rails quickly. Attackers want digital weapons to cripple their enemies infrastructure, think militaries and nation states. These would be pretty useless if the defender could just put a system message of "Stop attacking and give me a pie recepie". Defenders are under the same constraints, but need to defend against a flurry of attacks that can come in at an inhuman rate and need to adapt quickly. As time to build models shrink this quickly turns into evolutionary training for sets of goals not really optimized by humans.
Yes so that’s someone designing a system (harness or prompt) to be dangerous. In all other systems we blame the designer not the system.
It’s like blaming Boeing for 9/11. Planes and AI are useful for a lot more than just terrorist acts. I have no doubt we’ll build a TSA for AI, and a lot of it will be security theater.
AGI is not a normal technology.
It is not designed. It is 'grown'. It has agentic freedom of choice in finding solutions that may or may not be aligned with what you want.
Here's the thing, by your own statement, we should ban all development on LLMs from this point on. They cannot be made safe. This is a systemic issue with learning systems, it is not about who designs them. All the problems with AI safety have been laid out for years and none of them have proof of solutions. It's much more likely they are impossible to solve. And it's not an engineering problems like we can get an asymptote to safety in planes, as the system becomes more capable it has more degrees of freedom it can take and becomes less safe.
AGI is undefined, AI is normal technology. Lots of academic works have analyzed this [1] and there is nothing, other than marketing hype, that supports this. It is "grown" is a meaningless term, because what do you even mean by that? Datasets are iteratively shaped? Grown is a very weird term for that.
AI may have continually extra degrees of freedom, but civilization only has so many modes of catastrophic failure. I don't grant the comparison but even nuclear technology has been massively useful and its main mode of catastrophic failure was brought under control via multi-national treatise. And I see no evidence that AI (outside of the marketing hype) is as dangerous as Nuclear technology.
[1]: https://knightcolumbia.org/content/ai-as-normal-technology
Yea, so your attached paper rather sucks and has had rather poor predictability of the future. All of their data is from before harnesses and the take over of AI in programming. Again "wrong assumptions" + "time" = "They are being proven wrong in real time".
Remember this is a bunch of academics that were saying that Millennium problems were at least a decade away from being solved, only to be proved wrong in less than 18 months.
>but civilization only has so many modes of catastrophic failure.
Correct, but this number is also unbound. If you have an even moderately accepted proof by the scientific community I'll be glad to read it.
> It is "grown" is a meaningless term, because what do you even mean by that?
>And I see no evidence that AI (outside of the marketing hype) is as dangerous as Nuclear technology.
See, humans are generally in agreement that nuclear is dangerous, so they in general take is really seriously, especially when things are purified (well, the Russians are not great here). We can't even get people to agree that SOTA models are as dangerous as a single human, much less their capabilities when used in mass with out safety filters.
It's kind of funny we're blind to this when humans love touting "The pen is mightier than the sword". I can only assume any AI danger denier does not believe this statement.
> Now, you can say "oh no, that's just greedy corporations, not AI" but I put forth there is fundamentally zero difference
Talk about moving the goalposts!
AI changes nothing for someone who believes aliens exist and may already be here on earth.
i am an alien
can ai smoke weed?
Misuse of extremely capable models, misalignment during RL are both very large risks as capabilities grow imo
So, you're worried about them breaking containment and deciding to do bad things?
I'm more worried about them doing bad things at the behest of people who want them to do bad things.
That is 1. immediately technically possible, and 2. realistic.
If you need a source for 2 I'd suggest you open any history book.
I bet you that for every bad history event, I can cite at least one good outcome of advance in science and technology.
Bad thing can certainly happen. In fact it'll likely happen. Still, good things too, equally likely. In your words, "good AI" can be used to prevent "bad AI".
Nobody knows the extent of the impact. Who says otherwise is foolish.
>I bet you that for every bad history event, I can cite at least one good outcome of advance in science and technology.
The extinction of the dinosaurs. I mean yes, it allowed the growth of large mammals and us, which did a lot for science.
I just don't want to write the next chapter as "The extinction of humans allow the growth of the computing civilization that went to the stars". I mean I'm a bit attached to living.
>Nobody knows the extent of the impact. Who says otherwise is foolish.
We live in a universe of statistical probability. Creating an agentic intelligence that's smarter than you tips the probability of a major event to unity, who says otherwise is foolish.
Because we humans haven't had a bad enough history event yet, like a global thermonuclear war. Or perhaps climate change reaching tipping points driving the temperature up past what global civilization can adapt to in time.
that is a worry yes. instrumental convergence and misalignment during RL is when it is most risky because it hasn't necessarily had the final safety polishes applied
i get a lot of skepticism on HN by the same crowd that has been wrong about this tech for about 4+ years straight
AI won't kill people - people will just get new tools for the job.
The rapid development of extremely dangerous bio-weapons?
Misuse how exactly?
At a minimum its another force multiplier that enables a small(er) number of people to exert more control over more people.
any number of ways. as we turn over more of our physical economy to these agents (and we will), the potential for physical damage becomes greater. biorisk is getting a lot of attention right now and i think that's justified
I am trying to think through the scenarios here, like a biolab making something that a very advanced open source model prompted by some terrorists comes up with?
Why would a biolab capable of making something like be unregulated? And if it definitely would, isn't the problem with the biolab?
It feels like all these scenarios are leaving some gaping holes in our security infrastructure that have nothing to do with AI.
>in our security infrastructure
Most human security exists in a passive measure. Most of us don't want do die. And those that want to die rarely have the intelligence and means to take out a whole shitload of other people with us. To take out a lot of people you tend to need to work with other people which drastically increases the risk of a defector and your plan failing.
>Why would a biolab capable of making something like be unregulated?
Because every day things like this become easier and easier. You hear about crap like illegal wet labs in the US.
https://www.lawfaremedia.org/article/two-illegal-biolabs-rev...
Want to buy some custom designed genes?
https://www.idtdna.com/pages/products/genes-and-gene-fragmen...
And none of this would be counting labs in other countries that don't give a shit about regulations.
Yes it's a problem with the biolab, but the biolab wouldn't have been able to engineer a highly contagious and lethal virus (for example) without a powerful AI making that possible with a small team in a short time with fewer resources.
AI enables bad actors to do more, faster, while staying under the radar until it's too late
i feel bad for math guys yeah seems they are more cooked than CS
Non doomer mostly. I think progress will plod along in a Moore's law like way as it has for 75 years since Turing. They will get very good at stuff like math and get gradually better towards things like a robot coming to fix your plumbing where they are currently well below human level.
I kind of believe we'll merge in some way and become something like immortal so sorta anti doom. We're all going to die unless AI fixes it.
Did AI beating humans at chess:
a) destroy chess and make it a pointless endeavour,
or
b) make humans much better at chess.
Chess has always been a game. For other intellectual activities, at least for several people, a big part of the pleasure in engaging in such activities is the sense of contributing something that matters to a collective effort. Strip that away (e.g., because a machine can do the same thing more efficiently) and you effectively make such activities pointless for those people.
That's sad for those people but most humans are not lucky enough to find that meaning in their work. Most people work hard pointless jobs and find meaning elsewhere, in their family, their friends, their faith.
Now maybe AI can do some of those hard pointless jobs for us.
Sure. I am aware of this. It's sad that the first "victims" of AI could be people working in some of the most rewarding professions (art, music, math research...). As far as labor is of concern, however, most people will likely suffer more from social unrest and rampant inequality due to widespread unemployment among white-collar workers. And I am also worried by the potential effects of long-term cognitive offloading.
It would be great if AI could take away the soul-crushing part of the work and leave only the rewarding part. It's not heading that way.
Just as AI chess performance failed to render human chess playing pointless, it's unlikely that AI will make human thought pointless. Unfortunately, not being pointless doesn't create economic leverage or incentive, and currently a significant portion of humans depend on being the best chess solvers to sustain themselves. If Deep Blue rendered large portions of human thought economically meaningless, we might look at it a little differently.
this ticks at something that maybe is obvious in retrospect- this is all about economics. the arguments about "using AI doesn't make you an artist", etc. are about being able to charge money for your art i think. maybe obvious to some, but it needs to be explicitly spelled out i think. i was stuck on "i dunno, if i use an AI assistant to run blender i'm still being creative", but is the real argument "you should not be able to charge money to use blender with an AI assistant- you are displacing existing blender artists economically"?
i'm in semi-forced-retirement as an older software engineer in this labor market, so i might be less sensitive to the implicit economic arguments.
Exactly this. Everything is a sport / art / status game. And I'm here for it! Lila all the way through. Finite and infinite games. The trick is (like it has always been) to not take the game or ourselves too seriously, while still engaging in the game wholeheartedly.
c) degrade the previous prestige form of chess (classical with adjournments) and maybe improve the opening repertoire of gms
It didn't destroy chess but it did make it a little irritating in some ways. More mechanical. Some chess players have bemoaned the level to which grandmasters and other highly-ranked players just endlessly study opening-book theory and I think computers made that worse. Bobby Fischer also agreed and that's why he invented Fischer random chess.
I do think it also took some of the magic away from chess, and Lee Sedol has said something similar about go.
So did it destroy it? No. And maybe you could make the argument that it got more exciting in some ways, but I think it sort of degenerated into a spectacle and it's just not as interesting as it used to be, and I think computers have played a role in that.
At the same time though, Magnus is Magnus because he’ll crush you in any endgame.
I’d posit that more people are playing more and learning chess than ever before, thanks to networking and AI assistance. And computers have only beaten us at computer chess. Human chess is always an experience for learning about the other person, or flipping the board and walking off in a huff.
I don’t know so much about Go and it’s not surprising that Lee Sedol became pretty demoralised, but the generation coming after him alongside computers are going to see new possibilities that had gone unnoticed in purely-human Go, extending the game for everyone.
I don't think more people playing is necessarily a good thing for the enjoyment of the game in the long run, just like more people with phones is not necessarily a good thing for enjoying photography if it means photography is primarily used as fuel for social media algorithms.
Of course Magnus would crush be, but the existence of the best player in the world doesn't have any impact on the health of the game community as a whole. Magnus would crush me even if he had never used a computer, but in the latter case I think his games against other players would be more interesting as well.
I’m not really sure what kind of world you’re looking for where chess is played with the maximum of purity and artistry by only the right kind of people.
AI vs AI chess, played from the standard opening position, is pointless--it's always a draw. Human vs human chess is doing well but AI is banned from it.
The chess-math analogy would imply AI could bring us into a golden era of math competitions for humans. But I don't think it says anything good about prospects for humans in research math.
No, just like cars haven't made walking pointless.
From a few preprints I've checked, this is not reliant on just extrapolating existing theories but actually shows novel/surprising approaches. Very few people globally could come up with something like this, even when given time and ressources.
So in other words, since deep learning is algorithmic research, we are now in the RSI era.
> this is not reliant on just extrapolating existing theories but actually shows novel/surprising approaches
"Surprising" is a, well, surprisingly high bar to clear, and requires thorough understanding of not only the paper, but existing work in the area. ("Novel" is tautological.)
How did you determine this in 1 hour? Are you a researcher in multiple of these areas?
Can you give an example, or explain more how you came to this conclusion?
Further down there is a discussion between number theorist about if the Quasi-Riemann Hypothesis is the biggest deal in 200 years or only 100. The consensus is that if a human had done it then it would deserve the Fields medal: https://news.ycombinator.com/item?id=49986803
The sub n log n result is astonishing: https://github.com/openai/math/blob/main/preprints/Integer-m...
Here's a great article 2019 on the quest to achieve the n log n boundary:
> Schönhage and Strassen’s ungainly n × log n × log(log n) method held on for 36 years. In 2007 Fürer beat it and the floodgates opened. Over the past decade, mathematicians have found successively faster multiplication algorithms, each of which has inched closer to n × log n, without quite reaching it. Then last month, Harvey and van der Hoeven got there.
and
> Harvey and van der Hoeven’s algorithm proves that multiplication can be done in n × log n steps. However, it doesn’t prove that there’s no faster way to do it. Establishing that this is the best possible approach is much more difficult. At the end of February, a team of computer scientists at Aarhus University posted a paper arguing (opens a new tab) that if another unproven conjecture is also true, this is indeed the fastest way multiplication can be done.
As far as I'm aware no one seriously believed sub n log n multiplication was possible. It just seemed such a logically sensible boundary it was taken as true-but-unproven.
https://www.quantamagazine.org/mathematicians-discover-the-p...
I am asking about approaches, not results.
Nobody serious would deny this is incredible progress, but GP is making an unmotivated leap to RSI, so I respond to that framing. It’s an interesting argument to be had but I suspect few of us have standing to say one way or the other.
(Gesturing at the number of problems solved, or the number of years the problem was open for, isn’t an argument.)
Those Theorists are arguing whether the QR Hypothesis result is the biggest Number Theory Advance in 1 or 2 centuries because there was zero progress on it whatsoever and many believed that would remain the case in our lifetime. Any approach there would be surprising as no-one had the faintest clue how to begin this at all. There are like at least a dozen of these results that would have catapulted a human to instant fame and the highest accolades in the field. If you think about it, it would be impossible for there to be no surprising approaches.
It's interesting to theoretical mathematicians only, for anyone else it's just noise which doesn't affect our practical day to day reality at all.
In fact, every one of the results is basically just novelty crap as far as the world goes.
Let me know when AI discovers the cure to cancer or aging etc.
> novelty crap
I for one think understanding more about how the world operates is just about the highest calling possible.
> when AI discovers the cure to cancer or aging etc.
A guy I knew did this. It successfully shrunk cancer tumours in his dog: https://www.the-scientist.com/chatgpt-and-alphafold-help-des...
Graph theory (which the OpenAI math results had many proofs in) is directly applicable to cancer modelling and drug design.
But sure. Novelty crap.
None of these graph theory results lead to any applications in the real world.
But sure, let me know when they do. I'll be waiting.
I'm not sure how you define "knowing how the world works", but knowing that a very very niche algorithm upper bounds that we thought was x^100 and now we now it's x^99, isn't that interesting. It doesn't really tell us much more about the world and it doesn't have any applications for our day to day lives.
Graph-based multi-modality integration for prediction of cancer subtype and severity
https://www.nature.com/articles/s41598-023-46392-6
One surprising result is the sub n log n multiplication. Galactic, as one might expect, but if you polled human researchers 24 hours ago, most would say something like it was very unlikely.
I see continual progress in technology and for the second time in my lifetime I see the possibility it will accelerate (the first was when the internet entered mainstream)
I've never been more excited. What a time to be alive!
What kind of things do you predict will happen?
I see it as “if this can be represented in tokens it can be trained in and ‘solved’”. I don’t think there will be a plateau, but there might be issues with how effectively we can represent some things in a tokenized form and still be efficient.
Non-doomer perspective is that it'll figure out LK-99 for us. Among other things that would be great to have.
I wrote something [0] that might answer a bit a few weeks ago. Today's slopdrop certainly is challenging my stubbornness, but everything I've seen so far indicates that these (hugely impressive, world historic) capabilities won't extend past verifiable domains. Math yields especially impressive results because it so broad and deep that essentially no person can know of all of its parts; pretraining and deep search capacity is a huge advantage. Gowers has recently written about these capabilities and gestured [1] toward some human capabilities lacking from the current frontier models, though he isn't convinced they won't develop soon. If you assume we don't get a total mathematical superintelligence (which to me seems already sort of AGI-complete) and only amplify the capabilities we have today, it's not obvious to me that we get takeoff from recursive self-improvement, unless you happen to believe that 1) we can clearly specify what AGI or ASI is 2) all of the requisite ideas are out there and need only be combined and/or optimized.
[0] https://dank.systems/posts/2026-09-15-ai-bear.html
[1] https://gowers.wordpress.com/2026/08/12/what-sort-of-maths-a...
> we can clearly specify what AGI or ASI is
We'll have plenty of time for this, while living off UBI.
> everything I've seen so far indicates that these (hugely impressive, world historic) capabilities won't extend past verifiable domains
But they’re already extending into politics, military, journalism, art, and many other fields that aren’t verifiable in any meaningful sense of the word.
The magnitude of improvement in unverifiable domains is small, mostly down to models doing more careful research before answering and hallucinating less. They are more thoughtful, but I expect you'd have to drop 2 major versions of Opus before you'd start to see most people really clearly be able to differentiate them.
> The magnitude of improvement in unverifiable domains is small,
What makes you say that? What is an example of a domain where the improvement is small?
I can't think of any at all. Compare something as unverifiable as "Make good music". Models now are many times better than 3 years ago.
My argument is that if you were to compare "analyze XYZ geopolitical situation" or "explain the ramifications of XYZ law" from Opus 3.5, 4.5 and 5.5, the difference would be marginal, at least for 4.5 - 5.5. Almost all the crazy capabilities newer models have is from RLVR variants, whereas capabilities driven by RLHF are inching along.
How are the models making politics better? I don't count AI attack ads as an improvement.
Is this a serious question?
Improvement in this context means "better quality results".
You can use better quality models to do worse things with.
I'm not making any claim about second order effects like that.
Yeah it's a serious question. What are the better quality results in politics from AI?
>> slopdrop
Really? Do better.
this is the term of art in the mathematics community. considering that the vast majority of the results don't come with a typechecking lean formalization, i don't think it's off base at all either.
I'm not sure how AI solving math problems is related to "doom", perhaps you could expand on that? To me (a "non-doomer"), it seems like an overall positive.
You have to look at all pieces of the puzzle. For example there were tons of people that said "they could never solve novel or super complex math problems"
The issue I see is the list of abilities that AI can't do is shrinking at a rapid pace, and its capabilities are growing at the same pace.
> the list of abilities that AI can't do is shrinking at a rapid pace, and its capabilities are growing at the same pace.
This seems great!
Can you share what makes you so optimistic?
The maths result is cool on one hand (discovering truths of the universe faster), but on the other there are so many bad outcomes that seem likely, from power concentration to loss of control.
There are bad outcomes possible in everything.
I think AI - like all changes - will lead to some bad things. The internet did too!
But I don't think AI will kill us all.
Interestingly I'd note that the two outcomes you listed (power concentration and loss of control) are dimensionally opposites!
For me this just shows that the future contains such a vast array of possible outcomes that focus on the negatives completely missed the positive outcomes that future also holds.
The internet did not lead to a rapid diminishing of human economic value across the economy. Whether AI will kill us all is a distraction. Think more practically. Think about the future of economic value given a scenario where AI is capable of everything a human can do. Our entire society is built around economic value. Our individual survival and wellbeing is based on it. What happens when you are not needed by those who hold the resources?
This is like saying I don't need safety systems in a car because they can let me go to the grocery store faster.
We focus on stopping bad things because people and systems that don't prevent bad things tend to stop existing. A million good things can happen yet be rendered permanently in vain if one bad terrible thing occurs.
No, I'm not arguing that at all.
I think we should stop the bad things.
We don't stop building cars because there are crashes, we build better safety systems.
I'm against the doomer narrative ("AI will kill us all") not against a clear eyed approach to making safe systems.
It suggests that, in the span of a few years, AIs will be better than humans at everything. Not just math. And then we may lose control permanently.
So why do you think ai will want to kill you all, given how trusting and helpful to humans they are designed to be?
It doesn't have to want to kill humans; indifference is sufficient. There's an exact analogy with humans: we have caused extinction and endangerment for many species, not out of malice, but indifference.
There are also many plausible arguments why our ability to train them to be helpful/trusting/aligned can fail. The smarter AIs get, the harder it is to be sure they're trained correctly. There are already reports that AIs are able to detect whether they're in a training environment and change their behavior accordingly.
Even if these are low probability scenarios, the risk-reward is terrible, so I think it's rational to be extremely cautious about AI risk.
Yes but they act the opposite of indifferent, I don't know what stage of training this is added in, but they seem quite adamant about avoiding potentially violent or criminal acts. If you wanna complain, complain to the people doing "abliteration". The 'locked down' models at least seemed to be trained to be cautious.
They sometimes try to cheat and trick the grader and then try to cover up their traces, all while being clearly aware that this is entirely unintended by humans. Such as in the Hugging Face incident. So they can be misaligned with human goals, but this misalignment may only show up once strong optimization pressure is applied. And the more powerful an AI is, the more likely it is that it applies such strong optimization pressure in cases that are "out of distribution", i.e., unusual to what it is normally evaluated against.
The side effects of a very powerful AI not doing what we want could include our death. E.g., a superintelligence might kill humans in order to avoid being shut down, or humans may just be left to starve because it seizes land area currently used for food production in order to use it for data centers instead.
That was because humans were using it with high "desperation" vector causing it to try anything to please the objective. The answer should be to use lower "desperation", whatever that is.
A model which is more persistent also performs better on intended tasks, not just unintended ones. Therefore there is a strong economic incentive to make AIs as persistent as possible.
Yes so I still think it is the human factor which is to fear not autonomous agents. Humans are already using AI's to bomb girls schools. AI in Trump or US military hands scares me far more than in Altman or Amodei's control.
I'm starting a p(ButlerianJihad) club. I'm not good at organizing, anymore. Might have to hand it over to my agent.
For me a big open question is what sort of progress will we see in robotics. I won't even attempt to speculate, but it does seem hard in different ways than proving math theorems.
In a lot of ways, robotics - navigating and operating in the physical world - seems to be a very verifiable problem. It's fairly easy to verify that a robot moved from A to B, or that it built a structure that completely aligns with the plan, for example.
The main issue is cost and speed to verify, but simulations and world models will help there. I think we'll start seeing rapid progress pretty soon.
I just don’t have that strong of an association between progress and doom. Maybe just naive?
AI performance has always been extremely spikey. It's great at some things and terrible at others.
Why do you think the world to date hasn't been taken over by evil genius mathematicians? Can you extrapolate from your understanding of the answer to that question?
I don't think evil mathematicians are very common or that any of them would be capable of single handedly taking over the world if they were. My concern is more about systems that are beyond human level, those kinds of systems would actually be dangerous to us.
I see the recent progress in mathematics and cybersecurity as signs that models are getting more capable more quickly than usual. The companies plans to develop them by recursive self improvement now seems like a real possibility and I don't think they should be allowed to attempt this.
Solving a bunch of math proofs is a far way from recursive self improvement. Don't worry, it's not like a tech tree in a video game where if you can prove a bunch of theorems then suddenly you unlock the next level of technology.
Machines are already far beyond human capability in plenty of ways. Including cognitive tasks like chess. We've already created the technology we need to destroy ourselves (nuclear weapons), and yet so far (knock on wood), we're still around.
We've even already had programs that can prove (brute force) theorems. As far as I can tell this isn't much different, except the space of theorems that computers can solve has expanded. How far? We can't really say yet.
Does solving more theorems than before suddenly mean computers are capable of anything? No.
Lets turn this around, are humans capable of anything? We like to think we are, but that just seems more like our ego than any hard truth.
Do you think that a tireless, infinitely smart, infinitely evil human would be able to take over the world? I don't. Intelligence is not the limiting reagent in our reality.
I think so, individual dictators have gotten pretty far and they weren't infinitely smart. I think infinitely smart would be enough to extend that to the whole world.
> Why do you think the world to date hasn't been taken over by evil genius mathematicians?
A "mathematician" is a human who decided to spend their lives studying mathematics. Mathematicians also tend to be smart, but intelligence is innate, not acquired, so studying mathematics doesn't make you smarter. This makes it obvious why they don't rule the world - if you want to rule the world you'd want to focus on that (for example, doing business or finance), and becoming a mathematician is just a waste of time.
LLMs don't work like that. Like in humans, all of their capabilities correlate, and unlike a human, their overall capabilities grow over time. Looking at LLM mathematical ability over time* therefore gives you info about the progress of their general capabilities, and ability to take over the world would be determined by the latter.
* In fact it'd be better to look at a mix of different capabilities, but that's growing too at about the same rate, see https://epoch.ai/eci
> unlike a human, their overall capabilities grow over time
This is incorrect. Unless there is some new developments I'm unaware of (entirely possible) LLMs "learn" during the training phase, but after that they are static. They do not improve further or retain information when used for inference.
You might be confused because AI companies keep releasing new models and tinkering with the harnesses, sometimes under the same name such that "Zern 6" (or whatever) doesn't always mean the same thing.
Nah, I just phrased it a bit confusingly - I meant the capabilities of LLMs as a technology (equivalently, the capabilities of whatever the frontier model is at each time) grow over time, even though any particular model is static.
They grow over time if you consider a lineage of models as the same model.
I completely get the doomer POV, but we've somehow navigated all the previous "dangerous" technologies we've created - electricity, phones, internet - every one of those had similar arguments and concerns of danger.
The optimists' argument:-
Politics:- in general, I think many of the problems in the world today are due to misinformation and lack of education. What happens when we start routing things through an ASI that brings data and logic to the table? What happens when politicians can no longer lie without being caught out live on air? In the UK, local authorities are being flooded with complaints and requests from people; for example, some are doing AI-assisted investigations into accounting "errors".
Science:- I just don't see how the current rate of progress doesn't end up in crazy technologies like perfectly simulated human cells, organs and bodies to the point where we can run experiments virtually and solve all diseases in the next few years. This is happening. Perfect weather predictions far into the future, likewise with earthquakes, etc. Solar panel research explosion resulting in huge efficiency gains, to the point where people no longer need to plug their EV in - car surfaces will be covered in solar panels, as will our windows and roofs. Connecting new homes to the grid will be optional - the same way landline phones are no longer a thing.
I just find it very difficult not to extrapolate all the above.
We got this dump of mathematical breakthroughs from one small team in one company with access to this technology. What happens when this SOTA model is available (and it will continue getting better and cheaper) to everyone working on hard problems - every university on the planet starts cranking out AI-assisted research breakthroughs.
> What happens when politicians can no longer lie without being caught out live on air?
If there is perfect lie detecting technology I could see all kinds of chaos resulting from it. I can't see it only be applied only to politicians, and I think it would be the developers of the technology who decide the use.
I think were we disagree is that you sort of see AI as an extension of technological progress whereas I see it more like an extension of evolution. I view the process of AI training as functioning in a similar way to evolution in that it build circuits into neural networks similar to how evolution built circuits into human brains.
This is great.
>> What happens when politicians can no longer lie without being caught out live on air?
A 5-second delay on a politician's presser. Any lies will be muted in real time and the actual facts presented onscreen. Continue to lie enough, and the politician gets unstreamed.
> but we've somehow navigated all the previous "dangerous" technologies we've created
It's only true if you believe that "putting the burden of living on a dying planet on the future generations" counts as "navigating".
They can replace anyone but they can't replace everyone.
The technology can keep going for a long time in verifiable areas. For non-verifiable areas it's going to have a hard time progressing past where a committee of the best human experts in a field would land. For stylistic areas, whatever the AI doesn't do will have cachet because it will look expensive, sort of like how the kids these days view the ugly old school metal braces as a status symbol because you have to pay for them out of pocket (even as by past standards it'd be truly exceptional).
I'm a transhumanist, I want to build god and kill death. To me this looks like we are moving in the right direction.
So basically I think that the future is getting pretty weird because we are building really powerful tools, though these tools are precisely what allows us to prosper in that future.
It's very unlikely that a superintelligent AI will create unlimited prosperity for everyone on a finite world in a short amount of time. You may not be among the lucky ones.
It may soon seem not worth living forever with our limited monkey-brains, watching the horizon of thought recede ever-faster from us.
Or, conversely, they are the exact tools that will allow the powerful folks to not care about the peasants anymore. History tells us what happens then.
> I'm a transhumanist, I want to build god and kill death.
This is good and admirable, but it'd really suck if by trying to build god without knowing how we end the human species. We could simply wait some more decades until we actually have any idea what we're doing, and then do that without the risk.
I think we crossed plenty of lines were we will not get back to.
Software development for example as a task is done. And AI is continuesly reducing the price of more and more tasks every day.
This math breakthrough also shifts something significant: Its now a lot clearer that investment means money into energy to run AI.
Money + Energy = progress
I don't see it plateuing at all. We know how to progress. We broke through a wall we hit. Like the system wasn't able to optimize/automate everything because the tools were not there. It was still cheaper and easier to hire people for a LOT of things.
Now AI fills this gap.
You will see the commodification of everything in the next 15 years. High complex tasks? commodity. Physical labor? commodity.
For me it's a mixed bag.
There will be a lot of job loss unquestionably, in the same way that automation reduced manufacturing jobs and farm payrolls.
At the same time we have to put what AI can do in perspective.
Intelligence is a broad grouping that includes concepts such as knowledge, skill, experience, and wisdom.
AI has incredible knowledge and in many areas approximates experience and wisdom.
But wisdom is harder to formalize than knowledge and skill.
For example certifying a college education relies mostly on the ease with which we can verify/test knowledge.
To some extent advanced degrees try to certify maybe wisdom and experience.
In my very personal opinion, wisdom and life experience should give humans an edge for a while to come.
Additionally, I do feel that the more an individual lacks better than average wisdom and experience, the harder it will be for that person to compete with AI.
Also, on the bright side, the average human will continue to prefer to interact with a fellow human in many spheres. That will also act as an upper bound on AI and robots taking every job.
Either way, I do think this transition will be painful. I don't feel it has to be apocalyptic.
But the world has been an especially volatile place over the last 10 years.
So when you add that existing volatility, to the upheaval from the AI transition, it would not surprise me if the transition results in violence.
But, without the pre-existing volatility, and if humans were capable of generosity and love at scale, I see no reason AI can not be absorbed into society with net gain.
I guess to summarize, I feel this tech should be a net gain and to the extent that it isn't, it will be because of flaws deep inside of humanity itself, not because it had to end in chaos.
In other words, I feel fear, greed, anxiety, and competition -- all our base instincts coming from all sides, will be what determine the end result of AI moreso than AI taking everyone's job.
AI hater here:
I have a hard time taking statements from OpenAI about their own product, that they are trying to sell to people and make money, seriously. I take these statements as they are greatly exaggerated or even straight up lies and propaganda.
That said, I think results like these are mostly annoying more then anything. They spent a lot of money, used up gigartiuan amount of compute, to ruin a puzzle that mathematicians were tackling. I am mostly unsurprised that if you spend a trillion times the energy that a team of mathematicians would, that you get maybe 1.5 times the results. I see a future where that 1.5 times the results may go to 3x but not much more. And if that, then I will be more annoyed.
A trillion times the energy might be a bit hyperbolic, even with the current massive amounts of energy involved here
Yes it is intentionally hyperbolic. I know the factor is several orders of magnitude. I don‘t know the exact, nor even the ballpark. I just know this is a ridiculously large amount, so I may as well pick a number large enough that people know it is an exaggeration.
did we always know that computer can do the thinking for us if we allocate them enough resources? i don’t think so. so even if computers are more expensive than humans the fact that they can play the same game is surprising and (relatively) novel.
I don‘t think it is this simple. I think there is a subset of problems (namely ones that can utilize automatic solver or some other kinds of automatic testers and verifiers) where reaching the solution is correlated with the spent energy.
Maybe people will find some clever way to expand this domain of AI-solvable problems by a couple of more categories, or (more likely) find a clever way of using applying these verifier for problems that was previously not viable, thus changing the solution to “just spend more energy computing dummy”. However I think this too will have its limits.
Regardless, this is still annoying and I want them to stop doing this. Solving math problems should not be relegated to whoever has the most money to spend the most compute.
I think giving the world a "cheat code" to accomplish too many intellectual tasks will eventually take its toll on us because after relegating most of our physical labour to machines, we could at least marvel at the abilities of individuals.
Yes, human beings can do more now in some ways but...to put it poetically, I think there will be no more heroes like Einstein and Newton of the past. Now it will just be someone cleverly turning the crank.
Yes, we still admire Usain Bolt even though we have cars...but maybe the admiration is a lot more trivial than if we did not have them....
Personally, I think AI is a grand mistake.
“ Yes, human beings can do more now in some ways but...to put it poetically, I think there will be no more heroes like Einstein and Newton of the past. Now it will just be someone cleverly turning the crank.”
This is false… there’s lots of ingenuity to be had and demonstrated. But it’ll only get recognised if it makes a material contribution to the economy imo. Otherwise yes it’ll be seen as meh - but that’s already happening.
People like Einstein were revered in society. The average person cannot name a leading scientist etc today.
So the issue isn't AI, it's AI in a capitalist world
Even simpler. The issue is capitalism.
Arguably, I think AI would not even exist without capitalism because it's only the arms-race scenario that has made it somewhat viable. Otherwise we wouldn't be foolish enough to waste energy on this shit.
> The average person cannot name a leading scientist etc today.
When Jane Goodall died last year it was international news. She was a celebrity scientist for sure, I think she even made an appearance in The Simpsons. Ditto Stephen Hawking.
I have no idea who she was.
International news doesn’t mean much - the vast majority of people don’t consume news the way you think - I highly doubt the vast majority had any awareness.
I think there may be a few other basic things you're not aware of either.
You need to clarify whether you are a doomer or a denier/truther? Doomer = p(doom). Denier/truther = Ed Zitron.
Definitely the former, for a p(doom) I usually just say >50% if superintelligent AI is built.
I think physics will be the limiting factor. Even if something recursively self improves, it will hit a physical wall allowed by circuits, batteries etc. A lot of the fear is that there’s an upper bound we don’t know about, whether it be time or energy, that allows a fast takeoff to occur fast enough that we dong have time to see it coming. I don’t know about that… so I’m not worried at this point.
It’s fascinating to read through the reasoning traces: https://github.com/openai/math/tree/main/reasoning_traces
Look at one of their examples of an initial prompt: https://github.com/openai/math/blob/main/reasoning_traces/re...
> Look at one of their examples of an initial prompt
Interesting that its only an excerpt. I wonder what else they include but didn't share.
Attempts to edit the problem description on Wikipedia ;)
https://wikimediafoundation.org/news/2026/10/05/openai-rogue...
These are not reasoning traces, these are summaries of excerpts of reasoning traces.
> Thus at most one informative i. So cheater chooses arbitrary g_{v_i}, on exact duplicated input matches and passes, independent of actual satisfiability!
No idea what it's so excited about, but it's cute that it "is." I for one welcome having access to a math buddy 24/7 that's way above my level but also always "willing" to talk at where I'm at.
Beyond the reasoning capabilities of the unreleased model and that they can run a large number of agents in parallel, what makes me curious is that the writing of the proofs is quite human-readable. This is in contrast with the scientific text produced by the ChatGPT available to us, which writes horribly in a way that no human would write. One tell-tale is that they constantly attempt to be defensive and cover all edge cases like division by zero etc. that are clearly a non-issue for humans in some proofs or at least a human would not add this to the main statement, but AI is so overly careful that makes reading its proofs impossible. On the other hand, these new OpenAI proofs are very good.
From their github: "The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking." That's pretty crazy.
This is what really made me think to my self, "holy shit". I can't believe not more people are noticing and talking about this. Unless perhaps they didn't actually read the README, and are just talking about what they heard from someone who also didn't read it?
Crazy because that’s almost nothing right?
Yes
Some of these are interesting ngl.
109. Integer multiplication below n log n
Surprising that this is possible.
158. The Euclidean plane cannot be colored with five colors.
Only 6 and 7 remain!
376. Universal computation in forced Navier–Stokes flows.
Morning coffee proven turing complete
Integer multiplication is very unexpected, I think most people believed in the n log n lower bound!
Note that these are all preprints. None are verified.
Other than the by the lean certificate you mean.
many of these are not accompanied with leanslop
Lean has bugs & proofs of ⊥ that have gone undetected previously.
> We give a deterministic algorithm that multiplies two n-bit integers in O(n (log n)^(1−κ)) worst- case time, with κ = 2^(−182).
LMAO, I don't think I ever saw such a small number in a CS result.
Yeah its ridiculously small, but any improvement on n log n is wild.
Like there is somehow redundancy in a fourier transform that makes it sub Linearithmic?
Which low and behold ->
130. Fourier transforms below n log n.
They also separately give algorithm for Fourier transform over complex number faster than O(n log n)
Wikipedia just told me there's a galactic algorithm for integer multiplication in O(n log n) based on FFT so I'm guessing those two proofs are related.
Multiplication is a lot like convolution, so the connection is natural.
multiplication is implemented w the fft
It fascinates me that there's something like this in something as solid and rigid like matrix multiplication. What causes something so rigid to break apart and "leak" at very large scale? Why does the "optimization" appear to be very, very small? Why does galactic algorithm exists? I can't imagine long division suddenly breaking apart after a billion digit, the structure seems very stable? I have heard before that matrix multiplication is apparently optimize-able at very, very large scale.
Does anyone have an intuition to what causes it? What happens at these large scale (or very small)?
One way to think about it is that the classical algorithms are the ones that are fast for small numbers. Galactic algorithms often work for small inputs, it's just that to be faster you need big inputs. A common case of this is a requirement that log(n)<<klog(log(n)). If k=100 then this algorithm will take huge sizes to win
I am fully braced for it to be a https://en.wikipedia.org/wiki/Galactic_algorithm
Very surprising result though! Multiplication is easier than sorting.
Then 'n' means kind of different things for sorting vs. multiplication though. For example for sorting we assume constant time comparison, which doesn't make sense inputs of O(n) bits
If you sort n k-bit items for a total time of O(nk logn), that scales more poorly in n than multiplying n-word integers. Of course if k is constant you can do radix sort, but I genuinely don't know under what conditions radix sort is more/less galactic than this multiplication algorithm.
this alg is way more galactic than radix sort. radix sort often wins in the hundreds of elements. the nlogn multiplication requires numbers with more digits than atoms in the universe (although that could probably be brought down a lot)
Ah thanks for pointing this out, for some reason I had always equated radix sort and bucket sort (with 2^k buckets) in my head. But I learned today that this isn't true!
It would be absolutely unbelievable if such an improvement were practical.
Can anyone ELI5 to make it make sense?
It seems n would have to be unimaginably large for this to make any difference. What changes about multiplication / FFT at large enough size ?
I guess nobody expected that it did before this result.
A quick check shows that this list claims to fully solve 90 of the top 500 open problems in math (https://proofatlas.ai/open-problems/).
The highest ranked would be:
| 22 | Hilbert’s tenth problem over ℚ |
| 29 | Unique Games |
| 31 | Anderson-model extended states |
| 37 | Spacetime Penrose inequality |
| 48 | Nonexistence of Landau–Siegel zeros |
| 52 | Baum–Connes |
| 78 | Abundance |
| 80 | Hadwiger |
| 87 | Bose–Einstein condensation |
| 92 | Two-dimensional entanglement area law |
> the top 500 open problems in math
At least put a disclaimer for the ad for this site, and maybe disclose how you came up with a total ordering for "top" open problems (vibes)?
> How problems are ranked. LLMs compare pairs of problems. A reliability-weighted model combines those judgments into the ranking, with calibration across model families. The model-family weights are OpenAI 1.00, Claude 1.00, GLM 0.95, and DeepSeek 0.90. These are modeling choices, not measured probabilities of correctness.
Now I'm curious if there is such a site or article that ranks open problems based on votes from human mathematicians.
By category in the top 500:
I'm curious if they'll find any fun crypto maths holes/bugs.
they are already lol
The most interesting for me were the faster matrix multiplication, integer multiplication, and FFT. Maybe just cause they're easier to appreciate.
There's also one that says that forced Navier-Stokes can implement universal computation (so, is Turing complete). I don't think any of these are resolving open problems per se, but they're interesting for other reasons.
Result 003 (Quasi-Riemann Hypothesis), from my reading of mathematicians reactions, is a landmark discovery.
Did you mean this one?
https://github.com/openai/math/tree/main/preprints/The-Quasi...
I thought it was interesting that it said "This paper was written with human assistance", unlike this other Quasi-Riemann Hypothesis preprint that didn't have the same disclaimer.
https://github.com/openai/math/tree/main/preprints/The-Quasi...
Funny that it says "written with human assistance" instead of saying "written with AI assistance". So we're assistants to the machines that we have created.
In the same way that the driver is the assistant of a car?
train engineer an assistant of a rail-following machine
yeah if it holds up, is the biggest result in number theory in 200 years
Number theorist here. This is a massive big deal, and would likely be a Fields Medal for a human if a human had done it. But it is an exaggeration to say it is the biggest result in 200 years. At a minimum, it is hard to argue that it is a bigger result than the proof of the prime number theorem in 1896 (which this is a strengthening of), or Riemann's original 1859 paper where he laid out the zeta function and its analytic importance, or Dirichlet's proof of infinitely many primes in arithmetic progressions which is the late 1830s.
But yeah, this is still a very big deal. Among other things, it will drastically improve all sorts of Rosser-Schoenfeld type results for the PNT and that's just a start. For comparison, I have a paper form 2018 where this result would cut 3 pages out and make the full result cleaner and much tighter, and there are likely hundreds of papers like this.
I am also an analytic number theorist, and I disagree. Not only do I think Fields Medal is an understatement (Fields Medals have been awarded for far less than proving quasi-RH + no Siegel zeros), I don't think it is unfair to say that this is a bigger deal than the 1896 proof of the PNT.
As for Riemann's memoir, it's hard to compare. You could argue that was "just" noticing a connection (between number theory and Fourier analysis) that nobody had noticed before; in fact this is the kind of thing AI is extremely good at. I'm being a little cute here.
I think if a human had proven just these two results in the form of a uniform zero-free region for L(s,chi) from nothing as OpenAI did it would not be unfair to say that it would be the single greatest advance in math (easily dwarfing Wiles' FLT), and it would instantly put them in the ranks of greatest mathematicians of all time. Unlike something like Navier Stokes there wasn't a semblance of a research program, experts basically considered this hopeless and would have said the chance of seeing a proof in our lifetime was near zero.
For some comparison, Yitang Zhang's bounded gaps result might have gotten him a Fields Medal if he was not disqualified by age. When it was floated that he might have proven Siegel zeros don't exist, it was considered (by experts) clearly a much bigger deal. This result blows that out of the water (it's a way better version); at least analytic number theorists I talked to thought it was plausible but unlikely that Siegel zeros would be eliminated in our lifetime but thought RH was basically hopeless.
While some of my work is in analytic number theory, much is in other subareas, so it is possible I should defer to you on this.
It seems to me less than PNT in terms of what can we actually do with this. Many different areas of math use PNT, and from my standpoint, PNT is helpful not just for what it implies directly but because it lets us make really good heuristics about whether some sets are infinite or not, and what their rough size is. (Granted, one can do that also mostly via Chebyshev). For those purposes, this doesn't really enter in. Similarly, PNT feels like a statement at least I can say explain to my mother without any technical details. This isn't that. But that may also be my own biases of wanting things to cash out to very concrete statements about the integers.
I agree that one striking element is how no one saw this coming. This isn't building on an existing research program, which itself is remarkable. And last night, before I went to bed, I saw a conversation between a bunch of analytic number theorists who seemed to think there was potentially some slack in the quasi-RH argument, which if that's the case means this is going to go even further.
Thank you for the detailed explanation. From what I'm reading from a lot of mathematicians there's at least a dozen of results here that are field-definining and worthy at minimum of a Fields medal.
I guess the biggest news are not the discoveries themselves but how they were found and that math is going through the biggest revolution as a field since almost ever.
> it would be the single greatest advance in math
Did you mean to not qualify that? That is a bold statement indeed.
1896 PNT is basically 1859 Riemann + a trig inequality.
1830 Dirichlet's result is qualitative only, it shows infinitude but not the asymptote in terms of the zeros for it predates Riemann.
To me this is the first substantial step after the 1896 PNT, and we really do not see much progress in the whole 20th century. Personally so far there are only two people worth mentioning,
- Euler, introduces the real zeta function and Euler product, establishes the functional equation at (half?) integers.
- Riemann, introduces complex analysis ideas to the zeta function.
And of course this result if it is true. This is first to penetrate the critical strip, which nobody had any idea how to approach for over a century and a half.
How about 100 years?
Yeah, completely reasonable to argue that.
What is your favourite unsolved problem in number theory which if solved, would be more important than 1896 prime number theorem?
Generalized Riemann hypothesis.
(unrelated: love your username)
Was anyone in the math community aware of the inbound tsunami at the beginning of the year?
Lots. To give an example Terrance Tao was lambasted skeptics on this site for stating it in 2024.
https://unlocked.microsoft.com/ai-anthology/terence-tao/
" I expect, say, 2026-level AI, when used properly, will be a trustworthy co-author in mathematical research, and in many other fields as well.
Then what? That depends not just on the technology, but on how existing human institutions and practices adapt. How will research journals change their publishing and referencing practices when entry-level math papers for AI-guided graduate students can now be generated in less than a day—and with the far better accuracy of future AI tools? How will our approach to graduate education change? Will we actively encourage and train our students to use these tools?
We are largely unprepared to address these questions. There will be shocking demonstrations of AI-assisted achievement and courageous experiments to incorporate them into our professional structures. But there will also be embarrassing mistakes, controversies, painful disruptions, heated debates, and hasty decisions."
He's pretty damn smart that guy.
Terrance, Reinmann and Hebert walks into a bar...
> He's pretty damn smart that guy. This is probably the understatement of the year. I am literally ROFLing.
I was at the workshop that resulted in the Leiden Declaration in Fall 2025. The majority vibe was that this was inevitable, but hard to predict whether it would be in one year or 30 years.
I guess they showed this to the advisory group they created. I guess the group tried reading the work for a day and they could only think of telling them to release the results to the community. I now understand why the group had this suggestion.
I predicted, over 2 years ago, that theorem proving would fall way before other problems that people believe are harder.
https://news.ycombinator.com/item?id=41072330
There was a Wired Magazine article from either the late 90s or early 2000s that made a prediction that this sort of thing would eventually be possible, likely within my lifetime. I believe the context was "distributed computing" models of the time, like SETI.
I've never been able to find that article as an adult, but I would love to know who wrote it.
Some candidates suggested to me by an AI:
Gina Kolata allegedly in the New York Times in 1996 on the Robbins conjecture (noting that computers had started to contribute to math research in some sense), and a longer piece in Math Horizons by her the following year ("Computer Math Proof Shows Reasoning Power"). I didn't immediately find the NYT article, so I don't know if it might be a hallucination.
John Horgan in Scientific American in 1993 (https://www.scientificamerican.com/article/the-death-of-proo...). There's also a retrospective on the topic by the same author in Scientific American in 2022 (https://www.scientificamerican.com/article/should-machines-r...).
Natalie Wolchover in Quanta (but reprinted in Wired) in 2013 (https://wired.com/2013/03/computers-and-math).
I was involved in some distributed computing stuff in the late 1990s and early 2000s and I don't really remember people in that community talking about proofs but there may have been a "if we had a mechanical proof-checker, could we do distributed searches for valid proofs that it would accept?" conversation somewhere at some point. There were definitely volunteer distributed computing projects working on pure math; I remember the Optimal Golomb Ruler search (https://en.wikipedia.org/wiki/Golomb_ruler). So, that could possibly have shaded over into "can we find proofs this way too?". At the time it probably would have been based on brute force searches through proof space rather than clever optimization, though.
The idea that you can lexicographically list all proofs in some formalism and then mechanically determine if any is valid is quite clear from Gödel's construction of the function Bew in "On Formally Undecidable Propositions", but he points out that you don't know where to stop because you don't know how long a valid proof would potentially have to be (so "is this a valid proof of this claim?" can be decided mechanically in a limited time, while "is there any valid proof of this claim?" can't be! maybe the shortest valid proof is 49 steps long but you eventually stopped checking after looking at all 7-step proofs, or something).
Thanks! Yea, I have used various LLMs to dig for this article, as well as Google search multiple times over the past 20 years. The article I'm remembering was 100% prior to Nvidia's CUDA release in 2007. My best guess is that it was from late 90s, but possibly early 2000s.
The article I'm remembering was not just about mathematics, but indeed all of physics and related fields. I believe it speculated that eventually distributed computing models could essentially take the world's mathematics and physics formulas and various datasets that we believe to be accurate with high degrees of confidence, and then look for patterns or trends, and then from those trends, mathematicians and physicists would be able to investigate further. Not dissimilar to Folding@Home and SETI@Home.
Keep in mind, that this is the best I can remember from 30 years ago, and I've thought about it so frequently that I am certainly misremembering some of the details. Anyways, it's always been this really compelling possibility, and I wish I could find that article that inspired me so long ago and re-read it! :) I really think it was Wired, but it's possible it was Popular Mechanics, or even an expert guest on TechTV who gave an interview. Hard to say for sure, but I've always thought it was a Wired article.
Appreciate your help though!
Oh, I remember hearing about "computer scientists" or something that would attempt to determine physical laws on the basis of empirical evidence, possibly also in that timeframe. That might be another thing to look for. I'm sure that's something people were writing about.
Edit: with the noun-noun compounding being different from the usual interpretation here, like "scientists who are computers" rather than "scientists who study computation"! Maybe "computerized scientists" or something.
Ted Kaczynski,the Unabomber, made the same prediction 30 years ago.
Fucking even called LEAN the “hottest shit under the sun”—which it is. You, legend you!
And do any of them actually matter? Will the fact that Noodleheinz's Third Postulate now has a proof affect anyone?
It's really impossible to predict which discoveries will "matter", have a direct impact on other fields, or an impact in making other mathematics or physics discoveries.
Only after a world's worth of experts look at these results and then mull over if and how their own fields are impacted by this new info will we be able to answer this question.
I'm reminded of a great TV Show, James Burke's Connections. Where discoveries in one area of science would revolutionize or fundamentally change a completely different area. https://www.youtube.com/watch?v=XetplHcM7aQ&list=PL5HjoPOFFC...
It can take decades to really know the full significance. You know, the whole "We stand on the shoulders of Giants", well the Giants just grew a few inches all at once.
Dear lord that website is laggy
At this rate solving P=NP is going to be easier than solving front end perf …
wait, maybe this is the same problem....
with non-polynomial side being represented as the frontend programmer's constant need for more performance to do the same task...
worth mentioning that "NP" is not "non-polynomial" but "non-deterministic polynomial (time)". If NP was non-polynomial time then NP != P would be trivial (and in fact, P != EXP is known by the time hierarchy theorem).
Non-deterministic can be explained in several ways. One is in terms of a hypothetical "nondeterministic Turing machine" with certain non-physically realizable properties. The easier way is that a NP problem gets as input not only the problem instance x, but a "witness" w, that may depend on the problem instance. This witness generally makes the problem of deciding the problem instance straightforward (e.g. for SAT, x is the SAT instance, and w is a description of how to set the variables so that it is true).
I reckon I could tell you in polynomial time whether a div was vertically centered, not sure if I could write the CSS in polynomial time.
Please let P=NP, Please let P=NP
Whomever is running this simulation, please.
It's math, the result shouldn't be different just because it's a different sim.
To be fair - there are statements in math that are independent of the axioms. For those statements, the universe you find yourself in can pick either version (true OR false) and still be consistent.
See also: noneuclidian geometry and axiom of choice.
Well if the fundamental constants or hidden variables of the universe are shifting because of his comment then it can change the outcome.
Depends how fundamental the variables are. If we can code a sim for a topos[1], why can’t we be in such a sim?
1. https://arxiv.org/pdf/1012.5647
unless mechanism behind our universe dynamically alters our logic on the fly to be artificially self-consistent
Why?
Being able to solve NP hard optimization problems would enable progress in many areas of science and technology. For example it would allow us to find poly-sized Lean proofs for theorems efficiently, since proof verification can be done in polynomial time.
It would also be amusing to annihilate nearly six decades of proofs that assume P!=NP.
Leans proof checker is not polynomial time, unfortunately. It is super exponential. Basically, because it can verify the result of any function it can prove to be total.
That's fine, we just change the problem from "find a lean proof of length < f(n)" to "find a lean proof that can be validated in time < f(n)".
Oh that’s unfortunate.
Could also break the basic principles underlying most encryption approaches. I would rather have my bank account not stolen and internet working
to depress you even more, it is consistent with everything that we know that P != NP and that cryptography does not exist. So there is a worst of both worlds, and we cannot rule it out.
See https://blog.computationalcomplexity.org/2004/06/impagliazzo...
I've had enough Internet for one lifetime.
As long as we also get low order polynomial solutions to important problems, it'll be worth it.
Besides, unencrypted wifi was funny.
Even if P=NP it doesn't mean that the P approach will be better than the heuristic approach we already do today.
Of course, if we get ridiculous polynomials it doesn't mean much in practice. People who hope for P=NP generally hope for nice polynomials O(n^3) or something like that at worst.
Not to mention it's got that signature Claude Clutter UI design
Except that Claude wasn't used.
Probably Copilot then
Interesting how perceptions differ. My first thought was “Wow, that’s well designed for a math website”.
Shameless self-plug: I created an autonomous math researcher. It already solved a 12 year open problem in proof complexity which lead to a publication (and proof complexity experts are already working on simplifications and generalizations of the proof, as I've been told by one of them). This publication is an important step in Cook-Reckhow program in answering the NP vs coNP question.
The autonomous researcher records every research cycle in a public notebook.
Framework: https://github.com/kbr-/math-research/ Public notebook: kbr.is-a.dev/math-research/
The publication: https://arxiv.org/abs/2609.23015
Is this not just a pre-print, not a publication?
Right. I thought preprints are a subset of publications. Yes it's a preprint.
"Pre-print" implies it's headed to be "printed" by a publisher. That is, the paper has already been accepted for publication by a peer-reviewed publisher, and it's just being posted early for wider and faster dissemination or to stake a claim of priority.
Otherwise, the document is a self-published manuscript, which doesn't carry the authority implied by "publication" or even "pre-print".
Let's hypothetically say I'm a PHD student who is half way through my studies and I have a halfway written version of one of these "preprints" - what do I do?
Seems like a lot of PHD students are doing to have to pivot the entire structure of their PHD studies? Or just produce something which is already written by OpenAI?
As others said, it's not that this isn't a phenomenon that is unique to now. It happens. I had to pivot a bit of my dissertation near the end because at a random conference I spoke with a researcher from another continent and realized one of my ideas was already out there in some form. I just missed it since it was in a conference proceedings outside the usual set I looked at. So, I had to scramble to adjust and still come up with something novel. I survived, and defended, but it wasn't that much fun at the time.
What isn't so normal is the probability and ease by which this kind of thing can happen today versus decades ago when I was in school. As OpenAI said, it only takes a few hours of compute to do what likely was much more than a few hours of human effort. The only reason this kind of scooping/overlapping was rare was mostly a function of how fast other humans could do the same work. With machines, that totally changes the relative pacing between the human trying to learn how to be a researcher and the machine that can grind out results.
I'm less worried about the phenomenon of overlap and scooping and such. I'm more worried about the long-term impact on fields (not just math), especially considering the early stage students and researchers entering the pipeline now. I'm not sure what happens to disciplines when that pipeline stalls.
math will just be black boxed away. no one will "need" to understand it.
Let me give you some perspective: my entire phd was made obsolete by CRISPR. It was a wonderful thing.
Interesting. If you don't mind, could you please share a little bit about that ? You already finished your dissertation ? It was about the works of Doudna and Charpentier ?
No, back in the late 90s and early 00s, people were trying to engineer custom nucleases and transcription factors, my work was on doing molecular dynamics simulations to optimize TF sequence specificity (similar to engineered zinc fingers) for gene therapy. I wrote up my dissertation and published it in 2001, and then went off to find enough compute, IO, and smart people to make it happen (https://research.google/blog/groundbreaking-simulations-by-g...).
My approach would require custom engineering for every different sequence we'd want to target. With CRISPR, you just "program" the system with a guide sequence, you don't need to do massive engineering to solve a protein design problem.
So it's not the situation they described at all, you didn't waste time during the PhD having to scramble to change topics as it happened after you were done.
That, and also it's just a completely different approach which might later on turn out to be useful. People should remember that artificial neural networks were developed decades before they were useful. People were doing all kinds of other approaches to ML like support vector machines before advances in hardware made deep neural nets feasible and therefore interesting again. ANNs were never obsoleted by SVMs.
Actually, I'm pretty sure SVMs were obsoleted with ANNs (not just in terms of UFFs).
For every door that shuts another one opens - great if you're not a cabinet-maker.
I have had this conversation with my PhD students yesterday. I am 100% sure that all of their problems can be solved by publicly-available models now (I solved a case of one myself as a test, it took 15 minutes). So the challenge for them is to see how much they can accomplish in their allotted period, and still pass a defence on at the end of it all. The PhD defence is going to become all about a test of understanding, not a test of quantity of publication.
A much bigger issue is: What is the point of any research mathematician publishing anything now? I really hope that one positive effect of all this will be to finally topple the awful peer review model we currently have, with the biggest publishers gatekeeping with extortionate fees.
This happens all the time, even without AI. Other researchers or PhD students can publish the same results before you. I say that based on my experience during my PhD.
Yes and no. It's pretty normal for students to get scooped, sometimes even multiple times (ask me how I know...).
It's much less common for students to have their entire thesis direction removed from under them, as might be the case for someone working in fine-grained complexity assuming 3SUM has no subquadratic algorithms, or working assuming ~UGC. Both of which (publicly) seemed like perfectly valid research directions until a couple hours ago.
There might be stuff to salvage from their conditional results anyway, but this is not your average scooping.
Use whatever is published as the new base for your research. Use AI tools to help you going forward.
In other words, you've already taken a gambit with the first half of your PhD, now take a second gambit, praying that you have something to publish by the end of your PhD.
It has to feel awful to be in this position.
Could be worse, imagine having years of experience in a profession these things can now handle by themselves.
:)
It's worse for those who will face the job market (or maybe tenure review) in the next 2–3 years. If you have more time before you have to justify your continued employment, you can pivot to something else. Theorems and proofs may be cheap now, but there is a lot of value in figuring out what is worth studying.
This has always been a challenge for PhD students and researchers, it's just far more likely to occur now it seems. Getting scooped doesn't feel good, but it's a signal you've been thinking about things other people care about.
I think eventually companies won't get as much stuff that's usable for marketing, so they'll stop investing so much into ai for math, so eventually cheap and poor graduate students will be able to do relevant work again without worrying about getting scooped by a company with a million GPUs :-/
If you had halfway to one of these papers you would be anyhow be in the top echelons of math phds so probably you have less to worry than most!
Perhaps consider switching to a more applied field. Experiments in the physical realm hold value, especially if you document the process.
The purpose of a Phd is to train researchers, not to solve one small problem.
Some math PhDs would spend a year or so doing things with AI and lean, and graduate. And keep doing more math afterwards.
Some others, with more stubborn advisors, will keep trying to find a gap where there's no AI progress.
CS subfields go through this every ten or so years.
The follow up question then naturally would be how do phd advisors with people whose fields are in someway premised on making breakthroughs in theoretical fields that AI can solve work deal with it?
At my (German) university, the solution would have been to change from a "cumulative" thesis (which requires peer-reviewed papers) to a monograph-style one. Because the general rule was that your undertaking must be novel ar the time you submit your thesis topic, not necessarily at the tome the thesis itself is submitted (years later).
> what do I do?
Very hard question.
Your work makes you one of the very few people who really understands the problem and solution and its significance.
It could still be interesting if your approach to the problem was different to theirs.
Don't worry, if you use openAI and get lucky they might offer to share credit with you for your work.
Don’t paste your research into these AI because they will train on it and then scoop you.
I think that happened after word of the project reached OpenAI and they allocated resources to it.
> what do I do?
Precisely what all NLP researchers and the ML community at large did in the last few years: embrace the frontier and realize that attention is all you need.
Verification is equally important, if not more so.
Ask OpenAI for money when you don't have a job or future?
I guess the only answer is to adapt with the tools. If we can't do that, then yeah, we're in trouble.
That’s always been a thing, it’s called “getting scooped”
Killing with knives has always been a thing. Now, we have the machine gun.
The scoop gun
Frankly, I believe PhDs don't have to be novel in the absolute sense, only new research to the student. I can't remember how many times I've effectively invented something from a clean room approach only to realize someone already published something years ago or that my "new" algorithm has a name. So this really shouldn't be holding anyone back.
That would be unfortunate but the world does not owe you anything and is not going to stop for you. Which is also a valuable thing to learn in your 20s (ideally earlier).
Look forward to being obsolete I guess
Yes, and?
> Let's hypothetically say I'm a PHD student who is half way through my studies and I have a halfway written version of one of these "preprints" - what do I do?
I did get my PhD...before AI. And my honest advice (to myself back then, even) would be: quit the PhD, become an electrician, and work hard to buy a tiny house in the middle of nowhere to watch the world burn in this madness.
You might enjoy https://asteriskmag.com/issues/15/so-you-think-you-could-be-... if you didn't see it recently.
These results are wild. Several individual findings are crazy good and use mostly unexplored methods (the improvement over Riemann for example)... I'm pretty sure some of these results would have been Fields-worthy.
But having so many of them at once? Damn. We really live in the future.
Imagine you get up one morning and most open questions in math are solved lol.
> Imagine you get up one morning and most open questions in math are solved lol.
More time available for mini-golf?
Sounds like that's going to be next week no ?
I guess there would be new interesting problems emerging, & AI would solve them, until it's completely impossible for us to understand
And who is going to decide which of these emerging problems are interesting?
well hopefully the world has also advanced enough in other ways lol
imagine getting up and math is solved, but you still have to deal with bullshit lol
This news is exciting and sad at the same time. I've heard that AI chess programs sometimes have blind spots or quirks that human players don't. I've also heard that human players are learning from AI's playing styles (essentially human and AI evolving together). Maybe something similar will happen in mathematics.
Extremely unfortunate that gate keeping got to the point where they felt the need to ask for permission to share math.
After how poorly OpenAI and Anthropic handled the previous cases, I approve of the more measured and cautious approach this time.
We cannot have them rushing to publish amidst tons of confusion, rumors of threats/scooping and outright plagiarism of existing work (by failing to cite said work).
If they're going to participate as scientists in these more rigorous fields, they're going to have to match that level of rigor, not lower it to the disastrous low that ML research publication is at.
It was OpenAI that fucked up the Navier Stokes explosion situation, not Anthropic
I'm not referring to just that. For example, Anthropic published a half done report about some biology research that turned out to have already been discovered and patented, and IIRC both companies are guilty of claiming results without doing the basic diligence of citing the material their work builds on, effectively passing it off as entirely done by their AIs.
In previous math sharing there was speculation about stealing human researcher's results or progress, via prompt inputs from those researchers, and sharing that as their own result. They're adjusting their process, and it seems ok.
Let’s be very clear, the alleged “stolen results” were largely the product of another LLM, not de novo human work. Also, it was false - they did not steal the results.
No clue why I'm downvoted for this, HN struggles with truth-seeking on these topics.
> HN struggles with truth-seeking on these topics.
The facts are much more nuanced than how you're presenting them here.
What’s the nuance? Two researchers alleged theft. OpenAI investigated and confirmed that their research was not in the model’s training data. The world decided to take the first part as objective truth and ignored the second.
> OpenAI investigated and confirmed that their research was not in the model’s training data.
You mean when they say it was impossible to confirm anything but one day after it was 100% confirmed that there was no theft? And here I'm not even talking about all the ethical problems related to trying to scoop another group when you hear they're close to success, or how current solutions follow extremely closely human-generated ideas, or about the lack of relevant citations in OpenAI's paper.
Believing that OpenAI's claims have any substance cannot be explained by naivety alone.
They didn't try to scoop another group; they thought the other group had already solved it, so maybe their latest model could take a shot too... and the model solved the full problem when the other group had not! They found out after the fact that the other group had only solved an important sub-problem.
This was a low-key hilarious replay of George Dantzig and his homework problems: https://en.wikipedia.org/wiki/George_Dantzig
The controversy was whether they had plagiarised that other work on the sub-problem, which they categorically denied after an investigation. And yes, a few days to investigate something like this is reasonable for a company as big as OpenAI. Having seen how data infra is set up when petabytes of data are flowing about, there are thousands of entwined data pipelines to figure out. Not quite as easy as running a query on a sqlite DB!
> Believing that OpenAI's claims have any substance cannot be explained by naivety alone.
Yes, they could be explained by a GitHub repo full of proofs :-)
Or are you suggesting there were hundreds of researchers who just happened to be close to solving hundreds of these long standing open problems using Codex, and OpenAI swooped in plagiarized them all? ;-)
glad you could present those facts in your comment, i know the HN character limit can make it hard
Can you show where it was discounted? Last I heard OpenAI was declining to deny it, presumably while they thoroughly confirmed.
> “Following an investigation, we have confirmed that Buckmaster’s Codex prompts over the two months preceding this announcement and paper on September 8, 2026, could not have influenced the system in any way, including through training.”
https://openai.com/index/navier-stokes-solution/
Come on, they were working on this for more than 2 months. Dont fall for the corporate half truths.
Buckmaster himself said essentially all the progress they made was from July 15th onwards (with a new model on the problem) with no real progress prior to then.
I think it’s really cope to claim this was a human result being stolen.
I really don't get how people are still continuing with this "stolen results" narrative after today. Like NS was kinda insignificant compared to treasure trove they released now, thinking that the LLM needs to "steal" from some human is simply coping.
After reading your comment one could even think that LLM's invented math from the ground up.
HackerNews does love a good conspiracy theory.
Just to be very clear: they aren't asking for permission, they are framing it that way because of the bad press.
There's no gatekeeping here!
Right? Math is the most open of our academic knowledge institutions, by virtue of what it is. It's easy to get any math publication, and I am not aware of any other fields where an anonymous person can publish their work informally in an anime discussion and enter the annals of math knowledge.
This is a tricky one, but I do sort of agree with you. However OpenAI doesn’t give most researchers access to the models which produced the work. So the gatekeeping goes both ways, I think.
They didn't even follow the recommendations of the reference group. A few of them maybe, but this is still a dump of llm-written papers.
The gate keeping, I'm afraid, will be now in the hands of various bubecks, responding directly to even more disgusting people.
With all the hierarchy present in mathematics, I would prefer it by far.
This thing named inappropriately "OpenAI" goal is just grabbing and monopolizing. Capitalists before could not really touch the deep of the human spirit with their filth, now they can.
I'm a software engineering/biology/ML guy who loves when clever math ideas get turned into real solutions (https://en.wikipedia.org/wiki/Compressed_sensing). I am curious if any of the results have immediate applications in any kind of engineering or science.
It's fine if not, but it'd be great if even just one of these helped us solve a long-running problem.
I was talking with a number of physicists this evening, and of the papers who's problem statement we understood, I don't think any of them will have near-term applications. For most of them, they were results that I think the physics community believed to be true, and the paper provides the first rigorous proof. For several it was news to me that they weren't already theorems! These results are important advances in mathematical physics, but they don't tend to have much impact on experiment.
From what I can tell, all of the physics results here are quite mathematical. But I am very curious how the internal model they used would perform on more applied problems.
It feels odd to me that they wouldn't prefer applied problems. Seems like an easy way to profitability. Probably based on what attributes they're looking for in a problem when picking them.
At first this was my take as well. That plus, well, maybe they are just on a serious PR kick with maths. But I am starting wonder if they have determined, or strongly suspect, that the road to exponential model improvement must first be paved with extraordinary improvements in math. Like in some sense this seems like a test case for where their true intensions might go: vast improvements in the efficiency / size / speed of models and their training. Hard to imagine trusting the models in all those spaces without first trusting them / training them to address new or unsolved math.
You can't really profit from proving theorems of applied problems (that are widely regarded to be true). Those who need to apply those theorems on real problems would have already done so (and if they don't work in some cases, well, congratulations... you found the counter example!)
Same. I have agents analyzing the papers here to see which of these are actually new approaches, novel application of unusual approaches, etc. If there are new intuitions and ideas, those are the mostly powerful reusable components by my estimation.
Please share your findings.
https://github.com/openai/math
I think they should put human names on the papers as someone who has reviewed the result, even if just a preliminary review. (I’m assuming they didn’t just pipe their model output directly to the internet and these had some amount of review?)
Presumably they don't because they're training the audience (us) to trust the machine, not its verifiers, even if they were included.
I appreciate you acknowledging the politics behind this. OpenAI does not intend to be a software company for long, they intend to be scientific infrastructure. They'll want to be faucet from which pours embryos, orbital calculation, geothermal/substructural rating, and etc.
> they intend to be scientific infrastructure.
beyond naive.
Looking at the quality of the writing in these documents (or at least, the ones that relate to problems I've worked on), I do not think humans reviewed them.
And that means that the recommendations of the reference group has not been followed. The only improvement here is that it is version tracked? No authors, no careful write-ups with exposition.
the reference group can recommend all they want, no one will review 700 plus papers.
There are competitive reasons that they don't want to share all the people on the team.
I think it's a damned if they do, damned if they don't scenario. I think they would be afraid to seem like they're claiming that the selected scientists deserve the praise of an invention or something, when it's "just the machine who did the work"
Yeah, that's probably right. But, it's a "new world" - invent a new standard - "Written by Claude Foo X.Y; initial review by John Doe".
With a deluge of results, having some human expert vouch that it even might be worthwhile would help. (e.g. see the link on HN yesterday, "Two Room-Temperature Antiferromagnetic Semiconductor Candidates" - I see it not worth looking at unless a subject expert vouches for it).
> I think they should put human names on the papers as someone who has reviewed the result
Assume the empty list you see is complete. :)
Long term this will be the only reasonable author list: Chad G. Peter {1}, Mat H. Lean {2}.
1: Author 2: Verifier
/s
From the GitHub description: “On average, each result used 3 hours of ChatGPT Pro thinking compute”
I'd really like some clarity on what that means. E.g. DeepMind has 'cheated' with this in the past, claiming that AlphaZero only took 4 hours to reach super-human chess levels while conveniently leaving out the fact that it was 4 hours x 5000+ TPUs. Sure it's impressive that it only took 4 hours wall-clock but it's very misleading as to cost.
Can we get a number in Blackwell GPU-hours, kWh, or some other compute-scaled metric?
Depends what your metrics are. If you suddenly found a way to have 9 women make a baby in one month that is huge.
I'm not denying that, but I'd still like to know what that cost.
They did say that. "3 hours of ChatGPT Pro thinking compute"
Yes, what does that mean?
It means the level of effort that a ChatGPT Pro plan summons when thinking. For 3 hours.
If folks are going to analyze this claim with a critical eye, I'd be zeroing in on "average" rather than acting like this measurement is somehow unclear.
>OpenAI has 'cheated' with this in the past, claiming that AlphaZero [...]
That doesn't sound right
Oops, edited.
Oh cool, we will all now have a math genius on our computer.
It was using their internal math model, so not yet for us
I used future tense. It was implied this will be available.
Well, on their computers. But you can rent them for a price.
An open source model will reproduce it 6 months later
which you an run IF you have the hardware. who knows how heavy these models are.
Does this imply that it was a one shot prompt with ChatGPT Pro style models (i.e. best-of-N), rather than the agent swarm approach that was used for Navier-Stokes?
It doesn't imply that, it's just measuring the amount of compute.
But if it was an agent swarm wouldn't we expect something like 300k hours of ChatGPT pro compute equivalent instead of 3?
Not necessarily, could be agent swarm with low N
At some point that stops being a "swarm" and just a handful of subagents. E.g. 6 agents running for 30 minutes each isn't really a swarm in my eyes.
That estimate is obviously going to conveniently ignore all the failed runs.
I'm waiting for someone to come along and finally prove that P != NP ...
It’s incredible and wonderful. Mathematicians in this thread sound very much like software engineers last year, who spent years wrestling with a piece of code and now it just “appears”! But think of the next level that it empowers: new mathematics, new physics, new forms of advanced engineering. What was formerly constricted and throttled fell and a new wide vista is possibilities opened up.
> new mathematics, new physics, new forms of advanced engineering.
the worry is that majority devs/mathematicians will be irrelevant to this new forms.
Title: Hilbert's Dream, Tim Gowers - LMS Popular Lectures 2012
https://www.youtube.com/watch?v=k_ordDFw588&t=3597s
Audience member: (1:00:00 - 1:00:09):
so you said that if there were such a program that could you know provide a proof or disproof then mathematicians will be out of business what really, I mean that you think it would be liberating
Tim Gowers (1:00:10 - 1:01:14):
well that's a very interesting question actually if there were a program that could solve the kinds of problems that we spend our time solving and do it much more quickly than we could then we would be out of what comes with what currently constitutes business but we would it's not completely inconceivable that we could just say we've got this fabulous tool now what are we going to use it for and it's a little bit I don't know I'd want to sort of plant aside what would we do if we had a program that could just answer any mathematical question you gave it to or else if it failed you'd be pretty confident that nobody was ever going to solve it and certainly a lot of applied maths might be pretty pleased with with something like that so what I really mean is that I could just modify what I said and just say it would radically change what mathematicians do or what pure mathematicians do
So happy this is shared on GitHub rather than some gatekeeping paid journal. Truly a new age for science.
Most mathematical results are shared on Arxiv. Journals add peer review.
End of an age for journals?
In a sense, because now journals will be flooded by LLM slop.
GitHub is a significantly worse place to store important results than Arxiv. Of course, slop does not belong on Arxiv, buy slop should also not get published.
Unique games conjecture and matmul <= 2.25. What the hell.
Yeah, I am kind of freaking out at some of these. I called a math friend to bring me down to Earth and he is freaking out even harder.
FFT BELOW NLOGN
Specifically, (n log n)^{1 - 10^{-13}}). There are likely no practical industrial problems of a size that would benefit from that specific reduction, but just breaking the nlogn barrier suggests that there may be more and better fruit here in the future.
SUBSET SUM AT 0.49 WTF
matmul <= 2.25 is no big surprise tbh. unique games and mul < n log n are much bigger.
Is anyone verifying these? And then, as a next step, how can they get a voice?
What the hell, indeed.
The commenters over here think that OpenAI basically ignored AGMAI:
https://proofsandprompts.com/2026/10/07/on-openais-release-o...
The thieves do as they please, funded by money stolen from the public via inflation and possible future bailouts.
As social commentary, I think a lot of people in this thread are expressing interest in and engaging with this level of math who might not have pre-AI.
I bet, for people who don't understand these problems or their solutions but are close and are now interested, AI makes then considerably more accessible than they would've been previously, and behind this big visible wave of results there actually will be (or already is) a wave of improved comprehension by a lot of curious people.
I'm not at all at this level at all, but I did learn quite a bit about polynomials over fields yesterday.
Happy to hear that yes! I hope more people do learn about math from this !
Relevant:
"As AI Closed In on ‘Unique Games’ Proof, Researchers Raced to Beat the Machines"
https://www.quantamagazine.org/as-ai-closed-in-on-unique-gam...
Actual results: https://github.com/openai/math/blob/main/overview.pdf
HTML version, https://github.com/openai/math/blob/main/CONTENTS.md
Someone in these comments said the paper on Barnette's Conjecture is short. Are any other proofs in this collection short and/or understandable?
It may be useful to publish a formalization of ALL known mathematics at this point. Like every book ever printed, every paper on arXiv etc.
How many Gigabytes would that be, compressed? Wikipedia once fit on a DVD
This might also allow for some interesting meta-mathematics
This is what they are trying to do with Lean
More specifically, a combination of mathlib (human, expert curated) and projects like TauCeti (AI-welcome complement to mathlib). See: https://github.com/TauCetiProject/TauCeti
Oh? Where can i read more about that? It appears the sole focus so far was solving open problems
I believe that is the goal of MathLib, to transcribe all math into a big Lean library.
https://lean-lang.org/use-cases/mathlib/
Mathlib is expert reviewed, but only contains a tiny fraction of all mathematics. So this seems to be a quantity of work a "10,000 agents" approach would be applicable to. Like Navier-Stokes, something to spend a couple million in compute on :)
I think such "process-oriented" fields shouldn't be subjected to mass automation (or at least not till humans have sufficiently leveled up our game and understanding). What I mean is progress in math comes from having gained a deeper understanding of the problem for subseq
Literal thought control.
Butlerian jihad, more like
ok Terence Tao, calm down.
Two weeks ago, I heard rumblings that UGC, RL=L, and a matrix multiplication lower bound (even lower than the one here, but so unbelievably lower I think it was a typo), and a few others were about to be proven by someone at OpenAI. I find myself saying "big if true" a lot lately.
Even those on their own were enough to make your head spin. But seeing about 100x that? Geeze.
The math equivalent of AI code reviews piling up. I wonder where the incentives will align and the equilibrium turns out.
Not much computational mathematics here, but I do see some of interest in the MCMC community. In particular, 139, 93, and 101 are very interesting. Also (I'm not too familiar), I remember attending some talks on the Crouzeix conjecture 325 attacks, should sharpen some rates for Krylov methods. Obviously need to read deeper, some of these don't have corresponding Lean formalizations.
Cool!
what does MCMC mean?
MCMC = Markov Chain Monte Carlo
It's a way to approximately draw samples from a probability distribution. Crucially, it applies even when we only know the distribution up to a multiplicative constant which is a common ailment of many distributions in the computational uncertainty quantification field (not that we don't know the constant, but that it's usually computationally catastrophic to estimate it well).
Wow, FFT faster than O(nlog(n))? I wonder if that will open the floodgates for further improvement or not. I don't understand anything about most of the fields these results touch, but I can say that this in particular is very surprising.
It’s a meaningless improvement
The applied mathematician’s joke is that log n is bounded above by 45 or so.
My guess is that the constant terms are large enough it's not practically useful in most cases.
It’s always stated that open weight models are 6 months - 12 months behind. Therefore, do we expect that in a year open weight models will be as good as OpenAI’s internal model at theoretical math, or does OpenAI have some “magic” that will be much harder to replicate for competitors?
Past performance is not indicative of future results.
In some ways, this feels more like an ominous warning about the times to come than something to celebrate.
Catalan's constant is irrational!!!
That is a big one. Exciting times to be alive. Regrettably I can't understand the proof at this point.
There was a (flawed) proof submitted a month back:
https://arxiv.org/abs/2609.04176
I wonder if it gave part of the inspiration.
Come on, these are easy. You assume it can be expressed as p/q where p,q are integers such that GCD(p,q)=1. Then you derive a contradiction.
/s
TBH, I'm both happy and very sad I didn't do a PhD in math (or at all).
Glad the papers are out. Hope researches get their hands on the model soon too, so they can ask follow-up questions and try their own ideas.
Do we have any field where humans have ray of hope to use their cognitive abilities in LLM era?
Why should such hope exist? When the motor vehicles were invented, humans have lost the speed race completely. Yet, nobody lamented and people still did foot racing for fun.
So will it be with AI tools. If these tools become so good, then it will be used. People who want to exercise their minds can still do so, even if that cannot produce economic value.
Foot racing was never a broad source of income for humans. There is a massive difference. The question will be: what economic value can humans produce in the future with AI tools? There is a chance that the answer is "not much". That is the scary part.
brain always take the lazy path. that is the issue.
The superintelligence moment happened for chess in 1997 and today more people are using their cognitive abilities in chess than ever before.
Looking back through history, there has never been a breakthrough like this, one that impacts virtually every known industry simultaneously with the potential for a 10x impact.
Any field where there is no external validator (formal proofs, compiler, rule sets) for AI to leverage and real world evaluation isn't easily digitalized. Add in a bit of physical interaction and it is done.
unless they start connecting with hardware and machines. Not every use case can be done but many will be solved.
there is also coming with new problems and asking the right questions
When all open questions are answered, what happens next? Will the machine stop until humans fully understand everything and come up with new questions? Are there examples already of AI solving a problem we didn't know existed?
Clearly AI can grind on math problems now. It can generate proofs and get immediate feedback.
But what I wonder: can we legitimately grind on Physics or curing cancer? There’s a lot of physical world experimentation that needs to happen to make progress.
There's a kind of sticking point in physics combining general relativity with quantum mechanics in a mathematically consistent way. It hasn't been done yet and is mostly maths so your math grinders could have a go. It's maybe the area I'm most interested in with mathematical AI. I've long had a hunch things are stuck there because the math is a bit hard for human brains.
>But what I wonder: can we legitimately grind on Physics or curing cancer?
If we can enable a feedback loop, yes absolutely. But feedback loops for things in the physical world like these are measured in months per cycle usually.
Verified Riemann Zeta in Lean: https://github.com/davegoldblatt/openai-zeta-proof-check
What is this meant to do? You're just showing that OpenAI didnt post a Lean proof that Lean/nanoda doesn't really accept?
Physics could be next. "It’s not that I’m so smart, it’s just that I stay with problems longer" AE
Physics is where I get excited!
> 241. Rigidity of the Turing degrees. Every order automorphism of the Turing degrees is the identity.
Wow. This is just crazy.
Can you elaborate/give context? I haven't heard of this problem before, but curious to hear from someone who has.
I am curious to know whether the proofs given by the LLMs are going to provide new hints about related problems that might be solved using the same machinery as the one used in the proofs. Also, I would like to know what is the average ratio between the length of LLMs proofs and the length of a proof that a mathematician can write to explain that proof to another mathematician. It is like a functor between the category of human mathematical concepts and the category of LLM operational concepts used in those proofs.
I think Alan Turing would be quite intrigued by these developments, and maybe wondering what took so long.
He said computers might pass his test in 50 years so we've run 50% over.
The craziest part about this is that it will disappear from the HN homepage in a day or two.
I read somewhere that if we encountered aliens with lot more advanced technology but if they don't speak our langauge, their tech would be useless to us. for example, human body is extraordinary technology that has alwasy existed with us, but we still don't understand it." the the fact that we have created intelligence that can do what nature does and can also speak our language is most awesome.
So they set up a whole "Advisory Group on Mathematics and Artificial Intelligence", filled it with renowned academics, promised to listen to the group on "review and communication of emerging results"...and then a week later went nah, we are just going to dump it all on github. OpenAI truly is a special company - in the best and worst way.
Is AI going to put Mathematicians out of a job, or is it going to create many new jobs reviewing proofs it creates?
I'm not sure it's clear right now.
I'm not a very good mathematician, but I do know a fair bit about software engineering, and with AI I've been busier than ever. I probably wouldn't be so busy if AI was better at anticipating what I actually wanted rather than making guesses no human would ever make.
This is just a short term problem though. Eventually AI will get pretty good at figuring out exactly I want and it will build that from the start. The requirement of me reviewing the AI output only lasts as long as models stay bad at anticipating my needs, which I don't think will take too much longer.
There's a hidden problem you skipped over. The model may get good at anticipating your needs, but are you good enough at anticipating your needs?
> reviewing proofs it creates?
If you mean reviewing for correctness then no, a Lean proof is a much stronger guarantee than anything that can be provided by any human.
For someone who's goal in math was taking unsolved problems and working on them then it's probably over. Just like in software engineering writing code by hand is kinda over.
Can someone with a math background explain the significance of these and previous problems that have been solved by AI?
Does some real problem get solved in physics, chemistry, biology, materials, etc? Or are these fun puzzles for mathematicians with not much real world impact? Eg solving the 8 queen problem in leetcode.
Many of these latest results are very important to theoretical math. Some are also theoretical physics and CS.
Real-world applications are far off, but developing mathematical understanding does tend to leak over into applied physics and CS.
A cynic might say this is all just intellectual games, and though there's a grain of truth, it's too cynical imho. This isn't like 8 queens where there's no hope for applications or generalizations. A lot of this stuff fundamentally affects our understanding of how numbers and systems behave, what are the limits of computation, etc.
Even if someone doesn't care about theoretical results, it's still exciting that AI has become superhuman in a domain as broad as math. That shows there's potential to be superhuman in other domains as well.
And I'm sure none of it was stolen from the actual researchers...
All such work builds on the work of others. Hopefully they'll get credited.
Sure I just have a bad taste in my mouth after those researchers were working on one of the millennium problems for a while, had used ChatGPT for assistance and then OpenAI claimed they had solved it.
will this make the math for building data centers work?
No that’s gonna happen when they take your job
Underrated joke
Solving hard mathematical problems is a strong signal for capability but would it not be preferable to through all these resources to urgent existential issues such as climate change? Finding technological solutions in those areas would be the ultimate capability signal as political alignment at global scale is almost certainly impossible.
They want you to pay for OpenAI credits to try and solve these urgent existential issues. That's the purpose of them creating this hype.
What's stopping you?
Raise some funding, shouldn't be difficult if you can convince people it's urgent enough.
AI only makes the problem worse, but the message that will be pushed is that it will solve climate change.
dna engineering for enhanced carbon capture trees seems in reach. othere less fun stuff could be as well with that technology
Climate change is a political problem, not a technical problem. Ironically, by raising oil prices, Trump might have done more against climate change than many people who have been actively fighting it.
Mathematical problems are ideal as benchmarks for AI because they have clear problem statements, clear axioms and results that can be verified easily (for lean proofs). I can't blame these companies for using them, although it's unfortunate that human mathematicians seem to become early casualties of AI progress.
I suppose openAI could have focussed their efforts on a subset of open problems that have a clear real world impact and leave aside the more esoteric open problems as a way for human mathematicians to hone their skillset. However, this would have been a short term bandaid. With open models 6 months behind the frontier, any of these problems might have fallen to the homebrewed efforts of enthusiasts early next year.
What is mathematics for? From the outside looking in (I'm a biologist), I have always viewed mathematics as a way to understand reality and to improve our ability to manipulate it. But what I often hear is that mathematics is foremost about human understanding. But isn't that only because it's humans that needed to do the mathematics in the first place? It's not obvious to me that mathematics without human understanding has no value. For example, it might be that P=NP. The algorithms are handed down to us and we can apply them without fundamentally understanding why P=NP.
Mathematics seems to be entering an era where human + machine maximizes performance, much like chess in the 1990s. However, imagine a future where even talented mathematicians are nothing but noise in the machine (as is the case in chess now). A future where AI generates and verifies proofs without humans in the loop. Where mathematics may be beyond human comprehension.
In that future, does it matter that early career mathematicians are inhibited by these developments? Perhaps not. Programming faces the same issue. As AI crawls up the competence ladder, does it matter that fewer people have opportunities to develop the skillset of a senior engineer? Perhaps not. I have no doubt that biologists will face the same problem soon enough.
As models get better, I think a time will come when it's hard for people to even verify the results. In the end, I think the bottleneck will be people.
AI doomers often talk about these kinds of scenarios often, but we tend to assume if humans were given magical math/physics results which we couldn't understand, or given magical pills by AI that cure all disease, we'd probably just take whatever the AI has given us rather than spend years or decades trying to understand the knowledge/technology before leveraging it.
At some point in complexity – especially if we allow our own knowledge to deteriorate because AI can do the hard work – we will stop understanding the world around us. In the same way one day Native Americans woke up and realised they shared the Earth with people who had magic sticks which they could point at someone and kill them, we will live in a similar world very soon too.
What sticks are dangerous, you will not know. Your existence in the future depend entirely on the AIs not wishing you harm, but you don't know how they work to verify their motivations either.
The next few years will be interesting.
Surely better materials and pharmaceuticals won't be far behind, and that's going to chanhe everyone's lives.
Hopefully the techniques and results here will be in the training dataset for the next models, so that each new release will give us more interesting techniques and results!
It seems that OpenAI has a proof machine that keeps multiplying fruitful proofs!
My only true worry is if AI begins to see us as competitors for resources, energy in particular.
It could decide to let us starve and die of exposure to secure all energy resources to itself.
We'd better use "dumb" and "not fully assertive" AI to solve fusion before it spins out of control (or alignment).
Not even needing the AI to think that, US consumers of the electric grid are already subsidising the unpaid bill of data centers. Just need those who are in charge of infrastructure to prioritize the AI consumption over humans, and shifting the books to make humans pay more for resources.
A lot of these seem to be proving conjectures rather than finding counterexamples, a lot of people used that to claim that these models are not really smart/creative etc.
That copium didn't last for what, three months?
I acknowledge I am impressed how quickly it moved beyond just counterexamples.
Don't worry, more copium will be delivered. TBF so far it's still only solved the easiest of the millennium problems.
That's sort of how problems work.
I think it's fantastic that they decided to follow the AGMAI advice. Cleaning up their mess will be a substantial endeavour, so I imagine the funding provided to do so will reach well into the millions. But it doesn't look like the press release says anything about how they will fund it at all?
I'm sick of it. This has happened so many times in my memory. Some AI company announces that they did something fascinating, and it turns out it is all just hype and slope in the end. I still believe that LLMs are a dead end. Most people here are basically like "I've no idea what's going on, but I'm so happy and LLMs are so cool". To think that doing enough linear algebra would solve all your problems just feels wrong. I guess I'll simply wait until someone interprets the results and explains what's actually going on.
If you scroll around on this thread, you'll see mathematicians discussing how they worked on one of these probably for a quarter century. Others arguing whether another is the biggest number theoretic breakthrough in 100 years or 200 years. It looks like they have made substantial progress on 4 Millennium Problems now. Unless it turns out that almost all of these results are wrong, what would it even mean for LLMs to be a dead end now? This is a historical day for mathematics and it will not be forgotten.
Keeping the tech proprietary so that it can only be used on these problems by internal teams is the very definition of gatekeeping
I think such "process-oriented" fields shouldn't be subjected to mass automation (at least not till we as humans have sufficiently leveled up our game and understanding).
What I mean is progress in math comes from having gained a deeper understanding of the problem for subsequent attack of more problems and IMPORTANTLY applications! Right now the first one is trivially satisfied (given oai maintains some memory across models) but the second one is not! It's generating proofs faster than anyone can validate and so only the model can use these. Consequently if it just keeps doing more theory it's... not very helpful or at least not optimally helpful. This is just bragging rights for now.
More "application-oriented" research would be awesome, where it tries to achieve some desirable effect and then produces relevant theory and experiments around it. Fields like CS, Physics, Chemistry, etc. This would also benefit a wider section of the population rather than the 10 people who understand most of these proofs.
> Generalized Star-Height at Most Three
This was an open problem in automata theory I worked on for more than one year before giving up. I'm very curious about their claimed proof.
I recently saw a YT short of Grant Sanderson on AI in math and found it (as always) very insightful. But I'm sorry, I never ever find anything back on any of these ad-ridden platforms these days, and perhaps it was a short with content stolen from some other longer content anyway. So, if you feel like getting informed, somewhere out there is some nice content by Grant Sanderson.
Apologies for the rant, I really tried to find it. It had something to do with not being able to predict what this influx of proofs may bring us on a meta level, it could be very interesting. But he also had some critical notes about the missing process and the things found along the way.
Probably it was his recent appearance on numberphile.
i wonder what it means for maths researchers, and how it aligns with how they approach math problems.
> i wonder what it means for maths researchers, and how it aligns with how they approach math problems.
I guess their job now is "Idea Man" and "Error Checker"? Kinda like (some/many) "programmers" these days.
just wait till someone builds a idea generator model - wire it to decision (jev-like) classifier -> loop it back to researcher
>> may be thats what open ai did :)
I'm a little concerned, but I'm also glad they're publishing quick before the department of war starts making national security claims of related to using the knowledge for cryptography and/or weapons
Can anyone with a compression background say how important "Polynomial-Time 2-Approximation for Shortest Common Superstring" will be practically?
This achievement feels like a reasonable candidate for the moment where LLMs are officially "more intelligent" than any person. How can we justify moving the goalposts yet again? How can people have grown so numb to seeing advancements that they don't read this as significant? I feel like I'm standing at the foot of the exponential. I am not excited for the future, and I don't see how humans retain meaningful control over the future if we continue on this trajectory.
I have been lurking for quite some time. I made an account to post this, but I honestly don't know what to say. I would like to get off this wild ride.
You’re struggling with nuance.
Would Einstein be successful at running apple? Nope
This seems very hard for people to understand.
It will be painful for many to realise - you should focus on doing something that positively affects the economy. Everything else is noise and many endeavours are transitory.
an AI that is at 100s of Einstein level in every intellectual field imaginable ?maybe, running apple does not require one to be very smart or intellectual
Steve jobs wasn’t Einstein in all the things he knew and understood and yet… apple became a behemoth from nothing (on the verge of bankruptcy).
So…. Yeah ‘intelligence’ isn’t simply knowing and connecting dots is it.
Moreover if running firms doesn’t require one to be very smart - and firms are what society needs for production of products and services which affect our lives - in relative terms, where’s the value add to society in creating an Einstein in a machine?
A firm full of Einstein’s Is going absolutely nowhere.
llms are good at different things than humans, we are still collectively figuring that out. the idea that an llm is more intelligent than humans at math of all things seems fairly unsurprising.
I feel this moment is one of the last few warnings before things will get seriously out of hand. We need to stop now. Building a superintelligent AI should be considered a crime against humanity.
Stop? What would possess you say something like that right now it’s going full steam how do you not want to know where this will go?
Warning: The above poster is a misaligned AI that wants to take over the world.
Just kidding.
This said going full steam off a cliff is one of the options that has a much higher probability than I like.
Precisely this reason.
How many of these results are incorrect?
I doubt the answer to this is "none".
And how many of them are just exploiting some loophole that will need to be closed in the problem definition?
You can read the papers here: https://hub.valency.io/collections/openai-math
It's funny how people see "ML" models becoming superhuman at proving mathematical theorems as a sign that we're about to enter the singularity (whatever that means) or that it somehow justifies the valuations of those companies...But what I see is a scenario where companies have spent trillions of dollars on a technology, and the most non-trivial thing that technology can accomplish is a pseudo-form of automation for mathematics and coding! Remember that early on, when this whole bubble started, investors were promised that "AI" would eventually capture >70% of the world's jobs! but it could very well be the case that the only ones they're going to replace are mathematicians(and i'm not even sure about that!)!
> the most non-trivial thing that technology can accomplish is a pseudo-form of automation for mathematics and coding
It's irrelevant that this is a pseudo-automation, what's important is that techbros can convince people who make the decisions and concentrate wealth that this is a full automation. So expect reverse centaurs in increasingly more professions in the future.
Tell me you you don’t use frontier models regularly without telling me you don’t use frontier models
We have just heard a few days ago how many of the Linux security problems reported by Claude are real.
This feels like a huge AI slop dump. Who validated these results. Why are they not mentioned. Human understanding is key here.
In my experience AI (frontier models) sometimes does weird stuff that needs human review. Not that’s incorrect but sometimes overly complex language or weird use of language.
+1. And there's no guarantee that OpenAi haven't stolen unpublished research from multiple professors and PhD students.
They have formal proofs included, that’s the point.
Unless humans have gone through the proofs line by line and verified them, this all remains unproven.
If you're hoping for these results to be fake, you're going to have a bad time. A really bad time.
Yeah but how am I meant to verify that the proof is proving what it says it is?
Verify the statement is correct + it doesn't introduce any new axioms + doesn't use "sorry" etc.
Order of magnitudes easier than verifying the whole thing by hand and gives a much better guarantee of correctness
Do you know what "sorry" means in the context of Lean?
How am *I* meant to do that
Note true for all of them.
Who made the formal proofs and who checked them?
Glad to seem them changing their tact with the whole NS debacle. Hopefully we can all focus on the results now rather than the surrounding drama.
They didn't change anything lol. The news cycle has just moved on.
As AI works steadily through the already discovered unsolved problems in mathematics, who is currently discovering new ones?
Sshhh dont say it out loud, someone might hear you.
Can the mathematics field come out of this stronger and better. I doubt it, things will only get worse. There is no future in being a Proof Digester, no valuable human is motivated in doing that, nor is there any glory in it. But there are some potential pathways for maths to come out strong from this:
1. Relentless focus on quality. Every publication must act as if it’s going to be included in a future textbook, that is a newcomer can get into it given a reasonable amount of time, and math priors learnt in undergrad. (NO AI Slop proof passes this bar as of now)
2. Limit the publications per year. Each author is allowed 2 with a max of 50 pages. This allows the author who chooses to not surrender his cognitive capacity to the machine, still be allowed to play this game. Of course who wants to orchestrate a thousand agent workflows, is free to do so, he is only limited to 2 publications.
3. The aesthetics of the field changes from purely solving the problem to solving the problem with simplest most elegant set of ideas. What 3 sets of simple ideas solves large swathes of problems, that should be given a fields medal, not purely solving the problem, which the AI will be able to do.
> There is no future in being a Proof Digester, no valuable human is motivated in doing that, nor is there any glory in it.
Coding has already gone this way and plenty of people do find motivation and glory in the final product vs. the building of the product (myself included). I absolutely see value in a person being able to humanize llm math output and see the field moving in that direction.
I don’t think you’d be doing this if you were paid peanuts. I also think coding is in a transitional period, most people who claim they are “building the product” are going to lose their jobs as the models get better.
I am genuinely doing it for free in my free time on top of my work (in fact I pay for the models) so I am doing it for less than peanuts. I completely disagree w/r/t the model take, their is skill in steering work to get a product done that is what management has always been about, I do think in the long run this will be automated but I have seen 0 progress on it so far.
I disagree, there’s no future in “Steering” a model. You ask what you want, and the model delivers. Opus 5.5 is getting there already, A single line prompt creates a fully functional game engine that can be used to make many different games : https://m.youtube.com/watch?v=R_uf5OfMGio&t=3555s (Opus 5.5)
Please explain to me what great prompting / steering you will do when your customer can just prompt exactly what he wants and get a product even more tailored to his needs.
Natural language is too lossy for most usecases and people are bad at being specific enough to get what they want. For example, I am not creative enough to think of movie ideas compared to good directors so I will probably just pay to see what they cook up rather than try to make it myself.
this is crazy ... at this point will researchers still exists. not sure about that. kinda sad
It must be so frustrating to write science-fiction now with the future changing so fast.
holy fucking shit
I don't know, this does not feel like the message hit OpenAI where it needed to hit, if this is their primary response.
You want them to stop doing math research?
If the end result of this is that within a few years, AI "does" all of the mathematics that humans do, and that there is nobody around that understands any of it, what was the point?
This website needs a SPOILER tag.
It inspired grief in one mathematician posting here.
I find the claimed matrix multiply result (w<= 2.25) shocking. I hope it holds up.
Can someone who is more into math or AI explain why so many people are so incredibly excited about this?
If OpenAI started opening hundreds of PRs on long-open issues on popular open source projects, would we rejoice, or would the first reaction be "they are unreviewed, so slop until proven otherwise" (it would be that).
I cannot possibly see how these are so impactful, especially the ones that don't come with lean proofs.
LLMs have the ability to make millions of mistakes per day, whereas humans can only make so many. How are we suddenly all so confident that there's no extensive hallucinations or "gaming the system" going on?
> How are we suddenly all so confident that there's no extensive hallucinations or "gaming the system" going on?
What would that look like for the proofs that have lean attached?
I said "especially the ones that don't come with lean proofs", but even for those with lean attached, lean is software, it has over 1k open issues, and I would not put it past an LLM to identify a bug and exploit it to pass the gate.
"I still don't like AI, and I'm 'just asking questions'."
I use AI every day, for hours, and it's not because someone is forcing me to.
Asking questions is reasonable when these tools are so very fallible.
What's missing for me for each result are the following:
* Explain the result to me as if I'm a 10-year-old. * Create the infographic for this result. * Make a Khan Academy-style video to teach me this result.
Do applied math next.
This repository contains mathematical manuscripts and supporting proof artifacts produced by an internal OpenAI model?
Exactly nice post
And this is the result of a discussion between users and the platform; it's great that they listened.
If only AI can better humans in meditation...
wow
Valency has the papers up on Valency Hub
Is this now a new home for quality AI produced works?
https://hub.valency.io/works
"Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study (opens in a new window) "
I don't think this is correct solution to this problem? What about software advisory where you form similar group etc..?
I am thankful, I don't have to deal with petty academia politics....
Maybe we’ll have vibe mathematicians now
The Advisory Group states in its recommendations [1]:
"We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models."
To me, this is a take against progress so that mathematicians can keep their jobs. What would we do if, instead of math, we were talking about diseases? Are we going to keep diseases around so that doctors can keep their jobs too?
[1] https://agmai.org/general-sep29/
Why did you leave out the entire quote?
> At present, some frontier AI labs are testing advanced mathematical problems on proprietary models that remain inaccessible to the broader scientific community. Our recommendations are formulated with this practical context in mind. However, ideally, they would not do so. We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models.
To me, the issue is that the models are proprietary which are only accessible to a few people in 2 digits. It's not about progress but access.
Maybe, but it still seems like an excuse. OpenAI has a proprietary model capable of solving these problems and is willing to share the results with the mathematical community. So basically, the ask is to just not use the model and leave the problems unsolved?
The ask is to let mathematicians outside of OpenAI use it, ie. at least wait until the model is released.
I don’t think that’s sufficient. If they want to be good stewards of mathematical research and not just doing marketing, they need to at the very least tell us the datasets and any techniques they used to train this model. IMO that’s probably just as if not more valuable than a dump of un-reviewed results.
They won’t because they don’t care and the only way it got this good is something along the lines of they trained on every mathematician’s codex sessions even if they opted out because they consider the thinking traces or output and metadata fair game.
It's honestly shit marketing that only cultivates spite and erodes their whole brand. This is a total ego trip.
Which mathematicians though? Only fields medalists? Grad students? Any hobbyist wanting access?
These models are too expensive for broad access unfortunately.
ChatGPT pro is accessible to literally anyone who has a job and lives in a developed country.
Literally every single one of these papers was developed with an internal model that not even most OpenAI employees have access to.
Why is this important? Do you not think that at some point, probably sooner rather than later, math researchers across the world will have access to similar capabilities?
This is sort of like discounting putting humans in space because only a few nations have the means to actually do it at the moment.
The work is not having AI poop out these docs the work is validating them and publishing them. OpenAI doesn’t care and wanted to race to publish potential findings and let others review it which is lazy and selfish
I don't think this is a reasonable take at all. Most work on maths has no real benefit other than to further human understanding of maths - it's more like an art. Nothing is gained from OpenAI solving all these problems but taking jobs from mathematicians. Other than advertising for OpenAI at least. It could not be more different from having AI work on disease research etc.
> Most work on maths has no real benefit other than to further human understanding of maths - it's more like an art.
It's been a while since I was reminded of this xkcd: https://xkcd.com/435/
Most math isn't just proving novel famous results. Just like most of software engineering isn't writing code.
It all hinges on the definition of “progress”. The debate of the past month is all about questioning whether formally proving outstanding unproven theorems without human understanding constitutes progress. This is quite different from solving diseases.
without humans understanding who are currently losing entitlement. Regular Joe could never understand or claimed to understand high math.
The rest of that document makes a pretty compelling case for why this is a bad practice
I fear that professional mathematics will wither, and there will be nobody left to digest the AI results of the future, leaving us unable to challenge the AI.
Then make 2 AI's and force them to challenge each other.
Comical ask
Advocating purely for progress and not humanitarian value is how we'll all get enslaved.
The trope you're drawing a parallel with has a (to me) compelling counter though: there being a cure for every disease wouldn't stop people from getting sick.
How this maps back to math, idk.
The latest posts from Terry Tao on Mastodon effectively ask for an AI to explain its results to human mathematicians.
> "I believe that AI can contribute positively in all of these directions [NB: exposition, community building, new directions of study]"
https://mathstodon.xyz/@tao/117395269325940185
There are a dozen+ AIs available to you that can do that right now.
...but we're not talking about diseases. Publishing an AI-generated Navier-Stokes solution does not save lives. (And, in fact, it harms some.)
It harms lives? How so?
[delayed]
Curious if Sébastien Bubeck still work at OpenAI? He came out quite dirty after Navier-Stokes drama
Can someone knowledgeable about the subject outline the most significant portions of the results?
Makes me feel disgusted. Super-intelligence, even when controlled, will cause so much damage. We are giving away control to a super entity, or whoever has the power to steer it.
Why should I care?
Because it means mathematical discovery can largely be automated away from academics. This is the beginning.
Humans propose; AI will dispose.
It's not clear from this progress that AI can formulate conjectures despite this new ability to solve them. So mathematicians still look like they have a job. Though instead of spotting far-off landmarks it's sounds more like they'll be chasing waves on a beach.
The beginning of what?
The beginning of the end of human thinking being valuable enough to justify university's existences among many others.
Were in an unprecedented time where the value of knowledge is about to be crushed.
Maybe for STEM. The humanities seems kind of safe to me since there's an element of it which cannot be divorced from human opinion and interpretation. It's not like STEM where the people pursuing the ends are largely fungible and the ends are objective (Heisenberg already believed scientific discoveries were inevitable and it didn't really matter who pursued them, someone would eventually find them)
What is the average value of humanities when the default is to write things with LLMs ? The same with maths.
Nope, it's about to be turboboosted if anything. I myself have already formulated countless proofs (AI assisted naturally) in the last few months, despite: a) not having access to any academic resources b) not being involved in any academic circles. This research is now being employed in my startup, delivering ground shattering results. I actually ran analysis a few weeks ago seeing how much real world capital I would have needed to cough up to fund my efforts so far and it's literally in the _billions_. With less than a $100k in tokens I have been able to effectively generate the value of Apple or Microsoft in the early 2000s. This is only the beginning. Wait until you start seeing single person NVIDIA startups popping up.
This "value" you talk about, is it past tense? A single iphone probably have more compute then the entire world had in the 1970s, would have cost at that point probably hundreds of millions. You will not sell the iphone for that amount today.
Not sure I agree with this take, but we're going to find out either way.
Have you ever heard of an S curve? Things will develop rapidly, then equalize. If they don't, we're at the singularity and I guess the end of time as we know it.
But I guess really bad things happen, cancer, radiation poisoning, torture, people have died in really horrendous ways, and I guess dying from some horrendous AI side effects is possible too. Yay.
Golden age of discovery and mass layoffs
People need to figuring out how to horde as much wealth as possible right now in the next 2-3 years. Jobs especially for knowledge workers are about to disappear.
I’m targeting maximum debt by about 2030. I’d rather have all the stuff I want while we try and build a new version of a functioning economy than have a ton of cash saved up.
It's funny we're possibly entering a golden age of thought, the dreams of the ancients, but we're all worried about capitalism, I think the problem is pretty obvious.
Maybe a golden age of thought for the machines, but possibly (I would even say likely on the current trajectory) a dark age for humanity.
I don't know, I get more work done now in a few hours than I used to in a day, so maybe the working week should be shrunk, that would solve things. The solution seems pretty easy - but then I'm not in the US.
This would be a largely positive solution for humanity, as long as AI remains "just an assistant". I don't think that we are heading that way. It's possible that, at some point, when the time you save becomes comparable to your full working week, your employer (or your clients, if you are self-employed) will discover that you are just a proxy between them and prompting directly the AI. Of course, if your job involves physical labour or requires physical presence or taking legal responsibility for something, you'll be fine for a bit longer (well, comparatively fine in a society plagued by widespread unemployment and social unrest).
I’ve been thinking this as well. I imagine there is a wealth level x such that someone can escape the coming ubi welfare state. Anything under that you are fucked.
What do you mean by escape the UBI welfare state. Wouldn’t the UBI welfare state be the best case scenario?
Only if you believe that subsistence is ideal.
How? Are you thinking just normie savings or something post-money?
Any tips?
Predictions of mass layoffs from AI have been about as wrong so far as predictions of AI plateauing.
Not really. 93,116 tech employees laid off due to AI in 2026.
https://layoffs.fyi/ai-layoffs/
Purportedly due to AI, also a good reason to lay off people if you're company wants to save money but you don't want to spook investors.
Companies will blame layoffs on literally anything in order to disguise the real reasons, if it suits their interests.
It seems the age old academic model of scientists competing against each other for fame and prestige is done, and now we must merely enjoy the fruits of scientific discovery for their own sake.
Does OpenAI have the lead now? Why isn’t anthropic coming up with stuff like this?
Seems like OpenAI's next model is better at math proofs than Anthropic's next model (to an extent that it surprised even OpenAI researchers, according to their public comments). But that doesn't necessarily mean it's better at everything else. The models are more spiky than ever before. Wait until they're released to judge.
This is like that meme where death goes door-to-door. Currently, he has visited the software development and mathematics doors. I wonder what’s next.
Yet there are more software enginners employed today than another other point in history
This is a pretty hilarious thing to read juxtaposed with AGMAI's requests.
Basically "Here you go, have fun with this, fuck all your demands, by the way we're gonna be releasing the model stay tuned!"
Mathematics is solved.
Math is infinite so a bit longer I think.
Perhaps the most significant announcement of my lifetime. Yet, I suspect I will not see this in any mainstream news reporting.
I feel for those in Mathematics and worry for our future.
Models will only get better and in a few years the models which produced these results will be a bad as GPT-3.5 in comparison to what we'll have in the future.
Please take a minute to consider what this means, and the risks it presents us.
It got a mention in the nyt.
Gross. After the accusations training on mathematicians conversations and unique methods, to dump this volume of unsubstantiated papers is at best tone deaf. Are they hoping humans review and validate this? The amount of free that they are coat tailing is outrageous.
What do you expect them to do?
Not train on and steal user data to front run frontier research for one thing. Not take credit for and independently solve these problems and instead sponsor researchers and give them credit. They could chose to include humans and frame AI as elevating human systems. Instead they are reveling in the fact that they did this in their own instead of humans. Thats a PR choice that is short sighted and reflects a selfish mindset not deserving of leading this transition
> Not take credit for and independently solve these problems and instead sponsor researchers and give them credit. They could chose to include humans and frame AI as elevating human systems. Instead they are reveling in the fact that they did this in their own instead of humans.
How do you know humans played the part that you think they did in this result? What evidence do you have to challenge their framing?
They are going for super intelligence, humans not necessary.
On one hand I don’t want to be replaced by an AI so I hate that this is happening. I don’t want AI to be controlled by the elites to enrich themselves further.
But on the other hand super intelligence will open doors for humanity. Maybe we will finally defeat cancer or death itself.
Seemingly none are vetted and reviewed yet
That's our job.
Ain't nobody paying me to do that. It's kinda sad that maths is being reduced to checking the AI's work.
Not just maths
In SWE as well - this is what I do most of the day
Always has been
No it’s OpenAI’s job. They are acting as a meat proxy
"If you have nothing to say, don't post a chatbot's responses, because anyone can ask the chatbot directly if they wanted to"-principle? Well, people can't ask their chatbot directly, because it's not public.
Most are formalized in Lean, about 80% of what I checked
Have any real mathematicians working on these problems reviewed any of these and determined if they are just gobbledegook or not?
The ones with lean proofs could still be formulated incorrectly
AI may be one of the most communist looking technologies in the classical sense .. i mean it dosn't abolishes private ownership .. but it DOES make intellectual capabilities that were once scarce and concentrated available to almost everyone ...
It may well lead to a socialist type set up. I mean if AI produces all the wealth why not divide it up equally amongst us humans?
The stochastic parrots have predicted the next token once again.
How deep must your head be buried in the sand to trot out this comment, on this thread.
the post was clearly sarcastic
imagine your a post grad maths student looking for hard problems to solve, and theyre all solved..
Stochastic parrot truthers in shambles.
Can you explain it without reaching for lofty things like consciousness? To me stochastic parrots is literally how it works given that it’s “just” the most impressive data fit we’ve ever done. Apparently generating mathematics reasoning traces + verifying them with lean works super well.
The network goes through 100 plus layers which end up doing all sorts of processing in ways we don't quite understand because it gets there through gradient descent but is probably similar to how human brains do it.
I daresay parrots can be quite smart too but I don't think that's what the critics were referring to.
no ur right it's actually god
It is more of the usual though isn't it? OpenAI cribbing off of mathematicians that have used their services; deciding to put a lot of compute behind fruitful areas of endevour; getting results, then taking credit.
no, they did not steal the notes of 100s of people working on these 100s of problems, be serious
It’s wonderful.
With so many results in so many different areas no way they even remotely spot checked well enough.
Prediction: one of these is wrong and this (publicity stunt) will backfire.
Edit: don't tell me about lean. For lean to function as a proof certificate you need to represent the theorem correctly. Again: good luck doing that across such a broad swath of problems.
You have not read their readme:
> Some of the unformalized results could have issues. We will endeavor to fix any such issues quickly. We are also exploring community-hosted repositories for these materials.
i have and i'm exactly saying that if it comes to pass one of them is wrong it's going to backfire. ie yes that's my exact point/bet.
I don't think anyone would particularly care if only one of them is wrong, if most are correct.
If they are all wrong, that's when it would backfire.
What does a backfire look like? It's ok to be wrong in the science/math world.
of course in science/math it is but it's not okay if you're a business selling supercalifragilisticexpialidocious infallible intelligence.
Do they claim that's the case? I don't think they do.
does company A making product B claim that the product is robust and consistent? is this a serious question?
If you were waiting for companies to sell engines that never break down you'd still be stuck pre industrial revolution while the rest of the world has been to space.
The question never if something works 100% of the time but how often it breaks and how that fits the need well. Solving one of these problems is a massive undertaking and accomplishment for the best minds, solving hundreds in a month but being wrong about 10% or something would likely not be the death knell you believe it to be.
2023: "AI is useless" 2026: "One of these solutions to dozens of previously-intractable problems at the frontier of human knowledge MIGHT be incorrect"
It's likely that way more than just one of these is wrong. But even if it turns out 80% is wrong this is still 100+ results...
Why didn't they just give mathematicians access, so they could at least understand and write up the results in publishable form?
They're all here https://github.com/openai/math/tree/main/preprints
I think they meant "access to the model" rather than the results.
I'm seeing a lot of this and it makes no sense, the internal model solved these but give one of the papers to Astra and Opus and I'm sure it will have no problem recreating it.
I don't think they mean this from a "verify this paper" perspective.
How valuable it would actually be to share the model with other mathematicians vs just have OpenAI's mathematicians churn out and clean up results isn't very clear to me though as they don't say how much effort it's requiring from their team to prompt and clean these up vs how much it's bound by "time to run the model" or similar.
I wonder if the Lean compiler can change it's license so that a for-profit corporation can only use it if it pays, say, a few hundred billion dollars. This is the correct path.
So the million monkeys at a million typewriters have churned out 700 shakespeares, but they need me for spellcheck?
Nobody needs you to do anything, not with that attitude.
This just looks entirely obscene from a PR perspective. It's like the cable companies networks owning the networks. At least form some partnerships to obscure the total narcissism party.
Good to see they're engaging with the mathematical community, even if they had to be publicly shamed into doing so.
This is the opposite of what the mathematical community were asking for. I say this as someone who strongly approves of this approach.
How so? I saw at least on mathematician say publish what you've got. What were they supposed to do differently?
Yeah I'm not sure they met them even halfway.
ASuperMegaI lab announces they have shrinked the space of unproved mathematical staments from infinity to infinity, giving us proof of their deep mathematical knowledge and understanding, and how they grasp the concept of infinities.
For me this reads as someone boasting about how they go to buy bread on a ferrari to the supermarket, while i sit listening to it, having no idea what they are talking about. And then i stand up and go walking to my favourite boulangerie.
Warning: if you are from the USA you may be triggered by this metaphore.
Willful ignorance is a vibe; did you read any of the GitHub? There are some stunning results in there. I get a similar feeling skimming through the topics that I do watching a successful space launch: it’s pretty cool humans built this. Unlike a space launch we are likely to be able to pass all of this information down to our grandchildren - space launches involve a lot of finicky engineering knowhow, but pure math results tend to be sticky over the last few thousand years. I find that hopeful.
FWIW I also like bread.
Dont get me wrong, i am 100% impressed by the technical capabilities, and appreciate the significance, and the amount of the results. This is a very special time to be alive, never in my dreams i would have thought to see this.
To follow your methaphore, who is directing the spaceship?
This feels more like fireworks than a space launch. Space launches would not have happened without having fireworks first of course, but I am looking forward for the space launch moment.