Yes, but there's still the possibility that it's either exploiting a bug in Lean, or that the theorem statement is not set up correctly (i.e. it's actually proved a different theorem).
My understanding is that the theorem statement is quite simple, so i guess the latter is not very likely, but the former is very much a possibility in a proof this large, and it will take some human eyeballs to go over it before convincing mathematicians.
They do a Comparator Challenge to validate that they actually solved the correct theorem from the result, which they copied from Google/DeepMind: https://github.com/openai/NavierStokesAndEuler/blob/f9e8bc5b... - this is valid for both Euler and NS.
Also, they validated with an external kernel from the Lean Kernel Arena. That way bugs in the Lean kernels were found in the past already, iirc.
Having said that, I strongly believe a positive result with the challenge above is the reason why they published it. I highly doubt anybody at OpenAI (or anywhere else) fully gets the proof after such a short time since publishing. This is also what Terry Tao criticized the most in my opinion.
Independent of the remaining drama [0], from my point of view, the proof is correct and an achievement.
Verification with Lean is a piece of empirical evidence that the proof is correct. The paper passing peer review would be another. But even together, those two would be insufficient to establish the claim.
While it's a convenient to assume that mathematics deals with logical statements, any attempt to evaluate those statements relies on physical processes with both known and unknown failure modes. There cannot be a test that establishes it unambiguously whether a claim is true or false. In all nontrivial situations, mathematical truth is based on expert consensus. When a new claim is made, people will try to raise and resolve objections, until a consensus emerges one way or another.
As for C++, all compilers are different. For any given compiler, there are valid C++ programs the compiler fails to compile and invalid programs it compiles without any errors or warnings. And now that I think of it, a new version of a compiler crashing with valid code earlier versions used to handle is the only class of compiler bugs I see with any regularity.
Formalized in Lean, just five months ago [0], resulted in discovery of bugs.
Just because Lean can compile it, does not mean it is safely proven. It is the start of a process to check whether something actually holds, not the end.
If that's the current burden of proof required in your world for maths then that's fine! 't'ain't in my world: I want to see peer reviewed and published. Surely that's not too much to ask. Its not perfect but generally works rather well for maths.
I'm not a sodding programmer so please don't assume everyone here is one. I'm not a mathematician either but I do have standards: Your counter argument is a poorly constructed and inappropriately deployed example of "proof by whataboutism".
In what world is peer review a higher standard than formal verification in Lean?
Not in the world mathematicians have been living in for the past decades at least. Nearly all big theorems that have been formalized so far had been published beforehand, and it was usually regarded as a step up in rigor. Wrong results get published in peer reviewed journals all the time.
I find it really strange the way “peer-reviewed” is used by the general public as some gold standard of truth. As a former academic who has been there, the process is extremely arbitrary and variable. Are people aware that the “peer” refers not to a community or a committee, but literally to one random guy or maybe a couple with zero accreditation? And that the journal editor can do whatever they want with this peer’s “review” including completely ignoring it?
The existence of solutions to problems does not prevent you from solving the problems yourself anyway, if your goal truly is personal development. You're free to go solve Navier-Stokes yourself right now. You're free to manually do all of the AI work in any field, actually. You won't get grant money or prestige (which are not a part of conceptual understanding and insight) but you will get all of the conceptual understanding and insight you're after. You have not been deprived of it.
Except most people that get actually good at a field do so through their day job. By removing mathematician as a viable job, most can't afford to spend the necessary hours to become a good mathematician at all. We're going back to only rich people or those with a rich patron having a chance of developing themselves.
Are they claiming that the only value in solving these problems was for their field's personal development process? I thought Navier-Stokes (and some of the other millenium prize problems) actually had implications for useful technology. It would be insane to demand that people avoid making progress on technology that can save lives or improve general quality of life, just to protect the sanctity of your karate belt system. Perhaps in lieu of open problems left to solve, mathematicians should be welcome to take up chess or sudoku to keep their minds spry.
No, there's most likely zero practical benefit of having found a pathological edge case in which the Navier-Stokes equations do not work. We're most assuredly not talking about "saving lives" here. Unless advanced aliens show up and tell us they'll destroy Earth unless a counterexample to the N-S equations is provided within 24 hours.
> I thought Navier-Stokes (and some of the other millenium prize problems) actually had implications for useful technology.
This is the core misunderstanding that the open letter is attempting to correct.
Developing a better understanding of the Navier-Stokes equations could have a number of implications for useful technology. They're fundamental to fluid dynamics, and turbulence in particular is something that many people feel we could work with more effectively if we better understood how and why it's generated. The Navier-Stokes smoothness problem is an interesting and long-standing benchmark for this understanding; we don't know why it should be so hard to answer, so we hoped that the process of developing a proof to the problem would produce more understanding. (We may still be able to extract this understanding after the fact, if OpenAI's proof is fully human-comprehensible.)
Simply knowing that there exists a finite-time blowup is not practically useful. We know that fluids in the real world don't produce random singularities, so the result can't really have much physical meaning. What it illustrates is that the Navier-Stokes equations fail to model physical fluids in some yet to be characterized way.
Many people (possibly a majority) use their phone for most of their computing, and perhaps a tablet to stream movies/tv. They use a laptop or desktop for work only, and see those as frustrating and clunky devices. There could be a lot of demand for "phone, but with a tablet screen when I want to watch Netflix" from the segment of the population that has zero overlap with people commenting on a tech forum.
I know someone exactly like that. He runs a very profitable side business (reselling) using only his smartphone to take pictures, upload them, message customers, emails etc. I'm always baffled how he doesn't at least have some small laptop to help him type out stuff when needed, but he's doing just fine.
I just got the z fold 8. I'm finding vibe coding much more appealing in tablet mode with the extra horizontal real estate. Chat on the left, browser preview on the right. Before, vibe coding away from my desk has always been a PITA. This is actually useful.
Huge outside the US where people live in small places and use public transport rather than living in mansions with big TVs and be in a car (no phone use) all day.
Your example rewritten in intelligent English (I was curious):
> Note: the potential for a console freeze was previously noted but ignored. handoff-4.3-done.html stated, "could not break console, but [will need fixed later if I'm wrong]."
One could imagine that a perfect writer might also append: "It could be worth looking into what caused that wrong assumption, to prevent similar cases in the future," at most.
Everything else seems to be bad attempts at relatable writing to invoke emotion (an exercise that we should really stop trying to train emotionless matrix weights to attempt).
> Everything else seems to be bad attempts at relatable writing to invoke emotion (an exercise that we should really stop trying to train emotionless matrix weights to attempt).
One of the things actual science fiction got wrong: to the extent that the thing AI does can be called "understanding", emotion is not unusually difficult for them to understand.
I think this was the biggest shock of the original ChatGPT for me. Just how completely unrobotic its voice was compared to everything we'd ever imagined in sci fi. Even that early version was also way more adept at understanding things like implication and sarcasm than any movie AI.
Me too. Almost every Sci-Fi AI proceeds from the premise that we will make something very obviously machine and then have to train it to seem more human. I was completely caught off guard by us taking the approach of distilling all available human output into a statistical model and using it to brute-force something resembling thought and personality through sheer data processing scale.
The unsurprising part once it was clear that approach was viable, was that humans wouldn’t be able to help but anthropomorphize it. I feel like the movie Ex Machina is more relevant than ever.
The "benefiting all humanity" charters were immediately demonstrated to be a ruse. The business model is to hook users into endlessly chatting with your new friend, thus increasing their sales. Yeah, it was surprising and disappointing.
I can buy that Meta's model is that, and IDK about Grok because I stay far far away from it, and I think OpenAI are throwing business ideas at the wall and seeing what sticks.
The clear exception here is Anthropic, who seem to mostly be selling to software developers, whose general reaction to the bot is "please talk less and just do the work, I have enough going on without having to read you yammering".
Before they really started to figure out instruction tuning, there were some wild moments. The AI Dungeon 2 "storyteller" would regularly "lose patience" with its users and roast them or even "hang up" on them.
I tried MUTHR from Alien, but had to keep toning it down because it took it too far, and eventually disabled it (in favour of caveman mode) because it didn't seem to be able to function properly talking like that.
This is the same as systemic bullshitting. Not the first time I've smelled it on fluffy LLM output. I think it's the result of the training trying to induce the LLMs to talk over users' heads even when they are professionals, to entice further use on the grounds of 'oh it's so smart I can't understand its genius train of thought', but I don't think eliciting language like that really taps into 'associations of smarter previous language users'. More likely it's 'associations of rampant bullshitters'.
Thank you so much for this link, this is extremely interesting and it really makes me wonder, consciously, I have never read this construction online before, but is this because I actually haven't seen it, or did I subconsciously write it off as an abbreviation or a typo. Maybe there is some analysis of how this construction gets used online, too, which might reveal some linguistic patterns of internet communities.
A data center allows the website you're whining on to function. If they're really as inconsequential to your life as you're implying, maybe you should boycott them by not using any websites or apps that rely on a data center to function.
There's probably a correlation between how serious the work someone is doing is and how likely they are to want to disable training usage. An opt-out might end up specifically jettisoning the best training data.
A decentralized inference network would be cool. Something that's set up so that I can run a model for personal use on beefy hardware, but also farm out the unused GPU time to the network, probably at much lower prices than normal providers since it would be slower and would lack data security guarantees.
There's probably some selection bias in that stat, too. Based purely on vibes, I wouldn't be surprised if startups pitching to a16z are using open Chinese models slightly more often than startups pitching to other VCs.
That doesn't really even diminish the contribution from Fable, if true. Droves of grad students have been provided the same sorts of non-trivial insights and turned up no results.
reply