19

I know people here are very skeptical of AI in general, and there is definitely a lot of hype, but I think the progress in the last decade has been incredible.

Here are some quotes

“In my field of quantum physics, it gives significantly more detailed and coherent responses” than did the company’s last model, GPT-4o, says Mario Krenn, leader of the Artificial Scientist Lab at the Max Planck Institute for the Science of Light in Erlangen, Germany.

Strikingly, o1 has become the first large language model to beat PhD-level scholars on the hardest series of questions — the ‘diamond’ set — in a test called the Graduate-Level Google-Proof Q&A Benchmark (GPQA)1. OpenAI says that its scholars scored just under 70% on GPQA Diamond, and o1 scored 78% overall, with a particularly high score of 93% in physics

OpenAI also tested o1 on a qualifying exam for the International Mathematics Olympiad. Its previous best model, GPT-4o, correctly solved only 13% of the problems, whereas o1 scored 83%.

Kyle Kabasares, a data scientist at the Bay Area Environmental Research Institute in Moffett Field, California, used o1 to replicate some coding from his PhD project that calculated the mass of black holes. “I was just in awe,” he says, noting that it took o1 about an hour to accomplish what took him many months.

Catherine Brownstein, a geneticist at Boston Children’s Hospital in Massachusetts, says the hospital is currently testing several AI systems, including o1-preview, for applications such as connecting the dots between patient characteristics and genes for rare diseases. She says o1 “is more accurate and gives options I didn’t think were possible from a chatbot”.

you are viewing a single comment's thread
view the rest of the comments
[-] InevitableSwing@hexbear.net 29 points 1 month ago

There is definitely a lot of hype.

I'm not being sarcastic when I say I have yet to see a single real world example where the AI does extraordinarily well and lives up to the hype. It's always the same.

It's brilliant!*

*When it's spoonfed in a non real world situation. Your results may vary. Void were prohibited.

OpenAI also tested o1 on a qualifying exam for the International Mathematics Olympiad. Its previous best model, GPT-4o, correctly solved only 13% of the problems, whereas o1 scored 83%.

Ah, I read an article on the Mathematics Olympiad. The NYT agrees!...

Move Over, Mathematicians, Here Comes AlphaProof

A.I. is getting good at math — and might soon make a worthy collaborator for humans.

The problem - as always - is the US media is shit. Comments on that article by randos are better and far more informative than that PR-hype article pretending to be journalism.

Major problem with this article: competition math problems use a standardized collection of solution techniques, it is known in advance that a solution exists, and that the solution can be obtained by a prepared competitor within a few hours.

“Applying known solutions to problems of bounded complexity” is exactly what machines always do and doesn’t compete with the frontier in any discipline.

---

Note in the caption of the figure that the problem had to be translated into a formalized statement in AlphaGeometry's own language (presumably by people). This is often the hardest part of solving one of these problems.

AI tech bros keep promising the moon and the stars. But then their AI doesn't deliver so tech bros lie even more about everything to get more funding. But things don't pan out again. And the churn continues. Tech bros promise the moon and the stars...

[-] Barx@hexbear.net 6 points 1 month ago

You have to admire the grift.

Shame it requires the energy use of entire countries and is a weapon for disciplining labor.

this post was submitted on 04 Oct 2024
19 points (85.2% liked)

technology

23308 readers
243 users here now

On the road to fully automated luxury gay space communism.

Spreading Linux propaganda since 2020

Rules:

founded 4 years ago
MODERATORS