663

submitted 2 years ago by L4s@lemmy.world to c/technology@lemmy.world

150 comments fedilink hide all child comments

Over just a few months, ChatGPT went from correctly answering a simple math problem 98% of the time to just 2%, study finds. Researchers found wild fluctuations—called drift—in the technology’s abi...::ChatGPT went from answering a simple math correctly 98% of the time to just 2%, over the course of a few months.

you are viewing a single comment's thread
view the rest of the comments

[-] Windex007@lemmy.world 152 points 2 years ago

It isn't and has never been a truth machine, and while it may have performed worse with the question "is 10777 prime" it may have performed better on "is 526713 prime"

ChatGPT generates responses that it believes would "look like" what a response "should look like" based on other things it has seen. People still very stubbornly refuse to accept that generating responses that "look appropriate" and "are right" are two completely different and unrelated things.

[-] deweydecibel@lemmy.world 17 points 2 years ago* (last edited 2 years ago)

In order for it to be correct, it would need humans employees to fact check it, which defeats its purpose.

[-] Windex007@lemmy.world 19 points 2 years ago

It really depends on the domain. Asking an AI to do anything that relies on a rigorous definition of correctness (math, coding, etc) then the kinds of model that chatGPT just isn't great for that kinda thing.

More "traditional" methods of language processing can handle some of these questions much better. Wolfram Alpha comes to mind. You could ask these questions plain text and you actually CAN be very certain of the correctness of the results.

I expect that an NLP that can extract and classify assertions within a text, and then feed those assertions into better "Oracle" systems like Wolfram Alpha (for math) could be used to kinda "fact check" things that systems like chatGPT spit out.

Like, it's cool fucking tech. I'm super excited about it. It solves pretty impressively and effiently a really hard problem of "how do I make something that SOUNDS good against an infinitely variable set of prompts?" What it is, is super fucking cool.

Considering how VC is flocking to anything even remotely related to chatGPT-ish things, I'm sure it won't be long before we see companies able to build "correctness" layers around systems like chatGPT using alternative techniques which actually do have the capacity to qualify assertions being made.

[-] datavoid@lemmy.ml 4 points 2 years ago

That's kind of the whole point of RLHF though

[-] oktoberpaard@feddit.nl 1 points 2 years ago

That’s not necessarily true: https://arstechnica.com/google/2023/06/googles-bard-ai-can-now-write-and-execute-code-to-answer-a-question/. If the question gets interpreted correctly and it manages to write working code to answer it, it could correctly answer questions that it has never seen before.

this post was submitted on 20 Jul 2023

663 points (97.6% liked)

Technology

84941 readers

1675 users here now

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related news or articles.
Be excellent to each other!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
Check for duplicates before posting, duplicates may be removed
Accounts 7 days and younger will have their posts automatically removed.

Approved Bots

founded 3 years ago

MODERATORS

L3s@lemmy.world

enu@lemmy.world

technopagan@lemmy.world

L4s@lemmy.world

L3s@hackingne.ws