719

If AI can now speak Italian, it can certainly replace us... (sopuli.xyz)

submitted 2 years ago by lseif@sopuli.xyz to c/programmerhumor@lemmy.ml

89 comments fedilink hide all child comments

you are viewing a single comment's thread
view the rest of the comments

[-] stingpie@lemmy.world 69 points 2 years ago

This might be happening because of the 'elegant' (incredibly hacky) way openai encodes multiple languages into their models. Instead of using all character sets, they use a modulo operator on each character, to make all Unicode characters represented by a small range of values. On the back end, it somehow detects which language is being spoken, and uses that character set for the response. Seeing as the last line seems to be the same mathematical expression as what you asked, my guess is that your equation just happened to perfectly match some sentence that would make sense in the weird language.

[-] PlexSheep@infosec.pub 32 points 2 years ago

Do you have a source for that? Seems like an internal detail a corpo wouldn't publish

[-] stingpie@lemmy.world 20 points 2 years ago

Can't find the exact source–I'm on mobile right now–but the code for the gpt-2 encoder uses a utf-8 to unicode look up table to shrink the vocab size. https://github.com/openai/gpt-2/blob/master/src/encoder.py

[-] crispy_kilt@feddit.de 3 points 2 years ago

Seriously? Python for massive amounts of data? It's a nice scripting language, but it's excruciatingly slow

[-] stingpie@lemmy.world 6 points 2 years ago

There are bindings in java and c++, but python is the industry standard for AI. The libraries for machine learning are actually written in c++, but use python language bindings. Python doesn't tend to slow things down since machine learning is gpu-bound anyway. There are also library specific programming languages which urges the user to make pythonic code that can be compiled into c++.

[-] NeatNit@discuss.tchncs.de 17 points 2 years ago

I suppose it's conceivable that there's a bug in converting between different representations of Unicode, but I'm not buying and of this "detected which language is being spoken" nonsense or the use of character sets. It would just use Unicode.

The modulo idea makes absolutely no sense, as LLMs use tokens, not characters, and there's soooooo many tokens. It would make no sense to make those tokens ambiguous.

[-] stingpie@lemmy.world 8 points 2 years ago

I completely agree that it's a stupid way of doing things, but it is how openai reduced the vocab size of gpt-2 & gpt-3. As far as I know–I have only read the comments in the source code– the conversion is done as a preprocessing step. Here's the code to gpt-2: https://github.com/openai/gpt-2/blob/master/src/encoder.py I did apparently make a mistake, as the vocab reduction is done through a lut instead of a simple mod.

this post was submitted on 12 Jun 2024

719 points (98.0% liked)

Programmer Humor

42245 readers

7 users here now

Post funny things about programming here! (Or just rant about your favourite programming language.)

Rules:

Posts must be relevant to programming, programmers, or computer science.
No NSFW content.
Jokes must be in good taste. No hate speech, bigotry, etc.

founded 6 years ago

MODERATORS

AgreeableLandscape@lemmy.ml

cat_programmer@lemmy.ml