139
submitted 1 year ago* (last edited 1 year ago) by cypherpunks@lemmy.ml to c/technology@beehaw.org
you are viewing a single comment's thread
view the rest of the comments
[-] Moonrise2473@feddit.it 4 points 1 year ago

Seems very improbable that they scraped a pirate website with forced registration and tight daily download limits (10 books a day max?) to get content that's often mislabeled and not presented in an homogeneous way.

Probably it's just using the excerpt from Amazon (which instead with paid API access is much more easy to access) as a prompt and build on it

[-] luciole@beehaw.org 3 points 1 year ago

There's been ongoing suspicions that pirated content was used to train popular LLMs simply because popular datasets used for training LLMs do include such content. The Washington Post did an article about it.

Google's C4 dataset used for research included illegal websites. What remains to be seen is if it was cleaned up before training Bard as we know it today. OpenAI as revealed nothing on its dataset.

this post was submitted on 10 Jul 2023
139 points (100.0% liked)

Technology

37747 readers
201 users here now

A nice place to discuss rumors, happenings, innovations, and challenges in the technology sphere. We also welcome discussions on the intersections of technology and society. If it’s technological news or discussion of technology, it probably belongs here.

Remember the overriding ethos on Beehaw: Be(e) Nice. Each user you encounter here is a person, and should be treated with kindness (even if they’re wrong, or use a Linux distro you don’t like). Personal attacks will not be tolerated.

Subcommunities on Beehaw:


This community's icon was made by Aaron Schneider, under the CC-BY-NC-SA 4.0 license.

founded 2 years ago
MODERATORS