Technology

59534 readers

3195 users here now

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related content.
Be excellent to each another!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, to ask if your bot can be added please contact us.
Check for duplicates before posting, duplicates may be removed

Approved Bots

founded 1 year ago

MODERATORS

Asking ChatGPT to Repeat Words ‘Forever’ Is Now a Terms of Service Violation (www.404media.co)

submitted 11 months ago by misk@sopuli.xyz to c/technology@lemmy.world

21 comments fedilink hide all child comments

you are viewing a single comment's thread
view the rest of the comments

[–] Sibbo@sopuli.xyz 1 points 11 months ago (1 children)

How can the training data be sensitive, if noone ever agreed to give their sensitive data to OpenAI?

[–] TWeaK@lemm.ee 1 points 11 months ago (1 children)

Exactly this. And how can an AI which "doesn't have the source material" in its database be able to recall such information?

[–] Jordan117@lemmy.world 1 points 11 months ago

IIRC based on the source paper the "verbatim" text is common stuff like legal boilerplate, shared code snippets, book jacket blurbs, alphabetical lists of countries, and other text repeated countless times across the web. It's the text equivalent of DALL-E "memorizing" a meme template or a stock image -- it doesn't mean all or even most of the training data is stored within the model, just that certain pieces of highly duplicated data have ascended to the level of concept and can be reproduced under unusual circumstances.