Technology

83500 readers

3690 users here now

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related news or articles.
Be excellent to each other!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
Check for duplicates before posting, duplicates may be removed
Accounts 7 days and younger will have their posts automatically removed.

Approved Bots

founded 2 years ago

MODERATORS

L3s@lemmy.world

enu@lemmy.world

technopagan@lemmy.world

L4s@lemmy.world

L3s@hackingne.ws

837

Meta admits using pirated books to train AI, but won't pay for it (www.techspot.com)

submitted 2 years ago by throws_lemy@lemmy.nz to c/technology@lemmy.world

164 comments fedilink hide all child comments

you are viewing a single comment's thread
view the rest of the comments

[–] archomrade@midwest.social 2 points 2 years ago (1 children)

Copyright is absolute. The rightsholder has complete and total right to dictate how it is copied.

Really and truly, this is not how this works. The exemptions granted by the office of the registrar are granting an exemption to copyright claims against fair uses. It isn't talking about whether the claim can be awarded damages, it's talking about the claim being exempt in entirety. You can think about copyright as an exemption to the first amendment right to free speech, and the exemption to copyright as describing where that 'right' does not apply. Copyright holders do not get to control the use of their work where fair use has been determined by the registrar, which is reconsidered every 3 years.

This is all just pedantry, though, and has no practical significance. Saying “fair use means copyright has not been infringed” doesn’t change anything.

True enough, but it seems like it's important for your understanding in how copyright works.

Or perhaps rather some kind of 3D array - which could just be considered an advanced form of database. But yeah, you’re right here, you win this pedantry round lol. 1-1.

I wasn't being pedantic, that distinction is important for how copyright is conceptualized. The AI model is the thing being considered for infringement, so it's important to note that the works being claimed within it do not exist as such within the model. The '3-d array' does not contain copyrighted works. You can think of it as a large metadata file, describing how to construct language as analyzed through the training data. The nature and purpose of the 'work' is night-and-day different from the works being claimed, and 'database' is a clear misrepresentation (possibly even intentionally so) of what it is.

Yeah I don’t want to go down the avenue of suing the AI itself for infringement.

That was exactly what you pivoted to in your comment here, i'm not sure why you're now saying you don't want to go down that avenue. I'm confused what you're arguing at this point.

Although, I was much happier replying to you before I just saw the downvotes you’ve apparently given me across the board. That’s a bit poor behaviour on your part, you shouldn’t downvote just because you disagree - and you can’t even say that I’m wrong as a justification when the whole thing is being heavily debated and adjudicated over whether it is right or wrong.

I've down-voted your comments because they contain inaccuracies and could be misleading to others. You shouldn't let my grading of your comments reflect my attitude towards you; i'm sure you're a fine individual. Downvotes don't mean anything on Lemmy anyway, i'm not sure 'spitting in your face' is a fair or accurate description, but I don't want to invalidate your feelings, so I apologize for making you feel that way as that wasn't my intent.

[–] TWeaK@lemm.ee 1 points 2 years ago

I’ve down-voted your comments because they contain inaccuracies and could be misleading to others. You shouldn’t let my grading of your comments reflect my attitude towards you; i’m sure you’re a fine individual. Downvotes don’t mean anything on Lemmy anyway, i’m not sure ‘spitting in your face’ is a fair or accurate description, but I don’t want to invalidate your feelings, so I apologize for making you feel that way as that wasn’t my intent.

No worries, you've been very respectable. My feelings weren't particularly hurt, I just felt the need to call it out.

Personally, I'm against downvoting things merely because they are wrong. If someone says something that's wrong, it may well be a commonly held misconception, and downvoting it also demotes any correction that has been given, which means other people who hold the misconception are less likely to be corrected.

And that's beside the fact that I don't really think I'm completely wrong here :o)

That was exactly what you pivoted to in your comment here, i’m not sure why you’re now saying you don’t want to go down that avenue. I’m confused what you’re arguing at this point.

To be a little more specific, I don't want to go down the route of blaming AI itself for copyright infringement. That is to say, whether or not AI is bound by laws the way that humans are. I think it is only worthwhile considering whether the AI developer and/or the users are infringing copyright through their creation or use of AI. In particular, I think the legal or philosophical question of whether AI is affected by laws in the same way humans are is pointless when we're just talking about LLM's and not a true Artificial Intelligence.

The ‘3-d array’ does not contain copyrighted works. You can think of it as a large metadata file, describing how to construct language as analyzed through the training data. The nature and purpose of the ‘work’ is night-and-day different from the works being claimed, and ‘database’ is a clear misrepresentation (possibly even intentionally so) of what it is.

Yes absolutely, the LLM itself does not include copyrighted works. That's not what I'm arguing. The two issues I take are with the database of information the LLM is trained on. This database does contain copyrighted works, AI developers admit that it does, but they claim it is fair use research. I disagree with this claim, to use the terminology from one of your links their "research" is not "scholarly" - it is commercial product development.

The other issue is that the LLM can reproduce copyrighted work. While I agree with you in some sense that the user of the LLM is instructing it to infringe copyright, and thus the user is responsible, in another sense I think the developer is also responsible because they have given the tool the capability to do this. This is perhaps not a strong argument, particularly when the developers have made efforts to fix these "bugs" as they come to light.

However my most important point is that the developers have infringed copyright by building a training database full of copyrighted works, which the LLM was then trained on. The LLM itself isn't copyright infringement, but they infringed copyright to develop it.