Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Adobe, Microsoft and Google have all done this now.

Thats ~ 4 trillion dollars of companies betting that the law will say anyone may train an AI model on any public data, and anyone may use the output of that AI without compensating owners of the training data.

When 4 trillion dollars is at stake, not only do you put the best lawyers on the case, but you also pay congress to change the law if things aren't heading your way.

I'm pretty sure now that the debate of AI ownership is a foregone conclusion - nobody owns AI outputs.



>nobody owns AI outputs

This would be fantastic imo. A new era of the commons.

>Adobe

I disagree here. Adobe has trained only on public domain and their own stock images. So why would adobe be against training on unlicensed data being an infringement? It would eliminate much of their competition...


> Adobe has trained only on public domain and their own stock images.

Adobe is lying. They are relying on general ignorance about the technology to get away with it.

Adobe has not shown how they train the text encoders in Firefly, or what images were used for the text-based conditioning (i.e. "text to image") part of their image generation model. They are almost certainly using CLIP or T5, which are trained on LAION2b, an image dataset with the very problems they are trying to address, C4 (a text dataset similarly encumbered) and similar.

bUt nO oNe eLsE hAs bRoUgHt tHiS uP. It's so arcane for non-practitioners. Talk about this directly with someone like Astropulse, who monetizes a Stable Diffusion model: no confusion, totally agrees with me. By comparison, I've pinged the Ars Technica journalist who just wrote about this issue: crickets. Posted to the Adobe forum: crickets. E-mailed them on their specific address for this: crickets. I have no idea why something so obvious has slipped by everyone's radar!


Would it be impossible to train their own text encoder on just the images they have? How many would one need?


I welcome anyone who works at Adobe to simply answer this question and put it to rest. There is absolutely nothing sensitive about the issue, unless it exposes them in a lie.

So no chance. I think it's a big fat lie. They'd have to have made some other scientific breakthrough, which they didn't.

Using information from https://openai.com/research/clip and https://github.com/mlfoundations/open_clip, it's possible to answer this question.

It's certainly not impossible, but it's impracticable. On 248m images (roughly the size of Adobe Stock), CLIP gets 37% on ImageNet, and on the 2000m from LAION, it performs 71-80%. And even with 2000m images, CLIP is substantially worse performing than the approach that Imagen uses for "text comprehension," which relies on essentially many billions more images and text tokens.


Interesting. I looked through the laion Datasets a bit and it was astonishing how bad the captions really are. Very very short captions if not completely wrong. Amazing to me that this even works at all. I wonder how much better clip etc would perform and be more efficient if they had probably tagged images, not just with the alt text. Maybe that's why dalle 3 is so good at following the prompts?


> I'm pretty sure now that the debate of AI ownership is a foregone conclusion - nobody owns AI outputs.

but at the same time, they put ToS that you may not train a new LLM using the output of their LLM...


Classic case of wanting their cake and eating it too. Although I don't think they'll be surprised if their TOS doesn't hold up in court either.


IANAL, but I believe that wouldn't just be simple copyright infringement, but a breach of contract.


The workaround is to use an intermediary so that you don’t have any contractual obligations to breach, and the intermediary stays far away from your downstream use, so they never breached the agreement either.


I would refine that:

Right now, I think it's more that they don't want US v TeensyStartup to be the case that sets precedent.

By stepping in with these indemnification clauses, they aren't betting $4T that they're sure to win. They're just reserving a (much smaller) open check to protect against losing because of somebody else's lawyers.

They may win, they may lose, but they want to make sure they're the ones who get to fight for it either way.


They could step in on a case by case basis even without this indemnification.


Who will be the first to type in "Mickey Mouse, digital art" and watch these companies take on $DIS?


"this indemnity only applies if you didn’t try to intentionally create or use generated output to infringe the rights of others"


If you somehow accidentally generate Mickey Mouse and decide to monetize it, how are you going to defend yourself that your prompt didn't include "Mickey Mouse" or something?


If it gets to in front of an actual judge, they would likely conclude that the average person should have been able to recognize the characteristic image of Mickey Mouse. So they won't even bother with asking whether the prompt actually contained those words or not.


Yeah, I guess that works for characters like Mickey Mouse. But for every Mickey Mouse there are hundreds or thousands of other characters that are less widely known.


Those very likely won't get in front of an actual judge, since the damages would be small enough to make a long court battle pointless, so the question would be moot.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: