News: 1653419121

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Google says it would release its photorealistic DALL-E 2 rival – but this AI is too prejudiced for you to use

(2022/05/24)


DALL·E 2 may have to cede its throne as the most impressive image-generating AI to Google, which has revealed its own text-to-image model called Imagen.

Like OpenAI's DALL·E 2, Google's system outputs images of stuff based on written prompts from users. Ask it for a vulture flying off with a laptop in its claws and you'll perhaps get just that, all generated on the fly.

A quick glance at Imagen's [1]website shows off some of the pictures it's created (and Google has carefully curated), such as a blue jay perched on a pile of macarons, a robot couple enjoying wine in front of the Eiffel Tower, or Imagen's own name sprouting from a book. According to the team, "human raters exceedingly prefer Imagen over all other models in both image-text alignment and image fidelity," but they would say that, wouldn't they.

[2]

Imagen comes from Google Research's Brain Team, who claim the AI achieved an unprecedented level of photorealism thanks to a combination of transformer and image diffusion models. When tested against similar models, such as [3]DALL·E 2 and VQ-GAN+CLIP, the team said Imagen blew the lot out of the water. DrawBench, a list of 200 prompts used to benchmark the models, was built in-house.

Imagen's work, with prompts ... Source: Google

Imagen's designers say that their key breakthrough was in the training stage of their model. Their work, the team said, shows how effective large, frozen pre-trained language models can be as text encoders. Scaling that language model, they found, had far more impact on performance than scaling Imagen's other components.

"Our observation … encourages future research directions on exploring even bigger language models as text encoders," the team wrote.

[4]

[5]

Unfortunately for those hoping to take a crack at Imagen, the team that created it said it isn't releasing its code nor a public demo, for several reasons.

For example, Imagen isn't good at generating human faces. In experiments with pictures including human faces, Imagen only received a 39.2 percent preference from human raters over reference images. When human faces were removed, that number jumped to 43.9 percent.

[6]

Unfortunately, Google didn't provide any Imagen-generated human pictures, so it's impossible to tell how they compare to those generated by platforms like [7]This Person Does Not Exist , which uses a general adversarial network to generate faces.

Aside from technical concerns, and more importantly, Imagen's creators found that it's a bit racist and sexist even though they tried to prevent such biases.

[8]OpenAI's DALL·E 2 generates AI images that are sometimes biased or NSFW

[9]Fake it until you make it: Can synthetic data help train your AI model?

[10]1,000-plus AI-generated LinkedIn faces uncovered

[11]AI really can't copyright the art it generates – US officials

Imagen showed "an overall bias towards generating images of people with lighter skin tones and … portraying different professions to align with Western gender stereotypes," the team wrote. Eliminating humans didn't help much, either: "Imagen encodes a range of social and cultural biases when generating images of activities, events and objects."

Like similar AIs, Imagen was trained on image-text pairs scraped from the internet into publicly available datasets like COCO and LAION-400M. The Imagen team said it filtered a subset of the data to remove noise and offensive content, though an audit of the LAION dataset "uncovered a wide range of inappropriate content including pornographic imagery, racist slurs, and harmful social stereotypes."

Bias in machine learning is a well-known issue: Twitter's [12]image cropping and Google's [13]computer vision are just a couple that have been singled out for playing into stereotypes that are coded into the data we produce.

[14]

"There are a multitude of data challenges that must be addressed before text-to-image models like Imagen can be safely integrated into user-facing applications … We strongly caution against the use of text-to-image generation methods for any user-facing tools without close care and attention to the contents of the training dataset," Imagen's creators said. ®

Get our [15]Tech Resources



[1] https://imagen.research.google/

[2] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=2&c=2Yo1VgiSn3WubkIkuGvmfMAAAAFc&t=ct%3Dns%26unitnum%3D2%26raptor%3Dcondor%26pos%3Dtop%26test%3D0

[3] https://www.theregister.com/2022/04/07/openai_dalle2_ai/

[4] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Yo1VgiSn3WubkIkuGvmfMAAAAFc&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[5] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Yo1VgiSn3WubkIkuGvmfMAAAAFc&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[6] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=4&c=44Yo1VgiSn3WubkIkuGvmfMAAAAFc&t=ct%3Dns%26unitnum%3D4%26raptor%3Dfalcon%26pos%3Dmid%26test%3D0

[7] https://this-person-does-not-exist.com/en

[8] https://www.theregister.com/2022/05/08/in_brief_ai/

[9] https://www.theregister.com/2022/04/18/fake_ai_data/

[10] https://www.theregister.com/2022/03/28/ai_fake_linkedin_faces/

[11] https://www.theregister.com/2022/02/22/ai_art_copyright/

[12] https://www.theregister.com/2021/08/11/defcon_twitter_ai/

[13] https://www.theregister.com/2020/04/13/ai_roundup/

[14] https://pubads.g.doubleclick.net/gampad/jump?co=1&iu=/6978/reg_software/aiml&sz=300x50%7C300x100%7C300x250%7C300x251%7C300x252%7C300x600%7C300x601&tile=3&c=33Yo1VgiSn3WubkIkuGvmfMAAAAFc&t=ct%3Dns%26unitnum%3D3%26raptor%3Deagle%26pos%3Dmid%26test%3D0

[15] https://whitepapers.theregister.com/



Not the usual tropes again...

Juillen 1

If you look for racism and sexism, you'll find it. Even if it's not there.

What the training set has done is find the general composition of the set, and render that. If that's what the contribution to the set is, that's what'll be reflected.

I'm always confounded by AI workers coming up with these exclamations after exposing learning mechanisms to a representative set of data and it doesn't come up with what they want. If you want to teach something to come up with answers you want, you need to put the effort into creating a curated set of information for it to learn from (this is something that every species on the planet has learned long ago, which is why they survive and in cases such as some Birds and Simians, develop their own actual cultures.

But, when you select only what you want to see, then you have to understand that it's a product of your own biases. Once you take off the limiters and let it see what's out there, it'll learn things you'd rather it didn't.

There are, of course, confounding issues (like contrast for some human gene adaptations, which make some faces harder to recognise for example), but those are technical hurdles which you can control for and work with (it's part of learning how to teach a learning system).

Evolution

Anonymous Coward

Is all about that for a given situation, some choices that are better than others.

All it would have taken for you to not be here is for any one of your ancestors, all the way back through your lineage to the first "living" organism, to have made a wrong choice or have been in the wrong place at the wrong time before it procreated the next chain in your lineage.

Re: Not the usual tropes again...

theOtherJT

In other words "Damn computer, seeing what's there rather than what we wish was there!"

That's the uncomfortable thing with AI. It really doesn't care who it upsets. It doesn't have the concept of upsetting. Or who, if it comes to it.

You get the hard cold facts of your dataset.

uncovered a wide range of inappropriate content

Ken Moorhouse

Maybe this is the epitome of AI, because this is what people think, at the rawest, most basic level.

Convolve the Bible, Qur'an, etc. into the mix and this would help rectify the "rules", to make the output more palatable for "family" consumption, but it would still be limited by who the human beings are that feed it. It also depends how many texts are fed into the mix (the etc. I mentioned at the beginning of this para).

Re: uncovered a wide range of inappropriate content

Anonymous Coward

So if AI is going to follow the rules it's going to be very interesting in America. For example the Republicans bill to keep transgender women and girls in Louisiana from competing in college and K-12 athletic teams has just won final legislative passage ... so the only way for AI to verify that transgender students are not competing will be to discard all clothing.

I tried the clone syscall on me, but it didn't work.
-- Mike Neuffer trying to fix a serious time problem