Google unveils its competitor to OpenAI’s text-to-image modelLeigh Mc Gowranon May 24, 2022 at 12:25 Silicon RepublicSilicon Republic

0

Google Research has created a competitor for OpenAI’s text-to-image system, with its own AI model that works using a similar diffusion method called Imagen.

Google’s research team said Imagen is a text-to-image model that has an “unprecedented degree of photorealism” and a deep level of language understanding.

Text-to-image AI models are able to understand the relationship between an image and the words used to describe it. Once a description is added, these systems can create multiple images based on how it interprets the text, combining different concepts, attributes and styles.

For example, if a certain image is designed as “a photo of”, it can be altered into something like “an oil painting” while retaining all the other key description words used to make the original image. Imagen’s team has shared a number of example images that the AI model has created.

OpenAI created the first version of its text-to-image model called DALL-E last year, but unveiled the improved model called DALL-E 2 last month which “generates more realistic and accurate images with four times greater resolution”.

In their research paper, the team behind Imagen suggests that scaling the pretrained text encoder size is more important than scaling the diffusion model size.

The research team also said that large pretrained frozen text encoders are very effective for the text-to-image task. As a result, the team said Imagen scored highly on the common objects in context dataset without ever being trained on it.

Google’s research team said it has also created a benchmark tool to assess and compare different text-to-image models called DrawBench.

When using their DrawBench model, Google’s team said human raters preferred Imagen over other models such as DALL-E 2 in side-by-side comparisons “both in terms of sample quality and image-text alignment”.

Concerns of misuse

Similarly to Open-AI, Google Research said there are several ethical challenges to be considered with text-to image research. The team said these models can affect society in “complex ways” and that the risk of misuse raises concerns in terms of creating open-source code and demos.

“The data requirements of text-to-image models have led researchers to rely heavily on large, mostly uncurated, web-scraped datasets,” Google Research said. “While this approach has enabled rapid algorithmic advances in recent years, datasets of this nature often reflect social stereotypes, oppressive viewpoints, and derogatory, or otherwise harmful, associations to marginalised identity groups.”

Google Research also said that its preliminary analysis of Imagen suggests that the model encodes a range of “social and cultural biases” when making images of activities, events and objects.

“We aim to make progress on several of these open challenges and limitations in future work,” Google Research said.

When Open-AI unveiled DALL-E 2 last month, concerns were raised that this sort of technology could help people spread disinformation online through the use of authentic-looking fake images.

10 things you need to know direct to your inbox every weekday. Sign up for the Daily Brief, Silicon Republic’s digest of essential sci-tech news.

The post Google unveils its competitor to OpenAI’s text-to-image model appeared first on Silicon Republic.

Leave a Comment