
Google has been in the tech industry for decades, and has also ventured into developing AIs to not only improve its products, but to also experiment with it.
For more than many times, researchers have showed how AIs can make computers increasingly smarter. With AIs, computers can be trained to understand the things beyond their own programming, in many ways that can even surprise their own creators.
Among the many things AIs can do, is creating images based on text input.
The idea here, is allowing users to give it any descriptive text, and let the AI turn that into an image.
OpenAI, for example, has what it calls the 'DALL·E', which is essentially the 'GPT' for images. And building on that AI, OpenAI has announced its successor, the 'DALL·E 2', and later the 'DALL·E 3'.
Google is taking on the OpenAI's product (as well as some other text-to-image generators) using what it calls 'Imagen'.
And this time, Google has finally released Imagen 3 to the public, initially to those on the U.S. through its ImageFX platform.
We’re introducing Imagen 3: our highest quality text-to-image generation model yet.
It produces visuals with incredible detail, realistic lighting and fewer distracting artifacts.
From quick sketches to very high-res imagery, here’s a look at what it can create. #GoogleIO pic.twitter.com/XMrQYGeSiO— Google DeepMind (@GoogleDeepMind) May 14, 2024
Imagen 3 was first announced back in May 2024, during Google I/O.
Then, in June, Google only released it to to select Vertex AI users.
This time, with Imagen 3 available to the public, users can clearly see how this AI tool is versatile and powerful.
Describing it, Google said that Imagen 3 is "a latent diffusion model that generates high quality images from text prompts."
In this case, not only that the upgraded AI image generator can create images based on user prompts, because users can also edit the image by highlighting an area and instructing the AI to make the desired changes.
Imagen 3 can generate high-quality visuals in a wide range of styles- from photorealistic landscapes to richly textured oil paintings or whimsical claymation scenes.
https://t.co/T7Iv8unW8Q— Google DeepMind (@GoogleDeepMind) May 16, 2024
"We describe our quality and responsibility evaluations. Imagen 3 is preferred over other state-of-the-art (SOTA) models at the time of evaluation."
"Imagen 3 is our highest quality text-to-image model. It generates an incredible level of detail, producing photorealistic, lifelike images with far fewer distracting visual artifacts than our prior models," said Google.
In a research paper (PDF) explaining the technology, Imagen is said to be able to understand descriptive prompts, such as:
"Photo of a felt puppet diorama scene of a tranquil nature scene of a secluded forest clearing with a large friendly, rounded robot is rendered in a risograph style. An owl sits on the robots shoulders and a fox at its feet. Soft washes of color, 5 color, and a light-filled palette create a sense of peace and serenity, inviting contemplation and the appreciation of natural beauty."
The tool has some restrictions.
According to the research paper, Google implements a multi-stage safety and quality filtering process in order to employ data cleaning and filtering methods in line with Google’s policies.

These methods include the removal of unsafe, violent, or low-quality images; the removal of AI-generated images to prevent the model from learning artifacts or biases that may be found in AI-generated mages, downranking similar images to minimize the risk of outputs overfitting training, synthetic captioning, and the filtering of unsafe captions.
A suite of evaluations are then made across the end-to-end lifecycle of model development and deployment.
They include the evaluation to prevent violence, hate, explicit sexualization, and over-sexualization.
The team also tries to ensure that the AI cannot be used to generating convincing-looking images of public figures or weapons.
In all, Imagen 3 is unlike Elon Musk's unhinged Grok-2, which can be made to generate images of politicians doing what they're not supposed to, and also generate copyrighted items.
But again, the nature of generative AI means that users may still be able to find workarounds to bypass this limitation by providing carefully prompts, for example.
















































































































































































































































































































































































