You can ask a Gemini Image model to generate and edit images using both text-only and text-and-file prompts. When you use Firebase AI Logic , you can make this request directly from your app.
Благодаря этой возможности вы можете делать, например, следующее:
Iteratively generate images through conversation with natural language, adjusting images while maintaining consistency and context.
Generate images with high-quality text rendering, including long strings of text.
Generate interleaved text-image output. For example, a blog post with text and images in a single turn. Previously, this required stringing together multiple models.
Generate images using Gemini's world knowledge and reasoning capabilities.
Вы можете настроить способ ответа для вашего запроса таким образом, чтобы модель генерировала только изображения или и изображения, и текст. Полный список поддерживаемых функций (вместе с примерами запросов) вы найдете далее на этой странице.
Jump to code for text-to-image Jump to code for interleaved text & images
Jump to code for image editing Jump to code for iterative image editing
| See other guides for additional options for working with images Analyze images Analyze images on-device Generate structured output |
Прежде чем начать
Click your Gemini API provider to view provider-specific content and code on this page. |
Если вы еще этого не сделали, пройдите руководство по началу работы , в котором описывается, как настроить проект Firebase, подключить приложение к Firebase, добавить SDK, инициализировать бэкэнд-сервис для выбранного вами поставщика API Gemini и создать экземпляр GenerativeModel .
Модели, поддерживающие эту возможность
-
gemini-3-pro-image(также известный как "Nano Banana Pro") -
gemini-3.1-flash-image(aka "Nano Banana 2") -
gemini-3.1-flash-lite-image(aka "Nano Banana 2 Lite")
Image-generating Gemini 2.5 models support this capability, but they're all deprecated.
Создание и редактирование изображений
You can generate and edit images using a Gemini model.
Создание изображений (ввод только текста)
| Прежде чем опробовать этот пример, выполните раздел «Перед началом работы » этого руководства, чтобы настроить свой проект и приложение. In that section, you'll also click a button for your chosen Gemini API provider so that you see provider-specific content on this page . |
You can ask a Gemini Image model to generate images by prompting with text.
Create a GenerativeModel instance, include the response modality of IMAGE in your model configuration, and call generateContent .
Быстрый
import FirebaseAILogic
// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a Gemini model that supports image output.
let generativeModel = FirebaseAI.firebaseAI(backend: .googleAI()).generativeModel(
modelName: "gemini-3.1-flash-image",
// Configure the model to respond with images only.
generationConfig: GenerationConfig(responseModalities: [.image])
)
// Provide a text prompt instructing the model to generate an image
let prompt = "Generate an image of the Eiffel tower with fireworks in the background."
// To generate an image, call `generateContent` with the text input
let response = try await model.generateContent(prompt)
// Handle the generated image
guard let inlineDataPart = response.inlineDataParts.first else {
fatalError("No image data in response.")
}
guard let uiImage = UIImage(data: inlineDataPart.data) else {
fatalError("Failed to convert data to UIImage.")
}
Kotlin
// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a Gemini model that supports image output.
val model = Firebase.ai(backend = GenerativeBackend.googleAI()).generativeModel(
modelName = "gemini-3.1-flash-image",
// Configure the model to respond with images only.
generationConfig = generationConfig {
responseModalities = listOf(ResponseModality.IMAGE) }
)
// Provide a text prompt instructing the model to generate an image
val prompt = "Generate an image of the Eiffel tower with fireworks in the background."
// To generate image output, call `generateContent` with the text input
val generatedImageAsBitmap = model.generateContent(prompt)
// Handle the generated image
.candidates.first().content.parts.filterIsInstance<ImagePart>().firstOrNull()?.image
Java
// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a Gemini model that supports image output.
GenerativeModel ai = FirebaseAI.getInstance(GenerativeBackend.googleAI()).generativeModel(
"gemini-3.1-flash-image",
// Configure the model to respond with images only.
new GenerationConfig.Builder()
.setResponseModalities(Arrays.asList(ResponseModality.IMAGE))
.build()
);
GenerativeModelFutures model = GenerativeModelFutures.from(ai);
// Provide a text prompt instructing the model to generate an image
Content prompt = new Content.Builder()
.addText("Generate an image of the Eiffel Tower with fireworks in the background.")
.build();
// To generate an image, call `generateContent` with the text input
ListenableFuture<GenerateContentResponse> response = model.generateContent(prompt);
Futures.addCallback(response, new FutureCallback<GenerateContentResponse>() {
@Override
public void onSuccess(GenerateContentResponse result) {
// iterate over all the parts in the first candidate in the result object
for (Part part : result.getCandidates().get(0).getContent().getParts()) {
if (part instanceof ImagePart) {
ImagePart imagePart = (ImagePart) part;
// The returned image as a bitmap
Bitmap generatedImageAsBitmap = imagePart.getImage();
break;
}
}
}
@Override
public void onFailure(Throwable t) {
t.printStackTrace();
}
}, executor);
Web
import { initializeApp } from "firebase/app";
import { getAI, getGenerativeModel, GoogleAIBackend, ResponseModality } from "firebase/ai";
// TODO(developer): Replace the following with your app's Firebase configuration
// See: https://firebase.google.com/docs/web/learn-more#config-object
const firebaseConfig = {
// ...
};
// Initialize FirebaseApp
const firebaseApp = initializeApp(firebaseConfig);
// Initialize the Gemini Developer API backend service.
const ai = getAI(firebaseApp, { backend: new GoogleAIBackend() });
// Create a `GenerativeModel` instance with a model that supports your use case
const model = getGenerativeModel(ai, {
model: "gemini-3.1-flash-image",
// Configure the model to respond with images only.
generationConfig: {
responseModalities: [ResponseModality.IMAGE],
},
});
// Provide a text prompt instructing the model to generate an image
const prompt = 'Generate an image of the Eiffel Tower with fireworks in the background.';
// To generate an image, call `generateContent` with the text input
const result = model.generateContent(prompt);
// Handle the generated image
try {
const inlineDataParts = result.response.inlineDataParts();
if (inlineDataParts?.[0]) {
const image = inlineDataParts[0].inlineData;
console.log(image.mimeType, image.data);
}
} catch (err) {
console.error('Prompt or candidate was blocked:', err);
}
Dart
import 'package:firebase_ai/firebase_ai.dart';
import 'package:firebase_core/firebase_core.dart';
import 'firebase_options.dart';
await Firebase.initializeApp(
options: DefaultFirebaseOptions.currentPlatform,
);
// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a Gemini model that supports image output.
final model = FirebaseAI.googleAI().generativeModel(
model: 'gemini-3.1-flash-image',
// Configure the model to respond with images only.
generationConfig: GenerationConfig(responseModalities: [ResponseModalities.image]),
);
// Provide a text prompt instructing the model to generate an image
final prompt = [Content.text('Generate an image of the Eiffel Tower with fireworks in the background.')];
// To generate an image, call `generateContent` with the text input
final response = await model.generateContent(prompt);
if (response.inlineDataParts.isNotEmpty) {
final imageBytes = response.inlineDataParts[0].bytes;
// Process the image
} else {
// Handle the case where no images were generated
print('Error: No images were generated.');
}
Единство
using Firebase;
using Firebase.AI;
// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a Gemini model that supports image output.
var model = FirebaseAI.GetInstance(FirebaseAI.Backend.GoogleAI()).GetGenerativeModel(
modelName: "gemini-3.1-flash-image",
// Configure the model to respond with images only.
generationConfig: new GenerationConfig(
responseModalities: new[] { ResponseModality.Image })
);
// Provide a text prompt instructing the model to generate an image
var prompt = "Generate an image of the Eiffel Tower with fireworks in the background.";
// To generate an image, call `GenerateContentAsync` with the text input
var response = await model.GenerateContentAsync(prompt);
var text = response.Text;
if (!string.IsNullOrWhiteSpace(text)) {
// Do something with the text
}
// Handle the generated image
var imageParts = response.Candidates.First().Content.Parts
.OfType<ModelContent.InlineDataPart>()
.Where(part => part.MimeType == "image/png");
foreach (var imagePart in imageParts) {
// Load the Image into a Unity Texture2D object
UnityEngine.Texture2D texture2D = new(2, 2);
if (texture2D.LoadImage(imagePart.Data.ToArray())) {
// Do something with the image
}
}
Создание чередующихся изображений и текста
| Прежде чем опробовать этот пример, выполните раздел «Перед началом работы » этого руководства, чтобы настроить свой проект и приложение. In that section, you'll also click a button for your chosen Gemini API provider so that you see provider-specific content on this page . |
Вы можете попросить модель Gemini Image сгенерировать изображения, чередующиеся с текстовыми ответами. Например, вы можете сгенерировать изображения того, как может выглядеть каждый шаг сгенерированного рецепта, вместе с инструкциями к этому шагу, и вам не нужно будет отправлять отдельные запросы модели или разным моделям.
Создайте экземпляр GenerativeModel , укажите в конфигурации модели варианты ответа TEXT и IMAGE и вызовите generateContent .
Быстрый
import FirebaseAILogic
// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a Gemini model that supports image output.
let generativeModel = FirebaseAI.firebaseAI(backend: .googleAI()).generativeModel(
modelName: "gemini-3.1-flash-image",
// Configure the model to respond with text and images.
generationConfig: GenerationConfig(responseModalities: [.text, .image])
)
// Provide a text prompt instructing the model to generate interleaved text and images
let prompt = """
Generate an illustrated recipe for a paella.
Create images to go alongside the text as you generate the recipe
"""
// To generate interleaved text and images, call `generateContent` with the text input
let response = try await model.generateContent(prompt)
// Handle the generated text and image
guard let candidate = response.candidates.first else {
fatalError("No candidates in response.")
}
for part in candidate.content.parts {
switch part {
case let textPart as TextPart:
// Do something with the generated text
let text = textPart.text
case let inlineDataPart as InlineDataPart:
// Do something with the generated image
guard let uiImage = UIImage(data: inlineDataPart.data) else {
fatalError("Failed to convert data to UIImage.")
}
default:
fatalError("Unsupported part type: \(part)")
}
}
Kotlin
// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a Gemini model that supports image output.
val model = Firebase.ai(backend = GenerativeBackend.googleAI()).generativeModel(
modelName = "gemini-3.1-flash-image",
// Configure the model to respond with text and images.
generationConfig = generationConfig {
responseModalities = listOf(ResponseModality.TEXT, ResponseModality.IMAGE) }
)
// Provide a text prompt instructing the model to generate interleaved text and images
val prompt = """
Generate an illustrated recipe for a paella.
Create images to go alongside the text as you generate the recipe
""".trimIndent()
// To generate interleaved text and images, call `generateContent` with the text input
val responseContent = model.generateContent(prompt).candidates.first().content
// The response will contain image and text parts interleaved
for (part in responseContent.parts) {
when (part) {
is ImagePart -> {
// ImagePart as a bitmap
val generatedImageAsBitmap: Bitmap? = part.asImageOrNull()
}
is TextPart -> {
// Text content from the TextPart
val text = part.text
}
}
}
Java
// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a Gemini model that supports image output.
GenerativeModel ai = FirebaseAI.getInstance(GenerativeBackend.googleAI()).generativeModel(
"gemini-3.1-flash-image",
// Configure the model to respond with text and images.
new GenerationConfig.Builder()
.setResponseModalities(Arrays.asList(ResponseModality.TEXT, ResponseModality.IMAGE))
.build()
);
GenerativeModelFutures model = GenerativeModelFutures.from(ai);
// Provide a text prompt instructing the model to generate interleaved text and images
Content prompt = new Content.Builder()
.addText("Generate an illustrated recipe for a paella.\n" +
"Create images to go alongside the text as you generate the recipe")
.build();
// To generate interleaved text and images, call `generateContent` with the text input
ListenableFuture<GenerateContentResponse> response = model.generateContent(prompt);
Futures.addCallback(response, new FutureCallback<GenerateContentResponse>() {
@Override
public void onSuccess(GenerateContentResponse result) {
Content responseContent = result.getCandidates().get(0).getContent();
// The response will contain image and text parts interleaved
for (Part part : responseContent.getParts()) {
if (part instanceof ImagePart) {
// ImagePart as a bitmap
Bitmap generatedImageAsBitmap = ((ImagePart) part).getImage();
} else if (part instanceof TextPart){
// Text content from the TextPart
String text = ((TextPart) part).getText();
}
}
}
@Override
public void onFailure(Throwable t) {
System.err.println(t);
}
}, executor);
Web
import { initializeApp } from "firebase/app";
import { getAI, getGenerativeModel, GoogleAIBackend, ResponseModality } from "firebase/ai";
// TODO(developer): Replace the following with your app's Firebase configuration
// See: https://firebase.google.com/docs/web/learn-more#config-object
const firebaseConfig = {
// ...
};
// Initialize FirebaseApp
const firebaseApp = initializeApp(firebaseConfig);
// Initialize the Gemini Developer API backend service.
const ai = getAI(firebaseApp, { backend: new GoogleAIBackend() });
// Create a `GenerativeModel` instance with a model that supports your use case
const model = getGenerativeModel(ai, {
model: "gemini-3.1-flash-image",
// Configure the model to respond with text and images.
generationConfig: {
responseModalities: [ResponseModality.TEXT, ResponseModality.IMAGE],
},
});
// Provide a text prompt instructing the model to generate interleaved text and images
const prompt = 'Generate an illustrated recipe for a paella.\n.' +
'Create images to go alongside the text as you generate the recipe';
// To generate interleaved text and images, call `generateContent` with the text input
const result = await model.generateContent(prompt);
// Handle the generated text and image
try {
const response = result.response;
if (response.candidates?.[0].content?.parts) {
for (const part of response.candidates?.[0].content?.parts) {
if (part.text) {
// Do something with the text
console.log(part.text)
}
if (part.inlineData) {
// Do something with the image
const image = part.inlineData;
console.log(image.mimeType, image.data);
}
}
}
} catch (err) {
console.error('Prompt or candidate was blocked:', err);
}
Dart
import 'package:firebase_ai/firebase_ai.dart';
import 'package:firebase_core/firebase_core.dart';
import 'firebase_options.dart';
await Firebase.initializeApp(
options: DefaultFirebaseOptions.currentPlatform,
);
// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a Gemini model that supports image output.
final model = FirebaseAI.googleAI().generativeModel(
model: 'gemini-3.1-flash-image',
// Configure the model to respond with text and images.
generationConfig: GenerationConfig(responseModalities: [ResponseModalities.text, ResponseModalities.image]),
);
// Provide a text prompt instructing the model to generate interleaved text and images
final prompt = [Content.text(
'Generate an illustrated recipe for a paella\n ' +
'Create images to go alongside the text as you generate the recipe'
)];
// To generate interleaved text and images, call `generateContent` with the text input
final response = await model.generateContent(prompt);
// Handle the generated text and image
final parts = response.candidates.firstOrNull?.content.parts
if (parts.isNotEmpty) {
for (final part in parts) {
if (part is TextPart) {
// Do something with text part
final text = part.text
}
if (part is InlineDataPart) {
// Process image
final imageBytes = part.bytes
}
}
} else {
// Handle the case where no images were generated
print('Error: No images were generated.');
}
Единство
using Firebase;
using Firebase.AI;
// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a Gemini model that supports image output.
var model = FirebaseAI.GetInstance(FirebaseAI.Backend.GoogleAI()).GetGenerativeModel(
modelName: "gemini-3.1-flash-image",
// Configure the model to respond with text and images.
generationConfig: new GenerationConfig(
responseModalities: new[] { ResponseModality.Text, ResponseModality.Image })
);
// Provide a text prompt instructing the model to generate interleaved text and images
var prompt = "Generate an illustrated recipe for a paella \n" +
"Create images to go alongside the text as you generate the recipe";
// To generate interleaved text and images, call `GenerateContentAsync` with the text input
var response = await model.GenerateContentAsync(prompt);
// Handle the generated text and image
foreach (var part in response.Candidates.First().Content.Parts) {
if (part is ModelContent.TextPart textPart) {
if (!string.IsNullOrWhiteSpace(textPart.Text)) {
// Do something with the text
}
} else if (part is ModelContent.InlineDataPart dataPart) {
if (dataPart.MimeType == "image/png") {
// Load the Image into a Unity Texture2D object
UnityEngine.Texture2D texture2D = new(2, 2);
if (texture2D.LoadImage(dataPart.Data.ToArray())) {
// Do something with the image
}
}
}
}
Редактирование изображений (ввод текста и изображений)
| Прежде чем опробовать этот пример, выполните раздел «Перед началом работы » этого руководства, чтобы настроить свой проект и приложение. In that section, you'll also click a button for your chosen Gemini API provider so that you see provider-specific content on this page . |
You can ask a Gemini Image model to edit images by prompting with text and one or more images.
Create a GenerativeModel instance, include the response modality of IMAGE in your model configuration, and call generateContent .
Быстрый
import FirebaseAILogic
// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a Gemini model that supports image output.
let generativeModel = FirebaseAI.firebaseAI(backend: .googleAI()).generativeModel(
modelName: "gemini-3.1-flash-image",
// Configure the model to respond with images only.
generationConfig: GenerationConfig(responseModalities: [.image])
)
// Provide an image for the model to edit
guard let image = UIImage(named: "scones") else { fatalError("Image file not found.") }
// Provide a text prompt instructing the model to edit the image
let prompt = "Edit this image to make it look like a cartoon"
// To edit the image, call `generateContent` with the image and text input
let response = try await model.generateContent(image, prompt)
// Handle the generated image
guard let inlineDataPart = response.inlineDataParts.first else {
fatalError("No image data in response.")
}
guard let uiImage = UIImage(data: inlineDataPart.data) else {
fatalError("Failed to convert data to UIImage.")
}
Kotlin
// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a Gemini model that supports image output.
val model = Firebase.ai(backend = GenerativeBackend.googleAI()).generativeModel(
modelName = "gemini-3.1-flash-image",
// Configure the model to respond with images only.
generationConfig = generationConfig {
responseModalities = listOf(ResponseModality.IMAGE) }
)
// Provide an image for the model to edit
val bitmap = BitmapFactory.decodeResource(context.resources, R.drawable.scones)
// Provide a text prompt instructing the model to edit the image
val prompt = content {
image(bitmap)
text("Edit this image to make it look like a cartoon")
}
// To edit the image, call `generateContent` with the prompt (image and text input)
val generatedImageAsBitmap = model.generateContent(prompt)
// Handle the generated text and image
.candidates.first().content.parts.filterIsInstance<ImagePart>().firstOrNull()?.image
Java
// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a Gemini model that supports image output.
GenerativeModel ai = FirebaseAI.getInstance(GenerativeBackend.googleAI()).generativeModel(
"gemini-3.1-flash-image",
// Configure the model to respond with images only.
new GenerationConfig.Builder()
.setResponseModalities(Arrays.asList(ResponseModality.IMAGE))
.build()
);
GenerativeModelFutures model = GenerativeModelFutures.from(ai);
// Provide an image for the model to edit
Bitmap bitmap = BitmapFactory.decodeResource(resources, R.drawable.scones);
// Provide a text prompt instructing the model to edit the image
Content promptcontent = new Content.Builder()
.addImage(bitmap)
.addText("Edit this image to make it look like a cartoon")
.build();
// To edit the image, call `generateContent` with the prompt (image and text input)
ListenableFuture<GenerateContentResponse> response = model.generateContent(promptcontent);
Futures.addCallback(response, new FutureCallback<GenerateContentResponse>() {
@Override
public void onSuccess(GenerateContentResponse result) {
// iterate over all the parts in the first candidate in the result object
for (Part part : result.getCandidates().get(0).getContent().getParts()) {
if (part instanceof ImagePart) {
ImagePart imagePart = (ImagePart) part;
Bitmap generatedImageAsBitmap = imagePart.getImage();
break;
}
}
}
@Override
public void onFailure(Throwable t) {
t.printStackTrace();
}
}, executor);
Web
import { initializeApp } from "firebase/app";
import { getAI, getGenerativeModel, GoogleAIBackend, ResponseModality } from "firebase/ai";
// TODO(developer): Replace the following with your app's Firebase configuration
// See: https://firebase.google.com/docs/web/learn-more#config-object
const firebaseConfig = {
// ...
};
// Initialize FirebaseApp
const firebaseApp = initializeApp(firebaseConfig);
// Initialize the Gemini Developer API backend service.
const ai = getAI(firebaseApp, { backend: new GoogleAIBackend() });
// Create a `GenerativeModel` instance with a model that supports your use case
const model = getGenerativeModel(ai, {
model: "gemini-3.1-flash-image",
// Configure the model to respond with images only.
generationConfig: {
responseModalities: [ResponseModality.IMAGE],
},
});
// Prepare an image for the model to edit
async function fileToGenerativePart(file) {
const base64EncodedDataPromise = new Promise((resolve) => {
const reader = new FileReader();
reader.onloadend = () => resolve(reader.result.split(',')[1]);
reader.readAsDataURL(file);
});
return {
inlineData: { data: await base64EncodedDataPromise, mimeType: file.type },
};
}
// Provide a text prompt instructing the model to edit the image
const prompt = "Edit this image to make it look like a cartoon";
const fileInputEl = document.querySelector("input[type=file]");
const imagePart = await fileToGenerativePart(fileInputEl.files[0]);
// To edit the image, call `generateContent` with the image and text input
const result = await model.generateContent([prompt, imagePart]);
// Handle the generated image
try {
const inlineDataParts = result.response.inlineDataParts();
if (inlineDataParts?.[0]) {
const image = inlineDataParts[0].inlineData;
console.log(image.mimeType, image.data);
}
} catch (err) {
console.error('Prompt or candidate was blocked:', err);
}
Dart
import 'package:firebase_ai/firebase_ai.dart';
import 'package:firebase_core/firebase_core.dart';
import 'firebase_options.dart';
await Firebase.initializeApp(
options: DefaultFirebaseOptions.currentPlatform,
);
// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a Gemini model that supports image output.
final model = FirebaseAI.googleAI().generativeModel(
model: 'gemini-3.1-flash-image',
// Configure the model to respond with images only.
generationConfig: GenerationConfig(responseModalities: [ResponseModalities.image]),
);
// Prepare an image for the model to edit
final image = await File('scones.jpg').readAsBytes();
final imagePart = InlineDataPart('image/jpeg', image);
// Provide a text prompt instructing the model to edit the image
final prompt = TextPart("Edit this image to make it look like a cartoon");
// To edit the image, call `generateContent` with the image and text input
final response = await model.generateContent([
Content.multi([prompt,imagePart])
]);
// Handle the generated image
if (response.inlineDataParts.isNotEmpty) {
final imageBytes = response.inlineDataParts[0].bytes;
// Process the image
} else {
// Handle the case where no images were generated
print('Error: No images were generated.');
}
Единство
using Firebase;
using Firebase.AI;
// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a Gemini model that supports image output.
var model = FirebaseAI.GetInstance(FirebaseAI.Backend.GoogleAI()).GetGenerativeModel(
modelName: "gemini-3.1-flash-image",
// Configure the model to respond with images only.
generationConfig: new GenerationConfig(
responseModalities: new[] { ResponseModality.Image })
);
// Prepare an image for the model to edit
var imageFile = System.IO.File.ReadAllBytes(System.IO.Path.Combine(
UnityEngine.Application.streamingAssetsPath, "scones.jpg"));
var image = ModelContent.InlineData("image/jpeg", imageFile);
// Provide a text prompt instructing the model to edit the image
var prompt = ModelContent.Text("Edit this image to make it look like a cartoon.");
// To edit the image, call `GenerateContent` with the image and text input
var response = await model.GenerateContentAsync(new [] { prompt, image });
var text = response.Text;
if (!string.IsNullOrWhiteSpace(text)) {
// Do something with the text
}
// Handle the generated image
var imageParts = response.Candidates.First().Content.Parts
.OfType<ModelContent.InlineDataPart>()
.Where(part => part.MimeType == "image/png");
foreach (var imagePart in imageParts) {
// Load the Image into a Unity Texture2D object
Texture2D texture2D = new Texture2D(2, 2);
if (texture2D.LoadImage(imagePart.Data.ToArray())) {
// Do something with the image
}
}
Редактируйте и повторяйте изображения с помощью многоходового чата.
| Прежде чем опробовать этот пример, выполните раздел «Перед началом работы » этого руководства, чтобы настроить свой проект и приложение. В этом разделе вам также нужно будет нажать кнопку для выбранного вами поставщика API Gemini , чтобы увидеть на этой странице контент, относящийся к данному поставщику . |
Используя многоходовый чат, вы можете взаимодействовать с моделью Gemini Image, обрабатывая изображения, которые она генерирует или которые предоставляете вы.
Создайте экземпляр GenerativeModel , укажите в конфигурации модели тип ответа IMAGE и вызовите методы startChat() и sendMessage() для отправки сообщений новым пользователям.
Быстрый
import FirebaseAILogic
// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a Gemini model that supports image output.
let generativeModel = FirebaseAI.firebaseAI(backend: .googleAI()).generativeModel(
modelName: "gemini-3.1-flash-image",
// Configure the model to respond with images only.
generationConfig: GenerationConfig(responseModalities: [.image])
)
// Initialize the chat
let chat = model.startChat()
guard let image = UIImage(named: "scones") else { fatalError("Image file not found.") }
// Provide an initial text prompt instructing the model to edit the image
let prompt = "Edit this image to make it look like a cartoon"
// To generate an initial response, send a user message with the image and text prompt
let response = try await chat.sendMessage(image, prompt)
// Inspect the generated image
guard let inlineDataPart = response.inlineDataParts.first else {
fatalError("No image data in response.")
}
guard let uiImage = UIImage(data: inlineDataPart.data) else {
fatalError("Failed to convert data to UIImage.")
}
// Follow up requests do not need to specify the image again
let followUpResponse = try await chat.sendMessage("But make it old-school line drawing style")
// Inspect the edited image after the follow up request
guard let followUpInlineDataPart = followUpResponse.inlineDataParts.first else {
fatalError("No image data in response.")
}
guard let followUpUIImage = UIImage(data: followUpInlineDataPart.data) else {
fatalError("Failed to convert data to UIImage.")
}
Kotlin
// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a Gemini model that supports image output.
val model = Firebase.ai(backend = GenerativeBackend.googleAI()).generativeModel(
modelName = "gemini-3.1-flash-image",
// Configure the model to respond with images only.
generationConfig = generationConfig {
responseModalities = listOf(ResponseModality.IMAGE) }
)
// Provide an image for the model to edit
val bitmap = BitmapFactory.decodeResource(context.resources, R.drawable.scones)
// Create the initial prompt instructing the model to edit the image
val prompt = content {
image(bitmap)
text("Edit this image to make it look like a cartoon")
}
// Initialize the chat
val chat = model.startChat()
// To generate an initial response, send a user message with the image and text prompt
var response = chat.sendMessage(prompt)
// Inspect the returned image
var generatedImageAsBitmap = response
.candidates.first().content.parts.filterIsInstance<ImagePart>().firstOrNull()?.image
// Follow up requests do not need to specify the image again
response = chat.sendMessage("But make it old-school line drawing style")
generatedImageAsBitmap = response
.candidates.first().content.parts.filterIsInstance<ImagePart>().firstOrNull()?.image
Java
// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a Gemini model that supports image output.
GenerativeModel ai = FirebaseAI.getInstance(GenerativeBackend.googleAI()).generativeModel(
"gemini-3.1-flash-image",
// Configure the model to respond with images only.
new GenerationConfig.Builder()
.setResponseModalities(Arrays.asList(ResponseModality.IMAGE))
.build()
);
GenerativeModelFutures model = GenerativeModelFutures.from(ai);
// Provide an image for the model to edit
Bitmap bitmap = BitmapFactory.decodeResource(resources, R.drawable.scones);
// Initialize the chat
ChatFutures chat = model.startChat();
// Create the initial prompt instructing the model to edit the image
Content prompt = new Content.Builder()
.setRole("user")
.addImage(bitmap)
.addText("Edit this image to make it look like a cartoon")
.build();
// To generate an initial response, send a user message with the image and text prompt
ListenableFuture<GenerateContentResponse> response = chat.sendMessage(prompt);
// Extract the image from the initial response
ListenableFuture<@Nullable Bitmap> initialRequest = Futures.transform(response, result -> {
for (Part part : result.getCandidates().get(0).getContent().getParts()) {
if (part instanceof ImagePart) {
ImagePart imagePart = (ImagePart) part;
return imagePart.getImage();
}
}
return null;
}, executor);
// Follow up requests do not need to specify the image again
ListenableFuture<GenerateContentResponse> modelResponseFuture = Futures.transformAsync(
initialRequest,
generatedImage -> {
Content followUpPrompt = new Content.Builder()
.addText("But make it old-school line drawing style")
.build();
return chat.sendMessage(followUpPrompt);
},
executor);
// Add a final callback to check the reworked image
Futures.addCallback(modelResponseFuture, new FutureCallback<GenerateContentResponse>() {
@Override
public void onSuccess(GenerateContentResponse result) {
for (Part part : result.getCandidates().get(0).getContent().getParts()) {
if (part instanceof ImagePart) {
ImagePart imagePart = (ImagePart) part;
Bitmap generatedImageAsBitmap = imagePart.getImage();
break;
}
}
}
@Override
public void onFailure(Throwable t) {
t.printStackTrace();
}
}, executor);
Web
import { initializeApp } from "firebase/app";
import { getAI, getGenerativeModel, GoogleAIBackend, ResponseModality } from "firebase/ai";
// TODO(developer): Replace the following with your app's Firebase configuration
// See: https://firebase.google.com/docs/web/learn-more#config-object
const firebaseConfig = {
// ...
};
// Initialize FirebaseApp
const firebaseApp = initializeApp(firebaseConfig);
// Initialize the Gemini Developer API backend service.
const ai = getAI(firebaseApp, { backend: new GoogleAIBackend() });
// Create a `GenerativeModel` instance with a model that supports your use case
const model = getGenerativeModel(ai, {
model: "gemini-3.1-flash-image",
// Configure the model to respond with images only.
generationConfig: {
responseModalities: [ResponseModality.IMAGE],
},
});
// Prepare an image for the model to edit
async function fileToGenerativePart(file) {
const base64EncodedDataPromise = new Promise((resolve) => {
const reader = new FileReader();
reader.onloadend = () => resolve(reader.result.split(',')[1]);
reader.readAsDataURL(file);
});
return {
inlineData: { data: await base64EncodedDataPromise, mimeType: file.type },
};
}
const fileInputEl = document.querySelector("input[type=file]");
const imagePart = await fileToGenerativePart(fileInputEl.files[0]);
// Provide an initial text prompt instructing the model to edit the image
const prompt = "Edit this image to make it look like a cartoon";
// Initialize the chat
const chat = model.startChat();
// To generate an initial response, send a user message with the image and text prompt
const result = await chat.sendMessage([prompt, imagePart]);
// Request and inspect the generated image
try {
const inlineDataParts = result.response.inlineDataParts();
if (inlineDataParts?.[0]) {
// Inspect the generated image
const image = inlineDataParts[0].inlineData;
console.log(image.mimeType, image.data);
}
} catch (err) {
console.error('Prompt or candidate was blocked:', err);
}
// Follow up requests do not need to specify the image again
const followUpResult = await chat.sendMessage("But make it old-school line drawing style");
// Request and inspect the returned image
try {
const followUpInlineDataParts = followUpResult.response.inlineDataParts();
if (followUpInlineDataParts?.[0]) {
// Inspect the generated image
const followUpImage = followUpInlineDataParts[0].inlineData;
console.log(followUpImage.mimeType, followUpImage.data);
}
} catch (err) {
console.error('Prompt or candidate was blocked:', err);
}
Dart
import 'package:firebase_ai/firebase_ai.dart';
import 'package:firebase_core/firebase_core.dart';
import 'firebase_options.dart';
await Firebase.initializeApp(
options: DefaultFirebaseOptions.currentPlatform,
);
// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a Gemini model that supports image output.
final model = FirebaseAI.googleAI().generativeModel(
model: 'gemini-3.1-flash-image',
// Configure the model to respond with images only.
generationConfig: GenerationConfig(responseModalities: [ResponseModalities.image]),
);
// Prepare an image for the model to edit
final image = await File('scones.jpg').readAsBytes();
final imagePart = InlineDataPart('image/jpeg', image);
// Provide an initial text prompt instructing the model to edit the image
final prompt = TextPart("Edit this image to make it look like a cartoon");
// Initialize the chat
final chat = model.startChat();
// To generate an initial response, send a user message with the image and text prompt
final response = await chat.sendMessage([
Content.multi([prompt,imagePart])
]);
// Inspect the returned image
if (response.inlineDataParts.isNotEmpty) {
final imageBytes = response.inlineDataParts[0].bytes;
// Process the image
} else {
// Handle the case where no images were generated
print('Error: No images were generated.');
}
// Follow up requests do not need to specify the image again
final followUpResponse = await chat.sendMessage([
Content.text("But make it old-school line drawing style")
]);
// Inspect the returned image
if (followUpResponse.inlineDataParts.isNotEmpty) {
final followUpImageBytes = response.inlineDataParts[0].bytes;
// Process the image
} else {
// Handle the case where no images were generated
print('Error: No images were generated.');
}
Единство
using Firebase;
using Firebase.AI;
// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a Gemini model that supports image output.
var model = FirebaseAI.GetInstance(FirebaseAI.Backend.GoogleAI()).GetGenerativeModel(
modelName: "gemini-3.1-flash-image",
// Configure the model to respond with images only.
generationConfig: new GenerationConfig(
responseModalities: new[] { ResponseModality.Image })
);
// Prepare an image for the model to edit
var imageFile = System.IO.File.ReadAllBytes(System.IO.Path.Combine(
UnityEngine.Application.streamingAssetsPath, "scones.jpg"));
var image = ModelContent.InlineData("image/jpeg", imageFile);
// Provide an initial text prompt instructing the model to edit the image
var prompt = ModelContent.Text("Edit this image to make it look like a cartoon.");
// Initialize the chat
var chat = model.StartChat();
// To generate an initial response, send a user message with the image and text prompt
var response = await chat.SendMessageAsync(new [] { prompt, image });
// Inspect the returned image
var imageParts = response.Candidates.First().Content.Parts
.OfType<ModelContent.InlineDataPart>()
.Where(part => part.MimeType == "image/png");
// Load the image into a Unity Texture2D object
UnityEngine.Texture2D texture2D = new(2, 2);
if (texture2D.LoadImage(imageParts.First().Data.ToArray())) {
// Do something with the image
}
// Follow up requests do not need to specify the image again
var followUpResponse = await chat.SendMessageAsync("But make it old-school line drawing style");
// Inspect the returned image
var followUpImageParts = followUpResponse.Candidates.First().Content.Parts
.OfType<ModelContent.InlineDataPart>()
.Where(part => part.MimeType == "image/png");
// Load the image into a Unity Texture2D object
UnityEngine.Texture2D followUpTexture2D = new(2, 2);
if (followUpTexture2D.LoadImage(followUpImageParts.First().Data.ToArray())) {
// Do something with the image
}
Предоставьте справочные изображения.
В моделях Gemini Image вы можете указать в задании эталонные изображения. Эти изображения могут включать в себя следующее:
Gemini 3.x Pro Image (
gemini-3-pro-image, также известный как "Nano Banana Pro")- До 6 высококачественных изображений объектов для включения в итоговое изображение.
- До 5 изображений персонажей для обеспечения единообразия их внешнего вида.
- До 3 изображений могут быть использованы в качестве стилистических ориентиров.
Образ прошивки Gemini 3.x (
gemini-3.1-flash-image, также известный как "Nano Banana 2"):- До 10 высококачественных изображений объектов для включения в итоговое изображение.
- До 4 изображений персонажей для обеспечения единообразия их внешнего вида.
Gemini 3.x Flash‑Lite Image (
gemini-3.1-flash-lite-image, также известный как "Nano Banana 2 Lite"):- До 10 высококачественных изображений объектов для включения в итоговое изображение.
- До 4 изображений персонажей для обеспечения единообразия их внешнего вида.
Изображение Gemini 2.5 Flash Image (
gemini-2.5-flash-image, также известное как "Nano Banana"):- До 3 изображений
Настройка генерации изображений
По умолчанию модели Gemini Image генерируют квадратные изображения (соотношение сторон 1:1) с разрешением 1024x1024. Вы можете настроить вывод генерируемых изображений, используя свойство imageConfig в generationConfig .
Например, вы можете настроить выходное изображение, используя перечисления соотношения сторон и размера изображения (например, соотношение сторон 16:9 и разрешение 2K, в результате чего получится изображение размером 2752x1536), следующим образом:
Быстрый
// ...
// Specify aspect ratio and image size (both optional) in an image config.
let imageConfig = ImageConfig(aspectRatio: .landscape16x9, imageSize: .size2K)
let generationConfig = GenerationConfig(
// Configure the model to respond with text (optional) and images (required).
// Include the text modality only if you want interleaved text and images in the response.
responseModalities: [.text, .image],
imageConfig: imageConfig
)
// Make sure you initialize your chosen Gemini API backend service
let model = FirebaseAI.firebaseAI().generativeModel(
modelName: "gemini-3.1-flash-image",
generationConfig: generationConfig
)
// ...
Kotlin
// ...
val config = generationConfig {
// Configure the model to respond with text (optional) and images (required).
// Include the text modality only if you want interleaved text and images in the response.
responseModalities = listOf(ResponseModality.TEXT, ResponseModality.IMAGE)
// Specify aspect ratio and image size (both optional) in an image config.
imageConfig = imageConfig {
aspectRatio = AspectRatio.LANDSCAPE_16x9
imageSize = ImageSize.SIZE_2K
}
}
// Make sure you initialize your chosen Gemini API backend service
val model = Firebase.ai.generativeModel(
modelName = "gemini-3.1-flash-image",
generationConfig = config
)
// ...
Java
// ...
GenerationConfig config = new GenerationConfig.Builder()
// Configure the model to respond with text (optional) and images (required).
// Include the text modality only if you want interleaved text and images in the response.
.setResponseModalities(Arrays.asList(ResponseModality.TEXT, ResponseModality.IMAGE))
// Specify aspect ratio and image size (both optional) in an image config.
.setImageConfig(
ImageConfig.builder()
.setAspectRatio(AspectRatio.LANDSCAPE_16x9)
.setImageSize(ImageSize.SIZE_2K)
.build()
)
.build();
// Make sure you initialize your chosen Gemini API backend service
GenerativeModel model = FirebaseAI.getInstance().generativeModel(
"gemini-3.1-flash-image",
config
);
// ...
Web
import {
getGenerativeModel,
ResponseModality,
ImageConfigAspectRatio,
ImageConfigImageSize
} from "firebase/ai";
// ...
const generationConfig = {
// Configure the model to respond with text (optional) and images (required).
// Include the text modality only if you want interleaved text and images in the response.
responseModalities: [ResponseModality.TEXT, ResponseModality.IMAGE],
// Specify aspect ratio and image size (both optional) in an image config.
imageConfig: {
aspectRatio: ImageConfigAspectRatio.LANDSCAPE_16x9,
imageSize: ImageConfigImageSize.SIZE_2K
}
};
// Make sure you initialize your chosen Gemini API backend service
const model = getGenerativeModel(ai, {
model: "gemini-3.1-flash-image",
generationConfig
});
// ...
Dart
// ...
final generationConfig = GenerationConfig(
// Configure the model to respond with text (optional) and images (required).
// Include the text modality only if you want interleaved text and images in the response.
responseModalities: [ResponseModalities.text, ResponseModalities.image],
// Specify aspect ratio and image size (both optional) in an image config.
imageConfig: ImageConfig(
aspectRatio: ImageAspectRatio.landscape16x9,
imageSize: ImageSize.size2K,
),
);
// Make sure you initialize your chosen Gemini API backend service
final model = FirebaseAI.instance.generativeModel(
model: 'gemini-3.1-flash-image,
generationConfig: generationConfig,
);
// ...
Единство
// ...
var generationConfig = new GenerationConfig(
// Configure the model to respond with text (optional) and images (required).
// Include the text modality only if you want interleaved text and images in the response.
responseModalities: new[] { ResponseModality.Text, ResponseModality.Image },
// Specify aspect ratio and image size (both optional) in an image config.
imageConfig: new ImageConfig(
aspectRatio: ImageConfig.AspectRatio.Landscape16x9,
imageSize: ImageConfig.ImageSize.Size2K)
);
// Make sure you initialize your chosen Gemini API backend service.
var model = FirebaseAI.GetInstance().GetGenerativeModel(
modelName: "gemini-3.1-flash-image",
generationConfig: generationConfig
);
// ...
Поддерживаемые соотношения сторон
Все модели Gemini Image поддерживают следующие соотношения сторон:
По умолчанию: 1:1 (квадрат)
1:1 , 1:4 , 1:8 , 2:3 , 3:2 , 3:4 , 4:1 , 4:3 , 4:5 , 5:4 , 8:1 , 9:16 , 16:9 , 21:9
Поддерживаемые размеры (разрешения) изображений
Поддерживаемые размеры изображений (разрешения) зависят от используемой модели.
| Модель Gemini Image | Поддерживаемые размеры (разрешения) |
|---|---|
Образ Gemini 3.x Progemini-3-pro-image("Nano Banana Pro") | По умолчанию: 1K (1024)1K (1024), 2K (2048), 4K (4096) |
Образ Gemini 3.x Flashgemini-3.1-flash-image("Нано-банан 2") | По умолчанию: 1K (1024)512 , 1K (1024), 2K (2048), 4K (4096) |
Gemini 3.x Flash‑Lite Imagegemini-3.1-flash-lite-image("Nano Banana 2 Lite") | По умолчанию: 1K (1024)512 (только API для разработчиков Gemini ) , 1K (1024) |
Изображение со вспышкой Gemini 2.5gemini-2.5-flash-image("Нано-банан") | Фиксировано на 1K (1024) |
Необходимо использовать суффикс K в верхнем регистре (например, 1K , 2K , 4K ). Значение 512 не использует суффикс K Суффикс k в нижнем регистре (например, 1k ) будет отклонен.
Поддерживаемые функции
Поддерживаются следующие функции, такие как режимы отображения, инструменты, ввод и языки.
Информацию о поддерживаемых соотношениях сторон и разрешениях для каждой модели см. в разделе «Настройка генерации изображений» ранее в этом руководстве.
Поддерживаемые режимы
Ниже перечислены поддерживаемые «режимы» для моделей изображений Gemini . Эти «режимы» не задаются явно в ваших запросах. Это скорее рекомендуемые шаблоны для распространенных сценариев использования. Для каждого режима в этом списке показан пример запроса и приведен пример кода ранее в этом руководстве.
Текстовая Изображение(я) (только текст для преобразования в изображение)
- Создайте изображение Эйфелевой башни с фейерверками на заднем плане.
Текстовая Изображение(я) (текст отображается внутри изображения)
- Создайте кинематографическое изображение большого здания с помощью гигантской проекции текста, нанесенной на фасад здания.
Текстовая Изображение(я) и текст (чередование)
Создайте иллюстрированный рецепт паэльи. Добавляйте изображения к тексту по мере создания рецепта.
Создайте историю о собаке в стиле 3D-мультфильма. Для каждой сцены создайте изображение.
Изображение(я) и текст Изображение(я) и текст (чередование)
- [изображение обставленной комнаты] + Какие еще цвета диванов подойдут для моего помещения? Можете обновить изображение?
Редактирование изображений (преобразование текста и изображения в изображение)
[изображение булочек] + Отредактируйте это изображение, чтобы оно выглядело как мультфильм.
[изображение кошки] + [изображение подушки] + Вышейте крестиком мою кошку на этой подушке.
Многоэтапная обработка изображений (чат)
- [изображение синей машины] + Превратите эту машину в кабриолет. Затем измените цвет на желтый.
Поддерживаемые инструменты
Поддержка функции "Заземление" с помощью
Gemini 3.x Pro Image (
gemini-3-pro-image, также известный как "Nano Banana Pro")Образ прошивки Gemini 3.x (
gemini-3.1-flash-image, также известный как "Nano Banana 2")
Другие поддерживаемые возможности
Поддерживается многомодальный ввод:
В качестве входных данных для изображений используются все модели Gemini Image.
Видеовход:
- Образ прошивки Gemini 3.x (
gemini-3.1-flash-image, также известный как "Nano Banana 2") - Gemini 3.x Flash‑Lite Image (
gemini-3.1-flash-lite-image, также известный как "Nano Banana 2 Lite")
- Образ прошивки Gemini 3.x (
Аудиовход: отсутствует в моделях Gemini Image.
Все модели Gemini Image поддерживают следующие функции:
- Создание изображений в формате PNG.
- Создание и редактирование изображений людей.
- Использование фильтров безопасности, обеспечивающих гибкий и менее ограничительный пользовательский опыт.
Поддержка генерации структурированного вывода (например, в формате JSON) :
- Gemini 3.x Pro Image (
gemini-3-pro-image, также известный как "Nano Banana Pro")
- Gemini 3.x Pro Image (
Поддерживаемые языки
Хотя модели Gemini Image могут принимать текстовые подсказки на многих языках , языки, перечисленные в этом разделе, обеспечат наилучшую производительность.
Поддерживаемые языки для текстовой подсказки:
- Модели изображений Gemini 3.x :
ar-EG,de-DE,EN,es-MX,fr-FR,hi-IN,id-ID,it-IT,ja-JP,ko-KR,pt-BR,ru-RU,uk-UA,vi-VN,zh-CN - Модель флэш-образа Gemini 2.5 :
EN,es-MX,ja-JP,zh-CN,hi-IN.
- Модели изображений Gemini 3.x :
Поддерживаемые языки для текста на сгенерированном изображении:
- Модели образов Gemini 3.x : те же языки, что и в списке выше.
- Модель изображения Gemini 2.5 Flash : только на английском языке.
Чтобы использовать определенный язык в сгенерированном изображении (даже без кода языка), просто укажите модели в вашем запросе (например, «Обновите эту инфографику, чтобы она была на испанском языке. Не изменяйте никакие другие элементы изображения» ).
Передовые методы
Ниже приведены рекомендации по использованию моделей изображений Gemini .
При создании изображения, содержащего текст, сначала создайте сам текст, а затем создайте изображение с этим текстом.
Генерация изображений может срабатывать не всегда. Кроме того, в следующих ситуациях генерация изображений или текста может работать не так, как ожидается:
Модель может генерировать только текст, но не изображение (особенно если подсказка неоднозначна). В этом случае
FinishReasonбудет иметь значениеNO_IMAGE.
Попробуйте явно указать, какие изображения нужны в качестве результата. Например, «сгенерировать изображение», «предоставлять изображения по мере необходимости», «обновить изображение».Генерация модели может остановиться на полпути.
Попробуйте еще раз или воспользуйтесь другим вариантом ответа.Модель может генерировать текст в виде изображения.
Попробуйте явно указать, какой текст должен быть получен. Например, «сгенерировать повествовательный текст вместе с иллюстрациями».Если запрос потенциально небезопасен, модель может не обработать его и вместо этого вернуть ответ, указывающий на невозможность создания небезопасных изображений. В этом случае
FinishReasonбудет иметьSTOP.