Gemini Live API позволяет взаимодействовать с Gemini в реальном времени с низкой задержкой, используя голос и видео. Модель двунаправленная.
Live API и его семейство моделей могут обрабатывать непрерывные потоки аудио, видео или текста и предоставлять мгновенные ответы, похожие на человеческую речь, создавая естественный разговорный интерфейс для ваших пользователей.
На этой странице рассказывается, как начать работу с самой распространенной функцией – потоковой передачей аудиовхода и аудиовыхода, но Live API поддерживает множество других функций и вариантов конфигурации.
Live API – это API с сохранением состояния, который создает подключение WebSocket, чтобы установить сеанс между клиентом и сервером Gemini. Подробную информацию можно найти в справочной документации по Live API (Gemini Developer API | Agent Platform Gemini API (formerly Vertex AI)).
Полезные ресурсы
Swift – приложение для быстрого запуска | Android – приложение для быстрого запуска | Веб-приложение – приложение для быстрого запуска | Flutter – приложение для быстрого запуска | Unity – скоро!
Попробуйте Gemini Live API в реальном развернутом приложении. Для этого откройте Flutter ИИ-студию в консоли Firebase.
Подготовка
Если вы ещё этого не сделали, выполните инструкции из руководства по началу работы. В нем рассказывается, как настроить проект Firebase, подключить к нему приложение, добавить SDK, инициализировать бэкенд-службу для выбранного поставщика Gemini API и создать экземпляр LiveModel.
Вы можете создать прототип с помощью запросов и Live API в Google AI Studio или Agent Studio.
Модели, поддерживающие эту функцию
Модели 3.x
Gemini Developer API
gemini-3.1-flash-live-preview
Несмотря на то что это предварительная версия модели, она доступна на бесплатном плане Gemini Developer API.
Agent Platform Gemini API (formerly Vertex AI)
Модели Gemini Live 3.x не поддерживаются
Модели 2.5
Хотя у модели могут быть разные названия в зависимости от поставщика API Gemini, ее функции остаются неизменными.
Gemini Developer API
gemini-2.5-flash-native-audio-preview-12-2025gemini-2.5-flash-native-audio-preview-09-2025
Несмотря на то что это предварительные модели, они доступны на бесплатном уровне Gemini Developer API.
Agent Platform Gemini API (formerly Vertex AI)
gemini-live-2.5-flash-native-audio(выпущен в декабре 2025 г.)gemini-live-2.5-flash-preview-native-audio-09-2025
При использовании Agent Platform Gemini API (formerly Vertex AI) модели Live API 2.5 не доступны в регионе
global.
Передача аудиовхода и аудиовыхода
|
Нажмите на поставщика Gemini API, чтобы посмотреть контент и код, относящиеся к нему. |
В примере ниже показана базовая реализация для отправки потокового аудиовхода и получения потокового аудиовыхода.
Дополнительные возможности и функции Live API описаны в разделе Что ещё можно сделать?.
Swift
Чтобы использовать Live API, создайте экземпляр LiveModel и задайте для параметра response modality значение audio.
import FirebaseAILogic
// Initialize the Gemini Developer API backend service.
// Create a `liveModel` instance with a model that supports the Live API.
let liveModel = FirebaseAI.firebaseAI(backend: .googleAI()).liveModel(
modelName: "GEMINI_LIVE_API_MODEL_NAME",
// Configure the model to respond with audio.
generationConfig: LiveGenerationConfig(
responseModalities: [.audio]
)
)
do {
let session = try await liveModel.connect()
// Load the audio file, or tap a microphone.
guard let audioFile = NSDataAsset(name: "audio.pcm") else {
fatalError("Failed to load audio file")
}
// Provide the audio data.
await session.sendAudioRealtime(audioFile.data)
var outputText = ""
for try await message in session.responses {
if case let .content(content) = message.payload {
content.modelTurn?.parts.forEach { part in
if let part = part as? InlineDataPart, part.mimeType.starts(with: "audio/pcm") {
// Handle 16bit pcm audio data at 24khz
playAudio(part.data)
}
}
// Optional: if you don't need to send more requests.
if content.isTurnComplete {
await session.close()
}
}
}
} catch {
fatalError(error.localizedDescription)
}
Kotlin
Чтобы использовать Live API, создайте экземпляр LiveModel и задайте для параметра response modality значение AUDIO.
// Initialize the Gemini Developer API backend service.
// Create a `liveModel` instance with a model that supports the Live API.
val liveModel = Firebase.ai(backend = GenerativeBackend.googleAI()).liveModel(
modelName = "GEMINI_LIVE_API_MODEL_NAME",
// Configure the model to respond with audio.
generationConfig = liveGenerationConfig {
responseModality = ResponseModality.AUDIO
}
)
val session = liveModel.connect()
// This is the recommended approach.
// However, you can create your own recorder and handle the stream.
session.startAudioConversation()
Java
Чтобы использовать Live API, создайте экземпляр LiveModel и задайте для параметра response modality значение AUDIO.
ExecutorService executor = Executors.newFixedThreadPool(1);
// Initialize the Gemini Developer API backend service.
// Create a `liveModel` instance with a model that supports the Live API.
LiveGenerativeModel lm = FirebaseAI.getInstance(GenerativeBackend.googleAI()).liveModel(
"GEMINI_LIVE_API_MODEL_NAME",
// Configure the model to respond with audio.
new LiveGenerationConfig.Builder()
.setResponseModality(ResponseModality.AUDIO)
.build()
);
LiveModelFutures liveModel = LiveModelFutures.from(lm);
ListenableFuture<LiveSession> sessionFuture = liveModel.connect();
Futures.addCallback(sessionFuture, new FutureCallback<LiveSession>() {
@Override
public void onSuccess(LiveSession ses) {
LiveSessionFutures session = LiveSessionFutures.from(ses);
session.startAudioConversation();
}
@Override
public void onFailure(Throwable t) {
// Handle exceptions
}
}, executor);
Web
Чтобы использовать Live API, создайте экземпляр LiveGenerativeModel и задайте для параметра response modality значение AUDIO.
import { initializeApp } from "firebase/app";
import { getAI, getLiveGenerativeModel, GoogleAIBackend, ResponseModality } from "firebase/ai";
// TODO(developer): Replace the following with your app's Firebase configuration
// See: https://firebase.google.com/docs/web/learn-more#config-object
const firebaseConfig = {
// ...
};
// Initialize FirebaseApp
const firebaseApp = initializeApp(firebaseConfig);
// Initialize the Gemini Developer API backend service.
const ai = getAI(firebaseApp, { backend: new GoogleAIBackend() });
// Create a `LiveGenerativeModel` instance with a model that supports the Live API.
const liveModel = getLiveGenerativeModel(ai, {
model: "GEMINI_LIVE_API_MODEL_NAME",
// Configure the model to respond with audio.
generationConfig: {
responseModalities: [ResponseModality.AUDIO],
},
});
const session = await liveModel.connect();
// Start the audio conversation.
const audioConversationController = await startAudioConversation(session);
// ... Later, to stop the audio conversation
// await audioConversationController.stop()
Dart
Чтобы использовать Live API, создайте экземпляр LiveGenerativeModel и задайте для параметра response modality значение audio.
import 'package:firebase_ai/firebase_ai.dart';
import 'package:firebase_core/firebase_core.dart';
import 'firebase_options.dart';
import 'package:your_audio_recorder_package/your_audio_recorder_package.dart';
late LiveModelSession _session;
final _audioRecorder = YourAudioRecorder();
await Firebase.initializeApp(
options: DefaultFirebaseOptions.currentPlatform,
);
// Initialize the Gemini Developer API backend service.
// Create a `liveGenerativeModel` instance with a model that supports the Live API.
final liveModel = FirebaseAI.googleAI().liveGenerativeModel(
model: 'GEMINI_LIVE_API_MODEL_NAME',
// Configure the model to respond with audio.
liveGenerationConfig: LiveGenerationConfig(
responseModalities: [ResponseModalities.audio],
),
);
_session = await liveModel.connect();
final audioRecordStream = _audioRecorder.startRecordingStream();
// Map the Uint8List stream to InlineDataPart stream.
final mediaChunkStream = audioRecordStream.map((data) {
return InlineDataPart('audio/pcm', data);
});
await _session.startMediaStream(mediaChunkStream);
// In a separate thread, receive the audio response from the model.
await for (final message in _session.receive()) {
// Process the received message.
}
Unity
Чтобы использовать Live API, создайте экземпляр LiveModel и задайте для параметра response modality значение Audio.
using Firebase;
using Firebase.AI;
async Task SendTextReceiveAudio() {
// Initialize the Gemini Developer API backend service.
// Create a `LiveModel` instance with a model that supports the Live API.
var liveModel = FirebaseAI.GetInstance(FirebaseAI.Backend.GoogleAI()).GetLiveModel(
modelName: "GEMINI_LIVE_API_MODEL_NAME",
// Configure the model to respond with audio.
liveGenerationConfig: new LiveGenerationConfig(
responseModalities: new[] { ResponseModality.Audio })
);
LiveSession session = await liveModel.ConnectAsync();
// Start a coroutine to send audio from the Microphone.
var recordingCoroutine = StartCoroutine(SendAudio(session));
// Start receiving the response.
await ReceiveAudio(session);
}
IEnumerator SendAudio(LiveSession liveSession) {
string microphoneDeviceName = null;
int recordingFrequency = 16000;
int recordingBufferSeconds = 2;
var recordingClip = Microphone.Start(microphoneDeviceName, true,
recordingBufferSeconds, recordingFrequency);
int lastSamplePosition = 0;
while (true) {
if (!Microphone.IsRecording(microphoneDeviceName)) {
yield break;
}
int currentSamplePosition = Microphone.GetPosition(microphoneDeviceName);
if (currentSamplePosition != lastSamplePosition) {
// The Microphone uses a circular buffer, so we need to check if the
// current position wrapped around to the beginning, and handle it accordingly.
int sampleCount;
if (currentSamplePosition > lastSamplePosition) {
sampleCount = currentSamplePosition - lastSamplePosition;
} else {
sampleCount = recordingClip.samples - lastSamplePosition + currentSamplePosition;
}
if (sampleCount > 0) {
// Get the audio chunk.
float[] samples = new float[sampleCount];
recordingClip.GetData(samples, lastSamplePosition);
// Send the data, discarding the resulting Task to avoid the warning.
_ = liveSession.SendAudioAsync(samples);
lastSamplePosition = currentSamplePosition;
}
}
// Wait for a short delay before reading the next sample from the Microphone.
const float MicrophoneReadDelay = 0.5f;
yield return new WaitForSeconds(MicrophoneReadDelay);
}
}
Queue<float> audioBuffer = new();
async Task ReceiveAudio(LiveSession liveSession) {
int sampleRate = 24000;
int channelCount = 1;
// Create a looping AudioClip to fill with the received audio data.
int bufferSamples = (int)(sampleRate * channelCount);
AudioClip clip = AudioClip.Create("StreamingPCM", bufferSamples, channelCount,
sampleRate, true, OnAudioRead);
// Attach the clip to an AudioSource and start playing it.
AudioSource audioSource = GetComponent<AudioSource>();
audioSource.clip = clip;
audioSource.loop = true;
audioSource.Play();
// Start receiving the response.
await foreach (var message in liveSession.ReceiveAsync()) {
// Process the received message
foreach (float[] pcmData in message.AudioAsFloat) {
lock (audioBuffer) {
foreach (float sample in pcmData) {
audioBuffer.Enqueue(sample);
}
}
}
}
}
// This method is called by the AudioClip to load audio data.
private void OnAudioRead(float[] data) {
int samplesToProvide = data.Length;
int samplesProvided = 0;
lock(audioBuffer) {
while (samplesProvided < samplesToProvide && audioBuffer.Count > 0) {
data[samplesProvided] = audioBuffer.Dequeue();
samplesProvided++;
}
}
while (samplesProvided < samplesToProvide) {
data[samplesProvided] = 0.0f;
samplesProvided++;
}
}
Цены и подсчет токенов
Информацию о ценах на модели Live API можно найти в документации выбранного вами поставщика Gemini API:Gemini Developer API | Agent Platform Gemini API (formerly Vertex AI).
Независимо от того, какого поставщика Gemini API вы используете, Live API не поддерживает Count Tokens API.
Что ещё ты умеешь делать?
Ознакомьтесь с полным набором возможностей для Live API, например с потоковой передачей различных входных данных (аудио, текста или видео + аудио).
Настройте реализацию, используя различные варианты конфигурации, например добавьте транскрипцию или задайте голос для ответа.
Узнайте, как управлять сеансами, в том числе обновлять контент во время сеанса, сжимать окно контекста, определять, когда сеанс подходит к концу, и возобновлять сеанс.
Чтобы повысить эффективность реализации, предоставьте модели доступ к инструментам, таким как вызов функций и Grounding с помощью
Google Search . Официальная документация по использованию инструментов с Live API скоро будет опубликована.Ознакомьтесь с ограничениями и требованиями для использования Live API, например с ограничениями на продолжительность сеанса, частоту запросов и поддерживаемые языки.