شروع به کار با Gemini Live API بااستفاده از Firebase AI Logic


Gemini Live API امکان تعاملات صوتی و تصویری هم‌زمان با تأخیر کم با مدل Gemini را فراهم می‌کند که دوطرفه است.

‫Live API و خانواده ویژه مدل‌های آن می‌توانند جاری‌سازی‌های پیوسته صوت، ویدیو، یا نوشتار را پردازش کنند تا پاسخ‌های گفتاری فوری و شبیه انسان ارائه دهند و تجربه مکالمه طبیعی برای کاربران شما ایجاد کنند.

این صفحه نحوه شروع کار با رایج‌ترین قابلیت — جاری‌سازی ورودی و خروجی صدا را توضیح می‌دهد، اما Live API از قابلیت‌های متفاوت و گزینه‌های پیکربندی بسیاری پشتیبانی می‌کند.

Live API یک «میانای برنامه‌سازی کاربردی» حالت‌دار است که برای ایجاد جلسه بین کارخواه و سرور Gemini، اتصال WebSocket ایجاد می‌کند. برای جزئیات، به Live API سند مرجع (Gemini Developer API | Agent Platform Gemini API (formerly Vertex AI)) مراجعه کنید.

رفتن به نمونه‌های کد

منابع مفید را بررسی کنید

قبل از شروع

اگر هنوز این کار را نکرده‌اید، راهنمای شروع به کار را تکمیل کنید. این راهنما نحوه راه‌اندازی پروژه Firebase، متصل کردن برنامه به Firebase، افزودن «کیت توسعه نرم‌افزار»، مقداردهی اولیه سرویس زیرینه برای ارائه‌دهنده Gemini API انتخابی، و ایجاد نمونه LiveModel را توضیح می‌دهد.

می‌توانید با پیام‌واره‌ها و Live API در Google AI Studio یا Agent  Studio نمونه اولیه بسازید.

مدل‌هایی که از این قابلیت پشتیبانی می‌کنند

  • مدل‌های ۳.x

    • Gemini Developer API

      • gemini-3.1-flash-live-preview

      اگرچه این مدل پیش‌نمایش است، اما در «سطح رایگان» Gemini Developer API دردسترس است.

    • Agent Platform Gemini API (formerly Vertex AI)

      عدم پشتیبانی از مدل‌های Gemini Live 3.x

  • ‫۲.۵ مدل

    اگرچه مدل بسته به ارائه‌دهنده Gemini API نام‌های مدل متفاوتی دارد، ویژگی‌های مدل یکسان است.

    • Gemini Developer API

      • gemini-2.5-flash-native-audio-preview-12-2025
      • gemini-2.5-flash-native-audio-preview-09-2025

      اگرچه این‌ها مدل‌های پیش‌نمایش هستند، اما در «سطح رایگان» Gemini Developer API دردسترس هستند.

    • Agent Platform Gemini API (formerly Vertex AI)

      • gemini-live-2.5-flash-native-audio (در دسامبر ۲۰۲۵ منتشر شد)
      • gemini-live-2.5-flash-preview-native-audio-09-2025

      هنگام استفاده از Agent Platform Gemini API (formerly Vertex AI)، مدل‌های Live API 2.5 در مکان global دردسترس نیستند.

جاری‌سازی ورودی و خروجی صوتی

روی ارائه‌دهنده Gemini API خود کلیک کنید تا محتوا و کد مخصوص ارائه‌دهنده را در این صفحه مشاهده کنید.

مثال زیر پیاده‌سازی پایه را برای ارسال ورودی صوتی جاری‌سازی‌شده و دریافت برونداد صوتی جاری‌سازی‌شده نشان می‌دهد.

برای گزینه‌ها و قابلیت‌های بیشتر برای Live API، بخش «چه کارهای دیگری می‌توانید انجام دهید؟» را در ادامه این صفحه مرور کنید.

Swift

برای استفاده از Live API، نمونه‌ای LiveModel ایجاد کنید و روش پاسخ را روی audio تنظیم کنید.


import FirebaseAILogic

// Initialize the Gemini Developer API backend service.
// Create a `liveModel` instance with a model that supports the Live API.
let liveModel = FirebaseAI.firebaseAI(backend: .googleAI()).liveModel(
  modelName: "GEMINI_LIVE_API_MODEL_NAME",
  // Configure the model to respond with audio.
  generationConfig: LiveGenerationConfig(
    responseModalities: [.audio]
  )
)

do {
  let session = try await liveModel.connect()

  // Load the audio file, or tap a microphone.
  guard let audioFile = NSDataAsset(name: "audio.pcm") else {
    fatalError("Failed to load audio file")
  }

  // Provide the audio data.
  await session.sendAudioRealtime(audioFile.data)

  var outputText = ""
  for try await message in session.responses {
    if case let .content(content) = message.payload {
      content.modelTurn?.parts.forEach { part in
        if let part = part as? InlineDataPart, part.mimeType.starts(with: "audio/pcm") {
          // Handle 16bit pcm audio data at 24khz
          playAudio(part.data)
        }
      }
      // Optional: if you don't need to send more requests.
      if content.isTurnComplete {
        await session.close()
      }
    }
  }
} catch {
  fatalError(error.localizedDescription)
}

Kotlin

برای استفاده از Live API، نمونه‌ای LiveModel ایجاد کنید و روش پاسخ را روی AUDIO تنظیم کنید.


// Initialize the Gemini Developer API backend service.
// Create a `liveModel` instance with a model that supports the Live API.
val liveModel = Firebase.ai(backend = GenerativeBackend.googleAI()).liveModel(
    modelName = "GEMINI_LIVE_API_MODEL_NAME",
    // Configure the model to respond with audio.
    generationConfig = liveGenerationConfig {
        responseModality = ResponseModality.AUDIO
   }
)

val session = liveModel.connect()

// This is the recommended approach.
// However, you can create your own recorder and handle the stream.
session.startAudioConversation()

Java

برای استفاده از Live API، نمونه‌ای LiveModel ایجاد کنید و روش پاسخ را روی AUDIO تنظیم کنید.


ExecutorService executor = Executors.newFixedThreadPool(1);
// Initialize the Gemini Developer API backend service.
// Create a `liveModel` instance with a model that supports the Live API.
LiveGenerativeModel lm = FirebaseAI.getInstance(GenerativeBackend.googleAI()).liveModel(
        "GEMINI_LIVE_API_MODEL_NAME",
        // Configure the model to respond with audio.
        new LiveGenerationConfig.Builder()
                .setResponseModality(ResponseModality.AUDIO)
                .build()
);
LiveModelFutures liveModel = LiveModelFutures.from(lm);

ListenableFuture<LiveSession> sessionFuture =  liveModel.connect();

Futures.addCallback(sessionFuture, new FutureCallback<LiveSession>() {
    @Override
    public void onSuccess(LiveSession ses) {
	 LiveSessionFutures session = LiveSessionFutures.from(ses);
        session.startAudioConversation();
    }
    @Override
    public void onFailure(Throwable t) {
        // Handle exceptions
    }
}, executor);

Web

برای استفاده از Live API، نمونه‌ای LiveGenerativeModel ایجاد کنید و روش پاسخ را روی AUDIO تنظیم کنید.


import { initializeApp } from "firebase/app";
import { getAI, getLiveGenerativeModel, GoogleAIBackend, ResponseModality } from "firebase/ai";

// TODO(developer): Replace the following with your app's Firebase configuration
// See: https://firebase.google.com/docs/web/learn-more#config-object
const firebaseConfig = {
  // ...
};

// Initialize FirebaseApp
const firebaseApp = initializeApp(firebaseConfig);

// Initialize the Gemini Developer API backend service.
const ai = getAI(firebaseApp, { backend: new GoogleAIBackend() });

// Create a `LiveGenerativeModel` instance with a model that supports the Live API.
const liveModel = getLiveGenerativeModel(ai, {
  model: "GEMINI_LIVE_API_MODEL_NAME",
  // Configure the model to respond with audio.
  generationConfig: {
    responseModalities: [ResponseModality.AUDIO],
  },
});

const session = await liveModel.connect();

// Start the audio conversation.
const audioConversationController = await startAudioConversation(session);

// ... Later, to stop the audio conversation
// await audioConversationController.stop()

Dart

برای استفاده از Live API، نمونه‌ای LiveGenerativeModel ایجاد کنید و روش پاسخ را روی audio تنظیم کنید.


import 'package:firebase_ai/firebase_ai.dart';
import 'package:firebase_core/firebase_core.dart';
import 'firebase_options.dart';
import 'package:your_audio_recorder_package/your_audio_recorder_package.dart';

late LiveModelSession _session;
final _audioRecorder = YourAudioRecorder();

await Firebase.initializeApp(
  options: DefaultFirebaseOptions.currentPlatform,
);

// Initialize the Gemini Developer API backend service.
// Create a `liveGenerativeModel` instance with a model that supports the Live API.
final liveModel = FirebaseAI.googleAI().liveGenerativeModel(
  model: 'GEMINI_LIVE_API_MODEL_NAME',
  // Configure the model to respond with audio.
  liveGenerationConfig: LiveGenerationConfig(
    responseModalities: [ResponseModalities.audio],
  ),
);

_session = await liveModel.connect();

final audioRecordStream = _audioRecorder.startRecordingStream();
// Map the Uint8List stream to InlineDataPart stream.
final mediaChunkStream = audioRecordStream.map((data) {
  return InlineDataPart('audio/pcm', data);
});
await _session.startMediaStream(mediaChunkStream);

// In a separate thread, receive the audio response from the model.
await for (final message in _session.receive()) {
   // Process the received message.
}

Unity

برای استفاده از Live API، نمونه‌ای LiveModel ایجاد کنید و روش پاسخ را روی Audio تنظیم کنید.


using Firebase;
using Firebase.AI;

async Task SendTextReceiveAudio() {
  // Initialize the Gemini Developer API backend service.
  // Create a `LiveModel` instance with a model that supports the Live API.
  var liveModel = FirebaseAI.GetInstance(FirebaseAI.Backend.GoogleAI()).GetLiveModel(
      modelName: "GEMINI_LIVE_API_MODEL_NAME",
      // Configure the model to respond with audio.
      liveGenerationConfig: new LiveGenerationConfig(
          responseModalities: new[] { ResponseModality.Audio })
    );

  LiveSession session = await liveModel.ConnectAsync();

  // Start a coroutine to send audio from the Microphone.
  var recordingCoroutine = StartCoroutine(SendAudio(session));

  // Start receiving the response.
  await ReceiveAudio(session);
}

IEnumerator SendAudio(LiveSession liveSession) {
  string microphoneDeviceName = null;
  int recordingFrequency = 16000;
  int recordingBufferSeconds = 2;

  var recordingClip = Microphone.Start(microphoneDeviceName, true,
                                       recordingBufferSeconds, recordingFrequency);

  int lastSamplePosition = 0;
  while (true) {
    if (!Microphone.IsRecording(microphoneDeviceName)) {
      yield break;
    }

    int currentSamplePosition = Microphone.GetPosition(microphoneDeviceName);

    if (currentSamplePosition != lastSamplePosition) {
      // The Microphone uses a circular buffer, so we need to check if the
      // current position wrapped around to the beginning, and handle it accordingly.
      int sampleCount;
      if (currentSamplePosition > lastSamplePosition) {
        sampleCount = currentSamplePosition - lastSamplePosition;
      } else {
        sampleCount = recordingClip.samples - lastSamplePosition + currentSamplePosition;
      }

      if (sampleCount > 0) {
        // Get the audio chunk.
        float[] samples = new float[sampleCount];
        recordingClip.GetData(samples, lastSamplePosition);

        // Send the data, discarding the resulting Task to avoid the warning.
        _ = liveSession.SendAudioAsync(samples);

        lastSamplePosition = currentSamplePosition;
      }
    }

    // Wait for a short delay before reading the next sample from the Microphone.
    const float MicrophoneReadDelay = 0.5f;
    yield return new WaitForSeconds(MicrophoneReadDelay);
  }
}

Queue<float> audioBuffer = new();

async Task ReceiveAudio(LiveSession liveSession) {
  int sampleRate = 24000;
  int channelCount = 1;

  // Create a looping AudioClip to fill with the received audio data.
  int bufferSamples = (int)(sampleRate * channelCount);
  AudioClip clip = AudioClip.Create("StreamingPCM", bufferSamples, channelCount,
                                    sampleRate, true, OnAudioRead);

  // Attach the clip to an AudioSource and start playing it.
  AudioSource audioSource = GetComponent<AudioSource>();
  audioSource.clip = clip;
  audioSource.loop = true;
  audioSource.Play();

  // Start receiving the response.
  await foreach (var message in liveSession.ReceiveAsync()) {
    // Process the received message
    foreach (float[] pcmData in message.AudioAsFloat) {
      lock (audioBuffer) {
        foreach (float sample in pcmData) {
          audioBuffer.Enqueue(sample);
        }
      }
    }
  }
}

// This method is called by the AudioClip to load audio data.
private void OnAudioRead(float[] data) {
  int samplesToProvide = data.Length;
  int samplesProvided = 0;

  lock(audioBuffer) {
    while (samplesProvided < samplesToProvide && audioBuffer.Count > 0) {
      data[samplesProvided] = audioBuffer.Dequeue();
      samplesProvided++;
    }
  }

  while (samplesProvided < samplesToProvide) {
    data[samplesProvided] = 0.0f;
    samplesProvided++;
  }
}



قیمت‌گذاری و شمارش داده‌واحد

می‌توانید اطلاعات قیمت‌گذاری مدل‌های Live API را در مستندات ارائه‌دهنده Gemini API انتخابی‌تان پیدا کنید: Gemini Developer API | Agent Platform Gemini API (formerly Vertex AI).

صرف‌نظر از ارائه‌دهنده Gemini API شما، Live API از «میانای برنامه‌سازی کاربردی شمارش نشان‌ها» پشتیبانی نمی‌کند.



چه کارهای دیگری می‌توانید انجام دهید؟

  • مجموعه کامل قابلیت‌ها برای Live API را بررسی کنید، مانند جاری‌سازی روش‌های مختلف ورودی (صوت، نوشتار، یا ویدیو + صوت).

  • پیاده‌سازی‌تان را بااستفاده از گزینه‌های پیکربندی مختلف مثل افزودن ترانویسی یا تنظیم صدای پاسخ سفارشی‌سازی کنید.

  • درباره مدیریت جلسه‌ها، ازجمله به‌روزرسانی محتوا در میان جلسه، فشرده کردن پنجره زمینه‌ای، تشخیص اینکه جلسه درحال اتمام است، و ازسر گرفتن جلسه بیشتر بدانید.

  • با دادن دسترسی مدل به ابزارهایی مثل فراخوانی تابع و «مستندسازی» با Google Search، پیاده‌سازی‌تان را پرتوان کنید. مستندات رسمی برای استفاده از ابزارها با Live API به‌زودی ارائه می‌شود!

  • درباره حدود و مشخصات استفاده از Live API مثل طول جلسه، حدود نرخ، زبان‌های پشتیبانی‌شده، و غیره بیشتر بدانید.